Short answer
Observability is the ability to understand a system’s internal state from its outputs. Metrics are numeric time series for dashboards and alerts, such as latency and error rate. Logs are detailed event records for investigating specific cases. Traces follow one request across services to show where time is spent. Together they answer what, where and why.
Practical guidance
- Alert on symptoms users feel — error rate, latency — not every internal signal.
- Use structured logs with request or trace ids.
- Define SLIs and SLOs, and use error budgets to balance speed and reliability.
How to answer it in an interview
- Mention the four golden signals: latency, traffic, errors, saturation.
- Describe an incident you debugged with traces.