Monitoring vs observability: logs, metrics, traces and SLOs

What monitoring and observability each answer, how logs, metrics and traces fit together with OpenTelemetry, and how to set SLOs and error budgets around student journeys.

8 min read
On this page 9 sections
  1. Monitoring vs observability
  2. Logs, metrics and traces
  3. OpenTelemetry
  4. SLIs, SLOs and SLAs
  5. Error budgets
  6. A worked example: the test-submission budget
  7. What to measure on a learning platform
  8. Key takeaways
  9. Frequently asked questions

Monitoring tells you that something you anticipated has gone wrong: an alert fires because the error rate or latency crossed a line you drew in advance. Observability is the ability to work out why, including for failures nobody predicted, by asking new questions of your logs, metrics and traces. A learning platform needs both, tied to service level objectives that measure what students actually feel: lectures that start quickly and don't buffer, and test submissions that always go through.

Monitoring vs observability

AspectMonitoringObservability
Question it answersIs something broken right now?Why is it broken, for whom, and since when?
Built fromPredefined dashboards, thresholds and health checksDetailed telemetry you can slice by any attribute after the fact
Good atKnown failure modes: disk full, error spike, service downNew failure modes: slow submits only for students on one app version in one city
Typical outputAlerts and status dashboardsAd hoc queries, traces of individual requests, correlated logs

The two aren't rivals. Google's SRE book frames the job as two questions, "what's broken, and why?", and recommends that monitoring start from four golden signals for each service: latency, traffic, errors and saturation. Monitoring those gives you the "what". Observability is having enough context in your telemetry to answer the "why" without shipping new logging code in the middle of an incident.

Logs, metrics and traces

SignalWhat it isWhat it answersCost driver
MetricsNumbers over time: counters, gauges, histogramsHow many, how fast, how full, and is it getting worse?Number of distinct label combinations
LogsTimestamped records of individual eventsWhat exactly happened in this request or this job?Volume and retention
TracesThe path of one request through every service, as a tree of timed spansWhere did this slow request spend its time?Volume, controlled by sampling

Picture a student pressing "Submit test". A metric tells you submit latency at the 95th percentile rose from 300 ms to 4 seconds at 10:02. A trace of one slow submission shows 3.6 seconds spent waiting for a database connection. Logs from that request carry the same request ID, so you can see the pool warnings around it. Each signal on its own is a clue; linked by shared IDs, they give you the answer.

Two practical rules. Keep high-cardinality values such as student or attempt IDs out of metric labels, where every unique value creates a new time series, and put them in logs and trace attributes instead. And keep personal data such as phone numbers out of logs altogether; a hashed or internal ID is enough to follow a request.

OpenTelemetry

OpenTelemetry is an open-source, vendor-neutral framework for generating, collecting and exporting traces, metrics and logs. It is a Cloud Native Computing Foundation project, formed from the merger of two earlier projects, OpenTracing and OpenCensus. It is deliberately not a backend: it produces and ships telemetry, while tools such as Prometheus and Jaeger, or a commercial service, store and query it. That separation means you instrument once and can change backends later.

For a Django app, zero-code instrumentation is the quickest start. It installs instrumentation for the libraries it finds, such as Django, the Postgres driver, Redis and Celery:

pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install

OTEL_SERVICE_NAME=web-app \
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317 \
opentelemetry-instrument python manage.py runserver --noreload

In production behind Gunicorn, initialise tracing in each worker after it forks; OpenTelemetry's Python documentation explains that its batch exporter isn't fork-safe and shows a post_fork hook. Across services, the W3C Trace Context standard's traceparent header carries the trace ID from your web app into the Celery tasks it queues, so one student's request and its background work appear as one trace.

SLIs, SLOs and SLAs

Google's SRE book defines the three terms precisely:

  • A service level indicator (SLI) is a precisely defined measurement of one aspect of the service, best expressed as good events divided by valid events: submissions saved divided by submissions attempted.

  • A service level objective (SLO) is a target for an SLI over a period: 99.95% of submissions saved over 30 days.

  • A service level agreement (SLA) is a contract that includes consequences for missing the objectives. If nothing happens when a target is missed, it's an SLO, not an SLA.

Choose SLIs from what users care about, not from what is easy to measure, and use percentiles rather than averages, because an average hides the slowest students entirely.

Error budgets

An SLO below 100% leaves an error budget: the amount of failure you can afford in the period. A 99.9% SLO over 30 days allows 0.1% of requests to fail, the equivalent of 43.2 minutes of complete downtime. The budget turns reliability into a shared decision. While there is budget left, teams ship; when it's spent, risky deploys pause and reliability work comes first.

Alert on how fast the budget is burning, not on every blip. Google's SRE workbook recommends, for a 99.9% SLO, paging when 2% of the monthly budget burns within an hour (a burn rate of 14.4) or 5% within six hours (a burn rate of 6), and raising a ticket when 10% burns over three days. Each alert also checks a short window, such as the last five minutes, so it stops firing once the problem is fixed.

A worked example: the test-submission budget

Say your platform handles 4,00,000 test submissions a month, with an SLO of 99.95% saved. The error budget is 200 failed submissions a month. Now a 20,000-student mock test starts at 10:00, and a database problem makes 1% of submissions fail for that hour: 200 failures. One bad morning has spent the whole month's budget. The numbers are illustrative, but the pattern is real for exam platforms: reliability is decided in a few peak hours, which is why error budgets pair naturally with deploy freezes before big tests and with the practices in our guide to zero-downtime deployment.

What to measure on a learning platform

Student journeySLI (good ÷ valid)Illustrative SLO over 30 daysMeasured where
Starting a lecturePlays that show the first frame within 2 seconds ÷ play attempts99%The player
Watching without stallsPlayback sessions with a rebuffering ratio under 1% ÷ sessions98%The player
Submitting a testSubmissions saved ÷ submissions attempted99.95%Server, plus client retry logs
Logging inOTPs delivered within 30 seconds ÷ OTPs requested99%SMS gateway delivery reports
Joining a live classJoins that start playing within 5 seconds ÷ join attempts99%The app

Behind these, keep the service and infrastructure signals that explain them: the golden signals per endpoint, queue length and age of the oldest job, database connections and replication lag, cache hit ratio, and CDN errors by region and ISP. Player measurements, covered in our guide to adaptive bitrate streaming, matter most for video, because server dashboards can look perfect while students in one city buffer. For how these layers fit together, see our guide to scaling a video learning platform, and for recovery targets when things go badly wrong, RTO and RPO.

Key takeaways

  • Monitoring answers "is it broken?"; observability answers "why?", including for failures you didn't predict.

  • Metrics, logs and traces each answer different questions; link them with shared request and trace IDs.

  • OpenTelemetry lets you instrument once and choose or change backends later.

  • Define SLIs from student journeys, set SLOs, and alert on error-budget burn rate rather than raw thresholds.

  • On exam platforms, a few peak hours decide the month's reliability, so protect them.

Frequently asked questions

What is observability in DevOps?

In DevOps, observability means building and running systems so that engineers can understand their internal state from the telemetry they emit: metrics, logs and traces. It supports the DevOps habit of shipping small changes often, because when something goes wrong after a deploy, teams can quickly see which change, service and users are affected, and roll back or fix forward with confidence.

What is observability and monitoring?

They are two layers of the same practice. Monitoring watches known indicators, such as error rate, latency and saturation, and alerts when they cross thresholds. Observability is the wider property of having rich enough, well-linked telemetry to investigate any problem, including new ones. In practice, monitoring tells the on-call engineer that something is wrong, and observability lets them find out why.

What is observability in microservices?

In a microservices system, one user request can pass through many services, queues and databases, so a failure's cause is often far from where it shows up. Observability there relies on distributed tracing, which follows each request across services using propagated trace IDs, alongside per-service metrics and correlated logs. Without it, teams end up guessing which of many services caused a slowdown.

What is the difference between SLI, SLO and SLA?

An SLI is a measurement, such as the percentage of test submissions saved successfully. An SLO is the target for that measurement over a period, such as 99.95% over 30 days. An SLA is a contract with users or customers that attaches consequences, such as service credits, to missing agreed targets. Internal SLOs are usually set stricter than any external SLA.

Share this article

Looking for something else?

Talk to Us