Skip to main content
Welcome. This lesson is the opening module for Domain 2: Observability Fundamentals. We’ll define observability in the context of platform engineering, explain the primary signals (metrics, logs, traces), and cover supporting concepts such as events, alerts, SLIs/SLOs/SLAs, and unified telemetry pipelines. Observability in platform engineering means building visibility into the platform by default. In large, distributed systems, observability enables real-time debugging, proactive operations, and data-driven decisions. Traditional monitoring tells you what broke; observability helps you understand why — and what to do next. On the certification exam expect scenario-based questions that ask you to choose the correct signal(s) (metrics, logs, traces, events) to determine the cause and impact of an issue. Foundation: The classical three pillars are metrics, logs, and distributed traces. Events and alerts add business and operational context that overlays these pillars.
A presentation slide titled "Foundation – Metrics, Logs, and Distributed Traces" showing a three-layer pyramid that defines Metrics (numerical measurements), Logs (timestamped events), and Traces (request journeys across services) with matching icons. A lower band also highlights 2025 extensions like Events and Alerts.
Key point: each pillar provides a different perspective on system health and performance. Use the right signal for the right purpose and combine signals to speed up root-cause analysis.

Metrics

Metrics are numeric signals used for trend analysis, alerting, and dashboards. They are compact, high-cardinality-friendly summaries that drive alerting and SLO measurement. Types of metrics:
  • Counters — monotonically increasing values (e.g., total requests).
  • Gauges — instantaneous values that can go up and down (e.g., memory usage).
  • Histograms — bucketed distributions (excellent for latency distribution and aggregation).
  • Summaries — local quantiles (percentile calculation per instance; not trivially aggregatable across instances).
Metric type overview: Example metric names and common naming patterns:
Prometheus is widely used in Kubernetes environments:
  • Prometheus scrapes metrics via a pull model from instrumented endpoints.
  • PromQL is the Prometheus query language for building alerts and dashboards.
  • Integration options include Pushgateway (for short-lived jobs — use cautiously), and exporting via the OpenTelemetry Collector.
  • Adopt label-based, multi-dimensional metrics (include environment, team, zone, etc.) to provide context for querying and aggregation.
Example labeled metrics:
A presentation slide titled "Metrics — Numbers That Tell the Story" showing four colorful cards for Counters, Gauges, Histograms, and Summaries with icons. Each card includes a brief description of the metric type (e.g., counters for always-increasing values, gauges for current state, histograms for distributions, summaries for quantiles).

Logs

Logs record time-stamped events and rich context useful for debugging and forensic analysis. Structured logs (JSON or key/value) are preferred because they are queryable and easy to correlate with traces and metrics. Common fields to include in structured logs:
  • timestamp, level (INFO/WARN/ERROR), service, trace_id, span_id, message
  • error details, user or tenant IDs, and application-specific metadata
Example structured log:
Log pipeline components: Structured logging that includes trace_id and span_id lets you jump from a log event into a distributed trace for faster root-cause isolation.
A presentation slide titled "Logs – Detailed Event Context for Debugging" showing a circular workflow for collecting, storing, and analyzing logs. It lists tools like Fluent Bit, OpenTelemetry Collector, Elasticsearch, Loki, Kibana and Grafana.

Distributed Tracing

Distributed tracing captures end-to-end request journeys across services. A trace is composed of spans; each span records an operation, timing, and span context (trace_id, span_id, parent relationship). Tracing helps identify latency hotspots, dependencies, and serialization of work. Typical tracing architecture:
  • Instrumentation libraries (language SDKs)
  • Local agents/collectors (Jaeger agent, OpenTelemetry Collector)
  • Back-end trace storage and query/UI (Jaeger, Tempo, commercial tracing services)
Example trace summary:
Use tracing when you need to examine end-to-end user journeys (frontend → API gateway → microservices → database) and pinpoint where latency accumulates or failures propagate.
A slide titled "Distributed Tracing – Following Requests Across Services" listing brief definitions for Trace, Span, Span Context, and Timing. On the right is a Jaeger architecture diagram showing application instrumentation, jaeger client/agent/collector, a Cassandra data store, and query/UI components with OpenTracing and language icons.

Events

Events are structured, business-context signals that can annotate telemetry and tie technical signals to business activities (sales, feature toggles, capacity changes). They help correlate user behavior and business outcomes with system telemetry. Example business event (YAML):
Events are useful for annotating dashboards, enriching traces/logs, and powering analytics and automated workflows.

Alerts

Alerts notify operators when telemetry indicates conditions that require attention. Effective alerts are actionable and include context such as runbooks, links to dashboards, and a representative trace or log sample.
Noisy alerts reduce operational effectiveness. Tune thresholds, use multi-signal triggers, and include contextual information (runbooks, traces, dashboards) to keep alerts actionable.
Example alert payload:
Design alerts to reduce noise:
  • Prefer multi-signal alerts (e.g., metric threshold + error-rate increase).
  • Use dynamic thresholds or anomaly detection where appropriate.
  • Include remediation steps and responsible on-call team.

SLIs, SLOs, and SLAs

  • SLI (Service Level Indicator): a measured metric (e.g., successful_requests / total_requests).
  • SLO (Service Level Objective): the target for an SLI (e.g., 99.5% availability over 30 days).
  • SLA (Service Level Agreement): contractual commitments that may map to SLOs and include penalties for breaches.
Example SLO and error budget:
Use SLOs and error budgets to balance reliability and feature velocity. Certification questions often test conceptual understanding of which SLIs/SLOs best match different service types.

OpenTelemetry and unified pipelines

OpenTelemetry (OTel) is the vendor-neutral standard for collecting metrics, logs, and traces. It supports auto-instrumentation and provides the OpenTelemetry Collector — a single pipeline for ingesting, processing, enriching, and exporting telemetry to multiple back ends (Prometheus, Jaeger/Tempo, Loki, commercial SaaS).
OpenTelemetry provides a vendor-neutral way to collect and route metrics, logs, and traces. The OTel Collector centralizes ingestion, processing, and export to Prometheus, Jaeger/Tempo, Loki, or commercial backends.
A presentation slide titled "OpenTelemetry – One Standard for All Observability Signals" showing three colored avatar icons labeled Swati, Alan, and Phuong with short role descriptions beneath each. A small "Sparkle Pony Ranch" tag appears above the avatars and a copyright notice is at the bottom.

Platform engineering best practices

  • Collect telemetry across application, infrastructure, tenant, and platform services.
  • Centralize processing and enrichment (OTel Collector) to reduce agent complexity.
  • Store and visualize data where it is most useful, balancing retention and cost.
  • Bake instrumentation into the platform: observability by default enables rapid debugging and automation.
A presentation slide titled "Built-in Observability – Platform Engineering Best Practice" showing a colorful circular arrow diagram that links four metric categories: Application Metrics, Infrastructure Metrics, Tenant Metrics, and Platform Services, each with brief example items.
Data flow: collection → processing → storage → analysis → visualization → action. This pipeline turns raw telemetry into actionable intelligence.
An infographic titled "Data Flow – From Application to Insight" showing a vertical pipeline with stages like Collection, Processing, Storage, Analysis, Visualization, and Action, each paired with simple icons. It illustrates the flow from raw data collection through processing and storage to analysis, visualization, and final action.
Service-level thinking turns data into reliability goals. SLIs measure; SLOs set targets; SLAs formalize commitments. Observability is essential to know whether you are meeting those commitments.
A presentation slide titled "Service-Level Objectives – Turning Data Into Reliability Goals" that lists SLI, SLO, and SLA with brief definitions. On the right is a mock observability dashboard showing overall health, deployment frequency, and performance charts.

Summary — Key takeaways

  • The core observability pillars are metrics, logs, and traces. Events and alerts add business and operational context.
  • Use metrics for trends, alerts, and SLOs; use logs for detailed event context; use traces for end-to-end latency and dependency analysis.
  • OpenTelemetry provides a vendor-neutral, unified way to collect and route telemetry.
  • Platform engineering should be SLO-driven: define SLIs, set SLOs, and manage error budgets to balance reliability and delivery speed.
  • Instrumentation and observability should be built into the platform to enable proactive operations, fast incident response, and efficient troubleshooting.
  • Be mindful of storage, cost, and signal quality. Reduce noisy alerts and favor contextual, multi-signal detection.
Further lessons in Domain 2 will expand on advanced topics: long-term storage trade-offs, cost-efficient retention strategies, advanced SLO frameworks, and incident response playbooks. Thanks for reading.

Watch Video