This guide covers: the reconciliation loop, controller/informer architecture, node health and self-healing, custom controllers/operators, and design best practices for building resilient controllers. Use this as a conceptual reference for exam prep and real-world platform design.
- Certification exams commonly test imperative vs. declarative approaches and scenarios where the reconciliation model is the correct answer.
- For platform teams (for example, “Sparkle Pony Ranch”), the Kubernetes API provides a single, consistent interface used by
kubectl, UIs, controllers, and custom automation. - Everything — from a
kubectllisting to an operator action — goes through the same API surface.
spec that declares desired state. Controllers (the control plane) continuously move the cluster toward that desired state and report observed state in a status section.
Example Deployment — desired state in spec, observed state in status:

- Observe current state (via API server and informers).
- Compare current vs desired state.
- Act to correct drift (create, update, or delete resources).
- Repeat until state converges.

- Controllers implement reconciliation logic: they define the business rules for a resource type.
- Informers watch the API server, maintain a local cache, and emit events on changes.
- Work queues buffer and de-duplicate keys so controllers can handle spikes and retries without being overwhelmed.
- Slow reaction is usually an informer/queue or processing issue, not a polling setting.
- Platform teams can extend Kubernetes by adding custom controllers to automate provisioning, SLO-based scaling, and other platform concerns.

Controller chains of responsibility
Controllers often work in chains. For example, a Deployment controller creates ReplicaSets; a ReplicaSet controller creates Pods. Each controller focuses on a resource type and relies on the API server and informers to coordinate.
Node controller and node health
The node controller manages node lifecycle and health signals:
- Tracks node heartbeats and readiness (via Node status and Lease objects from kubelet).
- Updates node status and conditions (e.g., Ready, MemoryPressure, DiskPressure).
- Tracks capacity and resource pressure that influence scheduling.

- Pod restart: kubelet restarts crashed containers.
- Pod replacement: ReplicaSets recreate pods to maintain desired replicas.
- Service recovery: Services/Endpoints update when pod health changes.
- Node drain/migration: workloads are relocated when nodes are removed or upgraded.

- Validates and authorizes all requests.
- Runs admission control plugins (mutating and validating webhooks).
- Persists cluster state to etcd.
Do not modify etcd directly. Controllers and automation should submit changes through the API server to ensure validation, admission control, and persistence workflows are honored.

- Watch the API server for changes.
- Maintain a local cache to minimize API traffic.
- Emit events and enqueue keys for controllers.
- Use work queues to handle spikes, retries, and deduplication.

- Transient failures: retries with backoff and idempotent operations.
- Rate limiting: respect API server throttle and implement retry logic.
- Resource conflicts: use optimistic concurrency (e.g.,
resourceVersion) and retry on conflict. - External dependency failures: graceful degradation and accurate status reporting.

- Make reconciliation idempotent (safe to re-run).
- Reconcile against the current observed state, not just event payloads.
- Report
statuson custom resources to expose progress and health. - Use informers and a local cache to reduce API load.
- Implement retries, exponential backoff, and graceful error handling.

- The Kubernetes API server is the consistent REST interface used by
kubectl, dashboards, controllers, and automation. - The reconciliation loop (observe → compare → act → repeat) is the fundamental pattern enabling declarative, self-healing systems.
- Controllers implement reconciliation logic; informers provide efficient event-driven caching and watches.
- Custom controllers and operators let you extend Kubernetes with domain-specific automation and operational expertise.
- Always interact with the API server and design controllers to be idempotent, observable, and resilient.

