- Heterogeneous compute: VMs, serverless, and containers require different allocation strategies.
- Millisecond-scale workloads: event-driven microservices need fast, predictable scaling.
- Cost control: scale precisely to avoid wasted spend.
- Sustainability & efficiency: reduce unnecessary resource consumption and carbon footprint.

- Normal traffic: ~100 spawns/sec.
- Event peak: ~2,000 spawns/sec (20× spike). Over-provisioning for peak wastes cost; under-provisioning harms UX. Platform engineers must balance peak capacity, cost, and SLA.

- Support peak load without degrading UX.
- Maintain cost/SLA balance.
- Design predictable scaling behavior for applications.
- Requests: scheduler uses these to place pods — a guaranteed floor.
- Limits: kubelet enforces these at runtime — a ceiling.
- QoS classes: Guaranteed, Burstable, BestEffort — influence eviction order under node pressure.

- OS/system daemons
- kubelet / container runtime / network plugins
- Application workloads

- Use realistic values guided by load testing/profiling.
- Prefer horizontal scaling (replicas) and leave some headroom for system processes (~20% is a common starting point).
Requests affect scheduling (the scheduler uses them). Limits affect runtime enforcement (the kubelet enforces them). This distinction is frequently tested on the CNPA exam.
Horizontal Pod Autoscaler (HPA)
- Uses metrics from the Metrics Server or external adapters.
- Can scale on CPU, memory, or custom/pod/external metrics.
- Samples at regular intervals and adjusts replicas between min/max.


- Default evaluation frequency is frequent (often ~15s), configurable.
- When multiple metrics exist, HPA uses the metric that recommends the highest replica count.
- Scale-up tends to be faster; scale-down is conservative to avoid flapping.

- VPA can recommend or apply new CPU/memory requests; applying changes may restart pods.
- Dynamic Resource Allocation (DRA) is an emerging approach that enables more flexible runtime changes (including specialized hardware like GPUs) in supported platforms without disruptive restarts.
- Observes unschedulable pods and adjusts node pool size.
- Scales down idle nodes while honoring PodDisruptionBudgets (PDBs).
- Requires cloud permissions (e.g., IAM) to create/delete instances.


- Pod requests may not match any available instance type — autoscaler cannot add an unsupported node type.
- Autoscaler respects PDBs and will not scale in if it would violate them.
- Missing or incorrect cloud permissions will block node operations.

- Scales on external signals: queue depth (Kafka, RabbitMQ), Prometheus, cloud metrics, or scheduled (cron) triggers.
- Can scale to zero for idle workloads (serverless-style behavior).
- Supports multi-trigger logic and uses the trigger that yields the highest replica recommendation.



- Karpenter chooses instance types and provisions nodes just-in-time, mixing Spot and On-Demand where appropriate.
- Reduces reliance on pre-defined Auto Scaling Groups and improves startup latency for nodes.
- Cloud providers provide managed variants (GKE Autopilot, AKS virtual nodes).
- Declarative claims for specialized hardware (pattern similar to PersistentVolumeClaims).
- Runtime flexibility in supported environments (less disruptive than restarts).


- The platform defines the resource class and parameters; workloads create claims and reference them from pods.
- Database autoscalers for stateful workloads.
- Batch schedulers and job managers (e.g., Volcano).
- Multi-cluster autoscalers (e.g., Admiralty).
- Predictive autoscalers using historical patterns (custom operators, Kueue-like controllers).

- Missing/incorrect metrics (e.g., HPA not scaling because Metrics Server is down).
- Pods stuck in Pending (requests exceed available node types, quotas, or constraints).
- Cluster Autoscaler issues (permissions/roles or wrong node pools).
- DRA/VPA recommendations not applied due to policies or constraints.
- Cloud credentials and IAM permissions for autoscalers that manage nodes.
- Health of metric pipeline (Metrics Server, Prometheus, external metrics adapter).
- PodDisruptionBudgets, resource quotas, and namespace/cluster policies.
- Pod scheduling constraints and node selectors/taints/tolerations.


- Monitor the scaling control plane (HPA, Cluster Autoscaler, Karpenter, metrics exporters).
- Test scaling behavior in staging/non-production using load tests.
- Document runbooks and common troubleshooting steps for operators and developers.
- Implement health checks for scaling components (startup, readiness, liveness probes).
- Use resource quotas and limit ranges to prevent noisy neighbors and runaway requests.

- Resource fundamentals: requests determine scheduling; limits determine runtime enforcement.
- Multi-layer scaling: HPA for stateless pods, Cluster Autoscaler/Karpenter for node capacity, KEDA for event-driven/business metrics, and DRA/VPA for dynamic resource sizing where supported.
- Troubleshooting usually involves metrics, permissions, scheduling constraints, or policy conflicts — validate these first.
- Test scaling behavior before production use and automate monitoring + runbooks to reduce incident MTTR.
Be sure autoscalers have correct cloud provider permissions and your metric pipeline is healthy. Missing IAM roles or a downed metrics server are common root causes of autoscaling failures.
- Kubernetes Concepts: Requests and Limits — https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/
- Horizontal Pod Autoscaler (HPA) — https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/
- Cluster Autoscaler — https://github.com/kubernetes/autoscaler/tree/master/cluster-autoscaler
- KEDA — https://keda.sh/
- Karpenter — https://karpenter.sh/
- Vertical Pod Autoscaler (VPA) — https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler