> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Resource Management and Scaling

> Guide to Kubernetes resource management, autoscaling layers, node provisioning, and dynamic resource allocation to balance performance, cost, and reliability for platform engineers.

Welcome — this lesson covers practical, exam-relevant concepts for Domain 4: intelligent resource management, capacity planning, and scaling. These are core skills for platform engineers and frequent topics on the CNPA certification: resource requests and limits, QoS classes, autoscaling layers, node provisioning, and emerging dynamic resource allocation techniques.

Why this matters

* Heterogeneous compute: VMs, serverless, and containers require different allocation strategies.
* Millisecond-scale workloads: event-driven microservices need fast, predictable scaling.
* Cost control: scale precisely to avoid wasted spend.
* Sustainability & efficiency: reduce unnecessary resource consumption and carbon footprint.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/platform-engineering-resource-management-infographic.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=b07eb70edf4f926f488ab7a99f5dc357" alt="An infographic titled &#x22;Platform Engineering Demands Intelligent Resource Management&#x22; that highlights four challenges—Multi-Cloud Complexity, Event-Driven Workloads, Cost Optimization, and Sustainability Goals—each shown with a colored icon and short explanatory text." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/platform-engineering-resource-management-infographic.jpg" />
</Frame>

Real-world example: Sparkle Pony Ranch

* Normal traffic: \~100 spawns/sec.
* Event peak: \~2,000 spawns/sec (20× spike).
  Over-provisioning for peak wastes cost; under-provisioning harms UX. Platform engineers must balance peak capacity, cost, and SLA.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/pony-spawning-scaling-peak-cost-waste.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=ca1a8c14f07df2a4d024192c5d7a32f1" alt="A presentation slide titled &#x22;The Scaling Challenge – Peak Pony Spawning Season&#x22; with three panels showing Normal Load (handles 100 spawns/sec), Peak Load (spikes to 2,000 spawns/sec), and Cost Waste (paying $15,000/month while 75% of capacity sits idle)." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/pony-spawning-scaling-peak-cost-waste.jpg" />
</Frame>

Platform goals (examples: Swati, Alan, Fong)

* Support peak load without degrading UX.
* Maintain cost/SLA balance.
* Design predictable scaling behavior for applications.

Kubernetes primitives: Requests, Limits, and QoS

* Requests: scheduler uses these to place pods — a guaranteed floor.
* Limits: kubelet enforces these at runtime — a ceiling.
* QoS classes: Guaranteed, Burstable, BestEffort — influence eviction order under node pressure.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/kubernetes-resource-requests-limits-qos.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=a529c92be1b8f216c6b5b9422f6fe3c6" alt="A presentation slide titled &#x22;Kubernetes Resource Requests and Limits&#x22; showing three colored panels. The panels summarize Requests (minimum guaranteed resources for pod scheduling), Limits (maximum resources a pod can consume), and QoS Classes (Guaranteed, Burstable, BestEffort)." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/kubernetes-resource-requests-limits-qos.jpg" />
</Frame>

Node-level resource behavior and eviction
When a node is constrained, kubelet applies eviction thresholds and uses QoS to choose which pods to evict first. Node resources include:

* OS/system daemons
* kubelet / container runtime / network plugins
* Application workloads

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/kubernetes-requests-limits-eviction-order-diagram.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=00b0abb5956b5c1881dd33dae5ce5b12" alt="A labeled layered diagram titled &#x22;Kubernetes Resource Requests and Limits&#x22; showing a Kubernetes Node with stacked colored blocks for Eviction Threshold, Applications, kubelet/Docker/kube-proxy/CNI, and OS system daemons. It illustrates resource allocation and eviction order within a Kubernetes node." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/kubernetes-requests-limits-eviction-order-diagram.jpg" />
</Frame>

Example deployment with requests and limits

* Use realistic values guided by load testing/profiling.
* Prefer horizontal scaling (replicas) and leave some headroom for system processes (\~20% is a common starting point).

```yaml theme={null}
apiVersion: apps/v1
kind: Deployment
metadata:
  name: pony-spawner
spec:
  selector:
    matchLabels:
      app: pony-spawner
  template:
    metadata:
      labels:
        app: pony-spawner
    spec:
      containers:
        - name: pony-spawner
          image: registry.spr.com/pony-spawner:v2.1.0
          resources:
            requests:
              cpu: "200m"     # 0.2 CPU guaranteed for scheduling
              memory: "256Mi" # 256Mi memory guaranteed
            limits:
              cpu: "500m"     # Max 0.5 CPU at runtime
              memory: "512Mi" # Max 512Mi memory at runtime
```

<Callout icon="lightbulb" color="#1CB2FE">
  Requests affect scheduling (the scheduler uses them). Limits affect runtime enforcement (the kubelet enforces them). This distinction is frequently tested on the CNPA exam.
</Callout>

Autoscaling layers — how they fit together
Kubernetes and the surrounding ecosystem implement multiple complementary autoscaling layers:

| Layer | Purpose | Examples / Notes |
| - | - | - |
| Pod-level (horizontal) | Scale replicas of stateless services | `HorizontalPodAutoscaler` (HPA) — CPU, memory, custom metrics |
| Pod-level (vertical) | Adjust container requests (resource size) | `VerticalPodAutoscaler` (VPA) / DRA recommendations |
| Node autoscaling | Scale node pools / VMs to match pod resource needs | `Cluster Autoscaler`, managed autoscalers |
| Node provisioners | Fast, workload-aware node provisioning | `Karpenter`, `GKE Autopilot`, `AKS virtual nodes` |
| Event-driven | Business/external/event triggers; can scale to zero | `KEDA` |
| Specialized/autoscalers | DB autoscalers, batch schedulers, multi-cluster | `Volcano`, `Admiralty`, predictive autoscalers |

Horizontal Pod Autoscaler (HPA)

* Uses metrics from the Metrics Server or external adapters.
* Can scale on CPU, memory, or custom/pod/external metrics.
* Samples at regular intervals and adjusts replicas between min/max.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/kubernetes-hpa-metrics-server-pod-scaling.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=512dfac6a3c621949d5e7d26dba2318a" alt="A diagram of Kubernetes Horizontal Pod Autoscaler: the Metrics Server feeds the Horizontal Pod Autoscaler (HPA), which then scales Pods. Dashed boxes and an arrow show scaling up/down additional pod replicas." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/kubernetes-hpa-metrics-server-pod-scaling.jpg" />
</Frame>

HPA scaling types: CPU, memory, and custom metrics

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/hpa-scaling-cpu-memory-custom-metrics.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=9c94ca1c0723d942f86d64926d667cff" alt="A presentation slide titled &#x22;HPA – Scaling Pods Based on Metrics&#x22; showing three scaling types: CPU-based, memory-based, and custom metrics. Each type has a short description about when to scale (CPU thresholds, memory consumption patterns, or business metrics like queue length)." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/hpa-scaling-cpu-memory-custom-metrics.jpg" />
</Frame>

HPA example (CPU + pod-level custom metric)

```yaml theme={null}
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: pony-spawner-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: pony-spawner
  minReplicas: 3
  maxReplicas: 50
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70
    - type: Pods
      pods:
        metric:
          name: pony_requests_per_second
        target:
          type: AverageValue
          averageValue: "30"
```

HPA behavior notes

* Default evaluation frequency is frequent (often \~15s), configurable.
* When multiple metrics exist, HPA uses the metric that recommends the highest replica count.
* Scale-up tends to be faster; scale-down is conservative to avoid flapping.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/sparkle-pony-hpa-15s-scaleup-5min.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=15a6af301eabf76fd44a010406542991" alt="A presentation slide titled &#x22;Sparkle Pony Ranch HPA Configuration&#x22; summarizing key HPA behaviors. It states HPA checks metrics every 15 seconds, uses the metric suggesting the highest number of replicas, and scales up immediately while scale-down has a 5-minute delay." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/sparkle-pony-hpa-15s-scaleup-5min.jpg" />
</Frame>

Vertical Pod Autoscaler (VPA) and Dynamic Resource Allocation (DRA)

* VPA can recommend or apply new CPU/memory requests; applying changes may restart pods.
* Dynamic Resource Allocation (DRA) is an emerging approach that enables more flexible runtime changes (including specialized hardware like GPUs) in supported platforms without disruptive restarts.

Cluster Autoscaler (node-level)

* Observes unschedulable pods and adjusts node pool size.
* Scales down idle nodes while honoring PodDisruptionBudgets (PDBs).
* Requires cloud permissions (e.g., IAM) to create/delete instances.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/0r3GTobZImleUlJh/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/cluster-autoscaler-node-pool-scaling-diagram.jpg?fit=max&auto=format&n=0r3GTobZImleUlJh&q=85&s=4b4d8ca71c0cf690288ea3c79edc573e" alt="A diagram titled &#x22;Cluster Autoscaler – Scale the Infrastructure Itself&#x22; showing how pending pods trigger expansion logic to add available node types. It illustrates a current cluster with node pools like Compute, AMD Compute, High Memory, and GPU being scaled by the autoscaler." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/cluster-autoscaler-node-pool-scaling-diagram.jpg" />
</Frame>

Cluster Autoscaler capabilities

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/0r3GTobZImleUlJh/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/cluster-autoscaler-node-management-multizone-costbalance.jpg?fit=max&auto=format&n=0r3GTobZImleUlJh&q=85&s=1e5d99c0ff98a7d0445475d79d28c1f0" alt="A presentation slide titled &#x22;Cluster Autoscaler – Scale the Infrastructure Itself&#x22; showing three colored cards listing key capabilities: Node Management, Multi-Zone Support, and Cost Balance with brief descriptions of each." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/cluster-autoscaler-node-management-multizone-costbalance.jpg" />
</Frame>

Common pitfalls for Cluster Autoscaler

* Pod requests may not match any available instance type — autoscaler cannot add an unsupported node type.
* Autoscaler respects PDBs and will not scale in if it would violate them.
* Missing or incorrect cloud permissions will block node operations.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/cluster-autoscaler-pitfalls-pod-pdb-iam.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=394d6eca461d2ac0f1225a195d6eac9f" alt="A presentation slide titled &#x22;Cluster Autoscaler – Scale the Infrastructure Itself&#x22; with a highlighted &#x22;Common Pitfalls&#x22; box. It lists three issues: pod requests may exceed available node types, the autoscaler respects Pod Disruption Budgets during scale‑in, and it requires proper IAM permissions to manage nodes." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/cluster-autoscaler-pitfalls-pod-pdb-iam.jpg" />
</Frame>

KEDA — event-driven autoscaling

* Scales on external signals: queue depth (Kafka, RabbitMQ), Prometheus, cloud metrics, or scheduled (cron) triggers.
* Can scale to zero for idle workloads (serverless-style behavior).
* Supports multi-trigger logic and uses the trigger that yields the highest replica recommendation.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/keda-beyond-cpu-memory-scaling.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=9283e3ea6dc1752aef30d2110094875f" alt="A slide titled &#x22;KEDA — Beyond CPU and Memory Scaling&#x22; showing three colored boxes: Message Queues, External Metrics, and Cron Scaling. Each box lists examples (Kafka/RabbitMQ/Azure Service Bus; Prometheus/CloudWatch/DataDog; and predictive time-based scaling)." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/keda-beyond-cpu-memory-scaling.jpg" />
</Frame>

KEDA use cases: business events, queue-driven processing, or time-based burst scaling.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/keda-multi-trigger-scaling-pony-services.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=b627f3befcc640ceb8d5c90ac120c140" alt="A presentation slide titled &#x22;KEDA – Scaling Pony Services on External Events&#x22; showing &#x22;KEDA's Multi-Trigger Logic.&#x22; It lists three points: use the highest replica count, triggers scale at different speeds, and allow combining triggers for advanced scaling." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/keda-multi-trigger-scaling-pony-services.jpg" />
</Frame>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/keda-scaling-pony-services-external-events.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=99cee94d62f5e2f78126b2223b97b380" alt="A presentation slide titled &#x22;KEDA – Scaling Pony Services on External Events.&#x22; It lists platform value points: enabling scaling from technical metrics (e.g., Kafka lag), supporting business metrics (e.g., pony queue depth), and abstracting complexity for developers with intelligent scaling." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/keda-scaling-pony-services-external-events.jpg" />
</Frame>

Karpenter and workload-aware node provisioning

* Karpenter chooses instance types and provisions nodes just-in-time, mixing Spot and On-Demand where appropriate.
* Reduces reliance on pre-defined Auto Scaling Groups and improves startup latency for nodes.
* Cloud providers provide managed variants (GKE Autopilot, AKS virtual nodes).

Dynamic Resource Allocation (DRA) — beyond CPU/memory
DRA extends resource management to GPUs, custom hardware, and finer-grained resource shares (in implementations that support fractional GPU sharing). It can allow:

* Declarative claims for specialized hardware (pattern similar to PersistentVolumeClaims).
* Runtime flexibility in supported environments (less disruptive than restarts).

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/dynamic-resource-allocation-beyond-cpu-memory.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=a27bdd581b9245442163c9980605dc2b" alt="A presentation slide titled &#x22;Dynamic Resource Allocation – Beyond CPU and Memory&#x22; highlighting three features. It lists GPU Management (dynamic/fractional GPU sharing), Custom Resources (network bandwidth, storage IOPS, specialized hardware), and Runtime Flexibility (modify allocations without pod restarts)." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/dynamic-resource-allocation-beyond-cpu-memory.jpg" />
</Frame>

Platform value for DRA

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/dynamic-resource-allocation-platform-value.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=c51840b84c55b8ca1fd8425ea6734c10" alt="A presentation slide titled &#x22;Dynamic Resource Allocation – Beyond CPU and Memory&#x22; showing a &#x22;Platform Value&#x22; box with two bullets: &#x22;Enables self-service access to specialized hardware&#x22; and &#x22;Replaces manual GPU allocation with declarative requests.&#x22; The slide also shows a faint gear/server illustration and a © Copyright KodeKloud note." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/dynamic-resource-allocation-platform-value.jpg" />
</Frame>

DRA example: resource claim + pod reference

* The platform defines the resource class and parameters; workloads create claims and reference them from pods.

```yaml theme={null}
apiVersion: resource.k8s.io/v1alpha2
kind: ResourceClaim
metadata:
  name: pony-ai-gpu
spec:
  resourceClassName: gpu-class
  parametersRef:
    apiGroup: gpu.example.com
    kind: GpuClaimParameters
    name: pony-processing-params
---
apiVersion: v1
kind: Pod
metadata:
  name: pony-ai-processor-pod
spec:
  containers:
    - name: pony-ai-processor
      image: registry.spr.com/pony-ai:v1.0.0
      resources:
        claims:
          - name: pony-ai-gpu
```

DRA benefits (where supported): fractional GPU sharing, automatic cleanup, and self-service access to expensive resources for complex workloads.

Ecosystem autoscalers for specialized workloads

* Database autoscalers for stateful workloads.
* Batch schedulers and job managers (e.g., Volcano).
* Multi-cluster autoscalers (e.g., Admiralty).
* Predictive autoscalers using historical patterns (custom operators, Kueue-like controllers).

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/ecosystem-autoscalers-specialized-workloads-slide.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=0b25e8f0ef5de87aacb836c12bda0676" alt="A presentation slide titled &#x22;Ecosystem Autoscalers for Specialized Workloads&#x22; showing four labeled boxes describing Database Autoscalers, Batch Job Scaling, Multi-Cluster Autoscaling, and Predictive Autoscaling with brief notes on each. The slide also includes a small copyright notice for KodeKloud." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/ecosystem-autoscalers-specialized-workloads-slide.jpg" />
</Frame>

Troubleshooting common scaling problems

* Missing/incorrect metrics (e.g., HPA not scaling because Metrics Server is down).
* Pods stuck in Pending (requests exceed available node types, quotas, or constraints).
* Cluster Autoscaler issues (permissions/roles or wrong node pools).
* DRA/VPA recommendations not applied due to policies or constraints.

When troubleshooting, verify:

* Cloud credentials and IAM permissions for autoscalers that manage nodes.
* Health of metric pipeline (Metrics Server, Prometheus, external metrics adapter).
* PodDisruptionBudgets, resource quotas, and namespace/cluster policies.
* Pod scheduling constraints and node selectors/taints/tolerations.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/scaling-troubleshooting-checklist.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=bf0f7e212620c8ab11803a2944e67e8c" alt="A presentation slide titled &#x22;Common Scaling Problems and Solutions&#x22; with a &#x22;Troubleshooting Checklist&#x22; header. Four colored boxes list issues—Resource Constraints, Metrics Availability, Policy Conflicts, and DRA Issues—each paired with a short recommended check or fix." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/scaling-troubleshooting-checklist.jpg" />
</Frame>

Common symptom examples

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/common-scaling-problems-hpa-pending-vpa.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=916f75ac246428210ed188de073d36b5" alt="A presentation slide titled &#x22;Common Scaling Problems and Solutions&#x22; showing three numbered rounded boxes listing: 01 HPA not scaling despite high CPU usage, 02 Pods stuck in Pending state, and 03 VPA recommendations not being applied. The slide has a simple white layout with a small &#x22;© Copyright KodeKloud&#x22; footer." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/common-scaling-problems-hpa-pending-vpa.jpg" />
</Frame>

Best practices for platform engineers

* Monitor the scaling control plane (HPA, Cluster Autoscaler, Karpenter, metrics exporters).
* Test scaling behavior in staging/non-production using load tests.
* Document runbooks and common troubleshooting steps for operators and developers.
* Implement health checks for scaling components (startup, readiness, liveness probes).
* Use resource quotas and limit ranges to prevent noisy neighbors and runaway requests.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/1hdMw9SCEvW2M34U/images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/platform-engineering-scaling-best-practices.jpg?fit=max&auto=format&n=1hdMw9SCEvW2M34U&q=85&s=5c4859f679d7a78af6a2733283c95bfc" alt="A slide titled &#x22;Common Scaling Problems and Solutions&#x22; / &#x22;Platform Engineering Best Practices&#x22; showing four numbered, colored boxes. They list best practices: set up monitoring for scaling infrastructure, test scaling in non-production, document troubleshooting procedures, and implement automated health checks for scaling components." width="1920" height="1080" data-path="images/Prep-Course-Certified-Cloud-Native-Platform-Engineering-Associate-CNPA/Domain-4-Platform-APIs-and-Provisioning-Infrastructure/Resource-Management-and-Scaling/platform-engineering-scaling-best-practices.jpg" />
</Frame>

Key takeaways

* Resource fundamentals: requests determine scheduling; limits determine runtime enforcement.
* Multi-layer scaling: HPA for stateless pods, Cluster Autoscaler/Karpenter for node capacity, KEDA for event-driven/business metrics, and DRA/VPA for dynamic resource sizing where supported.
* Troubleshooting usually involves metrics, permissions, scheduling constraints, or policy conflicts — validate these first.
* Test scaling behavior before production use and automate monitoring + runbooks to reduce incident MTTR.

<Callout icon="warning" color="#FF6B6B">
  Be sure autoscalers have correct cloud provider permissions and your metric pipeline is healthy. Missing IAM roles or a downed metrics server are common root causes of autoscaling failures.
</Callout>

Further reading and references

* Kubernetes Concepts: Requests and Limits — [https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/)
* Horizontal Pod Autoscaler (HPA) — [https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/](https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/)
* Cluster Autoscaler — [https://github.com/kubernetes/autoscaler/tree/master/cluster-autoscaler](https://github.com/kubernetes/autoscaler/tree/master/cluster-autoscaler)
* KEDA — [https://keda.sh/](https://keda.sh/)
* Karpenter — [https://karpenter.sh/](https://karpenter.sh/)
* Vertical Pod Autoscaler (VPA) — [https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler](https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler)

Intelligent resource management lets teams scale reliably, control costs, and deliver consistent user experiences. Thanks for reading this lesson.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/certified-cloud-native-platform-engineering-associate-cnpa/module/7b8d1069-510d-48c6-b656-7573a193aeff/lesson/d943f08b-3449-46da-8575-b8ebd21f8c61" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.