
Photo: FrancezcoPezce, CC0
Kubernetes Cost Optimization: Save 40-70% on Your Cloud Bill
Right-sizing, spot instances, autoscaling, and resource quotas — real strategies that saved companies millions on Kubernetes costs in 2026.
The average Kubernetes cluster wastes 49% of allocated resources (CNCF 2025 survey). Here are the specific strategies that actually reduce costs, with real numbers from production environments.
Key Takeaways
- 49% of Kubernetes CPU and 42% of memory is wasted (CNCF 2025)
- Right-sizing alone saves 30-40% on average compute costs
- Spot/preemptible instances save 60-80% for fault-tolerant workloads
- Companies implementing FinOps save $2.4M annually per $10M spend
- Vertical Pod Autoscaler reduces over-provisioning by 35%
The Real Cost Problem
According to the CNCF 2025 Kubernetes Survey (4,000+ respondents):
- 49% of allocated CPU is never used
- 42% of allocated memory is never used
- Average waste: $150K/month for a mid-size cluster (500 pods)
- Time to optimize: Most teams spend 2-4 weeks on initial optimization
Why This Happens
| Root Cause | % of Teams | Impact |
|---|---|---|
| No resource requests/limits set | 34% | Over-provisioning |
| Static over-provisioning | 28% | "Just in case" padding |
| No autoscaling configured | 22% | Fixed capacity |
| Development clusters left running | 16% | Night/weekend waste |
Strategy 1: Right-Sizing (30-40% Savings)
How to Identify Over-Provisioned Pods
# Install Goldilocks (free, open-source)
kubectl apply -f https://github.com/FairwindsOps/goldilocks/releases/latest/download/goldilocks.yaml
# View recommendations in dashboard
kubectl port-forward -n goldilocks svc/goldilocks 8080:80
Goldilocks shows actual vs. requested resources for every pod. Most teams find:
- CPU requests are 2-5x higher than actual usage
- Memory requests are 1.5-3x higher than actual usage
Real Results
Shopify (case study, KubeCon 2025):
- Analyzed 15,000 pods across 200 microservices
- Found average CPU utilization of 12% (requests were 8x actual usage)
- Right-sized over 3 months
- Result: $1.8M annual savings (35% reduction in compute costs)
Spotify (published results):
- Implemented vertical pod autoscaler across all services
- Reduced average pod CPU request from 500m to 200m
- Result: 40% cost reduction, zero performance degradation
Implementation
# Example: Before right-sizing
resources:
requests:
cpu: "2000m" # 2 CPU cores requested
memory: "4Gi" # 4GB memory requested
limits:
cpu: "4000m"
memory: "8Gi"
# After right-sizing (based on actual usage data)
resources:
requests:
cpu: "500m" # Actual usage: 200-400m
memory: "1Gi" # Actual usage: 500MB-1.2GB
limits:
cpu: "1000m"
memory: "2Gi"
Strategy 2: Spot/Preemptible Instances (60-80% Savings)
Cloud Provider Pricing (2026)
| Provider | On-Demand (c6i.xlarge) | Spot Price | Savings |
|---|---|---|---|
| AWS | $0.17/hr | $0.05/hr | 71% |
| Google Cloud | $0.16/hr | $0.048/hr | 70% |
| Azure | $0.168/hr | $0.05/hr | 70% |
When to Use Spot Instances
Good candidates (fault-tolerant):
- Batch processing jobs
- CI/CD pipelines
- Development/staging environments
- Stateless web services with autoscaling
- ML training jobs
Bad candidates (critical path):
- Databases
- Single-instance services
- Services with stateful sessions
- Real-time payment processing
Spot Instance Best Practices
# Use multiple instance types for resilience
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: node.kubernetes.io/capacity-type
operator: In
values: ["spot"]
tolerations:
- key: "kubernetes.azure.com/scalesetpriority"
operator: "Equal"
value: "spot"
effect: "NoSchedule"
Netflix reports running 70% of their compute on spot instances for non-critical workloads, saving approximately $200M annually.
Strategy 3: Cluster Autoscaler (20-30% Savings)
How It Works
The Cluster Autoscaler adjusts the number of nodes based on pod scheduling requirements:
- Scale up: When pods can't be scheduled due to insufficient resources
- Scale down: When nodes are underutilized for 10+ minutes
Configuration
apiVersion: autoscaling.k8s.io/v1
kind: ClusterAutoscaler
metadata:
name: cluster-autoscaler
spec:
scaleDownDelayAfterAdd: 10m
scaleDownUnneededTime: 10m
maxNodeProvisionTime: 15m
expander: least-waste # Choose node group with least wasted resources
Real Results
Airbnb (KubeCon 2025):
- Implemented cluster autoscaler with custom expander
- Node count fluctuates from 500 (off-peak) to 2,000 (peak)
- Result: 35% reduction in node costs
- Peak-to-trough ratio: 4:1
Strategy 4: Vertical Pod Autoscaler (VPA)
VPA vs HPA
| Feature | VPA | HPA |
|---|---|---|
| Adjusts | CPU/memory requests | Number of replicas |
| When | Right-sizing | Traffic spikes |
| Downtime | Requires restart | No restart |
| Best for | Steady-state optimization | Burst handling |
VPA Configuration
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: my-app-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
updatePolicy:
updateMode: "Auto"
resourcePolicy:
containerPolicies:
- containerName: "*"
minAllowed:
cpu: "100m"
memory: "256Mi"
maxAllowed:
cpu: "4000m"
memory: "8Gi"
Strategy 5: Namespace Resource Quotas
Prevent Cost Runaway
apiVersion: v1
kind: ResourceQuota
metadata:
name: team-quota
namespace: team-backend
spec:
hard:
requests.cpu: "20"
requests.memory: "40Gi"
limits.cpu: "40"
limits.memory: "80Gi"
pods: "50"
persistentvolumeclaims: "10"
This prevents any single team from consuming more than their allocated resources.
Strategy 6: Karpenter (AWS) or Node Auto-Provisioning (GKE)
Karpenter (AWS)
Karpenter is a node provisioner that automatically selects optimal instance types:
- Considers: Price, availability, workload requirements
- Consolidation: Moves pods to fewer nodes when possible
- Result: AWS reports 50-60% cost reduction vs default Cluster Autoscaler
Node Auto-Provisioning (GKE)
Google's equivalent:
- Automatically creates node pools based on workload patterns
- Uses spot instances for batch workloads
- Result: 30-40% cost reduction reported by Google Cloud customers
Cost Monitoring Tools
| Tool | Type | Cost |
|---|---|---|
| Kubecost | Commercial | Free tier + $449/mo |
| OpenCost | Open-source | Free |
| Goldilocks | Open-source | Free |
| CloudHealth | Commercial | Custom pricing |
| Infracost | IaC cost estimation | Free tier + $20/mo |
OpenCost Setup
# Install OpenCost
kubectl apply --server-side -f https://raw.githubusercontent.com/opencost/opencost/develop/kubernetes/opencost.yaml
# View cost dashboard
kubectl port-forward -n opencost svc/opencost 9090:9090
Real Company Results
| Company | Strategy | Savings | Timeline |
|---|---|---|---|
| Shopify | Right-sizing + Spot | $1.8M/year | 3 months |
| Airbnb | Autoscaler + VPA | 35% reduction | 2 months |
| Netflix | Spot instances (70%) | $200M/year | Ongoing |
| Spotify | VPA + Right-sizing | 40% reduction | 6 months |
| DoorDash | Karpenter + Spot | 50% reduction | 4 months |
Implementation Checklist
- Week 1: Install OpenCost/Goldilocks, identify waste
- Week 2: Right-size top 20% of pods (by cost)
- Week 3: Enable spot instances for non-critical workloads
- Week 4: Configure Cluster Autoscaler with optimal settings
- Week 5: Implement VPA for steady-state optimization
- Week 6: Set up namespace quotas and cost alerts
- Ongoing: Monthly cost reviews, continuous optimization
Frequently Asked Questions
How much can I save with Kubernetes optimization?
Based on real company data, most teams save 30-50% on their Kubernetes costs within 2-3 months. The biggest wins come from right-sizing (30-40%) and spot instances (60-80% for eligible workloads). Companies implementing comprehensive FinOps practices save an average of $2.4M per $10M in annual cloud spend.
What is the best Kubernetes cost optimization tool?
For most teams, OpenCost (free, open-source) is the best starting point. It provides real-time cost visibility per pod, namespace, and label. Pair it with Goldilocks for right-sizing recommendations. If you need more features (automated optimization, multi-cloud), Kubecost is the most popular commercial option.
Will cost optimization affect performance?
When done correctly, no. Right-sizing is based on actual usage data -- you're removing resources that were never used. Spot instances can affect availability if not configured properly, but with proper Pod Disruption Budgets and multi-AZ deployment, the impact is minimal. Start with non-critical workloads and expand gradually.
Conclusion
Kubernetes cost optimization isn't a one-time project -- it's an ongoing practice. The biggest wins come from three strategies: right-sizing (remove waste), spot instances (use cheap capacity), and autoscaling (match demand). Companies implementing these strategies consistently save 30-50% on their cloud bills within 2-3 months. Start with OpenCost to measure your current waste, then systematically address the top offenders.
Was this article helpful?
Stay in the loop
Get the latest tech news and AI insights delivered to your inbox. No spam, unsubscribe anytime.
TechVeb Team
Your trusted source for the latest in technology, AI innovations, and digital trends. We bring you in-depth analysis, expert reviews, and comprehensive guides.
Learn more about us →Continue Reading
View all →
Ansible DevOps Guide 2026: Playbooks & Automation
Master Ansible automation for DevOps in 2026. Learn infrastructure playbooks, roles, inventory management, and configuration best practices.

AWS Guide for Beginners (2026): EC2, S3, Lambda & RDS
Master Amazon Web Services in 2026. Learn EC2, S3, Lambda, and RDS with practical examples in this complete beginner's guide to AWS cloud computing.

Microsoft Azure for Beginners: Complete 2026 Guide
Learn Microsoft Azure cloud fundamentals in 2026. Explore virtual machines, App Service, serverless Azure Functions, and enterprise integration easily.

Azure vs AWS vs GCP (2026): Best Cloud Comparison
Compare Azure, AWS, and GCP in 2026. Explore pricing, features, AI capabilities, and key strengths to choose the right cloud provider for your business.

CI/CD Pipeline Best Practices for 2026: Full Guide
Master CI/CD pipeline best practices in 2026. Compare GitHub Actions, GitLab CI, Jenkins, and CircleCI to boost speed, security, and release velocity.

Cloud Cost Optimization: How to Cut Cloud Bills by 60%
Learn proven cloud cost optimization strategies for 2026. Reduce your AWS, Azure, and GCP bills by up to 60% with right-sizing, spot instances, and FinOps.