Kubernetes makes compute feel free right up until the invoice arrives. Last year I walked into a SaaS spending $310k a month with utilization under 20%. Three months later the bill was down 29% and reliability was up. No re-architecture, no heroics — just four levers, pulled in order.
1. Rightsize requests before anything else
Over-provisioned requests are the silent majority of Kubernetes waste. We sampled actual usage over two weeks, set requests at p95 plus headroom, and enforced them with a policy that warns for a sprint before it blocks. Developers adjusted once and never thought about it again.
2. Make limits boring
Limits without load testing are superstition. We load-tested the top twenty services, set limits from data, and deleted the copy-pasted values that had throttled the checkout service every Black Friday for three years.
- Requests from measured p95, reviewed quarterly — not from vibes.
- Limits from load tests, with alerts on throttling before users notice.
- Burstable classes for batch, guaranteed only where latency pays for it.
3. Autoscale both directions
Everyone configures scale-up; almost nobody tunes scale-down. Karpenter with consolidation enabled, plus scheduled scaling for the nightly batch window, removed a third of our nodes between 8pm and 6am.
4. Showback per namespace, then chargeback
We published a weekly cost-per-team report with trend lines and named the top five idle resources. Social pressure did more in a month than a year of tickets. When finance later asked for chargeback, the data was already trusted.
Total: 29% off the monthly burn, zero incidents caused by the program, and a platform team that now reviews cost in the same standup as latency.
- Kubernetes
- FinOps
- Autoscaling
- EKS