Kubernetes makes compute feel free right up until the invoice arrives. Last year I walked into a SaaS spending $310k a month with utilization under 20%. Three months later the bill was down 29% and reliability was up. No re-architecture, no heroics — just four levers, pulled in order.

1. Rightsize requests before anything else

Over-provisioned requests are the silent majority of Kubernetes waste. We sampled actual usage over two weeks, set requests at p95 plus headroom, and enforced them with a policy that warns for a sprint before it blocks. Developers adjusted once and never thought about it again.

2. Make limits boring

Limits without load testing are superstition. We load-tested the top twenty services, set limits from data, and deleted the copy-pasted values that had throttled the checkout service every Black Friday for three years.

  • Requests from measured p95, reviewed quarterly — not from vibes.
  • Limits from load tests, with alerts on throttling before users notice.
  • Burstable classes for batch, guaranteed only where latency pays for it.
You can't optimize what you can't attribute. Showback comes before savings, always.

3. Autoscale both directions

Everyone configures scale-up; almost nobody tunes scale-down. Karpenter with consolidation enabled, plus scheduled scaling for the nightly batch window, removed a third of our nodes between 8pm and 6am.

4. Showback per namespace, then chargeback

We published a weekly cost-per-team report with trend lines and named the top five idle resources. Social pressure did more in a month than a year of tickets. When finance later asked for chargeback, the data was already trusted.

Total: 29% off the monthly burn, zero incidents caused by the program, and a platform team that now reviews cost in the same standup as latency.

  • Kubernetes
  • FinOps
  • Autoscaling
  • EKS