Kubernetes Cost Management: What Actually Drives the Bill
Ask most platform teams why the cluster costs what it does, and the honest answer is nobody knows. That's the problem to fix before any optimization tool helps.
Part 1 of 3 in the Kubernetes Cost Management series. Read the series overview.
Every platform team I’ve worked with can tell me their monthly cloud bill to the dollar. Almost none of them can tell me why it’s that number instead of half that number. Kubernetes made deploying and scaling applications easy. It did not make it easy to know what any of that costs, and the gap between those two things is where most of the waste lives.
A 2023 Sysdig report put an actual number on it: 69% of requested CPU in the average Kubernetes environment goes unused. Reserved, paid for, sitting idle. That’s not a rounding error in anyone’s budget.
Cost management, stripped down to what it actually is, means making sure your environment is efficient: catching the places where resources are over-allocated or under-allocated, then adjusting to match what the workload actually needs. Two examples I run into constantly:
A workload requests 1 GiB of memory and 1 CPU. It uses 500 MiB and 0.5 CPU. That gap, multiplied across a few hundred pods, is real money sitting idle for no reason.
A 6-node cluster, each node 4 CPUs and 16 GB RAM, running 10 microservices at 40% CPU and 60% memory utilization. That cluster can lose a node and nobody notices, except finance.
Where the money actually goes
Cluster cost breaks into three categories, and they don’t behave the same way:
Resource allocation costs. CPU, GPU, memory, disk, load balancers, static IPs. You pay for what you’ve reserved, not what you use, which is exactly why over-provisioning stays invisible until someone finally adds it up.
Resource usage costs. Network bandwidth and anything else billed on actual consumption instead of allocation.
Overhead costs. Control plane fees, node licensing (OpenShift is the obvious one), your observability stack’s subscription, and the platform team’s own time keeping the thing running. None of this shows up on a per-workload line, so it’s the easiest category to lose track of entirely.

Cluster cost, broken down by category.
Containers are the smallest unit of allocation, and that’s where I’d start looking first. An over-allocated container is the root of most allocation waste further up the stack. I’ve seen teams request 8 CPUs and 16 GB of memory for a service that needs 4 CPUs and 8 GB, because that’s what the last service on the team needed and nobody revisited the number. In a managed cluster that’s straightforward overpayment. On-prem, it’s worse: it forces you to buy hardware you didn’t need.
Storage costs come from the same instinct, applied to volumes instead of compute. Persistent volumes that outlive the workload that created them are the classic case: a team spins up a temporary environment, it writes a large log volume, the environment gets torn down, and the volume doesn’t. Nobody’s watching it, so it just accrues cost month after month until someone doing a storage audit finds it by accident.
Network cost is the one that catches people the most, because it’s invisible until a multi-AZ deployment is already live in production. Egress to the internet, for your APIs and customer-facing traffic, is usually the biggest line item. Cross-zone traffic between workloads is the sneaky one: nodes get spread across availability zones for resilience, and every inter-workload call that crosses a zone boundary gets billed, with no single dashboard telling you which workload is responsible.
Kafka on Kubernetes is the clearest version of this I’ve seen. Brokers and their partitions end up spread across AZs. Every time a producer or consumer talks to a broker in a different zone, that’s a cross-AZ charge. Then Kafka’s own replication, which exists for fault tolerance, does the same thing again for every replica sync. Nobody designs a Kafka deployment thinking about availability-zone billing, but by the time it’s running in production at real traffic, that cross-zone chatter is a cost line of its own.
Why teams lose track of it
Four problems come up over and over, roughly in this order of how much damage they do.
No visibility into what’s driving cost. Cloud billing tells you what you spent in total. It doesn’t tell you which workload spent it. Platform teams end up asking two questions they can’t answer: how efficiently is each workload actually using what it’s been given, and what does this specific workload or team cost us. Kubernetes’ own elasticity makes this worse. Workloads and nodes scale up and down constantly, so a bill spike doesn’t come with an obvious culprit attached. Add in dormant workloads still running for no reason, and you’ve got spend nobody can explain.
No feedback loop on rightsizing. Rightsizing, at the workload and cluster level, is the single most effective lever you have. The problem isn’t that teams don’t rightsize. It’s that they do it once, guess generously to “play it safe,” and never come back to check if the guess was right. T-shirt sizing gives engineers a starting point, not a correction mechanism. Without usage data feeding back into that decision, “which workloads are actually over-provisioned” stays an unanswered question indefinitely.
No accountability. This is the one enterprise customers bring up unprompted. One I worked with put it about as plainly as it gets: the platform team was absorbing most of the cost, and they wanted a transparent, fair chargeback model, because once a team is accountable for its own share, it starts taking optimization seriously. That’s the whole argument for chargeback in one sentence. But you can’t build a chargeback model on top of a visibility problem you haven’t solved. Attribution has to come first, or the “fair” part of a fair chargeback model is fiction.
Multi-cloud and hybrid complexity. This doesn’t introduce a new problem so much as it multiplies the first three across more billing systems, each with its own pricing model and its own blind spots.
What this means in practice
None of this is really an argument for buying a new tool. It’s an argument for treating cost the way you’d treat latency or error rate: a signal you watch continuously, not a spreadsheet exercise someone does once a quarter after finance asks a question nobody can answer on the spot. Visibility has to come first. Rightsizing and accountability only work once you actually have it.
This article was originally published on the Randoli blog.
