Kubernetes Cost Management, Part 2: From Visibility to Continuous Optimization
Knowing a cluster costs $20,000 a month doesn't tell you what to do about it. Optimization has to run as a feedback loop, not a periodic cleanup.
Part 2 of 3 in the Kubernetes Cost Management series. Read the series overview.
From Cost Visibility to Continuous Optimization
In Part 1 of this series, we looked at where Kubernetes costs come from and why managing them becomes difficult as environments grow.
Understanding the bill is an important first step. But knowing that a cluster costs $20,000 a month doesn’t necessarily tell you what to do about it. The harder questions are:
- Why does it cost $20,000?
- Which applications and workloads are responsible for that cost?
- Are they using the resources being allocated to them?
- If we reduce those resources, what happens to the application?
- And if we find an optimization opportunity today, how do we prevent the same inefficiency from appearing again six months from now?
This is where I think we need to change how we approach Kubernetes cost optimization. It shouldn’t be a periodic exercise where someone looks at the cloud bill, finds a few expensive resources and starts cutting.
Kubernetes environments don’t stay static. Applications change. Traffic patterns change. New workloads get deployed. Resource requests get increased to solve performance problems. Clusters scale. New node pools appear. Teams change the way applications are deployed. Something that was appropriately sized six months ago may be significantly over-provisioned today.
For that reason, I think Kubernetes cost optimization needs to operate as a continuous feedback loop:
Observe → Identify → Recommend → Change → Validate → Repeat
Let’s look at what that means in practice.
Start by Making Cost Actionable
Most organizations already have some visibility into their cloud spend. AWS, Azure and GCP all provide detailed billing information, and FinOps teams may already have dashboards showing cloud spend by account, service or business unit.
But cloud cost visibility and Kubernetes cost visibility are not the same thing. A cloud bill might tell you that a particular account spent $80,000 on compute last month. Even if you narrow that down to an EKS, AKS or GKE cluster, you still haven’t answered the question an engineering team needs answered: which workloads are responsible for that cost?
For Kubernetes, we need to be able to move through several levels of attribution:
Cloud Account → Cluster → Namespace → Workload → Pod → Container
And eventually we need to connect those resources to organizational concepts:
Team → Application → Product → Cost Center
That is what makes cost actionable. Suppose Kubernetes spend increases 25% over a two-week period. Knowing that the increase came from three clusters is useful. Knowing that most of the increase came from two workloads is much more useful. Knowing that one of those workloads doubled its CPU requests after a deployment gives an engineer somewhere to start investigating.
This is where cost starts moving beyond reporting and becomes another form of operational telemetry.
Establish a Baseline Before Optimizing
Once you can attribute cost, the temptation is to immediately start looking for things to reduce. Before doing that, establish a baseline.
Consider a container requesting 2 CPUs but currently using 500m. It looks dramatically oversized. But what does “currently” mean? The last five minutes? The last hour? The last seven days? What happens every weekday at 9 AM? What happened during the end-of-month processing cycle? What happened during the last traffic spike?
Point-in-time utilization isn’t enough to make a good rightsizing decision. Averages aren’t necessarily enough either. For a workload, we want to understand the relationship between:
Requested Resources → Actual Usage → Workload Behavior → Cost
over a meaningful period of time.
The same principle applies at the cluster level. If a 20-node cluster spends most of the day at 35% utilization, that could represent a significant optimization opportunity. Or perhaps those nodes are needed for a predictable peak every evening. Those are two very different situations. Before optimizing, understand what normal looks like.
Start With the Workload, Not the Node
When investigating Kubernetes cost, one of the first places I look is resource requests.
Let’s take a simple example. Assume an application has 20 pods. Each pod requests 2 CPUs but normally consumes around 500m. At first glance, you might say the application is using only 25% of the CPU allocated to it. But the impact goes beyond unused CPU.
Kubernetes scheduling decisions are based on resource requests. The scheduler needs to find nodes capable of satisfying those requests regardless of whether the application ultimately consumes all of the requested capacity. If enough workloads are significantly over-requesting CPU or memory, Kubernetes needs more node capacity to schedule them, and if you use node autoscaling, that can ultimately result in additional infrastructure being provisioned.
So there is a relationship that is easy to overlook:
Container Requests → Pod Requests → Scheduling → Node Capacity → Infrastructure Cost
This is why I prefer to think about Kubernetes optimization from the workload outward. Rightsize the workloads first. Then look at how efficiently those workloads fit onto the infrastructure underneath them. Vertical Pod Autoscaler is the tool most teams reach for at this stage.
Rightsizing Isn’t About Finding the Smallest Number
If a container requests 4 CPUs but normally consumes less than 1 CPU, it can be tempting to recommend changing the request to 1 CPU. Mathematically, that might provide the largest saving. Operationally, it may be a terrible recommendation.
Production applications need headroom. Traffic spikes happen. Garbage collection creates temporary resource pressure. Batch workloads behave differently from APIs. Some applications have predictable daily peaks while others experience sudden bursts. Memory also behaves very differently from CPU: a CPU-constrained application may become slower, while an application that exceeds available memory can get killed.
The objective of rightsizing therefore isn’t:
How little can we give this workload?
It is:
What is an appropriate resource allocation based on how this workload actually behaves?
That requires historical data and context. For some applications, p95 utilization over a meaningful period plus reasonable headroom might provide a useful starting point. For others, p99 may matter. Some workloads may need to be evaluated around known business events rather than a generic 30-day average. There isn’t one formula that works for every workload, and that is why I am cautious when I see very large “potential savings” numbers generated purely from utilization data. A potential saving isn’t the same thing as a realized saving.
For a concrete, mode-by-mode approach to rolling this out safely, see VPA Recommendation Mode: A Practical Rightsizing Workflow, or for rightsizing beyond VPA (container, pod, namespace, and cluster levels), see Rightsizing Kubernetes Workloads: A Practical Guide.
A Rightsizing Recommendation Is Only the Beginning
Generating a recommendation is relatively easy. Let’s say the recommendation is:
Reduce CPU request from 2 cores to 1 core.
The important questions start after that recommendation is generated:
- What happens if we make the change?
- Does latency increase, or does CPU throttling increase?
- Do pods restart, or does the HPA start behaving differently?
- Does application throughput change?
- Does Kubernetes schedule the workload more efficiently, and does the cluster eventually release capacity?
- And ultimately: did we actually reduce cost?
This is where I think cost optimization and observability need to come together. Cost data can tell you where an optimization opportunity may exist. Metrics, logs, traces and application performance data help you determine whether the change is safe. And after the change is made, those same signals help determine whether the optimization worked.
A cost optimization system shouldn’t stop at “you could save $500 per month.” The feedback loop should end when we can say “we made the change, application behavior remained within our expected parameters, and we actually reduced infrastructure cost.” That is a much higher bar, but it is also much closer to how engineers actually want to operate production systems.
Workload Optimization Has a Horizontal Dimension Too
Rightsizing isn’t limited to CPU and memory requests. There is another question worth asking: do we need this many replicas?
Imagine a deployment running 20 pods where each pod is reasonably rightsized. If the application could meet its availability and performance requirements with 10 or 12 replicas for most of the day, optimizing CPU requests alone misses a significant part of the opportunity.
This is where horizontal scaling enters the picture. Horizontal Pod Autoscaling can increase and decrease replicas as demand changes, but just as with vertical rightsizing, the quality of the outcome depends on how scaling is configured. A workload with an overly aggressive minimum replica count may maintain substantial unused capacity, while a workload with an overly conservative scaling policy may create performance problems during spikes.
Cost optimization therefore needs to consider both:
- Vertical efficiency: How much CPU and memory does each instance need?
- Horizontal efficiency: How many instances do we need?
Those decisions eventually affect the infrastructure required underneath them.
Then Optimize the Cluster
Once workload requests and replica counts reasonably reflect actual application requirements, cluster optimization becomes much more meaningful.
Suppose a cluster currently runs 20 nodes. After rightsizing workloads and improving scaling, the same workloads may fit comfortably on 16 nodes. Now we have moved from theoretical savings to an opportunity to remove actual infrastructure.
But simply looking at average cluster CPU utilization isn’t enough. You need to consider CPU and memory together, along with pod placement, node types, availability requirements, affinity rules, taints, topology constraints and intentional spare capacity.
The type of node also matters. A cluster may have plenty of CPU available while being constrained by memory. Adding more of the same instance type may solve the scheduling problem, but it may not be the most economical solution, a different CPU-to-memory ratio could fit the workloads better. The same problem becomes even more pronounced with GPU workloads, where expensive accelerators can sit significantly underutilized.
Cluster optimization is therefore not simply “how do I remove nodes?” It is “what infrastructure configuration best fits the workloads I actually need to run?”
Autoscaling Doesn’t Automatically Mean Cost Optimization
It is easy to assume that enabling autoscaling solves this problem. Kubernetes gives us several mechanisms: Horizontal Pod Autoscaling can change replica counts, Vertical Pod Autoscaling can adjust resource requests, and node autoscaling can add and remove infrastructure as scheduling requirements change.
These are powerful mechanisms, but automation doesn’t automatically produce efficiency. Consider workloads that consistently request four times more CPU than they need. A node autoscaler still has to provision enough infrastructure to satisfy those requests, it doesn’t know that the application developer was overly conservative when setting them. In other words, automation can automate inefficiency just as effectively as it can automate efficiency. Autoscaling works best when the signals and resource requirements driving it reasonably reflect the behavior of the applications.
There is also another trap: optimizing purely for utilization. A cluster running at 95% utilization might look fantastic from a cost perspective. It might also be one traffic spike away from a very bad day. Some spare capacity is intentional, production systems may need room for traffic bursts, node failures, rolling deployments or workloads being rescheduled. So I don’t think 100% utilization should ever be the objective, the goal should be intentional utilization.
If 20% of a cluster is unused, we should know why. If that capacity exists because the application needs to absorb sudden traffic spikes, that may be entirely reasonable. If it exists because workloads haven’t been rightsized in two years, that is a different problem.
Find the Waste Nobody Is Looking At
Not every optimization opportunity requires sophisticated modeling. Kubernetes environments accumulate things. A development environment gets created for a project and nobody removes it. A test workload stops being used but continues running. A persistent volume remains after its application disappears. A namespace that was supposed to exist for two weeks survives for two years. An application stops receiving meaningful traffic but still runs six replicas around the clock.
Individually, these may not attract attention. Across dozens or hundreds of clusters, they can add up to meaningful spend. Continuously identifying dormant and idle workloads is therefore one of the simpler ways to surface potential savings, and OpenCost is a free, low-effort way to start seeing this data.
The important word, however, is potential. Dormant doesn’t necessarily mean unnecessary. A disaster recovery service may be intentionally idle. A standby workload may exist specifically for a failure scenario. A development environment might only be needed once a month. The job of the system is to identify something worth investigating, the decision still needs context.
This distinction matters because one of the fastest ways to lose engineering teams’ trust in a cost optimization program is to repeatedly tell them that necessary capacity is “waste.”
Don’t Wait for the Monthly Cloud Bill
There is another part of cost optimization that I think should look more like traditional observability. We don’t wait until the end of the month to discover that application latency increased. Why should we wait until the end of the month to discover that an application’s infrastructure cost increased 40%?
Cost is a signal. Suppose a service normally costs around $300 per day. After a deployment, it starts costing $500 per day. That doesn’t necessarily mean something is wrong, perhaps traffic increased significantly, or a deliberate architecture change required additional capacity. But it is worth knowing that the change occurred.
The interesting question is not simply “did cost increase?” It is “why did cost increase?”
- Was there a deployment?
- Did replica count increase, or did resource requests change?
- Did traffic increase, or did a new workload appear?
- Did infrastructure pricing change?
This is where connecting cost data to operational context becomes particularly valuable. Instead of treating cost anomalies as something a FinOps team investigates separately, they can become part of the same operational picture engineers already use to understand changes in their systems.
Cost Without Ownership Is Just a Report
Eventually, Kubernetes cost optimization stops being purely a technical problem. The platform team can identify that a workload appears oversized, but the platform team may not know whether that workload can safely be changed. The application team usually does. That means costs need owners.
At a technical level, Kubernetes gives us useful dimensions such as namespaces, workloads and labels. Those can then be mapped to organizational structures:
Workload → Application → Team → Business Unit
This doesn’t mean every organization needs to immediately implement chargeback. For many organizations, showback is the more useful starting point: give engineering teams visibility into what their applications cost, how that cost changes, which workloads are driving it, and where there appear to be optimization opportunities.
Now the conversation changes. Instead of the platform team saying:
“You need to reduce your CPU requests.”
the discussion becomes:
“This application is requesting significantly more CPU than it typically consumes. Here is what that capacity costs, here is how the workload has behaved over the last 30 days, and here is the estimated impact of reducing it.”
That is a much more useful engineering conversation. More mature organizations can take the same allocation model further into formal chargeback and unit economics, such as cost per service, tenant or other business-relevant unit.
Closing the Loop With Automation
So far, most of the optimization process still assumes a human is doing the work. The system identifies an opportunity, an engineer investigates it, someone decides what to change, the change is implemented, then somebody checks whether it worked. There is good reason for that caution, especially in production environments.
But this is also where I think AI agents have the potential to change the cost optimization workflow. Consider a rightsizing recommendation: instead of simply telling an engineer that a workload appears oversized, an SRE agent could investigate further. It could look at historical resource usage, examine application performance, check whether there were recent deployments or unusual traffic patterns, and determine whether the proposed change fits within an organization’s established policies. And rather than simply producing another recommendation on a dashboard, it could initiate an approved remediation workflow.
That doesn’t have to mean giving an AI agent unrestricted permission to change production. There are several levels of automation:
- At one level, the agent can investigate and recommend.
- At another, it can prepare the change, perhaps by creating a pull request, and wait for an engineer to approve it.
- For well-understood, low-risk scenarios, an organization might allow an agent to execute a predefined runbook or remediation workflow automatically.
The important part is that the process shouldn’t stop when the change is made. The agent can continue observing the workload after remediation: did resource utilization improve, did latency or error rates change, did the workload remain healthy, was the expected saving actually realized? If not, the workflow can flag the result for investigation or, where appropriate, invoke a rollback process.
Now our feedback loop becomes:
Detect → Investigate → Recommend → Approve → Remediate → Validate
This is where cost optimization starts looking much more like an SRE practice.
From Cost Optimization to a FinOps Practice
Not every organization needs to implement everything we’ve discussed on day one.
If you are starting with nothing more than a monthly cloud bill, being able to accurately understand Kubernetes cost by cluster, namespace and workload is already a significant step forward. The next step might be identifying obvious waste and rightsizing opportunities. Then you can introduce continuous monitoring, budgets and cost alerts. As the practice matures, you can establish ownership, showback and potentially chargeback. Eventually, parts of the optimization and remediation process can become increasingly automated.
This progression matters. Trying to automate remediation before you trust your cost allocation or rightsizing recommendations probably isn’t a good idea. Similarly, implementing chargeback when teams don’t trust how their costs are calculated is likely to create more arguments than accountability.
FinOps describes this progression through a maturity model, often expressed as Crawl, Walk, Run. I think that model is particularly useful for Kubernetes because the technical and organizational maturity have to develop together. We’ll explore that in the next part of this series, including what Crawl, Walk and Run actually look like for Kubernetes and where automation and AI agents fit into that progression.
If you’d rather not build this maturity model piece by piece, Cost & FinOps is the Randoli capability built around exactly this: attribution, rightsizing, and chargeback in one place. Curious how that compares to a dedicated cost tool? See Randoli vs Kubecost.
Kubernetes Cost Optimization Is a Feedback Loop
If there is one idea I would take away from this article, it is that Kubernetes cost optimization isn’t a one-time project. It’s a feedback loop:
Observe → Identify → Recommend → Change → Validate → Repeat
Observe cost and resource consumption. Identify where inefficiencies exist. Make recommendations using historical data and application context. Apply changes with the appropriate controls. Validate the effect on cost, performance and reliability. Then repeat as the environment changes.
That is why I think Kubernetes cost optimization should be approached more like an SRE practice than an annual budgeting exercise. We don’t optimize application performance once a year and assume the system will remain optimized, we continuously observe it because the system continuously changes. Cost is no different. And increasingly, we have an opportunity to close more of that loop through automation, while still maintaining the controls required for production environments.
The goal isn’t simply to make Kubernetes cheaper. It’s to continuously understand the relationship between cost, capacity, performance and reliability, and use that information to make better engineering decisions. That’s what sustainable Kubernetes cost optimization looks like.
Related reading
- Kubernetes Cost Management, Part 1: What Actually Drives the Bill: the cost factors and challenges this article builds on.
- A Guide to Kubernetes VPA and VPA Recommendation Mode: A Practical Rightsizing Workflow: the tooling behind the workload-level rightsizing discussed above.
- Rightsizing Kubernetes Workloads: A Practical Guide: rightsizing across container, pod, namespace, and cluster levels.
- 5 Tips for Monitoring Kubernetes Spend with OpenCost: a free starting point for the visibility this article assumes.
- Cost & FinOps: attribution, rightsizing, and chargeback without assembling this yourself.
This article was originally published on the Randoli blog.
