You are paying for GPUs that never did the work.
Your utilization panel reports that a kernel was resident, not that the chip did anything with it. On real hardware we measured that gap at one hundred times. We measure what fraction of your GPU spend produced work, join it to your AWS bill, and rank what is recoverable.
1/40 multiprocessors active Every dark block is silicon on the invoice doing nothing. The dashboard for this exact moment read 20 percent.
DCGM_FI_DEV_GPU_UTIL
DCGM_FI_PROF_SM_ACTIVE
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE
The reported number was one hundred times the measured one. Not a broken setup: the metric answers a different question than the one everybody asks it. All 242 samples, all six fields, including the control run that saturates the die, are in the full measurement record.
Nobody in your org can answer what fraction of the GPU bill produced work.
Not because your team is careless. Because answering it needs two datasets joined that almost nobody has both halves of, and the number everyone reaches for instead is measuring something else entirely.
The metric is the wrong metric
DCGM_FI_DEV_GPU_UTIL reports whether any kernel was resident during the sample window. It says nothing about how much of the chip that kernel used. One tiny kernel on one of forty multiprocessors scores identically to a workload filling all forty.
The bill hides the expensive machines
Commitment-covered resources report $0 unblended, so your most expensive GPU node vanishes from a naive cost ranking. Untagged spend has no owner to ask. Amortized cost per resource is the only view that survives contact with finance.
So capacity planning is a guess
Every request for more GPUs gets approved against a number that can be wrong by two orders of magnitude, and no one in the room has the instrument to say so. The gap compounds every quarter the fleet grows.
Instrument the hardware. Join it to the bill. Recover what the evidence supports.
Measure
DCGM profiling telemetry per workload, exported to Prometheus: multiprocessor activity, warp occupancy, tensor pipe activity, memory bandwidth. Real work, sampled per second, attributable to a pod. Most clusters have never had this turned on, so step one is usually building the instrument before reading it.
Attribute
A CUR 2.0 export at hourly granularity with resource IDs, normalized and amortized, then joined to the cluster through provider IDs. The output is cost per GPU resource, per team, per namespace, per model, and the share of it that carries no owner at all.
Recover
A remediation sequence ordered by dollars per unit of disruption. Requests right-sized against measured p95 plus headroom, time-slicing or MIG where workloads are small enough to share, bin-packing and autoscaler tuning so scale-in actually happens, idle reclamation with owner notification. Executed with your engineers, in your change windows.
Every number we hand you carries its evidence.
The industry norm is one confident figure that does not survive its first contact with a skeptical staff engineer. Ours is lower and it holds up in the room.
Unproven dollars are excluded from the recoverable total, and the exclusion is enforced in the tooling rather than promised in a footnote. Under client pressure, discipline is the thing that bends. Code does not.
This is a narrow practice, on purpose.
Worth a conversation
- Above $50,000 a month in GPU spend. That is the gate. Total cloud spend and headcount are secondary.
- Training or inference on Kubernetes: EKS, GKE, or self-managed.
- No dedicated FinOps function, or one person doing it alongside another job.
- An engineering leader who can approve a five-figure invoice without a procurement cycle.
Not worth either of our time
- Under $20,000 a month in total cloud spend. The fee would not clear the finding.
- Fully managed inference only, such as Bedrock or serverless SageMaker. There is no cluster to instrument.
- A FinOps team already exporting GPU telemetry to Prometheus. You have done the hard part.
- Anyone whose opening question is what percentage of savings we take.
Four ways in. Fixed fee, every one.
Never a percentage of savings. Contingency pricing gives the person measuring your infrastructure a financial stake in the size of the number, which is the one incentive you do not want in the room.
Spend Snapshot
Three numbers, live on one call: amortized GPU spend per resource for the last full month, the share carrying no owner tag, and how many provisioned GPUs have zero pods requesting one. That third number needs no new instrumentation and is usually the one that lands.
Cost Attribution Report
Twelve to eighteen pages. Cost per GPU resource, team, namespace and model. Ranked recoverable spend with a dollar figure and a confidence rating on every line, the evidence gaps named explicitly, and the instrumentation plan if you do not have telemetry yet.
Recovery Sprint
Priced from the findings of the report, scoped to the specific items named in the SOW. Telemetry deployed, requests right-sized against measured p95, sharing where it fits, autoscaler tuned so scale-in happens, idle reclamation with owner notification.
Cost Watch
Right-sizing that runs once at deployment is not right-sizing. Quarterly attribution against the same baseline, drift analysis, alerting on new unowned GPU spend and on allocation without utilization, commitment review, standing engineering channel.
Forty-five minutes and read-only access is enough to know whether there is anything here.
One month of billing data and fifteen minutes describing your cluster. If the numbers come back clean, we will say so on the call and that is the end of it.
- Direct
- 301-802-1073
- Kam Kheri
- Response time
- Same business day