Nameplate Analytics LLC Book a snapshot
GPU cost measurement · AWS + Kubernetes

You are paying for GPUs that never did the work.

Your utilization panel reports that a kernel was resident, not that the chip did anything with it. On real hardware we measured that gap at one hundred times. We measure what fraction of your GPU spend produced work, join it to your AWS bill, and rank what is recoverable.

Telemetry capture Tesla T4 · g4dn.xlarge driver 595.91.07 DCGM @ 1 Hz 242 samples 2026-09-04
Streaming multiprocessors workload: 64-element add, looped

1/40 multiprocessors active Every dark block is silicon on the invoice doing nothing. The dashboard for this exact moment read 20 percent.

Reported · GPUTLwhat your panel shows
20.0%

DCGM_FI_DEV_GPU_UTIL

Measured · SMACTmultiprocessors working
0.2%

DCGM_FI_PROF_SM_ACTIVE

Measured · TENSOtensor cores working
0.0%

DCGM_FI_PROF_PIPE_TENSOR_ACTIVE

The reported number was one hundred times the measured one. Not a broken setup: the metric answers a different question than the one everybody asks it. All 242 samples, all six fields, including the control run that saturates the die, are in the full measurement record.

The problem

Nobody in your org can answer what fraction of the GPU bill produced work.

Not because your team is careless. Because answering it needs two datasets joined that almost nobody has both halves of, and the number everyone reaches for instead is measuring something else entirely.

01

The metric is the wrong metric

DCGM_FI_DEV_GPU_UTIL reports whether any kernel was resident during the sample window. It says nothing about how much of the chip that kernel used. One tiny kernel on one of forty multiprocessors scores identically to a workload filling all forty.

02

The bill hides the expensive machines

Commitment-covered resources report $0 unblended, so your most expensive GPU node vanishes from a naive cost ranking. Untagged spend has no owner to ask. Amortized cost per resource is the only view that survives contact with finance.

03

So capacity planning is a guess

Every request for more GPUs gets approved against a number that can be wrong by two orders of magnitude, and no one in the room has the instrument to say so. The gap compounds every quarter the fleet grows.

The method

Instrument the hardware. Join it to the bill. Recover what the evidence supports.

01

Measure

DCGM profiling telemetry per workload, exported to Prometheus: multiprocessor activity, warp occupancy, tensor pipe activity, memory bandwidth. Real work, sampled per second, attributable to a pod. Most clusters have never had this turned on, so step one is usually building the instrument before reading it.

02

Attribute

A CUR 2.0 export at hourly granularity with resource IDs, normalized and amortized, then joined to the cluster through provider IDs. The output is cost per GPU resource, per team, per namespace, per model, and the share of it that carries no owner at all.

03

Recover

A remediation sequence ordered by dollars per unit of disruption. Requests right-sized against measured p95 plus headroom, time-slicing or MIG where workloads are small enough to share, bin-packing and autoscaler tuning so scale-in actually happens, idle reclamation with owner notification. Executed with your engineers, in your change windows.

The difference

Every number we hand you carries its evidence.

The industry norm is one confident figure that does not survive its first contact with a skeptical staff engineer. Ours is lower and it holds up in the room.

PROVEN
Your billing data alone establishes it. No assumption required and nothing to argue about.
LIKELY
Billing data plus exactly one assumption, written out in full so your engineers can accept or reject it on the spot.
UNPROVEN
A real cost and a well-founded suspicion, waiting on data that does not exist in your environment yet.

Unproven dollars are excluded from the recoverable total, and the exclusion is enforced in the tooling rather than promised in a footnote. Under client pressure, discipline is the thing that bends. Code does not.

Fit

This is a narrow practice, on purpose.

Worth a conversation

  • Above $50,000 a month in GPU spend. That is the gate. Total cloud spend and headcount are secondary.
  • Training or inference on Kubernetes: EKS, GKE, or self-managed.
  • No dedicated FinOps function, or one person doing it alongside another job.
  • An engineering leader who can approve a five-figure invoice without a procurement cycle.

Not worth either of our time

  • Under $20,000 a month in total cloud spend. The fee would not clear the finding.
  • Fully managed inference only, such as Bedrock or serverless SageMaker. There is no cluster to instrument.
  • A FinOps team already exporting GPU telemetry to Prometheus. You have done the hard part.
  • Anyone whose opening question is what percentage of savings we take.
Engagements

Four ways in. Fixed fee, every one.

Never a percentage of savings. Contingency pricing gives the person measuring your infrastructure a financial stake in the size of the number, which is the one incentive you do not want in the room.

01

Spend Snapshot

Three numbers, live on one call: amortized GPU spend per resource for the last full month, the share carrying no owner tag, and how many provisioned GPUs have zero pods requesting one. That third number needs no new instrumentation and is usually the one that lands.

No charge45 minutes
02

Cost Attribution Report

Twelve to eighteen pages. Cost per GPU resource, team, namespace and model. Ranked recoverable spend with a dollar figure and a confidence rating on every line, the evidence gaps named explicitly, and the instrumentation plan if you do not have telemetry yet.

$18,000 – $25,000two weeks
03

Recovery Sprint

Priced from the findings of the report, scoped to the specific items named in the SOW. Telemetry deployed, requests right-sized against measured p95, sharing where it fits, autoscaler tuned so scale-in happens, idle reclamation with owner notification.

$45,000 – $70,000six to eight weeks
04

Cost Watch

Right-sizing that runs once at deployment is not right-sizing. Quarterly attribution against the same baseline, drift analysis, alerting on new unowned GPU spend and on allocation without utilization, commitment review, standing engineering channel.

$12,000 – $15,000per quarter
Start here

Forty-five minutes and read-only access is enough to know whether there is anything here.

One month of billing data and fifteen minutes describing your cluster. If the numbers come back clean, we will say so on the call and that is the end of it.

LinkedIn
Kam Kheri
Response time
Same business day