Nameplate Analytics

GPU cost attribution

Northwind Systems, Inc.

Report date
2026-09-08
Data period
2026-08-01 to 2026-08-30
Accounts
2 in scope
Findings
9

1 What we are willing to claim

Every finding in this report carries a confidence rating. Only findings the data supports are added to the recoverable total. The rest are printed in full and left out of the number, because a figure that does not survive a skeptical engineer is worth less than no figure at all.

$163k
recoverable per month, proven and likely findings only
$1.95M
recoverable per year at the same run rate
$665k
observed but not claimed, excluded from the figures to the left
$107k$665kProven recoverable$107k per monthLikely recoverable$55k per month · one stated assumption eachNot claimed$665k per month · excluded from the total
The hatched segment is money we are not claiming. It is shown at the same scale as the rest so the proportion is visible rather than described.

Total amortized spend across the period was $485k, of which $476k was GPU compute. Amortized cost is used throughout, not unblended cost: AWS writes an unblended cost of zero on reserved-instance-covered line items, so a ranking by unblended cost makes the most expensive machine in an account disappear.

2 How to read a finding

PROVEN

The billing data proves both the cost and the waste. No assumption about how the workload behaves is involved.

LIKELY

The billing data proves the cost. The waste follows from one stated assumption, printed with the finding, which you can accept or reject line by line.

UNPROVEN

There is a cost and a suspicion. The data that would settle it does not exist yet. These dollars are listed and then excluded from the recoverable total.

3 Findings

PROVEN

$107k per month recoverable across 3 findings.

23 of 63 provisioned GPUs have no pod requesting them, $100,846/mo

$101k
observed per month
$101k
recoverable per month
3
resources
2/5
disruption to change

What to do.Scale the affected node groups down to the number of GPUs actually requested. Where a node group exists for burst capacity, set its floor to zero and let the cluster autoscaler bring it up on a pending pod.

nodeinstanceinstanceGPUsGPUs requested$ / month
ip-10-0-001-10.ec2.internali-000000000000001001p5.48xlarge80$71,734
ip-10-0-000-10.ec2.internali-000000000000001000p4d.24xlarge80$14,825
ip-10-0-002-10.ec2.internali-000000000000001002g6e.xlarge10$842

10 tagged non-production instances run continuously, $6,495/mo

$6,495
observed per month
$4,562
recoverable per month
10
resources
2/5
disruption to change

What to do.Stop non-production instances outside working hours with an EventBridge schedule or an autoscaler floor of zero. GPU instances should be stopped rather than resized, because the accelerator is the entire cost.

resourceaccountinstancegpudays always on$ / monthwhy it was classified this way
i-00000000000000100b111111111111g4dn.12xlargeyes30$2,854Environment tag = dev
i-00000000000000100a111111111111g5.xlargeyes30$734Environment tag = dev
i-00000000000000100d111111111111g5.xlargeyes30$734Environment tag = dev
i-000000000000009007111111111111c5.9xlargeno30$692Environment tag = dev
i-00000000000000100c111111111111g4dn.xlargeyes30$384Environment tag = dev
i-000000000000009001111111111111m5.4xlargeno30$347Environment tag = dev
i-000000000000009010111111111111m5.4xlargeno30$347Environment tag = dev
i-00000000000000900d111111111111r5.2xlargeno30$228Environment tag = dev
i-000000000000009004111111111111m6i.xlargeno30$87Environment tag = dev
i-000000000000009013111111111111m6i.xlargeno30$87Environment tag = dev

Commitment purchased and not consumed: $1,216/mo of Reserved Instance capacity and $760/mo of Savings Plan commitment

$1,976
observed per month
$1,976
recoverable per month
2
resources
2/5
disruption to change

What to do.Sell unused Standard RIs on the Reserved Instance Marketplace where the term allows, or move workloads onto the committed instance families so the commitment is consumed. Savings Plan commitments cannot be sold; the recovery is to shift eligible usage onto them before the term ends.

resourceaccountcommitmentusage typeunused $ / month
arn:aws:ec2:us-east-1:ri/r-abc111111111111RIFeeRIFee$1,216
arn:aws:savingsplans::sp/sp-xyz111111111111SavingsPlanRecurringFeeRIFee$760

LIKELY

$55k per month recoverable across 3 findings.

14 production GPU instances run continuously at on-demand rates, $232,885/mo

$233k
observed per month
$47k
recoverable per month
14
resources
3/5
disruption to change

Assumption.The recoverable figure assumes a 20% rate reduction from a one-year no-upfront Compute Savings Plan, and assumes this capacity is still wanted twelve months from now. Verify the current discount for these families against AWS pricing before committing; if the workload is being migrated or retired inside the term, this finding does not apply.

What to do.Size a Compute Savings Plan to the observed floor of GPU usage, not to the peak. Commit to the level that has been running every day of the period and leave the rest on demand.

resourceaccountinstanceacceleratordays always on$ / month
i-000000000000001001111111111111p5.48xlargeNVIDIA H10030$71,734
i-000000000000001005111111111111p5.48xlargeNVIDIA H10030$71,734
i-000000000000001004111111111111p4d.24xlargeNVIDIA A10030$23,911
i-000000000000001024111111111111p4d.24xlargeNVIDIA A10030$23,911
i-00000000000000101b111111111111p3.8xlargeNVIDIA V10030$8,930
i-00000000000000101d111111111111p3.8xlargeNVIDIA V10030$8,930
i-000000000000001007111111111111g5.12xlargeNVIDIA A10G30$4,138
i-00000000000000101f111111111111g5.12xlargeNVIDIA A10G30$4,138
i-000000000000001023111111111111g5.12xlargeNVIDIA A10G30$4,138
i-000000000000001027111111111111g5.12xlargeNVIDIA A10G30$4,138
i-00000000000000101a111111111111p3.2xlargeNVIDIA V10030$2,233
i-00000000000000101c111111111111p3.2xlargeNVIDIA V10030$2,233
i-00000000000000101e111111111111g6e.xlargeNVIDIA L40S30$1,358
i-000000000000001022111111111111g6e.xlargeNVIDIA L40S30$1,358

9 untagged instances read as non-production and run continuously, $11,916/mo

$12k
observed per month
$8,369
recoverable per month
9
resources
2/5
disruption to change

Assumption.These resources carry no Environment tag. Each is classified from its account, its name, or its instance family, and the signals are listed per resource below. The finding holds if the client confirms the classification; strike any row that is wrong and the total falls by that row's cost.

What to do.Confirm the classification, tag the resources so this stops being an inference, then schedule them off outside working hours.

resourceaccountinstancegpudays always on$ / monthwhy it was classified this way
i-00000000000000100e111111111111g4dn.12xlargeyes30$2,854Name tag contains 'dev'
i-000000000000001011111111111111g4dn.12xlargeyes30$2,854Name tag contains 'dev'
i-000000000000001014222222222222g4dn.12xlargeyes30$2,854account name contains 'sandbox'
i-000000000000001010111111111111g5.xlargeyes30$734Name tag contains 'dev'; single-accelerator g5 instance, the shape of a development box
i-000000000000001013222222222222g5.xlargeyes30$734account name contains 'sandbox'; single-accelerator g5 instance, the shape of a development box
i-000000000000001016111111111111g5.xlargeyes30$734Name tag contains 'qa'; single-accelerator g5 instance, the shape of a development box
i-00000000000000100f111111111111g4dn.xlargeyes30$384Name tag contains 'dev'; single-accelerator g4dn instance, the shape of a development box
i-000000000000001012222222222222g4dn.xlargeyes30$384account name contains 'sandbox'; single-accelerator g4dn instance, the shape of a development box
i-000000000000001015111111111111g4dn.xlargeyes30$384Name tag contains 'qa'; single-accelerator g4dn instance, the shape of a development box

4 EBS volumes and snapshots billed for the full period with no associated compute, $320/mo

$320
observed per month
$320
recoverable per month
4
resources
1/5
disruption to change

Assumption.A CUR does not carry EBS attachment state, so this is inferred from a volume billing every day of the period while nothing in the account consumed compute hours against it. Confirm with `aws ec2 describe-volumes --filters Name=status,Values=available` before deleting anything. Snapshots that are the only copy of a dataset are not waste regardless of what the billing says.

What to do.Snapshot then delete volumes confirmed as available. For snapshots, apply a lifecycle policy rather than deleting by hand.

resourceaccountusage typeregiondays billed$ / month
arn:aws:ec2:us-east-1:111111111111:volume/vol-00000000000000000111111111111USE1-EBS:VolumeUsage.gp3us-east-130$80
arn:aws:ec2:us-east-1:111111111111:volume/vol-00000000000000001111111111111USE1-EBS:VolumeUsage.gp3us-east-130$80
arn:aws:ec2:us-east-1:111111111111:volume/vol-00000000000000002111111111111USE1-EBS:VolumeUsage.gp3us-east-130$80
arn:aws:ec2:us-east-1:111111111111:volume/vol-00000000000000003111111111111USE1-EBS:VolumeUsage.gp3us-east-130$80

UNPROVEN

$665k per month observed under these findings, all of it excluded from the recoverable total.

$22,326/mo on NVIDIA V100 (p3), superseded silicon

$22k
observed per month
$0
claimed as recoverable
4
resources
4/5
disruption to change

What is missing.Whether a newer family is cheaper for these workloads depends on model size, precision, memory footprint and batch size. None of that is in a Cost and Usage Report. Closing this gap needs one benchmark run of the real workload on a current-generation instance, measuring throughput per dollar rather than utilization. That is a day of work per distinct workload and it is the only thing that will settle it.

What to do.Benchmark the largest single workload on a current-generation instance before changing anything. Do not migrate a fleet on the strength of a spec sheet.

resourceaccountinstanceaccelerator$ / month
i-00000000000000101b111111111111p3.8xlargeNVIDIA V100$8,930
i-00000000000000101d111111111111p3.8xlargeNVIDIA V100$8,930
i-00000000000000101a111111111111p3.2xlargeNVIDIA V100$2,233
i-00000000000000101c111111111111p3.2xlargeNVIDIA V100$2,233

$482,302/mo of GPU spend has no utilization telemetry (100% of GPU spend)

$482k
observed per month
$0
claimed as recoverable
40
resources
1/5
disruption to change

What is missing.No per-GPU utilization data exists for these instances, so no statement can be made about whether the silicon did work. Closing this gap means running the NVIDIA DCGM exporter with profiling metrics enabled and scraping DCGM_FI_PROF_SM_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE into Prometheus. Note that DCGM_FI_DEV_GPU_UTIL, the metric on most default dashboards, does not close this gap: it reports whether a kernel was resident, not how much of the chip it used. On a T4 measured on 2026-09-04, a workload reporting 20% GPU utilization was using 0.2% of the streaming multiprocessors.

What to do.Deploy the DCGM exporter with container labelling enabled (--container-labels), which is off by default and is what makes per-pod attribution possible. Two weeks of data is enough to size the next engagement.

resourceaccountinstanceaccelerator$ / month
i-000000000000001001111111111111p5.48xlargeNVIDIA H100$71,734
i-000000000000001005111111111111p5.48xlargeNVIDIA H100$71,734
i-000000000000001025111111111111p5.48xlargeNVIDIA H100$51,649
i-000000000000001009111111111111p5.48xlargeNVIDIA H100$51,649
i-000000000000001021111111111111p5.48xlargeNVIDIA H100$44,475
i-000000000000001019111111111111p5.48xlargeNVIDIA H100$26,900
i-000000000000001004111111111111p4d.24xlargeNVIDIA A100$23,911
i-000000000000001024111111111111p4d.24xlargeNVIDIA A100$23,911
i-000000000000001000111111111111p4d.24xlargeNVIDIA A100$14,825
i-000000000000001008111111111111p4d.24xlargeNVIDIA A100$14,825
i-000000000000001020111111111111p4d.24xlargeNVIDIA A100$14,825
i-00000000000000101b111111111111p3.8xlargeNVIDIA V100$8,930
i-00000000000000101d111111111111p3.8xlargeNVIDIA V100$8,930
i-000000000000001018111111111111p4d.24xlargeNVIDIA A100$6,456
i-000000000000001023111111111111g5.12xlargeNVIDIA A10G$4,138
i-000000000000001007111111111111g5.12xlargeNVIDIA A10G$4,138
i-00000000000000101f111111111111g5.12xlargeNVIDIA A10G$4,138
i-000000000000001027111111111111g5.12xlargeNVIDIA A10G$4,138
i-000000000000001003111111111111g5.12xlargeNVIDIA A10G$2,980
i-000000000000001014222222222222g4dn.12xlargeNVIDIA T4$2,854
i-000000000000001011111111111111g4dn.12xlargeNVIDIA T4$2,854
i-00000000000000100e111111111111g4dn.12xlargeNVIDIA T4$2,854
i-00000000000000100b111111111111g4dn.12xlargeNVIDIA T4$2,854
i-00000000000000101a111111111111p3.2xlargeNVIDIA V100$2,233
i-00000000000000101c111111111111p3.2xlargeNVIDIA V100$2,233

Showing the 25 largest of 40 resources. The full list is in the machine-readable findings file that accompanies this report.

$160,357/mo of GPU capacity is claimed by a pod but has no utilization data

$160k
observed per month
$0
claimed as recoverable
12
resources
1/5
disruption to change

What is missing.A pod holding a GPU proves the accelerator is reserved, not that it is busy. Whether these GPUs are doing work needs DCGM_FI_PROF_SM_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE scraped per pod, which requires the DCGM exporter with --container-labels enabled. That flag is off by default, which is why these metrics are usually missing even in clusters that already run the exporter.

What to do.Deploy the DCGM exporter with container labelling, scrape for two weeks, then re-run this audit. Nothing should be resized before that data exists.

4 Cluster allocation

How many of the accelerators you are paying for has anything asked for. This needs no new instrumentation: if a node advertises eight GPUs and no pod requests one, utilization is zero by definition.

40 of 63 GPUs requested (63%)23 provisioned and unclaimed
Read from .status.capacity on each node and from the GPU requests of every non-terminated pod. 3 of 12 GPU nodes have no pod requesting an accelerator.
$101k
per month on GPUs nothing has requested
$160k
per month on GPUs a pod is holding
12
nodes matched to a line in the bill

Cost by namespace

platform$89kresearch$54kinference$17k
Each node's cost is split by the share of its GPUs a namespace requested. The unrequested share is not charged to anyone, which is the point of the figure above.

5 What this report cannot tell you

Each gap below is a question the billing data cannot answer, the dollars sitting behind it, and the specific measurement that would close it.

Question left openObserved per monthWhat would close it
$482,302/mo of GPU spend has no utilization telemetry (100% of GPU spend)$482kNo per-GPU utilization data exists for these instances, so no statement can be made about whether the silicon did work. Closing this gap means running the NVIDIA DCGM exporter with profiling metrics enabled and scraping DCGM_FI_PROF_SM_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE into Prometheus. Note that DCGM_FI_DEV_GPU_UTIL, the metric on most default dashboards, does not close this gap: it reports whether a kernel was resident, not how much of the chip it used. On a T4 measured on 2026-09-04, a workload reporting 20% GPU utilization was using 0.2% of the streaming multiprocessors.
$160,357/mo of GPU capacity is claimed by a pod but has no utilization data$160kA pod holding a GPU proves the accelerator is reserved, not that it is busy. Whether these GPUs are doing work needs DCGM_FI_PROF_SM_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE scraped per pod, which requires the DCGM exporter with --container-labels enabled. That flag is off by default, which is why these metrics are usually missing even in clusters that already run the exporter.
$22,326/mo on NVIDIA V100 (p3), superseded silicon$22kWhether a newer family is cheaper for these workloads depends on model size, precision, memory footprint and batch size. None of that is in a Cost and Usage Report. Closing this gap needs one benchmark run of the real workload on a current-generation instance, measuring throughput per dollar rather than utilization. That is a day of work per distinct workload and it is the only thing that will settle it.

None of the $665k per month above is included in the recoverable figure. Closing these gaps is the work of the next engagement, and the honest reason to do it is that the answer might be that nothing is wrong.

6 Order of work

Ordered by recoverable dollars per unit of disruption, not by dollars alone. The first item is the one that returns the most for the least argument.

gpu zero allocation$101kdisruption 2/5uncommitted steady gpu$47kdisruption 3/5always on nonprod inferred$8,369disruption 2/5always on nonprod tagged$4,562disruption 2/5unused commitment$1,976disruption 2/5orphaned storage$320disruption 1/5
Bar length is recoverable dollars per month. The disruption score at the right is the effort and risk of making the change, from 1 (delete an orphan) to 5 (re-architect a pipeline).
  1. 23 of 63 provisioned GPUs have no pod requesting them, $100,846/mo

    $101k per month, disruption 2 of 5. Scale the affected node groups down to the number of GPUs actually requested. Where a node group exists for burst capacity, set its floor to zero and let the cluster autoscaler bring it up on a pending pod.

  2. 14 production GPU instances run continuously at on-demand rates, $232,885/mo

    $47k per month, disruption 3 of 5. Size a Compute Savings Plan to the observed floor of GPU usage, not to the peak. Commit to the level that has been running every day of the period and leave the rest on demand.

  3. 9 untagged instances read as non-production and run continuously, $11,916/mo

    $8,369 per month, disruption 2 of 5. Confirm the classification, tag the resources so this stops being an inference, then schedule them off outside working hours.

  4. 10 tagged non-production instances run continuously, $6,495/mo

    $4,562 per month, disruption 2 of 5. Stop non-production instances outside working hours with an EventBridge schedule or an autoscaler floor of zero. GPU instances should be stopped rather than resized, because the accelerator is the entire cost.

  5. Commitment purchased and not consumed: $1,216/mo of Reserved Instance capacity and $760/mo of Savings Plan commitment

    $1,976 per month, disruption 2 of 5. Sell unused Standard RIs on the Reserved Instance Marketplace where the term allows, or move workloads onto the committed instance families so the commitment is consumed. Savings Plan commitments cannot be sold; the recovery is to shift eligible usage onto them before the term ends.

  6. 4 EBS volumes and snapshots billed for the full period with no associated compute, $320/mo

    $320 per month, disruption 1 of 5. Snapshot then delete volumes confirmed as available. For snapshots, apply a lifecycle policy rather than deleting by hand.

7 Method

Costs are amortized. On-demand line items use unblended cost; reserved-instance-covered usage uses reservation/EffectiveCost; Savings-Plan-covered usage uses savingsPlan/SavingsPlanEffectiveCost; unused commitment is read from the RIFee and SavingsPlanRecurringFee rows. Savings Plan negation rows are excluded so that covered usage is not counted twice.

Period figures are scaled to a 30.4-day month. Where a resource is identified as non-production without a tag saying so, the signals used are printed next to the resource and the finding is rated LIKELY, never PROVEN.