Every finding in this report carries a confidence rating. Only findings the data supports are added to the recoverable total. The rest are printed in full and left out of the number, because a figure that does not survive a skeptical engineer is worth less than no figure at all.
$163k
recoverable per month, proven and likely findings only
$1.95M
recoverable per year at the same run rate
$665k
observed but not claimed, excluded from the figures to the left
The hatched segment is money we are not claiming. It is shown at the same scale as the rest so the proportion is visible rather than described.
Total amortized spend across the period was $485k, of which $476k was GPU compute. Amortized cost is used throughout, not unblended cost: AWS writes an unblended cost of zero on reserved-instance-covered line items, so a ranking by unblended cost makes the most expensive machine in an account disappear.
2 How to read a finding
PROVEN
The billing data proves both the cost and the waste. No assumption about how the workload behaves is involved.
LIKELY
The billing data proves the cost. The waste follows from one stated assumption, printed with the finding, which you can accept or reject line by line.
UNPROVEN
There is a cost and a suspicion. The data that would settle it does not exist yet. These dollars are listed and then excluded from the recoverable total.
3 Findings
PROVEN
$107k per month recoverable across 3 findings.
23 of 63 provisioned GPUs have no pod requesting them, $100,846/mo
$101k
observed per month
$101k
recoverable per month
3
resources
2/5
disruption to change
What to do.Scale the affected node groups down to the number of GPUs actually requested. Where a node group exists for burst capacity, set its floor to zero and let the cluster autoscaler bring it up on a pending pod.
node
instance
instance
GPUs
GPUs requested
$ / month
ip-10-0-001-10.ec2.internal
i-000000000000001001
p5.48xlarge
8
0
$71,734
ip-10-0-000-10.ec2.internal
i-000000000000001000
p4d.24xlarge
8
0
$14,825
ip-10-0-002-10.ec2.internal
i-000000000000001002
g6e.xlarge
1
0
$842
10 tagged non-production instances run continuously, $6,495/mo
$6,495
observed per month
$4,562
recoverable per month
10
resources
2/5
disruption to change
What to do.Stop non-production instances outside working hours with an EventBridge schedule or an autoscaler floor of zero. GPU instances should be stopped rather than resized, because the accelerator is the entire cost.
resource
account
instance
gpu
days always on
$ / month
why it was classified this way
i-00000000000000100b
111111111111
g4dn.12xlarge
yes
30
$2,854
Environment tag = dev
i-00000000000000100a
111111111111
g5.xlarge
yes
30
$734
Environment tag = dev
i-00000000000000100d
111111111111
g5.xlarge
yes
30
$734
Environment tag = dev
i-000000000000009007
111111111111
c5.9xlarge
no
30
$692
Environment tag = dev
i-00000000000000100c
111111111111
g4dn.xlarge
yes
30
$384
Environment tag = dev
i-000000000000009001
111111111111
m5.4xlarge
no
30
$347
Environment tag = dev
i-000000000000009010
111111111111
m5.4xlarge
no
30
$347
Environment tag = dev
i-00000000000000900d
111111111111
r5.2xlarge
no
30
$228
Environment tag = dev
i-000000000000009004
111111111111
m6i.xlarge
no
30
$87
Environment tag = dev
i-000000000000009013
111111111111
m6i.xlarge
no
30
$87
Environment tag = dev
Commitment purchased and not consumed: $1,216/mo of Reserved Instance capacity and $760/mo of Savings Plan commitment
$1,976
observed per month
$1,976
recoverable per month
2
resources
2/5
disruption to change
What to do.Sell unused Standard RIs on the Reserved Instance Marketplace where the term allows, or move workloads onto the committed instance families so the commitment is consumed. Savings Plan commitments cannot be sold; the recovery is to shift eligible usage onto them before the term ends.
resource
account
commitment
usage type
unused $ / month
arn:aws:ec2:us-east-1:ri/r-abc
111111111111
RIFee
RIFee
$1,216
arn:aws:savingsplans::sp/sp-xyz
111111111111
SavingsPlanRecurringFee
RIFee
$760
LIKELY
$55k per month recoverable across 3 findings.
14 production GPU instances run continuously at on-demand rates, $232,885/mo
$233k
observed per month
$47k
recoverable per month
14
resources
3/5
disruption to change
Assumption.The recoverable figure assumes a 20% rate reduction from a one-year no-upfront Compute Savings Plan, and assumes this capacity is still wanted twelve months from now. Verify the current discount for these families against AWS pricing before committing; if the workload is being migrated or retired inside the term, this finding does not apply.
What to do.Size a Compute Savings Plan to the observed floor of GPU usage, not to the peak. Commit to the level that has been running every day of the period and leave the rest on demand.
resource
account
instance
accelerator
days always on
$ / month
i-000000000000001001
111111111111
p5.48xlarge
NVIDIA H100
30
$71,734
i-000000000000001005
111111111111
p5.48xlarge
NVIDIA H100
30
$71,734
i-000000000000001004
111111111111
p4d.24xlarge
NVIDIA A100
30
$23,911
i-000000000000001024
111111111111
p4d.24xlarge
NVIDIA A100
30
$23,911
i-00000000000000101b
111111111111
p3.8xlarge
NVIDIA V100
30
$8,930
i-00000000000000101d
111111111111
p3.8xlarge
NVIDIA V100
30
$8,930
i-000000000000001007
111111111111
g5.12xlarge
NVIDIA A10G
30
$4,138
i-00000000000000101f
111111111111
g5.12xlarge
NVIDIA A10G
30
$4,138
i-000000000000001023
111111111111
g5.12xlarge
NVIDIA A10G
30
$4,138
i-000000000000001027
111111111111
g5.12xlarge
NVIDIA A10G
30
$4,138
i-00000000000000101a
111111111111
p3.2xlarge
NVIDIA V100
30
$2,233
i-00000000000000101c
111111111111
p3.2xlarge
NVIDIA V100
30
$2,233
i-00000000000000101e
111111111111
g6e.xlarge
NVIDIA L40S
30
$1,358
i-000000000000001022
111111111111
g6e.xlarge
NVIDIA L40S
30
$1,358
9 untagged instances read as non-production and run continuously, $11,916/mo
$12k
observed per month
$8,369
recoverable per month
9
resources
2/5
disruption to change
Assumption.These resources carry no Environment tag. Each is classified from its account, its name, or its instance family, and the signals are listed per resource below. The finding holds if the client confirms the classification; strike any row that is wrong and the total falls by that row's cost.
What to do.Confirm the classification, tag the resources so this stops being an inference, then schedule them off outside working hours.
resource
account
instance
gpu
days always on
$ / month
why it was classified this way
i-00000000000000100e
111111111111
g4dn.12xlarge
yes
30
$2,854
Name tag contains 'dev'
i-000000000000001011
111111111111
g4dn.12xlarge
yes
30
$2,854
Name tag contains 'dev'
i-000000000000001014
222222222222
g4dn.12xlarge
yes
30
$2,854
account name contains 'sandbox'
i-000000000000001010
111111111111
g5.xlarge
yes
30
$734
Name tag contains 'dev'; single-accelerator g5 instance, the shape of a development box
i-000000000000001013
222222222222
g5.xlarge
yes
30
$734
account name contains 'sandbox'; single-accelerator g5 instance, the shape of a development box
i-000000000000001016
111111111111
g5.xlarge
yes
30
$734
Name tag contains 'qa'; single-accelerator g5 instance, the shape of a development box
i-00000000000000100f
111111111111
g4dn.xlarge
yes
30
$384
Name tag contains 'dev'; single-accelerator g4dn instance, the shape of a development box
i-000000000000001012
222222222222
g4dn.xlarge
yes
30
$384
account name contains 'sandbox'; single-accelerator g4dn instance, the shape of a development box
i-000000000000001015
111111111111
g4dn.xlarge
yes
30
$384
Name tag contains 'qa'; single-accelerator g4dn instance, the shape of a development box
4 EBS volumes and snapshots billed for the full period with no associated compute, $320/mo
$320
observed per month
$320
recoverable per month
4
resources
1/5
disruption to change
Assumption.A CUR does not carry EBS attachment state, so this is inferred from a volume billing every day of the period while nothing in the account consumed compute hours against it. Confirm with `aws ec2 describe-volumes --filters Name=status,Values=available` before deleting anything. Snapshots that are the only copy of a dataset are not waste regardless of what the billing says.
What to do.Snapshot then delete volumes confirmed as available. For snapshots, apply a lifecycle policy rather than deleting by hand.
$665k per month observed under these findings, all of it excluded from the recoverable total.
$22,326/mo on NVIDIA V100 (p3), superseded silicon
$22k
observed per month
$0
claimed as recoverable
4
resources
4/5
disruption to change
What is missing.Whether a newer family is cheaper for these workloads depends on model size, precision, memory footprint and batch size. None of that is in a Cost and Usage Report. Closing this gap needs one benchmark run of the real workload on a current-generation instance, measuring throughput per dollar rather than utilization. That is a day of work per distinct workload and it is the only thing that will settle it.
What to do.Benchmark the largest single workload on a current-generation instance before changing anything. Do not migrate a fleet on the strength of a spec sheet.
resource
account
instance
accelerator
$ / month
i-00000000000000101b
111111111111
p3.8xlarge
NVIDIA V100
$8,930
i-00000000000000101d
111111111111
p3.8xlarge
NVIDIA V100
$8,930
i-00000000000000101a
111111111111
p3.2xlarge
NVIDIA V100
$2,233
i-00000000000000101c
111111111111
p3.2xlarge
NVIDIA V100
$2,233
$482,302/mo of GPU spend has no utilization telemetry (100% of GPU spend)
$482k
observed per month
$0
claimed as recoverable
40
resources
1/5
disruption to change
What is missing.No per-GPU utilization data exists for these instances, so no statement can be made about whether the silicon did work. Closing this gap means running the NVIDIA DCGM exporter with profiling metrics enabled and scraping DCGM_FI_PROF_SM_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE into Prometheus. Note that DCGM_FI_DEV_GPU_UTIL, the metric on most default dashboards, does not close this gap: it reports whether a kernel was resident, not how much of the chip it used. On a T4 measured on 2026-09-04, a workload reporting 20% GPU utilization was using 0.2% of the streaming multiprocessors.
What to do.Deploy the DCGM exporter with container labelling enabled (--container-labels), which is off by default and is what makes per-pod attribution possible. Two weeks of data is enough to size the next engagement.
resource
account
instance
accelerator
$ / month
i-000000000000001001
111111111111
p5.48xlarge
NVIDIA H100
$71,734
i-000000000000001005
111111111111
p5.48xlarge
NVIDIA H100
$71,734
i-000000000000001025
111111111111
p5.48xlarge
NVIDIA H100
$51,649
i-000000000000001009
111111111111
p5.48xlarge
NVIDIA H100
$51,649
i-000000000000001021
111111111111
p5.48xlarge
NVIDIA H100
$44,475
i-000000000000001019
111111111111
p5.48xlarge
NVIDIA H100
$26,900
i-000000000000001004
111111111111
p4d.24xlarge
NVIDIA A100
$23,911
i-000000000000001024
111111111111
p4d.24xlarge
NVIDIA A100
$23,911
i-000000000000001000
111111111111
p4d.24xlarge
NVIDIA A100
$14,825
i-000000000000001008
111111111111
p4d.24xlarge
NVIDIA A100
$14,825
i-000000000000001020
111111111111
p4d.24xlarge
NVIDIA A100
$14,825
i-00000000000000101b
111111111111
p3.8xlarge
NVIDIA V100
$8,930
i-00000000000000101d
111111111111
p3.8xlarge
NVIDIA V100
$8,930
i-000000000000001018
111111111111
p4d.24xlarge
NVIDIA A100
$6,456
i-000000000000001023
111111111111
g5.12xlarge
NVIDIA A10G
$4,138
i-000000000000001007
111111111111
g5.12xlarge
NVIDIA A10G
$4,138
i-00000000000000101f
111111111111
g5.12xlarge
NVIDIA A10G
$4,138
i-000000000000001027
111111111111
g5.12xlarge
NVIDIA A10G
$4,138
i-000000000000001003
111111111111
g5.12xlarge
NVIDIA A10G
$2,980
i-000000000000001014
222222222222
g4dn.12xlarge
NVIDIA T4
$2,854
i-000000000000001011
111111111111
g4dn.12xlarge
NVIDIA T4
$2,854
i-00000000000000100e
111111111111
g4dn.12xlarge
NVIDIA T4
$2,854
i-00000000000000100b
111111111111
g4dn.12xlarge
NVIDIA T4
$2,854
i-00000000000000101a
111111111111
p3.2xlarge
NVIDIA V100
$2,233
i-00000000000000101c
111111111111
p3.2xlarge
NVIDIA V100
$2,233
Showing the 25 largest of 40 resources. The full list is in the machine-readable findings file that accompanies this report.
$160,357/mo of GPU capacity is claimed by a pod but has no utilization data
$160k
observed per month
$0
claimed as recoverable
12
resources
1/5
disruption to change
What is missing.A pod holding a GPU proves the accelerator is reserved, not that it is busy. Whether these GPUs are doing work needs DCGM_FI_PROF_SM_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE scraped per pod, which requires the DCGM exporter with --container-labels enabled. That flag is off by default, which is why these metrics are usually missing even in clusters that already run the exporter.
What to do.Deploy the DCGM exporter with container labelling, scrape for two weeks, then re-run this audit. Nothing should be resized before that data exists.
4 Cluster allocation
How many of the accelerators you are paying for has anything asked for. This needs no new instrumentation: if a node advertises eight GPUs and no pod requests one, utilization is zero by definition.
Read from .status.capacity on each node and from the GPU requests of every non-terminated pod. 3 of 12 GPU nodes have no pod requesting an accelerator.
$101k
per month on GPUs nothing has requested
$160k
per month on GPUs a pod is holding
12
nodes matched to a line in the bill
Cost by namespace
Each node's cost is split by the share of its GPUs a namespace requested. The unrequested share is not charged to anyone, which is the point of the figure above.
5 What this report cannot tell you
Each gap below is a question the billing data cannot answer, the dollars sitting behind it, and the specific measurement that would close it.
Question left open
Observed per month
What would close it
$482,302/mo of GPU spend has no utilization telemetry (100% of GPU spend)
$482k
No per-GPU utilization data exists for these instances, so no statement can be made about whether the silicon did work. Closing this gap means running the NVIDIA DCGM exporter with profiling metrics enabled and scraping DCGM_FI_PROF_SM_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE into Prometheus. Note that DCGM_FI_DEV_GPU_UTIL, the metric on most default dashboards, does not close this gap: it reports whether a kernel was resident, not how much of the chip it used. On a T4 measured on 2026-09-04, a workload reporting 20% GPU utilization was using 0.2% of the streaming multiprocessors.
$160,357/mo of GPU capacity is claimed by a pod but has no utilization data
$160k
A pod holding a GPU proves the accelerator is reserved, not that it is busy. Whether these GPUs are doing work needs DCGM_FI_PROF_SM_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE scraped per pod, which requires the DCGM exporter with --container-labels enabled. That flag is off by default, which is why these metrics are usually missing even in clusters that already run the exporter.
$22,326/mo on NVIDIA V100 (p3), superseded silicon
$22k
Whether a newer family is cheaper for these workloads depends on model size, precision, memory footprint and batch size. None of that is in a Cost and Usage Report. Closing this gap needs one benchmark run of the real workload on a current-generation instance, measuring throughput per dollar rather than utilization. That is a day of work per distinct workload and it is the only thing that will settle it.
None of the $665k per month above is included in the recoverable figure. Closing these gaps is the work of the next engagement, and the honest reason to do it is that the answer might be that nothing is wrong.
6 Order of work
Ordered by recoverable dollars per unit of disruption, not by dollars alone. The first item is the one that returns the most for the least argument.
Bar length is recoverable dollars per month. The disruption score at the right is the effort and risk of making the change, from 1 (delete an orphan) to 5 (re-architect a pipeline).
23 of 63 provisioned GPUs have no pod requesting them, $100,846/mo
$101k per month, disruption 2 of 5. Scale the affected node groups down to the number of GPUs actually requested. Where a node group exists for burst capacity, set its floor to zero and let the cluster autoscaler bring it up on a pending pod.
14 production GPU instances run continuously at on-demand rates, $232,885/mo
$47k per month, disruption 3 of 5. Size a Compute Savings Plan to the observed floor of GPU usage, not to the peak. Commit to the level that has been running every day of the period and leave the rest on demand.
9 untagged instances read as non-production and run continuously, $11,916/mo
$8,369 per month, disruption 2 of 5. Confirm the classification, tag the resources so this stops being an inference, then schedule them off outside working hours.
10 tagged non-production instances run continuously, $6,495/mo
$4,562 per month, disruption 2 of 5. Stop non-production instances outside working hours with an EventBridge schedule or an autoscaler floor of zero. GPU instances should be stopped rather than resized, because the accelerator is the entire cost.
Commitment purchased and not consumed: $1,216/mo of Reserved Instance capacity and $760/mo of Savings Plan commitment
$1,976 per month, disruption 2 of 5. Sell unused Standard RIs on the Reserved Instance Marketplace where the term allows, or move workloads onto the committed instance families so the commitment is consumed. Savings Plan commitments cannot be sold; the recovery is to shift eligible usage onto them before the term ends.
4 EBS volumes and snapshots billed for the full period with no associated compute, $320/mo
$320 per month, disruption 1 of 5. Snapshot then delete volumes confirmed as available. For snapshots, apply a lifecycle policy rather than deleting by hand.
7 Method
Costs are amortized. On-demand line items use unblended cost; reserved-instance-covered usage uses reservation/EffectiveCost; Savings-Plan-covered usage uses savingsPlan/SavingsPlanEffectiveCost; unused commitment is read from the RIFee and SavingsPlanRecurringFee rows. Savings Plan negation rows are excluded so that covered usage is not counted twice.
Period figures are scaled to a 30.4-day month. Where a resource is identified as non-production without a tag saying so, the signals used are printed next to the resource and the finding is rated LIKELY, never PROVEN.
30 days of data; period figures scaled to a 30.4-day month by a factor of 1.013
3 parquet file(s), cur2 schema, duckdb backend
Cluster allocation data was supplied, so zero-allocation GPU cost is reported as PROVEN rather than inferred from billing alone.