The person who scopes the work is the one who runs the instrument.
Nameplate Analytics is a specialist practice, not a staffing pyramid. There is no junior consultant assigned after the sale and no template report with your logo dropped into the header. You are buying the judgment of the engineer who does the measuring.
Kam Kheri
Seven years building and running cloud infrastructure at federal scale. Currently supporting production cloud infrastructure for a national public research program across AWS, GCP and Kubernetes: Terraform, Ansible, Puppet, and the monitoring stacks this practice depends on. Prometheus, Grafana, the ELK and Splunk side, and the DCGM exporter that produced the measurement on this site.
Alongside that, technical lead on a nine-engineer team building a personnel and security onboarding platform from scratch to replace a commercial Workday deployment. C#/.NET 8, Next.js, TypeScript, multi-schema PostgreSQL. That work is the reason the tooling behind this practice is engineered rather than scripted: gpuaudit has a schema layer, an amortization layer, a cluster allocation module and a regression suite, because that is how you build something whose output a client will act on.
Federal and government clients throughout. Active Top Secret clearance, which is worth mentioning here for one specific reason: you are being asked to grant read access to your billing account and your cluster, and it is fair to want to know who is on the other end of that.
Generic cloud cost work is finished. This is the part nobody instrumented.
Rightsizing, idle instances, commitment coverage. That work is commoditized, several tools do it for free, and the practitioners surveyed in the 2026 State of FinOps report say the obvious waste is already gone. A firm selling you that today is selling you a report you could generate yourself.
The question almost nobody can answer is what fraction of GPU spend produced work, because answering it requires joining billing data to hardware telemetry that most organizations have never turned on. Granular AI and GPU monitoring is the single most-requested unmet tooling capability in that same survey, and 98 percent of FinOps teams now manage AI spend, up from 31 percent two years earlier.
So the practice is deliberately narrow: GPU cost attribution for teams running machine learning on Kubernetes. We say no to work outside that, including work we could technically do, because a specialist who takes generalist engagements stops being a specialist.
Four commitments, and we will hold to them against our own interest.
Fixed fee, never contingency
A percentage of savings gives the person measuring your infrastructure a stake in the size of the number. We will not take that structure even when a client offers it.
Confidence on every finding
Proven, likely, or unproven. Unproven dollars stay out of the recoverable total, and the exclusion is enforced in the tooling because under pressure discipline is what bends.
We will tell you when there is nothing
If the snapshot comes back clean we say so on the call and stop. A bad engagement costs more than no engagement, for both sides.
The measurement is reproducible
The Terraform, the workload scripts and the raw 242-sample log are public on GitHub. Rebuild the instance, run the sampler, and check the number against ours. Nothing on this site rests on a figure you have to take on faith.
Nothing is touched without a change window
Reports touch nothing at all. Sprints run in your repositories, through your review process, with your engineer as the merge authority.
Scope stays where it was written
The SOW names the specific items agreed at the end of the previous tier. Anything outside it gets repriced in writing rather than absorbed and resented.
You are granting a stranger read access to your billing account. Here is exactly what that means.
What we ask for
- Read-only, always. Billing exports and cluster reads. No write permissions at any tier that produces a document.
- Scoped and time-limited credentials, revoked at the end of the engagement and confirmed in writing.
- Federated short-lived credentials where your identity provider supports it, rather than long-lived access keys.
- A mutual NDA signed before any access is granted, not after.
What we never do
- Install agents on your production nodes.
- Touch production traffic routing, at any tier.
- Subcontract, offshore, or grant a third party access to your environment.
- Retain your data after close-out. The deliverable is yours; the working copies are destroyed.
The full security and data handling document, covering credential storage, incident notification and subprocessor policy, goes out with the mutual NDA. Ask for it before the first call if it needs to clear your security team. It is written to be read by one, not to be waved at one.
Start with the free snapshot, or just ask a question.
Forty-five minutes and read-only access to one month of billing data is enough to know whether there is anything here. If there is not, we will say so on the call.
- Direct line
- 301-802-1073
- Kam Kheri
- Registered office
- 7901 4th St N, Ste 300
St. Petersburg, FL 33702 - Practice location
- Miami, Florida
- Response time
- Same business day