Inference cost audit

You self-host open models on vLLM or SGLang. GPU rental prices rose about a fifth in September, and both engines put out a release every two weeks that can change defaults you never set on purpose. The audit measures the traffic your deployment serves within your latency targets today, finds configuration changes that serve the same traffic on fewer GPUs, and leaves you the benchmark harness so you can rerun every number.

About me

I'm Rishabh Sinha, and LLM serving infrastructure is the work I do. My merged contributions span 14 open-source serving projects. Of those, kvcached holds the largest share, and I help triage issues there. I wrote servetune, the open-source harness the audit runs on.

Scope

A sample of your real traffic, or a synthetic profile matched to its length distribution, runs against your current setup and then against candidate changes one at a time. Candidates usually include FP8 weights and FP8 KV cache, prefill budget and batch limits, speculative decoding, KV offload, and replica sizing. Prefix caching gets special attention because gateways and chat templates often defeat it without anyone noticing.

Each candidate gets three numbers at your SLO: capacity, cost per million output tokens, and replicas needed at your peak. Before you roll out anything that can change model output, I mark it for a check on your own eval set. On a sample agent workload on one H100, one engine flag raised capacity within the same latency targets from 5.4 to 8.2 requests per second, and turning prefix caching off cut it to 2.8. See a sample report.

Deliverables

You receive a written report ranking changes by measured savings, the exact flags in rollout order, the harness with raw results, and a 60 minute readout call.

Process

After a short intake you share configs and metrics, plus either a traffic sample or its rate and length statistics. Benchmarks run on a staging replica you provide or on rented GPUs of the same type, billed at cost. Apart from the readout everything is asynchronous, and nothing needs production access.

Pricing

Config review

$2,500 one week

Ranked recommendations from your configs and traffic statistics, plus a list of what changed in the engine releases since your version, without benchmark runs.

Full audit

$9,500 two weeks

Adds measured results on your hardware type. If the measured saving at your SLO comes in under 15 percent, you pay the config review price instead.

Release re-tune

$3,500 per month

Each new vLLM or SGLang release gets checked against your configuration. Anything likely to move cost by more than a few percent gets re-benchmarked, and you receive a short note on what to change. Up to 8 hours of async help a month, and no on-call.

Limits

Production on-call is not part of any package, and there is no platform to buy. If a managed endpoint would be cheaper for your traffic, the report will say so.

Email me for a 20 minute call See the intake questions

Email: rsinha17@terpmail.umd.edu