Back to the inference cost audit ยท servetune on GitHub

Sample inference cost audit: Qwen3-8B, agent-style traffic, 1x H100

Headline. Of the tested configurations, fp8_model_kv_16k serves 8.40 requests/s within the SLO against 5.36 for the current setup, which cuts cost per million output tokens from $0.45 to $0.29 (-36%). At the stated peak of 40 requests/s that is 5 replicas instead of 8, about $5,891 a month.

Setup

Results, cheapest first

p95 time to first token vs delivered throughputTTFT SLO 1500 msdelivered requests / sp95 TTFT (ms)0.004.01,5757.93,15011.94,72515.86,300baselinebatched_16kfp8_kvfp8_modelfp8_model_kv_16kno_prefix_cache
ConfigCapacity at SLO (req/s)vs currentOutput tok/s at capacity$ per 1M output tokvs currentp95 TTFT / TPOT at capacity (ms)
fp8_model_kv_16k8.40+57%2,583$0.289-36%140 / 13.1
fp8_kv8.24+54%2,533$0.295-35%125 / 13.8
baseline5.36+0%1,649$0.453+0%564 / 43.3
batched_16k5.35-0%1,644$0.455+0%579 / 44.0
no_prefix_cache2.84-47%873$0.856+89%436 / 23.0
fp8_modelnone passedn/an/an/an/asee notes

What each config changes

Method and limits

servetune reports capacity as the highest tested request rate that met both SLOs with every request completed. True capacity lies between that rate and the next one tested. Synthetic prompts come from the engine's own random dataset with the stated length ranges and shared prefix. In a client audit, servetune replays the client's own traffic instead. Quality effects of quantization are outside this run's scope, so check any FP8 change on your own eval set before rollout.