Intake: inference cost audit

Answer what you can. Ranges are fine. Nothing here needs secrets or production access.

Models and engine

  1. Which models do you serve (name, size, precision, any fine-tunes)?
  2. Engine and version (vLLM, SGLang, other), and how you launch it (paste the command or Helm values).
  3. Any gateway or router in front (vLLM router, llm-d, LiteLLM, Envoy, a custom proxy)?

Hardware and spend

  1. GPU type and count per replica, number of replicas, cloud or on-prem.
  2. What you pay per GPU-hour, or your monthly inference bill.
  3. Autoscaling: on or off, and what it scales on.

Traffic

  1. Peak and average requests per second.
  2. Typical input and output lengths (median and p95 if you have them).
  3. How much of each prompt is shared across requests (system prompt, tool schemas, RAG preamble), roughly in tokens.
  4. Streaming or not, and whether requests come from an OpenAI-style, Anthropic-style or custom API.
  5. Can you share a sample of anonymized request metadata (timestamps, input and output token counts, no content)?

Targets

  1. Latency targets: time to first token and time per output token, at p50 and p95.
  2. Quality guardrails: which eval set, and whether quantized weights or KV cache are acceptable in principle.

Logistics

  1. Can you provide a staging replica of the same GPU type, or should benchmarks run on rented GPUs at cost?
  2. Who reads the report, and what decision does it feed (budget, capacity plan, migration)?

Send your answers to rsinha17@terpmail.umd.edu. Ranges and rough numbers are fine.

Back to the inference cost audit