Intake: inference cost audit
Answer what you can. Ranges are fine. Nothing here needs secrets or production access.
Models and engine
- Which models do you serve (name, size, precision, any fine-tunes)?
- Engine and version (vLLM, SGLang, other), and how you launch it (paste the command or Helm values).
- Any gateway or router in front (vLLM router, llm-d, LiteLLM, Envoy, a custom proxy)?
Hardware and spend
- GPU type and count per replica, number of replicas, cloud or on-prem.
- What you pay per GPU-hour, or your monthly inference bill.
- Autoscaling: on or off, and what it scales on.
Traffic
- Peak and average requests per second.
- Typical input and output lengths (median and p95 if you have them).
- How much of each prompt is shared across requests (system prompt, tool schemas, RAG preamble), roughly in tokens.
- Streaming or not, and whether requests come from an OpenAI-style, Anthropic-style or custom API.
- Can you share a sample of anonymized request metadata (timestamps, input and output token counts, no content)?
Targets
- Latency targets: time to first token and time per output token, at p50 and p95.
- Quality guardrails: which eval set, and whether quantized weights or KV cache are acceptable in principle.
Logistics
- Can you provide a staging replica of the same GPU type, or should benchmarks run on rented GPUs at cost?
- Who reads the report, and what decision does it feed (budget, capacity plan, migration)?
Send your answers to rsinha17@terpmail.umd.edu. Ranges and rough numbers are fine.