Files
hermes-skills/mlops/llm-inference-benchmarking/references/aiperf-capacity-curve-analysis.md
T

99 lines
3.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Capacity Curve Analysis When to Stop Scaling Concurrency
## The Core Question
"Shall we test conc=200?" The answer depends on extrapolating the existing capacity curve, not on gut feeling.
## Extrapolation Rule
Given TTFT p99 values at conc=10, 50, 100, estimate conc=200 using a **linear degradation model** (the simplest that fits observed data):
```
TTFT_p99(conc) ≈ base + k × conc
Where k is derived from the steepest interval:
k = (TTFT_p99(conc=100) - TTFT_p99(conc=50)) / (100 - 50)
```
If the extrapolated TTFT p99 at conc=200 exceeds a **user-acceptable threshold** (typically 5000ms = 5s for interactive applications, 10000ms = 10s for batch), conc=200 is **operationally irrelevant**.
## Worked Example (from this session)
### gpt-oss-120b
| Conc | TTFT p99 |
|------|----------|
| 10 | 978 ms |
| 25 | 1924 ms |
| 50 | 3575 ms |
| 100 | 7310 ms |
Extrapolation to conc=200:
- k = (7310 - 3575) / 50 = **74.7 ms/conc**
- TTFT_p99(200) = 7310 + 74.7 × 100 = **14,780 ms**
**Verdict: Unusable. Users won't wait 15 seconds for the first token.**
### gemma-4-31b-it
| Conc | TTFT p99 |
|------|----------|
| 10 | 381 ms |
| 50 | 1042 ms |
| 100 | 2060 ms |
Extrapolation:
- k = (2060 - 1042) / 50 = **20.4 ms/conc**
- TTFT_p99(200) = 2060 + 20.4 × 100 = **4,100 ms**
**Verdict: Borderline. 4.1s is acceptable for some batch use cases, unacceptable for chat.**
### moonshotai-kimi-k2.6
| Conc | TTFT p99 |
|------|----------|
| 10 | 695 ms |
| 25 | 1178 ms |
| 50 | 1989 ms |
| 100 | 3926 ms |
Extrapolation:
- k = (3926 - 1989) / 50 = **38.7 ms/conc**
- TTFT_p99(200) = 3926 + 38.7 × 100 = **7,796 ms**
**Verdict: Unusable for interactive use.**
## Alternative to Higher Concurrency: Longer Inputs
When conc=100 is already the practical ceiling, the more valuable test is **long-context prefill stress**:
| Test | What it measures | When to use |
|------|-----------------|-------------|
| conc=10→100 | Queueing / scheduling efficiency | Find max parallel users |
| ISL=512 → 1024 → 2048 → 4096 | Memory bandwidth / KV-cache limits | RAG, document analysis, long-context agents |
A model that degrades gracefully at conc=100 but collapses at ISL=2048 has a **memory-bandwidth bottleneck**, not a scheduling bottleneck. The fix is faster memory (HBM3e), not more GPU compute.
## Breakpoint Detection Heuristic
Formally define the "breakpoint" as the highest concurrency where **all three** conditions hold:
1. TTFT p99 < 5000 ms (5 seconds)
2. ITL p99 < 200 ms (readable streaming, ~5 tok/s per user)
3. Output throughput still scales linearly with concurrency (no saturation)
From observed data:
| Model | Breakpoint (conc) | Limiting factor |
|-------|-------------------|-----------------|
| gemma-4-31b-it | ~100 | TTFT ≈ 2s still OK |
| moonshotai-kimi-k2.6 | ~75 | TTFT ≈ 3.5s approaching limit |
| gpt-oss-120b | ~40 | TTFT ≈ 3.5s at conc=50 |
| qwen3.6-27b-nvfp4 | ~40 | TTFT ≈ 3s at conc=50 |
## Operational Recommendation
- Set **hard concurrency limits** at the breakpoint, not at the hardware maximum
- Implement **queueing + load balancing** beyond the breakpoint
- Test **LLM-as-a-Service tiers**:
- Tier 1 (Premium): conc ≤ 10, guaranteed < 1s TTFT
- Tier 2 (Standard): conc ≤ 50, guaranteed < 3s TTFT
- Tier 3 (Batch): conc ≤ 100, best-effort, queueing OK