110 lines
4.4 KiB
Markdown
110 lines
4.4 KiB
Markdown
# Prefill Throughput Bias from Cold-Start Requests
|
||
|
||
## Problem
|
||
|
||
AIPerf's per-run `profile_export_aiperf.json` aggregates ALL requests, including the very first one which has to establish a fresh HTTP connection. This biases the `prefill_throughput_per_user` metric significantly:
|
||
|
||
| Request | Input Tokens | Prefill tok/s | TTFT | Connection |
|
||
|---------|-------------|---------------|------|------------|
|
||
| 0 | 7 | **4 tok/s** | 1699 ms | DNS (1002ms) + TCP (1144ms) |
|
||
| 1 | 371 | 2003 tok/s | 185 ms | Reused ✓ |
|
||
| 2 | 676 | 3596 tok/s | 188 ms | Reused ✓ |
|
||
| 5 | 1488 | 7571 tok/s | 197 ms | Reused ✓ |
|
||
|
||
The first request's DNS lookup + TCP handshake (2+ seconds) is counted toward "prefill time", deflating the metric. The **true** prefill throughput (requests 1+) is ~2K–7.5K tok/s, not the ~4K aggregate.
|
||
|
||
## Impact
|
||
|
||
- **Prefill metrics**: Inflated by ~1.5–2× for short runs (<20 requests). Improves as sample size grows.
|
||
- **TTFT p99**: May be driven by the warm-up request's elevated TTFT.
|
||
- **Low-concurrency runs** (conc=1, single session) are most affected because only ONE session exists and its first request is always cold.
|
||
|
||
## Mitigation
|
||
|
||
### Option A: Discard First Request (Recommended for Prescriptive Benchmarks)
|
||
|
||
Parse `profile_export.jsonl` (line-delimited JSON, one request per line) and exclude `session_num=0` / `turn_index=0`:
|
||
|
||
```python
|
||
import json
|
||
|
||
metrics = []
|
||
with open("profile_export.jsonl") as f:
|
||
for line in f:
|
||
req = json.loads(line)
|
||
# Skip first turn of first session (cold-start)
|
||
if req["metadata"]["turn_index"] == 0 and req["metadata"]["session_num"] == 0:
|
||
continue
|
||
m = req["metrics"]
|
||
metrics.append({
|
||
"ttft": m["time_to_first_token"]["value"],
|
||
"prefill": m["prefill_throughput_per_user"]["value"],
|
||
})
|
||
|
||
# Now metrics[] excludes the warm-up request
|
||
```
|
||
|
||
### Option B: Use Connection Warm-Up Before Benchmarking
|
||
|
||
Send 1–2 throwaway requests to establish the TCP + TLS + HTTP/2 connection pool before the measured phase. AIPerf does not have a built-in warm-up phase, so this must be done externally:
|
||
|
||
```bash
|
||
# Pre-warm the connection
|
||
venv/bin/aiperf profile --config warmup.yaml \
|
||
--concurrency 1 --request-count 1 \
|
||
--artifact-dir /tmp/warmup
|
||
|
||
# Now run the real benchmark — connections are hot
|
||
venv/bin/aiperf profile --config benchmark.yaml \
|
||
--concurrency 1 --request-count 50 \
|
||
--artifact-dir results/benchmark
|
||
```
|
||
|
||
The warmup run should use the **same endpoint, model, and streaming=true** as the benchmark to ensure the connection pool is pre-established.
|
||
|
||
### Option C: Accept Bias for Comparative Benchmarks
|
||
|
||
If comparing **N models on the same infra with the same settings**, the cold-start bias applies equally to all. The absolute numbers are slightly off, but **rankings and relative ratios remain valid**. Use this only when:
|
||
- Absolute SLA numbers are not the goal
|
||
- All runs use identical connection topology (same client → same endpoint)
|
||
- Sample size is large enough that one cold request is <5% of total
|
||
|
||
## When It Matters Most
|
||
|
||
| Scenario | Severity | Why |
|
||
|----------|----------|-----|
|
||
| conc=1, sessions=1 | 🔴 High | First request is 100% of traffic |
|
||
| conc=5, sessions=5 | 🟡 Medium | 1 of 5 sessions starts cold |
|
||
| conc=50, sessions=50 | 🟢 Low | Many sessions, cold-start amortized |
|
||
| Capacity (10→100+) | 🟢 Low | Many sessions, cold-start amortized |
|
||
| Steady-state (10min) | 🟢 Low | Thousands of requests, negligible |
|
||
|
||
## Per-Request Data Location
|
||
|
||
Raw per-request data is in `profile_export.jsonl` (NOT `profile_export_aiperf.json` which is aggregated). Each line is a JSON object with:
|
||
|
||
```json
|
||
{
|
||
"metadata": {
|
||
"session_num": 0,
|
||
"turn_index": 0,
|
||
"worker_id": "worker_766ae0c2",
|
||
"conversation_id": "...",
|
||
"was_cancelled": false
|
||
},
|
||
"metrics": {
|
||
"time_to_first_token": {"value": 1699.2, "unit": "ms"},
|
||
"prefill_throughput_per_user": {"value": 4.1, "unit": "tokens/sec/user"},
|
||
"request_latency": {"value": 26372.4, "unit": "ms"},
|
||
"http_req_connection_reused": {"value": 0, "unit": "ratio"}
|
||
}
|
||
}
|
||
```
|
||
|
||
Key indicators of a cold request:
|
||
- `"http_req_connection_reused": {"value": 0}` — fresh connection
|
||
- `"http_req_dns_lookup"` > 500ms — DNS resolution happened
|
||
- `"http_req_connecting"` > 500ms — TCP/TLS handshake happened
|
||
|
||
A reused connection has `"http_req_connection_reused": {"value": 1}` and near-zero DNS/connect times.
|