4.8 KiB
Long-Context Prefill Test – Memory Bandwidth Bottleneck
Purpose
Isolate the memory-bandwidth prefill bottleneck (relevant for RAG, document analysis, long-context agents). While capacity tests scale concurrency with short inputs, this test scales input length at fixed low concurrency to find where KV-cache limits hit.
When to Use
- The capacity curve shows conc=100 is already the practical ceiling
- The real workload involves long documents (1K–8K tokens) not many short chats
- You need to answer: "Can this model handle our RAG pipeline with 4K context windows?"
AIPerf Config
Use type: "synthetic" with fixed ISL. The prompt_template must include a variable that expands to the target length. AIPerf's synthetic engine automatically generates filler text to reach the target token count.
Template: 1024-token input
schema_version: "2.0"
benchmark:
models:
items:
- name: "vllm/REPLACE-MODEL-NAME"
endpoint:
urls: ["https://ENDPOINT/v1"]
type: "chat"
streaming: true
datasets:
default:
- name: "synthetic_long"
type: "synthetic"
prompt_template:
content: "Summarize the following text concisely: {{text}}"
variables:
- name: "text"
values: [
"artificial intelligence",
"quantum computing",
"blockchain technology",
"climate change",
"machine learning",
"renewable energy"
]
strategy:
type: "fixed"
max_total_tokens: 1152 # input + output must fit
input_tokens: 1024 # target input length
output_tokens: 128 # short output to focus on prefill
phases:
- type: "concurrency"
name: "prefill_stress"
duration: 300 # 5 minutes, time-based
sessions: 10 # fixed low concurrency
tokenizer:
name: "builtin"
Scaling input length
Create a sweep by varying input_tokens across runs:
| Run | input_tokens | max_total_tokens | What it tests |
|---|---|---|---|
| 1 | 512 | 640 | Baseline (standard context) |
| 2 | 1024 | 1152 | Long context (RAG documents) |
| 3 | 2048 | 2176 | Very long context |
| 4 | 4096 | 4224 | Extreme context (edge case) |
Constraint: max_total_tokens >= input_tokens + output_tokens + 16 (AIPerf padding). The model's own context window (max_model_len) is the hard ceiling—check the endpoint's /v1/models response.
CLI Execution
for islen in 512 1024 2048; do
model="vllm/gemma-4-31b-it"
model_safe=$(echo "$model" | tr '/.' '_')
# Generate per-run config or use sed to modify input_tokens
sed "s/input_tokens: 1024/input_tokens: ${islen}/; s/max_total_tokens: 1152/max_total_tokens: $((islen + 128 + 16))/" \
long-context-template.yaml > /tmp/run.yaml
venv/bin/aiperf profile \
--config /tmp/run.yaml \
--api-key "$API_KEY" \
--concurrency 10 \
--ui none \
--no-gpu-telemetry \
--no-server-metrics \
--artifact-dir "results/longcontext/${model_safe}/isl${islen}"
done
Metrics to Extract
From profile_export_aiperf.json:
| Metric | Key | Unit | Interpretation |
|---|---|---|---|
| Prefill throughput | prefill_throughput_per_user.avg |
tok/s/user | How fast a single user's long prompt is processed |
| TTFT | time_to_first_token.avg / .p99 |
ms | Time until first token at long context |
| Input length | input_sequence_length.avg |
tokens | Actual achieved input length (verify synthetic hit target) |
| Output/User | output_token_throughput_per_user.avg |
tok/s/user | Generation speed (should be stable; if it drops, KV-cache eviction) |
Expected Behavior
Healthy system:
- Prefill throughput degrades gracefully with longer inputs (sub-linear)
- TTFT scales roughly linearly with input length:
TTFT ≈ input_tokens / prefill_tps - Output/User stays constant (generation is independent of input length)
Unhealthy system (KV-cache limits):
- Prefill throughput collapses at a threshold length (e.g., 2048 tokens)
- TTFT becomes super-linear (quadratic or worse)
- Output/User drops (KV-cache thrashing or CPU offloading)
Interpreting Results
Calculate prefill time per 1K tokens:
Prefill time per 1K = 1000 / prefill_throughput_per_user (seconds)
From observed conc=100 data (short-input baseline):
| Model | Prefill/User @ 512 tok | Est. per 1K |
|---|---|---|
| gemma-4-31b-it | 505 tok/s | ~2.0s |
| moonshotai-kimi-k2.6 | 276 tok/s | ~3.6s |
| gpt-oss-120b | 137 tok/s | ~7.3s |
If the long-context test at 1024 tokens shows significantly worse than double these estimates, the model/server has non-linear KV-cache scaling—a hard limit for RAG workloads.