Files

4.8 KiB
Raw Permalink Blame History

Long-Context Prefill Test Memory Bandwidth Bottleneck

Purpose

Isolate the memory-bandwidth prefill bottleneck (relevant for RAG, document analysis, long-context agents). While capacity tests scale concurrency with short inputs, this test scales input length at fixed low concurrency to find where KV-cache limits hit.

When to Use

  • The capacity curve shows conc=100 is already the practical ceiling
  • The real workload involves long documents (1K8K tokens) not many short chats
  • You need to answer: "Can this model handle our RAG pipeline with 4K context windows?"

AIPerf Config

Use type: "synthetic" with fixed ISL. The prompt_template must include a variable that expands to the target length. AIPerf's synthetic engine automatically generates filler text to reach the target token count.

Template: 1024-token input

schema_version: "2.0"
benchmark:
  models:
    items:
      - name: "vllm/REPLACE-MODEL-NAME"
  endpoint:
    urls: ["https://ENDPOINT/v1"]
    type: "chat"
    streaming: true
  datasets:
    default:
      - name: "synthetic_long"
        type: "synthetic"
        prompt_template:
          content: "Summarize the following text concisely: {{text}}"
        variables:
          - name: "text"
            values: [
              "artificial intelligence",
              "quantum computing",
              "blockchain technology",
              "climate change",
              "machine learning",
              "renewable energy"
            ]
    strategy:
      type: "fixed"
      max_total_tokens: 1152    # input + output must fit
      input_tokens: 1024        # target input length
      output_tokens: 128        # short output to focus on prefill
  phases:
    - type: "concurrency"
      name: "prefill_stress"
      duration: 300             # 5 minutes, time-based
      sessions: 10              # fixed low concurrency
  tokenizer:
    name: "builtin"

Scaling input length

Create a sweep by varying input_tokens across runs:

Run input_tokens max_total_tokens What it tests
1 512 640 Baseline (standard context)
2 1024 1152 Long context (RAG documents)
3 2048 2176 Very long context
4 4096 4224 Extreme context (edge case)

Constraint: max_total_tokens >= input_tokens + output_tokens + 16 (AIPerf padding). The model's own context window (max_model_len) is the hard ceiling—check the endpoint's /v1/models response.

CLI Execution

for islen in 512 1024 2048; do
  model="vllm/gemma-4-31b-it"
  model_safe=$(echo "$model" | tr '/.' '_')
  
  # Generate per-run config or use sed to modify input_tokens
  sed "s/input_tokens: 1024/input_tokens: ${islen}/; s/max_total_tokens: 1152/max_total_tokens: $((islen + 128 + 16))/" \
    long-context-template.yaml > /tmp/run.yaml
  
  venv/bin/aiperf profile \
    --config /tmp/run.yaml \
    --api-key "$API_KEY" \
    --concurrency 10 \
    --ui none \
    --no-gpu-telemetry \
    --no-server-metrics \
    --artifact-dir "results/longcontext/${model_safe}/isl${islen}"
done

Metrics to Extract

From profile_export_aiperf.json:

Metric Key Unit Interpretation
Prefill throughput prefill_throughput_per_user.avg tok/s/user How fast a single user's long prompt is processed
TTFT time_to_first_token.avg / .p99 ms Time until first token at long context
Input length input_sequence_length.avg tokens Actual achieved input length (verify synthetic hit target)
Output/User output_token_throughput_per_user.avg tok/s/user Generation speed (should be stable; if it drops, KV-cache eviction)

Expected Behavior

Healthy system:

  • Prefill throughput degrades gracefully with longer inputs (sub-linear)
  • TTFT scales roughly linearly with input length: TTFT ≈ input_tokens / prefill_tps
  • Output/User stays constant (generation is independent of input length)

Unhealthy system (KV-cache limits):

  • Prefill throughput collapses at a threshold length (e.g., 2048 tokens)
  • TTFT becomes super-linear (quadratic or worse)
  • Output/User drops (KV-cache thrashing or CPU offloading)

Interpreting Results

Calculate prefill time per 1K tokens:

Prefill time per 1K = 1000 / prefill_throughput_per_user (seconds)

From observed conc=100 data (short-input baseline):

Model Prefill/User @ 512 tok Est. per 1K
gemma-4-31b-it 505 tok/s ~2.0s
moonshotai-kimi-k2.6 276 tok/s ~3.6s
gpt-oss-120b 137 tok/s ~7.3s

If the long-context test at 1024 tokens shows significantly worse than double these estimates, the model/server has non-linear KV-cache scaling—a hard limit for RAG workloads.