Files
hermes-skills/mlops/llm-inference-benchmarking/references/aiperf-long-input-benchmarking.md
T

110 lines
3.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# AIPerf Long-Input and Extreme-Context Benchmarking
Session: 2026-06-18 — Full suite benchmark of 4 models (gemma-4-31b-it, gpt-oss-120b, moonshotai/kimi-k2.6, qwen3.6-27b-nvfp4) on ai.noris.de
## Key Finding: Long Inputs > High Concurrency for Bottleneck Isolation
When evaluating maximum node capacity, testing conc=200+ is often operationally irrelevant because TTFT p99 exceeds acceptable thresholds (7s+ at conc=100). The real infrastructure bottleneck for RAG / long-context workloads is **memory-bandwidth-bound prefill** with long inputs.
Recommended alternative: **Long-Input Prefill Test**
- Concurrency: Fixed low (e.g., 10) to avoid queueing artifacts
- Input length: 1024 / 2048 / 4096 tokens (double/triple the standard 512)
- Output length: Short (128) to isolate prefill from generation
- Duration: 300s to amortize cold-start
## Extreme Context Test (10k Input / 10k Output)
For stress-testing context-window limits and sustained generation capacity:
- **Input:** 10,000 tokens
- **Output:** 10,000 tokens
- **Concurrency:** 1 (sequentially, because runtime is 2001000 seconds per request)
- **Requests:** 10 per model
- **Purpose:** Validates the full KV-cache lifecycle and sustained decode performance
**Expected runtime:** At 1050 tok/s output, one request takes 2001000 seconds. Schedule accordingly.
**Template:**
```yaml
schema_version: "2.0"
benchmark:
models:
items:
- name: "vllm/MODEL_NAME"
endpoint:
urls: ["https://ai.noris.de/v1"]
type: "chat"
streaming: true
datasets:
- type: "synthetic"
name: "main"
entries: 10
isl:
type: "fixed"
value: 10000
osl:
type: "fixed"
value: 10000
phases:
- type: "concurrency"
name: "profiling"
requests: 10
```
## Correct AIPerf 0.10.0 Synthetic Dataset Config (Long Input)
```yaml
schema_version: "2.0"
benchmark:
models:
items:
- name: "vllm/gemma-4-31b-it"
endpoint:
urls: ["https://ai.noris.de/v1"]
type: "chat"
streaming: true
datasets:
- name: "synthetic_long"
type: "synthetic"
prompt_template:
content: "Analyze the following topic in depth and provide detailed insights."
strategy:
type: "fixed"
max_total_tokens: 1152
input_tokens: 1024
output_tokens: 128
phases:
- type: "concurrency"
name: "profiling"
duration: 300
```
**Critical YAML structure rules learned:**
- `datasets` MUST be a **list** (`- name: ...`), NOT a dict with `default:` key
- `name` goes directly under the list item, NOT nested under a `synthetic:` key
- `prompt_template` and `strategy` are siblings to `name` and `type`
## Benchmark Results at conc=100 (Standard 512-in / 128-out)
| Model | Input tok/s (total) | Output tok/s (total) | Total tok/s | TTFT p99 |
|-------|--------------------|---------------------|-------------|----------|
| moonshotai-kimi-k2.6 | 5,115 | 1,270 | 6,385 | 3,926ms |
| gemma-4-31b-it | 4,385 | 1,083 | 5,468 | 2,060ms |
| qwen3.6-27b-nvfp4 | 4,191 | 992 | 5,183 | 7,391ms |
| gpt-oss-120b | 3,300 | 804 | 4,104 | 7,310ms |
**Prefill degradation @ conc=100:** All models lose ~80% of per-user prefill speed vs conc=10.
## MSP Service-Tier Perspective (User Preference)
When presenting for infrastructure/commercial decisions, structure per **service tier** not model comparison:
- Single-User Experience (conc=1): TTFT p99, Output/User, SLA class
- Multi-Tenant Load (conc=10→100): Throughput scaling, degradation curve
- Steady-State: Long-run stability, drift detection
- Service Recommendation: Positioning (Entry/Premium), suitable use cases
## Related Files
- `templates/aiperf-long-input-config.yaml` — Working starter template
- `references/aiperf-suite-orchestrator-pattern.md` — Full orchestration pattern
- `references/aiperf-token-metrics-explained.md` — Token metric unit reference