Files
hermes-skills/mlops/llm-inference-benchmarking/references/aiperf-suite-orchestrator-pattern.md
T

121 lines
3.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# AIPerf Suite Orchestrator Pattern
Reference: building a 12-run benchmark suite (4 models × 3 scenarios) against ai.noris.de.
## Directory Convention
```
aiperf-benchmark-suite/
├── configs/
│ ├── gemma-4-31b-it/
│ │ ├── sales.yaml # ShareGPT dataset, requests: 50, sessions: 1
│ │ ├── capacity.yaml # Synthetic 512/128, requests: 100
│ │ └── steady.yaml # ShareGPT, duration: 600
│ ├── gpt-oss-120b/...
│ └── ... (one dir per model)
├── scripts/
│ ├── run-suite.sh # Orchestrator
│ ├── generate-summary.py # Markdown table from JSON
│ └── plot-results.py # Matplotlib charts + CSV
├── results/
│ └── YYYY-MM-DD-HHMMSS/
│ ├── gemma-4-31b-it/
│ │ ├── sales/conc1/
│ │ ├── sales/conc5/
│ │ ├── sales/conc10/
│ │ ├── capacity/conc10/
│ │ ├── capacity/conc25/
│ │ ├── capacity/conc50/
│ │ ├── capacity/conc100/
│ │ └── steady/
│ └── ...
└── venv/ # AIPerf installed here
```
## Orchestrator Structure
The orchestrator is a bash script that:
1. **Sets absolute paths**`AIPERF="$BASE_DIR/venv/bin/aiperf"` (never rely on `$PATH`)
2. **Validates env** — checks `NORIS_API_KEY` and connectivity
3. **Loops models × scenarios**
4. **Per scenario, loops concurrency levels** (for count-based scenarios)
5. **Calls AIPerf with CLI overrides** for each sub-run
6. **Pauses between runs** (`sleep 5-30`) to avoid rate limits
7. **Generates summary + plots** after all runs complete
8. **Continues on error** (`|| true`) — one failed sub-run doesn't kill the suite
## Critical: Venv Path Fix
When running in background (cron, detached shell, Telegram-triggered runs), `$PATH` does NOT include the venv. The orchestrator MUST:
```bash
AIPERF="$BASE_DIR/venv/bin/aiperf"
if [ ! -x "$AIPERF" ]; then
echo "ERROR: AIPerf not found at $AIPERF"
exit 1
fi
# Then call "$AIPERF" profile ... everywhere
```
Failure mode: `aiperf: command not found` on every run, producing empty results silently.
## CLI Override Pattern for Suite Runs
The YAML configs define the base structure. CLI flags sweep parameters:
```bash
# Sales scenario: sweep concurrencies
for conc in 1 5 10; do
"$AIPERF" profile --config "$CONFIG" \
--concurrency "$conc" \
--request-count 50 \
--artifact-dir "results/sales/conc${conc}"
done
# Capacity scenario: stepped ramp
for conc in 10 25 50 100; do
"$AIPERF" profile --config "$CONFIG" \
--concurrency "$conc" \
--request-count 100 \
--artifact-dir "results/capacity/conc${conc}"
done
# Steady scenario: time-based
"$AIPERF" profile --config "$CONFIG" \
--concurrency 20 \
--artifact-dir "results/steady"
```
## Pre-Download Strategy for ShareGPT
First sales run downloads ShareGPT from HuggingFace (~25 min). To avoid this during production:
```bash
# One-time warm-up (any model config with public dataset works)
venv/bin/aiperf profile --config configs/any-model/sales.yaml \
--concurrency 1 --request-count 1 \
--artifact-dir /tmp/warmup --no-gpu-telemetry --no-server-metrics
```
Subsequent runs use `.cache/aiperf/datasets/` instantly.
## Post-Processing Sequence
After the orchestrator finishes all runs:
```bash
# 1. Markdown summary
python3 scripts/generate-summary.py results/YYYY-MM-DD-HHMMSS
# 2. Plots + CSV
python3 scripts/plot-results.py results/YYYY-MM-DD-HHMMSS
```
Outputs:
- `results/.../README.md` human-readable markdown table
- `results/.../plots/summary.csv` Excel / Google Sheets import
- `results/.../plots/capacity_curve.png` throughput vs concurrency
- `results/.../plots/latency_curve.png` TTFT p99 vs concurrency
- `results/.../plots/sales_comparison.png` low-concurrency latency