Initial commit: Hermes Agent Skills collection

This commit is contained in:
Debian
2026-07-12 19:02:59 +00:00
commit e9cc106625
789 changed files with 233126 additions and 0 deletions
@@ -0,0 +1,133 @@
# Long-Context Prefill Test Memory Bandwidth Bottleneck
## Purpose
Isolate the **memory-bandwidth prefill bottleneck** (relevant for RAG, document analysis, long-context agents). While capacity tests scale concurrency with short inputs, this test scales **input length** at fixed low concurrency to find where KV-cache limits hit.
## When to Use
- The capacity curve shows conc=100 is already the practical ceiling
- The real workload involves **long documents** (1K8K tokens) not many short chats
- You need to answer: "Can this model handle our RAG pipeline with 4K context windows?"
## AIPerf Config
Use `type: "synthetic"` with `fixed` ISL. The `prompt_template` must include a variable that expands to the target length. AIPerf's synthetic engine automatically generates filler text to reach the target token count.
### Template: 1024-token input
```yaml
schema_version: "2.0"
benchmark:
models:
items:
- name: "vllm/REPLACE-MODEL-NAME"
endpoint:
urls: ["https://ENDPOINT/v1"]
type: "chat"
streaming: true
datasets:
default:
- name: "synthetic_long"
type: "synthetic"
prompt_template:
content: "Summarize the following text concisely: {{text}}"
variables:
- name: "text"
values: [
"artificial intelligence",
"quantum computing",
"blockchain technology",
"climate change",
"machine learning",
"renewable energy"
]
strategy:
type: "fixed"
max_total_tokens: 1152 # input + output must fit
input_tokens: 1024 # target input length
output_tokens: 128 # short output to focus on prefill
phases:
- type: "concurrency"
name: "prefill_stress"
duration: 300 # 5 minutes, time-based
sessions: 10 # fixed low concurrency
tokenizer:
name: "builtin"
```
### Scaling input length
Create a sweep by varying `input_tokens` across runs:
| Run | input_tokens | max_total_tokens | What it tests |
|-----|-------------|------------------|---------------|
| 1 | 512 | 640 | Baseline (standard context) |
| 2 | 1024 | 1152 | Long context (RAG documents) |
| 3 | 2048 | 2176 | Very long context |
| 4 | 4096 | 4224 | Extreme context (edge case) |
**Constraint:** `max_total_tokens >= input_tokens + output_tokens + 16` (AIPerf padding). The model's own context window (`max_model_len`) is the hard ceiling—check the endpoint's `/v1/models` response.
## CLI Execution
```bash
for islen in 512 1024 2048; do
model="vllm/gemma-4-31b-it"
model_safe=$(echo "$model" | tr '/.' '_')
# Generate per-run config or use sed to modify input_tokens
sed "s/input_tokens: 1024/input_tokens: ${islen}/; s/max_total_tokens: 1152/max_total_tokens: $((islen + 128 + 16))/" \
long-context-template.yaml > /tmp/run.yaml
venv/bin/aiperf profile \
--config /tmp/run.yaml \
--api-key "$API_KEY" \
--concurrency 10 \
--ui none \
--no-gpu-telemetry \
--no-server-metrics \
--artifact-dir "results/longcontext/${model_safe}/isl${islen}"
done
```
## Metrics to Extract
From `profile_export_aiperf.json`:
| Metric | Key | Unit | Interpretation |
|--------|-----|------|----------------|
| Prefill throughput | `prefill_throughput_per_user.avg` | tok/s/user | How fast a single user's long prompt is processed |
| TTFT | `time_to_first_token.avg` / `.p99` | ms | Time until first token at long context |
| Input length | `input_sequence_length.avg` | tokens | Actual achieved input length (verify synthetic hit target) |
| Output/User | `output_token_throughput_per_user.avg` | tok/s/user | Generation speed (should be stable; if it drops, KV-cache eviction) |
## Expected Behavior
Healthy system:
- Prefill throughput **degrades gracefully** with longer inputs (sub-linear)
- TTFT scales roughly linearly with input length: `TTFT ≈ input_tokens / prefill_tps`
- Output/User stays constant (generation is independent of input length)
Unhealthy system (KV-cache limits):
- Prefill throughput **collapses** at a threshold length (e.g., 2048 tokens)
- TTFT becomes super-linear (quadratic or worse)
- Output/User drops (KV-cache thrashing or CPU offloading)
## Interpreting Results
Calculate **prefill time per 1K tokens**:
```
Prefill time per 1K = 1000 / prefill_throughput_per_user (seconds)
```
From observed conc=100 data (short-input baseline):
| Model | Prefill/User @ 512 tok | Est. per 1K |
|-------|----------------------|-------------|
| gemma-4-31b-it | 505 tok/s | ~2.0s |
| moonshotai-kimi-k2.6 | 276 tok/s | ~3.6s |
| gpt-oss-120b | 137 tok/s | ~7.3s |
If the long-context test at 1024 tokens shows significantly worse than double these estimates, the model/server has **non-linear KV-cache scaling**—a hard limit for RAG workloads.