3.5 KiB
Noris Provider: OpenAI Endpoint (vLLM)
Endpoint Status
Noris currently runs vLLM with OpenAI-compatible API only. The Anthropic-compatible endpoint does not work.
| Format | Base URL | Status |
|---|---|---|
| OpenAI | https://ai.noris.de/v1 |
✅ Working |
| Anthropic | https://ai.noris.de/anthropic |
❌ HTTP 405 (Method Not Allowed) |
Only endpoint that works: POST https://ai.noris.de/v1/chat/completions
Correct Hermes Config
providers:
noris:
type: openai
base_url: https://ai.noris.de/v1
api_key: sk-bf-8779d73c-6a51-49d6-a22f-00a0290ca6a7
Critical: Use https://, NOT http:// — HTTP returns 503. The path must be exactly https://ai.noris.de/v1 (no trailing /chat/completions; Hermes appends that).
Prompt Caching Status: NOT FUNCTIONAL
Test result (2026-06-18): cache_control blocks are silently ignored by the vLLM backend. vLLM may do transparent prefix caching internally, but the API does not expose cache hit/miss counters.
- Anthropic-style
cache_control: {type: "ephemeral"}→ not supported (Anthropic endpoint is 405 anyway) - OpenAI-style prefix caching → vLLM handles this automatically, but no API visibility
Why Neither Caching Approach Works via API
The Anthropic endpoint (/anthropic) returns HTTP 405 for all tested paths (/, /v1/messages). The OpenAI endpoint (/v1/chat/completions) runs vLLM 0.22.1, which does not expose cache_read_input_tokens or similar in the usage object. Any client-side cache accounting will always show zero.
Cost / Speed Implications
- No explicit caching benefit visible to client. Every request carries full prompt cost.
- No speed gain measurable via API. vLLM may internally cache KV-cache prefixes, but this is opaque.
- No harm sending
cache_control(it is ignored), but also no benefit.
Recommendation
Always use the OpenAI endpoint (type: openai, base_url: https://ai.noris.de/v1). Do not attempt Anthropic format with Noris.
Verification Command (OpenAI)
API_KEY="$(grep -A1 'providers:' ~/.hermes/config.yaml | grep api_key | head -n1 | sed 's/.*: //')"
curl -s -X POST https://ai.noris.de/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${API_KEY}" \
-d '{
"model": "vllm/qwen3.6-27b-nvfp4",
"max_tokens": 50,
"messages": [{"role": "user", "content": "Say hi"}]
}' | python3 -c "import sys,json; d=json.load(sys.stdin); print('Model:', d.get('model')); print('Usage:', d.get('usage')); print('Content:', d.get('choices',[{}])[0].get('message',{}).get('content'))"
Noris Virtual Key Routing: vllm/ Prefix Required
Noris uses virtual-key routing — the vllm/ prefix in model names is not optional, it is how the provider routes your request to the correct inference backend. Omitting it causes "virtual key not found" errors even when the API key is valid.
| ❌ Wrong | ✅ Right |
|---|---|
moonshotai/kimi-k2.6 |
vllm/moonshotai/kimi-k2.6 |
This applies everywhere: model.model, model.default, cronjob model fields, auxiliary overrides, and inline /model commands.
Hermes itself strips the vllm/ prefix when calling the provider, but Noris requires it in the model identifier. If you see HTTP 401 "virtual key not found" with a valid API key, check that the model name includes the vllm/ prefix.
Quick verification
hermes config get model.model # should start with "vllm/"
hermes cron list # check model column for missing prefix