Files

3.5 KiB

Noris Provider: OpenAI Endpoint (vLLM)

Endpoint Status

Noris currently runs vLLM with OpenAI-compatible API only. The Anthropic-compatible endpoint does not work.

Format Base URL Status
OpenAI https://ai.noris.de/v1 Working
Anthropic https://ai.noris.de/anthropic HTTP 405 (Method Not Allowed)

Only endpoint that works: POST https://ai.noris.de/v1/chat/completions

Correct Hermes Config

providers:
  noris:
    type: openai
    base_url: https://ai.noris.de/v1
    api_key: sk-bf-8779d73c-6a51-49d6-a22f-00a0290ca6a7

Critical: Use https://, NOT http:// — HTTP returns 503. The path must be exactly https://ai.noris.de/v1 (no trailing /chat/completions; Hermes appends that).

Prompt Caching Status: NOT FUNCTIONAL

Test result (2026-06-18): cache_control blocks are silently ignored by the vLLM backend. vLLM may do transparent prefix caching internally, but the API does not expose cache hit/miss counters.

  • Anthropic-style cache_control: {type: "ephemeral"} → not supported (Anthropic endpoint is 405 anyway)
  • OpenAI-style prefix caching → vLLM handles this automatically, but no API visibility

Why Neither Caching Approach Works via API

The Anthropic endpoint (/anthropic) returns HTTP 405 for all tested paths (/, /v1/messages). The OpenAI endpoint (/v1/chat/completions) runs vLLM 0.22.1, which does not expose cache_read_input_tokens or similar in the usage object. Any client-side cache accounting will always show zero.

Cost / Speed Implications

  • No explicit caching benefit visible to client. Every request carries full prompt cost.
  • No speed gain measurable via API. vLLM may internally cache KV-cache prefixes, but this is opaque.
  • No harm sending cache_control (it is ignored), but also no benefit.

Recommendation

Always use the OpenAI endpoint (type: openai, base_url: https://ai.noris.de/v1). Do not attempt Anthropic format with Noris.

Verification Command (OpenAI)

API_KEY="$(grep -A1 'providers:' ~/.hermes/config.yaml | grep api_key | head -n1 | sed 's/.*: //')"

curl -s -X POST https://ai.noris.de/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ${API_KEY}" \
  -d '{
    "model": "vllm/qwen3.6-27b-nvfp4",
    "max_tokens": 50,
    "messages": [{"role": "user", "content": "Say hi"}]
  }' | python3 -c "import sys,json; d=json.load(sys.stdin); print('Model:', d.get('model')); print('Usage:', d.get('usage')); print('Content:', d.get('choices',[{}])[0].get('message',{}).get('content'))"

Noris Virtual Key Routing: vllm/ Prefix Required

Noris uses virtual-key routing — the vllm/ prefix in model names is not optional, it is how the provider routes your request to the correct inference backend. Omitting it causes "virtual key not found" errors even when the API key is valid.

Wrong Right
moonshotai/kimi-k2.6 vllm/moonshotai/kimi-k2.6

This applies everywhere: model.model, model.default, cronjob model fields, auxiliary overrides, and inline /model commands.

Hermes itself strips the vllm/ prefix when calling the provider, but Noris requires it in the model identifier. If you see HTTP 401 "virtual key not found" with a valid API key, check that the model name includes the vllm/ prefix.

Quick verification

hermes config get model.model   # should start with "vllm/"
hermes cron list                # check model column for missing prefix