# Noris Provider: OpenAI Endpoint (vLLM) ## Endpoint Status Noris currently runs **vLLM with OpenAI-compatible API only**. The Anthropic-compatible endpoint does **not work**. | Format | Base URL | Status | |--------|----------|--------| | OpenAI | `https://ai.noris.de/v1` | ✅ Working | | Anthropic | `https://ai.noris.de/anthropic` | ❌ HTTP 405 (Method Not Allowed) | **Only endpoint that works:** `POST https://ai.noris.de/v1/chat/completions` ## Correct Hermes Config ```yaml providers: noris: type: openai base_url: https://ai.noris.de/v1 api_key: sk-bf-8779d73c-6a51-49d6-a22f-00a0290ca6a7 ``` **Critical:** Use `https://`, **NOT** `http://` — HTTP returns 503. The path must be exactly `https://ai.noris.de/v1` (no trailing `/chat/completions`; Hermes appends that). ## Prompt Caching Status: NOT FUNCTIONAL **Test result (2026-06-18):** `cache_control` blocks are silently ignored by the vLLM backend. vLLM may do transparent prefix caching internally, but the API does not expose cache hit/miss counters. - Anthropic-style `cache_control: {type: "ephemeral"}` → not supported (Anthropic endpoint is 405 anyway) - OpenAI-style prefix caching → vLLM handles this automatically, but no API visibility ## Why Neither Caching Approach Works via API The Anthropic endpoint (`/anthropic`) returns HTTP 405 for all tested paths (`/`, `/v1/messages`). The OpenAI endpoint (`/v1/chat/completions`) runs vLLM 0.22.1, which does **not** expose `cache_read_input_tokens` or similar in the `usage` object. Any client-side cache accounting will always show zero. ## Cost / Speed Implications - **No explicit caching benefit visible to client.** Every request carries full prompt cost. - **No speed gain measurable via API.** vLLM may internally cache KV-cache prefixes, but this is opaque. - **No harm** sending `cache_control` (it is ignored), but also no benefit. ## Recommendation Always use the OpenAI endpoint (`type: openai`, `base_url: https://ai.noris.de/v1`). Do not attempt Anthropic format with Noris. ## Verification Command (OpenAI) ```bash API_KEY="$(grep -A1 'providers:' ~/.hermes/config.yaml | grep api_key | head -n1 | sed 's/.*: //')" curl -s -X POST https://ai.noris.de/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${API_KEY}" \ -d '{ "model": "vllm/qwen3.6-27b-nvfp4", "max_tokens": 50, "messages": [{"role": "user", "content": "Say hi"}] }' | python3 -c "import sys,json; d=json.load(sys.stdin); print('Model:', d.get('model')); print('Usage:', d.get('usage')); print('Content:', d.get('choices',[{}])[0].get('message',{}).get('content'))" ``` ## Noris Virtual Key Routing: `vllm/` Prefix Required Noris uses **virtual-key routing** — the `vllm/` prefix in model names is **not optional**, it is how the provider routes your request to the correct inference backend. Omitting it causes "virtual key not found" errors even when the API key is valid. | ❌ Wrong | ✅ Right | |---------|----------| | `moonshotai/kimi-k2.6` | `vllm/moonshotai/kimi-k2.6` | **This applies everywhere:** `model.model`, `model.default`, cronjob `model` fields, auxiliary overrides, and inline `/model` commands. Hermes itself strips the `vllm/` prefix when calling the provider, but Noris requires it in the model identifier. If you see **HTTP 401 "virtual key not found"** with a valid API key, check that the model name includes the `vllm/` prefix. ### Quick verification ```bash hermes config get model.model # should start with "vllm/" hermes cron list # check model column for missing prefix ```