Initial commit: Hermes Agent Skills collection
This commit is contained in:
@@ -0,0 +1,79 @@
|
||||
# Noris Provider: OpenAI Endpoint (vLLM)
|
||||
|
||||
## Endpoint Status
|
||||
|
||||
Noris currently runs **vLLM with OpenAI-compatible API only**. The Anthropic-compatible endpoint does **not work**.
|
||||
|
||||
| Format | Base URL | Status |
|
||||
|--------|----------|--------|
|
||||
| OpenAI | `https://ai.noris.de/v1` | ✅ Working |
|
||||
| Anthropic | `https://ai.noris.de/anthropic` | ❌ HTTP 405 (Method Not Allowed) |
|
||||
|
||||
**Only endpoint that works:** `POST https://ai.noris.de/v1/chat/completions`
|
||||
|
||||
## Correct Hermes Config
|
||||
|
||||
```yaml
|
||||
providers:
|
||||
noris:
|
||||
type: openai
|
||||
base_url: https://ai.noris.de/v1
|
||||
api_key: sk-bf-8779d73c-6a51-49d6-a22f-00a0290ca6a7
|
||||
```
|
||||
|
||||
**Critical:** Use `https://`, **NOT** `http://` — HTTP returns 503. The path must be exactly `https://ai.noris.de/v1` (no trailing `/chat/completions`; Hermes appends that).
|
||||
|
||||
## Prompt Caching Status: NOT FUNCTIONAL
|
||||
|
||||
**Test result (2026-06-18):** `cache_control` blocks are silently ignored by the vLLM backend. vLLM may do transparent prefix caching internally, but the API does not expose cache hit/miss counters.
|
||||
|
||||
- Anthropic-style `cache_control: {type: "ephemeral"}` → not supported (Anthropic endpoint is 405 anyway)
|
||||
- OpenAI-style prefix caching → vLLM handles this automatically, but no API visibility
|
||||
|
||||
## Why Neither Caching Approach Works via API
|
||||
|
||||
The Anthropic endpoint (`/anthropic`) returns HTTP 405 for all tested paths (`/`, `/v1/messages`). The OpenAI endpoint (`/v1/chat/completions`) runs vLLM 0.22.1, which does **not** expose `cache_read_input_tokens` or similar in the `usage` object. Any client-side cache accounting will always show zero.
|
||||
|
||||
## Cost / Speed Implications
|
||||
|
||||
- **No explicit caching benefit visible to client.** Every request carries full prompt cost.
|
||||
- **No speed gain measurable via API.** vLLM may internally cache KV-cache prefixes, but this is opaque.
|
||||
- **No harm** sending `cache_control` (it is ignored), but also no benefit.
|
||||
|
||||
## Recommendation
|
||||
|
||||
Always use the OpenAI endpoint (`type: openai`, `base_url: https://ai.noris.de/v1`). Do not attempt Anthropic format with Noris.
|
||||
|
||||
## Verification Command (OpenAI)
|
||||
|
||||
```bash
|
||||
API_KEY="$(grep -A1 'providers:' ~/.hermes/config.yaml | grep api_key | head -n1 | sed 's/.*: //')"
|
||||
|
||||
curl -s -X POST https://ai.noris.de/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer ${API_KEY}" \
|
||||
-d '{
|
||||
"model": "vllm/qwen3.6-27b-nvfp4",
|
||||
"max_tokens": 50,
|
||||
"messages": [{"role": "user", "content": "Say hi"}]
|
||||
}' | python3 -c "import sys,json; d=json.load(sys.stdin); print('Model:', d.get('model')); print('Usage:', d.get('usage')); print('Content:', d.get('choices',[{}])[0].get('message',{}).get('content'))"
|
||||
```
|
||||
|
||||
## Noris Virtual Key Routing: `vllm/` Prefix Required
|
||||
|
||||
Noris uses **virtual-key routing** — the `vllm/` prefix in model names is **not optional**, it is how the provider routes your request to the correct inference backend. Omitting it causes "virtual key not found" errors even when the API key is valid.
|
||||
|
||||
| ❌ Wrong | ✅ Right |
|
||||
|---------|----------|
|
||||
| `moonshotai/kimi-k2.6` | `vllm/moonshotai/kimi-k2.6` |
|
||||
|
||||
**This applies everywhere:** `model.model`, `model.default`, cronjob `model` fields, auxiliary overrides, and inline `/model` commands.
|
||||
|
||||
Hermes itself strips the `vllm/` prefix when calling the provider, but Noris requires it in the model identifier. If you see **HTTP 401 "virtual key not found"** with a valid API key, check that the model name includes the `vllm/` prefix.
|
||||
|
||||
### Quick verification
|
||||
|
||||
```bash
|
||||
hermes config get model.model # should start with "vllm/"
|
||||
hermes cron list # check model column for missing prefix
|
||||
```
|
||||
Reference in New Issue
Block a user