Initial commit: Hermes Agent Skills collection
This commit is contained in:
@@ -0,0 +1,86 @@
|
||||
# Noris: vLLM Reasoning Format Behavior
|
||||
|
||||
## Observed Behavior (ai.noris.de, vLLM 0.22.1)
|
||||
|
||||
Models on the Noris vLLM cluster that have chain-of-thought reasoning enabled return **only a `reasoning` field**, with `content: null`:
|
||||
|
||||
```json
|
||||
{
|
||||
"choices": [{
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": null,
|
||||
"reasoning": "Here's a thinking process:\n\n1. **Analyze User Input:** ...",
|
||||
"reasoning_details": [{
|
||||
"index": 0,
|
||||
"type": "reasoning.text",
|
||||
"text": "Here's a thinking process:\n\n1. **Analyze User Input:** ..."
|
||||
}]
|
||||
}
|
||||
}],
|
||||
"usage": {
|
||||
"prompt_tokens": 21,
|
||||
"completion_tokens": 100,
|
||||
"total_tokens": 121
|
||||
},
|
||||
"system_fingerprint": "vllm-0.22.1-a8c0978c",
|
||||
"extra_fields": {
|
||||
"request_type": "chat_completion",
|
||||
"provider": "vllm",
|
||||
"latency": 1738,
|
||||
"chunk_index": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## What This Means for Clients
|
||||
|
||||
| Field | Value | Implication |
|
||||
|-------|-------|-------------|
|
||||
| `content` | `null` | No final assistant message text |
|
||||
| `reasoning` | Non-empty string | Chain-of-thought thinking process |
|
||||
| `reasoning_details` | Array of blocks | Token-level reasoning decomposition |
|
||||
| `finish_reason` | `"length"` | Fixed length responses even if reasoning completes mid-sentence |
|
||||
|
||||
## Models Affected
|
||||
|
||||
| Model | Reasoning Field |
|
||||
|-------|-----------------|
|
||||
| `vllm/qwen3.6-27b-nvfp4` | ✅ reasoning-only |
|
||||
| `vllm/moonshotai/kimi-k2.6` | Unknown (not tested for this pattern) |
|
||||
| `vllm/gpt-oss-120b` | Unknown |
|
||||
| `vllm/gemma-4-31b-it` | Unknown |
|
||||
|
||||
## Client Handling
|
||||
|
||||
When integrating with Noris vLLM reasoning models, clients must:
|
||||
|
||||
1. **Check `reasoning` before `content`** — `content` will be `null`
|
||||
2. **Extract text from `reasoning` field** for user-facing output
|
||||
3. **Account for `finish_reason: "length"`** — reasoning may be truncated mid-thought
|
||||
4. **Consider latency** — reasoning models are slower (~890ms–1700ms for 50–100 tokens on Noris)
|
||||
|
||||
## Example: Minimal curl test with reasoning extraction
|
||||
|
||||
```bash
|
||||
curl -sS -X POST "https://ai.noris.de/v1/chat/completions" \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer sk-bf-..." \
|
||||
-d '{
|
||||
"model": "vllm/qwen3.6-27b-nvfp4",
|
||||
"max_tokens": 100,
|
||||
"messages": [{"role": "user", "content": "What is prompt caching?"}]
|
||||
}' | python3 -c "
|
||||
import sys, json
|
||||
d = json.load(sys.stdin)
|
||||
msg = d['choices'][0]['message']
|
||||
print('Content:', msg.get('content'))
|
||||
print('Has reasoning:', 'reasoning' in msg)
|
||||
print('Reasoning preview:', msg.get('reasoning', 'N/A')[:200])
|
||||
print('Usage:', d['usage'])
|
||||
"
|
||||
```
|
||||
|
||||
## OpenAI API Spec Compliance
|
||||
|
||||
This is **non-standard** OpenAI-format behavior. Per OpenAI spec, `content` should contain the assistant's message text. vLLM with reasoning enabled diverges by putting the model output in `reasoning` and leaving `content` null. Clients expecting standard OpenAI responses will see empty assistant messages unless they handle this field.
|
||||
Reference in New Issue
Block a user