2.9 KiB
2.9 KiB
Noris: vLLM Reasoning Format Behavior
Observed Behavior (ai.noris.de, vLLM 0.22.1)
Models on the Noris vLLM cluster that have chain-of-thought reasoning enabled return only a reasoning field, with content: null:
{
"choices": [{
"message": {
"role": "assistant",
"content": null,
"reasoning": "Here's a thinking process:\n\n1. **Analyze User Input:** ...",
"reasoning_details": [{
"index": 0,
"type": "reasoning.text",
"text": "Here's a thinking process:\n\n1. **Analyze User Input:** ..."
}]
}
}],
"usage": {
"prompt_tokens": 21,
"completion_tokens": 100,
"total_tokens": 121
},
"system_fingerprint": "vllm-0.22.1-a8c0978c",
"extra_fields": {
"request_type": "chat_completion",
"provider": "vllm",
"latency": 1738,
"chunk_index": 0
}
}
What This Means for Clients
| Field | Value | Implication |
|---|---|---|
content |
null |
No final assistant message text |
reasoning |
Non-empty string | Chain-of-thought thinking process |
reasoning_details |
Array of blocks | Token-level reasoning decomposition |
finish_reason |
"length" |
Fixed length responses even if reasoning completes mid-sentence |
Models Affected
| Model | Reasoning Field |
|---|---|
vllm/qwen3.6-27b-nvfp4 |
✅ reasoning-only |
vllm/moonshotai/kimi-k2.6 |
Unknown (not tested for this pattern) |
vllm/gpt-oss-120b |
Unknown |
vllm/gemma-4-31b-it |
Unknown |
Client Handling
When integrating with Noris vLLM reasoning models, clients must:
- Check
reasoningbeforecontent—contentwill benull - Extract text from
reasoningfield for user-facing output - Account for
finish_reason: "length"— reasoning may be truncated mid-thought - Consider latency — reasoning models are slower (~890ms–1700ms for 50–100 tokens on Noris)
Example: Minimal curl test with reasoning extraction
curl -sS -X POST "https://ai.noris.de/v1/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-bf-..." \
-d '{
"model": "vllm/qwen3.6-27b-nvfp4",
"max_tokens": 100,
"messages": [{"role": "user", "content": "What is prompt caching?"}]
}' | python3 -c "
import sys, json
d = json.load(sys.stdin)
msg = d['choices'][0]['message']
print('Content:', msg.get('content'))
print('Has reasoning:', 'reasoning' in msg)
print('Reasoning preview:', msg.get('reasoning', 'N/A')[:200])
print('Usage:', d['usage'])
"
OpenAI API Spec Compliance
This is non-standard OpenAI-format behavior. Per OpenAI spec, content should contain the assistant's message text. vLLM with reasoning enabled diverges by putting the model output in reasoning and leaving content null. Clients expecting standard OpenAI responses will see empty assistant messages unless they handle this field.