16 KiBLFS
OpenAI-Compatible Model API Reference
This reference is for SkillsBench/OpenCode smoke tests and harness runs. It intentionally covers only OpenAI-compatible API format. Do not put API keys in this file; keys live in .env.
Default policy in this file: use the strongest available reasoning mode for every model. If a provider does not expose an explicit reasoning-strength knob, use the model's native thinking mode and preserve returned reasoning fields across multi-turn/tool-call turns.
Current Smoke-Tested Models
| Target | Model env | Base URL env | Highest reasoning default |
|---|---|---|---|
| Kimi | KIMI_MODEL=kimi-k2.6 |
KIMI_BASE_URL |
extra_body={"thinking": {"type": "enabled"}}; omit sampling overrides. |
| MiniMax | MINIMAX_MODEL=MiniMax-M2.7 |
MINIMAX_BASE_URL |
Native Interleaved Thinking; add extra_body={"reasoning_split": True} and preserve reasoning_details. |
| Qwen Plus | QWEN_PLUS_MODEL=qwen3.6-plus |
QWEN_BASE_URL |
extra_body={"enable_thinking": True, "preserve_thinking": True}; omit thinking_budget to avoid capping thinking. |
| Qwen Max Preview | QWEN_MAX_PREVIEW_MODEL=qwen3.6-max-preview |
QWEN_BASE_URL |
Same as Qwen Plus; Max Preview is preferred for hardest reasoning/coding tasks. |
| GLM | GLM_MODEL=glm-5.1 |
GLM_BASE_URL |
extra_body={"thinking": {"type": "enabled"}}; use a large max_tokens when cost allows. |
| DeepSeek Pro | DEEPSEEK_MODEL=deepseek-v4-pro |
DEEPSEEK_BASE_URL |
reasoning_effort="max" plus extra_body={"thinking": {"type": "enabled"}}. |
| DeepSeek Flash | DEEPSEEK_FLASH_MODEL=deepseek-v4-flash |
DEEPSEEK_BASE_URL |
Same as DeepSeek Pro. |
| Xiaomi MiMo Pro | XIAOMI_MIMO_PRO_MODEL=mimo-v2.5-pro |
XIAOMI_BASE_URL |
Native reasoning model; include extra_body={"thinking": {"type": "enabled"}} if the endpoint accepts it. |
| Xiaomi MiMo | XIAOMI_MIMO_MODEL=mimo-v2.5 |
XIAOMI_BASE_URL |
Same as Xiaomi MiMo Pro. |
| Doubao Pro | DOUBAO_SEED_2_PRO_CHAT_MODEL=ep-... |
DOUBAO_VOLCES_BASE_URL |
extra_body={"thinking": {"type": "enabled"}}; use your custom endpoint ID, not the public alias. |
| Doubao Lite | DOUBAO_SEED_2_LITE_CHAT_MODEL=ep-... |
DOUBAO_VOLCES_BASE_URL |
Same as Doubao Pro. |
| HY3 Preview | HY3_MODEL=hy3-preview |
HUNYUAN_BASE_URL |
HY3 is a fast/slow-thinking model; use the HY3 model directly and set extra_body={"enable_enhancement": True} for Hunyuan-compatible enhancement. |
Shared Helper
Use this helper for standalone API tests. It always returns the highest known reasoning settings for the provider.
def highest_reasoning_kwargs(target: str) -> dict:
if target == "kimi":
return {
"extra_body": {"thinking": {"type": "enabled"}},
"max_tokens": 32768,
}
if target.startswith("minimax"):
return {
"extra_body": {"reasoning_split": True},
"max_tokens": 65536,
}
if target.startswith("qwen"):
return {
"extra_body": {
"enable_thinking": True,
"preserve_thinking": True,
},
"max_tokens": 65536,
}
if target == "glm":
return {
"extra_body": {"thinking": {"type": "enabled"}},
"max_tokens": 131072,
"temperature": 1.0,
}
if target.startswith("deepseek"):
return {
"reasoning_effort": "max",
"extra_body": {"thinking": {"type": "enabled"}},
"max_tokens": 65536,
}
if target.startswith("xiaomi"):
return {
"extra_body": {"thinking": {"type": "enabled"}},
"max_tokens": 131072,
}
if target.startswith("doubao"):
return {
"extra_body": {"thinking": {"type": "enabled"}},
"max_tokens": 65536,
}
if target == "hy3":
return {
"extra_body": {"enable_enhancement": True},
"max_tokens": 65536,
"temperature": 0.9,
"top_p": 1.0,
}
return {}
Generic Call
import os
from openai import OpenAI
target = "deepseek-pro"
client = OpenAI(
api_key=os.environ["PROVIDER_API_KEY"],
base_url=os.environ["PROVIDER_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ["PROVIDER_MODEL"],
messages=[{"role": "user", "content": "Solve a hard reasoning problem."}],
**highest_reasoning_kwargs(target),
)
message = response.choices[0].message
print(message.content)
For model/harness work, preserve the full assistant message object when continuing a conversation, especially for MiniMax, Qwen, DeepSeek, Kimi, GLM, and tool-call runs. Stripping reasoning_content, reasoning_details, or provider-specific thinking fields can break later tool-call turns.
Docker Runner Apt Mirror
run_opencode_bench_docker_open_models.py stages task Dockerfiles under the run root instead of editing tasks/ directly. During staging it can add a small apt network block before the first task install step:
- writes
/etc/apt/apt.conf.d/99skillsbench-networkwith IPv4, timeout, and retry settings; - rewrites Ubuntu
archive.ubuntu.comandsecurity.ubuntu.comsources to a configured mirror when one is available; - records the resolved mirror and apt settings in
manifest.json.
The default mirror mode is auto. On this GCP VM it detects the region from /etc/resolv.conf and uses:
http://us-central1.gce.archive.ubuntu.com/ubuntu/
Useful flags:
--ubuntu-apt-mirror auto
--ubuntu-apt-mirror http://us-central1.gce.archive.ubuntu.com/ubuntu/
--ubuntu-apt-mirror none
--no-apt-force-ipv4
--apt-http-timeout-sec 30
--apt-retries 3
Tool-Call Format Notes
smoke_open_model_access.py --tool-call was run against all default Chat Completions targets using a single function tool:
{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
},
"required": ["location"]
}
}
}
All tested providers returned the same top-level OpenAI Chat Completions tool-call shape:
{
"choices": [
{
"message": {
"role": "assistant",
"content": "",
"tool_calls": [
{
"id": "...",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"location\": \"San Francisco, CA\"}"
}
}
]
},
"finish_reason": "tool_calls"
}
]
}
Observed results:
| Target | Tool-call shape | Arguments JSON | Reasoning/thinking location in tool-call response |
|---|---|---|---|
| Kimi | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| MiniMax | message.tool_calls[] |
Valid JSON string | Inline <think>...</think> in message.content by default |
| Qwen Plus | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| Qwen Max Preview | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| GLM | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| DeepSeek Pro | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| DeepSeek Flash | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| Xiaomi MiMo Pro | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| Xiaomi MiMo | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| Doubao Pro | message.tool_calls[] |
Valid JSON string | message.reasoning_content |
| Doubao Lite | message.tool_calls[] |
Valid JSON string | No extra reasoning field observed in this smoke run |
| HY3 Preview | message.tool_calls[] |
Valid JSON string | No extra reasoning field observed in this smoke run |
The harness should not assume content contains the useful assistant state during tool calls. Many reasoning models returned content="" with tool_calls plus hidden/sidecar reasoning fields. Store and replay the full assistant message object returned by the SDK, not a hand-built {role, content, tool_calls} subset.
Provider-specific caveats:
- MiniMax without
reasoning_split=Trueplaces thinking in visible<think>...</think>content. Withreasoning_split=True, preservereasoning_detailsif present. - Kimi, Qwen, GLM, DeepSeek, Xiaomi, and Doubao Pro may return
reasoning_content; preserve it across tool-call turns. - DeepSeek V4 thinking-mode tool calls are especially sensitive to preserving
reasoning_content; losing it can cause later turns to fail or drift. - Doubao Lite and HY3 used standard
tool_callsin the smoke run but did not expose sidecar reasoning fields in that response. Still preserve the full message, because behavior can vary by prompt and model revision. function.argumentswas always a JSON string in the tested targets. Parse it withjson.loads; do not expect a native object.- Whitespace in
function.argumentsdiffers by provider ({"location":"..."}vs{"location": "..."}), so compare parsed JSON, not raw strings.
Tool-call smoke command:
uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py \
--tool-call \
--timeout 25 \
--hard-timeout 35 \
--max-tokens 512 \
--concurrency 8 \
--json-out open_model_tool_call_smoke.json
Provider Examples
Kimi K2.6
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["KIMI_API_KEY"],
base_url=os.environ["KIMI_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ.get("KIMI_MODEL", "kimi-k2.6"),
messages=[{"role": "user", "content": "Think deeply and solve this."}],
extra_body={"thinking": {"type": "enabled"}},
max_tokens=32768,
)
print(response.choices[0].message.content)
Current account note: the Kimi key smoke-tested successfully against https://api.moonshot.ai/v1. The .cn endpoint returned 401 for this key.
MiniMax M2.7
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MINIMAX_API_KEY"],
base_url=os.environ["MINIMAX_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ.get("MINIMAX_MODEL", "MiniMax-M2.7"),
messages=[{"role": "user", "content": "Think through this coding task carefully."}],
extra_body={"reasoning_split": True},
max_tokens=65536,
)
message = response.choices[0].message
print(message.content)
For multi-turn or tool-call conversations, append the full message object back into history, including reasoning_details.
Qwen 3.6
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["QWEN_API_KEY"],
base_url=os.environ["QWEN_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ.get("QWEN_MAX_PREVIEW_MODEL", "qwen3.6-max-preview"),
messages=[{"role": "user", "content": "Solve this difficult agentic coding task."}],
extra_body={
"enable_thinking": True,
"preserve_thinking": True,
},
max_tokens=65536,
)
print(response.choices[0].message.content)
Do not set thinking_budget when you want the provider default maximum chain-of-thought budget. Set it only when you intentionally want to cap reasoning tokens.
GLM 5.1
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["GLM_API_KEY"],
base_url=os.environ["GLM_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ.get("GLM_MODEL", "glm-5.1"),
messages=[{"role": "user", "content": "Plan and solve this long-horizon task."}],
extra_body={"thinking": {"type": "enabled"}},
max_tokens=131072,
temperature=1.0,
)
print(response.choices[0].message.content)
DeepSeek V4
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEEPSEEK_API_KEY"],
base_url=os.environ["DEEPSEEK_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ.get("DEEPSEEK_MODEL", "deepseek-v4-pro"),
messages=[{"role": "user", "content": "Solve this hard task with maximum reasoning."}],
reasoning_effort="max",
extra_body={"thinking": {"type": "enabled"}},
max_tokens=65536,
)
print(response.choices[0].message.content)
For OpenCode/tool-call harnesses, preserve reasoning_content in assistant messages between turns.
Xiaomi MiMo V2.5
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["XIAOMI_API_KEY"],
base_url=os.environ["XIAOMI_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ.get("XIAOMI_MIMO_PRO_MODEL", "mimo-v2.5-pro"),
messages=[{"role": "user", "content": "Solve this long-horizon agent task."}],
extra_body={"thinking": {"type": "enabled"}},
max_tokens=131072,
)
print(response.choices[0].message.content)
The Token Plan endpoint smoke-tested successfully with mimo-v2.5-pro and mimo-v2.5. If thinking.enabled is rejected by the live endpoint, omit it; the endpoint still returned reasoning_content during smoke tests.
Doubao Seed 2.0
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DOUBAO_SEED_2_PRO_API_KEY"],
base_url=os.environ["DOUBAO_VOLCES_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ["DOUBAO_SEED_2_PRO_CHAT_MODEL"],
messages=[{"role": "user", "content": "Use deep thinking for this task."}],
extra_body={"thinking": {"type": "enabled"}},
max_tokens=65536,
)
print(response.choices[0].message.content)
Current account note: public model aliases returned InvalidEndpointOrModel.ModelIDAccessDisabled; use the custom ep-... endpoint IDs in .env.
HY3 Preview
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["HUNYUAN_API_KEY"],
base_url=os.environ["HUNYUAN_BASE_URL"],
)
response = client.chat.completions.create(
model=os.environ.get("HY3_MODEL", "hy3-preview"),
messages=[{"role": "user", "content": "Think carefully and solve this agent task."}],
extra_body={"enable_enhancement": True},
max_tokens=65536,
temperature=0.9,
top_p=1.0,
)
print(response.choices[0].message.content)
Current account endpoint:
https://api.lkeap.cloud.tencent.com/plan/v3
The Anthropic-compatible HY endpoint is intentionally omitted from this file.
Smoke Tests
Default API-only smoke test:
cd ~/skillsbench
uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py \
--timeout 14 \
--hard-timeout 18 \
--max-tokens 8 \
--concurrency 8 \
--json-out open_model_access_smoke.json
Single-model smoke:
uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py --only hy3
uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py --only qwen-max-preview
uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py --only deepseek-pro
Mapping Into OpenCode/BenchFlow
For one-off harness tests, run a small Docker-backed bench eval run through the
open-model runner:
uv run python experiments/scripts/open_model_scripts/run_opencode_bench_docker_open_models.py run \
--target deepseek-pro \
--task court-form-filling \
--condition without-skills \
--concurrency 1
The runner injects an OpenCode @ai-sdk/openai-compatible provider config for
each target, so the command path stays on BenchFlow's bench eval run Docker backend
instead of any alternate runner path.
Sources Checked
- Kimi K2.6 docs: https://platform.kimi.ai/docs/api/models-overview and https://platform.moonshot.ai/docs/guide/use-kimi-k2-thinking-model.en-US
- MiniMax M2.7 docs: https://platform.minimax.io/docs/guides/text-m2-function-call and https://platform.minimax.io/docs/api-reference/api-overview
- Qwen OpenAI-compatible API: https://docs.qwencloud.com/api-reference/chat/openai-chat and https://qwen.ai/blog?id=qwen3.6
- GLM 5.1 docs: https://docs.z.ai/guides/llm/glm-5.1
- DeepSeek thinking mode and agent config: https://api-docs.deepseek.com/guides/thinking_mode and https://api-docs.deepseek.com/quick_start/agent_integrations/deepcode
- Xiaomi MiMo public/API pages: https://mimo.mi.com/ and https://platform.xiaomimimo.com/token-plan
- Volcengine/Doubao thinking parameter reference: https://www.volcengine.com/docs/84313/1350013 and https://www.volcengine.com/docs/82379/1494384
- Tencent Hunyuan OpenAI-compatible API and HY3 announcement/model card: https://cloud.tencent.com/document/product/1729/111007, https://www.tencent.com/en-us/articles/2202320.html, https://huggingface.co/tencent/Hy3-preview/blob/main/README.md