# OpenAI-Compatible Model API Reference This reference is for SkillsBench/OpenCode smoke tests and harness runs. It intentionally covers only OpenAI-compatible API format. Do not put API keys in this file; keys live in `.env`. Default policy in this file: use the strongest available reasoning mode for every model. If a provider does not expose an explicit reasoning-strength knob, use the model's native thinking mode and preserve returned reasoning fields across multi-turn/tool-call turns. ## Current Smoke-Tested Models | Target | Model env | Base URL env | Highest reasoning default | | --- | --- | --- | --- | | Kimi | `KIMI_MODEL=kimi-k2.6` | `KIMI_BASE_URL` | `extra_body={"thinking": {"type": "enabled"}}`; omit sampling overrides. | | MiniMax | `MINIMAX_MODEL=MiniMax-M2.7` | `MINIMAX_BASE_URL` | Native Interleaved Thinking; add `extra_body={"reasoning_split": True}` and preserve `reasoning_details`. | | Qwen Plus | `QWEN_PLUS_MODEL=qwen3.6-plus` | `QWEN_BASE_URL` | `extra_body={"enable_thinking": True, "preserve_thinking": True}`; omit `thinking_budget` to avoid capping thinking. | | Qwen Max Preview | `QWEN_MAX_PREVIEW_MODEL=qwen3.6-max-preview` | `QWEN_BASE_URL` | Same as Qwen Plus; Max Preview is preferred for hardest reasoning/coding tasks. | | GLM | `GLM_MODEL=glm-5.1` | `GLM_BASE_URL` | `extra_body={"thinking": {"type": "enabled"}}`; use a large `max_tokens` when cost allows. | | DeepSeek Pro | `DEEPSEEK_MODEL=deepseek-v4-pro` | `DEEPSEEK_BASE_URL` | `reasoning_effort="max"` plus `extra_body={"thinking": {"type": "enabled"}}`. | | DeepSeek Flash | `DEEPSEEK_FLASH_MODEL=deepseek-v4-flash` | `DEEPSEEK_BASE_URL` | Same as DeepSeek Pro. | | Xiaomi MiMo Pro | `XIAOMI_MIMO_PRO_MODEL=mimo-v2.5-pro` | `XIAOMI_BASE_URL` | Native reasoning model; include `extra_body={"thinking": {"type": "enabled"}}` if the endpoint accepts it. | | Xiaomi MiMo | `XIAOMI_MIMO_MODEL=mimo-v2.5` | `XIAOMI_BASE_URL` | Same as Xiaomi MiMo Pro. | | Doubao Pro | `DOUBAO_SEED_2_PRO_CHAT_MODEL=ep-...` | `DOUBAO_VOLCES_BASE_URL` | `extra_body={"thinking": {"type": "enabled"}}`; use your custom endpoint ID, not the public alias. | | Doubao Lite | `DOUBAO_SEED_2_LITE_CHAT_MODEL=ep-...` | `DOUBAO_VOLCES_BASE_URL` | Same as Doubao Pro. | | HY3 Preview | `HY3_MODEL=hy3-preview` | `HUNYUAN_BASE_URL` | HY3 is a fast/slow-thinking model; use the HY3 model directly and set `extra_body={"enable_enhancement": True}` for Hunyuan-compatible enhancement. | ## Shared Helper Use this helper for standalone API tests. It always returns the highest known reasoning settings for the provider. ```python def highest_reasoning_kwargs(target: str) -> dict: if target == "kimi": return { "extra_body": {"thinking": {"type": "enabled"}}, "max_tokens": 32768, } if target.startswith("minimax"): return { "extra_body": {"reasoning_split": True}, "max_tokens": 65536, } if target.startswith("qwen"): return { "extra_body": { "enable_thinking": True, "preserve_thinking": True, }, "max_tokens": 65536, } if target == "glm": return { "extra_body": {"thinking": {"type": "enabled"}}, "max_tokens": 131072, "temperature": 1.0, } if target.startswith("deepseek"): return { "reasoning_effort": "max", "extra_body": {"thinking": {"type": "enabled"}}, "max_tokens": 65536, } if target.startswith("xiaomi"): return { "extra_body": {"thinking": {"type": "enabled"}}, "max_tokens": 131072, } if target.startswith("doubao"): return { "extra_body": {"thinking": {"type": "enabled"}}, "max_tokens": 65536, } if target == "hy3": return { "extra_body": {"enable_enhancement": True}, "max_tokens": 65536, "temperature": 0.9, "top_p": 1.0, } return {} ``` ## Generic Call ```python import os from openai import OpenAI target = "deepseek-pro" client = OpenAI( api_key=os.environ["PROVIDER_API_KEY"], base_url=os.environ["PROVIDER_BASE_URL"], ) response = client.chat.completions.create( model=os.environ["PROVIDER_MODEL"], messages=[{"role": "user", "content": "Solve a hard reasoning problem."}], **highest_reasoning_kwargs(target), ) message = response.choices[0].message print(message.content) ``` For model/harness work, preserve the full assistant message object when continuing a conversation, especially for MiniMax, Qwen, DeepSeek, Kimi, GLM, and tool-call runs. Stripping `reasoning_content`, `reasoning_details`, or provider-specific thinking fields can break later tool-call turns. ## Docker Runner Apt Mirror `run_opencode_bench_docker_open_models.py` stages task Dockerfiles under the run root instead of editing `tasks/` directly. During staging it can add a small apt network block before the first task install step: - writes `/etc/apt/apt.conf.d/99skillsbench-network` with IPv4, timeout, and retry settings; - rewrites Ubuntu `archive.ubuntu.com` and `security.ubuntu.com` sources to a configured mirror when one is available; - records the resolved mirror and apt settings in `manifest.json`. The default mirror mode is `auto`. On this GCP VM it detects the region from `/etc/resolv.conf` and uses: ```text http://us-central1.gce.archive.ubuntu.com/ubuntu/ ``` Useful flags: ```bash --ubuntu-apt-mirror auto --ubuntu-apt-mirror http://us-central1.gce.archive.ubuntu.com/ubuntu/ --ubuntu-apt-mirror none --no-apt-force-ipv4 --apt-http-timeout-sec 30 --apt-retries 3 ``` ## Tool-Call Format Notes `smoke_open_model_access.py --tool-call` was run against all default Chat Completions targets using a single function tool: ```json { "type": "function", "function": { "name": "get_weather", "parameters": { "type": "object", "properties": { "location": {"type": "string"} }, "required": ["location"] } } } ``` All tested providers returned the same top-level OpenAI Chat Completions tool-call shape: ```json { "choices": [ { "message": { "role": "assistant", "content": "", "tool_calls": [ { "id": "...", "type": "function", "function": { "name": "get_weather", "arguments": "{\"location\": \"San Francisco, CA\"}" } } ] }, "finish_reason": "tool_calls" } ] } ``` Observed results: | Target | Tool-call shape | Arguments JSON | Reasoning/thinking location in tool-call response | | --- | --- | --- | --- | | Kimi | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | MiniMax | `message.tool_calls[]` | Valid JSON string | Inline `...` in `message.content` by default | | Qwen Plus | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | Qwen Max Preview | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | GLM | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | DeepSeek Pro | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | DeepSeek Flash | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | Xiaomi MiMo Pro | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | Xiaomi MiMo | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | Doubao Pro | `message.tool_calls[]` | Valid JSON string | `message.reasoning_content` | | Doubao Lite | `message.tool_calls[]` | Valid JSON string | No extra reasoning field observed in this smoke run | | HY3 Preview | `message.tool_calls[]` | Valid JSON string | No extra reasoning field observed in this smoke run | The harness should not assume `content` contains the useful assistant state during tool calls. Many reasoning models returned `content=""` with `tool_calls` plus hidden/sidecar reasoning fields. Store and replay the full assistant message object returned by the SDK, not a hand-built `{role, content, tool_calls}` subset. Provider-specific caveats: - MiniMax without `reasoning_split=True` places thinking in visible `...` content. With `reasoning_split=True`, preserve `reasoning_details` if present. - Kimi, Qwen, GLM, DeepSeek, Xiaomi, and Doubao Pro may return `reasoning_content`; preserve it across tool-call turns. - DeepSeek V4 thinking-mode tool calls are especially sensitive to preserving `reasoning_content`; losing it can cause later turns to fail or drift. - Doubao Lite and HY3 used standard `tool_calls` in the smoke run but did not expose sidecar reasoning fields in that response. Still preserve the full message, because behavior can vary by prompt and model revision. - `function.arguments` was always a JSON string in the tested targets. Parse it with `json.loads`; do not expect a native object. - Whitespace in `function.arguments` differs by provider (`{"location":"..."}` vs `{"location": "..."}`), so compare parsed JSON, not raw strings. Tool-call smoke command: ```bash uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py \ --tool-call \ --timeout 25 \ --hard-timeout 35 \ --max-tokens 512 \ --concurrency 8 \ --json-out open_model_tool_call_smoke.json ``` ## Provider Examples ### Kimi K2.6 ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ["KIMI_API_KEY"], base_url=os.environ["KIMI_BASE_URL"], ) response = client.chat.completions.create( model=os.environ.get("KIMI_MODEL", "kimi-k2.6"), messages=[{"role": "user", "content": "Think deeply and solve this."}], extra_body={"thinking": {"type": "enabled"}}, max_tokens=32768, ) print(response.choices[0].message.content) ``` Current account note: the Kimi key smoke-tested successfully against `https://api.moonshot.ai/v1`. The `.cn` endpoint returned 401 for this key. ### MiniMax M2.7 ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ["MINIMAX_API_KEY"], base_url=os.environ["MINIMAX_BASE_URL"], ) response = client.chat.completions.create( model=os.environ.get("MINIMAX_MODEL", "MiniMax-M2.7"), messages=[{"role": "user", "content": "Think through this coding task carefully."}], extra_body={"reasoning_split": True}, max_tokens=65536, ) message = response.choices[0].message print(message.content) ``` For multi-turn or tool-call conversations, append the full `message` object back into history, including `reasoning_details`. ### Qwen 3.6 ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ["QWEN_API_KEY"], base_url=os.environ["QWEN_BASE_URL"], ) response = client.chat.completions.create( model=os.environ.get("QWEN_MAX_PREVIEW_MODEL", "qwen3.6-max-preview"), messages=[{"role": "user", "content": "Solve this difficult agentic coding task."}], extra_body={ "enable_thinking": True, "preserve_thinking": True, }, max_tokens=65536, ) print(response.choices[0].message.content) ``` Do not set `thinking_budget` when you want the provider default maximum chain-of-thought budget. Set it only when you intentionally want to cap reasoning tokens. ### GLM 5.1 ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ["GLM_API_KEY"], base_url=os.environ["GLM_BASE_URL"], ) response = client.chat.completions.create( model=os.environ.get("GLM_MODEL", "glm-5.1"), messages=[{"role": "user", "content": "Plan and solve this long-horizon task."}], extra_body={"thinking": {"type": "enabled"}}, max_tokens=131072, temperature=1.0, ) print(response.choices[0].message.content) ``` ### DeepSeek V4 ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ["DEEPSEEK_API_KEY"], base_url=os.environ["DEEPSEEK_BASE_URL"], ) response = client.chat.completions.create( model=os.environ.get("DEEPSEEK_MODEL", "deepseek-v4-pro"), messages=[{"role": "user", "content": "Solve this hard task with maximum reasoning."}], reasoning_effort="max", extra_body={"thinking": {"type": "enabled"}}, max_tokens=65536, ) print(response.choices[0].message.content) ``` For OpenCode/tool-call harnesses, preserve `reasoning_content` in assistant messages between turns. ### Xiaomi MiMo V2.5 ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ["XIAOMI_API_KEY"], base_url=os.environ["XIAOMI_BASE_URL"], ) response = client.chat.completions.create( model=os.environ.get("XIAOMI_MIMO_PRO_MODEL", "mimo-v2.5-pro"), messages=[{"role": "user", "content": "Solve this long-horizon agent task."}], extra_body={"thinking": {"type": "enabled"}}, max_tokens=131072, ) print(response.choices[0].message.content) ``` The Token Plan endpoint smoke-tested successfully with `mimo-v2.5-pro` and `mimo-v2.5`. If `thinking.enabled` is rejected by the live endpoint, omit it; the endpoint still returned `reasoning_content` during smoke tests. ### Doubao Seed 2.0 ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ["DOUBAO_SEED_2_PRO_API_KEY"], base_url=os.environ["DOUBAO_VOLCES_BASE_URL"], ) response = client.chat.completions.create( model=os.environ["DOUBAO_SEED_2_PRO_CHAT_MODEL"], messages=[{"role": "user", "content": "Use deep thinking for this task."}], extra_body={"thinking": {"type": "enabled"}}, max_tokens=65536, ) print(response.choices[0].message.content) ``` Current account note: public model aliases returned `InvalidEndpointOrModel.ModelIDAccessDisabled`; use the custom `ep-...` endpoint IDs in `.env`. ### HY3 Preview ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ["HUNYUAN_API_KEY"], base_url=os.environ["HUNYUAN_BASE_URL"], ) response = client.chat.completions.create( model=os.environ.get("HY3_MODEL", "hy3-preview"), messages=[{"role": "user", "content": "Think carefully and solve this agent task."}], extra_body={"enable_enhancement": True}, max_tokens=65536, temperature=0.9, top_p=1.0, ) print(response.choices[0].message.content) ``` Current account endpoint: ```text https://api.lkeap.cloud.tencent.com/plan/v3 ``` The Anthropic-compatible HY endpoint is intentionally omitted from this file. ## Smoke Tests Default API-only smoke test: ```bash cd ~/skillsbench uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py \ --timeout 14 \ --hard-timeout 18 \ --max-tokens 8 \ --concurrency 8 \ --json-out open_model_access_smoke.json ``` Single-model smoke: ```bash uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py --only hy3 uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py --only qwen-max-preview uv run python experiments/scripts/open_model_scripts/smoke_open_model_access.py --only deepseek-pro ``` ## Mapping Into OpenCode/BenchFlow For one-off harness tests, run a small Docker-backed `bench eval run` through the open-model runner: ```bash uv run python experiments/scripts/open_model_scripts/run_opencode_bench_docker_open_models.py run \ --target deepseek-pro \ --task court-form-filling \ --condition without-skills \ --concurrency 1 ``` The runner injects an OpenCode `@ai-sdk/openai-compatible` provider config for each target, so the command path stays on BenchFlow's `bench eval run` Docker backend instead of any alternate runner path. ## Sources Checked - Kimi K2.6 docs: https://platform.kimi.ai/docs/api/models-overview and https://platform.moonshot.ai/docs/guide/use-kimi-k2-thinking-model.en-US - MiniMax M2.7 docs: https://platform.minimax.io/docs/guides/text-m2-function-call and https://platform.minimax.io/docs/api-reference/api-overview - Qwen OpenAI-compatible API: https://docs.qwencloud.com/api-reference/chat/openai-chat and https://qwen.ai/blog?id=qwen3.6 - GLM 5.1 docs: https://docs.z.ai/guides/llm/glm-5.1 - DeepSeek thinking mode and agent config: https://api-docs.deepseek.com/guides/thinking_mode and https://api-docs.deepseek.com/quick_start/agent_integrations/deepcode - Xiaomi MiMo public/API pages: https://mimo.mi.com/ and https://platform.xiaomimimo.com/token-plan - Volcengine/Doubao thinking parameter reference: https://www.volcengine.com/docs/84313/1350013 and https://www.volcengine.com/docs/82379/1494384 - Tencent Hunyuan OpenAI-compatible API and HY3 announcement/model card: https://cloud.tencent.com/document/product/1729/111007, https://www.tencent.com/en-us/articles/2202320.html, https://huggingface.co/tencent/Hy3-preview/blob/main/README.md