Pre-release Harn is pre-1.0 — the language, standard library, and CLI may change between releases. See the release notes

Provider support recommendations

This page aggregates Harn's provider/model catalog, runtime capability rules, small curated notes, and optional harn eval coding-agent benchmark summaries. Regenerate with make gen-provider-support and verify with make check-provider-support.

No benchmark summary is baked into this checked-in page. To layer local empirical results, run harn provider catalog support --empirical .harn-runs/coding-agent-bench/latest/summary.json.

Credential variables#

Set one of a provider's variables to an API key, or to a harn-secret://namespace/name reference. Harn reads the variables left to right and uses the first one that is set.

A star marks the short list Harn names first in setup messages and pickers. The other providers work exactly the same way. To see which variables are already set on your machine, run harn doctor.

ProviderCatalog idCredential variables
★ AnthropicanthropicANTHROPIC_API_KEY
★ OpenAIopenaiOPENAI_API_KEY
★ Google GeminigeminiGEMINI_API_KEY or GOOGLE_API_KEY
★ OpenRouteropenrouterOPENROUTER_API_KEY
★ GroqgroqGROQ_API_KEY
★ DeepSeekdeepseekDEEPSEEK_API_KEY
★ Ollamaollamanone — runs without a key
AtlasatlasATLAS_API_KEY or ATLASCLOUD_API_KEY
Azure Openaiazure_openaiAZURE_OPENAI_API_KEY or AZURE_OPENAI_AD_TOKEN or AZURE_OPENAI_BEARER_TOKEN
BasetenbasetenBASETEN_API_KEY
Bedrockbedrockresolved by the platform credential chain
CerebrascerebrasCEREBRAS_API_KEY
Cloudflare Ai Gatewaycloudflare_ai_gatewayCLOUDFLARE_API_TOKEN or CLOUDFLARE_API_KEY
CoherecohereCOHERE_API_KEY
DashscopedashscopeDASHSCOPE_API_KEY
DeepinfradeepinfraDEEPINFRA_API_KEY or DEEPINFRA_TOKEN
FireworksfireworksFIREWORKS_API_KEY
FlexaiflexaiFLEXAI_API_KEY
FriendlifriendliFRIENDLI_API_KEY or FRIENDLI_TOKEN
Github Modelsgithub_modelsGITHUB_MODELS_TOKEN
HuggingfacehuggingfaceHF_TOKEN or HUGGINGFACE_API_KEY
HunyuanhunyuanHUNYUAN_API_KEY or TENCENT_HUNYUAN_API_KEY
HyperbolichyperbolicHYPERBOLIC_API_KEY
InceptioninceptionINCEPTION_API_KEY
Llamacppllamacppnone — runs without a key
Locallocalnone — runs without a key
MetametaMETA_API_KEY
MinimaxminimaxMINIMAX_API_KEY
MistralmistralMISTRAL_API_KEY
Mlxmlxnone — runs without a key
MoonshotmoonshotMOONSHOT_API_KEY or MOONSHOT_AI_API_KEY or KIMI_API_KEY
NebiusnebiusNEBIUS_API_KEY
NvidianvidiaNVIDIA_API_KEY or NIM_API_KEY
ParasailparasailPARASAIL_API_KEY
QianfanqianfanQIANFAN_API_KEY or BAIDU_QIANFAN_API_KEY
SambanovasambanovaSAMBANOVA_API_KEY
SiliconflowsiliconflowSILICONFLOW_API_KEY
Tgitginone — runs without a key
TogethertogetherTOGETHER_AI_API_KEY
Vercel AI Gatewayvercel_ai_gatewayAI_GATEWAY_API_KEY or VERCEL_AI_GATEWAY_API_KEY
VertexvertexVERTEX_AI_ACCESS_TOKEN or GOOGLE_OAUTH_ACCESS_TOKEN or GOOGLE_APPLICATION_CREDENTIALS
Vllmvllmnone — runs without a key
Volcengine Arkvolcengine_arkARK_API_KEY or VOLCENGINE_ARK_API_KEY or VOLCENGINE_API_KEY
XaixaiXAI_API_KEY
ZaizaiZAI_API_KEY or ZHIPU_API_KEY

Capability comparison#

ProviderEndpoint styleRecommended selectorTool modeNative toolsText toolsStructured outputReasoning knobsCacheBatchServing tiersUsage confidenceEmpirical
AnthropicAnthropic Messages APIhaikunativeyesyesnative / native_jsonenabledyesYes (50%)fast:premiumhighnot_recorded
AtlasOpenAI-compatible chat completionsatlastextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
Azure OpenaiOpenAI-compatible chat completionsazure_openai:gpt-*nativeyesyesnone / native_jsonnonenoYes (50%)noneprovider_defaultnot_recorded
BasetenOpenAI-compatible chat completionsbaseten:baseten/deepseek-ai/DeepSeek-V4-Flash-0731nativeyesyesnative / native_jsoneffort,reasoning_effortyesNononehighnot_recorded
BedrockAWS Bedrock Conversebedrock:anthropic.claude-sonnet-4-5-20250929-v1:0nativeyesyesnone / xml_taggednonenoYesnoneprovider_defaultnot_recorded
CerebrasOpenAI-compatible chat completionscerebras/gpt-oss-120bnativeyesyesnative / native_jsoneffort,reasoning_effortnoNononehighnot_recorded
Cloudflare Ai GatewayOpenAI-compatible chat completionscloudflare_ai_gatewaytextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
CohereOpenAI-compatible chat completionscohere:command-a-plus-05-2026nativeyesyesnative / native_jsonadaptivenoNononehighnot_recorded
DashscopeOpenAI-compatible chat completionsdashscope:dashscope/qwen3-coder-nextnativeyesyesnative / delimiteddisable_directive:/no_think,enabledyesNononehighnot_recorded
DeepinfraOpenAI-compatible chat completionsdeepinfra:deepinfra/Qwen/Qwen3-235B-A22B-Instruct-2507nativeyesyesnative / native_jsonnonenoNononehighnot_recorded
DeepSeekOpenAI-compatible chat completionsdeepseek:deepseek-v4-flashnativeyesyesnative / native_jsoneffort,enabled,reasoning_effortyesNononehighnot_recorded
FireworksOpenAI-compatible chat completionsfireworks:accounts/fireworks/models/gpt-oss-120btextnoyesnone / native_jsoneffort,reasoning_effortnoYes (50%)nonehighnot_recorded
FlexaiOpenAI-compatible chat completionsflexaitextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
FriendliOpenAI-compatible chat completionsfriendlitextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
Gemini APIGemini generateContentgemini:gemini-3.5-flash-litenativeyesyesnative / native_jsonadaptive,effort,enabled,reasoning_effortyesYes (50%)flex:discounted, priority:premiummediumnot_recorded
Github ModelsOpenAI-compatible chat completionsgithub_modelstextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
GroqOpenAI-compatible chat completionsgroq:qwen/qwen3.6-27bnativeyesyesnative / native_jsontoggleyesYes (50%)nonehighnot_recorded
Hugging Face Inference ProvidersOpenAI-compatible chat completions through the HF routerhuggingface-qwen3-codernativeyesyesnative / delimitednonenoNononemediumnot_recorded
HunyuanOpenAI-compatible chat completionshunyuantextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
HyperbolicOpenAI-compatible chat completionshyperbolictextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
InceptionOpenAI-compatible chat completionsinception:mercury-2nativeyesyesnative / native_jsoneffort,reasoning_effortnoNononehighnot_recorded
llama.cpp serverOpenAI-compatible llama-serverllamacpp-qwen3.6-q4nativeyesyesnative / delimiteddisable_directive:/no_think,enablednoNononemediumnot_recorded
OpenAI-compatible local serverOpenAI-compatible chat completionslocal-gemma4textyesyesnative / delimitedenablednoNononelownot_recorded
MetaOpenAI-compatible chat completionsmeta:muse-spark-1.2-contributornativeyesyesnative / native_jsonenabledyesNononehighnot_recorded
MinimaxOpenAI-compatible chat completionsminimax:MiniMax-M2.5-highspeednativeyesyesdelimited / delimitedenabledyesNononehighnot_recorded
Mistral via OpenRouterOpenAI-compatible chat completions through OpenRouteropenrouter:mistralai/mistral-small-2603nativeyesyesnative / native_jsonnoneyesNononemediumnot_recorded
MLX OpenAI-compatible serverOpenAI-compatible MLX servermlx-qwen3.6nativeyesyesnative / delimiteddisable_directive:/no_think,enablednoNononemediumnot_recorded
MoonshotOpenAI-compatible chat completionsmoonshot:moonshot/kimi-k2.6nativeyesyesnative / native_jsonenabledyesNononehighnot_recorded
NebiusOpenAI-compatible chat completionsnebiustextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
NvidiaOpenAI-compatible chat completionsnvidia:nvidia/minimax-m3nativeyesyesdelimited / delimitedadaptiveyesNononehighnot_recorded
OllamaOllama native chat APIdevstral-small-2textnoyesformat_kw / delimitednonenoNononehighnot_recorded
OpenAIOpenAI chat completions / Responses-compatible routesopenai:gpt-5.4-mininativeyesyesnative / native_jsoneffort,reasoning_effort,reasoning_noneyesYes (50%)fast:premium, flex:discountedhighnot_recorded
OpenRouterOpenAI-compatible chat completionsopenrouter:google/gemini-2.5-flashnativeyesyesnative / native_jsoneffort,enabled,reasoning_effortyesNononehighnot_recorded
ParasailOpenAI-compatible chat completionsparasailtextnoyesnone / nonenonenoYesnoneprovider_defaultnot_recorded
QianfanOpenAI-compatible chat completionsqianfantextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
SambanovaOpenAI-compatible chat completionssambanova:sambanova/gpt-oss-120btextnoyesnative / native_jsoneffort,reasoning_effortnoNononehighnot_recorded
SiliconflowOpenAI-compatible chat completionssiliconflowtextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
TgiOpenAI-compatible chat completionstgitextnoyesnone / nonenonenoNononelocal_zero_costnot_recorded
TogetherOpenAI-compatible chat completionstogether:openai/gpt-oss-20bnativeyesyesnative / native_jsoneffort,reasoning_effortnoYes (50%)nonehighnot_recorded
Vercel AI GatewayOpenAI-compatible chat completionsvercel_ai_gateway:vercel/openai/gpt-5.4-nanonativeyesyesnative / native_jsoneffort,reasoning_effort,reasoning_noneyesNononehighnot_recorded
VertexGemini generateContentvertex:vertex/gemini-2.5-flashnativeyesyesnone / native_jsonnonenoNononeprovider_defaultnot_recorded
VllmOpenAI-compatible chat completionsvllmtextnoyesnone / nonenonenoNononelocal_zero_costnot_recorded
Volcengine ArkOpenAI-compatible chat completionsvolcengine_arktextnoyesnone / nonenonenoNononeprovider_defaultnot_recorded
XaiOpenAI-compatible chat completionsxai:grok-build-0.1nativeyesyesnative / native_jsonadaptiveyesYesnonehighnot_recorded
ZaiOpenAI-compatible chat completionszai:glm-4.6nativeyesyesnative / native_jsonenabledyesNononehighnot_recorded

Anthropic#

  • catalog provider: anthropic
  • recommended route: haiku (claude-haiku-4-5-20251001)
  • endpoint style: Anthropic Messages API
  • recommended Harn options:
provider = "anthropic"
model = "haiku"
tool_format = "native"
structured_output_mode = "xml_tagged"

Notes:

  • Native tools, prompt caching, file upload, and XML-oriented scaffolding are first-class in Harn capability data.
  • Claude 4.7 rows use adaptive thinking; older Claude 4 rows use explicit thinking controls where supported.

Caveats:

  • Strict JSON output is modeled as tool-use or XML-tagged output rather than OpenAI-style response_json_schema.

MCP notes:

  • No provider-specific MCP connector is required; Harn exposes MCP tools through the runtime tool registry.

Cerebras#

  • catalog provider: cerebras
  • recommended route: cerebras/gpt-oss-120b (gpt-oss-120b)
  • endpoint style: OpenAI-compatible chat completions
  • recommended Harn options:
provider = "cerebras"
model = "gpt-oss-120b"
tool_format = "native"
structured_output_mode = "native_json"

Notes:

  • Harn catalogs Cerebras public serverless rows separately from dedicated-endpoint weights so clients do not present unprovisioned enterprise endpoints as one-click routes.
  • Use the slash-prefixed selector form (cerebras/<model>) when a single string must carry both provider and model identity; Harn strips the prefix before sending the provider-native model id.

Caveats:

  • Preview rows such as zai-glm-4.7 may be discontinued by Cerebras on short notice; pin public production workloads to non-preview rows unless the caller opts into preview behavior.

MCP notes:

  • MCP tools are normalized through Harn tool definitions before they become OpenAI-compatible Cerebras tool schemas.

Gemini API#

  • catalog provider: gemini
  • recommended route: gemini:gemini-3.5-flash-lite (gemini-3.5-flash-lite)
  • endpoint style: Gemini generateContent
  • recommended Harn options:
provider = "gemini"
model = "gemini-3.5-flash-lite"
tool_format = "native"
structured_output_mode = "native_json"

Notes:

  • Harn lowers native tools to Gemini function declarations and maps function responses back into the transcript.
  • Gemini response usage maps cached-content token counts when the provider reports them.

Caveats:

  • Harn does not create Gemini context-cache resources yet; cache accounting is therefore observational.

MCP notes:

  • MCP tools are regular Harn runtime tools before they become Gemini function declarations.

Hugging Face Inference Providers#

  • catalog provider: huggingface
  • recommended route: huggingface-qwen3-coder (Qwen/Qwen3-Coder-480B-A35B-Instruct)
  • endpoint style: OpenAI-compatible chat completions through the HF router
  • recommended Harn options:
provider = "huggingface"
model = "huggingface-qwen3-coder"
tool_format = "native"

Notes:

  • The Hugging Face router uses the OpenAI-compatible chat completions API and supports tools and streaming for chat-completion models.
  • Qwen3-Coder 480B A35B is the recommended HF router coding row because its model card publishes native 262K context, agentic coding focus, and a function-call-oriented format.

Caveats:

  • Router availability, latency, and pricing depend on the selected upstream provider; run readiness and tool probes before promoting a route into production defaults.
  • Do not infer reasoning controls from other Qwen rows: this Qwen3-Coder model card says the model is non-thinking.

MCP notes:

  • MCP tools are rendered as OpenAI-compatible tool definitions on this route.

Inception#

  • catalog provider: inception
  • recommended route: inception:mercury-2 (mercury-2)
  • endpoint style: OpenAI-compatible chat completions
  • recommended Harn options:
provider = "inception"
model = "mercury-2"
tool_format = "native"
structured_output_mode = "native_json"

Caveats:

  • Inception documents OpenAI-compatible tool use for Mercury 2; no Harn parity probe has run yet.

llama.cpp server#

  • catalog provider: llamacpp
  • recommended route: llamacpp-qwen3.6-q4 (qwen3.6-35b-a3b-ud-q4-k-xl)
  • endpoint style: OpenAI-compatible llama-server
  • recommended Harn options:
provider = "llamacpp"
model = "llamacpp-qwen3.6-q4"
tool_format = "native"
thinking = "off"

Notes:

  • llama.cpp gets its own provider so Harn can model Qwen chat-template and thinking behavior separately from generic local OpenAI-compatible servers.

Caveats:

  • Run both provider readiness and tool probes after changing GGUF, context, KV-cache, or chat-template settings.
  • 2026-08-19 CUDA receipt, llama-server b9994-14d3ba45f on an RTX 5090 host, Qwen3.6-35B-A3B-UD-Q4_K_XL, n_ctx 65536, chat template sha256 55d4931433fe, harn 0.10.105. Independent replication of the 2026-08-18 Metal receipt (llama-server b10360-48d22e295), which superseded the 2026-07-19 #5162 CUDA receipt. That reversal is now ATTRIBUTED, where the Metal receipt could only say hardware, runtime build, and revision all differed. This arm holds hardware at CUDA and still finds native working, so it was not hardware; and the #5162 arm was never a native arm at all. The mechanism is the one this row already states about the Metal arm: with native_tools = false no tools array reaches the wire and the arm cannot be measured. #5162's surviving invocation and stderr show it passed a tool_format override, which clears the FORMAT gate only, and no capability override. A static read of the July revision 38067db78 confirms the two gates are independent: the format gate is cleared by override_reason, a separate gate rewrites native to json on a text_only route with no override argument, and the tools array is attached only when the RESOLVED format is native, so no branch at that revision could put tools on the wire for this row. #5162 therefore measured the half-cleared state and never measured native. Stated at its true strength: the invocation and its output were read and the July-era code has no path that could have served a tools array, but no July request body survives to be read, and an unrecorded capability override in that lane cannot be excluded by observation, only judged unlikely because such an override is a deliberate act that the Metal receipt's author documented explicitly when they used it. Six coding-agent fixtures, two replicates, three forced formats, 36 runs, pre-registered validity gate held with skipped_runs = 0 in every arm. Completion: native 10/12, tagged text 10/12, fenced JSON 8/12, cell-for-cell identical to the Metal receipt. Six of the eight failures are no-tool-diagnosis failing 0/2 in ALL THREE arms; that is the grader defect harn#6843, which is fixed on the Metal receipt's branch and NOT on this one, so it reproduces as expected and is not a capability read. On the ten measurable cells: native 10/10, text 10/10, json 8/10. NO CLAIM of a native completion advantage is made: 12 runs per arm cannot separate 10/12 from 8/12, and the generated overlay's preferred_tool_format = native is the classifier's TIEBREAK DEFAULT rather than an empirical result, while its confidence = high means only that sample size and replicate count cleared a floor. What the receipt does establish is that the native channel returns parseable tool calls on this route under CUDA: 57 requests carried a tools array of up to 6 tools and 63 native tool calls came back on the structured channel, counted twice by independent paths (an intercepting proxy on the wire and the per-run rows) which agree exactly, and not read from a parsed_tool_calls field. CONSTRAINED DECODING IS NOT INVOKED: zero of the 57 tool-bearing requests carried grammar, json_schema, response_format, or any guided_* field. All 29 response_format requests carried NO tools and are repair/output-contract turns, reproducing on CUDA the correction the Metal receipt made about its own arm; the arm is therefore native over a plain tools channel with the repair-turn contract active. THROUGHPUT, which neither prior receipt measured: excluding 22 of 225 degenerate sub-20-token completions whose per-second rates are division artifacts, tool-bearing requests ran a median 247.1 tok/s (p10 244.8, p90 248.8, n=55) against 247.4 tok/s (p10 233.9, p90 250.0, n=148) without tools. The 30-40% constrained-decoding tax seen on grammar-bearing routes does not appear here, consistent with no grammar being applied. Prompt cost measured on the wire: native median 2444 system-prompt characters against 8473 for the text channel, which must teach its dialect every turn. Secondary, over 12 runs per arm: wall 41.0s native / 35.9s text / 49.4s json; iterations 57 / 66 / 68; tool calls 63 / 56 / 52; REJECTED tool calls 8 / 4 / 3, the most on native. BLIND SPOTS: one quant, one host, one llama.cpp build, where the July sweep covered three quants; two replicates is the classifier's floor and not convergence grade; the 29 repair turns cannot be attributed to particular arms by request shape, only shown to carry no tools; the server held a warm 21k-token prompt cache at the start and that was not controlled for, though the rig interleaves formats; and the #5162 attribution above rests on that run's invocation, its stderr, and a static read of its revision, NOT on any surviving request body, because none does. server_parser = none.

Local setup:

  • Run harn models install local-qwen3.6-gguf for the recommended download and launch commands.

OpenAI-compatible local server#

  • catalog provider: local
  • recommended route: local-gemma4 (gemma-4-26b-a4b-it)
  • endpoint style: OpenAI-compatible chat completions
  • recommended Harn options:
provider = "local"
model = "local-gemma4"
tool_format = "text"

Notes:

  • Use this generic provider when a local server speaks OpenAI chat completions but does not need a provider-specific quirk profile.

Caveats:

  • Prefer llamacpp or mlx when those runtimes are known, because their capability rows can encode template-specific behavior.

MCP notes:

  • MCP tools are exposed as OpenAI-compatible tool definitions unless the route is configured to prefer Harn text tools.

Local setup:

  • Set LOCAL_LLM_BASE_URL and either LOCAL_LLM_MODEL or an explicit Harn model selector, then run harn provider ready local.

Mistral via OpenRouter#

  • catalog provider: openrouter
  • recommended route: openrouter:mistralai/mistral-small-2603 (mistralai/mistral-small-2603)
  • endpoint style: OpenAI-compatible chat completions through OpenRouter
  • recommended Harn options:
provider = "openrouter"
model = "mistralai/mistral-small-2603"
tool_format = "native"

Notes:

  • Harn catalogs hosted Mistral routes through OpenRouter today, so endpoint and auth behavior are OpenAI-compatible.
  • Use this row for Mistral-family recommendation surfaces until a direct Mistral provider is cataloged.

Caveats:

  • Provider-native behavior depends on the OpenRouter model route; run the coding-agent benchmark before promoting it to a default for critical harnesses.

MCP notes:

  • MCP tools are rendered as OpenAI-compatible tool definitions on this route.

MLX OpenAI-compatible server#

  • catalog provider: mlx
  • recommended route: mlx-qwen3.6 (unsloth/Qwen3.6-35B-A3B-UD-MLX-4bit)
  • endpoint style: OpenAI-compatible MLX server
  • recommended Harn options:
provider = "mlx"
model = "mlx-qwen3.6"
tool_format = "native"

Notes:

  • MLX routes use native tools only after the served identity and tool probe match the cataloged model.

Caveats:

  • mlx_lm.server flags vary by release; launch first, then verify with harn provider ready mlx.

Local setup:

  • Run harn models install mlx-qwen3.6 for the venv, download, launch, and verification commands.

Ollama#

  • catalog provider: ollama
  • recommended route: devstral-small-2 (devstral-small-2:24b)
  • endpoint style: Ollama native chat API
  • recommended Harn options:
provider = "ollama"
model = "devstral-small-2"
tool_format = "text"
thinking = "off"

Notes:

  • Local Ollama model quality varies by template and quantization; Harn defaults known fragile routes to the text-tool contract.
  • Use harn provider tool-probe receipts to promote aliases from unknown to native/text/disabled on a machine.

Caveats:

  • Some Ollama native tool parsers reject otherwise valid text-mode model output; the capability table records those routes as text-only.

MCP notes:

  • MCP tools are local Harn tools; prefer the text tool contract unless a probe proves native calls work for the installed model.

Local setup:

  • Run harn models install devstral-small-2, then verify with harn provider ready ollama --model <model>. For local qwen3.x, use the llamacpp provider (e.g. local-qwen3.6) — Ollama's qwen3.5-family tool-call parser 500s on text-tool output.

OpenAI#

  • catalog provider: openai
  • recommended route: openai:gpt-5.4-mini (gpt-5.4-mini)
  • endpoint style: OpenAI chat completions / Responses-compatible routes
  • recommended Harn options:
provider = "openai"
model = "gpt-5.4-mini"
tool_format = "native"
structured_output_mode = "native_json"

Notes:

  • OpenAI-family routes default to native tool calls and native JSON structured output when the model row supports tools.
  • Reasoning models use developer-role instructions and reasoning-summary transcript projection where the capability row declares it.

Caveats:

  • Use explicit reasoning effort only on reasoning rows; non-reasoning chat models should keep thinking disabled.

MCP notes:

  • Hosted MCP behavior is normalized through Harn tool definitions; provider-side hosted tools remain a separate provider feature.

Sambanova#

  • catalog provider: sambanova
  • recommended route: sambanova:sambanova/gpt-oss-120b (sambanova/gpt-oss-120b)
  • endpoint style: OpenAI-compatible chat completions
  • recommended Harn options:
provider = "sambanova"
model = "sambanova/gpt-oss-120b"
tool_format = "text"
structured_output_mode = "native_json"

Caveats:

  • 2026-06-24 Harn agent-loop (gpt-oss-120b, zig-feat, tool grounding present): SambaNova native ended with a provider/tool-protocol failure (Harmony empty tool_calls / reasoning-channel-only class). Text/heredoc is the clean pay-per-token channel. See vLLM #22578/#44216, SGLang #8976/#10738, openai/harmony #68.