Provider detection
FIM One uses LiteLLM as a universal adapter. The_resolve_litellm_model() function in core/model/openai_compatible.py maps the user’s LLM_BASE_URL + LLM_MODEL to a LiteLLM model identifier with a provider prefix. The prefix determines how LiteLLM routes the request — native API protocol (Anthropic Messages API, Gemini, etc.) or generic OpenAI-compatible /v1/chat/completions.
Resolution order:
- Explicit provider (from DB
ModelConfig.providerfield) — highest priority. If the provider matches a known domain in the URL, noapi_baseis returned (LiteLLM routes natively). Otherwise,api_baseis set to the relay URL. - Domain match against
KNOWN_DOMAINS— official API endpoints are recognized by hostname. - URL path hint against
PATH_PROVIDER_HINTS— common on relay platforms like UniAPI where/claudeor/anthropicin the path indicates the upstream protocol. - Fallback —
openai/prefix (generic OpenAI-compatible).
When the provider prefix is a native protocol (anthropic, gemini, etc.) and the URL is not the official endpoint, LiteLLM uses the native protocol but sends requests to the relay’s
api_base. This means provider-specific behaviors — including the Bedrock prefill issue described below — apply regardless of whether the request goes to the official API or through a relay.
tool_choice — the four modes
Thetool_choice parameter is standardized via the OpenAI format. LiteLLM translates it to each provider’s native protocol before sending the request.
The distinction between
"auto" and forced ({"type":"function",...}) is the crux of every compatibility issue in FIM One. These two modes are used by completely different subsystems with different requirements.
Where tool_choice is used
Two subsystems usetool_choice, and they use it in fundamentally different ways.
ReAct engine — tool_choice=“auto”
The ReAct loop needs the model to decide each iteration: call a tool, or give a final answer. Only"auto" makes sense here — the model freely chooses between producing tool_calls or text content. This is compatible with all providers, all models, and all modes including extended thinking.
The ReAct engine uses native function calling (_run_native) when abilities["tool_call"] = True, falling back to JSON-in-content mode (_run_json) otherwise. Both modes use "auto" — the difference is whether tools are passed via the tools parameter or described in the system prompt. See ReAct Engine — Dual-mode execution for details.
structured_llm_call — tool_choice=forced
One-shot structured extraction (schema annotation, DAG planning, plan analysis). Forces the model to call a specific virtual function, guaranteeing structured JSON output. This is the call site that triggers provider-specific errors.structured_llm_call implements a 3-level degradation chain. The three rungs are named after the code, and their numbering is not the same as the four guarantee tiers described in the section below:
The critical design difference: structured_llm_call’s fallback is runtime — it dynamically tries each level and catches exceptions to fall through. The ReAct engine’s mode selection is build-time — it checks _native_mode_active once at the start and commits to one mode for the entire loop. This means structured_llm_call can recover from provider-specific 400 errors transparently, while ReAct relies on the mode being correctly chosen upfront.
Structured output: the four guarantee tiers
“Structured output” is one phrase covering four different mechanisms. They all produce JSON and they do not offer the same guarantee. What separates them is where the constraint lives: inside the decoder, which cannot emit a token the schema forbids, or inside the prompt, which can only ask.
The mechanism behind T1 and T2 is constrained decoding. The provider compiles the schema into a grammar and, at each decoding step, masks out every token that would make the output unparseable against it. A required field cannot be skipped because the token that would close the object is not in the allowed set until that field has been emitted. This is a different kind of claim from “the model was asked nicely and complied,” and it is why T1/T2 hold at temperature 1 and on small models, where prompt compliance does not.
T1 and T2 are not interchangeable
They give the same guarantee through different channels, and the channel matters in three ways. The turn shape differs. T1 returns a tool-call turn; the model has decided to act. T2 returns an ordinary assistant message; the model has answered in a shape. When the caller wants one extraction and not an agent loop, T2 says so directly instead of inventing a virtual function for the model to “call”. Forcing T1 collides with thinking. Getting a schema-gated payload out of T1 usually means forcing the choice, eitherrequired or a named function. Several providers reject a forced tool choice while extended thinking is active. Table B below records which ones. T2 has no such conflict: it is a response constraint, not a tool constraint, so it is the only schema-gated path that survives thinking being on. For a deployment standardised on reasoning models, that is the practical argument for T2 rather than a stylistic one.
T1 costs a round of schema plumbing that T2 does not. A virtual function needs a name, a description, and a decision about whether the model may decline to call it.
What T2 costs
Constrained decoding is not free and the schema you already have is usually not the schema it accepts.- A schema subset. OpenAI’s strict mode wants
additionalProperties: falseon every object and every property listed inrequired; optionality is expressed as a union withnullrather than by omission fromrequired. Root must be an object. There are caps on total properties and nesting depth. Most hand-written schemas need editing before they are eligible. - Grammar compilation. The first request carrying a new schema pays a one-off compile latency on the provider side. Reusing a stable schema amortises it; generating a fresh schema per request does not.
- A new failure surface. A model that will not answer returns a refusal rather than a schema-shaped object, so the caller needs a branch for it.
What no tier guarantees
Every tier constrains form. None constrains truth. A T2 response can be perfectly schema-valid and factually wrong, and an enum-constrained field will return one of the allowed values even when none of them is correct, because the decoder’s job is to keep the output inside the grammar rather than to know the answer. Schema gating removes parse failures and field-shape bugs. It does not remove the need to check what the value says.Where FIM One sits
Five consequences follow, and they are the honest limitations of the current design.
- The chain has no schema-gated rung to fall back to. When Level 1 fails, the next rung down is
json_object, which guarantees only that the text parses. There is no intermediate step that still enforces fields. - Level 1’s guarantee is weaker than its name suggests. Without
strict, native function calling is the pre-2024 non-strict variety: the model usually honours the schema and is not prevented from omitting a required field, inventing a key, or returning a string where a number was declared. - On
anthropic/routes, Level 2 is not really T3. The Anthropic Messages API has noresponse_format, so LiteLLM emulates JSON mode by injecting an assistant prefill. That is a prompt-layer device, which puts the effective guarantee between T4 and T3 rather than at T3. The Bedrock prefill trap documented below is this emulation failing loudly; the quieter cost is that the guarantee was never the one the parameter name implies. - Nothing validates the result against the schema afterwards.
jsonschemais not a dependency. Whatever checking happens is per-call-site, inside the optionalparse_fn, and its rigour varies by call site. - Failure is absorbed rather than surfaced. Nearly every caller passes
default_value, so an exhausted chain returns a plausible-looking object instead of raising.StructuredCallResult.level_usedrecords which rung produced the value, but no call site reads it, and on thedefault_valuepath it reportsplain_textfor a chain that in fact produced nothing.
structured_llm_call logs one line per completed call at INFO, emitted only after the data has survived parse_fn, so it names a rung that genuinely succeeded:
WARNING with level=none outcome=default_value. Grepping these two lines over a day of traffic is the only way to tell whether the missing T2 rung costs anything in a given deployment, since the default_value design means the symptom is a mediocre answer rather than an error.
Provider support for each tier
This table documents upstream APIs, not FIM One code paths. Every other table on this page names the function that implements the behaviour. This one cannot, because FIM One uses no T2 feature from any provider and sets
strict on none of them. It is here so that a decision to implement T2 starts from what is actually available. Vendor capability moves quickly and support often lands on some models in a family before others; re-check the vendor’s own documentation before relying on a row.
Three patterns are worth naming, because they explain the shape of the table rather than just restating it.
T3 is universal, T2 is not. Every provider FIM One supports offers
json_object, and roughly half offer a schema tier. A portable structured-output path therefore has to be built on T3 or below, which is the reason FIM One’s chain looks the way it does. Adding T2 means adding a per-model capability flag, not a global switch.
Always-on thinking tends to close the T1 door. GLM, Kimi’s thinking models, and deepseek-reasoner all restrict or reject a forced tool choice, and Anthropic rejects it while thinking is active. MiniMax is the counterexample. Where the door is closed, T2 is the only remaining schema-gated option, which is why the missing rung matters more on a Chinese-model deployment than on an OpenAI one.
Self-hosting inverts the usual ordering. Constrained decoding is a property of the serving stack, so it is available on a local checkpoint whose instruction-following is far weaker than a frontier model’s. The models that need schema gating most are the ones most able to get it.
The Bedrock prefill trap
Whenresponse_format={"type":"json_object"} is passed for a model resolved with the anthropic/ prefix, LiteLLM internally injects an assistant prefill message to simulate JSON mode. The Anthropic Messages API has no native response_format parameter, so LiteLLM approximates it by prepending an opening brace as assistant content:
role: "assistant" — they call this “assistant message prefill” and throw:
- The model is resolved with the
anthropic/prefix (via domain match or URL path hint). response_format={"type":"json_object"}is passed (the json_mode code path instructured_llm_call).- The actual backend is AWS Bedrock (which rejects prefill).
json_mode_enabled flag described below eliminates the wasted Level 2 call.
The fix: json_mode_enabled
A per-modeljson_mode_enabled flag controls whether Level 2 (json_mode) is ever attempted:
- DB-configured models: toggle in Admin → Models → Advanced settings. The flag is stored on
ModelProviderModel.json_mode_enabled(defaultTRUE). - ENV-configured models: set
LLM_JSON_MODE_ENABLED=falsein your environment. - Effect: when disabled,
abilities["json_mode"]returnsFalse→response_formatis never passed → no prefill → Bedrock works. The degradation chain becomesnative_fc → plain_text, skipping the doomed json_mode call entirely. - Cost of the skip: the model still returns valid JSON in practice, because the system prompt asks for it and
extract_json()parses free-form content reliably on modern models. What is lost is the guarantee rather than the output: the chain now ends at T4, where nothing but the prompt constrains the result. On ananthropic/route that loss is smaller than it looks, since the emulated JSON mode was never a decoder-level guarantee either.
Thinking models + forced tool_choice
Several providers reject a forcedtool_choice while extended thinking is active, on the grounds that pinning a specific function call contradicts the model’s freedom to reason first:
structured_llm_call resolves the conflict on its own by passing reasoning_effort=None on the native-FC level, which turns thinking off for that one call (structured.py::_call_llm). Structured output needs schema compliance, not deep reasoning, so disabling thinking there is both correct and cheaper.
Where thinking cannot be switched off through the API, native_fc fails with a 400 on every structured call and costs roughly ten seconds before the chain falls through to json_mode. Kimi is the common case: with thinking on only auto is supported, and a forced tool choice requires turning thinking off, which Moonshot exposes only through the model id (kimi-k2 has it off, kimi-k2.5 and kimi-k2-thinking have it on). FIM One has no parameter that flips it, so the remedy is the tool_choice_enabled flag below.
The fix: tool_choice_enabled
A per-modeltool_choice_enabled flag controls whether Level 1 (native_fc) is ever attempted:
- DB-configured models: toggle in Admin → Models → Advanced → “Native Function Calling”. The flag is stored on
ModelProviderModel.tool_choice_enabled(defaultTRUE). - ENV-configured models: set
LLM_TOOL_CHOICE_ENABLED=falsein your environment. - Effect: when disabled,
abilities["tool_choice"]returnsFalse→ the degradation chain starts from Level 2 (json_mode) or Level 3 (plain_text), skipping native_fc entirely. This eliminates the ~10s penalty per structured call for incompatible models. - ReAct agent unaffected:
tool_choice_enabledonly controls forced tool selection instructured_llm_call. The ReAct engine usestool_choice="auto"(model freely decides), which works with all models regardless of this setting.
tool_choice_enabled and tool_call are separate ability flags. tool_call (always True for OpenAICompatibleLLM) gates whether tools are passed to the model at all — disabling it would break the ReAct agent. tool_choice only gates whether forced tool selection is attempted for structured output extraction.tool_choice="auto" is unaffected by thinking mode. The ReAct engine uses "auto" exclusively, so agent execution works with thinking enabled.
Provider migration note: Some third-party relays silently drop unsupported parameters like
reasoning_effort (drop_params=True), so thinking is never activated even when configured. When migrating to a provider that properly supports thinking (Bedrock, direct Anthropic API), the reasoning_effort=None in native_fc ensures consistent behavior. No user action is needed — structured output works identically across all providers.Provider Capability Matrix
This section is the authoritative record of what each provider supports and what FIM One does about it. Every row names the function that implements the behaviour, so any claim here can be checked against the code. Other pages link here instead of repeating the data; when the code changes, this section changes with it. A row describes a provider’s protocol, not a single model. Where models inside one family differ (DeepSeek chat against reasoner, Kimi with thinking on against off), the cell says so.Table A: Protocol routing
How a configuredbase_url plus model becomes a LiteLLM call, and what happens when the first choice of interface is unavailable.
How GPT-5.x picks a protocol.
FIM_GPT5_RESPONSES_MODE selects it: native (the default) talks /v1/responses directly through litellm.aresponses, bridge uses LiteLLM’s chat-completions translation, and off forces plain chat completions. The native path exists because the bridge is lossy in the one place that matters: it discards the reasoning items, so a GPT-5.x agent re-derives its chain of thought on every tool round. Talking the protocol directly lets those items be replayed. A call that explicitly passes reasoning_effort=None, which is what structured_llm_call and the finish-signal probes do, stays on chat completions, because a call that wants no thinking has no reasoning state to preserve. The exception is models that cannot switch reasoning off (gpt-6.1-*, gpt-6-astra): their chat completions reject function tools at any reasoning_effort, so those calls stay on /v1/responses and run at the lowest effort, low. For the same reason off, or an endpoint without /v1/responses, leaves these models without tool calls.
Two properties of that native request are load-bearing and easy to get wrong:
store=falsekeeps the conversation stateless upstream, andinclude=["reasoning.encrypted_content"]asks for the encrypted payload to be returned. Without the include, the reasoning items arrive empty and replay silently becomes a no-op.- A replayed reasoning item must have its server-side
idstripped. Withstore=falsenothing is persisted upstream, so echoing the id back earnsItem with id 'rs_...' not found. Items are not persisted when 'store' is set to false. Theencrypted_contentblob carries the state on its own, so dropping the id costs nothing (sanitize_reasoning_item).
anthropic/-routed relay it inherits Anthropic protocol behaviour, including LiteLLM’s json-mode assistant prefill, which newer Bedrock versions reject. Reached through an OpenAI-compatible gateway it resolves as openai/, no prefill is injected, and json_mode_enabled can stay on.
Table B: Conflicts and workarounds
FIM One emits three of the fourtool_choice states: auto from the ReAct loop (react.py::_run_native), a named function from structured output (structured.py::_call_llm), and none when the finish-signal answer replays the tools payload. No call site emits required; that column records the provider’s constraint on non-auto tool choice, which applies to required and to a named function alike.
Table C: Thinking protocol
LLM_REASONING_EFFORT accepts low, medium and high; any other value is read as unset (deps.py::_reasoning_effort). What FIM One then puts on the wire is per-provider, and that is what this table records. The replay column is the return value of reasoning_replay_policy, which is a small closed set of four states rather than a per-provider list.
unsupported and informational_only produce the same bytes on the wire: both strip reasoning_content and signature from outgoing history. They differ in intent, so a model that clearly does reason but lands in unsupported is a gap in the fragment table rather than a live bug.
Relay/proxy gotchas
Third-party gateways fail in ways a direct provider does not, and most of those failures are silent. Each row below pairs the symptom with its mechanism and with what FIM One already does about it.Support boundary. FIM One guarantees the behaviour documented on this page for first-party endpoints: OpenAI’s own API, Anthropic, Google, and any vendor serving its own models directly. Third-party relays are supported on a best-effort basis and are not covered by that guarantee, because what a relay does to a request is outside our control and frequently outside its own documentation. A relay can drop a parameter, rewrite history, strip a cache breakpoint, or answer a protocol it only partially implements, and in most of those cases it returns a
200 rather than an error.This is a statement about what we promise, not a restriction on what runs. FIM One does not maintain an allowlist of approved hosts, and nothing here is gated on a domain. Capability is decided by what an endpoint actually does: a missing route answers 404 and is remembered, an ignored include yields empty reasoning items and the replay becomes a no-op, and a rejected request falls back for that call. Probing the endpoint is more accurate than inferring its capabilities from its hostname, and it is the only approach that keeps working for Azure OpenAI, enterprise gateways, and self-hosted proxies that implement the protocol correctly.If a relay misbehaves in a way the fallbacks do not catch, pin the protocol yourself with FIM_GPT5_RESPONSES_MODE (bridge or off) or the per-model tool_choice_enabled and json_mode_enabled toggles, and reproduce against the first-party endpoint before filing it as a FIM One bug.Recommended per-model configuration
Bothtool_choice_enabled and json_mode_enabled can be toggled per model in Admin → Models → Advanced settings. The defaults, both TRUE, are correct for most providers; only adjust when you see errors or wasted latency. Which providers need an adjustment is recorded in Table B above, and the per-model view an operator fills in lives in Model Management.
ENV-level overrides apply to all models configured via environment variables (not admin UI):
Reasoning effort and thinking configuration
FIM One exposes two env vars for controlling extended thinking / reasoning:
Two behaviours follow automatically once thinking is on, and neither needs user configuration:
- Temperature is handled for you. On an
anthropic/route with thinking active,_build_request_kwargspinstemperatureto 1.0, which is what Bedrock demands. Models that reject sampling parameters outright (Opus 4.7 and 4.8, Fable 5, Mythos 5) havetemperatureremoved from the request entirely, thinking or not. Do not setLLM_TEMPERATURE=1by hand for this. - GPT-5.x keeps tools and reasoning together where it can. FIM One probes the Responses bridge first for GPT-5.x, because that is the only surface where the two combine. An endpoint with no usable
/v1/responsesroute falls back to chat completions, the verdict is cached per endpoint, and on that path a request carryingtoolssends an explicitreasoning_effortofnone. Omitting the field is not equivalent, since the server default is notnone.
Defensive parsing for structured output
Even with native_fc working correctly, the structured output pipeline includes a defensive parsing layer to handle edge cases from any provider or compatibility layer. The DAG planner’s_dict_to_steps parser handles three common edge cases:
-
Single object instead of array. Some models return
{"steps": {"id": "1", "task": "..."}}(a single step object) instead of{"steps": [{"id": "1", "task": "..."}]}(an array). The parser detects this by checking foridortaskkeys and wraps the object in a list. -
Double-encoded JSON string. When structured output falls through to json_mode (which lacks schema enforcement), some providers return the
stepsvalue as a JSON string rather than a native array — e.g.,{"steps": "[{\"id\": \"1\", ...}]"}. This string may also contain literal newlines (from the model’s formatting) that break standardjson.loads. The parser usesextract_json_value()(which includes_repair_json_strings) to handle:- Literal newlines inside JSON string values
- Invalid escape sequences (common with LaTeX or code content)
- Other serialization quirks from compatibility layers
-
Missing
stepswrapper. The model may return a single step as the top-level object without thestepswrapper key. The parser detectsidandtaskat the root level and wraps accordingly.
Under normal operation, native_fc returns properly structured tool call arguments and these edge cases do not arise. The defensive parsers exist as a safety net for custom
BaseLLM subclasses, unusual provider behaviors, or fallback scenarios where structured output degrades to json_mode or plain_text.Prompt caching (cross-provider)
FIM One implements Anthropic’s explicit prompt caching viacache_control breakpoints and simultaneously benefits every other provider’s automatic prefix caching through the Prompt Section Registry. The goal is a single prompt-assembly path that works across all providers without per-call prompt shape divergence.
Architecture
Thefim_one.core.prompt module exposes three primitives:
PromptSection— a named fragment with either a staticcontent: stror a dynamiccontent: CallablePromptRegistry— a memoized store (static sections render once, dynamic sections re-render per call)DYNAMIC_BOUNDARY— a sentinel marker the registry inserts between the last static section and the first dynamic one, so callers can split the rendered prompt at the cache breakpoint
- Static prefix (~95% of the prompt) — identity, core guidelines, tool descriptions
- Dynamic suffix — current datetime, per-request language directive, handoff context
Capability detection
fim_one.core.prompt.caching.is_cache_capable(model_id) returns True when the model id contains any of: claude, anthropic, bedrock/anthropic, vertex_ai/claude. These providers receive two role="system" messages with cache_control: {"type": "ephemeral"} on the first (static) message.
Every other provider receives a single concatenated system message with no cache_control field — necessary because non-Anthropic endpoints either reject the field or silently drop it, and sending it through some relays causes 400 unknown parameter errors.
Cross-provider coverage
The
PromptRegistry benefits every provider with auto prefix caching “for free” — by keeping the static portion byte-identical across calls (current datetime lives in the dynamic suffix, not prefix), every auto-caching provider’s hash matches and hits their cache. This is why the Registry is a foundational modelless win even before considering Anthropic-specific cache_control.
Observability
Everychat/* response’s done_payload now includes:
TurnProfiler emits a structured log line per turn: turn_cache summary | model=claude-sonnet-4-6 | read_tokens=1067 | create_tokens=0 | saved_input_tokens=961 (~90%). This also functions as a relay honesty probe — if you route through an API relay, compare actual billed input vs read_tokens to detect whether the relay strips cache_control or keeps the 0.10× discount.
No dollar estimate is returned at the LLM layer — pricing and relay markup are applied above, so the LLM layer only returns objective token counts.
Multi-turn cache ROI
Measured on Claude 4 ReAct turns with the default agent prompt:
A 10-iteration ReAct run with 10 tools saves ~8,640 input tokens per turn after the first (9 cache hits × 1067 tokens × 90%). Anthropic charges 1.25× for cache write on the first call, so the breakeven is at the second call — single-shot queries do not benefit.
Reasoning replay policy (modelless correctness)
Extended thinking / reasoning blocks behave differently across providers. A uniform serialization policy breaks both protocol contracts and automatic prefix caches.fim_one.core.prompt.reasoning.reasoning_replay_policy(model_id) returns one of four values and gates ChatMessage.to_openai_dict(replay_policy=...) in OpenAICompatibleLLM._build_request_kwargs().
Four policies
anthropic_thinking— Claude family (includinganthropic/,bedrock/anthropic,vertex_ai/claude). Thinking blocks MUST be replayed withsignatureattached; Anthropic rejects subsequent turns if the signature is missing or altered.informational_only— models that emit CoT but do NOT expect replay: DeepSeek reasoning mode (deepseek-reasoneron V3.2, and the olderdeepseek-r1and R1-Distill ids the fragment table still matches), Qwen QwQ, Gemini flash-thinking, OpenAI o1 / o3 / o4. Their documentation explicitly says “do not sendreasoning_contentback in message history”. Sending it anyway:- Violates the provider contract (may start rejecting in future versions)
- Silently invalidates their automatic prefix cache — message bytes mutate on every turn, breaking the hash
openai_responses— GPT-5.x, matched on thegpt-5fragment. Its reasoning state is not text but a sequence of opaque items carrying encrypted payloads, and only/v1/responseshas a slot for them. On that protocol the items are replayed verbatim, which is what keeps the model’s chain of thought alive across tool rounds. The readable summary is still dropped from outgoing requests, so on the chat-completions fallback this behaves exactly likeinformational_only. Checked before the informational fragments, whose genericreasoningentry would otherwise swallow proxy-tagged GPT-5 ids.unsupported— the catch-all: models with no reasoning capability (GPT-4o, Gemini 1.5, Mistral, Llama), and reasoning models whose id matches no fragment (GLM, MiniMax, Kimi, Doubao). No field should be replayed either way, so this policy puts the same bytes on the wire asinformational_only. It is also the safe default for unknown model ids.
reasoning_content and the opaque reasoning_items are independent fields on ChatMessage. to_openai_dict() never serialises the items at all, so they are structurally incapable of leaking onto a chat-completions request, whatever the policy says.
Enforcement
All policy evaluation happens in one place (_build_request_kwargs). ChatMessage.to_openai_dict(replay_policy=None) preserves the A3 permissive default so uncoordinated callers don’t regress. The cross-provider test matrix lives in tests/test_reasoning_replay_policy.py with reverse assertions proving that non-Anthropic requests do NOT leak reasoning_content.
For users
Both feature and bug behavior is automatic — you don’t need to configure anything. Workflow implications:- If you switch agents between Claude and DeepSeek in the same conversation, history is stored with thinking blocks intact; on the next turn, the outgoing message shape adapts per the current model.
- If you use a proxy / custom
BaseLLMsubclass, make sure its model id is recognizable (contains one of the fragments) or the defaultunsupportedpolicy will apply — which is safe but means Claude behind an unusual proxy might lose thinking replay. Add the model-id fragment to_CACHE_CAPABLE_MODEL_FRAGMENTS(incore/prompt/caching.py) and/or the reasoning policy lookup.
Troubleshooting
“This model does not support assistant message prefill” Bedrock + json_mode. Two fixes: (1) setLLM_JSON_MODE_ENABLED=false or disable JSON Mode in the admin model settings; or (2) if your Bedrock provider offers an OpenAI-compatible /v1/chat/completions endpoint, switch to that — FIM One resolves it as openai/ and the prefill injection never occurs.
“Thinking may not be enabled when tool_choice forces tool use” / “tool_choice ‘specified’ is incompatible with thinking enabled”
For Anthropic models, structured_llm_call disables thinking for native_fc calls automatically. Where thinking cannot be turned off through the API, such as kimi-k2.5 and kimi-k2-thinking or deepseek-reasoner, disable “Native Function Calling” in the model’s Advanced settings, or set LLM_TOOL_CHOICE_ENABLED=false globally. The degradation chain will skip native_fc and extract structured output via json_mode or plain_text instead. Check Table B of the Provider Capability Matrix before assuming a thinking model has this problem; MiniMax does not.
“DAG pipeline failed: LLM ‘steps’ is not an array”
The LLM returned the steps field as a string or single object instead of an array. This typically means structured output fell through to json_mode (which lacks schema enforcement). Check the log for structured_llm_call: level=xxx — if it shows json_mode instead of native_fc, native_fc is failing silently. If using a custom BaseLLM subclass, verify it accepts the reasoning_effort kwarg.
ReAct falls back to JSON mode unexpectedly
Check that the model’s abilities["tool_call"] is True. This is always True for OpenAICompatibleLLM, but a custom BaseLLM subclass might override it. Verify with the model detail endpoint in the admin API.
structured_llm_call exhausts all levels and raises StructuredOutputError
The model failed to produce parseable JSON at any level. This is rare with modern models. Check: (1) the schema is valid JSON Schema, (2) the model has enough max_tokens to produce the full response, (3) the system prompt is not contradicting the schema instructions. The DAG planner and analyzer both provide default_value fallbacks, so this error only propagates from call sites that explicitly omit defaults.