How FIM One Uses Models
FIM One has three model roles:
Fast and Reasoning fall back to General if not configured. For production deployments, splitting into at least two models (General + Fast) gives the best cost/quality balance.
These roles can be configured via ENV variables or through the admin UI’s Model Groups feature, which allows one-click switching between model sets. See Model Management for the full admin UI guide.
Quick Selection Matrix
This table lists the combinations we recommend, not the full set of providers FIM One supports. xAI (Grok), ByteDance Doubao, Mistral and any OpenAI-compatible relay all work; see the Provider Capability Matrix for the complete list and how each one is routed.
Structured Output Compatibility
FIM One’s DAG planner needs the model to return valid structured JSON. Internally, it tries three extraction levels in order:- Native Function Calling — forces the model to output JSON matching a schema via the tool-call API. Most reliable.
- JSON Mode — requests
response_format: json_object. Guarantees valid JSON, but does not enforce schema compliance. - Plain Text Extraction — parses JSON from free-form text as a last resort.
tool_choice) give the best planning reliability. If a model only reaches Level 2, its output quality depends on how well it follows prompt instructions — weaker models may produce valid JSON that doesn’t match the expected structure.
This table is a selection aid. The authoritative, code-anchored version, including which
tool_choice states each provider accepts and what FIM One does when one is rejected, is the Provider Capability Matrix.
Thinking / Reasoning Compatibility
Different providers implement “thinking” (chain-of-thought reasoning) in fundamentally different ways. This matters because thinking mode can conflict with tool calling, and the output appears in different places depending on the provider. FIM One handles all of these transparently — this table helps you understand what’s happening under the hood.Key Concepts
- Opt-in — thinking is off by default; you enable it via an API parameter (e.g.,
reasoning_effort). Can be selectively disabled per call. - Always-on — the model always thinks; no API parameter to turn it off. You’d need to switch to a non-thinking model variant to avoid it.
- Model-level — thinking is determined by which model ID you choose (e.g.,
deepseek-reasonervsdeepseek-chat), not by a parameter.
Compatibility Matrix
The per-provider detail behind this table, including how each
effort level is translated and whether reasoning is replayed on later turns, lives in Table C of the Provider Capability Matrix.
How FIM One Handles Each Case
API-levelreasoning_content (Claude, DeepSeek): The reasoning field is read directly from the API response and displayed in the UI Reasoning panel. No post-processing needed.
<think> tags in content (MiniMax, Qwen, QwQ, and other open-source derivatives): FIM One automatically strips <think>...</think> tags from the content field and reroutes the thinking text to the Reasoning panel. This works for both streaming and non-streaming responses.
Forced FC + thinking conflicts are per-provider, not a property of thinking models in general. Claude rejects the combination but its thinking is opt-in, so FIM One turns thinking off for that one call by passing reasoning_effort=None and native function calling proceeds. Kimi also rejects it, and its thinking is selected by model id rather than by a parameter, so the fix there is to disable Native Function Calling for the thinking models. MiniMax thinks on every call and accepts forced function calling anyway, which is why no workaround applies to it.
Fallback chain: If forced function calling fails for any reason, FIM One falls back automatically: native FC → JSON mode → plain text extraction. This three-tier approach ensures planning works even with providers that have partial tool-calling support.
If you’re using a model that always thinks (MiniMax M2.7,
deepseek-reasoner) as your Main LLM, the thinking output will appear in every agent iteration’s Reasoning panel. This is normal — it doesn’t affect functionality, and you get to see the model’s reasoning process.Provider Details
OpenAI
The most battle-tested option. OpenAI models have the best native function calling (tool-calling) support, which directly impacts agent reliability. The GPT-5 family (August 2025+) is a major generational leap over GPT-4. Recommended models:- Main:
gpt-5.4(latest flagship, Mar 2026 — 1M+ context, computer use) oro3(best reasoning accuracy) - Fast:
gpt-5.4-mini(4.50 per MTok) orgpt-5.4-nano(cheapest at 1.25 per MTok) - Budget Fast:
gpt-5-mini(2.00) andgpt-5-nano(0.40) remain available at lower prices - Legacy:
gpt-4.1(still in API, 1M context, good for coding)
LLM_REASONING_EFFORT=medium — works natively with o-series and GPT-5.x models. GPT-5.4 supports reasoning_effort with levels none, low, medium, high, xhigh. The o-series requires max_completion_tokens instead of max_tokens, which LiteLLM handles automatically. Note: /v1/chat/completions rejects GPT-5.x requests that combine tools with reasoning, so FIM One talks the Responses API directly for GPT-5.x, where both work together. That path also carries each turn’s encrypted reasoning into the next one, so an agent building a multi-step answer keeps what it already worked out instead of re-deriving it at every tool call. Endpoints without a /v1/responses route fall back to chat completions with an explicit reasoning_effort: "none" during agent tool-use steps, and FIM_GPT5_RESPONSES_MODE can force either fallback by hand. Other model families on OpenAI-compatible endpoints stay on chat completions: they gain nothing from Responses, and proxy shims for it can badly buffer streaming. GPT-5.4 requires temperature=1, which FIM One handles automatically via LiteLLM’s parameter filtering (drop_params).
Anthropic (Claude)
Claude excels at nuanced reasoning and complex multi-step tasks. FIM One connects via LiteLLM, which routes Anthropic models through their native API automatically. The current generation is Claude 4.6 (February 2026). Recommended models:- Main:
claude-sonnet-4-6(best balance of capability and cost — 15 per MTok) - Fast:
claude-haiku-4-5(fast and cheap — 5 per MTok) - Premium:
claude-opus-4-6(most capable, 128K max output — 25 per MTok)
https://api.anthropic.com/v1/
Opus 4.6 and Sonnet 4.6 have a 1M context window (GA since March 13, 2026 — no beta header needed). Haiku 4.5 has a 200K context window.
Reasoning: Set LLM_REASONING_EFFORT=medium. LiteLLM routes Anthropic models through the native API, so reasoning_content (extended thinking) is fully returned and visible in the UI “thinking” step. Claude 4.6 and newer use Adaptive Thinking (thinking: {type: "adaptive"} plus output_config.effort) in place of a manual budget_tokens, which FIM One emits directly rather than relying on LiteLLM’s mapping. Anthropic requires temperature=1 while thinking is active, and the system enforces that for you: the request builder pins the value on Anthropic routes, and on models that reject sampling parameters outright it removes temperature altogether. Do not set LLM_TEMPERATURE=1 by hand. See Extended Thinking for details.
Google Gemini
Gemini models offer strong performance at competitive pricing via Google’s OpenAI-compatible endpoint. The 3.x generation (late 2025+) is a major leap — Gemini 3 Flash outperforms 2.5 Pro while being 3x faster. Note:gemini-3-pro-preview was shut down March 9, 2026 — use gemini-3.1-pro-preview instead.
Recommended models:
- Stable (GA):
gemini-2.5-pro(main) +gemini-2.5-flash(fast) — production-ready - Latest (Preview):
gemini-3.1-pro-preview(main) +gemini-3-flash-preview(fast) +gemini-3.1-flash-lite-preview(budget fast) — best performance, but preview status
https://generativelanguage.googleapis.com/v1beta/openai/
Reasoning: reasoning_effort is supported on the compatibility endpoint — set LLM_REASONING_EFFORT=medium and it works out of the box.
DeepSeek
DeepSeek offers the best cost/performance ratio in the market. V3.2 (December 2025) unified the chat and reasoning lineages into a single model, with incredibly low pricing. Model IDs (both backed by V3.2):deepseek-chat— general purpose (non-thinking mode)deepseek-reasoner— chain-of-thought reasoning mode, returnsreasoning_content
https://api.deepseek.com
Pricing: 0.42 per MTok (cache hit: $0.028) — by far the cheapest frontier-class API.
Output limits: deepseek-chat max output is 8K tokens (must set explicitly via max_tokens). deepseek-reasoner max output is 64K tokens (includes chain-of-thought).
V4 expected April 2026: trillion-parameter multimodal model with 1M context window. Expect new model IDs when it launches.
Chinese Domestic Models
All major Chinese model providers expose OpenAI-compatible endpoints. These are particularly strong for Chinese-language tasks and offer competitive local pricing.Qwen / 通义千问 (Alibaba Cloud)
Qwen 3.5 (February 2026) is the latest generation — the 397B MoE flagship outperforms GPT-5.2 on MMLU-Pro. Strongest Chinese language support and cheapest frontier-class pricing (~$0.11/MTok input).- Base URL (China):
https://dashscope.aliyuncs.com/compatible-mode/v1 - Base URL (Global):
https://dashscope-intl.aliyuncs.com/compatible-mode/v1 - Main:
qwen3.5-plus(flagship, 1M context, 0.66 per MTok) orqwen3-max(256K, strongest) - Fast:
qwen3.5-flash(0.22 per MTok) orqwen-turbo(0.08 per MTok) - Reasoning:
qwen3-maxwithenable_thinking: trueparameter (there is no separateqwen3-max-thinkingmodel ID)
ChatGLM / 智谱
GLM-4.7 and GLM-5 (2026) are the latest models. GLM-5 is the 745B MoE flagship approaching Claude Opus-level on coding/agent tasks.- Base URL (Domestic):
https://open.bigmodel.cn/api/paas/v4 - Base URL (Z.AI International):
https://api.z.ai/api/paas/v4 - Main:
glm-4.7(strong coding, 2.20 on Z.AI) - Fast:
glm-4.7-flash(free tier!) orglm-4.7-flashx(0.40, higher throughput) - Reasoning:
glm-5(745B MoE flagship, 3.20)
tool_choice is not supported — only "auto" works.
MiniMax
MiniMax M2.7 (March 18, 2026) is the latest model, open-weight and scores 80.2% on SWE-Bench. M2.5 remains available as a fast/budget option. MiniMax provides two separate API endpoints for different regions:- Base URL (Global/海外版):
https://api.minimax.io/v1— for users outside mainland China - Base URL (China/国内版):
https://api.minimaxi.com/v1— for users in mainland China (note the extraiinminimaxi) - Main:
MiniMax-M2.7 - Fast:
MiniMax-M2.5 - Speed:
MiniMax-M2.7-highspeed(2x cost, lower latency)
Kimi / 月之暗面 (Moonshot)
Kimi K2.5 (January 2026) has 256K context and strong coding performance (76.8% SWE-Bench among open-source models).- Base URL (Global):
https://api.moonshot.ai/v1 - Base URL (China):
https://api.moonshot.cn/v1 - Main:
kimi-k2.5 - Fast:
kimi-k2(non-thinking, function calling works) - Reasoning:
kimi-k2-thinking(2.00 per MTok)
tool_choice only works when thinking mode is off. When thinking is enabled, only "auto" is supported.
Local Models (Ollama)
Run models entirely on your own hardware — no API key needed, fully offline. Ollama exposes an OpenAI-compatible endpoint out of the box. The open-source landscape has changed dramatically — Qwen 3.5, Llama 4, and GPT-OSS (OpenAI’s first open-weight models) are all available. Base URL:http://localhost:11434/v1
Recommended models by VRAM:
Best for tool-calling: Qwen 3/3.5 (32B+), GLM-4.7, GPT-OSS, Mistral — these have explicit function-calling training. Models with 14B+ parameters are the minimum for reliable tool calling; 32B+ is strongly preferred.
Third-Party Relay Platforms
Many users access multiple model providers through a single relay (proxy) service. FIM One automatically detects the correct API protocol based on URL path patterns — just fill in theLLM_BASE_URL and it works.
How It Works
When your base URL points to a third-party relay, FIM One inspects the URL path to determine which protocol to use:
Resolution order: Explicit DB provider field > domain match (official APIs) > URL path hint (relay platforms) > OpenAI compatible fallback.
Example: One Relay, Three Protocols
With a single relay account, you can access different providers by simply changing the base URL path:Step-by-Step: How Path Detection Works
Here’s a concrete example showing what happens internally when you configure a relay:- FIM One sees
/claudein the URL path → detects Anthropic native protocol - Model is prefixed as
anthropic/claude-sonnet-4-6for LiteLLM routing - Requests use Anthropic’s
/v1/messagesformat withx-api-keyauth header reasoning_effort=mediumis translated to Anthropic’s nativethinkingparameter (not OpenAI’sreasoning_effort)
Why This Matters
- Anthropic native endpoint gives you proper
reasoning_contentsupport (extended thinking visible in the UI), correct tool-calling format, andx-api-keyauthentication — features lost when using OpenAI-compatible translation. - Google native endpoint gives you native Gemini parameters and
x-goog-api-keyauthentication. - OpenAI compatible is the universal fallback and works with any relay, but provider-specific features (like extended thinking output) may be unavailable.
If your relay platform uses non-standard path conventions (e.g., no
/claude or /anthropic in the URL), FIM One falls back to OpenAI compatible protocol — which works for most use cases. For full native protocol support, you can set the provider field explicitly via the admin model configuration UI.Relays are best-effort. FIM One’s documented behaviour is guaranteed for first-party endpoints, meaning OpenAI, Anthropic, Google, and other vendors serving their own models directly. Relays work and are widely used, including for models that are only reachable that way, but what a relay does to a request is outside our control, so they carry no such guarantee. Nothing is blocked by hostname: capability is probed per endpoint, and unsupported behaviour falls back on its own. Before reporting a model-layer bug, reproduce it against the first-party endpoint.
Configuration Strategy
Main vs Fast: When to Split
- Split when your main model is expensive or slow (e.g.,
gpt-5.4+gpt-5.4-nano). DAG mode runs many parallel steps — using a cheaper fast model saves significant cost. - Same model when your model is already cheap (e.g.,
deepseek-chatfor both). The overhead of managing two models isn’t worth it.
When to Enable Reasoning
- Enable for complex analytical tasks, multi-step planning, and tasks requiring careful judgment
- Disable (default) for routine tasks, simple Q&A, and cost-sensitive deployments
- Reasoning typically increases cost 2-5x per request —
mediumeffort is a good starting point
Context Window Sizing
SetLLM_CONTEXT_SIZE to match your model’s actual window:
For local models, set both
LLM_CONTEXT_SIZE and LLM_MAX_OUTPUT_TOKENS explicitly — defaults assume cloud-scale context windows that local models cannot support.