The thesis
No single model wins every job. Model name alone is not the right routing unit. Base model, product or API surface, available tools, deployment requirements, and the consequences of an incorrect result together shape real behavior. This guide is editorial and should be validated against the intended workload before standardizing.
Selected current models
This is a selected lineup, not a complete catalog. Omitting a model is preferable to publishing an unverified entry.
OpenAI GA
GPT-5.6 is generally available across ChatGPT, Codex, and the OpenAI API in three tiers: GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. Product availability and selectable reasoning effort vary by plan and surface. Native input includes text and images. Primary native output is text. Other media generation may use separate models or tools. Programmatic Tool Calling is available; multi-agent execution through the Responses API is beta. OpenAI's ultra setting coordinates four agents in parallel by default for supported workloads.[1]
Anthropic GA Limited access
Claude Fable 5 and Claude Mythos 5 share the same underlying model. Fable includes Anthropic's deployed safety classifiers, while Mythos provides approved organizations with controlled access under different safeguard conditions. Fable may refuse selected requests; API refusals may return a refusal stop reason. In supported Anthropic environments, an optional fallback (beta, platform-dependent) can rerun a refused request using another model such as Opus 4.8. Anthropic reports that more than 95 percent of Fable sessions do not trigger fallback. Fable 5 requires 30-day retention and is not eligible for zero-data-retention treatment; verify Mythos retention requirements separately for the selected provider.[2]
Google GA
Gemini 3.6 Flash supports multiple input types. Its primary generated output is text unless a separate media-generation model or tool is used. It supports Computer Use as a built-in client-side tool through supported Gemini API and Gemini Enterprise environments. Public Gemini products do not reproduce Google Search, AI Mode, or AI Overviews. Actual search surfaces must be evaluated directly.[3]
xAI GA
Grok 4.5 is available through xAI's API and the Grok product. The base model is distinct from X Search, live web retrieval, and other external tools; those capabilities are tool-mediated rather than intrinsic to the model.[4]
DeepSeek Preview
DeepSeek describes V4 Pro and V4 Flash as preview releases. Documented modalities are text input and text output through the public API. The announced retirement deadline for the legacy deepseek-chat and deepseek-reasoner aliases passed on July 24, 2026 at 15:59 UTC. Verify the current model ID before deployment.[5]
Cohere GA
Command A+ has a 128K input context and is positioned for private or sovereign deployment, enterprise retrieval, multilingual applications, citation-oriented responses, controlled enterprise infrastructure, and tool-using enterprise workflows. It is not a universal general-purpose peer of the largest frontier models and should not be grouped with million-token-context models on the basis of context size.[6]
Meta US public preview
Muse Spark 1.1 is available through Meta's public-preview model API for eligible developers in the United States. It is not open-weight and should not be positioned as a normal production default.[7]
Moonshot (Kimi) Verify
Kimi K3 is available via API; weight availability must be confirmed from Moonshot's current model page on the deployment date. Kimi K3 is a very large model. Open-weight availability does not imply practical self-deployment for a typical organization.[8]
Alibaba (Qwen) GA
Qwen3.7-Plus and Qwen3.7-Max are separate hosted models with distinct capabilities and modalities documented individually by Alibaba Cloud. Neither hosted tier should be described as open-weight unless that specific hosted model is explicitly documented as open-weight.[9]
Mistral GA
Mistral Small 4 is Apache 2.0, open-weight, 256K context, with text and image input. Mistral OCR 4 remains a document-ingestion specialist rather than a general frontier competitor.[10]
Apple Platform
Apple AFM 3 is an Apple ecosystem model deployment rather than a conventional public model API. Availability depends on Apple operating-system, device, feature, and Private Cloud Compute rollout.[11]
Routing table by job class
| Job class | What matters most | Selected candidates | Important constraint |
|---|---|---|---|
| Fast drafting and transformation | Responsiveness, clean text output | GPT-5.6 Luna, Gemini 3.6 Flash, Qwen3.7-Plus | Confirm reasoning effort and output format for the surface in use. |
| Research and synthesis | Handling of long or mixed material | GPT-5.6 Terra, Claude Fable 5, Gemini 3.6 Pro-tier | Test with representative source material before standardizing. |
| Evidence-grounded analysis | Grounding to supplied sources and citation behavior | Claude Fable 5, Cohere Command A+, GPT-5.6 Terra | Grounding depends on retrieval configuration, not model choice alone. |
| High-impact decision support | Interpretability, reviewability, controls | Claude Fable 5, GPT-5.6 Sol | A named human must be accountable for the final decision. |
| Code and repository work | Repository comprehension and tool use | GPT-5.6 (Codex), Claude Fable 5 | Validate on the target stack; results vary by language and repo size. |
| Tool-using agents | Reliable tool invocation and recovery | GPT-5.6 (Responses API, beta multi-agent), Claude Fable 5 | Multi-agent execution features are beta on some surfaces. |
| Computer use | Screen understanding and action safety | Gemini 3.6 Flash (Computer Use tool) | Sandbox and rate-limit before granting real environment access. |
| Document ingestion | OCR fidelity and layout preservation | Mistral OCR 4 | OCR quality varies by document type and language. |
| Classification and evaluation | Consistency across a labeled sample | GPT-5.6 Luna, Qwen3.7-Plus, Mistral Small 4 | Evaluate against a held-out labeled sample before scaling. |
| Media production | Native or tool-based media generation | Provider-specific media models | Native model output is primarily text; media generation typically uses separate models or tools. |
| High-throughput text work | Sustainable operating profile | GPT-5.6 Luna, Gemini 3.6 Flash, Mistral Small 4 | Test at expected concurrency; effort settings affect behavior. |
| Private or sovereign deployment | Data residency and infrastructure control | Cohere Command A+, Mistral Small 4 (open-weight) | Confirm supported regions and infrastructure options with the provider. |
| On-device interaction | Platform integration | Apple AFM 3 | Availability depends on OS, device, and PCC rollout. |
| AEO and SXO support | Structured-data and content review | GPT-5.6 Terra, Claude Fable 5, Gemini 3.6 Flash | Public chatbots do not reproduce proprietary search and recommendation systems. |
Performance varies by model version, endpoint, effort setting, tools, context, and workload. Validate the complete operating configuration before standardizing.
A note on cost
Token price alone does not determine operating cost. Tool use, retries, review burden, latency, and deployment requirements can materially change the final cost of a workflow.
A note on context length
Advertised context size does not guarantee dependable performance across long or complex material. Test representative documents before deployment.
A note on governance
High-impact uses require approved data handling, appropriate organizational controls, and a named human accountable for the final decision. Model capability alone does not authorize autonomous action.
A note on AEO and SXO work
AI models can support content review, structured-data work, document interpretation, and comparison of publicly observable answer experiences. Public chatbots do not reproduce proprietary search and recommendation systems.
A note on benchmarks
This guide intentionally avoids a fixed benchmark leaderboard because scores depend on model version, tools, reasoning effort, scaffolding, and test conditions. Benchmark figures are meaningful only when the benchmark version, exact model version, date, harness, tool use, reasoning effort, and whether the result is vendor-reported or independent are all disclosed together.
Sources
- OpenAI. GPT-5.6 model and API documentation. platform.openai.com/docs/models
- Anthropic. Claude model overview and safeguards documentation. docs.anthropic.com/en/docs/about-claude/models
- Google. Gemini API models and Computer Use documentation. ai.google.dev/gemini-api/docs/models
- xAI. Grok API model documentation. docs.x.ai/docs/models
- DeepSeek. Model and deprecation notices. api-docs.deepseek.com
- Cohere. Command A+ model documentation. docs.cohere.com/docs/models
- Meta. Muse Spark preview access documentation. ai.meta.com
- Moonshot. Kimi K3 model page. platform.moonshot.ai
- Alibaba Cloud. Qwen3.7-Plus and Qwen3.7-Max documentation. help.aliyun.com/zh/model-studio
- Mistral. Mistral Small 4 and OCR 4 documentation. docs.mistral.ai
- Apple. Apple Intelligence Foundation Models. machinelearning.apple.com