Available Models

Prev Next

Available Models

AI CTRL provides access to models from four providers: OpenAI, Anthropic, Google, and Perplexity. This page groups them by the four types described in Choosing a Model.

For what each model costs to run, see the AI CTRL Ratecard.

Your list will be shorter than this one. Your administrators choose which models appear in your dropdown, and most deployments enable a curated handful rather than everything available. If something here isn't in your list, ask them: it's a deliberate choice, not an oversight.

Within each type below, models are listed roughly from most economical to most expensive. GPT-5 mini is the standard baseline that credit allocations are measured against.

Everyday / lightweight

General questions, drafting, summarizing, rewriting, quick lookups. Fast and inexpensive: these handle more than people expect, so try one first and move up only when the answer isn't good enough.

Model Provider Best for
GPT-5 nano OpenAI Summaries, tagging, and high-volume triage
Gemini 2.0 Flash-Lite Google Simple input–output and minimum spend per request
Gemini 2.0 Flash Google Balanced multimodal assistant
Gemini 2.5 Flash-Lite Google Routing, translation, and extremely high request volume
GPT-4.1 nano OpenAI High-volume classification, routing, and low-latency tasks
GPT-4o mini OpenAI Budget multimodal chat, vision, and simple extraction
GPT-5.4 Nano OpenAI Snippets, small fixes, classification, and routing
GPT-5 mini ★ OpenAI General chat, analysis, and light coding
Claude Haiku 3 Anthropic Maximum economy and minimum latency at high scale
Gemini 3.1 Flash-Lite Google High-throughput tasks where cost and speed both matter
Gemini 2.5 Flash Google Chat, summarization, and extraction at large scale
GPT-4.1 mini OpenAI Cost-effective coding, extraction, and everyday assistants
GPT-5.4 Mini OpenAI Quick iterations, lighter coding, and tool-heavy workflows
Claude Haiku 4.5 Anthropic High-volume assistants and chat with strong answer quality
Gemini 3.5 Flash Google Fast reasoning, multimodal agents, and large-scale production workloads
GPT-5.6 Luna OpenAI Classification, summarization, lightweight chat, and high-throughput automation
GPT-4o OpenAI Vision and general multimodal assistants
Gemini 3.6 Flash Google Agentic and multimodal workloads: stronger coding and knowledge work than 3.5 Flash, at lower cost
Gemini 3.7 Flash Google Fast everyday work at large context
GPT-5.6 Terra OpenAI Everyday professional work: the balanced default where Sol is overkill

Reasoning

Multi-step problems, analysis, technical writing, and code. Slower and more verbose than a short question needs, and the most expensive per credit: reserve them for work that needs the thinking.

Model Provider Best for
o4-mini OpenAI Cost-efficient math, logic, and structured reasoning
GPT-5 OpenAI Reasoning, coding, and general-purpose assistants
GPT-5.1 OpenAI Everyday coding and agent-style tasks
GPT-5.1 Codex OpenAI Economical coding, implementation, and code review
GPT-5.2 OpenAI Coding, tool use, and multimodal workflows
GPT-5.2 Codex OpenAI IDE-style edits, debugging, and repo-aware changes
GPT-5.3 Codex OpenAI Agentic coding, refactors, and multi-file software work
o3 OpenAI STEM, logic, and careful step-by-step analysis
Gemini 3 Pro Google Complex analysis, coding, and multimodal assistants (preview)
Gemini 3.1 Pro Google Agents, coding, and multimodal reasoning (preview)
GPT-5.4 OpenAI Coding, agents, math, and broad analysis
Claude Sonnet 4 Anthropic Implementation, documentation, and general software tasks
Claude Sonnet 5 Anthropic Agents, coding, and near-Opus quality at lower cost
Claude Opus 4.5 Anthropic Elite reasoning, research depth, and strategic analysis
Claude Opus 4.7 Anthropic Stronger coding, vision, and complex multi-step tasks
Claude Opus 4.8 Anthropic Most capable Opus — complex reasoning and long-horizon agentic work
Claude Opus 5 Anthropic Complex agentic coding, autonomous workflows, and advanced research where accuracy outweighs cost
GPT-5.5 OpenAI Flagship for coding, agents, computer use, and professional work
Claude Opus 4 Anthropic Broad flagship-quality coding, analysis, and writing
Claude Opus 4.1 Anthropic Premium coding, instruction following, and high-stakes work
GPT-5 Pro OpenAI Deep analysis and planning with extended thinking
GPT-5.2 Pro OpenAI Very demanding reasoning, long-form analysis, complex decisions
GPT-5.4 Pro OpenAI Hardest problems, maximum precision, long-running work
GPT-5.5 Pro OpenAI Highest-precision reasoning, audits, and high-stakes professional work
GPT-5.6 Sol OpenAI Agentic multi-step tasks, research synthesis, and high-stakes technical workflows

Large context

Long documents, large data sets, and anything used with the Structured Data Tool. Use when the volume is genuinely large, not by default.

Model Provider Best for
Gemini 3 Flash Google Fast multimodal work and thinking-style tasks at large context
Gemini 2.5 Pro Google Complex code, deep analysis, and very long context
GPT-4.1 OpenAI Strong coding and long-context instruction following
Claude Sonnet 4.5 Anthropic Product engineering and coding across long sessions, up to 1M with opt-in context
Claude Sonnet 4.6 Anthropic Coding, agent planning, and drafting with strong long-context reasoning
Claude Opus 4.6 Anthropic Hardest agents, complex coding, and very long documents

Several newer models also handle very long context and are listed above under their primary use rather than repeated here: the GPT-5.6 family (Luna, Terra, and Sol) at 1.05M tokens, and Claude Opus 5 and Gemini 3.6 Flash at 1M. Any of them will take a large document.

Specialized

Single-purpose models. Availability depends on your organization's configuration.

Model Provider Best for
Sonar Perplexity Web-grounded answers
Sonar Reasoning Perplexity Structured analysis grounded in live search results
Sonar Pro Perplexity Multi-step research questions, broad sources, nuanced synthesis
Sonar Deep Research Perplexity Long reports, due diligence, and literature-style synthesis
Sonar Reasoning Pro Perplexity Deep multi-step reasoning over retrieved web evidence
o4-mini Deep Research OpenAI Faster, more affordable deep research
o3 Deep Research OpenAI Long reports, due diligence, and literature-style synthesis
DALL-E 3 OpenAI Creative illustrations, concepts, and marketing-style imagery
GPT Image 1.5 OpenAI Strong image quality, variations, and multimodal workflows at scale
GPT Image 2 OpenAI Latest image generation with rich detail, edits, and complex scenes

The Smart Model Router

Not a model: a dispatcher. It reads your question and assigns whichever model suits it, choosing only from those you already have access to. For most people doing most things, this is the right default.

Model retirements

Providers retire their own models on their own schedules, and a retired model is removed from AI CTRL when that happens. Expedient gives notice ahead of any retirement and works with your administrators through the change.

If you've built a custom model, check what it runs on. A custom model is locked to a single base model: when that base model retires, the custom model stops working until it's pointed at a replacement. Review anything your team depends on whenever a retirement is announced.

New models

New models are added roughly every two weeks, typically two to four weeks after a provider's public release, once they've been tested against our security requirements. They arrive restricted by default, so your administrators decide who gets access to each one.

Related articles