Models
Every model. One coworker.
Your coworker runs on 190+ models from every major lab. 70+ of them have open weights. You choose which ones your team may use.
In the catalogue today
- OpenAI47 models
- GPT-48K
- GPT-4 Turbo128K
- GPT-4.11M
- GPT-4.1 mini1M
- GPT-4.1 nano1M
- GPT-4o128K
- GPT-4o mini128K
- GPT-5400K
- GPT-5-Codex400K
- GPT-5 Mini400K
- GPT-5 Nano400K
- GPT-5 Pro400K
- GPT-5.1400K
- GPT-5.1-Codex400K
- GPT 5.1 Codex Max400K
- GPT-5.1 Codex mini400K
- GPT-5.1 Instant128K
- GPT 5.1 Thinking400K
- GPT-5.2400K
- GPT-5.2 Chat128K
- GPT-5.2-Codex400K
- GPT-5.2 Pro400K
- GPT-5.3 Chat128K
- GPT-5.3 Chat (latest)128K
- GPT-5.3 Codex400K
- GPT-5.3 Codex Spark128K
- GPT-5.41M
- GPT-5.4 mini400K
- GPT-5.4 nano400K
- GPT-5.4 Pro1M
- GPT-5.51M
- GPT-5.5 Pro1M
- GPT-5.61M
- GPT-5.6 Luna1M
- GPT-5.6 Sol1M
- GPT-5.6 Terra1M
- GPT OSS 20B131K
- GPT OSS 120B131K
- gpt-oss-safeguard-20b131K
- GPT-Realtime-2.1128K
- o1200K
- o1-pro200K
- o3200K
- o3-deep-research200K
- o3-mini200K
- o3-pro200K
- o4-mini200K
- Alibaba23 models
- Qwen 3 Coder 30B A3B Instruct262K
- Qwen 3 Max Thinking256K
- Qwen 3.5 Flash1M
- Qwen 3.5 Plus1M
- Qwen 3.6 27B256K
- Qwen 3.6 Max Preview240K
- Qwen 3.6 Plus1M
- Qwen 3.7 Max991K
- Qwen 3.7 Plus1M
- Qwen 3.32B128K
- Qwen3-14B40K
- Qwen3-30B-A3B40K
- Qwen3 235B A22B Instruct 2507262K
- Qwen3 235B A22B Thinking 2507131K
- Qwen3 Coder 480B A35B Instruct262K
- Qwen3 Coder Next256K
- Qwen3 Coder Plus1M
- Qwen3 Max262K
- Qwen3 Max Preview262K
- Qwen3 Next 80B A3B Instruct131K
- Qwen3 Next 80B A3B Thinking131K
- Qwen3 VL Instruct131K
- Qwen3 VL Thinking131K
- Anthropic16 models
- Claude Fable 51M
- Claude Haiku 3200K
- Claude Haiku 4.5200K
- Claude Opus 4200K
- Claude Opus 4.1200K
- Claude Opus 4.5200K
- Claude Opus 4.61M
- Claude Opus 4.71M
- Claude Opus 4.81M
- Claude Opus 4.8 (Fast)1M
- Claude Opus 51M
- Claude Opus 5 (Fast)1M
- Claude Sonnet 41M
- Claude Sonnet 4.51M
- Claude Sonnet 4.61M
- Claude Sonnet 51M
- Z.ai16 models
- GLM-4.5131K
- GLM-4.5-Air131K
- GLM-4.5-Flash131K
- GLM-4.5V64K
- GLM-4.6204K
- GLM-4.6V128K
- GLM-4.6V-Flash128K
- GLM-4.7204K
- GLM-4.7-Flash200K
- GLM-4.7-FlashX200K
- GLM-5204K
- GLM-5-Turbo200K
- GLM-5.1200K
- GLM-5.21M
- GLM 5.2 Fast1M
- GLM-5V-Turbo200K
- Google12 models
- Gemini 2.5 Flash1M
- Gemini 2.5 Flash Lite1M
- Gemini 2.5 Pro1M
- Gemini 3 Flash1M
- Gemini 3 Pro Preview1M
- Gemini 3.1 Flash Lite1M
- Gemini 3.1 Pro Preview1M
- Gemini 3.5 Flash1M
- Gemini 3.5 Flash Lite1M
- Gemini 3.6 Flash1M
- Gemma 4 26B A4B IT262K
- Gemma 4 31B IT262K
- Mistral12 models
- Codestral (latest)256K
- Devstral 2256K
- Devstral Small 2256K
- Magistral Medium (latest)128K
- Magistral Small128K
- Ministral 3B (latest)128K
- Ministral 8B (latest)128K
- Mistral Medium 3.1128K
- Mistral Medium Latest256K
- Mistral Nemo128K
- Mistral Small (latest)32K
- Pixtral 12B128K
- xAI11 models
- Grok 4.1 Fast Non-Reasoning1M
- Grok 4.1 Fast Reasoning1M
- Grok 4.31M
- Grok 4.5500K
- Grok 4.20 Beta Non-Reasoning2M
- Grok 4.20 Beta Reasoning2M
- Grok 4.20 Multi-Agent2M
- Grok 4.20 Multi Agent Beta2M
- Grok 4.20 Non-Reasoning2M
- Grok 4.20 Reasoning2M
- Grok Build 0.1256K
- MiniMax8 models
- MiniMax M2205K
- MiniMax M2.1204K
- MiniMax M2.1 Lightning204K
- MiniMax M2.5204K
- MiniMax M2.5 High Speed204K
- Minimax M2.7204K
- MiniMax M2.7 High Speed204K
- MiniMax M31M
- Moonshot AI8 models
- Kimi K2 Instruct131K
- Kimi K2 Thinking216K
- Kimi K2.5262K
- Kimi K2.6262K
- Kimi K2.7 Code256K
- Kimi K2.7 Code High Speed262K
- Kimi K31M
- Kimi K3 Fast1M
- DeepSeek7 models
- DeepSeek-R1128K
- DeepSeek V3 0324163K
- DeepSeek-V3.1163K
- DeepSeek V3.1 Terminus131K
- DeepSeek V3.2 Thinking128K
- DeepSeek V4 Flash1M
- DeepSeek V4 Pro1M
- Meta6 models
- Llama 3.1 8B Instruct128K
- Llama 3.1 70B Instruct128K
- Llama-3.3-70B-Instruct128K
- Llama-4-Maverick-17B-128E-Instruct-FP8128K
- Llama-4-Scout-17B-16E-Instruct-FP8128K
- Muse Spark 1.11M
- Amazon3 models
- Nova Lite300K
- Nova Micro128K
- Nova Pro300K
- Kwaipilot3 models
- Kat Coder Air V2.5256K
- Kat Coder Pro V2256K
- Kat Coder Pro V2.5256K
- NVIDIA3 models
- Nemotron 3 Ultra1M
- Nvidia Nemotron Nano 9B V2131K
- Nvidia Nemotron Nano 12B V2 VL131K
- ByteDance2 models
- Seed 1.6256K
- Seed 1.8256K
- Inception2 models
- Mercury 2128K
- Mercury Coder Small Beta32K
- Perplexity2 models
- Sonar127K
- Sonar Pro200K
- Poolside2 models
- Laguna S 2.11M
- Laguna S 2.1 Free256K
- StepFun2 models
- Step 3.7 Flash256K
- StepFun 3.5 Flash262K
- Xiaomi2 models
- MiMo M2.51M
- MiMo V2.5 Pro1M
- Arcee AI1 model
- Trinity Large Thinking262K
- Cohere1 model
- Command A256K
- InclusionAI1 model
- Ling 3.0 Flash256K
- Sakana AI1 model
- Fugu Ultra1M
- Tencent1 model
- Hy3262K
- Thinking Machines1 model
- Inkling256K
From the models.dev catalogue, counted on 25 August 2026. Every one of them calls tools. 150 think step by step. 119 read images.
One job, one engine
A different model for each job
The email tagger does not need the model that writes your board pack. Stop paying top rate for easy work.
The jobThe engine you setCost
- Tag 400 support emailsHigh volume, one right answerMinistral 8B (latest)Mistral
- Draft each replyYour tone, your refund rulesClaude Sonnet 5Anthropic
- Read the 90-page contractOne long document, one passGemini 3 Pro PreviewGoogle
- Write the board packHard, and the board reads itClaude Opus 5Anthropic
Four jobs, four engines, one coworker. Change any row and nothing else moves.Cheapest to dearest
Controls
Choose the model for the job
Let Alfera recommend a model, or choose one for the conversation.
Choose a model
Compare the available models in the picker.
Use a recommendation
Let Alfera choose a model for the work.
Switch when you need
Choose another model from the conversation composer.
Bring your own keys
Use your provider contract and the model bill goes to you, under your own agreement.
Default
- Claude Sonnet 5Anthropicdefault
Approved
- GPT-5.6OpenAI
- Gemini 3 Pro PreviewGoogle
- Llama 3.1 70B InstructMeta
- Ministral 8B (latest)Mistral
Blocked
- Every other modelblocked
Provider keys
Open weights
Open models, no lock-in
70+ of the models we ship have open weights. You can run them cheaper, and you can leave any lab that changes its terms.
- Llama
- Qwen
- DeepSeek
- GLM
- Kimi
- Mistral
- Gemma
- gpt-oss
The job is repetitive
Send it to an open model. The same work lands, and the bill is smaller.
You swap the engine
Your memory, your permissions and your connected tools stay where they were.
A lab rewrites its terms
You approve a different model and carry on. Nobody has to migrate anything.
Any models, including open-source
Start on the default. Move to any model your team approves. Your memory and permissions stay put.