Skip to content

Models

Models are the AI Alfred can use. Your model choice affects answer quality, speed, usage cost, supported inputs, and how much context Alfred can consider.

You set your preferred model from the Settings Panel on the right of the Dashboard.

Use a stronger model for hard reasoning, coding, detailed writing, complex file analysis, and high-stakes planning.

Use a faster or lower-cost model for quick questions, drafts, simple summaries, and high-volume work.

Most people are best served by picking one capable model for daily use and switching to a cheaper one when they know they’re going to have a long or repetitive conversation.

Model availability depends on your plan. Plans are covered in Patreon.

ModelBest forInputsAvailability
GPT-5.6-TerraAdvanced reasoning, professional writing, complex analysisText, imagesElite and above
GPT-5.6-LunaLighter, faster work that still needs solid reasoningText, imagesPower and above
Claude-SonnetCoding, careful reasoning, long structured writingText, imagesElite and above
Claude-HaikuFast everyday work with good reasoningText, imagesPower and above
Gemini-ProLarge context, complex analysis, multimodal workText, images, audio, videoElite and above
Gemini-FlashFast, cost-effective general workText, images, audio, videoAll plans
Gemini-Flash-LiteLowest-cost conservation modeText, images, audio, videoAll plans
Grok-4.6General work with a more direct answering styleText, imagesPower and above
Grok-4.3The previous Grok generation, still availableText, imagesPower and above
DeepSeek-ProHigher-quality text-only workTextAll plans
DeepSeek-FlashFast text-only workTextAll plans
Minimax-M3General work with image and video inputText, images, videoPower and above
MiMo-V2.5Cheap multimodal workText, images, audio, videoAll plans
MiMo-V2.5-ProCheap text-only work with a bit more depthTextAll plans
GLM-5.2Text-only work, strong at structured outputTextPower and above

Model names, availability and the lineup itself change over time as new models release and old ones get sunset. The Dashboard model picker is the source of truth for your account because it only shows what your current plan can use. If a model is locked, the picker shows the plan you’d need for it.

Alfred also uses a few models internally that you don’t select, for example image generation models and the models his agents run on. Those don’t appear in the picker.

Every model has a context limit, which is how much of a conversation it can consider at once. Most models Alfred offers carry a large context, but a longer conversation still costs more per message because the whole thing gets sent along each time.

If a conversation gets very long, start a new one. It’s cheaper, and answers are usually better because Alfred isn’t carrying a pile of unrelated history. You can also turn on Memory Compression in the Settings Panel to keep long conversations lighter.

Your selected model is a preference, not a guarantee. When your usage gets high, Alfred may temporarily switch to a more cost-effective model to keep your daily allowance available for longer.

You may see Alfred mention that he switched models. This usually means you’re approaching a usage threshold. It does not mean your subscription changed, and it does not remove your preferred model setting.

Some models can work with images, audio, video, or large context better than others. Models that can’t process images are given a Vision Agent that mediates vision tasks on their behalf, so uploading an image to a text-only model still works, it just goes through an extra step.

Text-only models tend to be the cheapest, so if you never send images they’re often the better pick.

More capable models use more allowance. Larger prompts, long files, web research, generated outputs, and tool use all add to it too.

For simple work a cheaper model is often the better experience, because it’s faster and lets you do more before reaching your limit.

See Usage for how allowance works.