Models and model selection
Choose from the current Chat catalog, inspect capabilities, and understand privacy and plan locks.
Open the model picker
The Chat composer model control is the source of truth for account and route availability. Selecting the model name beside the composer controls opens the picker. The picker groups rows by model family, supports provider filters and search, and opens an information view with the provider, release label, context and output fields, published prices when present, and capability indicators. A capability describes the selected catalog route; it does not guarantee identical behavior on every provider route.
Capabilities and access costs
Each table lists one model per row, grouped alphabetically by provider and model name. The icons describe the published catalog capabilities. ZDR means zero data retention. Privacy policies can differ between routes to the same model; an asterisk identifies policies available on only some routes.
Privacy icons: solid = all routes; * = some routes; slashed = not offered; ? = not published. Other icons appear when the catalog advertises that capability. Focus or hover over an icon for its label.
Access cost shows the published upstream API rates used to access these models, in US dollars. Token rates are per one million tokens; audio duration and character rates show their own units. These are model access rates, separate from AI Integrator subscription prices. Tiered context pricing, routing, caching, and additional tool charges can change the cost of a request. Open More rates for published cache and tier details.
Rates and capabilities come from the AI Gateway model catalog, refreshed hourly. A missing rate is marked Not published. Default identifies the application's configured default for that model type; account settings can override it.
Language models (LLM)
Language models power Chat conversations, including the Claude, GPT, Gemini, and other families listed below. The model picker determines availability for the account and its privacy settings. Installed Code runtimes have their own model selections.
| Model | Capabilities | Access cost USD |
|---|---|---|
| Alibaba / Qwen | ||
Qwen 3.6 Plusalibaba/qwen3.6-plus | Input$0.50 / 1M tokens Output$3.00 / 1M tokens TieredMore rates
| |
Qwen 3.7 Flashalibaba/qwen3.7-flash | Input$0.03 / 1M tokens Output$0.13 / 1M tokens TieredMore rates
| |
Qwen 3.7 Maxalibaba/qwen3.7-max | Input$2.50 / 1M tokens Output$7.50 / 1M tokens More rates
| |
Qwen 3.7 Plusalibaba/qwen3.7-plus | Input$0.40 / 1M tokens Output$1.60 / 1M tokens TieredMore rates
| |
Qwen 3.8 Flashalibaba/qwen3.8-flash | Input$0.16 / 1M tokens Output$0.47 / 1M tokens More rates
| |
Qwen 3.8 Maxalibaba/qwen3.8-max | Input$2.00 / 1M tokens Output$6.00 / 1M tokens More rates
| |
| Anthropic | ||
Claude Fable 5.1anthropic/claude-fable-5.1 | Input$10.00 / 1M tokens Output$50.00 / 1M tokens More rates
| |
Claude Haiku 4.5anthropic/claude-haiku-4.5 | Input$1.00 / 1M tokens Output$5.00 / 1M tokens More rates
| |
Claude Opus 4.5anthropic/claude-opus-4.5 | Input$5.00 / 1M tokens Output$25.00 / 1M tokens More rates
| |
Claude Opus 4.6anthropic/claude-opus-4.6 | Input$5.00 / 1M tokens Output$25.00 / 1M tokens More rates
| |
Claude Opus 4.7anthropic/claude-opus-4.7 | Input$5.00 / 1M tokens Output$25.00 / 1M tokens More rates
| |
Claude Opus 4.8anthropic/claude-opus-4.8 | Input$5.00 / 1M tokens Output$25.00 / 1M tokens More rates
| |
Claude Opus 5anthropic/claude-opus-5 | Input$5.00 / 1M tokens Output$25.00 / 1M tokens More rates
| |
Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Input$3.00 / 1M tokens Output$15.00 / 1M tokens TieredMore rates
| |
Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Input$3.00 / 1M tokens Output$15.00 / 1M tokens More rates
| |
Claude Sonnet 5Defaultanthropic/claude-sonnet-5 | Input$2.00 / 1M tokens Output$10.00 / 1M tokens More rates
| |
| DeepSeek | ||
DeepSeek V3.2deepseek/deepseek-v3.2 | * | Input$0.62 / 1M tokens Output$1.85 / 1M tokens Varies by provider |
DeepSeek V3.2 Thinkingdeepseek/deepseek-v3.2-thinking | * | Input$0.62 / 1M tokens Output$1.85 / 1M tokens Varies by provider |
DeepSeek V4 Flashdeepseek/deepseek-v4-flash | ** | Input$0.13 / 1M tokens Output$0.26 / 1M tokens Varies by providerMore rates
|
DeepSeek V4 Prodeepseek/deepseek-v4-pro | ** | Input$0.66 / 1M tokens Output$1.98 / 1M tokens Varies by providerMore rates
|
Gemini 2.5 Flashgoogle/gemini-2.5-flash | * | Input$0.30 / 1M tokens Output$2.50 / 1M tokens More rates
|
Gemini 2.5 Flash Litegoogle/gemini-2.5-flash-lite | * | Input$0.10 / 1M tokens Output$0.40 / 1M tokens More rates
|
Gemini 2.5 Progoogle/gemini-2.5-pro | * | Input$1.25 / 1M tokens Output$10.00 / 1M tokens TieredMore rates
|
Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | * | Input$0.25 / 1M tokens Output$1.50 / 1M tokens More rates
|
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | * | Input$2.00 / 1M tokens Output$12.00 / 1M tokens TieredMore rates
|
Gemini 3.5 Flashgoogle/gemini-3.5-flash | * | Input$1.50 / 1M tokens Output$9.00 / 1M tokens More rates
|
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | * | Input$0.30 / 1M tokens Output$2.50 / 1M tokens More rates
|
Gemini 3.6 Flashgoogle/gemini-3.6-flash | * | Input$0.75 / 1M tokens Output$3.75 / 1M tokens More rates
|
Gemini 3.7 Flashgoogle/gemini-3.7-flash | * | Input$0.75 / 1M tokens Output$3.75 / 1M tokens More rates
|
Gemini 3.8 Flashgoogle/gemini-3.8-flash | * | Input$0.75 / 1M tokens Output$3.75 / 1M tokens More rates
|
| Meta | ||
Muse Glimmer 30Bmeta/muse-glimmer-30b | Input$0.35 / 1M tokens Output$1.50 / 1M tokens Varies by providerMore rates
| |
Muse Spark 1.1meta/muse-spark-1.1 | Input$1.25 / 1M tokens Output$4.25 / 1M tokens More rates
| |
Muse Spark 1.2meta/muse-spark-1.2 | Input$1.25 / 1M tokens Output$4.25 / 1M tokens More rates
| |
Muse Spark 1.2 Contributormeta/muse-spark-1.2-contributor | Input$0.10 / 1M tokens Output$0.20 / 1M tokens More rates
| |
Muse Spark 1.3meta/muse-spark-1.3 | Input$1.25 / 1M tokens Output$4.25 / 1M tokens More rates
| |
Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor | Input$0.10 / 1M tokens Output$0.20 / 1M tokens More rates
| |
| MiniMax | ||
MiniMax M2.7minimax/minimax-m2.7 | * | Input$0.30 / 1M tokens Output$1.20 / 1M tokens More rates
|
MiniMax M3minimax/minimax-m3 | ** | Input$0.30 / 1M tokens Output$1.20 / 1M tokens Varies by providerMore rates
|
| Mistral | ||
Devstral 2mistral/devstral-2 | Input$0.40 / 1M tokens Output$2.00 / 1M tokens | |
Devstral Small 2mistral/devstral-small-2 | Input$0.10 / 1M tokens Output$0.30 / 1M tokens | |
Ministral 14Bmistral/ministral-14b | Input$0.20 / 1M tokens Output$0.20 / 1M tokens | |
Mistral Large 3mistral/mistral-large-3 | Input$0.50 / 1M tokens Output$1.50 / 1M tokens | |
Mistral Medium Latestmistral/mistral-medium-3.5 | Input$1.50 / 1M tokens Output$7.50 / 1M tokens | |
| Moonshot AI | ||
Kimi K2.5moonshotai/kimi-k2.5 | * | Input$0.60 / 1M tokens Output$3.00 / 1M tokens More rates
|
Kimi K2.6moonshotai/kimi-k2.6 | * | Input$0.95 / 1M tokens Output$4.00 / 1M tokens More rates
|
Kimi K2.7 Codemoonshotai/kimi-k2.7-code | Input$0.95 / 1M tokens Output$4.00 / 1M tokens Varies by providerMore rates
| |
Kimi K3moonshotai/kimi-k3 | ** | Input$3.00 / 1M tokens Output$15.00 / 1M tokens Varies by providerMore rates
|
| NVIDIA | ||
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | Input$0.60 / 1M tokens Output$2.40 / 1M tokens Varies by providerMore rates
| |
| OpenAI | ||
GPT 5.4openai/gpt-5.4 | * | Input$2.50 / 1M tokens Output$15.00 / 1M tokens TieredMore rates
|
GPT 5.4 Miniopenai/gpt-5.4-mini | * | Input$0.75 / 1M tokens Output$4.50 / 1M tokens More rates
|
GPT 5.4 Nanoopenai/gpt-5.4-nano | * | Input$0.20 / 1M tokens Output$1.25 / 1M tokens More rates
|
GPT 5.5openai/gpt-5.5 | * | Input$5.00 / 1M tokens Output$30.00 / 1M tokens Tiered · Varies by providerMore rates
|
GPT 5.6 Lunaopenai/gpt-5.6-luna | * | Input$0.20 / 1M tokens Output$1.20 / 1M tokens TieredMore rates
|
GPT 5.6 Solopenai/gpt-5.6-sol | * | Input$2.00 / 1M tokens Output$10.00 / 1M tokens Tiered · Varies by providerMore rates
|
GPT 5.6 Terraopenai/gpt-5.6-terra | * | Input$2.00 / 1M tokens Output$12.00 / 1M tokens TieredMore rates
|
GPT OSS 120Bopenai/gpt-oss-120b | Input$0.10 / 1M tokens Output$0.50 / 1M tokens Varies by provider | |
GPT-4.1openai/gpt-4.1 | * | Input$2.00 / 1M tokens Output$8.00 / 1M tokens More rates
|
GPT-4.1 miniopenai/gpt-4.1-mini | * | Input$0.40 / 1M tokens Output$1.60 / 1M tokens More rates
|
GPT-4.1 nanoopenai/gpt-4.1-nano | * | Input$0.10 / 1M tokens Output$0.40 / 1M tokens Varies by providerMore rates
|
GPT-4oopenai/gpt-4o | * | Input$2.50 / 1M tokens Output$10.00 / 1M tokens More rates
|
GPT-4o miniopenai/gpt-4o-mini | * | Input$0.15 / 1M tokens Output$0.60 / 1M tokens More rates
|
GPT-6 Astraopenai/gpt-6-astra | * | Input$10.00 / 1M tokens Output$50.00 / 1M tokens TieredMore rates
|
o3openai/o3 | Input$2.00 / 1M tokens Output$8.00 / 1M tokens More rates
| |
o4-miniopenai/o4-mini | * | Input$1.10 / 1M tokens Output$4.40 / 1M tokens More rates
|
| Thinking Machines | ||
Inklingthinkingmachines/inkling | ** | Input$1.00 / 1M tokens Output$4.05 / 1M tokens Varies by providerMore rates
|
Inkling Smallthinkingmachines/inkling-small | ** | Input$0.50 / 1M tokens Output$1.20 / 1M tokens Varies by providerMore rates
|
| xAI | ||
Grok 4.1 Fast Non-Reasoningspacexai/grok-4.1-fast-non-reasoning | Input$0.20 / 1M tokens Output$0.50 / 1M tokens More rates
| |
Grok 4.1 Fast Reasoningspacexai/grok-4.1-fast-reasoning | Input$0.20 / 1M tokens Output$0.50 / 1M tokens More rates
| |
Grok 4.3spacexai/grok-4.3 | Input$1.25 / 1M tokens Output$2.50 / 1M tokens TieredMore rates
| |
Grok 4.5spacexai/grok-4.5 | Input$2.00 / 1M tokens Output$6.00 / 1M tokens TieredMore rates
| |
Grok 4.6spacexai/grok-4.6 | Input$2.00 / 1M tokens Output$6.00 / 1M tokens TieredMore rates
| |
Grok 4.20 Non-Reasoningspacexai/grok-4.20-non-reasoning | Input$1.25 / 1M tokens Output$2.50 / 1M tokens Tiered · Varies by providerMore rates
| |
Grok 4.20 Reasoningspacexai/grok-4.20-reasoning | Input$1.25 / 1M tokens Output$2.50 / 1M tokens Tiered · Varies by providerMore rates
| |
| Z.ai | ||
GLM 4.7 Flashzai/glm-4.7-flash | ** | Input$0.07 / 1M tokens Output$0.40 / 1M tokens |
GLM 5zai/glm-5 | ** | Input$1.00 / 1M tokens Output$3.20 / 1M tokens |
GLM 5.1zai/glm-5.1 | ** | Input$1.40 / 1M tokens Output$4.40 / 1M tokens More rates
|
GLM 5.2zai/glm-5.2 | ** | Input$0.80 / 1M tokens Output$2.55 / 1M tokens Varies by providerMore rates
|
GLM 5.3zai/glm-5.3 | ** | Input$1.40 / 1M tokens Output$4.40 / 1M tokens Varies by providerMore rates
|
GLM 5.3 Flashzai/glm-5.3-flash | ** | Input$0.15 / 1M tokens Output$0.50 / 1M tokens Varies by providerMore rates
|
Speech to text (STT)
Transcription models turn recorded speech into text. Choose a speech-to-text model in Settings → Speech & Voice. The Live audio icon identifies a streaming audio route in the catalog; each application's recording controls determine how it is used.
| Model | Capabilities | Access cost USD |
|---|---|---|
| Fish Audio | ||
Transcribe-1fish-audio/transcribe-1 | Catalog rate$0.00 · listed as free | |
Gemini 3.5 Transcribegoogle/gemini-3.5-transcribe | Input$2.00 / 1M tokens Audio input$2.00 / 1M tokens Output$12.00 / 1M tokens | |
Gemini 3.5 Transcribe Livegoogle/gemini-3.5-transcribe-live | Audio$0.009 / minute | |
| OpenAI | ||
GPT-4o mini Transcribeopenai/gpt-4o-mini-transcribe | Input$1.25 / 1M tokens Audio input$1.25 / 1M tokens Output$5.00 / 1M tokens | |
GPT-4o Transcribeopenai/gpt-4o-transcribe | Input$2.50 / 1M tokens Audio input$2.50 / 1M tokens Output$10.00 / 1M tokens | |
gpt-realtime-whisperopenai/gpt-realtime-whisper | Audio$0.01704 / minute | |
Whisperopenai/whisper-1 | Audio$0.006 / minute | |
| xAI | ||
Grok STTDefaultspacexai/grok-stt | Audio$0.00168 / minute | |
Text to speech (TTS)
Speech models generate audio for read aloud. Choose the model and an available voice in Settings → Speech & Voice. Character-based rates apply to the input text.
| Model | Capabilities | Access cost USD |
|---|---|---|
| Fish Audio | ||
S1fish-audio/s1 | Catalog rate$0.00 · listed as free | |
S2 Profish-audio/s2-pro | Catalog rate$0.00 · listed as free | |
S2.1 Profish-audio/s2.1-pro | Catalog rate$0.00 · listed as free | |
| OpenAI | ||
TTS-1Defaultopenai/tts-1 | Text$15.00 / 1M characters | |
TTS-1 HDopenai/tts-1-hd | Text$30.00 / 1M characters | |
| xAI | ||
Grok TTSspacexai/grok-tts | Text$15.00 / 1M characters | |
Embedding models
Embedding models support retrieval across memory, conversation history, and indexed documents. Their rates apply to the text being embedded. Changing the embedding model can require existing content to be indexed again; see Memory.
| Model | Capabilities | Access cost USD |
|---|---|---|
| Alibaba / Qwen | ||
Qwen3 Embedding 4BDefaultalibaba/qwen3-embedding-4b | Input$0.02 / 1M tokens | |
| Amazon | ||
Titan Text Embeddings V2amazon/titan-embed-text-v2 | Input$0.02 / 1M tokens | |
Text Embedding 005google/text-embedding-005 | Input$0.025 / 1M tokens | |
| OpenAI | ||
text-embedding-3-smallopenai/text-embedding-3-small | * | Input$0.02 / 1M tokens |
| Voyage AI | ||
Voyage 3.5voyage/voyage-3.5 | Input$0.06 / 1M tokens | |
Voyage 3.5 Litevoyage/voyage-3.5-lite | Input$0.02 / 1M tokens | |
Voyage 4voyage/voyage-4 | Input$0.06 / 1M tokens | |
Voyage 4 Largevoyage/voyage-4-large | Input$0.12 / 1M tokens | |
Voyage 4 Litevoyage/voyage-4-lite | Input$0.02 / 1M tokens | |
Reranking models
Reranking models score retrieved results for relevance. These catalog entries are separate from the Chat model picker; inclusion in the catalog does not mean every retrieval uses a reranker.
| Model | Capabilities | Access cost USD |
|---|---|---|
| Cohere | ||
Cohere Rerank 3.5cohere/rerank-v3.5 | Not published | |
Reasoning and privacy locks
The reasoning slider folds down to None for models without reasoning support. Privacy filters can require Zero Data Retention or no training; rows without a compatible route are locked. Premium rows can require Pro or Max. Fast is a separate route when the catalog publishes one and can have different metering. See Usage, Privacy, Voice, and Settings.