Skip to content
AI IntegratorDocumentation
Documentation/Configuration
Configuration

Models and model selection

Choose from the current Chat catalog, inspect capabilities, and understand privacy and plan locks.

Open the model picker

The Chat composer model control is the source of truth for account and route availability. Selecting the model name beside the composer controls opens the picker. The picker groups rows by model family, supports provider filters and search, and opens an information view with the provider, release label, context and output fields, published prices when present, and capability indicators. A capability describes the selected catalog route; it does not guarantee identical behavior on every provider route.

Capabilities and access costs

Each table lists one model per row, grouped alphabetically by provider and model name. The icons describe the published catalog capabilities. ZDR means zero data retention. Privacy policies can differ between routes to the same model; an asterisk identifies policies available on only some routes.

Tool useReasoningImage inputFile / PDF inputPrompt cachingLive audioZero data retentionNo training

Privacy icons: solid = all routes; * = some routes; slashed = not offered; ? = not published. Other icons appear when the catalog advertises that capability. Focus or hover over an icon for its label.

Access cost shows the published upstream API rates used to access these models, in US dollars. Token rates are per one million tokens; audio duration and character rates show their own units. These are model access rates, separate from AI Integrator subscription prices. Tiered context pricing, routing, caching, and additional tool charges can change the cost of a request. Open More rates for published cache and tier details.

Rates and capabilities come from the AI Gateway model catalog, refreshed hourly. A missing rate is marked Not published. Default identifies the application's configured default for that model type; account settings can override it.

Language models (LLM)

Language models power Chat conversations, including the Claude, GPT, Gemini, and other families listed below. The model picker determines availability for the account and its privacy settings. Installed Code runtimes have their own model selections.

ModelCapabilitiesAccess cost USD
Alibaba / Qwen
Qwen 3.6 Plusalibaba/qwen3.6-plus
Input$0.50 / 1M tokens
Output$3.00 / 1M tokens
Tiered
More rates
Cached input
$0.10 / 1M tokens
Cache write
$0.625 / 1M tokens
Input, 0–256,000 context tokens
$0.50 / 1M tokens
Input, 256,000+ context tokens
$2.00 / 1M tokens
Output, 0–256,000 context tokens
$3.00 / 1M tokens
Output, 256,000+ context tokens
$6.00 / 1M tokens
Cached input, 0–256,000 context tokens
$0.10 / 1M tokens
Cached input, 256,000+ context tokens
$0.20 / 1M tokens
Cache write, 0–256,000 context tokens
$0.625 / 1M tokens
Cache write, 256,000+ context tokens
$2.50 / 1M tokens
Qwen 3.7 Flashalibaba/qwen3.7-flash
Input$0.03 / 1M tokens
Output$0.13 / 1M tokens
Tiered
More rates
Cached input
$0.006 / 1M tokens
Cache write
$0.038 / 1M tokens
Input, 0–32,000 context tokens
$0.03 / 1M tokens
Input, 32,000–256,000 context tokens
$0.10 / 1M tokens
Input, 256,000+ context tokens
$0.20 / 1M tokens
Output, 0–32,000 context tokens
$0.13 / 1M tokens
Output, 32,000–256,000 context tokens
$0.40 / 1M tokens
Output, 256,000+ context tokens
$0.80 / 1M tokens
Cached input, 0–32,000 context tokens
$0.006 / 1M tokens
Cached input, 32,000–256,000 context tokens
$0.02 / 1M tokens
Cached input, 256,000+ context tokens
$0.04 / 1M tokens
Cache write, 0–32,000 context tokens
$0.038 / 1M tokens
Cache write, 32,000–256,000 context tokens
$0.125 / 1M tokens
Cache write, 256,000+ context tokens
$0.25 / 1M tokens
Qwen 3.7 Maxalibaba/qwen3.7-max
Input$2.50 / 1M tokens
Output$7.50 / 1M tokens
More rates
Cached input
$0.50 / 1M tokens
Cache write
$3.125 / 1M tokens
Qwen 3.7 Plusalibaba/qwen3.7-plus
Input$0.40 / 1M tokens
Output$1.60 / 1M tokens
Tiered
More rates
Cached input
$0.08 / 1M tokens
Cache write
$0.50 / 1M tokens
Input, 0–256,000 context tokens
$0.40 / 1M tokens
Input, 256,000+ context tokens
$1.20 / 1M tokens
Output, 0–256,000 context tokens
$1.60 / 1M tokens
Output, 256,000+ context tokens
$4.80 / 1M tokens
Cached input, 0–256,000 context tokens
$0.08 / 1M tokens
Cached input, 256,000+ context tokens
$0.24 / 1M tokens
Cache write, 0–256,000 context tokens
$0.50 / 1M tokens
Cache write, 256,000+ context tokens
$1.50 / 1M tokens
Qwen 3.8 Flashalibaba/qwen3.8-flash
Input$0.16 / 1M tokens
Output$0.47 / 1M tokens
More rates
Cached input
$0.016 / 1M tokens
Cache write
$0.20 / 1M tokens
Qwen 3.8 Maxalibaba/qwen3.8-max
Input$2.00 / 1M tokens
Output$6.00 / 1M tokens
More rates
Cached input
$0.25 / 1M tokens
Cache write
$2.50 / 1M tokens
Anthropic
Claude Fable 5.1anthropic/claude-fable-5.1
Input$10.00 / 1M tokens
Output$50.00 / 1M tokens
More rates
Cached input
$0.25 / 1M tokens
Cache write
$12.50 / 1M tokens
Web search
$10.00 / request
Claude Haiku 4.5anthropic/claude-haiku-4.5
Input$1.00 / 1M tokens
Output$5.00 / 1M tokens
More rates
Cached input
$0.10 / 1M tokens
Cache write
$1.25 / 1M tokens
Web search
$10.00 / request
Claude Opus 4.5anthropic/claude-opus-4.5
Input$5.00 / 1M tokens
Output$25.00 / 1M tokens
More rates
Cached input
$0.50 / 1M tokens
Cache write
$6.25 / 1M tokens
Web search
$10.00 / request
Claude Opus 4.6anthropic/claude-opus-4.6
Input$5.00 / 1M tokens
Output$25.00 / 1M tokens
More rates
Cached input
$0.50 / 1M tokens
Cache write
$6.25 / 1M tokens
Web search
$10.00 / request
Claude Opus 4.7anthropic/claude-opus-4.7
Input$5.00 / 1M tokens
Output$25.00 / 1M tokens
More rates
Cached input
$0.50 / 1M tokens
Cache write
$6.25 / 1M tokens
Web search
$10.00 / request
Claude Opus 4.8anthropic/claude-opus-4.8
Input$5.00 / 1M tokens
Output$25.00 / 1M tokens
More rates
Cached input
$0.50 / 1M tokens
Cache write
$6.25 / 1M tokens
Web search
$10.00 / request
Claude Opus 5anthropic/claude-opus-5
Input$5.00 / 1M tokens
Output$25.00 / 1M tokens
More rates
Cached input
$0.50 / 1M tokens
Cache write
$6.25 / 1M tokens
Web search
$10.00 / request
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Input$3.00 / 1M tokens
Output$15.00 / 1M tokens
Tiered
More rates
Cached input
$0.30 / 1M tokens
Cache write
$3.75 / 1M tokens
Web search
$10.00 / request
Input, 0–200,001 context tokens
$3.00 / 1M tokens
Input, 200,001+ context tokens
$6.00 / 1M tokens
Output, 0–200,001 context tokens
$15.00 / 1M tokens
Output, 200,001+ context tokens
$22.50 / 1M tokens
Cached input, 0–200,001 context tokens
$0.30 / 1M tokens
Cached input, 200,001+ context tokens
$0.60 / 1M tokens
Cache write, 0–200,001 context tokens
$3.75 / 1M tokens
Cache write, 200,001+ context tokens
$7.50 / 1M tokens
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Input$3.00 / 1M tokens
Output$15.00 / 1M tokens
More rates
Cached input
$0.30 / 1M tokens
Cache write
$3.75 / 1M tokens
Web search
$10.00 / request
Claude Sonnet 5Defaultanthropic/claude-sonnet-5
Input$2.00 / 1M tokens
Output$10.00 / 1M tokens
More rates
Cached input
$0.20 / 1M tokens
Cache write
$2.50 / 1M tokens
Web search
$10.00 / request
DeepSeek
DeepSeek V3.2deepseek/deepseek-v3.2
Input$0.62 / 1M tokens
Output$1.85 / 1M tokens
Varies by provider
DeepSeek V3.2 Thinkingdeepseek/deepseek-v3.2-thinking
Input$0.62 / 1M tokens
Output$1.85 / 1M tokens
Varies by provider
DeepSeek V4 Flashdeepseek/deepseek-v4-flash
Input$0.13 / 1M tokens
Output$0.26 / 1M tokens
Varies by provider
More rates
Cached input
$0.028 / 1M tokens
DeepSeek V4 Prodeepseek/deepseek-v4-pro
Input$0.66 / 1M tokens
Output$1.98 / 1M tokens
Varies by provider
More rates
Cached input
$0.022 / 1M tokens
Google
Gemini 2.5 Flashgoogle/gemini-2.5-flash
Input$0.30 / 1M tokens
Output$2.50 / 1M tokens
More rates
Cached input
$0.03 / 1M tokens
Web search
$35.00 / request
Gemini 2.5 Flash Litegoogle/gemini-2.5-flash-lite
Input$0.10 / 1M tokens
Output$0.40 / 1M tokens
More rates
Cached input
$0.01 / 1M tokens
Web search
$35.00 / request
Gemini 2.5 Progoogle/gemini-2.5-pro
Input$1.25 / 1M tokens
Output$10.00 / 1M tokens
Tiered
More rates
Cached input
$0.125 / 1M tokens
Web search
$35.00 / request
Input, 0–200,001 context tokens
$1.25 / 1M tokens
Input, 200,001+ context tokens
$2.50 / 1M tokens
Output, 0–200,001 context tokens
$10.00 / 1M tokens
Output, 200,001+ context tokens
$15.00 / 1M tokens
Cached input, 0–200,001 context tokens
$0.125 / 1M tokens
Cached input, 200,001+ context tokens
$0.25 / 1M tokens
Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite
Input$0.25 / 1M tokens
Output$1.50 / 1M tokens
More rates
Cached input
$0.03 / 1M tokens
Web search
$14.00 / request
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Input$2.00 / 1M tokens
Output$12.00 / 1M tokens
Tiered
More rates
Cached input
$0.20 / 1M tokens
Web search
$14.00 / request
Input, 0–200,001 context tokens
$2.00 / 1M tokens
Input, 200,001+ context tokens
$4.00 / 1M tokens
Output, 0–200,001 context tokens
$12.00 / 1M tokens
Output, 200,001+ context tokens
$18.00 / 1M tokens
Cached input, 0–200,001 context tokens
$0.20 / 1M tokens
Cached input, 200,001+ context tokens
$0.40 / 1M tokens
Gemini 3.5 Flashgoogle/gemini-3.5-flash
Input$1.50 / 1M tokens
Output$9.00 / 1M tokens
More rates
Cached input
$0.15 / 1M tokens
Web search
$14.00 / request
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Input$0.30 / 1M tokens
Output$2.50 / 1M tokens
More rates
Cached input
$0.03 / 1M tokens
Web search
$14.00 / request
Gemini 3.6 Flashgoogle/gemini-3.6-flash
Input$0.75 / 1M tokens
Output$3.75 / 1M tokens
More rates
Cached input
$0.075 / 1M tokens
Web search
$14.00 / request
Gemini 3.7 Flashgoogle/gemini-3.7-flash
Input$0.75 / 1M tokens
Output$3.75 / 1M tokens
More rates
Cached input
$0.075 / 1M tokens
Web search
$14.00 / request
Gemini 3.8 Flashgoogle/gemini-3.8-flash
Input$0.75 / 1M tokens
Output$3.75 / 1M tokens
More rates
Cached input
$0.075 / 1M tokens
Web search
$14.00 / request
Meta
Muse Glimmer 30Bmeta/muse-glimmer-30b
Input$0.35 / 1M tokens
Output$1.50 / 1M tokens
Varies by provider
More rates
Cached input
$0.04 / 1M tokens
Muse Spark 1.1meta/muse-spark-1.1
Input$1.25 / 1M tokens
Output$4.25 / 1M tokens
More rates
Cached input
$0.15 / 1M tokens
Muse Spark 1.2meta/muse-spark-1.2
Input$1.25 / 1M tokens
Output$4.25 / 1M tokens
More rates
Cached input
$0.15 / 1M tokens
Muse Spark 1.2 Contributormeta/muse-spark-1.2-contributor
Input$0.10 / 1M tokens
Output$0.20 / 1M tokens
More rates
Cached input
$0.002 / 1M tokens
Muse Spark 1.3meta/muse-spark-1.3
Input$1.25 / 1M tokens
Output$4.25 / 1M tokens
More rates
Cached input
$0.15 / 1M tokens
Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor
Input$0.10 / 1M tokens
Output$0.20 / 1M tokens
More rates
Cached input
$0.002 / 1M tokens
MiniMax
MiniMax M2.7minimax/minimax-m2.7
Input$0.30 / 1M tokens
Output$1.20 / 1M tokens
More rates
Cached input
$0.06 / 1M tokens
Cache write
$0.375 / 1M tokens
MiniMax M3minimax/minimax-m3
Input$0.30 / 1M tokens
Output$1.20 / 1M tokens
Varies by provider
More rates
Cached input
$0.06 / 1M tokens
Mistral
Devstral 2mistral/devstral-2
Input$0.40 / 1M tokens
Output$2.00 / 1M tokens
Devstral Small 2mistral/devstral-small-2
Input$0.10 / 1M tokens
Output$0.30 / 1M tokens
Ministral 14Bmistral/ministral-14b
Input$0.20 / 1M tokens
Output$0.20 / 1M tokens
Mistral Large 3mistral/mistral-large-3
Input$0.50 / 1M tokens
Output$1.50 / 1M tokens
Mistral Medium Latestmistral/mistral-medium-3.5
Input$1.50 / 1M tokens
Output$7.50 / 1M tokens
Moonshot AI
Kimi K2.5moonshotai/kimi-k2.5
Input$0.60 / 1M tokens
Output$3.00 / 1M tokens
More rates
Cached input
$0.10 / 1M tokens
Kimi K2.6moonshotai/kimi-k2.6
Input$0.95 / 1M tokens
Output$4.00 / 1M tokens
More rates
Cached input
$0.16 / 1M tokens
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
Input$0.95 / 1M tokens
Output$4.00 / 1M tokens
Varies by provider
More rates
Cached input
$0.16 / 1M tokens
Kimi K3moonshotai/kimi-k3
Input$3.00 / 1M tokens
Output$15.00 / 1M tokens
Varies by provider
More rates
Cached input
$0.30 / 1M tokens
NVIDIA
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
Input$0.60 / 1M tokens
Output$2.40 / 1M tokens
Varies by provider
More rates
Cached input
$0.12 / 1M tokens
OpenAI
GPT 5.4openai/gpt-5.4
Input$2.50 / 1M tokens
Output$15.00 / 1M tokens
Tiered
More rates
Cached input
$0.25 / 1M tokens
Web search
$10.00 / request
Input, 0–272,000 context tokens
$2.50 / 1M tokens
Input, 272,000+ context tokens
$5.00 / 1M tokens
Output, 0–272,000 context tokens
$15.00 / 1M tokens
Output, 272,000+ context tokens
$22.50 / 1M tokens
Cached input, 0–272,000 context tokens
$0.25 / 1M tokens
Cached input, 272,000+ context tokens
$0.50 / 1M tokens
GPT 5.4 Miniopenai/gpt-5.4-mini
Input$0.75 / 1M tokens
Output$4.50 / 1M tokens
More rates
Cached input
$0.075 / 1M tokens
Web search
$10.00 / request
GPT 5.4 Nanoopenai/gpt-5.4-nano
Input$0.20 / 1M tokens
Output$1.25 / 1M tokens
More rates
Cached input
$0.02 / 1M tokens
Web search
$10.00 / request
GPT 5.5openai/gpt-5.5
Input$5.00 / 1M tokens
Output$30.00 / 1M tokens
Tiered · Varies by provider
More rates
Cached input
$0.50 / 1M tokens
Web search
$10.00 / request
Input, 0–272,000 context tokens
$5.00 / 1M tokens
Input, 272,000+ context tokens
$10.00 / 1M tokens
Output, 0–272,000 context tokens
$30.00 / 1M tokens
Output, 272,000+ context tokens
$45.00 / 1M tokens
Cached input, 0–272,000 context tokens
$0.50 / 1M tokens
Cached input, 272,000+ context tokens
$1.00 / 1M tokens
GPT 5.6 Lunaopenai/gpt-5.6-luna
Input$0.20 / 1M tokens
Output$1.20 / 1M tokens
Tiered
More rates
Cached input
$0.02 / 1M tokens
Cache write
$0.25 / 1M tokens
Web search
$10.00 / request
Input, 0–272,000 context tokens
$0.20 / 1M tokens
Input, 272,000+ context tokens
$0.40 / 1M tokens
Output, 0–272,000 context tokens
$1.20 / 1M tokens
Output, 272,000+ context tokens
$1.80 / 1M tokens
Cached input, 0–272,000 context tokens
$0.02 / 1M tokens
Cached input, 272,000+ context tokens
$0.04 / 1M tokens
Cache write, 0–272,000 context tokens
$0.25 / 1M tokens
Cache write, 272,000+ context tokens
$0.50 / 1M tokens
GPT 5.6 Solopenai/gpt-5.6-sol
Input$2.00 / 1M tokens
Output$10.00 / 1M tokens
Tiered · Varies by provider
More rates
Cached input
$0.20 / 1M tokens
Cache write
$2.50 / 1M tokens
Web search
$10.00 / request
Input, 0–272,000 context tokens
$2.00 / 1M tokens
Input, 272,000+ context tokens
$4.00 / 1M tokens
Output, 0–272,000 context tokens
$10.00 / 1M tokens
Output, 272,000+ context tokens
$15.00 / 1M tokens
Cached input, 0–272,000 context tokens
$0.20 / 1M tokens
Cached input, 272,000+ context tokens
$0.40 / 1M tokens
Cache write, 0–272,000 context tokens
$2.50 / 1M tokens
Cache write, 272,000+ context tokens
$5.00 / 1M tokens
GPT 5.6 Terraopenai/gpt-5.6-terra
Input$2.00 / 1M tokens
Output$12.00 / 1M tokens
Tiered
More rates
Cached input
$0.20 / 1M tokens
Cache write
$2.50 / 1M tokens
Web search
$10.00 / request
Input, 0–272,000 context tokens
$2.00 / 1M tokens
Input, 272,000+ context tokens
$4.00 / 1M tokens
Output, 0–272,000 context tokens
$12.00 / 1M tokens
Output, 272,000+ context tokens
$18.00 / 1M tokens
Cached input, 0–272,000 context tokens
$0.20 / 1M tokens
Cached input, 272,000+ context tokens
$0.40 / 1M tokens
Cache write, 0–272,000 context tokens
$2.50 / 1M tokens
Cache write, 272,000+ context tokens
$5.00 / 1M tokens
GPT OSS 120Bopenai/gpt-oss-120b
Input$0.10 / 1M tokens
Output$0.50 / 1M tokens
Varies by provider
GPT-4.1openai/gpt-4.1
Input$2.00 / 1M tokens
Output$8.00 / 1M tokens
More rates
Cached input
$0.50 / 1M tokens
Web search
$10.00 / request
GPT-4.1 miniopenai/gpt-4.1-mini
Input$0.40 / 1M tokens
Output$1.60 / 1M tokens
More rates
Cached input
$0.10 / 1M tokens
Web search
$10.00 / request
GPT-4.1 nanoopenai/gpt-4.1-nano
Input$0.10 / 1M tokens
Output$0.40 / 1M tokens
Varies by provider
More rates
Cached input
$0.025 / 1M tokens
Web search
$10.00 / request
GPT-4oopenai/gpt-4o
Input$2.50 / 1M tokens
Output$10.00 / 1M tokens
More rates
Cached input
$1.25 / 1M tokens
Web search
$10.00 / request
GPT-4o miniopenai/gpt-4o-mini
Input$0.15 / 1M tokens
Output$0.60 / 1M tokens
More rates
Cached input
$0.075 / 1M tokens
Web search
$10.00 / request
GPT-6 Astraopenai/gpt-6-astra
Input$10.00 / 1M tokens
Output$50.00 / 1M tokens
Tiered
More rates
Cached input
$1.00 / 1M tokens
Cache write
$12.50 / 1M tokens
Web search
$10.00 / request
Input, 0–272,001 context tokens
$10.00 / 1M tokens
Input, 272,001+ context tokens
$20.00 / 1M tokens
Output, 0–272,001 context tokens
$50.00 / 1M tokens
Output, 272,001+ context tokens
$75.00 / 1M tokens
Cached input, 0–272,001 context tokens
$1.00 / 1M tokens
Cached input, 272,001+ context tokens
$2.00 / 1M tokens
Cache write, 0–272,001 context tokens
$12.50 / 1M tokens
Cache write, 272,001+ context tokens
$25.00 / 1M tokens
o3openai/o3
Input$2.00 / 1M tokens
Output$8.00 / 1M tokens
More rates
Cached input
$0.50 / 1M tokens
Web search
$10.00 / request
o4-miniopenai/o4-mini
Input$1.10 / 1M tokens
Output$4.40 / 1M tokens
More rates
Cached input
$0.275 / 1M tokens
Web search
$10.00 / request
Thinking Machines
Inklingthinkingmachines/inkling
Input$1.00 / 1M tokens
Output$4.05 / 1M tokens
Varies by provider
More rates
Cached input
$0.17 / 1M tokens
Inkling Smallthinkingmachines/inkling-small
Input$0.50 / 1M tokens
Output$1.20 / 1M tokens
Varies by provider
More rates
Cached input
$0.10 / 1M tokens
xAI
Grok 4.1 Fast Non-Reasoningspacexai/grok-4.1-fast-non-reasoning
Input$0.20 / 1M tokens
Output$0.50 / 1M tokens
More rates
Cached input
$0.05 / 1M tokens
Grok 4.1 Fast Reasoningspacexai/grok-4.1-fast-reasoning
Input$0.20 / 1M tokens
Output$0.50 / 1M tokens
More rates
Cached input
$0.05 / 1M tokens
Grok 4.3spacexai/grok-4.3
Input$1.25 / 1M tokens
Output$2.50 / 1M tokens
Tiered
More rates
Cached input
$0.20 / 1M tokens
Web search
$5.00 / request
Input, 0–200,001 context tokens
$1.25 / 1M tokens
Input, 200,001+ context tokens
$2.50 / 1M tokens
Output, 0–200,001 context tokens
$2.50 / 1M tokens
Output, 200,001+ context tokens
$5.00 / 1M tokens
Cached input, 0–200,001 context tokens
$0.20 / 1M tokens
Cached input, 200,001+ context tokens
$0.40 / 1M tokens
Grok 4.5spacexai/grok-4.5
Input$2.00 / 1M tokens
Output$6.00 / 1M tokens
Tiered
More rates
Cached input
$0.30 / 1M tokens
Web search
$5.00 / request
Input, 0–200,001 context tokens
$2.00 / 1M tokens
Input, 200,001+ context tokens
$4.00 / 1M tokens
Output, 0–200,001 context tokens
$6.00 / 1M tokens
Output, 200,001+ context tokens
$12.00 / 1M tokens
Cached input, 0–200,001 context tokens
$0.30 / 1M tokens
Cached input, 200,001+ context tokens
$0.60 / 1M tokens
Grok 4.6spacexai/grok-4.6
Input$2.00 / 1M tokens
Output$6.00 / 1M tokens
Tiered
More rates
Cached input
$0.50 / 1M tokens
Web search
$5.00 / request
Input, 0–200,001 context tokens
$2.00 / 1M tokens
Input, 200,001+ context tokens
$4.00 / 1M tokens
Output, 0–200,001 context tokens
$6.00 / 1M tokens
Output, 200,001+ context tokens
$12.00 / 1M tokens
Cached input, 0–200,001 context tokens
$0.50 / 1M tokens
Cached input, 200,001+ context tokens
$1.00 / 1M tokens
Grok 4.20 Non-Reasoningspacexai/grok-4.20-non-reasoning
Input$1.25 / 1M tokens
Output$2.50 / 1M tokens
Tiered · Varies by provider
More rates
Cached input
$0.20 / 1M tokens
Web search
$5.00 / request
Input, 0–200,001 context tokens
$1.25 / 1M tokens
Input, 200,001+ context tokens
$2.50 / 1M tokens
Output, 0–200,001 context tokens
$2.50 / 1M tokens
Output, 200,001+ context tokens
$5.00 / 1M tokens
Cached input, 0–200,001 context tokens
$0.20 / 1M tokens
Cached input, 200,001+ context tokens
$0.40 / 1M tokens
Grok 4.20 Reasoningspacexai/grok-4.20-reasoning
Input$1.25 / 1M tokens
Output$2.50 / 1M tokens
Tiered · Varies by provider
More rates
Cached input
$0.20 / 1M tokens
Web search
$5.00 / request
Input, 0–200,001 context tokens
$1.25 / 1M tokens
Input, 200,001+ context tokens
$2.50 / 1M tokens
Output, 0–200,001 context tokens
$2.50 / 1M tokens
Output, 200,001+ context tokens
$5.00 / 1M tokens
Cached input, 0–200,001 context tokens
$0.20 / 1M tokens
Cached input, 200,001+ context tokens
$0.40 / 1M tokens
Z.ai
GLM 4.7 Flashzai/glm-4.7-flash
Input$0.07 / 1M tokens
Output$0.40 / 1M tokens
GLM 5zai/glm-5
Input$1.00 / 1M tokens
Output$3.20 / 1M tokens
GLM 5.1zai/glm-5.1
Input$1.40 / 1M tokens
Output$4.40 / 1M tokens
More rates
Cached input
$0.26 / 1M tokens
GLM 5.2zai/glm-5.2
Input$0.80 / 1M tokens
Output$2.55 / 1M tokens
Varies by provider
More rates
Cached input
$0.16 / 1M tokens
GLM 5.3zai/glm-5.3
Input$1.40 / 1M tokens
Output$4.40 / 1M tokens
Varies by provider
More rates
Cached input
$0.14 / 1M tokens
GLM 5.3 Flashzai/glm-5.3-flash
Input$0.15 / 1M tokens
Output$0.50 / 1M tokens
Varies by provider
More rates
Cached input
$0.03 / 1M tokens

Speech to text (STT)

Transcription models turn recorded speech into text. Choose a speech-to-text model in Settings → Speech & Voice. The Live audio icon identifies a streaming audio route in the catalog; each application's recording controls determine how it is used.

ModelCapabilitiesAccess cost USD
Fish Audio
Transcribe-1fish-audio/transcribe-1
Catalog rate$0.00 · listed as free
Google
Gemini 3.5 Transcribegoogle/gemini-3.5-transcribe
Input$2.00 / 1M tokens
Audio input$2.00 / 1M tokens
Output$12.00 / 1M tokens
Gemini 3.5 Transcribe Livegoogle/gemini-3.5-transcribe-live
Audio$0.009 / minute
OpenAI
GPT-4o mini Transcribeopenai/gpt-4o-mini-transcribe
Input$1.25 / 1M tokens
Audio input$1.25 / 1M tokens
Output$5.00 / 1M tokens
GPT-4o Transcribeopenai/gpt-4o-transcribe
Input$2.50 / 1M tokens
Audio input$2.50 / 1M tokens
Output$10.00 / 1M tokens
gpt-realtime-whisperopenai/gpt-realtime-whisper
Audio$0.01704 / minute
Whisperopenai/whisper-1
Audio$0.006 / minute
xAI
Grok STTDefaultspacexai/grok-stt
Audio$0.00168 / minute

Text to speech (TTS)

Speech models generate audio for read aloud. Choose the model and an available voice in Settings → Speech & Voice. Character-based rates apply to the input text.

ModelCapabilitiesAccess cost USD
Fish Audio
S1fish-audio/s1
Catalog rate$0.00 · listed as free
S2 Profish-audio/s2-pro
Catalog rate$0.00 · listed as free
S2.1 Profish-audio/s2.1-pro
Catalog rate$0.00 · listed as free
OpenAI
TTS-1Defaultopenai/tts-1
Text$15.00 / 1M characters
TTS-1 HDopenai/tts-1-hd
Text$30.00 / 1M characters
xAI
Grok TTSspacexai/grok-tts
Text$15.00 / 1M characters

Embedding models

Embedding models support retrieval across memory, conversation history, and indexed documents. Their rates apply to the text being embedded. Changing the embedding model can require existing content to be indexed again; see Memory.

ModelCapabilitiesAccess cost USD
Alibaba / Qwen
Qwen3 Embedding 4BDefaultalibaba/qwen3-embedding-4b
Input$0.02 / 1M tokens
Amazon
Titan Text Embeddings V2amazon/titan-embed-text-v2
Input$0.02 / 1M tokens
Google
Text Embedding 005google/text-embedding-005
Input$0.025 / 1M tokens
OpenAI
text-embedding-3-smallopenai/text-embedding-3-small
Input$0.02 / 1M tokens
Voyage AI
Voyage 3.5voyage/voyage-3.5
Input$0.06 / 1M tokens
Voyage 3.5 Litevoyage/voyage-3.5-lite
Input$0.02 / 1M tokens
Voyage 4voyage/voyage-4
Input$0.06 / 1M tokens
Voyage 4 Largevoyage/voyage-4-large
Input$0.12 / 1M tokens
Voyage 4 Litevoyage/voyage-4-lite
Input$0.02 / 1M tokens

Reranking models

Reranking models score retrieved results for relevance. These catalog entries are separate from the Chat model picker; inclusion in the catalog does not mean every retrieval uses a reranker.

ModelCapabilitiesAccess cost USD
Cohere
Cohere Rerank 3.5cohere/rerank-v3.5
Not published

Reasoning and privacy locks

The reasoning slider folds down to None for models without reasoning support. Privacy filters can require Zero Data Retention or no training; rows without a compatible route are locked. Premium rows can require Pro or Max. Fast is a separate route when the catalog publishes one and can have different metering. See Usage, Privacy, Voice, and Settings.

Documentation