Catalog
Every model. One rate card.
glm-5.3-prime
Zhipu AI · 1M context
qwen3.8-max-prime
Qwen · 1M context
gpt-6-luna-pro
OpenAI · 1.1M context
gpt-6-luna
OpenAI · 1.1M context
gpt-6-sol-pro
OpenAI · 1.1M context
grok-4.7
xAI · 500K context
qwen3.8-omni-flash
Qwen · 1M context
glm-5.3-flashx
Zhipu AI · 1M context
deepseek-v4.1-flash
DeepSeek · 1M context
qwen3.8-max-0902
Qwen · 1M context
gemini-3.8-flash
Google · 1M context
glm-5.3-flash
Zhipu AI · 1.3M context
deepseek-v4-flash-vision-exp
DeepSeek · 1M context
gemini-3.7-flash
Google · 1M context
deepseek-v4-pro-0813
DeepSeek · 1M context
grok-4.6
xAI · 500K context
gemini-3.6-flash
Google · 1M context
kimi-k3
Moonshot AI · 1M context
grok-4.5
xAI · 500K context
kimi-k2.7-code
Moonshot AI · 262K context
minimax-m3
MiniMax · 1M context
mistral-medium-3-5
Mistral · 262K context
kimi-k2.6
Moonshot AI · 262K context
minimax-m2.7
MiniMax · 205K context
mistral-small-2603
Mistral · 262K context
minimax-m2.5
MiniMax · 205K context
devstral-2512
Mistral · 262K context
llama-4-maverick
Meta · 1M context
llama-4-scout
Meta · 1.3M context
llama-3.3-70b-instruct
Meta · 131K context
You pay for tokens. That is the whole bill.
With a prepaid key, each call is charged its exact token count at the posted rate. With pay-per-call you sign a ceiling (the full input plus max_tokens of output) and the unused part returns to your wallet. Repeated deterministic requests are answered from cache at half the charge.
Rates shown are indicative public list prices while the lane is in preview. The OpenAI-shaped listing will be served at GET /v1/models; see the models reference.