GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 LunaOpens in new tab, served with reasoning.mode set to pro for higher-quality responses on complex tasks.
Cost note: pro mode spends far more reasoning tokens per request, so a typical request costs several times more than the same request on GPT-5.6 Luna and takes much longer to complete. It is intended for hard, high-stakes problems where the extra accuracy justifies the cost. For everyday coding, agentic, and chat workloads, use GPT-5.6 LunaOpens in new tab instead.
Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-modeOpens in new tab
| $0.20 | $1.20 | $0.02 | 8.98s | 90 tps | ||
| $0.20 | $1.20 | $0.02 | 10.75s | 111 tps | ||
Not used in Standard routing:Why these endpoints are not used | ||||||
Flex | $0.10 | $0.60 | $0.01 | 58.07s | 78 tps | |
| $0.22 | $1.32 | $0.022 | 8.75s | 109 tps | ||
Fast | $0.40 | $2.40 | $0.04 | 5.41s | 185 tps | |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.