MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office work, reaching fluency in generating and operating Word, Excel, and Powerpoint files, context switching between diverse software environments, and working across different agent and human teams. Scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp, M2.5 is also more token efficient than previous generations, having been trained to optimize its actions and output through planning.
| $0.27 | $0.95 | $0.03 | 0.76s | 35 tps | ||
10% off | $0.30$0.27 | $1.20$1.08 | $0.03$0.027 | 0.96s | 70 tps | |
| $0.295 | $1.20 | $0.06 | 1.44s | 79 tps | ||
| $0.30 | $1.20 | $0.06 | 0.78s | 25 tps | ||
| $0.30 | $1.20 | $0.06 | 0.31s | 163 tps | ||
| $0.30 | $1.20 | $0.03 | 1.16s | 11 tps | ||
| $0.30 | $1.20 | $0.03 | 1.02s | 8 tps | ||
| $0.60 | $2.40 | $0.06 | 1.30s | 48 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.