| 1 | GGPT-5.5 OpenAI's latest flagship model with 1050K context, leading AI reasoning and coding capabilities | OpenAI | 98.0 | 1050 | 5 | 30 | 91.50 | 91.40 | 1420 |
| 2 | CClaude Fable 5 Anthropic's most capable model, positioned above Opus tier. Public Mythos-class model with 1M context, scoring >10% higher than Claude Opus 4.8 on key benchmarks. Adaptive thinking only. | Anthropic | 97.0 | 1000 | 10 | 50 | — | — | — |
| 3 | GGemini 3.5 Pro Google DeepMind's flagship model with 1000K context window, comprehensive capabilities | Google DeepMind | 96.0 | 1000 | 1.50 | 9 | 91 | 89.50 | 1400 |
| 4 | GGPT-5.6 Sol OpenAI's flagship GPT-5.6 model. HealthBench Professional 60.5%, High on cybersecurity and biosecurity evals. New API Fast mode: 2.5x faster than Standard at 2x price, no change in intelligence (2026-07-30). | OpenAI | 96.0 | 1050 | 4 | 24 | 0 | — | 0 |
| 5 | GGrok 4.6 xAI's flagship model released Aug 12, 2026 (now operating as SpaceXAI). Scores 61 on the AA Intelligence Index, level with GPT-5.6 Sol. Strong in agentic work: GDPval-AA v2 Elo 1753, Terminal-Bench v2.1 88.4%, τ³-Banking 50.7%. Priced at $2/$6 per 1M tokens with 500K context; ~$0.84 per task, on par with Kimi K3. | xAI | 95.0 | 500 | 2 | 6 | 0 | — | 1753 |
| 6 | CClaude Opus 4.8 Anthropic's most advanced flagship model with industry-leading reasoning and safety features | Anthropic | 95.0 | 1000 | 5 | 25 | 0 | — | 0 |
| 7 | CClaude Opus 5 Anthropic's latest Opus flagship, within 0.5% of Fable 5 frontier intelligence at half the cost, $5/$25 per M tokens, #1 on AA Intelligence Leaderboard | Anthropic | 94.0 | 1000 | 5 | 25 | 89.20 | — | 1320 |
| 8 | KKimi K3 Moonshot AI's latest flagship model. 2.8T-parameter MoE (16/896 experts active), 1M context window, native vision. AA Intelligence Index #4, trailing only Fable 5 and GPT-5.6 Sol. Open weights releasing July 27. | Moonshot AI | 94.0 | 1049 | 3 | 15 | — | — | — |
| 9 | GGPT-5.6 Terra Balanced GPT-5.6 model for everyday work. API price cut 20% on 2026-07-30 ($2.00/M input, $12.00/M output, cached input $0.20/M). | OpenAI | 93.5 | 1050 | 2 | 12 | 0 | — | 0 |
| 10 | GGemini 3.7 Flash Google's Flash-tier workhorse model released Aug 13, 2026 for coding and agents, intro-priced at half of 3.6 Flash's original rate ($0.75/$3.75 per 1M tokens). FrontierCode 1.1: 43.6%, DeepSWE v1.1: 65.3%. | Google DeepMind | 93.0 | 1049 | 0.38 | 1.88 | 0 | — | 0 |
| 11 | CClaude Opus 4.7 Anthropic's premium reasoning model, 1000K context, for complex tasks and enterprise use | Anthropic | 93.0 | 1000 | 5 | 25 | 91 | 92.50 | 1400 |
| 12 | MMuse Spark 1.2 Meta Muse Spark 1.2 multimodal model with 1M context, successor to Muse Spark 1.1 | Meta | 92.0 | 1024 | 1.25 | 4.25 | 0 | — | 0 |
| 13 | CClaude Opus 5 Fast Fast-mode variant of Claude Opus 5 with identical capabilities, ~2.5x faster output at 2x base pricing ($10/$50 per M tokens), 1M context, ideal for latency-sensitive long-horizon agentic work | Anthropic | 92.0 | 1000 | 10 | 50 | 89.20 | — | 1320 |
| 14 | GGemini 3.6 Flash Google's latest Gemini Flash series flagship for coding, knowledge work, and multimodal tasks. 17% less output tokens, DeepSWE 49%, OSWorld-Verified 83.0%. API $1.50/$7.50. | Google DeepMind | 92.0 | 1049 | 1.50 | 7.50 | 0 | — | 0 |
| 15 | MMuse Spark 1.1 Meta's flagship multimodal reasoning model with 1M-token context, built for agentic tasks, coding, and computer use | Meta | 92.0 | 1024 | 1.25 | 4.25 | 0 | — | 0 |
| 16 | CClaude Sonnet 5 Anthropic Claude mid-range model, balancing performance and cost | Anthropic | 92.0 | 1000 | 2 | 10 | — | — | 1312 |
| 17 | QQwen 3.8 2.4T A95B Qwen 3.8 flagship MoE model with 2.4T total params and 95B active, 1M context window | Alibaba (Qwen) | 91.0 | 1049 | 2 | 6 | 0 | — | 0 |
| 18 | GGLM-5.2 Zhipu AI's latest flagship model with strong bilingual understanding and reasoning performance | Z.ai (Zhipu AI) | 91.0 | 1000 | 1.40 | 4.40 | — | — | — |
| 19 | SSakana Fugu Ultra Fugu Ultra is Sakana AI's flagship multi-agent orchestration model. Rather than a single monolithic model, it dynamically orchestrates a pool of expert models to tackle complex multi-step tasks. Benchmarks competitive with Fable 5 and Mythos Preview. 1M context window, text+image input. | Sakana AI | 91.0 | 1000 | 5 | 30 | — | — | — |
| 20 | GGPT-5.6 Luna Fastest, most affordable GPT-5.6 model. API price cut 80% on 2026-07-30 ($0.20/M input, $1.20/M output), near-frontier performance at roughly 6 cents on the dollar per task. Supports tools and multi-step workflows. | OpenAI | 91.0 | 1050 | 0.20 | 1.20 | 0 | — | 0 |
| 21 | GGLM-5.3 Zhipu AI's post-training RL model with open-source SOTA coding and emergent cyber capabilities | Zhipu AI (Z.ai) | 90.0 | 1000 | 1.40 | 4.40 | 0 | — | 1250 |
| 22 | QQwen3.8 Max Alibaba's newest flagship in the Qwen family: 2.4T-parameter MoE (95B active), focused on coding and agentic cowork long-horizon tasks. Open weights Qwen3.8-2.4T-A95B (BF16/FP8) released on Hugging Face Aug 13, 2026 — the largest open-weight model ever; the open version is text-only with mandatory thinking and 262K native context. PaperBench 93.0, the highest published score. API $2/$6 per M tokens. | Alibaba (Qwen) | 90.0 | 1000 | 2 | 6 | 0 | — | 0 |
| 23 | CClaude Opus 4.6 Anthropic's previous-generation flagship model with strong reasoning and long-context capabilities | Anthropic | 90.0 | 1000 | 5 | 25 | 0 | — | 0 |
| 24 | GGPT-5.4 Pro OpenAI's professional-grade model offering advanced reasoning for enterprise workloads | OpenAI | 90.0 | 1050 | 30 | 180 | — | — | — |
| 25 | GGemini 3.5 Flash Google lightweight flagship model with built-in Computer Use, function calling, Search/Maps Grounding — ideal for agent scenarios | Google DeepMind | 90.0 | 1049 | 1.50 | 9 | 92.30 | 86.80 | 1370 |
| 26 | GGPT-5.5 Instant GPT-5.5 low-latency version, fast responses ideal for chat scenarios | OpenAI | 88.0 | 922 | 0.75 | 3 | 89.50 | 88.20 | 1350 |
| 27 | DDeepSeek V4 Flash DeepSeek-V4-Flash-0731 (2026-07-31): 284B MoE (13B active), 1M context, 384K max output. Major agent capability upgrade - Terminal Bench 2.1 82.7, Toolathlon 70.3, Cybergym 76.7, DeepSWE 54.4, far exceeding V4-Pro-Preview. ARC Prize verified (2026-08-07): 89.0% on ARC-AGI-1, 61.4% on ARC-AGI-2 at max effort ($0.02/$0.04 per task). Native Responses API with Codex adaptation. $0.14/M input, $0.28/M output. | DeepSeek | 88.0 | 1311 | 0.22 | 0.66 | 0 | — | 0 |
| 28 | VVibeThinker-3B 3B dense reasoning model. AIME26: 94.3, based on Qwen2.5. Uses Spectrum-to-Signal post-training. No tool calling support, focused on math and code reasoning. | Weibo AI | 88.0 | 32 | — | — | — | — | — |
| 29 | DDeepSeek V4 Pro DeepSeek's flagship reasoning model. Official release deepseek-v4-pro-0813 shipped Aug 13, 2026, replacing V4-Pro-Preview. 1M context, up to 384K output, $0.435/$0.87 per 1M tokens. Big agent gains: Terminal Bench 2.1 87.9, DeepSWE 62.7, Cybergym 83.3 (preview: 72.1/12.8/52.7). Natively supports the Responses API (Codex-compatible). | DeepSeek | 87.0 | 1049 | 0.66 | 1.98 | — | — | — |
| 30 | DDeepSeek V4 Flash Vision DeepSeek V4 Flash Vision experimental, multimodal input support, extremely low pricing | DeepSeek | 86.0 | 1049 | 0.22 | 0.66 | 0 | — | 0 |
| 31 | KKimi K2.7 Code Moonshot AI's specialized coding model based on Kimi K2.6 architecture | Moonshot AI | 85.0 | 256 | 0.74 | 3.50 | — | — | — |
| 32 | QQwen3.7 Max Alibaba Qwen's most powerful model, 1000K context, MoE architecture | Alibaba (Qwen) | 85.0 | 1000 | 1.25 | 3.75 | 87 | 87 | 1300 |
| 33 | GGemini 3.1 Pro Google DeepMind previous-gen flagship model with 1049K context | Google DeepMind | 85.0 | 1049 | 2 | 12 | 87.50 | 85 | 1300 |
| 34 | GGrok 4.5 xAI's latest coding and agentic model based on 1.5T V9 foundation, 500K context, $2/$6/M tokens, coding performance competitive with GPT-5.5 | xAI | 83.0 | 500 | 2 | 6 | 0 | — | 0 |
| 35 | QQwen 3.8 27B Qwen 3.8 27B open-weight model with 1M context, strong reasoning capabilities | Alibaba (Qwen) | 83.0 | 1000 | 0.40 | 3 | 0 | — | 0 |
| 36 | SSeed 2.1 Turbo ByteDance Seed 2.1 Turbo model, 262K context, fast speed and low pricing | ByteDance | 82.0 | 262 | 0.50 | 2.50 | 0 | — | 0 |
| 37 | GGrok 4.20 Multi-Agent xAI multi-agent reasoning model built on Grok 4.20, supports multi-agent collaborative orchestration with 2M context, ideal for complex task decomposition and parallel execution | xAI | 82.0 | 2000 | 1.25 | 2.50 | 86 | 85 | 1275 |
| 38 | CCursor Composer 2.5 Cursor AI IDE built-in Agent coding mode, 256K context | Cursor | 82.0 | 256 | 0 | 0 | 85 | 86 | 1260 |
| 39 | KKimi K2.6 Moonshot AI's latest Kimi model, 262K context, strong Agent capabilities | Moonshot AI | 82.0 | 262 | 0.68 | 3.42 | 85.50 | 84.50 | 1280 |
| 40 | GGPT-5.4 OpenAI GPT-5.4 flagship model with 1050K context | OpenAI | 82.0 | 1050 | 2.50 | 15 | 88.20 | 87.50 | 1320 |
| 41 | GGrok 4.20 xAI's latest model featuring multi-agent coordination and enhanced reasoning | xAI | 81.0 | 2000 | 1.25 | 2.50 | — | — | — |
| 42 | CClaude Sonnet 4.6 Anthropic Claude Sonnet series, mid-to-high-end reasoning model | Anthropic | 80.0 | 1000 | 3 | 15 | 86.50 | 88 | 1280 |
| 43 | NNemotron 3.5 Lightning NVIDIA Nemotron 3.5 Lightning efficient inference model, 262K context, ultra-low latency | NVIDIA | 80.0 | 262 | 0.08 | 0.20 | 0 | — | 0 |
| 44 | MMiniMax M3 MiniMax's latest flagship model with competitive performance across reasoning tasks | MiniMax | 80.0 | 1000 | 0.30 | 1.20 | — | — | — |
| 45 | WWindsurf SWE-1.6 Windsurf full-stack AI coding assistant, 200K context | Windsurf (Codeium) | 80.0 | 200 | 0 | 0 | 0 | 0 | 0 |
| 46 | GGrok 4.3 xAI Grok latest flagship with 1000K context | xAI | 80.0 | 1000 | 1.25 | 2.50 | 86 | 85 | 1270 |
| 47 | MMiMo-V2.5 Pro Xiaomi MiMo-V2.5 Pro on-device LLM | Xiaomi | 78.0 | 1000 | 0.44 | 0.88 | 85 | 84 | 1260 |
| 48 | QQwen3.8-27B Qwen3.8 series open model (released Aug 14, 2026, Apache 2.0): 27B dense multimodal model with native image/video understanding, 262K native context (extendable to 1M via YaRN), Gated DeltaNet hybrid architecture, thinking mode on by default and toggleable per request. SWE-bench Pro 61.7, QwenSWEBench 79.0, DeepSWE 1.1 42.2, LiveCodeBench v6 90.3, OSWorld 84.3 - beats Claude Opus 4.6 Max on multiple coding and agent benchmarks. Runs locally on consumer hardware (~48 tps q4km on a 4090). | Alibaba (Qwen) | 78.0 | 1000 | 0.45 | 3.20 | 0 | — | 0 |
| 49 | LLongCat-2.0 1.6 trillion parameter MoE model from Meituan (LongCat), ~48B activated per token, trained on domestic AI ASIC superpods, 1M context window, MIT license | Meituan | 78.0 | 1049 | 0 | 0 | — | — | — |
| 50 | MMuse Glimmer Meta Superintelligence Labs' open 30B on-device agentic model (Apache 2.0), distilled from Muse Spark, quantized to under 20GB for single-GPU local runs, with multimodal input and tool calling. | Meta (Superintelligence Labs) | 78.0 | 128 | 0 | 0 | 0 | — | 0 |
| 51 | QQwen3.6 Plus Alibaba Qwen3.6 Plus mid-to-high-end model, 1000K context | Alibaba (Qwen) | 76.0 | 1000 | 0.33 | 1.95 | 84 | 84 | 1250 |
| 52 | GGPT-4o OpenAI GPT-4o multimodal model with 128K context | OpenAI | 75.0 | 128 | 2.50 | 10 | 88.70 | 90.20 | 1287 |
| 53 | GGemini 3.5 Flash-Lite Google's fastest 3.5-series model at 350 tokens/s, designed for low-latency, high-throughput agentic search and document processing. Priced $0.30/$2.50. SWE-Bench Pro 54.2%, OSWorld 74.0%. 1M context. | Google DeepMind | 75.0 | 1000 | 0.30 | 2.50 | 0 | — | 0 |
| 54 | QQwen3.7 Plus Alibaba's flagship Qwen model with improved instruction following and tool use | Alibaba (Qwen) | 75.0 | 1000 | 0.32 | 1.28 | — | — | — |
| 55 | XXiaomi-Robotics-1 Xiaomi's embodied foundation model, 100K+ hours real-world pretraining, natural language instructions, 1600+ scenario generalization | Xiaomi | 75.0 | 128 | 0 | 0 | 0 | — | 0 |
| 56 | GGLM-5.1 Zhipu AI GLM series previous-gen flagship, 200K context | 智谱AI (Zhipu) | 75.0 | 200 | 0.40 | 1.20 | 83 | 82 | 1240 |
| 57 | GGemini Robotics 2 Google DeepMind's next-gen embodied intelligence family: Gemini Robotics 2 (VLA whole-body control), Gemini Robotics ER 2 (embodied reasoning, on AI Studio) and On-Device 2 (on-device). Fine dexterity, multi-robot collaboration, adapts to new robot bodies in hours. | Google DeepMind | 74.0 | 128 | 0 | 0 | 0 | — | 0 |
| 58 | CCursor Composer 2 Cursor Composer 2 AI coding assistant | Cursor | 72.0 | 256 | 0 | 0 | 82 | 82 | 1220 |
| 59 | MMiMo-V2.5 Xiaomi MiMo-V2.5 on-device LLM | Xiaomi | 72.0 | 1049 | 0.15 | 0.29 | — | — | — |
| 60 | MMiniMax-M2.7 MiniMax M2.7 chat model | MiniMax | 72.0 | 205 | 0.28 | 1.20 | 82 | 81 | 1220 |
| 61 | KKimi K2.5 Moonshot AI Kimi K2.5 chat model, 262K context | Moonshot AI | 72.0 | 262 | 0.40 | 1.90 | 82 | 82 | 1220 |
| 62 | GGemini 3 Flash Google DeepMind lightweight Gemini model with 1000K context | Google DeepMind | 70.0 | 1000 | 0.15 | 0.60 | 82 | 80.50 | 1220 |
| 63 | GGLM-5 Zhipu AI GLM-5 model, previous-gen flagship | 智谱AI (Zhipu) | 70.0 | 200 | 0.30 | 0.90 | 81 | 79 | 1210 |
| 64 | QQwen3.5 397B Alibaba Qwen3.5 397B parameter large model | Alibaba (Qwen) | 68.0 | 262 | 0.45 | 1.35 | 80.50 | 80.50 | 1200 |
| 65 | IInkling Open-weight MoE model by Thinking Machines Lab, 975B total / 41B active params. Multimodal (text, image, audio), 1M context. AA Intelligence Index 41 - leading US open weights model. Strong agent performance: Elo 1238 on GDPval-AA v2, beating Kimi K2.6 and DeepSeek V4 Flash. | Thinking Machines Lab | 68.0 | 1049 | 1 | 4.05 | 0 | — | 1238 |
| 66 | GGPT-5.4 Mini OpenAI's compact model in the GPT-5.4 series, optimized for efficiency | OpenAI | 67.0 | 400 | 0.75 | 4.50 | — | — | — |
| 67 | QQwen3 Coder 480B A35B Qwen's most powerful open-source coding model. 480B MoE with 35B active params, native 256K context (YaRN scalable to 1M). Strong SWE-Bench performance. Apache 2.0 licensed. Ships with Qwen Code CLI. | Alibaba (Qwen) | 66.0 | 256 | 0.22 | 1.80 | — | — | — |
| 68 | IInkling Small Open-weights MoE from Thinking Machines Lab, 276B total/12B active, comparable performance to Inkling at a quarter of the size. Native audio+image reasoning, variable thinking effort, up to 1M context. Highest-scoring open-weight model on ARC Prize. | Thinking Machines Lab | 66.0 | 524 | 0 | 0 | 0 | — | 0 |
| 69 | GGemini 2.5 Pro Google DeepMind previous-gen Gemini high-end model | Google DeepMind | 65.0 | 1000 | 0.35 | 1.40 | 80.50 | 78 | 1180 |
| 70 | GGrok 3 xAI previous-gen Grok model with 1000K context | xAI | 65.0 | 1000 | 0.15 | 0.60 | 80 | 80 | 1180 |
| 71 | HHunyuan Hy3 Preview Tencent Hunyuan Hy3 Preview model, 256K context | Tencent Hunyuan | 65.0 | 256 | 0.06 | 0.18 | 79 | 78 | 1180 |
| 72 | SSeed 2.0 Code ByteDance Seed 2.0 Code model optimized for frontend development, multilingual coding, and agentic coding | ByteDance Seed | 65.0 | 262 | 0.50 | 3 | 0 | — | 0 |
| 73 | PPoolside Laguna S 2.1 Poolside coding-specialized agent model, 118B-A8B MoE, 1M context, Terminal-Bench 70.2%, DeepSWE 40.4%, focused on autonomous long-horizon engineering work | Poolside | 65.0 | 1000 | 0.10 | 0.20 | — | — | — |
| 74 | DDeepSeek V3.2 DeepSeek's capable mid-range model offering strong performance at lower cost | DeepSeek | 63.0 | 164 | 0.23 | 0.34 | 0 | — | 0 |
| 75 | NNemotron 3 Ultra NVIDIA's most capable open-weight model with strong benchmark performance | NVIDIA | 62.0 | 1000 | 0.50 | 2.20 | — | — | — |
| 76 | GGPT-Live-1 OpenAI real-time voice conversation model, supports interruption and continuation, simultaneous voice chat and reasoning | OpenAI | 62.0 | 32 | 0 | 0 | 0 | — | 0 |
| 77 | CClaude 4.5 Haiku Anthropic lightweight fast model with 200K context | Anthropic | 60.0 | 200 | 0.80 | 4 | 78 | 75 | 1150 |
| 78 | GGemini 3.1 Flash Lite Google's lightweight Gemini model optimized for speed and efficiency on edge devices | Google DeepMind | 60.0 | 1049 | 0.25 | 1.50 | 0 | — | 0 |
| 79 | CCodestral 2508 Mistral AI's specialized coding model, optimized for code generation, completion, and refactoring across multiple programming languages | Mistral AI | 60.0 | 256 | 0.30 | 0.90 | 0 | — | 0 |
| 80 | SStep 3.7 Flash StepFun's latest multimodal MoE model with 196B parameter language backbone and vision encoder for native image/video understanding | StepFun | 60.0 | 256 | 0.20 | 1.15 | 78 | 75 | — |
| 81 | GGPT-5.4 Nano OpenAI's compact model in the GPT-5.4 lineup, designed for rapid inference and cost-efficient deployment | OpenAI | 58.0 | 400 | 0.20 | 1.25 | — | — | — |
| 82 | LLaguna XS 2.1 Poolside Laguna XS 2.1 is a 33B-A3B MoE coding agent model optimized for local deployment. Builds on XS.2 with improved SWE-bench Multilingual (63.1%) and stronger terminal-style task performance. Supports 256K context, runs locally in vLLM/SGLang. | Poolside | 58.0 | 256 | 0.06 | 0.12 | — | — | — |
| 83 | NNex AGI Nex-N2-Pro Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active / 397B total parameters. Built on Qwen3.5 architecture, supports text and image input, 262K context window. Extremely cost-effective at $0.25/M input tokens. | Nex AGI | 58.0 | 262 | 0.25 | 1 | — | — | — |
| 84 | CCursor Composer 1.5 Cursor Composer 1.5 earlier version | Cursor | 58.0 | 200 | 0 | 0 | 76 | 74 | 1150 |
| 85 | QQwen3.7 Flash Alibaba's lightweight vision-language reasoning model in the Qwen3.7 family, 1M context, accepts image/video input, suited for multimodal agents, visual coding, search and computer interaction, priced at $0.03/$0.13 per M tokens | Alibaba (Qwen) | 58.0 | 1000 | 0.03 | 0.13 | 0 | — | 0 |
| 86 | DDeepSeek R1 DeepSeek R1 reasoning-enhanced model, 128K context, MIT open-source | DeepSeek | 55.0 | 128 | 0.55 | 2.19 | 78.50 | 78.50 | 1100 |
| 87 | NNova 2.0 Pro Amazon Nova 2.0 Pro enterprise-grade model | Amazon | 55.0 | 256 | 0.80 | 3.20 | 76 | 72 | 1120 |
| 88 | NNemotron 3 Super NVIDIA Nemotron 3 Super open-source model | NVIDIA | 55.0 | 1000 | 0.14 | 0.42 | 76 | 74 | 1120 |
| 89 | SSolar Pro 4 Upstage Solar Pro4 cost-efficient LLM with 524K context, ultra-low pricing | Upstage | 55.0 | 524 | 0.03 | 0.12 | 0 | — | 0 |
| 90 | MMistral Large 2512 Mistral AI's flagship large language model with strong reasoning and multilingual capabilities | Mistral AI | 55.0 | 262 | 0.50 | 1.50 | 0 | — | 0 |
| 91 | MMistral Medium 3.5 Mistral AI's mid-range model balancing performance and efficiency for production deployments | Mistral AI | 55.0 | 262 | 1.50 | 7.50 | 0 | — | 0 |
| 92 | CCohere North Mini Code Cohere North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse MoE with 30B total / 3B active parameters, optimized for end-to-end coding tasks. 256K context window, text input only. | Cohere | 55.0 | 256 | 0 | 0 | — | — | — |
| 93 | SStep 3.5 Flash StepFun Step 3.5 Flash fast model | StepFun (阶跃星辰) | 55.0 | 256 | 0.03 | 0.09 | 75 | 72 | 1100 |
| 94 | DDoubao Seed Code ByteDance Doubao Seed Code coding model | 字节跳动 (ByteDance) | 55.0 | 256 | 0.10 | 0.30 | 76 | 74 | 1120 |
| 95 | MMistral Large 3 Mistral AI previous-gen flagship with 256K context | Mistral AI | 50.0 | 256 | 0.30 | 0.90 | 75 | 70 | 1100 |
| 96 | GGrok Build 0.1 xAI's coding-focused model trained for agentic software engineering workflows, supports text+image input | xAI | 50.0 | 256 | 1 | 2 | — | — | — |
| 97 | CCommand A+ Cohere Command A+ enhanced enterprise-grade generation model, 128K context | Cohere | 48.0 | 128 | 0 | 0 | 0 | — | 0 |
| 98 | LLlama 4 Maverick Meta Llama 4 Maverick open-source model with 1000K context | Meta | 45.0 | 1000 | 0.17 | 0.50 | 72 | 72 | 1080 |
| 99 | EERNIE 5.0 Thinking Baidu ERNIE 5.0 Thinking reasoning-enhanced | 百度 (Baidu) | 45.0 | 128 | 0.25 | 0.75 | 70 | 68 | 1050 |
| 100 | GGranite 4.1 8B IBM's open-source 8B dense enterprise model, matching 32B MoE performance, 131K context, designed for enterprise tasks | IBM | 40.0 | 131 | 0.05 | 0.10 | — | — | — |