LinkWord
Home
Directory
Articles
AI models
Tools
Pixel Plaza
Settings
ContactRSSFriend linksSubmit site
Privacy Policy·Disclaimer
陕ICP备2025083618号-2

Hot channels

AI ToolsDeveloper ToolsProductivity ToolsEntertainment & MediaJobs & Careers
DirectoryArticlesTools
Back to home

AI model leaderboard · LinkWord

Compare chat, image, and video models by composite score and category-specific metrics. Rankings are editorially maintained for reference.

RankModelProvider
ScoreEditorial composite score; higher ranks higher
Context (K)Context window size (thousand tokens)
Input $Input token price (USD per 1M tokens)
Output $Output token price (USD per 1M tokens)
MMLUMassive Multitask Language Understanding accuracy (%)
HumanEvalCode generation benchmark pass rate (%)
EloArena Elo from human preference battles; higher is stronger
1
G
GPT-5.5
OpenAI's latest flagship model with 1050K context, leading AI reasoning and coding capabilities
OpenAI98.0105053091.5091.401420
2
C
Claude Fable 5
Anthropic's most capable model, positioned above Opus tier. Public Mythos-class model with 1M context, scoring >10% higher than Claude Opus 4.8 on key benchmarks. Adaptive thinking only.
Anthropic97.010001050———
3
G
Gemini 3.5 Pro
Google DeepMind's flagship model with 1000K context window, comprehensive capabilities
Google DeepMind96.010001.5099189.501400
4
G
GPT-5.6 Sol
OpenAI's flagship GPT-5.6 model. HealthBench Professional 60.5%, High on cybersecurity and biosecurity evals. New API Fast mode: 2.5x faster than Standard at 2x price, no change in intelligence (2026-07-30).
OpenAI96.010504240—0
5
G
Grok 4.6
xAI's flagship model released Aug 12, 2026 (now operating as SpaceXAI). Scores 61 on the AA Intelligence Index, level with GPT-5.6 Sol. Strong in agentic work: GDPval-AA v2 Elo 1753, Terminal-Bench v2.1 88.4%, τ³-Banking 50.7%. Priced at $2/$6 per 1M tokens with 500K context; ~$0.84 per task, on par with Kimi K3.
xAI95.0500260—1753
6
C
Claude Opus 4.8
Anthropic's most advanced flagship model with industry-leading reasoning and safety features
Anthropic95.010005250—0
7
C
Claude Opus 5
Anthropic's latest Opus flagship, within 0.5% of Fable 5 frontier intelligence at half the cost, $5/$25 per M tokens, #1 on AA Intelligence Leaderboard
Anthropic94.0100052589.20—1320
8
K
Kimi K3
Moonshot AI's latest flagship model. 2.8T-parameter MoE (16/896 experts active), 1M context window, native vision. AA Intelligence Index #4, trailing only Fable 5 and GPT-5.6 Sol. Open weights releasing July 27.
Moonshot AI94.01049315———
9
G
GPT-5.6 Terra
Balanced GPT-5.6 model for everyday work. API price cut 20% on 2026-07-30 ($2.00/M input, $12.00/M output, cached input $0.20/M).
OpenAI93.510502120—0
10
G
Gemini 3.7 Flash
Google's Flash-tier workhorse model released Aug 13, 2026 for coding and agents, intro-priced at half of 3.6 Flash's original rate ($0.75/$3.75 per 1M tokens). FrontierCode 1.1: 43.6%, DeepSWE v1.1: 65.3%.
Google DeepMind93.010490.381.880—0
11
C
Claude Opus 4.7
Anthropic's premium reasoning model, 1000K context, for complex tasks and enterprise use
Anthropic93.010005259192.501400
12
M
Muse Spark 1.2
Meta Muse Spark 1.2 multimodal model with 1M context, successor to Muse Spark 1.1
Meta92.010241.254.250—0
13
C
Claude Opus 5 Fast
Fast-mode variant of Claude Opus 5 with identical capabilities, ~2.5x faster output at 2x base pricing ($10/$50 per M tokens), 1M context, ideal for latency-sensitive long-horizon agentic work
Anthropic92.01000105089.20—1320
14
G
Gemini 3.6 Flash
Google's latest Gemini Flash series flagship for coding, knowledge work, and multimodal tasks. 17% less output tokens, DeepSWE 49%, OSWorld-Verified 83.0%. API $1.50/$7.50.
Google DeepMind92.010491.507.500—0
15
M
Muse Spark 1.1
Meta's flagship multimodal reasoning model with 1M-token context, built for agentic tasks, coding, and computer use
Meta92.010241.254.250—0
16
C
Claude Sonnet 5
Anthropic Claude mid-range model, balancing performance and cost
Anthropic92.01000210——1312
17
Q
Qwen 3.8 2.4T A95B
Qwen 3.8 flagship MoE model with 2.4T total params and 95B active, 1M context window
Alibaba (Qwen)91.01049260—0
18
G
GLM-5.2
Zhipu AI's latest flagship model with strong bilingual understanding and reasoning performance
Z.ai (Zhipu AI)91.010001.404.40———
19
S
Sakana Fugu Ultra
Fugu Ultra is Sakana AI's flagship multi-agent orchestration model. Rather than a single monolithic model, it dynamically orchestrates a pool of expert models to tackle complex multi-step tasks. Benchmarks competitive with Fable 5 and Mythos Preview. 1M context window, text+image input.
Sakana AI91.01000530———
20
G
GPT-5.6 Luna
Fastest, most affordable GPT-5.6 model. API price cut 80% on 2026-07-30 ($0.20/M input, $1.20/M output), near-frontier performance at roughly 6 cents on the dollar per task. Supports tools and multi-step workflows.
OpenAI91.010500.201.200—0
21
G
GLM-5.3
Zhipu AI's post-training RL model with open-source SOTA coding and emergent cyber capabilities
Zhipu AI (Z.ai)90.010001.404.400—1250
22
Q
Qwen3.8 Max
Alibaba's newest flagship in the Qwen family: 2.4T-parameter MoE (95B active), focused on coding and agentic cowork long-horizon tasks. Open weights Qwen3.8-2.4T-A95B (BF16/FP8) released on Hugging Face Aug 13, 2026 — the largest open-weight model ever; the open version is text-only with mandatory thinking and 262K native context. PaperBench 93.0, the highest published score. API $2/$6 per M tokens.
Alibaba (Qwen)90.01000260—0
23
C
Claude Opus 4.6
Anthropic's previous-generation flagship model with strong reasoning and long-context capabilities
Anthropic90.010005250—0
24
G
GPT-5.4 Pro
OpenAI's professional-grade model offering advanced reasoning for enterprise workloads
OpenAI90.0105030180———
25
G
Gemini 3.5 Flash
Google lightweight flagship model with built-in Computer Use, function calling, Search/Maps Grounding — ideal for agent scenarios
Google DeepMind90.010491.50992.3086.801370
26
G
GPT-5.5 Instant
GPT-5.5 low-latency version, fast responses ideal for chat scenarios
OpenAI88.09220.75389.5088.201350
27
D
DeepSeek V4 Flash
DeepSeek-V4-Flash-0731 (2026-07-31): 284B MoE (13B active), 1M context, 384K max output. Major agent capability upgrade - Terminal Bench 2.1 82.7, Toolathlon 70.3, Cybergym 76.7, DeepSWE 54.4, far exceeding V4-Pro-Preview. ARC Prize verified (2026-08-07): 89.0% on ARC-AGI-1, 61.4% on ARC-AGI-2 at max effort ($0.02/$0.04 per task). Native Responses API with Codex adaptation. $0.14/M input, $0.28/M output.
DeepSeek88.013110.220.660—0
28
V
VibeThinker-3B
3B dense reasoning model. AIME26: 94.3, based on Qwen2.5. Uses Spectrum-to-Signal post-training. No tool calling support, focused on math and code reasoning.
Weibo AI88.032—————
29
D
DeepSeek V4 Pro
DeepSeek's flagship reasoning model. Official release deepseek-v4-pro-0813 shipped Aug 13, 2026, replacing V4-Pro-Preview. 1M context, up to 384K output, $0.435/$0.87 per 1M tokens. Big agent gains: Terminal Bench 2.1 87.9, DeepSWE 62.7, Cybergym 83.3 (preview: 72.1/12.8/52.7). Natively supports the Responses API (Codex-compatible).
DeepSeek87.010490.661.98———
30
D
DeepSeek V4 Flash Vision
DeepSeek V4 Flash Vision experimental, multimodal input support, extremely low pricing
DeepSeek86.010490.220.660—0
31
K
Kimi K2.7 Code
Moonshot AI's specialized coding model based on Kimi K2.6 architecture
Moonshot AI85.02560.743.50———
32
Q
Qwen3.7 Max
Alibaba Qwen's most powerful model, 1000K context, MoE architecture
Alibaba (Qwen)85.010001.253.7587871300
33
G
Gemini 3.1 Pro
Google DeepMind previous-gen flagship model with 1049K context
Google DeepMind85.0104921287.50851300
34
G
Grok 4.5
xAI's latest coding and agentic model based on 1.5T V9 foundation, 500K context, $2/$6/M tokens, coding performance competitive with GPT-5.5
xAI83.0500260—0
35
Q
Qwen 3.8 27B
Qwen 3.8 27B open-weight model with 1M context, strong reasoning capabilities
Alibaba (Qwen)83.010000.4030—0
36
S
Seed 2.1 Turbo
ByteDance Seed 2.1 Turbo model, 262K context, fast speed and low pricing
ByteDance82.02620.502.500—0
37
G
Grok 4.20 Multi-Agent
xAI multi-agent reasoning model built on Grok 4.20, supports multi-agent collaborative orchestration with 2M context, ideal for complex task decomposition and parallel execution
xAI82.020001.252.5086851275
38
C
Cursor Composer 2.5
Cursor AI IDE built-in Agent coding mode, 256K context
Cursor82.02560085861260
39
K
Kimi K2.6
Moonshot AI's latest Kimi model, 262K context, strong Agent capabilities
Moonshot AI82.02620.683.4285.5084.501280
40
G
GPT-5.4
OpenAI GPT-5.4 flagship model with 1050K context
OpenAI82.010502.501588.2087.501320
41
G
Grok 4.20
xAI's latest model featuring multi-agent coordination and enhanced reasoning
xAI81.020001.252.50———
42
C
Claude Sonnet 4.6
Anthropic Claude Sonnet series, mid-to-high-end reasoning model
Anthropic80.0100031586.50881280
43
N
Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning efficient inference model, 262K context, ultra-low latency
NVIDIA80.02620.080.200—0
44
M
MiniMax M3
MiniMax's latest flagship model with competitive performance across reasoning tasks
MiniMax80.010000.301.20———
45
W
Windsurf SWE-1.6
Windsurf full-stack AI coding assistant, 200K context
Windsurf (Codeium)80.020000000
46
G
Grok 4.3
xAI Grok latest flagship with 1000K context
xAI80.010001.252.5086851270
47
M
MiMo-V2.5 Pro
Xiaomi MiMo-V2.5 Pro on-device LLM
Xiaomi78.010000.440.8885841260
48
Q
Qwen3.8-27B
Qwen3.8 series open model (released Aug 14, 2026, Apache 2.0): 27B dense multimodal model with native image/video understanding, 262K native context (extendable to 1M via YaRN), Gated DeltaNet hybrid architecture, thinking mode on by default and toggleable per request. SWE-bench Pro 61.7, QwenSWEBench 79.0, DeepSWE 1.1 42.2, LiveCodeBench v6 90.3, OSWorld 84.3 - beats Claude Opus 4.6 Max on multiple coding and agent benchmarks. Runs locally on consumer hardware (~48 tps q4km on a 4090).
Alibaba (Qwen)78.010000.453.200—0
49
L
LongCat-2.0
1.6 trillion parameter MoE model from Meituan (LongCat), ~48B activated per token, trained on domestic AI ASIC superpods, 1M context window, MIT license
Meituan78.0104900———
50
M
Muse Glimmer
Meta Superintelligence Labs' open 30B on-device agentic model (Apache 2.0), distilled from Muse Spark, quantized to under 20GB for single-GPU local runs, with multimodal input and tool calling.
Meta (Superintelligence Labs)78.0128000—0
51
Q
Qwen3.6 Plus
Alibaba Qwen3.6 Plus mid-to-high-end model, 1000K context
Alibaba (Qwen)76.010000.331.9584841250
52
G
GPT-4o
OpenAI GPT-4o multimodal model with 128K context
OpenAI75.01282.501088.7090.201287
53
G
Gemini 3.5 Flash-Lite
Google's fastest 3.5-series model at 350 tokens/s, designed for low-latency, high-throughput agentic search and document processing. Priced $0.30/$2.50. SWE-Bench Pro 54.2%, OSWorld 74.0%. 1M context.
Google DeepMind75.010000.302.500—0
54
Q
Qwen3.7 Plus
Alibaba's flagship Qwen model with improved instruction following and tool use
Alibaba (Qwen)75.010000.321.28———
55
X
Xiaomi-Robotics-1
Xiaomi's embodied foundation model, 100K+ hours real-world pretraining, natural language instructions, 1600+ scenario generalization
Xiaomi75.0128000—0
56
G
GLM-5.1
Zhipu AI GLM series previous-gen flagship, 200K context
智谱AI (Zhipu)75.02000.401.2083821240
57
G
Gemini Robotics 2
Google DeepMind's next-gen embodied intelligence family: Gemini Robotics 2 (VLA whole-body control), Gemini Robotics ER 2 (embodied reasoning, on AI Studio) and On-Device 2 (on-device). Fine dexterity, multi-robot collaboration, adapts to new robot bodies in hours.
Google DeepMind74.0128000—0
58
C
Cursor Composer 2
Cursor Composer 2 AI coding assistant
Cursor72.02560082821220
59
M
MiMo-V2.5
Xiaomi MiMo-V2.5 on-device LLM
Xiaomi72.010490.150.29———
60
M
MiniMax-M2.7
MiniMax M2.7 chat model
MiniMax72.02050.281.2082811220
61
K
Kimi K2.5
Moonshot AI Kimi K2.5 chat model, 262K context
Moonshot AI72.02620.401.9082821220
62
G
Gemini 3 Flash
Google DeepMind lightweight Gemini model with 1000K context
Google DeepMind70.010000.150.608280.501220
63
G
GLM-5
Zhipu AI GLM-5 model, previous-gen flagship
智谱AI (Zhipu)70.02000.300.9081791210
64
Q
Qwen3.5 397B
Alibaba Qwen3.5 397B parameter large model
Alibaba (Qwen)68.02620.451.3580.5080.501200
65
I
Inkling
Open-weight MoE model by Thinking Machines Lab, 975B total / 41B active params. Multimodal (text, image, audio), 1M context. AA Intelligence Index 41 - leading US open weights model. Strong agent performance: Elo 1238 on GDPval-AA v2, beating Kimi K2.6 and DeepSeek V4 Flash.
Thinking Machines Lab68.0104914.050—1238
66
G
GPT-5.4 Mini
OpenAI's compact model in the GPT-5.4 series, optimized for efficiency
OpenAI67.04000.754.50———
67
Q
Qwen3 Coder 480B A35B
Qwen's most powerful open-source coding model. 480B MoE with 35B active params, native 256K context (YaRN scalable to 1M). Strong SWE-Bench performance. Apache 2.0 licensed. Ships with Qwen Code CLI.
Alibaba (Qwen)66.02560.221.80———
68
I
Inkling Small
Open-weights MoE from Thinking Machines Lab, 276B total/12B active, comparable performance to Inkling at a quarter of the size. Native audio+image reasoning, variable thinking effort, up to 1M context. Highest-scoring open-weight model on ARC Prize.
Thinking Machines Lab66.0524000—0
69
G
Gemini 2.5 Pro
Google DeepMind previous-gen Gemini high-end model
Google DeepMind65.010000.351.4080.50781180
70
G
Grok 3
xAI previous-gen Grok model with 1000K context
xAI65.010000.150.6080801180
71
H
Hunyuan Hy3 Preview
Tencent Hunyuan Hy3 Preview model, 256K context
Tencent Hunyuan65.02560.060.1879781180
72
S
Seed 2.0 Code
ByteDance Seed 2.0 Code model optimized for frontend development, multilingual coding, and agentic coding
ByteDance Seed65.02620.5030—0
73
P
Poolside Laguna S 2.1
Poolside coding-specialized agent model, 118B-A8B MoE, 1M context, Terminal-Bench 70.2%, DeepSWE 40.4%, focused on autonomous long-horizon engineering work
Poolside65.010000.100.20———
74
D
DeepSeek V3.2
DeepSeek's capable mid-range model offering strong performance at lower cost
DeepSeek63.01640.230.340—0
75
N
Nemotron 3 Ultra
NVIDIA's most capable open-weight model with strong benchmark performance
NVIDIA62.010000.502.20———
76
G
GPT-Live-1
OpenAI real-time voice conversation model, supports interruption and continuation, simultaneous voice chat and reasoning
OpenAI62.032000—0
77
C
Claude 4.5 Haiku
Anthropic lightweight fast model with 200K context
Anthropic60.02000.80478751150
78
G
Gemini 3.1 Flash Lite
Google's lightweight Gemini model optimized for speed and efficiency on edge devices
Google DeepMind60.010490.251.500—0
79
C
Codestral 2508
Mistral AI's specialized coding model, optimized for code generation, completion, and refactoring across multiple programming languages
Mistral AI60.02560.300.900—0
80
S
Step 3.7 Flash
StepFun's latest multimodal MoE model with 196B parameter language backbone and vision encoder for native image/video understanding
StepFun60.02560.201.157875—
81
G
GPT-5.4 Nano
OpenAI's compact model in the GPT-5.4 lineup, designed for rapid inference and cost-efficient deployment
OpenAI58.04000.201.25———
82
L
Laguna XS 2.1
Poolside Laguna XS 2.1 is a 33B-A3B MoE coding agent model optimized for local deployment. Builds on XS.2 with improved SWE-bench Multilingual (63.1%) and stronger terminal-style task performance. Supports 256K context, runs locally in vLLM/SGLang.
Poolside58.02560.060.12———
83
N
Nex AGI Nex-N2-Pro
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active / 397B total parameters. Built on Qwen3.5 architecture, supports text and image input, 262K context window. Extremely cost-effective at $0.25/M input tokens.
Nex AGI58.02620.251———
84
C
Cursor Composer 1.5
Cursor Composer 1.5 earlier version
Cursor58.02000076741150
85
Q
Qwen3.7 Flash
Alibaba's lightweight vision-language reasoning model in the Qwen3.7 family, 1M context, accepts image/video input, suited for multimodal agents, visual coding, search and computer interaction, priced at $0.03/$0.13 per M tokens
Alibaba (Qwen)58.010000.030.130—0
86
D
DeepSeek R1
DeepSeek R1 reasoning-enhanced model, 128K context, MIT open-source
DeepSeek55.01280.552.1978.5078.501100
87
N
Nova 2.0 Pro
Amazon Nova 2.0 Pro enterprise-grade model
Amazon55.02560.803.2076721120
88
N
Nemotron 3 Super
NVIDIA Nemotron 3 Super open-source model
NVIDIA55.010000.140.4276741120
89
S
Solar Pro 4
Upstage Solar Pro4 cost-efficient LLM with 524K context, ultra-low pricing
Upstage55.05240.030.120—0
90
M
Mistral Large 2512
Mistral AI's flagship large language model with strong reasoning and multilingual capabilities
Mistral AI55.02620.501.500—0
91
M
Mistral Medium 3.5
Mistral AI's mid-range model balancing performance and efficiency for production deployments
Mistral AI55.02621.507.500—0
92
C
Cohere North Mini Code
Cohere North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse MoE with 30B total / 3B active parameters, optimized for end-to-end coding tasks. 256K context window, text input only.
Cohere55.025600———
93
S
Step 3.5 Flash
StepFun Step 3.5 Flash fast model
StepFun (阶跃星辰)55.02560.030.0975721100
94
D
Doubao Seed Code
ByteDance Doubao Seed Code coding model
字节跳动 (ByteDance)55.02560.100.3076741120
95
M
Mistral Large 3
Mistral AI previous-gen flagship with 256K context
Mistral AI50.02560.300.9075701100
96
G
Grok Build 0.1
xAI's coding-focused model trained for agentic software engineering workflows, supports text+image input
xAI50.025612———
97
C
Command A+
Cohere Command A+ enhanced enterprise-grade generation model, 128K context
Cohere48.0128000—0
98
L
Llama 4 Maverick
Meta Llama 4 Maverick open-source model with 1000K context
Meta45.010000.170.5072721080
99
E
ERNIE 5.0 Thinking
Baidu ERNIE 5.0 Thinking reasoning-enhanced
百度 (Baidu)45.01280.250.7570681050
100
G
Granite 4.1 8B
IBM's open-source 8B dense enterprise model, matching 32B MoE performance, 131K context, designed for enterprise tasks
IBM40.01310.050.10———

Disclaimer

This leaderboard is editorially curated and updated by LinkWord. Composite scores, benchmark figures, pricing, and capability metrics are provided for browsing and rough comparison only—not as purchase advice, performance guarantees, or endorsements. Data may lag official vendor releases; results vary by evaluation method, model version, and use case. Third-party trademarks belong to their respective owners. You assume all risk from relying on this page; if you cite or republish, please attribute the source.