Mercury 2.5 benchmarked: 770 tokens per second but below-average intelligence on Artificial Analysis

Mercury 2.5, released in September 2026, is one of the most interesting speed-focused models on Artificial Analysis right now. It posts 770.4 output tokens per second, ranking #2 of 175 models for speed, but its Intelligence Index score of 12 sits below the median of 13. For teams choosing between throughput and accuracy, this trade-off deserves a closer look.

Speed at a glance: why 770 tokens per second matters

Latency-sensitive applications such as streaming chat, coding copilots, and real-time translation benefit directly from high output speed. At 770 tokens per second, Mercury 2.5 can deliver long answers faster than most competitors, which improves perceived responsiveness and throughput in multi-user systems.

The model supports a 260k-token context window, roughly 390 A4 pages. That is large enough for long-form document processing, legal review, and extended coding sessions. For applications that chunk prompts into large contexts, the speed advantage compounds across multiple turns.

Intelligence trade-off: below-average benchmark score

Artificial Analysis Intelligence Index places Mercury 2.5 at 12 out of 175 models, below the median of 13. In benchmark tasks, it generated 35M tokens, far below the median of 85M. That suggests the model prefers shorter outputs and may struggle with problems that require extended reasoning chains.

Pricing is moderate: $0.25 per 1M input tokens and $0.75 per 1M output tokens. With a 90% cache discount, repeated requests can drop to $0.06 per Intelligence Index task. That makes it price-competitive among similarly priced models, but buyers should still compare intelligence-per-dollar carefully.

Who should use Mercury 2.5?

  • High-volume chatbot operators: speed reduces queue time under load.
  • Streaming and real-time UIs: faster token delivery improves UX.
  • Budget-conscious experimentation: moderate pricing plus caching keeps costs predictable.

Teams that need deep reasoning, complex math, or long-form analysis may still prefer higher-intelligence models despite slower speeds. Mercury 2.5 is best treated as a throughput specialist rather than a reasoning powerhouse.

Market context

The model comes from Inception and was released in September 2026. Its positioning resembles other speed-first entrants: sacrifice some intelligence for dramatically better latency and lower per-request cost. As inference hardware improves, this segment may grow, especially for edge deployments and consumer-facing assistants.

Developers should benchmark Mercury 2.5 against their actual task distribution. A model that is fast but occasionally shallow may still win if most user requests are short and latency matters more than depth.

Bottom line

Mercury 2.5 is a clear example of the specialization trend in LLMs. It is not the smartest model, but it is one of the fastest at a reasonable price. For latency-critical workloads, it is worth testing; for deep reasoning, look elsewhere.

Related reading:

📤 Share this article
Weibo |
Twitter |
LinkedIn

📬 Subscribe to AI News

Daily AI tool reviews and usage tips


发表评论