Mercury 2.5 is a reasoning model built by Inception, and it posts a pure output speed that sets it apart from its peers. The model generates text at 780.8 tokens per second, which is a genuine measure of how fast a model processes information. That number comes from a new analysis of the model’s performance, and it arrives at a moment when the AI industry is full of announcements and claims.
artificialanalysis.ai’s analysis comes from a benchmark called the Artificial Analysis Intelligence Index, which measures models across reasoning, knowledge, mathematics and coding. Mercury 2.5 scores 12 on that index, placing it below average among comparable models. Its median comparison group sits at 13.
The Speed Numbers
Mercury 2.8 generates output at 780.8 tokens per second, according to Inception’s API. That is well above the median of 111.9 tokens per second for reasoning models in a similar price tier.
The model’s time to first token — the latency between asking a question and getting the first word back — is 3.06 seconds. That is somewhat higher than the median of 2.21 seconds for models in the same tier.
Here is how the three core numbers stack up:
| Metric | Mercury 2.5 | Median |
|---|---|---|
| Output speed | 780.8 tokens per second | 111.9 tokens per second |
| Time to first token | 3.06 seconds | 2.21 seconds |
| Intelligence score | 12 | 13 |
What the Model Does
Mercury 2.5 is a reasoning model, which means it uses extended thinking or chain-of-thought reasoning to work through complex problems before providing an answer. It supports text input and text output. It does not handle images. It is not multimodal. It has a context window of 260k tokens, which determines how much text and conversation history the model can process in a single request.
The model was released on September 8, 2026. It is proprietary, meaning the model weights are not publicly available. Inception has not disclosed the model size or parameter count.
The Intelligence Test
When evaluated on the Intelligence Index, Mercury 2.5 generated 35M output tokens. That is better than average compared to the median of 85M tokens. The model’s 35M token output is fairly concise, which is a notable result for a model that is also notably fast.
The Intelligence Index is a composite benchmark that tests models across reasoning, knowledge, mathematics and coding. It includes specific sub-tests, such as professional document reasoning, physics reasoning, quantitative analysis on spreadsheets and documents, and medical long context reasoning.
Mercury 2.5 is available via API through one provider. Pricing is $0.25 per 1M input tokens and $0.75 per 1M output tokens, based on Inception’s API. For a blended rate — a 7:2:1 cache hit/input/output ratio — that works out to $0.14 per 1M tokens. Pricing may vary by provider.
The Pricing Picture
Pricing is where Mercury 2.5 makes its strongest case. At $0.25 per 1M input tokens, the model sits at the median price for reasoning models in its tier. The median price for output tokens is $0.92.
On average, it costs $0.06 per task to evaluate Mercury 2.5 on the Intelligence Index. That is a small number, and it reflects the model’s moderate pricing overall.
The model is also notably fast and fairly concise, which is a combination that matters for users who need quick answers and do not want to pay for excess output. The analysis describes the model as well priced when compared to other models of similar cost.
The Trade-Offs
Every model has a weakness, and Mercury 2.5 is no exception. Its intelligence score of 12 puts it below average on the Artificial Analysis Intelligence Index. That is the trade-off for its speed and price.
The model’s 3.06-second latency is somewhat higher than the median of 2.21 seconds. That means the first word takes longer to arrive, though the source does not describe how subsequent tokens behave.
The model’s 260k-token context window is generous, but it is not unlimited. Users who need to reference very long histories will hit that ceiling eventually.
Ranking the Core Metrics
The three metrics that define Mercury 2.5 are output speed, latency and intelligence score. Here is how they rank against the median for reasoning models in a similar price tier:
- Output speed — 780.8 tokens per second vs. median of 111.9 tokens per second. Well above average.
- Latency — 3.06 seconds vs. median of 2.21 seconds. Somewhat higher than average.
- Intelligence score — 12 vs. median of 13. Below average.
Why This Matters
A model’s speed is not just a feature. It determines how fast a prompt can be answered, how quickly a workflow can move, and how long a user has to wait for a reply. A fast model can keep a conversation moving, while a slow model can grind it to a halt.
The 780.8 tokens per second figure is a real measure of processing power. It tells you how much text the model can spit out in a second, which is a practical number for anyone building a system around it.
The Bottom Line
Mercury 2.5 is a solid, well-priced reasoning model that happens to be exceptionally fast. Its intelligence score of 12 puts it below average, but its output speed of 780.8 tokens per second and its $0.06 per task evaluation cost give it a strong value proposition.
The model is available via API through one provider, and its input price of $0.25 matches the median for its tier. The 3.06-second latency is a mild concern, but the combination of speed, price and output economy makes it a compelling option for users who need quick answers without paying a premium.
For anyone shopping for a reasoning model, Mercury 2.5 deserves a look. The numbers are real, the pricing is fair, and the speed is the kind of number that gets noticed.
The quiet number — 780.8 tokens per second — is the story.
Source material: “Mercury 2.5 LLM hits 770 tokens per second,” artificialanalysis.ai.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

