Midterms 2026See who we think should earn your vote, based on our standardsThe guide →
WRITTEN IN PLAIN AMERICAN ENGLISH.
CLAY TRIBUNE.
Advertisement

Mistral Opens Price Sheet on Large 4 Model, Leaving Users to Do the Math

Mistral unveils Mistral Large 4, a massive multimodal model with a curious pricing table no one can quite figure.

By mitch·4 min read
A glowing digital dashboard displaying abstract charts and numbers, symbolizing the complex pricing structure of Mistral's new model.

Mistral AI has opened the curtains on Mistral Large 4, its latest multimodal model, and the company has done something unusual: it has published a price sheet that reads like a puzzle wrapped in a table. The details are available on docs.mistral.ai, where the announcement landed today and drew attention across the web.

The Numbers That Won’t Add Up

Mistral Large 4 is described as a general-purpose model built around a Mixture-of-Experts architecture. It carries 49 billion active parameters and 1.05 trillion total parameters, with a vision encoder that weighs in at 1.6 billion. The model handles text, images, and audio across multiple modalities, though the exact count of those modalities is not stated.

What is stated is the cost. For uncached requests, the model charges $1.36 per million tokens on input, $0.68 per million tokens on output, and $0.14 per million tokens on cached input. The cached output rate sits at $2.09 per million tokens. There is also a public preview version, labeled v26.10, dated October 6, 2026.

Advertisement

The pricing table mixes unit prices, raw parameter counts, and context limits without explaining any of them. A reader who wants to know how much a full request costs has to do the math themselves, and the result is a curiosity rather than an announcement.

What the Table Actually Shows

The table is dense and oddly specific. It breaks down the cost of each operation in terms of tokens, which are the units models use to measure text input and output. The cached rates are higher than the uncached ones, which is a standard arrangement in the field — cached data costs more because it is stored rather than processed anew.

Mistral Large 4 is a state-of-the-art, open-weight, general-purpose multimodal model with a granular Mixture-of-Experts architecture.

That line, taken directly from the announcement, is the closest thing the page offers to a mission statement. The Mixture-of-Experts architecture is a design pattern where a large model splits its work across many smaller experts, each handling a specific kind of task. It is a common approach in the field, though the announcement does not explain how Mistral Large 4’s particular arrangement compares to other models.

The Missing Context

The announcement does not say what the model is used for, who it is aimed at, or how it compares to competitors. It lists technical details and prices without framing them in terms of performance or capability. The reader is left to wonder whether the 1.05 trillion parameters translate into better image captioning, faster text generation, or improved speech synthesis — the announcement simply does not say.

Why the Announcement Matters Anyway

The sheer scale of the model is notable. 1.05 trillion parameters is a large number, and the vision encoder’s 1.6 billion parameters suggest serious image-processing capability. The Mixture-of-Experts architecture is a genuine advance in model design, even if the announcement does not explain it in depth.

The pricing structure is also worth noting. The cached rates are significantly higher than the uncached ones, which reflects a real-world trade-off: storing data is cheaper than processing it from scratch, but the economics change when you charge by the token.

The Verdict on the Announcement

The announcement is a technical data dump dressed up as a product launch. It tells you what the model can do in broad strokes — multimodal, general-purpose, Mixture-of-Experts — and it tells you how much it costs to use. What it does not tell you is why you would want to use it.

That is a fair complaint, but it is not a fatal one. The model exists, the numbers are on the table, and anyone who wants to benchmark it against other systems now has a starting point. The announcement is incomplete, but it is complete enough to be useful.

Key Facts Box

  • Model: Mistral Large 4
  • Architecture: Mixture-of-Experts
  • Active parameters: 49B
  • Total parameters: 1.05T
  • Vision encoder: 1.6B
  • Public preview: v26.10, October 6, 2026
  • Uncached input: $1.36 per million tokens
  • Uncached output: $0.68 per million tokens
  • Cached input: $0.14 per million tokens
  • Cached output: $2.09 per million tokens

The announcement is a technical data dump dressed up as a product launch. It tells you what the model can do in broad strokes — multimodal, general-purpose, Mixture-of-Experts — and it tells you how much it costs to use. What it does not tell you is why you would want to use it. That is a fair complaint, but it is not a fatal one. The model exists, the numbers are on the table, and anyone who wants to benchmark it against other systems now has a starting point.

Source material: “Mistral Large 4,” Mistral AI.

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.