Mistral Large 4: What a 1T Open Weight Model Costs You
Mistral Large 4 ships 1T parameters at $1.36/$4.18 per million tokens. What that price gap means for heavy AI users weighing open weight models.
Searches for “open weight model” are up 601 percent in three months. That is not a niche curiosity anymore. It is a buying signal from people who pay real money for tokens and want a second exit besides the closed frontier they already rent.
On October 6, 2026, Mistral AI gave that search a destination. The French lab launched the public preview of Mistral Large 4, nicknamed Le Chonk, a 1 trillion parameter natively multimodal model that it positions as an alternative to both closed American models and the open Chinese wave. The headline number for anyone tracking the “open weight model” trend is not the parameter count. It is the price: $1.36 per million input tokens and $4.18 per million output tokens.
That is the conversation. Not what ML4 scores on a benchmark, but what it does to the arithmetic of a heavy user’s monthly bill. This is a breakdown of where ML4 lands, what the open weight promise actually delivers, and how to test it without burning a weekend.
What Mistral Large 4 actually is
ML4 is a 1 trillion parameter model with 49 billion active parameters, trained from scratch on 3,800 to 4,000 Nvidia Grace Blackwell GPUs in Mistral’s own European datacenters. It is natively multimodal, spanning coding, agentic workflows, visual grounding and document intelligence, and it ships with a stated fluency in more than 160 languages.
The nickname Le Chonk is a self-aware joke about the parameter count. The substance is elsewhere: Mistral says ML4 outperforms every open weight model developed in the US or Europe, reaches state of the art among open models on cybersecurity, finance and law workloads, and in some domains like visual grounding surpasses frontier closed models.
The one caveat that matters for planning: ML4 is not fully open weight yet. Today it is served through a guarded public endpoint on Mistral Studio. The weights drop at the end of October, three weeks out, after red-teaming with cybersecurity partners and state authorities. So the model you can call today is the hosted preview. The model you can self-host is the promise.
That distinction splits the “open weight ai” pitch into two phases, and each phase has a different cost profile. Anyone planning a migration on October 7 should plan for both.
The pricing math that matters for heavy users
Fifty cents per million tokens is not the interesting figure. The interesting figure is the ratio between the closed frontier you probably use and the open option sitting next to it.
A heavy user spending $300 to $500 a month on a frontier subscription or API is not paying for intelligence. They are paying for a bundle: quota, reliability, the newest weights, and the convenience of not running anything. ML4 at $1.36 in and $4.18 out attacks one leg of that bundle directly. Compare it against frontier API pricing that for comparable multimodal work often runs several times higher on input and two to four times higher on output.
The lesson from the “best open source llm” searches and the “ai token cost” surge (up 54 percent) is the same: people are no longer asking whether an open model is good enough in the abstract. They are asking whether the per token delta justifies the operational work.
Here is the honest version of the trade. A closed frontier API charges a premium for zero maintenance, guaranteed uptime, and the newest capability the day it ships. An open weight deployment trades that premium for a fixed infrastructure bill and a variable labor bill. The token cost goes down. The operational cost goes up. Whether the net wins depends entirely on how steady your volume is.
Steady volume favors open weights. Spiky, unpredictable, and latency-critical volume still favors the closed API. ML4 does not change that rule. It changes the size of the gap on the open side.
Open weight versus open source: the distinction that trips planning
A lot of the confusion in the “open weight ai” trend comes from blending two different things.
Open source, in the strict software sense, means the training data, the code and the weights are available for you to inspect, modify and redistribute. Very few frontier models qualify at the training data level.
Open weight means the trained parameters are downloadable and runnable. You can host them, fine tune them, and crucially, you cannot be unplugged mid-incident. Mistral’s own framing leans hard on this: in cybersecurity, a provider level refusal can block legitimate vulnerability research, and losing access to a capability during an incident is itself a risk.
For a heavy user, open weight is the property that matters, not open source. You rarely need to retrain the model. You need the guarantee that the endpoint answering your requests is one you control and cannot have removed underneath you.
This is the same logic that pushed the open weight model trend up 601 percent. It is less about ideology and more about continuity risk and cost predictability.
Where ML4 lands against the rest of the field
Mistral trained ML4 on two to three times less compute than its Chinese competitors and significantly less than the closed labs, and it says so loudly. The strategic claim is a “third way”: not a closed model that can be switched off, not an open model that arrives with geopolitical strings, but a European open weight option with enterprise backing from ASML and Samsung.
That positioning is not just marketing. It answers a real question for enterprises and institutions who want open weights but have procurement rules about where the weights come from.
Against the confirmed landscape, ML4 slots in as the largest open weight multimodal model outside China. Against ChatGPT, Claude and Gemini, it is the cheaper hosted option that will soon become a self-hostable one. Against the open Chinese wave, it is the same performance tier with a European compute lineage and a security-cleared release path.
Best open source llm lists will need updating, but the practical comparison is not leaderboards. It is your own workload, priced per task.
What to do this month before the weights drop
The weights arrive at the end of October. Use the preview window to make the migration decision with data instead of vibes. A few moves that pay off.
Instrument one workflow, not everything. Pick a single repetitive job you already run, one with steady volume and closed answer evaluation. Log the tokens in and out per task on your current provider. That baseline is the only number that makes the ML4 comparison real.
Price per completed task, not per token. The per million rate is where the pitch lives. The number that hits your card is cost per completed task, including retries and failed generations. A model that is cheaper per token but needs an extra retry is not cheaper.
Shadow test on the preview endpoint. Route a copy of the live traffic to ML4 and compare output quality on your own criteria. Do not trust a benchmark you did not run.
Model the self-host case separately. Once the weights are out, the cost shifts from per token to GPU hours plus engineering time. Estimate the GPU bill at your actual concurrency and add a realistic hours figure. Many teams find the crossover sits higher than they assumed.
Keep a closed lane open. The point of an open weight hedge is optionality, not replacement. The teams that handle provider shocks best are the ones that already have a second path wired up, even if it carries only ten percent of traffic.
The bottom line for a $300-plus a month user
Mistral Large 4 does not make your bill cheaper today. It makes the cheaper option more credible, and that is the part that changes behavior over the next quarter.
The open weight model trend is rising because heavy users have started treating provider risk and token cost as line items they can actually renegotiate. ML4 puts a 1 trillion parameter multimodal model at $1.36 and $4.18 per million tokens, with weights three weeks away and a European control story attached. For anyone paying off-plan prices every month, that is worth the two hours it takes to instrument one workflow and run the shadow test.
The closed frontier still owns the top of the capability curve and the easiest path to production. But the open weight side just got a bigger, multimodal, cheaper entrant, and the search volume says a lot of you are already doing the math.
Now available
Stop guessing your AI limits
The Mac app and web dashboard watch your Claude, ChatGPT, Gemini and more, and warn you before quotas hit.