Mistral is opening its newest flagship AI model to enterprise testing before the full release is finished.
The company opened Large 4’s public API preview on Oct. 6 for tasks spanning multiple steps and applications. The company says downloadable weights and further technical details will follow by the end of October.
Enterprise teams now have to judge whether those capabilities translate into better work inside their own systems. Performance, control, and cost will determine how far the new model gets past the testing stage.
Workplace tests move past chatbot prompts
Mistral Large 4, which the company has nicknamed Le Chonk, is being tested on jobs that take several steps instead of one prompt and one response. Company results cover work across business applications, software development, and security.
- Business software. Large 4 scored 59.9% on AutomationBench, which covers 657 workflows using applications such as Gmail and Salesforce. The test measures whether an AI system can carry out a sequence of tasks across workplace software.
- Coding. Results reached 49.8% on the Coding Agent Index and 61.7% on DeepSWE v1.1. Both evaluations test how well an AI system works through software-development jobs that involve several steps.
- Cybersecurity. Large 4 scored 82% on an evaluation involving reproduction of a real software vulnerability followed by a repair. The company also reported 93.3% resistance on a test that tries to trick an AI system into following malicious instructions.
Professional work is part of the release as well. The tool can edit spreadsheets and documents and work with visual material. Its AA-Briefcase result reached 1,393 Elo on an evaluation of extended knowledge-work assignments.
Security scores alone do not establish whether an AI-generated fix will behave safely in production. Research into AI-generated vulnerability fixes found that closing an exploit did not guarantee the patch preserved normal software behavior. Companies testing Large 4 for security work should check generated fixes against expected application behavior before allowing automated changes in production.
Planned downloadable weights would enable self-hosting
Large 4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in data centers across Europe. Preview API traffic runs on infrastructure operated by the AI company, which also plans a regional deployment governed by European law. Regional infrastructure is part of its AI strategy
Downloadable weights will let security organizations run the model in a private cloud or on their own premises under internal policies. Companies handling sensitive code or proprietary documents could keep greater control over where AI work runs. Operating a model internally also brings hardware, staffing and maintenance costs that a hosted API keeps with the provider.
Some information needed to plan an internal deployment, like architecture details and post-training information, are still pending. Teams tracking the open-weight AI market can test the hosted version now, but sizing their own infrastructure will depend on the remaining technical details.
What eWeek found: Large 4 expands context while raising API prices
eWeek compared Large 3 and Large 4. Large 4 offers about 3.9 times the context capacity, while regular input and output prices rise about 172% and 179%. The table uses the launch announcement’s 49 billion active parameters; Mistral’s model card lists 52 billion, implying an increase of about 27% rather than 20% over Large 3.
| Measure | Large 3 | Large 4 | Difference |
| Total parameters | 675 billion | 1.05 trillion | About 56% higher |
| Active parameters per token | 41 billion | 49 billion | About 20% higher |
| Maximum context | 256,000 tokens | 1 million tokens | About 3.9x |
| Regular input price per 1M tokens | $0.50 | $1.36 | About 172% higher |
| Regular output price per 1M tokens | $1.50 | $4.18 | About 179% higher |
A mixture-of-experts model uses only part of its full parameter pool for each token it processes. Active parameters indicate how much of that network participates at each step. Large 4 therefore grows substantially in total size without a matching increase in active parameters, although parameter counts alone cannot establish speed, memory use, or serving cost.
Context capacity rises from 256,000 tokens to 1 million. A team could send a longer contract or a larger code repository in one request before reaching the limit. Bigger prompts can also mean more billable input tokens, so access to a larger context window can increase spending when teams use it heavily.
A two-week launch discount temporarily lowers Large 4 pricing to $0.68 per million input tokens and $2.09 per million output tokens. Teams considering an enterprise model upgrade should run the same representative task through both versions, then compare completion rate and final API cost with token use and latency.
Even during the launch discount, Large 4’s input price is 36% higher and its output price about 39% higher than Large 3’s. Enterprise teams should compare cost per successfully completed task: fewer retries or calls could offset the higher rates, but that needs to be demonstrated on representative company workloads.
More AI news: Wikimedia says heavy traffic from suspected OpenAI agents strained its infrastructure, though it has not established that the agents caused the outage.


