presentofai

Mistral launches Mistral Large 4, a 1T-parameter open-weight model in preview

TL;DR

Mistral releases Large 4, a 1-trillion-parameter open-weight model that beats several top open-source rivals and keeps training running continuously to produce even larger versions.

What happened

  • Mistral Large 4 launched October 6, 2026, in public preview on Mistral's cloud platform, with weights to follow later in October.
  • 1 trillion total parameters, mixture-of-experts architecture activating only 49 billion at a time for hardware efficiency.
  • Trained on 3,800 Grace Blackwell chips (each pairing two Nvidia Blackwell GPUs with one CPU), generating 33 billion tokens per day during rollouts.
  • Mistral's custom software stack runs tens of thousands of rollouts in parallel, asynchronously from the training workflow to avoid bottlenecks.
  • Training has not stopped: Mistral is running the same workflow continuously, expecting larger, more capable variants in coming months.

Why it matters

  • Open-weight at 1T parameters is a significant threshold, giving enterprises and researchers access to frontier-class scale without vendor lock-in.
  • Outperforms Qwen3.8 Max and DeepSeek V4 Pro on AutomationBench and AA-Briefcase, the latter covering multi-week knowledge-work tasks, signaling real enterprise utility.
  • Scored 82% on the AA Cyber Index for patching open-source projects, ahead of open-source rivals, making it a credible tool for security teams.
  • Beats GPT-6 Astra by 1% on the Dense200 vision benchmark, a rare open-source win over a frontier closed model.
  • The continuous training pipeline means Large 4 is a platform, not a one-off release: a full series of use-case-optimized models is planned, compressing Mistral's future release cadence.

What to watch next

  • Weight release timing: public weights later in October will determine how fast the open-source community fine-tunes and deploys at scale.
  • Benchmark gaps on coding: Large 4 still trails frontier models like Astra on popular coding tests; watch whether the continuous training run closes that gap in the next iteration.
  • Competitive response from DeepSeek and Qwen: both were beaten on knowledge-work benchmarks, and their next releases will signal whether Mistral's efficiency architecture is a durable advantage.

Originally published on Present of AI, a daily source-linked AI news timeline. Read the full timeline or browse the open dataset.