Mistral AI has released Leanstral 1.5, a powerful new open-source model designed to bring formal verification — the mathematical proof that software behaves exactly as intended — to every developer.

Unlike conventional AI coding assistants that generate code prone to bugs, Leanstral 1.5 specializes in Lean 4, a functional programming language and proof assistant. It can write, check, and debug formal proofs that guarantee software correctness at the deepest level.

Math Benchmarks Shattered

Leanstral 1.5 achieves remarkable results on rigorous mathematical benchmarks. It completely saturates miniF2F (100% on both validation and test sets), solves 587 out of 672 problems from the prestigious Putnam Mathematical Competition, and sets a new state-of-the-art on FATE-H (87%) and FATE-X (34%), which test graduate and PhD-level abstract algebra.

Perhaps most impressive is its cost efficiency. Each Putnam problem solved costs approximately $4, compared to an estimated $300 or more for competing approaches running on massive compute budgets.

Real-World Bug Catching

Beyond benchmarks, Leanstral 1.5 demonstrated practical value by automatically discovering bugs in real software. An automated pipeline translated Rust code into Lean, had Leanstral infer correctness properties, and attempted to prove them. Across 57 open-source repositories, the system flagged 47 violated properties, with 11 pointing to genuine bugs — 5 of them previously unreported on GitHub.

One example: Leanstral caught an integer overflow in the sign function for zigzag decoding of the `datrs/varinteger` library. On input `Std.U64.MAX`, the expression `(value + 1)` overflowed, causing crashes in debug mode and silent corruption in release — an edge case that traditional testing and fuzzing would typically miss.

Open and Accessible

Leanstral 1.5 is released under the permissive Apache-2.0 license. Despite having 119 billion total parameters, it uses a Mixture-of-Experts architecture with only 6 billion active parameters per forward pass, making it efficient to run. The weights are available on Hugging Face, and Mistral offers a free API endpoint.

The model is trained through a three-stage process: mid-training on mathematical and code data, supervised fine-tuning, and reinforcement learning with a novel CISPO algorithm. It operates in two environments — a multi-turn theorem-proving loop with Lean compiler feedback, and a full code agent environment where it can edit files, run bash commands, and use the Lean language server.

Implications for Software Reliability

Formal verification has long been considered too expensive and specialized for mainstream development. Leanstral 1.5 challenges that notion by making proof engineering practical and cost-effective. As AI systems take on increasingly critical roles in healthcare, finance, and infrastructure, the ability to mathematically guarantee software behavior becomes not just valuable but essential.

With Leanstral 1.5, Mistral is betting that the future of coding isn't just about generating more code — it's about proving that code is correct before it ever runs.