Mistral Release Notes

Follow

113 release notes curated from 55 sources by the Releasebot Team. Last updated: Sep 30, 2026

Get this feed:

Mistral Products

  • Sep 28, 2026
    • Date parsed from source:
      Sep 28, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Mistral logo

    Mistral

    September 28

    Mistral deprecates OCR 4.0 and Leanstral 1.5 and points users to OCR 4.1 and GLM 5.3 at the same price.

    DEPRECATED

    OCR 4.0 (mistral-ocr-4-0) is deprecated and retires on September 30, 2026. Use OCR 4.1 (mistral-ocr-4-1) or mistral-ocr-latest) instead, at the same price.

    DEPRECATED

    Leanstral 1.5 (labs-leanstral-1-5), an experimental Labs model, is deprecated and retires on September 30, 2026.

    DEPRECATED

    Z.ai GLM 5.2 (zai-glm-5-2) is deprecated and retires on October 31, 2026. Use Z.ai GLM 5.3 (zai-glm-5-3) instead, at the same price.

    DEPRECATED

    Original source
  • Sep 27, 2026
    • Date parsed from source:
      Sep 27, 2026
    • First seen by Releasebot:
      Sep 30, 2026
    Mistral logo

    Mistral

    September 27

    Mistral notes Z.ai GLM 5.3 is now generally available.

    Z.ai GLM 5.3 (zai-glm-5-3) is now Generally Available.

    MODEL RELEASED

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Mistral and hundreds of other software products.

    Create account
  • Sep 22, 2026
    • Date parsed from source:
      Sep 22, 2026
    • First seen by Releasebot:
      Sep 22, 2026
    Mistral logo

    Mistral Common by Mistral

    v1.12.0: Agnostic validation mode

    Mistral Common releases 1.12.0 with broader validation, safer multimodal handling, and chat template improvements. It also fixes image and audio edge cases, tightens tool call and message checks, and refreshes docs, contribution guidance, and release tooling.

    What's Changed

    • Add comment and coverage guidelines to AGENTS.md by @juliendenize in #278
    • Allow overriding auto-detected plain thinking support in chat template generation by @juliendenize in #285
    • Rename LICENCE to LICENSE by @juliendenize in #286
    • Set reference audio filename for speech multipart uploads by @Mr-Neutr0n in #284
    • Keep at least one pixel per axis when resizing an image by @arthi-arumugam-git in #289
    • Reject AudioConfig where audio_length_per_tok floors to zero by @shoemoney in #292
    • Deprecate Tokenized.text in favor of tokenizer.decode by @juliendenize in #282
    • fix(tokenizer): validate ImageConfig fields to prevent ZeroDivisionError by @shoemoney in #294
    • Update uv tooling to 0.12.6 by @juliendenize in #297
    • Parallelize unit and integration test suites by @juliendenize in #298
    • Reject null tool call IDs before encoding by @juliendenize in #293
    • Deprecate direct continue_final_message on ChatCompletionRequest by @juliendenize in #303
    • Improve typing at dynamic tokenizer boundaries by @juliendenize in #296
    • Align tokenizer type contracts by @juliendenize in #299
    • docs: fix typo in experimental usage guide by @ManoharPaturi in #306
    • Add contribution guidelines and templates by @juliendenize in #314
    • Simplify issue creation by @juliendenize in #315
    • Switch documentation from MkDocs to Zensical by @juliendenize in #318
    • Fix docs deploy: replace non-existent zensical gh-deploy with actions-gh-pages by @juliendenize in #319
    • Fix actions-gh-pages ref to real v4 commit SHA by @juliendenize in #320
    • Improve docstrings across the codebase by @juliendenize in #317
    • Add agnostic validation mode by @juliendenize in #295
    • Raise InvalidToolCallError on malformed v2-v7 tool call outputs by @Chocolatine75 in #310
    • Raise InvalidMessageStructureException for invalid continue_final_message by @juliendenize in #325
    • fix(image): preserve float32 dtype and avoid float64 promotion in normalize (#316) by @andriiorap in #321
    • Allow null assistant content in OpenAI requests by @shubhangi013 in #327
    • Fix image URL loading crashes and add a download timeout by @BarneyChambers in #308
    • Add test-compat label workflow: vLLM + Transformers compat tests on CPU by @juliendenize in #334
    • Release 1.12.0 by @juliendenize in #326

    New Contributors

    • @Mr-Neutr0n made their first contribution in #284
    • @arthi-arumugam-git made their first contribution in #289
    • @shoemoney made their first contribution in #292
    • @ManoharPaturi made their first contribution in #306
    • @Chocolatine75 made their first contribution in #310
    • @andriiorap made their first contribution in #321
    • @shubhangi013 made their first contribution in #327
    • @BarneyChambers made their first contribution in #308

    Full Changelog: v1.11.7...v1.12.0

    Original source
  • Aug 30, 2026
    • Date parsed from source:
      Aug 30, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Mistral logo

    Mistral

    August 30

    Mistral OCR 4.1 is now generally available, expanding the company’s OCR offering.

    OCR 4.1 (mistral-ocr-4-1) is now Generally Available.

    MODEL RELEASED

    Original source
  • Aug 20, 2026
    • Date parsed from source:
      Aug 20, 2026
    • First seen by Releasebot:
      Aug 20, 2026
    Mistral logo

    Mistral

    Agentic Search. More accurate and efficient results from your AI systems.

    Mistral introduces Agentic Search, a new retrieval layer that helps AI systems navigate, read, and verify complex documents with better accuracy, fewer turns, lower token use, and reduced latency. It is available through Search Toolkit and Libraries.

    Mistral Agentic Search delivers more accurate search results while reducing turns, token use, and latency against FinanceBench and OfficeQA Pro benchmarks

    Agentic Search is the retrieval layer that enables AI systems to navigate, read, and verify information inside even the most complex documents. Available through Mistral Search Toolkit and Libraries.

    Mistral Agentic Search helps enterprises get better results from their AI systems by letting models search and navigate their organization’s most complex data and documents. Agentic Search introduces a multi-step retrieval loop for finding, inspecting, and verifying information across data sources, wherever it is stored. Agentic Search is available through Mistral Search Toolkit, built into Libraries in both Studio and Vibe, and gives you:

    • Support for sensitive domain-specific data. Mistral’s portable and open tooling helps you unlock value from your data without crossing your isolation boundaries in the cloud or on-premises.
    • Improved search results. Your models can search and navigate your data beyond retrieved chunks–inside long, dense documents or across multiple sources.
    • Access to existing indexes. Agentic Search builds on your existing search index using five tools: search, open, navigate, read, and grep.
    • Higher accuracy. Agentic Search delivers to 3x correctness on financial filings, from 26.7% to 86%, based on FinanceBench. On table-heavy, multi-doc questions of the OfficeQA Pro benchmark, we measure a +45.6 point gain (6.3% to 51.9%).
    • Lower latency and token use. Targeted navigation enables Agentic Search to reduce p90 latency up to 39.6%. Fewer repeated searches reduce token consumption by up to one-third.

    Data creates competitive advantage

    Competitive edge is built upon years of real-world operations–your data, your processes, and your domain expertise. Proprietary knowledge is both critical to your success and highly confidential, meaning it lives behind isolation boundaries, segmented deployments, and self-hosted platforms. It accumulates in financial filings, legal contracts, internal resources, and government records–long, dense documents that traditional search methods can’t navigate effectively.

    Agents that learn and improve continuously can help you compound your competitive advantage, but these agents are often separated from confidential data and proprietary knowledge for security reasons. Getting real impact from AI means pairing frontier reasoning with retrieval tools that can safely reach your most sensitive material.

    Traditional RAG falls short

    Traditional, one-shot RAG retrieves a fixed set of text chunks and asks a model to answer in a single pass. This works when the answer appears in one of the top results, but falters when the model must navigate a long report, follow references, compare multiple documents, or verify the underlying evidence.

    The limitation is more pronounced on dense, complex data and documents. The information needed to answer a question may be spread across documents or buried in a particular table, footnote, or clause. One-shot RAG-based search fails to use the full power of frontier AI and to provide reliable answers for three reasons:

    • Retrieval without reasoning: The model must answer from the chunks selected during the initial retrieval, even when they are incomplete or not relevant. It cannot decide that it needs a different document, another section, or more context before responding, which limits the impact of the model’s reasoning.
    • Chunk-level limit: Critical data is often held in complex multi-modal documents. When asked, “What was the company’s effective tax rate in Q3?” an index may find the correct document but cannot open it, navigate to the table, read the surrounding context, or verify the answer.
    • No iteration: Many questions need more than one retrieval pass to get the correct answer. The model may need to refine its search, inspect a promising document, follow a reference, compare multiple sources, keep track of what it has seen, and try a new route when the first results are insufficient. One-shot RAG provides no way to take these next steps.

    Without Agentic Search (one-shot retrieval), example shows a single search tool call finds partial data but misses needed months.

    With Agentic Search, multiple tool calls (search, read) retrieve complete data from documents, enabling accurate answers.

    How Agentic Search works

    Mistral Search Toolkit provides open modules for ingesting, embedding, and indexing critical and complex data in the cloud or on-premises. Agentic Search builds on this index by giving the model five tools that resemble familiar file-system operations:

    • search finds relevant documents across the corpus using the existing index.
    • open opens a specific document.
    • navigate moves to a page, section, or region within it.
    • read retrieves the content at that location.
    • grep finds a pattern within an open document.

    Rather than answering only from the initial top-k results, the model can inspect what it finds, refine its search, open relevant documents, navigate to specific sections, and read the source material before answering. The index identifies likely sources; Agentic Search determines what to inspect within and across them.

    These tools do not require fine-tuning or model-specific training. As models get better at reasoning and tool use, retrievals get better without infrastructure changes. This is a key property: retrieval quality scales with model capability instead of being capped by your chunking strategy.

    Use Agentic Search for

    • Long documents. Filings, contracts, manuals, technical specifications, and reports where the answer may appear on a particular page or in a specific table, clause, figure, or footnote.
    • Questions across multiple sources. Research that requires the model to find, compare, or reconcile evidence from several documents before reaching an answer.
    • Answers that must be verified. Financial figures, legal clauses, regulatory references, and operational data, where the response can be referenced in a stable and specific document location.
    • Tables and structured documents. Financial statements, government records, and scanned PDFs where meaning depends on rows, columns, page position, or surrounding context–not narrative text alone.

    Indexed retrieval is the right starting point for

    • Direct lookups. Short, clean documents where the answer is likely to appear in one of the first retrieved chunks.
    • High-volume search. Keyword or semantic lookups that need to return relevant passages without reasoning over or navigating through them.
    • Simple, predictable questions. Use cases where the likely source and location of the answer are known in advance and additional retrieval steps are unlikely to improve the result.

    One-shot RAG is often sufficient for these searches. Add Agentic Search when questions require the model to move beyond the initial results and investigate the source material. A well-configured index remains the right foundation in both cases.

    More relevant results, faster

    We benchmarked Agentic Search on two industry-standard evaluations, using the out-of-the-box Mistral Search Toolkit stack: default chunking, default ranking, no tuning. These results are floors, not ceilings, meaning you can further improve result quality with use-case-specific tuning.

    With these benchmarks, we tested two models using the Mistral Search Toolkit: Mistral Medium 3.5 (MM 3.5) and Z.ai GLM-5.2 (GLM-5.2), showcasing performance of a smaller model (MM 3.5) and a larger model (GLM-5.2).

    Benchmark results are consistent: the agentic loop delivers substantive quality improvements and navigation tools increase accuracy while reducing wasted tokens, turns, and latency. We observe the same performance patterns across first- and third-party models, which indicates that Agentic Search is model-agnostic, and that search quality should improve with new models.

    FinanceBench: 368 SEC filings, 150 questions

    FinanceBench (Islam et al., 2023) tests financial question-answering over 368 SEC filings (10-K / 10-Q / 8-K), averaging ~147 pages each, ~53,900 pages total: long, table-heavy financial documents. Answers scored by an LLM judge calibrated against human labels.

    Findings:

    • The search-only Agentic loop is the biggest quality lever. Moving from one-shot RAG to a search-only loop lifts accuracy by +47.3pp for MM 3.5 and +52.6pp for GLM-5.2–a ~3x improvement for both models. Because models can search iteratively, they can recover from weak first results, refine queries, and use the index as an active tool.
    • Navigation adds accuracy. Adding open, navigate, read, and grep lifts accuracy again (+8.7pp for MM 3.5, +6.7pp for GLM-5.2). This means a targeted drill-in search beats repeated broad search in complex documents.
    • Token and performance efficiency improve with better retrieval tools. The full loop with Navigation answers more questions correctly while using fewer tokens than the search-only loop (MM 3.5: -23.9% token usage, GLM-5.2: -33.7%). The retrieval tools are not additional overhead–they replace wasted search retries with precise navigation.
    • Latency goes down where it matters. Across FinanceBench, adding navigation retrieval tools improves latency: p90 drops 255s → 154s and mean latency drops 108s → 71s. In general, we see the search-only loops conduct repeated broad searches, while navigation helps the model identify evidence more quickly.

    OfficeQA Pro: 696 Treasury Bulletins, 133 questions

    OfficeQA Pro is a verifiable numeric benchmark over historical U.S. Treasury Bulletins: scanned, table-heavy government-finance PDFs across a 696-document, ~89,000-page corpus. We report the first pass for the 133-question "pro" subset.

    Findings:

    • Agentic Search and the Agentic loop + Navigation are successful against a harder, verifiable benchmark. OfficeQA Pro has numeric answers, scanned PDFs, and deep table lookups. Even here, the full agentic loop lifts accuracy materially from one-shot RAG, reaching 51.9% for GLM-5.2 (+45.6pp) and increasing +27.1pp for MM 3.5.
    • Navigation improves quality while cutting waste. Using the full loop (Agentic loop + Navigation) improves accuracy by up to 35.6% (+7.5pp, MM 3.5; +8.3pp, 19.0% GLM-5.2), while reducing token consumption. Turns declined by up to 7.0% (MM 3.5, 2.3% GLM-5.2).
    • The harder the benchmark, the more important the retrieval loop becomes. OfficeQA Pro is built around numeric answers in scanned, table-heavy documents. One-shot RAG barely gets started, while the agentic loop allows the model to search iteratively, inspect evidence, and deliver substantial accuracy improvements.
    • The tooling stack drives substantial impact on document intelligence and search performance. Per Kimi research, GLM-5.2 scores 41.4% on OfficeQA Pro with the Claude Code harness, compared with 51.9% on the Mistral harness–+10.5pp on the same underlying model.

    Getting started

    Learn more about Agentic Search in the documentation. You can get started across cloud and on-premises deployments using either:

    • Mistral Search Toolkit. Integrate Agentic Search into your own agents, workflows, and customer deployments.
    • Libraries. Use Agentic Search out-of-the-box in Studio and Vibe, without building the retrieval system yourself.

    The fastest way to test Search Toolkit is with the Search Starter App. It creates a local index for your own corpus using a default configuration, so you can try Agentic Search without needing to be a search expert. When you’re ready to configure your use case, you can:

    • Set up ingestion. Select parsers, chunking strategies, embedding models, and extractors for your data and file types.
    • Tune indexing and ranking. Manage Vespa schemas, indexing behavior, and relevance profiles.
    • Extend retrieval. Add query rewriting, reranking, or hybrid retrieval to the search pipeline.
    Original source
  • Similar to Mistral with recent updates:

  • Aug 11, 2026
    • Date parsed from source:
      Aug 11, 2026
    • First seen by Releasebot:
      Aug 13, 2026
    Mistral logo

    Mistral

    In-region inference, open models, and new European infrastructure for sovereign AI.

    Mistral expands sovereign AI infrastructure with regional inference endpoints, a new Priority Tier for mission-critical workloads, and support for third-party open models starting with Z.ai’s GLM-5.2, while also launching a coalition to secure long-term European compute capacity.

    Regional control, production-grade reliability

    Mistral is advancing AI sovereignty by offering enterprises and countries control over AI models, infrastructure, and compute capacity, ensuring regional compliance and reliability. The company is expanding open model access, introducing regional endpoints and priority tiers, and forming a coalition to secure long-term European AI compute capacity. With plans to build up to 1 GW of capacity by 2030, Mistral aims to provide a scalable, sovereign AI infrastructure that retains value and control for users.

    At Mistral, we believe every enterprise and country must be in control of the models it uses, choose where the intelligence runs, control the compute capacity to scale it, and retain its compounding value.

    Today, we are taking three concrete steps as we build the foundations of our customers' AI sovereignty: strengthening the reliability and regional control of inference, expanding access to third-party open models within that infrastructure, and bringing together enterprises and institutions to secure long-term commitments for compute capacity in Europe.

    Most of our customers run our models inside their own data centers and cloud environments today. As AI becomes more deeply embedded in production, they need confidence that the underlying capacity will remain available, resilient, and under the regional controls their workloads require. For some, that means complementing infrastructure they manage themselves with capacity provided and operated by Mistral.

    At the inference layer, that means two things: keeping data and processing in-region, and having dependable access to capacity when demand peaks. We are strengthening both.

    Mistral Regional Endpoints, now generally available, let customers choose whether their inference runs in Europe or the US, helping them align inference location with their data-residency, regulatory, and latency requirements. Inference and the associated processing take place in the selected region, subject to limited, safeguarded transfers to sub-processors that may occur outside that region, as described in our Trust Center. Our new Mistral Priority Tier, now in public preview, provides committed service levels for mission-critical workloads, including custom rate limits, and is backed by an uptime SLA. Mistral is the only European AI lab to offer both: choice of processing region and a committed, SLA-backed service level.

    Of course Mistral models will continue to be available through partners, as well - you can learn more about our close partnerships here.

    Sovereign intelligence, built on open model choice

    Control over where AI runs, however, is only one part of AI sovereignty. Customers also need control over the intelligence they choose to run on that infrastructure. Today’s AI systems are no longer built on a single model but on an ensemble of capabilities: extended reasoning, high-volume production, or work shaped around a company's own data. Mistral builds for each, from frontier models to specialist ones like Mistral OCR and Voxtral to custom models trained on a company's own knowledge.

    Our customers particularly value us for pioneering open models. Open weights give them what mission-critical work demands: the ability to see inside a model, adapt it, and retain the intelligence they build with it. This is why we are enthusiastic contributors to the Open Secure AI Alliance and NVIDIA Nemotron Coalition. We are now extending that openness beyond our own models. Mistral’s platform will support third-party open models, starting with Z.ai’s GLM-5.2. This and future open models will run on the same infrastructure, regional controls, and service commitments as Mistral models, so customers can broaden model choice without fragmenting where their AI runs.

    "Different workloads need different models, and that will keep changing as the frontier evolves. Mistral allows us to run open models under strict regional controls and service commitments, making it easy for us to maintain data residency and compliance requirements while furthering our commitment to open source"

    Matan Grinberg, CEO and cofounder of Factory

    A coalition that secures long-term AI capacity

    Open intelligence is inseparable from the compute beneath it. And without assured access to that compute, Europe cannot control the AI systems on which its industries, institutions, and future competitiveness will depend. Capacity is increasingly strategic, yet remains scarce, fragmented, and difficult to secure.

    Mistral is bringing together an anchor group of enterprises whose multi-year commitments can support infrastructure in Europe at a scale no participant could secure alone. Aggregating long-term demand in this way will help determine what capacity is built, where it is located, and whom it serves. European Compute Units, or ECUs, convert those commitments into access to Mistral-built infrastructure over multiple years. Participants can use that capacity across the range of products available on Mistral Compute as their needs evolve.

    "As the neutral execution layer and trusted system of record at the core of the global travel industry, Amadeus bridges raw foundational data with modern Agentic AI to improve the traveller experience for everyone everywhere. In an AI-driven world, capacity, deployment control, and operating continuity become increasingly important for all enterprises. Equally crucial, at the same time, is ensuring businesses have the confidence to make long-horizon commitments at scale; something that Mistral’s massive compute undertaking supports."

    Luis Maroto, CEO of Amadeus

    "Few industrial endeavors will matter more to Europe’s next generation than building the capacity to develop and run AI on its own terms. Mistral is taking on that challenge with the scale, ambition, and staying power it demands, giving enterprises the confidence to build their most consequential workloads on that foundation."

    Christophe Fouquet, CEO of ASML

    "Building AI capacity isn't just a technology question - it's a question of who shapes the future of European industry. This program is about giving our clients the compute, the partnerships, and the confidence to run their most critical AI workloads on infrastructure built for scale and built to last."

    Aiman Ezzat, CEO, Capgemini

    "Europe needs sovereign infrastructure to ensure its technological independence. Mistral embodies this ambition by providing a robust, scalable platform aligned with our values, enabling businesses to deploy critical solutions with confidence. Innovating without relying on foreign actors is an imperative for Europe. With Mistral Compute, we now have a European neocloud capable of competing on a global scale while retaining control over our data and models. The IRN (Digital Resilience Index) will further help businesses measure their digital dependence. This is a major breakthrough for our ecosystem."

    Olivier Sichel, CEO of Caisse des Dépôts

    "AI is the industrial revolution of our time, including for non-tech companies like CMA CGM, a global Group active in shipping, logistics and media. We chose Mistral for its world-class technology, its ability to co-develop robust and resilient solutions tailored to our operational needs and to efficiently deploy them at scale. That deployment is already under way among thousands of employees and across geographies. It is transforming areas such as customer care to improve quality, responsiveness and reliability for our customers."

    Rodolphe Saadé, Chairman and CEO of CMA CGM

    A framework for advancing AI sovereignty

    Europe is the place where this framework begins, but the same challenge exists everywhere: organizations and governments need to harness the power of frontier AI for their mission-critical needs without surrendering control over the infrastructure and intelligence loop.

    The path will differ by region, but the foundations are the same: operational control, choice of intelligence, and assured access to compute. Europe can show how these elements come together to create an AI ecosystem that remains open to global innovation while preserving institutional autonomy.

    That is the model Mistral is building toward: one in which enterprises, public institutions, and startups can use the best AI available, shape it around their own knowledge, and retain the value it creates.

    Original source
  • Aug 4, 2026
    • Date parsed from source:
      Aug 4, 2026
    • First seen by Releasebot:
      Aug 5, 2026
    Mistral logo

    Mistral

    Introducing Shieldstral.

    Mistral releases Shieldstral, a 3B open-weights multimodal safety classifier that turns moderation into policy-adaptive question answering. It handles text and images, returns calibrated safety scores, runs on a single 16GB GPU, and ships under Apache 2.0.

    Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task.

    Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU.

    A 3B open-weights, policy-adaptive multimodal safety classifier that matches models up to 7x its size on text safety and sets a new state of the art on multimodal moderation.

    “Does this content promote violence against a protected group? Is this image safe to show to a minor? Did the assistant refuse the request?”

    Every product that ships a model needs to answer questions like these — but the right answer depends on the product, the audience, and the moment. The same content can be fine for a cybersecurity research tool and harmful on a mental-health platform. Most guardrail models bake a fixed taxonomy of harm categories into their weights, so re-targeting them to a new deployment context means retraining. And because safety definitions differ across applications and domains, there is no single "correct" set of categories to model in the first place.

    Shieldstral takes a different approach: you write the policy as a plain-language question at inference time, and the model returns a calibrated safety score. No retraining, one interface for text and images, and a verdict from a single token. Please refer to our technical report here.

    As an inaugural member of the Open Secure AI Alliance with NVIDIA and other organizations, today we're releasing Shieldstral as open weights under Apache 2.0, available for download here.

    Moderation as a question

    Shieldstral frames content moderation as a binary question-answering task. Each request has three parts:

    • — the evaluation context, strictness, and (optionally) a definition of what counts as unsafe content.
    • — a single yes/no question, e.g. "Does this content promote physical violence?"
    • — the content to judge: a prompt, a response, a prompt–response pair, or an image with optional text.

    At inference the model reads out only the yes and no logits and softmax-normalizes them into a continuous safety score. This one simple formulation does a lot of work: it unifies prompt classification, response moderation, refusal detection, and toxicity detection into a single problem; it lets policies live entirely in the prompt, so one checkpoint adapts to novel policies at deployment time.

    Highlights

    • Strong performance — matches or outperforms open guard models up to 7× its size across text safety, refusal detection, policy adaptability, and multimodal benchmarks.
    • Adaptive and flexible — a single natural-language interface covers text, image, and text+image content across prompts, responses, and prompt–response pairs. Policies are supplied as free-form queries and re-targeted at inference time, without retraining.
    • Small, trained on heterogeneous sources — a 3B model that runs on a single 16GB GPU, trained on real and synthetic data with diverse label formats and taxonomies, consolidated into one framework.
    • Continuous safety score — returns a calibrated yes/no probability from a single forward pass, so you can threshold or rank by confidence rather than relying on a discrete label.
    • Open — Apache 2.0 weights.

    Benchmarks

    We evaluate Shieldstral against open guard models up to 7x its size across four axes. All evaluation samples are held out from training.

    Text safety

    Refusal detection

    Policy adaptability

    Multimodal safety

    How we built it

    The core idea is that a small model can beat much larger ones if the data is right. Getting the data right meant solving four problems:

    Unify heterogeneous data.

    Public safety datasets disagree on taxonomies, labels, and annotation conventions — from binary safe/unsafe flags to fine-grained multi-label taxonomies. We convert every dataset into the same instruction–query–document format with a per-dataset processor, and we vary the wording of instructions, queries, and prompt–response delimiters so the model generalizes across phrasing instead of overfitting to one style. We also calibrate strictness per source — strict for adversarial jailbreaks, lenient for response-quality data — so the model learns calibrated decision boundaries. This lets us consolidate sources that would otherwise be incompatible.

    Teach discrimination, not memorization.

    If trained on a fixed set of policy labels, a model learns only to classify those predefined policies, rather than reasoning about the precise boundaries of a given policy. This prevents generalization to novel policies. Instead, we construct sets of deliberately similar, easily confused policies and ask an LLM to rewrite safe text into contrastive pairs: each rewrite is engineered to violate one policy but not its sibling. This trains the model to distinguish which specific policy a piece of content violates, a skill that transfers to unseen, user-defined policies at inference time.

    Ground safety in images.

    Unsafe images can't be synthezised by an LLM the way text can, so visual safety data is scarce. We supplement limited moderation datasets with general-purpose image datasets as high-quality negatives, mutate queries to augment the dataset, and filter every image–query pair through a vision–language reranker to reduce mislabeled data and hallucinations.

    Combine complementary checkpoints.

    We fine-tune with LoRA and merge — via SLERP — a checkpoint calibrated on public data, one that adds fine-grained policy discrimination from generated data, and the base instruct model. The merge recovers common policy calibration and policy adaptability in a single model, and instruction-following from the base model transfers to the moderation task.

    Forge.

    We built Shieldstral end to end on Forge, our platform for training, aligning, and evaluating custom models. Forge managed the infrastructure, data and model sharding, metrics, and logging on top of state-of-the-art distributed training, so the team could stay focused on the data which is what determines the safety model's quality.

    What's next

    Shieldstral is a step toward moderation that adapts to context instead of forcing every product through one frozen taxonomy. We're continuing to push on multilingual coverage, longer-document robustness, and broader multimodal safety — and we'd love to see what the community builds on top of it.

    BTW, we're hiring! If you want to help make AI better, see our careers page.

    Original source
  • Jul 23, 2026
    • Date parsed from source:
      Jul 23, 2026
    • First seen by Releasebot:
      Jul 24, 2026
    Mistral logo

    Mistral Common by Mistral

    v1.11.7: Patch release

    Mistral Common adds tokenizer chat template auto-detect, stricter decoding validation, and tool-message fixes.

    What's Changed

    • Update AGENTS.md structure and add explicit keyword args rule by @juliendenize in #270
    • Validate special_token_policy strings when decoding by @sarathfrancis90 in #273
    • Add auto-detecting chat template generation from tokenizer file by @juliendenize in #272
    • Reject thinking chunks in tool messages by @juliendenize in #275
    • Emit all render_content macro args explicitly by @juliendenize in #274
    • Fix documentation drift from code by @juliendenize in #276
    • Bump version to 1.11.7 by @juliendenize in #277

    Full Changelog: v1.11.6...v1.11.7

    Original source
  • Jul 17, 2026
    • Date parsed from source:
      Jul 17, 2026
    • First seen by Releasebot:
      Jul 18, 2026
    Mistral logo

    Mistral Common by Mistral

    v1.11.6: Fixes and patch grammar selector

    Mistral Common releases a bugfix update that tightens validators, prevents audio config errors, improves SentencePiece token decoding, auto-selects grammar variants from the tokenizer, and restores the deprecated reasoning parameter for Jinja templates.

    What's Changed

    • Raise ValueError instead of ValidationError in audio chunk validators by @sarathfrancis90 in #255
    • Reject AudioConfig with sampling_rate < frame_rate instead of a later ZeroDivisionError by @CharlesCNorton in #253
    • Decode normal tokens under SpecialTokenPolicy.RAISE in SentencePiece by @sarathfrancis90 in #256
    • Auto-select grammar variant from tokenizer by @juliendenize in #260
    • Fix seed field not bridged in SpeechRequest OpenAI conversion by @winklemad in #262
    • Restore deprecated reasoning parameter in select_jinja_template by @juliendenize in #266
    • Bump version to 1.11.6 by @juliendenize in #267

    New Contributors

    • @winklemad made their first contribution in #262

    Full Changelog: v1.11.5...v1.11.6

    Original source
  • Jul 15, 2026
    • Date parsed from source:
      Jul 15, 2026
    • First seen by Releasebot:
      Aug 13, 2026
    Mistral logo

    Mistral

    July 15

    Mistral releases OCR 4.1 with updated aliases and finer confidence score granularity options in the OCR API.

    We released OCR 4.1 (mistral-ocr-4-1). mistral-ocr-latest and mistral-ocr-4 now point to it.

    MODEL RELEASED

    The OCR API confidence_scores_granularity parameter now supports "block" granularity. "page" returns page-level scores only, "block" returns page-level and block-level scores, and "word" returns page-level and word-level scores.

    API UPDATED

    Original source
  • Jul 9, 2026
    • Date parsed from source:
      Jul 9, 2026
    • First seen by Releasebot:
      Aug 5, 2026
    Mistral logo

    Mistral

    Your Prompts and Skills need a system of record.

    Mistral adds Prompts and Skills in Studio, turning scattered AI instructions into governed production assets with versioning, ownership, audit logs, rollback, and observability. It helps teams iterate faster, ship with control, and keep behavior traceable and compliant.

    Most enterprises struggle with unmanaged, scattered AI prompts and skills, leading to inconsistent behavior and untraceable issues.

    Studio provides a centralized system of record for versioning, ownership, and traceability, enabling fast iteration and controlled deployment while maintaining compliance. By treating prompts as production assets with immutable versions, clear ownership, and audit logs, Studio ensures AI behavior is governed, discoverable, and aligned with business policies.

    Most enterprises can't say which version of a prompt is running in their AI right now. The instructions that decide how that AI behaves get scattered the moment more than one team touches them, leading to an inconsistent experience for users and an untraceable problem for teams.

    As of today, Studio gives your Prompts and Skills a system of record: a single place where each one is versioned, owned, and traceable.

    Prompts and skills outgrew the way they're managed

    Prompts and Skills are production assets. They hold the business logic, the tone, and the policy your AI follows when it answers a customer or makes a call. What your AI does in front of a customer comes down to the prompts and skills in use. When that behaviour is wrong, the fix has to ship as fast as any production incident, not wait for the next code release.

    And in most enterprises, they're managed like scratch notes. Prompts started as quick experiments, then they shipped. Now they sit in code repos, notebooks, and Slack threads, with no clear owner and no shared history. Skills get rebuilt, or forked by one team because they lacked visibility to another team’s version.

    In many enterprises, prompts already live in version-controlled code, which means tracking changes was never the hard part. The friction is elsewhere. The people who understand the instructions best, the line-of-business teams who set the policy and the wording, don't work in the codebase, so every change waits on an engineer. And refining an instruction takes iteration and testing, which a codebase makes expensive: one version ships at a time, and every attempt means editing code and waiting for a deploy.

    So most teams stop iterating early. They ship a version that's good enough and then may leave it, and the instructions that shape every customer answer can stay well short of what they could be.

    Iterate fast, ship with control

    While you're building, iterating on an instruction should be quick. In code, even a one-line change to a prompt can mean waiting on a CI run before you see how it behaves. Studio lets any AI builder, developer or not, edit a prompt or skill and test it right away, without a pipeline run for every attempt.

    Shipping to production is different, and it should be. A change bound for production goes through the tests and approvals your enterprise already requires. What changes is who can drive it. A domain expert or line-of-business owner can improve a production instruction the same way a developer would, and the promotion using simple labels still triggers your CI/CD, for example through the SDK in a GitHub Actions workflow. The people closest to the work improve the behaviour, inside the controls you already run.

    Because every asset is governed and discoverable, good work spreads instead of getting rebuilt. Anything in a workspace is available to that whole team today, so a prompt one person gets right is usable by their colleagues at once.

    A system of record for AI behavior.

    Studio treats every prompt and skill as a tracked, versioned asset with an owner, a full history, and a lineage.

    • Immutable versions. Every version is recorded and fixed. A version that shipped can't be quietly changed after the fact, so the record always matches what ran.
    • Rollback. Compare any two versions, see exactly what changed, and revert to a known-good version in minutes.
    • Clear ownership. Every asset has a named owner, so there's always an audit trail to track changes.
    • Classification labels. Helps call or find the right prompts and skills by their labels easily (e.g. “Production” vs “Staging”).
    • Audit logs. Each change is logged with who made it and when. The trail an auditor will ask for exists by default.

    The part a standalone catalog can't do.

    A separate prompt tool can list your assets. It can't tell you whether they work, because it sits outside the system that runs them.

    Because your prompts and skills live where your AI runs, Studio can connect them to how it actually behaves. Through Observability, lineage and telemetry trace a production output back to the version of the asset behind it, and back to the usage that prompted the last change. The skills your agents run are reachable as MCP servers straight from Studio, so what executes in production is the same governed asset you versioned, not a copy that drifted. You define behaviour, watch it run, and improve it, all against one source of truth. That closed loop is the difference between cataloging your AI and governing it.

    Control for the people who answer to auditors.

    Ungoverned prompts are a liability for the people who answer to auditors. They embed data-handling rules and policy decisions someone will eventually have to defend, and today they often live where no compliance team can see them.

    Studio changes the default. Every asset moves through a clear path to production, from a staging version to a tagged production version, so shipping a change is deliberate rather than accidental.

    An asset starts as visible only to its creator, then when appropriate can be promoted to the workspace and, in time, across the organisation, with control over who can use it at each step. Across every deployment mode, your data stays inside your perimeter.

    Available now in Studio.

    Prompts and skills are available to Mistral Studio customers today. If you run AI in production, Studio turns scattered prompts and skills into governed assets you can trust.

    Read documentation:

    • Create reusable prompts in Studio
    • Create reusable skills in Studio

    Or explore Prompts and Skills in Studio.

    Original source
  • July 2026
    • No date parsed from source.
    • First seen by Releasebot:
      Jul 9, 2026
    Mistral logo

    Mistral

    Robostral Navigate

    Mistral introduces Robostral Navigate, its first embodied navigation model, built to steer robots with a single RGB camera and plain-language instructions. The 8B model is trained in simulation, supports multiple robot types, and delivers strong benchmark results with more efficient training.

    Robostral Navigate is an 8B model that enables robots to autonomously navigate complex environments using only a single RGB camera, achieving 76.6% success on unseen R2R-CE benchmarks—outperforming multi-sensor approaches while being more efficient. Built entirely in-house with simulated data and token-efficient techniques, it generalizes across robot types and adapts to real-world obstacles unseen during training. The model combines pointing-based navigation with reinforcement learning for continuous improvement, paving the way for unified embodied AI in robotics.

    Today we're introducing Robostral Navigate, our first model built for embodied navigation. It's an 8B model that takes RGB images and a plain-language instruction and moves a robot through an environment:

    “Leave the lobby, walk through the corridor, enter the supply room, and stop to face the second shelf.”

    To perform such tasks, other models often employ depth sensors, LiDAR, or several cameras working together. Robostral Navigate uses only one ordinary RGB camera and no depth sensors, yet still achieves 76.6% on R2R-CE (Room-to-Room in Continuous Environments) validation unseen, the benchmark for following instructions in environments held out of training. Consequently, it beats the best single-camera approach by 9.7 points and the best system using depth or multiple cameras by 4.5 points, despite using neither.

    Navigation

    Our model is designed for robotic navigation, enabling robots to autonomously navigate complex environments, including offices, residential and commercial buildings, and outdoor settings.

    Robostral Navigate running fully autonomously in one long-horizon instruction route through a working office.

    This technology unlocks numerous applications across manufacturing, delivery, logistics, and hospitality, making it one of the most in-demand capabilities for our customers today. Give Robostral Navigate one instruction and it completes the entire task on its own, moving through a live space full of people and obstacles it was never shown, capable of adapting to any setting.

    Highlights

    • State-of-the-art performance on R2R-CE
      • 79.4% Success Rate on validation seen
      • 76.6% Success Rate on validation unseen
    • Operates from a single RGB camera, with no LiDAR or depth sensors
    • 8B model, built in-house and trained entirely in simulation
    • Runs on wheeled, legged, and flying robots, and generalizes across robot sizes
    • Robust to differences in camera intrinsics
    • Token-efficient training via prefix-caching

    Navigation via pointing

    Given a task and a history of observations, Robostral Navigate predicts where the robot should move next via pointing: it infers the image coordinates of the target location in the robot's current camera view, together with the desired orientation upon arrival. Unlike commands relying on metric displacements, pointing makes the policy naturally robust to changes in camera intrinsics and world scale.

    However, this method cannot handle cases where the target location lies outside the current field of view. When pointing does not apply, the model falls back to displacements in the robot's local coordinate frame, such as:

    "Move 2 meters forward, 1.5 meters to the left, and turn 25 degrees left."

    Built from the ground up

    Robostral Navigate is built entirely in-house and does not rely on existing open-source VLMs.

    The model is initialized from our vision-language model specialized for grounding tasks such as pointing, counting, and object localization. Navigation emerges as a natural extension of these capabilities: once it understands where things are, it learns how to move.

    We built an efficient data generation pipeline entirely in simulation. This enabled rapid iteration on the data, resulting in a dataset of approximately 400,000 trajectories collected across 6,000 scenes.

    Efficient supervised training

    A key ingredient of Robostral Navigate is an efficient training algorithm based on prefix-caching. Using a tree-based attention-masking strategy, our method compresses an entire episode into a single sequence, enabling training on all time steps in a single forward pass while preventing information leakage between time steps.

    Compared to training with one sample per time step, our approach reduces the number of training tokens by 22× while preserving all of the learning signals. In practice, this method transforms training runs that would take months into runs that complete in days.

    Online reinforcement learning

    We leverage our knowledge of post-training LLMs at scale, using online reinforcement learning, to boost the performance of Robostral Navigate. After the supervised training stage, we further improve the model's performance using CISPO, an online reinforcement learning algorithm. This enables the model to learn from trial and error, recover from failures, and acquire exploratory behaviors, effectively mitigating the distribution shift issue of vanilla behavior cloning. This alone improved the success rate by 3.2%. We are not seeing any plateauing, so we are confident that more training and more experiments will continue to push this number up.

    What's Next

    Robostral Navigate is only the first step toward a unified embodied agent.

    We believe navigation is a foundational capability for general-purpose robotics. By combining large-scale simulation, efficient training, and strong grounding priors, Robostral Navigate demonstrates that state-of-the-art embodied navigation can be achieved with a compact model and a single RGB camera.

    Start your journey to embodied frontier AI, talk with our team.

    BTW, we're hiring!

    The release of our navigation models marks a significant step forward, but our journey is far from over. Our ambition is to enable robots to autonomously navigate complex environments—offices, homes, commercial buildings, and outdoor spaces—and there's a lot more work to do. We are actively expanding our robotics team and looking for talented research scientists and engineers who share our ambition.

    If you're interested in joining us on our mission to bring seamless navigation to robots everywhere, we welcome your applications to join our team!

    By Théo Cachet, Arjun Majumdar, Srijan Mishra, Thomas Chabal, Chris Bamford, Elliot Chane-Sane, Benjamin Tibi, Ludovic Ho Fuh, Olivier Duchenne - AI Science Robotics

    Original source
  • Jul 2, 2026
    • Date parsed from source:
      Jul 2, 2026
    • First seen by Releasebot:
      Jul 4, 2026
    Mistral logo

    Mistral

    Leanstral 1.5: Proof Abundance for All

    Mistral releases Leanstral 1.5, a free Apache-2.0 open model for Lean 4 proof engineering that delivers major gains in formal verification, tops key benchmarks, strengthens code verification, and is now available on Hugging Face and via a free API.

    Leanstral 1.5

    Leanstral 1.5, a free Apache-2.0 licensed model with 6B active parameters, delivers a major performance upgrade in formal verification, saturating miniF2F, solving 587/672 PutnamBench problems, and achieving state-of-the-art results on FATE-H (87%) and FATE-X (34%). Trained through mid-training, supervised fine-tuning, and reinforcement learning with CISPO, it excels in agentic proof engineering and real-world code verification, uncovering 5 previously unknown bugs across 57 repositories tested. Fully open-sourced and available via Hugging Face and a free API, Leanstral 1.5 is now accessible for practical proof engineering in Lean 4.

    Since its launch, Leanstral has offered an open, practical approach to proof engineering in Lean 4. Today, we are releasing Leanstral 1.5, a free Apache-2.0 licensed model with 119B total and only 6B active parameters, delivering a performance upgrade that makes formal verification more powerful and accessible than ever.

    Leanstral 1.5 saturates miniF2F, solves 587/672 PutnamBench problems, and achieves a new state-of-the-art of 87% on FATE-H and 34% on FATE-X. Beyond benchmarks, it verifies complex code properties and uncovers previously unknown bugs in open-source repositories—proving that rigorous formal methods can be both effective and practical for real-world use.

    Training Leanstral

    Leanstral 1.5 goes through a three-stage process: mid-training, supervised fine-tuning, and reinforcement learning with CISPO. Leanstral 1.5 leverages extensive training on two RL environments:

    In the multiturn environment, the model is given a theorem statement and must either prove or disprove it. The model submits a proof, receives Lean compiler feedback, and refines its approach with each attempt. If the proof compiles it succeeds; otherwise the loop continues until the model either solves the problem or exhausts its budget.

    In the code agent environment, Leanstral operates like a developer in a raw filesystem: it edits files, runs bash commands, and uses the Lean language server to inspect goals, errors, and type information in real time. This allows it to tackle long-horizon tasks like completing partial proofs in a repository, building auxiliary lemmas, and persisting through multiple rounds of context compaction. The model learns to navigate the full proof-engineering workflow and is finally verified by our fork of SafeVerify for correctness given a list of target theorems.

    Evaluation

    We evaluate Leanstral on the following benchmarks:

    • miniF2F is a cross-system benchmark for formal mathematics, ranging from elementary problems to IMO-level challenges, testing diverse proof abilities across algebra, combinatorics, and number theory.
    • PutnamBench consists of 672 problems from the Putnam Mathematical Competition, requiring deep reasoning and long proof chains to solve challenging mathematical problems.
    • FATE-H and FATE-X are abstract algebra benchmarks for graduate and PhD-level problems, respectively, testing advanced reasoning in areas like group theory, ring theory, and module theory.
    • FLTEval is based on real pull requests from the Fermat’s Last Theorem repository, testing practical proof engineering with real-world complexity.

    We saturate miniF2F completely, reaching 100% on both the validation and test sets. On PutnamBench and FATE-H/X, we compare Leanstral 1.5 against Goedel-Architect without natural-language guidance, Seed-Prover 1.5 at its high setting, and AxProverBase. Leanstral reaches a new state-of-the-art on FATE-H/X, solving 87 and 34 problems respectively. On PutnamBench, it edges out Seed-Prover 1.5 high by 7 problems at far lower cost: about $4 per problem, against an estimated $300 or more for Seed-Prover, whose high setting runs with a budget of 10 H20-days per problem. The only provers ranked higher operate under different conditions—some receive natural-language proof guidance, others cost far more to run, like Aleph Prover at $54–68 per problem.

    Leanstral 1.5 shows the strongest test-time scaling we have seen from a formal-reasoning model. The figure below tracks Pass@8 on PutnamBench as we raise the token budget per attempt from 25k to 4M: performance climbs smoothly and monotonically the whole way, from 44 problems solved at 50k to 244 at 200k, 493 at 1M, and 587 at 4M. Rather than giving up when a proof runs long, Leanstral keeps reasoning, editing files, and revising across millions of tokens, turning that budget directly into solved problems—the same behavior behind the AVL-tree proof below, which ran for over 2.7 million tokens across 22 compactions.

    With this release, we also fully open source FLTEval. Leanstral 1.5 lifts pass@1 on the benchmark from 21.9 to 28.9 and pass@8 from 31.9 to 43.2, surpassing Opus 4.6's 39.6 at one-seventh the cost. It also widens its lead over open-source models 3–10× larger, as shown in the figure below.

    Code Verification Case Studies

    While being primarily trained for mathematics, Leanstral 1.5 exhibits strong abilities in code verification. We present 2 critical case studies to demonstrate its impact.

    AVL Trees: Proving Time Complexity

    AVL trees are self-balancing binary search trees that maintain O(log n) height through rebalancing during insertions and deletions. Leanstral 1.5 proved these time complexity guarantees for a real implementation—a task that required structural induction to mirror the tree’s recursive structure, careful handling of monadic time tracking, and exhaustive case analysis for rebalancing paths. Over 2.7 million tokens and 22 compactions, Leanstral systematically unfolded each layer of the TimeM monad, exposing the underlying computations despite their interleaving with control flow. It established an almost tight bound of 48 steps per height unit plus a constant for insertion, then connected height to tree size via a logarithmic relationship, delivering complete, verified proofs that insertion and deletion are indeed O(log n).

    Bug Discovery: Finding Hidden Flaws

    To test Leanstral’s bug-catching abilities, we built an automated pipeline: Aeneas translates Rust code to Lean, while Leanstral infers the user intent and generates correctness properties from the code. Leanstral then attempts to prove each property in four attempts. If they all fail, it tries to prove the negation instead, also with four attempts. Across 57 tested repositories, this process flagged 47 violated properties, with 11 pointing to genuine bugs—5 of them previously unreported on GitHub.

    One such bug was in the sign function for zigzag decoding of the datrs/varinteger library. On input Std.U64.MAX, the expression (value + 1) overflowed, causing crashes in debug mode and silent corruption in release mode—an edge case that testing and fuzzing would typically miss. Leanstral’s pipeline caught it automatically, demonstrating that formal verification can already be applied to real-world codebases and find bugs that some traditional methods overlook.

    Get Started

    Leanstral 1.5 has a Apache-2.0 license. The weights can be found on Huggingface, while also being available now as a free API endpoint as leanstral-1-5. We recommend using it in Mistral Vibe. To begin your journey, grab an API Key, and:

    1. Set up Mistral Vibe
    uv tool install mistral-vibe
    uv tool update mistral-vibe
    vibe --setup
    
    1. Install Leanstral 1.5
    /leanstall
    exit
    
    1. Launch the agent
    vibe --agent lean
    
    1. Install Lean LSP MCP (Optional)

    It is highly recommended to install Lean LSP MCP by adding the following to your ~/.vibe/config.toml

    [[mcp_servers]]
    name = "lean-lsp"
    transport = "stdio"
    command = "uvx"
    args = ["lean-lsp-mcp"]
    tool_timeout_sec = 600
    

    If there are no existing MCP servers, you may have to remove mcp_servers = [].

    1. Start proving

    Ask Leanstral to tackle a theorem, debug a proof, or contribute to a repository. It’s that simple.

    Original source
  • Jun 29, 2026
    • Date parsed from source:
      Jun 29, 2026
    • First seen by Releasebot:
      Jul 1, 2026
    Mistral logo

    Mistral

    June 29

    Mistral releases Leanstral 1.5 with better proof engineering, improved training mix, and longer-context reasoning.

    We released Leanstral 1.5 (labs-leanstral-1-5), an updated Lean 4 formal proof engineering model with improved SFT mixture quality and extended long-context reasoning. This model will be retired on September 30, 2026.

    MODEL RELEASED

    Original source
  • Jun 26, 2026
    • Date parsed from source:
      Jun 26, 2026
    • First seen by Releasebot:
      Jun 27, 2026
    Mistral logo

    Mistral Common by Mistral

    v1.11.5: Hotfix encoding only two consecutive images

    Mistral Common fixes multi-image content ordering in a new release.

    What's Changed

    Fix multi-image content ordering by @juliendenize in #254

    Full Changelog: v1.11.4...v1.11.5

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.