Cohere Release Notes

Follow

81 release notes curated from 22 sources by the Releasebot Team. Last updated: Jul 7, 2026

Get this feed:
  • Jul 7, 2026
    • Date parsed from source:
      Jul 7, 2026
    • First seen by Releasebot:
      Jul 7, 2026
    Cohere logo

    Cohere

    Meet Cohere Transcribe Arabic

    Cohere releases Cohere Transcribe Arabic, an open-source speech-to-text model for Arabic speakers with strong accuracy across major dialects and speech patterns. It is optimized for production inference, available via the V2 Audio Transcriptions API, open weights on Hugging Face, and Model Vault deployment.

    Today we are releasing Cohere Transcribe Arabic.

    This open-source speech-to-text model is a fine-tune of Cohere Transcribe using Arabic speech data. It lets Arabic speakers transcribe their voice with unmatched accuracy and support for regional dialects or speech patterns.

    It is currently the most accurate open-source Arabic ASR model available today and is optimized for production inference and throughput.

    Technical Details

    • Model Name: cohere-transcribe-arabic-07-2026
    • Size: 2B
    • Architecture: conformer-based encoder-decoder
    • Languages supported: Arabic (all major dialects), English (including English spoken with an Arabic accent)
    • License: Apache 2.0

    Availability

    Cohere Transcribe Arabic is available through the V2 Audio Transcriptions API and as open weights on Hugging Face. For production use, Model Vault deployment is also supported.

    For more details, see the model documentation.

    Original source
  • Jul 7, 2026
    • Date parsed from source:
      Jul 7, 2026
    • First seen by Releasebot:
      Jul 7, 2026
    Cohere logo

    Cohere

    Meet Cohere Transcribe Arabic

    Cohere releases Transcribe Arabic, an open-source Arabic speech recognition model with state-of-the-art transcription accuracy, strong dialect and code-switching handling, and enterprise-ready throughput. It’s available under Apache 2.0 via Hugging Face, the Cohere API, and Model Vault.

    The world’s most accurate open-source model for Arabic speech recognition.

    Key takeaways

    • Transcribe Arabic achieves the highest accuracy for Arabic-language transcription among open-weights models.
    • It is designed to capture the nuances and dialectical richness of Arabic, while being fit for enterprise speech applications.
    • Human reviewers preferred Cohere Transcribe Arabic to Whisper in 96% of tests.
    • This is a proud advance in the region’s sovereign AI capabilities, bringing frontier performance to millions of Arabic-speakers.

    Today, we are releasing Cohere Transcribe Arabic as an open-source model. Based on our 2B frontier Automatic Speech Recognition (ASR) model released earlier this year, Cohere Transcribe Arabic is built for the realities of Arabic in business and developer settings: dialect variation, bilingual Arabic-English speech, code-switching, and domain-specific vocabulary.

    It is the most accurate, open-source Arabic speech-to-text model to date, outperforming leading alternatives, including Whisper and OmniASR, across dialects and common speech patterns. It also delivers substantial gains over Cohere Transcribe on both Arabic and bilingual Arabic-English audio.

    Cohere Transcribe Arabic is available under the Apache 2.0 license. Developers can download the weights and read our quickstart implementations on Hugging Face, or access the hosted model through the Cohere API or Model Vault.

    One model, many voices

    Arabic is a remarkably rich language in both script and speech. More than 300 million people speak Arabic as their mother tongue, across roughly 30 recognized varieties shaped by distinct cultural, regional, and historical contexts. Saudi Arabia alone is home to three major dialect groups and many more linguistic subgroups.

    This diversity, however, heavily complicates efforts towards normalisation. While Modern Standard Arabic (MSA) provides a de facto common written standard, everyday speech varies significantly across dialects. Morphological differences, regional pronunciation, and code-switching — the use of Arabic and non-Arabic vocabulary in the same conversation, often in professional settings — make a single, uniform approach to communication difficult.

    The challenge is especially stark in the development of natural language technology, such as ASR. How do you train a model that preserves dialectal nuance while remaining useful beyond a particular market? The result has been a frontier-language gap: Arabic remains under-served by state-of-the-art AI systems while English continues to dominate model development and evaluation.

    To help narrow that gap, Cohere embarked on a simple mission: build an enterprise-ready solution that lets Arabic users speak in their natural voice.

    We started with Cohere Transcribe (launched in March with leading English-language accuracy and broad multilingual coverage) and trained it extensively on data spanning Arabic dialects, professional language, code-switching, and varied acoustic conditions.

    The result is a new state-of-the-art solution in how Arabic speech is captured, ready for production use and openly available to all.

    Unrivalled accuracy

    Cohere Transcribe Arabic achieves the lowest average word error rate (WER) of any open-source model on the Hugging Face Arabic ASR Leaderboard, with a WER of 25.87. This is a 2.45-point improvement over the previous leader, Meta’s OmniASR-LLM-7B, and an 11-point improvement over OpenAI’s Whisper Large V3.

    The gains are broad-based. Cohere Transcribe Arabic delivers the best overall WER and ranks first on four of the six composite task sets. On Casablanca, for example, which evaluates conversational Arabic across eight dialects, it improves on OmniASR by nearly six points. On Common Voice, a crowd-sourced dataset covering 25 dialects, it reduces WER by more than two points from a low previous base.

    The performance carried through to human evaluations with native Arabic speakers. Evaluators assessed transcription quality across three dimensions:

    1. Overall accuracy: How well the transcript captured both the form and semantic meaning of the audio
    2. Dialect faithfulness: Whether the model preserved the speaker’s dialect rather than converging toward Modern Standard Arabic
    3. Robustness to code-switching: Whether English terms were correctly transcribed in Latin script

    Cohere Transcribe Arabic scored highest on all three dimensions compared with Whisper and Cohere Transcribe. In head-to-head evaluations, it was preferred over Whisper in 95.8% of tests.

    The model also improved significantly on Cohere Transcribe for English spoken with an Arabic accent, covering many workplace and second-language English use cases. Human evaluators preferred Cohere Transcribe Arabic over Cohere Transcribe in 77.2% of tests, and found it broadly comparable to Whisper on these inputs, with Cohere Transcribe Arabic preferred in 52.6% of tests.

    Enterprise-ready

    Cohere Transcribe Arabic is built for high-throughput serving in production environments, where performance under concurrent demand matters as much as model quality.

    We optimized the system around vLLM to handle high-volume speech workloads, even when audio inputs vary in length. Further updates to the runtime and model stack helped yield up to 2x higher throughput.

    Overall, Cohere Transcribe Arabic achieves an RTFx (real-time factor multiple) score of 525 versus 146 for Whisper Large V3 and 66 for omniASR 7B-LLM.

    See for yourself

    Read the following examples to see how Cohere Transcribe Arabic successfully handles code-switching and dialectic nuance within an enterprise setting.

    Example 1: A common workplace dialogue about internal processes and human resources

    Where Cohere Transcribe Arabic wins:

    • Maintains the Gulf dialect: Cohere Transcribe Arabic preserves regional phrasing such as انا ابغي, ما عندي خبرة, من وين ابدا and وايش هي القوايم, while Whisper standardizes these into less speaker-faithful forms.
    • Preserves bilinguality: Cohere Transcribe Arabic captures hybrid workplace vocabulary such as “annual leave” and “HRIS” exactly as spoken. By contrast, Cohere Transcribe either mistranslates these terms, for example rendering HRIS as اتش ار اي اس, or misses them entirely.

    Example 2: Text involving more technical and domain-specific language

    Where Cohere Transcribe Arabic wins:

    • Maintains technical accuracy: Cohere Transcribe Arabic keeps enterprise terminology intact, including “AI initiatives”, “customer experience”, and “network optimization”, rather than translating or approximating these terms in Arabic.
    • Semantic completeness: Cohere Transcribe Arabic retains the speaker’s original ask, such as لناجحة use cases من ال examples اعطني, instead of changing it into a less natural or semantically altered phrase, as Whisper does with ...وقد اعطينا ايضا مثالات.

    Making speech sovereign

    Everyone deserves sovereignty when it comes to AI — no matter their mother tongue. It begins with control: over data, infrastructure, and your models.

    Cohere Transcribe Arabic runs efficiently on consumer hardware, with no reliance on external APIs or cloud services, and is available under a permissive license. We’re giving developers open access to state-of-the-art speech AI in their own language and deployable on their own terms.

    What’s next? We’re expanding enterprise-grade transcription capabilities and bringing support for new speech-powered applications for North users. This is all in collaboration with our regional partners and fast-growing local developer communities.

    If you’d like to know how Cohere can make your AI capabilities more sovereign, please get in touch.

    Start transcribing

    Model weights are available today on Hugging Face. You can test the model beforehand in our Hugging Face Space.

    Cohere Transcribe Arabic is also available through the Cohere API with free access subject to rate limits. Get a production key and refer to the API documentation to get started.

    For managed production deployment without rate limits, provision a dedicated Model Vault from your Cohere dashboard. Pricing is calculated per instance-hour, with discounted plans available for longer-term commitments.

    Let us know

    We want to see what developers build with Cohere Transcribe Arabic. Share your projects with us on X (@Cohere), Discord, or Hugging Face. These are also the best places to provide feedback and discuss new features or integrations.

    Key contributors: Shaun Cassini, Sebastian Vincent, Xiaolu Lu, Julian Mack, Dhruti Joshi, Pierre Richemond.

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Cohere and hundreds of other software products.

    Create account
  • Jun 9, 2026
    • Date parsed from source:
      Jun 9, 2026
    • First seen by Releasebot:
      Jun 10, 2026
    Cohere logo

    Cohere

    Introducing North Mini Code: Cohere’s first model for developers

    Cohere launches North Mini Code, its first open-source agentic coding model for developers. The 30B MoE model is built for efficient code generation, software engineering, and terminal tasks, with strong performance, fast throughput, and availability on Hugging Face, Model Vault, Cohere API, and OpenRouter.

    Small, efficient, and open-source — our first agentic coding model, built for the sovereign developer ecosystem.

    Today we're launching North Mini Code open-source. A mixture-of-experts (MoE) model, North Mini Code is Cohere's first agentic coding model, and the inaugural member of our next generation of powerful models.

    At 30B total parameters with just 3B active, North Mini Code delivers strong software development performance without demanding extensive hardware to match. Efficient by design, it's built to run where you need it.

    Freely available under an Apache 2.0 license, North Mini Code advances Cohere’s mission to make sovereign AI a practical reality, giving developers direct access to agentic coding capabilities. We're building in the open, because the future of AI should be shaped by the people running, testing, and improving it.

    Download the weights on Hugging Face, or deploy in a dedicated, managed inference environment on Model Vault. Alternatively, try it for free in your harness of choice on OpenCode or with a Cohere API key. Share what you build and tag @ Cohere on X or Discord, or engage with us on Reddit.

    Snapshot

    Model: North-Mini-Code-1.0

    License: Apache 2.0

    Model size: 30B total; 3B active

    Context length: 256K total context; 64K max generation

    Optimized for: Code generation, agentic software engineering, and terminal tasks

    Availability: Hugging Face (Weights), Cohere API, Cohere Model Vault, OpenRouter

    Hardware (minimum): 1× H100 @ FP8

    Agentic coding capabilities

    North Mini Code achieves competitive scores across benchmarks against models of this size class, demonstrating strong performance in real-world software engineering tasks.

    Image 1: North Mini Code’s performance in agentic software engineering and terminal tasks, along with complex code generation benchmarks, compared to leading open-source models of a similar size. ¹ ²

    North Mini Code’s benchmark scores translate to a 33.4 on the Artificial Analysis Coding Index, a competitive position among similarly sized models.

    The speed advantage for developer tasks

    North Mini Code is designed for speed and efficiency, with a strong focus on minimizing total cost of ownership as we continue to refine and scale the model.

    In our testing, North Mini Code achieved up to 2.8x higher output throughput than Devstral Small 2 under identical concurrency levels and hardware configurations. In practical terms, that translates to nearly three times the work rate, enabling faster iteration while reducing computational overhead.

    North Mini Code also demonstrated a 30% advantage in inter-token latency, a metric that reflects the consistency and pacing of token generation. Time-to-first-token (TTFT) performance was more closely matched between the two models, with Devstral Small 2 maintaining a slight edge across the tested conditions.

    Image 2: North Mini Code’s output speed and latency compared to Devstral Small 2, across high and low concurrencies, in internal tests using coding prompts.

    Sovereign open models for developers

    North Mini Code is our first open-source model for developers. As coding agents transform software engineering, developers need control and flexibility over their agentic coding infrastructure.

    North Mini Code represents a step forward in small agentic coding models that can accomplish tasks that matter to developers. Specifically, it is built for agentic workflows, including understanding and orchestrating sub-agents, mapping systems architecture, and running code reviews. Deploy on-prem or locally, on your own terms.

    Community feedback will directly shape our roadmap as we expand the ecosystem toward more open and sovereign developer models. Try North Mini Code when you need freedom from vendor constraints, and help us build what's next.

    What’s next?

    North Mini Code launches as the first—but certainly not the last—of Cohere's new generation of powerful models, designed for a more sovereign open-source ecosystem.

    We're committed to increasing our capabilities, with community input informing what comes next.

    Getting started

    Help us build a complete sovereign AI ecosystem for software development by trying North Mini Code. North Mini Code is available for free on Hugging Face and Model Vault—our fully managed inference platform. We've specifically trained it for compatibility with OpenCode, but it works with most coding agents.

    Share what you build and tag @ Cohere on X or Discord, or engage with us on Reddit to help shape the future of sovereign models.

    Visit our documentation for detailed model specs, deployment guides, and cookbooks to get started.

    Footnotes

    1 We used publicly reported scores for competitor models either from original reports or Artificial Analysis Intelligence Index where available. Additionally, Gemma 4’s scores for agentic coding tasks were reported by Qwen team. For the benchmark results that any public report is missing denoted by (*) in Image 1, we run internally with recommended model configuration.

    2 We evaluated North Mini Code using “SWE-agent” harness for SWE-Bench Verified and SWE-Bench Pro, and a simple ReAct harness employing a single terminal-use tool for Terminal Bench v2. For Terminal Bench Hard, we used Terminus-2 harness for both North Mini Code and the other models that are evaluated internally.

    Original source
  • Jun 9, 2026
    • Date parsed from source:
      Jun 9, 2026
    • First seen by Releasebot:
      Jun 9, 2026
    Cohere logo

    Cohere

    Announcing Cohere's North-Mini-Code-1.0

    Cohere releases North-Mini-Code-1.0, its first agentic coding model, built as a 30B total, 3B active Mixture of Experts model for local-friendly coding workloads. It is available through the Chat V2 API, as open weights on Hugging Face, and with Model Vault deployment support.

    We're pleased to announce the release of North-Mini-Code-1.0, Cohere's first agentic coding model. It is a 30 billion total / 3 billion active parameter Mixture of Experts model trained specifically for agentic coding, with a small enough active footprint to run on local hardware.

    Technical Details

    Model Name: north-mini-code-1-0

    Context Length: 256K input, 64K output

    License: Apache 2.0

    Availability

    North-Mini-Code-1.0 is available through the Chat V2 API and as open weights on Hugging Face. For production use, Model Vault deployment is also supported.

    For more details, see the model documentation.

    Original source
  • May 20, 2026
    • Date parsed from source:
      May 20, 2026
    • First seen by Releasebot:
      Jun 1, 2026
    Cohere logo

    Cohere

    Introducing Command A+: Making sovereign agentic capabilities available to all

    Cohere releases Command A+, an open-source enterprise LLM for complex reasoning, multimodal and multilingual agentic tasks. The model runs efficiently on as little as two H100 GPUs, is freely available under Apache 2.0, and brings faster inference with broader language coverage.

    Our fastest and most powerful language model yet.

    Command A+ is an open-source enterprise workhorse built for complex reasoning, multimodal and multilingual agentic tasks — all while running on as little as two H100 GPUs.

    Today, we’re releasing Command A+ open-source. A mixture-of-experts (MoE) model, Command A+ is an efficient, versatile, and privately deployable LLM built for high-performance agentic tasks with minimal compute overhead.

    Born from a year of deploying North with our customers, it surpasses every previous generation in the Command series and unifies their capabilities into a single scalable model.

    Now freely available under an Apache 2.0 license, Command A+ advances Cohere’s mission to make sovereign AI a technological reality — giving developers direct access to enterprise-grade agentic capabilities across experimentation, deployment, and production workflows.

    Visit Hugging Face to download the weights - available in several near lossless quantizations - and read our implementation guides. For a dedicated, managed inference environment, deploy Command A+ in Model Vault today.

    Snapshot:

    • Model: command-a-plus-05-2026
    • License: Apache 2.0
    • Architecture: Sparse / MoE
    • Model size: 218B total; 25B active
    • Context length: 128K input context; 64K max generation
    • Input modalities: Text, image, tool use
    • Output modalities: Text, reasoning, tool use
    • Languages: Supports 48 languages.
    • Optimized for: Reasoning, agentic workflows, RAG, multilingual, multimodal document processing
    • Supported frameworks: vLLM, Transformers
    • Hardware (minimum): 1× B200 @ W4A4, 2× H100s @ W4A4

    Northwards:

    For the past year, North — Cohere’s integrated enterprise workspace for building and deploying agentic AI — has been the driving force behind much of our innovation. Through that work, we set out to build a unified model for customers that simplifies deployment, can run locally, and synthesizes capabilities from across the Command family.

    The work is already paying off. Read how our customers have been using North to transform their operations.

    However, sovereign AI is much bigger than Cohere. Empowering engineers with models that they can run, control, and adapt themselves is the most acute challenge facing this generation of AI.

    We’ve optimized Command A+ for practical, developer-focused use, including support for low-bit quantization, efficient inference, and integration across open inference frameworks. AI independence for all.

    We can’t wait to see what the community builds.

    Command, consolidated:

    Command A+ outperforms previous Command A models in key dimensions of enterprise workloads, including multimodal understanding, retrieval, long-horizon, and complex reasoning.

    Compared with Command A Reasoning, 𝜏²-Bench Telecom scores improved from 37% to 85%, with agentic coding performance on Terminal-Bench Hard reaching 25% from 3%. Gains were also achieved on non-agentic reasoning, instruction following, and other code generation tasks.

    Command A+ performs strongly within North applications, reflecting its original design goals. Agentic Question Answering accuracy and spreadsheet analysis quality improved by 20% and 32% over Command A Reasoning, respectively. Memory performance — testing North’s skill in reasoning across conversations and stored data — scored 54% with Command A+ compared to 39% with Command A Reasoning.

    For multimodal understanding and reasoning, Command A+ achieved 63% on MMMU Pro and 75.1% on MMMU, (compared with 65.3% for Command A Vision for the latter). MathVista scores increased from 73.5% to 80.6%, and CharXiv reasoning improved from 46.9% to 52.7%, reflecting broad gains across document understanding tasks.

    Command A+ significantly expands multilingual capability, broadening language coverage from 23 to 48 languages and recording gains in machine translation and multilingual reasoning.

    Command A+ achieved a score of 37 on the Artificial Analysis Intelligence Index, outperforming other leading open models, reflecting its strength as a general-purpose model for enterprise agentic workflows.

    Efficiency at scale:

    Efficiency is a core constraint in enterprise AI deployment. It determines whether a language model can be deployed practically at scale by shaping the compute, memory, latency, power, and infrastructure required to serve it reliably and cost-effectively.

    We engineered Command A+ to be extremely hardware efficient. The model is available today on Hugging Face in 16-bit (BF16), 8-bit (FP8), and 4-bit (W4A4) quantizations, with imperceptible differences in quality. In practice, this enables Command A+ to run on as little as two NVIDIA H100s or a single NVIDIA Blackwell GPU, with virtually no quality degradation.

    Command A+ is also our fastest model to date, having 218B total and 25B active parameters compared to Command A Reasoning’s 111B dense architecture. At the same quantization and concurrency levels, it delivers up to 63% higher Output Tokens per Second (TOPS), and reduces Time To First token (TTFT) by up to 17%. The W4A4 quantization contributes an additional 47% increase in speed and a further 13% reduction in latency.

    We’re also using speculative decoding to accelerate text generation without impacting output quality. We optimized the approach specifically for the model’s MoE architecture, delivering an additional 1.5-1.6x inference speedup for both text and multimodal inputs.

    Command A+ is the first model to use our latest tokenizer, delivering substantial compression improvements over its predecessor. Fewer tokens are now required to generate the same response, reducing a major driver of inference cost. Notably, these gains extend to major non-European languages, which are often underrepresented during tokenizer training. Tokenization efficiency improved by 20% for Arabic, 16% for Korean, and 18% for Japanese.

    Fujitsu believes Command A+’s mixture-of-experts architecture and strong agentic performance align well with our commitment to deliver innovative, sovereign AI solutions through Takane and the Kozuchi Enterprise AI Factory. We look forward to leveraging its capabilities to accelerate secure, scalable AI adoption for our customers. — Vivek Mahajan, Corporate Executive Officer, Corporate Vice President, CTO, in charge of System Platform, Fujitsu Limited

    What’s next?

    Progress in sovereign AI today depends on advancing three fronts simultaneously: performance, security, and cost. At Cohere, we are investing across all three — both in our models and in the domain-specific capabilities that power North.

    That means improving reasoning, multimodal understanding, and coding performance, while ensuring models remain fit to run entirely within customer environments. The goal is not just stronger benchmarks, but systems that can support enterprise-wide transformation under real operational constraints.

    We have already applied this approach across our other model families — including Embed, Rerank, and Transcribe — where we have achieved state-of-the-art performance alongside efficient, cost-aware inference.

    Getting started:

    Command A+ is available today on Hugging Face, as well as through Model Vault. You can also try the model for free on our Space or with a Cohere API key.

    Visit our documentation for detailed model specs, deployment guides, and cookbooks to get started.

    Original source
  • Similar to Cohere with recent updates:

  • May 20, 2026
    • Date parsed from source:
      May 20, 2026
    • First seen by Releasebot:
      May 20, 2026
    Cohere logo

    Cohere

    Announcing Cohere's Command A+

    Cohere releases Command A+, its latest Command A family model, bringing vision, reasoning, translation, and agentic task support together in one MoE model. It adds broader multilingual coverage, faster production deployment, and availability through Cohere’s standard API endpoints.

    We're pleased to announce the release of Command A+, the last model in the Command A family of models, combining support for vision inputs, reasoning capabilities, translation capabilities, and agentic tasks all within the same model. It is also notably our first Mixture of Experts (MoE) model with 25 billion active parameters ands 218 billion total parameters.

    Key Features

    • Agentic Applications: With notable performance increases in tool use and agentic tasks, Command A+ is the strongest agentic model in the Command family.
    • Expanded Multilingual Support: With 48 languages supported, including all official EU languages, this more than doubles the support of languages from our prior models.
    • Efficient & Fast: With as few as 1 x B200 or 2 x H100s required to deploy the model, and up to 110% throughput increase and 30% decrease in latency over Command A Reasoning, the model is designed for production-grade deployments.

    Technical Details

    • Model Name: command-a-plus-05-2026
    • Context Length: 128K input, 64K output
    • Languages covered: English, Arabic, Bulgarian, Bengali, Catalan, Czech, Danish, German, Greek, Spanish, Estonian, Persian, Finnish, Filipino, French, Irish, Hebrew, Hindi, Croatian, Hungarian, Indonesian, Icelandic, Italian, Japanese, Korean, Lithuanian, Latvian, Malay, Maltese, Dutch, Norwegian, Punjabi, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Serbian, Swedish, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Vietnamese, Chinese.
    • License: Apache 2.0

    Availability

    Command A+ (command-a-plus-05-2026) is now available for all Cohere users through our standard API endpoints. For enterprise customers, private deployment options are available to ensure maximum security and control over your translation workflows.

    For more detailed information about Command A+, including technical specifications and implementation examples, visit our model documentation.

    Original source
  • Apr 4, 2026
    • Date parsed from source:
      Apr 4, 2026
    • First seen by Releasebot:
      Apr 6, 2026
    Cohere logo

    Cohere

    Retirement of Embed v2.0 and Aya Expanse / Vision 8B

    Cohere retires several embedding and chat models and points users to newer replacements.

    Effective April 4, 2026, the following models are no longer available. Requests using these model IDs will fail.

    Retired models

    • embed-english-v2.0
    • embed-english-light-v2.0
    • embed-multilingual-v2.0
    • c4ai-aya-expanse-8b
    • c4ai-aya-vision-8b

    We recommend these replacements:

    Embedding tasks

    • embed-english-v3.0
    • embed-multilingual-v3.0
    • embed-v4.0

    Chat tasks

    • command-r7b-12-2024
    • command-a-03-2025
    • command-a-reasoning-08-2025

    For the full announcement and lifecycle context, see the Deprecations page. For questions or assistance, contact [email protected].

    Original source
  • Mar 26, 2026
    • Date parsed from source:
      Mar 26, 2026
    • First seen by Releasebot:
      Jun 1, 2026
    Cohere logo

    Cohere

    Introducing Cohere Transcribe: a new state-of-the-art in open-source speech recognition

    Cohere launches Transcribe, an open-source ASR model for accurate speech-to-text across 14 languages. It offers strong real-world transcription performance, best-in-class throughput, and flexible deployment through Hugging Face, API access, and Model Vault.

    Cohere is announcing Transcribe, a state-of-the-art automatic speech recognition (ASR) model that is open source and available today for download.

    Speech is rapidly becoming a core modality for AI-enabled workloads and automations — from meeting transcription and speech analytics to real-time customer support agents.

    Our objective was straightforward: push the frontier of dedicated ASR model accuracy under practical conditions. The model was trained from scratch with a deliberate focus on minimizing word error rate (WER), while keeping production readiness top-of-mind. In other words, not just a research artifact, but a system designed for everyday use.

    Cohere Transcribe reflects that intent. It is available for open-source use with full infrastructure control, maintains a manageable inference footprint suitable for practical GPU and local utilization, delivers best-in-class serving efficiency, and is also available via Model Vault — Cohere’s secure, fully managed model inference platform.

    Cohere Transcribe currently ranks #1 for accuracy on HuggingFace’s Open ASR Leaderboard, setting a new benchmark for real-world transcription performance.

    This marks our zero-to-one in bringing high-performance speech recognition into enterprise AI workflows. Read on to learn more.

    Model overview

    Name: cohere-transcribe-03-2026
    Architecture: conformer-based encoder-decoder
    Input: audio waveform → log-Mel spectrogram
    Output: transcribed text
    Model size: 2B
    Model: a large Conformer encoder extracts acoustic representations, followed by a lightweight Transformer decoder for token generation
    Training objective: standard supervised cross-entropy on output tokens; trained from scratch
    Languages: trained on 14 languages:

    • European: English, French, German, Italian, Spanish, Portuguese, Greek, Dutch, Polish
    • APAC: Chinese (Mandarin), Japanese, Korean, Vietnamese
    • MENA: Arabic

    License: Apache 2.0

    Image 1: Cohere Transcribe is an open-weights Conformer ASR model converting speech audio into text across 14 supported languages.

    Model performance

    Accuracy

    Cohere Transcribe is the latest standard for English speech recognition accuracy. It leads the HuggingFace Open ASR Leaderboard with an average word error rate of just 5.42%, outperforming all open- and closed-source dedicated ASR alternatives, including Whisper Large v3, ElevenLabs Scribe v2, and Qwen3-ASR-1.7B. This captures the model’s versatile capability across real-world speech tasks, such as robustness to multiple-speaker environments, boardroom-style acoustics (e.g. AMI dataset), and diverse accents (e.g. Voxpopuli dataset).

    [Table of model WER scores omitted for brevity]

    Image 2: the Hugging Face Open ASR Leaderboard as of 03.26.2026. This is a widely used, standardized benchmark evaluating automatic speech recognition systems across curated datasets using word error rate (WER) as the primary metric, computed over normalized reference-hypothesis alignments, where lower WER indicates higher transcription fidelity. See the live leaderboard here.

    Critically, these gains aren’t limited to benchmark datasets. We see the same state-of-the-art performance carried over into human evaluations, where trained reviewers assess transcription quality across real-world audio for accuracy, coherence, and usability. Consistency across both evaluation methods reinforces that Cohere Transcribe’s performance translates reliably from controlled tests to practical enterprise settings.

    Image 3: human preference evaluation of model transcripts in English. In a pairwise comparison, annotators were asked to express preferences for generations which primarily preserved meaning - but also avoided hallucination, correctly identified named entities, and provided verbatim transcripts with appropriate formatting. A score of 50% or higher indicates that Cohere Transcribe was preferred on average in the head-to-head comparison.

    Image 4: human evaluation of ASR accuracy for a selection of supported languages. A score of 50% or higher indicates that Cohere Transcribe was preferred on average in the head-to-head comparison.

    Throughput

    In production settings, ASR systems must operate under strict latency and throughput constraints; even if accurate, slow or resource-intensive transcription can directly impact user experience, operational efficiency, and cost.

    Transcribe extends the Pareto frontier, delivering state-of-the-art accuracy (low WER) while sustaining best-in-class throughput (high RTFx) within the 1B+ parameter model cohort.

    Image 5: throughput (RTFx) vs accuracy (WER) plot for leading models larger than 1B in size. RTFx (real-time factor multiple) measures how fast an audio model processes its input relative to real time.

    “We’re genuinely impressed with what Cohere has built with Transcribe. The speed is exceptional — turning minutes of audio into usable transcripts in seconds — and it immediately unlocks new possibilities for real-time products and workflows. In our testing, the model handled everyday speech very well and delivered strong, reliable transcription quality. The overall experience has been smooth and easy to work with. We’re excited to be partnering with Cohere and to continue exploring what we can build with this technology.”

    — Paige Dickie, Vice-President, Radical Ventures

    Zero to one, and beyond.

    We are working towards deeper integration of Cohere Transcribe with North, Cohere’s AI agent orchestration platform. With planned updates, Cohere Transcribe will evolve from a high-accuracy transcription model into a broader foundation for enterprise speech intelligence.

    Getting started.

    Cohere Transcribe is now available for download on Hugging Face. Follow the setup instructions to run the model locally, or even in edge environments.

    You can also access Cohere Transcribe via our API for free, low-setup experimentation subject to rate limits. See the documentation for usage details and integration guidance.

    For production deployment without rate limits, provision a dedicated Model Vault. This enables low-latency, private cloud inference without having to manage infrastructure. Pricing is calculated per hour-instance, with discounted plans for longer-term commitments. Contact our team to discuss your requirements.

    Key contributors: Julian Mack (Member of Technical Staff), Ekagra Ranjan (Member of Technical Staff), Cassie Cao (Product Manager), Bharat Venkitesh (Manager of Technical Staff), Pierre Harvey Richemond (Manager of Technical Staff).

    Original source
  • Mar 26, 2026
    • Date parsed from source:
      Mar 26, 2026
    • First seen by Releasebot:
      Mar 26, 2026
    Cohere logo

    Cohere

    Announcing the Cohere Transcribe model

    Cohere releases Cohere Transcribe, its first transcription model for audio-in, text-out speech recognition. It supports multiple languages and is available now through the Audio Transcriptions API, with free experimentation and production deployment options.

    We're pleased to announce the release of Cohere Transcribe, our first transcription model.
    Cohere Transcribe specializes in audio-in, text-out, automatic speech recognition (ASR).

    Technical details

    • Model name: cohere-transcribe-03-2026
    • Input: Audio waveform
    • Output: Text
    • Languages covered: English, German, French, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Vietnamese, Chinese, Arabic, Japanese, Korean.
    • License: Apache 2.0
    • API endpoint: Audio Transcriptions API

    Getting started

    The model is available immediately through Cohere's Audio Transcriptions API endpoint.

    You can start transcribing audio using the following example query:

    import cohere
    co = cohere.ClientV2()
    response = co.audio.transcriptions.create(
    model="cohere-transcribe-03-2026",
    language="en",
    file=open("./sample.wav", "rb"),
    )
    print(response)
    

    Availability

    You can access Cohere Transcribe via our API for free, low-setup experimentation subject to rate limits. See the Different Types of API Keys and Rate Limits page for usage details and integration guidance.

    For production deployment without rate limits, provision a dedicated Model Vault. This enables low-latency, private cloud inference without having to manage infrastructure. Pricing is calculated per hour-instance, with discounted plans for longer-term commitments.

    Contact our team to discuss your requirements.

    Original source
  • Jan 28, 2026
    • Date parsed from source:
      Jan 28, 2026
    • First seen by Releasebot:
      Jun 1, 2026
    Cohere logo

    Cohere

    Introducing Model Vault: Your private platform for secure and scalable model inference

    Cohere launches Model Vault, a fully isolated SaaS platform for secure, high-performance model serving and scaling. It offloads inference operations, adds real-time monitoring, and supports flexible deployment for North and standalone customers.

    Key Contributors

    Manoj Govindassamy, Maxime Brunet, Jeremy Pekmez, Inna Shteinbuk, Elliott Choi

    Model Vault simplifies serving and scaling Cohere models, so teams can focus on building, not infrastructure.

    Cohere is launching Model Vault, a dedicated, fully isolated SaaS platform for customers to run Cohere models securely, at scale, and with guaranteed performance.

    Model Vault combines the reduced operational overhead of fully managed SaaS with the security and performance advantages of self-hosting. North users deploy their application within a secure VPC while offloading the maintenance-heavy model inference and scaling to their secure, cloud-based Model Vault. The result: lower cost of ownership and faster enterprise adoption.

    Model Vault marks a milestone in Cohere’s journey to deliver transformative AI that solves the real-world challenges faced by enterprise customers today. Read on to learn more about what’s on offer, or get started now.

    The convenience-control tradeoff

    At the heart of enterprise AI transformation lies a fundamental constraint: how to balance speed and scale in adoption with the infrastructure control and compliance requirements of a mature, modern-day organization.

    This is increasingly a problem of inference. As enterprises work to incorporate models into more workflows, teams, and products, inference is gradually replacing training as the dominant AI workload. McKinsey now projects that inference will account for the majority of AI compute by the end of the decade, even as per-unit costs continue to fall. For businesses, this shifts the core concern from funding episodic training runs to accounting for inference — a recurring cost that scales directly with increased adoption.

    For most enterprises, available deployment options clearly reflect this tension. Multi-tenant SaaS platforms optimize for speed and operational simplicity by amortizing infrastructure costs across multiple customers. However, the lack of workload isolation inherent to shared environments introduces problems like noisy-neighbour effects, enforced rate limits during peak usage, and unpredictable latency. These platforms also tend to limit model configurability, provide little visibility into workload-level performance, and often fall short of the compliance standards of highly regulated enterprises.

    Self-hosted deployments — whether on-premises or within a customer-managed VPC — address many of these limitations by offering greater control over infrastructure, performance characteristics, and security boundaries. The capital and operational burden underpinning that control, however, can be prohibitively costly, particularly for teams looking to scale.

    With North, Cohere’s enterprise platform for building agentic AI applications, this problem of infrastructure becomes even more acute. Model serving still requires provisioning and managing hardware, but agentic workloads by their design are bursty, multifaceted, and unpredictable.

    A new framework for scalable, secure AI deployment

    Model Vault is the latest addition to Cohere’s suite of secure deployment options. It is designed to address this tradeoff between convenience and control by transferring the operational overheads of model inference to a secure, Cohere-managed cloud environment.

    The result is a SaaS deployment model that removes the constraints of shared infrastructure, such as resource contention and unpredictable performance, while preserving the speed, elasticity, and ease of use associated with multi-tenant SaaS.

    To be clear, there is no single deployment approach that fits every organization. For some enterprises, multi-tenant SaaS or self-hosted deployments will remain the right choice, depending on internal capabilities, data strategy, and regulatory requirements. Industries with highly standardized workflows, for example, may meet compliance needs through application-level controls alone and continue to favor shared SaaS platforms.

    But for many enterprises, infrastructure management has become a growing barrier to production-grade agentic AI. They want to scale model inference capacity without scaling operational load. Model Vault is built for these teams.

    Decouple inference from development

    Model Vault is a Cohere-managed solution. That means we assume the full operational burden of production inference: deploying models, managing upgrades and dependencies, provisioning and scaling capacity, and ensuring performance and availability. We deliver Model Vault users 99.9%+ guaranteed availability, backed by production-grade SLOs and latency performance that matches or beats leading managed AI/ML deployment platforms.

    This materially lowers the total cost of ownership by removing the need for customers to procure, provision, and operate GPU-backed inference infrastructure. North deployments can now run entirely on CPU-only environments, while the highest overhead ML infrastructure work is done within your Model Vault.

    Model Vault lets ML teams scale workloads elastically without pre-allocating capacity or absorbing the risk of idle GPUs. Inference capacity expands and contracts with demand, eliminating the tension between overprovisioning for peak performance and underprovisioning that degrades latency and availability.

    In simple terms, Model Vault lets Cohere absorb the operational complexity of model serving, freeing your most valuable asset — your engineers — to focus on moving agentic AI applications from experimentation to production. Inference is not your differentiator.

    Throughout, enterprise control is retained where it matters most. North customers keep full ownership of the control plane, covering agent logic, workflow orchestration, conversation states, data storage, pipelines. They have full architectural flexibility over which data and workloads they choose to host on-prem or on their existing VPCs, and which can be handled by Model Vault.

    No sharing. No limits. No surprises.

    Model Vault is deployed as a logically isolated virtual private cloud. No infrastructure components are shared across customers, ensuring strong isolation for performance, reliability, and security. This includes dedicated network load balancers, reverse proxies, serving middleware, inference servers, and the underlying GPU accelerators.

    Teams can now use Model Vault to run unlimited production workloads without competing for scarce capacity. The resources you provision are reserved for you — no more rate limits to maintain platform stability. At the same time, Model Vault dynamically scales inference capacity according to customizable logic, so performance remains consistent even for bursty agentic inference.

    Observe and optimize in real time

    Model Vault gives MLOps teams unique visibility into how inference workloads behave in production. Through a dedicated, real-time dashboard, teams can track request patterns, latency, token throughput, and resource utilization, enabling them to quickly pinpoint bottlenecks and continuously optimize for efficiency, predictability, and performance as demand evolves.

    Setting up a new Vault is quick and easy. Select your desired model and performance tier (S, M, L, XL), the replica range, and click Create Vault. You can now run Cohere models in your system.

    Manage your Vault directory and all deployed models from your Cohere dashboard.

    Model Vault also equips teams with real-time monitoring of model usage and performance.

    Getting started

    Model Vault is designed for quick and intuitive onboarding, with minimal manual setup. It can be integrated into your existing North deployment, or serve as a private inference platform for Cohere models used individually.

    Here’s how to create your Model Vault in minutes:

    Without North:

    1. Go to dashboard.cohere.com. Create an account and log in.
    2. Find Vaults in the side bar, then select “+ New Vault.”
    3. Name your Vault.
    4. Select your Cohere model, and the performance tier.
    5. Specify the desired range of replicas for each selected model. This defines your scaling limits.
    6. Confirm your Vault.

    That’s it! Your infrastructure is provisioned and you’ll instantly receive your unique subdomain, API endpoint, and access to your private monitoring dashboard.

    With North:

    1. Follow the steps above to create a dedicated Model Vault for your organization.
    2. Make sure you select the models that you need for your North deployment.
    3. Once you have your API endpoint, update the Helm chart in your North deployment, so that the desired models point to that endpoint.
    4. After setup, customers can manage the primary and secondary inference endpoints for Command models in North directly from the admin interface.
    5. Your North deployment is now configured with Model Vault.

    Supported models

    Model Vault currently supports all of Cohere’s latest state-of-the-art embedding, reranker, and generative models. This includes bundled offerings for our North customers (see table below).

    What’s more, Rerank and Embed models, which underpin Cohere’s state-of-the-art AI search and retrieval capabilities, can also be self-served through your Model Vault. That means on-demand access for seamless setup and faster experimentation. Customers interested in self-serve access to our Command models or integrated model bundles can request to join the waitlist.

    Model Vault pricing plans are flexible and can be tailored to your team’s workload requirements and budgeting preferences. Speak with one of our team to find what works best for you.

    You can also read our documentation for full, step-by-step instructions, as well as our model specs and how-to guides for building agentic AI applications in North.

    Build fast. Stay in control.

    Start shipping secure, high-performance enterprise AI with Model Vault today. Get started.

    Original source
  • Dec 11, 2025
    • Date parsed from source:
      Dec 11, 2025
    • First seen by Releasebot:
      Jun 1, 2026
    Cohere logo

    Cohere

    Introducing Rerank 4: Cohere’s most powerful reranker yet

    Cohere releases Rerank 4, its most advanced reranker for enterprise AI search, with stronger retrieval relevance, lower latency, 32K context, multilingual support across 100+ languages, self-learning customization, and flexible deployment on Cohere Platform, SageMaker AI, and Microsoft Foundry.

    Key Contributors

    Clifton Poth (Member of Technical Staff), Fabian David Schmidt (Member of Technical Staff), Martin Hentschel (Member of Technical Staff), Daniel Simig (Manager of Technical Staff), Nils Reimers (VP of Embeddings & Search), Elliott Choi (Director of Product)

    Rerank 4 is the most advanced set of reranker models available today, purpose-built to meet the realities and challenges of enterprise AI search. It delivers best-in-class retrieval, outperforming the likes of MongoDB’s Voyage models and ElasticSearch’s Jina rerankers in overall search relevance, as well as improved latency, flexible deployment options, deep customizability and robust multilingual performance. Designed for business-critical applications across key industries and domains, Rerank 4 sets a new standard for accuracy and adaptability in enterprise search.

    Why reranking matters

    Rerankers significantly enhance the accuracy of enterprise AI search by refining initial retrieval results. After a fast but broad candidate-generation step using methods like BM25 or bi-encoder embeddings, the system is left with documents, passages, or product listings that approximate relevance but often miss the nuance of the user’s intent. Rerank 4 addresses this gap using a cross-encoder architecture that processes queries and candidates jointly, capturing subtle semantic relationships and reordering results to surface the most relevant items. This approach balances speed and accuracy, delivering high-quality results without evaluating the entire corpus.

    Rerank 4 is a key component of North, Cohere’s agentic AI platform, which combines intelligent search (Embed, Rerank), large language models (our Command series), and customizable AI agents to automate tasks and accelerate decision-making. It integrates seamlessly into existing AI search solutions, including hybrid, vector, and keyword-based systems, with minimal code changes. Beyond improving retrieval-augmented generation (RAG) pipelines, Rerank 4 is critical in agentic AI systems, providing distilled, high-quality information for stronger reasoning and context-aware agent performance. By filtering out irrelevant content before it reaches the generative model, Rerank 4 reduces token usage and minimizes the number of costly retries an agent might otherwise need to “get it right,” sparing the user both pennies and latency. This is especially impactful for agentic AI, where complex, multi-step interactions can quickly drive up model calls and saturate context windows.

    Rerank 4 Fast & Pro

    Rerank 4 boasts a 32K context window - the largest of our Rerank series to date and a four-fold increase over the previous generation. This enables the model to handle longer documents, evaluate multiple passages simultaneously, and capture relationships across sections that shorter windows would miss. This expanded capacity therefore improves ranking accuracy for realistic document types and increases confidence in the relevance of retrieved results.

    To support different search workloads, Rerank 4 is available in two versions. Fast is a smaller model designed for use cases where both speed and accuracy are important. Pro, meanwhile, is optimized for tasks that require deeper reasoning, analysis and ironclad precision.

    Rerank 4 Fast - where speed wins:

    1. E-commerce. Prospective customers benefit from high relevance product recommendations as well as the option to explore online catalogues with semantic search. Here, faster results equal a better shopping experience and higher conversion rates.
    2. Programming. During a project, engineers need to quickly and conveniently reference documentation - like design specs - without derailing their creativity or collaboration.
    3. Customer service. Enterprise help desk agents need to efficiently triage incoming tickets to resolve cases, prevent backlogs, and keep their customers happy.

    Rerank 4 Pro - where depth matters:

    1. Finance. When generating risk models, analysts rely on troves of market reports, regulatory filings, and transaction histories. More retrieval time means richer data analysis and more accurate scenarios modeled.
    2. Healthcare. Clinicians routinely review patient records and clinical trial reports to identify appropriate courses of treatment. Accuracy and completeness of information can have devastating consequences for patient welfare.
    3. Manufacturing. Quality engineers consult technical manuals and production logs to diagnose deviations in highly complex processes that led to product defects. More informed decision-making helps productivity, safety, and end customer satisfaction.

    Superior performance in enterprise-specific and multilingual tasks

    Our customers need to know that their semantic search is not just performant, but works for the queries, workflows, languages, and data types that constitute their everyday work. This means the model has to understand the nuances of each user and the problem they're trying to solve.

    Benchmarking against industry standards confirms that Rerank 4 outperforms all current competitive alternatives in a broad array of enterprise domains, such as finance, healthcare, manufacturing and others.

    Rerank 4 also leads the pack in multilingual performance. Users can deploy Rerank in over 100 world languages, including state-of-the-art retrieval in 10 major business languages (Arabic, Chinese, French, German, Hindi, Japanese, Korean, Portuguese, Russian, and Spanish).

    Self Learning

    Semantic search is an investment for your company. Like any human-executed task, semantic search will see greater returns once it ‘learns’ the means to do its job more efficiently. On designing Rerank 4, we asked ourselves: how can we bake this concept into our product offerings?

    Rerank 4 is our first reranking model with self-learning capability. In partnership with Cohere, users can customize their model for certain use cases - such as those they encounter most frequently - without the need for further annotated data. This might involve stating preferences for particular content types, for use of prescribed language or terminology, or directing the model to specific document corpora.

    An example: imagine a Consumer Loan Specialist at a commercial bank. Their job is to evaluate borrowing applications and guide customers through the process of acquiring a loan. Each case requires frequent references to internal documentation relating to, say, approval criteria, product details, regulatory standards, or even company policies. Fast retrieval is key for customer satisfaction - but accuracy can’t be compromised either. At least not within this narrow domain.

    Rerank 4 moves towards overcoming this tradeoff in speed and quality through Self Learning. With automated adaptation to domain-specific use cases, the accuracy of Rerank 4 Fast iteratively improves to become competitive with much larger market alternatives. In many of our tests, Rerank 4 Fast began to converge and even surpass out-of-the-box Pro performance. Self-Learning Rerank will soon bring continuous RAG optimization to North, advancing precision and robustness in domain‑specific workflows.

    Looking further, we also explored how Rerank 4’s self-learning capability performs on entirely new search domains. Using healthcare-focused datasets that mimic a clinician’s need to retrieve patient-specific information - not just expertise from a given medical discipline - we found that enabling Self Learning produced consistent, substantial gains. The result: a clear and significant boost in retrieval quality for Rerank 4 Fast, across the board.

    Getting started

    Rerank 4 is now available on Cohere’s Platform, Amazon SageMaker AI (Fast & Pro) and Microsoft Foundry, with additional platform support coming soon.

    Click here for pricing options, and consult our docs for learn more about Rerank 4 setup.

    Rerank 4 can also be deployed into any Virtual Private Cloud (VPC) or on-premise environment. To learn more about options for large enterprise deployments, please contact our sales team.

    Original source
  • Dec 11, 2025
    • Date parsed from source:
      Dec 11, 2025
    • First seen by Releasebot:
      Mar 20, 2026
    Cohere logo

    Cohere

    Cohere's Rerank v4.0 Model is Here!

    Cohere releases Rerank 4.0, its newest foundational ranking model with two variants for quality or speed, multilingual and JSON document support, and a 32k token context window for longer, more capable re-ranking.

    We're pleased to announce the release of Rerank 4.0 our newest and most performant foundational model for ranking.

    Technical Details

    • Two model variants available:rerank-v4.0-pro: Optimized for state-of-the-art quality and complex use-cases
    • rerank-v4.0-fast: Optimized for low latency and high throughput use-cases
    • Multilingual support: Re-rank both English and non-English documents
    • Semi-structured data support: Re-rank JSON documents
    • Extended context length: 32k token context window

    Example Query

    import cohere
    co = cohere.ClientV2()
    query = "What is the capital of the United States?"
    docs = [
    "Carson City is the capital city of the American state of Nevada. At the 2010 United States Census, Carson City had a population of 55,274.",
    "The Commonwealth of the Northern Mariana Islands is a group of islands in the Pacific Ocean that are a political division controlled by the United States. Its capital is Saipan.",
    "Charlotte Amalie is the capital and largest city of the United States Virgin Islands. It has about 20,000 people. The city is on the island of Saint Thomas.",
    "Washington, D.C. (also known as simply Washington or D.C., and officially as the District of Columbia) is the capital of the United States. It is a federal district. The President of the USA and many major national government offices are in the territory. This makes it the political center of the United States of America.",
    "Capital punishment has existed in the United States since before the United States was a country. As of 2017, capital punishment is legal in 30 of the 50 states. The federal government (including the United States military) also uses capital punishment.",
    ]
    results = co.rerank(
    model="rerank-v4.0-pro", query=query, documents=docs, top_n=5
    )
    
    Original source
  • Sep 16, 2025
    • Date parsed from source:
      Sep 16, 2025
    • First seen by Releasebot:
      Mar 20, 2026
    Cohere logo

    Cohere

    Announcing Major Command Deprecations

    Cohere deprecates several legacy models, fine-tuning options, API endpoints, and UI apps, while steering users to newer command-r models and command-a for replacements.

    As part of our ongoing commitment to delivering advanced AI solutions, we are deprecating the following models, features, and API endpoints:

    Deprecated Models

    • command-r-03-2024 (and the alias command-r)
    • command-r-plus-04-2024 (and the alias command-r-plus)
    • command-light
    • command
    • summarize (Refer to the migration guide for alternatives).

    For command model replacements, we recommend you use command-r-08-2024, command-r-plus-08-2024, or command-a-03-2025 (which is the strongest-performing model across domains) instead.

    Retired Fine-Tuning Capabilities

    All fine-tuning options via dashboard and API for models including command-light, command, command-r, classify, and rerank are being retired. Previously fine-tuned models will no longer be accessible.

    Deprecated Features and API Endpoints

    • /v1/connectors (Managed connectors for RAG)
    • /v1/chat parameters: connectors, search_queries_only
    • /v1/generate (Legacy generative endpoint)
    • /v1/summarize (Legacy summarization endpoint)
    • /v1/classify
    • Slack App integration
    • Coral Web UI (chat.cohere.com and coral.cohere.com)

    For questions, reach out to [email protected]

    Original source
  • Aug 28, 2025
    • Date parsed from source:
      Aug 28, 2025
    • First seen by Releasebot:
      Jun 1, 2026
    Cohere logo

    Cohere

    Command A Translate: Secure translation for global enterprises

    Cohere introduces Command A Translate, a secure enterprise translation model with private deployment options, fine-tuning support, and coverage across 23 business languages. It also adds Deep Translation, an agentic multi-step approach for even higher-quality results.

    The new industry standard for secure, enterprise-ready machine translation.

    Today, we’re introducing Command A Translate, a state-of-the-art model specifically designed for high-quality translation tasks. It consistently outperforms all other models, including GPT-5, DeepSeek-V3, DeepL Pro’s LLM, and Google Translate. We can also push its quality even higher through Deep Translation, our new agentic approach that uses a multi-step process to refine translations.

    Command A Translate delivers industry-leading performance while offering enterprises full control of their data through private deployment options. Translation frequently involves handling sensitive business documents, such as contracts, financial reports, and materials with confidential customer information that shouldn’t be exposed to consumer services. For global companies operating across multiple languages, it ensures translations remain precise, reliable, and highly secure.

    Deep Translation Elevates Leading Translation Quality

    Command A Translate is built for enterprises that prioritize high-quality translations to reduce rewrites, lower costs, and enable teams across geographies to work together seamlessly. The model excels at key translation benchmarks, including speech, news, social media, and literary text. It achieves best-in-class results across a wide range of business languages.

    On the WMT24++ dataset, Command A Translate achieved a best-in-class xCometXL score of 83.8, and a score of 84.4 with Deep Translation. This shows the average xCometXL score for the 23 languages covered by Command A Translate on the English to L2 datasets of WMT24++. *DeepL Pro does not cover Hindi and Persian and the numbers for those languages were estimated through nearest neighbor imputation.

    Performance can be further enhanced through our innovative agentic approach: Deep Translation. Deep Translation achieves unmatched quality through iterative reasoning for more complex translation use cases. It allows the model to gradually refine translations through multiple steps improving fluency and naturalness, ensuring the output reads as if originally written in the target language. This is especially valuable for use cases with the highest quality bar, like translating legal documents. Deep Translation is available today by request, reach out to learn more.

    To meet the needs of global enterprises, the model supports translation across 23 widely used business languages: English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Arabic, Chinese, Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew, and Persian.

    Privately deployable and fine-tunable

    For enterprises needing to translate sensitive information without sending it to an external service, Command A Translate can be privately hosted behind your secured firewalls. This allows businesses to break down internal language silos while maintaining complete control over their data.

    Due to its efficient design, Command A Translate can be deployed with minimal hardware requirements. For low footprint deployments, enterprises can bring it to production on one GPU (H100/A100) using 4-bit quantization without a noticeable decrease in translation quality (less than 0.5 xCometXL points).

    For those organizations with unique needs, from industry specific translations to supporting new languages, we offer fine-tuning and customization support for production use cases.

    Why secure translation matters

    Enterprises rely on translation for some of their most sensitive and business-critical documents. They cannot risk data leakage, compliance violations, or misunderstandings. Mistranslated documents can reduce trust and have strategic implications. Entire industries depend on translation as a core service, from localization providers to companies producing subtitles, advertising, and news at a global scale.

    At Cohere, we use Command A Translate internally to ensure we have the highest quality multilingual data possible. Early testers like leading language and content solutions company RWS have confirmed Command A Translate’s strong performance across complex translation tasks.

    "RWS and the Language Weaver team have been working closely with Cohere to evaluate their Command A Translate model, optimized for automated translation tasks. We rigorously tested it across 23 languages and multiple domains, comparing it against other leading providers using both automated metrics and human reviews from RWS’s professional linguists. Command A Translate consistently demonstrated world-class quality in every comparison, seamlessly handling complex challenges like numbers, names, and polysemous words with precision. We look forward to continuing this collaboration to deliver innovative AI-powered translation applications for customers." – Dragos Munteanu, Vice President of Research and Development, RWS

    Availability

    Command A Translate is available today on the Cohere platform and for research use on Hugging Face. If you are interested in private or on-prem deployments, please contact our sales team for bespoke pricing.

    Original source
  • Aug 28, 2025
    • Date parsed from source:
      Aug 28, 2025
    • First seen by Releasebot:
      Mar 20, 2026
    Cohere logo

    Cohere

    Announcing Cohere's Command A Translate Model

    Cohere launches Command A Translate, its first machine translation model, bringing accurate, fluent translations across 23 languages with long-context support, strong deployment efficiency, and secure options for sensitive data. It is now available through Cohere’s Chat API and standard endpoints.

    We're excited to announce the release of Command A Translate, Cohere's first machine translation model. It achieves state-of-the-art performance at producing accurate, fluent translations across 23 languages.

    Key Features

    • 23 supported languages: English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Chinese, Arabic, Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew, and Persian
    • 111 billion parameters for superior translation quality
    • 16K token context length (8K input + 8K output) for handling longer texts
    • Optimized for deployment on 1-2 GPUs (A100s/H100s)
    • Secure deployment options for sensitive data translation

    Getting Started

    The model is available immediately through Cohere's Chat API endpoint. You can start translating text with simple prompts or integrate it programmatically into your applications.

    from cohere import ClientV2
    co = ClientV2(api_key="<YOUR API KEY>")
    response = co.chat(
    model="command-a-translate-08-2025",
    messages=[
    {
    "role": "user",
    "content": "Translate this text to Spanish: Hello, how are you?",
    }
    ],
    )
    

    Availability

    Command A Translate (command-a-translate-08-2025) is now available for all Cohere users through our standard API endpoints. For enterprise customers, private deployment options are available to ensure maximum security and control over your translation workflows.

    For more detailed information about Command A Translate, including technical specifications and implementation examples, visit our model documentation.

    Original source
Releasebot

Curated by the Releasebot team

Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.

Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.