AI Voice and Speech Release Notes

Release notes for AI voice synthesis, text-to-speech and audio generation tools

Get this feed:

Products (12)

Latest AI Voice and Speech Updates

  • Sep 1, 2026
    • Date parsed from source:
      Sep 1, 2026
    • First seen by Releasebot:
      Sep 2, 2026
    Plaud logo

    Plaud

    Plaud One AI Earbuds Sell Out U.S. Pre-Sale in One Day

    Plaud releases Plaud One Explorer Edition, wearable AI earbuds built to capture conversations and turn them into action with Plaud Intelligence and Plaud Agent. The limited U.S. pre-sale sold out in one day, with shipments expected in Q4 2026.

    Good things move fast. Plaud One Explorer Edition sold out its U.S. pre-sale inventory just one day after launch.

    On August 27th, we introduced Plaud One Explorer Edition, a new pair of wearable AI earbuds designed for the agent era. Plaud One captures conversations, understands their context and intent, and turns what was said into action through Plaud Intelligence and Plaud Agent.

    Within a day, our first U.S. pre-sale inventory was gone.

    The response extended far beyond sales. Plaud One featured across Forbes, Bloomberg, TechCrunch, The Verge and other technology and business media, while the launch took Times Square by storm. CNET described Plaud One as "a reinvention of headphones for the AI note-taking age."

    "A reinvention of headphones for the AI note-taking age"

    • Katie Collins, CNET

    TechCrunch highlighted its eSIM-enabled case and ability to connect conversations to AI agents, while The Verge spotlit the innovation from previous form factors into AI-powered earbuds.

    What is Plaud One?

    Plaud One is wearable AI built around conversation.

    Every era of computing has had a defining interface. We typed on keyboards. We tapped on screens. As AI agents become capable of doing more on our behalf, we believe the next interface will be even more natural: conversation.

    Plaud One can be worn as earbuds or used through its standalone case to capture phone calls, online meetings, in-person conversations and ideas on the go. A built-in eSIM with 4G LTE connectivity allows the device to stay connected without relying on a nearby phone.

    But recording is only the starting point.

    Plaud Intelligence is designed to understand the context and intent inside conversations. Plaud Agent can then use that context across connected tools including Gmail, Google Calendar, Notion and Slack to help create follow-ups, documents, presentations and other finished work.

    In other words: capture the conversation, understand what matters, then get things done.

    Why AI earbuds?

    Some of the most important information in our lives never begins as a document or prompt. It starts as something someone says.

    A customer makes a request during a call. A decision gets made in a meeting. An idea surfaces while walking between appointments. A next step is agreed to over coffee.

    Traditional AI usually enters the process later, after someone remembers what happened and types it into a screen.

    Plaud One is designed to move AI closer to the moment where that context is actually created.

    For a sales professional, that could mean capturing a client conversation and turning it into a follow-up. For an executive, it could mean asking for the context from a previous meeting before walking into the next one. For a consultant, it could mean transforming hours of discussion into a structured findings report.

    That shift—from AI you go to, toward AI that can work from the context of your real-world conversations—is at the heart of what Plaud One represents.

    One day, a lot of momentum

    The reaction to Plaud One suggests people are ready to explore what that future could look like.

    The U.S. pre-sale sold out in one day, following a launch that generated attention across some of the world’s largest technology and business publications around the world.

    TechCrunch called attention to Plaud One’s eSIM-enabled case and AI agent capabilities. Forbes explored how the earbuds can move from capturing conversations to acting on them. The Verge highlighted the new form factor and its combination of recording, transcription, summarization and agent capabilities.

    And in New York, Plaud One showed up a little larger than an earbud usually does: right in the heart of Times Square.

    It was a big first day for a product built around a very simple idea:

    In the agent era, just speak.

    What happens next for Plaud One?

    Plaud One Explorer Edition is a limited release designed to put the next generation of wearable AI into the hands of early users and learn from how they use it in the real world.

    The Explorer Edition is priced at $249.99, with shipments expected to begin in Q4 2026. Each Explorer Edition includes $200 in Plaud Credits for Plaud Intelligence agent capabilities and other AI features.

    The first batch moved fast. We’re just getting started.

    Sign up for the Plaud One waitlist.

    Original source
  • Aug 29, 2026
    • Date parsed from source:
      Aug 29, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Speechify logo

    Speechify

    API: watermark verification with no credential

    Speechify adds a public watermark verification API that checks whether an audio clip carries a Speechify watermark without any credential. The new endpoint returns a simple true or false response, while the existing detect API still provides the confidence score for authenticated use.

    API: watermark verification with no credential

    A new endpoint answers whether a clip carries a Speechify watermark without any credential.

    POST /v1/audio/watermark/verify takes the same audio upload as POST /v1/audio/watermark/detect and returns a bare {"watermarked": true|false} — no key, no session, nothing to authenticate.

    It is the API half of the public tool at speechify.ai/detect, and it exists because California’s AI Transparency Act (BPC 22757.2) requires a detection tool that is publicly accessible and invokable without visiting a website.

    verify answers; detect measures. The new verb deliberately omits the confidence score its sibling returns — a public score is a gradient to optimise against. Keep using POST /v1/audio/watermark/detect with an API key when you want the score.

    Because verify takes no credential, it is rate-limited per client address and draws on a shared platform budget, so expect 429 under sustained automated use. Nothing changes for an existing integration: POST /v1/audio/watermark/detect’s path, request and response are untouched.

    Original source
  • All of your release notes in one feed

    Join Releasebot and get updates from Plaud and hundreds of other software products.

    Create account
  • Aug 28, 2026
    • Date parsed from source:
      Aug 28, 2026
    • First seen by Releasebot:
      Aug 29, 2026
    Wispr Flow logo

    Wispr Flow

    For Teams & Enterprise:

    Wispr Flow adds admin portal controls for Notetaker and new self-serve Growth plans for teams.

    Admins can now turn Notetaker on or off for their entire organization from the admin portal with a single toggle, enforced across every member's account. There's a confirmation step before disabling, so it never happens by accident.

    New self-serve Growth plans:

    Teams can now buy a new Growth plan straight from the admin portal, choosing dictation-only or bundling in unlimited Notetaker, with monthly or annual billing shown per currency. Upgrade your whole team in a couple of clicks with most of the features of our Enterprise plan, no sales call required.

    Original source
  • Aug 27, 2026
    • Date parsed from source:
      Aug 27, 2026
    • First seen by Releasebot:
      Aug 27, 2026
    Cartesia logo

    Cartesia

    Introducing Sonic-3.6

    Cartesia releases Sonic-3.6 with major gains in naturalness, multilingual quality, and speed, adding 61 locales, 500+ preset voices, Odia and Urdu, Hinglish support, stronger accent cloning, and enterprise deployment options. Now generally available in the playground and API.

    At Cartesia, we’re deeply invested in making foundational advances to AI architectures and algorithms, because that’s what drives step changes in model capabilities.

    Sonic-3.6 is the culmination of our latest advances in architectures that learn efficiently from multilingual audio across pre- and post-training. Sonic-3.6 makes huge strides compared to Sonic-3.5 in naturalness and quality, with listeners preferring it in up to 93% of blind head-to-head tests across fifteen locales.

    At the start of this year, we made a bet that instead of incrementally adjusting the existing paradigm, the way to do the best research in the world was to rethink everything from first principles. Since then, we’ve rebuilt our data, model architecture, training and evaluation from the ground up with novel ideas, and we’re seeing those efforts pay off.

    Sonic-3.6 lands just two months after Sonic-3.5, and it once again takes #1 on the Artificial Analysis leaderboard across both the controlled and provider voice boards. On the controlled board, the incumbent model we beat is Sonic-3.5. Overtaking our own model while no other provider has closed the gap is a testament to our accelerating research velocity.

    Closing the gap to natural speech

    For anyone running voice at scale, naturalness decides whether a customer stays on the line. A voice that sounds even slightly synthetic makes callers lose trust and ask for a human. We built Sonic-3.6 to sound natural and native enough to keep the conversation moving across 44 languages.

    In US English, listeners chose Sonic-3.6 over Eleven v3 92% of the time in blind head-to-head tests, thanks to better intonation, higher voice quality, and context-driven emotionality. Sonic-3.6 reads a line the way it’s meant to be heard and adapts delivery to the conversation’s context the way a person does.

    ¿Hablas español?

    We’ve always viewed English as table stakes, but sounding native in every other language or accent is the real frontier that decides whether voice AI stays a mostly-English technology or becomes something the whole world can actually talk to. Most models are intelligible abroad, but few are convincing enough to get the accent, tone, and local pronunciation of a name or a place right.

    We tested Sonic-3.6 the way a customer would judge it: by ear, in each language, against leading competitors. Native-speaker panels preferred Sonic-3.6 head-to-head in every core language we tested.

    What these improvements unlock across languages:

    • Broader reach: 61 locales (11 of them Indic) and 500+ preset voices now live, plus two new languages: Odia and Urdu, both launching above 96% transcript accuracy.
    • Hinglish support: Code-switch between Hindi and English in a single generation using transcripts in Devanagari, Latin script, or a mix.
    • Improved accent adherence for instant voice clones: Sonic-3.6 holds onto a speaker’s accent far better, including less widely spoken ones.
    • Locale-awareness: Set the optional locale field so that dates, times, and numbers come out the way a local would say them.

    Building for production

    Naturalness may win in the demo, but we know the live call is the real test. Built for enterprise workloads at scale, Sonic-3.6:

    • Replies under 90ms and generates nearly 2x faster than v3 Conversational (132 vs 68 characters/sec).
    • Runs at scale across cloud, on-prem, and localized endpoints, backed by a 99.9% uptime SLA.
    • Powers millions of production calls today with the reliability that volume demands.
    • Deploys in your own cloud through Baseten, Amazon SageMaker, or Together AI, so audio never has to leave your infrastructure.
    • Meets enterprise compliance, with SOC 2, PCI, HIPAA, and GDPR covered.
    • Fits your brand with hands-on help from our team to develop and find your custom voice.

    Hear it for yourself

    Sonic-3.6 is now generally available in the playground or via the API.

    Try it soon, though. At the rate we’re going, we might have a new one out before you finish integrating.

    Original source
  • Aug 26, 2026
    • Date parsed from source:
      Aug 26, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Speechify logo

    Speechify

    API: simba-3.2 voice cloning is now self-serve for every workspace

    Speechify expands cloned personal voices to synthesize on simba-3.2 for every workspace with no enablement step, ending the limited release and removing the per-workspace allow-list. GET /v1/voices now shows simba-3.2 on cloned voices, while non-English cloned voices still use simba-3.0.

    Cloned (personal) voices synthesize on simba-3.2 for every workspace, with no enablement step. The limited release announced on 2026-08-06 is over and the per-workspace allow-list behind it is gone; you no longer need to contact us.

    Nothing else changes. The request and response are identical to a stock-voice call, and simba-3.2 remains English only, so a cloned voice with a non-English locale still returns 400 — use simba-3.0 for those.

    GET /v1/voices now names simba-3.2 on your cloned voices without any per-workspace condition, and driving a picker off each voice’s models array remains the right pattern. Cloning on simba-3.0, simba-english, and simba-multilingual is unchanged.

    Original source
  • Aug 25, 2026
    • Date parsed from source:
      Aug 25, 2026
    • First seen by Releasebot:
      Sep 1, 2026
    Speechify logo

    Speechify

    API: Projects — group resources, scope credentials, and attribute spend

    Speechify adds Projects in the public API, letting workspaces group resources, scope credentials, and track spend in one place. The new project endpoints are additive and opt-in, with support for managing projects, moving agents, and filtering lists by project.

    API: Projects — group resources, scope credentials, and attribute spend

    The /v1/projects endpoints are now in the public API reference. A project groups the resources you create inside a workspace and the spend you incur from them, so one workspace can run several environments or several end customers without splitting into separate accounts.

    Every workspace has an implicit Default project: any resource with no project lives there, and nothing you already send changes — the surface is additive and opt-in.

    What a project groups:

    Kind Belongs to a project Agents, knowledge bases, tools, audio assets Yes, and can be moved later Phone numbers and SIP trunks Yes API keys and service accounts Yes — a pin fixed when the credential is created Vault credentials and webhook endpoints One project, or workspace-wide Conversations, callers, batch calls, test runs, memories Yes, frozen at creation and never re-attributed Usage and spend Attributed through the calling credential’s pin Cloned voices From the pin on the creating credential — except a consent-verified clone, which is always workspace-wide The public voice and model catalog No — workspace-wide

    Manage the lifecycle with POST/GET/PATCH/DELETE /v1/projects and .../{project_id}, plus archive, unarchive, restore, teardown, stats, audit, promote, and the members sub-tree (grant / revoke access). Move an existing agent with POST /v1/agents/{agent_id}/move.

    Filtering: every list endpoint that takes a project accepts a project_id query parameter. Omit it to get everything you can reach, pass a proj_... id for one project, or pass the literal default for the implicit Default project. On lists whose rows can be workspace-wide — credentials, webhook endpoints and cloned voices — that literal is shared instead, because an absent project there means workspace-wide rather than Default.

    Names are unique per workspace, case-insensitively; a workspace holds at most 100 live projects, and at the cap the create is refused with 409 project_limit_reached.

    A project is a filter and a grouping, not a security boundary — it scopes what a credential may reach, what a scoped member sees, and where spend lands. If one team must be unable to see another’s data at all, use a separate workspace. See Projects.

    Original source
  • Aug 24, 2026
    • Date parsed from source:
      Aug 24, 2026
    • First seen by Releasebot:
      Aug 27, 2026
    Eleven Labs logo

    Eleven Labs

    August 24, 2026

    Eleven Labs releases ElevenLabs CLI v1.0.0 and expands ElevenAgents with generally available procedures, conversation triage tickets, and realtime context usage events. The update also brings broader SDK and widget improvements across JavaScript, Python, Swift, React, React Native, and client packages.

    ElevenLabs CLI

    The ElevenLabs CLI v1.0.0 is now available. Every ElevenLabs API operation is available as a subcommand, with JSON, table, YAML and CSV output, automatic pagination and shell completion.

    The CLI also provides local configuration workflows for ElevenAgents. Store agent, tool and test configurations as files, synchronize them with push and pull commands, manage branches, run tests, install ElevenLabs UI components and select a data residency region.

    Procedures

    Procedures are now generally available in ElevenAgents. A procedure contains task-specific instructions and a trigger that determines when they apply. During a conversation, the agent loads the relevant procedure, allowing one agent to handle distinct tasks without placing every instruction in its system prompt.

    Use free-form procedures when the agent can adapt the wording or order of instructions. Use structured procedures when steps must run in a defined order. Both types can be used alongside workflows on the same agent.

    ElevenAgents

    Conversation triage tickets: Added APIs for creating tickets about agent performance, creating manual follow-up tickets, listing tickets, assigning workspace members, updating status and assignee, and adding ticket-level or turn-level comments.

    Conversation observability: Realtime clients can receive a context_usage event after each completed agent turn. The event reports the model, prompt token count and model context limit.

    SDK Releases

    JavaScript SDK

    v2.65.0 - Added clients and types for conversation triage tickets and lightweight conversation summaries. Conversation APIs add procedure, invalid-tool-call and sort filters; procedure APIs add agent version selection; topic APIs add evaluation detail controls and frustration sorting. The release also adds knowledge base refresh frequency, test environment and alerting integration fields, plus realtime Dubbing message types.

    Python SDK

    v2.65.0 - Added clients and types for conversation triage tickets and lightweight conversation summaries. Conversation APIs add procedure, invalid-tool-call and sort filters; procedure APIs add agent version selection; topic APIs add evaluation detail controls and frustration sorting. The release also adds knowledge base refresh frequency, test environment and alerting integration fields, and preserves repeated multipart fields for Speech to Text keyterms and Dubbing webhook_ids.

    Swift SDK

    v3.3.0 - Added typed client error parsing, expects_response handling and type-safe tool results. WebRTC startup now becomes ready when the agent joins instead of waiting for audio-track subscription. The release also hardens URL handling, suppresses unknown-event noise and removes realtime-thread allocation from software mute processing.

    Packages

    @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected] and @elevenlabs/[email protected] - WebRTC output capture now uses self-hosted AudioWorklet paths under strict content security policies. Worklet caching includes the requested source, preventing an earlier inline URL from replacing a self-hosted path. React Native setup errors now point to @elevenlabs/react-native.

    @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected] and @elevenlabs/[email protected] - Added optional webRtc.iceTransportPolicy. Set it to relay to restrict WebRTC ICE candidates to TURN relays on networks that block direct UDP traffic.

    @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected] and @elevenlabs/[email protected] - Added the onContextUsage callback and ContextUsageEvent type. React exposes the callback through useConversation. React Native now resolves through its package export condition instead of browser-global detection.

    @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected], @elevenlabs/[email protected] and @elevenlabs/[email protected] - Added first-message rich content with attributed button responses, onExternalAgentDisconnected and onMCPToolApprovalRequest. WebRTC now uses the selected input device, preserves mute state when switching microphones and fails setup if initiation data cannot be sent. The widget adds concurrency queue status, optional language selection on the collapsed trigger and first-message rendering for agents that support both text and voice.

    API

    Original source
  • Aug 22, 2026
    • Date parsed from source:
      Aug 22, 2026
    • First seen by Releasebot:
      Aug 23, 2026
    Eleven Labs logo

    Eleven Labs

    August 22, 2026

    Eleven Labs deprecates its local MCP server and moves users to a hosted MCP server with OAuth sign-in.

    MCP server

    Local MCP server deprecated: The local ElevenLabs MCP server and MCP player are deprecated in favor of the hosted MCP server. Both repositories are archived and will no longer receive updates. The hosted server is available at https://api.elevenlabs.io/v1/mcp, signs in with your ElevenLabs account through OAuth and requires no local installation or API key. See the hosted MCP server documentation for setup instructions for Claude, Cursor and other MCP clients.

    Original source
  • August 2026
    • No date parsed from source.
    • First seen by Releasebot:
      Aug 22, 2026
    Speechify logo

    Speechify

    API: Simba 1.6 retired at version 2026-09-21, switched off 2026-11-21

    Speechify announces the retirement of Simba 1.6 models, with simba-english and simba-multilingual removed from selection at API version 2026-09-21 and fully switched off on 2026-11-21. The update guides users to move to Simba 3 and notes API and model-list changes.

    API: Simba 1.6 retired at version 2026-09-21, switched off 2026-11-21

    simba-english and simba-multilingual - the Simba 1.6 pair - are being withdrawn in two steps:

    • From API version 2026-09-21 they are no longer selectable. Naming either returns 400 with the error code model_retired.
    • On 2026-11-21 both models are switched off. From that date they are unreachable on every API version, including a workspace pinned below the retirement.

    If you use either model, you have until 2026-11-21 to migrate, and for most integrations that is a one-line model change. Pinning your workspace’s API version to a date before 2026-09-21 keeps things working in the meantime with no code change at all - but it is a migration window, not an exemption, and it ends on the same day for everyone.

    Why the shorter notice. Our usual sunset window is 12 months. Simba 1.6 is a previous-generation family: it cannot serve /v1/audio/stream/with-timestamps at all and runs at roughly 2.5x the time-to-first-byte of the streaming-native models. Consolidating onto the Simba 3 fleet is what lets us keep improving latency and quality for everyone, and we did not want to spend a year running two stacks to do it.

    Where to go.

    From To Notes simba-english simba-3.2 English. Lower time-to-first-byte, richer expressivity, streaming-native. Serves a curated stock roster plus your own cloned voices. simba-english simba-3.0 English, if you need a voice outside the simba-3.2 roster. Accepts every catalog voice. simba-multilingual simba-3.0 English, de-DE, es-ES, es-MX, fr-FR, it-IT, pt-BR. Streaming-native, and a cloned voice speaks all of them from one voice ID.

    An omitted model is unaffected: it already resolves to simba-3.0.

    If you synthesize outside Simba 3.0’s seven locales, simba-3.0 on its own is not a like-for-like replacement - and we are not asking you to drop those languages. Broader multilingual coverage on the current model generation is planned to be available before this date, and we will confirm the model and the timing directly rather than leave you to read it off a changelog. Talk to us so we can line your migration up with it.

    If that coverage is not in your hands by 2026-11-21, we move the shutdown date rather than cut the languages off. That is the commitment; the date is not.

    What changed at this version.

    POST /v1/audio/speech and both /v1/audio/stream routes reject the two ids. GET /v1/audio/models returns only the models you can actually call, and each voice’s models array in GET /v1/voices does the same - so a picker driven off either endpoint stays correct without special-casing.

    Read the dates off the API. While your workspace is pinned below 2026-09-21, both models still appear in GET /v1/audio/models carrying retired_at: "2026-09-21" and sunset_at: "2026-11-21". Surface sunset_at in your own tooling if you have a deadline to track.

    See the API Versioning guide for how to read and set your workspace’s pinned version.

    Original source
  • August 2026
    • No date parsed from source.
    • First seen by Releasebot:
      Aug 22, 2026
    Wispr Flow logo

    Wispr Flow

    A Sign-In Flow That Actually Works

    Wispr Flow improves Android sign-in and reporting with smoother, more reliable flows. Google sign-in now opens in a secure in-app browser tab and returns users straight to Flow, while interrupted logins can be resumed. Issue, feedback, and transcription reports also keep sending in the background.

    Signing in with Google on Android now happens in a secure in-app browser tab that returns you straight to Wispr Flow. No more landing on a random browser tab, an old page, or the Wispr Flow website after picking your account. If your sign-in gets interrupted, a resume prompt lets you pick up right where you left off.

    Find it: automatic, the sign-in screen the first time you open Flow or sign back in.

    Help Center: Login Issues with Wispr Flow

    Reports That Never Get Lost

    Issue reports, feedback, and transcription reports now keep sending even if you leave the screen. You'll get a confirmation wherever you are in the app, and the form clearly shows when it's busy: a spinner, a dimmed Send button, and a short note explaining why it's temporarily locked.

    Find it: Dashboard menu > Report an issue or Share feedback, and Home > any transcript > Report.

    Original source