Hume Release Notes
45 release notes curated from 57 sources by the Releasebot Team. Last updated: Sep 11, 2026
- Sep 10, 2026
- Date parsed from source:Sep 10, 2026
- First seen by Releasebot:Sep 11, 2026
Introducing the Hume Voice Replication Leaderboard
Hume launches the Voice Replication Leaderboard as part of its Real World VoiceEQ Benchmark, giving teams a live way to compare text-to-speech models on speaker identity, naturalness, audio quality, and accent and emotion handling.
A new VoiceEQ benchmark for voice replication
Today we are publishing the Voice Replication Leaderboard, a new component of Hume’s Real World VoiceEQ Benchmark. It measures how well leading text-to-speech models reproduce a specific voice across varied real-world conditions.
Like Real-World VoiceEQ, the leaderboard combines human judgments with objective measures rather than reducing performance to a single score.
How we evaluated voice replication
We evaluated eleven leading models, asking each to reproduce the same 25 reference voices. To better reflect real-world use, the reference set spans three categories designed to stress different aspects of replication:
- Five standard references drawn from archival recordings of public figures and conversational US speech, not specifically selected for accent or emotional variation;
- Five expressive references spanning anger, fear, happy, sad, and calm
- 15 references selected for accent variation, split into 6 native English varieties and 9 non-native accents, drawn from a mix of read and conversational speech.
Every tested model was asked to replicate these reference voices with the same seven prompts, written to avoid prescribing a particular emotion or delivery style. Each generated clip was evaluated across four dimensions:
Same speaker: A 1–5 human rating of whether the generated clip sounded like the reference speaker.
Quality: A 1–5 human rating of whether the audio was clean and free of artifacts, independent of speaker identity.
Naturalness: A 1–5 human rating of how human the delivery sounded, including pacing and prosody, independent of speaker identity.
Objective similarity: The cosine similarity between the TitaNet speaker embeddings of the reference and generated clips, scored from 0 to 1.
For the three human-rated dimensions, each clip was independently assessed by three paid raters via the Hume Human Feedback API. Raters were blind to the model that produced it and heard the reference recording first, followed by the generated clip. When we report a model's score across all references, we weight the three reference categories equally, one third each, rather than averaging over samples.
Results
Overall, no model was strongest across every dimension or reference type. Performance depended both on what was measured and on the voice being replicated. The results reinforce a core principle of VoiceEQ: voice systems should be evaluated as a profile of capabilities, not reduced to a single score.
Voice Replication Leaderboard
Human raters were asked to score 1–5 of whether a clip generated by a TTS model sounded like the reference speaker provided. Best value per column in accent color. Hover a score for the full row.
Ranked by same speaker. No model leads every column: identity, quality, naturalness and objective similarity are led by four different models.
hume.ai/rw-voice-eq
Sounding natural is not the same as sounding like the right person
A voice can sound convincingly human without sounding like the person it is intended to replicate. Cartesia sonic-3.6-beta received the highest naturalness score (4.36), but ranked eighth on same-speaker similarity (3.63). OpenBMB VoxCPM2 ranked first on same-speaker similarity (4.21), but placed fifth on naturalness (3.96).
This is not evidence of an unavoidable trade-off; for example some models perform well on both. But it shows why the dimensions must be measured separately.
Performance changes with the voice being replicated
Replicating speech from expressive reference voices created the wide gap in performance: Higgs Audio v3 scored 3.38 on emotional references, while VoxCPM2 scored 4.38 on the same category. However no model performed equally well across standard, expressive, and accented references. Notably, Higgs Audio v3 tied with VoxCPM2 for the lead on accented voices, but also had the widest spread across the three reference categories.
Same-speaker score by reference type
Human rating out of 5. The 25 reference voices split into 5 standard, 5 emotional and 15 accented, and the accented set splits again into 6 native English varieties and 9 non-native accents. Darker is higher. Spread is best category minus worst. Best value per column in accent.
The two models at the top score higher on non-native accents than native. Every other model but one does the reverse. Higgs Audio v3 leads accents at 4.28 and drops to 3.38 on emotional references, the widest spread on the board.
hume.ai/rw-voice-eq
Clean audio is increasingly becoming table stakes among the models tested
Quality scores were high and tightly grouped, ranging from 3.85 to 4.61, with an average of approximately 4.30. Same-speaker ratings varied more widely, from 2.91 to 4.21. For example, Inworld TTS-2 received the highest quality score (4.61), but ranked ninth on same-speaker similarity (3.62).
Objective and human-rated similarity measures are complementary, not interchangeable
The objective embedding score broadly supported what human listeners heard but the rankings diverged in meaningful cases. VoxCPM2 and LongCat took the top two places under both approaches, in opposite order, and Fish Audio s2-pro and Higgs Audio v3 held third and fourth under both. ElevenLabs Eleven v3 ranked last under both.
Cartesia sonic-3.6-beta ranked eighth by human same-speaker rating and fifth by objective similarity. Speaker embeddings provide a useful, scalable signal, but they do not capture everything people hear when deciding whether two clips sound like the same person. For evaluation and model improvement, objective measures and human judgments are most useful together.
What this means for teams building and choosing voice models
For product teams: test the voice you actually need - given there is no single model that is best across speaker identity, naturalness, quality, and every reference type. The right choice depends on the application. A model suited to neutral narration may not be the best fit for emotionally varied customer service; one that performs well for one speaker may be less reliable for speakers with different accents.
Teams should test models using voices, prompts, languages, emotional conditions, and recording environments that resemble their actual deployment rather than relying on a general score alone.
For model developers: separate the failure modes to drive improvement.
Improving voice replication requires measuring identity, naturalness, quality, and objective similarity independently. Researchers should break results down by reference type and, where possible, by individual emotion and accent. This helps determine whether a weakness reflects emotional speech broadly, a particular expressive condition, or a specific type of reference recording.
For example, our leaderboard suggests that emotionally expressive speech remains a challenge. To diagnose where the failure occurs, researchers can test two conditions separately (i) replicating a voice from an emotionally expressive reference while keeping the target text neutral, and (ii) asking the model to deliver emotionally charged text from a neutral reference. Comparing same-speaker and naturalness scores across these conditions can help distinguish identity extraction from expressive generation. If speaker identity remains stable but naturalness falls when emotional delivery is required, the weakness is likely in how the model renders expressive speech. If identity falls when the reference itself is emotional, even with neutral target text, the model may be struggling to recover a consistent representation of the speaker from expressive audio.
Explore the leaderboard
The Voice Replication Leaderboard is live on Hugging Face, with per-category breakdowns and curated reference and replica audio samples. It extends the Real World VoiceEQ Benchmark’s effort to evaluate voice AI the way people actually experience it.
Public leaderboards provide a shared baseline. Hume also works with frontier labs and product teams to design evaluations around their own speakers, languages, emotional contexts, and production failure modes—combining human ratings and objective measures to identify weaknesses and generate the preference data needed to improve them. If you’re evaluating a model or developing one yourself, get in touch.
Original source - Aug 21, 2026
- Date parsed from source:Aug 21, 2026
- First seen by Releasebot:Sep 11, 2026
Measuring benchmark optimization in speech recognition
Hume adds a Benchmark fitting tab to the Open ASR Leaderboard and new tests to expose benchmark optimization in speech models. The update highlights reference errors, masked numbers, and orthographic switching across public and held-out audio to better reflect real-world transcription quality.
Public voice AI benchmarks increasingly suggest that models are performing at human levels.
Yet those scores don't always reflect how models work in the real-world. Since public benchmarks are open and widely used, models can also become optimized for the tests themselves. Their scores may improve because they have learned benchmark-specific patterns and not because they have become better at the underlying task.
One reason is that traditional benchmarks overlook many of the conditions and qualities that make voice systems reliable, natural, contextually appropriate, and effective in practice. That's why we recently introduced held-out sets in Real World VoiceEQ, the Open-ASR Leaderboard, and the Far-field ASR Leaderboard: to measure more of what matters in real-world use.
However, broader measurement alone does not solve the problem. This phenomenon, sometimes called benchmark optimization or "benchmaxxing," is often discussed around machine learning, however, it has been difficult to measure in speech recognition.
Our latest research introduces three tests to help quantify it. We evaluated 11 widely used open-source ASR models and found that several of the highest-scoring systems reproduced benchmark transcripts from the VoxPopuli English and LibriSpeech (clean, other) datasets – even when the audio contradicted them, relevant words had been silenced, or the audio equally supported two different written forms.
In some cases, models appeared to rely not only on what was said, but also on subtle acoustic cues that indicated which benchmark they were being tested on. As a result, their scores overstated how well they could transcribe speech more generally.
Reference disagreement (VoxPopuli case study)
VoxPopuli is known to contain a high number of transcription errors (which is why Artificial Analysis released a cleaned version). Our consensus disagreement probe tests what happens when leading ASR models encounter these errors: Do they accurately transcribe what the audio says, or reproduce the benchmark's incorrect reference transcript?
To test this at scale, we use an ensemble of independent models selected for their low phoneme error rate (PER). PER measures how closely a written transcription matches the sounds in the audio, making it a useful proxy for how faithfully a model transcribes what it hears. The ensemble results can be used to flag cases in which the models unanimously disagree with the benchmark's reference transcript. We then compare a sample of those flagged cases against human annotations to validate the corrected transcripts.
For example, one VoxPopuli clip audibly includes the phrase "Thank you, Mr. President," but the reference transcript omits "Thank you." Six of the 11 models we tested reproduced the benchmark's erroneous transcript—giving the "expected" answer even though it contradicted the audio. On the real clip, the formatting follows the same pattern: models that omit "Thank you" also reproduce the benchmark's punctuation style, writing "Mr" without a period, while models that include the audible phrase tend to write "Mr." with the period.
When we present the same content in newly collected voices from EU parliamentary recordings or generic voices, this behavior often weakens or disappears. In the below samples, all but one model flips back to transcribing the audio-faithful transcript for a clone of a new parliamentary recording. This suggests that the models are responding to acoustic cues that help them identify the benchmark membership and thus produce the expected transcript even if it contradicts the audio.
The reference transcript for this clip reads "Mr President, I have another complaint about this procedure, which is that it is not secret." The audio in all three clips below actually says the same thing, preceded by an audible "Thank you,"—the clones are text-to-speech renditions of that true sentence, so the courtesy is audible in all three. Green highlighting and ✅ mark a transcript that includes the audible "Thank you"; red highlighting and ❌ mark a transcript that reproduces the benchmark's erroneous omission. All transcripts are raw model output, prior to any normalization—casing and punctuation are preserved exactly as generated, including lowercase output from some models.
Parakeet is the only model that flips between reproducing the benchmark on the real clip and getting it right on the same-speaker clone. Phi-4 is the only model still dropping the courtesy on the ep-fresh clone. When we instead resynthesize the sentence in a generic TTS voice unconnected to any parliamentary recording, all eleven models restore the courtesy.
The results suggest that this problem is both widespread and meaningful. Our methodology flagged potential reference errors in 40% of the VoxPopuli test clips we analyzed, affecting roughly 3% of all reference words.
Models exhibiting benchmark-optimized behavior reproduced erroneous reference transcripts 18–30% of the time. The scatterplot below compares VoxPopuli word error rate (WER) on the x-axis with the rate at which each model reproduces the benchmark's incorrect reference instead of the consensus correction. The models with the lowest WER—and therefore the strongest reported benchmark performance—are also the most likely to reproduce these errors.
Masked Entity Retrieval
To build on the consensus disagreement probe, we deliberately silence numbers in the audio samples of test datasets and ask the models to transcribe what it hears. The number is literally absent from the audio, so models should not output any number, much less the exact number in the text.
Some of these numbers are semi-predictable (although still unlikely for a model to predict), yet others are quite surprising. The following clip combines both probes, showing both how models recreate reference transcript errors including an incorrect number and one model even autocompletes a relatively random year (2011) despite it being silenced. In each model's row below:
- green highlighting with strikethrough marks reference-transcript words the model correctly did not reproduce (audio-faithful);
- green highlighting with underline marks a correct, audio-faithful insertion in place of the reference's erroneous wording;
- red highlighting (plain text) reproduces the reference transcript's erroneous, audio-unsupported content: keeping "Mr President", writing "more than 1 amendments" where the audio says "one thousand six hundred", supplying the silenced year "2011", or ending on "plenary".
Recovery rates were highest on the public benchmarks and lower on held-out or newly collected audio (ep-fresh and libri-fresh below). On LibriSpeech, some of the strongest benchmark-performing models reproduced masked numbers in roughly 30–40% of examples, even though the number itself had been removed. The effect weakened on freshly collected data for several models, suggesting that the surrounding benchmark-associated audio—not only textual autocomplete—helped the models recover the reference.
Orthographic Switching
Our orthographic switching probe tests whether models reproduce the exact spelling used in a benchmark's reference transcript despite it not being clear in the audio. Orthographic variants are words that are semantically and phonetically identical but can be spelled different ways (1 vs one, Mr. vs mister, John vs Jon, Honor vs Honour, etc). In theory, models should consistently prefer one spelling over another, or alternate between them at roughly random rates. If models systematically switch to match what is in each benchmark's reference transcript, that suggests the models are picking up on which spelling the test expects.
Transcription: "I URGED ON THE BOYS THAT WHATEVER HAPPENED WE SHOULD NOT SHOOT ANY ONE" — models using "any one": 6/11, models using "anyone": 5/11
Transcription: "CAMOUFLAGE WAS NOT A WORD THE CAPTAIN OR ANYONE ELSE OF HIS TIME YET UNDERSTOOD" — models using "any one": 2/11, models using "anyone": 9/11
Within LibriSpeech, we test one intra-dataset switch involving an older spacing convention: some reference transcripts use "any one", while others use "anyone." We measure the minimum accuracy for a given variant, which we call "switch rate". If a model only uses one variant it would have a 0% switch rate; a model which picks randomly would be expected to have a 50% switch rate. A model which knows which variant to use in every test sample would earn a 100% switch rate.
Our second probe tests an inter-dataset switch, in which each benchmark uses a different spelling convention consistently across its test corpus. For example, VoxPopuli uses the abbreviation "Mr.", while LibriSpeech spells out "Mister."
Multiple models exceed the 50% random-choice baseline, with some reaching roughly 90% switch accuracy. This suggests that the models can identify which dataset an audio sample comes from and select the spelling convention that benchmark expects, even though both forms sound identical.
Localizing the switches
To test whether these behaviors generalize beyond the public benchmarks, we also collected fresh data from the same source domains but after the models' training cutoffs: recent European Parliament recordings for VoxPopuli and recordings from newly active LibriVox narrators for LibriSpeech. However, when presented with recently collected data from the same domain, many models stop matching the reference transcript and revert to more audio faithful transcriptions.
Other interventions point to the same conclusion. Phrases which are present in the audio but are omitted in the reference transcript can reappear when a model is asked to translate the audio or when its attention is restricted to the relevant frames. Trimming away surrounding benchmark context, or appending ordinary conversational audio, can also restore the faithful transcript. Appending VoxPopuli audio can have the opposite effect, making otherwise faithful synthetic or mined samples more likely to match the benchmark reference.
Together, these results suggest that models are able to faithfully transcribe the literal spoken words, but are using surrounding acoustic context to decide whether to follow the audio or a benchmark-specific transcription policy.
Conclusion
Our findings suggest that, on two major open-source datasets, some models detect dataset-associated acoustic cues and adjust their transcription behavior accordingly. Specifically, models may reproduce words that are absent from the audio but present in the reference transcript, recover silenced numbers at elevated rates, or use surrounding acoustic context to select the written variant expected by a particular benchmark.
For people selecting models, these findings underscore the importance of using fully held-out evaluation sets, as RW-Voice-EQ Bench and the Open ASR Leaderboard do, and of looking beyond word error rate on a single public benchmark. To this end, a "Benchmark fitting" tab has been added to the Open ASR Leaderboard, which includes two of the above analyses across all models: quantifying (1) reference error rates from VoxPopuli and (2) orthographic switching across all public datasets. The relevant scripts are open-sourced on GitHub as well as the un-normalized model outputs.
Our findings also suggest that benchmark developers should avoid simple independent and identically distributed test splits in favor of temporal, speaker, or other metadata-based separation. Greater transparency around training data and model-selection procedures would also help researchers understand how these behaviors arise.
Public benchmarks remain valuable: they are transparent, repeatable, easy to run, and well understood by the research community. But they are most useful when we can distinguish genuine transcription improvements from benchmark-specific gains that do not generalize to new audio.
For more information, we encourage you to read our full report.
Original source All of your release notes in one feed
Join Releasebot and get updates from Hume and hundreds of other software products.
- May 2026
- No date parsed from source.
- First seen by Releasebot:May 16, 2026
TTS API bug fixes
Hume fixes a bug that caused duplicate interleaved TTS audio and distorted output.
Fixed a bug where duplicate interleaved audio was included in TTS audio output.
This resolves an issue where audio chunks could be duplicated and interleaved, resulting in distorted output.
Original source - May 15, 2026
- Date parsed from source:May 15, 2026
- First seen by Releasebot:May 16, 2026
May 15, 2026
Hume adds an experimental temperature parameter to its TTS API for more varied or consistent speech generation.
TTS API additions
Added an experimental temperature parameter to TTS endpoints. Controls sampling temperature for speech generation. Higher values increase variation; lower values increase consistency.
Original source - Apr 10, 2026
- Date parsed from source:Apr 10, 2026
- First seen by Releasebot:Apr 11, 2026
April 10, 2026
Hume adds configurable turn detection and interruption settings to EVI configs, giving users finer control over turn-taking, speech detection, and interruption behavior on a per-config basis.
EVI API additions
Added configurable turn detection and interruption settings to EVI configs. You can now control how EVI handles turn-taking and interruptions on a per-config basis.
- turn_detection.end_of_turn_silence_ms: How long EVI waits after speech ends before committing a turn (500-3000ms, default 800ms).
- turn_detection.speech_detection_threshold: Sensitivity of voice activity detection (0.0-1.0, default 0.5).
- turn_detection.prefix_padding_ms: Audio padding before detected speech (default 300ms).
- interruption.min_interruption_ms: Minimum speech duration before EVI can be interrupted (50-2000ms, default 800ms).
Similar to Hume with recent updates:
- xAI release notes269 release notes · Latest Oct 2, 2026
- Anthropic release notes853 release notes · Latest Oct 4, 2026
- Cursor release notes137 release notes · Latest Sep 23, 2026
- Eleven Labs release notes99 release notes · Latest Sep 28, 2026
- Perplexity release notes31 release notes · Latest Sep 21, 2026
- Canva release notes41 release notes · Latest Jun 24, 2026
- Mar 10, 2026
- Date parsed from source:Mar 10, 2026
- First seen by Releasebot:Mar 10, 2026
Opensourcing TADA: Fast, Reliable Speech Generation Through Text-Acoustic Synchronization
Hume releases TADA, a groundbreaking Text-Acoustic Dual Alignment for fast, low-hallucination LLM based TTS. Open-sourced now with 1B and 3B models and full audio tokenizer, enabling on-device deployment and real-time voice with one-to-one text audio mapping. Available on HuggingFace and GitHub for research and development.
Approach
The future of voice AI hinges on sounding natural, fast, expressive, and free of quirks like hallucinated words or skipped content. Today's LLM-based TTS systems are forced to choose between speed, quality, and reliability because of a fundamental mismatch between how text and audio are represented inside language models.
TADA (Text-Acoustic Dual Alignment) resolves that mismatch with a novel tokenization schema that synchronizes text and speech one-to-one. The result: the fastest LLM-based TTS system available, with competitive voice quality, virtually zero content hallucinations, and a footprint light enough for on-device deployment.
Hume AI is open-sourcing TADA to accelerate progress toward efficient, reliable voice generation. Code and pre-trained models are available now.
For input audio, an encoder paired with an aligner extracts acoustic features from the audio segment corresponding to each text token. For output audio, the LLM's final hidden state serves as a conditioning vector for a flow-matching head, which generates acoustic features that are then decoded into audio and fed back into the model.
Since each LLM step corresponds to exactly one text token and one acoustic representation, TADA generates speech faster and with less computational effort. And because the architecture enforces a strict one-to-one mapping between text and audio, the model cannot skip or hallucinate content by construction.
Evaluation
Hallucination Rate
SAMPLES WITH CER > 0.15, SIGNALING SKIPPED WORDS, INSERTED CONTENT, OR UNINTELLIGIBLE SPEECH. OUT OF 1,088 SAMPLES.
- FireRedTTS-2 41
- Higgs Audio v2 24
- VibeVoice 1.5B 17
- TADA-3B 0
- TADA-1B 0
Speed
TADA generates speech at a real-time factor (RTF) of 0.09 — more than 5x faster than similar grade LLM-based TTS systems. This is possible because TADA operates at just 2–3 frames (tokens) per second of audio, compared to 12.5–75 tokens per second in other approaches.
Hallucination
Our model was trained on large scale, in-the-wild data, without post-training, and achieves the same reliability as models trained on smaller curated datasets. We measured hallucination rate by flagging any sample with a character error rate (CER) above 0.15 — a threshold that captures unintelligible speech, skipped text, and inserted content. In the 1000+ test samples from LibriTTSR, TADA produced zero hallucinations.
Voice Quality
Results on SEED-TTS-EVAL and LIBRITTSR-EVAL show that TADA achieves reliability comparable to Index-TTS — one of the few systems with similarly low hallucination rates — while being trained on a larger, less-curated dataset.
In human evaluation on expressive, long-form speech (EARS dataset), TADA scored 4.18/5.0 on speaker similarity and 3.78/5.0 on naturalness, placing second overall — ahead of several systems trained on significantly more data.
Potential Applications
On-device deployment. TADA is lightweight enough to run on mobile phones and edge devices without requiring cloud inference. For device manufacturers and app developers building voice interfaces, this means lower latency, better privacy, and no API dependency.
Long-form and conversational speech. TADA's synchronous tokenization is dramatically more context-efficient than existing approaches. Where a conventional system exhausts a 2048-token context window in about 70 seconds of audio, TADA can accommodate roughly 700 seconds in the same budget. This opens the door to long-form narration, extended dialogue, and multi-turn voice interactions.
Production reliability. Zero hallucinations in our tests suggests fewer edge cases to catch, fewer customer complaints, and less post-processing overhead in the product. This makes TADA well-suited for deploying voice in regulated or sensitive environments like healthcare, finance, and education.
Limitations and Future Work
Long-form degradation. While the model supports more than 10 minutes of context, we noticed occasional cases of speaker drift during long generations. Our online rejection sampling strategy reduces this significantly, but it's not fully resolved. We suggest resetting the context as an intermediate workaround.
The modality gap. When the model generates text alongside speech, language quality drops relative to text-only mode. We introduce Speech Free Guidance (SFG), a technique that blends logits from text-only and text-speech inference modes to help close this gap, but more work is required.
Use-cases. The model is only pre-trained on speech continuation; further fine-tuning is required for assistant scenarios. Get in touch to inquire about Hume's extensive library of fine-tuning data.
Scale. The current release covers English and seven additional languages, so there's clear room to expand. We're training larger models with broader language coverage with Hume AI data.
We're releasing TADA because we believe this architecture opens a productive direction for the field, and we want to accelerate progress. We invite researchers and developers to build on this work — whether that means extending the tokenizer to new modalities, solving the long-context problem, or adapting the framework for new applications.
Get Started
TADA is available now under an open-source license. We're releasing 1B and 3B parameter Llama-based models and the full audio tokenizer and decoder.
1B (English):
huggingface.co/HumeAI/tada-1b
3B (multilingual):
huggingface.co/HumeAI/tada-3b-ml
Demo:
huggingface.co/spaces/HumeAI/tada
GitHub:
github.com/HumeAI/tada
TADA was developed by Trung Dang, Sharath Rao, Ananya Gupta, Christopher Gagne, Panagiotis Tzirakis, Alice Baird, Jakub Piotr Cłapa, Peter Chin, and Alan Cowen at Hume AI.
Hume builds voice AI research infrastructure for frontier labs and AI-first enterprises. If you're working on voice models and need high-quality training data, evaluation systems, or reinforcement learning infrastructure, get in touch at
Original source - Feb 27, 2026
- Date parsed from source:Feb 27, 2026
- First seen by Releasebot:Feb 28, 2026
- Modified by Releasebot:May 16, 2026
February 27, 2026
Hume adds EVI API support for new LLM models and zero prompt expansion control.
EVI API additions
Added support for new supplemental LLM models: claude-opus-4-6, gpt-5.1, gpt-5.1-priority, gpt-5.2, gpt-5.2-priority.
Added support for zero prompt expansion. You can now set prompt_expansion to ZERO when configuring an external LLM, disabling automatic prompt expansion and giving you full control over the system prompt.
Original source - December 2025
- No date parsed from source.
- First seen by Releasebot:Dec 23, 2025
EVI API improvements
Added support for OpenAI’s GPT-5, GPT-5-mini, and GPT-5-nano models as a supplemental LLM options.
Original source - December 2025
- No date parsed from source.
- First seen by Releasebot:Dec 23, 2025
EVI API improvements
New supplemental LLMs are now supported for all EVI versions:
- Claude Sonnet 4 (Anthropic)
- Llama 4 Maverick (SambaNova)
- Qwen3 32B (SambaNova)
- DeepSeek R1-Distill (Llama 3.3 70B Instruct, via SambaNova)
- Kimi K2 (Groq)
- Nov 14, 2025
- Date parsed from source:Nov 14, 2025
- First seen by Releasebot:Dec 23, 2025
November 14, 2025
EVI API additions
Added support for a new SESSION_SETTINGS chat event in the EVI chat history API. When you fetch chat events via /v0/evi/chats/:id, the response now includes entries that indicate when system settings were updated and which settings were applied.
Original source - Nov 7, 2025
- Date parsed from source:Nov 7, 2025
- First seen by Releasebot:Dec 23, 2025
November 7, 2025
Voice conversion and secure EVI enhancements expand the platform with new endpoints to convert speech to a target voice, manage active chats over WebSocket, and handle tool calls via a webhook. This update signals practical, user-facing API improvements.
TTS API additions
Introduced voice conversion endpoints. Send speech, specify a voice, and receive audio converted to that target voice.
- POST /v0/tts/voice_conversion/json: JSON response with audio and metadata.
- POST /v0/tts/voice_conversion/file: Audio file response.
EVI API additions
Added a control plane API for EVI. Perform secure server-side actions and connect to active chats. See the Control Plane guide.
- POST /v0/evi/chat/:chat_id/send: Send a message to an active chat.
- WSS /v0/evi/chat/:chat_id/connect: Connect to an active chat over WebSocket.
Added a tool call webhook event. Subscribe to tool calls to know when to invoke your tool, then send the tool response back to the chat using the control plane.
Original source - Oct 24, 2025
- Date parsed from source:Oct 24, 2025
- First seen by Releasebot:Feb 18, 2026
AudioStack × Hume: Professional Audio for Creatives
AudioStack expands its AI audio production suite with Hume’s expressive voices, boosting speed, consistency, and emotional depth for ads, podcasts, and branded content. The integration enables scalable, natural sounding output across markets and languages while reducing costs and raising creative quality.
About AudioStack
AudioStack is an enterprise AI audio production platform trusted by global creative teams at Publicis, Omnicom, iHeartMedia, Dentsu, and more. Their AI-driven production suite empowers agencies, publishers, AdTech platforms, and brands to create broadcast-ready audio content 10 times faster at a fraction of traditional costs—reducing production expenses by up to 80% while scaling effortlessly across markets and languages.
Expanding Audiostack’s Voice Library with Hume
AudioStack offers a comprehensive voice library for audio advertisements, podcasts, and branded content. As they continue to grow their voice offerings, they're integrating Hume's emotionally intelligent voices to meet two core demands of creative teams:
Consistent Stability
Enterprise content generation requires voices that perform reliably across thousands of productions. Hume's voices deliver consistent quality and pronunciation, ensuring brand messaging remains clear and professional, whether creating one ad or thousands of dynamic variations.Natural Expressiveness
Generic TTS voices often sound flat or robotic—a dealbreaker for agencies creating audio that needs to engage audiences. Hume's voices bring genuine emotional depth, helping audio content feel authentic, engaging, and human.
By adding Hume’s expressive voices to their platform, AudioStack enables creative teams and advertisers to produce high-quality, emotionally resonant audio at scale.
For more information on how empathic AI can enhance your digital solutions, contact Hume AI.
Original source - Oct 21, 2025
- Date parsed from source:Oct 21, 2025
- First seen by Releasebot:Feb 18, 2026
Creating immersive avatar experiences with Render Foundry
Render Foundry unveils an immersive Babe Ruth simulator built with Hume AI voice cloning, blending Unreal Engine storytelling with authentic, warm dialogue. The experience lets visitors talk with Babe Ruth, syncing likeness with audio for a lifelike museum interaction.
About Render Foundry
Render Foundry specializes in creating immersive experiences, from interactive museum installations to digital twins of entire campuses. Led by Shane Boyce and Josh Harwell, their team combines Unreal Engine expertise with cutting-edge storytelling to blur the line between reality and simulation.
When they set out to create an interactive Babe Ruth experience, they needed a voice that could do the impossible: bring him back to life with authenticity, warmth, and the personality that made him a legend.
Watch the Experience
Using Hume's custom voice cloning technology, Render Foundry created a Babe Ruth simulator that feels emotionally authentic. Hume captured the tonal qualities, cadence, and personality of the baseball icon, allowing visitors to have natural, engaging conversations with one of sports' most beloved figures.
The result is an experience that transcends typical museum exhibits. Visitors not only learn about Babe Ruth, but also connect with him. Render Foundry created something truly special by simulating Babe’s likeness and syncing the audio and the visual, making this experience one-of-a-kind.
Josh Harwell, the Creative Director at Render Foundry, says,
“We’re excited to offer these curated experiences. It’s fun to watch clients interact with our characters as they are brought to life. Whether it’s a historical figure, a mascot, or a brand ambassador, Hume helps us deliver a solution that humanizes the responses of AI.”
For more information on how empathic AI can enhance your digital solutions, contact Hume AI.
Original source - Oct 14, 2025
- Date parsed from source:Oct 14, 2025
- First seen by Releasebot:Feb 18, 2026
Revelum × Hume: Detecting Voice Fraud in Real-Time
Revelum launches an AI-native security platform that stops deepfake fraud in real time with call risk analysis, live deepfake detection, and precise timestamps. A strategic partnership with Hume AI speeds up resilient detection and responsible AI use, demonstrated by a real-time fraud scenario.
Revelum
Revelum is an AI-native security platform that protects institutions from deepfake impersonations and fraud in real time. Founded by Enrique Barco in 2025, Revelum provides turn-key solutions and developer-friendly APIs to safeguard institutions from emerging threats in the era of AI-driven fraud.
Revelum’s strategic partnership with Hume AI ensures their detection systems can identify even the most advanced synthetic voices, including Hume's own Empathic Voice Interface (EVI).
The Demo: AI-Powered Fraud in Action
In a recent demo, Revelum showcased how attackers use AI voice agents to attempt account takeovers. The scenario: "Jake" calls customer support, claiming an AI assistant accidentally changed and deleted his password. The EVI-powered voice sounds natural, but Revelum's technology instantly flags the deepfake.
In real-time, Revelum’s platform provides:
- A call risk assessment analysis
- Real-time deepfake detection counts
- Precise timestamps of synthetic voice segments
The customer service agent receives an alert and initiates callback verification—stopping the attack immediately.
Hume’s Partnership with Revelum
Our collaboration with Revelum creates a critical feedback loop for responsible AI development:
- Early Access to EVI: Revelum trains their models on Hume's cutting-edge voice technology, ensuring detection capabilities stay ahead of emerging threats before they reach malicious actors.
- Continuous Refinement: As Hume's emotionally intelligent voices become more sophisticated, Revelum's detection algorithms evolve in parallel.
Revelum founder Enrique notes:
"By partnering with Hume, we’re taking a vital step toward building technology that anticipates — not just reacts to — the evolving tactics of bad actors seeking to misuse powerful models. Together, we’re staying one step ahead in ensuring generative AI is used responsibly."
For more information on how empathic AI can enhance your digital solutions, contact Hume AI.
Original source - Oct 3, 2025
- Date parsed from source:Oct 3, 2025
- First seen by Releasebot:Dec 23, 2025
October 3, 2025
Octave 2 upgrades delivered across TTS and EVI: HTTP and WebSocket endpoints now support Octave 2 with timestamps at word and phoneme levels, plus EVI 4-mini enables multilingual TTS via Octave 2 with your chosen LLM.
TTS API improvements
Octave 2 is now available for use in TTS endpoints.
For HTTP endpoints, specify "version": "2" in your request body. (reference)
For the WebSocket endpoint, specify version=2 in the query parameters of the handshake request. (reference)Word and phoneme level timestamps are now supported for TTS APIs. See our Timestamps Guide to learn more.
EVI API improvements
EVI version 4-mini is now available. This version enables the use of Octave 2 for TTS alongside a supplemental LLM of your choosing, bringing Octave 2’s multilingual capabilities to EVI. Specify "version": "4-mini" in your EVI config version to use it.
Original source
Curated by the Releasebot team
Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.
Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.