Events
Voice AI Space “Global Night” mixer across five cities, a no-talks, no-demos networking night for founders, engineers, and researchers (Aug 20, SF + London + Paris + Barcelona + Toronto, Voice AI Space)
Dex and OpenAI dig into production voice AI engineering at a London meetup, with talks on turn-taking, interruption handling, and latency from speakers at each company, followed by networking. (Aug 19, London, LUMA)
Top Updates 💪
NVIDIA releases NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model (MarkTechPost)
Krisp launches Voice Isolation v2.5 that improves STT word error rates (46% drop) in the most challenging multi-speaker cases (Krisp Blog)
Deepgram launches Flux TTS, a conversation-native text-to-speech model that tracks tone across turns and starts output in as low as 80ms. (Deepgram)
Alibaba launches CosyVoice Studio, China’s first full-stack AI voice productivity platform built on its benchmark-topping Qwen-Audio model. (Pandaily)
Soniox ships TTS v2, one model for narration, creative work, and voice agents with cloning and expressive control across 60+ languages. (Techgenyz)
Omilia raises $67M Series B to expand its voice-first agentic CX platform, having grown live ARR more than 10x to over $60M. (Slator)
Apple is in talks to pay publishers up to nine figures to feed current news into its revamped AI-powered Siri via a pay-per-query model. (WSJ)
Regal ties its voice AI agents into Five9, letting call events trigger AI follow-up calls and texts through Five9’s AI Agent Connect program. (SiliconAngle)
Sarvam partners with HP India to bring its Kivi voice assistant to laptops, letting users dictate and act across apps in Indian languages. (Sarvam)
Amazon’s retail chief says conversational AI will reshape shopping, comparing the shift to the move from printed catalogues to search bars. (IT Brief)
Assort Health unveils Referrals, an AI agent that processes and triages every patient referral via voice, text, and email, automating ~90% end-to-end. (PR Newswire)
India’s IISc releases SraVaani, the first multilingual Indian speech model trained on 65 languages and dialects, open-sourced under MIT. (The Hindu)
Zoom combines communications, AI, and workflow into one agentic CX platform, turning calls and meetings into completed business outcomes. (Current Analysis)
Bose bets on licensing and tiny AI, with CEO Lila Snyder detailing a pivot to software and wearables on The Verge’s Decoder podcast. (The Verge)
Resemble AI goes all in on deepfake detection, halting new voice AI sales after its report found 15,700+ people victimized in six months. (Resemble AI)
Engineering Corner 😎
Sierra shows how its agents navigate IVR systems, handling multi-level menus, silence, and human detection out of the box. (Sierra)
OpenAI just changed how Voice AI is built (BlogGeekMe)
A dev builds an on-device real-time translator on macOS 26 using Apple’s new speech, translation, and LLM APIs for live subtitles. (dev.to)
AssemblyAI walks through Whisper speaker diarization, showing how modern models attribute speech from segments as short as 250ms. (AssemblyAI)
FirstBranch turns Congress hearings into a search engine, built in three days with transcripts, embeddings, clustering, and a map. (Felix Haba)
The SLT 2026 SmartGlasses Challenge benchmarks egocentric multi-talker speech recognition on 106 hours of four-channel wearable audio. (arXiv)
A dev.to guide covers batch audio transcription in Node.js, using async webhooks to process long recordings without live streaming. (dev.to)
A benchmark tests local speech-to-text with Foundry Local and C#, measuring cold-start, real-time factor, memory, and accuracy on-device. (C# Corner)
YazSes is an offline hold-to-talk voice dictation tool, a privacy-first cross-platform system that runs transcription fully on-device. (GitHub)

