Events đ€
Vapi hosts âEnterprise voice AI after hoursâ at HumanX Amsterdam (Sep 23, Amsterdam, LUMA)
AssemblyAI runs âBuild Night: Create Your Own Dictation Appâ (Sep 24, San Francisco, LUMA)
Google DeepMind hosts âGemini Audio | At Nightâ (Sep 24, San Francisco, DeepMind)
Top Updates đȘ
Google launches Gemini 3.8 Live, voice model that reason while speaking and tops the Artificial Analysis Speech-to-Speech index. (Google)
xAI ships Grok Voice Transcribe 2.0, calling it twice as accurate as v1 and ranking first among streaming models on the Artificial Analysis leaderboard. (xAI)
ServiceNow integrates Krispâs Voice Isolation to power its Voice Agents (LinkedIn)
Speechmatics launches Agent STT, a model built to catch the high-consequence errors, like a wrong digit or missed negation, that derail voice agents. (Business Insider)
Apple launches Siri AI, a rebuilt assistant with personal context, onscreen awareness, and systemwide app actions, rolling out in beta in English. (Apple)
StepFun ships StepAudio 3 Realtime, a think-while-speaking duplex model that posts 98.9 on the Full-Duplex voice benchmark. (AI Weekly)
Wispr Flow debuts Canto, its first speech model, cutting noisy real-world error rates from over 30% to single digits (Wispr Flow)
DeepL Voice now preserves your voice, carrying tone, rhythm, and intonation across real-time translation in 30+ languages on Zoom, Teams, and Meet. (PR Newswire)
Microsoft adds real-time voice agents to Copilot Studio, letting organizations build speech-to-speech agents first via Dynamics 365 Contact Center. (Microsoft)
Superhuman acquires AI notetaker Fathom, folding bot-free meeting capture into its email, calendar, and agent suite. (citybiz)
Treble raises $18M to expand its voice simulation platform, generating synthetic acoustic data for model training, robotics, and consumer hardware. (TechCrunch)
Google pushes AI for every language, extending its speech stack toward the 1,000 most-spoken languages via cross-lingual transfer. (Google)
Radisys launches the V.AI ecosystem, bundling ElevenLabs, Hiya, Speechmatics, and others so telcos can deploy voice AI inside their own networks. (Korea Newswire)
Zoom pivots from video calls to agentic work, positioning ZoomMate to search across systems and finish follow-up tasks, not just summarize. (WebProNews)
WhatsApp tests speech-to-text messaging, letting users dictate a message and send it as text, processed on-device and offline. (The News)
Goodlord warns AI voice spoofing is gaming tenant referencing, with fraud rings using real-time voice tools to bypass reference callbacks. (Goodlord)
Research shows scammers clone local accents, exploiting familiarity to lower listenersâ guard as AI voice scams surged over 1,200%. (Android Headlines)
Google DeepMind argues AGI will be spoken, not typed, making the case for one natively multimodal speech-to-speech model over cascaded pipelines. (BigGo)
FineVoice launches Aunio, an AI audio production agent that coordinates voices, music, and sound effects into publish-ready audio end-to-end. (PR Newswire)
RingCentral rolls out AIR Pro for Healthcare, an agentic voice AI with prebuilt patient workflows and native integration to 100+ EHR systems. (RingCentral)
ImpactFactory.ai ships version 1.5, adding a full-duplex conversational agent that guides users through marketing tasks. (EIN Presswire)
Fierce argues telcos must embrace voice AI or get left behind, as Alianzaâs Crux pitches operators a path beyond being a dumb pipe. (Fierce Network)
Engineering Corner đ
Google details building real-time voice apps with Gemini Audio, its developer guide to the new Live and Extended Thinking models. (Google)
A benchmark comparison of speech-to-text APIs finds Meta Muse leading streaming WER at 3.1% while undercutting Googleâs pricing. (Tech Insider)
Nari Labs tops Covalâs voice benchmarks, taking #1 STT latency and #1 TTS word error rate in the mid-September run. (Nari Labs)
HackerNoon explains Grok Voice Realtime, xAIâs audio-to-audio model with tools, search, and WebSocket streaming for voice agents. (HackerNoon)
A tutorial builds a Vapi + ElevenLabs appointment agent that books slots and sends reminders, wired to Google Calendar and Twilio via n8n. (dev.to)
A guide to fine-tuning NVIDIA Nemotron 3.5 ASR walks through adapting the open model to a language, domain, or accent with NeMo. (dev.to)
A paper documents building a production Greek-English recognizer, evaluating 23 training runs against nine WER, LID, and hallucination gates. (arXiv)


Two more updates worth considering adding: Deepgram Nova-3 Pharma STT model and India Endpoint both launched.