Events
Agora runs a hands-on Voice AI workshop in San Francisco (Aug 13, San Francisco, Voice AI Space)
Top Updates 💪
Microsoft tests MAI-Realtime, its first native full-duplex voice model that listens while speaking across 17 languages, cutting Copilot’s reliance on OpenAI. (Startup Fortune)
ByteDance launches SeedRealtime, a full-duplex model that watches, listens, and speaks at once, now live in the Doubao app’s phone-call feature. (ByteDance)
Wispr Flow launches Notetaker, a Mac meeting assistant that captures system audio bot-free, cleans transcripts, and auto-creates tasks. (9to5Mac)
Omilia raises $67M Series B to expand its voice-first agentic CX platform, having grown live ARR more than 10x to over $60M since Series A. (CMSWire)
Five9 lands a ~$100M contract with a Fortune 100 financial services firm migrating off on-prem, won in partnership with Google Cloud. (Pulse 2.0)
Salesforce launches Japanese Agentforce Voice using Kotoba’s Koto model for low-latency speech that handles honorifics and context-dependent nuance. (IBTimes)
Sierra lands CarMax for inbound sales call AI, boosting call resolution and cutting unresolved calls since its May deployment. (CMSWire)
Yellow.ai goes public via a $550M SPAC merger, aiming to buy legacy BPOs and convert them into AI-native operations. (CMSWire)
8x8 posts record revenue as AI adoption surged 121% year over year, with over 2,900 AI agents built on its platform. (IT Brief)
3CLogic wins a major hospital system to modernize its IT service desk with voice AI embedded natively in ServiceNow. (PR Newswire)
Granola faces a class-action lawsuit alleging its bot-free notetaker records meetings without all-party consent and trains AI on them by default. (Computerworld)
ElevenLabs deepens its Tokyo push, adapting its voice platform for the Japanese market as APAC expansion accelerates. (btrax)
Orvera AI is shortlisted by Everest Group in its voice AI agents for CXM spotlight, cited for strength in banking, insurance, and healthcare. (MarTech Series)
Telnyx makes the case for owned voice AI infrastructure, arguing single-backbone control of telephony, compute, and inference beats stitched multi-vendor stacks. (Under30CEO)
Last Week
The 2026 State of Voice in CX
Voice AI is scaling faster than contact center operating models and it shows.
Engineering Corner 😎
OpenAI details how it built GPT-Live, a full-duplex realtime system with a delegation layer that hands hard tasks to GPT-5.5 while keeping you talking. (OpenAI)
Google Creative Lab ships an offline Gemma Translator running Gemma 4 E2B on a Raspberry Pi 5, with full source and 3D-print files released. (AI Weekly)
Cisco IT shares how it modernized voice security with AI, cutting toll fraud 70% and manual investigation effort 60% with behavioral analytics. (Cisco)
AWS publishes a serverless real-time voice AI pattern for enterprise sales coaching, built on native services for live transcription and analytics. (AWS)
Microsoft shows live speech-to-text with Foundry Local and C#, streaming raw PCM audio to a local Nemotron model with no API key or cloud. (Microsoft)
A dev.to tutorial turns a phone call into a formatted email in 80 lines of Python, cleaning up AI voice memos with the Telnyx API. (dev.to)
A dev.to explainer breaks down modern voice cloning, walking through the speaker encoder, synthesis model, and vocoder pipeline. (dev.to)
A new paper introduces QazAVSR, a Kazakh audio-visual speech recognition model built on HuBERT and ViT with a 57-hour dataset. (MDPI)
A study adapts foundation models for Turkic speech-to-speech translation, comparing fine-tuned cascade and direct end-to-end approaches. (MDPI)


