Events 🎤
Decagon runs Decagon Dialogues 2026 (Oct 1, San Francisco, Decagon)
Speechmatics and Tuner co-host a voice AI meetup (Oct 1, London, LUMA)
Top Updates 💪
Meta unveils Ray-Ban Meta Audio, camera-free AI glasses at $349 that pair earbud audio with the Muse assistant, translation, and calls. (Meta)
Krisp publishes an open Voice Isolation benchmark, testing 11 STT configurations on real audio with a second voice, where word error rate drops 73%. (Hugging Face)
Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS (Google)
Qualcomm announces Snapdragon Sound Elite Gen 2, an audio chip that runs on-device AI and Wi-Fi in earbuds and glasses without a phone. (CNET)
SoundHound introduces OASYS Edge, fully embedded agentic voice AI that runs locally in vehicles and smart devices even offline. (GlobeNewswire)
Sarvam AI releases Saaras V4, a speech-to-text model covering all 22 Indian languages plus English with sub-150ms streaming. (Analytics Insight)
Inworld ships Realtime TTS-2, giving developers natural-language control over delivery in live interactions, with a Flash variant at 25ms time-to-first-byte. (Entrepreneur)
Sela raises $21M for voice AI agents that help originate over $1B in mortgage loans a month, used by six of the ten largest independent mortgage banks. (Yahoo Finance)
Dextr AI raises $6.7M for a hospitality agent platform whose Daisy voice agent books reservations and takes payments in 90+ languages. (WebWire)
Krisp brings real-time voice AI to contact centers, installing at the OS level to add translation, accent conversion and voice security on any softphone. (PC Tech Magazine)
A 60-org coalition backed by the Gates Foundation commits to bringing AI to 3.4 billion people in their own languages within five years. (Gates Foundation)
A Washington court rules patients can’t access ambient AI recordings, treating AI scribe audio as exempt administrative documentation. (Clinical Advisor)
Gizmodo argues AI assistants belong in your ears, not on your face, making the case for earbuds over camera glasses on privacy and social grounds. (Gizmodo)
Telecoms.com argues voice AI inference belongs inside the network, where edge deployment cuts latency and adds native deepfake defense. (Telecoms.com)
QSR Web says voice AI has rewritten the drive-thru, as chains scale order-taking agents to ease staffing and keep the line moving. (QSR Web)
Engineering Corner 😎
NVIDIA releases Nemotron 3 Diarization, a 100M-parameter open-weight model that tracks up to eight speakers in real time and tops Diarization-Bench. (HackerNoon)
AWS shows deploying Qwen3-TTS on SageMaker, with 3-second voice cloning and instruction-driven timbre and emotion across 10 languages. (AWS)
AWS packages WhisperX on SageMaker for word-level, speaker-labeled transcription with alignment, diarization, and subtitle output. (AWS)
Pariveda built real-time clinical voice notes for Henry Schein One on Amazon Nova and Bedrock, turning chairside talk into structured notes. (AWS)
A dev.to guide builds sub-400ms voice agents for enterprise telephony, covering VAD tuning, streaming STT, and turn-level latency budgets. (dev.to)
SayScroll launches a voice-activated AI teleprompter that scrolls to a speaker’s cadence and pauses on hesitation across 60+ languages. (Dispatch)

