Events
Deepgram hosts a fireside chat on the future of voice agents, evaluating agents beyond WER and whether cascaded STT-LLM-TTS still holds up. (Jul 23, London, Voice AI Space)
AssemblyAI demos Universal-3.5 Pro Realtime, its new model with context carryover and conversation memory, then opens up for a fireside chat with production voice agent builders. (Jul 23, SF, Luma)
Top Updates 💪
Alibaba launches Qwen Audio 3.0 with real-time voice that can proactively use external tools, covering 113 languages for ASR and 36 for TTS. (KuCoin)
Meta patents an AI wearable that continuously analyzes voice to track the user’s emotional state, raising concerns under the EU AI Act’s emotion-inference ban. (The Next Web)
PwC and OpenAI launch agentic customer service solutions combining PwC’s CX expertise with OpenAI’s multimodal voice and digital agent APIs. (CX Today)
Google Voice adds Gemini AI notes that auto-summarize calls with key points and action items, plus new standalone plans starting at $10/mo. (WebProNews)
Google quietly opted users into AI training on voice queries and uploaded media via a new Search Services History setting, with no opt-in required. (Fox News)
Rime raises $24M Series A to build enterprise speech-to-speech models, powering nearly 100M phone calls monthly for Mayo Clinic, Dialpad, and others. (Rime)
LALAL.AI launches Lynx, a neural network built for speech denoising that is 6x smaller than its flagship model while matching output quality. (Slator)
Sber’s GigaChat adds emotion detection and can process audio up to three hours long with speaker separation, timestamps, and segment summaries. (BusinessNewsThisWeek)
DoorDash, ObserveAI, and AWS scale AI-powered quality evaluation across 19,000 agents, automating nearly 100% of interaction reviews. (PR Newswire)
Samsung adds cloud transcription to its Voice Recorder app, giving users a choice between on-device privacy and cloud-powered accuracy. (WebProNews)
Aina raises $5.5M to build hardware that controls AI agents rather than just recording, with its first product Dune already shipping to early adopters. (TechCrunch)
Chen Institute and Science honor neuroscientist Sergey Stavisky for an AI speech neuroprosthesis that decodes brain activity into spoken words at 97.5% accuracy. (PR Newswire)
New research finds that AI voice phishing works because of persuasive scripting, not vocal realism, as 70% of targets detect the synthetic voice but comply anyway. (Help Net Security)
Instadesk, Huawei, and iFlytek open a joint AI customer experience lab in Uzbekistan, combining multilingual ASR/TTS with Ascend cloud infrastructure. (Manila Times)
VoicePing 3.0 launches with real-time translation, AI dubbing, new ASR and MT models, and MCP/API access for enterprise multilingual workflows. (PR Newswire)
A study flags five risks in clinical AI scribes: inconsistent consent, weak performance on accented speech, background noise, missing human review, and unclear accountability. (ResultSense)
Telcos are sitting on a voice AI opportunity bigger than their own cost savings, argues an analysis that says the real play is selling voice infrastructure to enterprises. (Sebastian Barros)
Engineering Corner 😎
Apple’s SpeechAnalyzer API outperforms Whisper Small in English benchmarks, running fully on-device on iOS 26 and macOS Tahoe. (Gigazine)
Cohere releases Transcribe Arabic, an open-source 2B-param ASR model that beats Meta’s 7B model on Arabic with multidialect and code-switching support. (Cohere)
Llamafile 0.10.4 ships transcribefile, a portable speech-to-text CLI built on transcribe.cpp that supports 16+ model families with GPU acceleration. (Phoronix)
A prototype tongue-reading system uses ultrasound and ML to decode silent speech from tongue movements, enabling voice input without making a sound. (Hackaday)
Adafruit demos VAD, STT, and TTS all running on a single RP2040 microcontroller, bringing a complete voice AI pipeline to a $4 chip. (Adafruit)
ReSpeaker Clip is an open-source wearable AI recorder with dual mics, BLE 5.3, Wi-Fi 6, and full SDK for building custom voice AI applications. (Seeed Studio)
Resemble AI explains how neural audio watermarking embeds inaudible signals during voice generation for traceability, ahead of the EU AI Act’s August 2 deadline. (Resemble AI)
A Python tutorial walks through building a real-time AI phone agent that quotes prices using tool calling and low-latency voice synthesis. (Low Latency Club)
Voice AI benchmarks hide a gap: systems scoring 600-800ms in the lab hit 2-4 seconds on real telephony, and accuracy drops from 51% to 26-38% under realistic audio. (Embedded Computing)

