Events
Agora Convo AI World mixer in Nashville (Aug 4–5, Nashville, LUMA)
Vapi and Deepgram go live on model tradeoffs, comparing Nova-3, Flux, and Aura-2 across transcription, turn detection, and latency. (Aug 6, Hybrid, LUMA)
Agora and Seeed Studio run a hands-on hardware workshop at Silicon Valley’s Robotics Fair (Aug 8, San Mateo, CA, Voice AI Space)
Top Updates 💪
xAI launches Grok Voice Think Fast 2.0, which reasons while speaking for smarter answers with no added latency, scoring 82.9 on the speech quality index. (xAI)
OpenAI ships GPT-Transcribe and GPT-Live-Transcribe, cutting transcription errors by 52% vs. Whisper with real-time and batch modes across 57 languages. (Slator)
Google gives Gemini for Mac voice control with Fn-key activation, intelligent dictation that strips filler words, and an opt-in screen-aware reasoning mode. (9to5Google)
Fish Audio raises $52M in seed funding after hitting $21M ARR and 8M+ users in its first year, with voice cloning and streaming across 80+ languages. (TechCrunch)
Encore AI raises $30M Series A to build voice agents that learn winning behaviors from customer call recordings using “interaction mining.” (TechCrunch)
Smallest AI raises $13M Series A to make voice agents indistinguishable from humans, bringing total funding to $21M with sub-second latency models. (TechCrunch)
PolyAI releases Dialog RSN-1, an audio-native model that reasons directly over raw call audio instead of transcripts, delivering sub-300ms responses. (SiliconAngle)
OpenAI expands GPT-Live to Edu, Business, and Enterprise plans globally, bringing full-duplex voice with background reasoning to organizations. (TheWinCentral)
Qwen Audio 3.0 Realtime Plus tops OpenAI on the Artificial Analysis Speech-to-Speech Index at 84.1% vs. GPT-Realtime-2.1’s 79.1%, a first for Alibaba. (BetaNews)
Sarvam AI announces a trillion-parameter model built in India, with pricing 5.5x cheaper than GPT-5.4 Mini for coding and research workloads. (Free Press Journal)
Krafton releases A.X K2 Raon-Speech, a 21B-parameter audio model that processes speech directly to preserve emotion and tone, ranked first among Korean models. (The Investor)
Boson AI unveils Higgs RealTime for full speech-to-speech processing that skips text conversion to preserve vocal nuance, led by ex-AWS scientist Alex Smola. (CryptoBriefing)
8x8 extends AI across its full CX platform beyond the contact center, adding conversation intelligence, agent development, and workforce management for every team. (CMSWire)
NICE and RingCentral expand into a bi-directional partnership, combining UCaaS, CCaaS, and AI in a single offering with mutual resale. (Telecom Reseller)
Parlance adds instant self-service config and AI call summaries for healthcare contact centers, now deployed across 400+ health systems. (PR Newswire)
Tata Communications launches voice AI for India’s 63M SMBs with speech-to-speech engagement, sub-500ms latency, and multilingual support via Tata Tele. (Inc42)
Tesla rolls out Grok voice assistant in India for the Model Y via OTA update, supporting Hindi and five other Indian languages for navigation and queries. (IndianWeb2)
Japan’s Justice Ministry backs voice rights protection from AI, concluding that publicity rights can address unauthorized voice cloning without new legislation. (Kyodo News)
Tysa, Conectys’ AI voice agent, marks one year of handling global customer conversations across multiple languages and use cases. (EIN Presswire)
Last Week Podcast
Engineering Corner 😎
HackerNoon breaks down voice-to-voice AI architectures, from codec tokenization and RVQ prediction to full-duplex designs like Moshi and GPT-Live. (HackerNoon)
A deep dive into ARK-ASR-3B’s architecture shows how combining a Whisper encoder with a Qwen decoder achieves 5.04% WER on the Open ASR Leaderboard. (HackerNoon)
Building SeMamba for speech enhancement walks through using Mamba state-space models to recover clean speech from noisy audio in a single forward pass. (Level Up Coding)
Deepgram integrates with AWS SageMaker AI via IAM temporary delegation, letting enterprises self-host Nova and Aura-2 with zero standing cross-account access. (AWS)
A Frontiers in AI paper compares Wav2Vec and Whisper for Telugu speech recognition, benchmarking pretrained models on low-resource Indian languages. (Frontiers)

