Top Updates 💪
Meta launches Muse Voice Transcribe, one real-time model for streaming ASR, diarization, and endpointing across 70+ languages at $0.18 an hour. (Meta)
Microsoft ships MAI-Transcribe-2, undercutting OpenAI, Google, and ElevenLabs on price at $0.10/hour (VentureBeat)
IBM releases Granite Speech 5.0 TurboCTC, a 470M open model transcribing 3.5 hours of speech per second by dropping the decoder. (Slator)
Phonely launches Alma, a voice LLM trained on 10M+ phone calls that runs sub-185ms and is 84% cheaper than GPT-4.1. (SiliconAngle)
Deepdub launches Phantom Z 3.4 Conversational, a multilingual TTS with 150ms time-to-first-audio built to survive real customer calls. (PR Newswire)
Google adds voice to Gmail, Docs, and Keep, letting users talk to their inbox and dictate notes via Gemini-powered Live modes. (TechTimes)
Google Translate adds background live translation on Android, keeping translation running when you switch apps or lock the screen. (WebProNews)
Genesys expands agentic voice AI at Xperience 2026, adding end-of-turn detection and native Deepgram and ElevenLabs integrations. (DestinationCRM)
Dialpad and Glean connect conversation insights to enterprise knowledge, bringing call intelligence into Glean’s knowledge platform. (Dialpad)
Digital Island partners with RingCentral as its first New Zealand partner to bring agentic voice AI to local businesses. (StockTitan)
Krisp CEO Davit Baghdasaryan on transforming enterprise communication, discussing how voice AI is reshaping the enterprise stack. (Analytics Insight)
Walmart is sued over AI voiceprints in a proposed Illinois class action alleging it built biometric templates from calls without BIPA consent. (Travel News)
SwitchBot launches the AI MindClip, a 16.8g clip-on wearable powered by Qwen that turns conversations into summaries and to-dos. (Incredible Things)
TikTok adds voice comments, letting users leave 60-second audio replies alongside comment polls and photo carousels. (TechMyMoney)
Tesla wires Grok Voice Think Fast 2.0 into the Summer Update, giving in-cabin voice full vehicle control across 100+ commands. (Not a Tesla App)
Engineering Corner 😎
Nokia introduces Automatic Contextual Audio Denoising, a task where the system infers scene context before deciding what counts as noise. (Nokia)
A TTFT-first benchmark ranks lowest-latency inference APIs for voice agents, finding Groq’s LPU leads on time-to-first-token. (MarkTechPost)
A hands-on comparison of GPT-Live, Gemini Live, and Grok Voice breaks down latency, pricing, and interruption handling across the three. (Tech Insider)
VoiceStudio is a fully local ElevenLabs alternative, bundling 16 TTS and 11 ASR engines with a local API and MCP in 646 languages. (dev.to)
Vowen 0.5.6 ships offline voice dictation, a free system-wide tool running local Whisper models across 99 languages with no cloud. (Warp2Search)

