Events
Vapi and Inngest build a voice agent live, taking it from demo to production with retries, long-running execution, and evals along the way. (Jul 29, Hybrid, Voice AI Space)
Temporal hosts a durable multimodal AI meetup with HeyGen, Vapi, and Modal, including a Vapi talk on scaling voice agents from one call to ten thousand. (Jul 29, SF, LUMA)
Top Updates 💪
OpenAI brings ChatGPT Voice to desktop with GPT-Live-powered agents that can control apps, dictate in any window, and work with Codex on Mac and Windows. (TechCrunch)
Anthropic upgrades Claude voice mode to Opus and Sonnet models with cross-app automation for Gmail, Slack, and Notion in 10 languages. (WebProNews)
Alibaba launches Qwen Audio 3.0 TTS in Flash and Plus tiers across 16 languages, reaching #1 on the Artificial Analysis TTS arena at 1,236 Elo. (MarkTechPost)
ByteDance releases Seed Audio 1.0, a unified model for voice, music, and sound effects with zero-shot voice cloning from a single reference clip. (ByteDance)
Deepgram puts Nova-3 on Snapdragon with on-device voice AI optimized for Qualcomm’s Hexagon NPU, eliminating the need for cloud connectivity. (Yahoo Finance)
Telli raises $15M seed from Redalpine, Y Combinator, and Cherry Ventures to build AI agents that replace traditional call center operations. (PYMNTS)
Valence AI raises $5M backed by SRI International to integrate real-time emotion detection into voice AI with US patents on emotional intelligence. (SRI)
Cast Insights raises $4.5M pre-seed for real-time speech intelligence across TV, radio, and podcasts, having processed over 2.3M hours of audio. (SiliconAngle)
Zoom launches real-time voice translation that converts spoken language live during meetings across five languages. (Zoom)
Zoom Scribe adds speech accessibility features for real-time captioning and transcription to improve meeting inclusivity. (Zoom)
Otter.ai introduces Live Assist, a real-time coaching agent that listens to calls and provides guidance customizable with team playbooks. (BusinessWire)
Level AI launches Latitude, a suite of 7 purpose-built CX models that the company claims are 49x cheaper than frontier LLMs for contact centers. (Smart Customer Service)
Apple rolls out Genius Bar Live Notes, using AI to transcribe and summarize in-store customer appointments in real time. (HotHardware)
Ukraine adds voice AI to its Diia government app, letting citizens access 170+ public services through ElevenLabs-powered spoken conversations. (Smart Cities World)
AudioCodes targets 40-50% voice AI growth in 2026 with its conversational AI segment on track to reach $25M, aiming for $80M by 2028. (Channel Insider)
Parlance nears 2 billion healthcare calls processed by its voice AI platform, marking a major deployment milestone in patient-facing automation. (PR Newswire)
WhisperAI surpasses 330,000 users and launches a transcription API with real-time STT, crossing seven-figure ARR in under a year. (TechBullion)
XMOS unveils VocalFusion XVF3620, an AI voice processor combining on-chip noise reduction, beamforming, and echo cancellation in a single device. (audioXpress)
Synaptics launches Astra SR80, an always-on edge AI audio processor for voice capture, biometric auth, and agentic AI devices. (audioXpress)
Smallest AI’s TTS ranks #1 for Hindi in blind listening tests, with Lightning v3.1 preferred 76% of the time over OpenAI’s GPT-4o-mini-TTS. (North Jersey)
MouthPad launches a $1,400 tongue-controlled trackpad alongside Vox, a $200 whisper-to-type wearable microphone for hands-free voice input. (Morningstar)
Sweekar AI pocket pet uses staged voice development from babbling to fluent speech as a core growth mechanic, debuted at CES 2026 by Takway AI. (audioXpress)
SoliderSound launches Go-Denoise, a desktop app using neural spectral processing to remove noise from speech and vocals in a single pass. (Monthly Mixing)
Forbes asks if AI is taking over audio as AI now hosts radio shows, narrates audiobooks, and generates podcasts, sparking debate over quality vs. human narration. (Forbes)
Parloa argues CX leaders should stop measuring voice AI by deflection alone, pushing for resolution rate, customer effort, and escalation quality as better metrics. (CX Today)
Engineering Corner 😎
Cue AI runs Gemma 4 locally for voice dictation, cutting latency 44% and dropping marginal inference cost to zero on desktop voice agents. (Google DeepMind)
Oracle builds a 24/7 healthcare voice agent using NVIDIA PersonaPlex and LiveKit on OCI, demonstrating full-duplex speech-to-speech for medical assistants. (Oracle)
MarkTechPost compares the best open ASR models of 2026, finding top models within one WER point while license and streaming become the real differentiators. (MarkTechPost)
The voice agent latency playbook argues that input accuracy is a latency feature since wrong transcripts turn one conversation turn into three. (HackerNoon)
AssemblyAI details the future of real-time STT with its Universal-3.5 Pro model, semantic turn detection, and a single-WebSocket Voice Agent API. (AssemblyAI)
A dev.to guide on voice agent turn-taking explains how to keep AI calls under 600ms end-to-end with VAD, barge-in handling, and streaming pipelines. (dev.to)
NVIDIA Nemotron 3.5 ASR runs 14x faster than Whisper on a CPU-only laptop via parakeet.cpp, transcribing a 1m46s clip in 32 seconds vs. Whisper’s 7m34s. (Medium)
OpenWhispr is an open-source voice dictation app supporting local Whisper and NVIDIA Parakeet models with zero telemetry, now at 2,100+ GitHub stars. (GitHub)
HackerNoon explains why TTS evaluation needs human ears, noting that WER and CER miss intonation and naturalness, requiring manual listening across checkpoints. (HackerNoon)
A developer narrates their blog with Kokoro, an 82M-parameter local TTS model that generates two hours of audio in 15 minutes on an M1 MacBook. (Bart de Goede)

