Events 🎤
AssemblyAI: gathering builders and founders for demos and networking on shipping production voice agents. (Sep 1, New York, LUMA)
Vapi and Cartesia: hands-on session on turn detection, interruption handling, latency, and locale support inside Cartesia’s Playground. (Sep 2, LUMA)
Top Updates 💪
Google releases Gemini 3.5 Transcribe, a speech-to-text model reporting 2.6% average WER across 85+ languages while stripping fillers and fixing slips. (Google)
SoundHound moves to acquire LivePerson, combining voice agentic AI with digital messaging to build an end-to-end omnichannel platform. (Constellation Research)
Krisp’s VIVA powers Telnyx’s voice AI stack, adding upstream voice isolation so agents transcribe more accurately and avoid false interruptions at 10M+ call minutes a day. (Krisp)
Ringg raises $15M to push its voice AI orchestration layer past the phone call into WhatsApp and browser agents. (Slator)
Plaud launches the One AI earbuds, a $249 wearable with an LTE-enabled case that records and transcribes meetings without a phone. (9to5Google)
Ex-Nothing founders launch Relay Q, a $150 dictation mic and macOS app running on Gemini for quiet, accurate voice-to-text. (Wired)
DeepL research finds live voice translation is the AI Americans want most, chosen over meeting summaries and email drafting. (PR Newswire)
Intron’s new voice model handles African code-switching, following speakers who change languages mid-sentence for clinics and contact centers. (TechCabal)
3CLogic releases an AI Agent Evaluator, auto-scoring 100% of voice AI interactions on resolution and quality without manual review. (PR Newswire)
Fireflies wants to turn its notetaker into an action-taker, with CEO Krish Ramineni pushing AI teammates that execute post-meeting work. (Economic Times)
Crexendo partners with Tresic to turn every customer conversation into AI-driven tasks and alerts for its 250+ service-provider licensees. (Newswire)
A blind Egyptian entrepreneur builds ScribeMe, an AI app giving running audio descriptions of surroundings, now used in 140 countries. (Times of India)
CX Today argues agentic CX exposes operational gaps, as ServiceNow, Salesforce, and Synthflow shift the question to running autonomy with accountability. (CX Today)
Unauthorized AI celebrity voices are spreading in promotion, raising right-of-publicity and false-endorsement fights as the NO FAKES Act advances. (Big News Network)
TechRadar dissects why ChatGPT’s new voice works, finding its “umms,” pauses, and breaths are what make it feel human. (TechRadar)
Engineering Corner 😎
Daily releases Pipecat PhoneLLM Alpha 1, an open-weights voice-agent LLM that matches GPT-5.6 Terra at 94% lower cost and faster time-to-first-token. (Daily)
Krisp ships an MCP upgrade for meeting data, adding smarter search, bulk AI tagging, bulk speaker correction, and AI-managed action items. (Krisp)
Google shows how to evaluate live voice agents in ADK, driving agents with simulated users that speak their turns as Gemini TTS audio. (Google Developers)
A guide to evaluating voice agents with LangSmith scores them on real outcomes: bookings made, tickets resolved, transfers landed. (dev.to)
A deep dive on the 2000ms vs 250ms architecture war explains why cascaded voice stacks still beat speech-to-speech in production. (Towards AI)
A complete guide to Gemini 3.5 Transcribe walks through verbatim timestamps vs. clean smart-mode output with runnable Colab code. (dev.to)
A tutorial builds a production audio transcription pipeline in Python, wiring up streaming STT with timestamps and diarization. (dev.to)
A paper fine-tunes NVIDIA Canary Flash for telephony ASR, adapting compact models to noisy contact-center audio and jargon. (arXiv)
TLive-Omni is an omni-modal model for live commerce, fusing speech, video, and text with timestamped tokens for real-time understanding. (arXiv)
HackerNoon shows how Audio8 TTS-0.1B brings voice cloning to smaller GPUs, a ~170M-parameter model with zero-shot cloning. (HackerNoon)
Vibe is a free open-source local transcription app for Windows, macOS, and Linux that runs Whisper-based models fully offline. (MajorGeeks)

