<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Voice AI Newsletter]]></title><description><![CDATA[Voice AI insights from Krisp's CEO]]></description><link>https://voice-ai-newsletter.krisp.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!YLgs!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F831a2f7e-d0a7-4e3d-87a8-c42c65d0b71c_1000x1000.png</url><title>Voice AI Newsletter</title><link>https://voice-ai-newsletter.krisp.ai</link></image><generator>Substack</generator><lastBuildDate>Mon, 10 Aug 2026 19:50:09 GMT</lastBuildDate><atom:link href="https://voice-ai-newsletter.krisp.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Krisp Technologies]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[krispai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[krispai@substack.com]]></itunes:email><itunes:name><![CDATA[Davit Baghdasaryan]]></itunes:name></itunes:owner><itunes:author><![CDATA[Davit Baghdasaryan]]></itunes:author><googleplay:owner><![CDATA[krispai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[krispai@substack.com]]></googleplay:email><googleplay:author><![CDATA[Davit Baghdasaryan]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Microsoft and ByteDance join the full-duplex race ]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/microsoft-and-bytedance-join-the</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/microsoft-and-bytedance-join-the</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 10 Aug 2026 14:02:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/80febbbf-2348-4780-b186-1d6f1715da79_2560x1440.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><span>Agora runs a hands-on Voice AI workshop in San Francisco</span> (Aug 13, San Francisco, <a href="https://www.voiceaispace.com/events/voice-ai-workshop---sf-bay-area"><span>Voice AI Space</span></a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Microsoft tests MAI-Realtime</strong>, its first native full-duplex voice model that listens while speaking across 17 languages, cutting Copilot&#8217;s reliance on OpenAI. (<a href="https://startupfortune.com/microsoft-quietly-tests-mai-realtime-a-voice-ai-that-talks-while-it-listens/">Startup Fortune</a>)</p></li><li><p><strong>ByteDance launches SeedRealtime</strong>, a full-duplex model that watches, listens, and speaks at once, now live in the Doubao app&#8217;s phone-call feature. (<a href="https://seed.bytedance.com/en/SeedRealtime">ByteDance</a>)</p></li><li><p><strong>Wispr Flow launches Notetaker</strong>, a Mac meeting assistant that captures system audio bot-free, cleans transcripts, and auto-creates tasks. (<a href="https://9to5mac.com/2026/08/05/wispr-flow-takes-on-ai-meeting-assistants-with-notetaker-its-first-product-beyond-dictation/">9to5Mac</a>)</p></li><li><p><strong>Omilia raises $67M Series B</strong> to expand its voice-first agentic CX platform, having grown live ARR more than 10x to over $60M since Series A. (<a href="https://www.cmswire.com/customer-experience/omilia-raises-67m-series-b-for-voice-ai-push/">CMSWire</a>)</p></li><li><p><strong>Five9 lands a ~$100M contract</strong> with a Fortune 100 financial services firm migrating off on-prem, won in partnership with Google Cloud. (<a href="https://pulse2.com/five9-lands-approximately-100-million-contract-as-voice-ai-agents-drive-contact-center-push/">Pulse 2.0</a>)</p></li><li><p><strong>Salesforce launches Japanese Agentforce Voice</strong> using Kotoba&#8217;s Koto model for low-latency speech that handles honorifics and context-dependent nuance. (<a href="https://jp.ibtimes.com/salesforce-launches-japan-agentforce-voice-kotoba-technology-103514">IBTimes</a>)</p></li><li><p><strong>Sierra lands CarMax</strong> for inbound sales call AI, boosting call resolution and cutting unresolved calls since its May deployment. (<a href="https://www.cmswire.com/customer-experience/sierra-lands-carmax-as-conversational-ais-enterprise-land-grab-widens/">CMSWire</a>)</p></li><li><p><strong>Yellow.ai goes public via a $550M SPAC merger</strong>, aiming to buy legacy BPOs and convert them into AI-native operations. (<a href="https://www.cmswire.com/contact-center/inside-yellow-ais-550m-ipo-target-and-vision-to-transform-legacy-call-centers-with-ai/">CMSWire</a>)</p></li><li><p><strong>8x8 posts record revenue</strong> as AI adoption surged 121% year over year, with over 2,900 AI agents built on its platform. (<a href="https://itbrief.com.au/story/8x8-revenue-hits-record-as-ai-adoption-more-than-doubles">IT Brief</a>)</p></li><li><p><strong>3CLogic wins a major hospital system</strong> to modernize its IT service desk with voice AI embedded natively in ServiceNow. (<a href="https://www.prnewswire.com/news-releases/major-multi-hospital-system-selects-3clogic-to-modernize-it-service-operations-with-voice-ai-for-servicenow-302841610.html">PR Newswire</a>)</p></li><li><p><strong>Granola faces a class-action lawsuit</strong> alleging its bot-free notetaker records meetings without all-party consent and trains AI on them by default. (<a href="https://www.computerworld.com/article/4206255/granola-lawsuit-raises-concerns-over-ai-note-taking-app-privacy.html">Computerworld</a>)</p></li><li><p><strong>ElevenLabs deepens its Tokyo push</strong>, adapting its voice platform for the Japanese market as APAC expansion accelerates. (<a href="https://blog.btrax.com/elevenlabs-tokyo/">btrax</a>)</p></li><li><p><strong>Orvera AI is shortlisted by Everest Group</strong> in its voice AI agents for CXM spotlight, cited for strength in banking, insurance, and healthcare. (<a href="https://martechseries.com/predictive-ai/ai-platforms-machine-learning/orvera-ai-was-included-among-the-shortlisted-providers-featured-in-everest-groups-tech-provider-spotlight-voice-ai-agents-in-customer-experience-management-cxm/">MarTech Series</a>)</p></li><li><p><strong>Telnyx makes the case for owned voice AI infrastructure</strong>, arguing single-backbone control of telephony, compute, and inference beats stitched multi-vendor stacks. (<a href="https://www.under30ceo.com/telnyx-explains-where-voice-ai-infrastructure-platforms-are-earning-their-keep/">Under30CEO</a>)</p></li></ul><div><hr></div><h2>Last Week</h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;67d96f29-bed9-44dd-a3e2-5b367defb1cd&quot;,&quot;caption&quot;:&quot;Voice AI is scaling faster than contact center operating models and it shows.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The 2026 State of Voice in CX&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:32916364,&quot;name&quot;:&quot;Davit Baghdasaryan&quot;,&quot;bio&quot;:&quot;CEO &amp; Co-Founder of Krisp, early pioneer in Voice AI.\n20+ years in engineering. 18 US patent applications, ex Twilion&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/23088dde-6cb0-44df-b220-5f22830cdd4c_1179x960.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-06T14:26:27.687Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abe31bf6-14d9-4d5f-95a3-210a57bce139_1744x800.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/p/the-2026-state-of-voice-in-cx&quot;,&quot;section_name&quot;:&quot;Articles&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:209508711,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:12,&quot;comment_count&quot;:0,&quot;publication_id&quot;:2073467,&quot;publication_name&quot;:&quot;Voice AI Newsletter&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!YLgs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F831a2f7e-d0a7-4e3d-87a8-c42c65d0b71c_1000x1000.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>OpenAI details how it built GPT-Live</strong>, a full-duplex realtime system with a delegation layer that hands hard tasks to GPT-5.5 while keeping you talking. (<a href="https://openai.com/index/continuous-voice-interaction-with-gpt-live/">OpenAI</a>)</p></li></ul><ul><li><p><strong>Google Creative Lab ships an offline Gemma Translator</strong> running Gemma 4 E2B on a Raspberry Pi 5, with full source and 3D-print files released. (<a href="https://aiweekly.co/alerts/google-creative-lab-ships-offline-gemma-translator-for-pi-5">AI Weekly</a>)</p></li><li><p><strong>Cisco IT shares how it modernized voice security with AI</strong>, cutting toll fraud 70% and manual investigation effort 60% with behavioral analytics. (<a href="https://blogs.cisco.com/cisco-on-cisco/how-cisco-it-modernized-voice-security-with-ai">Cisco</a>)</p></li><li><p><strong>AWS publishes a serverless real-time voice AI pattern</strong> for enterprise sales coaching, built on native services for live transcription and analytics. (<a href="https://aws.amazon.com/blogs/industries/serverless-real-time-voice-ai-on-aws-a-pattern-for-enterprise-sales-coaching/">AWS</a>)</p></li><li><p><strong>Microsoft shows live speech-to-text with Foundry Local and C#</strong>, streaming raw PCM audio to a local Nemotron model with no API key or cloud. (<a href="https://devblogs.microsoft.com/dotnet/foundry-local-live-speech-to-text-csharp/">Microsoft</a>)</p></li><li><p><strong>A dev.to tutorial turns a phone call into a formatted email</strong> in 80 lines of Python, cleaning up AI voice memos with the Telnyx API. (<a href="https://dev.to/harpreetseehra/from-phone-call-to-formatted-email-in-80-lines-of-python-ai-voice-memo-cleanup-with-telnyx-32ga">dev.to</a>)</p></li><li><p><strong>A dev.to explainer breaks down modern voice cloning</strong>, walking through the speaker encoder, synthesis model, and vocoder pipeline. (<a href="https://dev.to/peter_zou_7b770f8ba45fd14/understanding-ai-voice-cloning-how-modern-voice-ai-actually-works-16db">dev.to</a>)</p></li><li><p><strong>A new paper introduces QazAVSR</strong>, a Kazakh audio-visual speech recognition model built on HuBERT and ViT with a 57-hour dataset. (<a href="https://www.mdpi.com/2078-2489/17/8/756">MDPI</a>)</p></li><li><p><strong>A study adapts foundation models for Turkic speech-to-speech translation</strong>, comparing fine-tuned cascade and direct end-to-end approaches. (<a href="https://www.mdpi.com/2504-2289/10/8/259">MDPI</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The 2026 State of Voice in CX]]></title><description><![CDATA[815 CX leaders, 12 countries. Here's what they told us.]]></description><link>https://voice-ai-newsletter.krisp.ai/p/the-2026-state-of-voice-in-cx</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/the-2026-state-of-voice-in-cx</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Thu, 06 Aug 2026 14:26:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/abe31bf6-14d9-4d5f-95a3-210a57bce139_1744x800.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4FSV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4FSV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4FSV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg" width="1200" height="503" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:503,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:121605,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4FSV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Voice AI is scaling faster than contact center operating models and it shows.</h3><p>New research by <a href="https://krisp.ai/">Krisp</a>, in partnership with <a href="https://ryanadvisory.com/">Ryan Strategic Advisory</a>, shows voice AI adoption is accelerating, while staffing models, technology stacks, and budgets stay rooted in the past.</p><p>Here are the five things that stood out most.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe to receive new posts and insights.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>1. The biggest barrier to global voice support is people, not technology</h2><p>Contact centers are struggling to find enough agents with the right language skills in the right markets.</p><p>40% named hiring and staffing as a top barrier to scaling voice support globally and another 37% cited the cost of supporting multiple languages.</p><p>Technology integration ranked far lower at 14%.</p><p>Enterprises often frame voice AI as a technology project. The stronger business case is in solving workforce constraints and enabling global expansion.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BODR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BODR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 424w, https://substackcdn.com/image/fetch/$s_!BODR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 848w, https://substackcdn.com/image/fetch/$s_!BODR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 1272w, https://substackcdn.com/image/fetch/$s_!BODR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BODR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png" width="1456" height="737" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:737,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127936,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BODR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 424w, https://substackcdn.com/image/fetch/$s_!BODR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 848w, https://substackcdn.com/image/fetch/$s_!BODR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 1272w, https://substackcdn.com/image/fetch/$s_!BODR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The global CX labor model may reach its limit before the technology does.</p><h2>2. Enterprises are using AI to make existing teams more effective</h2><p>Agent Assist and Speech Analytics received the highest impact scores at 5.0 out of 6, followed by real-time coaching at 4.9.</p><p>Multilingual AI and Voice Translation ranked close behind at 4.8 and 4.7.</p><p>The near-term priority is clear: help current agents work faster and handle more complex conversations. Language AI goes further by expanding the customers and markets those agents can support.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uqUC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uqUC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 424w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 848w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 1272w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uqUC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png" width="1456" height="634" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:634,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:128378,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uqUC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 424w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 848w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 1272w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Agent Assist may win the first budget, but language AI could create the larger economic shift.</p><h2>3. Voice AI adoption is approaching a tipping point</h2><p>Planned adoption matches or exceeds current usage across every AI category measured.</p><p>AI Voice Translation shows the largest jump: 28% currently use it, while 60% plan to adopt it within the next 6&#8211;12 months.</p><p>Planned adoption also reaches 44% for AI Noise Cancellation, 42% for Agent Assist, and 39% for Accent Conversion.</p><p>The market has moved from exploring voice AI to deciding what to deploy first.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!R4e-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!R4e-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 424w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 848w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 1272w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!R4e-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png" width="1456" height="887" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:887,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:165952,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!R4e-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 424w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 848w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 1272w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Companies still treating voice AI as an innovation project are already behind the buying cycle.</p><h2>4. AI adoption is rising faster than AI budgets</h2><p>56% of enterprises allocate more than half of their voice CX budget to onshore agents.</p><p>Meanwhile, 65% spend less than 10% on AI agents, and 55% spend less than 10% on AI translation.</p><p>The biggest share of spending stays tied to the model facing the most pressure. AI can change the economics by helping the same workforce support more customers, markets, and languages.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wDw9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wDw9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 424w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 848w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 1272w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wDw9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png" width="1456" height="916" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:916,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:109699,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wDw9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 424w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 848w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 1272w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The real AI budget may already exist but it&#8217;s sitting inside labor spend.</p><h2>5. The market changed dramatically since 2025</h2><p>Reliance on human translation fell from 65% in 2025 to 35% in 2026.</p><p>Over the same period:</p><ul><li><p>Voice Translation intent rose from 36% to 60%</p></li><li><p>Accent Conversion adoption and intent grew from 24% to 61%</p></li><li><p>AI Noise Cancellation reached 69%</p></li><li><p>Speech Analytics reached 80%</p></li></ul><p>Voice AI is moving into the core contact center stack faster than most operating models can adapt. The bigger risk is building around assumptions that are already outdated.</p><h2>Voice AI is moving faster than most operating models can adapt</h2><p>The market has entered a new phase: adoption is accelerating, budgets are starting to shift, and the old assumptions around staffing, language coverage, and delivery are breaking down.</p><p>The advantage will go to companies that act on those signals early, before fragmented pilots become fragmented infrastructure.</p><p><strong>Read the full 2026 State of Voice in CX report for the complete findings and industry implications.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://resources.krisp.ai/hubfs/WhitePaper/2026%20State%20of%20Voice-VAI.pdf&quot;,&quot;text&quot;:&quot;Read the full report&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://resources.krisp.ai/hubfs/WhitePaper/2026%20State%20of%20Voice-VAI.pdf"><span>Read the full report</span></a></p><p></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[xAI launches Grok Voice Think v2, Fish Audio raises $52M]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/xai-launches-grok-voice-think-v2</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/xai-launches-grok-voice-think-v2</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 03 Aug 2026 14:01:43 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/16517ae1-e167-464a-9141-60fbfcc3463b_1496x1002.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p>Agora Convo AI World mixer in Nashville (Aug 4&#8211;5, Nashville, <a href="https://luma.com/5g2d76uj">LUMA</a>)</p></li><li><p>Vapi and Deepgram go live on model tradeoffs, comparing Nova-3, Flux, and Aura-2 across transcription, turn detection, and latency. (Aug 6, Hybrid, <a href="https://luma.com/vapi-44bz">LUMA</a>)</p></li><li><p>Agora and Seeed Studio run a hands-on hardware workshop at Silicon Valley&#8217;s Robotics Fair (Aug 8, San Mateo, CA, <a href="https://www.voiceaispace.com/events/agora--seeed-studio-physical-ai-workshop-build-customize-and-deploy-an-ai-agent-to-seeed-studio-respeaker-flex-using-agora-mybot">Voice AI Space</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>xAI launches Grok Voice Think Fast 2.0</strong>, which reasons while speaking for smarter answers with no added latency, scoring 82.9 on the speech quality index. (<a href="https://x.ai/news/grok-voice-think-fast-2">xAI</a>)</p></li><li><p><strong>OpenAI ships GPT-Transcribe and GPT-Live-Transcribe</strong>, cutting transcription errors by 52% vs. Whisper with real-time and batch modes across 57 languages. (<a href="https://slator.com/openai-two-new-models-transcription/">Slator</a>)</p></li><li><p><strong>Google gives Gemini for Mac voice control</strong> with Fn-key activation, intelligent dictation that strips filler words, and an opt-in screen-aware reasoning mode. (<a href="https://9to5google.com/2026/07/29/gemini-mac-voice-control/">9to5Google</a>)</p></li><li><p><strong>Fish Audio raises $52M in seed funding</strong> after hitting $21M ARR and 8M+ users in its first year, with voice cloning and streaming across 80+ languages. (<a href="https://techcrunch.com/2026/07/28/fish-audio-raises-50m-seed-to-build-ai-voice-models-for-creators-and-enterprises/">TechCrunch</a>)</p></li><li><p><strong>Encore AI raises $30M Series A</strong> to build voice agents that learn winning behaviors from customer call recordings using &#8220;interaction mining.&#8221; (<a href="https://techcrunch.com/2026/07/29/encore-ai-raises-30m-to-build-ai-agents-that-learn-from-customer-calls/">TechCrunch</a>)</p></li><li><p><strong>Smallest AI raises $13M Series A</strong> to make voice agents indistinguishable from humans, bringing total funding to $21M with sub-second latency models. (<a href="https://techcrunch.com/2026/07/31/smallest-ai-raises-13m-to-build-ultra-fast-voice-ai-that-sounds-genuinely-human/">TechCrunch</a>)</p></li><li><p><strong>PolyAI releases Dialog RSN-1</strong>, an audio-native model that reasons directly over raw call audio instead of transcripts, delivering sub-300ms responses. (<a href="https://siliconangle.com/2026/07/30/polyai-launches-new-real-time-voice-conversation-model-make-ai-driven-calls-human/">SiliconAngle</a>)</p></li><li><p><strong>OpenAI expands GPT-Live to Edu, Business, and Enterprise</strong> plans globally, bringing full-duplex voice with background reasoning to organizations. (<a href="https://thewincentral.com/gpt-live-chatgpt-voice-edu-business-enterprise/">TheWinCentral</a>)</p></li><li><p><strong>Qwen Audio 3.0 Realtime Plus tops OpenAI</strong> on the Artificial Analysis Speech-to-Speech Index at 84.1% vs. GPT-Realtime-2.1&#8217;s 79.1%, a first for Alibaba. (<a href="https://betanews.com/article/qwen-audio-3-vs-openai-speech-benchmark/">BetaNews</a>)</p></li><li><p><strong>Sarvam AI announces a trillion-parameter model</strong> built in India, with pricing 5.5x cheaper than GPT-5.4 Mini for coding and research workloads. (<a href="https://www.freepressjournal.in/tech/sarvam-bets-big-on-indias-ai-future-with-trillion-parameter-model-will-be-priced-five-times-cheaper-than-global-rivals">Free Press Journal</a>)</p></li><li><p><strong>Krafton releases A.X K2 Raon-Speech</strong>, a 21B-parameter audio model that processes speech directly to preserve emotion and tone, ranked first among Korean models. (<a href="https://www.theinvestor.co.kr/article/10824774">The Investor</a>)</p></li><li><p><strong>Boson AI unveils Higgs RealTime</strong> for full speech-to-speech processing that skips text conversion to preserve vocal nuance, led by ex-AWS scientist Alex Smola. (<a href="https://cryptobriefing.com/boson-ai-higgs-realtime-voice-model/">CryptoBriefing</a>)</p></li><li><p><strong>8x8 extends AI across its full CX platform</strong> beyond the contact center, adding conversation intelligence, agent development, and workforce management for every team. (<a href="https://www.cmswire.com/contact-center/8x8-extends-ai-across-its-cx-platform-to-all-teams/">CMSWire</a>)</p></li><li><p><strong>NICE and RingCentral expand into a bi-directional partnership</strong>, combining UCaaS, CCaaS, and AI in a single offering with mutual resale. (<a href="https://telecomreseller.com/2026/07/29/ringcentral-nice/">Telecom Reseller</a>)</p></li><li><p><strong>Parlance adds instant self-service config and AI call summaries</strong> for healthcare contact centers, now deployed across 400+ health systems. (<a href="https://www.prnewswire.com/news-releases/parlance-gives-healthcare-contact-centers-instant-configuration-control-and-ai-briefed-agent-transfers-302839592.html">PR Newswire</a>)</p></li><li><p><strong>Tata Communications launches voice AI for India&#8217;s 63M SMBs</strong> with speech-to-speech engagement, sub-500ms latency, and multilingual support via Tata Tele. (<a href="https://inc42.com/buzz/tata-communications-expands-voice-ai-push-eyes-smb-adoption/">Inc42</a>)</p></li><li><p><strong>Tesla rolls out Grok voice assistant in India</strong> for the Model Y via OTA update, supporting Hindi and five other Indian languages for navigation and queries. (<a href="https://www.indianweb2.com/2026/08/tesla-model-y-india-launch-adds-grok-ai.html">IndianWeb2</a>)</p></li><li><p><strong>Japan&#8217;s Justice Ministry backs voice rights protection</strong> from AI, concluding that publicity rights can address unauthorized voice cloning without new legislation. (<a href="https://english.kyodonews.net/articles/-/80955">Kyodo News</a>)</p></li><li><p><strong>Tysa, Conectys&#8217; AI voice agent, marks one year</strong> of handling global customer conversations across multiple languages and use cases. (<a href="https://www.einpresswire.com/article/923836012/tysa-turns-one-conectys-ai-voice-agent-marks-a-year-of-global-conversation">EIN Presswire</a>)</p></li></ul><div><hr></div><h2>Last Week Podcast</h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;2841f5d1-52a1-42d9-aea6-4094ef1add61&quot;,&quot;caption&quot;:&quot;In the Future of Voice AI series of interviews, I ask three questions to my guests: - What problems do you currently see in Enterprise Voice AI? - How does your company solve these problems? - What solutions do you envision in the next 5 years?&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;AI, Fraud, and the Human Contact Center | Geoff Burbridge (Principal + Founder @ Human Edge Advisory)&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:32916364,&quot;name&quot;:&quot;Davit Baghdasaryan&quot;,&quot;bio&quot;:&quot;CEO &amp; Co-Founder of Krisp, early pioneer in Voice AI.\n20+ years in engineering. 18 US patent applications, ex Twilion&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/23088dde-6cb0-44df-b220-5f22830cdd4c_1179x960.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-30T14:25:31.976Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3c2468c4-ac27-42c4-b7ee-68fb7eb87ca3_1165x776.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/p/ai-fraud-and-the-human-contact-center&quot;,&quot;section_name&quot;:&quot;Podcast&quot;,&quot;video_upload_id&quot;:&quot;beed0608-56ff-4a2c-b3d4-76fecf415753&quot;,&quot;id&quot;:207828532,&quot;type&quot;:&quot;podcast&quot;,&quot;reaction_count&quot;:13,&quot;comment_count&quot;:0,&quot;publication_id&quot;:2073467,&quot;publication_name&quot;:&quot;Voice AI Newsletter&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!YLgs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F831a2f7e-d0a7-4e3d-87a8-c42c65d0b71c_1000x1000.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>HackerNoon breaks down voice-to-voice AI architectures</strong>, from codec tokenization and RVQ prediction to full-duplex designs like Moshi and GPT-Live. (<a href="https://hackernoon.com/how-modern-voice-to-voice-ai-models-work">HackerNoon</a>)</p></li></ul><ul><li><p><strong>A deep dive into ARK-ASR-3B&#8217;s architecture</strong> shows how combining a Whisper encoder with a Qwen decoder achieves 5.04% WER on the Open ASR Leaderboard. (<a href="https://hackernoon.com/inside-ark-asr-3bs-whisper-and-qwen-architecture">HackerNoon</a>)</p></li><li><p><strong>Building SeMamba for speech enhancement</strong> walks through using Mamba state-space models to recover clean speech from noisy audio in a single forward pass. (<a href="https://levelup.gitconnected.com/from-noise-to-clarity-building-semamba-for-speech-enhancement-2e534e62dd7e">Level Up Coding</a>)</p></li><li><p><strong>Deepgram integrates with AWS SageMaker AI</strong> via IAM temporary delegation, letting enterprises self-host Nova and Aura-2 with zero standing cross-account access. (<a href="https://aws.amazon.com/blogs/machine-learning/deepgram-enhances-amazon-sagemaker-ai-support-with-aws-iam-temporary-delegation/">AWS</a>)</p></li><li><p><strong>A Frontiers in AI paper compares Wav2Vec and Whisper</strong> for Telugu speech recognition, benchmarking pretrained models on low-resource Indian languages. (<a href="https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2026.1878977/full">Frontiers</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI, Fraud, and the Human Contact Center | Geoff Burbridge (Principal + Founder @ Human Edge Advisory)]]></title><description><![CDATA[Watch now | In the Future of Voice AI series of interviews, I ask three questions to my guests: - What problems do you currently see in Enterprise Voice AI? - How does your company solve these problems? - What solutions do you envision in the next 5 years?]]></description><link>https://voice-ai-newsletter.krisp.ai/p/ai-fraud-and-the-human-contact-center</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/ai-fraud-and-the-human-contact-center</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Thu, 30 Jul 2026 14:25:31 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207828532/b65adac707fa211808bf12cef85fe568.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<pre><code><code>In the Future of Voice AI series of interviews, I ask three questions to my guests:

- What problems do you currently see in Enterprise Voice AI?
- How does your company solve these problems?
- What solutions do you envision in the next 5 years?</code></code></pre><p>This episode&#8217;s guest is <a href="https://www.linkedin.com/in/geoffreyburbridge/">Geoff Burbridge</a>, Principal + Founder at <a href="https://www.humanedgeadvisory.co/">Human Edge Advisory</a>.</p><p>Geoff Burbridge is Principal + Founder of Human Edge Advisory, where he helps executive teams, boards, and founders transform organizations at the intersection of people, process, and technology. With more than 30 years of executive leadership experience at Bank of America, USAA, Truist, and Capital One, Geoff advises organizations on AI readiness, operating model design, and enterprise transformation, helping leaders balance innovation with the security, governance, and trust required by today's enterprises. He is a keynote speaker, executive moderator, and board advisor who believes that as technology continues to evolve, the greatest competitive advantage will always be the Human Edge.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gd_K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gd_K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 424w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 848w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gd_K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png" width="1200" height="1200" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1200,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:303874,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/207828532?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gd_K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 424w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 848w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.youtube.com/@futureofvoiceai&quot;,&quot;text&quot;:&quot;Listen on YouTube&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.youtube.com/@futureofvoiceai"><span>Listen on YouTube</span></a></p><h3><strong>Recap Video</strong></h3><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;3a642a58-3752-461e-ab23-95ac30014d0c&quot;,&quot;duration&quot;:null}"></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive weekly updates.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><strong>Takeaways</strong></h3><ul><li><p>AI&#8217;s biggest contact center impact may be stopping fraud, not replacing agents.</p></li><li><p>The first AI strategy will often be wrong because the market, not the roadmap, reveals the real use case.</p></li><li><p>BPOs may emerge as AI winners, not casualties, because they can apply automation across massive service operations.</p></li><li><p>The industry has already retreated from &#8220;replace the agent&#8221; to &#8220;make the agent better.&#8221;</p></li><li><p>AI adoption will stall when companies automate tasks but ignore how the workforce itself must change.</p></li><li><p>The real AI risk is not failure, but destroying customer confidence before the technology earns their trust.</p></li><li><p>Fraudsters face don&#8217;t the same rules as enterprises, giving attackers a speed advantage that regulation alone cannot solve.</p></li><li><p>Voice fraud is becoming an AI-versus-AI fight that humans cannot manage at machine speed.</p></li><li><p>Most contact centers already have the signals to stop fraud but cannot connect them fast enough to act.</p></li><li><p>Authentication failure is weak protection when attackers can rapidly change devices, numbers, identities, and voices.</p></li><li><p>The customer call companies want to eliminate may be the exact interaction that stops a major fraud loss.</p></li><li><p>Human trust makes voice service powerful, but it also makes voice one of the easiest channels to exploit.</p></li><li><p>Deepfake scams succeed by creating urgency before the victim has time to question what sounds real.</p></li><li><p>The best fraud systems will not just detect known attacks; they will identify new attack patterns as they form.</p></li><li><p>Agentic AI is more valuable for absorbing demand spikes than replacing every customer conversation.</p></li><li><p>The best automation targets are processes that already reach the right outcome almost every time.</p></li><li><p>Repeatedly failed journeys need more human judgment, not another layer of automation.</p></li><li><p>AI coaching can scale practice without the pressure and resistance that make traditional role-play fail.</p></li><li><p>The intersectionality of AI is where its value compounds, connecting fraud signals, customer context, automation, and human judgment in real time.</p></li><li><p>AI&#8217;s real power comes from connecting fraud signals, customer context, agent judgment, and automation in the same moment.</p></li><li><p>Contact centers overcomplicate service when customers fundamentally only care about four things: answering the phone, being nice, resolution, and ease.</p></li><li><p>In Geoff&#8217;s experience, authentic service earns trust because customers believe people, not scripts.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[ChatGPT Voice Hits Desktop, Claude Voice Goes Full Opus]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/chatgpt-voice-hits-desktop-claude</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/chatgpt-voice-hits-desktop-claude</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 27 Jul 2026 14:02:05 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/de5511dc-b81f-4af4-80d6-41f829aa1efd_1574x876.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>Vapi and Inngest build a voice agent live,</strong> taking it from demo to production with retries, long-running execution, and evals along the way. (Jul 29, Hybrid, <a href="https://www.voiceaispace.com/events/from-voice-demo-to-production-agent-building-reliable-ai-workflows">Voice AI Space</a>)</p></li><li><p><strong>Temporal hosts a durable multimodal AI meetup with HeyGen, Vapi, and Modal,</strong> including a Vapi talk on scaling voice agents from one call to ten thousand. (Jul 29, SF, <a href="https://luma.com/durable-ai-july">LUMA</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>OpenAI brings ChatGPT Voice to desktop</strong> with GPT-Live-powered agents that can control apps, dictate in any window, and work with Codex on Mac and Windows. (<a href="https://techcrunch.com/2026/07/24/openais-new-voice-mode-makes-it-to-the-chatgpt-desktop-app/">TechCrunch</a>)</p></li><li><p><strong>Anthropic upgrades Claude voice mode</strong> to Opus and Sonnet models with cross-app automation for Gmail, Slack, and Notion in 10 languages. (<a href="https://www.webpronews.com/anthropic-arms-claude-voice-mode-with-opus-and-sonnet-brains/">WebProNews</a>)</p></li><li><p><strong>Alibaba launches Qwen Audio 3.0 TTS</strong> in Flash and Plus tiers across 16 languages, reaching #1 on the Artificial Analysis TTS arena at 1,236 Elo. (<a href="https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/">MarkTechPost</a>)</p></li><li><p><strong>ByteDance releases Seed Audio 1.0</strong>, a unified model for voice, music, and sound effects with zero-shot voice cloning from a single reference clip. (<a href="https://seed.bytedance.com/en/seedaudio1_0">ByteDance</a>)</p></li><li><p><strong>Deepgram puts Nova-3 on Snapdragon</strong> with on-device voice AI optimized for Qualcomm&#8217;s Hexagon NPU, eliminating the need for cloud connectivity. (<a href="https://finance.yahoo.com/technology/ai/articles/deepgram-delivers-real-time-voice-123000272.html">Yahoo Finance</a>)</p></li><li><p><strong>Telli raises $15M seed</strong> from Redalpine, Y Combinator, and Cherry Ventures to build AI agents that replace traditional call center operations. (<a href="https://www.pymnts.com/news/investment-tracker/2026/startup-telli-raises-15-million-to-replace-call-centers-with-ai-agents/">PYMNTS</a>)</p></li><li><p><strong>Valence AI raises $5M</strong> backed by SRI International to integrate real-time emotion detection into voice AI with US patents on emotional intelligence. (<a href="https://www.sri.com/press/story/sri-backed-valence-ai-raises-5m-to-integrate-emotional-intelligence-into-the-trust-stack/">SRI</a>)</p></li><li><p><strong>Cast Insights raises $4.5M pre-seed</strong> for real-time speech intelligence across TV, radio, and podcasts, having processed over 2.3M hours of audio. (<a href="https://siliconangle.com/2026/07/23/real-time-speech-intelligence-startup-cast-insights-raises-4-5m-track-worlds-tv-radio-podcasts/">SiliconAngle</a>)</p></li><li><p><strong>Zoom launches real-time voice translation</strong> that converts spoken language live during meetings across five languages. (<a href="https://www.zoom.com/en/blog/voice-translation-zoom/">Zoom</a>)</p></li><li><p><strong>Zoom Scribe adds speech accessibility</strong> features for real-time captioning and transcription to improve meeting inclusivity. (<a href="https://www.zoom.com/en/blog/zoom-scribe-speech-accessibility/">Zoom</a>)</p></li><li><p><strong>Otter.ai introduces Live Assist</strong>, a real-time coaching agent that listens to calls and provides guidance customizable with team playbooks. (<a href="https://www.businesswire.com/news/home/20260721446216/en/Otter.ai-Introduces-Live-Assist-The-First-Live-Coaching-Agent-for-Every-Call">BusinessWire</a>)</p></li><li><p><strong>Level AI launches Latitude</strong>, a suite of 7 purpose-built CX models that the company claims are 49x cheaper than frontier LLMs for contact centers. (<a href="http://www.smartcustomerservice.com/Articles/News-Briefs/Level-AI-Launches-Latitude-a-Suite-of-7-AI-Models-175782.aspx">Smart Customer Service</a>)</p></li><li><p><strong>Apple rolls out Genius Bar Live Notes</strong>, using AI to transcribe and summarize in-store customer appointments in real time. (<a href="https://hothardware.com/news/apple-genius-bar-ai-recording-tool">HotHardware</a>)</p></li><li><p><strong>Ukraine adds voice AI to its Diia government app</strong>, letting citizens access 170+ public services through ElevenLabs-powered spoken conversations. (<a href="https://www.smartcitiesworld.net/news/ukraine-adds-voice-ai-to-national-government-assistant">Smart Cities World</a>)</p></li><li><p><strong>AudioCodes targets 40-50% voice AI growth</strong> in 2026 with its conversational AI segment on track to reach $25M, aiming for $80M by 2028. (<a href="https://www.channelinsider.com/ai/building-channel-revenue/audiocodes-voice-ai-opportunity/">Channel Insider</a>)</p></li><li><p><strong>Parlance nears 2 billion healthcare calls</strong> processed by its voice AI platform, marking a major deployment milestone in patient-facing automation. (<a href="https://www.prnewswire.com/news-releases/2-billion-calls-parlance-voice-ai-nears-healthcare-milestone-302832861.html">PR Newswire</a>)</p></li><li><p><strong>WhisperAI surpasses 330,000 users</strong> and launches a transcription API with real-time STT, crossing seven-figure ARR in under a year. (<a href="https://techbullion.com/whisperai-surpasses-330000-professionals-worldwide-and-launches-advanced-whisperai-transcription-api/">TechBullion</a>)</p></li><li><p><strong>XMOS unveils VocalFusion XVF3620</strong>, an AI voice processor combining on-chip noise reduction, beamforming, and echo cancellation in a single device. (<a href="https://audioxpress.com/news/xmos-unveils-vocalfusion-xvf3620-with-advanced-ai-voice-processing">audioXpress</a>)</p></li><li><p><strong>Synaptics launches Astra SR80</strong>, an always-on edge AI audio processor for voice capture, biometric auth, and agentic AI devices. (<a href="https://audioxpress.com/article/human-centric-devices-and-the-rise-of-always-on-edge-ai-audio">audioXpress</a>)</p></li><li><p><strong>Smallest AI&#8217;s TTS ranks #1 for Hindi</strong> in blind listening tests, with Lightning v3.1 preferred 76% of the time over OpenAI&#8217;s GPT-4o-mini-TTS. (<a href="https://www.northjersey.com/press-release/story/217025/smallest-ais-tts-ranked-the-top-real-time-voice-for-hindi-customer-support/">North Jersey</a>)</p></li><li><p><strong>MouthPad launches a $1,400 tongue-controlled trackpad</strong> alongside Vox, a $200 whisper-to-type wearable microphone for hands-free voice input. (<a href="https://www.morningstar.com/news/pr-newswire/20260722sf10265/mouthpad-the-first-ever-tongue-trackpad-launches-to-the-public-alongside-new-wearable-vox">Morningstar</a>)</p></li><li><p><strong>Sweekar AI pocket pet uses staged voice development</strong> from babbling to fluent speech as a core growth mechanic, debuted at CES 2026 by Takway AI. (<a href="https://audioxpress.com/news/sweekar-ai-pocket-pet-turns-voice-interaction-into-a-growth-mechanism">audioXpress</a>)</p></li><li><p><strong>SoliderSound launches Go-Denoise</strong>, a desktop app using neural spectral processing to remove noise from speech and vocals in a single pass. (<a href="https://mixing.co.kr/en/42216">Monthly Mixing</a>)</p></li><li><p><strong>Forbes asks if AI is taking over audio</strong> as AI now hosts radio shows, narrates audiobooks, and generates podcasts, sparking debate over quality vs. human narration. (<a href="https://www.forbes.com/sites/frankracioppi/2026/07/22/is-ai-taking-over-audio---radio-podcasting-audiobooks/">Forbes</a>)</p></li><li><p><strong>Parloa argues CX leaders should stop measuring voice AI by deflection alone</strong>, pushing for resolution rate, customer effort, and escalation quality as better metrics. (<a href="https://www.cxtoday.com/contact-center/why-cx-leaders-should-stop-measuring-voice-ai-by-deflection-alone-parloa-cs-0228/">CX Today</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Cue AI runs Gemma 4 locally for voice dictation</strong>, cutting latency 44% and dropping marginal inference cost to zero on desktop voice agents. (<a href="https://deepmind.google/models/gemma/gemmaverse/cue-ai/">Google DeepMind</a>)</p></li></ul><ul><li><p><strong>Oracle builds a 24/7 healthcare voice agent</strong> using NVIDIA PersonaPlex and LiveKit on OCI, demonstrating full-duplex speech-to-speech for medical assistants. (<a href="https://blogs.oracle.com/cloud-infrastructure/healthcare-agent-nvidia-personaplex-livekit">Oracle</a>)</p></li><li><p><strong>MarkTechPost compares the best open ASR models of 2026</strong>, finding top models within one WER point while license and streaming become the real differentiators. (<a href="https://www.marktechpost.com/2026/07/23/best-open-speech-recognition-asr-models-in-2026-wer-languages-latency-and-license-compared/">MarkTechPost</a>)</p></li><li><p><strong>The voice agent latency playbook</strong> argues that input accuracy is a latency feature since wrong transcripts turn one conversation turn into three. (<a href="https://hackernoon.com/the-voice-agent-latency-playbook-stt-turn-detection-and-the-tradeoffs-nobody-talks-about">HackerNoon</a>)</p></li><li><p><strong>AssemblyAI details the future of real-time STT</strong> with its Universal-3.5 Pro model, semantic turn detection, and a single-WebSocket Voice Agent API. (<a href="https://www.assemblyai.com/blog/realtime-future-of-speech-to-text">AssemblyAI</a>)</p></li><li><p><strong>A dev.to guide on voice agent turn-taking</strong> explains how to keep AI calls under 600ms end-to-end with VAD, barge-in handling, and streaming pipelines. (<a href="https://dev.to/jackm-singularity/voice-agent-turn-taking-stop-live-ai-calls-from-talking-over-users-590b">dev.to</a>)</p></li><li><p><strong>NVIDIA Nemotron 3.5 ASR runs 14x faster than Whisper</strong> on a CPU-only laptop via parakeet.cpp, transcribing a 1m46s clip in 32 seconds vs. Whisper&#8217;s 7m34s. (<a href="https://medium.com/@hellorahulk/running-nvidia-nemotron-3-5-asr-locally-with-parakeet-cpp-and-how-it-beat-whisper-on-my-laptop-42105b504307">Medium</a>)</p></li><li><p><strong>OpenWhispr is an open-source voice dictation app</strong> supporting local Whisper and NVIDIA Parakeet models with zero telemetry, now at 2,100+ GitHub stars. (<a href="https://github.com/OpenWhispr/openwhispr">GitHub</a>)</p></li><li><p><strong>HackerNoon explains why TTS evaluation needs human ears</strong>, noting that WER and CER miss intonation and naturalness, requiring manual listening across checkpoints. (<a href="https://hackernoon.com/in-text-to-speech-your-ears-always-matter-more-than-your-metrics">HackerNoon</a>)</p></li><li><p><strong>A developer narrates their blog with Kokoro</strong>, an 82M-parameter local TTS model that generates two hours of audio in 15 minutes on an M1 MacBook. (<a href="https://bart.degoe.de/narrating-your-blog-with-kokoro-a-local-opensource-model/">Bart de Goede</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[What CCW 2026 told us about Voice AI in CX]]></title><description><![CDATA[CCW Vegas had no shortage of Voice AI.]]></description><link>https://voice-ai-newsletter.krisp.ai/p/what-ccw-2026-told-us-about-voice</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/what-ccw-2026-told-us-about-voice</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Thu, 23 Jul 2026 14:30:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bscE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bscE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bscE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bscE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bscE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bscE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bscE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9575439,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/206995171?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bscE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bscE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bscE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bscE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>CCW Vegas had no shortage of Voice AI. What stood out was how little patience buyers have left for vague promises.</p><p>The conversation has shifted from <em><strong>what can AI do?</strong></em> to <em><strong>what has it actually done?</strong></em></p><p><span>This was the year the industry stopped debating AI and started auditing it. </span>Here is what stood out most.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>1. AI has entered its proof era</h2><p>The AI market was crowded, but not always convincing. Many vendors claimed leadership, while case studies were thin and demos were limited. Buyers weren&#8217;t short on options, they were short on proof.</p><p>The companies that stood out immediately show what their solutions change: lower handle time, faster resolution, lower cost, or less work for the agent.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZEft!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZEft!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZEft!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7762708,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/206995171?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZEft!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Takeaway:</strong> In a crowded market, credibility comes from showing results buyers can see, measure, and trust &#8212; and showing them upfront.</p><h2>2. The best AI removes friction, not people</h2><p>Despite all the talk about Voice AI eating CX, the industry is still growing. CCW felt bigger, busier, and more confident than ever.</p><p>The strongest signal from both BPOs and enterprises was that human-led and AI-powered CX will grow together.</p><p>Damola Adamolekun, CEO of Red Lobster, offered one of the clearest operating principles of the week in his keynote:</p><blockquote><h4>&#8220;Focus on the changes with the highest impact and lowest effort.&#8221;</h4></blockquote><p><strong>Takeaway:</strong> AI&#8217;s near-term value is quick wins and removing the friction that holds companies and people back.</p><h2>3. Voice Translation is changing the economics of language support</h2><p>Multilingual conversations weren&#8217;t about features; they were about replacing a broken model: customer needs support, agent dials an interpreter, everyone waits. Interpreters are expensive, hard to staff, and regulated language requirements keep expanding.</p><p>The strongest proof came from production. On the main stage, Automated Health Systems shared how deploying Voice Translation reduced average multilingual calls from more than <strong>40 minutes to just 9</strong>, expanding language access while dramatically reducing operational costs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A4fr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A4fr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A4fr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8612140,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/206995171?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A4fr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Takeaway:</strong> Translation isn&#8217;t just improving access &#8212; it&#8217;s changing the cost, speed, and scale of multilingual service.</p><h2>4. Voice AI is CX infrastructure</h2><p>The show floor made the shift hard to miss.</p><p>Cisco and Microsoft led with agentic workflows. Jabra focused more on analytics. Poly and Epos weren&#8217;t there.</p><p>At the same time, Voice AI companies had a much larger presence than in prior years.</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;bfe2d5a1-7d80-45bd-b9e8-bcf4aaa924bb&quot;,&quot;duration&quot;:null}"></div><p>The value is moving into what happens inside the conversation: improving audio, which improves everything downstream&#8212; translation, transcription, agent guidance, detecting risk, and powering AI interactions in real time.</p><p><strong>Takeaway:</strong> Voice AI is becoming core infrastructure that the rest of CX depends on.</p><h2>5. Voice security is urgent</h2><p>Deepfake detection drew real interest, including follow-up requests from major enterprises.</p><p>The risk is growing fast:</p><ul><li><p>Deepfake fraud attempts grew <strong>2,137% in three years</strong></p></li><li><p>AI-related cybercrime losses topped <strong>$893 million in 2025</strong></p></li><li><p>Humans detect synthetic voices correctly only about <strong>60% of the time</strong></p></li></ul><p>Most fraud controls still sit before or after the call, but most voice attacks happen during it.</p><p>As Geoff Burbridge, Founder of Human Edge Advisory, put it:</p><blockquote><p>&#8220;AI is being used by bad actors out there, and newsflash, they&#8217;re not restrained. They&#8217;re not worried about regulation.&#8221;</p></blockquote><p><strong>Takeaway:</strong> Voice security is moving from future risk to an urgent buying requirement.</p><p><strong>Sneak peek:</strong> Geoff joins the <em>Voice AI Podcast</em> next week to break down where synthetic voice fraud is hitting hardest, how real-time detection works in production, and what companies need to do now.</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;bbb86917-86e0-48c8-a04f-d10ecfacaf54&quot;,&quot;duration&quot;:null}"></div><h2>A milestone worth sharing</h2><p>Krisp was named <strong>Disruptive Technology of the Year</strong> at the 2026 CCW Excellence Awards.</p><p><span>We&#8217;re honored by the recognition, but more importantly, it reflects where the industry is heading. </span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bgWL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bgWL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bgWL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png" width="260" height="260" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1080,&quot;width&quot;:1080,&quot;resizeWidth&quot;:260,&quot;bytes&quot;:130971,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/206995171?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bgWL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The technologies earning attention today aren&#8217;t chasing novelty. They&#8217;re solving real operational problems and delivering measurable results. </span></p><p><span>That was the story across the show floor, and one we&#8217;re proud to be part of.</span></p><h2>What comes next</h2><p>By CCW 2027:</p><ul><li><p><strong>Execution starts separating leaders from followers:</strong> Organizations that deployed AI in 2025-2026 will have results to show:</p><ul><li><p>Those that redesigned around the technology will have compounding returns.</p></li><li><p>Those that layered it on top of old structures might be starting over.</p></li></ul></li><li><p><strong><span>AI autonomy:</span></strong><span> the Organizations moving fastest are starting with agent assist, proving value in production, and expanding from there. Deploying autonomous AI before you&#8217;ve built trust into the system is how you lose both.</span></p></li><li><p><strong><span>Voice security becomes a procurement requirement:</span></strong><span> Operators asking questions now will have policies and vendor requirements in place, while those not paying attention will be responding to incidents.</span></p></li><li><p><strong><span>Multilingual support consolidates:</span></strong><span> Operators want fewer vendors and end-to-end accountability across language access, accent support, and translation.</span></p></li><li><p><strong><span>Agent experience reaches a tipping point:</span></strong><span> The industry can&#8217;t sustain current attrition rates. Companies that reduce friction and lower cognitive load, not just cut headcount, will have a structural talent advantage.</span></p></li><li><p><strong><span>Voice gets the engineering investment it has always deserved:</span></strong><span> The channel, technology, and infrastructure layer are finally getting serious attention. Long overdue.</span></p></li><li><p><strong>AI pressure will widen the BPO gap:</strong> BPOs that modernize their delivery model will pull ahead of those that don&#8217;t.</p></li></ul><p>The market wants proof, fewer points of failure, and technology that improves the conversation in real time. By next year, the winners will be the ones with the clearest outcomes.</p><p>See you there.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Rime Raises $24M, Meta Patents Voice Emotion Tracking]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/rime-raises-24m-meta-patents-voice</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/rime-raises-24m-meta-patents-voice</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 20 Jul 2026 14:02:26 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a40a3425-3b67-4604-acb5-7abd4cfcab4f_850x425.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>Deepgram hosts a fireside chat on the future of voice agents,</strong> evaluating agents beyond WER and whether cascaded STT-LLM-TTS still holds up. (Jul 23, London, <a href="https://www.voiceaispace.com/events/the-future-of-voice-agents-a-fireside-chat">Voice AI Space</a>)</p></li><li><p><strong>AssemblyAI demos Universal-3.5 Pro Realtime,</strong> its new model with context carryover and conversation memory, then opens up for a fireside chat with production voice agent builders. (Jul 23, SF, <a href="https://luma.com/m0thk5ai">Luma</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Alibaba launches Qwen Audio 3.0</strong> with real-time voice that can proactively use external tools, covering 113 languages for ASR and 36 for TTS. (<a href="https://www.kucoin.com/news/flash/aliyun-launches-qwen-audio-3-0-realtime-voice-ai-can-proactively-use-external-tools">KuCoin</a>)</p></li><li><p><strong>Meta patents an AI wearable</strong> that continuously analyzes voice to track the user&#8217;s emotional state, raising concerns under the EU AI Act&#8217;s emotion-inference ban. (<a href="https://thenextweb.com/news/meta-patent-mood-tracking-voice-emotion-ai">The Next Web</a>)</p></li><li><p><strong>PwC and OpenAI launch agentic customer service</strong> solutions combining PwC&#8217;s CX expertise with OpenAI&#8217;s multimodal voice and digital agent APIs. (<a href="https://www.cxtoday.com/ai-automation-in-cx/pwc-openai-agentic-customer-service/">CX Today</a>)</p></li></ul><ul><li><p><strong>Google Voice adds Gemini AI notes</strong> that auto-summarize calls with key points and action items, plus new standalone plans starting at $10/mo. (<a href="https://www.webpronews.com/google-voice-brings-gemini-ai-notes-to-calls-with-new-standalone-plans-at-lower-cost/">WebProNews</a>)</p></li><li><p><strong>Google quietly opted users</strong> into AI training on voice queries and uploaded media via a new Search Services History setting, with no opt-in required. (<a href="https://www.foxnews.com/tech/google-may-use-your-photos-voice-train-ai">Fox News</a>)</p></li><li><p><strong>Rime raises $24M Series A</strong> to build enterprise speech-to-speech models, powering nearly 100M phone calls monthly for Mayo Clinic, Dialpad, and others. (<a href="https://www.rime.ai/resources/rime-series-a-announcement">Rime</a>)</p></li></ul><ul><li><p><strong>LALAL.AI launches Lynx,</strong> a neural network built for speech denoising that is 6x smaller than its flagship model while matching output quality. (<a href="https://slator.com/lalal-ai-lynx-launch/">Slator</a>)</p></li><li><p><strong>Sber&#8217;s GigaChat adds emotion detection</strong> and can process audio up to three hours long with speaker separation, timestamps, and segment summaries. (<a href="https://businessnewsthisweek.com/business/gigachat-the-ai-assistant-detects-emotions-and-finds-content-in-long-audio-files/">BusinessNewsThisWeek</a>)</p></li><li><p><strong>DoorDash, ObserveAI, and AWS</strong> scale AI-powered quality evaluation across 19,000 agents, automating nearly 100% of interaction reviews. (<a href="https://www.prnewswire.com/news-releases/doordash-observeai-and-aws-partner-to-scale-customer-centric-ai-across-19-000-agents-302824595.html">PR Newswire</a>)</p></li><li><p><strong>Samsung adds cloud transcription</strong> to its Voice Recorder app, giving users a choice between on-device privacy and cloud-powered accuracy. (<a href="https://www.webpronews.com/samsung-voice-recorders-cloud-transcription-upgrade-signals-shift-in-on-device-ai-limits/">WebProNews</a>)</p></li><li><p><strong>Aina raises $5.5M</strong> to build hardware that controls AI agents rather than just recording, with its first product Dune already shipping to early adopters. (<a href="https://techcrunch.com/2026/07/16/ultrahumans-former-hardware-vp-raises-5-5m-for-devices-that-control-ai-agents-not-just-record-you/">TechCrunch</a>)</p></li><li><p><strong>Chen Institute and Science honor</strong> neuroscientist Sergey Stavisky for an AI speech neuroprosthesis that decodes brain activity into spoken words at 97.5% accuracy. (<a href="https://www.prnewswire.com/news-releases/ai-gives-people-back-their-own-voice-chen-institute-and-science-prize-honors-neuroscientist-sergey-stavisky-302827712.html">PR Newswire</a>)</p></li><li><p><strong>New research finds</strong> that AI voice phishing works because of persuasive scripting, not vocal realism, as 70% of targets detect the synthetic voice but comply anyway. (<a href="https://www.helpnetsecurity.com/2026/07/17/research-ai-voice-phishing/">Help Net Security</a>)</p></li><li><p><strong>Instadesk, Huawei, and iFlytek</strong> open a joint AI customer experience lab in Uzbekistan, combining multilingual ASR/TTS with Ascend cloud infrastructure. (<a href="https://www.manilatimes.net/2026/07/16/tmt-newswire/pr-newswire/instadesk-unveils-ai-customer-experience-lab-with-huawei-and-iflytek-in-central-asia/2385796">Manila Times</a>)</p></li><li><p><strong>VoicePing 3.0 launches</strong> with real-time translation, AI dubbing, new ASR and MT models, and MCP/API access for enterprise multilingual workflows. (<a href="https://www.prnewswire.com/news-releases/voiceping-releases-voiceping-3-0-for-enterprise-multilingual-communication-302823495.html">PR Newswire</a>)</p></li><li><p><strong>A study flags five risks</strong> in clinical AI scribes: inconsistent consent, weak performance on accented speech, background noise, missing human review, and unclear accountability. (<a href="https://www.resultsense.com/news/2026-07-17-clinical-ai-scribes-risks/">ResultSense</a>)</p></li><li><p><strong>Telcos are sitting on a voice AI opportunity</strong> bigger than their own cost savings, argues an analysis that says the real play is selling voice infrastructure to enterprises. (<a href="https://sebastianbarros.substack.com/p/the-telco-ai-voice-is-bigger-than">Sebastian Barros</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Apple&#8217;s SpeechAnalyzer API</strong> outperforms Whisper Small in English benchmarks, running fully on-device on iOS 26 and macOS Tahoe. (<a href="https://gigazine.net/gsc_news/en/20260714-apple-speech-analyzer-benchmark/#gsc.tab=0">Gigazine</a>)</p></li><li><p><strong>Cohere releases Transcribe Arabic,</strong> an open-source 2B-param ASR model that beats Meta&#8217;s 7B model on Arabic with multidialect and code-switching support. (<a href="https://cohere.com/blog/transcribe-arabic">Cohere</a>)</p></li><li><p><strong>Llamafile 0.10.4 ships transcribefile,</strong> a portable speech-to-text CLI built on transcribe.cpp that supports 16+ model families with GPU acceleration. (<a href="https://www.phoronix.com/news/Llamafile-0.10.4">Phoronix</a>)</p></li><li><p><strong>A prototype tongue-reading system</strong> uses ultrasound and ML to decode silent speech from tongue movements, enabling voice input without making a sound. (<a href="https://hackaday.com/2026/07/12/speak-silently-with-an-ultrasound-probe/">Hackaday</a>)</p></li><li><p><strong>Adafruit demos VAD, STT, and TTS</strong> all running on a single RP2040 microcontroller, bringing a complete voice AI pipeline to a $4 chip. (<a href="https://blog.adafruit.com/2026/07/15/voice-activity-detection-speech-to-text-and-text-to-speech-all-on-rp2040/">Adafruit</a>)</p></li><li><p><strong>ReSpeaker Clip</strong> is an open-source wearable AI recorder with dual mics, BLE 5.3, Wi-Fi 6, and full SDK for building custom voice AI applications. (<a href="https://www.seeedstudio.com/blog/2026/07/13/respeaker-clip-an-open-wearable-ai-recorder-for-building-voice-ai-applications/">Seeed Studio</a>)</p></li><li><p><strong>Resemble AI explains</strong> how neural audio watermarking embeds inaudible signals during voice generation for traceability, ahead of the EU AI Act&#8217;s August 2 deadline. (<a href="https://www.resemble.ai/resources/where-neural-audio-watermarking-fits-in-audio-security">Resemble AI</a>)</p></li><li><p><strong>A Python tutorial walks through</strong> building a real-time AI phone agent that quotes prices using tool calling and low-latency voice synthesis. (<a href="https://lowlatencyclub.ai/blog/posts/ai-price-quote-phone-agent-python">Low Latency Club</a>)</p></li><li><p><strong>Voice AI benchmarks hide a gap:</strong> systems scoring 600-800ms in the lab hit 2-4 seconds on real telephony, and accuracy drops from 51% to 26-38% under realistic audio. (<a href="https://embeddedcomputing.com/technology/ai-machine-learning/computer-vision-speech-processing/beyond-the-lab-why-voice-ai-must-be-benchmarked-under-real-world-conditions">Embedded Computing</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[GPT-Live Goes Full-Duplex, Taco Bell expands voice AI]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/gpt-live-goes-full-duplex-taco-bell</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/gpt-live-goes-full-duplex-taco-bell</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 13 Jul 2026 14:01:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/309dcc6d-4530-47b3-91e4-9e19e7019d6f_2408x1290.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong><a href="http://DeepLearning.AI">DeepLearning.AI</a> Voice AI Hackathon: </strong>In-person hackathon with Sabre + Vocal Bridge. Teams build voice agents that book real trips end-to-end, judged by Andrew Ng. (Jul 18, Mountain View, <a href="https://luma.com/fmypremp">Luma</a>)</p></li><li><p><strong>Voice Agent Evalathon: </strong>Okareo x Telnyx virtual challenge to red-team a voice agent&#8217;s reasoning, execution, and stability. (Jul 15, Virtual, <a href="https://www.meetup.com/ai-agent-simulation-reliability-group/events/315320251/">Meetup</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>OpenAI launches GPT-Live,</strong> a full-duplex voice model that listens and speaks simultaneously, powering a more natural ChatGPT Voice experience. (<a href="https://openai.com/index/introducing-gpt-live/">OpenAI</a>)</p></li><li><p><strong>xAI releases 21 multilingual flagship voices</strong> for Grok, expanding its voice lineup across multiple languages. (<a href="https://x.ai/news/new-flagship-voices">xAI</a>)</p></li><li><p><strong>Gradium raises $100M seed</strong> backed by NVIDIA, making the Paris-based Kyutai spinout one of the largest seed rounds ever for ultra-low latency voice AI. (<a href="https://techcrunch.com/2026/07/09/paris-based-ai-voice-startup-gradium-raises-100m-seed-backed-by-nvidia/">TechCrunch</a>)</p></li><li><p><strong>Rylo AI raises $85M</strong> to scale its real-time captioning platform for deaf and hard-of-hearing users, rebranded from Nagish. (<a href="https://www.alleywatch.com/2026/07/rylo-ai-communication-accessibility-deaf-hard-of-hearing-real-time-captioning-platform-tomer-aharoni/">AlleyWatch</a>)</p></li><li><p><strong>Five9 launches next-gen Voice AI Agents</strong> built on a purpose-built agentic architecture (<a href="http://insidermonkey.com/blog/five9-inc-fivn-reveals-next-gen-voice-ai-agents-1797773/">Insider Monkey</a>)</p></li><li><p><strong>Taco Bell expands voice AI</strong> to 890+ drive-thrus across 38 states, powered by Omilia. (<a href="https://www.fermag.com/articles/taco-bell-expands-voice-ai-at-us-drive-thrus/">FER Magazine</a>)</p></li><li><p><strong>Omilia launches Lexis,</strong> a native generative TTS engine for enterprise CX with sub-45ms latency. (<a href="https://aithority.com/it-and-devops/cloud/omilia-launches-the-only-native-voice-in-enterprise-cx/">AIthority</a>)</p></li><li><p><strong>Dell AI Factory partners with Deepgram</strong> and Penguin Solutions to deliver enterprise-grade real-time voice AI infrastructure. (<a href="https://www.businessinsider.com/sc/dell-ai-factory-powers-real-time-voice-ai">Business Insider</a>)</p></li><li><p><strong>SoundHound&#8217;s OASYS platform</strong> wins &#8220;Agentic AI Company of the Year&#8221; as Q1 revenue hits $44M, up 52% year-over-year. (<a href="https://www.nasdaq.com/articles/soundhound-trends-voice-ai-oasys-and-rising-agentic-demand-today">Nasdaq</a>)</p></li><li><p><strong>CallTower partners with Sestek</strong> to add conversational AI, voice biometrics, and real-time translation to its CX portfolio. (<a href="https://telecomreseller.com/2026/07/10/calltower-expands-ai-for-cx-portfolio-through-strategic-partnership-with-sestek/">Telecom Reseller</a>)</p></li><li><p><strong>Whispp raises $5M</strong> to scale its on-device AI that reconstructs speech for people with voice disorders in real time. (<a href="https://pulse2.com/whispp-raises-e5-million-to-scale-real-time-on-device-voice-reconstruction-ai/amp/">Pulse 2.0</a>)</p></li><li><p><strong>AI voice agents boosted specialty care enrollment 340%</strong> in a peer-reviewed clinical study by RadiantGraph. (<a href="https://www.prnewswire.com/news-releases/study-finds-ai-voice-agents-increased-specialty-care-program-enrollment-rates-340-in-real-world-clinical-setting-302818900.html">PR Newswire</a>)</p></li><li><p><strong>A $25M deepfake scam at Arup</strong> used AI-generated executives on a video call, becoming a landmark case for corporate voice fraud. (<a href="https://financefeeds.com/the-25-million-ai-deepfake-scam-that-changed-corporate-security/">FinanceFeeds</a>)</p></li><li><p><strong>Reality Defender warns</strong> that autonomous AI callers can now bypass contact center defenses at scale, posing a new voice fraud threat. (<a href="https://www.biometricupdate.com/202607/the-agentic-caller-always-rings-at-scale-reality-defender-explores-new-ai-voice-threat">Biometric Update</a>)</p></li><li><p><strong>Hydaway launches RealityChek,</strong> a streaming audio deepfake detector for enterprises that flags synthetic speech in real time. (<a href="https://www.biometricupdate.com/202607/hydaway-introduces-real-time-enterprise-audio-deepfake-detection">Biometric Update</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>OpenAI ships GPT-Realtime-2.1-mini</strong> with reasoning and tool use support at 6x lower cost and 25% reduced latency. (<a href="https://www.marktechpost.com/2026/07/06/openai-gpt-realtime-2-1-mini-reasoning-realtime-api/">Marktechpost</a>)</p></li><li><p><strong>AssemblyAI details how Universal-3.5 Pro</strong> handles noisy audio, sharing techniques for improving transcription accuracy on hard recordings. (<a href="https://www.assemblyai.com/blog/async-transcription-accuracy-hard-audio">AssemblyAI</a>)</p></li><li><p><strong>Flock Safety explains its audio detection system</strong> that identifies gunshots and crashes in real time using acoustic sensors and AI classification. (<a href="https://www.flocksafety.com/blog/how-flocks-audio-detection-works">Flock Safety</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[xAI Ships Voice Agent Builder, Krisp named 2026 Disruptive Technology]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/xai-ships-voice-agent-builder-krisp</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/xai-ships-voice-agent-builder-krisp</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 06 Jul 2026 14:02:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/68229207-8166-4e14-a810-1bea7989b6b4_1426x740.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>Real-Time Observability for Production Voice AI</strong> Regal walks through its Observability Dashboard with real catch-and-fix examples from production voice agents. (Jul 9, SF, <a href="https://www.voiceaispace.com/events/do-you-know-what-your-ai-agent-is-doing-real-time-observability-for-production-voice-ai">Voice AI Space</a>)</p></li><li><p><strong>AI Tinkerers San Francisco: July GTM Engineering Track w/ Attio</strong> Builder-only, no-pitch meetup with live GTM engineering demos - a solid room for voice agent developers. (Jul 8, SF, <a href="https://sf.aitinkerers.org/p/ai-tinkerers-san-francisco-july-gtm-engineering-track-w-attio">AI Tinkerers</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>xAI launches Voice Agent Builder</strong>, a no-code platform for production voice agents with telephony and guardrails at $0.05/min. (<a href="https://x.ai/news/grok-voice-agent-builder">xAI</a>)</p></li></ul><ul><li><p><strong>Krisp named 2026 Disruptive Technology</strong> of the Year by CMP Research for its voice AI infrastructure. (<a href="https://krisp.ai/blog/krisp-named-2026-disruptive-technology-of-the-year/">Krisp Blog</a>)</p></li><li><p><strong>ElevenLabs explores a $22B tender offer</strong>, doubling its valuation from the $11B Series D five months ago. (<a href="https://www.techinasia.com/news/voice-ai-startup-elevenlabs-seeks-22b-sources">Tech in Asia</a>)</p></li><li><p><strong>Pocket raises $11M</strong> from Accel and YC for its $129 AI note-taking puck that has shipped 130K units. (<a href="https://techcrunch.com/2026/06/29/pocket-raises-11m-in-bet-on-rising-demand-for-ai-note-taking-devices/">TechCrunch</a>)</p></li><li><p><strong>Lucida AI raises &#8364;6.1M</strong> seed for its speech-to-speech language coaching platform, now at 3M users. (<a href="https://www.eu-startups.com/2026/06/uk-speech-ai-startup-lucida-ai-lands-e6-1-million-to-develop-speech-native-ai-for-global-communication/">EU-Startups</a>)</p></li><li><p><strong>US senators revive the AI Labeling Act</strong>, a bipartisan bill requiring AI-generated audio and video to carry disclosure labels. (<a href="https://www.musicbusinessworldwide.com/us-senators-revive-bill-to-force-ai-generated-audio-video-and-images-to-carry-labels/">Music Business Worldwide</a>)</p></li><li><p><strong>Retell AI launches Conductor</strong>, a graph-native review interface with an AI copilot for production voice agents. (<a href="https://aithority.com/machine-learning/voice-ai-startup-retell-ai-launches-conductor-featuring-the-first-ever-graph-native-review-interface-for-production-voice-agents/">AIthority</a>)</p></li><li><p><strong>Syntiant and Vibe partner</strong> to bring voice-enabled AI to smart workspace hardware using edge AI chips. (<a href="https://www.manilatimes.net/2026/06/30/tmt-newswire/globenewswire/syntiant-and-vibe-collaborate-to-advance-voice-enabled-ai-driven-workspace-experiences/2375782">GlobeNewsWire</a>)</p></li><li><p><strong>HealthLynked launches</strong> an AI healthcare platform with 24/7 scheduling and medical office voice agents. (<a href="https://www.globenewswire.com/news-release/2026/06/29/3318905/0/en/healthlynked-launches-ai-powered-healthcare-communication-platform-featuring-24-7-appointment-scheduling-and-medical-office-ai-agents.html">GlobeNewsWire</a>)</p></li><li><p><strong>RevComm launches MiiTel for Retail</strong>, extending its voice AI analytics to in-store customer conversations. (<a href="https://jp.ibtimes.com/revcomm-launches-miitel-retail-store-voice-ai-102266">IBTimes JP</a>)</p></li><li><p><strong>Patient trust is the biggest barrier</strong> to healthcare voice AI, not the technology, argues a Forbes analysis. (<a href="https://www.forbes.com/councils/forbestechcouncil/2026/06/30/the-hardest-problem-in-healthcare-voice-ai-isnt-the-technology-its-patient-trust/">Forbes</a>)</p></li><li><p><strong>Voices similar to our own are more persuasive</strong>, finds new research raising concerns about companies weaponizing stored voice data. (<a href="https://nautil.us/the-dangers-of-ai-voice-clones-1282420">Nautilus</a>)</p></li><li><p><strong>Synthflow deployed voice AI in one day</strong> for Nellis Auction, now handling 80% of each customer interaction automatically. (<a href="https://cxm.world/customer-experience/the-phone-line-that-hung-up-on-customers-and-the-voice-ai-that-fixed-it-in-a-day/">CXM World</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Mozilla AI releases transcribe-cpp</strong>, an open-source C/C++ STT library with ggml runtime and GPU acceleration. (<a href="https://www.startuphub.ai/ai-news/technology/2026/mozilla-ai-unveils-transcribe-cpp">StartupHub</a>)</p></li></ul><ul><li><p><strong>ViiTorVoice-NAR goes open source</strong> with word-level TTS editing that swaps individual words without regenerating surrounding audio. (<a href="https://www.techtimes.com/articles/319524/20260702/text-speech-ai-edits-single-words-mid-recording-viitorvoice-goes-open-source.htm">TechTimes</a>)</p></li><li><p><strong>Higgs TTS 2 3B</strong> from BosonAI is a 5.8B-param TTS model trained on 10M+ hours with zero-shot voice cloning. (<a href="https://hackernoon.com/the-higgs-tts-2-3b-base-model-a-text-to-speech-foundation-model">HackerNoon</a>)</p></li><li><p><strong>Vowen 0.4.8 released</strong>, a free offline voice productivity app using Whisper-based local transcription. (<a href="https://www.warp2search.net/story/vowen-048-released/">Warp2Search</a>)</p></li><li><p><strong>WhisTam</strong> is a Whisper-based framework for Tamil dialect speech recognition, ranking 2nd at DravidianLangTech@ACL 2026. (<a href="https://aclanthology.org/2026.dravidianlangtech-1.70/">ACL Anthology</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[$200M Pours Into Voice AI, OpenAI Bidi-1 Leaks]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/200m-pours-into-voice-ai-openai-bidi</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/200m-pours-into-voice-ai-openai-bidi</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 29 Jun 2026 14:03:12 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e90060d3-0692-46ad-84f2-adf95038b653_1670x928.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>AI Engineer World&#8217;s Fair</strong> is a flagship AI engineering conference with a dedicated Voice &amp; Realtime AI miniconference featured this year (Jun 29-Jul 2, SF | <a href="https://www.ai.engineer/worldsfair/2026">AI Engineer</a>)</p></li><li><p><strong>Low Latency Lounge by Deepgram</strong> is an invite only evening for engineers building the fastest AI in the stack. Together AI and Runware are cohosting (Jun 30, SF | <a href="https://luma.com/low-latency-lounge">LUMA</a>)</p></li><li><p><strong>Real-Time Voice AI &#215; Device Builders Meetup</strong> &#8220;Give Voice to Robots!&#8221; Runs alongside IVS Kyoto (Jul 2, Kyoto | <a href="https://www.voiceaispace.com/events/-real-time-voice-ai--device-builders-meetup-kyoto">Voice AI Space</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>AssemblyAI launches</strong> Universal-3.5 Pro Realtime, the first streaming STT model that takes the agent&#8217;s question as input (<a href="https://www.assemblyai.com/blog/universal-3-5-pro-realtime">AssemblyAI Blog</a>)</p></li><li><p><strong>Five9 launches</strong> Voice AI Agents and AI Agent Studio at CCW, bringing agentic CX to enterprise contact centers. (<a href="https://www.cxtoday.com/ai-automation-in-cx/five9-voice-ai-agents-agentic-cx-launch/">CX Today</a>)</p></li><li><p><strong>Krisp launches</strong> Voice Security for deepfake detection and fraud detection for contact centers. (<a href="https://www.cxtoday.com/security-privacy-compliance/krisp-expands-contact-center-ai-platform-with-voice-security-and-speech-analytics/">CX Today</a>)</p></li><li><p><strong>CallMiner launches</strong> real-time AI guidance that lets contact center agents initiate AI assistance on demand with human-in-the-loop controls. (<a href="https://www.businesswire.com/news/home/20260622825087/en/CallMiner-Enhances-Real-Time-Agent-Performance-and-Customer-Experience-with-New-AI-Capabilities">BusinessWire</a>)</p></li><li><p><strong>Assort Health raises</strong> $120M Series C led by Menlo Ventures at a $1.2B valuation to scale its voice AI agent platform across healthcare. (<a href="https://www.fiercehealthcare.com/ai-and-machine-learning/assort-health-scores-120m-series-c-scale-voice-ai-agent-platform-healthcare">Fierce Healthcare</a>)</p></li><li><p><strong>Prosper AI raises</strong> $30M Series A led by a16z to scale its autonomous patient journey platform, reporting 5x revenue growth in six months. (<a href="https://hackernoon.com/prosper-ai-raises-$30m-led-by-a16z-to-scale-autonomous-patient-journey-platform">HackerNoon</a>)</p></li><li><p><strong>Coval raises</strong> $28M Series A led by Norwest to advance its voice AI evaluation and testing platform, founded by an ex-Waymo engineer. (<a href="https://pulse2.com/coval-raises-28-million-series-a-to-advance-voice-ai-evaluation-platform/">Pulse2</a>)</p></li><li><p><strong>Kotoba Technologies raises</strong> $10M seed led by Kindred Ventures for its real-time East Asian voice translation platform with sub-2s latency. (<a href="https://gamesbeat.com/kotoba-technologies-raises-10m-for-real-time-voice-ai-platform-in-east-asia/">VentureBeat</a>)</p></li><li><p><strong>Valence AI raises</strong> $5M seed and secures US patents on real-time emotional detection from live speech. (<a href="https://www.prnewswire.com/news-releases/valence-ai-raises-5-million-secures-us-patents-on-real-time-emotional-detection-from-live-speech-302808293.html">PR Newswire</a>)</p></li><li><p><strong>TELUS Digital partners</strong> with ElevenLabs as a preferred implementation partner to scale voice AI alongside frontline customer care teams. (<a href="https://www.prnewswire.com/news-releases/telus-digital-and-elevenlabs-partner-to-scale-voice-ai-alongside-frontline-customer-care-teams-882149628.html">PR Newswire</a>)</p></li><li><p><strong>OpenAI&#8217;s GPT-Bidi-1</strong> leaks as a full-duplex voice model that can listen and speak simultaneously, enabling true bidirectional conversation. (<a href="https://cryptobriefing.com/openai-chatgpt-bidi-1-voice-model/">Crypto Briefing</a>)</p></li><li><p><strong>Conduent unveils</strong> a next-gen CX platform with real-time translation across 90+ languages to accelerate agent performance. (<a href="https://www.news.conduent.com/news/conduent-introduces-ai-powered-next-generation-cx-platform-to-expand-global-customer-reach-and-accelerate-agent-performance">Conduent</a>)</p></li><li><p><strong>Speechify brings</strong> free voice typing to all iPhone and Mac users, adding AI-powered dictation across every app. (<a href="https://9to5mac.com/2026/06/23/speechify-brings-voice-typing-to-all-iphone-and-mac-users/">9to5Mac</a>)</p></li><li><p><strong>Modulate launches</strong> an AI music detection API with 95% precision across 76 genres to help platforms verify AI-generated music. (<a href="https://www.morningstar.com/news/accesswire/1181688msn/modulate-launches-ai-music-detection-api-to-help-platforms-verify-ai-generated-music-at-scale">Morningstar</a>)</p></li><li><p><strong>ByteDance releases</strong> Seed Audio 1.0, a unified model that generates speech, music, and ambient sound from a single architecture. (<a href="https://www.citybuzz.co/2026/06/25/seed-audio-1-0-launches-unified-ai-audio-generation-for-speech-music-and-ambient-sound/">CityBuzz</a>)</p></li><li><p><strong>Amazon launches</strong> Alexa Plus Hindi beta in India, targeting 600M+ Hindi speakers with its upgraded AI assistant. (<a href="https://thenextweb.com/news/amazon-alexa-plus-india-hindi-beta-testing">The Next Web</a>)</p></li><li><p><strong>ElevenLabs adopts</strong> Google&#8217;s SynthID watermarking to tag all AI-generated speech, making synthetic voices easier to detect. (<a href="https://www.digitaltrends.com/cool-tech/ai-voices-are-getting-harder-to-spot-this-elevenlabs-feature-could-change-that/">Digital Trends</a>)</p></li><li><p><strong>Shure says</strong> audio quality is now the critical bottleneck for AI-powered meetings, and microphone clarity drives everything. (<a href="https://www.inavateonthenet.net/features/article/shure-says-audio-is-now-critical-to-ai-meetings--and-clarity-is-everything">InAVate</a>)</p></li><li><p><strong>Attention Labs launches</strong> SAA, a selective auditory attention layer that lets voice AI detect when it is being directly addressed. (<a href="https://www.dispatch.com/press-release/story/203592/attention-labs-launches-saa-the-engagement-control-layer-that-lets-voice-ai-know-when-it-is-being-addressed/">Dispatch</a>)</p></li><li><p><strong>Deepgram and Fortanix</strong> partner to run voice AI on-premises with NVIDIA confidential computing, keeping audio data encrypted during processing. (<a href="https://radioinfo.com.au/audioinfo/technology-news/private-voice-ai-introduced-to-better-protect-your-audio/">RadioInfo</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Gradium releases</strong> STT-Translate and S2S-Translate, real-time speech translation models that beat GPT Realtime Translate on accuracy and latency. (<a href="https://www.marktechpost.com/2026/06/24/gradium-launches-stt-translate-and-s2s-translate-real-time-speech-translation-models-beating-gpt-realtime-translate-on-accuracy-and-latency/">MarkTechPost</a>)</p></li></ul><ul><li><p><strong>AWS publishes</strong> a full tutorial on building a healthcare appointment agent with Amazon Nova 2 Sonic and Bedrock AgentCore. (<a href="https://aws.amazon.com/blogs/machine-learning/build-a-healthcare-appointment-agent-with-amazon-nova-2-sonic/">AWS Blog</a>)</p></li><li><p><strong>AssemblyAI shares</strong> four techniques for prompting Claude to build production-ready voice agents in about 30 seconds. (<a href="https://www.assemblyai.com/blog/prompting-claude-build-voice-agents">AssemblyAI Blog</a>)</p></li><li><p><strong>Deepgram discusses</strong> voice AI infrastructure and the path to production-grade agents on the Telecom Reseller podcast. (<a href="https://telecomreseller.com/2026/06/24/deepgram-on-voice-ai-infrastructure-and-the-road-to-production-grade-agents-podcast/">Telecom Reseller</a>)</p></li><li><p><strong>ACL 2026 publishes</strong> 10 voice AI papers covering noise-robust ASR, accented speech recognition, environment-aware TTS, controllable speech synthesis, multi-speaker diarization, and multilingual translation. (<a href="https://aclanthology.org/volumes/2026.acl-long/">ACL Anthology</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Soniox launches v5, Bland raises $50M, Mistral ships Voxtral Transcribe 2 and more]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/bland-raises-50m-soniox-launches</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/bland-raises-50m-soniox-launches</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 22 Jun 2026 14:02:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1b7f88f1-7c0e-4241-afaf-7a4cb9b5dace_1298x828.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>Voice AI Meetup Madrid</strong> is a small gathering hosted by Deepgram, AWS, and Pipecat for founders and engineers building with voice AI in Spain (Jun 23, Madrid | <a href="https://www.pipecat.ai/events">Pipecat</a>)</p></li><li><p><strong>Boba-thon</strong> is a hands-on AI build night by AI Valley &#215; Workato - teams form, prototype AI workflows, and demo by end of night. (Jun 25, San Francisco | <a href="https://www.voiceaispace.com/events/boba-thon">Voice AI Space</a>)</p></li><li><p><strong>UK &amp; Ireland Speech Workshop</strong> brings together speech science researchers and industry builders around advances in healthcare speech tech. (Jun 22-24, London | <a href="https://sites.google.com/view/ukis2026/home">UKIS2026</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Bland raises</strong> $50M Series C led by Dell Technologies, bringing its total funding past $100M. (<a href="https://fortune.com/2026/06/16/voice-ai-bland-50-million-after-being-rejected-by-180-investors/">Fortune</a>)</p></li><li><p><strong>Soniox launches</strong> <strong>v5</strong> Real-Time and Async, a speech model that turns live conversations into structured, speaker-aware intelligence. (<a href="https://soniox.com/blog/soniox-v5-real-time">Soniox</a>)</p></li></ul><ul><li><p><strong>Google launches</strong> a $99 Gemini-powered Home Speaker, its first standalone smart speaker since the Nest Audio in 2020. (<a href="https://techcrunch.com/2026/06/17/google-bets-on-gemini-to-reinvent-the-smart-home-speaker/">TechCrunch</a>)</p></li><li><p><strong>Plaud crosses</strong> $100M ARR in two years, making it the fastest hardware-led AI company to hit that milestone. (<a href="https://itbrief.com.au/story/plaud-says-arr-jumps-to-usd-100-million-in-two-years">ITBrief</a>)</p></li><li><p><strong>Respond.io raises</strong> $62.5M Series B to expand its AI-powered customer messaging platform into North America and Europe. (<a href="https://martechseries.com/sales-marketing/messaging/respond-io-raises-62-5m-series-b-to-scale-ai-powered-customer-conversations-into-north-america-and-europe/">MarTech Series</a>)</p></li><li><p><strong>Poland invests</strong> $11M in ElevenLabs and launches AI Lab Poland to grow its national AI ecosystem. (<a href="https://mezha.ua/en/news/polshcha-vkladaye-11-mln-u-rozvitok-shi-startapu-elevenlabs-312384/amp/">Mezha</a>)</p></li><li><p><strong>Mistral ships</strong> Voxtral Transcribe 2, an open-source on-device ASR model with batch transcription at $0.003 per minute. (<a href="https://mistral.ai/news/voxtral-transcribe-2/">Mistral</a>)</p></li><li><p><strong>Gnani AI launches</strong> Prisma v2.5, ranking first in 8 of 9 Indian language ASR benchmarks against Sarvam and ElevenLabs. (<a href="https://www.medianama.com/2026/06/223-gnani-ai-prisma-v2-5-speech-recognition-model-better-accuracy-sarvam/">MediaNama</a>)</p></li><li><p><strong>Tencent Cloud and Inworld AI</strong> partner to integrate sub-130ms TTS into Tencent&#8217;s real-time communication infrastructure. (<a href="https://en.prnasia.com/releases/apac/tencent-cloud-and-inworld-ai-announce-strategic-partnership-to-deliver-a-one-stop-lifelike-realtime-voice-ai-solution-537364.shtml">PR Newswire Asia</a>)</p></li><li><p><strong>Tencent Cloud and Soniox</strong> partner to bring multilingual speech-to-text across 200+ countries via Tencent RTC. (<a href="https://futurecio.tech/tencent-cloud-and-soniox-partner-to-elevate-enterprise-voice-ai/">FutureCIO</a>)</p></li><li><p><strong>DeepL acquires</strong> Mixhalo&#8217;s ultra-low-latency audio team and technology to scale its real-time voice translation product. (<a href="https://www.wallstreet-online.de/nachricht/21013238-eqs-news-deepl-expands-into-silicon-valley-adds-mixhalo-team-and-technology-to-accelerate-voice-ai-at-scale">PR Newswire</a>)</p></li><li><p><strong>TELUS Digital and Cresta</strong> partner to deliver AI agents alongside human agents in enterprise contact centers. (<a href="https://www.prnewswire.com/news-releases/telus-digital-and-cresta-partner-to-deliver-ai-agents-and-augment-human-agents-to-elevate-customer-experience-857014786.html">PR Newswire</a>)</p></li><li><p><strong>Parloa becomes</strong> the first agentic AI provider on Alvaria&#8217;s outbound platform, targeting regulated industries. (<a href="https://www.prnewswire.com/news-releases/parloa-and-alvaria-set-to-revolutionize-proactive-support-with-industry-first-in-agentic-cx-302802254.html">PR Newswire</a>)</p></li><li><p><strong>LiveKit Inference</strong> now defaults to zero data retention, meaning prompts and audio are never stored by any model provider. (<a href="https://x.com/livekit/status/2067319738926006387?s=20">LiveKit</a>)</p></li><li><p><strong>AI fraud cost</strong> $442B globally in 2025 as voice clones now fool even experts, per an INTERPOL report. (<a href="https://www.techtimes.com/articles/318458/20260616/ai-fraud-cost-world-442-billion-last-year-voice-clones-now-fool-even-experts.htm">TechTimes</a>)</p></li><li><p><strong>UC study finds</strong> vocal similarity alone drives persuasion, with listeners complying more when a speaker&#8217;s voice matches theirs. (<a href="https://www.uc.edu/news/articles/2026/06/ai-voice-cloning-vocal-similarity-uc-study.html">UC News</a>)</p></li><li><p><strong>AI voice clones</strong> are up to 20% more intelligible than real humans in noisy environments, a new JASA study shows. (<a href="https://www.psypost.org/ai-voice-clones-are-easier-to-understand-in-noisy-environments-than-real-humans/">PsyPost</a>)</p></li><li><p><strong>India&#8217;s telecom layer</strong> needs rebuilding for voice AI to scale, with traditional infrastructure adding 300-500ms of latency. (<a href="https://inc42.com/buzz/why-rebuilding-telecom-infrastructure-is-critical-for-next-wave-of-voice-ai/">Inc42</a>)</p></li><li><p><strong>Multilingual voice AI</strong> is India&#8217;s next big opportunity, with 600M+ vernacular users driving enterprise demand. (<a href="https://www.expresscomputer.in/guest-blogs/why-multilingual-voice-ai-is-indias-next-big-opportunity/136033/">Express Computer</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>TowardsAI tutorial</strong> on using Gemini streaming TTS to make voice apps feel instant. (<a href="https://pub.towardsai.net/gemini-streaming-tts-how-developers-can-make-ai-voice-apps-feel-instant-01ef246f398e">Towards AI</a>)</p></li></ul><ul><li><p><strong>Dev.to walkthrough</strong> of building a voice AI platform with 28 modules in Python. (<a href="https://dev.to/ryanwinston_134/building-a-voice-ai-platform-with-28-modules-in-python-4hbm">Dev.to</a>)</p></li><li><p><strong>CTO field report</strong> on testing 184 AI text-to-speech models across quality, latency, and cost. (<a href="https://dev.to/gentleforge/i-tested-184-ai-text-to-speech-models-a-ctos-field-report-20h6">Dev.to</a>)</p></li><li><p><strong>Dev.to tutorial</strong> on simple text-to-speech in Python using PythonAIBrain. (<a href="https://dev.to/divyanshusinha136/text-to-speech-in-python-made-simple-with-pythonaibrain-2bc8">Dev.to</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Krisp Voice Translation v3, New Siri AI and more]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/krisp-voice-translation-v3-new-siri</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/krisp-voice-translation-v3-new-siri</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 15 Jun 2026 14:03:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/69211f54-fa73-41cd-bf6a-05bae5057589_1450x730.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Top Updates &#128170;</h2><ul><li><p><strong>Krisp ships</strong> Voice Translation v3 with 96% accuracy in 61 languages and opens a self-serve developer API. (<a href="https://krisp.ai/blog/krisp-launches-v3-real-time-voice-translation/">Krisp</a> | <a href="https://krisp.ai/blog/introducing-voice-translation-api/">Krisp</a>)</p></li><li><p><strong>Apple launches</strong> Siri AI at WWDC with multi-turn conversations and a standalone app powered by Gemini. (<a href="https://www.apple.com/newsroom/2026/06/apple-introduces-siri-ai-a-profoundly-more-capable-and-personal-assistant/">Apple</a>)</p></li></ul><ul><li><p><strong>Google launches</strong> Gemini 3.5 Live Translate for real-time speech translation across 70+ languages. (<a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/">Google</a>)</p></li><li><p><strong>Mistral is raising</strong> ~&#8364;3B at a &#8364;20B valuation, nearly doubling since its September round. (<a href="https://techcrunch.com/2026/06/12/mistral-is-rumored-to-be-raising-e3b-at-e20-valuation/">TechCrunch</a>)</p></li><li><p><strong>Equal AI raises</strong> $30M Series B to scale India&#8217;s voice-first AI assistant across a billion smartphones. (<a href="https://www.livemint.com/companies/start-ups/equal-ai-funding-series-b-prosus-ventures-tomales-bay-capital-consumer-ai-voice-ai-11781241081404.html">LiveMint</a>)</p></li><li><p><strong>NICE makes</strong> agentic AI the native architecture of its CX platform at NICE World 2026. (<a href="https://www.cmswire.com/contact-center/nice-makes-its-move-at-nice-world-2026-agentic-ai-is-now-the-architecture/">CMSWire</a>)</p></li><li><p><strong>Microsoft launches</strong> MAI-Voice-2, a TTS model supporting 10 languages and zero-shot voice cloning. (<a href="https://www.blockchain-council.org/ai/introducing-mai-voice-2/">Blockchain Council</a>)</p></li><li><p><strong>AI voice scams</strong> surged 1,210% in 2025, needing just 3 seconds of audio to clone any voice. (<a href="https://www.foxnews.com/tech/ai-voice-scams-clone-familys-voice">Fox News</a>)</p></li><li><p><strong>Google will save</strong> search images and audio by default for AI model training. (<a href="https://www.theverge.com/tech/947836/google-search-privacy-settings-images-audio">The Verge</a>)</p></li><li><p><strong>AI ambient scribes</strong> cut physician burnout by 21 percentage points in a Mass General Brigham study. (<a href="https://www.medicaldaily.com/ai-ambient-scribe-physician-burnout-mass-general-brigham-ucla-study-2026-475610">Medical Daily</a>)</p></li><li><p><strong>MindBio delivers</strong> AI voice kiosks that detect intoxication and fatigue from speech patterns. (<a href="https://www.streetwisereports.com/article/2026/06/12/mindbio-therapeutics-advances-ai-voice-tech-for-workplace-safety-in-growing-biotech-and-ai-detector-markets.html">StreetWise Reports</a>)</p></li><li><p><strong>Top Gear asks</strong> whether AI voice control in cars is the next big thing or a waste of time. (<a href="https://www.topgear.com/car-news/electric/ai-voice-control-cars-next-big-thing-or-a-complete-waste-time">Top Gear</a>)</p></li><li><p><strong>WSJ reports</strong> the job AI was supposed to kill now needs more humans than ever. (<a href="https://www.wsj.com/tech/ai/the-job-that-ai-was-supposed-to-kill-needs-more-humans-than-ever-0771e4cf">WSJ</a>)</p></li><li><p><strong>Voicegain hires</strong> a VP of Sales to push voice AI into healthcare call centers. (<a href="https://www.prweb.com/releases/voicegain-appoints-tracy-puleo-as-vice-president-of-sales-to-accelerate-voice-ai-growth-in-healthcare-call-centers-302793704.html">PRWeb</a>)</p></li><li><p><strong>Speechmatics named</strong> HackerNoon&#8217;s Company of the Week for speech AI innovation. (<a href="https://hackernoon.com/meet-speechmatics-hackernoon-company-of-the-week">HackerNoon</a>)</p></li><li><p><strong>Voice AI adoption</strong> crosses an enterprise threshold in contact centers with measurable ROI. (<a href="https://www.cxtoday.com/contact-center/why-voice-ai-adoption-is-accelerating-in-2026/">CXToday</a>)</p></li><li><p><strong>India positions</strong> itself as the world&#8217;s CX leader as voice AI reshapes its call center industry. (<a href="https://www.expresscomputer.in/news/indias-voice-ai-opportunity-from-the-worlds-call-centre-to-the-worlds-cx-leader/135823/">Express Computer</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Kyutai shows</strong> how RL post-training improves turn-taking and backchanneling in full-duplex voice models. (<a href="https://kyutai.org/blog/2026-06-10-interactivity">Kyutai</a>)</p></li></ul><ul><li><p><strong>Treble and Hugging Face</strong> launch FFASR, the first open benchmark for far-field speech recognition. (<a href="https://www.newsfilecorp.com/release/300719/Treble-Technologies-and-Hugging-Face-Address-Voice-AIs-Unspoken-Dilemma-With-Groundbreaking-Benchmark-of-ASR-Models">Newsfilecorp</a>)</p></li><li><p><strong>Red Hat publishes</strong> a guide to building a local voice agent with OpenShift AI. (<a href="https://developers.redhat.com/articles/2026/06/08/build-local-voice-agent-red-hat-openshift-ai">Red Hat Developer</a>)</p></li><li><p><strong>DrivenData announces</strong> winners of &#8220;On Top of Pasketti,&#8221; a children&#8217;s speech recognition challenge. (<a href="https://drivendata.co/blog/on-top-of-pasketti-winners">DrivenData</a>)</p></li><li><p><strong>Dev.to tutorial</strong> on extracting conversation intelligence from audio beyond simple dictation. (<a href="https://dev.to/nfc/beyond-dictation-how-to-extract-true-conversation-intelligence-from-audio-in-seconds-21ep">Dev.to</a>)</p></li><li><p><strong>Dev.to tutorial</strong> on building voice agents that send follow-up emails via Nylas. (<a href="https://dev.to/qasim157/voice-agents-that-follow-up-by-email-5ej6">Dev.to</a>)</p></li><li><p><strong>Blog tutorial</strong> covers building an ElevenLabs + n8n voice AI sales agent end to end. (<a href="https://whoisalfaz.me/blog/elevenlabs-n8n-voice-ai-sales-agent/">whoisalfaz.me</a>)</p></li><li><p><strong>ParseJargon paper</strong> introduces real-time jargon translation for online meetings using LLMs. (<a href="https://arxiv.org/abs/2508.10239">arXiv</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Grok Powers Vapi, Gemma 4 Brings Audio to Your Laptop]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/grok-powers-vapi-gemma-4-brings-audio</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/grok-powers-vapi-gemma-4-brings-audio</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 08 Jun 2026 14:02:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/569c37ee-e821-43ec-bcb5-a157d010e893_1896x1054.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Top Updates &#128170;</h2><ul><li><p><strong>xAI brings Grok</strong> TTS and STT to Vapi, letting developers build voice agents with Grok&#8217;s speech models. (<a href="https://x.ai/news/grok-vapi">xAI</a>)</p></li><li><p><strong>Sesame launches</strong> its iOS app with four conversational voice agents, built by the co-founders of Oculus. (<a href="https://techcrunch.com/2026/05/28/sesame-the-conversational-ai-startup-from-oculus-founders-launches-its-ios-app/">TechCrunch</a>)</p></li></ul><ul><li><p><strong>Google releases</strong> Gemma 4 12B, an open-source multimodal model with native audio that runs on a 16GB laptop. (<a href="https://venturebeat.com/technology/googles-new-open-source-gemma-4-12b-analyzes-audio-video-and-runs-entirely-locally-on-a-typical-16gb-enterprise-laptop">VentureBeat</a>)</p></li><li><p><strong>AethexAI raises $3M</strong> to build voice AI infrastructure for Africa and the Middle East. (<a href="https://techcrunch.com/2026/06/03/these-two-founders-left-goldman-and-meta-to-build-voice-ai-for-markets-everyone-else-overlooked/">TechCrunch</a>)</p></li><li><p><strong>Aircall acquires</strong> Piper AI to add revenue intelligence and sales automation to its voice platform. (<a href="https://www.webpronews.com/aircall-buys-piper-ai-in-bid-to-own-the-full-sales-revenue-cycle/">WebProNews</a>)</p></li><li><p><strong>8x8 launches Pulse</strong>, a conversational intelligence tool that turns calls and chats into actionable business insights. (<a href="https://www.businesswire.com/news/home/20260603379786/en/8x8-Introduces-8x8-Pulse-Conversational-Intelligence-Built-for-Where-Decisions-Are-Made">BusinessWire</a>)</p></li><li><p><strong>Google rolls out</strong> real-time deepfake voice detection on Android to catch AI scam calls as they happen. (<a href="https://www.techbuzz.ai/articles/google-deploys-ai-to-detect-deepfake-voice-scams-in-real-time">TechBuzz</a>)</p></li><li><p><strong>Microsoft Edge adds</strong> on-device speech recognition and translation APIs powered by local AI models. (<a href="https://blogs.windows.com/msedgedev/2026/06/02/expanding-on-device-ai-in-microsoft-edge-new-models-and-apis-for-the-web/">Microsoft</a>)</p></li><li><p><strong>McDonald&#8217;s pilots</strong> ArchIQ, a voice AI drive-thru that handles 90% of orders without human help. (<a href="https://www.theedadvocate.org/how-mcdonalds-ai-drive-thru-system-could-change-fast-food-forever/">TheEdAdvocate</a>)</p></li><li><p><strong>Peak XV eyes</strong> a $10M round in Ringg AI as Indian voice agent startups gain momentum. (<a href="https://m.economictimes.com/tech/funding/peak-xv-in-talks-to-back-ringg-ai-sources-say-as-voice-ai-gains-attention/articleshow/131488657.cms">Economic Times</a>)</p></li><li><p><strong>Deepgram partners</strong> with Fortanix to run voice AI on-premises using NVIDIA confidential computing. (<a href="https://itnerd.blog/2026/06/01/deepgram-delivers-private-voice-ai-to-regulated-industries-with-on-premises-deployments-powered-by-fortanix-confidential-ai-and-nvidia-confidential-computing/">ITNerd</a>)</p></li><li><p><strong>Americans lost $893M</strong> to AI scams last year, with voice cloning attacks leading the surge. (<a href="https://www.the-independent.com/news/world/americas/crime/ai-scams-americans-lost-millions-b2984788.html">The Independent</a>)</p></li><li><p><strong>Equity demands</strong> Fish Audio remove unauthorized AI clones of performers&#8217; voices from its platform. (<a href="https://www.equity.org.uk/news/2026/equity-demands-fish-audio-removes-unauthorised-ai-voices">Equity</a>)</p></li><li><p><strong>Sarvam AI opens</strong> its multilingual voice agents platform to the public, covering 11 Indian languages. (<a href="https://letsdatascience.com/news/sarvam-ai-opens-voice-agents-platform-to-public-a32ad441">LetDataScience</a>)</p></li><li><p><strong>Ubuntu plans</strong> to ship AI-powered speech-to-text across all text fields in the OS. (<a href="https://www.omgubuntu.co.uk/2026/06/ubuntu-speech-to-text-ai/amp">OMG Ubuntu</a>)</p></li><li><p><strong>ENCO debuts</strong> EnSpeak, a real-time voice-to-voice translation system for live venues and classrooms. (<a href="https://ravepubs.com/enco-brings-real-time-voice-translation-to-proav-with-infocomm-debut-of-enspeak/">RavePubs</a>)</p></li><li><p><strong>Broadvoice launches</strong> GoEngage and AI Analyst, adding speech-to-speech voice AI to its contact center. (<a href="http://www.smartcustomerservice.com/Articles/News-Briefs/Broadvoice-Launches-GoEngage-and-AI-Analyst-175105.aspx">SmartCustomerService</a>)</p></li><li><p><strong>ElevenLabs opens</strong> a pop-up store in NYC where every part of the experience is run by a voice agent. (<a href="https://letsdatascience.com/news/elevenlabs-runs-nyc-pop-up-featuring-voice-agents-a46404fb">LetDataScience</a>)</p></li><li><p><strong>RingCentral leads</strong> the G2 Summer 2026 AI VoIP category with 137 product badges. (<a href="https://www.ringcentral.com/us/en/blog/ringcental-leads-g2-summer-2026-ai-voip-category/">RingCentral</a>)</p></li><li><p><strong>In2ition AI launches</strong> Iris, an always-on AI companion that joins live meetings instead of just transcribing them. (<a href="https://www.prnewswire.com/news-releases/in2ition-ai-launches-iris-the-always-on-ai-companion-that-participates-in-conversations-instead-of-analyzing-them-after-they-end-302789131.html">PRNewswire</a>)</p></li><li><p><strong>Astreya integrates</strong> 3CLogic voice AI into its ServiceNow-based IT service desk. (<a href="https://www.prnewswire.com/news-releases/astreya-expands-ai-first-service-desk-with-3clogic-integration-unifying-voice-ai-and-itsm-on-servicenow-302788520.html">PRNewswire</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>NVIDIA publishes</strong> a fine-tuning guide for Nemotron 3.5 ASR, its 600M-param streaming model covering 40 languages. (<a href="https://huggingface.co/blog/nvidia/fine-tuning-nemotron-35-asr">Hugging Face</a>)</p></li></ul><ul><li><p><strong>Higgs Audio v3</strong> is a 4B-param chat-native TTS model supporting 102 languages with zero-shot voice cloning. (<a href="https://www.lmsys.org/blog/2026-06-04-higgs-audio-v3-tts/">LMSYS</a>)</p></li><li><p><strong>MisoTTS</strong> is an 8B emotive TTS model with open weights that claims 110ms latency. (<a href="https://www.marktechpost.com/2026/06/04/miso-labs-releases-misotts-an-8b-emotive-text-to-speech-model-with-open-weights/">MarkTechPost</a>)</p></li><li><p><strong>Audio-Interaction</strong> is a 3B open-source model that listens nonstop and decides every 0.4 seconds whether to speak. (<a href="https://the-decoder.com/new-open-source-voice-model-listens-nonstop-and-decides-every-0-4-seconds-whether-to-speak-or-stay-silent/">The Decoder</a>)</p></li><li><p><strong>pyannote.ai&#8217;s Bredin</strong> explains how speaker diarization makes voice AI understand conversations, not just transcribe them. (<a href="https://www.startuphub.ai/ai-news/ai-research/2026/pyannoteai-s-bredin-on-building-conversational-voice-ai">StartupHub</a>)</p></li><li><p><strong>HackerNoon walks through</strong> how to transfer an AI voice agent to a human without losing context. (<a href="https://hackernoon.com/the-warm-handoff-how-to-transfer-an-ai-voice-agent-to-a-human-without-losing-context">HackerNoon</a>)</p></li><li><p><strong>HackerNoon lists</strong> the 7 best voice agent testing platforms for 2026. (<a href="https://hackernoon.com/7-of-the-best-voice-agent-testing-platforms-in-2026">HackerNoon</a>)</p></li><li><p><strong>TechStartups breaks down</strong> how speech datasets for AI are built, what they contain, and where they fail. (<a href="https://techstartups.com/2026/06/01/speech-datasets-for-ai-what-they-contain-how-theyre-built-and-where-they-break/">TechStartups</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Anthropic's Trillion-Dollar Moment]]></title><description><![CDATA[Voice AI weekly digest: massive funding, major platform partnerships, translation breakthroughs, and a growing push into wearables, healthcare, and enterprise software.]]></description><link>https://voice-ai-newsletter.krisp.ai/p/anthropics-trillion-dollar-moment</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/anthropics-trillion-dollar-moment</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 01 Jun 2026 13:45:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/aac0e4f8-a458-4a67-a9bd-5157801ebcb1_1836x1088.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Top Updates &#128170;</h2><ul><li><p>Anthropic closes a Series H near a $965B valuation, landing alongside its Claude Opus 4.8 launch. (<a href="https://techcrunch.com/2026/05/28/anthropic-raises-65-billion-nears-1t-valuation-ahead-of-ipo/">TechCrunch</a>)</p></li><li><p>Parloa deploys its $350M war chest into partnerships with SAP, Microsoft, OpenAI, Five9, and Epic. (<a href="https://thenextweb.com/news/parloa-turns-its-350-million-war-chest-into-a-partnership-web-spanning-sap-microsoft-and-openai">The Next Web</a>)</p></li><li><p>Exclusive: Krisp scales its infra deployment paradigm (<a href="https://www.youtube.com/watch?v=09plsaCZAAU">AIM Network</a>)</p></li><li><p>Greenhouse acquires Ezra AI Labs, folding a voice-AI interviewer into its hiring platform. (<a href="https://www.prnewswire.com/news-releases/greenhouse-completes-acquisition-of-ezra-ai-labs-bringing-conversational-ai-to-the-hiring-process-302782372.html">PR Newswire</a>)</p></li><li><p>Alibaba Updates Speech Translation Model, Triples Language Coverage (<a href="https://slator.com/alibaba-speech-translation-model-triples-language-coverage/">Slator</a>)</p></li><li><p>StepFun ships StepAudio 2.5 Realtime, an end-to-end speech LLM with roleplay RLHF and paralinguistic perception. (<a href="https://www.marktechpost.com/2026/05/24/stepfun-releases-stepaudio-2-5-realtime-an-end-to-end-voice-model-with-roleplay-specific-rlhf-and-paralinguistic-comprehension/">MarkTechPost</a>)</p></li><li><p>COLDI launches a turnkey platform for integrated AI voice agents aimed at lead management. (<a href="https://www.prnewswire.com/news-releases/coldi-unveils-turnkey-platform-for-integrated-ai-voice-agents-302781678.html">PR Newswire</a>)</p></li><li><p>What the Language Solutions and AI Market Should Take Away From Google I/O (<a href="https://slator.com/language-solutions-ai-market-take-aways-google/">Slator</a>)</p></li><li><p>Palabra.ai crosses $1M ARR, a 17x six-month climb for its real-time speech-to-speech translator. (<a href="https://aithority.com/uncategorized/palabra-ai-real-time-ai-voice-translator-hits-1m-arr-grows-17x-in-six-months/">AiThority</a>)</p></li><li><p>iFlytek debuts 40g AI glasses with an on-device GlassClaw agent and live translation in 122 languages. (<a href="https://longbridge.com/en/news/287989104">Longbridge</a>)</p></li><li><p>iFLYTEK unveils AI Recorder S6 with long-range voice recording and smart summaries (<a href="https://markets.financialcontent.com/stocks/article/abnewswire-2026-5-29-iflytek-unveils-ai-recorder-s6-with-long-range-voice-recording-smart-summaries-and-enterprise-grade-data-security">FinancialContent</a>)</p></li><li><p>An ElevenLabs-linked deal licenses Stan Lee&#8217;s voice and likeness for AI-narrated audiobooks and comics. (<a href="https://kotaku.com/stan-lee-marvel-voice-likeness-rights-ai-elevenlabs-2000699882">Kotaku</a>)</p></li><li><p>What Apple&#8217;s New AI Glasses Mean for the Future of Wearables. (<a href="https://www.geeky-gadgets.com/apple-ai-glasses-features/">Geeky Gadgets</a>)</p></li><li><p>A new study shows inaudible audio commands can hijack AI voice models unheard by humans. (<a href="https://decrypt.co/369042/inaudible-audio-attacks-hijack-ai-voice-models">Decrypt</a>)</p></li><li><p>AI Studios Launches Context-Aware Expressive TTS with 1,000+ AI Voices (<a href="https://markets.businessinsider.com/news/stocks/ai-studios-launches-context-aware-expressive-tts-with-1-000-ai-voices-1036193321">Business Insider</a>)</p></li><li><p>What healthcare organizations need to get right about AI transcription. (<a href="https://natlawreview.com/article/ai-transcription-tools-health-care-what-house-counsel-needs-get-right">National Law Review</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p>OmniVoice Studio ships as a local, open-source ElevenLabs alternative with cloning, dubbing, diarization, and an MCP server. (<a href="https://www.marktechpost.com/2026/05/26/meet-omnivoice-studio-a-local-open-source-alternative-to-elevenlabs/">MarkTechPost</a>)</p></li><li><p>A field guide to production voice agents tackles sub-300ms latency with LiveKit and WebRTC. (<a href="https://dev.to/dishant_sethi/building-production-voice-ai-agents-latency-architecture-and-what-nobody-tells-you-3jhj">dev.to</a>)</p></li><li><p>A walkthrough adds Gemma 4 speech recognition to a .NET desktop app via a llama-server sidecar. (<a href="https://dev.to/mdemin729/adding-gemma-4-speech-recognition-to-a-net-desktop-app-the-llama-server-sidecar-that-survived-298j">dev.to</a>)</p></li><li><p>Vaani pairs speech recognition with Indian Sign Language on Android using MediaPipe. (<a href="https://dev.to/kinara2020/vaani-ai-making-communication-more-inclusive-with-speech-recognition-and-indian-sign-language-1i55">dev.to</a>)</p></li><li><p>FlowSpeech offers context-aware TTS with controllable emotion, pacing, and pauses across 30+ voices. (<a href="https://flowspeech.io/">flowspeech.io</a>)</p></li><li><p>Vowen runs fully offline STT on Windows and macOS, free and privacy-first. (<a href="https://www.majorgeeks.com/files/details/vowen.html">MajorGeeks</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Google I/O Goes Voice-First, Corti Beats OpenAI on Medical STT]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/google-io-goes-voice-first-corti</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/google-io-goes-voice-first-corti</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 25 May 2026 14:03:05 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/515b6cef-3dc9-4faa-be44-a239892b0b3d_1290x966.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Top Updates &#128170;</h2><ul><li><p><strong>Google adds voice</strong> to Gmail, Docs and Keep letting users search their inbox and dictate by speaking instead of typing. (<a href="https://techcrunch.com/2026/05/19/you-can-now-talk-to-your-gmail-inbox-as-seen-at-google-io-2026/">TechCrunch</a>)</p></li></ul><ul><li><p><strong>Google unveils</strong> audio-powered smart glasses at I/O 2026, taking on Meta in the wearable AI race. (<a href="https://techcrunch.com/2026/05/19/google-takes-a-page-out-of-metas-book-announces-new-audio-powered-smart-glasses-at-io-2026/">TechCrunch</a>)</p></li><li><p><strong>Spotify launches</strong> an ElevenLabs-powered tool that lets authors create audiobooks from text. (<a href="https://techcrunch.com/2026/05/21/spotify-launches-an-elevenlabs-powered-audiobook-creation-tool/">TechCrunch</a>)</p></li><li><p><strong>Corti&#8217;s Symphony model</strong> outperforms OpenAI&#8217;s Whisper on medical terminology accuracy for speech-to-text. (<a href="https://venturebeat.com/technology/cortis-new-symphony-for-speech-to-text-model-beats-openai-at-medical-terminology-accuracy-highlighting-the-value-of-specialized-ai">VentureBeat</a>)</p></li><li><p><strong>Zoom opens</strong> its AI Translator and Summarizer as standalone APIs for third-party developers. (<a href="https://news.zoom.com/zoom-mcp-expanded-capabilities/">Zoom</a> | <a href="https://slator.com/zoom-ai-services-translator-summarizer/">Slator</a>)</p></li><li><p><strong>Twilio shares surged</strong> 60% as voice AI adoption accelerates across its communications platform. (<a href="https://sebastianbarros.substack.com/p/twilio-shares-surged-60-on-voice">Sebastian Barros</a>)</p></li><li><p><strong>Zendesk expands</strong> its AI agents across ChatGPT, Gemini, voice and messaging channels. (<a href="https://www.techradar.com/pro/zendesk-expands-ai-agents-across-chatgpt-gemini-voice-and-messaging">TechRadar</a>)</p></li><li><p><strong>Kardome ships</strong> its voice AI in LG OLED TVs, reaching mass-market consumers for the first time. (<a href="https://audioxpress.com/news/kardome-voice-ai-reaches-mass-market-with-lg-oled-tv-deployments">AudioXpress</a>)</p></li><li><p><strong>Amazon&#8217;s Alexa</strong> can now generate full podcast episodes on any topic you ask for. (<a href="https://www.techbuzz.ai/articles/amazon-s-new-alexa-powered-feature-can-generate-podcast-episodes">TechBuzz</a>)</p></li><li><p><strong>Alibaba releases</strong> Qwen3.5 LiveTranslate Flash, a real-time interpreter covering 60 languages at 2.8-second latency. (<a href="https://www.marktechpost.com/2026/05/20/alibaba-qwen-team-introduces-qwen3-5-livetranslate-flash-real-time-multimodal-interpretation-across-60-languages-at-2-8-second-latency/">MarkTechPost</a>)</p></li><li><p><strong>NTSB shuts down</strong> its public docket after people used AI to recreate dead pilots&#8217; voices from spectrograms. (<a href="https://www.engadget.com/2180049/people-used-ai-to-recreate-the-voices-of-pilots-killed-in-a-plane-crash/">Engadget</a>)</p></li><li><p><strong>Columbia researchers</strong> pass the first human trial of a brain-controlled hearing system that isolates one speaker in noise. (<a href="https://www.medscape.com/viewarticle/brain-controlled-hearing-system-passes-first-human-trial-2026a1000gcc">Medscape</a>)</p></li><li><p><strong>iProov launches</strong> a deepfake detection system designed specifically for enterprise video calls. (<a href="https://financefeeds.com/iproov-launches-deepfake-detection-system-for-enterprise-video-calls/">FinanceFeeds</a>)</p></li><li><p><strong>Halsa Global launches</strong> Voice IQ, a Salesforce-native conversational AI for enterprise sales. (<a href="https://www.newswire.com/news/halsa-global-launches-voice-iq-a-salesforce-native-conversational-22784148">Newswire</a>)</p></li><li><p><strong>Korean tech firms</strong> double down on voice AI with localized models and in-car assistants. (<a href="https://www.koreatimes.co.kr/business/tech-science/20260522/tech-firms-double-down-on-voice-ai-as-next-battleground-emerges">Korea Times</a>)</p></li><li><p><strong>Tamber launches</strong> its AI music creation platform after raising $5M from Adobe Ventures. (<a href="https://www.musicbusinessworldwide.com/after-raising-5m-adobe-backed-tamber-officially-launches-its-ai-music-making-platform/">Music Business Worldwide</a>)</p></li><li><p><strong>TalkSign launches</strong> Palm 1.0 and Echo 1.0, AI models for sign language recognition and generation. (<a href="https://techcabal.com/2026/05/20/talksign-launches-ai-powered-palm-1-0-and-echo-1-0/">TechCabal</a>)</p></li><li><p><strong>CMU research shows</strong> adding audio cues like typing sounds makes AI feel more human but also more rude. (<a href="https://techxplore.com/news/2026-05-audio-cues-ai-human-users.html">TechXplore</a>)</p></li><li><p><strong>Office workers shift</strong> from typing to voice dictation as AI transcription apps go mainstream. (<a href="https://theweek.com/tech/the-changing-sounds-of-the-office">The Week</a>)</p></li><li><p><strong>Synthflow AI handles</strong> over 5 million calls a month as call centres move to voice AI at scale. (<a href="https://tech.eu/2026/05/22/the-call-centre-enters-the-voice-ai-era/">Tech.eu</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>AWS publishes</strong> a guide to building real-time voice apps with SageMaker AI and vLLM using bidirectional streaming. (<a href="https://aws.amazon.com/blogs/machine-learning/build-real-time-voice-applications-with-amazon-sagemaker-ai-and-vllm/">AWS Blog</a>)</p></li></ul><ul><li><p><strong>VoiceBox</strong> is an open-source voice cloning app that runs locally from 3 seconds of audio with no cloud uploads. (<a href="https://www.techtimes.com/articles/316850/20260519/voicebox-clones-any-voice-3-seconds-audio-runs-locally-free-has-no-consent-lock.htm">TechTimes</a>)</p></li><li><p><strong>Vowen</strong> is a free offline voice dictation tool for Windows and macOS that transcribes speech system-wide. (<a href="https://www.majorgeeks.com/files/details/vowen.html">MajorGeeks</a>)</p></li><li><p><strong>NoteSnip</strong> turns video transcripts into source-grounded AI study notes across YouTube, podcasts and PDFs. (<a href="https://dev.to/_993f2d61f0282f6943ea3/from-video-transcripts-to-source-grounded-ai-notes-a-practical-look-at-notesnip-33in">Dev.to</a>)</p></li><li><p><strong>IEEE Spectrum covers</strong> how Maori researchers are building indigenous AI voice models to preserve te reo Maori. (<a href="https://spectrum.ieee.org/indigenous-ai-voice-models-maori">IEEE Spectrum</a>)</p></li><li><p><strong>Memeburn ranks</strong> the best AI voice generators of 2026 by use case, from cloning to e-learning. (<a href="https://memeburn.com/best-ai-voice-generator/">Memeburn</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Customer Service Hiring Is Surging. So Is Voice AI]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/customer-service-hiring-is-surging</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/customer-service-hiring-is-surging</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 18 May 2026 14:03:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bV3W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Customer service job postings are up ~8% YoY. More voice AI doesn&#8217;t mean fewer human agents - it means more conversations.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bV3W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bV3W!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png 424w, https://substackcdn.com/image/fetch/$s_!bV3W!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png 848w, https://substackcdn.com/image/fetch/$s_!bV3W!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!bV3W!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bV3W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png" width="1248" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1248,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:950139,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/198210285?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bV3W!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png 424w, https://substackcdn.com/image/fetch/$s_!bV3W!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png 848w, https://substackcdn.com/image/fetch/$s_!bV3W!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!bV3W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33f49e81-df6b-43c2-b40e-ba17d750ad1b_1248x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Top Updates &#128170;</h2><ul><li><p>Vapi raises $50M for its voice AI agent platform, now valued at $500M. (<a href="https://techcrunch.com/2026/05/12/vapi-hits-500m-valuation-as-amazon-ring-chose-its-ai-platform-over-40-rivals/">TechCrunch</a>)</p></li></ul><ul><li><p>Thinking Machines previews voice+video models that can listen and talk at the same time. (<a href="https://venturebeat.com/technology/thinking-machines-shows-off-preview-of-near-realtime-ai-voice-and-video-conversation-with-new-interaction-models">VentureBeat</a>)</p></li><li><p>Wispr seeks $260M at a $2B valuation for its voice dictation app. (<a href="https://www.bloomberg.com/news/articles/2026-05-12/ai-dictation-startup-wispr-in-funding-talks-at-2-billion-value">Bloomberg</a>)</p></li><li><p>OpenAI acquires Weights.gg, a voice cloning startup, and folds the team internally. (<a href="https://www.itvoice.in/openai-quietly-acquires-voice-cloning-startup-weights-gg-to-boost-audio-ai-capabilities">ITVoice</a>)</p></li><li><p>Medicare will reimburse AI voice agents that manage chronic care patients. (<a href="https://www.webpronews.com/medicares-quiet-bet-on-ai-agents-that-could-reshape-chronic-care/">WebProNews</a>)</p></li><li><p>Better.com&#8217;s voice agent handles 35% of mortgage calls without human involvement. (<a href="https://www.pymnts.com/artificial-intelligence-2/2026/better-coms-ai-agent-resolved-35-of-mortgage-calls-alone/">PYMNTS</a>)</p></li><li><p>Bajaj Finance replaces 1,500 calling agents with 10 AI voice bots. (<a href="https://techstory.in/10-ai-bots-replace-1500-employees-at-bajaj-finance-as-automation-wave-intensifies/">TechStory</a>)</p></li><li><p>Rivian rolls out a voice assistant across its R1 and R2 vehicles. (<a href="https://insideevs.com/news/795539/rivian-assistant-launch-r1-r2-2026/">InsideEVs</a>)</p></li><li><p>Quiq adds voice AI to its platform and rebrands for enterprise scale. (<a href="https://customerservicemanager.com/quiq-expands-voice-ai-and-rebrands-to-focus-on-scaled-enterprise-deployments/">CSM Magazine</a>)</p></li><li><p>ElevenLabs signs McConaughey, Caine, and Minnelli for AI voice partnerships. (<a href="https://deadline.com/2026/05/elevenlabs-mati-staniszewski-matthew-mcconaughey-ai-audio-1236900840/">Deadline</a>)</p></li><li><p>Activate invests in ElevenLabs to help grow its India business. (<a href="https://www.businesstoday.in/technology/story/activate-invests-in-elevenlabs-bets-big-on-indias-voice-ai-opportunity-531498-2026-05-14">BusinessToday</a>)</p></li><li><p>RingCentral named Leader by IDC, Omdia, and Metrigy for customer engagement. (<a href="https://www.ringcentral.com/us/en/blog/leader-analyst-reports-future-cus**tomer-engagement/">RingCentral Blog</a>)</p></li><li><p>Smallest AI runs its TTS on Tenstorrent chips at 4x lower cost. (<a href="https://m.thewire.in/article/ptiprnews/smallest-ai-and-tenstorrent-partnership-democratises-voice-ai-4x-reduction-in-cost-through-hardware-acceleration/amp">The Wire</a>)</p></li><li><p>MindBio detects intoxication from voice alone using AI speech analysis. (<a href="https://markets.businessinsider.com/news/stocks/networknews-audio-announces-audio-press-release-apr-discussing-combining-artificial-intelligence-with-speech-analysis-to-detect-intoxication-1036163396">BusinessInsider</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p>AWS adds Qwen3 speech models to SageMaker JumpStart for TTS and ASR. (<a href="https://aws.amazon.com/about-aws/whats-new/2026/05/speech-models-on-sagemaker-jumpstart/">AWS</a>)</p></li><li><p>Foundry Local v1.1 adds live speech-to-text that runs entirely on-device. (<a href="https://devblogs.microsoft.com/foundry/foundry-local-v1-1/">Microsoft DevBlogs</a>)</p></li></ul><ul><li><p>Supertone open-sources Supertonic v3, an on-device TTS supporting 31 languages. (<a href="https://www.marktechpost.com/2026/05/15/supertone-releases-supertonic-v3-on-device-text-to-speech-model-with-31-language-support-fewer-reading-failures-and-expression-tags/">MarkTechPost</a>)</p></li><li><p>Coval publishes open TTS benchmarks comparing speed and accuracy across major providers. (<a href="https://benchmarks.coval.ai/tts">Coval</a>)</p></li><li><p>OpenMOSS gets a C++ port for easy local deployment without Python. (<a href="https://startupfortune.com/openmoss-gets-a-c-port-as-local-voice-ai-chases-easier-deployment/">StartupFortune</a>)</p></li><li><p>ThirdReality ships a $70 open-source voice assistant for Home Assistant. (<a href="https://www.prweb.com/releases/thirdreality-launches-voice--music-assistant-dev-edition-302767779.html">PRWeb</a>)</p></li><li><p>Monologue adds CLI and MCP support for piping voice dictation into AI agents. (<a href="https://www.macstories.net/notes/monologue-notes-cli/">MacStories</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://resources.krisp.ai/fullband-2025&quot;,&quot;text&quot;:&quot;Register now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://resources.krisp.ai/fullband-2025"><span>Register now</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Updates from Krisp, OpenAI, ServiceNow and much more!]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/updates-from-krisp-openai-servicenow</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/updates-from-krisp-openai-servicenow</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 11 May 2026 14:00:36 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ae2e60e3-d0bb-4432-bdf7-ccf410b092a9_1370x774.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Top Updates &#128170;</h2><ul><li><p><strong>Krisp launches VIVA 2.0</strong> with Turn Prediction v3 and a first-of-its-kind Interrupt Prediction model, all running on CPU with no transcription required. (<a href="https://krisp.ai/blog/viva-2-0-ai-infrastructure-for-voice-ai-agents/">Krisp Blog</a>)</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;56ef9704-29e8-4a52-b708-d3d1e89ea776&quot;,&quot;duration&quot;:null}"></div></li><li><p><strong>OpenAI launches three real-time audio models</strong> for its API: GPT-Realtime-2 with GPT-5-class reasoning, GPT-Realtime-Translate for live translation across 70+ languages, and GPT-Realtime-Whisper for streaming speech-to-text. (<a href="https://www.reuters.com/business/media-telecom/openai-unveils-three-audio-models-real-time-voice-tasks-2026-05-07/">Reuters</a>)</p></li></ul><ul><li><p><strong>Twilio unveils a Conversation Layer at SIGNAL 2026</strong> with persistent Memory, Orchestrator, Intelligence, and open-source Agent Connect for plugging in any AI provider. (<a href="https://martech.org/twilio-launches-conversation-layer-to-unify-ai-and-human-interactions/">MarTech</a>)</p></li><li><p><strong>Inworld ships Realtime TTS-2,</strong> a frontier voice model that reads user emotion and tone in real time and adapts pacing, softness, and empathy mid-conversation. (<a href="https://www.morningstar.com/news/business-wire/20260505096579/inworld-launches-new-frontier-voice-model-that-gives-ai-agents-contextual-empathy">BusinessWire</a>)</p></li><li><p><strong>ServiceNow unveils Otto,</strong> a unified conversational AI layer combining Now Assist, Moveworks, and voice agents across every department and system. (<a href="https://theaieconomy.substack.com/p/servicenow-otto-conversational-ai-enterprise">The AI Economy</a>)</p></li><li><p><strong>SoundHound launches OASYS,</strong> a self-learning agentic platform that auto-builds, orchestrates, and improves voice AI agents from documentation and transcripts. (<a href="https://www.globenewswire.com/news-release/2026/05/05/3287821/0/en/soundhound-ai-introduces-oasys-the-world-s-first-self-learning-orchestrated-agentic-ai-platform-where-ai-builds-ai.html">GlobeNewsWire</a>)</p></li><li><p><strong>ElevenLabs adds BlackRock, NVIDIA, and Jamie Foxx</strong> to its $550M+ Series D as annualized revenue crosses $500M, up from $350M at the end of 2025. (<a href="https://techcrunch.com/2026/05/05/elevenlabs-lists-blackrock-jamie-foxx-and-eva-longoria-as-new-investors/">TechCrunch</a>)</p></li><li><p><strong>Greenhouse acquires Ezra AI Labs</strong> to bring voice AI interviewing into its ATS as applications per recruiter have spiked over 400% since 2023. (<a href="https://www.prnewswire.com/news-releases/greenhouse-has-entered-into-a-definitive-agreement-to-acquire-ezra-ai-labs-bringing-conversational-ai-to-the-hiring-process-302762658.html">PR Newswire</a>)</p></li><li><p><strong>Ethos raises $22.75M from a16z</strong> for an expert network that onboards 35K people per week through voice AI interviews. (<a href="https://techcrunch.com/2026/05/06/ethos-raises-22-75m-from-a16z-for-its-expert-network-with-voice-onboarding/">TechCrunch</a>)</p></li><li><p><strong>8x8 launches AI Studio</strong> in early availability, letting teams describe needs in plain language and deploy voice and digital AI agents without adding vendors. (<a href="https://www.cmswire.com/contact-center/8x8-expands-cx-platform-with-ai/">CMSWire</a>)</p></li><li><p><strong>Wispr Flow bets on India</strong> as its fastest-growing market with Hinglish dictation support, 2.5M downloads, and 100% month-over-month growth. (<a href="https://techcrunch.com/2026/05/09/voice-ai-in-india-is-hard-wispr-flow-is-betting-on-it-anyway/">TechCrunch</a>)</p></li><li><p><strong>ElevenLabs powers SpoonLabs&#8217; audio novels,</strong> cutting production time from months to hours and launching PodNovel across Korea, Japan, and Taiwan. (<a href="https://www.digitaltoday.co.kr/en/view/52978/elevenlabs-supplies-voice-ai-solution-to-spoonlabs-audio-platform">DigitalToday</a>)</p></li><li><p><strong>eGain launches AI Agent IVA,</strong> a knowledge-powered virtual agent that replaces IVR dial trees with natural conversation and 24/7 voice support. (<a href="https://www.globenewswire.com/news-release/2026/05/06/3288531/0/en/egain-launches-ai-agent-iva-to-deliver-accurate-conversational-customer-service.html">GlobeNewsWire</a>)</p></li><li><p><strong>Gnani.ai hires eight senior execs</strong> after its $10M Series B, processing over 30M voice AI calls daily for 200+ enterprise customers in India. (<a href="https://www.businesstoday.in/technology/story/gnaniai-hires-senior-executives-across-bfsi-product-and-ai-delivery-after-10-million-fundingg-530008-2026-05-06">BusinessToday</a>)</p></li><li><p><strong>Vobiz.ai raises $1M seed</strong> to build AI-native telephony infrastructure in India with DID provisioning, low-latency SIP trunking, and LLM audio streaming. (<a href="https://www.techinasia.com/news/indian-startup-vobiz-ai-secures-1m-for-voice-ai">Tech in Asia</a>)</p></li><li><p><strong>Twinnin targets $3M seed round</strong> for its voice and face cloning marketplace where actors license digital likenesses to studios, backed by Google and NVIDIA. (<a href="https://deadline.com/2026/05/ai-plaform-twinnin-funding-round-3-million-signs-up-twins-1236882734/">Deadline</a>)</p></li><li><p><strong>BCM One partners with TD Synnex</strong> to bring Pure IP voice services and SkySwitch UCaaS to the MSP channel through the distributor&#8217;s partner network. (<a href="https://www.crn.com/news/channel-news/2026/bcm-one-td-synnex-partnership-helps-msps-cash-in-on-voice-ai-opportunity">CRN</a>)</p></li><li><p><strong>AI note-taking earbuds go mainstream</strong> as Viaim and Mobvoi ship wireless earbuds that record, transcribe, and summarize meetings entirely on-device. (<a href="https://www.howtogeek.com/ai-note-taking-earbuds-record-and-summarize-meetings/">How-To Geek</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>OpenAI publishes its WebRTC infrastructure playbook,</strong> detailing a split relay + transceiver architecture that routes voice AI sessions for 900M+ weekly users at 300-500ms latency. (<a href="https://openai.com/index/delivering-low-latency-voice-ai-at-scale/">OpenAI Blog</a>)</p></li></ul><ul><li><p><strong>TypeWhisper open-sources Mac dictation</strong> with 10 ASR engines including WhisperKit, Parakeet, Apple SpeechAnalyzer, Groq, and xAI Grok STT, all running locally. (<a href="https://github.com/TypeWhisper/typewhisper-mac">GitHub</a>)</p></li><li><p><strong>Dictee ships offline voice dictation for Linux</strong> as a KDE Plasma 6 plasmoid with Rust backend, 4 ASR engines, and NVIDIA Parakeet via ONNX Runtime. (<a href="https://github.com/rcspam/dictee">GitHub</a>)</p></li><li><p><strong>TTS models for Indian languages:</strong> a dev survey covering Hindi, Tamil, Bengali, and Telugu with architecture comparisons and demo links. (<a href="https://dev.to/vinodsrajpurohit/tts-models-for-indian-languages-the-tech-giving-bharat-a-voice-1ij7">dev.to</a>)</p></li><li><p><strong>Build a voice agent with LiveKit + AssemblyAI</strong> using Universal-3 Pro Streaming STT with function calling and MCP integration. (<a href="https://dev.to/martschweiger/build-a-voice-agent-with-livekit-and-assemblyais-voice-agent-api-3mnm">dev.to</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://resources.krisp.ai/fullband-2025&quot;,&quot;text&quot;:&quot;Register now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://resources.krisp.ai/fullband-2025"><span>Register now</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Voice Agents Go Mainstream]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/voice-agents-go-mainstream</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/voice-agents-go-mainstream</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 04 May 2026 13:17:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d31b576a-8c1c-46fe-be99-d396c55480dc_1896x1052.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Three important Voice AI events this week:</p><ul><li><p><a href="https://signal.twilio.com/">Twilio Signal</a> - May 6-7 in SF</p></li><li><p><a href="https://cerebralvalleyvoice.com/">Cerebral Valley Voice Summit</a> - May 6 in SF</p></li><li><p><a href="https://luma.com/SpeechAImeetup">NVIDIA Developer Meetup</a> | Building and Evaluating Real-time Voice Agents - May 7 in SF</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>xAI launches Custom Voices,</strong> a voice cloning API that creates a voice ID from 120 seconds of audio with speaker verification, plus 80+ built-in voices across 28 languages. (<a href="https://venturebeat.com/technology/xai-launches-grok-4-3-at-an-aggressively-low-price-and-a-new-fast-powerful-voice-cloning-suite">VentureBeat</a>)</p></li><li><p><strong>Microsoft ships real-time voice agents in Copilot Studio,</strong> now GA in Dynamics 365 Contact Center with low-latency speech-to-speech, interruptions, and mid-call language switching. (<a href="https://www.microsoft.com/en-us/microsoft-copilot/blog/copilot-studio/extend-ai-voice-support-introducing-real-time-voice-agents-in-microsoft-copilot-studio/">Microsoft Blog</a>)</p></li><li><p><strong>Amazon adds &#8220;Join the Chat&#8221; to product pages,</strong> letting shoppers ask voice or text questions during AI audio summaries and get real-time conversational answers. (<a href="https://techcrunch.com/2026/04/28/amazon-launches-an-ai-powered-audio-qa-experience-on-product-pages/">TechCrunch</a>)</p></li><li><p><strong>Otter.ai pivots from notetaker to Conversational Knowledge Engine,</strong> launching MCP connectors, AI Chat, and desktop app to turn meeting data into agentic workflows. (<a href="https://www.businesswire.com/news/home/20260428313206/en/Otter.ai-Evolves-from-AI-Notetaker-to-Create-%24100B-Enterprise-Conversational-Knowledge-Engine-Market">BusinessWire</a>)</p></li><li><p><strong>Deepgram launches Flux Multilingual</strong> with 10 languages and mid-call language switching, plus model-based turn detection under 400ms. (<a href="https://siliconangle.com/2026/04/29/deepgram-expands-flux-10-languages-mid-call-switching-voice-agents/">SiliconANGLE</a>)</p></li><li><p><strong>Twilio Q1 voice revenue hits a 19-quarter high,</strong> up 20% YoY with Conversational Intelligence and Branded Calling both growing over 100%. (<a href="https://thenextweb.com/news/twilio-q1-2026-voice-ai-revenue">The Next Web</a>)</p></li><li><p><strong>NordVPN adds AI voice deepfake detector</strong> to its Chrome extension, analyzing acoustic patterns in real time without recording or interpreting content. (<a href="https://betanews.com/article/nordvpn-adds-ai-voice-detector-to-its-chrome-extension/">BetaNews</a>)</p></li><li><p><strong>Audion raises $15M</strong> to bring AI-powered contextual audio ad targeting to the U.S., processing 500K hours of audio weekly for brands like Apple and Nike. (<a href="https://www.axios.com/2026/04/27/audion-audio-adtech-raises-us">Axios</a>)</p></li><li><p><strong>3CLogic launches outbound voice AI agents</strong> with multimodal voice+digital capabilities and an automated LLM-powered QA engine for scoring every AI interaction. (<a href="https://www.prnewswire.com/news-releases/3clogic-accelerates-enterprise-roi-with-new-outbound-voice-ai-agents-multimodal-voice-ai-capabilities-and-automated-ai-agent-evaluations-302753788.html">PR Newswire</a>)</p></li><li><p><strong>AI-generated podcasts are booming</strong> on Spotify, Apple, and YouTube, with AI hosts that sound convincingly human raising questions about disclosure. (<a href="https://www.inc.com/moses-jeanfrancois/ai-generated-podcasts-boom-on-audio-platforms-are-you-listening-to-one/91338876">Inc</a>)</p></li><li><p><strong>Tells launches AI voice agents on existing SMS numbers</strong> with a single toggle, adding sub-second-latency voice to any business texting line without a new number or integration. (<a href="https://aithority.com/machine-learning/tells-launches-ai-voice-agents-on-existing-sms-numbers-with-one-click/">AIthority</a>)</p></li><li><p><strong>SpeakON ships a MagSafe AI dictation accessory</strong> that turns iPhone voice input into formatted, tone-adapted text with translation across 12 languages. (<a href="https://9to5mac.com/2026/04/27/key-takeaways-after-testing-out-speakon-an-ai-powered-dictation-iphone-accessory/">9to5Mac</a>)</p></li><li><p><strong>Docplanner&#8217;s voice AI agent &#8220;Noa Booking&#8221;</strong> doubles doctor appointment bookings vs traditional call centers, built on Twilio ConversationRelay. (<a href="https://www.healthtechdigital.com/docplanner-expands-patient-access-with-voice-ai-agent-powered-by-twilio/">Health Tech Digital</a>)</p></li><li><p><strong>Lumeris adds native audio to its Tom platform</strong> using Gemini&#8217;s speech-to-speech capabilities for real-time, empathetic patient conversations in primary care. (<a href="https://hitconsultant.net/2026/04/22/lumeris-native-audio-tom-google-gemini-primary-care/">HIT Consultant</a>)</p></li><li><p><strong>Ablio launches AI-powered interpretation</strong> with hybrid human+AI model, combining ASR, neural translation, and TTS for live multilingual events on Zoom and Teams. (<a href="https://aithority.com/machine-learning/ablio-launches-ai-powered-interpretation-platform-with-hybrid-human-ai-model/">AIthority</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Sakana AI introduces KAME,</strong> a tandem speech-to-speech architecture that lets a backend LLM inject knowledge in real time while the front-end keeps talking with near-zero latency. (<a href="https://pub.sakana.ai/kame/">Sakana AI</a>)</p></li></ul><ul><li><p><strong>NVIDIA releases Nemotron 3 Nano Omni,</strong> an open 30B-A3B multimodal model unifying vision, audio, and language with 9x higher throughput than competing omni models. (<a href="https://blogs.nvidia.com/blog/nemotron-3-nano-omni-multimodal-ai-agents/">NVIDIA Blog</a>)</p></li><li><p><strong>OpenMOSS releases MOSS-Audio,</strong> an open-source foundation model for speech, sound, music understanding, and time-aware audio reasoning in 4B and 8B variants. (<a href="https://www.marktechpost.com/2026/04/27/openmoss-releases-moss-audio-an-open-source-foundation-model-for-speech-sound-music-and-time-aware-audio-reasoning/">MarkTechPost</a>)</p></li><li><p><strong>Async publishes open TTS benchmark</strong> revealing major accuracy gaps when streaming models handle phone numbers, dates, and prices in production. (<a href="https://podnews.net/press-release/async-ai-voice-benchmark">Podnews</a>)</p></li><li><p><strong>Speaker diarization explained:</strong> how AI knows who said what, from spectral embeddings to clustering. (<a href="https://dev.to/quillhub/speaker-diarization-explained-how-ai-knows-who-said-what-9fi">dev.to</a>)</p></li><li><p><strong>Laravel AI SDK tutorial:</strong> add TTS and voice to your app in 20 minutes. (<a href="https://dev.to/hafiz619/laravel-ai-sdk-add-text-to-speech-and-voice-to-your-app-in-20-minutes-35fb">dev.to</a>)</p></li><li><p><strong>Hobbyist builds a C-3PO head</strong> with real-time voice interaction using off-the-shelf speech models. (<a href="https://letsdatascience.com/news/hobbyist-builds-c-3po-head-with-real-time-voice-63356d06">Let&#8217;s Data Science</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://resources.krisp.ai/fullband-2025&quot;,&quot;text&quot;:&quot;Register now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://resources.krisp.ai/fullband-2025"><span>Register now</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Voice AI's Consolidation Begins]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/voice-ais-consolidation-begins</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/voice-ais-consolidation-begins</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 27 Apr 2026 14:01:59 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/788195ac-4ca6-4e56-a622-1d5168942cab_1662x1230.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Two important Voice AI events in the coming weeks:</p><ul><li><p><a href="https://signal.twilio.com/">Twilio Signal</a> - May 6-7 in SF</p></li><li><p><a href="https://cerebralvalleyvoice.com/">Cerebral Valley Voice Summit</a> - May 6 in SF</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>xAI launches Grok Voice Think Fast 1.0,</strong> ranking #1 on the tau-voice Bench for full-duplex voice agents and already powering Starlink support with a 20% sales conversion rate. (<a href="https://x.ai/news/grok-voice-think-fast-1">xAI</a>)</p></li><li><p><strong>Anker unveils THUS,</strong> the first compute-in-memory AI audio chip, claiming 150x more on-device AI power for noise cancellation in its upcoming Soundcore earbuds. (<a href="https://www.theverge.com/tech/916463/anker-thus-chip-announcement">The Verge</a>)</p></li><li><p><strong>SoundHound acquires LivePerson for $43M,</strong> combining voice agentic AI with LivePerson&#8217;s digital messaging platform that handles one billion customer messages per month. (<a href="https://www.globenewswire.com/news-release/2026/04/21/3278086/0/en/soundhound-ai-to-acquire-liveperson-combining-proprietary-voice-agentic-ai-and-digital-messaging-to-create-a-world-leading-end-to-end-omnichannel-conversational-ai-platform.html">GlobeNewswire</a>)</p></li><li><p><strong>Krisp Voice AI SDK won double Webby Awards</strong> for Technical Achievement (<a href="https://www.linkedin.com/feed/update/urn:li:activity:7452374372713046016/">LinkedIn</a>)</p></li><li><p><strong>Speechmatics delivers on-device STT for Adobe Premiere,</strong> transcribing an hour of video in 55 seconds offline with accuracy within 5% of cloud. (<a href="https://www.tvtechnology.com/production/adobe-and-speechmatics-deliver-cloud-grade-on-device-speech-recognition-for-premiere">TV Technology</a>)</p></li><li><p><strong>Nothing launches Essential Voice,</strong> an AI dictation tool that cleans filler words and formats speech-to-text system-wide in 100+ languages. (<a href="https://techcrunch.com/2026/04/24/nothing-introduces-an-ai-powered-dictation-tool/">TechCrunch</a>)</p></li><li><p><strong>Synthflow AI and 8x8 partner</strong> to embed no-code voice AI agents directly into the 8x8 Contact Center platform across 30+ languages. (<a href="https://venturebeat.com/business/synthflow-ai-and-8x8-enter-strategic-partnership-to-deliver-next-generation-agentic-ai">VentureBeat</a>)</p></li><li><p><strong>Google Meet AI note-taking now works for in-person meetings,</strong> generating transcripts, summaries, and action items from face-to-face conversations via mobile. (<a href="https://lifehacker.com/tech/google-meet-can-now-take-notes-during-in-person-meetings">Lifehacker</a>)</p></li><li><p><strong>Xiaomi releases MiMo v2.5 TTS and open-sources MiMo v2.5 ASR,</strong> a full voice pipeline with voice cloning, voice design, and dialect-aware recognition for the agent era. (<a href="https://www.gizmochina.com/2026/04/24/xiaomi-introduces-mimo-v2-5-tts-and-asr-as-a-full-voice-pipeline-for-the-agent-era/">Gizmochina</a>)</p></li><li><p><strong>Volkswagen will ship voice AI in all China-built cars</strong> starting H2 2026, using on-device LLMs from Tencent, Alibaba, and Baidu. (<a href="https://www.cnbc.com/2026/04/21/volkswagen-voice-ai-chinese-cars-automaker.html">CNBC</a>)</p></li><li><p><strong>Newo appoints new CEO after $25M Series A</strong> to scale partner-led voice AI infrastructure for MSPs, VoIP providers, and software platforms serving SMBs. (<a href="https://www.globenewswire.com/news-release/2026/04/21/3277867/0/en/jason-luo-appointed-ceo-of-newo-to-accelerate-partner-led-growth-in-voice-ai-infrastructure-following-25m-series-a.html">GlobeNewswire</a>)</p></li><li><p><strong>Ericsson embeds AI calling and fraud detection into IMS,</strong> partnering with Hiya for real-time spam blocking as 86% of unknown calls go unanswered. (<a href="https://www.ericsson.com/en/blog/2026/4/ai-voice-in-telecom-powering-calls-and-securing-networks">Ericsson Blog</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Streaming TTS models fail over 60% of sentences</strong> containing phone numbers, dates, and prices due to 5-20x less context than batch mode. (<a href="https://www.technology.org/2026/04/24/streaming-tts-models-fail-over-60-of-sentences-containing-numbers-dates-and-prices/">Technology.org</a>)</p></li></ul><ul><li><p><strong>AI neck sensor turns silent speech into voice</strong> by reading microscopic throat muscle movements with a CNN+transformer pipeline from POSTECH. (<a href="https://www.digitaltrends.com/wearables/ai-powered-neck-sensor-can-turn-silent-speech-into-audible-voice/">Digital Trends</a>)</p></li><li><p><strong>AWS guide to cost-effective multilingual transcription</strong> at scale using NVIDIA Parakeet TDT and AWS Batch. (<a href="https://aws.amazon.com/blogs/machine-learning/cost-effective-multilingual-audio-transcription-at-scale-with-parakeet-tdt-and-aws-batch/">AWS Blog</a>)</p></li><li><p><strong>Ghost Pepper:</strong> open-source browser extension for real-time voice transcription and LLM-powered responses. (<a href="https://matthartman.github.io/ghost-pepper/">GitHub</a>)</p></li><li><p><strong>Mimi Codec deep-dive</strong> on its layered audio compression design for neural speech coding. (<a href="https://letsdatascience.com/news/mimi-codec-reveals-layered-audio-compression-design-4ea7aa4a">LetsDDataScience</a>)</p></li><li><p><strong>AssemblyAI showcases configurable STT</strong> with tunable turn-taking, medical mode for streaming, and real-time speaker labeling. (<a href="https://www.tipranks.com/news/private-companies/assemblyai-showcases-configurable-speech-to-text-features-for-voice-ai-developers">TipRanks</a>)</p></li></ul><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://resources.krisp.ai/fullband-2025&quot;,&quot;text&quot;:&quot;Register now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://resources.krisp.ai/fullband-2025"><span>Register now</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Everyone Wants a Voice Platform]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/everyone-wants-a-voice-platform</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/everyone-wants-a-voice-platform</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 20 Apr 2026 14:04:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/673d0805-1f8f-4fae-a65a-cce1decc4a2d_1520x754.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Top Updates &#128170;</h2><ul><li><p><strong>xAI ships standalone Grok STT and TTS APIs</strong> with streaming transcription at $0.20/hr and expressive TTS with inline emotion tags across 20 languages. (<a href="https://x.ai/news/grok-stt-and-tts-apis">xAI</a>)</p></li><li><p><strong>Google launches Gemini 3.1 Flash TTS</strong> with 200+ audio tags for fine-grained voice control, multi-speaker dialogue, and SynthID watermarking across 70+ languages. (<a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-tts/">Google Blog</a>)</p></li></ul><ul><li><p><strong>Starlink customer support is now Grok-powered,</strong> with a voice AI agent handling sales, troubleshooting, and account setup on the phone line. (<a href="https://au.pcmag.com/networking/117104/hi-this-is-ai-starlinks-customer-support-now-features-grok-voice-chatbot">PCMag</a>)</p></li><li><p><strong>Cloudflare adds real-time voice to its Agents SDK,</strong> enabling voice-enabled agents over WebSockets in ~30 lines of server code on Durable Objects. (<a href="https://blog.cloudflare.com/voice-agents/">Cloudflare Blog</a>)</p></li><li><p><strong>DeepL launches voice-to-voice translation</strong> for meetings with Zoom and Teams add-ons, plus a developer API for custom use cases like call centers. (<a href="https://techcrunch.com/2026/04/16/deepl-known-for-text-translation-now-wants-to-translate-your-voice/">TechCrunch</a>)</p></li><li><p><strong>Phonely raises $16M Series A</strong> for AI phone agents that drove $10M+ in insurance policy sales for a single customer this year. (<a href="https://www.axios.com/pro/enterprise-software-deals/2026/04/15/voice-ai-startup-phonely-16-million">Axios</a>)</p></li><li><p><strong>Krisp launches British English accent conversion,</strong> letting offshore agents in India, Philippines, and beyond sound local for UK-facing programs in real time. (<a href="https://cxm.world/customer-experience/krisp-expands-accent-conversion-to-british-english-targeting-uk-facing-offshore-operations/">CXM World</a>)</p></li><li><p><strong>interface.ai launches Nexus,</strong> a fully agentic CCaaS platform for credit unions that eliminates hold queues by keeping AI in the conversation with human backup. (<a href="https://www.globenewswire.com/news-release/2026/04/14/3273785/0/en/interface-ai-Launches-Nexus-The-World-s-First-Fully-Agentic-CCaaS-Platform-That-Ends-the-Era-of-Hold-Queues-Transfers-and-Binary-Call-Routing-for-Credit-Unions-and-Community-Banks.html">GlobeNewswire</a>)</p></li><li><p><strong>ConverseNow partners with Deliverect</strong> to pipe voice AI phone and drive-thru orders into unified restaurant order management across thousands of locations. (<a href="http://www.prnewswire.com/news-releases/conversenow-and-deliverect-announce-partnership-to-bring-voice-ai-ordering-into-unified-restaurant-order--menu-management-302743431.html">PR Newswire</a>)</p></li><li><p><strong>ENCO unveils enSpeak at NAB Show,</strong> adding real-time voice translation to its captioning workflow so viewers can hear live broadcasts in their preferred language. (<a href="https://content-technology.com/nabshow/enco-enspeak-adds-real-time-voice-translation-to-captioning/">Content + Technology</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>NVIDIA releases Audio Flamingo Next (AF-Next),</strong> an open large audio-language model that understands speech, sound, and music with 30-minute context and timestamp-grounded reasoning. (<a href="https://www.marktechpost.com/2026/04/14/nvidia-and-the-university-of-maryland-researchers-released-audio-flamingo-next-af-next-a-super-powerful-and-open-large-audio-language-model/">MarkTechPost</a>)</p></li></ul><ul><li><p><strong>MOSS-TTS-Nano-100M</strong> brings multilingual voice cloning to CPUs with a 100M-param model that streams 48kHz audio in 20 languages. (<a href="https://hackernoon.com/moss-tts-nano-100m-brings-multilingual-voice-cloning-to-cpus">HackerNoon</a>)</p></li><li><p><strong>Hands-on VibeVoice tutorial</strong> covering speaker-aware ASR, real-time TTS, and speech-to-speech pipelines with code. (<a href="https://www.marktechpost.com/2026/04/12/a-hands-on-coding-tutorial-for-microsoft-vibevoice-covering-speaker-aware-asr-real-time-tts-and-speech-to-speech-pipelines/">MarkTechPost</a>)</p></li><li><p><strong>Build a real-time voice agent with Pipecat,</strong> step-by-step guide to streaming STT/TTS pipelines. (<a href="https://hackernoon.com/how-to-build-a-real-time-voice-agent-with-pipecat">HackerNoon</a>)</p></li><li><p><strong>Build an AI medical scribe</strong> using voice agents for clinical documentation. (<a href="https://hackernoon.com/how-to-build-an-ai-medical-scribe-with-voice-agents">HackerNoon</a>)</p></li><li><p><strong>Diction:</strong> self-hosted STT setup guide as an open alternative to Wispr Flow. (<a href="https://dev.to/omachala/how-to-set-up-diction-the-self-hosted-speech-to-text-alternative-to-wispr-flow-20km">dev.to</a>)</p></li><li><p><strong>Deepgram and Modulate benchmarked</strong> against real-world audio conditions. (<a href="https://hackernoon.com/how-deepgram-and-modulate-benchmark-against-real-world-audio">HackerNoon</a>)</p></li></ul><div><hr></div><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://resources.krisp.ai/fullband-2025&quot;,&quot;text&quot;:&quot;Register now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://resources.krisp.ai/fullband-2025"><span>Register now</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>