<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Voice AI Newsletter]]></title><description><![CDATA[Voice AI insights from Krisp's CEO]]></description><link>https://voice-ai-newsletter.krisp.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!YLgs!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F831a2f7e-d0a7-4e3d-87a8-c42c65d0b71c_1000x1000.png</url><title>Voice AI Newsletter</title><link>https://voice-ai-newsletter.krisp.ai</link></image><generator>Substack</generator><lastBuildDate>Sat, 10 Oct 2026 13:02:13 GMT</lastBuildDate><atom:link href="https://voice-ai-newsletter.krisp.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Krisp Technologies]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[krispai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[krispai@substack.com]]></itunes:email><itunes:name><![CDATA[Davit Baghdasaryan]]></itunes:name></itunes:owner><itunes:author><![CDATA[Davit Baghdasaryan]]></itunes:author><googleplay:owner><![CDATA[krispai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[krispai@substack.com]]></googleplay:email><googleplay:author><![CDATA[Davit Baghdasaryan]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Architecture Behind Customer Service Voice Agents | Sam Krut (Co-Founder & President, Flip)]]></title><description><![CDATA[Watch now (28 mins) | In the Future of Voice AI series of interviews, I ask three questions to my guests: - What problems do you currently see in Enterprise Voice AI? - How does your company solve these problems? - What solutions do you envision in the next 5 years?]]></description><link>https://voice-ai-newsletter.krisp.ai/p/the-architecture-behind-customer</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/the-architecture-behind-customer</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Thu, 08 Oct 2026 14:27:02 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/219015990/fd516b754fb369002a41e723914f781f.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<pre><code><code>In the Future of Voice AI series of interviews, I ask three questions to my guests:

- What problems do you currently see in Enterprise Voice AI?
- How does your company solve these problems?
- What solutions do you envision in the next 5 years?</code></code></pre><p>This episode&#8217;s guest is <a href="https://www.linkedin.com/in/sam-krut/">Sam Krut</a>, Co-Founder and President of <a href="https://flipcx.com/">Flip</a>.</p><p><span>Sam Krut is co-founder and President of Flip, the conversational AI platform helping leading brands transform how they serve customers over the phone. Since co-founding Flip, Sam has worked closely with clients across retail, healthcare, and transportation to deploy AI across some of their highest-volume and most complex customer and patient interactions. He brings a firsthand perspective on how conversational AI is evolving beyond call deflection and cost savings to resolve customer needs, drive business outcomes, and reshape the role of the contact center.</span></p><p>Flip is the conversational Voice AI platform purpose built for retail, healthcare, and transportation. Flip's leading phone AI autonomously resolves customer service calls end-to-end, replacing outdated IVR systems with fast, natural, and brand-consistent experiences at scale. With 300+ million calls processed, 250+ enterprise clients, and 80+ native integrations, Flip is the trusted Voice AI partner for brands including Under Armour, TrustCare, and Curb Mobility. Founded in 2018, Flip is headquartered in New York, NY and has offices in LA and the UK.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8_Ze!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8_Ze!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png 424w, https://substackcdn.com/image/fetch/$s_!8_Ze!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png 848w, https://substackcdn.com/image/fetch/$s_!8_Ze!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!8_Ze!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8_Ze!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png" width="1200" height="1200" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1200,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:417747,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/219015990?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8_Ze!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png 424w, https://substackcdn.com/image/fetch/$s_!8_Ze!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png 848w, https://substackcdn.com/image/fetch/$s_!8_Ze!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!8_Ze!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93412566-d14c-485e-b695-96fd74f700c6_1200x1200.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.youtube.com/@futureofvoiceai&quot;,&quot;text&quot;:&quot;Listen on YouTube&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.youtube.com/@futureofvoiceai"><span>Listen on YouTube</span></a></p><h3><strong>Recap Video</strong></h3><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;1faca13d-47e4-4b43-af93-4073d44e1b71&quot;,&quot;duration&quot;:null}"></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive weekly updates.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><strong>Takeaways</strong></h3><ul><li><p>Vertical depth is becoming a real moat in voice AI because repeatable workflows and integrations matter more than a generic agent that can talk about anything.</p></li><li><p>Voice AI&#8217;s biggest competitive advantage may shift from model quality to system access, because an agent that can&#8217;t act is still just a better IVR.</p></li><li><p>70&#8211;90% automation changes the contact center from a human operation supported by AI into an AI operation with humans handling the exceptions.</p></li><li><p>The real enterprise AI problem is not knowledge retrieval; it is turning undocumented human judgment, exceptions, and workarounds into something machines can reliably execute.</p></li><li><p>Better models alone will not unlock higher automation if the underlying business still runs on weak APIs, fragmented systems, and manual processes.</p></li><li><p>The market may be overbuilding general-purpose voice agents when the strongest economics come from solving the same narrow workflows repeatedly across an industry.</p></li><li><p>&#8220;Containment&#8221; can hide failure because keeping someone away from a human does not matter if the customer has to call back.</p></li><li><p>Outcome-based pricing could become a forcing function for the industry, because vendors only make more money when the customer actually gets more value.</p></li><li><p>The push for one giant model may reverse as smaller models get capable enough to handle speed-critical tasks while larger models handle context and reasoning.</p></li><li><p>Speech-to-speech is still easier to demo than deploy because natural conversation means little if the system loses context or cannot reliably use tools.</p></li><li><p>Voice AI may be closer to technical readiness than organizational readiness, with enterprises now setting the ceiling on automation more than the technology does.</p></li><li><p>If AI agents need the same tools, permissions, and process updates as employees, companies will eventually have to manage them more like digital workers than software features.</p></li><li><p>The rise and fall of dedicated &#8220;AI&#8221; roles would be a sign of success, not failure, because mature AI disappears into the teams and workflows that already own the business outcome.</p></li><li><p>BPOs face a business-model shift more than an extinction event, as value moves away from supplying labor and toward process expertise, exception handling, and AI-enabled operations.</p></li><li><p>As conversation quality improves, voice AI differentiation will move down the stack into latency, orchestration, integrations, and workflow execution.</p></li><li><p>The industry&#8217;s biggest adoption problem may soon be reputation debt: years of bad automated phone experiences have trained customers to expect failure before the conversation even starts.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[ElevenLabs V4 and Microsoft V2 top the charts]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/elevenlabs-v4-and-microsoft-v2-top</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/elevenlabs-v4-and-microsoft-v2-top</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 05 Oct 2026 14:01:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!D30d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/company/krisp-ai/">Krisp</a> publishes an open <a href="https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark">Voice Isolation benchmark</a>, testing 11 STT configurations on real audio with a second voice, where word error rate drops 73%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D30d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D30d!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png 424w, https://substackcdn.com/image/fetch/$s_!D30d!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png 848w, https://substackcdn.com/image/fetch/$s_!D30d!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png 1272w, https://substackcdn.com/image/fetch/$s_!D30d!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D30d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png" width="1456" height="1330" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1330,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4864498,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/218760233?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D30d!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png 424w, https://substackcdn.com/image/fetch/$s_!D30d!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png 848w, https://substackcdn.com/image/fetch/$s_!D30d!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png 1272w, https://substackcdn.com/image/fetch/$s_!D30d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445c8e9b-a681-4b21-9042-f2f7a1d22d11_3284x3000.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Events &#127908;</h2><ul><li><p><strong>Vonage and Deepgram open Deepgram&#8217;s SF HQ for &#8220;Voice AI in the Wild&#8221;</strong> (Oct 6, San Francisco, <a href="https://partiful.com/e/Gp49nL2GPStlf9AlRpms">Partiful</a>)</p></li><li><p><strong>Agora runs &#8220;From Prototype to Production: Scaling Voice AI&#8221;</strong> (Oct 6, San Francisco, <a href="https://partiful.com/e/IL3sKweGcfohlDRJGFkk">Partiful</a>)</p></li><li><p><strong>AssemblyAI hosts a hardware hackathon</strong> (Oct 7, San Francisco, <a href="https://luma.com/w9e4qgol">Luma</a>)</p></li><li><p><strong>Aqua Voice, Speechmatics, and LiveKit tackle &#8220;Solving voice as an interface&#8221;</strong> (Oct 7, San Francisco, <a href="https://partiful.com/e/Qfb44oJOo4cYvr64J7Kw">Partiful</a>)</p></li><li><p><strong>Coval and Cartesia host &#8220;Voices in the Room&#8221;</strong> (Oct 8, San Francisco, <a href="https://partiful.com/e/xn5Zxz9za7u4vEsGd5cf">Partiful</a>)</p></li><li><p><strong>Agora brings EdTech leaders together on voice-first learning</strong> (Oct 8, London, <a href="https://luma.com/cjumnkfs">Luma</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>ElevenLabs launches Eleven v4 and v4 Turbo</strong>, adding automatic expression control, 10-second voice cloning, and 90+ languages (<a href="https://techcrunch.com/2026/09/28/elevenlabs-new-v4-speech-model-supports-more-expression-control-and-90-languages/">TechCrunch</a>)</p></li></ul><ul><li><p><strong>Microsoft ships MAI-Transcribe-2-Streaming</strong> plus MAI-Voice 2.1 models, topping the streaming AA-WER benchmark at 2.5% (<a href="https://www.unite.ai/microsoft-launches-mai-transcribe-2-streaming-and-two-mai-voice-models/">Unite.AI</a>)</p></li><li><p><strong>Alibaba releases Qwen-Audio 3.1 Realtime</strong>, a full-duplex voice model with 262K context and tool calling, cutting voice API prices up to ~85%. (<a href="https://www.marktechpost.com/2026/09/28/alibaba-qwen-releases-qwen-audio-3-1-realtime-a-full-duplex-voice-model-trained-to-think-act-and-decide-when-to-speak/">MarkTechPost</a>)</p></li><li><p><strong>Inworld acquires Ultravox</strong> (formerly Fixie) to add a real-time voice-agent platform with turn-taking and interruption handling to its speech stack. (<a href="https://www.morningstar.com/news/business-wire/20261001868570/ai-research-lab-inworld-acquires-voice-agent-platform-ultravox">Morningstar</a>)</p></li><li><p><strong>Wemnal integrates</strong> Krisp VIVA to improve its voice AI agents (<a href="https://www.linkedin.com/posts/better-turn-taking-more-accurate-speech-share-7508935547386720257-dcvl/?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAhkcvABZW_UaGpyFWVKSyJuZA1bZWc0-4s">LinkedIn</a>) </p></li><li><p><strong>Modulate raises $25M</strong> for voice models and an analysis suite spanning transcription, emotion, deepfake and AI-music detection, and agent policy enforcement. (<a href="https://techcrunch.com/2026/09/28/modulate-raises-25m-for-its-voice-models-and-analysis-suite/">TechCrunch</a>)</p></li><li><p><strong>Instinct raises $1B at a $10B valuation</strong> for a personal AI agent that calls and texts on users&#8217; behalf, quadrupling its valuation in a month. (<a href="https://siliconangle.com/2026/09/28/everyday-personal-ai-assistant-startup-instinct-raises-1b-at-10b-valuation/">SiliconAngle</a>)</p></li><li><p><strong>Relay raises $36M</strong> to turn the frontline push-to-talk radio into an AI assistant that logs and acts on what workers say. (<a href="https://www.pymnts.com/technology/2026/relay-raises-36-million-to-turn-the-frontline-radio-into-an-ai-assistant/">PYMNTS</a>)</p></li><li><p><strong>Presto raises $10M</strong> to expand its drive-thru voice AI across enterprise QSRs, now live in hundreds of locations. (<a href="https://finance.yahoo.com/technology/ai/articles/presto-raises-10-million-accelerate-115800931.html">Yahoo Finance</a>)</p></li><li><p><strong>Deepslate raises $8.8M</strong> to build a Europe-hosted speech-to-speech model it bills as the fastest on Artificial Analysis at 440ms. (<a href="https://techfundingnews.com/deepslate-raises-8-8m-seed-to-build-a-speech-to-speech-ai-model-in-europe/">Tech Funding News</a>)</p></li><li><p><strong>Klang raises &#8364;1.32M</strong> and open-sources Pianissimo, a Swedish speech-to-text model that runs locally and builds on NVIDIA Parakeet. (<a href="https://www.eu-startups.com/2026/09/helsingborgs-conversation-ai-startup-klang-raises-e1-32-million-and-releases-open-speech-to-text-model-for-swedish/">EU-Startups</a>)</p></li><li><p><strong>ElevenLabs doubles its valuation to $22B</strong>, letting employees cash out via a $300M tender offer six months after its $11B raise. (<a href="https://techcrunch.com/2026/09/30/ai-voice-startup-elevenlabs-doubles-valuation-to-22b/">TechCrunch</a>)</p></li><li><p><strong>AI could bring </strong>10x more Voice work to India<strong> (</strong><a href="https://www.businessworld.in/article/ai-could-bring-10x-more-voice-work-to-india-as-accent-barriers-fall-krisp-ceo-626331">BusinessWorld</a><strong>)</strong></p></li><li><p><strong>Fireflies launches Fireflies Talk</strong>, free unlimited voice dictation across apps on Mac and Windows in 100+ languages. (<a href="https://www.globenewswire.com/news-release/2026/09/29/3370905/0/en/fireflies-ai-launches-fireflies-talk-adding-free-unlimited-voice-dictation-to-the-workday.html">GlobeNewswire</a>)</p></li><li><p><strong>Tavus previews Griffin</strong>, a full-duplex video-to-video model that sees, hears, and talks back, with 48% of testers thinking it was human. (<a href="https://therundownai.beehiiv.com/p/tavus-ai-looks-listens-and-talks-back-live">The Rundown AI</a>)</p></li><li><p><strong>Suno launches Speech beta</strong>, generating synthetic voices with matched background music in one track from a prompt or script. (<a href="https://www.theverge.com/ai-artificial-intelligence/1003925/suno-speech-ai-voice-feature-beta-availability">The Verge</a>)</p></li><li><p><strong>Mobvoi announces the TicNote Watch</strong>, a $249 smartwatch that records, transcribes, and summarizes conversations from the wrist offline. (<a href="https://www.prnewswire.com/news-releases/mobvoi-announces-ticnote-watch-the-ai-note-taker-that-captures-your-day-from-your-wrist-302894236.html">PR Newswire</a>)</p></li><li><p><strong>EssilorLuxottica launches Nuance Audio Plus</strong>, second-gen open-ear hearing glasses adding iPhone calls and longer battery, from $829. (<a href="https://www.wareable.com/wearable-tech/nuance-audio-plus-hearing-glasses-second-generation-announcement-availability-price">Wareable</a>)</p></li><li><p><strong>Starkey unveils Omega AI+</strong>, its most advanced hearing aid, built on a new Gen AI neuro processor and quad-DNN architecture. (<a href="https://fortune.com/2026/10/01/billion-dollar-tech-company-starkey-unveils-its-newest-hearing-aid-with-ai-accessible-enough-for-all-generations-to-use/">Fortune</a>)</p></li><li><p><strong>Logitech launches the Zone Vibe Pro</strong>, a boomless headset with AI call noise reduction filtering up to 93% of background sound. (<a href="https://www.cravingtech.com/logitech-zone-vibe-pro-brings-anc-ai-call-noise-reduction-and-66-hour-battery-life.html">CravingTech</a>)</p></li><li><p><strong>RingCentral partners with NovelVox</strong> to wire contact centers into core banking platforms like Fiserv, FIS, and Jack Henry. (<a href="https://www.ringcentral.com/us/en/blog/bridging-the-gap-between-contact-centers-core-financial-platforms-the-ringcentral-novelvox-partnership-advantage/">RingCentral</a>)</p></li><li><p><strong>An AI voice test could screen for Type 2 diabetes</strong>, flagging risk from 20 seconds of speech in a 21,000-person study. (<a href="https://www.forbes.com/sites/fionariley/2026/09/29/ai-voice-test-could-screen-for-type-2-diabetes-research-finds/">Forbes</a>)</p></li><li><p><strong>Deepfake defense draws new capital</strong>, with voice-detection checks forecast to near $5.5B by 2028 as fraud screening scales. (<a href="https://www.biometricupdate.com/202609/deepfake-defense-draws-new-capital-as-voice-detection-checks-near-5-5b-by-2028">Biometric Update</a>)</p></li><li><p><strong>Telecom Reseller argues contact centers should own their voice layer</strong>, with Gladia&#8217;s CEO framing voice AI as core infrastructure, not an add-on. (<a href="https://telecomreseller.com/2026/10/01/why-contact-centers-need-to-own-your-voice-with-voice-ai/">Telecom Reseller</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>FermionResearch releases Phonon-2</strong>, the most accurate open English ASR model under 900MB at 5.21% WER, in a 164MB download. (<a href="https://huggingface.co/FermionResearch/Phonon-2">Hugging Face</a>)</p></li></ul><ul><li><p><strong>A comparison of OpenAI, Gemini, and Qwen realtime voice APIs</strong> finds up to a 213x gap in per-minute cost across the three. (<a href="https://tech-insider.org/openai-vs-google-vs-qwen-voice-ai-apis-2026/">Tech Insider</a>)</p></li><li><p><strong>whisper-local ships offline push-to-talk dictation</strong>, a free faster-whisper tool that types at the cursor in any app, fully on-device. (<a href="https://pypi.org/project/whisper-local/">PyPI</a>)</p></li><li><p><strong>A paper proposes inference-time target-speaker unlearning in ASR</strong>, letting a speaker opt out of meeting transcription without leaving the call. (<a href="https://arxiv.org/abs/2609.30439">arXiv</a>)</p></li><li><p><strong>Linux Journal shows adding speech synthesis to scripts and systems</strong>, using TTS APIs to voice Linux services without hosting a model. (<a href="https://www.linuxjournal.com/content/speech-synthesis-linux-adding-voice-scripts-and-systems">Linux Journal</a>)</p></li><li><p><strong>A breakdown of voice-agent latency</strong> explains why response delay kills calls and how to cut it through the STT, LLM, and TTS stack. (<a href="https://ai.plainenglish.io/ai-voice-agent-latency-why-response-delay-kills-calls-and-how-to-fix-it-f559931ca780">AI in Plain English</a>)</p></li><li><p><strong>A dev essay on why AI voices wear thin after three minutes</strong> dissects the &#8220;30-second trap&#8221; of long-form synthetic narration. (<a href="https://dev.to/julianbrown/why-ai-voices-sound-incredible-for-30-seconds-and-unbearable-after-three-minutes-1mfi">dev.to</a>)</p></li><li><p><strong>Perficient briefs on NotebookLM and Wispr Flow</strong>, pairing source-grounded research with voice dictation for content workflows. (<a href="https://blogs.perficient.com/ai-tools-briefing-notebooklm-and-wispr-flow-ai-powered-research-and-content-creation/">Perficient</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Voice AI moves to Earbuds and Glasses]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/voice-ai-moves-to-earbuds-and-glasses</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/voice-ai-moves-to-earbuds-and-glasses</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 28 Sep 2026 14:01:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ee3b26df-3113-4625-8ed1-67656727502f_1392x780.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events &#127908;</h2><ul><li><p><strong>Decagon runs Decagon Dialogues 2026</strong> (Oct 1, San Francisco, <a href="https://decagon.ai/decagon-dialogues-2026">Decagon</a>)</p></li><li><p><strong>Speechmatics and Tuner co-host a voice AI meetup</strong> (Oct 1, London, <a href="https://luma.com/pyhvutqe">LUMA</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Meta unveils Ray-Ban Meta Audio</strong>, camera-free AI glasses at $349 that pair earbud audio with the Muse assistant, translation, and calls. (<a href="https://www.meta.com/blog/meta-connect-2026-everything-we-announced/">Meta</a>)</p></li><li><p><strong>Krisp publishes an open Voice Isolation benchmark</strong>, testing 11 STT configurations on real audio with a second voice, where word error rate drops 73%. (<a href="https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark">Hugging Face</a>)</p></li><li><p><strong>Google launches </strong>Gemini 3.8 Flash TTS and Flash-Lite TTS (<a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/">Google</a>)</p></li><li><p><strong>Qualcomm announces Snapdragon Sound Elite Gen 2</strong>, an audio chip that runs on-device AI and Wi-Fi in earbuds and glasses without a phone. (<a href="https://www.cnet.com/tech/services-and-software/qualcomm-snapdragon-sound-elite-gen-2-chip-headphones-ai-rebrand/">CNET</a>)</p></li><li><p><strong>SoundHound introduces OASYS Edge</strong>, fully embedded agentic voice AI that runs locally in vehicles and smart devices even offline. (<a href="https://www.globenewswire.com/news-release/2026/09/24/3368388/0/en/soundhound-ai-introduces-oasys-edge-bringing-fully-embedded-agentic-voice-ai-to-vehicles-and-smart-devices.html">GlobeNewswire</a>)</p></li><li><p><strong>Sarvam AI releases Saaras V4</strong>, a speech-to-text model covering all 22 Indian languages plus English with sub-150ms streaming. (<a href="https://www.analyticsinsight.net/artificial-intelligence/sarvam-ai-saaras-v4-how-the-multilingual-speech-model-is-built-for-indias-diverse-languages">Analytics Insight</a>)</p></li><li><p><strong>Inworld ships Realtime TTS-2</strong>, giving developers natural-language control over delivery in live interactions, with a Flash variant at 25ms time-to-first-byte. (<a href="https://www.entrepreneur.com/business-news/tech/how-realtime-voice-ai-is-giving-developers-more-control-over-delivery">Entrepreneur</a>)</p></li><li><p><strong>Sela raises $21M</strong> for voice AI agents that help originate over $1B in mortgage loans a month, used by six of the ten largest independent mortgage banks. (<a href="https://finance.yahoo.com/technology/ai/articles/sela-raises-21m-voice-ai-121150993.html">Yahoo Finance</a>)</p></li><li><p><strong>Dextr AI raises $6.7M</strong> for a hospitality agent platform whose Daisy voice agent books reservations and takes payments in 90+ languages. (<a href="https://www.webwire.com/ViewPressRel.asp?aId=360910">WebWire</a>)</p></li><li><p><strong>Krisp brings real-time voice AI to contact centers</strong>, installing at the OS level to add translation, accent conversion and voice security on any softphone. (<a href="https://pctechmag.com/2026/09/krisp-real-time-voice-ai-built-for-contact-centers-and-bpos/">PC Tech Magazine</a>)</p></li><li><p><strong>A 60-org coalition backed by the Gates Foundation</strong> commits to bringing AI to 3.4 billion people in their own languages within five years. (<a href="https://www.gatesfoundation.org/ideas/progress/ai-for-good/language-commitment">Gates Foundation</a>)</p></li><li><p><strong>A Washington court rules patients can&#8217;t access ambient AI recordings</strong>, treating AI scribe audio as exempt administrative documentation. (<a href="https://www.clinicaladvisor.com/news/ambient-ai-audio-recordings-patient-access-ruling/">Clinical Advisor</a>)</p></li><li><p><strong>Gizmodo argues AI assistants belong in your ears, not on your face</strong>, making the case for earbuds over camera glasses on privacy and social grounds. (<a href="https://gizmodo.com/ai-assistants-belong-in-your-ears-not-on-your-face-2000814863">Gizmodo</a>)</p></li><li><p><strong>Telecoms.com argues voice AI inference belongs inside the network</strong>, where edge deployment cuts latency and adds native deepfake defense. (<a href="https://www.telecoms.com/ai/why-voice-ai-inference-belongs-inside-the-network">Telecoms.com</a>)</p></li><li><p><strong>QSR Web says voice AI has rewritten the drive-thru</strong>, as chains scale order-taking agents to ease staffing and keep the line moving. (<a href="https://www.qsrweb.com/blogs/voice-ai-has-rewritten-the-drive-thru-the-future-of-qsr/">QSR Web</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>NVIDIA releases Nemotron 3 Diarization</strong>, a 100M-parameter open-weight model that tracks up to eight speakers in real time and tops Diarization-Bench. (<a href="https://hackernoon.com/nemotron-3-diarization-heres-what-you-need-to-know">HackerNoon</a>)</p></li><li><p><strong>AWS shows deploying Qwen3-TTS on SageMaker</strong>, with 3-second voice cloning and instruction-driven timbre and emotion across 10 languages. (<a href="https://aws.amazon.com/blogs/machine-learning/deploying-real-time-personalized-speech-with-qwen3-tts-on-amazon-sagemaker-ai/">AWS</a>)</p></li><li><p><strong>AWS packages WhisperX on SageMaker</strong> for word-level, speaker-labeled transcription with alignment, diarization, and subtitle output. (<a href="https://aws.amazon.com/blogs/machine-learning/speaker-labeled-transcription-with-whisperx-on-sagemaker-ai/">AWS</a>)</p></li><li><p><strong>Pariveda built real-time clinical voice notes for Henry Schein One</strong> on Amazon Nova and Bedrock, turning chairside talk into structured notes. (<a href="https://aws.amazon.com/blogs/publicsector/how-pariveda-built-real-time-clinical-voice-notes-for-henry-schein-one-using-amazon-nova-and-amazon-bedrock-agentcore/">AWS</a>)</p></li><li><p><strong>A dev.to guide builds sub-400ms voice agents for enterprise telephony</strong>, covering VAD tuning, streaming STT, and turn-level latency budgets. (<a href="https://dev.to/amit_sharma_3883bddd761cd/building-sub-400ms-real-time-voice-ai-agents-for-enterprise-telephony-4bki">dev.to</a>)</p></li><li><p><strong>SayScroll launches a voice-activated AI teleprompter</strong> that scrolls to a speaker&#8217;s cadence and pauses on hesitation across 60+ languages. (<a href="https://www.dispatch.com/press-release/story/242843/sayscroll-launches-voice-activated-ai-teleprompter-to-eliminate-manual-scrolling-for-video-creators/">Dispatch</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Gemini 3.8 Live, Grok Transcribe 2.0 and much more this week]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/gemini-38-live-grok-transcribe-20</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/gemini-38-live-grok-transcribe-20</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:03:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1120285c-ae59-487a-a383-f5152926c639_1600x1116.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events &#127908;</h2><ul><li><p><strong>Vapi hosts &#8220;Enterprise voice AI after hours&#8221; at HumanX Amsterdam</strong> (Sep 23, Amsterdam, <a href="https://luma.com/d7jixwna">LUMA</a>)</p></li><li><p><strong>AssemblyAI runs &#8220;Build Night: Create Your Own Dictation App&#8221;</strong>  (Sep 24, San Francisco, <a href="https://luma.com/xwnkujzr">LUMA</a>)</p></li><li><p><strong>Google DeepMind hosts &#8220;Gemini Audio | At Night&#8221;</strong> (Sep 24, San Francisco, <a href="https://rsvp.withgoogle.com/events/gemini-audio-at-night">DeepMind</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Google launches Gemini 3.8 Live</strong>, voice model that reason while speaking and tops the Artificial Analysis Speech-to-Speech index. (<a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/">Google</a>)</p></li><li><p><strong>xAI ships Grok Voice Transcribe 2.0</strong>, calling it twice as accurate as v1 and ranking first among streaming models on the Artificial Analysis leaderboard. (<a href="https://x.ai/news/grok-voice-transcribe-2">xAI</a>)</p></li><li><p><strong>ServiceNow</strong> integrates Krisp&#8217;s Voice Isolation to power its Voice Agents (<a href="https://www.linkedin.com/feed/update/urn:li:activity:7505996610229927936/">LinkedIn</a>)</p></li><li><p><strong>Speechmatics launches Agent STT</strong>, a model built to catch the high-consequence errors, like a wrong digit or missed negation, that derail voice agents. (<a href="https://markets.businessinsider.com/news/stocks/speechmatics-launches-agent-stt-for-the-speech-errors-that-derail-voice-agents-1036553575">Business Insider</a>)</p></li><li><p><strong>Apple launches Siri AI</strong>, a rebuilt assistant with personal context, onscreen awareness, and systemwide app actions, rolling out in beta in English. (<a href="https://www.apple.com/newsroom/2026/09/siri-ai-a-profoundly-more-capable-and-personal-assistant-is-here/">Apple</a>)</p></li><li><p><strong>StepFun ships StepAudio 3 Realtime</strong>, a think-while-speaking duplex model that posts 98.9 on the Full-Duplex voice benchmark. (<a href="https://aiweekly.co/alerts/stepaudio-3-realtime-posts-989-on-full-duplex-voice-bench">AI Weekly</a>)</p></li><li><p><strong>Wispr Flow debuts Canto</strong>, its first speech model, cutting noisy real-world error rates from over 30% to single digits (<a href="https://wisprflow.ai/canto">Wispr Flow</a>)</p></li><li><p><strong>DeepL Voice now preserves your voice</strong>, carrying tone, rhythm, and intonation across real-time translation in 30+ languages on Zoom, Teams, and Meet. (<a href="https://www.prnewswire.com/news-releases/deepl-voice-now-preserves-your-voice-in-real-time-multilingual-conversations-302878386.html">PR Newswire</a>)</p></li><li><p><strong>Microsoft adds real-time voice agents to Copilot Studio</strong>, letting organizations build speech-to-speech agents first via Dynamics 365 Contact Center. (<a href="https://www.microsoft.com/en-us/copilot/blog/copilot-studio/extend-ai-voice-support-introducing-real-time-voice-agents-in-microsoft-copilot-studio/">Microsoft</a>)</p></li><li><p><strong>Superhuman acquires AI notetaker Fathom</strong>, folding bot-free meeting capture into its email, calendar, and agent suite. (<a href="https://www.citybiz.co/article/902680/superhuman-acquires-fathom-to-add-ai-meeting-intelligence-across-productivity-platform/">citybiz</a>)</p></li><li><p><strong>Treble raises $18M</strong> to expand its voice simulation platform, generating synthetic acoustic data for model training, robotics, and consumer hardware. (<a href="https://techcrunch.com/2026/09/16/iceland-based-treble-raises-18-million-for-its-voice-simulation-platform/">TechCrunch</a>)</p></li><li><p><strong>Google pushes AI for every language</strong>, extending its speech stack toward the 1,000 most-spoken languages via cross-lingual transfer. (<a href="https://blog.google/innovation-and-ai/technology/ai/ai-for-every-language/">Google</a>)</p></li><li><p><strong>Radisys launches the V.AI ecosystem</strong>, bundling ElevenLabs, Hiya, Speechmatics, and others so telcos can deploy voice AI inside their own networks. (<a href="https://www.newswire.co.kr/newsRead.php?no=1042464">Korea Newswire</a>)</p></li><li><p><strong>Zoom pivots from video calls to agentic work</strong>, positioning ZoomMate to search across systems and finish follow-up tasks, not just summarize. (<a href="https://www.webpronews.com/zooms-high-stakes-ai-pivot-from-video-calls-to-agentic-workflows-that-finish-the-job/">WebProNews</a>)</p></li><li><p><strong>WhatsApp tests speech-to-text messaging</strong>, letting users dictate a message and send it as text, processed on-device and offline. (<a href="https://www.thenews.com.pk/latest/1416849-whatsapp-tests-new-feature-that-turns-speech-into-text">The News</a>)</p></li><li><p><strong>Goodlord warns AI voice spoofing is gaming tenant referencing</strong>, with fraud rings using real-time voice tools to bypass reference callbacks. (<a href="https://blog.goodlord.co/ai-voice-spoofing-fraud-rings-tenant-referencing">Goodlord</a>)</p></li><li><p><strong>Research shows scammers clone local accents</strong>, exploiting familiarity to lower listeners&#8217; guard as AI voice scams surged over 1,200%. (<a href="https://www.androidheadlines.com/2026/09/scammers-using-voice-ai-mimic-local-accents-research.html">Android Headlines</a>)</p></li><li><p><strong>Google DeepMind argues AGI will be spoken, not typed</strong>, making the case for one natively multimodal speech-to-speech model over cascaded pipelines. (<a href="https://finance.biggo.com/news/3ac6a027e3382396">BigGo</a>)</p></li><li><p><strong>FineVoice launches Aunio</strong>, an AI audio production agent that coordinates voices, music, and sound effects into publish-ready audio end-to-end. (<a href="https://www.prnewswire.com/news-releases/finevoice-launches-aunio-an-ai-audio-production-agent-for-end-to-end-audio-creation-302882335.html">PR Newswire</a>)</p></li><li><p><strong>RingCentral rolls out AIR Pro for Healthcare</strong>, an agentic voice AI with prebuilt patient workflows and native integration to 100+ EHR systems. (<a href="https://www.ringcentral.com/us/en/blog/conversational-ai-healthcare/">RingCentral</a>)</p></li><li><p><strong>ImpactFactory.ai ships version 1.5</strong>, adding a full-duplex conversational agent that guides users through marketing tasks. (<a href="https://www.einpresswire.com/article/942886146/impactfactory-ai-releases-version-1-5-with-full-duplex-conversational-ai-agent">EIN Presswire</a>)</p></li><li><p><strong>Fierce argues telcos must embrace voice AI or get left behind</strong>, as Alianza&#8217;s Crux pitches operators a path beyond being a dumb pipe. (<a href="https://www.fierce-network.com/wireless/telcos-can-get-onboard-voice-ai-or-get-left-behind-again">Fierce Network</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Google details building real-time voice apps with Gemini Audio</strong>, its developer guide to the new Live and Extended Thinking models. (<a href="https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/">Google</a>)</p></li></ul><ul><li><p><strong>A benchmark comparison of speech-to-text APIs</strong> finds Meta Muse leading streaming WER at 3.1% while undercutting Google&#8217;s pricing. (<a href="https://tech-insider.org/speech-to-text-api-comparison-2026/">Tech Insider</a>)</p></li><li><p><strong>Nari Labs tops Coval&#8217;s voice benchmarks</strong>, taking #1 STT latency and #1 TTS word error rate in the mid-September run. (<a href="https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/">Nari Labs</a>)</p></li><li><p><strong>HackerNoon explains Grok Voice Realtime</strong>, xAI&#8217;s audio-to-audio model with tools, search, and WebSocket streaming for voice agents. (<a href="https://hackernoon.com/grok-voice-realtime-xais-audio-to-audio-model-explained">HackerNoon</a>)</p></li><li><p><strong>A tutorial builds a Vapi + ElevenLabs appointment agent</strong> that books slots and sends reminders, wired to Google Calendar and Twilio via n8n. (<a href="https://dev.to/samchenreviews/voice-ai-for-appointment-booking-build-a-vapi-elevenlabs-agent-that-books-slots-and-sends-2mge">dev.to</a>)</p></li><li><p><strong>A guide to fine-tuning NVIDIA Nemotron 3.5 ASR</strong> walks through adapting the open model to a language, domain, or accent with NeMo. (<a href="https://dev.to/judy_miranttie/how-to-fine-tune-nvidia-nemotron-35-asr-for-your-language-domain-or-accent-2n8b">dev.to</a>)</p></li><li><p><strong>A paper documents building a production Greek-English recognizer</strong>, evaluating 23 training runs against nine WER, LID, and hallucination gates. (<a href="https://arxiv.org/abs/2609.13498">arXiv</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Apple’s Watch normalizes always-on listening and more this week!]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/apples-watch-normalizes-always-on</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/apples-watch-normalizes-always-on</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 14 Sep 2026 14:00:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/edd5a095-a772-4b50-b81e-a67072297ca9_1702x948.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events &#127908;</h2><ul><li><p><strong>Audio Layer hosts &#8220;Voice x Robotics,&#8221;</strong> an SF meetup with live robot demos (Sep 15, San Francisco, <a href="https://luma.com/uxmg18ib">LUMA</a>)</p></li><li><p><strong>StepFun unveils StepAudio 3 at a live SF meetup,</strong> with real-time demos and a builders&#8217; panel (Sep 16, San Francisco, <a href="https://luma.com/bkpe5h92">LUMA</a>)</p></li><li><p><strong>Regal hosts Regal Rise,</strong> its flagship conference on rebuilding customer interactions around AI voice agents. (Sep 17, NYC + virtual, <a href="https://luma.com/regal-rise-2026">LUMA</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Apple&#8217;s Watch normalizes always-on listening</strong>, with Siri Recap ambiently noting conversations through the day and drawing fresh privacy scrutiny. (<a href="https://www.cnet.com/tech/services-and-software/apple-watch-audio-intelligence-ai-always-listening-commentary/">CNET</a>)</p></li><li><p><strong>Meta Muse Voice Transcribe is built for always-on AI glasses</strong>, with an 80ms listen-or-write loop that raises bystander-consent questions as it diarizes whole rooms. (<a href="https://shattered.io/meta-muse-voice-transcribe-ai-glasses-2026/">shattered.io</a>)</p></li><li><p><strong>OpenAI brings GPT-Live-1 to the API</strong> at $0.05 a minute, giving developers its full-duplex voice model plus 12 new real-time voices. (<a href="https://openai.com/index/introducing-gpt-live-1-in-the-api/">OpenAI</a>)</p></li></ul><ul><li><p><strong>Microsoft ships Azure AI Speech LLM 2607</strong>, cutting latency up to 3x with better mixed-language accuracy for enterprise transcription. (<a href="https://devblogs.microsoft.com/foundry/announcing-azure-ai-speech-llm-2607/">Microsoft</a>)</p></li><li><p><strong>Listen Labs walks away from a $1.5B round</strong> amid talks for Salesforce to acquire the voice-AI research startup for about $2B. (<a href="https://techcrunch.com/2026/09/09/ai-research-startup-listen-labs-scrubbed-a-1-5b-funding-round-for-salesforce-talks/">TechCrunch</a>)</p></li><li><p><strong>Rime partners with Dialpad</strong> as the premier text-to-speech provider behind all of Dialpad&#8217;s agentic voice AI products. (<a href="https://www.webwire.com/ViewPressRel.asp?aId=360266">WebWire</a>)</p></li><li><p><strong>PolyAI ships Wren</strong>, a coding agent inside Dialog Studio that drafts a first version of a voice agent from a natural-language description. (<a href="https://aimagazine.com/news/polyais-wren-the-coding-agent-behind-your-dialogue-agents">AI Magazine</a>)</p></li><li><p><strong>Observe.AI launches Performance Agents</strong>, connecting conversation analysis, coaching, and performance measurement for contact centers. (<a href="https://www.cmswire.com/customer-experience/observeai-launches-ai-coaching-agents-for-contact-centers/">CMSWire</a>)</p></li><li><p><strong>Amazon Prime Video adds AI lip-sync dubbing</strong>, altering on-screen mouth movements to match English dialog on a German-language series. (<a href="https://www.nytimes.com/2026/09/10/arts/television/amazon-prime-video-ai-translation-dub.html">NYT</a>)</p></li><li><p><strong>Bodhan AI, IIT Madras, and NVIDIA release open Indic models</strong> covering speech, translation, text-to-speech, and OCR across 20+ languages. (<a href="https://www.freepressjournal.in/education/bodhan-ai-iit-madras-nvidia-launch-open-ai-models-for-indic-languages-covering-speech-translation-text-to-speech-ocr">Free Press Journal</a>)</p></li><li><p><strong>Krisp reframed as a contact center platform</strong>, extending from noise cancellation and notes into QA, compliance, translation, and voice security. (<a href="https://www.theglobeandmail.com/investing/markets/markets-news/ACN/4497319/beyond-noise-cancellation-and-meeting-notes-how-krisp-became-a-contact-center-platform/">The Globe and Mail</a>)</p></li><li><p><strong>Blue Machines AI launches Aurora</strong>, a BFSI-native speech-to-text model built for India&#8217;s code-switching, noisy telephone conversations. (<a href="https://economictimes.indiatimes.com/ai/ai-insights/blue-machines-ai-unveils-speech-to-text-model-built-for-indias-bfsi-sector/articleshow/133878588.cms">Economic Times</a>)</p></li><li><p><strong>AudioCodes bets big on voice AI</strong> as conversational AI grows over 50% YoY, even as its cash reserves decline. (<a href="http://insidermonkey.com/blog/audiocodes-audc-bets-big-on-voice-ai-while-cash-dwindles-1828542/">Insider Monkey</a>)</p></li><li><p><strong>SAINS and 8nabler sign an MoU</strong> to expand voice AI in Sarawak public services, after handling 4,000+ ICT support calls autonomously. (<a href="https://www.theborneopost.com/2026/09/11/sains-8nabler-sign-mou-to-expand-voice-ai-for-public-services/">Borneo Post</a>)</p></li><li><p><strong>Viaim launches the Rise AI agent earbuds</strong> at IFA 2026, a &#8220;Listen, Think, Act&#8221; hearable that turns conversations into project context. (<a href="https://klgadgetguy.com/viaim-rise-ai-agent-earbuds-launch-ifa-2026/">KLGadgetGuy</a>)</p></li><li><p><strong>SCMP rebuilds its article narration in-house</strong>, replacing external providers with its own AI text-to-speech for natural voice. (<a href="https://www.scmp.com/announcements/article/3367019/high-quality-natural-audio-listen-scmp-articles-improved-narration">SCMP</a>)</p></li><li><p><strong>A new breed of conversational AI is reviving voice in customer service</strong>, as agentic systems move past scripted bots. (<a href="https://itcblogs.currentanalysis.com/2026/09/10/a-new-breed-of-conversational-ai-is-leading-to-voice-resurgence-in-customer-service/">Current Analysis</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>ARTPARK and IISc release the VAANI Noise Event Dataset</strong>, real-world spontaneous speech with timestamped noise across 100+ Indian languages. (<a href="https://huggingface.co/blog/ARTPARK-IISc/a-real-world-dataset-for-noise-robust-speech-ai">Hugging Face</a>)</p></li></ul><ul><li><p><strong>A dev.to walkthrough drives macOS workflows with Gemini 3.5 Transcribe</strong>, combining spoken instructions with on-screen context. (<a href="https://dev.to/alifar/google-gemini-35-transcribe-brings-voice-driven-workflows-to-macos-5gdn">dev.to</a>)</p></li><li><p><strong>Candor-LR is a new audio-visual speech recognition benchmark</strong> built from 1,656 natural dyadic video calls with overlapping, unscripted speech. (<a href="https://arxiv.org/abs/2609.10394">arXiv</a>)</p></li><li><p><strong>A paper tackles Korean ASR error correction</strong> for call centers, introducing a large dialogue-level benchmark for text-based post-editing. (<a href="https://arxiv.org/abs/2609.09889">arXiv</a>)</p></li><li><p><strong>A Frontiers study fine-tunes Wav2Vec2 and Whisper for Telugu</strong>, benchmarking speech-to-text on a low-resource Indian language. (<a href="https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2026.1883049/full">Frontiers</a>)</p></li><li><p><strong>HackerNoon covers Dr. Sunday David Ubur&#8217;s affective architecture</strong>, arguing STT discards the acoustic cues that carry human intent. (<a href="https://hackernoon.com/why-speech-recognition-misses-human-context-dr-sunday-david-uburs-affective-architecture">HackerNoon</a>)</p></li><li><p><strong>A &#8220;self-listening&#8221; paper targets voice AI interruptions</strong>, giving agents a way to monitor their own output to handle barge-in. (<a href="https://aiweekly.co/alerts/self-listening-paper-targets-voice-ai-anchor-interruption">AI Weekly</a>)</p></li><li><p><strong>Audio8 releases TTS Preview 0.6B</strong>, a compact multilingual model with zero-shot voice cloning across 11 languages. (<a href="https://hackernoon.com/audio8-tts-preview-06b-the-multilingual-text-to-speech-model-that-supports-speech-generation">HackerNoon</a>)</p></li><li><p><strong>Aqua Voice targets professionals with the Avalon model</strong>, a voice-to-text tool that beats Whisper and Scribe on technical and coding vocab. (<a href="https://dynamicbusiness.com/ai-tools/aqua-voice-voice-to-text-for-professionals.html">Dynamic Business</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[STT wars continue: new models from Meta, MS, IBM and more]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/stt-wars-continue-new-models-from</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/stt-wars-continue-new-models-from</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 07 Sep 2026 14:01:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/58b0dd9d-6b14-4f78-bbc3-f4787d791076_1890x1058.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Top Updates &#128170;</h2><ul><li><p><strong>Meta launches Muse Voice Transcribe</strong>, one real-time model for streaming ASR, diarization, and endpointing across 70+ languages at $0.18 an hour. (<a href="https://research.meta.ai/blog/introducing-muse-voice-transcribe">Meta</a>)</p></li><li><p><strong>Microsoft ships MAI-Transcribe-2</strong>, undercutting OpenAI, Google, and ElevenLabs on price at $0.10/hour (<a href="https://venturebeat.com/infrastructure/microsoft-ais-mai-transcribe-2-undercuts-openai-google-and-elevenlabs-on-price-and-speed">VentureBeat</a>)</p></li><li><p><strong>IBM releases Granite Speech 5.0 TurboCTC</strong>, a 470M open model transcribing 3.5 hours of speech per second by dropping the decoder. (<a href="https://slator.com/ibm-granite-speech-5-transcription-model/">Slator</a>)</p></li><li><p><strong>Phonely launches Alma</strong>, a voice LLM trained on 10M+ phone calls that runs sub-185ms and is 84% cheaper than GPT-4.1. (<a href="https://siliconangle.com/2026/09/01/phonely-launches-alma-a-voice-ai-model-trained-on-10m-phone-conversations/">SiliconAngle</a>)</p></li><li><p><strong>Deepdub launches Phantom Z 3.4 Conversational</strong>, a multilingual TTS with 150ms time-to-first-audio built to survive real customer calls. (<a href="https://www.prnewswire.com/il/news-releases/deepdub-launches-phantom-z-3-4-conversational-multilingual-text-to-speech-built-to-survive-real-customers-not-just-demos-302869134.html">PR Newswire</a>)</p></li><li><p><strong>Google adds voice to Gmail, Docs, and Keep</strong>, letting users talk to their inbox and dictate notes via Gemini-powered Live modes. (<a href="https://www.techtimes.com/articles/326590/20260904/google-workspace-launches-voice-ai-docs-live-costs-4x-more-gmail-live.htm">TechTimes</a>)</p></li><li><p><strong>Google Translate adds background live translation</strong> on Android, keeping translation running when you switch apps or lock the screen. (<a href="https://www.webpronews.com/google-translates-new-background-mode-lets-live-translation-run-free-on-android/">WebProNews</a>)</p></li><li><p><strong>Genesys expands agentic voice AI at Xperience 2026</strong>, adding end-of-turn detection and native Deepgram and ElevenLabs integrations. (<a href="https://www.destinationcrm.com/Articles/ReadArticle.aspx?ArticleID=176451">DestinationCRM</a>)</p></li><li><p><strong>Dialpad and Glean connect conversation insights to enterprise knowledge</strong>, bringing call intelligence into Glean&#8217;s knowledge platform. (<a href="https://www.dialpad.com/press/dialpad-and-glean-connect-customer-conversation-insights-with-enterprise-knowledge/">Dialpad</a>)</p></li><li><p><strong>Digital Island partners with RingCentral</strong> as its first New Zealand partner to bring agentic voice AI to local businesses. (<a href="https://www.stocktitan.net/news/RNG/digital-island-and-ring-central-partner-to-bring-agentic-voice-ai-i5l20dh35jjh.html">StockTitan</a>)</p></li><li><p><strong>Krisp CEO Davit Baghdasaryan on transforming enterprise communication</strong>, discussing how voice AI is reshaping the enterprise stack. (<a href="https://www.analyticsinsight.net/amp/story/podcast/davit-baghdasaryan-on-how-voice-ai-is-transforming-enterprise-communication">Analytics Insight</a>)</p></li><li><p><strong>Walmart is sued over AI voiceprints</strong> in a proposed Illinois class action alleging it built biometric templates from calls without BIPA consent. (<a href="https://travel-news.co.uk/travel/walmart-biometric-voiceprint-lawsuit-filed-over-illinois-customer-calls/38251/">Travel News</a>)</p></li><li><p><strong>SwitchBot launches the AI MindClip</strong>, a 16.8g clip-on wearable powered by Qwen that turns conversations into summaries and to-dos. (<a href="https://incrediblethings.com/artificial-intelligence/switchbot-ai-mindclip-wearable-ai-assistant/">Incredible Things</a>)</p></li><li><p><strong>TikTok adds voice comments</strong>, letting users leave 60-second audio replies alongside comment polls and photo carousels. (<a href="https://techmymoney.com/2026/09/05/tiktok-voice-comments-add-60-second-audio-replies/">TechMyMoney</a>)</p></li><li><p><strong>Tesla wires Grok Voice Think Fast 2.0 into the Summer Update</strong>, giving in-cabin voice full vehicle control across 100+ commands. (<a href="https://www.notateslaapp.com/news/4638/tesla-now-uses-grok-think-fast-20-with-summer-update">Not a Tesla App</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Nokia introduces Automatic Contextual Audio Denoising</strong>, a task where the system infers scene context before deciding what counts as noise. (<a href="https://www.nokia.com/blog/when-noise-depends-on-context-introducing-automatic-contextual-audio-denoising/">Nokia</a>)</p></li></ul><ul><li><p><strong>A TTFT-first benchmark ranks lowest-latency inference APIs</strong> for voice agents, finding Groq&#8217;s LPU leads on time-to-first-token. (<a href="https://www.marktechpost.com/2026/08/30/lowest-latency-inference-apis-for-voice-and-realtime-agents-a-time-to-first-token-ttft-first-benchmark/">MarkTechPost</a>)</p></li><li><p><strong>A hands-on comparison of GPT-Live, Gemini Live, and Grok Voice</strong> breaks down latency, pricing, and interruption handling across the three. (<a href="https://tech-insider.org/ie/gpt-live-vs-gemini-live-vs-grok-voice-2026/">Tech Insider</a>)</p></li><li><p><strong>VoiceStudio is a fully local ElevenLabs alternative</strong>, bundling 16 TTS and 11 ASR engines with a local API and MCP in 646 languages. (<a href="https://dev.to/sloves/why-debpalashvoicestudio-is-gaining-attention-local-voice-ai-without-a-hosted-api-303d">dev.to</a>)</p></li><li><p><strong>Vowen 0.5.6 ships offline voice dictation</strong>, a free system-wide tool running local Whisper models across 99 languages with no cloud. (<a href="https://www.warp2search.net/story/vowen-056-released/">Warp2Search</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Google launches Gemini 3.5 Transcribe, SoundHound Buys LivePerson and more]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/google-launches-gemini-35-transcribe</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/google-launches-gemini-35-transcribe</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 31 Aug 2026 14:03:40 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/45767125-d026-4be8-a308-a191c020b54c_2628x1306.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events &#127908;</h2><ul><li><p><strong>AssemblyAI:</strong> gathering builders and founders for demos and networking on shipping production voice agents. (Sep 1, New York, <a href="https://luma.com/v4fbpqof">LUMA</a>)</p></li><li><p><strong>Vapi and Cartesia:</strong> hands-on session on turn detection, interruption handling, latency, and locale support inside Cartesia&#8217;s Playground. (Sep 2, <a href="https://luma.com/jnb3cfym">LUMA</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Google releases Gemini 3.5 Transcribe</strong>, a speech-to-text model reporting 2.6% average WER across 85+ languages while stripping fillers and fixing slips. (<a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/">Google</a>)</p></li><li><p><strong>SoundHound moves to acquire LivePerson</strong>, combining voice agentic AI with digital messaging to build an end-to-end omnichannel platform. (<a href="https://www.constellationr.com/insights/news/soundhound-ai-liveperson-deal-may-bolster-its-ai-agent-ambitions">Constellation Research</a>)</p></li><li><p><strong>Krisp&#8217;s VIVA powers Telnyx&#8217;s voice AI stack</strong>, adding upstream voice isolation so agents transcribe more accurately and avoid false interruptions at 10M+ call minutes a day. (<a href="https://www.linkedin.com/posts/krisphq_10m-call-minutes-a-day-leaves-very-little-activity-7499107308380119041-b4nx/">Krisp</a>)</p></li><li><p><strong>Ringg raises $15M</strong> to push its voice AI orchestration layer past the phone call into WhatsApp and browser agents. (<a href="https://slator.com/ringg-raises-15m-funding-voice-ai-orchestration/">Slator</a>)</p></li><li><p><strong>Plaud launches the One AI earbuds</strong>, a $249 wearable with an LTE-enabled case that records and transcribes meetings without a phone. (<a href="https://9to5google.com/2026/08/27/plaud-one-is-a-pair-of-air-recording-earbuds/">9to5Google</a>)</p></li><li><p><strong>Ex-Nothing founders launch Relay Q</strong>, a $150 dictation mic and macOS app running on Gemini for quiet, accurate voice-to-text. (<a href="https://www.wired.com/story/relay-q-voice-to-text-ai-app/">Wired</a>)</p></li><li><p><strong>DeepL research finds live voice translation is the AI Americans want most</strong>, chosen over meeting summaries and email drafting. (<a href="https://www.prnewswire.com/news-releases/live-voice-translation-is-the-ai-americans-want-most-ahead-of-email-drafting-and-meeting-notes-new-deepl-research-finds-302858925.html">PR Newswire</a>)</p></li><li><p><strong>Intron&#8217;s new voice model handles African code-switching</strong>, following speakers who change languages mid-sentence for clinics and contact centers. (<a href="https://techcabal.com/2026/08/26/intron-voice-ai/">TechCabal</a>)</p></li><li><p><strong>3CLogic releases an AI Agent Evaluator</strong>, auto-scoring 100% of voice AI interactions on resolution and quality without manual review. (<a href="https://www.prnewswire.com/news-releases/3clogic-releases-ai-agent-evaluator-to-automate-qa-and-scoring-of-voice-ai-agents-302858764.html">PR Newswire</a>)</p></li><li><p><strong>Fireflies wants to turn its notetaker into an action-taker</strong>, with CEO Krish Ramineni pushing AI teammates that execute post-meeting work. (<a href="https://economictimes.indiatimes.com/ai/ai-insights/fireflies-ceo-krish-ramineni-wants-to-turn-its-ai-notetaker-into-an-action-taker/articleshow/133459364.cms">Economic Times</a>)</p></li><li><p><strong>Crexendo partners with Tresic</strong> to turn every customer conversation into AI-driven tasks and alerts for its 250+ service-provider licensees. (<a href="https://www.newswire.com/news/crexendo-partners-with-tresic-to-turn-every-customer-conversation-into-ai">Newswire</a>)</p></li><li><p><strong>A blind Egyptian entrepreneur builds ScribeMe</strong>, an AI app giving running audio descriptions of surroundings, now used in 140 countries. (<a href="https://timesofindia.indiatimes.com/world/rest-of-world/a-blind-egyptian-entrepreneur-built-an-ai-app-that-describes-the-world-through-audio-helping-visually-impaired-users-experience-wildlife-objects-and-surrounding-scenes-through-their-phones/articleshow/133463516.cms">Times of India</a>)</p></li><li><p><strong>CX Today argues agentic CX exposes operational gaps</strong>, as ServiceNow, Salesforce, and Synthflow shift the question to running autonomy with accountability. (<a href="https://www.cxtoday.com/ai-automation-in-cx/servicenow-salesforce-and-synthflow-expose-the-operational-issues-behind-agentic-cx/">CX Today</a>)</p></li><li><p><strong>Unauthorized AI celebrity voices are spreading in promotion</strong>, raising right-of-publicity and false-endorsement fights as the NO FAKES Act advances. (<a href="https://www.bignewsnetwork.com/news/279268203/when-celebrities-never-said-it-the-funny-rise-of-unauthorized-ai-voices-in-promotion">Big News Network</a>)</p></li><li><p><strong>TechRadar dissects why ChatGPT&#8217;s new voice works</strong>, finding its &#8220;umms,&#8221; pauses, and breaths are what make it feel human. (<a href="https://www.techradar.com/ai-platforms-assistants/chatgpt/i-spent-20-minutes-counting-everything-i-hate-about-chatgpts-new-voice-and-accidentally-discovered-why-it-works">TechRadar</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Daily releases Pipecat PhoneLLM Alpha 1</strong>, an open-weights voice-agent LLM that matches GPT-5.6 Terra at 94% lower cost and faster time-to-first-token. (<a href="https://www.daily.co/blog/announcing-pipecat-phonellm-alpha-1/">Daily</a>)</p></li><li><p><strong>Krisp ships an MCP upgrade for meeting data</strong>, adding smarter search, bulk AI tagging, bulk speaker correction, and AI-managed action items. (<a href="https://www.linkedin.com/feed/update/urn:li:activity:7498034718899142656/">Krisp</a>)</p></li><li><p><strong>Google shows how to evaluate live voice agents in ADK</strong>, driving agents with simulated users that speak their turns as Gemini TTS audio. (<a href="https://developers.googleblog.com/how-to-evaluate-live-voice-agents-in-adk/">Google Developers</a>)</p></li><li><p><strong>A guide to evaluating voice agents with LangSmith</strong> scores them on real outcomes: bookings made, tickets resolved, transfers landed. (<a href="https://dev.to/focused_dot_io/how-to-evaluate-voice-agents-with-langsmith-focused-ogc">dev.to</a>)</p></li><li><p><strong>A deep dive on the 2000ms vs 250ms architecture war</strong> explains why cascaded voice stacks still beat speech-to-speech in production. (<a href="https://pub.towardsai.net/2000ms-vs-250ms-the-hidden-architecture-war-behind-every-voice-ai-product-0ffe2afffe9b">Towards AI</a>)</p></li><li><p><strong>A complete guide to Gemini 3.5 Transcribe</strong> walks through verbatim timestamps vs. clean smart-mode output with runnable Colab code. (<a href="https://dev.to/googleai/stop-wrestling-with-asr-the-complete-guide-to-gemini-35-transcribe-1m6i">dev.to</a>)</p></li><li><p><strong>A tutorial builds a production audio transcription pipeline in Python</strong>, wiring up streaming STT with timestamps and diarization. (<a href="https://dev.to/smallestai/how-to-build-a-production-ready-audio-transcription-pipeline-in-python-2of9">dev.to</a>)</p></li><li><p><strong>A paper fine-tunes NVIDIA Canary Flash for telephony ASR</strong>, adapting compact models to noisy contact-center audio and jargon. (<a href="https://arxiv.org/abs/2608.24916">arXiv</a>)</p></li><li><p><strong>TLive-Omni is an omni-modal model for live commerce</strong>, fusing speech, video, and text with timestamped tokens for real-time understanding. (<a href="https://arxiv.org/abs/2608.20958">arXiv</a>)</p></li><li><p><strong>HackerNoon shows how Audio8 TTS-0.1B brings voice cloning to smaller GPUs</strong>, a ~170M-parameter model with zero-shot cloning. (<a href="https://hackernoon.com/how-audio8-tts-01b-brings-voice-cloning-to-smaller-gpus">HackerNoon</a>)</p></li><li><p><strong>Vibe is a free open-source local transcription app</strong> for Windows, macOS, and Linux that runs Whisper-based models fully offline. (<a href="http://majorgeeks.com/files/details/vibe.html">MajorGeeks</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Wispr hits $2B, Krisp and 8x8 partner, Cartesia Tops the Arena]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/wispr-hits-2b-krisp-and-8x8-partner</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/wispr-hits-2b-krisp-and-8x8-partner</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 24 Aug 2026 14:03:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/374ef0bc-e6be-4715-b136-9c80b8c6e750_1792x1024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Top Updates &#128170;</h2><ul><li><p><strong>Wispr raises $280M at a $2B valuation</strong> as it moves beyond dictation, previewing Canto, a speech model that cuts noisy-condition error rates sharply. (<a href="https://siliconangle.com/2026/08/17/wispr-raises-280m-to-power-up-natural-speech-to-text-using-ai/">SiliconAngle</a>)</p></li></ul><ul><li><p><strong>Cartesia ships Sonic 3.6</strong>, a streaming TTS model that now tops both Artificial Analysis speech arenas with sub-90ms time-to-first-audio across 44 languages. (<a href="https://slator.com/cartesia-launches-text-to-speech-model/">Slator</a>)</p></li><li><p><strong>Adobe Firefly adds AI music, speech, and sound effects</strong>, folding audio generation into its creative studio with commercial usage rights. (<a href="https://blog.adobe.com/en/publish/2026/08/20/adobe-firefly-expands-its-creative-ai-studio-generate-music-speech-and-sound-effects-in-one-place">Adobe</a>)</p></li><li><p><strong>HappyRobot raises $150M at a $1.2B valuation</strong> to run freight operations by voice, automating broker and carrier calls end to end. (<a href="https://forkast.news/happyrobot-wants-1-2-billion-to-run-your-freight-by-voice/">Forkast</a>)</p></li><li><p><strong>Krisp joins the 8x8 Technology Partner Ecosystem</strong>, bringing its noise cancellation and accent conversion to the 8x8 CX platform. (<a href="https://martechseries.com/predictive-ai/ai-platforms-machine-learning/krisp-joins-8x8-technology-partner-ecosystem-to-deliver-clearer-voice-experiences/">MarTech Series</a>)</p></li><li><p><strong>Murf AI launches Falcon 2</strong>, a voice model priced at $0.01/min that ranks ahead of OpenAI Realtime and ElevenLabs Flash on naturalness. (<a href="https://www.storyboard18.com/brand-marketing/bengaluru-startup-murf-ais-falcon-2-takes-on-openai-elevenlabs-with-lower-cost-voice-ai-108272.htm">Storyboard18</a>)</p></li><li><p><strong>Sarvam&#8217;s Saaras V3 outperforms OpenAI and ElevenLabs</strong> on Indian-language and Indian-accented speech benchmarks. (<a href="https://cryptobriefing.com/sarvam-ai-outperforms-openai-elevenlabs-indian-voice/">CryptoBriefing</a>)</p></li><li><p><strong>Breeze Blue unveils Breeze TTS 2</strong>, a real-time model for games and digital companions that tops the TTS Voice Design benchmark. (<a href="https://www.cincinnati.com/press-release/story/110150/breeze-blue-unveils-breeze-tts-2-real-time-flagship-voice-ai-for-interactive-media/">Cincinnati.com</a>)</p></li><li><p><strong>AudioStack adds Gradium&#8217;s voices</strong> to its AI audio platform, expanding regional accent coverage for media and advertising customers. (<a href="https://slator.com/audiostack-adds-gradium-voices-ai-audio-platform/">Slator</a>)</p></li><li><p><strong>AudioCodes and AnywhereNow earn Microsoft Teams certification</strong>, with AudioCodes&#8217; Voca CIC among the first certified Teams voice agents. (<a href="https://www.speechtechmag.com/Articles/ReadArticle.aspx?ArticleID=176180">SpeechTech</a>)</p></li><li><p><strong>Reddit tests turning text posts into AI-narrated videos</strong>, adding a &#8220;Play&#8221; toggle that voices threads to capture content it once lost to TikTok. (<a href="https://www.theverge.com/tech/981289/reddit-ai-text-video-posts">The Verge</a>)</p></li><li><p><strong>Calendly launches Notetaker and the Callie assistant</strong>, focusing AI on the follow-up work that happens after a meeting ends. (<a href="https://autogpt.net/calendly-ai-meeting-notetaker-callie-assistant/">AutoGPT</a>)</p></li><li><p><strong>Boya unveils the Notra Neo AI note taker</strong>, a $149 device with dual mics and 30dB noise reduction for meetings, calls, and Bluetooth audio. (<a href="https://chinagadgetsreviews.com/the-next-gen-ai-note-taker-is-taking-shape-boya-notra-neo.html">China Gadgets Reviews</a>)</p></li><li><p><strong>JPMorgan Chase warns AI is making social engineering easier</strong>, as attackers clone voices from short clips to impersonate staff. (<a href="https://www.jpmorganchase.com/about/technology/blog/ai-is-making-social-engineering-easier-than-ever">JPMorgan Chase</a>)</p></li><li><p><strong>The ABA examines the voice fraud threat to banking</strong>, noting detection still degrades under real phone-line noise and compression. (<a href="https://bankingjournal.aba.com/2026/08/the-voice-fraud-threat-to-banking/">ABA Banking Journal</a>)</p></li><li><p><strong>Logitech says audio accuracy is the key</strong> as voice dictation spreads through offices, with one in five AI interactions now voice-based. (<a href="https://itbrief.com.au/story/logitech-says-audio-accuracy-key-as-ai-voice-dictation-grows">IT Brief</a>)</p></li></ul><div><hr></div><h2>Last Week Podcast </h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;e3a95f18-8739-4d82-82a3-65f49a25a317&quot;,&quot;caption&quot;:&quot;In the Future of Voice AI series of interviews, I ask three questions to my guests: - What problems do you currently see in Enterprise Voice AI? - How does your company solve these problems? - What solutions do you envision in the next 5 years?&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Voice AI for Frontline Safety | Arpan Podduturi (VP of Product, Samsara)&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:32916364,&quot;name&quot;:&quot;Davit Baghdasaryan&quot;,&quot;bio&quot;:&quot;CEO &amp; Co-Founder of Krisp, early pioneer in Voice AI.\n20+ years in engineering. 18 US patent applications, ex Twilion&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/23088dde-6cb0-44df-b220-5f22830cdd4c_1179x960.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-20T14:25:39.461Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8220eb53-915a-41c8-b1c3-1ebc847f677f_1165x776.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/p/voice-ai-for-frontline-safety-arpan&quot;,&quot;section_name&quot;:&quot;Podcast&quot;,&quot;video_upload_id&quot;:&quot;0593f5ce-bb8a-4a1c-b55d-0cd5d6d1cce7&quot;,&quot;id&quot;:211250413,&quot;type&quot;:&quot;podcast&quot;,&quot;reaction_count&quot;:15,&quot;comment_count&quot;:1,&quot;publication_id&quot;:2073467,&quot;publication_name&quot;:&quot;Voice AI Newsletter&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!YLgs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F831a2f7e-d0a7-4e3d-87a8-c42c65d0b71c_1000x1000.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>superwhisper open-sources S1-mini</strong>, a 462MB text normalizer that turns raw ASR transcripts into clean written text on a laptop CPU. (<a href="https://www.marktechpost.com/2026/08/20/meet-s1-mini-superwhispers-462-mb-open-weights-text-normalizer-that-turns-raw-asr-transcripts-into-clean-written-text/amp/">MarkTechPost</a>)</p></li></ul><ul><li><p><strong>Nari Labs pushes the Qwen3-TTS speed-cost frontier</strong>, hitting sub-50ms time-to-first-audio at ~$2 per 1M characters on a single H100. (<a href="https://nari-labs.com/blog/qwen3-tts-speed-cost-frontier/">Nari Labs</a>)</p></li><li><p><strong>A HackerNoon team builds the same STT app cloud and local</strong>, comparing cost, latency, and accuracy trade-offs of each path. (<a href="https://hackernoon.com/cloud-or-local-speech-to-text-we-built-the-same-app-both-ways-heres-what-we-learnt">HackerNoon</a>)</p></li><li><p><strong>A dev.to post tests 4 TTS engines on 12,000 live healthcare calls</strong>, reporting which one patients responded to best in production. (<a href="https://dev.to/autor_tech/we-tested-4-text-to-speech-engines-on-12000-live-healthcare-calls-heres-which-one-patients-5b88">dev.to</a>)</p></li><li><p><strong>A guide on streaming ASR vs Whisper on mobile</strong> explains when to switch to a streaming model for sub-second, interactive voice. (<a href="https://dev.to/voxrtio/streaming-asr-vs-whisper-on-mobile-when-to-switch-5cm7">dev.to</a>)</p></li><li><p><strong>A developer explains why WhatsApp voice notes break transcription</strong>, from background noise to code-switching and mixed languages. (<a href="https://dev.to/talha_hussain/why-whatsapp-voice-notes-break-general-purpose-transcription-4nfp">dev.to</a>)</p></li><li><p><strong>A dev builds a free no-signup audio-and-video-to-text tool</strong>, sharing the pipeline behind browser-based transcription. (<a href="https://dev.to/nadiakesslerdev/show-dev-i-built-a-free-no-signup-tool-that-turns-audio-and-video-into-text-4f45">dev.to</a>)</p></li><li><p><strong>A developer builds a browser voice-to-writing tool</strong>, arguing typing slows down thinking and dictation keeps ideas flowing. (<a href="https://dev.to/hussain_jatoi/i-built-a-browser-voice-to-writing-tool-because-typing-slows-down-thinking-4f9a">dev.to</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Voice AI for Frontline Safety | Arpan Podduturi (VP of Product, Samsara)]]></title><description><![CDATA[Watch now | In the Future of Voice AI series of interviews, I ask three questions to my guests: - What problems do you currently see in Enterprise Voice AI? - How does your company solve these problems? - What solutions do you envision in the next 5 years?]]></description><link>https://voice-ai-newsletter.krisp.ai/p/voice-ai-for-frontline-safety-arpan</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/voice-ai-for-frontline-safety-arpan</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Thu, 20 Aug 2026 14:25:39 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211250413/895ae8c62c4ccc119b3d37a0114c7d8a.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<pre><code><code>In the Future of Voice AI series of interviews, I ask three questions to my guests:

- What problems do you currently see in Enterprise Voice AI?
- How does your company solve these problems?
- What solutions do you envision in the next 5 years?</code></code></pre><p>This episode&#8217;s guest is <a href="https://www.linkedin.com/in/arpanpodduturi/">Arpan Podduturi</a>, VP of Product at <a href="https://www.samsara.com/">Samsara</a>.</p><p><span>Arpan Podduturi is Vice President of Product at Samsara, where he leads strategy and development of physical AI products that help protect the frontline workers who make up 80% of the global workforce. Before joining Samsara, Arpan served as Vice President of Product at Shopify, overseeing Shop App and agentic shopping. Prior to Shopify, Arpan led product management for Etsy&#8217;s seller businesses and mobile products as well as ESPN&#8217;s flagship mobile app&#8212;the world&#8217;s most downloaded sports app. He has also held product and design roles at Disney, Viacom, and the National Hockey League, and founded his own direct-to-consumer brand.</span></p><p>Samsara is the pioneer of the Connected Operations&#174; Platform, which is an open platform that connects the people, devices, and systems of some of the world's most complex operations, allowing them to develop actionable insights and improve their operations. With tens of thousands of customers across North America and Europe, Samsara is a proud technology partner to the people who keep our global economy running, including the world's leading organizations across industries in transportation, construction, wholesale and retail trade, field services, logistics, manufacturing, utilities and energy, government, healthcare and education, food and beverage, and others. The company's mission is to increase the safety, efficiency, and sustainability of the operations that power the global economy.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LJb3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LJb3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LJb3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LJb3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LJb3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LJb3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg" width="1200" height="1200" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1200,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:397614,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/211250413?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LJb3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LJb3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LJb3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LJb3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a79e179-ffd4-43ec-9a6c-ce570aa76279_1200x1200.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.youtube.com/@futureofvoiceai&quot;,&quot;text&quot;:&quot;Listen on YouTube&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.youtube.com/@futureofvoiceai"><span>Listen on YouTube</span></a></p><h3><strong>Recap Video</strong></h3><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;f4655ecc-ffe6-4670-9673-758aa29fde5a&quot;,&quot;duration&quot;:null}"></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive weekly updates.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><strong>Takeaways</strong></h3><ul><li><p>The biggest overlooked AI market may be the 80% of workers who cannot sit at a screen while doing their jobs.</p></li><li><p>Samsara&#8217;s vision starts with giving every driver hands-free access to a superintelligent assistant inside the truck.</p></li><li><p>Voice AI can provide safety coverage when human dispatchers are unavailable, including during peak drowsiness hours from 11 p.m. to 4 a.m.</p></li><li><p>A voice agent that keeps a tired driver alert until the next rest stop turns AI from a productivity tool into potentially life-saving infrastructure.</p></li><li><p>Replacing unexplained warning beeps with clear spoken guidance can improve behavior because workers understand what happened and what to do next.</p></li><li><p>Automating routine calls does more than reduce labor; it makes personalized guidance possible across thousands of vehicles at once.</p></li><li><p>Geofenced voice agents can deliver instructions at the exact moment they become relevant, removing the need for dispatchers to track and call every driver.</p></li><li><p>The clearest AI business cases come from narrow workflows where every minute saved maps directly to lower costs or higher margins.</p></li><li><p>Proprietary operational data is what makes a voice agent useful at work, not simply access to a powerful language model.</p></li><li><p>Highly structured voice conversations currently outperform open-ended ones, challenging the belief that more freedom always creates a better AI experience.</p></li><li><p>Building many effective agents creates a new problem when they lack shared context and begin competing for the same user&#8217;s attention.</p></li><li><p>Agent orchestration will become as important as speech quality because a safety warning must take priority over a routine notification.</p></li><li><p>Technical performance alone will not drive adoption when workers may see an unexpected talking camera as surveillance, interference, or disrespect.</p></li><li><p>Personalization must include how much guidance each worker wants, since one driver may welcome an AI companion while another feels patronized.</p></li><li><p>The next major shift is from agents built around management workflows to one assistant built around the frontline worker&#8217;s full day.</p></li><li><p>The end state is a hands-free work assistant that unifies routes, safety policies, job instructions, and conversation history without forcing workers to check a phone.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Exciting updates from Nvidia, Krisp, Deepgram, Soniox and more]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/exciting-updates-from-nvidia-krisp</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/exciting-updates-from-nvidia-krisp</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 17 Aug 2026 14:00:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a6918910-5e50-476e-9547-7e58634d94ed_1744x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>Voice AI Space &#8220;Global Night&#8221; mixer across five cities,</strong> a no-talks, no-demos networking night for founders, engineers, and researchers (Aug 20, SF + London + Paris + Barcelona + Toronto, <a href="https://www.voiceaispace.com/events/voice-ai-community-mixer-sf">Voice AI Space</a>)</p></li><li><p><strong>Dex and OpenAI dig into production voice AI engineering at a London meetup,</strong> with talks on turn-taking, interruption handling, and latency from speakers at each company, followed by networking. (Aug 19, London, <a href="https://luma.com/dex-foeb">LUMA</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>NVIDIA releases NemotronLabs VoiceChat 11B</strong>, an open full-duplex speech-to-speech model  (<a href="https://www.marktechpost.com/2026/08/09/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech-model-with-450-ms-turn-taking-and-live-tool-calling/">MarkTechPost</a>)</p></li></ul><ul><li><p><strong>Krisp launches Voice Isolation v2.5</strong> that improves STT word error rates (46% drop) in the most challenging multi-speaker cases <strong>(</strong><a href="https://krisp.ai/blog/voice-isolation-2-5/">Krisp Blog</a><strong>)</strong></p></li><li><p><strong>Deepgram launches Flux TTS</strong>, a conversation-native text-to-speech model that tracks tone across turns and starts output in as low as 80ms. (<a href="https://deepgram.com/learn/introducing-flux-tts-conversation-native-text-to-speech-for-real-time-voice-agents">Deepgram</a>)</p></li><li><p><strong>Alibaba launches CosyVoice Studio</strong>, China&#8217;s first full-stack AI voice productivity platform built on its benchmark-topping Qwen-Audio model. (<a href="https://pandaily.com/alibaba-cosyvoice-studio-full-stack-voice-ai-platform-aug2026">Pandaily</a>)</p></li><li><p><strong>Soniox ships TTS v2</strong>, one model for narration, creative work, and voice agents with cloning and expressive control across 60+ languages. (<a href="https://techgenyz.com/soniox-tts-v2-60-languages-realistic-voice-cloning/">Techgenyz</a>)</p></li><li><p><strong>Omilia raises $67M Series B</strong> to expand its voice-first agentic CX platform, having grown live ARR more than 10x to over $60M. (<a href="https://slator.com/voice-ai-omilia-series-b/">Slator</a>)</p></li><li><p><strong>Apple is in talks to pay publishers</strong> up to nine figures to feed current news into its revamped AI-powered Siri via a pay-per-query model. (<a href="https://www.wsj.com/business/media/apple-in-talks-to-pay-publishers-to-improve-ai-powered-siri-0641f64b">WSJ</a>)</p></li><li><p><strong>Regal ties its voice AI agents into Five9</strong>, letting call events trigger AI follow-up calls and texts through Five9&#8217;s AI Agent Connect program. (<a href="https://siliconangle.com/2026/08/13/regal-ties-voice-ai-agents-five9s-contact-center-platform/">SiliconAngle</a>)</p></li><li><p><strong>Sarvam partners with HP India</strong> to bring its Kivi voice assistant to laptops, letting users dictate and act across apps in Indian languages. (<a href="https://www.sarvam.ai/partnerships/hp">Sarvam</a>)</p></li><li><p><strong>Amazon&#8217;s retail chief says conversational AI will reshape shopping</strong>, comparing the shift to the move from printed catalogues to search bars. (<a href="https://itbrief.co.uk/story/amazon-chief-says-conversational-ai-will-reshape-shopping">IT Brief</a>)</p></li><li><p><strong>Assort Health unveils Referrals</strong>, an AI agent that processes and triages every patient referral via voice, text, and email, automating ~90% end-to-end. (<a href="https://www.prnewswire.com/news-releases/assort-health-unveils-referrals-the-first-ai-agent-that-handles-processing-triage-and-continuous-patient-conversations-for-every-referral-302848934.html">PR Newswire</a>)</p></li><li><p><strong>India&#8217;s IISc releases SraVaani</strong>, the first multilingual Indian speech model trained on 65 languages and dialects, open-sourced under MIT. (<a href="https://www.thehindu.com/news/national/karnataka/first-multilingual-indian-speech-recognition-model-trained-on-65-languages-and-dialects-released/article71341129.ece">The Hindu</a>)</p></li><li><p><strong>Zoom combines communications, AI, and workflow</strong> into one agentic CX platform, turning calls and meetings into completed business outcomes. (<a href="https://itcblogs.currentanalysis.com/2026/08/13/zoom-is-delivering-better-cx-by-combining-communications-ai-and-workflow-into-one-platform/">Current Analysis</a>)</p></li><li><p><strong>Bose bets on licensing and tiny AI</strong>, with CEO Lila Snyder detailing a pivot to software and wearables on The Verge&#8217;s Decoder podcast. (<a href="https://www.theverge.com/podcast/975732/bose-ceo-lila-snyder-ai-wearables-licensing-headphones-audio">The Verge</a>)</p></li><li><p><strong>Resemble AI goes all in on deepfake detection</strong>, halting new voice AI sales after its report found 15,700+ people victimized in six months. (<a href="https://www.resemble.ai/resources/why-a-former-voice-ai-company-went-all-in-on-deepfake-detection">Resemble AI</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Sierra shows how its agents navigate IVR systems</strong>, handling multi-level menus, silence, and human detection out of the box. (<a href="https://sierra.ai/blog/navigating-ivr-systems">Sierra</a>)</p></li><li><p><strong>OpenAI just changed how Voice AI is built (</strong><a href="https://bloggeek.me/gpt-live-full-duplex-voice-ai/">BlogGeekMe</a><strong>)</strong></p></li></ul><ul><li><p><strong>A dev builds an on-device real-time translator</strong> on macOS 26 using Apple&#8217;s new speech, translation, and LLM APIs for live subtitles. (<a href="https://dev.to/toffy/apple-quietly-shipped-everything-you-need-to-build-a-real-time-translator-so-i-built-one-9ce">dev.to</a>)</p></li><li><p><strong>AssemblyAI walks through Whisper speaker diarization</strong>, showing how modern models attribute speech from segments as short as 250ms. (<a href="https://www.assemblyai.com/blog/whisper-speaker-diarization">AssemblyAI</a>)</p></li><li><p><strong>FirstBranch turns Congress hearings into a search engine</strong>, built in three days with transcripts, embeddings, clustering, and a map. (<a href="https://www.felixhaba.com/writing/firstbranch-conversational-intelligence-for-congress/">Felix Haba</a>)</p></li><li><p><strong>The SLT 2026 SmartGlasses Challenge</strong> benchmarks egocentric multi-talker speech recognition on 106 hours of four-channel wearable audio. (<a href="https://arxiv.org/abs/2608.12034">arXiv</a>)</p></li><li><p><strong>A dev.to guide covers batch audio transcription in Node.js</strong>, using async webhooks to process long recordings without live streaming. (<a href="https://dev.to/viggoknight2318/batch-audio-transcription-apis-in-nodejs-async-webhooks-for-long-recordings-in-2026-2j5e">dev.to</a>)</p></li><li><p><strong>A benchmark tests local speech-to-text with Foundry Local and C#</strong>, measuring cold-start, real-time factor, memory, and accuracy on-device. (<a href="https://www.c-sharpcorner.com/article/benchmarking-local-speech-to-text-with-foundry-local-and-c-sharp/">C# Corner</a>)</p></li><li><p><strong>YazSes is an offline hold-to-talk voice dictation tool</strong>, a privacy-first cross-platform system that runs transcription fully on-device. (<a href="https://github.com/MSKazemi/yazses">GitHub</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Microsoft and ByteDance join the full-duplex race ]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/microsoft-and-bytedance-join-the</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/microsoft-and-bytedance-join-the</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 10 Aug 2026 14:02:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/80febbbf-2348-4780-b186-1d6f1715da79_2560x1440.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><span>Agora runs a hands-on Voice AI workshop in San Francisco</span> (Aug 13, San Francisco, <a href="https://www.voiceaispace.com/events/voice-ai-workshop---sf-bay-area"><span>Voice AI Space</span></a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Microsoft tests MAI-Realtime</strong>, its first native full-duplex voice model that listens while speaking across 17 languages, cutting Copilot&#8217;s reliance on OpenAI. (<a href="https://startupfortune.com/microsoft-quietly-tests-mai-realtime-a-voice-ai-that-talks-while-it-listens/">Startup Fortune</a>)</p></li><li><p><strong>ByteDance launches SeedRealtime</strong>, a full-duplex model that watches, listens, and speaks at once, now live in the Doubao app&#8217;s phone-call feature. (<a href="https://seed.bytedance.com/en/SeedRealtime">ByteDance</a>)</p></li><li><p><strong>Wispr Flow launches Notetaker</strong>, a Mac meeting assistant that captures system audio bot-free, cleans transcripts, and auto-creates tasks. (<a href="https://9to5mac.com/2026/08/05/wispr-flow-takes-on-ai-meeting-assistants-with-notetaker-its-first-product-beyond-dictation/">9to5Mac</a>)</p></li><li><p><strong>Omilia raises $67M Series B</strong> to expand its voice-first agentic CX platform, having grown live ARR more than 10x to over $60M since Series A. (<a href="https://www.cmswire.com/customer-experience/omilia-raises-67m-series-b-for-voice-ai-push/">CMSWire</a>)</p></li><li><p><strong>Five9 lands a ~$100M contract</strong> with a Fortune 100 financial services firm migrating off on-prem, won in partnership with Google Cloud. (<a href="https://pulse2.com/five9-lands-approximately-100-million-contract-as-voice-ai-agents-drive-contact-center-push/">Pulse 2.0</a>)</p></li><li><p><strong>Salesforce launches Japanese Agentforce Voice</strong> using Kotoba&#8217;s Koto model for low-latency speech that handles honorifics and context-dependent nuance. (<a href="https://jp.ibtimes.com/salesforce-launches-japan-agentforce-voice-kotoba-technology-103514">IBTimes</a>)</p></li><li><p><strong>Sierra lands CarMax</strong> for inbound sales call AI, boosting call resolution and cutting unresolved calls since its May deployment. (<a href="https://www.cmswire.com/customer-experience/sierra-lands-carmax-as-conversational-ais-enterprise-land-grab-widens/">CMSWire</a>)</p></li><li><p><strong>Yellow.ai goes public via a $550M SPAC merger</strong>, aiming to buy legacy BPOs and convert them into AI-native operations. (<a href="https://www.cmswire.com/contact-center/inside-yellow-ais-550m-ipo-target-and-vision-to-transform-legacy-call-centers-with-ai/">CMSWire</a>)</p></li><li><p><strong>8x8 posts record revenue</strong> as AI adoption surged 121% year over year, with over 2,900 AI agents built on its platform. (<a href="https://itbrief.com.au/story/8x8-revenue-hits-record-as-ai-adoption-more-than-doubles">IT Brief</a>)</p></li><li><p><strong>3CLogic wins a major hospital system</strong> to modernize its IT service desk with voice AI embedded natively in ServiceNow. (<a href="https://www.prnewswire.com/news-releases/major-multi-hospital-system-selects-3clogic-to-modernize-it-service-operations-with-voice-ai-for-servicenow-302841610.html">PR Newswire</a>)</p></li><li><p><strong>Granola faces a class-action lawsuit</strong> alleging its bot-free notetaker records meetings without all-party consent and trains AI on them by default. (<a href="https://www.computerworld.com/article/4206255/granola-lawsuit-raises-concerns-over-ai-note-taking-app-privacy.html">Computerworld</a>)</p></li><li><p><strong>ElevenLabs deepens its Tokyo push</strong>, adapting its voice platform for the Japanese market as APAC expansion accelerates. (<a href="https://blog.btrax.com/elevenlabs-tokyo/">btrax</a>)</p></li><li><p><strong>Orvera AI is shortlisted by Everest Group</strong> in its voice AI agents for CXM spotlight, cited for strength in banking, insurance, and healthcare. (<a href="https://martechseries.com/predictive-ai/ai-platforms-machine-learning/orvera-ai-was-included-among-the-shortlisted-providers-featured-in-everest-groups-tech-provider-spotlight-voice-ai-agents-in-customer-experience-management-cxm/">MarTech Series</a>)</p></li><li><p><strong>Telnyx makes the case for owned voice AI infrastructure</strong>, arguing single-backbone control of telephony, compute, and inference beats stitched multi-vendor stacks. (<a href="https://www.under30ceo.com/telnyx-explains-where-voice-ai-infrastructure-platforms-are-earning-their-keep/">Under30CEO</a>)</p></li></ul><div><hr></div><h2>Last Week</h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;67d96f29-bed9-44dd-a3e2-5b367defb1cd&quot;,&quot;caption&quot;:&quot;Voice AI is scaling faster than contact center operating models and it shows.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The 2026 State of Voice in CX&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:32916364,&quot;name&quot;:&quot;Davit Baghdasaryan&quot;,&quot;bio&quot;:&quot;CEO &amp; Co-Founder of Krisp, early pioneer in Voice AI.\n20+ years in engineering. 18 US patent applications, ex Twilion&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/23088dde-6cb0-44df-b220-5f22830cdd4c_1179x960.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-06T14:26:27.687Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abe31bf6-14d9-4d5f-95a3-210a57bce139_1744x800.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/p/the-2026-state-of-voice-in-cx&quot;,&quot;section_name&quot;:&quot;Articles&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:209508711,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:12,&quot;comment_count&quot;:0,&quot;publication_id&quot;:2073467,&quot;publication_name&quot;:&quot;Voice AI Newsletter&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!YLgs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F831a2f7e-d0a7-4e3d-87a8-c42c65d0b71c_1000x1000.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>OpenAI details how it built GPT-Live</strong>, a full-duplex realtime system with a delegation layer that hands hard tasks to GPT-5.5 while keeping you talking. (<a href="https://openai.com/index/continuous-voice-interaction-with-gpt-live/">OpenAI</a>)</p></li></ul><ul><li><p><strong>Google Creative Lab ships an offline Gemma Translator</strong> running Gemma 4 E2B on a Raspberry Pi 5, with full source and 3D-print files released. (<a href="https://aiweekly.co/alerts/google-creative-lab-ships-offline-gemma-translator-for-pi-5">AI Weekly</a>)</p></li><li><p><strong>Cisco IT shares how it modernized voice security with AI</strong>, cutting toll fraud 70% and manual investigation effort 60% with behavioral analytics. (<a href="https://blogs.cisco.com/cisco-on-cisco/how-cisco-it-modernized-voice-security-with-ai">Cisco</a>)</p></li><li><p><strong>AWS publishes a serverless real-time voice AI pattern</strong> for enterprise sales coaching, built on native services for live transcription and analytics. (<a href="https://aws.amazon.com/blogs/industries/serverless-real-time-voice-ai-on-aws-a-pattern-for-enterprise-sales-coaching/">AWS</a>)</p></li><li><p><strong>Microsoft shows live speech-to-text with Foundry Local and C#</strong>, streaming raw PCM audio to a local Nemotron model with no API key or cloud. (<a href="https://devblogs.microsoft.com/dotnet/foundry-local-live-speech-to-text-csharp/">Microsoft</a>)</p></li><li><p><strong>A dev.to tutorial turns a phone call into a formatted email</strong> in 80 lines of Python, cleaning up AI voice memos with the Telnyx API. (<a href="https://dev.to/harpreetseehra/from-phone-call-to-formatted-email-in-80-lines-of-python-ai-voice-memo-cleanup-with-telnyx-32ga">dev.to</a>)</p></li><li><p><strong>A dev.to explainer breaks down modern voice cloning</strong>, walking through the speaker encoder, synthesis model, and vocoder pipeline. (<a href="https://dev.to/peter_zou_7b770f8ba45fd14/understanding-ai-voice-cloning-how-modern-voice-ai-actually-works-16db">dev.to</a>)</p></li><li><p><strong>A new paper introduces QazAVSR</strong>, a Kazakh audio-visual speech recognition model built on HuBERT and ViT with a 57-hour dataset. (<a href="https://www.mdpi.com/2078-2489/17/8/756">MDPI</a>)</p></li><li><p><strong>A study adapts foundation models for Turkic speech-to-speech translation</strong>, comparing fine-tuned cascade and direct end-to-end approaches. (<a href="https://www.mdpi.com/2504-2289/10/8/259">MDPI</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The 2026 State of Voice in CX]]></title><description><![CDATA[815 CX leaders, 12 countries. Here's what they told us.]]></description><link>https://voice-ai-newsletter.krisp.ai/p/the-2026-state-of-voice-in-cx</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/the-2026-state-of-voice-in-cx</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Thu, 06 Aug 2026 14:26:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/abe31bf6-14d9-4d5f-95a3-210a57bce139_1744x800.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4FSV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4FSV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4FSV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg" width="1200" height="503" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:503,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:121605,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4FSV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4FSV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85251e86-416f-4796-a859-567a29de3907_1200x503.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Voice AI is scaling faster than contact center operating models and it shows.</h3><p>New research by <a href="https://krisp.ai/">Krisp</a>, in partnership with <a href="https://ryanadvisory.com/">Ryan Strategic Advisory</a>, shows voice AI adoption is accelerating, while staffing models, technology stacks, and budgets stay rooted in the past.</p><p>Here are the five things that stood out most.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe to receive new posts and insights.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>1. The biggest barrier to global voice support is people, not technology</h2><p>Contact centers are struggling to find enough agents with the right language skills in the right markets.</p><p>40% named hiring and staffing as a top barrier to scaling voice support globally and another 37% cited the cost of supporting multiple languages.</p><p>Technology integration ranked far lower at 14%.</p><p>Enterprises often frame voice AI as a technology project. The stronger business case is in solving workforce constraints and enabling global expansion.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BODR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BODR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 424w, https://substackcdn.com/image/fetch/$s_!BODR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 848w, https://substackcdn.com/image/fetch/$s_!BODR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 1272w, https://substackcdn.com/image/fetch/$s_!BODR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BODR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png" width="1456" height="737" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:737,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127936,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BODR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 424w, https://substackcdn.com/image/fetch/$s_!BODR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 848w, https://substackcdn.com/image/fetch/$s_!BODR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 1272w, https://substackcdn.com/image/fetch/$s_!BODR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8439b5e9-e006-4cb1-8b16-3ebeeed29c20_1917x970.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The global CX labor model may reach its limit before the technology does.</p><h2>2. Enterprises are using AI to make existing teams more effective</h2><p>Agent Assist and Speech Analytics received the highest impact scores at 5.0 out of 6, followed by real-time coaching at 4.9.</p><p>Multilingual AI and Voice Translation ranked close behind at 4.8 and 4.7.</p><p>The near-term priority is clear: help current agents work faster and handle more complex conversations. Language AI goes further by expanding the customers and markets those agents can support.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uqUC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uqUC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 424w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 848w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 1272w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uqUC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png" width="1456" height="634" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:634,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:128378,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uqUC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 424w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 848w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 1272w, https://substackcdn.com/image/fetch/$s_!uqUC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff80548e7-91d0-463a-aee6-a77b5ef94d47_1915x834.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Agent Assist may win the first budget, but language AI could create the larger economic shift.</p><h2>3. Voice AI adoption is approaching a tipping point</h2><p>Planned adoption matches or exceeds current usage across every AI category measured.</p><p>AI Voice Translation shows the largest jump: 28% currently use it, while 60% plan to adopt it within the next 6&#8211;12 months.</p><p>Planned adoption also reaches 44% for AI Noise Cancellation, 42% for Agent Assist, and 39% for Accent Conversion.</p><p>The market has moved from exploring voice AI to deciding what to deploy first.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!R4e-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!R4e-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 424w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 848w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 1272w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!R4e-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png" width="1456" height="887" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:887,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:165952,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!R4e-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 424w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 848w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 1272w, https://substackcdn.com/image/fetch/$s_!R4e-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13e68aee-955d-41bb-8fc7-aecef1810ed2_1897x1156.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Companies still treating voice AI as an innovation project are already behind the buying cycle.</p><h2>4. AI adoption is rising faster than AI budgets</h2><p>56% of enterprises allocate more than half of their voice CX budget to onshore agents.</p><p>Meanwhile, 65% spend less than 10% on AI agents, and 55% spend less than 10% on AI translation.</p><p>The biggest share of spending stays tied to the model facing the most pressure. AI can change the economics by helping the same workforce support more customers, markets, and languages.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wDw9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wDw9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 424w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 848w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 1272w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wDw9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png" width="1456" height="916" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:916,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:109699,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/209508711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wDw9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 424w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 848w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 1272w, https://substackcdn.com/image/fetch/$s_!wDw9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d71b836-1ad7-4ee8-b063-c7218f36b146_1902x1197.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The real AI budget may already exist but it&#8217;s sitting inside labor spend.</p><h2>5. The market changed dramatically since 2025</h2><p>Reliance on human translation fell from 65% in 2025 to 35% in 2026.</p><p>Over the same period:</p><ul><li><p>Voice Translation intent rose from 36% to 60%</p></li><li><p>Accent Conversion adoption and intent grew from 24% to 61%</p></li><li><p>AI Noise Cancellation reached 69%</p></li><li><p>Speech Analytics reached 80%</p></li></ul><p>Voice AI is moving into the core contact center stack faster than most operating models can adapt. The bigger risk is building around assumptions that are already outdated.</p><h2>Voice AI is moving faster than most operating models can adapt</h2><p>The market has entered a new phase: adoption is accelerating, budgets are starting to shift, and the old assumptions around staffing, language coverage, and delivery are breaking down.</p><p>The advantage will go to companies that act on those signals early, before fragmented pilots become fragmented infrastructure.</p><p><strong>Read the full 2026 State of Voice in CX report for the complete findings and industry implications.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://resources.krisp.ai/hubfs/WhitePaper/2026%20State%20of%20Voice-VAI.pdf&quot;,&quot;text&quot;:&quot;Read the full report&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://resources.krisp.ai/hubfs/WhitePaper/2026%20State%20of%20Voice-VAI.pdf"><span>Read the full report</span></a></p><p></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[xAI launches Grok Voice Think v2, Fish Audio raises $52M]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/xai-launches-grok-voice-think-v2</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/xai-launches-grok-voice-think-v2</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 03 Aug 2026 14:01:43 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/16517ae1-e167-464a-9141-60fbfcc3463b_1496x1002.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p>Agora Convo AI World mixer in Nashville (Aug 4&#8211;5, Nashville, <a href="https://luma.com/5g2d76uj">LUMA</a>)</p></li><li><p>Vapi and Deepgram go live on model tradeoffs, comparing Nova-3, Flux, and Aura-2 across transcription, turn detection, and latency. (Aug 6, Hybrid, <a href="https://luma.com/vapi-44bz">LUMA</a>)</p></li><li><p>Agora and Seeed Studio run a hands-on hardware workshop at Silicon Valley&#8217;s Robotics Fair (Aug 8, San Mateo, CA, <a href="https://www.voiceaispace.com/events/agora--seeed-studio-physical-ai-workshop-build-customize-and-deploy-an-ai-agent-to-seeed-studio-respeaker-flex-using-agora-mybot">Voice AI Space</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>xAI launches Grok Voice Think Fast 2.0</strong>, which reasons while speaking for smarter answers with no added latency, scoring 82.9 on the speech quality index. (<a href="https://x.ai/news/grok-voice-think-fast-2">xAI</a>)</p></li><li><p><strong>OpenAI ships GPT-Transcribe and GPT-Live-Transcribe</strong>, cutting transcription errors by 52% vs. Whisper with real-time and batch modes across 57 languages. (<a href="https://slator.com/openai-two-new-models-transcription/">Slator</a>)</p></li><li><p><strong>Google gives Gemini for Mac voice control</strong> with Fn-key activation, intelligent dictation that strips filler words, and an opt-in screen-aware reasoning mode. (<a href="https://9to5google.com/2026/07/29/gemini-mac-voice-control/">9to5Google</a>)</p></li><li><p><strong>Fish Audio raises $52M in seed funding</strong> after hitting $21M ARR and 8M+ users in its first year, with voice cloning and streaming across 80+ languages. (<a href="https://techcrunch.com/2026/07/28/fish-audio-raises-50m-seed-to-build-ai-voice-models-for-creators-and-enterprises/">TechCrunch</a>)</p></li><li><p><strong>Encore AI raises $30M Series A</strong> to build voice agents that learn winning behaviors from customer call recordings using &#8220;interaction mining.&#8221; (<a href="https://techcrunch.com/2026/07/29/encore-ai-raises-30m-to-build-ai-agents-that-learn-from-customer-calls/">TechCrunch</a>)</p></li><li><p><strong>Smallest AI raises $13M Series A</strong> to make voice agents indistinguishable from humans, bringing total funding to $21M with sub-second latency models. (<a href="https://techcrunch.com/2026/07/31/smallest-ai-raises-13m-to-build-ultra-fast-voice-ai-that-sounds-genuinely-human/">TechCrunch</a>)</p></li><li><p><strong>PolyAI releases Dialog RSN-1</strong>, an audio-native model that reasons directly over raw call audio instead of transcripts, delivering sub-300ms responses. (<a href="https://siliconangle.com/2026/07/30/polyai-launches-new-real-time-voice-conversation-model-make-ai-driven-calls-human/">SiliconAngle</a>)</p></li><li><p><strong>OpenAI expands GPT-Live to Edu, Business, and Enterprise</strong> plans globally, bringing full-duplex voice with background reasoning to organizations. (<a href="https://thewincentral.com/gpt-live-chatgpt-voice-edu-business-enterprise/">TheWinCentral</a>)</p></li><li><p><strong>Qwen Audio 3.0 Realtime Plus tops OpenAI</strong> on the Artificial Analysis Speech-to-Speech Index at 84.1% vs. GPT-Realtime-2.1&#8217;s 79.1%, a first for Alibaba. (<a href="https://betanews.com/article/qwen-audio-3-vs-openai-speech-benchmark/">BetaNews</a>)</p></li><li><p><strong>Sarvam AI announces a trillion-parameter model</strong> built in India, with pricing 5.5x cheaper than GPT-5.4 Mini for coding and research workloads. (<a href="https://www.freepressjournal.in/tech/sarvam-bets-big-on-indias-ai-future-with-trillion-parameter-model-will-be-priced-five-times-cheaper-than-global-rivals">Free Press Journal</a>)</p></li><li><p><strong>Krafton releases A.X K2 Raon-Speech</strong>, a 21B-parameter audio model that processes speech directly to preserve emotion and tone, ranked first among Korean models. (<a href="https://www.theinvestor.co.kr/article/10824774">The Investor</a>)</p></li><li><p><strong>Boson AI unveils Higgs RealTime</strong> for full speech-to-speech processing that skips text conversion to preserve vocal nuance, led by ex-AWS scientist Alex Smola. (<a href="https://cryptobriefing.com/boson-ai-higgs-realtime-voice-model/">CryptoBriefing</a>)</p></li><li><p><strong>8x8 extends AI across its full CX platform</strong> beyond the contact center, adding conversation intelligence, agent development, and workforce management for every team. (<a href="https://www.cmswire.com/contact-center/8x8-extends-ai-across-its-cx-platform-to-all-teams/">CMSWire</a>)</p></li><li><p><strong>NICE and RingCentral expand into a bi-directional partnership</strong>, combining UCaaS, CCaaS, and AI in a single offering with mutual resale. (<a href="https://telecomreseller.com/2026/07/29/ringcentral-nice/">Telecom Reseller</a>)</p></li><li><p><strong>Parlance adds instant self-service config and AI call summaries</strong> for healthcare contact centers, now deployed across 400+ health systems. (<a href="https://www.prnewswire.com/news-releases/parlance-gives-healthcare-contact-centers-instant-configuration-control-and-ai-briefed-agent-transfers-302839592.html">PR Newswire</a>)</p></li><li><p><strong>Tata Communications launches voice AI for India&#8217;s 63M SMBs</strong> with speech-to-speech engagement, sub-500ms latency, and multilingual support via Tata Tele. (<a href="https://inc42.com/buzz/tata-communications-expands-voice-ai-push-eyes-smb-adoption/">Inc42</a>)</p></li><li><p><strong>Tesla rolls out Grok voice assistant in India</strong> for the Model Y via OTA update, supporting Hindi and five other Indian languages for navigation and queries. (<a href="https://www.indianweb2.com/2026/08/tesla-model-y-india-launch-adds-grok-ai.html">IndianWeb2</a>)</p></li><li><p><strong>Japan&#8217;s Justice Ministry backs voice rights protection</strong> from AI, concluding that publicity rights can address unauthorized voice cloning without new legislation. (<a href="https://english.kyodonews.net/articles/-/80955">Kyodo News</a>)</p></li><li><p><strong>Tysa, Conectys&#8217; AI voice agent, marks one year</strong> of handling global customer conversations across multiple languages and use cases. (<a href="https://www.einpresswire.com/article/923836012/tysa-turns-one-conectys-ai-voice-agent-marks-a-year-of-global-conversation">EIN Presswire</a>)</p></li></ul><div><hr></div><h2>Last Week Podcast</h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;2841f5d1-52a1-42d9-aea6-4094ef1add61&quot;,&quot;caption&quot;:&quot;In the Future of Voice AI series of interviews, I ask three questions to my guests: - What problems do you currently see in Enterprise Voice AI? - How does your company solve these problems? - What solutions do you envision in the next 5 years?&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;AI, Fraud, and the Human Contact Center | Geoff Burbridge (Principal + Founder @ Human Edge Advisory)&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:32916364,&quot;name&quot;:&quot;Davit Baghdasaryan&quot;,&quot;bio&quot;:&quot;CEO &amp; Co-Founder of Krisp, early pioneer in Voice AI.\n20+ years in engineering. 18 US patent applications, ex Twilion&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/23088dde-6cb0-44df-b220-5f22830cdd4c_1179x960.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-30T14:25:31.976Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3c2468c4-ac27-42c4-b7ee-68fb7eb87ca3_1165x776.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/p/ai-fraud-and-the-human-contact-center&quot;,&quot;section_name&quot;:&quot;Podcast&quot;,&quot;video_upload_id&quot;:&quot;beed0608-56ff-4a2c-b3d4-76fecf415753&quot;,&quot;id&quot;:207828532,&quot;type&quot;:&quot;podcast&quot;,&quot;reaction_count&quot;:13,&quot;comment_count&quot;:0,&quot;publication_id&quot;:2073467,&quot;publication_name&quot;:&quot;Voice AI Newsletter&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!YLgs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F831a2f7e-d0a7-4e3d-87a8-c42c65d0b71c_1000x1000.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>HackerNoon breaks down voice-to-voice AI architectures</strong>, from codec tokenization and RVQ prediction to full-duplex designs like Moshi and GPT-Live. (<a href="https://hackernoon.com/how-modern-voice-to-voice-ai-models-work">HackerNoon</a>)</p></li></ul><ul><li><p><strong>A deep dive into ARK-ASR-3B&#8217;s architecture</strong> shows how combining a Whisper encoder with a Qwen decoder achieves 5.04% WER on the Open ASR Leaderboard. (<a href="https://hackernoon.com/inside-ark-asr-3bs-whisper-and-qwen-architecture">HackerNoon</a>)</p></li><li><p><strong>Building SeMamba for speech enhancement</strong> walks through using Mamba state-space models to recover clean speech from noisy audio in a single forward pass. (<a href="https://levelup.gitconnected.com/from-noise-to-clarity-building-semamba-for-speech-enhancement-2e534e62dd7e">Level Up Coding</a>)</p></li><li><p><strong>Deepgram integrates with AWS SageMaker AI</strong> via IAM temporary delegation, letting enterprises self-host Nova and Aura-2 with zero standing cross-account access. (<a href="https://aws.amazon.com/blogs/machine-learning/deepgram-enhances-amazon-sagemaker-ai-support-with-aws-iam-temporary-delegation/">AWS</a>)</p></li><li><p><strong>A Frontiers in AI paper compares Wav2Vec and Whisper</strong> for Telugu speech recognition, benchmarking pretrained models on low-resource Indian languages. (<a href="https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2026.1878977/full">Frontiers</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI, Fraud, and the Human Contact Center | Geoff Burbridge (Principal + Founder @ Human Edge Advisory)]]></title><description><![CDATA[Watch now | In the Future of Voice AI series of interviews, I ask three questions to my guests: - What problems do you currently see in Enterprise Voice AI? - How does your company solve these problems? - What solutions do you envision in the next 5 years?]]></description><link>https://voice-ai-newsletter.krisp.ai/p/ai-fraud-and-the-human-contact-center</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/ai-fraud-and-the-human-contact-center</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Thu, 30 Jul 2026 14:25:31 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207828532/b65adac707fa211808bf12cef85fe568.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<pre><code><code>In the Future of Voice AI series of interviews, I ask three questions to my guests:

- What problems do you currently see in Enterprise Voice AI?
- How does your company solve these problems?
- What solutions do you envision in the next 5 years?</code></code></pre><p>This episode&#8217;s guest is <a href="https://www.linkedin.com/in/geoffreyburbridge/">Geoff Burbridge</a>, Principal + Founder at <a href="https://www.humanedgeadvisory.co/">Human Edge Advisory</a>.</p><p>Geoff Burbridge is Principal + Founder of Human Edge Advisory, where he helps executive teams, boards, and founders transform organizations at the intersection of people, process, and technology. With more than 30 years of executive leadership experience at Bank of America, USAA, Truist, and Capital One, Geoff advises organizations on AI readiness, operating model design, and enterprise transformation, helping leaders balance innovation with the security, governance, and trust required by today's enterprises. He is a keynote speaker, executive moderator, and board advisor who believes that as technology continues to evolve, the greatest competitive advantage will always be the Human Edge.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gd_K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gd_K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 424w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 848w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gd_K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png" width="1200" height="1200" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1200,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:303874,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/207828532?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gd_K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 424w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 848w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!gd_K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a2956bc-2d9a-47cf-9699-560783a2a968_1200x1200.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.youtube.com/@futureofvoiceai&quot;,&quot;text&quot;:&quot;Listen on YouTube&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.youtube.com/@futureofvoiceai"><span>Listen on YouTube</span></a></p><h3><strong>Recap Video</strong></h3><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;3a642a58-3752-461e-ab23-95ac30014d0c&quot;,&quot;duration&quot;:null}"></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive weekly updates.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><strong>Takeaways</strong></h3><ul><li><p>AI&#8217;s biggest contact center impact may be stopping fraud, not replacing agents.</p></li><li><p>The first AI strategy will often be wrong because the market, not the roadmap, reveals the real use case.</p></li><li><p>BPOs may emerge as AI winners, not casualties, because they can apply automation across massive service operations.</p></li><li><p>The industry has already retreated from &#8220;replace the agent&#8221; to &#8220;make the agent better.&#8221;</p></li><li><p>AI adoption will stall when companies automate tasks but ignore how the workforce itself must change.</p></li><li><p>The real AI risk is not failure, but destroying customer confidence before the technology earns their trust.</p></li><li><p>Fraudsters face don&#8217;t the same rules as enterprises, giving attackers a speed advantage that regulation alone cannot solve.</p></li><li><p>Voice fraud is becoming an AI-versus-AI fight that humans cannot manage at machine speed.</p></li><li><p>Most contact centers already have the signals to stop fraud but cannot connect them fast enough to act.</p></li><li><p>Authentication failure is weak protection when attackers can rapidly change devices, numbers, identities, and voices.</p></li><li><p>The customer call companies want to eliminate may be the exact interaction that stops a major fraud loss.</p></li><li><p>Human trust makes voice service powerful, but it also makes voice one of the easiest channels to exploit.</p></li><li><p>Deepfake scams succeed by creating urgency before the victim has time to question what sounds real.</p></li><li><p>The best fraud systems will not just detect known attacks; they will identify new attack patterns as they form.</p></li><li><p>Agentic AI is more valuable for absorbing demand spikes than replacing every customer conversation.</p></li><li><p>The best automation targets are processes that already reach the right outcome almost every time.</p></li><li><p>Repeatedly failed journeys need more human judgment, not another layer of automation.</p></li><li><p>AI coaching can scale practice without the pressure and resistance that make traditional role-play fail.</p></li><li><p>The intersectionality of AI is where its value compounds, connecting fraud signals, customer context, automation, and human judgment in real time.</p></li><li><p>AI&#8217;s real power comes from connecting fraud signals, customer context, agent judgment, and automation in the same moment.</p></li><li><p>Contact centers overcomplicate service when customers fundamentally only care about four things: answering the phone, being nice, resolution, and ease.</p></li><li><p>In Geoff&#8217;s experience, authentic service earns trust because customers believe people, not scripts.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[ChatGPT Voice Hits Desktop, Claude Voice Goes Full Opus]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/chatgpt-voice-hits-desktop-claude</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/chatgpt-voice-hits-desktop-claude</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 27 Jul 2026 14:02:05 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/de5511dc-b81f-4af4-80d6-41f829aa1efd_1574x876.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>Vapi and Inngest build a voice agent live,</strong> taking it from demo to production with retries, long-running execution, and evals along the way. (Jul 29, Hybrid, <a href="https://www.voiceaispace.com/events/from-voice-demo-to-production-agent-building-reliable-ai-workflows">Voice AI Space</a>)</p></li><li><p><strong>Temporal hosts a durable multimodal AI meetup with HeyGen, Vapi, and Modal,</strong> including a Vapi talk on scaling voice agents from one call to ten thousand. (Jul 29, SF, <a href="https://luma.com/durable-ai-july">LUMA</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>OpenAI brings ChatGPT Voice to desktop</strong> with GPT-Live-powered agents that can control apps, dictate in any window, and work with Codex on Mac and Windows. (<a href="https://techcrunch.com/2026/07/24/openais-new-voice-mode-makes-it-to-the-chatgpt-desktop-app/">TechCrunch</a>)</p></li><li><p><strong>Anthropic upgrades Claude voice mode</strong> to Opus and Sonnet models with cross-app automation for Gmail, Slack, and Notion in 10 languages. (<a href="https://www.webpronews.com/anthropic-arms-claude-voice-mode-with-opus-and-sonnet-brains/">WebProNews</a>)</p></li><li><p><strong>Alibaba launches Qwen Audio 3.0 TTS</strong> in Flash and Plus tiers across 16 languages, reaching #1 on the Artificial Analysis TTS arena at 1,236 Elo. (<a href="https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/">MarkTechPost</a>)</p></li><li><p><strong>ByteDance releases Seed Audio 1.0</strong>, a unified model for voice, music, and sound effects with zero-shot voice cloning from a single reference clip. (<a href="https://seed.bytedance.com/en/seedaudio1_0">ByteDance</a>)</p></li><li><p><strong>Deepgram puts Nova-3 on Snapdragon</strong> with on-device voice AI optimized for Qualcomm&#8217;s Hexagon NPU, eliminating the need for cloud connectivity. (<a href="https://finance.yahoo.com/technology/ai/articles/deepgram-delivers-real-time-voice-123000272.html">Yahoo Finance</a>)</p></li><li><p><strong>Telli raises $15M seed</strong> from Redalpine, Y Combinator, and Cherry Ventures to build AI agents that replace traditional call center operations. (<a href="https://www.pymnts.com/news/investment-tracker/2026/startup-telli-raises-15-million-to-replace-call-centers-with-ai-agents/">PYMNTS</a>)</p></li><li><p><strong>Valence AI raises $5M</strong> backed by SRI International to integrate real-time emotion detection into voice AI with US patents on emotional intelligence. (<a href="https://www.sri.com/press/story/sri-backed-valence-ai-raises-5m-to-integrate-emotional-intelligence-into-the-trust-stack/">SRI</a>)</p></li><li><p><strong>Cast Insights raises $4.5M pre-seed</strong> for real-time speech intelligence across TV, radio, and podcasts, having processed over 2.3M hours of audio. (<a href="https://siliconangle.com/2026/07/23/real-time-speech-intelligence-startup-cast-insights-raises-4-5m-track-worlds-tv-radio-podcasts/">SiliconAngle</a>)</p></li><li><p><strong>Zoom launches real-time voice translation</strong> that converts spoken language live during meetings across five languages. (<a href="https://www.zoom.com/en/blog/voice-translation-zoom/">Zoom</a>)</p></li><li><p><strong>Zoom Scribe adds speech accessibility</strong> features for real-time captioning and transcription to improve meeting inclusivity. (<a href="https://www.zoom.com/en/blog/zoom-scribe-speech-accessibility/">Zoom</a>)</p></li><li><p><strong>Otter.ai introduces Live Assist</strong>, a real-time coaching agent that listens to calls and provides guidance customizable with team playbooks. (<a href="https://www.businesswire.com/news/home/20260721446216/en/Otter.ai-Introduces-Live-Assist-The-First-Live-Coaching-Agent-for-Every-Call">BusinessWire</a>)</p></li><li><p><strong>Level AI launches Latitude</strong>, a suite of 7 purpose-built CX models that the company claims are 49x cheaper than frontier LLMs for contact centers. (<a href="http://www.smartcustomerservice.com/Articles/News-Briefs/Level-AI-Launches-Latitude-a-Suite-of-7-AI-Models-175782.aspx">Smart Customer Service</a>)</p></li><li><p><strong>Apple rolls out Genius Bar Live Notes</strong>, using AI to transcribe and summarize in-store customer appointments in real time. (<a href="https://hothardware.com/news/apple-genius-bar-ai-recording-tool">HotHardware</a>)</p></li><li><p><strong>Ukraine adds voice AI to its Diia government app</strong>, letting citizens access 170+ public services through ElevenLabs-powered spoken conversations. (<a href="https://www.smartcitiesworld.net/news/ukraine-adds-voice-ai-to-national-government-assistant">Smart Cities World</a>)</p></li><li><p><strong>AudioCodes targets 40-50% voice AI growth</strong> in 2026 with its conversational AI segment on track to reach $25M, aiming for $80M by 2028. (<a href="https://www.channelinsider.com/ai/building-channel-revenue/audiocodes-voice-ai-opportunity/">Channel Insider</a>)</p></li><li><p><strong>Parlance nears 2 billion healthcare calls</strong> processed by its voice AI platform, marking a major deployment milestone in patient-facing automation. (<a href="https://www.prnewswire.com/news-releases/2-billion-calls-parlance-voice-ai-nears-healthcare-milestone-302832861.html">PR Newswire</a>)</p></li><li><p><strong>WhisperAI surpasses 330,000 users</strong> and launches a transcription API with real-time STT, crossing seven-figure ARR in under a year. (<a href="https://techbullion.com/whisperai-surpasses-330000-professionals-worldwide-and-launches-advanced-whisperai-transcription-api/">TechBullion</a>)</p></li><li><p><strong>XMOS unveils VocalFusion XVF3620</strong>, an AI voice processor combining on-chip noise reduction, beamforming, and echo cancellation in a single device. (<a href="https://audioxpress.com/news/xmos-unveils-vocalfusion-xvf3620-with-advanced-ai-voice-processing">audioXpress</a>)</p></li><li><p><strong>Synaptics launches Astra SR80</strong>, an always-on edge AI audio processor for voice capture, biometric auth, and agentic AI devices. (<a href="https://audioxpress.com/article/human-centric-devices-and-the-rise-of-always-on-edge-ai-audio">audioXpress</a>)</p></li><li><p><strong>Smallest AI&#8217;s TTS ranks #1 for Hindi</strong> in blind listening tests, with Lightning v3.1 preferred 76% of the time over OpenAI&#8217;s GPT-4o-mini-TTS. (<a href="https://www.northjersey.com/press-release/story/217025/smallest-ais-tts-ranked-the-top-real-time-voice-for-hindi-customer-support/">North Jersey</a>)</p></li><li><p><strong>MouthPad launches a $1,400 tongue-controlled trackpad</strong> alongside Vox, a $200 whisper-to-type wearable microphone for hands-free voice input. (<a href="https://www.morningstar.com/news/pr-newswire/20260722sf10265/mouthpad-the-first-ever-tongue-trackpad-launches-to-the-public-alongside-new-wearable-vox">Morningstar</a>)</p></li><li><p><strong>Sweekar AI pocket pet uses staged voice development</strong> from babbling to fluent speech as a core growth mechanic, debuted at CES 2026 by Takway AI. (<a href="https://audioxpress.com/news/sweekar-ai-pocket-pet-turns-voice-interaction-into-a-growth-mechanism">audioXpress</a>)</p></li><li><p><strong>SoliderSound launches Go-Denoise</strong>, a desktop app using neural spectral processing to remove noise from speech and vocals in a single pass. (<a href="https://mixing.co.kr/en/42216">Monthly Mixing</a>)</p></li><li><p><strong>Forbes asks if AI is taking over audio</strong> as AI now hosts radio shows, narrates audiobooks, and generates podcasts, sparking debate over quality vs. human narration. (<a href="https://www.forbes.com/sites/frankracioppi/2026/07/22/is-ai-taking-over-audio---radio-podcasting-audiobooks/">Forbes</a>)</p></li><li><p><strong>Parloa argues CX leaders should stop measuring voice AI by deflection alone</strong>, pushing for resolution rate, customer effort, and escalation quality as better metrics. (<a href="https://www.cxtoday.com/contact-center/why-cx-leaders-should-stop-measuring-voice-ai-by-deflection-alone-parloa-cs-0228/">CX Today</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Cue AI runs Gemma 4 locally for voice dictation</strong>, cutting latency 44% and dropping marginal inference cost to zero on desktop voice agents. (<a href="https://deepmind.google/models/gemma/gemmaverse/cue-ai/">Google DeepMind</a>)</p></li></ul><ul><li><p><strong>Oracle builds a 24/7 healthcare voice agent</strong> using NVIDIA PersonaPlex and LiveKit on OCI, demonstrating full-duplex speech-to-speech for medical assistants. (<a href="https://blogs.oracle.com/cloud-infrastructure/healthcare-agent-nvidia-personaplex-livekit">Oracle</a>)</p></li><li><p><strong>MarkTechPost compares the best open ASR models of 2026</strong>, finding top models within one WER point while license and streaming become the real differentiators. (<a href="https://www.marktechpost.com/2026/07/23/best-open-speech-recognition-asr-models-in-2026-wer-languages-latency-and-license-compared/">MarkTechPost</a>)</p></li><li><p><strong>The voice agent latency playbook</strong> argues that input accuracy is a latency feature since wrong transcripts turn one conversation turn into three. (<a href="https://hackernoon.com/the-voice-agent-latency-playbook-stt-turn-detection-and-the-tradeoffs-nobody-talks-about">HackerNoon</a>)</p></li><li><p><strong>AssemblyAI details the future of real-time STT</strong> with its Universal-3.5 Pro model, semantic turn detection, and a single-WebSocket Voice Agent API. (<a href="https://www.assemblyai.com/blog/realtime-future-of-speech-to-text">AssemblyAI</a>)</p></li><li><p><strong>A dev.to guide on voice agent turn-taking</strong> explains how to keep AI calls under 600ms end-to-end with VAD, barge-in handling, and streaming pipelines. (<a href="https://dev.to/jackm-singularity/voice-agent-turn-taking-stop-live-ai-calls-from-talking-over-users-590b">dev.to</a>)</p></li><li><p><strong>NVIDIA Nemotron 3.5 ASR runs 14x faster than Whisper</strong> on a CPU-only laptop via parakeet.cpp, transcribing a 1m46s clip in 32 seconds vs. Whisper&#8217;s 7m34s. (<a href="https://medium.com/@hellorahulk/running-nvidia-nemotron-3-5-asr-locally-with-parakeet-cpp-and-how-it-beat-whisper-on-my-laptop-42105b504307">Medium</a>)</p></li><li><p><strong>OpenWhispr is an open-source voice dictation app</strong> supporting local Whisper and NVIDIA Parakeet models with zero telemetry, now at 2,100+ GitHub stars. (<a href="https://github.com/OpenWhispr/openwhispr">GitHub</a>)</p></li><li><p><strong>HackerNoon explains why TTS evaluation needs human ears</strong>, noting that WER and CER miss intonation and naturalness, requiring manual listening across checkpoints. (<a href="https://hackernoon.com/in-text-to-speech-your-ears-always-matter-more-than-your-metrics">HackerNoon</a>)</p></li><li><p><strong>A developer narrates their blog with Kokoro</strong>, an 82M-parameter local TTS model that generates two hours of audio in 15 minutes on an M1 MacBook. (<a href="https://bart.degoe.de/narrating-your-blog-with-kokoro-a-local-opensource-model/">Bart de Goede</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[What CCW 2026 told us about Voice AI in CX]]></title><description><![CDATA[CCW Vegas had no shortage of Voice AI.]]></description><link>https://voice-ai-newsletter.krisp.ai/p/what-ccw-2026-told-us-about-voice</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/what-ccw-2026-told-us-about-voice</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Thu, 23 Jul 2026 14:30:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bscE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bscE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bscE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bscE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bscE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bscE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bscE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9575439,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/206995171?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bscE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bscE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bscE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bscE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849dff6e-c35c-4532-8ddf-789e58aa77d3_8192x5464.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>CCW Vegas had no shortage of Voice AI. What stood out was how little patience buyers have left for vague promises.</p><p>The conversation has shifted from <em><strong>what can AI do?</strong></em> to <em><strong>what has it actually done?</strong></em></p><p><span>This was the year the industry stopped debating AI and started auditing it. </span>Here is what stood out most.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>1. AI has entered its proof era</h2><p>The AI market was crowded, but not always convincing. Many vendors claimed leadership, while case studies were thin and demos were limited. Buyers weren&#8217;t short on options, they were short on proof.</p><p>The companies that stood out immediately show what their solutions change: lower handle time, faster resolution, lower cost, or less work for the agent.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZEft!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZEft!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZEft!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7762708,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/206995171?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZEft!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ZEft!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a37d29-646f-47ee-93e6-18bbeb9717b4_8192x5464.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Takeaway:</strong> In a crowded market, credibility comes from showing results buyers can see, measure, and trust &#8212; and showing them upfront.</p><h2>2. The best AI removes friction, not people</h2><p>Despite all the talk about Voice AI eating CX, the industry is still growing. CCW felt bigger, busier, and more confident than ever.</p><p>The strongest signal from both BPOs and enterprises was that human-led and AI-powered CX will grow together.</p><p>Damola Adamolekun, CEO of Red Lobster, offered one of the clearest operating principles of the week in his keynote:</p><blockquote><h4>&#8220;Focus on the changes with the highest impact and lowest effort.&#8221;</h4></blockquote><p><strong>Takeaway:</strong> AI&#8217;s near-term value is quick wins and removing the friction that holds companies and people back.</p><h2>3. Voice Translation is changing the economics of language support</h2><p>Multilingual conversations weren&#8217;t about features; they were about replacing a broken model: customer needs support, agent dials an interpreter, everyone waits. Interpreters are expensive, hard to staff, and regulated language requirements keep expanding.</p><p>The strongest proof came from production. On the main stage, Automated Health Systems shared how deploying Voice Translation reduced average multilingual calls from more than <strong>40 minutes to just 9</strong>, expanding language access while dramatically reducing operational costs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A4fr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A4fr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A4fr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8612140,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/206995171?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A4fr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 424w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 848w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!A4fr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F664732b3-92d3-46a9-8b97-1b16b8339f83_8192x5464.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Takeaway:</strong> Translation isn&#8217;t just improving access &#8212; it&#8217;s changing the cost, speed, and scale of multilingual service.</p><h2>4. Voice AI is CX infrastructure</h2><p>The show floor made the shift hard to miss.</p><p>Cisco and Microsoft led with agentic workflows. Jabra focused more on analytics. Poly and Epos weren&#8217;t there.</p><p>At the same time, Voice AI companies had a much larger presence than in prior years.</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;bfe2d5a1-7d80-45bd-b9e8-bcf4aaa924bb&quot;,&quot;duration&quot;:null}"></div><p>The value is moving into what happens inside the conversation: improving audio, which improves everything downstream&#8212; translation, transcription, agent guidance, detecting risk, and powering AI interactions in real time.</p><p><strong>Takeaway:</strong> Voice AI is becoming core infrastructure that the rest of CX depends on.</p><h2>5. Voice security is urgent</h2><p>Deepfake detection drew real interest, including follow-up requests from major enterprises.</p><p>The risk is growing fast:</p><ul><li><p>Deepfake fraud attempts grew <strong>2,137% in three years</strong></p></li><li><p>AI-related cybercrime losses topped <strong>$893 million in 2025</strong></p></li><li><p>Humans detect synthetic voices correctly only about <strong>60% of the time</strong></p></li></ul><p>Most fraud controls still sit before or after the call, but most voice attacks happen during it.</p><p>As Geoff Burbridge, Founder of Human Edge Advisory, put it:</p><blockquote><p>&#8220;AI is being used by bad actors out there, and newsflash, they&#8217;re not restrained. They&#8217;re not worried about regulation.&#8221;</p></blockquote><p><strong>Takeaway:</strong> Voice security is moving from future risk to an urgent buying requirement.</p><p><strong>Sneak peek:</strong> Geoff joins the <em>Voice AI Podcast</em> next week to break down where synthetic voice fraud is hitting hardest, how real-time detection works in production, and what companies need to do now.</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;bbb86917-86e0-48c8-a04f-d10ecfacaf54&quot;,&quot;duration&quot;:null}"></div><h2>A milestone worth sharing</h2><p>Krisp was named <strong>Disruptive Technology of the Year</strong> at the 2026 CCW Excellence Awards.</p><p><span>We&#8217;re honored by the recognition, but more importantly, it reflects where the industry is heading. </span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bgWL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bgWL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bgWL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png" width="260" height="260" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1080,&quot;width&quot;:1080,&quot;resizeWidth&quot;:260,&quot;bytes&quot;:130971,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://voice-ai-newsletter.krisp.ai/i/206995171?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bgWL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!bgWL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F746e9329-11e0-452a-b350-6940dbc1be3b_1080x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The technologies earning attention today aren&#8217;t chasing novelty. They&#8217;re solving real operational problems and delivering measurable results. </span></p><p><span>That was the story across the show floor, and one we&#8217;re proud to be part of.</span></p><h2>What comes next</h2><p>By CCW 2027:</p><ul><li><p><strong>Execution starts separating leaders from followers:</strong> Organizations that deployed AI in 2025-2026 will have results to show:</p><ul><li><p>Those that redesigned around the technology will have compounding returns.</p></li><li><p>Those that layered it on top of old structures might be starting over.</p></li></ul></li><li><p><strong><span>AI autonomy:</span></strong><span> the Organizations moving fastest are starting with agent assist, proving value in production, and expanding from there. Deploying autonomous AI before you&#8217;ve built trust into the system is how you lose both.</span></p></li><li><p><strong><span>Voice security becomes a procurement requirement:</span></strong><span> Operators asking questions now will have policies and vendor requirements in place, while those not paying attention will be responding to incidents.</span></p></li><li><p><strong><span>Multilingual support consolidates:</span></strong><span> Operators want fewer vendors and end-to-end accountability across language access, accent support, and translation.</span></p></li><li><p><strong><span>Agent experience reaches a tipping point:</span></strong><span> The industry can&#8217;t sustain current attrition rates. Companies that reduce friction and lower cognitive load, not just cut headcount, will have a structural talent advantage.</span></p></li><li><p><strong><span>Voice gets the engineering investment it has always deserved:</span></strong><span> The channel, technology, and infrastructure layer are finally getting serious attention. Long overdue.</span></p></li><li><p><strong>AI pressure will widen the BPO gap:</strong> BPOs that modernize their delivery model will pull ahead of those that don&#8217;t.</p></li></ul><p>The market wants proof, fewer points of failure, and technology that improves the conversation in real time. By next year, the winners will be the ones with the clearest outcomes.</p><p>See you there.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Voice AI Newsletter! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Rime Raises $24M, Meta Patents Voice Emotion Tracking]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/rime-raises-24m-meta-patents-voice</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/rime-raises-24m-meta-patents-voice</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 20 Jul 2026 14:02:26 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a40a3425-3b67-4604-acb5-7abd4cfcab4f_850x425.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>Deepgram hosts a fireside chat on the future of voice agents,</strong> evaluating agents beyond WER and whether cascaded STT-LLM-TTS still holds up. (Jul 23, London, <a href="https://www.voiceaispace.com/events/the-future-of-voice-agents-a-fireside-chat">Voice AI Space</a>)</p></li><li><p><strong>AssemblyAI demos Universal-3.5 Pro Realtime,</strong> its new model with context carryover and conversation memory, then opens up for a fireside chat with production voice agent builders. (Jul 23, SF, <a href="https://luma.com/m0thk5ai">Luma</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>Alibaba launches Qwen Audio 3.0</strong> with real-time voice that can proactively use external tools, covering 113 languages for ASR and 36 for TTS. (<a href="https://www.kucoin.com/news/flash/aliyun-launches-qwen-audio-3-0-realtime-voice-ai-can-proactively-use-external-tools">KuCoin</a>)</p></li><li><p><strong>Meta patents an AI wearable</strong> that continuously analyzes voice to track the user&#8217;s emotional state, raising concerns under the EU AI Act&#8217;s emotion-inference ban. (<a href="https://thenextweb.com/news/meta-patent-mood-tracking-voice-emotion-ai">The Next Web</a>)</p></li><li><p><strong>PwC and OpenAI launch agentic customer service</strong> solutions combining PwC&#8217;s CX expertise with OpenAI&#8217;s multimodal voice and digital agent APIs. (<a href="https://www.cxtoday.com/ai-automation-in-cx/pwc-openai-agentic-customer-service/">CX Today</a>)</p></li></ul><ul><li><p><strong>Google Voice adds Gemini AI notes</strong> that auto-summarize calls with key points and action items, plus new standalone plans starting at $10/mo. (<a href="https://www.webpronews.com/google-voice-brings-gemini-ai-notes-to-calls-with-new-standalone-plans-at-lower-cost/">WebProNews</a>)</p></li><li><p><strong>Google quietly opted users</strong> into AI training on voice queries and uploaded media via a new Search Services History setting, with no opt-in required. (<a href="https://www.foxnews.com/tech/google-may-use-your-photos-voice-train-ai">Fox News</a>)</p></li><li><p><strong>Rime raises $24M Series A</strong> to build enterprise speech-to-speech models, powering nearly 100M phone calls monthly for Mayo Clinic, Dialpad, and others. (<a href="https://www.rime.ai/resources/rime-series-a-announcement">Rime</a>)</p></li></ul><ul><li><p><strong>LALAL.AI launches Lynx,</strong> a neural network built for speech denoising that is 6x smaller than its flagship model while matching output quality. (<a href="https://slator.com/lalal-ai-lynx-launch/">Slator</a>)</p></li><li><p><strong>Sber&#8217;s GigaChat adds emotion detection</strong> and can process audio up to three hours long with speaker separation, timestamps, and segment summaries. (<a href="https://businessnewsthisweek.com/business/gigachat-the-ai-assistant-detects-emotions-and-finds-content-in-long-audio-files/">BusinessNewsThisWeek</a>)</p></li><li><p><strong>DoorDash, ObserveAI, and AWS</strong> scale AI-powered quality evaluation across 19,000 agents, automating nearly 100% of interaction reviews. (<a href="https://www.prnewswire.com/news-releases/doordash-observeai-and-aws-partner-to-scale-customer-centric-ai-across-19-000-agents-302824595.html">PR Newswire</a>)</p></li><li><p><strong>Samsung adds cloud transcription</strong> to its Voice Recorder app, giving users a choice between on-device privacy and cloud-powered accuracy. (<a href="https://www.webpronews.com/samsung-voice-recorders-cloud-transcription-upgrade-signals-shift-in-on-device-ai-limits/">WebProNews</a>)</p></li><li><p><strong>Aina raises $5.5M</strong> to build hardware that controls AI agents rather than just recording, with its first product Dune already shipping to early adopters. (<a href="https://techcrunch.com/2026/07/16/ultrahumans-former-hardware-vp-raises-5-5m-for-devices-that-control-ai-agents-not-just-record-you/">TechCrunch</a>)</p></li><li><p><strong>Chen Institute and Science honor</strong> neuroscientist Sergey Stavisky for an AI speech neuroprosthesis that decodes brain activity into spoken words at 97.5% accuracy. (<a href="https://www.prnewswire.com/news-releases/ai-gives-people-back-their-own-voice-chen-institute-and-science-prize-honors-neuroscientist-sergey-stavisky-302827712.html">PR Newswire</a>)</p></li><li><p><strong>New research finds</strong> that AI voice phishing works because of persuasive scripting, not vocal realism, as 70% of targets detect the synthetic voice but comply anyway. (<a href="https://www.helpnetsecurity.com/2026/07/17/research-ai-voice-phishing/">Help Net Security</a>)</p></li><li><p><strong>Instadesk, Huawei, and iFlytek</strong> open a joint AI customer experience lab in Uzbekistan, combining multilingual ASR/TTS with Ascend cloud infrastructure. (<a href="https://www.manilatimes.net/2026/07/16/tmt-newswire/pr-newswire/instadesk-unveils-ai-customer-experience-lab-with-huawei-and-iflytek-in-central-asia/2385796">Manila Times</a>)</p></li><li><p><strong>VoicePing 3.0 launches</strong> with real-time translation, AI dubbing, new ASR and MT models, and MCP/API access for enterprise multilingual workflows. (<a href="https://www.prnewswire.com/news-releases/voiceping-releases-voiceping-3-0-for-enterprise-multilingual-communication-302823495.html">PR Newswire</a>)</p></li><li><p><strong>A study flags five risks</strong> in clinical AI scribes: inconsistent consent, weak performance on accented speech, background noise, missing human review, and unclear accountability. (<a href="https://www.resultsense.com/news/2026-07-17-clinical-ai-scribes-risks/">ResultSense</a>)</p></li><li><p><strong>Telcos are sitting on a voice AI opportunity</strong> bigger than their own cost savings, argues an analysis that says the real play is selling voice infrastructure to enterprises. (<a href="https://sebastianbarros.substack.com/p/the-telco-ai-voice-is-bigger-than">Sebastian Barros</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Apple&#8217;s SpeechAnalyzer API</strong> outperforms Whisper Small in English benchmarks, running fully on-device on iOS 26 and macOS Tahoe. (<a href="https://gigazine.net/gsc_news/en/20260714-apple-speech-analyzer-benchmark/#gsc.tab=0">Gigazine</a>)</p></li><li><p><strong>Cohere releases Transcribe Arabic,</strong> an open-source 2B-param ASR model that beats Meta&#8217;s 7B model on Arabic with multidialect and code-switching support. (<a href="https://cohere.com/blog/transcribe-arabic">Cohere</a>)</p></li><li><p><strong>Llamafile 0.10.4 ships transcribefile,</strong> a portable speech-to-text CLI built on transcribe.cpp that supports 16+ model families with GPU acceleration. (<a href="https://www.phoronix.com/news/Llamafile-0.10.4">Phoronix</a>)</p></li><li><p><strong>A prototype tongue-reading system</strong> uses ultrasound and ML to decode silent speech from tongue movements, enabling voice input without making a sound. (<a href="https://hackaday.com/2026/07/12/speak-silently-with-an-ultrasound-probe/">Hackaday</a>)</p></li><li><p><strong>Adafruit demos VAD, STT, and TTS</strong> all running on a single RP2040 microcontroller, bringing a complete voice AI pipeline to a $4 chip. (<a href="https://blog.adafruit.com/2026/07/15/voice-activity-detection-speech-to-text-and-text-to-speech-all-on-rp2040/">Adafruit</a>)</p></li><li><p><strong>ReSpeaker Clip</strong> is an open-source wearable AI recorder with dual mics, BLE 5.3, Wi-Fi 6, and full SDK for building custom voice AI applications. (<a href="https://www.seeedstudio.com/blog/2026/07/13/respeaker-clip-an-open-wearable-ai-recorder-for-building-voice-ai-applications/">Seeed Studio</a>)</p></li><li><p><strong>Resemble AI explains</strong> how neural audio watermarking embeds inaudible signals during voice generation for traceability, ahead of the EU AI Act&#8217;s August 2 deadline. (<a href="https://www.resemble.ai/resources/where-neural-audio-watermarking-fits-in-audio-security">Resemble AI</a>)</p></li><li><p><strong>A Python tutorial walks through</strong> building a real-time AI phone agent that quotes prices using tool calling and low-latency voice synthesis. (<a href="https://lowlatencyclub.ai/blog/posts/ai-price-quote-phone-agent-python">Low Latency Club</a>)</p></li><li><p><strong>Voice AI benchmarks hide a gap:</strong> systems scoring 600-800ms in the lab hit 2-4 seconds on real telephony, and accuracy drops from 51% to 26-38% under realistic audio. (<a href="https://embeddedcomputing.com/technology/ai-machine-learning/computer-vision-speech-processing/beyond-the-lab-why-voice-ai-must-be-benchmarked-under-real-world-conditions">Embedded Computing</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[GPT-Live Goes Full-Duplex, Taco Bell expands voice AI]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/gpt-live-goes-full-duplex-taco-bell</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/gpt-live-goes-full-duplex-taco-bell</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 13 Jul 2026 14:01:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/309dcc6d-4530-47b3-91e4-9e19e7019d6f_2408x1290.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong><a href="http://DeepLearning.AI">DeepLearning.AI</a> Voice AI Hackathon: </strong>In-person hackathon with Sabre + Vocal Bridge. Teams build voice agents that book real trips end-to-end, judged by Andrew Ng. (Jul 18, Mountain View, <a href="https://luma.com/fmypremp">Luma</a>)</p></li><li><p><strong>Voice Agent Evalathon: </strong>Okareo x Telnyx virtual challenge to red-team a voice agent&#8217;s reasoning, execution, and stability. (Jul 15, Virtual, <a href="https://www.meetup.com/ai-agent-simulation-reliability-group/events/315320251/">Meetup</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>OpenAI launches GPT-Live,</strong> a full-duplex voice model that listens and speaks simultaneously, powering a more natural ChatGPT Voice experience. (<a href="https://openai.com/index/introducing-gpt-live/">OpenAI</a>)</p></li><li><p><strong>xAI releases 21 multilingual flagship voices</strong> for Grok, expanding its voice lineup across multiple languages. (<a href="https://x.ai/news/new-flagship-voices">xAI</a>)</p></li><li><p><strong>Gradium raises $100M seed</strong> backed by NVIDIA, making the Paris-based Kyutai spinout one of the largest seed rounds ever for ultra-low latency voice AI. (<a href="https://techcrunch.com/2026/07/09/paris-based-ai-voice-startup-gradium-raises-100m-seed-backed-by-nvidia/">TechCrunch</a>)</p></li><li><p><strong>Rylo AI raises $85M</strong> to scale its real-time captioning platform for deaf and hard-of-hearing users, rebranded from Nagish. (<a href="https://www.alleywatch.com/2026/07/rylo-ai-communication-accessibility-deaf-hard-of-hearing-real-time-captioning-platform-tomer-aharoni/">AlleyWatch</a>)</p></li><li><p><strong>Five9 launches next-gen Voice AI Agents</strong> built on a purpose-built agentic architecture (<a href="http://insidermonkey.com/blog/five9-inc-fivn-reveals-next-gen-voice-ai-agents-1797773/">Insider Monkey</a>)</p></li><li><p><strong>Taco Bell expands voice AI</strong> to 890+ drive-thrus across 38 states, powered by Omilia. (<a href="https://www.fermag.com/articles/taco-bell-expands-voice-ai-at-us-drive-thrus/">FER Magazine</a>)</p></li><li><p><strong>Omilia launches Lexis,</strong> a native generative TTS engine for enterprise CX with sub-45ms latency. (<a href="https://aithority.com/it-and-devops/cloud/omilia-launches-the-only-native-voice-in-enterprise-cx/">AIthority</a>)</p></li><li><p><strong>Dell AI Factory partners with Deepgram</strong> and Penguin Solutions to deliver enterprise-grade real-time voice AI infrastructure. (<a href="https://www.businessinsider.com/sc/dell-ai-factory-powers-real-time-voice-ai">Business Insider</a>)</p></li><li><p><strong>SoundHound&#8217;s OASYS platform</strong> wins &#8220;Agentic AI Company of the Year&#8221; as Q1 revenue hits $44M, up 52% year-over-year. (<a href="https://www.nasdaq.com/articles/soundhound-trends-voice-ai-oasys-and-rising-agentic-demand-today">Nasdaq</a>)</p></li><li><p><strong>CallTower partners with Sestek</strong> to add conversational AI, voice biometrics, and real-time translation to its CX portfolio. (<a href="https://telecomreseller.com/2026/07/10/calltower-expands-ai-for-cx-portfolio-through-strategic-partnership-with-sestek/">Telecom Reseller</a>)</p></li><li><p><strong>Whispp raises $5M</strong> to scale its on-device AI that reconstructs speech for people with voice disorders in real time. (<a href="https://pulse2.com/whispp-raises-e5-million-to-scale-real-time-on-device-voice-reconstruction-ai/amp/">Pulse 2.0</a>)</p></li><li><p><strong>AI voice agents boosted specialty care enrollment 340%</strong> in a peer-reviewed clinical study by RadiantGraph. (<a href="https://www.prnewswire.com/news-releases/study-finds-ai-voice-agents-increased-specialty-care-program-enrollment-rates-340-in-real-world-clinical-setting-302818900.html">PR Newswire</a>)</p></li><li><p><strong>A $25M deepfake scam at Arup</strong> used AI-generated executives on a video call, becoming a landmark case for corporate voice fraud. (<a href="https://financefeeds.com/the-25-million-ai-deepfake-scam-that-changed-corporate-security/">FinanceFeeds</a>)</p></li><li><p><strong>Reality Defender warns</strong> that autonomous AI callers can now bypass contact center defenses at scale, posing a new voice fraud threat. (<a href="https://www.biometricupdate.com/202607/the-agentic-caller-always-rings-at-scale-reality-defender-explores-new-ai-voice-threat">Biometric Update</a>)</p></li><li><p><strong>Hydaway launches RealityChek,</strong> a streaming audio deepfake detector for enterprises that flags synthetic speech in real time. (<a href="https://www.biometricupdate.com/202607/hydaway-introduces-real-time-enterprise-audio-deepfake-detection">Biometric Update</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>OpenAI ships GPT-Realtime-2.1-mini</strong> with reasoning and tool use support at 6x lower cost and 25% reduced latency. (<a href="https://www.marktechpost.com/2026/07/06/openai-gpt-realtime-2-1-mini-reasoning-realtime-api/">Marktechpost</a>)</p></li><li><p><strong>AssemblyAI details how Universal-3.5 Pro</strong> handles noisy audio, sharing techniques for improving transcription accuracy on hard recordings. (<a href="https://www.assemblyai.com/blog/async-transcription-accuracy-hard-audio">AssemblyAI</a>)</p></li><li><p><strong>Flock Safety explains its audio detection system</strong> that identifies gunshots and crashes in real time using acoustic sensors and AI classification. (<a href="https://www.flocksafety.com/blog/how-flocks-audio-detection-works">Flock Safety</a>)</p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[xAI Ships Voice Agent Builder, Krisp named 2026 Disruptive Technology]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/xai-ships-voice-agent-builder-krisp</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/xai-ships-voice-agent-builder-krisp</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 06 Jul 2026 14:02:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/68229207-8166-4e14-a810-1bea7989b6b4_1426x740.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>Real-Time Observability for Production Voice AI</strong> Regal walks through its Observability Dashboard with real catch-and-fix examples from production voice agents. (Jul 9, SF, <a href="https://www.voiceaispace.com/events/do-you-know-what-your-ai-agent-is-doing-real-time-observability-for-production-voice-ai">Voice AI Space</a>)</p></li><li><p><strong>AI Tinkerers San Francisco: July GTM Engineering Track w/ Attio</strong> Builder-only, no-pitch meetup with live GTM engineering demos - a solid room for voice agent developers. (Jul 8, SF, <a href="https://sf.aitinkerers.org/p/ai-tinkerers-san-francisco-july-gtm-engineering-track-w-attio">AI Tinkerers</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>xAI launches Voice Agent Builder</strong>, a no-code platform for production voice agents with telephony and guardrails at $0.05/min. (<a href="https://x.ai/news/grok-voice-agent-builder">xAI</a>)</p></li></ul><ul><li><p><strong>Krisp named 2026 Disruptive Technology</strong> of the Year by CMP Research for its voice AI infrastructure. (<a href="https://krisp.ai/blog/krisp-named-2026-disruptive-technology-of-the-year/">Krisp Blog</a>)</p></li><li><p><strong>ElevenLabs explores a $22B tender offer</strong>, doubling its valuation from the $11B Series D five months ago. (<a href="https://www.techinasia.com/news/voice-ai-startup-elevenlabs-seeks-22b-sources">Tech in Asia</a>)</p></li><li><p><strong>Pocket raises $11M</strong> from Accel and YC for its $129 AI note-taking puck that has shipped 130K units. (<a href="https://techcrunch.com/2026/06/29/pocket-raises-11m-in-bet-on-rising-demand-for-ai-note-taking-devices/">TechCrunch</a>)</p></li><li><p><strong>Lucida AI raises &#8364;6.1M</strong> seed for its speech-to-speech language coaching platform, now at 3M users. (<a href="https://www.eu-startups.com/2026/06/uk-speech-ai-startup-lucida-ai-lands-e6-1-million-to-develop-speech-native-ai-for-global-communication/">EU-Startups</a>)</p></li><li><p><strong>US senators revive the AI Labeling Act</strong>, a bipartisan bill requiring AI-generated audio and video to carry disclosure labels. (<a href="https://www.musicbusinessworldwide.com/us-senators-revive-bill-to-force-ai-generated-audio-video-and-images-to-carry-labels/">Music Business Worldwide</a>)</p></li><li><p><strong>Retell AI launches Conductor</strong>, a graph-native review interface with an AI copilot for production voice agents. (<a href="https://aithority.com/machine-learning/voice-ai-startup-retell-ai-launches-conductor-featuring-the-first-ever-graph-native-review-interface-for-production-voice-agents/">AIthority</a>)</p></li><li><p><strong>Syntiant and Vibe partner</strong> to bring voice-enabled AI to smart workspace hardware using edge AI chips. (<a href="https://www.manilatimes.net/2026/06/30/tmt-newswire/globenewswire/syntiant-and-vibe-collaborate-to-advance-voice-enabled-ai-driven-workspace-experiences/2375782">GlobeNewsWire</a>)</p></li><li><p><strong>HealthLynked launches</strong> an AI healthcare platform with 24/7 scheduling and medical office voice agents. (<a href="https://www.globenewswire.com/news-release/2026/06/29/3318905/0/en/healthlynked-launches-ai-powered-healthcare-communication-platform-featuring-24-7-appointment-scheduling-and-medical-office-ai-agents.html">GlobeNewsWire</a>)</p></li><li><p><strong>RevComm launches MiiTel for Retail</strong>, extending its voice AI analytics to in-store customer conversations. (<a href="https://jp.ibtimes.com/revcomm-launches-miitel-retail-store-voice-ai-102266">IBTimes JP</a>)</p></li><li><p><strong>Patient trust is the biggest barrier</strong> to healthcare voice AI, not the technology, argues a Forbes analysis. (<a href="https://www.forbes.com/councils/forbestechcouncil/2026/06/30/the-hardest-problem-in-healthcare-voice-ai-isnt-the-technology-its-patient-trust/">Forbes</a>)</p></li><li><p><strong>Voices similar to our own are more persuasive</strong>, finds new research raising concerns about companies weaponizing stored voice data. (<a href="https://nautil.us/the-dangers-of-ai-voice-clones-1282420">Nautilus</a>)</p></li><li><p><strong>Synthflow deployed voice AI in one day</strong> for Nellis Auction, now handling 80% of each customer interaction automatically. (<a href="https://cxm.world/customer-experience/the-phone-line-that-hung-up-on-customers-and-the-voice-ai-that-fixed-it-in-a-day/">CXM World</a>)</p></li></ul><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Mozilla AI releases transcribe-cpp</strong>, an open-source C/C++ STT library with ggml runtime and GPU acceleration. (<a href="https://www.startuphub.ai/ai-news/technology/2026/mozilla-ai-unveils-transcribe-cpp">StartupHub</a>)</p></li></ul><ul><li><p><strong>ViiTorVoice-NAR goes open source</strong> with word-level TTS editing that swaps individual words without regenerating surrounding audio. (<a href="https://www.techtimes.com/articles/319524/20260702/text-speech-ai-edits-single-words-mid-recording-viitorvoice-goes-open-source.htm">TechTimes</a>)</p></li><li><p><strong>Higgs TTS 2 3B</strong> from BosonAI is a 5.8B-param TTS model trained on 10M+ hours with zero-shot voice cloning. (<a href="https://hackernoon.com/the-higgs-tts-2-3b-base-model-a-text-to-speech-foundation-model">HackerNoon</a>)</p></li><li><p><strong>Vowen 0.4.8 released</strong>, a free offline voice productivity app using Whisper-based local transcription. (<a href="https://www.warp2search.net/story/vowen-048-released/">Warp2Search</a>)</p></li><li><p><strong>WhisTam</strong> is a Whisper-based framework for Tamil dialect speech recognition, ranking 2nd at DravidianLangTech@ACL 2026. (<a href="https://aclanthology.org/2026.dravidianlangtech-1.70/">ACL Anthology</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[$200M Pours Into Voice AI, OpenAI Bidi-1 Leaks]]></title><description><![CDATA[Voice AI weekly digest]]></description><link>https://voice-ai-newsletter.krisp.ai/p/200m-pours-into-voice-ai-openai-bidi</link><guid isPermaLink="false">https://voice-ai-newsletter.krisp.ai/p/200m-pours-into-voice-ai-openai-bidi</guid><dc:creator><![CDATA[Davit Baghdasaryan]]></dc:creator><pubDate>Mon, 29 Jun 2026 14:03:12 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e90060d3-0692-46ad-84f2-adf95038b653_1670x928.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Events </h2><ul><li><p><strong>AI Engineer World&#8217;s Fair</strong> is a flagship AI engineering conference with a dedicated Voice &amp; Realtime AI miniconference featured this year (Jun 29-Jul 2, SF | <a href="https://www.ai.engineer/worldsfair/2026">AI Engineer</a>)</p></li><li><p><strong>Low Latency Lounge by Deepgram</strong> is an invite only evening for engineers building the fastest AI in the stack. Together AI and Runware are cohosting (Jun 30, SF | <a href="https://luma.com/low-latency-lounge">LUMA</a>)</p></li><li><p><strong>Real-Time Voice AI &#215; Device Builders Meetup</strong> &#8220;Give Voice to Robots!&#8221; Runs alongside IVS Kyoto (Jul 2, Kyoto | <a href="https://www.voiceaispace.com/events/-real-time-voice-ai--device-builders-meetup-kyoto">Voice AI Space</a>)</p></li></ul><h2>Top Updates &#128170;</h2><ul><li><p><strong>AssemblyAI launches</strong> Universal-3.5 Pro Realtime, the first streaming STT model that takes the agent&#8217;s question as input (<a href="https://www.assemblyai.com/blog/universal-3-5-pro-realtime">AssemblyAI Blog</a>)</p></li><li><p><strong>Five9 launches</strong> Voice AI Agents and AI Agent Studio at CCW, bringing agentic CX to enterprise contact centers. (<a href="https://www.cxtoday.com/ai-automation-in-cx/five9-voice-ai-agents-agentic-cx-launch/">CX Today</a>)</p></li><li><p><strong>Krisp launches</strong> Voice Security for deepfake detection and fraud detection for contact centers. (<a href="https://www.cxtoday.com/security-privacy-compliance/krisp-expands-contact-center-ai-platform-with-voice-security-and-speech-analytics/">CX Today</a>)</p></li><li><p><strong>CallMiner launches</strong> real-time AI guidance that lets contact center agents initiate AI assistance on demand with human-in-the-loop controls. (<a href="https://www.businesswire.com/news/home/20260622825087/en/CallMiner-Enhances-Real-Time-Agent-Performance-and-Customer-Experience-with-New-AI-Capabilities">BusinessWire</a>)</p></li><li><p><strong>Assort Health raises</strong> $120M Series C led by Menlo Ventures at a $1.2B valuation to scale its voice AI agent platform across healthcare. (<a href="https://www.fiercehealthcare.com/ai-and-machine-learning/assort-health-scores-120m-series-c-scale-voice-ai-agent-platform-healthcare">Fierce Healthcare</a>)</p></li><li><p><strong>Prosper AI raises</strong> $30M Series A led by a16z to scale its autonomous patient journey platform, reporting 5x revenue growth in six months. (<a href="https://hackernoon.com/prosper-ai-raises-$30m-led-by-a16z-to-scale-autonomous-patient-journey-platform">HackerNoon</a>)</p></li><li><p><strong>Coval raises</strong> $28M Series A led by Norwest to advance its voice AI evaluation and testing platform, founded by an ex-Waymo engineer. (<a href="https://pulse2.com/coval-raises-28-million-series-a-to-advance-voice-ai-evaluation-platform/">Pulse2</a>)</p></li><li><p><strong>Kotoba Technologies raises</strong> $10M seed led by Kindred Ventures for its real-time East Asian voice translation platform with sub-2s latency. (<a href="https://gamesbeat.com/kotoba-technologies-raises-10m-for-real-time-voice-ai-platform-in-east-asia/">VentureBeat</a>)</p></li><li><p><strong>Valence AI raises</strong> $5M seed and secures US patents on real-time emotional detection from live speech. (<a href="https://www.prnewswire.com/news-releases/valence-ai-raises-5-million-secures-us-patents-on-real-time-emotional-detection-from-live-speech-302808293.html">PR Newswire</a>)</p></li><li><p><strong>TELUS Digital partners</strong> with ElevenLabs as a preferred implementation partner to scale voice AI alongside frontline customer care teams. (<a href="https://www.prnewswire.com/news-releases/telus-digital-and-elevenlabs-partner-to-scale-voice-ai-alongside-frontline-customer-care-teams-882149628.html">PR Newswire</a>)</p></li><li><p><strong>OpenAI&#8217;s GPT-Bidi-1</strong> leaks as a full-duplex voice model that can listen and speak simultaneously, enabling true bidirectional conversation. (<a href="https://cryptobriefing.com/openai-chatgpt-bidi-1-voice-model/">Crypto Briefing</a>)</p></li><li><p><strong>Conduent unveils</strong> a next-gen CX platform with real-time translation across 90+ languages to accelerate agent performance. (<a href="https://www.news.conduent.com/news/conduent-introduces-ai-powered-next-generation-cx-platform-to-expand-global-customer-reach-and-accelerate-agent-performance">Conduent</a>)</p></li><li><p><strong>Speechify brings</strong> free voice typing to all iPhone and Mac users, adding AI-powered dictation across every app. (<a href="https://9to5mac.com/2026/06/23/speechify-brings-voice-typing-to-all-iphone-and-mac-users/">9to5Mac</a>)</p></li><li><p><strong>Modulate launches</strong> an AI music detection API with 95% precision across 76 genres to help platforms verify AI-generated music. (<a href="https://www.morningstar.com/news/accesswire/1181688msn/modulate-launches-ai-music-detection-api-to-help-platforms-verify-ai-generated-music-at-scale">Morningstar</a>)</p></li><li><p><strong>ByteDance releases</strong> Seed Audio 1.0, a unified model that generates speech, music, and ambient sound from a single architecture. (<a href="https://www.citybuzz.co/2026/06/25/seed-audio-1-0-launches-unified-ai-audio-generation-for-speech-music-and-ambient-sound/">CityBuzz</a>)</p></li><li><p><strong>Amazon launches</strong> Alexa Plus Hindi beta in India, targeting 600M+ Hindi speakers with its upgraded AI assistant. (<a href="https://thenextweb.com/news/amazon-alexa-plus-india-hindi-beta-testing">The Next Web</a>)</p></li><li><p><strong>ElevenLabs adopts</strong> Google&#8217;s SynthID watermarking to tag all AI-generated speech, making synthetic voices easier to detect. (<a href="https://www.digitaltrends.com/cool-tech/ai-voices-are-getting-harder-to-spot-this-elevenlabs-feature-could-change-that/">Digital Trends</a>)</p></li><li><p><strong>Shure says</strong> audio quality is now the critical bottleneck for AI-powered meetings, and microphone clarity drives everything. (<a href="https://www.inavateonthenet.net/features/article/shure-says-audio-is-now-critical-to-ai-meetings--and-clarity-is-everything">InAVate</a>)</p></li><li><p><strong>Attention Labs launches</strong> SAA, a selective auditory attention layer that lets voice AI detect when it is being directly addressed. (<a href="https://www.dispatch.com/press-release/story/203592/attention-labs-launches-saa-the-engagement-control-layer-that-lets-voice-ai-know-when-it-is-being-addressed/">Dispatch</a>)</p></li><li><p><strong>Deepgram and Fortanix</strong> partner to run voice AI on-premises with NVIDIA confidential computing, keeping audio data encrypted during processing. (<a href="https://radioinfo.com.au/audioinfo/technology-news/private-voice-ai-introduced-to-better-protect-your-audio/">RadioInfo</a>)</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://voice-ai-newsletter.krisp.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Engineering Corner &#128526;</h2><ul><li><p><strong>Gradium releases</strong> STT-Translate and S2S-Translate, real-time speech translation models that beat GPT Realtime Translate on accuracy and latency. (<a href="https://www.marktechpost.com/2026/06/24/gradium-launches-stt-translate-and-s2s-translate-real-time-speech-translation-models-beating-gpt-realtime-translate-on-accuracy-and-latency/">MarkTechPost</a>)</p></li></ul><ul><li><p><strong>AWS publishes</strong> a full tutorial on building a healthcare appointment agent with Amazon Nova 2 Sonic and Bedrock AgentCore. (<a href="https://aws.amazon.com/blogs/machine-learning/build-a-healthcare-appointment-agent-with-amazon-nova-2-sonic/">AWS Blog</a>)</p></li><li><p><strong>AssemblyAI shares</strong> four techniques for prompting Claude to build production-ready voice agents in about 30 seconds. (<a href="https://www.assemblyai.com/blog/prompting-claude-build-voice-agents">AssemblyAI Blog</a>)</p></li><li><p><strong>Deepgram discusses</strong> voice AI infrastructure and the path to production-grade agents on the Telecom Reseller podcast. (<a href="https://telecomreseller.com/2026/06/24/deepgram-on-voice-ai-infrastructure-and-the-road-to-production-grade-agents-podcast/">Telecom Reseller</a>)</p></li><li><p><strong>ACL 2026 publishes</strong> 10 voice AI papers covering noise-robust ASR, accented speech recognition, environment-aware TTS, controllable speech synthesis, multi-speaker diarization, and multilingual translation. (<a href="https://aclanthology.org/volumes/2026.acl-long/">ACL Anthology</a>)</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://voice-ai-newsletter.krisp.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get the most important news in Voice AI delivered directly to your inbox every week</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>