Top Updates 💪
Wispr raises $280M at a $2B valuation as it moves beyond dictation, previewing Canto, a speech model that cuts noisy-condition error rates sharply. (SiliconAngle)
Cartesia ships Sonic 3.6, a streaming TTS model that now tops both Artificial Analysis speech arenas with sub-90ms time-to-first-audio across 44 languages. (Slator)
Adobe Firefly adds AI music, speech, and sound effects, folding audio generation into its creative studio with commercial usage rights. (Adobe)
HappyRobot raises $150M at a $1.2B valuation to run freight operations by voice, automating broker and carrier calls end to end. (Forkast)
Krisp joins the 8x8 Technology Partner Ecosystem, bringing its noise cancellation and accent conversion to the 8x8 CX platform. (MarTech Series)
Murf AI launches Falcon 2, a voice model priced at $0.01/min that ranks ahead of OpenAI Realtime and ElevenLabs Flash on naturalness. (Storyboard18)
Sarvam’s Saaras V3 outperforms OpenAI and ElevenLabs on Indian-language and Indian-accented speech benchmarks. (CryptoBriefing)
Breeze Blue unveils Breeze TTS 2, a real-time model for games and digital companions that tops the TTS Voice Design benchmark. (Cincinnati.com)
AudioStack adds Gradium’s voices to its AI audio platform, expanding regional accent coverage for media and advertising customers. (Slator)
AudioCodes and AnywhereNow earn Microsoft Teams certification, with AudioCodes’ Voca CIC among the first certified Teams voice agents. (SpeechTech)
Reddit tests turning text posts into AI-narrated videos, adding a “Play” toggle that voices threads to capture content it once lost to TikTok. (The Verge)
Calendly launches Notetaker and the Callie assistant, focusing AI on the follow-up work that happens after a meeting ends. (AutoGPT)
Boya unveils the Notra Neo AI note taker, a $149 device with dual mics and 30dB noise reduction for meetings, calls, and Bluetooth audio. (China Gadgets Reviews)
JPMorgan Chase warns AI is making social engineering easier, as attackers clone voices from short clips to impersonate staff. (JPMorgan Chase)
The ABA examines the voice fraud threat to banking, noting detection still degrades under real phone-line noise and compression. (ABA Banking Journal)
Logitech says audio accuracy is the key as voice dictation spreads through offices, with one in five AI interactions now voice-based. (IT Brief)
Last Week Podcast
Engineering Corner 😎
superwhisper open-sources S1-mini, a 462MB text normalizer that turns raw ASR transcripts into clean written text on a laptop CPU. (MarkTechPost)
Nari Labs pushes the Qwen3-TTS speed-cost frontier, hitting sub-50ms time-to-first-audio at ~$2 per 1M characters on a single H100. (Nari Labs)
A HackerNoon team builds the same STT app cloud and local, comparing cost, latency, and accuracy trade-offs of each path. (HackerNoon)
A dev.to post tests 4 TTS engines on 12,000 live healthcare calls, reporting which one patients responded to best in production. (dev.to)
A guide on streaming ASR vs Whisper on mobile explains when to switch to a streaming model for sub-second, interactive voice. (dev.to)
A developer explains why WhatsApp voice notes break transcription, from background noise to code-switching and mixed languages. (dev.to)
A dev builds a free no-signup audio-and-video-to-text tool, sharing the pipeline behind browser-based transcription. (dev.to)
A developer builds a browser voice-to-writing tool, arguing typing slows down thinking and dictation keeps ideas flowing. (dev.to)

