Tag
4 articles
This article explains Google's new Gemini 3.5 Transcribe speech-to-text model, detailing its dual-endpoint architecture, technical mechanisms, and implications for developers building voice agents and transcription systems.
Google's new Gemini 3.5 Transcribe supports 85 languages and features real-time auto-correction of verbal stumbles with a 4.0% word error rate.
This article explains speech-to-text technology and how Microsoft's new MAI-Transcribe-1.5 model improves speed, accuracy, and language support for converting spoken words into text.
Microsoft's new MAI-Transcribe-1 model runs 2.5x faster than its predecessor and costs just $0.36 per audio hour, making it a powerful tool for transcription across 25 languages.