An API powered by Google's AI technology allows you to accurately convert speech into text. You can accurately caption your content, provide a better user experience with products using voice commands, and gain insight from customer interactions to improve your service. Google's deep learning neural network algorithms are the most advanced in automatic speech recognition (ASR). Speech-to-Text allows for experimentation, creation, management, and customization of custom resources. You can deploy speech recognition wherever you need it, whether it's in the cloud using the API or on-premises using Speech-to-Text O-Prem. You can customize speech recognition to translate domain-specific terms or rare words. Automated conversion of spoken numbers into addresses, years and currencies. Our user interface makes it easy to experiment with your speech audio.
Learn more

Fathom is an AI meeting assistant that helps users capture, summarize, search, and act on meetings with less manual work. The platform creates accurate transcripts, instant summaries, action items, and follow-up notes so users can focus on live conversations instead of taking notes. Fathom supports both traditional meeting capture and bot-free capture through its desktop app. Teams can use Fathom as a shared source of truth across customer calls, internal meetings, strategy sessions, and project conversations. Ask Fathom lets users search across meetings and ask questions about conversations, decisions, commitments, risks, and next steps. The platform also supports topic monitoring so important moments and signals are easier to find. Fathom syncs meeting notes, insights, and action items into tools such as Slack, Salesforce, HubSpot, Notion, Asana, Gmail, Zoom, Google Meet, Microsoft Teams, ChatGPT, Claude, Zapier, and API or MCP workflows. It supports security and compliance needs with SOC 2 Type II, GDPR, HIPAA compliance, SSO, and SCIM. By combining AI notetaking, bot-free capture, transcripts, summaries, integrations, search, and workflow automation, Fathom helps teams move from meetings to execution faster.
Learn more
Wispr Flow
Wispr Flow is an AI-powered voice dictation platform that helps users write faster by speaking instead of typing. The app works across Mac, Windows, iPhone, and Android and can be used inside everyday applications for messages, emails, documents, code, notes, and workflows. Wispr Flow transcribes natural speech and automatically turns it into clearer, more polished writing by removing filler words, correcting mistakes, and improving structure. The platform is designed to help users create, code, message, and write at the speed of thought, with positioning around being four times faster than typing. AI Auto Edits help transform unstructured spoken thoughts into formatted, readable text without requiring manual cleanup. A personal dictionary helps Flow learn names, technical terms, company words, and other unique vocabulary. Snippet shortcuts let individuals and teams speak short cues that expand into frequently used formatted text. Wispr Flow also supports more than 100 languages and automatically detects language changes during dictation. By combining voice-to-text, AI rewriting, cross-app support, personal vocabulary, snippets, and multilingual transcription, Wispr Flow helps users turn speech into usable writing anywhere they work.
Learn more
Google AI Edge Eloquent
Google AI Edge Eloquent is a sophisticated dictation application powered by artificial intelligence that converts spoken language into refined, professional text directly on mobile devices. Utilizing Google's cutting-edge Gemma technology, it effectively closes the gap between unrefined speech and well-crafted written communication, surpassing conventional speech-to-text applications that merely capture every utterance and mistake as they are spoken. The app intelligently discards filler words like “ums” and “uhs” as well as mid-sentence corrections, ensuring that the resulting text reflects the user’s intended message with clarity and precision. It provides real-time transcription while users speak, followed by a smart text enhancement process after recording is halted, and can generate various output formats, including concise bullet points, formal prose, and both shorter and longer adaptations. Operating primarily on-device through efficient AI Edge runtimes, it ensures quick responsiveness without needing a server connection, thus facilitating complete offline functionality. This innovative approach allows users to maintain their focus on the content rather than the mechanics of dictation.
Learn more