An API powered by Google's AI technology allows you to accurately convert speech into text. You can accurately caption your content, provide a better user experience with products using voice commands, and gain insight from customer interactions to improve your service. Google's deep learning neural network algorithms are the most advanced in automatic speech recognition (ASR). Speech-to-Text allows for experimentation, creation, management, and customization of custom resources. You can deploy speech recognition wherever you need it, whether it's in the cloud using the API or on-premises using Speech-to-Text O-Prem. You can customize speech recognition to translate domain-specific terms or rare words. Automated conversion of spoken numbers into addresses, years and currencies. Our user interface makes it easy to experiment with your speech audio.
Learn more

Fathom is an AI meeting assistant that helps users capture, summarize, search, and act on meetings with less manual work. The platform creates accurate transcripts, instant summaries, action items, and follow-up notes so users can focus on live conversations instead of taking notes. Fathom supports both traditional meeting capture and bot-free capture through its desktop app. Teams can use Fathom as a shared source of truth across customer calls, internal meetings, strategy sessions, and project conversations. Ask Fathom lets users search across meetings and ask questions about conversations, decisions, commitments, risks, and next steps. The platform also supports topic monitoring so important moments and signals are easier to find. Fathom syncs meeting notes, insights, and action items into tools such as Slack, Salesforce, HubSpot, Notion, Asana, Gmail, Zoom, Google Meet, Microsoft Teams, ChatGPT, Claude, Zapier, and API or MCP workflows. It supports security and compliance needs with SOC 2 Type II, GDPR, HIPAA compliance, SSO, and SCIM. By combining AI notetaking, bot-free capture, transcripts, summaries, integrations, search, and workflow automation, Fathom helps teams move from meetings to execution faster.
Learn more
FluidVoice
FluidVoice is a free and open-source dictation application for macOS that combines local speech recognition with an on-device AI model known as Fluid-1, which enhances the quality of dictation. By using a single hotkey, users can dictate text into virtually any input field across various applications such as email, documents, chat interfaces, terminals, code editors, and more, with the text being displayed almost instantaneously. The application relies on local speech models that function offline, allowing for secure dictation without needing an internet connection, while optional AI post-processing can utilize Fluid Intelligence, OpenAI, Groq, or other custom providers. Fluid-1 improves initial dictation by refining rough entries, correcting formatting, capitalization, dates, names, and numbers, and it adjusts the tone according to the currently active application, all while preserving the speaker's intended meaning. Users have the flexibility to develop personalized prompts tailored to different applications, and with modes like Write Mode, Command Mode, and Direct Dictation, transitioning between tasks is seamless. Furthermore, FluidVoice is capable of supporting over 40 languages, leveraging various models such as Nemotron Speech 3.5, Parakeet Flash, Parakeet TDT versions 2 and 3, Cohere Transcribe, Apple Speech, and Whisper, thus catering to a diverse user base and enhancing accessibility in dictation across different linguistic backgrounds. This versatility makes FluidVoice an essential tool for those seeking effective and efficient dictation solutions.
Learn more
Apple Dictation
Apple Dictation is an integrated feature in macOS that allows users to input text by speaking in any application where text can be entered. Once the feature is activated, users can initiate it using the Microphone key, a personalized keyboard shortcut, or by selecting Start Dictation, and can terminate it with the Escape key, the same shortcut, or the Microphone key again. For those using Apple silicon Macs, it is possible to continue typing while dictating, with Dictation resuming listening once keyboard activity ceases. Users can issue spoken commands to insert emoji, punctuation marks, new lines, and new paragraphs, and the system can automatically add commas, periods, and question marks in supported languages. Dictation is capable of handling text of any length and will cease automatically after 30 seconds of no detected speech input. If macOS is unsure about a particular word, it highlights the text for users to either choose an alternative or make a manual correction. Additionally, users can enable multiple Dictation languages and switch between them during their speech, while also having the option to select their preferred microphone or allow macOS to make that choice automatically. This flexibility makes Apple Dictation a powerful tool for enhancing productivity and accessibility across macOS applications.
Learn more