An API powered by Google's AI technology allows you to accurately convert speech into text. You can accurately caption your content, provide a better user experience with products using voice commands, and gain insight from customer interactions to improve your service. Google's deep learning neural network algorithms are the most advanced in automatic speech recognition (ASR). Speech-to-Text allows for experimentation, creation, management, and customization of custom resources. You can deploy speech recognition wherever you need it, whether it's in the cloud using the API or on-premises using Speech-to-Text O-Prem. You can customize speech recognition to translate domain-specific terms or rare words. Automated conversion of spoken numbers into addresses, years and currencies. Our user interface makes it easy to experiment with your speech audio.
Learn more

Fathom is an AI meeting assistant that helps users capture, summarize, search, and act on meetings with less manual work. The platform creates accurate transcripts, instant summaries, action items, and follow-up notes so users can focus on live conversations instead of taking notes. Fathom supports both traditional meeting capture and bot-free capture through its desktop app. Teams can use Fathom as a shared source of truth across customer calls, internal meetings, strategy sessions, and project conversations. Ask Fathom lets users search across meetings and ask questions about conversations, decisions, commitments, risks, and next steps. The platform also supports topic monitoring so important moments and signals are easier to find. Fathom syncs meeting notes, insights, and action items into tools such as Slack, Salesforce, HubSpot, Notion, Asana, Gmail, Zoom, Google Meet, Microsoft Teams, ChatGPT, Claude, Zapier, and API or MCP workflows. It supports security and compliance needs with SOC 2 Type II, GDPR, HIPAA compliance, SSO, and SCIM. By combining AI notetaking, bot-free capture, transcripts, summaries, integrations, search, and workflow automation, Fathom helps teams move from meetings to execution faster.
Learn more
Dictation Speech to Text
You now have the ability to enhance speech recognition by adding personalized words! You can find this feature in the setup under manage custom words. The Dictation Speech to Text feature allows you to dictate, record, translate, and transcribe text, eliminating the need for manual typing. It utilizes cutting-edge voice recognition technology, primarily designed for converting speech into text and facilitating translation for messaging. Forget about typing; simply use your voice to dictate and translate! Almost all messaging applications can be adjusted to work seamlessly with the 'Dictation Speech to Text' function. This tool employs the integrated speech recognition engine for accurate results. Supporting over 40 languages, Dictation Speech to Text provides three text zones, marked by language flags, enabling you to set different languages in your preferences. This setup allows for effortless switching between various language projects with a single click. Translation is incredibly simple—just tap the translation button! Additionally, you can choose your desired target language for translation in the app's settings, making the process even more user-friendly and efficient.
Learn more
Spokenly
Spokenly is an innovative dictation application powered by AI, available for Mac, iPhone, Windows, and Linux, designed to convert spoken words into clear, punctuated text in any working environment. By simply holding a shortcut, users can speak naturally and then release to insert the transcription directly at the cursor across various platforms including browsers, email, chat applications, word processors, IDEs, terminals, and more. This versatile app accommodates over 100 languages, supporting mixed-language dictation, and provides both local and cloud-based speech-to-text models. Users can utilize on-device models like Whisper and Parakeet for offline operation, while cloud services from companies such as OpenAI, Deepgram, Groq, Soniox, and ElevenLabs can be accessed for enhanced accuracy or real-time transcription needs. Additionally, the Local Only Mode ensures that voice data remains solely on the device, preventing any network interactions. The application features modes that allow users to save different transcription models, select AI providers, set prompts, and choose output styles tailored for specific tasks. Furthermore, the AI Instructions feature enables users to eliminate filler words, correct grammar and punctuation, summarize, rewrite, translate, or reformat the dictated text, enhancing the overall functionality and user experience of the app. With its extensive capabilities, Spokenly stands out as a comprehensive solution for anyone looking to streamline their dictation process.
Learn more