An API powered by Google's AI technology allows you to accurately convert speech into text. You can accurately caption your content, provide a better user experience with products using voice commands, and gain insight from customer interactions to improve your service. Google's deep learning neural network algorithms are the most advanced in automatic speech recognition (ASR). Speech-to-Text allows for experimentation, creation, management, and customization of custom resources. You can deploy speech recognition wherever you need it, whether it's in the cloud using the API or on-premises using Speech-to-Text O-Prem. You can customize speech recognition to translate domain-specific terms or rare words. Automated conversion of spoken numbers into addresses, years and currencies. Our user interface makes it easy to experiment with your speech audio.
Learn more

Most AI video tools hand you a black box: closed weights, a subscription, and no way to see what is happening under the hood. LTX takes the opposite approach. Built by Lightricks, LTX is an open foundation model that generates and simulates across video, audio, and the physical world, and it puts the weights, the code, and the control in your hands.
At the center of the model is LTX-2.5, a 22B-parameter dual-stream diffusion transformer that produces native 4K video at up to 50 frames per second, with audio and video generated together in a single pass rather than stitched together afterward. Artificial Analysis, an independent benchmarking group, currently ranks LTX among the top three AI video models in the world.
You choose how you want to use it. Download the open weights and run LTX-2.5 on your own hardware. License the model for on-premise deployment backed by enterprise support. Or build directly on LTX Studio, the production suite that turns the model into a full creative workflow. Companies like ElevenLabs, Asteria Film Co., Magnopus, and NVIDIA already rely on LTX for their own work.
LTX is not built for one-off social clips. It is infrastructure for teams that generate motion, audio, and physical environments as part of their own products and pipelines.
Learn more
Echo Live
Echo Live brings live voice transformation and creative audio tools to Windows and Mac. Use it to experiment with a new voice in a gaming session, add reactions to Discord conversations or prepare sounds for your next stream. AI voice changing, a customizable effects workspace, a soundboard and text-to-speech are available in the same desktop app.
For live voices, explore AI models and route the processed microphone signal into your chat or streaming setup. Processing happens on your computer, with results and responsiveness depending on the chosen model, processing mode and hardware. The guided setup walks you through microphone configuration and your first voice.
Voice Lab is for building your own sound. Combine pitch and formant adjustments with harmonizer, reverb and delay in custom effect chains. For quick audio cues, import clips into soundboards and assign keyboard shortcuts so reactions and recurring stream sounds are close at hand.
The text-to-speech workspace turns typed scripts into spoken clips with downloadable speech models. Replay results from history or explore optional RVC processing for character voices. The languages a speech model can generate depend on that model, not on the language selected for the menus.
The interface can be set to English, German, Spanish, French, Portuguese, Japanese, Korean, Traditional Chinese, Simplified Chinese, Polish, Russian, Italian or Turkish. Some text uses an English fallback.
Echo Live is available for Windows 10/11 x64 and Apple Silicon Mac. Download it at voicechanger.live and create a free Echo account to get started. A free version is available, with optional Pro features for users who want to upgrade.
Learn more
Spokenly
Spokenly is an innovative dictation application powered by AI, available for Mac, iPhone, Windows, and Linux, designed to convert spoken words into clear, punctuated text in any working environment. By simply holding a shortcut, users can speak naturally and then release to insert the transcription directly at the cursor across various platforms including browsers, email, chat applications, word processors, IDEs, terminals, and more. This versatile app accommodates over 100 languages, supporting mixed-language dictation, and provides both local and cloud-based speech-to-text models. Users can utilize on-device models like Whisper and Parakeet for offline operation, while cloud services from companies such as OpenAI, Deepgram, Groq, Soniox, and ElevenLabs can be accessed for enhanced accuracy or real-time transcription needs. Additionally, the Local Only Mode ensures that voice data remains solely on the device, preventing any network interactions. The application features modes that allow users to save different transcription models, select AI providers, set prompts, and choose output styles tailored for specific tasks. Furthermore, the AI Instructions feature enables users to eliminate filler words, correct grammar and punctuation, summarize, rewrite, translate, or reformat the dictated text, enhancing the overall functionality and user experience of the app. With its extensive capabilities, Spokenly stands out as a comprehensive solution for anyone looking to streamline their dictation process.
Learn more