An API powered by Google's AI technology allows you to accurately convert speech into text. You can accurately caption your content, provide a better user experience with products using voice commands, and gain insight from customer interactions to improve your service. Google's deep learning neural network algorithms are the most advanced in automatic speech recognition (ASR). Speech-to-Text allows for experimentation, creation, management, and customization of custom resources. You can deploy speech recognition wherever you need it, whether it's in the cloud using the API or on-premises using Speech-to-Text O-Prem. You can customize speech recognition to translate domain-specific terms or rare words. Automated conversion of spoken numbers into addresses, years and currencies. Our user interface makes it easy to experiment with your speech audio.
Learn more

Most AI video tools hand you a black box: closed weights, a subscription, and no way to see what is happening under the hood. LTX takes the opposite approach. Built by Lightricks, LTX is an open foundation model that generates and simulates across video, audio, and the physical world, and it puts the weights, the code, and the control in your hands.
At the center of the model is LTX-2.5, a 22B-parameter dual-stream diffusion transformer that produces native 4K video at up to 50 frames per second, with audio and video generated together in a single pass rather than stitched together afterward. Artificial Analysis, an independent benchmarking group, currently ranks LTX among the top three AI video models in the world.
You choose how you want to use it. Download the open weights and run LTX-2.5 on your own hardware. License the model for on-premise deployment backed by enterprise support. Or build directly on LTX Studio, the production suite that turns the model into a full creative workflow. Companies like ElevenLabs, Asteria Film Co., Magnopus, and NVIDIA already rely on LTX for their own work.
LTX is not built for one-off social clips. It is infrastructure for teams that generate motion, audio, and physical environments as part of their own products and pipelines.
Learn more
Echo Live
Echo Live brings live voice transformation and creative audio tools to Windows and Mac. Use it to experiment with a new voice in a gaming session, add reactions to Discord conversations or prepare sounds for your next stream. AI voice changing, a customizable effects workspace, a soundboard and text-to-speech are available in the same desktop app.
For live voices, explore AI models and route the processed microphone signal into your chat or streaming setup. Processing happens on your computer, with results and responsiveness depending on the chosen model, processing mode and hardware. The guided setup walks you through microphone configuration and your first voice.
Voice Lab is for building your own sound. Combine pitch and formant adjustments with harmonizer, reverb and delay in custom effect chains. For quick audio cues, import clips into soundboards and assign keyboard shortcuts so reactions and recurring stream sounds are close at hand.
The text-to-speech workspace turns typed scripts into spoken clips with downloadable speech models. Replay results from history or explore optional RVC processing for character voices. The languages a speech model can generate depend on that model, not on the language selected for the menus.
The interface can be set to English, German, Spanish, French, Portuguese, Japanese, Korean, Traditional Chinese, Simplified Chinese, Polish, Russian, Italian or Turkish. Some text uses an English fallback.
Echo Live is available for Windows 10/11 x64 and Apple Silicon Mac. Download it at voicechanger.live and create a free Echo account to get started. A free version is available, with optional Pro features for users who want to upgrade.
Learn more
Voxtral
Voxtral models represent cutting-edge open-source systems designed for speech understanding, available in two sizes: a larger 24 B variant aimed at production-scale use and a smaller 3 B variant suitable for local and edge applications, both of which are provided under the Apache 2.0 license. These models excel in delivering precise transcription while featuring inherent semantic comprehension, accommodating long-form contexts of up to 32 K tokens and incorporating built-in question-and-answer capabilities along with structured summarization. They automatically detect languages across a range of major tongues and enable direct function-calling to activate backend workflows through voice commands. Retaining the textual strengths of their Mistral Small 3.1 architecture, Voxtral can process audio inputs of up to 30 minutes for transcription tasks and up to 40 minutes for comprehension, consistently surpassing both open-source and proprietary competitors in benchmarks like LibriSpeech, Mozilla Common Voice, and FLEURS. Users can access Voxtral through downloads on Hugging Face, API endpoints, or by utilizing private on-premises deployments, and the model also provides options for domain-specific fine-tuning along with advanced features tailored for enterprise needs, thus enhancing its applicability across various sectors.
Learn more