Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
The Gemini Live API is an advanced preview feature designed to facilitate low-latency, bidirectional interactions through voice and video with the Gemini system. This innovation allows users to engage in conversations that feel natural and human-like, while also enabling them to interrupt the model's responses via voice commands. In addition to handling text inputs, the model is capable of processing audio and video, yielding both text and audio outputs. Recent enhancements include the introduction of two new voice options and support for 30 additional languages, along with the ability to configure the output language as needed. Furthermore, users can adjust image resolution settings (66/256 tokens), decide on turn coverage (whether to send all inputs continuously or only during user speech), and customize interruption preferences. Additional features encompass voice activity detection, new client events for signaling the end of a turn, token count tracking, and a client event for marking the end of the stream. The system also supports text streaming, along with configurable session resumption that retains session data on the server for up to 24 hours, and the capability for extended sessions utilizing a sliding context window for better conversation continuity. Overall, Gemini Live API enhances interaction quality, making it more versatile and user-friendly.
Description
ai|coustics is a platform powered by AI technology that aims to enhance both audio and video recordings by improving speech intelligibility and removing unwanted background noise. The platform features an intuitive web application that allows users to upload their files for enhancement, along with an API and SDK that enable developers to incorporate real-time audio processing into their own software and hardware solutions. Two main AI models drive its functionality: Finch, which excels in noise reduction, and Lark, which recovers lost frequencies and adds richness for a studio-quality listening experience. Supporting more than 40 file formats such as MP3, MP4, WAV, and MOV, ai|coustics also offers batch processing options to streamline workflow. With a user base exceeding 500,000, including prominent organizations such as BosePark, Bayerischer Rundfunk, and Sieve, ai|coustics serves a diverse range of clients. The platform is especially advantageous for podcasters, content creators, educators, and developers aiming to provide superior audio quality across multiple channels. Furthermore, its versatility makes it an essential tool for anyone looking to elevate their audio production standards.
API Access
Has API
Yes
API Access
Has API
No
Integrations
Daily
Yes
LiveKit
Yes
Agora
Yes
C++
No
Firebase
Yes
Fishjam
Yes
Gemini
Yes
Gemini 3 Pro Image
Yes
Gemini 3.1 Flash Image
Yes
Gemini 3.1 Flash Live
Yes
Integrations
Daily
Yes
LiveKit
Yes
Agora
No
C++
Yes
Firebase
No
Fishjam
No
Gemini
No
Gemini 3 Pro Image
No
Gemini 3.1 Flash Image
No
Gemini 3.1 Flash Live
No
Pricing Details
No price information available.
Free Trial
No
Free Version
No
Pricing Details
$149 / month
There are three tiers of monthly subscriptions with increasing minutes allowance, and a custom enterprise option.
Free Trial
Yes
Free Version
No
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Deployment
Web-Based
No
On-Premises
Yes
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Customer Support
Business Hours
Yes
Live Rep (24/7)
No
Online Support
Yes
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Types of Training
Training Docs
Yes
Webinars
Yes
Live Training (Online)
Yes
In Person
Yes
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
Yes
In Person
No
Vendor Details
Company Name
Founded
1998
Country
United States
Website
ai.google.dev/gemini-api/docs/live
Vendor Details
Company Name
ai-coustics
Founded
2021
Country
Germany
Website
ai-coustics.com