Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
An EU-based company offers an inference API compatible with OpenAI and Anthropic models. Their premier model operates on dedicated GPUs located in EIA data centres and ensures that no data is retained, as all prompts and completions are processed solely in memory—meaning they are neither stored nor logged, and are not utilized for training purposes. Additionally, users have access to routed open models from various third-party providers using the same key, which are also clearly marked. The service includes a Data Processing Agreement (DPA) and an invoice from the EU entity. Notable features include streaming capabilities, tool calling, structured output, a publicly available DPA and sub-processor list, as well as a pricing model based on token usage. During a measurement conducted on the live system in August 2026, the service demonstrated a capacity of processing 176 tokens per second per stream, with the first token being generated in just 0.3 seconds, highlighting its efficiency and speed. Such performance metrics are critical for developers seeking reliable and rapid AI solutions in their applications.
Description
WebLLM serves as a robust inference engine for language models that operates directly in web browsers, utilizing WebGPU technology to provide hardware acceleration for efficient LLM tasks without needing server support. This platform is fully compatible with the OpenAI API, which allows for smooth incorporation of features such as JSON mode, function-calling capabilities, and streaming functionalities. With native support for a variety of models, including Llama, Phi, Gemma, RedPajama, Mistral, and Qwen, WebLLM proves to be adaptable for a wide range of artificial intelligence applications. Users can easily upload and implement custom models in MLC format, tailoring WebLLM to fit particular requirements and use cases. The integration process is made simple through package managers like NPM and Yarn or via CDN, and it is enhanced by a wealth of examples and a modular architecture that allows for seamless connections with user interface elements. Additionally, the platform's ability to support streaming chat completions facilitates immediate output generation, making it ideal for dynamic applications such as chatbots and virtual assistants, further enriching user interaction. This versatility opens up new possibilities for developers looking to enhance their web applications with advanced AI capabilities.
API Access
Has API
No
API Access
Has API
Yes
Screenshots View All
No images available
Integrations
Alpaca
No
Codestral
No
Codestral Mamba
No
Dolly
No
Gemma
No
JSON
No
Llama 2
No
Llama 3
No
Llama 3.1
No
Llama 3.2
No
Integrations
Alpaca
Yes
Codestral
Yes
Codestral Mamba
Yes
Dolly
Yes
Gemma
Yes
JSON
Yes
Llama 2
Yes
Llama 3
Yes
Llama 3.1
Yes
Llama 3.2
Yes
Pricing Details
$0.04 per 1M input tokens
Free Trial
No
Free Version
Yes
Pricing Details
Free
Free Trial
No
Free Version
Yes
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
No
In Person
No
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
No
In Person
No
Vendor Details
Company Name
Heabsy
Founded
2014
Country
Slovakia
Website
heabsy.com
Vendor Details
Company Name
WebLLM
Website
webllm.mlc.ai/