Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
EmbeddingGemma 2 is a versatile and lightweight multimodal embedding model that facilitates the mapping of text, code, images, video, and audio into a unified embedding space, which is useful for applications such as search, retrieval, classification, routing, and RAG. It is constructed on the Gemma 4 framework and distributed under the Apache 2.0 license, featuring an impressive 740 million parameters while being fine-tuned for efficient on-device inference. Its flexible architecture allows for the use of only 270 million parameters for text-centric tasks, and it includes additional vision and audio encoders for comprehensive multimodal capabilities. Furthermore, the innovative Matryoshka Representation Learning technique enables developers to compress output vectors from 768 dimensions down to 512, 256, or even 128 dimensions, effectively reducing the storage and memory demands for local vector databases. The model is equipped with an 8K-token context window, providing the capability to handle up to 5.5 minutes of audio, 29 images, 58 video frames, or various combinations of these inputs seamlessly on local hardware. This adaptability makes it particularly valuable for developers seeking to enhance their applications with rich multimedia integration.
Description
On June 23, 2025, Microsoft unveiled Mu, an innovative 330-million-parameter encoder–decoder language model specifically crafted to enhance the agent experience within Windows environments by effectively translating natural language inquiries into function calls for Settings, all processed on-device via NPUs at a remarkable speed of over 100 tokens per second while ensuring impressive accuracy. By leveraging Phi Silica optimizations, Mu’s encoder–decoder design employs a fixed-length latent representation that significantly reduces both computational demands and memory usage, achieving a 47 percent reduction in first-token latency and a decoding speed that is 4.7 times greater on Qualcomm Hexagon NPUs when compared to other decoder-only models. Additionally, the model benefits from hardware-aware tuning techniques, which include a thoughtful 2/3–1/3 split of encoder and decoder parameters, shared weights for input and output embeddings, Dual LayerNorm, rotary positional embeddings, and grouped-query attention, allowing for swift inference rates exceeding 200 tokens per second on devices such as the Surface Laptop 7, along with sub-500 ms response times for settings-related queries. This combination of features positions Mu as a groundbreaking advancement in on-device language processing capabilities.
API Access
Has API
Yes
API Access
Has API
No
Integrations
No details available.
Integrations
No details available.
Pricing Details
No price information available.
Free Trial
No
Free Version
No
Pricing Details
No price information available.
Free Trial
No
Free Version
No
Deployment
Web-Based
Yes
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
No
Mac
No
Linux
No
Chromebook
No
Deployment
Web-Based
No
On-Premises
No
iPhone App
No
iPad App
No
Android App
No
Windows
Yes
Mac
No
Linux
No
Chromebook
No
Customer Support
Business Hours
No
Live Rep (24/7)
No
Online Support
Yes
Customer Support
Business Hours
Yes
Live Rep (24/7)
No
Online Support
Yes
Types of Training
Training Docs
Yes
Webinars
No
Live Training (Online)
No
In Person
No
Types of Training
Training Docs
Yes
Webinars
Yes
Live Training (Online)
Yes
In Person
Yes
Vendor Details
Company Name
Founded
1998
Country
United States
Website
blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/
Vendor Details
Company Name
Microsoft
Founded
1975
Country
United States
Website
blogs.windows.com/windowsexperience/2025/06/23/introducing-mu-language-model-and-how-it-enabled-the-agent-in-windows-settings/
Product Features
Product Features
Alternatives
No Alternatives