|
|
Run in Google Colab
|
|
|
View source on GitHub
|
EmbeddingGemma 2 is a lightweight, open-source 740M parameter model designed for unified text and multimodal embeddings. Built on the Gemma 4 decoder architecture, it maps inputs into a unified 768-dimensional vector space.
Because of its compact footprint, you can run demanding semantic workflows (e.g. Retrieval-Augmented Generation (RAG), semantic search, and zero-shot classification) locally and offline on consumer GPUs or CPUs.
What this guide covers:
- Installing dependencies and loading EmbeddingGemma 2
- Generating text embeddings, computing semantic similarity, and applying task-specific prompts
- Building downstream applications (RAG and zero-shot classification)
- Optimizing storage and inference speed with Matryoshka Representation Learning (MRL)
Looking for image, audio, or video embeddings? Check out Multimodal Embeddings with EmbeddingGemma 2 guide.
Setup and Installation
Install sentence-transformers and transformers.
Sentence Transformers is a Python framework for using and training embedding and reranker models. For more details on the framework, see the Sentence Transformers documentation.
pip install -U sentence-transformers transformersModel Initialization
Instantiate EmbeddingGemma 2 directly using the SentenceTransformer class.
from sentence_transformers import SentenceTransformer
model_id = "google/embeddinggemma-2"
model = SentenceTransformer(model_id)
print(f"device: {model.device}")
print(f"dtype: {model.dtype}")
print(model)
print("Total number of parameters in the model:", sum(p.numel() for p in model.parameters()))
Loading weights: 0%| | 0/1376 [00:00<?, ?it/s]
device: cuda:0
dtype: torch.bfloat16
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}, 'image': {'method': 'forward', 'method_output_name': 'last_hidden_state'}, 'audio': {'method': 'forward', 'method_output_name': 'last_hidden_state'}, 'video': {'method': 'forward', 'method_output_name': 'last_hidden_state'}, 'message': {'method': 'forward', 'method_output_name': 'last_hidden_state', 'format': 'structured'} }, 'module_output_name': 'token_embeddings', 'architecture': 'EmbeddingGemma2Model'})
(1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
(2): Normalize({'module_input_name': 'sentence_embedding', 'module_output_name': 'sentence_embedding'})
)
Total number of parameters in the model: 744371488
Tip (Text-Only Memory Optimization): By default,
SentenceTransformerloads all modality encoders (740M parameters). If your workflow only uses text, you can passconfig_kwargs={"vision_config": None, "audio_config": None}to skip loading the vision and audio encoders (270M parameters).
Note on Precision: EmbeddingGemma 2 activations are incompatible with
float16. Usebfloat16(default on supported GPUs) orfloat32.
Text Embeddings and Task Prompts
Learn how to encode text into dense vector representations, measure semantic similarity between inputs, and apply task-specific prompt prefixes to optimize embedding quality.
Generating Text Embeddings
An embedding is a dense numerical vector that captures the semantic meaning of an input. Let's encode three words—["king", "queen", "car"]—to inspect their 768-dimensional representations.
words = ["king", "queen", "car"]
# Generate embeddings using model.encode()
embeddings = model.encode(words)
print(embeddings)
for idx, embedding in enumerate(embeddings):
print(f"Embedding {idx+1} (shape): {embedding.shape}")
[[ 0.03222656 0.00285339 0.02160645 ... 0.012146 0.03149414 -0.0067749 ] [ 0.00787354 0.01361084 0.02502441 ... -0.01269531 0.02612305 0.01361084] [-0.04052734 -0.00860596 0.04492188 ... 0.00964355 -0.01501465 -0.00726318]] Embedding 1 (shape): (768,) Embedding 2 (shape): (768,) Embedding 3 (shape): (768,)
The key property of embeddings is the distance between vectors in semantic space. Because "king" and "queen" share semantic characteristics as royal titles, their vectors are positioned close together, whereas "car" lies further away.
Determining Semantic Similarity
We can quantify how semantically similar two inputs are by computing the cosine similarity between their embeddings using model.similarity().
Below are three sentence pairs demonstrating high, medium, and low semantic similarity:
- High Similarity:
- Sentence A:
"The chef prepared a delicious meal for the guests." - Sentence B:
"A tasty dinner was cooked by the chef for the visitors." - Reasoning: Both sentences describe the exact same event using active vs. passive voice.
- Sentence A:
- Medium Similarity:
- Sentence A:
"She is an expert in machine learning." - Sentence B:
"He has a deep interest in artificial intelligence." - Reasoning: Related domains (ML is a subfield of AI), but different subjects and levels of expertise.
- Sentence A:
- Low Similarity:
- Sentence A:
"The weather in Tokyo is sunny today." - Sentence B:
"I need to buy groceries for the week." - Reasoning: Completely unrelated topics with no semantic overlap.
- Sentence A:
sentence_high = [
"The chef prepared a delicious meal for the guests.",
"A tasty dinner was cooked by the chef for the visitors.",
]
sentence_medium = [
"She is an expert in machine learning.",
"He has a deep interest in artificial intelligence.",
]
sentence_low = [
"The weather in Tokyo is sunny today.",
"I need to buy groceries for the week.",
]
for pair in [sentence_high, sentence_medium, sentence_low]:
print("🙋♂️", pair)
pair_embeddings = model.encode(pair)
score = model.similarity(pair_embeddings[0], pair_embeddings[1]).item()
print(f" `-> 🤖 similarity score: {score:.4f}")
🙋♂️ ['The chef prepared a delicious meal for the guests.', 'A tasty dinner was cooked by the chef for the visitors.'] `-> 🤖 similarity score: 0.9184 🙋♂️ ['She is an expert in machine learning.', 'He has a deep interest in artificial intelligence.'] `-> 🤖 similarity score: 0.7835 🙋♂️ ['The weather in Tokyo is sunny today.', 'I need to buy groceries for the week.'] `-> 🤖 similarity score: 0.6880
Note on interpreting scores: Cosine similarity values for text embeddings tend to skew higher than human intuition might suggest, often rarely dropping below 0.5 in practice. For instance, even completely unrelated sentences (like the "Low Similarity" example below) can yield a similarity score of roughly 0.6880. Keep this higher baseline in mind when evaluating semantic overlap.
Optimizing Embeddings with Task Prompts
To achieve optimal performance with EmbeddingGemma 2 on text tasks, prepend an instructional prompt (task prefix) to your input. Task prompts condition the model for specific objectives (e.g. sentence similarity, asymmetric search retrieval, or classification) and help separate queries from documents.
Prefix Notation & Usage
We offer two types of task prefixes, depending on how embeddings are used in the task. There are two types of tasks:
- Asymmetric Tasks (e.g. retrieval): Use a query prefix for queries and a document prefix for corpus items.
- Symmetric Tasks (e.g. classification, similarity): Apply the same task prefix to all inputs being compared.
| Use Case | Task Type | Query Task Instruction | Document Task Instruction (use none if no title) |
|---|---|---|---|
| Web / document search | Asymmetric | task: search result | query: {query} |
title: {title} | text: {content} |
| Question answering | Asymmetric | task: question answering | query: {question} |
title: {title} | text: {passage} |
| Fact checking | Asymmetric | task: fact checking | query: {claim} |
title: {title} | text: {evidence} |
| Code search | Asymmetric | task: code retrieval | query: {query} |
title: {title or filename} | text: {code} |
| Text classification | Symmetric | task: classification | query: {content} |
N/A |
| Clustering | Symmetric | task: clustering | query: {content} |
N/A |
| Measuring similarity | Symmetric | task: sentence similarity | query: {content} |
N/A |
Please note: Prefixes are used on text inputs only; when encoding images, video, or audio, pass them directly.
You can apply prompts during inference in three ways:
Using
prompt_name(Recommended): Select a built-in task prompt registered inmodel.prompts:embeddings = model.encode(sentences, prompt_name="STS")Using
prompt: Pass a custom prefix string directly tomodel.encode():embeddings = model.encode(sentences, prompt="task: sentence similarity | query: ")Using the Default Prompt: If
default_prompt_nameis configured on the model, callingmodel.encode(sentences)applies it automatically.
Let's inspect all built-in task prompts and re-evaluate our sentence similarity pairs using prompt_name="STS":
print("Available task prompts:")
for name, prefix in model.prompts.items():
print(f" {name}: {prefix!r}")
print("-" * 80)
for pair in [sentence_high, sentence_medium, sentence_low]:
print("🙋♂️", pair)
pair_embeddings = model.encode(pair, prompt_name="STS")
score = model.similarity(pair_embeddings[0], pair_embeddings[1]).item()
print(f" `-> 🤖 STS score: {score:.4f}")
Available task prompts: query: 'task: search result | query: ' document: 'title: none | text: ' BitextMining: 'task: search result | query: ' Classification: 'task: classification | query: ' Clustering: 'task: clustering | query: ' CodeRetrieval: 'task: code retrieval | query: ' Document: 'title: none | text: ' FactChecking: 'task: fact checking | query: ' InstructionRetrieval: 'task: code retrieval | query: ' MultilabelClassification: 'task: classification | query: ' PairClassification: 'task: sentence similarity | query: ' QuestionAnswering: 'task: question answering | query: ' Reranking: 'task: search result | query: ' Retrieval: 'task: search result | query: ' Retrieval-document: 'title: none | text: ' Retrieval-query: 'task: search result | query: ' STS: 'task: sentence similarity | query: ' SearchQuery: 'task: search result | query: ' SentenceSimilarity: 'task: sentence similarity | query: ' Summarization: 'task: sentence similarity | query: ' -------------------------------------------------------------------------------- 🙋♂️ ['The chef prepared a delicious meal for the guests.', 'A tasty dinner was cooked by the chef for the visitors.'] `-> 🤖 STS score: 0.9555 🙋♂️ ['She is an expert in machine learning.', 'He has a deep interest in artificial intelligence.'] `-> 🤖 STS score: 0.8628 🙋♂️ ['The weather in Tokyo is sunny today.', 'I need to buy groceries for the week.'] `-> 🤖 STS score: 0.7340
Downstream Applications
Put EmbeddingGemma 2 embeddings into practice to power common real-world workflows, such as document retrieval for Retrieval-Augmented Generation (RAG) and zero-shot text classification.
Retrieval-Augmented Generation (RAG)
In asymmetric retrieval and RAG pipelines, queries are typically short questions while documents are longer passages. Use dedicated retrieval prompts to optimize both sides of the search index:
- Search Queries: Use
prompt_name="Retrieval-query"(applies"task: search result | query: "). - Documents: Use
prompt_name="Retrieval-document"(applies"title: none | text: "), or pass a custompromptstring to include document titles ("title: <Document Title> | text: "). Including titles significantly improves retrieval accuracy.
query = "How do I reduce memory usage when running EmbeddingGemma 2?"
documents = [
{
"title": "Model Initialization and Memory Optimization",
"text": "You can disable the vision or audio encoders by passing vision_config=None or audio_config=None inside config_kwargs.",
},
{
"title": "Matryoshka Representation Learning",
"text": "MRL allows truncating 768-dimensional embeddings to 512 or 256 dimensions for faster vector search and smaller indexes.",
},
{
"title": "Audio Processing Specifications",
"text": "Audio waveforms are automatically resampled to 16 kHz before feature extraction.",
},
]
# Encode query with Retrieval-query prompt
query_emb = model.encode(query, prompt_name="Retrieval-query")
# Encode documents with their titles embedded in the prompt
doc_embeddings = [
model.encode(doc["text"], prompt=f"title: {doc['title']} | text: ")
for doc in documents
]
scores = model.similarity(query_emb, doc_embeddings)[0]
for doc, score in zip(documents, scores):
print(f"[Score: {score:.4f}] {doc['title']}")
[Score: 0.6855] Model Initialization and Memory Optimization [Score: 0.6386] Matryoshka Representation Learning [Score: 0.5653] Audio Processing Specifications
For a complete end-to-end RAG tutorial, see the RAG with EmbeddingGemma notebook in the Gemma Cookbook.
Zero-Shot Classification
Zero-shot text classification assigns inputs to predefined categories without task-specific training by selecting the candidate label with the highest embedding similarity to the input text.
When choosing a task prompt for zero-shot classification, consider whether your target labels match standard classification datasets:
- Custom or Arbitrary Categories (
prompt_name="Clustering"): For domain-specific or ad-hoc labels like"Billing Issue","Technical Support", and"Sales Inquiry", the"Classification"prompt does not work well because the training dataset does not include those specific labels. Instead, useprompt_name="Clustering"to group inputs and candidate labels by semantic topic similarity. - Standard Classification Tasks (
prompt_name="Classification"): Common classification prompts—such as sentiment analysis ("Positive","Neutral","Negative")—work well withprompt_name="Classification"because they align with the classification tasks seen during training.
Custom Category Routing with Clustering
Use prompt_name="Clustering" to route incoming support tickets to custom department categories:
labels = ["Billing Issue", "Technical Support", "Sales Inquiry"]
tickets = [
"Excuse me, the app freezes on the login screen. It won't work even when I try to reset my password.",
"I would like to inquire about your enterprise plan pricing and features for a team of 50 people.",
]
# Encode labels and incoming tickets using the Clustering prompt
label_embeddings = model.encode(labels, prompt_name="Clustering")
ticket_embeddings = model.encode(tickets, prompt_name="Clustering")
# Compute similarity matrix and find the best-matching label for each ticket
similarities = model.similarity(ticket_embeddings, label_embeddings)
print("Similarity matrix:\n", similarities)
predicted_indices = similarities.argmax(dim=1)
for ticket, pred_idx in zip(tickets, predicted_indices):
print(f"\n🙋♂️ Ticket: {ticket}")
print(f" `-> 🤖 Predicted Category: {labels[pred_idx]}")
Similarity matrix:
tensor([[0.8704, 0.8904, 0.8394],
[0.8426, 0.8566, 0.8758]])
🙋♂️ Ticket: Excuse me, the app freezes on the login screen. It won't work even when I try to reset my password.
`-> 🤖 Predicted Category: Technical Support
🙋♂️ Ticket: I would like to inquire about your enterprise plan pricing and features for a team of 50 people.
`-> 🤖 Predicted Category: Sales Inquiry
Sentiment Analysis with Classification
For standard sentiment analysis, encode the common sentiment labels and input sentences using prompt_name="Classification":
labels = ["Positive", "Neutral", "Negative"]
sentences = [
# Positive
"The sunset over the ridge was absolutely breathtaking.",
"This is hands down the best coffee I have ever tasted.",
# Neutral
"The package arrived on Tuesday afternoon.",
"The meeting is scheduled for 10:00 AM in Conference Room B.",
# Negative
"The battery dies after barely two hours of use.",
"It felt like an utter waste of time and money.",
]
# Encode labels and incoming sentences using the Classification prompt
label_embeddings = model.encode(labels, prompt_name="Classification")
sentence_embeddings = model.encode(sentences, prompt_name="Classification")
# Compute similarity matrix and find the best-matching label for each sentence
similarities = model.similarity(sentence_embeddings, label_embeddings)
print("Similarity matrix:\n", similarities)
predicted_indices = similarities.argmax(dim=1)
for sentence, pred_idx in zip(sentences, predicted_indices):
print(f"\n🙋♂️ Sentence: {sentence}")
print(f" `-> 🤖 Predicted Sentiment: {labels[pred_idx]}")
Similarity matrix:
tensor([[0.8177, 0.6642, 0.6423],
[0.8731, 0.6198, 0.6375],
[0.6581, 0.8198, 0.6425],
[0.6115, 0.7675, 0.6093],
[0.6273, 0.6680, 0.7666],
[0.6103, 0.6236, 0.8660]])
🙋♂️ Sentence: The sunset over the ridge was absolutely breathtaking.
`-> 🤖 Predicted Sentiment: Positive
🙋♂️ Sentence: This is hands down the best coffee I have ever tasted.
`-> 🤖 Predicted Sentiment: Positive
🙋♂️ Sentence: The package arrived on Tuesday afternoon.
`-> 🤖 Predicted Sentiment: Neutral
🙋♂️ Sentence: The meeting is scheduled for 10:00 AM in Conference Room B.
`-> 🤖 Predicted Sentiment: Neutral
🙋♂️ Sentence: The battery dies after barely two hours of use.
`-> 🤖 Predicted Sentiment: Negative
🙋♂️ Sentence: It felt like an utter waste of time and money.
`-> 🤖 Predicted Sentiment: Negative
Matryoshka Representation Learning (MRL)
EmbeddingGemma 2 is trained with Matryoshka Representation Learning (MRL), which concentrates the most important semantic information in the earliest dimensions of the output vector.
Instead of deploying separate models for different latency or storage budgets, you can truncate the full 768-dimensional embedding to smaller prefixes—such as 512, 256 or 128 dimensions—while preserving strong semantic quality. When truncating dimensions, always set normalize_embeddings=True so that the truncated vectors are re-normalized to unit length.
Let's compare word similarity scores across 768, 512, and 128 dimensions:
words = ["king", "queen", "car"]
def evaluate_mrl(embeddings, sim_fn_name):
print(f"Embedding shape: {embeddings.shape} | Similarity function: {sim_fn_name}")
sims = model.similarity(embeddings[0], embeddings[1:])
for idx, word in enumerate(words[1:]):
print(f" 🙋♂️ 'king' vs. '{word}' -> 🤖 score: {sims[0][idx].item():.4f}")
# 1. Full 768-dimensional embeddings
emb_768 = model.encode(words, prompt_name="STS", normalize_embeddings=True)
evaluate_mrl(emb_768, model.similarity_fn_name)
Embedding shape: (3, 768) | Similarity function: cosine 🙋♂️ 'king' vs. 'queen' -> 🤖 score: 0.9128 🙋♂️ 'king' vs. 'car' -> 🤖 score: 0.8354
Truncating to 512 Dimensions
Pass truncate_dim=512 directly to model.encode() to slice and re-normalize vectors on the fly:
# 2. Truncate to 512 dimensions at inference time
emb_512 = model.encode(
words,
prompt_name="STS",
truncate_dim=512,
normalize_embeddings=True,
)
evaluate_mrl(emb_512, model.similarity_fn_name)
Embedding shape: (3, 512) | Similarity function: cosine 🙋♂️ 'king' vs. 'queen' -> 🤖 score: 0.9124 🙋♂️ 'king' vs. 'car' -> 🤖 score: 0.8334
Truncating to 128 Dimensions with Dot-Product Similarity
For maximum speed in high-throughput vector databases, you can configure truncate_dim=128 and similarity_fn_name="dot" at model initialization. Because normalized embeddings have unit length, dot product is mathematically equivalent to cosine similarity while being faster to compute:
# 3. Initialize model with default 128-dim truncation and dot-product similarity
fast_model = SentenceTransformer(
model_id,
truncate_dim=128,
similarity_fn_name="dot",
device=model.device,
)
emb_128 = fast_model.encode(words, prompt_name="STS", normalize_embeddings=True)
evaluate_mrl(emb_128, fast_model.similarity_fn_name)
Loading weights: 0%| | 0/1376 [00:00<?, ?it/s] Embedding shape: (3, 128) | Similarity function: dot 🙋♂️ 'king' vs. 'queen' -> 🤖 score: 0.9470 🙋♂️ 'king' vs. 'car' -> 🤖 score: 0.8957
Summary and Next Steps
You now know how to generate text embeddings using EmbeddingGemma 2 and the sentence-transformers library. You also explored how to apply task-specific prompts, build zero-shot classifiers and RAG retrievers, and compress embeddings using Matryoshka Representation Learning (MRL).
Run in Google Colab
View source on GitHub