EmbeddingGemma 2 launched on October 6, 2026 under Apache 2.0. It is a sub-1B model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space.
This article covers the architecture, the benchmarks, and runnable scripts to provide measured results.
| Specification | EmbeddingGemma 2 |
| Base model | Gemma 4 |
| License | Apache 2.0 |
| Output dimension | 768, truncatable to 512, 256, 128 |
| Context window | 8,192 tokens, shared across modalities |
| Modalities | Text, code, images, video, audio |
| Parameters | 270M text-only to 740M full multimodal |
| Released | October 6, 2026 |
The usual way to search across mixed content is to run one model per content type and stitch the results together afterwards. That means separate indexes, separate score ranges, and no reliable way to compare a photo against a paragraph.
EmbeddingGemma 2 removes that problem by sending every content type through one backbone. Each modality gets its own front-end encoder, but the output always lands in the same 768-dimensional space.

| Encoder | Size | Handles | Built on |
|---|---|---|---|
| Text and code | 270M, the base | Plain text and source code, up to 8,192 tokens | An adapted Gemma 4 decoder |
| Vision | +170M | Photos, charts, slides, PDF pages, and video frames | A dedicated vision encoder |
| Audio | +300M | Speech and ambient sound, fed in as raw audio | A dedicated audio encoder |
The text encoder is the base and is always present. Vision and audio are additions on top of it, which is where the four parameter counts come from.

One checkpoint, four ways to load it.
You are not choosing between four different models. There is one set of weights, and you decide at load time which encoders get read into memory. Skipped encoders cost nothing, on disk or in RAM.
Because the output space is identical in all four cases, a query embedded with the 270M text setup can be matched against documents embedded with the full 740M model. Two practical consequences:
If you run Gemma 4 as the generator in an on-device RAG stack, the two models share a text tokenizer and the same audio encoder design. That overlap is loaded once rather than twice, which matters when the constraint is phone memory rather than server memory.
| Benchmark | Result |
| MTEB (Code) | 14% higher than EmbeddingGemma 1 |
| Multilingual text | Retains EmbeddingGemma 1 accuracy |
| Image, video, audio retrieval | New capability, not present in version 1 |
Google published the MTEB (Code) gain as a percentage improvement rather than absolute scores, so there is no single number to compare against other models directly.
| Dimension | Text and code | Image, video, speech | Storage per 1M vectors |
| 768d | baseline | baseline | 1,465 MB |
| 512d | no stated loss | no stated loss | 977 MB |
| 256d | near baseline | roughly 95% | 488 MB |
| 128d | roughly 90% | roughly 75% | 244 MB |
Two things to read off that table. Storage figures assume bfloat16, and the media columns fall away faster than text does. Google flags 128d multimodal as the case to test before you ship it, and the 75% figure explains why.
Results from the five scripts below, on a small test corpus. These are measurements on one corpus, not benchmark scores.
| Measurement | Result |
| 256d versus 768d ranking | Identical top-1 and MRR |
| 128d ranking | Top-1 held, MRR fell from 0.667 to 0.656 |
| Mean similarity at 128d | Rose 12.6% versus 768d despite identical ranking |
| Image cross-modal margin | 0.1154 over the nearest wrong query |
| Audio cross-modal margin | 0.0137 over the nearest wrong query |
| Default output precision | float32, not bfloat16, so indexes are double the quoted size |
First install sentence transformer and other relvant libraries using:
pip install -U sentence-transformers[image,audio,video] transformers
Requires sentence-transformers 6.1.0 or later.
Reports Python version, library version, available modalities and hardware before you download 740M parameters.
File: 00_check_setup.py
import sys
import importlib
import shutil
def ok(b):
return "OK " if b else "-- "
print("=" * 66)
print("EmbeddingGemma 2 setup check")
print("=" * 66)
print(f"\nPython {sys.version.split()[0]} (3.9+ needed)")
# --- core library -------------------------------------------------
try:
import sentence_transformers as st
ver = st.__version__
major, minor = (int(x) for x in ver.split(".")[:2])
good = (major, minor) >= (6, 1)
print(f"{ok(good)}sentence-transformers {ver} (need 6.1.0+)")
if not good:
print(" pip install -U sentence-transformers")
except ImportError:
print("-- sentence-transformers NOT INSTALLED")
print(
" pip install -U sentence-transformers[image,audio,video] "
"transformers"
)
sys.exit(1)
# --- which modalities can actually run ---------------------------
print("\nModality support:")
mods = {
"text / code": [], # always available
"images": ["PIL"],
"audio": ["soundfile", "librosa"],
"video": ["decord"],
}
for name, deps in mods.items():
missing = [
d for d in deps
if importlib.util.find_spec(d) is None
]
print(f" {ok(not missing)}{name:<14}", end="")
print("" if not missing else f"missing: {', '.join(missing)}")
if shutil.which("ffmpeg") is None:
print(" -- ffmpeg not on PATH (needed for some video/audio decoding)")
# --- hardware ----------------------------------------------------
print("\nHardware:")
try:
import torch
if torch.cuda.is_available():
g = torch.cuda.get_device_properties(0)
print(
f" OK CUDA: {g.name}, "
f"{g.total_memory / 1024**3:.1f} GB"
)
elif (
getattr(torch.backends, "mps", None)
and torch.backends.mps.is_available()
):
print(" OK Apple MPS available")
else:
print(
" -- CPU only. Works, but expect slow indexing "
"on large corpora."
)
except ImportError:
print(" -- torch not importable")
# --- what the four configs will cost you in RAM ------------------
print("\nConfigurations you can load:")
for label, params, cfg in [
(
"text + code",
"270M",
'{"vision_config": None, "audio_config": None}',
),
(
"text + image + video",
"440M",
'{"audio_config": None}',
),
(
"text + audio",
"570M",
'{"vision_config": None}',
),
(
"everything",
"740M",
'{} (default)',
),
]:
print(f" {params:>5} {label:<22} config_kwargs={cfg}")
print(
"\nAll four load from the same checkpoint and share one vector space."
)
print("Embeddings made with one config are comparable with another.")
Output:

Queries and documents use different task prompts. encode_query() and encode_document() apply the correct prompt automatically. Google’s guide shows prompt_name passed to encode(); the dedicated methods are harder to get wrong.
File: 01_text_and_code.py
from sentence_transformers import SentenceTransformer
MODEL_ID = "google/embeddinggemma-2"
print("Loading text-only config (270M)...")
model = SentenceTransformer(
MODEL_ID,
config_kwargs={
"vision_config": None,
"audio_config": None,
},
)
print(
f"Loaded. Embedding dimension: "
f"{model.get_sentence_embedding_dimension()}\n"
)
# A small corpus mixing prose and code so you can see the difference
# task prompts make on each.
DOCS = [
"The northern lights are caused by charged particles from the sun striking the upper atmosphere.",
"Aurora borealis is most visible near the magnetic poles during periods of high solar activity.",
"To reverse a linked list in place, walk the list while rewiring each node's next pointer.",
(
"def reverse_list(head):\n"
" prev = None\n"
" while head:\n"
" head.next, prev, head = prev, head, head.next\n"
" return prev"
),
"Postgres VACUUM reclaims storage occupied by dead tuples left behind by updates and deletes.",
"A B-tree index stores keys in sorted order and supports range scans efficiently.",
]
QUERIES = [
"what causes the northern lights",
"how do I reverse a linked list",
"why does my database table keep growing",
]
# The important bit: queries and documents get DIFFERENT prompts.
# sentence-transformers 6.1 ships dedicated helpers for this.
doc_emb = model.encode_document(DOCS)
print(f"Encoded {len(DOCS)} documents -> {doc_emb.shape}\n")
for q in QUERIES:
q_emb = model.encode_query(q)
sims = model.similarity(q_emb, doc_emb)[0]
best = int(sims.argmax())
print(f"Q: {q}")
print(f" score {sims[best]:.4f} -> {DOCS[best][:72]}...")
# Show the runner-up so you can see the margin.
order = sims.argsort(descending=True)
second = int(order[1])
print(
f" runner-up {sims[second]:.4f} "
f"margin {sims[best] - sims[second]:.4f}\n"
)
print(
"If the margin is small, your corpus has near-duplicates "
"or the query is ambiguous."
)
print(
"That number is more useful than the raw score "
"for debugging retrieval."
)
Output:

All three queries retrieved the correct document. Margins:
| Query | Top score | Margin |
| what causes the northern lights | 0.8525 | 0.1299 |
| how do I reverse a linked list | 0.8485 | 0.0599 |
| why does my database table keep growing | 0.6673 | 0.0407 |
Media is passed as a dictionary keyed by modality with no prompt. Only the text query gets a task prompt.
File: 02_multimodal.py
from sentence_transformers import SentenceTransformer
MODEL_ID = "google/embeddinggemma-2"
print("Loading full multimodal model (740M)...")
model = SentenceTransformer(MODEL_ID)
# Media is passed as a dict keyed by modality, with NO prompt.
# Only the text query gets a task prompt.
image_emb = model.encode({"image": "data/sunset_beach.jpg"})
audio_emb = model.encode({"audio": "data/ocean_waves.wav"})
for query in [
"ocean waves at sunset",
"a busy city street",
"someone playing piano",
]:
q = model.encode_query(query)
s_img = float(model.similarity(q, image_emb)[0][0])
s_aud = float(model.similarity(q, audio_emb)[0][0])
print(f"\n{query!r}")
print(f" vs image : {s_img:+.4f}")
print(f" vs audio : {s_aud:+.4f}")
# Interleaved: ONE embedding covering text, photo and video together.
# The markers say where each media item sits inside the text.
listing_emb = model.encode(
{
"text": (
"Waterproof trail shoe. <|image|> "
"Grip test on wet rock: <|video|>"
),
"image": "data/trail_shoe.jpg",
"video": "data/grip_test.mp4",
}
)
q = model.encode_query("waterproof trail shoes")
print(
f"\ninterleaved product listing : "
f"{float(model.similarity(q, listing_emb)[0][0]):+.4f}"
)
print(
"\nThe control to watch: scores for unrelated queries "
"should sit clearly lower."
)
print(
"If 'a busy city street' scores close to 'ocean waves at sunset' "
"on the same"
)
print(
"image, the embedding is not discriminating and retrieval "
"will be noisy."
)
Output:


Measures what truncation costs on your corpus rather than on Google’s.
File: 03_matryoshka.py
import numpy as np
from sentence_transformers import SentenceTransformer
MODEL_ID = "google/embeddinggemma-2"
DIMS = [768, 512, 256, 128]
model = SentenceTransformer(
MODEL_ID,
config_kwargs={
"vision_config": None,
"audio_config": None,
},
)
# Replace these with YOUR corpus and YOUR queries plus known-correct answers.
DOCS = [
"Postgres VACUUM reclaims storage from dead tuples after updates and deletes.",
"A B-tree index stores keys sorted and supports efficient range scans.",
"Connection pooling reuses database connections to avoid per-request handshakes.",
"Write-ahead logging records changes before they are applied to data files.",
"Table partitioning splits one large table into smaller physical pieces.",
"The northern lights are caused by charged particles from the sun.",
"Linked list reversal rewires each node's next pointer while walking the list.",
"A hash index supports equality lookups but not range queries.",
]
# (query, index of the correct document)
GOLD = [
("why does my table keep growing after deletes", 0),
("speed up queries over a date range", 1),
("too many database connections being opened", 2),
("how is durability guaranteed on crash", 3),
]
print(
f"{'dim':>5} {'storage/1M':>12} {'top-1':>7} "
f"{'MRR':>7} {'mean score':>11}"
)
print("-" * 52)
results = {}
for d in DIMS:
doc_emb = model.encode_document(
DOCS,
truncate_dim=d,
normalize_embeddings=True,
)
hits, rr, scores = 0, [], []
for q, gold_idx in GOLD:
q_emb = model.encode_query(
q,
truncate_dim=d,
normalize_embeddings=True,
)
sims = model.similarity(q_emb, doc_emb)[0].cpu().numpy()
order = np.argsort(-sims)
rank = int(np.where(order == gold_idx)[0][0]) + 1
hits += rank == 1
rr.append(1.0 / rank)
scores.append(float(sims[gold_idx]))
mb = 1_000_000 * d * 2 / 1024**2 # bfloat16
results[d] = (
hits / len(GOLD),
float(np.mean(rr)),
float(np.mean(scores)),
)
print(
f"{d:>5} {mb:>9.0f} MB "
f"{hits / len(GOLD):>7.2f} "
f"{np.mean(rr):>7.3f} "
f"{np.mean(scores):>11.4f}"
)
print(
"\nPick the smallest dimension where top-1 and MRR "
"still hold on YOUR data."
)
print(
"Google's guidance: 256d keeps most text quality, "
"128d drops to roughly 90%."
)
print(
"Your corpus may behave differently, which is the entire "
"point of measuring it."
)
Output:


| Dimension | Mean score | Change vs 768d |
| 768d | 0.7232 | baseline |
| 512d | 0.7291 | +0.8% |
| 256d | 0.7393 | +2.2% |
| 128d | 0.8145 | +12.6% |
Code retrieval is where version 2 improved most. This run indexed the sentence-transformers library against itself, so the correct answers are checkable.
File: 04_code_search.py
import pathlib
import numpy as np
from sentence_transformers import SentenceTransformer
ROOT = pathlib.Path("./your-repo") # <-- change this
MAX_FILES = 300
DIM = 256 # 3x storage saving, keeps most text/code quality
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={
"vision_config": None,
"audio_config": None,
},
)
def chunks(path, max_lines=60):
"""Split a file into function-sized pieces. Crude but good enough to index."""
try:
lines = path.read_text(errors="ignore").splitlines()
except Exception:
return
buf, start = [], 1
for i, ln in enumerate(lines, 1):
if ln.startswith(("def ", "class ")) and buf:
yield start, "\n".join(buf)
buf, start = [ln], i
else:
buf.append(ln)
if len(buf) >= max_lines:
yield start, "\n".join(buf)
buf, start = [], i + 1
if buf:
yield start, "\n".join(buf)
records = []
for f in list(ROOT.rglob("*.py"))[:MAX_FILES]:
for line_no, body in chunks(f):
if body.strip():
records.append(
{
"file": str(f.relative_to(ROOT)),
"line": line_no,
"code": body,
}
)
print(f"Indexing {len(records)} chunks from {ROOT} ...")
emb = model.encode_document(
[r["code"] for r in records],
truncate_dim=DIM,
normalize_embeddings=True,
batch_size=32,
show_progress_bar=True,
)
mb = emb.nbytes / 1024**2
print(f"Index built: {emb.shape}, {mb:.1f} MB in memory\n")
for query in [
"where are retry attempts with exponential backoff handled",
"code that validates a JWT token",
"function that writes results to a CSV file",
]:
q = model.encode_query(
query,
truncate_dim=DIM,
normalize_embeddings=True,
)
sims = (emb @ q.T).ravel()
print(f"Q: {query}")
for rank, i in enumerate(np.argsort(-sims)[:3], 1):
head = records[i]["code"].strip().splitlines()[0][:62]
print(
f" {rank}. {sims[i]:.4f} "
f"{records[i]['file']}:{records[i]['line']} {head}"
)
print()
Output:

| Your data | config_kwargs | Parameters |
| Text and code only | {'vision_config': None, 'audio_config': None} |
270M |
| Text, images, video | {'audio_config': None} |
440M |
| Text and audio | {'vision_config': None} |
570M |
| Everything | default | 740M |
Disabled encoders are never loaded, so the saving applies to weights and peak memory.
Decide this by what breaks first. If your constraint is recall, stay high. If it is memory or query latency, come down and measure what you lost.
| Dimension | Pick it when |
| 768d or 512d | You are searching across media, or a missed result costs more than the extra storage |
| 256d | Storage is a real constraint. A sensible default for text-heavy indexes. |
| 128d | The index is large and text-only, or this is a cheap first pass before a re-ranker |
Our Step 3 run found 256d identical to 768d on ranking, so starting at 256d and moving up only if you measure a loss is the cheaper order to work in.
One checkpoint under Apache 2.0 covers five content types in one vector space, loads only the encoders you need, and trades vector size against quality with one parameter.
Run Step 3 first on your own corpus. It decides your index size, memory budget and hosting cost, and no published benchmark can give you that number.
Some of the images have been sourced from the developer guide, the model weights, the Gemma documentation, and the Embedding Draw Challenge.