EmbeddingGemma 2 โ€” Gewell

This is the EmbeddingGemma 2 bundle used by Gewell, a single-GPU inference engine specialized for NVIDIA Blackwell (sm_120/sm_120a). It is a lossless BF16 repack of google/embeddinggemma-2 into Gewell's native embedding layout, split into separately loadable text, vision and audio components. It is not a Transformers, Sentence Transformers, vLLM, or llama.cpp checkpoint.

Bundle contents

File Bytes SHA-256
text.safetensors 542,056,504 84bfdbbc9d5fe6dced9fd3b572329416fe2794c9c9e3a157346bc29b7c7ff361
vision.safetensors 335,543,360 2a146e1979e1fa9ac660aa871a2b405b92a0224b938ab4573dced85b1e43d600
audio.safetensors 611,314,696 ca511bf78d93dfc5b9d72bd61574e36087188a357e86f6fdbe65e034ae595017
config.json 4,455 b8f1e9931b57fbc054acdb445c41765d55b0074c58d145fa82839941ad1b5bb3
processor_config.json 1,788 168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c
tokenizer.json 32,170,510 4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4

The components hold 413 text, 211 vision/bridge and 752 audio/bridge tensors (1,376 total). Tensor values are copied without conversion; native startup validates the configuration, tokenizer, processor contract and tensor inventory of every loaded component.

Download and launch

Follow the Gewell build instructions, then download this repository:

MODEL_DIR=/absolute/path/to/EmbeddingGemma-2-Gewell

hf download LeDissolution/EmbeddingGemma-2-Gewell --local-dir "$MODEL_DIR"

build/gewell serve-embeddings "$MODEL_DIR" --vision --audio

The server listens on 127.0.0.1:6311 and serves the OpenAI-compatible POST /v1/embeddings endpoint under the model ID google/embeddinggemma-2:

curl -s http://127.0.0.1:6311/v1/embeddings \
  -H 'Content-Type: application/json' \
  -d '{"model":"google/embeddinggemma-2","input":["task: search result | query: What causes auroras?","title: none | text: Charged particles from the sun cause auroras."],"dimensions":768}'

Text is always loaded. --vision adds image and video input; --audio adds audio input. Mixed image/video and audio requests need both. Explicit GPU allocations at the default 8192-token capacity, excluding CUDA/cuBLAS/cuDNN overhead:

Options Weight bytes Scratch bytes
text only 274,210,816 167,773,696
--vision 610,017,280 178,782,528
--vision --audio 1,223,507,968 1,028,441,648

The token-embedding table stays in host memory. Inputs support up to 8192 tokens including BOS/EOS and 128/256/512/768 output dimensions. Task prefixes are not inserted automatically; supply them as described in the upstream model card.

The engine was developed and tested on an RTX PRO 6000 Blackwell. See the model guide, CLI reference and HTTP API documentation for request formats, media limits and serving options.

Scope and provenance

This repository was created with

python3 tools/extract_embeddinggemma2.py --snapshot SNAPSHOT --vision --audio --output MODEL_DIR

from google/embeddinggemma-2 at revision 914f7f89142e33e77833254d9c9b90c3cef7303b. Extraction is deterministic; you can reproduce these files from the upstream snapshot instead of downloading them.

Use is subject to the upstream EmbeddingGemma 2 license and terms.

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for LeDissolution/EmbeddingGemma-2-Gewell

Finetuned
(45)
this model