EmbeddingGemma 2 โ Gewell
This is the EmbeddingGemma 2 bundle used by
Gewell, a single-GPU inference
engine specialized for NVIDIA Blackwell (sm_120/sm_120a). It is a lossless
BF16 repack of google/embeddinggemma-2 into Gewell's native embedding layout,
split into separately loadable text, vision and audio components. It is not a
Transformers, Sentence Transformers, vLLM, or llama.cpp checkpoint.
Bundle contents
| File | Bytes | SHA-256 |
|---|---|---|
text.safetensors |
542,056,504 | 84bfdbbc9d5fe6dced9fd3b572329416fe2794c9c9e3a157346bc29b7c7ff361 |
vision.safetensors |
335,543,360 | 2a146e1979e1fa9ac660aa871a2b405b92a0224b938ab4573dced85b1e43d600 |
audio.safetensors |
611,314,696 | ca511bf78d93dfc5b9d72bd61574e36087188a357e86f6fdbe65e034ae595017 |
config.json |
4,455 | b8f1e9931b57fbc054acdb445c41765d55b0074c58d145fa82839941ad1b5bb3 |
processor_config.json |
1,788 | 168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c |
tokenizer.json |
32,170,510 | 4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4 |
The components hold 413 text, 211 vision/bridge and 752 audio/bridge tensors (1,376 total). Tensor values are copied without conversion; native startup validates the configuration, tokenizer, processor contract and tensor inventory of every loaded component.
Download and launch
Follow the Gewell build instructions, then download this repository:
MODEL_DIR=/absolute/path/to/EmbeddingGemma-2-Gewell
hf download LeDissolution/EmbeddingGemma-2-Gewell --local-dir "$MODEL_DIR"
build/gewell serve-embeddings "$MODEL_DIR" --vision --audio
The server listens on 127.0.0.1:6311 and serves the OpenAI-compatible
POST /v1/embeddings endpoint under the model ID google/embeddinggemma-2:
curl -s http://127.0.0.1:6311/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{"model":"google/embeddinggemma-2","input":["task: search result | query: What causes auroras?","title: none | text: Charged particles from the sun cause auroras."],"dimensions":768}'
Text is always loaded. --vision adds image and video input; --audio adds
audio input. Mixed image/video and audio requests need both. Explicit GPU
allocations at the default 8192-token capacity, excluding CUDA/cuBLAS/cuDNN
overhead:
| Options | Weight bytes | Scratch bytes |
|---|---|---|
| text only | 274,210,816 | 167,773,696 |
--vision |
610,017,280 | 178,782,528 |
--vision --audio |
1,223,507,968 | 1,028,441,648 |
The token-embedding table stays in host memory. Inputs support up to 8192 tokens including BOS/EOS and 128/256/512/768 output dimensions. Task prefixes are not inserted automatically; supply them as described in the upstream model card.
The engine was developed and tested on an RTX PRO 6000 Blackwell. See the model guide, CLI reference and HTTP API documentation for request formats, media limits and serving options.
Scope and provenance
This repository was created with
python3 tools/extract_embeddinggemma2.py --snapshot SNAPSHOT --vision --audio --output MODEL_DIR
from google/embeddinggemma-2 at revision
914f7f89142e33e77833254d9c9b90c3cef7303b. Extraction is deterministic; you
can reproduce these files from the upstream snapshot instead of downloading
them.
Use is subject to the upstream EmbeddingGemma 2 license and terms.
- Downloads last month
- 16
Model tree for LeDissolution/EmbeddingGemma-2-Gewell
Base model
google/embeddinggemma-2