| --- |
| license: mit |
| language: |
| - en |
| library_name: pytorch |
| pipeline_tag: feature-extraction |
| tags: |
| - sentence-embeddings |
| - embeddings |
| - distillation |
| - matryoshka |
| - mteb |
| - permissive |
| --- |
| |
| # open-ogma-micro |
|
|
| **A 2.51M-parameter English sentence-embedding model with a fully permissive |
| license chain** β MIT-licensed teachers, ODC-BY corpora, MIT weights. The |
| `open-ogma` family is the permissive successor to the `axiotic/ogma-*` family |
| (which is CC-BY-NC-4.0 due to its training-data mix). |
|
|
| - **Architecture**: 2-layer transformer trunk (d_model=128, d_output=128, |
| mean pooling, task tokens `[QRY]/[DOC]/[SYM]`) + two projection heads: |
| `proj_small` 128β384 and `proj_large` 128β1024. |
| - **Default encode path**: `proj_small` (384d) β the head used for the headline |
| MTEB row. **High-fidelity mode**: `proj_large` (1024d) scores **+0.010 mean / |
| +0.023 retrieval** over the default β opt in when output width is not a |
| constraint. |
| - **Matryoshka**: both heads and the trunk were trained with nested-slice |
| losses at [32, 64, 128] β truncate any output to its leading dims and |
| re-normalise for smaller embeddings. |
|
|
| ## MTEB(eng, v2) β full 41-task benchmark |
|
|
| | output | dim | mean | Class | Clust | Pair | Rerank | Retr | STS | Summ | |
| |---|---:|---:|---:|---:|---:|---:|---:|---:|---:| |
| | `proj_large` (opt-in) | 1024 | **0.5202** | β | β | β | β | 0.2823 | β | β | |
| | `proj_small` (default) | 384 | 0.5104 | β | β | β | β | 0.2589 | β | β | |
| | trunk (pre-head) | 128 | 0.5057 | β | β | β | β | 0.2566 | β | β | |
|
|
| Reference points measured with the identical harness (subprocess-per-task, |
| mteb 2.12/2.18; harness validated by reproducing `minishlab/potion-base-8M` |
| at 0.5328 vs its official 53.33): teacher `bge-small-en-v1.5` (33M) scores |
| 0.6368 and `bge-large-en-v1.5` (335M) 0.6528 β this model retains ~82% of |
| the small teacher's mean at 7.5% of its parameters. |
|
|
| In-training small-mteb (20-task subset) best: **0.5271**. |
|
|
| ## Recipe |
|
|
| Distilled for 100B tokens (bs=128 Γ seq 1024) from an **ensemble of two MIT |
| teachers** β `BAAI/bge-small-en-v1.5` (384d) and `BAAI/bge-large-en-v1.5` |
| (1024d) β each supervised through its own projection head with equal loss |
| weight (0.5/0.5), on a 10.6M-document permissive blend of C4 and FineWeb-Edu |
| (contamination-flagged against MTEB test sets, 0.11% hits marked). Loss: |
| matryoshka-weighted distillation (Ξ±) + in-batch contrastive (Ξ²) on a |
| 1.0β0.7/0.3 schedule; task-token mix {QRY 0.25, DOC 0.25, SYM 0.5}; AdamW |
| lr 5e-4 cosine, wd 0.01. |
|
|
| ## License chain (verified) |
|
|
| | component | license | |
| |---|---| |
| | teacher `BAAI/bge-small-en-v1.5` | MIT | |
| | teacher `BAAI/bge-large-en-v1.5` | MIT | |
| | corpus C4 (allenai/c4) | ODC-BY | |
| | corpus FineWeb-Edu (HuggingFaceFW) | ODC-BY | |
| | tokenizer (ogma sentencepiece, trained in-project) | MIT | |
| | **these weights** | **MIT** | |
|
|
| ODC-BY requires attribution for the source corpora, which this card provides. |
| No non-commercial or share-alike terms anywhere in the chain. |
|
|
| ## Usage |
|
|
| ```python |
| # clone this repo, then: |
| from ogma_libre import OgmaLibre |
| model = OgmaLibre.from_repo("path/to/open-ogma-micro") |
| |
| embs = model.encode(["A quick brown fox."]) # (N, 384) default |
| embs = model.encode(["A quick brown fox."], head="large1024") # (N, 1024) best quality |
| embs = model.encode(["A quick brown fox."], head="base") # (N, 128) trunk |
| |
| # Retrieval convention: encode queries with task="qry", documents with task="doc". |
| q = model.encode(["what is a fox?"], task="qry") |
| d = model.encode(["The fox is a small canid."], task="doc") |
| ``` |
|
|
| Files: `model.safetensors` (weights), `config.json` (architecture + heads), |
| `tokenizer/ogma_sp.model` (sentencepiece), `ogma/` (vendored model code), |
| `ogma_libre.py` (loader). |
|
|
| ## Sibling |
|
|
| `axiotic/open-ogma-small` β the 8.96M-parameter sibling (same recipe, small trunk): 0.5843 @1024d. |
|
|