File size: 8,220 Bytes
bff06c9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
---
license: apache-2.0
library_name: sentence-transformers
pipeline_tag: sentence-similarity
tags:
  - sentence-transformers
  - embeddings
  - legal
  - retrieval
  - feature-extraction
  - mteb
  - mleb
  - legal-tech
  - case-law
  - contracts
  - judicialmind
---

# GreenLeaf Law Embed Tiny

**by [JudicialMind](https://judicialmind.ai)***The AI-Native Legal Intelligence System*

A 596M-parameter dense embedding model purpose-built for legal text retrieval. Scores **78.29% on MLEB-12** and **64.38% on MTEB(Law, v1)** — competitive with models 13x larger.

## Why GreenLeaf?

Most embedding models treat legal text like any other English prose. GreenLeaf doesn't. Every design decision — from bidirectional attention to jurisdiction-aware training data curation — targets the specific challenges of legal retrieval: long documents, precise citations, cross-references, and domain-specific terminology.

Built by [JudicialMind](https://judicialmind.ai) as part of our open research initiative to advance legal AI.

## Highlights

| | |
|---|---|
| **Parameters** | 596M |
| **Embedding Size** | 1024 |
| **Context** | 1024 tokens (supports up to 32,768) |
| **Attention** | Bidirectional — every token sees every other token |
| **Pooling** | Mean |
| **Precision** | bfloat16 (default), int8 and binary quantization built in |
| **License** | Apache 2.0 |

## Performance

### MLEB-12 — Massive Legal Embedding Benchmark

| Dataset | Score |
|---------|-------|
| Legal RAG Bench | 54.16 |
| Bar Exam QA | 68.38 |
| SCALR | 73.04 |
| ECHR Retrieval | 41.27 |
| Singaporean Judicial Keywords | 86.63 |
| GDPR Holdings Retrieval | 93.43 |
| UK Legislative Long Titles | 95.88 |
| Australian Tax Guidance | 78.66 |
| Irish Legislative Summaries | 91.92 |
| Contractual Clause Retrieval | 91.29 |
| License TL;DR Retrieval | 72.54 |
| Consumer Contracts QA | 92.29 |
| **Overall** | **78.29** |

### MTEB(Law, v1)

| Task | Score |
|------|-------|
| AILACasedocs | 40.73 |
| AILAStatutes | 58.68 |
| GerDaLIRSmall | 38.51 |
| LeCaRDv2 | 69.52 |
| LegalBenchConsumerContractsQA | 85.86 |
| LegalBenchCorporateLobbying | 94.94 |
| LegalQuAD | 58.62 |
| LegalSummarization | 68.22 |
| **Overall** | **64.38** |

### How It Compares

| Rank | Model | MLEB | Size |
|------|-------|------|------|
| 1 | Kanon 2 Embedder | 81.9% | — |
| 2 | Voyage 4 Large | 81.1% | 7B+ |
| 3 | Voyage 4 | 79.6% | — |
| **4** | **GreenLeaf Law Embed Tiny** | **78.3%** | **0.6B** |
| 5 | Voyage 4 Lite | 76.4% | — |
| 6 | Qwen3 Embedding 8B | 75.9% | 8B |
| 7 | Gemini Embedding 001 | 72.1% | — |
| 8 | Jina v5 Text Small | 71.0% | — |
| 9 | OpenAI Text Embedding 3 Large | 70.8% | — |

Best accuracy-per-parameter on the leaderboard. Runs on a laptop.

### Matryoshka Dimension Truncation

GreenLeaf embeddings are trained with Matryoshka Representation Learning — truncate to any dimension without retraining. Halving to 512d costs only **1.34%** accuracy; 128d retains **91.6%** of full performance.

| Dataset | 1024d | 512d | 256d | 128d |
|---------|-------|------|------|------|
| ECHR Retrieval | 40.77 | 41.35 | 38.18 | 33.58 |
| Legal RAG Bench | 54.38 | 49.93 | 39.98 | 34.71 |
| SCALR | 72.77 | 70.22 | 72.50 | 68.23 |
| Consumer Contracts QA | 92.14 | 92.27 | 91.13 | 88.66 |
| Singaporean Judicial Keywords | 86.07 | 84.30 | 82.63 | 79.52 |
| Australian Tax Guidance | 78.69 | 77.93 | 76.22 | 73.61 |
| Contractual Clause Retrieval | 91.20 | 90.05 | 87.71 | 83.31 |
| Irish Legislative Summaries | 91.29 | 90.22 | 89.51 | 86.90 |
| GDPR Holdings Retrieval | 93.53 | 93.11 | 92.32 | 89.96 |
| License TL;DR Retrieval | 73.91 | 72.83 | 70.44 | 71.32 |
| UK Legislative Long Titles | 95.88 | 95.62 | 96.99 | 96.73 |
| Bar Exam QA | 68.30 | 65.03 | 58.79 | 53.62 |
| **MLEB-12 Mean** | **78.24** | **76.90** | **74.70** | **71.68** |
| **Drop vs 1024d** | — | **-1.34%** | **-3.54%** | **-6.56%** |

```python
# Truncate at inference — no retraining needed
emb_512 = model.encode(texts)[:, :512]
emb_256 = model.encode(texts)[:, :256]
```

## Quick Start

### SentenceTransformers

```python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "judicialmind/greenleaf-law-embed-tiny",
    trust_remote_code=True,
)

texts = [
    "The plaintiff filed a motion for summary judgment.",
    "Contract terms must be interpreted in good faith.",
]

# Default: raw float embeddings (bfloat16)
embeddings = model.encode(texts)  # Shape: (2, 1024)

# Optional: quantize for storage / speed
emb_int8 = model.encode(texts, quantization="int8")
emb_binary = model.encode(texts, quantization="binary")
```

### Text Embeddings Inference (TEI)

```bash
docker run --gpus all --shm-size 1g -p 8080:80 \
  ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 \
  --model-id judicialmind/greenleaf-law-embed-tiny --dtype float32
```

```bash
curl http://0.0.0.0:8080/embed \
  -H "Content-Type: application/json" \
  -d '{"inputs": ["Legal text here"], "normalize": false}'
```

### ONNX

```python
import onnxruntime as ort
from transformers import AutoTokenizer
import numpy as np

tokenizer = AutoTokenizer.from_pretrained(
    "judicialmind/greenleaf-law-embed-tiny", trust_remote_code=True
)
session = ort.InferenceSession("onnx/model.onnx")

texts = ["Legal document text here"]
inputs = tokenizer(texts, padding=True, truncation=True, return_tensors="np")
outputs = session.run(
    [o.name for o in session.get_outputs()],
    {"input_ids": inputs["input_ids"].astype(np.int64),
     "attention_mask": inputs["attention_mask"].astype(np.int64)},
)
```

## Architecture

GreenLeaf uses a bidirectional transformer encoder derived from the Qwen3 architecture. The key modification: causal masking is removed across all 28 layers, so every token attends to the full input sequence in both directions.

This matters for legal text. A holding on page 12 can change the meaning of a defined term on page 1. A causal model can't see that. GreenLeaf can.

| Property | Value |
|----------|-------|
| Base architecture | Qwen3 |
| Hidden layers | 28 |
| Hidden size | 1024 |
| Attention heads | 16 (8 KV heads) |
| Intermediate size | 3072 |
| Vocabulary | 151,936 |
| Max position embeddings | 32,768 |
| Sliding window | None (full-context) |
| KV cache | Disabled |
| Tie embeddings | Yes |

### Quantization

Three output modes, selectable at inference:

| Mode | Output | Storage | Use case |
|------|--------|---------|----------|
| `None` (default) | float32 per dim | 4 KB/doc | Maximum accuracy |
| `"int8"` | int8 per dim | 1 KB/doc | 4x compression, minimal loss |
| `"binary"` | ±1 per dim | 128 B/doc | 32x compression, fast Hamming search |
| `"ubinary"` | packed bits | 128 B/doc | Same as binary, pre-packed |

## Training

Trained on a large-scale proprietary legal corpus spanning multiple jurisdictions (US, UK, EU, AU, SG, IE) and document types (case law, contracts, legislation, regulatory filings, tax guidance). Contrastive learning with hard negative mining.

## About JudicialMind

[JudicialMind](https://judicialmind.ai) builds an orchestrated system of specialized AI agents that work like a legal team — every agent with a domain expertise, composed into directed workflows that run end-to-end legal operations: intake, research, drafting, negotiation, dispute resolution.

We serve AmLaw firms, in-house legal teams, solo & boutique practices, and courts/ADR bodies.

**Open artifacts:**
- 📦 [judicialmind/legal-training-dataset](https://huggingface.co/datasets/judicialmind/legal-training-dataset) — 3.69M annotated query-passage pairs across 35 languages
- 🤖 More models coming soon — retrieval, reranking, and legal-reasoning models trained on our corpus

**Connect:**
- 🌐 [judicialmind.ai](https://judicialmind.ai)
- 📧 surya@judicialmind.ai
- 🧑‍⚖️ [Early access](https://judicialmind.ai/waitlist)

## Limitations

- 1024-token default context. Longer documents should be chunked (32k context available with `--max-length` override).
- Specialist model. For general-purpose embedding, a general-domain model will perform better on non-legal text.



## License

Apache 2.0. Free for commercial and research use.