File size: 6,980 Bytes
595ee34 b4b9404 595ee34 b4b9404 595ee34 b4b9404 595ee34 b4b9404 595ee34 b4b9404 595ee34 b4b9404 595ee34 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 | ---
license: apache-2.0
base_model: intfloat/multilingual-e5-small
library_name: multilingual-embedding
pipeline_tag: sentence-similarity
tags:
- retrieval
- cross-lingual
- legal
- eu-law
- eurlex
- lora
language:
- de
- en
- es
- fr
- it
---
# embed-eulaw-multi
A cross-lingual retriever for European Union law. It is a LoRA adapter over
`intfloat/multilingual-e5-small` that finds the same provision of EU law across five
languages: German, English, Spanish, French and Italian.
Give it a provision in one language and it retrieves the matching provision in another. It is
a retriever, not a reranker, and not a question-answering model.
## The number
Measured on a document-isolated held out split, 5,509 queries against a 5,509-passage pool:
| | recall@1 | 95% CI | hits |
|---|---|---|---|
| base `intfloat/multilingual-e5-small` | 0.2919 | [0.2800, 0.3040] | 1,608 / 5,509 |
| **+ this adapter** | **0.6480** | [0.6353, 0.6605] | 3,570 / 5,509 |
| delta | **+0.3561** | | **+122.0%** |
The two intervals do not overlap.
Per language after adaptation: de 0.6534, en 0.6480, es 0.6400, fr 0.6480, it 0.6510. The
spread is flat, so no single language is carrying the average. The largest gain is German,
which starts weakest on the base at 0.1352.
## What this number does and does not establish
**It is in distribution.** The adapter was trained and scored on the same corpus and the same
task. It measures how much adaptation helps where it was trained. It is not evidence of
transfer to other legal text, other domains, or other languages, and we have not measured
that. A sibling model of ours gained substantially in distribution and then transferred flat
out of origin, so please do not assume this one will not do the same.
**It is one training run.** We report the draw count, which is one. The training loop is only
partly seeded: the data shuffle and dropout draw from a configured seed, but the LoRA down
projection is initialised at attach time, before that seed is set. A re-run would start from a
different adapter and we cannot promise it lands on the same number. That is not a claim it
would move. We have not measured this adapter's run to run spread.
**Passages were scored under a 256 token cap.** We have not measured how often a relevant
passage was truncated at that cap. The number above is the number under that setting.
**Two earlier figures are retracted.** Development runs reported +126.3% and +128.79% on a
split that held out texts rather than whole documents. Those measured a different quantity and
are withdrawn. They are not larger versions of +122.0%, the training volumes are not matched,
and no difference between them may be read as the cost of anything.
## Using it
Install the package, then download the adapter and load it from the local directory:
```bash
pip install 'quanfire-multilingual-embedding[neural]'
hf download quanfire-ai/embed-eulaw-multi --revision v1.0.0 --local-dir embed-eulaw-multi
```
```python
from multilingual_embedding.embedding.neural.adapter import load_adapter
encoder, meta = load_adapter("embed-eulaw-multi")
texts = [
"The processing of personal data shall be lawful only if the data subject has given consent.",
"Die Verarbeitung personenbezogener Daten ist nur rechtmaessig, wenn die betroffene Person ihre Einwilligung erteilt hat.",
"Member States shall ensure that fishing vessels exceeding 12 metres carry a satellite tracking device.",
]
vectors = encoder.encode_batch(texts)
```
`load_adapter` returns the encoder together with its metadata, because the metadata carries
the prefixes needed to use the model correctly. `encode_batch` takes a list of strings;
`encode` takes a single string.
On the three sentences above, the English and German expressions of the same provision sit at
cosine 0.835, while the unrelated fisheries article sits at 0.379.
### Base revision
This adapter's manifest predates our base pinning convention, so `meta.checkpoint_revision`
reads back as `None`. The base revision it was built against is
`614241f622f53c4eeff9890bdc4f31cfecc418b3`, which we establish from the training host's model
cache rather than from a field the run recorded: that snapshot is the only one present there
and it predates the run by a month. Pass it explicitly if you need the base held still, since
an upstream repository can change what sits behind a name:
```python
encoder, meta = load_adapter(
"embed-eulaw-multi",
revision="614241f622f53c4eeff9890bdc4f31cfecc418b3",
)
```
## How it was trained
LoRA rank 32, alpha 64, applied to the query and value projections. 589,824 adapter
parameters. One epoch over 19,966 aligned pairs, learning rate 1e-4, batch size 256, mean
pooling, 384 dimensions, max sequence length 256, normalised outputs.
The training pairs are not mined and not machine translated. The Union publishes the same act
as professionally aligned expressions in many languages, so the German text and the French
text of one regulation are two authentic expressions of the same law, aligned by the Union's
own structure: the same CELEX identifier, the same article position. The alignment was done by
the people who wrote the law.
The held out split isolates whole documents. No held out provision shares a document with
anything in training.
**The exact training tree cannot be handed over as a commit.** The working copy of the
training module on the build host was written after this adapter was, so the file on disk is
provably not the file that ran. What we can hand over is the pair of commits that introduced
the document level split, both verifiably in effect before this run started. The evaluation
provenance is intact and independently checkable; it is the training tree that is not
recoverable. No claim here asserts reproducible training.
## Provenance and licence
The corpus is EU legal text from EUR-Lex, the Union's official legal database, taken through
its official channels only. No mirrors and no third party re-hosts.
EUR-Lex legal texts are reusable under Creative Commons Attribution 4.0, granted by Commission
Decision 2011/833/EU. Commercial reuse is named in the grant, not inferred. Database metadata
is CC0.
Only the Official Journal is treated as the authentic source. Material that travels alongside
it is out of scope and is not in the corpus: no third party inserts, no publisher marks, and
no international accounting standards, which carry their own rights and would need separate
clearance.
**The weights are Apache-2.0. The corpus is not redistributed.** The model is
non-reconstructive: it is a 589,824 parameter adapter that emits vectors, not text, and it
cannot reproduce its training corpus. Attribution for the underlying legal text is to the
European Union via EUR-Lex.
## Scope
Five languages, EU legal provisions, in distribution. It has not been evaluated on national
law, on case law, on the other nineteen EU languages, or on any non legal text.
|