semantic-mapper / EVAL_RESULTS.md
bogdanraduta's picture
Add EVAL_RESULTS.md
7975cfa verified
|
Raw
History Blame Contribute Delete
2.57 kB
# FlowX Semantic Mapper (4B) — Evaluation Results
Held-out evaluation of the released model (`flowx-semantic-mapper-4b-v2`). All numbers are on a
**document-disjoint held-out set of 112 records** spanning EN/FR/DE/RO across banking,
insurance, logistics, and labor. Decoding: greedy (temperature 0), `enable_thinking=False`.
Metrics are exact set-match precision/recall/F1 unless noted; `concepts` also reports a
fuzzy/semantic token-match F1, and `entities` reports field-accuracy over actor/action/object.
## Headline (held-out, n=112)
| Field | Precision | Recall | F1 |
| --- | --- | --- | --- |
| JSON validity | | | 1.00 |
| all-3-facets present | | | 1.00 |
| domain_tags (free-text) | 0.552 | 0.570 | 0.555 |
| concepts (controlled 252-vocab, exact) | 0.547 | 0.558 | 0.541 |
| concepts (fuzzy / semantic match) | 0.596 | 0.590 | 0.578 |
| entities (field-accuracy) | | | 0.530 |
## The controlled-vocabulary result (before / after)
The central design change of this release was replacing open free-text `concepts` (1062
unique labels, 88% seen exactly once, unlearnable and unmeasurable) with a **controlled
252-concept taxonomy** (`concept_taxonomy.yaml`), then expanding the corpus with real
regulatory text.
| Concepts F1 | value | held-out |
| --- | --- | --- |
| open free-text vocabulary (v1) | 0.24 (exact) / 0.37 (fuzzy) | n=35 |
| controlled vocab + data expansion (this release) | **0.54** (exact) / **0.58** (fuzzy) | n=112 |
Concepts F1 more than doubled, and `entities` went from **unmeasurable** (the facet was empty
in the old held-out) to **0.53**, on a held-out set that is 3x larger and multilingual.
## Honest reading / caveats
- **Home-field note.** The held-out shares the annotation lineage of the training data
(real regulatory text, frontier-model-assisted ontology annotation to the controlled vocab).
Read these as a strong format-and-consistency signal; validate on your own corpus before
production claims.
- **Concepts are partial (~0.54).** The model surfaces many but not all applicable concepts.
- **`domain_tags` are free-text on purpose** (recall + human browsing) and are honestly noisier
than the controlled `concepts`.
- **English is strongest;** FR/DE/RO are supported and materially improved by the real-data
expansion, but remain a bit weaker than EN.
## Reproduction
Held-out set: `mlx_data/ontology_cv2/valid.jsonl` (built document-disjoint from the training
split). Scorer: `eval_mapper_4b.py <model_path> mlx_data/ontology_cv2/valid.jsonl`.
_Author: Bogdan Răduță, Head of Research, FlowX.AI._