shl-recommender / eval /TUNING_LOG.md
Eshit's picture
HF Space deploy snapshot
5733f37
|
Raw
History Blame Contribute Delete
3.94 kB
# Retrieval Tuning Log β€” Mean Recall@10
Metric: **Mean Recall@10** = mean over the 10 public traces of
`|expected ∩ top-10| / |expected|`, where the labeled shortlist is mapped to
catalog ids. Measured by `python eval/recall_eval.py [--assemble] [--router]`.
Two query modes are reported:
- **user-text** (deterministic): query = concatenation of the trace's user
messages. Reproducible; isolates retrieval changes from LLM noise.
- **router** (live): query = the router's synthesized `search_query` (one LLM
call, temperature 0; ~Β±0.01 run-to-run).
| Step | Change | Query | Mean Recall@10 | Ξ” |
|------|--------|-------|---------------:|----:|
| 0 | **Baseline** hybrid dense+BM25+RRF (n=50, rrf_k=60) | user-text | **0.5012** | β€” |
| a | Sweep RRF constant Γ— candidate N | user-text | 0.48–0.52 | ~0 |
| d | **test_type-aware assembly** (inject OPQ32r `P` + Verify G+ `A` defaults) | user-text | **0.6931** | **+0.1919** |
| b | Enrich embedded text (name emphasis; job_levels already present) | user-text | 0.6931 | 0 |
| c | Router `search_query` synthesis | router | 0.7131 | +0.0200 |
| cβ€² | Router prompt: "include every distinct skill" | router | ~0.72 (0.713–0.727) | +0.01 |
**Baseline 0.5012 β†’ shipped ~0.72** (router query + assembly). Groundedness is
1.0 by construction (URLs looked up from `catalog.json`; non-catalog ids dropped).
---
## (a) RRF constant + dense/BM25 N β€” *flat, kept defaults*
Swept `rrf_k ∈ {10,20,40,60,100}` Γ— `n ∈ {20,30,50,80,120,200,370}`. The whole
grid sits in 0.48–0.52; the single best cell (n=30, rrf_k=20 β†’ 0.5155) is not
robust (neighbouring cells drop to 0.4955), i.e. noise. **Kept n=50, rrf_k=60.**
Fusion params are not the bottleneck β€” the misses are missing *documents*, not
mis-ranked ones.
## (b) Enrich embedded text β€” *no runway, flat*
CLAUDE.md suggested adding `job_levels` + `competencies`. `job_levels` is already
in `assessment_text`; the catalog has **no** competencies/keywords field beyond
the test-type `keys` (already included). Tested name-emphasis (name embedded
twice) to lift crowded short-named tests (`sql-new`, `ms-excel-new`): **0.6931 β†’
0.6931, no change.** Reverted.
## (c) Router search_query synthesis β€” *small, real gain*
Using the router's synthesized query instead of raw user text: **0.6931 β†’
0.7131**. Adding "include EVERY distinct skill/technology, each as its own term"
to the router prompt (targets the C9 Java-flood that drops `sql-new`/`docker-new`)
nudged to ~0.72. Small but real; kept.
## (d) test_type-aware assembly β€” *the lever, +0.19*
The labeled shortlists are **batteries**: a role-skills spine + two recurring
defaults the user rarely names β€” a personality measure (**OPQ32r**, `P`, in 7/10
finals) and a cognitive measure (**SHL Verify Interactive G+**, `A`). After
retrieval we reserve slots and guarantee these defaults for hiring-style requests,
unless the user opts out (`app/assembly.py`). This recovered the single largest
miss class: **0.5012 β†’ 0.6931**. Opt-out uses proximity matching so "keep Verify
G+" alongside "drop the OPQ" only drops personality (see `test_assembly.py`).
---
## What didn't work
- **Per-message / per-skill multi-query fusion** (RRF across sub-queries):
**0.486**, *worse* than a single query β€” sub-queries pulled unrelated items and
diluted the exact-name signal. Abandoned.
- **RRF/N tuning** and **text enrichment**: both flat (above).
## Remaining misses (not recoverable by retrieval alone)
Near-duplicate product families where the short canonical variant loses to
`-365`/`-essentials` siblings (C8 `ms-excel-new`), multi-skill JDs where one
skill family floods the pool (C9 `sql-new`, `docker-new`), and second personality
*reports* (C1 `opq-universal-competency-report`). These are ranking/diversity
problems; the next lever would be per-family de-duplication, not more fusion
tuning. **Returns have flattened β€” stopped here.**