File size: 5,778 Bytes
f744235
 
 
 
 
 
 
 
 
ba60dfa
f744235
98ecbb1
f744235
 
 
 
 
0c09e81
98ecbb1
 
 
f744235
0c09e81
98ecbb1
 
 
 
 
 
f744235
98ecbb1
 
 
 
 
 
 
0c09e81
 
f744235
 
 
0c09e81
f744235
 
 
0c09e81
 
 
 
f744235
98ecbb1
0c09e81
 
 
 
f744235
0c09e81
98ecbb1
 
 
 
 
 
 
 
0c09e81
 
 
98ecbb1
 
 
 
0c09e81
98ecbb1
 
 
0c09e81
 
 
ba60dfa
98ecbb1
 
 
 
 
 
 
 
 
0c09e81
98ecbb1
0c09e81
 
98ecbb1
 
 
 
 
 
 
f744235
 
98ecbb1
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
---
license: apache-2.0
language:
- multilingual
tags:
- cross-encoder
- reranker
- feed-ranking
- enterprise-feed
- learning-to-rank
pipeline_tag: text-classification
base_model: nreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large
library_name: transformers
---

# Third-Pass Feed Ranker (job-title Γ— feed-post relevance)

A lightweight **cross-encoder** that scores how relevant an enterprise social-feed post is to a
viewer, given only the viewer's **job title** and the **post text**. It is a *third-pass reranker*:
it re-scores a small candidate slate (~20 items) from earlier passes to surface a genuinely
job-relevant post that was buried below the top slot.

- **Input:** `job_title` (query) + `post_text` (passage) β†’ single relevance score (higher = more relevant)
- **Base:** `nreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large` β€” a **generic** distilled
  multilingual MiniLM-L12 (~117M params), *not* a search reranker. Fine-tuning defines relevance
  purely from this task's data, without a general-search "dense-technical-text = relevant" prior.
- **Runtime:** server-class **CPU** at scale (~120 pairs/sec on 8 CPU threads; INT8/ONNX 3–6Γ— more).
- **Trained with a listwise ranking objective** (optimizes which item wins the slate), pointwise
  inference unchanged.

## What it does β€” and what it deliberately avoids
- **Ranks by MEANING, not keywords.** Relevance is decided by substance (e.g. "patient-monitoring
  escalation protocol" β†’ a clinical role). Masking role words barely changes the ranking, and the
  role keyword is trained to be uncorrelated with the label (neither a shortcut nor a penalty).
- **Not fooled by dense technical jargon.** Training injects dense-technical posts as hard negatives
  for non-technical roles, so a machine-learning or software post does **not** score as broadly
  relevant to, say, a nurse or a chef β€” it only wins for roles it actually fits.

## Usage β€” scoring + the gate
```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
name = "FDS-Iterations/third-pass-feed-ranker"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name).eval()

def scores(title, posts):
    enc = tok([title]*len(posts), posts, truncation=True, max_length=160, padding=True, return_tensors="pt")
    with torch.no_grad():
        return model(**enc).logits.squeeze(-1).tolist()

TAU = 1.0    # promotion margin; calibrate per deployment (see below)
def third_pass(title, slate):
    s = scores(title, slate)
    challenger = max(range(1, len(slate)), key=lambda i: s[i])
    return ("promote", challenger) if s[challenger] - s[0] > TAU else ("no_change", None)
```
**The model only scores.** The gate (which item to move to slot 1, and the Ο„ threshold) is *your*
logic β€” none of it is in the weights. **Ο„ is distribution-sensitive**: on realistic feeds a value
around **0.5–1.5** trades recovery vs. false-promotion; re-calibrate on a sample of your own
no-relevance feeds. Higher Ο„ is more conservative.

## Evaluation
Measured on an **external hold-out of 157 job titles that never appear in training** (adjacent
real-world variants of trained roles + high-volume roles the training taxonomy under-covers), each
with substance gems buried in 20-item feeds:

| metric | value |
|---|---|
| relevant post surfaced to #1 in its feed (novel titles) | **~60%** |
| role's gem beats a dense-technical distractor | **~85%** |
| realistic-density recovery (buried gem promoted) | **~0.64** @ low Ο„ |
| false-promotion on no-relevance feeds | **~0.3%** |

On a deliberately **adversarial worst-case slate** (~9 simultaneous strong competitors, including
posts relevant to *other* roles), exact-#1 recovery drops to ~0.3 β€” a stress bound, not a
realistic-feed number.

## Intended use & limitations
- Re-ranking short enterprise-feed candidate slates by job-title relevance; abstains on role-less
  titles.
- **Trained entirely on SYNTHETIC data** (LLM-generated posts + synthetic slates). **Validate on
  your own data before production.**
- **Fine role-discrimination is limited.** It reliably separates a relevant post from ordinary
  filler, but distinguishing a role's exact post from a *closely adjacent* role's post is near the
  capacity ceiling of a 117M model.
- **Out-of-distribution phrasing is the weak axis.** Terse status fragments, log/ticket snippets,
  and atypical wording score more noisily than well-formed posts β€” for all roles.
- **Very high-volume roles absent from the training taxonomy** (a few common titles) generalize at
  reduced magnitude; adding them to training stabilizes them.
- Cross-lingual mixing (non-English title vs English-only feed) is weaker than same-language feeds.
- Relevance is title-driven; recency/importance beyond relevance must live in your decision logic.

## Training & method
Job titles derived from a public occupation taxonomy; LLM-generated posts where relevance is by
**substance** and the role is never named; adversarial 20-item slates with held-out roles and posts,
in both well-formed and terse registers. Objective = pointwise relevance MSE **plus a listwise
ranking loss** that pushes the relevant post to rank #1 within its slate. Two anti-bias measures:
the role keyword is **rebalanced** to be label-uncorrelated, and **dense-technical hard negatives**
are injected into non-technical roles' slates so technical vocabulary is not a global relevance
signal. Headline metrics are measured on titles **never seen in training**.

## License & attribution
Apache-2.0. Inherits from `nreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large` β€” verify its
license carries through. Training posts were generated with a Qwen model; review the applicable terms.