File size: 21,170 Bytes
184d40c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cc68a12
 
25910d6
 
 
cc68a12
25910d6
 
 
 
cc68a12
25910d6
 
cc68a12
25910d6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
184d40c
 
 
 
 
8141a5d
 
 
 
 
 
 
 
 
 
 
 
be92d5f
 
 
 
 
 
 
 
43f97a4
 
 
 
 
 
 
 
 
be92d5f
7af8059
 
 
 
 
184d40c
 
 
be92d5f
 
 
 
184d40c
 
 
 
558fffa
d409fec
 
 
 
 
 
 
 
184d40c
 
 
 
 
 
 
 
 
d409fec
 
184d40c
d409fec
 
 
 
 
 
 
 
184d40c
d409fec
184d40c
d409fec
184d40c
 
 
 
 
 
d409fec
 
 
184d40c
 
d409fec
558fffa
 
 
 
184d40c
d409fec
 
184d40c
 
 
d409fec
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
558fffa
 
 
 
 
 
 
 
 
 
 
 
 
d409fec
 
558fffa
 
 
 
 
184d40c
 
 
 
 
 
d409fec
184d40c
 
 
 
 
 
 
 
 
 
 
be92d5f
 
 
 
184d40c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7af8059
 
 
 
 
 
 
 
 
184d40c
 
7af8059
 
 
 
 
 
184d40c
 
d409fec
 
 
 
 
 
 
 
 
 
 
 
 
 
184d40c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d409fec
184d40c
 
d409fec
184d40c
 
d409fec
184d40c
 
 
be92d5f
184d40c
be92d5f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
184d40c
 
 
 
 
 
 
 
 
 
7af8059
 
 
 
 
 
184d40c
be92d5f
 
 
 
 
 
 
 
 
 
 
184d40c
 
7af8059
 
 
 
 
184d40c
 
 
 
8141a5d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
184d40c
 
 
 
 
 
 
be92d5f
 
 
 
184d40c
 
 
 
 
d409fec
184d40c
 
d409fec
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
"""Relevance reranking for search results, using Laya (convaiinnovations).

Why this exists
---------------
The SearXNG bridge races several engines and hands back the first set that
looks big enough. Nothing in that path asks whether the results are *about* the
query, so the fastest engine wins β€” and a small independent index is fast.
Measured on production, 2026-10-02, from the route's own `source` field:

    "what is the capital of Peru"            -> mwmbl -> NPR transcripts about
        mummy lice, an article about mosquitoes in Delhi, "The Capital Cycle".
        Eight results. None of them answers the question.
    "Peru capital city"                     -> mwmbl -> a martial arts studio
        in Wellington, an NScale railroad modellers' club.
    "Donald Trump AI regulations policies"  -> bing  -> "Donald Trump -
        Wikipedia", AP's Trump hub, Reuters' Trump hub.

None of that is an outage, so nothing raised an alarm: HTTP 200, `degraded:
false`, a full result page. It is the failure mode where the search path is
technically healthy and completely useless, and a count-based "is the set big
enough" check cannot see it β€” a lexical overlap check cannot either, because
"Viceroyalty of Peru" really does contain the words "Peru" and "capital" and
still does not answer the question.

A cross-encoder can, because it reads the query and the passage together and
answers "does this passage answer this question" rather than "do these strings
overlap". Laya is the right size for that: a non-autoregressive System 1
decision model, 421M params in English (ModernBERT-large) and 322M for 100+
languages (mmBERT-base), one forward pass per document, Apache-2.0, calibrated
probabilities via RLCD. It never generates text, so there is nothing to parse
and nothing to hallucinate.

Design constraints, all of them learned the hard way
----------------------------------------------------
* **Fail open, always.** If the model is missing, slow, or the import breaks,
  the caller gets its results back untouched. Search must not depend on a
  component that is described as "kind of dumb" and can be swapped out.
* **Never return an empty page.** If the model rejects every result, that is
  not a confident judgement, it is a model that did not understand the query β€”
  so the original set is returned and `starved` is set. That flag is the signal
  to retune DROP_BELOW, not a silent success.
  An earlier version of this file had a floor of 3 surviving results, on the
  theory that a short page is a bad page. That was wrong in the exact case
  this module exists for: the mwmbl result set for "what is the capital of Peru"
  is eight results of which perhaps two are worth showing, and a floor of 3
  throws those two away and serves the eight instead. A short honest page beats
  a full wrong one. The only floor that is genuinely wrong is zero.
* **Log every score.** The threshold is not a guess to be shipped and admired;
  it is a number to be chosen from the distribution below. RECENT keeps the
  last N (query, scores) so the threshold can be tuned from real traffic.
* **Deadline-bounded.** Reranking is an enhancement, never a reason to make a
  user wait: the caller passes a budget and this module stops at it.
"""

import os
import threading
import time
from collections import deque

# ── configuration ───────────────────────────────────────────────────────────
# Off by default so a bad deploy cannot take search down; the route treats a
# disabled reranker as "no opinion" and returns what SearXNG gave it.
ENABLED = os.environ.get("AEL_RERANK", "1") not in ("0", "false", "no")

# "english" is 421M ModernBERT-large, 512 tokens, the strongest.
# "multilingual" is 322M mmBERT-base, ~2x faster, 100+ languages. The route
# sends whichever it wants; the default here is the stronger English one
# because Ael's corpus is overwhelmingly English, and non-English queries are
# still scored correctly by it far more often than not.
MODEL = os.environ.get("AEL_RERANK_MODEL", "english")

# Drop a result the model thinks this unlikely to answer the query.
#
# This went 0.5 -> 0.35 -> 0.15, and the direction is the finding: **the model is
# a good ranker of snippets and a bad dropper of them.** Measured on production,
# 2026-10-02:
#
# Ranked correctly, repeatedly. "capital of Peru" put "Lima" (0.838) above a
# Nat Geo photo of a Peruvian mysticete (0.042); "who won the 1876 presidential
# election" put the Wikipedia election article above a Reddit thread about a
# video game achievement. Reordering is where it earns its keep.
#
# Dropped answers it should not have. At 0.35 the threshold removed the correct
# result for four questions in a row and kept a generic page instead:
#
#   "Unix program that replaced man with compressed pages" -> mandb gone,
#       "Linux/Unix Tutorial" kept
#   "first artificial satellite to orbit a body other than Earth" -> Luna gone,
#       a Hackaday post kept
#   "TLD delegated to a donut shop" -> .donuts gone, ".radio" kept
#   "IUPAC systematic name of caffeine" -> the chemistry pages kept, the
#       answer-bearing ones dropped
#
# The cause is not a bad threshold, it is what the model is reading: a title and
# a snippet, never the page. A snippet about mandb does not contain the words
# "compressed manual pages", so the passage looks like a non-answer even when
# the page is the answer. No threshold fixes that β€” the useful behaviour is to
# *rank* on what the snippet says and let the reader open the page, which is
# what the tool prompt already tells the model to do.
#
# So: drop only what the model actively rejects, and let ordering do the rest.
# Measured unrelated results sit at 0.004-0.098, so 0.15 removes the clearly
# irrelevant while leaving the ambiguous ones for the reader. The result is
# deliberately a longer page than the 0.35 experiment produced β€” a research
# tool that returns one confident result is a worse research tool.
DROP_BELOW = float(os.environ.get("AEL_RERANK_DROP", "0.15"))

# Never return an empty page. See the starvation note above β€” this is a floor of
# one, not a target: a short page is fine, an empty one is not.
MIN_RESULTS = 1

# …but one result is not a page you can research with. When the threshold would
# leave fewer than this, the *highest-scoring* results are kept instead of the
# original set: the floor must not resurrect the junk, so it tops the list back
# up from the model's own ranking rather than falling back to engine order.
#
# This is not the "floor of 3" the module's own docstring warns about. That one
# fell back to the whole original set, so a set of eight with two worth showing
# served all eight. Topping up from the ranking cannot do that: the results that
# come back are still the ones the model rated highest, just including a couple
# more of them.
MIN_KEEP = int(os.environ.get("AEL_RERANK_MIN_KEEP", "2"))

# How many results the model is allowed to judge. A 421M cross-encoder on a
# shared CPU does not score eight documents inside an 800ms budget, and the
# first version's answer to that was to throw the whole set away β€” which is why
# a reranker that was installed, loaded, healthy and simply too slow reported
# itself identically to one that had never run. Judging the first N and leaving
# the rest in engine order is strictly better than judging none: the survivors
# are reordered by relevance and the rejects are dropped, and the untested tail
# keeps the order the race gave it.
#
# 10, from measurement: production runs at ~150ms a document on the Space's
# CPU, and the race hands over 8. At the old cap of 6 the last two results were
# never judged, and an unjudged result is the one thing that re-admits exactly
# the junk this stage exists to remove β€” "capital of Peru" still came back with
# a Granta short story and a Nat Geo photo in it. The cap is now above a full
# result set, so the tail is normally empty; the budget below is what actually
# bounds the work, and this is only a backstop against a pathologically long set.
MAX_SCORED = int(os.environ.get("AEL_RERANK_MAX", "10"))

# How much text the model is asked to judge. The model's context is 512 tokens;
# a title plus a full snippet can exceed it and the model raises rather than
# truncating, which used to surface as a silent per-query failure.
PASSAGE_MAX_CHARS = int(os.environ.get("AEL_RERANK_PASSAGE_CHARS", "1200"))

# Recent (query, scores) for threshold tuning. Bounded so a long-lived Space
# cannot grow without limit.
RECENT = deque(maxlen=int(os.environ.get("AEL_RERANK_LOG", "200")))
# Wall-clock of recent reranks, same bound. This is the tuning data for
# MAX_SCORED and the budget: the scores say what to keep, this says what it
# costs to judge it.
RECENT_MS = deque(maxlen=int(os.environ.get("AEL_RERANK_LOG", "200")))

_router = None
_router_lock = threading.Lock()
_load_error = None
_load_message = None
_loading = False
_load_failed_at = 0.0

# A failed load used to latch for the lifetime of the container, which turned one
# flaky checkpoint download into a permanently disabled reranker β€” the Space
# would report "no opinion" for the rest of its life and nobody would know why.
# Failures are retried on a cooldown instead.
RETRY_AFTER_S = float(os.environ.get("AEL_RERANK_RETRY_S", "60"))


def _get_router():
    """Load the checkpoint once, lazily. Returns None if it cannot be loaded.

    Deliberately never raises: a Space that cannot afford the model still
    serves search, and the route needs to be able to ask "any opinion?"
    without handling an exception.
    """
    global _router, _load_error, _load_message, _loading, _load_failed_at
    if _router is not None:
        return _router
    if _load_error is not None and (time.time() - _load_failed_at) < RETRY_AFTER_S:
        return None
    if _loading:
        # Never block a search on a checkpoint download. The caller's deadline is
        # ~1s; loading 421M parameters takes minutes the first time and is not
        # something to discover inside a user's request. `warm()` exists so this
        # is the rare path rather than the normal one.
        return None
    with _router_lock:
        if _router is not None or _loading:
            return _router
        _loading = True
        try:
            from laya import Router

            # max_loaded=1: one checkpoint resident, which is all this uses.
            _router = Router(max_loaded=1)
            _router.preload([MODEL])
            _load_error = None
            _load_message = None
            print(f"[rerank] {MODEL} ready", flush=True)
        except Exception as e:  # noqa: BLE001 - reported, never raised
            _load_error = f"{type(e).__name__}"
            _load_failed_at = time.time()
            # The message matters and nothing else about this exception does:
            # "the reranker is off" is a decision, "the reranker broke because
            # <this>" is a bug report. The caller redacts before logging it.
            _load_message = f"{type(e).__name__}: {e}"[:300]
            print(f"[rerank] disabled: cannot load laya ({_load_error}): {e}", flush=True)
        finally:
            _loading = False
    return _router


def warm():
    """Start loading the checkpoint now, in the background.

    Called once when the Space boots. Without this the first search request of
    every container lifetime paid the full download-and-load cost inside an
    800ms budget, so the reranker could never once have applied on a cold
    container β€” and the only symptom was a timeout, which reads like a slow
    network. Warming moves the cost to the ~2 minutes after a deploy, where
    nobody is waiting.

    Never raises and never blocks: a Space that cannot afford the model still
    serves search.
    """
    if not ENABLED or _router is not None or _loading:
        return
    t = threading.Thread(target=_get_router, name="rerank-warm", daemon=True)
    t.start()


def _no_opinion():
    """Why this module has nothing to say, for the caller.

    A reranker that is silently disabled is the worst outcome available: the
    route's fail-open path cannot tell "the model declined" from "the model is
    not installed", and both look like a working search that quietly does not
    rerank. The first version of this stage returned `ok: false` and nothing
    else, and that ambiguity cost a full session of debugging the Space's
    build when the Space had been fine the whole time β€” the app simply had the
    URL wrong and never asked.
    """
    if not ENABLED:
        return "disabled by config"
    if _loading:
        return "still loading"
    if _load_error:
        return _load_message or _load_error
    return None


def status():
    """For the /status route: is this thing actually working right now."""
    router = _get_router()
    return {
        "enabled": ENABLED,
        "ready": router is not None,
        "loading": _loading,
        "model": MODEL,
        "drop_below": DROP_BELOW,
        "min_results": MIN_RESULTS,
        "load_error": _load_error,
        "samples_logged": len(RECENT),
    }


def recent_scores(limit=20):
    """Recent (query, scores) pairs, newest first β€” the threshold-tuning data."""
    items = list(RECENT)[-limit:][::-1]
    ms = list(RECENT_MS)[-limit:][::-1]
    return [
        {"q": q, "scores": s, "ms": m} for (q, s), m in zip(items, ms)
    ]


def _question(query):
    """The typed question. `noul` returns the calibrated probability of YES.

    The query is repeated in the instructions rather than passed as context
    because that is the documented shape: state is the passage, questions carry
    the ask. Asking in the imperative ("does this passage contain the answer")
    is what makes it a relevance judgement and not a topic match.
    """
    return {
        "relevant": {
            "type": "noul",
            "instructions": (
                "Does this passage contain the information needed to answer the "
                "web search query below?\n"
                f"Query: {query}\n"
                "Answer no if the passage is about a related topic but does not "
                "actually address the query, and no if it is a navigation page, "
                "a category listing, or a site homepage."
            ),
        }
    }


def _passage(result):
    """The text the model judges. Title carries most of the topical signal;
    snippet is what we have of the page without fetching it.

    Truncated to PASSAGE_MAX_CHARS. Two reasons, and only the first was
    obvious: a snippet plus title can run past the model's context, and the
    model raises rather than truncating β€” which surfaced as a per-query
    `ok: false` with no reason, reproducible only for certain queries. The
    second is latency: a document costs ~150ms on the Space's CPU and most of
    that is attention over the input, so the cap pays for itself twice.
    """
    title = str(result.get("title") or "").strip()
    snippet = str(result.get("snippet") or "").strip()
    text = (title + "\n" + snippet).strip()
    if len(text) > PASSAGE_MAX_CHARS:
        # Title first: it survives the cut, and it is the part carrying the
        # topical signal.
        text = title[:PASSAGE_MAX_CHARS] + "\n" + snippet[: max(0, PASSAGE_MAX_CHARS - len(title))]
    return text


def _stamp(out):
    """Attach the *current* reason at the moment of return.

    Stamping once, at the top, is wrong in both directions and both happened:
    a reason captured before the load was attempted is `None` on exactly the
    path that needs one, and a reason captured before a *successful* retry
    survives into an `ok: true` response. A field that is always populated is a
    field everyone learns to ignore, which is the failure this whole mechanism
    exists to prevent.
    """
    out["no_opinion"] = _no_opinion()
    return out


def rerank(query, results, budget_ms=800):
    """Score, reorder and optionally drop `results` for `query`.

    Returns a dict; never raises. On any problem `results` comes back in its
    original order with `ok: false`, because a reranker that degrades search
    is worse than no reranker.
    """
    out = {
        "ok": False,
        "results": results,
        "model": MODEL,
        "drop_below": DROP_BELOW,
        "dropped": 0,
        "starved": False,
        "scores": [],
        "no_opinion": None,
    }
    if not ENABLED or not results:
        return _stamp(out)
    router = _get_router()
    if router is None:
        return _stamp(out)

    t0 = time.time()
    scored = []
    judged = 0
    for i, r in enumerate(results):
        if judged >= MAX_SCORED:
            # Enough judged to reorder by. The tail is kept, unscored, at the
            # end β€” we have no opinion about it and will not pretend to.
            scored.append((None, i, r))
            continue
        if judged and (time.time() - t0) * 1000 > budget_ms:
            # Over budget, but not empty-handed. A partial judgement is worth
            # more than none: the documents we did score get reordered and the
            # ones the model rejected get dropped. Only a run that scored
            # *nothing* throws its work away.
            print(
                f"[rerank] over budget ({budget_ms}ms) after {judged} of {len(results)}"
                f" β€” reranking what was judged",
                flush=True,
            )
            break
        passage = _passage(r)
        if not passage:
            # No text to judge. Keep it, unscored, and let the ordering stand.
            scored.append((None, i, r))
            continue
        try:
            res = router.predict(passage, _question(query))
            p = float(res["answers"]["relevant"]["noul"])
        except Exception as e:  # noqa: BLE001
            print(f"[rerank] score failed for result {i}: {type(e).__name__}: {e}", flush=True)
            # Fail open, whole set untouched β€” and say why. A scoring failure
            # used to be the same three words as a reranker that was not
            # installed, which is how a per-query failure hid for a day.
            out["results"] = results
            out["no_opinion"] = f"scoring failed: {type(e).__name__}: {e}"[:300]
            return out
        scored.append((p, i, r))
        judged += 1

    # Anything the loop never reached keeps its place at the end, unscored.
    if len(scored) < len(results):
        seen_idx = {i for _, i, _ in scored}
        for i, r in enumerate(results):
            if i not in seen_idx:
                scored.append((None, i, r))

    out["ms"] = int((time.time() - t0) * 1000)
    out["judged"] = judged

    if not any(p is not None for p, _, _ in scored):
        # Every result was textless, so the model was never asked anything.
        # That is a different thing from the model declining, and it is the
        # kind of difference that is invisible until it is named.
        out["no_opinion"] = "no result carried any text to judge"
        return out

    # Highest score first. Unscored results keep their relative position at the
    # end rather than being dropped: we have no opinion about them.
    scored.sort(key=lambda t: (t[0] is None, -(t[0] if t[0] is not None else 0.0), t[1]))

    judged_rows = [row for row in scored if row[0] is not None]
    below = [row for row in judged_rows if row[0] < DROP_BELOW]
    above = [row for row in judged_rows if row[0] >= DROP_BELOW]
    unscored = [row for row in scored if row[0] is None]

    if len(above) + len(unscored) < MIN_KEEP and below and above:
        # Top up from the model's own ranking β€” never from the original order.
        # Guarded on `above` being non-empty: when *nothing* cleared the bar
        # that is not a thin page, it is a model that did not understand the
        # query, and it is reported as starvation below instead of being
        # disguised as two confident results.
        need = MIN_KEEP - (len(above) + len(unscored))
        above = above + below[:need]
        below = below[need:]

    kept = [r for _, _, r in above] + [r for _, _, r in unscored]
    dropped = len(below)

    out["scores"] = [
        {"title": str(r.get("title") or "")[:80], "score": p}
        for p, _, r in scored
        if p is not None
    ]
    RECENT.append((query[:120], [p for p, _, _ in scored if p is not None]))
    # The duration is the number this stage actually needs tuning against, and
    # it is the one that was missing: a reranker that is too slow for its own
    # budget and a reranker that is not installed look identical from outside.
    RECENT_MS.append(out["ms"])

    if dropped and len(kept) < MIN_RESULTS:
        # The model rejected everything, which is a model that did not
        # understand the query rather than a set with nothing in it. Say so
        # instead of pretending.
        return _stamp(dict(out, results=results, starved=True))

    out.update(ok=True, results=kept, dropped=dropped)
    return _stamp(out)