Respite-API / rerank.py
cazyundee's picture
sync: from ael backend/space/
25910d6
Raw History Blame Contribute Delete
21.2 kB
"""Relevance reranking for search results, using Laya (convaiinnovations).
Why this exists
---------------
The SearXNG bridge races several engines and hands back the first set that
looks big enough. Nothing in that path asks whether the results are *about* the
query, so the fastest engine wins — and a small independent index is fast.
Measured on production, 2026-10-02, from the route's own `source` field:
"what is the capital of Peru" -> mwmbl -> NPR transcripts about
mummy lice, an article about mosquitoes in Delhi, "The Capital Cycle".
Eight results. None of them answers the question.
"Peru capital city" -> mwmbl -> a martial arts studio
in Wellington, an NScale railroad modellers' club.
"Donald Trump AI regulations policies" -> bing -> "Donald Trump -
Wikipedia", AP's Trump hub, Reuters' Trump hub.
None of that is an outage, so nothing raised an alarm: HTTP 200, `degraded:
false`, a full result page. It is the failure mode where the search path is
technically healthy and completely useless, and a count-based "is the set big
enough" check cannot see it — a lexical overlap check cannot either, because
"Viceroyalty of Peru" really does contain the words "Peru" and "capital" and
still does not answer the question.
A cross-encoder can, because it reads the query and the passage together and
answers "does this passage answer this question" rather than "do these strings
overlap". Laya is the right size for that: a non-autoregressive System 1
decision model, 421M params in English (ModernBERT-large) and 322M for 100+
languages (mmBERT-base), one forward pass per document, Apache-2.0, calibrated
probabilities via RLCD. It never generates text, so there is nothing to parse
and nothing to hallucinate.
Design constraints, all of them learned the hard way
----------------------------------------------------
* **Fail open, always.** If the model is missing, slow, or the import breaks,
the caller gets its results back untouched. Search must not depend on a
component that is described as "kind of dumb" and can be swapped out.
* **Never return an empty page.** If the model rejects every result, that is
not a confident judgement, it is a model that did not understand the query —
so the original set is returned and `starved` is set. That flag is the signal
to retune DROP_BELOW, not a silent success.
An earlier version of this file had a floor of 3 surviving results, on the
theory that a short page is a bad page. That was wrong in the exact case
this module exists for: the mwmbl result set for "what is the capital of Peru"
is eight results of which perhaps two are worth showing, and a floor of 3
throws those two away and serves the eight instead. A short honest page beats
a full wrong one. The only floor that is genuinely wrong is zero.
* **Log every score.** The threshold is not a guess to be shipped and admired;
it is a number to be chosen from the distribution below. RECENT keeps the
last N (query, scores) so the threshold can be tuned from real traffic.
* **Deadline-bounded.** Reranking is an enhancement, never a reason to make a
user wait: the caller passes a budget and this module stops at it.
"""
import os
import threading
import time
from collections import deque
# ── configuration ───────────────────────────────────────────────────────────
# Off by default so a bad deploy cannot take search down; the route treats a
# disabled reranker as "no opinion" and returns what SearXNG gave it.
ENABLED = os.environ.get("AEL_RERANK", "1") not in ("0", "false", "no")
# "english" is 421M ModernBERT-large, 512 tokens, the strongest.
# "multilingual" is 322M mmBERT-base, ~2x faster, 100+ languages. The route
# sends whichever it wants; the default here is the stronger English one
# because Ael's corpus is overwhelmingly English, and non-English queries are
# still scored correctly by it far more often than not.
MODEL = os.environ.get("AEL_RERANK_MODEL", "english")
# Drop a result the model thinks this unlikely to answer the query.
#
# This went 0.5 -> 0.35 -> 0.15, and the direction is the finding: **the model is
# a good ranker of snippets and a bad dropper of them.** Measured on production,
# 2026-10-02:
#
# Ranked correctly, repeatedly. "capital of Peru" put "Lima" (0.838) above a
# Nat Geo photo of a Peruvian mysticete (0.042); "who won the 1876 presidential
# election" put the Wikipedia election article above a Reddit thread about a
# video game achievement. Reordering is where it earns its keep.
#
# Dropped answers it should not have. At 0.35 the threshold removed the correct
# result for four questions in a row and kept a generic page instead:
#
# "Unix program that replaced man with compressed pages" -> mandb gone,
# "Linux/Unix Tutorial" kept
# "first artificial satellite to orbit a body other than Earth" -> Luna gone,
# a Hackaday post kept
# "TLD delegated to a donut shop" -> .donuts gone, ".radio" kept
# "IUPAC systematic name of caffeine" -> the chemistry pages kept, the
# answer-bearing ones dropped
#
# The cause is not a bad threshold, it is what the model is reading: a title and
# a snippet, never the page. A snippet about mandb does not contain the words
# "compressed manual pages", so the passage looks like a non-answer even when
# the page is the answer. No threshold fixes that — the useful behaviour is to
# *rank* on what the snippet says and let the reader open the page, which is
# what the tool prompt already tells the model to do.
#
# So: drop only what the model actively rejects, and let ordering do the rest.
# Measured unrelated results sit at 0.004-0.098, so 0.15 removes the clearly
# irrelevant while leaving the ambiguous ones for the reader. The result is
# deliberately a longer page than the 0.35 experiment produced — a research
# tool that returns one confident result is a worse research tool.
DROP_BELOW = float(os.environ.get("AEL_RERANK_DROP", "0.15"))
# Never return an empty page. See the starvation note above — this is a floor of
# one, not a target: a short page is fine, an empty one is not.
MIN_RESULTS = 1
# …but one result is not a page you can research with. When the threshold would
# leave fewer than this, the *highest-scoring* results are kept instead of the
# original set: the floor must not resurrect the junk, so it tops the list back
# up from the model's own ranking rather than falling back to engine order.
#
# This is not the "floor of 3" the module's own docstring warns about. That one
# fell back to the whole original set, so a set of eight with two worth showing
# served all eight. Topping up from the ranking cannot do that: the results that
# come back are still the ones the model rated highest, just including a couple
# more of them.
MIN_KEEP = int(os.environ.get("AEL_RERANK_MIN_KEEP", "2"))
# How many results the model is allowed to judge. A 421M cross-encoder on a
# shared CPU does not score eight documents inside an 800ms budget, and the
# first version's answer to that was to throw the whole set away — which is why
# a reranker that was installed, loaded, healthy and simply too slow reported
# itself identically to one that had never run. Judging the first N and leaving
# the rest in engine order is strictly better than judging none: the survivors
# are reordered by relevance and the rejects are dropped, and the untested tail
# keeps the order the race gave it.
#
# 10, from measurement: production runs at ~150ms a document on the Space's
# CPU, and the race hands over 8. At the old cap of 6 the last two results were
# never judged, and an unjudged result is the one thing that re-admits exactly
# the junk this stage exists to remove — "capital of Peru" still came back with
# a Granta short story and a Nat Geo photo in it. The cap is now above a full
# result set, so the tail is normally empty; the budget below is what actually
# bounds the work, and this is only a backstop against a pathologically long set.
MAX_SCORED = int(os.environ.get("AEL_RERANK_MAX", "10"))
# How much text the model is asked to judge. The model's context is 512 tokens;
# a title plus a full snippet can exceed it and the model raises rather than
# truncating, which used to surface as a silent per-query failure.
PASSAGE_MAX_CHARS = int(os.environ.get("AEL_RERANK_PASSAGE_CHARS", "1200"))
# Recent (query, scores) for threshold tuning. Bounded so a long-lived Space
# cannot grow without limit.
RECENT = deque(maxlen=int(os.environ.get("AEL_RERANK_LOG", "200")))
# Wall-clock of recent reranks, same bound. This is the tuning data for
# MAX_SCORED and the budget: the scores say what to keep, this says what it
# costs to judge it.
RECENT_MS = deque(maxlen=int(os.environ.get("AEL_RERANK_LOG", "200")))
_router = None
_router_lock = threading.Lock()
_load_error = None
_load_message = None
_loading = False
_load_failed_at = 0.0
# A failed load used to latch for the lifetime of the container, which turned one
# flaky checkpoint download into a permanently disabled reranker — the Space
# would report "no opinion" for the rest of its life and nobody would know why.
# Failures are retried on a cooldown instead.
RETRY_AFTER_S = float(os.environ.get("AEL_RERANK_RETRY_S", "60"))
def _get_router():
"""Load the checkpoint once, lazily. Returns None if it cannot be loaded.
Deliberately never raises: a Space that cannot afford the model still
serves search, and the route needs to be able to ask "any opinion?"
without handling an exception.
"""
global _router, _load_error, _load_message, _loading, _load_failed_at
if _router is not None:
return _router
if _load_error is not None and (time.time() - _load_failed_at) < RETRY_AFTER_S:
return None
if _loading:
# Never block a search on a checkpoint download. The caller's deadline is
# ~1s; loading 421M parameters takes minutes the first time and is not
# something to discover inside a user's request. `warm()` exists so this
# is the rare path rather than the normal one.
return None
with _router_lock:
if _router is not None or _loading:
return _router
_loading = True
try:
from laya import Router
# max_loaded=1: one checkpoint resident, which is all this uses.
_router = Router(max_loaded=1)
_router.preload([MODEL])
_load_error = None
_load_message = None
print(f"[rerank] {MODEL} ready", flush=True)
except Exception as e: # noqa: BLE001 - reported, never raised
_load_error = f"{type(e).__name__}"
_load_failed_at = time.time()
# The message matters and nothing else about this exception does:
# "the reranker is off" is a decision, "the reranker broke because
# <this>" is a bug report. The caller redacts before logging it.
_load_message = f"{type(e).__name__}: {e}"[:300]
print(f"[rerank] disabled: cannot load laya ({_load_error}): {e}", flush=True)
finally:
_loading = False
return _router
def warm():
"""Start loading the checkpoint now, in the background.
Called once when the Space boots. Without this the first search request of
every container lifetime paid the full download-and-load cost inside an
800ms budget, so the reranker could never once have applied on a cold
container — and the only symptom was a timeout, which reads like a slow
network. Warming moves the cost to the ~2 minutes after a deploy, where
nobody is waiting.
Never raises and never blocks: a Space that cannot afford the model still
serves search.
"""
if not ENABLED or _router is not None or _loading:
return
t = threading.Thread(target=_get_router, name="rerank-warm", daemon=True)
t.start()
def _no_opinion():
"""Why this module has nothing to say, for the caller.
A reranker that is silently disabled is the worst outcome available: the
route's fail-open path cannot tell "the model declined" from "the model is
not installed", and both look like a working search that quietly does not
rerank. The first version of this stage returned `ok: false` and nothing
else, and that ambiguity cost a full session of debugging the Space's
build when the Space had been fine the whole time — the app simply had the
URL wrong and never asked.
"""
if not ENABLED:
return "disabled by config"
if _loading:
return "still loading"
if _load_error:
return _load_message or _load_error
return None
def status():
"""For the /status route: is this thing actually working right now."""
router = _get_router()
return {
"enabled": ENABLED,
"ready": router is not None,
"loading": _loading,
"model": MODEL,
"drop_below": DROP_BELOW,
"min_results": MIN_RESULTS,
"load_error": _load_error,
"samples_logged": len(RECENT),
}
def recent_scores(limit=20):
"""Recent (query, scores) pairs, newest first — the threshold-tuning data."""
items = list(RECENT)[-limit:][::-1]
ms = list(RECENT_MS)[-limit:][::-1]
return [
{"q": q, "scores": s, "ms": m} for (q, s), m in zip(items, ms)
]
def _question(query):
"""The typed question. `noul` returns the calibrated probability of YES.
The query is repeated in the instructions rather than passed as context
because that is the documented shape: state is the passage, questions carry
the ask. Asking in the imperative ("does this passage contain the answer")
is what makes it a relevance judgement and not a topic match.
"""
return {
"relevant": {
"type": "noul",
"instructions": (
"Does this passage contain the information needed to answer the "
"web search query below?\n"
f"Query: {query}\n"
"Answer no if the passage is about a related topic but does not "
"actually address the query, and no if it is a navigation page, "
"a category listing, or a site homepage."
),
}
}
def _passage(result):
"""The text the model judges. Title carries most of the topical signal;
snippet is what we have of the page without fetching it.
Truncated to PASSAGE_MAX_CHARS. Two reasons, and only the first was
obvious: a snippet plus title can run past the model's context, and the
model raises rather than truncating — which surfaced as a per-query
`ok: false` with no reason, reproducible only for certain queries. The
second is latency: a document costs ~150ms on the Space's CPU and most of
that is attention over the input, so the cap pays for itself twice.
"""
title = str(result.get("title") or "").strip()
snippet = str(result.get("snippet") or "").strip()
text = (title + "\n" + snippet).strip()
if len(text) > PASSAGE_MAX_CHARS:
# Title first: it survives the cut, and it is the part carrying the
# topical signal.
text = title[:PASSAGE_MAX_CHARS] + "\n" + snippet[: max(0, PASSAGE_MAX_CHARS - len(title))]
return text
def _stamp(out):
"""Attach the *current* reason at the moment of return.
Stamping once, at the top, is wrong in both directions and both happened:
a reason captured before the load was attempted is `None` on exactly the
path that needs one, and a reason captured before a *successful* retry
survives into an `ok: true` response. A field that is always populated is a
field everyone learns to ignore, which is the failure this whole mechanism
exists to prevent.
"""
out["no_opinion"] = _no_opinion()
return out
def rerank(query, results, budget_ms=800):
"""Score, reorder and optionally drop `results` for `query`.
Returns a dict; never raises. On any problem `results` comes back in its
original order with `ok: false`, because a reranker that degrades search
is worse than no reranker.
"""
out = {
"ok": False,
"results": results,
"model": MODEL,
"drop_below": DROP_BELOW,
"dropped": 0,
"starved": False,
"scores": [],
"no_opinion": None,
}
if not ENABLED or not results:
return _stamp(out)
router = _get_router()
if router is None:
return _stamp(out)
t0 = time.time()
scored = []
judged = 0
for i, r in enumerate(results):
if judged >= MAX_SCORED:
# Enough judged to reorder by. The tail is kept, unscored, at the
# end — we have no opinion about it and will not pretend to.
scored.append((None, i, r))
continue
if judged and (time.time() - t0) * 1000 > budget_ms:
# Over budget, but not empty-handed. A partial judgement is worth
# more than none: the documents we did score get reordered and the
# ones the model rejected get dropped. Only a run that scored
# *nothing* throws its work away.
print(
f"[rerank] over budget ({budget_ms}ms) after {judged} of {len(results)}"
f" — reranking what was judged",
flush=True,
)
break
passage = _passage(r)
if not passage:
# No text to judge. Keep it, unscored, and let the ordering stand.
scored.append((None, i, r))
continue
try:
res = router.predict(passage, _question(query))
p = float(res["answers"]["relevant"]["noul"])
except Exception as e: # noqa: BLE001
print(f"[rerank] score failed for result {i}: {type(e).__name__}: {e}", flush=True)
# Fail open, whole set untouched — and say why. A scoring failure
# used to be the same three words as a reranker that was not
# installed, which is how a per-query failure hid for a day.
out["results"] = results
out["no_opinion"] = f"scoring failed: {type(e).__name__}: {e}"[:300]
return out
scored.append((p, i, r))
judged += 1
# Anything the loop never reached keeps its place at the end, unscored.
if len(scored) < len(results):
seen_idx = {i for _, i, _ in scored}
for i, r in enumerate(results):
if i not in seen_idx:
scored.append((None, i, r))
out["ms"] = int((time.time() - t0) * 1000)
out["judged"] = judged
if not any(p is not None for p, _, _ in scored):
# Every result was textless, so the model was never asked anything.
# That is a different thing from the model declining, and it is the
# kind of difference that is invisible until it is named.
out["no_opinion"] = "no result carried any text to judge"
return out
# Highest score first. Unscored results keep their relative position at the
# end rather than being dropped: we have no opinion about them.
scored.sort(key=lambda t: (t[0] is None, -(t[0] if t[0] is not None else 0.0), t[1]))
judged_rows = [row for row in scored if row[0] is not None]
below = [row for row in judged_rows if row[0] < DROP_BELOW]
above = [row for row in judged_rows if row[0] >= DROP_BELOW]
unscored = [row for row in scored if row[0] is None]
if len(above) + len(unscored) < MIN_KEEP and below and above:
# Top up from the model's own ranking — never from the original order.
# Guarded on `above` being non-empty: when *nothing* cleared the bar
# that is not a thin page, it is a model that did not understand the
# query, and it is reported as starvation below instead of being
# disguised as two confident results.
need = MIN_KEEP - (len(above) + len(unscored))
above = above + below[:need]
below = below[need:]
kept = [r for _, _, r in above] + [r for _, _, r in unscored]
dropped = len(below)
out["scores"] = [
{"title": str(r.get("title") or "")[:80], "score": p}
for p, _, r in scored
if p is not None
]
RECENT.append((query[:120], [p for p, _, _ in scored if p is not None]))
# The duration is the number this stage actually needs tuning against, and
# it is the one that was missing: a reranker that is too slow for its own
# budget and a reranker that is not installed look identical from outside.
RECENT_MS.append(out["ms"])
if dropped and len(kept) < MIN_RESULTS:
# The model rejected everything, which is a model that did not
# understand the query rather than a set with nothing in it. Say so
# instead of pretending.
return _stamp(dict(out, results=results, starved=True))
out.update(ok=True, results=kept, dropped=dropped)
return _stamp(out)