Spaces:
Running on Zero
Running on Zero
Download rerank.py from cazyundee/Respite-API: direct link, hf CLI and curl.
- Browser
- Download file 21.2 kB
-
https://huggingface.co/spaces/cazyundee/Respite-API/resolve/main/rerank.py
- Command line
-
hf download hf://spaces/cazyundee/Respite-API/rerank.py
-
curl -L -o rerank.py https://huggingface.co/spaces/cazyundee/Respite-API/resolve/main/rerank.py
21.2 kB
| """Relevance reranking for search results, using Laya (convaiinnovations). | |
| Why this exists | |
| --------------- | |
| The SearXNG bridge races several engines and hands back the first set that | |
| looks big enough. Nothing in that path asks whether the results are *about* the | |
| query, so the fastest engine wins — and a small independent index is fast. | |
| Measured on production, 2026-10-02, from the route's own `source` field: | |
| "what is the capital of Peru" -> mwmbl -> NPR transcripts about | |
| mummy lice, an article about mosquitoes in Delhi, "The Capital Cycle". | |
| Eight results. None of them answers the question. | |
| "Peru capital city" -> mwmbl -> a martial arts studio | |
| in Wellington, an NScale railroad modellers' club. | |
| "Donald Trump AI regulations policies" -> bing -> "Donald Trump - | |
| Wikipedia", AP's Trump hub, Reuters' Trump hub. | |
| None of that is an outage, so nothing raised an alarm: HTTP 200, `degraded: | |
| false`, a full result page. It is the failure mode where the search path is | |
| technically healthy and completely useless, and a count-based "is the set big | |
| enough" check cannot see it — a lexical overlap check cannot either, because | |
| "Viceroyalty of Peru" really does contain the words "Peru" and "capital" and | |
| still does not answer the question. | |
| A cross-encoder can, because it reads the query and the passage together and | |
| answers "does this passage answer this question" rather than "do these strings | |
| overlap". Laya is the right size for that: a non-autoregressive System 1 | |
| decision model, 421M params in English (ModernBERT-large) and 322M for 100+ | |
| languages (mmBERT-base), one forward pass per document, Apache-2.0, calibrated | |
| probabilities via RLCD. It never generates text, so there is nothing to parse | |
| and nothing to hallucinate. | |
| Design constraints, all of them learned the hard way | |
| ---------------------------------------------------- | |
| * **Fail open, always.** If the model is missing, slow, or the import breaks, | |
| the caller gets its results back untouched. Search must not depend on a | |
| component that is described as "kind of dumb" and can be swapped out. | |
| * **Never return an empty page.** If the model rejects every result, that is | |
| not a confident judgement, it is a model that did not understand the query — | |
| so the original set is returned and `starved` is set. That flag is the signal | |
| to retune DROP_BELOW, not a silent success. | |
| An earlier version of this file had a floor of 3 surviving results, on the | |
| theory that a short page is a bad page. That was wrong in the exact case | |
| this module exists for: the mwmbl result set for "what is the capital of Peru" | |
| is eight results of which perhaps two are worth showing, and a floor of 3 | |
| throws those two away and serves the eight instead. A short honest page beats | |
| a full wrong one. The only floor that is genuinely wrong is zero. | |
| * **Log every score.** The threshold is not a guess to be shipped and admired; | |
| it is a number to be chosen from the distribution below. RECENT keeps the | |
| last N (query, scores) so the threshold can be tuned from real traffic. | |
| * **Deadline-bounded.** Reranking is an enhancement, never a reason to make a | |
| user wait: the caller passes a budget and this module stops at it. | |
| """ | |
| import os | |
| import threading | |
| import time | |
| from collections import deque | |
| # ── configuration ─────────────────────────────────────────────────────────── | |
| # Off by default so a bad deploy cannot take search down; the route treats a | |
| # disabled reranker as "no opinion" and returns what SearXNG gave it. | |
| ENABLED = os.environ.get("AEL_RERANK", "1") not in ("0", "false", "no") | |
| # "english" is 421M ModernBERT-large, 512 tokens, the strongest. | |
| # "multilingual" is 322M mmBERT-base, ~2x faster, 100+ languages. The route | |
| # sends whichever it wants; the default here is the stronger English one | |
| # because Ael's corpus is overwhelmingly English, and non-English queries are | |
| # still scored correctly by it far more often than not. | |
| MODEL = os.environ.get("AEL_RERANK_MODEL", "english") | |
| # Drop a result the model thinks this unlikely to answer the query. | |
| # | |
| # This went 0.5 -> 0.35 -> 0.15, and the direction is the finding: **the model is | |
| # a good ranker of snippets and a bad dropper of them.** Measured on production, | |
| # 2026-10-02: | |
| # | |
| # Ranked correctly, repeatedly. "capital of Peru" put "Lima" (0.838) above a | |
| # Nat Geo photo of a Peruvian mysticete (0.042); "who won the 1876 presidential | |
| # election" put the Wikipedia election article above a Reddit thread about a | |
| # video game achievement. Reordering is where it earns its keep. | |
| # | |
| # Dropped answers it should not have. At 0.35 the threshold removed the correct | |
| # result for four questions in a row and kept a generic page instead: | |
| # | |
| # "Unix program that replaced man with compressed pages" -> mandb gone, | |
| # "Linux/Unix Tutorial" kept | |
| # "first artificial satellite to orbit a body other than Earth" -> Luna gone, | |
| # a Hackaday post kept | |
| # "TLD delegated to a donut shop" -> .donuts gone, ".radio" kept | |
| # "IUPAC systematic name of caffeine" -> the chemistry pages kept, the | |
| # answer-bearing ones dropped | |
| # | |
| # The cause is not a bad threshold, it is what the model is reading: a title and | |
| # a snippet, never the page. A snippet about mandb does not contain the words | |
| # "compressed manual pages", so the passage looks like a non-answer even when | |
| # the page is the answer. No threshold fixes that — the useful behaviour is to | |
| # *rank* on what the snippet says and let the reader open the page, which is | |
| # what the tool prompt already tells the model to do. | |
| # | |
| # So: drop only what the model actively rejects, and let ordering do the rest. | |
| # Measured unrelated results sit at 0.004-0.098, so 0.15 removes the clearly | |
| # irrelevant while leaving the ambiguous ones for the reader. The result is | |
| # deliberately a longer page than the 0.35 experiment produced — a research | |
| # tool that returns one confident result is a worse research tool. | |
| DROP_BELOW = float(os.environ.get("AEL_RERANK_DROP", "0.15")) | |
| # Never return an empty page. See the starvation note above — this is a floor of | |
| # one, not a target: a short page is fine, an empty one is not. | |
| MIN_RESULTS = 1 | |
| # …but one result is not a page you can research with. When the threshold would | |
| # leave fewer than this, the *highest-scoring* results are kept instead of the | |
| # original set: the floor must not resurrect the junk, so it tops the list back | |
| # up from the model's own ranking rather than falling back to engine order. | |
| # | |
| # This is not the "floor of 3" the module's own docstring warns about. That one | |
| # fell back to the whole original set, so a set of eight with two worth showing | |
| # served all eight. Topping up from the ranking cannot do that: the results that | |
| # come back are still the ones the model rated highest, just including a couple | |
| # more of them. | |
| MIN_KEEP = int(os.environ.get("AEL_RERANK_MIN_KEEP", "2")) | |
| # How many results the model is allowed to judge. A 421M cross-encoder on a | |
| # shared CPU does not score eight documents inside an 800ms budget, and the | |
| # first version's answer to that was to throw the whole set away — which is why | |
| # a reranker that was installed, loaded, healthy and simply too slow reported | |
| # itself identically to one that had never run. Judging the first N and leaving | |
| # the rest in engine order is strictly better than judging none: the survivors | |
| # are reordered by relevance and the rejects are dropped, and the untested tail | |
| # keeps the order the race gave it. | |
| # | |
| # 10, from measurement: production runs at ~150ms a document on the Space's | |
| # CPU, and the race hands over 8. At the old cap of 6 the last two results were | |
| # never judged, and an unjudged result is the one thing that re-admits exactly | |
| # the junk this stage exists to remove — "capital of Peru" still came back with | |
| # a Granta short story and a Nat Geo photo in it. The cap is now above a full | |
| # result set, so the tail is normally empty; the budget below is what actually | |
| # bounds the work, and this is only a backstop against a pathologically long set. | |
| MAX_SCORED = int(os.environ.get("AEL_RERANK_MAX", "10")) | |
| # How much text the model is asked to judge. The model's context is 512 tokens; | |
| # a title plus a full snippet can exceed it and the model raises rather than | |
| # truncating, which used to surface as a silent per-query failure. | |
| PASSAGE_MAX_CHARS = int(os.environ.get("AEL_RERANK_PASSAGE_CHARS", "1200")) | |
| # Recent (query, scores) for threshold tuning. Bounded so a long-lived Space | |
| # cannot grow without limit. | |
| RECENT = deque(maxlen=int(os.environ.get("AEL_RERANK_LOG", "200"))) | |
| # Wall-clock of recent reranks, same bound. This is the tuning data for | |
| # MAX_SCORED and the budget: the scores say what to keep, this says what it | |
| # costs to judge it. | |
| RECENT_MS = deque(maxlen=int(os.environ.get("AEL_RERANK_LOG", "200"))) | |
| _router = None | |
| _router_lock = threading.Lock() | |
| _load_error = None | |
| _load_message = None | |
| _loading = False | |
| _load_failed_at = 0.0 | |
| # A failed load used to latch for the lifetime of the container, which turned one | |
| # flaky checkpoint download into a permanently disabled reranker — the Space | |
| # would report "no opinion" for the rest of its life and nobody would know why. | |
| # Failures are retried on a cooldown instead. | |
| RETRY_AFTER_S = float(os.environ.get("AEL_RERANK_RETRY_S", "60")) | |
| def _get_router(): | |
| """Load the checkpoint once, lazily. Returns None if it cannot be loaded. | |
| Deliberately never raises: a Space that cannot afford the model still | |
| serves search, and the route needs to be able to ask "any opinion?" | |
| without handling an exception. | |
| """ | |
| global _router, _load_error, _load_message, _loading, _load_failed_at | |
| if _router is not None: | |
| return _router | |
| if _load_error is not None and (time.time() - _load_failed_at) < RETRY_AFTER_S: | |
| return None | |
| if _loading: | |
| # Never block a search on a checkpoint download. The caller's deadline is | |
| # ~1s; loading 421M parameters takes minutes the first time and is not | |
| # something to discover inside a user's request. `warm()` exists so this | |
| # is the rare path rather than the normal one. | |
| return None | |
| with _router_lock: | |
| if _router is not None or _loading: | |
| return _router | |
| _loading = True | |
| try: | |
| from laya import Router | |
| # max_loaded=1: one checkpoint resident, which is all this uses. | |
| _router = Router(max_loaded=1) | |
| _router.preload([MODEL]) | |
| _load_error = None | |
| _load_message = None | |
| print(f"[rerank] {MODEL} ready", flush=True) | |
| except Exception as e: # noqa: BLE001 - reported, never raised | |
| _load_error = f"{type(e).__name__}" | |
| _load_failed_at = time.time() | |
| # The message matters and nothing else about this exception does: | |
| # "the reranker is off" is a decision, "the reranker broke because | |
| # <this>" is a bug report. The caller redacts before logging it. | |
| _load_message = f"{type(e).__name__}: {e}"[:300] | |
| print(f"[rerank] disabled: cannot load laya ({_load_error}): {e}", flush=True) | |
| finally: | |
| _loading = False | |
| return _router | |
| def warm(): | |
| """Start loading the checkpoint now, in the background. | |
| Called once when the Space boots. Without this the first search request of | |
| every container lifetime paid the full download-and-load cost inside an | |
| 800ms budget, so the reranker could never once have applied on a cold | |
| container — and the only symptom was a timeout, which reads like a slow | |
| network. Warming moves the cost to the ~2 minutes after a deploy, where | |
| nobody is waiting. | |
| Never raises and never blocks: a Space that cannot afford the model still | |
| serves search. | |
| """ | |
| if not ENABLED or _router is not None or _loading: | |
| return | |
| t = threading.Thread(target=_get_router, name="rerank-warm", daemon=True) | |
| t.start() | |
| def _no_opinion(): | |
| """Why this module has nothing to say, for the caller. | |
| A reranker that is silently disabled is the worst outcome available: the | |
| route's fail-open path cannot tell "the model declined" from "the model is | |
| not installed", and both look like a working search that quietly does not | |
| rerank. The first version of this stage returned `ok: false` and nothing | |
| else, and that ambiguity cost a full session of debugging the Space's | |
| build when the Space had been fine the whole time — the app simply had the | |
| URL wrong and never asked. | |
| """ | |
| if not ENABLED: | |
| return "disabled by config" | |
| if _loading: | |
| return "still loading" | |
| if _load_error: | |
| return _load_message or _load_error | |
| return None | |
| def status(): | |
| """For the /status route: is this thing actually working right now.""" | |
| router = _get_router() | |
| return { | |
| "enabled": ENABLED, | |
| "ready": router is not None, | |
| "loading": _loading, | |
| "model": MODEL, | |
| "drop_below": DROP_BELOW, | |
| "min_results": MIN_RESULTS, | |
| "load_error": _load_error, | |
| "samples_logged": len(RECENT), | |
| } | |
| def recent_scores(limit=20): | |
| """Recent (query, scores) pairs, newest first — the threshold-tuning data.""" | |
| items = list(RECENT)[-limit:][::-1] | |
| ms = list(RECENT_MS)[-limit:][::-1] | |
| return [ | |
| {"q": q, "scores": s, "ms": m} for (q, s), m in zip(items, ms) | |
| ] | |
| def _question(query): | |
| """The typed question. `noul` returns the calibrated probability of YES. | |
| The query is repeated in the instructions rather than passed as context | |
| because that is the documented shape: state is the passage, questions carry | |
| the ask. Asking in the imperative ("does this passage contain the answer") | |
| is what makes it a relevance judgement and not a topic match. | |
| """ | |
| return { | |
| "relevant": { | |
| "type": "noul", | |
| "instructions": ( | |
| "Does this passage contain the information needed to answer the " | |
| "web search query below?\n" | |
| f"Query: {query}\n" | |
| "Answer no if the passage is about a related topic but does not " | |
| "actually address the query, and no if it is a navigation page, " | |
| "a category listing, or a site homepage." | |
| ), | |
| } | |
| } | |
| def _passage(result): | |
| """The text the model judges. Title carries most of the topical signal; | |
| snippet is what we have of the page without fetching it. | |
| Truncated to PASSAGE_MAX_CHARS. Two reasons, and only the first was | |
| obvious: a snippet plus title can run past the model's context, and the | |
| model raises rather than truncating — which surfaced as a per-query | |
| `ok: false` with no reason, reproducible only for certain queries. The | |
| second is latency: a document costs ~150ms on the Space's CPU and most of | |
| that is attention over the input, so the cap pays for itself twice. | |
| """ | |
| title = str(result.get("title") or "").strip() | |
| snippet = str(result.get("snippet") or "").strip() | |
| text = (title + "\n" + snippet).strip() | |
| if len(text) > PASSAGE_MAX_CHARS: | |
| # Title first: it survives the cut, and it is the part carrying the | |
| # topical signal. | |
| text = title[:PASSAGE_MAX_CHARS] + "\n" + snippet[: max(0, PASSAGE_MAX_CHARS - len(title))] | |
| return text | |
| def _stamp(out): | |
| """Attach the *current* reason at the moment of return. | |
| Stamping once, at the top, is wrong in both directions and both happened: | |
| a reason captured before the load was attempted is `None` on exactly the | |
| path that needs one, and a reason captured before a *successful* retry | |
| survives into an `ok: true` response. A field that is always populated is a | |
| field everyone learns to ignore, which is the failure this whole mechanism | |
| exists to prevent. | |
| """ | |
| out["no_opinion"] = _no_opinion() | |
| return out | |
| def rerank(query, results, budget_ms=800): | |
| """Score, reorder and optionally drop `results` for `query`. | |
| Returns a dict; never raises. On any problem `results` comes back in its | |
| original order with `ok: false`, because a reranker that degrades search | |
| is worse than no reranker. | |
| """ | |
| out = { | |
| "ok": False, | |
| "results": results, | |
| "model": MODEL, | |
| "drop_below": DROP_BELOW, | |
| "dropped": 0, | |
| "starved": False, | |
| "scores": [], | |
| "no_opinion": None, | |
| } | |
| if not ENABLED or not results: | |
| return _stamp(out) | |
| router = _get_router() | |
| if router is None: | |
| return _stamp(out) | |
| t0 = time.time() | |
| scored = [] | |
| judged = 0 | |
| for i, r in enumerate(results): | |
| if judged >= MAX_SCORED: | |
| # Enough judged to reorder by. The tail is kept, unscored, at the | |
| # end — we have no opinion about it and will not pretend to. | |
| scored.append((None, i, r)) | |
| continue | |
| if judged and (time.time() - t0) * 1000 > budget_ms: | |
| # Over budget, but not empty-handed. A partial judgement is worth | |
| # more than none: the documents we did score get reordered and the | |
| # ones the model rejected get dropped. Only a run that scored | |
| # *nothing* throws its work away. | |
| print( | |
| f"[rerank] over budget ({budget_ms}ms) after {judged} of {len(results)}" | |
| f" — reranking what was judged", | |
| flush=True, | |
| ) | |
| break | |
| passage = _passage(r) | |
| if not passage: | |
| # No text to judge. Keep it, unscored, and let the ordering stand. | |
| scored.append((None, i, r)) | |
| continue | |
| try: | |
| res = router.predict(passage, _question(query)) | |
| p = float(res["answers"]["relevant"]["noul"]) | |
| except Exception as e: # noqa: BLE001 | |
| print(f"[rerank] score failed for result {i}: {type(e).__name__}: {e}", flush=True) | |
| # Fail open, whole set untouched — and say why. A scoring failure | |
| # used to be the same three words as a reranker that was not | |
| # installed, which is how a per-query failure hid for a day. | |
| out["results"] = results | |
| out["no_opinion"] = f"scoring failed: {type(e).__name__}: {e}"[:300] | |
| return out | |
| scored.append((p, i, r)) | |
| judged += 1 | |
| # Anything the loop never reached keeps its place at the end, unscored. | |
| if len(scored) < len(results): | |
| seen_idx = {i for _, i, _ in scored} | |
| for i, r in enumerate(results): | |
| if i not in seen_idx: | |
| scored.append((None, i, r)) | |
| out["ms"] = int((time.time() - t0) * 1000) | |
| out["judged"] = judged | |
| if not any(p is not None for p, _, _ in scored): | |
| # Every result was textless, so the model was never asked anything. | |
| # That is a different thing from the model declining, and it is the | |
| # kind of difference that is invisible until it is named. | |
| out["no_opinion"] = "no result carried any text to judge" | |
| return out | |
| # Highest score first. Unscored results keep their relative position at the | |
| # end rather than being dropped: we have no opinion about them. | |
| scored.sort(key=lambda t: (t[0] is None, -(t[0] if t[0] is not None else 0.0), t[1])) | |
| judged_rows = [row for row in scored if row[0] is not None] | |
| below = [row for row in judged_rows if row[0] < DROP_BELOW] | |
| above = [row for row in judged_rows if row[0] >= DROP_BELOW] | |
| unscored = [row for row in scored if row[0] is None] | |
| if len(above) + len(unscored) < MIN_KEEP and below and above: | |
| # Top up from the model's own ranking — never from the original order. | |
| # Guarded on `above` being non-empty: when *nothing* cleared the bar | |
| # that is not a thin page, it is a model that did not understand the | |
| # query, and it is reported as starvation below instead of being | |
| # disguised as two confident results. | |
| need = MIN_KEEP - (len(above) + len(unscored)) | |
| above = above + below[:need] | |
| below = below[need:] | |
| kept = [r for _, _, r in above] + [r for _, _, r in unscored] | |
| dropped = len(below) | |
| out["scores"] = [ | |
| {"title": str(r.get("title") or "")[:80], "score": p} | |
| for p, _, r in scored | |
| if p is not None | |
| ] | |
| RECENT.append((query[:120], [p for p, _, _ in scored if p is not None])) | |
| # The duration is the number this stage actually needs tuning against, and | |
| # it is the one that was missing: a reranker that is too slow for its own | |
| # budget and a reranker that is not installed look identical from outside. | |
| RECENT_MS.append(out["ms"]) | |
| if dropped and len(kept) < MIN_RESULTS: | |
| # The model rejected everything, which is a model that did not | |
| # understand the query rather than a set with nothing in it. Say so | |
| # instead of pretending. | |
| return _stamp(dict(out, results=results, starved=True)) | |
| out.update(ok=True, results=kept, dropped=dropped) | |
| return _stamp(out) | |