Rei v1

Rei is a small, non-generative decision model for research agents. It never writes text. You give it a state (JSON, such as a search result) and typed questions, and it returns one of the answers you defined, with calibrated probabilities. It's built for the decisions a research agent makes around web search: which results are relevant, which pages to open, which sources to trust, when results disagree, and when to search again.

  • 271M parameters: embeddinggemma-2's text encoder (the 134M-parameter vocabulary table is frozen) plus a small scoring head.
  • 13 languages: English, Chinese, Icelandic, Spanish, Hindi, Arabic, Portuguese, French, Russian, Japanese, German, Korean, Indonesian.
  • Trained on 82,511 decisions about real Brave Search results, labeled by a DeepSeek V4.1 Flash ensemble.
  • About 26 ms per search result (two questions) on a GPU (RTX 5090). It also runs on CPU, much more slowly.

Usage

rei.py in this repo is all you need besides four common packages:

pip install torch "transformers>=5.19" safetensors huggingface_hub
hf download gnukeith/rei-v1 rei.py --local-dir .
from rei import Rei, QUESTIONS

rei = Rei.from_pretrained("gnukeith/rei-v1")  # GPU if available, otherwise CPU
result = {"title": "...", "url": "...", "site": "...", "snippet": "..."}
rei.decide(
    {"research_question": "...", "search_query": "...", "result": result},
    {
        "relevance": QUESTIONS["relevance"],
        "open": QUESTIONS["worth_opening"],
        "paywall": {"type": "noul", "statement": "The page is likely behind a paywall."},  # your own question
    },
)

rei.decide_many([(state, questions), ...]) answers many states (say, every result of a search) in one batched pass.

Output

Real output for a hand-written Date A Live wiki result, on the question "Who voices Kurumi Tokisaki in Japanese and English?":

{
 "relevance": {"score": 3, "probabilities": {"0": 0.013, "1": 0.040, "2": 0.294, "3": 0.653}, "confidence": 0.653},
 "open":      {"noul": 0.968},
 "source":    {"choice": "reference", "probabilities": {"reference": 0.836, "community": 0.036, "...": "..."}, "confidence": 0.836}
}
  • choice: one of your option keys, with a probability for each.
  • score: a value on your scale, with a probability for each value.
  • noul: the probability that your statement is true.

Built-in tasks

The built-in tasks (QUESTIONS in rei.py) are what Rei knows best. Give each the state shape it was trained on. A result is {title, url, site, snippet}, optionally with extra_snippets and age.

Task Type State
relevance score 0-3 research_question, search_query, result
worth_opening noul research_question, search_query, result
source_type choice of 9 result
needs_fresh noul research_question
search_type choice of 8 research_question
next_step choice of 4 research_question, search_query, results (up to 5)
same_info noul research_question, result_a, result_b
claim_check choice of 3 claim, source (title, url, site, text)

Your own questions work too, but less reliably (72% agreement on held-out custom questions). Put what's being judged in the state and keep the question about the decision. A claim belongs in the state as claim, not inside the question text.

Evaluation: agreement with the teacher

Accuracy against the teacher ensemble's answer on 3,968 test items. Test items come from research questions that never appeared in training. This measures how well Rei reproduces the teacher's judgement, not ground truth.

Test: 76.7% against a 42.8% majority baseline. Calibration error (ECE) is 0.031. Validation: 79.0%.

Task items accuracy majority baseline
claim_check 1,085 81.4% 33.9%
custom questions 421 72.0% 15.0%
needs_fresh 181 82.9% 60.2%
next_step 177 76.3% 56.5%
relevance 1,041 68.5% (mean error 0.33 points) 48.2%
same_info 180 88.3% 86.1%
search_type 173 76.9% 31.8%
source_type 347 80.1% 18.2%
worth_opening 363 79.6% 78.5%

The noul tasks are mostly one class, so accuracy understates them. AUC (how well P(true) ranks true above false): needs_fresh 0.91, same_info 0.83, custom 0.83, worth_opening 0.76.

By domain: anime 78.4%, programming 78.0%, academic 77.8%, general 74.4%.

By language: Chinese 80.7%, French 80.4%, German 80.3%, Spanish 79.8%, Arabic 79.5%, Japanese 78.4%, English 77.0%, Hindi 75.7%, Korean 75.7%, Portuguese 75.5%, Russian 75.0%, Indonesian 74.4%, Icelandic 69.6%.

Benchmark: Brave + Claude, with and without Rei

Does Rei help a Claude research pipeline? Held-out research questions were answered from their real Brave results, about 18 per question (snippets, from a query plus a cross-check query). Four pipelines were compared:

  • all: the model reads every result.
  • Rei filter: the model reads only results Rei rates relevant, with near-duplicates removed.
  • Rei's top 6: Rei's six best results.
  • Brave's top 6: the search engine's first six.

A blind Claude Opus 5.5 judge compared answers pairwise, in both orders, with every result in view. Score is the share of judgments won, ties counting half; 0.50 means no difference. 95% intervals are in brackets. Opus: 80 questions. Sonnet and Haiku: 60.

Opus 5.5 Sonnet 5.5 Haiku 5.5
Rei's top 6 vs Brave's top 6 (same cost) 0.67 (0.58-0.76) 0.65 (0.54-0.77) 0.71 (0.60-0.82)
Rei filter vs all results 0.35 (0.27-0.43) 0.40 (0.30-0.51) 0.38 (0.27-0.48)
Rei filter vs Opus 5.5 on all results - 0.08 (0.03-0.13) 0.01 (0.00-0.03)
Answer cost per question: all / Rei filter / 6 results $0.077 / $0.071 / $0.046 $0.029 / $0.026 / $0.016 $0.0020 / $0.0018 / $0.0012

What it shows:

  • Rei picks better sources than search ranking. At the same cost, answers built from Rei's top 6 beat answers from Brave's top 6 for every model size, and the gain is largest for the smallest model. Most of the gain is completeness (0.67-0.73), with smaller gains in correctness and grounding.
  • With snippets, reading everything still wins. About 18 snippets are only ~9k tokens, and the model's own output is most of the cost. Rei's filter saves 8-10% but loses some completeness. Any 6-result budget loses to reading everything (Opus: 0.01).
  • Rei doesn't let a smaller model replace Opus here. Sonnet and Haiku with Rei's sources lose clearly to Opus with everything. The judge is Opus itself, which may favour its own style. The within-model rows don't have that issue.

Not tested yet: the expensive step in deep research, opening full pages (5-20k tokens each), where choosing 6 pages instead of 18 would cut about two thirds of the reading cost. Rei's advantage at a fixed budget suggests that's where it pays off.

Speed

per result per question (~18 results, 2 questions each)
GPU, RTX 5090 26 ms 0.46 s (8 GiB at the default batch size)
CPU, 16 threads ~1.2 s ~17-28 s

Each candidate answer is encoded together with the state, so cost grows with the number of candidates and the state's length.

Training

  • Data: 82,511 decisions in four domains: general research, anime and manga, programming, and academic (high school to postdoc). The states are real Brave Search results.
    • The research questions span a grid of domains, intents and difficulties, with 20% cross-lingual searches.
    • Hard cases were constructed from real data: off-topic results from related questions, likely-duplicate pairs, results that disagree, and six kinds of generated claims.
  • Labels: three DeepSeek V4.1 Flash votes with shuffled options.
    • A split vote gets two more votes plus a maximum-effort judge.
    • The judge re-checks 15% of unanimous items and disagrees on 2-3% of them.
    • Labels keep the vote distribution. Generated items are kept only when the ensemble independently reaches the intended answer.
  • Training:
    • 3 epochs on 74,855 items, 2 h 55 min on one RTX 5090.
    • Soft-label cross-entropy, with rare answers weighted up.
    • Paraphrased questions, and randomly renamed option keys so Rei reads the descriptions.
    • Every question type, including noul, is scored as a choice between candidates through one head.
    • Temperature scaling on validation.
  • Not used: Claude outputs (the benchmark answers are evaluation only). The training data itself isn't released.

Limitations

  • Rei reproduces a teacher's judgement (DeepSeek V4.1 Flash), errors included. It doesn't know facts; it judges what the state says.
  • Icelandic trails the other languages by 5-11 points.
  • Custom questions, especially ones unlike the training tasks, are much less reliable than the built-in tasks.
  • On CPU, scoring a full result list takes tens of seconds.
  • The benchmark covers snippets only, with an LLM judge, on 60-80 questions per model.

License

Apache-2.0, the license of the base model google/embeddinggemma-2.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gnukeith/rei-v1

Finetuned
(34)
this model