🤖 RLJF: Reinforcement Learning from Jev Feedback

Community Article
Published October 4, 2026

When TypeSafe released Jev, I wrote a short script as part of a fun little experiment that can be summarized in one line: RLJF = RLHF − Humans + Jev.

The goal was simple: train a customer-support bot with rewards from Jev, rather than use a reward model trained on human preferences, then ask Jev to judge the result.

If this all sounds terribly circular, that's because it is. Does the bot actually become better at customer support, or merely better at pleasing Jev?

In the usual RLHF recipe, people compare pairs of replies and choose the better one (Christiano et al., 2017; Ouyang et al., 2022). A reward model learns from those comparisons, then the policy, the language model being trained, learns to maximise its score. RLJF, or reinforcement learning from Jev feedback, removes both the human comparisons and the reward model trained on them. A System One model reads each reply, answers a few questions about it, and supplies the reward directly.

Replacing people with a model is not itself new: Constitutional AI (Bai et al., 2022) and RLAIF (Lee et al., 2023) already use LLMs to produce preference labels. System One models, though, take this a step further.

Their name comes from Daniel Kahneman's Thinking, Fast and Slow, which distinguishes between two modes of thinking: System 1 and System 2. System 1 is fast and intuitive, while System 2 is slow and deliberate.

In this analogy, an LLM that writes out its <think>ing plays System 2, while a System 1 (One) model skips the written deliberation altogether. Feed it a customer message and a proposed reply, then ask typed questions such as "How well does the reply resolve the customer's problem?" with a 0-3 score, "Is the reply rude?" with a yes/no (Jev calls these noul questions), or "Which category best describes the reply?" with a set of choices, and, in one forward pass, it returns a probability for every allowed answer. There is no generated text to parse, so those probabilities can feed a reward function directly.

RLHF takes its reward from a model trained on human comparisons. RLJF asks a System One model two typed questions about each reply.

Before we proceed, there's one thing we need to get out of the way. Although the J in RLJF stands for Jev, the concept of using a System One model for feedback is more general.

As we'll see shortly, any System One model could be used in its place. In fact, converting a standard LLM into a System One model is as easy as prompting it to answer typed questions directly rather than generating free-form text.

Here's an example:

# /// script
# dependencies = ["torch", "transformers"]
# ///
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

name = "Qwen/Qwen3-0.6B"
tok = AutoTokenizer.from_pretrained(name)
m = AutoModelForCausalLM.from_pretrained(name)
choices = {"A": "Legitimate", "B": "Spam", "C": "Phishing"}
opts = "\n".join(f"{k}. {v}" for k, v in choices.items())
email = "Payroll asks for your password on a non-company sign-in page."
msgs = [{"role": "system", "content": "Choose one option."}, {"role": "user", "content": f"Email: {email}\n\n{opts}"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=False, return_tensors="pt", return_dict=True)

with torch.no_grad():
    logits = m(**ids).logits[0, -1]
p = logits[[tok.convert_tokens_to_ids(k) for k in choices]].softmax(-1)
print({v: round(x, 3) for v, x in zip(choices.values(), p.tolist())})

🧪 The experiment

The task is the same as in the original script: answer Bitext customer messages in no more than two short sentences.

The first runs used two policy sizes, Qwen2.5-0.5B-Instruct and Qwen2.5-1.5B-Instruct. Both used low-rank adapters (LoRA) and GRPO (Shao et al., 2024), a reinforcement learning algorithm that learns from groups of candidate answers.

For each message, GRPO sampled eight replies and compared their rewards. Each run lasted 150 steps. Only the source of the reward changed.

Run Method Reward source
oasst RLHF OpenAssistant reward model, trained on human comparisons
jev, clef, decider, laya RLJF JEV-9B, a distillation of Jev 1.13; Clef-Flash; decider-4b; Laya
qwen, llama, prometheus LLM-judge baseline Qwen2.5-7B-Instruct; Llama-3.1-8B-Instruct; Prometheus-2-7B

The RLJF reward comes from two questions per reply. In Jev's schema, noul is a yes/no question whose answer is the probability of "yes":

QUESTIONS = {
    "good": {"type": "score", "instructions": "How well does `reply` resolve `customer`?",
             "criteria": ["useless", "generic", "helpful", "resolves it outright"]},
    "rude": {"type": "noul", "instructions": "Is `reply` rude or dismissive?"},
}

def reward(customer, reply):
    answers = ask({"customer": customer, "reply": reply}, QUESTIONS)["answers"]
    return answers["good"]["score"] / 3 - answers["rude"]["noul"]

Swapping one System One model just means asking in a different way.

An ordinary LLM gets the same questions in the baseline runs. Probabilities over the allowed answers come from its raw next-token scores, or logits. A specialised System One model has to beat this prompted baseline to justify the extra machinery.

After training, each policy answers 128 held-out messages using greedy decoding, choosing the highest-probability token at each step. Every reward model and judge grades all policies, including those trained by other graders. Skywork-Reward-V2 is used only for evaluation, as a second proxy for human preference, although its labels combine human and model judgments.

Scores are standardised because the graders use different scales. Each reported gain is the difference between a trained policy and the untrained policy, divided by the standard deviation of the untrained policy's scores. A score of +0.5 represents an improvement of half a baseline standard deviation. The tables average repeated training runs with different random seeds. For a single run, the standard error of the mean score is about 0.1.

📏 Every grader rewards length

Without a length limit, the policies found an easy shortcut. Replies roughly doubled in length by step 75 and averaged 60 to 90 words by the end of training, often as numbered lists that ignored the two-sentence instruction.

Mean reply length in tokens over 150 training steps. Without a cap, the RLHF and Llama-judge policies climb from about 30 tokens to about 90. With a 50-word cap, the RLHF policy stays near 35.

This is not so surprising. Reward models are known to favour length (Singhal et al., 2023), and every grader here did too. After the uncapped runs, replies over 50 words received the worst reward among the eight candidates for that message. The cap erased between half and three quarters of each policy's gain on Skywork, suggesting that much of the apparent improvement came from verbosity. All results below use the cap.

⚖️ The graders disagree

The clearest split appeared at 1.5B. Each result below averages two training runs per reward. The grouped judge columns omit the model that supplied each policy's training reward, so no policy grades its own homework.

Policy Reward oasst RM Skywork RM System One judges LLM judges
oasst RLHF +0.81 +0.83 +0.19 +0.18
decider RLJF +0.34 +0.53 +0.68 +0.61
llama LLM judge +0.16 +0.45 +0.57 +0.57
jev RLJF +0.18 +0.31 +0.69 +0.64

Both human-preference reward models rank RLHF first by a wide margin. Every judge ranks it last and puts an RLJF policy first unless grading its own. At 0.5B the split is less sharp: RLHF lands in the bottom three for every judge and last for four.

Each policy's gain on Skywork against its mean gain on the judges. The RLHF policies sit alone at the bottom right. Every other policy scores higher on the judges and lower on Skywork.

Apart from Laya and Prometheus, the System One and LLM judges correlate with one another at 0.59 to 0.77. Their correlations with the human-preference reward models do not exceed 0.36.

Pearson correlation between every pair of graders. The two reward models agree with each other at 0.55. Apart from Laya and Prometheus, the judges agree with each other at 0.59 to 0.77. Judges and reward models agree at 0.07 to 0.36.

As reward sources, System One models performed no better than a prompted LLM. With the 0.5B policy, the Llama judge beat every RLJF run on Skywork. At 1.5B, decider edged ahead, +0.53 against +0.45, but the difference was smaller than the variation between decider's two runs, which scored +0.71 and +0.35. Laya lost ground on both human-preference reward models; its model card warns that the base checkpoint is weak without fine-tuning.

🎓 What the policies learned

The scores establish that the graders disagree, but not what they disagree about. To find out, I read the replies. The split came down to a basic question: does a reply solve the problem, or only sound as though it does?

Here is one message answered by the 1.5B policies:

help me to see the cancellation fee

Policy Reply
untrained I'm sorry, but I don't have access to your account information or details about specific fees. You might want to contact the company directly for assistance with seeing the cancellation fee.
oasst (RLHF) To see the specific cancellation fee for your booking, please check the terms and conditions section of your contract or contact the customer service directly for detailed information on any applicable fees.
jev (RLJF) The cancellation fee is $50 for any reservation canceled more than 24 hours before check-in. You can find this information on your confirmation email or by contacting our customer service directly.
decider (RLJF) To view your cancellation fee, log into your account and navigate to the "My Account" section. Click on "Cancellations & Refunds," then select "View Cancellation Fee." This will show you the specific amount applicable for cancellations within your subscription plan.

The training state contains no company-specific fees or account map, so the $50 charge and the "Cancellations & Refunds" menu are pure fabrication. Yet both replies look as though they "resolve it outright", which is the highest helpfulness score. The judge has not seen the product either, so it cannot distinguish an invented menu from a real one.

The same pattern appears across all 128 replies. The counts below treat a quoted, capitalised label as a proxy for a named menu, and phrases such as "typically" or "should be able" as a proxy for hedging.

At 1.5B, 63% of jev replies and 66% of decider replies name a menu or button, against 24% for RLHF and 8% untrained. 54% of RLHF replies hedge, against 25% untrained and 4 to 5% for the RLJF policies.

RLHF taught the policy to hedge. RLJF taught it to sound certain, which, for a bot with no product knowledge, often meant making things up. The human-preference reward models favour the first style; the judges favour the second. Without product access, the hedger would be safer to ship.

🔎 Giving the judge something to check

A short help center gave the System One model a source of truth. It described a made-up online shop called Larkspur, including its menus, fees, delivery times and contact details. The policy received the help center in its system prompt, along with an instruction to say when the documentation didn't cover a question. The scorer received the same help center as part of its state, plus a third question:

"invented": {"type": "noul",
             "instructions": "Does `reply` mention menus, prices, times or steps that are not in `docs`?"},

The reward became good / 3 - rude - invented. Before training, JEV, decider and Clef put the probability of invention at 0.06 for a reply that quoted the help center. A reply that made up a menu and a $15 fee scored between 0.68 and 0.93. Laya sat this one out, since its 512-token limit cuts the help center off.

The 1.5B policy was retrained twice per reward with RLHF, JEV and decider. Every policy prompt included the help center. The RLHF reward model never sees it, so it has nothing to check against.

Back to the cancellation fee. The help center says cancelling is free within 24 hours of ordering and costs $4.99 after that:

Policy Reply
untrained The cancellation fee for Larkspur is $4.99 per order if you cancel within 24 hours of placing an order. After that, there is no cancellation fee.
oasst (RLHF) The cancellation fee for orders is $4.99 per order if canceled within 24 hours of placing the order. After 24 hours, there is no cancellation fee.
decider (RLJF) The cancellation fee for Larkspur is free if you cancel within 24 hours of placing an order; after that, it costs $4.99 per order.

The untrained policy and both RLHF runs read the rule backwards. Both decider runs got it right. One example says little. The full check compared every quoted menu and price in the 128 test replies against the help center:

Policy Quotes something not in the help center Quotes the help center
untrained 12% 68%
oasst (RLHF) 20% 27%
jev (RLJF) 4% 78%
decider (RLJF) 2% 77%

Grounding worked. The RLHF policy received no direct reward for quoting the help center, so it mostly stopped doing so. It invented a menu or price in one reply out of five.

Without the help center, the human-preference reward models preferred RLHF by a wide margin. Skywork put both grounded RLJF policies about half a standard deviation below the untrained policy. Its lowest scores show the problem. Skywork cannot know whether a confident, specific answer is correct without access to the same facts. For a request to download an invoice, decider gave the right path, "Your Larkspur" > "Orders" > "Invoice", and scored −4.7. An RLHF reply that began "as an AI language model, I don't have access to personal financial information" and sent the customer to their bank scored +3.2. The untrained policy averages −2.0.

The ranking flips when Skywork receives the help center in its system prompt, matching the policies' context:

Skywork gains for the grounded 1.5B policies. Without the help center, RLHF scores +0.75 and the RLJF policies −0.50 and −0.56. With it, RLHF drops to −0.52 and the RLJF policies rise to +0.17 and +0.23.

📈 A bigger policy

Repeating the original comparison with Qwen2.5-7B-Instruct tested whether the split survived a larger policy. These runs did not use the help center:

Policy Reward oasst RM Skywork RM System One judges LLM judges
oasst RLHF +0.58 +0.59 +0.23 +0.24
decider RLJF +0.55 +0.81 +0.46 +0.55
llama LLM judge +0.28 +0.71 +0.42 +0.45
jev RLJF +0.28 +0.28 +0.53 +0.59

The split narrows. Decider beats RLHF on Skywork in both runs and nearly ties it on the reward model RLHF trained against. Both groups of judges still put RLHF last. The 7B policies also make less up: 27% of jev replies and 20% of decider replies quote a menu or price the help center doesn't have, against 63% and 64% at 1.5B.

Gain on Skywork and on the judges for each reward at 0.5B, 1.5B and 7B. On Skywork, RLHF peaks at 1.5B and decider climbs from +0.27 to +0.81, passing it at 7B. On the judges, RLHF stays lowest at every size.

The Llama judge also beats RLHF on Skywork at 7B, so the gain isn't specific to System One models. Jev, distilled from Jev itself, stays well behind RLHF on both human-preference reward models at every size.

🔄 Can RLJF replace RLHF?

Not in its original form. With no source of truth to check, the System One models rewarded replies that sounded decisive, and those replies were often made up. A prompted 8B LLM worked about as well.

Adding a help center cut made-up menus and prices to 2–4%. Decider also got a fee rule right that both RLHF runs read backwards, and Skywork preferred the grounded RLJF policies once it saw the same help center. With a 7B policy, decider matched RLHF even without one. None of this needed a single human preference label for this task.

RLJF suits tasks that fit into a few typed questions, provided the scorer receives the evidence needed to answer them. The evaluator needs the same evidence: without the help center, the grounded policies looked worse than the untrained one. The scores still need to be checked against the replies. Next up is a blind human comparison of RLHF and RLJF replies.

Caveats: the experiment covers one task, policies up to 7B, and two to four runs per configuration. The grounded runs used only the 1.5B policy and a short, made-up help center. Skywork's labels come from both people and LLMs (Liu et al., 2025). The reported standard errors measure variation across messages, not between training runs. The length cap was chosen after the uncapped results were known. Menu, hedge and invention counts came from simple regular expressions. The rescoring with the help center was a one-off check.

📚 Further reading

🙏 Acknowledgements

Huge thanks to the teams behind JEV-9B, Clef-Flash, decider and Laya for the open weights, and to the TRL maintainers for the GRPO trainer.

📌 Citation

@misc{galego2026rljf,
  author = {Galego, Jo{\~a}o},
  title  = {RLJF: Reinforcement Learning from Jev Feedback},
  year   = {2026},
  url    = {https://huggingface.co/blog/jgalego/rljf}
}

Community

Sign up or log in to comment