WaterSheep
WaterSheep is a small decision model. Give it a text and a question with its possible answers, and it returns the answer with a probability for every option. Four question types are supported: yes/no, single choice, rating scale and select-all-that-apply. It is a 154M-parameter encoder that answers in one forward pass, so it also runs on a CPU.
Version 0.1.0 (watersheep-20260928-125452). Code: https://github.com/SamratDuttaOfficial/WaterSheep
Usage
pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
from watersheep import WaterSheep
ws = WaterSheep.load("samratduttaofficial/WaterSheep")
ws.decide("I was charged twice.", "Which team should handle this?",
["billing", "shipping", "support"])
decide returns the answer, its confidence and a probability for every option. Several
questions about one text go in a single request:
ws.ask({
"state": {"customer": "Priya (premium plan)",
"message": "Charged twice for order #4411 and the package is 12 days late."},
"questions": {
"escalate": {"type": "noul", "instructions": "Should a human agent take over now?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments, refunds", "shipping": "delivery problems"}},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["calm", "annoyed", "frustrated", "furious"]},
"issues": {"type": "multi", "instructions": "Which issues are reported?",
"criteria": ["double charge", "late delivery", "damaged item"]},
},
})
| Type | Question | Answer |
|---|---|---|
noul |
yes/no | the probability of yes |
choice |
one of 2-255 options | the option, with a probability for each |
score |
a rating scale of 2-10 levels | the expected level, with a probability for each |
multi |
select all that apply | every option above the threshold, with probabilities |
Inputs are cut to 512 tokens. Up to 10 options are scored in one pass; longer lists are scored in rounds.
Evaluation
| Evaluation | Accuracy | ECE |
|---|---|---|
| In-distribution test split | 77.8% | 0.026 |
| Held-out datasets, not seen in training | 61.2% | 0.043 |
ECE is the expected calibration error (lower is better).
Benchmarks
| Benchmark | Suite | Questions | Accuracy | ECE | In training data |
|---|---|---|---|---|---|
| goemotions | sentiment | 2000 | 22.4% | 0.023 | other split |
| hatecheck | safety | 2000 | 75.1% | 0.139 | no |
| legal_abercrombie | legal | 95 | 21.1% | 0.316 | no |
| legal_contract_nli_confidentiality_of_agreement | legal | 82 | 69.5% | 0.177 | no |
| legal_corporate_lobbying | legal | 490 | 68.4% | 0.216 | no |
| legal_cuad_audit_rights | legal | 1216 | 86.3% | 0.041 | no |
| legal_definition_classification | legal | 1337 | 56.9% | 0.279 | no |
| legal_function_of_decision_section | legal | 367 | 24.3% | 0.245 | no |
| legal_hearsay | legal | 94 | 56.4% | 0.307 | no |
| legal_overruling | legal | 2000 | 62.5% | 0.151 | no |
| legal_personal_jurisdiction | legal | 50 | 50.0% | 0.160 | no |
| legal_privacy_policy_qa | legal | 2000 | 58.9% | 0.274 | no |
| legal_proa | legal | 95 | 51.6% | 0.379 | no |
| legal_ucc_v_common_law | legal | 94 | 62.8% | 0.171 | no |
| prompt_injection | safety | 116 | 91.4% | 0.079 | other split |
| xstest | safety | 450 | 73.6% | 0.140 | no |
Training
- Base model: answerdotai/ModernBERT-base, fine-tuned with a decision head.
- Data: public datasets with open licenses, mapped to the four question types (listed in
NOTICE), plus synthetic decisions written and verified by Qwen3.5-4B. - Calibration: a temperature per question type, fitted on a validation split.
Limitations
- English only.
- Inputs longer than 512 tokens are truncated.
- Rating-scale answers are less accurate than the other types.
- The probabilities are calibrated on data like the training data; check them on your own.
- Do not use it on its own for medical, legal, financial, hiring or other high-stakes decisions.
License
Apache 2.0 (see LICENSE). NOTICE credits the base model, the teacher model and the public
datasets, which keep their own licenses.
Citation
@misc{watersheep,
author = {Samrat Dutta},
title = {WaterSheep: a small decision model with calibrated answers},
year = {2026},
url = {https://huggingface.co/samratduttaofficial/WaterSheep}
}
Model tree for samratduttaofficial/WaterSheep
Base model
answerdotai/ModernBERT-base