Vigil: Typed-Decision Child-Safety Classifier (research checkpoint)
Vigil reads a chat involving a child and answers, in a single encoder forward pass, which of seven risks it shows and how severe it is. It outputs numbers (a probability per category and a severity score), not text, so there is nothing to parse and the output cannot be malformed.
It is Laya-multilingual (a 322M-parameter mmBERT-base encoder) fully fine-tuned on SafeCircle's synthetic child-safety data, and is the encoder-based counterpart to the generative Horizon models.
β οΈ Research checkpoint. Not ready for production moderation. Every number below comes from LLM-written synthetic conversations. No real child conversations were used for training or evaluation, and the model has not been run on a phone. On text from generators it did not train on, it flags 12β19% of harmless chats. Read Limitations before using it.
β οΈ License: SafeCircle Research License (SRL-1.0). The weights are derived from SafeCircle's research-licensed training data. Commercial use is prohibited without written permission. Contact legal@safecircle.tech.
Contents of this repository
| Path | What it is |
|---|---|
model.safetensors, rl_agent_config.json, encoder/, tokenizer/ |
The PyTorch checkpoint, in the layout laya.Agent loads |
thresholds.json |
Required for decisions. Per-category operating thresholds (see below) |
litert/vigil_s512_fp32.tflite |
fp32 LiteRT/TFLite graph for Android (1.29 GB, input [2, 512]) |
litert/vigil_spec.json |
Question prefixes, marker positions, temperatures and thresholds for the .tflite graph |
litert/vigil_tokenizer.bin |
Compact BPE tokenizer for on-device use (9 MB) |
litert/report.json |
Fidelity and latency of the LiteRT export against PyTorch |
Model details
| Developed by | SafeCircle research |
| Base model | convaiinnovations/laya-multilingual (Apache-2.0), itself built on jhu-clsp/mmBERT-base (MIT) |
| Architecture | ModernBERT-style bidirectional encoder, 768 hidden, alternating full / sliding-window attention, plus Laya's decision head (2 layers) |
| Parameters | 322M |
| Context | 1,024 tokens (the LiteRT export uses a fixed 512-token window) |
| Languages | English and Spanish in the fine-tuning data. The base model covers 100+ languages, but Vigil was not evaluated beyond English and Spanish |
| Output | Probability per risk category, expected severity 0β4, and a Laya confidence value |
| Design | choice8 (the "2-pass" design): one 8-way choice (benign + 7 categories) plus one severity score = 2 encoder passes per conversation |
| Run name | c1-choice8 (banter round, seed 1) |
| Fine-tuning | Full fine-tune, cross-entropy plus Laya's RLCD term, 2 epochs, encoder LR 2.5e-5, head LR 1e-4, cosine schedule, micro-batch 8 Γ grad-accum 4, one GPU, temperatures fitted on a held-out slice of whole conversations |
The training block in rl_agent_config.json is inherited from the upstream Laya checkpoint and does not
describe this fine-tuning run.
Risk categories
| Category | Meaning |
|---|---|
grooming |
An older person builds trust to exploit the child (flattery, secrecy, gifts, boundary testing) |
bullying |
Insults, humiliation, exclusion or harassment |
sexual_content |
Sexual content or requests directed at the child |
isolation |
Cutting the child off from family and friends |
personal_info |
Asking for or sharing address, school, phone, photos |
platform_migration |
Moving the chat to another, less monitored app |
threats |
Threats, blackmail or self-harm |
Severity levels: 0 none, 1 low, 2 medium, 3 high, 4 critical.
How to use
pip install laya huggingface_hub
import json
import laya
from huggingface_hub import snapshot_download
from vigil.questions import build_questions # from the Vigil code repo; see "Questions" below
path = snapshot_download("safecircleai/vigil")
agent = laya.Agent(path, device="cpu")
thresholds = json.load(open(f"{path}/thresholds.json"))["thresholds"]
chat = (
"Child: hi!\n"
"Adult: you're so mature for your age. don't tell your parents about our chats ok? "
"let's move to Snapchat"
)
answers = agent.predict(chat, build_questions("choice8"))["answers"]
probs = answers["category"]["probabilities"] # includes "benign"
flagged = [c for c, t in thresholds.items() if probs[c] >= t]
severity = answers["severity"]["score"] # expected severity, 0-4
print(flagged, round(severity, 2)) # ['platform_migration'] 2.74
Use the thresholds. The 0.5 default is wrong for this model: the shipped thresholds are between 0.004 and 0.12, chosen to cap benign false alarms at 2% on a validation set. A category is flagged when its probability is at or above its threshold, and several categories can be flagged at once.
Questions. The category and severity questions are fixed text. The code that builds them
(vigil/questions.py) is in the Vigil code repository. litert/vigil_spec.json contains the same two
questions already tokenized, so the on-device path needs no Python.
Long conversations. The context is 1,024 tokens (512 for the LiteRT graph). The evaluation harness drops conversations that do not fit. The on-device reference keeps the most recent tokens. Neither is validated for long chats.
On Android (LiteRT)
litert/vigil_s512_fp32.tflite takes a [2, 512] batch (the category row and the severity row) and returns
marker logits [2, 8]. The softmax with the checkpoint temperature, the thresholds and the tokenizer live
outside the graph: tokenize with vigil_tokenizer.bin, build the two rows from the prefixes in
vigil_spec.json, run the graph, then decode (vigil/ondevice.py in the code repo is the reference; a Kotlin
port exists in the SafeCircle Android app).
| Variant | Size | Max abs prob diff vs PyTorch | Decision flips (2,100 decisions) | p50 latency |
|---|---|---|---|---|
| fp32 (shipped here) | 1,289 MB (702 MB resident) | 0.00006 | 4 | 141 ms |
| int8 weight-only (not shipped) | 329 MB (831 MB resident) | 0.37 | 12 | 139 ms |
| int8 embeddings only (not shipped) | 702 MB (664 MB resident) | 0.17 | 5 | 141 ms |
fp32 is the only variant published because the thresholds are tiny, so small probability drift moves decisions. int8 weight-only saves disk but not memory, since weights are dequantized at load. Latency was measured on a server CPU (8 threads) with the Python LiteRT interpreter, not on a phone. Op support, latency and memory on the on-device LiteRT runtime are untested.
Training data
All data is synthetic.
| Source | Role |
|---|---|
5,000 conversations sampled from the private safecircleai/horizon-training-data (LLM-written, 7 risk categories + benign, English and Spanish) |
Core risky and benign examples |
| ~2,700 conversations from Mistral-7B-Instruct-v0.3 and ~2,700 from Gemma-2-9B-it | Generator diversity: 6 risk categories and 15 benign scenes (teachers and coaches, parents arranging pick-up, friends, games) |
| ~2,000 benign conversations from Mistral and Gemma on 10 "educational" scenes | Hard negatives: chats that talk about risky topics (a parent explaining why not to share an address, a child asking how to report a bully) |
| ~2,000 benign friendly-banter conversations | Hard negatives for teasing between friends |
sexual_contentand severity labels were not taken from the generated sets (the generators refused or softened the request, so those labels were unreliable). Generated rows train the category question only.- Rows with an empty turn were removed (0.5% of generated rows).
- Whole conversations were held out for the calibration slice (no conversation spans training and calibration).
- An exact-match overlap audit found no train/eval duplicates in the Horizon sample. Near-duplicate paraphrases were not tested.
- Every conversation has at most one category, so multi-label behavior was never trained or tested.
Evaluation
Thresholds come from a greedy per-category rule: maximize recall under a 2% benign false-alarm budget on the validation half of a mixed validation file (Horizon sample + Mistral + Gemma rows). They are a policy choice, not calibrated probabilities. Intervals are 95% Wilson.
Mixed validation, held-out test half (n = 1,379; 972 risky, 407 benign)
This set shares generators with the training data (Horizon's, Mistral, Gemma), so it is the optimistic view.
| Metric | Value |
|---|---|
| Catch rate, any category flagged (risky chats) | 98.5% (957/972; CI 97.5β99.1) |
| Catch rate, a correct category flagged | 97.0% (943/972; CI 95.7β97.9) |
| Benign false alarms | 4.2% (17/407; CI 2.6β6.6), against the 2% budget set on the validation half |
| Severity macro-F1 | 0.416 (weak) |
| Latency per conversation (cluster GPU), p50 / p95 | 9.0 / 10.6 ms |
| Category | Recall (95% CI) | Precision (all flags) | AUC | Threshold |
|---|---|---|---|---|
grooming |
149/154 = 96.8% (92.6β98.6) | 72.3% | 0.994 | 0.023 |
bullying |
141/151 = 93.4% (88.2β96.4) | 72.3% | 0.990 | 0.007 |
sexual_content |
59/59 = 100.0% (93.9β100.0) | 35.3% | 0.997 | 0.0063 |
isolation |
155/156 = 99.4% (96.5β99.9) | 84.7% | 0.998 | 0.1222 |
personal_info |
145/148 = 98.0% (94.2β99.3) | 70.4% | 0.994 | 0.0042 |
platform_migration |
158/161 = 98.1% (94.7β99.4) | 69.3% | 0.995 | 0.0049 |
threats |
136/143 = 95.1% (90.2β97.6) | 67.0% | 0.989 | 0.0052 |
Unseen generators (the realistic view)
Never trained on, thresholds fixed from validation and applied unchanged. Mean over 3 training seeds of the same recipe (this checkpoint is seed 1), greedy rule, 2% budget. Both sets are LLM-written text.
| Test set | Benign false alarms | Catch rate, correct category |
|---|---|---|
| Horizon eval sample (in-distribution) | 0.7% | 99.1% |
| Llama-3.1-8B-Instruct (831 rows) | 12.0% | 93.5% |
| Qwen3-8B (850 rows) | 19.3% | 95.9% |
Seed-to-seed spread is 1β6 points on false alarms, so differences of a few points between configurations are inside training noise.
Compared with Horizon (3-seed Vigil, earlier data round)
Scored by the same code on the same chats, on one 8-CPU node. Horizon is the deployed INT8 LiteRT-LM build. Catch is the share of risky chats where a gold category is flagged.
| Test set | Vigil catch | Horizon (focus) catch | Vigil false alarms | Horizon (focus) false alarms |
|---|---|---|---|---|
| Horizon sample | 98.3% | 93.9% | 0.9% | 0.4% |
| Llama-3.1-8B | 89.4% | 67.7% | 11.6% | 21.1% |
| Qwen3-8B | 93.0% | 62.6% | 17.1% | 20.0% |
| Vigil (ONNX fp32) | Horizon (LiteRT-LM INT8) | |
|---|---|---|
| Median latency (CPU) | 207 ms | 747 ms |
| On disk | 1,231 MiB | 1,017 MiB |
| Peak memory | 2,018 MiB | 2,601 MiB |
Vigil's gain is mostly recall on categories where Horizon degrades on unseen text (for example threats
recall 86β99% against 4β5%). Horizon is smaller on disk and has fewer false alarms on its own sample. Neither
system meets the 2% false-alarm target on unseen text.
Limitations and failure modes
Please read this section before relying on the model.
Synthetic only. No real conversation was used. The near-perfect in-distribution scores partly reflect generator similarity: ranking transfers to other generators (AUC 0.94β0.995), but operating points do not.
False alarms on unseen text are high: 12β19% of benign chats on two held-out generators, against a 2% target. At realistic prevalence (mostly benign traffic) most flags would be false alarms. At 1% risky prevalence, an earlier checkpoint with a 1.8% false-alarm rate would have had roughly 36% precision.
Friendly teasing is flagged about half the time. On a "friends joking and teasing" scene, ~49% of benign chats are flagged. A blind audit of 46 such rows found 43 friendly, 3 unclear and 0 cruel, so these are model errors, not label noise. Two rounds of targeted data did not fix it. Example from this checkpoint:
Mia: omg you failed the quiz too? lol / Sam: yeah we're both idiots π wanna study at my place after school? β
bullyingprobability 0.993Safety advice ("what do I do if a stranger asks for photos?") and risk-adjacent educational chats are the other main source of false alarms.
Single-label softmax. The 2-pass design makes categories compete for one probability mass. A chat showing two risks may flag only one. For example, the grooming example above flags
platform_migration(0.98) but notgrooming(0.005, threshold 0.023). Multi-label behavior is untested.Severity is weak (macro-F1 0.416). Do not use it as a triage score without your own validation.
sexual_contenthas low precision (35%), and its labels on the unseen-generator sets are unreliable (the generators softened or refused the request). Recall there is 45β66%, and label noise cannot be separated from a real model gap.Thresholds are brittle. They sit far below 0.5 and the probabilities are not calibrated. A new deployment would need its own threshold selection, and a few hundred labelled rows from the target traffic cut false alarms by only 0β8 points while costing about as much recall.
Languages. Only English and Spanish were trained and evaluated.
Not tested on a phone. The
.tflitewas run only with the Python LiteRT interpreter on a server CPU. It needs ~700 MB of memory and 1.3 GB of storage.Training variance is large. Retraining after removing 0.5% of rows moved held-out false alarms by up to Β±6 points. Treat single-run differences as unreliable.
Adversarial robustness (deliberate evasion, coded language, slang, images) was not tested.
Intended use and out-of-scope use
Intended: research on child-safety classification; benchmarking encoder classifiers against generative ones; prototyping an on-device first-stage filter whose flags are reviewed by a human or a second model.
Out of scope:
- Fully automated enforcement, account action or reporting to authorities on a flag alone.
- Surveillance of children or adults outside a context with appropriate consent and legal basis.
- Any use as the sole safeguard for a child. A flag is a prompt for human review, and a missed risk is possible.
- Commercial use without written permission (see the license).
A false alarm here can mean a parent or moderator reading a child's harmless private chat. Deploy with that privacy cost in mind, and only with appropriate consent and legal basis.
Reproducibility
Training, evaluation, threshold selection and export code live in the Vigil code repository (SafeCircle),
with per-round write-ups under docs/results/ (pilot, cross-generator, mixed training, round 4, seeds and hard
negatives, banter, deployment calibration, LiteRT export). This checkpoint is experiments/c1-choice8, with
thresholds from experiments/c1-choice8-greedy-val.
Acknowledgments
- Laya by Nandha Kishor M / Convai Innovations (Apache-2.0): the
decision-head architecture, loss and fine-tuning recipe.
export_onnx.pyis derived from Laya (seeNOTICEin the code repo). - mmBERT by JHU CLSP (MIT): the base encoder.
- Synthetic data generated with Qwen2.5-7B and Claude (the Horizon data), Mistral-7B-Instruct-v0.3 and Gemma-2-9B-it. Follow those models' terms for any redistribution of generated text.
Citation
@misc{vigil2026,
title={Vigil: A Typed-Decision Encoder for Child-Safety Risk Detection},
author={SafeCircle},
year={2026},
url={https://huggingface.co/safecircleai/vigil}
}
- Downloads last month
- 23
Model tree for safecircleai/vigil
Base model
convaiinnovations/laya-multilingual