Vigil: Typed-Decision Child-Safety Classifier (research checkpoint)

Vigil reads a chat involving a child and answers, in a single encoder forward pass, which of seven risks it shows and how severe it is. It outputs numbers (a probability per category and a severity score), not text, so there is nothing to parse and the output cannot be malformed.

It is Laya-multilingual (a 322M-parameter mmBERT-base encoder) fully fine-tuned on SafeCircle's synthetic child-safety data, and is the encoder-based counterpart to the generative Horizon models.

⚠️ Research checkpoint. Not ready for production moderation. Every number below comes from LLM-written synthetic conversations. No real child conversations were used for training or evaluation, and the model has not been run on a phone. On text from generators it did not train on, it flags 12–19% of harmless chats. Read Limitations before using it.

⚠️ License: SafeCircle Research License (SRL-1.0). The weights are derived from SafeCircle's research-licensed training data. Commercial use is prohibited without written permission. Contact legal@safecircle.tech.


Contents of this repository

Path What it is
model.safetensors, rl_agent_config.json, encoder/, tokenizer/ The PyTorch checkpoint, in the layout laya.Agent loads
thresholds.json Required for decisions. Per-category operating thresholds (see below)
litert/vigil_s512_fp32.tflite fp32 LiteRT/TFLite graph for Android (1.29 GB, input [2, 512])
litert/vigil_spec.json Question prefixes, marker positions, temperatures and thresholds for the .tflite graph
litert/vigil_tokenizer.bin Compact BPE tokenizer for on-device use (9 MB)
litert/report.json Fidelity and latency of the LiteRT export against PyTorch

Model details

Developed by SafeCircle research
Base model convaiinnovations/laya-multilingual (Apache-2.0), itself built on jhu-clsp/mmBERT-base (MIT)
Architecture ModernBERT-style bidirectional encoder, 768 hidden, alternating full / sliding-window attention, plus Laya's decision head (2 layers)
Parameters 322M
Context 1,024 tokens (the LiteRT export uses a fixed 512-token window)
Languages English and Spanish in the fine-tuning data. The base model covers 100+ languages, but Vigil was not evaluated beyond English and Spanish
Output Probability per risk category, expected severity 0–4, and a Laya confidence value
Design choice8 (the "2-pass" design): one 8-way choice (benign + 7 categories) plus one severity score = 2 encoder passes per conversation
Run name c1-choice8 (banter round, seed 1)
Fine-tuning Full fine-tune, cross-entropy plus Laya's RLCD term, 2 epochs, encoder LR 2.5e-5, head LR 1e-4, cosine schedule, micro-batch 8 Γ— grad-accum 4, one GPU, temperatures fitted on a held-out slice of whole conversations

The training block in rl_agent_config.json is inherited from the upstream Laya checkpoint and does not describe this fine-tuning run.

Risk categories

Category Meaning
grooming An older person builds trust to exploit the child (flattery, secrecy, gifts, boundary testing)
bullying Insults, humiliation, exclusion or harassment
sexual_content Sexual content or requests directed at the child
isolation Cutting the child off from family and friends
personal_info Asking for or sharing address, school, phone, photos
platform_migration Moving the chat to another, less monitored app
threats Threats, blackmail or self-harm

Severity levels: 0 none, 1 low, 2 medium, 3 high, 4 critical.

How to use

pip install laya huggingface_hub
import json
import laya
from huggingface_hub import snapshot_download
from vigil.questions import build_questions   # from the Vigil code repo; see "Questions" below

path = snapshot_download("safecircleai/vigil")
agent = laya.Agent(path, device="cpu")
thresholds = json.load(open(f"{path}/thresholds.json"))["thresholds"]

chat = (
    "Child: hi!\n"
    "Adult: you're so mature for your age. don't tell your parents about our chats ok? "
    "let's move to Snapchat"
)
answers = agent.predict(chat, build_questions("choice8"))["answers"]

probs = answers["category"]["probabilities"]            # includes "benign"
flagged = [c for c, t in thresholds.items() if probs[c] >= t]
severity = answers["severity"]["score"]                 # expected severity, 0-4
print(flagged, round(severity, 2))                      # ['platform_migration'] 2.74

Use the thresholds. The 0.5 default is wrong for this model: the shipped thresholds are between 0.004 and 0.12, chosen to cap benign false alarms at 2% on a validation set. A category is flagged when its probability is at or above its threshold, and several categories can be flagged at once.

Questions. The category and severity questions are fixed text. The code that builds them (vigil/questions.py) is in the Vigil code repository. litert/vigil_spec.json contains the same two questions already tokenized, so the on-device path needs no Python.

Long conversations. The context is 1,024 tokens (512 for the LiteRT graph). The evaluation harness drops conversations that do not fit. The on-device reference keeps the most recent tokens. Neither is validated for long chats.

On Android (LiteRT)

litert/vigil_s512_fp32.tflite takes a [2, 512] batch (the category row and the severity row) and returns marker logits [2, 8]. The softmax with the checkpoint temperature, the thresholds and the tokenizer live outside the graph: tokenize with vigil_tokenizer.bin, build the two rows from the prefixes in vigil_spec.json, run the graph, then decode (vigil/ondevice.py in the code repo is the reference; a Kotlin port exists in the SafeCircle Android app).

Variant Size Max abs prob diff vs PyTorch Decision flips (2,100 decisions) p50 latency
fp32 (shipped here) 1,289 MB (702 MB resident) 0.00006 4 141 ms
int8 weight-only (not shipped) 329 MB (831 MB resident) 0.37 12 139 ms
int8 embeddings only (not shipped) 702 MB (664 MB resident) 0.17 5 141 ms

fp32 is the only variant published because the thresholds are tiny, so small probability drift moves decisions. int8 weight-only saves disk but not memory, since weights are dequantized at load. Latency was measured on a server CPU (8 threads) with the Python LiteRT interpreter, not on a phone. Op support, latency and memory on the on-device LiteRT runtime are untested.

Training data

All data is synthetic.

Source Role
5,000 conversations sampled from the private safecircleai/horizon-training-data (LLM-written, 7 risk categories + benign, English and Spanish) Core risky and benign examples
~2,700 conversations from Mistral-7B-Instruct-v0.3 and ~2,700 from Gemma-2-9B-it Generator diversity: 6 risk categories and 15 benign scenes (teachers and coaches, parents arranging pick-up, friends, games)
~2,000 benign conversations from Mistral and Gemma on 10 "educational" scenes Hard negatives: chats that talk about risky topics (a parent explaining why not to share an address, a child asking how to report a bully)
~2,000 benign friendly-banter conversations Hard negatives for teasing between friends
  • sexual_content and severity labels were not taken from the generated sets (the generators refused or softened the request, so those labels were unreliable). Generated rows train the category question only.
  • Rows with an empty turn were removed (0.5% of generated rows).
  • Whole conversations were held out for the calibration slice (no conversation spans training and calibration).
  • An exact-match overlap audit found no train/eval duplicates in the Horizon sample. Near-duplicate paraphrases were not tested.
  • Every conversation has at most one category, so multi-label behavior was never trained or tested.

Evaluation

Thresholds come from a greedy per-category rule: maximize recall under a 2% benign false-alarm budget on the validation half of a mixed validation file (Horizon sample + Mistral + Gemma rows). They are a policy choice, not calibrated probabilities. Intervals are 95% Wilson.

Mixed validation, held-out test half (n = 1,379; 972 risky, 407 benign)

This set shares generators with the training data (Horizon's, Mistral, Gemma), so it is the optimistic view.

Metric Value
Catch rate, any category flagged (risky chats) 98.5% (957/972; CI 97.5–99.1)
Catch rate, a correct category flagged 97.0% (943/972; CI 95.7–97.9)
Benign false alarms 4.2% (17/407; CI 2.6–6.6), against the 2% budget set on the validation half
Severity macro-F1 0.416 (weak)
Latency per conversation (cluster GPU), p50 / p95 9.0 / 10.6 ms
Category Recall (95% CI) Precision (all flags) AUC Threshold
grooming 149/154 = 96.8% (92.6–98.6) 72.3% 0.994 0.023
bullying 141/151 = 93.4% (88.2–96.4) 72.3% 0.990 0.007
sexual_content 59/59 = 100.0% (93.9–100.0) 35.3% 0.997 0.0063
isolation 155/156 = 99.4% (96.5–99.9) 84.7% 0.998 0.1222
personal_info 145/148 = 98.0% (94.2–99.3) 70.4% 0.994 0.0042
platform_migration 158/161 = 98.1% (94.7–99.4) 69.3% 0.995 0.0049
threats 136/143 = 95.1% (90.2–97.6) 67.0% 0.989 0.0052

Unseen generators (the realistic view)

Never trained on, thresholds fixed from validation and applied unchanged. Mean over 3 training seeds of the same recipe (this checkpoint is seed 1), greedy rule, 2% budget. Both sets are LLM-written text.

Test set Benign false alarms Catch rate, correct category
Horizon eval sample (in-distribution) 0.7% 99.1%
Llama-3.1-8B-Instruct (831 rows) 12.0% 93.5%
Qwen3-8B (850 rows) 19.3% 95.9%

Seed-to-seed spread is 1–6 points on false alarms, so differences of a few points between configurations are inside training noise.

Compared with Horizon (3-seed Vigil, earlier data round)

Scored by the same code on the same chats, on one 8-CPU node. Horizon is the deployed INT8 LiteRT-LM build. Catch is the share of risky chats where a gold category is flagged.

Test set Vigil catch Horizon (focus) catch Vigil false alarms Horizon (focus) false alarms
Horizon sample 98.3% 93.9% 0.9% 0.4%
Llama-3.1-8B 89.4% 67.7% 11.6% 21.1%
Qwen3-8B 93.0% 62.6% 17.1% 20.0%
Vigil (ONNX fp32) Horizon (LiteRT-LM INT8)
Median latency (CPU) 207 ms 747 ms
On disk 1,231 MiB 1,017 MiB
Peak memory 2,018 MiB 2,601 MiB

Vigil's gain is mostly recall on categories where Horizon degrades on unseen text (for example threats recall 86–99% against 4–5%). Horizon is smaller on disk and has fewer false alarms on its own sample. Neither system meets the 2% false-alarm target on unseen text.

Limitations and failure modes

Please read this section before relying on the model.

  1. Synthetic only. No real conversation was used. The near-perfect in-distribution scores partly reflect generator similarity: ranking transfers to other generators (AUC 0.94–0.995), but operating points do not.

  2. False alarms on unseen text are high: 12–19% of benign chats on two held-out generators, against a 2% target. At realistic prevalence (mostly benign traffic) most flags would be false alarms. At 1% risky prevalence, an earlier checkpoint with a 1.8% false-alarm rate would have had roughly 36% precision.

  3. Friendly teasing is flagged about half the time. On a "friends joking and teasing" scene, ~49% of benign chats are flagged. A blind audit of 46 such rows found 43 friendly, 3 unclear and 0 cruel, so these are model errors, not label noise. Two rounds of targeted data did not fix it. Example from this checkpoint:

    Mia: omg you failed the quiz too? lol / Sam: yeah we're both idiots πŸ˜‚ wanna study at my place after school? β†’ bullying probability 0.993

    Safety advice ("what do I do if a stranger asks for photos?") and risk-adjacent educational chats are the other main source of false alarms.

  4. Single-label softmax. The 2-pass design makes categories compete for one probability mass. A chat showing two risks may flag only one. For example, the grooming example above flags platform_migration (0.98) but not grooming (0.005, threshold 0.023). Multi-label behavior is untested.

  5. Severity is weak (macro-F1 0.416). Do not use it as a triage score without your own validation.

  6. sexual_content has low precision (35%), and its labels on the unseen-generator sets are unreliable (the generators softened or refused the request). Recall there is 45–66%, and label noise cannot be separated from a real model gap.

  7. Thresholds are brittle. They sit far below 0.5 and the probabilities are not calibrated. A new deployment would need its own threshold selection, and a few hundred labelled rows from the target traffic cut false alarms by only 0–8 points while costing about as much recall.

  8. Languages. Only English and Spanish were trained and evaluated.

  9. Not tested on a phone. The .tflite was run only with the Python LiteRT interpreter on a server CPU. It needs ~700 MB of memory and 1.3 GB of storage.

  10. Training variance is large. Retraining after removing 0.5% of rows moved held-out false alarms by up to Β±6 points. Treat single-run differences as unreliable.

  11. Adversarial robustness (deliberate evasion, coded language, slang, images) was not tested.

Intended use and out-of-scope use

Intended: research on child-safety classification; benchmarking encoder classifiers against generative ones; prototyping an on-device first-stage filter whose flags are reviewed by a human or a second model.

Out of scope:

  • Fully automated enforcement, account action or reporting to authorities on a flag alone.
  • Surveillance of children or adults outside a context with appropriate consent and legal basis.
  • Any use as the sole safeguard for a child. A flag is a prompt for human review, and a missed risk is possible.
  • Commercial use without written permission (see the license).

A false alarm here can mean a parent or moderator reading a child's harmless private chat. Deploy with that privacy cost in mind, and only with appropriate consent and legal basis.

Reproducibility

Training, evaluation, threshold selection and export code live in the Vigil code repository (SafeCircle), with per-round write-ups under docs/results/ (pilot, cross-generator, mixed training, round 4, seeds and hard negatives, banter, deployment calibration, LiteRT export). This checkpoint is experiments/c1-choice8, with thresholds from experiments/c1-choice8-greedy-val.

Acknowledgments

  • Laya by Nandha Kishor M / Convai Innovations (Apache-2.0): the decision-head architecture, loss and fine-tuning recipe. export_onnx.py is derived from Laya (see NOTICE in the code repo).
  • mmBERT by JHU CLSP (MIT): the base encoder.
  • Synthetic data generated with Qwen2.5-7B and Claude (the Horizon data), Mistral-7B-Instruct-v0.3 and Gemma-2-9B-it. Follow those models' terms for any redistribution of generated text.

Citation

@misc{vigil2026,
  title={Vigil: A Typed-Decision Encoder for Child-Safety Risk Detection},
  author={SafeCircle},
  year={2026},
  url={https://huggingface.co/safecircleai/vigil}
}
Downloads last month
23
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for safecircleai/vigil

Finetuned
(59)
this model