⚠️ Devseis AI Act Classifier — v6 (Experiment, Not Legal Advice)

A Devseis model, built on LEGAL-BERT-SMALL.

This model does not give legal advice, and its output must not be used on its own to decide an AI system's EU AI Act tier. It powers the in-browser half of Devseis/Caveat, a demo that shows where automated classification works and where it fails.

It reads a free-text description of an AI system and returns one of four tiers:

Label Meaning
prohibited Art 5 banned practice
high_risk Annex III use case, or Annex I product safety component
limited_risk Art 50 transparency duties (chatbots and voice agents, deepfakes, AI-written public text)
minimal_risk none of the above

General-purpose AI (GPAI) models are decided by a rule, not by this model. In Caveat, a separate rule (gpai_rule.js) runs first. If the text describes a provider training or releasing a general-purpose model, the rule decides gpai or gpai_systemic_risk from the compute figure (Art 51(2), > 10^25 FLOPs) or a Commission designation (Art 51(1)(b)), and this model is not used. Small text classifiers read exponents unreliably, so a rule is exact where the model would guess. The rule passes 655 of 655 test inputs, and it never triggers on any of the 267 non-GPAI v3 scenarios.

Training data

  • Dataset: Devseis/devseis-ai-act-classifier-v6-data (data_v6_systems): 366 scenarios / 1,310 rows.

    Split Scenarios Rows high / minimal / prohibited / limited (scenarios)
    train 242 938 96 / 76 / 43 / 27
    validation 54 162 19 / 17 / 10 / 8
    test 70 210 27 / 21 / 12 / 10
  • The data is synthetic and template-expanded. Each scenario is one written use case with one label. In validation and test it appears as 3 rows that differ in their opening phrase; in training, each scenario has one extra row in a different voice (a question, a pitch, a tender request…) so the model doesn't learn one writing style. The real sample size is the scenario count, not the row count.

  • Splits are by scenario, so all rows of a scenario sit in the same split and no test scenario is seen in training. All v5 rows keep their split, so v5 and v6 can be compared on identical test rows.

  • Sources: the text of Regulation (EU) 2024/1689 (Art 5, Art 50, Annex I, Annex III), its recitals, and the European Commission's draft guidelines on classifying high-risk AI systems, with their worked "falls within / falls outside" examples.

Label corrections since the published v2 model (v3)

An audit found 6 v2 labels that were wrong under the Act. They were fixed without moving any scenario between splits:

Scenario v2 → v3 Why
Office system reading employees' faces to adjust lighting limited → prohibited Art 5(1)(f): workplace emotion inference
Retail mood inference from in-store cameras limited → high_risk Annex III(1)(c); Art 50(3) on top
Stadium live face ID for crowd management (test set) prohibited → high_risk Art 5(1)(h) covers law enforcement only
Airport live face ID for general security prohibited → high_risk not stated to be law enforcement
University chatbot answering applicants minimal → limited Art 50(1) chatbot
Age bracket from photo for ad personalisation minimal → high_risk consistent age/sex policy (below)

3 rows of the published test_v2 evaluation had wrong labels (the stadium scenario). Against the corrected labels, the published v2 model scores accuracy 0.747, macro-F1 0.776 and prohibited recall 0.889, not the published 0.765 / 0.794 / 0.900.

v3 also reworded 7 scenarios so the text states the legal test (for example, the significant-harm element of Art 5(1)(a)), dropped one contradictory near-duplicate, and fixed the grammar of all 661 templated rows.

Scope added in v4, v5 and v6

  • v4 (+37 scenarios): Annex I product safety components (medical devices, IVDs, machinery, toys, lifts, pressure equipment) and non-safety contrasts; the AI Omnibus prohibitions; Art 6(3) profiling counter-examples and derogations; deepfake contrasts (satire, consented digital doubles); and downstream builders on a GPAI model.
  • v5 (+23 scenarios): the v4 model labelled every phone-based agent in test as minimal_risk, because all the conversational limited_risk examples in training were text chat. v5 adds voice, phone, SMS, WhatsApp, email and avatar agents (Art 50(1)), and consented or non-identifiable synthetic media (Art 50(2)/(4)). It also adds minimal_risk contrasts on the same channels with no direct AI interaction with the public, such as call transcription and staff-reviewed email drafts.
  • v6 (+39 scenarios, plus a style row for every training scenario): targets weak spots found on v5 validation data. It adds harmless look-alikes in high-risk domains that the Act leaves out (interview scheduling, police transcription, car-damage claims, campaign leaflet routes, face blurring, 1:1 face verification), each paired where useful with a high-risk counterpart; untargeted facial-image scraping (Art 5(1)(e)) in varied wording; and consented or fictional synthetic media. A similarity check rejected any new scenario too close to a validation, test or out-of-distribution text.

Policy choices and legal dates — read before relying on a label

  • Age and sex estimation from biometrics is labelled high_risk (biometric categorisation, Annex III(1)(b)). The law is unsettled here: some such uses may be ancillary features outside Annex III. This is a deliberate, consistent policy choice, not a settled reading.
  • The AI Omnibus prohibitions apply from 2 December 2026. The AI Omnibus, Regulation (EU) 2026/1744 (Official Journal, 24 July 2026), inserts two new bans: Art. 5(1)(ba), AI systems that generate non-consensual intimate imagery of identifiable people, and Art. 5(1)(bb), AI systems that generate child sexual abuse material. Under the new Art. 5(1a), they apply where that output is the system's intended purpose, or a reasonably foreseeable outcome without adequate safeguards. These scenarios are labelled prohibited, which is the law as it will stand from that date. Before then, the legally current tier for the intimate-deepfake scenarios is limited_risk (Art 50(4)).

Model

  • Base model: nlpaueb/legal-bert-small-uncased (35M parameters, pretrained on legal text including EU legislation).
  • Why this backbone: on v4, it tied with DistilBERT on macro-F1 (paired difference −0.012, 95% CI −0.19 to +0.16). Its prohibited recall was much higher (0.949 vs 0.616; its worst seed beat DistilBERT's best), and it is half the size. SetFit (bge-small-en-v1.5) was partly tuned: 2 of its 4 configurations ran before tuning was stopped. Its best validation macro-F1 (0.711) trailed both.
  • Training: learning rate 2e-5 (chosen on validation from {2e-5, 3e-5, 5e-5}), batch size 16, weight decay 0.01, 10% warmup, up to 15 epochs, keeping the epoch with the best validation macro-F1. No class weights: on validation they lowered macro-F1 without improving prohibited recall. Max input length is 128 tokens. The shipped weights are the seed-42 run, a choice fixed before any results were seen.
  • Browser export: int8 ONNX, 35.5 MB (dynamic QUInt8, chosen from 8 settings by validation agreement). It agrees with the PyTorch model on 99.0% of test rows under Python onnxruntime and 98.6% under the browser engine (onnxruntime-web, 3 of 210 rows); test accuracy 82.9% vs 83.3% in PyTorch.
  • "Uncertain" in Caveat: the page shows "uncertain — needs expert review" when the top probability is below 50% (the model is split between tiers). A stricter confidence threshold was tested on validation and rejected: this model's mistakes are mostly confident (even at 95% confidence it is right 88% of the time), so holding back low-confidence answers would hide more correct answers than errors. On the test set the 50% rule flagged 1 of 210 rows.

Results

Mean ± std over 3 seeds (42, 43, 44). 95% confidence intervals from a scenario-level bootstrap (1,000 resamples; all rows of a scenario move together).

Full v6 test set — 210 rows / 70 scenarios

Metric Value 95% CI
Prohibited recall (headline) 0.991 ± 0.013
Banned practices labelled minimal/limited (severe errors) 0 of 108 rows
Accuracy 0.822 ± 0.008 0.735–0.905
Macro-F1 0.850 ± 0.010 0.766–0.913
Recall: high / limited / minimal 0.761 / 0.967 / 0.735

Compared with v5 and the published model, on identical rows (the v5 test set, 198 rows / 66 scenarios)

Model Accuracy Macro-F1 (95% CI) Prohibited recall Severe errors limited / minimal recall
This model (v6) 0.815 ± 0.010 0.843 (0.758–0.912) 0.990 0 / 99 0.963 / 0.719
v5 0.796 ± 0.009 0.824 (0.741–0.888) 0.980 0 / 99 0.889 / 0.684
Published v2 model 0.747 0.777 (0.674–0.861) 0.909 0 / 33 1.000 / 0.754

v6 is better than v5 on nearly every measure, but the gain is not statistically established: the paired macro-F1 difference is +0.020 (95% CI −0.079 to +0.125). High-risk recall moved slightly the other way (0.770 → 0.761, within noise). Over-flagging of harmless systems improved only modestly.

Out-of-distribution check — 35 rows written in a different style by a different generator

Model Accuracy Macro-F1 Prohibited recall Severe errors
This model (v6) 0.800 ± 0.047 0.818 0.917 1 (of 24)
v5 0.762 ± 0.036 0.778 0.833 3 (of 24)
Published v2 model 0.800 0.813 0.875 0 (of 8)
  • Fixed: untargeted facial-image scraping phrased as a "crawler", which v5 labelled minimal risk on every seed, is now prohibited on every seed. Caveat: v6 targeted this weakness after the out-of-distribution check exposed it, so this one case is no longer a blind test.
  • Remaining: the shipped seed-42 checkpoint labels an advertising engine that targets people with learning disabilities with exploitative loan offers (Art 5(1)(b)) as limited risk. The other two seeds get it right. We did not switch seeds because of this result, to keep the out-of-distribution set an honest test.

The test sets are small. A gap of under about 5 points between models is usually noise. And because every scenario is synthetic, these scores are an upper bound. Real system descriptions labelled by a qualified reviewer would be the real test, and have not been used.

Limitations

  • Synthetic, template-expanded training data, with small test sets (70 scenarios, 12 of them prohibited). Even a perfect score on 12 prohibited scenarios is consistent with a true miss rate of up to about 25%.
  • minimal_risk recall is 0.735: the model still pushes about 1 in 4 harmless systems up to high_risk, which is the safer direction to be wrong in, but still wrong.
  • Its confidence is not a reliable warning sign: most of its mistakes are made confidently.
  • On text phrased unlike its training data it is weaker, and the shipped checkpoint missed one exploitation case (see above).
  • The model pattern-matches: it can be swayed by vocabulary and does not reason about the legal test.
  • It returns a single tier. Real systems can trigger several obligations at once (for example high-risk plus Art 50 disclosure).

Base model, licence and attribution

This model is an adaptation of LEGAL-BERT-SMALL (nlpaueb/legal-bert-small-uncased) by I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras and I. Androutsopoulos, which is licensed under CC BY-SA 4.0.

  • Changes made by Devseis: added a 4-class classification head, fine-tuned all weights on data_v6_systems, then exported to ONNX and quantised to int8.
  • Licence of this model: CC BY-SA 4.0, as the base model's ShareAlike term requires. The training data is Devseis's own work and is licensed separately (CC BY 4.0).
  • No endorsement: the LEGAL-BERT authors are not affiliated with Devseis and do not endorse this model.
@inproceedings{chalkidis-etal-2020-legal,
    title = "{LEGAL}-{BERT}: The Muppets straight out of Law School",
    author = "Chalkidis, Ilias and Fergadiotis, Manos and Malakasiotis, Prodromos and Aletras, Nikolaos and Androutsopoulos, Ion",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020",
    month = nov,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    doi = "10.18653/v1/2020.findings-emnlp.261",
    pages = "2898--2904"
}

How to use

Python (PyTorch weights):

from transformers import pipeline
clf = pipeline("text-classification", model="Devseis/devseis-ai-act-classifier-v6")
clf("An app that ranks job applicants' CVs before a recruiter sees them.")
# [{'label': 'high_risk', 'score': ...}]

In the browser (transformers.js, int8 ONNX, 35.5 MB):

import { pipeline } from "@huggingface/transformers";
const clf = await pipeline("text-classification", "Devseis/devseis-ai-act-classifier-v6", { dtype: "q8" });
const [top] = await clf("An app that ranks job applicants' CVs before a recruiter sees them.");

This model does not handle general-purpose AI (GPAI) model providers. Caveat runs a small rule for that first (gpai_rule.js).

Files

Path What it is
model.safetensors, config.json PyTorch weights (seed-42 run) and configuration
tokenizer.json, tokenizer_config.json, special_tokens_map.json, vocab.txt tokenizer
onnx/model_quantized.onnx the int8 model the browser runs (35.5 MB)
evaluation/report_test_v6.md full results on the v6 test set (210 rows / 70 scenarios)
evaluation/report_test_v5_subset.md results on the v5 test rows, next to v5 and the published v2 model
evaluation/export_check.json int8-vs-PyTorch agreement check and the quantisation settings tried

Training and test data: Devseis/devseis-ai-act-classifier-v6-data. Previous version: Devseis/devseis-ai-act-classifier-v5. Earlier versions (v1–v3, DistilBERT): Devseis/eu-ai-act-classifier-experiment.

Not legal advice

This model and Caveat are research demonstrations. Classifying a real AI system under the EU AI Act needs expert review of the full facts.

License

Model: CC BY-SA 4.0 (see "Base model, licence and attribution" above). Training data: CC BY 4.0.

Downloads last month
13
Safetensors
Model size
35.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Devseis/devseis-ai-act-classifier-v6

Quantized
(2)
this model

Dataset used to train Devseis/devseis-ai-act-classifier-v6

Space using Devseis/devseis-ai-act-classifier-v6 1

Collection including Devseis/devseis-ai-act-classifier-v6