EU AI Act risk check experiment, by Devseis
Your email app opens with a draft to Devseis — nothing is sent until you press send.
How good is it, honestly?
Measured on 70 example systems the model never saw while learning. Tap any for a plain explanation.
Out of all test cases that truly belong to each tier, the share the model got right. Average of 3 training runs.
Same 66 test systems for everyone. The ranges show how much the score could move by chance () — they overlap, so these gaps are not proven.
| Version | Banned caught | Balanced score | Accuracy | Chatbots & voice agents |
|---|---|---|---|---|
| v6 — running now legal-text model, 242 training cases, more writing styles | 99% | 0.84 (0.76–0.91) | 82% | 96% |
| v5 (previous) same model, before the weak-spot fixes | 98% | 0.82 (0.74–0.89) | 80% | 89% |
| Original live model general model, v2 data | 91% | 0.78 (0.67–0.86) | 75% | 100% |
"Balanced score" = , which punishes a model for ignoring any one tier.
About 1 in 4 truly harmless systems gets pushed up to "high risk" (down from 1 in 3). Safer than the reverse, but still wrong.
On 35 cases written in a different style it scored 80% (up from 76%), and once called an exploitative ad engine aimed at people with learning disabilities "limited risk" — a real miss. ()
Only 12 banned examples in the test. Even a perfect score there can't rule out missing 1 in 4 in real life.
It reacts to words it learned, not to the legal test. It gives one tier, even when several rules apply at once.
A small spots descriptions of big general-purpose AI models and decides those from the training-compute figure (). Tested on 655 cases, all correct.
Devseis's own model: the small legal-text model legal-bert-small, by Devseis on 938 examples. Runs entirely in your browser — nothing is sent anywhere.
A general AI given the EU AI Act rules as instructions and asked to reason. Not trained on this task, never independently checked. Runs on Cloudflare's free tier.
Check it yourself: model & card · training & test data · full evaluation reports · earlier versions
The model is useful for a first look. The decisions that matter need someone who can ask questions.
From the Official Journal texts of the AI Act (Regulation (EU) 2024/1689, Art 111 and 113) and the AI Omnibus (Regulation (EU) 2026/1744). Dates the Omnibus changed are marked.
Simplified overview, not legal advice. Which dates apply to your system depends on its role and category.
Built by Devseis
Devseis is a Luxembourg-based consultancy and accredited training provider, founded and directed by Dr. Muhammad Umer Wasim (PhD). We help SMEs, startups and technology companies assess and reduce cyber and compliance risk, and explain it in plain language for management.
How this was built
What we tried, what happened, and what each step taught us — including the mistakes.