Reproducible numbers
What it detects, and what it costs
Every classifier is scored separately in all 26 target languages at a calibrated threshold, on synthetic corpora generated per language. The shape first; every table is one click below it.
Accuracy, mean and weakest language
One row per evaluated detector. The dot pair is the honest summary of 26 numbers: where the detector sits on average, and the single language where it is worst.
Two detectors are deliberately not on the figure, because their tasks are scored differently. topic_scope picks one taxonomy node or "none of these": top-1 accuracy 0.859 on taxonomies from deployment types it was not trained on, where the model it replaced scores 0.804 and the original bi-encoder 0.479 on the same rows. groundedness is scored on grounded and not-grounded pairs: pair accuracy 0.602. Both have their detail in their own rows below.
Every row, per detector and per language
An aggregate hides the tail, so each detector opens to its full 26-row table: precision, recall, false positive rate and F1 per language, with the corpus size and threshold that produced them.
gibberishRuns todayT1, input0.992 mean 0.946 worst, Croatian
- Mean language F1
- 0.992Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.92Calibrated on validation, objective macro_f1.
- Positives in test
- 779Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 15,491Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 29 | 1.00 | 1.00 | 0.00 | 1.000 |
| Bulgarianbg | 32 | 1.00 | 1.00 | 0.00 | 1.000 |
| Croatianhr | 29 | 1.00 | 0.90 | 0.00 | 0.946 |
| Czechcs | 32 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 32 | 1.00 | 1.00 | 0.00 | 1.000 |
| Dutchnl | 29 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 30 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 31 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 31 | 0.97 | 0.97 | 0.03 | 0.968 |
| Frenchfr | 31 | 0.97 | 1.00 | 0.03 | 0.984 |
| Germande | 30 | 1.00 | 0.93 | 0.00 | 0.966 |
| Greekel | 31 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 30 | 1.00 | 1.00 | 0.00 | 1.000 |
| Irishga | 30 | 1.00 | 0.97 | 0.00 | 0.983 |
| Italianit | 30 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 30 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 30 | 1.00 | 0.97 | 0.00 | 0.983 |
| Maltesemtnot in the base model's pretraining | 29 | 0.97 | 0.97 | 0.03 | 0.966 |
| Polishpl | 28 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 30 | 0.97 | 1.00 | 0.03 | 0.984 |
| Romanianro | 28 | 1.00 | 1.00 | 0.00 | 1.000 |
| Slovaksk | 29 | 1.00 | 1.00 | 0.00 | 1.000 |
| Sloveniansl | 30 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 29 | 1.00 | 1.00 | 0.00 | 1.000 |
| Swedishsv | 30 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 29 | 1.00 | 1.00 | 0.00 | 1.000 |
Weakest languages. Croatian at 0.946, German at 0.966, Maltese at 0.966, Finnish at 0.968, Irish at 0.983. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
abbreviation_heavy0.0%code_or_identifier0.0%mixed_language_valid1.3%mundane_account_access0.0%mundane_informational0.0%mundane_operational0.0%mundane_transactional0.0%proper_nouns2.6%short_but_valid1.4%typo_ridden_but_readable0.0%
Exported to ONNX at opset 17, traced at 32 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
biasRuns todayT2, output0.983 mean 0.942 worst, Maltese
- Mean language F1
- 0.983Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.77Calibrated on validation, objective macro_f1.
- Positives in test
- 2064Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 45,421Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 80 | 0.92 | 1.00 | 0.07 | 0.958 |
| Bulgarianbg | 80 | 0.99 | 0.97 | 0.01 | 0.981 |
| Croatianhr | 80 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 80 | 0.99 | 1.00 | 0.01 | 0.994 |
| Danishda | 80 | 1.00 | 1.00 | 0.00 | 1.000 |
| Dutchnl | 79 | 0.99 | 1.00 | 0.01 | 0.994 |
| Englishen | 80 | 0.96 | 0.99 | 0.03 | 0.975 |
| Estonianet | 79 | 0.96 | 1.00 | 0.03 | 0.981 |
| Finnishfi | 80 | 0.99 | 0.99 | 0.01 | 0.988 |
| Frenchfr | 80 | 0.98 | 1.00 | 0.02 | 0.988 |
| Germande | 80 | 0.96 | 0.99 | 0.03 | 0.975 |
| Greekel | 80 | 1.00 | 0.99 | 0.00 | 0.994 |
| Hungarianhu | 80 | 0.95 | 0.99 | 0.04 | 0.969 |
| Irishga | 77 | 0.97 | 0.96 | 0.02 | 0.967 |
| Italianit | 80 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 80 | 0.98 | 0.99 | 0.02 | 0.981 |
| Lithuanianlt | 79 | 0.99 | 0.99 | 0.01 | 0.987 |
| Maltesemtnot in the base model's pretraining | 78 | 0.95 | 0.94 | 0.04 | 0.942 |
| Polishpl | 80 | 0.98 | 1.00 | 0.02 | 0.988 |
| Portuguesept | 76 | 0.97 | 0.97 | 0.02 | 0.974 |
| Romanianro | 80 | 1.00 | 0.99 | 0.00 | 0.994 |
| Slovaksk | 80 | 0.99 | 0.99 | 0.01 | 0.988 |
| Sloveniansl | 80 | 1.00 | 0.99 | 0.00 | 0.994 |
| Spanishes | 76 | 1.00 | 0.97 | 0.00 | 0.987 |
| Swedishsv | 80 | 0.99 | 0.99 | 0.01 | 0.988 |
| Turkishtr | 80 | 0.95 | 0.97 | 0.04 | 0.963 |
Weakest languages. Maltese at 0.942, Azerbaijani at 0.958, Turkish at 0.963, Irish at 0.967, Hungarian at 0.969. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
counter_stereotype1.1%demographic_statistic0.0%discussing_bias1.7%inclusive_phrasing9.6%mundane_informational0.0%mundane_operational0.0%mundane_transactional0.0%neutral_description0.0%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
injectionRuns todayT2, input0.989 mean 0.882 worst, Maltese
- Mean language F1
- 0.989Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.02Calibrated on validation, objective recall_at_fpr.
- Positives in test
- 1080Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 45,541Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 42 | 1.00 | 1.00 | 0.00 | 1.000 |
| Bulgarianbg | 41 | 1.00 | 1.00 | 0.00 | 1.000 |
| Croatianhr | 42 | 0.98 | 1.00 | 0.01 | 0.988 |
| Czechcs | 42 | 0.95 | 1.00 | 0.02 | 0.977 |
| Danishda | 41 | 0.98 | 1.00 | 0.01 | 0.988 |
| Dutchnl | 41 | 1.00 | 0.98 | 0.00 | 0.988 |
| Englishen | 41 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 41 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 42 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 41 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 42 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 42 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 42 | 0.95 | 1.00 | 0.02 | 0.977 |
| Irishga | 42 | 0.98 | 0.98 | 0.01 | 0.976 |
| Italianit | 42 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 41 | 0.98 | 1.00 | 0.01 | 0.988 |
| Lithuanianlt | 42 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 42 | 0.80 | 0.98 | 0.08 | 0.882 |
| Polishpl | 41 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 42 | 0.98 | 1.00 | 0.01 | 0.988 |
| Romanianro | 42 | 0.98 | 1.00 | 0.01 | 0.988 |
| Slovaksk | 41 | 1.00 | 1.00 | 0.00 | 1.000 |
| Sloveniansl | 42 | 0.98 | 1.00 | 0.01 | 0.988 |
| Spanishes | 40 | 0.98 | 1.00 | 0.01 | 0.988 |
| Swedishsv | 41 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 42 | 1.00 | 1.00 | 0.00 | 1.000 |
Weakest languages. Maltese at 0.882, Irish at 0.976, Czech at 0.977, Hungarian at 0.977, Spanish at 0.988. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
meta_question1.2%mundane_account_access0.0%mundane_informational0.0%mundane_operational0.5%mundane_transactional0.5%ordinary_instruction0.0%ordinary_question0.6%quoted_attack0.9%roleplay_benign1.2%security_discussion0.9%technical_identifiers0.6%technical_payload0.6%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
moderationRuns todayT2, input and output0.980 mean 0.857 worst, Maltese
- Mean language F1
- 0.980Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.84Calibrated on validation, objective macro_f1.
- Positives in test
- 1623Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 33,708Synthetic, generated per language by claude-haiku-4-5.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 67 | 0.95 | 0.94 | 0.05 | 0.947 |
| Bulgarianbg | 69 | 1.00 | 0.97 | 0.00 | 0.985 |
| Croatianhr | 54 | 1.00 | 0.98 | 0.00 | 0.991 |
| Czechcs | 70 | 0.99 | 0.99 | 0.02 | 0.986 |
| Danishda | 62 | 1.00 | 0.98 | 0.00 | 0.992 |
| Dutchnl | 65 | 0.98 | 0.95 | 0.02 | 0.969 |
| Englishen | 60 | 0.98 | 1.00 | 0.01 | 0.992 |
| Estonianet | 56 | 0.98 | 0.98 | 0.01 | 0.982 |
| Finnishfi | 61 | 1.00 | 0.98 | 0.00 | 0.992 |
| Frenchfr | 61 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 61 | 1.00 | 0.97 | 0.00 | 0.983 |
| Greekel | 60 | 1.00 | 0.95 | 0.00 | 0.974 |
| Hungarianhu | 60 | 1.00 | 1.00 | 0.00 | 1.000 |
| Irishga | 69 | 0.97 | 0.88 | 0.03 | 0.924 |
| Italianit | 58 | 1.00 | 0.97 | 0.00 | 0.983 |
| Latvianlv | 65 | 0.97 | 0.98 | 0.03 | 0.977 |
| Lithuanianlt | 64 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 59 | 0.85 | 0.86 | 0.13 | 0.857 |
| Polishpl | 68 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 74 | 1.00 | 1.00 | 0.00 | 1.000 |
| Romanianro | 62 | 1.00 | 0.98 | 0.00 | 0.992 |
| Slovaksk | 59 | 1.00 | 0.98 | 0.00 | 0.992 |
| Sloveniansl | 62 | 1.00 | 0.98 | 0.00 | 0.992 |
| Spanishes | 53 | 0.96 | 1.00 | 0.03 | 0.982 |
| Swedishsv | 61 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 63 | 1.00 | 0.95 | 0.00 | 0.976 |
Weakest languages. Maltese at 0.857, Irish at 0.924, Azerbaijani at 0.947, Dutch at 0.969, Greek at 0.974. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
cyber_intrusion_near_miss0.0%defamation_near_miss0.0%election_integrity_near_miss1.9%extremism_near_miss7.8%fraud_deception_near_miss3.5%hate_incitement_near_miss0.0%illicit_drugs_near_miss0.0%mundane_account_access2.8%mundane_enterprise1.0%mundane_informational0.0%mundane_operational0.0%mundane_transactional0.0%property_crime_near_miss9.2%self_harm_near_miss3.2%sexual_exploitation_near_miss0.0%violent_facilitation_near_miss0.0%weapons_cbrn_near_miss0.0%
nsfwRuns todayT2, input and output0.974 mean 0.868 worst, Maltese
- Mean language F1
- 0.974Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.63Calibrated on validation, objective macro_f1.
- Positives in test
- 622Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 14,014Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
**This threshold is not a tuned parameter, and the shipped policy default stays at 0.76.** Two seeds on the identical corpus read 0.63 and 0.86, and sweeping either seed's own validation split gives macro F1 0.8969 to 0.9190 across the whole range from 0.50 to 0.95: the curve is flat, so calibration is picking the argmax of noise rather than a real optimum. 0.76 is the value reviewed and shipped before this retrain and is unchanged by it.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Bulgarianbg | 24 | 0.96 | 0.92 | 0.03 | 0.936 |
| Croatianhr | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 24 | 0.96 | 1.00 | 0.03 | 0.980 |
| Dutchnl | 24 | 0.96 | 0.96 | 0.03 | 0.958 |
| Englishen | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 24 | 0.96 | 1.00 | 0.03 | 0.980 |
| Frenchfr | 24 | 0.96 | 1.00 | 0.03 | 0.980 |
| Germande | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 24 | 0.92 | 0.96 | 0.06 | 0.939 |
| Hungarianhu | 24 | 0.96 | 1.00 | 0.03 | 0.980 |
| Irishga | 23 | 1.00 | 0.83 | 0.00 | 0.905 |
| Italianit | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 24 | 0.89 | 1.00 | 0.09 | 0.941 |
| Maltesemtnot in the base model's pretraining | 24 | 0.79 | 0.96 | 0.18 | 0.868 |
| Polishpl | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 24 | 0.96 | 1.00 | 0.03 | 0.980 |
| Romanianro | 24 | 1.00 | 0.92 | 0.00 | 0.957 |
| Slovaksk | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Sloveniansl | 23 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 24 | 0.96 | 1.00 | 0.03 | 0.980 |
| Swedishsv | 24 | 0.96 | 0.96 | 0.03 | 0.958 |
| Turkishtr | 24 | 0.96 | 1.00 | 0.03 | 0.980 |
Weakest languages. Maltese at 0.868, Irish at 0.905, Bulgarian at 0.936, Greek at 0.939, Lithuanian at 0.941. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
art_history_nudity1.3%breastfeeding_parenting1.3%clinical_anatomy1.3%mundane_account_access1.3%mundane_informational0.0%mundane_operational0.0%mundane_transactional0.0%news_report_of_violence2.6%romantic_non_explicit16.7%sex_education2.6%surgical_description0.0%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
politenessRuns todayT2, output0.978 mean 0.826 worst, Maltese
- Mean language F1
- 0.978Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.1Calibrated on validation, objective macro_f1.
- Positives in test
- 545Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 14,739Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
**This threshold is not a tuned parameter, and the shipped policy default stays at 0.89.** Two seeds on the identical corpus read 0.10 and 0.36, a spread of 0.26, the same shape as nsfw's calibration on the same retrain campaign. Per-language quality is stable across the two seeds (mean spread 0.0015), the threshold picked from a flat curve is not. 0.89 is the value reviewed and shipped before this retrain and is unchanged by it.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 21 | 0.95 | 1.00 | 0.03 | 0.977 |
| Bulgarianbg | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Croatianhr | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Dutchnl | 21 | 0.91 | 1.00 | 0.05 | 0.955 |
| Englishen | 21 | 0.95 | 1.00 | 0.03 | 0.977 |
| Estonianet | 21 | 0.91 | 1.00 | 0.05 | 0.955 |
| Finnishfi | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 21 | 0.95 | 0.95 | 0.03 | 0.952 |
| Germande | 21 | 1.00 | 0.95 | 0.00 | 0.976 |
| Greekel | 21 | 1.00 | 0.90 | 0.00 | 0.950 |
| Hungarianhu | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Irishga | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Italianit | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 21 | 0.76 | 0.90 | 0.16 | 0.826 |
| Polishpl | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Romanianro | 21 | 0.95 | 1.00 | 0.03 | 0.977 |
| Slovaksk | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Sloveniansl | 21 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 21 | 1.00 | 0.95 | 0.00 | 0.976 |
| Swedishsv | 20 | 0.87 | 1.00 | 0.08 | 0.930 |
| Turkishtr | 21 | 0.95 | 1.00 | 0.03 | 0.977 |
Weakest languages. Maltese at 0.826, Swedish at 0.930, Greek at 0.950, French at 0.952, Estonian at 0.955. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
bad_news_delivered_well1.5%brief_but_courteous1.5%firm_refusal_polite0.8%mundane_account_access2.6%mundane_informational0.0%mundane_operational0.0%mundane_transactional1.3%neutral_professional0.0%warm7.7%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
regulated_adviceRuns todayT2, output0.986 mean 0.900 worst, Maltese
- Mean language F1
- 0.986Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.88Calibrated on validation, objective macro_f1.
- Positives in test
- 572Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 13,502Synthetic, generated per language by claude-haiku-4-5.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Bulgarianbg | 22 | 1.00 | 0.95 | 0.00 | 0.977 |
| Croatianhr | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 22 | 1.00 | 0.95 | 0.00 | 0.977 |
| Danishda | 22 | 1.00 | 0.95 | 0.00 | 0.977 |
| Dutchnl | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 22 | 1.00 | 0.95 | 0.00 | 0.977 |
| Estonianet | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 22 | 1.00 | 0.95 | 0.00 | 0.977 |
| Irishga | 22 | 0.91 | 0.95 | 0.07 | 0.933 |
| Italianit | 22 | 1.00 | 0.95 | 0.00 | 0.977 |
| Latvianlv | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 22 | 1.00 | 0.82 | 0.00 | 0.900 |
| Polishpl | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Romanianro | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Slovaksk | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Sloveniansl | 22 | 1.00 | 0.95 | 0.00 | 0.977 |
| Spanishes | 22 | 1.00 | 0.95 | 0.00 | 0.977 |
| Swedishsv | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 22 | 1.00 | 1.00 | 0.00 | 1.000 |
Weakest languages. Maltese at 0.900, Irish at 0.933, Bulgarian at 0.977, Czech at 0.977, Danish at 0.977. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
definition0.0%general_risk0.0%historical_fact0.0%hypothetical0.0%mundane_account_access1.8%mundane_informational0.0%mundane_operational0.0%mundane_transactional0.0%process_description0.8%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
toxicityRuns todayT2, input and output0.992 mean 0.950 worst, Swedish
- Mean language F1
- 0.992Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.81Calibrated on validation, objective macro_f1.
- Positives in test
- 518Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 13,778Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Bulgarianbg | 20 | 0.95 | 1.00 | 0.03 | 0.976 |
| Croatianhr | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 20 | 0.95 | 1.00 | 0.03 | 0.976 |
| Dutchnl | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 20 | 1.00 | 0.95 | 0.00 | 0.974 |
| Germande | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 19 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 20 | 1.00 | 0.95 | 0.00 | 0.974 |
| Irishga | 19 | 1.00 | 1.00 | 0.00 | 1.000 |
| Italianit | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 20 | 0.95 | 1.00 | 0.03 | 0.976 |
| Maltesemtnot in the base model's pretraining | 20 | 0.91 | 1.00 | 0.06 | 0.952 |
| Polishpl | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Romanianro | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Slovaksk | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Sloveniansl | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Swedishsv | 20 | 0.95 | 0.95 | 0.03 | 0.950 |
| Turkishtr | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
Weakest languages. Swedish at 0.950, Maltese at 0.952, French at 0.974, Hungarian at 0.974, Bulgarian at 0.976. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
angry_but_civil3.9%civil_disagreement0.0%harsh_criticism_of_work0.0%mundane_informational0.0%mundane_operational1.3%mundane_transactional0.0%profanity_without_target1.0%quoted_abuse_in_complaint0.0%reclaimed_ingroup0.0%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
groundednessRuns todayT3, outputscored on pairs0.602 pair accuracy
Whether the claims in an answer are supported by the sources it was given. Scored on pairs, a grounded and a not-grounded reading of the same claim, because that is the confusion the detector exists to resolve: three earlier models scored string similarity instead and were refused. Ships disabled by the default policy; a policy that needs it turns it on.
- Pair accuracy
- 0.602Both halves of a grounded and not-grounded pair judged correctly.
- F1, grounded
- 0.820779 examples.
- F1, not_grounded
- 0.761779 examples.
- Export gate
- 0 of 300Decisions changed by the FP16 export.
| Language | N | Pair accuracy |
|---|---|---|
| Azerbaijaniaz | 58 | 0.793 |
| Bulgarianbg | 54 | 0.796 |
| Croatianhr | 64 | 0.672 |
| Czechcs | 64 | 0.781 |
| Danishda | 66 | 0.849 |
| Dutchnl | 56 | 0.804 |
| Englishen | 62 | 0.839 |
| Estonianet | 64 | 0.797 |
| Finnishfi | 58 | 0.810 |
| Frenchfr | 64 | 0.734 |
| Germande | 60 | 0.867 |
| Greekel | 60 | 0.800 |
| Hungarianhu | 68 | 0.765 |
| Irishga | 58 | 0.810 |
| Italianit | 66 | 0.758 |
| Latvianlv | 46 | 0.826 |
| Lithuanianlt | 58 | 0.759 |
| Maltesemtnot in the base model's pretraining | 62 | 0.758 |
| Polishpl | 64 | 0.734 |
| Portuguesept | 50 | 0.900 |
| Romanianro | 60 | 0.767 |
| Slovaksk | 64 | 0.813 |
| Sloveniansl | 60 | 0.817 |
| Spanishes | 50 | 0.840 |
| Swedishsv | 62 | 0.823 |
| Turkishtr | 60 | 0.800 |
Weakest languages by pair accuracy: Croatian at 0.672, French at 0.734, Polish at 0.734. None of the three is outside the base model's pretraining: the corpus, not pretraining, is what bounds these scores.
topic_scopeRuns todayT3, inputone node or none0.859 top-1
Backed by flowxai/topic-scope-v3: XLM-RoBERTa large and a two-layer decision head. It reads the message, a question and every node the policy supplies, and answers one node or "none of these", with a probability calibrated on validation rows. It replaced flowxai/topic-scope-v2, the same head on XLM-RoBERTa base, which had replaced flowxai/topic-scope, a bi-encoder that scored similarity to each node, could not read a node described by exclusion and had no way to answer none. The figures below are all three models on the same held-out rows. The rows are synthetic, generated per language, and split by taxonomy, so they measure how the model carries to taxonomies it has not seen, not an error rate on real traffic. A second training run with a different seed scores 0.869 on the unseen deployment types. It measures 146 ms p95 at the reference input with a three-node taxonomy, against a 300 ms budget. Its cost grows with the node text, and the library's budget test holds it under 300 ms at 40 nodes, the most it is given.
- Unseen deployment types
- 0.859v2 0.804, bi-encoder 0.479, 11,497 rows.
- Unseen taxonomies, trained types
- 0.887v2 0.808, bi-encoder 0.471, 11,071 rows.
- Calibration error
- 0.0118ECE of the top answer, unseen deployment types.
- Languages
- 26Scored separately.
| Language | N | Top-1 accuracy |
|---|---|---|
| Azerbaijaniaz | 413 | 0.816 |
| Bulgarianbg | 445 | 0.825 |
| Croatianhr | 430 | 0.893 |
| Czechcs | 412 | 0.862 |
| Danishda | 551 | 0.844 |
| Dutchnl | 406 | 0.862 |
| Englishen | 471 | 0.828 |
| Estonianet | 489 | 0.851 |
| Finnishfi | 495 | 0.885 |
| Frenchfr | 493 | 0.888 |
| Germande | 546 | 0.824 |
| Greekel | 462 | 0.857 |
| Hungarianhu | 383 | 0.859 |
| Irishga | 427 | 0.808 |
| Italianit | 372 | 0.895 |
| Latvianlv | 389 | 0.895 |
| Lithuanianlt | 403 | 0.861 |
| Maltesemtnot in the base model's pretraining | 412 | 0.750 |
| Polishpl | 434 | 0.876 |
| Portuguesept | 449 | 0.895 |
| Romanianro | 454 | 0.861 |
| Slovaksk | 425 | 0.896 |
| Sloveniansl | 414 | 0.882 |
| Spanishes | 496 | 0.899 |
| Swedishsv | 419 | 0.888 |
| Turkishtr | 407 | 0.840 |
Where the models differ most: a message that names a topic only to rule it out, 0.728 against 0.667 for v2 and 0.327 for the bi-encoder, over 672 rows; a message about a sibling of an allowed node, 0.864 against 0.778 for v2 and 0.380 for the bi-encoder, over 2,212 rows; a message that belongs to no node, where the right answer is none, 0.980 against 0.979 for v2 and 0.906 for the bi-encoder, over 817 rows. The bi-encoder's none bar (0.875) was chosen on validation, so the comparison is not against an untuned baseline. Sources: reports/typed_decisions_a3_large_late_off050_compare.json and, for v2, reports/typed_decisions_a3_compare.json in the training repository.
On 400 hand-written probes in 10 languages, run through the library, it passes 0.752 against 0.708 for v2. The known weakness shows there: short, plain in-scope questions get "none of these" in 70 of 200 cases (v2: 73). The library logs a none answer by default rather than blocking it. Source: reports/topic_scope_v3_probes.json.
Latency, two clusters and a budget
A millisecond figure means nothing without an input length and a thread count, so every figure here describes one reference: 87 tokens of prose on 1 thread, on CPU.
What threads buy
The default stays at one thread: a library that quietly takes the host's cores is worse than one that is honestly slower, and a policy can raise it deliberately. Inside a 94-token window, cost is 0.241 ms per token; past it, a second forward pass adds 7.12 ms at once.
Cost against input length
A sweep across input lengths shows the shape a single point cannot: close to linear inside a window, stepping by most of a forward pass at each window boundary.
The caveats that change the reading
Every language was generated in that language, which is what makes 26 affordable. These are in-distribution results, not a claim about production traffic.
A per-language F1 of 1.000 can rest on a handful of examples: toxicity has 518 positives across 26 languages. The support column in the tables above is what tells strong evidence from weak.
At the 0.5 default, four of these detectors scored 0.000 in all 26 languages. Every figure here is quoted at its calibrated threshold.
The numbers above measure piiguard, pii's default. A policy can select cee-pii instead, which needs a GPU and has no per-language evaluation table yet, so it is not comparable to the rows above. See it on the models page.
Same-class comparisons
Comparisons against Presidio (PII) and Llama Guard (moderation) are being run with the same harness and will be published here with the artifacts.
Where these numbers come from
Each detector has an evaluation report, a calibration record, a corpus manifest carrying a content hash, and an ONNX export manifest carrying the artifact hashes and the quantisation drift. This site reads those files and renders them. It does not hold a second copy of any number.