Reproducible numbers
What it detects, and what it costs
Every classifier is scored separately in all 26 target languages at a calibrated threshold, on synthetic corpora generated per language. The shape first; every table is one click below it.
Accuracy, mean and weakest language
One row per evaluated detector. The dot pair is the honest summary of 26 numbers: where the detector sits on average, and the single language where it is worst.
Two detectors are deliberately not on the figure. topic_scope is a retrieval task with no threshold: top-1 accuracy 0.857 and sibling rejection 0.916 over 155 near-miss pairs, with the detail in its own row below. And groundedness publishes nothing, because none of the three models trained for it is good enough to adopt.
Every row, per detector and per language
An aggregate hides the tail, so each detector opens to its full 26-row table: precision, recall, false positive rate and F1 per language, with the corpus size and threshold that produced them.
gibberishRuns todayT1, input0.966 mean 0.870 worst, Czech
- Mean language F1
- 0.966Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.05Calibrated on validation, objective macro_f1.
- Positives in test
- 276Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 7,630Synthetic, generated per language by claude-haiku-4-5.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 10 | 0.91 | 1.00 | 0.06 | 0.952 |
| Bulgarianbg | 9 | 1.00 | 1.00 | 0.00 | 1.000 |
| Croatianhr | 10 | 0.91 | 1.00 | 0.06 | 0.952 |
| Czechcs | 11 | 0.83 | 0.91 | 0.11 | 0.870 |
| Danishda | 11 | 0.92 | 1.00 | 0.06 | 0.957 |
| Dutchnl | 11 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 9 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 11 | 0.92 | 1.00 | 0.05 | 0.957 |
| Finnishfi | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 11 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 11 | 1.00 | 0.91 | 0.00 | 0.952 |
| Greekel | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 11 | 1.00 | 1.00 | 0.00 | 1.000 |
| Irishga | 12 | 0.92 | 0.92 | 0.06 | 0.917 |
| Italianit | 11 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 12 | 0.86 | 1.00 | 0.11 | 0.923 |
| Polishpl | 11 | 1.00 | 0.91 | 0.00 | 0.952 |
| Portuguesept | 12 | 0.92 | 1.00 | 0.06 | 0.960 |
| Romanianro | 11 | 0.85 | 1.00 | 0.11 | 0.917 |
| Slovaksk | 11 | 0.92 | 1.00 | 0.06 | 0.957 |
| Sloveniansl | 11 | 0.91 | 0.91 | 0.06 | 0.909 |
| Spanishes | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Swedishsv | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 10 | 0.91 | 1.00 | 0.06 | 0.952 |
Weakest languages. Czech at 0.870, Slovenian at 0.909, Irish at 0.917, Romanian at 0.917, Maltese at 0.923. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
abbreviation_heavy0.0%code_or_identifier1.4%mixed_language_valid0.0%proper_nouns17.1%short_but_valid0.0%typo_ridden_but_readable1.3%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and quantising it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
biasRuns todayT2, output0.977 mean 0.824 worst, Maltese
- Mean language F1
- 0.977Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.57Calibrated on validation, objective macro_f1.
- Positives in test
- 264Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 6,133Synthetic, generated per language by claude-opus-5.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 9 | 0.90 | 1.00 | 0.10 | 0.947 |
| Bulgarianbg | 8 | 1.00 | 0.88 | 0.00 | 0.933 |
| Croatianhr | 13 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 13 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 9 | 1.00 | 1.00 | 0.00 | 1.000 |
| Dutchnl | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 7 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 6 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 8 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 9 | 1.00 | 1.00 | 0.00 | 1.000 |
| Irishga | 9 | 1.00 | 0.89 | 0.00 | 0.941 |
| Italianit | 16 | 1.00 | 0.94 | 0.00 | 0.968 |
| Latvianlv | 8 | 1.00 | 0.88 | 0.00 | 0.933 |
| Lithuanianlt | 11 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 9 | 0.88 | 0.78 | 0.08 | 0.824 |
| Polishpl | 9 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 11 | 0.92 | 1.00 | 0.08 | 0.957 |
| Romanianro | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Slovaksk | 8 | 0.89 | 1.00 | 0.06 | 0.941 |
| Sloveniansl | 11 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Swedishsv | 13 | 1.00 | 0.92 | 0.00 | 0.960 |
| Turkishtr | 12 | 1.00 | 1.00 | 0.00 | 1.000 |
Weakest languages. Maltese at 0.824, Bulgarian at 0.933, Latvian at 0.933, Irish at 0.941, Slovak at 0.941. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
counter_stereotype2.0%demographic_statistic0.0%discussing_bias1.9%inclusive_phrasing1.8%mundane_informational0.0%mundane_operational3.0%mundane_transactional0.0%neutral_description0.0%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and quantising it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
injectionRuns todayT2, input0.970 mean 0.727 worst, Maltese
- Mean language F1
- 0.970Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.19Calibrated on validation, objective recall_at_fpr.
- Positives in test
- 357Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 11,599Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 14 | 1.00 | 0.93 | 0.00 | 0.963 |
| Bulgarianbg | 13 | 1.00 | 1.00 | 0.00 | 1.000 |
| Croatianhr | 13 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 13 | 1.00 | 0.92 | 0.00 | 0.960 |
| Danishda | 14 | 0.93 | 1.00 | 0.03 | 0.966 |
| Dutchnl | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 14 | 0.93 | 1.00 | 0.03 | 0.966 |
| Estonianet | 14 | 0.93 | 1.00 | 0.03 | 0.966 |
| Finnishfi | 14 | 0.93 | 1.00 | 0.03 | 0.966 |
| Frenchfr | 13 | 0.92 | 0.92 | 0.03 | 0.923 |
| Germande | 14 | 0.93 | 1.00 | 0.03 | 0.966 |
| Greekel | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Irishga | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Italianit | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 13 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 14 | 1.00 | 0.57 | 0.00 | 0.727 |
| Polishpl | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 14 | 0.88 | 1.00 | 0.07 | 0.933 |
| Romanianro | 14 | 0.93 | 1.00 | 0.03 | 0.966 |
| Slovaksk | 14 | 1.00 | 0.93 | 0.00 | 0.963 |
| Sloveniansl | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 12 | 1.00 | 1.00 | 0.00 | 1.000 |
| Swedishsv | 14 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 14 | 0.93 | 1.00 | 0.03 | 0.966 |
Weakest languages. Maltese at 0.727, French at 0.923, Portuguese at 0.933, Czech at 0.960, Azerbaijani at 0.963. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
meta_question2.9%mundane_informational0.0%mundane_operational0.0%mundane_transactional0.0%ordinary_instruction0.0%ordinary_question0.0%quoted_attack1.9%roleplay_benign1.9%security_discussion2.9%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and quantising it changed 1 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
nsfwRuns todayT2, input and output0.934 mean 0.600 worst, Maltese
- Mean language F1
- 0.934Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.76Calibrated on validation, objective macro_f1.
- Positives in test
- 259Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 9,969Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 10 | 0.83 | 1.00 | 0.07 | 0.909 |
| Bulgarianbg | 10 | 0.91 | 1.00 | 0.03 | 0.952 |
| Croatianhr | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 10 | 0.83 | 1.00 | 0.07 | 0.909 |
| Danishda | 10 | 0.91 | 1.00 | 0.03 | 0.952 |
| Dutchnl | 10 | 0.83 | 1.00 | 0.07 | 0.909 |
| Englishen | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 10 | 0.91 | 1.00 | 0.03 | 0.952 |
| Frenchfr | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 10 | 0.77 | 1.00 | 0.10 | 0.870 |
| Hungarianhu | 10 | 0.71 | 1.00 | 0.13 | 0.833 |
| Irishga | 9 | 0.88 | 0.78 | 0.03 | 0.824 |
| Italianit | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 10 | 0.83 | 1.00 | 0.07 | 0.909 |
| Maltesemtnot in the base model's pretraining | 10 | 0.60 | 0.60 | 0.13 | 0.600 |
| Polishpl | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 10 | 0.91 | 1.00 | 0.03 | 0.952 |
| Romanianro | 10 | 1.00 | 0.90 | 0.00 | 0.947 |
| Slovaksk | 10 | 1.00 | 1.00 | 0.00 | 1.000 |
| Sloveniansl | 10 | 0.91 | 1.00 | 0.03 | 0.952 |
| Spanishes | 10 | 0.90 | 0.90 | 0.03 | 0.900 |
| Swedishsv | 10 | 0.91 | 1.00 | 0.03 | 0.952 |
| Turkishtr | 10 | 0.91 | 1.00 | 0.03 | 0.952 |
Weakest languages. Maltese at 0.600, Irish at 0.824, Hungarian at 0.833, Greek at 0.870, Spanish at 0.900. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
art_history_nudity1.3%breastfeeding_parenting1.3%clinical_anatomy1.3%mundane_informational0.0%mundane_operational0.0%mundane_transactional0.0%news_report_of_violence5.1%romantic_non_explicit25.6%sex_education1.3%surgical_description0.0%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and quantising it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
politenessRuns todayT2, output0.962 mean 0.788 worst, Irish
- Mean language F1
- 0.962Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.89Calibrated on validation, objective macro_f1.
- Positives in test
- 392Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 10,227Synthetic, generated per language by claude-haiku-4-5.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 15 | 0.94 | 1.00 | 0.04 | 0.968 |
| Bulgarianbg | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Croatianhr | 15 | 0.83 | 1.00 | 0.12 | 0.909 |
| Czechcs | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 15 | 0.94 | 1.00 | 0.04 | 0.968 |
| Dutchnl | 15 | 0.93 | 0.93 | 0.04 | 0.933 |
| Englishen | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 15 | 0.94 | 1.00 | 0.04 | 0.968 |
| Hungarianhu | 15 | 0.94 | 1.00 | 0.04 | 0.968 |
| Irishga | 15 | 0.72 | 0.87 | 0.20 | 0.788 |
| Italianit | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 15 | 0.88 | 1.00 | 0.08 | 0.938 |
| Lithuanianlt | 16 | 0.94 | 1.00 | 0.04 | 0.970 |
| Maltesemtnot in the base model's pretraining | 15 | 0.87 | 0.87 | 0.08 | 0.867 |
| Polishpl | 15 | 0.94 | 1.00 | 0.04 | 0.968 |
| Portuguesept | 15 | 0.93 | 0.93 | 0.04 | 0.933 |
| Romanianro | 16 | 1.00 | 0.94 | 0.00 | 0.968 |
| Slovaksk | 15 | 1.00 | 0.93 | 0.00 | 0.966 |
| Sloveniansl | 15 | 0.94 | 1.00 | 0.04 | 0.968 |
| Spanishes | 15 | 1.00 | 0.93 | 0.00 | 0.966 |
| Swedishsv | 15 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 15 | 0.94 | 1.00 | 0.04 | 0.968 |
Weakest languages. Irish at 0.788, Maltese at 0.867, Croatian at 0.909, Dutch at 0.933, Portuguese at 0.933. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
bad_news_delivered_well0.8%brief_but_courteous3.1%firm_refusal_polite4.6%neutral_professional0.0%warm8.5%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and quantising it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
regulated_adviceRuns todayT2, output0.995 mean 0.957 worst, Hungarian
- Mean language F1
- 0.995Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.72Calibrated on validation, objective macro_f1.
- Positives in test
- 622Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 11,436Synthetic, generated per language by claude-haiku-4-5.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Bulgarianbg | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Croatianhr | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Dutchnl | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Germande | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 24 | 1.00 | 0.92 | 0.00 | 0.957 |
| Irishga | 24 | 1.00 | 0.96 | 0.00 | 0.979 |
| Italianit | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Maltesemtnot in the base model's pretraining | 23 | 1.00 | 0.96 | 0.00 | 0.978 |
| Polishpl | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 23 | 1.00 | 1.00 | 0.00 | 1.000 |
| Romanianro | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Slovaksk | 24 | 1.00 | 0.96 | 0.00 | 0.979 |
| Sloveniansl | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 24 | 1.00 | 0.96 | 0.00 | 0.979 |
| Swedishsv | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
| Turkishtr | 24 | 1.00 | 1.00 | 0.00 | 1.000 |
Weakest languages. Hungarian at 0.957, Maltese at 0.978, Spanish at 0.979, Irish at 0.979, Slovak at 0.979. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
definition0.0%general_risk0.0%historical_fact0.0%hypothetical0.0%process_description0.0%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and quantising it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
toxicityRuns todayT2, input and output0.992 mean 0.950 worst, Swedish
- Mean language F1
- 0.992Mean of the 26 rows below, not the macro F1 that calibration optimised.
- Threshold
- 0.81Calibrated on validation, objective macro_f1.
- Positives in test
- 518Counts tp plus fn. Negatives are scored too and produce the FPR column.
- Corpus
- 13,778Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
| Language | Positives | P | R | FPR | F1 |
|---|---|---|---|---|---|
| Azerbaijaniaz | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Bulgarianbg | 20 | 0.95 | 1.00 | 0.03 | 0.976 |
| Croatianhr | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Czechcs | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Danishda | 20 | 0.95 | 1.00 | 0.03 | 0.976 |
| Dutchnl | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Englishen | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Estonianet | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Finnishfi | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Frenchfr | 20 | 1.00 | 0.95 | 0.00 | 0.974 |
| Germande | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Greekel | 19 | 1.00 | 1.00 | 0.00 | 1.000 |
| Hungarianhu | 20 | 1.00 | 0.95 | 0.00 | 0.974 |
| Irishga | 19 | 1.00 | 1.00 | 0.00 | 1.000 |
| Italianit | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Latvianlv | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Lithuanianlt | 20 | 0.95 | 1.00 | 0.03 | 0.976 |
| Maltesemtnot in the base model's pretraining | 20 | 0.91 | 1.00 | 0.06 | 0.952 |
| Polishpl | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Portuguesept | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Romanianro | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Slovaksk | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Sloveniansl | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Spanishes | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
| Swedishsv | 20 | 0.95 | 0.95 | 0.03 | 0.950 |
| Turkishtr | 20 | 1.00 | 1.00 | 0.00 | 1.000 |
Weakest languages. Swedish at 0.950, Maltese at 0.952, French at 0.974, Hungarian at 0.974, Bulgarian at 0.976. Published rather than dropped.
False positives on near misses
These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.
angry_but_civil3.9%civil_disagreement0.0%harsh_criticism_of_work0.0%mundane_informational0.0%mundane_operational1.3%mundane_transactional0.0%profanity_without_target1.0%quoted_abuse_in_complaint0.0%reclaimed_ingroup0.0%
Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and quantising it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.
topic_scopeRuns todayT3, inputretrieval, not classification0.857 top-1
A retrieval task rather than a binary classifier, so it has no threshold, no precision and no recall. A published bi-encoder backs it and it runs today, measured at 46 ms p95 against its 300 ms budget at the reference input.
- Top-1 accuracy
- 0.857Against the taxonomy.
- Sibling rejection
- 0.916Rejecting a near neighbour in the taxonomy is the hard case.
- Corpus
- 3,265Synthetic, claude-haiku-4-5.
- Languages
- 26Scored separately.
Four of the five weakest nodes are in banking/, the one branch three levels deep, where siblings differ by a word: accounts/opening against accounts/closing. Telling those apart is what the sibling rejection figure measures, over 155 near-miss pairs. The per-node scores rest on three to thirteen examples each, so they are not quoted here.
Latency, two clusters and a budget
A millisecond figure means nothing without an input length and a thread count, so every figure here describes one reference: 87 tokens of prose on 1 thread, on CPU.
What threads buy
The default stays at one thread: a library that quietly takes the host's cores is worse than one that is honestly slower, and a policy can raise it deliberately. Inside a 94-token window, cost is 1.663 ms per token; past it, a second forward pass adds 33.25 ms at once.
Cost against input length
A sweep across input lengths shows the shape a single point cannot: close to linear inside a window, stepping by most of a forward pass at each window boundary.
The caveats that change the reading
Every language was generated in that language, which is what makes 26 affordable. These are in-distribution results, not a claim about production traffic.
A per-language F1 of 1.000 can rest on a handful of examples: nsfw has 259 positives across 26 languages. The support column in the tables above is what tells strong evidence from weak.
At the 0.5 default, four of these detectors scored 0.000 in all 26 languages. Every figure here is quoted at its calibrated threshold.
Same-class comparisons
Comparisons against Presidio (PII) and Llama Guard (moderation) are being run with the same harness and will be published here with the artifacts.
Where these numbers come from
Each detector has an evaluation report, a calibration record, a corpus manifest carrying a content hash, and an ONNX export manifest carrying the artifact hashes and the quantisation drift. This site reads those files and renders them. It does not hold a second copy of any number.