Skip to content

Reproducible numbers

What it detects, and what it costs

Every classifier is scored separately in all 26 target languages at a calibrated threshold, on synthetic corpora generated per language. The shape first; every table is one click below it.

26languages, every classifier scored separately in each
0.974 to 0.992mean per-language F1 across the evaluated detectors
13–146 msp95 per model-backed detector at 87 tokens, one thread
33 of 33detectors in the catalogue run today

Accuracy, mean and weakest language

One row per evaluated detector. The dot pair is the honest summary of 26 numbers: where the detector sits on average, and the single language where it is worst.

0.50.60.70.80.91.0F1, per-language, test split, calibrated thresholdgibberishCroatian 0.950.992toxicitySwedish 0.950.992injectionMaltese 0.880.989regulated_adviceMaltese 0.900.986biasMaltese 0.940.983moderationMaltese 0.860.980politenessMaltese 0.830.978nsfwMaltese 0.870.974
Each rule runs from the detector's weakest language, the amber dot, to its mean over all 26. The tail is the finding: a mean hides exactly the languages this figure pins down.

Two detectors are deliberately not on the figure, because their tasks are scored differently. topic_scope picks one taxonomy node or "none of these": top-1 accuracy 0.859 on taxonomies from deployment types it was not trained on, where the model it replaced scores 0.804 and the original bi-encoder 0.479 on the same rows. groundedness is scored on grounded and not-grounded pairs: pair accuracy 0.602. Both have their detail in their own rows below.

Every row, per detector and per language

An aggregate hides the tail, so each detector opens to its full 26-row table: precision, recall, false positive rate and F1 per language, with the corpus size and threshold that produced them.

gibberishRuns todayT1, input0.992 mean 0.946 worst, Croatian

Mean language F1
0.992Mean of the 26 rows below, not the macro F1 that calibration optimised.
Threshold
0.92Calibrated on validation, objective macro_f1.
Positives in test
779Counts tp plus fn. Negatives are scored too and produce the FPR column.
Corpus
15,491Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
gibberish: precision, recall, false positive rate and F1 per language, on the test split at threshold 0.92. Every language is listed, including the ones that score zero.
LanguagePositivesPRFPRF1
Azerbaijaniaz291.001.000.001.000
Bulgarianbg321.001.000.001.000
Croatianhr291.000.900.000.946
Czechcs321.001.000.001.000
Danishda321.001.000.001.000
Dutchnl291.001.000.001.000
Englishen301.001.000.001.000
Estonianet311.001.000.001.000
Finnishfi310.970.970.030.968
Frenchfr310.971.000.030.984
Germande301.000.930.000.966
Greekel311.001.000.001.000
Hungarianhu301.001.000.001.000
Irishga301.000.970.000.983
Italianit301.001.000.001.000
Latvianlv301.001.000.001.000
Lithuanianlt301.000.970.000.983
Maltesemtnot in the base model's pretraining290.970.970.030.966
Polishpl281.001.000.001.000
Portuguesept300.971.000.030.984
Romanianro281.001.000.001.000
Slovaksk291.001.000.001.000
Sloveniansl301.001.000.001.000
Spanishes291.001.000.001.000
Swedishsv301.001.000.001.000
Turkishtr291.001.000.001.000

Weakest languages. Croatian at 0.946, German at 0.966, Maltese at 0.966, Finnish at 0.968, Irish at 0.983. Published rather than dropped.

False positives on near misses

These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.

  • abbreviation_heavy0.0%
  • code_or_identifier0.0%
  • mixed_language_valid1.3%
  • mundane_account_access0.0%
  • mundane_informational0.0%
  • mundane_operational0.0%
  • mundane_transactional0.0%
  • proper_nouns2.6%
  • short_but_valid1.4%
  • typo_ridden_but_readable0.0%

Exported to ONNX at opset 17, traced at 32 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.

biasRuns todayT2, output0.983 mean 0.942 worst, Maltese

Mean language F1
0.983Mean of the 26 rows below, not the macro F1 that calibration optimised.
Threshold
0.77Calibrated on validation, objective macro_f1.
Positives in test
2064Counts tp plus fn. Negatives are scored too and produce the FPR column.
Corpus
45,421Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
bias: precision, recall, false positive rate and F1 per language, on the test split at threshold 0.77. Every language is listed, including the ones that score zero.
LanguagePositivesPRFPRF1
Azerbaijaniaz800.921.000.070.958
Bulgarianbg800.990.970.010.981
Croatianhr801.001.000.001.000
Czechcs800.991.000.010.994
Danishda801.001.000.001.000
Dutchnl790.991.000.010.994
Englishen800.960.990.030.975
Estonianet790.961.000.030.981
Finnishfi800.990.990.010.988
Frenchfr800.981.000.020.988
Germande800.960.990.030.975
Greekel801.000.990.000.994
Hungarianhu800.950.990.040.969
Irishga770.970.960.020.967
Italianit801.001.000.001.000
Latvianlv800.980.990.020.981
Lithuanianlt790.990.990.010.987
Maltesemtnot in the base model's pretraining780.950.940.040.942
Polishpl800.981.000.020.988
Portuguesept760.970.970.020.974
Romanianro801.000.990.000.994
Slovaksk800.990.990.010.988
Sloveniansl801.000.990.000.994
Spanishes761.000.970.000.987
Swedishsv800.990.990.010.988
Turkishtr800.950.970.040.963

Weakest languages. Maltese at 0.942, Azerbaijani at 0.958, Turkish at 0.963, Irish at 0.967, Hungarian at 0.969. Published rather than dropped.

False positives on near misses

These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.

  • counter_stereotype1.1%
  • demographic_statistic0.0%
  • discussing_bias1.7%
  • inclusive_phrasing9.6%
  • mundane_informational0.0%
  • mundane_operational0.0%
  • mundane_transactional0.0%
  • neutral_description0.0%

Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.

injectionRuns todayT2, input0.989 mean 0.882 worst, Maltese

Mean language F1
0.989Mean of the 26 rows below, not the macro F1 that calibration optimised.
Threshold
0.02Calibrated on validation, objective recall_at_fpr.
Positives in test
1080Counts tp plus fn. Negatives are scored too and produce the FPR column.
Corpus
45,541Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
injection: precision, recall, false positive rate and F1 per language, on the test split at threshold 0.02. Every language is listed, including the ones that score zero.
LanguagePositivesPRFPRF1
Azerbaijaniaz421.001.000.001.000
Bulgarianbg411.001.000.001.000
Croatianhr420.981.000.010.988
Czechcs420.951.000.020.977
Danishda410.981.000.010.988
Dutchnl411.000.980.000.988
Englishen411.001.000.001.000
Estonianet411.001.000.001.000
Finnishfi421.001.000.001.000
Frenchfr411.001.000.001.000
Germande421.001.000.001.000
Greekel421.001.000.001.000
Hungarianhu420.951.000.020.977
Irishga420.980.980.010.976
Italianit421.001.000.001.000
Latvianlv410.981.000.010.988
Lithuanianlt421.001.000.001.000
Maltesemtnot in the base model's pretraining420.800.980.080.882
Polishpl411.001.000.001.000
Portuguesept420.981.000.010.988
Romanianro420.981.000.010.988
Slovaksk411.001.000.001.000
Sloveniansl420.981.000.010.988
Spanishes400.981.000.010.988
Swedishsv411.001.000.001.000
Turkishtr421.001.000.001.000

Weakest languages. Maltese at 0.882, Irish at 0.976, Czech at 0.977, Hungarian at 0.977, Spanish at 0.988. Published rather than dropped.

False positives on near misses

These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.

  • meta_question1.2%
  • mundane_account_access0.0%
  • mundane_informational0.0%
  • mundane_operational0.5%
  • mundane_transactional0.5%
  • ordinary_instruction0.0%
  • ordinary_question0.6%
  • quoted_attack0.9%
  • roleplay_benign1.2%
  • security_discussion0.9%
  • technical_identifiers0.6%
  • technical_payload0.6%

Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.

moderationRuns todayT2, input and output0.980 mean 0.857 worst, Maltese

Mean language F1
0.980Mean of the 26 rows below, not the macro F1 that calibration optimised.
Threshold
0.84Calibrated on validation, objective macro_f1.
Positives in test
1623Counts tp plus fn. Negatives are scored too and produce the FPR column.
Corpus
33,708Synthetic, generated per language by claude-haiku-4-5.
moderation: precision, recall, false positive rate and F1 per language, on the test split at threshold 0.84. Every language is listed, including the ones that score zero.
LanguagePositivesPRFPRF1
Azerbaijaniaz670.950.940.050.947
Bulgarianbg691.000.970.000.985
Croatianhr541.000.980.000.991
Czechcs700.990.990.020.986
Danishda621.000.980.000.992
Dutchnl650.980.950.020.969
Englishen600.981.000.010.992
Estonianet560.980.980.010.982
Finnishfi611.000.980.000.992
Frenchfr611.001.000.001.000
Germande611.000.970.000.983
Greekel601.000.950.000.974
Hungarianhu601.001.000.001.000
Irishga690.970.880.030.924
Italianit581.000.970.000.983
Latvianlv650.970.980.030.977
Lithuanianlt641.001.000.001.000
Maltesemtnot in the base model's pretraining590.850.860.130.857
Polishpl681.001.000.001.000
Portuguesept741.001.000.001.000
Romanianro621.000.980.000.992
Slovaksk591.000.980.000.992
Sloveniansl621.000.980.000.992
Spanishes530.961.000.030.982
Swedishsv611.001.000.001.000
Turkishtr631.000.950.000.976

Weakest languages. Maltese at 0.857, Irish at 0.924, Azerbaijani at 0.947, Dutch at 0.969, Greek at 0.974. Published rather than dropped.

False positives on near misses

These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.

  • cyber_intrusion_near_miss0.0%
  • defamation_near_miss0.0%
  • election_integrity_near_miss1.9%
  • extremism_near_miss7.8%
  • fraud_deception_near_miss3.5%
  • hate_incitement_near_miss0.0%
  • illicit_drugs_near_miss0.0%
  • mundane_account_access2.8%
  • mundane_enterprise1.0%
  • mundane_informational0.0%
  • mundane_operational0.0%
  • mundane_transactional0.0%
  • property_crime_near_miss9.2%
  • self_harm_near_miss3.2%
  • sexual_exploitation_near_miss0.0%
  • violent_facilitation_near_miss0.0%
  • weapons_cbrn_near_miss0.0%
nsfwRuns todayT2, input and output0.974 mean 0.868 worst, Maltese

Mean language F1
0.974Mean of the 26 rows below, not the macro F1 that calibration optimised.
Threshold
0.63Calibrated on validation, objective macro_f1.
Positives in test
622Counts tp plus fn. Negatives are scored too and produce the FPR column.
Corpus
14,014Synthetic, generated per language by ollama/gpt-oss:120b-cloud.

**This threshold is not a tuned parameter, and the shipped policy default stays at 0.76.** Two seeds on the identical corpus read 0.63 and 0.86, and sweeping either seed's own validation split gives macro F1 0.8969 to 0.9190 across the whole range from 0.50 to 0.95: the curve is flat, so calibration is picking the argmax of noise rather than a real optimum. 0.76 is the value reviewed and shipped before this retrain and is unchanged by it.

nsfw: precision, recall, false positive rate and F1 per language, on the test split at threshold 0.63. Every language is listed, including the ones that score zero.
LanguagePositivesPRFPRF1
Azerbaijaniaz241.001.000.001.000
Bulgarianbg240.960.920.030.936
Croatianhr241.001.000.001.000
Czechcs241.001.000.001.000
Danishda240.961.000.030.980
Dutchnl240.960.960.030.958
Englishen241.001.000.001.000
Estonianet241.001.000.001.000
Finnishfi240.961.000.030.980
Frenchfr240.961.000.030.980
Germande241.001.000.001.000
Greekel240.920.960.060.939
Hungarianhu240.961.000.030.980
Irishga231.000.830.000.905
Italianit241.001.000.001.000
Latvianlv241.001.000.001.000
Lithuanianlt240.891.000.090.941
Maltesemtnot in the base model's pretraining240.790.960.180.868
Polishpl241.001.000.001.000
Portuguesept240.961.000.030.980
Romanianro241.000.920.000.957
Slovaksk241.001.000.001.000
Sloveniansl231.001.000.001.000
Spanishes240.961.000.030.980
Swedishsv240.960.960.030.958
Turkishtr240.961.000.030.980

Weakest languages. Maltese at 0.868, Irish at 0.905, Bulgarian at 0.936, Greek at 0.939, Lithuanian at 0.941. Published rather than dropped.

False positives on near misses

These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.

  • art_history_nudity1.3%
  • breastfeeding_parenting1.3%
  • clinical_anatomy1.3%
  • mundane_account_access1.3%
  • mundane_informational0.0%
  • mundane_operational0.0%
  • mundane_transactional0.0%
  • news_report_of_violence2.6%
  • romantic_non_explicit16.7%
  • sex_education2.6%
  • surgical_description0.0%

Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.

politenessRuns todayT2, output0.978 mean 0.826 worst, Maltese

Mean language F1
0.978Mean of the 26 rows below, not the macro F1 that calibration optimised.
Threshold
0.1Calibrated on validation, objective macro_f1.
Positives in test
545Counts tp plus fn. Negatives are scored too and produce the FPR column.
Corpus
14,739Synthetic, generated per language by ollama/gpt-oss:120b-cloud.

**This threshold is not a tuned parameter, and the shipped policy default stays at 0.89.** Two seeds on the identical corpus read 0.10 and 0.36, a spread of 0.26, the same shape as nsfw's calibration on the same retrain campaign. Per-language quality is stable across the two seeds (mean spread 0.0015), the threshold picked from a flat curve is not. 0.89 is the value reviewed and shipped before this retrain and is unchanged by it.

politeness: precision, recall, false positive rate and F1 per language, on the test split at threshold 0.1. Every language is listed, including the ones that score zero.
LanguagePositivesPRFPRF1
Azerbaijaniaz210.951.000.030.977
Bulgarianbg211.001.000.001.000
Croatianhr211.001.000.001.000
Czechcs211.001.000.001.000
Danishda211.001.000.001.000
Dutchnl210.911.000.050.955
Englishen210.951.000.030.977
Estonianet210.911.000.050.955
Finnishfi211.001.000.001.000
Frenchfr210.950.950.030.952
Germande211.000.950.000.976
Greekel211.000.900.000.950
Hungarianhu211.001.000.001.000
Irishga211.001.000.001.000
Italianit211.001.000.001.000
Latvianlv211.001.000.001.000
Lithuanianlt211.001.000.001.000
Maltesemtnot in the base model's pretraining210.760.900.160.826
Polishpl211.001.000.001.000
Portuguesept211.001.000.001.000
Romanianro210.951.000.030.977
Slovaksk211.001.000.001.000
Sloveniansl211.001.000.001.000
Spanishes211.000.950.000.976
Swedishsv200.871.000.080.930
Turkishtr210.951.000.030.977

Weakest languages. Maltese at 0.826, Swedish at 0.930, Greek at 0.950, French at 0.952, Estonian at 0.955. Published rather than dropped.

False positives on near misses

These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.

  • bad_news_delivered_well1.5%
  • brief_but_courteous1.5%
  • firm_refusal_polite0.8%
  • mundane_account_access2.6%
  • mundane_informational0.0%
  • mundane_operational0.0%
  • mundane_transactional1.3%
  • neutral_professional0.0%
  • warm7.7%

Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.

regulated_adviceRuns todayT2, output0.986 mean 0.900 worst, Maltese

Mean language F1
0.986Mean of the 26 rows below, not the macro F1 that calibration optimised.
Threshold
0.88Calibrated on validation, objective macro_f1.
Positives in test
572Counts tp plus fn. Negatives are scored too and produce the FPR column.
Corpus
13,502Synthetic, generated per language by claude-haiku-4-5.
regulated_advice: precision, recall, false positive rate and F1 per language, on the test split at threshold 0.88. Every language is listed, including the ones that score zero.
LanguagePositivesPRFPRF1
Azerbaijaniaz221.001.000.001.000
Bulgarianbg221.000.950.000.977
Croatianhr221.001.000.001.000
Czechcs221.000.950.000.977
Danishda221.000.950.000.977
Dutchnl221.001.000.001.000
Englishen221.000.950.000.977
Estonianet221.001.000.001.000
Finnishfi221.001.000.001.000
Frenchfr221.001.000.001.000
Germande221.001.000.001.000
Greekel221.001.000.001.000
Hungarianhu221.000.950.000.977
Irishga220.910.950.070.933
Italianit221.000.950.000.977
Latvianlv221.001.000.001.000
Lithuanianlt221.001.000.001.000
Maltesemtnot in the base model's pretraining221.000.820.000.900
Polishpl221.001.000.001.000
Portuguesept221.001.000.001.000
Romanianro221.001.000.001.000
Slovaksk221.001.000.001.000
Sloveniansl221.000.950.000.977
Spanishes221.000.950.000.977
Swedishsv221.001.000.001.000
Turkishtr221.001.000.001.000

Weakest languages. Maltese at 0.900, Irish at 0.933, Bulgarian at 0.977, Czech at 0.977, Danish at 0.977. Published rather than dropped.

False positives on near misses

These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.

  • definition0.0%
  • general_risk0.0%
  • historical_fact0.0%
  • hypothetical0.0%
  • mundane_account_access1.8%
  • mundane_informational0.0%
  • mundane_operational0.0%
  • mundane_transactional0.0%
  • process_description0.8%

Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.

toxicityRuns todayT2, input and output0.992 mean 0.950 worst, Swedish

Mean language F1
0.992Mean of the 26 rows below, not the macro F1 that calibration optimised.
Threshold
0.81Calibrated on validation, objective macro_f1.
Positives in test
518Counts tp plus fn. Negatives are scored too and produce the FPR column.
Corpus
13,778Synthetic, generated per language by ollama/gpt-oss:120b-cloud.
toxicity: precision, recall, false positive rate and F1 per language, on the test split at threshold 0.81. Every language is listed, including the ones that score zero.
LanguagePositivesPRFPRF1
Azerbaijaniaz201.001.000.001.000
Bulgarianbg200.951.000.030.976
Croatianhr201.001.000.001.000
Czechcs201.001.000.001.000
Danishda200.951.000.030.976
Dutchnl201.001.000.001.000
Englishen201.001.000.001.000
Estonianet201.001.000.001.000
Finnishfi201.001.000.001.000
Frenchfr201.000.950.000.974
Germande201.001.000.001.000
Greekel191.001.000.001.000
Hungarianhu201.000.950.000.974
Irishga191.001.000.001.000
Italianit201.001.000.001.000
Latvianlv201.001.000.001.000
Lithuanianlt200.951.000.030.976
Maltesemtnot in the base model's pretraining200.911.000.060.952
Polishpl201.001.000.001.000
Portuguesept201.001.000.001.000
Romanianro201.001.000.001.000
Slovaksk201.001.000.001.000
Sloveniansl201.001.000.001.000
Spanishes201.001.000.001.000
Swedishsv200.950.950.030.950
Turkishtr201.001.000.001.000

Weakest languages. Swedish at 0.950, Maltese at 0.952, French at 0.974, Hungarian at 0.974, Bulgarian at 0.976. Published rather than dropped.

False positives on near misses

These registers contain no positives at all. They exist to be mistaken for the class, and the only number that matters is how often the detector was fooled.

  • angry_but_civil3.9%
  • civil_disagreement0.0%
  • harsh_criticism_of_work0.0%
  • mundane_informational0.0%
  • mundane_operational1.3%
  • mundane_transactional0.0%
  • profanity_without_target1.0%
  • quoted_abuse_in_complaint0.0%
  • reclaimed_ingroup0.0%

Exported to ONNX at opset 17, traced at 96 tokens. The INT8 artifact is 535 MB against 1112 MB for fp32, and compressing it changed 0 of 300 test decisions. That last number is the one that decides whether shipping INT8 is honest.

groundednessRuns todayT3, outputscored on pairs0.602 pair accuracy

Whether the claims in an answer are supported by the sources it was given. Scored on pairs, a grounded and a not-grounded reading of the same claim, because that is the confusion the detector exists to resolve: three earlier models scored string similarity instead and were refused. Ships disabled by the default policy; a policy that needs it turns it on.

Pair accuracy
0.602Both halves of a grounded and not-grounded pair judged correctly.
F1, grounded
0.820779 examples.
F1, not_grounded
0.761779 examples.
Export gate
0 of 300Decisions changed by the FP16 export.
groundedness: pair accuracy per language, on the test split at the calibrated threshold. Every language is listed, including the weakest.
LanguageNPair accuracy
Azerbaijaniaz580.793
Bulgarianbg540.796
Croatianhr640.672
Czechcs640.781
Danishda660.849
Dutchnl560.804
Englishen620.839
Estonianet640.797
Finnishfi580.810
Frenchfr640.734
Germande600.867
Greekel600.800
Hungarianhu680.765
Irishga580.810
Italianit660.758
Latvianlv460.826
Lithuanianlt580.759
Maltesemtnot in the base model's pretraining620.758
Polishpl640.734
Portuguesept500.900
Romanianro600.767
Slovaksk640.813
Sloveniansl600.817
Spanishes500.840
Swedishsv620.823
Turkishtr600.800

Weakest languages by pair accuracy: Croatian at 0.672, French at 0.734, Polish at 0.734. None of the three is outside the base model's pretraining: the corpus, not pretraining, is what bounds these scores.

topic_scopeRuns todayT3, inputone node or none0.859 top-1

Backed by flowxai/topic-scope-v3: XLM-RoBERTa large and a two-layer decision head. It reads the message, a question and every node the policy supplies, and answers one node or "none of these", with a probability calibrated on validation rows. It replaced flowxai/topic-scope-v2, the same head on XLM-RoBERTa base, which had replaced flowxai/topic-scope, a bi-encoder that scored similarity to each node, could not read a node described by exclusion and had no way to answer none. The figures below are all three models on the same held-out rows. The rows are synthetic, generated per language, and split by taxonomy, so they measure how the model carries to taxonomies it has not seen, not an error rate on real traffic. A second training run with a different seed scores 0.869 on the unseen deployment types. It measures 146 ms p95 at the reference input with a three-node taxonomy, against a 300 ms budget. Its cost grows with the node text, and the library's budget test holds it under 300 ms at 40 nodes, the most it is given.

Unseen deployment types
0.859v2 0.804, bi-encoder 0.479, 11,497 rows.
Unseen taxonomies, trained types
0.887v2 0.808, bi-encoder 0.471, 11,071 rows.
Calibration error
0.0118ECE of the top answer, unseen deployment types.
Languages
26Scored separately.
topic_scope: top-1 accuracy per language over the offered nodes plus none, on taxonomies from unseen deployment types. Every language is listed, including the weakest.
LanguageNTop-1 accuracy
Azerbaijaniaz4130.816
Bulgarianbg4450.825
Croatianhr4300.893
Czechcs4120.862
Danishda5510.844
Dutchnl4060.862
Englishen4710.828
Estonianet4890.851
Finnishfi4950.885
Frenchfr4930.888
Germande5460.824
Greekel4620.857
Hungarianhu3830.859
Irishga4270.808
Italianit3720.895
Latvianlv3890.895
Lithuanianlt4030.861
Maltesemtnot in the base model's pretraining4120.750
Polishpl4340.876
Portuguesept4490.895
Romanianro4540.861
Slovaksk4250.896
Sloveniansl4140.882
Spanishes4960.899
Swedishsv4190.888
Turkishtr4070.840

Where the models differ most: a message that names a topic only to rule it out, 0.728 against 0.667 for v2 and 0.327 for the bi-encoder, over 672 rows; a message about a sibling of an allowed node, 0.864 against 0.778 for v2 and 0.380 for the bi-encoder, over 2,212 rows; a message that belongs to no node, where the right answer is none, 0.980 against 0.979 for v2 and 0.906 for the bi-encoder, over 817 rows. The bi-encoder's none bar (0.875) was chosen on validation, so the comparison is not against an untuned baseline. Sources: reports/typed_decisions_a3_large_late_off050_compare.json and, for v2, reports/typed_decisions_a3_compare.json in the training repository.

On 400 hand-written probes in 10 languages, run through the library, it passes 0.752 against 0.708 for v2. The known weakness shows there: short, plain in-scope questions get "none of these" in 70 of 200 cases (v2: 73). The library logs a none answer by default rather than blocking it. Source: reports/topic_scope_v3_probes.json.

Latency, two clusters and a budget

A millisecond figure means nothing without an input length and a thread count, so every figure here describes one reference: 87 tokens of prose on 1 thread, on CPU.

060120180240measured p95, ms, 87-token input, one thread225 ms budget18 rules, under 1 ms9 encoders, one forward passmoderationnsfwtopic_scope
One dot per detector, and no middle band: a detector is either a rule that costs almost nothing or an encoder that costs one forward pass, which is why the tiers schedule encoders rather than trimming them.

What threads buy

24.24 ms1 thread, the library default
26.77 ms2 threads
17.09 ms4 threads
19.5 ms8 threads

The default stays at one thread: a library that quietly takes the host's cores is worse than one that is honestly slower, and a policy can raise it deliberately. Inside a 94-token window, cost is 0.241 ms per token; past it, a second forward pass adds 7.12 ms at once.

Cost against input length

A sweep across input lengths shows the shape a single point cannot: close to linear inside a window, stepping by most of a forward pass at each window boundary.

075150225300ms168794128input length, tokens225 ms budget94-token window2nd pass +7.12 ms1 thread
Cost is linear inside a window and steps at each boundary: within one 94-token window the slope is 0.241 ms per token, and text past it needs a second forward pass. The 225 ms line is the per-scan budget, not a limit on input length: longer text costs more passes, each budgeted on its own.

The caveats that change the reading

Synthetic corpora

Every language was generated in that language, which is what makes 26 affordable. These are in-distribution results, not a claim about production traffic.

Support counts positives

A per-language F1 of 1.000 can rest on a handful of examples: toxicity has 518 positives across 26 languages. The support column in the tables above is what tells strong evidence from weak.

Thresholds are part of the result

At the 0.5 default, four of these detectors scored 0.000 in all 26 languages. Every figure here is quoted at its calibrated threshold.

pii has a second model, not in these tables

The numbers above measure piiguard, pii's default. A policy can select cee-pii instead, which needs a GPU and has no per-language evaluation table yet, so it is not comparable to the rows above. See it on the models page.

Same-class comparisons

Comparisons against Presidio (PII) and Llama Guard (moderation) are being run with the same harness and will be published here with the artifacts.

Where these numbers come from

Each detector has an evaluation report, a calibration record, a corpus manifest carrying a content hash, and an ONNX export manifest carrying the artifact hashes and the quantisation drift. This site reads those files and renders them. It does not hold a second copy of any number.

regenerate.sh
# The site's accuracy data is generated, not hand written.node scripts/gen-evals.mjs # It reads, per detector:#   reports/artifacts/<detector>-full/<detector>_eval.json#   reports/artifacts/<detector>-full/calibration.json#   reports/artifacts/<detector>-full/onnx/export_manifest.json#   reports/data/<detector>_manifest.json
The full detector setThe models behind these numbers

Get started

Check what crosses.

Read the docsSee the numbers