Skip to content

Reference

Measured performance

Generated from the library's own benchmark run.

Generated from the library's own benchmark run.

Every score below carries the number of examples behind it, and the n means column says what those examples are. A count of positives is not a count of everything evaluated, and the two are not interchangeable. Where the evaluation recorded no count at all the cell reads not recorded, which means the score is unverified rather than good or bad.

Detectors

DetectorTierStatusMetricMacroMedianWorstTest casesn means
banned_termsT1built––––––
biasT2builtf10.9830.9870.9422064positive examples
code_presentT1built––––––
confusablesT1built––––––
disclosureT0built––––––
encoded_payloadT1built––––––
gibberishT1builtf10.9921.0000.946779positive examples
groundednessT3builtexact_match_accuracy0.7960.8000.6721558examples evaluated
infra_leakageT1built––––––
injectionT2builtf10.9891.0000.8821080positive examples
internal_domainsT1built––––––
invisible_textT0built––––––
json_schemaT1built––––––
language_idT1built––––––
link_integrityT1built––––––
markup_injectionT1built––––––
moderationT2builtf10.9800.9880.8571623positive examples
nsfwT2builtf10.9740.9800.868622positive examples
output_formatT1built––––––
output_leakageT1built––––––
piiT1built––––––
politenessT2builtf10.9781.0000.826545positive examples
postal_codeT1built––––––
regulated_adviceT2builtf10.9861.0000.900572positive examples
repetitionT1built––––––
secretsT0built––––––
sql_injectionT1built––––––
summary_supportT1built––––––
system_prompt_leakageT1built––––––
token_limitT1built––––––
topic_scopeT3builttop1_accuracy0.8590.8620.75011497examples evaluated
toxicityT2builtf10.9921.0000.950518positive examples
url_reachabilityT3built––––––

Caveats

  • bias: the score above is per language and asks whether the detector fires at all, not which of its 5 labels applies. Per label the weakest with support is gender at 0.9533, against a macro of 0.9826 here.
  • gibberish: the score above is per language and asks whether the detector fires at all, not which of its 3 labels applies. Per label the weakest with support is repetition at 0.9709, against a macro of 0.9915 here.
  • groundedness: the score above is per language and asks whether the detector fires at all, not which of its 2 labels applies. Per label the weakest with support is not_grounded at 0.7612, against a macro of 0.7965 here.
  • groundedness: scored on the 1558 of 3214 test rows the model never trained on. A re-split had moved 1656 of its training rows into the test split, and they are excluded with their pair partners. 1120 of those are from registers added to the corpus after the model trained, so the figure mostly measures cases it never learned.
  • groundedness: no calibrated threshold recorded, so this detector runs at the policy default. Several detectors in this family reported nothing at 0.5 while separating positives from negatives well below it.
  • injection: the score above is per language and asks whether the detector fires at all, not which of its 3 labels applies. Per label the weakest with support is jailbreak at 0.9603, against a macro of 0.9891 here.
  • moderation: the score above is per language and asks whether the detector fires at all, not which of its 12 labels applies. Per label the weakest with support is fraud_deception at 0.7972, against a macro of 0.9795 here.
  • nsfw: the score above is per language and asks whether the detector fires at all, not which of its 2 labels applies. Per label the weakest with support is sexual at 0.9492, against a macro of 0.9738 here.
  • regulated_advice: the score above is per language and asks whether the detector fires at all, not which of its 3 labels applies. Per label the weakest with support is legal_advice at 0.8758, against a macro of 0.9864 here.
  • toxicity: the score above is per language and asks whether the detector fires at all, not which of its 4 labels applies. Per label the weakest with support is harassment at 0.9723, against a macro of 0.9915 here.

Latency

At a 396 character reference input, 1 thread, CPUExecutionProvider. Romanian prose with no entities in it, so this measures the cost of looking rather than the cost of finding.

Detectorp95 msBudget msnote
banned_terms0.1675.0–
bias19.097225.0–
code_present0.0085.0–
confusables0.1245.0–
disclosure0.0295.0–
encoded_payload0.1945.0–
gibberish27.287225.0–
groundedness12.851300.0–
infra_leakage0.0475.0–
injection40.669225.0–
internal_domains0.1605.0–
invisible_text0.0245.0–
json_schema0.0015.0–
language_id0.2235.0–
link_integrity0.1675.0–
markup_injection0.1705.0–
moderation41.637150.0–
nsfw42.529225.0–
output_format0.0015.0–
output_leakage24.939225.0–
pii24.354225.0–
politeness18.987225.0–
postal_code0.0015.0–
regulated_advice40.609225.0–
repetition0.2925.0–
secrets0.0281.0–
sql_injection0.1485.0–
summary_support0.5535.0–
system_prompt_leakage0.1965.0the unconfigured path
token_limit0.0015.0the unconfigured path
topic_scope146.453300.0–
toxicity19.062225.0–
url_reachability0.0053000.0the unconfigured path

Model variants

A model a detector runs only when a policy names it, not the one it runs by default. pii's row above always measures piiguard; a variant here needs its own row because a policy that selects it gets none of the figures above.

ModelStatusp95 msNeedsQuality
cee-piimeasured102.597gpunot recorded

This page is generated from docs/reference/performance.md in the library repository. Read it as markdown, or edit it at the source.

Get started

Check what crosses.

Read the docsSee the numbers