Skip to content

Reference

Measured performance

Generated from the library's own benchmark run.

Generated from the library's own benchmark run.

Every score below carries the number of examples behind it, and the n means column says what those examples are. A count of positives is not a count of everything evaluated, and the two are not interchangeable. Where the evaluation recorded no count at all the cell reads not recorded, which means the score is unverified rather than good or bad.

Detectors

DetectorTierStatusMetricMacroMedianWorstTest casesn means
banned_termsT1built
biasT2builtf10.9771.0000.824264positive examples
code_presentT1built
disclosureT0built
encoded_payloadT1built
gibberishT1builtf10.9660.9580.870276positive examples
groundednessT3not built
injectionT2builtf10.9700.9830.727357positive examples
internal_domainsT1built
invisible_textT0built
json_schemaT1built
language_idT1built
markup_injectionT1built
moderationT2not built
nsfwT2builtf10.9340.9520.600259positive examples
output_formatT1built
output_leakageT1built
piiT1built
politenessT2builtf10.9620.9680.788392positive examples
postal_codeT1built
regulated_adviceT2builtf10.9951.0000.957622positive examples
repetitionT1built
secretsT0built
sql_injectionT1built
summary_supportT1built
system_prompt_leakageT1built
token_limitT1built
topic_scopeT3builttop1_accuracy0.8650.8570.375175examples evaluated
toxicityT2builtf10.9921.0000.950518positive examples
url_reachabilityT3built

Caveats

  • bias: 12 of 26 languages have fewer than 10 positive examples: az, bg, da, de, el, fi, ga, hu, lv, mt, pl, sk. Their individual scores are indicative rather than measured.
  • gibberish: 2 of 26 languages have fewer than 10 positive examples: bg, en. Their individual scores are indicative rather than measured.
  • nsfw: 1 of 26 languages have fewer than 10 positive examples: ga. Their individual scores are indicative rather than measured.
  • topic_scope: 26 of 26 languages have fewer than 10 examples evaluated: az, bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv, tr. Their individual scores are indicative rather than measured.
  • topic_scope: no calibrated threshold recorded, so this detector runs at the policy default. Several detectors in this family reported nothing at 0.5 while separating positives from negatives well below it.

Latency

At a 396 character reference input, 1 thread, CPUExecutionProvider. Romanian prose with no entities in it, so this measures the cost of looking rather than the cost of finding.

Detectorp95 msBudget msnote
banned_terms0.1685.0
bias21.327225.0
code_present0.0095.0
disclosure0.0385.0
encoded_payload0.2055.0the unconfigured path
gibberish27.528225.0
injection27.622225.0
internal_domains0.2345.0
invisible_text0.0385.0
json_schema0.0015.0
language_id0.3625.0
markup_injection0.2395.0
nsfw27.544225.0
output_format0.0015.0
output_leakage27.991225.0
pii27.636225.0
politeness30.319225.0
postal_code0.0025.0
regulated_advice30.430225.0
repetition0.4975.0
secrets0.0481.0
sql_injection0.2385.0
summary_support0.9215.0
system_prompt_leakage0.3185.0the unconfigured path
token_limit0.0015.0the unconfigured path
topic_scope46.298300.0
toxicity27.224225.0
url_reachability0.0073000.0the unconfigured path

This page is generated from docs/reference/performance.md in the library repository. Read it as markdown, or edit it at the source.