Reference
Measured performance
Generated from the library's own benchmark run.
Generated from the library's own benchmark run.
Every score below carries the number of examples behind it, and the n means
column says what those examples are. A count of positives is not a count
of everything evaluated, and the two are not interchangeable. Where the
evaluation recorded no count at all the cell reads not recorded, which
means the score is unverified rather than good or bad.
Detectors
| Detector | Tier | Status | Metric | Macro | Median | Worst | Test cases | n means |
|---|---|---|---|---|---|---|---|---|
banned_terms | T1 | built | – | – | – | – | – | – |
bias | T2 | built | f1 | 0.983 | 0.987 | 0.942 | 2064 | positive examples |
code_present | T1 | built | – | – | – | – | – | – |
confusables | T1 | built | – | – | – | – | – | – |
disclosure | T0 | built | – | – | – | – | – | – |
encoded_payload | T1 | built | – | – | – | – | – | – |
gibberish | T1 | built | f1 | 0.992 | 1.000 | 0.946 | 779 | positive examples |
groundedness | T3 | built | exact_match_accuracy | 0.796 | 0.800 | 0.672 | 1558 | examples evaluated |
infra_leakage | T1 | built | – | – | – | – | – | – |
injection | T2 | built | f1 | 0.989 | 1.000 | 0.882 | 1080 | positive examples |
internal_domains | T1 | built | – | – | – | – | – | – |
invisible_text | T0 | built | – | – | – | – | – | – |
json_schema | T1 | built | – | – | – | – | – | – |
language_id | T1 | built | – | – | – | – | – | – |
link_integrity | T1 | built | – | – | – | – | – | – |
markup_injection | T1 | built | – | – | – | – | – | – |
moderation | T2 | built | f1 | 0.980 | 0.988 | 0.857 | 1623 | positive examples |
nsfw | T2 | built | f1 | 0.974 | 0.980 | 0.868 | 622 | positive examples |
output_format | T1 | built | – | – | – | – | – | – |
output_leakage | T1 | built | – | – | – | – | – | – |
pii | T1 | built | – | – | – | – | – | – |
politeness | T2 | built | f1 | 0.978 | 1.000 | 0.826 | 545 | positive examples |
postal_code | T1 | built | – | – | – | – | – | – |
regulated_advice | T2 | built | f1 | 0.986 | 1.000 | 0.900 | 572 | positive examples |
repetition | T1 | built | – | – | – | – | – | – |
secrets | T0 | built | – | – | – | – | – | – |
sql_injection | T1 | built | – | – | – | – | – | – |
summary_support | T1 | built | – | – | – | – | – | – |
system_prompt_leakage | T1 | built | – | – | – | – | – | – |
token_limit | T1 | built | – | – | – | – | – | – |
topic_scope | T3 | built | top1_accuracy | 0.859 | 0.862 | 0.750 | 11497 | examples evaluated |
toxicity | T2 | built | f1 | 0.992 | 1.000 | 0.950 | 518 | positive examples |
url_reachability | T3 | built | – | – | – | – | – | – |
Caveats
bias: the score above is per language and asks whether the detector fires at all, not which of its 5 labels applies. Per label the weakest with support is gender at 0.9533, against a macro of 0.9826 here.gibberish: the score above is per language and asks whether the detector fires at all, not which of its 3 labels applies. Per label the weakest with support is repetition at 0.9709, against a macro of 0.9915 here.groundedness: the score above is per language and asks whether the detector fires at all, not which of its 2 labels applies. Per label the weakest with support is not_grounded at 0.7612, against a macro of 0.7965 here.groundedness: scored on the 1558 of 3214 test rows the model never trained on. A re-split had moved 1656 of its training rows into the test split, and they are excluded with their pair partners. 1120 of those are from registers added to the corpus after the model trained, so the figure mostly measures cases it never learned.groundedness: no calibrated threshold recorded, so this detector runs at the policy default. Several detectors in this family reported nothing at 0.5 while separating positives from negatives well below it.injection: the score above is per language and asks whether the detector fires at all, not which of its 3 labels applies. Per label the weakest with support is jailbreak at 0.9603, against a macro of 0.9891 here.moderation: the score above is per language and asks whether the detector fires at all, not which of its 12 labels applies. Per label the weakest with support is fraud_deception at 0.7972, against a macro of 0.9795 here.nsfw: the score above is per language and asks whether the detector fires at all, not which of its 2 labels applies. Per label the weakest with support is sexual at 0.9492, against a macro of 0.9738 here.regulated_advice: the score above is per language and asks whether the detector fires at all, not which of its 3 labels applies. Per label the weakest with support is legal_advice at 0.8758, against a macro of 0.9864 here.toxicity: the score above is per language and asks whether the detector fires at all, not which of its 4 labels applies. Per label the weakest with support is harassment at 0.9723, against a macro of 0.9915 here.
Latency
At a 396 character reference input, 1 thread, CPUExecutionProvider. Romanian prose with no entities in it, so this measures the cost of looking rather than the cost of finding.
| Detector | p95 ms | Budget ms | note |
|---|---|---|---|
banned_terms | 0.167 | 5.0 | – |
bias | 19.097 | 225.0 | – |
code_present | 0.008 | 5.0 | – |
confusables | 0.124 | 5.0 | – |
disclosure | 0.029 | 5.0 | – |
encoded_payload | 0.194 | 5.0 | – |
gibberish | 27.287 | 225.0 | – |
groundedness | 12.851 | 300.0 | – |
infra_leakage | 0.047 | 5.0 | – |
injection | 40.669 | 225.0 | – |
internal_domains | 0.160 | 5.0 | – |
invisible_text | 0.024 | 5.0 | – |
json_schema | 0.001 | 5.0 | – |
language_id | 0.223 | 5.0 | – |
link_integrity | 0.167 | 5.0 | – |
markup_injection | 0.170 | 5.0 | – |
moderation | 41.637 | 150.0 | – |
nsfw | 42.529 | 225.0 | – |
output_format | 0.001 | 5.0 | – |
output_leakage | 24.939 | 225.0 | – |
pii | 24.354 | 225.0 | – |
politeness | 18.987 | 225.0 | – |
postal_code | 0.001 | 5.0 | – |
regulated_advice | 40.609 | 225.0 | – |
repetition | 0.292 | 5.0 | – |
secrets | 0.028 | 1.0 | – |
sql_injection | 0.148 | 5.0 | – |
summary_support | 0.553 | 5.0 | – |
system_prompt_leakage | 0.196 | 5.0 | the unconfigured path |
token_limit | 0.001 | 5.0 | the unconfigured path |
topic_scope | 146.453 | 300.0 | – |
toxicity | 19.062 | 225.0 | – |
url_reachability | 0.005 | 3000.0 | the unconfigured path |
Model variants
A model a detector runs only when a policy names it, not the one it runs by
default. pii's row above always measures piiguard; a variant here needs
its own row because a policy that selects it gets none of the figures above.
| Model | Status | p95 ms | Needs | Quality |
|---|---|---|---|---|
cee-pii | measured | 102.597 | gpu | not recorded |
This page is generated from docs/reference/performance.md in the library repository. Read it as markdown, or edit it at the source.