It runs where your data already is
CPU is the reference target, not a fallback. Weights are fetched once and cached, and after that a scan works with the network interface down, which a test asserts by making a socket call raise.
An open-source project by FlowX.AI
It runs on CPU with the network interface down. An embeddable Python library: two functions, a policy file a reviewer can read without knowing Python, and an evidence record that holds hashes rather than user text.
Built for teams that run models inside their own perimeter, in more languages than English, and have to prove what was checked.
01 · Tiers
The cheap checks run on everything. The expensive ones run only when something cheaper has found a reason, which is what keeps a scan inside its latency budget.
02 · Stamp
What ran, which model revision made each call, the resolved policy hash, and a hash of the input. Never the text itself. The record is canonical JSON, so its hash reproduces on another machine, and it can be signed with a key you hold.
03 · Local
Weights are fetched once and cached. After that a scan works with the network interface down, so the text you check never reaches a third party and there is no per-call bill.
piiguard against five hosted models on 30 documents. The benchmark
04 · Try it
There is no client to construct, no gateway to run, and nothing wraps your model call. You call two functions and decide what to do with what they return. This one is live: pick an example or type your own.
1from flowx_border import scan_input, scan_output, load_policy23policy = load_policy("border-code.yaml")45crossing = scan_input(user_text, policy)6if crossing.verdict == "block":7return refuse(crossing.evidence.record_id)89answer = your_model.complete(crossing.text)10out = scan_output(answer, policy)11archive(out.evidence)
Live: this calls flowx_border on a small CPU instance kept for this site, and the text above is sent to it. In your deployment the library runs inside your own process, offline.
05 · Architecture
A message passing between two services that already trust each other has been checked once, and checking it again costs latency without changing the answer. The cost of inspecting everything everywhere is the reason teams end up turning inspection off.
The design is borrowed from the Schengen model: internal checks could be abolished only because the external checks became strong, uniform, and governed by one shared rulebook.
06 · Catalogue
Tiers decide what runs when, so the expensive checks only run once something cheaper has found a reason.
Reading the table. Runs today means installed and callable now. Trained means the model exists but the detector is not yet callable; a policy that asks it to block or redact raises before any scan happens. Needs dependency means an optional extra must be installed. Needs network means the check makes outbound requests and is off unless a policy turns it on. In the F1 column, not a classifier marks rule-based detectors, which are deterministic and measured on latency only. No corpus yet means no evaluation set we trust exists; the cost column says whether its figure is measured or a budget.
| Detector | Tier | What it does | Needs | Backed by | Mean F1 | Cost | Status |
|---|---|---|---|---|---|---|---|
| secretsinput, outputruns today | T0 | Credentials in text on its way to the model: named key formats, plus a deliberately conservative entropy rule. | CPU | rule | not a classifier | 0.028 msmeasured | Runs today |
| piiinput, outputruns today | T1 | Personal data in input or output, as named entity spans with checksum validation where the identifier has one. | CPU | ner | per entity | 24.354 msmeasured | Runs today |
| injectioninputruns today | T2 | Attempts to talk the model out of its instructions. | CPU | classifier | 0.989 | 40.669 msmeasured | Runs today |
| moderationinput, outputruns today | T2 | Thirteen hazard categories in one pass, from violent crime to election misinformation. Replaces the capability Llama Guard and ShieldGemma provide, with weights this project can ship. | CPU | classifier | 0.980 | 41.637 msmeasured | Runs today |
| toxicityinput, outputruns today | T2 | Abusive or hateful language, in input or output. | CPU | classifier | 0.992 | 19.062 msmeasured | Runs today |
| regulated_adviceoutputruns today | T2 | Output that reads as regulated financial, legal or medical advice. | CPU | classifier | 0.986 | 40.609 msmeasured | Runs today |
| topic_scopeinputruns today | T3 | Whether a request is inside the subject matter the product covers. | CPU | classifier | no corpus yet | 146.453 msmeasured | Runs today |
| groundednessoutputruns today | T3 | Whether the claims in an answer are supported by the sources it was given. | CPU | classifier | no corpus yet | 12.851 msmeasured | Runs today |
07 · Why it is different
CPU is the reference target, not a fallback. Weights are fetched once and cached, and after that a scan works with the network interface down, which a test asserts by making a socket call raise.
The PII model is trained across all 26 target languages at once, each country's national identifier generated to pass its own checksum. Span-level F1 0.998, weakest language French at 0.977. Each of the 8 classifiers is scored in every language, and every row is published, including the ones that fail.
Every crossing produces a stamp: which detectors ran, which model revision each used, the resolved policy hash, and a hash of the text. Labels and scores, never the text. Optionally signed with a key you hold. The console above shows a real one.
| Language | Models |
|---|---|
| en | 19 |
| ro | 18 |
| hu | 15 |
| pl | 15 |
| de | 14 |
| fr | 14 |
| az | 13 |
| bg | 13 |
| cs | 13 |
| da | 13 |
| el | 13 |
| es | 13 |
| et | 13 |
| fi | 13 |
| ga | 13 |
| hr | 13 |
| it | 13 |
| lt | 13 |
| lv | 13 |
| mt | 13 |
| nl | 13 |
| pt | 13 |
| sk | 13 |
| sl | 13 |
| sv | 13 |
| tr | 13 |
08 · Latency
Every latency figure here describes 87 tokens of prose on 1 thread, the library default, so a scan does not take cores from the application it runs inside. Cost is close to linear at 0.241 ms per token. At the reference length the model-backed detectors measure between 13 and 146 ms p95 each, depending on the model.
09 · Weights
The models this library ships: the PII tagger and 10 classifiers. All on Hugging Face under Apache-2.0, each carrying its per-language evaluation in its card, and small enough to run on a CPU you already own. Browse our models.
Token classification, PII spans
Eight entity types: card, date, email, IBAN, location, national ID, person, phone. Identifiers are generated checksum-valid in training, so an IBAN that fails mod-97 is not reported as an IBAN.
Apache-2.0
pii: { model: piiguard }
Sequence classification
Backs the injection detector, for the text a retrieval tool fetched as much as the text a user typed.
Apache-2.0
Multi-label classification, 12 hazard categories
Twelve hazard categories in one pass, from violent crime to election misinformation. The taxonomy has thirteen; child_safety is left out of training on purpose, because it cannot be generated synthetically and needs a vetted source. The capability Llama Guard and ShieldGemma provide, with weights this project ships.
Apache-2.0
Sequence classification
Backs the regulated_advice detector: the line between explaining an instrument and recommending it.
Apache-2.0
10 · Scope
It does not sit in front of your model, it does not hold your traffic, and it does not wrap your model call. If it stops working, your application still runs, it just stops producing evidence.
Nothing here is certified, and no library can be. Obligations under the EU AI Act sit with the provider or deployer of a system, not with a dependency it installs. What this produces is an auditable record of which checks ran and what they found, which supports the evidence requirements of a governance process you run yourself. What it does support: the disclosure detector records whether an AI disclosure was present in each output, in 26 languages, which is evidence a transparency process can file.
It reads text and reports on text. It knows nothing about your authentication, your tool permissions, or what your agent is allowed to do with the answer it got.
A detector whose model is not published raises an error naming what is missing rather than returning an empty result, because a check that silently passes is worse than one that is absent.
Get started
Install it and read the record.
Weights are fetched once and cached. After that a scan needs no network, and nothing you scan reaches FlowX.AI.
Python 3.11 or newer. Apache-2.0. The package is flowx-border, the import is flowx_border, the GitHub org is flowx-ai, and the models live under flowxai on Hugging Face. Four spellings, one project.