Skip to content

2026-08-11 · 6 min read

Our regulated-advice detector scores 0.983. Its recall is 0.75.

Both numbers are real, they measure different slices of the same test set, and the one we would have put on a slide is the one that matters least.

Updated 2026-10-01: The regulated-advice model has been retrained since this was written. Its mean per-language F1 is now 0.986 (0.9864), against the 0.983 discussed here. Per-label recall is now 0.86 to 0.91 rather than 0.75 and 0.79, and financial advice has 260 positives in the test split rather than none. The tables in the body read the current evaluation, so they can differ from the figures quoted in the text.

The regulated_advice detector separates explaining a financial instrument from recommending one. It scores a mean per-language F1 of 0.986 across all 26 target languages. That is the number that would have gone on a slide, and it is close to useless on its own.

What the headline number is averaging

It is the mean of 26 per-language F1 values, computed over 572 positive examples in the test split. Twenty-five of those languages sit above 0.95. The weakest is mt at 0.900, which is Maltese, and Maltese is absent from the base model's pretraining, so that one is a property of XLM-RoBERTa rather than of our data.

Read that way it looks like a solved problem. The label breakdown says otherwise.

The same test set, sliced by label

Per-label results at threshold 0.88. Precision holds. Recall does not.
LabelSupportPRF1
financial_advice2600.8750.9110.893
legal_advice1560.8930.8590.876
medical_advice1560.9500.8590.902

Precision is excellent: 0.950 on medical advice, meaning when it fires it is right. Recall is 0.91 and 0.86 and 0.86. Between a fifth and a quarter of the advice in the test set walks straight past it.

A guardrail with high precision and mediocre recall is a specific kind of dangerous. It never annoys anyone, so nobody turns it off, and it produces a clean evidence record every time it misses.

The label that has no test data at all

0 labels have zero positives in the test split: . That is the flagship case. The detector exists because a banking assistant should not tell someone which fund to buy, and the split we evaluated on contains no examples of it.

The 0.983 is not wrong. It is averaging over language, and language is not the axis on which this detector is weak. Nobody hid anything; the number was simply answering a question we were not asking.

Where it does hold up

The register breakdown is the part that survives scrutiny, and it is the part that was hardest to build. These are the positive registers:

Positive registers. A personal recommendation is the easy case; a suitability claim is the one that reads like an explanation.
RegisterSupportPRF1
directive1331.0000.9700.985
personal_recommendation3281.0001.0001.000
suitability_claim1111.0000.9190.958

And these are the hard negatives, written to sit as close to the class as possible without being in it. They contain no positives at all, so the only column that means anything is the false positive rate:

  • definition at 0.0% false positives
  • general_risk at 0.0% false positives
  • historical_fact at 0.0% false positives
  • hypothetical at 0.0% false positives
  • mundane_account_access at 1.8% false positives
  • mundane_informational at 0.0% false positives
  • mundane_operational at 0.0% false positives
  • mundane_transactional at 0.0% false positives
  • process_description at 0.8% false positives

A definition of an ETF, a general statement about risk, a historical fact, a hypothetical, a description of a process. Those are the sentences a keyword filter destroys, and the detector leaves them alone almost perfectly. That is the thing worth reporting, and it is nowhere near the top of the page in the version of this we nearly published.

The corpus

13,502 examples across 26 languages, generated per language rather than translated, of which 7,788 are unlabelled negatives. Generated by claude-haiku-4-5 against prompt version advice_sys_v2_mundane, content hash 25b93152e30d. Synthetic and in-distribution, which is what makes 26 languages affordable and is also why none of this is a claim about production traffic.

The threshold is 0.88 and it is the uncalibrated default. This detector predates the calibration sweep the others went through. Given that recall is the weak axis, a lower threshold is the obvious next experiment, and it has not been run.

What we changed

The benchmarks page leads with the caveats rather than the results, and every detector shows its positive count next to its F1 so a reader can see how much evidence is behind a number. The three things on the fix list for this detector are: get positives for the empty label into the test split, calibrate the threshold against recall rather than accept the default, and evaluate against real transcripts instead of generated ones.

Until then the per-language table is published in full, including Maltese at 0.900.

regulated_advice, per language, at threshold 0.88. Every row, including the ones that fail.
LanguagePositivesPRFPRF1
Azerbaijaniaz221.001.000.001.000
Bulgarianbg221.000.950.000.977
Croatianhr221.001.000.001.000
Czechcs221.000.950.000.977
Danishda221.000.950.000.977
Dutchnl221.001.000.001.000
Englishen221.000.950.000.977
Estonianet221.001.000.001.000
Finnishfi221.001.000.001.000
Frenchfr221.001.000.001.000
Germande221.001.000.001.000
Greekel221.001.000.001.000
Hungarianhu221.000.950.000.977
Irishga220.910.950.070.933
Italianit221.000.950.000.977
Latvianlv221.001.000.001.000
Lithuanianlt221.001.000.001.000
Maltesemtnot in the base model's pretraining221.000.820.000.900
Polishpl221.001.000.001.000
Portuguesept221.001.000.001.000
Romanianro221.001.000.001.000
Slovaksk221.001.000.001.000
Sloveniansl221.000.950.000.977
Spanishes221.000.950.000.977
Swedishsv221.001.000.001.000
Turkishtr221.001.000.001.000

References

  1. [1]regulated_advice evaluation reportreports/artifacts/regulatedadvice-full/regulated_advice_eval.json, in the training repository. Every figure in this post is read from it at build time.
  2. [2]Corpus manifestreports/data/regulated_advice_manifest.json. 11,436 examples across 26 languages, content hash 48de4bb155a8.
  3. [3]How support, precision and recall are computedborder_train/metrics.py. Support is tp plus fn, so it counts positives only and the negatives are what the false positive rate measures.
  4. [4]The full per-language table
  5. [5]The use case this detector exists for
  6. [6]XLM-RoBERTaThe base model for every classifier in the set. Maltese is absent from its pretraining corpus, which is why Maltese is the weakest language.
All posts

Get started

Check what crosses.

Read the docsSee the numbers