22 KiB
Local sensitive-column classifier: TypeScript or Python
Date: 2026-09-02 Scope: library/runtime choice for a local classifier without a generative LLM. Synthetic data generation and model-boundary policy for description generation are deliberately out of scope.
Recommendation
Implement the mandatory classifier inside the TypeScript backend, with no learned model in its mandatory path. The agreed policy is dominated by deterministic work: type/declared-size rules, bounded string normalization, patterns, checksums, exact dictionaries, context terms, and "one match makes the column sensitive". TypeScript is fully adequate for that work and keeps the classifier in the process which already owns the catalog, source connectors, run lifecycle, cancellation, and the human override.
Use small, focused dependencies rather than a general NLP framework:
validatorfor syntax/checksum-oriented recognizers such as email, IP, MAC, IBAN, BIC, payment cards, and locale-aware Italian tax IDs, passports, identity cards and VAT numbers; its official API exposes these validators individually, so unused validators need not be imported (validator.js source and API).libphonenumber-js/maxfor international telephone parsing and digit-pattern validation. The project'smaxmetadata is intentionally more precise than its default/minimal metadata (libphonenumber-js validation documentation).- ThothII-owned recognizers for identifiers not covered adequately by a selected validator (for example the Italian driving licence) and organization-specific identifiers, each with a version, tests, context words, and checksum/semantic validation where the identifier defines one.
- A safe regular-expression boundary for administrator-authored rules.
node-re2offers a mostlyRegExp-compatible, ReDoS-resistant engine and a set API for matching many patterns, but is a native addon that downloads or builds a binary during installation. That packaging cost must be tested against the project images before adoption (node-re2 README). If arbitrary regular expressions are not exposed, ThothII can instead keep a reviewed built-in pattern set and bound every inspected value.
Do not add NLP.js, spaCy, Presidio, or a transformer merely to implement regexes and dictionaries. They do not make deterministic identifiers more semantically understandable. In particular, NLP.js documents useful Italian tokenization plus enum/regex/built-in entities, but its advertised built-ins are mostly the same formatted values already covered above; it does not advertise a ready-made Italian PII model for arbitrary people or clinical concepts (NLP.js official README).
Add an optional, non-generative statistical NER pack based on
fastino/gliner2-privacy-filter-PII-multi.
The selected model and its GLiNER2 runtime are Apache-2.0, its mDeBERTa base is MIT, and the model
is trained across seven languages including Italian. Run the official implementation in one warmed,
CPU-only Python worker; do not spawn it once per column or table. It returns entity evidence only.
The TypeScript classifier remains the sole owner of the final column assessment and invokes NER
only for bounded text that deterministic rules did not resolve and only while the scan deadline has
time remaining.
This is a positive Q29 decision, not a placeholder for later library selection. The production gate
is limited to pinning artifacts, producing an SBOM/license inventory, and measuring the selected
model on PSD fixtures and target CPUs. If the optional pack is absent or misses its deadline, the
normal coverage rule produces unknown; ThothII does not substitute a generative LLM.
Why the TypeScript baseline is sufficient
The classifier is not being asked to infer an open-ended privacy judgment from prose. It executes a versioned policy and produces evidence. Its mandatory recognizers can be expressed as:
| Rule family | Implementation | Model needed? |
|---|---|---|
declared text/CLOB/unbounded string or declared maximum >500 |
catalog metadata rule | no |
| an observed value longer than 500 characters | length rule, ideally detected in SQL before transferring the value | no |
| email, IP/MAC, URL, UUID, payment card, IBAN/BIC, phone | pattern plus format/checksum validator | no |
| Italian fiscal code/VAT/passport/identity card/driving licence | locale-aware validator where available, otherwise a reviewed country recognizer; pattern, context and checksum where defined | no |
| credentials and technical secrets | anchored prefixes, token shapes, entropy/character-class heuristics, and organization allow/deny rules | no |
| hospital- or customer-specific terms/codes | normalized exact dictionary or phrase matching | no |
| person/location/organization in otherwise unstructured short text | statistical NER is useful but probabilistic | yes, optional |
| clinical meaning in unstructured short Italian text | domain NER or a governed terminology; generic NER is not enough | optional and domain-specific |
This is not speculative parity with a Python implementation. Presidio's own supported-entity table shows that its global email, IBAN, credit-card and similar recognizers use pattern matching, validation, context and checksums, while its Italian fiscal-code/VAT/passport/identity-card/licence coverage is likewise rule-based (Presidio supported entities). Those mechanisms are portable; Presidio provides tested implementations and orchestration, not a unique Python-language capability.
For large dictionaries, exact normalized token/phrase matching is still deterministic NLP and does
not require a learned model. The policy format should describe the rule rather than the library,
for example kind: email, kind: checksum_id, kind: regex, kind: phrase_set, and
kind: max_length. This makes a future engine change possible without changing persisted policy or
assessment evidence.
What Python/Presidio would materially add
Presidio is the strongest Python option if ThothII later needs a broader recognizer ecosystem. Its
Analyzer composes rule-based recognizers, regexes, validation, NER and contextual enhancement; it
also exposes custom recognizers, batch processing and decision tracing
(Presidio Analyzer architecture). Its current
catalog includes Italian identifiers and its YAML registry supports custom patterns and language-
specific context. Recent source also contains a NoOpNlpEngine, so a deterministic-only Presidio
deployment can avoid loading spaCy; this removes model cost but not the Python runtime/process
boundary
(Presidio NoOpNlpEngine integration).
presidio-structured is conceptually close because it maps tabular columns/JSON keys to detected
entities, but its documented strategies select the most common, highest-confidence, or a mixed
entity. ThothII would still need its own deadline/sampling/unknown semantics and human-override
lifecycle; Presidio also lists improved mixed free-text/structured support as future work
(Presidio Structured).
That benefit does not currently justify using it as ThothII's mandatory engine:
- The backend currently has no PII/NLP dependencies and is an ESM TypeScript service
(
backend/package.json); the existing Python environment belongs to the separate harness/core layer and also has no Presidio or spaCy dependency (harness/pyproject.toml). A Python analyzer therefore means a new ownership and deployment boundary, not simply another import. - Presidio is Python-based and can be run as a package or REST container. Its REST API intentionally has no built-in authentication, so a sidecar would have to remain on a protected internal network and be guarded by infrastructure (Presidio FAQ, deployment).
- Presidio defaults to English. Adding Italian requires both an Italian NLP pipeline and adapted language-specific recognizers/context; language-agnostic regexes alone do not solve that setup (Presidio multilingual guidance).
- Presidio itself warns that automated detection cannot guarantee finding all sensitive data and documents an unavoidable false-positive/false-negative trade-off (Presidio FAQ). It cannot turn a negative sample into proof that a column is safe.
Learned NLP is not an LLM, but it is still a model
A spaCy NER pipeline does not generate text and is not an LLM. It is a statistical token classifier:
it predicts spans such as PERSON, LOCATION and ORGANIZATION. spaCy explicitly notes that the
predictions depend on the training examples and will not always be correct
(spaCy NER documentation). That makes it
a useful additional positive detector, not the authority which declares a sampled column safe.
spaCy publishes Italian CPU pipelines with an NER component, but they are trained on news/media
data rather than clinical notes. More importantly, all three official packages (it_core_news_sm,
md, and lg) are licensed CC BY-NC-SA 3.0, so they are a no-go for a product which can be
used commercially; using the MIT-licensed spaCy runtime does not change the weights' license
(spaCy Italian pipelines,
sm metadata,
md metadata,
lg metadata). A general Italian model also recognizes
names and places, not all sensitive clinical meaning. Training a suitable model would require a
labeled, representative and legally usable corpus; Presidio's own model guidance likewise states
that a labeled PII dataset is required to train a new model
(Presidio spaCy/Stanza guidance).
Consequently, the absence of a local LLM is irrelevant to sensitivity assessment. Installations with a local LLM may let description generation receive broader source content under a separate model-boundary policy, but the sensitivity classifier does not call that LLM. The optional NER pack is a distinct local classifier capability and is never inferred from the description-model catalog.
Licensing and redistribution gate
Framework code, model weights, and training data are three separate licensing layers. A permissive
runtime does not grant ThothII permission to bundle weights trained from restricted data. For an
open-source product that may be deployed commercially, the default rule should therefore be:
reject NC, research-only, custom restrictive, missing, or ambiguous model licenses. Pinning an
approved artifact must also pin its model card, upstream model, training datasets, required notices,
and checksums. This is a technical compatibility screen, not legal advice.
| Candidate | Framework code | Weights and training-data evidence | Commercial bundling in ThothII | Decision |
|---|---|---|---|---|
| Presidio, deterministic recognizers only | MIT; its license permits use and redistribution with the notice (Presidio license) | No general Italian model is required with NoOpNlpEngine; third-party dependencies still retain their notices |
Compatible, but adds a Python process without adding unique detection capability for the agreed baseline | GO legally; NO-GO architecturally for v1 |
| spaCy runtime | MIT (spaCy license) | Runtime only; no permission is conferred on downloaded pipelines | Compatible if used with separately approved weights | GO for code only |
Official spaCy Italian sm/md/lg |
spaCy code is MIT | Every current package declares CC BY-NC-SA 3.0; the metadata traces part of training to UD Italian ISDT under the same non-commercial license (sm, md, lg) |
NC excludes the intended commercial deployments; SA adds redistribution obligations |
NO-GO |
| Stanza runtime | Apache-2.0 (Stanza license) | Stanford says its language packs are ODC-By only to the extent it owns the rights and explicitly tells users to inspect training-data licenses. Italian NER uses FBK's KIND data; KIND annotations are CC BY-NC 4.0 or CC BY-NC-SA 4.0 (Stanza model licensing note, Italian NER source, KIND license) | The Italian model's non-commercial source data fails the product requirement | GO for code; NO-GO for Italian weights |
| Flair runtime and official multilingual NER | MIT (Flair repository and license) | The official ner-multi model card neither includes Italian among its four trained languages nor declares a model license; a maintainer issue asking whether bundled weights are MIT was closed without an answer (model card, license issue) |
Runtime is usable, but these weights are unsuitable and their redistribution terms are not established | GO for code; NO-GO for these weights |
osiria/flare-it-ner |
Runs in Transformers/PyTorch; the card labels the weights MIT | Trained on CC BY 4.0 WikiNER plus an additional manually annotated custom dataset whose redistribution terms are not identified in the model card (model card, WikiNER dataset) | The weights look promising, but an MIT label alone does not resolve the undocumented custom training-data rights | NO-GO until provenance is explicit |
| Transformers.js / ONNX runtime | Apache-2.0 (Transformers.js license) | ONNX conversion does not replace the source model's license or cure training-data restrictions; each exact model needs its own audit | Compatible runtime, but there is no blanket approval for arbitrary Hub/ONNX weights | GO for code; weights case by case |
Laibniz/italian-ner-pii-browser-uncased |
Packaged for Transformers.js as quantized ONNX; model card says Apache-2.0 | It is a conversion of an Osiria model and therefore inherits the same undocumented custom-data provenance (model card) | CPU/browser packaging does not cure the upstream licensing gap | NO-GO |
Ar86Bat/multilang-pii-ner |
XLM-R/Transformers; model card labels weights MIT | Trained on ai4privacy/open-pii-masking-500k-ai4privacy, whose card declares CC BY 4.0 (model card, dataset card) |
Commercial redistribution appears compatible with MIT notice plus CC BY attribution; an ONNX export, dependency audit, and PSD evaluation are still required | CONDITIONAL GO as optional PII NER |
urchade/gliner_multi_pii-v1 |
GLiNER code and model are Apache-2.0 | Its published multilingual synthetic dataset, including Italian, is Apache-2.0 (model card, dataset card) | The cleanest publicly inspectable license chain, but the published SPY comparison reports substantially lower recall than the selected Fastino model | GO legally; reserve candidate |
fastino/gliner2-privacy-filter-PII-multi |
GLiNER2 code is Apache-2.0; the direct local dependencies are permissively licensed | Weights are Apache-2.0 and the mDeBERTa base is MIT. The authors report a purpose-built synthetic corpus spanning seven languages, including Italian; the corpus itself is not published, so training is not fully reproducible (model card, paper, base model) | Free local/commercial use and redistribution are allowed subject to Apache notices. It has the best reported exact-span F1 and recall among the compared open detectors on SPY, although that evaluation is not Italian-specific | SELECTED optional NER pack |
openai/privacy-filter |
Code and weights are Apache-2.0 and it runs locally through Python or Transformers.js | The publisher explicitly permits commercial deployment, but documents the model as primarily English and warns that non-English performance may drop (model card, source) | Clean license and direct TypeScript path, but insufficient Italian evidence and a larger one-billion-parameter artifact make it a weaker fit | GO legally; NO-GO as primary Italian detector |
| OpenMed Italian PII models | Cards label the weights Apache-2.0 | They are trained on ai4privacy/pii-masking-400k, whose license limits use to academic/non-commercial purposes unless a commercial agreement is obtained (representative model, dataset license) |
An Apache label on derivative weights does not remove the explicit upstream non-commercial restriction | NO-GO |
The search therefore found a usable, free-of-charge, permissively licensed solution. Fastino GLiNER2-PII is selected because it combines an Apache code/weight grant, explicit Italian training, CPU operation, configurable entity labels, and materially higher published recall than the older Urchade model. This does not prove Italian clinical accuracy: the published SPY evaluation is English and the training corpus is synthetic. Those are quality and reproducibility limitations, not a reason to reintroduce a generative LLM into the classifier.
An implementation smoke test found a packaging incompatibility in the selected upstream versions:
the checkpoint was written by Transformers 5.8, while gliner2[local]==2.0.0 requires
transformers>=4.38,<5. Every published checkpoint revision has the same tokenizer metadata, and
upstream issue 145 reports the identical failure.
ThothII therefore pins Transformers 4.57.6 and performs the minimal documented-key conversion in a
temporary local view, without changing the downloaded model or its checksums. This compatibility
shim was accepted only after an offline CPU smoke test detected Italian names, dates, locations,
and usernames; any unexpected metadata shape fails closed. A future upstream fix must replace,
not silently stack on, this shim.
Fit with current ThothII design
The current backend already owns source sampling and the human-owned Sensitive Data Flag; source
values are transient and existing description sampling is bounded (PROJECT_STATE.md and
backend/src/catalog/description-source-sampler.ts). The new
classifier should therefore be a backend module behind the single effective sensitivity-decision
point already being designed, while preserving these agreed semantics:
sensitive,non_sensitive, andunknownare assessment outcomes;unknownleaves the human-owned Sensitive Data Flag unchanged. Its effect on later description generation is out of scope here.- Any positive rule match stops the column scan and yields
sensitive. - A negative time-bounded sample yields
unknown, nevernon_sensitive. - The administrator may set the final binary flag either way; the automated result remains a proposal with rule/version/coverage evidence and no stored source value.
- Learned NER, if present, can add positive evidence but cannot convert
unknowntonon_sensitive.
The five-second scan budget should be enforced at the database-read orchestration layer and not inside a recognizer library. An in-process deterministic engine maximizes the portion of that budget available to database reads and avoids model cold-start/IPC costs. It should stream bounded batches, check cancellation/deadline between batches, and stop at the first match.
Acceptance gate for enabling the selected NER pack
Build a versioned, value-free fixture corpus representing at least: Italian/general identifiers, false-positive lookalikes, bounded short notes with people/places, credential formats, null/high- cardinality columns, and PSD-specific clinical codes. Measure per category rather than one aggregate score. Pin the GLiNER2 package and model revision/checksum, build an SBOM including transitive dependencies, retain required notices, prohibit model downloads at inference time, and measure warm latency and memory on the target CPU. The detector is enabled only where those checks pass.
The acceptance test may tune label descriptions and per-label confidence thresholds, but it does
not reopen the architecture or silently select another Hub model. If quality or performance is
unacceptable, the optional NER pack remains disabled and unresolved sampled columns remain
unknown. Replacing the chosen model or adding a local-LLM fallback requires a new documented
decision.