feat: refine metadata catalog workflows
This commit is contained in:
@@ -0,0 +1,218 @@
|
||||
# Local sensitive-column classifier: TypeScript or Python
|
||||
|
||||
**Date:** 2026-09-02
|
||||
**Scope:** library/runtime choice for a local classifier without a generative LLM. Synthetic data generation and
|
||||
model-boundary policy for description generation are deliberately out of scope.
|
||||
|
||||
## Recommendation
|
||||
|
||||
Implement the mandatory classifier **inside the TypeScript backend**, with no learned model in its
|
||||
mandatory path. The agreed policy is dominated by deterministic work:
|
||||
type/declared-size rules, bounded string normalization, patterns, checksums, exact dictionaries,
|
||||
context terms, and "one match makes the column sensitive". TypeScript is fully adequate for that
|
||||
work and keeps the classifier in the process which already owns the catalog, source connectors,
|
||||
run lifecycle, cancellation, and the human override.
|
||||
|
||||
Use small, focused dependencies rather than a general NLP framework:
|
||||
|
||||
- `validator` for syntax/checksum-oriented recognizers such as email, IP, MAC, IBAN, BIC, payment
|
||||
cards, and locale-aware Italian tax IDs, passports, identity cards and VAT numbers; its official
|
||||
API exposes these validators individually, so unused validators need not be imported
|
||||
([validator.js source and API](https://github.com/validatorjs/validator.js/)).
|
||||
- `libphonenumber-js/max` for international telephone parsing and digit-pattern validation. The
|
||||
project's `max` metadata is intentionally more precise than its default/minimal metadata
|
||||
([libphonenumber-js validation documentation](https://github.com/catamphetamine/libphonenumber-js/blob/master/README.md#isvalid-boolean)).
|
||||
- ThothII-owned recognizers for identifiers not covered adequately by a selected validator (for
|
||||
example the Italian driving licence) and organization-specific identifiers, each with a version,
|
||||
tests, context words, and checksum/semantic validation where the identifier defines one.
|
||||
- A safe regular-expression boundary for administrator-authored rules. `node-re2` offers a mostly
|
||||
`RegExp`-compatible, ReDoS-resistant engine and a set API for matching many patterns, but is a
|
||||
native addon that downloads or builds a binary during installation. That packaging cost must be
|
||||
tested against the project images before adoption
|
||||
([node-re2 README](https://github.com/uhop/node-re2/blob/master/README.md)). If arbitrary regular
|
||||
expressions are not exposed, ThothII can instead keep a reviewed built-in pattern set and bound
|
||||
every inspected value.
|
||||
|
||||
Do **not** add NLP.js, spaCy, Presidio, or a transformer merely to implement regexes and dictionaries.
|
||||
They do not make deterministic identifiers more semantically understandable. In particular, NLP.js
|
||||
documents useful Italian tokenization plus enum/regex/built-in entities, but its advertised built-ins
|
||||
are mostly the same formatted values already covered above; it does not advertise a ready-made
|
||||
Italian PII model for arbitrary people or clinical concepts
|
||||
([NLP.js official README](https://github.com/axa-group/nlp.js/)).
|
||||
|
||||
Add an **optional, non-generative statistical NER pack** based on
|
||||
[`fastino/gliner2-privacy-filter-PII-multi`](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi).
|
||||
The selected model and its GLiNER2 runtime are Apache-2.0, its mDeBERTa base is MIT, and the model
|
||||
is trained across seven languages including Italian. Run the official implementation in one warmed,
|
||||
CPU-only Python worker; do not spawn it once per column or table. It returns entity evidence only.
|
||||
The TypeScript classifier remains the sole owner of the final column assessment and invokes NER
|
||||
only for bounded text that deterministic rules did not resolve and only while the scan deadline has
|
||||
time remaining.
|
||||
|
||||
This is a positive Q29 decision, not a placeholder for later library selection. The production gate
|
||||
is limited to pinning artifacts, producing an SBOM/license inventory, and measuring the selected
|
||||
model on PSD fixtures and target CPUs. If the optional pack is absent or misses its deadline, the
|
||||
normal coverage rule produces `unknown`; ThothII does not substitute a generative LLM.
|
||||
|
||||
## Why the TypeScript baseline is sufficient
|
||||
|
||||
The classifier is not being asked to infer an open-ended privacy judgment from prose. It executes a
|
||||
versioned policy and produces evidence. Its mandatory recognizers can be expressed as:
|
||||
|
||||
| Rule family | Implementation | Model needed? |
|
||||
| --- | --- | --- |
|
||||
| declared `text`/CLOB/unbounded string or declared maximum `>500` | catalog metadata rule | no |
|
||||
| an observed value longer than 500 characters | length rule, ideally detected in SQL before transferring the value | no |
|
||||
| email, IP/MAC, URL, UUID, payment card, IBAN/BIC, phone | pattern plus format/checksum validator | no |
|
||||
| Italian fiscal code/VAT/passport/identity card/driving licence | locale-aware validator where available, otherwise a reviewed country recognizer; pattern, context and checksum where defined | no |
|
||||
| credentials and technical secrets | anchored prefixes, token shapes, entropy/character-class heuristics, and organization allow/deny rules | no |
|
||||
| hospital- or customer-specific terms/codes | normalized exact dictionary or phrase matching | no |
|
||||
| person/location/organization in otherwise unstructured short text | statistical NER is useful but probabilistic | yes, optional |
|
||||
| clinical meaning in unstructured short Italian text | domain NER or a governed terminology; generic NER is not enough | optional and domain-specific |
|
||||
|
||||
This is not speculative parity with a Python implementation. Presidio's own supported-entity table
|
||||
shows that its global email, IBAN, credit-card and similar recognizers use pattern matching,
|
||||
validation, context and checksums, while its Italian fiscal-code/VAT/passport/identity-card/licence
|
||||
coverage is likewise rule-based
|
||||
([Presidio supported entities](https://presidio.dataprivacystack.org/supported_entities/)). Those
|
||||
mechanisms are portable; Presidio provides tested implementations and orchestration, not a unique
|
||||
Python-language capability.
|
||||
|
||||
For large dictionaries, exact normalized token/phrase matching is still deterministic NLP and does
|
||||
not require a learned model. The policy format should describe the rule rather than the library,
|
||||
for example `kind: email`, `kind: checksum_id`, `kind: regex`, `kind: phrase_set`, and
|
||||
`kind: max_length`. This makes a future engine change possible without changing persisted policy or
|
||||
assessment evidence.
|
||||
|
||||
## What Python/Presidio would materially add
|
||||
|
||||
Presidio is the strongest Python option if ThothII later needs a broader recognizer ecosystem. Its
|
||||
Analyzer composes rule-based recognizers, regexes, validation, NER and contextual enhancement; it
|
||||
also exposes custom recognizers, batch processing and decision tracing
|
||||
([Presidio Analyzer architecture](https://presidio.dataprivacystack.org/analyzer/)). Its current
|
||||
catalog includes Italian identifiers and its YAML registry supports custom patterns and language-
|
||||
specific context. Recent source also contains a `NoOpNlpEngine`, so a deterministic-only Presidio
|
||||
deployment can avoid loading spaCy; this removes model cost but not the Python runtime/process
|
||||
boundary
|
||||
([Presidio `NoOpNlpEngine` integration](https://github.com/data-privacy-stack/presidio/blob/main/presidio-analyzer/presidio_analyzer/recognizer_registry/recognizer_registry.py)).
|
||||
|
||||
`presidio-structured` is conceptually close because it maps tabular columns/JSON keys to detected
|
||||
entities, but its documented strategies select the most common, highest-confidence, or a mixed
|
||||
entity. ThothII would still need its own deadline/sampling/`unknown` semantics and human-override
|
||||
lifecycle; Presidio also lists improved mixed free-text/structured support as future work
|
||||
([Presidio Structured](https://presidio.dataprivacystack.org/structured/)).
|
||||
|
||||
That benefit does not currently justify using it as ThothII's mandatory engine:
|
||||
|
||||
- The backend currently has no PII/NLP dependencies and is an ESM TypeScript service
|
||||
(`backend/package.json`); the existing Python environment belongs to the separate harness/core
|
||||
layer and also has no Presidio or spaCy dependency (`harness/pyproject.toml`). A Python analyzer therefore means a new
|
||||
ownership and deployment boundary, not simply another import.
|
||||
- Presidio is Python-based and can be run as a package or REST container. Its REST API intentionally
|
||||
has no built-in authentication, so a sidecar would have to remain on a protected internal network
|
||||
and be guarded by infrastructure
|
||||
([Presidio FAQ, deployment](https://presidio.dataprivacystack.org/faq/#how-can-i-deploy-presidio-into-my-environment)).
|
||||
- Presidio defaults to English. Adding Italian requires both an Italian NLP pipeline and adapted
|
||||
language-specific recognizers/context; language-agnostic regexes alone do not solve that setup
|
||||
([Presidio multilingual guidance](https://github.com/data-privacy-stack/presidio/blob/main/docs/analyzer/languages.md)).
|
||||
- Presidio itself warns that automated detection cannot guarantee finding all sensitive data and
|
||||
documents an unavoidable false-positive/false-negative trade-off
|
||||
([Presidio FAQ](https://presidio.dataprivacystack.org/faq/#what-is-presidio)). It cannot turn a
|
||||
negative sample into proof that a column is safe.
|
||||
|
||||
## Learned NLP is not an LLM, but it is still a model
|
||||
|
||||
A spaCy NER pipeline does not generate text and is not an LLM. It is a statistical token classifier:
|
||||
it predicts spans such as `PERSON`, `LOCATION` and `ORGANIZATION`. spaCy explicitly notes that the
|
||||
predictions depend on the training examples and will not always be correct
|
||||
([spaCy NER documentation](https://spacy.io/usage/linguistic-features#named-entities)). That makes it
|
||||
a useful additional **positive detector**, not the authority which declares a sampled column safe.
|
||||
|
||||
spaCy publishes Italian CPU pipelines with an NER component, but they are trained on news/media
|
||||
data rather than clinical notes. More importantly, all three official packages (`it_core_news_sm`,
|
||||
`md`, and `lg`) are licensed **CC BY-NC-SA 3.0**, so they are a no-go for a product which can be
|
||||
used commercially; using the MIT-licensed spaCy runtime does not change the weights' license
|
||||
([spaCy Italian pipelines](https://spacy.io/models/it),
|
||||
[`sm` metadata](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_sm-3.8.0.json),
|
||||
[`md` metadata](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_md-3.8.0.json),
|
||||
[`lg` metadata](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_lg-3.8.0.json)). A general Italian model also recognizes
|
||||
names and places, not all sensitive clinical meaning. Training a suitable model would require a
|
||||
labeled, representative and legally usable corpus; Presidio's own model guidance likewise states
|
||||
that a labeled PII dataset is required to train a new model
|
||||
([Presidio spaCy/Stanza guidance](https://github.com/data-privacy-stack/presidio/blob/main/docs/analyzer/nlp_engines/spacy_stanza.md#training-your-own-model)).
|
||||
|
||||
Consequently, the absence of a local LLM is irrelevant to sensitivity assessment. Installations
|
||||
with a local LLM may let description generation receive broader source content under a separate
|
||||
model-boundary policy, but **the sensitivity classifier does not call that LLM**. The optional NER
|
||||
pack is a distinct local classifier capability and is never inferred from the description-model
|
||||
catalog.
|
||||
|
||||
## Licensing and redistribution gate
|
||||
|
||||
Framework code, model weights, and training data are three separate licensing layers. A permissive
|
||||
runtime does not grant ThothII permission to bundle weights trained from restricted data. For an
|
||||
open-source product that may be deployed commercially, the default rule should therefore be:
|
||||
**reject `NC`, research-only, custom restrictive, missing, or ambiguous model licenses**. Pinning an
|
||||
approved artifact must also pin its model card, upstream model, training datasets, required notices,
|
||||
and checksums. This is a technical compatibility screen, not legal advice.
|
||||
|
||||
| Candidate | Framework code | Weights and training-data evidence | Commercial bundling in ThothII | Decision |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Presidio, deterministic recognizers only | MIT; its license permits use and redistribution with the notice ([Presidio license](https://github.com/data-privacy-stack/presidio/blob/main/LICENSE)) | No general Italian model is required with `NoOpNlpEngine`; third-party dependencies still retain their notices | Compatible, but adds a Python process without adding unique detection capability for the agreed baseline | **GO legally; NO-GO architecturally for v1** |
|
||||
| spaCy runtime | MIT ([spaCy license](https://github.com/explosion/spaCy/blob/master/LICENSE)) | Runtime only; no permission is conferred on downloaded pipelines | Compatible if used with separately approved weights | **GO for code only** |
|
||||
| Official spaCy Italian `sm`/`md`/`lg` | spaCy code is MIT | Every current package declares CC BY-NC-SA 3.0; the metadata traces part of training to UD Italian ISDT under the same non-commercial license ([`sm`](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_sm-3.8.0.json), [`md`](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_md-3.8.0.json), [`lg`](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_lg-3.8.0.json)) | `NC` excludes the intended commercial deployments; `SA` adds redistribution obligations | **NO-GO** |
|
||||
| Stanza runtime | Apache-2.0 ([Stanza license](https://github.com/stanfordnlp/stanza/blob/main/LICENSE)) | Stanford says its language packs are ODC-By only to the extent it owns the rights and explicitly tells users to inspect training-data licenses. Italian NER uses FBK's KIND data; KIND annotations are CC BY-NC 4.0 or CC BY-NC-SA 4.0 ([Stanza model licensing note](https://stanfordnlp.github.io/stanza/performance.html), [Italian NER source](https://stanfordnlp.github.io/stanza/ner_models.html), [KIND license](https://github.com/dhfbk/KIND#license)) | The Italian model's non-commercial source data fails the product requirement | **GO for code; NO-GO for Italian weights** |
|
||||
| Flair runtime and official multilingual NER | MIT ([Flair repository and license](https://github.com/flairNLP/flair)) | The official `ner-multi` model card neither includes Italian among its four trained languages nor declares a model license; a maintainer issue asking whether bundled weights are MIT was closed without an answer ([model card](https://huggingface.co/flair/ner-multi), [license issue](https://github.com/flairNLP/flair/issues/1487)) | Runtime is usable, but these weights are unsuitable and their redistribution terms are not established | **GO for code; NO-GO for these weights** |
|
||||
| `osiria/flare-it-ner` | Runs in Transformers/PyTorch; the card labels the weights MIT | Trained on CC BY 4.0 WikiNER plus an additional manually annotated custom dataset whose redistribution terms are not identified in the model card ([model card](https://huggingface.co/osiria/flare-it-ner), [WikiNER dataset](https://figshare.com/articles/dataset/Learning_multilingual_named_entity_recognition_from_Wikipedia/5462500)) | The weights look promising, but an MIT label alone does not resolve the undocumented custom training-data rights | **NO-GO until provenance is explicit** |
|
||||
| Transformers.js / ONNX runtime | Apache-2.0 ([Transformers.js license](https://github.com/huggingface/transformers.js/blob/main/LICENSE)) | ONNX conversion does not replace the source model's license or cure training-data restrictions; each exact model needs its own audit | Compatible runtime, but there is no blanket approval for arbitrary Hub/ONNX weights | **GO for code; weights case by case** |
|
||||
| `Laibniz/italian-ner-pii-browser-uncased` | Packaged for Transformers.js as quantized ONNX; model card says Apache-2.0 | It is a conversion of an Osiria model and therefore inherits the same undocumented custom-data provenance ([model card](https://huggingface.co/Laibniz/italian-ner-pii-browser-uncased)) | CPU/browser packaging does not cure the upstream licensing gap | **NO-GO** |
|
||||
| `Ar86Bat/multilang-pii-ner` | XLM-R/Transformers; model card labels weights MIT | Trained on `ai4privacy/open-pii-masking-500k-ai4privacy`, whose card declares CC BY 4.0 ([model card](https://huggingface.co/Ar86Bat/multilang-pii-ner), [dataset card](https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy)) | Commercial redistribution appears compatible with MIT notice plus CC BY attribution; an ONNX export, dependency audit, and PSD evaluation are still required | **CONDITIONAL GO as optional PII NER** |
|
||||
| `urchade/gliner_multi_pii-v1` | GLiNER code and model are Apache-2.0 | Its published multilingual synthetic dataset, including Italian, is Apache-2.0 ([model card](https://huggingface.co/urchade/gliner_multi_pii-v1), [dataset card](https://huggingface.co/datasets/urchade/synthetic-pii-ner-mistral-v1)) | The cleanest publicly inspectable license chain, but the published SPY comparison reports substantially lower recall than the selected Fastino model | **GO legally; reserve candidate** |
|
||||
| `fastino/gliner2-privacy-filter-PII-multi` | GLiNER2 code is Apache-2.0; the direct local dependencies are permissively licensed | Weights are Apache-2.0 and the mDeBERTa base is MIT. The authors report a purpose-built synthetic corpus spanning seven languages, including Italian; the corpus itself is not published, so training is not fully reproducible ([model card](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi), [paper](https://arxiv.org/abs/2605.09973), [base model](https://huggingface.co/microsoft/mdeberta-v3-base)) | Free local/commercial use and redistribution are allowed subject to Apache notices. It has the best reported exact-span F1 and recall among the compared open detectors on SPY, although that evaluation is not Italian-specific | **SELECTED optional NER pack** |
|
||||
| `openai/privacy-filter` | Code and weights are Apache-2.0 and it runs locally through Python or Transformers.js | The publisher explicitly permits commercial deployment, but documents the model as primarily English and warns that non-English performance may drop ([model card](https://huggingface.co/openai/privacy-filter), [source](https://github.com/openai/privacy-filter)) | Clean license and direct TypeScript path, but insufficient Italian evidence and a larger one-billion-parameter artifact make it a weaker fit | **GO legally; NO-GO as primary Italian detector** |
|
||||
| OpenMed Italian PII models | Cards label the weights Apache-2.0 | They are trained on `ai4privacy/pii-masking-400k`, whose license limits use to academic/non-commercial purposes unless a commercial agreement is obtained ([representative model](https://huggingface.co/OpenMed/OpenMed-PII-Italian-BioClinicalBERT-Base-110M-v1), [dataset license](https://huggingface.co/datasets/ai4privacy/pii-masking-400k/blob/main/LICENSE)) | An Apache label on derivative weights does not remove the explicit upstream non-commercial restriction | **NO-GO** |
|
||||
|
||||
The search therefore found a usable, free-of-charge, permissively licensed solution. Fastino
|
||||
GLiNER2-PII is selected because it combines an Apache code/weight grant, explicit Italian training,
|
||||
CPU operation, configurable entity labels, and materially higher published recall than the older
|
||||
Urchade model. This does not prove Italian clinical accuracy: the published SPY evaluation is
|
||||
English and the training corpus is synthetic. Those are quality and reproducibility limitations,
|
||||
not a reason to reintroduce a generative LLM into the classifier.
|
||||
|
||||
## Fit with current ThothII design
|
||||
|
||||
The current backend already owns source sampling and the human-owned `Sensitive Data Flag`; source
|
||||
values are transient and existing description sampling is bounded (`PROJECT_STATE.md` and
|
||||
`backend/src/catalog/description-source-sampler.ts`). The new
|
||||
classifier should therefore be a backend module behind the single effective sensitivity-decision
|
||||
point already being designed, while preserving these agreed semantics:
|
||||
|
||||
1. `sensitive`, `non_sensitive`, and `unknown` are assessment outcomes; `unknown` leaves the
|
||||
human-owned Sensitive Data Flag unchanged. Its effect on later description generation is out of
|
||||
scope here.
|
||||
2. Any positive rule match stops the column scan and yields `sensitive`.
|
||||
3. A negative time-bounded sample yields `unknown`, never `non_sensitive`.
|
||||
4. The administrator may set the final binary flag either way; the automated result remains a
|
||||
proposal with rule/version/coverage evidence and no stored source value.
|
||||
5. Learned NER, if present, can add positive evidence but cannot convert `unknown` to
|
||||
`non_sensitive`.
|
||||
|
||||
The five-second scan budget should be enforced at the database-read orchestration layer and not
|
||||
inside a recognizer library. An in-process deterministic engine maximizes the portion of that budget
|
||||
available to database reads and avoids model cold-start/IPC costs. It should stream bounded batches,
|
||||
check cancellation/deadline between batches, and stop at the first match.
|
||||
|
||||
## Acceptance gate for enabling the selected NER pack
|
||||
|
||||
Build a versioned, value-free fixture corpus representing at least: Italian/general identifiers,
|
||||
false-positive lookalikes, bounded short notes with people/places, credential formats, null/high-
|
||||
cardinality columns, and PSD-specific clinical codes. Measure per category rather than one aggregate
|
||||
score. Pin the GLiNER2 package and model revision/checksum, build an SBOM including transitive
|
||||
dependencies, retain required notices, prohibit model downloads at inference time, and measure warm
|
||||
latency and memory on the target CPU. The detector is enabled only where those checks pass.
|
||||
|
||||
The acceptance test may tune label descriptions and per-label confidence thresholds, but it does
|
||||
not reopen the architecture or silently select another Hub model. If quality or performance is
|
||||
unacceptable, the optional NER pack remains disabled and unresolved sampled columns remain
|
||||
`unknown`. Replacing the chosen model or adding a local-LLM fallback requires a new documented
|
||||
decision.
|
||||
Reference in New Issue
Block a user