feat: refine metadata catalog workflows
This commit is contained in:
@@ -0,0 +1,67 @@
|
||||
# Use one Installation Model Catalog with runtime projections
|
||||
|
||||
ThothII currently declares model availability independently in installation
|
||||
`metadataGeneration`, Pi configuration, and workspace `llm_policy` and embedding settings. The
|
||||
Installation Model Catalog in `thothii-installation.yaml` becomes the sole authored authority for
|
||||
session, metadata-generation, and embedding models, including their allowed usages and per-usage
|
||||
defaults. Workspace descriptors retain database identity and scope plus Evidence concerns, but no
|
||||
model policy or selection; installation-local Database Bindings remain separate from them.
|
||||
|
||||
Pi, the backend metadata-generation helper, and the embedding runtime consume generated Model
|
||||
Runtime Projections of that catalog. Runtime Model Selection stores only a canonical catalog model
|
||||
identity and use-specific controls such as thinking level; endpoint, provider, capabilities, and
|
||||
credential references remain catalog facts, while secret values remain in protected secret stores.
|
||||
The metadata-generation helper continues to use LiteLLM independently of Pi, as established by
|
||||
ADR-0009: unifying model declaration does not unify execution lifecycles.
|
||||
|
||||
The migration is intentionally fail-closed. Workspace schema v4 removes `llm_policy` and the
|
||||
entire redundant `semantic_index`; collection identity is derived from the workspace identity,
|
||||
while vector-store and embedding facts come from the installation. A deterministic migration
|
||||
rewrites existing descriptors. The
|
||||
installation loader replaces `metadataGeneration` with `modelCatalog` and rejects the legacy form
|
||||
with an actionable migration error rather than keeping two live sources. Existing Pi
|
||||
`models.json` and enabled-model settings become generated artifacts and are never edited as
|
||||
authoritative configuration.
|
||||
|
||||
Catalog identities use the canonical `provider/model` form; runtime-specific upstream names are
|
||||
adapter facts, not additional ThothII identities. Installation descriptors use schema version 2
|
||||
and model-free workspace descriptors use schema version 4. Removing a model never substitutes it
|
||||
inside an existing session: an unresolvable resume fails explicitly. Published semantic indexes
|
||||
record the embedding identity and dimensions that produced them and require explicit
|
||||
reprocessing when those facts change.
|
||||
|
||||
Model eligibility is expressed by the presence of a `session` or `metadataGeneration` block,
|
||||
without a duplicate usages list. The installation declares one active embedding identity and its
|
||||
dimensions rather than a selectable embedding catalog. Runtime projections are regenerated
|
||||
deterministically and atomically at start, so they require no persisted digest and are excluded
|
||||
from installation backups; restore regenerates them from the validated installation descriptor.
|
||||
|
||||
A provider owns one endpoint, one explicit authentication mode, and only the runtime adapters it
|
||||
needs. Authentication is either a protected secret-environment reference, Pi-owned authentication
|
||||
for session-only built-in models, or explicit keyless operation for an explicit endpoint; models
|
||||
cannot override it. Pi built-in model facts are not copied into the installation. A session default
|
||||
and the single embedding definition are required, while metadata generation and its default may be
|
||||
omitted together. The schema deliberately excludes unused abstractions and future properties until
|
||||
runtime behavior requires them.
|
||||
|
||||
The catalog session default replaces `PI_PROVIDER`, `PI_MODEL`, and persisted installation model
|
||||
defaults as configuration sources. A user or session selection is only a canonical catalog
|
||||
reference, and an existing session keeps that reference without silently switching models. A
|
||||
generated Compose projection supplies the catalog-derived embedding values and projection mounts
|
||||
to every affected service, so neither the base Compose files nor `operator.env` repeat model facts.
|
||||
|
||||
Workspace v3-to-v4 migration is a deterministic removal of `llm_policy` and `semantic_index`.
|
||||
Installation migration instead inspects the legacy installation metadata block and both Pi model
|
||||
files: it emits a v2 candidate only when their identities and settings can be reconciled without
|
||||
guessing. Conflicts produce an actionable report and leave every source untouched.
|
||||
|
||||
## Considered Options
|
||||
|
||||
- A separate catalog file referenced by the installation was rejected because it adds path,
|
||||
permission, backup, and atomic-update coordination without a current need for cross-installation
|
||||
sharing.
|
||||
- Pi `models.json` was rejected as the authority because it is a Pi-specific projection that does
|
||||
not express all ThothII usages, built-in providers, metadata-generation controls, or embedding
|
||||
facts.
|
||||
- Transitional dual reading was rejected because it would preserve the configuration discrepancy
|
||||
this decision is intended to eliminate.
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
status: accepted
|
||||
---
|
||||
|
||||
# Assess sensitive columns locally from source content
|
||||
|
||||
The Sensitive Data Flag remains a human-owned boolean. An explicit, selection-scoped sensitivity
|
||||
analysis may propose changes by inspecting both catalog metadata and source values, but it never
|
||||
writes the flag. The administrator may accept, reject, or reverse every proposal.
|
||||
|
||||
One TypeScript `SensitivityClassifier` is the only component allowed to produce the column-level
|
||||
assessment `sensitive`, `non_sensitive`, or `unknown`. It applies a versioned Sensitive Data Policy
|
||||
and consumes values through database-independent streaming adapters. Database-specific code may
|
||||
read and normalize bounded values, but it may not decide sensitivity.
|
||||
|
||||
The classifier first applies deterministic metadata rules, value validators, checksums,
|
||||
dictionaries, and length rules. A single validated sensitive match makes the whole column
|
||||
`sensitive`; any textual value longer than 500 characters is such a match. It attempts a complete
|
||||
scan, but after five seconds per table it continues by sampling within the remaining run budget. A
|
||||
completed scan with no finding may produce `non_sensitive`; an incomplete scan with no finding
|
||||
produces `unknown`.
|
||||
|
||||
Ambiguous text may additionally be sent to an optional local NER detector only while time remains.
|
||||
The detector runs on CPU, receives no tools or network access, does not persist source values, and
|
||||
returns evidence rather than the column decision. The initial supported detector is
|
||||
[`fastino/gliner2-privacy-filter-PII-multi`](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi),
|
||||
used through the Apache-2.0 GLiNER2 Python library with a pinned model revision. Its model weights
|
||||
and GLiNER2 code are Apache-2.0, and its mDeBERTa base model is MIT. It is trained for seven
|
||||
languages including Italian and can run on CPU without using the installation's GPUs.
|
||||
|
||||
The NER detector is an optional installation asset because its weights and runtime are materially
|
||||
larger than the deterministic TypeScript engine. It is invoked only for otherwise unresolved text,
|
||||
never for values already classified by a decisive rule. If it is disabled, unavailable, times out,
|
||||
or returns no qualifying evidence before the deadline, the classifier follows the same coverage
|
||||
rule and may return `unknown`.
|
||||
|
||||
No generative LLM, internal or external, participates in sensitivity assessment. A fallback to an
|
||||
installation-local LLM is unnecessary while a permissively licensed local NER implementation is
|
||||
available, and would reintroduce queue latency, non-deterministic judgments, prompt-injection
|
||||
surface, and contention with normal inference. Introducing such a fallback would require a new
|
||||
decision based on evidence that the NER path is unusable.
|
||||
|
||||
This decision supersedes ADR-0011 only where that ADR assigns draft sensitivity suggestions to an
|
||||
AI using structural metadata. ADR-0011's human authority, transient draft, explicit scope, and
|
||||
Sensitive Data Flag remain in force. How description generation uses the flag is outside this
|
||||
decision.
|
||||
|
||||
The selected model's published evaluation is not an Italian production acceptance test. Before
|
||||
enabling the NER profile by default, ThothII must pin the artifacts, generate a dependency/license
|
||||
inventory, and pass a CPU benchmark plus a labeled Italian corpus representative of the target
|
||||
databases. Failure of those gates disables NER; it does not silently select another model or an
|
||||
LLM.
|
||||
@@ -0,0 +1,245 @@
|
||||
# Installation Model Catalog
|
||||
|
||||
Status: accepted design; implementation not started.
|
||||
|
||||
## Outcome
|
||||
|
||||
`thothii-installation.yaml` is the only operator-authored source for models used by interactive
|
||||
sessions, metadata generation, and embedding. Runtime-specific files are deterministic projections,
|
||||
not additional configuration sources. Workspace descriptors contain database and Evidence concerns
|
||||
and no model, provider, allowlist, default, embedding, or vector-store configuration.
|
||||
|
||||
This design does not merge execution lifecycles. Pi continues to run interactive sessions, the
|
||||
short-lived LiteLLM helper continues to perform metadata generation, and the internal Ollama service
|
||||
continues to provide embeddings. They share model declaration, not execution machinery.
|
||||
|
||||
## Canonical installation shape
|
||||
|
||||
The following example covers all currently required cases: a Pi built-in model, an authenticated
|
||||
custom endpoint, a keyless internal endpoint, metadata generation, and the single embedding model.
|
||||
|
||||
```yaml
|
||||
schemaVersion: 2
|
||||
profile: server
|
||||
projectDirectory: /srv/thothii
|
||||
envFile: /srv/thothii/operator.env
|
||||
|
||||
workspaceRepository:
|
||||
remote: git@git.example.com:organization/workspaces.git
|
||||
branch: main
|
||||
access: ssh
|
||||
|
||||
modelCatalog:
|
||||
defaults:
|
||||
session: zai/glm-5.3
|
||||
metadataGeneration: local-qwen/qwen3.6-35b-a3b
|
||||
|
||||
embedding:
|
||||
id: ollama/qwen3-embedding:0.6b
|
||||
dimensions: 1024
|
||||
|
||||
providers:
|
||||
deepseek:
|
||||
authentication:
|
||||
mode: pi_auth
|
||||
session:
|
||||
mode: pi_builtin
|
||||
models:
|
||||
deepseek-v4-pro:
|
||||
session: {}
|
||||
deepseek-v4-flash:
|
||||
session: {}
|
||||
|
||||
zai:
|
||||
endpoint:
|
||||
baseUrl: https://api.z.ai/api/coding/paas/v4
|
||||
authentication:
|
||||
mode: secret_env
|
||||
apiKeyEnv: ZAI_API_KEY
|
||||
session:
|
||||
mode: openai_compatible
|
||||
metadataGeneration:
|
||||
litellmProvider: openai
|
||||
models:
|
||||
glm-5.3:
|
||||
label: GLM-5.3
|
||||
session:
|
||||
reasoning: true
|
||||
contextWindow: 200000
|
||||
maxTokens: 131072
|
||||
metadataGeneration: {}
|
||||
|
||||
local-qwen:
|
||||
endpoint:
|
||||
baseUrl: https://ml-aritmolab.policlinicosandonato.it/v1
|
||||
authentication:
|
||||
mode: none
|
||||
session:
|
||||
mode: openai_compatible
|
||||
metadataGeneration:
|
||||
litellmProvider: openai
|
||||
models:
|
||||
qwen3.6-35b-a3b:
|
||||
label: Qwen3.6 35B A3B
|
||||
session:
|
||||
reasoning: false
|
||||
contextWindow: 131072
|
||||
maxTokens: 16384
|
||||
compatibility:
|
||||
supportsDeveloperRole: false
|
||||
supportsReasoningEffort: false
|
||||
supportsStore: false
|
||||
maxTokensField: max_tokens
|
||||
metadataGeneration:
|
||||
disableThinking: true
|
||||
|
||||
authentication:
|
||||
configDirectory: /srv/thothii/auth-canonical
|
||||
runtimeProjection:
|
||||
directory: /srv/thothii/auth-runtime
|
||||
uid: 10001
|
||||
gid: 10001
|
||||
```
|
||||
|
||||
The catalog uses maps instead of repeated IDs. The canonical identity of a model is always derived
|
||||
as `<provider-key>/<model-key>`. `upstreamModel` may be added to a model only when the endpoint uses
|
||||
a different identifier. `label` is optional and falls back to the canonical identity.
|
||||
|
||||
Model eligibility is not repeated in an `usages` array. A `session` block makes the model eligible
|
||||
for sessions; a `metadataGeneration` block makes it eligible for metadata generation. The embedding
|
||||
is a single required installation value rather than a list plus default.
|
||||
|
||||
## Provider and authentication rules
|
||||
|
||||
A provider owns one endpoint, one authentication mode, and zero or one adapter for each runtime.
|
||||
Model entries cannot override provider endpoint or credentials. If the same upstream service needs
|
||||
different endpoints or credentials, the installation declares two provider identities.
|
||||
|
||||
Supported session modes are intentionally closed:
|
||||
|
||||
- `pi_builtin`: Pi already owns the model's technical descriptor; the model's `session` block is
|
||||
empty and ThothII does not copy context-window or compatibility facts.
|
||||
- `openai_compatible`: ThothII generates a Pi custom-provider descriptor; each session model supplies
|
||||
the technical values required by Pi.
|
||||
|
||||
Metadata generation uses the provider-level `litellmProvider`. A model-level
|
||||
`metadataGeneration.disableThinking: true` is permitted only for an explicit compatible endpoint.
|
||||
There is no generic adapter or plugin abstraction in schema version 2.
|
||||
|
||||
Exactly one provider authentication mode is allowed:
|
||||
|
||||
- `secret_env` requires an approved API-key environment reference present in the protected secret
|
||||
bundle. Secret values never enter YAML, generated files, logs, arguments, or API responses.
|
||||
- `pi_auth` is valid only for session-only `pi_builtin` providers and resolves through Pi's protected
|
||||
authentication projection.
|
||||
- `none` is valid only for an explicit endpoint. Runtime projections may supply a fixed non-secret
|
||||
compatibility placeholder when a client library requires a non-empty key.
|
||||
|
||||
## Defaults and selections
|
||||
|
||||
`defaults.session` and `embedding` are required. `defaults.metadataGeneration` is required exactly
|
||||
when at least one model has a `metadataGeneration` block; metadata generation may otherwise be
|
||||
absent and its UI controls are disabled.
|
||||
|
||||
`modelCatalog.defaults.session` is the only configured session-model default. `PI_PROVIDER`,
|
||||
`PI_MODEL`, and provider/model fields in installation-default settings are removed. A user choice is
|
||||
a Model Selection containing only the canonical model identity and runtime controls such as thinking
|
||||
level. A session manifest pins the selected canonical identity.
|
||||
|
||||
Removing the currently selected model causes new-session selection to fall back to the catalog
|
||||
default with an explicit administrative warning. An existing session is never silently moved to a
|
||||
different model; resume fails with `model_unavailable` when its pinned identity can no longer be
|
||||
resolved.
|
||||
|
||||
## Generated runtime projections
|
||||
|
||||
Before Compose starts, `tht` strictly validates schema version 2 and generates installation-local
|
||||
artifacts below `deploy/<installation-id>/generated/`:
|
||||
|
||||
- a normalized catalog JSON consumed defensively by the backend;
|
||||
- Pi `models.json` for custom providers;
|
||||
- Pi `settings.json`, combining fixed product settings with the session-eligible canonical IDs;
|
||||
- a Compose override that mounts the projections and supplies embedding identity and dimensions to
|
||||
core, preprocessing, and `embedding-model-init`.
|
||||
|
||||
Generation is deterministic and published only after every candidate artifact validates. A failed
|
||||
generation aborts start before Compose is invoked. `tht doctor` recomputes expected bytes and reports
|
||||
differences; no digest manifest or separate apply command exists. When projection bytes change,
|
||||
`tht start` recreates the affected services so they cannot continue with an older bind mount.
|
||||
|
||||
Generated projections are not backed up. Restore validates the canonical installation descriptor,
|
||||
regenerates every projection, and only then starts services. Base Compose files and `operator.env`
|
||||
must contain no model identities, defaults, endpoints, or dimensions.
|
||||
|
||||
## Workspace schema v4
|
||||
|
||||
Workspace schema v4 removes both top-level `llm_policy` and `semantic_index`. The entire latter
|
||||
block is redundant today: its engine and distance are product constants, its collection duplicates
|
||||
the workspace ID, and its model and dimensions are installation facts.
|
||||
|
||||
The runtime derives:
|
||||
|
||||
- Qdrant collection identity from the workspace ID;
|
||||
- engine and distance from the supported product contract;
|
||||
- embedding identity and dimensions from the Installation Model Catalog.
|
||||
|
||||
The published index generation records the canonical embedding identity and dimensions that created
|
||||
it. A mismatch makes the index explicitly incompatible and requires operator-triggered
|
||||
preprocessing. No existing index is deleted or rebuilt automatically.
|
||||
|
||||
The v3-to-v4 workspace migration is deterministic: set `workspace.schema_version` to `4`, remove
|
||||
`llm_policy`, and remove `semantic_index`. It does not alter database, Evidence, diagnostics, or
|
||||
binding data.
|
||||
|
||||
## Installation migration
|
||||
|
||||
Legacy installation migration must inspect all three former sources:
|
||||
|
||||
1. `metadataGeneration` in `thothii-installation.yaml`;
|
||||
2. `deploy/pi/models.json`;
|
||||
3. `deploy/pi/settings.json`.
|
||||
|
||||
The migrator emits a version-2 candidate only when it can reconcile identities, endpoints,
|
||||
credentials, and runtime-specific facts without guessing. Ambiguous aliases such as `glm-53`,
|
||||
`zai/glm-5.3`, and `openai/glm-5.3` are not silently equated. A conflict produces a field-level
|
||||
report and leaves every input unchanged for operator resolution.
|
||||
|
||||
After migration, the strict loader rejects `metadataGeneration`, workspace `llm_policy`, workspace
|
||||
`semantic_index`, legacy Pi source files, unknown fields, duplicate YAML keys, invalid defaults, and
|
||||
incompatible authentication/adapter combinations with an actionable `migration_required` or
|
||||
validation error.
|
||||
|
||||
## Final simplicity audit
|
||||
|
||||
The accepted design removes every configuration duplication that can be removed without inference:
|
||||
|
||||
- one authored installation file instead of an installation block plus two Pi files;
|
||||
- one canonical `provider/model` identity instead of display IDs and runtime IDs;
|
||||
- per-use blocks instead of a duplicated usages list;
|
||||
- one embedding entry instead of a selectable embedding catalog;
|
||||
- one catalog session default instead of environment and settings defaults;
|
||||
- no model or vector-store fields in workspace descriptors;
|
||||
- provider-level credentials instead of per-model credentials;
|
||||
- no generic runtime-plugin abstraction;
|
||||
- no persisted digest, apply command, or backup of generated projections.
|
||||
|
||||
The remaining generated files are necessary boundary adapters, not configuration concepts. Making
|
||||
the backend parse the authoring YAML independently would remove one file but restore two semantic
|
||||
validators. Hard-coding embedding values in Compose would remove one projection but restore a model
|
||||
source outside the catalog. Inferring authentication from missing fields would save one YAML key but
|
||||
turn a safe explicit choice into ambiguity. These apparent simplifications are therefore rejected.
|
||||
|
||||
No further reduction was found that preserves one authority, strict validation, explicit security,
|
||||
session determinism, and model-free workspaces.
|
||||
|
||||
## Implementation surface
|
||||
|
||||
Implementation must update the host `tht` installation loader, setup and lifecycle projection,
|
||||
doctor, backup/restore, Compose mounts and embedding inputs, backend catalog/settings/session model
|
||||
resolution, workspace schema and migration, runtime rendering and diagnostics, frontend workspace
|
||||
drafts and model filtering, examples, fixtures, and documentation. Existing session manifests remain
|
||||
readable and keep their pinned provider/model identity; only resume resolution changes to the new
|
||||
catalog.
|
||||
|
||||
This document authorizes design only. Software implementation begins only after a separate explicit
|
||||
request.
|
||||
@@ -0,0 +1,218 @@
|
||||
# Local sensitive-column classifier: TypeScript or Python
|
||||
|
||||
**Date:** 2026-09-02
|
||||
**Scope:** library/runtime choice for a local classifier without a generative LLM. Synthetic data generation and
|
||||
model-boundary policy for description generation are deliberately out of scope.
|
||||
|
||||
## Recommendation
|
||||
|
||||
Implement the mandatory classifier **inside the TypeScript backend**, with no learned model in its
|
||||
mandatory path. The agreed policy is dominated by deterministic work:
|
||||
type/declared-size rules, bounded string normalization, patterns, checksums, exact dictionaries,
|
||||
context terms, and "one match makes the column sensitive". TypeScript is fully adequate for that
|
||||
work and keeps the classifier in the process which already owns the catalog, source connectors,
|
||||
run lifecycle, cancellation, and the human override.
|
||||
|
||||
Use small, focused dependencies rather than a general NLP framework:
|
||||
|
||||
- `validator` for syntax/checksum-oriented recognizers such as email, IP, MAC, IBAN, BIC, payment
|
||||
cards, and locale-aware Italian tax IDs, passports, identity cards and VAT numbers; its official
|
||||
API exposes these validators individually, so unused validators need not be imported
|
||||
([validator.js source and API](https://github.com/validatorjs/validator.js/)).
|
||||
- `libphonenumber-js/max` for international telephone parsing and digit-pattern validation. The
|
||||
project's `max` metadata is intentionally more precise than its default/minimal metadata
|
||||
([libphonenumber-js validation documentation](https://github.com/catamphetamine/libphonenumber-js/blob/master/README.md#isvalid-boolean)).
|
||||
- ThothII-owned recognizers for identifiers not covered adequately by a selected validator (for
|
||||
example the Italian driving licence) and organization-specific identifiers, each with a version,
|
||||
tests, context words, and checksum/semantic validation where the identifier defines one.
|
||||
- A safe regular-expression boundary for administrator-authored rules. `node-re2` offers a mostly
|
||||
`RegExp`-compatible, ReDoS-resistant engine and a set API for matching many patterns, but is a
|
||||
native addon that downloads or builds a binary during installation. That packaging cost must be
|
||||
tested against the project images before adoption
|
||||
([node-re2 README](https://github.com/uhop/node-re2/blob/master/README.md)). If arbitrary regular
|
||||
expressions are not exposed, ThothII can instead keep a reviewed built-in pattern set and bound
|
||||
every inspected value.
|
||||
|
||||
Do **not** add NLP.js, spaCy, Presidio, or a transformer merely to implement regexes and dictionaries.
|
||||
They do not make deterministic identifiers more semantically understandable. In particular, NLP.js
|
||||
documents useful Italian tokenization plus enum/regex/built-in entities, but its advertised built-ins
|
||||
are mostly the same formatted values already covered above; it does not advertise a ready-made
|
||||
Italian PII model for arbitrary people or clinical concepts
|
||||
([NLP.js official README](https://github.com/axa-group/nlp.js/)).
|
||||
|
||||
Add an **optional, non-generative statistical NER pack** based on
|
||||
[`fastino/gliner2-privacy-filter-PII-multi`](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi).
|
||||
The selected model and its GLiNER2 runtime are Apache-2.0, its mDeBERTa base is MIT, and the model
|
||||
is trained across seven languages including Italian. Run the official implementation in one warmed,
|
||||
CPU-only Python worker; do not spawn it once per column or table. It returns entity evidence only.
|
||||
The TypeScript classifier remains the sole owner of the final column assessment and invokes NER
|
||||
only for bounded text that deterministic rules did not resolve and only while the scan deadline has
|
||||
time remaining.
|
||||
|
||||
This is a positive Q29 decision, not a placeholder for later library selection. The production gate
|
||||
is limited to pinning artifacts, producing an SBOM/license inventory, and measuring the selected
|
||||
model on PSD fixtures and target CPUs. If the optional pack is absent or misses its deadline, the
|
||||
normal coverage rule produces `unknown`; ThothII does not substitute a generative LLM.
|
||||
|
||||
## Why the TypeScript baseline is sufficient
|
||||
|
||||
The classifier is not being asked to infer an open-ended privacy judgment from prose. It executes a
|
||||
versioned policy and produces evidence. Its mandatory recognizers can be expressed as:
|
||||
|
||||
| Rule family | Implementation | Model needed? |
|
||||
| --- | --- | --- |
|
||||
| declared `text`/CLOB/unbounded string or declared maximum `>500` | catalog metadata rule | no |
|
||||
| an observed value longer than 500 characters | length rule, ideally detected in SQL before transferring the value | no |
|
||||
| email, IP/MAC, URL, UUID, payment card, IBAN/BIC, phone | pattern plus format/checksum validator | no |
|
||||
| Italian fiscal code/VAT/passport/identity card/driving licence | locale-aware validator where available, otherwise a reviewed country recognizer; pattern, context and checksum where defined | no |
|
||||
| credentials and technical secrets | anchored prefixes, token shapes, entropy/character-class heuristics, and organization allow/deny rules | no |
|
||||
| hospital- or customer-specific terms/codes | normalized exact dictionary or phrase matching | no |
|
||||
| person/location/organization in otherwise unstructured short text | statistical NER is useful but probabilistic | yes, optional |
|
||||
| clinical meaning in unstructured short Italian text | domain NER or a governed terminology; generic NER is not enough | optional and domain-specific |
|
||||
|
||||
This is not speculative parity with a Python implementation. Presidio's own supported-entity table
|
||||
shows that its global email, IBAN, credit-card and similar recognizers use pattern matching,
|
||||
validation, context and checksums, while its Italian fiscal-code/VAT/passport/identity-card/licence
|
||||
coverage is likewise rule-based
|
||||
([Presidio supported entities](https://presidio.dataprivacystack.org/supported_entities/)). Those
|
||||
mechanisms are portable; Presidio provides tested implementations and orchestration, not a unique
|
||||
Python-language capability.
|
||||
|
||||
For large dictionaries, exact normalized token/phrase matching is still deterministic NLP and does
|
||||
not require a learned model. The policy format should describe the rule rather than the library,
|
||||
for example `kind: email`, `kind: checksum_id`, `kind: regex`, `kind: phrase_set`, and
|
||||
`kind: max_length`. This makes a future engine change possible without changing persisted policy or
|
||||
assessment evidence.
|
||||
|
||||
## What Python/Presidio would materially add
|
||||
|
||||
Presidio is the strongest Python option if ThothII later needs a broader recognizer ecosystem. Its
|
||||
Analyzer composes rule-based recognizers, regexes, validation, NER and contextual enhancement; it
|
||||
also exposes custom recognizers, batch processing and decision tracing
|
||||
([Presidio Analyzer architecture](https://presidio.dataprivacystack.org/analyzer/)). Its current
|
||||
catalog includes Italian identifiers and its YAML registry supports custom patterns and language-
|
||||
specific context. Recent source also contains a `NoOpNlpEngine`, so a deterministic-only Presidio
|
||||
deployment can avoid loading spaCy; this removes model cost but not the Python runtime/process
|
||||
boundary
|
||||
([Presidio `NoOpNlpEngine` integration](https://github.com/data-privacy-stack/presidio/blob/main/presidio-analyzer/presidio_analyzer/recognizer_registry/recognizer_registry.py)).
|
||||
|
||||
`presidio-structured` is conceptually close because it maps tabular columns/JSON keys to detected
|
||||
entities, but its documented strategies select the most common, highest-confidence, or a mixed
|
||||
entity. ThothII would still need its own deadline/sampling/`unknown` semantics and human-override
|
||||
lifecycle; Presidio also lists improved mixed free-text/structured support as future work
|
||||
([Presidio Structured](https://presidio.dataprivacystack.org/structured/)).
|
||||
|
||||
That benefit does not currently justify using it as ThothII's mandatory engine:
|
||||
|
||||
- The backend currently has no PII/NLP dependencies and is an ESM TypeScript service
|
||||
(`backend/package.json`); the existing Python environment belongs to the separate harness/core
|
||||
layer and also has no Presidio or spaCy dependency (`harness/pyproject.toml`). A Python analyzer therefore means a new
|
||||
ownership and deployment boundary, not simply another import.
|
||||
- Presidio is Python-based and can be run as a package or REST container. Its REST API intentionally
|
||||
has no built-in authentication, so a sidecar would have to remain on a protected internal network
|
||||
and be guarded by infrastructure
|
||||
([Presidio FAQ, deployment](https://presidio.dataprivacystack.org/faq/#how-can-i-deploy-presidio-into-my-environment)).
|
||||
- Presidio defaults to English. Adding Italian requires both an Italian NLP pipeline and adapted
|
||||
language-specific recognizers/context; language-agnostic regexes alone do not solve that setup
|
||||
([Presidio multilingual guidance](https://github.com/data-privacy-stack/presidio/blob/main/docs/analyzer/languages.md)).
|
||||
- Presidio itself warns that automated detection cannot guarantee finding all sensitive data and
|
||||
documents an unavoidable false-positive/false-negative trade-off
|
||||
([Presidio FAQ](https://presidio.dataprivacystack.org/faq/#what-is-presidio)). It cannot turn a
|
||||
negative sample into proof that a column is safe.
|
||||
|
||||
## Learned NLP is not an LLM, but it is still a model
|
||||
|
||||
A spaCy NER pipeline does not generate text and is not an LLM. It is a statistical token classifier:
|
||||
it predicts spans such as `PERSON`, `LOCATION` and `ORGANIZATION`. spaCy explicitly notes that the
|
||||
predictions depend on the training examples and will not always be correct
|
||||
([spaCy NER documentation](https://spacy.io/usage/linguistic-features#named-entities)). That makes it
|
||||
a useful additional **positive detector**, not the authority which declares a sampled column safe.
|
||||
|
||||
spaCy publishes Italian CPU pipelines with an NER component, but they are trained on news/media
|
||||
data rather than clinical notes. More importantly, all three official packages (`it_core_news_sm`,
|
||||
`md`, and `lg`) are licensed **CC BY-NC-SA 3.0**, so they are a no-go for a product which can be
|
||||
used commercially; using the MIT-licensed spaCy runtime does not change the weights' license
|
||||
([spaCy Italian pipelines](https://spacy.io/models/it),
|
||||
[`sm` metadata](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_sm-3.8.0.json),
|
||||
[`md` metadata](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_md-3.8.0.json),
|
||||
[`lg` metadata](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_lg-3.8.0.json)). A general Italian model also recognizes
|
||||
names and places, not all sensitive clinical meaning. Training a suitable model would require a
|
||||
labeled, representative and legally usable corpus; Presidio's own model guidance likewise states
|
||||
that a labeled PII dataset is required to train a new model
|
||||
([Presidio spaCy/Stanza guidance](https://github.com/data-privacy-stack/presidio/blob/main/docs/analyzer/nlp_engines/spacy_stanza.md#training-your-own-model)).
|
||||
|
||||
Consequently, the absence of a local LLM is irrelevant to sensitivity assessment. Installations
|
||||
with a local LLM may let description generation receive broader source content under a separate
|
||||
model-boundary policy, but **the sensitivity classifier does not call that LLM**. The optional NER
|
||||
pack is a distinct local classifier capability and is never inferred from the description-model
|
||||
catalog.
|
||||
|
||||
## Licensing and redistribution gate
|
||||
|
||||
Framework code, model weights, and training data are three separate licensing layers. A permissive
|
||||
runtime does not grant ThothII permission to bundle weights trained from restricted data. For an
|
||||
open-source product that may be deployed commercially, the default rule should therefore be:
|
||||
**reject `NC`, research-only, custom restrictive, missing, or ambiguous model licenses**. Pinning an
|
||||
approved artifact must also pin its model card, upstream model, training datasets, required notices,
|
||||
and checksums. This is a technical compatibility screen, not legal advice.
|
||||
|
||||
| Candidate | Framework code | Weights and training-data evidence | Commercial bundling in ThothII | Decision |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Presidio, deterministic recognizers only | MIT; its license permits use and redistribution with the notice ([Presidio license](https://github.com/data-privacy-stack/presidio/blob/main/LICENSE)) | No general Italian model is required with `NoOpNlpEngine`; third-party dependencies still retain their notices | Compatible, but adds a Python process without adding unique detection capability for the agreed baseline | **GO legally; NO-GO architecturally for v1** |
|
||||
| spaCy runtime | MIT ([spaCy license](https://github.com/explosion/spaCy/blob/master/LICENSE)) | Runtime only; no permission is conferred on downloaded pipelines | Compatible if used with separately approved weights | **GO for code only** |
|
||||
| Official spaCy Italian `sm`/`md`/`lg` | spaCy code is MIT | Every current package declares CC BY-NC-SA 3.0; the metadata traces part of training to UD Italian ISDT under the same non-commercial license ([`sm`](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_sm-3.8.0.json), [`md`](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_md-3.8.0.json), [`lg`](https://raw.githubusercontent.com/explosion/spacy-models/master/meta/it_core_news_lg-3.8.0.json)) | `NC` excludes the intended commercial deployments; `SA` adds redistribution obligations | **NO-GO** |
|
||||
| Stanza runtime | Apache-2.0 ([Stanza license](https://github.com/stanfordnlp/stanza/blob/main/LICENSE)) | Stanford says its language packs are ODC-By only to the extent it owns the rights and explicitly tells users to inspect training-data licenses. Italian NER uses FBK's KIND data; KIND annotations are CC BY-NC 4.0 or CC BY-NC-SA 4.0 ([Stanza model licensing note](https://stanfordnlp.github.io/stanza/performance.html), [Italian NER source](https://stanfordnlp.github.io/stanza/ner_models.html), [KIND license](https://github.com/dhfbk/KIND#license)) | The Italian model's non-commercial source data fails the product requirement | **GO for code; NO-GO for Italian weights** |
|
||||
| Flair runtime and official multilingual NER | MIT ([Flair repository and license](https://github.com/flairNLP/flair)) | The official `ner-multi` model card neither includes Italian among its four trained languages nor declares a model license; a maintainer issue asking whether bundled weights are MIT was closed without an answer ([model card](https://huggingface.co/flair/ner-multi), [license issue](https://github.com/flairNLP/flair/issues/1487)) | Runtime is usable, but these weights are unsuitable and their redistribution terms are not established | **GO for code; NO-GO for these weights** |
|
||||
| `osiria/flare-it-ner` | Runs in Transformers/PyTorch; the card labels the weights MIT | Trained on CC BY 4.0 WikiNER plus an additional manually annotated custom dataset whose redistribution terms are not identified in the model card ([model card](https://huggingface.co/osiria/flare-it-ner), [WikiNER dataset](https://figshare.com/articles/dataset/Learning_multilingual_named_entity_recognition_from_Wikipedia/5462500)) | The weights look promising, but an MIT label alone does not resolve the undocumented custom training-data rights | **NO-GO until provenance is explicit** |
|
||||
| Transformers.js / ONNX runtime | Apache-2.0 ([Transformers.js license](https://github.com/huggingface/transformers.js/blob/main/LICENSE)) | ONNX conversion does not replace the source model's license or cure training-data restrictions; each exact model needs its own audit | Compatible runtime, but there is no blanket approval for arbitrary Hub/ONNX weights | **GO for code; weights case by case** |
|
||||
| `Laibniz/italian-ner-pii-browser-uncased` | Packaged for Transformers.js as quantized ONNX; model card says Apache-2.0 | It is a conversion of an Osiria model and therefore inherits the same undocumented custom-data provenance ([model card](https://huggingface.co/Laibniz/italian-ner-pii-browser-uncased)) | CPU/browser packaging does not cure the upstream licensing gap | **NO-GO** |
|
||||
| `Ar86Bat/multilang-pii-ner` | XLM-R/Transformers; model card labels weights MIT | Trained on `ai4privacy/open-pii-masking-500k-ai4privacy`, whose card declares CC BY 4.0 ([model card](https://huggingface.co/Ar86Bat/multilang-pii-ner), [dataset card](https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy)) | Commercial redistribution appears compatible with MIT notice plus CC BY attribution; an ONNX export, dependency audit, and PSD evaluation are still required | **CONDITIONAL GO as optional PII NER** |
|
||||
| `urchade/gliner_multi_pii-v1` | GLiNER code and model are Apache-2.0 | Its published multilingual synthetic dataset, including Italian, is Apache-2.0 ([model card](https://huggingface.co/urchade/gliner_multi_pii-v1), [dataset card](https://huggingface.co/datasets/urchade/synthetic-pii-ner-mistral-v1)) | The cleanest publicly inspectable license chain, but the published SPY comparison reports substantially lower recall than the selected Fastino model | **GO legally; reserve candidate** |
|
||||
| `fastino/gliner2-privacy-filter-PII-multi` | GLiNER2 code is Apache-2.0; the direct local dependencies are permissively licensed | Weights are Apache-2.0 and the mDeBERTa base is MIT. The authors report a purpose-built synthetic corpus spanning seven languages, including Italian; the corpus itself is not published, so training is not fully reproducible ([model card](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi), [paper](https://arxiv.org/abs/2605.09973), [base model](https://huggingface.co/microsoft/mdeberta-v3-base)) | Free local/commercial use and redistribution are allowed subject to Apache notices. It has the best reported exact-span F1 and recall among the compared open detectors on SPY, although that evaluation is not Italian-specific | **SELECTED optional NER pack** |
|
||||
| `openai/privacy-filter` | Code and weights are Apache-2.0 and it runs locally through Python or Transformers.js | The publisher explicitly permits commercial deployment, but documents the model as primarily English and warns that non-English performance may drop ([model card](https://huggingface.co/openai/privacy-filter), [source](https://github.com/openai/privacy-filter)) | Clean license and direct TypeScript path, but insufficient Italian evidence and a larger one-billion-parameter artifact make it a weaker fit | **GO legally; NO-GO as primary Italian detector** |
|
||||
| OpenMed Italian PII models | Cards label the weights Apache-2.0 | They are trained on `ai4privacy/pii-masking-400k`, whose license limits use to academic/non-commercial purposes unless a commercial agreement is obtained ([representative model](https://huggingface.co/OpenMed/OpenMed-PII-Italian-BioClinicalBERT-Base-110M-v1), [dataset license](https://huggingface.co/datasets/ai4privacy/pii-masking-400k/blob/main/LICENSE)) | An Apache label on derivative weights does not remove the explicit upstream non-commercial restriction | **NO-GO** |
|
||||
|
||||
The search therefore found a usable, free-of-charge, permissively licensed solution. Fastino
|
||||
GLiNER2-PII is selected because it combines an Apache code/weight grant, explicit Italian training,
|
||||
CPU operation, configurable entity labels, and materially higher published recall than the older
|
||||
Urchade model. This does not prove Italian clinical accuracy: the published SPY evaluation is
|
||||
English and the training corpus is synthetic. Those are quality and reproducibility limitations,
|
||||
not a reason to reintroduce a generative LLM into the classifier.
|
||||
|
||||
## Fit with current ThothII design
|
||||
|
||||
The current backend already owns source sampling and the human-owned `Sensitive Data Flag`; source
|
||||
values are transient and existing description sampling is bounded (`PROJECT_STATE.md` and
|
||||
`backend/src/catalog/description-source-sampler.ts`). The new
|
||||
classifier should therefore be a backend module behind the single effective sensitivity-decision
|
||||
point already being designed, while preserving these agreed semantics:
|
||||
|
||||
1. `sensitive`, `non_sensitive`, and `unknown` are assessment outcomes; `unknown` leaves the
|
||||
human-owned Sensitive Data Flag unchanged. Its effect on later description generation is out of
|
||||
scope here.
|
||||
2. Any positive rule match stops the column scan and yields `sensitive`.
|
||||
3. A negative time-bounded sample yields `unknown`, never `non_sensitive`.
|
||||
4. The administrator may set the final binary flag either way; the automated result remains a
|
||||
proposal with rule/version/coverage evidence and no stored source value.
|
||||
5. Learned NER, if present, can add positive evidence but cannot convert `unknown` to
|
||||
`non_sensitive`.
|
||||
|
||||
The five-second scan budget should be enforced at the database-read orchestration layer and not
|
||||
inside a recognizer library. An in-process deterministic engine maximizes the portion of that budget
|
||||
available to database reads and avoids model cold-start/IPC costs. It should stream bounded batches,
|
||||
check cancellation/deadline between batches, and stop at the first match.
|
||||
|
||||
## Acceptance gate for enabling the selected NER pack
|
||||
|
||||
Build a versioned, value-free fixture corpus representing at least: Italian/general identifiers,
|
||||
false-positive lookalikes, bounded short notes with people/places, credential formats, null/high-
|
||||
cardinality columns, and PSD-specific clinical codes. Measure per category rather than one aggregate
|
||||
score. Pin the GLiNER2 package and model revision/checksum, build an SBOM including transitive
|
||||
dependencies, retain required notices, prohibit model downloads at inference time, and measure warm
|
||||
latency and memory on the target CPU. The detector is enabled only where those checks pass.
|
||||
|
||||
The acceptance test may tune label descriptions and per-label confidence thresholds, but it does
|
||||
not reopen the architecture or silently select another Hub model. If quality or performance is
|
||||
unacceptable, the optional NER pack remains disabled and unresolved sampled columns remain
|
||||
`unknown`. Replacing the chosen model or adding a local-LLM fallback requires a new documented
|
||||
decision.
|
||||
Reference in New Issue
Block a user