Files
ThothII/docs/adr/0014-assess-sensitive-columns-locally-from-source-content.md

4.0 KiB

status
status
superseded by ADR-0015

Assess sensitive columns locally from source content

The Sensitive Data Flag remains a human-owned boolean. An explicit, selection-scoped sensitivity analysis may propose changes by inspecting both catalog metadata and source values, but it never writes the flag. The administrator may accept, reject, or reverse every proposal.

One TypeScript SensitivityClassifier is the only component allowed to produce the column-level assessment sensitive, non_sensitive, or unknown. It applies a versioned Sensitive Data Policy and consumes values through database-independent streaming adapters. Database-specific code may read and normalize bounded values, but it may not decide sensitivity.

The classifier first applies deterministic metadata rules, value validators, checksums, dictionaries, and length rules. A single validated sensitive match makes the whole column sensitive; any textual value longer than 500 characters is such a match. It attempts a complete scan, but after five seconds per table it continues by sampling within the remaining run budget. A completed scan with no finding may produce non_sensitive; an incomplete scan with no finding produces unknown.

Ambiguous text may additionally be sent to an optional local NER detector only while time remains. By default it receives at most two candidates per table and shares a ten-second allowance across the entire analysis run. The detector runs on CPU, receives no tools or network access (enforced inside the worker with a fail-closed seccomp network-syscall filter), does not persist source values, and returns evidence rather than the column decision. The initial supported detector is fastino/gliner2-privacy-filter-PII-multi, used through the Apache-2.0 GLiNER2 Python library with a pinned model revision. Its model weights and GLiNER2 code are Apache-2.0, and its mDeBERTa base model is MIT. It is trained for seven languages including Italian and can run on CPU without using the installation's GPUs.

The selected checkpoint currently carries Transformers 5 tokenizer metadata while the released GLiNER2 2.0.0 runtime requires Transformers 4. ThothII may bridge only that known key rename in a temporary view of the immutable, checksummed artifact; unexpected or ambiguous metadata fails closed. Removing the compatibility bridge requires an offline smoke test against a corrected, pinned upstream release.

The NER detector is an optional installation asset because its weights and runtime are materially larger than the deterministic TypeScript engine. It is invoked only for otherwise unresolved text, never for values already classified by a decisive rule. If it is disabled, unavailable, times out, or returns no qualifying evidence before the deadline, the classifier follows the same coverage rule and may return unknown.

No generative LLM, internal or external, participates in sensitivity assessment. A fallback to an installation-local LLM is unnecessary while a permissively licensed local NER implementation is available, and would reintroduce queue latency, non-deterministic judgments, prompt-injection surface, and contention with normal inference. Introducing such a fallback would require a new decision based on evidence that the NER path is unusable.

This decision supersedes ADR-0011 only where that ADR assigns draft sensitivity suggestions to an AI using structural metadata. ADR-0011's human authority, transient draft, explicit scope, and Sensitive Data Flag remain in force. How description generation uses the flag is outside this decision.

The selected model's published evaluation is not an Italian production acceptance test. Before enabling the NER profile by default, ThothII must pin the artifacts, generate a dependency/license inventory, and pass a CPU benchmark plus a labeled Italian corpus representative of the target databases. Failure of those gates disables NER; it does not silently select another model or an LLM.