Local sensitivity analysis¶
Database Management can assess selected columns without sending their metadata or contents to a
generative model. The feature is advisory: it creates a transient review draft, while the catalog's
Sensitive Data Flag changes only when an administrator explicitly saves a choice. The administrator
may set either value, including overriding a sensitive proposal.
Default policy¶
SensitivityClassifier is the only column-level decision point. The versioned sensitivity-v4
policy combines:
- a structural exclusion for declared
bigintprimary-key columns and undeclaredbigintcolumns following the exactpknaming convention; their values are non-informative identifiers and are therefore not inspected as possible sensitive content. The evidence distinguishes declared constraints from convention-based inference; - normalized column-name rules for direct identifiers, credentials, and health data;
- validated content rules for email, Italian fiscal code and VAT, passport, identity-card and driving-licence identifiers, phone numbers, IBAN/BIC, payment-card checksums, IP/MAC addresses, URLs, UUIDs, access keys, private-key markers, sensitive keys inside bounded recursive JSON, and a reviewed Italian clinical-term dictionary;
- a conservative length rule: any observed textual value longer than 500 characters makes the entire column sensitive.
One decisive value is enough to classify the column as sensitive and removes it from subsequent
passes. Binary or otherwise uninspectable column types are also proposed as sensitive, because
their contents cannot be cleared by the textual rules. A completed analysis has only two draft
outcomes: sensitive and non_sensitive. Empty or all-null columns are non_sensitive with
no_values coverage; a sampled column with no match is non_sensitive with explicit sampled
coverage. The administrator remains free to reverse either proposal before saving it.
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
REST run_query adapters project at most 501 characters per value, use only SELECT, and never
persist source values. Tables proven to contain at most 1,000 rows are fully scanned. Larger tables
are processed breadth-first so every table gets the cheapest pass before any table gets a deeper
one:
- inspect up to 300 non-null values per unresolved column;
- inspect up to 700 additional values, reaching a 1,000-value target;
- for unresolved text, JSON, and XML columns only, inspect up to 2,000 additional values, reaching a 3,000-value target.
At most two tables are scanned concurrently, and the database adapter groups at most 25 columns in
one source query. Each probe or value query has a five-second statement timeout; PostgreSQL-wire
reads run in a read-only transaction and always end with rollback. Sampling is bounded and
repeatable for a policy version. If a randomized sample is empty or reaches its query timeout, the
adapter tries one sequential bounded sample; if that also times out, the source error fails the run
and returns no review instead of manufacturing unknown decisions.
There is no global sixty-second analysis deadline. Work is bounded by sample counts, per-query
timeouts, and early column exits. The operation is interrupted only when its request connection is
aborted or the backend restarts. Historical or interrupted run counters named unknown represent
columns that were not processed; unknown is not a sensitivity-v4 column assessment.
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted. Neither path stores values, matched spans, prompts, or free-form model output.
Optional CPU-only GLiNER2 evidence¶
The deterministic engine works without Python NER. An installation may opt into
fastino/gliner2-privacy-filter-PII-multi for unresolved short text. It runs in a persistent local
Python worker, adds sanitized evidence, and never becomes a second decision point. The worker:
- loads a local model directory only and forces Hugging Face/Transformers offline mode;
- starts warming in the background when the backend starts; an analysis never waits for warm-up and skips NER until the worker is ready, so loading cannot consume the run's NER allowance;
- hides CUDA and HIP devices and loads weights with
map_location="cpu"; - starts with a scrubbed environment, then installs a fail-closed seccomp filter that denies network syscalls before accepting source text (the Python socket API is disabled as defense in depth);
- receives at most two 500-character candidates per table by default, selected breadth-first across unresolved columns, and shares a ten-second NER allowance across the whole run;
- returns only column ID, normalized label, and confidence; source text and entity spans are not returned or stored;
- is skipped on timeout, startup failure, invalid output, or absent configuration. No LLM fallback is selected.
The pinned model revision is c153999da5f4c509df4322b0c6a1baf3d2c284d7. GLiNER2 and the model
are Apache-2.0; the published mDeBERTa base is MIT. The optional runtime pins
gliner2[local]==2.0.0, transformers==4.57.6, and the CPU-only PyTorch wheel
torch==2.14.0+cpu. It lives in /opt/sensitivity-ner, is not installed in the default core image,
and does not install CUDA packages.
The current upstream checkpoint was saved by Transformers 5.8 even though GLiNER2 2.0.0 officially
requires Transformers <5; the resulting tokenizer error is independently reported in
GLiNER2 issue 145. At startup ThothII leaves the
pinned model directory unchanged and creates a temporary symlink view that maps the checkpoint's
extra_special_tokens list to the Transformers 4 name additional_special_tokens. Any other or
ambiguous shape fails closed and leaves the optional NER unavailable. The offline CPU smoke test
must remain part of every dependency or model revision update.
Prepare and enable the optional profile¶
Download happens during explicit installation, never during inference:
./scripts/fetch-sensitivity-ner-model.sh /absolute/path/to/gliner2-pii
The script builds the separate thothii-core:sensitivity-ner image, downloads the exact revision, and writes
MODEL_SHA256SUMS. Keep the model directory outside the repository. Then set:
export THOTH_ENABLE_SENSITIVITY_NER=1
export THT_SENSITIVITY_NER_MODEL_DIR=/absolute/path/to/gliner2-pii
export THT_SENSITIVITY_NER_THREADS=2
./scripts/run-stack.sh
For an operator-managed Compose invocation, include deploy/compose.sensitivity-ner.yaml after the
base and installation overlays. The core build argument INSTALL_SENSITIVITY_NER=true installs the
optional Python dependencies into their isolated virtualenv. The model mount is read-only. Values
above eight threads are rejected; start with two so classification cannot contend heavily with
other CPU workloads.
Acceptance on a real database¶
Run the first evaluation in shadow mode: read the source with its existing read-only role, do not save proposed flags, and report only aggregate counts, rule IDs, coverage, and timings. Never copy matched values into test output. Use a separately approved, labeled Italian corpus to calculate precision and recall; source values must remain inside the authorized environment.
Inside the configured core runtime, the non-mutating command is:
npm run sensitivity:shadow -- <workspace-id>
It reads catalog metadata and source values but emits one aggregate JSON object with no database, table, column, or source-value detail. It neither creates an analysis run nor updates a flag.
Enabling NER by default requires all of these gates:
- the pinned artifact and
MODEL_SHA256SUMSare archived with the installation inventory; - the Python dependency/license inventory contains only redistribution-compatible licenses;
- the CPU benchmark stays within the configured NER allowance and does not use a GPU;
- the labeled Italian evaluation meets thresholds approved by the product owner.
If a gate fails, leave NER disabled. The deterministic policy remains available and produces the binary draft from its scan coverage; no content is sent to an internal or external LLM.
Benchmarks from a particular installation are not a guarantee for another database or machine. NER remains opt-in until an approved evaluation establishes that additional findings justify their false-positive rate and operational cost. Keep benchmark and release records with the installation's technical evidence, separate from this operator procedure.