feat: classify sensitive columns locally

This commit is contained in:
Codex
2026-09-03 02:11:13 +02:00
parent 7b87e95427
commit f114d0065a
57 changed files with 4038 additions and 1149 deletions
+10 -7
View File
@@ -137,14 +137,17 @@ Generated descriptions can be requested for selected tables, selected columns, e
target, or targets with a missing generated description. The backend accepts one installation-wide
run and processes targets sequentially. Every catalog column has a **Sensitive** flag, which defaults
to `false`, including after a newly discovered column is synchronized. Before generation, an
administrator can ask the configured model to suggest flags from structural metadata only (database,
schema, table and column names, data types, nullability, primary keys, and foreign keys). Suggestions
remain an unsaved draft until a human reviews and saves them.
administrator can run local sensitivity analysis over the selected database, tables, or columns.
The analysis combines structural metadata with bounded read-only inspection of source values. It
uses no generative AI and no installation-catalog model. Assessments remain an unsaved draft until
a human reviews and saves them; the reviewer may reverse any proposal.
The page exposes separate histories for description generation and sensitive-field suggestions.
Sensitive-suggestion history stores the selected model, scope, status, aggregate counts, timestamps,
and sanitized events. It does not store the proposed per-column flags, prompts, raw model output, or
provider diagnostics; closing an unsaved review still discards that draft.
The page exposes separate histories for description generation and sensitivity analysis. Analysis
history stores the local policy version, scope, status, aggregate `sensitive`, `non_sensitive`, and
`unknown` counts, timestamps, and sanitized events. It does not store source values, per-column
proposals, NER spans, or worker diagnostics; closing an unsaved review discards that draft. The
rules, time bounds, and optional CPU-only NER profile are documented in
[Local sensitivity analysis](sensitivity-analysis.md).
For a column with `sensitive=false`, the worker may read at most five source rows and five
representative non-null values through a read-only connector. For `sensitive=true`, the source query
+128
View File
@@ -0,0 +1,128 @@
# Local sensitivity analysis
Database Management can assess selected columns without sending their metadata or contents to a
generative model. The feature is advisory: it creates a transient review draft, while the catalog's
Sensitive Data Flag changes only when an administrator explicitly saves a choice. The administrator
may set either value, including overriding a `sensitive` proposal.
## Default policy
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v1`
policy combines:
- normalized column-name rules for direct identifiers, credentials, and health data;
- validated content rules for email, Italian fiscal code and VAT, passport, identity-card and
driving-licence identifiers, phone numbers, IBAN/BIC, payment-card checksums, IP/MAC addresses,
URLs, UUIDs, access keys, private-key markers, sensitive keys inside bounded recursive JSON, and
a reviewed Italian clinical-term dictionary;
- a conservative length rule: any observed textual value longer than 500 characters makes the
entire column sensitive.
One decisive value is enough to classify the column as `sensitive`. A complete scan with no match
may classify it as `non_sensitive`. Empty, all-null, binary/uninspectable, interrupted, and sampled
no-match columns are `unknown`; an `unknown` draft preserves the current human flag.
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
persist source values. A full scan gets five seconds per table. If it cannot finish, the adapter uses
a bounded repeatable sample within the sixty-second request deadline. PostgreSQL-wire reads run in a
read-only transaction and always end with rollback. A REST scan can claim complete coverage only
when it finishes in one request; multi-request pagination has no shared source transaction and is
therefore conservatively reported as sampled.
The HTTP operation stops waiting at sixty seconds. The same expiring signal is checked before and
after catalog selection, source access, progress writes, and every table. If it expires after a run
has been created, that run is finalized as `interrupted` and all not-decisively-processed columns
are counted as `unknown`; no review payload is returned from the timed-out request.
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
Neither path stores values, matched spans, prompts, or free-form model output.
## Optional CPU-only GLiNER2 evidence
The deterministic engine works without Python NER. An installation may opt into
`fastino/gliner2-privacy-filter-PII-multi` for unresolved short text. It runs in a persistent local
Python worker, adds sanitized evidence, and never becomes a second decision point. The worker:
- loads a local model directory only and forces Hugging Face/Transformers offline mode;
- starts warming in the background when the backend starts; an analysis never waits for warm-up
and skips NER until the worker is ready, so loading cannot consume the run's NER allowance;
- hides CUDA and HIP devices and loads weights with `map_location="cpu"`;
- starts with a scrubbed environment, then installs a fail-closed seccomp filter that denies
network syscalls before accepting source text (the Python socket API is disabled as defense in depth);
- receives at most two 500-character candidates per table by default, selected breadth-first
across unresolved columns, and shares a ten-second NER allowance across the whole run;
- returns only column ID, normalized label, and confidence; source text and entity spans are not
returned or stored;
- is skipped on timeout, startup failure, invalid output, or absent configuration. No LLM fallback
is selected.
The pinned model revision is `c153999da5f4c509df4322b0c6a1baf3d2c284d7`. GLiNER2 and the model
are Apache-2.0; the published mDeBERTa base is MIT. The optional runtime pins
`gliner2[local]==2.0.0`, `transformers==4.57.6`, and the CPU-only PyTorch wheel
`torch==2.14.0+cpu`. It lives in `/opt/sensitivity-ner`, is not installed in the default core image,
and does not install CUDA packages.
The current upstream checkpoint was saved by Transformers 5.8 even though GLiNER2 2.0.0 officially
requires Transformers `<5`; the resulting tokenizer error is independently reported in
[GLiNER2 issue 145](https://github.com/fastino-ai/GLiNER2/issues/145). At startup ThothII leaves the
pinned model directory unchanged and creates a temporary symlink view that maps the checkpoint's
`extra_special_tokens` list to the Transformers 4 name `additional_special_tokens`. Any other or
ambiguous shape fails closed and leaves the optional NER unavailable. The offline CPU smoke test
must remain part of every dependency or model revision update.
## Prepare and enable the optional profile
Download happens during explicit installation, never during inference:
```bash
./scripts/fetch-sensitivity-ner-model.sh /absolute/path/to/gliner2-pii
```
The script builds the separate `thothii-core:sensitivity-ner` image, downloads the exact revision, and writes
`MODEL_SHA256SUMS`. Keep the model directory outside the repository. Then set:
```bash
export THOTH_ENABLE_SENSITIVITY_NER=1
export THT_SENSITIVITY_NER_MODEL_DIR=/absolute/path/to/gliner2-pii
export THT_SENSITIVITY_NER_THREADS=2
./scripts/run-stack.sh
```
For an operator-managed Compose invocation, include `deploy/compose.sensitivity-ner.yaml` after the
base and installation overlays. The core build argument `INSTALL_SENSITIVITY_NER=true` installs the
optional Python dependencies into their isolated virtualenv. The model mount is read-only. Values
above eight threads are rejected; start with two so classification cannot contend heavily with
other CPU workloads.
## Acceptance on a real database
Run the first evaluation in shadow mode: read the source with its existing read-only role, do not
save proposed flags, and report only aggregate counts, rule IDs, coverage, and timings. Never copy
matched values into test output. Use a separately approved, labeled Italian corpus to calculate
precision and recall; raw PSD values must remain inside the authorized environment.
Inside the configured core runtime, the non-mutating command is:
```bash
npm run sensitivity:shadow -- psd-clinical
```
It reads catalog metadata and source values but emits one aggregate JSON object with no database,
table, column, or source-value detail. It neither creates an analysis run nor updates a flag.
Enabling NER by default requires all of these gates:
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
3. the CPU benchmark stays within the configured deadlines and does not use a GPU;
4. the labeled Italian evaluation meets thresholds approved by the product owner.
If a gate fails, leave NER disabled. The deterministic policy remains available and unresolved
columns remain `unknown` rather than being sent to an internal or external LLM.
The first aggregate PSD shadow comparison is recorded in
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
local CPU runner, NER found additional entities but reduced total coverage inside the 60-second
deadline, so the accepted setting remains disabled by default.