feat: classify sensitive columns locally
This commit is contained in:
@@ -137,14 +137,17 @@ Generated descriptions can be requested for selected tables, selected columns, e
|
||||
target, or targets with a missing generated description. The backend accepts one installation-wide
|
||||
run and processes targets sequentially. Every catalog column has a **Sensitive** flag, which defaults
|
||||
to `false`, including after a newly discovered column is synchronized. Before generation, an
|
||||
administrator can ask the configured model to suggest flags from structural metadata only (database,
|
||||
schema, table and column names, data types, nullability, primary keys, and foreign keys). Suggestions
|
||||
remain an unsaved draft until a human reviews and saves them.
|
||||
administrator can run local sensitivity analysis over the selected database, tables, or columns.
|
||||
The analysis combines structural metadata with bounded read-only inspection of source values. It
|
||||
uses no generative AI and no installation-catalog model. Assessments remain an unsaved draft until
|
||||
a human reviews and saves them; the reviewer may reverse any proposal.
|
||||
|
||||
The page exposes separate histories for description generation and sensitive-field suggestions.
|
||||
Sensitive-suggestion history stores the selected model, scope, status, aggregate counts, timestamps,
|
||||
and sanitized events. It does not store the proposed per-column flags, prompts, raw model output, or
|
||||
provider diagnostics; closing an unsaved review still discards that draft.
|
||||
The page exposes separate histories for description generation and sensitivity analysis. Analysis
|
||||
history stores the local policy version, scope, status, aggregate `sensitive`, `non_sensitive`, and
|
||||
`unknown` counts, timestamps, and sanitized events. It does not store source values, per-column
|
||||
proposals, NER spans, or worker diagnostics; closing an unsaved review discards that draft. The
|
||||
rules, time bounds, and optional CPU-only NER profile are documented in
|
||||
[Local sensitivity analysis](sensitivity-analysis.md).
|
||||
|
||||
For a column with `sensitive=false`, the worker may read at most five source rows and five
|
||||
representative non-null values through a read-only connector. For `sensitive=true`, the source query
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
# Local sensitivity analysis
|
||||
|
||||
Database Management can assess selected columns without sending their metadata or contents to a
|
||||
generative model. The feature is advisory: it creates a transient review draft, while the catalog's
|
||||
Sensitive Data Flag changes only when an administrator explicitly saves a choice. The administrator
|
||||
may set either value, including overriding a `sensitive` proposal.
|
||||
|
||||
## Default policy
|
||||
|
||||
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v1`
|
||||
policy combines:
|
||||
|
||||
- normalized column-name rules for direct identifiers, credentials, and health data;
|
||||
- validated content rules for email, Italian fiscal code and VAT, passport, identity-card and
|
||||
driving-licence identifiers, phone numbers, IBAN/BIC, payment-card checksums, IP/MAC addresses,
|
||||
URLs, UUIDs, access keys, private-key markers, sensitive keys inside bounded recursive JSON, and
|
||||
a reviewed Italian clinical-term dictionary;
|
||||
- a conservative length rule: any observed textual value longer than 500 characters makes the
|
||||
entire column sensitive.
|
||||
|
||||
One decisive value is enough to classify the column as `sensitive`. A complete scan with no match
|
||||
may classify it as `non_sensitive`. Empty, all-null, binary/uninspectable, interrupted, and sampled
|
||||
no-match columns are `unknown`; an `unknown` draft preserves the current human flag.
|
||||
|
||||
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
|
||||
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
|
||||
persist source values. A full scan gets five seconds per table. If it cannot finish, the adapter uses
|
||||
a bounded repeatable sample within the sixty-second request deadline. PostgreSQL-wire reads run in a
|
||||
read-only transaction and always end with rollback. A REST scan can claim complete coverage only
|
||||
when it finishes in one request; multi-request pagination has no shared source transaction and is
|
||||
therefore conservatively reported as sampled.
|
||||
|
||||
The HTTP operation stops waiting at sixty seconds. The same expiring signal is checked before and
|
||||
after catalog selection, source access, progress writes, and every table. If it expires after a run
|
||||
has been created, that run is finalized as `interrupted` and all not-decisively-processed columns
|
||||
are counted as `unknown`; no review payload is returned from the timed-out request.
|
||||
|
||||
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
|
||||
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
|
||||
Neither path stores values, matched spans, prompts, or free-form model output.
|
||||
|
||||
## Optional CPU-only GLiNER2 evidence
|
||||
|
||||
The deterministic engine works without Python NER. An installation may opt into
|
||||
`fastino/gliner2-privacy-filter-PII-multi` for unresolved short text. It runs in a persistent local
|
||||
Python worker, adds sanitized evidence, and never becomes a second decision point. The worker:
|
||||
|
||||
- loads a local model directory only and forces Hugging Face/Transformers offline mode;
|
||||
- starts warming in the background when the backend starts; an analysis never waits for warm-up
|
||||
and skips NER until the worker is ready, so loading cannot consume the run's NER allowance;
|
||||
- hides CUDA and HIP devices and loads weights with `map_location="cpu"`;
|
||||
- starts with a scrubbed environment, then installs a fail-closed seccomp filter that denies
|
||||
network syscalls before accepting source text (the Python socket API is disabled as defense in depth);
|
||||
- receives at most two 500-character candidates per table by default, selected breadth-first
|
||||
across unresolved columns, and shares a ten-second NER allowance across the whole run;
|
||||
- returns only column ID, normalized label, and confidence; source text and entity spans are not
|
||||
returned or stored;
|
||||
- is skipped on timeout, startup failure, invalid output, or absent configuration. No LLM fallback
|
||||
is selected.
|
||||
|
||||
The pinned model revision is `c153999da5f4c509df4322b0c6a1baf3d2c284d7`. GLiNER2 and the model
|
||||
are Apache-2.0; the published mDeBERTa base is MIT. The optional runtime pins
|
||||
`gliner2[local]==2.0.0`, `transformers==4.57.6`, and the CPU-only PyTorch wheel
|
||||
`torch==2.14.0+cpu`. It lives in `/opt/sensitivity-ner`, is not installed in the default core image,
|
||||
and does not install CUDA packages.
|
||||
|
||||
The current upstream checkpoint was saved by Transformers 5.8 even though GLiNER2 2.0.0 officially
|
||||
requires Transformers `<5`; the resulting tokenizer error is independently reported in
|
||||
[GLiNER2 issue 145](https://github.com/fastino-ai/GLiNER2/issues/145). At startup ThothII leaves the
|
||||
pinned model directory unchanged and creates a temporary symlink view that maps the checkpoint's
|
||||
`extra_special_tokens` list to the Transformers 4 name `additional_special_tokens`. Any other or
|
||||
ambiguous shape fails closed and leaves the optional NER unavailable. The offline CPU smoke test
|
||||
must remain part of every dependency or model revision update.
|
||||
|
||||
## Prepare and enable the optional profile
|
||||
|
||||
Download happens during explicit installation, never during inference:
|
||||
|
||||
```bash
|
||||
./scripts/fetch-sensitivity-ner-model.sh /absolute/path/to/gliner2-pii
|
||||
```
|
||||
|
||||
The script builds the separate `thothii-core:sensitivity-ner` image, downloads the exact revision, and writes
|
||||
`MODEL_SHA256SUMS`. Keep the model directory outside the repository. Then set:
|
||||
|
||||
```bash
|
||||
export THOTH_ENABLE_SENSITIVITY_NER=1
|
||||
export THT_SENSITIVITY_NER_MODEL_DIR=/absolute/path/to/gliner2-pii
|
||||
export THT_SENSITIVITY_NER_THREADS=2
|
||||
./scripts/run-stack.sh
|
||||
```
|
||||
|
||||
For an operator-managed Compose invocation, include `deploy/compose.sensitivity-ner.yaml` after the
|
||||
base and installation overlays. The core build argument `INSTALL_SENSITIVITY_NER=true` installs the
|
||||
optional Python dependencies into their isolated virtualenv. The model mount is read-only. Values
|
||||
above eight threads are rejected; start with two so classification cannot contend heavily with
|
||||
other CPU workloads.
|
||||
|
||||
## Acceptance on a real database
|
||||
|
||||
Run the first evaluation in shadow mode: read the source with its existing read-only role, do not
|
||||
save proposed flags, and report only aggregate counts, rule IDs, coverage, and timings. Never copy
|
||||
matched values into test output. Use a separately approved, labeled Italian corpus to calculate
|
||||
precision and recall; raw PSD values must remain inside the authorized environment.
|
||||
|
||||
Inside the configured core runtime, the non-mutating command is:
|
||||
|
||||
```bash
|
||||
npm run sensitivity:shadow -- psd-clinical
|
||||
```
|
||||
|
||||
It reads catalog metadata and source values but emits one aggregate JSON object with no database,
|
||||
table, column, or source-value detail. It neither creates an analysis run nor updates a flag.
|
||||
|
||||
Enabling NER by default requires all of these gates:
|
||||
|
||||
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
|
||||
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
|
||||
3. the CPU benchmark stays within the configured deadlines and does not use a GPU;
|
||||
4. the labeled Italian evaluation meets thresholds approved by the product owner.
|
||||
|
||||
If a gate fails, leave NER disabled. The deterministic policy remains available and unresolved
|
||||
columns remain `unknown` rather than being sent to an internal or external LLM.
|
||||
|
||||
The first aggregate PSD shadow comparison is recorded in
|
||||
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
|
||||
local CPU runner, NER found additional entities but reduced total coverage inside the 60-second
|
||||
deadline, so the accepted setting remains disabled by default.
|
||||
Reference in New Issue
Block a user