129 lines
7.4 KiB
Markdown
129 lines
7.4 KiB
Markdown
# Local sensitivity analysis
|
|
|
|
Database Management can assess selected columns without sending their metadata or contents to a
|
|
generative model. The feature is advisory: it creates a transient review draft, while the catalog's
|
|
Sensitive Data Flag changes only when an administrator explicitly saves a choice. The administrator
|
|
may set either value, including overriding a `sensitive` proposal.
|
|
|
|
## Default policy
|
|
|
|
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v1`
|
|
policy combines:
|
|
|
|
- normalized column-name rules for direct identifiers, credentials, and health data;
|
|
- validated content rules for email, Italian fiscal code and VAT, passport, identity-card and
|
|
driving-licence identifiers, phone numbers, IBAN/BIC, payment-card checksums, IP/MAC addresses,
|
|
URLs, UUIDs, access keys, private-key markers, sensitive keys inside bounded recursive JSON, and
|
|
a reviewed Italian clinical-term dictionary;
|
|
- a conservative length rule: any observed textual value longer than 500 characters makes the
|
|
entire column sensitive.
|
|
|
|
One decisive value is enough to classify the column as `sensitive`. A complete scan with no match
|
|
may classify it as `non_sensitive`. Empty, all-null, binary/uninspectable, interrupted, and sampled
|
|
no-match columns are `unknown`; an `unknown` draft preserves the current human flag.
|
|
|
|
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
|
|
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
|
|
persist source values. A full scan gets five seconds per table. If it cannot finish, the adapter uses
|
|
a bounded repeatable sample within the sixty-second request deadline. PostgreSQL-wire reads run in a
|
|
read-only transaction and always end with rollback. A REST scan can claim complete coverage only
|
|
when it finishes in one request; multi-request pagination has no shared source transaction and is
|
|
therefore conservatively reported as sampled.
|
|
|
|
The HTTP operation stops waiting at sixty seconds. The same expiring signal is checked before and
|
|
after catalog selection, source access, progress writes, and every table. If it expires after a run
|
|
has been created, that run is finalized as `interrupted` and all not-decisively-processed columns
|
|
are counted as `unknown`; no review payload is returned from the timed-out request.
|
|
|
|
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
|
|
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
|
|
Neither path stores values, matched spans, prompts, or free-form model output.
|
|
|
|
## Optional CPU-only GLiNER2 evidence
|
|
|
|
The deterministic engine works without Python NER. An installation may opt into
|
|
`fastino/gliner2-privacy-filter-PII-multi` for unresolved short text. It runs in a persistent local
|
|
Python worker, adds sanitized evidence, and never becomes a second decision point. The worker:
|
|
|
|
- loads a local model directory only and forces Hugging Face/Transformers offline mode;
|
|
- starts warming in the background when the backend starts; an analysis never waits for warm-up
|
|
and skips NER until the worker is ready, so loading cannot consume the run's NER allowance;
|
|
- hides CUDA and HIP devices and loads weights with `map_location="cpu"`;
|
|
- starts with a scrubbed environment, then installs a fail-closed seccomp filter that denies
|
|
network syscalls before accepting source text (the Python socket API is disabled as defense in depth);
|
|
- receives at most two 500-character candidates per table by default, selected breadth-first
|
|
across unresolved columns, and shares a ten-second NER allowance across the whole run;
|
|
- returns only column ID, normalized label, and confidence; source text and entity spans are not
|
|
returned or stored;
|
|
- is skipped on timeout, startup failure, invalid output, or absent configuration. No LLM fallback
|
|
is selected.
|
|
|
|
The pinned model revision is `c153999da5f4c509df4322b0c6a1baf3d2c284d7`. GLiNER2 and the model
|
|
are Apache-2.0; the published mDeBERTa base is MIT. The optional runtime pins
|
|
`gliner2[local]==2.0.0`, `transformers==4.57.6`, and the CPU-only PyTorch wheel
|
|
`torch==2.14.0+cpu`. It lives in `/opt/sensitivity-ner`, is not installed in the default core image,
|
|
and does not install CUDA packages.
|
|
|
|
The current upstream checkpoint was saved by Transformers 5.8 even though GLiNER2 2.0.0 officially
|
|
requires Transformers `<5`; the resulting tokenizer error is independently reported in
|
|
[GLiNER2 issue 145](https://github.com/fastino-ai/GLiNER2/issues/145). At startup ThothII leaves the
|
|
pinned model directory unchanged and creates a temporary symlink view that maps the checkpoint's
|
|
`extra_special_tokens` list to the Transformers 4 name `additional_special_tokens`. Any other or
|
|
ambiguous shape fails closed and leaves the optional NER unavailable. The offline CPU smoke test
|
|
must remain part of every dependency or model revision update.
|
|
|
|
## Prepare and enable the optional profile
|
|
|
|
Download happens during explicit installation, never during inference:
|
|
|
|
```bash
|
|
./scripts/fetch-sensitivity-ner-model.sh /absolute/path/to/gliner2-pii
|
|
```
|
|
|
|
The script builds the separate `thothii-core:sensitivity-ner` image, downloads the exact revision, and writes
|
|
`MODEL_SHA256SUMS`. Keep the model directory outside the repository. Then set:
|
|
|
|
```bash
|
|
export THOTH_ENABLE_SENSITIVITY_NER=1
|
|
export THT_SENSITIVITY_NER_MODEL_DIR=/absolute/path/to/gliner2-pii
|
|
export THT_SENSITIVITY_NER_THREADS=2
|
|
./scripts/run-stack.sh
|
|
```
|
|
|
|
For an operator-managed Compose invocation, include `deploy/compose.sensitivity-ner.yaml` after the
|
|
base and installation overlays. The core build argument `INSTALL_SENSITIVITY_NER=true` installs the
|
|
optional Python dependencies into their isolated virtualenv. The model mount is read-only. Values
|
|
above eight threads are rejected; start with two so classification cannot contend heavily with
|
|
other CPU workloads.
|
|
|
|
## Acceptance on a real database
|
|
|
|
Run the first evaluation in shadow mode: read the source with its existing read-only role, do not
|
|
save proposed flags, and report only aggregate counts, rule IDs, coverage, and timings. Never copy
|
|
matched values into test output. Use a separately approved, labeled Italian corpus to calculate
|
|
precision and recall; raw PSD values must remain inside the authorized environment.
|
|
|
|
Inside the configured core runtime, the non-mutating command is:
|
|
|
|
```bash
|
|
npm run sensitivity:shadow -- psd-clinical
|
|
```
|
|
|
|
It reads catalog metadata and source values but emits one aggregate JSON object with no database,
|
|
table, column, or source-value detail. It neither creates an analysis run nor updates a flag.
|
|
|
|
Enabling NER by default requires all of these gates:
|
|
|
|
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
|
|
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
|
|
3. the CPU benchmark stays within the configured deadlines and does not use a GPU;
|
|
4. the labeled Italian evaluation meets thresholds approved by the product owner.
|
|
|
|
If a gate fails, leave NER disabled. The deterministic policy remains available and unresolved
|
|
columns remain `unknown` rather than being sent to an internal or external LLM.
|
|
|
|
The first aggregate PSD shadow comparison is recorded in
|
|
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
|
|
local CPU runner, NER found additional entities but reduced total coverage inside the 60-second
|
|
deadline, so the accepted setting remains disabled by default.
|