145 lines
8.5 KiB
Markdown
145 lines
8.5 KiB
Markdown
# Local sensitivity analysis
|
|
|
|
Database Management can assess selected columns without sending their metadata or contents to a
|
|
generative model. The feature is advisory: it creates a transient review draft, while the catalog's
|
|
Sensitive Data Flag changes only when an administrator explicitly saves a choice. The administrator
|
|
may set either value, including overriding a `sensitive` proposal.
|
|
|
|
## Default policy
|
|
|
|
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v2`
|
|
policy combines:
|
|
|
|
- normalized column-name rules for direct identifiers, credentials, and health data;
|
|
- validated content rules for email, Italian fiscal code and VAT, passport, identity-card and
|
|
driving-licence identifiers, phone numbers, IBAN/BIC, payment-card checksums, IP/MAC addresses,
|
|
URLs, UUIDs, access keys, private-key markers, sensitive keys inside bounded recursive JSON, and
|
|
a reviewed Italian clinical-term dictionary;
|
|
- a conservative length rule: any observed textual value longer than 500 characters makes the
|
|
entire column sensitive.
|
|
|
|
One decisive value is enough to classify the column as `sensitive` and removes it from subsequent
|
|
passes. Binary or otherwise uninspectable column types are also proposed as `sensitive`, because
|
|
their contents cannot be cleared by the textual rules. A completed analysis has only two draft
|
|
outcomes: `sensitive` and `non_sensitive`. Empty or all-null columns are `non_sensitive` with
|
|
`no_values` coverage; a sampled column with no match is `non_sensitive` with explicit sampled
|
|
coverage. The administrator remains free to reverse either proposal before saving it.
|
|
|
|
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
|
|
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
|
|
persist source values. Tables proven to contain at most 1,000 rows are fully scanned. Larger tables
|
|
are processed breadth-first so every table gets the cheapest pass before any table gets a deeper
|
|
one:
|
|
|
|
1. inspect up to 300 non-null values per unresolved column;
|
|
2. inspect up to 700 additional values, reaching a 1,000-value target;
|
|
3. for unresolved text, JSON, and XML columns only, inspect up to 2,000 additional values, reaching
|
|
a 3,000-value target.
|
|
|
|
At most two tables are scanned concurrently, and the database adapter groups at most 25 columns in
|
|
one source query. Each probe or value query has a five-second statement timeout; PostgreSQL-wire
|
|
reads run in a read-only transaction and always end with rollback. Sampling is bounded and
|
|
repeatable for a policy version. If a randomized sample is empty or reaches its query timeout, the
|
|
adapter tries one sequential bounded sample; if that also times out, the source error fails the run
|
|
and returns no review instead of manufacturing `unknown` decisions.
|
|
|
|
There is no global sixty-second analysis deadline. Work is bounded by sample counts, per-query
|
|
timeouts, and early column exits. The operation is interrupted only when its request connection is
|
|
aborted or the backend restarts. Historical or interrupted run counters named `unknown` represent
|
|
columns that were not processed; `unknown` is not a `sensitivity-v2` column assessment.
|
|
|
|
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
|
|
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
|
|
Neither path stores values, matched spans, prompts, or free-form model output.
|
|
|
|
## Optional CPU-only GLiNER2 evidence
|
|
|
|
The deterministic engine works without Python NER. An installation may opt into
|
|
`fastino/gliner2-privacy-filter-PII-multi` for unresolved short text. It runs in a persistent local
|
|
Python worker, adds sanitized evidence, and never becomes a second decision point. The worker:
|
|
|
|
- loads a local model directory only and forces Hugging Face/Transformers offline mode;
|
|
- starts warming in the background when the backend starts; an analysis never waits for warm-up
|
|
and skips NER until the worker is ready, so loading cannot consume the run's NER allowance;
|
|
- hides CUDA and HIP devices and loads weights with `map_location="cpu"`;
|
|
- starts with a scrubbed environment, then installs a fail-closed seccomp filter that denies
|
|
network syscalls before accepting source text (the Python socket API is disabled as defense in depth);
|
|
- receives at most two 500-character candidates per table by default, selected breadth-first
|
|
across unresolved columns, and shares a ten-second NER allowance across the whole run;
|
|
- returns only column ID, normalized label, and confidence; source text and entity spans are not
|
|
returned or stored;
|
|
- is skipped on timeout, startup failure, invalid output, or absent configuration. No LLM fallback
|
|
is selected.
|
|
|
|
The pinned model revision is `c153999da5f4c509df4322b0c6a1baf3d2c284d7`. GLiNER2 and the model
|
|
are Apache-2.0; the published mDeBERTa base is MIT. The optional runtime pins
|
|
`gliner2[local]==2.0.0`, `transformers==4.57.6`, and the CPU-only PyTorch wheel
|
|
`torch==2.14.0+cpu`. It lives in `/opt/sensitivity-ner`, is not installed in the default core image,
|
|
and does not install CUDA packages.
|
|
|
|
The current upstream checkpoint was saved by Transformers 5.8 even though GLiNER2 2.0.0 officially
|
|
requires Transformers `<5`; the resulting tokenizer error is independently reported in
|
|
[GLiNER2 issue 145](https://github.com/fastino-ai/GLiNER2/issues/145). At startup ThothII leaves the
|
|
pinned model directory unchanged and creates a temporary symlink view that maps the checkpoint's
|
|
`extra_special_tokens` list to the Transformers 4 name `additional_special_tokens`. Any other or
|
|
ambiguous shape fails closed and leaves the optional NER unavailable. The offline CPU smoke test
|
|
must remain part of every dependency or model revision update.
|
|
|
|
## Prepare and enable the optional profile
|
|
|
|
Download happens during explicit installation, never during inference:
|
|
|
|
```bash
|
|
./scripts/fetch-sensitivity-ner-model.sh /absolute/path/to/gliner2-pii
|
|
```
|
|
|
|
The script builds the separate `thothii-core:sensitivity-ner` image, downloads the exact revision, and writes
|
|
`MODEL_SHA256SUMS`. Keep the model directory outside the repository. Then set:
|
|
|
|
```bash
|
|
export THOTH_ENABLE_SENSITIVITY_NER=1
|
|
export THT_SENSITIVITY_NER_MODEL_DIR=/absolute/path/to/gliner2-pii
|
|
export THT_SENSITIVITY_NER_THREADS=2
|
|
./scripts/run-stack.sh
|
|
```
|
|
|
|
For an operator-managed Compose invocation, include `deploy/compose.sensitivity-ner.yaml` after the
|
|
base and installation overlays. The core build argument `INSTALL_SENSITIVITY_NER=true` installs the
|
|
optional Python dependencies into their isolated virtualenv. The model mount is read-only. Values
|
|
above eight threads are rejected; start with two so classification cannot contend heavily with
|
|
other CPU workloads.
|
|
|
|
## Acceptance on a real database
|
|
|
|
Run the first evaluation in shadow mode: read the source with its existing read-only role, do not
|
|
save proposed flags, and report only aggregate counts, rule IDs, coverage, and timings. Never copy
|
|
matched values into test output. Use a separately approved, labeled Italian corpus to calculate
|
|
precision and recall; raw PSD values must remain inside the authorized environment.
|
|
|
|
Inside the configured core runtime, the non-mutating command is:
|
|
|
|
```bash
|
|
npm run sensitivity:shadow -- psd-clinical
|
|
```
|
|
|
|
It reads catalog metadata and source values but emits one aggregate JSON object with no database,
|
|
table, column, or source-value detail. It neither creates an analysis run nor updates a flag.
|
|
|
|
Enabling NER by default requires all of these gates:
|
|
|
|
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
|
|
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
|
|
3. the CPU benchmark stays within the configured NER allowance and does not use a GPU;
|
|
4. the labeled Italian evaluation meets thresholds approved by the product owner.
|
|
|
|
If a gate fails, leave NER disabled. The deterministic policy remains available and produces the
|
|
binary draft from its scan coverage; no content is sent to an internal or external LLM.
|
|
|
|
The first aggregate PSD shadow comparison is recorded in
|
|
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
|
|
local CPU runner, NER found additional entities but reduced total coverage under the superseded
|
|
global deadline, so the accepted setting remains disabled by default pending a new v2 benchmark.
|
|
The deterministic progressive PSD run is recorded in
|
|
[`2026-09-03-psd-progressive-sensitivity-shadow.md`](../reports/2026-09-03-psd-progressive-sensitivity-shadow.md):
|
|
it assessed all 2,275 columns with zero `unknown` decisions and left NER disabled.
|