feat: sample sensitive columns progressively
This commit is contained in:
@@ -7,7 +7,7 @@ may set either value, including overriding a `sensitive` proposal.
|
||||
|
||||
## Default policy
|
||||
|
||||
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v1`
|
||||
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v2`
|
||||
policy combines:
|
||||
|
||||
- normalized column-name rules for direct identifiers, credentials, and health data;
|
||||
@@ -18,22 +18,35 @@ policy combines:
|
||||
- a conservative length rule: any observed textual value longer than 500 characters makes the
|
||||
entire column sensitive.
|
||||
|
||||
One decisive value is enough to classify the column as `sensitive`. A complete scan with no match
|
||||
may classify it as `non_sensitive`. Empty, all-null, binary/uninspectable, interrupted, and sampled
|
||||
no-match columns are `unknown`; an `unknown` draft preserves the current human flag.
|
||||
One decisive value is enough to classify the column as `sensitive` and removes it from subsequent
|
||||
passes. Binary or otherwise uninspectable column types are also proposed as `sensitive`, because
|
||||
their contents cannot be cleared by the textual rules. A completed analysis has only two draft
|
||||
outcomes: `sensitive` and `non_sensitive`. Empty or all-null columns are `non_sensitive` with
|
||||
`no_values` coverage; a sampled column with no match is `non_sensitive` with explicit sampled
|
||||
coverage. The administrator remains free to reverse either proposal before saving it.
|
||||
|
||||
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
|
||||
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
|
||||
persist source values. A full scan gets five seconds per table. If it cannot finish, the adapter uses
|
||||
a bounded repeatable sample within the sixty-second request deadline. PostgreSQL-wire reads run in a
|
||||
read-only transaction and always end with rollback. A REST scan can claim complete coverage only
|
||||
when it finishes in one request; multi-request pagination has no shared source transaction and is
|
||||
therefore conservatively reported as sampled.
|
||||
persist source values. Tables proven to contain at most 1,000 rows are fully scanned. Larger tables
|
||||
are processed breadth-first so every table gets the cheapest pass before any table gets a deeper
|
||||
one:
|
||||
|
||||
The HTTP operation stops waiting at sixty seconds. The same expiring signal is checked before and
|
||||
after catalog selection, source access, progress writes, and every table. If it expires after a run
|
||||
has been created, that run is finalized as `interrupted` and all not-decisively-processed columns
|
||||
are counted as `unknown`; no review payload is returned from the timed-out request.
|
||||
1. inspect up to 300 non-null values per unresolved column;
|
||||
2. inspect up to 700 additional values, reaching a 1,000-value target;
|
||||
3. for unresolved text, JSON, and XML columns only, inspect up to 2,000 additional values, reaching
|
||||
a 3,000-value target.
|
||||
|
||||
At most two tables are scanned concurrently, and the database adapter groups at most 25 columns in
|
||||
one source query. Each probe or value query has a five-second statement timeout; PostgreSQL-wire
|
||||
reads run in a read-only transaction and always end with rollback. Sampling is bounded and
|
||||
repeatable for a policy version. If a randomized sample is empty or reaches its query timeout, the
|
||||
adapter tries one sequential bounded sample; if that also times out, the source error fails the run
|
||||
and returns no review instead of manufacturing `unknown` decisions.
|
||||
|
||||
There is no global sixty-second analysis deadline. Work is bounded by sample counts, per-query
|
||||
timeouts, and early column exits. The operation is interrupted only when its request connection is
|
||||
aborted or the backend restarts. Historical or interrupted run counters named `unknown` represent
|
||||
columns that were not processed; `unknown` is not a `sensitivity-v2` column assessment.
|
||||
|
||||
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
|
||||
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
|
||||
@@ -116,13 +129,16 @@ Enabling NER by default requires all of these gates:
|
||||
|
||||
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
|
||||
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
|
||||
3. the CPU benchmark stays within the configured deadlines and does not use a GPU;
|
||||
3. the CPU benchmark stays within the configured NER allowance and does not use a GPU;
|
||||
4. the labeled Italian evaluation meets thresholds approved by the product owner.
|
||||
|
||||
If a gate fails, leave NER disabled. The deterministic policy remains available and unresolved
|
||||
columns remain `unknown` rather than being sent to an internal or external LLM.
|
||||
If a gate fails, leave NER disabled. The deterministic policy remains available and produces the
|
||||
binary draft from its scan coverage; no content is sent to an internal or external LLM.
|
||||
|
||||
The first aggregate PSD shadow comparison is recorded in
|
||||
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
|
||||
local CPU runner, NER found additional entities but reduced total coverage inside the 60-second
|
||||
deadline, so the accepted setting remains disabled by default.
|
||||
local CPU runner, NER found additional entities but reduced total coverage under the superseded
|
||||
global deadline, so the accepted setting remains disabled by default pending a new v2 benchmark.
|
||||
The deterministic progressive PSD run is recorded in
|
||||
[`2026-09-03-psd-progressive-sensitivity-shadow.md`](../reports/2026-09-03-psd-progressive-sensitivity-shadow.md):
|
||||
it assessed all 2,275 columns with zero `unknown` decisions and left NER disabled.
|
||||
|
||||
Reference in New Issue
Block a user