feat: sample sensitive columns progressively
This commit is contained in:
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: accepted
|
||||
status: superseded by ADR-0015
|
||||
---
|
||||
|
||||
# Assess sensitive columns locally from source content
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
---
|
||||
status: accepted
|
||||
---
|
||||
|
||||
# Use progressive sampling for sensitive columns
|
||||
|
||||
The `sensitivity-v1` wall-clock policy produced too many `unknown` assessments: a global
|
||||
sixty-second deadline coupled the outcome of one column to table order, source latency, and optional
|
||||
NER cost. Those outcomes were not useful for description-generation gating, because they did not
|
||||
provide a usable draft Sensitive Data Flag.
|
||||
|
||||
`sensitivity-v2` bounds database effort by inspected values rather than by one global clock. Tables
|
||||
proven to contain at most 1,000 rows are fully scanned. Larger tables are processed breadth-first in
|
||||
three passes: 300 values per unresolved column, 700 additional values to reach 1,000, then 2,000
|
||||
additional values to reach 3,000 for unresolved text, JSON, and XML columns. One positive rule or
|
||||
NER finding is enough to stop later work for that column. At most two tables are scanned
|
||||
concurrently. Source queries contain at most 25 columns and each has a five-second statement timeout.
|
||||
An empty or timed-out randomized sample gets one sequential bounded retry; two timeouts fail the run.
|
||||
|
||||
A completed v2 analysis returns only `sensitive` or `non_sensitive`. Sampled no-match, empty, and
|
||||
all-null columns are proposed as `non_sensitive`, with coverage reported independently so the human
|
||||
reviewer can judge the strength of the proposal. Binary or otherwise uninspectable column types are
|
||||
proposed as `sensitive`. A source failure fails the analysis and returns no review; it is not
|
||||
converted into `unknown`. The administrator can still set either final value.
|
||||
|
||||
The HTTP operation has no global analysis deadline. It is canceled when the client disconnects or
|
||||
the backend restarts. Historic and interrupted run records retain the database field named
|
||||
`unknown` for compatibility, where it counts unprocessed columns rather than a v2 assessment.
|
||||
|
||||
This decision supersedes ADR-0014 only for scan effort, coverage semantics, and the assessment
|
||||
domain. ADR-0014 remains authoritative for the single local TypeScript decision point, the absence
|
||||
of generative LLMs, transient human-reviewed drafts, the 500-character rule, and optional CPU-only
|
||||
NER evidence.
|
||||
@@ -107,16 +107,17 @@ explicitly unlocked; it never resumes automatically.
|
||||
Sensitivity analysis is a synchronous administrative request and does not use the installation
|
||||
model catalog. Database-specific adapters stream bounded normalized values from read-only source
|
||||
connections; the TypeScript `SensitivityClassifier` is the single decision point for
|
||||
`sensitive | non_sensitive | unknown`. Deterministic rules run first. A complete scan is attempted
|
||||
for at most five seconds per table, then the adapter samples within the sixty-second request budget.
|
||||
An optional offline GLiNER2 worker may add NER evidence on CPU for unresolved short text, but it
|
||||
cannot make or persist the decision itself.
|
||||
`sensitive | non_sensitive`. Deterministic rules run first. Tables up to 1,000 rows are fully
|
||||
scanned; larger tables use breadth-first targets of 300, 1,000, and 3,000 values, with the last pass
|
||||
limited to text-like columns. Source queries have five-second limits, but the request has no global
|
||||
analysis deadline. An optional offline GLiNER2 worker may add NER evidence on CPU for unresolved
|
||||
short text, but it cannot make or persist the decision itself.
|
||||
|
||||
Each attempt has its own durable run and ordered sanitized events, separate from Description
|
||||
Generation because its lifecycle and counters differ. The run records the local policy version,
|
||||
coverage aggregates, and sanitized rule identifiers. Proposed flags, source values, NER spans, and
|
||||
worker diagnostics remain transient. Only an explicit administrator save changes the human-owned
|
||||
Sensitive Data Flag.
|
||||
Generation because its lifecycle and counters differ. The run records the local policy version and
|
||||
aggregate decision counts. Coverage, rule identifiers, proposed flags, source values, NER spans,
|
||||
and worker diagnostics remain transient. Only an explicit administrator save changes the
|
||||
human-owned Sensitive Data Flag.
|
||||
|
||||
## Main backend classes
|
||||
|
||||
|
||||
@@ -7,7 +7,7 @@ may set either value, including overriding a `sensitive` proposal.
|
||||
|
||||
## Default policy
|
||||
|
||||
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v1`
|
||||
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v2`
|
||||
policy combines:
|
||||
|
||||
- normalized column-name rules for direct identifiers, credentials, and health data;
|
||||
@@ -18,22 +18,35 @@ policy combines:
|
||||
- a conservative length rule: any observed textual value longer than 500 characters makes the
|
||||
entire column sensitive.
|
||||
|
||||
One decisive value is enough to classify the column as `sensitive`. A complete scan with no match
|
||||
may classify it as `non_sensitive`. Empty, all-null, binary/uninspectable, interrupted, and sampled
|
||||
no-match columns are `unknown`; an `unknown` draft preserves the current human flag.
|
||||
One decisive value is enough to classify the column as `sensitive` and removes it from subsequent
|
||||
passes. Binary or otherwise uninspectable column types are also proposed as `sensitive`, because
|
||||
their contents cannot be cleared by the textual rules. A completed analysis has only two draft
|
||||
outcomes: `sensitive` and `non_sensitive`. Empty or all-null columns are `non_sensitive` with
|
||||
`no_values` coverage; a sampled column with no match is `non_sensitive` with explicit sampled
|
||||
coverage. The administrator remains free to reverse either proposal before saving it.
|
||||
|
||||
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
|
||||
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
|
||||
persist source values. A full scan gets five seconds per table. If it cannot finish, the adapter uses
|
||||
a bounded repeatable sample within the sixty-second request deadline. PostgreSQL-wire reads run in a
|
||||
read-only transaction and always end with rollback. A REST scan can claim complete coverage only
|
||||
when it finishes in one request; multi-request pagination has no shared source transaction and is
|
||||
therefore conservatively reported as sampled.
|
||||
persist source values. Tables proven to contain at most 1,000 rows are fully scanned. Larger tables
|
||||
are processed breadth-first so every table gets the cheapest pass before any table gets a deeper
|
||||
one:
|
||||
|
||||
The HTTP operation stops waiting at sixty seconds. The same expiring signal is checked before and
|
||||
after catalog selection, source access, progress writes, and every table. If it expires after a run
|
||||
has been created, that run is finalized as `interrupted` and all not-decisively-processed columns
|
||||
are counted as `unknown`; no review payload is returned from the timed-out request.
|
||||
1. inspect up to 300 non-null values per unresolved column;
|
||||
2. inspect up to 700 additional values, reaching a 1,000-value target;
|
||||
3. for unresolved text, JSON, and XML columns only, inspect up to 2,000 additional values, reaching
|
||||
a 3,000-value target.
|
||||
|
||||
At most two tables are scanned concurrently, and the database adapter groups at most 25 columns in
|
||||
one source query. Each probe or value query has a five-second statement timeout; PostgreSQL-wire
|
||||
reads run in a read-only transaction and always end with rollback. Sampling is bounded and
|
||||
repeatable for a policy version. If a randomized sample is empty or reaches its query timeout, the
|
||||
adapter tries one sequential bounded sample; if that also times out, the source error fails the run
|
||||
and returns no review instead of manufacturing `unknown` decisions.
|
||||
|
||||
There is no global sixty-second analysis deadline. Work is bounded by sample counts, per-query
|
||||
timeouts, and early column exits. The operation is interrupted only when its request connection is
|
||||
aborted or the backend restarts. Historical or interrupted run counters named `unknown` represent
|
||||
columns that were not processed; `unknown` is not a `sensitivity-v2` column assessment.
|
||||
|
||||
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
|
||||
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
|
||||
@@ -116,13 +129,16 @@ Enabling NER by default requires all of these gates:
|
||||
|
||||
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
|
||||
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
|
||||
3. the CPU benchmark stays within the configured deadlines and does not use a GPU;
|
||||
3. the CPU benchmark stays within the configured NER allowance and does not use a GPU;
|
||||
4. the labeled Italian evaluation meets thresholds approved by the product owner.
|
||||
|
||||
If a gate fails, leave NER disabled. The deterministic policy remains available and unresolved
|
||||
columns remain `unknown` rather than being sent to an internal or external LLM.
|
||||
If a gate fails, leave NER disabled. The deterministic policy remains available and produces the
|
||||
binary draft from its scan coverage; no content is sent to an internal or external LLM.
|
||||
|
||||
The first aggregate PSD shadow comparison is recorded in
|
||||
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
|
||||
local CPU runner, NER found additional entities but reduced total coverage inside the 60-second
|
||||
deadline, so the accepted setting remains disabled by default.
|
||||
local CPU runner, NER found additional entities but reduced total coverage under the superseded
|
||||
global deadline, so the accepted setting remains disabled by default pending a new v2 benchmark.
|
||||
The deterministic progressive PSD run is recorded in
|
||||
[`2026-09-03-psd-progressive-sensitivity-shadow.md`](../reports/2026-09-03-psd-progressive-sensitivity-shadow.md):
|
||||
it assessed all 2,275 columns with zero `unknown` decisions and left NER disabled.
|
||||
|
||||
@@ -0,0 +1,49 @@
|
||||
# PSD progressive sensitivity shadow evaluation
|
||||
|
||||
Date: 2026-09-03
|
||||
|
||||
This report records an aggregate, non-mutating evaluation of `sensitivity-v2` against the PSD
|
||||
workspace. The configured connector accessed the source data warehouse with its read-only role.
|
||||
The shadow command did not create an analysis run, update local catalog metadata, or save Sensitive
|
||||
Data Flags. No database, table, column, source value, or matched span was emitted.
|
||||
|
||||
The final post-fix run used the local Docker CPU environment, without NER. It inspected all 2,275
|
||||
catalog columns through the progressive 300, 1,000, and text-only 3,000-value policy.
|
||||
|
||||
| Sensitive | Non-sensitive | Unknown decisions | Analysis time |
|
||||
| ---: | ---: | ---: | ---: |
|
||||
| 343 | 1,932 | 0 | 50,082 ms |
|
||||
|
||||
Coverage was reported independently from the decision:
|
||||
|
||||
| Metadata decision | Complete scan | Sampled | No observed values |
|
||||
| ---: | ---: | ---: | ---: |
|
||||
| 39 | 0 | 2,128 | 108 |
|
||||
|
||||
The sampled no-match population comprised 1,337 columns ending after the 1,000-value target and
|
||||
487 text-like columns ending after the 3,000-value target. Positive findings were:
|
||||
|
||||
| Rule | Columns |
|
||||
| --- | ---: |
|
||||
| `pii.phone_number` | 202 |
|
||||
| `text.over_500_characters` | 52 |
|
||||
| `metadata.health` | 34 |
|
||||
| `health.clinical_term` | 25 |
|
||||
| `pii.italian_vat` | 8 |
|
||||
| `pii.email` | 7 |
|
||||
| `metadata.direct_identifier` | 5 |
|
||||
| `pii.uuid` | 5 |
|
||||
| `financial.payment_card` | 3 |
|
||||
| `pii.italian_fiscal_code` | 2 |
|
||||
|
||||
Three successful v2 diagnostic runs produced the same decisions and aggregate rule counts. Their
|
||||
times ranged from 49,253 to 130,818 ms, showing that source load still affects latency even though
|
||||
it no longer changes the outcome through a global deadline. An earlier run exposed an intermittent
|
||||
randomized-query timeout. The adapter now retries that case once with a sequential bounded query;
|
||||
a regression test covers the fallback, while two consecutive timeouts still fail the whole analysis
|
||||
instead of creating `unknown` decisions.
|
||||
|
||||
This is a coverage and operational benchmark, not a precision/recall acceptance test. In
|
||||
particular, the 202 phone-number findings and every other rule family still require human review or
|
||||
a separately approved labeled corpus before their false-positive rate can be measured. NER remains
|
||||
disabled by default.
|
||||
Reference in New Issue
Block a user