feat: sample sensitive columns progressively

This commit is contained in:
Codex
2026-09-03 10:25:05 +02:00
parent f114d0065a
commit 8e778b9edb
24 changed files with 1001 additions and 581 deletions
@@ -1,5 +1,5 @@
---
status: accepted
status: superseded by ADR-0015
---
# Assess sensitive columns locally from source content
@@ -0,0 +1,33 @@
---
status: accepted
---
# Use progressive sampling for sensitive columns
The `sensitivity-v1` wall-clock policy produced too many `unknown` assessments: a global
sixty-second deadline coupled the outcome of one column to table order, source latency, and optional
NER cost. Those outcomes were not useful for description-generation gating, because they did not
provide a usable draft Sensitive Data Flag.
`sensitivity-v2` bounds database effort by inspected values rather than by one global clock. Tables
proven to contain at most 1,000 rows are fully scanned. Larger tables are processed breadth-first in
three passes: 300 values per unresolved column, 700 additional values to reach 1,000, then 2,000
additional values to reach 3,000 for unresolved text, JSON, and XML columns. One positive rule or
NER finding is enough to stop later work for that column. At most two tables are scanned
concurrently. Source queries contain at most 25 columns and each has a five-second statement timeout.
An empty or timed-out randomized sample gets one sequential bounded retry; two timeouts fail the run.
A completed v2 analysis returns only `sensitive` or `non_sensitive`. Sampled no-match, empty, and
all-null columns are proposed as `non_sensitive`, with coverage reported independently so the human
reviewer can judge the strength of the proposal. Binary or otherwise uninspectable column types are
proposed as `sensitive`. A source failure fails the analysis and returns no review; it is not
converted into `unknown`. The administrator can still set either final value.
The HTTP operation has no global analysis deadline. It is canceled when the client disconnects or
the backend restarts. Historic and interrupted run records retain the database field named
`unknown` for compatibility, where it counts unprocessed columns rather than a v2 assessment.
This decision supersedes ADR-0014 only for scan effort, coverage semantics, and the assessment
domain. ADR-0014 remains authoritative for the single local TypeScript decision point, the absence
of generative LLMs, transient human-reviewed drafts, the 500-character rule, and optional CPU-only
NER evidence.
+9 -8
View File
@@ -107,16 +107,17 @@ explicitly unlocked; it never resumes automatically.
Sensitivity analysis is a synchronous administrative request and does not use the installation
model catalog. Database-specific adapters stream bounded normalized values from read-only source
connections; the TypeScript `SensitivityClassifier` is the single decision point for
`sensitive | non_sensitive | unknown`. Deterministic rules run first. A complete scan is attempted
for at most five seconds per table, then the adapter samples within the sixty-second request budget.
An optional offline GLiNER2 worker may add NER evidence on CPU for unresolved short text, but it
cannot make or persist the decision itself.
`sensitive | non_sensitive`. Deterministic rules run first. Tables up to 1,000 rows are fully
scanned; larger tables use breadth-first targets of 300, 1,000, and 3,000 values, with the last pass
limited to text-like columns. Source queries have five-second limits, but the request has no global
analysis deadline. An optional offline GLiNER2 worker may add NER evidence on CPU for unresolved
short text, but it cannot make or persist the decision itself.
Each attempt has its own durable run and ordered sanitized events, separate from Description
Generation because its lifecycle and counters differ. The run records the local policy version,
coverage aggregates, and sanitized rule identifiers. Proposed flags, source values, NER spans, and
worker diagnostics remain transient. Only an explicit administrator save changes the human-owned
Sensitive Data Flag.
Generation because its lifecycle and counters differ. The run records the local policy version and
aggregate decision counts. Coverage, rule identifiers, proposed flags, source values, NER spans,
and worker diagnostics remain transient. Only an explicit administrator save changes the
human-owned Sensitive Data Flag.
## Main backend classes
+34 -18
View File
@@ -7,7 +7,7 @@ may set either value, including overriding a `sensitive` proposal.
## Default policy
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v1`
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v2`
policy combines:
- normalized column-name rules for direct identifiers, credentials, and health data;
@@ -18,22 +18,35 @@ policy combines:
- a conservative length rule: any observed textual value longer than 500 characters makes the
entire column sensitive.
One decisive value is enough to classify the column as `sensitive`. A complete scan with no match
may classify it as `non_sensitive`. Empty, all-null, binary/uninspectable, interrupted, and sampled
no-match columns are `unknown`; an `unknown` draft preserves the current human flag.
One decisive value is enough to classify the column as `sensitive` and removes it from subsequent
passes. Binary or otherwise uninspectable column types are also proposed as `sensitive`, because
their contents cannot be cleared by the textual rules. A completed analysis has only two draft
outcomes: `sensitive` and `non_sensitive`. Empty or all-null columns are `non_sensitive` with
`no_values` coverage; a sampled column with no match is `non_sensitive` with explicit sampled
coverage. The administrator remains free to reverse either proposal before saving it.
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
persist source values. A full scan gets five seconds per table. If it cannot finish, the adapter uses
a bounded repeatable sample within the sixty-second request deadline. PostgreSQL-wire reads run in a
read-only transaction and always end with rollback. A REST scan can claim complete coverage only
when it finishes in one request; multi-request pagination has no shared source transaction and is
therefore conservatively reported as sampled.
persist source values. Tables proven to contain at most 1,000 rows are fully scanned. Larger tables
are processed breadth-first so every table gets the cheapest pass before any table gets a deeper
one:
The HTTP operation stops waiting at sixty seconds. The same expiring signal is checked before and
after catalog selection, source access, progress writes, and every table. If it expires after a run
has been created, that run is finalized as `interrupted` and all not-decisively-processed columns
are counted as `unknown`; no review payload is returned from the timed-out request.
1. inspect up to 300 non-null values per unresolved column;
2. inspect up to 700 additional values, reaching a 1,000-value target;
3. for unresolved text, JSON, and XML columns only, inspect up to 2,000 additional values, reaching
a 3,000-value target.
At most two tables are scanned concurrently, and the database adapter groups at most 25 columns in
one source query. Each probe or value query has a five-second statement timeout; PostgreSQL-wire
reads run in a read-only transaction and always end with rollback. Sampling is bounded and
repeatable for a policy version. If a randomized sample is empty or reaches its query timeout, the
adapter tries one sequential bounded sample; if that also times out, the source error fails the run
and returns no review instead of manufacturing `unknown` decisions.
There is no global sixty-second analysis deadline. Work is bounded by sample counts, per-query
timeouts, and early column exits. The operation is interrupted only when its request connection is
aborted or the backend restarts. Historical or interrupted run counters named `unknown` represent
columns that were not processed; `unknown` is not a `sensitivity-v2` column assessment.
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
@@ -116,13 +129,16 @@ Enabling NER by default requires all of these gates:
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
3. the CPU benchmark stays within the configured deadlines and does not use a GPU;
3. the CPU benchmark stays within the configured NER allowance and does not use a GPU;
4. the labeled Italian evaluation meets thresholds approved by the product owner.
If a gate fails, leave NER disabled. The deterministic policy remains available and unresolved
columns remain `unknown` rather than being sent to an internal or external LLM.
If a gate fails, leave NER disabled. The deterministic policy remains available and produces the
binary draft from its scan coverage; no content is sent to an internal or external LLM.
The first aggregate PSD shadow comparison is recorded in
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
local CPU runner, NER found additional entities but reduced total coverage inside the 60-second
deadline, so the accepted setting remains disabled by default.
local CPU runner, NER found additional entities but reduced total coverage under the superseded
global deadline, so the accepted setting remains disabled by default pending a new v2 benchmark.
The deterministic progressive PSD run is recorded in
[`2026-09-03-psd-progressive-sensitivity-shadow.md`](../reports/2026-09-03-psd-progressive-sensitivity-shadow.md):
it assessed all 2,275 columns with zero `unknown` decisions and left NER disabled.
@@ -0,0 +1,49 @@
# PSD progressive sensitivity shadow evaluation
Date: 2026-09-03
This report records an aggregate, non-mutating evaluation of `sensitivity-v2` against the PSD
workspace. The configured connector accessed the source data warehouse with its read-only role.
The shadow command did not create an analysis run, update local catalog metadata, or save Sensitive
Data Flags. No database, table, column, source value, or matched span was emitted.
The final post-fix run used the local Docker CPU environment, without NER. It inspected all 2,275
catalog columns through the progressive 300, 1,000, and text-only 3,000-value policy.
| Sensitive | Non-sensitive | Unknown decisions | Analysis time |
| ---: | ---: | ---: | ---: |
| 343 | 1,932 | 0 | 50,082 ms |
Coverage was reported independently from the decision:
| Metadata decision | Complete scan | Sampled | No observed values |
| ---: | ---: | ---: | ---: |
| 39 | 0 | 2,128 | 108 |
The sampled no-match population comprised 1,337 columns ending after the 1,000-value target and
487 text-like columns ending after the 3,000-value target. Positive findings were:
| Rule | Columns |
| --- | ---: |
| `pii.phone_number` | 202 |
| `text.over_500_characters` | 52 |
| `metadata.health` | 34 |
| `health.clinical_term` | 25 |
| `pii.italian_vat` | 8 |
| `pii.email` | 7 |
| `metadata.direct_identifier` | 5 |
| `pii.uuid` | 5 |
| `financial.payment_card` | 3 |
| `pii.italian_fiscal_code` | 2 |
Three successful v2 diagnostic runs produced the same decisions and aggregate rule counts. Their
times ranged from 49,253 to 130,818 ms, showing that source load still affects latency even though
it no longer changes the outcome through a global deadline. An earlier run exposed an intermittent
randomized-query timeout. The adapter now retries that case once with a sequential bounded query;
a regression test covers the fallback, while two consecutive timeouts still fail the whole analysis
instead of creating `unknown` decisions.
This is a coverage and operational benchmark, not a precision/recall acceptance test. In
particular, the 202 phone-number findings and every other rule family still require human review or
a separately approved labeled corpus before their false-positive rate can be measured. NER remains
disabled by default.