feat: sample sensitive columns progressively

This commit is contained in:
Codex
2026-09-03 10:25:05 +02:00
parent f114d0065a
commit 8e778b9edb
24 changed files with 1001 additions and 581 deletions
@@ -1,5 +1,5 @@
---
status: accepted
status: superseded by ADR-0015
---
# Assess sensitive columns locally from source content
@@ -0,0 +1,33 @@
---
status: accepted
---
# Use progressive sampling for sensitive columns
The `sensitivity-v1` wall-clock policy produced too many `unknown` assessments: a global
sixty-second deadline coupled the outcome of one column to table order, source latency, and optional
NER cost. Those outcomes were not useful for description-generation gating, because they did not
provide a usable draft Sensitive Data Flag.
`sensitivity-v2` bounds database effort by inspected values rather than by one global clock. Tables
proven to contain at most 1,000 rows are fully scanned. Larger tables are processed breadth-first in
three passes: 300 values per unresolved column, 700 additional values to reach 1,000, then 2,000
additional values to reach 3,000 for unresolved text, JSON, and XML columns. One positive rule or
NER finding is enough to stop later work for that column. At most two tables are scanned
concurrently. Source queries contain at most 25 columns and each has a five-second statement timeout.
An empty or timed-out randomized sample gets one sequential bounded retry; two timeouts fail the run.
A completed v2 analysis returns only `sensitive` or `non_sensitive`. Sampled no-match, empty, and
all-null columns are proposed as `non_sensitive`, with coverage reported independently so the human
reviewer can judge the strength of the proposal. Binary or otherwise uninspectable column types are
proposed as `sensitive`. A source failure fails the analysis and returns no review; it is not
converted into `unknown`. The administrator can still set either final value.
The HTTP operation has no global analysis deadline. It is canceled when the client disconnects or
the backend restarts. Historic and interrupted run records retain the database field named
`unknown` for compatibility, where it counts unprocessed columns rather than a v2 assessment.
This decision supersedes ADR-0014 only for scan effort, coverage semantics, and the assessment
domain. ADR-0014 remains authoritative for the single local TypeScript decision point, the absence
of generative LLMs, transient human-reviewed drafts, the 500-character rule, and optional CPU-only
NER evidence.