2.1 KiB
status
| status |
|---|
| accepted |
Use progressive sampling for sensitive columns
The sensitivity-v1 wall-clock policy produced too many unknown assessments: a global
sixty-second deadline coupled the outcome of one column to table order, source latency, and optional
NER cost. Those outcomes were not useful for description-generation gating, because they did not
provide a usable draft Sensitive Data Flag.
sensitivity-v2 bounds database effort by inspected values rather than by one global clock. Tables
proven to contain at most 1,000 rows are fully scanned. Larger tables are processed breadth-first in
three passes: 300 values per unresolved column, 700 additional values to reach 1,000, then 2,000
additional values to reach 3,000 for unresolved text, JSON, and XML columns. One positive rule or
NER finding is enough to stop later work for that column. At most two tables are scanned
concurrently. Source queries contain at most 25 columns and each has a five-second statement timeout.
An empty or timed-out randomized sample gets one sequential bounded retry; two timeouts fail the run.
A completed v2 analysis returns only sensitive or non_sensitive. Sampled no-match, empty, and
all-null columns are proposed as non_sensitive, with coverage reported independently so the human
reviewer can judge the strength of the proposal. Binary or otherwise uninspectable column types are
proposed as sensitive. A source failure fails the analysis and returns no review; it is not
converted into unknown. The administrator can still set either final value.
The HTTP operation has no global analysis deadline. It is canceled when the client disconnects or
the backend restarts. Historic and interrupted run records retain the database field named
unknown for compatibility, where it counts unprocessed columns rather than a v2 assessment.
This decision supersedes ADR-0014 only for scan effort, coverage semantics, and the assessment domain. ADR-0014 remains authoritative for the single local TypeScript decision point, the absence of generative LLMs, transient human-reviewed drafts, the 500-character rule, and optional CPU-only NER evidence.