Files
ThothII/docs/adr/0015-use-progressive-sampling-for-sensitive-columns.md
T

2.1 KiB

status
status
accepted

Use progressive sampling for sensitive columns

The sensitivity-v1 wall-clock policy produced too many unknown assessments: a global sixty-second deadline coupled the outcome of one column to table order, source latency, and optional NER cost. Those outcomes were not useful for description-generation gating, because they did not provide a usable draft Sensitive Data Flag.

sensitivity-v2 bounds database effort by inspected values rather than by one global clock. Tables proven to contain at most 1,000 rows are fully scanned. Larger tables are processed breadth-first in three passes: 300 values per unresolved column, 700 additional values to reach 1,000, then 2,000 additional values to reach 3,000 for unresolved text, JSON, and XML columns. One positive rule or NER finding is enough to stop later work for that column. At most two tables are scanned concurrently. Source queries contain at most 25 columns and each has a five-second statement timeout. An empty or timed-out randomized sample gets one sequential bounded retry; two timeouts fail the run.

A completed v2 analysis returns only sensitive or non_sensitive. Sampled no-match, empty, and all-null columns are proposed as non_sensitive, with coverage reported independently so the human reviewer can judge the strength of the proposal. Binary or otherwise uninspectable column types are proposed as sensitive. A source failure fails the analysis and returns no review; it is not converted into unknown. The administrator can still set either final value.

The HTTP operation has no global analysis deadline. It is canceled when the client disconnects or the backend restarts. Historic and interrupted run records retain the database field named unknown for compatibility, where it counts unprocessed columns rather than a v2 assessment.

This decision supersedes ADR-0014 only for scan effort, coverage semantics, and the assessment domain. ADR-0014 remains authoritative for the single local TypeScript decision point, the absence of generative LLMs, transient human-reviewed drafts, the 500-character rule, and optional CPU-only NER evidence.