feat: sample sensitive columns progressively

This commit is contained in:
Codex
2026-09-03 10:25:05 +02:00
parent f114d0065a
commit 8e778b9edb
24 changed files with 1001 additions and 581 deletions
+13 -12
View File
@@ -98,16 +98,17 @@ column. The KPI strip reads installation-wide or selected-database aggregates fr
description history, and sensitive-field review/history use the production APIs in right-side
drawers rather than prototype fixtures; closing a history drawer does not stop its background run.
Sensitive-field review is now driven by the versioned local `sensitivity-v1` policy, not by a
Sensitive-field review is now driven by the versioned local `sensitivity-v2` policy, not by a
catalog model. The backend reads selected source tables through read-only, database-specific
adapters and makes every `sensitive | non_sensitive | unknown` decision in the TypeScript
`SensitivityClassifier`. A single validated match protects the column; a full scan is limited to
five seconds per table before sampling and the whole request to sixty seconds. Draft assessments
remain transient until an administrator explicitly saves them. Optional GLiNER2 evidence is
CPU-only, offline, opt-in, and never replaces the deterministic decision point; see
`docs/operations/sensitivity-analysis.md`. The aggregate PSD shadow comparison kept NER disabled by
default because its extra findings did not offset the coverage lost to inference within the global
deadline; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`.
adapters and makes every `sensitive | non_sensitive` draft decision in the TypeScript
`SensitivityClassifier`. A single validated match protects the column. Tables up to 1,000 rows are
fully scanned; larger tables use breadth-first 300, 1,000, and text-only 3,000-value targets, with a
five-second limit per source query and no global request deadline. Source failures fail the run
instead of yielding `unknown`; coverage remains visible separately from the proposal. Draft
assessments remain transient until an administrator explicitly saves them. Optional GLiNER2
evidence is CPU-only, offline, opt-in, and never replaces the deterministic decision point; see
`docs/operations/sensitivity-analysis.md`. The earlier v1 PSD shadow comparison kept NER disabled by
default; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`. A v2 PSD benchmark is still due.
Physical membership, source
comments, column types/default/nullability/PK positions, and constraint-level ordered FK pairs are
@@ -169,9 +170,9 @@ AI Description Generation uses the catalog's human-owned Sensitive Data Flag. Th
`false`, including for newly synchronized columns. An administrator may request a local sensitivity
analysis for one selected database, selected tables, or selected columns. One deterministic
TypeScript classifier combines metadata, bounded source-content rules, and optional CPU-only NER;
no generative model decides the result. Its `sensitive`, `non_sensitive`, or `unknown` assessments
remain an unsaved draft until the human reviews and saves any chosen flag changes, including a
downgrade to non-sensitive.
no generative model decides the result. Its `sensitive` or `non_sensitive` assessments remain an
unsaved draft until the human reviews and saves any chosen flag changes, including a downgrade to
non-sensitive. Coverage is reported separately; interrupted history may count unprocessed columns.
Each started analysis records a separate Sensitivity Analysis Run with aggregate counters and safe
ordered events. This operational history never stores per-column assessments, source values,
matched spans, prompts, or free-form diagnostics; reloading still discards an unsaved review draft.