feat: classify sensitive columns locally

This commit is contained in:
Codex
2026-09-03 02:11:13 +02:00
parent 7b87e95427
commit f114d0065a
57 changed files with 4038 additions and 1149 deletions
+20 -9
View File
@@ -98,6 +98,17 @@ column. The KPI strip reads installation-wide or selected-database aggregates fr
description history, and sensitive-field review/history use the production APIs in right-side
drawers rather than prototype fixtures; closing a history drawer does not stop its background run.
Sensitive-field review is now driven by the versioned local `sensitivity-v1` policy, not by a
catalog model. The backend reads selected source tables through read-only, database-specific
adapters and makes every `sensitive | non_sensitive | unknown` decision in the TypeScript
`SensitivityClassifier`. A single validated match protects the column; a full scan is limited to
five seconds per table before sampling and the whole request to sixty seconds. Draft assessments
remain transient until an administrator explicitly saves them. Optional GLiNER2 evidence is
CPU-only, offline, opt-in, and never replaces the deterministic decision point; see
`docs/operations/sensitivity-analysis.md`. The aggregate PSD shadow comparison kept NER disabled by
default because its extra findings did not offset the coverage lost to inference within the global
deadline; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`.
Physical membership, source
comments, column types/default/nullability/PK positions, and constraint-level ordered FK pairs are
projections of the external schema. They cannot be created, renamed, or structurally edited by
@@ -155,15 +166,15 @@ Semantic aliases, value descriptions, synonyms, and concepts remain deferred to
slices.
AI Description Generation uses the catalog's human-owned Sensitive Data Flag. The flag defaults to
`false`, including for newly synchronized columns. An administrator may request an AI proposal based
only on structural metadata for one selected database, selected tables, or selected columns. The
backend divides large scopes into deterministic model requests of at most ten columns, also bounded
by helper message size, and combines their results, but the proposal remains an unsaved draft until
the human reviews and saves it.
Each started suggestion attempt records a separate Sensitive Data Suggestion Run with aggregate
counters and safe ordered events. This operational history never stores per-column proposals,
prompts, raw model output, or provider diagnostics; reloading still discards an unsaved review
draft.
`false`, including for newly synchronized columns. An administrator may request a local sensitivity
analysis for one selected database, selected tables, or selected columns. One deterministic
TypeScript classifier combines metadata, bounded source-content rules, and optional CPU-only NER;
no generative model decides the result. Its `sensitive`, `non_sensitive`, or `unknown` assessments
remain an unsaved draft until the human reviews and saves any chosen flag changes, including a
downgrade to non-sensitive.
Each started analysis records a separate Sensitivity Analysis Run with aggregate counters and safe
ordered events. This operational history never stores per-column assessments, source values,
matched spans, prompts, or free-form diagnostics; reloading still discards an unsaved review draft.
For unprotected columns, up to five source rows and five representative non-null values may be sent
transiently to the configured model provider. Protected columns are omitted from source reads and
replaced in the prompt by deterministic plausible values derived only from their metadata. Existing