feat: complete catalog sensitivity enhancements

This commit is contained in:
Codex
2026-09-04 15:11:18 +02:00
parent b891246664
commit 7b1d69a65b
72 changed files with 33866 additions and 358 deletions
+6 -2
View File
@@ -7,9 +7,13 @@ may set either value, including overriding a `sensitive` proposal.
## Default policy
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v2`
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v4`
policy combines:
- a structural exclusion for declared `bigint` primary-key columns and undeclared `bigint`
columns following the exact `pk` naming convention; their values are non-informative identifiers
and are therefore not inspected as possible sensitive content. The evidence distinguishes
declared constraints from convention-based inference;
- normalized column-name rules for direct identifiers, credentials, and health data;
- validated content rules for email, Italian fiscal code and VAT, passport, identity-card and
driving-licence identifiers, phone numbers, IBAN/BIC, payment-card checksums, IP/MAC addresses,
@@ -46,7 +50,7 @@ and returns no review instead of manufacturing `unknown` decisions.
There is no global sixty-second analysis deadline. Work is bounded by sample counts, per-query
timeouts, and early column exits. The operation is interrupted only when its request connection is
aborted or the backend restarts. Historical or interrupted run counters named `unknown` represent
columns that were not processed; `unknown` is not a `sensitivity-v2` column assessment.
columns that were not processed; `unknown` is not a `sensitivity-v4` column assessment.
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.