feat: sample sensitive columns progressively

This commit is contained in:
Codex
2026-09-03 10:25:05 +02:00
parent f114d0065a
commit 8e778b9edb
24 changed files with 1001 additions and 581 deletions
@@ -0,0 +1,49 @@
# PSD progressive sensitivity shadow evaluation
Date: 2026-09-03
This report records an aggregate, non-mutating evaluation of `sensitivity-v2` against the PSD
workspace. The configured connector accessed the source data warehouse with its read-only role.
The shadow command did not create an analysis run, update local catalog metadata, or save Sensitive
Data Flags. No database, table, column, source value, or matched span was emitted.
The final post-fix run used the local Docker CPU environment, without NER. It inspected all 2,275
catalog columns through the progressive 300, 1,000, and text-only 3,000-value policy.
| Sensitive | Non-sensitive | Unknown decisions | Analysis time |
| ---: | ---: | ---: | ---: |
| 343 | 1,932 | 0 | 50,082 ms |
Coverage was reported independently from the decision:
| Metadata decision | Complete scan | Sampled | No observed values |
| ---: | ---: | ---: | ---: |
| 39 | 0 | 2,128 | 108 |
The sampled no-match population comprised 1,337 columns ending after the 1,000-value target and
487 text-like columns ending after the 3,000-value target. Positive findings were:
| Rule | Columns |
| --- | ---: |
| `pii.phone_number` | 202 |
| `text.over_500_characters` | 52 |
| `metadata.health` | 34 |
| `health.clinical_term` | 25 |
| `pii.italian_vat` | 8 |
| `pii.email` | 7 |
| `metadata.direct_identifier` | 5 |
| `pii.uuid` | 5 |
| `financial.payment_card` | 3 |
| `pii.italian_fiscal_code` | 2 |
Three successful v2 diagnostic runs produced the same decisions and aggregate rule counts. Their
times ranged from 49,253 to 130,818 ms, showing that source load still affects latency even though
it no longer changes the outcome through a global deadline. An earlier run exposed an intermittent
randomized-query timeout. The adapter now retries that case once with a sequential bounded query;
a regression test covers the fallback, while two consecutive timeouts still fail the whole analysis
instead of creating `unknown` decisions.
This is a coverage and operational benchmark, not a precision/recall acceptance test. In
particular, the 202 phone-number findings and every other rule family still require human review or
a separately approved labeled corpus before their false-positive rate can be measured. NER remains
disabled by default.