Files
ThothII/docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md
2026-09-03 10:59:09 +02:00

64 lines
3.3 KiB
Markdown

# PSD progressive sensitivity shadow evaluation
Date: 2026-09-03
This report records an aggregate, non-mutating evaluation of `sensitivity-v2` against the PSD
workspace. The configured connector accessed the source data warehouse with its read-only role.
The shadow command did not create an analysis run, update local catalog metadata, or save Sensitive
Data Flags. No database, table, column, source value, or matched span was emitted.
The final post-fix runs used the local Docker CPU environment. They inspected all 2,275 catalog
columns through the progressive 300, 1,000, and text-only 3,000-value policy. The second profile
enabled the checksummed GLiNER2 artifact on CPU with two threads, at most two candidates per table,
and the shared ten-second inference allowance.
| Profile | Sensitive | Non-sensitive | Unknown decisions | NER findings | Warm analysis time | Cold wall time |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Deterministic policy | 343 | 1,932 | 0 | 0 | 50,082 ms | not measured |
| CPU NER | 361 | 1,914 | 0 | 18 | 61,349 ms | about 69,750 ms |
The warm analysis measurement starts after the persistent worker has loaded the model, matching
ordinary backend operation. The cold wall measurement also includes one-shot container startup and
model warm-up, which added about 8.4 seconds in this environment. CPU NER added 11,267 ms, or about
22.5%, to the comparable warm analysis and proposed 18 additional sensitive columns.
Coverage was reported independently from the decision:
| Metadata decision | Complete scan | Sampled | No observed values |
| ---: | ---: | ---: | ---: |
| 39 | 0 | 2,128 | 108 |
Without NER, the sampled no-match population comprised 1,337 columns ending after the 1,000-value
target and 487 text-like columns ending after the 3,000-value target. With NER, 18 of the former
received `ner.entity` evidence. Deterministic positive findings were:
| Rule | Columns |
| --- | ---: |
| `pii.phone_number` | 202 |
| `text.over_500_characters` | 52 |
| `metadata.health` | 34 |
| `health.clinical_term` | 25 |
| `pii.italian_vat` | 8 |
| `pii.email` | 7 |
| `metadata.direct_identifier` | 5 |
| `pii.uuid` | 5 |
| `financial.payment_card` | 3 |
| `pii.italian_fiscal_code` | 2 |
The CPU NER profile added `ner.entity` evidence for 18 columns. The model checksum manifest
contained eight entries and all eight verified before the run. The worker used the CPU-only image;
no GPU was exposed to the detector.
Three successful v2 diagnostic runs produced the same decisions and aggregate rule counts. Their
times ranged from 49,253 to 130,818 ms, showing that source load still affects latency even though
it no longer changes the outcome through a global deadline. An earlier run exposed an intermittent
randomized-query timeout. The adapter now retries that case once with a sequential bounded query;
a regression test covers the fallback, while two consecutive timeouts still fail the whole analysis
instead of creating `unknown` decisions.
This is a coverage and operational benchmark, not a precision/recall acceptance test. In
particular, the 202 phone-number findings, the 18 NER findings, and every other rule family still
require human review or a separately approved labeled corpus before precision and recall can be
measured. The v2 timing now supports operational planning, but does not by itself justify enabling
NER by default.