From b8912466643c049be53509affea264faaf1ce76c Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 3 Sep 2026 10:59:09 +0200 Subject: [PATCH] docs: add PSD CPU NER benchmark --- PROJECT_STATE.md | 4 ++- docs/operations/sensitivity-analysis.md | 6 ++-- ...9-03-psd-progressive-sensitivity-shadow.md | 34 +++++++++++++------ 3 files changed, 31 insertions(+), 13 deletions(-) diff --git a/PROJECT_STATE.md b/PROJECT_STATE.md index 658da211..5560f398 100644 --- a/PROJECT_STATE.md +++ b/PROJECT_STATE.md @@ -108,7 +108,9 @@ instead of yielding `unknown`; coverage remains visible separately from the prop assessments remain transient until an administrator explicitly saves them. Optional GLiNER2 evidence is CPU-only, offline, opt-in, and never replaces the deterministic decision point; see `docs/operations/sensitivity-analysis.md`. The earlier v1 PSD shadow comparison kept NER disabled by -default; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`. A v2 PSD benchmark is still due. +default; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`. The v2 comparison completed all +2,275 columns: CPU NER added 18 sensitive proposals and increased warm runtime from 50.1 to 61.3 +seconds; see `docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md`. Physical membership, source comments, column types/default/nullability/PK positions, and constraint-level ordered FK pairs are diff --git a/docs/operations/sensitivity-analysis.md b/docs/operations/sensitivity-analysis.md index 3c6a95f1..ebf7f295 100644 --- a/docs/operations/sensitivity-analysis.md +++ b/docs/operations/sensitivity-analysis.md @@ -138,7 +138,9 @@ binary draft from its scan coverage; no content is sent to an internal or extern The first aggregate PSD shadow comparison is recorded in [`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the local CPU runner, NER found additional entities but reduced total coverage under the superseded -global deadline, so the accepted setting remains disabled by default pending a new v2 benchmark. +global deadline. The v2 benchmark removed that confounder: CPU NER added 18 sensitive proposals and +increased the warm analysis time from 50.1 to 61.3 seconds. It remains opt-in until a labeled Italian +evaluation establishes that the additional findings justify their false-positive rate and cost. The deterministic progressive PSD run is recorded in [`2026-09-03-psd-progressive-sensitivity-shadow.md`](../reports/2026-09-03-psd-progressive-sensitivity-shadow.md): -it assessed all 2,275 columns with zero `unknown` decisions and left NER disabled. +both deterministic and CPU-NER profiles assessed all 2,275 columns with zero `unknown` decisions. diff --git a/docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md b/docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md index 34890527..70e389cc 100644 --- a/docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md +++ b/docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md @@ -7,12 +7,20 @@ workspace. The configured connector accessed the source data warehouse with its The shadow command did not create an analysis run, update local catalog metadata, or save Sensitive Data Flags. No database, table, column, source value, or matched span was emitted. -The final post-fix run used the local Docker CPU environment, without NER. It inspected all 2,275 -catalog columns through the progressive 300, 1,000, and text-only 3,000-value policy. +The final post-fix runs used the local Docker CPU environment. They inspected all 2,275 catalog +columns through the progressive 300, 1,000, and text-only 3,000-value policy. The second profile +enabled the checksummed GLiNER2 artifact on CPU with two threads, at most two candidates per table, +and the shared ten-second inference allowance. -| Sensitive | Non-sensitive | Unknown decisions | Analysis time | -| ---: | ---: | ---: | ---: | -| 343 | 1,932 | 0 | 50,082 ms | +| Profile | Sensitive | Non-sensitive | Unknown decisions | NER findings | Warm analysis time | Cold wall time | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | +| Deterministic policy | 343 | 1,932 | 0 | 0 | 50,082 ms | not measured | +| CPU NER | 361 | 1,914 | 0 | 18 | 61,349 ms | about 69,750 ms | + +The warm analysis measurement starts after the persistent worker has loaded the model, matching +ordinary backend operation. The cold wall measurement also includes one-shot container startup and +model warm-up, which added about 8.4 seconds in this environment. CPU NER added 11,267 ms, or about +22.5%, to the comparable warm analysis and proposed 18 additional sensitive columns. Coverage was reported independently from the decision: @@ -20,8 +28,9 @@ Coverage was reported independently from the decision: | ---: | ---: | ---: | ---: | | 39 | 0 | 2,128 | 108 | -The sampled no-match population comprised 1,337 columns ending after the 1,000-value target and -487 text-like columns ending after the 3,000-value target. Positive findings were: +Without NER, the sampled no-match population comprised 1,337 columns ending after the 1,000-value +target and 487 text-like columns ending after the 3,000-value target. With NER, 18 of the former +received `ner.entity` evidence. Deterministic positive findings were: | Rule | Columns | | --- | ---: | @@ -36,6 +45,10 @@ The sampled no-match population comprised 1,337 columns ending after the 1,000-v | `financial.payment_card` | 3 | | `pii.italian_fiscal_code` | 2 | +The CPU NER profile added `ner.entity` evidence for 18 columns. The model checksum manifest +contained eight entries and all eight verified before the run. The worker used the CPU-only image; +no GPU was exposed to the detector. + Three successful v2 diagnostic runs produced the same decisions and aggregate rule counts. Their times ranged from 49,253 to 130,818 ms, showing that source load still affects latency even though it no longer changes the outcome through a global deadline. An earlier run exposed an intermittent @@ -44,6 +57,7 @@ a regression test covers the fallback, while two consecutive timeouts still fail instead of creating `unknown` decisions. This is a coverage and operational benchmark, not a precision/recall acceptance test. In -particular, the 202 phone-number findings and every other rule family still require human review or -a separately approved labeled corpus before their false-positive rate can be measured. NER remains -disabled by default. +particular, the 202 phone-number findings, the 18 NER findings, and every other rule family still +require human review or a separately approved labeled corpus before precision and recall can be +measured. The v2 timing now supports operational planning, but does not by itself justify enabling +NER by default.