docs: add PSD CPU NER benchmark

This commit is contained in:
Codex
2026-09-03 10:59:09 +02:00
parent 8e778b9edb
commit b891246664
3 changed files with 31 additions and 13 deletions
+3 -1
View File
@@ -108,7 +108,9 @@ instead of yielding `unknown`; coverage remains visible separately from the prop
assessments remain transient until an administrator explicitly saves them. Optional GLiNER2
evidence is CPU-only, offline, opt-in, and never replaces the deterministic decision point; see
`docs/operations/sensitivity-analysis.md`. The earlier v1 PSD shadow comparison kept NER disabled by
default; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`. A v2 PSD benchmark is still due.
default; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`. The v2 comparison completed all
2,275 columns: CPU NER added 18 sensitive proposals and increased warm runtime from 50.1 to 61.3
seconds; see `docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md`.
Physical membership, source
comments, column types/default/nullability/PK positions, and constraint-level ordered FK pairs are
+4 -2
View File
@@ -138,7 +138,9 @@ binary draft from its scan coverage; no content is sent to an internal or extern
The first aggregate PSD shadow comparison is recorded in
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
local CPU runner, NER found additional entities but reduced total coverage under the superseded
global deadline, so the accepted setting remains disabled by default pending a new v2 benchmark.
global deadline. The v2 benchmark removed that confounder: CPU NER added 18 sensitive proposals and
increased the warm analysis time from 50.1 to 61.3 seconds. It remains opt-in until a labeled Italian
evaluation establishes that the additional findings justify their false-positive rate and cost.
The deterministic progressive PSD run is recorded in
[`2026-09-03-psd-progressive-sensitivity-shadow.md`](../reports/2026-09-03-psd-progressive-sensitivity-shadow.md):
it assessed all 2,275 columns with zero `unknown` decisions and left NER disabled.
both deterministic and CPU-NER profiles assessed all 2,275 columns with zero `unknown` decisions.
@@ -7,12 +7,20 @@ workspace. The configured connector accessed the source data warehouse with its
The shadow command did not create an analysis run, update local catalog metadata, or save Sensitive
Data Flags. No database, table, column, source value, or matched span was emitted.
The final post-fix run used the local Docker CPU environment, without NER. It inspected all 2,275
catalog columns through the progressive 300, 1,000, and text-only 3,000-value policy.
The final post-fix runs used the local Docker CPU environment. They inspected all 2,275 catalog
columns through the progressive 300, 1,000, and text-only 3,000-value policy. The second profile
enabled the checksummed GLiNER2 artifact on CPU with two threads, at most two candidates per table,
and the shared ten-second inference allowance.
| Sensitive | Non-sensitive | Unknown decisions | Analysis time |
| ---: | ---: | ---: | ---: |
| 343 | 1,932 | 0 | 50,082 ms |
| Profile | Sensitive | Non-sensitive | Unknown decisions | NER findings | Warm analysis time | Cold wall time |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Deterministic policy | 343 | 1,932 | 0 | 0 | 50,082 ms | not measured |
| CPU NER | 361 | 1,914 | 0 | 18 | 61,349 ms | about 69,750 ms |
The warm analysis measurement starts after the persistent worker has loaded the model, matching
ordinary backend operation. The cold wall measurement also includes one-shot container startup and
model warm-up, which added about 8.4 seconds in this environment. CPU NER added 11,267 ms, or about
22.5%, to the comparable warm analysis and proposed 18 additional sensitive columns.
Coverage was reported independently from the decision:
@@ -20,8 +28,9 @@ Coverage was reported independently from the decision:
| ---: | ---: | ---: | ---: |
| 39 | 0 | 2,128 | 108 |
The sampled no-match population comprised 1,337 columns ending after the 1,000-value target and
487 text-like columns ending after the 3,000-value target. Positive findings were:
Without NER, the sampled no-match population comprised 1,337 columns ending after the 1,000-value
target and 487 text-like columns ending after the 3,000-value target. With NER, 18 of the former
received `ner.entity` evidence. Deterministic positive findings were:
| Rule | Columns |
| --- | ---: |
@@ -36,6 +45,10 @@ The sampled no-match population comprised 1,337 columns ending after the 1,000-v
| `financial.payment_card` | 3 |
| `pii.italian_fiscal_code` | 2 |
The CPU NER profile added `ner.entity` evidence for 18 columns. The model checksum manifest
contained eight entries and all eight verified before the run. The worker used the CPU-only image;
no GPU was exposed to the detector.
Three successful v2 diagnostic runs produced the same decisions and aggregate rule counts. Their
times ranged from 49,253 to 130,818 ms, showing that source load still affects latency even though
it no longer changes the outcome through a global deadline. An earlier run exposed an intermittent
@@ -44,6 +57,7 @@ a regression test covers the fallback, while two consecutive timeouts still fail
instead of creating `unknown` decisions.
This is a coverage and operational benchmark, not a precision/recall acceptance test. In
particular, the 202 phone-number findings and every other rule family still require human review or
a separately approved labeled corpus before their false-positive rate can be measured. NER remains
disabled by default.
particular, the 202 phone-number findings, the 18 NER findings, and every other rule family still
require human review or a separately approved labeled corpus before precision and recall can be
measured. The v2 timing now supports operational planning, but does not by itself justify enabling
NER by default.