docs: add PSD CPU NER benchmark
This commit is contained in:
+3
-1
@@ -108,7 +108,9 @@ instead of yielding `unknown`; coverage remains visible separately from the prop
|
||||
assessments remain transient until an administrator explicitly saves them. Optional GLiNER2
|
||||
evidence is CPU-only, offline, opt-in, and never replaces the deterministic decision point; see
|
||||
`docs/operations/sensitivity-analysis.md`. The earlier v1 PSD shadow comparison kept NER disabled by
|
||||
default; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`. A v2 PSD benchmark is still due.
|
||||
default; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`. The v2 comparison completed all
|
||||
2,275 columns: CPU NER added 18 sensitive proposals and increased warm runtime from 50.1 to 61.3
|
||||
seconds; see `docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md`.
|
||||
|
||||
Physical membership, source
|
||||
comments, column types/default/nullability/PK positions, and constraint-level ordered FK pairs are
|
||||
|
||||
@@ -138,7 +138,9 @@ binary draft from its scan coverage; no content is sent to an internal or extern
|
||||
The first aggregate PSD shadow comparison is recorded in
|
||||
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
|
||||
local CPU runner, NER found additional entities but reduced total coverage under the superseded
|
||||
global deadline, so the accepted setting remains disabled by default pending a new v2 benchmark.
|
||||
global deadline. The v2 benchmark removed that confounder: CPU NER added 18 sensitive proposals and
|
||||
increased the warm analysis time from 50.1 to 61.3 seconds. It remains opt-in until a labeled Italian
|
||||
evaluation establishes that the additional findings justify their false-positive rate and cost.
|
||||
The deterministic progressive PSD run is recorded in
|
||||
[`2026-09-03-psd-progressive-sensitivity-shadow.md`](../reports/2026-09-03-psd-progressive-sensitivity-shadow.md):
|
||||
it assessed all 2,275 columns with zero `unknown` decisions and left NER disabled.
|
||||
both deterministic and CPU-NER profiles assessed all 2,275 columns with zero `unknown` decisions.
|
||||
|
||||
@@ -7,12 +7,20 @@ workspace. The configured connector accessed the source data warehouse with its
|
||||
The shadow command did not create an analysis run, update local catalog metadata, or save Sensitive
|
||||
Data Flags. No database, table, column, source value, or matched span was emitted.
|
||||
|
||||
The final post-fix run used the local Docker CPU environment, without NER. It inspected all 2,275
|
||||
catalog columns through the progressive 300, 1,000, and text-only 3,000-value policy.
|
||||
The final post-fix runs used the local Docker CPU environment. They inspected all 2,275 catalog
|
||||
columns through the progressive 300, 1,000, and text-only 3,000-value policy. The second profile
|
||||
enabled the checksummed GLiNER2 artifact on CPU with two threads, at most two candidates per table,
|
||||
and the shared ten-second inference allowance.
|
||||
|
||||
| Sensitive | Non-sensitive | Unknown decisions | Analysis time |
|
||||
| ---: | ---: | ---: | ---: |
|
||||
| 343 | 1,932 | 0 | 50,082 ms |
|
||||
| Profile | Sensitive | Non-sensitive | Unknown decisions | NER findings | Warm analysis time | Cold wall time |
|
||||
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||
| Deterministic policy | 343 | 1,932 | 0 | 0 | 50,082 ms | not measured |
|
||||
| CPU NER | 361 | 1,914 | 0 | 18 | 61,349 ms | about 69,750 ms |
|
||||
|
||||
The warm analysis measurement starts after the persistent worker has loaded the model, matching
|
||||
ordinary backend operation. The cold wall measurement also includes one-shot container startup and
|
||||
model warm-up, which added about 8.4 seconds in this environment. CPU NER added 11,267 ms, or about
|
||||
22.5%, to the comparable warm analysis and proposed 18 additional sensitive columns.
|
||||
|
||||
Coverage was reported independently from the decision:
|
||||
|
||||
@@ -20,8 +28,9 @@ Coverage was reported independently from the decision:
|
||||
| ---: | ---: | ---: | ---: |
|
||||
| 39 | 0 | 2,128 | 108 |
|
||||
|
||||
The sampled no-match population comprised 1,337 columns ending after the 1,000-value target and
|
||||
487 text-like columns ending after the 3,000-value target. Positive findings were:
|
||||
Without NER, the sampled no-match population comprised 1,337 columns ending after the 1,000-value
|
||||
target and 487 text-like columns ending after the 3,000-value target. With NER, 18 of the former
|
||||
received `ner.entity` evidence. Deterministic positive findings were:
|
||||
|
||||
| Rule | Columns |
|
||||
| --- | ---: |
|
||||
@@ -36,6 +45,10 @@ The sampled no-match population comprised 1,337 columns ending after the 1,000-v
|
||||
| `financial.payment_card` | 3 |
|
||||
| `pii.italian_fiscal_code` | 2 |
|
||||
|
||||
The CPU NER profile added `ner.entity` evidence for 18 columns. The model checksum manifest
|
||||
contained eight entries and all eight verified before the run. The worker used the CPU-only image;
|
||||
no GPU was exposed to the detector.
|
||||
|
||||
Three successful v2 diagnostic runs produced the same decisions and aggregate rule counts. Their
|
||||
times ranged from 49,253 to 130,818 ms, showing that source load still affects latency even though
|
||||
it no longer changes the outcome through a global deadline. An earlier run exposed an intermittent
|
||||
@@ -44,6 +57,7 @@ a regression test covers the fallback, while two consecutive timeouts still fail
|
||||
instead of creating `unknown` decisions.
|
||||
|
||||
This is a coverage and operational benchmark, not a precision/recall acceptance test. In
|
||||
particular, the 202 phone-number findings and every other rule family still require human review or
|
||||
a separately approved labeled corpus before their false-positive rate can be measured. NER remains
|
||||
disabled by default.
|
||||
particular, the 202 phone-number findings, the 18 NER findings, and every other rule family still
|
||||
require human review or a separately approved labeled corpus before precision and recall can be
|
||||
measured. The v2 timing now supports operational planning, but does not by itself justify enabling
|
||||
NER by default.
|
||||
|
||||
Reference in New Issue
Block a user