2.2 KiB
PSD sensitivity shadow evaluation
Date: 2026-09-02
This report records an aggregate, non-mutating evaluation of sensitivity-v1 against the PSD
workspace. The source data warehouse was accessed through the configured read-only connector. The
shadow command did not create an analysis run, update catalog metadata, or save Sensitive Data
Flags. No database, table, column, source value, matched span, or free-form diagnostic was emitted.
The runner was the local Docker arm64 CPU environment connected to the PSD source; this was not a
benchmark of the PSD production server. Both runs used the same 2,275 catalog columns and a
60-second analysis deadline.
| Profile | Sensitive | Non-sensitive | Unknown | NER findings | Analysis time |
|---|---|---|---|---|---|
| Deterministic policy | 57 | 8 | 2,210 | 0 | 60,017 ms |
| CPU NER, pre-warmed, two candidates/table, 10 s shared allowance | 68 | 0 | 2,207 | 11 | 60,022 ms |
The deterministic run produced findings from metadata, phone-number, and Italian clinical-term rules. The optional NER run identified eleven additional unresolved text candidates, but its inference time reduced the source coverage reached before the global deadline. The number of definitive non-sensitive assessments consequently fell from eight to zero, so this broad shadow run does not justify enabling NER by default.
The separate offline synthetic Italian smoke test succeeded with a full_name finding at high
confidence. The image dependency check reported no broken requirements, and PyTorch reported
cuda=False, no CUDA runtime, and zero GPU devices. A defense-in-depth test retained a raw socket
constructor before Python-level blocking and confirmed that the worker's seccomp filter still
rejected the socket syscall with EPERM.
Acceptance outcome
- Keep the deterministic TypeScript policy enabled by default.
- Keep GLiNER2 available only through the explicit CPU-only installation profile.
- Do not enable NER by default for PSD on the basis of this shadow run.
- Reconsider the PSD setting only after a benchmark on the actual target CPU and a labeled Italian corpus demonstrate a useful precision/recall gain without unacceptable coverage loss.