Files
ThothII/docs/reports/2026-09-02-psd-sensitivity-shadow.md

2.2 KiB

PSD sensitivity shadow evaluation

Date: 2026-09-02

This report records an aggregate, non-mutating evaluation of sensitivity-v1 against the PSD workspace. The source data warehouse was accessed through the configured read-only connector. The shadow command did not create an analysis run, update catalog metadata, or save Sensitive Data Flags. No database, table, column, source value, matched span, or free-form diagnostic was emitted.

The runner was the local Docker arm64 CPU environment connected to the PSD source; this was not a benchmark of the PSD production server. Both runs used the same 2,275 catalog columns and a 60-second analysis deadline.

Profile Sensitive Non-sensitive Unknown NER findings Analysis time
Deterministic policy 57 8 2,210 0 60,017 ms
CPU NER, pre-warmed, two candidates/table, 10 s shared allowance 68 0 2,207 11 60,022 ms

The deterministic run produced findings from metadata, phone-number, and Italian clinical-term rules. The optional NER run identified eleven additional unresolved text candidates, but its inference time reduced the source coverage reached before the global deadline. The number of definitive non-sensitive assessments consequently fell from eight to zero, so this broad shadow run does not justify enabling NER by default.

The separate offline synthetic Italian smoke test succeeded with a full_name finding at high confidence. The image dependency check reported no broken requirements, and PyTorch reported cuda=False, no CUDA runtime, and zero GPU devices. A defense-in-depth test retained a raw socket constructor before Python-level blocking and confirmed that the worker's seccomp filter still rejected the socket syscall with EPERM.

Acceptance outcome

  • Keep the deterministic TypeScript policy enabled by default.
  • Keep GLiNER2 available only through the explicit CPU-only installation profile.
  • Do not enable NER by default for PSD on the basis of this shadow run.
  • Reconsider the PSD setting only after a benchmark on the actual target CPU and a labeled Italian corpus demonstrate a useful precision/recall gain without unacceptable coverage loss.