feat: classify sensitive columns locally

This commit is contained in:
Codex
2026-09-03 02:11:13 +02:00
parent 7b87e95427
commit f114d0065a
57 changed files with 4038 additions and 1149 deletions
@@ -21,13 +21,21 @@ completed scan with no finding may produce `non_sensitive`; an incomplete scan w
produces `unknown`.
Ambiguous text may additionally be sent to an optional local NER detector only while time remains.
The detector runs on CPU, receives no tools or network access, does not persist source values, and
returns evidence rather than the column decision. The initial supported detector is
By default it receives at most two candidates per table and shares a ten-second allowance across
the entire analysis run. The detector runs on CPU, receives no tools or network access (enforced
inside the worker with a fail-closed seccomp network-syscall filter), does not
persist source values, and returns evidence rather than the column decision. The initial supported detector is
[`fastino/gliner2-privacy-filter-PII-multi`](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi),
used through the Apache-2.0 GLiNER2 Python library with a pinned model revision. Its model weights
and GLiNER2 code are Apache-2.0, and its mDeBERTa base model is MIT. It is trained for seven
languages including Italian and can run on CPU without using the installation's GPUs.
The selected checkpoint currently carries Transformers 5 tokenizer metadata while the released
GLiNER2 2.0.0 runtime requires Transformers 4. ThothII may bridge only that known key rename in a
temporary view of the immutable, checksummed artifact; unexpected or ambiguous metadata fails
closed. Removing the compatibility bridge requires an offline smoke test against a corrected,
pinned upstream release.
The NER detector is an optional installation asset because its weights and runtime are materially
larger than the deterministic TypeScript engine. It is invoked only for otherwise unresolved text,
never for values already classified by a decisive rule. If it is disabled, unavailable, times out,
+13 -5
View File
@@ -104,11 +104,19 @@ or second orchestration subsystem. A target receives at most one provider retry;
exhausted technical batches fail the run. Stale work is marked interrupted at startup and must be
explicitly unlocked; it never resumes automatically.
Sensitive-field suggestion generation remains a synchronous administrative request, but each
attempt has its own durable run and ordered sanitized events. This history is separate from
Description Generation because its lifecycle and counters differ. Only execution metadata and
aggregate counts are stored; proposed flags, prompts, raw model output, and provider diagnostics
remain transient.
Sensitivity analysis is a synchronous administrative request and does not use the installation
model catalog. Database-specific adapters stream bounded normalized values from read-only source
connections; the TypeScript `SensitivityClassifier` is the single decision point for
`sensitive | non_sensitive | unknown`. Deterministic rules run first. A complete scan is attempted
for at most five seconds per table, then the adapter samples within the sixty-second request budget.
An optional offline GLiNER2 worker may add NER evidence on CPU for unresolved short text, but it
cannot make or persist the decision itself.
Each attempt has its own durable run and ordered sanitized events, separate from Description
Generation because its lifecycle and counters differ. The run records the local policy version,
coverage aggregates, and sanitized rule identifiers. Proposed flags, source values, NER spans, and
worker diagnostics remain transient. Only an explicit administrator save changes the human-owned
Sensitive Data Flag.
## Main backend classes
+10 -7
View File
@@ -137,14 +137,17 @@ Generated descriptions can be requested for selected tables, selected columns, e
target, or targets with a missing generated description. The backend accepts one installation-wide
run and processes targets sequentially. Every catalog column has a **Sensitive** flag, which defaults
to `false`, including after a newly discovered column is synchronized. Before generation, an
administrator can ask the configured model to suggest flags from structural metadata only (database,
schema, table and column names, data types, nullability, primary keys, and foreign keys). Suggestions
remain an unsaved draft until a human reviews and saves them.
administrator can run local sensitivity analysis over the selected database, tables, or columns.
The analysis combines structural metadata with bounded read-only inspection of source values. It
uses no generative AI and no installation-catalog model. Assessments remain an unsaved draft until
a human reviews and saves them; the reviewer may reverse any proposal.
The page exposes separate histories for description generation and sensitive-field suggestions.
Sensitive-suggestion history stores the selected model, scope, status, aggregate counts, timestamps,
and sanitized events. It does not store the proposed per-column flags, prompts, raw model output, or
provider diagnostics; closing an unsaved review still discards that draft.
The page exposes separate histories for description generation and sensitivity analysis. Analysis
history stores the local policy version, scope, status, aggregate `sensitive`, `non_sensitive`, and
`unknown` counts, timestamps, and sanitized events. It does not store source values, per-column
proposals, NER spans, or worker diagnostics; closing an unsaved review discards that draft. The
rules, time bounds, and optional CPU-only NER profile are documented in
[Local sensitivity analysis](sensitivity-analysis.md).
For a column with `sensitive=false`, the worker may read at most five source rows and five
representative non-null values through a read-only connector. For `sensitive=true`, the source query
+128
View File
@@ -0,0 +1,128 @@
# Local sensitivity analysis
Database Management can assess selected columns without sending their metadata or contents to a
generative model. The feature is advisory: it creates a transient review draft, while the catalog's
Sensitive Data Flag changes only when an administrator explicitly saves a choice. The administrator
may set either value, including overriding a `sensitive` proposal.
## Default policy
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v1`
policy combines:
- normalized column-name rules for direct identifiers, credentials, and health data;
- validated content rules for email, Italian fiscal code and VAT, passport, identity-card and
driving-licence identifiers, phone numbers, IBAN/BIC, payment-card checksums, IP/MAC addresses,
URLs, UUIDs, access keys, private-key markers, sensitive keys inside bounded recursive JSON, and
a reviewed Italian clinical-term dictionary;
- a conservative length rule: any observed textual value longer than 500 characters makes the
entire column sensitive.
One decisive value is enough to classify the column as `sensitive`. A complete scan with no match
may classify it as `non_sensitive`. Empty, all-null, binary/uninspectable, interrupted, and sampled
no-match columns are `unknown`; an `unknown` draft preserves the current human flag.
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
persist source values. A full scan gets five seconds per table. If it cannot finish, the adapter uses
a bounded repeatable sample within the sixty-second request deadline. PostgreSQL-wire reads run in a
read-only transaction and always end with rollback. A REST scan can claim complete coverage only
when it finishes in one request; multi-request pagination has no shared source transaction and is
therefore conservatively reported as sampled.
The HTTP operation stops waiting at sixty seconds. The same expiring signal is checked before and
after catalog selection, source access, progress writes, and every table. If it expires after a run
has been created, that run is finalized as `interrupted` and all not-decisively-processed columns
are counted as `unknown`; no review payload is returned from the timed-out request.
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
Neither path stores values, matched spans, prompts, or free-form model output.
## Optional CPU-only GLiNER2 evidence
The deterministic engine works without Python NER. An installation may opt into
`fastino/gliner2-privacy-filter-PII-multi` for unresolved short text. It runs in a persistent local
Python worker, adds sanitized evidence, and never becomes a second decision point. The worker:
- loads a local model directory only and forces Hugging Face/Transformers offline mode;
- starts warming in the background when the backend starts; an analysis never waits for warm-up
and skips NER until the worker is ready, so loading cannot consume the run's NER allowance;
- hides CUDA and HIP devices and loads weights with `map_location="cpu"`;
- starts with a scrubbed environment, then installs a fail-closed seccomp filter that denies
network syscalls before accepting source text (the Python socket API is disabled as defense in depth);
- receives at most two 500-character candidates per table by default, selected breadth-first
across unresolved columns, and shares a ten-second NER allowance across the whole run;
- returns only column ID, normalized label, and confidence; source text and entity spans are not
returned or stored;
- is skipped on timeout, startup failure, invalid output, or absent configuration. No LLM fallback
is selected.
The pinned model revision is `c153999da5f4c509df4322b0c6a1baf3d2c284d7`. GLiNER2 and the model
are Apache-2.0; the published mDeBERTa base is MIT. The optional runtime pins
`gliner2[local]==2.0.0`, `transformers==4.57.6`, and the CPU-only PyTorch wheel
`torch==2.14.0+cpu`. It lives in `/opt/sensitivity-ner`, is not installed in the default core image,
and does not install CUDA packages.
The current upstream checkpoint was saved by Transformers 5.8 even though GLiNER2 2.0.0 officially
requires Transformers `<5`; the resulting tokenizer error is independently reported in
[GLiNER2 issue 145](https://github.com/fastino-ai/GLiNER2/issues/145). At startup ThothII leaves the
pinned model directory unchanged and creates a temporary symlink view that maps the checkpoint's
`extra_special_tokens` list to the Transformers 4 name `additional_special_tokens`. Any other or
ambiguous shape fails closed and leaves the optional NER unavailable. The offline CPU smoke test
must remain part of every dependency or model revision update.
## Prepare and enable the optional profile
Download happens during explicit installation, never during inference:
```bash
./scripts/fetch-sensitivity-ner-model.sh /absolute/path/to/gliner2-pii
```
The script builds the separate `thothii-core:sensitivity-ner` image, downloads the exact revision, and writes
`MODEL_SHA256SUMS`. Keep the model directory outside the repository. Then set:
```bash
export THOTH_ENABLE_SENSITIVITY_NER=1
export THT_SENSITIVITY_NER_MODEL_DIR=/absolute/path/to/gliner2-pii
export THT_SENSITIVITY_NER_THREADS=2
./scripts/run-stack.sh
```
For an operator-managed Compose invocation, include `deploy/compose.sensitivity-ner.yaml` after the
base and installation overlays. The core build argument `INSTALL_SENSITIVITY_NER=true` installs the
optional Python dependencies into their isolated virtualenv. The model mount is read-only. Values
above eight threads are rejected; start with two so classification cannot contend heavily with
other CPU workloads.
## Acceptance on a real database
Run the first evaluation in shadow mode: read the source with its existing read-only role, do not
save proposed flags, and report only aggregate counts, rule IDs, coverage, and timings. Never copy
matched values into test output. Use a separately approved, labeled Italian corpus to calculate
precision and recall; raw PSD values must remain inside the authorized environment.
Inside the configured core runtime, the non-mutating command is:
```bash
npm run sensitivity:shadow -- psd-clinical
```
It reads catalog metadata and source values but emits one aggregate JSON object with no database,
table, column, or source-value detail. It neither creates an analysis run nor updates a flag.
Enabling NER by default requires all of these gates:
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
3. the CPU benchmark stays within the configured deadlines and does not use a GPU;
4. the labeled Italian evaluation meets thresholds approved by the product owner.
If a gate fails, leave NER disabled. The deterministic policy remains available and unresolved
columns remain `unknown` rather than being sent to an internal or external LLM.
The first aggregate PSD shadow comparison is recorded in
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
local CPU runner, NER found additional entities but reduced total coverage inside the 60-second
deadline, so the accepted setting remains disabled by default.
@@ -0,0 +1,37 @@
# PSD sensitivity shadow evaluation
Date: 2026-09-02
This report records an aggregate, non-mutating evaluation of `sensitivity-v1` against the PSD
workspace. The source data warehouse was accessed through the configured read-only connector. The
shadow command did not create an analysis run, update catalog metadata, or save Sensitive Data
Flags. No database, table, column, source value, matched span, or free-form diagnostic was emitted.
The runner was the local Docker `arm64` CPU environment connected to the PSD source; this was not a
benchmark of the PSD production server. Both runs used the same 2,275 catalog columns and a
60-second analysis deadline.
| Profile | Sensitive | Non-sensitive | Unknown | NER findings | Analysis time |
| --- | ---: | ---: | ---: | ---: | ---: |
| Deterministic policy | 57 | 8 | 2,210 | 0 | 60,017 ms |
| CPU NER, pre-warmed, two candidates/table, 10 s shared allowance | 68 | 0 | 2,207 | 11 | 60,022 ms |
The deterministic run produced findings from metadata, phone-number, and Italian clinical-term
rules. The optional NER run identified eleven additional unresolved text candidates, but its
inference time reduced the source coverage reached before the global deadline. The number of
definitive non-sensitive assessments consequently fell from eight to zero, so this broad shadow
run does not justify enabling NER by default.
The separate offline synthetic Italian smoke test succeeded with a `full_name` finding at high
confidence. The image dependency check reported no broken requirements, and PyTorch reported
`cuda=False`, no CUDA runtime, and zero GPU devices. A defense-in-depth test retained a raw socket
constructor before Python-level blocking and confirmed that the worker's seccomp filter still
rejected the socket syscall with `EPERM`.
## Acceptance outcome
- Keep the deterministic TypeScript policy enabled by default.
- Keep GLiNER2 available only through the explicit CPU-only installation profile.
- Do not enable NER by default for PSD on the basis of this shadow run.
- Reconsider the PSD setting only after a benchmark on the actual target CPU and a labeled Italian
corpus demonstrate a useful precision/recall gain without unacceptable coverage loss.
@@ -0,0 +1,36 @@
# Optional sensitivity NER license inventory
This inventory covers the isolated `/opt/sensitivity-ner` Python environment built from
`backend/python/sensitivity-ner-requirements.txt` on 2 September 2026. It is a technical release
gate, not legal advice. Every dependency is version-locked; changing any version requires
regenerating this inventory and rerunning the offline CPU smoke test.
The optional runtime also dynamically links Debian's `libseccomp2` (LGPL-2.1-only) solely to
install its kernel-enforced network syscall filter; no libseccomp source is incorporated into ThothII.
No dependency or selected model uses a non-commercial, research-only, source-available, GPL, or
AGPL license. MPL-2.0, PSF-2.0, and the permissive composite licenses below allow free-of-charge and
commercial use, but distributors must still preserve their applicable notices and license texts.
| License family | Locked packages |
| --- | --- |
| Apache-2.0 | `accelerate==1.14.0`, `gliner2==2.0.0`, `hf-xet==1.6.0`, `huggingface-hub==0.36.2`, `peft==0.20.0`, `requests==2.34.2`, `safetensors==0.8.0`, `tokenizers==0.22.2`, `transformers==4.57.6` |
| MIT | `annotated-types==0.8.0`, `charset-normalizer==3.5.1`, `filelock==3.32.5`, `pydantic==2.13.5`, `pydantic-core==2.46.5`, `PyYAML==6.0.3`, `typing-inspection==0.4.4`, `urllib3==2.7.0` |
| BSD-2/3-Clause | `fsspec==2026.7.0`, `idna==3.19`, `Jinja2==3.1.6`, `MarkupSafe==3.0.3`, `mpmath==1.3.0`, `networkx==3.6.1`, `psutil==7.2.2`, `sympy==1.14.0` |
| MPL-2.0 or mixed MPL/MIT | `certifi==2026.7.22`, `tqdm==4.70.0` |
| PSF-2.0 | `typing-extensions==4.16.0` |
| Composite permissive | `numpy==2.5.2` (BSD-3-Clause, 0BSD, MIT, Zlib, CC0), `packaging==26.3` (Apache-2.0 or BSD-2-Clause), `regex==2026.9.3` (Apache-2.0 and CNRI-Python), `torch==2.14.0+cpu` (Apache-2.0, LLVM exception, BSD, BSL-1.0, MIT) |
The selected `fastino/gliner2-privacy-filter-PII-multi` weights at revision
`c153999da5f4c509df4322b0c6a1baf3d2c284d7` are marked Apache-2.0 in the
[model card](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi). Its published
`microsoft/mdeberta-v3-base` base model is MIT. The Fastino training corpus is described as
synthetic but is not published, so the training process is not independently reproducible.
Before distributing the optional image or model pack:
1. retain the upstream license and notice files for all packaged wheels, system libraries, and weights;
2. archive `THOTHII_MODEL_REVISION` and the verified `MODEL_SHA256SUMS` beside the model;
3. verify that `pip check` succeeds in the isolated environment;
4. compare the installed distribution/version set with this inventory;
5. repeat the licensing review if an upstream artifact or dependency changes.
@@ -179,6 +179,16 @@ Urchade model. This does not prove Italian clinical accuracy: the published SPY
English and the training corpus is synthetic. Those are quality and reproducibility limitations,
not a reason to reintroduce a generative LLM into the classifier.
An implementation smoke test found a packaging incompatibility in the selected upstream versions:
the checkpoint was written by Transformers 5.8, while `gliner2[local]==2.0.0` requires
`transformers>=4.38,<5`. Every published checkpoint revision has the same tokenizer metadata, and
[upstream issue 145](https://github.com/fastino-ai/GLiNER2/issues/145) reports the identical failure.
ThothII therefore pins Transformers 4.57.6 and performs the minimal documented-key conversion in a
temporary local view, without changing the downloaded model or its checksums. This compatibility
shim was accepted only after an offline CPU smoke test detected Italian names, dates, locations,
and usernames; any unexpected metadata shape fails closed. A future upstream fix must replace,
not silently stack on, this shim.
## Fit with current ThothII design
The current backend already owns source sampling and the human-owned `Sensitive Data Flag`; source
@@ -84,7 +84,7 @@ stato ripreso nella descrizione. Deve quindi essere ripetuta sul comportamento p
- scope di sincronizzazione `tables`, `columns`, `relationships` e `all`;
- run durevoli, conferma delle differenze distruttive, cancellazione, recovery, eventi SSE e
fallback polling;
- Sensitive Data Flag, suggerimenti AI strutturali, review draft e storico operativo;
- Sensitive Data Flag, analisi locale di metadati e contenuti, review draft e storico operativo;
- scope di generazione `selected_columns`, `selected_tables`, `all` e `missing`;
- campionamento read-only, valori sintetici per colonne protette, batching, retry, stop, Unlock,
storico e consolidamento;
@@ -211,17 +211,17 @@ riapplicare le migrazioni e ripartire dal database reale.
| ID | P | Livello | Scenario | Risultato atteso |
| --- | --- | --- | --- | --- |
| PRV-01 | P0 | Repository/API | Prima sincronizzazione, re-sync e aggiunta di una colonna. | Il default è `false`; il valore umano delle colonne esistenti è preservato; la nuova colonna è esplicitamente da riesaminare ma non riceve uno stato audit inventato. |
| PRV-02 | P0 | API | Suggerimento per un database, tabelle selezionate e colonne selezionate; database multipli, target duplicati o mancanti. | Il provider riceve esattamente le colonne dello scope. Input ambigui sono rifiutati prima della chiamata e nessun flag cambia. |
| PRV-03 | P0 | Contratto | Ispezione del messaggio al classifier. | Sono presenti solo database, schema, tabella, colonna, tipo, nullabilità, PK e FK. Non compaiono righe, valori, commenti, descrizioni, flag corrente o segreti. |
| PRV-04 | P1 | Contratto | `wide_entity`, identificatori lunghi e limite byte. | Ordine deterministico, batch massimi di dieci colonne e rispetto del limite messaggio; una singola colonna non rappresentabile fallisce prima del provider con errore sicuro. |
| PRV-05 | P0 | API | Risposta valida, fenced/prosa, JSON malformato, target mancante/duplicato/ignoto e provider failure. | Ogni colonna richiesta compare una sola volta. Una classificazione invalida viene ritentata una volta; dopo esaurimento si ottiene errore sanitizzato e nessuna modifica. |
| PRV-02 | P0 | API | Analisi per un database, tabelle selezionate e colonne selezionate; database multipli, target duplicati o mancanti. | Solo i target dello scope raggiungono l'adapter read-only. Input ambigui sono rifiutati prima della lettura e nessun flag cambia. |
| PRV-03 | P0 | Unit/Contratto | Valori con email, codice fiscale italiano valido, IBAN, carta con Luhn, chiave privata, chiave JSON sensibile e testo oltre 500 caratteri. | Un solo riscontro validato rende l'intera colonna `sensitive`; l'evidenza contiene solo rule ID e conteggi sanitizzati, mai il valore. |
| PRV-04 | P0 | Integrazione | Scansione completa oltre cinque secondi, timeout PostgreSQL e budget globale di sessanta secondi. | L'adapter passa al campionamento, ripristina la transazione dopo `statement_timeout`, resta read-only e non supera la deadline. Copertura incompleta senza match produce `unknown`. |
| PRV-05 | P0 | Unit/API | Tabella vuota, colonna all-null, binario non ispezionabile, scan completo senza match e scan incompleto senza match. | Gli esiti sono rispettivamente `unknown`, `unknown`, `unknown`, `non_sensitive` e `unknown`; `unknown` conserva la scelta umana corrente. |
| PRV-06 | P0 | UI/API | Apertura draft, modifica manuale, chiusura/reload e salvataggio. | La proposta non è persistita prima di Save; reload la scarta. Il reviewer può invertire scelte; si salvano solo colonne cambiate con versione ottimistica; un conflitto richiede reload. |
| PRV-07 | P1 | Repository/UI | Tentativi completati, falliti e attivi al restart. | Ogni tentativo ha un run distinto con scope, modello, contatori ed eventi sanitizzati; startup marca `interrupted` i run attivi. Storico newest-first senza target ID, proposte, prompt, output grezzo o diagnostica provider. |
| PRV-07 | P1 | Repository/UI | Tentativi completati, falliti e attivi al restart. | Ogni tentativo ha un run distinto con scope, engine `local`, versione policy, tre contatori ed eventi sanitizzati; startup marca `interrupted` i run attivi. Storico newest-first senza target ID, proposte, valori o diagnostica worker. |
| PRV-08 | P0 | Integrazione | Generazione descrizioni su target con canary protetti. | Le colonne protette sono assenti dalla proiezione SQL, non semplicemente filtrate dopo la lettura. Se non rimangono colonne leggibili non viene eseguita una `SELECT`. Nessun canary protetto esce dal processo. |
| PRV-09 | P0 | Contratto/Integrazione | Tabella mista con colonne sensibili e pubbliche. | Per le sensibili il prompt contiene valori plausibili, deterministici e limitati derivati dai soli metadati, nello stesso formato dei campioni e senza etichettarli al modello come sintetici. Per le pubbliche: massimo cinque righe e cinque valori rappresentativi, valori troncati e transazione read-only chiusa con rollback. |
| PRV-10 | P1 | API | Cambio `false→true→false` dopo una descrizione già generata. | Il testo esistente non viene rigenerato retroattivamente. Solo le generazioni future cambiano fonte del contesto; tornando `false` il campionamento reale torna eleggibile. |
| PRV-11 | P0 | API/UI | Utente senza `database.manage`, modello non configurato, catalogo/provider indisponibile e richiesta interrotta. | Controlli nascosti/disabilitati in UI e rifiuto server-side; errori non espongono dettagli. Un tentativo fallito compare nello storico senza trasformarsi in audit della decisione umana. |
| PRV-12 | P1 | L2 | Corpus strutturale etichettato con identificatori personali, credenziali/token, salute, finanza, localizzazione e controlli non sensibili/ambigui, in inglese e italiano. | Si misurano precisione, recall e falsi negativi per modello. I campi critici mancati sono sottoposti al product owner; la soglia quantitativa va ratificata prima di diventare gate, perché il classifier è advisory e human-in-the-loop. |
| PRV-11 | P0 | API/UI | Utente senza `database.manage`, sorgente/adapter indisponibile, NER assente o in timeout e richiesta interrotta. | Controlli nascosti/disabilitati in UI e rifiuto server-side; errori non espongono dettagli. Il NER opzionale degrada alle regole/coverage senza selezionare un LLM. |
| PRV-12 | P1 | L2 | Corpus etichettato con identificatori personali, credenziali/token, salute, finanza, localizzazione e controlli non sensibili/ambigui, in inglese e italiano. | Si misurano precisione, recall, falsi negativi, copertura e latenza separatamente per policy deterministica e NER CPU. La soglia va ratificata prima di abilitare NER per default; il classifier resta advisory e human-in-the-loop. |
| PRV-13 | P0 | UI/E2E | Modifica di un flag nella review senza Save e tentativo immediato di generare descrizioni. | Gate di rilascio da formalizzare: la generazione deve essere bloccata finché il draft non è salvato o scartato. In alternativa la UI deve dichiarare inequivocabilmente che verrà usato il valore persistito; non è accettabile mostrare “protetto” e campionare come non protetto. |
## 9. Casi di test — generazione e consolidamento dei commenti
@@ -301,9 +301,9 @@ un valore protetto non può mai esserlo.
| Area | Evidenza automatica già presente | Gap principale |
| --- | --- | --- |
| Snapshot e sincronizzazione | `backend/test/catalog-schema-introspector.test.ts`, `catalog-schema-routes.test.ts`, `catalog-table-introspector.test.ts`, `catalog-repository.integration.test.ts` | Introspezione `pg_catalog` realmente end-to-end, parità live dei tre trasporti e un unico E2E con re-scan distruttivo. |
| Privacy | `catalog-description-generation-routes.test.ts`, `catalog-description-generation-worker.test.ts`, `catalog-description-source-sampler.test.ts`, `catalog-synthetic-sample-value.test.ts` | Prova canary integrata query→prompt→API/log e benchmark reale post-ADR 0011. |
| Privacy | `catalog-sensitivity-classifier.test.ts`, `catalog-sensitivity-value-source.test.ts`, `catalog-local-ner-detector.test.ts` e i test di route/review | Prova shadow PSD con report solo aggregato, corpus italiano etichettato e benchmark NER CPU post-ADR 0014. |
| Generazione | `catalog-description-generation-routes.test.ts`, `catalog-description-generation-worker.test.ts`, `catalog-description-generation.integration.test.ts` e test del helper | Accettazione reale aggiornata, cancellazione di una query PostgreSQL bloccata e integrazione ermetica fino all'endpoint LiteLLM locale. |
| UI | `DatabaseManagementPage.test.tsx`, `DescriptionGenerationDrawer.test.tsx`, `SensitiveDataSuggestionHistoryDrawer.test.tsx` | L'E2E Playwright corrente verifica soprattutto il layout, non il workflow funzionale. |
| UI | `DatabaseManagementPage.test.tsx`, `DescriptionGenerationDrawer.test.tsx`, `SensitivityAnalysisHistoryDrawer.test.tsx` | L'E2E Playwright corrente verifica soprattutto il layout, non il workflow funzionale. |
Nuovi asset consigliati:
@@ -320,7 +320,7 @@ Nuovi asset consigliati:
### Wave 1 — contratto rapido
- parser snapshot, introspector, scope e diff;
- classifier strutturale, batching e validazione output;
- classifier locale, validatori/checksum, copertura e fallback al campionamento;
- sampler, valori sintetici, prompt bounds e parser descrizioni;
- autorizzazione, redazione e race del coordinator.
@@ -340,8 +340,8 @@ Nuovi asset consigliati:
### Wave 4 — accettazione L2
- provider reale sul database collegato dopo classificazione e review dei flag;
- benchmark PRV-12 e rubric GEN-15;
- analisi shadow sul database collegato e review dei flag, senza scritture in sorgente;
- benchmark NER CPU PRV-12 e rubric GEN-15 per la generazione descrizioni;
- scansione finale di canary e segreti;
- approvazione del product owner.
@@ -418,7 +418,7 @@ sorgente reale.
mostrare la disclosure anche per tabelle/colonne selezionate; il piano considera entrambi P0.
- L'AbortSignal corrente va provato contro una query PostgreSQL realmente bloccata: la sola
cancellazione del helper non dimostra che la lettura sorgente sia interrompibile.
- Lo storico dei Sensitive Data Suggestion Run è operativo, non un audit delle decisioni umane.
- Lo storico dei Sensitivity Analysis Run è operativo, non un audit delle decisioni umane.
- La policy privacy non è ancora applicata allo schema-linking/LSH; nessun risultato di questo piano
deve essere presentato come copertura di quel percorso.
- Un provider reale resta non deterministico: il rilascio deve dipendere dai gate tecnici e dalla