feat: classify sensitive columns locally
This commit is contained in:
@@ -21,13 +21,21 @@ completed scan with no finding may produce `non_sensitive`; an incomplete scan w
|
||||
produces `unknown`.
|
||||
|
||||
Ambiguous text may additionally be sent to an optional local NER detector only while time remains.
|
||||
The detector runs on CPU, receives no tools or network access, does not persist source values, and
|
||||
returns evidence rather than the column decision. The initial supported detector is
|
||||
By default it receives at most two candidates per table and shares a ten-second allowance across
|
||||
the entire analysis run. The detector runs on CPU, receives no tools or network access (enforced
|
||||
inside the worker with a fail-closed seccomp network-syscall filter), does not
|
||||
persist source values, and returns evidence rather than the column decision. The initial supported detector is
|
||||
[`fastino/gliner2-privacy-filter-PII-multi`](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi),
|
||||
used through the Apache-2.0 GLiNER2 Python library with a pinned model revision. Its model weights
|
||||
and GLiNER2 code are Apache-2.0, and its mDeBERTa base model is MIT. It is trained for seven
|
||||
languages including Italian and can run on CPU without using the installation's GPUs.
|
||||
|
||||
The selected checkpoint currently carries Transformers 5 tokenizer metadata while the released
|
||||
GLiNER2 2.0.0 runtime requires Transformers 4. ThothII may bridge only that known key rename in a
|
||||
temporary view of the immutable, checksummed artifact; unexpected or ambiguous metadata fails
|
||||
closed. Removing the compatibility bridge requires an offline smoke test against a corrected,
|
||||
pinned upstream release.
|
||||
|
||||
The NER detector is an optional installation asset because its weights and runtime are materially
|
||||
larger than the deterministic TypeScript engine. It is invoked only for otherwise unresolved text,
|
||||
never for values already classified by a decisive rule. If it is disabled, unavailable, times out,
|
||||
|
||||
@@ -104,11 +104,19 @@ or second orchestration subsystem. A target receives at most one provider retry;
|
||||
exhausted technical batches fail the run. Stale work is marked interrupted at startup and must be
|
||||
explicitly unlocked; it never resumes automatically.
|
||||
|
||||
Sensitive-field suggestion generation remains a synchronous administrative request, but each
|
||||
attempt has its own durable run and ordered sanitized events. This history is separate from
|
||||
Description Generation because its lifecycle and counters differ. Only execution metadata and
|
||||
aggregate counts are stored; proposed flags, prompts, raw model output, and provider diagnostics
|
||||
remain transient.
|
||||
Sensitivity analysis is a synchronous administrative request and does not use the installation
|
||||
model catalog. Database-specific adapters stream bounded normalized values from read-only source
|
||||
connections; the TypeScript `SensitivityClassifier` is the single decision point for
|
||||
`sensitive | non_sensitive | unknown`. Deterministic rules run first. A complete scan is attempted
|
||||
for at most five seconds per table, then the adapter samples within the sixty-second request budget.
|
||||
An optional offline GLiNER2 worker may add NER evidence on CPU for unresolved short text, but it
|
||||
cannot make or persist the decision itself.
|
||||
|
||||
Each attempt has its own durable run and ordered sanitized events, separate from Description
|
||||
Generation because its lifecycle and counters differ. The run records the local policy version,
|
||||
coverage aggregates, and sanitized rule identifiers. Proposed flags, source values, NER spans, and
|
||||
worker diagnostics remain transient. Only an explicit administrator save changes the human-owned
|
||||
Sensitive Data Flag.
|
||||
|
||||
## Main backend classes
|
||||
|
||||
|
||||
@@ -137,14 +137,17 @@ Generated descriptions can be requested for selected tables, selected columns, e
|
||||
target, or targets with a missing generated description. The backend accepts one installation-wide
|
||||
run and processes targets sequentially. Every catalog column has a **Sensitive** flag, which defaults
|
||||
to `false`, including after a newly discovered column is synchronized. Before generation, an
|
||||
administrator can ask the configured model to suggest flags from structural metadata only (database,
|
||||
schema, table and column names, data types, nullability, primary keys, and foreign keys). Suggestions
|
||||
remain an unsaved draft until a human reviews and saves them.
|
||||
administrator can run local sensitivity analysis over the selected database, tables, or columns.
|
||||
The analysis combines structural metadata with bounded read-only inspection of source values. It
|
||||
uses no generative AI and no installation-catalog model. Assessments remain an unsaved draft until
|
||||
a human reviews and saves them; the reviewer may reverse any proposal.
|
||||
|
||||
The page exposes separate histories for description generation and sensitive-field suggestions.
|
||||
Sensitive-suggestion history stores the selected model, scope, status, aggregate counts, timestamps,
|
||||
and sanitized events. It does not store the proposed per-column flags, prompts, raw model output, or
|
||||
provider diagnostics; closing an unsaved review still discards that draft.
|
||||
The page exposes separate histories for description generation and sensitivity analysis. Analysis
|
||||
history stores the local policy version, scope, status, aggregate `sensitive`, `non_sensitive`, and
|
||||
`unknown` counts, timestamps, and sanitized events. It does not store source values, per-column
|
||||
proposals, NER spans, or worker diagnostics; closing an unsaved review discards that draft. The
|
||||
rules, time bounds, and optional CPU-only NER profile are documented in
|
||||
[Local sensitivity analysis](sensitivity-analysis.md).
|
||||
|
||||
For a column with `sensitive=false`, the worker may read at most five source rows and five
|
||||
representative non-null values through a read-only connector. For `sensitive=true`, the source query
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
# Local sensitivity analysis
|
||||
|
||||
Database Management can assess selected columns without sending their metadata or contents to a
|
||||
generative model. The feature is advisory: it creates a transient review draft, while the catalog's
|
||||
Sensitive Data Flag changes only when an administrator explicitly saves a choice. The administrator
|
||||
may set either value, including overriding a `sensitive` proposal.
|
||||
|
||||
## Default policy
|
||||
|
||||
`SensitivityClassifier` is the only column-level decision point. The versioned `sensitivity-v1`
|
||||
policy combines:
|
||||
|
||||
- normalized column-name rules for direct identifiers, credentials, and health data;
|
||||
- validated content rules for email, Italian fiscal code and VAT, passport, identity-card and
|
||||
driving-licence identifiers, phone numbers, IBAN/BIC, payment-card checksums, IP/MAC addresses,
|
||||
URLs, UUIDs, access keys, private-key markers, sensitive keys inside bounded recursive JSON, and
|
||||
a reviewed Italian clinical-term dictionary;
|
||||
- a conservative length rule: any observed textual value longer than 500 characters makes the
|
||||
entire column sensitive.
|
||||
|
||||
One decisive value is enough to classify the column as `sensitive`. A complete scan with no match
|
||||
may classify it as `non_sensitive`. Empty, all-null, binary/uninspectable, interrupted, and sampled
|
||||
no-match columns are `unknown`; an `unknown` draft preserves the current human flag.
|
||||
|
||||
Source reads are database-specific, but decisions are database-independent. PostgreSQL direct and
|
||||
REST `run_query` adapters project at most 501 characters per value, use only `SELECT`, and never
|
||||
persist source values. A full scan gets five seconds per table. If it cannot finish, the adapter uses
|
||||
a bounded repeatable sample within the sixty-second request deadline. PostgreSQL-wire reads run in a
|
||||
read-only transaction and always end with rollback. A REST scan can claim complete coverage only
|
||||
when it finishes in one request; multi-request pagination has no shared source transaction and is
|
||||
therefore conservatively reported as sampled.
|
||||
|
||||
The HTTP operation stops waiting at sixty seconds. The same expiring signal is checked before and
|
||||
after catalog selection, source access, progress writes, and every table. If it expires after a run
|
||||
has been created, that run is finalized as `interrupted` and all not-decisively-processed columns
|
||||
are counted as `unknown`; no review payload is returned from the timed-out request.
|
||||
|
||||
History stores only the policy version, aggregate outcomes, timestamps, and fixed operational
|
||||
events. Sanitized rule IDs are returned in the transient review and shadow report, not persisted.
|
||||
Neither path stores values, matched spans, prompts, or free-form model output.
|
||||
|
||||
## Optional CPU-only GLiNER2 evidence
|
||||
|
||||
The deterministic engine works without Python NER. An installation may opt into
|
||||
`fastino/gliner2-privacy-filter-PII-multi` for unresolved short text. It runs in a persistent local
|
||||
Python worker, adds sanitized evidence, and never becomes a second decision point. The worker:
|
||||
|
||||
- loads a local model directory only and forces Hugging Face/Transformers offline mode;
|
||||
- starts warming in the background when the backend starts; an analysis never waits for warm-up
|
||||
and skips NER until the worker is ready, so loading cannot consume the run's NER allowance;
|
||||
- hides CUDA and HIP devices and loads weights with `map_location="cpu"`;
|
||||
- starts with a scrubbed environment, then installs a fail-closed seccomp filter that denies
|
||||
network syscalls before accepting source text (the Python socket API is disabled as defense in depth);
|
||||
- receives at most two 500-character candidates per table by default, selected breadth-first
|
||||
across unresolved columns, and shares a ten-second NER allowance across the whole run;
|
||||
- returns only column ID, normalized label, and confidence; source text and entity spans are not
|
||||
returned or stored;
|
||||
- is skipped on timeout, startup failure, invalid output, or absent configuration. No LLM fallback
|
||||
is selected.
|
||||
|
||||
The pinned model revision is `c153999da5f4c509df4322b0c6a1baf3d2c284d7`. GLiNER2 and the model
|
||||
are Apache-2.0; the published mDeBERTa base is MIT. The optional runtime pins
|
||||
`gliner2[local]==2.0.0`, `transformers==4.57.6`, and the CPU-only PyTorch wheel
|
||||
`torch==2.14.0+cpu`. It lives in `/opt/sensitivity-ner`, is not installed in the default core image,
|
||||
and does not install CUDA packages.
|
||||
|
||||
The current upstream checkpoint was saved by Transformers 5.8 even though GLiNER2 2.0.0 officially
|
||||
requires Transformers `<5`; the resulting tokenizer error is independently reported in
|
||||
[GLiNER2 issue 145](https://github.com/fastino-ai/GLiNER2/issues/145). At startup ThothII leaves the
|
||||
pinned model directory unchanged and creates a temporary symlink view that maps the checkpoint's
|
||||
`extra_special_tokens` list to the Transformers 4 name `additional_special_tokens`. Any other or
|
||||
ambiguous shape fails closed and leaves the optional NER unavailable. The offline CPU smoke test
|
||||
must remain part of every dependency or model revision update.
|
||||
|
||||
## Prepare and enable the optional profile
|
||||
|
||||
Download happens during explicit installation, never during inference:
|
||||
|
||||
```bash
|
||||
./scripts/fetch-sensitivity-ner-model.sh /absolute/path/to/gliner2-pii
|
||||
```
|
||||
|
||||
The script builds the separate `thothii-core:sensitivity-ner` image, downloads the exact revision, and writes
|
||||
`MODEL_SHA256SUMS`. Keep the model directory outside the repository. Then set:
|
||||
|
||||
```bash
|
||||
export THOTH_ENABLE_SENSITIVITY_NER=1
|
||||
export THT_SENSITIVITY_NER_MODEL_DIR=/absolute/path/to/gliner2-pii
|
||||
export THT_SENSITIVITY_NER_THREADS=2
|
||||
./scripts/run-stack.sh
|
||||
```
|
||||
|
||||
For an operator-managed Compose invocation, include `deploy/compose.sensitivity-ner.yaml` after the
|
||||
base and installation overlays. The core build argument `INSTALL_SENSITIVITY_NER=true` installs the
|
||||
optional Python dependencies into their isolated virtualenv. The model mount is read-only. Values
|
||||
above eight threads are rejected; start with two so classification cannot contend heavily with
|
||||
other CPU workloads.
|
||||
|
||||
## Acceptance on a real database
|
||||
|
||||
Run the first evaluation in shadow mode: read the source with its existing read-only role, do not
|
||||
save proposed flags, and report only aggregate counts, rule IDs, coverage, and timings. Never copy
|
||||
matched values into test output. Use a separately approved, labeled Italian corpus to calculate
|
||||
precision and recall; raw PSD values must remain inside the authorized environment.
|
||||
|
||||
Inside the configured core runtime, the non-mutating command is:
|
||||
|
||||
```bash
|
||||
npm run sensitivity:shadow -- psd-clinical
|
||||
```
|
||||
|
||||
It reads catalog metadata and source values but emits one aggregate JSON object with no database,
|
||||
table, column, or source-value detail. It neither creates an analysis run nor updates a flag.
|
||||
|
||||
Enabling NER by default requires all of these gates:
|
||||
|
||||
1. the pinned artifact and `MODEL_SHA256SUMS` are archived with the installation inventory;
|
||||
2. the Python dependency/license inventory contains only redistribution-compatible licenses;
|
||||
3. the CPU benchmark stays within the configured deadlines and does not use a GPU;
|
||||
4. the labeled Italian evaluation meets thresholds approved by the product owner.
|
||||
|
||||
If a gate fails, leave NER disabled. The deterministic policy remains available and unresolved
|
||||
columns remain `unknown` rather than being sent to an internal or external LLM.
|
||||
|
||||
The first aggregate PSD shadow comparison is recorded in
|
||||
[`2026-09-02-psd-sensitivity-shadow.md`](../reports/2026-09-02-psd-sensitivity-shadow.md). On the
|
||||
local CPU runner, NER found additional entities but reduced total coverage inside the 60-second
|
||||
deadline, so the accepted setting remains disabled by default.
|
||||
@@ -0,0 +1,37 @@
|
||||
# PSD sensitivity shadow evaluation
|
||||
|
||||
Date: 2026-09-02
|
||||
|
||||
This report records an aggregate, non-mutating evaluation of `sensitivity-v1` against the PSD
|
||||
workspace. The source data warehouse was accessed through the configured read-only connector. The
|
||||
shadow command did not create an analysis run, update catalog metadata, or save Sensitive Data
|
||||
Flags. No database, table, column, source value, matched span, or free-form diagnostic was emitted.
|
||||
|
||||
The runner was the local Docker `arm64` CPU environment connected to the PSD source; this was not a
|
||||
benchmark of the PSD production server. Both runs used the same 2,275 catalog columns and a
|
||||
60-second analysis deadline.
|
||||
|
||||
| Profile | Sensitive | Non-sensitive | Unknown | NER findings | Analysis time |
|
||||
| --- | ---: | ---: | ---: | ---: | ---: |
|
||||
| Deterministic policy | 57 | 8 | 2,210 | 0 | 60,017 ms |
|
||||
| CPU NER, pre-warmed, two candidates/table, 10 s shared allowance | 68 | 0 | 2,207 | 11 | 60,022 ms |
|
||||
|
||||
The deterministic run produced findings from metadata, phone-number, and Italian clinical-term
|
||||
rules. The optional NER run identified eleven additional unresolved text candidates, but its
|
||||
inference time reduced the source coverage reached before the global deadline. The number of
|
||||
definitive non-sensitive assessments consequently fell from eight to zero, so this broad shadow
|
||||
run does not justify enabling NER by default.
|
||||
|
||||
The separate offline synthetic Italian smoke test succeeded with a `full_name` finding at high
|
||||
confidence. The image dependency check reported no broken requirements, and PyTorch reported
|
||||
`cuda=False`, no CUDA runtime, and zero GPU devices. A defense-in-depth test retained a raw socket
|
||||
constructor before Python-level blocking and confirmed that the worker's seccomp filter still
|
||||
rejected the socket syscall with `EPERM`.
|
||||
|
||||
## Acceptance outcome
|
||||
|
||||
- Keep the deterministic TypeScript policy enabled by default.
|
||||
- Keep GLiNER2 available only through the explicit CPU-only installation profile.
|
||||
- Do not enable NER by default for PSD on the basis of this shadow run.
|
||||
- Reconsider the PSD setting only after a benchmark on the actual target CPU and a labeled Italian
|
||||
corpus demonstrate a useful precision/recall gain without unacceptable coverage loss.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Optional sensitivity NER license inventory
|
||||
|
||||
This inventory covers the isolated `/opt/sensitivity-ner` Python environment built from
|
||||
`backend/python/sensitivity-ner-requirements.txt` on 2 September 2026. It is a technical release
|
||||
gate, not legal advice. Every dependency is version-locked; changing any version requires
|
||||
regenerating this inventory and rerunning the offline CPU smoke test.
|
||||
|
||||
The optional runtime also dynamically links Debian's `libseccomp2` (LGPL-2.1-only) solely to
|
||||
install its kernel-enforced network syscall filter; no libseccomp source is incorporated into ThothII.
|
||||
|
||||
No dependency or selected model uses a non-commercial, research-only, source-available, GPL, or
|
||||
AGPL license. MPL-2.0, PSF-2.0, and the permissive composite licenses below allow free-of-charge and
|
||||
commercial use, but distributors must still preserve their applicable notices and license texts.
|
||||
|
||||
| License family | Locked packages |
|
||||
| --- | --- |
|
||||
| Apache-2.0 | `accelerate==1.14.0`, `gliner2==2.0.0`, `hf-xet==1.6.0`, `huggingface-hub==0.36.2`, `peft==0.20.0`, `requests==2.34.2`, `safetensors==0.8.0`, `tokenizers==0.22.2`, `transformers==4.57.6` |
|
||||
| MIT | `annotated-types==0.8.0`, `charset-normalizer==3.5.1`, `filelock==3.32.5`, `pydantic==2.13.5`, `pydantic-core==2.46.5`, `PyYAML==6.0.3`, `typing-inspection==0.4.4`, `urllib3==2.7.0` |
|
||||
| BSD-2/3-Clause | `fsspec==2026.7.0`, `idna==3.19`, `Jinja2==3.1.6`, `MarkupSafe==3.0.3`, `mpmath==1.3.0`, `networkx==3.6.1`, `psutil==7.2.2`, `sympy==1.14.0` |
|
||||
| MPL-2.0 or mixed MPL/MIT | `certifi==2026.7.22`, `tqdm==4.70.0` |
|
||||
| PSF-2.0 | `typing-extensions==4.16.0` |
|
||||
| Composite permissive | `numpy==2.5.2` (BSD-3-Clause, 0BSD, MIT, Zlib, CC0), `packaging==26.3` (Apache-2.0 or BSD-2-Clause), `regex==2026.9.3` (Apache-2.0 and CNRI-Python), `torch==2.14.0+cpu` (Apache-2.0, LLVM exception, BSD, BSL-1.0, MIT) |
|
||||
|
||||
The selected `fastino/gliner2-privacy-filter-PII-multi` weights at revision
|
||||
`c153999da5f4c509df4322b0c6a1baf3d2c284d7` are marked Apache-2.0 in the
|
||||
[model card](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi). Its published
|
||||
`microsoft/mdeberta-v3-base` base model is MIT. The Fastino training corpus is described as
|
||||
synthetic but is not published, so the training process is not independently reproducible.
|
||||
|
||||
Before distributing the optional image or model pack:
|
||||
|
||||
1. retain the upstream license and notice files for all packaged wheels, system libraries, and weights;
|
||||
2. archive `THOTHII_MODEL_REVISION` and the verified `MODEL_SHA256SUMS` beside the model;
|
||||
3. verify that `pip check` succeeds in the isolated environment;
|
||||
4. compare the installed distribution/version set with this inventory;
|
||||
5. repeat the licensing review if an upstream artifact or dependency changes.
|
||||
@@ -179,6 +179,16 @@ Urchade model. This does not prove Italian clinical accuracy: the published SPY
|
||||
English and the training corpus is synthetic. Those are quality and reproducibility limitations,
|
||||
not a reason to reintroduce a generative LLM into the classifier.
|
||||
|
||||
An implementation smoke test found a packaging incompatibility in the selected upstream versions:
|
||||
the checkpoint was written by Transformers 5.8, while `gliner2[local]==2.0.0` requires
|
||||
`transformers>=4.38,<5`. Every published checkpoint revision has the same tokenizer metadata, and
|
||||
[upstream issue 145](https://github.com/fastino-ai/GLiNER2/issues/145) reports the identical failure.
|
||||
ThothII therefore pins Transformers 4.57.6 and performs the minimal documented-key conversion in a
|
||||
temporary local view, without changing the downloaded model or its checksums. This compatibility
|
||||
shim was accepted only after an offline CPU smoke test detected Italian names, dates, locations,
|
||||
and usernames; any unexpected metadata shape fails closed. A future upstream fix must replace,
|
||||
not silently stack on, this shim.
|
||||
|
||||
## Fit with current ThothII design
|
||||
|
||||
The current backend already owns source sampling and the human-owned `Sensitive Data Flag`; source
|
||||
|
||||
@@ -84,7 +84,7 @@ stato ripreso nella descrizione. Deve quindi essere ripetuta sul comportamento p
|
||||
- scope di sincronizzazione `tables`, `columns`, `relationships` e `all`;
|
||||
- run durevoli, conferma delle differenze distruttive, cancellazione, recovery, eventi SSE e
|
||||
fallback polling;
|
||||
- Sensitive Data Flag, suggerimenti AI strutturali, review draft e storico operativo;
|
||||
- Sensitive Data Flag, analisi locale di metadati e contenuti, review draft e storico operativo;
|
||||
- scope di generazione `selected_columns`, `selected_tables`, `all` e `missing`;
|
||||
- campionamento read-only, valori sintetici per colonne protette, batching, retry, stop, Unlock,
|
||||
storico e consolidamento;
|
||||
@@ -211,17 +211,17 @@ riapplicare le migrazioni e ripartire dal database reale.
|
||||
| ID | P | Livello | Scenario | Risultato atteso |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| PRV-01 | P0 | Repository/API | Prima sincronizzazione, re-sync e aggiunta di una colonna. | Il default è `false`; il valore umano delle colonne esistenti è preservato; la nuova colonna è esplicitamente da riesaminare ma non riceve uno stato audit inventato. |
|
||||
| PRV-02 | P0 | API | Suggerimento per un database, tabelle selezionate e colonne selezionate; database multipli, target duplicati o mancanti. | Il provider riceve esattamente le colonne dello scope. Input ambigui sono rifiutati prima della chiamata e nessun flag cambia. |
|
||||
| PRV-03 | P0 | Contratto | Ispezione del messaggio al classifier. | Sono presenti solo database, schema, tabella, colonna, tipo, nullabilità, PK e FK. Non compaiono righe, valori, commenti, descrizioni, flag corrente o segreti. |
|
||||
| PRV-04 | P1 | Contratto | `wide_entity`, identificatori lunghi e limite byte. | Ordine deterministico, batch massimi di dieci colonne e rispetto del limite messaggio; una singola colonna non rappresentabile fallisce prima del provider con errore sicuro. |
|
||||
| PRV-05 | P0 | API | Risposta valida, fenced/prosa, JSON malformato, target mancante/duplicato/ignoto e provider failure. | Ogni colonna richiesta compare una sola volta. Una classificazione invalida viene ritentata una volta; dopo esaurimento si ottiene errore sanitizzato e nessuna modifica. |
|
||||
| PRV-02 | P0 | API | Analisi per un database, tabelle selezionate e colonne selezionate; database multipli, target duplicati o mancanti. | Solo i target dello scope raggiungono l'adapter read-only. Input ambigui sono rifiutati prima della lettura e nessun flag cambia. |
|
||||
| PRV-03 | P0 | Unit/Contratto | Valori con email, codice fiscale italiano valido, IBAN, carta con Luhn, chiave privata, chiave JSON sensibile e testo oltre 500 caratteri. | Un solo riscontro validato rende l'intera colonna `sensitive`; l'evidenza contiene solo rule ID e conteggi sanitizzati, mai il valore. |
|
||||
| PRV-04 | P0 | Integrazione | Scansione completa oltre cinque secondi, timeout PostgreSQL e budget globale di sessanta secondi. | L'adapter passa al campionamento, ripristina la transazione dopo `statement_timeout`, resta read-only e non supera la deadline. Copertura incompleta senza match produce `unknown`. |
|
||||
| PRV-05 | P0 | Unit/API | Tabella vuota, colonna all-null, binario non ispezionabile, scan completo senza match e scan incompleto senza match. | Gli esiti sono rispettivamente `unknown`, `unknown`, `unknown`, `non_sensitive` e `unknown`; `unknown` conserva la scelta umana corrente. |
|
||||
| PRV-06 | P0 | UI/API | Apertura draft, modifica manuale, chiusura/reload e salvataggio. | La proposta non è persistita prima di Save; reload la scarta. Il reviewer può invertire scelte; si salvano solo colonne cambiate con versione ottimistica; un conflitto richiede reload. |
|
||||
| PRV-07 | P1 | Repository/UI | Tentativi completati, falliti e attivi al restart. | Ogni tentativo ha un run distinto con scope, modello, contatori ed eventi sanitizzati; startup marca `interrupted` i run attivi. Storico newest-first senza target ID, proposte, prompt, output grezzo o diagnostica provider. |
|
||||
| PRV-07 | P1 | Repository/UI | Tentativi completati, falliti e attivi al restart. | Ogni tentativo ha un run distinto con scope, engine `local`, versione policy, tre contatori ed eventi sanitizzati; startup marca `interrupted` i run attivi. Storico newest-first senza target ID, proposte, valori o diagnostica worker. |
|
||||
| PRV-08 | P0 | Integrazione | Generazione descrizioni su target con canary protetti. | Le colonne protette sono assenti dalla proiezione SQL, non semplicemente filtrate dopo la lettura. Se non rimangono colonne leggibili non viene eseguita una `SELECT`. Nessun canary protetto esce dal processo. |
|
||||
| PRV-09 | P0 | Contratto/Integrazione | Tabella mista con colonne sensibili e pubbliche. | Per le sensibili il prompt contiene valori plausibili, deterministici e limitati derivati dai soli metadati, nello stesso formato dei campioni e senza etichettarli al modello come sintetici. Per le pubbliche: massimo cinque righe e cinque valori rappresentativi, valori troncati e transazione read-only chiusa con rollback. |
|
||||
| PRV-10 | P1 | API | Cambio `false→true→false` dopo una descrizione già generata. | Il testo esistente non viene rigenerato retroattivamente. Solo le generazioni future cambiano fonte del contesto; tornando `false` il campionamento reale torna eleggibile. |
|
||||
| PRV-11 | P0 | API/UI | Utente senza `database.manage`, modello non configurato, catalogo/provider indisponibile e richiesta interrotta. | Controlli nascosti/disabilitati in UI e rifiuto server-side; errori non espongono dettagli. Un tentativo fallito compare nello storico senza trasformarsi in audit della decisione umana. |
|
||||
| PRV-12 | P1 | L2 | Corpus strutturale etichettato con identificatori personali, credenziali/token, salute, finanza, localizzazione e controlli non sensibili/ambigui, in inglese e italiano. | Si misurano precisione, recall e falsi negativi per modello. I campi critici mancati sono sottoposti al product owner; la soglia quantitativa va ratificata prima di diventare gate, perché il classifier è advisory e human-in-the-loop. |
|
||||
| PRV-11 | P0 | API/UI | Utente senza `database.manage`, sorgente/adapter indisponibile, NER assente o in timeout e richiesta interrotta. | Controlli nascosti/disabilitati in UI e rifiuto server-side; errori non espongono dettagli. Il NER opzionale degrada alle regole/coverage senza selezionare un LLM. |
|
||||
| PRV-12 | P1 | L2 | Corpus etichettato con identificatori personali, credenziali/token, salute, finanza, localizzazione e controlli non sensibili/ambigui, in inglese e italiano. | Si misurano precisione, recall, falsi negativi, copertura e latenza separatamente per policy deterministica e NER CPU. La soglia va ratificata prima di abilitare NER per default; il classifier resta advisory e human-in-the-loop. |
|
||||
| PRV-13 | P0 | UI/E2E | Modifica di un flag nella review senza Save e tentativo immediato di generare descrizioni. | Gate di rilascio da formalizzare: la generazione deve essere bloccata finché il draft non è salvato o scartato. In alternativa la UI deve dichiarare inequivocabilmente che verrà usato il valore persistito; non è accettabile mostrare “protetto” e campionare come non protetto. |
|
||||
|
||||
## 9. Casi di test — generazione e consolidamento dei commenti
|
||||
@@ -301,9 +301,9 @@ un valore protetto non può mai esserlo.
|
||||
| Area | Evidenza automatica già presente | Gap principale |
|
||||
| --- | --- | --- |
|
||||
| Snapshot e sincronizzazione | `backend/test/catalog-schema-introspector.test.ts`, `catalog-schema-routes.test.ts`, `catalog-table-introspector.test.ts`, `catalog-repository.integration.test.ts` | Introspezione `pg_catalog` realmente end-to-end, parità live dei tre trasporti e un unico E2E con re-scan distruttivo. |
|
||||
| Privacy | `catalog-description-generation-routes.test.ts`, `catalog-description-generation-worker.test.ts`, `catalog-description-source-sampler.test.ts`, `catalog-synthetic-sample-value.test.ts` | Prova canary integrata query→prompt→API/log e benchmark reale post-ADR 0011. |
|
||||
| Privacy | `catalog-sensitivity-classifier.test.ts`, `catalog-sensitivity-value-source.test.ts`, `catalog-local-ner-detector.test.ts` e i test di route/review | Prova shadow PSD con report solo aggregato, corpus italiano etichettato e benchmark NER CPU post-ADR 0014. |
|
||||
| Generazione | `catalog-description-generation-routes.test.ts`, `catalog-description-generation-worker.test.ts`, `catalog-description-generation.integration.test.ts` e test del helper | Accettazione reale aggiornata, cancellazione di una query PostgreSQL bloccata e integrazione ermetica fino all'endpoint LiteLLM locale. |
|
||||
| UI | `DatabaseManagementPage.test.tsx`, `DescriptionGenerationDrawer.test.tsx`, `SensitiveDataSuggestionHistoryDrawer.test.tsx` | L'E2E Playwright corrente verifica soprattutto il layout, non il workflow funzionale. |
|
||||
| UI | `DatabaseManagementPage.test.tsx`, `DescriptionGenerationDrawer.test.tsx`, `SensitivityAnalysisHistoryDrawer.test.tsx` | L'E2E Playwright corrente verifica soprattutto il layout, non il workflow funzionale. |
|
||||
|
||||
Nuovi asset consigliati:
|
||||
|
||||
@@ -320,7 +320,7 @@ Nuovi asset consigliati:
|
||||
### Wave 1 — contratto rapido
|
||||
|
||||
- parser snapshot, introspector, scope e diff;
|
||||
- classifier strutturale, batching e validazione output;
|
||||
- classifier locale, validatori/checksum, copertura e fallback al campionamento;
|
||||
- sampler, valori sintetici, prompt bounds e parser descrizioni;
|
||||
- autorizzazione, redazione e race del coordinator.
|
||||
|
||||
@@ -340,8 +340,8 @@ Nuovi asset consigliati:
|
||||
|
||||
### Wave 4 — accettazione L2
|
||||
|
||||
- provider reale sul database collegato dopo classificazione e review dei flag;
|
||||
- benchmark PRV-12 e rubric GEN-15;
|
||||
- analisi shadow sul database collegato e review dei flag, senza scritture in sorgente;
|
||||
- benchmark NER CPU PRV-12 e rubric GEN-15 per la generazione descrizioni;
|
||||
- scansione finale di canary e segreti;
|
||||
- approvazione del product owner.
|
||||
|
||||
@@ -418,7 +418,7 @@ sorgente reale.
|
||||
mostrare la disclosure anche per tabelle/colonne selezionate; il piano considera entrambi P0.
|
||||
- L'AbortSignal corrente va provato contro una query PostgreSQL realmente bloccata: la sola
|
||||
cancellazione del helper non dimostra che la lettura sorgente sia interrompibile.
|
||||
- Lo storico dei Sensitive Data Suggestion Run è operativo, non un audit delle decisioni umane.
|
||||
- Lo storico dei Sensitivity Analysis Run è operativo, non un audit delle decisioni umane.
|
||||
- La policy privacy non è ancora applicata allo schema-linking/LSH; nessun risultato di questo piano
|
||||
deve essere presentato come copertura di quel percorso.
|
||||
- Un provider reale resta non deterministico: il rilascio deve dipendere dai gate tecnici e dalla
|
||||
|
||||
Reference in New Issue
Block a user