Compare commits

..
Author SHA1 Message Date
Codex ad744f0212 docs: add guarded server upgrade runbook
Publish documentation / publish (push) Successful in 1m20s
2026-09-04 17:37:26 +02:00
Codex eba6148511 test: stabilize pre-deployment gates
Remove the redundant timing-dependent native Argon2 concurrency test while retaining native vector coverage and deterministic limiter coverage. Refresh stale deployment and browser contracts, make release scripts portable across Bash/macOS, and update production dependency locks for resolved security advisories.
2026-09-04 16:15:35 +02:00
Codex 7b1d69a65b feat: complete catalog sensitivity enhancements 2026-09-04 15:11:18 +02:00
Codex b891246664 docs: add PSD CPU NER benchmark 2026-09-03 10:59:09 +02:00
Codex 8e778b9edb feat: sample sensitive columns progressively 2026-09-03 10:25:05 +02:00
Codex f114d0065a feat: classify sensitive columns locally 2026-09-03 02:11:13 +02:00
Codex 7b87e95427 fix: refresh catalog after hidden sync completion 2026-09-02 23:23:00 +02:00
Codex a50475d687 chore: ignore generated deployment projections 2026-09-02 20:35:08 +02:00
Codex a6a5bf2036 fix: harden model catalog projections 2026-09-02 19:25:01 +02:00
Codex ce4c31a6fb docs: align restore guidance with workspace v4 2026-09-02 18:47:57 +02:00
Codex 538dc8ef56 test: name workspace schema v4 gate 2026-09-02 18:46:49 +02:00
Codex 7b7927bfe5 feat: unify installation model catalog 2026-09-02 18:45:33 +02:00
278 changed files with 43604 additions and 6412 deletions
+2
View File
@@ -35,6 +35,7 @@ config/ca-chain.pem
# ThothII deployment configuration and secret values (keep only the README tracked)
deploy/.env
deploy/env/local.env
deploy/compose.connector-secrets.local.yaml
deploy/compose.psd-local.yaml
deploy/workspaces/psd.yaml
@@ -45,6 +46,7 @@ deploy/secrets/*
# Per-installation configuration generated by `tht setup` (examples stay tracked).
deploy/*/thothii-installation.yaml
deploy/*/operator.env
deploy/*/generated/
deploy/*/secrets/*
!deploy/*/secrets/.gitkeep
!deploy/*/secrets/*.example
+7 -6
View File
@@ -83,7 +83,8 @@ frontend (React/SSE) → backend (Fastify) → pi --mode rpc → tht/harness →
`SseHub` fans them out over SSE to the browser. The separate PostgreSQL catalog stores database
metadata and sequential AI description-generation runs. Description generation samples the DWH
through read-only connectors and calls a short-lived Python LiteLLM helper; it does not use Pi or
expose a public CLI command. App settings still live in `backend/data/settings.json`.
expose a public CLI command. Sessions, metadata generation, and embedding resolve models from the
generated Installation Model Catalog; `thothii-installation.yaml` is its only authored source.
- **Human-in-the-loop gate contract.** The model proposes; a human reviewer decides at gates
via widgets (`reviewer_select` = single pick — a chosen option carrying a `decision` payload
@@ -99,11 +100,11 @@ frontend (React/SSE) → backend (Fastify) → pi --mode rpc → tht/harness →
- **`--json` output must be pristine** (only valid JSON on stdout) — used as a machine contract.
- **UI strings are English; document *content* stays the workspace language** (Italian for
`psd`) because it's the real data. Only chrome/labels are English.
- **Workspaces** (`harness/workspaces/*.yaml`) set the DB target and **absolute**
`paths.sessions/artifacts/indexes` — for `psd` these point at a *separate, uncommitted* repo
(`tht-workspace-psd/`). Secrets live ONLY in `harness/.env` (gitignored).
- **Settings are global** (`backend/data/settings.json`: workspace/provider/model/thinking);
the New-session form is question-only.
- **Workspace schema v4** defines database and Evidence concerns only. Embedding/model facts come
from the installation catalog. The legacy `harness/workspaces/*.yaml` runtime snapshots still use
absolute session/artifact/index paths; secrets stay in `harness/.env` (gitignored).
- **Settings are global** (`backend/data/settings.json`: workspace/thinking). Provider/model choices
are ephemeral canonical catalog selections pinned into the session manifest.
- **Resume**: a resumable session re-enters at its last incomplete phase. The backend refuses
resume with 409 when `finalized` or `archived`, and `PiProcessManager.spawnFor` must send
`/riprendi-sessione <id>` (resume mode) vs `/nuova-domanda` (new) — sending the wrong prompt
+25 -3
View File
@@ -267,6 +267,21 @@ ridefinirne provider, endpoint o capacità.
dell'Installation Model Catalog richiesta da uno specifico runtime. Può essere rigenerata
integralmente dalla configurazione dell'installazione.
## Distribuzione del prodotto
**Customer-Hosted Installation** — Un'installazione eseguita interamente nel trust boundary
controllato dall'organizzazione cliente, inclusi eventuali tenant cloud privati. Credenziali,
domande, prompt, metadati e risultati non attraversano quel boundary.
_Avoid_: on-premise deployment, self-managed deployment
**Community Edition** — La distribuzione open source utilizzabile gratuitamente anche in
produzione e capace di eseguire il workflow fondamentale completo.
_Avoid_: free tier, trial edition
**Enterprise Edition** — La distribuzione con licenza commerciale che aggiunge governance
organizzativa, esercizio production-grade e industrializzazione alla Community Edition.
_Avoid_: paid tier, pro edition
## Catalogo dei metadati
**Workspace Database** — Il database che appartiene a un solo workspace e non può essere
@@ -420,6 +435,12 @@ diventa Catalog Metadata.
**Sensitive Data Flag** — La classificazione binaria umana applicata a una Catalog Column. Può
essere impostata liberamente dall'amministratore anche in contrasto con una valutazione automatica.
**Sensitivity Reason** — La motivazione sanificata persistita insieme al Sensitive Data Flag
quando l'amministratore salva una Sensitivity Review Draft. È Catalog Metadata della colonna, non
history della run; viene rimossa quando il flag torna non-sensitive e può essere assente per una
classificazione manuale priva di valutazione locale.
_Avoid_: AI reasoning, source evidence
**Local Sensitivity Assessment** — La valutazione locale, non autoritativa e priva di LLM di una
Catalog Column, basata su metadati e contenuto sorgente, con esito `sensitive`, `non_sensitive`
oppure `unknown`.
@@ -451,12 +472,13 @@ _Avoid_: Sensitive Data Suggestion Run, AI analysis
**Sensitivity Review Draft** — La proposta transitoria che associa alle colonne selezionate una
Local Sensitivity Assessment e le relative evidenze sanificate. Non modifica il Sensitive Data Flag
finché l'amministratore non salva le proprie decisioni e viene scartata al reload.
né la Sensitivity Reason finché l'amministratore non salva le proprie decisioni e viene scartata al
reload.
_Avoid_: automatic flag
**Sensitivity Analysis Event** — Una riga testuale ordinata e sanificata che registra l'avvio,
l'esito o l'errore di una Sensitivity Analysis Run senza conservare contenuti sorgente, output grezzi
del detector o proposte per colonna.
l'avanzamento per fase e batch, l'esito o l'errore di una Sensitivity Analysis Run senza conservare
contenuti sorgente, output grezzi del detector o proposte per colonna.
_Avoid_: Sensitive Data Suggestion Event
**Introspection Capability** — Una categoria di struttura fisica che una Database Binding
+57 -12
View File
@@ -1,12 +1,17 @@
# ThothII — Project State
Last updated: 2026-08-31.
Last updated: 2026-09-04.
This file is the short operational snapshot. Stable commands and the architecture mental model
live in `AGENTS.md`; current design and runtime contracts live under `docs/architecture/`,
`docs/contracts/`, `docs/adr/`, and `docs/evidence.md`. Superseded plans and reports are
available from Git history rather than duplicated in the working tree.
The guarded server migration from a legacy checkout to the schema-v2 installation, Gitea source,
Authentik, internal catalog/embedding services, and the PSD workspace repository is documented in
`docs/operations/server-upgrade-gitea-workspace-v2.md`. Treat its operator gates and rollback
requirements as mandatory; do not replace the running server stack in place.
## Current product shape
ThothII is a human-in-the-loop datamart builder with three independently built layers:
@@ -59,10 +64,28 @@ tht --installation /absolute/path/thothii-installation.yaml workspace preprocess
These commands use the profile-gated `workspace-maintenance` service. The former standalone
preprocessing Compose fixtures are retired.
Workspace descriptors use schema v3. For PSD, workspace content and runtime roots point to the
Workspace descriptors use schema v4 and contain only database, Evidence, diagnostics, and binding
concerns; model, provider, embedding, and vector-store configuration is installation-owned. For
PSD, workspace content and runtime roots point to the
separate uncommitted repository `/Users/mp/projects/tht-workspace-psd`. Secrets remain outside
Git and are supplied only through installation-local protected files.
## Installation Model Catalog
`thothii-installation.yaml` schema version 2 is the only operator-authored source for session,
metadata-generation, and embedding models. The host `tht` lifecycle validates `modelCatalog` and
regenerates the backend catalog, Pi `models.json`/`settings.json`, and Compose override under the
installation-local `generated/` directory. Those projections are replaceable runtime adapters:
they are not edited, backed up, or treated as configuration.
Session and metadata defaults use canonical `provider/model` IDs. Provider authentication declares
one explicit mode (`secret_env`, `pi_auth`, or `none`); `secret_env` names a protected bundle key.
The backend settings store now owns only the selected workspace and thinking level. Existing v1
installations use the explicit catalog migration command; schema-v3 workspace descriptors are
converted deterministically in their curator-owned repository before commit. Strict runtime loading
does not silently infer or merge legacy sources. ADR 0013 and
`docs/plans/2026-09-02-installation-model-catalog.md` record the decision and implementation.
## Database management
The database, table, and authoritative physical-schema catalog slices are implemented. Database
@@ -80,6 +103,23 @@ column. The KPI strip reads installation-wide or selected-database aggregates fr
description history, and sensitive-field review/history use the production APIs in right-side
drawers rather than prototype fixtures; closing a history drawer does not stop its background run.
Sensitive-field review is now driven by the versioned local `sensitivity-v4` policy, not by a
catalog model. The backend reads selected source tables through read-only, database-specific
adapters and makes every `sensitive | non_sensitive` draft decision in the TypeScript
`SensitivityClassifier`. A single validated match protects the column. Tables up to 1,000 rows are
fully scanned; larger tables use breadth-first 300, 1,000, and text-only 3,000-value targets, with a
five-second limit per source query and no global request deadline. Source failures fail the run
instead of yielding `unknown`; coverage remains visible separately from the proposal. Draft
assessments remain transient until an administrator explicitly saves them. Optional GLiNER2
evidence is CPU-only, offline, opt-in, and never replaces the deterministic decision point; see
`docs/operations/sensitivity-analysis.md`. The earlier v1 PSD shadow comparison kept NER disabled by
default; see `docs/reports/2026-09-02-psd-sensitivity-shadow.md`. The v2 comparison completed all
2,275 columns: CPU NER added 18 sensitive proposals and increased warm runtime from 50.1 to 61.3
seconds; see `docs/reports/2026-09-03-psd-progressive-sensitivity-shadow.md`.
Version 4 excludes declared `bigint` primary-key columns and conventionally named `pk bigint`
columns before source inspection, reporting both as non-informative structural identifiers while
distinguishing declared constraints from inferred roles.
Physical membership, source
comments, column types/default/nullability/PK positions, and constraint-level ordered FK pairs are
projections of the external schema. They cannot be created, renamed, or structurally edited by
@@ -137,15 +177,19 @@ Semantic aliases, value descriptions, synonyms, and concepts remain deferred to
slices.
AI Description Generation uses the catalog's human-owned Sensitive Data Flag. The flag defaults to
`false`, including for newly synchronized columns. An administrator may request an AI proposal based
only on structural metadata for one selected database, selected tables, or selected columns. The
backend divides large scopes into deterministic model requests of at most ten columns, also bounded
by helper message size, and combines their results, but the proposal remains an unsaved draft until
the human reviews and saves it.
Each started suggestion attempt records a separate Sensitive Data Suggestion Run with aggregate
counters and safe ordered events. This operational history never stores per-column proposals,
prompts, raw model output, or provider diagnostics; reloading still discards an unsaved review
draft.
`false`, including for newly synchronized columns. An administrator may request a local sensitivity
analysis for one selected database, selected tables, or selected columns. One deterministic
TypeScript classifier combines metadata, bounded source-content rules, and optional CPU-only NER;
no generative model decides the result. Its `sensitive` or `non_sensitive` assessments remain an
unsaved draft until the human reviews and saves any chosen flag changes, including a downgrade to
non-sensitive. Coverage is reported separately; interrupted history may count unprocessed columns.
Each started analysis records a separate Sensitivity Analysis Run with aggregate counters and safe
ordered events. The progress drawer opens before the synchronous request completes, polls the run,
and displays sanitized source-scan and local-NER phase/batch activity while classification is in
progress. This operational history never stores per-column assessments, source values,
matched spans, prompts, or free-form diagnostics. Saving a sensitive decision persists a sanitized
Sensitivity Reason as column Catalog Metadata alongside the human-owned flag; clearing the flag
clears that reason. Reloading still discards an unsaved review draft.
For unprotected columns, up to five source rows and five representative non-null values may be sent
transiently to the configured model provider. Protected columns are omitted from source reads and
replaced in the prompt by deterministic plausible values derived only from their metadata. Existing
@@ -164,7 +208,8 @@ available only when no local start, worker, or helper is live. Runs remain inspe
live SSE log with ordered polling fallback; there is no automatic resume or user-facing generation
CLI. ADRs 0009–0010 record the runtime and source-sampling decisions.
Metadata-generation setup accepts the protected `DEEPSEEK_API_KEY` and `ZAI_API_KEY` references.
The Installation Model Catalog accepts the protected `DEEPSEEK_API_KEY` and `ZAI_API_KEY`
references for metadata-generation providers.
It also accepts a model with no secret reference only when its OpenAI-compatible endpoint is
explicit; this covers the VPN-only AritmoLab Qwen 3.6 server without creating a fake operator
credential. The Python client supplies only its fixed non-secret compatibility placeholder.
+25 -35
View File
@@ -106,15 +106,18 @@ a remote user's partial list. The isolated deployment exercise is
`./scripts/verify-workspace-install-docs.sh --profile local` or `--profile server`.
<!-- workspace-descriptor-contract:start -->
Schema v3 is the only accepted workspace descriptor. Schema v1 and v2 workspace descriptors are
rejected before activation. Candidate snapshot validation therefore makes activation or a pull fail
atomically while the prior valid snapshot remains active. There is no in-product migrator or
automatic conversion. A repository must already contain reviewed v3 descriptors. One workspace
owns one Qdrant collection;
Schema v4 is the only accepted workspace descriptor. Schema v1, v2, and v3 workspace descriptors
are rejected before activation. Candidate snapshot validation therefore makes activation or a pull
fail atomically while the prior valid snapshot remains active. One workspace owns one Qdrant collection;
schema, Evidence, and Memory records share that collection and stay separated by indexed payload
`kind`.
<!-- workspace-descriptor-contract:end -->
<!-- non-workspace-migration:start -->
Convert a v3 descriptor before publication by setting `workspace.schema_version` to `4` and
removing `llm_policy` and `semantic_index`; no database or Evidence field changes.
<!-- non-workspace-migration:end -->
For NL→SQL runtime sessions, connector `ssh_tunnel` bindings remain diagnostic-only: their bounded
probe cleans up the loopback forward and returns `workspace_not_activatable`; session creation is
rejected before persistence. Database management is a separate boundary and supports a strict
@@ -176,10 +179,10 @@ secret files, upstream-auth checks, and a fail-closed `503` assertion for its de
unavailable disposable session endpoint. No real provider, database credential, or repository
secret is required.
For a clean server bind, `scripts/prepare-server-pi-state.sh` creates the hidden regular
`agent/auth.json`, `agent/models.json`, and `agent/settings.json` mount targets atomically before
Compose. The server smoke starts from an empty Pi-state root and applies this same preflight; the
real protected/tracked sources remain separate read-only mounts. Deterministic fixture tests render
For a clean server bind, `scripts/prepare-server-pi-state.sh` creates the hidden regular Pi agent
mount targets atomically before Compose. The auth target receives the protected credential bind;
the model and settings targets receive generated read-only projections. The server smoke starts
from an empty Pi-state root and applies this same preflight. Deterministic fixture tests render
both profiles, verify that bindings stay on `core`, check mount readability, and run the production
workspace resolver. Wrong-service, wrong-value, and broken-secret-mount mutations must fail.
@@ -189,7 +192,7 @@ an independent 32-minute outer timeout and does not retry a failed command.
Current release status (2026-08-05): clean-root render/setup and the production runtime-binding
resolver contracts are green. The server fixture supplies all four private trusted claims,
including exact non-admin value `0`, and a focused test proves nginx normalization produces the
accepted non-admin backend principal. Canonical schema-v3 registry descriptors now pass through
accepted non-admin backend principal. Canonical schema-v4 registry descriptors now pass through
one backend-owned, secret-safe runtime handoff for inventory and session execution; canonical
identity and durable session/artifact/index roots are retained. The fresh update-only smoke passed
bad-candidate mutation, automatic `rolled_back` compensation, exact prior-image restoration,
@@ -269,7 +272,7 @@ Compose project name by passing `--confirm-project`:
The restore script stops `qdrant`, validates the exact labeled target, stages the current volume
contents for rollback, extracts the requested archive into the volume, and then returns the
service to its prior running state. It restores semantic storage only. Before reopening write
traffic, the workspace registry must already be at a reviewed v3 descriptor revision compatible
traffic, the workspace registry must already be at a reviewed v4 descriptor revision compatible
with the restored collection; then run backend health checks and a known retrieval query. The
helper does not restore descriptors, rename collections, or reconcile an incompatible collection
contract.
@@ -295,14 +298,12 @@ Copy `deploy/secrets/thothii.secrets.example` to a protected host file, include
keys, and set its absolute path as `THT_SECRETS_FILE` in the operator env. Keep Pi's native
provider auth in the separate protected file named by `PI_AUTH_FILE`.
Description Generation is configured independently in the protected installation descriptor under
`metadataGeneration`. Set `THT_INSTALLATION_CONFIG_SOURCE` to that exact host file; Compose mounts
it read-only into `core` and supplies the fixed runtime `THT_INSTALLATION_CONFIG_FILE` path. Each
keyed model stores only an audited `apiKeyEnv` reference. The referenced value stays in the secret
bundle; a model may omit `apiKeyEnv` only when it declares an explicit endpoint that accepts
unauthenticated requests. The browser receives only model IDs, labels, and the configured default.
Configuration changes take effect after restart and do not use Pi settings or workspace
`llm_policy`.
Interactive sessions, Description Generation, and embedding share the protected installation
descriptor's `modelCatalog`. Set `THT_INSTALLATION_CONFIG_SOURCE` to that exact host file; `tht`
validates it and generates the runtime catalog, Pi adapters, and Compose override before startup.
Each authenticated provider stores only an audited `apiKeyEnv` reference; the referenced value stays
in the secret bundle. A provider may use `authentication.mode: none` only with an explicit keyless
endpoint. The browser receives only eligible model IDs, labels, and the catalog default.
Before enabling Description Generation, approve the selected model provider for bounded source-data
disclosure. Every catalog column has a **Sensitive** flag that defaults to `false`. Administrators can
@@ -333,22 +334,11 @@ the host/secret-manager materialization and add a reviewed Compose override that
does not create that mount. The frontend remains on loopback; the authenticated host proxy is the
only public listener.
Set the selected model provider in application settings (or `PI_PROVIDER`). For each Pi spawn the
backend validates and reads `THT_MODEL_API_KEY` from the bundle, then exposes its value only as the provider's
recognized child variable (for example `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, or
`ZAI_API_KEY`). Neither the generic file path nor deprecated `PI_PROVIDER_API_KEY` is inherited by
Pi. Local providers such as Ollama require no model key.
`THT_MODEL_API_KEY` supports Pi providers whose authentication is exactly one key:
`ant-ling`, `anthropic`, `cerebras`, `deepseek`, `fireworks`, `github-copilot`, `google`
(including the `gemini` alias), `google-vertex` when using its API-key mode, `groq`,
`huggingface`, `kimi-coding`, `minimax`, `minimax-cn`, `mistral`, `moonshotai`,
`moonshotai-cn`, `nvidia`, `openai`, `opencode`, `opencode-go`, `openrouter`, `together`,
`vercel-ai-gateway`, `xai`, the four `xiaomi*` providers, `zai`, and `zai-coding-cn`.
Compound providers are deliberately unsupported: `amazon-bedrock`, `azure-openai-responses`,
`cloudflare-workers-ai`, and `cloudflare-ai-gateway` require multiple credential/configuration
values. Selecting one fails before Pi starts; ambient AWS, Azure, and Cloudflare credentials are
still scrubbed. Supporting them requires a future dedicated provider-specific configuration.
For each Pi spawn, the backend resolves the selected canonical provider/model in the runtime catalog,
reads exactly that provider's declared `apiKeyEnv` value from the bundle, and exposes only that key
to the child. Ambient provider credentials and secret-bundle paths are scrubbed. Providers needing a
compound credential bundle remain unsupported until the catalog gains an explicit generic contract
for them.
## User-owned session server cutover
+84 -19
View File
@@ -12,14 +12,17 @@
"@types/pg": "^8.20.3",
"fastify": "^5.0.0",
"kysely": "^0.29.5",
"libphonenumber-js": "1.13.12",
"openid-client": "6.8.5",
"pg": "^8.22.0",
"validator": "13.15.35",
"yaml": "^2.9.0",
"zod": "^4.4.3"
},
"devDependencies": {
"@testcontainers/postgresql": "^12.1.0",
"@types/node": "24.13.3",
"@types/validator": "13.15.10",
"tsx": "^4.19.0",
"typescript": "^5.6.0",
"vitest": "^2.1.0"
@@ -1329,6 +1332,13 @@
"dev": true,
"license": "MIT"
},
"node_modules/@types/validator": {
"version": "13.15.10",
"resolved": "https://registry.npmjs.org/@types/validator/-/validator-13.15.10.tgz",
"integrity": "sha512-T8L6i7wCuyoK8A/ZeLYt1+q0ty3Zb9+qbSSvrIVitzT3YjZqkTZ40IbRsPanlB4h1QB3JVL1SYCdR6ngtFYcuA==",
"dev": true,
"license": "MIT"
},
"node_modules/@vitest/expect": {
"version": "2.1.9",
"resolved": "https://registry.npmjs.org/@vitest/expect/-/expect-2.1.9.tgz",
@@ -2458,9 +2468,9 @@
}
},
"node_modules/fast-uri": {
"version": "3.1.5",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-3.1.5.tgz",
"integrity": "sha512-gHwA1O9LDIcKunMKhObS/HimwtehO1nPUECKAu5TpKgaO19fcWEl4bliWe1jWxVFvIXztJjjQ4L8XQ1EU9f7Jw==",
"version": "3.1.7",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-3.1.7.tgz",
"integrity": "sha512-dOvZVzjdZdz7phd9v6jCbwxrBW3fK6n8Rc0CtdmM4bumzMnxywBYhuph6J819RRw/ku+rLbelwfMunktuzVVHg==",
"funding": [
{
"type": "github",
@@ -2474,9 +2484,9 @@
"license": "BSD-3-Clause"
},
"node_modules/fastify": {
"version": "5.8.5",
"resolved": "https://registry.npmjs.org/fastify/-/fastify-5.8.5.tgz",
"integrity": "sha512-Yqptv59pQzPgQUSIm87hMqHJmdkb1+GPxdE6vW6FRyVE9G86mt7rOghitiU4JHRaTyDUk9pfeKmDeu70lAwM4Q==",
"version": "5.12.3",
"resolved": "https://registry.npmjs.org/fastify/-/fastify-5.12.3.tgz",
"integrity": "sha512-reZ8wce5VNCcufIt9AVtzZa3L4u1j8esikn7OEgHWLVpRpL5R7Y2+Xzj70OUkv5zDfzUAxXZT6cu4Rt0zr3EKA==",
"funding": [
{
"type": "github",
@@ -2495,11 +2505,11 @@
"@fastify/proxy-addr": "^5.0.0",
"abstract-logging": "^2.0.1",
"avvio": "^9.0.0",
"fast-json-stringify": "^6.0.0",
"find-my-way": "^9.0.0",
"fast-json-stringify": "^7.0.0",
"find-my-way": "^9.6.0",
"light-my-request": "^6.0.0",
"pino": "^9.14.0 || ^10.1.0",
"process-warning": "^5.0.0",
"process-warning": "^5.1.0",
"rfdc": "^1.3.1",
"secure-json-parse": "^4.0.0",
"semver": "^7.6.0",
@@ -2522,6 +2532,46 @@
],
"license": "MIT"
},
"node_modules/fastify/node_modules/fast-json-stringify": {
"version": "7.0.1",
"resolved": "https://registry.npmjs.org/fast-json-stringify/-/fast-json-stringify-7.0.1.tgz",
"integrity": "sha512-eRSayARSbbwlBjpP4vnTTIRD5QPcIrmihPxDeN1DtKnHPg66UuJLx+8hlK1kaFdjvzyQ/dzALoi4vwAQ+T+iZA==",
"funding": [
{
"type": "github",
"url": "https://github.com/sponsors/fastify"
},
{
"type": "opencollective",
"url": "https://opencollective.com/fastify"
}
],
"license": "MIT",
"dependencies": {
"@fastify/merge-json-schemas": "^0.2.0",
"ajv": "^8.12.0",
"ajv-formats": "^3.0.1",
"fast-uri": "^4.0.0",
"json-schema-ref-resolver": "^3.0.0",
"rfdc": "^1.2.0"
}
},
"node_modules/fastify/node_modules/fast-uri": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-4.1.4.tgz",
"integrity": "sha512-dODXrIxlS9JSdgAnhIUKOosKV1oMtU2VtVw87QRaHzyl5jxO290Ii5tEZfCfzfWNHi3jKWwBSdQj0qIyshdZdQ==",
"funding": [
{
"type": "github",
"url": "https://github.com/sponsors/fastify"
},
{
"type": "opencollective",
"url": "https://opencollective.com/fastify"
}
],
"license": "BSD-3-Clause"
},
"node_modules/fastq": {
"version": "1.20.1",
"resolved": "https://registry.npmjs.org/fastq/-/fastq-1.20.1.tgz",
@@ -2824,6 +2874,12 @@
"safe-buffer": "~5.1.0"
}
},
"node_modules/libphonenumber-js": {
"version": "1.13.12",
"resolved": "https://registry.npmjs.org/libphonenumber-js/-/libphonenumber-js-1.13.12.tgz",
"integrity": "sha512-uLVeV1c9OTk6qkdqnj+mpMD+ZdnZ0szVyWu58HwMmpwkHA1gCEkyjd3veZQXDnuw9KEwSRjcc9B1pS9XKIN1fA==",
"license": "MIT"
},
"node_modules/light-my-request": {
"version": "6.6.0",
"resolved": "https://registry.npmjs.org/light-my-request/-/light-my-request-6.6.0.tgz",
@@ -2971,9 +3027,9 @@
"optional": true
},
"node_modules/nanoid": {
"version": "3.3.15",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.15.tgz",
"integrity": "sha512-y7Wygv/7mEOvxTuEQDB8StXdMRBWf1kR/tlhAzBRUFkB2jfcLOAxO/SHmOO2zgz1pVgK29/kyupn059/bCHdjA==",
"version": "3.3.18",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.18.tgz",
"integrity": "sha512-DTg4MJbGMWkfi6VZFdNt2/caMbQy4Ou+Op/hJQvGEWcnVfoA1QA+xzRKAzw9jD6+GVOOeYr/mIcuDSdug6F6+w==",
"dev": true,
"funding": [
{
@@ -3225,9 +3281,9 @@
"license": "MIT"
},
"node_modules/postcss": {
"version": "8.5.15",
"resolved": "https://registry.npmjs.org/postcss/-/postcss-8.5.15.tgz",
"integrity": "sha512-FfR8sjd4em2T6fb3I2MwAJU7HWVMr9zba+enmQeeWFfCbm+UOC/0X4DS8XtpUTMwWMGbjKYP7xjfNekzyGmB3A==",
"version": "8.5.28",
"resolved": "https://registry.npmjs.org/postcss/-/postcss-8.5.28.tgz",
"integrity": "sha512-RRuzqDtt5Y9h3quz5hWhK+TPnsmVs6WwSU6LkJMeY4HstUEDuYTG8UJSdawMRzmzAtV+KEoG8N3Qg2qLy5vM/A==",
"dev": true,
"funding": [
{
@@ -3245,7 +3301,7 @@
],
"license": "MIT",
"dependencies": {
"nanoid": "^3.3.12",
"nanoid": "^3.3.18",
"picocolors": "^1.1.1",
"source-map-js": "^1.2.1"
},
@@ -3310,9 +3366,9 @@
"license": "MIT"
},
"node_modules/process-warning": {
"version": "5.0.0",
"resolved": "https://registry.npmjs.org/process-warning/-/process-warning-5.0.0.tgz",
"integrity": "sha512-a39t9ApHNx2L4+HBnQKqxxHNs1r7KF+Intd8Q/g1bUh6q0WIp9voPXJ/x0j+ZL45KF1pJd9+q2jLIRMfvEshkA==",
"version": "5.1.0",
"resolved": "https://registry.npmjs.org/process-warning/-/process-warning-5.1.0.tgz",
"integrity": "sha512-jQSaVHsPgtyw60e1rQ/A+/ArPEj/S8pS/vFnyGa/gYFXrKk/6RuDkoqVDQ5NI5MmS01698ltlAk0NoDBNLujRw==",
"funding": [
{
"type": "github",
@@ -4121,6 +4177,15 @@
"dev": true,
"license": "MIT"
},
"node_modules/validator": {
"version": "13.15.35",
"resolved": "https://registry.npmjs.org/validator/-/validator-13.15.35.tgz",
"integrity": "sha512-TQ5pAGhd5whStmqWvYF4OjQROlmv9SMFVt37qoCBdqRffuuklWYQlCNnEs2ZaIBD1kZRNnikiZOS1eqgkar0iw==",
"license": "MIT",
"engines": {
"node": ">= 0.10"
}
},
"node_modules/vite": {
"version": "5.4.21",
"resolved": "https://registry.npmjs.org/vite/-/vite-5.4.21.tgz",
+6 -1
View File
@@ -7,9 +7,11 @@
"prebuild": "node scripts/clean-dist.mjs",
"build": "tsc -p tsconfig.json",
"catalog:migrate": "node dist/catalog/migrate.js",
"sensitivity:shadow": "node dist/catalog/sensitivity-shadow.js",
"test": "vitest run",
"start": "node dist/server.js",
"test:schema-v3-verifier": "python3 -I -B scripts/test_revision_state_policy.py && node --test scripts/verify-workspace-descriptor-files.test.mjs scripts/revision-state-policy.test.mjs"
"test:schema-v4-verifier": "python3 -I -B scripts/test_revision_state_policy.py && node --test scripts/verify-workspace-descriptor-files.test.mjs scripts/revision-state-policy.test.mjs",
"test:schema-v3-verifier": "npm run test:schema-v4-verifier"
},
"dependencies": {
"@fastify/cookie": "11.1.2",
@@ -18,14 +20,17 @@
"@types/pg": "^8.20.3",
"fastify": "^5.0.0",
"kysely": "^0.29.5",
"libphonenumber-js": "1.13.12",
"openid-client": "6.8.5",
"pg": "^8.22.0",
"validator": "13.15.35",
"yaml": "^2.9.0",
"zod": "^4.4.3"
},
"devDependencies": {
"@testcontainers/postgresql": "^12.1.0",
"@types/node": "24.13.3",
"@types/validator": "13.15.10",
"tsx": "^4.19.0",
"typescript": "^5.6.0",
"vitest": "^2.1.0"
@@ -0,0 +1,8 @@
34448b82c17d60fec9b65b1f093c115ddbaadc04beb1b0140b6bfed2e012a930 ./.gitattributes
4d9344c58a2a2ea4bb4ff4f7c611a853cf413205fc10d0cace564eba06f73828 ./README.md
180f0a10d1d5ed5ce3318db0bcb0b1b7780d79a52f0a8fc3acbd27f74536d0e4 ./THOTHII_MODEL_REVISION
164f17362bcf9d114067d3465e7374bfdd79ce6b605acb745de5a49dabb9595c ./config.json
f27dd63cc43a248d2566f0b6ad7a115db353676ce0561dcbca45bac766464c1a ./encoder_config/config.json
0280f6f39f6012da50b6640bad438d9b7e763a1b0102094115d1b710c4dd79b6 ./model.safetensors
f6df10ec83bea993035b2dd7c39345a3d4fcf23421c2adb6cb4ffc1e6d1bc4b5 ./tokenizer.json
233beed1f1095cccfc7907cde31a8d90a0c6aa4fdfaf6493f8e55fd162e81ae6 ./tokenizer_config.json
@@ -0,0 +1,34 @@
# Optional offline CPU pack. Fully version-locked in its own venv; not part of the base image.
--extra-index-url https://download.pytorch.org/whl/cpu
accelerate==1.14.0
annotated-types==0.8.0
certifi==2026.7.22
charset-normalizer==3.5.1
filelock==3.32.5
fsspec==2026.7.0
gliner2[local]==2.0.0
hf-xet==1.6.0
huggingface-hub==0.36.2
idna==3.19
Jinja2==3.1.6
MarkupSafe==3.0.3
mpmath==1.3.0
networkx==3.6.1
numpy==2.5.2
packaging==26.3
peft==0.20.0
psutil==7.2.2
pydantic==2.13.5
pydantic-core==2.46.5
PyYAML==6.0.3
regex==2026.9.3
requests==2.34.2
safetensors==0.8.0
sympy==1.14.0
tokenizers==0.22.2
torch==2.14.0+cpu
tqdm==4.70.0
transformers==4.57.6
typing-extensions==4.16.0
typing-inspection==0.4.4
urllib3==2.7.0
+301
View File
@@ -0,0 +1,301 @@
"""Offline, CPU-only JSONL worker for optional sensitivity NER evidence."""
from __future__ import annotations
import argparse
import contextlib
import ctypes
import errno
import hashlib
import json
import os
import socket
import sys
import tempfile
from pathlib import Path
from typing import Any
PII_LABELS = [
"person",
"full_name",
"first_name",
"middle_name",
"last_name",
"date_of_birth",
"email",
"phone_number",
"address",
"street_address",
"city",
"state_or_region",
"postal_code",
"country",
"government_id",
"national_id_number",
"passport_number",
"drivers_license_number",
"license_number",
"tax_id",
"tax_number",
"bank_account",
"account_number",
"routing_number",
"iban",
"payment_card",
"card_number",
"card_expiry",
"card_cvv",
"username",
"ip_address",
"account_id",
"sensitive_account_id",
"password",
"secret",
"api_key",
"access_token",
"recovery_code",
"sensitive_date",
"document_date",
"expiration_date",
"transaction_date",
]
_MODEL_COMPAT_DIRECTORY: tempfile.TemporaryDirectory[str] | None = None
_EXPECTED_MODEL_REVISION = "c153999da5f4c509df4322b0c6a1baf3d2c284d7"
def _arguments() -> argparse.Namespace:
parser = argparse.ArgumentParser(add_help=False)
parser.add_argument("--model", required=True)
parser.add_argument("--threads", type=int, default=2)
return parser.parse_args()
def _disable_network() -> None:
libc = ctypes.CDLL(None, use_errno=True)
libc.prctl.argtypes = [
ctypes.c_int,
ctypes.c_ulong,
ctypes.c_ulong,
ctypes.c_ulong,
ctypes.c_ulong,
]
libc.prctl.restype = ctypes.c_int
if libc.prctl(38, 1, 0, 0, 0) != 0: # PR_SET_NO_NEW_PRIVS
raise RuntimeError("cannot enable no-new-privileges for network isolation")
try:
seccomp = ctypes.CDLL("libseccomp.so.2", use_errno=True)
except OSError as error:
raise RuntimeError("libseccomp is required for network isolation") from error
seccomp.seccomp_init.argtypes = [ctypes.c_uint32]
seccomp.seccomp_init.restype = ctypes.c_void_p
seccomp.seccomp_syscall_resolve_name.argtypes = [ctypes.c_char_p]
seccomp.seccomp_syscall_resolve_name.restype = ctypes.c_int
seccomp.seccomp_rule_add.argtypes = [
ctypes.c_void_p,
ctypes.c_uint32,
ctypes.c_int,
ctypes.c_uint,
]
seccomp.seccomp_rule_add.restype = ctypes.c_int
seccomp.seccomp_load.argtypes = [ctypes.c_void_p]
seccomp.seccomp_load.restype = ctypes.c_int
seccomp.seccomp_release.argtypes = [ctypes.c_void_p]
seccomp.seccomp_release.restype = None
allow = 0x7FFF0000 # SCMP_ACT_ALLOW
deny = 0x00050000 | errno.EPERM # SCMP_ACT_ERRNO(EPERM)
filter_context = seccomp.seccomp_init(allow)
if not filter_context:
raise RuntimeError("cannot initialize network syscall filter")
try:
for syscall in (
"socket",
"connect",
"sendto",
"sendmsg",
"sendmmsg",
"bind",
"listen",
"accept",
"accept4",
):
syscall_number = seccomp.seccomp_syscall_resolve_name(syscall.encode("ascii"))
if syscall_number < 0:
raise RuntimeError(f"cannot resolve network syscall: {syscall}")
if seccomp.seccomp_rule_add(filter_context, deny, syscall_number, 0) != 0:
raise RuntimeError(f"cannot block network syscall: {syscall}")
if seccomp.seccomp_load(filter_context) != 0:
raise RuntimeError("cannot activate network syscall filter")
finally:
seccomp.seccomp_release(filter_context)
def blocked(*_args: Any, **_kwargs: Any) -> Any:
raise PermissionError(errno.EPERM, "network disabled")
socket.socket = blocked # type: ignore[assignment]
socket.create_connection = blocked # type: ignore[assignment]
def _verify_model(path: Path) -> None:
revision_path = path / "THOTHII_MODEL_REVISION"
try:
revision = revision_path.read_text(encoding="utf-8").strip()
except OSError as error:
raise RuntimeError("model revision marker is unavailable") from error
if revision != _EXPECTED_MODEL_REVISION:
raise RuntimeError("model revision is not approved")
manifest_path = Path(__file__).with_name("sensitivity-ner-model-sha256.txt")
try:
manifest = manifest_path.read_text(encoding="utf-8").splitlines()
except OSError as error:
raise RuntimeError("model checksum manifest is unavailable") from error
for line in manifest:
checksum, separator, relative_name = line.partition(" ")
if not separator or len(checksum) != 64 or not relative_name.startswith("./"):
raise RuntimeError("model checksum manifest is invalid")
relative_path = Path(relative_name[2:])
if relative_path.is_absolute() or ".." in relative_path.parts:
raise RuntimeError("model checksum path is invalid")
model_file = path / relative_path
if not model_file.is_file() or model_file.is_symlink():
raise RuntimeError("approved model file is unavailable")
digest = hashlib.sha256()
with model_file.open("rb") as stream:
for chunk in iter(lambda: stream.read(1024 * 1024), b""):
digest.update(chunk)
if digest.hexdigest() != checksum:
raise RuntimeError("approved model checksum does not match")
def _transformers4_model_path(path: Path) -> Path:
"""Adapt tokenizer metadata emitted by Transformers 5 without changing pinned weights.
GLiNER2 2.0.0 officially requires Transformers <5, while current Fastino checkpoints were
saved by Transformers 5.8.0. Transformers 4 calls the same list
``additional_special_tokens``; Transformers 5 renamed it to ``extra_special_tokens`` and
changed its type. Keep the downloaded model immutable and create a temporary symlink view
containing only the compatibility metadata needed by the supported GLiNER2 dependency set.
"""
tokenizer_path = path / "tokenizer_config.json"
try:
tokenizer = json.loads(tokenizer_path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError) as error:
raise RuntimeError("invalid tokenizer configuration") from error
extra_tokens = tokenizer.get("extra_special_tokens")
if extra_tokens is None:
return path
if not isinstance(extra_tokens, list) or not all(isinstance(token, str) for token in extra_tokens):
raise RuntimeError("unsupported extra_special_tokens configuration")
if "additional_special_tokens" in tokenizer:
raise RuntimeError("ambiguous special-token configuration")
global _MODEL_COMPAT_DIRECTORY
_MODEL_COMPAT_DIRECTORY = tempfile.TemporaryDirectory(prefix="thothii-ner-model-")
compatible_path = Path(_MODEL_COMPAT_DIRECTORY.name)
for child in path.iterdir():
if child.name == tokenizer_path.name:
continue
(compatible_path / child.name).symlink_to(child, target_is_directory=child.is_dir())
tokenizer["additional_special_tokens"] = tokenizer.pop("extra_special_tokens")
(compatible_path / tokenizer_path.name).write_text(
json.dumps(tokenizer, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
return compatible_path
def _load_model(model_path: str, threads: int) -> Any:
path = Path(model_path).resolve(strict=True)
if not path.is_dir():
raise RuntimeError("model path must be a local directory")
_verify_model(path)
os.environ["CUDA_VISIBLE_DEVICES"] = ""
os.environ["HIP_VISIBLE_DEVICES"] = ""
os.environ["HF_HUB_OFFLINE"] = "1"
os.environ["TRANSFORMERS_OFFLINE"] = "1"
import torch
from gliner2 import AutoExtractor
torch.set_num_threads(max(1, min(threads, 8)))
torch.set_num_interop_threads(1)
compatible_path = _transformers4_model_path(path)
with contextlib.redirect_stdout(sys.stderr):
model = AutoExtractor.from_pretrained(str(compatible_path), map_location="cpu")
_disable_network()
return model
def _request(value: Any) -> tuple[str, list[dict[str, str]]]:
if not isinstance(value, dict) or not isinstance(value.get("id"), str):
raise ValueError("invalid request")
candidates = value.get("candidates")
if not isinstance(candidates, list) or not 1 <= len(candidates) <= 128:
raise ValueError("invalid candidates")
parsed: list[dict[str, str]] = []
for candidate in candidates:
if not isinstance(candidate, dict):
raise ValueError("invalid candidate")
column_id = candidate.get("columnId")
text = candidate.get("text")
if not isinstance(column_id, str) or not isinstance(text, str) or not 1 <= len(text) <= 500:
raise ValueError("invalid candidate")
parsed.append({"columnId": column_id, "text": text})
return value["id"], parsed
def _detect(model: Any, candidates: list[dict[str, str]]) -> list[dict[str, Any]]:
evidence: list[dict[str, Any]] = []
for candidate in candidates:
result = model.extract_entities(
candidate["text"],
PII_LABELS,
threshold=0.5,
include_confidence=True,
)
entities = result.get("entities", {}) if isinstance(result, dict) else {}
best: tuple[str, float] | None = None
if isinstance(entities, dict):
for label, matches in entities.items():
if label not in PII_LABELS or not isinstance(matches, list):
continue
for match in matches:
if not isinstance(match, dict):
continue
confidence = match.get("confidence")
if not isinstance(confidence, (int, float)) or not 0 <= confidence <= 1:
continue
if best is None or confidence > best[1]:
best = (label, float(confidence))
if best is not None:
evidence.append(
{
"columnId": candidate["columnId"],
"label": best[0],
"confidence": best[1],
}
)
return evidence
def main() -> int:
args = _arguments()
model = _load_model(args.model, args.threads)
print(json.dumps({"ready": True}, separators=(",", ":")), flush=True)
for line in sys.stdin:
request_id = "invalid"
try:
request_id, candidates = _request(json.loads(line))
response = {"id": request_id, "ok": True, "evidence": _detect(model, candidates)}
except Exception:
response = {"id": request_id, "ok": False, "error": "detection_failed"}
print(json.dumps(response, separators=(",", ":")), flush=True)
return 0
if __name__ == "__main__":
raise SystemExit(main())
+1 -6
View File
@@ -1195,13 +1195,8 @@ export async function executeChecks({ checks, failAt, recorder } = {}) {
function baseWorkspace(id, evidenceSource) {
return {
workspace: { schema_version: 3, id, name: `P1 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P1 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "postgres", schema: "public", supported_transports: ["postgres_direct"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } },
};
}
+1 -1
View File
@@ -109,7 +109,7 @@ async function validateDistFiles(repo,files){const dist=join(repo,"backend","dis
export async function readManualOwnership({repositoryRoot=defaultRepositoryRoot}={}){const repo=realpathSync(repositoryRoot),root=fixedManualRoot(repo);noSymlinkExisting(repo,root);let rootEntry,ownershipEntry;try{rootEntry=await lstat(root);ownershipEntry=await lstat(join(root,"ownership.json"));}catch{throw new Error("manual ownership is missing");}if(!rootEntry.isDirectory()||rootEntry.isSymbolicLink()||await realpath(root)!==root||!ownershipEntry.isFile()||ownershipEntry.isSymbolicLink())throw new Error("manual ownership is unsafe");let value;try{value=JSON.parse(await readFile(join(root,"ownership.json"),"utf8"));}catch{throw new Error("manual ownership is malformed");}const baseValid=value.schemaVersion===1&&value.kind==="p1-manual-acceptance"&&HEX64.test(value.nonce??"")&&value.repositoryRoot===repo&&value.root===root&&value.status==="PENDING"&&["PREPARING","READY"].includes(value.stage)&&value.listener?.host===HOST&&value.listener?.port===PORT&&value.listener?.state==="stopped"&&typeof value.createdAt==="string"&&validEntrypoint(value.entrypoint,repo)&&validDistManifest(value.distManifest,root)&&JSON.stringify(value.resources)===JSON.stringify([root,{kind:"fastify",host:HOST,port:PORT}]);const readyLog=value.backendLog?.path===join(root,"logs/backend.log")&&Number.isSafeInteger(value.backendLog?.dev)&&Number.isSafeInteger(value.backendLog?.ino);if(!baseValid||(value.stage==="READY"?!readyLog:value.backendLog!==null))throw new Error("manual ownership identity mismatch");return value;}
async function run(executable,argv,options={}){return await exec(executable,argv,{...options,maxBuffer:2*1024*1024,encoding:"utf8"});}
function descriptor(id,source){return{workspace:{schema_version:3,id,name:`P1 ${id}`,language:"en"},dwh:{engine:"postgres",database:"postgres",schema:"public",supported_transports:["postgres_direct"]},semantic_index:{vector_store:{engine:"qdrant",collection:id,dimensions:1024,distance:"cosine"},embedding:{provider:"ollama_internal",model:"qwen3-embedding:0.6b",dimensions:1024}},llm_policy:{allowed:["zai/glm-5.2"]},evidence:{source,policy:{max_chunk_chars:4000,retain_published_generations:3}}};}
function descriptor(id,source){return{workspace:{schema_version:4,id,name:`P1 ${id}`,language:"en"},dwh:{engine:"postgres",database:"postgres",schema:"public",supported_transports:["postgres_direct"]},evidence:{source,policy:{max_chunk_chars:4000,retain_published_generations:3}}};}
function descriptors(){return[descriptor("p1-filesystem",{type:"filesystem",uri:"workspace-content/p1-filesystem/evidence",patterns:["**/*.md"],max_bytes:10485760}),descriptor("p1-http",{type:"http",uris:["https://evidence.example.test/guide.md"],authentication:"signed_urls_file",connect_timeout_ms:1250,read_timeout_ms:30001,max_bytes:12345,max_redirects:2,allow_private_hosts:false,max_cache_bytes:67890}),descriptor("p1-s3",{type:"s3",uri:"s3://p1-evidence/published/",endpoint_url:"https://s3.example.test/",region:"eu-west-1",credentials:"static_files",trusted_endpoint:true,allow_private_endpoint:false,allow_insecure_endpoint:false,max_bytes:12345,max_objects:33,max_pages:4,page_size:5})];}
function quote(value){return `'${String(value).replaceAll("'",`'"'"'`)}'`;}
async function checkPrerequisites(repo){for(const path of ["scripts/p1-acceptance.sh","scripts/test-p1-acceptance.sh","backend/scripts/p1-acceptance.mjs","backend/dist/server.js"]){try{await access(join(repo,path));}catch{throw new Error(`Task 8 prerequisite is missing: ${path}`);}}for(const command of ["node","npm","git","curl","unzip","zipinfo","lsof","python3"]){try{await run(command,[command==="unzip"||command==="lsof"?"-v":command==="zipinfo"?"-h":"--version"]);}catch{throw new Error(`missing prerequisite: ${command}`);}}const tht=join(repo,"harness",".venv","bin","tht");try{await access(tht,constants.X_OK);}catch{throw new Error("missing prerequisite: harness/.venv/bin/tht");}}
@@ -403,7 +403,7 @@ test("generated render command validates saved responses and owned snapshot befo
});
const renderSnapshotYaml=`workspace:
schema_version: 3
schema_version: 4
id: p1-filesystem
name: P1 filesystem
language: en
@@ -412,11 +412,6 @@ dwh:
database: postgres
schema: public
supported_transports: [postgres_direct]
semantic_index:
vector_store: {engine: qdrant, collection: p1-filesystem, dimensions: 1024, distance: cosine}
embedding: {provider: ollama_internal, model: qwen3-embedding:0.6b, dimensions: 1024}
llm_policy:
allowed: [zai/glm-5.2]
evidence:
source: {type: filesystem, uri: workspace-content/p1-filesystem/evidence, patterns: ["**/*.md"], max_bytes: 10485760}
policy: {max_chunk_chars: 4000, retain_published_generations: 3}
+2 -7
View File
@@ -16,7 +16,7 @@ async function fixture() {
await writeFile(join(root,"installation/base.yaml"),"{}\n");
const secret=join(root,"fixture-secrets/dwh-password"); await writeFile(secret,"not-inspected",{mode:0o600});
await writeFile(snapshot,`workspace:
schema_version: 3
schema_version: 4
id: p1-filesystem
name: P1 filesystem
language: en
@@ -25,11 +25,6 @@ dwh:
database: postgres
schema: public
supported_transports: [postgres_direct]
semantic_index:
vector_store: {engine: qdrant, collection: p1-filesystem, dimensions: 1024, distance: cosine}
embedding: {provider: ollama_internal, model: qwen3-embedding:0.6b, dimensions: 1024}
llm_policy:
allowed: [zai/glm-5.2]
evidence:
source: {type: filesystem, uri: workspace-content/p1-filesystem/evidence, patterns: ["**/*.md"], max_bytes: 10485760}
policy: {max_chunk_chars: 4000, retain_published_generations: 3}
@@ -61,7 +56,7 @@ test("renderer refuses snapshot manifest head, digest, and expected-digest tampe
test("renderer refuses a missing or malformed snapshot manifest",async()=>{ const f=await fixture(); const output=join(f.root,"rendered/nomanifest.yaml"); await rm(f.manifestPath); await assert.rejects(call(f,{outputPath:output}),/snapshot manifest.*(missing|unbounded|unsafe)/); await writeFile(f.manifestPath,"{not json"); await assert.rejects(call(f,{outputPath:output}),/snapshot manifest.*malformed/); await assert.rejects(lstat(output)); assert.deepEqual(await runtimeLeases(f),[]); });
test("renderer rejects a regular snapshot replacement against its manifest",async()=>{ const f=await fixture(); const output=join(f.root,"rendered/replaced.yaml"); await assert.rejects(call(f,{outputPath:output,beforePublish:async()=>{await writeFile(f.snapshot,"workspace:\n schema_version: 3\n id: p1-filesystem\n name: replaced\n")}}),/snapshot content changed/); await assert.rejects(lstat(output)); });
test("renderer rejects a regular snapshot replacement against its manifest",async()=>{ const f=await fixture(); const output=join(f.root,"rendered/replaced.yaml"); await assert.rejects(call(f,{outputPath:output,beforePublish:async()=>{await writeFile(f.snapshot,"workspace:\n schema_version: 4\n id: p1-filesystem\n name: replaced\n")}}),/snapshot content changed/); await assert.rejects(lstat(output)); });
test("renderer anchors publication when rendered parent is concurrently swapped", async()=>{
const f=await fixture(),output=join(f.root,"rendered/raced.yaml"),moved=join(f.root,"rendered-moved"),outside=join(f.repo,"outside-rendered"); await mkdir(outside);
+1 -6
View File
@@ -328,13 +328,8 @@ async function tht(ctx, argv, options = {}) {
function namespace(id) { return id.toUpperCase().replaceAll("-", "_"); }
function baseWorkspace(id, evidenceSource) {
return {
workspace: { schema_version: 3, id, name: `P1.1 ${id}`, description: `Catalog entry for ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P1.1 ${id}`, description: `Catalog entry for ${id}`, language: "en" },
dwh: { engine: "postgres", database: "postgres", schema: "public", supported_transports: ["postgres_direct"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } },
};
}
+1 -6
View File
@@ -81,13 +81,8 @@ async function git(executable, argv, options = {}) {
function namespace(id) { return id.toUpperCase().replaceAll("-", "_"); }
function baseWorkspace(id, evidenceSource) {
return {
workspace: { schema_version: 3, id, name: `P1.1 ${id}`, description: `Catalog entry for ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P1.1 ${id}`, description: `Catalog entry for ${id}`, language: "en" },
dwh: { engine: "postgres", database: "postgres", schema: "public", supported_transports: ["postgres_direct"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } },
};
}
+1 -6
View File
@@ -375,16 +375,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+1 -6
View File
@@ -376,16 +376,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+1 -6
View File
@@ -379,16 +379,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+1 -6
View File
@@ -374,16 +374,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
+1 -6
View File
@@ -374,16 +374,11 @@ function installationProjectName(installationPath) {
function baseWorkspace(id, { dwhBaseUrl, evidenceSource }) {
return {
workspace: { schema_version: 3, id, name: `P2 ${id}`, language: "en" },
workspace: { schema_version: 4, id, name: `P2 ${id}`, language: "en" },
dwh: { engine: "postgres", database: "warehouse", schema: "dw", supported_transports: ["rest_api"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
diagnostics: {
dwh_rest: { method: "POST", path: "/rpc/ping", auth: "x-api-key", response: { database: "database", schema: "schema" } },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
...(evidenceSource ? { evidence: { source: evidenceSource, policy: { max_chunk_chars: 4000, retain_published_generations: 3 } } } : {}),
};
}
@@ -22,6 +22,7 @@ const reviewedExpandableBlocks = new Map([
{ sha256: "37f18ce7ce93cb8b84f3b3708462cc16d50fdc7bab22836c382dbacf8382f05f", rationale: "Generates the reviewed synthetic tht installer artifact." },
]],
["scripts/test-server-pi-state-topology.sh", [
{ sha256: "435c769b8cbd7b834f56fdabddb86ba04fb404dd0d8a6b7c21719a8b0f7cf011", rationale: "Generates the reviewed model-catalog projection override for the isolated server topology test." },
{ sha256: "6ae9567db53d6cd45a2c19c98acaf45f382450b157ea7d6f6d35125f68c50947", rationale: "Generates the isolated server topology test environment, including its installation descriptor and authentication configuration root." },
]],
["scripts/test-vector-backup-restore-safety.sh", [
@@ -36,11 +37,10 @@ const reviewedExpandableBlocks = new Map([
{ sha256: "b903e5dae953ae1372f1a5276f12a92ed3dd632b897f3afe5e00c646d90a1b42", rationale: "Same reviewed block in the repository-required CRLF checkout representation." },
]],
["scripts/unified-deployment-smoke.sh", [
{ sha256: "1d60bf140165a8fabfa0c3729e776136904717e67becf3e0ab68c70d8e37847e", rationale: "Generates reviewed Task 13 runtime configuration." },
{ sha256: "36d3d8a2362dbdc4fad90948d6c227586d749f56b9a4bc5b6b5a91bcbec6407b", rationale: "Generates the reviewed local Task 13 Compose override." },
{ sha256: "c556f7d910d0788e219b042957e6b307cb9925b43920c680535d0d3a6dcbdb25", rationale: "Generates the reviewed local Task 13 installation descriptor." },
{ sha256: "b6c0826151b2c8b955399d1abf5b691cc8fe6b6454b17da000dde7ba3bc55d2d", rationale: "Generates the reviewed local Task 13 Compose override with normalized catalog mounts." },
{ sha256: "24f69d12b8554aa2bebba455be99fde3e60743eef5a40fa2ef5b29397a477c03", rationale: "Generates the reviewed local Task 13 installation descriptor with its model catalog." },
{ sha256: "526006fa6d48a8080b3834723630c64de5005a67243e944ebf1da15212b4d654", rationale: "Generates the reviewed server Task 13 Compose override." },
{ sha256: "c57ae2205c21ead0c2015a353aaabb948fa4ddd9b78a2cdcdb71f48cf2db742d", rationale: "Generates the reviewed projected-auth server Task 13 installation descriptor." },
{ sha256: "406ccead1967f642225c946fc4a23fe5b019c9764cc5153e1125876ade16ec90", rationale: "Generates the reviewed projected-auth server Task 13 installation descriptor with its model catalog." },
]],
["scripts/vector-backup.sh", [
{ sha256: "571899db49dfdcec8107fbe1e0a86a61e7581979d3c4c248c20546843e275bcf", rationale: "Generates the reviewed backup manifest inside the helper command." },
@@ -103,6 +103,7 @@ function isPolicyImplementationException(label, category) {
]);
if (implementations.has(label)) return true;
if (category === "migration-marker" && new Set([
"backend/src/workspaces/schema.ts",
"scripts/workspace_descriptor_doc_contract.py",
"scripts/test_workspace_descriptor_doc_contract.py",
"backend/scripts/clean-dist.test.mjs",
@@ -173,7 +174,7 @@ function validateWorkspaceSource(source, label, { requireWorkspace, expandable =
try {
parseWorkspaceYaml(source);
} catch (error) {
throw new Error(`${label}: workspace descriptor is not valid schema v3: ${error instanceof Error ? error.message : String(error)}`);
throw new Error(`${label}: workspace descriptor is not valid schema v4: ${error instanceof Error ? error.message : String(error)}`);
}
return true;
}
@@ -33,20 +33,20 @@ function bashN(root, path) {
function replaceWorkspaceKeys(source, workspaceKey, schemaLine) {
return source
.replace(/^workspace:$/m, workspaceKey)
.replace(/^ schema_version: 3$/m, schemaLine);
.replace(/^ schema_version: 4$/m, schemaLine);
}
test("production parser accepts semantic v3 with quoted Unicode/tagged keys and spacing", async (t) => {
test("production parser accepts semantic v4 with quoted Unicode/tagged keys and spacing", async (t) => {
const root = await fixture(t);
const unicode = replaceWorkspaceKeys(
canonicalDescriptor,
'"\\u0077orkspace" :',
' "\\u0073chema_version" : 3',
' "\\u0073chema_version" : 4',
);
const tagged = replaceWorkspaceKeys(
canonicalDescriptor,
"!!str workspace :",
" !!str schema_version : 3",
" !!str schema_version : 4",
);
await put(root, "deploy/workspaces/unicode.yaml", unicode);
await put(root, "deploy/workspaces/tagged.yaml", tagged);
@@ -59,14 +59,15 @@ test("production parser accepts semantic v3 with quoted Unicode/tagged keys and
});
});
test("production parser rejects fancy keys with every non-v3 or ambiguous value", async (t) => {
test("production parser rejects fancy keys with every non-v4 or ambiguous value", async (t) => {
const invalid = [
["unicode-v2", '"\\u0077orkspace" :', ' "\\u0073chema_version" : 2'],
["tagged-leading-zero", "!!str workspace :", " !!str schema_version : 02"],
["hexadecimal", "workspace :", " schema_version : 0x2"],
["multiline", "workspace :", " schema_version : >\n 3"],
["duplicate", "workspace :", " schema_version : 3\n schema_version: 3"],
["inline", "workspace: { schema_version: 3 }", " schema_version: 3"],
["unicode-v3", '"\\u0077orkspace" :', ' "\\u0073chema_version" : 3'],
["tagged-leading-zero", "!!str workspace :", " !!str schema_version : 03"],
["hexadecimal", "workspace :", " schema_version : 0x3"],
["multiline", "workspace :", " schema_version : >\n 4"],
["duplicate", "workspace :", " schema_version : 4\n schema_version: 4"],
["inline", "workspace: { schema_version: 4 }", " schema_version: 4"],
];
for (const [name, workspaceKey, schemaLine] of invalid) {
await t.test(name, async () => {
@@ -120,7 +121,7 @@ test("PowerShell embedded workspace mappings are rejected while bundle-only stri
const root = await fixture(t);
const source = [
"$workspace = @'",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 0x2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 0x2").trimEnd(),
"'@",
'$bundle = @"',
"bundle:",
@@ -137,7 +138,7 @@ test("PowerShell embedded workspace mappings are rejected while bundle-only stri
test("workspace descriptor family entries require a top-level workspace", async (t) => {
const root = await fixture(t);
await put(root, "scripts/fixtures/workspace-registry-future.yaml", "bundle:\n schema_version: 3\n");
await put(root, "scripts/fixtures/workspace-registry-future.yaml", "bundle:\n schema_version: 4\n");
await assert.rejects(
verifyEntries({
root,
@@ -179,7 +180,7 @@ test("script scalar workspace remains a bundle even with descriptor-like sibling
test("standalone descriptor files require workspace to be a mapping", async (t) => {
const root = await fixture(t);
const path = "scripts/fixtures/workspace-registry-scalar.yaml";
await put(root, path, "workspace: analytics\nschema_version: 3\n");
await put(root, path, "workspace: analytics\nschema_version: 4\n");
await assert.rejects(
verifyEntries({ root, entries: [entry("workspace_descriptor", path)] }),
/workspace.*mapping/i,
@@ -193,11 +194,11 @@ test("Bash extractor supports hyphen, digit, escaped delimiters, and tab strippi
name: "hyphen-v2",
opener: "cat <<'WORKSPACE-YAML'",
delimiter: "WORKSPACE-YAML",
descriptor: canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2"),
descriptor: canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2"),
rejected: true,
},
{
name: "digit-v3",
name: "digit-v4",
opener: "cat <<2YAML",
delimiter: "2YAML",
descriptor: canonicalDescriptor,
@@ -207,11 +208,11 @@ test("Bash extractor supports hyphen, digit, escaped delimiters, and tab strippi
name: "escaped-v2",
opener: "cat <<WORKSPACE\\-YAML",
delimiter: "WORKSPACE-YAML",
descriptor: canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2"),
descriptor: canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2"),
rejected: true,
},
{
name: "tab-strip-v3",
name: "tab-strip-v4",
opener: "cat <<-'TAB-YAML'",
delimiter: "\tTAB-YAML",
descriptor: canonicalDescriptor.split("\n").map((line) => `\t${line}`).join("\n"),
@@ -272,7 +273,7 @@ test("non-stripping heredoc close requires an exact physical delimiter line", as
"#!/usr/bin/env bash",
"cat <<'---'",
"--- ",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2").trimEnd(),
"---",
"",
].join("\n");
@@ -313,7 +314,7 @@ test("double-quoted non-special backslash is preserved in the delimiter", async
"#!/usr/bin/env bash",
'cat <<"\\---"',
"---",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2").trimEnd(),
"\\---",
"",
].join("\n");
@@ -355,7 +356,7 @@ test("split heredoc operator continuation cannot bypass v2 validation", async (t
"#!/usr/bin/env bash",
"cat <\\",
"<'YAML'",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2").trimEnd(),
"YAML",
"",
].join("\n");
@@ -424,7 +425,7 @@ test("PowerShell comment backslash cannot hide a following v2 here-string", asyn
const source = [
"# harmless PowerShell comment \\",
"$workspace = @'",
canonicalDescriptor.replace(" schema_version: 3", " schema_version: 2").trimEnd(),
canonicalDescriptor.replace(" schema_version: 4", " schema_version: 2").trimEnd(),
"'@",
"",
].join("\n");
@@ -435,7 +436,7 @@ test("PowerShell comment backslash cannot hide a following v2 here-string", asyn
);
});
test("PowerShell dialect accepts normal v3 and non-workspace bundle here-strings", async (t) => {
test("PowerShell dialect accepts normal v4 and non-workspace bundle here-strings", async (t) => {
const root = await fixture(t);
const path = "scripts/powershell-valid-smoke.ps1";
const source = [
@@ -494,10 +495,10 @@ test("PowerShell cast and concatenation openers cannot hide embedded descriptors
test("expandable YAML interpolation that can hide a workspace descriptor fails closed", async (t) => {
const root = await fixture(t);
const cases = [
["braced-key", "${key}:\n schema_version: 3"],
["plain-key", "$key:\n schema_version: 3"],
["quoted-key", '"$key" :\n schema_version: 3'],
["subexpression-key", "$($key):\n schema_version: 3"],
["braced-key", "${key}:\n schema_version: 4"],
["plain-key", "$key:\n schema_version: 4"],
["quoted-key", '"$key" :\n schema_version: 4'],
["subexpression-key", "$($key):\n schema_version: 4"],
["version", "workspace:\n schema_version: $version"],
];
for (const [name, body] of cases) {
@@ -564,7 +565,7 @@ test("unmarked expandable Bash YAML cannot generate descriptor keys or values at
"key=workspace",
"cat <<YAML",
generatedKey,
" schema_version: 3",
" schema_version: 4",
"YAML",
"",
].join("\n");
@@ -600,14 +601,14 @@ test("an in-band marker cannot authorize expandable content", async (t) => {
for (const [path, source] of [
["scripts/fake-marker.sh", [
"#!/usr/bin/env bash",
"# schema-v3-only: expandable-nonworkspace",
"# schema-v4-only: expandable-nonworkspace",
"cat <<YAML",
"${DESCRIPTOR}",
"YAML",
"",
].join("\n")],
["scripts/fake-marker.ps1", [
"# schema-v3-only: expandable-nonworkspace",
"# schema-v4-only: expandable-nonworkspace",
'$yaml = @"',
"$descriptor",
'"@',
+65 -16
View File
@@ -3,6 +3,7 @@ import cors from "@fastify/cors";
import cookie from "@fastify/cookie";
import rateLimit from "@fastify/rate-limit";
import { dirname, isAbsolute, join } from "node:path";
import { fileURLToPath } from "node:url";
import { tmpdir } from "node:os";
import type { AppConfig } from "./config.js";
import { ThtRunner } from "./tht/tht-runner.js";
@@ -22,7 +23,8 @@ import { isUsableAuthenticationSecret } from "./auth/secret-policy.js";
import { secretValue } from "./config/secret-bundle.js";
import { sessionRoutes } from "./routes/sessions.js";
import { sqlRoutes } from "./routes/sql.js";
import { metaRoutes, type ListModelsFn } from "./routes/meta.js";
import { metaRoutes } from "./routes/meta.js";
import type { ListModelsFn } from "./pi/list-models.js";
import { settingsRoutes, effectiveSettings } from "./routes/settings.js";
import { createPiModelLister } from "./pi/list-models.js";
import { createPiManagement, type PiManagementService } from "./pi/management.js";
@@ -31,7 +33,11 @@ import { ReadinessManager } from "./runtime/readiness-manager.js";
import { MaintenanceBarrier } from "./runtime/maintenance-gate.js";
import { WorkspaceRegistry } from "./workspaces/registry.js";
import { createProductionWorkspaceDiagnoser } from "./workspaces/diagnostics.js";
import { workspaceRoutes, type WorkspaceDiagnoser } from "./routes/workspaces.js";
import {
workspaceRoutes,
type WorkspaceDatabaseTester,
type WorkspaceDiagnoser,
} from "./routes/workspaces.js";
import { piManagementRoutes } from "./routes/pi-management.js";
import { supportsSessionRuntime } from "./workspaces/bindings.js";
import { resolveRuntimeBindingsWithWorkspaceSecrets } from "./workspaces/secret-requirements.js";
@@ -56,8 +62,11 @@ import { metadataGenerationModelRoutes } from "./routes/metadata-generation-mode
import { catalogDescriptionConsolidationRoutes } from "./routes/catalog-description-consolidation.js";
import { PythonModelCompleter, type ModelCompleter } from "./catalog/model-completer.js";
import { DescriptionGenerationWorker } from "./catalog/description-generation-worker.js";
import { SensitiveDataSuggester } from "./catalog/sensitive-data-suggester.js";
import { SensitiveDataSuggestionRunner } from "./catalog/sensitive-data-suggestion-runner.js";
import { SensitivityAnalysisService } from "./catalog/sensitivity-analysis-service.js";
import { SensitivityAnalysisRunner } from "./catalog/sensitivity-analysis-runner.js";
import { SensitivityClassifier, type LocalNerDetector, type SensitivityValueSource } from "./catalog/sensitivity-classifier.js";
import { ConcreteSensitivityValueSource } from "./catalog/sensitivity-value-source.js";
import { PythonLocalNerDetector } from "./catalog/local-ner-detector.js";
import {
ConcreteDescriptionSourceSampler,
type DescriptionSourceSampler,
@@ -66,6 +75,7 @@ import { catalogDescriptionGenerationRoutes } from "./routes/catalog-description
import { CatalogLogicalRelationshipService } from "./catalog/logical-relationship-service.js";
import { catalogLogicalRelationshipRoutes } from "./routes/catalog-logical-relationships.js";
import { EffectiveRelationshipSnapshotProvider } from "./catalog/effective-relationship-snapshot.js";
import { loadRuntimeModelCatalog, type RuntimeModelCatalog } from "./models/runtime-model-catalog.js";
export interface BuildAppDeps {
thtRunner?: ThtRunner;
@@ -77,6 +87,7 @@ export interface BuildAppDeps {
hub?: SseHub;
workspaceRegistry?: WorkspaceRegistry;
workspaceDiagnoser?: WorkspaceDiagnoser;
workspaceDatabaseTester?: WorkspaceDatabaseTester;
workspaceSecretStore?: WorkspaceSecretStore;
catalogRepository?: CatalogRepository;
catalogService?: CatalogService;
@@ -88,8 +99,11 @@ export interface BuildAppDeps {
catalogSyncWorker?: CatalogSyncWorker;
catalogOperationCoordinator?: CatalogOperationCoordinator;
metadataGenerationModels?: MetadataGenerationModels;
runtimeModelCatalog?: RuntimeModelCatalog;
modelCompleter?: ModelCompleter;
descriptionSourceSampler?: DescriptionSourceSampler;
sensitivityValueSource?: SensitivityValueSource;
localNerDetector?: LocalNerDetector;
workspaceRuntimeSupport?: (workspace: WorkspaceDescriptor) => boolean;
maintenanceBarrier?: MaintenanceBarrier;
piManagement?: PiManagementService;
@@ -155,17 +169,22 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
semanticRuntime: {
internalQdrantUrl: config.internalQdrantUrl,
internalEmbeddingUrl: config.internalEmbeddingUrl,
internalEmbeddingId: config.internalEmbeddingId,
internalEmbeddingModel: config.internalEmbeddingModel,
internalEmbeddingDimensions: config.internalEmbeddingDimensions,
},
});
const mgr = deps?.mgr ?? new PiProcessManager(config, deps?.spawnFn ? { spawnFn: deps.spawnFn } : undefined);
const hub = deps?.hub ?? new SseHub();
const workspaceRegistry = deps?.workspaceRegistry ?? new WorkspaceRegistry(config.workspaceRegistry);
const catalogRepository = deps?.catalogRepository ?? createCatalogRepository(config.catalogDatabase);
const catalogOperationCoordinator = deps?.catalogOperationCoordinator ?? new CatalogOperationCoordinator();
const runtimeModelCatalog = deps?.runtimeModelCatalog ?? loadRuntimeModelCatalog(config.modelCatalogFile);
const mgr = deps?.mgr ?? new PiProcessManager(config, {
...(deps?.spawnFn ? { spawnFn: deps.spawnFn } : {}),
modelCatalog: runtimeModelCatalog,
});
const metadataGenerationModels = deps?.metadataGenerationModels ?? loadMetadataGenerationModels({
installationFile: config.installationConfigFile,
catalogFile: config.modelCatalogFile,
secretsFile: config.secretsFile,
});
const modelCompleter = deps?.modelCompleter ?? new PythonModelCompleter({
@@ -186,12 +205,24 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
catalogOperationCoordinator,
descriptionSourceSampler,
);
const sensitiveDataSuggester = new SensitiveDataSuggester(
const sensitivityValueSource = deps?.sensitivityValueSource
?? new ConcreteSensitivityValueSource(catalogPostgresAccess, workspaceSecretStore);
const configuredNerWorker = config.sensitivityNer?.workerScript
?? fileURLToPath(new URL("../python/sensitivity_ner_worker.py", import.meta.url));
const localNerDetector = deps?.localNerDetector ?? (config.sensitivityNer
? new PythonLocalNerDetector({
pythonExecutable: config.sensitivityNer.pythonExecutable,
workerScript: configuredNerWorker,
modelPath: config.sensitivityNer.modelPath,
cwd: dirname(configuredNerWorker),
threads: config.sensitivityNer.threads,
})
: undefined);
const sensitiveDataSuggester = new SensitivityAnalysisService(
catalogRepository,
metadataGenerationModels,
modelCompleter,
new SensitivityClassifier(sensitivityValueSource, localNerDetector),
);
const sensitiveDataSuggestionRunner = new SensitiveDataSuggestionRunner(
const sensitivityAnalysisRunner = new SensitivityAnalysisRunner(
catalogRepository,
sensitiveDataSuggester,
);
@@ -204,6 +235,10 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
catalogPostgresAccess,
catalogOperationCoordinator,
);
const workspaceDatabaseTester = deps?.workspaceDatabaseTester ?? (async (workspaceId: string) => {
const database = await catalogRepository.getByWorkspace(workspaceId);
return database ? catalogService.test(database) : undefined;
});
const catalogTableService = deps?.catalogTableService ?? new CatalogTableService(catalogRepository);
const catalogLogicalRelationshipService = deps?.catalogLogicalRelationshipService
?? new CatalogLogicalRelationshipService(catalogRepository);
@@ -227,16 +262,25 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
);
app.addHook("onReady", async () => { await catalogSyncWorker.initialize(); });
app.addHook("onReady", async () => { await descriptionGenerationWorker.initialize(); });
app.addHook("onReady", async () => { await sensitiveDataSuggestionRunner.initialize(); });
app.addHook("onReady", async () => { await sensitivityAnalysisRunner.initialize(); });
if (localNerDetector?.warmup) {
app.addHook("onReady", async () => {
void localNerDetector.warmup?.().catch(() => undefined);
});
}
if (!deps?.catalogRepository && catalogRepository.close) {
app.addHook("onClose", async () => { await catalogRepository.close?.(); });
}
app.addHook("onClose", async () => { await catalogSyncWorker.stop(); });
app.addHook("onClose", async () => { await descriptionGenerationWorker.stop(); });
if (localNerDetector?.close) {
app.addHook("onClose", async () => { await localNerDetector.close?.(); });
}
const workspaceDiagnoser = deps?.workspaceDiagnoser
?? createProductionWorkspaceDiagnoser(config.workspaceDiagnosticTimeoutMs, undefined, {
internalQdrantUrl: config.internalQdrantUrl,
internalEmbeddingUrl: config.internalEmbeddingUrl,
internalEmbeddingId: config.internalEmbeddingId,
internalEmbeddingModel: config.internalEmbeddingModel,
internalEmbeddingDimensions: config.internalEmbeddingDimensions,
});
@@ -259,6 +303,7 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
);
const listModels = deps?.listModels ?? createPiModelLister(config, {
modelCatalog: runtimeModelCatalog,
warn: (detail) => app.log.warn(
{ component: "pi-model-list", detail },
"Pi enabled-model configuration warning",
@@ -271,7 +316,7 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
const getSettings = async (principal: PrincipalContext): Promise<Settings> => {
if (deps?.getSettings) return await deps.getSettings(principal);
const stored = loadSettings(config);
const effective = effectiveSettings(config, stored);
const effective = effectiveSettings(config, stored, runtimeModelCatalog);
// In the registry system the legacy `harness/workspaces/*.yaml` default is obsolete: when no
// installation workspace is pinned, default to the first active registry workspace.
if (!stored.workspace) {
@@ -284,7 +329,9 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
}
return effective;
};
const piManagement = deps?.piManagement ?? createPiManagement(config, { listModels });
const piManagement = deps?.piManagement ?? createPiManagement(config, {
modelCatalog: runtimeModelCatalog,
});
const maintenanceBarrier = deps?.maintenanceBarrier ?? new MaintenanceBarrier(config.maintenanceFile);
const localRegistryResolver = deps?.localUserRegistry === undefined
@@ -408,6 +455,7 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
dwhPrecheck: config.dwhPrecheck,
legacyWorkspaceMode: config.legacyWorkspaceMode,
workspaceRuntimeSupport,
modelCatalog: runtimeModelCatalog,
maintenanceBarrier,
effectiveRelationships,
});
@@ -439,13 +487,14 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
return maintenanceBarrier.status();
});
sqlRoutes(app, { tht: tht as ThtRunner, getSettings, workspaceRegistry });
metaRoutes(app, { harnessDir: config.harnessDir, listModels });
metaRoutes(app, { harnessDir: config.harnessDir, modelCatalog: runtimeModelCatalog });
workspaceRoutes(app, {
registry: workspaceRegistry,
config: config.workspaceRegistry,
diagnose: workspaceDiagnoser,
authDiagnoser,
secretStore: workspaceSecretStore,
testDatabaseConnection: workspaceDatabaseTester,
});
catalogDatabaseRoutes(app, { repository: catalogRepository, service: catalogService, operations: catalogOperationCoordinator });
catalogTableRoutes(app, {
@@ -470,9 +519,9 @@ export function buildApp(config: AppConfig, deps?: BuildAppDeps): FastifyInstanc
catalogDescriptionGenerationRoutes(app, {
repository: catalogRepository,
worker: descriptionGenerationWorker,
sensitiveDataSuggestionRunner,
sensitivityAnalysisRunner,
});
settingsRoutes(app, { cfg: config, listModels, getSettings });
settingsRoutes(app, { cfg: config, getSettings });
piManagementRoutes(app, { service: piManagement });
return app;
+254
View File
@@ -0,0 +1,254 @@
import { randomUUID } from "node:crypto";
import { spawn, type ChildProcessWithoutNullStreams } from "node:child_process";
import { tmpdir } from "node:os";
import { z } from "zod";
import type {
LocalNerCandidate,
LocalNerDetector,
LocalNerEvidence,
} from "./sensitivity-classifier.js";
const MAX_LINE_BYTES = 64 * 1024;
const candidateSchema = z.object({
columnId: z.uuid(),
text: z.string().min(1).max(500),
}).strict();
const workerMessageSchema = z.union([
z.object({ ready: z.literal(true) }).strict(),
z.object({
id: z.uuid(),
ok: z.literal(true),
evidence: z.array(z.object({
columnId: z.uuid(),
label: z.string().min(1).max(80),
confidence: z.number().min(0).max(1),
}).strict()).max(1_000),
}).strict(),
z.object({ id: z.uuid(), ok: z.literal(false), error: z.string().min(1).max(80) }).strict(),
]);
export class LocalNerUnavailableError extends Error {
constructor() {
super("local NER is unavailable");
this.name = "LocalNerUnavailableError";
}
}
interface PendingRequest {
resolve: (value: readonly LocalNerEvidence[]) => void;
reject: (error: Error) => void;
timer: ReturnType<typeof setTimeout>;
signal: AbortSignal;
cancel: () => void;
}
/** Persistent JSONL adapter for the optional, CPU-only Python NER worker. */
export class PythonLocalNerDetector implements LocalNerDetector {
private child?: ChildProcessWithoutNullStreams;
private ready?: Promise<void>;
private readyResolve?: () => void;
private readyReject?: (error: Error) => void;
private workerReady = false;
private stdout = "";
private readonly pending = new Map<string, PendingRequest>();
constructor(private readonly options: {
pythonExecutable: string;
workerScript: string;
modelPath: string;
cwd: string;
threads?: number;
startupTimeoutMs?: number;
}) {}
async warmup(): Promise<void> {
await this.ensureStarted();
}
isReady(): boolean {
return this.workerReady
&& this.child !== undefined
&& this.child.exitCode === null
&& this.child.signalCode === null;
}
async detect(
candidates: readonly LocalNerCandidate[],
signal: AbortSignal,
deadline: number,
): Promise<readonly LocalNerEvidence[]> {
const parsed = z.array(candidateSchema).min(1).max(128).parse(candidates);
if (signal.aborted || deadline <= Date.now()) throw new LocalNerUnavailableError();
await this.ensureStartedWithin(signal, deadline);
if (!this.child || this.child.exitCode !== null || this.child.signalCode !== null) {
throw new LocalNerUnavailableError();
}
const id = randomUUID();
return await new Promise<readonly LocalNerEvidence[]>((resolve, reject) => {
const fail = () => {
this.finishPending(id);
reject(new LocalNerUnavailableError());
this.stopWorker();
};
const timer = setTimeout(fail, Math.max(1, Math.floor(deadline - Date.now())));
const cancel = fail;
const pending: PendingRequest = { resolve, reject, timer, signal, cancel };
this.pending.set(id, pending);
signal.addEventListener("abort", cancel, { once: true });
this.child!.stdin.write(`${JSON.stringify({ id, candidates: parsed })}\n`, (error) => {
if (error) fail();
});
});
}
async close(): Promise<void> {
const child = this.child;
if (!child || child.exitCode !== null || child.signalCode !== null) return;
await new Promise<void>((resolve) => {
child.once("close", () => resolve());
child.kill("SIGTERM");
setTimeout(() => {
if (child.exitCode === null && child.signalCode === null) child.kill("SIGKILL");
}, 250).unref();
});
}
private async ensureStarted(): Promise<void> {
if (this.ready) return await this.ready;
this.ready = new Promise<void>((resolve, reject) => {
this.readyResolve = resolve;
this.readyReject = reject;
});
const threads = String(this.options.threads ?? 2);
const inheritedRuntimeEnvironment = Object.fromEntries([
"PATH", "SystemRoot", "WINDIR", "PATHEXT", "TMPDIR", "TEMP", "TMP", "LANG", "LC_ALL",
].flatMap((name) => process.env[name] === undefined ? [] : [[name, process.env[name]!]]));
const child = spawn(this.options.pythonExecutable, [
"-I",
"-B",
this.options.workerScript,
"--model",
this.options.modelPath,
"--threads",
threads,
], {
cwd: this.options.cwd,
stdio: ["pipe", "pipe", "pipe"],
env: {
...inheritedRuntimeEnvironment,
HOME: process.env.HOME ?? tmpdir(),
CUDA_VISIBLE_DEVICES: "",
HIP_VISIBLE_DEVICES: "",
HF_HUB_OFFLINE: "1",
HF_HUB_DISABLE_TELEMETRY: "1",
TRANSFORMERS_OFFLINE: "1",
TOKENIZERS_PARALLELISM: "false",
PYTHONNOUSERSITE: "1",
OMP_NUM_THREADS: threads,
MKL_NUM_THREADS: threads,
OPENBLAS_NUM_THREADS: threads,
HTTP_PROXY: "",
HTTPS_PROXY: "",
ALL_PROXY: "",
NO_PROXY: "*",
},
});
this.child = child;
child.stdout.setEncoding("utf8");
child.stdout.on("data", (chunk: string) => this.receive(chunk));
child.stderr.resume();
child.once("error", () => this.failWorker());
child.once("close", () => this.failWorker());
const startupTimer = setTimeout(() => this.failWorker(), this.options.startupTimeoutMs ?? 120_000);
startupTimer.unref();
try {
await this.ready;
} finally {
clearTimeout(startupTimer);
}
}
private async ensureStartedWithin(signal: AbortSignal, deadline: number): Promise<void> {
const started = this.ensureStarted();
await new Promise<void>((resolve, reject) => {
let settled = false;
const finish = (error?: Error, stopWorker = false) => {
if (settled) return;
settled = true;
clearTimeout(timer);
signal.removeEventListener("abort", cancel);
if (stopWorker) this.failWorker();
if (error) reject(error);
else resolve();
};
const cancel = () => finish(new LocalNerUnavailableError(), true);
const timer = setTimeout(cancel, Math.max(1, Math.floor(deadline - Date.now())));
signal.addEventListener("abort", cancel, { once: true });
void started.then(
() => finish(),
() => finish(new LocalNerUnavailableError()),
);
});
}
private receive(chunk: string): void {
this.stdout += chunk;
if (Buffer.byteLength(this.stdout, "utf8") > MAX_LINE_BYTES) {
this.failWorker();
return;
}
let newline: number;
while ((newline = this.stdout.indexOf("\n")) >= 0) {
const line = this.stdout.slice(0, newline);
this.stdout = this.stdout.slice(newline + 1);
if (!line) continue;
try {
const message = workerMessageSchema.parse(JSON.parse(line));
if ("ready" in message) {
this.workerReady = true;
this.readyResolve?.();
this.readyResolve = undefined;
this.readyReject = undefined;
continue;
}
const pending = this.pending.get(message.id);
if (!pending) continue;
this.finishPending(message.id);
if (message.ok) pending.resolve(message.evidence);
else pending.reject(new LocalNerUnavailableError());
} catch {
this.failWorker();
return;
}
}
}
private finishPending(id: string): void {
const pending = this.pending.get(id);
if (!pending) return;
clearTimeout(pending.timer);
pending.signal.removeEventListener("abort", pending.cancel);
this.pending.delete(id);
}
private stopWorker(): void {
const child = this.child;
if (child && child.exitCode === null && child.signalCode === null) child.kill("SIGTERM");
}
private failWorker(): void {
const error = new LocalNerUnavailableError();
this.readyReject?.(error);
this.readyResolve = undefined;
this.readyReject = undefined;
for (const [id, pending] of this.pending) {
this.finishPending(id);
pending.reject(error);
}
this.stopWorker();
this.child = undefined;
this.ready = undefined;
this.workerReady = false;
this.stdout = "";
}
}
+96 -57
View File
@@ -35,10 +35,10 @@ import {
type DescriptionGenerationRun,
type DescriptionGenerationRunUpdate,
type DescriptionGenerationScope,
type SensitiveDataSuggestionEvent,
type SensitiveDataSuggestionRun,
type SensitiveDataSuggestionRunUpdate,
type SensitiveDataSuggestionScope,
type SensitivityAnalysisEvent,
type SensitivityAnalysisRun,
type SensitivityAnalysisRunUpdate,
type SensitivityAnalysisScope,
type TableSyncRepositoryResult,
type WorkspaceDatabase,
} from "./types.js";
@@ -56,8 +56,8 @@ export class MemoryCatalogRepository implements CatalogRepository {
private readonly logicalRelationships = new Map<string, CatalogLogicalRelationship>();
private readonly descriptionGenerationRuns = new Map<string, DescriptionGenerationRun>();
private readonly descriptionGenerationEvents = new Map<string, DescriptionGenerationEvent[]>();
private readonly sensitiveDataSuggestionRuns = new Map<string, SensitiveDataSuggestionRun>();
private readonly sensitiveDataSuggestionEvents = new Map<string, SensitiveDataSuggestionEvent[]>();
private readonly sensitivityAnalysisRuns = new Map<string, SensitivityAnalysisRun>();
private readonly sensitivityAnalysisEvents = new Map<string, SensitivityAnalysisEvent[]>();
private readonly syncRuns = new Map<string, CatalogSyncRun>();
private readonly syncEvents = new Map<string, CatalogSyncEvent[]>();
@@ -83,8 +83,10 @@ export class MemoryCatalogRepository implements CatalogRepository {
const tableIds = new Set(tables.map((table) => table.id));
const columns = [...this.columns.values()]
.filter((column) => tableIds.has(column.tableId));
const relationships = [...this.relationships.values()]
.filter((relationship) => selectedDatabaseIds.has(relationship.databaseId));
const relationships = [
...this.relationships.values(),
...this.logicalRelationships.values(),
].filter((relationship) => selectedDatabaseIds.has(relationship.databaseId));
return createCatalogMetrics(databaseId, {
tables: tables.length,
@@ -183,10 +185,10 @@ export class MemoryCatalogRepository implements CatalogRepository {
this.descriptionGenerationRuns.delete(runId);
this.descriptionGenerationEvents.delete(runId);
}
for (const [runId, run] of this.sensitiveDataSuggestionRuns) {
for (const [runId, run] of this.sensitivityAnalysisRuns) {
if (run.databaseId !== id) continue;
this.sensitiveDataSuggestionRuns.delete(runId);
this.sensitiveDataSuggestionEvents.delete(runId);
this.sensitivityAnalysisRuns.delete(runId);
this.sensitivityAnalysisEvents.delete(runId);
}
return this.records.delete(id);
}
@@ -255,6 +257,7 @@ export class MemoryCatalogRepository implements CatalogRepository {
description: string | null,
generatedDescription: string | null,
sensitive?: boolean,
sensitivityReason?: string | null,
): Promise<CatalogColumn | undefined> {
const current = await this.getColumn(databaseId, tableId, columnId);
if (!current || current.version !== expectedVersion) return undefined;
@@ -263,6 +266,9 @@ export class MemoryCatalogRepository implements CatalogRepository {
description,
generatedDescription,
sensitive: sensitive ?? current.sensitive,
sensitivityReason: sensitive === false
? null
: sensitivityReason === undefined ? current.sensitivityReason : sensitivityReason,
version: current.version + 1,
updatedAt: new Date().toISOString(),
};
@@ -276,10 +282,26 @@ export class MemoryCatalogRepository implements CatalogRepository {
targetIds: readonly string[],
): Promise<CatalogDescriptionConsolidationCounts | undefined> {
const selectedTargetIds = [...new Set(targetIds)];
if (!this.records.has(databaseId) || selectedTargetIds.length === 0) {
if (!this.records.has(databaseId)) {
return undefined;
}
const now = new Date().toISOString();
if (target === "database_columns") {
const targets = [...this.columns.values()].filter((column) => (
this.tables.get(column.tableId)?.databaseId === databaseId
));
const copied = targets.filter((column) => Boolean(column.generatedDescription?.trim()));
for (const column of copied) {
this.columns.set(column.id, {
...column,
description: column.generatedDescription,
version: column.version + 1,
updatedAt: now,
});
}
return { copied: copied.length, skipped: targets.length - copied.length };
}
if (selectedTargetIds.length === 0) return undefined;
if (target === "tables") {
const targets = selectedTargetIds.map((id) => this.tables.get(id));
if (targets.some((table) => !table || table.databaseId !== databaseId)) return undefined;
@@ -433,93 +455,96 @@ export class MemoryCatalogRepository implements CatalogRepository {
.map((event) => structuredClone(event));
}
async createSensitiveDataSuggestionRun(
async createSensitivityAnalysisRun(
databaseId: string,
scope: SensitiveDataSuggestionScope,
modelId: string,
): Promise<SensitiveDataSuggestionRun> {
scope: SensitivityAnalysisScope,
origin: { engine: "llm"; modelId: string } | { engine: "local"; policyVersion: string },
): Promise<SensitivityAnalysisRun> {
const now = new Date().toISOString();
const run: SensitiveDataSuggestionRun = {
const run: SensitivityAnalysisRun = {
id: randomUUID(),
databaseId,
scope,
modelId,
engine: origin.engine,
modelId: origin.engine === "llm" ? origin.modelId : null,
policyVersion: origin.engine === "local" ? origin.policyVersion : null,
status: "running",
total: 0,
suggestedSensitive: 0,
suggestedNonSensitive: 0,
inputTokens: 0,
cacheReadTokens: 0,
outputTokens: 0,
suggestedNonSensitive: 0,
unknown: 0,
inputTokens: 0,
cacheReadTokens: 0,
outputTokens: 0,
createdAt: now,
startedAt: now,
updatedAt: now,
finishedAt: null,
errorSummary: null,
};
this.sensitiveDataSuggestionRuns.set(run.id, run);
this.sensitivityAnalysisRuns.set(run.id, run);
return structuredClone(run);
}
async getSensitiveDataSuggestionRun(
async getSensitivityAnalysisRun(
runId: string,
): Promise<SensitiveDataSuggestionRun | undefined> {
const run = this.sensitiveDataSuggestionRuns.get(runId);
): Promise<SensitivityAnalysisRun | undefined> {
const run = this.sensitivityAnalysisRuns.get(runId);
return run ? structuredClone(run) : undefined;
}
async listSensitiveDataSuggestionRuns(limit = 50): Promise<SensitiveDataSuggestionRun[]> {
return [...this.sensitiveDataSuggestionRuns.values()]
async listSensitivityAnalysisRuns(limit = 50): Promise<SensitivityAnalysisRun[]> {
return [...this.sensitivityAnalysisRuns.values()]
.sort((a, b) => b.createdAt.localeCompare(a.createdAt) || b.id.localeCompare(a.id))
.slice(0, limit)
.map((run) => structuredClone(run));
}
async interruptActiveSensitiveDataSuggestionRuns(
async interruptActiveSensitivityAnalysisRuns(
errorSummary: string,
): Promise<SensitiveDataSuggestionRun[]> {
const interrupted: SensitiveDataSuggestionRun[] = [];
for (const run of this.sensitiveDataSuggestionRuns.values()) {
): Promise<SensitivityAnalysisRun[]> {
const interrupted: SensitivityAnalysisRun[] = [];
for (const run of this.sensitivityAnalysisRuns.values()) {
if (run.status !== "running") continue;
const now = new Date().toISOString();
const updated: SensitiveDataSuggestionRun = {
const updated: SensitivityAnalysisRun = {
...run,
status: "interrupted",
updatedAt: now,
finishedAt: now,
errorSummary,
};
this.sensitiveDataSuggestionRuns.set(run.id, updated);
this.sensitivityAnalysisRuns.set(run.id, updated);
interrupted.push(structuredClone(updated));
}
return interrupted;
}
async updateSensitiveDataSuggestionRun(
async updateSensitivityAnalysisRun(
runId: string,
update: SensitiveDataSuggestionRunUpdate,
): Promise<SensitiveDataSuggestionRun | undefined> {
const current = this.sensitiveDataSuggestionRuns.get(runId);
update: SensitivityAnalysisRunUpdate,
): Promise<SensitivityAnalysisRun | undefined> {
const current = this.sensitivityAnalysisRuns.get(runId);
if (!current) return undefined;
const updated = {
...current,
...structuredClone(update),
updatedAt: new Date().toISOString(),
};
this.sensitiveDataSuggestionRuns.set(runId, updated);
this.sensitivityAnalysisRuns.set(runId, updated);
return structuredClone(updated);
}
async appendSensitiveDataSuggestionEvent(
async appendSensitivityAnalysisEvent(
runId: string,
level: SensitiveDataSuggestionEvent["level"],
level: SensitivityAnalysisEvent["level"],
message: string,
): Promise<SensitiveDataSuggestionEvent> {
if (!this.sensitiveDataSuggestionRuns.has(runId)) {
throw new CatalogConflictError("Sensitive Data Suggestion Run does not exist");
): Promise<SensitivityAnalysisEvent> {
if (!this.sensitivityAnalysisRuns.has(runId)) {
throw new CatalogConflictError("Sensitivity Analysis Run does not exist");
}
const events = this.sensitiveDataSuggestionEvents.get(runId) ?? [];
const event: SensitiveDataSuggestionEvent = {
const events = this.sensitivityAnalysisEvents.get(runId) ?? [];
const event: SensitivityAnalysisEvent = {
runId,
sequence: events.length + 1,
level,
@@ -527,15 +552,15 @@ export class MemoryCatalogRepository implements CatalogRepository {
createdAt: new Date().toISOString(),
};
events.push(event);
this.sensitiveDataSuggestionEvents.set(runId, events);
this.sensitivityAnalysisEvents.set(runId, events);
return structuredClone(event);
}
async listSensitiveDataSuggestionEvents(
async listSensitivityAnalysisEvents(
runId: string,
afterSequence = 0,
): Promise<SensitiveDataSuggestionEvent[]> {
return (this.sensitiveDataSuggestionEvents.get(runId) ?? [])
): Promise<SensitivityAnalysisEvent[]> {
return (this.sensitivityAnalysisEvents.get(runId) ?? [])
.filter((event) => event.sequence > afterSequence)
.map((event) => structuredClone(event));
}
@@ -684,18 +709,22 @@ export class MemoryCatalogRepository implements CatalogRepository {
const tables = [...this.tables.values()].filter((table) => selected.has(table.databaseId));
const tableIds = new Set(tables.map((table) => table.id));
const columns = [...this.columns.values()].filter((column) => tableIds.has(column.tableId));
const relationships = [...this.relationships.values()]
const physicalRelationships = [...this.relationships.values()]
.filter((relationship) => selected.has(relationship.databaseId));
const logicalRelationships = [...this.logicalRelationships.values()]
.filter((relationship) => selected.has(relationship.databaseId));
const relationshipCount = physicalRelationships.length + logicalRelationships.length;
if (target === "tables") {
for (const table of tables) this.deleteTable(table.id);
this.markCatalogIncomplete(selectedDatabaseIds);
return { tables: tables.length, columns: columns.length, relationships: relationships.length };
return { tables: tables.length, columns: columns.length, relationships: relationshipCount };
}
for (const relationship of relationships) this.relationships.delete(relationship.id);
for (const relationship of physicalRelationships) this.relationships.delete(relationship.id);
for (const relationship of logicalRelationships) this.logicalRelationships.delete(relationship.id);
for (const databaseId of selectedDatabaseIds) this.refreshForeignKeyFlags(databaseId);
this.markCatalogIncomplete(selectedDatabaseIds);
return { tables: 0, columns: 0, relationships: relationships.length };
return { tables: 0, columns: 0, relationships: relationshipCount };
}
async deleteTableMetadata(
@@ -733,14 +762,23 @@ export class MemoryCatalogRepository implements CatalogRepository {
return { tables: 0, columns: columns.length, relationships: 0 };
}
const relationships = [...this.relationships.values()].filter((relationship) => (
const physicalRelationships = [...this.relationships.values()].filter((relationship) => (
relationship.databaseId === databaseId
&& (selected.has(relationship.sourceTableId) || selected.has(relationship.targetTableId))
));
for (const relationship of relationships) this.relationships.delete(relationship.id);
const logicalRelationships = [...this.logicalRelationships.values()].filter((relationship) => (
relationship.databaseId === databaseId
&& (selected.has(relationship.sourceTableId) || selected.has(relationship.targetTableId))
));
for (const relationship of physicalRelationships) this.relationships.delete(relationship.id);
for (const relationship of logicalRelationships) this.logicalRelationships.delete(relationship.id);
this.refreshForeignKeyFlags(databaseId);
this.markCatalogIncomplete([databaseId]);
return { tables: 0, columns: 0, relationships: relationships.length };
return {
tables: 0,
columns: 0,
relationships: physicalRelationships.length + logicalRelationships.length,
};
}
async planSchemaSync(
@@ -924,6 +962,7 @@ export class MemoryCatalogRepository implements CatalogRepository {
description: null,
generatedDescription: null,
sensitive: false,
sensitivityReason: null,
lastSyncedDatabaseVersion: expectedDatabaseVersion,
lastSyncedAt: now,
version: 1,
+41 -162
View File
@@ -1,63 +1,5 @@
import {
closeSync, constants, fstatSync, lstatSync, openSync, readFileSync,
type Stats,
} from "node:fs";
import { parseAllDocuments } from "yaml";
import { z } from "zod";
import {
loadSecretBundle,
METADATA_GENERATION_SECRET_KEYS,
} from "../config/secret-bundle.js";
const MAX_INSTALLATION_BYTES = 1024 * 1024;
const RUNTIME_INSTALLATION_FILE = "/run/thothii-installation/thothii-installation.yaml";
const modelId = z.string().regex(/^[a-z][a-z0-9._-]{0,63}$/);
const apiKeyEnvironment = z.enum(METADATA_GENERATION_SECRET_KEYS);
const endpointSchema = z.object({
baseUrl: z.string().min(1).max(2048).refine((value) => {
try {
const url = new URL(value);
return (url.protocol === "http:" || url.protocol === "https:")
&& url.username === "" && url.password === "" && url.search === "" && url.hash === "";
} catch {
return false;
}
}),
apiVersion: z.string().regex(/^[A-Za-z0-9][A-Za-z0-9._-]{0,127}$/).optional(),
}).strict();
const configuredModelSchema = z.object({
id: modelId,
label: z.string().min(1).max(128).refine((value) => value.trim() === value && !/\p{Cc}/u.test(value)),
litellm: z.object({
provider: z.string().regex(/^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/),
model: z.string().regex(/^[A-Za-z0-9][A-Za-z0-9._:/-]{0,255}$/),
disableThinking: z.literal(true).optional(),
endpoint: endpointSchema.optional(),
}).strict(),
apiKeyEnv: apiKeyEnvironment.optional(),
}).strict().superRefine((value, context) => {
if (value.apiKeyEnv === undefined && value.litellm.endpoint === undefined) {
context.addIssue({
code: z.ZodIssueCode.custom,
path: ["apiKeyEnv"],
message: "keyless models require an explicit endpoint",
});
}
if (value.litellm.disableThinking === true && value.litellm.endpoint === undefined) {
context.addIssue({
code: z.ZodIssueCode.custom,
path: ["litellm", "disableThinking"],
message: "thinking may be disabled only for an explicit endpoint",
});
}
});
const metadataGenerationSchema = z.object({
default: modelId.optional(),
models: z.array(configuredModelSchema).max(64).default([]),
}).strict();
const installationSchema = z.object({
metadataGeneration: metadataGenerationSchema.optional(),
}).passthrough();
import { loadSecretBundle } from "../config/secret-bundle.js";
import { loadRuntimeModelCatalog } from "../models/runtime-model-catalog.js";
export interface MetadataGenerationModelChoice {
id: string;
@@ -86,7 +28,6 @@ export class MetadataGenerationModelUnavailableError extends Error {
}
}
/** The complete interface callers need: safe discovery plus fail-closed runtime resolution. */
export interface MetadataGenerationModels {
catalog(): MetadataGenerationModelCatalog;
resolve(selection: string): ResolvedMetadataGenerationModel;
@@ -96,140 +37,78 @@ class RestartLoadedMetadataGenerationModels implements MetadataGenerationModels
readonly #models: ReadonlyMap<string, ResolvedMetadataGenerationModel>;
readonly #catalog: MetadataGenerationModelCatalog;
constructor(
models: ReadonlyMap<string, ResolvedMetadataGenerationModel> = new Map(),
defaultModel: string | null = null,
choices: MetadataGenerationModelChoice[] = [],
) {
constructor(models: ReadonlyMap<string, ResolvedMetadataGenerationModel>, defaultModel: string | null) {
this.#models = models;
this.#catalog = {
models: choices.map((choice) => ({ ...choice })),
models: [...models.values()].map(({ id }) => ({ id, label: id })),
default: defaultModel,
};
}
catalog(): MetadataGenerationModelCatalog {
return {
models: this.#catalog.models.map((choice) => ({ ...choice })),
default: this.#catalog.default,
};
return { models: this.#catalog.models.map((choice) => ({ ...choice })), default: this.#catalog.default };
}
resolve(selection: string): ResolvedMetadataGenerationModel {
const model = typeof selection === "string" ? this.#models.get(selection) : undefined;
const model = this.#models.get(selection);
if (!model) throw new MetadataGenerationModelUnavailableError();
return model;
}
}
function invalid(message = "metadata-generation configuration is invalid"): Error {
function invalid(message = "metadata-generation runtime catalog is invalid"): Error {
return new Error(message);
}
function protectedInstallationStat(file: string, info: Stats): boolean {
const mode = info.mode & 0o777;
if (!info.isFile() || info.isSymbolicLink() || info.nlink !== 1
|| info.size < 1 || info.size > MAX_INSTALLATION_BYTES) return false;
if (file === RUNTIME_INSTALLATION_FILE && info.uid === 0 && mode === 0o444) return true;
return info.uid === (process.getuid?.() ?? info.uid) && (mode === 0o400 || mode === 0o600);
}
function readProtectedInstallation(file: string): string {
let descriptor: number | undefined;
try {
const before = lstatSync(file);
if (!protectedInstallationStat(file, before)) throw new Error("unavailable");
descriptor = openSync(file, constants.O_RDONLY | constants.O_NOFOLLOW);
const opened = fstatSync(descriptor);
if (!protectedInstallationStat(file, opened)
|| before.dev !== opened.dev || before.ino !== opened.ino) throw new Error("unavailable");
const source = readFileSync(descriptor, "utf8");
const after = fstatSync(descriptor);
const current = lstatSync(file);
if (!protectedInstallationStat(file, after) || !protectedInstallationStat(file, current)
|| opened.dev !== after.dev || opened.ino !== after.ino
|| opened.dev !== current.dev || opened.ino !== current.ino) throw new Error("unavailable");
return source;
} finally {
if (descriptor !== undefined) try { closeSync(descriptor); } catch { /* sanitized below */ }
}
}
function readInstallation(file: string): unknown {
try {
const documents = parseAllDocuments(readProtectedInstallation(file), { uniqueKeys: true });
if (documents.length !== 1) throw invalid("metadata-generation installation must contain one YAML document");
const document = documents[0];
if (document.errors.length > 0 || document.warnings.length > 0) {
throw invalid("metadata-generation installation contains invalid YAML");
}
return document.toJSON();
} catch (error) {
if (error instanceof Error && error.message.startsWith("metadata-generation")) throw error;
throw invalid("metadata-generation installation is unavailable");
}
}
export function loadMetadataGenerationModels(options: {
installationFile?: string;
catalogFile?: string;
secretsFile?: string;
}): MetadataGenerationModels {
if (!options.installationFile) return new RestartLoadedMetadataGenerationModels();
const installation = installationSchema.safeParse(readInstallation(options.installationFile));
if (!installation.success) throw invalid();
const configured = installation.data.metadataGeneration;
if (!configured || configured.models.length === 0) {
if (configured?.default !== undefined) throw invalid("metadata-generation default does not identify a configured model");
return new RestartLoadedMetadataGenerationModels();
}
if (!configured.default) throw invalid("metadata-generation default is required when models are configured");
const catalog = loadRuntimeModelCatalog(options.catalogFile);
const configured = catalog.metadataModels();
if (configured.length === 0) return new RestartLoadedMetadataGenerationModels(new Map(), null);
const seen = new Set<string>();
for (const model of configured.models) {
if (seen.has(model.id)) throw invalid(`metadata-generation model id "${model.id}" is duplicated`);
seen.add(model.id);
}
if (!seen.has(configured.default)) {
throw invalid(`metadata-generation default "${configured.default}" is not configured`);
}
const requiresSecrets = configured.models.some((model) => model.apiKeyEnv !== undefined);
const requiresSecrets = configured.some((model) => model.authentication.mode === "secret_env");
let secrets: ReadonlyMap<string, string> = new Map();
if (requiresSecrets) {
if (!options.secretsFile) throw invalid("metadata-generation keyed models require THT_SECRETS_FILE");
try {
secrets = loadSecretBundle(options.secretsFile);
} catch {
throw invalid("metadata-generation secrets are unavailable");
}
try { secrets = loadSecretBundle(options.secretsFile); }
catch { throw invalid("metadata-generation secrets are unavailable"); }
}
const models = new Map<string, ResolvedMetadataGenerationModel>();
for (const configuredModel of configured.models) {
const labels = new Map<string, string>();
for (const configuredModel of configured) {
const adapter = configuredModel.metadataAdapter;
if (!adapter || configuredModel.authentication.mode === "pi_auth") throw invalid();
const apiKeyEnv = configuredModel.authentication.apiKeyEnv;
let apiKey: string | undefined;
if (configuredModel.apiKeyEnv !== undefined) {
apiKey = secrets.get(configuredModel.apiKeyEnv);
if (!apiKey) {
throw invalid(`metadata-generation model "${configuredModel.id}" secret "${configuredModel.apiKeyEnv}" is missing`);
}
if (configuredModel.authentication.mode === "secret_env") {
if (!apiKeyEnv) throw invalid();
apiKey = secrets.get(apiKeyEnv);
if (!apiKey) throw invalid(`metadata-generation model "${configuredModel.id}" secret "${apiKeyEnv}" is missing`);
if (apiKey.length > 16 * 1024 || /\s/u.test(apiKey)) {
throw invalid(`metadata-generation model "${configuredModel.id}" secret "${configuredModel.apiKeyEnv}" is unusable`);
throw invalid(`metadata-generation model "${configuredModel.id}" secret "${apiKeyEnv}" is unusable`);
}
}
labels.set(configuredModel.id, configuredModel.label);
models.set(configuredModel.id, Object.freeze({
id: configuredModel.id,
provider: configuredModel.litellm.provider,
model: configuredModel.litellm.model,
...(configuredModel.litellm.disableThinking === true ? { disableThinking: true as const } : {}),
...(configuredModel.litellm.endpoint === undefined
? {}
: { endpoint: Object.freeze({ ...configuredModel.litellm.endpoint }) }),
...(configuredModel.apiKeyEnv === undefined
? {}
: { apiKeyEnv: configuredModel.apiKeyEnv, apiKey }),
provider: adapter.litellmProvider,
model: configuredModel.upstreamModel,
...(configuredModel.metadataGeneration?.disableThinking === true
? { disableThinking: true as const } : {}),
...(configuredModel.endpoint ? { endpoint: Object.freeze({ ...configuredModel.endpoint }) } : {}),
...(apiKeyEnv ? { apiKeyEnv, apiKey } : {}),
}));
}
return new RestartLoadedMetadataGenerationModels(
models,
configured.default,
configured.models.map(({ id, label }) => ({ id, label })),
);
const result = new RestartLoadedMetadataGenerationModels(models, catalog.defaultMetadataGeneration);
const safe = result.catalog();
return {
catalog: () => ({
default: safe.default,
models: safe.models.map((choice) => ({ ...choice, label: labels.get(choice.id) ?? choice.id })),
}),
resolve: (selection) => result.resolve(selection),
};
}
+8 -2
View File
@@ -9,9 +9,12 @@ import * as catalogSchemaSyncMigration from "./migrations/003_catalog_schema_syn
import * as catalogRuntimeSequencePrivilegesMigration from "./migrations/004_catalog_runtime_sequence_privileges.js";
import * as descriptionGenerationRunsMigration from "./migrations/005_description_generation_runs.js";
import * as sensitiveDataFlagMigration from "./migrations/006_sensitive_data_flag.js";
import * as sensitiveDataSuggestionRunsMigration from "./migrations/007_sensitive_data_suggestion_runs.js";
import * as sensitivityAnalysisRunsMigration from "./migrations/007_sensitive_data_suggestion_runs.js";
import * as catalogLogicalRelationshipsMigration from "./migrations/008_catalog_logical_relationships.js";
import * as aiTokenUsageMigration from "./migrations/009_ai_token_usage.js";
import * as canonicalModelIdsMigration from "./migrations/010_canonical_model_ids.js";
import * as localSensitivityAnalysisMigration from "./migrations/011_local_sensitivity_analysis.js";
import * as sensitivityReasonMigration from "./migrations/012_sensitivity_reason.js";
const connectionString = process.env.THT_CATALOG_MIGRATOR_DATABASE_URL;
const host = process.env.THT_CATALOG_DB_HOST;
@@ -43,9 +46,12 @@ const provider: MigrationProvider = {
"004_catalog_runtime_sequence_privileges": catalogRuntimeSequencePrivilegesMigration,
"005_description_generation_runs": descriptionGenerationRunsMigration,
"006_sensitive_data_flag": sensitiveDataFlagMigration,
"007_sensitive_data_suggestion_runs": sensitiveDataSuggestionRunsMigration,
"007_sensitive_data_suggestion_runs": sensitivityAnalysisRunsMigration,
"008_catalog_logical_relationships": catalogLogicalRelationshipsMigration,
"009_ai_token_usage": aiTokenUsageMigration,
"010_canonical_model_ids": canonicalModelIdsMigration,
"011_local_sensitivity_analysis": localSensitivityAnalysisMigration,
"012_sensitivity_reason": sensitivityReasonMigration,
};
},
};
@@ -0,0 +1,28 @@
import { sql, type Kysely } from "kysely";
import type { CatalogDatabase } from "../repository.js";
const canonicalModelPattern = "^[a-z][a-z0-9._-]{0,63}/[A-Za-z0-9][A-Za-z0-9._:-]{0,255}$";
const legacyModelPattern = "^[a-z][a-z0-9._-]{0,63}$";
const historicalOrCanonicalModelPattern = `(${legacyModelPattern})|(${canonicalModelPattern})`;
export async function up(db: Kysely<CatalogDatabase>): Promise<void> {
await sql.raw(`alter table description_generation_runs
drop constraint description_generation_runs_model_id_check,
add constraint description_generation_runs_model_id_check
check (model_id ~ '${historicalOrCanonicalModelPattern}')`).execute(db);
await sql.raw(`alter table sensitive_data_suggestion_runs
drop constraint sensitive_data_suggestion_runs_model_id_check,
add constraint sensitive_data_suggestion_runs_model_id_check
check (model_id ~ '${historicalOrCanonicalModelPattern}')`).execute(db);
}
export async function down(db: Kysely<CatalogDatabase>): Promise<void> {
await sql.raw(`alter table sensitive_data_suggestion_runs
drop constraint sensitive_data_suggestion_runs_model_id_check,
add constraint sensitive_data_suggestion_runs_model_id_check
check (model_id ~ '${legacyModelPattern}')`).execute(db);
await sql.raw(`alter table description_generation_runs
drop constraint description_generation_runs_model_id_check,
add constraint description_generation_runs_model_id_check
check (model_id ~ '${legacyModelPattern}')`).execute(db);
}
@@ -0,0 +1,42 @@
import { sql, type Kysely } from "kysely";
import type { CatalogDatabase } from "../repository.js";
export async function up(db: Kysely<CatalogDatabase>): Promise<void> {
await sql.raw(`alter table sensitive_data_suggestion_runs
alter column model_id drop not null,
add column engine text not null default 'llm',
add column policy_version text,
add column unknown integer not null default 0,
drop constraint sensitive_data_suggestion_runs_counters_check,
add constraint sensitive_data_suggestion_runs_counters_check
check (total >= 0
and suggested_sensitive >= 0
and suggested_non_sensitive >= 0
and unknown >= 0
and suggested_sensitive + suggested_non_sensitive + unknown <= total),
add constraint sensitive_data_suggestion_runs_engine_check
check (engine in ('llm', 'local')),
add constraint sensitive_data_suggestion_runs_origin_check
check ((engine = 'llm' and model_id is not null and policy_version is null)
or (engine = 'local' and model_id is null
and policy_version ~ '^[a-z][a-z0-9._-]{0,63}$'))`).execute(db);
}
export async function down(db: Kysely<CatalogDatabase>): Promise<void> {
await sql.raw(`alter table sensitive_data_suggestion_runs
drop constraint sensitive_data_suggestion_runs_origin_check,
drop constraint sensitive_data_suggestion_runs_engine_check,
drop constraint sensitive_data_suggestion_runs_counters_check`).execute(db);
await sql.raw(`update sensitive_data_suggestion_runs
set model_id = coalesce(model_id, 'local/sensitivity-v1')`).execute(db);
await sql.raw(`alter table sensitive_data_suggestion_runs
drop column unknown,
drop column policy_version,
drop column engine,
alter column model_id set not null,
add constraint sensitive_data_suggestion_runs_counters_check
check (total >= 0
and suggested_sensitive >= 0
and suggested_non_sensitive >= 0
and suggested_sensitive + suggested_non_sensitive <= total)`).execute(db);
}
@@ -0,0 +1,12 @@
import type { Kysely } from "kysely";
import type { CatalogDatabase } from "../repository.js";
export async function up(db: Kysely<CatalogDatabase>): Promise<void> {
await db.schema.alterTable("catalog_columns")
.addColumn("sensitivity_reason", "text")
.execute();
}
export async function down(db: Kysely<CatalogDatabase>): Promise<void> {
await db.schema.alterTable("catalog_columns").dropColumn("sensitivity_reason").execute();
}
+132 -67
View File
@@ -45,10 +45,10 @@ import {
type DescriptionGenerationScope,
type ObservedCatalogTable,
type ObservedSchemaSnapshot,
type SensitiveDataSuggestionEvent,
type SensitiveDataSuggestionRun,
type SensitiveDataSuggestionRunUpdate,
type SensitiveDataSuggestionScope,
type SensitivityAnalysisEvent,
type SensitivityAnalysisRun,
type SensitivityAnalysisRunUpdate,
type SensitivityAnalysisScope,
type TableSyncRepositoryResult,
type WorkspaceDatabase,
} from "./types.js";
@@ -118,6 +118,7 @@ interface CatalogColumnTable {
description: string | null;
generatedDescription: string | null;
sensitive: Generated<boolean>;
sensitivityReason: Generated<string | null>;
lastSyncedDatabaseVersion: number | null;
lastSyncedAt: Timestamp | null;
version: Generated<number>;
@@ -189,15 +190,18 @@ interface DescriptionGenerationEventTable {
createdAt: Timestamp;
}
interface SensitiveDataSuggestionRunTable {
interface SensitivityAnalysisRunTable {
id: string;
databaseId: string;
scope: SensitiveDataSuggestionScope;
modelId: string;
status: SensitiveDataSuggestionRun["status"];
scope: SensitivityAnalysisScope;
engine: SensitivityAnalysisRun["engine"];
modelId: string | null;
policyVersion: string | null;
status: SensitivityAnalysisRun["status"];
total: number;
suggestedSensitive: number;
suggestedNonSensitive: number;
unknown: number;
inputTokens: number;
cacheReadTokens: number;
outputTokens: number;
@@ -208,10 +212,10 @@ interface SensitiveDataSuggestionRunTable {
errorSummary: string | null;
}
interface SensitiveDataSuggestionEventTable {
interface SensitivityAnalysisEventTable {
runId: string;
sequence: number;
level: SensitiveDataSuggestionEvent["level"];
level: SensitivityAnalysisEvent["level"];
message: string;
createdAt: Timestamp;
}
@@ -264,8 +268,9 @@ export interface CatalogDatabase {
catalogLogicalRelationships: CatalogLogicalRelationshipTable;
descriptionGenerationRuns: DescriptionGenerationRunTable;
descriptionGenerationEvents: DescriptionGenerationEventTable;
sensitiveDataSuggestionRuns: SensitiveDataSuggestionRunTable;
sensitiveDataSuggestionEvents: SensitiveDataSuggestionEventTable;
// Legacy physical table names retained for migration and storage compatibility.
sensitiveDataSuggestionRuns: SensitivityAnalysisRunTable;
sensitiveDataSuggestionEvents: SensitivityAnalysisEventTable;
catalogSyncRuns: CatalogSyncRunTable;
catalogSyncEvents: CatalogSyncEventTable;
}
@@ -356,6 +361,7 @@ function serializeColumn(row: Selectable<CatalogColumnTable>, foreignKeyCount =
description: row.description,
generatedDescription: row.generatedDescription,
sensitive: row.sensitive,
sensitivityReason: row.sensitivityReason,
lastSyncedDatabaseVersion: row.lastSyncedDatabaseVersion,
lastSyncedAt: row.lastSyncedAt === null ? null : new Date(row.lastSyncedAt).toISOString(),
version: row.version,
@@ -402,9 +408,9 @@ function serializeDescriptionGenerationEvent(
return { ...row, createdAt: new Date(row.createdAt).toISOString() };
}
function serializeSensitiveDataSuggestionRun(
row: Selectable<SensitiveDataSuggestionRunTable>,
): SensitiveDataSuggestionRun {
function serializeSensitivityAnalysisRun(
row: Selectable<SensitivityAnalysisRunTable>,
): SensitivityAnalysisRun {
const stamp = (value: Date | string | null) => value === null ? null : new Date(value).toISOString();
return {
...row,
@@ -415,9 +421,9 @@ function serializeSensitiveDataSuggestionRun(
};
}
function serializeSensitiveDataSuggestionEvent(
row: Selectable<SensitiveDataSuggestionEventTable>,
): SensitiveDataSuggestionEvent {
function serializeSensitivityAnalysisEvent(
row: Selectable<SensitivityAnalysisEventTable>,
): SensitivityAnalysisEvent {
return { ...row, createdAt: new Date(row.createdAt).toISOString() };
}
@@ -512,10 +518,18 @@ export class KyselyCatalogRepository implements CatalogRepository {
relationship_metrics AS (
SELECT
count(*)::int AS relationships,
max(catalog_relationships.updated_at) AS updated_at
FROM catalog_relationships
INNER JOIN selected_databases
ON selected_databases.id = catalog_relationships.database_id
max(relationship.updated_at) AS updated_at
FROM (
SELECT catalog_relationships.updated_at
FROM catalog_relationships
INNER JOIN selected_databases
ON selected_databases.id = catalog_relationships.database_id
UNION ALL
SELECT catalog_logical_relationships.updated_at
FROM catalog_logical_relationships
INNER JOIN selected_databases
ON selected_databases.id = catalog_logical_relationships.database_id
) AS relationship
)
SELECT
(SELECT count(*)::int FROM selected_databases) AS "databaseCount",
@@ -724,6 +738,7 @@ export class KyselyCatalogRepository implements CatalogRepository {
description: string | null,
generatedDescription: string | null,
sensitive?: boolean,
sensitivityReason?: string | null,
): Promise<CatalogColumn | undefined> {
const belongs = await this.db.selectFrom("catalogTables").select("id")
.where("id", "=", tableId).where("databaseId", "=", databaseId).executeTakeFirst();
@@ -732,6 +747,9 @@ export class KyselyCatalogRepository implements CatalogRepository {
description,
generatedDescription,
...(sensitive === undefined ? {} : { sensitive }),
...(sensitive === false
? { sensitivityReason: null }
: sensitivityReason === undefined ? {} : { sensitivityReason }),
version: sql`version + 1`,
updatedAt: sql`now()`,
}).where("id", "=", columnId).where("tableId", "=", tableId)
@@ -745,7 +763,7 @@ export class KyselyCatalogRepository implements CatalogRepository {
targetIds: readonly string[],
): Promise<CatalogDescriptionConsolidationCounts | undefined> {
const selectedTargetIds = [...new Set(targetIds)];
if (selectedTargetIds.length === 0) return undefined;
if (target !== "database_columns" && selectedTargetIds.length === 0) return undefined;
return await this.db.transaction().execute(async (trx) => {
const database = await trx.selectFrom("workspaceDatabases").select("id")
.where("id", "=", databaseId).forUpdate().executeTakeFirst();
@@ -780,24 +798,30 @@ export class KyselyCatalogRepository implements CatalogRepository {
const rows = tableRows.length === 0 ? [] : await trx.selectFrom("catalogColumns")
.select(["id", "generatedDescription"])
.where("tableId", "in", tableRows.map((table) => table.id))
.where("id", "in", selectedTargetIds)
.$if(target === "columns", (query) => query.where("id", "in", selectedTargetIds))
.orderBy("id")
.forUpdate()
.execute();
if (rows.length !== selectedTargetIds.length) return undefined;
if (target === "columns" && rows.length !== selectedTargetIds.length) return undefined;
const copiedIds = rows
.filter((row) => Boolean(row.generatedDescription?.trim()))
.map((row) => row.id);
if (copiedIds.length > 0) {
await trx.updateTable("catalogColumns").set({
let update = trx.updateTable("catalogColumns").set({
description: sql`generated_description`,
version: sql`version + 1`,
updatedAt: sql`now()`,
}).where("id", "in", copiedIds).execute();
});
update = target === "database_columns"
? update
.where("tableId", "in", tableRows.map((table) => table.id))
.where(sql<boolean>`nullif(btrim(generated_description), '') is not null`)
: update.where("id", "in", copiedIds);
await update.execute();
}
return {
copied: copiedIds.length,
skipped: selectedTargetIds.length - copiedIds.length,
skipped: rows.length - copiedIds.length,
};
});
}
@@ -938,52 +962,55 @@ export class KyselyCatalogRepository implements CatalogRepository {
return rows.map(serializeDescriptionGenerationEvent);
}
async createSensitiveDataSuggestionRun(
async createSensitivityAnalysisRun(
databaseId: string,
scope: SensitiveDataSuggestionScope,
modelId: string,
): Promise<SensitiveDataSuggestionRun> {
scope: SensitivityAnalysisScope,
origin: { engine: "llm"; modelId: string } | { engine: "local"; policyVersion: string },
): Promise<SensitivityAnalysisRun> {
const row = await this.db.insertInto("sensitiveDataSuggestionRuns").values({
id: randomUUID(),
databaseId,
scope,
modelId,
engine: origin.engine,
modelId: origin.engine === "llm" ? origin.modelId : null,
policyVersion: origin.engine === "local" ? origin.policyVersion : null,
status: "running",
total: 0,
suggestedSensitive: 0,
suggestedNonSensitive: 0,
unknown: 0,
inputTokens: 0,
cacheReadTokens: 0,
outputTokens: 0,
finishedAt: null,
errorSummary: null,
}).returningAll().executeTakeFirstOrThrow();
return serializeSensitiveDataSuggestionRun(row);
return serializeSensitivityAnalysisRun(row);
}
async getSensitiveDataSuggestionRun(
async getSensitivityAnalysisRun(
runId: string,
): Promise<SensitiveDataSuggestionRun | undefined> {
): Promise<SensitivityAnalysisRun | undefined> {
const row = await this.db.selectFrom("sensitiveDataSuggestionRuns")
.selectAll()
.where("id", "=", runId)
.executeTakeFirst();
return row ? serializeSensitiveDataSuggestionRun(row) : undefined;
return row ? serializeSensitivityAnalysisRun(row) : undefined;
}
async listSensitiveDataSuggestionRuns(limit = 50): Promise<SensitiveDataSuggestionRun[]> {
async listSensitivityAnalysisRuns(limit = 50): Promise<SensitivityAnalysisRun[]> {
const rows = await this.db.selectFrom("sensitiveDataSuggestionRuns")
.selectAll()
.orderBy("createdAt", "desc")
.orderBy("id", "desc")
.limit(limit)
.execute();
return rows.map(serializeSensitiveDataSuggestionRun);
return rows.map(serializeSensitivityAnalysisRun);
}
async interruptActiveSensitiveDataSuggestionRuns(
async interruptActiveSensitivityAnalysisRuns(
errorSummary: string,
): Promise<SensitiveDataSuggestionRun[]> {
): Promise<SensitivityAnalysisRun[]> {
const rows = await this.db.updateTable("sensitiveDataSuggestionRuns")
.set({
status: "interrupted",
@@ -994,34 +1021,34 @@ export class KyselyCatalogRepository implements CatalogRepository {
.where("status", "=", "running")
.returningAll()
.execute();
return rows.map(serializeSensitiveDataSuggestionRun);
return rows.map(serializeSensitivityAnalysisRun);
}
async updateSensitiveDataSuggestionRun(
async updateSensitivityAnalysisRun(
runId: string,
update: SensitiveDataSuggestionRunUpdate,
): Promise<SensitiveDataSuggestionRun | undefined> {
update: SensitivityAnalysisRunUpdate,
): Promise<SensitivityAnalysisRun | undefined> {
const values: any = { ...update, updatedAt: sql`now()` };
const row = await this.db.updateTable("sensitiveDataSuggestionRuns")
.set(values)
.where("id", "=", runId)
.returningAll()
.executeTakeFirst();
return row ? serializeSensitiveDataSuggestionRun(row) : undefined;
return row ? serializeSensitivityAnalysisRun(row) : undefined;
}
async appendSensitiveDataSuggestionEvent(
async appendSensitivityAnalysisEvent(
runId: string,
level: SensitiveDataSuggestionEvent["level"],
level: SensitivityAnalysisEvent["level"],
message: string,
): Promise<SensitiveDataSuggestionEvent> {
): Promise<SensitivityAnalysisEvent> {
return await this.db.transaction().execute(async (trx) => {
const run = await trx.selectFrom("sensitiveDataSuggestionRuns")
.select("id")
.where("id", "=", runId)
.forUpdate()
.executeTakeFirst();
if (!run) throw new CatalogConflictError("Sensitive Data Suggestion Run does not exist");
if (!run) throw new CatalogConflictError("Sensitivity Analysis Run does not exist");
const current = await trx.selectFrom("sensitiveDataSuggestionEvents")
.select(sql<number>`coalesce(max(sequence), 0)::int`.as("sequence"))
.where("runId", "=", runId)
@@ -1032,21 +1059,21 @@ export class KyselyCatalogRepository implements CatalogRepository {
level,
message,
}).returningAll().executeTakeFirstOrThrow();
return serializeSensitiveDataSuggestionEvent(row);
return serializeSensitivityAnalysisEvent(row);
});
}
async listSensitiveDataSuggestionEvents(
async listSensitivityAnalysisEvents(
runId: string,
afterSequence = 0,
): Promise<SensitiveDataSuggestionEvent[]> {
): Promise<SensitivityAnalysisEvent[]> {
const rows = await this.db.selectFrom("sensitiveDataSuggestionEvents")
.selectAll()
.where("runId", "=", runId)
.where("sequence", ">", afterSequence)
.orderBy("sequence")
.execute();
return rows.map(serializeSensitiveDataSuggestionEvent);
return rows.map(serializeSensitivityAnalysisEvent);
}
async listRelationships(databaseId: string): Promise<CatalogPhysicalRelationship[]> {
@@ -1254,16 +1281,25 @@ export class KyselyCatalogRepository implements CatalogRepository {
if (databases.length !== selectedDatabaseIds.length) return undefined;
if (target === "relationships") {
const count = await trx.selectFrom("catalogRelationships")
const physicalCount = await trx.selectFrom("catalogRelationships")
.select(sql<number>`count(*)::int`.as("count"))
.where("databaseId", "in", selectedDatabaseIds).executeTakeFirst();
const logicalCount = await trx.selectFrom("catalogLogicalRelationships")
.select(sql<number>`count(*)::int`.as("count"))
.where("databaseId", "in", selectedDatabaseIds).executeTakeFirst();
await trx.deleteFrom("catalogLogicalRelationships")
.where("databaseId", "in", selectedDatabaseIds).execute();
await trx.deleteFrom("catalogRelationships")
.where("databaseId", "in", selectedDatabaseIds).execute();
await trx.updateTable("workspaceDatabases").set({
schemaSyncedVersion: null,
schemaSyncedAt: null,
}).where("id", "in", selectedDatabaseIds).execute();
return { tables: 0, columns: 0, relationships: Number(count?.count ?? 0) };
return {
tables: 0,
columns: 0,
relationships: Number(physicalCount?.count ?? 0) + Number(logicalCount?.count ?? 0),
};
}
const tableCount = await trx.selectFrom("catalogTables")
@@ -1273,7 +1309,10 @@ export class KyselyCatalogRepository implements CatalogRepository {
.innerJoin("catalogTables", "catalogTables.id", "catalogColumns.tableId")
.select(sql<number>`count(*)::int`.as("count"))
.where("catalogTables.databaseId", "in", selectedDatabaseIds).executeTakeFirst();
const relationshipCount = await trx.selectFrom("catalogRelationships")
const physicalRelationshipCount = await trx.selectFrom("catalogRelationships")
.select(sql<number>`count(*)::int`.as("count"))
.where("databaseId", "in", selectedDatabaseIds).executeTakeFirst();
const logicalRelationshipCount = await trx.selectFrom("catalogLogicalRelationships")
.select(sql<number>`count(*)::int`.as("count"))
.where("databaseId", "in", selectedDatabaseIds).executeTakeFirst();
await trx.deleteFrom("catalogTables")
@@ -1285,7 +1324,8 @@ export class KyselyCatalogRepository implements CatalogRepository {
return {
tables: Number(tableCount?.count ?? 0),
columns: Number(columnCount?.count ?? 0),
relationships: Number(relationshipCount?.count ?? 0),
relationships: Number(physicalRelationshipCount?.count ?? 0)
+ Number(logicalRelationshipCount?.count ?? 0),
};
});
}
@@ -1319,13 +1359,34 @@ export class KyselyCatalogRepository implements CatalogRepository {
return { tables: 0, columns: Number(count?.count ?? 0), relationships: 0 };
}
const count = await trx.selectFrom("catalogRelationships")
const physicalCount = await trx.selectFrom("catalogRelationships")
.select(sql<number>`count(*)::int`.as("count"))
.where("databaseId", "=", databaseId)
.where((eb) => eb.or([
eb("sourceTableId", "in", selectedTableIds),
eb("targetTableId", "in", selectedTableIds),
])).executeTakeFirst();
const selectedColumnIds = (await trx.selectFrom("catalogColumns")
.select("id")
.where("tableId", "in", selectedTableIds)
.execute()).map((column) => column.id);
const logicalCount = selectedColumnIds.length === 0
? undefined
: await trx.selectFrom("catalogLogicalRelationships")
.select(sql<number>`count(*)::int`.as("count"))
.where("databaseId", "=", databaseId)
.where((eb) => eb.or([
eb("sourceColumnId", "in", selectedColumnIds),
eb("targetColumnId", "in", selectedColumnIds),
])).executeTakeFirst();
if (selectedColumnIds.length > 0) {
await trx.deleteFrom("catalogLogicalRelationships")
.where("databaseId", "=", databaseId)
.where((eb) => eb.or([
eb("sourceColumnId", "in", selectedColumnIds),
eb("targetColumnId", "in", selectedColumnIds),
])).execute();
}
await trx.deleteFrom("catalogRelationships")
.where("databaseId", "=", databaseId)
.where((eb) => eb.or([
@@ -1336,7 +1397,11 @@ export class KyselyCatalogRepository implements CatalogRepository {
schemaSyncedVersion: null,
schemaSyncedAt: null,
}).where("id", "=", databaseId).execute();
return { tables: 0, columns: 0, relationships: Number(count?.count ?? 0) };
return {
tables: 0,
columns: 0,
relationships: Number(physicalCount?.count ?? 0) + Number(logicalCount?.count ?? 0),
};
});
}
@@ -1792,13 +1857,13 @@ export class UnavailableCatalogRepository implements CatalogRepository {
async updateDescriptionGenerationRun(): Promise<DescriptionGenerationRun | undefined> { return this.fail(); }
async appendDescriptionGenerationEvent(): Promise<DescriptionGenerationEvent> { return this.fail(); }
async listDescriptionGenerationEvents(): Promise<DescriptionGenerationEvent[]> { return this.fail(); }
async createSensitiveDataSuggestionRun(): Promise<SensitiveDataSuggestionRun> { return this.fail(); }
async getSensitiveDataSuggestionRun(): Promise<SensitiveDataSuggestionRun | undefined> { return this.fail(); }
async listSensitiveDataSuggestionRuns(): Promise<SensitiveDataSuggestionRun[]> { return this.fail(); }
async interruptActiveSensitiveDataSuggestionRuns(): Promise<SensitiveDataSuggestionRun[]> { return this.fail(); }
async updateSensitiveDataSuggestionRun(): Promise<SensitiveDataSuggestionRun | undefined> { return this.fail(); }
async appendSensitiveDataSuggestionEvent(): Promise<SensitiveDataSuggestionEvent> { return this.fail(); }
async listSensitiveDataSuggestionEvents(): Promise<SensitiveDataSuggestionEvent[]> { return this.fail(); }
async createSensitivityAnalysisRun(): Promise<SensitivityAnalysisRun> { return this.fail(); }
async getSensitivityAnalysisRun(): Promise<SensitivityAnalysisRun | undefined> { return this.fail(); }
async listSensitivityAnalysisRuns(): Promise<SensitivityAnalysisRun[]> { return this.fail(); }
async interruptActiveSensitivityAnalysisRuns(): Promise<SensitivityAnalysisRun[]> { return this.fail(); }
async updateSensitivityAnalysisRun(): Promise<SensitivityAnalysisRun | undefined> { return this.fail(); }
async appendSensitivityAnalysisEvent(): Promise<SensitivityAnalysisEvent> { return this.fail(); }
async listSensitivityAnalysisEvents(): Promise<SensitivityAnalysisEvent[]> { return this.fail(); }
async listRelationships(): Promise<CatalogPhysicalRelationship[]> { return this.fail(); }
async listLogicalRelationships(): Promise<CatalogLogicalRelationship[]> { return this.fail(); }
async getLogicalRelationshipContext(): Promise<CatalogLogicalRelationshipContext | undefined> { return this.fail(); }
@@ -1,254 +0,0 @@
import { z } from "zod";
import type { MetadataGenerationModels } from "./metadata-generation-models.js";
import type { ModelCompleter, ModelCompletionMessage, ModelCompletionResult, ModelCompletionUsage } from "./model-completer.js";
import type {
CatalogColumn,
CatalogRepository,
CatalogTable,
SensitiveDataSuggestionScope,
} from "./types.js";
export type { SensitiveDataSuggestionScope } from "./types.js";
// The helper accepts at most 64 KiB per message. Keep the same safety margin used by
// Description Generation so UTF-8 structural metadata never reaches that hard limit.
const MAX_USER_MESSAGE_BYTES = 60 * 1024;
// Preserve ThothAI's proven completion granularity: small batches keep generation time and
// structured-output accuracy predictable even when the helper byte limit would allow much more.
const MAX_COLUMNS_PER_BATCH = 10;
const responseSchema = z.object({
suggestions: z.array(z.object({
columnId: z.uuid(),
sensitive: z.boolean(),
}).strict()),
}).strict();
interface StructuralColumn {
columnId: string;
tableId: string;
table: string;
column: string;
dataType: string;
nullable: boolean;
primaryKey: boolean;
foreignKey: boolean;
version: number;
currentSensitive: boolean;
}
export interface SensitiveDataSuggestion {
columnId: string;
tableId: string;
tableName: string;
columnName: string;
version: number;
currentSensitive: boolean;
sensitive: boolean;
}
export class SensitiveDataSuggestionTargetNotFoundError extends Error {
constructor(readonly target: "database" | "table" | "column") {
super(`${target} not found`);
this.name = "SensitiveDataSuggestionTargetNotFoundError";
}
}
export class SensitiveDataSuggestionDuplicateTargetIdsError extends Error {
constructor() {
super("sensitive-data suggestion target IDs must be unique");
this.name = "SensitiveDataSuggestionDuplicateTargetIdsError";
}
}
export class SensitiveDataSuggestionNoEligibleColumnsError extends Error {
constructor(readonly scope: SensitiveDataSuggestionScope) {
super("selected scope has no catalog columns");
this.name = "SensitiveDataSuggestionNoEligibleColumnsError";
}
}
export class SensitiveDataSuggestionPayloadTooLargeError extends Error {
constructor() {
super("sensitive-data suggestion structural metadata is too large");
this.name = "SensitiveDataSuggestionPayloadTooLargeError";
}
}
export class SensitiveDataSuggestionInvalidResponseError extends Error {
constructor() {
super("sensitive-data suggestion response is invalid");
this.name = "SensitiveDataSuggestionInvalidResponseError";
}
}
function userContent(
database: { databaseName: string; schema: string },
columns: readonly StructuralColumn[],
): string {
return JSON.stringify({
database: database.databaseName,
schema: database.schema,
columns: columns.map((column) => ({
columnId: column.columnId,
table: column.table,
column: column.column,
dataType: column.dataType,
nullable: column.nullable,
primaryKey: column.primaryKey,
foreignKey: column.foreignKey,
})),
});
}
function batchesFor(
database: { databaseName: string; schema: string },
columns: readonly StructuralColumn[],
): StructuralColumn[][] {
const batches: StructuralColumn[][] = [];
let current: StructuralColumn[] = [];
for (const column of columns) {
if (current.length === MAX_COLUMNS_PER_BATCH) {
batches.push(current);
current = [];
}
const candidate = [...current, column];
if (Buffer.byteLength(userContent(database, candidate), "utf8") <= MAX_USER_MESSAGE_BYTES) {
current = candidate;
continue;
}
if (current.length === 0) throw new SensitiveDataSuggestionPayloadTooLargeError();
batches.push(current);
current = [column];
if (Buffer.byteLength(userContent(database, current), "utf8") > MAX_USER_MESSAGE_BYTES) {
throw new SensitiveDataSuggestionPayloadTooLargeError();
}
}
if (current.length > 0) batches.push(current);
return batches;
}
function structuralColumn(table: CatalogTable, column: CatalogColumn): StructuralColumn {
return {
columnId: column.id,
tableId: table.id,
table: table.name,
column: column.name,
dataType: column.dataType,
nullable: column.isNullable,
primaryKey: column.isPrimaryKey,
foreignKey: column.isForeignKey,
version: column.version,
currentSensitive: column.sensitive,
};
}
const systemMessage: ModelCompletionMessage = {
role: "system",
content: [
"Classify whether each database column is likely to contain sensitive source values.",
"Use only the supplied structural metadata. Return strict JSON with this exact shape:",
'{"suggestions":[{"columnId":"uuid","sensitive":true}]}',
"Return every supplied column exactly once. Do not add explanations or markdown.",
].join("\n"),
};
export class SensitiveDataSuggester {
constructor(
private readonly repository: CatalogRepository,
private readonly models: MetadataGenerationModels,
private readonly completer: ModelCompleter,
) {}
private async selectColumns(
databaseId: string,
scope: SensitiveDataSuggestionScope,
targetIds: readonly string[],
): Promise<StructuralColumn[]> {
if (new Set(targetIds).size !== targetIds.length) {
throw new SensitiveDataSuggestionDuplicateTargetIdsError();
}
const tables = await this.repository.listTables(databaseId);
const tableIds = new Set(targetIds);
const selectedTables = scope === "selected_tables"
? tables.filter((table) => tableIds.has(table.id))
: tables;
if (scope === "selected_tables" && selectedTables.length !== targetIds.length) {
throw new SensitiveDataSuggestionTargetNotFoundError("table");
}
const columns = (await Promise.all(selectedTables.map(async (table) => (
(await this.repository.listColumns(databaseId, table.id)).map((column) => (
structuralColumn(table, column)
))
)))).flat();
const columnIds = new Set(targetIds);
const selectedColumns = scope === "selected_columns"
? columns.filter((column) => columnIds.has(column.columnId))
: columns;
if (scope === "selected_columns" && selectedColumns.length !== targetIds.length) {
throw new SensitiveDataSuggestionTargetNotFoundError("column");
}
if (selectedColumns.length === 0) {
throw new SensitiveDataSuggestionNoEligibleColumnsError(scope);
}
return selectedColumns;
}
async suggest(
databaseId: string,
modelId: string,
scope: SensitiveDataSuggestionScope,
targetIds: readonly string[],
signal: AbortSignal,
onPrepared?: (total: number) => void | Promise<void>,
onProgress?: (processed: number, suggestions: readonly SensitiveDataSuggestion[]) => void | Promise<void>,
onUsage?: (usage: ModelCompletionUsage) => void | Promise<void>,
): Promise<readonly SensitiveDataSuggestion[]> {
const database = await this.repository.get(databaseId);
if (!database) throw new SensitiveDataSuggestionTargetNotFoundError("database");
const columns = await this.selectColumns(databaseId, scope, targetIds);
await onPrepared?.(columns.length);
const model = this.models.resolve(modelId);
const suggestions: SensitiveDataSuggestion[] = [];
for (const batch of batchesFor(database, columns)) {
let received: Map<string, { columnId: string; sensitive: boolean }> | undefined;
for (let attempt = 0; attempt < 2 && !received; attempt += 1) {
const completion = await this.completer.complete({
model,
signal,
messages: [systemMessage, { role: "user", content: userContent(database, batch) }],
});
const result: ModelCompletionResult = typeof completion === "string"
? { content: completion, usage: { input: 0, cacheRead: 0, output: 0 } }
: completion;
await onUsage?.(result.usage);
const content = result.content;
try {
const parsed = responseSchema.parse(JSON.parse(content));
const expected = new Set(batch.map((column) => column.columnId));
const candidate = new Map(parsed.suggestions.map((suggestion) => [suggestion.columnId, suggestion]));
if (candidate.size !== parsed.suggestions.length
|| candidate.size !== expected.size
|| [...candidate.keys()].some((columnId) => !expected.has(columnId))) {
throw new SensitiveDataSuggestionInvalidResponseError();
}
received = candidate;
} catch {
if (attempt === 1) throw new SensitiveDataSuggestionInvalidResponseError();
}
}
suggestions.push(...batch.map((column) => ({
columnId: column.columnId,
tableId: column.tableId,
tableName: column.table,
columnName: column.column,
version: column.version,
currentSensitive: column.currentSensitive,
sensitive: received!.get(column.columnId)!.sensitive,
})));
await onProgress?.(suggestions.length, suggestions.slice(-batch.length));
}
return suggestions;
}
}
@@ -1,136 +0,0 @@
import type {
SensitiveDataSuggestion,
} from "./sensitive-data-suggester.js";
import {
SensitiveDataSuggester,
SensitiveDataSuggestionTargetNotFoundError,
} from "./sensitive-data-suggester.js";
import type {
CatalogRepository,
SensitiveDataSuggestionRun,
SensitiveDataSuggestionScope,
} from "./types.js";
import type { ModelCompletionUsage } from "./model-completer.js";
const interruptedMessage = "Sensitive-field suggestion generation was interrupted by backend restart.";
const failedMessage = "Sensitive-field suggestion generation failed.";
export interface SensitiveDataSuggestionRunResult {
suggestions: readonly SensitiveDataSuggestion[];
run: SensitiveDataSuggestionRun;
}
export class SensitiveDataSuggestionRunner {
constructor(
private readonly repository: CatalogRepository,
private readonly suggester: SensitiveDataSuggester,
) {}
async initialize(): Promise<void> {
if (!(await this.repository.available())) return;
const interrupted = await this.repository.interruptActiveSensitiveDataSuggestionRuns(
interruptedMessage,
);
for (const run of interrupted) {
await this.repository.appendSensitiveDataSuggestionEvent(
run.id,
"warning",
interruptedMessage,
);
}
}
async run(
databaseId: string,
modelId: string,
scope: SensitiveDataSuggestionScope,
targetIds: readonly string[],
signal: AbortSignal,
): Promise<SensitiveDataSuggestionRunResult> {
if (!(await this.repository.get(databaseId))) {
throw new SensitiveDataSuggestionTargetNotFoundError("database");
}
const started = await this.repository.createSensitiveDataSuggestionRun(
databaseId,
scope,
modelId,
);
try {
await this.repository.appendSensitiveDataSuggestionEvent(
started.id,
"info",
"Sensitive-field suggestion generation started.",
);
const suggestions = await this.suggester.suggest(
databaseId,
modelId,
scope,
targetIds,
signal,
async (total) => {
const prepared = await this.repository.updateSensitiveDataSuggestionRun(started.id, {
total,
});
if (!prepared) throw new Error("Sensitive Data Suggestion Run disappeared");
},
async (processed, batch) => {
const suggestedSensitive = batch.filter((suggestion) => suggestion.sensitive).length;
const suggestedNonSensitive = batch.length - suggestedSensitive;
const current = await this.repository.getSensitiveDataSuggestionRun(started.id);
if (!current) throw new Error("Sensitive Data Suggestion Run disappeared");
const progress = await this.repository.updateSensitiveDataSuggestionRun(started.id, {
suggestedSensitive: current.suggestedSensitive + suggestedSensitive,
suggestedNonSensitive: current.suggestedNonSensitive + suggestedNonSensitive,
});
if (!progress) throw new Error("Sensitive Data Suggestion Run disappeared");
await this.repository.appendSensitiveDataSuggestionEvent(
started.id,
"info",
`Classified ${processed} of ${progress.total} columns.`,
);
},
async (usage: ModelCompletionUsage) => {
const current = await this.repository.getSensitiveDataSuggestionRun(started.id);
if (!current) throw new Error("Sensitive Data Suggestion Run disappeared");
await this.repository.updateSensitiveDataSuggestionRun(started.id, {
inputTokens: current.inputTokens + usage.input,
cacheReadTokens: current.cacheReadTokens + usage.cacheRead,
outputTokens: current.outputTokens + usage.output,
});
},
);
const suggestedSensitive = suggestions.filter((suggestion) => suggestion.sensitive).length;
const suggestedNonSensitive = suggestions.length - suggestedSensitive;
await this.repository.appendSensitiveDataSuggestionEvent(
started.id,
"info",
`Sensitive-field suggestion generation completed for ${suggestions.length} column${
suggestions.length === 1 ? "" : "s"
}.`,
);
const completed = await this.repository.updateSensitiveDataSuggestionRun(started.id, {
status: "completed",
total: suggestions.length,
suggestedSensitive,
suggestedNonSensitive,
finishedAt: new Date().toISOString(),
errorSummary: null,
});
if (!completed) throw new Error("Sensitive Data Suggestion Run disappeared");
return { suggestions, run: completed };
} catch (error) {
await this.repository.updateSensitiveDataSuggestionRun(started.id, {
status: "failed",
finishedAt: new Date().toISOString(),
errorSummary: failedMessage,
}).catch(() => undefined);
await this.repository.appendSensitiveDataSuggestionEvent(
started.id,
"error",
failedMessage,
).catch(() => undefined);
throw error;
}
}
}
@@ -0,0 +1,178 @@
import type {
SensitivityReviewItem,
} from "./sensitivity-analysis-service.js";
import {
SENSITIVITY_POLICY_VERSION,
SensitivityAnalysisInterruptedError,
SensitivityAnalysisService,
SensitivityAnalysisTargetNotFoundError,
} from "./sensitivity-analysis-service.js";
import type {
CatalogRepository,
SensitivityAnalysisRun,
SensitivityAnalysisScope,
} from "./types.js";
const interruptedMessage = "Local sensitivity analysis was interrupted by backend restart.";
const interruptedDuringRunMessage = "Local sensitivity analysis was interrupted before completion.";
const failedMessage = "Local sensitivity analysis failed.";
function ensureActive(signal: AbortSignal): void {
if (signal.aborted) throw new SensitivityAnalysisInterruptedError();
}
export interface SensitivityAnalysisRunResult {
suggestions: readonly SensitivityReviewItem[];
run: SensitivityAnalysisRun;
}
export class SensitivityAnalysisRunner {
constructor(
private readonly repository: CatalogRepository,
private readonly analysis: SensitivityAnalysisService,
) {}
async initialize(): Promise<void> {
if (!(await this.repository.available())) return;
const interrupted = await this.repository.interruptActiveSensitivityAnalysisRuns(
interruptedMessage,
);
for (const run of interrupted) {
await this.repository.appendSensitivityAnalysisEvent(
run.id,
"warning",
interruptedMessage,
);
}
}
async run(
databaseId: string,
scope: SensitivityAnalysisScope,
targetIds: readonly string[],
signal: AbortSignal,
): Promise<SensitivityAnalysisRunResult> {
ensureActive(signal);
const database = await this.repository.get(databaseId);
ensureActive(signal);
if (!database) {
throw new SensitivityAnalysisTargetNotFoundError("database");
}
ensureActive(signal);
const started = await this.repository.createSensitivityAnalysisRun(
databaseId,
scope,
{ engine: "local", policyVersion: SENSITIVITY_POLICY_VERSION },
);
let preparedTotal = 0;
let processedSensitive = 0;
let processedNonSensitive = 0;
try {
ensureActive(signal);
await this.repository.appendSensitivityAnalysisEvent(
started.id,
"info",
"Local sensitivity analysis started.",
);
ensureActive(signal);
const suggestions = await this.analysis.analyze(
databaseId,
scope,
targetIds,
signal,
async (total) => {
ensureActive(signal);
preparedTotal = total;
const prepared = await this.repository.updateSensitivityAnalysisRun(started.id, {
total,
});
ensureActive(signal);
if (!prepared) throw new Error("Sensitivity Analysis Run disappeared");
},
async (processed, batch) => {
ensureActive(signal);
const suggestedSensitive = batch.filter(
(suggestion) => suggestion.assessment === "sensitive",
).length;
const suggestedNonSensitive = batch.filter(
(suggestion) => suggestion.assessment === "non_sensitive",
).length;
const current = await this.repository.getSensitivityAnalysisRun(started.id);
ensureActive(signal);
if (!current) throw new Error("Sensitivity Analysis Run disappeared");
const progress = await this.repository.updateSensitivityAnalysisRun(started.id, {
suggestedSensitive: current.suggestedSensitive + suggestedSensitive,
suggestedNonSensitive: current.suggestedNonSensitive + suggestedNonSensitive,
});
if (!progress) throw new Error("Sensitivity Analysis Run disappeared");
processedSensitive += suggestedSensitive;
processedNonSensitive += suggestedNonSensitive;
ensureActive(signal);
await this.repository.appendSensitivityAnalysisEvent(
started.id,
"info",
`Assessed ${processed} of ${progress.total} columns locally.`,
);
ensureActive(signal);
},
async (message) => {
ensureActive(signal);
await this.repository.appendSensitivityAnalysisEvent(
started.id,
"info",
message,
);
ensureActive(signal);
},
);
ensureActive(signal);
const suggestedSensitive = suggestions.filter(
(suggestion) => suggestion.assessment === "sensitive",
).length;
const suggestedNonSensitive = suggestions.filter(
(suggestion) => suggestion.assessment === "non_sensitive",
).length;
await this.repository.appendSensitivityAnalysisEvent(
started.id,
"info",
`Local sensitivity analysis completed for ${suggestions.length} column${
suggestions.length === 1 ? "" : "s"
}.`,
);
ensureActive(signal);
const completed = await this.repository.updateSensitivityAnalysisRun(started.id, {
status: "completed",
total: suggestions.length,
suggestedSensitive,
suggestedNonSensitive,
unknown: 0,
finishedAt: new Date().toISOString(),
errorSummary: null,
});
ensureActive(signal);
if (!completed) throw new Error("Sensitivity Analysis Run disappeared");
return { suggestions, run: completed };
} catch (error) {
const interrupted = signal.aborted || error instanceof SensitivityAnalysisInterruptedError;
const message = interrupted ? interruptedDuringRunMessage : failedMessage;
await this.repository.updateSensitivityAnalysisRun(started.id, {
status: interrupted ? "interrupted" : "failed",
...(interrupted ? {
total: preparedTotal,
suggestedSensitive: processedSensitive,
suggestedNonSensitive: processedNonSensitive,
unknown: Math.max(0, preparedTotal - processedSensitive - processedNonSensitive),
} : {}),
finishedAt: new Date().toISOString(),
errorSummary: message,
}).catch(() => undefined);
await this.repository.appendSensitivityAnalysisEvent(
started.id,
interrupted ? "warning" : "error",
message,
).catch(() => undefined);
throw error;
}
}
}
@@ -0,0 +1,184 @@
import type {
SensitivityClassifier,
SensitivityColumnAssessment,
SensitivityEvidence,
SensitivityNerBudget,
} from "./sensitivity-classifier.js";
import type {
CatalogColumn,
CatalogRepository,
CatalogTable,
SensitivityAnalysisScope,
} from "./types.js";
export type { SensitivityAnalysisScope } from "./types.js";
export const SENSITIVITY_POLICY_VERSION = "sensitivity-v4";
interface SelectedColumn {
table: CatalogTable;
column: CatalogColumn;
}
export interface SensitivityReviewItem {
columnId: string;
tableId: string;
tableName: string;
columnName: string;
version: number;
currentSensitive: boolean;
sensitive: boolean;
assessment: SensitivityColumnAssessment["assessment"];
evidence: readonly SensitivityEvidence[];
observedValues: number;
coverage: SensitivityColumnAssessment["coverage"];
}
export class SensitivityAnalysisTargetNotFoundError extends Error {
constructor(readonly target: "database" | "table" | "column") {
super(`${target} not found`);
this.name = "SensitivityAnalysisTargetNotFoundError";
}
}
export class SensitivityAnalysisDuplicateTargetIdsError extends Error {
constructor() {
super("sensitivity analysis target IDs must be unique");
this.name = "SensitivityAnalysisDuplicateTargetIdsError";
}
}
export class SensitivityAnalysisInterruptedError extends Error {
constructor() {
super("sensitivity analysis interrupted");
this.name = "SensitivityAnalysisInterruptedError";
}
}
function ensureActive(signal: AbortSignal): void {
if (signal.aborted) throw new SensitivityAnalysisInterruptedError();
}
export class SensitivityAnalysisNoEligibleColumnsError extends Error {
constructor(readonly scope: SensitivityAnalysisScope) {
super("selected scope has no catalog columns");
this.name = "SensitivityAnalysisNoEligibleColumnsError";
}
}
/** Selection and table orchestration around the single SensitivityClassifier decision module. */
export class SensitivityAnalysisService {
constructor(
private readonly repository: CatalogRepository,
private readonly classifier: SensitivityClassifier,
private readonly options: { nerBudgetMs?: number } = {},
) {}
private async selectColumns(
databaseId: string,
scope: SensitivityAnalysisScope,
targetIds: readonly string[],
signal: AbortSignal,
): Promise<readonly SelectedColumn[]> {
ensureActive(signal);
if (new Set(targetIds).size !== targetIds.length) {
throw new SensitivityAnalysisDuplicateTargetIdsError();
}
const tables = await this.repository.listTables(databaseId);
ensureActive(signal);
const tableIds = new Set(targetIds);
const selectedTables = scope === "selected_tables"
? tables.filter((table) => tableIds.has(table.id))
: tables;
if (scope === "selected_tables" && selectedTables.length !== targetIds.length) {
throw new SensitivityAnalysisTargetNotFoundError("table");
}
const columns = (await Promise.all(selectedTables.map(async (table) => (
(await this.repository.listColumns(databaseId, table.id)).map((column) => ({ table, column }))
)))).flat();
ensureActive(signal);
const columnIds = new Set(targetIds);
const selectedColumns = scope === "selected_columns"
? columns.filter(({ column }) => columnIds.has(column.id))
: columns;
if (scope === "selected_columns" && selectedColumns.length !== targetIds.length) {
throw new SensitivityAnalysisTargetNotFoundError("column");
}
if (selectedColumns.length === 0) {
throw new SensitivityAnalysisNoEligibleColumnsError(scope);
}
return selectedColumns;
}
async analyze(
databaseId: string,
scope: SensitivityAnalysisScope,
targetIds: readonly string[],
signal: AbortSignal,
onPrepared?: (total: number) => void | Promise<void>,
onProgress?: (processed: number, suggestions: readonly SensitivityReviewItem[]) => void | Promise<void>,
onActivity?: (message: string) => void | Promise<void>,
): Promise<readonly SensitivityReviewItem[]> {
const configuredNerBudget = this.options.nerBudgetMs ?? 10_000;
const nerBudget: SensitivityNerBudget = {
remainingMs: Number.isFinite(configuredNerBudget) && configuredNerBudget >= 0
? configuredNerBudget
: 10_000,
};
ensureActive(signal);
const database = await this.repository.get(databaseId);
ensureActive(signal);
if (!database) throw new SensitivityAnalysisTargetNotFoundError("database");
const selected = await this.selectColumns(databaseId, scope, targetIds, signal);
await onPrepared?.(selected.length);
ensureActive(signal);
const byTable = new Map<string, SelectedColumn[]>();
for (const item of selected) {
const items = byTable.get(item.table.id) ?? [];
items.push(item);
byTable.set(item.table.id, items);
}
const tableTargets = [...byTable.values()].map((items) => {
const first = items[0]!;
return {
database,
table: first.table,
columns: items.map(({ column }) => column),
};
});
const assessments = await this.classifier.assess(
tableTargets,
signal,
nerBudget,
onActivity,
);
ensureActive(signal);
const assessmentById = new Map(assessments.map((assessment) => [
assessment.columnId,
assessment,
]));
const suggestions: SensitivityReviewItem[] = [];
for (const items of byTable.values()) {
ensureActive(signal);
const batch = items.map(({ table, column }) => {
const assessment = assessmentById.get(column.id)!;
return {
columnId: column.id,
tableId: table.id,
tableName: table.name,
columnName: column.name,
version: column.version,
currentSensitive: column.sensitive,
sensitive: assessment.proposedSensitive,
assessment: assessment.assessment,
evidence: assessment.evidence,
observedValues: assessment.observedValues,
coverage: assessment.coverage,
};
});
suggestions.push(...batch);
await onProgress?.(suggestions.length, batch);
ensureActive(signal);
}
return suggestions;
}
}
@@ -0,0 +1,535 @@
import type { CatalogColumn, CatalogTable, WorkspaceDatabase } from "./types.js";
import { findPhoneNumbersInText } from "libphonenumber-js/max";
import validator from "validator";
export type SensitivityAssessment = "sensitive" | "non_sensitive";
export interface SensitivityEvidence {
kind: "metadata" | "content" | "length" | "ner" | "coverage" | "type";
ruleId: string;
label?: string;
confidence?: number;
}
export interface SensitivityValueObservation {
columnId: string;
value: string | null;
characterLength: number | null;
}
export interface SensitivityScanCoverage {
kind: "complete" | "sampled";
observedValues: number;
}
export interface SensitivityTableScan {
batches: readonly (readonly SensitivityValueObservation[])[];
coverage: SensitivityScanCoverage;
}
export interface SensitivityScanRequest {
database: WorkspaceDatabase;
table: CatalogTable;
columns: readonly CatalogColumn[];
valuesPerColumn: number;
sampleOffset: number;
sampleSeed: number;
queryTimeoutMs: number;
fullScanThreshold?: number;
}
export interface SensitivityValueSource {
scanTable(
request: SensitivityScanRequest,
consume: (batch: readonly SensitivityValueObservation[]) => void | Promise<void>,
signal: AbortSignal,
): Promise<SensitivityScanCoverage>;
}
export interface LocalNerCandidate {
columnId: string;
text: string;
}
export interface LocalNerEvidence {
columnId: string;
label: string;
confidence: number;
}
export interface SensitivityNerBudget {
remainingMs: number;
}
/** Optional local detector. It returns evidence only; it never decides a column assessment. */
export interface LocalNerDetector {
warmup?(): Promise<void>;
isReady?(): boolean;
detect(
candidates: readonly LocalNerCandidate[],
signal: AbortSignal,
deadline: number,
): Promise<readonly LocalNerEvidence[]>;
close?(): Promise<void>;
}
export interface SensitivityColumnAssessment {
columnId: string;
assessment: SensitivityAssessment;
proposedSensitive: boolean;
evidence: readonly SensitivityEvidence[];
observedValues: number;
coverage: "metadata" | "complete" | "sampled" | "no_values";
}
export interface SensitivityTableTarget {
database: WorkspaceDatabase;
table: CatalogTable;
columns: readonly CatalogColumn[];
}
const EMAIL = /(?<![\p{L}\p{N}._%+-])[\p{L}\p{N}._%+-]+@[\p{L}\p{N}.-]+\.[\p{L}]{2,63}(?![\p{L}\p{N}._%+-])/giu;
const DIRECT_IDENTIFIER_NAMES = new Set([
"address", "birth_date", "codice_fiscale", "date_of_birth", "dob", "email", "e_mail",
"bic", "first_name", "fiscal_code", "full_name", "iban", "indirizzo", "last_name", "mobile",
"nome", "passport", "phone", "surname", "swift", "swift_code", "tax_id", "telefono",
]);
const CREDENTIAL_NAME = /(?:^|_)(?:api_key|credential|password|passwd|private_key|pwd|secret|token)(?:_|$)/u;
const HEALTH_NAME = /(?:^|_)(?:anamnesi|clinical|diagnos(?:i|is)|health|medical|patient|patologia|therapy|terapia)(?:_|$)/u;
const CLINICAL_TERM = /(?:^|[^\p{L}])(?:allergi[ae]|anamnesi|carcinoma|chemioterapia|diabete|diagnos[ei]|epatite|farmac[io]|gravidanza|hiv|metastasi|neoplasia|patologia|radioterapia|referto|terapia|tumore)(?:$|[^\p{L}])/iu;
const UNSUPPORTED_BINARY_TYPE = /(?:^|\s)(?:binary|blob|bytea|image|varbinary)(?:\s|$|\()/iu;
const DEEP_TEXT_TYPE = /(?:^|\s)(?:char|character|citext|clob|json|jsonb|nchar|nvarchar|string|text|varchar|xml)(?:\s|$|\()/iu;
const MAX_NER_CANDIDATES_PER_REQUEST = 128;
const MAX_CONCURRENT_TABLE_SCANS = 2;
export const SENSITIVITY_SAMPLE_PHASES = [
{ targetValuesPerColumn: 300, additionalValuesPerColumn: 300, sampleSeed: 37, deepTextOnly: false },
{ targetValuesPerColumn: 1_000, additionalValuesPerColumn: 700, sampleSeed: 73, deepTextOnly: false },
{ targetValuesPerColumn: 3_000, additionalValuesPerColumn: 2_000, sampleSeed: 109, deepTextOnly: true },
] as const;
function normalizedName(value: string): string {
return value.normalize("NFKD")
.replace(/[\u0300-\u036f]/g, "")
.replace(/([a-z0-9])([A-Z])/g, "$1_$2")
.toLocaleLowerCase("en-US")
.replace(/[^a-z0-9]+/g, "_")
.replace(/^_+|_+$/g, "");
}
function boundedCount(value: number | undefined, fallback: number, maximum: number): number {
return value === undefined || !Number.isSafeInteger(value)
? fallback
: Math.max(1, Math.min(value, maximum));
}
function metadataEvidence(column: CatalogColumn): SensitivityEvidence | undefined {
const ruleId = sensitiveNameRule(column.name);
return ruleId ? { kind: "metadata", ruleId } : undefined;
}
function nonSensitiveStructuralEvidence(column: CatalogColumn): SensitivityEvidence | undefined {
if (column.dataType.trim().toLowerCase() !== "bigint") return undefined;
if (column.isPrimaryKey || column.primaryKeyPosition !== null) {
return {
kind: "type",
ruleId: "type.bigint_primary_key_non_informative",
label: "non-informative bigint primary key",
};
}
if (normalizedName(column.name) === "pk") {
return {
kind: "metadata",
ruleId: "metadata.bigint_pk_identifier_non_informative",
label: "non-informative conventional bigint primary-key identifier",
};
}
return undefined;
}
function sensitiveNameRule(value: string): string | undefined {
const name = normalizedName(value);
if (DIRECT_IDENTIFIER_NAMES.has(name)) {
return "metadata.direct_identifier";
}
if (CREDENTIAL_NAME.test(name)) {
return "metadata.credential";
}
if (HEALTH_NAME.test(name)) {
return "metadata.health";
}
return undefined;
}
const ITALIAN_FISCAL_CODE = /(?<![A-Z0-9])[A-Z]{6}[0-9LMNPQRSTUV]{2}[ABCDEHLMPRST][0-9LMNPQRSTUV]{2}[A-Z][0-9LMNPQRSTUV]{3}[A-Z](?![A-Z0-9])/giu;
const FISCAL_ODD: Record<string, number> = {
"0": 1, "1": 0, "2": 5, "3": 7, "4": 9, "5": 13, "6": 15, "7": 17, "8": 19, "9": 21,
A: 1, B: 0, C: 5, D: 7, E: 9, F: 13, G: 15, H: 17, I: 19, J: 21,
K: 2, L: 4, M: 18, N: 20, O: 11, P: 3, Q: 6, R: 8, S: 12, T: 14,
U: 16, V: 10, W: 22, X: 25, Y: 24, Z: 23,
};
function validItalianFiscalCode(candidate: string): boolean {
const value = candidate.toUpperCase();
if (value.length !== 16) return false;
let sum = 0;
for (let index = 0; index < 15; index += 1) {
const character = value[index]!;
if (index % 2 === 0) sum += FISCAL_ODD[character] ?? -1000;
else sum += /\d/u.test(character) ? Number(character) : character.charCodeAt(0) - 65;
}
return String.fromCharCode(65 + (sum % 26)) === value[15];
}
function validIban(candidate: string): boolean {
const value = candidate.replace(/\s+/gu, "").toUpperCase();
if (!/^[A-Z]{2}\d{2}[A-Z0-9]{11,30}$/u.test(value)) return false;
const rearranged = value.slice(4) + value.slice(0, 4);
let remainder = 0;
for (const character of rearranged) {
const digits = /\d/u.test(character) ? character : String(character.charCodeAt(0) - 55);
for (const digit of digits) remainder = (remainder * 10 + Number(digit)) % 97;
}
return remainder === 1;
}
function validPaymentCard(candidate: string): boolean {
const digits = candidate.replace(/[ -]/gu, "");
if (!/^\d{13,19}$/u.test(digits) || /^(\d)\1+$/u.test(digits)) return false;
let sum = 0;
let double = false;
for (let index = digits.length - 1; index >= 0; index -= 1) {
let digit = Number(digits[index]);
if (double) {
digit *= 2;
if (digit > 9) digit -= 9;
}
sum += digit;
double = !double;
}
return sum % 10 === 0;
}
function jsonHasSensitiveKey(value: string): boolean {
const trimmed = value.trim();
if (!(trimmed.startsWith("{") || trimmed.startsWith("["))) return false;
try {
const pending: Array<{ value: unknown; depth: number }> = [{ value: JSON.parse(trimmed), depth: 0 }];
let visited = 0;
while (pending.length > 0 && visited < 1_000) {
const item = pending.pop()!;
visited += 1;
if (item.depth > 8 || item.value === null || typeof item.value !== "object") continue;
if (Array.isArray(item.value)) {
for (const child of item.value) pending.push({ value: child, depth: item.depth + 1 });
continue;
}
for (const [key, child] of Object.entries(item.value)) {
if (sensitiveNameRule(key)) return true;
pending.push({ value: child, depth: item.depth + 1 });
}
}
} catch {
return false;
}
return false;
}
function contentEvidence(value: string): SensitivityEvidence | undefined {
if (/-----BEGIN (?:[A-Z0-9]+ )?PRIVATE KEY-----/u.test(value)) {
return { kind: "content", ruleId: "credential.private_key" };
}
if (/(?:^|[^A-Z0-9])AKIA[A-Z0-9]{16}(?![A-Z0-9])/u.test(value)
|| /(?:^|[^A-Za-z0-9_])gh[pousr]_[A-Za-z0-9_]{30,}(?![A-Za-z0-9_])/u.test(value)
|| /(?:^|[^A-Za-z0-9_-])eyJ[A-Za-z0-9_-]{5,}\.[A-Za-z0-9_-]{5,}\.[A-Za-z0-9_-]{5,}(?![A-Za-z0-9_-])/u.test(value)) {
return { kind: "content", ruleId: "credential.access_key" };
}
if (/(?:^|[^\p{L}\p{N}_])(?:api[_ -]?key|access[_ -]?token|password|passwd|pwd|secret)\s*[:=]\s*[^\s,;]{4,}/iu.test(value)) {
return { kind: "content", ruleId: "credential.key_value" };
}
if (CLINICAL_TERM.test(value)) return { kind: "content", ruleId: "health.clinical_term" };
for (const match of value.matchAll(EMAIL)) {
if (validator.isEmail(match[0])) return { kind: "content", ruleId: "pii.email" };
}
for (const match of value.matchAll(ITALIAN_FISCAL_CODE)) {
if (validItalianFiscalCode(match[0])) {
return { kind: "content", ruleId: "pii.italian_fiscal_code" };
}
}
for (const match of value.matchAll(/\b(?:passaporto|passport)(?:\s+(?:numero|number|n\.?))?\s*[:#-]?\s*([A-Z0-9]{9})\b/giu)) {
if (validator.isPassportNumber(match[1]!, "IT")) {
return { kind: "content", ruleId: "pii.passport_number" };
}
}
for (const match of value.matchAll(/\bC[A-Z]\d{5}[A-Z]{2}\b/giu)) {
if (validator.isIdentityCard(match[0], "IT")) {
return { kind: "content", ruleId: "pii.identity_card" };
}
}
if (/\b(?:patente(?:\s+di\s+guida)?|driving\s+licen[cs]e)(?:\s+(?:numero|number|n\.?))?\s*[:#-]?\s*[A-Z0-9]{8,12}\b/iu.test(value)) {
return { kind: "content", ruleId: "pii.drivers_license_number" };
}
for (const match of value.matchAll(/(?<![A-Z0-9])[A-Z]{2}\d{2}(?:\s?[A-Z0-9]){11,30}(?![A-Z0-9])/giu)) {
if (validIban(match[0])) return { kind: "content", ruleId: "financial.iban" };
}
for (const match of value.matchAll(/(?<!\d)(?:\d[ -]?){13,19}(?!\d)/gu)) {
if (validPaymentCard(match[0])) {
return { kind: "content", ruleId: "financial.payment_card" };
}
}
for (const match of value.matchAll(/(?<![A-Z0-9])[A-Z]{6}[A-Z0-9]{2}(?:[A-Z0-9]{3})?(?![A-Z0-9])/giu)) {
const before = value.slice(Math.max(0, (match.index ?? 0) - 24), match.index ?? 0);
if (/\b(?:bic|swift)\s*[:=-]?\s*$/iu.test(before) && validator.isBIC(match[0])) {
return { kind: "content", ruleId: "financial.bic" };
}
}
for (const match of value.matchAll(/(?<!\d)(?:IT[ .-]?)?\d{11}(?!\d)/giu)) {
const candidate = match[0].replace(/[ .-]/gu, "");
if (validator.isVAT(candidate.replace(/^IT/iu, ""), "IT")) {
return { kind: "content", ruleId: "pii.italian_vat" };
}
}
for (const match of value.matchAll(/(?<![A-F0-9])(?:[A-F0-9]{2}[:-]){5}[A-F0-9]{2}(?![A-F0-9])/giu)) {
if (validator.isMACAddress(match[0])) {
return { kind: "content", ruleId: "network.mac_address" };
}
}
for (const match of value.matchAll(/(?<![A-F0-9:.])[A-F0-9:.]{3,45}(?![A-F0-9:.])/giu)) {
if (validator.isIP(match[0])) return { kind: "content", ruleId: "network.ip_address" };
}
for (const match of value.matchAll(/(?<![A-F0-9-])[0-9A-F]{8}-[0-9A-F]{4}-[1-8][0-9A-F]{3}-[89AB][0-9A-F]{3}-[0-9A-F]{12}(?![A-F0-9-])/giu)) {
if (validator.isUUID(match[0])) return { kind: "content", ruleId: "pii.uuid" };
}
for (const match of value.matchAll(/\b(?:https?|ftp):\/\/[^\s<>"']+/giu)) {
const candidate = match[0].replace(/[.,;:!?\])}]+$/u, "");
if (validator.isURL(candidate, { require_protocol: true })) {
return { kind: "content", ruleId: "network.url" };
}
}
if (findPhoneNumbersInText(value, "IT").some((match) => match.number.isValid())) {
return { kind: "content", ruleId: "pii.phone_number" };
}
if (jsonHasSensitiveKey(value)) {
return { kind: "content", ruleId: "pii.json_sensitive_key" };
}
return undefined;
}
interface ColumnState {
column: CatalogColumn;
evidence: SensitivityEvidence[];
nonSensitiveEvidence?: SensitivityEvidence;
observedValues: number;
nerCandidates: string[];
coverage: "metadata" | "complete" | "sampled" | "no_values";
sampledTarget: number;
}
/** Sole decision module for local column-level sensitivity assessments. */
export class SensitivityClassifier {
constructor(
private readonly values: SensitivityValueSource,
private readonly detector?: LocalNerDetector,
private readonly options: {
queryTimeoutMs?: number;
nerConfidenceThreshold?: number;
maxNerValuesPerColumn?: number;
maxNerCandidatesPerTable?: number;
now?: () => number;
} = {},
) {}
async assess(
targets: readonly SensitivityTableTarget[],
signal: AbortSignal,
sharedNerBudget?: SensitivityNerBudget,
onActivity?: (message: string) => void | Promise<void>,
): Promise<readonly SensitivityColumnAssessment[]> {
const now = this.options.now ?? Date.now;
const maxNerValuesPerColumn = boundedCount(this.options.maxNerValuesPerColumn, 8, 8);
const states = new Map<string, ColumnState>();
for (const target of targets) {
for (const column of target.columns) {
const nonSensitiveEvidence = nonSensitiveStructuralEvidence(column);
const metadataMatch = nonSensitiveEvidence ? undefined : metadataEvidence(column);
const binary = !nonSensitiveEvidence && UNSUPPORTED_BINARY_TYPE.test(column.dataType);
states.set(column.id, {
column,
evidence: metadataMatch
? [metadataMatch]
: binary
? [{ kind: "type", ruleId: "type.binary_uninspectable" }]
: [],
...(nonSensitiveEvidence ? { nonSensitiveEvidence } : {}),
observedValues: 0,
nerCandidates: [],
coverage: nonSensitiveEvidence || metadataMatch || binary ? "metadata" : "no_values",
sampledTarget: 0,
});
}
}
const completeTables = new Set<string>();
for (const [phaseIndex, phase] of SENSITIVITY_SAMPLE_PHASES.entries()) {
for (let offset = 0; offset < targets.length; offset += MAX_CONCURRENT_TABLE_SCANS) {
signal.throwIfAborted();
const batchNumber = Math.floor(offset / MAX_CONCURRENT_TABLE_SCANS) + 1;
const batchCount = Math.ceil(targets.length / MAX_CONCURRENT_TABLE_SCANS);
await onActivity?.(
`Scanning source data: pass ${phaseIndex + 1} of ${SENSITIVITY_SAMPLE_PHASES.length}, table batch ${batchNumber} of ${batchCount}.`,
);
signal.throwIfAborted();
const peerController = new AbortController();
const scanSignal = AbortSignal.any([signal, peerController.signal]);
try {
await Promise.all(targets.slice(offset, offset + MAX_CONCURRENT_TABLE_SCANS).map(async (target) => {
if (completeTables.has(target.table.id)) return;
const columns = target.columns.filter((column) => {
const state = states.get(column.id)!;
return state.evidence.length === 0 && !state.nonSensitiveEvidence
&& (!phase.deepTextOnly || DEEP_TEXT_TYPE.test(column.dataType));
});
if (columns.length === 0) return;
const coverage = await this.values.scanTable({
...target,
columns,
valuesPerColumn: phase.additionalValuesPerColumn,
sampleOffset: phase.targetValuesPerColumn - phase.additionalValuesPerColumn,
sampleSeed: phase.sampleSeed,
queryTimeoutMs: this.options.queryTimeoutMs ?? 5_000,
...(phaseIndex === 0 ? { fullScanThreshold: 1_000 } : {}),
}, (batch) => {
for (const item of batch) {
if (item.value === null) continue;
const state = states.get(item.columnId);
if (!state || state.evidence.length > 0) continue;
state.observedValues += 1;
if ((item.characterLength ?? item.value.length) > 500) {
state.evidence.push({ kind: "length", ruleId: "text.over_500_characters" });
continue;
}
const match = contentEvidence(item.value);
if (match) {
state.evidence.push(match);
continue;
}
if (state.nerCandidates.length < maxNerValuesPerColumn
&& !state.nerCandidates.includes(item.value)) {
state.nerCandidates.push(item.value);
}
}
}, scanSignal);
for (const column of columns) {
const state = states.get(column.id)!;
state.sampledTarget = Math.max(state.sampledTarget, phase.targetValuesPerColumn);
state.coverage = coverage.kind === "complete"
? "complete"
: state.observedValues === 0 ? "no_values" : "sampled";
}
if (coverage.kind === "complete") completeTables.add(target.table.id);
}));
} catch (error) {
peerController.abort(error);
throw error;
}
}
}
const nerBudget = sharedNerBudget ?? { remainingMs: 10_000 };
if (this.detector && (this.detector.isReady?.() ?? true) && !signal.aborted
&& nerBudget.remainingMs > 0) {
const maxCandidates = boundedCount(this.options.maxNerCandidatesPerTable, 2, 1_024);
const threshold = this.options.nerConfidenceThreshold ?? 0.8;
for (const [targetIndex, target] of targets.entries()) {
signal.throwIfAborted();
if (nerBudget.remainingMs <= 0) break;
const candidates: LocalNerCandidate[] = [];
candidateSelection: for (let valueIndex = 0; valueIndex < maxNerValuesPerColumn; valueIndex += 1) {
for (const column of target.columns) {
const state = states.get(column.id)!;
if (state.evidence.length > 0 || state.nonSensitiveEvidence) continue;
const text = state.nerCandidates[valueIndex];
if (text === undefined) continue;
candidates.push({ columnId: column.id, text });
if (candidates.length >= maxCandidates) break candidateSelection;
}
}
if (candidates.length === 0) continue;
await onActivity?.(
`Running local entity detection: table ${targetIndex + 1} of ${targets.length}.`,
);
signal.throwIfAborted();
const startedAt = now();
const deadline = startedAt + nerBudget.remainingMs;
try {
for (let offset = 0; offset < candidates.length; offset += MAX_NER_CANDIDATES_PER_REQUEST) {
if (signal.aborted || now() >= deadline) break;
try {
const detected = await this.detector.detect(
candidates.slice(offset, offset + MAX_NER_CANDIDATES_PER_REQUEST),
signal,
deadline,
);
for (const item of detected) {
const state = states.get(item.columnId);
if (!state || state.evidence.length > 0 || !Number.isFinite(item.confidence)
|| item.confidence < threshold || item.confidence > 1) continue;
const label = normalizedName(item.label).slice(0, 80);
if (!label) continue;
state.evidence.push({
kind: "ner",
ruleId: "ner.entity",
label,
confidence: item.confidence,
});
}
} catch {
// NER is optional: deterministic findings and scan coverage remain authoritative.
break;
}
}
} finally {
nerBudget.remainingMs = Math.max(0, nerBudget.remainingMs - Math.max(1, now() - startedAt));
}
}
}
return targets.flatMap((target) => target.columns.map((column) => {
const state = states.get(column.id)!;
const sensitive = state.evidence.length > 0;
const coverage = state.nonSensitiveEvidence
? "metadata"
: state.observedValues === 0 && !sensitive ? "no_values" : state.coverage;
const coverageEvidence: SensitivityEvidence[] = sensitive
? state.evidence
: state.nonSensitiveEvidence
? [state.nonSensitiveEvidence]
: [{
kind: "coverage",
ruleId: coverage === "complete"
? "coverage.complete"
: coverage === "no_values"
? "coverage.no_values"
: `coverage.sampled_${state.sampledTarget}`,
}];
return {
columnId: column.id,
assessment: sensitive ? "sensitive" : "non_sensitive",
proposedSensitive: sensitive,
evidence: coverageEvidence,
observedValues: state.observedValues,
coverage,
};
}));
}
/** Convenience for focused callers and rule-level tests. Production orchestration uses assess(). */
async assessTable(
target: SensitivityTableTarget,
signal: AbortSignal,
_retiredRunDeadline?: number,
nerBudget?: SensitivityNerBudget,
): Promise<readonly SensitivityColumnAssessment[]> {
return await this.assess([target], signal, nerBudget);
}
}
+100
View File
@@ -0,0 +1,100 @@
import { dirname } from "node:path";
import { fileURLToPath } from "node:url";
import { loadConfig } from "../config.js";
import { WorkspaceSecretStore } from "../workspaces/secret-store.js";
import { PythonLocalNerDetector } from "./local-ner-detector.js";
import { ConcreteCatalogPostgresAccess } from "./postgres-access.js";
import { createCatalogRepository } from "./repository.js";
import {
SENSITIVITY_POLICY_VERSION,
SensitivityAnalysisService,
} from "./sensitivity-analysis-service.js";
import { SensitivityClassifier } from "./sensitivity-classifier.js";
import { ConcreteSensitivityValueSource } from "./sensitivity-value-source.js";
const WORKSPACE_ID = /^[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?$/u;
async function main(): Promise<void> {
const workspaceId = process.argv[2];
if (!workspaceId || !WORKSPACE_ID.test(workspaceId)) {
process.stderr.write("Usage: sensitivity-shadow <workspace-id>\n");
process.exitCode = 2;
return;
}
let detector: PythonLocalNerDetector | undefined;
let stage = "configuration";
try {
const config = loadConfig(process.env);
stage = "catalog";
const repository = createCatalogRepository(config.catalogDatabase);
if (!(await repository.available())) throw new Error("catalog unavailable");
const database = await repository.getByWorkspace(workspaceId);
if (!database) throw new Error("database unavailable");
stage = "source";
const secretStore = new WorkspaceSecretStore({
root: config.workspaceSecretStoreRoot,
runtimeRoot: config.workspaceSecretRuntimeRoot,
installationId: config.workspaceRegistry.installationId,
});
const access = new ConcreteCatalogPostgresAccess(secretStore, {
connectTimeoutMs: config.workspaceDiagnosticTimeoutMs,
});
const source = new ConcreteSensitivityValueSource(access, secretStore);
if (config.sensitivityNer) {
const workerScript = config.sensitivityNer.workerScript
?? fileURLToPath(new URL("../../python/sensitivity_ner_worker.py", import.meta.url));
detector = new PythonLocalNerDetector({
pythonExecutable: config.sensitivityNer.pythonExecutable,
workerScript,
modelPath: config.sensitivityNer.modelPath,
cwd: dirname(workerScript),
threads: config.sensitivityNer.threads,
});
try {
await detector.warmup();
} catch {
await detector.close();
detector = undefined;
}
}
const startedAt = Date.now();
stage = "analysis";
const suggestions = await new SensitivityAnalysisService(
repository,
new SensitivityClassifier(source, detector),
).analyze(database.id, "all", [], new AbortController().signal);
const assessments = { sensitive: 0, nonSensitive: 0 };
const coverage = { metadata: 0, complete: 0, sampled: 0, noValues: 0 };
const rules = new Map<string, number>();
for (const suggestion of suggestions) {
if (suggestion.assessment === "sensitive") assessments.sensitive += 1;
else assessments.nonSensitive += 1;
if (suggestion.coverage === "no_values") coverage.noValues += 1;
else coverage[suggestion.coverage] += 1;
for (const evidence of suggestion.evidence) {
rules.set(evidence.ruleId, (rules.get(evidence.ruleId) ?? 0) + 1);
}
}
process.stdout.write(`${JSON.stringify({
ok: true,
policyVersion: SENSITIVITY_POLICY_VERSION,
nerEnabled: detector !== undefined,
total: suggestions.length,
assessments,
coverage,
rules: Object.fromEntries([...rules].sort(([left], [right]) => left.localeCompare(right))),
elapsedMs: Date.now() - startedAt,
})}\n`);
} catch {
process.stdout.write(`${JSON.stringify({
ok: false,
code: `sensitivity_shadow_${stage}_failed`,
})}\n`);
process.exitCode = 1;
} finally {
await detector?.close();
}
}
await main();
@@ -0,0 +1,291 @@
import { readFile } from "node:fs/promises";
import type { WorkspaceSecretStore } from "../workspaces/secret-store.js";
import { CATALOG_SECRET_IDS } from "./secrets.js";
import type { CatalogPostgresAccess } from "./postgres-access.js";
import type {
SensitivityScanCoverage,
SensitivityScanRequest,
SensitivityValueObservation,
SensitivityValueSource,
} from "./sensitivity-classifier.js";
import { CatalogConnectorError, type CatalogColumn } from "./types.js";
const MAX_VALUE_CHARACTERS = 501;
const MAX_COLUMNS_PER_QUERY = 25;
const SAMPLE_OVERSCAN_FACTOR = 10;
function quoteIdentifier(identifier: string): string {
return `"${identifier.replaceAll('"', '""')}"`;
}
function chunks<T>(items: readonly T[], size: number): T[][] {
const result: T[][] = [];
for (let offset = 0; offset < items.length; offset += size) {
result.push(items.slice(offset, offset + size));
}
return result;
}
function tableReference(request: SensitivityScanRequest): string {
return `${quoteIdentifier(request.database.schema)}.${quoteIdentifier(request.table.name)}`;
}
function samplePercentage(valuesPerColumn: number): number {
if (valuesPerColumn <= 300) return 30;
if (valuesPerColumn <= 700) return 70;
return 100;
}
function flatValueQuery(
request: SensitivityScanRequest,
columns: readonly CatalogColumn[],
options: { complete: boolean; randomized: boolean },
): string {
const projections = columns.map((column) => quoteIdentifier(column.name)).join(", ");
const perColumnLimit = options.complete
? request.fullScanThreshold ?? request.valuesPerColumn
: request.valuesPerColumn;
const rowLimit = Math.max(perColumnLimit, perColumnLimit * SAMPLE_OVERSCAN_FACTOR);
const sample = options.complete
? `SELECT ${projections} FROM ${tableReference(request)}`
: [
`SELECT ${projections} FROM ${tableReference(request)}`,
...(options.randomized
? [`TABLESAMPLE SYSTEM (${samplePercentage(request.valuesPerColumn)}) REPEATABLE (${request.sampleSeed})`]
: []),
`LIMIT ${rowLimit} OFFSET ${request.sampleOffset}`,
].join(" ");
const values = columns.map((column, index) => {
const identifier = quoteIdentifier(column.name);
return [
`(${index}, LEFT((sampled.${identifier})::text, ${MAX_VALUE_CHARACTERS}),`,
`CASE WHEN sampled.${identifier} IS NULL THEN NULL`,
`ELSE char_length((sampled.${identifier})::text) END)`,
].join(" ");
}).join(", ");
return [
`WITH sampled AS MATERIALIZED (${sample}),`,
"ranked AS (",
"SELECT value.__column_index, value.__value, value.__length,",
"row_number() OVER (PARTITION BY value.__column_index) AS __rank",
"FROM sampled",
`CROSS JOIN LATERAL (VALUES ${values}) AS value(__column_index, __value, __length)`,
"WHERE value.__value IS NOT NULL",
")",
"SELECT __column_index, __value, __length FROM ranked",
`WHERE __rank <= ${perColumnLimit}`,
].join(" ");
}
function observations(
columns: readonly CatalogColumn[],
rows: readonly Record<string, unknown>[],
): SensitivityValueObservation[] {
return rows.flatMap((row) => {
const index = Number(row.__column_index);
const column = Number.isSafeInteger(index) && index >= 0 ? columns[index] : undefined;
if (!column || row.__value === null || row.__value === undefined) return [];
const value = String(row.__value);
const parsedLength = row.__length === null || row.__length === undefined
? null
: Number(row.__length);
return [{
columnId: column.id,
value,
characterLength: parsedLength !== null && Number.isSafeInteger(parsedLength) && parsedLength >= 0
? parsedLength
: value.length,
}];
});
}
function cancelled(error: unknown): boolean {
return Boolean(error && typeof error === "object" && "code" in error && error.code === "57014");
}
/**
* Database-specific sampling adapter. Policy stays in SensitivityClassifier; this module only
* produces bounded, normalized non-null observations without persisting or logging values.
*/
export class ConcreteSensitivityValueSource implements SensitivityValueSource {
constructor(
private readonly access: CatalogPostgresAccess,
private readonly secretStore?: Pick<WorkspaceSecretStore, "materialize">,
) {}
async scanTable(
request: SensitivityScanRequest,
consume: (batch: readonly SensitivityValueObservation[]) => void | Promise<void>,
signal: AbortSignal,
): Promise<SensitivityScanCoverage> {
if (request.columns.length === 0) return { kind: "complete", observedValues: 0 };
if (request.database.binding.transport === "rest_api") {
return await this.scanRest(request, consume, signal);
}
return await this.scanPostgres(request, consume, signal);
}
private async scanPostgres(
request: SensitivityScanRequest,
consume: (batch: readonly SensitivityValueObservation[]) => void | Promise<void>,
signal: AbortSignal,
): Promise<SensitivityScanCoverage> {
const client = await this.access.connect(request.database, signal);
let transactionOpen = false;
let savepointSequence = 0;
let observedValues = 0;
try {
signal.throwIfAborted();
await client.query("BEGIN TRANSACTION READ ONLY", []);
transactionOpen = true;
await client.query("SELECT set_config('statement_timeout', $1, true)", [
`${Math.max(1, Math.floor(request.queryTimeoutMs))}ms`,
]);
const boundedQuery = async (sql: string): Promise<Array<Record<string, unknown>> | undefined> => {
signal.throwIfAborted();
savepointSequence += 1;
const savepoint = `sensitivity_scan_${savepointSequence}`;
await client.query(`SAVEPOINT ${savepoint}`, []);
try {
return (await client.query(sql, [])).rows;
} catch (error) {
if (!cancelled(error)) throw error;
await client.query(`ROLLBACK TO SAVEPOINT ${savepoint}`, []);
return undefined;
} finally {
await client.query(`RELEASE SAVEPOINT ${savepoint}`, []).catch(() => undefined);
}
};
let complete = false;
if (request.fullScanThreshold !== undefined) {
const probe = await boundedQuery(
`SELECT 1 AS __present FROM ${tableReference(request)} LIMIT ${request.fullScanThreshold + 1}`,
);
complete = probe !== undefined && probe.length <= request.fullScanThreshold;
}
for (const columnChunk of chunks(request.columns, MAX_COLUMNS_PER_QUERY)) {
signal.throwIfAborted();
let rows = await boundedQuery(flatValueQuery(request, columnChunk, {
complete,
randomized: !complete,
}));
if (rows === undefined && complete) {
complete = false;
rows = await boundedQuery(flatValueQuery(request, columnChunk, {
complete: false,
randomized: true,
}));
}
if (!complete && (rows === undefined || rows.length === 0)) {
rows = await boundedQuery(flatValueQuery(request, columnChunk, {
complete: false,
randomized: false,
}));
}
if (rows === undefined) throw new CatalogConnectorError("Sensitivity sample query timed out");
const batch = observations(columnChunk, rows);
observedValues += batch.length;
if (batch.length > 0) await consume(batch);
}
return { kind: complete ? "complete" : "sampled", observedValues };
} catch (error) {
if (error instanceof CatalogConnectorError) throw error;
throw new CatalogConnectorError("Sensitivity source scan failed");
} finally {
if (transactionOpen) await client.query("ROLLBACK", []).catch(() => undefined);
await client.end().catch(() => undefined);
}
}
private async scanRest(
request: SensitivityScanRequest,
consume: (batch: readonly SensitivityValueObservation[]) => void | Promise<void>,
signal: AbortSignal,
): Promise<SensitivityScanCoverage> {
if (!this.secretStore) throw new CatalogConnectorError("REST sensitivity scanning is not configured");
const auth = request.database.binding.restAuth ?? "bearer";
const materialized = this.secretStore.materialize(
request.database.workspaceId,
auth === "none" ? [] : [CATALOG_SECRET_IDS.apiKey],
);
let observedValues = 0;
try {
const headers: Record<string, string> = { "content-type": "application/json" };
if (auth !== "none") {
const credentialFile = materialized.files.get(CATALOG_SECRET_IDS.apiKey);
if (!credentialFile) throw new CatalogConnectorError("REST API key is not configured");
const credential = (await readFile(credentialFile, "utf8")).trim();
if (auth === "bearer") headers.authorization = `Bearer ${credential}`;
else headers["x-api-key"] = credential;
}
const baseUrl = request.database.binding.baseUrl?.replace(/\/+$/u, "");
if (!baseUrl) throw new CatalogConnectorError("Database binding is incomplete");
const runQuery = async (sql: string): Promise<Array<Record<string, unknown>> | undefined> => {
const timeout = AbortSignal.timeout(Math.max(1, Math.floor(request.queryTimeoutMs)));
try {
const response = await fetch(`${baseUrl}/rpc/run_query`, {
method: "POST",
headers,
body: JSON.stringify({ query_text: sql }),
signal: AbortSignal.any([signal, timeout]),
});
if (!response.ok) throw new CatalogConnectorError("REST sensitivity source scan failed");
const body: unknown = await response.json();
if (!Array.isArray(body)
|| body.some((row) => !row || typeof row !== "object" || Array.isArray(row))) {
throw new CatalogConnectorError("REST sensitivity source response is invalid");
}
return body as Array<Record<string, unknown>>;
} catch (error) {
if (signal.aborted) throw error;
if (timeout.aborted) return undefined;
throw error;
}
};
let complete = false;
if (request.fullScanThreshold !== undefined) {
const probe = await runQuery(
`SELECT 1 AS __present FROM ${tableReference(request)} LIMIT ${request.fullScanThreshold + 1}`,
);
complete = probe !== undefined && probe.length <= request.fullScanThreshold;
}
let requestCount = request.fullScanThreshold === undefined ? 0 : 1;
for (const columnChunk of chunks(request.columns, MAX_COLUMNS_PER_QUERY)) {
signal.throwIfAborted();
let rows = await runQuery(flatValueQuery(request, columnChunk, {
complete,
randomized: !complete,
}));
requestCount += 1;
if (rows === undefined && complete) {
complete = false;
rows = await runQuery(flatValueQuery(request, columnChunk, {
complete: false,
randomized: true,
}));
requestCount += 1;
}
if (!complete && (rows === undefined || rows.length === 0)) {
rows = await runQuery(flatValueQuery(request, columnChunk, {
complete: false,
randomized: false,
}));
requestCount += 1;
}
if (rows === undefined) throw new CatalogConnectorError("REST sensitivity sample query timed out");
const batch = observations(columnChunk, rows);
observedValues += batch.length;
if (batch.length > 0) await consume(batch);
}
// Multiple HTTP requests cannot share a source snapshot, so only one-request reads are complete.
return { kind: complete && requestCount === 1 ? "complete" : "sampled", observedValues };
} catch (error) {
if (error instanceof CatalogConnectorError) throw error;
throw new CatalogConnectorError("REST sensitivity source scan failed");
} finally {
materialized.release();
}
}
}
+4 -9
View File
@@ -39,7 +39,10 @@ function safeFailure(error: unknown): { code: string; message: string } {
};
}
if (error instanceof CatalogConnectorError) {
return { code: "schema_introspection_failed", message: "The database schema could not be read safely." };
return {
code: "schema_introspection_failed",
message: "The database schema could not be read. Check the connection and credentials, then try again.",
};
}
return { code: "schema_sync_failed", message: "Schema synchronization failed." };
}
@@ -64,7 +67,6 @@ export class CatalogSyncWorker {
}
async start(database: WorkspaceDatabase, scope: CatalogSyncScope, tableIds: readonly string[]): Promise<CatalogSyncRun> {
this.assertReady(database);
const uniqueTableIds = [...new Set(tableIds)];
if (scope === "columns") {
const tables = await Promise.all(uniqueTableIds.map((tableId) => this.repository.getTable(database.id, tableId)));
@@ -177,7 +179,6 @@ export class CatalogSyncWorker {
if (!database || database.version !== claimed.requestedDatabaseVersion) {
throw new CatalogConflictError("Database binding changed before synchronization started");
}
this.assertReady(database);
const progress: CatalogSchemaScanProgress = async (phase, counts) => {
await this.checkCancelled(runId);
await this.repository.updateSyncRun(runId, {
@@ -277,12 +278,6 @@ export class CatalogSyncWorker {
}
}
private assertReady(database: WorkspaceDatabase): void {
if (database.connectionStatus !== "reachable" || database.testedVersion !== database.version) {
throw new CatalogConflictError("Test the current database binding before synchronizing its schema");
}
}
private assertCapability(scope: CatalogSyncScope, snapshot: ObservedSchemaSnapshot): void {
const required = scope === "all" ? ["tables", "columns", "relationships"] as const : [scope] as const;
for (const name of required) {
+32 -26
View File
@@ -102,6 +102,7 @@ export interface CatalogColumn {
description: string | null;
generatedDescription: string | null;
sensitive: boolean;
sensitivityReason: string | null;
lastSyncedDatabaseVersion: number | null;
lastSyncedAt: string | null;
version: number;
@@ -195,7 +196,7 @@ export interface CatalogLogicalRelationshipCandidate {
export type CatalogDatabaseMetadataDeleteTarget = "tables" | "relationships";
export type CatalogTableMetadataDeleteTarget = "columns" | "relationships";
export type CatalogDescriptionTarget = "tables" | "columns";
export type CatalogDescriptionTarget = "tables" | "columns" | "database_columns";
export interface CatalogMetadataDeleteCounts {
tables: number;
@@ -266,18 +267,21 @@ export interface DescriptionGenerationEvent {
createdAt: string;
}
export type SensitiveDataSuggestionScope = "all" | "selected_tables" | "selected_columns";
export type SensitiveDataSuggestionStatus = "running" | "completed" | "failed" | "interrupted";
export type SensitivityAnalysisScope = "all" | "selected_tables" | "selected_columns";
export type SensitivityAnalysisStatus = "running" | "completed" | "failed" | "interrupted";
export interface SensitiveDataSuggestionRun {
export interface SensitivityAnalysisRun {
id: string;
databaseId: string;
scope: SensitiveDataSuggestionScope;
modelId: string;
status: SensitiveDataSuggestionStatus;
scope: SensitivityAnalysisScope;
engine: "llm" | "local";
modelId: string | null;
policyVersion: string | null;
status: SensitivityAnalysisStatus;
total: number;
suggestedSensitive: number;
suggestedNonSensitive: number;
unknown: number;
inputTokens: number;
cacheReadTokens: number;
outputTokens: number;
@@ -288,11 +292,12 @@ export interface SensitiveDataSuggestionRun {
errorSummary: string | null;
}
export interface SensitiveDataSuggestionRunUpdate {
status?: SensitiveDataSuggestionStatus;
export interface SensitivityAnalysisRunUpdate {
status?: SensitivityAnalysisStatus;
total?: number;
suggestedSensitive?: number;
suggestedNonSensitive?: number;
unknown?: number;
finishedAt?: string | null;
errorSummary?: string | null;
inputTokens?: number;
@@ -300,7 +305,7 @@ export interface SensitiveDataSuggestionRunUpdate {
outputTokens?: number;
}
export interface SensitiveDataSuggestionEvent {
export interface SensitivityAnalysisEvent {
runId: string;
sequence: number;
level: "info" | "warning" | "error";
@@ -459,6 +464,7 @@ export interface CatalogRepository {
description: string | null,
generatedDescription: string | null,
sensitive?: boolean,
sensitivityReason?: string | null,
): Promise<CatalogColumn | undefined>;
consolidateGeneratedDescriptions(
databaseId: string,
@@ -491,29 +497,29 @@ export interface CatalogRepository {
runId: string,
afterSequence?: number,
): Promise<DescriptionGenerationEvent[]>;
createSensitiveDataSuggestionRun(
createSensitivityAnalysisRun(
databaseId: string,
scope: SensitiveDataSuggestionScope,
modelId: string,
): Promise<SensitiveDataSuggestionRun>;
getSensitiveDataSuggestionRun(runId: string): Promise<SensitiveDataSuggestionRun | undefined>;
listSensitiveDataSuggestionRuns(limit?: number): Promise<SensitiveDataSuggestionRun[]>;
interruptActiveSensitiveDataSuggestionRuns(
scope: SensitivityAnalysisScope,
origin: { engine: "llm"; modelId: string } | { engine: "local"; policyVersion: string },
): Promise<SensitivityAnalysisRun>;
getSensitivityAnalysisRun(runId: string): Promise<SensitivityAnalysisRun | undefined>;
listSensitivityAnalysisRuns(limit?: number): Promise<SensitivityAnalysisRun[]>;
interruptActiveSensitivityAnalysisRuns(
errorSummary: string,
): Promise<SensitiveDataSuggestionRun[]>;
updateSensitiveDataSuggestionRun(
): Promise<SensitivityAnalysisRun[]>;
updateSensitivityAnalysisRun(
runId: string,
update: SensitiveDataSuggestionRunUpdate,
): Promise<SensitiveDataSuggestionRun | undefined>;
appendSensitiveDataSuggestionEvent(
update: SensitivityAnalysisRunUpdate,
): Promise<SensitivityAnalysisRun | undefined>;
appendSensitivityAnalysisEvent(
runId: string,
level: SensitiveDataSuggestionEvent["level"],
level: SensitivityAnalysisEvent["level"],
message: string,
): Promise<SensitiveDataSuggestionEvent>;
listSensitiveDataSuggestionEvents(
): Promise<SensitivityAnalysisEvent>;
listSensitivityAnalysisEvents(
runId: string,
afterSequence?: number,
): Promise<SensitiveDataSuggestionEvent[]>;
): Promise<SensitivityAnalysisEvent[]>;
listRelationships(databaseId: string): Promise<CatalogPhysicalRelationship[]>;
listLogicalRelationships(databaseId: string): Promise<CatalogLogicalRelationship[]>;
getLogicalRelationshipContext(databaseId: string): Promise<CatalogLogicalRelationshipContext | undefined>;
+68 -2
View File
@@ -32,6 +32,13 @@ export interface AppConfig {
piManagementTimeoutMs: number;
secretsFile?: string;
installationConfigFile?: string;
modelCatalogFile?: string;
sensitivityNer?: {
pythonExecutable: string;
modelPath: string;
workerScript?: string;
threads: number;
};
piAuthFile?: string;
secretFiles: Readonly<Record<string, string | undefined>>;
modelApiKeyFile?: string;
@@ -50,6 +57,7 @@ export interface AppConfig {
workspaceSecretRuntimeRoot: string;
internalQdrantUrl: string;
internalEmbeddingUrl: string;
internalEmbeddingId: string;
internalEmbeddingModel: string;
internalEmbeddingDimensions: number;
}
@@ -189,6 +197,21 @@ function positiveDimension(value: string | undefined, fallback: number): number
return parsed;
}
function internalEmbeddingIdentity(
identityValue: string | undefined,
modelValue: string | undefined,
): { id: string; model: string } {
const id = identityValue ?? "ollama/qwen3-embedding:0.6b";
if (!/^ollama\/[A-Za-z0-9][A-Za-z0-9._:-]{0,255}$/.test(id)) {
throw new Error("internal embedding identity configuration is invalid");
}
const model = id.slice(id.indexOf("/") + 1);
if (modelValue !== undefined && modelValue !== model) {
throw new Error("internal embedding model does not match its canonical identity");
}
return { id, model };
}
function catalogDatabase(env: Record<string, string | undefined>): CatalogConnectionConfig | undefined {
const value = env.THT_CATALOG_DATABASE_URL;
if (value !== undefined) {
@@ -330,6 +353,42 @@ export function loadConfig(
|| installationConfigFile.includes("\0")
|| !path.isAbsolute(installationConfigFile)
)) throw new Error("installation configuration is invalid");
const modelCatalogFile = env.THT_MODEL_CATALOG_FILE;
if (modelCatalogFile !== undefined && (
modelCatalogFile.trim() !== modelCatalogFile
|| modelCatalogFile.length === 0
|| modelCatalogFile.includes("\0")
|| !path.isAbsolute(modelCatalogFile)
)) throw new Error("runtime model catalog configuration is invalid");
const sensitivityNerModelPath = env.THT_SENSITIVITY_NER_MODEL_PATH;
const sensitivityNerPython = env.THT_SENSITIVITY_NER_PYTHON;
const sensitivityNerWorker = env.THT_SENSITIVITY_NER_WORKER;
for (const [value, label] of [
[sensitivityNerModelPath, "model path"],
[sensitivityNerPython, "Python executable"],
[sensitivityNerWorker, "worker path"],
] as const) {
if (value !== undefined && (
value.length === 0 || value.trim() !== value || value.includes("\0") || !path.isAbsolute(value)
)) throw new Error(`sensitivity NER ${label} configuration is invalid`);
}
if (sensitivityNerModelPath === undefined && (
sensitivityNerPython !== undefined
|| sensitivityNerWorker !== undefined
|| env.THT_SENSITIVITY_NER_THREADS !== undefined
)) throw new Error("sensitivity NER settings require a model path");
const sensitivityNerThreads = Number(env.THT_SENSITIVITY_NER_THREADS ?? 2);
if (!Number.isSafeInteger(sensitivityNerThreads) || sensitivityNerThreads < 1 || sensitivityNerThreads > 8) {
throw new Error("sensitivity NER thread configuration is invalid");
}
const sensitivityNer = sensitivityNerModelPath === undefined
? undefined
: {
modelPath: sensitivityNerModelPath,
pythonExecutable: sensitivityNerPython ?? "/opt/sensitivity-ner/bin/python",
...(sensitivityNerWorker ? { workerScript: sensitivityNerWorker } : {}),
threads: sensitivityNerThreads,
};
const piAuthFile = env.THT_PI_AUTH_FILE;
if (piAuthFile !== undefined && (
piAuthFile.trim() !== piAuthFile || piAuthFile.length === 0 || piAuthFile.includes("\0")
@@ -391,6 +450,10 @@ export function loadConfig(
"internal embedding URL",
["embedding", "localhost"],
);
const internalEmbedding = internalEmbeddingIdentity(
env.THT_INTERNAL_EMBEDDING_ID,
env.THT_INTERNAL_EMBEDDING_MODEL,
);
return {
host: env.HOST ?? "127.0.0.1",
port: Number(env.PORT ?? 8787),
@@ -404,7 +467,7 @@ export function loadConfig(
publicExposure,
sessionStorage,
catalogDatabase: catalogDatabase(env),
defaults: { provider: env.PI_PROVIDER, model: env.PI_MODEL, thinking: env.PI_THINKING },
defaults: { thinking: env.PI_THINKING },
maxPiProcesses: Number(env.MAX_PI_PROCESSES ?? 4),
settingsFile,
maintenanceFile: env.THT_MAINTENANCE_FILE ?? path.join(path.dirname(settingsFile), "maintenance.json"),
@@ -413,6 +476,8 @@ export function loadConfig(
piManagementTimeoutMs: piManagementTimeout(env.PI_MANAGEMENT_TIMEOUT_MS),
secretsFile,
installationConfigFile,
modelCatalogFile,
sensitivityNer,
piAuthFile,
secretFiles,
modelApiKeyFile,
@@ -425,7 +490,8 @@ export function loadConfig(
workspaceSecretRuntimeRoot,
internalQdrantUrl,
internalEmbeddingUrl,
internalEmbeddingModel: env.THT_INTERNAL_EMBEDDING_MODEL ?? "qwen3-embedding:0.6b",
internalEmbeddingId: internalEmbedding.id,
internalEmbeddingModel: internalEmbedding.model,
internalEmbeddingDimensions: positiveDimension(env.THT_INTERNAL_EMBEDDING_DIMENSIONS, 1024),
};
}
+1 -1
View File
@@ -8,7 +8,7 @@ import {
isUsableAuthenticationSecret,
} from "../auth/secret-policy.js";
/** Credential names that metadata-generation model entries may reference. */
/** Credential names that Installation Model Catalog providers may reference. */
export const METADATA_GENERATION_SECRET_KEYS = Object.freeze([
"THT_METADATA_API_KEY", "ANTHROPIC_API_KEY", "AZURE_API_KEY", "GEMINI_API_KEY",
"DEEPSEEK_API_KEY", "OPENAI_API_KEY", "OPENROUTER_API_KEY", "ZAI_API_KEY",
+161
View File
@@ -0,0 +1,161 @@
import {
closeSync, constants, fstatSync, lstatSync, openSync, readFileSync, type Stats,
} from "node:fs";
import { z } from "zod";
const MAX_CATALOG_BYTES = 1024 * 1024;
const RUNTIME_CATALOG_FILE = "/run/thothii-model-catalog/catalog.json";
const canonicalId = z.string().regex(/^[a-z][a-z0-9._-]{0,63}\/[A-Za-z0-9][A-Za-z0-9._:-]{0,255}$/);
const secretBundleKey = /^[A-Z][A-Z0-9_]{0,63}$/;
const endpointSchema = z.object({
baseUrl: z.string().url(),
apiVersion: z.string().optional(),
}).strict();
const authenticationSchema = z.object({
mode: z.enum(["secret_env", "pi_auth", "none"]),
apiKeyEnv: z.string().optional(),
}).strict();
const runtimeModelSchema = z.object({
id: canonicalId,
provider: z.string().min(1),
model: z.string().min(1),
label: z.string().min(1),
upstreamModel: z.string().min(1),
endpoint: endpointSchema.optional(),
authentication: authenticationSchema,
sessionAdapter: z.object({ mode: z.enum(["pi_builtin", "openai_compatible"]) }).strict().optional(),
metadataAdapter: z.object({ litellmProvider: z.string().min(1) }).strict().optional(),
session: z.object({
reasoning: z.boolean(),
input: z.array(z.string()).optional(),
cost: z.object({
input: z.number(), output: z.number(), cacheRead: z.number(), cacheWrite: z.number(),
}).strict().optional(),
contextWindow: z.number().int().positive().optional(),
maxTokens: z.number().int().positive().optional(),
compatibility: z.object({
supportsDeveloperRole: z.boolean(),
supportsReasoningEffort: z.boolean(),
supportsStore: z.boolean(),
maxTokensField: z.string().optional(),
}).strict().optional(),
}).strict().optional(),
metadataGeneration: z.object({ disableThinking: z.boolean() }).strict().optional(),
}).strict();
const catalogSchema = z.object({
schemaVersion: z.literal(1),
defaultSession: canonicalId,
defaultMetadataGeneration: canonicalId.optional(),
embedding: z.object({ id: canonicalId, dimensions: z.number().int().positive() }).strict(),
models: z.array(runtimeModelSchema).max(64),
}).strict();
export type RuntimeModel = z.infer<typeof runtimeModelSchema>;
export interface RuntimeModelCatalog {
readonly defaultSession: string | null;
readonly defaultMetadataGeneration: string | null;
readonly embedding: Readonly<{ id: string; dimensions: number }> | null;
sessionModels(): readonly RuntimeModel[];
metadataModels(): readonly RuntimeModel[];
hasSession(id: string): boolean;
}
class RestartLoadedRuntimeModelCatalog implements RuntimeModelCatalog {
readonly defaultSession: string | null;
readonly defaultMetadataGeneration: string | null;
readonly embedding: Readonly<{ id: string; dimensions: number }> | null;
readonly #sessions: readonly RuntimeModel[];
readonly #metadata: readonly RuntimeModel[];
readonly #sessionIds: ReadonlySet<string>;
constructor(catalog?: z.infer<typeof catalogSchema>) {
this.defaultSession = catalog?.defaultSession ?? null;
this.defaultMetadataGeneration = catalog?.defaultMetadataGeneration ?? null;
this.embedding = catalog ? Object.freeze({ ...catalog.embedding }) : null;
this.#sessions = Object.freeze((catalog?.models ?? []).filter((model) => model.session !== undefined));
this.#metadata = Object.freeze((catalog?.models ?? []).filter((model) => model.metadataGeneration !== undefined));
this.#sessionIds = new Set(this.#sessions.map((model) => model.id));
}
sessionModels(): readonly RuntimeModel[] { return this.#sessions.map((model) => ({ ...model })); }
metadataModels(): readonly RuntimeModel[] { return this.#metadata.map((model) => ({ ...model })); }
hasSession(id: string): boolean { return this.#sessionIds.has(id); }
}
function protectedCatalogStat(file: string, info: Stats): boolean {
const mode = info.mode & 0o777;
if (!info.isFile() || info.isSymbolicLink() || info.nlink !== 1
|| info.size < 1 || info.size > MAX_CATALOG_BYTES) return false;
if (file === RUNTIME_CATALOG_FILE && info.uid === 0 && (mode === 0o444 || mode === 0o644)) return true;
return info.uid === (process.getuid?.() ?? info.uid) && (mode === 0o400 || mode === 0o600 || mode === 0o644);
}
function readProtectedCatalog(file: string): unknown {
let descriptor: number | undefined;
try {
const before = lstatSync(file);
if (!protectedCatalogStat(file, before)) throw new Error("runtime model catalog is unavailable");
descriptor = openSync(file, constants.O_RDONLY | constants.O_NOFOLLOW);
const opened = fstatSync(descriptor);
if (!protectedCatalogStat(file, opened)
|| before.dev !== opened.dev || before.ino !== opened.ino) throw new Error("runtime model catalog is unavailable");
const source = readFileSync(descriptor, "utf8");
const after = fstatSync(descriptor);
const current = lstatSync(file);
if (!protectedCatalogStat(file, after) || !protectedCatalogStat(file, current)
|| opened.dev !== after.dev || opened.ino !== after.ino
|| opened.dev !== current.dev || opened.ino !== current.ino) throw new Error("runtime model catalog is unavailable");
return JSON.parse(source);
} catch {
throw new Error("runtime model catalog is unavailable");
} finally {
if (descriptor !== undefined) try { closeSync(descriptor); } catch { /* sanitized above */ }
}
}
export function loadRuntimeModelCatalog(file?: string): RuntimeModelCatalog {
if (!file) return new RestartLoadedRuntimeModelCatalog();
const parsed = catalogSchema.safeParse(readProtectedCatalog(file));
if (!parsed.success) throw new Error("runtime model catalog is invalid");
if (parsed.data.models.some((model) => !validRuntimeModel(model))) {
throw new Error("runtime model catalog is invalid");
}
const ids = new Set(parsed.data.models.map((model) => model.id));
if (ids.size !== parsed.data.models.length) throw new Error("runtime model catalog contains duplicate models");
const sessions = parsed.data.models.filter((model) => model.session !== undefined).map((model) => model.id);
const metadata = parsed.data.models.filter((model) => model.metadataGeneration !== undefined).map((model) => model.id);
if (!sessions.includes(parsed.data.defaultSession)) throw new Error("runtime model catalog session default is invalid");
if ((metadata.length > 0) !== (parsed.data.defaultMetadataGeneration !== undefined)
|| (parsed.data.defaultMetadataGeneration !== undefined
&& !metadata.includes(parsed.data.defaultMetadataGeneration))) {
throw new Error("runtime model catalog metadata default is invalid");
}
return new RestartLoadedRuntimeModelCatalog(parsed.data);
}
function validRuntimeModel(model: RuntimeModel): boolean {
if ((model.session !== undefined) !== (model.sessionAdapter !== undefined)) return false;
if ((model.metadataGeneration !== undefined) !== (model.metadataAdapter !== undefined)) return false;
switch (model.authentication.mode) {
case "secret_env":
return model.authentication.apiKeyEnv !== undefined
&& secretBundleKey.test(model.authentication.apiKeyEnv);
case "pi_auth":
return model.authentication.apiKeyEnv === undefined
&& model.metadataGeneration === undefined
&& model.sessionAdapter?.mode === "pi_builtin";
case "none":
return model.authentication.apiKeyEnv === undefined && model.endpoint !== undefined;
}
}
export function splitCanonicalModelId(id: string): { provider: string; model: string } {
const slash = id.indexOf("/");
if (slash <= 0 || slash === id.length - 1) throw new Error("model identity is invalid");
return { provider: id.slice(0, slash), model: id.slice(slash + 1) };
}
+8 -6
View File
@@ -8,8 +8,8 @@ import { join } from "node:path";
import { loadConfig, type AppConfig } from "./config.js";
import { rolesToPermissions } from "./auth/config.js";
import type { PrincipalContext } from "./auth/principal.js";
import { createPiModelLister } from "./pi/list-models.js";
import { createPiManagement } from "./pi/management.js";
import { loadRuntimeModelCatalog } from "./models/runtime-model-catalog.js";
import { effectiveSettings } from "./routes/settings.js";
import { MaintenanceBarrier } from "./runtime/maintenance-gate.js";
import { loadSettings } from "./settings/settings-store.js";
@@ -19,7 +19,7 @@ import { WorkspaceSecretStore } from "./workspaces/secret-store.js";
type OperatorAction = "maintenance-activate" | "maintenance-deactivate" | "maintenance-status"
| "session-inventory" | "workflow-doctor" | "workspace-integrity"
| "pi-options" | "pi-test" | "effective-settings";
| "pi-test" | "effective-settings";
const lifecyclePrincipal: PrincipalContext = {
issuer: "tht-operator-command",
@@ -110,9 +110,11 @@ export async function runOperatorAction(
if (action === "session-inventory") return await sessionInventory(config);
if (action === "workflow-doctor") return await workflowDiagnostics(config);
if (action === "workspace-integrity") return await workspaceIntegrity(config);
if (action === "effective-settings") return effectiveSettings(config, loadSettings(config));
const service = createPiManagement(config, { listModels: createPiModelLister(config) });
if (action === "pi-options") return await service.options();
const modelCatalog = loadRuntimeModelCatalog(config.modelCatalogFile);
if (action === "effective-settings") {
return effectiveSettings(config, loadSettings(config), modelCatalog);
}
const service = createPiManagement(config, { modelCatalog });
if (action === "pi-test") return await service.test();
throw new Error("unsupported operator action");
}
@@ -121,7 +123,7 @@ async function main(): Promise<void> {
const action = process.argv[2] as OperatorAction | undefined;
if (!action || ![
"maintenance-activate", "maintenance-deactivate", "maintenance-status", "session-inventory",
"workflow-doctor", "workspace-integrity", "pi-options", "pi-test", "effective-settings",
"workflow-doctor", "workspace-integrity", "pi-test", "effective-settings",
].includes(action)) throw new Error("invalid operator action");
const result = await runOperatorAction(action, loadConfig(process.env));
process.stdout.write(`${JSON.stringify(result)}\n`);
+20 -2
View File
@@ -10,6 +10,7 @@ import {
readConfiguredPiAgentFile,
validateDeclarativePiConfig,
} from "./managed-config.js";
import type { RuntimeModelCatalog } from "../models/runtime-model-catalog.js";
export interface PiModel {
provider: string;
@@ -18,6 +19,8 @@ export interface PiModel {
reasoning: boolean;
}
export type ListModelsFn = () => Promise<PiModel[]>;
interface Opts {
spawnFn?: (
command: string,
@@ -28,6 +31,7 @@ interface Opts {
nowMs?: () => number;
loadEnabledModels?: () => PiEnabledModelsResult;
readModelsStore?: () => string | undefined;
modelCatalog?: RuntimeModelCatalog;
warn?: (message: string) => void;
}
@@ -36,7 +40,7 @@ interface Opts {
* configured) via an ephemeral `pi --mode rpc` process. Result is cached for
* `ttlMs`. The returned function rejects on timeout/error; callers degrade.
*/
export function createPiModelLister(cfg: AppConfig, opts: Opts = {}): () => Promise<PiModel[]> {
export function createPiModelLister(cfg: AppConfig, opts: Opts = {}): ListModelsFn {
const ttlMs = opts.ttlMs ?? 60_000;
const now = opts.nowMs ?? (() => Date.now());
const spawnFn = opts.spawnFn ?? nodeSpawn;
@@ -80,9 +84,23 @@ export function createPiModelLister(cfg: AppConfig, opts: Opts = {}): () => Prom
const byCompositeId = new Map(
available.map((model) => [`${model.provider}/${model.id}`, model]),
);
const catalogByPiId = new Map(
(opts.modelCatalog?.sessionModels() ?? []).map((model) => [
`${model.provider}/${model.upstreamModel}`,
model,
]),
);
const models = enabled.ids.flatMap((id) => {
const model = byCompositeId.get(id);
return model ? [model] : [];
if (!model) return [];
const catalogModel = catalogByPiId.get(id);
return [catalogModel ? {
...model,
provider: catalogModel.provider,
id: catalogModel.model,
name: catalogModel.label,
reasoning: catalogModel.session?.reasoning ?? model.reasoning,
} : model];
});
if (models.length === 0) opts.warn?.("No Pi-enabled models are currently available");
cache = { at: now(), models };
+27 -92
View File
@@ -2,12 +2,11 @@ import { execFile as nodeExecFile } from "node:child_process";
import { promisify } from "node:util";
import type { AppConfig } from "../config.js";
import { secretValue } from "../config/secret-bundle.js";
import { loadSettings, type Settings } from "../settings/settings-store.js";
import {
loadSettings,
saveSettings,
type Settings,
} from "../settings/settings-store.js";
import type { PiModel } from "./list-models.js";
splitCanonicalModelId,
type RuntimeModelCatalog,
} from "../models/runtime-model-catalog.js";
import {
configuredPiProviderApiKey,
PI_MANAGED_CONFIG_ERROR_MESSAGE,
@@ -45,13 +44,6 @@ export interface PiStatus {
message?: string;
}
export interface PiOptions {
providers: string[];
models: Array<{ provider: string; id: string }>;
reasoning: PiReasoning[];
checkedAt: string;
}
export interface PiTestResult {
ready: boolean;
checkedAt: string;
@@ -76,15 +68,13 @@ export type PiExecFile = (
export interface PiManagementService {
status(): Promise<PiStatus>;
options(): Promise<PiOptions>;
configure(value: PiInstallationConfig): Promise<PiInstallationConfig & { updatedAt: string }>;
test(): Promise<PiTestResult>;
logs(): Promise<PiLogs>;
}
export class PiManagementError extends Error {
constructor(
public readonly code: "pi_management_invalid_config" | "pi_management_unavailable" | "pi_management_write_failed",
public readonly code: "pi_management_unavailable",
message: string,
) {
super(message);
@@ -93,10 +83,9 @@ export class PiManagementError extends Error {
interface PiManagementDeps {
execute?: PiExecFile;
listModels: () => Promise<PiModel[]>;
modelCatalog: RuntimeModelCatalog;
smokeProvider?: PiProviderSmoke;
readSettings?: () => Settings;
saveSettings?: (settings: Settings) => Settings;
readLogs?: () => string | Promise<string>;
credentialStatus?: (provider: string | undefined) => PiCredentialStatus;
now?: () => Date;
@@ -111,54 +100,36 @@ export function createPiManagement(config: AppConfig, deps: PiManagementDeps): P
};
const execute = deps.execute ?? defaultExecFile;
const readSettings = deps.readSettings ?? (() => loadSettings(config));
const persistSettings = deps.saveSettings ?? ((settings) => saveSettings(config, settings));
const readLogs = deps.readLogs ?? (() => diagnostics.join("\n"));
const smokeProvider = deps.smokeProvider ?? createPiProviderSmoke(config);
const smokeProvider = deps.smokeProvider ?? createPiProviderSmoke(config, {
modelCatalog: deps.modelCatalog,
});
const credentialStatus = deps.credentialStatus ?? ((provider: string | undefined) => {
try {
const model = deps.modelCatalog.defaultSession
? deps.modelCatalog.sessionModels().find((entry) => entry.id === deps.modelCatalog.defaultSession)
: undefined;
const credentialName = model?.authentication.mode === "secret_env"
? model.authentication.apiKeyEnv
: undefined;
const configuredApiKey = configuredPiProviderApiKey(
readConfiguredPiAgentFile("models.json", true),
provider,
) ?? (credentialName ? `$${credentialName}` : undefined);
return piProviderCredentialStatus({
provider,
authProviders: loadPiAuthProviders(),
resolveCredentialValue: () => secretValue(config, "THT_MODEL_API_KEY"),
resolveCredentialValue: () => credentialName
? secretValue(config, credentialName)
: config.modelCatalogFile ? undefined : secretValue(config, "THT_MODEL_API_KEY"),
credentialFile: config.modelApiKeyFile,
configuredApiKey: configuredPiProviderApiKey(
readConfiguredPiAgentFile("models.json", true),
provider,
),
configuredApiKey,
});
} catch {
return "missing";
}
});
const closedOptions = async (): Promise<Omit<PiOptions, "checkedAt">> => {
let listed: PiModel[];
try {
listed = await deps.listModels();
} catch (error) {
if (isPiManagedConfigError(error)) {
throw new PiManagementError("pi_management_unavailable", PI_MANAGED_CONFIG_ERROR_MESSAGE);
}
throw new PiManagementError("pi_management_unavailable", "Pi model choices are unavailable");
}
const models: Array<{ provider: string; id: string }> = [];
const providers: string[] = [];
const seenModels = new Set<string>();
const seenProviders = new Set<string>();
for (const model of listed) {
if (!isChoice(model?.provider) || !isChoice(model?.id)) continue;
const key = `${model.provider}\u0000${model.id}`;
if (seenModels.has(key)) continue;
seenModels.add(key);
models.push({ provider: model.provider, id: model.id });
if (!seenProviders.has(model.provider)) {
seenProviders.add(model.provider);
providers.push(model.provider);
}
}
return { providers, models, reasoning: [...REASONING_CHOICES] };
};
const version = async (timeoutMs = config.piManagementTimeoutMs): Promise<string> => {
let output: { stdout: string; stderr: string };
try {
@@ -180,12 +151,12 @@ export function createPiManagement(config: AppConfig, deps: PiManagementDeps): P
const installationConfig = (): PiInstallationConfig => {
const settings = readSettings();
const provider = config.defaults.provider ?? settings.provider;
const model = config.defaults.model ?? settings.model;
const reasoning = config.defaults.thinking ?? settings.thinking;
const selected = deps.modelCatalog.defaultSession
? splitCanonicalModelId(deps.modelCatalog.defaultSession)
: undefined;
return {
...(isChoice(provider) ? { provider } : {}),
...(isChoice(model) ? { model } : {}),
...(selected ? selected : {}),
...(isReasoning(reasoning) ? { reasoning } : {}),
};
};
@@ -206,28 +177,6 @@ export function createPiManagement(config: AppConfig, deps: PiManagementDeps): P
}
},
async options(): Promise<PiOptions> {
const choices = await closedOptions();
return { ...choices, checkedAt: now().toISOString() };
},
async configure(value: PiInstallationConfig): Promise<PiInstallationConfig & { updatedAt: string }> {
if (!isInstallationConfig(value)) {
throw new PiManagementError("pi_management_invalid_config", "Pi installation configuration is invalid");
}
const choices = await closedOptions();
if (!choices.models.some((model) => model.provider === value.provider && model.id === value.model)) {
throw new PiManagementError("pi_management_invalid_config", "Pi provider and model must be selected from available choices");
}
try {
persistSettings({ ...readSettings(), provider: value.provider, model: value.model, thinking: value.reasoning });
} catch {
throw new PiManagementError("pi_management_write_failed", "Pi installation configuration could not be saved");
}
addDiagnostic("Pi installation defaults updated");
return { ...value, updatedAt: now().toISOString() };
},
async test(): Promise<PiTestResult> {
const checkedAt = now().toISOString();
const deadline = Date.now() + config.piManagementTimeoutMs;
@@ -293,24 +242,10 @@ async function defaultExecFile(command: string, args: string[], options: PiExecF
return { stdout: String(result.stdout), stderr: String(result.stderr) };
}
function isChoice(value: unknown): value is string {
return typeof value === "string" && value.length > 0 && value.length <= 128 && value.trim() === value
&& /^[A-Za-z0-9][A-Za-z0-9._/-]*$/u.test(value);
}
function isReasoning(value: unknown): value is PiReasoning {
return typeof value === "string" && (REASONING_CHOICES as readonly string[]).includes(value);
}
function isInstallationConfig(value: unknown): value is Required<PiInstallationConfig> {
if (!value || typeof value !== "object" || Array.isArray(value)) return false;
const candidate = value as Record<string, unknown>;
if (Object.keys(candidate).length !== 3 || Object.keys(candidate).some((key) => !["provider", "model", "reasoning"].includes(key))) {
return false;
}
return isChoice(candidate.provider) && isChoice(candidate.model) && isReasoning(candidate.reasoning);
}
function isTimeout(error: unknown): boolean {
return Boolean(
error && typeof error === "object" && (
+40 -12
View File
@@ -11,6 +11,10 @@ import {
configuredPiProviderApiKey,
createPiRuntimeAgentSnapshot,
} from "./managed-config.js";
import {
loadRuntimeModelCatalog,
type RuntimeModelCatalog,
} from "../models/runtime-model-catalog.js";
export interface SessionRuntime {
rpc: RpcClient;
@@ -42,23 +46,32 @@ export class PiProcessManager {
private runtimes = new Map<string, SessionRuntime>();
private agentSnapshotCleanups = new WeakMap<ChildProcessWithoutNullStreams, () => void>();
private spawnFn: (
sessionId: string, author: string, provider: string | undefined, principal?: PrincipalContext,
runtimeConfigPath?: string,
sessionId: string, author: string, provider: string | undefined, model: string | undefined,
principal?: PrincipalContext, runtimeConfigPath?: string,
) => ChildProcessWithoutNullStreams;
private loadAuthProviders: (agentDir: string) => ReadonlySet<string>;
private modelCatalog: RuntimeModelCatalog;
private modelCatalogConfigured: boolean;
constructor(
private cfg: AppConfig,
opts?: { spawnFn?: SpawnFn; authProviders?: (agentDir: string) => ReadonlySet<string> },
opts?: {
spawnFn?: SpawnFn;
authProviders?: (agentDir: string) => ReadonlySet<string>;
modelCatalog?: RuntimeModelCatalog;
},
) {
this.modelCatalog = opts?.modelCatalog ?? loadRuntimeModelCatalog(cfg.modelCatalogFile);
this.modelCatalogConfigured = cfg.modelCatalogFile !== undefined
|| this.modelCatalog.defaultSession !== null;
this.loadAuthProviders = opts?.authProviders
?? ((agentDir) => loadPiAuthProviders({ agentDir }));
if (opts?.spawnFn) {
this.spawnFn = (sessionId, author, provider, principal, runtimeConfigPath) =>
this.spawnPi(opts.spawnFn!, sessionId, author, provider, principal, runtimeConfigPath);
this.spawnFn = (sessionId, author, provider, model, principal, runtimeConfigPath) =>
this.spawnPi(opts.spawnFn!, sessionId, author, provider, model, principal, runtimeConfigPath);
} else {
this.spawnFn = (sessionId, author, provider, principal, runtimeConfigPath) =>
this.spawnPi(nodeSpawn, sessionId, author, provider, principal, runtimeConfigPath);
this.spawnFn = (sessionId, author, provider, model, principal, runtimeConfigPath) =>
this.spawnPi(nodeSpawn, sessionId, author, provider, model, principal, runtimeConfigPath);
}
}
@@ -71,7 +84,7 @@ export class PiProcessManager {
private spawnPi(
spawnFn: SpawnFn, sessionId: string, author: string, provider: string | undefined,
principal?: PrincipalContext, runtimeConfigPath?: string,
model: string | undefined, principal?: PrincipalContext, runtimeConfigPath?: string,
): ChildProcessWithoutNullStreams {
// This is the final shared boundary for createFor(), spawnFor(), and resume(). Validate
// before auth-provider inspection, then make Pi consume the exact copied bytes rather than
@@ -79,12 +92,23 @@ export class PiProcessManager {
const agent = createPiRuntimeAgentSnapshot();
let child: ChildProcessWithoutNullStreams | undefined;
try {
const catalogModel = provider && model
? this.modelCatalog.sessionModels()
.find((entry) => entry.provider === provider && entry.model === model)
: undefined;
const credentialName = catalogModel?.authentication.mode === "secret_env"
? catalogModel.authentication.apiKeyEnv
: undefined;
const projectedApiKey = configuredPiProviderApiKey(agent.models, provider)
?? (credentialName ? `$${credentialName}` : undefined);
const env = buildPiChildEnv({
provider,
authProviders: this.loadAuthProviders(agent.agentDir),
credentialValue: secretValue(this.cfg, "THT_MODEL_API_KEY"),
credentialValue: credentialName
? secretValue(this.cfg, credentialName)
: this.modelCatalogConfigured ? undefined : secretValue(this.cfg, "THT_MODEL_API_KEY"),
credentialFile: this.cfg.modelApiKeyFile,
configuredApiKey: configuredPiProviderApiKey(agent.models, provider),
configuredApiKey: projectedApiKey,
additions: { THT_SESSION: sessionId, THT_AUTHOR: author },
});
env.PI_CODING_AGENT_DIR = agent.agentDir;
@@ -162,9 +186,10 @@ export class PiProcessManager {
}
const author = o.author ?? "dev@local";
const provider = canonicalPiProvider(o.provider ?? this.cfg.defaults.provider);
const model = o.model ?? this.cfg.defaults.model;
let child: ChildProcessWithoutNullStreams;
try {
child = this.spawnFn(sessionId, author, provider, o.principal, o.runtimeConfig?.path);
child = this.spawnFn(sessionId, author, provider, model, o.principal, o.runtimeConfig?.path);
} catch (error) {
o.runtimeConfig?.release();
throw error;
@@ -235,8 +260,11 @@ export class PiProcessManager {
const thinking = o.thinking ?? this.cfg.defaults.thinking;
if (provider && model) {
const upstreamModel = this.modelCatalog.sessionModels()
.find((entry) => entry.provider === provider && entry.model === model)
?.upstreamModel ?? model;
const response = await rt.rpc.request(
{ type: "set_model", provider, modelId: model } as object & { type: string },
{ type: "set_model", provider, modelId: upstreamModel } as object & { type: string },
);
rt.bridge.setContextWindow(response?.data?.contextWindow);
}
+21 -3
View File
@@ -17,6 +17,10 @@ import {
readConfiguredPiAgentFile,
validateDeclarativePiConfig,
} from "./managed-config.js";
import {
loadRuntimeModelCatalog,
type RuntimeModelCatalog,
} from "../models/runtime-model-catalog.js";
const SMOKE_PROMPT = "Provider health check. Reply with exactly OK.";
const SMOKE_ARGS = [
@@ -49,6 +53,7 @@ interface ProviderSmokeOptions {
authProviders?: () => ReadonlySet<string>;
readAuthStore?: () => string;
readModelsStore?: () => string | undefined;
modelCatalog?: RuntimeModelCatalog;
}
export function createPiProviderSmoke(
@@ -69,12 +74,25 @@ export function createPiProviderSmoke(
const configuredModels = options.readModelsStore
? options.readModelsStore()
: readConfiguredPiAgentFile("models.json", true);
const catalog = options.modelCatalog ?? loadRuntimeModelCatalog(config.modelCatalogFile);
const catalogConfigured = config.modelCatalogFile !== undefined
|| catalog.defaultSession !== null;
const catalogModel = catalog.sessionModels()
.find((entry) => entry.provider === canonicalProvider && entry.model === model);
const upstreamModel = catalogModel?.upstreamModel ?? model;
const credentialName = catalogModel?.authentication.mode === "secret_env"
? catalogModel.authentication.apiKeyEnv
: undefined;
const projectedApiKey = configuredPiProviderApiKey(configuredModels, canonicalProvider)
?? (credentialName ? `$${credentialName}` : undefined);
const env = buildPiChildEnv({
provider: canonicalProvider,
authProviders: configuredAuthProviders,
credentialValue: secretValue(config, "THT_MODEL_API_KEY"),
credentialValue: credentialName
? secretValue(config, credentialName)
: catalogConfigured ? undefined : secretValue(config, "THT_MODEL_API_KEY"),
credentialFile: config.modelApiKeyFile,
configuredApiKey: configuredPiProviderApiKey(configuredModels, canonicalProvider),
configuredApiKey: projectedApiKey,
});
clearPrincipalEnvironment(env);
delete env.THT_DATA_ROOT;
@@ -113,7 +131,7 @@ export function createPiProviderSmoke(
const capabilityGuard = failOnUnexpectedCapabilities(rpc);
const turn = async (): Promise<void> => {
requireSuccessfulResponse(await rpc.request({
type: "set_model", provider: canonicalProvider, modelId: model,
type: "set_model", provider: canonicalProvider, modelId: upstreamModel,
} as object & { type: string }));
requireSuccessfulResponse(await rpc.request({
type: "set_thinking_level", level: reasoning,
@@ -9,10 +9,13 @@ import {
} from "../catalog/types.js";
const idSchema = z.uuid();
const consolidationSchema = z.object({
target: z.enum(["tables", "columns"]),
targetIds: z.array(idSchema).min(1).max(10_000),
}).strict();
const consolidationSchema = z.discriminatedUnion("target", [
z.object({
target: z.enum(["tables", "columns"]),
targetIds: z.array(idSchema).min(1).max(10_000),
}).strict(),
z.object({ target: z.literal("database_columns") }).strict(),
]);
function manage(request: FastifyRequest, reply: FastifyReply) {
return isPrincipalContext(requirePermission(request, reply, "database.manage"));
@@ -49,7 +52,7 @@ export function catalogDescriptionConsolidationRoutes(
try {
const databaseId = idSchema.parse((request.params as { databaseId?: unknown }).databaseId);
const input = consolidationSchema.parse(request.body);
const targetIds = [...new Set(input.targetIds)];
const targetIds = "targetIds" in input ? [...new Set(input.targetIds)] : [];
const result = await deps.operations.run(
databaseId,
async () => await deps.repository.consolidateGeneratedDescriptions(
@@ -11,38 +11,35 @@ import {
type DescriptionGenerationWorker,
} from "../catalog/description-generation-worker.js";
import { MetadataGenerationModelUnavailableError } from "../catalog/metadata-generation-models.js";
import { ModelCompletionProviderError } from "../catalog/model-completer.js";
import {
SensitiveDataSuggestionDuplicateTargetIdsError,
SensitiveDataSuggestionInvalidResponseError,
SensitiveDataSuggestionNoEligibleColumnsError,
SensitiveDataSuggestionPayloadTooLargeError,
SensitiveDataSuggestionTargetNotFoundError,
} from "../catalog/sensitive-data-suggester.js";
import type { SensitiveDataSuggestionRunner } from "../catalog/sensitive-data-suggestion-runner.js";
SensitivityAnalysisDuplicateTargetIdsError,
SensitivityAnalysisInterruptedError,
SensitivityAnalysisNoEligibleColumnsError,
SensitivityAnalysisTargetNotFoundError,
} from "../catalog/sensitivity-analysis-service.js";
import type { SensitivityAnalysisRunner } from "../catalog/sensitivity-analysis-runner.js";
import {
CatalogOperationInProgressError,
CatalogConnectorError,
CatalogUnavailableError,
DescriptionGenerationRunActiveError,
type CatalogRepository,
type DescriptionGenerationEvent,
type DescriptionGenerationRun,
type SensitiveDataSuggestionEvent,
type SensitiveDataSuggestionRun,
type SensitivityAnalysisEvent,
type SensitivityAnalysisRun,
} from "../catalog/types.js";
const idSchema = z.uuid();
const modelIdSchema = z.string().regex(/^[a-z][a-z0-9._-]{0,63}$/);
const modelIdSchema = z.string().regex(/^[a-z][a-z0-9._-]{0,63}\/[A-Za-z0-9][A-Za-z0-9._:-]{0,255}$/);
const selectedTargetIdsSchema = z.array(idSchema).min(1);
const suggestionSchema = z.discriminatedUnion("scope", [
z.object({ modelId: modelIdSchema, scope: z.literal("all") }).strict(),
z.object({ scope: z.literal("all") }).strict(),
z.object({
modelId: modelIdSchema,
scope: z.literal("selected_tables"),
targetIds: selectedTargetIdsSchema,
}).strict(),
z.object({
modelId: modelIdSchema,
scope: z.literal("selected_columns"),
targetIds: selectedTargetIdsSchema,
}).strict(),
@@ -112,7 +109,7 @@ function publicRun(run: DescriptionGenerationRun) {
};
}
function publicSensitiveDataSuggestionEvent(event: SensitiveDataSuggestionEvent) {
function publicSensitivityAnalysisEvent(event: SensitivityAnalysisEvent) {
return {
runId: event.runId,
sequence: event.sequence,
@@ -122,16 +119,19 @@ function publicSensitiveDataSuggestionEvent(event: SensitiveDataSuggestionEvent)
};
}
function publicSensitiveDataSuggestionRun(run: SensitiveDataSuggestionRun) {
function publicSensitivityAnalysisRun(run: SensitivityAnalysisRun) {
return {
id: run.id,
databaseId: run.databaseId,
scope: run.scope,
engine: run.engine,
modelId: run.modelId,
policyVersion: run.policyVersion,
status: run.status,
total: run.total,
suggestedSensitive: run.suggestedSensitive,
suggestedNonSensitive: run.suggestedNonSensitive,
unknown: run.unknown,
inputTokens: run.inputTokens,
cacheReadTokens: run.cacheReadTokens,
outputTokens: run.outputTokens,
@@ -232,16 +232,10 @@ function safeSuggestionError(reply: FastifyReply, error: unknown) {
if (error instanceof CatalogUnavailableError) {
return reply.code(503).send({
code: "catalog_unavailable",
message: "The database catalog is unavailable, so no sensitive-field suggestions were prepared.",
message: "The database catalog is unavailable, so no sensitivity assessments were prepared.",
});
}
if (error instanceof MetadataGenerationModelUnavailableError) {
return reply.code(409).send({
code: "metadata_generation_model_unavailable",
message: "The selected metadata-generation model is unavailable.",
});
}
if (error instanceof SensitiveDataSuggestionTargetNotFoundError) {
if (error instanceof SensitivityAnalysisTargetNotFoundError) {
const code = error.target === "database"
? "database_not_found"
: error.target === "table"
@@ -254,45 +248,39 @@ function safeSuggestionError(reply: FastifyReply, error: unknown) {
: "One or more selected Catalog Columns were not found in this database.";
return reply.code(404).send({ code, message });
}
if (error instanceof SensitiveDataSuggestionDuplicateTargetIdsError) {
if (error instanceof SensitivityAnalysisDuplicateTargetIdsError) {
return reply.code(400).send({
code: "sensitive_data_suggestion_target_ids_duplicate",
message: "Each selected table or column must appear only once.",
});
}
if (error instanceof SensitiveDataSuggestionNoEligibleColumnsError) {
if (error instanceof SensitivityAnalysisNoEligibleColumnsError) {
return reply.code(409).send({
code: "sensitive_data_suggestion_no_columns",
message: "The selected scope contains no Catalog Columns to classify.",
message: "The selected scope contains no Catalog Columns to assess.",
});
}
if (error instanceof SensitiveDataSuggestionPayloadTooLargeError) {
return reply.code(413).send({
code: "sensitive_data_suggestion_payload_too_large",
message: "The selected structural metadata cannot be divided into safe LLM requests.",
if (error instanceof SensitivityAnalysisInterruptedError) {
return reply.code(499).send({
code: "sensitivity_analysis_interrupted",
message: "Sensitivity analysis was interrupted before completion. No assessments were applied.",
});
}
if (error instanceof SensitiveDataSuggestionInvalidResponseError) {
if (error instanceof CatalogConnectorError) {
return reply.code(502).send({
code: "sensitive_data_suggestion_invalid_response",
message: "The LLM returned an incomplete or invalid classification. No suggestions were applied.",
});
}
if (error instanceof ModelCompletionProviderError) {
return reply.code(502).send({
code: "sensitive_data_suggestion_provider_unavailable",
message: "The selected LLM service could not complete the request. No suggestions were applied.",
code: "sensitivity_source_unavailable",
message: "The source values could not be inspected safely. No assessments were applied.",
});
}
if (error instanceof z.ZodError) {
return reply.code(400).send({
code: "sensitive_data_suggestion_request_invalid",
message: "Choose a database, one or more tables, or one or more columns to classify.",
message: "Choose a database, one or more tables, or one or more columns to assess.",
});
}
return reply.code(500).send({
code: "sensitive_data_suggestion_failed",
message: "Sensitive-field suggestions failed before review. No changes were applied.",
message: "Local sensitivity analysis failed before review. No changes were applied.",
});
}
@@ -300,18 +288,18 @@ function safeSuggestionHistoryError(reply: FastifyReply, error: unknown) {
if (error instanceof CatalogUnavailableError) {
return reply.code(503).send({
code: "catalog_unavailable",
message: "Sensitive Data Suggestion history is unavailable because the database catalog is unavailable.",
message: "Sensitivity Analysis history is unavailable because the database catalog is unavailable.",
});
}
if (error instanceof z.ZodError) {
return reply.code(400).send({
code: "sensitive_data_suggestion_history_request_invalid",
message: "Sensitive Data Suggestion history parameters are invalid.",
message: "Sensitivity Analysis history parameters are invalid.",
});
}
return reply.code(500).send({
code: "sensitive_data_suggestion_history_failed",
message: "Sensitive Data Suggestion history could not be loaded.",
message: "Sensitivity Analysis history could not be loaded.",
});
}
@@ -320,7 +308,7 @@ export function catalogDescriptionGenerationRoutes(
deps: {
repository: CatalogRepository;
worker: DescriptionGenerationWorker;
sensitiveDataSuggestionRunner: SensitiveDataSuggestionRunner;
sensitivityAnalysisRunner: SensitivityAnalysisRunner;
},
): void {
app.post("/catalog/databases/:databaseId/sensitive-data-suggestions", async (request, reply) => {
@@ -328,16 +316,25 @@ export function catalogDescriptionGenerationRoutes(
try {
const databaseId = idSchema.parse((request.params as { databaseId?: unknown }).databaseId);
const input = suggestionSchema.parse(request.body);
const result = await deps.sensitiveDataSuggestionRunner.run(
databaseId,
input.modelId,
input.scope,
"targetIds" in input ? input.targetIds : [],
new AbortController().signal,
);
const controller = new AbortController();
const abort = () => controller.abort();
request.raw.once("aborted", abort);
reply.raw.once("close", abort);
let result;
try {
result = await deps.sensitivityAnalysisRunner.run(
databaseId,
input.scope,
"targetIds" in input ? input.targetIds : [],
controller.signal,
);
} finally {
request.raw.off("aborted", abort);
reply.raw.off("close", abort);
}
return {
suggestions: result.suggestions,
run: publicSensitiveDataSuggestionRun(result.run),
run: publicSensitivityAnalysisRun(result.run),
};
} catch (error) {
return safeSuggestionError(reply, error);
@@ -348,8 +345,8 @@ export function catalogDescriptionGenerationRoutes(
if (!manage(request, reply)) return reply;
try {
const { limit } = historyQuerySchema.parse(request.query);
return (await deps.repository.listSensitiveDataSuggestionRuns(limit))
.map(publicSensitiveDataSuggestionRun);
return (await deps.repository.listSensitivityAnalysisRuns(limit))
.map(publicSensitivityAnalysisRun);
} catch (error) {
return safeSuggestionHistoryError(reply, error);
}
@@ -359,12 +356,12 @@ export function catalogDescriptionGenerationRoutes(
if (!manage(request, reply)) return reply;
try {
const runId = idSchema.parse((request.params as { runId?: unknown }).runId);
const run = await deps.repository.getSensitiveDataSuggestionRun(runId);
const run = await deps.repository.getSensitivityAnalysisRun(runId);
if (!run) return reply.code(404).send({
code: "sensitive_data_suggestion_run_not_found",
message: "Sensitive Data Suggestion Run was not found.",
message: "Sensitivity Analysis Run was not found.",
});
return publicSensitiveDataSuggestionRun(run);
return publicSensitivityAnalysisRun(run);
} catch (error) {
return safeSuggestionHistoryError(reply, error);
}
@@ -375,14 +372,14 @@ export function catalogDescriptionGenerationRoutes(
try {
const runId = idSchema.parse((request.params as { runId?: unknown }).runId);
const { after } = eventQuerySchema.parse(request.query);
if (!(await deps.repository.getSensitiveDataSuggestionRun(runId))) {
if (!(await deps.repository.getSensitivityAnalysisRun(runId))) {
return reply.code(404).send({
code: "sensitive_data_suggestion_run_not_found",
message: "Sensitive Data Suggestion Run was not found.",
message: "Sensitivity Analysis Run was not found.",
});
}
return (await deps.repository.listSensitiveDataSuggestionEvents(runId, after))
.map(publicSensitiveDataSuggestionEvent);
return (await deps.repository.listSensitivityAnalysisEvents(runId, after))
.map(publicSensitivityAnalysisEvent);
} catch (error) {
return safeSuggestionHistoryError(reply, error);
}
+18 -2
View File
@@ -19,8 +19,14 @@ const metadataSchema = z.object({
description: z.string().max(20_000).nullable().optional(),
generatedDescription: z.string().max(20_000).nullable().optional(),
sensitive: z.boolean().optional(),
sensitivityReason: z.string().max(2_000).nullable().optional(),
}).strict().refine((value) => (
"description" in value || "generatedDescription" in value || "sensitive" in value
"description" in value
|| "generatedDescription" in value
|| "sensitive" in value
|| "sensitivityReason" in value
)).refine((value) => (
value.sensitivityReason == null || value.sensitive === true
));
const createRunSchema = z.object({
version: z.number().int().positive(),
@@ -56,7 +62,10 @@ function safeError(reply: FastifyReply, error: unknown) {
return reply.code(409).send({ code: "schema_sync_conflict", message: error.message });
}
if (error instanceof CatalogConnectorError) {
return reply.code(502).send({ code: "schema_introspection_failed", message: "The database schema could not be read safely." });
return reply.code(502).send({
code: "schema_introspection_failed",
message: "The database schema could not be read. Check the connection and credentials, then try again.",
});
}
if (error instanceof z.ZodError) {
return reply.code(400).send({ code: "schema_request_invalid", message: "Schema request is invalid." });
@@ -113,6 +122,12 @@ export function catalogSchemaRoutes(
if (current.version !== input.version) {
return reply.code(409).send({ code: "column_stale", message: "Column metadata changed. Reload and try again." });
}
const nextSensitive = input.sensitive ?? current.sensitive;
const nextSensitivityReason = nextSensitive
? ("sensitivityReason" in input
? normalized(input.sensitivityReason ?? null)
: current.sensitivityReason)
: null;
const updated = await deps.repository.updateColumnMetadata(
databaseId,
tableId,
@@ -123,6 +138,7 @@ export function catalogSchemaRoutes(
? normalized(input.generatedDescription ?? null)
: current.generatedDescription,
input.sensitive,
nextSensitivityReason,
);
if (!updated) return reply.code(409).send({ code: "column_stale", message: "Column metadata changed. Reload and try again." });
return updated;
+10 -15
View File
@@ -1,10 +1,8 @@
import { readdirSync } from "node:fs";
import { join } from "node:path";
import type { FastifyInstance } from "fastify";
import type { PiModel } from "../pi/list-models.js";
import { isPrincipalContext, requirePermission } from "../auth/authorization.js";
export type ListModelsFn = () => Promise<PiModel[]>;
import type { RuntimeModelCatalog } from "../models/runtime-model-catalog.js";
/**
* List YAML workspace configs found in <harnessDir>/workspaces/*.yaml.
@@ -25,20 +23,17 @@ export function listWorkspaces(harnessDir: string): { name: string; file: string
export function metaRoutes(
app: FastifyInstance,
deps: { harnessDir: string; listModels?: ListModelsFn },
deps: { harnessDir: string; modelCatalog: RuntimeModelCatalog },
): void {
app.get("/models", async (request, reply) => {
if (!isPrincipalContext(requirePermission(request, reply, "session.use"))) return reply;
const fn = deps.listModels ?? (async () => []);
try {
return { models: await fn() };
} catch (error) {
app.log.warn({
component: "pi-model-list",
errorType: error instanceof Error ? error.name : typeof error,
}, "Pi model listing failed");
// Graceful fallback: Pi may not be running; don't crash the server.
return { models: [] as PiModel[] };
}
return {
models: deps.modelCatalog.sessionModels().map((entry) => ({
provider: entry.provider,
id: entry.model,
name: entry.label,
reasoning: entry.session?.reasoning ?? false,
})),
};
});
}
+1 -9
View File
@@ -11,13 +11,6 @@ export function piManagementRoutes(
deps: { service: PiManagementService },
): void {
app.get("/pi-management/status", async (request, reply) => run(request, reply, deps, () => deps.service.status()));
app.get("/pi-management/options", async (request, reply) => run(request, reply, deps, () => deps.service.options()));
app.put("/pi-management/config", async (request, reply) => run(
request,
reply,
deps,
() => deps.service.configure((request.body ?? {}) as Record<string, unknown>),
));
app.post("/pi-management/test", async (request, reply) => run(request, reply, deps, () => deps.service.test()));
app.get("/pi-management/logs", async (request, reply) => run(request, reply, deps, () => deps.service.logs()));
}
@@ -38,8 +31,7 @@ async function run<T>(
return await action();
} catch (error) {
if (error instanceof PiManagementError) {
const statusCode = error.code === "pi_management_invalid_config" ? 400 : 503;
return reply.code(statusCode).send({ code: error.code, error: error.message });
return reply.code(503).send({ code: error.code, error: error.message });
}
return reply.code(503).send({ code: "pi_management_unavailable", error: "Pi management is unavailable" });
}
+18 -9
View File
@@ -6,12 +6,13 @@ import type { Settings } from "../settings/settings-store.js";
import { getPrincipal } from "../auth/auth.js";
import type { PrincipalContext } from "../auth/principal.js";
import type { ReadinessManager } from "../runtime/readiness-manager.js";
import type { ListModelsFn } from "./meta.js";
import type { ListModelsFn } from "../pi/list-models.js";
import type { WorkspaceRegistry } from "../workspaces/registry.js";
import { validateOperationalWorkspace, type WorkspaceDescriptor } from "../workspaces/schema.js";
import type { MaintenanceBarrier } from "../runtime/maintenance-gate.js";
import { hasPermission, isPrincipalContext, requirePermission } from "../auth/authorization.js";
import type { EffectiveRelationshipSnapshotProvider } from "../catalog/effective-relationship-snapshot.js";
import { splitCanonicalModelId, type RuntimeModelCatalog } from "../models/runtime-model-catalog.js";
const BOOTSTRAP_FAILURE_MESSAGE =
"Session startup failed. Check configuration and connectivity, then Resume the session.";
@@ -43,6 +44,7 @@ export function sessionRoutes(
maintenanceBarrier: MaintenanceBarrier;
/** Optional only for narrow route-test stubs and installations without a Catalog database. */
effectiveRelationships?: EffectiveRelationshipSnapshotProvider;
modelCatalog: RuntimeModelCatalog;
},
) {
const lifecycleTails = new Map<string, Promise<void>>();
@@ -358,7 +360,6 @@ export function sessionRoutes(
let workspaceId: string | undefined;
let workspaceRevision: string | undefined;
let workspaceDescriptor: WorkspaceDescriptor | undefined;
let allowedModels: readonly string[] | undefined;
if (requestedWorkspaceId) {
try {
const registry = d.workspaceRegistry as Partial<WorkspaceRegistry>;
@@ -378,7 +379,6 @@ export function sessionRoutes(
workspaceId = resolved.revision.id;
workspaceRevision = resolved.revision.commit;
workspaceDescriptor = resolved.workspace;
allowedModels = resolved.workspace.llm_policy.allowed;
} catch {
return reply.code(409).send({
error: WORKSPACE_REVISION_UNAVAILABLE_MESSAGE,
@@ -386,12 +386,17 @@ export function sessionRoutes(
});
}
}
const provider = b.provider ?? s.provider;
const model = b.model ?? s.model;
const requestedCanonical = b.provider && b.model ? `${b.provider}/${b.model}` : undefined;
let selectedCanonical = requestedCanonical ?? d.modelCatalog.defaultSession;
let modelWarning: string | undefined;
if (selectedCanonical && d.modelCatalog.defaultSession && !d.modelCatalog.hasSession(selectedCanonical)) {
selectedCanonical = d.modelCatalog.defaultSession;
modelWarning = `Configured model ${requestedCanonical ?? "selection"} is unavailable; using ${selectedCanonical}.`;
}
const selected = selectedCanonical ? splitCanonicalModelId(selectedCanonical) : undefined;
const provider = selected?.provider ?? b.provider;
const model = selected?.model ?? b.model;
const thinking = b.thinking ?? s.thinking;
if (allowedModels && provider && model && !allowedModels.includes(`${provider}/${model}`)) {
return reply.code(400).send({ error: "Selected model is not allowed by this workspace." });
}
// A persisted session is resumable without keeping Pi alive. New work replaces every
// runtime owned by this principal, while runtimes belonging to other users remain intact.
// Optional chaining preserves the deliberately narrow manager stubs used by route tests.
@@ -490,7 +495,7 @@ export function sessionRoutes(
),
() => d.mgr.start(id, rt, runtimeOptions),
);
return { id };
return { id, ...(modelWarning ? { warning: modelWarning } : {}) };
} finally {
if (revisionLease && !manifestPersisted) {
await revisionLease.abort().catch((error: unknown) => {
@@ -598,6 +603,10 @@ export function sessionRoutes(
provider?: string; model?: string; thinking?: string;
workspace_id?: string; workspace_revision?: string;
};
const savedCanonical = saved.provider && saved.model ? `${saved.provider}/${saved.model}` : "";
if (d.modelCatalog.defaultSession && (!savedCanonical || !d.modelCatalog.hasSession(savedCanonical))) {
return reply.code(503).send({ error: MODEL_UNAVAILABLE_MESSAGE, code: "model_unavailable" });
}
let workspaceConfigPath: string;
let workspaceDescriptor: WorkspaceDescriptor | undefined;
try {
+16 -22
View File
@@ -1,17 +1,27 @@
import type { FastifyInstance } from "fastify";
import type { AppConfig } from "../config.js";
import type { Settings } from "../settings/settings-store.js";
import { listWorkspaces, type ListModelsFn } from "./meta.js";
import { listWorkspaces } from "./meta.js";
import type { PrincipalContext } from "../auth/principal.js";
import { isPrincipalContext, requirePermission } from "../auth/authorization.js";
import {
splitCanonicalModelId,
type RuntimeModelCatalog,
} from "../models/runtime-model-catalog.js";
/** Merge stored settings over env/first-workspace defaults. */
export function effectiveSettings(cfg: AppConfig, stored: Settings): Settings {
/** Merge only workspace and runtime-thinking preferences; model defaults belong to modelCatalog. */
export function effectiveSettings(
cfg: AppConfig,
stored: Settings,
modelCatalog?: RuntimeModelCatalog,
): Settings {
const workspaces = listWorkspaces(cfg.harnessDir);
const selected = modelCatalog?.defaultSession
? splitCanonicalModelId(modelCatalog.defaultSession)
: undefined;
return {
workspace: stored.workspace ?? workspaces[0]?.name,
provider: cfg.defaults.provider ?? stored.provider,
model: cfg.defaults.model ?? stored.model,
...(selected ?? {}),
thinking: cfg.defaults.thinking ?? stored.thinking,
};
}
@@ -19,7 +29,7 @@ export function effectiveSettings(cfg: AppConfig, stored: Settings): Settings {
export function settingsRoutes(
app: FastifyInstance,
deps: {
cfg: AppConfig; listModels: ListModelsFn;
cfg: AppConfig;
getSettings: (principal: PrincipalContext) => Promise<Settings>;
},
): void {
@@ -37,22 +47,6 @@ export function settingsRoutes(
const principal = requirePermission(req, reply, "settings.manage");
if (!isPrincipalContext(principal)) return principal;
const b = (req.body ?? {}) as Settings;
if (b.model) {
let available: { provider: string; id: string }[] = [];
try {
available = await deps.listModels();
} catch {
available = [];
}
// Only validate when Pi gave us a non-empty list; otherwise allow (degraded).
if (available.length > 0 && !available.some(
(candidate) => candidate.provider === b.provider && candidate.id === b.model,
)) {
return reply.code(400).send({
error: `Unknown model: ${b.provider ?? "unknown"}/${b.model}`,
});
}
}
try {
// Retain this endpoint as a validating compatibility surface for older clients, but do
// not write anonymous users' choices to shared server storage.
+48 -5
View File
@@ -16,23 +16,33 @@ import {
type WorkspaceDescriptor,
} from "../workspaces/schema.js";
import type { RuntimeBindings } from "../workspaces/runtime-renderer.js";
import type { ConnectorDiagnostics } from "../workspaces/diagnostics.js";
import type {
ConnectorDiagnostics,
Diagnostic,
WorkspaceDiagnosticOptions,
} from "../workspaces/diagnostics.js";
import { isPrincipalContext, requirePermission } from "../auth/authorization.js";
import type { AuthDiagnoser } from "../auth/diagnostics.js";
import { decodeAuthDiagnostics, type AuthDiagnostics } from "../auth/group-catalog.js";
import type { WorkspaceDatabase } from "../catalog/types.js";
export type WorkspaceDiagnoser = (
workspace: WorkspaceDescriptor,
bindings: RuntimeBindings,
options: { writeProbe: boolean },
options: WorkspaceDiagnosticOptions,
) => Promise<ConnectorDiagnostics>;
export type WorkspaceDatabaseTester = (
workspaceId: string,
) => Promise<WorkspaceDatabase | undefined>;
interface WorkspaceRoutesDeps {
registry: WorkspaceRegistry;
config: WorkspaceRegistryConfig;
diagnose: WorkspaceDiagnoser;
authDiagnoser: AuthDiagnoser;
secretStore: WorkspaceSecretStore;
testDatabaseConnection: WorkspaceDatabaseTester;
}
const workspaceId = z.string().regex(/^[a-z][a-z0-9-]{2,62}$/);
@@ -56,6 +66,20 @@ const SAFE_MESSAGES = {
semantic_index_incompatible: "Semantic index is incompatible with this workspace.",
} as const;
const catalogConnectionUnavailable = (): Diagnostic => ({
level: "error",
code: "connector_unavailable",
field: "dwh",
message: "The configured database could not be reached or authenticated.",
});
const catalogConnectionMissing = (): Diagnostic => ({
level: "error",
code: "binding_missing",
field: "dwh",
message: "Configure this workspace in Database Management before testing connections.",
});
function authenticationReport(value: unknown): AuthDiagnostics {
const report = decodeAuthDiagnostics(value);
if (!report) throw new Error("invalid authentication diagnostic report");
@@ -238,14 +262,33 @@ export function workspaceRoutes(app: FastifyInstance, deps: WorkspaceRoutesDeps)
deps.secretStore,
);
try {
const [workspaceDiagnostics, inspectedAuthentication] = await Promise.all([
deps.diagnose(operational, lease.bindings, { writeProbe: false }),
const [workspaceDiagnostics, testedDatabase, inspectedAuthentication] = await Promise.all([
deps.diagnose(operational, lease.bindings, {
writeProbe: false,
skipDwh: true,
}),
deps.testDatabaseConnection(id),
deps.authDiagnoser.inspect({ live: true }),
]);
const authentication = authenticationReport(inspectedAuthentication);
const catalogConnectionReady = testedDatabase?.connectionStatus === "reachable";
const catalogConnectionDiagnostic = !testedDatabase
? catalogConnectionMissing()
: catalogConnectionReady
? undefined
: catalogConnectionUnavailable();
const diagnostics = catalogConnectionDiagnostic
? [
...workspaceDiagnostics.diagnostics.filter(({ code }) => code !== "binding_ok"),
catalogConnectionDiagnostic,
]
: workspaceDiagnostics.diagnostics;
return {
...workspaceDiagnostics,
activatable: workspaceDiagnostics.activatable && authentication.ready,
activatable: workspaceDiagnostics.activatable
&& catalogConnectionReady
&& authentication.ready,
diagnostics,
authentication,
};
} finally {
+8 -1
View File
@@ -13,6 +13,7 @@ import type { AppConfig } from "../config.js";
export interface Settings {
workspace?: string;
/** Legacy input fields are ignored when loading/evaluating installation settings. */
provider?: string;
model?: string;
thinking?: string;
@@ -37,7 +38,13 @@ export function loadSettings(cfg: AppConfig): Settings {
try {
const raw = readFileSync(cfg.settingsFile, "utf8");
const parsed = JSON.parse(raw);
if (parsed && typeof parsed === "object") return parsed as Settings;
if (parsed && typeof parsed === "object") {
const value = parsed as Record<string, unknown>;
return {
...(typeof value.workspace === "string" ? { workspace: value.workspace } : {}),
...(typeof value.thinking === "string" ? { thinking: value.thinking } : {}),
};
}
return {};
} catch {
return {};
+5 -1
View File
@@ -598,7 +598,11 @@ export class ThtRunner {
} catch {
return { ok: false, code: "workspace_not_activatable" };
}
const collection = descriptor.semantic_index.vector_store;
const collection = {
collection: descriptor.workspace.id,
dimensions: this.cfg.semanticRuntime.internalEmbeddingDimensions,
distance: "cosine" as const,
};
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), Math.max(1, timeoutSec) * 1000);
try {
+1
View File
@@ -238,6 +238,7 @@ function createProductionService(): WorkspacePreprocessingService {
});
return new WorkspacePreprocessingService({
dataRoot: config.dataRoot ?? "/data",
embeddingDimensions: config.internalEmbeddingDimensions,
httpPrivateHostAllowlist: (process.env.THT_EVIDENCE_PRIVATE_HOST_ALLOWLIST ?? "")
.split(",").map((value) => value.trim()).filter((value) => value.length > 0),
acquireActiveRuntime: async (workspaceId) => {
+4 -4
View File
@@ -59,12 +59,12 @@ function safeSecretFilePath(path: string, secretRoots: readonly string[]): strin
function requireSupportedDescriptor(workspace: unknown): void {
if (typeof workspace !== "object" || workspace === null) {
throw new Error("Workspace bindings support only workspace schema version 3");
throw new Error("Workspace bindings support only workspace schema version 4");
}
const metadata = Reflect.get(workspace, "workspace");
if (typeof metadata !== "object" || metadata === null
|| Reflect.get(metadata, "schema_version") !== 3) {
throw new Error("Workspace bindings support only workspace schema version 3");
|| Reflect.get(metadata, "schema_version") !== 4) {
throw new Error("Workspace bindings support only workspace schema version 4");
}
}
@@ -155,7 +155,7 @@ export function resolveEvidenceBinding(
return { values, missing };
}
/** Resolve the complete schema-v3 runtime binding set. */
/** Resolve the complete schema-v4 runtime binding set. */
export function resolveRuntimeBindings(
workspace: WorkspaceDescriptor,
env: NodeJS.ProcessEnv,
+3 -3
View File
@@ -137,12 +137,12 @@ function evidenceVariables(
function requireSupportedDescriptor(workspace: unknown): void {
if (typeof workspace !== "object" || workspace === null) {
throw new Error("Installation contract supports only workspace schema version 3");
throw new Error("Installation contract supports only workspace schema version 4");
}
const metadata = Reflect.get(workspace, "workspace");
if (typeof metadata !== "object" || metadata === null
|| Reflect.get(metadata, "schema_version") !== 3) {
throw new Error("Installation contract supports only workspace schema version 3");
|| Reflect.get(metadata, "schema_version") !== 4) {
throw new Error("Installation contract supports only workspace schema version 4");
}
}
+79 -60
View File
@@ -130,6 +130,11 @@ export interface DiagnosticAdapters {
probeEmbedding(request: EmbeddingDiagnosticRequest): Promise<EmbeddingDiagnosticResult>;
}
export interface WorkspaceDiagnosticOptions {
writeProbe: boolean;
skipDwh?: boolean;
}
export const DEFAULT_WORKSPACE_DIAGNOSTIC_TIMEOUT_MS = 5_000;
async function secretPresent(file: string): Promise<boolean> {
@@ -406,12 +411,12 @@ function numericBinding(binding: Record<string, string>, name: string): number |
function requireSupportedDescriptor(workspace: unknown): void {
if (typeof workspace !== "object" || workspace === null) {
throw new Error("Workspace diagnoser supports only workspace schema version 3");
throw new Error("Workspace diagnoser supports only workspace schema version 4");
}
const metadata = Reflect.get(workspace, "workspace");
if (typeof metadata !== "object" || metadata === null
|| Reflect.get(metadata, "schema_version") !== 3) {
throw new Error("Workspace diagnoser supports only workspace schema version 3");
|| Reflect.get(metadata, "schema_version") !== 4) {
throw new Error("Workspace diagnoser supports only workspace schema version 4");
}
}
@@ -421,6 +426,7 @@ async function diagnoseValidatedWorkspace(
adapters: DiagnosticAdapters,
timeoutMs: number,
semanticRuntime: SemanticRuntimeConfig,
skipDwh: boolean,
): Promise<ConnectorDiagnostics> {
const evidenceField = descriptor.evidence?.source.type === "http"
? "evidence.source.authentication"
@@ -432,85 +438,93 @@ async function diagnoseValidatedWorkspace(
variable,
}));
const diagnostics = [
...[...bindings.dwh.missing].sort().map((field) => diagnosticError("binding_missing", field)),
...(skipDwh
? []
: [...bindings.dwh.missing].sort().map((field) => diagnosticError("binding_missing", field))),
...evidenceDiagnostics,
];
if (diagnostics.length > 0) return { activatable: false, diagnostics };
const dwhTimeout = boundedTimeout(descriptor.dwh.timeout_ms, timeoutMs);
let activatable = true;
const dwhValues = bindings.dwh.values;
const dwhField = (suffix: string) => bindingName(descriptor, suffix);
const dwhResource = { database: descriptor.dwh.database, schema: descriptor.dwh.schema };
let dwhRequest: ConnectorDiagnosticRequest | undefined;
if (bindings.dwh.transport === "rest_api") {
const diagnostic = descriptor.diagnostics?.dwh_rest;
const baseUrl = dwhValues[dwhField("BASE_URL")];
if (diagnostic && baseUrl) {
const credentialFile = diagnostic.auth === "none" ? undefined : dwhValues[dwhField("API_KEY_FILE")];
if (diagnostic.auth === "none" || credentialFile !== undefined) {
if (!skipDwh) {
const dwhTimeout = boundedTimeout(descriptor.dwh.timeout_ms, timeoutMs);
const dwhValues = bindings.dwh.values;
const dwhField = (suffix: string) => bindingName(descriptor, suffix);
const dwhResource = { database: descriptor.dwh.database, schema: descriptor.dwh.schema };
let dwhRequest: ConnectorDiagnosticRequest | undefined;
if (bindings.dwh.transport === "rest_api") {
const diagnostic = descriptor.diagnostics?.dwh_rest;
const baseUrl = dwhValues[dwhField("BASE_URL")];
if (diagnostic && baseUrl) {
const credentialFile = diagnostic.auth === "none" ? undefined : dwhValues[dwhField("API_KEY_FILE")];
if (diagnostic.auth === "none" || credentialFile !== undefined) {
dwhRequest = {
role: "dwh",
transport: "rest_api",
baseUrl,
credentialFile,
tlsCaFile: dwhValues[dwhField("TLS_CA_FILE")],
resource: dwhResource,
timeoutMs: dwhTimeout,
signal: new AbortController().signal,
diagnostic,
};
}
}
} else if (bindings.dwh.transport === "postgres_direct") {
const host = dwhValues[dwhField("HOST")];
const port = numericBinding(dwhValues, dwhField("PORT"));
const user = dwhValues[dwhField("USER")];
const credentialFile = dwhValues[dwhField("PASSWORD_FILE")];
if (host && port && user && credentialFile) {
dwhRequest = {
role: "dwh",
transport: "rest_api",
baseUrl,
transport: "postgres_direct",
host,
port,
user,
credentialFile,
tlsCaFile: dwhValues[dwhField("TLS_CA_FILE")],
resource: dwhResource,
timeoutMs: dwhTimeout,
signal: new AbortController().signal,
diagnostic,
};
}
}
} else if (bindings.dwh.transport === "postgres_direct") {
const host = dwhValues[dwhField("HOST")];
const port = numericBinding(dwhValues, dwhField("PORT"));
const user = dwhValues[dwhField("USER")];
const credentialFile = dwhValues[dwhField("PASSWORD_FILE")];
if (host && port && user && credentialFile) {
dwhRequest = {
role: "dwh",
transport: "postgres_direct",
host,
port,
user,
credentialFile,
tlsCaFile: dwhValues[dwhField("TLS_CA_FILE")],
resource: dwhResource,
timeoutMs: dwhTimeout,
signal: new AbortController().signal,
};
if (!dwhRequest) {
diagnostics.push(diagnosticError("workspace_not_activatable"));
return { activatable: false, diagnostics };
}
}
if (!dwhRequest) {
diagnostics.push(diagnosticError("workspace_not_activatable"));
return { activatable: false, diagnostics };
}
try {
const dwhResult = await withTimeout(dwhTimeout, (signal) => adapters.probeConnector({
...dwhRequest,
signal,
timeoutMs: dwhTimeout,
}));
if (!hasRequiredConnectorChecks(dwhResult, dwhRequest.resource)) {
try {
const dwhResult = await withTimeout(dwhTimeout, (signal) => adapters.probeConnector({
...dwhRequest,
signal,
timeoutMs: dwhTimeout,
}));
if (!hasRequiredConnectorChecks(dwhResult, dwhRequest.resource)) {
diagnostics.push(diagnosticError("connector_unavailable"));
activatable = false;
}
} catch {
diagnostics.push(diagnosticError("connector_unavailable"));
activatable = false;
}
} catch {
diagnostics.push(diagnosticError("connector_unavailable"));
activatable = false;
}
try {
const vector = await withTimeout(timeoutMs, (signal) => adapters.inspectQdrant({
baseUrl: semanticRuntime.internalQdrantUrl,
collection: descriptor.semantic_index.vector_store.collection,
collection: descriptor.workspace.id,
timeoutMs,
signal,
}));
const expected = descriptor.semantic_index.vector_store;
const expected = {
collection: descriptor.workspace.id,
dimensions: semanticRuntime.internalEmbeddingDimensions,
distance: "cosine",
};
if (vector.collection !== expected.collection
|| vector.dimensions !== expected.dimensions
|| vector.distance !== expected.distance) {
@@ -529,10 +543,8 @@ async function diagnoseValidatedWorkspace(
timeoutMs,
signal,
}));
if (semanticRuntime.internalEmbeddingModel !== descriptor.semantic_index.embedding.model
|| semanticRuntime.internalEmbeddingDimensions !== descriptor.semantic_index.embedding.dimensions
|| !embedding.available
|| embedding.dimensions !== descriptor.semantic_index.embedding.dimensions) {
if (!embedding.available
|| embedding.dimensions !== semanticRuntime.internalEmbeddingDimensions) {
diagnostics.push(diagnosticError("semantic_index_incompatible"));
activatable = false;
}
@@ -569,11 +581,18 @@ export function createWorkspaceDiagnoser(
return async function diagnose(
workspace: WorkspaceDescriptor,
bindings: RuntimeBindings,
_options: { writeProbe: boolean },
diagnosticOptions: WorkspaceDiagnosticOptions,
): Promise<ConnectorDiagnostics> {
requireSupportedDescriptor(workspace);
const descriptor = validateWorkspaceDescriptor(workspace);
return await diagnoseValidatedWorkspace(descriptor, bindings, adapters, timeoutMs, semanticRuntime);
return await diagnoseValidatedWorkspace(
descriptor,
bindings,
adapters,
timeoutMs,
semanticRuntime,
diagnosticOptions.skipDwh ?? false,
);
};
}
+6 -3
View File
@@ -2,7 +2,7 @@ import { createHash } from "node:crypto";
import { normalize } from "node:path";
export interface CanonicalEffectiveConfig {
schemaVersion: 1;
schemaVersion: 2;
dwh: CanonicalDwhConfig;
vector: CanonicalVectorConfig;
embedding: CanonicalEmbeddingConfig;
@@ -27,6 +27,7 @@ export interface CanonicalVectorConfig {
}
export interface CanonicalEmbeddingConfig {
id: string;
model: string;
dimensions: number;
}
@@ -137,8 +138,10 @@ function buildEmbeddingConfig(rendered: Record<string, unknown>): CanonicalEmbed
if (!embeddings) {
throw new TypeError("effective config is missing embedding resources");
}
const model = requireString(embeddings, "model");
return {
model: requireString(embeddings, "model"),
id: optionalString(embeddings, "id") ?? `ollama/${model}`,
model,
dimensions: requireNumber(embeddings, "dimensions"),
};
}
@@ -166,7 +169,7 @@ export function buildCanonicalEffectiveConfig(renderedConfig: unknown): Canonica
throw new TypeError("effective config requires a rendered configuration object");
}
return {
schemaVersion: 1,
schemaVersion: 2,
dwh: buildDwhConfig(rendered),
vector: buildVectorConfig(rendered),
embedding: buildEmbeddingConfig(rendered),
@@ -77,6 +77,7 @@ export interface WorkspacePreprocessingServiceDeps {
{ ok: true } | { ok: false; code: "workspace_not_activatable" | "semantic_index_incompatible" }
>;
httpPrivateHostAllowlist?: readonly string[];
embeddingDimensions?: number;
}
interface RunScope {
@@ -118,7 +119,7 @@ export class WorkspacePreprocessingService {
async vectorInspect(options: { workspaceId: string }): Promise<WorkspaceOperationResult> {
const runtime = await this.deps.acquireActiveRuntime(options.workspaceId);
const collection = runtime.workspace.semantic_index.vector_store.collection;
const collection = runtime.workspace.workspace.id;
const res = await fetch(`${runtime.configLease.semanticQdrantUrl}/collections/${encodeURIComponent(collection)}`, { method: "GET" });
if (!res.ok) return baseResult(runtime, "vector inspect", "failed", "semantic_index_incompatible", { warnings: ["collection unavailable"] });
const body = await res.json() as any;
@@ -132,7 +133,7 @@ export class WorkspacePreprocessingService {
async vectorRebuild(options: { workspaceId: string; collection?: string; confirm?: string; destroy?: boolean }): Promise<WorkspaceOperationResult> {
const runtime = await this.deps.acquireActiveRuntime(options.workspaceId);
const collection = runtime.workspace.semantic_index.vector_store.collection;
const collection = runtime.workspace.workspace.id;
if (options.collection !== collection || options.confirm !== collection || options.destroy !== true) {
return baseResult(runtime, "vector rebuild", "failed", "semantic_index_incompatible", { warnings: ["rebuild requires exact confirmation and --destroy"] });
}
@@ -143,8 +144,8 @@ export class WorkspacePreprocessingService {
const recreated = await reconcileCollection({
baseUrl: runtime.configLease.semanticQdrantUrl,
collection,
dimensions: runtime.workspace.semantic_index.vector_store.dimensions,
distance: runtime.workspace.semantic_index.vector_store.distance,
dimensions: this.deps.embeddingDimensions ?? 1024,
distance: "cosine",
mode: "self_heal",
});
if (!recreated.ok) return baseResult(runtime, "vector rebuild", "failed", "semantic_index_incompatible", { warnings: ["collection recreate failed"] });
@@ -400,6 +401,8 @@ export class WorkspacePreprocessingService {
catalogBlob: runtime.catalogBlob,
configDigest: runtime.configLease.configDigest,
bindingDigest: runtime.configLease.bindingDigest,
embeddingId: runtime.configLease.effectiveConfig.embedding.id,
embeddingDimensions: runtime.configLease.effectiveConfig.embedding.dimensions,
});
return { runtime, state, job };
}
+13 -3
View File
@@ -44,11 +44,13 @@ export interface BeginPreprocessingJobOptions {
catalogBlob: string;
configDigest: string;
bindingDigest: string;
embeddingId: string;
embeddingDimensions: number;
runId?: string;
}
export interface PreprocessingJobState {
schemaVersion: 1;
schemaVersion: 2;
runId: string;
operation: string;
workspaceId: string;
@@ -57,6 +59,8 @@ export interface PreprocessingJobState {
catalogBlob: string;
configDigest: string;
bindingDigest: string;
embeddingId: string;
embeddingDimensions: number;
completedStages: string[];
childRuns: Record<string, string>;
status: "active" | "succeeded" | "blocked" | "failed";
@@ -146,7 +150,7 @@ function decodeJob(value: unknown): PreprocessingJobState {
}
const record = value as Record<string, unknown>;
if (
record.schemaVersion !== 1
record.schemaVersion !== 2
|| typeof record.runId !== "string"
|| typeof record.operation !== "string"
|| typeof record.workspaceId !== "string"
@@ -155,6 +159,8 @@ function decodeJob(value: unknown): PreprocessingJobState {
|| typeof record.catalogBlob !== "string"
|| typeof record.configDigest !== "string"
|| typeof record.bindingDigest !== "string"
|| typeof record.embeddingId !== "string"
|| typeof record.embeddingDimensions !== "number"
|| !Array.isArray(record.completedStages)
|| typeof record.childRuns !== "object" || record.childRuns === null || Array.isArray(record.childRuns)
|| !["active", "succeeded", "blocked", "failed"].includes(String(record.status))
@@ -275,6 +281,8 @@ export class PreprocessingStateStore {
|| existing.catalogBlob !== options.catalogBlob
|| existing.configDigest !== options.configDigest
|| existing.bindingDigest !== options.bindingDigest
|| existing.embeddingId !== options.embeddingId
|| existing.embeddingDimensions !== options.embeddingDimensions
) {
throw new PreprocessingStateError(
"preprocessing_resume_mismatch",
@@ -295,7 +303,7 @@ export class PreprocessingStateStore {
}
}
const job: PreprocessingJobState = {
schemaVersion: 1,
schemaVersion: 2,
runId,
operation: options.operation,
workspaceId: this.options.workspaceId,
@@ -304,6 +312,8 @@ export class PreprocessingStateStore {
catalogBlob: options.catalogBlob,
configDigest: options.configDigest,
bindingDigest: options.bindingDigest,
embeddingId: options.embeddingId,
embeddingDimensions: options.embeddingDimensions,
completedStages: [],
childRuns: {},
status: "active",
+1 -1
View File
@@ -560,7 +560,7 @@ export class WorkspaceRegistry {
});
}
}
const collection = workspace.semantic_index.vector_store.collection;
const collection = workspace.workspace.id;
const owner = collectionOwners.get(collection);
if (owner !== undefined) {
throw new Error(`duplicate qdrant collection ownership: ${collection} (${owner}, ${id})`);
+24 -6
View File
@@ -34,6 +34,7 @@ export interface RuntimeInstallationOverlay {
export interface SemanticRuntimeConfig {
internalQdrantUrl: string;
internalEmbeddingUrl: string;
internalEmbeddingId?: string;
internalEmbeddingModel: string;
internalEmbeddingDimensions: number;
}
@@ -41,6 +42,7 @@ export interface SemanticRuntimeConfig {
export const DEFAULT_SEMANTIC_RUNTIME: SemanticRuntimeConfig = {
internalQdrantUrl: "http://qdrant:6333",
internalEmbeddingUrl: "http://embedding:11434",
internalEmbeddingId: "ollama/qwen3-embedding:0.6b",
internalEmbeddingModel: "qwen3-embedding:0.6b",
internalEmbeddingDimensions: 1024,
};
@@ -202,16 +204,16 @@ function placeholderConnection(identity: { database: string; schema: string }):
function requireSupportedDescriptor(workspace: unknown): void {
if (typeof workspace !== "object" || workspace === null) {
throw new Error("Runtime renderer supports only workspace schema version 3");
throw new Error("Runtime renderer supports only workspace schema version 4");
}
const metadata = Reflect.get(workspace, "workspace");
if (typeof metadata !== "object" || metadata === null
|| Reflect.get(metadata, "schema_version") !== 3) {
throw new Error("Runtime renderer supports only workspace schema version 3");
|| Reflect.get(metadata, "schema_version") !== 4) {
throw new Error("Runtime renderer supports only workspace schema version 4");
}
}
/** Render the schema-v3 compatibility fields consumed by the current Python harness. */
/** Render the schema-v4 compatibility fields consumed by the current Python harness. */
export function renderRuntimeConfig(
workspace: WorkspaceDescriptor,
bindings: RuntimeBindings,
@@ -263,16 +265,32 @@ export function renderRuntimeConfig(
...(installation.profile === undefined ? {} : { profile: installation.profile }),
language: descriptor.workspace.language,
database,
semantic_index: descriptor.semantic_index,
semantic_index: {
vector_store: {
engine: "qdrant",
collection: descriptor.workspace.id,
dimensions: semanticRuntime.internalEmbeddingDimensions,
distance: "cosine",
},
embedding: {
provider: "ollama_internal",
id: semanticRuntime.internalEmbeddingId
?? `ollama/${semanticRuntime.internalEmbeddingModel}`,
model: semanticRuntime.internalEmbeddingModel,
dimensions: semanticRuntime.internalEmbeddingDimensions,
},
},
resources: {
vector: {
engine: "qdrant",
base_url: semanticRuntime.internalQdrantUrl,
collection: descriptor.semantic_index.vector_store.collection,
collection: descriptor.workspace.id,
},
embeddings: {
provider: "ollama_internal",
base_url: semanticRuntime.internalEmbeddingUrl,
id: semanticRuntime.internalEmbeddingId
?? `ollama/${semanticRuntime.internalEmbeddingModel}`,
model: semanticRuntime.internalEmbeddingModel,
dimensions: semanticRuntime.internalEmbeddingDimensions,
},
+29 -67
View File
@@ -35,7 +35,7 @@ export interface CanonicalDiagnostics {
}
interface WorkspaceMetadata {
schema_version: 3;
schema_version: 4;
id: string;
name: string;
description?: string;
@@ -51,31 +51,12 @@ interface WorkspaceDwh {
supported_transports: DwhTransport[];
}
interface WorkspaceBase<TVectorStore> {
interface WorkspaceBase {
workspace: WorkspaceMetadata;
dwh: WorkspaceDwh;
semantic_index: {
vector_store: TVectorStore;
embedding: {
provider: "ollama_internal";
model: "qwen3-embedding:0.6b";
dimensions: 1024;
};
};
llm_policy: {
default?: `${string}/${string}`;
allowed: `${string}/${string}`[];
};
diagnostics?: Pick<CanonicalDiagnostics, "dwh_rest">;
}
interface QdrantVectorStore {
engine: "qdrant";
collection: string;
dimensions: 1024;
distance: "cosine";
}
export interface EvidencePolicy {
max_chunk_chars: number;
retain_published_generations: number;
@@ -120,12 +101,12 @@ export interface WorkspaceEvidence {
policy: EvidencePolicy;
}
export interface WorkspaceV3 extends WorkspaceBase<QdrantVectorStore> {
export interface WorkspaceV4 extends WorkspaceBase {
evidence?: WorkspaceEvidence;
}
export type CanonicalWorkspace = WorkspaceV3;
export type WorkspaceDescriptor = WorkspaceV3;
export type CanonicalWorkspace = WorkspaceV4;
export type WorkspaceDescriptor = WorkspaceV4;
const workspaceId = z.string().regex(/^[a-z][a-z0-9-]{2,62}$/, {
message: "workspace id must match ^[a-z][a-z0-9-]{2,62}$",
@@ -135,9 +116,6 @@ const identifier = z.string().regex(/^[A-Za-z_][A-Za-z0-9_]*$/, {
});
const port = z.number().int().min(1).max(65_535);
const timeoutMs = z.number().int().positive();
const modelReference = z.string().regex(/^[^/\s]+\/[^/\s]+$/, {
message: "model must use provider/model syntax",
});
function isOriginRelativeDiagnosticPath(value: string): boolean {
return /^\/(?!\/)[^\\\u0000-\u001F\u007F?#]*$/.test(value) && !/%5c/i.test(value);
@@ -166,21 +144,6 @@ const dwhSchema = z.object({
timeout_ms: timeoutMs.optional(),
supported_transports: z.array(z.enum(DWH_TRANSPORTS)).min(1),
}).strict();
const internalEmbeddingSchema = z.object({
provider: z.literal("ollama_internal"),
model: z.literal("qwen3-embedding:0.6b"),
dimensions: z.literal(1024),
}).strict();
const qdrantVectorStoreSchema = z.object({
engine: z.literal("qdrant"),
collection: workspaceId,
dimensions: z.literal(1024),
distance: z.literal("cosine"),
}).strict();
const llmPolicySchema = z.object({
default: modelReference.optional(),
allowed: z.array(modelReference).min(1),
}).strict();
const positiveSafeInteger = z.number().int().safe().positive();
const nonnegativeSafeInteger = z.number().int().safe().nonnegative();
@@ -366,7 +329,6 @@ function unique<T>(values: readonly T[], context: z.RefinementCtx, path: Propert
function workspaceInvariants(workspace: any, context: z.RefinementCtx): void {
unique(workspace.dwh.supported_transports, context, ["dwh", "supported_transports"]);
unique(workspace.llm_policy.allowed, context, ["llm_policy", "allowed"]);
if (workspace.evidence?.source.type === "filesystem") {
const expected = `${workspace.workspace.id}/evidence`;
@@ -379,20 +341,6 @@ function workspaceInvariants(workspace: any, context: z.RefinementCtx): void {
}
}
if (workspace.semantic_index.vector_store.dimensions !== workspace.semantic_index.embedding.dimensions) {
context.addIssue({
code: "custom",
path: ["semantic_index", "embedding", "dimensions"],
message: "embedding dimensions must match vector store dimensions",
});
}
if (workspace.llm_policy.default && !workspace.llm_policy.allowed.includes(workspace.llm_policy.default)) {
context.addIssue({
code: "custom",
path: ["llm_policy", "default"],
message: "LLM default must be included in the allowlist",
});
}
if (workspace.diagnostics?.dwh_rest && !workspace.dwh.supported_transports.includes("rest_api")) {
context.addIssue({
code: "custom",
@@ -402,23 +350,18 @@ function workspaceInvariants(workspace: any, context: z.RefinementCtx): void {
}
}
const WorkspaceV3Schema = z.object({
const WorkspaceV4Schema = z.object({
dwh: dwhSchema,
llm_policy: llmPolicySchema,
evidence: workspaceEvidenceSchema.optional(),
diagnostics: z.object({
dwh_rest: dwhRestDiagnostic.optional(),
}).strict().optional(),
workspace: z.object({
schema_version: z.literal(3), id: workspaceId, name: z.string().trim().min(1),
schema_version: z.literal(4), id: workspaceId, name: z.string().trim().min(1),
description: z.string().trim().min(1).optional(), language: z.enum(["en", "it"]),
}).strict(),
semantic_index: z.object({
vector_store: qdrantVectorStoreSchema,
embedding: internalEmbeddingSchema,
}).strict(),
}).strict().superRefine(workspaceInvariants);
const WorkspaceDescriptorSchema = WorkspaceV3Schema;
const WorkspaceDescriptorSchema = WorkspaceV4Schema;
export function parseWorkspaceYaml(source: string): WorkspaceDescriptor {
const documents = parseAllDocuments(source, { uniqueKeys: true });
@@ -431,6 +374,25 @@ export function parseWorkspaceYaml(source: string): WorkspaceDescriptor {
return validateWorkspaceDescriptor(document.toJSON());
}
/** Deterministically removes the two installation-owned v3 blocks without altering workspace data. */
export function migrateWorkspaceV3Yaml(source: string): string {
const documents = parseAllDocuments(source, { uniqueKeys: true });
if (documents.length !== 1) throw new Error("Workspace YAML must contain exactly one document");
const document = documents[0];
if (document.errors.length > 0 || document.warnings.length > 0) {
throw new Error("Invalid workspace YAML");
}
const value = document.toJSON() as Record<string, unknown>;
const metadata = value.workspace as Record<string, unknown> | undefined;
if (!metadata || metadata.schema_version !== 3) {
throw new Error("Workspace migration requires schema version 3");
}
metadata.schema_version = 4;
delete value.semantic_index;
delete value.llm_policy;
return serializeWorkspaceYaml(validateWorkspaceDescriptor(value));
}
export function validateWorkspaceDescriptor(workspace: unknown): WorkspaceDescriptor {
return WorkspaceDescriptorSchema.parse(workspace) as WorkspaceDescriptor;
}
@@ -439,11 +401,11 @@ export function isCanonicalWorkspace(workspace: unknown): workspace is Canonical
return WorkspaceDescriptorSchema.safeParse(workspace).success;
}
export function isOperationalWorkspace(workspace: unknown): workspace is WorkspaceV3 {
export function isOperationalWorkspace(workspace: unknown): workspace is WorkspaceV4 {
return WorkspaceDescriptorSchema.safeParse(workspace).success;
}
export function validateOperationalWorkspace(workspace: unknown): WorkspaceV3 {
export function validateOperationalWorkspace(workspace: unknown): WorkspaceV4 {
return validateWorkspaceDescriptor(workspace);
}
+1 -1
View File
@@ -19,4 +19,4 @@ export type WorkspaceErrorCode =
| "workspace_stale" | "git_unavailable" | "git_auth_failed" | "git_non_fast_forward"
| "connector_unavailable" | "semantic_index_incompatible";
export type { WorkspaceV3 } from "./schema.js";
export type { WorkspaceV4 } from "./schema.js";
+3 -16
View File
@@ -5,14 +5,12 @@ import { createLocalAuthFixture } from "./auth-test-fixtures.js";
function fakeService(): PiManagementService {
return {
status: vi.fn(async () => ({ ready: true })),
options: vi.fn(async () => ({ providers: [], models: [], reasoning: [], checkedAt: "2026-08-17T00:00:00.000Z" })),
configure: vi.fn(async (value) => ({ ...value, updatedAt: "2026-08-17T00:00:00.000Z" })),
test: vi.fn(async () => ({ ready: true, checkedAt: "2026-08-17T00:00:00.000Z" })),
logs: vi.fn(async () => ({ lines: [] })),
};
}
test("a local HTTPS cookie session authorizes Pi writes through an untrusted internal HTTP hop", async () => {
test("a local HTTPS cookie session authorizes the Pi smoke check through an untrusted internal HTTP hop", async () => {
const service = fakeService();
const fixture = await createLocalAuthFixture(
{ piManagement: service },
@@ -25,31 +23,21 @@ test("a local HTTPS cookie session authorizes Pi writes through an untrusted int
expect(fixture.publicUrl).toBe("HTTPS://thothii.example.test");
const proxyHeaders = fixture.sessionHeaders({ host: "127.0.0.1:8080" });
const configured = await fixture.app.inject({
method: "PUT",
url: "/pi-management/config",
headers: proxyHeaders,
payload: { provider: "zai", model: "glm-5.2", reasoning: "high" },
});
const smoke = await fixture.app.inject({
method: "POST",
url: "/pi-management/test",
headers: proxyHeaders,
});
expect(configured.statusCode).toBe(200);
expect(smoke.statusCode).toBe(200);
expect(service.configure).toHaveBeenCalledTimes(1);
expect(service.test).toHaveBeenCalledTimes(1);
fixture.resetDownstreamHits();
vi.mocked(service.configure).mockClear();
vi.mocked(service.test).mockClear();
const wrongOrigin = await fixture.app.inject({
method: "PUT",
url: "/pi-management/config",
method: "POST",
url: "/pi-management/test",
headers: fixture.sessionHeaders({ host: "127.0.0.1:8080", origin: "https://evil.example" }),
payload: { provider: "zai", model: "glm-5.2", reasoning: "high" },
});
const wrongCsrf = await fixture.app.inject({
method: "POST",
@@ -62,7 +50,6 @@ test("a local HTTPS cookie session authorizes Pi writes through an untrusted int
expect(response.json()).toEqual({ code: "csrf_failed", error: "Request origin validation failed" });
}
expect(fixture.downstreamHits()).toBe(0);
expect(service.configure).not.toHaveBeenCalled();
expect(service.test).not.toHaveBeenCalled();
} finally {
await fixture.close();
-13
View File
@@ -416,19 +416,6 @@ test("only two Argon2 verifications run concurrently and excess login attempts f
expect((await second).statusCode).toBe(401);
});
test("the real native asynchronous Argon2 verifier holds two permits and releases them after completion", async () => {
const { app } = await createLocalApp();
const first = login(app, { password: `${password}!` });
const second = login(app, { password: `${password}!` });
await new Promise<void>((resolve) => setImmediate(resolve));
const excess = await login(app, { password: `${password}!` });
expect(excess.statusCode).toBe(429);
await expect(first).resolves.toMatchObject({ statusCode: 401 });
await expect(second).resolves.toMatchObject({ statusCode: 401 });
await expect(login(app, { password: `${password}!` })).resolves.toMatchObject({ statusCode: 401 });
});
test("a verifier failure is sanitized and releases its concurrency permit", async () => {
let attempts = 0;
const user = {
+70 -7
View File
@@ -11,6 +11,7 @@ import type { ObservedSchemaSnapshot } from "../src/catalog/types.js";
import { WorkspaceSecretStore } from "../src/workspaces/secret-store.js";
import type { WorkspaceRegistry, WorkspaceRevision } from "../src/workspaces/registry.js";
import type { WorkspaceDescriptor } from "../src/workspaces/schema.js";
import type { WorkspaceDiagnoser } from "../src/routes/workspaces.js";
const roots: string[] = [];
afterEach(() => {
@@ -19,16 +20,11 @@ afterEach(() => {
});
const workspace: WorkspaceDescriptor = {
workspace: { schema_version: 3, id: "psd-clinical", name: "Policlinico San Donato", language: "it" },
workspace: { schema_version: 4, id: "psd-clinical", name: "Policlinico San Donato", language: "it" },
dwh: {
engine: "postgres", database: "warehouse", schema: "datawarehouse", port: 5432,
supported_transports: ["postgres_direct", "rest_api"],
},
semantic_index: {
vector_store: { engine: "qdrant", collection: "psd", dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
diagnostics: { dwh_rest: { method: "GET", path: "/health", auth: "bearer", response: { database: "database", schema: "schema" } } },
};
const revision: WorkspaceRevision = { id: "psd-clinical", commit: "a".repeat(40), blob: "b".repeat(40), snapshotPath: "/tmp/psd.yaml" };
@@ -38,6 +34,7 @@ function setup(
catalogDependencies: {
catalogOperationCoordinator?: CatalogOperationCoordinator;
catalogPostgresAccess?: CatalogPostgresAccess;
workspaceDiagnoser?: WorkspaceDiagnoser;
} = {},
workspaceDescriptor: WorkspaceDescriptor = workspace,
) {
@@ -61,7 +58,7 @@ function setup(
workspaceRegistry: registry,
workspaceSecretStore: secretStore,
catalogRepository: repository,
workspaceDiagnoser: vi.fn(),
workspaceDiagnoser: vi.fn(async () => ({ activatable: true, diagnostics: [] })),
...catalogDependencies,
});
return { app, secretStore, repository };
@@ -280,6 +277,72 @@ test("rejects a connection test while another catalog operation owns the databas
}
});
test("workspace and database tests use the same current catalog database binding", async () => {
const connect = vi.fn(async () => ({
query: vi.fn(async () => ({
rows: [{ database: "warehouse", schema: "datawarehouse" }],
})),
end: vi.fn(async () => undefined),
}));
const diagnose: WorkspaceDiagnoser = vi.fn(async () => ({
activatable: true,
diagnostics: [{
level: "info",
code: "binding_ok",
message: "Installation bindings and diagnostics succeeded.",
}],
}));
const { app } = setup({
THT_WS_PSD_CLINICAL_DWH_TRANSPORT: "postgres_direct",
THT_WS_PSD_CLINICAL_DWH_HOST: "legacy-db.internal",
THT_WS_PSD_CLINICAL_DWH_PORT: "5432",
THT_WS_PSD_CLINICAL_DWH_USER: "legacy-reader",
}, {
catalogPostgresAccess: { connect } as CatalogPostgresAccess,
workspaceDiagnoser: diagnose,
});
const created = (await app.inject({
method: "POST",
url: "/catalog/databases",
payload: {
...direct,
binding: { ...direct.binding, host: "current-db.internal", username: "current-reader" },
},
})).json();
const databaseTest = await app.inject({
method: "POST",
url: `/catalog/databases/${created.id}/test`,
payload: { version: created.version },
});
const workspaceTest = await app.inject({
method: "POST",
url: "/workspaces/psd-clinical/test",
payload: {},
});
expect(databaseTest.statusCode).toBe(200);
expect(workspaceTest.statusCode).toBe(200);
expect(connect).toHaveBeenCalledTimes(2);
expect(connect.mock.calls.map(([database]) => database)).toEqual([
expect.objectContaining({
databaseName: "warehouse",
schema: "datawarehouse",
binding: expect.objectContaining({ host: "current-db.internal", username: "current-reader" }),
}),
expect.objectContaining({
databaseName: "warehouse",
schema: "datawarehouse",
binding: expect.objectContaining({ host: "current-db.internal", username: "current-reader" }),
}),
]);
expect(diagnose).toHaveBeenCalledWith(
workspace,
expect.any(Object),
{ writeProbe: false, skipDwh: true },
);
});
test("returns exact global and per-database fleet metrics", async () => {
const { app, repository } = setup();
const database = await repository.create(direct);
@@ -14,6 +14,7 @@ import {
type ModelCompletionRequest,
} from "../src/catalog/model-completer.js";
import { CatalogOperationCoordinator } from "../src/catalog/operation-coordinator.js";
import type { SensitivityValueSource } from "../src/catalog/sensitivity-classifier.js";
import type {
CatalogDatabaseClient,
CatalogPostgresAccess,
@@ -24,7 +25,7 @@ import type { WorkspaceDescriptor } from "../src/workspaces/schema.js";
const workspace: WorkspaceDescriptor = {
workspace: {
schema_version: 3,
schema_version: 4,
id: "psd-clinical",
name: "Policlinico San Donato",
language: "it",
@@ -36,11 +37,6 @@ const workspace: WorkspaceDescriptor = {
port: 5432,
supported_transports: ["postgres_direct"],
},
semantic_index: {
vector_store: { engine: "qdrant", collection: "psd", dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
};
const revision: WorkspaceRevision = {
id: "psd-clinical",
@@ -49,7 +45,7 @@ const revision: WorkspaceRevision = {
snapshotPath: "/tmp/psd.yaml",
};
const configuredModel: ResolvedMetadataGenerationModel = {
id: "openai-mini",
id: "openai/gpt-4.1-mini",
provider: "openai",
model: "gpt-4.1-mini",
apiKeyEnv: "OPENAI_API_KEY",
@@ -74,6 +70,16 @@ async function setup(
sample: vi.fn(async () => []),
},
catalogPostgresAccess?: CatalogPostgresAccess,
sensitivityValueSource: SensitivityValueSource = {
scanTable: vi.fn(async (request, consume) => {
await consume(request.columns.map((column) => ({
columnId: column.id,
value: "ordinary",
characterLength: 8,
})));
return { kind: "complete", observedValues: 1 };
}),
},
) {
const repository = new MemoryCatalogRepository();
const database = await repository.create({
@@ -120,10 +126,11 @@ async function setup(
catalogOperationCoordinator: operations,
metadataGenerationModels: models(),
modelCompleter,
sensitivityValueSource,
...(descriptionSourceSampler ? { descriptionSourceSampler } : {}),
...(catalogPostgresAccess ? { catalogPostgresAccess } : {}),
});
return { app, repository, database, table, column, operations };
return { app, repository, database, table, column, operations, sensitivityValueSource };
}
async function waitForTerminalRun(app: ReturnType<typeof buildApp>, runId: string) {
@@ -141,22 +148,15 @@ async function waitForTerminalRun(app: ReturnType<typeof buildApp>, runId: strin
throw new Error(`Description Generation Run ${runId} did not finish`);
}
test("suggests sensitive flags from structural metadata without persisting them", async () => {
const modelCompleter = {
complete: vi.fn(async () => JSON.stringify({
suggestions: [{ columnId: expect.any(String), sensitive: true }],
})),
};
test("assesses sensitive flags locally without persisting them or calling an LLM", async () => {
const modelCompleter: ModelCompleter = { complete: vi.fn(async () => "unused") };
const { app, repository, database, table, column } = await setup(modelCompleter);
modelCompleter.complete.mockResolvedValueOnce(JSON.stringify({
suggestions: [{ columnId: column.id, sensitive: true }],
}));
try {
const response = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: { modelId: configuredModel.id, scope: "all" },
payload: { scope: "all" },
});
expect(response.statusCode).toBe(200);
@@ -165,11 +165,14 @@ test("suggests sensitive flags from structural metadata without persisting them"
run: {
databaseId: database.id,
scope: "all",
modelId: configuredModel.id,
engine: "local",
modelId: null,
policyVersion: "sensitivity-v4",
status: "completed",
total: 1,
suggestedSensitive: 1,
suggestedNonSensitive: 0,
unknown: 0,
errorSummary: null,
},
suggestions: [{
@@ -180,6 +183,9 @@ test("suggests sensitive flags from structural metadata without persisting them"
version: column.version,
currentSensitive: false,
sensitive: true,
assessment: "sensitive",
evidence: [{ kind: "metadata", ruleId: "metadata.direct_identifier" }],
coverage: "metadata",
}],
});
expect(await repository.getColumn(database.id, column.tableId, column.id))
@@ -212,49 +218,69 @@ test("suggests sensitive flags from structural metadata without persisting them"
runId: responseBody.run.id,
sequence: 1,
level: "info",
message: "Sensitive-field suggestion generation started.",
message: "Local sensitivity analysis started.",
},
{
runId: responseBody.run.id,
sequence: 2,
level: "info",
message: "Classified 1 of 1 columns.",
message: "Scanning source data: pass 1 of 3, table batch 1 of 1.",
},
{
runId: responseBody.run.id,
sequence: 3,
level: "info",
message: "Sensitive-field suggestion generation completed for 1 column.",
message: "Scanning source data: pass 2 of 3, table batch 1 of 1.",
},
{
runId: responseBody.run.id,
sequence: 4,
level: "info",
message: "Scanning source data: pass 3 of 3, table batch 1 of 1.",
},
{
runId: responseBody.run.id,
sequence: 5,
level: "info",
message: "Assessed 1 of 1 columns locally.",
},
{
runId: responseBody.run.id,
sequence: 6,
level: "info",
message: "Local sensitivity analysis completed for 1 column.",
},
]);
const request = modelCompleter.complete.mock.calls[0]![0] as ModelCompletionRequest;
const prompt = request.messages.map((message) => message.content).join("\n");
expect(prompt).toContain("patients");
expect(prompt).toContain("birth_date");
expect(prompt).toContain("date");
expect(prompt).not.toContain("Patient date of birth");
expect(prompt).not.toContain("test-provider-secret");
expect(modelCompleter.complete).not.toHaveBeenCalled();
} finally {
await app.close();
}
});
test("limits sensitive-data suggestions to the selected tables or columns", async () => {
const modelCompleter: ModelCompleter = {
complete: vi.fn(async (request) => {
const payload = JSON.parse(request.messages.find((message) => message.role === "user")!.content) as {
columns: Array<{ columnId: string; column: string }>;
};
return JSON.stringify({
suggestions: payload.columns.map((column) => ({
columnId: column.columnId,
sensitive: column.column.includes("name") || column.column.includes("note"),
})),
});
}),
};
const { app, repository, database } = await setup(modelCompleter);
test("does not impose a global HTTP deadline on sensitivity analysis", async () => {
const timeout = vi.spyOn(AbortSignal, "timeout");
const { app, repository, database } = await setup({ complete: vi.fn(async () => "unused") });
try {
const response = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: { scope: "all" },
});
expect(response.statusCode).toBe(200);
expect(timeout).not.toHaveBeenCalled();
expect(await repository.listSensitivityAnalysisRuns()).toHaveLength(1);
} finally {
timeout.mockRestore();
await app.close();
}
});
test("limits sensitivity analysis to the selected tables or columns", async () => {
const modelCompleter: ModelCompleter = { complete: vi.fn(async () => "unused") };
const { app, repository, database, sensitivityValueSource } = await setup(modelCompleter);
await repository.applySchemaSync(database.id, database.version, "all", [], {
schemaVersion: 1,
capabilities: { tables: "available", columns: "available", relationships: "available" },
@@ -286,7 +312,6 @@ test("limits sensitive-data suggestions to the selected tables or columns", asyn
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: {
modelId: configuredModel.id,
scope: "selected_tables",
targetIds: [visits.id, patients.id],
},
@@ -306,7 +331,6 @@ test("limits sensitive-data suggestions to the selected tables or columns", asyn
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: {
modelId: configuredModel.id,
scope: "selected_columns",
targetIds: [clinicalNote.id, status.id],
},
@@ -318,44 +342,41 @@ test("limits sensitive-data suggestions to the selected tables or columns", asyn
expect.objectContaining({ tableId: visits.id, columnId: clinicalNote.id, sensitive: true }),
]));
const prompts = vi.mocked(modelCompleter.complete).mock.calls.map(([request]) => (
JSON.parse(request.messages.find((message) => message.role === "user")!.content) as {
columns: Array<{ columnId: string }>;
}
));
expect(prompts[0]!.columns.map((column) => column.columnId).sort()).toEqual(
[...patientColumns, ...visitColumns].map((column) => column.id).sort(),
);
expect(prompts[0]!.columns.map((column) => column.columnId)).not.toContain(billingColumns[0]!.id);
expect(prompts[1]!.columns.map((column) => column.columnId).sort()).toEqual(
[status.id, clinicalNote.id].sort(),
const scannedColumnIds = vi.mocked(sensitivityValueSource.scanTable).mock.calls.flatMap(
([request]) => request.columns.map((column) => column.id),
);
expect(scannedColumnIds).toEqual([status.id, status.id]);
expect(scannedColumnIds).not.toContain(patientColumns.find(
(column) => column.name === "patient_name",
)!.id);
expect(scannedColumnIds).not.toContain(clinicalNote.id);
expect(scannedColumnIds).not.toContain(billingColumns[0]!.id);
expect(modelCompleter.complete).not.toHaveBeenCalled();
} finally {
await app.close();
}
});
test("explains invalid sensitive-data suggestion selections without calling the model", async () => {
test("explains invalid sensitivity-analysis selections without reading source values", async () => {
const modelCompleter: ModelCompleter = { complete: vi.fn(async () => "unused") };
const { app, database, table } = await setup(modelCompleter);
const { app, database, table, sensitivityValueSource } = await setup(modelCompleter);
try {
const empty = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: { modelId: configuredModel.id, scope: "selected_tables", targetIds: [] },
payload: { scope: "selected_tables", targetIds: [] },
});
expect(empty.statusCode).toBe(400);
expect(empty.json()).toEqual({
code: "sensitive_data_suggestion_request_invalid",
message: "Choose a database, one or more tables, or one or more columns to classify.",
message: "Choose a database, one or more tables, or one or more columns to assess.",
});
const duplicate = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: {
modelId: configuredModel.id,
scope: "selected_tables",
targetIds: [table.id, table.id],
},
@@ -370,7 +391,6 @@ test("explains invalid sensitive-data suggestion selections without calling the
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: {
modelId: configuredModel.id,
scope: "selected_tables",
targetIds: ["00000000-0000-4000-8000-000000000001"],
},
@@ -385,7 +405,6 @@ test("explains invalid sensitive-data suggestion selections without calling the
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: {
modelId: configuredModel.id,
scope: "selected_columns",
targetIds: ["00000000-0000-4000-8000-000000000002"],
},
@@ -396,205 +415,7 @@ test("explains invalid sensitive-data suggestion selections without calling the
message: "One or more selected Catalog Columns were not found in this database.",
});
expect(modelCompleter.complete).not.toHaveBeenCalled();
} finally {
await app.close();
}
});
test("batches sensitive-data suggestions for schemas larger than one helper message", async () => {
const maxHelperMessageBytes = 64 * 1024;
const seenColumnIds: string[] = [];
const modelCompleter: ModelCompleter = {
complete: vi.fn(async (request) => {
const userMessage = request.messages.find((message) => message.role === "user")!;
expect(Buffer.byteLength(userMessage.content, "utf8")).toBeLessThanOrEqual(maxHelperMessageBytes);
const payload = JSON.parse(userMessage.content) as {
columns: Array<{ columnId: string; column: string }>;
};
expect(payload.columns.length).toBeLessThanOrEqual(10);
seenColumnIds.push(...payload.columns.map((column) => column.columnId));
return JSON.stringify({
suggestions: payload.columns.map((column) => ({
columnId: column.columnId,
sensitive: column.column.endsWith("_private"),
})),
});
}),
};
const { app, repository, database } = await setup(modelCompleter);
const columnCount = 900;
await repository.applySchemaSync(database.id, database.version, "all", [], {
schemaVersion: 1,
capabilities: { tables: "available", columns: "available", relationships: "available" },
tables: [{ name: "wide_table", sourceComment: null }],
columns: Array.from({ length: columnCount }, (_, index) => ({
tableName: "wide_table",
name: `field_${index.toString().padStart(4, "0")}${index % 10 === 0 ? "_private" : ""}`,
ordinalPosition: index + 1,
dataType: "character varying(255)",
isNullable: true,
defaultExpression: null,
primaryKeyPosition: null,
sourceComment: null,
})),
relationships: [],
});
const wideTable = (await repository.listTables(database.id)).find((table) => table.name === "wide_table")!;
const expectedColumnIds = (await repository.listColumns(database.id, wideTable.id)).map((column) => column.id);
try {
const response = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: { modelId: configuredModel.id, scope: "all" },
});
expect(response.statusCode).toBe(200);
const suggestions = response.json().suggestions as Array<{
columnName: string;
currentSensitive: boolean;
sensitive: boolean;
}>;
expect(suggestions).toHaveLength(columnCount);
expect(suggestions).toEqual(expect.arrayContaining([
expect.objectContaining({ columnName: "field_0000_private", currentSensitive: false, sensitive: true }),
expect.objectContaining({ columnName: "field_0001", currentSensitive: false, sensitive: false }),
]));
expect(vi.mocked(modelCompleter.complete).mock.calls.length).toBeGreaterThan(1);
expect(seenColumnIds.slice().sort()).toEqual(expectedColumnIds.slice().sort());
expect(new Set(seenColumnIds).size).toBe(columnCount);
} finally {
await app.close();
}
});
test("retries one invalid sensitive-data classification before returning the review draft", async () => {
const modelCompleter: ModelCompleter = {
complete: vi.fn(async () => "unused"),
};
const { app, database, column } = await setup(modelCompleter);
vi.mocked(modelCompleter.complete)
.mockResolvedValueOnce("not-json")
.mockResolvedValueOnce(JSON.stringify({
suggestions: [{ columnId: column.id, sensitive: true }],
}));
try {
const response = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: { modelId: configuredModel.id, scope: "all" },
});
expect(response.statusCode).toBe(200);
expect(response.json().suggestions).toEqual([
expect.objectContaining({ columnId: column.id, sensitive: true }),
]);
expect(modelCompleter.complete).toHaveBeenCalledTimes(2);
} finally {
await app.close();
}
});
test.each(["malformed", "incomplete", "duplicate"] as const)(
"fails safely when sensitive-data suggestions are %s",
async (kind) => {
const modelCompleter: ModelCompleter = {
complete: vi.fn(async () => "unused"),
};
const { app, repository, database, column } = await setup(modelCompleter);
const rawResponse = kind === "malformed"
? "RAW_PROVIDER_RESPONSE_DO_NOT_EXPOSE_{"
: kind === "incomplete"
? JSON.stringify({ suggestions: [] })
: JSON.stringify({
suggestions: [
{ columnId: column.id, sensitive: true },
{ columnId: column.id, sensitive: true },
],
});
vi.mocked(modelCompleter.complete).mockResolvedValueOnce(rawResponse);
try {
const response = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: { modelId: configuredModel.id, scope: "all" },
});
expect(response.statusCode).toBe(502);
expect(response.json()).toEqual({
code: "sensitive_data_suggestion_invalid_response",
message: "The LLM returned an incomplete or invalid classification. No suggestions were applied.",
});
expect(response.body).not.toContain(rawResponse);
expect(await repository.getColumn(database.id, column.tableId, column.id))
.toMatchObject({ sensitive: false });
} finally {
await app.close();
}
},
);
test("explains a sensitive-data suggestion provider failure without exposing provider details", async () => {
const modelCompleter: ModelCompleter = {
complete: vi.fn(async () => {
throw new ModelCompletionProviderError();
}),
};
const { app, repository, database, column } = await setup(modelCompleter);
try {
const response = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sensitive-data-suggestions`,
payload: { modelId: configuredModel.id, scope: "all" },
});
expect(response.statusCode).toBe(502);
expect(response.json()).toEqual({
code: "sensitive_data_suggestion_provider_unavailable",
message: "The selected LLM service could not complete the request. No suggestions were applied.",
});
expect(response.body).not.toContain("model completion failed");
expect(await repository.getColumn(database.id, column.tableId, column.id))
.toMatchObject({ sensitive: false });
const history = await app.inject({
method: "GET",
url: "/catalog/sensitive-data-suggestion-runs",
});
expect(history.statusCode).toBe(200);
const [failedRun] = history.json();
expect(failedRun).toMatchObject({
databaseId: database.id,
status: "failed",
total: 1,
suggestedSensitive: 0,
suggestedNonSensitive: 0,
errorSummary: "Sensitive-field suggestion generation failed.",
});
const events = await app.inject({
method: "GET",
url: `/catalog/sensitive-data-suggestion-runs/${failedRun.id}/events-list`,
});
expect(events.statusCode).toBe(200);
expect(events.json()).toMatchObject([
{
runId: failedRun.id,
sequence: 1,
level: "info",
message: "Sensitive-field suggestion generation started.",
},
{
runId: failedRun.id,
sequence: 2,
level: "error",
message: "Sensitive-field suggestion generation failed.",
},
]);
expect(events.body).not.toContain("model completion failed");
expect(sensitivityValueSource.scanTable).not.toHaveBeenCalled();
} finally {
await app.close();
}
@@ -705,7 +526,7 @@ test("generates one selected Catalog Column from a single JSON code fence", asyn
expect(completionRequest.messages[0]?.content).toContain('{"results":[');
expect(completionRequest.messages[1]?.content).toContain(`"targetId":"${column.id}"`);
expect(completionRequest.messages[1]?.content).not.toMatch(/source rows|samples|example values/i);
expect(start.body).not.toMatch(/test-provider-secret|gpt-4\.1|openai\/gpt|Catalog metadata/);
expect(start.body).not.toMatch(/test-provider-secret|Catalog metadata/);
resolveCompletion(`\`\`\`json\n${JSON.stringify({
results: [{
@@ -2909,7 +2730,7 @@ test("validates selected targets and resolves every requested target before laun
const unknownModel = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/description-generation-runs`,
payload: { modelId: "unknown-model", scope: "selected_columns", targetIds: [column.id] },
payload: { modelId: "openai/unknown-model", scope: "selected_columns", targetIds: [column.id] },
});
expect(unknownModel.statusCode).toBe(409);
expect(unknownModel.json().code).toBe("metadata_generation_model_unavailable");
@@ -14,6 +14,9 @@ import { up as upDescriptionGeneration } from "../src/catalog/migrations/005_des
import { up as upSensitiveDataFlag } from "../src/catalog/migrations/006_sensitive_data_flag.js";
import { up as upSensitiveSuggestionRuns } from "../src/catalog/migrations/007_sensitive_data_suggestion_runs.js";
import { up as upAiTokenUsage } from "../src/catalog/migrations/009_ai_token_usage.js";
import { up as upCanonicalModelIds } from "../src/catalog/migrations/010_canonical_model_ids.js";
import { up as upLocalSensitivityAnalysis } from "../src/catalog/migrations/011_local_sensitivity_analysis.js";
import { up as upSensitivityReason } from "../src/catalog/migrations/012_sensitivity_reason.js";
import { KyselyCatalogRepository, type CatalogDatabase } from "../src/catalog/repository.js";
import { loadConfig } from "../src/config.js";
import type { WorkspaceRegistry } from "../src/workspaces/registry.js";
@@ -50,6 +53,9 @@ test.skipIf(!dockerAvailable)("Fastify persists Description Generation success a
await upDescriptionGeneration(db);
await upSensitiveSuggestionRuns(db);
await upAiTokenUsage(db);
await upCanonicalModelIds(db);
await upLocalSensitivityAnalysis(db);
await upSensitivityReason(db);
const repository = new KyselyCatalogRepository(db);
const database = await repository.create({
workspaceId: "psd-clinical",
@@ -158,9 +164,9 @@ test.skipIf(!dockerAvailable)("Fastify persists Description Generation success a
}),
};
const models: MetadataGenerationModels = {
catalog: () => ({ models: [{ id: "openai-mini", label: "OpenAI Mini" }], default: "openai-mini" }),
catalog: () => ({ models: [{ id: "openai/gpt-4.1-mini", label: "OpenAI Mini" }], default: "openai/gpt-4.1-mini" }),
resolve: () => ({
id: "openai-mini",
id: "openai/gpt-4.1-mini",
provider: "openai",
model: "gpt-4.1-mini",
apiKeyEnv: "OPENAI_API_KEY",
@@ -205,12 +211,12 @@ test.skipIf(!dockerAvailable)("Fastify persists Description Generation success a
method: "POST",
url: `/catalog/databases/${database.id}/description-generation-runs`,
payload: {
modelId: "openai-mini",
modelId: "openai/gpt-4.1-mini",
scope: "selected_columns",
targetIds: [status.id, birthDate.id],
},
});
expect(successfulStart.statusCode).toBe(202);
expect(successfulStart.statusCode, successfulStart.body).toBe(202);
expect(await terminalRun(app, successfulStart.json().id)).toMatchObject({
status: "completed",
total: 2,
@@ -233,7 +239,7 @@ test.skipIf(!dockerAvailable)("Fastify persists Description Generation success a
const tableStart = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/description-generation-runs`,
payload: { modelId: "openai-mini", scope: "selected_tables", targetIds: [table.id] },
payload: { modelId: "openai/gpt-4.1-mini", scope: "selected_tables", targetIds: [table.id] },
});
expect(tableStart.statusCode).toBe(202);
expect(await terminalRun(app, tableStart.json().id)).toMatchObject({
@@ -253,7 +259,7 @@ test.skipIf(!dockerAvailable)("Fastify persists Description Generation success a
const failedStart = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/description-generation-runs`,
payload: { modelId: "openai-mini", scope: "selected_columns", targetIds: [status.id] },
payload: { modelId: "openai/gpt-4.1-mini", scope: "selected_columns", targetIds: [status.id] },
});
expect(failedStart.statusCode).toBe(202);
const failedRun = await terminalRun(app, failedStart.json().id);
@@ -283,7 +289,7 @@ test.skipIf(!dockerAvailable)("Fastify persists Description Generation success a
const allStart = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/description-generation-runs`,
payload: { modelId: "openai-mini", scope: "all" },
payload: { modelId: "openai/gpt-4.1-mini", scope: "all" },
});
expect(allStart.statusCode).toBe(202);
const allRun = await terminalRun(app, allStart.json().id);
@@ -327,7 +333,7 @@ test.skipIf(!dockerAvailable)("Fastify persists Description Generation success a
const missingStart = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/description-generation-runs`,
payload: { modelId: "openai-mini", scope: "missing" },
payload: { modelId: "openai/gpt-4.1-mini", scope: "missing" },
});
expect(missingStart.statusCode).toBe(202);
expect(await terminalRun(app, missingStart.json().id)).toMatchObject({
@@ -0,0 +1,131 @@
import { existsSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { afterEach, expect, test, vi } from "vitest";
import { PythonLocalNerDetector } from "../src/catalog/local-ner-detector.js";
const roots: string[] = [];
afterEach(() => {
vi.unstubAllEnvs();
for (const root of roots.splice(0)) rmSync(root, { recursive: true, force: true });
});
test("keeps a CPU-only local worker warm and returns sanitized evidence", async () => {
vi.stubEnv("THT_MODEL_API_KEY", "must-not-reach-worker");
const root = mkdtempSync(join(tmpdir(), "thothii-local-ner-"));
roots.push(root);
const helper = join(root, "fake_ner_worker.py");
writeFileSync(helper, `
import json
import os
import pathlib
import sys
root = pathlib.Path.cwd()
root.joinpath("runtime.json").write_text(json.dumps({
"argv": sys.argv,
"cuda": os.environ.get("CUDA_VISIBLE_DEVICES"),
"hip": os.environ.get("HIP_VISIBLE_DEVICES"),
"offline": os.environ.get("HF_HUB_OFFLINE"),
"inherited_secret": os.environ.get("THT_MODEL_API_KEY"),
"pid": os.getpid(),
}), encoding="utf-8")
print(json.dumps({"ready": True}), flush=True)
for line in sys.stdin:
request = json.loads(line)
root.joinpath("request.json").write_text(json.dumps(request), encoding="utf-8")
print(json.dumps({
"id": request["id"],
"ok": True,
"evidence": [{
"columnId": request["candidates"][0]["columnId"],
"label": "person",
"confidence": 0.93,
}],
}), flush=True)
`, "utf8");
const detector = new PythonLocalNerDetector({
pythonExecutable: "python3",
workerScript: helper,
modelPath: join(root, "pinned-model"),
cwd: root,
threads: 2,
startupTimeoutMs: 5_000,
});
const candidate = {
columnId: "33333333-3333-4333-8333-333333333333",
text: "Dimesso Mario Rossi",
};
try {
expect(detector.isReady()).toBe(false);
await detector.warmup();
expect(detector.isReady()).toBe(true);
expect(existsSync(join(root, "request.json"))).toBe(false);
await expect(detector.detect(
[candidate],
new AbortController().signal,
Date.now() + 5_000,
)).resolves.toEqual([{
columnId: candidate.columnId,
label: "person",
confidence: 0.93,
}]);
const firstRuntime = JSON.parse(readFileSync(join(root, "runtime.json"), "utf8"));
expect(firstRuntime).toMatchObject({
cuda: "",
hip: "",
offline: "1",
inherited_secret: null,
});
expect(JSON.stringify(firstRuntime.argv)).not.toContain(candidate.text);
expect(JSON.parse(readFileSync(join(root, "request.json"), "utf8")).candidates).toEqual([candidate]);
await detector.detect([candidate], new AbortController().signal, Date.now() + 5_000);
const secondRuntime = JSON.parse(readFileSync(join(root, "runtime.json"), "utf8"));
expect(secondRuntime.pid).toBe(firstRuntime.pid);
} finally {
await detector.close();
}
});
test("bounds worker startup by the caller deadline", async () => {
const root = mkdtempSync(join(tmpdir(), "thothii-local-ner-deadline-"));
roots.push(root);
const helper = join(root, "slow_ner_worker.py");
writeFileSync(helper, `
import json
import sys
import time
time.sleep(2)
print(json.dumps({"ready": True}), flush=True)
for line in sys.stdin:
request = json.loads(line)
print(json.dumps({"id": request["id"], "ok": True, "evidence": []}), flush=True)
`, "utf8");
const detector = new PythonLocalNerDetector({
pythonExecutable: "python3",
workerScript: helper,
modelPath: join(root, "pinned-model"),
cwd: root,
startupTimeoutMs: 5_000,
});
const startedAt = Date.now();
try {
await expect(detector.detect(
[{
columnId: "33333333-3333-4333-8333-333333333333",
text: "Dimesso Mario Rossi",
}],
new AbortController().signal,
startedAt + 50,
)).rejects.toThrow("local NER is unavailable");
expect(Date.now() - startedAt).toBeLessThan(1_000);
} finally {
await detector.close();
}
});
@@ -1,4 +1,5 @@
import { spawnSync } from "node:child_process";
import { randomUUID } from "node:crypto";
import { PostgreSqlContainer } from "@testcontainers/postgresql";
import { CamelCasePlugin, Kysely, PostgresDialect, sql } from "kysely";
import { Pool } from "pg";
@@ -14,6 +15,9 @@ import { up as upSensitiveDataFlag } from "../src/catalog/migrations/006_sensiti
import { up as upSensitiveSuggestionRuns } from "../src/catalog/migrations/007_sensitive_data_suggestion_runs.js";
import { up as upLogicalRelationships } from "../src/catalog/migrations/008_catalog_logical_relationships.js";
import { up as upAiTokenUsage } from "../src/catalog/migrations/009_ai_token_usage.js";
import { up as upCanonicalModelIds } from "../src/catalog/migrations/010_canonical_model_ids.js";
import { up as upLocalSensitivityAnalysis } from "../src/catalog/migrations/011_local_sensitivity_analysis.js";
import { up as upSensitivityReason } from "../src/catalog/migrations/012_sensitivity_reason.js";
const dockerAvailable = spawnSync("docker", ["info"], { stdio: "ignore" }).status === 0;
@@ -32,6 +36,42 @@ test.skipIf(!dockerAvailable)("PostgreSQL migration enforces one database per wo
await upDescriptionGeneration(db);
await upSensitiveSuggestionRuns(db);
await upAiTokenUsage(db);
const historicalDatabaseId = randomUUID();
await db.insertInto("workspaceDatabases").values({
id: historicalDatabaseId,
workspaceId: "migration-history",
engine: "postgres",
databaseName: "warehouse",
schemaName: "public",
}).execute();
await db.insertInto("descriptionGenerationRuns").values({
id: randomUUID(), databaseId: historicalDatabaseId, scope: "all",
modelId: "openai-mini", language: "en", status: "completed", total: 1,
processed: 1, generated: 1,
}).execute();
const historicalSuggestionRunId = randomUUID();
await db.insertInto("sensitiveDataSuggestionRuns").values({
id: historicalSuggestionRunId, databaseId: historicalDatabaseId, scope: "all",
modelId: "openai-mini", status: "completed", total: 1,
suggestedSensitive: 1,
}).execute();
await upCanonicalModelIds(db);
await upLocalSensitivityAnalysis(db);
await upSensitivityReason(db);
await expect(db.selectFrom("sensitiveDataSuggestionRuns")
.select(["engine", "modelId", "policyVersion", "unknown"])
.where("id", "=", historicalSuggestionRunId)
.executeTakeFirstOrThrow()).resolves.toMatchObject({
engine: "llm",
modelId: "openai-mini",
policyVersion: null,
unknown: 0,
});
await expect(db.insertInto("descriptionGenerationRuns").values({
id: randomUUID(), databaseId: historicalDatabaseId, scope: "all",
modelId: "openai/gpt-5-mini", language: "en", status: "completed", total: 1,
processed: 1, generated: 1,
}).execute()).resolves.toBeDefined();
await sql`CREATE ROLE thothii_catalog_runtime`.execute(db);
await upRuntimeSequencePrivileges(db);
const sequencePrivilege = await sql<{ allowed: boolean }>`
@@ -104,7 +144,11 @@ test.skipIf(!dockerAvailable)("PostgreSQL migration enforces one database per wo
patientName.description,
patientName.generatedDescription,
true,
)).toMatchObject({ sensitive: true });
"Local assessment matched content rule pii.person_name.",
)).toMatchObject({
sensitive: true,
sensitivityReason: "Local assessment matched content rule pii.person_name.",
});
expect(await repository.getCatalogMetrics(created.id)).toEqual({
scope: "database",
databaseId: created.id,
@@ -149,6 +193,7 @@ test.skipIf(!dockerAvailable)("PostgreSQL migration enforces one database per wo
)).toMatchObject({ updated: 1 });
expect(await repository.getColumn(created.id, patients.id, patientName.id)).toMatchObject({
sensitive: true,
sensitivityReason: "Local assessment matched content rule pii.person_name.",
sourceComment: "Sensitive patient name",
});
@@ -243,10 +288,33 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository performs scoped metadata cl
};
await repository.applySchemaSync(database.id, database.version, "all", [], snapshot);
const patients = (await repository.listTables(database.id)).find((table) => table.name === "patients")!;
const context = (await repository.getLogicalRelationshipContext(database.id))!;
const source = context.endpoints.find((endpoint) => (
endpoint.tableName === "visits" && endpoint.columnName === "id"
))!;
const target = context.endpoints.find((endpoint) => (
endpoint.tableName === "patients" && endpoint.columnName === "id"
))!;
const generatedCandidate = { sourceColumnId: source.columnId, targetColumnId: target.columnId };
expect(await repository.deleteTableMetadata(database.id, [patients.id], "relationships"))
.toEqual({ tables: 0, columns: 0, relationships: 1 });
await expect(repository.insertGeneratedLogicalRelationships(database.id, [generatedCandidate]))
.resolves.toBe(1);
await expect(repository.getCatalogMetrics(database.id))
.resolves.toMatchObject({ relationships: 2 });
expect(await repository.deleteDatabaseMetadata([database.id], "relationships"))
.toEqual({ tables: 0, columns: 0, relationships: 2 });
expect(await repository.listRelationships(database.id)).toEqual([]);
expect(await repository.listLogicalRelationships(database.id)).toEqual([]);
await expect(repository.getCatalogMetrics(database.id))
.resolves.toMatchObject({ relationships: 0 });
await repository.applySchemaSync(database.id, database.version, "relationships", [], snapshot);
await expect(repository.insertGeneratedLogicalRelationships(database.id, [generatedCandidate]))
.resolves.toBe(1);
expect(await repository.deleteTableMetadata(database.id, [patients.id], "relationships"))
.toEqual({ tables: 0, columns: 0, relationships: 2 });
expect(await repository.listRelationships(database.id)).toEqual([]);
expect(await repository.listLogicalRelationships(database.id)).toEqual([]);
expect(await repository.listColumns(database.id, patients.id)).toHaveLength(2);
expect((await repository.get(database.id))?.schemaSyncedVersion).toBeUndefined();
@@ -260,13 +328,24 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository performs scoped metadata cl
await repository.applySchemaSync(database.id, database.version, "columns", [patients.id], snapshot);
await repository.applySchemaSync(database.id, database.version, "relationships", [], snapshot);
const refreshedContext = (await repository.getLogicalRelationshipContext(database.id))!;
const refreshedSource = refreshedContext.endpoints.find((endpoint) => (
endpoint.tableName === "visits" && endpoint.columnName === "id"
))!;
const refreshedTarget = refreshedContext.endpoints.find((endpoint) => (
endpoint.tableName === "patients" && endpoint.columnName === "id"
))!;
await expect(repository.insertGeneratedLogicalRelationships(database.id, [{
sourceColumnId: refreshedSource.columnId,
targetColumnId: refreshedTarget.columnId,
}])).resolves.toBe(1);
expect(await repository.deleteDatabaseMetadata([
database.id,
"99999999-9999-4999-8999-999999999999",
], "tables")).toBeUndefined();
expect(await repository.listTables(database.id)).toHaveLength(2);
expect(await repository.deleteDatabaseMetadata([database.id], "tables"))
.toEqual({ tables: 2, columns: 4, relationships: 1 });
.toEqual({ tables: 2, columns: 4, relationships: 2 });
expect(await repository.get(database.id)).toBeDefined();
expect(await repository.listTables(database.id)).toEqual([]);
expect(await repository.listRelationships(database.id)).toEqual([]);
@@ -391,6 +470,9 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists description and se
await upDescriptionGeneration(db);
await upSensitiveSuggestionRuns(db);
await upAiTokenUsage(db);
await upCanonicalModelIds(db);
await upLocalSensitivityAnalysis(db);
await upSensitivityReason(db);
const repository = new KyselyCatalogRepository(db);
const firstDatabase = await repository.create({
workspaceId: "generation-one",
@@ -428,14 +510,14 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists description and se
const run = await repository.createDescriptionGenerationRun(
firstDatabase.id,
"selected_columns",
"openai-mini",
"openai/gpt-4.1-mini",
"it",
1,
);
expect(run).toMatchObject({
databaseId: firstDatabase.id,
scope: "selected_columns",
modelId: "openai-mini",
modelId: "openai/gpt-4.1-mini",
language: "it",
status: "queued",
total: 1,
@@ -450,7 +532,7 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists description and se
await expect(repository.createDescriptionGenerationRun(
secondDatabase.id,
"selected_columns",
"openai-mini",
"openai/gpt-4.1-mini",
"en",
1,
)).rejects.toThrow("A description generation run is already active");
@@ -461,7 +543,7 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists description and se
await expect(repository.createDescriptionGenerationRun(
secondDatabase.id,
"selected_columns",
"openai-mini",
"openai/gpt-4.1-mini",
"en",
1,
)).rejects.toThrow("A description generation run is already active");
@@ -505,7 +587,7 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists description and se
const next = await repository.createDescriptionGenerationRun(
secondDatabase.id,
"missing",
"openai-mini",
"openai/gpt-4.1-mini",
"en",
1,
);
@@ -533,7 +615,7 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists description and se
const allRun = await repository.createDescriptionGenerationRun(
firstDatabase.id,
"all",
"openai-mini",
"openai/gpt-4.1-mini",
"it",
2,
);
@@ -559,10 +641,10 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists description and se
]);
expect(await repository.getActiveDescriptionGenerationRun()).toBeUndefined();
const suggestionRun = await repository.createSensitiveDataSuggestionRun(
const suggestionRun = await repository.createSensitivityAnalysisRun(
firstDatabase.id,
"selected_columns",
"openai-mini",
{ engine: "local", policyVersion: "sensitivity-v1" },
);
expect(suggestionRun).toMatchObject({
databaseId: firstDatabase.id,
@@ -570,43 +652,49 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists description and se
total: 0,
suggestedSensitive: 0,
suggestedNonSensitive: 0,
unknown: 0,
engine: "local",
modelId: null,
policyVersion: "sensitivity-v1",
startedAt: expect.any(String),
});
await repository.appendSensitiveDataSuggestionEvent(
await repository.appendSensitivityAnalysisEvent(
suggestionRun.id,
"info",
"Sensitive-field suggestion generation started.",
);
await repository.appendSensitiveDataSuggestionEvent(
await repository.appendSensitivityAnalysisEvent(
suggestionRun.id,
"info",
"Sensitive-field suggestion generation completed for 2 columns.",
);
expect(await repository.updateSensitiveDataSuggestionRun(suggestionRun.id, {
expect(await repository.updateSensitivityAnalysisRun(suggestionRun.id, {
status: "completed",
total: 2,
suggestedSensitive: 1,
suggestedNonSensitive: 1,
suggestedNonSensitive: 0,
unknown: 1,
finishedAt: new Date().toISOString(),
})).toMatchObject({
status: "completed",
total: 2,
suggestedSensitive: 1,
suggestedNonSensitive: 1,
suggestedNonSensitive: 0,
unknown: 1,
});
expect(await repository.listSensitiveDataSuggestionEvents(suggestionRun.id, 1)).toEqual([
expect(await repository.listSensitivityAnalysisEvents(suggestionRun.id, 1)).toEqual([
expect.objectContaining({ sequence: 2, level: "info" }),
]);
expect((await repository.listSensitiveDataSuggestionRuns(1))[0]).toMatchObject({
expect((await repository.listSensitivityAnalysisRuns(1))[0]).toMatchObject({
id: suggestionRun.id,
});
const interruptedSuggestionRun = await repository.createSensitiveDataSuggestionRun(
const interruptedSuggestionRun = await repository.createSensitivityAnalysisRun(
secondDatabase.id,
"all",
"openai-mini",
{ engine: "local", policyVersion: "sensitivity-v1" },
);
expect(await repository.interruptActiveSensitiveDataSuggestionRuns(
expect(await repository.interruptActiveSensitivityAnalysisRuns(
"Sensitive-field suggestion generation was interrupted by backend restart.",
)).toEqual([
expect.objectContaining({
@@ -657,6 +745,8 @@ test.skipIf(!dockerAvailable)("PostgreSQL repository persists logical relationsh
const target = context.endpoints.find((item) => item.tableName === "users" && item.columnName === "id")!;
const created = await repository.insertLogicalRelationship(database.id, source.columnId, target.columnId, false);
expect(created).toMatchObject({ origin: "manual", status: "active" });
expect(await repository.getCatalogMetrics(database.id))
.toMatchObject({ relationships: 1 });
await expect(repository.insertLogicalRelationship(database.id, source.columnId, target.columnId, true))
.resolves.toBeUndefined();
+126 -10
View File
@@ -7,7 +7,11 @@ import { loadConfig } from "../src/config.js";
import { MemoryCatalogRepository } from "../src/catalog/memory-repository.js";
import { CatalogOperationCoordinator } from "../src/catalog/operation-coordinator.js";
import type { CatalogSchemaIntrospector } from "../src/catalog/schema-introspector.js";
import type { CatalogSyncRun, ObservedSchemaSnapshot } from "../src/catalog/types.js";
import {
CatalogConnectorError,
type CatalogSyncRun,
type ObservedSchemaSnapshot,
} from "../src/catalog/types.js";
import { WorkspaceSecretStore } from "../src/workspaces/secret-store.js";
import type { WorkspaceRegistry, WorkspaceRevision } from "../src/workspaces/registry.js";
import type { WorkspaceDescriptor } from "../src/workspaces/schema.js";
@@ -16,13 +20,8 @@ const roots: string[] = [];
afterEach(() => { for (const root of roots.splice(0)) rmSync(root, { recursive: true, force: true }); });
const workspace: WorkspaceDescriptor = {
workspace: { schema_version: 3, id: "psd-clinical", name: "Policlinico San Donato", language: "it" },
workspace: { schema_version: 4, id: "psd-clinical", name: "Policlinico San Donato", language: "it" },
dwh: { engine: "postgres", database: "warehouse", schema: "datawarehouse", port: 5432, supported_transports: ["postgres_direct"] },
semantic_index: {
vector_store: { engine: "qdrant", collection: "psd", dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
};
const revision: WorkspaceRevision = { id: "psd-clinical", commit: "a".repeat(40), blob: "b".repeat(40), snapshotPath: "/tmp/psd.yaml" };
@@ -169,6 +168,32 @@ test("synchronizes a full physical schema and derives primary and foreign key fl
expect((await repository.get(database.id))?.schemaSyncedVersion).toBe(database.version);
});
test("attempts synchronization after a failed connection test and reports the live access failure", async () => {
const { app, repository, database, scan } = await setup();
await repository.recordTest(database.id, database.version, {
connectionStatus: "failed",
testedVersion: database.version,
lastTestedAt: new Date().toISOString(),
lastErrorCode: "connector_unavailable",
lastErrorMessage: "The database connector could not be reached or authenticated.",
});
scan.mockRejectedValueOnce(new CatalogConnectorError("upstream credentials must not escape"));
const started = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/sync-runs`,
payload: { version: database.version, scope: "all", tableIds: [] },
});
expect(started.statusCode).toBe(202);
const failed = await waitFor(repository, started.json().id, "failed");
expect(scan).toHaveBeenCalledOnce();
expect(failed).toMatchObject({
errorCode: "schema_introspection_failed",
errorMessage: "The database schema could not be read. Check the connection and credentials, then try again.",
});
});
test("synchronizes columns for every catalog table when no table selection is supplied", async () => {
const { app, repository, database, setObserved } = await setup();
const tablesRun = await app.inject({
@@ -252,13 +277,18 @@ test("keeps generated descriptions editable and preserves them across synchroniz
});
const sensitiveOnly = await app.inject({
method: "PATCH", url: `/catalog/databases/${database.id}/tables/${patients.id}/columns/${idColumn.id}`,
payload: { version: editedColumn.json().version, sensitive: true },
payload: {
version: editedColumn.json().version,
sensitive: true,
sensitivityReason: "Local assessment matched content rule pii.email.",
},
});
expect(sensitiveOnly.statusCode).toBe(200);
expect(sensitiveOnly.json()).toMatchObject({
description: "Reviewed key",
generatedDescription: "Generated key draft",
sensitive: true,
sensitivityReason: "Local assessment matched content rule pii.email.",
});
const emptyPatch = await app.inject({
method: "PATCH", url: `/catalog/databases/${database.id}/tables/${patients.id}/columns/${idColumn.id}`,
@@ -273,6 +303,7 @@ test("keeps generated descriptions editable and preserves them across synchroniz
description: "Reviewed key",
generatedDescription: "Generated key draft",
sensitive: true,
sensitivityReason: "Local assessment matched content rule pii.email.",
});
});
@@ -363,6 +394,67 @@ test("consolidates non-empty generated column descriptions and preserves curated
expect(scan).not.toHaveBeenCalled();
});
test("consolidates generated descriptions for every column in a database", async () => {
const { app, repository, database, scan } = await setup();
await seedCatalog(repository, database);
const tables = await repository.listTables(database.id);
const patients = tables.find((table) => table.name === "patients")!;
const visits = tables.find((table) => table.name === "visits")!;
const patientId = (await repository.listColumns(database.id, patients.id))[0]!;
const visitColumns = await repository.listColumns(database.id, visits.id);
const visitId = visitColumns.find((column) => column.name === "id")!;
const visitPatientId = visitColumns.find((column) => column.name === "patient_id")!;
await repository.updateColumnMetadata(
database.id,
patients.id,
patientId.id,
patientId.version,
"Curated patient identifier",
"Generated patient identifier",
);
await repository.updateColumnMetadata(
database.id,
visits.id,
visitId.id,
visitId.version,
"Curated visit identifier",
"Generated visit identifier",
);
await repository.updateColumnMetadata(
database.id,
visits.id,
visitPatientId.id,
visitPatientId.version,
"Keep curated patient reference",
"",
);
const response = await app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/descriptions/consolidate`,
payload: { target: "database_columns" },
});
expect(response.statusCode).toBe(200);
expect(response.json()).toEqual({ copied: 2, skipped: 1 });
expect(await repository.getColumn(database.id, patients.id, patientId.id)).toMatchObject({
description: "Generated patient identifier",
generatedDescription: "Generated patient identifier",
version: patientId.version + 2,
});
expect(await repository.getColumn(database.id, visits.id, visitId.id)).toMatchObject({
description: "Generated visit identifier",
generatedDescription: "Generated visit identifier",
version: visitId.version + 2,
});
expect(await repository.getColumn(database.id, visits.id, visitPatientId.id)).toMatchObject({
description: "Keep curated patient reference",
generatedDescription: "",
version: visitPatientId.version + 1,
});
expect(scan).not.toHaveBeenCalled();
});
test("rejects description consolidation while the Workspace Database is reserved", async () => {
const { app, repository, database, operations } = await setup();
await seedCatalog(repository, database);
@@ -437,9 +529,14 @@ test("strictly validates description consolidation database and target ids", asy
url: `/catalog/databases/${database.id}/descriptions/consolidate`,
payload: { target: "tables", targetIds: [table.id], unexpected: true },
}),
app.inject({
method: "POST",
url: `/catalog/databases/${database.id}/descriptions/consolidate`,
payload: { target: "database_columns", targetIds: [table.id] },
}),
]);
expect(responses.map((response) => response.statusCode)).toEqual([400, 400, 400]);
expect(responses.map((response) => response.statusCode)).toEqual([400, 400, 400, 400]);
for (const response of responses) {
expect(response.json()).toEqual({
code: "description_consolidation_invalid",
@@ -529,6 +626,18 @@ test("deletes relationships for selected databases without deleting their tables
const { app, repository, database } = await setup();
await seedCatalog(repository, database);
const tables = await repository.listTables(database.id);
const context = (await repository.getLogicalRelationshipContext(database.id))!;
const source = context.endpoints.find((endpoint) => (
endpoint.tableName === "visits" && endpoint.columnName === "id"
))!;
const target = context.endpoints.find((endpoint) => (
endpoint.tableName === "patients" && endpoint.columnName === "id"
))!;
await expect(repository.insertGeneratedLogicalRelationships(database.id, [{
sourceColumnId: source.columnId,
targetColumnId: target.columnId,
}])).resolves.toBe(1);
await expect(repository.getCatalogMetrics(database.id)).resolves.toMatchObject({ relationships: 2 });
const response = await app.inject({
method: "POST",
@@ -537,10 +646,12 @@ test("deletes relationships for selected databases without deleting their tables
});
expect(response.statusCode).toBe(200);
expect(response.json()).toEqual({ tables: 0, columns: 0, relationships: 1 });
expect(response.json()).toEqual({ tables: 0, columns: 0, relationships: 2 });
expect(await repository.listTables(database.id)).toHaveLength(2);
expect(await repository.listColumns(database.id, tables[0]!.id)).not.toEqual([]);
expect(await repository.listRelationships(database.id)).toEqual([]);
expect(await repository.listLogicalRelationships(database.id)).toEqual([]);
await expect(repository.getCatalogMetrics(database.id)).resolves.toMatchObject({ relationships: 0 });
expect((await repository.get(database.id))?.schemaSyncedVersion).toBeUndefined();
});
@@ -676,6 +787,11 @@ test("rebuilds generated relationships and returns the exact summary", async ()
});
expect(first.statusCode).toBe(200);
expect(first.json()).toEqual({ added: 1, alreadyPresent: 0, excluded: 0, ambiguous: 0 });
const metrics = await app.inject({
method: "GET", url: `/catalog/metrics?databaseId=${database.id}`,
});
expect(metrics.statusCode).toBe(200);
expect(metrics.json()).toMatchObject({ relationships: 1 });
const generated = (await repository.listLogicalRelationships(database.id))[0]!;
await app.inject({
@@ -0,0 +1,265 @@
import { expect, test, vi } from "vitest";
import {
SensitivityAnalysisInterruptedError,
SensitivityAnalysisService,
} from "../src/catalog/sensitivity-analysis-service.js";
import { SensitivityAnalysisRunner } from "../src/catalog/sensitivity-analysis-runner.js";
import type { SensitivityClassifier } from "../src/catalog/sensitivity-classifier.js";
import type {
CatalogColumn,
CatalogRepository,
CatalogTable,
SensitivityAnalysisRun,
WorkspaceDatabase,
} from "../src/catalog/types.js";
const database = {
id: "11111111-1111-4111-8111-111111111111",
workspaceId: "psd-clinical",
engine: "postgres",
databaseName: "warehouse",
schema: "public",
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
connectionStatus: "reachable",
binding: { transport: "postgres_direct", host: "db.internal", port: 5432, username: "reader" },
} satisfies WorkspaceDatabase;
function catalogTable(id: string, name: string): CatalogTable {
return {
id,
databaseId: database.id,
name,
sourceComment: null,
description: null,
generatedDescription: null,
lastSyncedDatabaseVersion: 1,
lastSyncedAt: "2026-09-02T08:00:00Z",
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
};
}
function catalogColumn(id: string, tableId: string, name: string): CatalogColumn {
return {
id,
tableId,
name,
ordinalPosition: 1,
dataType: "text",
isNullable: true,
defaultExpression: null,
primaryKeyPosition: null,
isPrimaryKey: false,
isForeignKey: false,
foreignKeyCount: 0,
sourceComment: null,
description: null,
generatedDescription: null,
sensitive: false,
lastSyncedDatabaseVersion: 1,
lastSyncedAt: "2026-09-02T08:00:00Z",
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
};
}
const running: SensitivityAnalysisRun = {
id: "22222222-2222-4222-8222-222222222222",
databaseId: database.id,
scope: "all",
engine: "local",
modelId: null,
policyVersion: "sensitivity-v4",
status: "running",
total: 0,
suggestedSensitive: 0,
suggestedNonSensitive: 0,
unknown: 0,
inputTokens: 0,
cacheReadTokens: 0,
outputTokens: 0,
createdAt: "2026-09-02T08:00:00Z",
startedAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
finishedAt: null,
errorSummary: null,
};
test("stops catalog selection when the request expires during a catalog read", async () => {
const controller = new AbortController();
const listTables = vi.fn();
const repository = {
get: vi.fn(async () => {
controller.abort();
return database;
}),
listTables,
} as unknown as CatalogRepository;
const classifier = { assess: vi.fn() } as unknown as SensitivityClassifier;
const analysis = new SensitivityAnalysisService(repository, classifier);
await expect(analysis.analyze(
database.id,
"all",
[],
controller.signal,
)).rejects.toBeInstanceOf(SensitivityAnalysisInterruptedError);
expect(listTables).not.toHaveBeenCalled();
expect(classifier.assess).not.toHaveBeenCalled();
});
test("classifies all selected tables in one breadth-first run and reports coverage", async () => {
const firstTable = catalogTable("33333333-3333-4333-8333-333333333333", "patients");
const secondTable = catalogTable("44444444-4444-4444-8444-444444444444", "encounters");
const firstColumn = catalogColumn(
"55555555-5555-4555-8555-555555555555",
firstTable.id,
"status",
);
const secondColumn = catalogColumn(
"66666666-6666-4666-8666-666666666666",
secondTable.id,
"note",
);
const repository = {
get: vi.fn(async () => database),
listTables: vi.fn(async () => [firstTable, secondTable]),
listColumns: vi.fn(async (_databaseId: string, tableId: string) => (
tableId === firstTable.id ? [firstColumn] : [secondColumn]
)),
} as unknown as CatalogRepository;
const assess = vi.fn(async (
_targets,
_signal,
_nerBudget,
onActivity?: (message: string) => void | Promise<void>,
) => {
await onActivity?.("Scanning source data: pass 1 of 3, table batch 1 of 1.");
return [
{
columnId: firstColumn.id,
assessment: "non_sensitive" as const,
proposedSensitive: false,
evidence: [{ kind: "coverage" as const, ruleId: "coverage.sampled_1000" }],
observedValues: 1_000,
coverage: "sampled" as const,
},
{
columnId: secondColumn.id,
assessment: "sensitive" as const,
proposedSensitive: true,
evidence: [{ kind: "content" as const, ruleId: "pii.email" }],
observedValues: 12,
coverage: "sampled" as const,
},
];
});
const classifier = { assess } as unknown as SensitivityClassifier;
const onPrepared = vi.fn();
const onProgress = vi.fn();
const onActivity = vi.fn();
const suggestions = await new SensitivityAnalysisService(repository, classifier).analyze(
database.id,
"all",
[],
new AbortController().signal,
onPrepared,
onProgress,
onActivity,
);
expect(assess).toHaveBeenCalledOnce();
expect(assess.mock.calls[0]![0]).toEqual([
{ database, table: firstTable, columns: [firstColumn] },
{ database, table: secondTable, columns: [secondColumn] },
]);
expect(onPrepared).toHaveBeenCalledWith(2);
expect(onActivity).toHaveBeenCalledWith(
"Scanning source data: pass 1 of 3, table batch 1 of 1.",
);
expect(onProgress.mock.calls.map(([processed]) => processed)).toEqual([1, 2]);
expect(suggestions).toEqual([
expect.objectContaining({ columnId: firstColumn.id, sensitive: false, coverage: "sampled" }),
expect.objectContaining({ columnId: secondColumn.id, sensitive: true, coverage: "sampled" }),
]);
});
test("persists classifier activity in the running analysis event log", async () => {
let persisted = running;
const appendEvent = vi.fn(async () => undefined);
const repository = {
get: vi.fn(async () => database),
createSensitivityAnalysisRun: vi.fn(async () => running),
getSensitivityAnalysisRun: vi.fn(async () => persisted),
updateSensitivityAnalysisRun: vi.fn(async (
_runId: string,
changes: Partial<SensitivityAnalysisRun>,
) => {
persisted = { ...persisted, ...changes };
return persisted;
}),
appendSensitivityAnalysisEvent: appendEvent,
} as unknown as CatalogRepository;
const analysis = {
analyze: vi.fn(async (
_databaseId,
_scope,
_targetIds,
_signal,
onPrepared,
_onProgress,
onActivity,
) => {
await onPrepared?.(0);
await onActivity?.("Scanning source data: pass 1 of 3, table batch 1 of 1.");
return [];
}),
} as unknown as SensitivityAnalysisService;
await new SensitivityAnalysisRunner(repository, analysis).run(
database.id,
"all",
[],
new AbortController().signal,
);
expect(appendEvent).toHaveBeenCalledWith(
running.id,
"info",
"Scanning source data: pass 1 of 3, table batch 1 of 1.",
);
});
test("marks a created run interrupted if the request deadline expires during persistence", async () => {
const controller = new AbortController();
const update = vi.fn(async (_runId: string, changes: Partial<SensitivityAnalysisRun>) => ({
...running,
...changes,
}));
const repository = {
get: vi.fn(async () => database),
createSensitivityAnalysisRun: vi.fn(async () => {
controller.abort();
return running;
}),
updateSensitivityAnalysisRun: update,
appendSensitivityAnalysisEvent: vi.fn(async () => undefined),
} as unknown as CatalogRepository;
const analysis = { analyze: vi.fn() } as unknown as SensitivityAnalysisService;
const runner = new SensitivityAnalysisRunner(repository, analysis);
await expect(runner.run(database.id, "all", [], controller.signal))
.rejects.toBeInstanceOf(SensitivityAnalysisInterruptedError);
expect(analysis.analyze).not.toHaveBeenCalled();
expect(update).toHaveBeenCalledWith(running.id, expect.objectContaining({
status: "interrupted",
total: 0,
unknown: 0,
errorSummary: "Local sensitivity analysis was interrupted before completion.",
}));
});
@@ -0,0 +1,706 @@
import { expect, test, vi } from "vitest";
import {
SensitivityClassifier,
type LocalNerDetector,
type SensitivityNerBudget,
type SensitivityTableScan,
type SensitivityValueSource,
} from "../src/catalog/sensitivity-classifier.js";
import type { CatalogColumn, CatalogTable, WorkspaceDatabase } from "../src/catalog/types.js";
import { CatalogConnectorError } from "../src/catalog/types.js";
const database = {
id: "11111111-1111-4111-8111-111111111111",
workspaceId: "psd-clinical",
engine: "postgres",
databaseName: "warehouse",
schema: "public",
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
connectionStatus: "reachable",
binding: { transport: "postgres_direct", host: "db.internal", port: 5432, username: "reader" },
} satisfies WorkspaceDatabase;
const table = {
id: "22222222-2222-4222-8222-222222222222",
databaseId: database.id,
name: "observations",
sourceComment: null,
description: null,
generatedDescription: null,
lastSyncedDatabaseVersion: 1,
lastSyncedAt: "2026-09-02T08:00:00Z",
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
} satisfies CatalogTable;
function column(overrides: Partial<CatalogColumn> = {}): CatalogColumn {
return {
id: "33333333-3333-4333-8333-333333333333",
tableId: table.id,
name: "note",
ordinalPosition: 1,
dataType: "character varying",
isNullable: true,
defaultExpression: null,
primaryKeyPosition: null,
isPrimaryKey: false,
isForeignKey: false,
foreignKeyCount: 0,
sourceComment: null,
description: null,
generatedDescription: null,
sensitive: false,
lastSyncedDatabaseVersion: 1,
lastSyncedAt: "2026-09-02T08:00:00Z",
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
...overrides,
};
}
function source(scan: SensitivityTableScan): SensitivityValueSource {
return { scanTable: vi.fn(async (_request, consume) => {
for (const batch of scan.batches) await consume(batch);
return scan.coverage;
}) };
}
test("reports source-scan activity before a long table scan completes", async () => {
let releaseScan!: () => void;
const scanGate = new Promise<void>((resolve) => {
releaseScan = resolve;
});
const scanTable = vi.fn(async () => {
await scanGate;
return { kind: "complete" as const, observedValues: 0 };
});
const activity = vi.fn();
const analysis = new SensitivityClassifier({ scanTable }).assess(
[{ database, table, columns: [column()] }],
new AbortController().signal,
undefined,
activity,
);
await vi.waitFor(() => expect(scanTable).toHaveBeenCalledOnce());
releaseScan();
await analysis;
expect(activity).toHaveBeenCalledWith(
"Scanning source data: pass 1 of 3, table batch 1 of 1.",
);
});
test("one email hidden in a generically named column makes the whole column sensitive", async () => {
const target = column();
const values = source({
batches: [[
{ columnId: target.id, value: "nessun contatto", characterLength: 16 },
{ columnId: target.id, value: "mario.rossi@example.it", characterLength: 23 },
]],
coverage: { kind: "complete", observedValues: 2 },
});
const classifier = new SensitivityClassifier(values);
const [assessment] = await classifier.assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
columnId: target.id,
assessment: "sensitive",
proposedSensitive: true,
evidence: [{ kind: "content", ruleId: "pii.email" }],
});
});
test("one text value longer than 500 characters makes the whole column sensitive", async () => {
const target = column({ name: "comment" });
const values = source({
batches: [[{ columnId: target.id, value: "x".repeat(501), characterLength: 743 }]],
coverage: { kind: "sampled", observedValues: 1 },
});
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "sensitive",
proposedSensitive: true,
evidence: [{ kind: "length", ruleId: "text.over_500_characters" }],
});
});
test("scans every table at 300 before advancing to 1,000 and 3,000 values", async () => {
const otherTable = { ...table, id: "77777777-7777-4777-8777-777777777777", name: "events" };
const first = column({ name: "status" });
const second = column({
id: "88888888-8888-4888-8888-888888888888",
tableId: otherTable.id,
name: "comment",
});
const calls: string[] = [];
const values: SensitivityValueSource = {
scanTable: vi.fn(async (request, consume) => {
calls.push(`${request.table.name}:${request.valuesPerColumn}:${request.sampleOffset}`);
await consume(request.columns.map((item) => ({
columnId: item.id,
value: "ordinary",
characterLength: 8,
})));
return { kind: "sampled", observedValues: request.columns.length };
}),
};
await new SensitivityClassifier(values).assess([
{ database, table, columns: [first] },
{ database, table: otherTable, columns: [second] },
], new AbortController().signal);
expect(calls).toEqual([
"observations:300:0",
"events:300:0",
"observations:700:300",
"events:700:300",
"observations:2000:1000",
"events:2000:1000",
]);
});
test("runs at most two table scans concurrently", async () => {
const targets = Array.from({ length: 3 }, (_, index) => {
const targetTable = {
...table,
id: `00000000-0000-4000-8000-${(index + 1).toString().padStart(12, "0")}`,
name: `table_${index + 1}`,
};
return {
database,
table: targetTable,
columns: [column({
id: `10000000-0000-4000-8000-${(index + 1).toString().padStart(12, "0")}`,
tableId: targetTable.id,
name: `attribute_${index + 1}`,
})],
};
});
let active = 0;
let maximum = 0;
const values: SensitivityValueSource = {
scanTable: vi.fn(async () => {
active += 1;
maximum = Math.max(maximum, active);
await Promise.resolve();
active -= 1;
return { kind: "complete", observedValues: 0 };
}),
};
await new SensitivityClassifier(values).assess(targets, new AbortController().signal);
expect(maximum).toBe(2);
expect(values.scanTable).toHaveBeenCalledTimes(3);
});
test("aborts a peer table scan when another concurrent source scan fails", async () => {
const otherTable = { ...table, id: "77777777-7777-4777-8777-777777777777", name: "events" };
const first = column({ name: "status" });
const second = column({
id: "88888888-8888-4888-8888-888888888888",
tableId: otherTable.id,
name: "comment",
});
let peerSignal: AbortSignal | undefined;
const failure = new CatalogConnectorError("source unavailable");
const values: SensitivityValueSource = {
scanTable: vi.fn(async (request, _consume, scanSignal) => {
if (request.table.id === table.id) {
await Promise.resolve();
throw failure;
}
peerSignal = scanSignal;
return await new Promise((_resolve, reject) => {
scanSignal.addEventListener("abort", () => reject(scanSignal.reason), { once: true });
});
}),
};
await expect(new SensitivityClassifier(values).assess([
{ database, table, columns: [first] },
{ database, table: otherTable, columns: [second] },
], new AbortController().signal)).rejects.toBe(failure);
expect(peerSignal?.aborted).toBe(true);
});
test("stops sampling a column as soon as one value is sensitive", async () => {
const target = column();
const values: SensitivityValueSource = {
scanTable: vi.fn(async (request, consume) => {
await consume([{ columnId: target.id, value: "mario.rossi@example.it", characterLength: 23 }]);
return { kind: "sampled", observedValues: 1 };
}),
};
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(values.scanTable).toHaveBeenCalledOnce();
expect(assessment).toMatchObject({ assessment: "sensitive", proposedSensitive: true });
});
test("stops non-text columns after the 1,000-value stage", async () => {
const target = column({ dataType: "integer", name: "sequence_number" });
const values: SensitivityValueSource = {
scanTable: vi.fn(async (request, consume) => {
await consume([{ columnId: target.id, value: "42", characterLength: 2 }]);
return { kind: "sampled", observedValues: 1 };
}),
};
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(vi.mocked(values.scanTable).mock.calls.map(([request]) => request.valuesPerColumn))
.toEqual([300, 700]);
expect(assessment).toMatchObject({
assessment: "non_sensitive",
proposedSensitive: false,
coverage: "sampled",
evidence: [{ kind: "coverage", ruleId: "coverage.sampled_1000" }],
});
});
test("complete coverage classifies benign and empty columns as non-sensitive", async () => {
const benign = column({ id: "44444444-4444-4444-8444-444444444444", name: "status" });
const empty = column({ id: "55555555-5555-4555-8555-555555555555", name: "optional_note" });
const humanProtected = column({
id: "66666666-6666-4666-8666-666666666666",
name: "category",
sensitive: true,
});
const values = source({
batches: [[
{ columnId: benign.id, value: "active", characterLength: 6 },
{ columnId: empty.id, value: null, characterLength: null },
{ columnId: humanProtected.id, value: "administrative", characterLength: 14 },
]],
coverage: { kind: "complete", observedValues: 1 },
});
const assessments = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [benign, empty, humanProtected] },
new AbortController().signal,
);
expect(assessments).toEqual([
expect.objectContaining({ columnId: benign.id, assessment: "non_sensitive", proposedSensitive: false }),
expect.objectContaining({
columnId: empty.id,
assessment: "non_sensitive",
proposedSensitive: false,
evidence: [{ kind: "coverage", ruleId: "coverage.no_values" }],
}),
expect.objectContaining({
columnId: humanProtected.id,
assessment: "non_sensitive",
proposedSensitive: false,
}),
]);
});
test("sampled coverage without a match proposes non-sensitive independently of the current flag", async () => {
const target = column({ sensitive: true });
const values = source({
batches: [[{ columnId: target.id, value: "ordinary", characterLength: 8 }]],
coverage: { kind: "sampled", observedValues: 1 },
});
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "non_sensitive",
proposedSensitive: false,
evidence: [{ kind: "coverage", ruleId: "coverage.sampled_3000" }],
});
});
test("an unavailable source fails the analysis instead of producing unknown decisions", async () => {
const unresolved = column();
const metadataMatch = column({
id: "44444444-4444-4444-8444-444444444444",
name: "codice_fiscale",
});
const values: SensitivityValueSource = {
scanTable: vi.fn(async () => {
throw new CatalogConnectorError("upstream detail must not escape");
}),
};
await expect(new SensitivityClassifier(values).assessTable(
{ database, table, columns: [unresolved, metadataMatch] },
new AbortController().signal,
)).rejects.toBeInstanceOf(CatalogConnectorError);
});
test("strong Italian PII metadata is sensitive even when the source column is empty", async () => {
const target = column({ name: "codice_fiscale" });
const values = source({
batches: [],
coverage: { kind: "complete", observedValues: 0 },
});
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "sensitive",
proposedSensitive: true,
evidence: [{ kind: "metadata", ruleId: "metadata.direct_identifier" }],
});
expect(values.scanTable).not.toHaveBeenCalled();
});
test("excludes bigint primary keys from content analysis as non-informative identifiers", async () => {
const target = column({
name: "id",
dataType: "bigint",
primaryKeyPosition: 1,
isPrimaryKey: true,
});
const values = source({
batches: [[{
columnId: target.id,
value: "3471234567",
characterLength: 10,
}]],
coverage: { kind: "complete", observedValues: 1 },
});
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "non_sensitive",
proposedSensitive: false,
evidence: [{
kind: "type",
ruleId: "type.bigint_primary_key_non_informative",
label: "non-informative bigint primary key",
}],
coverage: "metadata",
});
expect(values.scanTable).not.toHaveBeenCalled();
});
test("infers an undeclared bigint column named pk as a non-informative primary-key identifier", async () => {
const target = column({
name: "pk",
dataType: "bigint",
primaryKeyPosition: null,
isPrimaryKey: false,
});
const values = source({
batches: [[{
columnId: target.id,
value: "3471234567",
characterLength: 10,
}]],
coverage: { kind: "complete", observedValues: 1 },
});
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "non_sensitive",
proposedSensitive: false,
evidence: [{
kind: "metadata",
ruleId: "metadata.bigint_pk_identifier_non_informative",
label: "non-informative conventional bigint primary-key identifier",
}],
coverage: "metadata",
});
expect(values.scanTable).not.toHaveBeenCalled();
});
test("still inspects phone-like values in bigint columns that are not primary keys", async () => {
const target = column({ name: "id", dataType: "bigint" });
const values = source({
batches: [[{
columnId: target.id,
value: "3471234567",
characterLength: 10,
}]],
coverage: { kind: "complete", observedValues: 1 },
});
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "sensitive",
proposedSensitive: true,
evidence: [{ kind: "content", ruleId: "pii.phone_number" }],
});
expect(values.scanTable).toHaveBeenCalledOnce();
});
test.each([
["RSSMRA85T10A562S", "pii.italian_fiscal_code"],
["IT60 X054 2811 1010 0000 0123 456", "financial.iban"],
["4111 1111 1111 1111", "financial.payment_card"],
["SWIFT DEUTDEFF500", "financial.bic"],
["Partita IVA 00743110157", "pii.italian_vat"],
["Passaporto YA1234567", "pii.passport_number"],
["Carta d'identità CA12345AA", "pii.identity_card"],
["Patente di guida U11234567A", "pii.drivers_license_number"],
["Chiamare +39 347 123 4567", "pii.phone_number"],
["Client 192.168.1.5", "network.ip_address"],
["Device 00:1B:44:11:3A:B7", "network.mac_address"],
["https://example.org/profiles/mario", "network.url"],
["550e8400-e29b-41d4-a716-446655440000", "pii.uuid"],
["AWS key AKIAIOSFODNN7EXAMPLE", "credential.access_key"],
["Diagnosi: carcinoma mammario con metastasi ossee", "health.clinical_term"],
["-----BEGIN PRIVATE KEY----- secret -----END PRIVATE KEY-----", "credential.private_key"],
['{"profile":{"email":"not yet supplied"}}', "pii.json_sensitive_key"],
] as const)("recognizes validated sensitive content without relying on the column name: %s", async (
value,
ruleId,
) => {
const target = column();
const values = source({
batches: [[{ columnId: target.id, value, characterLength: value.length }]],
coverage: { kind: "complete", observedValues: 1 },
});
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "sensitive",
evidence: [{ kind: "content", ruleId }],
});
});
test("does not make a malformed email decisive", async () => {
const target = column();
const [assessment] = await new SensitivityClassifier(source({
batches: [[{
columnId: target.id,
value: "contatto a@b..com non valido",
characterLength: 28,
}]],
coverage: { kind: "complete", observedValues: 1 },
})).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "non_sensitive",
evidence: [{ kind: "coverage", ruleId: "coverage.complete" }],
});
});
test("finds a valid email after a malformed candidate in the same value", async () => {
const target = column();
const [assessment] = await new SensitivityClassifier(source({
batches: [[{
columnId: target.id,
value: "contatto a@b..com; indirizzo valido mario.rossi@example.it",
characterLength: 58,
}]],
coverage: { kind: "complete", observedValues: 1 },
})).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "sensitive",
evidence: [{ kind: "content", ruleId: "pii.email" }],
});
});
test("optional local NER evidence can make otherwise ambiguous Italian text sensitive", async () => {
const target = column();
const values = source({
batches: [[{ columnId: target.id, value: "Dimesso Mario Rossi", characterLength: 19 }]],
coverage: { kind: "sampled", observedValues: 1 },
});
const detector: LocalNerDetector = {
detect: vi.fn(async () => [{ columnId: target.id, label: "person_name", confidence: 0.91 }]),
};
const [assessment] = await new SensitivityClassifier(values, detector).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(detector.detect).toHaveBeenCalledWith(
[{ columnId: target.id, text: "Dimesso Mario Rossi" }],
expect.any(AbortSignal),
expect.any(Number),
);
expect(assessment).toMatchObject({
assessment: "sensitive",
evidence: [{ kind: "ner", ruleId: "ner.entity", label: "person_name", confidence: 0.91 }],
});
});
test("does not wait for an optional NER worker that is still warming", async () => {
const target = column();
const detector: LocalNerDetector = {
isReady: () => false,
detect: vi.fn(async () => [{ columnId: target.id, label: "person", confidence: 0.99 }]),
};
const [assessment] = await new SensitivityClassifier(source({
batches: [[{ columnId: target.id, value: "Dimesso Mario Rossi", characterLength: 19 }]],
coverage: { kind: "sampled", observedValues: 1 },
}), detector).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(detector.detect).not.toHaveBeenCalled();
expect(assessment).toMatchObject({
assessment: "non_sensitive",
evidence: [{ kind: "coverage", ruleId: "coverage.sampled_3000" }],
});
});
test("bounds each optional NER request when an installation raises the per-table work limit", async () => {
const columns = Array.from({ length: 17 }, (_, index) => column({
id: `00000000-0000-4000-8000-${(index + 1).toString(16).padStart(12, "0")}`,
name: `attribute_${index + 1}`,
ordinalPosition: index + 1,
}));
const observations = columns.flatMap((item, columnIndex) => Array.from(
{ length: 8 },
(_, valueIndex) => ({
columnId: item.id,
value: `ordinary-${columnIndex}-${valueIndex}`,
characterLength: 13,
}),
));
const detector: LocalNerDetector = { detect: vi.fn(async () => []) };
await new SensitivityClassifier(source({
batches: [observations],
coverage: { kind: "complete", observedValues: 8 },
}), detector, { maxNerCandidatesPerTable: 136 }).assessTable(
{ database, table, columns },
new AbortController().signal,
);
expect(detector.detect).toHaveBeenCalledTimes(2);
expect(vi.mocked(detector.detect).mock.calls.map(([candidates]) => candidates.length)).toEqual([
128,
8,
]);
});
test("limits default NER work to two candidates spread across a wide table", async () => {
const columns = Array.from({ length: 10 }, (_, index) => column({
id: `10000000-0000-4000-8000-${(index + 1).toString(16).padStart(12, "0")}`,
name: `attribute_${index + 1}`,
ordinalPosition: index + 1,
}));
const detector: LocalNerDetector = { detect: vi.fn(async () => []) };
await new SensitivityClassifier(source({
batches: [columns.flatMap((item, columnIndex) => [0, 1].map((valueIndex) => ({
columnId: item.id,
value: `ordinary-${columnIndex}-${valueIndex}`,
characterLength: 13,
})))],
coverage: { kind: "complete", observedValues: 2 },
}), detector).assessTable(
{ database, table, columns },
new AbortController().signal,
);
expect(detector.detect).toHaveBeenCalledOnce();
const submitted = vi.mocked(detector.detect).mock.calls[0]![0];
expect(submitted).toHaveLength(2);
expect(new Set(submitted.map((candidate) => candidate.columnId)).size).toBe(2);
});
test("shares a bounded NER time allowance across tables in one analysis run", async () => {
const target = column();
const values = source({
batches: [[{ columnId: target.id, value: "Dimesso Mario Rossi", characterLength: 19 }]],
coverage: { kind: "sampled", observedValues: 1 },
});
const detector: LocalNerDetector = {
detect: vi.fn(async () => {
await new Promise((resolve) => setTimeout(resolve, 20));
return [];
}),
};
const classifier = new SensitivityClassifier(values, detector);
const nerBudget: SensitivityNerBudget = { remainingMs: 1 };
await classifier.assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
Date.now() + 1_000,
nerBudget,
);
await classifier.assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
Date.now() + 1_000,
nerBudget,
);
expect(detector.detect).toHaveBeenCalledOnce();
expect(nerBudget.remainingMs).toBe(0);
});
test("uninterpretable binary content is protected conservatively without scanning", async () => {
const target = column({ dataType: "bytea" });
const values = source({
batches: [[{ columnId: target.id, value: "\\xdeadbeef", characterLength: 10 }]],
coverage: { kind: "complete", observedValues: 1 },
});
const [assessment] = await new SensitivityClassifier(values).assessTable(
{ database, table, columns: [target] },
new AbortController().signal,
);
expect(assessment).toMatchObject({
assessment: "sensitive",
proposedSensitive: true,
evidence: [{ kind: "type", ruleId: "type.binary_uninspectable" }],
});
expect(values.scanTable).not.toHaveBeenCalled();
});
@@ -0,0 +1,305 @@
import { mkdtempSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { expect, test, vi } from "vitest";
import type { CatalogDatabaseClient, CatalogPostgresAccess } from "../src/catalog/postgres-access.js";
import { CATALOG_SECRET_IDS } from "../src/catalog/secrets.js";
import { ConcreteSensitivityValueSource } from "../src/catalog/sensitivity-value-source.js";
import {
CatalogConnectorError,
type CatalogColumn,
type CatalogTable,
type WorkspaceDatabase,
} from "../src/catalog/types.js";
import type { WorkspaceSecretStore } from "../src/workspaces/secret-store.js";
const database = {
id: "11111111-1111-4111-8111-111111111111",
workspaceId: "psd-clinical",
engine: "postgres",
databaseName: "warehouse",
schema: 'clinical"data',
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
connectionStatus: "reachable",
binding: { transport: "postgres_direct", host: "db.internal", port: 5432, username: "reader" },
} satisfies WorkspaceDatabase;
const table = {
id: "22222222-2222-4222-8222-222222222222",
databaseId: database.id,
name: 'patient"facts',
sourceComment: null,
description: null,
generatedDescription: null,
lastSyncedDatabaseVersion: 1,
lastSyncedAt: "2026-09-02T08:00:00Z",
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
} satisfies CatalogTable;
function column(id: string, name: string): CatalogColumn {
return {
id,
tableId: table.id,
name,
ordinalPosition: 1,
dataType: "text",
isNullable: true,
defaultExpression: null,
primaryKeyPosition: null,
isPrimaryKey: false,
isForeignKey: false,
foreignKeyCount: 0,
sourceComment: null,
description: null,
generatedDescription: null,
sensitive: false,
lastSyncedDatabaseVersion: 1,
lastSyncedAt: "2026-09-02T08:00:00Z",
version: 1,
createdAt: "2026-09-02T08:00:00Z",
updatedAt: "2026-09-02T08:00:00Z",
};
}
function request(columns: readonly CatalogColumn[], overrides: Record<string, unknown> = {}) {
return {
database,
table,
columns,
valuesPerColumn: 300,
sampleOffset: 0,
sampleSeed: 37,
queryTimeoutMs: 5_000,
fullScanThreshold: 1_000,
...overrides,
};
}
test("uses bounded read-only PostgreSQL sampling for tables above 1,000 rows", async () => {
const note = column("33333333-3333-4333-8333-333333333333", "note");
const contact = column("44444444-4444-4444-8444-444444444444", 'contact"value');
const query = vi.fn(async (sql: string) => {
if (sql.startsWith("SELECT 1 AS __present")) {
return { rows: Array.from({ length: 1_001 }, () => ({ __present: 1 })) };
}
if (sql.startsWith("WITH sampled")) {
return { rows: [
{ __column_index: 0, __value: "ordinary", __length: "8" },
{ __column_index: 1, __value: "mario.rossi@example.it", __length: 23 },
] };
}
return { rows: [] };
});
const end = vi.fn(async () => undefined);
const access: CatalogPostgresAccess = {
connect: vi.fn(async () => ({ query, end }) as CatalogDatabaseClient),
};
const consume = vi.fn();
await expect(new ConcreteSensitivityValueSource(access).scanTable(
request([note, contact]),
consume,
new AbortController().signal,
)).resolves.toEqual({ kind: "sampled", observedValues: 2 });
expect(query.mock.calls[0]).toEqual(["BEGIN TRANSACTION READ ONLY", []]);
expect(query).toHaveBeenCalledWith("SELECT set_config('statement_timeout', $1, true)", ["5000ms"]);
const sampleSql = query.mock.calls.map(([sql]) => String(sql)).find((sql) => sql.startsWith("WITH sampled"));
expect(sampleSql).toContain('FROM "clinical""data"."patient""facts" TABLESAMPLE SYSTEM (30)');
expect(sampleSql).toContain("REPEATABLE (37)");
expect(sampleSql).toContain("LIMIT 3000 OFFSET 0");
expect(sampleSql).toContain("CROSS JOIN LATERAL");
expect(sampleSql).toContain("WHERE __rank <= 300");
expect(consume).toHaveBeenCalledWith([
{ columnId: note.id, value: "ordinary", characterLength: 8 },
{ columnId: contact.id, value: "mario.rossi@example.it", characterLength: 23 },
]);
expect(query.mock.calls.at(-1)).toEqual(["ROLLBACK", []]);
expect(end).toHaveBeenCalledOnce();
});
test("fully scans a table when the 1,001-row probe proves it is small", async () => {
const note = column("33333333-3333-4333-8333-333333333333", "note");
const query = vi.fn(async (sql: string) => {
if (sql.startsWith("SELECT 1 AS __present")) return { rows: [{ __present: 1 }] };
if (sql.startsWith("WITH sampled")) {
return { rows: [{ __column_index: 0, __value: "ordinary", __length: 8 }] };
}
return { rows: [] };
});
const access: CatalogPostgresAccess = {
connect: vi.fn(async () => ({ query, end: vi.fn(async () => undefined) }) as CatalogDatabaseClient),
};
const consume = vi.fn();
await expect(new ConcreteSensitivityValueSource(access).scanTable(
request([note]),
consume,
new AbortController().signal,
)).resolves.toEqual({ kind: "complete", observedValues: 1 });
const valueSql = query.mock.calls.map(([sql]) => String(sql)).find((sql) => sql.startsWith("WITH sampled"));
expect(valueSql).not.toContain("TABLESAMPLE");
expect(valueSql).toContain("WHERE __rank <= 1000");
expect(consume).toHaveBeenCalledWith([
{ columnId: note.id, value: "ordinary", characterLength: 8 },
]);
});
test("falls back to sampling when the small-table probe reaches its query timeout", async () => {
const note = column("33333333-3333-4333-8333-333333333333", "note");
const query = vi.fn(async (sql: string) => {
if (sql.startsWith("SELECT 1 AS __present")) {
throw Object.assign(new Error("statement timeout"), { code: "57014" });
}
if (sql.startsWith("WITH sampled")) {
return { rows: [{ __column_index: 0, __value: "sample", __length: 6 }] };
}
return { rows: [] };
});
const access: CatalogPostgresAccess = {
connect: vi.fn(async () => ({ query, end: vi.fn(async () => undefined) }) as CatalogDatabaseClient),
};
const consume = vi.fn();
await expect(new ConcreteSensitivityValueSource(access).scanTable(
request([note]),
consume,
new AbortController().signal,
)).resolves.toEqual({ kind: "sampled", observedValues: 1 });
expect(query.mock.calls.map(([sql]) => String(sql))).toContain(
"ROLLBACK TO SAVEPOINT sensitivity_scan_1",
);
});
test("limits each source query to at most 25 columns", async () => {
const columns = Array.from({ length: 26 }, (_, index) => column(
`00000000-0000-4000-8000-${(index + 1).toString().padStart(12, "0")}`,
`attribute_${index + 1}`,
));
const query = vi.fn(async (sql: string) => {
if (sql.startsWith("SELECT 1 AS __present")) {
return { rows: Array.from({ length: 1_001 }, () => ({ __present: 1 })) };
}
if (sql.startsWith("WITH sampled")) {
return { rows: [{ __column_index: 0, __value: "ordinary", __length: 8 }] };
}
return { rows: [] };
});
const access: CatalogPostgresAccess = {
connect: vi.fn(async () => ({ query, end: vi.fn(async () => undefined) }) as CatalogDatabaseClient),
};
await new ConcreteSensitivityValueSource(access).scanTable(
request(columns),
vi.fn(),
new AbortController().signal,
);
expect(query.mock.calls.filter(([sql]) => String(sql).startsWith("WITH sampled"))).toHaveLength(2);
});
test("scans a REST run_query binding without PostgreSQL-wire access", async () => {
const root = mkdtempSync(join(tmpdir(), "tht-sensitivity-rest-"));
const credentialFile = join(root, "api-key");
writeFileSync(credentialFile, "test-api-key\n", { mode: 0o600 });
const release = vi.fn();
const secretStore = {
materialize: vi.fn(() => ({
files: new Map([[CATALOG_SECRET_IDS.apiKey, credentialFile]]),
release,
})),
} as unknown as WorkspaceSecretStore;
const fetchMock = vi.fn(async () => new Response(JSON.stringify([
{ __column_index: 0, __value: "mario.rossi@example.it", __length: 23 },
]), { status: 200, headers: { "content-type": "application/json" } }));
vi.stubGlobal("fetch", fetchMock);
const access: CatalogPostgresAccess = {
connect: vi.fn(async () => { throw new Error("PostgreSQL access must not be used"); }),
};
const values = new ConcreteSensitivityValueSource(access, secretStore);
const restDatabase: WorkspaceDatabase = {
...database,
binding: {
transport: "rest_api",
baseUrl: "https://dwh.example.test/root/",
restPath: "/health",
restAuth: "x-api-key",
},
};
const note = column("33333333-3333-4333-8333-333333333333", "note");
const consume = vi.fn();
try {
await expect(values.scanTable(request([note], {
database: restDatabase,
fullScanThreshold: undefined,
}), consume, new AbortController().signal)).resolves.toEqual({
kind: "sampled",
observedValues: 1,
});
expect(access.connect).not.toHaveBeenCalled();
expect(fetchMock).toHaveBeenCalledWith(
"https://dwh.example.test/root/rpc/run_query",
expect.objectContaining({
method: "POST",
headers: { "content-type": "application/json", "x-api-key": "test-api-key" },
}),
);
const body = JSON.parse(String(fetchMock.mock.calls[0]![1]!.body));
expect(body.query_text).toContain('FROM "clinical""data"."patient""facts" TABLESAMPLE SYSTEM (30)');
expect(consume).toHaveBeenCalledWith([
{ columnId: note.id, value: "mario.rossi@example.it", characterLength: 23 },
]);
expect(release).toHaveBeenCalledOnce();
} finally {
vi.unstubAllGlobals();
rmSync(root, { recursive: true, force: true });
}
});
test("falls back to a sequential bounded sample when randomized sampling times out", async () => {
const note = column("33333333-3333-4333-8333-333333333333", "note");
const query = vi.fn(async (sql: string) => {
if (sql.startsWith("WITH sampled") && sql.includes("TABLESAMPLE")) {
throw Object.assign(new Error("raw source detail"), { code: "57014" });
}
if (sql.startsWith("WITH sampled")) {
return { rows: [{ __column_index: 0, __value: "ordinary", __length: 8 }] };
}
return { rows: [] };
});
const access: CatalogPostgresAccess = {
connect: vi.fn(async () => ({ query, end: vi.fn(async () => undefined) }) as CatalogDatabaseClient),
};
await expect(new ConcreteSensitivityValueSource(access).scanTable(
request([note], { fullScanThreshold: undefined }),
vi.fn(),
new AbortController().signal,
)).resolves.toEqual({ kind: "sampled", observedValues: 1 });
expect(query.mock.calls.filter(([sql]) => String(sql).startsWith("WITH sampled"))).toHaveLength(2);
});
test("fails explicitly when both randomized and sequential sample queries time out", async () => {
const note = column("33333333-3333-4333-8333-333333333333", "note");
const query = vi.fn(async (sql: string) => {
if (sql.startsWith("WITH sampled")) {
throw Object.assign(new Error("raw source detail"), { code: "57014" });
}
return { rows: [] };
});
const access: CatalogPostgresAccess = {
connect: vi.fn(async () => ({ query, end: vi.fn(async () => undefined) }) as CatalogDatabaseClient),
};
await expect(new ConcreteSensitivityValueSource(access).scanTable(
request([note], { fullScanThreshold: undefined }),
vi.fn(),
new AbortController().signal,
)).rejects.toEqual(new CatalogConnectorError("Sensitivity sample query timed out"));
});
+8 -9
View File
@@ -13,16 +13,11 @@ const roots: string[] = [];
afterEach(() => { for (const root of roots.splice(0)) rmSync(root, { recursive: true, force: true }); });
const workspace: WorkspaceDescriptor = {
workspace: { schema_version: 3, id: "psd-clinical", name: "Policlinico San Donato", language: "it" },
workspace: { schema_version: 4, id: "psd-clinical", name: "Policlinico San Donato", language: "it" },
dwh: {
engine: "postgres", database: "warehouse", schema: "datawarehouse", port: 5432,
supported_transports: ["postgres_direct", "rest_api"],
},
semantic_index: {
vector_store: { engine: "qdrant", collection: "psd", dimensions: 1024, distance: "cosine" },
embedding: { provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024 },
},
llm_policy: { allowed: ["zai/glm-5.2"] },
diagnostics: { dwh_rest: { method: "GET", path: "/health", auth: "bearer", response: { database: "database", schema: "schema" } } },
};
const revision: WorkspaceRevision = {
@@ -108,7 +103,7 @@ test("requires an exact deletion confirmation before applying the atomic diff",
]);
});
test("refuses synchronization until the current binding has passed its connection test", async () => {
test("starts synchronization without requiring a prior connection test", async () => {
const { app, repository, database } = await setup();
await repository.update(database.id, database.version, {
workspaceId: database.workspaceId,
@@ -122,6 +117,10 @@ test("refuses synchronization until the current binding has passed its connectio
url: `/catalog/databases/${database.id}/sync-runs`,
payload: { version: database.version + 1, scope: "tables", tableIds: [] },
});
expect(response.statusCode).toBe(409);
expect(response.json()).toMatchObject({ code: "schema_sync_conflict" });
expect(response.statusCode).toBe(202);
expect(response.json()).toMatchObject({
databaseId: database.id,
scope: "tables",
state: "queued",
});
});
+46
View File
@@ -71,6 +71,7 @@ test("loadConfig keeps local development defaults", () => {
workspaceSecretRuntimeRoot: "/tmp/thothii-workspace-secrets",
internalQdrantUrl: "http://qdrant:6333",
internalEmbeddingUrl: "http://embedding:11434",
internalEmbeddingId: "ollama/qwen3-embedding:0.6b",
internalEmbeddingModel: "qwen3-embedding:0.6b",
internalEmbeddingDimensions: 1024,
authMode: "none",
@@ -80,6 +81,35 @@ test("loadConfig keeps local development defaults", () => {
expect(loadConfig({}).dataRoot).toBeUndefined();
});
test("loadConfig keeps local NER disabled unless an absolute model path is configured", () => {
expect(loadConfig({}).sensitivityNer).toBeUndefined();
expect(loadConfig({
THT_SENSITIVITY_NER_MODEL_PATH: "/models/gliner2-pii",
}).sensitivityNer?.pythonExecutable).toBe("/opt/sensitivity-ner/bin/python");
expect(loadConfig({
THT_SENSITIVITY_NER_MODEL_PATH: "/models/gliner2-pii",
THT_SENSITIVITY_NER_PYTHON: "/opt/sensitivity-ner/bin/python",
THT_SENSITIVITY_NER_WORKER: "/app/backend/python/sensitivity_ner_worker.py",
THT_SENSITIVITY_NER_THREADS: "3",
}).sensitivityNer).toEqual({
modelPath: "/models/gliner2-pii",
pythonExecutable: "/opt/sensitivity-ner/bin/python",
workerScript: "/app/backend/python/sensitivity_ner_worker.py",
threads: 3,
});
});
test("loadConfig rejects ambiguous or unsafe local NER configuration", () => {
expect(() => loadConfig({ THT_SENSITIVITY_NER_MODEL_PATH: "fastino/model" }))
.toThrow("sensitivity NER model path configuration is invalid");
expect(() => loadConfig({
THT_SENSITIVITY_NER_MODEL_PATH: "/models/gliner2-pii",
THT_SENSITIVITY_NER_THREADS: "0",
})).toThrow("sensitivity NER thread configuration is invalid");
expect(() => loadConfig({ THT_SENSITIVITY_NER_PYTHON: "/opt/ner/bin/python" }))
.toThrow("sensitivity NER settings require a model path");
});
test("loadConfig allows none and mock only outside production when auth.yaml is absent", () => {
const originalNodeEnvironment = process.env.NODE_ENV;
delete process.env.NODE_ENV;
@@ -198,6 +228,22 @@ test("loadConfig accepts only the allowed internal semantic runtime hosts", () =
.toThrow(/internal.*embedding|invalid/i);
});
test("loadConfig derives the embedding runtime model from its canonical catalog identity", () => {
expect(loadConfig({
THT_INTERNAL_EMBEDDING_ID: "ollama/nomic-embed-text",
THT_INTERNAL_EMBEDDING_MODEL: "nomic-embed-text",
})).toMatchObject({
internalEmbeddingId: "ollama/nomic-embed-text",
internalEmbeddingModel: "nomic-embed-text",
});
expect(() => loadConfig({
THT_INTERNAL_EMBEDDING_ID: "ollama/nomic-embed-text",
THT_INTERNAL_EMBEDDING_MODEL: "different-model",
})).toThrow("does not match its canonical identity");
expect(() => loadConfig({ THT_INTERNAL_EMBEDDING_ID: "not-canonical" }))
.toThrow("embedding identity configuration is invalid");
});
test("loadConfig enables the legacy workspace request only through explicit local mode", () => {
expect(loadConfig({ THT_LEGACY_WORKSPACE_MODE: "local" }).legacyWorkspaceMode).toBe(true);
+6 -2
View File
@@ -11,6 +11,7 @@ import {
const semanticRuntime = {
internalQdrantUrl: "http://qdrant:6333",
internalEmbeddingUrl: "http://embedding:11434",
internalEmbeddingId: "ollama/qwen3-embedding:0.6b",
internalEmbeddingModel: "qwen3-embedding:0.6b",
internalEmbeddingDimensions: 1024,
};
@@ -102,11 +103,11 @@ function restRendered(): Record<string, unknown> {
}
const directCanonical =
`{"schemaVersion":1,"dwh":{` +
`{"schemaVersion":2,"dwh":{` +
`"engine":"postgres","database":"postgres","schema":"datawarehouse",` +
`"transport":"postgres_direct","host":"dwh.internal","port":5432,"user":"thoth_reader"},` +
`"vector":{"collection":"psd-clinical","dimensions":1024,"distance":"cosine"},` +
`"embedding":{"model":"qwen3-embedding:0.6b","dimensions":1024},` +
`"embedding":{"id":"ollama/qwen3-embedding:0.6b","model":"qwen3-embedding:0.6b","dimensions":1024},` +
`"roots":{"artifacts":"/data/sessions/psd-clinical/artifacts",` +
`"indexes":"/data/sessions/psd-clinical/indexes"}}`;
@@ -192,6 +193,9 @@ test("DWH-affecting changes alter the effective config identity", () => {
const changedCollection = { ...base, resources: { ...base.resources, vector: { ...(base.resources as Record<string, any>).vector, collection: "other" } } };
expect(effectiveConfigIdentity("psd-clinical", changedCollection)).not.toBe(identityBefore);
const changedEmbeddingIdentity = { ...base, resources: { ...base.resources, embeddings: { ...(base.resources as Record<string, any>).embeddings, model: "other-embedding" } } };
expect(effectiveConfigIdentity("psd-clinical", changedEmbeddingIdentity)).not.toBe(identityBefore);
const changedTransport = restRendered();
expect(effectiveConfigIdentity("psd-clinical", changedTransport)).not.toBe(identityBefore);
});
+31
View File
@@ -6,6 +6,7 @@ import { join } from "node:path";
import path from "node:path";
import { createPiModelLister } from "../src/pi/list-models.js";
import { loadConfig } from "../src/config.js";
import type { RuntimeModel, RuntimeModelCatalog } from "../src/models/runtime-model-catalog.js";
const FAKE = path.resolve("../harness/tests/fake_pi/fake_pi_rpc.mjs");
@@ -44,6 +45,36 @@ test("createPiModelLister returns mapped PiModel[] from get_available_models", a
}
});
test("catalog listing translates upstream Pi IDs back to canonical model keys", async () => {
const script = scriptWith([
{ provider: "local", id: "qwen2.5:7b", name: "Upstream label", reasoning: false },
]);
const model: RuntimeModel = {
id: "local/qwen", provider: "local", model: "qwen", label: "Catalog Qwen",
upstreamModel: "qwen2.5:7b", endpoint: { baseUrl: "http://ollama:11434/v1" },
authentication: { mode: "none" }, sessionAdapter: { mode: "openai_compatible" },
session: { reasoning: true, contextWindow: 32768, maxTokens: 8192 },
};
const modelCatalog: RuntimeModelCatalog = {
defaultSession: model.id, defaultMetadataGeneration: null, embedding: null,
sessionModels: () => [model], metadataModels: () => [], hasSession: (id) => id === model.id,
};
try {
const lister = createPiModelLister(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
...noManagedModels,
modelCatalog,
loadEnabledModels: enabled("local/qwen2.5:7b"),
spawnFn: () => spawn("node", [FAKE, script]) as any,
});
await expect(lister()).resolves.toEqual([{
provider: "local", id: "qwen", name: "Catalog Qwen", reasoning: true,
}]);
} finally {
rmSync(path.dirname(script), { recursive: true, force: true });
}
});
test("createPiModelLister caches within ttl (spawns once for two calls)", async () => {
const script = scriptWith([{ provider: "zai", id: "glm-5.2", name: "GLM 5.2", reasoning: true }]);
try {
+85 -247
View File
@@ -1,15 +1,12 @@
import { chmodSync, mkdtempSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { afterEach, expect, test, vi } from "vitest";
import { buildApp } from "../src/app.js";
import { MemoryCatalogRepository } from "../src/catalog/memory-repository.js";
import { afterEach, expect, test } from "vitest";
import {
loadMetadataGenerationModels,
MetadataGenerationModelUnavailableError,
} from "../src/catalog/metadata-generation-models.js";
import { loadConfig } from "../src/config.js";
import type { WorkspaceRegistry } from "../src/workspaces/registry.js";
import { loadRuntimeModelCatalog, splitCanonicalModelId } from "../src/models/runtime-model-catalog.js";
const roots: string[] = [];
@@ -17,260 +14,101 @@ afterEach(() => {
for (const root of roots.splice(0)) rmSync(root, { recursive: true, force: true });
});
function metadataConfiguration(
metadataGeneration: string,
secrets = "OPENAI_API_KEY=raw-provider-secret\n",
) {
const root = mkdtempSync(join(tmpdir(), "thothii-metadata-models-"));
function runtimeCatalog(overrides: Record<string, unknown> = {}, secrets = "OPENAI_API_KEY=raw-provider-secret\n") {
const root = mkdtempSync(join(tmpdir(), "thothii-runtime-models-"));
roots.push(root);
const installationFile = join(root, "thothii-installation.yaml");
const catalogFile = join(root, "catalog.json");
const secretsFile = join(root, "thothii.secrets");
writeFileSync(installationFile, metadataGeneration, { mode: 0o600 });
writeFileSync(secretsFile, secrets, { mode: 0o600 });
chmodSync(installationFile, 0o600);
chmodSync(secretsFile, 0o600);
return { installationFile, secretsFile };
}
function appFor(installationFile: string, secretsFile: string) {
const config = loadConfig({
NODE_ENV: "test",
THT_HARNESS_DIR: "/missing",
THT_INSTALLATION_CONFIG_FILE: installationFile,
THT_SECRETS_FILE: secretsFile,
PI_PROVIDER: "unrelated-pi-provider",
PI_MODEL: "unrelated-pi-model",
});
return buildApp(config, {
thtRunner: {} as never,
workspaceRegistry: { list: vi.fn(async () => []) } as unknown as WorkspaceRegistry,
workspaceDiagnoser: vi.fn(),
catalogRepository: new MemoryCatalogRepository(),
});
}
test("exposes only safe metadata-generation choices and their configured default", async () => {
const { installationFile, secretsFile } = metadataConfiguration(`metadataGeneration:
default: openai-mini
models:
- id: openai-mini
label: OpenAI Mini
litellm:
provider: openai
model: gpt-4.1-mini
endpoint:
baseUrl: https://api.openai.example/v1
apiVersion: "2026-08-01"
apiKeyEnv: OPENAI_API_KEY
`);
const app = appFor(installationFile, secretsFile);
const response = await app.inject({ method: "GET", url: "/catalog/metadata-generation/models" });
expect(response.statusCode).toBe(200);
expect(response.json()).toEqual({
models: [{ id: "openai-mini", label: "OpenAI Mini" }],
default: "openai-mini",
});
expect(response.body).not.toMatch(/openai\/gpt|gpt-4\.1|api\.openai|OPENAI_API_KEY|raw-provider-secret/);
await app.close();
});
test("rejects an unprotected installation descriptor", () => {
const { installationFile, secretsFile } = metadataConfiguration(`metadataGeneration:
default: openai-mini
models:
- id: openai-mini
label: OpenAI Mini
litellm: {provider: openai, model: gpt-4.1-mini}
apiKeyEnv: OPENAI_API_KEY
`);
chmodSync(installationFile, 0o644);
expect(() => loadMetadataGenerationModels({ installationFile, secretsFile }))
.toThrow("metadata-generation installation is unavailable");
});
test("returns an empty safe catalog when no metadata-generation model is configured", async () => {
const { installationFile, secretsFile } = metadataConfiguration("profile: local\n");
const app = appFor(installationFile, secretsFile);
const response = await app.inject({ method: "GET", url: "/catalog/metadata-generation/models" });
expect(response.statusCode).toBe(200);
expect(response.json()).toEqual({ models: [], default: null });
await app.close();
});
test("resolves only a configured selection for the later generation boundary", () => {
const { installationFile, secretsFile } = metadataConfiguration(`metadataGeneration:
default: openai-mini
models:
- id: openai-mini
label: OpenAI Mini
litellm:
provider: openai
model: gpt-4.1-mini
endpoint: {baseUrl: https://api.openai.example/v1, apiVersion: "2026-08-01"}
apiKeyEnv: OPENAI_API_KEY
`);
const models = loadMetadataGenerationModels({ installationFile, secretsFile });
expect(models.resolve("openai-mini")).toEqual({
id: "openai-mini",
provider: "openai",
model: "gpt-4.1-mini",
endpoint: { baseUrl: "https://api.openai.example/v1", apiVersion: "2026-08-01" },
apiKeyEnv: "OPENAI_API_KEY",
apiKey: "raw-provider-secret",
});
expect(() => models.resolve("unknown-model")).toThrow(MetadataGenerationModelUnavailableError);
});
test("loads DeepSeek models, GLM, and an explicit keyless Qwen endpoint from installation setup", () => {
const { installationFile, secretsFile } = metadataConfiguration(`metadataGeneration:
default: glm-53
models:
- id: deepseek-v4-pro
label: DeepSeek V4 Pro
litellm: {provider: deepseek, model: deepseek-v4-pro}
apiKeyEnv: DEEPSEEK_API_KEY
- id: deepseek-v4-flash
label: DeepSeek V4 Flash
litellm: {provider: deepseek, model: deepseek-v4-flash}
apiKeyEnv: DEEPSEEK_API_KEY
- id: glm-53
label: GLM 5.3
litellm:
provider: openai
model: glm-5.3
endpoint: {baseUrl: https://api.z.ai/api/coding/paas/v4}
apiKeyEnv: ZAI_API_KEY
- id: qwen-36
label: Qwen 3.6
litellm:
provider: openai
model: qwen3.6-35b-a3b
disableThinking: true
endpoint: {baseUrl: https://models.internal.example/v1}
`, "DEEPSEEK_API_KEY=deepseek-secret\nZAI_API_KEY=zai-secret\n");
const models = loadMetadataGenerationModels({ installationFile, secretsFile });
expect(models.catalog()).toEqual({
const catalog = {
schemaVersion: 1,
defaultSession: "zai/glm-5.3",
defaultMetadataGeneration: "zai/glm-5.3",
embedding: { id: "ollama/qwen3-embedding:0.6b", dimensions: 1024 },
models: [
{ id: "deepseek-v4-pro", label: "DeepSeek V4 Pro" },
{ id: "deepseek-v4-flash", label: "DeepSeek V4 Flash" },
{ id: "glm-53", label: "GLM 5.3" },
{ id: "qwen-36", label: "Qwen 3.6" },
{
id: "zai/glm-5.3", provider: "zai", model: "glm-5.3", label: "GLM 5.3",
upstreamModel: "glm-5.3", endpoint: { baseUrl: "https://api.z.ai/v1" },
authentication: { mode: "secret_env", apiKeyEnv: "OPENAI_API_KEY" },
sessionAdapter: { mode: "openai_compatible" },
metadataAdapter: { litellmProvider: "openai" },
session: { reasoning: true, contextWindow: 200000, maxTokens: 131072 },
metadataGeneration: { disableThinking: false },
},
{
id: "deepseek/deepseek-v4-pro", provider: "deepseek", model: "deepseek-v4-pro",
label: "DeepSeek V4 Pro", upstreamModel: "deepseek-v4-pro",
authentication: { mode: "pi_auth" }, sessionAdapter: { mode: "pi_builtin" },
session: { reasoning: false },
},
],
default: "glm-53",
...overrides,
};
writeFileSync(catalogFile, JSON.stringify(catalog), { mode: 0o600 });
writeFileSync(secretsFile, secrets, { mode: 0o600 });
chmodSync(catalogFile, 0o600);
chmodSync(secretsFile, 0o600);
return { catalogFile, secretsFile };
}
test("loads session default and safe metadata choices from the normalized runtime catalog", () => {
const { catalogFile, secretsFile } = runtimeCatalog();
const runtime = loadRuntimeModelCatalog(catalogFile);
const metadata = loadMetadataGenerationModels({ catalogFile, secretsFile });
expect(runtime.defaultSession).toBe("zai/glm-5.3");
expect(runtime.hasSession("deepseek/deepseek-v4-pro")).toBe(true);
expect(metadata.catalog()).toEqual({
models: [{ id: "zai/glm-5.3", label: "GLM 5.3" }],
default: "zai/glm-5.3",
});
expect(models.resolve("deepseek-v4-pro")).toMatchObject({
apiKeyEnv: "DEEPSEEK_API_KEY",
apiKey: "deepseek-secret",
});
expect(models.resolve("qwen-36")).toEqual({
id: "qwen-36",
provider: "openai",
model: "qwen3.6-35b-a3b",
disableThinking: true,
endpoint: { baseUrl: "https://models.internal.example/v1" },
expect(metadata.resolve("zai/glm-5.3")).toEqual({
id: "zai/glm-5.3", provider: "openai", model: "glm-5.3",
endpoint: { baseUrl: "https://api.z.ai/v1" },
apiKeyEnv: "OPENAI_API_KEY", apiKey: "raw-provider-secret",
});
expect(() => metadata.resolve("zai/missing")).toThrow(MetadataGenerationModelUnavailableError);
});
test("loads an explicit keyless endpoint without a secret bundle", () => {
const { installationFile } = metadataConfiguration(`metadataGeneration:
default: qwen-36
models:
- id: qwen-36
label: Qwen 3.6
litellm:
provider: openai
model: qwen3.6-35b-a3b
disableThinking: true
endpoint: {baseUrl: https://models.internal.example/v1}
`);
test("returns empty catalogs when no runtime projection is configured", () => {
expect(loadRuntimeModelCatalog().defaultSession).toBeNull();
expect(loadMetadataGenerationModels({}).catalog()).toEqual({ models: [], default: null });
});
expect(loadMetadataGenerationModels({ installationFile }).resolve("qwen-36")).toEqual({
id: "qwen-36",
provider: "openai",
model: "qwen3.6-35b-a3b",
disableThinking: true,
endpoint: { baseUrl: "https://models.internal.example/v1" },
test("rejects a drifted default and an unprotected projection", () => {
const drifted = runtimeCatalog({ defaultSession: "zai/missing" });
expect(() => loadRuntimeModelCatalog(drifted.catalogFile)).toThrow("session default is invalid");
const unprotected = runtimeCatalog();
chmodSync(unprotected.catalogFile, 0o666);
expect(() => loadRuntimeModelCatalog(unprotected.catalogFile)).toThrow("runtime model catalog is unavailable");
});
test("rejects authentication semantics that cannot come from the installation catalog", () => {
const invalid = runtimeCatalog({
defaultMetadataGeneration: undefined,
models: [{
id: "zai/glm-5.3",
provider: "zai",
model: "glm-5.3",
label: "GLM 5.3",
upstreamModel: "glm-5.3",
authentication: { mode: "secret_env" },
sessionAdapter: { mode: "pi_builtin" },
session: { reasoning: true },
}],
});
expect(() => loadRuntimeModelCatalog(invalid.catalogFile)).toThrow("runtime model catalog is invalid");
});
test.each([
["invalid YAML", "metadataGeneration: [\n", "OPENAI_API_KEY=secret\n", /invalid YAML/],
["duplicate ids", `metadataGeneration:
default: openai-mini
models:
- {id: openai-mini, label: One, litellm: {provider: openai, model: gpt-4.1-mini}, apiKeyEnv: OPENAI_API_KEY}
- {id: openai-mini, label: Two, litellm: {provider: openai, model: gpt-4.1}, apiKeyEnv: OPENAI_API_KEY}
`, "OPENAI_API_KEY=secret\n", /model id "openai-mini" is duplicated/],
["missing default", `metadataGeneration:
models:
- {id: openai-mini, label: One, litellm: {provider: openai, model: gpt-4.1-mini}, apiKeyEnv: OPENAI_API_KEY}
`, "OPENAI_API_KEY=secret\n", /default is required/],
["unknown default", `metadataGeneration:
default: absent
models:
- {id: openai-mini, label: One, litellm: {provider: openai, model: gpt-4.1-mini}, apiKeyEnv: OPENAI_API_KEY}
`, "OPENAI_API_KEY=secret\n", /default "absent" is not configured/],
["malformed settings", `metadataGeneration:
default: openai-mini
models:
- {id: openai-mini, label: One, litellm: {provider: "open ai", model: gpt-4.1-mini}, apiKeyEnv: OPENAI_API_KEY}
`, "OPENAI_API_KEY=secret\n", /configuration is invalid/],
["malformed endpoint", `metadataGeneration:
default: openai-mini
models:
- id: openai-mini
label: One
litellm: {provider: openai, model: gpt-4.1-mini, endpoint: {baseUrl: not-a-url}}
apiKeyEnv: OPENAI_API_KEY
`, "OPENAI_API_KEY=secret\n", /configuration is invalid/],
["keyless hosted model without endpoint", `metadataGeneration:
default: openai-mini
models:
- {id: openai-mini, label: One, litellm: {provider: openai, model: gpt-4.1-mini}}
`, "", /configuration is invalid/],
["disable thinking without endpoint", `metadataGeneration:
default: openai-mini
models:
- id: openai-mini
label: One
litellm: {provider: openai, model: gpt-4.1-mini, disableThinking: true}
apiKeyEnv: OPENAI_API_KEY
`, "OPENAI_API_KEY=secret\n", /configuration is invalid/],
["unallowed secret reference", `metadataGeneration:
default: openai-mini
models:
- {id: openai-mini, label: One, litellm: {provider: openai, model: gpt-4.1-mini}, apiKeyEnv: THT_DWH_API_KEY}
`, "THT_DWH_API_KEY=secret\n", /configuration is invalid/],
["missing referenced secret", `metadataGeneration:
default: openai-mini
models:
- {id: openai-mini, label: One, litellm: {provider: openai, model: gpt-4.1-mini}, apiKeyEnv: OPENAI_API_KEY}
`, "THT_DWH_API_KEY=secret\n", /secret "OPENAI_API_KEY" is missing/],
["unusable referenced secret", `metadataGeneration:
default: openai-mini
models:
- {id: openai-mini, label: One, litellm: {provider: openai, model: gpt-4.1-mini}, apiKeyEnv: OPENAI_API_KEY}
`, "OPENAI_API_KEY=secret with whitespace\n", /secret "OPENAI_API_KEY" is unusable/],
] as const)("rejects %s metadata-generation configuration", (_name, yaml, secrets, expected) => {
const { installationFile, secretsFile } = metadataConfiguration(yaml, secrets);
expect(() => loadMetadataGenerationModels({ installationFile, secretsFile })).toThrow(expected);
test("fails closed for missing or unusable provider secrets", () => {
const missing = runtimeCatalog({}, "THT_DWH_API_KEY=other\n");
expect(() => loadMetadataGenerationModels(missing)).toThrow('secret "OPENAI_API_KEY" is missing');
const unusable = runtimeCatalog({}, "OPENAI_API_KEY=contains whitespace\n");
expect(() => loadMetadataGenerationModels(unusable)).toThrow('secret "OPENAI_API_KEY" is unusable');
});
test("rejects a missing secret-bundle declaration for configured models", () => {
const { installationFile } = metadataConfiguration(`metadataGeneration:
default: openai-mini
models:
- {id: openai-mini, label: One, litellm: {provider: openai, model: gpt-4.1-mini}, apiKeyEnv: OPENAI_API_KEY}
`);
expect(() => loadMetadataGenerationModels({ installationFile }))
.toThrow("metadata-generation keyed models require THT_SECRETS_FILE");
test("splits canonical session identities without provider aliases", () => {
expect(splitCanonicalModelId("zai/glm-5.3")).toEqual({ provider: "zai", model: "glm-5.3" });
expect(() => splitCanonicalModelId("glm-5.3")).toThrow("model identity is invalid");
});
+18 -108
View File
@@ -1,13 +1,13 @@
import { mkdtempSync, readFileSync, readdirSync, rmSync } from "node:fs";
import { mkdtempSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { expect, test, vi } from "vitest";
import { loadConfig } from "../src/config.js";
import {
PiManagementError,
createPiManagement,
type PiExecFile,
} from "../src/pi/management.js";
import type { RuntimeModelCatalog } from "../src/models/runtime-model-catalog.js";
function configFor(settingsFile = join(mkdtempSync(join(tmpdir(), "tht-pi-management-")), "settings.json")) {
return loadConfig({
@@ -18,10 +18,14 @@ function configFor(settingsFile = join(mkdtempSync(join(tmpdir(), "tht-pi-manage
});
}
const supportedModels = [
{ provider: "zai", id: "glm-5.2", name: "GLM 5.2", reasoning: true },
{ provider: "deepseek", id: "deepseek-v4", name: "DeepSeek V4", reasoning: true },
];
const modelCatalog: RuntimeModelCatalog = {
defaultSession: "zai/glm-5.2",
defaultMetadataGeneration: null,
embedding: { id: "ollama/qwen3-embedding:0.6b", dimensions: 1024 },
sessionModels: () => [],
metadataModels: () => [],
hasSession: (id) => id === "zai/glm-5.2",
};
function successfulExec(calls: Array<{ command: string; args: string[]; timeout: number }>): PiExecFile {
return async (command, args, options) => {
@@ -36,7 +40,7 @@ test("status parses only a Pi version from a fixed execFile argument array", asy
const calls: Array<{ command: string; args: string[]; timeout: number }> = [];
const service = createPiManagement(configFor(), {
execute: successfulExec(calls),
listModels: async () => supportedModels,
modelCatalog,
readSettings: () => ({ provider: "zai", model: "glm-5.2", thinking: "medium" }),
credentialStatus: () => "missing",
now: () => new Date("2026-08-05T10:00:00.000Z"),
@@ -63,7 +67,7 @@ test.each(["present", "missing"] as const)(
const checkedProviders: Array<string | undefined> = [];
const service = createPiManagement(configFor(), {
execute: successfulExec([]),
listModels: async () => supportedModels,
modelCatalog,
readSettings: () => ({ provider: "zai", model: "glm-5.2", thinking: "medium" }),
credentialStatus: (provider) => {
checkedProviders.push(provider);
@@ -85,100 +89,6 @@ test.each(["present", "missing"] as const)(
},
);
// Catches an options response that leaks provider metadata or lets callers choose model IDs that
// Pi did not explicitly enable for this installation.
test("options expose only closed provider, model, and reasoning choices", async () => {
const service = createPiManagement(configFor(), {
execute: successfulExec([]),
listModels: async () => supportedModels,
now: () => new Date("2026-08-05T10:00:00.000Z"),
});
await expect(service.options()).resolves.toEqual({
providers: ["zai", "deepseek"],
models: [
{ provider: "zai", id: "glm-5.2" },
{ provider: "deepseek", id: "deepseek-v4" },
],
reasoning: ["low", "medium", "high"],
checkedAt: "2026-08-05T10:00:00.000Z",
});
});
// Catches raw managed models.json validation details being collapsed into an ambiguous model-list
// failure or escaping through the Pi Management options API.
test("options report invalid managed model configuration with a stable sanitized error", async () => {
const service = createPiManagement(configFor(), {
execute: successfulExec([]),
listModels: async () => {
throw Object.assign(
new Error("!sensitive-command /private/models.json raw-secret"),
{ code: "PI_MANAGED_CONFIG_INVALID" },
);
},
});
let caught: unknown;
try {
await service.options();
} catch (error) {
caught = error;
}
expect(caught).toMatchObject<PiManagementError>({
code: "pi_management_unavailable",
message: "Pi provider/model configuration is invalid",
});
expect(String(caught)).not.toMatch(/sensitive|private|models\.json|secret/i);
});
// Catches configuration writes that accept whitespace, unknown choices, or extra free-form fields
// before reaching the durable installation settings file.
test("config rejects invalid free-form values before writing settings", async () => {
const directory = mkdtempSync(join(tmpdir(), "tht-pi-management-invalid-"));
try {
let writes = 0;
const service = createPiManagement(configFor(join(directory, "settings.json")), {
execute: successfulExec([]),
listModels: async () => supportedModels,
readSettings: () => ({}),
saveSettings: () => { writes += 1; return {}; },
});
await expect(service.configure({
provider: "zai ", model: "glm-5.2", reasoning: "medium", unexpected: "value",
} as any)).rejects.toMatchObject<PiManagementError>({ code: "pi_management_invalid_config" });
expect(writes).toBe(0);
} finally {
rmSync(directory, { recursive: true, force: true });
}
});
// Catches a non-atomic implementation that can leave partial settings or temporary files after a
// normal installation-default update.
test("config validates closed choices and atomically persists non-secret defaults", async () => {
const directory = mkdtempSync(join(tmpdir(), "tht-pi-management-write-"));
const settingsFile = join(directory, "settings.json");
try {
const service = createPiManagement(configFor(settingsFile), {
execute: successfulExec([]),
listModels: async () => supportedModels,
now: () => new Date("2026-08-05T10:00:00.000Z"),
});
await expect(service.configure({
provider: "zai", model: "glm-5.2", reasoning: "high",
})).resolves.toEqual({
provider: "zai", model: "glm-5.2", reasoning: "high", updatedAt: "2026-08-05T10:00:00.000Z",
});
expect(JSON.parse(readFileSync(settingsFile, "utf8"))).toEqual({
provider: "zai", model: "glm-5.2", thinking: "high",
});
expect(readdirSync(directory)).toEqual(["settings.json"]);
} finally {
rmSync(directory, { recursive: true, force: true });
}
});
// Catches a hung Pi smoke check that leaves an operator waiting indefinitely or returns raw child
// diagnostics containing provider credentials.
test("smoke uses the configured timeout and reports a sanitized timeout", async () => {
@@ -188,7 +98,7 @@ test("smoke uses the configured timeout and reports a sanitized timeout", async
calls.push({ command, args, timeout: options.timeout });
throw Object.assign(new Error("provider token=raw-provider-token"), { code: "ETIMEDOUT" });
},
listModels: async () => supportedModels,
modelCatalog,
readSettings: () => ({ provider: "zai", model: "glm-5.2", thinking: "medium" }),
now: () => new Date("2026-08-05T10:00:00.000Z"),
});
@@ -210,7 +120,7 @@ test("smoke exercises the configured provider and model", async () => {
const providerChecks: unknown[] = [];
const service = createPiManagement(configFor(), {
execute: successfulExec([]),
listModels: async () => supportedModels,
modelCatalog,
readSettings: () => ({ provider: "zai", model: "glm-5.2", thinking: "medium" }),
smokeProvider: async (request) => { providerChecks.push(request); },
now: () => new Date("2026-08-05T10:00:00.000Z"),
@@ -230,7 +140,7 @@ test("smoke exercises the configured provider and model", async () => {
test("smoke fails closed and sanitizes configured-provider authentication errors", async () => {
const service = createPiManagement(configFor(), {
execute: successfulExec([]),
listModels: async () => supportedModels,
modelCatalog,
readSettings: () => ({ provider: "zai", model: "glm-5.2", thinking: "medium" }),
smokeProvider: async () => {
throw new Error('401 {"token":"raw-expired-token","output":"raw-provider-output"}');
@@ -252,7 +162,7 @@ test("smoke fails closed and sanitizes configured-provider authentication errors
test("smoke reports invalid managed provider configuration with a stable sanitized error", async () => {
const service = createPiManagement(configFor(), {
execute: successfulExec([]),
listModels: async () => supportedModels,
modelCatalog,
readSettings: () => ({ provider: "zai", model: "glm-5.2", thinking: "medium" }),
smokeProvider: async () => {
throw Object.assign(
@@ -284,7 +194,7 @@ test("smoke applies one deadline across version and a hung provider turn", async
() => resolve({ stdout: "pi 0.80.3\n", stderr: "" }),
500,
)),
listModels: async () => supportedModels,
modelCatalog,
readSettings: () => ({ provider: "zai", model: "glm-5.2", thinking: "medium" }),
smokeProvider: async ({ timeoutMs }) => {
providerTimeouts.push(timeoutMs);
@@ -320,7 +230,7 @@ test("logs keep only the latest 200 redacted lines", async () => {
source[201] = "THT_MODEL_API_KEY=raw-env-secret";
const service = createPiManagement(configFor(), {
execute: successfulExec([]),
listModels: async () => supportedModels,
modelCatalog,
readLogs: () => source.join("\n"),
now: () => new Date("2026-08-05T10:00:00.000Z"),
});
+96
View File
@@ -8,6 +8,7 @@ import {
} from "node:fs";
import { tmpdir } from "node:os";
import { PiProcessManager } from "../src/pi/pi-process-manager.js";
import type { RuntimeModelCatalog, RuntimeModel } from "../src/models/runtime-model-catalog.js";
import { loadConfig } from "../src/config.js";
import {
PI_MANAGED_CONFIG_ERROR_MESSAGE,
@@ -641,6 +642,53 @@ test("session Pi spawn reads the single secret bundle and scrubs its path", asyn
}
});
test("session Pi spawn resolves the selected catalog credential from the secret bundle", () => {
const root = mkdtempSync(path.join(tmpdir(), "thothii-catalog-credential-"));
const agentDir = path.join(root, "agent");
mkdirSync(agentDir, { mode: 0o700 });
writeFileSync(path.join(agentDir, "auth.json"), "{}\n", { mode: 0o600 });
writeFileSync(path.join(agentDir, "models.json"), '{"providers":{}}\n', { mode: 0o600 });
const secret = path.join(root, "thothii.secrets");
writeFileSync(secret, "ZAI_API_KEY=catalog-secret\nTHT_MODEL_API_KEY=legacy-secret\n", { mode: 0o600 });
const model: RuntimeModel = {
id: "openai/test-model",
provider: "openai",
model: "test-model",
label: "Test model",
upstreamModel: "test-model",
authentication: { mode: "secret_env", apiKeyEnv: "ZAI_API_KEY" },
sessionAdapter: { mode: "pi_builtin" },
session: { reasoning: false },
};
const modelCatalog: RuntimeModelCatalog = {
defaultSession: model.id,
defaultMetadataGeneration: null,
embedding: null,
sessionModels: () => [model],
metadataModels: () => [],
hasSession: (id) => id === model.id,
};
const calls: any[][] = [];
const child = recordingChild();
child.stderr.resume = () => {};
vi.stubEnv("PI_CODING_AGENT_DIR", agentDir);
const mgr = new PiProcessManager(loadConfig({ THT_SECRETS_FILE: secret }), {
modelCatalog,
authProviders: () => new Set(),
spawnFn: (...args: any[]) => { calls.push(args); return child as any; },
});
try {
mgr.createFor("catalog-credential", { provider: "openai", model: "test-model" });
expect(calls[0][2].env.ZAI_API_KEY).toBe("catalog-secret");
expect(calls[0][2].env).not.toHaveProperty("OPENAI_API_KEY");
expect(calls[0][2].env).not.toHaveProperty("THT_MODEL_API_KEY");
} finally {
mgr.teardown("catalog-credential");
vi.unstubAllEnvs();
rmSync(root, { recursive: true, force: true });
}
});
test.each([["OpenAI", "openai"], ["gemini", "google"]])(
"set_model uses canonical packaged provider ID for %s", async (provider, canonical) => {
const secret = path.resolve(__dirname, `.canonical-key-${process.pid}-${provider}`);
@@ -670,6 +718,54 @@ test.each([["OpenAI", "openai"], ["gemini", "google"]])(
},
);
test("set_model translates a canonical catalog key to its upstream Pi model ID", async () => {
const root = mkdtempSync(path.join(tmpdir(), "thothii-upstream-model-"));
const agentDir = path.join(root, "agent");
mkdirSync(agentDir, { mode: 0o700 });
writeFileSync(path.join(agentDir, "models.json"), JSON.stringify({
providers: {
local: {
baseUrl: "http://ollama:11434/v1", apiKey: "local",
models: [{ id: "qwen2.5:7b" }],
},
},
}), { mode: 0o600 });
vi.stubEnv("PI_CODING_AGENT_DIR", agentDir);
const child = recordingChild();
child.stderr.resume = () => {};
child.stdin.write = (data: unknown) => {
const request = JSON.parse(String(data));
child._writes.push(String(data));
if (request.id) {
queueMicrotask(() => child.stdout.emit("data", `${JSON.stringify({
type: "response", id: request.id, success: true,
})}\n`));
}
return true;
};
const model: RuntimeModel = {
id: "local/qwen", provider: "local", model: "qwen", label: "Qwen",
upstreamModel: "qwen2.5:7b", endpoint: { baseUrl: "http://ollama:11434/v1" },
authentication: { mode: "none" }, sessionAdapter: { mode: "openai_compatible" },
session: { reasoning: false, contextWindow: 32768, maxTokens: 8192 },
};
const modelCatalog: RuntimeModelCatalog = {
defaultSession: model.id, defaultMetadataGeneration: null, embedding: null,
sessionModels: () => [model], metadataModels: () => [], hasSession: (id) => id === model.id,
};
const mgr = new PiProcessManager(loadConfig({ PI_BIN: "/usr/local/bin/pi" }), {
modelCatalog, authProviders: () => new Set(), spawnFn: () => child as any,
});
try {
await mgr.spawnFor("upstream-model", { provider: "local", model: "qwen" });
expect(child._writes.join("")).toContain('"modelId":"qwen2.5:7b"');
} finally {
mgr.teardown("upstream-model");
vi.unstubAllEnvs();
rmSync(root, { recursive: true, force: true });
}
});
test.each(["installation-local", "private-compatible"])(
"provider %s configured with a literal apiKey spawns without a managed key",
async (provider) => {
+60 -3
View File
@@ -1,9 +1,11 @@
import { EventEmitter } from "node:events";
import { existsSync, readFileSync, readdirSync } from "node:fs";
import { dirname } from "node:path";
import { existsSync, mkdtempSync, readFileSync, readdirSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { dirname, join } from "node:path";
import { afterEach, expect, test, vi } from "vitest";
import { loadConfig } from "../src/config.js";
import { createPiProviderSmoke } from "../src/pi/provider-smoke.js";
import type { RuntimeModel, RuntimeModelCatalog } from "../src/models/runtime-model-catalog.js";
afterEach(() => vi.unstubAllEnvs());
@@ -44,6 +46,50 @@ function successfulProviderChild() {
const MANAGED_CONFIG_ERROR = "Pi provider/model configuration is invalid";
test("provider smoke resolves the selected catalog credential from the secret bundle", async () => {
const root = mkdtempSync(join(tmpdir(), "thothii-smoke-catalog-credential-"));
const secret = join(root, "thothii.secrets");
writeFileSync(secret, "ZAI_API_KEY=catalog-secret\nTHT_MODEL_API_KEY=legacy-secret\n", { mode: 0o600 });
const model: RuntimeModel = {
id: "openai/test-model",
provider: "openai",
model: "test-model",
label: "Test model",
upstreamModel: "test-model",
authentication: { mode: "secret_env", apiKeyEnv: "ZAI_API_KEY" },
sessionAdapter: { mode: "pi_builtin" },
session: { reasoning: false },
};
const modelCatalog: RuntimeModelCatalog = {
defaultSession: model.id,
defaultMetadataGeneration: null,
embedding: null,
sessionModels: () => [model],
metadataModels: () => [],
hasSession: (id) => id === model.id,
};
let spawnEnv: NodeJS.ProcessEnv | undefined;
const smoke = createPiProviderSmoke(loadConfig({ THT_SECRETS_FILE: secret }), {
modelCatalog,
authProviders: () => new Set(),
readModelsStore: () => undefined,
spawnFn: (_command, _args, options) => {
spawnEnv = options.env;
return successfulProviderChild();
},
});
try {
await expect(smoke({
provider: "openai", model: "test-model", reasoning: "medium", timeoutMs: 750,
})).resolves.toBeUndefined();
expect(spawnEnv?.ZAI_API_KEY).toBe("catalog-secret");
expect(spawnEnv).not.toHaveProperty("OPENAI_API_KEY");
expect(spawnEnv).not.toHaveProperty("THT_MODEL_API_KEY");
} finally {
rmSync(root, { recursive: true, force: true });
}
});
// Catches an isolated smoke agent that copies auth.json but drops the selected custom
// provider/model from models.json, causing set_model to fail before the real request.
test("provider smoke reaches the selected custom provider from an isolated models.json", async () => {
@@ -254,11 +300,22 @@ test("provider smoke makes one configured request from an isolated no-capability
});
}
});
const smokeModel: RuntimeModel = {
id: "zai/catalog-glm", provider: "zai", model: "catalog-glm", label: "GLM",
upstreamModel: "glm-5.2", authentication: { mode: "pi_auth" },
sessionAdapter: { mode: "pi_builtin" }, session: { reasoning: true },
};
const smokeCatalog: RuntimeModelCatalog = {
defaultSession: smokeModel.id, defaultMetadataGeneration: null, embedding: null,
sessionModels: () => [smokeModel], metadataModels: () => [],
hasSession: (id) => id === smokeModel.id,
};
const smoke = createPiProviderSmoke(loadConfig({
THT_HARNESS_DIR: "/app/harness",
PI_BIN: "/usr/local/bin/pi",
THT_DATA_ROOT: "/mounted-workflow-state",
}), {
modelCatalog: smokeCatalog,
spawnFn: (...args) => {
spawns.push(args);
expect(args[2].cwd).not.toBe("/app/harness");
@@ -278,7 +335,7 @@ test("provider smoke makes one configured request from an isolated no-capability
});
await expect(smoke({
provider: "zai", model: "glm-5.2", reasoning: "medium", timeoutMs: 750,
provider: "zai", model: "catalog-glm", reasoning: "medium", timeoutMs: 750,
})).resolves.toBeUndefined();
expect(spawns).toHaveLength(1);
expect(spawns[0][0]).toBe("/usr/local/bin/pi");
+1 -8
View File
@@ -2,18 +2,11 @@ import { expect, test } from "vitest";
import { ReadinessManager } from "../src/runtime/readiness-manager.js";
const workspace = {
workspace: { schema_version: 3, id: "psd", name: "PSD", language: "it" },
workspace: { schema_version: 4, id: "psd", name: "PSD", language: "it" },
dwh: {
engine: "postgres", database: "warehouse", schema: "public",
supported_transports: ["postgres_direct"],
},
semantic_index: {
vector_store: { engine: "qdrant", collection: "psd", dimensions: 1024, distance: "cosine" },
embedding: {
provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024,
},
},
llm_policy: { allowed: ["zai/glm-5.2"] },
} as const;
function deferred<T>() {
+1 -13
View File
@@ -15,7 +15,7 @@ afterEach(() => {
});
const validYaml = `workspace:
schema_version: 3
schema_version: 4
id: psd-clinical
name: Policlinico San Donato
language: it
@@ -24,18 +24,6 @@ dwh:
database: postgres
schema: datawarehouse
supported_transports: [postgres_direct]
semantic_index:
vector_store:
engine: qdrant
collection: psd-clinical
dimensions: 1024
distance: cosine
embedding:
provider: ollama_internal
model: qwen3-embedding:0.6b
dimensions: 1024
llm_policy:
allowed: [zai/glm-5.2]
`;
async function git(cwd: string, args: string[]): Promise<string> {
+1 -13
View File
@@ -20,7 +20,7 @@ async function git(cwd: string, args: string[]): Promise<string> {
}
const descriptor = `workspace:
schema_version: 3
schema_version: 4
id: research
name: Research
language: en
@@ -29,18 +29,6 @@ dwh:
database: analytics
schema: mart
supported_transports: [postgres_direct]
semantic_index:
vector_store:
engine: qdrant
collection: research
dimensions: 1024
distance: cosine
embedding:
provider: ollama_internal
model: qwen3-embedding:0.6b
dimensions: 1024
llm_policy:
allowed: [zai/glm-5.2]
evidence:
source:
type: filesystem
+9 -27
View File
@@ -14,11 +14,6 @@ function fakeService(): PiManagementService {
config: { provider: "zai", model: "glm-5.2", reasoning: "medium" },
checkedAt: "2026-08-05T10:00:00.000Z",
})),
options: vi.fn(async () => ({
providers: ["zai"], models: [{ provider: "zai", id: "glm-5.2" }],
reasoning: ["low", "medium", "high"], checkedAt: "2026-08-05T10:00:00.000Z",
})),
configure: vi.fn(async (value) => ({ ...value, updatedAt: "2026-08-05T10:00:00.000Z" })),
test: vi.fn(async () => ({ ready: true, checkedAt: "2026-08-05T10:00:00.000Z" })),
logs: vi.fn(async () => ({ lines: ["Pi smoke check succeeded"], checkedAt: "2026-08-05T10:00:00.000Z" })),
};
@@ -106,24 +101,17 @@ test("loopback-only AUTH_MODE=none may read the sanitized Pi status", async () =
});
// A local implicit administrator has pi.manage, but a browser origin still cannot borrow that
// authority to mutate local configuration or trigger provider work.
// authority to trigger provider work.
test("loopback-only management rejects cross-origin writes for its local administrator", async () => {
const service = fakeService();
const app = appWith(service);
try {
const configured = await app.inject({
method: "PUT", url: "/pi-management/config",
headers: { host: "127.0.0.1:8080", origin: "https://evil.example" },
payload: { provider: "zai", model: "glm-5.2", reasoning: "high" },
});
const smoke = await app.inject({
method: "POST", url: "/pi-management/test",
headers: { host: "127.0.0.1:8080", origin: "https://evil.example" },
});
expect(configured.statusCode).toBe(403);
expect(smoke.statusCode).toBe(403);
expect(service.configure).not.toHaveBeenCalled();
expect(service.test).not.toHaveBeenCalled();
} finally {
await app.close();
@@ -132,14 +120,13 @@ test("loopback-only management rejects cross-origin writes for its local adminis
// Catches an origin guard that also blocks the same-origin Docker frontend or non-browser local
// lifecycle clients that do not send Origin.
test("loopback-only management preserves same-origin frontend and origin-less local writes", async () => {
test("loopback-only management preserves same-origin and origin-less smoke checks", async () => {
const service = fakeService();
const app = appWith(service);
try {
const sameOrigin = await app.inject({
method: "PUT", url: "/pi-management/config",
method: "POST", url: "/pi-management/test",
headers: { host: "127.0.0.1:8080", origin: "http://127.0.0.1:8080" },
payload: { provider: "zai", model: "glm-5.2", reasoning: "high" },
});
const lifecycleClient = await app.inject({ method: "POST", url: "/pi-management/test" });
@@ -150,25 +137,20 @@ test("loopback-only management preserves same-origin frontend and origin-less lo
}
});
// Catches route wiring that bypasses closed service validation or gives the browser a Docker/image
// lifecycle endpoint rather than only installation-default configuration and diagnostics.
test("trusted admins receive only configuration, smoke, options, and log endpoints", async () => {
// Catches a regression that reintroduces a browser-writable provider/model source.
test("trusted admins receive only status, smoke, and log endpoints", async () => {
const service = fakeService();
const app = appWith(service, exposedServerEnv);
try {
const options = await app.inject({ method: "GET", url: "/pi-management/options", headers: adminHeaders });
const configured = await app.inject({
method: "PUT", url: "/pi-management/config", headers: adminHeaders,
payload: { provider: "zai", model: "glm-5.2", reasoning: "high" },
});
const status = await app.inject({ method: "GET", url: "/pi-management/status", headers: adminHeaders });
const smoke = await app.inject({ method: "POST", url: "/pi-management/test", headers: adminHeaders });
const logs = await app.inject({ method: "GET", url: "/pi-management/logs", headers: adminHeaders });
expect(options.statusCode).toBe(200);
expect(configured.statusCode).toBe(200);
expect(configured.json()).toMatchObject({ provider: "zai", model: "glm-5.2", reasoning: "high" });
expect(status.statusCode).toBe(200);
expect(smoke.statusCode).toBe(200);
expect(logs.statusCode).toBe(200);
expect((await app.inject({ method: "GET", url: "/pi-management/options", headers: adminHeaders })).statusCode).toBe(404);
expect((await app.inject({ method: "PUT", url: "/pi-management/config", headers: adminHeaders })).statusCode).toBe(404);
expect(app.printRoutes()).not.toContain("update");
expect(app.printRoutes()).not.toContain("rollback");
} finally {
+45 -44
View File
@@ -18,25 +18,25 @@ const SCRIPT = path.resolve("../harness/tests/fake_pi/scripts/f1_disambiguation.
function operationalWorkspace(id = "default") {
return {
workspace: { schema_version: 3, id, name: id, language: "en" },
workspace: { schema_version: 4, id, name: id, language: "en" },
dwh: {
engine: "postgres", database: "warehouse", schema: "public",
supported_transports: ["postgres_direct"],
},
semantic_index: {
vector_store: {
engine: "qdrant", collection: id, dimensions: 1024, distance: "cosine",
},
embedding: {
provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024,
},
},
llm_policy: {
allowed: ["zai/glm-5.2", "deepseek/deepseek-v4-pro", "local-qwen/qwen3.6-35b-a3b"],
},
} as const;
}
function sessionCatalog(defaultSession = "zai/glm-5.2", available = [defaultSession]) {
return {
defaultSession,
defaultMetadataGeneration: null,
embedding: { id: "ollama/qwen3-embedding:0.6b", dimensions: 1024 },
sessionModels: () => [],
metadataModels: () => [],
hasSession: (id: string) => available.includes(id),
} as any;
}
const defaultWorkspaceRegistry = {
list: vi.fn(async () => [{
id: "default", commit: "e".repeat(40), blob: "f".repeat(40),
@@ -495,8 +495,8 @@ test("creates a session from the active immutable workspace revision", async ()
workspaceRegistry: {
read: vi.fn(async () => ({
workspace: {
workspace: { schema_version: 2, id: "psd-clinical", name: "PSD", language: "it" },
dwh: {}, semantic_index: {}, llm_policy: { allowed: ["zai/glm-5.2"] },
workspace: { schema_version: 4, id: "psd-clinical", name: "PSD", language: "it" },
dwh: {},
},
revision: {
id: "psd-clinical", commit: "a".repeat(40), blob: "b".repeat(40),
@@ -589,22 +589,11 @@ test("rejects an SSH-only workspace before persisting or starting a session", as
workspaceRegistry: {
acquireSessionRevision: vi.fn(async () => ({
workspace: {
workspace: { schema_version: 2, id: "ssh-workspace", name: "SSH", language: "en" },
workspace: { schema_version: 4, id: "ssh-workspace", name: "SSH", language: "en" },
dwh: {
engine: "postgres", database: "postgres", schema: "public",
supported_transports: ["ssh_tunnel"],
},
semantic_index: {
vector_store: {
engine: "pgvector", database: "postgres", schema: "vectors",
collection: "documents", dimensions: 768, distance: "cosine",
supported_transports: ["ssh_tunnel"],
},
embedding: {
provider: "ollama_compatible", model: "nomic-embed-text", dimensions: 768,
},
},
llm_policy: { allowed: ["zai/glm-5.2"] },
},
revision: {
id: "ssh-workspace", commit: "a".repeat(40), blob: "b".repeat(40),
@@ -632,7 +621,7 @@ test("hands a revision lease to retention only after the session manifest is dur
const markPersisted = vi.fn(async () => {});
const abort = vi.fn(async () => {});
const acquireSessionRevision = vi.fn(async () => ({
workspace: { llm_policy: { allowed: ["zai/glm-5.2"] } },
workspace: operationalWorkspace("leased"),
revision: {
id: "leased", commit: "a".repeat(40), blob: "b".repeat(40),
snapshotPath: `/data/workspace-registry/snapshots/${"a".repeat(40)}/leased.yaml`,
@@ -670,7 +659,7 @@ test("creates a session from the configured default workspace revision when work
const sessionNew = vi.fn(async () => ({ id: "default-pinned" }));
const registry = {
read: vi.fn(async (id: string) => ({
workspace: { llm_policy: { allowed: ["zai/glm-5.2"] } },
workspace: operationalWorkspace(id),
revision: {
id, commit: "c".repeat(40), blob: "d".repeat(40),
snapshotPath: `/data/workspace-registry/snapshots/${"c".repeat(40)}/${id}.yaml`,
@@ -758,7 +747,7 @@ test("session lifecycle locates a B session when installation default is A", asy
listModels: async () => [{ provider: "zai", id: "glm-5.2", name: "GLM 5.2", reasoning: true }],
workspaceRegistry: {
read: async (id: string) => ({
workspace: { llm_policy: { allowed: ["zai/glm-5.2"] } },
workspace: operationalWorkspace(id),
revision: { id, commit: "b".repeat(40), blob: "d".repeat(40), snapshotPath: bPath },
}),
list: async () => [
@@ -793,7 +782,7 @@ test("session lifecycle locates a B session when installation default is A", asy
expect(runtimeOptions).toEqual(["/runtime/1.yaml", "/runtime/2.yaml"]);
});
test("POST /sessions usa i settings (workspace/provider/model/thinking) e crea+avvia", async () => {
test("POST /sessions uses the catalog default with workspace/thinking settings and starts", async () => {
const modelKey = path.join(os.tmpdir(), `thoth-model-key-${process.pid}`);
writeFileSync(modelKey, "test-model-key", { mode: 0o600 });
chmodSync(modelKey, 0o600);
@@ -808,7 +797,8 @@ test("POST /sessions usa i settings (workspace/provider/model/thinking) e crea+a
sessionNew: async (o: any) => { sessionNewArg = o; return { id: "s1" }; },
sessionList: async () => [{ id: "s1" }],
} as any,
getSettings: () => ({ workspace: "w", provider: "zai", model: "glm-5.2", thinking: "high" }),
getSettings: () => ({ workspace: "w", thinking: "high" }),
runtimeModelCatalog: sessionCatalog(),
listModels: async () => [
{ provider: "zai", id: "glm-5.2", name: "GLM 5.2", reasoning: true },
],
@@ -2565,32 +2555,43 @@ test("POST /sessions proceeds when ollamaEnsure succeeds", async () => {
expect(ensureWs).toContain(`/snapshots/${"e".repeat(40)}/psd.yaml`);
});
test("POST /sessions rejects an unavailable saved model before persisting a session", async () => {
test("POST /sessions falls back from a stale requested model to the catalog default", async () => {
let created = 0;
let persisted: any;
const runtime = { bridge: { onClientEvent: () => {} } };
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: {
sessionNew: async () => { created += 1; return { id: "must-not-exist" }; },
sessionNew: async (options: any) => { created += 1; persisted = options; return { id: "fallback" }; },
searchPack: async () => {},
} as any,
readiness: { ensure: async () => ({ ok: true }) } as any,
getSettings: () => ({
workspace: "psd",
provider: "deepseek",
model: "deepseek-v4-pro",
thinking: "medium",
}) as any,
getSettings: () => ({ workspace: "psd", thinking: "medium" }) as any,
runtimeModelCatalog: sessionCatalog(),
listModels: async () => [
{ provider: "zai", id: "glm-5.2", name: "GLM 5.2", reasoning: true },
],
mgr: {
teardownForPrincipal: () => [],
createFor: () => runtime,
get: () => runtime,
configure: async () => {},
start: () => {},
} as any,
});
const res = await app.inject({ method: "POST", url: "/sessions", payload: { question: "q" } });
const res = await app.inject({
method: "POST",
url: "/sessions",
payload: { question: "q", provider: "deepseek", model: "deepseek-v4-pro" },
});
expect(res.statusCode).toBe(503);
expect(res.statusCode).toBe(200);
expect(res.json()).toEqual({
error: "Selected model is unavailable. Check Pi authentication and model settings, then try again.",
code: "model_unavailable",
id: "fallback",
warning: "Configured model deepseek/deepseek-v4-pro is unavailable; using zai/glm-5.2.",
});
expect(created).toBe(0);
expect(persisted).toMatchObject({ provider: "zai", model: "glm-5.2" });
expect(created).toBe(1);
});
test("POST /sessions marks a persisted session failed when runtime construction throws", async () => {
+16 -16
View File
@@ -21,7 +21,7 @@ function appWithTmpSettings(extraEnv: Record<string, string> = {}, deps = {}) {
return { app, dir };
}
test("GET /settings returns effective defaults (env provider/model/thinking, first workspace)", async () => {
test("GET /settings returns thinking and the first workspace without legacy model defaults", async () => {
const { app, dir } = appWithTmpSettings({ PI_PROVIDER: "zai", PI_MODEL: "glm-5.2", PI_THINKING: "medium" }, {
listModels: async () => [{ provider: "zai", id: "glm-5.2", name: "GLM 5.2", reasoning: true }],
});
@@ -29,8 +29,8 @@ test("GET /settings returns effective defaults (env provider/model/thinking, fir
const res = await app.inject({ method: "GET", url: "/settings" });
expect(res.statusCode).toBe(200);
const body = res.json();
expect(body.provider).toBe("zai");
expect(body.model).toBe("glm-5.2");
expect(body).not.toHaveProperty("provider");
expect(body).not.toHaveProperty("model");
expect(body.thinking).toBe("medium");
expect(typeof body.workspace).toBe("string"); // first workspace from ../harness/workspaces
} finally {
@@ -62,7 +62,9 @@ test("PUT /settings does not persist personal workspace or LLM choices", async (
});
expect(put.statusCode).toBe(200);
const got = await app.inject({ method: "GET", url: "/settings" });
expect(got.json()).toMatchObject({ provider: "zai", model: "glm-5.2", thinking: "medium" });
expect(got.json()).toMatchObject({ thinking: "medium" });
expect(got.json()).not.toHaveProperty("provider");
expect(got.json()).not.toHaveProperty("model");
expect(got.json()).not.toMatchObject({ workspace: "psd", thinking: "high" });
} finally {
rmSync(dir, { recursive: true, force: true });
@@ -114,7 +116,7 @@ test("settings no longer read or write principal-specific preferences", async ()
expect(preferences.size).toBe(0);
});
test("GET /settings retains complete legacy installation defaults without seeding a private profile", async () => {
test("GET /settings drops legacy installation model fields without seeding a private profile", async () => {
let preferences: Record<string, unknown> = {};
const writes: Record<string, unknown>[] = [];
const runner = {
@@ -135,9 +137,7 @@ test("GET /settings retains complete legacy installation defaults without seedin
const first = await app.inject({ method: "GET", url: "/settings" });
const second = await app.inject({ method: "GET", url: "/settings" });
const expected = {
workspace: "local", provider: "local-qwen", model: "qwen3.6-35b-a3b", thinking: "low",
};
const expected = { workspace: "local", thinking: "low" };
expect(first.statusCode).toBe(200);
expect(first.json()).toEqual(expected);
expect(second.json()).toEqual(expected);
@@ -167,15 +167,13 @@ test("GET /settings ignores stale private preferences in favor of installation d
const response = await app.inject({ method: "GET", url: "/settings" });
expect(response.statusCode).toBe(200);
expect(response.json()).toEqual({
workspace: "local", provider: "local-qwen", model: "qwen3.6-35b-a3b", thinking: "low",
});
expect(response.json()).toEqual({ workspace: "local", thinking: "low" });
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
test("PUT /settings rejects an unknown model when a model list is available", async () => {
test("PUT /settings ignores a legacy unknown model because the catalog owns model validity", async () => {
const { app, dir } = appWithTmpSettings({}, {
listModels: async () => [{ provider: "zai", id: "glm-5.2", name: "GLM 5.2", reasoning: true }],
});
@@ -184,14 +182,14 @@ test("PUT /settings rejects an unknown model when a model list is available", as
method: "PUT", url: "/settings",
payload: { workspace: "psd", provider: "zai", model: "does-not-exist", thinking: "low" },
});
expect(put.statusCode).toBe(400);
expect(put.json()).toMatchObject({ error: expect.stringMatching(/model/i) });
expect(put.statusCode).toBe(200);
expect(put.json()).not.toHaveProperty("model");
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
test("PUT /settings validates provider and model as one composite identifier", async () => {
test("PUT /settings ignores legacy provider/model pairs", async () => {
const { app, dir } = appWithTmpSettings({}, {
listModels: async () => [
{ provider: "provider-a", id: "shared-id", name: "A", reasoning: false },
@@ -204,7 +202,7 @@ test("PUT /settings validates provider and model as one composite identifier", a
workspace: "psd", provider: "provider-b", model: "shared-id", thinking: "low",
},
});
expect(wrongProvider.statusCode).toBe(400);
expect(wrongProvider.statusCode).toBe(200);
const exactPair = await app.inject({
method: "PUT", url: "/settings",
@@ -213,6 +211,8 @@ test("PUT /settings validates provider and model as one composite identifier", a
},
});
expect(exactPair.statusCode).toBe(200);
expect(exactPair.json()).not.toHaveProperty("provider");
expect(exactPair.json()).not.toHaveProperty("model");
} finally {
rmSync(dir, { recursive: true, force: true });
}
+25 -33
View File
@@ -1,4 +1,4 @@
import { test, expect, vi } from "vitest";
import { test, expect } from "vitest";
import Fastify from "fastify";
import { buildApp } from "../src/app.js";
import { loadConfig } from "../src/config.js";
@@ -156,12 +156,23 @@ test("registry-backed SQL preview resolves and uses the session's pinned runtime
// Workspace registry route coverage lives in routes-workspaces.test.ts. `/workspaces` no longer
// reads legacy harness files: the Git registry is the single shared source of truth.
test("GET /models returns {models:[...]} from injected listModels stub", async () => {
test("GET /models returns session choices from the installation model catalog", async () => {
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: {} as any,
listModels: async () => [
{ provider: "zai", id: "glm-5.2", name: "GLM 5.2", reasoning: true },
],
runtimeModelCatalog: {
defaultSession: "zai/glm-5.2",
defaultMetadataGeneration: null,
embedding: { id: "ollama/qwen3-embedding:0.6b", dimensions: 1024 },
sessionModels: () => [{
id: "zai/glm-5.2", provider: "zai", model: "glm-5.2", label: "GLM 5.2",
upstreamModel: "glm-5.2", authentication: { mode: "pi_auth" },
sessionAdapter: { mode: "pi_builtin" }, session: { reasoning: true },
}],
metadataModels: () => [],
hasSession: (id: string) => id === "zai/glm-5.2",
},
// Runtime introspection is a health gate for starting a session, not a second catalog.
listModels: async () => { throw new Error("Pi is unavailable"); },
});
const res = await app.inject({ method: "GET", url: "/models" });
@@ -172,36 +183,17 @@ test("GET /models returns {models:[...]} from injected listModels stub", async (
});
});
test("GET /models returns {models:[]} when listModels throws (graceful fallback)", async () => {
test("GET /models returns an empty list when the catalog has no session models", async () => {
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: {} as any,
listModels: async () => { throw new Error("Pi not running"); },
});
const res = await app.inject({ method: "GET", url: "/models" });
expect(res.statusCode).toBe(200);
expect(res.json()).toEqual({ models: [] });
});
test("GET /models logs a sanitized warning when listing fails", async () => {
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: {} as any,
listModels: async () => { throw new Error("credential-value-must-not-appear"); },
});
const warn = vi.spyOn(app.log, "warn");
const res = await app.inject({ method: "GET", url: "/models" });
expect(res.json()).toEqual({ models: [] });
expect(JSON.stringify(warn.mock.calls)).not.toContain("credential-value-must-not-appear");
expect(warn).toHaveBeenCalled();
});
test("GET /models with empty listModels stub returns empty array", async () => {
const app = buildApp(loadConfig({ THT_HARNESS_DIR: "../harness" }), {
thtRunner: {} as any,
listModels: async () => [],
runtimeModelCatalog: {
defaultSession: null,
defaultMetadataGeneration: null,
embedding: null,
sessionModels: () => [],
metadataModels: () => [],
hasSession: () => false,
},
});
const res = await app.inject({ method: "GET", url: "/models" });
+71 -18
View File
@@ -11,10 +11,12 @@ import { WorkspaceRegistry, type WorkspaceRevision } from "../src/workspaces/reg
import { serializeWorkspaceYaml, type CanonicalWorkspace } from "../src/workspaces/schema.js";
import { WorkspaceSecretStore } from "../src/workspaces/secret-store.js";
import type { AuthDiagnoser, AuthDiagnostics } from "../src/auth/diagnostics.js";
import type { WorkspaceDatabase } from "../src/catalog/types.js";
import type { WorkspaceDatabaseTester } from "../src/routes/workspaces.js";
const workspace: CanonicalWorkspace = {
workspace: {
schema_version: 3,
schema_version: 4,
id: "psd-clinical",
name: "Policlinico San Donato",
description: "Clinical analytics workspace",
@@ -26,20 +28,6 @@ const workspace: CanonicalWorkspace = {
schema: "datawarehouse",
supported_transports: ["postgres_direct"],
},
semantic_index: {
vector_store: {
engine: "qdrant",
collection: "psd-clinical",
dimensions: 1024,
distance: "cosine",
},
embedding: {
provider: "ollama_internal",
model: "qwen3-embedding:0.6b",
dimensions: 1024,
},
},
llm_policy: { allowed: ["zai/glm-5.2"] },
};
const revision: WorkspaceRevision = {
@@ -77,12 +65,35 @@ const readyAuthentication: AuthDiagnostics = {
checks: [{ level: "info", code: "auth_ready", message: "Authentication is ready." }],
};
const reachableWorkspaceDatabase: WorkspaceDatabase = {
id: "db-psd-clinical",
workspaceId: "psd-clinical",
engine: "postgres",
databaseName: "warehouse",
schema: "datawarehouse",
version: 1,
createdAt: "2026-01-01T00:00:00.000Z",
updatedAt: "2026-01-01T00:00:00.000Z",
binding: {
transport: "postgres_direct",
host: "current-db.internal",
port: 5432,
username: "current-reader",
},
connectionStatus: "reachable",
testedVersion: 1,
lastTestedAt: "2026-01-01T00:00:00.000Z",
};
function appFor(
registry: RegistryFake,
diagnose = vi.fn(async () => ({ activatable: true, diagnostics: [] })),
secretStore = testSecretStore(),
env: Record<string, string> = {},
authDiagnoser: AuthDiagnoser = { inspect: vi.fn(async () => readyAuthentication) },
workspaceDatabaseTester: WorkspaceDatabaseTester = vi.fn(
async () => reachableWorkspaceDatabase,
),
) {
return buildApp(loadConfig({
THT_HARNESS_DIR: "/missing-harness",
@@ -94,6 +105,7 @@ function appFor(
workspaceDiagnoser: diagnose,
workspaceSecretStore: secretStore,
authDiagnoser,
workspaceDatabaseTester,
} as any);
}
@@ -194,7 +206,7 @@ test("lists workspace summaries and reads a validated immutable workspace", asyn
expect(read.json()).toEqual({ workspace, revision });
});
test("validates a schema v3 workspace without mutating the repository", async () => {
test("validates a schema v4 workspace without mutating the repository", async () => {
const app = appFor(registryFake());
const response = await app.inject({
@@ -284,7 +296,7 @@ test.each([1, 2])("rejects schema v%s at the validation boundary with a sanitize
expect(response.body).not.toMatch(/migration_required|schema version/i);
});
test("runs diagnostics for a schema v3 workspace", async () => {
test("runs diagnostics for a schema v4 workspace", async () => {
const diagnose = vi.fn(async () => ({ activatable: true, diagnostics: [] }));
const app = appFor(registryFake(), diagnose);
@@ -297,7 +309,48 @@ test("runs diagnostics for a schema v3 workspace", async () => {
expect(diagnose).toHaveBeenCalledWith(workspace, {
dwh: expect.objectContaining({ transport: "postgres_direct" }),
evidence: { missing: [], values: {} },
}, { writeProbe: false });
}, { writeProbe: false, skipDwh: true });
});
test("reports a missing Database Management configuration without using the legacy DWH test", async () => {
const diagnose = vi.fn(async () => ({
activatable: true,
diagnostics: [{
level: "info" as const,
code: "binding_ok" as const,
message: "Installation bindings and diagnostics succeeded.",
}],
}));
const workspaceDatabaseTester = vi.fn(async () => undefined);
const app = appFor(
registryFake(),
diagnose,
testSecretStore(),
{},
{ inspect: vi.fn(async () => readyAuthentication) },
workspaceDatabaseTester,
);
const response = await app.inject({
method: "POST", url: "/workspaces/psd-clinical/test", payload: {},
});
expect(response.statusCode).toBe(200);
expect(response.json()).toMatchObject({
activatable: false,
diagnostics: [{
level: "error",
code: "binding_missing",
field: "dwh",
message: "Configure this workspace in Database Management before testing connections.",
}],
});
expect(diagnose).toHaveBeenCalledWith(
workspace,
expect.anything(),
{ writeProbe: false, skipDwh: true },
);
expect(workspaceDatabaseTester).toHaveBeenCalledWith("psd-clinical");
});
test("reports runtime secret requirements without returning stored values", async () => {
+6 -6
View File
@@ -27,9 +27,9 @@ test("saveSettings writes the file and loadSettings reads it back", () => {
const dir = mkdtempSync(join(tmpdir(), "tht-set-"));
try {
const cfg = cfgWith(join(dir, "nested", "settings.json"));
const saved = saveSettings(cfg, { workspace: "psd", provider: "zai", model: "glm-5.2", thinking: "medium" });
expect(saved.model).toBe("glm-5.2");
expect(loadSettings(cfg)).toEqual({ workspace: "psd", provider: "zai", model: "glm-5.2", thinking: "medium" });
const saved = saveSettings(cfg, { workspace: "psd", thinking: "medium" });
expect(saved.thinking).toBe("medium");
expect(loadSettings(cfg)).toEqual({ workspace: "psd", thinking: "medium" });
} finally {
rmSync(dir, { recursive: true, force: true });
}
@@ -55,11 +55,11 @@ test("saveSettings restores the previous file when post-rename directory durabil
const dir = mkdtempSync(join(tmpdir(), "tht-set-transaction-"));
try {
const cfg = cfgWith(join(dir, "settings.json"));
saveSettings(cfg, { provider: "old", model: "old-model", thinking: "low" });
saveSettings(cfg, { thinking: "low" });
let syncs = 0;
expect(() => saveSettings(
cfg,
{ provider: "new", model: "new-model", thinking: "high" },
{ thinking: "high" },
{
syncDirectory(directory: string) {
syncs += 1;
@@ -69,7 +69,7 @@ test("saveSettings restores the previous file when post-rename directory durabil
},
},
)).toThrow(/directory fsync failure/);
expect(loadSettings(cfg)).toEqual({ provider: "old", model: "old-model", thinking: "low" });
expect(loadSettings(cfg)).toEqual({ thinking: "low" });
expect(syncs).toBeGreaterThanOrEqual(2);
} finally {
rmSync(dir, { recursive: true, force: true });
+1 -8
View File
@@ -8,18 +8,11 @@ const keywordIndexes = [
];
const workspace: CanonicalWorkspace = {
workspace: { schema_version: 3, id: "psd", name: "PSD", language: "it" },
workspace: { schema_version: 4, id: "psd", name: "PSD", language: "it" },
dwh: {
engine: "postgres", database: "warehouse", schema: "public",
supported_transports: ["postgres_direct"],
},
semantic_index: {
vector_store: { engine: "qdrant", collection: "psd", dimensions: 1024, distance: "cosine" },
embedding: {
provider: "ollama_internal", model: "qwen3-embedding:0.6b", dimensions: 1024,
},
},
llm_policy: { allowed: ["zai/glm-5.2"] },
};
function runner(request: (...args: any[]) => Promise<any>) {
@@ -19,12 +19,13 @@ afterEach(() => {
const semanticRuntime = {
internalQdrantUrl: "http://qdrant:6333",
internalEmbeddingUrl: "http://embedding:11434",
internalEmbeddingId: "ollama/qwen3-embedding:0.6b",
internalEmbeddingModel: "qwen3-embedding:0.6b",
internalEmbeddingDimensions: 1024,
};
const baseWorkspace = parseWorkspaceYaml(`workspace:
schema_version: 3
schema_version: 4
id: psd-clinical
name: Runtime Lease
language: en
@@ -33,21 +34,9 @@ dwh:
database: analytics
schema: mart
supported_transports: [postgres_direct]
semantic_index:
vector_store:
engine: qdrant
collection: psd-clinical
dimensions: 1024
distance: cosine
embedding:
provider: ollama_internal
model: qwen3-embedding:0.6b
dimensions: 1024
llm_policy:
allowed: [zai/glm-5.2]
`);
const filesystemWorkspace = parseWorkspaceYaml(`${baseWorkspace ? '' : ''}workspace:
schema_version: 3
schema_version: 4
id: fs-workspace
name: Filesystem
language: en
@@ -56,25 +45,13 @@ dwh:
database: analytics
schema: mart
supported_transports: [postgres_direct]
semantic_index:
vector_store:
engine: qdrant
collection: fs-workspace
dimensions: 1024
distance: cosine
embedding:
provider: ollama_internal
model: qwen3-embedding:0.6b
dimensions: 1024
llm_policy:
allowed: [zai/glm-5.2]
evidence:
source:
type: filesystem
uri: fs-workspace/evidence
`);
const privateHttpWorkspace = parseWorkspaceYaml(`workspace:
schema_version: 3
schema_version: 4
id: http-workspace
name: Http
language: en
@@ -83,18 +60,6 @@ dwh:
database: analytics
schema: mart
supported_transports: [postgres_direct]
semantic_index:
vector_store:
engine: qdrant
collection: http-workspace
dimensions: 1024
distance: cosine
embedding:
provider: ollama_internal
model: qwen3-embedding:0.6b
dimensions: 1024
llm_policy:
allowed: [zai/glm-5.2]
evidence:
source:
type: http
@@ -125,7 +90,7 @@ function runtime(workspace = baseWorkspace, workspaceId = workspace.workspace.id
bindingDigest: "sha256:bindings",
semanticQdrantUrl: "http://qdrant:6333",
effectiveConfig: {
schemaVersion: 1,
schemaVersion: 2,
dwh: {
engine: "postgres",
database: "analytics",
@@ -136,7 +101,7 @@ function runtime(workspace = baseWorkspace, workspaceId = workspace.workspace.id
user: "reader",
},
vector: { collection: workspaceId, dimensions: 1024, distance: "cosine" },
embedding: { model: "qwen3-embedding:0.6b", dimensions: 1024 },
embedding: { id: "ollama/qwen3-embedding:0.6b", model: "qwen3-embedding:0.6b", dimensions: 1024 },
roots: { artifacts: "/data/artifacts", indexes: "/data/indexes" },
},
effectiveConfigIdentity: "workspace://psd-clinical@v1:" + "d".repeat(64),

Some files were not shown because too many files have changed in this diff Show More