feat: complete catalog-driven preprocessing
Publish documentation / publish (push) Successful in 2m12s

This commit is contained in:
Codex
2026-09-06 17:49:35 +02:00
parent 8707ae1d46
commit cffa60772e
141 changed files with 5898 additions and 3015 deletions
@@ -0,0 +1,63 @@
# Use PostgreSQL metadata for all core consumers
The Metadata Catalog in PostgreSQL is the sole authority for every fact about a Workspace Database
used by the core. This completes the cutover anticipated by ADR-0001 and ADR-0003 and supersedes the
ADR-0012 statements that kept workspace annotations authoritative for descriptions and materialized
a relationship-only snapshot per session. The Workspace Descriptor retains workspace identity and
Evidence scope, but contains no database identity, binding, structure, description, or relationship
data.
## Decision
All downstream consumers receive one deterministic, versioned **Catalog Metadata Snapshot** produced
from PostgreSQL. It contains the complete cataloged table and column structure, sensitivity metadata,
publishable descriptions, and the active Effective Relationship Map. Description precedence is
`Description`, then `Generated Description`, then the observed source comment; generated text is
therefore publishable without prior consolidation. Qdrant stores first-class table, column, and
relationship records, and a relationship hit contributes to both endpoint tables.
The native host CLI exposes one mutating preprocessing operation. The Node
`workspace-maintenance` process acquires a PostgreSQL lock shared with catalog synchronization,
description generation, and administrative catalog writes; it refuses to run while one of those
operations is active. It reads the catalog once, writes the snapshot to a temporary immutable JSON
file, and passes only that file to the Python harness. Catalog credentials are excluded from the
child environment. The harness never reads the Metadata Catalog directly.
Preprocessing is deliberately offline: no session uses the core while it is running, nobody starts
the core until it completes, and concurrent catalog modification is forbidden. PostgreSQL owns a
minimal `running | succeeded | failed` Preprocessing State, the input fingerprint, the Metadata
Content Revision, and the last successfully processed revision. Every relevant catalog mutation
increments the revision and makes the prior success stale in the same transaction. Core admission
requires a current `succeeded` state, but the system does not implement session draining, per-session
schema pinning, coexisting preprocessing generations, retention, or rollback.
Each run may overwrite the previous derived state. It replaces the Qdrant `schema` and `evidence`
slices while preserving `memory` and `solved_question`, rebuilds the current LSH and other required
artifacts, verifies the result, and marks success only at the end. A failure leaves the workspace
unready; rerunning starts from a clean derived state. Existing internal DWH or Evidence staging may
remain an implementation detail with minimum retention, but it is not an application-level
generation contract. A workspace without Evidence is valid and publishes an explicit absent state.
`physical.yaml`, `annotations.yaml`, the YAML FK review flow, and public partial preprocessing
commands are removed from the runtime contract. No legacy metadata is imported. Concepts, synonyms,
notes, annotation evidence, and the manual `eligible` override are retired; eligibility is computed
deterministically. The DWH is queried during preprocessing only for derived samples and LSH values,
not as an alternative metadata authority. The existing **Catalog Schema Snapshot** name remains
reserved for the DWH-to-catalog synchronization RPC.
## Considered Options
- Direct PostgreSQL access from every consumer would duplicate catalog queries and expose storage
details and credentials to the harness.
- Immutable application-level generations and an atomic composite release would permit concurrent
core use and rollback, but those capabilities are unnecessary under the accepted offline operating
constraints.
- Retaining or importing workspace annotations would preserve legacy semantic fields but would keep
an obsolete second authority despite there being no production data to migrate.
## Consequences
Every runtime metadata consumer uses the Catalog snapshot, not only Qdrant indexing. Read-only
inspection may remain, but only the complete preprocessing command can mutate derived state. The
workspace-preprocessing contract and tests describe this cutover directly; superseded partial
command designs remain available only in Git history.
@@ -0,0 +1,48 @@
# Separate reference vectors from runtime memory
Schema, relationships, and Evidence are deterministic projections of PostgreSQL Catalog metadata
and the pinned workspace Evidence revision. Memory and solved questions are produced by core use and
have a different lifecycle. Keeping both classes of records in one Qdrant collection makes a full
cleanup of derived preprocessing data unsafe because it can also erase durable runtime knowledge.
## Decision
Each workspace has two physical Qdrant collections behind the existing logical vector-store port:
- `<workspace>-reference` contains `schema_table`, `schema_column`, `schema_relationship`, and
`evidence`;
- `<workspace>-memory` contains `memory` and `solved_question`.
Callers continue to use logical record groups. The Qdrant adapter owns physical routing and merges
results when a logical search spans both collections. BM25 configuration belongs only to the
reference collection.
The public preprocessing clear operation deletes the reference collection, LSH, Evidence corpus,
private Catalog snapshot, and derived checkpoints. It never deletes the memory collection, session
artifacts, Catalog metadata, or source data. PostgreSQL records `derived_data_cleared`, which projects
to a required preprocessing state and prevents core admission until a complete run succeeds.
On the first preprocessing write or clear after this split, the adapter copies `memory` and
`solved_question` points from an existing legacy workspace-named collection into the new memory
collection before retiring the legacy collection. There is no general metadata migration, history,
generation retention, or rollback mechanism.
LSH ownership includes workspace ID, Catalog database ID, Metadata Content Revision, effective
configuration fingerprint, and input fingerprint. Reuse fails closed when any identity differs.
## Considered Options
- Deleting only Schema/Evidence points from one shared collection would keep memory but retain a
coupled physical lifecycle and make collection-level repair or recreation unsafe.
- Rebuilding memory after clear is impossible because preprocessing does not own its source of
truth.
- Keeping multiple preprocessing generations would add retention and rollback behavior that the
accepted offline operating model does not require.
## Consequences
Preprocessing can now clear and recreate all of its vector output without touching runtime memory.
Diagnostics and effective configuration must validate two collections. Operators get one clear
command and one inline-confirmed sidebar action; after either succeeds, preprocessing is required.
The split remains encapsulated in the vector adapter, so workflow and retrieval callers do not need
to know physical Qdrant collection names.
+16 -10
View File
@@ -48,8 +48,8 @@ A session is a directory under `sessions/` (the workspace defines the path): `se
- `PiProcessManager` runs one Pi child process per session and bridges its RPC stream.
- `SessionBridge` maps Pi RPC events to client events (`ui_request` / `text_delta` / `info`).
- `SseHub` distributes these events to the browser over SSE.
- `CatalogService` joins authoritative YAML workspace identities with installation-local database
configurations stored in PostgreSQL through Kysely.
- `CatalogService` joins authoritative YAML workspace identities and Evidence settings with the
sole database authority: installation-local PostgreSQL Metadata Catalog records.
- `CatalogTableService` reconciles persisted Catalog Tables with a successful external schema scan;
`ConcreteCatalogTableIntrospector` isolates direct PostgreSQL, typed REST, and SSH-tunnel access.
- `DescriptionGenerationWorker` admits one installation-wide run and processes its table or column
@@ -60,8 +60,9 @@ A session is a directory under `sessions/` (the workspace defines the path): `se
Application settings keep only workspace and thinking preferences in `backend/data/settings.json`;
provider/model defaults come from the generated Installation Model Catalog. Session state remains in
harness phase documents. PostgreSQL stores only the administrative database catalog, bindings, observed tables,
curated and generated descriptions, description-generation runs, and sanitized run events.
harness phase documents. PostgreSQL stores the administrative database catalog, bindings, observed tables,
curated and generated descriptions, description-generation runs, sanitized run events, and the
authoritative preprocessing state/revisions.
Connector secrets remain write-only in the encrypted workspace secret store; model credentials
remain in the protected installation secret bundle. Catalog SSH support is limited to connection
tests, table synchronization, and bounded description-generation sampling; it does not change the
@@ -86,18 +87,23 @@ only `curated/**/*.md` from the immutable revision root to preprocessing. Source
evaluation data remain available for traceability. The runtime never modifies, stages, commits, or
publishes the authoring repository.
Before indexing, the curated corpus from the pinned revision is validated. The shared Qdrant
collection keeps the unnamed dense vector used by Schema and Memory. Evidence preprocessing may
add only the sparse `bm25` vector with `idf`, without deleting, renaming, or recreating the collection.
`workspace preprocess evidence` and the Evidence part of `workspace preprocess run` are the only public
operations that perform this upgrade.
Before indexing, the curated corpus from the pinned revision is validated. Each workspace has two
physical Qdrant collections with different lifecycles. `reference` contains Schema, relationships,
and Evidence and may be replaced or cleared by preprocessing; `memory` contains `memory` and
`solved_question` records and is never preprocessing output. Only the `reference` collection has the
sparse `bm25` vector with `idf`, and only the Evidence stage writes sparse values.
The Administration control can clear the replaceable reference collection, LSH, corpus, and derived
checkpoints. The operation preserves the memory collection and makes preprocessing required before
the core can admit a new session.
## Recurring points of attention
- `tht -c`/`--config` is a **per-command** option. It must follow the subcommand, never precede it (`ThtRunner.buildArgv` enforces this).
- `--json` output must be plain JSON on stdout. It is a machine-readable contract.
- UI strings are in English. Document *content* stays in the workspace language because it is the actual data; only chrome and labels are in English.
- Each workspace defines its DWH target and working directories. Secrets remain in protected installation files, not in the workspace repository.
- Each workspace defines identity and optional Evidence only. The PostgreSQL Metadata Catalog defines
its DWH target and binding; secrets remain in the protected workspace secret store.
- Settings are global (`backend/data/settings.json`: workspace/thinking); provider/model choices are
canonical catalog references selected per session.
- **Resume**: a resumable session returns to its last incomplete phase. The backend rejects resume with 409 when `finalized` or `archived`; `PiProcessManager.spawnFor` must send `/riprendi-sessione <id>` for resume and `/nuova-domanda` for a new session. The wrong prompt silently turns a resume into a new question.
+28 -121
View File
@@ -1,128 +1,35 @@
# `.tht-dwh` — DWH generations, `OWNER.json`, ACTIVE, and fingerprints
# Internal Catalog-bound DWH artifacts
> Operator contract. P3 makes the effective DWH/preprocessing configuration reproducible and
> versioned across the operator CLI and the application sessions, and documents what `.tht-dwh`
> is so operators can reason about why a rerun is instant or why it takes minutes.
`.tht-dwh`, `OWNER.json`, and the `ACTIVE` generation pointer are private preprocessing artifacts,
not an operator-facing generation or rollback contract. Complete preprocessing materializes the
physical schema from the immutable PostgreSQL Catalog Metadata Snapshot and samples only eligible
DWH values to build LSH. It never obtains schema metadata from workspace YAML or by introspecting the
DWH in the harness.
## What `.tht-dwh` is
`OWNER.json` binds the artifact root to all of:
`.tht-dwh` is the workspace-local directory that stores the **prepared snapshots of the data
warehouse structure** (the catalog `physical.yaml` plus the LSH hashes used for fuzzy search).
ThothII does not re-read the whole database for every question: it prepares it once, stores the
result here, and reuses it. The directory lives under the workspace runtime root, for example:
- workspace ID;
- Catalog database ID;
- Metadata Content Revision;
- effective configuration fingerprint;
- input fingerprint.
```text
/data/sessions/<workspace-id>/.tht-dwh/
The complete binding must match before artifacts can be reused. This prevents an LSH built from one
database, workspace, Catalog revision, or effective configuration from being associated with
another. PostgreSQL separately owns readiness through `running | succeeded | failed`, the processed
revision, and the preprocessing input fingerprint.
The internal generation is overwritten by normal preprocessing and is not retained for rollback.
Recovery is always:
```sh
tht --installation <absolute>/thothii-installation.yaml \
workspace preprocess run --workspace <workspace-id>
```
## Immutable generations
The clear operation removes `.tht-dwh` together with the LSH, private Catalog snapshot, Evidence
corpus, and derived checkpoints. The next run recreates them from PostgreSQL and the pinned workspace
revision.
Each preparation run produces a **generation**: an immutable directory containing the catalog and
the LSH artifacts for one exact "effective configuration" (see fingerprints below). Generations
are never modified in place; a new run writes a new generation, and an `ACTIVE` pointer selects
which generation the workspace currently uses. Keeping the old generations makes rollback and
diagnosis safe.
## `OWNER.json`
Every generation root contains an `OWNER.json` that records who owns it:
```json
{
"workspace_id": "<workspace-id>",
"config_fingerprint": "sha256:<64 hex>",
"input_fingerprint": "sha256:<64 hex>"
}
```
- `config_fingerprint` is the digest of the **canonical effective configuration** (see below).
- `input_fingerprint` is the digest of the **logical configuration identity**.
Before reusing a generation, the harness compares the current canonical identity with the one in
`OWNER.json`. If they differ, the generation is **refused** (never silently reused) and a new one
is produced. This is what protects ThothII from using artifacts prepared for a different database,
endpoint, user, schema, or index contract.
The reader is compatible with the historical schema-v1 `OWNER.json` (same three keys, `sha256:`
values) so existing installations keep working; new writes use the versioned computation. There is
no automatic in-place reinterpretation: operators regenerate explicitly when a root is old.
## The canonical effective configuration and the logical identity
The **canonical effective configuration** is the non-secret subset of the rendered runtime
configuration that determines whether a prepared DWH generation is still valid:
```json
{
"schemaVersion": 1,
"dwh": {
"engine": "postgres",
"database": "<database>",
"schema": "<schema>",
"transport": "postgres_direct | rest_api | ...",
"host": "<host>",
"port": 5432,
"baseUrl": "<base-url>",
"user": "<user>"
},
"vector": { "collection": "<collection>", "dimensions": 1024, "distance": "cosine" },
"embedding": { "model": "<model>", "dimensions": 1024 },
"roots": { "artifacts": "<abs-path>", "indexes": "<abs-path>" }
}
```
Deliberately **excluded** (their change must not invalidate a DWH generation):
- `session_storage` and `runtime_identity` (a content-only Git commit or an Evidence-only change
must not force a full database re-introspection);
- Evidence source/policy (P6 materialization and Evidence preprocessing are separate);
- memory, search, and execution settings;
- **all credentials** (passwords, API keys, signed URLs, and secret-file paths).
The **logical configuration identity** is:
```text
workspace://<workspace-id>@v1:<sha256 of the canonical effective configuration>
```
It is the same for the operator CLI and for application sessions, because both derive it from the
same rendered configuration. That is the guarantee that the work prepared by `tht` is exactly
what the sessions will consume.
## Why a rerun can be instant or take minutes
- Same canonical identity (e.g., only Evidence files changed) → the generation is reused → the
DWH step is `unchanged` and fast.
- Changed canonical identity (different database, address, user, schema, collection, model, or
artifact/index roots) → the old generation is refused → ThothII re-introspects and writes a new
generation → the step takes as long as the first preparation.
## Safe migration, regeneration, and recovery
- **Migration**: existing schema-v1 `OWNER.json` roots are readable; to switch them to the
versioned identity, run a normal regeneration (explicit `--refresh`/new run). No automatic
in-place rewrite.
- **Regeneration**: a new run produces a new immutable generation and moves `ACTIVE`; the previous
generations remain for rollback.
- **Recovery**: if the active generation is corrupt or owned by another configuration, ThothII
fails closed (never mixes artifacts) and tells the operator to regenerate; the old generations
are still available for inspection.
## Memory root
P3 also gives each workspace an explicit **workspace-global memory root**:
```text
/data/sessions/<workspace-id>/memory/
```
All memory commands, locks, the canonical JSONL registry, and the Qdrant projection use this root
when present. A guarded migration copies and verifies exactly one legacy canonical JSONL from the
old `artifacts/memory` location under the workspace lock and rebuilds the projection; conflicting
legacy registries fail closed. There is no in-place reinterpretation.
## Revision-scoped search records
Schema and Evidence records in the Qdrant collection include the pinned `workspace_revision`, so
searches never mix descriptions or documents from different versions of the workspace. Memory and
solved-question records remain workspace-wide on purpose.
Workspace-global `memory` and Qdrant `solved_question` data are not preprocessing output and remain
preserved both when reference data is overwritten and when preprocessing is cleared.
+11 -5
View File
@@ -5,6 +5,10 @@ workspace descriptor. Evidence is optional: a valid v4 descriptor without it rem
When present, `evidence` is strict: it contains `source` and a defaulted strict `policy`; every
source variant and the policy reject unknown keys.
The version numbers are intentionally separate: the Evidence descriptor supports v1/v2, while the
latest Curated Evidence Unit format is v3. There is no Evidence descriptor v3/v4 and no Curated
Evidence Unit v4.
```mermaid
flowchart LR
REGISTRY["Workspace registry"] --> DESCRIPTOR["Evidence descriptor"]
@@ -228,13 +232,15 @@ authoring repository.
P1.1 performs no acquisition, extraction, preprocessing/indexing, embeddings, Qdrant writes,
active-snapshot retention, or GC.
## Operator validation
## Runtime publication
After the runtime configuration is rendered or acquired, validate it with the exact per-command
option ordering:
After the curator has published a valid revision, the installation operator publishes it only as
part of complete workspace preprocessing:
```sh
tht config check -c <path>
tht --installation <absolute>/thothii-installation.yaml \
workspace preprocess run --workspace <workspace-id>
```
Stop after validation. P2/P6 later owns preprocessing and materialization.
This command also consumes the PostgreSQL Catalog snapshot and rebuilds schema/LSH output. There is
no public Evidence-only preprocessing command.
+149 -160
View File
@@ -1,193 +1,182 @@
# Workspace preprocessing CLI contract
`tht` is the only supported host entrypoint for workspace preprocessing.
`tht` exposes one complete, one-shot mutating preprocessing operation. It must succeed before the
workspace can be used by the core.
```mermaid
flowchart LR
OP["Operator"] --> INVOKE["tht workspace preprocess"]
INVOKE --> VALIDATE["Validate descriptor\nand paths"]
VALIDATE --> SOURCE["Read source and\ncurated workspace"]
SOURCE --> NORMALIZE["Normalize and chunk"]
NORMALIZE --> INDEX["Update vector and\nBM25 indexes"]
INDEX --> VERIFY["Verify collection\nand generation"]
VERIFY --> READY["Generation ready"]
VALIDATE -->|"invalid"| STOP["Exit with diagnostic"]
```
## Invocation
## Public commands
```text
tht --installation <absolute>/thothii-installation.yaml workspace inspect
--workspace <id> [--json]
tht --installation <absolute>/thothii-installation.yaml workspace preprocess dwh
--workspace <id> [--resume <32hex>] [--json]
tht --installation <absolute>/thothii-installation.yaml workspace schema suggest-fks
--workspace <id>
[--from-sql <regular-file>]... [--assume <column=table>]...
[--output <new-file>] [--json]
tht --installation <absolute>/thothii-installation.yaml workspace schema check
--workspace <id>
[--annotations <regular-file> --reviewed-candidates <sha256:hex>]
[--json]
tht --installation <absolute>/thothii-installation.yaml workspace schema accept
--workspace <id> --run <32hex> --yes [--json]
tht --installation <absolute>/thothii-installation.yaml workspace index-schema
--workspace <id> [--json]
tht --installation <absolute>/thothii-installation.yaml workspace preprocess evidence
--workspace <id> [--dry-run] [--resume <32hex>] [--json]
tht --installation <absolute>/thothii-installation.yaml workspace preprocess run
--workspace <id> [--resume <32hex>] [--json]
tht --installation <absolute>/thothii-installation.yaml workspace vector inspect
--workspace <id> [--json]
tht --installation <absolute>/thothii-installation.yaml workspace vector rebuild
--workspace <id> --collection <name> --confirm <name> --destroy [--json]
tht --installation <absolute>/thothii-installation.yaml workspace preprocess clear
--workspace <id> [--json]
```
## Qdrant collection lifecycle (P4)
There are no public partial commands for DWH introspection, LSH, FK suggestions, schema indexing,
Evidence indexing, or Qdrant rebuild. `preprocess run` does not accept `--resume`, `--dry-run`, a
generation identifier, or a rollback option. Re-running it replaces the preceding derived output.
`preprocess clear` removes all replaceable preprocessing output and preserves runtime memory.
- `workspace vector inspect` reports the descriptor-owned Qdrant collection contract
(name, dimensions, distance, keyword indexes) **without mutation**.
- `workspace vector rebuild` deletes and recreates the descriptor-owned collection
with the exact contract (1024 dimensions, cosine distance, the 8 required keyword
payload indexes) under guards:
- `--collection <name>` must equal the descriptor's `semantic_index.vector_store.collection`;
- `--confirm <name>` must equal `--collection` (exact repetition);
- `--destroy` is required to confirm the destructive operation;
- the operator refuses any other combination with exit code 2 (usage).
- Self-heal at session admission: a missing collection is created and missing
keyword indexes are added by the shared collection manager; incompatible
dimensions/distance/index types are never mutated (`semantic_index_incompatible`).
- The operator path (`workspace-maintenance.js vector-inspect|vector-rebuild`)
performs the guarded rebuild; rebuild state is written before deletion and the
collection is verified after recreation. No prefix matching or global Qdrant
mutation is performed.
## Sources of truth
## Additive BM25 for Evidence
- The Workspace Descriptor v4 contains workspace identity and optional Evidence configuration.
It contains no database identity, connection, table, column, description, sensitivity, or
relationship data.
- PostgreSQL Metadata Catalog is the sole database authority. It owns the workspace/database
association, installation-local binding, tables, columns, descriptions, sensitivity flags,
physical foreign keys, and active logical relationships.
- The workspace Git revision remains authoritative for Evidence. Evidence Descriptor v1/v2 and
Curated Evidence Unit v3 are unchanged; there is no Evidence v4.
- Only `workspace preprocess evidence` (and the Evidence portion of `workspace preprocess run`)
may add the named sparse vector `bm25` with Qdrant modifier `idf`.
- The upgrade uses Qdrant's additive named-vector operation. It preserves the existing unnamed
dense vector and never deletes, renames, or rebuilds the shared collection.
- Session readiness remains read-only with respect to BM25. Schema, Memory, and solved-question
records therefore continue to use their existing dense-only points during and after an Evidence
upgrade.
- A missing `bm25` is added and reread before Evidence preprocessing starts. An existing definition
other than `modifier: idf` fails as `semantic_index_incompatible` without any collection mutation.
If a later Evidence candidate fails, the compatible additive schema remains in place; it does not
make the dense-only records unavailable.
No metadata is imported from legacy workspace YAML or `physical.yaml`/`annotations.yaml`.
## Preconditions and lock
The maintenance process obtains the PostgreSQL preprocessing lease for the workspace. Acquisition
fails unless:
- the workspace has a Catalog database;
- its latest schema synchronization matches the current database configuration version;
- no catalog sync, description-generation, or sensitivity-analysis run is active;
- no other preprocessing run is active.
While the state is `running`, PostgreSQL rejects database binding changes and every write to the
catalog tables, columns, physical relationships, relationship columns, and logical relationships.
It also rejects the start of the three conflicting background operations. The accepted operating
model assumes no core session is running and nobody attempts core admission during preprocessing;
there is therefore no drain protocol or session pinning.
## Complete pipeline
```mermaid
flowchart LR
CLI["workspace preprocess run"] --> LOCK["Acquire PostgreSQL lease"]
LOCK --> SNAP["Build Catalog Metadata Snapshot"]
SNAP --> LSH["Sample DWH and rebuild LSH"]
SNAP --> SCHEMA["Replace schema vectors"]
SCHEMA --> REL["Index relationships as first-class records"]
LSH --> EVIDENCE["Validate and replace Evidence vectors"]
REL --> EVIDENCE
EVIDENCE --> VERIFY["Verify complete result"]
VERIFY --> READY["Commit succeeded state"]
```
## Curated FK annotations (P5)
The backend reads PostgreSQL and writes one private `Catalog Metadata Snapshot` JSON file. The
Python child receives the snapshot path and the Catalog-derived DWH runtime binding, but never the
Metadata Catalog credentials. It does not query the Catalog.
- The canonical curated annotations file is `<workspace-id>/schema/annotations.yaml`, a regular
Git blob at the same commit as the descriptor. Absence is compatible (empty canonical set +
warning); symlinks, trees/gitlinks, oversized (>16 MiB), non-UTF-8, and malformed objects are
refused at activation.
- Activation synchronizes the blob to the immutable revision-qualified root
`/data/sessions/<id>/revisions/<commit>/artifacts/mschema/annotations.yaml` with a restrictive
mode and an adjacent ownership manifest `{ workspace, commit, blobId, contentDigest,
destination }`. Pinned runtimes resolve annotations from `paths.annotations_root`.
- `workspace schema accept --run <id> --yes` is the only human FK review primitive: after
commit/push/pull, it reads the current synced blob, validates it with the harness parser against
the physical schema and the recorded candidate digest, and records
`{ reviewedCandidatesDigest, annotationsDigest, workspaceRevision, blobId }`. `--yes` is
required; an empty file, an unknown run, a malformed blob, or a non-matching candidate fails
closed without recording a review. `schema check` alone is not evidence of human review.
- `preprocess run` continues only with the exact accepted blob digest and a compatible reusable
DWH binding; otherwise it records a new review checkpoint.
The snapshot contains the complete table and column structure, sensitivity, and effective active
relationships. Effective descriptions use this precedence:
## Commit-addressed Evidence materialization (P6)
1. curated `description`;
2. `generatedDescription`;
3. PostgreSQL `sourceComment`.
- Filesystem Evidence `<workspace-id>/evidence` is materialized from the exact pinned Git commit
into the immutable revision content root `<registry>/snapshots/<commit>/<id>/evidence` at
activation, with a sibling bounded manifest `<id>/evidence.manifest.json` whose digest is chained
into `snapshot.json`.
- Materialization uses fixed Git plumbing (`ls-tree -r -z` + `cat-file blob`) and refuses symlinks
and gitlinks at any depth, traversal/absolute/duplicate/cross-namespace paths, and non-regular
modes. Installation-local limits bound entry count (default 4096), total bytes (64 MiB),
per-file bytes (8 MiB), path bytes (4096), and manifest bytes (1 MiB); a size-sum preflight runs
before any bytes are written and no partial root is published.
- `preprocess evidence` and `preprocess run` operate directly on the materialized root; the
temporary `evidence_materialization_required` stop is retired (the code remains only for
pre-P6 compatibility). For `evidence.schema_version: 2`, runtime acquisition receives exactly
`curated/**/*.md`; `source/` and support files remain in the materialized tree for traceability.
HTTP/S3 Evidence is unchanged.
- The curator validates Evidence before merge. Preprocessing validates the pinned curated corpus
again before it constructs a candidate generation, so an invalid revision is never indexed.
- The runtime writes only its immutable materialized snapshot and derived index state. It never
writes, stages, commits, or pushes the workspace authoring repository.
- Materialized roots are retained with their commit-addressed snapshot directory and removed only
when the revision becomes unreferenced.
The DWH is queried only for derived value samples used by LSH. Sensitive columns are never sampled.
Eligibility is computed deterministically; legacy annotation concepts, synonyms, notes, Evidence
links, and manual eligibility overrides are not part of the snapshot.
## Validation
Qdrant receives first-class `schema_table`, `schema_column`, and `schema_relationship` records. A
relationship search hit promotes both endpoint tables. Each workspace has two physical collections:
- `--installation` is mandatory and absolute.
- `--workspace` is mandatory exactly once and must match `[a-z][a-z0-9-]{2,62}`.
- `--resume` values must be 32 lowercase hex characters.
- `--json` may be supplied once.
- `schema suggest-fks`
- allows at most 32 `--from-sql` files;
- each SQL file must be a canonical regular file, UTF-8, non-symlink, max 1 MiB;
- total SQL ingress must not exceed 16 MiB;
- allows at most 256 `--assume` values, each `column=table`, max 256 bytes;
- `--output` must name a new canonical path; existing targets are refused.
- `schema check`
- `--annotations` and `--reviewed-candidates` are all-or-nothing;
- annotations must be UTF-8, canonical, non-symlink, max 16 MiB;
- `--reviewed-candidates` must match `sha256:<64 lowercase hex>`.
- `schema accept`
- `--run` is mandatory and must be 32 lowercase hex characters;
- `--yes` is mandatory and may be supplied once;
- `--annotations`/`--reviewed-candidates`/`--from-sql`/`--assume` are not accepted.
- Unknown flags, passthrough separators, and shell fragments are rejected before Docker runs.
- `<workspace>-reference` contains `schema_table`, `schema_column`, `schema_relationship`, and
`evidence` records;
- `<workspace>-memory` contains `memory` and `solved_question` records.
A rerun replaces the relevant records in `reference` and leaves `memory` untouched. A workspace
without Evidence is valid and completes with an explicit warning.
The workspace LSH generation is bound to the workspace ID, Catalog database ID, Metadata Content
Revision, effective configuration fingerprint, and input fingerprint. A mismatch cannot reuse a
generation produced for another database or revision.
## Clear lifecycle
`workspace preprocess clear` is intentionally narrower than deleting all semantic data. It:
1. copies any `memory` and `solved_question` records still present in the pre-split workspace
collection to `<workspace>-memory`, then retires that legacy collection (a normal preprocessing
write performs the same one-time cutover if clear was not invoked first);
2. deletes `<workspace>-reference`;
3. removes the active LSH generation, Evidence corpus, private Catalog snapshot, and derived job
checkpoints;
4. records `derived_data_cleared` in PostgreSQL so readiness projects to `required`.
It never deletes `<workspace>-memory`, session artifacts, the workspace repository, Catalog metadata,
or source database data. There is no clear history or rollback. The next complete preprocessing run
recreates the reference collection and all local derived artifacts from their authoritative sources.
## PostgreSQL preprocessing state
PostgreSQL is authoritative for:
- `preprocessing_status`: `running | succeeded | failed`;
- current `metadata_content_revision`;
- last successful `preprocessed_metadata_revision`;
- `preprocessing_input_fingerprint`;
- start/finish timestamps and a bounded failure code.
Every relevant Catalog mutation increments `metadata_content_revision` and marks the prior result
stale in the same transaction. The input fingerprint covers both the immutable workspace Git
revision (including revision-pinned Evidence) and the effective Catalog-derived DWH/semantic
configuration.
Core admission requires all of the following:
- status is `succeeded`;
- processed and current Metadata Content Revision are equal;
- stored and current preprocessing input fingerprints are equal.
Failure leaves the workspace unavailable to the core. There is no rollback: fix the cause and run
the complete command again.
## Administration sidebar
The Administration sidebar exposes the same one-shot operation for the workspace selected in the
composer. It reads `GET /workspaces/:workspaceId/preprocessing`, invokes
`POST /workspaces/:workspaceId/preprocessing` for a run, and invokes
`DELETE /workspaces/:workspaceId/preprocessing` for clear. Both mutations delegate to the complete
workspace preprocessing service and do not expose partial stages.
The compact control has five states: `ready`, `required`, `running`, `blocked`, and `failed`.
**Run again** is available for `ready`, **Run** for `required`, and **Retry** for `failed`. **Clear**
appears to the left of that action whenever replaceable data can be cleared. It opens an inline
confirmation that names both the data removed and the memory retained. All run actions invoke the
same complete, idempotent preprocessing operation. A blocked state
names the current unmet prerequisite and links it to an operator action, but has no run diagnostic
because no preprocessing run started. A failed state shows only the latest bounded diagnostic:
failed stage, safe error code, and finish time. PostgreSQL overwrites that diagnostic when the next
run starts or finishes; there is no preprocessing history or rollback UI. The full service log is
available to operators with `docker compose logs core`.
Once a non-ready state is known, **New session** is disabled. Core admission remains the
authoritative enforcement boundary if the browser has not loaded the state yet.
## Container boundary
`tht` resolves the selected `core` image from the rendered installation, converts it to an immutable local image ID, writes a one-shot final override that pins both `core` and `workspace-maintenance` to that ID with `pull_policy: never`, and runs only:
The host CLI resolves the selected core image to an immutable local image ID and runs only:
```text
docker compose run --rm --no-deps --no-TTY --name <owned-name> workspace-maintenance <fixed-command>
docker compose run --rm --no-deps --no-TTY --name <owned-name> \
workspace-maintenance preprocess-run
docker compose run --rm --no-deps --no-TTY --name <owned-name> \
workspace-maintenance preprocess-clear
```
The request is streamed as one schema-versioned JSON document over stdin. Public stdout is always one schema-versioned JSON result; human mode is rendered from an allowlisted subset of that same result.
The request is one bounded schema-versioned JSON document on stdin. The maintenance process strips
all `THT_CATALOG_*` variables before spawning Python. JSON mode keeps stdout pristine.
## Public JSON result
## Result and exit codes
```json
{
"schemaVersion": 1,
"status": "succeeded|unchanged|dry_run|blocked|failed",
"code": "ok|workspace_not_found|workspace_not_activatable|binding_missing|preprocessing_conflict|preprocessing_resume_mismatch|manual_review_required|evidence_materialization_required|effective_config_mismatch|semantic_index_incompatible|annotation_invalid|egress_policy_refused",
"workspaceId": "abc",
"workspaceRevision": "1234567890abcdef1234567890abcdef12345678",
"descriptorBlob": "sha256:<64 lowercase hex>",
"operation": "inspect|preprocess-dwh|schema-suggest-fks|schema-check|index-schema|preprocess-evidence|preprocess-run",
"runId": "<optional 32hex>",
"childRuns": {"stage": "<optional 32hex>"},
"completedStages": ["stage"],
"counts": {"name": 1},
"artifactIdentities": [{"kind": "fk_candidates", "digest": "sha256:<64 lowercase hex>"}],
"warnings": ["safe warning"]
}
```
The public result includes `schemaVersion`, `status`, `code`, workspace/revision identity,
`operation`, completed stages, safe counts, artifact digests, and the input fingerprint. It never
contains credentials, connection strings, sampled values, SQL, or raw child errors.
`tht --json` parses the operator stdout strictly and re-encodes only the public fields above.
`evidence_materialization_required` is retained for pre-P6 compatibility; since P6, filesystem
Evidence is materialized at activation and preprocesses directly.
## Exit codes
- `0`: `succeeded`, `unchanged`, or `dry_run`
- `3`: `blocked`
- `2`: host-side grammar or local file safety failure
- `1`: operational failure or operator-reported `failed`
- `0`: preprocessing run or clear succeeded;
- `1`: operational or preprocessing failure;
- `2`: invalid host-side invocation.
+15 -13
View File
@@ -112,18 +112,22 @@ flowchart TD
VAL -->|errors or review items| FIX["Author corrections\nand review"]
FIX --> PREP
VAL -->|publishable| COMMIT["Commit del repository\nauthoring clone"]
COMMIT --> ING["tht preprocess evidence\nnormalization and chunking"]
COMMIT --> ING["tht workspace preprocess run\nnormalization and chunking"]
ING --> BM25["BM25 index"]
ING --> VEC["Embeddings and vector store"]
BM25 --> GEN["Candidate generation"]
VEC --> GEN
GEN --> ACT["Active generation"]
GEN --> ACT["Replace active Evidence slice"]
ACT --> RUNTIME["Evidence retrieval in the workflow"]
```
Preparation can restructure changed sources, but it does not publish by itself. `prepare` produces a proposal and can identify the document involved in an error. `validate` does not write or publish. The curator publishes the revision. The runtime reads a complete, validated revision, then the pipeline creates a versioned generation. Activation is atomic, and a previous generation remains available under the retention policy.
Preparation can restructure changed sources, but it does not publish by itself. `prepare` produces a proposal and can identify the document involved in an error. `validate` does not write or publish. The curator publishes the revision. The complete workspace preprocessing command reads that exact revision and replaces the active Evidence slice. There is no application-level rollback; rerun complete preprocessing after correcting a failure.
Runtime retrieval is hybrid. The dense branch uses embeddings, the BM25 branch uses lexical search, and deterministic fusion orders the results. The published unit keeps its provenance, which the model must cite when it uses the Evidence.
Runtime retrieval is hybrid. The dense branch uses embeddings, the BM25 branch uses lexical search,
and deterministic fusion orders the results. Evidence fragments live in the workspace `reference`
collection together with Schema and relationships; runtime Memory and solved questions live in a
separate `memory` collection. The published unit keeps its provenance, which the model must cite when
it uses the Evidence.
## Author responsibilities
@@ -191,17 +195,15 @@ tht evidence evaluate <workspace-root> --config <workspace-config> --generation
tht evidence resolve <workspace-root> evidence:<id> --retire
tht evidence resolve <workspace-root> evidence:<id> --source source/domain/nuovo.md
# Materialize and index a versioned generation.
tht preprocess evidence --config <workspace-config>
# Dry run and resume a job when supported by the configuration.
tht preprocess evidence --config <workspace-config> --dry-run
tht preprocess evidence --config <workspace-config> --resume <run-id>
# Materialize Catalog-derived schema/LSH and the revision-pinned Evidence in one run.
tht --installation <absolute>/thothii-installation.yaml \
workspace preprocess run --workspace <workspace-id>
```
`evidence prepare`, `evidence migrate`, `evidence validate`, and `evidence resolve` require the
repository path. `preprocess evidence` uses the workspace configuration because it needs the
embedding, vector store, retention policy, and artifact directory.
`evidence prepare`, `evidence migrate`, `evidence validate`, and `evidence resolve` are authoring
operations and require the repository path. Runtime publication is available only through the
complete host-side `workspace preprocess run`; there is no public Evidence-only preprocessing
command.
Exit codes are part of the operating contract: `evidence validate` returns `0` when the corpus is publishable, `1` for validation errors, and `3` when only review items or orphaned units remain. With `--json`, stdout must contain valid JSON only.
+7 -2
View File
@@ -100,9 +100,14 @@ Preprocessing runs through the native host CLI and the installation descriptor:
```sh
tht --installation /percorso/assoluto/thothii-installation.yaml \
workspace preprocess evidence --workspace <workspace-id>
workspace preprocess run --workspace <workspace-id>
```
Per rimuovere soltanto gli indici e gli artifact ricostruibili, preservando Memory e domande risolte:
```sh
tht --installation /percorso/assoluto/thothii-installation.yaml \
workspace preprocess dwh --workspace <workspace-id>
workspace preprocess clear --workspace <workspace-id>
```
The CLI runs the profile-gated `workspace-maintenance` service. See the
+9 -9
View File
@@ -1,12 +1,12 @@
# Database management
Database Management is an administrative catalog for an external PostgreSQL schema. It is separate
from workspace preprocessing and, today, does not change the DWH binding used by the NL→SQL
session workflow.
Database Management is the administrative PostgreSQL Metadata Catalog for an external PostgreSQL
schema. Its binding and metadata are the sole database source used by workspace preprocessing and
the NL→SQL session workflow.
## What the catalog owns
For each YAML workspace, an administrator may configure at most one Metadata Catalog binding. It holds
For each workspace identity, an administrator may configure at most one Metadata Catalog binding. It holds
the database name, schema, connection binding, write-only encrypted secrets, observed physical
schema, optional curated descriptions, generated descriptions, and durable operation history.
@@ -170,11 +170,11 @@ starting generation. Changing a flag affects future generations only; existing g
descriptions are not regenerated. Real and substituted samples remain transient and are not persisted
or returned to the browser.
The database-level **Copy generated descriptions to all columns** action applies every non-empty
AI-generated column description to the corresponding curated **Description** field in one atomic
operation. It skips empty generated descriptions, reports copied and skipped counts, and retains the
generated text. Because this can replace reviewed descriptions, the interface requires explicit
confirmation before applying it.
The database-level **Copy generated description to descriptions** action applies every non-empty
AI-generated table and column description to the corresponding curated **Description** field in one
atomic operation. It skips empty generated descriptions, reports aggregate copied and skipped
counts, and retains the generated text. Because this can replace reviewed descriptions, the
interface requires explicit confirmation before applying it.
The decisions behind this surface are [ADRs 0001–0011](../adr/0001-postgres-metadata-catalog.md)
and the detailed acceptance record is
@@ -0,0 +1,221 @@
# Server handoff — preprocessing complete
Questo documento è un'istruzione operativa per il Codex eseguito sul server ThothII. Quando
l'operatore chiede di applicarlo, eseguire i passi nell'ordine indicato e consegnare il report
finale. La release autorizzata è il tag annotato `260906-preprocessing-complete`.
## Obiettivo verificabile
Al termine devono essere vere tutte queste condizioni:
- il checkout server è esattamente il commit puntato dal tag;
- il binario host `tht` proviene dallo stesso checkout;
- la migrazione PostgreSQL del catalogo `013_catalog_preprocessing_state` è terminata con exit 0
prima dell'avvio del nuovo Core;
- immagini `core` e `frontend` sono state ricostruite e i servizi sono healthy;
- il workspace PSD ha completato il preprocessing catalog-driven;
- Schema, relazioni ed Evidence sono in `<workspace>-reference`, mentre `memory` e
`solved_question` restano in `<workspace>-memory`;
- login e nuovo caricamento aprono la superficie Core.
## Confini operativi
- Conservare descriptor, env, secret, catalogo PostgreSQL, Qdrant, Memory, sessioni e repository
workspace esistenti.
- Non stampare descriptor completi, env, secret, connection string o log non sanitizzati.
- Non usare `reset`, rebase, stash, prune, `down --volumes` o cancellazioni manuali.
- Non eseguire `workspace preprocess clear`: l'upgrade richiede soltanto il normale `run`, che è
idempotente, sostituisce i dati derivati e conserva la Memory.
- Fermarsi davanti a worktree sporco, tag non verificabile, branch divergente, descriptor ambiguo,
operazioni attive, migrazione fallita o servizi non healthy. Riportare l'evidenza redatta senza
tentare correzioni distruttive.
## 1. Risolvere i target e fare l'inventario
Posizionarsi nel checkout server e valorizzare percorsi assoluti reali:
```bash
THTII_REPO=$(git rev-parse --show-toplevel)
THTII_INSTALLATION=/percorso/assoluto/thothii-installation.yaml
THTII_RELEASE_TAG=260906-preprocessing-complete
THTII_WORKSPACE_ID=psd-clinical
cd "$THTII_REPO"
```
Se il descriptor attivo non è identificabile univocamente dai comandi già usati sul server,
chiederne il percorso all'operatore. Non cercare secret e non scegliere un file per somiglianza.
Eseguire l'inventario in sola lettura:
```bash
git status --short --branch
git branch --show-current
git rev-parse HEAD
git remote -v
test -f "$THTII_INSTALLATION"
tht --installation "$THTII_INSTALLATION" status
```
Registrare il vecchio hash. Il branch deve essere `main`, il worktree deve essere pulito e `origin`
deve puntare al Gitea autorizzato `https://git.tylconsulting.it/mptyl/ThothII.git`. Verificare che
non siano in corso sessioni Core o preprocessing; in caso contrario fermarsi.
**Completamento:** checkout, descriptor, installation Compose e workspace ID sono identificati
senza ambiguità; il server è inattivo dal punto di vista applicativo e il worktree è pulito.
## 2. Integrare l'esatta release
```bash
cd "$THTII_REPO"
git fetch --prune origin main
git fetch origin tag "$THTII_RELEASE_TAG"
test "$(git cat-file -t "$THTII_RELEASE_TAG")" = tag
THTII_RELEASE_COMMIT=$(git rev-parse "$THTII_RELEASE_TAG^{commit}")
git switch main
git merge --ff-only "$THTII_RELEASE_COMMIT"
test "$(git rev-parse HEAD)" = "$THTII_RELEASE_COMMIT"
git status --short --branch
```
Il merge deve essere fast-forward e il worktree deve restare pulito. Se `origin/main` contiene
commit successivi al tag, non integrarli in questa esecuzione: il tag è il confine della release.
**Completamento:** `HEAD` coincide byte per byte con il commit del tag annotato.
## 3. Installare il CLI della release e validare la configurazione
```bash
cd "$THTII_REPO"
./scripts/install-tht.sh
tht version --json
tht --installation "$THTII_INSTALLATION" update --check-only
```
Leggere strutturalmente dal descriptor i valori `projectDirectory`, `envFile`, `profile` e la lista
ordinata `overrides`, senza mostrare contenuti protetti. Devono descrivere il checkout corrente e il
profilo `server`. Il descriptor e l'env esistenti sono configurazione autorevole: non rigenerarli e
non sostituirli con gli esempi del repository.
Se esiste `.tht/<compose-project>/current-image.yaml`, fermarsi e segnalarlo: un pin immagine
esplicito renderebbe ambiguo il deploy dal checkout.
**Completamento:** il nuovo `tht` è installato, il descriptor corrente è valido e nessun pin
immagine impedisce la ricostruzione.
## 4. Preparare il comando Compose dell'installation
Costruire un array shell `THTII_COMPOSE` con lo stesso ordine usato da `tht`:
1. `docker compose`;
2. `--project-name thothii-<prime 12 cifre sha256 del percorso canonico del descriptor>`;
3. `--project-directory <projectDirectory>`;
4. `--env-file <envFile>`;
5. `-f <projectDirectory>/compose.yaml`;
6. `-f <projectDirectory>/deploy/compose.server.yaml`;
7. ogni override dichiarato, nello stesso ordine;
8. `-f <directory descriptor>/generated/compose.models.yaml`;
9. `-f <projectDirectory>/deploy/compose.auth-runtime-projection.yaml` quando il descriptor ha
`authentication.runtimeProjection`.
Per una installation server con Git SSH e senza altri overlay, la forma è:
```bash
THTII_INSTALLATION=$(realpath "$THTII_INSTALLATION")
THTII_INSTALL_DIR=$(dirname "$THTII_INSTALLATION")
THTII_ENV=/percorso/assoluto/letto-da-envFile
THTII_COMPOSE_PROJECT="thothii-$(printf '%s' "$THTII_INSTALLATION" | sha256sum | cut -c1-12)"
THTII_COMPOSE=(
docker compose
--project-name "$THTII_COMPOSE_PROJECT"
--project-directory "$THTII_REPO"
--env-file "$THTII_ENV"
-f "$THTII_REPO/compose.yaml"
-f "$THTII_REPO/deploy/compose.server.yaml"
-f "$THTII_REPO/deploy/compose.git-ssh.yaml"
-f "$THTII_INSTALL_DIR/generated/compose.models.yaml"
-f "$THTII_REPO/deploy/compose.auth-runtime-projection.yaml"
)
"${THTII_COMPOSE[@]}" config --quiet
```
Adattare l'array agli override realmente dichiarati. Non usare il blocco di esempio se il
descriptor differisce.
**Completamento:** `config --quiet` termina con exit 0 usando esattamente l'identità e i file della
installation già attiva.
## 5. Costruire, migrare e riavviare
La build può avvenire mentre i vecchi container sono ancora in esecuzione: i container mantengono
il loro image ID. Fermare Core e frontend soltanto dopo una build riuscita.
```bash
"${THTII_COMPOSE[@]}" build core frontend
"${THTII_COMPOSE[@]}" stop frontend core
"${THTII_COMPOSE[@]}" up --detach catalog-db qdrant embedding
"${THTII_COMPOSE[@]}" run --rm catalog-migrate
"${THTII_COMPOSE[@]}" run --rm embedding-model-init
tht --installation "$THTII_INSTALLATION" start
tht --installation "$THTII_INSTALLATION" status
tht --installation "$THTII_INSTALLATION" doctor --json
```
`catalog-migrate` deve terminare con exit 0. `embedding-model-init` è one-shot: `exited (0)` è il
suo stato corretto. Se un passo fallisce dopo lo stop, lasciare i container e i volumi disponibili
per diagnosi e raccogliere soltanto:
```bash
tht --installation "$THTII_INSTALLATION" status
tht --installation "$THTII_INSTALLATION" logs
```
**Completamento:** migrazione completata, servizi persistenti healthy e job embedding terminato con
exit 0.
## 6. Eseguire una volta il preprocessing completo
Ispezionare prima le precondizioni:
```bash
tht --installation "$THTII_INSTALLATION" \
workspace inspect --workspace "$THTII_WORKSPACE_ID" --json
```
Lo stato atteso dopo l'upgrade è `required`. Se è `blocked`, correggere soltanto il prerequisito
indicato. Se è `running`, fermarsi perché esiste un'operazione concorrente. Se è `failed`, leggere
il solo ultimo diagnostico e i log sanitizzati del Core prima di ritentare.
Quando le precondizioni sono soddisfatte, eseguire:
```bash
tht --installation "$THTII_INSTALLATION" \
workspace preprocess run --workspace "$THTII_WORKSPACE_ID" --json
tht --installation "$THTII_INSTALLATION" \
workspace inspect --workspace "$THTII_WORKSPACE_ID" --json
```
Il secondo inspect deve riportare il workspace pronto e la revisione metadati processata uguale a
quella corrente. Il run legge tabelle, colonne, descrizioni, sensibilità e foreign key dal Catalogo
PostgreSQL; il workspace YAML fornisce soltanto identità ed eventuale Evidence. Un workspace senza
Evidence è valido. Il run effettua anche il cutover dall'eventuale collezione Qdrant legacy,
preservando `memory` e `solved_question`.
**Completamento:** preprocessing riuscito una volta, Core ammesso e Memory preservata.
## 7. Accettazione e report
Verificare, senza creare una sessione reale se non richiesto dall'operatore:
- `git rev-parse HEAD` uguale a `git rev-parse 260906-preprocessing-complete^{commit}`;
- worktree pulito;
- `tht doctor --json` con esito positivo;
- `catalog-db`, Qdrant, embedding, Core e frontend healthy;
- ultimo `workspace inspect` pronto;
- login o hard refresh posizionati sulla pagina Core;
- **Administration → Preprocessing** mostra `Status: Ready` e **Run again**;
- **New session** è abilitato.
Consegnare un report breve con vecchio hash, nuovo hash, tag, versione `tht`, exit della migrazione,
stato servizi e risultato preprocessing. Redigere URL interni, nomi utente e qualsiasi dato
operativo sensibile.
@@ -629,14 +629,15 @@ tht --installation "$THTII_NEW_INSTALLATION" workspace preprocess run \
--workspace psd-clinical --json
```
Se il preprocess restituisce un run ID interrotto, usare il suo `--resume RUN` soltanto dopo aver
diagnosticato la causa. Non lanciare run paralleli. Il preprocess DWH materializza gli artifact
runtime; il preprocess Evidence ricostruisce l'indice Qdrant con l'embedding fisso. Le nuove sessioni
pinzano la revisione Git attiva del workspace.
Se il preprocessing fallisce, correggere il Catalog o la configurazione e rilanciare lo stesso
comando dall'inizio: non esistono resume o rollback. Il comando completo materializza LSH e vettori
di schema dal Metadata Catalog e ricostruisce la slice Evidence. Le nuove sessioni pinzano la
revisione Git attiva del workspace.
Non copiare nel nuovo descriptor i legacy `harness/workspaces/*.yaml` come catalogo authored e non
trasferire vecchi indici vettoriali incompatibili. Il repository workspace schema v4 definisce
database ed Evidence; binding, segreti e selezione attiva restano locali all'installation.
identità ed Evidence; database, metadati, binding, segreti e selezione attiva restano locali
all'installation.
**Esito richiesto:** `workspace inspect` mostra repository, branch e revisione PSD corretti; test
DWH diretto, schema sync e preprocess completo terminano con successo; Qdrant contiene la nuova
+27 -20
View File
@@ -7,7 +7,7 @@ boundary between what can be published and what can be used by an installation.
| Role | Owns | Does not own |
| --- | --- | --- |
| Curator | `thoth-workspaces.yaml`, `<id>/workspace.yaml`, Evidence, and curated schema annotations | installation secrets or active runtime bindings |
| Curator | `thoth-workspaces.yaml`, `<id>/workspace.yaml`, and Evidence | database metadata, installation secrets, or active runtime bindings |
| Installation operator | Git source, selected workspace, write-only runtime secrets, validation, connectivity, and preprocessing | commits or pushes to the workspace repository |
| Reviewer | NL→SQL decisions in a pinned session | workspace publication or preprocessing |
@@ -20,14 +20,14 @@ Schema v4 is the only accepted workspace descriptor. Schema v1, v2, and v3 works
are rejected before activation. Each catalog entry must have a matching descriptor at
`<id>/workspace.yaml` in the same Git commit. The application validates a complete candidate
revision and activates it atomically; invalid content leaves the preceding active revision in
place.
place. A v4 descriptor contains only workspace identity and optional Evidence configuration; it
does not contain a database or database metadata.
<!-- workspace-descriptor-contract:end -->
<!-- non-workspace-migration:start -->
To convert a v3 descriptor before committing it, set `workspace.schema_version` to `4`, remove
`llm_policy`, and remove `semantic_index`. Database, Evidence, diagnostics, and binding data remain
unchanged. Validate the resulting v4 repository revision before activation; ThothII never rewrites
the curator-owned repository during pull.
Create a clean v4 descriptor with `workspace` and optional `evidence`. Do not import the old DWH,
diagnostics, annotation, `llm_policy`, or `semantic_index` blocks. Configure the database in
Database Management. ThothII never rewrites the curator-owned repository during pull.
<!-- non-workspace-migration:end -->
## Operator sequence
@@ -41,29 +41,36 @@ the curator-owned repository during pull.
connections**. The workspace connection test uses that same current database configuration for
DWH connectivity and also checks the workspace Evidence and installation semantic services.
4. Select it as the installation workspace before creating sessions.
5. Use the host CLI for preprocessing. It dispatches a profile-gated maintenance service and
returns a single structured result; `--json` keeps stdout machine-readable.
5. Expand **Administration** in the right sidebar and run **Preprocessing**. The button is available
when all prerequisites are satisfied: it shows **Run** when preprocessing is required,
**Run again** when the workspace is already current, and **Retry** after a failure. The same
operation is available from the host CLI for unattended administration. **Clear**, immediately
to the left, removes only replaceable Schema/Evidence vectors, LSH, corpus, and checkpoints after
an inline confirmation; it preserves Memory and solved questions. `--json` keeps CLI stdout
machine-readable.
```sh
INSTALLATION=/absolute/path/thothii-installation.yaml
WORKSPACE=example-workspace
tht --installation "$INSTALLATION" workspace inspect --workspace "$WORKSPACE" --json
tht --installation "$INSTALLATION" workspace preprocess dwh --workspace "$WORKSPACE" --json
tht --installation "$INSTALLATION" workspace preprocess evidence --workspace "$WORKSPACE" --json
tht --installation "$INSTALLATION" workspace preprocess run --workspace "$WORKSPACE" --json
tht --installation "$INSTALLATION" workspace preprocess clear --workspace "$WORKSPACE" --json
```
For the full DWH → review → schema-index → Evidence chain, run
`workspace preprocess run`. It may stop with `manual_review_required` when FK candidates need a
curator decision. Publish the reviewed annotations, update the repository, then accept that exact
run and resume it:
The command snapshots tables, columns, descriptions, sensitivity, and active relationships from
PostgreSQL, samples eligible DWH values for LSH, and replaces the schema/Evidence vector slices.
It is rerunnable but not resumable and has no rollback. Catalog sync and description generation
remain separate operations and must already be complete.
```sh
tht --installation "$INSTALLATION" workspace schema accept \
--workspace "$WORKSPACE" --run <run-id> --yes --json
tht --installation "$INSTALLATION" workspace preprocess run \
--workspace "$WORKSPACE" --resume <run-id> --json
```
After clear, the sidebar reports **Required** and the core rejects new sessions until a complete run
succeeds. Clear can be repeated safely: an already absent reference collection or derived path is a
no-op, and the separate Memory collection is never a cleanup target.
The sidebar retains no run history. If the current prerequisite blocks a start, it explains what
must be completed and correctly reports that there is no run log. If the last run failed, it shows
only that run's safe stage, error code, and finish time; use `docker compose logs core` for the
corresponding service log.
The contract gives exact validation, exit code, and JSON rules in
[Workspace preprocessing CLI](../contracts/workspace-preprocessing-cli.md). For Evidence source
@@ -102,6 +102,8 @@ Description and never writes comments to the external Workspace Database.
Database, requested scope, selected model identifier, workspace language, progress counters,
timestamps, and an optional final error summary.
- Ordered Description Generation Events store timestamp, severity, and safe human-readable text.
When an event identifies a target, it uses the object type and qualified physical name, such as
`Column "patients.birth_date"` or `Table "patients"`; catalog UUIDs remain internal identifiers.
No durable per-target jobs, model invocation rows, prompt snapshots, sample snapshots, leases,
heartbeats, registry revisions, or provenance chains are introduced.
- Starting a run schedules an in-process background loop and returns the run immediately. The API
+13 -21
View File
@@ -71,7 +71,7 @@ va creato un corpus Evidence indipendente scollegato dal catalogo dei workspace.
permette di sceglierlo.
- Workspace ID: `psd-evidence-lab`.
- Directory: `psd-evidence-lab/`.
- Qdrant collection: `psd-evidence-lab`.
- Qdrant collections: `psd-evidence-lab-reference` and `psd-evidence-lab-memory`.
- Evidence URI: `psd-evidence-lab/evidence`.
- Database target: lo stesso DWH read-only di PSD, configurato però come binding del nuovo
workspace.
@@ -84,7 +84,7 @@ va creato un corpus Evidence indipendente scollegato dal catalogo dei workspace.
Il repository di authoring usato dall'umano deve essere un clone diverso dal checkout read-only
gestito dall'installazione.
Collection e workspace distinti isolano i dati, ma non i fault ai servizi. REL-05/06/10-13 e le
Collections e workspace distinti isolano i dati, ma non i fault ai servizi. REL-05/06/10-13 e le
prove di outage richiedono un Compose project lab con `dataRoot`, endpoint Qdrant/Ollama, volumi e
porte propri, oppure proxy di fault scoped esclusivamente al lab. Non fermare né corrompere i servizi
condivisi con PSD. Se questo isolamento non è disponibile, la campagna fault va marcata `NOT RUN` e
@@ -332,20 +332,11 @@ Con `--json`, stdout deve contenere un solo JSON valido.
```bash
"$OPERATOR_THT" --installation "$INSTALLATION" workspace inspect \
--workspace "$WS" --json
"$OPERATOR_THT" --installation "$INSTALLATION" workspace vector inspect \
--workspace "$WS" --json
"$OPERATOR_THT" --installation "$INSTALLATION" workspace preprocess evidence \
--workspace "$WS" --dry-run --json
"$OPERATOR_THT" --installation "$INSTALLATION" workspace preprocess evidence \
"$OPERATOR_THT" --installation "$INSTALLATION" workspace preprocess run \
--workspace "$WS" --json
```
Per un resume controllato:
```bash
"$OPERATOR_THT" --installation "$INSTALLATION" workspace preprocess evidence \
--workspace "$WS" --resume <run-id-32hex> --json
```
Il comando è completo, sovrascrivibile e non supporta dry-run, resume o rollback.
### 9.3 Probe tipizzata tecnica, non superficie host
@@ -654,8 +645,9 @@ separate e non possono essere dichiarate coperte dal solo PASS dell'evaluation.
## 16. Inventario vettoriale corretto
`workspace vector inspect` verifica oggi collection, dimensioni e distanza; non è un inventario dei
record. Per il test esaustivo l'agente deve fornire un helper read-only con questa interfaccia minima:
`workspace inspect` espone lo stato operativo del preprocessing, ma non è un inventario dei record
vettoriali. Per il test esaustivo l'agente deve fornire un helper read-only con questa interfaccia
minima:
```bash
"$EVIDENCE_INVENTORY" capture --installation "$INSTALLATION" --workspace "$WS" \
@@ -824,16 +816,16 @@ nel collaudo manuale.
Prima di G1 e dopo G6 confrontare:
- punti e nearest-neighbour canary di `schema_table` e `schema_column`;
- punti e nearest-neighbour canary di `schema_table`, `schema_column` e `schema_relationship`;
- punti `memory` e `solved_question` eventualmente presenti;
- dimensioni, distanza e dense unnamed della collection;
- query di controllo in `psd-clinical`;
- accesso DWH rigorosamente read-only;
- settings globali ripristinati esattamente al baseline e sessioni PSD non modificate.
`delete_generation`/retention Evidence non deve cancellare altri record kind. `vector rebuild
--destroy` non fa parte del percorso nominale; se usato per provare restore/cleanup deve puntare
alla collection lab, ripetere esattamente il suo nome nei guard e avere autorizzazione esplicita.
`delete_generation`/retention Evidence non deve cancellare altri record kind. Il preprocessing
completo può sostituire soltanto le slice derivate di schema ed Evidence; non deve cancellare
`memory` o `solved_question`.
## 20. Adapter e kind estesi (P2)
@@ -906,8 +898,8 @@ verbale se falliscono:
invariato e causare un full re-embedding;
- se una document generation viene riusata ma i suoi punti esistono solo sotto la revisione vecchia,
il filtro Qdrant revision-pinned può renderla irrecuperabile: inventory e retrieval devono fallire;
- `workspace vector rebuild` può ricreare il dense senza ripristinare subito BM25: non usarlo come
recovery nominale e, se testato, verificare il contratto completo dopo il rebuild;
- non esiste un rebuild Qdrant pubblico separato: il ripristino nominale è sempre una nuova
esecuzione completa di `workspace preprocess run`;
- il writer lock/conflitto va provato live, perché la sola presenza delle primitive nei test non
dimostra che il percorso produttivo le acquisisca;
- non esiste un comando pubblico per inventory, elenco job o riattivazione di una generazione