feat: implement metadata catalog database management

This commit is contained in:
Codex
2026-08-27 22:43:54 +02:00
parent 705af3aeb2
commit 79c4c925b5
86 changed files with 12566 additions and 135 deletions
+49 -3
View File
@@ -1,6 +1,6 @@
# ThothII — Project State
Last updated: 2026-08-26.
Last updated: 2026-08-27.
This file is the short operational snapshot. Stable commands and the architecture mental model
live in `AGENTS.md`; current design and runtime contracts live under `docs/architecture/`,
@@ -16,8 +16,9 @@ frontend (React/SSE) → backend (Fastify) → pi --mode rpc → tht/harness →
```
The harness owns the deterministic eight-phase NL→SQL workflow and all session persistence.
The backend is a process/RPC/SSE bridge without a database of its own. The frontend renders the
review gates and keeps the live transcript in memory. See
The backend remains a process/RPC/SSE bridge for sessions and now also owns an isolated PostgreSQL
metadata catalog for administrative database configuration. The frontend renders the review gates
and keeps the live transcript in memory. See
`docs/architecture/components.md` for the detailed component and data-flow map.
## Evidence restructuring — accepted
@@ -62,6 +63,51 @@ Workspace descriptors use schema v3. For PSD, workspace content and runtime root
separate uncommitted repository `/Users/mp/projects/tht-workspace-psd`. Secrets remain outside
Git and are supplied only through installation-local protected files.
## Database management
The database, table, and authoritative physical-schema catalog slices are implemented. Database
management opens a responsive AG Grid master-detail surface, lists every YAML workspace, creates
at most one PostgreSQL database configuration per workspace, edits direct PostgreSQL, REST API, or
SSH-tunnel installation bindings, replaces write-only encrypted secrets, and tests supported
connector bindings.
Configured databases use pure hierarchical navigation through `Overview`, `Tables`, and
`Relationships`; a selected table has `Overview` and `Columns`. Physical membership, source
comments, column types/default/nullability/PK positions, and constraint-level ordered FK pairs are
immutable projections of the external schema. Curated and generated descriptions are editable;
generated descriptions start null and AI generation/consolidation is deferred.
Schema refresh is one durable asynchronous engine with database-table, database-column,
selected-table-column, relationship, and full-database actions. Database-level menus expose the
table, all-column, relationship, and full scopes separately; selecting tables exposes column
synchronization for that subset. Runs have one-active-job-per-database exclusion, leases and
restart recovery, atomic apply, destructive-diff confirmation with re-scan, cancellation before
apply, retained history, and a live SSE log with polling fallback. Null metadata renders blank
rather than as a placeholder.
Direct PostgreSQL and strict known-host-verified OpenSSH use `pg_catalog`. REST bindings use the
typed full-snapshot `POST /rpc/schema_snapshot` contract when available. Servers such as the
current PSD endpoint that exposes only `POST /rpc/run_query` use one catalog-owned read-only query
to return the exact same strict v1 snapshot in a single round trip. Both paths remain fail-closed:
an absent capability, query error, partial result, or invalid snapshot applies no catalog changes.
SSH is not yet enabled for NL→SQL session runtime.
The catalog runs in the internal `catalog-db` PostgreSQL service. Kysely migrations are an explicit
one-shot `catalog-migrate` operation; `scripts/run-stack.sh` runs it before local startup. Runtime
sessions still consume the existing workspace configuration in this slice: database-management
records do not yet change the NL→SQL handoff. The accepted design is recorded in
`docs/plans/2026-08-26-metadata-catalog-from-thothai.md`, the snapshot contract under
`docs/contracts/`, and ADRs 0001–0007.
Semantic aliases, value descriptions, synonyms, concepts, AI metadata generation/consolidation,
and logical relationships remain deferred to their dedicated slices.
Integration of the completed metadata catalog with core schema-linking is explicitly deferred
until the database, table, column, relationship, and synchronization slices are complete. At that
point the next required design gate is to compare the catalog snapshot with the current DWH
preprocessing/schema-linking contracts and plan the cutover; this follow-up must not be treated as
optional cleanup or silently omitted.
## Active deployment work and manual gates
### PSD server deployment program