feat: consolidate database management work

Add catalog-owned logical relationships and runtime snapshots, extend the database-management UI and validation coverage, and document the updated operational workflow.

Keep active sensitive-generation status in a tooltip and indicator, and update the layout E2E to follow the history action in its new database-scoped location.
This commit is contained in:
Codex
2026-09-01 14:46:55 +02:00
parent f586152636
commit 076c9742c5
73 changed files with 6966 additions and 610 deletions
+25 -11
View File
@@ -1,6 +1,6 @@
# ThothII — Project State
Last updated: 2026-08-27.
Last updated: 2026-08-31.
This file is the short operational snapshot. Stable commands and the architecture mental model
live in `AGENTS.md`; current design and runtime contracts live under `docs/architecture/`,
@@ -90,6 +90,18 @@ relationships. Curated and generated descriptions are editable; generated descri
and Database Management can generate or consolidate them for selected tables, selected columns,
all targets, or only targets whose Generated Description is missing.
Relationship Management is now reachable directly from each configured Fleet database. One
Relationship Map shows read-only Physical Relationships together with Generated and Manual Logical
Relationships, with Active, Excluded, and All filters. Administrators can add a single-column
relationship, run deterministic name/PK/type inference, exclude or restore a logical relationship,
or delete it permanently. Exclusion retains a tombstone that a rebuild cannot reactivate; permanent
deletion allows a later rebuild to infer the same endpoints again. Inference uses no LLM, embedding,
or source values. It supports normalized table-qualified names, unique non-generic PK names,
composite-PK source columns, and the `*time_key -> dim_time.<single PK>` warehouse convention while
ignoring bare generic names. Explicit table/column metadata cleanup remains a destructive boundary: it removes
the attached logical relationships and exclusions and requires a full schema synchronization before
inference or runtime publication can continue.
The previous Database Management renderer remains a temporary comparison fallback for development
and staging only: `?db-ui=legacy` is honored in Vite development or when
`VITE_DB_MANAGEMENT_LEGACY=true`; it is not a production presentation. The standalone Fleet Ledger
@@ -115,13 +127,15 @@ SSH is not yet enabled for NL→SQL session runtime.
The catalog runs in the internal `catalog-db` PostgreSQL service. Kysely migrations are an explicit
one-shot `catalog-migrate` operation; `scripts/run-stack.sh` runs it before local startup. Runtime
sessions still consume the existing workspace configuration in this slice: database-management
records do not yet change the NL→SQL handoff. The accepted design is recorded in
sessions now consume the Catalog's active effective relationship map through an immutable JSON
snapshot tied to the runtime-config lease. The harness uses that snapshot as its exclusive
relationship source while retaining Git-pinned annotations for descriptive metadata; legacy
runtimes without a snapshot keep the previous merge behavior. The accepted design is recorded in
`docs/plans/2026-08-26-metadata-catalog-from-thothai.md`, the snapshot contract under
`docs/contracts/`, and ADRs 0001–0011.
`docs/contracts/`, and ADRs 0001–0012.
Semantic aliases, value descriptions, synonyms, concepts, and logical relationships remain
deferred to their dedicated slices.
Semantic aliases, value descriptions, synonyms, and concepts remain deferred to their dedicated
slices.
AI Description Generation uses the catalog's human-owned Sensitive Data Flag. The flag defaults to
`false`, including for newly synchronized columns. An administrator may request an AI proposal based
@@ -158,11 +172,11 @@ credential. The Python client supplies only its fixed non-secret compatibility p
The AritmoLab entry also sets `disableThinking: true`, mapped to the endpoint's chat-template flag,
because its default reasoning prose would violate the worker's exact JSON response contract.
Integration of the completed metadata catalog with core schema-linking is explicitly deferred
until the database, table, column, relationship, and synchronization slices are complete. At that
point the next required design gate is to compare the catalog snapshot with the current DWH
preprocessing/schema-linking contracts and plan the cutover; this follow-up must not be treated as
optional cleanup or silently omitted.
Logical relationship integration with core schema-linking is complete: session creation and resume
materialize the active physical/generated/manual map, retrieval-pack generation and Pi receive the
same runtime config, and snapshot validation fails closed on a declared missing, invalid, or orphaned
endpoint. Broader publication of other Catalog metadata to schema-linking remains a separate future
slice.
**Deferred follow-up — Sensitive Data Policy in schema-linking.** The policy is first delivered
and tested in catalog description generation. Its enforcement for core schema-linking remains