Files
ThothII/docs/operations/database-management.md
T
Codex 076c9742c5 feat: consolidate database management work
Add catalog-owned logical relationships and runtime snapshots, extend the database-management UI and validation coverage, and document the updated operational workflow.

Keep active sensitive-generation status in a tooltip and indicator, and update the layout E2E to follow the history action in its new database-scoped location.
2026-09-01 14:46:55 +02:00

121 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Database management
Database Management is an administrative catalog for an external PostgreSQL schema. It is separate
from workspace preprocessing and, today, does not change the DWH binding used by the NL→SQL
session workflow.
## What the catalog owns
For each YAML workspace, an administrator may configure at most one Metadata Catalog binding. It holds
the database name, schema, connection binding, write-only encrypted secrets, observed physical
schema, optional curated descriptions, generated descriptions, and durable operation history.
It does **not** become the external source of truth. Tables, columns, types, defaults,
nullability, primary-key positions, and ordered foreign-key pairs are observations of the source
schema and cannot be manually created, renamed, or structurally edited. Descriptions are the
editable metadata.
## Navigate the Fleet Ledger surface
Database Management opens Fleet Ledger inside the normal application shell. Only one data grid is
shown at a time: choose a database to see its tables, choose a table to see its columns, or open the
database's relationships view. Use the emphasized back control or breadcrumb to return to the parent
grid.
The KPI strip reports tables, columns, sensitive columns, relationships, and description coverage.
It uses `GET /catalog/metrics` without `databaseId` for installation totals and with `databaseId` for
the current database. Choose a selection-scoped operation from the action selector and then press
**Run**; unavailable operations remain listed with an explanation. Row-specific actions are the icon
controls in the final column, and each navigation or action icon has an immediate conceptual tooltip.
An unconfigured workspace exposes **Configure catalog** directly on its row; there is no global
database-creation action and the selected workspace cannot be changed in the configuration form.
The master grid keeps three independent states visible:
- **Revision / Evidence** comes from the active immutable workspace revision. Filesystem Evidence is
materialized with that revision; remote Evidence is reported as configured-but-unverified or as
requiring credentials.
- **NL→SQL runtime** is calculated from the workspace DWH/Evidence requirements and runtime secret
store. It also reports transports, such as SSH, that are diagnostic-only and unsupported by sessions.
- **Metadata Catalog** reports whether the installation-local catalog configuration exists, then shows
its separately versioned connection-test or synchronization state.
Configuration, object details, metadata editors, synchronization history, description history,
sensitive-field review, and suggestion-run history open in right-side drawers backed by the
production catalog APIs.
Closing a history drawer does not cancel a durable background run. Existing permission checks,
dirty/busy navigation guards, stale-state handling, and write-only secret behavior continue to
apply.
For temporary comparison in development or staging, add `?db-ui=legacy`; the parameter is honored
only by Vite development or an environment explicitly configured with
`VITE_DB_MANAGEMENT_LEGACY=true`. The separate prototype on port `5173` is not the application and
remains available only until the integrated Fleet Ledger surface passes owner acceptance.
## Configure and test a database
1. Open **Database Management** and find the repository workspace marked **Not configured**.
2. Choose **Configure catalog** on that row. Configure its PostgreSQL catalog binding with
`postgres_direct`, `rest_api`, or `ssh_tunnel` and
complete the binding fields that the chosen transport requires.
3. Enter secrets only when replacing them. They remain write-only and are never returned by the
application.
4. Run **Test connection** before any synchronization.
SSH uses a private key, optional key passphrase, mandatory `known_hosts`, and optional PostgreSQL
TLS CA/server name. REST prefers `POST /rpc/schema_snapshot`; when it is absent, the catalog may
use the same strict v1 snapshot through one read-only `POST /rpc/run_query`. An unavailable
capability, malformed snapshot, or connector error applies no catalog changes. See the
[schema snapshot contract](../contracts/catalog-schema-snapshot.md).
## Synchronize authoritative schema metadata
Start a synchronization from a database or a selected table set. The available scopes are tables,
columns, relationships, and all. One database can have only one active catalog operation at a
time; cleanup shares this exclusion.
The run scans first and publishes a durable operation. If it detects a destructive difference, it
requires confirmation and re-scans before applying. You can cancel before apply; completed and
failed runs remain in history. The live log is delivered over SSE with a polling fallback.
Explicit cleanup is different from source synchronization: administrators can clear selected
table/relationship or column/relationship catalog metadata without changing the external source,
the connection binding, or secrets. Deleting a table cascades to its columns and relationships.
## Generate and consolidate descriptions
Generated descriptions can be requested for selected tables, selected columns, every eligible
target, or targets with a missing generated description. The backend accepts one installation-wide
run and processes targets sequentially. Every catalog column has a **Sensitive** flag, which defaults
to `false`, including after a newly discovered column is synchronized. Before generation, an
administrator can ask the configured model to suggest flags from structural metadata only (database,
schema, table and column names, data types, nullability, primary keys, and foreign keys). Suggestions
remain an unsaved draft until a human reviews and saves them.
The page exposes separate histories for description generation and sensitive-field suggestions.
Sensitive-suggestion history stores the selected model, scope, status, aggregate counts, timestamps,
and sanitized events. It does not store the proposed per-column flags, prompts, raw model output, or
provider diagnostics; closing an unsaved review still discards that draft.
For a column with `sensitive=false`, the worker may read at most five source rows and five
representative non-null values through a read-only connector. For `sensitive=true`, the source query
does not request that column's values; deterministic plausible values derived only from its name and
type take their place in the model prompt. The prompt does not identify those values as synthetic, so
the model can still describe the field as if it had received representative data.
Each successful result is persisted immediately. Stop terminates the active helper but retains
earlier results. A helper has at most one provider retry; three consecutively exhausted technical
batches fail the run. Stale queued/running work is marked interrupted at startup and can be
unlocked only when no local worker/helper is live. There is no automatic resume and no public
description-generation CLI.
Review generated text before copying it into the curated **Description** field. Because the flag
defaults to `false`, an administrator must review the classification and mark protected fields before
starting generation. Changing a flag affects future generations only; existing generated or curated
descriptions are not regenerated. Real and substituted samples remain transient and are not persisted
or returned to the browser.
The decisions behind this surface are [ADRs 0001–0011](../adr/0001-postgres-metadata-catalog.md)
and the detailed acceptance record is
[AI catalog description generation acceptance](../testing/2026-08-29-ai-catalog-description-generation-acceptance.md).