3.9 KiB
Database management
Database Management is an administrative catalog for an external PostgreSQL schema. It is separate from workspace preprocessing and, today, does not change the DWH binding used by the NL→SQL session workflow.
What the catalog owns
For each YAML workspace, an administrator may create at most one database configuration. It holds the database name, schema, connection binding, write-only encrypted secrets, observed physical schema, optional curated descriptions, generated descriptions, and durable operation history.
It does not become the external source of truth. Tables, columns, types, defaults, nullability, primary-key positions, and ordered foreign-key pairs are observations of the source schema and cannot be manually created, renamed, or structurally edited. Descriptions are the editable metadata.
Configure and test a database
- Open Database Management and choose a workspace.
- Create its PostgreSQL configuration. Choose
postgres_direct,rest_api, orssh_tunneland complete the binding fields that the chosen transport requires. - Enter secrets only when replacing them. They remain write-only and are never returned by the application.
- Run Test connection before any synchronization.
SSH uses a private key, optional key passphrase, mandatory known_hosts, and optional PostgreSQL
TLS CA/server name. REST prefers POST /rpc/schema_snapshot; when it is absent, the catalog may
use the same strict v1 snapshot through one read-only POST /rpc/run_query. An unavailable
capability, malformed snapshot, or connector error applies no catalog changes. See the
schema snapshot contract.
Synchronize authoritative schema metadata
Start a synchronization from a database or a selected table set. The available scopes are tables, columns, relationships, and all. One database can have only one active catalog operation at a time; cleanup shares this exclusion.
The run scans first and publishes a durable operation. If it detects a destructive difference, it requires confirmation and re-scans before applying. You can cancel before apply; completed and failed runs remain in history. The live log is delivered over SSE with a polling fallback.
Explicit cleanup is different from source synchronization: administrators can clear selected table/relationship or column/relationship catalog metadata without changing the external source, the connection binding, or secrets. Deleting a table cascades to its columns and relationships.
Generate and consolidate descriptions
Generated descriptions can be requested for selected tables, selected columns, every eligible target, or targets with a missing generated description. The backend accepts one installation-wide run and processes targets sequentially. It reads at most five source rows and five representative non-null values per relevant source through a read-only connector, then sends that bounded sample transiently to the configured model provider.
Each successful result is persisted immediately. Stop terminates the active helper but retains earlier results. A helper has at most one provider retry; three consecutively exhausted technical batches fail the run. Stale queued/running work is marked interrupted at startup and can be unlocked only when no local worker/helper is live. There is no automatic resume and no public description-generation CLI.
Review generated text before copying it into the curated Description field. The sampling rule is a deliberate data-disclosure boundary: do not use this facility for fields whose values must not be sent to the configured provider until a Sensitive Data Policy is in place.
The decisions behind this surface are ADRs 0001–0010 and the detailed acceptance record is AI catalog description generation acceptance.