docs: reorganize operational documentation

This commit is contained in:
Codex
2026-08-29 20:08:48 +02:00
parent d504b1def1
commit 0ce05869cf
10 changed files with 371 additions and 337 deletions
+67
View File
@@ -0,0 +1,67 @@
# Database management
Database Management is an administrative catalog for an external PostgreSQL schema. It is separate
from workspace preprocessing and, today, does not change the DWH binding used by the NL→SQL
session workflow.
## What the catalog owns
For each YAML workspace, an administrator may create at most one database configuration. It holds
the database name, schema, connection binding, write-only encrypted secrets, observed physical
schema, optional curated descriptions, generated descriptions, and durable operation history.
It does **not** become the external source of truth. Tables, columns, types, defaults,
nullability, primary-key positions, and ordered foreign-key pairs are observations of the source
schema and cannot be manually created, renamed, or structurally edited. Descriptions are the
editable metadata.
## Configure and test a database
1. Open **Database Management** and choose a workspace.
2. Create its PostgreSQL configuration. Choose `postgres_direct`, `rest_api`, or `ssh_tunnel` and
complete the binding fields that the chosen transport requires.
3. Enter secrets only when replacing them. They remain write-only and are never returned by the
application.
4. Run **Test connection** before any synchronization.
SSH uses a private key, optional key passphrase, mandatory `known_hosts`, and optional PostgreSQL
TLS CA/server name. REST prefers `POST /rpc/schema_snapshot`; when it is absent, the catalog may
use the same strict v1 snapshot through one read-only `POST /rpc/run_query`. An unavailable
capability, malformed snapshot, or connector error applies no catalog changes. See the
[schema snapshot contract](../contracts/catalog-schema-snapshot.md).
## Synchronize authoritative schema metadata
Start a synchronization from a database or a selected table set. The available scopes are tables,
columns, relationships, and all. One database can have only one active catalog operation at a
time; cleanup shares this exclusion.
The run scans first and publishes a durable operation. If it detects a destructive difference, it
requires confirmation and re-scans before applying. You can cancel before apply; completed and
failed runs remain in history. The live log is delivered over SSE with a polling fallback.
Explicit cleanup is different from source synchronization: administrators can clear selected
table/relationship or column/relationship catalog metadata without changing the external source,
the connection binding, or secrets. Deleting a table cascades to its columns and relationships.
## Generate and consolidate descriptions
Generated descriptions can be requested for selected tables, selected columns, every eligible
target, or targets with a missing generated description. The backend accepts one installation-wide
run and processes targets sequentially. It reads at most five source rows and five representative
non-null values per relevant source through a read-only connector, then sends that bounded sample
transiently to the configured model provider.
Each successful result is persisted immediately. Stop terminates the active helper but retains
earlier results. A helper has at most one provider retry; three consecutively exhausted technical
batches fail the run. Stale queued/running work is marked interrupted at startup and can be
unlocked only when no local worker/helper is live. There is no automatic resume and no public
description-generation CLI.
Review generated text before copying it into the curated **Description** field. The sampling rule
is a deliberate data-disclosure boundary: do not use this facility for fields whose values must
not be sent to the configured provider until a Sensitive Data Policy is in place.
The decisions behind this surface are [ADRs 0001–0010](../adr/0001-postgres-metadata-catalog.md)
and the detailed acceptance record is
[AI catalog description generation acceptance](../testing/2026-08-29-ai-catalog-description-generation-acceptance.md).
+65
View File
@@ -0,0 +1,65 @@
# Workspace operations
A workspace is curator-owned Git content plus installation-local runtime bindings. It is the
boundary between what can be published and what can be used by an installation.
## Roles and ownership
| Role | Owns | Does not own |
| --- | --- | --- |
| Curator | `thoth-workspaces.yaml`, `<id>/workspace.yaml`, Evidence, and curated schema annotations | installation secrets or active runtime bindings |
| Installation operator | Git source, selected workspace, write-only runtime secrets, validation, connectivity, and preprocessing | commits or pushes to the workspace repository |
| Reviewer | NL→SQL decisions in a pinned session | workspace publication or preprocessing |
The root catalog has `schema_version: 1` and an ordered list of workspace identities. Each entry
must have a matching schema-v3 descriptor at `<id>/workspace.yaml` in the same Git commit. The
application validates a complete candidate revision and activates it atomically; invalid content
leaves the preceding active revision in place.
## Operator sequence
1. Curate and push a complete repository revision. Do not put DWH passwords, API keys, private
keys, or signed URLs in this repository.
2. In the application, update the workspace repository. This fetches and validates the candidate;
it never edits the remote repository.
3. Select the workspace. Supply or replace its write-only runtime secrets, then run **Validate
workspace source** and **Test workspace connections**.
4. Select it as the installation workspace before creating sessions.
5. Use the host CLI for preprocessing. It dispatches a profile-gated maintenance service and
returns a single structured result; `--json` keeps stdout machine-readable.
```sh
INSTALLATION=/absolute/path/thothii-installation.yaml
WORKSPACE=example-workspace
tht --installation "$INSTALLATION" workspace inspect --workspace "$WORKSPACE" --json
tht --installation "$INSTALLATION" workspace preprocess dwh --workspace "$WORKSPACE" --json
tht --installation "$INSTALLATION" workspace preprocess evidence --workspace "$WORKSPACE" --json
```
For the full DWH → review → schema-index → Evidence chain, run
`workspace preprocess run`. It may stop with `manual_review_required` when FK candidates need a
curator decision. Publish the reviewed annotations, update the repository, then accept that exact
run and resume it:
```sh
tht --installation "$INSTALLATION" workspace schema accept \
--workspace "$WORKSPACE" --run <run-id> --yes --json
tht --installation "$INSTALLATION" workspace preprocess run \
--workspace "$WORKSPACE" --resume <run-id> --json
```
The contract gives exact validation, exit code, and JSON rules in
[Workspace preprocessing CLI](../contracts/workspace-preprocessing-cli.md). For Evidence source
forms and the schema-v3 descriptor contract, see
[Workspace Evidence v3](../contracts/workspace-evidence-v3.md).
## Transport and revision rules
Runtime sessions support direct PostgreSQL and REST bindings. SSH tunnel bindings are diagnostic
only for this path, so they cannot admit an NL→SQL session. Database Management has its own
strict known-host SSH path for connection tests and schema synchronization.
Every new session pins the active Git revision. Snapshot cleanup retains revisions still
referenced by unarchived sessions. A later pull can prepare a future session but cannot alter a
resume.