feat: add AI catalog description generation

This commit is contained in:
Codex
2026-08-29 16:42:56 +02:00
parent b0afba81ca
commit 376dd5a09d
76 changed files with 14860 additions and 102 deletions
@@ -0,0 +1,32 @@
---
status: accepted
---
# Use one sequential description-generation run
AI description generation is an infrequent, administrator-triggered catalog operation. The
backend therefore owns one installation-wide asynchronous run and processes its model requests
sequentially. A second generation start is rejected while that run is active. The existing
catalog-operation coordinator also reserves the target Workspace Database so synchronization,
cleanup, and catalog edits cannot overlap the run.
Each model completion is performed by a short-lived internal Python process using LiteLLM. The
backend sends a structured request on stdin, reads pristine structured output from stdout, and
keeps diagnostics on stderr. This helper is neither an HTTP service nor a user-facing CLI. Using
Pi for this operation would couple a deterministic batch task to interactive session lifecycle and
gate behavior without adding product value; the helper reuses Python already present in the core
image while keeping the integration small.
Installation model entries normally reference a protected API key. A reference may be omitted
only for an explicit unauthenticated endpoint; the helper uses a fixed non-secret client
placeholder because OpenAI-compatible SDKs require a non-empty client value even when the server
ignores authentication.
For an explicitly configured Qwen-compatible endpoint, `disableThinking: true` maps to the narrow
chat-template option that prevents reasoning prose from surrounding the required JSON result.
Persistence is deliberately limited to one run record, ordered text events, and the Generated
Description written to each target as soon as it succeeds. There are no durable per-target jobs,
leases, invocation records, automatic resume, or distributed locks. On backend startup, any run
still recorded as queued or running becomes interrupted. An administrator continues by starting
Generate Missing, and may use Unlock only when no live generation process exists. Cancellation
stops the loop and terminates the current helper process.
@@ -0,0 +1,15 @@
---
status: accepted
---
# Allow bounded real source samples for description generation
Description generation may send up to five real rows and up to five distinct non-null example
values for relevant columns to the configured model provider, preserving the useful behavior of
ThothAI. Samples are read only for the current request, treated as untrusted data, bounded before
prompt construction, and never stored in generation runs, events, or catalog metadata.
This first slice makes that behavior explicit in the UI and operator documentation. A follow-up
Sensitive Data Policy is required to classify protected fields and exclude or anonymize their
values before model calls. Until that policy exists, administrators must regard generation as a
controlled disclosure of sampled source data to the selected provider.
+1 -1
View File
@@ -33,7 +33,7 @@ The production role expansion from `backend/src/auth/config.ts` is exact:
| Role | Permissions |
|---|---|
| `user` | `session.use` |
| `admin` | `session.use`, `session.read_all`, `session.manage_all`, `settings.manage`, `workspace.manage`, `workspace.secrets.manage`, `pi.manage`, `auth.diagnostics.read` |
| `admin` | `session.use`, `session.read_all`, `session.manage_all`, `settings.manage`, `workspace.manage`, `workspace.secrets.manage`, `database.manage`, `pi.manage`, `auth.diagnostics.read` |
`admin` therefore includes the ordinary `session.use` permission. No other role or permission
label is part of the production catalog.
+18 -2
View File
@@ -22,7 +22,9 @@ flowchart LR
THT --> VDB["Qdrant / vector store"]
BE --> CFG["settings.json\nworkspace registry"]
BE --> CAT["catalog-db\nPostgreSQL + Kysely"]
BE -->|catalog Test + Table Sync| DWH
BE -->|catalog Test + Sync + bounded AI sampling| DWH
BE -->|one request per subprocess| LLMHELPER["LiteLLM helper\nPython, short-lived"]
LLMHELPER -->|configured model| PROVIDER["AI provider"]
FE -.->|renders widgets| EXT
```
@@ -31,7 +33,7 @@ Dipendenze principali:
| Module | Depends on | Responsibility |
| --- | --- | --- |
| `frontend/` | Backend REST and SSE APIs | UI, gate widgets, and in-memory transcript |
| `backend/src/` | Pi, `tht`, configuration, workspace registry, catalog PostgreSQL, and read-only DWH connectors | Transport, session lifecycle, catalog CRUD, connection tests, table introspection, and APIs |
| `backend/src/` | Pi, `tht`, configuration, workspace registry, catalog PostgreSQL, read-only DWH connectors, and the internal LiteLLM helper | Transport, session lifecycle, catalog CRUD, connection tests, table introspection, sequential AI description generation, and APIs |
| `harness/.pi/` | Pi and `tht phase` | Workflow orchestration and human-in-the-loop gates |
| `harness/tht/` | Filesystem, DWH, and vector store | Persistence, CLI, Evidence, schema, and preprocessing |
| workspace repository | `source/`, `curated/`, manifest, and artifacts | Versioned Evidence source and session output |
@@ -65,6 +67,20 @@ sequenceDiagram
The backend uses `ThtRunner` for CLI subprocesses, `PiProcessManager` for one Pi process per session, `SessionBridge` to adapt RPC events, and `SseHub` to distribute them to clients.
## Catalog description generation
Catalog description generation is a backend-owned administrative operation, separate from the
Pi session workflow and from the public `tht` CLI. The frontend starts one run for selected catalog
tables or columns. A single installation-wide worker processes targets sequentially, reads at most
the configured bounded sample from the source DWH through its read-only connection, and invokes a
short-lived Python LiteLLM helper once per target. The selected model comes from the installation
descriptor; its API key remains in the protected installation secret bundle.
Each result is written immediately to `Generated Description`. Run state and sanitized activity
events are stored in `catalog-db` and exposed to the drawer through REST and SSE. An administrator
may later copy selected generated descriptions into `Description`. There is no parallel run queue,
automatic retry policy, or second orchestration subsystem.
## Main backend classes
The diagram shows the classes that form the bridge between the browser, Pi, and `tht`. Fastify routes receive requests and delegate to these services.
+12 -3
View File
@@ -12,6 +12,8 @@ flowchart LR
USER["Reviewer"] --> FE["Frontend\nReact and SSE"]
FE --> BE["Backend\nFastify"]
BE --> CATALOG["Metadata catalog\nPostgreSQL"]
BE --> MODEL["Configured AI model\nvia short-lived LiteLLM helper"]
BE -->|bounded read-only samples| DWH
BE --> PI["Pi\nRPC per sessione"]
PI --> THT["tht and harness\nworkflow and persistence"]
THT --> DWH["DWH\nread only"]
@@ -50,12 +52,19 @@ A session is a directory under `sessions/` (the workspace defines the path): `se
configurations stored in PostgreSQL through Kysely.
- `CatalogTableService` reconciles persisted Catalog Tables with a successful external schema scan;
`ConcreteCatalogTableIntrospector` isolates direct PostgreSQL, typed REST, and SSH-tunnel access.
- `DescriptionGenerationWorker` admits one installation-wide run and processes its table or column
targets sequentially.
- `PostgresDescriptionSourceSampler` reads bounded real rows and five representative examples from
the configured DWH connection; `ProcessModelCompleter` invokes the short-lived Python LiteLLM
helper with the installation-selected model.
Application settings remain in `backend/data/settings.json`; session state remains in harness phase
documents. PostgreSQL stores only the administrative database catalog, bindings, observed tables,
and curated descriptions. Connector secrets remain write-only in the encrypted workspace secret
store. Catalog SSH support is limited to connection tests and table synchronization; it does not
change the session runtime binding contract.
curated and generated descriptions, description-generation runs, and sanitized run events.
Connector secrets remain write-only in the encrypted workspace secret store; model credentials
remain in the protected installation secret bundle. Catalog SSH support is limited to connection
tests, table synchronization, and bounded description-generation sampling; it does not change the
session runtime binding contract.
## Human-in-the-loop gate contract
@@ -1,4 +1,5 @@
# Copy this file to an operator-controlled path named exactly thothii-installation.yaml.
# Copy this file to a protected operator-controlled path named exactly thothii-installation.yaml
# and set mode 0600 (or 0400) before using it as THT_INSTALLATION_CONFIG_SOURCE.
# Replace every absolute placeholder. Select exactly one Git transport override.
profile: local
projectDirectory: "/absolute/path/to/ThothII"
@@ -7,6 +8,17 @@ workspaceRepository:
remote: git@git.example.com:organization/workspaces.git
branch: main
access: ssh
metadataGeneration:
default: openai-mini
models:
- id: openai-mini
label: OpenAI Mini
litellm:
provider: openai
model: gpt-4.1-mini
apiKeyEnv: OPENAI_API_KEY
# apiKeyEnv may be omitted only for an explicit endpoint that accepts
# unauthenticated requests.
authentication:
configDirectory: "/absolute/path/to/thothii-auth"
overrides:
@@ -1,4 +1,5 @@
# Copy this file to a protected operator path named exactly thothii-installation.yaml.
# Copy this file to a protected operator path named exactly thothii-installation.yaml
# and set mode 0600 (or 0400) before using it as THT_INSTALLATION_CONFIG_SOURCE.
# Replace every absolute placeholder. Select exactly one Git transport override.
profile: server
projectDirectory: "/absolute/path/to/ThothII"
@@ -7,6 +8,19 @@ workspaceRepository:
remote: git@git.example.com:organization/workspaces.git
branch: main
access: ssh
metadataGeneration:
default: openai-mini
models:
- id: openai-mini
label: OpenAI Mini
litellm:
provider: openai
model: gpt-4.1-mini
endpoint:
baseUrl: https://api.openai.example/v1
apiKeyEnv: OPENAI_API_KEY
# apiKeyEnv may be omitted only for an explicit endpoint that accepts
# unauthenticated requests.
authentication:
# Root-operated source of truth; it is never mounted into core.
configDirectory: "/srv/example/thothii/auth-canonical"
+17
View File
@@ -42,6 +42,10 @@ cp deploy/env/local.env.example deploy/env/local.env
cp deploy/secrets/thothii.secrets.example deploy/secrets/thothii.secrets
chmod 600 deploy/secrets/thothii.secrets
# Copy docs/install/examples/thothii-installation.local.yaml to a protected operator path,
# replace every placeholder, chmod it 600, then set that exact path as
# THT_INSTALLATION_CONFIG_SOURCE in deploy/env/local.env.
docker compose --env-file deploy/env/local.env \
-f compose.yaml -f deploy/compose.local.yaml up --build -d
```
@@ -50,6 +54,7 @@ Set these values in `deploy/env/local.env`:
- `PI_AUTH_FILE`
- `THT_SECRETS_FILE`
- `THT_INSTALLATION_CONFIG_SOURCE` (the exact protected host `thothii-installation.yaml`)
- `THT_WORKSPACE_GIT_REMOTE`
- DWH endpoint
- LLM endpoint
@@ -64,8 +69,20 @@ The documented and supported bundle keys are:
```dotenv
THT_MODEL_API_KEY=...
THT_DWH_API_KEY=...
OPENAI_API_KEY=...
```
`THT_MODEL_API_KEY` remains Pi-only. A metadata-generation model instead references one audited
bundle name from `deploy/secrets/README.md` through `metadataGeneration.models[].apiKeyEnv`.
`apiKeyEnv` may be omitted only for an explicit endpoint that accepts unauthenticated requests;
hosted/default endpoints remain keyed.
Provider/model/endpoint settings stay in the protected installation descriptor; raw keys do not.
Compose mounts exactly `THT_INSTALLATION_CONFIG_SOURCE` into `core` as a read-only config and sets
the backend-only runtime path `THT_INSTALLATION_CONFIG_FILE` to
`/run/thothii-installation/thothii-installation.yaml`. Do not set the runtime path in the host env.
Descriptor and bundle changes are loaded only after application restart.
A private PEM CA remains outside the bundle and must be mounted through a reviewed Compose override.
## Preprocessing
@@ -0,0 +1,249 @@
# AI-generated descriptions for Catalog Tables and Catalog Columns
## Problem Statement
ThothII already stores a Generated Description separately from the curated Description for Catalog
Tables and Catalog Columns, but administrators cannot populate it with AI. ThothAI provides the
useful core workflow—generate table and column comments from schema context and small real-data
samples—but its execution, configuration, and interaction model cannot be copied directly into
ThothII.
Administrators need an asynchronous workflow integrated into Database Management. They must be
able to choose an installation-approved model, generate descriptions for selected or missing
targets, observe understandable progress, stop or recover a stuck operation, review generated
text, and explicitly consolidate it. The solution must retain ThothAI's practical simplicity and
must not introduce a general job platform, model gateway, distributed scheduler, or competing
user-facing CLI.
## Solution
Add Description Generation to Database Management as one installation-wide, sequential background
run owned by the Fastify backend. The browser starts a run and remains responsive while the backend
processes bounded requests one at a time. Each completion is delegated to a short-lived internal
Python helper using LiteLLM. Models, their default, and any API-key secret references are declared in
application setup YAML and are independent of both workspaces and Pi configuration.
Each valid result is written immediately to the target's Generated Description. A minimal run row
and ordered text events provide status, counters, history, and a live log. A stopped or crashed run
is not resumed automatically; completed results remain in place and Generate Missing supplies the
simple recovery path. An Unlock action marks a stale recorded run interrupted only when no helper
or backend generation loop is alive.
Prompts use catalog context and, when available, no more than five real rows and five representative
non-null examples. Samples are transient and never logged or persisted. A valid inability to infer
a description produces a standard application-localized value such as `Non generabile`; provider,
timeout, and response-validation failures remain technical errors.
Generated text remains separate from Description until an administrator uses the existing
checkbox selection and Actions control to consolidate it. Consolidation retains Generated
Description and never writes comments to the external Workspace Database.
## User Stories
1. As an installation operator, I want to declare the models allowed for metadata generation in setup YAML, so that model availability is controlled centrally.
2. As an installation operator, I want to declare one default metadata-generation model, so that administrators begin with a safe operational choice.
3. As an installation operator, I want each model to reference its own API-key secret, so that credentials are not stored in workspaces or browser-visible settings.
4. As an installation operator, I want metadata-generation models to remain independent of Pi models, so that changing this workflow cannot disrupt the core NL-to-SQL experience.
5. As an installation operator, I want invalid model setup to fail validation clearly, so that the application does not start with ambiguous provider behavior.
6. As a Catalog Administrator, I want generation controls to explain when no model is configured, so that I know why the action is unavailable.
7. As a Catalog Administrator, I want to select an approved model from a selector initialized to the setup default, so that I control which model performs the work.
8. As a Catalog Administrator, I want to generate descriptions for selected Catalog Tables, so that I can work on a focused part of the catalog.
9. As a Catalog Administrator, I want to generate descriptions for selected Catalog Columns, so that I can work on individual fields without regenerating a whole table.
10. As a Catalog Administrator, I want to generate all eligible descriptions for a Workspace Database, so that I can initialize a catalog in one operation.
11. As a Catalog Administrator, I want to generate only missing descriptions, so that I can continue interrupted work without replacing completed proposals.
12. As a Catalog Administrator, I want a full run to process Catalog Columns before their Catalog Tables, so that table descriptions can benefit from column descriptions.
13. As a Catalog Administrator, I want generation to run asynchronously after I start it, so that the browser remains usable and progress is not tied to one HTTP request.
14. As a Catalog Administrator, I want only one Description Generation Run active in the installation, so that provider traffic and operational behavior remain predictable.
15. As a Catalog Administrator, I want a second start attempt to return a clear conflict, so that I cannot accidentally overlap generation runs.
16. As a Catalog Administrator, I want synchronization, cleanup, consolidation, and edits for the target Workspace Database blocked during generation, so that the simple sequential run sees stable catalog state.
17. As a Catalog Administrator, I want to see the run's model, scope, status, counters, and timestamps, so that I understand what is happening.
18. As a Catalog Administrator, I want a chronological text log, so that I can follow completed targets and diagnose errors.
19. As a Catalog Administrator, I want live log updates with a polling fallback, so that temporary SSE problems do not hide run progress.
20. As a Catalog Administrator, I want completed and interrupted runs to remain inspectable, so that I can understand prior activity.
21. As a Catalog Administrator, I want to stop an active run, so that I can halt an incorrect or unexpectedly costly operation.
22. As a Catalog Administrator, I want stopping a run to terminate its current model helper and prevent later targets from starting, so that stop has prompt operational effect.
23. As a Catalog Administrator, I want valid results completed before a stop or failure to remain saved, so that useful work is not discarded.
24. As a Catalog Administrator, I want a run left active by a backend restart to become interrupted, so that the UI does not claim nonexistent work is still running.
25. As a Catalog Administrator, I want to unlock a stale active run when no generation process is alive, so that an erroneous recorded lock cannot block future work.
26. As a Catalog Administrator, I want Unlock rejected while a live generation process exists, so that recovery cannot create an overlapping run.
27. As a Catalog Administrator, I want Generate Missing to continue after interruption, so that recovery does not require a special resume mechanism.
28. As a Catalog Administrator, I want one retry for a transient model failure, so that a brief provider fault does not immediately lose a batch.
29. As a Catalog Administrator, I want the run to fail after three consecutive technical failures, so that a broken provider does not generate an unbounded stream of attempts.
30. As a Catalog Administrator, I want a successful request to reset the consecutive-failure count, so that isolated errors do not prematurely stop a useful run.
31. As a Catalog Administrator, I want a completed-with-errors result when isolated batches fail but the run reaches its end, so that partial problems remain visible.
32. As a Catalog Administrator, I want no automatic fallback to a different model, so that the selected model remains truthful and predictable.
33. As a Catalog Administrator, I want malformed or ambiguous model output rejected without writing it, so that descriptions cannot be assigned to the wrong target.
34. As a Catalog Administrator, I want an inability to infer a description represented by standard localized text, so that every valid outcome is understandable in the workspace language.
35. As a Catalog Administrator, I want technical failures kept distinct from non-generatable outcomes, so that provider problems are not mistaken for catalog knowledge.
36. As a Catalog Administrator, I want generated prose written in the workspace language, so that it matches the catalog's intended audience.
37. As a Catalog Administrator, I want generated text stored separately from curated Description, so that AI output remains a reviewable proposal.
38. As a Catalog Administrator, I want to edit a Generated Description manually, so that I can improve a proposal before consolidation.
39. As a Catalog Administrator, I want to select one or more tables or columns and run “Move generated description to Description” from the existing Actions control, so that review remains integrated into the current grids.
40. As a Catalog Administrator, I want consolidation to retain the Generated Description, so that I can still see the proposal from which the curated text was copied.
41. As a Catalog Administrator, I want selected records without a Generated Description skipped and reported, so that the bulk action does not erase curated text.
42. As a Catalog Administrator, I want consolidation and generation to modify only the Metadata Catalog, so that no external database comment is changed.
43. As a Catalog Administrator, I want prompts to use schema facts and existing catalog text, so that generated descriptions are grounded in available metadata.
44. As a Catalog Administrator, I want prompts to use at most five real source rows and five representative values when available, so that the model has useful examples without unbounded disclosure.
45. As a Catalog Administrator, I want to be warned that real source samples are sent to the selected provider, so that I can make an informed disclosure decision.
46. As a Catalog Administrator, I want sampled rows and values excluded from persistence and logs, so that operational history does not become a secondary data store.
47. As a security operator, I want API keys, prompts, samples, and complete provider payloads redacted from logs, so that diagnostics do not leak secrets or source data.
48. As a support operator, I want concise per-target and per-batch event messages, so that failures can be diagnosed without provider-specific internals.
49. As an authorized administrator, I want all generation, cancellation, unlock, and consolidation actions protected by database-management permission, so that ordinary users cannot mutate catalog metadata.
50. As an unauthorized user, I want generation controls hidden or disabled and API calls rejected, so that frontend visibility is not treated as authorization.
51. As an operator, I want setup changes to take effect after an application restart, so that configuration lifecycle remains simple and explicit.
52. As a product owner, I want the first release to avoid queues, parallel calls, distributed locks, and automatic resume, so that effort remains focused on generating and reviewing useful descriptions.
## Implementation Decisions
- The Fastify backend owns one installation-wide Description Generation Run and its sequential
processing loop. It does not delegate lifecycle ownership to Pi or Python.
- A Description Generation Run has one of `queued`, `running`, `completed`,
`completed_with_errors`, `cancelled`, `failed`, or `interrupted`. It stores the Workspace
Database, requested scope, selected model identifier, workspace language, progress counters,
timestamps, and an optional final error summary.
- Ordered Description Generation Events store timestamp, severity, and safe human-readable text.
No durable per-target jobs, model invocation rows, prompt snapshots, sample snapshots, leases,
heartbeats, registry revisions, or provenance chains are introduced.
- Starting a run schedules an in-process background loop and returns the run immediately. The API
exposes start, run/history lookup, event listing and streaming, cancellation, stale-run unlock,
and the safe list of configured model choices. There are no retry-item or resume endpoints.
- One in-memory generation manager enforces the installation-wide active-run rule. The existing
Catalog Operation Coordinator reserves the target Workspace Database for the duration of the
run, without being generalized into a new operation framework.
- On backend startup, persisted `queued` or `running` Description Generation Runs become
`interrupted`. The application performs no automatic replay or resume.
- Unlock succeeds only when no live generation loop or helper child exists. It marks the stale run
interrupted and releases the local reservation; it is not a distributed lock recovery protocol.
- Each model completion uses a short-lived Python helper backed by LiteLLM. Structured input is
supplied over stdin, structured output alone is emitted on stdout, diagnostics use stderr, and
the helper can be terminated by cancellation.
- A completion request contains no more than ten targets. Requests run one at a time. The helper
performs at most one retry for a transient technical failure.
- Three consecutive model-request failures fail the run. A successful request resets that count.
Isolated exhausted failures may be logged and skipped, producing `completed_with_errors` if the
run later reaches its end.
- Every valid generated or non-generatable result is applied immediately to Generated Description.
Earlier writes are retained after cancellation, interruption, or later failure.
- A response must identify requested targets unambiguously and classify each returned result as
generated or non-generatable. Duplicate, unknown, missing, or malformed mappings cause a
technical request failure and no result from that ambiguous response is applied.
- The parser also tolerates one JSON object enclosed by one complete `json` code fence, because
some supported models add that formatting despite the prompt. Any prose outside the fence,
multiple payloads, or malformed/ambiguous mappings remain invalid.
- The application supplies localized standard non-generatable text. Provider wording is not used
as the standard value, and technical errors never write that value.
- A full-database run generates eligible Catalog Columns before Catalog Tables. Generate Missing
excludes targets whose Generated Description is already non-empty; all-generation may replace
existing generated proposals only after the initiating action makes that scope explicit.
- Model choices are declared under a metadata-generation section in installation setup YAML. Each
choice has a stable identifier, display label, LiteLLM provider/model settings, optional endpoint
settings, and an optional environment-secret reference for its API key. The reference may be
omitted only when an explicit endpoint is configured for unauthenticated access. One identifier
is the default.
- An explicit endpoint may set `disableThinking: true`; the helper translates it only to the
Qwen-compatible chat-template switch needed to keep the response within the strict JSON contract.
- Metadata-generation setup is separate from application settings for Pi and from workspace
`llm_policy`. Raw keys never enter setup YAML, the catalog database, API responses, process
arguments, or event text. Configuration reload is restart-only.
- If setup defines no usable model, the safe model-list response is empty and the UI disables
generation with an explanation. The backend still rejects direct generation attempts.
- Prompt construction treats schema names, comments, descriptions, and values as untrusted data.
It requests output in the workspace language and separates instructions from catalog content.
- A request may contain up to five real source rows and up to five representative distinct,
non-null values for relevant columns. Inputs are bounded before prompt construction and are not
persisted or logged.
- The UI discloses that real data can be sent to the selected provider. A future Sensitive Data
Policy will classify values and exclude or anonymize protected data; that policy is not silently
approximated in this slice.
- The generation UI reuses Database Management's table and column selections, model selector,
Actions control, run drawer conventions, SSE delivery, and polling fallback where practical.
Visual parity with Catalog Sync Run logs is not required.
- The consolidation action copies each selected, non-empty Generated Description into Description
in a catalog transaction, retains Generated Description, skips empty proposals, and reports
copied and skipped counts. It never writes to the external Workspace Database.
- Generation, cancellation, unlock, and consolidation require the existing database-management
permission and are validated by the backend independently of UI state.
- No user-facing generation CLI is added. The Python process is an internal completion adapter,
not an operator surface or a long-lived service.
## Testing Decisions
- Tests assert externally observable behavior rather than private loop structure, process timing,
or LiteLLM implementation details.
- The primary and highest test seam is the Fastify catalog API with a test PostgreSQL catalog and
an injected fake Model Completer. It verifies complete paths through authorization, run
persistence, sequential processing, event delivery, Generated Description updates, and final
status without contacting a real provider.
- API tests cover each generation scope, column-before-table order, the ten-target request bound,
model validation, one-active-run conflict, target-database exclusion, cancellation, startup
interruption, Unlock safeguards, Generate Missing, partial success, consecutive failure
handling, non-generatable localization, malformed responses, redacted events, and permissions.
- Catalog repository integration tests verify the migration, run and event ordering, active-run
constraint, immediate description writes, history queries, startup interruption, and bulk
consolidation behavior against PostgreSQL.
- The Python helper has a small black-box contract suite using a simulated LiteLLM adapter. It
verifies stdin/stdout framing, pristine stdout, stderr diagnostics, normalized success and
failure output, one transient retry, secret redaction, and termination behavior.
- Setup-validation tests cover duplicate model identifiers, missing or unknown defaults, malformed
provider settings, missing secret references, safe public model projection, and strict separation
from Pi and workspace model settings.
- Database Management tests use the existing browser-level component seam with MSW. They verify
model selection and default, selected/all/missing actions, disabled state without models, running
progress and logs, polling recovery, cancellation, Unlock visibility, terminal summaries,
generated-text refresh, and selected consolidation with copied/skipped counts.
- Existing Catalog Sync Run route, repository, SSE, and drawer tests are prior art for asynchronous
status and event behavior. Existing catalog table/column editing and Database Management tests
are prior art for optimistic catalog updates, permissions, selection, and action controls.
- One required manual acceptance gate, outside deterministic CI, uses the installation's configured
default model and a disposable PostgreSQL database containing only invented data. Its application
credentials are read-only. It generates Italian text for one Catalog Column and one Catalog
Table, verifies their Generated Description, inspects the safe activity log, confirms that no key
or sample value is exposed, and consolidates one selected result. If the configured secret is not
available, acceptance stops without exposing or requesting the key in conversation.
- Delivery includes a strict MkDocs build executed through repository-managed, reproducible
documentation dependencies rather than globally installed Python packages. A readable direct
dependency file is retained, a complete transitive lock is generated with `uv`, and one canonical
repository command performs the strict build from that lock.
- Successful real-provider acceptance is recorded in a short sanitized report under
`docs/testing/`. It identifies the model and checks performed but contains no credentials,
prompts, source samples, complete provider payloads, or generated database values.
- No tests are added for worker queues, parallel generation, distributed locking, multi-replica
recovery, automatic resume, cost accounting, or model fallback because those behaviors are out
of scope.
## Out of Scope
- Reusing Pi to execute Description Generation or changing Pi's model configuration.
- A shared Installation Model Registry, model gateway, long-lived Python sidecar, or provider
management platform.
- A user-facing generation CLI.
- Parallel model calls, worker queues, adaptive rate limiting, distributed locks, leases,
heartbeats, automatic resume, or multi-replica execution.
- Durable target jobs, invocation history, prompts, samples, token usage, cost accounting,
provenance chains, target snapshots, or advanced retention controls.
- Automatic retry or resume of individual targets beyond one technical helper retry and a new
Generate Missing run.
- Automatic fallback to a different model.
- Writing generated text into comments of the external Workspace Database.
- Generating logical relationships or other catalog metadata beyond Catalog Table and Catalog
Column descriptions.
- Implementing the Sensitive Data Policy. Its definition and exclusion/anonymization behavior are
a required follow-up improvement.
- Generalizing the log viewer across unrelated metadata operations. That broader concern remains
related to Gitea issue #2.
## Further Notes
- The design deliberately follows ThothAI's proven simple workflow while adapting it to ThothII's
asynchronous browser interaction, setup ownership, and existing Generated Description model.
- The source-sampling disclosure is a release requirement, not merely documentation for operators.
- `completed_with_errors` is reserved for a run that reaches the end after isolated technical
failures. Three consecutive failures end the run as `failed`.
- Successful values are their own recovery record: after interruption, Generate Missing naturally
skips them without needing replay state.
- The implementation is available without a feature flag once the catalog migration and valid
setup are present. With no configured model, the feature remains visibly unavailable rather than
partially initialized.
- The final manual gate is intentionally narrow: one real-provider run covers one Catalog Column
and one Catalog Table, generated Italian text, safe events, and one consolidation. Automated
tests remain the evidence for All, Missing, Stop, restart interruption, Unlock, and failure paths.
@@ -0,0 +1,173 @@
# AI catalog description generation
Status: simplified design, API, persistence, test seams, and delivery tickets accepted.
## Objective
Bring ThothAI's useful AI comment-generation workflow into the ThothII Metadata Catalog without
turning it into a general job platform. Administrators can generate editable descriptions for
catalog tables and columns, inspect progress, stop a run, recover a stale run, and explicitly copy
approved generated text into the curated Description field.
The implementation is UI/API only. There is no user-facing generation command.
## ThothAI behavior retained
- Generate descriptions for selected tables, selected columns, missing descriptions, or all
eligible targets.
- Generate columns before their containing table when running the full workflow, so table prompts
can benefit from the resulting column descriptions.
- Process bounded batches of at most ten targets, one model request at a time.
- Include schema context, existing catalog text, up to five real source rows, and up to five
representative non-null values when available.
- Keep generated text separate from the curated Description until an administrator consolidates
it.
- Use the existing table and column checkboxes plus the Actions selector to copy Generated
Description into Description for one or more selected records. The generated value is retained.
- Store a localized standard value such as `Non generabile` when a valid model response says that
a description cannot be inferred.
Unlike ThothAI, every generation action is asynchronous from the browser's perspective and exposes
a persistent, readable activity log.
## Minimal architecture
The Fastify backend owns the run lifecycle and sequential loop. It starts one short-lived Python
helper for each model completion. The helper uses LiteLLM, accepts structured input on stdin,
returns structured output on stdout, and writes diagnostics only to stderr.
This is preferred over reusing Pi. Pi remains the interactive NL-to-SQL orchestration surface,
whereas description generation is a bounded batch transformation with no conversational state or
human gate. A LiteLLM helper avoids inventing a Pi session protocol for a task that needs one
request and one structured response.
There is no Python daemon, model gateway, queue service, worker pool, or generation CLI. Python is
already a core implementation language in ThothII's harness and core image; this helper does not
introduce a new runtime family.
## Run lifecycle and exclusion
- At most one Description Generation Run may be queued or running in the installation.
- Start returns immediately after creating the run and scheduling the in-process backend loop.
- Requests are sequential; there is no parallel provider traffic.
- The target Workspace Database is reserved through the existing in-memory catalog-operation
coordinator. Synchronization, cleanup, consolidation, and direct catalog edits for that database
are rejected while generation is active.
- A second generation start is rejected with a conflict response.
- Stop terminates the current helper process, stops further targets, and marks the run cancelled.
- Backend startup marks any queued or running generation row interrupted. It does not resume work.
- Generate Missing is the normal manual continuation mechanism because successful values were
already saved.
- Unlock is available only when the backend has no live generation process; it marks a stale
recorded run interrupted and clears the local reservation.
This is intentionally a single-process policy. Multi-replica coordination is out of scope.
## Persistence
Persist only:
- a Description Generation Run with database, scope, selected model, language, status, counters,
timestamps, and an optional final error summary;
- ordered Description Generation Events containing timestamp, level, and human-readable text;
- each successful or non-generatable result directly in the target's Generated Description.
Do not add per-target job rows, invocation history, prompt or sample snapshots, provider cost
accounting, leases, heartbeats, registry revisions, or generated-description provenance. The event
log is operational evidence, not a replay mechanism.
## Model setup
Selectable models and their default belong to application setup YAML, not to a workspace. Each
entry supplies a stable display identifier, LiteLLM provider/model information, optional endpoint
settings, and—unless that explicit endpoint is unauthenticated—a reference to an installation
secret containing the API key. Keyless entries without an explicit endpoint are invalid. Raw keys must not be
stored in the YAML, database, frontend, events, or process arguments.
An explicit endpoint may opt into `disableThinking: true` when its Qwen-compatible chat template
would otherwise place reasoning text around the required JSON result.
This metadata-generation configuration is independent of the existing Pi provider/model settings
and workspace `llm_policy`. A setup change takes effect after application restart. If no model is
configured, generation controls are disabled with an explanatory message.
The browser receives only the selectable identifiers and labels. The selected value defaults to
the setup default and is validated again by the backend when a run starts.
## Prompt inputs and outputs
Targets are grouped in model requests of at most ten. Prompts distinguish instructions from
untrusted schema names, comments, descriptions, and sampled values. A response must map every
returned result to a requested target and classify it as generated or non-generatable. Missing,
duplicate, unknown, or malformed target results make that request a technical failure rather than
silently writing ambiguous text.
One complete `json` code fence around the object is tolerated for model compatibility; prose
outside it, multiple payloads, and ambiguous mappings are still rejected.
For a complete database run, eligible columns are processed before tables. A table request can use
the current Generated Description or Description of its columns. The output language is the
workspace language; the standard non-generatable text is localized by the application rather than
trusted to arbitrary model wording.
Up to five source rows and five representative examples may be sent to the provider and are never
persisted. Delivery must call out this disclosure. A follow-up Sensitive Data Policy will define
which values are excluded or anonymized.
## Errors, retry, and logs
The helper performs at most one retry for a transient technical provider failure. A final failed
request produces an error event and increments the consecutive-error count. The run stops as
failed after three consecutive technical failures; any successful request resets the count. There
is no automatic fallback to another model.
Valid non-generatable outcomes are results, not technical errors. Successful results from earlier
requests remain stored when a later request fails or the run is stopped.
The UI shows status, counters, selected model, start/end times, and a chronological text log. Live
delivery may reuse the existing SSE infrastructure with polling as fallback; exact visual parity
with synchronization logs is not required. Logs must not contain API keys, prompts, source sample
values, or full provider payloads.
## Explicitly deferred complexity
- shared model registry or cutover of Pi configuration;
- long-lived Python sidecar or internal HTTP model gateway;
- generic catalog-operation kernel;
- durable target items, invocation records, target snapshots, or provenance chains;
- distributed locks, leases, heartbeats, worker queues, automatic resume, or multi-replica support;
- parallel calls, adaptive rate limiting, cost estimation, advanced metrics, or model fallback;
- user-facing generation CLI;
- automatic writeback to comments in the external database;
- Sensitive Data Policy implementation, which remains a required improvement after this slice.
## Delivery tracking
The accepted specification is Gitea issue #4 and the implementation is split into issues #5–#11.
Each ticket is a bounded vertical slice with explicit Gitea dependencies. Implementation proceeds
from the unblocked frontier, using a fresh subagent context for each ticket; integration and final
verification remain centralized so later slices cannot silently reopen the deferred platform
features above.
Issue #4 remains open until delivery completes four final gates: the stale Compose service-set
contract is corrected in its own commit; documentation dependencies are repository-managed and a
strict MkDocs build passes; one narrow real-provider acceptance run succeeds against non-sensitive
test data; and a separate, non-blocking Sensitive Data Policy design ticket is linked as required
follow-up work.
The documentation toolchain retains a readable direct-dependency input, adds a complete lock
generated with `uv`, and exposes one canonical strict-build command. Real-provider acceptance uses
the installation's configured default model and a disposable PostgreSQL database seeded only with
invented values and accessed read-only by the application. A missing protected model secret stops
the gate without disclosing it. The successful gate is captured in a sanitized report under
`docs/testing/` without prompts, samples, full generated values, payloads, or credentials.
Delivery is organized as four reviewable commits: the stale Compose contract correction, the
reproducible documentation toolchain, the AI-description feature, and—only after acceptance—the
sanitized acceptance report. A failed real-provider gate does not invalidate already verified
commits, but issue #4 remains open and no acceptance report claims success. Application defects are
fixed and reverified; missing configuration or provider unavailability is recorded and retried.
After every gate passes, the existing `codex/db-management` branch is pushed to its configured
origin without introducing a new pull-request workflow, then issue #4 is closed with links to the
delivery evidence. The separate Sensitive Data Policy issue is created as non-blocking follow-up,
linked to #4, and labeled `enhancement` plus `ready-for-human` because its design requires a future
`grill-with-docs` before agent implementation.