10 KiB
AI catalog description generation
Status: simplified design, API, persistence, test seams, and delivery tickets accepted.
Objective
Bring ThothAI's useful AI comment-generation workflow into the ThothII Metadata Catalog without turning it into a general job platform. Administrators can generate editable descriptions for catalog tables and columns, inspect progress, stop a run, recover a stale run, and explicitly copy approved generated text into the curated Description field.
The implementation is UI/API only. There is no user-facing generation command.
ThothAI behavior retained
- Generate descriptions for selected tables, selected columns, missing descriptions, or all eligible targets.
- Generate columns before their containing table when running the full workflow, so table prompts can benefit from the resulting column descriptions.
- Process bounded batches of at most ten targets, one model request at a time.
- Include schema context, existing catalog text, up to five real source rows, and up to five representative non-null values when available.
- Keep generated text separate from the curated Description until an administrator consolidates it.
- Use the existing table and column checkboxes plus the Actions selector to copy Generated Description into Description for one or more selected records. The generated value is retained.
- Store a localized standard value such as
Non generabilewhen a valid model response says that a description cannot be inferred.
Unlike ThothAI, every generation action is asynchronous from the browser's perspective and exposes a persistent, readable activity log.
Minimal architecture
The Fastify backend owns the run lifecycle and sequential loop. It starts one short-lived Python helper for each model completion. The helper uses LiteLLM, accepts structured input on stdin, returns structured output on stdout, and writes diagnostics only to stderr.
This is preferred over reusing Pi. Pi remains the interactive NL-to-SQL orchestration surface, whereas description generation is a bounded batch transformation with no conversational state or human gate. A LiteLLM helper avoids inventing a Pi session protocol for a task that needs one request and one structured response.
There is no Python daemon, model gateway, queue service, worker pool, or generation CLI. Python is already a core implementation language in ThothII's harness and core image; this helper does not introduce a new runtime family.
Run lifecycle and exclusion
- At most one Description Generation Run may be queued or running in the installation.
- Start returns immediately after creating the run and scheduling the in-process backend loop.
- Requests are sequential; there is no parallel provider traffic.
- The target Workspace Database is reserved through the existing in-memory catalog-operation coordinator. Synchronization, cleanup, consolidation, and direct catalog edits for that database are rejected while generation is active.
- A second generation start is rejected with a conflict response.
- Stop terminates the current helper process, stops further targets, and marks the run cancelled.
- Backend startup marks any queued or running generation row interrupted. It does not resume work.
- Generate Missing is the normal manual continuation mechanism because successful values were already saved.
- Unlock is available only when the backend has no live generation process; it marks a stale recorded run interrupted and clears the local reservation.
This is intentionally a single-process policy. Multi-replica coordination is out of scope.
Persistence
Persist only:
- a Description Generation Run with database, scope, selected model, language, status, counters, timestamps, and an optional final error summary;
- ordered Description Generation Events containing timestamp, level, and human-readable text;
- each successful or non-generatable result directly in the target's Generated Description.
Do not add per-target job rows, invocation history, prompt or sample snapshots, provider cost accounting, leases, heartbeats, registry revisions, or generated-description provenance. The event log is operational evidence, not a replay mechanism.
Model setup
Selectable models and their default belong to application setup YAML, not to a workspace. Each
entry supplies a stable display identifier, LiteLLM provider/model information, optional endpoint
settings, and—unless that explicit endpoint is unauthenticated—a reference to an installation
secret containing the API key. Keyless entries without an explicit endpoint are invalid. Raw keys must not be
stored in the YAML, database, frontend, events, or process arguments.
An explicit endpoint may opt into disableThinking: true when its Qwen-compatible chat template
would otherwise place reasoning text around the required JSON result.
This metadata-generation configuration is independent of the existing Pi provider/model settings
and workspace llm_policy. A setup change takes effect after application restart. If no model is
configured, generation controls are disabled with an explanatory message.
The browser receives only the selectable identifiers and labels. The selected value defaults to the setup default and is validated again by the backend when a run starts.
Prompt inputs and outputs
Targets are grouped in model requests of at most ten. Prompts distinguish instructions from
untrusted schema names, comments, descriptions, and sampled values. A response must map every
returned result to a requested target and classify it as generated or non-generatable. Missing,
duplicate, unknown, or malformed target results make that request a technical failure rather than
silently writing ambiguous text.
One complete json code fence around the object is tolerated for model compatibility; prose
outside it, multiple payloads, and ambiguous mappings are still rejected.
For a complete database run, eligible columns are processed before tables. A table request can use the current Generated Description or Description of its columns. The output language is the workspace language; the standard non-generatable text is localized by the application rather than trusted to arbitrary model wording.
Up to five source rows and five representative examples may be sent to the provider and are never persisted. Delivery must call out this disclosure. A follow-up Sensitive Data Policy will define which values are excluded or anonymized.
Errors, retry, and logs
The helper performs at most one retry for a transient technical provider failure. A final failed request produces an error event and increments the consecutive-error count. The run stops as failed after three consecutive technical failures; any successful request resets the count. There is no automatic fallback to another model.
Valid non-generatable outcomes are results, not technical errors. Successful results from earlier requests remain stored when a later request fails or the run is stopped.
The UI shows status, counters, selected model, start/end times, and a chronological text log. Live delivery may reuse the existing SSE infrastructure with polling as fallback; exact visual parity with synchronization logs is not required. Logs must not contain API keys, prompts, source sample values, or full provider payloads.
Explicitly deferred complexity
- shared model registry or cutover of Pi configuration;
- long-lived Python sidecar or internal HTTP model gateway;
- generic catalog-operation kernel;
- durable target items, invocation records, target snapshots, or provenance chains;
- distributed locks, leases, heartbeats, worker queues, automatic resume, or multi-replica support;
- parallel calls, adaptive rate limiting, cost estimation, advanced metrics, or model fallback;
- user-facing generation CLI;
- automatic writeback to comments in the external database;
- Sensitive Data Policy implementation, which remains a required improvement after this slice.
Delivery tracking
The accepted specification is Gitea issue #4 and the implementation is split into issues #5–#11. Each ticket is a bounded vertical slice with explicit Gitea dependencies. Implementation proceeds from the unblocked frontier, using a fresh subagent context for each ticket; integration and final verification remain centralized so later slices cannot silently reopen the deferred platform features above.
Issue #4 remains open until delivery completes four final gates: the stale Compose service-set contract is corrected in its own commit; documentation dependencies are repository-managed and a strict MkDocs build passes; one narrow real-provider acceptance run succeeds against non-sensitive test data; and a separate, non-blocking Sensitive Data Policy design ticket is linked as required follow-up work.
The documentation toolchain retains a readable direct-dependency input, adds a complete lock
generated with uv, and exposes one canonical strict-build command. Real-provider acceptance uses
the installation's configured default model and a disposable PostgreSQL database seeded only with
invented values and accessed read-only by the application. A missing protected model secret stops
the gate without disclosing it. The successful gate is captured in a sanitized report under
docs/testing/ without prompts, samples, full generated values, payloads, or credentials.
Delivery is organized as four reviewable commits: the stale Compose contract correction, the reproducible documentation toolchain, the AI-description feature, and—only after acceptance—the sanitized acceptance report. A failed real-provider gate does not invalidate already verified commits, but issue #4 remains open and no acceptance report claims success. Application defects are fixed and reverified; missing configuration or provider unavailability is recorded and retried.
After every gate passes, the existing codex/db-management branch is pushed to its configured
origin without introducing a new pull-request workflow, then issue #4 is closed with links to the
delivery evidence. The separate Sensitive Data Policy issue is created as non-blocking follow-up,
linked to #4, and labeled enhancement plus ready-for-human because its design requires a future
grill-with-docs before agent implementation.