feat: add AI catalog description generation

This commit is contained in:
Codex
2026-08-29 16:42:56 +02:00
parent b0afba81ca
commit 376dd5a09d
76 changed files with 14860 additions and 102 deletions
@@ -0,0 +1,173 @@
# AI catalog description generation
Status: simplified design, API, persistence, test seams, and delivery tickets accepted.
## Objective
Bring ThothAI's useful AI comment-generation workflow into the ThothII Metadata Catalog without
turning it into a general job platform. Administrators can generate editable descriptions for
catalog tables and columns, inspect progress, stop a run, recover a stale run, and explicitly copy
approved generated text into the curated Description field.
The implementation is UI/API only. There is no user-facing generation command.
## ThothAI behavior retained
- Generate descriptions for selected tables, selected columns, missing descriptions, or all
eligible targets.
- Generate columns before their containing table when running the full workflow, so table prompts
can benefit from the resulting column descriptions.
- Process bounded batches of at most ten targets, one model request at a time.
- Include schema context, existing catalog text, up to five real source rows, and up to five
representative non-null values when available.
- Keep generated text separate from the curated Description until an administrator consolidates
it.
- Use the existing table and column checkboxes plus the Actions selector to copy Generated
Description into Description for one or more selected records. The generated value is retained.
- Store a localized standard value such as `Non generabile` when a valid model response says that
a description cannot be inferred.
Unlike ThothAI, every generation action is asynchronous from the browser's perspective and exposes
a persistent, readable activity log.
## Minimal architecture
The Fastify backend owns the run lifecycle and sequential loop. It starts one short-lived Python
helper for each model completion. The helper uses LiteLLM, accepts structured input on stdin,
returns structured output on stdout, and writes diagnostics only to stderr.
This is preferred over reusing Pi. Pi remains the interactive NL-to-SQL orchestration surface,
whereas description generation is a bounded batch transformation with no conversational state or
human gate. A LiteLLM helper avoids inventing a Pi session protocol for a task that needs one
request and one structured response.
There is no Python daemon, model gateway, queue service, worker pool, or generation CLI. Python is
already a core implementation language in ThothII's harness and core image; this helper does not
introduce a new runtime family.
## Run lifecycle and exclusion
- At most one Description Generation Run may be queued or running in the installation.
- Start returns immediately after creating the run and scheduling the in-process backend loop.
- Requests are sequential; there is no parallel provider traffic.
- The target Workspace Database is reserved through the existing in-memory catalog-operation
coordinator. Synchronization, cleanup, consolidation, and direct catalog edits for that database
are rejected while generation is active.
- A second generation start is rejected with a conflict response.
- Stop terminates the current helper process, stops further targets, and marks the run cancelled.
- Backend startup marks any queued or running generation row interrupted. It does not resume work.
- Generate Missing is the normal manual continuation mechanism because successful values were
already saved.
- Unlock is available only when the backend has no live generation process; it marks a stale
recorded run interrupted and clears the local reservation.
This is intentionally a single-process policy. Multi-replica coordination is out of scope.
## Persistence
Persist only:
- a Description Generation Run with database, scope, selected model, language, status, counters,
timestamps, and an optional final error summary;
- ordered Description Generation Events containing timestamp, level, and human-readable text;
- each successful or non-generatable result directly in the target's Generated Description.
Do not add per-target job rows, invocation history, prompt or sample snapshots, provider cost
accounting, leases, heartbeats, registry revisions, or generated-description provenance. The event
log is operational evidence, not a replay mechanism.
## Model setup
Selectable models and their default belong to application setup YAML, not to a workspace. Each
entry supplies a stable display identifier, LiteLLM provider/model information, optional endpoint
settings, and—unless that explicit endpoint is unauthenticated—a reference to an installation
secret containing the API key. Keyless entries without an explicit endpoint are invalid. Raw keys must not be
stored in the YAML, database, frontend, events, or process arguments.
An explicit endpoint may opt into `disableThinking: true` when its Qwen-compatible chat template
would otherwise place reasoning text around the required JSON result.
This metadata-generation configuration is independent of the existing Pi provider/model settings
and workspace `llm_policy`. A setup change takes effect after application restart. If no model is
configured, generation controls are disabled with an explanatory message.
The browser receives only the selectable identifiers and labels. The selected value defaults to
the setup default and is validated again by the backend when a run starts.
## Prompt inputs and outputs
Targets are grouped in model requests of at most ten. Prompts distinguish instructions from
untrusted schema names, comments, descriptions, and sampled values. A response must map every
returned result to a requested target and classify it as generated or non-generatable. Missing,
duplicate, unknown, or malformed target results make that request a technical failure rather than
silently writing ambiguous text.
One complete `json` code fence around the object is tolerated for model compatibility; prose
outside it, multiple payloads, and ambiguous mappings are still rejected.
For a complete database run, eligible columns are processed before tables. A table request can use
the current Generated Description or Description of its columns. The output language is the
workspace language; the standard non-generatable text is localized by the application rather than
trusted to arbitrary model wording.
Up to five source rows and five representative examples may be sent to the provider and are never
persisted. Delivery must call out this disclosure. A follow-up Sensitive Data Policy will define
which values are excluded or anonymized.
## Errors, retry, and logs
The helper performs at most one retry for a transient technical provider failure. A final failed
request produces an error event and increments the consecutive-error count. The run stops as
failed after three consecutive technical failures; any successful request resets the count. There
is no automatic fallback to another model.
Valid non-generatable outcomes are results, not technical errors. Successful results from earlier
requests remain stored when a later request fails or the run is stopped.
The UI shows status, counters, selected model, start/end times, and a chronological text log. Live
delivery may reuse the existing SSE infrastructure with polling as fallback; exact visual parity
with synchronization logs is not required. Logs must not contain API keys, prompts, source sample
values, or full provider payloads.
## Explicitly deferred complexity
- shared model registry or cutover of Pi configuration;
- long-lived Python sidecar or internal HTTP model gateway;
- generic catalog-operation kernel;
- durable target items, invocation records, target snapshots, or provenance chains;
- distributed locks, leases, heartbeats, worker queues, automatic resume, or multi-replica support;
- parallel calls, adaptive rate limiting, cost estimation, advanced metrics, or model fallback;
- user-facing generation CLI;
- automatic writeback to comments in the external database;
- Sensitive Data Policy implementation, which remains a required improvement after this slice.
## Delivery tracking
The accepted specification is Gitea issue #4 and the implementation is split into issues #5–#11.
Each ticket is a bounded vertical slice with explicit Gitea dependencies. Implementation proceeds
from the unblocked frontier, using a fresh subagent context for each ticket; integration and final
verification remain centralized so later slices cannot silently reopen the deferred platform
features above.
Issue #4 remains open until delivery completes four final gates: the stale Compose service-set
contract is corrected in its own commit; documentation dependencies are repository-managed and a
strict MkDocs build passes; one narrow real-provider acceptance run succeeds against non-sensitive
test data; and a separate, non-blocking Sensitive Data Policy design ticket is linked as required
follow-up work.
The documentation toolchain retains a readable direct-dependency input, adds a complete lock
generated with `uv`, and exposes one canonical strict-build command. Real-provider acceptance uses
the installation's configured default model and a disposable PostgreSQL database seeded only with
invented values and accessed read-only by the application. A missing protected model secret stops
the gate without disclosing it. The successful gate is captured in a sanitized report under
`docs/testing/` without prompts, samples, full generated values, payloads, or credentials.
Delivery is organized as four reviewable commits: the stale Compose contract correction, the
reproducible documentation toolchain, the AI-description feature, and—only after acceptance—the
sanitized acceptance report. A failed real-provider gate does not invalidate already verified
commits, but issue #4 remains open and no acceptance report claims success. Application defects are
fixed and reverified; missing configuration or provider unavailability is recorded and retried.
After every gate passes, the existing `codex/db-management` branch is pushed to its configured
origin without introducing a new pull-request workflow, then issue #4 is closed with links to the
delivery evidence. The separate Sensitive Data Policy issue is created as non-blocking follow-up,
linked to #4, and labeled `enhancement` plus `ready-for-human` because its design requires a future
`grill-with-docs` before agent implementation.