Files
ThothII/docs/plans/2026-08-28-ai-catalog-description-generation-spec.md
T
Codex cffa60772e
Publish documentation / publish (push) Successful in 2m12s
feat: complete catalog-driven preprocessing
2026-09-06 17:49:35 +02:00

23 KiB

AI-generated descriptions for Catalog Tables and Catalog Columns

Problem Statement

ThothII already stores a Generated Description separately from the curated Description for Catalog Tables and Catalog Columns, but administrators cannot populate it with AI. ThothAI provides the useful core workflow—generate table and column comments from schema context and small real-data samples—but its execution, configuration, and interaction model cannot be copied directly into ThothII.

Administrators need an asynchronous workflow integrated into Database Management. They must be able to choose an installation-approved model, generate descriptions for selected or missing targets, observe understandable progress, stop or recover a stuck operation, review generated text, and explicitly consolidate it. The solution must retain ThothAI's practical simplicity and must not introduce a general job platform, model gateway, distributed scheduler, or competing user-facing CLI.

Solution

Add Description Generation to Database Management as one installation-wide, sequential background run owned by the Fastify backend. The browser starts a run and remains responsive while the backend processes bounded requests one at a time. Each completion is delegated to a short-lived internal Python helper using LiteLLM. Models, their default, and any API-key secret references are declared in application setup YAML and are independent of both workspaces and Pi configuration.

Each valid result is written immediately to the target's Generated Description. A minimal run row and ordered text events provide status, counters, history, and a live log. A stopped or crashed run is not resumed automatically; completed results remain in place and Generate Missing supplies the simple recovery path. An Unlock action marks a stale recorded run interrupted only when no helper or backend generation loop is alive.

Prompts use catalog context and, when available, no more than five real rows and five representative non-null examples. Samples are transient and never logged or persisted. A valid inability to infer a description produces a standard application-localized value such as Non generabile; provider, timeout, and response-validation failures remain technical errors.

Generated text remains separate from Description until an administrator uses the existing checkbox selection and Actions control to consolidate it. Consolidation retains Generated Description and never writes comments to the external Workspace Database.

User Stories

  1. As an installation operator, I want to declare the models allowed for metadata generation in setup YAML, so that model availability is controlled centrally.
  2. As an installation operator, I want to declare one default metadata-generation model, so that administrators begin with a safe operational choice.
  3. As an installation operator, I want each model to reference its own API-key secret, so that credentials are not stored in workspaces or browser-visible settings.
  4. As an installation operator, I want metadata-generation models to remain independent of Pi models, so that changing this workflow cannot disrupt the core NL-to-SQL experience.
  5. As an installation operator, I want invalid model setup to fail validation clearly, so that the application does not start with ambiguous provider behavior.
  6. As a Catalog Administrator, I want generation controls to explain when no model is configured, so that I know why the action is unavailable.
  7. As a Catalog Administrator, I want to select an approved model from a selector initialized to the setup default, so that I control which model performs the work.
  8. As a Catalog Administrator, I want to generate descriptions for selected Catalog Tables, so that I can work on a focused part of the catalog.
  9. As a Catalog Administrator, I want to generate descriptions for selected Catalog Columns, so that I can work on individual fields without regenerating a whole table.
  10. As a Catalog Administrator, I want to generate all eligible descriptions for a Workspace Database, so that I can initialize a catalog in one operation.
  11. As a Catalog Administrator, I want to generate only missing descriptions, so that I can continue interrupted work without replacing completed proposals.
  12. As a Catalog Administrator, I want a full run to process Catalog Columns before their Catalog Tables, so that table descriptions can benefit from column descriptions.
  13. As a Catalog Administrator, I want generation to run asynchronously after I start it, so that the browser remains usable and progress is not tied to one HTTP request.
  14. As a Catalog Administrator, I want only one Description Generation Run active in the installation, so that provider traffic and operational behavior remain predictable.
  15. As a Catalog Administrator, I want a second start attempt to return a clear conflict, so that I cannot accidentally overlap generation runs.
  16. As a Catalog Administrator, I want synchronization, cleanup, consolidation, and edits for the target Workspace Database blocked during generation, so that the simple sequential run sees stable catalog state.
  17. As a Catalog Administrator, I want to see the run's model, scope, status, counters, and timestamps, so that I understand what is happening.
  18. As a Catalog Administrator, I want a chronological text log, so that I can follow completed targets and diagnose errors.
  19. As a Catalog Administrator, I want live log updates with a polling fallback, so that temporary SSE problems do not hide run progress.
  20. As a Catalog Administrator, I want completed and interrupted runs to remain inspectable, so that I can understand prior activity.
  21. As a Catalog Administrator, I want to stop an active run, so that I can halt an incorrect or unexpectedly costly operation.
  22. As a Catalog Administrator, I want stopping a run to terminate its current model helper and prevent later targets from starting, so that stop has prompt operational effect.
  23. As a Catalog Administrator, I want valid results completed before a stop or failure to remain saved, so that useful work is not discarded.
  24. As a Catalog Administrator, I want a run left active by a backend restart to become interrupted, so that the UI does not claim nonexistent work is still running.
  25. As a Catalog Administrator, I want to unlock a stale active run when no generation process is alive, so that an erroneous recorded lock cannot block future work.
  26. As a Catalog Administrator, I want Unlock rejected while a live generation process exists, so that recovery cannot create an overlapping run.
  27. As a Catalog Administrator, I want Generate Missing to continue after interruption, so that recovery does not require a special resume mechanism.
  28. As a Catalog Administrator, I want one retry for a transient model failure, so that a brief provider fault does not immediately lose a batch.
  29. As a Catalog Administrator, I want the run to fail after three consecutive technical failures, so that a broken provider does not generate an unbounded stream of attempts.
  30. As a Catalog Administrator, I want a successful request to reset the consecutive-failure count, so that isolated errors do not prematurely stop a useful run.
  31. As a Catalog Administrator, I want a completed-with-errors result when isolated batches fail but the run reaches its end, so that partial problems remain visible.
  32. As a Catalog Administrator, I want no automatic fallback to a different model, so that the selected model remains truthful and predictable.
  33. As a Catalog Administrator, I want malformed or ambiguous model output rejected without writing it, so that descriptions cannot be assigned to the wrong target.
  34. As a Catalog Administrator, I want an inability to infer a description represented by standard localized text, so that every valid outcome is understandable in the workspace language.
  35. As a Catalog Administrator, I want technical failures kept distinct from non-generatable outcomes, so that provider problems are not mistaken for catalog knowledge.
  36. As a Catalog Administrator, I want generated prose written in the workspace language, so that it matches the catalog's intended audience.
  37. As a Catalog Administrator, I want generated text stored separately from curated Description, so that AI output remains a reviewable proposal.
  38. As a Catalog Administrator, I want to edit a Generated Description manually, so that I can improve a proposal before consolidation.
  39. As a Catalog Administrator, I want to select one or more tables or columns and run “Move generated description to Description” from the existing Actions control, so that review remains integrated into the current grids.
  40. As a Catalog Administrator, I want consolidation to retain the Generated Description, so that I can still see the proposal from which the curated text was copied.
  41. As a Catalog Administrator, I want selected records without a Generated Description skipped and reported, so that the bulk action does not erase curated text.
  42. As a Catalog Administrator, I want consolidation and generation to modify only the Metadata Catalog, so that no external database comment is changed.
  43. As a Catalog Administrator, I want prompts to use schema facts and existing catalog text, so that generated descriptions are grounded in available metadata.
  44. As a Catalog Administrator, I want prompts to use at most five real source rows and five representative values when available, so that the model has useful examples without unbounded disclosure.
  45. As a Catalog Administrator, I want to be warned that real source samples are sent to the selected provider, so that I can make an informed disclosure decision.
  46. As a Catalog Administrator, I want sampled rows and values excluded from persistence and logs, so that operational history does not become a secondary data store.
  47. As a security operator, I want API keys, prompts, samples, and complete provider payloads redacted from logs, so that diagnostics do not leak secrets or source data.
  48. As a support operator, I want concise per-target and per-batch event messages, so that failures can be diagnosed without provider-specific internals.
  49. As an authorized administrator, I want all generation, cancellation, unlock, and consolidation actions protected by database-management permission, so that ordinary users cannot mutate catalog metadata.
  50. As an unauthorized user, I want generation controls hidden or disabled and API calls rejected, so that frontend visibility is not treated as authorization.
  51. As an operator, I want setup changes to take effect after an application restart, so that configuration lifecycle remains simple and explicit.
  52. As a product owner, I want the first release to avoid queues, parallel calls, distributed locks, and automatic resume, so that effort remains focused on generating and reviewing useful descriptions.

Implementation Decisions

  • The Fastify backend owns one installation-wide Description Generation Run and its sequential processing loop. It does not delegate lifecycle ownership to Pi or Python.
  • A Description Generation Run has one of queued, running, completed, completed_with_errors, cancelled, failed, or interrupted. It stores the Workspace Database, requested scope, selected model identifier, workspace language, progress counters, timestamps, and an optional final error summary.
  • Ordered Description Generation Events store timestamp, severity, and safe human-readable text. When an event identifies a target, it uses the object type and qualified physical name, such as Column "patients.birth_date" or Table "patients"; catalog UUIDs remain internal identifiers. No durable per-target jobs, model invocation rows, prompt snapshots, sample snapshots, leases, heartbeats, registry revisions, or provenance chains are introduced.
  • Starting a run schedules an in-process background loop and returns the run immediately. The API exposes start, run/history lookup, event listing and streaming, cancellation, stale-run unlock, and the safe list of configured model choices. There are no retry-item or resume endpoints.
  • One in-memory generation manager enforces the installation-wide active-run rule. The existing Catalog Operation Coordinator reserves the target Workspace Database for the duration of the run, without being generalized into a new operation framework.
  • On backend startup, persisted queued or running Description Generation Runs become interrupted. The application performs no automatic replay or resume.
  • Unlock succeeds only when no live generation loop or helper child exists. It marks the stale run interrupted and releases the local reservation; it is not a distributed lock recovery protocol.
  • Each model completion uses a short-lived Python helper backed by LiteLLM. Structured input is supplied over stdin, structured output alone is emitted on stdout, diagnostics use stderr, and the helper can be terminated by cancellation.
  • A completion request contains no more than ten targets. Requests run one at a time. The helper performs at most one retry for a transient technical failure.
  • Three consecutive model-request failures fail the run. A successful request resets that count. Isolated exhausted failures may be logged and skipped, producing completed_with_errors if the run later reaches its end.
  • Every valid generated or non-generatable result is applied immediately to Generated Description. Earlier writes are retained after cancellation, interruption, or later failure.
  • A response must identify requested targets unambiguously and classify each returned result as generated or non-generatable. Duplicate, unknown, missing, or malformed mappings cause a technical request failure and no result from that ambiguous response is applied.
  • The parser also tolerates one JSON object enclosed by one complete json code fence, because some supported models add that formatting despite the prompt. Any prose outside the fence, multiple payloads, or malformed/ambiguous mappings remain invalid.
  • The application supplies localized standard non-generatable text. Provider wording is not used as the standard value, and technical errors never write that value.
  • A full-database run generates eligible Catalog Columns before Catalog Tables. Generate Missing excludes targets whose Generated Description is already non-empty; all-generation may replace existing generated proposals only after the initiating action makes that scope explicit.
  • Model choices are declared under a metadata-generation section in installation setup YAML. Each choice has a stable identifier, display label, LiteLLM provider/model settings, optional endpoint settings, and an optional environment-secret reference for its API key. The reference may be omitted only when an explicit endpoint is configured for unauthenticated access. One identifier is the default.
  • An explicit endpoint may set disableThinking: true; the helper translates it only to the Qwen-compatible chat-template switch needed to keep the response within the strict JSON contract.
  • Metadata-generation setup is separate from application settings for Pi and from workspace llm_policy. Raw keys never enter setup YAML, the catalog database, API responses, process arguments, or event text. Configuration reload is restart-only.
  • If setup defines no usable model, the safe model-list response is empty and the UI disables generation with an explanation. The backend still rejects direct generation attempts.
  • Prompt construction treats schema names, comments, descriptions, and values as untrusted data. It requests output in the workspace language and separates instructions from catalog content.
  • A request may contain up to five real source rows and up to five representative distinct, non-null values for relevant columns. Inputs are bounded before prompt construction and are not persisted or logged.
  • The UI discloses that real data can be sent to the selected provider. A future Sensitive Data Policy will classify values and exclude or anonymize protected data; that policy is not silently approximated in this slice.
  • The generation UI reuses Database Management's table and column selections, model selector, Actions control, run drawer conventions, SSE delivery, and polling fallback where practical. Visual parity with Catalog Sync Run logs is not required.
  • The consolidation action copies each selected, non-empty Generated Description into Description in a catalog transaction, retains Generated Description, skips empty proposals, and reports copied and skipped counts. It never writes to the external Workspace Database.
  • Generation, cancellation, unlock, and consolidation require the existing database-management permission and are validated by the backend independently of UI state.
  • No user-facing generation CLI is added. The Python process is an internal completion adapter, not an operator surface or a long-lived service.

Testing Decisions

  • Tests assert externally observable behavior rather than private loop structure, process timing, or LiteLLM implementation details.
  • The primary and highest test seam is the Fastify catalog API with a test PostgreSQL catalog and an injected fake Model Completer. It verifies complete paths through authorization, run persistence, sequential processing, event delivery, Generated Description updates, and final status without contacting a real provider.
  • API tests cover each generation scope, column-before-table order, the ten-target request bound, model validation, one-active-run conflict, target-database exclusion, cancellation, startup interruption, Unlock safeguards, Generate Missing, partial success, consecutive failure handling, non-generatable localization, malformed responses, redacted events, and permissions.
  • Catalog repository integration tests verify the migration, run and event ordering, active-run constraint, immediate description writes, history queries, startup interruption, and bulk consolidation behavior against PostgreSQL.
  • The Python helper has a small black-box contract suite using a simulated LiteLLM adapter. It verifies stdin/stdout framing, pristine stdout, stderr diagnostics, normalized success and failure output, one transient retry, secret redaction, and termination behavior.
  • Setup-validation tests cover duplicate model identifiers, missing or unknown defaults, malformed provider settings, missing secret references, safe public model projection, and strict separation from Pi and workspace model settings.
  • Database Management tests use the existing browser-level component seam with MSW. They verify model selection and default, selected/all/missing actions, disabled state without models, running progress and logs, polling recovery, cancellation, Unlock visibility, terminal summaries, generated-text refresh, and selected consolidation with copied/skipped counts.
  • Existing Catalog Sync Run route, repository, SSE, and drawer tests are prior art for asynchronous status and event behavior. Existing catalog table/column editing and Database Management tests are prior art for optimistic catalog updates, permissions, selection, and action controls.
  • One required manual acceptance gate, outside deterministic CI, uses the installation's configured default model and a disposable PostgreSQL database containing only invented data. Its application credentials are read-only. It generates Italian text for one Catalog Column and one Catalog Table, verifies their Generated Description, inspects the safe activity log, confirms that no key or sample value is exposed, and consolidates one selected result. If the configured secret is not available, acceptance stops without exposing or requesting the key in conversation.
  • Delivery includes a strict MkDocs build executed through repository-managed, reproducible documentation dependencies rather than globally installed Python packages. A readable direct dependency file is retained, a complete transitive lock is generated with uv, and one canonical repository command performs the strict build from that lock.
  • Successful real-provider acceptance is recorded in a short sanitized report under docs/testing/. It identifies the model and checks performed but contains no credentials, prompts, source samples, complete provider payloads, or generated database values.
  • No tests are added for worker queues, parallel generation, distributed locking, multi-replica recovery, automatic resume, cost accounting, or model fallback because those behaviors are out of scope.

Out of Scope

  • Reusing Pi to execute Description Generation or changing Pi's model configuration.
  • A shared Installation Model Registry, model gateway, long-lived Python sidecar, or provider management platform.
  • A user-facing generation CLI.
  • Parallel model calls, worker queues, adaptive rate limiting, distributed locks, leases, heartbeats, automatic resume, or multi-replica execution.
  • Durable target jobs, invocation history, prompts, samples, token usage, cost accounting, provenance chains, target snapshots, or advanced retention controls.
  • Automatic retry or resume of individual targets beyond one technical helper retry and a new Generate Missing run.
  • Automatic fallback to a different model.
  • Writing generated text into comments of the external Workspace Database.
  • Generating logical relationships or other catalog metadata beyond Catalog Table and Catalog Column descriptions.
  • Implementing the Sensitive Data Policy. Its definition and exclusion/anonymization behavior are a required follow-up improvement.
  • Generalizing the log viewer across unrelated metadata operations. That broader concern remains related to Gitea issue #2.

Further Notes

  • The design deliberately follows ThothAI's proven simple workflow while adapting it to ThothII's asynchronous browser interaction, setup ownership, and existing Generated Description model.
  • The source-sampling disclosure is a release requirement, not merely documentation for operators.
  • completed_with_errors is reserved for a run that reaches the end after isolated technical failures. Three consecutive failures end the run as failed.
  • Successful values are their own recovery record: after interruption, Generate Missing naturally skips them without needing replay state.
  • The implementation is available without a feature flag once the catalog migration and valid setup are present. With no configured model, the feature remains visibly unavailable rather than partially initialized.
  • The final manual gate is intentionally narrow: one real-provider run covers one Catalog Column and one Catalog Table, generated Italian text, safe events, and one consolidation. Automated tests remain the evidence for All, Missing, Stop, restart interruption, Unlock, and failure paths.