250 lines
22 KiB
Markdown
250 lines
22 KiB
Markdown
# AI-generated descriptions for Catalog Tables and Catalog Columns
|
|
|
|
## Problem Statement
|
|
|
|
ThothII already stores a Generated Description separately from the curated Description for Catalog
|
|
Tables and Catalog Columns, but administrators cannot populate it with AI. ThothAI provides the
|
|
useful core workflow—generate table and column comments from schema context and small real-data
|
|
samples—but its execution, configuration, and interaction model cannot be copied directly into
|
|
ThothII.
|
|
|
|
Administrators need an asynchronous workflow integrated into Database Management. They must be
|
|
able to choose an installation-approved model, generate descriptions for selected or missing
|
|
targets, observe understandable progress, stop or recover a stuck operation, review generated
|
|
text, and explicitly consolidate it. The solution must retain ThothAI's practical simplicity and
|
|
must not introduce a general job platform, model gateway, distributed scheduler, or competing
|
|
user-facing CLI.
|
|
|
|
## Solution
|
|
|
|
Add Description Generation to Database Management as one installation-wide, sequential background
|
|
run owned by the Fastify backend. The browser starts a run and remains responsive while the backend
|
|
processes bounded requests one at a time. Each completion is delegated to a short-lived internal
|
|
Python helper using LiteLLM. Models, their default, and any API-key secret references are declared in
|
|
application setup YAML and are independent of both workspaces and Pi configuration.
|
|
|
|
Each valid result is written immediately to the target's Generated Description. A minimal run row
|
|
and ordered text events provide status, counters, history, and a live log. A stopped or crashed run
|
|
is not resumed automatically; completed results remain in place and Generate Missing supplies the
|
|
simple recovery path. An Unlock action marks a stale recorded run interrupted only when no helper
|
|
or backend generation loop is alive.
|
|
|
|
Prompts use catalog context and, when available, no more than five real rows and five representative
|
|
non-null examples. Samples are transient and never logged or persisted. A valid inability to infer
|
|
a description produces a standard application-localized value such as `Non generabile`; provider,
|
|
timeout, and response-validation failures remain technical errors.
|
|
|
|
Generated text remains separate from Description until an administrator uses the existing
|
|
checkbox selection and Actions control to consolidate it. Consolidation retains Generated
|
|
Description and never writes comments to the external Workspace Database.
|
|
|
|
## User Stories
|
|
|
|
1. As an installation operator, I want to declare the models allowed for metadata generation in setup YAML, so that model availability is controlled centrally.
|
|
2. As an installation operator, I want to declare one default metadata-generation model, so that administrators begin with a safe operational choice.
|
|
3. As an installation operator, I want each model to reference its own API-key secret, so that credentials are not stored in workspaces or browser-visible settings.
|
|
4. As an installation operator, I want metadata-generation models to remain independent of Pi models, so that changing this workflow cannot disrupt the core NL-to-SQL experience.
|
|
5. As an installation operator, I want invalid model setup to fail validation clearly, so that the application does not start with ambiguous provider behavior.
|
|
6. As a Catalog Administrator, I want generation controls to explain when no model is configured, so that I know why the action is unavailable.
|
|
7. As a Catalog Administrator, I want to select an approved model from a selector initialized to the setup default, so that I control which model performs the work.
|
|
8. As a Catalog Administrator, I want to generate descriptions for selected Catalog Tables, so that I can work on a focused part of the catalog.
|
|
9. As a Catalog Administrator, I want to generate descriptions for selected Catalog Columns, so that I can work on individual fields without regenerating a whole table.
|
|
10. As a Catalog Administrator, I want to generate all eligible descriptions for a Workspace Database, so that I can initialize a catalog in one operation.
|
|
11. As a Catalog Administrator, I want to generate only missing descriptions, so that I can continue interrupted work without replacing completed proposals.
|
|
12. As a Catalog Administrator, I want a full run to process Catalog Columns before their Catalog Tables, so that table descriptions can benefit from column descriptions.
|
|
13. As a Catalog Administrator, I want generation to run asynchronously after I start it, so that the browser remains usable and progress is not tied to one HTTP request.
|
|
14. As a Catalog Administrator, I want only one Description Generation Run active in the installation, so that provider traffic and operational behavior remain predictable.
|
|
15. As a Catalog Administrator, I want a second start attempt to return a clear conflict, so that I cannot accidentally overlap generation runs.
|
|
16. As a Catalog Administrator, I want synchronization, cleanup, consolidation, and edits for the target Workspace Database blocked during generation, so that the simple sequential run sees stable catalog state.
|
|
17. As a Catalog Administrator, I want to see the run's model, scope, status, counters, and timestamps, so that I understand what is happening.
|
|
18. As a Catalog Administrator, I want a chronological text log, so that I can follow completed targets and diagnose errors.
|
|
19. As a Catalog Administrator, I want live log updates with a polling fallback, so that temporary SSE problems do not hide run progress.
|
|
20. As a Catalog Administrator, I want completed and interrupted runs to remain inspectable, so that I can understand prior activity.
|
|
21. As a Catalog Administrator, I want to stop an active run, so that I can halt an incorrect or unexpectedly costly operation.
|
|
22. As a Catalog Administrator, I want stopping a run to terminate its current model helper and prevent later targets from starting, so that stop has prompt operational effect.
|
|
23. As a Catalog Administrator, I want valid results completed before a stop or failure to remain saved, so that useful work is not discarded.
|
|
24. As a Catalog Administrator, I want a run left active by a backend restart to become interrupted, so that the UI does not claim nonexistent work is still running.
|
|
25. As a Catalog Administrator, I want to unlock a stale active run when no generation process is alive, so that an erroneous recorded lock cannot block future work.
|
|
26. As a Catalog Administrator, I want Unlock rejected while a live generation process exists, so that recovery cannot create an overlapping run.
|
|
27. As a Catalog Administrator, I want Generate Missing to continue after interruption, so that recovery does not require a special resume mechanism.
|
|
28. As a Catalog Administrator, I want one retry for a transient model failure, so that a brief provider fault does not immediately lose a batch.
|
|
29. As a Catalog Administrator, I want the run to fail after three consecutive technical failures, so that a broken provider does not generate an unbounded stream of attempts.
|
|
30. As a Catalog Administrator, I want a successful request to reset the consecutive-failure count, so that isolated errors do not prematurely stop a useful run.
|
|
31. As a Catalog Administrator, I want a completed-with-errors result when isolated batches fail but the run reaches its end, so that partial problems remain visible.
|
|
32. As a Catalog Administrator, I want no automatic fallback to a different model, so that the selected model remains truthful and predictable.
|
|
33. As a Catalog Administrator, I want malformed or ambiguous model output rejected without writing it, so that descriptions cannot be assigned to the wrong target.
|
|
34. As a Catalog Administrator, I want an inability to infer a description represented by standard localized text, so that every valid outcome is understandable in the workspace language.
|
|
35. As a Catalog Administrator, I want technical failures kept distinct from non-generatable outcomes, so that provider problems are not mistaken for catalog knowledge.
|
|
36. As a Catalog Administrator, I want generated prose written in the workspace language, so that it matches the catalog's intended audience.
|
|
37. As a Catalog Administrator, I want generated text stored separately from curated Description, so that AI output remains a reviewable proposal.
|
|
38. As a Catalog Administrator, I want to edit a Generated Description manually, so that I can improve a proposal before consolidation.
|
|
39. As a Catalog Administrator, I want to select one or more tables or columns and run “Move generated description to Description” from the existing Actions control, so that review remains integrated into the current grids.
|
|
40. As a Catalog Administrator, I want consolidation to retain the Generated Description, so that I can still see the proposal from which the curated text was copied.
|
|
41. As a Catalog Administrator, I want selected records without a Generated Description skipped and reported, so that the bulk action does not erase curated text.
|
|
42. As a Catalog Administrator, I want consolidation and generation to modify only the Metadata Catalog, so that no external database comment is changed.
|
|
43. As a Catalog Administrator, I want prompts to use schema facts and existing catalog text, so that generated descriptions are grounded in available metadata.
|
|
44. As a Catalog Administrator, I want prompts to use at most five real source rows and five representative values when available, so that the model has useful examples without unbounded disclosure.
|
|
45. As a Catalog Administrator, I want to be warned that real source samples are sent to the selected provider, so that I can make an informed disclosure decision.
|
|
46. As a Catalog Administrator, I want sampled rows and values excluded from persistence and logs, so that operational history does not become a secondary data store.
|
|
47. As a security operator, I want API keys, prompts, samples, and complete provider payloads redacted from logs, so that diagnostics do not leak secrets or source data.
|
|
48. As a support operator, I want concise per-target and per-batch event messages, so that failures can be diagnosed without provider-specific internals.
|
|
49. As an authorized administrator, I want all generation, cancellation, unlock, and consolidation actions protected by database-management permission, so that ordinary users cannot mutate catalog metadata.
|
|
50. As an unauthorized user, I want generation controls hidden or disabled and API calls rejected, so that frontend visibility is not treated as authorization.
|
|
51. As an operator, I want setup changes to take effect after an application restart, so that configuration lifecycle remains simple and explicit.
|
|
52. As a product owner, I want the first release to avoid queues, parallel calls, distributed locks, and automatic resume, so that effort remains focused on generating and reviewing useful descriptions.
|
|
|
|
## Implementation Decisions
|
|
|
|
- The Fastify backend owns one installation-wide Description Generation Run and its sequential
|
|
processing loop. It does not delegate lifecycle ownership to Pi or Python.
|
|
- A Description Generation Run has one of `queued`, `running`, `completed`,
|
|
`completed_with_errors`, `cancelled`, `failed`, or `interrupted`. It stores the Workspace
|
|
Database, requested scope, selected model identifier, workspace language, progress counters,
|
|
timestamps, and an optional final error summary.
|
|
- Ordered Description Generation Events store timestamp, severity, and safe human-readable text.
|
|
No durable per-target jobs, model invocation rows, prompt snapshots, sample snapshots, leases,
|
|
heartbeats, registry revisions, or provenance chains are introduced.
|
|
- Starting a run schedules an in-process background loop and returns the run immediately. The API
|
|
exposes start, run/history lookup, event listing and streaming, cancellation, stale-run unlock,
|
|
and the safe list of configured model choices. There are no retry-item or resume endpoints.
|
|
- One in-memory generation manager enforces the installation-wide active-run rule. The existing
|
|
Catalog Operation Coordinator reserves the target Workspace Database for the duration of the
|
|
run, without being generalized into a new operation framework.
|
|
- On backend startup, persisted `queued` or `running` Description Generation Runs become
|
|
`interrupted`. The application performs no automatic replay or resume.
|
|
- Unlock succeeds only when no live generation loop or helper child exists. It marks the stale run
|
|
interrupted and releases the local reservation; it is not a distributed lock recovery protocol.
|
|
- Each model completion uses a short-lived Python helper backed by LiteLLM. Structured input is
|
|
supplied over stdin, structured output alone is emitted on stdout, diagnostics use stderr, and
|
|
the helper can be terminated by cancellation.
|
|
- A completion request contains no more than ten targets. Requests run one at a time. The helper
|
|
performs at most one retry for a transient technical failure.
|
|
- Three consecutive model-request failures fail the run. A successful request resets that count.
|
|
Isolated exhausted failures may be logged and skipped, producing `completed_with_errors` if the
|
|
run later reaches its end.
|
|
- Every valid generated or non-generatable result is applied immediately to Generated Description.
|
|
Earlier writes are retained after cancellation, interruption, or later failure.
|
|
- A response must identify requested targets unambiguously and classify each returned result as
|
|
generated or non-generatable. Duplicate, unknown, missing, or malformed mappings cause a
|
|
technical request failure and no result from that ambiguous response is applied.
|
|
- The parser also tolerates one JSON object enclosed by one complete `json` code fence, because
|
|
some supported models add that formatting despite the prompt. Any prose outside the fence,
|
|
multiple payloads, or malformed/ambiguous mappings remain invalid.
|
|
- The application supplies localized standard non-generatable text. Provider wording is not used
|
|
as the standard value, and technical errors never write that value.
|
|
- A full-database run generates eligible Catalog Columns before Catalog Tables. Generate Missing
|
|
excludes targets whose Generated Description is already non-empty; all-generation may replace
|
|
existing generated proposals only after the initiating action makes that scope explicit.
|
|
- Model choices are declared under a metadata-generation section in installation setup YAML. Each
|
|
choice has a stable identifier, display label, LiteLLM provider/model settings, optional endpoint
|
|
settings, and an optional environment-secret reference for its API key. The reference may be
|
|
omitted only when an explicit endpoint is configured for unauthenticated access. One identifier
|
|
is the default.
|
|
- An explicit endpoint may set `disableThinking: true`; the helper translates it only to the
|
|
Qwen-compatible chat-template switch needed to keep the response within the strict JSON contract.
|
|
- Metadata-generation setup is separate from application settings for Pi and from workspace
|
|
`llm_policy`. Raw keys never enter setup YAML, the catalog database, API responses, process
|
|
arguments, or event text. Configuration reload is restart-only.
|
|
- If setup defines no usable model, the safe model-list response is empty and the UI disables
|
|
generation with an explanation. The backend still rejects direct generation attempts.
|
|
- Prompt construction treats schema names, comments, descriptions, and values as untrusted data.
|
|
It requests output in the workspace language and separates instructions from catalog content.
|
|
- A request may contain up to five real source rows and up to five representative distinct,
|
|
non-null values for relevant columns. Inputs are bounded before prompt construction and are not
|
|
persisted or logged.
|
|
- The UI discloses that real data can be sent to the selected provider. A future Sensitive Data
|
|
Policy will classify values and exclude or anonymize protected data; that policy is not silently
|
|
approximated in this slice.
|
|
- The generation UI reuses Database Management's table and column selections, model selector,
|
|
Actions control, run drawer conventions, SSE delivery, and polling fallback where practical.
|
|
Visual parity with Catalog Sync Run logs is not required.
|
|
- The consolidation action copies each selected, non-empty Generated Description into Description
|
|
in a catalog transaction, retains Generated Description, skips empty proposals, and reports
|
|
copied and skipped counts. It never writes to the external Workspace Database.
|
|
- Generation, cancellation, unlock, and consolidation require the existing database-management
|
|
permission and are validated by the backend independently of UI state.
|
|
- No user-facing generation CLI is added. The Python process is an internal completion adapter,
|
|
not an operator surface or a long-lived service.
|
|
|
|
## Testing Decisions
|
|
|
|
- Tests assert externally observable behavior rather than private loop structure, process timing,
|
|
or LiteLLM implementation details.
|
|
- The primary and highest test seam is the Fastify catalog API with a test PostgreSQL catalog and
|
|
an injected fake Model Completer. It verifies complete paths through authorization, run
|
|
persistence, sequential processing, event delivery, Generated Description updates, and final
|
|
status without contacting a real provider.
|
|
- API tests cover each generation scope, column-before-table order, the ten-target request bound,
|
|
model validation, one-active-run conflict, target-database exclusion, cancellation, startup
|
|
interruption, Unlock safeguards, Generate Missing, partial success, consecutive failure
|
|
handling, non-generatable localization, malformed responses, redacted events, and permissions.
|
|
- Catalog repository integration tests verify the migration, run and event ordering, active-run
|
|
constraint, immediate description writes, history queries, startup interruption, and bulk
|
|
consolidation behavior against PostgreSQL.
|
|
- The Python helper has a small black-box contract suite using a simulated LiteLLM adapter. It
|
|
verifies stdin/stdout framing, pristine stdout, stderr diagnostics, normalized success and
|
|
failure output, one transient retry, secret redaction, and termination behavior.
|
|
- Setup-validation tests cover duplicate model identifiers, missing or unknown defaults, malformed
|
|
provider settings, missing secret references, safe public model projection, and strict separation
|
|
from Pi and workspace model settings.
|
|
- Database Management tests use the existing browser-level component seam with MSW. They verify
|
|
model selection and default, selected/all/missing actions, disabled state without models, running
|
|
progress and logs, polling recovery, cancellation, Unlock visibility, terminal summaries,
|
|
generated-text refresh, and selected consolidation with copied/skipped counts.
|
|
- Existing Catalog Sync Run route, repository, SSE, and drawer tests are prior art for asynchronous
|
|
status and event behavior. Existing catalog table/column editing and Database Management tests
|
|
are prior art for optimistic catalog updates, permissions, selection, and action controls.
|
|
- One required manual acceptance gate, outside deterministic CI, uses the installation's configured
|
|
default model and a disposable PostgreSQL database containing only invented data. Its application
|
|
credentials are read-only. It generates Italian text for one Catalog Column and one Catalog
|
|
Table, verifies their Generated Description, inspects the safe activity log, confirms that no key
|
|
or sample value is exposed, and consolidates one selected result. If the configured secret is not
|
|
available, acceptance stops without exposing or requesting the key in conversation.
|
|
- Delivery includes a strict MkDocs build executed through repository-managed, reproducible
|
|
documentation dependencies rather than globally installed Python packages. A readable direct
|
|
dependency file is retained, a complete transitive lock is generated with `uv`, and one canonical
|
|
repository command performs the strict build from that lock.
|
|
- Successful real-provider acceptance is recorded in a short sanitized report under
|
|
`docs/testing/`. It identifies the model and checks performed but contains no credentials,
|
|
prompts, source samples, complete provider payloads, or generated database values.
|
|
- No tests are added for worker queues, parallel generation, distributed locking, multi-replica
|
|
recovery, automatic resume, cost accounting, or model fallback because those behaviors are out
|
|
of scope.
|
|
|
|
## Out of Scope
|
|
|
|
- Reusing Pi to execute Description Generation or changing Pi's model configuration.
|
|
- A shared Installation Model Registry, model gateway, long-lived Python sidecar, or provider
|
|
management platform.
|
|
- A user-facing generation CLI.
|
|
- Parallel model calls, worker queues, adaptive rate limiting, distributed locks, leases,
|
|
heartbeats, automatic resume, or multi-replica execution.
|
|
- Durable target jobs, invocation history, prompts, samples, token usage, cost accounting,
|
|
provenance chains, target snapshots, or advanced retention controls.
|
|
- Automatic retry or resume of individual targets beyond one technical helper retry and a new
|
|
Generate Missing run.
|
|
- Automatic fallback to a different model.
|
|
- Writing generated text into comments of the external Workspace Database.
|
|
- Generating logical relationships or other catalog metadata beyond Catalog Table and Catalog
|
|
Column descriptions.
|
|
- Implementing the Sensitive Data Policy. Its definition and exclusion/anonymization behavior are
|
|
a required follow-up improvement.
|
|
- Generalizing the log viewer across unrelated metadata operations. That broader concern remains
|
|
related to Gitea issue #2.
|
|
|
|
## Further Notes
|
|
|
|
- The design deliberately follows ThothAI's proven simple workflow while adapting it to ThothII's
|
|
asynchronous browser interaction, setup ownership, and existing Generated Description model.
|
|
- The source-sampling disclosure is a release requirement, not merely documentation for operators.
|
|
- `completed_with_errors` is reserved for a run that reaches the end after isolated technical
|
|
failures. Three consecutive failures end the run as `failed`.
|
|
- Successful values are their own recovery record: after interruption, Generate Missing naturally
|
|
skips them without needing replay state.
|
|
- The implementation is available without a feature flag once the catalog migration and valid
|
|
setup are present. With no configured model, the feature remains visibly unavailable rather than
|
|
partially initialized.
|
|
- The final manual gate is intentionally narrow: one real-provider run covers one Catalog Column
|
|
and one Catalog Table, generated Italian text, safe events, and one consolidation. Automated
|
|
tests remain the evidence for All, Missing, Stop, restart interruption, Unlock, and failure paths.
|