feat: add AI catalog description generation

This commit is contained in:
Codex
2026-08-29 16:42:56 +02:00
parent b0afba81ca
commit 376dd5a09d
76 changed files with 14860 additions and 102 deletions
+34 -4
View File
@@ -79,7 +79,8 @@ hand, but administrators can explicitly clear catalog tables, columns, or relati
touching the source database, binding, configuration, or secrets. Table deletion cascades through
columns and relationships; table-scoped relationship cleanup includes incoming and outgoing
relationships. Curated and generated descriptions are editable; generated descriptions start null
and AI generation/consolidation is deferred.
and Database Management can generate or consolidate them for selected tables, selected columns,
all targets, or only targets whose Generated Description is missing.
Schema refresh is one durable asynchronous engine with database-table, database-column,
selected-table-column, relationship, and full-database actions. Database-level menus expose the
@@ -103,10 +104,36 @@ one-shot `catalog-migrate` operation; `scripts/run-stack.sh` runs it before loca
sessions still consume the existing workspace configuration in this slice: database-management
records do not yet change the NL→SQL handoff. The accepted design is recorded in
`docs/plans/2026-08-26-metadata-catalog-from-thothai.md`, the snapshot contract under
`docs/contracts/`, and ADRs 0001–0008.
`docs/contracts/`, and ADRs 0001–0010.
Semantic aliases, value descriptions, synonyms, concepts, AI metadata generation/consolidation,
and logical relationships remain deferred to their dedicated slices.
Semantic aliases, value descriptions, synonyms, concepts, and logical relationships remain
deferred to their dedicated slices.
AI Description Generation preserves ThothAI's use of real source samples: up to five source rows
and five representative non-null example values may be sent transiently to the configured model
provider. The UI and operator documentation disclose this behavior. A required follow-up
improvement is a Sensitive Data Policy that classifies protected fields and excludes or anonymizes
their values before model calls.
The accepted AI-description design is recorded in
`docs/plans/2026-08-28-ai-catalog-description-generation.md`, with the formal specification in the
adjacent `-spec.md` document and Gitea issue #4. Gitea issues #5–#11 deliver the implementation.
The runtime deliberately keeps ThothAI's simple operating model: one installation-wide sequential
run owned by the backend, one short-lived Python/LiteLLM completion helper per request, and
persistence limited to the run, its safe ordered text events, and each Generated Description as
soon as it succeeds. The helper performs at most one provider retry and never falls back to another
model. Stop terminates the current helper and retains prior results; three consecutive exhausted
technical batches fail the run. Startup marks stale queued/running work interrupted, and Unlock is
available only when no local start, worker, or helper is live. Runs remain inspectable through a
live SSE log with ordered polling fallback; there is no automatic resume or user-facing generation
CLI. ADRs 0009–0010 record the runtime and source-sampling decisions.
Metadata-generation setup accepts the protected `DEEPSEEK_API_KEY` and `ZAI_API_KEY` references.
It also accepts a model with no secret reference only when its OpenAI-compatible endpoint is
explicit; this covers the VPN-only AritmoLab Qwen 3.6 server without creating a fake operator
credential. The Python client supplies only its fixed non-secret compatibility placeholder.
The AritmoLab entry also sets `disableThinking: true`, mapped to the endpoint's chat-template flag,
because its default reasoning prose would violate the worker's exact JSON response contract.
Integration of the completed metadata catalog with core schema-linking is explicitly deferred
until the database, table, column, relationship, and synchronization slices are complete. At that
@@ -159,6 +186,9 @@ from an automated PASS.
- The P1.1 workspace-directory registry and P2–P6 preprocessing workstreams are implemented and
have automated coverage.
- Evidence restructuring has a real PSD acceptance PASS as recorded above.
- AI Description Generation has automated coverage across installation setup, model selection,
generation/consolidation scopes, bounded sampling, cancellation/recovery, history, SSE/polling,
and the LiteLLM helper boundary.
- L2 tests requiring real providers or remote databases remain opt-in.
- Server deployment, release, and owner-operated acceptance steps remain pending wherever the
referenced runbooks require explicit approval.