feat: add AI catalog description generation

This commit is contained in:
Codex
2026-08-29 16:42:56 +02:00
parent b0afba81ca
commit 376dd5a09d
76 changed files with 14860 additions and 102 deletions
@@ -0,0 +1,32 @@
---
status: accepted
---
# Use one sequential description-generation run
AI description generation is an infrequent, administrator-triggered catalog operation. The
backend therefore owns one installation-wide asynchronous run and processes its model requests
sequentially. A second generation start is rejected while that run is active. The existing
catalog-operation coordinator also reserves the target Workspace Database so synchronization,
cleanup, and catalog edits cannot overlap the run.
Each model completion is performed by a short-lived internal Python process using LiteLLM. The
backend sends a structured request on stdin, reads pristine structured output from stdout, and
keeps diagnostics on stderr. This helper is neither an HTTP service nor a user-facing CLI. Using
Pi for this operation would couple a deterministic batch task to interactive session lifecycle and
gate behavior without adding product value; the helper reuses Python already present in the core
image while keeping the integration small.
Installation model entries normally reference a protected API key. A reference may be omitted
only for an explicit unauthenticated endpoint; the helper uses a fixed non-secret client
placeholder because OpenAI-compatible SDKs require a non-empty client value even when the server
ignores authentication.
For an explicitly configured Qwen-compatible endpoint, `disableThinking: true` maps to the narrow
chat-template option that prevents reasoning prose from surrounding the required JSON result.
Persistence is deliberately limited to one run record, ordered text events, and the Generated
Description written to each target as soon as it succeeds. There are no durable per-target jobs,
leases, invocation records, automatic resume, or distributed locks. On backend startup, any run
still recorded as queued or running becomes interrupted. An administrator continues by starting
Generate Missing, and may use Unlock only when no live generation process exists. Cancellation
stops the loop and terminates the current helper process.
@@ -0,0 +1,15 @@
---
status: accepted
---
# Allow bounded real source samples for description generation
Description generation may send up to five real rows and up to five distinct non-null example
values for relevant columns to the configured model provider, preserving the useful behavior of
ThothAI. Samples are read only for the current request, treated as untrusted data, bounded before
prompt construction, and never stored in generation runs, events, or catalog metadata.
This first slice makes that behavior explicit in the UI and operator documentation. A follow-up
Sensitive Data Policy is required to classify protected fields and exclude or anonymize their
values before model calls. Until that policy exists, administrators must regard generation as a
controlled disclosure of sampled source data to the selected provider.