feat: add AI catalog description generation
This commit is contained in:
@@ -22,7 +22,9 @@ flowchart LR
|
||||
THT --> VDB["Qdrant / vector store"]
|
||||
BE --> CFG["settings.json\nworkspace registry"]
|
||||
BE --> CAT["catalog-db\nPostgreSQL + Kysely"]
|
||||
BE -->|catalog Test + Table Sync| DWH
|
||||
BE -->|catalog Test + Sync + bounded AI sampling| DWH
|
||||
BE -->|one request per subprocess| LLMHELPER["LiteLLM helper\nPython, short-lived"]
|
||||
LLMHELPER -->|configured model| PROVIDER["AI provider"]
|
||||
FE -.->|renders widgets| EXT
|
||||
```
|
||||
|
||||
@@ -31,7 +33,7 @@ Dipendenze principali:
|
||||
| Module | Depends on | Responsibility |
|
||||
| --- | --- | --- |
|
||||
| `frontend/` | Backend REST and SSE APIs | UI, gate widgets, and in-memory transcript |
|
||||
| `backend/src/` | Pi, `tht`, configuration, workspace registry, catalog PostgreSQL, and read-only DWH connectors | Transport, session lifecycle, catalog CRUD, connection tests, table introspection, and APIs |
|
||||
| `backend/src/` | Pi, `tht`, configuration, workspace registry, catalog PostgreSQL, read-only DWH connectors, and the internal LiteLLM helper | Transport, session lifecycle, catalog CRUD, connection tests, table introspection, sequential AI description generation, and APIs |
|
||||
| `harness/.pi/` | Pi and `tht phase` | Workflow orchestration and human-in-the-loop gates |
|
||||
| `harness/tht/` | Filesystem, DWH, and vector store | Persistence, CLI, Evidence, schema, and preprocessing |
|
||||
| workspace repository | `source/`, `curated/`, manifest, and artifacts | Versioned Evidence source and session output |
|
||||
@@ -65,6 +67,20 @@ sequenceDiagram
|
||||
|
||||
The backend uses `ThtRunner` for CLI subprocesses, `PiProcessManager` for one Pi process per session, `SessionBridge` to adapt RPC events, and `SseHub` to distribute them to clients.
|
||||
|
||||
## Catalog description generation
|
||||
|
||||
Catalog description generation is a backend-owned administrative operation, separate from the
|
||||
Pi session workflow and from the public `tht` CLI. The frontend starts one run for selected catalog
|
||||
tables or columns. A single installation-wide worker processes targets sequentially, reads at most
|
||||
the configured bounded sample from the source DWH through its read-only connection, and invokes a
|
||||
short-lived Python LiteLLM helper once per target. The selected model comes from the installation
|
||||
descriptor; its API key remains in the protected installation secret bundle.
|
||||
|
||||
Each result is written immediately to `Generated Description`. Run state and sanitized activity
|
||||
events are stored in `catalog-db` and exposed to the drawer through REST and SSE. An administrator
|
||||
may later copy selected generated descriptions into `Description`. There is no parallel run queue,
|
||||
automatic retry policy, or second orchestration subsystem.
|
||||
|
||||
## Main backend classes
|
||||
|
||||
The diagram shows the classes that form the bridge between the browser, Pi, and `tht`. Fastify routes receive requests and delegate to these services.
|
||||
|
||||
Reference in New Issue
Block a user