12 KiB
Components, modules, and flows
This page complements the architecture overview with the module structure and flows through ThothII. The diagrams describe the current code, not a future architecture.
Modules and dependencies
The frontend communicates with the backend through REST and SSE. The backend does not own session
persistence: it starts Pi, invokes the tht CLI, and forwards events. It does own the separate
installation-local database catalog. The harness contains the workflow, the Python CLI, and
adapters for the DWH and vector store. Its Memory module also owns the authoritative
PostgreSQL archive of cards, links, dependencies and pending Qdrant projections.
Administrative API calls use the same harness service as workflow producers and recall.
flowchart LR
FE["frontend/\nReact + Vite"] -->|REST + SSE| BE["backend/\nFastify + TypeScript"]
BE -->|RPC stdin/stdout| PI["Pi\n--mode rpc"]
BE -->|subprocess\nJSON stdout| THT["harness/tht\nCLI Python"]
PI --> EXT["harness/.pi/extensions/\ntht-gate.js"]
EXT --> SKILL["harness/.pi/skills/\ntht-sessione"]
EXT --> THT
THT --> FS["Sessions and artifacts\nworkspace repository"]
THT --> DWH["DWH\nread-only"]
THT --> VDB["Qdrant / vector store"]
THT --> MEM["thoth_memory\nPostgreSQL Memory archive"]
BE --> CFG["settings.json\nworkspace + thinking"]
BE --> MODELS["generated runtime catalog\nfrom installation YAML"]
BE --> CAT["catalog-db\nPostgreSQL + Kysely"]
BE -->|catalog Test + Sync + bounded AI sampling| DWH
BE -->|one request per subprocess| LLMHELPER["LiteLLM helper\nPython, short-lived"]
LLMHELPER -->|catalog-selected model| PROVIDER["AI provider"]
FE -.->|renders widgets| EXT
Dipendenze principali:
| Module | Depends on | Responsibility |
|---|---|---|
frontend/ |
Backend REST and SSE APIs | UI, gate widgets, and in-memory transcript |
backend/src/ |
Pi, tht, configuration, workspace registry, catalog PostgreSQL, read-only DWH connectors, and the internal LiteLLM helper |
Transport, session lifecycle, catalog CRUD, connection tests, table introspection, sequential AI description generation, and APIs |
harness/.pi/ |
Pi and tht phase |
Workflow orchestration and human-in-the-loop gates |
harness/tht/ |
Filesystem, DWH, and vector store | Persistence, CLI, Evidence, schema, and preprocessing |
| workspace repository | source/, curated/, manifest, and artifacts |
Versioned Evidence source and session output |
Session sequence
The shared frontend is wrapped by ShellProvider (full preferences or a
replaceable portal presentation adapter), then AuthGate (backend identity),
then AppShell. Omics-specific DOM details belong only to OmicsPortalAdapter;
credentials and principal validation belong to the server, never that adapter.
Full/embedded do not duplicate the session workflow below. See
rendering architecture and
upstream identity for both boundaries.
The main path starts with a user question and ends with an SSE event. Reviewer decisions use the same channel and are persisted by the harness.
sequenceDiagram
actor U as User or reviewer
participant FE as Frontend
participant BE as Backend
participant PI as Pi RPC
participant THT as CLI tht
participant WS as Workspace
participant DWH as DWH
U->>FE: Send question or gate decision
FE->>BE: POST session / risposta widget
BE->>PI: RPC input o prompt di resume
PI->>THT: phase/session/evidence commands
THT->>WS: Read and write phase artifacts
THT->>DWH: Introspection or read-only query
DWH-->>THT: Schema, results, or diagnostics
THT-->>PI: JSON and phase state
PI-->>BE: RPC events and widget descriptor
BE-->>FE: SSE text_delta, info, ui_request
FE-->>U: Text, artifact, or review request
The backend uses ThtRunner for CLI subprocesses, PiProcessManager for one Pi process per session, SessionBridge to adapt RPC events, and SseHub to distribute them to clients.
Database Management frontend
AppShell mounts Fleet Ledger as the default Database Management presentation. The controller keeps
the existing React Query, AG Grid, permission, dirty-state, synchronization, SSE, and polling
contracts; Fleet Ledger changes the information architecture without introducing a second catalog
client. It shows exactly one grid at a time along the database → table → column hierarchy, with
relationships as a sibling database view. An emphasized back control and breadcrumb move to the
parent view.
Selected-row operations are exposed through a single action selector and explicit Run control. Row-specific actions remain icon controls in the pinned final column. Configuration, metadata editing, synchronization, description generation, and sensitive-field review/history use the real catalog state and open in right-side drawers. A drawer can close independently of a durable run.
The KPI strip calls GET /catalog/metrics: omitting databaseId returns installation-wide catalog
aggregates, while supplying it scopes the same aggregate contract to the selected database. The
previous renderer is reachable only as a temporary development/staging comparison with
?db-ui=legacy when Vite development or VITE_DB_MANAGEMENT_LEGACY=true enables it. The standalone
prototype on port 5173 remains outside AppShell only until the integrated surface is accepted.
Catalog description generation
Catalog description generation is a backend-owned administrative operation, separate from the
Pi session workflow and from the public tht CLI. The frontend starts one run for selected catalog
tables or columns. A single installation-wide worker processes targets sequentially, reads at most
the configured bounded sample from the source DWH through its read-only connection, and invokes a
short-lived Python LiteLLM helper once per target. The selected model comes from the installation
descriptor; its API key remains in the protected installation secret bundle.
Each result is written immediately to Generated Description. Run state and sanitized activity
events are stored in catalog-db and exposed to the drawer through REST and SSE. An administrator
may later copy selected generated descriptions into Description. There is no parallel run queue
or second orchestration subsystem. A target receives at most one provider retry; three consecutive
exhausted technical batches fail the run. Stale work is marked interrupted at startup and must be
explicitly unlocked; it never resumes automatically.
Sensitivity analysis is a synchronous administrative request and does not use the installation
model catalog. Database-specific adapters stream bounded normalized values from read-only source
connections; the TypeScript SensitivityClassifier is the single decision point for
sensitive | non_sensitive. Deterministic rules run first. Tables up to 1,000 rows are fully
scanned; larger tables use breadth-first targets of 300, 1,000, and 3,000 values, with the last pass
limited to text-like columns. Source queries have five-second limits, but the request has no global
analysis deadline. An optional offline GLiNER2 worker may add NER evidence on CPU for unresolved
short text, but it cannot make or persist the decision itself.
Each attempt has its own durable run and ordered sanitized events, separate from Description Generation because its lifecycle and counters differ. The run records the local policy version and aggregate decision counts. Before each potentially long source-scan batch and local-NER table pass, the classifier emits a sanitized activity event so the polling progress drawer remains visibly active while the synchronous analysis request is pending. Coverage, rule identifiers, proposed flags, source values, NER spans, and worker diagnostics remain transient in run history. Only an explicit administrator save changes the human-owned Sensitive Data Flag; saving a sensitive result also persists its sanitized Sensitivity Reason as column Catalog Metadata, while clearing the flag removes that reason.
Main backend classes
The diagram shows the classes that form the bridge between the browser, Pi, and tht. Fastify routes receive requests and delegate to these services.
classDiagram
class ThtRunner {
+buildArgv(command, args) string[]
+run(args) Promise~ThtResult~
+sessionShow(id) Promise~unknown~
}
class PiProcessManager {
-runtimes Map
+spawnFor(sessionId, mode) SessionRuntime
+resume(sessionId, tht) Promise~SessionRuntime~
+stop(sessionId) Promise~void~
}
class SessionBridge {
+handleRpcEvent(event) ClientEvent
+handleUiResponse(response) Promise~void~
}
class SseHub {
+subscribe(sessionId) AsyncIterable
+publish(sessionId, event) void
+close(sessionId) void
}
class SessionRoutes {
+createSession(request) Response
+resumeSession(id) Response
+postInput(id, input) Response
}
class WorkspaceRegistry {
+list() Workspace[]
+resolve(id) Workspace
}
class SettingsStore {
+get() Settings
+update(patch) Settings
}
class App {
+buildApp() FastifyInstance
}
App --> SessionRoutes
App --> WorkspaceRegistry
App --> SettingsStore
SessionRoutes --> PiProcessManager
SessionRoutes --> ThtRunner
SessionRoutes --> SseHub
PiProcessManager --> SessionBridge
PiProcessManager --> ThtRunner
SessionBridge --> SseHub
Python modules in the tht CLI
The CLI consists of Typer commands and domain modules. cli/ turns arguments into operations; evidence/, session/, db/, adapters/, and the other packages contain the application logic.
flowchart TB
MAIN["tht/cli/__init__.py"] --> CMD["tht/cli/*_cmd.py"]
CMD --> CONFIG["config.py\nworkspace.py\npaths.py"]
CMD --> SESSION["session_cmd.py\nsession/"]
CMD --> EVIDENCE["evidence_cmd.py\nevidence/"]
CMD --> PRE["preprocess_cmd.py\nevidence/corpus/"]
CMD --> PHASE["phase_cmd.py\nphase.py\nworkflow.py"]
CMD --> SQL["sql_cmd.py\ndb/\nrest/"]
EVIDENCE --> ACQ["evidence/acquisition.py\nadapters/ filesystem/http/s3"]
EVIDENCE --> CANON["evidence/canonical.py\ncontracts.py\nmodel.py"]
EVIDENCE --> AUTHOR["evidence/authoring.py"]
PRE --> PIPE["evidence/corpus/pipeline.py\nchunk.py normalize.py store.py"]
PRE --> VECTOR["adapters/vector/qdrant.py"]
PRE --> DWH["jobs/dwh_pipeline.py\nadapters/dwh/"]
SESSION --> REPO["session/filesystem_repository.py\npostgres_repository.py"]
PHASE --> LEDGER["decisions.py\nreview_decisions"]
The operator command tht in tools/tht/ is separate from the harness Python CLI. The former handles installation, lifecycle, authentication, and workspaces; the latter runs the workflow and data operations.
Eight-phase workflow and gates
The source of truth is harness/workflow.yaml. The current phase is computed from the decision ledger, not from a manually updated field.
flowchart LR
F1["F1\nChiarimento"] --> F2["F2\nMemoria"]
F2 --> F3["F3\nRiscrittura"]
F3 --> F4["F4\nSchema linking\nreviewer_decide"]
F4 --> F5["F5\nSintesi"]
F5 --> F6["F6\nCTE\nauto o skip"]
F6 --> F7["F7\nSQL finale\nreviewer_confirm"]
F7 --> F8["F8\nDatamart\nreviewer_decide"]
F1 -.->|reviewer_confirm| F1
F3 -.->|reviewer_confirm| F3
F4 -.->|decisioni su tabelle, colonne, evidence| F4
F6 -.->|cte_approved o cte_rejected| F6
F7 -.->|sql_approved o sql_rejected| F7
F8 -.->|datamart_requested o declined| F8
| Fase | Nome | Avanzamento | Artefatti principali |
|---|---|---|---|
| F1 | chiarimento | kind:phase |
decisioni di chiarimento |
| F2 | memoria | automatico se vuota | decisioni memoria |
| F3 | riscrittura | kind:phase |
question.md |
| F4 | schema linking | reviewer_decide |
schema_linking.json |
| F5 | sintesi | kind:phase |
verifica dello schema linking |
| F6 | CTE | automatico, oppure skip | cte_plan.json, ctes/, cte_tests.json |
| F7 | SQL finale | kind:phase dopo sql_approved |
sql_final.sql |
| F8 | datamart | reviewer_decide |
decisione su richiesta o rifiuto |
Un reviewer_select con decisione incorporata può confermare direttamente. Un reviewer_decide registra le scelte multiple. Un reviewer_confirm conferma un artefatto o la chiusura della fase. Il modello propone; il revisore decide e il ledger registrato è la fonte dello stato.