Files

248 lines
12 KiB
Markdown

# Components, modules, and flows
This page complements the [architecture overview](overview.md) with the module structure and flows through ThothII. The diagrams describe the current code, not a future architecture.
## Modules and dependencies
The frontend communicates with the backend through REST and SSE. The backend does not own session
persistence: it starts Pi, invokes the `tht` CLI, and forwards events. It does own the separate
installation-local database catalog. The harness contains the workflow, the Python CLI, and
adapters for the DWH and vector store. Its Memory module also owns the authoritative
PostgreSQL archive of cards, links, dependencies and pending Qdrant projections.
Administrative API calls use the same harness service as workflow producers and recall.
```mermaid
flowchart LR
FE["frontend/\nReact + Vite"] -->|REST + SSE| BE["backend/\nFastify + TypeScript"]
BE -->|RPC stdin/stdout| PI["Pi\n--mode rpc"]
BE -->|subprocess\nJSON stdout| THT["harness/tht\nCLI Python"]
PI --> EXT["harness/.pi/extensions/\ntht-gate.js"]
EXT --> SKILL["harness/.pi/skills/\ntht-sessione"]
EXT --> THT
THT --> FS["Sessions and artifacts\nworkspace repository"]
THT --> DWH["DWH\nread-only"]
THT --> VDB["Qdrant / vector store"]
THT --> MEM["thoth_memory\nPostgreSQL Memory archive"]
BE --> CFG["settings.json\nworkspace + thinking"]
BE --> MODELS["generated runtime catalog\nfrom installation YAML"]
BE --> CAT["catalog-db\nPostgreSQL + Kysely"]
BE -->|catalog Test + Sync + bounded AI sampling| DWH
BE -->|one request per subprocess| LLMHELPER["LiteLLM helper\nPython, short-lived"]
LLMHELPER -->|catalog-selected model| PROVIDER["AI provider"]
FE -.->|renders widgets| EXT
```
Dipendenze principali:
| Module | Depends on | Responsibility |
| --- | --- | --- |
| `frontend/` | Backend REST and SSE APIs | UI, gate widgets, and in-memory transcript |
| `backend/src/` | Pi, `tht`, configuration, workspace registry, catalog PostgreSQL, read-only DWH connectors, and the internal LiteLLM helper | Transport, session lifecycle, catalog CRUD, connection tests, table introspection, sequential AI description generation, and APIs |
| `harness/.pi/` | Pi and `tht phase` | Workflow orchestration and human-in-the-loop gates |
| `harness/tht/` | Filesystem, DWH, and vector store | Persistence, CLI, Evidence, schema, and preprocessing |
| workspace repository | `source/`, `curated/`, manifest, and artifacts | Versioned Evidence source and session output |
## Session sequence
The shared frontend is wrapped by `ShellProvider` (full preferences or a
replaceable portal presentation adapter), then `AuthGate` (backend identity),
then `AppShell`. Omics-specific DOM details belong only to `OmicsPortalAdapter`;
credentials and principal validation belong to the server, never that adapter.
Full/embedded do not duplicate the session workflow below. See
[rendering architecture](application-shell.md) and
[upstream identity](../install/authentication-upstream.md) for both boundaries.
The main path starts with a user question and ends with an SSE event. Reviewer decisions use the same channel and are persisted by the harness.
```mermaid
sequenceDiagram
actor U as User or reviewer
participant FE as Frontend
participant BE as Backend
participant PI as Pi RPC
participant THT as CLI tht
participant WS as Workspace
participant DWH as DWH
U->>FE: Send question or gate decision
FE->>BE: POST session / risposta widget
BE->>PI: RPC input o prompt di resume
PI->>THT: phase/session/evidence commands
THT->>WS: Read and write phase artifacts
THT->>DWH: Introspection or read-only query
DWH-->>THT: Schema, results, or diagnostics
THT-->>PI: JSON and phase state
PI-->>BE: RPC events and widget descriptor
BE-->>FE: SSE text_delta, info, ui_request
FE-->>U: Text, artifact, or review request
```
The backend uses `ThtRunner` for CLI subprocesses, `PiProcessManager` for one Pi process per session, `SessionBridge` to adapt RPC events, and `SseHub` to distribute them to clients.
## Database Management frontend
`AppShell` mounts Fleet Ledger as the default Database Management presentation. The controller keeps
the existing React Query, AG Grid, permission, dirty-state, synchronization, SSE, and polling
contracts; Fleet Ledger changes the information architecture without introducing a second catalog
client. It shows exactly one grid at a time along the database → table → column hierarchy, with
relationships as a sibling database view. An emphasized back control and breadcrumb move to the
parent view.
Selected-row operations are exposed through a single action selector and explicit **Run** control.
Row-specific actions remain icon controls in the pinned final column. Configuration, metadata
editing, synchronization, description generation, and sensitive-field review/history use the real
catalog state and open in right-side drawers. A drawer can close independently of a durable run.
The KPI strip calls `GET /catalog/metrics`: omitting `databaseId` returns installation-wide catalog
aggregates, while supplying it scopes the same aggregate contract to the selected database. The
previous renderer is reachable only as a temporary development/staging comparison with
`?db-ui=legacy` when Vite development or `VITE_DB_MANAGEMENT_LEGACY=true` enables it. The standalone
prototype on port `5173` remains outside `AppShell` only until the integrated surface is accepted.
## Catalog description generation
Catalog description generation is a backend-owned administrative operation, separate from the
Pi session workflow and from the public `tht` CLI. The frontend starts one run for selected catalog
tables or columns. A single installation-wide worker processes targets sequentially, reads at most
the configured bounded sample from the source DWH through its read-only connection, and invokes a
short-lived Python LiteLLM helper once per target. The selected model comes from the installation
descriptor; its API key remains in the protected installation secret bundle.
Each result is written immediately to `Generated Description`. Run state and sanitized activity
events are stored in `catalog-db` and exposed to the drawer through REST and SSE. An administrator
may later copy selected generated descriptions into `Description`. There is no parallel run queue
or second orchestration subsystem. A target receives at most one provider retry; three consecutive
exhausted technical batches fail the run. Stale work is marked interrupted at startup and must be
explicitly unlocked; it never resumes automatically.
Sensitivity analysis is a synchronous administrative request and does not use the installation
model catalog. Database-specific adapters stream bounded normalized values from read-only source
connections; the TypeScript `SensitivityClassifier` is the single decision point for
`sensitive | non_sensitive`. Deterministic rules run first. Tables up to 1,000 rows are fully
scanned; larger tables use breadth-first targets of 300, 1,000, and 3,000 values, with the last pass
limited to text-like columns. Source queries have five-second limits, but the request has no global
analysis deadline. An optional offline GLiNER2 worker may add NER evidence on CPU for unresolved
short text, but it cannot make or persist the decision itself.
Each attempt has its own durable run and ordered sanitized events, separate from Description
Generation because its lifecycle and counters differ. The run records the local policy version and
aggregate decision counts. Before each potentially long source-scan batch and local-NER table pass,
the classifier emits a sanitized activity event so the polling progress drawer remains visibly
active while the synchronous analysis request is pending. Coverage, rule identifiers, proposed flags, source values, NER spans,
and worker diagnostics remain transient in run history. Only an explicit administrator save changes
the human-owned Sensitive Data Flag; saving a sensitive result also persists its sanitized
Sensitivity Reason as column Catalog Metadata, while clearing the flag removes that reason.
## Main backend classes
The diagram shows the classes that form the bridge between the browser, Pi, and `tht`. Fastify routes receive requests and delegate to these services.
```mermaid
classDiagram
class ThtRunner {
+buildArgv(command, args) string[]
+run(args) Promise~ThtResult~
+sessionShow(id) Promise~unknown~
}
class PiProcessManager {
-runtimes Map
+spawnFor(sessionId, mode) SessionRuntime
+resume(sessionId, tht) Promise~SessionRuntime~
+stop(sessionId) Promise~void~
}
class SessionBridge {
+handleRpcEvent(event) ClientEvent
+handleUiResponse(response) Promise~void~
}
class SseHub {
+subscribe(sessionId) AsyncIterable
+publish(sessionId, event) void
+close(sessionId) void
}
class SessionRoutes {
+createSession(request) Response
+resumeSession(id) Response
+postInput(id, input) Response
}
class WorkspaceRegistry {
+list() Workspace[]
+resolve(id) Workspace
}
class SettingsStore {
+get() Settings
+update(patch) Settings
}
class App {
+buildApp() FastifyInstance
}
App --> SessionRoutes
App --> WorkspaceRegistry
App --> SettingsStore
SessionRoutes --> PiProcessManager
SessionRoutes --> ThtRunner
SessionRoutes --> SseHub
PiProcessManager --> SessionBridge
PiProcessManager --> ThtRunner
SessionBridge --> SseHub
```
## Python modules in the `tht` CLI
The CLI consists of Typer commands and domain modules. `cli/` turns arguments into operations; `evidence/`, `session/`, `db/`, `adapters/`, and the other packages contain the application logic.
```mermaid
flowchart TB
MAIN["tht/cli/__init__.py"] --> CMD["tht/cli/*_cmd.py"]
CMD --> CONFIG["config.py\nworkspace.py\npaths.py"]
CMD --> SESSION["session_cmd.py\nsession/"]
CMD --> EVIDENCE["evidence_cmd.py\nevidence/"]
CMD --> PRE["preprocess_cmd.py\nevidence/corpus/"]
CMD --> PHASE["phase_cmd.py\nphase.py\nworkflow.py"]
CMD --> SQL["sql_cmd.py\ndb/\nrest/"]
EVIDENCE --> ACQ["evidence/acquisition.py\nadapters/ filesystem/http/s3"]
EVIDENCE --> CANON["evidence/canonical.py\ncontracts.py\nmodel.py"]
EVIDENCE --> AUTHOR["evidence/authoring.py"]
PRE --> PIPE["evidence/corpus/pipeline.py\nchunk.py normalize.py store.py"]
PRE --> VECTOR["adapters/vector/qdrant.py"]
PRE --> DWH["jobs/dwh_pipeline.py\nadapters/dwh/"]
SESSION --> REPO["session/filesystem_repository.py\npostgres_repository.py"]
PHASE --> LEDGER["decisions.py\nreview_decisions"]
```
The operator command `tht` in `tools/tht/` is separate from the harness Python CLI. The former handles installation, lifecycle, authentication, and workspaces; the latter runs the workflow and data operations.
## Eight-phase workflow and gates
The source of truth is `harness/workflow.yaml`. The current phase is computed from the decision ledger, not from a manually updated field.
```mermaid
flowchart LR
F1["F1\nChiarimento"] --> F2["F2\nMemoria"]
F2 --> F3["F3\nRiscrittura"]
F3 --> F4["F4\nSchema linking\nreviewer_decide"]
F4 --> F5["F5\nSintesi"]
F5 --> F6["F6\nCTE\nauto o skip"]
F6 --> F7["F7\nSQL finale\nreviewer_confirm"]
F7 --> F8["F8\nDatamart\nreviewer_decide"]
F1 -.->|reviewer_confirm| F1
F3 -.->|reviewer_confirm| F3
F4 -.->|decisioni su tabelle, colonne, evidence| F4
F6 -.->|cte_approved o cte_rejected| F6
F7 -.->|sql_approved o sql_rejected| F7
F8 -.->|datamart_requested o declined| F8
```
| Fase | Nome | Avanzamento | Artefatti principali |
| --- | --- | --- | --- |
| F1 | chiarimento | `kind:phase` | decisioni di chiarimento |
| F2 | memoria | automatico se vuota | decisioni memoria |
| F3 | riscrittura | `kind:phase` | `question.md` |
| F4 | schema linking | `reviewer_decide` | `schema_linking.json` |
| F5 | sintesi | `kind:phase` | verifica dello schema linking |
| F6 | CTE | automatico, oppure skip | `cte_plan.json`, `ctes/`, `cte_tests.json` |
| F7 | SQL finale | `kind:phase` dopo `sql_approved` | `sql_final.sql` |
| F8 | datamart | `reviewer_decide` | decisione su richiesta o rifiuto |
Un `reviewer_select` con decisione incorporata può confermare direttamente. Un `reviewer_decide` registra le scelte multiple. Un `reviewer_confirm` conferma un artefatto o la chiusura della fase. Il modello propone; il revisore decide e il ledger registrato è la fonte dello stato.