feat: complete catalog-driven preprocessing
Publish documentation / publish (push) Successful in 2m12s

This commit is contained in:
Codex
2026-09-06 17:49:35 +02:00
parent 8707ae1d46
commit cffa60772e
141 changed files with 5898 additions and 3015 deletions
+15 -13
View File
@@ -112,18 +112,22 @@ flowchart TD
VAL -->|errors or review items| FIX["Author corrections\nand review"]
FIX --> PREP
VAL -->|publishable| COMMIT["Commit del repository\nauthoring clone"]
COMMIT --> ING["tht preprocess evidence\nnormalization and chunking"]
COMMIT --> ING["tht workspace preprocess run\nnormalization and chunking"]
ING --> BM25["BM25 index"]
ING --> VEC["Embeddings and vector store"]
BM25 --> GEN["Candidate generation"]
VEC --> GEN
GEN --> ACT["Active generation"]
GEN --> ACT["Replace active Evidence slice"]
ACT --> RUNTIME["Evidence retrieval in the workflow"]
```
Preparation can restructure changed sources, but it does not publish by itself. `prepare` produces a proposal and can identify the document involved in an error. `validate` does not write or publish. The curator publishes the revision. The runtime reads a complete, validated revision, then the pipeline creates a versioned generation. Activation is atomic, and a previous generation remains available under the retention policy.
Preparation can restructure changed sources, but it does not publish by itself. `prepare` produces a proposal and can identify the document involved in an error. `validate` does not write or publish. The curator publishes the revision. The complete workspace preprocessing command reads that exact revision and replaces the active Evidence slice. There is no application-level rollback; rerun complete preprocessing after correcting a failure.
Runtime retrieval is hybrid. The dense branch uses embeddings, the BM25 branch uses lexical search, and deterministic fusion orders the results. The published unit keeps its provenance, which the model must cite when it uses the Evidence.
Runtime retrieval is hybrid. The dense branch uses embeddings, the BM25 branch uses lexical search,
and deterministic fusion orders the results. Evidence fragments live in the workspace `reference`
collection together with Schema and relationships; runtime Memory and solved questions live in a
separate `memory` collection. The published unit keeps its provenance, which the model must cite when
it uses the Evidence.
## Author responsibilities
@@ -191,17 +195,15 @@ tht evidence evaluate <workspace-root> --config <workspace-config> --generation
tht evidence resolve <workspace-root> evidence:<id> --retire
tht evidence resolve <workspace-root> evidence:<id> --source source/domain/nuovo.md
# Materialize and index a versioned generation.
tht preprocess evidence --config <workspace-config>
# Dry run and resume a job when supported by the configuration.
tht preprocess evidence --config <workspace-config> --dry-run
tht preprocess evidence --config <workspace-config> --resume <run-id>
# Materialize Catalog-derived schema/LSH and the revision-pinned Evidence in one run.
tht --installation <absolute>/thothii-installation.yaml \
workspace preprocess run --workspace <workspace-id>
```
`evidence prepare`, `evidence migrate`, `evidence validate`, and `evidence resolve` require the
repository path. `preprocess evidence` uses the workspace configuration because it needs the
embedding, vector store, retention policy, and artifact directory.
`evidence prepare`, `evidence migrate`, `evidence validate`, and `evidence resolve` are authoring
operations and require the repository path. Runtime publication is available only through the
complete host-side `workspace preprocess run`; there is no public Evidence-only preprocessing
command.
Exit codes are part of the operating contract: `evidence validate` returns `0` when the corpus is publishable, `1` for validation errors, and `3` when only review items or orphaned units remain. With `--json`, stdout must contain valid JSON only.