diff --git a/docs/research/2026-08-23-thothii-metadata-publication-qdrant-seams.md b/docs/research/2026-08-23-thothii-metadata-publication-qdrant-seams.md new file mode 100644 index 00000000..51cc13ef --- /dev/null +++ b/docs/research/2026-08-23-thothii-metadata-publication-qdrant-seams.md @@ -0,0 +1,255 @@ +# ThothII metadata publication and Qdrant revision seams + +**Research question:** Which existing workspace, snapshot, preprocessing, Qdrant, +session-pinning, and runtime read-only contracts constrain metadata publication without changing +the NL→SQL workflow? + +## Conclusion + +The lowest-impact publication seam is **not** a new writer inside the NL→SQL workflow and it is +not a direct CRUD-to-Qdrant path. The existing core already consumes two canonical schema inputs: +an introspected `PhysicalSchema` and a curated `Annotations` document. An explicit publication +operation should project the approved, workspace-selected subset of the new metadata catalog into +those inputs, bind the projection to a new Git `workspace_revision`, activate that immutable +revision, and then invoke the existing `workspace index-schema` preprocessing operation. Qdrant +remains a derived, rebuildable projection. + +This preserves the existing runtime path: + +```text +approved catalog data + -> workspace Git revision (physical schema selection + annotations) + -> immutable registry snapshot + -> revision-bound runtime configuration + -> existing tht vector index-schema + -> Qdrant records filtered by workspace_id + workspace_revision + -> existing retrieval_pack / schema render / F4 review +``` + +The complete database inventory may remain in the Metadata Catalog. Only the explicit Core Schema +Selection should be projected into the workspace artifacts used by the SQL workflow. ThothII does +not currently model such a table-level selection in its descriptor, so that projection contract is +new work. + +## 1. Current authority and publication boundary + +The workspace descriptor identifies one PostgreSQL database and one physical schema, plus one +workspace-owned Qdrant collection; it has no table or column allowlist +([`backend/src/workspaces/schema.ts:160-178`](../../backend/src/workspaces/schema.ts#L160-L178)). +Consequently, a full-database metadata catalog and the subset eligible for the SQL core cannot be +represented as the same current descriptor object. + +The existing public contract makes the Git workspace repository curator-owned. Changes occur in a +separate authoring clone followed by installation pull, and the API does not write workspace, +schema, or Evidence paths +([`docs/contracts/workspace-evidence-v3.md:148-156`](../contracts/workspace-evidence-v3.md#L148-L156)). +Therefore a browser CRUD service cannot silently make its PostgreSQL state authoritative for the +core without either: + +1. exporting/committing a deterministic workspace projection through the existing curator flow; + or +2. deliberately replacing this Git-authority contract. + +The first option preserves current architecture and session reproducibility. + +Registry activation validates every descriptor at one exact commit, validates and synchronizes +the commit's `schema/annotations.yaml`, and rejects duplicate ownership of a Qdrant collection +([`backend/src/workspaces/registry.ts:531-576`](../../backend/src/workspaces/registry.ts#L531-L576)). +It writes the candidate snapshot into a staging directory, records file digests in +`snapshot.json`, atomically renames the directory, and only then moves active state +([`backend/src/workspaces/registry.ts:582-629`](../../backend/src/workspaces/registry.ts#L582-L629)). +That is the existing atomic publication boundary to reuse. + +## 2. Canonical schema inputs already consumed by the core + +The harness separates source facts from curated semantics: + +- `PhysicalSchema` contains database/schema identity and tables; table facts include comments, + columns, physical foreign keys, and indexes; columns contain type, nullability, primary-key, + default, source comment, examples, and eligibility + ([`harness/tht/mschema/models.py:23-64`](../../harness/tht/mschema/models.py#L23-L64)). +- `Annotations` contains curated table descriptions, concepts and notes, column descriptions, + synonyms, concepts, evidence, notes and eligibility overrides, plus logical foreign keys + ([`harness/tht/mschema/models.py:67-88`](../../harness/tht/mschema/models.py#L67-L88)). + +Rendering already gives annotations precedence over source comments and merges physical and +logical foreign keys. It also applies column eligibility before producing M-Schema context +([`harness/tht/mschema/render.py:9-43`](../../harness/tht/mschema/render.py#L9-L43), +[`harness/tht/mschema/render.py:46-80`](../../harness/tht/mschema/render.py#L46-L80)). This makes +`Annotations` the natural narrow projection target for approved descriptions, synonyms, logical +relationships, and eligibility from the new catalog. + +There are two current compatibility gaps: + +- `schema_records()` embeds table descriptions/concepts and column descriptions/synonyms/examples, + but does not include physical or logical foreign keys in vector record content + ([`harness/tht/vectorstore/records.py:101-132`](../../harness/tht/vectorstore/records.py#L101-L132)). + Relationships still reach the model through deterministic schema rendering, not through schema + candidate embeddings. +- The F4 widget loads columns with `tht schema columns` + ([`harness/.pi/extensions/tht-gate.js:1220-1254`](../../harness/.pi/extensions/tht-gate.js#L1220-L1254)), + but that command currently returns only `PhysicalSchema.comment` values and does not merge + `Annotations` + ([`harness/tht/cli/schema_cmd.py:592-621`](../../harness/tht/cli/schema_cmd.py#L592-L621)). + A publisher that writes only annotations would improve vector search and rendered M-Schema but + not the table/column descriptions displayed by this existing reviewer widget. Fixing the command + to use the existing merged description helpers would preserve the workflow shape while closing + the inconsistency. + +## 3. Existing preprocessing seam + +`workspace index-schema` is already the supported operator seam. It creates a revision-bound +runtime, checks collection compatibility, and runs the harness command +`vector index-schema --json` +([`backend/src/workspaces/preprocessing-service.ts:312-330`](../../backend/src/workspaces/preprocessing-service.ts#L312-L330)). +The full preprocessing operation performs DWH preparation, FK suggestion/review, schema indexing, +and optional Evidence preprocessing as separate resumable stages +([`backend/src/workspaces/preprocessing-service.ts:367-450`](../../backend/src/workspaces/preprocessing-service.ts#L367-L450)). +Metadata-only publication should normally use the narrow `index-schema` operation after its +workspace artifacts are valid, rather than coupling catalog CRUD to the full pipeline. + +The harness indexer reads the active immutable physical schema and revision-specific annotations, +constructs schema records, embeds only changed content, and writes them through the configured +vector adapter +([`harness/tht/cli/vector_cmd.py:100-140`](../../harness/tht/cli/vector_cmd.py#L100-L140)). Its +machine result carries the physical/annotation artifact digests, workspace revision, collection, +and counts +([`harness/tht/cli/vector_cmd.py:157-179`](../../harness/tht/cli/vector_cmd.py#L157-L179)). This is +the right place to retain publication evidence and audit linkage. + +Preprocessing state is already revision- and binding-aware. A resumed job must match operation, +workspace revision, descriptor/catalog blobs, runtime config and binding digests or it fails with +`preprocessing_resume_mismatch` +([`backend/src/workspaces/preprocessing-state.ts:262-313`](../../backend/src/workspaces/preprocessing-state.ts#L262-L313)). + +There is an important operational gate: every preprocessing operation calls +`assertSessionInventoryCompatible`; a non-finalized, non-archived session pinned to another +revision blocks preprocessing +([`backend/src/workspaces/preprocessing-service.ts:459-472`](../../backend/src/workspaces/preprocessing-service.ts#L459-L472), +[`backend/src/workspaces/preprocessing-state.ts:363-378`](../../backend/src/workspaces/preprocessing-state.ts#L363-L378)). +A catalog publication UX must expose this as a pending/blocking condition rather than report a +generic indexing failure. + +## 4. Qdrant identity, payload and collection contracts + +The collection contract is fixed at the descriptor's dimensions/distance and eight keyword payload +indexes: `content_hash`, `document_id`, `kind`, `record_key`, `record_kind`, +`vector_generation`, `workspace_id`, and `workspace_revision` +([`backend/src/workspaces/qdrant-collection.ts:1-20`](../../backend/src/workspaces/qdrant-collection.ts#L1-L20)). +`self_heal` may create a missing compatible collection or indexes; `require_existing` only validates +and refuses an incompatible collection +([`backend/src/workspaces/qdrant-collection.ts:71-104`](../../backend/src/workspaces/qdrant-collection.ts#L71-L104)). +The publication path must use this shared manager instead of inventing collection setup. + +Schema payloads include both the workspace and workspace revision, record identity, semantic kind, +content and content hash +([`harness/tht/vectorstore/records.py:30-49`](../../harness/tht/vectorstore/records.py#L30-L49)). +Qdrant point IDs for schema and Evidence also include the revision, and upserts use the same +revision in the payload +([`harness/tht/adapters/vector/qdrant.py:42-46`](../../harness/tht/adapters/vector/qdrant.py#L42-L46), +[`harness/tht/adapters/vector/qdrant.py:196-220`](../../harness/tht/adapters/vector/qdrant.py#L196-L220)). +Reads always filter by `workspace_id`; schema and Evidence reads additionally filter by the bound +`workspace_revision`, while memory and solved-question records intentionally remain +workspace-wide +([`harness/tht/adapters/vector/qdrant.py:291-308`](../../harness/tht/adapters/vector/qdrant.py#L291-L308)). + +This means an approved semantic change needs a new workspace revision if old sessions must retain +their previous view. Directly overwriting points under the same Git revision would mutate the +meaning of that supposedly immutable revision; adding an independent catalog-publication version +would require changing the current runtime filter contract. + +### Qdrant synchronization defect to resolve before catalog publication + +The current incremental synchronizer compares canonical records by content hash and upserts +changed records, but it has no deletion step +([`harness/tht/cli/vector_cmd.py:67-89`](../../harness/tht/cli/vector_cmd.py#L67-L89)). More +importantly, `QdrantVectorStore.existing_hashes()` filters by workspace and record kind but does not +apply `_revision_filter()` +([`harness/tht/adapters/vector/qdrant.py:174-194`](../../harness/tht/adapters/vector/qdrant.py#L174-L194)), +even though point IDs and reads are revision-scoped. Therefore a same-key/same-content record from +an older revision may be classified as unchanged and never written under the new revision. This is +an implementation defect/risk inferred directly from the two code paths, and it should be fixed +and regression-tested before the catalog relies on `index-schema` for multi-revision publication. + +## 5. Session pinning and why the SQL workflow can remain unchanged + +Normal session creation acquires an immutable registry revision, passes its snapshot path, +workspace ID and commit to `tht session new`, and only releases the retention lease after the +manifest has been persisted +([`backend/src/routes/sessions.ts:341-367`](../../backend/src/routes/sessions.ts#L341-L367), +[`backend/src/routes/sessions.ts:423-440`](../../backend/src/routes/sessions.ts#L423-L440)). The +manifest stores `workspace_id` and `workspace_revision` next to database/schema identity +([`harness/tht/session/store.py:101-132`](../../harness/tht/session/store.py#L101-L132)). Resume and +saved-SQL paths reopen that exact retained snapshot rather than the current installation default +([`backend/src/routes/sessions.ts:184-198`](../../backend/src/routes/sessions.ts#L184-L198), +[`backend/src/routes/sql.ts:24-43`](../../backend/src/routes/sql.ts#L24-L43)). + +The runtime renderer places the same workspace revision in `runtime_identity`, points the harness +at the descriptor-owned Qdrant collection, and supplies the internal embedding service +([`backend/src/workspaces/runtime-renderer.ts:248-277`](../../backend/src/workspaces/runtime-renderer.ts#L248-L277)). +The vector adapter is constructed directly from that configuration, including the revision +([`harness/tht/adapters/factory.py:27-42`](../../harness/tht/adapters/factory.py#L27-L42)). + +At session bootstrap the backend invokes the existing `search pack` command +([`backend/src/routes/sessions.ts:465-470`](../../backend/src/routes/sessions.ts#L465-L470)). That +command queries schema records with the existing schema kinds, ranks tables and persists the +candidate list +([`harness/tht/cli/search_cmd.py:291-310`](../../harness/tht/cli/search_cmd.py#L291-L310)). The Pi +extension reads the persisted retrieval pack through the CLI +([`harness/.pi/extensions/tht-gate.js:61-75`](../../harness/.pi/extensions/tht-gate.js#L61-L75)), +and F4 starts from those candidates while loading full table/column context through existing schema +commands. Thus catalog publication can improve the inputs without changing phases, gate semantics, +or persisted session artifacts. + +## 6. DWH reuse across metadata-only revisions + +DWH preparation uses immutable generations selected by an `ACTIVE` pointer +([`docs/contracts/tht-dwh.md:18-24`](../contracts/tht-dwh.md#L18-L24)). The effective-configuration +identity deliberately excludes `runtime_identity`, so a content-only Git revision does not force a +database re-introspection +([`docs/contracts/tht-dwh.md:74-90`](../contracts/tht-dwh.md#L74-L90)). This is the key enabling +property for metadata publication: a new revision can carry updated curated annotations/Core Schema +Selection, reuse the compatible physical DWH generation, and rebuild only the revision-scoped +schema projection. + +## 7. Constraints for the Metadata Catalog design + +The following should be treated as requirements for the architecture map: + +1. **Separate full inventory from core projection.** PostgreSQL may hold the complete database + catalog, drafts and AI-generated text. Only an explicit Core Schema Selection and approved + semantic fields are exported to core artifacts/Qdrant. +2. **Publish; do not live-link.** CRUD changes are not visible to the SQL workflow until an explicit, + audited publication succeeds. A failed projection, Git activation or Qdrant index operation + leaves the previous revision active. +3. **Use a new Git revision as the publication identity.** This preserves current snapshot, + retention, session resume and Qdrant filtering semantics. A separate mutable catalog revision + cannot be safely introduced without changing runtime reads. +4. **Reuse canonical formats.** Project physical facts/Core Schema Selection into a validated + `PhysicalSchema` view and semantic edits into `Annotations`; keep Mermaid and long-form database + documentation outside Qdrant unless a separate, explicit record kind and retrieval policy is + designed. +5. **Reuse the operator boundary.** Trigger `workspace index-schema`, observe its schema-versioned + result and persist its artifact identities. Do not call Qdrant from browser CRUD handlers. +6. **Keep runtime read-only.** The session/Pi process continues to read the pinned snapshot, + retrieval pack and Qdrant projection. AI generation and catalog writes belong to the separate + management control plane. +7. **Surface publication gates.** UI status must distinguish Git activation, incompatible/missing + collection, resumable-session revision conflict, embedding failure and completed publication. +8. **Repair and test cross-revision synchronization first.** Scope `existing_hashes()` to the bound + revision and define deletion/garbage-collection semantics before depending on repeated catalog + publication. +9. **Close the annotation display gap without changing the workflow.** Make `schema columns` read + the same merged descriptions used by M-Schema/vector rendering, so the current F4 widget sees + the approved catalog text. + +## Decision summary for the Wayfinder map + +- Keep the Metadata Catalog as a separate management subsystem and source of editable metadata. +- Keep Qdrant derived and revision-scoped; it is not the catalog database or source of truth. +- Publish through immutable workspace revisions plus the existing preprocessing service. +- Preserve the current NL→SQL workflow, Pi extension, phase model and session artifacts. +- Add a new, explicit Core Schema Selection contract because no table/column selection exists in + the current descriptor. +- Treat direct same-revision Qdrant writes, implicit CRUD publication, and bypassing the Git + snapshot boundary as rejected integration paths.