docs: revalidate metadata catalog research

This commit is contained in:
Codex
2026-08-31 14:54:52 +02:00
parent 16b7477861
commit 866aee4249
3 changed files with 345 additions and 253 deletions
@@ -4,21 +4,51 @@
session-pinning, and runtime read-only contracts constrain metadata publication without changing
the NL→SQL workflow?
**Last verified:** 2026-08-31
**Validity:** Active architectural research; the catalog-to-core integration described below is
still **deferred**, not implemented.
**Current authorities:** `PROJECT_STATE.md`,
[`workspace-evidence-v3.md`](../contracts/workspace-evidence-v3.md),
[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md),
[`tht-dwh.md`](../contracts/tht-dwh.md),
[`ADR-0001`](../adr/0001-postgres-metadata-catalog.md), and
[`ADR-0004`](../adr/0004-fastify-kysely-metadata-catalog.md). These sources and current code
override this research note if they diverge.
## Revalidation against the current implementation
| Finding | Status on 2026-08-31 | Current evidence and consequence |
| --- | --- | --- |
| PostgreSQL metadata authority in the existing Fastify backend | **Adopted/current** | ADR-0001 and ADR-0004 are implemented; the catalog is the management-plane authority. |
| Git workspace revision, immutable registry snapshot, and revision-pinned runtime | **Adopted/current** | Registry activation and the Workspace Evidence v3 contract still provide the publication boundary for runtime-owned files. |
| Qdrant read isolation for schema and Evidence by `workspace_id` + `workspace_revision` | **Adopted/current** | `QdrantVectorStore.search()` applies `_revision_filter()` to schema/Evidence; Memory and solved questions intentionally remain workspace-wide (`harness/tht/adapters/vector/qdrant.py`). |
| Physical schema as a file in the Git publication | **Superseded clarification** | `physical.yaml` is owned by the immutable `.tht-dwh` generation selected by `ACTIVE`, not by the Git workspace revision ([DWH contract](../contracts/tht-dwh.md)). Git currently supplies the revision-pinned curated `schema/annotations.yaml`. |
| Catalog-to-core publisher | **Open/deferred** | No current route or service projects catalog records into core artifacts. `PROJECT_STATE.md` explicitly defers the schema-linking integration (`PROJECT_STATE.md:160-171`). |
| Explicit Core Schema Selection | **Open/deferred** | The workspace descriptor selects one database and one physical schema, but has no table/column allowlist (`backend/src/workspaces/schema.ts`); no catalog selection is handed to `tht`. |
| Cross-revision hash lookup used by schema synchronization | **Open defect** | Reads are revision-filtered, but `existing_hashes()` is not. A same-key/same-content point from an older revision can suppress the required upsert into the new revision (`harness/tht/adapters/vector/qdrant.py`, `harness/tht/cli/vector_cmd.py`). |
| Deletion/GC | **Current for Evidence; open for schema** | Evidence has generation inventory, retention, compensation, and exact-generation deletion. `sync_canonical_records()` never deletes schema records absent from the new canonical set, and no revision-retention GC exists for schema points. |
| Annotation consumption | **Current but incomplete** | M-Schema rendering and schema embeddings consume `Annotations`; `tht schema columns` still returns only physical comments, so F4 does not display annotation descriptions (`harness/tht/cli/schema_cmd.py`, `harness/.pi/extensions/tht-gate.js`). |
## Conclusion
The lowest-impact publication seam is **not** a new writer inside the NL→SQL workflow and it is
not a direct CRUD-to-Qdrant path. The existing core already consumes two canonical schema inputs:
an introspected `PhysicalSchema` and a curated `Annotations` document. An explicit publication
operation should project the approved, workspace-selected subset of the new metadata catalog into
those inputs, bind the projection to a new Git `workspace_revision`, activate that immutable
The lowest-impact **proposed** publication seam is not a new writer inside the NL→SQL workflow and
is not a direct CRUD-to-Qdrant path. The existing core consumes an introspected `PhysicalSchema`
from the active immutable DWH generation and a curated, Git-revision-pinned `Annotations`
document. A future explicit publication operation should project the approved catalog subset into
a new Core Schema Selection contract and the curated annotations, activate a new immutable Git
revision, and then invoke the existing `workspace index-schema` preprocessing operation. Qdrant
remains a derived, rebuildable projection.
remains a derived, rebuildable projection. None of this catalog-to-core handoff exists yet;
`PROJECT_STATE.md:160-164` deliberately defers it to the next design
gate.
This preserves the existing runtime path:
The proposed flow would preserve the existing runtime path:
```text
approved catalog data
-> workspace Git revision (physical schema selection + annotations)
-> explicit Core Schema Selection + curated annotations projection (future)
-> workspace Git revision (annotations and selection contract; physical.yaml stays DWH-owned)
-> immutable registry snapshot
-> revision-bound runtime configuration
-> existing tht vector index-schema
@@ -26,23 +56,23 @@ approved catalog data
-> existing retrieval_pack / schema render / F4 review
```
The complete database inventory may remain in the Metadata Catalog. Only the explicit Core Schema
Selection should be projected into the workspace artifacts used by the SQL workflow. ThothII does
not currently model such a table-level selection in its descriptor, so that projection contract is
new work.
The complete database inventory remains in the Metadata Catalog. Only an explicit Core Schema
Selection and approved semantic fields should be projected into the artifacts used by the SQL
workflow. ThothII does not currently model or publish that table/column-level selection, so both
the projection contract and its operational publisher are new work.
## 1. Current authority and publication boundary
The workspace descriptor identifies one PostgreSQL database and one physical schema, plus one
workspace-owned Qdrant collection; it has no table or column allowlist
([`backend/src/workspaces/schema.ts:160-178`](../../backend/src/workspaces/schema.ts#L160-L178)).
(`backend/src/workspaces/schema.ts`).
Consequently, a full-database metadata catalog and the subset eligible for the SQL core cannot be
represented as the same current descriptor object.
The existing public contract makes the Git workspace repository curator-owned. Changes occur in a
separate authoring clone followed by installation pull, and the API does not write workspace,
schema, or Evidence paths
([`docs/contracts/workspace-evidence-v3.md:148-156`](../contracts/workspace-evidence-v3.md#L148-L156)).
([Workspace Evidence v3 contract](../contracts/workspace-evidence-v3.md)).
Therefore a browser CRUD service cannot silently make its PostgreSQL state authoritative for the
core without either:
@@ -54,10 +84,12 @@ The first option preserves current architecture and session reproducibility.
Registry activation validates every descriptor at one exact commit, validates and synchronizes
the commit's `schema/annotations.yaml`, and rejects duplicate ownership of a Qdrant collection
([`backend/src/workspaces/registry.ts:531-576`](../../backend/src/workspaces/registry.ts#L531-L576)).
(`backend/src/workspaces/registry.ts`,
`backend/src/workspaces/git-repository.ts`,
`backend/src/workspaces/annotations-sync.ts`).
It writes the candidate snapshot into a staging directory, records file digests in
`snapshot.json`, atomically renames the directory, and only then moves active state
([`backend/src/workspaces/registry.ts:582-629`](../../backend/src/workspaces/registry.ts#L582-L629)).
(`backend/src/workspaces/registry.ts`).
That is the existing atomic publication boundary to reuse.
## 2. Canonical schema inputs already consumed by the core
@@ -67,15 +99,14 @@ The harness separates source facts from curated semantics:
- `PhysicalSchema` contains database/schema identity and tables; table facts include comments,
columns, physical foreign keys, and indexes; columns contain type, nullability, primary-key,
default, source comment, examples, and eligibility
([`harness/tht/mschema/models.py:23-64`](../../harness/tht/mschema/models.py#L23-L64)).
(`harness/tht/mschema/models.py`).
- `Annotations` contains curated table descriptions, concepts and notes, column descriptions,
synonyms, concepts, evidence, notes and eligibility overrides, plus logical foreign keys
([`harness/tht/mschema/models.py:67-88`](../../harness/tht/mschema/models.py#L67-L88)).
(`harness/tht/mschema/models.py`).
Rendering already gives annotations precedence over source comments and merges physical and
logical foreign keys. It also applies column eligibility before producing M-Schema context
([`harness/tht/mschema/render.py:9-43`](../../harness/tht/mschema/render.py#L9-L43),
[`harness/tht/mschema/render.py:46-80`](../../harness/tht/mschema/render.py#L46-L80)). This makes
(`harness/tht/mschema/render.py`). This makes
`Annotations` the natural narrow projection target for approved descriptions, synonyms, logical
relationships, and eligibility from the new catalog.
@@ -83,14 +114,14 @@ There are two current compatibility gaps:
- `schema_records()` embeds table descriptions/concepts and column descriptions/synonyms/examples,
but does not include physical or logical foreign keys in vector record content
([`harness/tht/vectorstore/records.py:101-132`](../../harness/tht/vectorstore/records.py#L101-L132)).
(`harness/tht/vectorstore/records.py`).
Relationships still reach the model through deterministic schema rendering, not through schema
candidate embeddings.
- The F4 widget loads columns with `tht schema columns`
([`harness/.pi/extensions/tht-gate.js:1220-1254`](../../harness/.pi/extensions/tht-gate.js#L1220-L1254)),
but that command currently returns only `PhysicalSchema.comment` values and does not merge
(`harness/.pi/extensions/tht-gate.js`), but that
command still returns only `PhysicalSchema.comment` values and does not merge
`Annotations`
([`harness/tht/cli/schema_cmd.py:592-621`](../../harness/tht/cli/schema_cmd.py#L592-L621)).
(`harness/tht/cli/schema_cmd.py`).
A publisher that writes only annotations would improve vector search and rendered M-Schema but
not the table/column descriptions displayed by this existing reviewer widget. Fixing the command
to use the existing merged description helpers would preserve the workflow shape while closing
@@ -101,32 +132,33 @@ There are two current compatibility gaps:
`workspace index-schema` is already the supported operator seam. It creates a revision-bound
runtime, checks collection compatibility, and runs the harness command
`vector index-schema --json`
([`backend/src/workspaces/preprocessing-service.ts:312-330`](../../backend/src/workspaces/preprocessing-service.ts#L312-L330)).
(`backend/src/workspaces/preprocessing-service.ts`,
[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md)).
The full preprocessing operation performs DWH preparation, FK suggestion/review, schema indexing,
and optional Evidence preprocessing as separate resumable stages
([`backend/src/workspaces/preprocessing-service.ts:367-450`](../../backend/src/workspaces/preprocessing-service.ts#L367-L450)).
(`backend/src/workspaces/preprocessing-service.ts`).
Metadata-only publication should normally use the narrow `index-schema` operation after its
workspace artifacts are valid, rather than coupling catalog CRUD to the full pipeline.
The harness indexer reads the active immutable physical schema and revision-specific annotations,
constructs schema records, embeds only changed content, and writes them through the configured
vector adapter
([`harness/tht/cli/vector_cmd.py:100-140`](../../harness/tht/cli/vector_cmd.py#L100-L140)). Its
(`harness/tht/cli/vector_cmd.py`). Its
machine result carries the physical/annotation artifact digests, workspace revision, collection,
and counts
([`harness/tht/cli/vector_cmd.py:157-179`](../../harness/tht/cli/vector_cmd.py#L157-L179)). This is
(`harness/tht/cli/vector_cmd.py`). This is
the right place to retain publication evidence and audit linkage.
Preprocessing state is already revision- and binding-aware. A resumed job must match operation,
workspace revision, descriptor/catalog blobs, runtime config and binding digests or it fails with
`preprocessing_resume_mismatch`
([`backend/src/workspaces/preprocessing-state.ts:262-313`](../../backend/src/workspaces/preprocessing-state.ts#L262-L313)).
(`backend/src/workspaces/preprocessing-state.ts`).
There is an important operational gate: every preprocessing operation calls
`assertSessionInventoryCompatible`; a non-finalized, non-archived session pinned to another
revision blocks preprocessing
([`backend/src/workspaces/preprocessing-service.ts:459-472`](../../backend/src/workspaces/preprocessing-service.ts#L459-L472),
[`backend/src/workspaces/preprocessing-state.ts:363-378`](../../backend/src/workspaces/preprocessing-state.ts#L363-L378)).
(`backend/src/workspaces/preprocessing-service.ts`,
`backend/src/workspaces/preprocessing-state.ts`).
A catalog publication UX must expose this as a pending/blocking condition rather than report a
generic indexing failure.
@@ -135,23 +167,22 @@ generic indexing failure.
The collection contract is fixed at the descriptor's dimensions/distance and eight keyword payload
indexes: `content_hash`, `document_id`, `kind`, `record_key`, `record_kind`,
`vector_generation`, `workspace_id`, and `workspace_revision`
([`backend/src/workspaces/qdrant-collection.ts:1-20`](../../backend/src/workspaces/qdrant-collection.ts#L1-L20)).
(`backend/src/workspaces/qdrant-collection.ts`).
`self_heal` may create a missing compatible collection or indexes; `require_existing` only validates
and refuses an incompatible collection
([`backend/src/workspaces/qdrant-collection.ts:71-104`](../../backend/src/workspaces/qdrant-collection.ts#L71-L104)).
(`backend/src/workspaces/qdrant-collection.ts`).
The publication path must use this shared manager instead of inventing collection setup.
Schema payloads include both the workspace and workspace revision, record identity, semantic kind,
content and content hash
([`harness/tht/vectorstore/records.py:30-49`](../../harness/tht/vectorstore/records.py#L30-L49)).
(`harness/tht/vectorstore/records.py`).
Qdrant point IDs for schema and Evidence also include the revision, and upserts use the same
revision in the payload
([`harness/tht/adapters/vector/qdrant.py:42-46`](../../harness/tht/adapters/vector/qdrant.py#L42-L46),
[`harness/tht/adapters/vector/qdrant.py:196-220`](../../harness/tht/adapters/vector/qdrant.py#L196-L220)).
(`harness/tht/adapters/vector/qdrant.py`).
Reads always filter by `workspace_id`; schema and Evidence reads additionally filter by the bound
`workspace_revision`, while memory and solved-question records intentionally remain
workspace-wide
([`harness/tht/adapters/vector/qdrant.py:291-308`](../../harness/tht/adapters/vector/qdrant.py#L291-L308)).
(`harness/tht/adapters/vector/qdrant.py:125-145`).
This means an approved semantic change needs a new workspace revision if old sessions must retain
their previous view. Directly overwriting points under the same Git revision would mutate the
@@ -160,43 +191,52 @@ would require changing the current runtime filter contract.
### Qdrant synchronization defect to resolve before catalog publication
The current incremental synchronizer compares canonical records by content hash and upserts
The current incremental schema synchronizer compares canonical records by content hash and upserts
changed records, but it has no deletion step
([`harness/tht/cli/vector_cmd.py:67-89`](../../harness/tht/cli/vector_cmd.py#L67-L89)). More
(`harness/tht/cli/vector_cmd.py:77-99`). More
importantly, `QdrantVectorStore.existing_hashes()` filters by workspace and record kind but does not
apply `_revision_filter()`
([`harness/tht/adapters/vector/qdrant.py:174-194`](../../harness/tht/adapters/vector/qdrant.py#L174-L194)),
(`harness/tht/adapters/vector/qdrant.py:266-286`),
even though point IDs and reads are revision-scoped. Therefore a same-key/same-content record from
an older revision may be classified as unchanged and never written under the new revision. This is
an implementation defect/risk inferred directly from the two code paths, and it should be fixed
and regression-tested before the catalog relies on `index-schema` for multi-revision publication.
Deletion/GC must be stated per record family. Evidence cleanup is implemented: the corpus pipeline
retains the configured number of published generations, protects active/job-referenced
generations, and calls exact-generation deletion for evicted or compensated generations
(`harness/tht/evidence/corpus/pipeline.py`,
`harness/tht/adapters/vector/qdrant.py:350-390`). Schema cleanup is
not implemented: `delete_kinds()` exists as a workspace-scoped adapter primitive, but the schema
synchronizer never calls it, it is not revision-scoped, and there is no retention policy for old
schema revisions. Publication therefore still needs exact current-revision deletion semantics and
separate safe GC for unleased historical schema revisions.
## 5. Session pinning and why the SQL workflow can remain unchanged
Normal session creation acquires an immutable registry revision, passes its snapshot path,
workspace ID and commit to `tht session new`, and only releases the retention lease after the
manifest has been persisted
([`backend/src/routes/sessions.ts:341-367`](../../backend/src/routes/sessions.ts#L341-L367),
[`backend/src/routes/sessions.ts:423-440`](../../backend/src/routes/sessions.ts#L423-L440)). The
(`backend/src/routes/sessions.ts`). The
manifest stores `workspace_id` and `workspace_revision` next to database/schema identity
([`harness/tht/session/store.py:101-132`](../../harness/tht/session/store.py#L101-L132)). Resume and
(`harness/tht/session/store.py`). Resume and
saved-SQL paths reopen that exact retained snapshot rather than the current installation default
([`backend/src/routes/sessions.ts:184-198`](../../backend/src/routes/sessions.ts#L184-L198),
[`backend/src/routes/sql.ts:24-43`](../../backend/src/routes/sql.ts#L24-L43)).
(`backend/src/routes/sessions.ts`,
`backend/src/routes/sql.ts`).
The runtime renderer places the same workspace revision in `runtime_identity`, points the harness
at the descriptor-owned Qdrant collection, and supplies the internal embedding service
([`backend/src/workspaces/runtime-renderer.ts:248-277`](../../backend/src/workspaces/runtime-renderer.ts#L248-L277)).
(`backend/src/workspaces/runtime-renderer.ts`).
The vector adapter is constructed directly from that configuration, including the revision
([`harness/tht/adapters/factory.py:27-42`](../../harness/tht/adapters/factory.py#L27-L42)).
(`harness/tht/adapters/factory.py`).
At session bootstrap the backend invokes the existing `search pack` command
([`backend/src/routes/sessions.ts:465-470`](../../backend/src/routes/sessions.ts#L465-L470)). That
(`backend/src/routes/sessions.ts`). That
command queries schema records with the existing schema kinds, ranks tables and persists the
candidate list
([`harness/tht/cli/search_cmd.py:291-310`](../../harness/tht/cli/search_cmd.py#L291-L310)). The Pi
(`harness/tht/cli/search_cmd.py`). The Pi
extension reads the persisted retrieval pack through the CLI
([`harness/.pi/extensions/tht-gate.js:61-75`](../../harness/.pi/extensions/tht-gate.js#L61-L75)),
(`harness/.pi/extensions/tht-gate.js`),
and F4 starts from those candidates while loading full table/column context through existing schema
commands. Thus catalog publication can improve the inputs without changing phases, gate semantics,
or persisted session artifacts.
@@ -204,17 +244,18 @@ or persisted session artifacts.
## 6. DWH reuse across metadata-only revisions
DWH preparation uses immutable generations selected by an `ACTIVE` pointer
([`docs/contracts/tht-dwh.md:18-24`](../contracts/tht-dwh.md#L18-L24)). The effective-configuration
([DWH contract](../contracts/tht-dwh.md)). The effective-configuration
identity deliberately excludes `runtime_identity`, so a content-only Git revision does not force a
database re-introspection
([`docs/contracts/tht-dwh.md:74-90`](../contracts/tht-dwh.md#L74-L90)). This is the key enabling
property for metadata publication: a new revision can carry updated curated annotations/Core Schema
Selection, reuse the compatible physical DWH generation, and rebuild only the revision-scoped
schema projection.
([DWH contract](../contracts/tht-dwh.md)). This is the key enabling property for future metadata
publication: a new Git revision can carry updated curated annotations and a future Core Schema
Selection contract, reuse the compatible DWH-owned `physical.yaml`, and rebuild only the
revision-scoped schema projection. Today only the curated annotations part of that statement exists.
## 7. Constraints for the Metadata Catalog design
The following should be treated as requirements for the architecture map:
The following remain proposed requirements for the deferred catalog-to-core design gate; they are
not claims about current implementation:
1. **Separate full inventory from core projection.** PostgreSQL may hold the complete database
catalog, drafts and AI-generated text. Only an explicit Core Schema Selection and approved
@@ -225,10 +266,10 @@ The following should be treated as requirements for the architecture map:
3. **Use a new Git revision as the publication identity.** This preserves current snapshot,
retention, session resume and Qdrant filtering semantics. A separate mutable catalog revision
cannot be safely introduced without changing runtime reads.
4. **Reuse canonical formats.** Project physical facts/Core Schema Selection into a validated
`PhysicalSchema` view and semantic edits into `Annotations`; keep Mermaid and long-form database
documentation outside Qdrant unless a separate, explicit record kind and retrieval policy is
designed.
4. **Respect artifact ownership.** Keep introspected physical facts in the immutable DWH generation;
project the future Core Schema Selection through an explicit contract and approved semantic
edits into `Annotations`. Keep Mermaid and long-form database documentation outside Qdrant
unless a separate record kind and retrieval policy is designed.
5. **Reuse the operator boundary.** Trigger `workspace index-schema`, observe its schema-versioned
result and persist its artifact identities. Do not call Qdrant from browser CRUD handlers.
6. **Keep runtime read-only.** The session/Pi process continues to read the pinned snapshot,
@@ -236,14 +277,15 @@ The following should be treated as requirements for the architecture map:
management control plane.
7. **Surface publication gates.** UI status must distinguish Git activation, incompatible/missing
collection, resumable-session revision conflict, embedding failure and completed publication.
8. **Repair and test cross-revision synchronization first.** Scope `existing_hashes()` to the bound
revision and define deletion/garbage-collection semantics before depending on repeated catalog
publication.
8. **Repair and test cross-revision schema synchronization first.** Scope `existing_hashes()` to
the bound revision, delete records removed from the current revision's canonical schema, and
define safe schema-revision GC. Reuse rather than duplicate the already implemented Evidence
generation GC.
9. **Close the annotation display gap without changing the workflow.** Make `schema columns` read
the same merged descriptions used by M-Schema/vector rendering, so the current F4 widget sees
the approved catalog text.
## Decision summary for the Wayfinder map
## Proposed decision summary for the Wayfinder map
- Keep the Metadata Catalog as a separate management subsystem and source of editable metadata.
- Keep Qdrant derived and revision-scoped; it is not the catalog database or source of truth.
@@ -253,3 +295,7 @@ The following should be treated as requirements for the architecture map:
the current descriptor.
- Treat direct same-revision Qdrant writes, implicit CRUD publication, and bypassing the Git
snapshot boundary as rejected integration paths.
This summary remains design input. The authoritative current state is that catalog-to-core
integration and Sensitive Data Policy enforcement in schema-linking are deferred in
`PROJECT_STATE.md:160-171`.