302 lines
19 KiB
Markdown
302 lines
19 KiB
Markdown
# ThothII metadata publication and Qdrant revision seams
|
|
|
|
**Research question:** Which existing workspace, snapshot, preprocessing, Qdrant,
|
|
session-pinning, and runtime read-only contracts constrain metadata publication without changing
|
|
the NL→SQL workflow?
|
|
|
|
**Last verified:** 2026-08-31
|
|
|
|
**Validity:** Active architectural research; the catalog-to-core integration described below is
|
|
still **deferred**, not implemented.
|
|
|
|
**Current authorities:** `PROJECT_STATE.md`,
|
|
[`workspace-evidence-v3.md`](../contracts/workspace-evidence-v3.md),
|
|
[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md),
|
|
[`tht-dwh.md`](../contracts/tht-dwh.md),
|
|
[`ADR-0001`](../adr/0001-postgres-metadata-catalog.md), and
|
|
[`ADR-0004`](../adr/0004-fastify-kysely-metadata-catalog.md). These sources and current code
|
|
override this research note if they diverge.
|
|
|
|
## Revalidation against the current implementation
|
|
|
|
| Finding | Status on 2026-08-31 | Current evidence and consequence |
|
|
| --- | --- | --- |
|
|
| PostgreSQL metadata authority in the existing Fastify backend | **Adopted/current** | ADR-0001 and ADR-0004 are implemented; the catalog is the management-plane authority. |
|
|
| Git workspace revision, immutable registry snapshot, and revision-pinned runtime | **Adopted/current** | Registry activation and the Workspace Evidence v3 contract still provide the publication boundary for runtime-owned files. |
|
|
| Qdrant read isolation for schema and Evidence by `workspace_id` + `workspace_revision` | **Adopted/current** | `QdrantVectorStore.search()` applies `_revision_filter()` to schema/Evidence; Memory and solved questions intentionally remain workspace-wide (`harness/tht/adapters/vector/qdrant.py`). |
|
|
| Physical schema as a file in the Git publication | **Superseded clarification** | `physical.yaml` is owned by the immutable `.tht-dwh` generation selected by `ACTIVE`, not by the Git workspace revision ([DWH contract](../contracts/tht-dwh.md)). Git currently supplies the revision-pinned curated `schema/annotations.yaml`. |
|
|
| Catalog-to-core publisher | **Open/deferred** | No current route or service projects catalog records into core artifacts. `PROJECT_STATE.md` explicitly defers the schema-linking integration (`PROJECT_STATE.md:160-171`). |
|
|
| Explicit Core Schema Selection | **Open/deferred** | The workspace descriptor selects one database and one physical schema, but has no table/column allowlist (`backend/src/workspaces/schema.ts`); no catalog selection is handed to `tht`. |
|
|
| Cross-revision hash lookup used by schema synchronization | **Open defect** | Reads are revision-filtered, but `existing_hashes()` is not. A same-key/same-content point from an older revision can suppress the required upsert into the new revision (`harness/tht/adapters/vector/qdrant.py`, `harness/tht/cli/vector_cmd.py`). |
|
|
| Deletion/GC | **Current for Evidence; open for schema** | Evidence has generation inventory, retention, compensation, and exact-generation deletion. `sync_canonical_records()` never deletes schema records absent from the new canonical set, and no revision-retention GC exists for schema points. |
|
|
| Annotation consumption | **Current but incomplete** | M-Schema rendering and schema embeddings consume `Annotations`; `tht schema columns` still returns only physical comments, so F4 does not display annotation descriptions (`harness/tht/cli/schema_cmd.py`, `harness/.pi/extensions/tht-gate.js`). |
|
|
|
|
## Conclusion
|
|
|
|
The lowest-impact **proposed** publication seam is not a new writer inside the NL→SQL workflow and
|
|
is not a direct CRUD-to-Qdrant path. The existing core consumes an introspected `PhysicalSchema`
|
|
from the active immutable DWH generation and a curated, Git-revision-pinned `Annotations`
|
|
document. A future explicit publication operation should project the approved catalog subset into
|
|
a new Core Schema Selection contract and the curated annotations, activate a new immutable Git
|
|
revision, and then invoke the existing `workspace index-schema` preprocessing operation. Qdrant
|
|
remains a derived, rebuildable projection. None of this catalog-to-core handoff exists yet;
|
|
`PROJECT_STATE.md:160-164` deliberately defers it to the next design
|
|
gate.
|
|
|
|
The proposed flow would preserve the existing runtime path:
|
|
|
|
```text
|
|
approved catalog data
|
|
-> explicit Core Schema Selection + curated annotations projection (future)
|
|
-> workspace Git revision (annotations and selection contract; physical.yaml stays DWH-owned)
|
|
-> immutable registry snapshot
|
|
-> revision-bound runtime configuration
|
|
-> existing tht vector index-schema
|
|
-> Qdrant records filtered by workspace_id + workspace_revision
|
|
-> existing retrieval_pack / schema render / F4 review
|
|
```
|
|
|
|
The complete database inventory remains in the Metadata Catalog. Only an explicit Core Schema
|
|
Selection and approved semantic fields should be projected into the artifacts used by the SQL
|
|
workflow. ThothII does not currently model or publish that table/column-level selection, so both
|
|
the projection contract and its operational publisher are new work.
|
|
|
|
## 1. Current authority and publication boundary
|
|
|
|
The workspace descriptor identifies one PostgreSQL database and one physical schema, plus one
|
|
workspace-owned Qdrant collection; it has no table or column allowlist
|
|
(`backend/src/workspaces/schema.ts`).
|
|
Consequently, a full-database metadata catalog and the subset eligible for the SQL core cannot be
|
|
represented as the same current descriptor object.
|
|
|
|
The existing public contract makes the Git workspace repository curator-owned. Changes occur in a
|
|
separate authoring clone followed by installation pull, and the API does not write workspace,
|
|
schema, or Evidence paths
|
|
([Workspace Evidence v3 contract](../contracts/workspace-evidence-v3.md)).
|
|
Therefore a browser CRUD service cannot silently make its PostgreSQL state authoritative for the
|
|
core without either:
|
|
|
|
1. exporting/committing a deterministic workspace projection through the existing curator flow;
|
|
or
|
|
2. deliberately replacing this Git-authority contract.
|
|
|
|
The first option preserves current architecture and session reproducibility.
|
|
|
|
Registry activation validates every descriptor at one exact commit, validates and synchronizes
|
|
the commit's `schema/annotations.yaml`, and rejects duplicate ownership of a Qdrant collection
|
|
(`backend/src/workspaces/registry.ts`,
|
|
`backend/src/workspaces/git-repository.ts`,
|
|
`backend/src/workspaces/annotations-sync.ts`).
|
|
It writes the candidate snapshot into a staging directory, records file digests in
|
|
`snapshot.json`, atomically renames the directory, and only then moves active state
|
|
(`backend/src/workspaces/registry.ts`).
|
|
That is the existing atomic publication boundary to reuse.
|
|
|
|
## 2. Canonical schema inputs already consumed by the core
|
|
|
|
The harness separates source facts from curated semantics:
|
|
|
|
- `PhysicalSchema` contains database/schema identity and tables; table facts include comments,
|
|
columns, physical foreign keys, and indexes; columns contain type, nullability, primary-key,
|
|
default, source comment, examples, and eligibility
|
|
(`harness/tht/mschema/models.py`).
|
|
- `Annotations` contains curated table descriptions, concepts and notes, column descriptions,
|
|
synonyms, concepts, evidence, notes and eligibility overrides, plus logical foreign keys
|
|
(`harness/tht/mschema/models.py`).
|
|
|
|
Rendering already gives annotations precedence over source comments and merges physical and
|
|
logical foreign keys. It also applies column eligibility before producing M-Schema context
|
|
(`harness/tht/mschema/render.py`). This makes
|
|
`Annotations` the natural narrow projection target for approved descriptions, synonyms, logical
|
|
relationships, and eligibility from the new catalog.
|
|
|
|
There are two current compatibility gaps:
|
|
|
|
- `schema_records()` embeds table descriptions/concepts and column descriptions/synonyms/examples,
|
|
but does not include physical or logical foreign keys in vector record content
|
|
(`harness/tht/vectorstore/records.py`).
|
|
Relationships still reach the model through deterministic schema rendering, not through schema
|
|
candidate embeddings.
|
|
- The F4 widget loads columns with `tht schema columns`
|
|
(`harness/.pi/extensions/tht-gate.js`), but that
|
|
command still returns only `PhysicalSchema.comment` values and does not merge
|
|
`Annotations`
|
|
(`harness/tht/cli/schema_cmd.py`).
|
|
A publisher that writes only annotations would improve vector search and rendered M-Schema but
|
|
not the table/column descriptions displayed by this existing reviewer widget. Fixing the command
|
|
to use the existing merged description helpers would preserve the workflow shape while closing
|
|
the inconsistency.
|
|
|
|
## 3. Existing preprocessing seam
|
|
|
|
`workspace index-schema` is already the supported operator seam. It creates a revision-bound
|
|
runtime, checks collection compatibility, and runs the harness command
|
|
`vector index-schema --json`
|
|
(`backend/src/workspaces/preprocessing-service.ts`,
|
|
[`workspace-preprocessing-cli.md`](../contracts/workspace-preprocessing-cli.md)).
|
|
The full preprocessing operation performs DWH preparation, FK suggestion/review, schema indexing,
|
|
and optional Evidence preprocessing as separate resumable stages
|
|
(`backend/src/workspaces/preprocessing-service.ts`).
|
|
Metadata-only publication should normally use the narrow `index-schema` operation after its
|
|
workspace artifacts are valid, rather than coupling catalog CRUD to the full pipeline.
|
|
|
|
The harness indexer reads the active immutable physical schema and revision-specific annotations,
|
|
constructs schema records, embeds only changed content, and writes them through the configured
|
|
vector adapter
|
|
(`harness/tht/cli/vector_cmd.py`). Its
|
|
machine result carries the physical/annotation artifact digests, workspace revision, collection,
|
|
and counts
|
|
(`harness/tht/cli/vector_cmd.py`). This is
|
|
the right place to retain publication evidence and audit linkage.
|
|
|
|
Preprocessing state is already revision- and binding-aware. A resumed job must match operation,
|
|
workspace revision, descriptor/catalog blobs, runtime config and binding digests or it fails with
|
|
`preprocessing_resume_mismatch`
|
|
(`backend/src/workspaces/preprocessing-state.ts`).
|
|
|
|
There is an important operational gate: every preprocessing operation calls
|
|
`assertSessionInventoryCompatible`; a non-finalized, non-archived session pinned to another
|
|
revision blocks preprocessing
|
|
(`backend/src/workspaces/preprocessing-service.ts`,
|
|
`backend/src/workspaces/preprocessing-state.ts`).
|
|
A catalog publication UX must expose this as a pending/blocking condition rather than report a
|
|
generic indexing failure.
|
|
|
|
## 4. Qdrant identity, payload and collection contracts
|
|
|
|
The collection contract is fixed at the descriptor's dimensions/distance and eight keyword payload
|
|
indexes: `content_hash`, `document_id`, `kind`, `record_key`, `record_kind`,
|
|
`vector_generation`, `workspace_id`, and `workspace_revision`
|
|
(`backend/src/workspaces/qdrant-collection.ts`).
|
|
`self_heal` may create a missing compatible collection or indexes; `require_existing` only validates
|
|
and refuses an incompatible collection
|
|
(`backend/src/workspaces/qdrant-collection.ts`).
|
|
The publication path must use this shared manager instead of inventing collection setup.
|
|
|
|
Schema payloads include both the workspace and workspace revision, record identity, semantic kind,
|
|
content and content hash
|
|
(`harness/tht/vectorstore/records.py`).
|
|
Qdrant point IDs for schema and Evidence also include the revision, and upserts use the same
|
|
revision in the payload
|
|
(`harness/tht/adapters/vector/qdrant.py`).
|
|
Reads always filter by `workspace_id`; schema and Evidence reads additionally filter by the bound
|
|
`workspace_revision`, while memory and solved-question records intentionally remain
|
|
workspace-wide
|
|
(`harness/tht/adapters/vector/qdrant.py:125-145`).
|
|
|
|
This means an approved semantic change needs a new workspace revision if old sessions must retain
|
|
their previous view. Directly overwriting points under the same Git revision would mutate the
|
|
meaning of that supposedly immutable revision; adding an independent catalog-publication version
|
|
would require changing the current runtime filter contract.
|
|
|
|
### Qdrant synchronization defect to resolve before catalog publication
|
|
|
|
The current incremental schema synchronizer compares canonical records by content hash and upserts
|
|
changed records, but it has no deletion step
|
|
(`harness/tht/cli/vector_cmd.py:77-99`). More
|
|
importantly, `QdrantVectorStore.existing_hashes()` filters by workspace and record kind but does not
|
|
apply `_revision_filter()`
|
|
(`harness/tht/adapters/vector/qdrant.py:266-286`),
|
|
even though point IDs and reads are revision-scoped. Therefore a same-key/same-content record from
|
|
an older revision may be classified as unchanged and never written under the new revision. This is
|
|
an implementation defect/risk inferred directly from the two code paths, and it should be fixed
|
|
and regression-tested before the catalog relies on `index-schema` for multi-revision publication.
|
|
|
|
Deletion/GC must be stated per record family. Evidence cleanup is implemented: the corpus pipeline
|
|
retains the configured number of published generations, protects active/job-referenced
|
|
generations, and calls exact-generation deletion for evicted or compensated generations
|
|
(`harness/tht/evidence/corpus/pipeline.py`,
|
|
`harness/tht/adapters/vector/qdrant.py:350-390`). Schema cleanup is
|
|
not implemented: `delete_kinds()` exists as a workspace-scoped adapter primitive, but the schema
|
|
synchronizer never calls it, it is not revision-scoped, and there is no retention policy for old
|
|
schema revisions. Publication therefore still needs exact current-revision deletion semantics and
|
|
separate safe GC for unleased historical schema revisions.
|
|
|
|
## 5. Session pinning and why the SQL workflow can remain unchanged
|
|
|
|
Normal session creation acquires an immutable registry revision, passes its snapshot path,
|
|
workspace ID and commit to `tht session new`, and only releases the retention lease after the
|
|
manifest has been persisted
|
|
(`backend/src/routes/sessions.ts`). The
|
|
manifest stores `workspace_id` and `workspace_revision` next to database/schema identity
|
|
(`harness/tht/session/store.py`). Resume and
|
|
saved-SQL paths reopen that exact retained snapshot rather than the current installation default
|
|
(`backend/src/routes/sessions.ts`,
|
|
`backend/src/routes/sql.ts`).
|
|
|
|
The runtime renderer places the same workspace revision in `runtime_identity`, points the harness
|
|
at the descriptor-owned Qdrant collection, and supplies the internal embedding service
|
|
(`backend/src/workspaces/runtime-renderer.ts`).
|
|
The vector adapter is constructed directly from that configuration, including the revision
|
|
(`harness/tht/adapters/factory.py`).
|
|
|
|
At session bootstrap the backend invokes the existing `search pack` command
|
|
(`backend/src/routes/sessions.ts`). That
|
|
command queries schema records with the existing schema kinds, ranks tables and persists the
|
|
candidate list
|
|
(`harness/tht/cli/search_cmd.py`). The Pi
|
|
extension reads the persisted retrieval pack through the CLI
|
|
(`harness/.pi/extensions/tht-gate.js`),
|
|
and F4 starts from those candidates while loading full table/column context through existing schema
|
|
commands. Thus catalog publication can improve the inputs without changing phases, gate semantics,
|
|
or persisted session artifacts.
|
|
|
|
## 6. DWH reuse across metadata-only revisions
|
|
|
|
DWH preparation uses immutable generations selected by an `ACTIVE` pointer
|
|
([DWH contract](../contracts/tht-dwh.md)). The effective-configuration
|
|
identity deliberately excludes `runtime_identity`, so a content-only Git revision does not force a
|
|
database re-introspection
|
|
([DWH contract](../contracts/tht-dwh.md)). This is the key enabling property for future metadata
|
|
publication: a new Git revision can carry updated curated annotations and a future Core Schema
|
|
Selection contract, reuse the compatible DWH-owned `physical.yaml`, and rebuild only the
|
|
revision-scoped schema projection. Today only the curated annotations part of that statement exists.
|
|
|
|
## 7. Constraints for the Metadata Catalog design
|
|
|
|
The following remain proposed requirements for the deferred catalog-to-core design gate; they are
|
|
not claims about current implementation:
|
|
|
|
1. **Separate full inventory from core projection.** PostgreSQL may hold the complete database
|
|
catalog, drafts and AI-generated text. Only an explicit Core Schema Selection and approved
|
|
semantic fields are exported to core artifacts/Qdrant.
|
|
2. **Publish; do not live-link.** CRUD changes are not visible to the SQL workflow until an explicit,
|
|
audited publication succeeds. A failed projection, Git activation or Qdrant index operation
|
|
leaves the previous revision active.
|
|
3. **Use a new Git revision as the publication identity.** This preserves current snapshot,
|
|
retention, session resume and Qdrant filtering semantics. A separate mutable catalog revision
|
|
cannot be safely introduced without changing runtime reads.
|
|
4. **Respect artifact ownership.** Keep introspected physical facts in the immutable DWH generation;
|
|
project the future Core Schema Selection through an explicit contract and approved semantic
|
|
edits into `Annotations`. Keep Mermaid and long-form database documentation outside Qdrant
|
|
unless a separate record kind and retrieval policy is designed.
|
|
5. **Reuse the operator boundary.** Trigger `workspace index-schema`, observe its schema-versioned
|
|
result and persist its artifact identities. Do not call Qdrant from browser CRUD handlers.
|
|
6. **Keep runtime read-only.** The session/Pi process continues to read the pinned snapshot,
|
|
retrieval pack and Qdrant projection. AI generation and catalog writes belong to the separate
|
|
management control plane.
|
|
7. **Surface publication gates.** UI status must distinguish Git activation, incompatible/missing
|
|
collection, resumable-session revision conflict, embedding failure and completed publication.
|
|
8. **Repair and test cross-revision schema synchronization first.** Scope `existing_hashes()` to
|
|
the bound revision, delete records removed from the current revision's canonical schema, and
|
|
define safe schema-revision GC. Reuse rather than duplicate the already implemented Evidence
|
|
generation GC.
|
|
9. **Close the annotation display gap without changing the workflow.** Make `schema columns` read
|
|
the same merged descriptions used by M-Schema/vector rendering, so the current F4 widget sees
|
|
the approved catalog text.
|
|
|
|
## Proposed decision summary for the Wayfinder map
|
|
|
|
- Keep the Metadata Catalog as a separate management subsystem and source of editable metadata.
|
|
- Keep Qdrant derived and revision-scoped; it is not the catalog database or source of truth.
|
|
- Publish through immutable workspace revisions plus the existing preprocessing service.
|
|
- Preserve the current NL→SQL workflow, Pi extension, phase model and session artifacts.
|
|
- Add a new, explicit Core Schema Selection contract because no table/column selection exists in
|
|
the current descriptor.
|
|
- Treat direct same-revision Qdrant writes, implicit CRUD publication, and bypassing the Git
|
|
snapshot boundary as rejected integration paths.
|
|
|
|
This summary remains design input. The authoritative current state is that catalog-to-core
|
|
integration and Sensitive Data Policy enforcement in schema-linking are deferred in
|
|
`PROJECT_STATE.md:160-171`.
|