docs: trace metadata publication and qdrant seams
This commit is contained in:
@@ -0,0 +1,255 @@
|
|||||||
|
# ThothII metadata publication and Qdrant revision seams
|
||||||
|
|
||||||
|
**Research question:** Which existing workspace, snapshot, preprocessing, Qdrant,
|
||||||
|
session-pinning, and runtime read-only contracts constrain metadata publication without changing
|
||||||
|
the NL→SQL workflow?
|
||||||
|
|
||||||
|
## Conclusion
|
||||||
|
|
||||||
|
The lowest-impact publication seam is **not** a new writer inside the NL→SQL workflow and it is
|
||||||
|
not a direct CRUD-to-Qdrant path. The existing core already consumes two canonical schema inputs:
|
||||||
|
an introspected `PhysicalSchema` and a curated `Annotations` document. An explicit publication
|
||||||
|
operation should project the approved, workspace-selected subset of the new metadata catalog into
|
||||||
|
those inputs, bind the projection to a new Git `workspace_revision`, activate that immutable
|
||||||
|
revision, and then invoke the existing `workspace index-schema` preprocessing operation. Qdrant
|
||||||
|
remains a derived, rebuildable projection.
|
||||||
|
|
||||||
|
This preserves the existing runtime path:
|
||||||
|
|
||||||
|
```text
|
||||||
|
approved catalog data
|
||||||
|
-> workspace Git revision (physical schema selection + annotations)
|
||||||
|
-> immutable registry snapshot
|
||||||
|
-> revision-bound runtime configuration
|
||||||
|
-> existing tht vector index-schema
|
||||||
|
-> Qdrant records filtered by workspace_id + workspace_revision
|
||||||
|
-> existing retrieval_pack / schema render / F4 review
|
||||||
|
```
|
||||||
|
|
||||||
|
The complete database inventory may remain in the Metadata Catalog. Only the explicit Core Schema
|
||||||
|
Selection should be projected into the workspace artifacts used by the SQL workflow. ThothII does
|
||||||
|
not currently model such a table-level selection in its descriptor, so that projection contract is
|
||||||
|
new work.
|
||||||
|
|
||||||
|
## 1. Current authority and publication boundary
|
||||||
|
|
||||||
|
The workspace descriptor identifies one PostgreSQL database and one physical schema, plus one
|
||||||
|
workspace-owned Qdrant collection; it has no table or column allowlist
|
||||||
|
([`backend/src/workspaces/schema.ts:160-178`](../../backend/src/workspaces/schema.ts#L160-L178)).
|
||||||
|
Consequently, a full-database metadata catalog and the subset eligible for the SQL core cannot be
|
||||||
|
represented as the same current descriptor object.
|
||||||
|
|
||||||
|
The existing public contract makes the Git workspace repository curator-owned. Changes occur in a
|
||||||
|
separate authoring clone followed by installation pull, and the API does not write workspace,
|
||||||
|
schema, or Evidence paths
|
||||||
|
([`docs/contracts/workspace-evidence-v3.md:148-156`](../contracts/workspace-evidence-v3.md#L148-L156)).
|
||||||
|
Therefore a browser CRUD service cannot silently make its PostgreSQL state authoritative for the
|
||||||
|
core without either:
|
||||||
|
|
||||||
|
1. exporting/committing a deterministic workspace projection through the existing curator flow;
|
||||||
|
or
|
||||||
|
2. deliberately replacing this Git-authority contract.
|
||||||
|
|
||||||
|
The first option preserves current architecture and session reproducibility.
|
||||||
|
|
||||||
|
Registry activation validates every descriptor at one exact commit, validates and synchronizes
|
||||||
|
the commit's `schema/annotations.yaml`, and rejects duplicate ownership of a Qdrant collection
|
||||||
|
([`backend/src/workspaces/registry.ts:531-576`](../../backend/src/workspaces/registry.ts#L531-L576)).
|
||||||
|
It writes the candidate snapshot into a staging directory, records file digests in
|
||||||
|
`snapshot.json`, atomically renames the directory, and only then moves active state
|
||||||
|
([`backend/src/workspaces/registry.ts:582-629`](../../backend/src/workspaces/registry.ts#L582-L629)).
|
||||||
|
That is the existing atomic publication boundary to reuse.
|
||||||
|
|
||||||
|
## 2. Canonical schema inputs already consumed by the core
|
||||||
|
|
||||||
|
The harness separates source facts from curated semantics:
|
||||||
|
|
||||||
|
- `PhysicalSchema` contains database/schema identity and tables; table facts include comments,
|
||||||
|
columns, physical foreign keys, and indexes; columns contain type, nullability, primary-key,
|
||||||
|
default, source comment, examples, and eligibility
|
||||||
|
([`harness/tht/mschema/models.py:23-64`](../../harness/tht/mschema/models.py#L23-L64)).
|
||||||
|
- `Annotations` contains curated table descriptions, concepts and notes, column descriptions,
|
||||||
|
synonyms, concepts, evidence, notes and eligibility overrides, plus logical foreign keys
|
||||||
|
([`harness/tht/mschema/models.py:67-88`](../../harness/tht/mschema/models.py#L67-L88)).
|
||||||
|
|
||||||
|
Rendering already gives annotations precedence over source comments and merges physical and
|
||||||
|
logical foreign keys. It also applies column eligibility before producing M-Schema context
|
||||||
|
([`harness/tht/mschema/render.py:9-43`](../../harness/tht/mschema/render.py#L9-L43),
|
||||||
|
[`harness/tht/mschema/render.py:46-80`](../../harness/tht/mschema/render.py#L46-L80)). This makes
|
||||||
|
`Annotations` the natural narrow projection target for approved descriptions, synonyms, logical
|
||||||
|
relationships, and eligibility from the new catalog.
|
||||||
|
|
||||||
|
There are two current compatibility gaps:
|
||||||
|
|
||||||
|
- `schema_records()` embeds table descriptions/concepts and column descriptions/synonyms/examples,
|
||||||
|
but does not include physical or logical foreign keys in vector record content
|
||||||
|
([`harness/tht/vectorstore/records.py:101-132`](../../harness/tht/vectorstore/records.py#L101-L132)).
|
||||||
|
Relationships still reach the model through deterministic schema rendering, not through schema
|
||||||
|
candidate embeddings.
|
||||||
|
- The F4 widget loads columns with `tht schema columns`
|
||||||
|
([`harness/.pi/extensions/tht-gate.js:1220-1254`](../../harness/.pi/extensions/tht-gate.js#L1220-L1254)),
|
||||||
|
but that command currently returns only `PhysicalSchema.comment` values and does not merge
|
||||||
|
`Annotations`
|
||||||
|
([`harness/tht/cli/schema_cmd.py:592-621`](../../harness/tht/cli/schema_cmd.py#L592-L621)).
|
||||||
|
A publisher that writes only annotations would improve vector search and rendered M-Schema but
|
||||||
|
not the table/column descriptions displayed by this existing reviewer widget. Fixing the command
|
||||||
|
to use the existing merged description helpers would preserve the workflow shape while closing
|
||||||
|
the inconsistency.
|
||||||
|
|
||||||
|
## 3. Existing preprocessing seam
|
||||||
|
|
||||||
|
`workspace index-schema` is already the supported operator seam. It creates a revision-bound
|
||||||
|
runtime, checks collection compatibility, and runs the harness command
|
||||||
|
`vector index-schema --json`
|
||||||
|
([`backend/src/workspaces/preprocessing-service.ts:312-330`](../../backend/src/workspaces/preprocessing-service.ts#L312-L330)).
|
||||||
|
The full preprocessing operation performs DWH preparation, FK suggestion/review, schema indexing,
|
||||||
|
and optional Evidence preprocessing as separate resumable stages
|
||||||
|
([`backend/src/workspaces/preprocessing-service.ts:367-450`](../../backend/src/workspaces/preprocessing-service.ts#L367-L450)).
|
||||||
|
Metadata-only publication should normally use the narrow `index-schema` operation after its
|
||||||
|
workspace artifacts are valid, rather than coupling catalog CRUD to the full pipeline.
|
||||||
|
|
||||||
|
The harness indexer reads the active immutable physical schema and revision-specific annotations,
|
||||||
|
constructs schema records, embeds only changed content, and writes them through the configured
|
||||||
|
vector adapter
|
||||||
|
([`harness/tht/cli/vector_cmd.py:100-140`](../../harness/tht/cli/vector_cmd.py#L100-L140)). Its
|
||||||
|
machine result carries the physical/annotation artifact digests, workspace revision, collection,
|
||||||
|
and counts
|
||||||
|
([`harness/tht/cli/vector_cmd.py:157-179`](../../harness/tht/cli/vector_cmd.py#L157-L179)). This is
|
||||||
|
the right place to retain publication evidence and audit linkage.
|
||||||
|
|
||||||
|
Preprocessing state is already revision- and binding-aware. A resumed job must match operation,
|
||||||
|
workspace revision, descriptor/catalog blobs, runtime config and binding digests or it fails with
|
||||||
|
`preprocessing_resume_mismatch`
|
||||||
|
([`backend/src/workspaces/preprocessing-state.ts:262-313`](../../backend/src/workspaces/preprocessing-state.ts#L262-L313)).
|
||||||
|
|
||||||
|
There is an important operational gate: every preprocessing operation calls
|
||||||
|
`assertSessionInventoryCompatible`; a non-finalized, non-archived session pinned to another
|
||||||
|
revision blocks preprocessing
|
||||||
|
([`backend/src/workspaces/preprocessing-service.ts:459-472`](../../backend/src/workspaces/preprocessing-service.ts#L459-L472),
|
||||||
|
[`backend/src/workspaces/preprocessing-state.ts:363-378`](../../backend/src/workspaces/preprocessing-state.ts#L363-L378)).
|
||||||
|
A catalog publication UX must expose this as a pending/blocking condition rather than report a
|
||||||
|
generic indexing failure.
|
||||||
|
|
||||||
|
## 4. Qdrant identity, payload and collection contracts
|
||||||
|
|
||||||
|
The collection contract is fixed at the descriptor's dimensions/distance and eight keyword payload
|
||||||
|
indexes: `content_hash`, `document_id`, `kind`, `record_key`, `record_kind`,
|
||||||
|
`vector_generation`, `workspace_id`, and `workspace_revision`
|
||||||
|
([`backend/src/workspaces/qdrant-collection.ts:1-20`](../../backend/src/workspaces/qdrant-collection.ts#L1-L20)).
|
||||||
|
`self_heal` may create a missing compatible collection or indexes; `require_existing` only validates
|
||||||
|
and refuses an incompatible collection
|
||||||
|
([`backend/src/workspaces/qdrant-collection.ts:71-104`](../../backend/src/workspaces/qdrant-collection.ts#L71-L104)).
|
||||||
|
The publication path must use this shared manager instead of inventing collection setup.
|
||||||
|
|
||||||
|
Schema payloads include both the workspace and workspace revision, record identity, semantic kind,
|
||||||
|
content and content hash
|
||||||
|
([`harness/tht/vectorstore/records.py:30-49`](../../harness/tht/vectorstore/records.py#L30-L49)).
|
||||||
|
Qdrant point IDs for schema and Evidence also include the revision, and upserts use the same
|
||||||
|
revision in the payload
|
||||||
|
([`harness/tht/adapters/vector/qdrant.py:42-46`](../../harness/tht/adapters/vector/qdrant.py#L42-L46),
|
||||||
|
[`harness/tht/adapters/vector/qdrant.py:196-220`](../../harness/tht/adapters/vector/qdrant.py#L196-L220)).
|
||||||
|
Reads always filter by `workspace_id`; schema and Evidence reads additionally filter by the bound
|
||||||
|
`workspace_revision`, while memory and solved-question records intentionally remain
|
||||||
|
workspace-wide
|
||||||
|
([`harness/tht/adapters/vector/qdrant.py:291-308`](../../harness/tht/adapters/vector/qdrant.py#L291-L308)).
|
||||||
|
|
||||||
|
This means an approved semantic change needs a new workspace revision if old sessions must retain
|
||||||
|
their previous view. Directly overwriting points under the same Git revision would mutate the
|
||||||
|
meaning of that supposedly immutable revision; adding an independent catalog-publication version
|
||||||
|
would require changing the current runtime filter contract.
|
||||||
|
|
||||||
|
### Qdrant synchronization defect to resolve before catalog publication
|
||||||
|
|
||||||
|
The current incremental synchronizer compares canonical records by content hash and upserts
|
||||||
|
changed records, but it has no deletion step
|
||||||
|
([`harness/tht/cli/vector_cmd.py:67-89`](../../harness/tht/cli/vector_cmd.py#L67-L89)). More
|
||||||
|
importantly, `QdrantVectorStore.existing_hashes()` filters by workspace and record kind but does not
|
||||||
|
apply `_revision_filter()`
|
||||||
|
([`harness/tht/adapters/vector/qdrant.py:174-194`](../../harness/tht/adapters/vector/qdrant.py#L174-L194)),
|
||||||
|
even though point IDs and reads are revision-scoped. Therefore a same-key/same-content record from
|
||||||
|
an older revision may be classified as unchanged and never written under the new revision. This is
|
||||||
|
an implementation defect/risk inferred directly from the two code paths, and it should be fixed
|
||||||
|
and regression-tested before the catalog relies on `index-schema` for multi-revision publication.
|
||||||
|
|
||||||
|
## 5. Session pinning and why the SQL workflow can remain unchanged
|
||||||
|
|
||||||
|
Normal session creation acquires an immutable registry revision, passes its snapshot path,
|
||||||
|
workspace ID and commit to `tht session new`, and only releases the retention lease after the
|
||||||
|
manifest has been persisted
|
||||||
|
([`backend/src/routes/sessions.ts:341-367`](../../backend/src/routes/sessions.ts#L341-L367),
|
||||||
|
[`backend/src/routes/sessions.ts:423-440`](../../backend/src/routes/sessions.ts#L423-L440)). The
|
||||||
|
manifest stores `workspace_id` and `workspace_revision` next to database/schema identity
|
||||||
|
([`harness/tht/session/store.py:101-132`](../../harness/tht/session/store.py#L101-L132)). Resume and
|
||||||
|
saved-SQL paths reopen that exact retained snapshot rather than the current installation default
|
||||||
|
([`backend/src/routes/sessions.ts:184-198`](../../backend/src/routes/sessions.ts#L184-L198),
|
||||||
|
[`backend/src/routes/sql.ts:24-43`](../../backend/src/routes/sql.ts#L24-L43)).
|
||||||
|
|
||||||
|
The runtime renderer places the same workspace revision in `runtime_identity`, points the harness
|
||||||
|
at the descriptor-owned Qdrant collection, and supplies the internal embedding service
|
||||||
|
([`backend/src/workspaces/runtime-renderer.ts:248-277`](../../backend/src/workspaces/runtime-renderer.ts#L248-L277)).
|
||||||
|
The vector adapter is constructed directly from that configuration, including the revision
|
||||||
|
([`harness/tht/adapters/factory.py:27-42`](../../harness/tht/adapters/factory.py#L27-L42)).
|
||||||
|
|
||||||
|
At session bootstrap the backend invokes the existing `search pack` command
|
||||||
|
([`backend/src/routes/sessions.ts:465-470`](../../backend/src/routes/sessions.ts#L465-L470)). That
|
||||||
|
command queries schema records with the existing schema kinds, ranks tables and persists the
|
||||||
|
candidate list
|
||||||
|
([`harness/tht/cli/search_cmd.py:291-310`](../../harness/tht/cli/search_cmd.py#L291-L310)). The Pi
|
||||||
|
extension reads the persisted retrieval pack through the CLI
|
||||||
|
([`harness/.pi/extensions/tht-gate.js:61-75`](../../harness/.pi/extensions/tht-gate.js#L61-L75)),
|
||||||
|
and F4 starts from those candidates while loading full table/column context through existing schema
|
||||||
|
commands. Thus catalog publication can improve the inputs without changing phases, gate semantics,
|
||||||
|
or persisted session artifacts.
|
||||||
|
|
||||||
|
## 6. DWH reuse across metadata-only revisions
|
||||||
|
|
||||||
|
DWH preparation uses immutable generations selected by an `ACTIVE` pointer
|
||||||
|
([`docs/contracts/tht-dwh.md:18-24`](../contracts/tht-dwh.md#L18-L24)). The effective-configuration
|
||||||
|
identity deliberately excludes `runtime_identity`, so a content-only Git revision does not force a
|
||||||
|
database re-introspection
|
||||||
|
([`docs/contracts/tht-dwh.md:74-90`](../contracts/tht-dwh.md#L74-L90)). This is the key enabling
|
||||||
|
property for metadata publication: a new revision can carry updated curated annotations/Core Schema
|
||||||
|
Selection, reuse the compatible physical DWH generation, and rebuild only the revision-scoped
|
||||||
|
schema projection.
|
||||||
|
|
||||||
|
## 7. Constraints for the Metadata Catalog design
|
||||||
|
|
||||||
|
The following should be treated as requirements for the architecture map:
|
||||||
|
|
||||||
|
1. **Separate full inventory from core projection.** PostgreSQL may hold the complete database
|
||||||
|
catalog, drafts and AI-generated text. Only an explicit Core Schema Selection and approved
|
||||||
|
semantic fields are exported to core artifacts/Qdrant.
|
||||||
|
2. **Publish; do not live-link.** CRUD changes are not visible to the SQL workflow until an explicit,
|
||||||
|
audited publication succeeds. A failed projection, Git activation or Qdrant index operation
|
||||||
|
leaves the previous revision active.
|
||||||
|
3. **Use a new Git revision as the publication identity.** This preserves current snapshot,
|
||||||
|
retention, session resume and Qdrant filtering semantics. A separate mutable catalog revision
|
||||||
|
cannot be safely introduced without changing runtime reads.
|
||||||
|
4. **Reuse canonical formats.** Project physical facts/Core Schema Selection into a validated
|
||||||
|
`PhysicalSchema` view and semantic edits into `Annotations`; keep Mermaid and long-form database
|
||||||
|
documentation outside Qdrant unless a separate, explicit record kind and retrieval policy is
|
||||||
|
designed.
|
||||||
|
5. **Reuse the operator boundary.** Trigger `workspace index-schema`, observe its schema-versioned
|
||||||
|
result and persist its artifact identities. Do not call Qdrant from browser CRUD handlers.
|
||||||
|
6. **Keep runtime read-only.** The session/Pi process continues to read the pinned snapshot,
|
||||||
|
retrieval pack and Qdrant projection. AI generation and catalog writes belong to the separate
|
||||||
|
management control plane.
|
||||||
|
7. **Surface publication gates.** UI status must distinguish Git activation, incompatible/missing
|
||||||
|
collection, resumable-session revision conflict, embedding failure and completed publication.
|
||||||
|
8. **Repair and test cross-revision synchronization first.** Scope `existing_hashes()` to the bound
|
||||||
|
revision and define deletion/garbage-collection semantics before depending on repeated catalog
|
||||||
|
publication.
|
||||||
|
9. **Close the annotation display gap without changing the workflow.** Make `schema columns` read
|
||||||
|
the same merged descriptions used by M-Schema/vector rendering, so the current F4 widget sees
|
||||||
|
the approved catalog text.
|
||||||
|
|
||||||
|
## Decision summary for the Wayfinder map
|
||||||
|
|
||||||
|
- Keep the Metadata Catalog as a separate management subsystem and source of editable metadata.
|
||||||
|
- Keep Qdrant derived and revision-scoped; it is not the catalog database or source of truth.
|
||||||
|
- Publish through immutable workspace revisions plus the existing preprocessing service.
|
||||||
|
- Preserve the current NL→SQL workflow, Pi extension, phase model and session artifacts.
|
||||||
|
- Add a new, explicit Core Schema Selection contract because no table/column selection exists in
|
||||||
|
the current descriptor.
|
||||||
|
- Treat direct same-revision Qdrant writes, implicit CRUD publication, and bypassing the Git
|
||||||
|
snapshot boundary as rejected integration paths.
|
||||||
Reference in New Issue
Block a user