test(evidence): harden owner gate acceptance

This commit is contained in:
2026-08-25 02:34:10 +02:00
parent fc83d29b58
commit 8542124f27
5 changed files with 375 additions and 24 deletions
+9 -5
View File
@@ -6,11 +6,15 @@
isolated fake-restructurer probe and the selected Evidence/L0 contract suite: **222 passed, 1
known pytest deprecation warning**. The probe used only a `mktemp` workspace, confirmed the
ThothII worktree status was unchanged, and did not read or write an external PSD authoring path.
- **Scope boundary:** the realistic malformed source fixture, hermetic runner, and
`docs/testing/evidence-restructuring-manual.md` are ThothII-only artifacts. The real Pi call,
PSD source inventory, migration, Git review, vector before/after counts, and human walkthrough
remain **PENDING in issue #47** until the owner authorizes the migration window and exact target
branch. No PSD migration is represented by this entry.
- **Read-only owner-gate package:** authorized inspection of the clean PSD repository at
`47516f85b4db4a67cfa8a86cea4cb2e7b98c5813` recorded the exact 36-source inventory and the
legacy `psd-clinical` Qdrant baseline (163 schema tables, 2,275 schema columns, 2 memory, 1
solved question) in `docs/testing/evidence-restructuring-manual.md`. The before/after Git
status remained clean; no secret was read and no external repository mutation occurred.
- **Scope boundary:** the real Pi call, PSD migration, Git review, authoritative pre/post vector
inspection, activation, and human walkthrough remain **PENDING in issue #47** until the owner
authorizes the migration window and exact target branch. No PSD migration is represented by
this entry.
## Modular workflow refactor candidate — live (2026-08-24)
+86 -12
View File
@@ -2,8 +2,9 @@
Status: **PENDING OWNER AUTHORIZATION**. This guide records the manual work that must
occur only after the owner authorizes a migration window and exact PSD target branch.
The automated runner is hermetic: it uses a fake restructurer and temporary inputs; it
does not inspect, write, stage, or migrate the PSD authoring repository.
The automated runner is hermetic: it uses a fake restructurer and temporary inputs.
The owner-gate inventory below is a separately authorized, read-only PSD snapshot;
it did not write, stage, branch, commit, migrate, activate, push, or read secrets.
## Recorded automated boundary
@@ -26,20 +27,93 @@ per changed source after the authorization gate and examine every proposed curat
## Owner-gate package (issue #46)
Before any PSD write, provide all of the following to the owner:
Read-only snapshot collected 2026-08-25:
- clean ThothII commit and the exact local-gate/acceptance output;
- PSD repository: `/Users/mp/projects/tht-workspace-psd`, clean before and after the
inspection; immutable pre-migration commit
`47516f85b4db4a67cfa8a86cea4cb2e7b98c5813`.
- proposed PSD branch name: `codex/evidence-restructuring-psd` (proposal only; no
external branch has been created);
- the exact, authorized-snapshot list of 36 source files to move;
- the pre-migration PSD commit and the rollback command
`git -C <authorized-psd-clone> reset --hard <pre-migration-commit>`;
- before counts and representative IDs for `schema_table`, `schema_column`, `memory`,
and `solved_question`, plus dense-search samples for Schema and Memory.
- rollback command for the owner to use only if the later authorized migration must be
undone:
The 36-path inventory, PSD commit, and before counts are intentionally blank until the
owner authorizes the external repository inspection. Recording or executing them is
**issue #47**, not this issue.
```bash
git -C /Users/mp/projects/tht-workspace-psd reset --hard 47516f85b4db4a67cfa8a86cea4cb2e7b98c5813
```
- exact source inventory (36 files; `evidence/README.md` is not a source):
```text
psd-clinical/evidence/00-glossario/coorti-universi-pazienti.md
psd-clinical/evidence/00-glossario/glossario-termini-analitici.md
psd-clinical/evidence/00-glossario/glossario-termini-clinici.md
psd-clinical/evidence/00-glossario/glossario-termini-dwh.md
psd-clinical/evidence/00-glossario/note-di-lettura.md
psd-clinical/evidence/00-glossario/tassonomia-eventi.md
psd-clinical/evidence/10-domini-clinici/ablazione.md
psd-clinical/evidence/10-domini-clinici/anagrafica-paziente.md
psd-clinical/evidence/10-domini-clinici/cardioversione-elettrica.md
psd-clinical/evidence/10-domini-clinici/chiusura-auricola.md
psd-clinical/evidence/10-domini-clinici/documenti-clinici.md
psd-clinical/evidence/10-domini-clinici/genetica-clinica.md
psd-clinical/evidence/10-domini-clinici/icd.md
psd-clinical/evidence/10-domini-clinici/ilr.md
psd-clinical/evidence/10-domini-clinici/pacemaker.md
psd-clinical/evidence/10-domini-clinici/pm-icd-altro.md
psd-clinical/evidence/10-domini-clinici/visite-cardiologiche-genetiche.md
psd-clinical/evidence/20-valori-enum/enum-flag-booleani.md
psd-clinical/evidence/20-valori-enum/enum-flag-note-sn.md
psd-clinical/evidence/20-valori-enum/enum-innesto.md
psd-clinical/evidence/20-valori-enum/enum-isteresi.md
psd-clinical/evidence/20-valori-enum/enum-tipo-intervento.md
psd-clinical/evidence/30-esempi-nlq/nlq-ablazione.md
psd-clinical/evidence/30-esempi-nlq/nlq-cardioversione.md
psd-clinical/evidence/30-esempi-nlq/nlq-device-pacemaker-icd.md
psd-clinical/evidence/30-esempi-nlq/nlq-percorso-paziente.md
psd-clinical/evidence/30-esempi-nlq/nlq-studio-elettrofisiologico.md
psd-clinical/evidence/40-mapping-semantico/catena-staging-integration-dwh.md
psd-clinical/evidence/40-mapping-semantico/matrice-viewpoint-dominio.md
psd-clinical/evidence/40-mapping-semantico/registry-testo-clinico-fact-clinical-event-text.md
psd-clinical/evidence/40-mapping-semantico/regole-classificazione-fact-dim-bridge.md
psd-clinical/evidence/40-mapping-semantico/trasformazioni-valori-mapping-colonne.md
psd-clinical/evidence/50-metadati-normalizzazione/normalizzatori-valori-dwh.md
psd-clinical/evidence/50-metadati-normalizzazione/normalizzazione-codici-paziente-medici.md
psd-clinical/evidence/50-metadati-normalizzazione/normalizzazione-device-cied.md
```
The read-only Qdrant observation for the `psd-clinical` collection on the legacy PSD
bind `127.0.0.1:6333` was 163 `schema_table`, 2,275 `schema_column`, 2 `memory`, and
1 `solved_question` point. Representative IDs are:
| Kind | Sample IDs |
| --- | --- |
| `schema_table` | `01bc2535-24d6-5722-a58b-64a122b90b36`, `024ddf60-c80d-5223-ac06-9247e7de7027`, `086e00c8-b4a9-5063-b82d-e5a67add29cd` |
| `schema_column` | `00126cc1-7564-521a-a084-c2d670263258`, `00365200-2c55-5bb4-86bb-87dd2d1bb529`, `004304bc-b543-5b2d-8b40-f18da8e82af7` |
| `memory` | `8d5cd772-563a-5e22-b764-2ca76cf6efca`, `db74457a-3de8-5b95-9a31-d28a1ecf8141` |
| `solved_question` | `2b6bb7d2-1a35-5f49-bd8d-b0cdcb98459a` |
`tht ... workspace vector inspect --json` was attempted read-only but was blocked by
the local maintenance image's missing production `auth.yaml` / `AUTH_MODE=upstream`.
The counts above therefore come from the Qdrant read-only scroll endpoint and must be
repeated through the successful `vector inspect` command immediately before the future
authorized preprocessing action. No secret was read to bypass that guard.
### Reproducible no-write record
The inspection used only these read operations; `git status --porcelain` was empty both
before and after, and `git diff --quiet` succeeded after:
```bash
git -C /Users/mp/projects/tht-workspace-psd rev-parse HEAD
git -C /Users/mp/projects/tht-workspace-psd status --porcelain
git -C /Users/mp/projects/tht-workspace-psd ls-tree -r --name-only HEAD -- psd-clinical/evidence
node -e 'fetch("http://127.0.0.1:6333/collections/psd-clinical/points/scroll", {method:"POST", headers:{"content-type":"application/json"}, body:JSON.stringify({limit:4096,with_payload:true,with_vector:false})}).then(r => r.json()).then(x => console.log(x.result.points.length))'
git -C /Users/mp/projects/tht-workspace-psd diff --quiet
```
Before any PSD write, provide the owner this inventory, SHA, rollback command, baseline
counts/IDs, clean ThothII commit, and exact local-gate output. Issue #47 still owns all
external mutations and manual acceptance.
## Manual acceptance after authorization (issue #47)
@@ -0,0 +1,261 @@
"""The configured candidate evaluator must gate publication on its exact generation."""
import hashlib
from pathlib import Path
from types import SimpleNamespace
import pytest
from pydantic import ValidationError
from tht.cli import preprocess_cmd
from tht.config import VectorConfig, load_config
from tht.evidence import CuratedEvidence, EvidenceManifest, dump_curated_markdown, dump_manifest
from tht.evidence.authoring import ManifestSource
from tht.evidence.contracts import AcquiredDocument, SourceObject
from tht.evidence.corpus.chunk import ChunkPolicy
from tht.evidence.corpus.pipeline import CorpusPipeline, PipelineError
from tht.evidence.corpus.store import CorpusStore
from tht.ports.vector import VectorCapabilities, VectorHealth
class Source:
def __init__(self, item: SourceObject, content: str) -> None:
self.item = item
self.content = content
def discover(self):
return [self.item]
def acquire(self, item):
assert item == self.item
return AcquiredDocument(source=item, content=self.content.encode(), media_type="text/markdown")
class Embedder:
def embed_documents(self, texts):
return [[0.25] * 1024 for _ in texts]
def embed_query(self, query):
assert query
return [0.25] * 1024
class ObservableStore(CorpusStore):
def __init__(self, root: Path) -> None:
super().__init__(root)
self.events: list[tuple[str, str]] = []
def publish(self, generation: str) -> str:
self.events.append(("publish", generation))
return super().publish(generation)
class ObservableVectors:
capabilities = VectorCapabilities(search=True, existing_hashes=True, upsert=True)
def __init__(self, store: ObservableStore) -> None:
self.store = store
self.records = []
self.searches: list[dict] = []
self.fail_evaluation = False
def health(self):
return VectorHealth(
ok=True,
expected_dimension=1024,
observed_dimensions=(1024,),
dimension_compatible=True,
)
def existing_hashes(self, _collection, _kinds):
return {entry.record.id: entry.content_hash for entry in self.records}
def upsert(self, _collection, records):
self.records.extend(records)
return len(records)
def delete_generation(self, _collection, generation, _workspace_id):
self.records = [
entry for entry in self.records
if entry.record.metadata["vector_generation"] != generation
]
return 0
def list_evidence_generations(self, _collection, _workspace_id):
return sorted({entry.record.metadata["vector_generation"] for entry in self.records})
def search(self, _collections, _embedding, **kwargs):
metadata_filter = kwargs["metadata_filter"]
generation = metadata_filter["vector_generation"]
self.searches.append({
"generation": generation,
"mode": kwargs["retrieval_mode"],
"active": self.store.active_generation(),
"query": kwargs["query_text"],
})
assert self.store.active_generation() != generation
if self.fail_evaluation:
return []
for entry in self.records:
metadata = entry.record.metadata
if metadata["vector_generation"] == generation:
return [SimpleNamespace(
id=entry.record.id,
similarity=1.0,
metadata={
"evidence_id": metadata["evidence_id"],
"evidence_kind": metadata["evidence_kind"],
},
)]
return []
def _workspace(tmp_path: Path):
workspace = tmp_path / "workspace"
source_text = "Pazienti con età inferiore a 18 anni.\n"
source_file = workspace / "evidence" / "source" / "notes.md"
curated_file = workspace / "evidence" / "curated" / "formula" / "fascia-pediatrica.md"
source_file.parent.mkdir(parents=True)
curated_file.parent.mkdir(parents=True)
source_file.write_text(source_text, encoding="utf-8")
digest = "sha256:" + hashlib.sha256(source_text.encode()).hexdigest()
evidence = CuratedEvidence.model_validate({
"schema_version": 1,
"id": "evidence:fascia-pediatrica",
"title": "Fascia pediatrica",
"kind": "formula",
"purposes": ["sql_generation"],
"applies_to": {"columns": ["clinical.patient.birth_date"]},
"language": "it",
"provenance": {
"source_file": "source/notes.md",
"source_sha256": digest,
"supporting_excerpts": [source_text.strip()],
},
"review_items": [],
"payload": {
"concept": "fascia pediatrica",
"columns": ["clinical.patient.birth_date"],
"sql": "CASE WHEN age < 18 THEN 'pediatrica' ELSE 'adulta' END",
},
})
curated = dump_curated_markdown(evidence)
curated_file.write_text(curated, encoding="utf-8")
(workspace / "evidence" / "manifest.yaml").write_text(dump_manifest(EvidenceManifest(
schema_version=1,
pipeline_version="evidence-authoring-v1",
sources={"source/notes.md": ManifestSource(sha256=digest, units=(evidence.id,))},
orphans=(),
)), encoding="utf-8")
(workspace / "evidence" / "evaluation.yaml").write_text(
"""
schema_version: 1
queries:
- id: lexical
query: CASE patient.birth_date
profile: lexical
purpose: sql_generation
expected: [evidence:fascia-pediatrica]
- id: semantic
query: Quali pazienti sono pediatrici?
profile: semantic
purpose: sql_generation
expected: [evidence:fascia-pediatrica]
- id: mixed
query: Formula per patient.birth_date pediatrica
profile: mixed
purpose: sql_generation
expected: [evidence:fascia-pediatrica]
""".strip(),
encoding="utf-8",
)
config = tmp_path / "workspace.yaml"
config.write_text(
f"""
runtime_identity:
workspace_id: psd-clinical
workspace_revision: {'a' * 40}
dwh:
type: postgres_direct
connection: {{database: analytics, schema: mart, user: reader, password: test-only}}
vectors:
type: qdrant
base_url: http://qdrant:6333
collection: psd-clinical
embeddings:
provider: ollama_internal
base_url: http://embedding:11434
model: qwen3-embedding:0.6b
dim: 1024
evidence:
schema_version: 2
sources:
- type: filesystem
root: {workspace / 'evidence'}
roots:
sessions: {tmp_path / 'sessions'}
artifacts: {tmp_path / 'artifacts'}
indexes: {tmp_path / 'indexes'}
""".strip(),
encoding="utf-8",
)
item = SourceObject(
source_id="fs:curated-formula",
uri=curated_file.as_uri(),
fingerprint="sha256:" + "b" * 64,
metadata={"relative_path": "curated/formula/fascia-pediatrica.md"},
)
return load_config(config), Source(item, curated), item
def _pipeline(store, source, vectors, evaluator):
return CorpusPipeline(
store=store,
sources=[source],
embedder=Embedder(),
vector_store=vectors,
embedding_model="test-model",
embedding_dimensions=1024,
chunk_policy=ChunkPolicy(version="semantic:v1", max_chars=4000),
pipeline_version="evidence-v1",
workspace_id="psd-clinical",
candidate_evaluator=evaluator,
)
def test_validated_corpus_evaluates_exact_inactive_generation_before_publication(tmp_path):
cfg, source, item = _workspace(tmp_path)
preprocess_cmd._validate_materialized_curated_corpus(cfg)
store = ObservableStore(tmp_path / "corpus")
vectors = ObservableVectors(store)
evaluator = preprocess_cmd._candidate_evaluator(cfg, vector_store=vectors, embedder=Embedder())
first = _pipeline(store, source, vectors, evaluator).run()
assert first.published is True
assert store.active_generation() == first.generation
assert len(vectors.searches) == 9
assert {call["mode"] for call in vectors.searches} == {"dense", "bm25", "fused"}
assert {call["generation"] for call in vectors.searches} == {first.generation}
assert {call["active"] for call in vectors.searches} == {None}
assert store.events == [("publish", first.generation)]
assert VectorConfig().max_chunk_chars == 4000
assert set(VectorConfig.model_fields) == {"max_chunk_chars", "retain_published_generations"}
for alias in ("chunk_size", "max_chunk_size", "max_fragment_chars"):
with pytest.raises(ValidationError, match="extra_forbidden"):
VectorConfig.model_validate({alias: 4000})
assert all(len(fragment.content) <= VectorConfig().max_chunk_chars for fragment in first.manifest.chunks)
vectors.fail_evaluation = True
changed = SourceObject(
source_id=item.source_id,
uri=item.uri,
fingerprint="sha256:" + "c" * 64,
metadata=item.metadata,
)
with pytest.raises(PipelineError, match="candidate retrieval evaluation failed"):
_pipeline(store, Source(changed, source.content), vectors, evaluator).run()
assert store.active_generation() == first.generation
assert store.events == [("publish", first.generation)]
assert {entry.record.metadata["vector_generation"] for entry in vectors.records} == {first.generation}
@@ -15,3 +15,9 @@ def test_poorly_structured_authoring_fixture_exercises_the_restructuring_boundar
assert "https://" in text # URL
assert "non chiarisce" in text # ambiguity
assert "CASE WHEN" in text # PostgreSQL expression candidate
def test_acceptance_runner_includes_the_integrated_candidate_publication_gate():
runner = Path(__file__).parents[2] / "scripts" / "evidence-restructuring-acceptance.sh"
assert "test_evidence_candidate_publication.py" in runner.read_text(encoding="utf-8")
+13 -7
View File
@@ -141,47 +141,52 @@ run_git("add", "evidence")
run_git("commit", "-qm", "baseline human-reviewed curated evidence")
baseline_formula = run_git("show", "HEAD:evidence/curated/formula/fascia-pediatrica.md").stdout
fixture_target.write_text(fixture_target.read_text(encoding="utf-8") + "\nNota revisionata.\n", encoding="utf-8")
second = prepare_workspace_evidence(workspace, restructurer=fake, git_status=lambda _: ())
second = prepare_workspace_evidence(workspace, restructurer=fake)
assert second.model_calls == 1 and second.changed == ("source/notes/poorly-structured.md",)
changed_paths = set(run_git("diff", "--name-only").stdout.splitlines())
assert "evidence/manifest.yaml" in changed_paths
assert "evidence/curated/formula/fascia-pediatrica.md" in changed_paths
assert run_git("show", "HEAD:evidence/curated/formula/fascia-pediatrica.md").stdout == baseline_formula
run_git("add", "evidence")
run_git("commit", "-qm", "proposed fake restructuring")
snapshot = {
path.relative_to(workspace): path.read_bytes()
for path in (workspace / "evidence").rglob("*") if path.is_file()
}
no_op = prepare_workspace_evidence(workspace, restructurer=fake, git_status=lambda _: ())
no_op = prepare_workspace_evidence(workspace, restructurer=fake)
assert no_op.model_calls == 0 and no_op.changed == ()
assert snapshot == {
path.relative_to(workspace): path.read_bytes()
for path in (workspace / "evidence").rglob("*") if path.is_file()
}
curated_path = workspace / "evidence" / "curated" / "formula" / "fascia-pediatrica.md"
curated_path.write_text(curated_path.read_text(encoding="utf-8") + "\n", encoding="utf-8")
try:
prepare_workspace_evidence(
workspace, restructurer=fake, git_status=lambda _: (" M evidence/curated/formula/fascia-pediatrica.md",),
)
prepare_workspace_evidence(workspace, restructurer=fake)
except EvidencePreparationError as failure:
assert failure.code == "authoring_worktree_dirty"
else:
raise AssertionError("dirty curated state was not refused")
run_git("checkout", "--", "evidence/curated")
manifest_path = workspace / "evidence" / "manifest.yaml"
manifest_path.write_text(
manifest_path.read_text(encoding="utf-8").replace("evidence-authoring-v1", "evidence-authoring-v2"),
encoding="utf-8",
)
run_git("add", "evidence/manifest.yaml")
run_git("commit", "-qm", "incompatible authoring pipeline")
incompatible = manifest_path.read_bytes()
try:
prepare_workspace_evidence(workspace, restructurer=fake, git_status=lambda _: ())
prepare_workspace_evidence(workspace, restructurer=fake)
except EvidencePreparationError as failure:
assert failure.code == "pipeline_upgrade_required"
else:
raise AssertionError("incompatible pipeline version was not refused")
assert manifest_path.read_bytes() == incompatible
upgrade = prepare_workspace_evidence(workspace, restructurer=fake, git_status=lambda _: (), upgrade=True)
upgrade = prepare_workspace_evidence(workspace, restructurer=fake, upgrade=True)
assert upgrade.model_calls == 2 and set(upgrade.changed) == {
"source/notes/independent.md", "source/notes/poorly-structured.md",
}
@@ -198,6 +203,7 @@ PY
cd "$repo_root/harness"
.venv/bin/pytest -q \
tests/test_evidence_restructuring_fixture.py \
tests/test_evidence_candidate_publication.py \
tests/test_evidence_authoring.py \
tests/test_evidence_canonical.py \
tests/test_evidence_pi_restructurer.py \