Embeddings timeout was 120s, causing multi-minute hangs when Ollama was
down during F4/F6/F7 solved-search. Now: connect_timeout=5s across all
HTTP clients (REST + Ollama), read_timeout reduced to 30s for embeddings,
and OllamaEmbeddings auto-restarts the server on ConnectionError before
degrading gracefully.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.
- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
`tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- PROJECT_STATE.md: section on three deployed levers (FK annotations, context-pack, recap v2)
with setup instructions for new workspaces; psd pre-configured
- README.md: one-time setup for workspace (tht schema suggest-fks); note that levers 2+3 auto-activate
- Updated 'How to run' with explicit command and prereq checklist
- Last-updated timestamp: 2026-07-08
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
extract.mjs now emits a schema-linking ui_request descriptor (columns rendered
from the model's suggested_columns, since the live catalog enrichment needs the
DWH) and parses the reviewer's "N tabelle" outcome; when neither the manifest
nor question.md is reachable, the NL question is approximated by de-slugging
the session id. replay.json regenerated from the 2026-07-06-175012 session
(18 answered gates, includes the F4 schema-linking gate).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.
- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
physical.yaml exists; --refresh forces the real re-introspection.
Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
introspect/--help/filesystem browsing; batch all searches in one turn);
F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
(maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
refresh bypass, corrupt-catalog fall-through, render fallback message) and
2 gate anti-bypass JS cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The original doc only covered search_similar; the write functions also need
solved_question in their kind whitelist for the memory table.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The FE derived 'working' purely as activeSession && !pendingWidget, so the
final workflow turn — the only one that ends without a follow-up gate —
left the spinner on forever (observed live: 21592s after F8 approve).
- SessionBridge maps Pi's agent_end -> SSE system_event {event: agent_end}
- PiProcessManager notifies the client (info error + synthetic agent_end)
when the child dies unexpectedly; expected teardowns stay silent
- sessionStore tracks agentActive (on: user entry/text_delta/ui_request,
off: agent_end); AppShell working now requires it; resume sets it
optimistically
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Extract shared statusBadgeClass() helper (src/viewers/statusBadge.ts) so
SqlViewer, CteResultViewer and PhaseSummaryViewer can't drift: ok/success/
passed/promoted render green (--success), warn renders amber (--warning),
error/failed stay destructive red, everything else stays neutral outline.
Previously CteResultViewer/PhaseSummaryViewer mapped "ok"/"promoted" to the
default badge variant, which is bg-primary (GSD red) — success states
rendered red.
- enrich.js: buildCteResultV2 now falls back preview.rows to [] instead of
null when last_test.preview_rows is missing (pre-upgrade ok records), and
CteResultViewer reads result.preview?.rows?.length with a null-safe
fallback so it degrades to the "No preview rows" empty state instead of
crashing.
- PhaseSummaryViewer: section.items is optional (model-authored sections can
be prose-only); render (section.items ?? []) instead of crashing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Regenerated demo fixture output with augment-schema-linking.mjs +
augment-review-gates.mjs applied, so the F6 cte_plan gate, CTE 1 cte_result
gate, and Fase 5 phase-summary gate carry v2 schema_version payloads and
render via CtePlanViewer/CteResultViewer/PhaseSummaryViewer in the replay
server. Regenerable; committed so the replay server serves it without a
build step.
Add cte-plan-fixture.json, cte-result-fixture.json, and
phase-summary-fixture.json — realistic Italian-language payloads for the
cardiology DWH replay session, conforming exactly to contracts.md (A/B/C).
Add augment-review-gates.mjs (pattern of augment-schema-linking.mjs): finds
the F6 cte_plan gate, first CTE 1 cte_result gate, and Fase 5 phase-summary
gate by title regex and swaps in the v2 artifact.data, so CtePlanViewer,
CteResultViewer, and PhaseSummaryViewer render in the offline replay.
Idempotent; exits 1 if an expected gate is missing.
Document the augment scripts and re-extract note in tools/replay/README.md.
Add CtePlanViewer, CteResultViewer, and PhaseSummaryViewer to render the
structured v2 payloads (schema_version: 2) the harness now emits for
artifact-gate widgets, per contracts.md. Extract PreviewGrid from
ResultsPanel as a reusable AG Grid component shared by CteResultViewer.
ArtifactView dispatches to the new viewers on a schema_version/shape
guard, falling back to the existing legacy renderers unchanged for
older sessions and replay fixtures.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:
- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
legacy-passthrough) and gate/enrich.js (index/description enrichment,
buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
(names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
payload C example; Discipline 7 reworded for gate-rebuilt cte_result
Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
WS1 of review-gates-v2: gives the JS gate (WS2) deterministic data to build the
cte_result v2 payload.
- CteTestRecord gains optional preview_rows (JSON-coerced, truncated cells);
test_cmd populates it from the bounded result rows.
- New read-only `tht cte info <name> --session <id> [--json]`: plan
index/total, persisted .sql, cte_plan_doc.json entry (if any), last
CteTestRecord. Exits 1 with a clean stderr message on missing
session/plan/name/sql.
- `tht cte plan --doc -` validates a chain-doc JSON (ctes[].name must match
--name, same order) and writes it to cte_plan_doc.json; cte_plan.json stays
a plain list[str] (load-bearing for tht.phase.next_cte). --doc is optional.
Regenerated demo fixture output of extract.mjs (WorkflowBar phase tags) with the
F4 gate-7 augment applied (tools/replay/augment-schema-linking.mjs, on the F4
branch). Regenerable; committed so the replay server serves it without a build step.
- ArtifactGateWidget: gate modal 90%→70% of viewport, larger heading, more padding.
- ArtifactView StructuredValue: depth-aware hierarchy (top-level hairline dividers,
warm bg-muted card surfaces + shadow for object-array items).
- index.css: .thot-label micro-labels switched to the mono register (SF Mono) with
tightened tracking, distinguishing labels/meta from sans body prose.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Wide tables (e.g. 116 columns, 3 suggested) made curation a scroll hunt. Now the
modal lists suggested columns first (stable sort on the descriptor flag, so
toggling never reorders rows) and adds a filter box matching name + description,
with a "shown/total" count. Selection stays keyed by column name, so the response
column order is unchanged (catalog order). Applies to read-only (excluded) tables too.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Bite-sized TDD tasks: schema-linking types + columns modal, gate widget with
staged per-table column selection, registry/summary wiring, and a replay
fixture generated from the real physical.yaml catalog. Offline-verifiable in
the replay server; harness wiring is Plan 2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reviewer curates, per promoted table, which columns to use over the full
catalog list (suggested pre-selected + bold); selection persists into
schema_linking.json + the decision ledger and softly guides SQL generation
(Option 1). Dedicated structured schema-linking gate widget (Approach A),
staged commit, inline gate + single columns modal. Hard SQL enforcement is
an explicit follow-up.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>