Reset the activity panel when starting a new question so the landing navigation is restored. Detect and persist the original question language, pass it through runtime and widget descriptors, and scope HITL controls to that language.
Validated with gate, session, backend and frontend tests, TypeScript checks, Ruff and strict docs build. Rebuilt and restarted local core/frontend; both healthy and serving HTTP successfully.
Implement approved specification #32 and tickets #33-#37. Keep host authentication server-verified and pin session interaction language. Compile scoped base selectors for browser compatibility and retain full gutters during CSS pruning.
Add PostgreSQL-backed memory, editable evidence with source review and activation, and human-approved archive repairs across the harness, API, and UI. Include migrations, deployment support, regression coverage, and validation documentation.
Refresh permissions from validated session roles so existing administrator logins can access newly deployed archive management features.
Where completeness is machine-detectable, the reviewer's last substantive
approval now closes the phase itself (same pattern as F3/F4/F8):
- F6: approving the LAST CTE of the plan (kind:"cte_result" with next_cte
now empty) advances the phase; a non-final CTE keeps the phase open and
names the next one.
- F7: kind:"sql" records sql_approved — which IS F7's only advance
prerequisite — and advances immediately.
Two reviewer interactions per session removed, both pure echoes. The
summary phase gate remains only where completeness is a human judgment
(F1, F2 with recorded memories, F5). SKILL.md states the rule and the
five self-closing gates; L1 tests cover last/non-last CTE and sql close.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit findings 5.1-5.3.
5.1 `phase reopen` now appends `phase_reopened` BEFORE the artifact
teardown: a crash between the two used to leave later-phase artifacts
deleted with the ledger still at the old phase (resume entered a phase
missing its artifacts). The inverse half-state — reopened with stale later
artifacts — is benign. Order locked by tests/test_phase_reopen_order.py.
5.2 New `tht decision add-batch --doc -`: N substantive decisions in ONE
atomic ledger write (meta types and cte_approved stay on `decision add`;
strictest min-phase enforced). reviewer_schema_linking now builds the
complete curation set and persists it with a single add-batch call — a
mid-loop failure can no longer leave the audit ledger half-written, and a
retry cannot duplicate the first K decisions.
5.3 The anti-bypass hook now also blocks BASH mutations of protected
state (`echo >> review_decisions.jsonl`, `sed -i` on the manifest,
`cat > tht-gate.js`, python open('w'), mv/rm/tee/…): FORBIDDEN only
covered tht subcommands and the write/edit hook only covered pi's own
tools. Read-only access (cat/grep/tail/ls) stays allowed.
Also: knownDecisionTypes is defensive — a workflow meta declaring NO
emits at all (older tht, minimal stubs) skips pre-validation instead of
rejecting every substantive type; with emits present, unknown types are
still rejected before the widget (new L1 test).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit findings 2.1 + 2.2 (high).
2.1 WorkflowBar's static phase list was fiction from F2 on (F2 "Schema
linking" vs memoria, F4 "SQL plan" vs schema_linking, …): every live
session showed the wrong phase name. Both maps now mirror
harness/workflow.yaml (F1 chiarimento … F8 datamart).
2.2 forceAdvance (6ee5bda) let reviewer_decide advance:true bypass the
phase gate on ANY phase, contradicting SKILL.md's "auto-advance only
empty F2 / skipped F6". reviewer_decide is back on advanceIfReady (exit-6
no-op) and tells the model to close via reviewer_confirm; forceAdvance
stays only where selection IS the approval by design: reviewer_schema_linking
(F4) and the F8 promotion close path. SKILL.md now names the three
self-closing gates (F3 rewrite_question, F4 schema-linking advance:true,
F8 memory_promote) so gate and skill state one contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit finding 1.1 (high): reviewer_memory_promote, rewrite_question,
write_schema_linking, write_cte_sql and write_final_sql still ran execute
without a top-level catch — the same unhandled-rejection class that killed
Pi in reviewer_schema_linking (fixed in 87cb806 for the four reviewer_*
tools). All 9 registered tools now share the pattern: any uncaught throw
becomes a textResult the model can react to.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The last live session crashed during reviewer_schema_linking: an uncaught
exception (likely from execFileSync in currentPhase/phaseMeta or from
cat.columns being undefined) rejected the async execute() Promise. Pi does
not catch rejected tool Promises — Node.js treats them as unhandled
rejections and kills the process.
Fix:
- Wrap all four reviewer tool execute bodies (select/decide/confirm/
schema_linking) in a top-level try/catch → returns a textResult on any
unexpected error instead of crashing Pi.
- reviewer_schema_linking: defensively re-parse `tables` if still a string
(belt-and-suspenders over prepareArguments), guard cat.columns before .map().
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Gate (tht-gate.js):
- Pre-validate decision types against workflow.yaml before showing reviewer widget
- Reject decisions emitted by later phases (min-phase check)
- Copy top-level `kind` into artifact when model forgets it (prevents loop)
- Force-advance on reviewer_decide/schema_linking when advance:true — skip
redundant reviewer_confirm gate
Backend:
- Emit agent_end on clean Pi exit (code 0 + bridge idle) instead of marking failed
Frontend:
- Strip <think> tags from transcript and activity panel
- Fix mermaid render with offscreen container + cleanup
- Graceful mermaid error: show source code instead of red error, fall back to table
Workflow:
- F2 now emits table_promoted and table_excluded (early schema linking decisions)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.
- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
`tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
- TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
- tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
same-name discovery + explicit --assume flag for multi-owner PKs
- mschema renders 【Foreign keys】 section populated; validation in merge.py
- SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic
Lever 2: Context-pack consolidation at kickoff (tht search pack)
- Single embedding of question, reused for schema + evidence + solved searches
- One command: tht search pack <question> --session <id> → retrieval_pack.md
- Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
- SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval
Lever 3: Phase-summary recap v2 auto-construction from session ledger
- tht session show --json includes full decisions ledger
- tht phase meta --json exports 'emits' (substantive decision types per phase)
- Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
- Model authors only summary + checks; recap table comes from persisted state (exact by construction)
- SKILL.md Disciplina 6: brief model output, gate fills the rest
Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.
- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
physical.yaml exists; --refresh forces the real re-introspection.
Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
introspect/--help/filesystem browsing; batch all searches in one turn);
F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
(maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
refresh bypass, corrupt-catalog fall-through, render fallback message) and
2 gate anti-bypass JS cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Extract shared statusBadgeClass() helper (src/viewers/statusBadge.ts) so
SqlViewer, CteResultViewer and PhaseSummaryViewer can't drift: ok/success/
passed/promoted render green (--success), warn renders amber (--warning),
error/failed stay destructive red, everything else stays neutral outline.
Previously CteResultViewer/PhaseSummaryViewer mapped "ok"/"promoted" to the
default badge variant, which is bg-primary (GSD red) — success states
rendered red.
- enrich.js: buildCteResultV2 now falls back preview.rows to [] instead of
null when last_test.preview_rows is missing (pre-upgrade ok records), and
CteResultViewer reads result.preview?.rows?.length with a null-safe
fallback so it degrades to the "No preview rows" empty state instead of
crashing.
- PhaseSummaryViewer: section.items is optional (model-authored sections can
be prose-only); render (section.items ?? []) instead of crashing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:
- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
legacy-passthrough) and gate/enrich.js (index/description enrichment,
buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
(names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
payload C example; Discipline 7 reworded for gate-rebuilt cte_result
Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>