Commit Graph
30 Commits
Author SHA1 Message Date
marcopan 694b7dd21f fix: harden reviewer workflow and memory handling 2026-07-23 12:34:23 +02:00
marcopan 2ff63d371f feat: harden workflow gates and expose token usage 2026-07-21 12:14:26 +02:00
marcopan 0cf09777f2 Fix session resume and PSD container configuration 2026-07-20 20:22:47 +02:00
marcopanandClaude Fable 5 83942c0b5c feat(harness): F6/F7 close on their last approval — no echo phase gate
Where completeness is machine-detectable, the reviewer's last substantive
approval now closes the phase itself (same pattern as F3/F4/F8):

- F6: approving the LAST CTE of the plan (kind:"cte_result" with next_cte
  now empty) advances the phase; a non-final CTE keeps the phase open and
  names the next one.
- F7: kind:"sql" records sql_approved — which IS F7's only advance
  prerequisite — and advances immediately.

Two reviewer interactions per session removed, both pure echoes. The
summary phase gate remains only where completeness is a human judgment
(F1, F2 with recorded memories, F5). SKILL.md states the rule and the
five self-closing gates; L1 tests cover last/non-last CTE and sql close.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 02:23:45 +02:00
marcopanandClaude Fable 5 3e072fe652 fix: realign workflow advance semantics and phase labels with workflow.yaml
Audit findings 2.1 + 2.2 (high).

2.1 WorkflowBar's static phase list was fiction from F2 on (F2 "Schema
linking" vs memoria, F4 "SQL plan" vs schema_linking, …): every live
session showed the wrong phase name. Both maps now mirror
harness/workflow.yaml (F1 chiarimento … F8 datamart).

2.2 forceAdvance (6ee5bda) let reviewer_decide advance:true bypass the
phase gate on ANY phase, contradicting SKILL.md's "auto-advance only
empty F2 / skipped F6". reviewer_decide is back on advanceIfReady (exit-6
no-op) and tells the model to close via reviewer_confirm; forceAdvance
stays only where selection IS the approval by design: reviewer_schema_linking
(F4) and the F8 promotion close path. SKILL.md now names the three
self-closing gates (F3 rewrite_question, F4 schema-linking advance:true,
F8 memory_promote) so gate and skill state one contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 01:28:18 +02:00
User c1cddaa667 refactor(harness): route workflow persistence through repositories 2026-07-16 18:01:29 +02:00
User 9d4e057426 fix: serialize decision ledger writes 2026-07-14 15:30:06 +02:00
User 333874a755 fix: make join approval atomic 2026-07-14 15:24:06 +02:00
User e228c6f2fd fix: harden phase summary open questions 2026-07-14 15:09:50 +02:00
User 27af6733c8 fix: make join review read only 2026-07-14 15:06:16 +02:00
User 1b1554b353 fix: auto-approve phase 3 rewrite 2026-07-14 13:51:39 +02:00
User c7d586aaf7 fix: clarify memory selection semantics 2026-07-14 13:13:59 +02:00
User 6dbf93fff9 feat: harden runtime readiness and session workflow 2026-07-14 10:27:25 +02:00
marcopan 08f0029793 fix: finalize session after memory promotion 2026-07-12 14:26:48 +02:00
marcopanandClaude Opus 4.6 893ad99594 fix(gate): auto-finalize session after last phase approval
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.

- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
  `tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-07-07 18:47:19 +02:00
marcopanandClaude Fable 5 e24b41b156 feat(opt): three efficiency levers for NL→SQL workflow
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
  - TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
  - tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
    same-name discovery + explicit --assume flag for multi-owner PKs
  - mschema renders 【Foreign keys】 section populated; validation in merge.py
  - SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic

Lever 2: Context-pack consolidation at kickoff (tht search pack)
  - Single embedding of question, reused for schema + evidence + solved searches
  - One command: tht search pack <question> --session <id> → retrieval_pack.md
  - Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
  - SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval

Lever 3: Phase-summary recap v2 auto-construction from session ledger
  - tht session show --json includes full decisions ledger
  - tht phase meta --json exports 'emits' (substantive decision types per phase)
  - Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
  - Model authors only summary + checks; recap table comes from persisted state (exact by construction)
  - SKILL.md Disciplina 6: brief model output, gate fills the rest

Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 17:43:08 +02:00
marcopanandClaude Fable 5 34aeda006e perf(f1): cache-guard schema introspect + F1 toolbox in skill (~-5 min per session)
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.

- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
  physical.yaml exists; --refresh forces the real re-introspection.
  Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
  introspect/--help/filesystem browsing; batch all searches in one turn);
  F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
  (maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
  refresh bypass, corrupt-catalog fall-through, render fallback message) and
  2 gate anti-bypass JS cases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 16:02:45 +02:00
marcopanandClaude Fable 5 8af0552dcc docs(skill): prescribe solved-question recall in F4/F6/F7; refresh PROJECT_STATE
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:54:42 +02:00
marcopanandClaude Fable 5 ae89957175 docs(skill): prescribe the F8 memory-promotion gate; drop optional D11 notes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:34:04 +02:00
marcopanandClaude Fable 5 3ad93cd02f feat(gate): deterministic v2 review-gate payloads (cte_plan/cte_result/phase)
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:

- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
  legacy-passthrough) and gate/enrich.js (index/description enrichment,
  buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
  (names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
  the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
  legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
  payload C example; Discipline 7 reworded for gate-rebuilt cte_result

Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 00:42:48 +02:00
marcopanandMarco Pancotti bd63e3194f docs(skill): Phase 4 uses reviewer_schema_linking; honor curated output columns 2026-07-06 23:25:00 +02:00
marcopan 78518210a0 docs(skill): document empty-memory auto-advance (no empty checklist) in Phase 2 2026-07-02 18:39:12 +02:00
marcopanandClaude Opus 4.8 3dadc6fbb5 docs(skill): F7 is two-step — kind:"sql" records, kind:"phase" advances
The cheat-sheet, Discipline 2, and Phase 7 said F7 closes with
reviewer_confirm kind:"sql" — but the gate's kind:"sql" only records
sql_approved:phase:7; F7 is not in _AUTO_ADVANCE_PHASES, so advancing to
F8 still needs reviewer_confirm kind:"phase" (mirrors F6). Same
record-vs-advance trap this branch removes elsewhere.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 19:55:23 +02:00
marcopanandClaude Opus 4.8 01de6f330f docs(skill): correct phase-advance contract + add phase cheat-sheet
F3/F4 close with reviewer_confirm kind:'phase' (they do NOT auto-advance);
advance:true only auto-advances F2-empty/F6-skip; rewrite_question belongs
to F3 not F1; F4 uses the new write_schema_linking tool with the documented
SchemaLinking shape. Adds a per-phase artifact/close cheat-sheet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 16:55:49 +02:00
marcopanandClaude Opus 4.8 cef9ae4368 feat(harness): single-select answers auto-confirm (reviewer_select persists)
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.

Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.

Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 18:01:50 +02:00
marcopanandClaude Opus 4.8 418187a4ad fix(backend): echo Pi's RPC id so reviewer gates unblock after answer
ctx.ui.input in `pi --mode rpc` correlates extension_ui_response on its own
top-level RPC id (crypto.randomUUID), not the descriptor id the gate carries
in `title`. SessionBridge replied with the descriptor id, so Pi silently
dropped the response and the model never resumed — every reviewer widget hung
after the human answered.

SessionBridge now stores Pi's top-level m.id (pendingPiId) and replies
extension_ui_response{ id: pendingPiId, value: <uiResponse> }; value still
carries the descriptor id so the gate's internal resp.id === descriptor.id
check still holds.

The fake-pi double had masked the bug by forcing m.id == descriptor.id; it now
mirrors real Pi (distinct randomUUID, correlate on it, drop unknown ids), with
a negative regression test. SKILL.md Phase 1 also now steers multi-answer
disambiguation to reviewer_decide (multiselect).

Tests: backend 67/67, tsc clean, fake-pi contract 2/2.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 10:43:49 +02:00
marcopan 57580676a3 docs(harness): add Phase 0 Resume cold-start procedure to the orchestrator skill 2026-06-29 12:41:27 +02:00
marcopanandClaude Opus 4.8 c4d130828f fix(harness): remediation difetti review — gate↔CLI, D15, D7/D6, D14, robustezza
Implementazione del piano di remediation progressiva sui difetti emersi
dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su
Postgres reale), 14 test JS del gate, ruff pulito.

Blocco 1 (CRITICA, integrazione gate↔CLI):
- phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa
  advance esplicito che applica i prerequisiti (prima non avanzava per le
  fasi a conferma umana).
- cte plan riceve i --name dal gate (param names); set-question con id
  posizionale; skill `tht search find`; nuovo comando `tht memory save-one`
  con dedup hash client-side in save_one_memory.

Blocco 2 (D15, stato post-rollback):
- campo `phase` su DecisionRecord + effective_decisions phase-aware per i
  subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla
  vista effective; finalize confronta col piano CTE effettivo, non glob;
  `decision add --retracts` + comando `decision retract`.

Blocco 3 (D7 read-only + D6 manifest):
- assert_read_only su tutti e quattro i codepath (direct + REST);
- manifest author/summary/updated_at/updated_by/schema_version popolati +
  helper touch_manifest sulle mutazioni.

Blocco 4-5 (D14a/D14b):
- decision_min_phase data-driven via `emits:` in workflow.yaml;
- formula evidence: status auto, search_formulas, gruppo CLI `tht formula`,
  `search find --kind formula`, load_evidence_dir salta i .sql.md.

Blocco 6 (robustezza):
- taskdoc slice promoted_tables + bound enforced; report escaping/bound +
  rsplit note; filtro kind reader REST/direct; conteggio upserted robusto;
  guard REST run_query non-list; LSH disallineato -> LshIndexError.

Blocco 7 (pulizia):
- dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only
  (README + connection.py).

Blocco 0 (parziale): test di compatibilità firma gate↔CLI
(tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da
disco (#23), parità eligibility REST/direct (#28), unificazione
reserved-labels (#30), memory_rejected da deselezione (#33).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 17:16:51 +02:00
marcopan 292048f777 feat(harness): language workspace param + skill riscritta in inglese (semantica completa)
Due cambiamenti interconnessi da user review:

1. language come parametro workspace (spec decisione 9):
   - Config.language (default 'en') + workspaces PSD con 'language: it'
   - Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
     sono nel workspace language; le istruzioni della skill restano in inglese
     (piu' affidabili per modelli piccoli, meno ambigue)

2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
   (autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
   - 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
     -> 'never advance a phase or record a decision without confirmation'
   - recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
     reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
     auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
     da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
     'leave ambiguity open', F1 passa lista completa non solo ultima
   - language contract esplicito (istruzioni EN, output nel workspace language)

Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).

Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
2026-06-27 14:50:30 +02:00
marcopan de61034a8d feat(harness): riscrittura skill tht-sessione (F1-F8 + 4 sottomoduli)
Skill ex-novo che riflette Thoth (non copia di ChironeWp3):
- vocabolario widget-descriptor (reviewer_select/decide/confirm) invece di 'dialog native'
- D11 save-one in F2 (upsert mirato vs full resync)
- D14a value_grounded (LSH multi-colonna non collassa) + D14b concept_formula in F4
- D13 free-text e D15 rollback nelle discipline trasversali
- memory vive SOLO nel vectordb (drop registry, spec 5): save-one/promote senza registry
- F8 datamart onesto (stub NotImplementedError)

Sottomoduli cte/memoria/rewriting/sql-generation portati adattando nsp->tht, con
i vincoli precisi trasferiti fedelmente (max 5 memorie, solo 3 tipi riusabili,
CTE solo WITH senza SELECT, dim_time join non aritmetica, sql_final pulito).

Verifica: zero residui nsp/chirone/psd nella skill; ogni 'tht <cmd>' citato e'
registrato (correzione: 'tht formula retrieve' era inesistente -> riformulato in
'ricerca nelle evidence'). Suite: 165 passed.
2026-06-27 14:24:20 +02:00