35 Commits
Author SHA1 Message Date
Codex bd416f7327 Fix new-question landing and question-language HITL
Publish documentation / publish (push) Successful in 34s
Reset the activity panel when starting a new question so the landing navigation is restored. Detect and persist the original question language, pass it through runtime and widget descriptors, and scope HITL controls to that language.

Validated with gate, session, backend and frontend tests, TypeScript checks, Ruff and strict docs build. Rebuilt and restarted local core/frontend; both healthy and serving HTTP successfully.
2026-09-21 19:47:22 +02:00
Codex d8a29bfbdd Add full shell, replaceable Omics adapter and bilingual interaction
Implement approved specification #32 and tickets #33-#37. Keep host authentication server-verified and pin session interaction language. Compile scoped base selectors for browser compatibility and retain full gutters during CSS pruning.
2026-09-13 14:26:39 +02:00
Codex 82e2c91f42 feat: implement memory and evidence administration with guided repairs
Publish documentation / publish (push) Successful in 1m27s
Add PostgreSQL-backed memory, editable evidence with source review and activation, and human-approved archive repairs across the harness, API, and UI. Include migrations, deployment support, regression coverage, and validation documentation.

Refresh permissions from validated session roles so existing administrator logins can access newly deployed archive management features.
2026-09-10 10:31:34 +02:00
marcopan cc30148b69 refactor(evidence): unify formulas with typed evidence 2026-08-25 01:23:04 +02:00
marcopan f1a9b567ba feat(evidence): contribute to semantic stages (#42) 2026-08-24 22:05:30 +02:00
marcopan 694b7dd21f fix: harden reviewer workflow and memory handling 2026-07-23 12:34:23 +02:00
marcopan 2ff63d371f feat: harden workflow gates and expose token usage 2026-07-21 12:14:26 +02:00
marcopan 0cf09777f2 Fix session resume and PSD container configuration 2026-07-20 20:22:47 +02:00
marcopanandClaude Fable 5 83942c0b5c feat(harness): F6/F7 close on their last approval — no echo phase gate
Where completeness is machine-detectable, the reviewer's last substantive
approval now closes the phase itself (same pattern as F3/F4/F8):

- F6: approving the LAST CTE of the plan (kind:"cte_result" with next_cte
  now empty) advances the phase; a non-final CTE keeps the phase open and
  names the next one.
- F7: kind:"sql" records sql_approved — which IS F7's only advance
  prerequisite — and advances immediately.

Two reviewer interactions per session removed, both pure echoes. The
summary phase gate remains only where completeness is a human judgment
(F1, F2 with recorded memories, F5). SKILL.md states the rule and the
five self-closing gates; L1 tests cover last/non-last CTE and sql close.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 02:23:45 +02:00
marcopanandClaude Fable 5 3e072fe652 fix: realign workflow advance semantics and phase labels with workflow.yaml
Audit findings 2.1 + 2.2 (high).

2.1 WorkflowBar's static phase list was fiction from F2 on (F2 "Schema
linking" vs memoria, F4 "SQL plan" vs schema_linking, …): every live
session showed the wrong phase name. Both maps now mirror
harness/workflow.yaml (F1 chiarimento … F8 datamart).

2.2 forceAdvance (6ee5bda) let reviewer_decide advance:true bypass the
phase gate on ANY phase, contradicting SKILL.md's "auto-advance only
empty F2 / skipped F6". reviewer_decide is back on advanceIfReady (exit-6
no-op) and tells the model to close via reviewer_confirm; forceAdvance
stays only where selection IS the approval by design: reviewer_schema_linking
(F4) and the F8 promotion close path. SKILL.md now names the three
self-closing gates (F3 rewrite_question, F4 schema-linking advance:true,
F8 memory_promote) so gate and skill state one contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 01:28:18 +02:00
User c1cddaa667 refactor(harness): route workflow persistence through repositories 2026-07-16 18:01:29 +02:00
User 9d4e057426 fix: serialize decision ledger writes 2026-07-14 15:30:06 +02:00
User 333874a755 fix: make join approval atomic 2026-07-14 15:24:06 +02:00
User e228c6f2fd fix: harden phase summary open questions 2026-07-14 15:09:50 +02:00
User 27af6733c8 fix: make join review read only 2026-07-14 15:06:16 +02:00
User 1b1554b353 fix: auto-approve phase 3 rewrite 2026-07-14 13:51:39 +02:00
User c7d586aaf7 fix: clarify memory selection semantics 2026-07-14 13:13:59 +02:00
User 6dbf93fff9 feat: harden runtime readiness and session workflow 2026-07-14 10:27:25 +02:00
marcopan 08f0029793 fix: finalize session after memory promotion 2026-07-12 14:26:48 +02:00
marcopanandClaude Opus 4.6 893ad99594 fix(gate): auto-finalize session after last phase approval
The model sometimes stops after receiving 'Fase approvata' without calling
`tht session finalize`, leaving the session open. Now the gate itself calls
finalize after advancing the max phase (F8), making session closure
deterministic regardless of model behavior.

- reviewer_confirm kind:phase: after phase advance at max_phase, gate calls
  `tht session finalize <session>` (best-effort with recovery message)
- SKILL.md updated: model no longer needs to call finalize itself
- Tests: 2 new JS tests (auto-finalize at max phase; no-finalize at non-max)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-07-07 18:47:19 +02:00
marcopanandClaude Fable 5 e24b41b156 feat(opt): three efficiency levers for NL→SQL workflow
Lever 1: Join-graph via FK logics in annotations + suggest-fks command
  - TableAnnotation.foreign_keys field stores curated logical FKs (DWH has no FK constraints)
  - tht schema suggest-fks: mine from approved SQL, heuristics (time_key → dim_time),
    same-name discovery + explicit --assume flag for multi-owner PKs
  - mschema renders 【Foreign keys】 section populated; validation in merge.py
  - SKILL.md F4 now reads FKs from mschema-text, no custom data_time_key logic

Lever 2: Context-pack consolidation at kickoff (tht search pack)
  - Single embedding of question, reused for schema + evidence + solved searches
  - One command: tht search pack <question> --session <id> → retrieval_pack.md
  - Graceful degradation when Ollama/vector store unreachable (exit 0, empty sections)
  - SKILL.md F1 prescribes as first call; reduces model thinking turns via pre-retrieval

Lever 3: Phase-summary recap v2 auto-construction from session ledger
  - tht session show --json includes full decisions ledger
  - tht phase meta --json exports 'emits' (substantive decision types per phase)
  - Gate appends deterministic 【Decisioni registrate in questa fase】 section (appendLedgerSection)
  - Model authors only summary + checks; recap table comes from persisted state (exact by construction)
  - SKILL.md Disciplina 6: brief model output, gate fills the rest

Tests: 358 Python (including 10 FK + 3 pack + 1 session-ledger tests) + 111 JS gate tests, all pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 17:43:08 +02:00
marcopanandClaude Fable 5 34aeda006e perf(f1): cache-guard schema introspect + F1 toolbox in skill (~-5 min per session)
Transcript analysis (session 2026-07-06-175012, GLM 5.2) showed F1 at 567s:
182s wasted on a useless `tht schema introspect` (re-introspecting the remote
DWH although physical.yaml was already materialized) plus ~220s of model
thinking inflated by ~7 exploratory turns (--help/find/cat). The actual
searches cost ~15s; reviewer gates (~145s, untouched) are the quality contract.

- schema_cmd.py: introspect now exits 0 with "OK (cache)" in ~1s when
  physical.yaml exists; --refresh forces the real re-introspection.
  Deterministic cross-model guarantee, verified live on psd (163 tables, 1.2s).
- SKILL.md: F1 toolbox (only `tht search find` + `tht schema render`; no
  introspect/--help/filesystem browsing; batch all searches in one turn);
  F4 step 1 is render-only with a one-shot introspect fallback.
- tht-gate.js: `tht schema introspect ... --refresh` added to FORBIDDEN
  (maintenance stays shell-only, never in-session).
- tests: 4 new pytest cases (cache hit placement proven with fake credentials,
  refresh bypass, corrupt-catalog fall-through, render fallback message) and
  2 gate anti-bypass JS cases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 16:02:45 +02:00
marcopanandClaude Fable 5 8af0552dcc docs(skill): prescribe solved-question recall in F4/F6/F7; refresh PROJECT_STATE
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:54:42 +02:00
marcopanandClaude Fable 5 ae89957175 docs(skill): prescribe the F8 memory-promotion gate; drop optional D11 notes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 13:34:04 +02:00
marcopanandClaude Fable 5 3ad93cd02f feat(gate): deterministic v2 review-gate payloads (cte_plan/cte_result/phase)
The tht-gate.js reviewer_confirm now builds structured v2 artifacts before the
widget so the reviewer approves gate-derived data, not raw model text:

- new pure modules gate/artifact-contracts.js (soft validators, {ok,errors},
  legacy-passthrough) and gate/enrich.js (index/description enrichment,
  buildCteResultV2 fusing thin model data with `tht cte info`, phase enrichment)
- cte_plan v2: validate + enrich + persist via `tht cte plan --name … --doc -`
  (names derived from data.ctes[]); legacy `names` param kept as fallback
- cte_result v2: rebuild from `tht cte next`/`tht cte info` (sql + preview from
  the persisted test record); null/error last_test -> actionable textResult
- phase v2: soft-validate + fill phase from meta + catalog descriptions
- prepareReviewerArguments coerces artifact.data too (GLM double-stringify);
  legacy markdown strings pass through unchanged
- SKILL.md: Phase 6 cte_plan payload A + thin cte_result guidance; Discipline 6
  payload C example; Discipline 7 reworded for gate-rebuilt cte_result

Legacy (non-v2) paths unchanged. TypeBox stays Type.Any() for artifact.data;
validation is soft (textResult) so models self-correct instead of looping.
All 102 gate JS tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 00:42:48 +02:00
marcopanandMarco Pancotti bd63e3194f docs(skill): Phase 4 uses reviewer_schema_linking; honor curated output columns 2026-07-06 23:25:00 +02:00
marcopan 78518210a0 docs(skill): document empty-memory auto-advance (no empty checklist) in Phase 2 2026-07-02 18:39:12 +02:00
marcopanandClaude Opus 4.8 3dadc6fbb5 docs(skill): F7 is two-step — kind:"sql" records, kind:"phase" advances
The cheat-sheet, Discipline 2, and Phase 7 said F7 closes with
reviewer_confirm kind:"sql" — but the gate's kind:"sql" only records
sql_approved:phase:7; F7 is not in _AUTO_ADVANCE_PHASES, so advancing to
F8 still needs reviewer_confirm kind:"phase" (mirrors F6). Same
record-vs-advance trap this branch removes elsewhere.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 19:55:23 +02:00
marcopanandClaude Opus 4.8 01de6f330f docs(skill): correct phase-advance contract + add phase cheat-sheet
F3/F4 close with reviewer_confirm kind:'phase' (they do NOT auto-advance);
advance:true only auto-advances F2-empty/F6-skip; rewrite_question belongs
to F3 not F1; F4 uses the new write_schema_linking tool with the documented
SchemaLinking shape. Adds a per-phase artifact/close cheat-sheet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 16:55:49 +02:00
marcopanandClaude Opus 4.8 cef9ae4368 feat(harness): single-select answers auto-confirm (reviewer_select persists)
F — reviewer_select options may now carry a `decision` payload {type, subject,
detail?, rationale?} plus an optional `advance`. Picking such an option IS the
confirmation: the gate persists it directly (tht decision add) and optionally
advances, with no redundant reviewer_decide/reviewer_confirm follow-up gate.
Options without a payload stay ask-only; back/exit/Other never persist.

Pure logic extracted + exported for unit tests: resolveSelectOutcome (classifies
the response) and decisionAddArgs (shared with reviewer_decide, DRY). Gate JS
suite 33/33 (gate_select_decision.test.js, +5); harness pytest 269 unchanged.

Contract docs updated together: reviewer_select tool description, SKILL.md
(widget summary, disciplines 2-3, Phase-1 single-pick), and the CLAUDE.md gate
note. Live verification (model truly emits reviewer_select+decision, decision in
review_decisions.jsonl, no follow-up gate) deferred to workstream G — it is
model-behavior-dependent.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 18:01:50 +02:00
marcopanandClaude Opus 4.8 418187a4ad fix(backend): echo Pi's RPC id so reviewer gates unblock after answer
ctx.ui.input in `pi --mode rpc` correlates extension_ui_response on its own
top-level RPC id (crypto.randomUUID), not the descriptor id the gate carries
in `title`. SessionBridge replied with the descriptor id, so Pi silently
dropped the response and the model never resumed — every reviewer widget hung
after the human answered.

SessionBridge now stores Pi's top-level m.id (pendingPiId) and replies
extension_ui_response{ id: pendingPiId, value: <uiResponse> }; value still
carries the descriptor id so the gate's internal resp.id === descriptor.id
check still holds.

The fake-pi double had masked the bug by forcing m.id == descriptor.id; it now
mirrors real Pi (distinct randomUUID, correlate on it, drop unknown ids), with
a negative regression test. SKILL.md Phase 1 also now steers multi-answer
disambiguation to reviewer_decide (multiselect).

Tests: backend 67/67, tsc clean, fake-pi contract 2/2.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 10:43:49 +02:00
marcopan 57580676a3 docs(harness): add Phase 0 Resume cold-start procedure to the orchestrator skill 2026-06-29 12:41:27 +02:00
marcopanandClaude Opus 4.8 c4d130828f fix(harness): remediation difetti review — gate↔CLI, D15, D7/D6, D14, robustezza
Implementazione del piano di remediation progressiva sui difetti emersi
dall'analisi dell'harness. Tutto verificato: 214 test Python (incl. L0 su
Postgres reale), 14 test JS del gate, ruff pulito.

Blocco 1 (CRITICA, integrazione gate↔CLI):
- phase advance: gate usa --auto + exit 6; reviewer_confirm kind:phase fa
  advance esplicito che applica i prerequisiti (prima non avanzava per le
  fasi a conferma umana).
- cte plan riceve i --name dal gate (param names); set-question con id
  posizionale; skill `tht search find`; nuovo comando `tht memory save-one`
  con dedup hash client-side in save_one_memory.

Blocco 2 (D15, stato post-rollback):
- campo `phase` su DecisionRecord + effective_decisions phase-aware per i
  subject "a nome" (cte_approved ecc.); _compute_promotions e finalize sulla
  vista effective; finalize confronta col piano CTE effettivo, non glob;
  `decision add --retracts` + comando `decision retract`.

Blocco 3 (D7 read-only + D6 manifest):
- assert_read_only su tutti e quattro i codepath (direct + REST);
- manifest author/summary/updated_at/updated_by/schema_version popolati +
  helper touch_manifest sulle mutazioni.

Blocco 4-5 (D14a/D14b):
- decision_min_phase data-driven via `emits:` in workflow.yaml;
- formula evidence: status auto, search_formulas, gruppo CLI `tht formula`,
  `search find --kind formula`, load_evidence_dir salta i .sql.md.

Blocco 6 (robustezza):
- taskdoc slice promoted_tables + bound enforced; report escaping/bound +
  rsplit note; filtro kind reader REST/direct; conteggio upserted robusto;
  guard REST run_query non-list; LSH disallineato -> LshIndexError.

Blocco 7 (pulizia):
- dead code gate e KIND_TO_TABLE morto rimossi; doc Postgres-only
  (README + connection.py).

Blocco 0 (parziale): test di compatibilità firma gate↔CLI
(tests/integration). Rinviati: fake-Pi runtime completo, artifact-gate da
disco (#23), parità eligibility REST/direct (#28), unificazione
reserved-labels (#30), memory_rejected da deselezione (#33).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 17:16:51 +02:00
marcopan 292048f777 feat(harness): language workspace param + skill riscritta in inglese (semantica completa)
Due cambiamenti interconnessi da user review:

1. language come parametro workspace (spec decisione 9):
   - Config.language (default 'en') + workspaces PSD con 'language: it'
   - Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
     sono nel workspace language; le istruzioni della skill restano in inglese
     (piu' affidabili per modelli piccoli, meno ambigue)

2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
   (autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
   - 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
     -> 'never advance a phase or record a decision without confirmation'
   - recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
     reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
     auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
     da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
     'leave ambiguity open', F1 passa lista completa non solo ultima
   - language contract esplicito (istruzioni EN, output nel workspace language)

Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).

Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
2026-06-27 14:50:30 +02:00
marcopan de61034a8d feat(harness): riscrittura skill tht-sessione (F1-F8 + 4 sottomoduli)
Skill ex-novo che riflette Thoth (non copia di ChironeWp3):
- vocabolario widget-descriptor (reviewer_select/decide/confirm) invece di 'dialog native'
- D11 save-one in F2 (upsert mirato vs full resync)
- D14a value_grounded (LSH multi-colonna non collassa) + D14b concept_formula in F4
- D13 free-text e D15 rollback nelle discipline trasversali
- memory vive SOLO nel vectordb (drop registry, spec 5): save-one/promote senza registry
- F8 datamart onesto (stub NotImplementedError)

Sottomoduli cte/memoria/rewriting/sql-generation portati adattando nsp->tht, con
i vincoli precisi trasferiti fedelmente (max 5 memorie, solo 3 tipi riusabili,
CTE solo WITH senza SELECT, dim_time join non aritmetica, sql_final pulito).

Verifica: zero residui nsp/chirone/psd nella skill; ogni 'tht <cmd>' citato e'
registrato (correzione: 'tht formula retrieve' era inesistente -> riformulato in
'ricerca nelle evidence'). Suite: 165 passed.
2026-06-27 14:24:20 +02:00