Unisce gli internals di Codex (secret-bundle, provider-credentials, auth upstream,
security hardening, CI multiarch) mantenendo le fix portal-specific:
- backend: configPath da THT_CONFIG (fix sessioni) + dataRoot di Codex; authMode 'upstream'
- Docker/compose: TENUTO il mio (verificato live: omics_network+alias, env_file, pi npm-g)
perche' il compose/Dockerfile/entrypoint di Codex sono accoppiati al suo modello
secret-bundle (tht doctor inesistente, secret-policy.sh). Adottabile in futuro.
- config.test.ts: preso Codex (superset)
Verificato: tsc clean, 132/132 vitest.
Bite-sized TDD tasks: schema-linking types + columns modal, gate widget with
staged per-table column selection, registry/summary wiring, and a replay
fixture generated from the real physical.yaml catalog. Offline-verifiable in
the replay server; harness wiring is Plan 2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reviewer curates, per promoted table, which columns to use over the full
catalog list (suggested pre-selected + bold); selection persists into
schema_linking.json + the decision ledger and softly guides SQL generation
(Option 1). Dedicated structured schema-linking gate widget (Approach A),
staged commit, inline gate + single columns modal. Hard SQL enforcement is
an explicit follow-up.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- present() = presentBlockingWidget: transport + no-limbo loop + response
validation; both branches share validateUiResponse; the id invariant must be
validated explicitly in the TUI branch (emitAndWait no longer runs there)
- TUI renderer uses numbered options + index parsing, never label mapping
(frontend already answers with ids); artifact-gate editor is viewer-only,
returned content ignored
- guards resolved: both TUI-only notices now ctx.mode === "tui"
- interactive-render.js as two layers (renderTuiDescriptor +
validateSyntheticResponse); add negative test cases + a real TUI smoke
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- #1 hasUI: real 0.80.3 comment is "true in TUI and RPC modes" (doc had the old
0.73.1 "false in print/RPC"). hasUI no longer distinguishes TUI from RPC, so it
is NOT a fallback for the interactive branch — only ctx.mode === "tui" is.
- #2 info/freetext: buildInfoRequest exists in builders.js but is NOT wired into
tht-gate.js (imports only select/multiselect/artifact-gate). info notices are
hand-written ctx.ui.notify; freetext is a control, not a descriptor. Table now
lists only the 3 real present() descriptors.
- #3 cite the REWRITE banner (tht-gate.js:8-10), not :227; id-match quoted verbatim.
- #4 new "Note implementative": 227 comment, 4/6 unguarded notify, and the
ctx.hasUI guard migration side-effect (now fires in RPC on 0.80.3).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Design doc for three coordinated harness fixes derived from the
2026-06-30-165708 session analysis: correct SKILL.md phase-advance
contract, fix the F6 cte_approved subject bug, and add a
tht-mediated schema_linking.json writer/validator.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tier 1 (clean-room first-turn harness, harness/scripts/model-matrix.mjs):
kickoff + resume chain in-turn on ALL available models — zai/glm-5.2,
deepseek/deepseek-v4-{pro,flash}, aritmolab/qwen3.6-35b-a3b, zai/glm-4.5-air.
The resume cold-start stall recurs on none (closes A's cross-model robustness).
aritmolab/gemma4-26b-a4b is a 404 at the endpoint (listed but not served) — an
availability gap classified as MODEL_ERROR, not a workflow issue.
Tier 2 (live, baseline zai/glm-5.2): F single-select auto-confirm verified
end-to-end — answering the first reviewer_select persisted a concept_clarified
decision (review_decisions.jsonl 0->1) with no follow-up confirmation gate.
Closes F's deferred live check.
No prompt hardening needed. Results in the G plan doc + memory. Throwaway psd
sessions used and deleted; real sessions untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Actionable scoping for the final workstream: a two-tier method (clean-room
first-turn harness generalised from the A2 repro-driver for cheap kickoff/resume
signals; sparing full Playwright F1 runs for the gate round-trip) over the
configured models (GLM 5.2 baseline, Deepseek V4, Qwen3.6, + breadth). Folds in
F's deferred single-select live check and A's resume-on-weaker-models check.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Active/Archive accordions; rename group (client-side reassign); central
area shows only last user entry + gate notify/info + active widget; the
verbose model stream moves to a left on-demand panel toggled by the WIP
icon (moved above the composer). Frontend-only; fine thinking-separation
deferred.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Hard-fail preflight at session create/restart: ensure Ollama up + warm the
configured embedding model, refuse the session if embeddings unavailable.
Parameterized ollama bin/start_cmd; tht ollama ensure command + backend 503.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the SKILL.md design contract (persisted state is the truth) and a
dedicated Resume correctness section: backend new/resume prompt mode,
missing cold-start procedure in the skill, end-to-end verification gate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Decisioni FE-1..FE-6: Vite+React SPA (no Next), TanStack Query+Zustand+hook SSE,
EventSource nativo (MVP auth=none), widget registry+fallback, test Vitest+RTL+MSW
+ Playwright e2e F1, build a slice con F1 come primo loop chiuso.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Lo spike Task 1 ha provato che ctx.sendRaw non esiste e pi.on(extension_ui_response)
non e' dispatchato. Aggiornati: Piano1 Task2 (mock ctx.ui), Task4 (rewrite gate a
ctx.ui.input con descriptor in title), Task10 (fake-pi-rpc shape nativa); Piano2
Task3/Task4 (SessionBridge decodifica title<->value). Contratto FE invariato.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Due piani separati (decomposizione concordata): il backend dipende da lavoro
harness non testato (gate RPC-ready, id injection, prereq CLI --json/--offset,
fake-Pi). Piano 1 rende l'harness pilotabile via RPC; Piano 2 costruisce il
backend Node/Fastify testato contro il fake-pi-rpc.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Decisioni BE-1..BE-7: un Pi per sessione attiva, SQL finale delegato a tht
(codepath unico, rischio D7 eliminato), resilienza via ricostruzione da disco +
re-emit del widget pendente, test con fake-Pi condiviso, backend pre-crea la
sessione, model/thinking/provider per-sessione persistiti, settings Pi MVP.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Due cambiamenti interconnessi da user review:
1. language come parametro workspace (spec decisione 9):
- Config.language (default 'en') + workspaces PSD con 'language: it'
- Generalizza Thoth oltre l'italiano: descrizioni tabelle/colonne ed evidence
sono nel workspace language; le istruzioni della skill restano in inglese
(piu' affidabili per modelli piccoli, meno ambigue)
2. Skill riscritta in INGLESE preservando la semantica COMPLETA dell'originale
(autocritica: la mia riscrittura precedente aveva perso ~10 vincoli precisi):
- 'promuovere' ambiguo (3 accezioni: phase advance / recommend / memory promote)
-> 'never advance a phase or record a decision without confirmation'
- recuperati vincoli persi: choice-is-confirmation (no reviewer_confirm dopo
reviewer_decide), reviewer_select SOLO per iterazione no-decision, messaggi
auto-contenuti obbligatori, artefatto = superficie di decisione (gate rilegge
da disco per CTE/SQL), candidati con provenienza+score non verita', opzione
'leave ambiguity open', F1 passa lista completa non solo ultima
- language contract esplicito (istruzioni EN, output nel workspace language)
Sottomoduli cte/memoria/rewriting/sql-generation in inglese, semantica tecnica
intatta (regole AV-SQL, dim_time trick, max 5 memorie solo 3 tipi riusabili).
Verifica: 0 residui nsp/chirone, tutti i tht <cmd> citati registrati, 165 passed.
- Task 3.2: nuovo Step 1b — drop funzioni registry da memory_cmd + tht/memory.py
(load/save/update/delete/promote/reusable_promotions); promotion (F5) riscritta
come upsert batch al vectordb; memory list/delete su vectordb (no registry).
- Skill F2/F5: drop riferimenti a 'memory promote + memory index' con registry.
- Onda 0b riscritta: repo workspace per-cliente separato (no copia in harness/),
LSH scarica-tutti-i-valori-distinti (no 'campiona'), indice nel repo workspace cliente.
Tre correzioni da user review:
5. Memory SOLO pgvector, niente registry. Il registry.jsonl di ChironeWp3 è vestigiale:
una volta che il metadata del vectordb ha subject/detail/rationale (decisione 6), il
registry non serve. Promotion (F5) = upsert diretto a vectordb. tht/memory.py non
porta le 6 funzioni registry. 'Cancellare' = metadata.status='superseded' (audit
trail; il writer e' upsert-only). Multi-workstation OK per costruzione.
7. Evidence: repo workspace separato per-cliente, NON dentro ThothII. ThothII e'
generico; un repo tht-workspace-<cliente>/ contiene evidence/ + workspace YAML +
indici LSH. Deploy = checkout ThothII + checkout workspace-cliente. Niente copia
in harness/.
8. LSH: scarica TUTTI i valori distinti (non 'campiona'), costruisce MinHash+LSH
dentro harness come preprocessing. Indice per-cliente (nel repo workspace cliente).
Sezione 3 punto 2 (F2) riscritta: save-one diretto, drop reference a promote/index
con registry. Arricchimento metadata ora 'obbligatorio' (non opzionale): senza
registry, il vectordb e' l'unica fonte. Residui nsp nei path spec corretti a tht.