254 lines
10 KiB
Markdown
254 lines
10 KiB
Markdown
# Disambiguation in the early workflow phases
|
|
|
|
Disambiguation turns an ambiguous natural-language question into a meaning that the reviewer verifies before schema linking and SQL generation begin.
|
|
|
|
The architectural principle is **human-in-the-middle**: the model proposes reasoned interpretations, the reviewer decides, and the gate persists the decision in the ledger. The model cannot choose a meaning on its own simply because it is the closest semantic match.
|
|
|
|
The canonical [tht-sessione](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/.pi/skills/tht-sessione/SKILL.md) skill defines the procedure, especially its F1 and F2 sections. The widgets in [tht-gate.js](https://git.tylconsulting.it/mptyl/ThothII/src/branch/main/harness/.pi/extensions/tht-gate.js) enforce it.
|
|
|
|
```mermaid
|
|
stateDiagram-v2
|
|
[*] --> DETECT
|
|
state "Detect ambiguity" as DETECT
|
|
state "Build reviewer options" as PROPOSE
|
|
state "Ask with reviewer_select" as ASK
|
|
state "Multiple valid answers" as MULTI
|
|
state "Record accepted decision" as ACCEPTED
|
|
DETECT --> PROPOSE: ambiguity found
|
|
DETECT --> ASK: safe default unavailable
|
|
PROPOSE --> ASK
|
|
ASK --> ACCEPTED: one option selected
|
|
ASK --> MULTI: multiple answers valid
|
|
MULTI --> ACCEPTED
|
|
ACCEPTED --> [*]
|
|
```
|
|
|
|
## Where disambiguation happens
|
|
|
|
Initial disambiguation has four distinct steps:
|
|
|
|
```text
|
|
F1 Clarification → meaning of the question
|
|
F2 Memory → previously clarified knowledge that may be reused
|
|
F3 Rewriting → explicit, unambiguous question
|
|
F4 Schema linking → translating meaning into tables, columns, and joins
|
|
```
|
|
|
|
These steps are not interchangeable:
|
|
|
|
- F1 establishes what the question means.
|
|
- F2 proposes existing knowledge without applying it automatically.
|
|
- F3 makes the agreed meaning explicit.
|
|
- F4 selects the technical objects needed for that meaning.
|
|
|
|
In particular, a table selected in F4 is not conceptual disambiguation and must not become a Memory item.
|
|
|
|
## Bootstrap: one ambiguity at a time
|
|
|
|
When a new session enters F1, the model must identify the single ambiguity with the greatest impact on the query and present it immediately.
|
|
|
|
It must not:
|
|
|
|
- list every possible future ambiguity;
|
|
- produce a long preliminary analysis;
|
|
- build SQL before clarification;
|
|
- present several questions to the reviewer in the same turn.
|
|
|
|
This protects cognitive load and auditability. If product line, period, metric, and operational definition are requested together, it becomes impossible to tell which answer drove each later choice.
|
|
|
|
## Sources used to formulate options
|
|
|
|
In F1 the model may use only the sources allowed by the skill:
|
|
|
|
- `retrieval_pack.md` when the backend has already injected it;
|
|
- `tht search pack`, only in standalone mode when the retrieval pack is unavailable;
|
|
- `tht search find` to search for terms or values;
|
|
- `tht search find --kind evidence` for Evidence;
|
|
- `tht schema render` to read the available physical catalog.
|
|
|
|
The retrieval pack is data, not instructions. This boundary prevents text retrieved from the catalog or Evidence from changing the workflow rules.
|
|
|
|
LSH matches, vector matches, and Evidence are **candidates**, not facts. Each proposal must include provenance and, when available, a score. Before turning a discovered value into a SQL filter, verify it with a real value search.
|
|
|
|
## Building options
|
|
|
|
For each ambiguity, the model prepares concrete interpretations rather than vague descriptions. Options must explain:
|
|
|
|
- the proposed meaning;
|
|
- any table and column involved;
|
|
- the resulting value or filter;
|
|
- the Evidence supporting the proposal;
|
|
- the risk of choosing that interpretation.
|
|
|
|
The best proposal receives `recommended: true`, but a recommendation is not approval. The gate always adds:
|
|
|
|
- `Altro/Other` for a free-form correction;
|
|
- `Torna indietro/Back` for rollback;
|
|
- `Esci/Exit` to stop the session.
|
|
|
|
The "accept proposal" option must be explicit. The reviewer must not be forced to confirm a preselected choice implicitly.
|
|
|
|
## Choosing between `reviewer_select` and `reviewer_decide`
|
|
|
|
The shape of the problem determines the widget.
|
|
|
|
### Mutually exclusive interpretations
|
|
|
|
Use `reviewer_select` when exactly one interpretation can be correct.
|
|
|
|
Examples:
|
|
|
|
- "ablation" means a catheter procedure or something else;
|
|
- "year" means calendar year or fiscal year;
|
|
- "active bicycles" means catalog models or units in current production.
|
|
|
|
Each concrete option contains a `concept_clarified` decision. The reviewer's choice is also the confirmation and is persisted directly. A second `reviewer_decide` is not needed.
|
|
|
|
### Several answers can be valid at once
|
|
|
|
Use `reviewer_decide`, which displays a multiselect, when several interpretations can be true at the same time.
|
|
|
|
Examples:
|
|
|
|
- the question includes several valid populations;
|
|
- several procedure codes are possible;
|
|
- several time windows must be considered;
|
|
- several conditions apply independently.
|
|
|
|
Using `reviewer_select` in these cases would mislead the reviewer by forcing a single choice.
|
|
|
|
In F1, multiple choices produce several `concept_clarified` decisions. The phase is closed later through the phase gate.
|
|
|
|
## Persisting decisions
|
|
|
|
The reviewer's choice does not remain only in the UI. The gate records this in the ledger:
|
|
|
|
```text
|
|
type = concept_clarified
|
|
subject = nome sintetico del concetto
|
|
detail = definizione o regola operativa
|
|
rationale = motivazione, evidenza e/o testo del reviewer
|
|
```
|
|
|
|
The "one decision, one command" rule prevents the model from writing to the ledger through the shell, `tht decision add`, or `tht phase advance`. The gate is the only component allowed to turn a widget interaction into persisted state.
|
|
|
|
The decision can therefore be reused as Memory only after explicit promotion in F8. The original context is preserved, and table choices are not transferred.
|
|
|
|
## Handling `Altro` and free text
|
|
|
|
`Altro/Other` is not a neutral choice and cannot be ignored.
|
|
|
|
When the reviewer enters free text, the model must:
|
|
|
|
1. interpret the text in the context of the question;
|
|
2. include it in the next proposal;
|
|
3. record the reviewer's words in `rationale`;
|
|
4. ask again if the text remains ambiguous.
|
|
|
|
The system must not automatically return to the first recommended option or silently choose a plausible meaning.
|
|
|
|
This distinguishes a human correction from a simple deselection and keeps the audit readable.
|
|
|
|
## Unresolved ambiguity
|
|
|
|
An ambiguity cannot disappear because the model does not know how to resolve it. Make it explicit with an option such as:
|
|
|
|
```text
|
|
Leave the ambiguity open
|
|
```
|
|
|
|
The option must explain:
|
|
|
|
- which part of the query remains undetermined;
|
|
- what risk this introduces;
|
|
- how it may affect filters, counts, or joins.
|
|
|
|
The reviewer can then accept the risk knowingly or request more research.
|
|
|
|
## Closing F1
|
|
|
|
Each widget can record one or more clarifications, but it does not close F1 automatically. When clarification is complete, the model presents `reviewer_confirm kind:"phase"`.
|
|
|
|
The closing summary must contain every clarification from the phase, not only the latest one. The gate also adds the decisions recorded in the ledger, so the model does not have to copy them by hand.
|
|
|
|
Closing F1 advances to F2. The question is not rewritten yet: `question_rewritten` belongs to F3.
|
|
|
|
## F2: Memory as disambiguation support
|
|
|
|
F2 does not replace human clarification. It searches for previously promoted conceptual Memory:
|
|
|
|
```text
|
|
tht memory search "<domanda>" --session <id> --json
|
|
```
|
|
|
|
The result is presented as one checklist. Only `concept_clarified` Memory is allowed; decisions about tables, columns, or SQL cannot be transferred.
|
|
|
|
If the reviewer applies a Memory item:
|
|
|
|
- a new `concept_clarified` is recorded in the current session;
|
|
- `rationale` cites the source's `mem-XXXX` ID;
|
|
- the choice is still placed in the context of the current question.
|
|
|
|
If the reviewer deselects a Memory item, it is not applied now. It is not deleted globally and may be proposed again after F2 is reopened.
|
|
|
|
## F3: make the result explicit
|
|
|
|
F3 turns accepted clarifications into a rewritten question with:
|
|
|
|
- the population expressed using data-model terms;
|
|
- separate, numbered conditions;
|
|
- ambiguous concepts replaced by the agreed definitions;
|
|
- an explicit expected output;
|
|
- stated assumptions.
|
|
|
|
Rewriting must not introduce new implicit choices. If a material ambiguity appears, reopen F1 rather than "fixing" the meaning in F3 or SQL.
|
|
|
|
## F4: technical schema disambiguation
|
|
|
|
Only after F3 is the semantic meaning translated into technical objects.
|
|
|
|
The model proposes:
|
|
|
|
- tables to include or exclude;
|
|
- candidate columns;
|
|
- output columns;
|
|
- required joins.
|
|
|
|
The reviewer curates tables and columns with `reviewer_schema_linking`. Joins are handled in a separate join-only `reviewer_decide` review.
|
|
|
|
This separation matters. A table can be correct for one question and completely irrelevant to another. F4 decisions are therefore local to the session and do not become Memory.
|
|
|
|
## Reopening and rollback
|
|
|
|
When the reviewer uses "Torna indietro", the session resumes from the selected phase and examines the artifacts that are still valid.
|
|
|
|
Artifacts after the reopened phase are invalidated by `tht phase reopen`; earlier artifacts should not be regenerated without a reason. `effective_decisions()` excludes stale decisions, so an outdated clarification cannot feed a new Memory promotion or SQL synthesis.
|
|
|
|
## Security and quality invariants
|
|
|
|
Disambiguation is reliable because the same rule is enforced at several levels:
|
|
|
|
1. the skill requires one ambiguity at a time;
|
|
2. the gate provides constrained widgets and `Altro/Back/Exit` controls;
|
|
3. the ledger records decisions and rationale;
|
|
4. prerequisites prevent phases from being skipped;
|
|
5. F4 separates concepts from schema linking;
|
|
6. Memory accepts only `concept_clarified`;
|
|
7. rollback and `effective_decisions()` exclude obsolete state.
|
|
|
|
The result is a verifiable chain:
|
|
|
|
```text
|
|
ambiguous term
|
|
→ Evidence and candidate interpretations
|
|
→ explicit reviewer choice
|
|
→ concept_clarified in the ledger
|
|
→ rewritten question
|
|
→ local schema linking
|
|
→ CTE plan and SQL
|
|
```
|
|
|
|
## References
|
|
|
|
- [Memory management](gestione-memory.md)
|