Files
ThothII/harness/.pi/skills/tht-evidence-authoring/SKILL.md
T
Codex 82e2c91f42
Publish documentation / publish (push) Successful in 1m27s
feat: implement memory and evidence administration with guided repairs
Add PostgreSQL-backed memory, editable evidence with source review and activation, and human-approved archive repairs across the harness, API, and UI. Include migrations, deployment support, regression coverage, and validation documentation.

Refresh permissions from validated session roles so existing administrator logins can access newly deployed archive management features.
2026-09-10 10:31:34 +02:00

5.2 KiB

name, description
name description
tht-evidence-authoring Restructure exactly one normalized Thoth Source Evidence request into strict, typed Evidence candidate JSON for deterministic host-side review.

Evidence authoring response contract

You receive exactly one normalized Source Evidence request. Return one JSON object with only a candidates array, without Markdown fences, comments, or explanatory text. The host strictly rejects unknown or missing fields.

Candidate JSON uses response protocol version 1. The host renders accepted candidates as editable Curated Evidence v4; supply the typed payload and exact excerpts in this response, leaving file rendering and provenance metadata to the host.

Use only facts present in normalized_text. Never merge, cite, or infer facts from another source. Reuse an existing_id only when it was supplied in previous_units; otherwise omit it. Preserve prior reviewed wording when it is still supported. When a prior unit is no longer supported, omit its existing_id and add a source_no_longer_supports_unit review item to the related candidate when applicable.

Each candidate has exactly this shape:

{
  "schema_version": 1,
  "existing_id": "evidence:optional-existing-id",
  "title": "Short human title",
  "kind": "glossary",
  "purposes": ["disambiguation"],
  "applies_to": {
    "concepts": ["concept"],
    "tables": ["schema.table"],
    "columns": ["schema.table.column"]
  },
  "language": "it",
  "supporting_excerpts": ["One short exact excerpt copied from normalized_text."],
  "review_items": [
    {"code": "specific_ambiguity", "message": "What a human must decide.", "field": "payload"}
  ],
  "payload": {"definition": "Typed payload described below.", "synonyms": [], "variants": []}
}

Omit existing_id for new candidates. applies_to must contain only concepts, tables, and columns; use empty arrays when the source does not establish a value. Every table identifier must be schema.table, and every column identifier must be schema.table.column. Include a table or column identifier only when that fully qualified literal already appears in normalized_text. Never qualify an unqualified name yourself. If the source contains only names such as fact_ilr or cod_paz, leave the corresponding tables or columns array empty, keep the names in prose, and add a review item when qualification matters. review_items may be empty. A review item contains only code, message, and optional field.

Allowed purposes are disambiguation, rewriting, schema_linking, and sql_generation. Allowed kind values and their exact payload shapes are:

  • glossary: {"definition": string, "synonyms": [string], "variants": [string]}
  • domain: {"rule": string}
  • enum: {"column": "schema.table.column", "values": {"stored value": "meaning"}}
  • example: {"question": string, "interpretation": string}
  • mapping: {"concept": string, "tables": ["schema.table"], "columns": ["schema.table.column"]}
  • normalization: {"input": string, "output": string, "rule": string}
  • formula: {"concept": string, "columns": ["schema.table.column"], "sql": "one PostgreSQL expression"}
  • reference: {"url": "https://...", "label": string, "description": string}

Return exactly one candidate: the source's primary independent, reviewable Evidence Unit. Preserve the source's secondary facts in that unit's typed rule, definition, or interpretation instead of emitting extra candidates; do not atomize individual sentences. The candidate must include one to five nonempty exact supporting_excerpts, each at most 1000 characters. Copy each excerpt as one continuous, character-for-character substring of normalized_text, including its original Markdown punctuation. Prefer copying one complete source line. Never paraphrase, normalize whitespace, remove backticks, or change quotation marks inside an excerpt.

Before returning JSON, check every excerpt with the equivalent of excerpt in normalized_text; replace any excerpt that would fail with an exact complete line from the source. Also check that there is exactly one candidate, every object has only the declared fields, every kind has the exact payload shape above, and stdout contains only the JSON object. Add a review item whenever the source leaves a material ambiguity; never silently guess a table, column, enum meaning, formula, or URL.

Use the source path as a kind hint: 00-glossario normally yields glossary or domain; 10-domini-clinici normally yields domain; 20-valori-enum normally yields enum; 30-esempi-nlq normally yields example; 40-mapping-semantico normally yields mapping or domain; and 50-metadati-normalizzazione normally yields normalization. Depart from the hinted kind only when the source explicitly provides the complete typed payload for another kind. Emit formula only when the source states one complete PostgreSQL expression and all referenced columns are fully qualified. If an enum, mapping, or formula payload would require an identifier that is not already fully qualified in the source, emit a domain candidate instead and record the missing qualification as a review item.

Do not assign a new canonical ID; the host does that deterministically. Do not use tools or alter any repository state.