feat(harness): port backend Onda 0 — vendor, lshindex, sqlcheck, execute, rest/exec, ctetest, report, datamart

8 moduli leaf portati verbatim da ChironeWp3 con rename psdwp3→tht:
- vendor/thoth_lsh (MinHash/LSH, leaf puro datasketch+tqdm) + VENDORED.md
- lshindex/ (build/save/load/query, dipende vendor + LshConfig)
- sqlcheck/ (validate_sql, leaf ExecutionConfig+mschema)
- execute/ + execute/warnings (run_controlled/explain, leaf sqlglot+sqlalchemy)
- rest/execute + rest/explain (REST variants, dipendono execute+rest.client)
- ctetest (CTE test records, leaf sqlglot+pydantic)
- report (validation report rendering, dipende execute+sqlcheck)
- datamart (stub NotImplementedError)

Verifica: import smoke catena completa OK, pytest 109 passed. Deps (datasketch, sqlglot,
sqlalchemy, pydantic, requests, tqdm) già in pyproject. VENDORED.md neutralizzato
(riferimenti PsdWp3→Thoth).
This commit is contained in:
2026-06-27 10:34:23 +02:00
parent fc5fbe6b65
commit ea6412fafc
12 changed files with 717 additions and 0 deletions
+30
View File
@@ -0,0 +1,30 @@
# Codice vendorizzato
## thoth_lsh.py
- Origine: `thoth_sqldb2` (pacchetto PyPI `thoth-dbmanager` 0.7.4), file `thoth_dbmanager/lsh/core.py`
- Copyright 2025 Marco Pancotti — Apache License 2.0
- Modifiche locali:
1. rimosso blocco di debug su colonna "doctype" in `create_lsh_index`;
2. `skip_column`: soglie parametrizzate (`max_total_chars`, `max_avg_length`) con default
identici agli originali (50000 / 20);
3. rimossa `query_lsh_index` (sostituita da `tht.lshindex.query_index`, che restituisce
anche gli score per il ranking spiegabile);
4. try/except generico rimosso da `create_lsh_index`: gli errori devono emergere.
5. `skip_column`: la guardia "name-like" (originariamente solo `"name"`) è stata estesa
a token bilingui EN/IT tramite la costante `NAME_LIKE_TOKENS`
(`name`, `nome`, `nominativ`, `denominazion`, `ragione_sociale`), sovrascrivibile col
parametro `name_tokens`. Necessario perché il datawarehouse Chirone è in italiano e le
colonne `nome`/`cognome`/`denominazione`/`ragione_sociale` non venivano riconosciute
dalla sola parola inglese `name`.
**Nota: `skip_column` e `NAME_LIKE_TOKENS` non sono più usati dal flusso Thoth.**
La selezione delle colonne da indicizzare è ora governata dal *principio di eleggibilità
delle colonne* (spec: `docs/superpowers/specs/2026-06-13-tht-column-eligibility-principle.md`).
L'esclusione dei testi larghi avviene a monte tramite il flag `eligible` persistito in
`physical.yaml` (impostato da `nsp schema introspect`), non tramite l'euristica sulla
lunghezza del vendorizzato. Le funzioni restano nel file per fedeltà alla sorgente upstream;
Thoth ha semplicemente smesso di importarle (vedi `tht/db/sampling.py`).
Le query di introspezione in `tht/db/introspect.py` sono ADATTATE (non copiate verbatim)
da `thoth_dbmanager/adapters/postgresql.py` della stessa versione.
View File
+82
View File
@@ -0,0 +1,82 @@
# Vendored from thoth_sqldb2 (thoth-dbmanager 0.7.4) — lsh/core.py
# Copyright 2025 Marco Pancotti — Apache License 2.0
# Modifiche locali documentate in VENDORED.md.
"""Core LSH (MinHash) per la ricerca di valori simili nei campi del database."""
import logging
from typing import Dict, List, Tuple
from datasketch import MinHash, MinHashLSH
from tqdm import tqdm
def create_minhash(signature_size: int, string: str, n_gram: int) -> MinHash:
m = MinHash(num_perm=signature_size)
for d in [string[i : i + n_gram] for i in range(len(string) - n_gram + 1)]:
m.update(d.encode("utf8"))
return m
# Token che, se presenti nel nome di una colonna, la marcano come "name-like":
# nomi propri di persone/enti, i cui valori sono utili al value-matching e quindi
# mai da escludere dall'indice LSH. Bilingue EN/IT (modifica locale, vedi VENDORED.md).
# Per substring coprono le forme flesse: "nome" -> cognome/soprannome,
# "name" -> surname/username, "denominazion" -> denominazione, ecc.
NAME_LIKE_TOKENS: tuple[str, ...] = (
"name",
"nome",
"nominativ",
"denominazion",
"ragione_sociale",
)
def skip_column(
column_name: str,
column_values: List[str],
max_total_chars: int = 50000,
max_avg_length: int = 20,
name_tokens: tuple[str, ...] = NAME_LIKE_TOKENS,
) -> bool:
lowered = column_name.lower()
if any(token in lowered for token in name_tokens):
return False
sum_of_lengths = sum(len(value) for value in column_values)
average_length = sum_of_lengths / len(column_values)
return (sum_of_lengths > max_total_chars) and (average_length > max_avg_length)
def jaccard_similarity(m1: MinHash, m2: MinHash) -> float:
return m1.jaccard(m2)
def create_lsh_index(
unique_values: Dict[str, Dict[str, List[str]]],
signature_size: int,
n_gram: int,
threshold: float,
verbose: bool = True,
) -> Tuple[MinHashLSH, Dict[str, Tuple[MinHash, str, str, str]]]:
lsh = MinHashLSH(threshold=threshold, num_perm=signature_size)
minhashes: Dict[str, Tuple[MinHash, str, str, str]] = {}
total = sum(
len(column_values)
for table_values in unique_values.values()
for column_values in table_values.values()
)
logging.info("Total unique values: %s", total)
progress_bar = tqdm(total=total, desc="Creating LSH") if verbose else None
for table_name, table_values in unique_values.items():
for column_name, column_values in table_values.items():
for idx, value in enumerate(column_values):
minhash = create_minhash(signature_size, value, n_gram)
minhash_key = f"{table_name}_{column_name}_{idx}"
minhashes[minhash_key] = (minhash, table_name, column_name, value)
lsh.insert(minhash_key, minhash)
if progress_bar:
progress_bar.update(1)
if progress_bar:
progress_bar.close()
return lsh, minhashes