Ports search/__init__.py (combined_search/RRF/aggregate) renamed psdwp3->nsp. New aggregate_lsh_multi (the D14a deviation): groups LSH hits by table keeping EVERY column where a value appears -- NOT collapsed to a single best column. The old _aggregate_lsh hid alternative groundings (e.g. 'ablazione' matching both a boolean flag and a free-text patologia field). aggregate_lsh_multi exposes all columns so the value-grounding widget lets the reviewer choose the anchor(s). Within one (table, column) the best-scored value is kept; columns ordered by score. value_grounded added to DecisionType (records the reviewer's anchor choice). L1: test_value_grounding (6 tests) -- multi-column exposure, grouping, within-column best-value, ordering, empty, and the value_grounded decision-type existence. Deferred: lshindex/ (needs vendor/thoth_lsh) and the L0 test_rrf.py land with the nsp lsh build command + index-building path; not needed for the pure L1 core here.
71 lines
2.7 KiB
Python
71 lines
2.7 KiB
Python
"""L1: value grounding -- multi-column LSH exposure (spec D14a).
|
|
|
|
The D14a deviation: a value cited in the question (e.g. 'ablazione') may match
|
|
MULTIPLE columns (a boolean flag, a free-text patologia field). The old behavior
|
|
collapsed matches to a single best column, hiding the alternative grounding.
|
|
aggregate_lsh_multi exposes every column where the value appears, grouped by table,
|
|
so the value-grounding widget can let the reviewer choose which column(s) anchor
|
|
the value.
|
|
"""
|
|
from nsp.search import aggregate_lsh_multi
|
|
|
|
|
|
def test_value_in_multiple_columns_returns_all():
|
|
hits = [
|
|
{"table": "t", "column": "c1", "value": "ablazione", "score": 0.9},
|
|
{"table": "t", "column": "c2", "value": "ablazione", "score": 0.7},
|
|
]
|
|
result = aggregate_lsh_multi(hits)
|
|
# NOT collapsed to single best -- both columns exposed
|
|
cols = {h["column"] for h in result["t"]}
|
|
assert cols == {"c1", "c2"}
|
|
|
|
|
|
def test_values_grouped_by_table():
|
|
hits = [
|
|
{"table": "pazienti", "column": "flag_abl", "value": "ablazione", "score": 0.9},
|
|
{"table": "ricoveri", "column": "procedura", "value": "ablazione", "score": 0.6},
|
|
]
|
|
result = aggregate_lsh_multi(hits)
|
|
assert set(result.keys()) == {"pazienti", "ricoveri"}
|
|
assert result["pazienti"][0]["column"] == "flag_abl"
|
|
assert result["ricoveri"][0]["column"] == "procedura"
|
|
|
|
|
|
def test_within_column_keeps_best_value():
|
|
# two hits on the SAME column: keep the best-scored value (no duplicate rows
|
|
# for one column), but the column still appears once.
|
|
hits = [
|
|
{"table": "t", "column": "c", "value": "ablazione", "score": 0.9},
|
|
{"table": "t", "column": "c", "value": "ablaz", "score": 0.5},
|
|
]
|
|
result = aggregate_lsh_multi(hits)
|
|
rows = result["t"]
|
|
assert len(rows) == 1
|
|
assert rows[0]["value"] == "ablazione" # best score kept
|
|
assert rows[0]["score"] == 0.9
|
|
|
|
|
|
def test_columns_ordered_by_score_desc_within_table():
|
|
hits = [
|
|
{"table": "t", "column": "low", "value": "x", "score": 0.3},
|
|
{"table": "t", "column": "high", "value": "x", "score": 0.95},
|
|
{"table": "t", "column": "mid", "value": "x", "score": 0.6},
|
|
]
|
|
result = aggregate_lsh_multi(hits)
|
|
cols = [h["column"] for h in result["t"]]
|
|
assert cols == ["high", "mid", "low"]
|
|
|
|
|
|
def test_empty_hits_returns_empty():
|
|
assert aggregate_lsh_multi([]) == {}
|
|
|
|
|
|
def test_value_grounded_decision_type_exists():
|
|
# D14a adds the value_grounded decision type so the gate can record the
|
|
# reviewer's choice of which column(s) anchor a cited value.
|
|
from nsp.decisions import DecisionType
|
|
import typing
|
|
args = typing.get_args(DecisionType)
|
|
assert "value_grounded" in args
|