Files
ThothII/presentation/slides/02-flow.md
T
Codex 0f1391e1b7 feat: AritmoLab AI presentation deck (work in progress)
Reveal.js deck + prototypes + speaker scripts for the talk
'Role of AI in the Analysis of Unstructured Clinical Databases'
(AritmoLab — Policlinico San Donato, 2-3 October 2026).

Lives on its own branch while in progress; not for main until ready.
2026-09-06 23:06:28 +02:00

112 lines
5.4 KiB
Markdown

# Slide 02 — The project at a glance (project flow diagram)
First slide of merit in ALL three content proposals (A/B/C). It absorbs the old
"platform at a glance" slide. Presenter: MP.
## On-slide (EN)
- Kicker: `What existed → what we built`
- Title: **The project at a glance**
- Diagram, two zones left→right:
- **What existed** (muted, dashed): Cardioref (cardiology records · procedures ·
letters) · Genetic data (labs · variants · nomenclature) · ECG (signals ·
device follow-up) · **Future sources** (dashed, "+" badge, muted label —
new subsystems can be connected as sources; label proposed 2026-09-06 over
"Next"/"Futures"/"Others")
- **What we built** (tinted): Staging (raw replica) → Integration (cleaned,
normalized, deduplicated) → Star schema (dimensional transformation) →
Datamarts (research-ready marts) → two endpoints in bordeaux:
**AritmoLab Data Warehouse** (queried for research) + **AritmoLab Portal**
(management & exploration)
- 🧠 badges with one-line captions:
- Integration: "AI reads the clinical text: pathologies, procedures,
drug-challenge outcomes"
- Datamarts: "AI builds them on demand from plain-English questions (ThothII)"
- On the existed→built arrow: "The mappings" (AI-drafted ingestion configs)
- Popup behaviour (Reveal deck): the contribution card opens NEAR the clicked
brain (not centered) and a red line ties the card to that brain; closing with
✕, click outside or Esc.
- Footnote 🧠: "Also at build time: the pipeline itself was co-engineered with
AI — mapping configs drafted by LLM agents, reviewed by humans, versioned in git."
## AI contribution popup texts (formatted draft 2026-09-06 — to discuss)
Structure per popup: lead line + bullets (bold label — detail; numbers
highlighted) + **"Positive effects" box** (bordered, always last — the payoff of
that AI use must stand out). Shown when clicking a 🧠 in the prototype (in
PPTX: click-to-appear boxes).
### 1. Ingestion — "The mappings" (brain on the existed→built arrow)
Lead: The mappings that turn three raw sources into clean data are
**crafted by AI** — **revised by humans**.
- **AI drafts** — LLM agents write the mappings from prompts and schema samples
- **What is a mapping?** — a plain instruction sheet that tells the system,
field by field, where each piece of data comes from and where it has to land
- **At scale** — 19,764 lines of instructions, all machine-validated
- **Humans revise** — every line is checked, then recorded in **git**, the
archive that keeps every version of the instructions and can undo any change
**Positive effects box:** Weeks of hand-writing became days of reviewing.
### 2. Integration — "Reading the clinical text"
Lead: A deterministic, bilingual (IT/EN) pattern library reads the free text
inside the nightly pipeline.
- **Pathologies** — 73,389 extracted from 58,438 discharge letters, into a
two-tier clinical ontology
- **Procedures ↔ pathologies** — linked at clause level
- **Drug-challenge tests** — 10,908 parsed (flecainide, ajmaline, adrenaline,
isoprenaline)
- **No black box** — semver rules, `pattern_version` on every row, ≥95% accuracy
gate on 100 manual reviews
**Positive effects box:** 58,438 letters of free text became an analysable research asset —
automatically, every night.
### 3. Datamarts — "Datamarts from plain English"
Lead: Researchers ask in plain English; the AI writes the SQL over the star schema.
- **Ask** — "how many Brugada patients had an effective ablation?"
- **Generate** — AI builds the SQL and assembles a curated datamart
- **Serve** — results flow to Superset for statistics and machine learning
- **In the loop** — the researcher reviews every proposed query before it runs
**Positive effects box:** A new research datamart in minutes instead of weeks of hand-written
SQL — with the researcher approving every query.
### OPEN — genetics (brain not placed until a/b decision)
(a) deterministic → keep no brain on Genetic data.
(b) AI upstream → brain caption candidate: "AI reconciles genetic records
upstream: variant nomenclature and lab reports normalized before ingestion"
(only if MP confirms this happens outside the ETL).
## Speaker script (EN, ~75s)
Here is the whole project on one slide. On the left, what already existed: three
hospital data sources — Cardioref, our electrophysiology records; the genetic
data coming from the labs; and the ECG signals. On the right, what we built in
AritmoLab: a staging replica, an integration layer where the data is cleaned and
normalized, a star-schema warehouse, and the research datamarts — delivered
through the data warehouse and a management portal.
Everywhere you see this brain symbol, that's where AI works for us. At
integration, AI reads the clinical text — pathologies, procedures, and
drug-challenge outcomes. At the very end, AI builds the datamarts on demand,
from plain-English questions. And even this pipeline was partly built with AI:
the mapping configurations were drafted by AI agents, then reviewed by humans
and versioned in git.
My colleague Dr. Paratico will show you exactly how the text reading works. But
first, let me show you why this is hard.
## Notes
- Sub-labels of the three sources are placeholders — MP to confirm the real
wording (esp. ECG: "signals · device follow-up").
- If the genetics a/b decision lands on (b), add a 🧠 on the Genetic data box
with the agreed caption.