Reveal.js deck + prototypes + speaker scripts for the talk 'Role of AI in the Analysis of Unstructured Clinical Databases' (AritmoLab — Policlinico San Donato, 2-3 October 2026). Lives on its own branch while in progress; not for main until ready.
112 lines
5.4 KiB
Markdown
112 lines
5.4 KiB
Markdown
# Slide 02 — The project at a glance (project flow diagram)
|
|
|
|
First slide of merit in ALL three content proposals (A/B/C). It absorbs the old
|
|
"platform at a glance" slide. Presenter: MP.
|
|
|
|
## On-slide (EN)
|
|
|
|
- Kicker: `What existed → what we built`
|
|
- Title: **The project at a glance**
|
|
- Diagram, two zones left→right:
|
|
- **What existed** (muted, dashed): Cardioref (cardiology records · procedures ·
|
|
letters) · Genetic data (labs · variants · nomenclature) · ECG (signals ·
|
|
device follow-up) · **Future sources** (dashed, "+" badge, muted label —
|
|
new subsystems can be connected as sources; label proposed 2026-09-06 over
|
|
"Next"/"Futures"/"Others")
|
|
- **What we built** (tinted): Staging (raw replica) → Integration (cleaned,
|
|
normalized, deduplicated) → Star schema (dimensional transformation) →
|
|
Datamarts (research-ready marts) → two endpoints in bordeaux:
|
|
**AritmoLab Data Warehouse** (queried for research) + **AritmoLab Portal**
|
|
(management & exploration)
|
|
- 🧠 badges with one-line captions:
|
|
- Integration: "AI reads the clinical text: pathologies, procedures,
|
|
drug-challenge outcomes"
|
|
- Datamarts: "AI builds them on demand from plain-English questions (ThothII)"
|
|
- On the existed→built arrow: "The mappings" (AI-drafted ingestion configs)
|
|
- Popup behaviour (Reveal deck): the contribution card opens NEAR the clicked
|
|
brain (not centered) and a red line ties the card to that brain; closing with
|
|
✕, click outside or Esc.
|
|
- Footnote 🧠: "Also at build time: the pipeline itself was co-engineered with
|
|
AI — mapping configs drafted by LLM agents, reviewed by humans, versioned in git."
|
|
## AI contribution popup texts (formatted draft 2026-09-06 — to discuss)
|
|
|
|
Structure per popup: lead line + bullets (bold label — detail; numbers
|
|
highlighted) + **"Positive effects" box** (bordered, always last — the payoff of
|
|
that AI use must stand out). Shown when clicking a 🧠 in the prototype (in
|
|
PPTX: click-to-appear boxes).
|
|
|
|
### 1. Ingestion — "The mappings" (brain on the existed→built arrow)
|
|
|
|
Lead: The mappings that turn three raw sources into clean data are
|
|
**crafted by AI** — **revised by humans**.
|
|
|
|
- **AI drafts** — LLM agents write the mappings from prompts and schema samples
|
|
- **What is a mapping?** — a plain instruction sheet that tells the system,
|
|
field by field, where each piece of data comes from and where it has to land
|
|
- **At scale** — 19,764 lines of instructions, all machine-validated
|
|
- **Humans revise** — every line is checked, then recorded in **git**, the
|
|
archive that keeps every version of the instructions and can undo any change
|
|
|
|
**Positive effects box:** Weeks of hand-writing became days of reviewing.
|
|
|
|
### 2. Integration — "Reading the clinical text"
|
|
|
|
Lead: A deterministic, bilingual (IT/EN) pattern library reads the free text
|
|
inside the nightly pipeline.
|
|
|
|
- **Pathologies** — 73,389 extracted from 58,438 discharge letters, into a
|
|
two-tier clinical ontology
|
|
- **Procedures ↔ pathologies** — linked at clause level
|
|
- **Drug-challenge tests** — 10,908 parsed (flecainide, ajmaline, adrenaline,
|
|
isoprenaline)
|
|
- **No black box** — semver rules, `pattern_version` on every row, ≥95% accuracy
|
|
gate on 100 manual reviews
|
|
|
|
**Positive effects box:** 58,438 letters of free text became an analysable research asset —
|
|
automatically, every night.
|
|
|
|
### 3. Datamarts — "Datamarts from plain English"
|
|
|
|
Lead: Researchers ask in plain English; the AI writes the SQL over the star schema.
|
|
|
|
- **Ask** — "how many Brugada patients had an effective ablation?"
|
|
- **Generate** — AI builds the SQL and assembles a curated datamart
|
|
- **Serve** — results flow to Superset for statistics and machine learning
|
|
- **In the loop** — the researcher reviews every proposed query before it runs
|
|
|
|
**Positive effects box:** A new research datamart in minutes instead of weeks of hand-written
|
|
SQL — with the researcher approving every query.
|
|
|
|
### OPEN — genetics (brain not placed until a/b decision)
|
|
|
|
(a) deterministic → keep no brain on Genetic data.
|
|
(b) AI upstream → brain caption candidate: "AI reconciles genetic records
|
|
upstream: variant nomenclature and lab reports normalized before ingestion"
|
|
(only if MP confirms this happens outside the ETL).
|
|
|
|
## Speaker script (EN, ~75s)
|
|
|
|
Here is the whole project on one slide. On the left, what already existed: three
|
|
hospital data sources — Cardioref, our electrophysiology records; the genetic
|
|
data coming from the labs; and the ECG signals. On the right, what we built in
|
|
AritmoLab: a staging replica, an integration layer where the data is cleaned and
|
|
normalized, a star-schema warehouse, and the research datamarts — delivered
|
|
through the data warehouse and a management portal.
|
|
|
|
Everywhere you see this brain symbol, that's where AI works for us. At
|
|
integration, AI reads the clinical text — pathologies, procedures, and
|
|
drug-challenge outcomes. At the very end, AI builds the datamarts on demand,
|
|
from plain-English questions. And even this pipeline was partly built with AI:
|
|
the mapping configurations were drafted by AI agents, then reviewed by humans
|
|
and versioned in git.
|
|
|
|
My colleague Dr. Paratico will show you exactly how the text reading works. But
|
|
first, let me show you why this is hard.
|
|
|
|
## Notes
|
|
|
|
- Sub-labels of the three sources are placeholders — MP to confirm the real
|
|
wording (esp. ECG: "signals · device follow-up").
|
|
- If the genetics a/b decision lands on (b), add a 🧠 on the Genetic data box
|
|
with the agreed caption.
|