feat: AritmoLab AI presentation deck (work in progress)

Reveal.js deck + prototypes + speaker scripts for the talk
'Role of AI in the Analysis of Unstructured Clinical Databases'
(AritmoLab — Policlinico San Donato, 2-3 October 2026).

Lives on its own branch while in progress; not for main until ready.
This commit is contained in:
Codex
2026-09-06 23:06:28 +02:00
parent cffa60772e
commit 0f1391e1b7
140 changed files with 36007 additions and 0 deletions
+44
View File
@@ -0,0 +1,44 @@
# Outline v3 (FINAL structure — proposal A) — Role of AI in the Analysis of Unstructured Clinical Databases
Deck: `deck/index.html` (Reveal.js, 15 slides, fonts embedded).
Presenters: **MP** = Dr. Marco Pancotti · **SP** = Dr.ssa Sara Paratico.
| # | Slide | Master | By | Time | Status |
|----|------------------------------------------------|--------------|----|------|--------|
| 01 | Title | title | MP | 40s | ✅ final |
| 02 | Where we started (4 systems + 2 missing) | interactive | MP | 65s | ✅ final |
| 03 | The project at a glance (flow + brain popups) | architecture | MP | 75s | ✅ final |
| 04 | What the text miner reads (numbers) | stats | SP | 60s | ✅ real data |
| 05 | The hard problems of clinical NLP | concept | SP | 75s | 🟡 draft |
| 06 | Building the clinical ontology | concept | SP | 75s | 🟡 draft |
| 07 | Trust the text: validation & human review | concept | SP | 75s | 🟡 draft |
| 08 | AI as co-engineer (mapping YAMLs) | artifact | MP | 70s | ✅ real |
| 09 | Genetic data — many sources, one patient | architecture | MP | 60s | ⚠️ wording TBD (a/b) |
| 10 | ThothII: ask the warehouse in plain English | concept | MP | 75s | 🟡 draft |
| 11 | From datamarts to predictive statistics & ML | concept | MP | 60s | 🟡 draft |
| 12 | Live demo cue (star schema · Superset · ThothII)| divider | MP | ~2min live | ✅ |
| 13 | Lessons learned | concept | MP | 60s | 🟡 draft |
| 14 | (merged into 15) | — | — | — | — |
| 15 | Thank you + Q&A (+ Substack pointer) | closing | MP | 25s | 🟡 draft |
Talk ≈ 13.4 min + ~2 min demo. 15 slides = user cap reached.
## Confirmed decisions
- Project named **AritmoLab** everywhere (internal codename never appears).
- Slide 02 (Where we started): map of legacy systems — Cardioref / Genetic data /
Omics Portal / ECG subsystems, each clickable → popup (usage + problem box);
two dashed "MISSING" cards (Health Intelligence, ML-ready datamarts).
- Slide 03 flow: 3 sources + "Future sources" (dashed) → staging → integration 🧠
→ star schema → datamarts 🧠 → AritmoLab (DWH + Portal); brain popups =
The mappings / Reading the clinical text / Datamarts from plain English.
- Text-analysis block (SP): slides 04–07. Scripts to be validated by SP.
- Speaker view (S): Tall layout customized — upcoming 25%, timer bottom-left,
notes 1.45em. Default: upcoming 20%, notes 80%.
- Fonts embedded (Source Sans 3 + Source Serif 4, OFL) — no Google dependency.
## OPEN
- Slide 09 genetics wording: (a) deterministic reconciliation [recommended] vs
(b) AI upstream (only if MP confirms it happens outside the ETL).
- Draft slides 05-07, 10-11, 13-15: review one by one.
- Dark theme of the whole deck: final pass.
- "twenty years" (slide 01 script + slide 02): confirm real data span.
+45
View File
@@ -0,0 +1,45 @@
# Slide 01 — Title
## On-slide (EN)
- Kicker: `AritmoLab · Policlinico San Donato`
- Title: **Role of AI in the Analysis of Unstructured Clinical Databases**
- Subtitle: How we turned a mix of structured and free-text clinical records into a research data platform — the AritmoLab platform experience
- Platform line (beside the AritmoLab pill): A clinical data platform hosting cardiology data and analysis tools
- Speakers (stacked, one per line):
- Dr. Marco Pancotti - MultiPhysixLab
- Dr. Sara Paratico - I.R.C.C.S. Policlinico San Donato
- Event (own line, italic): "Multidimensional Characterization of Cardiac Arrhythmias: Role of
Electrocardiology in the Artificial Intelligence Era"
- Venue & date (own line): San Donato Milanese, Milan, Italy · 2–3 October 2026
## Visual
Clean title master — no diagram (the project flow diagram lives on the first
slide of merit, see 02-flow.md). Title + subtitle breathe; below: pill +
platform one-liner, stacked speakers, event (italic) + venue/date.
## Speaker script (EN, ~45s)
Good morning everyone, and thank you for being here. I'm Marco Pancotti, and with
my colleague Dr. Sara Paratico we work on the clinical data platform of AritmoLab
at Policlinico San Donato.
Today I want to talk about a problem that every hospital knows very well. The most
valuable clinical information we have — the diagnosis, the patient's history, the
outcome of a test — is written in plain free text, inside systems that were never
designed to make that text usable.
Over the next twelve minutes I'll show you how we used artificial intelligence, in
three different roles, to transform twenty years of cardiology records into a
database that researchers can actually query — and that clinicians can verify.
This is the story of the AritmoLab platform. My colleague Dr. Sara Paratico will
join me to show how we read the clinical text itself.
## Notes
- Event: "Multidimensional Characterization of Cardiac Arrhythmias: Role of
Electrocardiology in the Artificial Intelligence Era", 2–3 October 2026.
- "twenty years" — verify against real data span (Cardioref history) before final.
- If the session chair is strict on time, cut the second paragraph and go straight
to "Over the next twelve minutes…".
+111
View File
@@ -0,0 +1,111 @@
# Slide 02 — The project at a glance (project flow diagram)
First slide of merit in ALL three content proposals (A/B/C). It absorbs the old
"platform at a glance" slide. Presenter: MP.
## On-slide (EN)
- Kicker: `What existed → what we built`
- Title: **The project at a glance**
- Diagram, two zones left→right:
- **What existed** (muted, dashed): Cardioref (cardiology records · procedures ·
letters) · Genetic data (labs · variants · nomenclature) · ECG (signals ·
device follow-up) · **Future sources** (dashed, "+" badge, muted label —
new subsystems can be connected as sources; label proposed 2026-09-06 over
"Next"/"Futures"/"Others")
- **What we built** (tinted): Staging (raw replica) → Integration (cleaned,
normalized, deduplicated) → Star schema (dimensional transformation) →
Datamarts (research-ready marts) → two endpoints in bordeaux:
**AritmoLab Data Warehouse** (queried for research) + **AritmoLab Portal**
(management & exploration)
- 🧠 badges with one-line captions:
- Integration: "AI reads the clinical text: pathologies, procedures,
drug-challenge outcomes"
- Datamarts: "AI builds them on demand from plain-English questions (ThothII)"
- On the existed→built arrow: "The mappings" (AI-drafted ingestion configs)
- Popup behaviour (Reveal deck): the contribution card opens NEAR the clicked
brain (not centered) and a red line ties the card to that brain; closing with
✕, click outside or Esc.
- Footnote 🧠: "Also at build time: the pipeline itself was co-engineered with
AI — mapping configs drafted by LLM agents, reviewed by humans, versioned in git."
## AI contribution popup texts (formatted draft 2026-09-06 — to discuss)
Structure per popup: lead line + bullets (bold label — detail; numbers
highlighted) + **"Positive effects" box** (bordered, always last — the payoff of
that AI use must stand out). Shown when clicking a 🧠 in the prototype (in
PPTX: click-to-appear boxes).
### 1. Ingestion — "The mappings" (brain on the existed→built arrow)
Lead: The mappings that turn three raw sources into clean data are
**crafted by AI** — **revised by humans**.
- **AI drafts** — LLM agents write the mappings from prompts and schema samples
- **What is a mapping?** — a plain instruction sheet that tells the system,
field by field, where each piece of data comes from and where it has to land
- **At scale** — 19,764 lines of instructions, all machine-validated
- **Humans revise** — every line is checked, then recorded in **git**, the
archive that keeps every version of the instructions and can undo any change
**Positive effects box:** Weeks of hand-writing became days of reviewing.
### 2. Integration — "Reading the clinical text"
Lead: A deterministic, bilingual (IT/EN) pattern library reads the free text
inside the nightly pipeline.
- **Pathologies** — 73,389 extracted from 58,438 discharge letters, into a
two-tier clinical ontology
- **Procedures ↔ pathologies** — linked at clause level
- **Drug-challenge tests** — 10,908 parsed (flecainide, ajmaline, adrenaline,
isoprenaline)
- **No black box** — semver rules, `pattern_version` on every row, ≥95% accuracy
gate on 100 manual reviews
**Positive effects box:** 58,438 letters of free text became an analysable research asset —
automatically, every night.
### 3. Datamarts — "Datamarts from plain English"
Lead: Researchers ask in plain English; the AI writes the SQL over the star schema.
- **Ask** — "how many Brugada patients had an effective ablation?"
- **Generate** — AI builds the SQL and assembles a curated datamart
- **Serve** — results flow to Superset for statistics and machine learning
- **In the loop** — the researcher reviews every proposed query before it runs
**Positive effects box:** A new research datamart in minutes instead of weeks of hand-written
SQL — with the researcher approving every query.
### OPEN — genetics (brain not placed until a/b decision)
(a) deterministic → keep no brain on Genetic data.
(b) AI upstream → brain caption candidate: "AI reconciles genetic records
upstream: variant nomenclature and lab reports normalized before ingestion"
(only if MP confirms this happens outside the ETL).
## Speaker script (EN, ~75s)
Here is the whole project on one slide. On the left, what already existed: three
hospital data sources — Cardioref, our electrophysiology records; the genetic
data coming from the labs; and the ECG signals. On the right, what we built in
AritmoLab: a staging replica, an integration layer where the data is cleaned and
normalized, a star-schema warehouse, and the research datamarts — delivered
through the data warehouse and a management portal.
Everywhere you see this brain symbol, that's where AI works for us. At
integration, AI reads the clinical text — pathologies, procedures, and
drug-challenge outcomes. At the very end, AI builds the datamarts on demand,
from plain-English questions. And even this pipeline was partly built with AI:
the mapping configurations were drafted by AI agents, then reviewed by humans
and versioned in git.
My colleague Dr. Paratico will show you exactly how the text reading works. But
first, let me show you why this is hard.
## Notes
- Sub-labels of the three sources are placeholders — MP to confirm the real
wording (esp. ECG: "signals · device follow-up").
- If the genetics a/b decision lands on (b), add a 🧠 on the Genetic data box
with the agreed caption.