feat: AritmoLab AI presentation deck (work in progress)
Reveal.js deck + prototypes + speaker scripts for the talk 'Role of AI in the Analysis of Unstructured Clinical Databases' (AritmoLab — Policlinico San Donato, 2-3 October 2026). Lives on its own branch while in progress; not for main until ready.
This commit is contained in:
@@ -0,0 +1,44 @@
|
||||
# Outline v3 (FINAL structure — proposal A) — Role of AI in the Analysis of Unstructured Clinical Databases
|
||||
|
||||
Deck: `deck/index.html` (Reveal.js, 15 slides, fonts embedded).
|
||||
Presenters: **MP** = Dr. Marco Pancotti · **SP** = Dr.ssa Sara Paratico.
|
||||
|
||||
| # | Slide | Master | By | Time | Status |
|
||||
|----|------------------------------------------------|--------------|----|------|--------|
|
||||
| 01 | Title | title | MP | 40s | ✅ final |
|
||||
| 02 | Where we started (4 systems + 2 missing) | interactive | MP | 65s | ✅ final |
|
||||
| 03 | The project at a glance (flow + brain popups) | architecture | MP | 75s | ✅ final |
|
||||
| 04 | What the text miner reads (numbers) | stats | SP | 60s | ✅ real data |
|
||||
| 05 | The hard problems of clinical NLP | concept | SP | 75s | 🟡 draft |
|
||||
| 06 | Building the clinical ontology | concept | SP | 75s | 🟡 draft |
|
||||
| 07 | Trust the text: validation & human review | concept | SP | 75s | 🟡 draft |
|
||||
| 08 | AI as co-engineer (mapping YAMLs) | artifact | MP | 70s | ✅ real |
|
||||
| 09 | Genetic data — many sources, one patient | architecture | MP | 60s | ⚠️ wording TBD (a/b) |
|
||||
| 10 | ThothII: ask the warehouse in plain English | concept | MP | 75s | 🟡 draft |
|
||||
| 11 | From datamarts to predictive statistics & ML | concept | MP | 60s | 🟡 draft |
|
||||
| 12 | Live demo cue (star schema · Superset · ThothII)| divider | MP | ~2min live | ✅ |
|
||||
| 13 | Lessons learned | concept | MP | 60s | 🟡 draft |
|
||||
| 14 | (merged into 15) | — | — | — | — |
|
||||
| 15 | Thank you + Q&A (+ Substack pointer) | closing | MP | 25s | 🟡 draft |
|
||||
|
||||
Talk ≈ 13.4 min + ~2 min demo. 15 slides = user cap reached.
|
||||
|
||||
## Confirmed decisions
|
||||
- Project named **AritmoLab** everywhere (internal codename never appears).
|
||||
- Slide 02 (Where we started): map of legacy systems — Cardioref / Genetic data /
|
||||
Omics Portal / ECG subsystems, each clickable → popup (usage + problem box);
|
||||
two dashed "MISSING" cards (Health Intelligence, ML-ready datamarts).
|
||||
- Slide 03 flow: 3 sources + "Future sources" (dashed) → staging → integration 🧠
|
||||
→ star schema → datamarts 🧠 → AritmoLab (DWH + Portal); brain popups =
|
||||
The mappings / Reading the clinical text / Datamarts from plain English.
|
||||
- Text-analysis block (SP): slides 04–07. Scripts to be validated by SP.
|
||||
- Speaker view (S): Tall layout customized — upcoming 25%, timer bottom-left,
|
||||
notes 1.45em. Default: upcoming 20%, notes 80%.
|
||||
- Fonts embedded (Source Sans 3 + Source Serif 4, OFL) — no Google dependency.
|
||||
|
||||
## OPEN
|
||||
- Slide 09 genetics wording: (a) deterministic reconciliation [recommended] vs
|
||||
(b) AI upstream (only if MP confirms it happens outside the ETL).
|
||||
- Draft slides 05-07, 10-11, 13-15: review one by one.
|
||||
- Dark theme of the whole deck: final pass.
|
||||
- "twenty years" (slide 01 script + slide 02): confirm real data span.
|
||||
@@ -0,0 +1,45 @@
|
||||
# Slide 01 — Title
|
||||
|
||||
## On-slide (EN)
|
||||
|
||||
- Kicker: `AritmoLab · Policlinico San Donato`
|
||||
- Title: **Role of AI in the Analysis of Unstructured Clinical Databases**
|
||||
- Subtitle: How we turned a mix of structured and free-text clinical records into a research data platform — the AritmoLab platform experience
|
||||
- Platform line (beside the AritmoLab pill): A clinical data platform hosting cardiology data and analysis tools
|
||||
- Speakers (stacked, one per line):
|
||||
- Dr. Marco Pancotti - MultiPhysixLab
|
||||
- Dr. Sara Paratico - I.R.C.C.S. Policlinico San Donato
|
||||
- Event (own line, italic): "Multidimensional Characterization of Cardiac Arrhythmias: Role of
|
||||
Electrocardiology in the Artificial Intelligence Era"
|
||||
- Venue & date (own line): San Donato Milanese, Milan, Italy · 2–3 October 2026
|
||||
|
||||
## Visual
|
||||
|
||||
Clean title master — no diagram (the project flow diagram lives on the first
|
||||
slide of merit, see 02-flow.md). Title + subtitle breathe; below: pill +
|
||||
platform one-liner, stacked speakers, event (italic) + venue/date.
|
||||
|
||||
## Speaker script (EN, ~45s)
|
||||
|
||||
Good morning everyone, and thank you for being here. I'm Marco Pancotti, and with
|
||||
my colleague Dr. Sara Paratico we work on the clinical data platform of AritmoLab
|
||||
at Policlinico San Donato.
|
||||
|
||||
Today I want to talk about a problem that every hospital knows very well. The most
|
||||
valuable clinical information we have — the diagnosis, the patient's history, the
|
||||
outcome of a test — is written in plain free text, inside systems that were never
|
||||
designed to make that text usable.
|
||||
|
||||
Over the next twelve minutes I'll show you how we used artificial intelligence, in
|
||||
three different roles, to transform twenty years of cardiology records into a
|
||||
database that researchers can actually query — and that clinicians can verify.
|
||||
This is the story of the AritmoLab platform. My colleague Dr. Sara Paratico will
|
||||
join me to show how we read the clinical text itself.
|
||||
|
||||
## Notes
|
||||
|
||||
- Event: "Multidimensional Characterization of Cardiac Arrhythmias: Role of
|
||||
Electrocardiology in the Artificial Intelligence Era", 2–3 October 2026.
|
||||
- "twenty years" — verify against real data span (Cardioref history) before final.
|
||||
- If the session chair is strict on time, cut the second paragraph and go straight
|
||||
to "Over the next twelve minutes…".
|
||||
@@ -0,0 +1,111 @@
|
||||
# Slide 02 — The project at a glance (project flow diagram)
|
||||
|
||||
First slide of merit in ALL three content proposals (A/B/C). It absorbs the old
|
||||
"platform at a glance" slide. Presenter: MP.
|
||||
|
||||
## On-slide (EN)
|
||||
|
||||
- Kicker: `What existed → what we built`
|
||||
- Title: **The project at a glance**
|
||||
- Diagram, two zones left→right:
|
||||
- **What existed** (muted, dashed): Cardioref (cardiology records · procedures ·
|
||||
letters) · Genetic data (labs · variants · nomenclature) · ECG (signals ·
|
||||
device follow-up) · **Future sources** (dashed, "+" badge, muted label —
|
||||
new subsystems can be connected as sources; label proposed 2026-09-06 over
|
||||
"Next"/"Futures"/"Others")
|
||||
- **What we built** (tinted): Staging (raw replica) → Integration (cleaned,
|
||||
normalized, deduplicated) → Star schema (dimensional transformation) →
|
||||
Datamarts (research-ready marts) → two endpoints in bordeaux:
|
||||
**AritmoLab Data Warehouse** (queried for research) + **AritmoLab Portal**
|
||||
(management & exploration)
|
||||
- 🧠 badges with one-line captions:
|
||||
- Integration: "AI reads the clinical text: pathologies, procedures,
|
||||
drug-challenge outcomes"
|
||||
- Datamarts: "AI builds them on demand from plain-English questions (ThothII)"
|
||||
- On the existed→built arrow: "The mappings" (AI-drafted ingestion configs)
|
||||
- Popup behaviour (Reveal deck): the contribution card opens NEAR the clicked
|
||||
brain (not centered) and a red line ties the card to that brain; closing with
|
||||
✕, click outside or Esc.
|
||||
- Footnote 🧠: "Also at build time: the pipeline itself was co-engineered with
|
||||
AI — mapping configs drafted by LLM agents, reviewed by humans, versioned in git."
|
||||
## AI contribution popup texts (formatted draft 2026-09-06 — to discuss)
|
||||
|
||||
Structure per popup: lead line + bullets (bold label — detail; numbers
|
||||
highlighted) + **"Positive effects" box** (bordered, always last — the payoff of
|
||||
that AI use must stand out). Shown when clicking a 🧠 in the prototype (in
|
||||
PPTX: click-to-appear boxes).
|
||||
|
||||
### 1. Ingestion — "The mappings" (brain on the existed→built arrow)
|
||||
|
||||
Lead: The mappings that turn three raw sources into clean data are
|
||||
**crafted by AI** — **revised by humans**.
|
||||
|
||||
- **AI drafts** — LLM agents write the mappings from prompts and schema samples
|
||||
- **What is a mapping?** — a plain instruction sheet that tells the system,
|
||||
field by field, where each piece of data comes from and where it has to land
|
||||
- **At scale** — 19,764 lines of instructions, all machine-validated
|
||||
- **Humans revise** — every line is checked, then recorded in **git**, the
|
||||
archive that keeps every version of the instructions and can undo any change
|
||||
|
||||
**Positive effects box:** Weeks of hand-writing became days of reviewing.
|
||||
|
||||
### 2. Integration — "Reading the clinical text"
|
||||
|
||||
Lead: A deterministic, bilingual (IT/EN) pattern library reads the free text
|
||||
inside the nightly pipeline.
|
||||
|
||||
- **Pathologies** — 73,389 extracted from 58,438 discharge letters, into a
|
||||
two-tier clinical ontology
|
||||
- **Procedures ↔ pathologies** — linked at clause level
|
||||
- **Drug-challenge tests** — 10,908 parsed (flecainide, ajmaline, adrenaline,
|
||||
isoprenaline)
|
||||
- **No black box** — semver rules, `pattern_version` on every row, ≥95% accuracy
|
||||
gate on 100 manual reviews
|
||||
|
||||
**Positive effects box:** 58,438 letters of free text became an analysable research asset —
|
||||
automatically, every night.
|
||||
|
||||
### 3. Datamarts — "Datamarts from plain English"
|
||||
|
||||
Lead: Researchers ask in plain English; the AI writes the SQL over the star schema.
|
||||
|
||||
- **Ask** — "how many Brugada patients had an effective ablation?"
|
||||
- **Generate** — AI builds the SQL and assembles a curated datamart
|
||||
- **Serve** — results flow to Superset for statistics and machine learning
|
||||
- **In the loop** — the researcher reviews every proposed query before it runs
|
||||
|
||||
**Positive effects box:** A new research datamart in minutes instead of weeks of hand-written
|
||||
SQL — with the researcher approving every query.
|
||||
|
||||
### OPEN — genetics (brain not placed until a/b decision)
|
||||
|
||||
(a) deterministic → keep no brain on Genetic data.
|
||||
(b) AI upstream → brain caption candidate: "AI reconciles genetic records
|
||||
upstream: variant nomenclature and lab reports normalized before ingestion"
|
||||
(only if MP confirms this happens outside the ETL).
|
||||
|
||||
## Speaker script (EN, ~75s)
|
||||
|
||||
Here is the whole project on one slide. On the left, what already existed: three
|
||||
hospital data sources — Cardioref, our electrophysiology records; the genetic
|
||||
data coming from the labs; and the ECG signals. On the right, what we built in
|
||||
AritmoLab: a staging replica, an integration layer where the data is cleaned and
|
||||
normalized, a star-schema warehouse, and the research datamarts — delivered
|
||||
through the data warehouse and a management portal.
|
||||
|
||||
Everywhere you see this brain symbol, that's where AI works for us. At
|
||||
integration, AI reads the clinical text — pathologies, procedures, and
|
||||
drug-challenge outcomes. At the very end, AI builds the datamarts on demand,
|
||||
from plain-English questions. And even this pipeline was partly built with AI:
|
||||
the mapping configurations were drafted by AI agents, then reviewed by humans
|
||||
and versioned in git.
|
||||
|
||||
My colleague Dr. Paratico will show you exactly how the text reading works. But
|
||||
first, let me show you why this is hard.
|
||||
|
||||
## Notes
|
||||
|
||||
- Sub-labels of the three sources are placeholders — MP to confirm the real
|
||||
wording (esp. ECG: "signals · device follow-up").
|
||||
- If the genetics a/b decision lands on (b), add a 🧠 on the Genetic data box
|
||||
with the agreed caption.
|
||||
Reference in New Issue
Block a user