<pclass="event">“Multidimensional Characterization of Cardiac Arrhythmias: Role of Electrocardiology in the Artificial Intelligence Era”</p>
<pclass="event-where">San Donato Milanese, Milan, Italy · 2–3 October 2026</p>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">01 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">01 / 13</span></footer>
<asideclass="notes">
Good morning everyone, and thank you for being here. I'm Marco Pancotti, and with my colleague Dr. Sara Paratico we work on the clinical data platform of AritmoLab at Policlinico San Donato.
Today I want to talk about a problem that every hospital knows very well. The most valuable clinical information we have — the diagnosis, the patient's history, the outcome of a test — is written in plain free text, inside systems that were never designed to make that text usable.
Over the next twelve minutes I'll show you how we used artificial intelligence, in three different roles, to transform twenty years of cardiology records into a database that researchers can actually query — and that clinicians can verify. My colleague Dr. Sara Paratico will join me to show how we read the clinical text itself.
Good morning. My name is Marco Pancotti, and I led the part of the PAMP-FA project dedicated to building the AritmoLab portal at Policlinico San Donato, with support from Sara Paratico, who will co-present with me today.
The scope of the project included a data warehouse and tools formachine learning and predictive statistics to support research by the Arrhythmology Unit, directed by Professor Pappone and Professor Locati.
Today we'll focus on the use of AI to extract structured information from clinical text, followed by a brief tour of the portal and ThothII, a tool for building datamarts from natural-language requests.
</aside>
</section>
@@ -52,7 +55,7 @@
<spanclass="band-section">Where we started</span>
</nav>
<divclass="sbody">
<h2>Four islands and four missing pieces</h2>
<h2>Where we started - Four islands and four missing pieces</h2>
<divclass="todo"data-missing="ci"style="margin-left:0"><spanclass="q">?</span><spanclass="tx"><b>Clinical Intelligence</b><span>dashboards on the clinical history of the Unit and its patients</span></span></div>
@@ -68,7 +71,7 @@
<spanclass="nm">Genetic data</span><spanclass="sc">DNA instruments — results typed into Excel by hand</span></button>
<spanclass="nm">ECG</span><spanclass="sc">paper strip + a CSV on request</span></button>
@@ -76,9 +79,20 @@
<pclass="hint"style="position:absolute; left:26%; width:48%; bottom:2px; margin:0; text-align:center">▸ click a system: how it was used — and what held it back · click a missing piece: what it would have given</p>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">02 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">02 / 13</span></footer>
<asideclass="notes">
Before I show you what we built, let me take you back twenty years. Cardioref, our electrophysiology database, had twenty years of clinical records: as a management tool for patients, daily activity and planning it was — and still is — excellent. But for research it was almost unusable, because most of the clinical knowledge was written in free text, outside the structured fields. The genetic data lived in its own system, which produced Excel sheets, with no connection to Cardioref. Our first Omics Portal was a promising idea, still embryonic and not integrated with anything. And a constellation of ECG subsystems that could only hand you a CSV file, on request. On top of that: no clinical intelligence on the historical data, no single datawarehouse as the source of truth, no ML-ready data for predictive statistics, and no professional portal to bring it all together. Four islands, four missing pieces. Click any system and I'll show you exactly what held it back. And here is what we built.
The starting point was four disconnected systems:
1 - <strong>Cardioref</strong> held twenty years of electrophysiology records and supported clinical operations, but much of the information needed for research was in free text.
2 - <strong>Genetic data</strong> were manually entered into Excel, without integration with Cardioref.
3 - The <strong>Aritmolab Portal</strong> was an early prototype, not yet integrated with the source systems.
4 - <strong>ECG</strong> data were available as paper records or CSV exports on request.
Whe lacked:
5 - some longitudinal <strong>clinical analytics</strong>
6 - a shared <strong>data warehouse</strong>,
7 - a set of structured datasets for model training,
8 - and a unified <strong>clinical portal</strong>.
</aside>
</section>
@@ -94,37 +108,72 @@
<divclass="sbody">
<divclass="kicker">The plan</div>
<h2>What we wanted to build</h2>
<pclass="lead">One platform, two destinations: the 360° patient portal and a research-ready warehouse. Built on open-source pillars, with AI coding agents wherever they helped. This talk follows the hardest piece — the unstructured data.</p>
<pclass="lead">One platform, two destinations: the 360° patient portal and a research-ready warehouse. Built on open-source pillars, with AI coding agents wherever they helped. Next, we’ll look at the hardest part: unstructured data.</p>
<divclass="flow">
<divclass="zone existed">
<buttontype="button"class="zone existed lens-target"id="plan-existed"data-lens="plan-existed"data-lens-title="What existed"data-lens-number="1"aria-label="Explore What existed"aria-expanded="false"aria-controls="slide-lens">
<spanclass="zlabel">What existed</span>
<divclass="fcol">
<divclass="fbox src"><b>Cardioref</b><span>cardiology records · procedures · letters</span></div>
<spanclass="fbox src future"><b><spanclass="plus">+</span>Future sources</b><span>new subsystems can be connected as sources</span></span>
</span>
<templateclass="lens-details">
<pclass="detail-intro">The starting material offered complementary views of the patient.</p>
<p><b>Clinical course · Cardioref</b>Visits, procedures and reports describe the course of care. Dates and narrative details give each event its clinical context.</p>
<p><b>Genetic findings</b>Laboratory results and variant descriptions add the genetic perspective, recorded separately from the clinical history.</p>
<p><b>Electrical activity · ECG</b>Tracings and exported signals document cardiac electrical activity, complementing the written account of the patient’s condition.</p>
<p><b>Different forms of evidence</b>Structured fields, free text and signals must retain their clinical meaning when connected. Future sources would extend this initial set.</p>
</template>
</button>
<divclass="fcol ai-hit">
<buttonclass="brain brain-btn"data-ai="ingestion"aria-label="AI contribution: the mappings">🧠</button>
<buttonclass="brain brain-btn"data-ai="ingestion"aria-label="AI contribution: the mappings"aria-pressed="false"><spanclass="ai-label"aria-hidden="true">AI</span><imgclass="artificial-brain"src="artificial-brain.svg"alt=""><spanclass="ai-number">6</span></button>
<divclass="fcol"style="flex:0.8; justify-content:center"><divclass="fbox"><b>Staging</b><span>raw replica of the sources</span></div></div>
<divclass="fcol"style="flex:0.8; justify-content:center"><buttontype="button"class="fbox lens-target"id="plan-staging"data-lens="plan-staging"data-lens-number="2"aria-label="Explore Staging"aria-expanded="false"aria-controls="slide-lens"><b>Staging</b><span>raw replica of the sources</span>
<templateclass="lens-details">
<pclass="detail-intro">The landing area: a faithful working copy of the data brought in from each source system.</p>
<p><b>What it contains</b>Original tables, identifiers, dates and clinical text, still in the source’s format.</p>
<p><b>Why it matters</b>Subsequent processing works on this copy. The original clinical system continues its daily work, and transformations can be checked against the imported data.</p>
<pclass="detail-example"><b>Clinical example</b>A Cardioref letter arrives with its original wording. “No syncope” is still text; this layer does not yet turn it into a clinical variable.</p>
</template>
</button></div>
<divclass="farrow">→</div>
<divclass="fcol ai"style="flex:1.25">
<buttonclass="brain brain-btn"data-ai="integration"aria-label="AI contribution: clinical text reading">🧠</button>
<pclass="detail-intro">The reconciliation layer: source records become consistent, connected clinical information.</p>
<p><b>What happens here</b>Formats and terminology are standardized, duplicates reconciled, and records linked through patient and event identifiers.</p>
<p><b>Where AI contributes</b>It extracts conditions, procedures and drug-challenge outcomes from clinical text. Negation and context matter; extraction quality needs validation.</p>
<pclass="detail-example"><b>Clinical example</b>“No syncope” is not a positive finding. A family history of Brugada must remain distinct from the patient’s own diagnosis.</p>
</template>
</button>
<divclass="fcap">AI reads the clinical text: pathologies, procedures, drug-challenge outcomes</div>
<divclass="fcol"style="flex:1.05; justify-content:center"><buttontype="button"class="fbox lens-target"id="plan-star-schema"data-lens="plan-star-schema"data-lens-number="4"aria-label="Explore Data warehouse"aria-expanded="false"aria-controls="slide-lens"><b>Data warehouse</b><span>star schema · organized for analysis</span>
<templateclass="lens-details">
<pclass="detail-intro">The shared analytical database: integrated clinical data reorganized for research across patients and over time.</p>
<p><b>How a star schema works</b>“Facts” represent events or measurements, such as a procedure or test. “Dimensions” describe their context, such as the patient, date and procedure type.</p>
<p><b>Why it matters</b>Researchers can filter, group and compare records using common definitions, without reconstructing every relationship from the original hospital tables.</p>
<pclass="detail-example"><b>Clinical question it can support</b>How many patients underwent a given procedure each year, and how does that distribution vary by age group?</p>
<pclass="detail-intro">Focused datasets derived from the warehouse for a specific research question, study or dashboard.</p>
<p><b>What they define</b>The cohort, time window, variables and level of detail: for example, one row per patient or one row per procedure.</p>
<p><b>How ThothII helps</b>A plain-English question becomes a proposed SQL query through a guided workflow with human review. The resulting dataset supports analysis and portal dashboards.</p>
<pclass="detail-example"><b>Illustrative study dataset</b>Patients who underwent a drug-challenge test, with test date, result and selected clinical characteristics. Its cohort and variable definitions must be agreed before interpreting results.</p>
</template>
</button>
<divclass="fcap">AI builds them on demand from plain-English questions (ThothII)</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">03 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">03 / 13</span></footer>
<asideclass="notes">
Here is the whole project on one slide. On the left, what already existed: three hospital data sources — Cardioref, our electrophysiology records; the genetic data coming from the labs; and the ECG signals. On the right, what we built in AritmoLab: a staging replica, an integration layer where the data is cleaned and normalized, a star-schema warehouse, and the research datamarts — delivered through the data warehouse and a management portal.
Everywhere you see the brain symbol, that's where AI works for us — click each brain and its contribution card opens, tied to the brain by a red line. At integration, AI reads the clinical text. At the end, AI builds the datamarts on demand from plain-English questions. And the mappings that bring the sources in were themselves drafted by AI, then revised by humans.
That was the plan. And to build it, we used AI everywhere it helped — drafting configs, reading clinical text, and at the very end, producing the datamarts themselves. This talk follows the hardest piece of that plan: the unstructured data. My colleague Dr. Paratico will show you exactly how the text reading works. But first, the first step of the climb: the mappings.
Here is the whole project on one slide.
On the left, what already existed, three hospital data sources:
(1) <strong>What existed</strong>: <strong>Cardioref</strong>, our electrophysiology records dataset, the <strong>genetic data</strong> coming from the labs, the <strong>ECG signals</strong>, and, in the future, other sources.
On the right, what we built in AritmoLab:
(2) a staging copy of all relevant Cardioref data, (3) an integration layer where the data is cleaned and normalized, (4) a star-schema warehouse, where the data are reorganized as facts and dimensions, and (5) research datamarts, generated from the data warehouse and presented through dashboards embedded in the portal.
Everywhere you see the brain symbol, that's where AI works for us.
6 - The <strong>mappings</strong> that bring the sources in were themselves drafted by AI, then revised by humans
7 - At <strong>integration</strong>, AI reads the clinical text.
8 - At the end, AI builds the <strong>datamarts</strong> on demand from plain-English questions.
That was the plan. And to build it, we used AI everywhere it helped — drafting configs, reading clinical text, and at the very end, producing the datamarts themselves.
Next, we’ll look at the hardest part: unstructured data. My colleague, Sara Paratico,do will show you exactly how the text reading works. But first, the first step of the climb: the mappings.
</aside>
</section>
@@ -207,9 +268,9 @@
</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">04 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">04 / 13</span></footer>
<asideclass="notes">
[DRAFT] Here is one real sentence from a discharge letter — in Italian, as the clinicians wrote it. In this single sentence there is a diagnosis (syncope), an ECG finding (ST elevation in V1–V3), and a drug-challenge result (flecainide positive for Brugada pattern). For a human cardiologist this is readable in two seconds. For a database, this is just a blob of text in a column.
Here is one real sentence from a discharge letter — in Italian, as the clinicians wrote it. In this single sentence there is a diagnosis (syncope), an ECG finding (ST elevation in V1–V3), and a drug-challenge result (flecainide positive for Brugada pattern). For a human cardiologist this is readable in two seconds. For a database, this is just a blob of text in a column.
The structured tables — demographics, procedures, dates — only tell half the story. The rest is locked inside these free-text fields, in every hospital system we have.
</aside>
</section>
@@ -253,7 +314,7 @@
<divclass="nitem"><b>Audited.</b> A new rule set goes live only after it proves ≥95% accuracy on 100 records reviewed by hand.</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">05 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">05 / 13</span></footer>
<asideclass="notes">
[DRAFT — Sara] I'll take you inside these numbers. We read 58,438 discharge letters — every single one, every night. From them the text miner extracted 73,389 pathology records, classified into our clinical ontology. It parsed 10,908 drug-challenge tests, distinguishing the therapy from the actual test. And at the end of the chain: 2,307 patients with a confirmed Brugada pattern — a cohort nobody could have built by hand.
This is not a one-off migration: the pipeline runs every night, and today it processes on average two hundred and forty letters a month. Four simple pieces do the work: a letter reader for the Italian and English text, a clinical matcher that recognises diagnoses, procedures and test outcomes, a context guard that keeps negations and family history apart, and the ontology sorter that files every finding into its clinical category.
@@ -306,7 +367,7 @@
</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">06 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">06 / 13</span></footer>
<asideclass="notes">
[DRAFT — Sara] Why can't we just search for "FA"? Because clinical text lies to naive search. "Fibrillazione atriale esclusa" contains the words of a diagnosis but negates it — so every rule scans a window around the match, looking for negation cues. "Padre con FA" is real atrial fibrillation — but in the father, not the patient: we record it as family history. Letters mix Italian and English, abbreviations collide ("TA" is blood pressure, not a therapy), and one field can contain five different statements — so we split the text into clauses first.
Each of these problems has a specific, versioned solution. Sara to expand with real corpus examples.
@@ -349,7 +410,7 @@
</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">07 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">07 / 13</span></footer>
<asideclass="notes">
[DRAFT — Sara] Extraction needs a target vocabulary — that's the ontology. Tier 1 holds the eleven arrhythmological categories that matter most for our research; Tier 2 holds five structural conditions. The order of the rules is itself clinical knowledge: Brugada patterns are tested before TV, because "substrato per TV" often appears in Brugada reports and would otherwise mask the diagnosis.
The ontology is versioned like software: a semantic version for the pattern library, stamped on every extracted row. When we add a synonym or fix a rule, the change is traceable — and the data can be rebuilt.
@@ -359,7 +420,7 @@
</section>
<sectionclass="arit center-v">
<sectionclass="arit center-v validation">
<headerclass="head">
<imgsrc="logo.png"alt="">
<spanclass="head-org">AritmoLab · Policlinico San Donato</span>
@@ -368,34 +429,74 @@
<spanclass="band-section">Text analysis</span>
</nav>
<divclass="sbody">
<!-- Verified against aritmolab/chirone-etl on Gitea, commit 2ab00106188f298c1a0c1e2c48b3a4dbc66c45b5.
Exact sources and interpretation limits: ../slides/08-validation-sources.md.
SC-003/004 are targets; ~893 refers to false-positive new-letter records, not distinct patients. -->
<divclass="kicker">Validation</div>
<h2>Trusting the text is a process, not a promise</h2>
<divclass="oitem"><spanclass="n">02</span><div><b>Clinical review</b><spanclass="d">— the extracted Brugada cohort was validated by clinicians on Superset dashboards</span></div></div>
<divclass="oitem"><spanclass="n">03</span><div><b>Feedback loop</b><spanclass="d">— ~893 Brugada false positives found and removed in pattern-library v1.3.1</span></div></div>
<divclass="oitem"><spanclass="n">04</span><div><b>Nothing is silently dropped</b><spanclass="d">— negated and family-attributed findings are stored too, filtered only at the mart layer</span></div></div>
<pclass="detail-intro">We separate correct classification from how often the rules find a condition.</p>
<p><b>Procedure classification</b>The specification sets a target of at least 95% agreement with a manual review of 100 randomly selected records: did the system assign the right procedure type?</p>
<p><b>Pathology coverage</b>The 85% target asks whether at least one recognized arrhythmia is extracted from a record, assuming most procedures concern known arrhythmias. It does not measure whether every diagnosis is correct or every disease is found.</p>
<pclass="detail-example"><b>How to interpret the numbers</b>These are acceptance targets. Clinical review is still needed to detect incorrect labels and missed findings; coverage alone cannot establish clinical accuracy.</p>
</template>
</button>
<buttontype="button"class="oitem lens-target"id="validation-review"data-lens="validation-review"data-lens-number="2"data-lens-placement="center"data-lens-caption="Validation · 02 / 04"aria-label="Explore Clinical criteria"aria-expanded="false"aria-controls="slide-lens"><spanclass="n">02</span><span><b>Clinical criteria</b><spanclass="d">— explicit criteria determine which patients belong in the analysis; a disease name alone is not enough</span></span>
<templateclass="lens-details">
<pclass="detail-intro">Finding a disease name is the first step. Inclusion in an analysis depends on what the text means and the criteria for that cohort.</p>
<p><b>Read the context</b>“Father with Brugada” concerns a relative; “test negative for Brugada” reports a negative result; “suspected Brugada” expresses uncertainty. None of these phrases alone establishes a diagnosis in the patient.</p>
<p><b>Make the selection explicit</b>Queries apply the cohort criteria to prepare the dataset. Superset displays the result. The diagnosis view first excludes negated findings and findings attributed to relatives.</p>
<pclass="detail-example"><b>A rule used in this project</b>For Brugada and long QT syndrome, that view also requires a positive provocative test or an ablation for the condition, at patient level. Other pathologies use only the first filter. These are project inclusion rules, not a universal diagnostic standard.</p>
<pclass="detail-intro">The 11 June 2026 audit identified about 893 false-positive Brugada records in the newer letter format.</p>
<p><b>The error and the fix</b>“Test alla flecainide negativo per sindrome di Brugada” was treated as an affirmed finding: the rules recognized “negato”, but missed “negativo”. Version 1.3.1 added “negativo”, “negativa” and “negativi” before and after the condition.</p>
<p><b>Make the correction repeatable</b>Regression tests require that this sentence still produces a Brugada finding, now marked as negated. A separate test checks that an affirmed Brugada diagnosis remains positive. Reprocessing applies the revised rules to historical letters.</p>
<pclass="detail-example"><b>What changed</b>The finding is reclassified, not erased. The ~893 figure counts affected records, not necessarily distinct patients; the cohort filters can now exclude those negated findings.</p>
</template>
</button>
<buttontype="button"class="oitem lens-target"id="validation-context"data-lens="validation-context"data-lens-number="4"data-lens-placement="center"data-lens-caption="Validation · 04 / 04"aria-label="Explore Nothing is silently dropped"aria-expanded="false"aria-controls="slide-lens"><spanclass="n">04</span><span><b>Nothing is silently dropped</b><spanclass="d">— negated and family-attributed findings are stored too, filtered only at the mart layer</span></span>
<templateclass="lens-details">
<pclass="detail-intro">A recognized finding can be kept even when it is negated or refers to someone else.</p>
<p><b>Store context alongside the finding</b>The extractor looks up to 60 characters before and after the match, stopping at a full stop, semicolon or line break. It records whether the finding is negated and whether it concerns the patient or a relative.</p>
<p><b>Filter when building the research dataset</b>Those flags travel through integration and the warehouse to the mentions dataset. The confirmed-diagnosis view excludes negated and family findings, while the underlying dataset remains available for other analyses.</p>
<pclass="detail-example"><b>Preservation has a defined scope</b>Repeated matches for the same pathology are consolidated, preferring an affirmed patient finding. “Suspected” or “to exclude” is not automatically treated as a negation. These limits remain explicit for clinical review.</p>
</template>
</button>
</div>
<divclass="roi roi-quiet"><bclass="k">Quality gate</b>An extraction is accepted only after 100 manual reviews confirm the agreed accuracy — the same standard for every new pattern release.</div>
<divclass="out-item"><spanclass="n">01</span><div><b>Machine speed</b><spanclass="d">twenty years of letters are read by rules, not by hand; when a pattern improves, the whole archive is rebuilt from the letters themselves</span></div></div>
<divclass="out-item"><spanclass="n">02</span><div><b>Audited precision</b><spanclass="d">every release clears the same gate before production: ≥95% procedure classification, ≥85% pathology detection</span></div></div>
<divclass="out-item"><spanclass="n">03</span><div><b>Deterministic by design</b><spanclass="d">the same letter always yields the same values: no LLM sampling, no randomness, every result traceable to a versioned pattern</span></div></div>
<divclass="roi roi-quiet"><bclass="k">Quality gate</b>Manual comparison checks procedure accuracy. Pathology coverage and clinical correctness are assessed separately.</div>
</div>
<buttontype="button"class="out-panel lens-target"id="validation-advantages"data-lens="validation-advantages"data-lens-number="5"data-lens-title="The advantages"data-lens-placement="center"data-lens-caption="Validation · Advantages"aria-label="Explore the advantages: speed, precision, determinism"aria-expanded="false"aria-controls="slide-lens">
<spanclass="out-item"><spanclass="n">01</span><span><b>Machine speed</b><spanclass="d">twenty years of letters are read by rules, not by hand; when a pattern improves, the whole archive can be reprocessed</span></span></span>
<spanclass="out-item"><spanclass="n">02</span><span><b>Explicit quality targets</b><spanclass="d">≥95% procedure accuracy on a manual sample; ≥85% pathology coverage, with clinical review of the resulting cohort</span></span></span>
<spanclass="out-item"><spanclass="n">03</span><span><b>Deterministic by design</b><spanclass="d">the same letter and rule version yield the same extraction: no LLM sampling, no randomness, results traceable to versioned rules</span></span></span>
<templateclass="lens-details">
<pclass="detail-intro">Speed, precision and determinism each address a different part of making clinical text usable for research.</p>
<p><b>01 · Speed: apply a correction across the archive</b>Once defined, the extraction rules process letters automatically. An improved rule can be applied again to historical text, without manually relabelling every record.</p>
<p><b>02 · Precision: make errors visible and correctable</b>Explicit quality targets, clinical context and cohort filters make the results inspectable. The Brugada correction shows how an audit can identify a recurring error, and regression tests can protect the fix. Quality still needs measurement and clinical review.</p>
<p><b>03 · Determinism: reproduce the extraction</b>With the same input, rule version and configuration, extraction produces the same findings. There is no language-model sampling at this step. Versioning explains why a result changes when the rules change.</p>
</template>
</button>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">08 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">08 / 13</span></footer>
<asideclass="notes">
[DRAFT — Sara] How do we know the extraction is right? Not by faith. Before any pattern release goes to production it passes agate: on 100 manually reviewed records, procedure classification must reach 95% accuracy, pathology detection 85%. The Brugada cohort itself was reviewed by clinicians on dashboards — and that review fed back into the library: version 1.3.1 alone removed roughly 900 false positives.
And we never throw information away: negated and family findings are stored, only filtered when the research mart needs them. Sara: adjust the numbers/wording to what you are comfortable presenting.
<p><buttontype="button"data-note-lens="validation-quality">1. Quality gates.</button> We set a 95% procedure-accuracy target, against 100 manual reviews. The 85% pathology target measures coverage, not diagnostic accuracy.</p>
<p><buttontype="button"data-note-lens="validation-review">2. Clinical criteria.</button> Finding a disease name in a letter is only the first step. We apply explicit clinical criteria to decide which patients belong in the analysis.</p>
<p><buttontype="button"data-note-lens="validation-feedback">3. Feedback loop.</button> Version 1.3.1 corrected missed negations in roughly 900 Brugada records, with regression tests protecting the fix.</p>
<p><buttontype="button"data-note-lens="validation-context">4. Nothing is silently dropped.</button> Negated and family findings retain their context. Research datasets filter them explicitly, so preserving a finding does not mean counting it as the patient’s diagnosis.</p>
<p><buttontype="button"data-note-lens="validation-advantages">5. The advantages.</button> We can reprocess the archive quickly, inspect the rules behind each result, and reproduce the same extraction with the same rules.</p>
<divclass="screenshot-tour"aria-label="AritmoLab tour, screenshots 1 to 7">
<buttontype="button"data-screenshot="home"data-screenshot-number="1"data-screenshot-title="Home page"aria-haspopup="dialog"><imgsrc="screenshots/1-HomePage.png"alt=""loading="lazy"><span><b>1</b> Home page</span></button>
<pclass="hint">[MP: 6-7 real screenshots of the portal, anonymized or demo data]</p>
<pclass="hint">Select a screen to enlarge · Follow the tour from 1 to 7</p>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">09 / 14</span></footer>
<asideclass="notes">[DRAFT] A quick tour of what AritmoLab looks like today — six screenshots: patient profile, genetics, dashboards, the 360-degree view. What began as four islands is now one platform the Unit uses every day. MP: replace with real screenshots and a spoken tour.</aside>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">09 / 13</span></footer>
<asideclass="notes">
<p><buttontype="button"data-note-screenshot="home">1. Home page.</button> The home page summarises around 57,000 patients, by age, sex and geographical origin. It is the starting point for exploring the archive.</p>
<p><buttontype="button"data-note-screenshot="patients">2. Patient list.</button> Search filters help us find a patient or study participant, open their record, or export the results.</p>
<p><buttontype="button"data-note-screenshot="profile">3. Patient profile.</button> The profile brings demographic and clinical fields together, with access to procedures, devices, diagnostic examinations and genetics.</p>
<p><buttontype="button"data-note-screenshot="history">4. Clinical history.</button> A dated timeline brings together clinical notes, discharge letters and procedures, retaining the source of each event.</p>
<p><buttontype="button"data-note-screenshot="procedure">5. Procedure details.</button> Here, an ablation record shows the treated arrhythmias, procedural details, recorded complications and conclusions.</p>
<p><buttontype="button"data-note-screenshot="dashboards">6. Dashboard catalogue.</button> We then move from individual records to dashboards covering departmental activity, procedures, devices and genetics.</p>
<p><buttontype="button"data-note-screenshot="brugada">7. Brugada dashboard.</button> The funnel separates text mentions, filtered mentions, confirmed cases and ablation outcomes. Other charts describe sex, age, annual diagnoses and the timing of pre- and post-assessments.</p>
</aside>
</section>
<!-- ============ 11 · CRISIS ============ -->
<sectionclass="arit">
<!-- ============ 10 · CRISIS ============ -->
<sectionclass="arit crisis">
<headerclass="head">
<imgsrc="logo.png"alt="">
<spanclass="head-org">AritmoLab · Policlinico San Donato</span>
@@ -429,52 +543,54 @@
<divclass="kicker">All good? Not yet.</div>
<h2>The warehouse speaks SQL. Research needs more.</h2>
<divclass="hgap"></div>
<ulclass="points">
<li><b>Clinicians and researchers don't write SQL</b> — a star schema is built for analysts, not for the ward</li>
<li><b>Health Intelligence</b> — to present the past, the Unit needs ready-made aggregates: volumes, incidence, outcomes</li>
<li><b>Predictive research</b> — multivariate analysis and ML models need clean, wide, cohort-shaped tables</li>
<li><b>Hand-made datamarts are the bottleneck</b> — even with coding agents, each one takes hours of careful work</li>
</ul>
<divclass="crisis-points"aria-label="Four barriers between the warehouse and research">
<buttontype="button"class="oitem lens-target"id="crisis-language"data-lens="crisis-language"data-lens-number="1"data-lens-title="From clinical question to SQL"data-lens-placement="center"data-lens-caption="The research gap · 01 / 04"aria-expanded="false"aria-controls="slide-lens"><spanclass="n">01</span><span><b>Clinical questions need translation</b><spanclass="d">Clinicians define the question; a query must express it across linked tables.</span></span>
<templateclass="lens-details">
<pclass="detail-intro">Knowing what to ask is different from knowing how the database stores the answer.</p>
<p><b>The clinical question</b>“How many patients with confirmed Brugada underwent an ablation?” requires agreed definitions of the cohort and procedure.</p>
<p><b>The engineering task</b>SQL must connect diagnoses, patients and procedures, apply the time window and count each patient once, even when several records describe the same person.</p>
<pclass="detail-example"><b>The bridge</b>Clinical expertise defines the meaning. A reviewed query turns that meaning into an explicit, checkable selection.</p>
</template>
</button>
<buttontype="button"class="oitem lens-target"id="crisis-intelligence"data-lens="crisis-intelligence"data-lens-number="2"data-lens-title="Health Intelligence: describe what happened"data-lens-placement="center"data-lens-caption="The research gap · 02 / 04"aria-expanded="false"aria-controls="slide-lens"><spanclass="n">02</span><span><b>Health Intelligence needs shared definitions</b><spanclass="d">Volumes, diagnoses and outcomes need consistent groups, periods and denominators.</span></span>
<templateclass="lens-details">
<pclass="detail-intro">A dashboard needs agreed indicators, not simply a chart drawn over raw records.</p>
<p><b>Define what is counted</b>Patients, admissions and procedures answer different questions. For an outcome percentage, specify which patients are eligible and the observation period.</p>
<p><b>Prepare comparable summaries</b>Aggregate by year, procedure or patient group using the same definitions. Keep missing information visible so that changes in documentation are not mistaken for changes in care.</p>
<pclass="detail-example"><b>Example</b>Annual ablation volumes describe activity. An outcome percentage also needs a defined denominator and follow-up window.</p>
</template>
</button>
<buttontype="button"class="oitem lens-target"id="crisis-prediction"data-lens="crisis-prediction"data-lens-number="3"data-lens-title="Predictive research: build the study table"data-lens-placement="center"data-lens-caption="The research gap · 03 / 04"aria-expanded="false"aria-controls="slide-lens"><spanclass="n">03</span><span><b>Predictive research needs a study dataset</b><spanclass="d">A defined cohort, consistent variables and outcomes measured over an agreed period.</span></span>
<templateclass="lens-details">
<pclass="detail-intro">Linked clinical records must become a table shaped around the study question.</p>
<p><b>Define one row</b>Choose the unit of analysis: for example, one patient or one procedure. Place the selected characteristics in columns, with consistent units and explicit handling of missing values.</p>
<p><b>Respect the timeline</b>Define when prediction would occur and when the outcome is assessed. Predictor variables must contain only information available at that prediction time.</p>
<pclass="detail-example"><b>Example</b>To study outcomes after ablation, separate pre-procedure characteristics from later observations. This prevents future information from leaking into the prediction.</p>
</template>
</button>
<buttontype="button"class="oitem lens-target"id="crisis-datamarts"data-lens="crisis-datamarts"data-lens-number="4"data-lens-title="Datamarts: the preparation bottleneck"data-lens-placement="center"data-lens-caption="The research gap · 04 / 04"aria-expanded="false"aria-controls="slide-lens"><spanclass="n">04</span><span><b>Hand-made datamarts are the bottleneck</b><spanclass="d">Each question requires selection, joins, checks and a reproducible dataset.</span></span>
<templateclass="lens-details">
<pclass="detail-intro">A datamart is a focused dataset prepared for a particular analysis.</p>
<p><b>More than writing SQL</b>Someone must agree the cohort, connect the sources, resolve duplicate records and check missing values and patient counts. A query can run successfully and still answer the wrong question.</p>
<p><b>Make the work repeatable</b>Keep the selection rules, query and checks together, so the dataset can be rebuilt when the data or study definition changes.</p>
<pclass="detail-example"><b>Where assistance helps</b>AI can draft the query. Clinicians and engineers still review its meaning and results before using the dataset.</p>
</template>
</button>
</div>
<divclass="roi"><bclass="k">The gap</b>Twenty years of data, one warehouse — and no fast road from a research question to an answer.</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">10 / 14</span></footer>
<asideclass="notes">[DRAFT] And here is the crisis of our story. The warehouse holds everything — but it speaks SQL. A clinician cannot query a star schema between two visits, and a researcher cannot build a study on raw tables. What research needs are datamarts: ready-made views for Health Intelligence, and clean cohort tables for multivariate analysis and machine learning. And building them by hand — even with coding agents helping — takes hours of careful, repetitive work for every single study. This is where our story got stuck.</aside>
</section>
<sectionclass="arit">
<headerclass="head">
<imgsrc="logo.png"alt="">
<spanclass="head-org">AritmoLab · Policlinico San Donato</span>
</header>
<navclass="band">
<spanclass="band-section">ThothII</span>
</nav>
<divclass="sbody">
<divclass="kicker">Role 3 — AI for analysis</div>
<h2>Ask the warehouse in plain English</h2>
<divclass="hgap"></div>
<divclass="quote">“How many Brugada patients had an effective ablation?”
<small>a researcher, in natural language</small>
</div>
<divclass="pipeline">
<divclass="stage ai"><b>NL → SQL</b><span>AI writes the query over the star schema</span><spanclass="tag">AI</span></div>
<divclass="arrow">→</div>
<divclass="stage"><b>Human review</b><span>the researcher checks every proposed query</span></div>
<divclass="stage"><b>Superset</b><span>statistics & dashboards</span></div>
</div>
<pclass="stat-note">The session is a conversation: refine the question, iterate on the datamart — every step reviewable.</p>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">11 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">10 / 13</span></footer>
<asideclass="notes">
[DRAFT] Third role: AI for the analysis itself. This is ThothII, our natural-language layer over the warehouse. A researcher asks "how many Brugada patients had an effective ablation?" in plain English; the AI writes the SQL over the star schema, proposes a datamart, and the researcher reviews the query before it runs. The result lands in Superset as statistics and dashboards.
The human stays in the loop — the AI proposes, the researcher approves. This is where the demo will go live in a couple of slides.
<p><strong>0. What is SQL?</strong><br>SQL means Structured Query Language. It tells a database what to select, connect and count.</p>
<p><buttontype="button"data-note-lens="crisis-language">1. Clinical questions.</button><br>“How many patients with confirmed Brugada underwent an ablation?” We must define confirmation and count each patient once.</p>
<p><buttontype="button"data-note-lens="crisis-intelligence">2. Health Intelligence.</button><br>Dashboards need agreed definitions, time periods and denominators to make comparisons meaningful.</p>
<p><buttontype="button"data-note-lens="crisis-prediction">3. Predictive research.</button><br>Study tables separate characteristics known before prediction from outcomes observed afterwards.</p>
<p><buttontype="button"data-note-lens="crisis-datamarts">4. Datamarts.</button><br>A datamart is an analysis dataset with explicit selection rules and repeatable checks.</p>
<p><strong>5. The gap.</strong><br>The gap is between clinical meaning and database instructions: storing data does not automatically make a question answerable.</p>
</aside>
</section>
<!-- ============ 13 · THOTHII TOUR ============ -->
<!-- ============ 11 · THOTHII TOUR ============ -->
<sectionclass="arit">
<headerclass="head">
<imgsrc="logo.png"alt="">
@@ -492,10 +608,10 @@
</div>
<pclass="hint">[MP: 6-7 real screenshots of ThothII — ask, review SQL, datamart, dashboard]</p>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">12 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">11 / 13</span></footer>
<asideclass="notes">[DRAFT] Let me walk you through ThothII: the researcher asks the question in plain English; the AI proposes the SQL; the query is reviewed; the datamart is assembled; the dashboard comes alive. Six screens, a few minutes. MP: real screenshots.</aside>
<pclass="hint">[Illustrative data, modeled on published arrhythmology predictors — Brugada focus]</p>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">13 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">12 / 13</span></footer>
<asideclass="notes">[DRAFT — next step: two SVG charts, illustrative data modeled on real literature] And this is the happy ending. From one datamart, the Unit can run a multivariate analysis — which factors truly drive arrhythmic risk in Brugada patients — and train machine-learning models on the same table. What used to take weeks of manual data preparation now takes minutes. The AI did not replace the researcher: it gave the researcher back their time.</aside>
</section>
@@ -535,7 +651,7 @@
<span>Dr. Sara Paratico - I.R.C.C.S. Policlinico San Donato</span>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">14 / 14</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">13 / 13</span></footer>
<asideclass="notes">
[DRAFT] Thank you for your attention. If you want to go deeper — the text miner, the mapping agents, ThothII — we are publishing the technical walkthroughs as videos on our Substack; the link is on the final version of this deck. And now, happy to take your questions.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.