<pclass="event">“Multidimensional Characterization of Cardiac Arrhythmias: Role of Electrocardiology in the Artificial Intelligence Era”</p>
<pclass="event">“Multidimensional Characterization of Cardiac Arrhythmias: Role of Electrocardiology in the Artificial Intelligence Era”</p>
<pclass="event-where">San Donato Milanese, Milan, Italy · 2–3 October 2026</p>
<pclass="event-where">San Donato Milanese, Milan, Italy · 2–3 October 2026</p>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">01 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">01 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
Good morning. My name is Marco Pancotti, and I led the part of the PAMP-FA project dedicated to building the AritmoLab portal at Policlinico San Donato, with support from Sara Paratico, who will co-present with me today.
Good morning. My name is Marco Pancotti, and I led the part of the PAMP-FA project dedicated to building the AritmoLab portal at Policlinico San Donato, with support from Sara Paratico, who will co-present with me today.
The scope of the project included a data warehouse and tools for machine learning and predictive statistics to support research by the Arrhythmology Unit, directed by Professor Pappone and Professor Locati.
The scope of the project included a data warehouse and tools for machine learning and predictive statistics to support research by the Arrhythmology Unit, directed by Professor Pappone and Professor Locati.
@@ -46,55 +46,7 @@
</section>
</section>
<sectionclass="arit">
<headerclass="head">
<imgsrc="logo.png"alt="">
<spanclass="head-org">AritmoLab · Policlinico San Donato</span>
</header>
<navclass="band">
<spanclass="band-section">Where we started</span>
</nav>
<divclass="sbody">
<h2>Where we started - Four islands and four missing pieces</h2>
<divclass="todo"data-missing="ci"style="margin-left:0"><spanclass="q">?</span><spanclass="tx"><b>Clinical Intelligence</b><span>dashboards on the clinical history of the Unit and its patients</span></span></div>
<divclass="todo"data-missing="dwh"style="margin-left:9%"><spanclass="q">?</span><spanclass="tx"><b>Datawarehouse</b><span>the single source of truth</span></span></div>
<divclass="todo"data-missing="ml"style="margin-left:4%"><spanclass="q">?</span><spanclass="tx"><b>ML-ready data</b><span>cohort tables, ready for model training</span></span></div>
<divclass="todo"data-missing="portal"style="margin-left:12%"><spanclass="q">?</span><spanclass="tx"><b>Professional Portal</b><span>AritmoLab — the grown-up Omics Portal</span></span></div>
<spanclass="nm">Cardioref</span><spanclass="sc"style="width:300px">visits · operations · reports — twenty years of management on structured forms</span></button>
<pclass="hint"style="position:absolute; left:26%; width:48%; bottom:2px; margin:0; text-align:center">▸ click a system: how it was used — and what held it back · click a missing piece: what it would have given</p>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">02 / 13</span></footer>
<asideclass="notes">
The starting point was four disconnected systems:
1 - <strong>Cardioref</strong> held twenty years of electrophysiology records and supported clinical operations, but much of the information needed for research was in free text.
2 - <strong>Genetic data</strong> were manually entered into Excel, without integration with Cardioref.
3 - The <strong>Aritmolab Portal</strong> was an early prototype, not yet integrated with the source systems.
4 - <strong>ECG</strong> data were available as paper records or CSV exports on request.
Whe lacked:
5 - some longitudinal <strong>clinical analytics</strong>
6 - a shared <strong>data warehouse</strong>,
7 - a set of structured datasets for model training,
8 - and a unified <strong>clinical portal</strong>.
</aside>
</section>
<sectionclass="arit flow-slide">
<sectionclass="arit flow-slide">
@@ -119,11 +71,11 @@
<spanclass="fbox src future"><b><spanclass="plus">+</span>Future sources</b><span>new subsystems can be connected as sources</span></span>
<spanclass="fbox src future"><b><spanclass="plus">+</span>Future sources</b><span>new subsystems can be connected as sources</span></span>
</span>
</span>
<templateclass="lens-details">
<templateclass="lens-details">
<pclass="detail-intro">The starting material offered complementary views of the patient.</p>
<pclass="detail-intro">Four disconnected systems held complementary views of the patient.</p>
<p><b>Clinical course · Cardioref</b>Visits, procedures and reports describe the course of care. Dates and narrative details give each event its clinical context.</p>
<p><b>Clinical course · Cardioref</b>Twenty years of visits, procedures and reports supported daily clinical work. Much of the information needed for research remained in free text.</p>
<p><b>Genetic findings</b>Laboratory results and variant descriptions add the genetic perspective, recorded separately from the clinical history.</p>
<p><b>Genetic findings</b>Laboratory results and variants were entered manually into Excel, without integration with Cardioref.</p>
<p><b>Electrical activity · ECG</b>Tracings and exported signals document cardiac electrical activity, complementing the written account of the patient’s condition.</p>
<p><b>Electrical activity · ECG</b>Separate device systems provided paper records or CSV exports on request, requiring manual collection.</p>
<p><b>Different forms of evidence</b>Structured fields, free text and signals must retain their clinical meaning when connected. Future sources would extend this initial set.</p>
<p><b>The early portal</b>The AritmoLab prototype, originally the Omics Portal, lacked integration with the sources. Connecting clinical history, genetics and ECG was the prerequisite for a shared patient view.</p>
</template>
</template>
</button>
</button>
<divclass="fcol ai-hit">
<divclass="fcol ai-hit">
@@ -157,9 +109,9 @@
<divclass="farrow">→</div>
<divclass="farrow">→</div>
<divclass="fcol"style="flex:1.05; justify-content:center"><buttontype="button"class="fbox lens-target"id="plan-star-schema"data-lens="plan-star-schema"data-lens-number="4"aria-label="Explore Data warehouse"aria-expanded="false"aria-controls="slide-lens"><b>Data warehouse</b><span>star schema · organized for analysis</span>
<divclass="fcol"style="flex:1.05; justify-content:center"><buttontype="button"class="fbox lens-target"id="plan-star-schema"data-lens="plan-star-schema"data-lens-number="4"aria-label="Explore Data warehouse"aria-expanded="false"aria-controls="slide-lens"><b>Data warehouse</b><span>star schema · organized for analysis</span>
<templateclass="lens-details">
<templateclass="lens-details">
<pclass="detail-intro">The shared analytical database: integrated clinical data reorganized for research across patients and over time.</p>
<pclass="detail-intro">A shared analytical database replaces separate exports and fragile spreadsheets with one source of truth.</p>
<p><b>How a star schema works</b>“Facts” represent events or measurements, such as a procedure or test. “Dimensions” describe their context, such as the patient, date and procedure type.</p>
<p><b>How a star schema works</b>“Facts” represent events or measurements, such as a procedure or test. “Dimensions” describe their context, such as the patient, date and procedure type.</p>
<p><b>Why it matters</b>Researchers can filter, group and compare records using common definitions, without reconstructing every relationship from the original hospital tables.</p>
<p><b>Clinical Intelligence</b>Common definitions make the Unit’s activity and patients’ histories comparable over time. Dashboards can show volumes, incidence and outcomes for management, audit and research.</p>
<pclass="detail-example"><b>Clinical question it can support</b>How many patients underwent a given procedure each year, and how does that distribution vary by age group?</p>
<pclass="detail-example"><b>Clinical question it can support</b>How many patients underwent a given procedure each year, and how does that distribution vary by age group?</p>
<pclass="detail-intro">Focused datasets derived from the warehouse for a specific research question, study or dashboard.</p>
<pclass="detail-intro">Reproducible cohorts provide the structured data previously missing for statistics and model training.</p>
<p><b>What they define</b>The cohort, time window, variables and level of detail: for example, one row per patient or one row per procedure.</p>
<p><b>What they define</b>The cohort, time window, variables and level of detail, such as one row per patient with features and outcomes.</p>
<p><b>How ThothII helps</b>A plain-English question becomes a proposed SQL query through a guided workflow with human review. The resulting dataset supports analysis and portal dashboards.</p>
<p><b>How ThothII helps</b>A plain-English question becomes proposed SQL through a guided workflow with human review, reducing manual dataset preparation.</p>
<pclass="detail-example"><b>Illustrative study dataset</b>Patients who underwent a drug-challenge test, with test date, result and selected clinical characteristics. Its cohort and variable definitions must be agreed before interpreting results.</p>
<p><b>The professional portal</b>AritmoLab brings integrated patient information and research dashboards into one place, extending the early Omics Portal.</p>
<pclass="detail-example"><b>Illustrative study dataset</b>Drug-challenge patients, test dates, results and clinical characteristics, with agreed cohort and variable definitions.</p>
</template>
</template>
</button>
</button>
<divclass="fcap">AI builds them on demand from plain-English questions (ThothII)</div>
<divclass="fcap">AI builds them on demand from plain-English questions (ThothII)</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">03 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">02 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
Here is the whole project on one slide.
<p>We had four goals: clinical dashboards to understand the Unit’s activity and patients’ histories, a shared data warehouse, structured cohorts for machine learning and predictive statistics, and a unified professional portal.</p>
<p>(1) We started with disconnected systems: twenty years of Cardioref records, much of their research value in free text, genetic results entered manually into Excel, ECGs on paper or exported on request, and an early portal without source integration. Each source answered a different need, but combining them for a clinical study meant collecting exports and rebuilding patient histories by hand.</p>
On the left, what already existed, three hospital data sources:
<p>(2) To connect them, staging preserves a working copy of the sources.</p>
(1)<strong>What existed</strong>: <strong>Cardioref</strong>, our electrophysiology records dataset, the <strong>genetic data</strong> coming from the labs, the <strong>ECG signals</strong>, and, in the future, other sources.
<p>(3) Integration then cleans, normalizes and links the records.</p>
<p>(4) The data warehouse organizes events and their context for analysis.</p>
On the right, what we built in AritmoLab:
<p>(5) From there, datamarts turn that shared information into research datasets and dashboards within AritmoLab.</p>
(2) a staging copy of all relevant Cardioref data, (3) an integration layer where the data is cleaned and normalized, (4) a star-schema warehouse, where the data are reorganized as facts and dimensions, and (5) research datamarts, generated from the data warehouse and presented through dashboards embedded in the portal.
<p>(6) The brain symbols show where AI contributes. It drafts the mappings, which humans review.</p>
<p>(7) During integration, AI extracts structured information from clinical text.</p>
Everywhere you see the brain symbol, that's where AI works for us.
<p>(8) ThothII then builds datamarts from plain-English questions, with human approval.</p>
6 - The <strong>mappings</strong> that bring the sources in were themselves drafted by AI, then revised by humans
<p>Open-source tools support the platform, and AI coding agents helped throughout development. The popups connect each component to the gap it addresses.</p>
7 - At <strong>integration</strong>, AI reads the clinical text.
<p>Next, we turn to the hardest part: unstructured data. Sara will explain how we extract clinical meaning and check its reliability.</p>
8 - At the end, AI builds the <strong>datamarts</strong> on demand from plain-English questions.
That was the plan. And to build it, we used AI everywhere it helped — drafting configs, reading clinical text, and at the very end, producing the datamarts themselves.
Next, we’ll look at the hardest part: unstructured data. My colleague, Sara Paratico,do will show you exactly how the text reading works. But first, the first step of the climb: the mappings.
</aside>
</aside>
</section>
</section>
@@ -268,10 +217,12 @@
</div>
</div>
</div>
</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">04 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">03 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
Here is one real sentence from a discharge letter — in Italian, as the clinicians wrote it. In this single sentence there is a diagnosis (syncope), an ECG finding (ST elevation in V1–V3), and a drug-challenge result (flecainide positive for Brugada pattern). For a human cardiologist this is readable in two seconds. For a database, this is just a blob of text in a column.
<p>Over twenty years, we have collected clinical information in discharge letters and procedure reports. Much remains in free text.</p>
The structured tables — demographics, procedures, dates — only tell half the story. The rest is locked inside these free-text fields, in every hospital system we have.
<p>Clinicians connect these details and form hypotheses. Here, one sentence combines syncope, ECG changes and a positive flecainide test. Research needs defined variables that preserve this meaning.</p>
<p>Our text miner extracts them in the integration layer, before the data enter the warehouse.</p>
<p>We face four challenges: capturing diverse clinical histories, interpreting context correctly, mapping different expressions to common terms, and validating accuracy. Reliable research and patient care depend on getting these details right.</p>
</aside>
</aside>
</section>
</section>
@@ -314,11 +265,11 @@
<divclass="nitem"><b>Audited.</b> A new rule set goes live only after it proves ≥95% accuracy on 100 records reviewed by hand.</div>
<divclass="nitem"><b>Audited.</b> A new rule set goes live only after it proves ≥95% accuracy on 100 records reviewed by hand.</div>
</div>
</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">05 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">04 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
[DRAFT — Sara] I'll take you inside these numbers. We read 58,438 discharge letters — every single one, every night. From them the text miner extracted 73,389 pathology records, classified into our clinical ontology. It parsed 10,908 drug-challenge tests, distinguishing the therapy from the actual test. And at the end of the chain: 2,307 patients with a confirmed Brugada pattern — a cohort nobody could have built by hand.
<p>Our text miner uses explicit clinical rules to identify diagnoses and procedures, including catheter ablation and drugchallenge tests, while checking their context.</p>
This is not a one-off migration: the pipeline runs every night, and today it processes on average two hundred and forty letters a month. Four simple pieces do the work: a letter reader for the Italian and English text, a clinical matcher that recognises diagnoses, procedures and test outcomes, a context guard that keeps negations and family history apart, and the ontology sorter that files every finding into its clinical category.
<p>These figures cover over 58,000 letters and around 73,000 records of clinical conditions. We identified almost 11,000 drug challenge test entries, including positive Brugada tests in 2,307 patients.</p>
The key word here is deterministic: no black box. Every number can be traced back to the rule that produced it.
<p>Each extracted record retains the rule version used. This lets us compare classifications with clinical review and investigate errors. For procedure classification, our target is at least 95 percent accuracy against 100 manually reviewed records.</p>
</aside>
</aside>
</section>
</section>
@@ -367,10 +318,11 @@
</div>
</div>
</div>
</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">06 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">05 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
[DRAFT — Sara] Why can't we just search for "FA"? Because clinical text lies to naive search. "Fibrillazione atriale esclusa" contains the words of a diagnosis but negates it — so every rule scans a window around the match, looking for negation cues. "Padre con FA" is real atrial fibrillation — but in the father, not the patient: we record it as family history. Letters mix Italian and English, abbreviations collide ("TA" is blood pressure, not a therapy), and one field can contain five different statements — so we split the text into clauses first.
<p>A report may describe atrial fibrillation in the patient, rule it out, or mention it in the father's history. The same term must lead to different classifications.</p>
Each of these problems has a specific, versioned solution. Sara to expand with real corpus examples.
<p>Our analyser checks for negation and family references, recording these attributes separately. It recognises Italian and English terms and common abbreviations. For example, "TA" may mean atrial tachycardia, but followed by a blood pressure value, it should not trigger that diagnosis.</p>
<p>When reports describe several events, the system separates the text into clauses. For drug challenge tests, this helps distinguish a positive result before ablation from a negative result afterwards, preserving each observation and its clinical context.</p>
</aside>
</aside>
</section>
</section>
@@ -410,12 +362,12 @@
</div>
</div>
</div>
</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">07 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">06 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
[DRAFT — Sara] Extraction needs a target vocabulary — that's the ontology. Tier 1 holds the eleven arrhythmological categories that matter most for our research; Tier 2 holds five structural conditions. The order of the rules is itself clinical knowledge: Brugada patterns are tested before TV, because "substrato per TV" often appears in Brugada reports and would otherwise mask the diagnosis.
<p>Our clinical ontology, a shared set of terms, helps identify arrhythmia-related diagnoses in the narrative.</p>
The ontology is versioned like software: a semantic version for the pattern library, stamped on every extracted row. When we add a synonym or fix a rule, the change is traceable — and the data can be rebuilt.
<p>The first tier covers conditions such as atrial fibrillation and Brugada. If none are found within a text field, the system checks a second tier for cardiovascular comorbidities, such as cardiomyopathy or heart failure, that may influence the patient's arrhythmic presentation.</p>
And this is the output of the whole work: reading a letter writes structured, quantitative values back onto the records that describe the procedures and implants it contains — and those records enter the warehouse generation exactly as if the clinicians had typed them during the visits. Text becomes data, indistinguishable from bedside data entry.
<p>Different expressions map to one concept: for example, "fibrillazione atriale" and "atrial fibrillation" receive the same label.</p>
Sara: review the category list and the Italian labels before final. Labels are now English-first (audience is mostly foreign) — please confirm the EN terms, especially PSVT as the umbrella for WPW/TPSV and AT/AVNRT for TA/TRN.
<p>The extracted records enter the warehouse alongside data clinicians entered in structured fields. We can then define patient groups and research datasets, retaining each finding's source and the version of the extraction rules.</p>
</aside>
</aside>
</section>
</section>
@@ -486,13 +438,13 @@
</button>
</button>
</div>
</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">08 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">07 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
<p><buttontype="button"data-note-lens="validation-quality">1. Quality gates.</button> We set a 95% procedure-accuracy target, against 100 manual reviews. The 85% pathology target measures coverage, not diagnostic accuracy.</p>
<p>(1) To check the extraction, we set a 95% procedure-accuracy target against 100 manual reviews. The 85% pathology target measures coverage, not diagnostic accuracy.</p>
<p><buttontype="button"data-note-lens="validation-review">2. Clinical criteria.</button> Finding a disease name in a letter is only the first step. We apply explicit clinical criteria to decide which patients belong in the analysis.</p>
<p>(2) Finding a disease name in a letter is only the first step. We also apply explicit clinical criteria to decide which patients belong in the analysis.</p>
<p><buttontype="button"data-note-lens="validation-feedback">3. Feedback loop.</button> Version 1.3.1 corrected missed negations in roughly 900 Brugada records, with regression tests protecting the fix.</p>
<p>(3) Review feeds back into the rules. Version 1.3.1 corrected missed negations in roughly 900 Brugada records, with regression tests protecting the fix.</p>
<p><buttontype="button"data-note-lens="validation-context">4. Nothing is silently dropped.</button> Negated and family findings retain their context. Research datasets filter them explicitly, so preserving a finding does not mean counting it as the patient’s diagnosis.</p>
<p>(4) Throughout this process, negated and family findings retain their context. Research datasets filter them explicitly, so preserving a finding does not mean counting it as the patient’s diagnosis.</p>
<p><buttontype="button"data-note-lens="validation-advantages">5. The advantages.</button> We can reprocess the archive quickly, inspect the rules behind each result, and reproduce the same extraction with the same rules.</p>
<p>(5) This lets us reprocess the archive quickly, inspect the rules behind each result, and reproduce the same extraction with the same rules.</p>
</aside>
</aside>
</section>
</section>
@@ -508,26 +460,29 @@
<divclass="sbody">
<divclass="sbody">
<divclass="kicker">The portal, today</div>
<divclass="kicker">The portal, today</div>
<h2>AritmoLab — a quick tour</h2>
<h2>AritmoLab — a quick tour</h2>
<divclass="screenshot-tour"data-screenshot-auto-openaria-label="AritmoLab tour, screenshots 1 to 7">
<divclass="screenshot-tour screenshot-tour-four"data-screenshot-auto-openaria-label="AritmoLab tour, screenshots 1 to 4">
<buttontype="button"data-screenshot="home"data-screenshot-number="1"data-screenshot-title="Home page"aria-haspopup="dialog"><imgsrc="screenshots/1-HomePage.png"alt=""loading="lazy"><span><b>1</b> Home page</span></button>
<buttontype="button"data-screenshot="home"data-screenshot-number="1"data-screenshot-title="Home page"aria-haspopup="dialog"><imgsrc="screenshots/1-HomePage.png"alt=""loading="lazy"><span><b>1</b> Home page</span></button>
<buttontype="button"data-screenshot="brugada"data-screenshot-number="4"data-screenshot-title="Brugada: one of many dashboards"aria-haspopup="dialog"><imgsrc="screenshots/7-Brugada-dashboard.png"alt=""loading="lazy"><span><b>4</b>Brugada: one of many dashboards</span></button>
<pclass="hint">Select a screen to enlarge · Follow the tour from 1 to 7</p>
<pclass="hint">Select a screen to enlarge · Follow the tour from 1 to 4</p>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">09 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">08 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
<p><buttontype="button"data-note-screenshot="home">1. Home page.</button> The home page summarises around 57,000 patients, by age, sex and geographical origin. It is the starting point for exploring the archive.</p>
<p><buttontype="button"data-note-screenshot="patients">2. Patient list.</button> Search filters help us find a patient or study participant, open their record, or export the results.</p>
<p>(1) AritmoLab is an extensive portal with many interconnected features. Our limited time prevents us from presenting it in detail, so we will show just four screens. Sara Paratico and I are available for a more in-depth presentation on request, either during the conference or afterwards through a remote connection.</p>
<p><buttontype="button"data-note-screenshot="profile">3. Patient profile.</button> The profile brings demographic and clinical fields together, with access to procedures, devices, diagnostic examinations and genetics.</p>
<p>The home page summarises around 57,000 patients by age, sex and geographical origin. It is the starting point for exploring the archive.</p>
<p><buttontype="button"data-note-screenshot="history">4. Clinical history.</button> A dated timeline brings together clinical notes, discharge letters and procedures, retaining the source of each event.</p>
</div></div>
<p><buttontype="button"data-note-screenshot="procedure">5. Procedure details.</button> Here, an ablation record shows the treated arrhythmias, procedural details, recorded complications and conclusions.</p>
<p><buttontype="button"data-note-screenshot="dashboards">6. Dashboard catalogue.</button> We then move from individual records to dashboards covering departmental activity, procedures, devices and genetics.</p>
<p>(2) From this overview, we can open a patient profile, which brings demographic and clinical information together, with access to procedures, devices, diagnostic examinations and genetics. It provides a single starting point for exploring an individual patient's record.</p>
<p><buttontype="button"data-note-screenshot="brugada">7. Brugada dashboard.</button> The funnel separates text mentions, filtered mentions, confirmed cases and ablation outcomes. Other charts describe sex, age, annual diagnoses and the timing of pre- and post-assessments.</p>
<p>(3) Moving from individual patient records to an overview of the department, the dashboard catalogue gives us access to clinical activity, procedures, devices and genetics.</p>
<p>(4) Among the many dashboards we could present, we chose Brugada as an example. The funnel separates text mentions, filtered mentions, confirmed cases and ablation outcomes. Other charts show sex, age, annual diagnoses and the timing of pre- and post-assessments.</p>
</div></div>
</aside>
</aside>
</section>
</section>
<!-- ============ 10 · CRISIS ============ -->
<!-- ============ 10 · CRISIS ============ -->
@@ -579,14 +534,14 @@
</div>
</div>
<divclass="roi"><bclass="k">The gap</b>Twenty years of data, one warehouse — and no fast road from a research question to an answer.</div>
<divclass="roi"><bclass="k">The gap</b>Twenty years of data, one warehouse — and no fast road from a research question to an answer.</div>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">10 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">09 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
<p><strong>0. What is SQL?</strong><br>SQL means Structured Query Language. It tells a database what to select, connect and count.</p>
<p>The warehouse speaks SQL, which means Structured Query Language. It tells a database what to select, connect and count.</p>
<p><buttontype="button"data-note-lens="crisis-language">1. Clinical questions.</button><br>“How many patients with confirmed Brugada underwent an ablation?” We must define confirmation and count each patient once.</p>
<p>(1) To answer a clinical question such as “How many patients with confirmed Brugada underwent an ablation?”, we must define confirmation and count each patient once.</p>
<p><buttontype="button"data-note-lens="crisis-intelligence">2. Health Intelligence.</button><br>Dashboards need agreed definitions, time periods and denominators to make comparisons meaningful.</p>
<p>(2) The same need for clarity applies to Health Intelligence: dashboards need agreed definitions, time periods and denominators to make comparisons meaningful.</p>
<p><buttontype="button"data-note-lens="crisis-prediction">3. Predictive research.</button><br>Study tables separate characteristics known before prediction from outcomes observed afterwards.</p>
<p>(3) For predictive research, study tables separate characteristics known before prediction from outcomes observed afterwards.</p>
<p><buttontype="button"data-note-lens="crisis-datamarts">4. Datamarts.</button><br>A datamart is an analysis dataset with explicit selection rules and repeatable checks.</p>
<p>(4) A datamart provides an analysis dataset with explicit selection rules and repeatable checks.</p>
<p><strong>5. The gap.</strong><br>The gap is between clinical meaning and database instructions: storing data does not automatically make a question answerable.</p>
<p>The gap is therefore between clinical meaning and database instructions: storing data does not automatically make a question answerable.</p>
</aside>
</aside>
</section>
</section>
@@ -600,136 +555,53 @@
<spanclass="band-section">A tour of ThothII</span>
<spanclass="band-section">A tour of ThothII</span>
</nav>
</nav>
<divclass="sbody">
<divclass="sbody">
<divclass="kicker">ThothII, step by step</div>
<divclass="kicker">Human in the Loop · Eight phases, six screens</div>
<h2>From question to datamart — the app</h2>
<h2>From clinical question to datamart, with human review</h2>
<divclass="screenshot-tour screenshot-tour-compact"data-screenshot-auto-openaria-label="ThothII tour, screenshots 1 to 12">
<divclass="screenshot-tour screenshot-tour-compact"data-screenshot-auto-openaria-label="ThothII tour, screenshots 1 to 6">
<buttontype="button"data-screenshot="thothii-disambiguation-1"data-screenshot-number="2"data-screenshot-title="Agreeing on the question"aria-haspopup="dialog">
<buttontype="button"data-screenshot="thothii-schema-final"data-screenshot-number="3"data-screenshot-title="Connecting to the data"aria-haspopup="dialog">
<imgsrc="screenshots/thothii/11-FinalStep.png?v=5"alt=""loading="lazy"><span><b>11</b> Final step</span>
</button>
<buttontype="button"data-screenshot="thothii-return"data-screenshot-number="12"data-screenshot-title="Back to start"aria-haspopup="dialog">
<imgsrc="screenshots/thothii/12-BackToStartingPoint.png?v=5"alt=""loading="lazy"><span><b>12</b> Back to start</span>
</button>
</button>
</div>
</div>
<pclass="hint">Select a screen to enlarge · Follow the tour from 1 to 12</p>
<pclass="hint"><strong>AI proposes. People review, correct and approve.</strong> · Select a screen to enlarge</p>
</div>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">11 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">10 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
<divdata-popup-notes="thothii-start"><h3>From a research question to SQL</h3><divdata-popup-notes-body><p>The task has changed. We are no longer extracting structured, coded values from the text within clinical records. We are now starting from a natural-language request to retrieve the records that contain the relevant values.</p><p>Research questions are often difficult to express clearly and unambiguously, and their clinical concepts do not map directly to the underlying database structure. As a result, writing the right SQL query can take hours, with repeated testing and refinement before it accurately reflects the intended research question.</p></div></div>
<divdata-popup-notes="thothii-disambiguation-1"><h3>Clarifying the question with the researcher</h3><divdata-popup-notes-body>
<p>(1) ThothII combines a wide range of capabilities in a guided, eight-phase workflow. A full presentation deserves at least thirty minutes. Here, we focus on the main steps.</p>
<p><strong>1. Workflow phase.</strong> At the top, the phase indicator shows where we are in the process. Here, phase one is highlighted: clarifying the research question.</p>
<p>It helps researchers turn clinical questions into checked SQL and reusable analysis datasets, bridging clinical meaning and database structure. AI proposes, and people review, correct and approve. This is Human in the Loop throughout the workflow, combining clinical judgment and database expertise.</p>
<p><strong>2. AI working time.</strong> The timer measures how long the AI has been working, excluding the time the human reviewer spends considering the options and making decisions.</p>
<p><strong>3. Model activity.</strong> On the left, we can follow the AI's running explanation of its analysis: the interpretations it considers, the evidence it consults, and the rationale for its proposals.</p>
<p><strong>4. Proposed interpretations.</strong> In the centre, the AI asks the researcher to choose between alternative interpretations. Here, the question is how to identify an ablation for atrial fibrillation in the database. This choice determines which patients enter the study.</p>
<p><strong>5. Reviewer actions.</strong> The reviewer can select a proposed option, or use <strong>Go back</strong> to revisit the previous step, <strong>Exit</strong> to stop the current workflow, or <strong>Other — specify</strong> to describe a different interpretation in free text.</p>
</div></div>
</div></div>
<divdata-popup-notes="thothii-disambiguation-final"><h3>From clarification to an agreed study question</h3><divdata-popup-notes-body>
<p><strong>Bringing the decisions together.</strong> The AI combines the original request with the researcher's answers into an explicit study definition. It records what we mean by each clinical concept, which patients to include, the time window, and the result we want.</p>
<p>(2) We begin by clarifying the question: here, what counts as an atrial fibrillation ablation? We then review relevant knowledge from earlier work before approving a precise reformulation: how many patients had their first recorded AF ablation between 2020 and 2023? We find the first procedure across the available history before filtering dates. The reviewer can challenge the interpretation.</p>
<p><strong>The clarified question.</strong> In this example: how many distinct patients had their first recorded ablation for atrial fibrillation between 2020 and 2023? We identify the procedure from the treated pathology, and find each patient's first AF ablation across the entire available database history before applying the date filter.</p>
<p><strong>Making the interpretation actionable.</strong> The summary links these choices to the relevant database fields and selection rules, and displays the checks completed during clarification. These agreed criteria guide the subsequent schema linking and SQL construction.</p>
<p><strong>Human confirmation.</strong> The researcher reviews the consolidated definition. <strong>Save and proceed</strong> confirms it and closes phase one; <strong>Reject</strong> sends it back for revision.</p>
</div></div>
</div></div>
<divdata-popup-notes="thothii-schema-1"><h3>Finding the data needed to answer the question</h3><divdata-popup-notes-body>
<p><strong>Linking clinical meaning to data.</strong> We are now in phase four, schema linking: identifying where the information needed to answer the agreed question is stored. Tables group related records; columns hold specific details, such as a patient identifier, a procedure date, or the condition treated.</p>
<p>(3) With the question agreed, we connect its concepts to tables, fields and relationships. We then review the mapping and selection rules together before building SQL. The reviewer checks which patients, procedures and dates will count.</p>
<p><strong>Proposed tables.</strong> The AI presents candidate tables and explains why each is relevant. Here, one table provides the ablation episodes and their dates; another identifies the pathology treated, allowing us to distinguish atrial fibrillation ablations from other procedures.</p>
<p><strong>Selecting the columns.</strong> The column controls let the reviewer inspect and adjust the fields selected for the query. We need the information required to identify patients, recognise the relevant procedures, and apply the agreed time criteria.</p>
<p><strong>Review before proceeding.</strong> The researcher and a technical reviewer can check this mapping together, adjust the selection, and confirm it. This establishes which data will support the SQL query; the relationships between the selected tables are reviewed next.</p>
</div></div>
</div></div>
<divdata-popup-notes="thothii-schema-final"><h3>Schema linking complete: from clinical meaning to database structure</h3><divdata-popup-notes-body>
<p><strong>What schema linking means.</strong> Schema linking connects the concepts in the clarified research question to the database: which tables contain the information, which columns represent each concept, and how records from different tables must be connected.</p>
<p>(4) Once these choices are clear, we build and test smaller query steps, called CTEs. Each has a purpose, SQL and sample results for review. Here, we inspect patient and procedure records. Successful execution alone does not establish clinical correctness, so the reviewer can accept or request changes.</p>
<p><strong>Our clinical example.</strong> Here, we connect the treated pathology to the corresponding ablation episode, associate that episode with its date and patient, and specify the rules for identifying each patient's first recorded AF ablation and applying the 2020–2023 window. The intended result is a count of distinct patients.</p>
<p><strong>What completion means.</strong> These choices are now consolidated into a documented mapping, with the relationships, selection rules, and validation checks shown for review. We have an explicit specification of where the answer will come from and how the relevant data fit together.</p>
<p><strong>The final approval.</strong> By selecting <strong>Save and proceed</strong>, the reviewer approves this mapping for the next stage: planning and building the SQL query. Query execution and verification of the resulting patient count still follow.</p>
</div></div>
</div></div>
<divdata-popup-notes="thothii-cte-plan"><h3>CTEs: building the query step by step</h3><divdata-popup-notes-body><p><strong>CTE stands for Common Table Expression.</strong> It is a named intermediate result within an SQL query.</p><p>CTEs break a complex research question into smaller, logical steps—for example, identifying AF ablations, finding the first procedure for each patient, and selecting the study population.</p><p>This makes the query easier to understand, check, and modify, helping us verify that each step reflects the intended clinical criteria.</p></div></div>
<divdata-popup-notes="thothii-cte-1"><h3>Reviewing a CTE: purpose, SQL, and results</h3><divdata-popup-notes-body>
<p>(5) We can now assemble and verify the final SQL against the agreed question. The reviewer approves it or requests changes, checking that the query answers the original clinical intent.</p>
<p><strong>Purpose and position.</strong> Each CTE is presented as a reviewable step. At the top, we see its name and position in the sequence: here, the first of three. A short explanation describes its clinical purpose and the reasoning behind the selection rules.</p>
<p><strong>The SQL implementation.</strong> The code shows how that purpose is translated into database operations. In this example, it connects the treated pathology to the ablation episode and selects AF ablations, returning the patient identifier, episode identifier, and procedure date.</p>
<p><strong>Tests and a data preview.</strong> The screen reports the test status and execution time, followed by a preview of the returned records. The ten rows shown are a limited preview, not the total study population. Successful execution still requires a check that the results make clinical sense.</p>
<p><strong>Human review.</strong> The reviewer can compare the explanation, code, and sample data before choosing <strong>Save and proceed</strong> or <strong>Reject</strong>. This makes each intermediate step inspectable before it contributes to the final query.</p>
<divdata-popup-notes="thothii-datamart"><h3>Reusing the query: daily datamarts or research on demand</h3><divdata-popup-notes-body>
<p>(6) Withthe query approved, we decide whether to produce a datamart: an analysis dataset that can be refreshed through the ETL pipeline. We also choose which clarifications to retain for future questions. The saved artifacts and review decisions document how the result was reached.</p>
<p>Once the query generation and review workflow is complete, the validated SQL query can be reused beyond the current session.</p>
<p><strong>Daily datamart generation.</strong> The query can be integrated into the broader ETL pipeline and scheduled to run every day. This allows the datamart to be rebuilt or refreshed systematically as new source data become available, using the same agreed selection rules.</p>
<p><strong>Reuse by researchers.</strong> Alternatively, the query can simply be saved as an SQL file and made available to researchers, who can inspect it, run it when needed, or adapt it for a subsequent study.</p>
<p>In both cases, the workflow produces a reusable query that captures the reviewed interpretation of the research question.</p>
</div></div>
<divdata-popup-notes="thothii-memory"><h3>Saving clinical clarifications for future queries</h3><divdata-popup-notes-body>
<p><strong>What we retain.</strong> Clarifying a research question can produce knowledge that is useful beyond the current study: for example, an agreed interpretation of a clinical term or how a procedure is represented in the local data.</p>
<p><strong>The researcher chooses.</strong> At the end of the workflow, ThothII presents candidate clarifications for reuse. The reviewer selects which ones to save as shared knowledge for the workspace.</p>
<p><strong>How this helps next time.</strong> When a related question is asked, ThothII can retrieve these clarifications and propose them to the reviewer. The reviewer checks whether they apply to the new question before using them. This helps avoid repeating the same clarification work while keeping each study's interpretation under human control.</p>
</div></div>
<divdata-popup-notes="thothii-finish"><h3>A saved session documents how the result was reached</h3><divdata-popup-notes-body>
<p>At the end of the workflow, ThothII finalizes a session that brings together the generated SQL statement and the documented process that led to it.</p>
<p>The session preserves the research question, its agreed interpretation, the mapping to the database, the intermediate query steps, and the recorded validation results.</p>
<p>Human involvement is captured through the recorded clarifications and review decisions: what the reviewer selected, approved, or rejected along the way. These records document how the interaction between the system and the researcher shaped the final query.</p>
<p>Researchers can revisit the session to inspect the SQL and understand the choices behind it. This provides a traceable record of how the original question became the final result.</p>
<p><strong>Time and cost for this example.</strong> The AI processing took less than five minutes, excluding the time spent on human review. Using DeepSeek V4 Flash, the estimated model cost for this run was about five cents.</p>
</div></div>
<divdata-popup-notes="thothii-return"><h3>Expert supervision and deployment options</h3><divdata-popup-notes-body>
<p><strong>Supervision requires knowledge of the data.</strong> The workflow must be supervised by someone who understands the meaning and content of the database tables and fields. AI can make plausible guesses about what a field represents or how records should be connected, and these guesses can become hallucinations. The reviewer's role is to challenge those assumptions and check them against the actual data and its clinical meaning.</p>
<p><strong>The researcher defines the dataset.</strong> Human judgement is essential to decide which information the study needs, which records meet the required quality criteria, and which time windows and temporal rules should apply. The responsible researcher must confirm these choices so that the resulting dataset addresses the research question. Successful SQL execution alone does not establish that the dataset is suitable for the study.</p>
<p><strong>Server or authorised workstation.</strong> At Policlinico San Donato, ThothII runs on a server. The application can also be installed on the PC of a qualified staff member who is authorised to access the database remotely. The application runs on that workstation, while the database server executes the SQL through the authorised connection.</p>
<p><strong>A local deployment option.</strong> Open models such as Qwen3.8-27B or Gemma 4 26B A4B are candidates for running the AI component on premises. Their suitability for this workflow should be assessed on representative research questions, including the quality of the generated SQL and the reliability of the review steps.</p>
<p><strong>Hardware within reach.</strong> With quantized models, a compact server equipped with a suitable GPU is a practical deployment option. A GPU budget of a few thousand euros is a planning target; the required memory and final cost depend on the model, context length, and number of concurrent sessions.</p>
<p><strong>Where the query runs.</strong> SQL execution takes place on the internal database server through controlled application code. The language model helps construct and review the query; the database engine executes it.</p>
<p><strong>Keeping model processing local.</strong> Hosting the session model on premises allows its prompts and responses to remain within the organisation. With an external model, the information included in prompts and tool results must be checked and filtered to prevent sensitive data from leaving the environment.</p>
<p><small>Model references: <ahref="https://huggingface.co/Qwen/Qwen3.8-27B"target="_blank"rel="noopener noreferrer">Qwen3.8-27B model card</a>; <ahref="https://ai.google.dev/gemma/docs/core"target="_blank"rel="noopener noreferrer">Gemma 4 models and memory requirements</a>.</small></p>
<spanclass="head-org">AritmoLab · Policlinico San Donato</span>
</header>
<navclass="band">
<spanclass="band-section">The happy ending</span>
</nav>
<divclass="sbody">
<divclass="kicker">From datamart to discovery</div>
<h2>One datamart — many questions answered</h2>
<divclass="charts">
<divclass="chart"><b>Multivariate analysis</b><span>[SVG chart — next step]</span></div>
<divclass="chart"><b>Machine learning</b><span>[SVG chart — next step]</span></div>
</div>
<pclass="hint">[Illustrative data, modeled on published arrhythmology predictors — Brugada focus]</p>
</div>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">12 / 13</span></footer>
<asideclass="notes">[DRAFT — next step: two SVG charts, illustrative data modeled on real literature] And this is the happy ending. From one datamart, the Unit can run a multivariate analysis — which factors truly drive arrhythmic risk in Brugada patients — and train machine-learning models on the same table. What used to take weeks of manual data preparation now takes minutes. The AI did not replace the researcher: it gave the researcher back their time.</aside>
</section>
<sectionclass="arit title-slide">
<sectionclass="arit title-slide">
<headerclass="head">
<headerclass="head">
<imgsrc="logo.png"alt="">
<imgsrc="logo.png"alt="">
@@ -742,15 +614,17 @@
<divclass="kicker">Questions?</div>
<divclass="kicker">Questions?</div>
<h1>Thank you</h1>
<h1>Thank you</h1>
<divclass="hgap"></div>
<divclass="hgap"></div>
<pclass="lead">Want the deep dives? The technical walkthroughs of the text miner, the mappingagents and ThothII are coming as videos on our Substack.</p>
<pclass="lead">We are available to explore AI in data reorganization and clinical text analysis,<br>and datamart generation from natural-language requests with ThothII.<br>Meet us during the conference or arrange a remote session.</p>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">13 / 13</span></footer>
<footerclass="foot"><spanclass="g">Role of AI in the Analysis of Unstructured Clinical Databases</span><spanclass="g">Dr. Marco Pancotti - MultiPhysixLab</span><spanclass="g">Dr. Sara Paratico - Gruppo San Donato</span><spanclass="g">San Donato Milanese, Milan, Italy · 2–3 October 2026</span><spanclass="g num">11 / 11</span></footer>
<asideclass="notes">
<asideclass="notes">
[DRAFT] Thank you for your attention. If you want to go deeper — the text miner, the mappingagents, ThothII — we are publishing the technical walkthroughs as videos on our Substack; the link is on the final version of this deck. And now, happy to take your questions.
Thank you for your attention. Sara Paratico and I are available for a closer look at how we use AI to reorganize data, analyse clinical text, and generate datamarts from natural-language requests with ThothII. We can discuss these topics and demonstrate the tools during the conference or remotely in the coming days. Please contact us at the email addresses on this slide.
</aside>
</aside>
</section>
</section>
@@ -780,10 +654,10 @@
});
});
// jump dropdown — one select per slide, in the red band
// jump dropdown — one select per slide, in the red band
constJUMP_TITLES=['Title','Where we started','What we wanted to build',
constJUMP_TITLES=['Title','What we wanted to build',
'The problem','What the text miner reads','Hard problems of clinical NLP','The clinical ontology',
'The problem','What the text miner reads','Hard problems of clinical NLP','The clinical ontology',
'Trust the text','AritmoLab today','All good? Not yet','A tour of ThothII',
'Trust the text','AritmoLab today','All good? Not yet','A tour of ThothII',
<divid="reading"tabindex="0"aria-label="Scrollable speaker notes"><divclass="slide-label"id="notes-number">SLIDE 01</div><h1id="notes-title">Preparing your presentation</h1><articleid="notes">The slide and its speaker notes will appear here.</article></div>
<divid="reading"tabindex="0"aria-label="Scrollable speaker notes"><articleid="notes">The slide and its speaker notes will appear here.</article></div>
<footer><span>Notes are visible only in this window.</span><span><kbd>←</kbd><kbd>→</kbd> Change slide · <kbd>Esc</kbd> Close popup</span></footer>
<footer><span>Notes are visible only in this window.</span><span><kbd>←</kbd><kbd>→</kbd> Change slide · <kbd>Esc</kbd> Close popup</span></footer>
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.