diff --git a/PROJECT_STATE.md b/PROJECT_STATE.md index 15c52b8f..f7aed2cf 100644 --- a/PROJECT_STATE.md +++ b/PROJECT_STATE.md @@ -1,6 +1,6 @@ # ThothII — Project State -Last updated: 2026-09-24 (slide 11 PNG numbering aligned with popups). +Last updated: 2026-09-24 (11-slide deck, consolidated project opening and final contacts). This file is the short operational snapshot. Stable commands and the architecture mental model live in `AGENTS.md`; current design and runtime contracts live under `docs/architecture/`, @@ -14,18 +14,42 @@ requirements as mandatory; do not replace the running server stack in place. ## Presentation publishing +The deck now has 11 slides. The cover remains; “Where we started” is folded into +“What we wanted to build”, now slide 2. Its speaker notes open with the four goals, +and its existing popups include the disconnected sources and missing capabilities. +Speaker notes now read as continuous prose without titles. On slide 2 and slides +7–10, parenthesized cues identify the corresponding popup button at the start of +its paragraph or dedicated popup text. The closing slide has no popup cues. +The console reading area also omits slide and popup headings; popup controls +remain below the preview. Edits preserve the content and keep the spoken text close +to its previous length; actual delivery time still depends on rehearsal. The former +penultimate “One datamart — many questions answered” slide is removed. The ThothII +tour leads directly to “Thank you”, with both contact emails in bold 32px type. + +Slide 8 now has four popups: Home page, Patient data (original PNG 3), +Dashboard catalogue (PNG 6), and Brugada dashboard (PNG 7), numbered 1–4. +Each has dedicated presenter notes; the opening explains the portal's breadth, +the limited time, and Marco and Sara's availability for an in-person or remote +follow-up. The overview uses a two-by-two grid; unused PNGs remain on disk. + Presentation PNG source (user instruction, 2026-09-23): all PNGs to import come from `/Users/mp/Desktop/ScreenshotPresentation/ThothIIScreens/`. This is the verified on-disk path the user refers to as `desktop/screenshot/presentation/thothIIscreens`. Copy supplied files unchanged into `presentation/deck/screenshots/thothii/`, preserving -their embedded annotations. The 12 slide 11 PNGs are numbered `01` through `12` in -popup order, in both the source folder and the deck. The first popup uses +their embedded annotations. The 12 source PNGs retain their `01` through `12` filenames. +The shortened slide 10 tour uses PNGs 01, 02, 05, 07, 08 and 09, numbered 1–6 +in the interface; unused images remain on disk. The 284-word presenter script +covers all eight phases, explicitly introduces Human in the Loop, and includes +the invitation to a longer presentation at the conference or remotely. +The planned duration is 2:50, to be confirmed by rehearsal. +Slide 11 prominently displays Marco Pancotti's and Sara Paratico's email addresses. +Local preview uses `python3 -m http.server 8000 --bind 127.0.0.1 --directory presentation`. +The first popup uses `01-StartingPoint.png`; popup 2 uses `02-Disambiguation01.png`. Refresh the image cache version when replacing a PNG. During the current editing session, refresh both operational browser windows after every modification, as explicitly requested. -Popup 3 uses `03-DisambiguationFinal.png`, popup 4 uses `04-SchemaLinking01.png`, -and popup 5 uses `05-CloseSchemaLinking.png`. Popups 6–12 are CTE planning, CTE 1, -Final SQL, Datamart, Saving memory, Final step, and Back to start. +Popup 3 uses `05-CloseSchemaLinking.png`, popup 4 uses `07-CTE01.png`, +popup 5 uses `08-FinalSQL.png`, and popup 6 uses `09-DatamartProduction.png`. The HTML references PNGs by relative URL: they are external static assets, not compiled or embedded. Keep `presentation/deck/screenshots/thothii/` with the deck; the Desktop source folder is not a runtime dependency. diff --git a/presentation/deck/index.html b/presentation/deck/index.html index 2d7b8134..62947ce1 100644 --- a/presentation/deck/index.html +++ b/presentation/deck/index.html @@ -8,7 +8,7 @@ - +
We had four goals: clinical dashboards to understand the Unit’s activity and patients’ histories, a shared data warehouse, structured cohorts for machine learning and predictive statistics, and a unified professional portal.
+(1) We started with disconnected systems: twenty years of Cardioref records, much of their research value in free text, genetic results entered manually into Excel, ECGs on paper or exported on request, and an early portal without source integration. Each source answered a different need, but combining them for a clinical study meant collecting exports and rebuilding patient histories by hand.
+(2) To connect them, staging preserves a working copy of the sources.
+(3) Integration then cleans, normalizes and links the records.
+(4) The data warehouse organizes events and their context for analysis.
+(5) From there, datamarts turn that shared information into research datasets and dashboards within AritmoLab.
+(6) The brain symbols show where AI contributes. It drafts the mappings, which humans review.
+(7) During integration, AI extracts structured information from clinical text.
+(8) ThothII then builds datamarts from plain-English questions, with human approval.
+Open-source tools support the platform, and AI coding agents helped throughout development. The popups connect each component to the gap it addresses.
+Next, we turn to the hardest part: unstructured data. Sara will explain how we extract clinical meaning and check its reliability.
Over twenty years, we have collected clinical information in discharge letters and procedure reports. Much remains in free text.
+Clinicians connect these details and form hypotheses. Here, one sentence combines syncope, ECG changes and a positive flecainide test. Research needs defined variables that preserve this meaning.
+Our text miner extracts them in the integration layer, before the data enter the warehouse.
+We face four challenges: capturing diverse clinical histories, interpreting context correctly, mapping different expressions to common terms, and validating accuracy. Reliable research and patient care depend on getting these details right.
Our text miner uses explicit clinical rules to identify diagnoses and procedures, including catheter ablation and drug challenge tests, while checking their context.
+These figures cover over 58,000 letters and around 73,000 records of clinical conditions. We identified almost 11,000 drug challenge test entries, including positive Brugada tests in 2,307 patients.
+Each extracted record retains the rule version used. This lets us compare classifications with clinical review and investigate errors. For procedure classification, our target is at least 95 percent accuracy against 100 manually reviewed records.
A report may describe atrial fibrillation in the patient, rule it out, or mention it in the father's history. The same term must lead to different classifications.
+Our analyser checks for negation and family references, recording these attributes separately. It recognises Italian and English terms and common abbreviations. For example, "TA" may mean atrial tachycardia, but followed by a blood pressure value, it should not trigger that diagnosis.
+When reports describe several events, the system separates the text into clauses. For drug challenge tests, this helps distinguish a positive result before ablation from a negative result afterwards, preserving each observation and its clinical context.
Our clinical ontology, a shared set of terms, helps identify arrhythmia-related diagnoses in the narrative.
+The first tier covers conditions such as atrial fibrillation and Brugada. If none are found within a text field, the system checks a second tier for cardiovascular comorbidities, such as cardiomyopathy or heart failure, that may influence the patient's arrhythmic presentation.
+Different expressions map to one concept: for example, "fibrillazione atriale" and "atrial fibrillation" receive the same label.
+The extracted records enter the warehouse alongside data clinicians entered in structured fields. We can then define patient groups and research datasets, retaining each finding's source and the version of the extraction rules.
We set a 95% procedure-accuracy target, against 100 manual reviews. The 85% pathology target measures coverage, not diagnostic accuracy.
-Finding a disease name in a letter is only the first step. We apply explicit clinical criteria to decide which patients belong in the analysis.
-Version 1.3.1 corrected missed negations in roughly 900 Brugada records, with regression tests protecting the fix.
-Negated and family findings retain their context. Research datasets filter them explicitly, so preserving a finding does not mean counting it as the patient’s diagnosis.
-We can reprocess the archive quickly, inspect the rules behind each result, and reproduce the same extraction with the same rules.
+(1) To check the extraction, we set a 95% procedure-accuracy target against 100 manual reviews. The 85% pathology target measures coverage, not diagnostic accuracy.
+(2) Finding a disease name in a letter is only the first step. We also apply explicit clinical criteria to decide which patients belong in the analysis.
+(3) Review feeds back into the rules. Version 1.3.1 corrected missed negations in roughly 900 Brugada records, with regression tests protecting the fix.
+(4) Throughout this process, negated and family findings retain their context. Research datasets filter them explicitly, so preserving a finding does not mean counting it as the patient’s diagnosis.
+(5) This lets us reprocess the archive quickly, inspect the rules behind each result, and reproduce the same extraction with the same rules.
Select a screen to enlarge · Follow the tour from 1 to 7
+Select a screen to enlarge · Follow the tour from 1 to 4
The home page summarises around 57,000 patients, by age, sex and geographical origin. It is the starting point for exploring the archive.
-Search filters help us find a patient or study participant, open their record, or export the results.
-The profile brings demographic and clinical fields together, with access to procedures, devices, diagnostic examinations and genetics.
-A dated timeline brings together clinical notes, discharge letters and procedures, retaining the source of each event.
-Here, an ablation record shows the treated arrhythmias, procedural details, recorded complications and conclusions.
-We then move from individual records to dashboards covering departmental activity, procedures, devices and genetics.
-The funnel separates text mentions, filtered mentions, confirmed cases and ablation outcomes. Other charts describe sex, age, annual diagnoses and the timing of pre- and post-assessments.
+(1) AritmoLab is an extensive portal with many interconnected features. Our limited time prevents us from presenting it in detail, so we will show just four screens. Sara Paratico and I are available for a more in-depth presentation on request, either during the conference or afterwards through a remote connection.
+The home page summarises around 57,000 patients by age, sex and geographical origin. It is the starting point for exploring the archive.
+(2) From this overview, we can open a patient profile, which brings demographic and clinical information together, with access to procedures, devices, diagnostic examinations and genetics. It provides a single starting point for exploring an individual patient's record.
+(3) Moving from individual patient records to an overview of the department, the dashboard catalogue gives us access to clinical activity, procedures, devices and genetics.
+(4) Among the many dashboards we could present, we chose Brugada as an example. The funnel separates text mentions, filtered mentions, confirmed cases and ablation outcomes. Other charts show sex, age, annual diagnoses and the timing of pre- and post-assessments.
+0. What is SQL?
SQL means Structured Query Language. It tells a database what to select, connect and count.
“How many patients with confirmed Brugada underwent an ablation?” We must define confirmation and count each patient once.
Dashboards need agreed definitions, time periods and denominators to make comparisons meaningful.
Study tables separate characteristics known before prediction from outcomes observed afterwards.
A datamart is an analysis dataset with explicit selection rules and repeatable checks.
5. The gap.
The gap is between clinical meaning and database instructions: storing data does not automatically make a question answerable.
The warehouse speaks SQL, which means Structured Query Language. It tells a database what to select, connect and count.
+(1) To answer a clinical question such as “How many patients with confirmed Brugada underwent an ablation?”, we must define confirmation and count each patient once.
+(2) The same need for clarity applies to Health Intelligence: dashboards need agreed definitions, time periods and denominators to make comparisons meaningful.
+(3) For predictive research, study tables separate characteristics known before prediction from outcomes observed afterwards.
+(4) A datamart provides an analysis dataset with explicit selection rules and repeatable checks.
+The gap is therefore between clinical meaning and database instructions: storing data does not automatically make a question answerable.
The task has changed. We are no longer extracting structured, coded values from the text within clinical records. We are now starting from a natural-language request to retrieve the records that contain the relevant values.
Research questions are often difficult to express clearly and unambiguously, and their clinical concepts do not map directly to the underlying database structure. As a result, writing the right SQL query can take hours, with repeated testing and refinement before it accurately reflects the intended research question.
1. Workflow phase. At the top, the phase indicator shows where we are in the process. Here, phase one is highlighted: clarifying the research question.
-2. AI working time. The timer measures how long the AI has been working, excluding the time the human reviewer spends considering the options and making decisions.
-3. Model activity. On the left, we can follow the AI's running explanation of its analysis: the interpretations it considers, the evidence it consults, and the rationale for its proposals.
-4. Proposed interpretations. In the centre, the AI asks the researcher to choose between alternative interpretations. Here, the question is how to identify an ablation for atrial fibrillation in the database. This choice determines which patients enter the study.
-5. Reviewer actions. The reviewer can select a proposed option, or use Go back to revisit the previous step, Exit to stop the current workflow, or Other — specify to describe a different interpretation in free text.
+(1) ThothII combines a wide range of capabilities in a guided, eight-phase workflow. A full presentation deserves at least thirty minutes. Here, we focus on the main steps.
+It helps researchers turn clinical questions into checked SQL and reusable analysis datasets, bridging clinical meaning and database structure. AI proposes, and people review, correct and approve. This is Human in the Loop throughout the workflow, combining clinical judgment and database expertise.
Bringing the decisions together. The AI combines the original request with the researcher's answers into an explicit study definition. It records what we mean by each clinical concept, which patients to include, the time window, and the result we want.
-The clarified question. In this example: how many distinct patients had their first recorded ablation for atrial fibrillation between 2020 and 2023? We identify the procedure from the treated pathology, and find each patient's first AF ablation across the entire available database history before applying the date filter.
-Making the interpretation actionable. The summary links these choices to the relevant database fields and selection rules, and displays the checks completed during clarification. These agreed criteria guide the subsequent schema linking and SQL construction.
-Human confirmation. The researcher reviews the consolidated definition. Save and proceed confirms it and closes phase one; Reject sends it back for revision.
+(2) We begin by clarifying the question: here, what counts as an atrial fibrillation ablation? We then review relevant knowledge from earlier work before approving a precise reformulation: how many patients had their first recorded AF ablation between 2020 and 2023? We find the first procedure across the available history before filtering dates. The reviewer can challenge the interpretation.
Linking clinical meaning to data. We are now in phase four, schema linking: identifying where the information needed to answer the agreed question is stored. Tables group related records; columns hold specific details, such as a patient identifier, a procedure date, or the condition treated.
-Proposed tables. The AI presents candidate tables and explains why each is relevant. Here, one table provides the ablation episodes and their dates; another identifies the pathology treated, allowing us to distinguish atrial fibrillation ablations from other procedures.
-Selecting the columns. The column controls let the reviewer inspect and adjust the fields selected for the query. We need the information required to identify patients, recognise the relevant procedures, and apply the agreed time criteria.
-Review before proceeding. The researcher and a technical reviewer can check this mapping together, adjust the selection, and confirm it. This establishes which data will support the SQL query; the relationships between the selected tables are reviewed next.
+(3) With the question agreed, we connect its concepts to tables, fields and relationships. We then review the mapping and selection rules together before building SQL. The reviewer checks which patients, procedures and dates will count.
What schema linking means. Schema linking connects the concepts in the clarified research question to the database: which tables contain the information, which columns represent each concept, and how records from different tables must be connected.
-Our clinical example. Here, we connect the treated pathology to the corresponding ablation episode, associate that episode with its date and patient, and specify the rules for identifying each patient's first recorded AF ablation and applying the 2020–2023 window. The intended result is a count of distinct patients.
-What completion means. These choices are now consolidated into a documented mapping, with the relationships, selection rules, and validation checks shown for review. We have an explicit specification of where the answer will come from and how the relevant data fit together.
-The final approval. By selecting Save and proceed, the reviewer approves this mapping for the next stage: planning and building the SQL query. Query execution and verification of the resulting patient count still follow.
+(4) Once these choices are clear, we build and test smaller query steps, called CTEs. Each has a purpose, SQL and sample results for review. Here, we inspect patient and procedure records. Successful execution alone does not establish clinical correctness, so the reviewer can accept or request changes.
CTE stands for Common Table Expression. It is a named intermediate result within an SQL query.
CTEs break a complex research question into smaller, logical steps—for example, identifying AF ablations, finding the first procedure for each patient, and selecting the study population.
This makes the query easier to understand, check, and modify, helping us verify that each step reflects the intended clinical criteria.
Purpose and position. Each CTE is presented as a reviewable step. At the top, we see its name and position in the sequence: here, the first of three. A short explanation describes its clinical purpose and the reasoning behind the selection rules.
-The SQL implementation. The code shows how that purpose is translated into database operations. In this example, it connects the treated pathology to the ablation episode and selects AF ablations, returning the patient identifier, episode identifier, and procedure date.
-Tests and a data preview. The screen reports the test status and execution time, followed by a preview of the returned records. The ten rows shown are a limited preview, not the total study population. Successful execution still requires a check that the results make clinical sense.
-Human review. The reviewer can compare the explanation, code, and sample data before choosing Save and proceed or Reject. This makes each intermediate step inspectable before it contributes to the final query.
+(5) We can now assemble and verify the final SQL against the agreed question. The reviewer approves it or requests changes, checking that the query answers the original clinical intent.
testo8
Once the query generation and review workflow is complete, the validated SQL query can be reused beyond the current session.
-Daily datamart generation. The query can be integrated into the broader ETL pipeline and scheduled to run every day. This allows the datamart to be rebuilt or refreshed systematically as new source data become available, using the same agreed selection rules.
-Reuse by researchers. Alternatively, the query can simply be saved as an SQL file and made available to researchers, who can inspect it, run it when needed, or adapt it for a subsequent study.
-In both cases, the workflow produces a reusable query that captures the reviewed interpretation of the research question.
-What we retain. Clarifying a research question can produce knowledge that is useful beyond the current study: for example, an agreed interpretation of a clinical term or how a procedure is represented in the local data.
-The researcher chooses. At the end of the workflow, ThothII presents candidate clarifications for reuse. The reviewer selects which ones to save as shared knowledge for the workspace.
-How this helps next time. When a related question is asked, ThothII can retrieve these clarifications and propose them to the reviewer. The reviewer checks whether they apply to the new question before using them. This helps avoid repeating the same clarification work while keeping each study's interpretation under human control.
-At the end of the workflow, ThothII finalizes a session that brings together the generated SQL statement and the documented process that led to it.
-The session preserves the research question, its agreed interpretation, the mapping to the database, the intermediate query steps, and the recorded validation results.
-Human involvement is captured through the recorded clarifications and review decisions: what the reviewer selected, approved, or rejected along the way. These records document how the interaction between the system and the researcher shaped the final query.
-Researchers can revisit the session to inspect the SQL and understand the choices behind it. This provides a traceable record of how the original question became the final result.
-Time and cost for this example. The AI processing took less than five minutes, excluding the time spent on human review. Using DeepSeek V4 Flash, the estimated model cost for this run was about five cents.
-Supervision requires knowledge of the data. The workflow must be supervised by someone who understands the meaning and content of the database tables and fields. AI can make plausible guesses about what a field represents or how records should be connected, and these guesses can become hallucinations. The reviewer's role is to challenge those assumptions and check them against the actual data and its clinical meaning.
-The researcher defines the dataset. Human judgement is essential to decide which information the study needs, which records meet the required quality criteria, and which time windows and temporal rules should apply. The responsible researcher must confirm these choices so that the resulting dataset addresses the research question. Successful SQL execution alone does not establish that the dataset is suitable for the study.
-Server or authorised workstation. At Policlinico San Donato, ThothII runs on a server. The application can also be installed on the PC of a qualified staff member who is authorised to access the database remotely. The application runs on that workstation, while the database server executes the SQL through the authorised connection.
-A local deployment option. Open models such as Qwen3.8-27B or Gemma 4 26B A4B are candidates for running the AI component on premises. Their suitability for this workflow should be assessed on representative research questions, including the quality of the generated SQL and the reliability of the review steps.
-Hardware within reach. With quantized models, a compact server equipped with a suitable GPU is a practical deployment option. A GPU budget of a few thousand euros is a planning target; the required memory and final cost depend on the model, context length, and number of concurrent sessions.
-Where the query runs. SQL execution takes place on the internal database server through controlled application code. The language model helps construct and review the query; the database engine executes it.
-Keeping model processing local. Hosting the session model on premises allows its prompts and responses to remain within the organisation. With an external model, the information included in prompts and tool results must be checked and filtered to prevent sensitive data from leaving the environment.
-Model references: Qwen3.8-27B model card; Gemma 4 models and memory requirements.
+(6) With the query approved, we decide whether to produce a datamart: an analysis dataset that can be refreshed through the ETL pipeline. We also choose which clarifications to retain for future questions. The saved artifacts and review decisions document how the result was reached.
- AritmoLab · Policlinico San Donato
- [Illustrative data, modeled on published arrhythmology predictors — Brugada focus]
-