Poster Session 3.W - Pharmaceutical Sciences and Health Technologies
Bereczki, Zoltán, MSc
E9BX9T
Department of Pharmacology and Pharmacotherapy, Semmelweis University
06307128401
bereczki.zoltan.andras@semmelweis.hu
Identifying Biomarkers Associated with Long COVID Subtypes Using Artificial Intelligence Models
Zoltán Bereczki1,2, Olivér Márton Balogh1,2, Dominika Lukovic3, Antonia Domanig3, Roland Molontay4, Péter Ferdinandy1,2,5, Bence Ágg1,2,5, Mariann Gyöngyösi3
1: Department of Pharmacology and Pharmacotherapy, Semmelweis University, H-1085 Budapest, Hungary
2: Center for Pharmacology and Drug Research & Development, Semmelweis University, H-1085 Budapest, Hungary
3: Division of Cardiology, Department of Internal Medicine II, Medical University of Vienna, 1090 Vienna, Austria
4: Institute of Biostatistics and Network Science, Semmelweis University, Üllői út 26., Budapest, H-1085, Hungary
5: Pharmahungary Group, H-6722 Szeged, Hungary
Poszter
Poster Session 3.W - Pharmaceutical Sciences and Health Technologies
English
Pharmaceutical Sciences and Health Technologies
Introduction
Long COVID multiorgan disease, a long-term consequence of COVID-19 infection with a prevalence of about 3-20% represents a major challenge to healthcare systems worldwide. In the absence of diagnostic biomarkers and specific therapies, early diagnosis is crucial to prevent disease progression. Accordingly, biomarkers and computational approaches for distinguishing between the clinical phenotypes of long COVID patients (cardiovascular, pulmonary or neurological) are still lacking.
Aims
Our aim was to identify potential biomarkers associated with long COVID subtypes through developing an artificial intelligence (AI) model trained to distinguish between these subtypes, using clinical data.
Methods
The dataset of 544 long COVID patients included 864 features. After preprocessing, features were selected by various missingness thresholds, and missing values in the remaining features were imputed using several methods. Decision-tree-based AI models (XGBoost, RandomForest, and CART) were trained on the resulting imputed datasets and evaluated using k-fold cross-validation with accuracy, weighted F1, and AUROC. Feature importance was assessed using SHapley Additive exPlanations (SHAP).
Result
After data preprocessing, 517 long COVID patients and 725 features were kept for further analysis. According to preliminary results, RandomForest model performed best (mean AUROC of 0.77 ± 0.03) with a maximum of 40% feature missingness. Feature importance analysis revealed potential biomarkers. For example, non-physiological cardiac abnormalities detected by either echocardiography or MRI, pulmonary symptoms, and absence of cardiovascular complaints are strongly associated with increased probability of cardiovascular, pulmonary, neurological subtype classification, respectively.
Conclusion
This is the first demonstration of identifying biomarker candidates characteristic of cardiovascular, pulmonary, and neurological subtypes of long COVID as a multi-organ disease, using a RandomForest-based AI model trained on clinical data from long COVID patients.
Funding
Project no. RRF-2.3.1-21-2022-00003 has been implemented with the support provided by the European Union. B.Z has been supported by a Semmelweis 250+ Excellence Fellowship. Gy.M has been supported by the Austrian Science Fund KLI 1064-B and Medical Scientific Fund of the Mayor of the City of Vienna 21176.
Semmelweis University
Bence Ágg
I do not give consent to the publication of my abstract on the website of the congress.
in doctoral studies before complex exam (PhD)
Szabad
elfogadva
poszter
nem rendelkezett róla
9757
13:36
13:39