Poster Session I. - A: Molecular Medicine
Mr. Szepesi-Nagy Istvan
K45SFS
HUN-REN RCNS Institute of Molecular Life Sciences
06304929480
szepesi.istvan@phd.semmelweis.hu
FragFlow: Automated Workflow for Large-Scale Quantitative Proteomics in High Performance Computing Environments
Istvan Szepesi-Nagy1, Roberta Borosta1, Zoltan Szabo2, Gabor E. Tusnady1, Lorinc S. Pongor3, Gergely Rona1
1: Insitute of Molecular Life Sciences, HUN-REN RCNS
2: Department of Medical Chemistry, Albert Szent-Györgyi Medical School, University of Szeged
3: Cancer Genomics and Epigenetics Core Group, Hungarian Centre of Excellence for Molecular Medicine (HCEMM)
Poszter
Poster Session I. - A: Molecular Medicine
English
Molecular Medicine
INTRODUCTION
Mass spectrometry-based proteomics generates large, complex datasets that often overwhelm desktop computational resources and require manual configuration for analysis. FragPipe (FP) delivers rapid peptide identification across diverse sample preparation and acquisition modes (DDA, DIA, TMT) but remains challenging to deploy at scale.
AIMS
Although some earlier studies have attempted to address the challenges of scalable proteomics workflows, the complexity of these solutions, coupled with the widespread adoption of FP, underscores the need for a dedicated workflow integration tailored specifically to FragPipe. Currently, utilizing FP in HPC environments requires multiple manual interventions, with deployment processes that are often elaborate and time-consuming, limiting its accessibility and scalability in large-scale analyses.
METHODS
We introduce FragFlow, a Nextflow‐based pipeline that containerizes FP, automates input manifest and workflow generation, manages tool dependencies and includes downstream data analysis options to enable reproducible, high‐performance analyses on HPC, cloud, and cluster environments.
RESULTS
Benchmarking against other available workflow-based solutions demonstrates that FragFlow matches quantitative accuracy while significantly reducing runtime and alleviating memory and I/O bottlenecks. We validate FragFlow across three representative datasets, DDA, DIA, and TMT, successfully recapitulating published biological signatures with minimal user intervention. By combining the sensitivity and speed of FragPipe with Nextflow’s orchestration, FragFlow democratizes large‐scale proteomics, making advanced MS data analysis accessible to researchers without extensive computational expertise.
CONCLUSION
FragFlow transforms FragPipe from a powerful yet desktop‐oriented application into a robust, reproducible, and user‐friendly HPC based pipeline capable of meeting the demands of modern, large‐scale proteomics. FragFlow enables rapid, reproducible extraction of biological insights from complex MS datasets, thereby accelerating discovery across diverse biomedical research areas.
FUNDING
I.S.N. is supported by the EKÖP-2024-124 New National Excellence Program of the Ministry for Culture and Innovation from the source of the National Research, Development and Innovation Fund.
Semmelweis University
Dr. Gergely Rona
I do not give consent to the publication of my abstract on the website of the congress.
in doctoral studies before complex exam (PhD)
Szabad
elfogadva
poszter
nem rendelkezett róla
8925
16:30
16:36