Contributions
227 contributions
Short papers
Oral and poster presentations
An Enhanced Pipeline for Multi-Omic Integration Based on Topological Data Analysis
Veronica Paparozzi, Anna Plaksienko, Marco Pedicini and Christine Nardini
An Enhanced Pipeline for Multi-Omic Integration Based on Topological Data Analysis
Veronica Paparozzi, Anna Plaksienko, Marco Pedicini and Christine Nardini
- Keywords
- multi-omicsnetworktopological data analysisharmonic persistent homologycancer biomarkers
An XAI-Driven Transformer for Non-Canonical Splice-Site Prediction
Ahmed Negm
An XAI-Driven Transformer for Non-Canonical Splice-Site Prediction
Ahmed Negm
- Keywords
- non-canonical splice sitesTransformerIntegrated Gradientsexplainable AI
Analysis of biological networks using Krylov subspace trajectories
H. Robert Frost
Analysis of biological networks using Krylov subspace trajectories
H. Robert Frost
- Keywords
- Krylov subspacebiological networksperturbation analysiscommunity detection
Genomic Centromere Profiling Tool using Centeny Maps for Centromere Characterization Across Genomes
Mauricio Orantes Bonilla, Matteo Tommaso Ungaro, Luca Corda, Andreas Giannis and Simona Giunta
Genomic Centromere Profiling Tool using Centeny Maps for Centromere Characterization Across Genomes
Mauricio Orantes Bonilla, Matteo Tommaso Ungaro, Luca Corda, Andreas Giannis and Simona Giunta
- Keywords
- Genomic centromere profiling pipelineCentenyCENP-B boxGenome-to-genome comparisonCentromere orgnaization
FusFun: a tool for the functional evaluation of gene fusions in RNA-seq–data
Francesca Miccolis, Martina Filieri, Elisa Ficarra and Marta Lovino
FusFun: a tool for the functional evaluation of gene fusions in RNA-seq–data
Francesca Miccolis, Martina Filieri, Elisa Ficarra and Marta Lovino
- Keywords
- fusion genesfunctional annotationRNA sequencingtranscriptomics
A Bayesian Adaptive Enrichment Design for Borrowing from Published Aggregate Summaries
Lara Maleyeff, Shirin Golchi and Erica E. M. Moodie
A Bayesian Adaptive Enrichment Design for Borrowing from Published Aggregate Summaries
Lara Maleyeff, Shirin Golchi and Erica E. M. Moodie
- Keywords
- Bayesian adaptive enrichmentnormalized power priorhistorical borrowingprecision medicinesubgroup identification
Incorporating Biologically Informed Masking into Protein Language Models for Pocket Identification
Luiza Gomes Ferreira
Incorporating Biologically Informed Masking into Protein Language Models for Pocket Identification
Luiza Gomes Ferreira
- Keywords
- Protein pocket identificationprotein language modelsmasked language modelingcomputational biology
Evaluating in-context learning with prompting regimes for transformer-based synthetic health tabular data generation
Amanda Bertgren, Fredrik Öhberg, Paolo Soda, Ulf Näslund, Patrik Wennberg and Christer Grönlund
Evaluating in-context learning with prompting regimes for transformer-based synthetic health tabular data generation
Amanda Bertgren, Fredrik Öhberg, Paolo Soda, Ulf Näslund, Patrik Wennberg and Christer Grönlund
- Keywords
- synthetic datatransformerpromptingtabular data
Network-based modelling of categorical data for personalized causal inference
Federico Castelletti and Laura Ferrini
Network-based modelling of categorical data for personalized causal inference
Federico Castelletti and Laura Ferrini
- Keywords
- Bayesian inferencecausal effectgraphical modelheterogeneitymixture model
Katz Kernel Based Stratification of Breast Cancer Patients
Francesco Cellitti, Lorenzo Di Rocco, Livia Perfetto, Veronica Lombardi and Umberto Ferraro Petrillo
Katz Kernel Based Stratification of Breast Cancer Patients
Francesco Cellitti, Lorenzo Di Rocco, Livia Perfetto, Veronica Lombardi and Umberto Ferraro Petrillo
- Keywords
- signal mechanistic networksgraph kernelgraph clustering
MAP-Net: An Interpretable Deep Learning Framework for Universal mRNA Manufacturability
Giorgio Ciano
MAP-Net: An Interpretable Deep Learning Framework for Universal mRNA Manufacturability
Giorgio Ciano
- Keywords
- mRNADeep LearningCodon OptimizationRNA Therapeutics
ChloroGuide: prediction of subplastid localization using protein language model
Filip Pietluch and Przemyslaw Gagat
ChloroGuide: prediction of subplastid localization using protein language model
Filip Pietluch and Przemyslaw Gagat
- Keywords
- ensemble learningESM-2protein language modelsprotein targetingsubplastid localization
PANDORA: a multiparametric framework for sensitive detection of insertional mutagenesis
Riccardo Pizzichemi
PANDORA: a multiparametric framework for sensitive detection of insertional mutagenesis
Riccardo Pizzichemi
- Keywords
- insertional mutagenesisintegration-site analysisgene therapyclonal trackingclonal dominance
Transporting principal causal effects across strata: A Bayesian causal inference approach
Veronica Ballerini, Falco Joannes Bargagli Stoffi and Francesca Dominici
Transporting principal causal effects across strata: A Bayesian causal inference approach
Veronica Ballerini, Falco Joannes Bargagli Stoffi and Francesca Dominici
- Keywords
- Covariate-shiftmediationprincipal stratificationenvironmental studiespublic health
A Neural Factorization Machine Approach to Drug–Target Interaction with Global Features
Mirco Perna and Simona Ester Rombo
A Neural Factorization Machine Approach to Drug–Target Interaction with Global Features
Mirco Perna and Simona Ester Rombo
- Keywords
- Drug-target associationsEmbeddingsDrug discovery
Cell Senescence Identification Using Single-Cell Transcriptomics Foundation Model
Yongqi Zhao, Fan Tong, Xiangwen Zheng, Yuanyuan Ma, Huajian Mao, Yuting Zhou, Peixiang Yang and Dongsheng Zhao
Cell Senescence Identification Using Single-Cell Transcriptomics Foundation Model
Yongqi Zhao, Fan Tong, Xiangwen Zheng, Yuanyuan Ma, Huajian Mao, Yuting Zhou, Peixiang Yang and Dongsheng Zhao
- Keywords
- Foundation ModelSingle-Cell TranscriptomicsCell Senescence
Dual-view Atomic and Fragment Graph Representation Learning for Drug Response Prediction
Duc-Khiem Doan and Duc-Hau Le
Dual-view Atomic and Fragment Graph Representation Learning for Drug Response Prediction
Duc-Khiem Doan and Duc-Hau Le
- Keywords
- Drug Response PredictionGraph Neural NetworksMolecular Fragment RepresentationDual-view Graph LearningPrecision Oncology
Evaluating SMT-Based Formal Inference for GRN Synthesis: From Data Density to Structural Rules
Ofri Caspi, Michal Greenberg, Eitan Tannenbaum and Hillel Kugler
Evaluating SMT-Based Formal Inference for GRN Synthesis: From Data Density to Structural Rules
Ofri Caspi, Michal Greenberg, Eitan Tannenbaum and Hillel Kugler
- Keywords
- Gene Regulatory NetworksFormal VerificationSatisfiability Modulo Theories
EXTRARNAS: A Framework for Extracting RNA Structures with Multiple Tools
Federico Di Petta, Piermichele Rosati, Piero Hierro Canchari, Michela Quadrini and Luca Tesei
EXTRARNAS: A Framework for Extracting RNA Structures with Multiple Tools
Federico Di Petta, Piermichele Rosati, Piero Hierro Canchari, Michela Quadrini and Luca Tesei
- Keywords
- RNA structure annotationbase pairsPDBDockerconsensus structure
PRIMA: a bidirectional state-space architecture and training approach for modelling protein-protein interactions
Arturo Fiorellini-Bernardis, Sebastien Boyer, Jean Quentin, Ghassene Jebali and Oliver Bent
PRIMA: a bidirectional state-space architecture and training approach for modelling protein-protein interactions
Arturo Fiorellini-Bernardis, Sebastien Boyer, Jean Quentin, Ghassene Jebali and Oliver Bent
- Keywords
- State-space ModelsProtein-Protein InteractionsLLMBinding Affinity Prediction
Optimizing RNA yield using deep neural networks coupled to massively parallel screening
Adrien Villain
Optimizing RNA yield using deep neural networks coupled to massively parallel screening
Adrien Villain
- Keywords
- RNA yieldMassively parallel screeningDeep neural networksManufacturability
Sequence-Derived Physicochemical Features and Stacked Learning for SUMOylation-Site Prediction
Junyan Li, Qiaobin Yao, Danda Rawat, Jiang Li and Shaolei Teng
Sequence-Derived Physicochemical Features and Stacked Learning for SUMOylation-Site Prediction
Junyan Li, Qiaobin Yao, Danda Rawat, Jiang Li and Shaolei Teng
- Keywords
- SUMOylationpost-translational modificationmachine learningfeature selectionensemble learning
Towards a unified framework: a data structures benchmark for representing genomic information
Matteo Tommaso Ungaro, Mauricio Orantes Bonilla, Francesca Querques, Evelyne Tassone and Simona Giunta
Towards a unified framework: a data structures benchmark for representing genomic information
Matteo Tommaso Ungaro, Mauricio Orantes Bonilla, Francesca Querques, Evelyne Tassone and Simona Giunta
- Keywords
- pangenomicsgenome graphssomatic mutationtranslational medicinehaplotype-resolved assembly
Toward Simulation-Informed Lipid Nanoparticle Design with Martini Coarse-Grained Models
Mariana Valério, Pablo Cardona Perez, Salma Maaquili, Mariia Borbuliak and Paulo C. T. Souza
Toward Simulation-Informed Lipid Nanoparticle Design with Martini Coarse-Grained Models
Mariana Valério, Pablo Cardona Perez, Salma Maaquili, Mariia Borbuliak and Paulo C. T. Souza
- Keywords
- Lipid nanoparticlesCoarse-grained molecular dynamicsMartini force fieldSimulation-derived descriptorsmRNA deliveryMachine learning
GPU-accelerated single-cell and spatial transcriptomics on NVIDIA DGX H100: a systematic benchmark of speed, scalability, and biological concordance
Luca Vedovelli, Corrado Lanera, Daniele Sabbatini and Dario Gregori
GPU-accelerated single-cell and spatial transcriptomics on NVIDIA DGX H100: a systematic benchmark of speed, scalability, and biological concordance
Luca Vedovelli, Corrado Lanera, Daniele Sabbatini and Dario Gregori
- Keywords
- GPU computingsingle-cell RNA-seqspatial transcriptomicsbenchmarkingrapids-singlecell
DualBioGraph: Patient-Held-Out Tumor--Normal Classification Using Dual Graph Neural Networks
Al Hamna Asif, Rana Abubakar, Amna Younus, Juan Jose Vegas Olmos and Filippo Cugini
DualBioGraph: Patient-Held-Out Tumor--Normal Classification Using Dual Graph Neural Networks
Al Hamna Asif, Rana Abubakar, Amna Younus, Juan Jose Vegas Olmos and Filippo Cugini
- Keywords
- single cellRNAseqgraph neural networksheterogeneous graphspatientvalidationGPU
A Multilayer Graph Diffusion Model for Spatial Transcriptomics Reveals Cell-Cell Propagation Dynamics.
Caterina Francesca Perri, Annamaria Defilippo, Federico Manuel Giorgi, Pietro Hiram Guzzi and Pierangelo Veltri
A Multilayer Graph Diffusion Model for Spatial Transcriptomics Reveals Cell-Cell Propagation Dynamics.
Caterina Francesca Perri, Annamaria Defilippo, Federico Manuel Giorgi, Pietro Hiram Guzzi and Pierangelo Veltri
- Keywords
- spatial transcriptomicsgraph diffusionmultilayer networksprobabilistic modellingkidney tissue
Interpretable Graph Neural Networks for Learning Drug-Gene Relationships and Inferring Phenotype Outcomes in Pharmacogenomics
Kgabe Ronald Molepo
Interpretable Graph Neural Networks for Learning Drug-Gene Relationships and Inferring Phenotype Outcomes in Pharmacogenomics
Kgabe Ronald Molepo
- Keywords
- pharmacogenomicsknowledge graphsrelational graph convolutional networks
Benchmarking a Spatial-Aware Foundation Model: An Empirical Evaluation of CellPLM in Transcriptomics
Juha Im, Haiping Liu and Hongpeng Zhou
Benchmarking a Spatial-Aware Foundation Model: An Empirical Evaluation of CellPLM in Transcriptomics
Juha Im, Haiping Liu and Hongpeng Zhou
- Keywords
- Single-cell RNA-seqSpatial TranscriptomicsFoundation ModelsTransformerTransfer Learning
Microbial Community Structure and Immune Integration in Lung Adenocarcinoma
Chiara Napoli, Francesco Bardozzo, Luigi Liguori, Claudio Angione, Annalisa Occhipinti and Roberto Tagliaferri
Microbial Community Structure and Immune Integration in Lung Adenocarcinoma
Chiara Napoli, Francesco Bardozzo, Luigi Liguori, Claudio Angione, Annalisa Occhipinti and Roberto Tagliaferri
- Keywords
- lung adenocarcinomamicrobiometranscriptomic
A Block Decomposed QUBO Workflow for Chromosome-Y Phylogeny Reconstruction
Giuliana Siddi Moreau, Riccardo Berutti, Manuela Profir, Lorenzo Pisani, Maria Laura Clemente and Lidia Leoni
A Block Decomposed QUBO Workflow for Chromosome-Y Phylogeny Reconstruction
Giuliana Siddi Moreau, Riccardo Berutti, Manuela Profir, Lorenzo Pisani, Maria Laura Clemente and Lidia Leoni
- Keywords
- QUBOADMM decompositionCounter-diabatic quantum optimisationPhylogeneticsY-chromosome
Hybrid Quantum-Classical Convolutional Neural Network for Genomic Sequence Classification
Giovanni Acampora, Angela Chiatto and Roberto Schiattarella
Hybrid Quantum-Classical Convolutional Neural Network for Genomic Sequence Classification
Giovanni Acampora, Angela Chiatto and Roberto Schiattarella
- Keywords
- Convolutional Neural NetworksGenomic Sequence ClassificationHybrid Quantum-Classical PipelineQuantum Machine LearningVariational Quantum Classifier
Quantum Kernel Estimation for the Discovery of Early Lung Cancer Detection
Hamed Javidi, Alex Zajichek, Hakan Doga, Laxmi Parida, Filippo Utro and Peter Mazzone
Quantum Kernel Estimation for the Discovery of Early Lung Cancer Detection
Hamed Javidi, Alex Zajichek, Hakan Doga, Laxmi Parida, Filippo Utro and Peter Mazzone
- Keywords
- quantum machine learningquantum kernel estimationcfDNA fragmentomicsDNA methylationlung cancer
Quantum-Enhanced Conformational Landscape Modelling of Amyloid-β42 via a Hybrid GNN-VQE/QAOA Architecture
Rajendra Reddy Bhavanam, Saran Boddu, Muthuraman Ramanathan and Likith Palakurthi
Quantum-Enhanced Conformational Landscape Modelling of Amyloid-β42 via a Hybrid GNN-VQE/QAOA Architecture
Rajendra Reddy Bhavanam, Saran Boddu, Muthuraman Ramanathan and Likith Palakurthi
- Keywords
- intrinsically disordered proteinsvariational quantum eigensolverquantum approximate optimisation algorithmgraph neural networkfree energy landscape
SOPHYSM: More Steps towards Digital Twins of Solid Tumours
Diego Cividini and Marco Antoniotti
SOPHYSM: More Steps towards Digital Twins of Solid Tumours
Diego Cividini and Marco Antoniotti
- Keywords
- SimulationCancerCell-lineagesDigital TwinImaging
A Framework for Diversity Constrained Design of Natural Product like Molecules
Aditi Vijil and Pin-Yu Chen
A Framework for Diversity Constrained Design of Natural Product like Molecules
Aditi Vijil and Pin-Yu Chen
- Keywords
- natural productsgenetic algorithmsdiversity constrained designdrug discovery
Bootstrap-Aggregated Method-of-Moments Estimation of Copula Correlation Under Generalized Gamma Dependent Censoring
Hyun-Soo Zhang and Chung Mo Nam
Bootstrap-Aggregated Method-of-Moments Estimation of Copula Correlation Under Generalized Gamma Dependent Censoring
Hyun-Soo Zhang and Chung Mo Nam
- Keywords
- Dependent censoringIdentifiabilityGeneralized gamma distributionCopulaSurvival analysis
MethylAging: A HiFi Long-Read Epigenetic Clock for Biological Age Prediction in an Emirati Blood Cohort Using Promoter-Aggregated Methylation Features
Dalia Barayan, Layan Hattab, Hoor Nadhari, Amna Ahmed, Mauricio Paton, Michael Olbrich, Antonello Maruotti, Mira Mousa, Habiba Alsafar and Andre Barreiros
MethylAging: A HiFi Long-Read Epigenetic Clock for Biological Age Prediction in an Emirati Blood Cohort Using Promoter-Aggregated Methylation Features
Dalia Barayan, Layan Hattab, Hoor Nadhari, Amna Ahmed, Mauricio Paton, Michael Olbrich, Antonello Maruotti, Mira Mousa, Habiba Alsafar and Andre Barreiros
- Keywords
- Epigenetic clockDNA methylationPacBio HiFiLong-read sequencingPromoter aggregationBorutaElastic netBiological ageEmirati cohortPrecision medicine
MicroEnsemble: A Patient-Anchor and Label-Anchor Contrastive Ensemble for Longitudinal Microbiome Classification
Amr Doma, Ahmed Wesal, Youssef Ellakany, Ghada El-Boghdady, Sama Mohamed, Yara Elshamy, Peter Salah and Tamer Basha
MicroEnsemble: A Patient-Anchor and Label-Anchor Contrastive Ensemble for Longitudinal Microbiome Classification
Amr Doma, Ahmed Wesal, Youssef Ellakany, Ghada El-Boghdady, Sama Mohamed, Yara Elshamy, Peter Salah and Tamer Basha
- Keywords
- microbiome classificationcontrastive learninginflammatory bowel diseaseensemble learning
Supervised NMF to unravel signature-class associations in microbiome composition: application to microbial signature of downy mildew epidemology in vineyards.
Alioune-Badara Diouf, Paola Fournier, Jean-Marc Frigerio, Corinne Vacher and Simon Labarthe
Supervised NMF to unravel signature-class associations in microbiome composition: application to microbial signature of downy mildew epidemology in vineyards.
Alioune-Badara Diouf, Paola Fournier, Jean-Marc Frigerio, Corinne Vacher and Simon Labarthe
- Keywords
- NMFmatrix factorizationsupervised NMFplant microbiomecrop epidemiosurveillance
Causal Effects on Latent Structures via Bayesian Nonparametric Factor Models
Roberta De Vito, Giovanni Parmigiani and Dafne Zorzetto
Causal Effects on Latent Structures via Bayesian Nonparametric Factor Models
Roberta De Vito, Giovanni Parmigiani and Dafne Zorzetto
- Keywords
- causal inferencelatent factorsBayesian additive regression treescumulative shrinkage processdependent Dirichlet process
A Landmarked Penalised Weibull Framework for Time-Updated Survival Modelling of Heterogeneous Clinical Outcomes in Multiple Sclerosis
Francisco José Aparicio Serrano, Ariadna Masot-Llima, Agustín Pappolla, Susana Otero-Romero, René Carvajal, Álvaro Cobo-Calvo, Manel Alberich, María Jesús Arévalo, Georgina Arrambide, Cristina Auger, Joaquín Castillo, Manuel Comabella, Ingrid Galán, Daniel Hernández, Carlos Nos, Jordi Río, Breogán Rodríguez-Acevedo, Jaume Sastre-Garriga, Ángela Vidal-Jordana, Ana Zabalza, Àlex Rovira, Xavier Montalban, Mar Tintoré, Marco Lorenzi, Deborah Pareto and Carmen Tur
A Landmarked Penalised Weibull Framework for Time-Updated Survival Modelling of Heterogeneous Clinical Outcomes in Multiple Sclerosis
Francisco José Aparicio Serrano, Ariadna Masot-Llima, Agustín Pappolla, Susana Otero-Romero, René Carvajal, Álvaro Cobo-Calvo, Manel Alberich, María Jesús Arévalo, Georgina Arrambide, Cristina Auger, Joaquín Castillo, Manuel Comabella, Ingrid Galán, Daniel Hernández, Carlos Nos, Jordi Río, Breogán Rodríguez-Acevedo, Jaume Sastre-Garriga, Ángela Vidal-Jordana, Ana Zabalza, Àlex Rovira, Xavier Montalban, Mar Tintoré, Marco Lorenzi, Deborah Pareto and Carmen Tur
- Keywords
- Survival AnalysisTime-Updated LandmarkingMultiple SclerosisTime-to-Event ModellingDisease Progression
Quantum-Assisted Clinical Risk Factor Selection for Diabetes Readmission Prediction
Floriano Caprio, Matteo Tortora, Paolo Soda and Sunil Gentyala
Quantum-Assisted Clinical Risk Factor Selection for Diabetes Readmission Prediction
Floriano Caprio, Matteo Tortora, Paolo Soda and Sunil Gentyala
- Keywords
- QAOAQUBOclinical feature selectiondiabetes readmissionmedical informatics
Quantum Machine Learning for Missense Variant Pathogenicity Prediction in the KCNQ Gene Family
Markel Garcia, Sara Capponi, Aitor Bergara and Aritz Leonardo
Quantum Machine Learning for Missense Variant Pathogenicity Prediction in the KCNQ Gene Family
Markel Garcia, Sara Capponi, Aitor Bergara and Aritz Leonardo
- Keywords
- Quantum ComputingQuantum Machine LearningVariant Pathogenicity PredictionChannelopathies
Temporal quantum feature maps for longitudinal biomedical data
Maria Demidik, Filippo Utro, Alexey Galda, Daniel Blankenberg, Karl Jansen and Laxmi Parida
Temporal quantum feature maps for longitudinal biomedical data
Maria Demidik, Filippo Utro, Alexey Galda, Daniel Blankenberg, Karl Jansen and Laxmi Parida
- Keywords
- quantum kernelslongitudinal datatemporal learningIQP circuitsbiomedical applications
Zero-shot phenotype prediction and screening of engineered immune receptor variants using contrastive learning-based workflow
Jie Shi, Yenho Chen, Kyle Daniels, Shangying Wang and Sara Capponi
Zero-shot phenotype prediction and screening of engineered immune receptor variants using contrastive learning-based workflow
Jie Shi, Yenho Chen, Kyle Daniels, Shangying Wang and Sara Capponi
- Keywords
- Machine LearningContrastive LearningCAR T cellsCell therapyco-stimulatory signaling domains
Diagnosing Rejection Collapse via Uncertainty Decomposition
Walter Endrizzi, Flavio Ragni, Stefano Bovo, Monica Moroni, Giuseppe Jurman and Venet Osmani
Diagnosing Rejection Collapse via Uncertainty Decomposition
Walter Endrizzi, Flavio Ragni, Stefano Bovo, Monica Moroni, Giuseppe Jurman and Venet Osmani
- Keywords
- UncertaintyRejection AnalysisStratificationPrognostic ModelingParkinson’s Disease
Federated vs. Centralized Learning: A Comparative Prediction Analysis for Hospitalization Cost
Shinto Pulickal Thomas, Jiregna Olani Kedida, Thierry Bossy, Juan Ramon Troncoso-Pastoriza, Alessandro Desideria, Dario Gregori and Corrado Lanera
Federated vs. Centralized Learning: A Comparative Prediction Analysis for Hospitalization Cost
Shinto Pulickal Thomas, Jiregna Olani Kedida, Thierry Bossy, Juan Ramon Troncoso-Pastoriza, Alessandro Desideria, Dario Gregori and Corrado Lanera
- Keywords
- Federated LearningGradient-Based OptimizationCost AssessmentMean Absolute Error
Lexical and Embedding Similarity for ICD to Hetionet Disease Alignment: A Controlled Multi-View Study
Amani Braham and Insaf Tnazefti
Lexical and Embedding Similarity for ICD to Hetionet Disease Alignment: A Controlled Multi-View Study
Amani Braham and Insaf Tnazefti
- Keywords
- Entity alignmentKnowledge graphsSemantic IntegrationBiomedical Informatics
LeProteine: Resource-Aware Protein Binder Design Pipelines on Blackwell GPU
Al Hamna Asif, Rana Abu Bakar, Filippo Cugini, Juan Jose Vegas Olmos and Amna Younus
LeProteine: Resource-Aware Protein Binder Design Pipelines on Blackwell GPU
Al Hamna Asif, Rana Abu Bakar, Filippo Cugini, Juan Jose Vegas Olmos and Amna Younus
- Keywords
- protein binder designDGX SparkBlackwell GPUreproducibilityhybrid executiongpucpu
NoduleFlow: Automatic 3D Longitudinal Lung Nodule Detection and Segmentation
Giulia Romoli, Filippo Ruffini, Francesco Di Feola, Valerio Guarrasi, Matteo Tortora and Paolo Soda
NoduleFlow: Automatic 3D Longitudinal Lung Nodule Detection and Segmentation
Giulia Romoli, Filippo Ruffini, Francesco Di Feola, Valerio Guarrasi, Matteo Tortora and Paolo Soda
- Keywords
- Lung cancerLongitudinal ImagingDetectionSegmentation
rSignificativity: an R package to compute significativity indexes for agreement values
Roberto Pagliarini and Alberto Casagrande
rSignificativity: an R package to compute significativity indexes for agreement values
Roberto Pagliarini and Alberto Casagrande
- Keywords
- Agreement measuresSignificativityMarginal SignificativityProbabilityR package
A comparative knowledge‑ and data‑driven machine learning analysis for H-CPAP failure prediction
Keying Qiao, Marco Masseroli, Luca Novelli, Matteo Bertoli, Roberto Cosentini, Eugenia Belotti, Isabelle Piazza, Chiara Baroni, Matteo Stenico, Fabiano Di Marco, Michele Capelli and Silvia Cascianelli
A comparative knowledge‑ and data‑driven machine learning analysis for H-CPAP failure prediction
Keying Qiao, Marco Masseroli, Luca Novelli, Matteo Bertoli, Roberto Cosentini, Eugenia Belotti, Isabelle Piazza, Chiara Baroni, Matteo Stenico, Fabiano Di Marco, Michele Capelli and Silvia Cascianelli
- Keywords
- Helmet-Continuous Positive Airway PressureRisk stratificationMachine Learning
ClinRAG-BiLSTM: An Agentic Cloud AI Framework for Explainable and Traceable Breast Cancer Decision Support
Adaleta Gicic and Dženana Đonko
ClinRAG-BiLSTM: An Agentic Cloud AI Framework for Explainable and Traceable Breast Cancer Decision Support
Adaleta Gicic and Dženana Đonko
- Keywords
- Breast cancerBayesian-BiLSTMRAGclinical decision supportexplainable AI
MedHE: Communication-Efficient Privacy-Preserving Federated Learning for Healthcare Text Classification
Farjana Yesmin
MedHE: Communication-Efficient Privacy-Preserving Federated Learning for Healthcare Text Classification
Farjana Yesmin
- Keywords
- Federated LearningHomomorphic EncryptionDifferential PrivacyHealthcare NLPGradient SparsificationCKKSMembership Inference AttackPrivacy-Preserving Machine Learning
NanoMambaNet: A Linear-Time CNN-Mamba Hybrid Architecture for Basecalling-Free Nanopore Pathogen Classification
Chayanika Bhattacharjee, Kolin Paul and Niranjan Sarode
NanoMambaNet: A Linear-Time CNN-Mamba Hybrid Architecture for Basecalling-Free Nanopore Pathogen Classification
Chayanika Bhattacharjee, Kolin Paul and Niranjan Sarode
- Keywords
- Nanopore SequencingState Space ModelsMambaContrastive LearningPathogen Identification
A Dependent Dirichlet Process Approach to Population Size Estimation
Lucia Gallucci
A Dependent Dirichlet Process Approach to Population Size Estimation
Lucia Gallucci
- Keywords
- Bayesian nonparametricscapture-recapturemissing data imputationmixture modelspopulation size imputation
Decoding HIV-1 Antibody Escape with Interpretable Protein Language Models
Fahsai Nakarin, Pawin Taechoyotin and Kayla G. Sprenger
Decoding HIV-1 Antibody Escape with Interpretable Protein Language Models
Fahsai Nakarin, Pawin Taechoyotin and Kayla G. Sprenger
- Keywords
- Protein Language ModelViral EscapeFitness LandscapeAntibodyExplainable AI
Comparative analysis of unsupervised multi-omics integration for pancreatic cancer subtype discovery
Gonçalo Costa-Ferreira, Marta B. Lopes and Eunice Carrasquinha
Comparative analysis of unsupervised multi-omics integration for pancreatic cancer subtype discovery
Gonçalo Costa-Ferreira, Marta B. Lopes and Eunice Carrasquinha
- Keywords
- Multi-omics integrationUnsupervised learningiClusterBayesMOFA+Pancreatic cancer
Functional Fuzzy Clustering of Longitudinal biological marker's trajectories
Silvia D'Elia, Valeria Paggi, Marco Alfò and Maria Brigida Ferraro
Functional Fuzzy Clustering of Longitudinal biological marker's trajectories
Silvia D'Elia, Valeria Paggi, Marco Alfò and Maria Brigida Ferraro
- Keywords
- Functional data analysisFuzzy clusteringLongitudinal trajectories
Multivariate Model-Based Landmarking Approach in Mixture Cure Models to Dynamic Prediction
Bianca Ferraro, Marco Alfò and Marta Cipriani
Multivariate Model-Based Landmarking Approach in Mixture Cure Models to Dynamic Prediction
Bianca Ferraro, Marco Alfò and Marta Cipriani
- Keywords
- mixture cure modelslandmarkingdynamic predictionmultivariate mixed-effects modelslongitudinal biomarker
Reducing the Size of Breastfeeding Reference Arms in Infant Nutrition Studies Using Prognostic Covariate Adjustment
Luca Lavalle
Reducing the Size of Breastfeeding Reference Arms in Infant Nutrition Studies Using Prognostic Covariate Adjustment
Luca Lavalle
- Keywords
- digital twinmachine learninginfant nutrition
Treatment persistence drives estimator performance in longitudinal causal inference based on observational data: A simulation study
Sergio Gaiotti, Sara Poletto, Enrico Longato, Erica Tavazzi and Martina Vettoretti
Treatment persistence drives estimator performance in longitudinal causal inference based on observational data: A simulation study
Sergio Gaiotti, Sara Poletto, Enrico Longato, Erica Tavazzi and Martina Vettoretti
- Keywords
- longitudinal causal inferencetime-varying confoundingtreatment persistence
An entropy based method for mapping mixed lineage next generation sequencing data
Gregory Fonseca
An entropy based method for mapping mixed lineage next generation sequencing data
Gregory Fonseca
- Keywords
- EntropyLineageSequence
Evaluating Locally-Deployable Large Language Models on Free-Text Data in an Italian Obstetric Context
Pierluigi Reali, Giulio Steyde, Gianluca Carta, Mark James Carman and Maria Gabriella Signorini
Evaluating Locally-Deployable Large Language Models on Free-Text Data in an Italian Obstetric Context
Pierluigi Reali, Giulio Steyde, Gianluca Carta, Mark James Carman and Maria Gabriella Signorini
- Keywords
- artificial intelligenceLLMunstructured datareal-world dataclinical notes
Integration of Large Language Models and Retrieval-Augmented Generation for Clinical Decision Support in Inflammatory Bowel Disease
Gabriele Rumi, Andrea Mazzei, Daniele Napolitano, Franco Scaldaferri, Giuseppe Cattaneo and Antonio Gasbarrini
Integration of Large Language Models and Retrieval-Augmented Generation for Clinical Decision Support in Inflammatory Bowel Disease
Gabriele Rumi, Andrea Mazzei, Daniele Napolitano, Franco Scaldaferri, Giuseppe Cattaneo and Antonio Gasbarrini
- Keywords
- Inflammatory bowel diseaseClinical NLPRetrieval-Augmented GenerationLarge Language ModelsClinical Decision Support
NASPER-NSV: A Temporal Neuro-Symbolic Verification Framework for Predictive Biomedical Intervention in Extreme Environments
Maharshi Patel Maharshi Patel
NASPER-NSV: A Temporal Neuro-Symbolic Verification Framework for Predictive Biomedical Intervention in Extreme Environments
Maharshi Patel Maharshi Patel
- Keywords
- neuro-symbolic AIpredictive medicinedigital twinbiomedical informatics
Leveraging hologenomic data for phenotypic prediction: potential and pitfalls
Solène Pety, Ingrid David, Andrea Rau and Mahendra Mariadassou
Leveraging hologenomic data for phenotypic prediction: potential and pitfalls
Solène Pety, Ingrid David, Andrea Rau and Mahendra Mariadassou
- Keywords
- holobiontmulti-omics integrationcompositional datamulti-generation
Longitudinal Microbiome Analysis in R: Review and Selection Pipeline
Irene Creus-Martí and Francisco José Santonja Gómez
Longitudinal Microbiome Analysis in R: Review and Selection Pipeline
Irene Creus-Martí and Francisco José Santonja Gómez
- Keywords
- Longitudinal dataCompositional dataR packagesMicrobiome
Predicting Alzheimer’s Progression Over Time Using Sheaf Neural Networks and Brain Graphs
Annamaria Defilippo, Giulia Avolio, Pierangelo Veltri, Pietro Liò and Pietro Hiram Guzzi
Predicting Alzheimer’s Progression Over Time Using Sheaf Neural Networks and Brain Graphs
Annamaria Defilippo, Giulia Avolio, Pierangelo Veltri, Pietro Liò and Pietro Hiram Guzzi
- Keywords
- Alzheimer's diseaseSheaf Neural NetworksTemporal graphsDisease progression forecasting
Prediction of treatment response in Multiple Sclerosis with Machine Learning models
Ariadna Masot-Llima, Francisco Aparicio-Serrano, Agustín Pappolla, Susana Otero-Romero, René Carvajal, Álvaro Cobo-Calvo, Manel Alberich, María Jesús Arévalo, Georgina Arrambide, Cristina Auger, Joaquín Castillo, Manuel Comabella, Ingrid Galán, Daniel Hernández, Carlos Nos, Jordi Río, Breogán Rodríguez-Acevedo, Jaume Sastre-Garriga, Ángela Vidal-Jordana, Ana Zabalza, Àlex Rovira, Xavier Montalban, Mar Tintoré, Xavier Lladó, Deborah Pareto and Carmen Tur
Prediction of treatment response in Multiple Sclerosis with Machine Learning models
Ariadna Masot-Llima, Francisco Aparicio-Serrano, Agustín Pappolla, Susana Otero-Romero, René Carvajal, Álvaro Cobo-Calvo, Manel Alberich, María Jesús Arévalo, Georgina Arrambide, Cristina Auger, Joaquín Castillo, Manuel Comabella, Ingrid Galán, Daniel Hernández, Carlos Nos, Jordi Río, Breogán Rodríguez-Acevedo, Jaume Sastre-Garriga, Ángela Vidal-Jordana, Ana Zabalza, Àlex Rovira, Xavier Montalban, Mar Tintoré, Xavier Lladó, Deborah Pareto and Carmen Tur
- Keywords
- Multiple SclerosisMachine LearningTreatment responsePredictive Models
Towards Federated and Explainable Quantum Machine Learning for Possible Epileptic Seizure Detection
Francesco Mercaldo, Hubert Schölnast, Oliver Eigner, Marta Petyx, Antonella Santone, Mario Cesarelli, Fabio Martinelli and Paul Tavolato
Towards Federated and Explainable Quantum Machine Learning for Possible Epileptic Seizure Detection
Francesco Mercaldo, Hubert Schölnast, Oliver Eigner, Marta Petyx, Antonella Santone, Mario Cesarelli, Fabio Martinelli and Paul Tavolato
- Keywords
- EpilepsyEEGQuantum Machine LearningFederated LearningExplainability
Estimating Oocyte Biological Age: An Explainable Deep Learning Approach for Assisted Reproduction
Andrea Piermartini, Donatella Firmani, Francesco Leotta, Claudio Manna and Silvia Raffi
Estimating Oocyte Biological Age: An Explainable Deep Learning Approach for Assisted Reproduction
Andrea Piermartini, Donatella Firmani, Francesco Leotta, Claudio Manna and Silvia Raffi
- Keywords
- image classificationexplainable AIassisted reproductive technologydeep learning
Growth-Aware CT Image Synthesis via Volume-Conditioned Diffusion Models
Francesco Di Feola, Lorenzo Carusone, Giulia Romoli, Filippo Ruffini, Massimiliano Mantegna, Valerio Guarrasi, Matteo Tortora and Paolo Soda
Growth-Aware CT Image Synthesis via Volume-Conditioned Diffusion Models
Francesco Di Feola, Lorenzo Carusone, Giulia Romoli, Filippo Ruffini, Massimiliano Mantegna, Valerio Guarrasi, Matteo Tortora and Paolo Soda
- Keywords
- Diffusion ModelsVolume ConditioningMedical Image SynthesisTumor Progression
A Sparse Rotationally Invariant Integrative Factor Model
Alejandra Avalos-Pacheco
A Sparse Rotationally Invariant Integrative Factor Model
Alejandra Avalos-Pacheco
- Keywords
- Factor RegressionData IntegrationSparsity
Informed Sparse Factorization for Psychometric Construct Recognition
Lorenzo Schiavon
Informed Sparse Factorization for Psychometric Construct Recognition
Lorenzo Schiavon
- Keywords
- Bayesian factorizationglobal-local shrinkageidentifiabilityorder invariancesparsity
Tensor decompositions for Bayesian factor models with dynamic loadings: an application to causal inference
Pantelis Samartsidis
Tensor decompositions for Bayesian factor models with dynamic loadings: an application to causal inference
Pantelis Samartsidis
- Keywords
- Factor analysisCausal inferenceTensor decomposition
Paraffinity - Semisynthetic Supervision for Automated Tissue Embedding Quality Control
Andrea Mazzei, Giorgia Sparapassi, Fabio Berrettoni, Davide Timperi, Dario Branco, Francesco Merolla, Stefania Staibano, Giuseppe Cattaneo and Gian Marco Di Domenico
Paraffinity - Semisynthetic Supervision for Automated Tissue Embedding Quality Control
Andrea Mazzei, Giorgia Sparapassi, Fabio Berrettoni, Davide Timperi, Dario Branco, Francesco Merolla, Stefania Staibano, Giuseppe Cattaneo and Gian Marco Di Domenico
- Keywords
- Digital PathologySynthetic DataQuality ControlDeep LearningParaffin Embedding
CellChat Hotspot: A Focused Lens on Tumor Microenvironment Communication
Dario Monaco, Mirea Dioguardi, Maria Dipalma, Eliseo Mattioli, Francesco Alfredo Zito, Francesco Giovannelli, Angela Ricco, Oronzo Brunetti, Antonella Argentiero and Simona De Summa
CellChat Hotspot: A Focused Lens on Tumor Microenvironment Communication
Dario Monaco, Mirea Dioguardi, Maria Dipalma, Eliseo Mattioli, Francesco Alfredo Zito, Francesco Giovannelli, Angela Ricco, Oronzo Brunetti, Antonella Argentiero and Simona De Summa
- Keywords
- Single-Cell RNA-seqCell-Cell CommunicationsCellChatTumor Microenvironment
Soft-GLCM: Stabilizing Differentiable Texture Analysis through Soft Quantization
Haoyue Chen, Giuseppe Salvaggio, Albert Comelli and Anthony Yezzi
Soft-GLCM: Stabilizing Differentiable Texture Analysis through Soft Quantization
Haoyue Chen, Giuseppe Salvaggio, Albert Comelli and Anthony Yezzi
- Keywords
- Texture AnalysisSoft-GLCMSoft QuantizationDifferentiable-GLCMRadiomics
Ontology-Aware Candidate Reranking for SNOMED CT Entity Linking in Clinical Notes
Luka Blašković, Nikola Tanković and Ivo Ipšić
Ontology-Aware Candidate Reranking for SNOMED CT Entity Linking in Clinical Notes
Luka Blašković, Nikola Tanković and Ivo Ipšić
- Keywords
- biomedical entity linkingSNOMED CTUMLSrerankingclinical NLP
Benchmarking differential abundance methods in respiratory microbiome: evaluating the impact of confounder adjustment
Céline Hosteins, John Barrera, Cristian Meza, Raphaël Enaud, Laurence Delhaes and Marta Avalos
Benchmarking differential abundance methods in respiratory microbiome: evaluating the impact of confounder adjustment
Céline Hosteins, John Barrera, Cristian Meza, Raphaël Enaud, Laurence Delhaes and Marta Avalos
- Keywords
- microbiomedifferential abundance analysisbenchmarkconfoundingrespiratory diseases
Microbiomes Under Human Pressure: An Integrated One Health Data Analysis Pipeline
Omar Orellana-Diaz, Gautier Chauvin, Luis Marti, Raphaël Enaud, Laurence Delhaes and Marta Avalos
Microbiomes Under Human Pressure: An Integrated One Health Data Analysis Pipeline
Omar Orellana-Diaz, Gautier Chauvin, Luis Marti, Raphaël Enaud, Laurence Delhaes and Marta Avalos
- Keywords
- Data analysis pipelineMachine learningMicrobiomeUrbanizationEnvironment
Assessing Community Stability in Weighted Single-Cell Networks
Valeria Policastro, Domenico Sgambato, Dario Righelli, Vincenzo della Gatta, Luisa Cutillo and Annamaria Carissimo
Assessing Community Stability in Weighted Single-Cell Networks
Valeria Policastro, Domenico Sgambato, Dario Righelli, Vincenzo della Gatta, Luisa Cutillo and Annamaria Carissimo
- Keywords
- community detectionweighted networksrobustnessclustering validationsingle cell
All Biases Are Not Equal: Implications for Equity in Emergency Department AI through a Gender Lens
Marta Avalos, Nadia Elorga-Castagnet, Lucía Gallego-Andrés, Cédric Gil-Jardiné, Ariel Guerra-Adames, Emilie Lesaine, Yamina Meziani, Marie Peyronnet and Andja Srebro
All Biases Are Not Equal: Implications for Equity in Emergency Department AI through a Gender Lens
Marta Avalos, Nadia Elorga-Castagnet, Lucía Gallego-Andrés, Cédric Gil-Jardiné, Ariel Guerra-Adames, Emilie Lesaine, Yamina Meziani, Marie Peyronnet and Andja Srebro
- Keywords
- Sex/Gender BiasEquityEthicsAI ActClinical AI Evaluation
Secondary structure design for efficient translation initiation
Tobias von der Haar and Corina Keller
Secondary structure design for efficient translation initiation
Tobias von der Haar and Corina Keller
- Keywords
- RNA Therapeuticssecondary structureribosomal scanningstart codon
Project-Aware Edited Nearest Neighbours for Label-Noise Filtering in SNP-Based Biogeographical Ancestry Inference
Cosimo Grazzini, Stefania Morelli, Elena Pilli, Giulia Cereda and Daniele Castellana
Project-Aware Edited Nearest Neighbours for Label-Noise Filtering in SNP-Based Biogeographical Ancestry Inference
Cosimo Grazzini, Stefania Morelli, Elena Pilli, Giulia Cereda and Daniele Castellana
- Keywords
- Biogeographical ancestrylabel-noise filteringProject-Aware ENNSNP-based classificationimperfect genomic data
SVC-Probe: A Framework for Evaluating Perturbation Generalization in Spatial Foundation-Model Embeddings
Jake Y. Chen, Huu Phong Nguyen, Fuad Al Abir and Ehsan Saghapour
SVC-Probe: A Framework for Evaluating Perturbation Generalization in Spatial Foundation-Model Embeddings
Jake Y. Chen, Huu Phong Nguyen, Fuad Al Abir and Ehsan Saghapour
- Keywords
- spatial virtual cellfoundation modelperturbation generalizationsubcellular proteomics
Toward Personalized Diabetic Retinopathy Screening: Deep Learning Fundus Image Analysis and Clinical Risk Factors
Flavio Ragni, Paolo Bresolin, Stefano Bovo, Madalina Cocu, Sandro Inchiostro, Federica Romanelli, Giulia Malfatti, Monica Moroni and Giuseppe Jurman
Toward Personalized Diabetic Retinopathy Screening: Deep Learning Fundus Image Analysis and Clinical Risk Factors
Flavio Ragni, Paolo Bresolin, Stefano Bovo, Madalina Cocu, Sandro Inchiostro, Federica Romanelli, Giulia Malfatti, Monica Moroni and Giuseppe Jurman
- Keywords
- diabetic retinopathypersonalized screeningdeep learning
LLM-Guided Integer Linear Programming for Antibody Library Design
Mikel Landajuela Larma
LLM-Guided Integer Linear Programming for Antibody Library Design
Mikel Landajuela Larma
- Keywords
- antibody library designinteger linear programminglarge language modelsmulti-objective optimizationstructure-guided design
Reinforcement Learning for Antibody Sequence Infilling
Chak Shing Lee, Conor F. Hayes, Denis Vashchenko and Mikel Landajuela Larma
Reinforcement Learning for Antibody Sequence Infilling
Chak Shing Lee, Conor F. Hayes, Denis Vashchenko and Mikel Landajuela Larma
- Keywords
- antibody designreinforcement learningsequence infillingoffline RLprotein language models
Extending Forge4Flame with Integrated Visualization and Customizable Airborne and Contact Transmission for Agent-Based Epidemic Simulations
Daniele Baccega, Irene Terrone, Matteo Costantino, Domiziana Porporato, Elisa Feyles, Irene Arduino, Rachele Francese, Manuela Donalisio, Marco Beccuti and Simone Pernice
Extending Forge4Flame with Integrated Visualization and Customizable Airborne and Contact Transmission for Agent-Based Epidemic Simulations
Daniele Baccega, Irene Terrone, Matteo Costantino, Domiziana Porporato, Elisa Feyles, Irene Arduino, Rachele Francese, Manuela Donalisio, Marco Beccuti and Simone Pernice
- Keywords
- Computational EpidemiologyDisease ModellingAgent-Based ModelGraphical User InterfaceFLAME GPU 2
Radiomics-Based Thyroid Nodule Classification with Label-Conditional Conformal Prediction
Laura Testa, Umberto Ferraro Petrillo and Pierpaolo Brutti
Radiomics-Based Thyroid Nodule Classification with Label-Conditional Conformal Prediction
Laura Testa, Umberto Ferraro Petrillo and Pierpaolo Brutti
- Keywords
- - Thyroid Nodule Classification- Ultrasound Imaging- Radiomics- Generalized Additive Model- Conformal Prediction
Discovering Heterogeneous Treatment Effects: A Bayesian Spike-and-Slab Approach
Dafne Zorzetto, Roberta De Vito and Falco J. Bargagli-Stoffi
Discovering Heterogeneous Treatment Effects: A Bayesian Spike-and-Slab Approach
Dafne Zorzetto, Roberta De Vito and Falco J. Bargagli-Stoffi
- Keywords
- causal inferencepublic healthhierarchical Bayesian modelingheterogeneous treatment effects
Self-Supervised Cytoplasmic Wave Detection Framework for In Vitro Fertilisation Embryo Time-Lapse
Attilio Di Vicino, Emanuel Di Nardo, Mario Cimmino, Riccardo Talevi, Sergio Chirico, Michele Di Capua and Angelo Ciaramella
Self-Supervised Cytoplasmic Wave Detection Framework for In Vitro Fertilisation Embryo Time-Lapse
Attilio Di Vicino, Emanuel Di Nardo, Mario Cimmino, Riccardo Talevi, Sergio Chirico, Michele Di Capua and Angelo Ciaramella
- Keywords
- contrastive learningIVF embryo selectioncytoplasmic wavesexplainable AIspatiotemporal representationembedding projection
Thyroid Nodule Segmentation in Ultrasound Images with Conformal Uncertainty Estimation
Francesco Teodori, Umberto Ferraro and Pierpaolo Brutti
Thyroid Nodule Segmentation in Ultrasound Images with Conformal Uncertainty Estimation
Francesco Teodori, Umberto Ferraro and Pierpaolo Brutti
- Keywords
- Thyroid Nodule SegmentationUltrasoundDeepLabV3+EfficientNet-B5Conformal Prediction
A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health
Oluwadara Adedeji, Michael Mayowa Farayola, Jeff Brozena, Irina Tal, Regina Connolly and Mark Matthews
A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health
Oluwadara Adedeji, Michael Mayowa Farayola, Jeff Brozena, Irina Tal, Regina Connolly and Mark Matthews
- Keywords
- Computational ethicsDigital phenotypingFinancial dataFinancial technologies
FLoRe: Fingerprint-based Long-Read Overlap Reconstruction
Manuel Sica, Rocco Zaccagnino and Rosalba Zizza
FLoRe: Fingerprint-based Long-Read Overlap Reconstruction
Manuel Sica, Rocco Zaccagnino and Rosalba Zizza
- Keywords
- De novo genome assemblylong-read overlap detectionseed-and-extendsequence fingerprintLyndon factorization
Privacy-preserving federated batch effect correction in omics data with ComBatFed
Yuliya Burankova, Mohammad Mahdi Kazemi, Linda Baumbach and Jan Baumbach
Privacy-preserving federated batch effect correction in omics data with ComBatFed
Yuliya Burankova, Mohammad Mahdi Kazemi, Linda Baumbach and Jan Baumbach
- Keywords
- federated learningbatch effect correctionComBatprivacy-preserving omicsmulti-center biomedical data analysis
Single-cell characterization of HERV-K transcriptional landscapes in acute and chronic HIV infection
Lucrezia Pierfederici, Elisabetta Lazzari, Gabriella Rozera, Lavinia Fabeni, Flavia Smoquina, Giulia Berno, Federica Forbici, Valentina Mazzotta, Andrea Antinori, Daniele Pietrucci, Daniele Maria Papetti, Fabrizio Maggi, Isabella Abbate and Giovanni Chillemi
Single-cell characterization of HERV-K transcriptional landscapes in acute and chronic HIV infection
Lucrezia Pierfederici, Elisabetta Lazzari, Gabriella Rozera, Lavinia Fabeni, Flavia Smoquina, Giulia Berno, Federica Forbici, Valentina Mazzotta, Andrea Antinori, Daniele Pietrucci, Daniele Maria Papetti, Fabrizio Maggi, Isabella Abbate and Giovanni Chillemi
- Keywords
- Human endogenous retrovirusesHERV-KHIV infectionsingle-cell transcriptomics
A Unified Framework for Bipartite Network Projection and Backbone Extraction, with an Application to Gene–Disease Associations
Panagiota Kontou, Ioannis Charatsiaris, Efstathios Voulgaris and Pantelis Bagos
A Unified Framework for Bipartite Network Projection and Backbone Extraction, with an Application to Gene–Disease Associations
Panagiota Kontou, Ioannis Charatsiaris, Efstathios Voulgaris and Pantelis Bagos
- Keywords
- bipartite networksprojection methodsbackbone extractiongene–disease associations
Bio-RL-GRN: Prior-Guided Reinforcement Learning for Gene Regulatory Network Inference
Tshepo Kitso Gobonamang, Okaile Marumo, Mavuna Sebapalo and Mpoeleng Dimane
Bio-RL-GRN: Prior-Guided Reinforcement Learning for Gene Regulatory Network Inference
Tshepo Kitso Gobonamang, Okaile Marumo, Mavuna Sebapalo and Mpoeleng Dimane
- Keywords
- GRNsreinforcement learningcomputational genomicsmulti-omics integrationexplainable AI
A Pipeline for Predicting Variant Associations with Complex Disorders via Functional Annotations
Francesco Gualdi, Zeno Darani, Daniele Malpetti, Marco Scutari and Francesca Mangili
A Pipeline for Predicting Variant Associations with Complex Disorders via Functional Annotations
Francesco Gualdi, Zeno Darani, Daniele Malpetti, Marco Scutari and Francesca Mangili
- Keywords
- genomicsartificial intelligencecomplex disordersgwas
Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causally Reliable Concept Bottleneck Models
Francesco De Santis, Daniele Malpetti, Francesco Gualdi and Francesca Mangili
Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causally Reliable Concept Bottleneck Models
Francesco De Santis, Daniele Malpetti, Francesco Gualdi and Francesca Mangili
- Keywords
- GWASNCDsInterpretability
Reliability-Aware Benchmarking of Self-Supervised Learning in Histopathology Under 1-10% Label Regimes
Alicem Koyun, Pranjali Changadeo Pingale and Sachin Mote
Reliability-Aware Benchmarking of Self-Supervised Learning in Histopathology Under 1-10% Label Regimes
Alicem Koyun, Pranjali Changadeo Pingale and Sachin Mote
- Keywords
- self-supervised learninghistopathologylabel efficiencycalibrationfederated learning
AutoEncoder based approach for the identification of genomic regions responsible for poorly described oligo-patient disorders
Joanna Szyda
AutoEncoder based approach for the identification of genomic regions responsible for poorly described oligo-patient disorders
Joanna Szyda
- Keywords
- anomaly detectionAutoEncoderchronic fatigue disorderDeep Learningwhole exome sequence
AI-assisted Single-Cell transcriptomics reveals persistent malignant CD4+ activity in Sézary Syndrome patient
Domenico Palumbo, Viola Melone, Luigi Palo, Carlo Ferravante, Giulia Salvatore, Cristina Cristofoletti, Maria Grazia Narducci and Roberta Tarallo
AI-assisted Single-Cell transcriptomics reveals persistent malignant CD4+ activity in Sézary Syndrome patient
Domenico Palumbo, Viola Melone, Luigi Palo, Carlo Ferravante, Giulia Salvatore, Cristina Cristofoletti, Maria Grazia Narducci and Roberta Tarallo
- Keywords
- Sézary syndromeCD4+Single CellAI
Machine learning identifies candidate microbiome features associated with rheumatoid arthritis risk despite limited external predictive performance
Raissa Silva, Rachel Audo, Axel Finckh, Claire Immediato Daien and Zubeyir Salis
Machine learning identifies candidate microbiome features associated with rheumatoid arthritis risk despite limited external predictive performance
Raissa Silva, Rachel Audo, Axel Finckh, Claire Immediato Daien and Zubeyir Salis
- Keywords
- Rheumatoid ArthritisMicrobiome dataMachine learning
Reinforcement Learning-Based Control Strategy for Boolean Gene Regulatory Networks
Michal Shoob, Hila Glazz, Avraham Raviv and Hillel Kugler
Reinforcement Learning-Based Control Strategy for Boolean Gene Regulatory Networks
Michal Shoob, Hila Glazz, Avraham Raviv and Hillel Kugler
- Keywords
- Gene Regulatory NetworksReinforcement LearningControl
Disease Classification from raw metagenomic data through Large Language Model embedding of DNA sequences
Gaspar Roy, Eugeni Belda, Yann Chevaleyre, Edi Prifti and Jean-Daniel Zucker
Disease Classification from raw metagenomic data through Large Language Model embedding of DNA sequences
Gaspar Roy, Eugeni Belda, Yann Chevaleyre, Edi Prifti and Jean-Daniel Zucker
- Keywords
- metagenomicsmicrobiomeLLMDeep Learningembedding
Expert Knowledge & Machine Understanding: Bridging Reactome’s Ontology with LLM Semantic Embeddings
Susanna Bravi, Riccardo De Luca, Rosa Sicilia, Christine Nardini and Mario Santoro
Expert Knowledge & Machine Understanding: Bridging Reactome’s Ontology with LLM Semantic Embeddings
Susanna Bravi, Riccardo De Luca, Rosa Sicilia, Christine Nardini and Mario Santoro
- Keywords
- knowledgebaseReactomeLLMhierarchyembedding
Identifying functional drivers of Hepatoblastoma outcomes via agent-based modeling and transcriptomics
Alessandro Ravoni, Yuanhua Liu, Stefano Cairo, Filippo Castiglione and Christine Nardini
Identifying functional drivers of Hepatoblastoma outcomes via agent-based modeling and transcriptomics
Alessandro Ravoni, Yuanhua Liu, Stefano Cairo, Filippo Castiglione and Christine Nardini
- Keywords
- Agent-based modelTranscriptomicsFactor analysisHepatoblastomaImmune system
Learned Indexing for Genome Exact Mapping: A Comparative Study
Zeinab Hassaan, Eman Hassan, Baker Mohammad and Lobna Said
Learned Indexing for Genome Exact Mapping: A Comparative Study
Zeinab Hassaan, Eman Hassan, Baker Mohammad and Lobna Said
- Keywords
- Genome analysislearned indexingmachine learninghuman genome
A study on the use of learning techniques and explainability methods on TMS-EEG data
Sara Nocco, Eleonora Arrigoni, Diego Corna, Francesca Gasparini, Alberto Pisoni, Leonor Josefina Romero Lauro and Aurora Saibene
A study on the use of learning techniques and explainability methods on TMS-EEG data
Sara Nocco, Eleonora Arrigoni, Diego Corna, Francesca Gasparini, Alberto Pisoni, Leonor Josefina Romero Lauro and Aurora Saibene
- Keywords
- TMS-EEGBrain ConnectivityEEGNetGraph Neural Networks (GNNs)
DirHopGNN: an efficient approach to k-hop GNNs
Francesco Madeddu and Mario Santoro
DirHopGNN: an efficient approach to k-hop GNNs
Francesco Madeddu and Mario Santoro
- Keywords
- GNNsk-hopKnowledge-GraphsScalable Graph LearningGraph Batching
CyChat: a conversational Cytoscape app for no-code, reproducible network analysis
Jeanine Liebold, Merle Stahl, Jan-Ole Schulze, Mohammad Mehdi Razavi, Gary D. Bader, Stefan Kurtz and Jan Baumbach
CyChat: a conversational Cytoscape app for no-code, reproducible network analysis
Jeanine Liebold, Merle Stahl, Jan-Ole Schulze, Mohammad Mehdi Razavi, Gary D. Bader, Stefan Kurtz and Jan Baumbach
- Keywords
- Cytoscapelarge language modelsautonomous agentsnetwork biologyreproducibility
iPerturb: An interpretable framework for predicting gene expression changes under perturbations
Khang Ta Gia and Linh Huynh Viet
iPerturb: An interpretable framework for predicting gene expression changes under perturbations
Khang Ta Gia and Linh Huynh Viet
- Keywords
- gene expression predictiongenetic perturbationgene regulatory networkbiologically informed neural networkinterpretable AI
LLMs Fail at Protein Sequence Annotation and Localization: A Comparative Study against a Tool-Augmented Agentic solution
Asia Leuzzi, Elisa Ficarra and Marta Lovino
LLMs Fail at Protein Sequence Annotation and Localization: A Comparative Study against a Tool-Augmented Agentic solution
Asia Leuzzi, Elisa Ficarra and Marta Lovino
- Keywords
- Agentic AILarge language modelsProtein sequence analysisRetrieval-augmented generationSequence localization
SpatialJEPA: JEPA-inspired graph-context distillation for spatially aware multiomics integration
Dylan Mann-Krzisnik and Yue Li
SpatialJEPA: JEPA-inspired graph-context distillation for spatially aware multiomics integration
Dylan Mann-Krzisnik and Yue Li
- Keywords
- spatial transcriptomicsself-supervised learningmultiomics integration
Modeling Overdispersed Zero-Inflated Longitudinal Count Data: A Beta-Binomial Mixed-Effects Approach
John Barrera, Ana Arribas-Gil, Dae-Jin Lee and Cristian Meza
Modeling Overdispersed Zero-Inflated Longitudinal Count Data: A Beta-Binomial Mixed-Effects Approach
John Barrera, Ana Arribas-Gil, Dae-Jin Lee and Cristian Meza
- Keywords
- zero-inflated beta-binomialmixed-effects modelsSAEM algorithmmicrobiomelongitudinal count data
How Much Architecture Does a Protein Transformer Need to Self-Discover?
Giansalvo Cirrincione, Elisa Ficarra and Marta Lovino
How Much Architecture Does a Protein Transformer Need to Self-Discover?
Giansalvo Cirrincione, Elisa Ficarra and Marta Lovino
- Keywords
- protein language modelsself-architecting transformersattention head growthablation studyPfam classification
Pixel-Wise Nuclei Localization in Cervical Cytology with Hierarchical Feature Representations
Ciro Russo, Alessandro Bria and Claudio Marrocco
Pixel-Wise Nuclei Localization in Cervical Cytology with Hierarchical Feature Representations
Ciro Russo, Alessandro Bria and Claudio Marrocco
- Keywords
- Cervical CytologyNuclei DetectionVision TransformersDense Object DetectionComputational Cytology
GreatORCA: a framework for biological laboratory data analysis and integration for computational modelling
Dora Tortarolo, Simone Pernice, Lorenzo Pace, Sandro Gepiro Contaldo, Giulia Burrone, Fabiana Clapero, Donatella Valdembri, Guido Serini, Federica Riccardo, Lidia Tarone, Marco Beccuti and Francesca Cordero
GreatORCA: a framework for biological laboratory data analysis and integration for computational modelling
Dora Tortarolo, Simone Pernice, Lorenzo Pace, Sandro Gepiro Contaldo, Giulia Burrone, Fabiana Clapero, Donatella Valdembri, Guido Serini, Federica Riccardo, Lidia Tarone, Marco Beccuti and Francesca Cordero
- Keywords
- Computational ModelsData AnalysisPetri NetsFAIR PrinciplesReproducibility
A quantum generative model for in silico clinical trials using scarce training datasets
Olatz Sanz, Reza Dastbasteh, Mikel Hernaez, Roberto Sanchez-Navarro, Maria Diez-Campelo, Felipe Prosper, Ana Alfonso-Pierola, Sara Capponi, Pedro Crespo Bofill and Josu Etxezarreta Martinez
A quantum generative model for in silico clinical trials using scarce training datasets
Olatz Sanz, Reza Dastbasteh, Mikel Hernaez, Roberto Sanchez-Navarro, Maria Diez-Campelo, Felipe Prosper, Ana Alfonso-Pierola, Sara Capponi, Pedro Crespo Bofill and Josu Etxezarreta Martinez
- Keywords
- in silicoMyelodysplastic Syndromequantum generative models
Interpretable Convolutional Neural Networks Reveal Sequence Determinants of SF3B1-Mutant Susceptible Cryptic Splice Junctions in Cancer
Luca Faretra, Francesco Napolitano, Massimo Pancione and Luigi Cerulo
Interpretable Convolutional Neural Networks Reveal Sequence Determinants of SF3B1-Mutant Susceptible Cryptic Splice Junctions in Cancer
Luca Faretra, Francesco Napolitano, Massimo Pancione and Luigi Cerulo
- Keywords
- SF3B1CancerAlternative SplicingConvolutional Neural NetworkBranchsite
GLD-VAE: A Guided Latent Diffusion Variational Autoencoder for Class Imbalance in Multi-Omics Data
Durga Parkhi, Srijay Deshpande and Animesh Acharjee
GLD-VAE: A Guided Latent Diffusion Variational Autoencoder for Class Imbalance in Multi-Omics Data
Durga Parkhi, Srijay Deshpande and Animesh Acharjee
- Keywords
- Generative AIMulti-omicsAI
Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis
Rafał Buler, Jakub Buler, Maciej Bobowicz and Michał Grochowski
Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis
Rafał Buler, Jakub Buler, Maciej Bobowicz and Michał Grochowski
- Keywords
- graph neural networksrelational inductive biasmultiple instance learningself-supervised learningskin cancer classificationcomputer visiondeep neural networks
Bridging Generative AI and Synthesis-on-Demand Chemical Spaces for Structure-based De Novo Drug Design
Cosmin Ichim, Viet-Khoa Tran-Nguyen and Pedro Ballester
Bridging Generative AI and Synthesis-on-Demand Chemical Spaces for Structure-based De Novo Drug Design
Cosmin Ichim, Viet-Khoa Tran-Nguyen and Pedro Ballester
- Keywords
- De Novo Drug DesignGenerative AIChemical SpacesDocking
Identification of Allosteric Binding Pockets and Ligand-Induced Network Rewiring in PHF6
Lisia Peqini, Alessio Ferrini, Lan Phan Bach Tun, Jacopo Miotto, Massimo Bellanda and Damiano Piovesan
Identification of Allosteric Binding Pockets and Ligand-Induced Network Rewiring in PHF6
Lisia Peqini, Alessio Ferrini, Lan Phan Bach Tun, Jacopo Miotto, Massimo Bellanda and Damiano Piovesan
- Keywords
- Virtual ScreeningADMEMD simulationBinding Free EnergyRIN analysis
A topological object-proposal pipeline for cellular fiber segmentation
Riccardo Ceccaroni, Valerio Reffo, Camilla Pezzini, Roberta Sartori, Marco Sandri, Pierpaolo Brutti and Davide Risso
A topological object-proposal pipeline for cellular fiber segmentation
Riccardo Ceccaroni, Valerio Reffo, Camilla Pezzini, Roberta Sartori, Marco Sandri, Pierpaolo Brutti and Davide Risso
- Keywords
- cellular fiber segmentationpersistent homologyGPU-accelerated algorithmH\&E imaginginstance segmentation
CudaMon: An R Package to Monitor NVIDIA GPUs, Showcased by Monitoring a GPU-accelerated Single-cell Analysis Workflow in R
Mohammad Amin Zadenoori, Riccardo Ceccaroni, Davide Risso and Gabriele Sales
CudaMon: An R Package to Monitor NVIDIA GPUs, Showcased by Monitoring a GPU-accelerated Single-cell Analysis Workflow in R
Mohammad Amin Zadenoori, Riccardo Ceccaroni, Davide Risso and Gabriele Sales
- Keywords
- GPU monitoringNVIDIA GPUsR packageperformance analysisreproducible computingsingle-cell analysis
OmicsGPT: Multi-Agent Bioinformatics Orchestration for Evidence-Grounded RNA-seq Interpretation
Patrick Roney and James Li
OmicsGPT: Multi-Agent Bioinformatics Orchestration for Evidence-Grounded RNA-seq Interpretation
Patrick Roney and James Li
- Keywords
- multi-agent AIRNA-seqbioinformaticsRAGoncology
Identifying relevant Facial Action Units associated with Depression for Generating Synthetic Images using LLMs
Jona Maximilian Kao and Michael Gertz
Identifying relevant Facial Action Units associated with Depression for Generating Synthetic Images using LLMs
Jona Maximilian Kao and Michael Gertz
- Keywords
- AU categorizationimage generations with LLMsdepression detectionsynthetic Data Generation
Manufacturability of mRNA Therapeutics: Native vs. N-1-Methyl-Pseudouridine
Roland Huber, Kuo Chieh Liao and Yue Wan
Manufacturability of mRNA Therapeutics: Native vs. N-1-Methyl-Pseudouridine
Roland Huber, Kuo Chieh Liao and Yue Wan
- Keywords
- mRNA manufacturabilityTranscription TerminationN-1-Methylpseudouridine
ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information
Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano and Marcello Maggiolini
ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information
Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano and Marcello Maggiolini
- Keywords
- Agentic AIRetrieval-Augmented GenerationClinical TrialsReActMedical Informatics
DQIS: Byzantine Fault Tolerance for Multi-Channel Immune Surveillance — A Quantum-Inspired Quorum-Based Framework Grounded on Melanoma, GBM, and PDAC Single-Cell RNA-Seq Data
Alex Morgan
DQIS: Byzantine Fault Tolerance for Multi-Channel Immune Surveillance — A Quantum-Inspired Quorum-Based Framework Grounded on Melanoma, GBM, and PDAC Single-Cell RNA-Seq Data
Alex Morgan
- Keywords
- Byzantine Fault Tolerancequorum logicimmune surveillancechannel independencesingle-cell RNA-seq
Exploring the potential of Bayesian Gaussian Mixture integration for Variational Autoencoders in the generation of synthetic tabular data
Samuele Magro, Massimo Bilancia and Barbara Cafarelli
Exploring the potential of Bayesian Gaussian Mixture integration for Variational Autoencoders in the generation of synthetic tabular data
Samuele Magro, Massimo Bilancia and Barbara Cafarelli
- Keywords
- Synthetic dataVariational AutoencoderGenerative Adversarial Network
Learning Invariant Biological Dynamics Across Environmentswith Physics-Aware Neural Differential Equations
Okaile Marumo, Mavuna Sebapalo, Tshepo Gobonamang and Pendukeni Phalayagae
Learning Invariant Biological Dynamics Across Environmentswith Physics-Aware Neural Differential Equations
Okaile Marumo, Mavuna Sebapalo, Tshepo Gobonamang and Pendukeni Phalayagae
- Keywords
- neural ordinary differential equationsinvariant learningbiological dynamicsphysics-constrained machine learningcross-environment generalisation
Analysis of Prompt Engineering for Drug Toxicity Prediction
Mia MacGregor, Aakash Welgamage Don and Mark Bartlett
Analysis of Prompt Engineering for Drug Toxicity Prediction
Mia MacGregor, Aakash Welgamage Don and Mark Bartlett
- Keywords
- Prompt Engineering Machine Learning Drug Toxicity
A new MRI-based feature for quantifying the Diffuse Low-Grade Glioma brain infiltration and discriminating patterns of patients
Jean-Marie Moureaux, Sophie Wantz-Mezieres, Cyril Brzenczek, Marie Blonski and Luc Taillandier
A new MRI-based feature for quantifying the Diffuse Low-Grade Glioma brain infiltration and discriminating patterns of patients
Jean-Marie Moureaux, Sophie Wantz-Mezieres, Cyril Brzenczek, Marie Blonski and Luc Taillandier
- Keywords
- clustering hierarchical kmeans brain diffuse low-grade glioma (DLGG)
Spatial host pathogen transcriptomics maps bacterial burden linked urothelial injury and repair in bladder infection
Luying Su, Gregory Fonseca and Hao Zhou
Spatial host pathogen transcriptomics maps bacterial burden linked urothelial injury and repair in bladder infection
Luying Su, Gregory Fonseca and Hao Zhou
- Keywords
- Urinary tract infection spatial transcriptomics Urothelial exfoliation
A ML-powered Multiscale Computational Platform Based on QSP and PBPK Modeling to Support the Development of mRNA-based Therapies
Elisa Pettinà, Frederick Abi Chahine, Elio Campanile, Stefano Giampiccolo and Luca Marchetti
A ML-powered Multiscale Computational Platform Based on QSP and PBPK Modeling to Support the Development of mRNA-based Therapies
Elisa Pettinà, Frederick Abi Chahine, Elio Campanile, Stefano Giampiccolo and Luca Marchetti
- Keywords
- mRNA therapies Physiologically Based Pharmacokinetics (PBPK) modeling Quantitative Systems Pharmacology (QSP) modeling Machine Learning (ML) Mathematical modeling
Functional Shapley Representations for Longitudinal Random Survival Forests using Hierarchical FPCA
Simon Grabner, Markus Löcher, Andrea Berghold and Bastian Pfeifer
Functional Shapley Representations for Longitudinal Random Survival Forests using Hierarchical FPCA
Simon Grabner, Markus Löcher, Andrea Berghold and Bastian Pfeifer
- Keywords
- Random Forestlongitudinal dataShapley valuesFPCA
Statistical Model Checking of Uncertain Continuous Time Markov Chain in Systems Biology
Lauren Bentley and Krishnendu Ghosh
Statistical Model Checking of Uncertain Continuous Time Markov Chain in Systems Biology
Lauren Bentley and Krishnendu Ghosh
- Keywords
- Random Forest longitudinal data Shapley values FPCA
On the adoption of Quantum Machine Learning for Privacy-Preserving and Explainable Diabetic Retinopathy Detection and Localisation
Fabio Martinelli, Antonella Santone, Mario Cesarelli, Marta Petyx and Francesco Mercaldo
On the adoption of Quantum Machine Learning for Privacy-Preserving and Explainable Diabetic Retinopathy Detection and Localisation
Fabio Martinelli, Antonella Santone, Mario Cesarelli, Marta Petyx and Francesco Mercaldo
- Keywords
- Diabetic Retinopathy Quantum Machine Learning Federated Learning Explainability Medical Imaging
NARCOD: Non-Arbitrarily Reproducible Clustering of transcriptOmics Data
Christophe Le Priol
NARCOD: Non-Arbitrarily Reproducible Clustering of transcriptOmics Data
Christophe Le Priol
- Keywords
- clustering reproducibility seed stochasticity transcriptomics
Multimodal Learning from Temporally Grounded Narrative Events
Sneha Jha
Multimodal Learning from Temporally Grounded Narrative Events
Sneha Jha
- Keywords
- multimodal fusionclinical temporal NLPcross-attentiontemporal generalization
Joint Bayesian Modelling for Neuroimaging Meta-Analysis
Silvia Montagna, Felice Lamberti and Saverio Ranciati
Joint Bayesian Modelling for Neuroimaging Meta-Analysis
Silvia Montagna, Felice Lamberti and Saverio Ranciati
- Keywords
- Neuroimaging meta-analysis Bayesian hierarchical modelling Latent factor models Spatial statistics
Comparative Deep Learning and Statistical Analysis for Spatial Prediction of Lymph Node Metastasis in Gastric Cancer
Gloria Santoro
Comparative Deep Learning and Statistical Analysis for Spatial Prediction of Lymph Node Metastasis in Gastric Cancer
Gloria Santoro
- Keywords
- Deep learning Predictive Modeling Spatial Heterogeneity Gastric Cancer
An unsupervised clustering analysis of breast cancer data derived from electronic health records enhanced through UMAP dimensionality reduction
Davide Chicco and Nicoletta Benvenuto
An unsupervised clustering analysis of breast cancer data derived from electronic health records enhanced through UMAP dimensionality reduction
Davide Chicco and Nicoletta Benvenuto
- Keywords
- clustering DBSCAN electronic health records dimensionality reduction UMAP breast cancer cancer
Privacy-Preserving Detection of Rare Disease-Associated Cell Subsets via Secure Multi-Party Computation
Seyma Selcan Magara, Esther Havemann, Debora Jutz, Ali Burak Unal and Mete Akgün
Privacy-Preserving Detection of Rare Disease-Associated Cell Subsets via Secure Multi-Party Computation
Seyma Selcan Magara, Esther Havemann, Debora Jutz, Ali Burak Unal and Mete Akgün
- Keywords
- clustering DBSCAN electronic health records dimensionality reduction UMAP breast cancer cancer
Gimme a rainy week life more: a computational study on the signal-to-noise ratio of epigenetic aging associations
Marco Roccetti
Gimme a rainy week life more: a computational study on the signal-to-noise ratio of epigenetic aging associations
Marco Roccetti
- Keywords
- Epigenetic Aging Biostatistical Significance Clinical Relevance Stochastic Lifespan Noise Structural Variance Decomposition
Posters
Poster contributions
CONTOUR-GUIDED SEGMENTATION AND X-RAY IMAGE VALIDATION FOR PNEUMONIA DETECTION USING MOBILE DEEP LEARNING IN LOW-RESOURCE SETTINGS
Gebregziabihier Nigusie, Birhane, Mizan-Tepi University* Marawan Elbatel, The Hong Kong University of Science and Technology Yohannes Ayana, Amplitude Ventures
CONTOUR-GUIDED SEGMENTATION AND X-RAY IMAGE VALIDATION FOR PNEUMONIA DETECTION USING MOBILE DEEP LEARNING IN LOW-RESOURCE SETTINGS
Gebregziabihier Nigusie, Birhane, Mizan-Tepi University* Marawan Elbatel, The Hong Kong University of Science and Technology Yohannes Ayana, Amplitude Ventures
- Keywords
- ROIEdge Density AnalysisMobileVitMobile Prediction
- Abstract
- Pneumonia is an infectious disease that affects the lungs and is commonly diagnosed using chest X-ray imaging. However, the availability of radiologists to interpret these images is often limited, especially in rural areas of developing countries. To mitigate this, research has explored computer vision models for automated pneumonia detection. Key challenges remain in accurately identifying the region of interest (ROI), deploying models on mobile devices, and verifying that input images are original and unaltered to maintain prediction reliability. In this study, we have developed a novel algorithm that validates X-ray images based on heuristic values and identifies anatomical regions through a contour-guided ROI algorithm. Experiments were conducted using both contour-guided ROI-applied images and original images on MobileNet and MobileViT architectures. Among the two models, MobileViT, a hybrid vision transformer achieved optimal performance, with an accuracy of 85.3% on contour-guided ROI images for three-class pneumonia classification, viral, bacterial, and normal cases. To enhance practical usability, we developed a mobile application for model deployment and easy access. The app automatically verifies whether the input is an original X-ray image and generates a summary of the prediction results with Processing a single image takes an average of 3 seconds and 1 MB of RAM on a mid-range device, making deployment feasible in resource-limited environments. This solution represents a valuable contribution toward improving pneumonia diagnosis in rural and resource-limited settings.
Explainable Deep-Learning on condition specific expression profiles reveals critical cytosines in gene regulation
Veerbhan, Kesarwani, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Academy of Scientific and Innovative Research (AcSIR), India. Anchit, Kumar, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Akanksha, Sharma, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Sagar, Gupta, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Academy of Scientific and Innovative Research (AcSIR), India. Kirti, Pandey, Integrative Plant AdaptOmics Lab (iPAL), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (IHBT), Palampur, Himachal Pradesh, India. Academy of Scientific and Innovative Research (AcSIR), India. Gaurav, Zinta, Integrative Plant AdaptOmics Lab (iPAL), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (IHBT), Palampur, Himachal Pradesh, India. Academy of Scientific and Innovative Research (AcSIR), India. Ravi, Shankar, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Academy of Scientific and Innovative Research (AcSIR), India.
Explainable Deep-Learning on condition specific expression profiles reveals critical cytosines in gene regulation
Veerbhan, Kesarwani, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Academy of Scientific and Innovative Research (AcSIR), India. Anchit, Kumar, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Akanksha, Sharma, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Sagar, Gupta, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Academy of Scientific and Innovative Research (AcSIR), India. Kirti, Pandey, Integrative Plant AdaptOmics Lab (iPAL), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (IHBT), Palampur, Himachal Pradesh, India. Academy of Scientific and Innovative Research (AcSIR), India. Gaurav, Zinta, Integrative Plant AdaptOmics Lab (iPAL), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (IHBT), Palampur, Himachal Pradesh, India. Academy of Scientific and Innovative Research (AcSIR), India. Ravi, Shankar, Studio of Computational Biology & Bioinformatics, The Himalayan Centre for High-throughput Computational Biology, (HiCHiCoB, A BIC supported by DBT, India), Biotechnology Division, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Palampur (HP), 176061, India. Academy of Scientific and Innovative Research (AcSIR), India.
- Keywords
- DNA methylationdeep learningResNetcritical cytosinesgene regulation
- Abstract
- Objective: DNA methylation at cytosines is a vital epigenetic regulator in plants, yet identifying which specific cytosines (defined here as "critical cytosines") exert the most significant influence on downstream gene expression remains a major challenge. Current computational tools focus primarily on predicting methylation from expression, neglecting the reverse relationship. This study aimed to develop an explainable deep-learning framework, CritiCal-C, to map condition-specific promoter methylation patterns to gene expression levels and identify the specific regulatory switches that drive transcriptional outcomes across diverse plant species. Methods: The study integrated a massive multi-omics dataset comprising 232 Whole-Genome Bisulfite Sequencing (WGBS) and 260 corresponding RNA-seq samples from Arabidopsis thaliana and Oryza sativa (rice) across 85 distinct biological conditions. We benchmarked multiple architectures, including XGBoost, CNN, and DenseNet, ultimately selecting a ResNet-9 model to effectively capture spatial patterns while mitigating overfitting. DNA sequences (2kb upstream) were one-hot encoded with a specialized five-channel representation to include methylated cytosine states ('M'). To identify critical cytosines, we utilized two explainability strategies: (1) systematic in-silico knockout analysis of individual cytosines, pairs, and 100bp windows to measure their impact on predictive accuracy, and (2) Gradient-weighted Class Activation Mapping (Grad-CAM) to visualize high-weight regulatory features. Results: The optimized ResNet-9 model achieved high predictive performance, with Spearman’s correlation values reaching 0.94 for A. thaliana and 0.95 for O. sativa. A key discovery was that GC content similarity in promoter regions, rather than sequence homology, is the primary determinant for successful cross-species expression prediction. Critical cytosines were found to be non-randomly distributed, significantly concentrating within the 1kb region upstream of the transcription start site (TSS). Wet-lab validation via targeted bisulfite sequencing of the heat-responsive gene AT1G20400 confirmed that six of seven computationally identified critical cytosines underwent dramatic hypermethylation under heat stress, directly correlating with a significant drop in gene expression. Conclusion: We established CritiCal-C as a pioneering resource for identifying functional epigenetic regulators. By providing an interpretable framework to decode the impact of specific cytosines, this work enables researchers to move beyond broad association studies toward targeted molecular interventions. The model’s successful application to additional species (Brassica rapa, Cucumis sativus, and Solanum tuberosum) demonstrates its universal applicability for precision agriculture and the study of plant environmental adaptation.
ProInterVal-BioXtal: A Web Server to Distinguish Biological from Crystallographic Interfaces Using Representation Learning and Graph Neural Networks
*Defne Alnigenis, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450, Turkey Damla Ovek Baydar, Norwegian Centre for Molecular Biosciences and Medicine, University of Oslo, Oslo, 0349, Norway Ozlem Keskin, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450, Turkey Attila Gursoy, Department of Computer Engineering, Koç University, Istanbul, 34450, Turkey
ProInterVal-BioXtal: A Web Server to Distinguish Biological from Crystallographic Interfaces Using Representation Learning and Graph Neural Networks
*Defne Alnigenis, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450, Turkey Damla Ovek Baydar, Norwegian Centre for Molecular Biosciences and Medicine, University of Oslo, Oslo, 0349, Norway Ozlem Keskin, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450, Turkey Attila Gursoy, Department of Computer Engineering, Koç University, Istanbul, 34450, Turkey
- Keywords
- protein-protein interactionsstructural bioinformaticscomputational biologygraph neural networksdeep learningrepresentation learningprotein-protein interfacescrystallographic artifactsAI in biology
- Abstract
- Accurate discrimination between biologically relevant protein–protein interfaces (PPIs) and crystallographic contacts is essential for reliable interpretation of macromolecular assemblies and their cellular functions. While X-ray crystallography remains a primary method for determining protein complex structures, crystallographic interfaces may be formed as a byproduct of the crystal packing. Therefore, robust computational approaches for interface annotation are needed. This study introduces ProInterVal-BioXtal, a web server that predicts class scores for PPIs, differentiating between biologically relevant and crystallographic interfaces. The method leverages protein representation learning and graph-based deep learning to capture structural and physicochemical features. Interface graphs derived from input complexes are processed by our novel graph-based contrastive model to learn interface representations which are used by a graph neural network for classification. The model is trained and validated on the MANY benchmark dataset comprising 5739 dimers with a balanced distribution of biological and crystal interfaces and evaluated on the DC benchmark dataset with curated interfaces of similar interface areas. On the DC test set, ProInterVal-BioXtal achieves 88% accuracy, 88% precision, and 85% F1 score, outperforming state-of-the-art methods including DeepRank-GNN, PRODIGY-CRYSTAL, EPPIC 3, PISA, and QSAlign. On an independent benchmark, the method achieves 83% accuracy and 0.91 AUC, surpassing DeepRank-GNN (AUC = 0.85). Through a user-friendly web interface, ProInterVal-BioXtal enables rapid predictions from PDB files or IDs, returning probabilistic classification scores. Our tool addresses a fundamental challenge in structural biology by utilizing graph-based protein representation which can detect complex interactions and dependencies. ProInterVal-BioXtal server is freely available at https://3dpath.ku.edu.tr/prointerval-bioxtal/
Piecewise Regression Mixture Models with Skewness
Getachew Dagne, University of South Florida
Piecewise Regression Mixture Models with Skewness
Getachew Dagne, University of South Florida
- Keywords
- Mixture modelsskewnessgrowth curve
- Abstract
- This presentation focuses on mixture regression models with piecewise growth curves for assessing longitudinal data that exhibit multiphasic features. Some longitudinal data may have features that include skewness, measurement errors in covariates, and heterogenous population. Within a heterogeneous population, there is a desire to identify differential effects of covariates on a response variable in the context of a mixture of subpopulations. Regression mixture models are key methods for assessing differential effects of covariates. In this presentation, we will discuss the findings of extending regression mixture models to incorporate skew-normal distribution, measurement errors, and piecewise growth mixture modeling for describing multiphasic trajectories over time and analyzing real data from a clinical study
Predicting Third-Generation Cephalosporin Resistance in Hospitalised Patients Using Machine Learning: Evidence from Tanzania's AMR Surveillance Data, 2021–2025
Antidius P. Rwehumbiza*, School of Public Health, Kilimanjaro Christian Medical Centre (KCMC) University, Tanzania M. Burke, School of Public Health, Kilimanjaro Christian Medical Centre (KCMC) University, Tanzania F. Muro, School of Public Health, Kilimanjaro Christian Medical Centre (KCMC) University, Tanzania E. Shewiyo School of Public Health, Kilimanjaro Christian Medical Centre (KCMC) University, Tanzania
Predicting Third-Generation Cephalosporin Resistance in Hospitalised Patients Using Machine Learning: Evidence from Tanzania's AMR Surveillance Data, 2021–2025
Antidius P. Rwehumbiza*, School of Public Health, Kilimanjaro Christian Medical Centre (KCMC) University, Tanzania M. Burke, School of Public Health, Kilimanjaro Christian Medical Centre (KCMC) University, Tanzania F. Muro, School of Public Health, Kilimanjaro Christian Medical Centre (KCMC) University, Tanzania E. Shewiyo School of Public Health, Kilimanjaro Christian Medical Centre (KCMC) University, Tanzania
- Keywords
- machine learningantimicrobial resistancethird-generation cephalosporinsTanzaniaantimicrobial stewardshippredictive modelling
- Abstract
- Objectives Antimicrobial resistance (AMR) poses a critical threat to universal health coverage across low- and middle-income countries. In Tanzania, third-generation cephalosporin (3GC) resistance among gram-negative bacteria exceeds 50%, fuelled by empirical prescribing and severely limited diagnostic capacity, driving excess mortality and escalating costs. This study aimed to develop, validate, and interpret machine learning (ML) models using Tanzania's national AMR surveillance data to predict 3GC resistance at the point of prescribing, directly addressing a diagnostic gap that undermines stewardship efforts across sub-Saharan Africa. Methods Deidentified AMR surveillance data (2021–2025) were analysed from hospitalised patients at two Tanzanian tertiary facilities: Kilimanjaro Christian Medical Centre (KCMC) and Mbeya Zonal Referral Hospital (MZRH). Inclusion was restricted to the first clinical isolate per patient per pathogen per sample type within 30 days; mixed isolates were excluded. Five ML algorithms, gradient boosting decision tree (GBDT), logistic regression, random forest, LASSO, and k-nearest neighbours, were trained on 70% of KCMC data using 10-fold cross-validation with SMOTE class balancing and grid-search hyperparameter optimisation in R 4.4.1. Model performance on the 30% test set was assessed by AUC-ROC, F1 score, accuracy, sensitivity, and specificity. External validation used MZRH data; calibration was evaluated via Brier score and Hosmer–Lemeshow test; feature importance was derived using SHAP values. Results Among 4,523 clinical isolates, 3,334 (73.7%) demonstrated 3GC resistance — underscoring the urgency of predictive tools in this setting. The XGboost model achieved the strongest overall performance: AUC-ROC 0.83 (95% CI: 0.85–0.91), with excellent calibration (Hosmer–Lemeshow p = 0.087; Brier score 0.186). The most influential predictors were organism type (SHAP 0.32) and ward of admission (SHAP 0.15), both clinically actionable at the point of empirical prescribing. In simulation, integrating the XBboost model into prescribing workflows reduced inappropriate 3GC use by 25% compared with standard empirical guidelines. Conclusion This study presents the first externally validated machine learning model for predicting 3GC resistance using nationally representative AMR surveillance data from Tanzania. Embedding the GBDT model within antimicrobial stewardship tools offers a scalable, AI-driven pathway to reduce broad-spectrum antibiotic use, inform national prescribing guidelines, and advance precision health equity across Africa. These findings carry direct relevance for LMICs seeking digital solutions to support UHC goals amid the global AMR crisis.
COTAN: scRNA-seq comprehensive workflow based on gene correlations
Silvia Giulia Galfrè* Marco Fantozzi Alina Sîrbu Irene Testa Matteo Tolloso Andrea Alberti Corrado Priami Francesco Moradin
COTAN: scRNA-seq comprehensive workflow based on gene correlations
Silvia Giulia Galfrè* Marco Fantozzi Alina Sîrbu Irene Testa Matteo Tolloso Andrea Alberti Corrado Priami Francesco Moradin
- Keywords
- scRNAseqgene-gene correlationlowly expressed genes
- Abstract
- Objective: scRNA-seq enables profiling of thousands of transcriptomes per experiment, but downstream analysis is complicated by extreme sparsity and strong cell-to-cell variability in detection efficiency. Standard workflows adapted from bulk RNA-seq typically apply normalization, log-transformation, and feature filtering (e.g., widely used toolkits such as Seurat and Scanpy). While effective in many settings, these steps can blur the information present in the data and systematically remove low-abundance genes, including transcription factors or other regulatory genes that are informative for cell identity and state. Robust analysis therefore benefits from methods that natively exploit the dominant zero/non-zero structure of UMI count matrices and avoid imputation-induced artifacts. Methods: We present an improved release of COTAN (CO-expression Tables ANalysis), an end-to-end workflow grounded in contingency-table statistics that models zero counts directly. COTAN estimates a cell-specific library-size parameter together with gene-specific parameters calibrated to observed zero frequencies. Using this cell-aware null model, COTAN computes a robust gene-gene co-expression matrix (coex) designed to reduce spurious associations driven by heterogeneous detection efficiency. We compared, in real data, the correlations detected by Seurat, Scanpy and Monocle, Scenic and CS-CORE. The theoretical model of coex can also be used to assess the differential expression of genes when partitioning cells into two subsets. We compared COTAN’s DEA with analogous marker-detecting functions available in Seurat, Scanpy, Monocle and Memento both for false-positive rate and for statistical power. We further introduce highly sensitive gene scoring via the Global Differentiation Index (GDI), derived from the lower tail of gene-gene association p-values, and the use of the upper tail of the GDI distribution to assess whether a cell group is transcriptomically homogeneous (Uniform Transcript, UT, clusters). Results: We validated COTAN workflow across a diverse benchmark of public scRNA-seq datasets spanning tissues, protocols, and sequencing efficiencies. On datasets with the presence of UMI, coex produced gene-gene association signals consistent with expected relationships among marker genes while strongly reducing spurious correlations. In marker detection benchmarks, DEA showed strong false-positive control in challenging within-cluster splits driven by library size and high statistical power in heterogeneous mixtures, with particularly strong sensitivity for low-expression genes. The GDI/UT framework detected subtle heterogeneity: increasing fractions of a second transcriptional population induced a clear rise in the upper tail of GDI, and UT status predicted whether the cell subset can be considered as formed just by a single cell population. Finally, we demonstrated clustering utility on a dataset of seven human lung cancer cell lines with known identities, where a coex-based data reduction supported recovery of fine transcriptional structure. Conclusions: By modeling zeros directly and correcting for cell-specific detection efficiency without log-transforms or imputation, the improved COTAN workflow provides a robust, biologically grounded approach for sparse scRNA-seq. It improves reliability of gene-gene correlation estimates, enhances detection of cluster markers (including low-abundance regulators), and offers an objective criterion for cluster homogeneity, supporting more reliable identification of cell types and states.
A Network Story: tumor educated platelets transcriptome and Glioblastoma
Stefano, Rinaldi, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza Alessandro, Taraborelli, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza *Mattia, Manna, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza Lorenzo, Farina, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza Manuela, Petti, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza
A Network Story: tumor educated platelets transcriptome and Glioblastoma
Stefano, Rinaldi, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza Alessandro, Taraborelli, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza *Mattia, Manna, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza Lorenzo, Farina, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza Manuela, Petti, Dipartimento di Ingegneria Informatica, Automatica e Gestionale Sapienza
- Keywords
- Tumor educated plateletsdisease moduleRNA-seqprotein-protein interactionnetwork analysis
- Abstract
- Motivation: Tumor-educated platelets (TEPs) are emerging as a promising liquid biopsy source for cancer detection and monitoring. However, it remains unclear whether their transcriptomic alterations reflect tumor-specific molecular mechanisms or broader systemic responses. In particular, their relationship with the molecular network underlying glioblastoma (GBM) has not been systematically investigated. Here, we test whether differentially expressed genes (DEGs) in TEPs from GBM patients are functionally connected to the GBM disease module in the human interactome. Objective: Assess whether transcriptomic alterations in TEPs from GBM patients are functionally and topologically associated with the GBM disease module in the human interactome. Methods: RNA-seq data from TEPs were obtained from GEO (GSE68086). Differential expression analysis comparing GBM samples to healthy controls (adjusted p-value < 0.01 and |log2FC| ≥ 0.5) identified 1401 DEGs. After Ensembl-to-symbol mapping and removal of duplicates, DEGs were mapped to a human protein–protein interaction (PPI) network derived from BioGRID, restricted to physical interactions. A curated set of 48 GBM seed genes from IntOGen defined the disease module. After removing 5 overlapping genes, the final set of 1298 non-overlapping DEGs were used for analysis. Network proximity was computed as the mean shortest-path from DEGs to disease genes. Statistical significance was assessed using 500 random gene sets of equal size sampled from the interactome after excluding both GBM disease genes and the observed DEGs. Random walk with restart (restart = 0.7) was performed on the full network using disease genes as seeds to prioritize DEGs by network connectivity. The top 50 ranked genes were analyzed using GO and Reactome enrichment. Results: DEGs showed strong topological proximity to GBM driver genes, with an observed mean distance of 1.41 compared to a null mean of 1.66 (SD = 0.013), corresponding to a Z-score of −18.69. The empirical p-value was < 0.002, indicating significantly closer proximity than expected by chance. Network propagation identified a subset of non-overlapping DEGs highly connected to the GBM module. Functional enrichment highlighted telomere maintenance and organization, DNA repair and metabolic processes, chromatin and chromosome organization, transcriptional and RNA regulation, cell cycle, protein folding, and cellular stress responses. Reactome analysis further supported enrichment in DNA repair, chaperone-mediated pathways, post-translational modification, transcriptional regulation, and innate immune/antiviral processes. Conclusion: TEP transcriptomic alterations in GBM are non-randomly distributed in the human interactome but localize in close proximity to the GBM disease module.
COMPUTATIONAL PREDICTION OF PATHOGENIC FMR1 VARIANTS IN FRAGILE X SYNDROME
Tamilinian, Vedha *
COMPUTATIONAL PREDICTION OF PATHOGENIC FMR1 VARIANTS IN FRAGILE X SYNDROME
Tamilinian, Vedha *
- Keywords
- FMR1FXSMLFMRPMissense mutation
- Abstract
- Fragile X Syndrome (FXS) is a genetic disorder that inhibits production of the FMRP protein, essential in synaptic plasticity and neural development. FXS is caused by more than 200 CGG repeats in the FMR1 gene. Recent findings have shown that FXS can also result from missense mutations in KH1 and KH2 domains of the FMR1 gene. However, there is limited research on identifying which variants are truly pathogenic, resulting in misdiagnoses. This project integrates factors including protein embeddings, evolutionary data, and structural features in developing a predictive model to improve classification of pathogenic FMR1 missense mutations in the KH1 and KH2 domains. Compiling these diverse features alongside a machine learning model, offers a comprehensive and scalable solution in classifying mutations that had previously uncertain significance. Data compilation begins with first collecting existing variants from ClinVar and generating every possible variant using the wild-type of FMRP. Then, values for each of the features were generated using protein embeddings from the ESM2 protein language model, structural predictions from NetSurfP, and evolutionary data from AlphaServer. The processed dataset was then used to train four supervised ML models: support vector machine, XGBoost, random forest, and logistic regression.To ensure that the model was truly learning rather than simply memorizing, 5-fold cross validation is also applied in each model’s development. Of these, XGBoost was the most accurate, with the accuracy was 0.94, pathogenic precision was 0.92, pathogenic recall was 0.95, and the F1-score was 0.94. A primary display of the success of the model is the mutation G266E in KH1, which was proven pathogenic in wet lab. With 171 labeled variants, the model classified 2223, of which 697 were pathogenic. Certain original amino acids, such as serine (S), proline (P), and asparagine (N), showed high risk when replaced. Other specific mutant amino acid replacements, including glutamic acid (E), alanine (A), and valine (V), were also associated with increased pathogenicity. The model also identified hotspots, or sequence positions where missense mutations are highly likely to be pathogenic: 6 in the KH1 domain and 8 in KH2. This model is the foundation for improved diagnosis for individuals who show the FXS phenotype but do not have the full mutation of CGG repeats. The model’s predicted probabilities of pathogenicity and its biologically interpretable insights provide clinicians and researchers with a benchmark for not only identifying FMR1 mutations in patients but also understanding their significance.
Modified Gompertz Function in Quantitative Evaluation of Endolysin Lytic Activity in Physiological Fluids.
Marek Adam, Harhala, Institute for Biology and Human Evolution, Faculty of Biology, Adam Mickiewicz University
Modified Gompertz Function in Quantitative Evaluation of Endolysin Lytic Activity in Physiological Fluids.
Marek Adam, Harhala, Institute for Biology and Human Evolution, Faculty of Biology, Adam Mickiewicz University
- Keywords
- endolysinantibacterialenzyme kineticsmodel fitting
- Abstract
- Modified Gompertz Function in Quantitative Evaluation of Endolysin Lytic Activity in Physiological Fluids. Objective: An accurate quantification of bacteriolytic activity is essential for evaluation and comparison of new antimicrobial therapies in the context of increasing antibiotic resistance. Endolysins, a bacteriophage-derived, fast, bacteriolytic enzymes are promising candidates for the treatment of antibiotic-resistant infections. However, evaluation of their lytic activity remains methodologically challenging. Presented study aimed to modify and apply a modified Gompertz- based statistical model for quantitative assessment of the lysis of metabolically active bacterial cells under biologically relevant conditions and concentrations. Methods: Bacteriolytic activity was measured using fluorescence-based assay previously researched. Disruption of live bacterial cells by bacteriolytic agent (endolysin) results in the release from now metabolically inactive cells their DNA to the reaction environment. Such DNA is bound by fluorescent dye SYTOX Green, present in a reaction environment, increasing its fluorescence signal. These measurements are fitted using the modified Gompertz function with model parameters interpreted biologically. The baseline fluorescence (lower asymptote), total lysis signal, lytic activity and the reaction delay are represented by parameters y, a, c and b, respectively. Results: The modified Gompertz function provided a significantly closer fit to the experimental data than traditionally used approaches. Modified parameters in the function improve fit of the function, especially when the reaction progress reaches over 50%. Presented modified Gompertz function enabled estimation of lytic activity from all experimental data points simultaneously accounting for baseline and maximum signals, and the delay often, but not always, observed in endolysins. The approach was tested and remained applicable across Cpl-1 and Pal endolysins and their variants. Approach was also tested in matrices such as serum, also in wide range of enzymatic and bacteria concentrations. Conclusion: The presented modified Gompertz function proved efficient, reproducible and physiologically compatible framework for the evaluation of whole, metabolically active Streptococcus pneumoniae cells. Also, the presented approach assigns to model parameters a biological component, e.g. proposed model requires cleavage of a critical number of chemical bonds in the bacterial cell-wall to produce a rupture sufficient for metabolic inactivation. The delay resulting from this aspect is represented by the model parameter b. Presented model improves the statistical descriptions of bacteriolytic assay between enzymes and in various concentrations and conditions. This modelling approach may improve the optimization of endolysin candidates, through modification or selection, particularly for medical use, especially under complex biological constraints.Methods: I present a mathematical modelling framework for quantitative assessment of bacteriolytic activity based on real-time fluorescence measurements and a modified Gompertz function. The method relies on fitting the modified Gompertz function to fluorescence emission data that correlate with the amount of metabolically inactive bacteria in the sample. The modified Gompertz function provides a significantly closer fit to experimental data compared to traditionally used approaches. In this model, parameter c represents the bacteriolytic activity rate and is defined as the quantitative lytic activity of the endolysin. Parameter b is associated with the delay in signal increase, commonly but not always, observed during endolysin mediated lysis. Parameters y and a describe the baseline signal and total fluorescence (i.e., the difference between baseline and complete bacterial lysis), respectively. Results: Presented approach proved effective in evaluating lysis of whole, metabolically active Streptococcus pneumoniae cells by the endolysins Cpl-1 and Pal. The method was also successfully applied to compare wild-type endolysins with engineered variants. The close fit to experimental data demonstrates the reproducibility and precision of the proposed modifications. Conclusions: The presented modified Gompertz function captures the kinetics inherent to lysis of intact bacterial cells by endolysins and aligns with the mechanistic model of endolysin action, which requires cleavage of a critical number of chemical bonds in the bacterial cell-wall to produce a rupture sufficient for metabolic inactivation. The model is robust under various physiological conditions, including the presence of plasma and varying bacterial concentrations. Presented model provides a sound, reproducible, and physiologically compatible method for evaluating the antimicrobial potential of bacteriolytic proteins. The approach is broadly applicable to the development of bacteriolytic agents, particularly for medical use, and offers quantitative support for protein optimization under complex biological constraints.
Exploring the role of marital status in Depression-Free Life Expectancy among older adults
Kokalla Erlini (1,2)*, Feraldi Alessandro (3), Cristina Giudici (3) 1 Department of Statistical Sciences, Sapienza University, Rome, Italy ² Faculty of Health, University of Vlore “Ismail Qemali”, Vlore, Albania 3 Department of Methods and Models for Economics, Territory and Finance, Sapienza University, Rome, Italy
Exploring the role of marital status in Depression-Free Life Expectancy among older adults
Kokalla Erlini (1,2)*, Feraldi Alessandro (3), Cristina Giudici (3) 1 Department of Statistical Sciences, Sapienza University, Rome, Italy ² Faculty of Health, University of Vlore “Ismail Qemali”, Vlore, Albania 3 Department of Methods and Models for Economics, Territory and Finance, Sapienza University, Rome, Italy
- Keywords
- Dep-FLEolder adultsmarital statusgenderrace
- Abstract
- Abstract Depression is one of the most common mental health conditions among older adults, affecting the overall wellbeing of this group. The prevalence of this condition is highly shaped by marital status and socio-economic inequalities, which may also influence total life expectancy among older adults. Objective: This study aims to examine Depression-Free Life Expectancy (Dep-FLE) by marital status (never married, married, divorced, separated, widowed), among older adults in the United States using multistate life table models and to assess how the association between marital status and Dep-FLE varies by gender and by race/ethnicity. Methods: This is a secondary data analysis using data from the RAND version of the U.S. Health and Retirement Study, waves 6-16. The analytical sample included 28 978 individuals aged 50 years and older, contributing 149 770 person-wave observations. Dep-FLE was estimated by using the multistate life table models based on transition probabilities between three states “Without depression”, “Depression” and “Death” as an absorbing state, derived from multinomial logistic regression. Results: Among participants, 58% (16 671) were female and 42% (12 307) male. In both genders, the majority were White (66.3%), married (61%), and with a medium level of education (53.1%). The overall prevalence of depression was 24.2%, with higher rates among women. Married individuals consistently showed the highest Dep-FLE in both genders (women: 27.9 years, men: 26.4 years). Across racial groups, married individuals showed the highest Dep-FLE, reaching approximately 27.3 years among White men and 28.9 years among White women, while separated/divorced and never married individuals had the lowest estimates (approximately 15-22 years). White older adults experienced the longest depression-free survival, whereas Black older adults showed the lowest values across most marital status categories. Although women had higher total life expectancy across all marital status categories, men spent greater proportion of their remaining years free from depression. Conclusions: Marital status strongly affects Depression-Free Life Expectancy among older adults. Married older adults experience longer and healthier lives, with more years lived free from depression compared with separated/divorced and never married individuals. Although women live longer than men, they experience a lower proportion of life free from depression, potentially reflecting greater psychological burden and gender- related health inequalities. Important racial disparities were also observed, with Black older adults showing shorter Dep-FLE across most marital status groups.
TensorPLS: an R package for tensor-aware PLS-based discriminant analysis of longitudinal multi-omics data
*Alessandro, Giordano, Telethon Institute of Genetics and Medicine (TIGEM), Pozzuoli, Italy Ana, Conesa, Institute for Integrative Systems Biology (I2SysBio), Parc Científic de la Universitat de València, Paterna, Valencia, Spain Leandro, Balzano-Nogueira, Institute for Integrative Systems Biology (I2SysBio), Parc Científic de la Universitat de València, Paterna, Valencia, Spain
TensorPLS: an R package for tensor-aware PLS-based discriminant analysis of longitudinal multi-omics data
*Alessandro, Giordano, Telethon Institute of Genetics and Medicine (TIGEM), Pozzuoli, Italy Ana, Conesa, Institute for Integrative Systems Biology (I2SysBio), Parc Científic de la Universitat de València, Paterna, Valencia, Spain Leandro, Balzano-Nogueira, Institute for Integrative Systems Biology (I2SysBio), Parc Científic de la Universitat de València, Paterna, Valencia, Spain
- Keywords
- Tensor dataPLS-DAMulti-omicsLongitudinal analysisFeature selection
- Abstract
- Objective Longitudinal and multi-omics studies increasingly generate data with a natural three-way structure, such as subjects × features × time. Although PLS-based discriminant models can be fitted on unfolded representations, biological interpretation requires recovering the original tensor organisation. We developed TensorPLS, an R package providing an end-to-end, tensor-aware workflow for three-dimensional omics data. The package integrates tensor construction, Tucker-3 imputation, PLS-based discriminant modelling, cross-validated performance assessment, VIP-driven feature prioritisation, and time-aware interpretation. Its objective is to support group discrimination, feature selection, model evaluation, and identification of the time points or data blocks that most strongly drive class separation. Methods TensorPLS converts input omics tables into aligned three-way tensors and handles missing values through iterative Tucker-3 imputation fitted directly on the tensor structure. For discriminant modelling, predictors are unfolded and analysed using an internal PLS core; model outputs are then mapped back to the original feature–time structure to enable tensor-aware interpretation. The package reports cross-validated Q² and explained variance R² to support component tuning and comparison of modelling configurations. A central feature of TensorPLS is the generation of complementary VIP representations. These identify feature–time pairs important for specific components, variables consistently relevant across time, and variables influential within selected temporal windows. The same VIP views form the basis for downstream feature selection, performed through percentile-based thresholding. Time-resolved VIP rankings can also be assessed by permutation-based robustness testing to distinguish stable feature–time signals from rankings arising under random label assignments. Results TensorPLS was evaluated on four longitudinal datasets derived from the TEDDY study, represented as subjects × variables × time tensors across five aligned time points relative to islet autoimmunity confirmation. The analysed modalities included whole-blood gene expression, plasma metabolomics, positive dietary biomarkers, and negative dietary biomarkers. Baseline models showed modality-specific differences in case–control discrimination, with the strongest signal observed for gene expression (R² = 0.91, Q² = 0.77), followed by GCTOFX metabolomics (R² = 0.84, Q² = 0.50), whereas dietary biomarker datasets showed weaker cross-validated performance. VIP-based feature selection substantially reduced dimensionality, for example from 21,285 to 816 variables in the gene expression dataset, and improved internal Q² values across all datasets. Mode-3 analyses identified modality-specific temporal windows contributing to discrimination, while VIP2D plots highlighted time-specific discriminant features and their robustness under permutation testing. Conclusion TensorPLS provides a flexible and reproducible R framework for analysing longitudinal multi-omics tensors. By combining tensor-native preprocessing, PLS-based modelling, VIP-driven feature selection, and temporal contribution analysis, the package enables users to move from high-dimensional three-way data to interpretable discriminant features and temporal signatures. These outputs support downstream biological interpretation, pathway analysis, and validation in independent cohorts.
Development of an AI-Based Assistive Technology for Neurodivergent Children: Enhancing Learning and Social Engagement in African Language Contexts
Abimbola R. Akinyemi
Development of an AI-Based Assistive Technology for Neurodivergent Children: Enhancing Learning and Social Engagement in African Language Contexts
Abimbola R. Akinyemi
- Keywords
- AI assistive technologyAutism Spectrum DisorderAfrican languagesneurodivergent childrenparticipatory design
- Abstract
- Development of an AI-Based Assistive Technology for Neurodivergent Children: Enhancing Learning and Social Engagement in African Language Contexts Objective: Children with Autism Spectrum Disorder (ASD) in sub-Saharan Africa face significant barriers to accessing evidence-based learning interventions. Existing tools are predominantly designed for English or other European languages, limiting accessibility for children who communicate primarily in African languages. In Rwanda, where Kinyarwanda is widely spoken, there is a critical absence of culturally and linguistically appropriate AI-supported learning technologies. Methods: This pilot study addresses this gap through the development of a voice-enabled AI assistive mobile chatbot tailored to African language contexts. A participatory design methodology was employed, engaging twenty children with ASD (aged 7–13), alongside caregivers and educators, as co-designers. This approach ensured contextual relevance, ethical sensitivity, and user-centered innovation. The system integrates African-language natural language processing resources, including datasets from community-driven initiatives such as Masakhane. Performance metrics, including word error rate and intent recognition accuracy, were evaluated under real-world conditions to ensure transparency in low-resource AI deployment. The pilot was conducted across home and school settings in Kigali, Rwanda. Each child participated in eight structured interaction sessions, enabling longitudinal observation of engagement patterns. A within-participant observational design was adopted, prioritizing feasibility and ethical considerations over controlled experimentation. Results: Findings indicate improvements in verbal initiation, turn-taking behavior, and sustained social interaction. Caregivers and educators reported high satisfaction with the system’s linguistic accessibility and cultural relevance. Conclusion: This study contributes to both AI and special education research by demonstrating the feasibility of localized, participatory AI solutions in low-resource settings. It also provides a replicable framework for developing inclusive assistive technologies for neurodivergent populations. The findings have implications for policy, educational practice, and the ethical deployment of AI in underserved communities
OMICS-query: An AI Agent-Based Chatbot for Bioinformatics Pipelines
Camilla Callierotti*, National Facility for Data Handling and Analysis, Fondazione Human Technopole, Milan, Italy Alberto Riva, National Facility for Data Handling and Analysis, Fondazione Human Technopole, Milan, Italy
OMICS-query: An AI Agent-Based Chatbot for Bioinformatics Pipelines
Camilla Callierotti*, National Facility for Data Handling and Analysis, Fondazione Human Technopole, Milan, Italy Alberto Riva, National Facility for Data Handling and Analysis, Fondazione Human Technopole, Milan, Italy
- Keywords
- Agentic AIBioinformatics pipelinesOmics data analysisLLMs
- Abstract
- Objective High-throughput bioinformatics workflows generate complex outputs whose downstream interpretation remains a major bottleneck. Researchers must navigate scattered files and write custom code to access their own data, follow-up questions require analyst intervention, and static reports cannot address iterative exploration. OMICS-query is an agent-based system for the interactive interrogation of multi-omics datasets through natural language, thanks to autonomous orchestration of multiple data analysis tools. Methods OMICS-query’s architecture combines a relational database with large language models (LLMs). Pipeline outputs (e.g., from nf-core/rnaseq) are ingested via an automated ETL workflow into a normalized relational schema. Full-featured analysis pipelines like the ones in nf-core produce datasets routinely exceeding millions of rows, far beyond any LLM's context window, and require complex operations that only a relational database can reliably execute. Following the ETL pipeline, the system leverages an agentic orchestration framework powered by a foundation model, where task-specific prompt templates guide reasoning across distinct modes (e.g., statistical inference, functional interpretation, visualization generation). Statistical computations are executed at the database level, while the LLM is restricted to interpretation, summarization, and visualization, enforcing a separation of concerns between deterministic computation and generative reasoning. The system supports iterative, context-aware querying via conversational state management across multi-step analytical workflows. Benchmarking was conducted on the nf-core/rnaseq pipeline, and evaluation metrics were defined at the task level, where a query was considered correct only if both the computational result and its biological interpretation were accurate. Results OMICS-query was benchmarked across multiple query categories, including metadata retrieval, quality control (QC), differential expression analysis, pathway enrichment, and visualization tasks. Evaluation was conducted on a curated query set (n=141), measuring both predictive performance and agentic execution reliability. The system achieved an overall accuracy of 93.6%, with precision 93.7%, recall 94.9%, and F1-score 94.3%, with near-perfect scores observed in differential expression queries, and strong results in pathway enrichment (90.2% accuracy) and visualization tasks (94.6% accuracy). It achieved a task completion rate of 93.6% and a tool-call success rate of 89.6%, indicating robust orchestration of external tools and database operations. Notably, the system demonstrated effective fault tolerance, with an error recovery rate of 88.9%, successfully handling intermediate failures during complex query execution. These results highlight the system’s ability to integrate structured data retrieval with LLM-based reasoning, maintaining high accuracy while supporting diverse query types ranging from statistical analysis to higher-level biological interpretation. The system also generates publication-ready visualizations and reports. Conclusion We show how agentic AI frameworks can democratize access to complex omics datasets, enabling researchers to independently explore their results, while reducing analyst support burden. The OMICS-query architecture is readily extensible to other nf-core pipelines and omics data types, offering a scalable solution for modern genomics and bioinformatics facilities.
Computational Framework for the Design of Second-Generation Pharmacochaperones Targeting Allosteric Pockets
Alessandra, Robello*, Department of Biochemical Sciences "A. Rossi Fanelli", Sapienza University of Rome, Rome, Italy. Francesco, Pirozzi, Department of Medicinal Chemistry and Technology, Sapienza University of Rome, Rome, Italy. Roberta, Astolfi, Department of Medicinal Chemistry and Technology, Sapienza University of Rome, Rome, Italy. Lidia, Giuliani, Department of Medicinal Chemistry and Technology, Sapienza University of Rome, Rome, Italy . Rino, Ragno, Department of Medicinal Chemistry and Technology, Sapienza University of Rome, Rome, Italy. Allegra, Via, Department of Biochemical Sciences "A. Rossi Fanelli", Sapienza University of Rome, Rome, Italy.
Computational Framework for the Design of Second-Generation Pharmacochaperones Targeting Allosteric Pockets
Alessandra, Robello*, Department of Biochemical Sciences "A. Rossi Fanelli", Sapienza University of Rome, Rome, Italy. Francesco, Pirozzi, Department of Medicinal Chemistry and Technology, Sapienza University of Rome, Rome, Italy. Roberta, Astolfi, Department of Medicinal Chemistry and Technology, Sapienza University of Rome, Rome, Italy. Lidia, Giuliani, Department of Medicinal Chemistry and Technology, Sapienza University of Rome, Rome, Italy . Rino, Ragno, Department of Medicinal Chemistry and Technology, Sapienza University of Rome, Rome, Italy. Allegra, Via, Department of Biochemical Sciences "A. Rossi Fanelli", Sapienza University of Rome, Rome, Italy.
- Keywords
- PharmacochaperonesAllosteric modulationGenerative molecular design
- Abstract
- Objective Protein misfolding caused by missense mutations is the underlying cause of rare diseases such as phenylketonuria, alkaptonuria, hereditary transthyretin amyloidosis, and lysosomal storage disorders. First-generation pharmacochaperones stabilized misfolded proteins by binding to the active site, but their therapeutic utility was limited by the difficulty of balancing enzymatic activity enhancement and inhibition. This study aims to investigate the binding mode of second-generation pharmacochaperones at allosteric pockets, characterizing the molecular determinants that govern their interaction with the target protein. A further objective is to develop novel drug candidates acting as pharmacochaperones, by simultaneously considering both ligand properties and protein structural features, to design molecules capable of selectively stabilizing misfolded proteins without interfering with the active site. Methods The proposed pipeline begins with a careful data preparation phase, in which a curated dataset of monomeric proteins with experimentally validated allosteric ligands is assembled and structurally refined to ensure physical consistency. Molecular docking is then performed to characterize the allosteric binding pocket geometry and to identify the preferred binding poses of known and candidate ligands, providing a structural basis for selectivity assessment. Subsequently, Molecular Interaction Fields (MIFs) will be calculated to create condition generative models for the de novo design of novel molecules complementary to the allosteric pocket and to quantitatively estimate ligand-pocket affinity, establishing an energy-based selectivity profile for each target. The pipeline will culminate in molecular dynamics (MD) simulations, aimed at dynamically validating the binding and exploring allosteric communication pathways between the ligand binding site and the active site. Results The expected outcomes of this study include a mechanistic understanding of how second-generation pharmacochaperones engage allosteric pockets at the atomic level; a set of quantitative descriptors capturing ligand-protein complementarity at allosteric versus orthosteric sites; a collection of novel computationally designed molecules with predicted pharmacochaperone activity; MIF-derived affinity estimates enabling the ranking and prioritization of candidate ligands; and dynamic insights into how allosteric binding modulates conformational stability and signal transmission to the active site. These results, currently in progress, are intended to provide a rational and generalizable basis for the discovery of effective therapeutic agents for rare conformational diseases. Conclusion The proposed framework integrates structural data preparation, docking-based pocket characterization, MIF-conditioned generative molecular design, and molecular dynamics validation into a modular and extensible pipeline. By jointly accounting for ligand and protein features, this approach addresses the core challenge of rationally designing pharmacochaperones that act selectively on allosteric sites, with potential applicability to a broad spectrum of rare diseases for which no effective treatments currently exist.
Integrative analysis of multi-modal single-cell data reveals chromosome alterations underlying critical disease development stages: a proof-of-principle in MDS
*Alice, Chiodi, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Department of Biomedical Sciences, Humanitas University, Pieve Emanuele, Milan, Italy Matteo, Zampini, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Department of Biomedical Sciences, Humanitas University, Pieve Emanuele, Milan, Italy Elena, Riva, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Nicolas, Sompairac, Comprehensive Cancer Center, King's College London, London, United Kingdom Giulia, Maggioni, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Rita, Antunes Dos Reis, Comprehensive Cancer Center, King's College London, London, United Kingdom Rosa, Andres Ejarque, Comprehensive Cancer Center, King's College London, London, United Kingdom Laura, Crisafulli, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Institute for Genetic and Biomedical Research, Milan Unit, CNR, Milan, Italy Nicolas, Derus, Department of Medical and Surgical Sciences (DIMEC), University of Bologna Martina, Tarozzi, Department of Medical and Surgical Sciences (DIMEC), University of Bologna Federico, Magnani, Department of Medical and Surgical Sciences (DIMEC), University of Bologna Claudia, Sala, Department of Medical and Surgical Sciences (DIMEC), University of Bologna; IRCCS Azienda Ospedaliero-Universitaria di Bologna S.Orsola, 40138, Bologna, Italy Elisabetta, Sauta, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Alessia, Campagna, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Gabriele, Todisco, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Department of Biomedical Sciences, Humanitas University, Pieve Emanuele, Milan, Italy Marta, Ubezio, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Antonio, Russo, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Gianluca, Asti, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Denise, Ventura, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Ivan, Ferrari, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Matteo, Brindisi, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Nicla, Manes, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Chiara, Milanesi, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Saverio, D'Amico, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Train SRL, Milan, Italy Gastone, Castellani, Department of Medical and Surgical Sciences (DIMEC), University of Bologna; IRCCS Azienda Ospedaliero-Universitaria di Bologna S.Orsola, 40138, Bologna, Italy Shahram, Kordasti, Comprehensive Cancer Center, King's College London, London, United Kingdom; DISCLIMO - Università Politecnica delle Marche, Hematology Department & Stem Cell Transplant Unit, Ancona, Italy Ettore, Mosca, Institute of Biomedical technologies, CNR, Milan, Italy Francesca, Ficara, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Institute for Genetic and Biomedical Research, Milan Unit, CNR, Milan, Italy Matteo, Della Porta, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Department of Biomedical Sciences, Humanitas University, Pieve Emanuele, Milan, Italy
Integrative analysis of multi-modal single-cell data reveals chromosome alterations underlying critical disease development stages: a proof-of-principle in MDS
*Alice, Chiodi, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Department of Biomedical Sciences, Humanitas University, Pieve Emanuele, Milan, Italy Matteo, Zampini, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Department of Biomedical Sciences, Humanitas University, Pieve Emanuele, Milan, Italy Elena, Riva, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Nicolas, Sompairac, Comprehensive Cancer Center, King's College London, London, United Kingdom Giulia, Maggioni, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Rita, Antunes Dos Reis, Comprehensive Cancer Center, King's College London, London, United Kingdom Rosa, Andres Ejarque, Comprehensive Cancer Center, King's College London, London, United Kingdom Laura, Crisafulli, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Institute for Genetic and Biomedical Research, Milan Unit, CNR, Milan, Italy Nicolas, Derus, Department of Medical and Surgical Sciences (DIMEC), University of Bologna Martina, Tarozzi, Department of Medical and Surgical Sciences (DIMEC), University of Bologna Federico, Magnani, Department of Medical and Surgical Sciences (DIMEC), University of Bologna Claudia, Sala, Department of Medical and Surgical Sciences (DIMEC), University of Bologna; IRCCS Azienda Ospedaliero-Universitaria di Bologna S.Orsola, 40138, Bologna, Italy Elisabetta, Sauta, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Alessia, Campagna, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Gabriele, Todisco, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Department of Biomedical Sciences, Humanitas University, Pieve Emanuele, Milan, Italy Marta, Ubezio, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Antonio, Russo, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Gianluca, Asti, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Denise, Ventura, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Ivan, Ferrari, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Matteo, Brindisi, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Nicla, Manes, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Chiara, Milanesi, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy Saverio, D'Amico, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Train SRL, Milan, Italy Gastone, Castellani, Department of Medical and Surgical Sciences (DIMEC), University of Bologna; IRCCS Azienda Ospedaliero-Universitaria di Bologna S.Orsola, 40138, Bologna, Italy Shahram, Kordasti, Comprehensive Cancer Center, King's College London, London, United Kingdom; DISCLIMO - Università Politecnica delle Marche, Hematology Department & Stem Cell Transplant Unit, Ancona, Italy Ettore, Mosca, Institute of Biomedical technologies, CNR, Milan, Italy Francesca, Ficara, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Institute for Genetic and Biomedical Research, Milan Unit, CNR, Milan, Italy Matteo, Della Porta, IRCCS Humanitas Research Hospital, Rozzano, Milano, Italy; Department of Biomedical Sciences, Humanitas University, Pieve Emanuele, Milan, Italy
- Keywords
- Single-celltranscriptomicsmulti-omics
- Abstract
- Objective: Myelodysplastic syndromes (MDS) are heterogeneous hematopoietic disorders characterized by clonal, transcriptional and microenvironmental changes over time, with an increased risk of progression to Acute Myeloid Leukaemia (AML). The disorder is heterogeneous in terms of mutations involved, time of evolution and response to treatment. We used a single cell multi-omics approach to investigate this heterogeneity. Methods: Two cohorts considered: evolution, with samples at diagnosis and AML evolution (20 patients); treatment, with samples at diagnosis, treatment, remission and relapse (13 patients). All samples: bone marrow, sequenced using CITE-Seq. Data analysis: CellRanger for preprocessing and alignment (GRCh38); Seurat v5 for quality filtering, normalization, dimensionality reduction (PCA), integration (Harmony, top 2’000 highly variable features used), clustering and differential expression analysis (both gene and protein expression, Wilcoxon rank-sum testing); celltype annotation with Azimuth and manual curation (marker genes and proteins, pathway enrichment). Longitudinal analysis: copy number variation (CNV), trajectory analysis, quantification of fold-enrichment of transcriptionally similar cells between timepoints through k-nearest neighbour analysis. Genomics data: mutational status on target genes through Genotyping of Transcriptome (GoT). Results: Concerning both cohorts, the analyses allowed the identification of the major cellular compartments involved in the disease, including hematopoietic stem and progenitor cells (HSPCs), myeloid populations, and components of the tumor microenvironment (T and NK cells). CNV profiles mirror the karyotype, supporting the robustness of the results and suggesting clone dynamics. Evolution cohort: comparisons between timepoints, through cell proportion analysis and transcriptional profiling, highlighted the presence of distinct evolutionary trajectories associated with changes in HSPC composition and differentiation states, with differences already detectable at diagnosis and able to predict the pathological fate. Longitudinal analysis of immune populations revealed substantial remodelling of the microenvironment over time, with evidence of altered cytotoxic programs, inflammatory activation, and acquisition of immunosuppressive features across different T and NK cell subsets. Treatment cohort: comparative analyses between pre- and post-treatment samples revealed marked inter-patient heterogeneity in response to therapy, with dynamic shifts in cellular composition and transcriptional programs associated with different clinical outcomes. Response was associated with enrichment of erythroid-associated populations, whereas relapse samples often showed persistence or re-expansion of cellular states already present at diagnosis, together with the emergence of additional transcriptional states in specific settings. Conclusions: The results obtained so far provide new insights in both tumor cells and the microenvironment of this heterogeneous disease, with findings from the evolution cohort highlighting the potential of the implemented approach. The next step will be the development of a method based on genomic and RNA trajectory analyses to model differences in disease evolution and treatment response.
Development and Internal Validation of a Multimodal Machine Learning Model Integrating Omics, Clinical and Psychosocial Data for Heart Failure Risk Prediction
* Andreea Eunice, Cosa, Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain Júlia, Perera Bel, Hospital del Mar Medical Research Institute (IMIM) (Bioinformatics Unit, MARGenomics & GRIB), Barcelona, Spain Sílvia, Jovells Vaqué, Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain Pau, Berenguer Molins, Hospital del Mar Medical Research Institute (IMIM) (Bioinformatics Unit, MARGenomics & GRIB), Barcelona, Spain Núria, José Bazán, Community Heart Failure Program, Cardiology Department, Bellvitge University Hospital, Hospitalet de Llobregat, Spain. Raúl, Ramos Polo, Community Heart Failure Program, Cardiology Department, Bellvitge University Hospital, L’Hospitalet de Llobregat, 08907 Barcelona, Spain. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain. Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Department of Clinical Sciences, School of Medicine, Universitat de Barcelona, 08007 Barcelona, Spain Josep, Francesch Manzano, Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain Marta, Tajes Orduña, Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain Àlex, Sánchez Pla, Centro de Investigación Biomédica en Red de Fragilidad y Envejecimiento Saludable (CIBERFES), Instituto de Salud Carlos III, Madrid 28029, Spain. Department of Genetics, Microbiology and Statistics, University of Barcelona, Barcelona 08028, Spain Josep, Comín-Colet, Community Heart Failure Program, Cardiology Department, Bellvitge University Hospital, L’Hospitalet de Llobregat, 08907 Barcelona, Spain. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain. Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Department of Clinical Sciences, School of Medicine, Universitat de Barcelona, 08007 Barcelona, Spain
Development and Internal Validation of a Multimodal Machine Learning Model Integrating Omics, Clinical and Psychosocial Data for Heart Failure Risk Prediction
* Andreea Eunice, Cosa, Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain Júlia, Perera Bel, Hospital del Mar Medical Research Institute (IMIM) (Bioinformatics Unit, MARGenomics & GRIB), Barcelona, Spain Sílvia, Jovells Vaqué, Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain Pau, Berenguer Molins, Hospital del Mar Medical Research Institute (IMIM) (Bioinformatics Unit, MARGenomics & GRIB), Barcelona, Spain Núria, José Bazán, Community Heart Failure Program, Cardiology Department, Bellvitge University Hospital, Hospitalet de Llobregat, Spain. Raúl, Ramos Polo, Community Heart Failure Program, Cardiology Department, Bellvitge University Hospital, L’Hospitalet de Llobregat, 08907 Barcelona, Spain. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain. Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Department of Clinical Sciences, School of Medicine, Universitat de Barcelona, 08007 Barcelona, Spain Josep, Francesch Manzano, Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain Marta, Tajes Orduña, Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain Àlex, Sánchez Pla, Centro de Investigación Biomédica en Red de Fragilidad y Envejecimiento Saludable (CIBERFES), Instituto de Salud Carlos III, Madrid 28029, Spain. Department of Genetics, Microbiology and Statistics, University of Barcelona, Barcelona 08028, Spain Josep, Comín-Colet, Community Heart Failure Program, Cardiology Department, Bellvitge University Hospital, L’Hospitalet de Llobregat, 08907 Barcelona, Spain. Bio-Heart Cardiovascular Diseases Research Group, Bellvitge Biomedical Research Institute (IDIBELL), L’Hospitalet de Llobregat, 08908 Barcelona, Spain. Centro de Investigación Biomédica en Red Enfermedades Cardiovasculares (CIBERCV), Instituto de Salud Carlos III, Madrid 28029. Department of Clinical Sciences, School of Medicine, Universitat de Barcelona, 08007 Barcelona, Spain
- Keywords
- Heart failureMachine learningBiomarkersRisk predictionMultimodal integration
- Abstract
- Objective The ORACLE multicentric study (ClinicalTrials.gov ID NCT05679713) aims to develop and validate a machine learning framework integrating omics, clinical and psychosocial data for heart failure (HF) risk prediction. Here, we report the development and internal validation of the model in the derivation cohort and assess the added value of multimodal integration over single-domain approaches. Methods We implemented a four-phase workflow including discovery, biomarker validation, model development, and external validation. Peripheral blood RNA-sequencing was performed in a discovery cohort of 60 patients (30 cases, 30 controls) to identify differentially expressed genes (DEGs). Candidate genes were further prioritized using penalized regression and feature selection approaches, complemented by literature support. Gene expression was subsequently evaluated by qPCR in an independent set of 119 patients to validate selected biomarkers. Machine learning models were trained on a dataset of 143 patients, combining genomic, clinical and psychosocial variables. Feature selection and model training were embedded within a 10-fold nested cross-validation framework to avoid information leakage. Model performance was evaluated using ROC-AUC and balanced accuracy, and compared across unimodal and multimodal feature sets. Feature importance analyses were conducted to identify key predictors. External validation in an independent cohort of 376 patients is currently underway. Results RNA-seq analysis identified 401 differentially expressed genes, which were reduced to 22 candidates through penalized regression and feature selection approaches. Subsequent qPCR validation resulted in a 6-gene biomarker panel. Models integrating genomic, clinical and psychosocial variables consistently improved predictive performance compared to those based on a single data domain. The best-performing model, based on penalized logistic regression, achieved an AUC of 0.79, outperforming tree-based approaches (AUC 0.75) and unimodal models. Across models, two genes from the biomarker panel were consistently ranked among the top predictors, alongside patient-reported quality of life and caregiver status, highlighting the complementary contribution of biological and psychosocial factors. Conclusion A multimodal machine learning approach integrating omics, clinical and psychosocial data improves HF risk prediction compared to single-domain models. The identification of a small gene panel, together with patient-centred variables, supports the value of multidimensional risk stratification. External validation is ongoing to assess generalizability.
Spatial GeoMx NGS profiling to map molecular drivers of pathological response in the tumor microenvironment of pleural mesothelioma patients receiving neoadjuvant chemo-immunotherapy: preliminary analysis from the CHIMERA multi-site trial
Angela, Grassi*, Clinical Research Unit, Veneto Institute of Oncology, IOV-IRCCS, Padua, Italy. Anna, Tosi, Immunology and Molecular Oncology Diagnostics, Veneto Institute of Oncology, IOV-IRCCS, Padua, Italy. Daniela, Scattolin, Medical Oncology 2, Veneto Institute of Oncology IOV-IRCCS, Padua, Italy; Department of Surgery, Oncology and Gastroenterology, University of Padova, Padua, Italy. Paolo Andrea, Zucali, Department of Oncology, Istituto di Ricovero e Cura a Carattere Scientifico (IRCCS) Humanitas Research Hospital, Rozzano, Milan, Italy. Federica, Grosso, Mesothelioma Unit, Azienda Ospedaliero-Universitaria SS. Antonio e Biagio e Cesare Arrigo, Alessandria. Giovanni Luca, Ceresoli, Humanitas Gavazzeni, Bergamo, Italy. Marcello, Tiseo, Department of Medicine and Surgery, University of Parma, Parma, Italy; Medical Oncology Unit, University Hospital of Parma, Parma, Italy. Alessandra, Baerz, Medical Oncology Department, CRO Aviano, National Cancer Institute, Istituto di Ricovero e Cura a Carattere Scientifico (IRCCS), Aviano, Italy. Fiorella, Calabrese, Department of Cardiac, Thoracic, Vascular Sciences and Public Health, University of Padova, Padua, Italy. Federica, Pezzuto, Department of Cardiac, Thoracic, Vascular Sciences and Public Health, University of Padova, Padua, Italy. Paola, Del Bianco, Clinical Research Unit, Veneto Institute of Oncology, IOV-IRCCS, Padua, Italy. Valentina, Guarneri, Department of Oncology, Veneto Institute of Oncology IOV-IRCCS, Padua, Italy. Department of Surgery, Oncology, and Gastroenterology, University of Padova, Padua, Italy. Antonio, Rosato, Department of Surgery, Oncology, and Gastroenterology, University of Padova, Padua, Italy; Immunology and Molecular Oncology Diagnostics, Veneto Institute of Oncology IOV-IRCCS, Padua, Italy. Gian Luca, De Salvo, Clinical Research Unit, Veneto Institute of Oncology, IOV-IRCCS, Padua, Italy. Giulia, Pasello, Medical Oncology 2, Veneto Institute of Oncology IOV-IRCCS, Padua, Italy; Department of Surgery, Oncology and Gastroenterology, University of Padova, Padua, Italy.
Spatial GeoMx NGS profiling to map molecular drivers of pathological response in the tumor microenvironment of pleural mesothelioma patients receiving neoadjuvant chemo-immunotherapy: preliminary analysis from the CHIMERA multi-site trial
Angela, Grassi*, Clinical Research Unit, Veneto Institute of Oncology, IOV-IRCCS, Padua, Italy. Anna, Tosi, Immunology and Molecular Oncology Diagnostics, Veneto Institute of Oncology, IOV-IRCCS, Padua, Italy. Daniela, Scattolin, Medical Oncology 2, Veneto Institute of Oncology IOV-IRCCS, Padua, Italy; Department of Surgery, Oncology and Gastroenterology, University of Padova, Padua, Italy. Paolo Andrea, Zucali, Department of Oncology, Istituto di Ricovero e Cura a Carattere Scientifico (IRCCS) Humanitas Research Hospital, Rozzano, Milan, Italy. Federica, Grosso, Mesothelioma Unit, Azienda Ospedaliero-Universitaria SS. Antonio e Biagio e Cesare Arrigo, Alessandria. Giovanni Luca, Ceresoli, Humanitas Gavazzeni, Bergamo, Italy. Marcello, Tiseo, Department of Medicine and Surgery, University of Parma, Parma, Italy; Medical Oncology Unit, University Hospital of Parma, Parma, Italy. Alessandra, Baerz, Medical Oncology Department, CRO Aviano, National Cancer Institute, Istituto di Ricovero e Cura a Carattere Scientifico (IRCCS), Aviano, Italy. Fiorella, Calabrese, Department of Cardiac, Thoracic, Vascular Sciences and Public Health, University of Padova, Padua, Italy. Federica, Pezzuto, Department of Cardiac, Thoracic, Vascular Sciences and Public Health, University of Padova, Padua, Italy. Paola, Del Bianco, Clinical Research Unit, Veneto Institute of Oncology, IOV-IRCCS, Padua, Italy. Valentina, Guarneri, Department of Oncology, Veneto Institute of Oncology IOV-IRCCS, Padua, Italy. Department of Surgery, Oncology, and Gastroenterology, University of Padova, Padua, Italy. Antonio, Rosato, Department of Surgery, Oncology, and Gastroenterology, University of Padova, Padua, Italy; Immunology and Molecular Oncology Diagnostics, Veneto Institute of Oncology IOV-IRCCS, Padua, Italy. Gian Luca, De Salvo, Clinical Research Unit, Veneto Institute of Oncology, IOV-IRCCS, Padua, Italy. Giulia, Pasello, Medical Oncology 2, Veneto Institute of Oncology IOV-IRCCS, Padua, Italy; Department of Surgery, Oncology and Gastroenterology, University of Padova, Padua, Italy.
- Keywords
- spatial transcriptomicsGeoMX Digital Spatial Profilerlinear-mixed effect modelspleural mesotheliomabiomarkers
- Abstract
- Objective The CHIMERA phase II multi-site trial evaluates neoadjuvant/perioperative chemo-immunotherapy (nCT-IO) in patients with stage I-IIIA epithelioid/biphasic pleural mesothelioma (PM). As a translational endpoint, this study leverages high-plex spatial transcriptomics (GeoMx NGS) to characterize the treatment-naive tumor microenvironment (TME). By applying compartment-specific differential expression analysis across tumor, stromal, CD3+, and CD68+ segments, we aim to identify biomarkers of pathological response and molecular drivers of therapy resistance to enable patient stratification and maximize clinical efficacy. Methods Treatment-naive tumor samples were analyzed using the NanoString GeoMx® Digital Spatial Profiler (DSP) across four morphologically defined segments: TUMOR, CD3+, CD68+, and OTHERS. Response was dichotomized as major pathological response (MPR; residual viable tumor, RVT≤10%) vs. non-MPR (RVT>10%). Transcriptomic data were pre-processed, quality-filtered, and normalized (Q3) using the GeomxTools Bioconductor package. While TUMOR and OTHERS analyses included all 6 MPR patients, the CD68+ and CD3+ cohorts were restricted to 4 and 3 cases, respectively, based on quality control (QC) and segment availability. Compartment-specific differential expression (DE) analysis was performed using Linear Mixed-Effects Models (LMM), treating the patient and run-batch as random intercepts to account for intra-patient correlation and technical variability. Multiple testing correction was applied using the Benjamini-Hochberg method, with significance set at a False Discovery Rate (FDR) < 0.05. Results Spatial transcriptomics was performed on 31 PM patients (6 MPR, 19%). DE was analyzed independently within each TME compartment to identify genes associated with therapeutic response. To ensure statistical robustness, the CD3+ segment was excluded from this report due to insufficient sample size in the MPR group (n=3). The stromal (OTHERS) segment showed the most extensive remodeling in non-MPR patients (25 genes, FDR<0.05), dominated by strong up-regulation of C7 (FC=10.81), SFRP2 (FC=8.76), and MGP (FC=4.34), alongside resistance markers like CXCL12 (FC=2.73) and USP33 (FC=1.36). In the TUMOR segment, USP33 was the sole statistically significant marker of resistance (FC=1.54, FDR=0.016). Conversely, MPR was characterized by specific metabolic signals, including TOMM40 enrichment in the CD68+ segment (FC=1.63, FDR=0.006) and mitochondrial factors like MIEF1 in the stroma (FC=1.25, FDR=0.031). Conclusion These preliminary data from the CHIMERA study highlight both the potential and the technical complexities of spatial transcriptomics in PM. In particular, GeoMx DSP NGS data analysis involves addressing several critical challenges, including (i) the optimization of segment-specific QC thresholds; (ii) signal recovery challenges from sparse regions of interest (ROIs) such as CD3+ and CD68+ populations; and (iii) the implementation of multi-level batch correction. Our biological results consistently link nCT-IO resistance to a fibrotic, cancer-associated fibroblast (CAF)-driven stromal program and a treatment-resistant tumor phenotype. These findings, although part of an ongoing study, demonstrate that spatial resolution is essential to elucidate resistance mechanisms occurring within the TME.
FLiCoN: Friendship Like differential Coexpression Network for the identification of Tumor-Educated Platelets Driver Genes in Glioma via structural Imbalance
Alessandro, Taraborelli, Dipartimento di Ingegneria informatica, automatica e gestionale -Università di Roma “Sapienza” Stefano, Rinaldi, Dipartimento di Ingegneria informatica, automatica e gestionale - Università di Roma “Sapienza” * Mattia, Manna, Dipartimento di Ingegneria informatica, automatica e gestionale - Università di Roma “Sapienza” Aurelia, Rughetti, Dipartimento di Medicina Sperimentale - Università di Roma “Sapienza” Lorenzo, Farina, Dipartimento di Ingegneria informatica, automatica e gestionale - Università di Roma “Sapienza” Manuela, Petti, Dipartimento di Ingegneria informatica, automatica e gestionale - Università di Roma “Sapienza”
FLiCoN: Friendship Like differential Coexpression Network for the identification of Tumor-Educated Platelets Driver Genes in Glioma via structural Imbalance
Alessandro, Taraborelli, Dipartimento di Ingegneria informatica, automatica e gestionale -Università di Roma “Sapienza” Stefano, Rinaldi, Dipartimento di Ingegneria informatica, automatica e gestionale - Università di Roma “Sapienza” * Mattia, Manna, Dipartimento di Ingegneria informatica, automatica e gestionale - Università di Roma “Sapienza” Aurelia, Rughetti, Dipartimento di Medicina Sperimentale - Università di Roma “Sapienza” Lorenzo, Farina, Dipartimento di Ingegneria informatica, automatica e gestionale - Università di Roma “Sapienza” Manuela, Petti, Dipartimento di Ingegneria informatica, automatica e gestionale - Università di Roma “Sapienza”
- Keywords
- Tumor Educated PlateletsGliomaStructural Balance TheoryDifferential CoexpressionRNA-Seq
- Abstract
- Objective: Tumor-educated platelets (TEPs) represent a pivotal resource for liquid biopsy, reflecting transcriptomic alterations induced by the tumor microenvironment. However, traditional transcriptomic analyses often focus on single-gene differential expression, overlooking high-order regulatory dynamics. While network-based approaches have emerged as powerful tools for capturing system-level properties, the signed nature of differential co-expression relationships remains largely underutilized and misinterpreted. This work proposes a computational framework to leverage Structural Balance Theory (SBT) by semantically adapting its principles to the biological context of co-expression networks. The goal is to identify driver genes associated with tumor conditions—specifically glioma—by analyzing network frustration elements. Methods: We introduce the Friendship-Like Differential Co-expression Network (FLiCoN) framework, which utilizes a Z-test module for correlation comparisons to quantify condition-specific shifts in gene-gene associations between glioma patients and healthy controls. The network is constructed through a semantic readaptation of edges: positive links (+1) denote friendship-like (functionally conserved) interactions, associated with low absolute Z-scores, while negative links (-1) represent states of tension or enmity, associated with high absolute Z-scores. We explored multiple edge selection strategies, employing thresholds based on quantiles of the absolute Z-score matrix to ensure network sparsification. This approach maintained global edge densities at 5% or 10% and fixed sign proportions positive:negative (80:20, 85:15, 90:10, 95:5) to assess both method sensitivity to thresholds and the robustness of the results. Each gene's contribution to structural imbalance was quantified using the Local Balance Index, restricted to cycles of length 3 (triangles). Statistical significance was assessed against a signed degree-preserving null model generated via a modified Maslov-Sneppen rewiring algorithm. Results: Analysis was conducted on the GSE183635 discovery cohort, comprising 114 glioma patients and 283 controls. Across the eight topological configurations generated through sparsification, the identified critical nodes exhibited exceptional consistency, with pairwise gene overlaps exceeding 50% in nearly all comparisons. This stability facilitated the extraction of a core signature of 50 recurrent driver genes. Biologically, these genes segregate into three macroscopic functional axes: platelet activation and adhesion dynamics, the sequestration of central nervous system-specific transcripts, and immune modulation within the tumor-circulation interface. Conclusion: This study establishes a new rationale for applying SBT in a bioinformatics context, specifically for differential co-expression analysis performed on tumor-educated platelets. We demonstrate how condition-induced network rewiring can be inspected through the local tension of higher-order structures. The methodological robustness observed across varying network densities and sign proportions validates the proposed semantic readaptation, while the identification of genes associated with coherent functional axes suggests high potential for this metric in biomarker discovery.
Quantum Annealing for De Novo Genome Assembly
(*) Manuel, Arcieri, Department of Computer Science, Sapienza University of Rome, Rome, Italy. Paolo, Zuliani, Department of Computer Science, Sapienza University of Rome, Rome, Italy. Giovanni, Carotenuto, Department of Ecological and Biological Sciences, University of Tuscia, Viterbo, Italy. Meryam, Carrus, Department of Ecological and Biological Sciences, University of Tuscia, Viterbo, Italy. Tiziana, Castrignanò, Department of Ecological and Biological Sciences, University of Tuscia, Viterbo, Italy.
Quantum Annealing for De Novo Genome Assembly
(*) Manuel, Arcieri, Department of Computer Science, Sapienza University of Rome, Rome, Italy. Paolo, Zuliani, Department of Computer Science, Sapienza University of Rome, Rome, Italy. Giovanni, Carotenuto, Department of Ecological and Biological Sciences, University of Tuscia, Viterbo, Italy. Meryam, Carrus, Department of Ecological and Biological Sciences, University of Tuscia, Viterbo, Italy. Tiziana, Castrignanò, Department of Ecological and Biological Sciences, University of Tuscia, Viterbo, Italy.
- Keywords
- Quantum annealingDe novo genome assemblyDe Bruijn graphD-WaveBioinformatics
- Abstract
- Objective: We present a hybrid classical-quantum algorithm for de novo genome assembly. The reconstruction of complete genomes from short sequencing reads remains computationally demanding, particularly in de novo assembly scenarios where no reference genome is available. Quantum computing, and specifically quantum annealing, has emerged as a promising solution for addressing such combinatorial optimisation challenges in bioinformatics. This work presents a novel algorithm designed for the D-Wave quantum annealer to perform de novo genome assembly of a small bacterial genome using the De Bruijn graph approach. Methods: Our method formulates the de novo genome assembly problem as an optimisation task compatible with the D-Wave quantum annealing architecture. Specifically, we construct a De Bruijn graph from short sequencing reads of a small bacterial genome, where k-mers derived from the reads serve as graph edges. The assembly problem is then encoded as a quadratic unconstrained binary optimisation (QUBO) problem suitable for execution on the D-Wave quantum annealer. This approach follows the general strategy of mapping genome assembly onto quantum annealing hardware. Our algorithm introduces novel modifications to improve the efficiency of the formulation and graph traversal for bacterial genome reconstruction on current-generation D-Wave devices. Results: Preliminary results demonstrate that our quantum annealing-based algorithm produces good-quality genome assemblies for the tested bacterial genomes. The assembled contigs show high accuracy and continuity, indicating that the D-Wave quantum annealer can effectively navigate the solution space of the De Bruijn graph-based assembly problem. Although the genomes under study are short (~5,000 bp), the results provide evidence that quantum annealing is a viable computational strategy for genome assembly, consistent with prior findings that quantum annealing devices can handle real genomic data. Conclusion: This work extends the application of quantum annealing to de novo bacterial genome assembly using the De Bruijn graph method, contributing to the growing body of evidence that quantum computing technologies can address high-complexity bioinformatics problems. Although the current study is limited to a small number of bacterial genomes, the promising assembly quality suggests that future generations of quantum annealing hardware with increased qubit counts and connectivity may enable assembly of larger and more complex genomes. This research represents a step toward leveraging quantum computing for practical genomics applications.
Toward Accessible and Reproducible Bioinformatics Workflows for Nanopore-Based Virus Surveillance
Balal Sadeghi
Toward Accessible and Reproducible Bioinformatics Workflows for Nanopore-Based Virus Surveillance
Balal Sadeghi
- Keywords
- Nanopore sequencing; virus surveillance; bioinformatics workflows; reproducibility; phylogenetic analysis; taxonomic classification
- Abstract
- Rapid virus surveillance is increasingly important for outbreak preparedness, diagnostics, and genomic monitoring. Nanopore sequencing offers major practical advantages for these applications because of its speed, portability, and flexibility. However, downstream bioinformatics analysis often remains difficult to deploy, reproduce, and adapt, particularly for users with limited computational expertise. There is therefore a clear need for integrated and user-friendly workflows that streamline key analytical steps such as read preprocessing, taxonomic classification, consensus genome reconstruction, phylogenetic analysis, and result reporting. Current practice often relies on fragmented tool combinations, complex configurations, and manual intervention, which can reduce accessibility and hinder routine use in both research and diagnostic settings. In this poster, we highlight the need for more accessible, modular, and reproducible bioinformatics solutions for nanopore-based virus surveillance and outline a workflow-oriented perspective aimed at reducing analytical complexity while preserving flexibility and scientific rigor. Such approaches could support broader adoption of genomic surveillance methods across laboratories and enable faster, more robust interpretation of viral sequencing data in routine and outbreak-related contexts.
Single-cell multi-omics integration for TCR–epitope binding inference in acute myeloid leukemia
Ludovica Celli*, Experimental Hematology Unit, IRCCS San Raffaele Hospital, Milan, Italy Francesca Marzuttini, Experimental Hematology Unit, IRCCS San Raffaele Hospital, Milan, Italy Chiara Bonini, Experimental Hematology Unit, IRCCS San Raffaele Hospital, Milan, Italy Ivan Merelli, Institute of Biomedical Technologies, National Research Council (ITB-CNR), Segrate (MI), Italy Eliana Ruggiero, Experimental Hematology Unit, IRCCS San Raffaele Hospital, Milan, Italy
Single-cell multi-omics integration for TCR–epitope binding inference in acute myeloid leukemia
Ludovica Celli*, Experimental Hematology Unit, IRCCS San Raffaele Hospital, Milan, Italy Francesca Marzuttini, Experimental Hematology Unit, IRCCS San Raffaele Hospital, Milan, Italy Chiara Bonini, Experimental Hematology Unit, IRCCS San Raffaele Hospital, Milan, Italy Ivan Merelli, Institute of Biomedical Technologies, National Research Council (ITB-CNR), Segrate (MI), Italy Eliana Ruggiero, Experimental Hematology Unit, IRCCS San Raffaele Hospital, Milan, Italy
- Keywords
- Single-cell multiomicsInnovative immunological tools and approachesAcute Myeloid LeukemiaT cell therapyMachine Learning
- Abstract
- Objective: Acute Myeloid Leukemia (AML) is a life-threatening hematological malignancy with poor prognosis, especially in elderly patients. Engineered T cells (eng-T cells), particularly those redirected with tumor-specific TCRs, represent a promising therapeutic strategy. However, identifying truly functional anti-tumor TCRs remains a major challenge due to the limited specificity of current epitope-binding assays and the incomplete validation of available TCR-epitope reference datasets. Most prediction tools rely mainly on CDR3β sequence similarity and are trained on datasets enriched for viral epitopes, limiting their applicability to tumor immunology. This project aims to develop a computational framework integrating single-cell multi-omics and TCR specificity prediction to improve prioritization of candidate anti-tumor TCRs in AML. Methods: We analyzed single-cell datasets integrating transcriptomics (scRNA-seq), TCR sequencing (VDJ), and epitope-binding signals obtained with dCODE Dextramers from peripheral CD8⁺ T cells of 10 AML patients post-HSCT to capture endogenous anti-tumor responses. Data integration and clonotype analysis were performed using Seurat, Harmony, and scRepertoire. Due to the high degree of non-specific and multipronged Dextramer binding, we implemented a refined strategy based on deviations in epitope-binding intensity distributions to improve TCR-target assignment beyond standard threshold-based methods. To infer antigen specificity, we combined GLIPH2 clustering with multiple TCR-epitope prediction tools, including TCRex, MixTCRpred, TCRGP, and a sequence generative model named GRIP (Generative Reconstruction of antIgen Peptides), based on an LSTM architecture with attention mechanisms. Predictions were integrated into a consensus prioritization score designed to: (i) identify TCRs clustering with viral-associated repertoires, allowing exclusion of likely bystander or cross-reactive clones, and (ii) detect convergence with previously reported tumor-reactive TCRs to support functional deorphanization and candidate prioritization for in vitro validation. Results: Our integrative pipeline identified 10 transcriptionally distinct CD8⁺ T-cell subsets in AML samples. Epitope-binding-positive cells were distributed across multiple clusters, with a substantial fraction of double-binding clonotypes (12.3%), supporting widespread TCR multipronged specificity. Notably, expanded clonotypes were significantly enriched within cytotoxic and exhausted T-cell states, highlighting the importance of integrative approaches combining epitope–TCR binding information and transcriptional profiling to better infer T-cell functional specificity. Cross-model comparison revealed substantial variability in predicted TCR-epitope associations, highlighting limitations in currently available validation datasets, particularly for tumor-associated antigens. Nevertheless, integrating multiple predictors enabled identification of recurrent specificity patterns and improved confidence in candidate selection. Our consensus scoring framework successfully filtered TCRs associated with viral-like signatures while enriching for putative tumor-reactive clonotypes prioritized for downstream functional validation. Conclusions: This study highlights the importance of integrative computational approaches for decoding anti-tumor T-cell responses in AML. By integrating single-cell multi-omics with refined Dextramer signal interpretation and complementary TCR-epitope prediction strategies, we developed a consensus scoring framework to prioritize candidate anti-tumor TCRs despite extensive cross-reactivity and limited availability of validated tumor-specific reference datasets
Towards Full Bioinformatics Automation at the National Facility for Genomics of Human Technopole
Luigi Antonio Lamparelli, Human Technopole* Daniel Carrillo Bautista, Human Technopole*
Towards Full Bioinformatics Automation at the National Facility for Genomics of Human Technopole
Luigi Antonio Lamparelli, Human Technopole* Daniel Carrillo Bautista, Human Technopole*
- Keywords
- automationsequencingdatabase
- Abstract
- Objective The National Facility for Genomics (NFG) at Human Technopole (HT) has supported cutting-edge research conducted in the institute since its establishment in 2021. In June 2024, NFG services were extended to the broader Italian scientific community through competitive open calls, allowing universities, hospitals, and other public research institutions across the country to access a diverse portfolio of services. Operating as a high-throughput production environment, the NFG manages complex workflows encompassing sample preparation, sequencing, data analysis, and result delivery. Given the diversity of protocols and platforms in use — each requiring tailored analytical approaches—a robust system is essential to ensure that bioinformatic analyses are conducted in a robust, automated and reproducible manner. Such a centralized system is using as backbone a complex relational database that captures both metadata and metrics. This database can be used for the automation of the processes or to organize information by projects, experiments, timeframes, and more—delivering substantial value for both day-to-day operations and the long-term strategic planning of the NFG. Methods Every project at the NFG begins in the laboratory, where samples are prepared and sequenced. Each sequencing run represents a discrete unit of work for the bioinformatics team. Ideally, once the run is started, a chain of automated and sequential processes is triggered, including sample sheet generation, basecalling and demultiplexing preprocessing, quality control, primary analysis execution, and finally metrics collection and database population. Currently, the system relies on Nextflow to manage pipelines, including both publicly available workflows from nf-core and custom in-house developments which also use custom Python packages. All custom code is version-controlled using GitLab, which also enables continuous integration and testing of each development via CI/CD pipelines. The workflows are executed on an HPC environment using Singularity, providing frozen, isolated environments that ensure reproducibility. Results The main goal of this initial phase of the NFG is to establish an unsupervised system that is both reliable and easy to expand when new elements are introduced by the laboratory. The proposed method interacts autonomously with the wet lab through automatic validation of sample sheets. It ensures reliability by tracking all metadata generated during sample sheet validation and maintaining it throughout the workflow. The system is also designed to be easily expandable, allowing new elements to be introduced by updating config files, without the need to modify the underlying code. Conclusions The NFG has established a centralized system automating complex workflows via Nextflow, Singularity, and rigorous version control. By integrating relational databases and autonomous wet-lab interactions, the facility provides scalable infrastructure for national genomic services. Future plans include event-triggered orchestration and a unified web application for all analytical results.
A tumor-agnostic network medicine framework reveals topological rewiring of systemic T-cell immunity in long-surviving cancer patients
*Davide, Capozzi, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Flavio, Valentino, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Lucrezia, Tuosto, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Angela, Asquino, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Angelica, Pace, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Alessio, Cirillo, Department of Radiological, Oncological and Pathological Science, “Sapienza” University of Rome, Rome, Italy Alain, Gelibter, Division of Oncology, Department of Radiological, Oncological and Pathological Science, Policlinico Umberto I, “Sapienza” University of Rome, Rome, Italy Andrea, Botticelli, Department of Radiological, Oncological and Pathological Science, “Sapienza” University of Rome, Rome, Italy Ilaria Grazia, Zizzari, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Chiara, Napoletano, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Paola, Paci, Department of Computer, Control, and Management Engineering Antonio Ruberti, Sapienza University of Rome, Italy Aurelia, Rughetti, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy
A tumor-agnostic network medicine framework reveals topological rewiring of systemic T-cell immunity in long-surviving cancer patients
*Davide, Capozzi, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Flavio, Valentino, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Lucrezia, Tuosto, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Angela, Asquino, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Angelica, Pace, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Alessio, Cirillo, Department of Radiological, Oncological and Pathological Science, “Sapienza” University of Rome, Rome, Italy Alain, Gelibter, Division of Oncology, Department of Radiological, Oncological and Pathological Science, Policlinico Umberto I, “Sapienza” University of Rome, Rome, Italy Andrea, Botticelli, Department of Radiological, Oncological and Pathological Science, “Sapienza” University of Rome, Rome, Italy Ilaria Grazia, Zizzari, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Chiara, Napoletano, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy Paola, Paci, Department of Computer, Control, and Management Engineering Antonio Ruberti, Sapienza University of Rome, Italy Aurelia, Rughetti, Laboratory of Tumor Immunology, Department of Experimental Medicine, Sapienza University of Rome, Rome, Italy
- Keywords
- Network MedicineTopological RewiringImmune Checkpoint InhibitorsFlow CytometryT-cell Memory
- Abstract
- Objective Although a particular subset of patients treated with immune-checkpoint inhibitors (ICIs) consistently shows long-term survival, the immune mechanisms behind such phenomena remain largely unknown. Standard analyses of flow-cytometric data often rely on univariate statistics, thus minimizing the underlying complexity of the immune system. We propose a network medicine framework that reliably uncovers non-obvious patterns among target T-cell subgroups in flow-cytometric data. Our objective is to investigate the topological rewiring of T-cell response in long survivors (LS) across different solid tumor cohorts, with the aim of uncovering a tumor-agnostic immune signature. Methods Peripheral blood mononuclear cells were collected from 93 patients with distinct solid tumors undergoing ICI therapy and 52 healthy donors and then analyzed using multiparametric flow-cytometry for a total of 31 T-cell populations. Patients were divided according to clinical parameters: LS (Overall-Survival > 18/24 months), Early Progressors (EP) (Progression-Free-Survival (PFS) ≤ 3 months), the remaining patients were defined as Intermediate (INT). To address the bounded, non-normal nature of cytometric percentages and ensure statistical rigor, we applied a logit transformation with limit of detection handling coupled with James-Stein shrinkage partial correlation, isolating direct cellular interdependencies. To robustly investigate immune signatures among subgroups, we implemented a differential overlay pipeline: network edges were first validated for stability via bootstrap resampling, followed by permutation testing to assess significant rewiring across clinical outcomes. Stable edges were then topologically classified (conserved, specific, or inverted). Finally, topological metrics (degree, betweenness, and Jaccard index) were integrated with sPLS-DA feature loadings to identify systemic master regulators. Results Comparison between clinical outcomes exposed a profound topological shift. The EP network is highly centralized around activation compartments (CD3+PD1+Effector, degree=11). Interestingly, Both LS and INT patients lose this activation centrality (degree drops to 5). However, while INT patients lose this signature, only the LS network undergoes extensive structural rewiring to establish a costimulatory, memory-driven topology. Here, the central hub is represented by CD3+CD137+Central memory (degree=12). Interestingly, Non-Suppressive Tregs act as a critical structural bottleneck specifically in the LS network, exhibiting the highest systemic betweenness (0.101). This transition is orchestrated by specific master regulators: notably, CD3+CD137+Naïve exhibits major discriminating power (sPLS-DA loading=0.919) and undergoes massive structural rewiring (Jaccard=0.083). This node-centric divergence is driven by statistically robust edge inversions, prominently the interactions between CD3+CD137+Effector memory RA+ and CD3+PD1+Naive (p=0.0015 in EP vs LS), and between CD3+Central memory and CD3+Effector memory (p=0.0065 in EP vs LS), proving that long-term survival relies on reversing specific immunological axes rather than global immune restoration. Conclusion From our results, long-term survival seems supported by rewired CD137+ memory T-cell networks. Our framework demonstrates that integrating edge inversions and nodal metrics from flow-cytometry data provides deep biological insights invisible to standard univariate analyses.
High-Performance Tensor-Based HSMM Decoding for Genome-Scale CpG Island Detection
Lorenzo, Piarulli, Sapienza University of Rome*. Elia, Belli, Sapienza University of Rome*. Daniele, De Sensi, Sapienza University of Rome.
High-Performance Tensor-Based HSMM Decoding for Genome-Scale CpG Island Detection
Lorenzo, Piarulli, Sapienza University of Rome*. Elia, Belli, Sapienza University of Rome*. Daniele, De Sensi, Sapienza University of Rome.
- Keywords
- HPCHidden Semi Markov ModelsGPUGenomics
- Abstract
- Hidden semi-Markov models (HSMMs) generalize classical hidden Markov models (HMMs) by replacing the implicit geometric state-duration assumption with explicit, arbitrary duration distributions, enabling more faithful modeling of temporal processes where segment lengths carry domain-specific meaning. Despite their superior expressiveness, HSMM algorithms have seen limited adoption in large-scale genomic applications because they introduce an additional computational factor proportional to the maximum state duration, raising complexity from O(TN²) to O(TN²D) and making whole-genome analysis intractable with existing sequential implementations. We leverage a high-performance library recently developed by our team, which reformulates the HSMM Viterbi algorithm into structured tensor operations, replacing the classical four-nested-loop computation with broadcasted tensor products over a three-dimensional structure. This reformulation exposes massive parallelism, mapping naturally onto SIMD units, multi-core CPUs, and, for the first time for HSMMs, GPUs. On standard benchmarks, the library achieves speedups of up to 14× on a single core, over 200× with multi-core parallelization, and over 570× on GPU relative to the state-of-the-art sequential baseline, reducing workloads from hours to seconds. To demonstrate the applicability of this library to genomic sequence analysis, we present its use on CpG island detection, a classical bioinformatics problem with direct relevance to gene regulation and epigenetic studies. CpG islands are genomic regions enriched in CpG dinucleotides, typically located at gene promoters and involved in methylation-mediated gene silencing. Existing detection methods include rule-based sliding-window approaches with fixed thresholds and HMM-based probabilistic segmentations. Both families have known limitations: rule-based methods are sensitive to arbitrary parameter choices and do not generalize across species, while HMMs impose geometric duration distributions that poorly reflect the empirically observed, multimodal length distribution of CpG islands. We formulate CpG island detection as a two-state HSMM with 16 emission symbols representing all possible dinucleotide pairs over the four-nucleotide DNA alphabet, and non-parametric duration distributions fitted from existing annotations. This formulation captures CpG dinucleotide enrichment through the emission model while separately encoding realistic segment-length constraints through explicit durations. We evaluate detection on human chromosomes 21 and 22, the standard benchmarking sequences used throughout the CpG island literature owing to their early complete sequencing and thorough annotation. Our analysis investigates whether the explicit duration modeling reveals CpG-enriched regions missed by existing methods or improves boundary precision for known islands. More broadly, this work opens the possibility of running HSMM-based genomic solvers efficiently on supercomputing infrastructure, enabling the bioinformatics community to deploy expressive probabilistic models at whole-genome scale without the computational compromises that have historically limited their adoption.
GLM-Based Parametric Survival Modeling Under Two-Phase Sampling: A Computational Framework with Application to High-Dimensional Methylation Data
*David, Soave, Wilfrid Laurer University Karina, Kwan, McGill University Celia, Greenwood, Lady Davis Institute for Medical Research
GLM-Based Parametric Survival Modeling Under Two-Phase Sampling: A Computational Framework with Application to High-Dimensional Methylation Data
*David, Soave, Wilfrid Laurer University Karina, Kwan, McGill University Celia, Greenwood, Lady Davis Institute for Medical Research
- Keywords
- Two-phase samplingCase-cohort designParametric survival modelingInverse probability weightingBiomarker risk prediction
- Abstract
- Objective Outcome-dependent two-phase designs, such as case-cohort and matched biomarker sub-studies, are widely used in large biomedical cohorts to reduce the cost of measuring expensive exposures. Standard analyses rely on weighted Cox regression, which leaves the baseline hazard unspecified and can complicate smooth absolute risk estimation. We propose a scalable parametric hazard modeling framework that extends the casebase representation to outcome-dependent two-phase sampling while retaining implementation within standard generalized linear model (GLM) software. Methods The proposed weighted casebase framework combines (i) an offset accounting for person-moment sampling and (ii) inverse-probability weights correcting for unequal inclusion probabilities under two-phase sampling. Estimation is performed using weighted logistic regression, enabling direct specification of parametric baseline hazards and smooth absolute risk prediction. We develop a computationally efficient variance estimator by aggregating dfbeta residuals at the individual level, avoiding repeated model refitting. Finite-sample performance was evaluated via simulation across cohort sizes up to 100,000 with binary, normal, and log-normal covariates. We further applied the method to a high-dimensional methylation biomarker sub-study within the Ontario Health Study, incorporating supervised principal components and cross-validation. Results In simulations, the weighted casebase estimator demonstrated bias, variance, and coverage properties comparable to weighted Cox regression across cohort sizes and covariate types. As inverse-probability weights increased, both approaches exhibited instability for continuous covariates, reflecting known sensitivity to extreme weights rather than features specific to the proposed framework. In the biomarker application, weighted casebase and weighted Cox models achieved similar predictive discrimination, while the casebase approach provided smooth hazard and absolute risk curves within a GLM workflow. Conclusion The weighted casebase framework offers a computationally efficient and theoretically grounded approach to parametric survival modeling under outcome-dependent two-phase sampling. By leveraging standard GLM software, it facilitates scalable implementation, smooth risk estimation, and integration into high-dimensional biomedical prediction pipelines.
A unified Medical Informatics black‑box framework for integrating environmental sampling and laboratory analytical data at the Luxembourg National Health Laboratory (LNS)
El Hassane, Ouaalaya*, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Matteo, Creta, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Giuseppe, Arena, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Françoise, Schaefers, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Nicolas, Joblin, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Cristiana, Costa Pereira, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Maria-Mirela, Ani, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Emilie, Hardy, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Maria, Torres Toda, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Ruth, Moeller, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Lorenzo, Favilli, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Radu, Duca, Department Health Protection, Luxembourg Health Directorate, Strassen, Luxembourg. An, van Nieuwenhuyse, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg.
A unified Medical Informatics black‑box framework for integrating environmental sampling and laboratory analytical data at the Luxembourg National Health Laboratory (LNS)
El Hassane, Ouaalaya*, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Matteo, Creta, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Giuseppe, Arena, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Françoise, Schaefers, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Nicolas, Joblin, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Cristiana, Costa Pereira, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Maria-Mirela, Ani, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Emilie, Hardy, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Maria, Torres Toda, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Ruth, Moeller, Medical Expertise and Data Intelligence Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Lorenzo, Favilli, Environmental Hygiene and Human Biological Monitoring Service, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg. Radu, Duca, Department Health Protection, Luxembourg Health Directorate, Strassen, Luxembourg. An, van Nieuwenhuyse, Department of Health Protection, Laboratoire National de Santé, Dudelange, Luxembourg.
- Keywords
- Medical InformaticsEnvironmental Health SurveillanceData IntegrationLaboratory Information SystemsDecision Support
- Abstract
- Objective Environmental health surveillance relies on the reliable collection, traceability, and interpretation of complex data generated across multiple operational layers, from field sampling to laboratory analysis. Within the Health Protection Department of the National Health Laboratory (LNS) in Luxembourg, environmental monitoring programs targeting chemical and microbiological exposures in indoor environments require close interaction between field and laboratory services. However, fragmentation between data collection tools and laboratory information systems can limit data quality, interoperability, and downstream data reuse. The objective of this work is to develop a unified Medical Informatics framework that creates a single global environmental health database by integrating field sampling activities with standardized laboratory analytical results. This work introduces a black‑box platform that automatically extracts, cleans, harmonizes, and corrects heterogeneous data collected in the field and generated in the laboratory, producing analysis‑ready datasets and initial descriptive outputs. Methods The framework was implemented within the Health Protection Department of LNS, supporting the workflows of the Medical Expertise and Data Intelligence (MEDI) service for environmental field sampling and the Environmental Hygiene and Human Biological Monitoring (EnvirOH) service for laboratory analysis. Nurses (trained in environmental sampling procedures) conducted standardized environmental sampling across multiple sites using a tablet‑based Application Form Analysis (AFA) system to capture structured data and contextual information. Environmental sampling covered multiple matrices and analytical domains. Sampling strategies were adapted to the targeted compounds or microorganisms and included active air sampling, dust collection, and surface sampling. Subsequently, the samples underwent a range of chemical and microbiological analyses to characterize indoor environmental exposures. Laboratory analyses were conducted using appropriate analytical techniques, while the GLIMS laboratory information management system was used for sample registration upon arrival, as well as for result validation and reporting. A central black‑box Medical Informatics layer automatically extracts, cleans, harmonizes, and corrects data from field and laboratory systems, integrates analytical metadata, and generates a unified, quality‑controlled database together with descriptive outputs and spatial heat maps. Results The system successfully consolidated end‑to‑end environmental sampling and laboratory analytical data into a single coherent dataset. Automation significantly reduced manual data handling, improved data consistency, and enhanced traceability across services. The generated descriptive outputs and heat maps enabled rapid visualization of exposure patterns and spatial variability within indoor environments. Conclusion This study demonstrates how a black‑box Medical Informatics framework can bridge field sampling and laboratory information systems to support integrated environmental health surveillance. By delivering both analysis‑ready datasets and early descriptive outputs, the proposed approach enhances data quality, interoperability, and decision support. It further provides a scalable and transferable model for national environmental surveillance within a One Health perspective
From hepatosome to interactome: network-based discovery of UBIAD1 as a lipid-metabolism target in hepatocellular carcinoma
Federica D’Annunzio*, Department of Translational and Precision Medicine, Sapienza University of Rome, Rome, Italy Matteo Pedrelli, Cardio Metabolic Unit, Department of Laboratory Medicine, and Department of Medicine (Huddinge), Karolinska Institutet, Stockholm, Sweden Giulia Fiscon, Department for the Promotion of Human Science and Quality of Life, San Raffaele Open University, Rome, Italy. Institute for Systems Analysis and Computer Science “Antonio Ruberti” National Research Council, Rome, Italy Paola Paci, Department of Computer, Control and Management Engineering, Sapienza University of Rome, Rome, Italy. Institute for Systems Analysis and Computer Science “Antonio Ruberti”, National Research Council, Rome, Italy Paolo Parini, Cardio Metabolic Unit, Department of Laboratory Medicine, and Department of Medicine (Huddinge), Karolinska Institutet, Stockholm, Sweden
From hepatosome to interactome: network-based discovery of UBIAD1 as a lipid-metabolism target in hepatocellular carcinoma
Federica D’Annunzio*, Department of Translational and Precision Medicine, Sapienza University of Rome, Rome, Italy Matteo Pedrelli, Cardio Metabolic Unit, Department of Laboratory Medicine, and Department of Medicine (Huddinge), Karolinska Institutet, Stockholm, Sweden Giulia Fiscon, Department for the Promotion of Human Science and Quality of Life, San Raffaele Open University, Rome, Italy. Institute for Systems Analysis and Computer Science “Antonio Ruberti” National Research Council, Rome, Italy Paola Paci, Department of Computer, Control and Management Engineering, Sapienza University of Rome, Rome, Italy. Institute for Systems Analysis and Computer Science “Antonio Ruberti”, National Research Council, Rome, Italy Paolo Parini, Cardio Metabolic Unit, Department of Laboratory Medicine, and Department of Medicine (Huddinge), Karolinska Institutet, Stockholm, Sweden
- Keywords
- Hepatocellular carcinomaNetwork Medicinecholesterol metabolism
- Abstract
- Objective: Hepatocellular carcinoma (HCC) is the third leading cause of cancer-related mortality worldwide, and dysregulated cholesterol metabolism is increasingly recognized as a driver of disease progression. Observational studies associate lipid-lowering drug use with reduced HCC risk and improved outcomes, yet the molecular mediators underlying these effects remain unclear. This study applied a Network Medicine framework to identify novel molecular targets linking cholesterol metabolism and HCC. Methods: SOAT2-only-HepG2 cells, an engineered cell line expressing only the cholesterol-esterifying enzyme SOAT2, were treated with atorvastatin, ezetimibe, or their combination (CT), followed by bulk RNA sequencing. SWIM (Paci et al., 2017) was applied to each treatment condition versus placebo controls to identify switch genes, defined as differentially expressed genes with specific topological properties associated with critical phenotypic transitions. Switch genes and their negatively correlated neighbors were then used to construct the hepatosome, a bipartite network linking genes to hepato-biliary diseases. CT switch genes were subsequently mapped onto the human protein-protein interaction network (PPI) to assess whether they formed a statistically significant disease module. Finally, the Random Walk with Restart (RWR) algorithm was applied to the interactome using known atorvastatin and ezetimibe targets as seed nodes to identify switch gene products in close network proximity to these drug targets. Results: The CT condition generated the most robust SWIM output: 28% of differentially expressed genes were classified as switch genes, nearly all uniformly downregulated and clustering within a single module. In the hepatosome, liver carcinoma emerged as the most highly connected disease node, with almost 80% of switch genes associated with hepatic tumor-related pathologies. Mapping onto the interactome revealed a statistically significant connected module. RWR analysis identified 70 switch genes in close network proximity to known atorvastatin and ezetimibe targets; UBIAD1 (UbiA Prenyltransferase Domain-Containing Protein 1) ranked 11th among 18,505 interactome proteins by modified z-score, identifying it as a putative novel target of these drugs. UBIAD1 was a switch gene across all treatment conditions, was linked to hepatic diseases within the hepatosome, and showed physical interactions with HMGCR and TPX2 in the PPI network. Despite physically interacting in the PPI network, UBIAD1 and HMGCR displayed a negative relationship in the SWIM correlation network. Drug-induced UBIAD1 downregulation was accompanied by compensatory HMGCR upregulation, consistent with an SREBP2-driven sterol feedback response. This pattern was replicated in RNA-seq data from liver biopsies of patients treated with simvastatin and ezetimibe (The Stockholm Study), supporting the translational relevance of this putative mechanism. Conclusion: By applying Network Medicine approaches in a unique pre-clinical hepatocyte model, we identified UBIAD1 as a high-priority node at the interface of cholesterol metabolism and HCC. The proposed UBIAD1-HMGCR inverse regulatory mechanism warrants functional validation and supports further investigation of statin-ezetimibe combinations as potential anticancer adjuvant therapies in HCC.
G⁴REP: A deep learning framework for the prediction of human RNA G-quadruplex-binding proteins
Serena, Rosignoli, Centre for Regenerative Medicine “Stefano Ferrari” Department of Life Sciences - University of Modena and Reggio Emilia. Sophie, Taraglio, Department of Biochemical Sciences “A. Rossi Fanelli” - Sapienza University of Rome. Francesco, Di Luzio, Department of Information Engineering, Electronics and Telecommunications - Sapienza University of Rome. * Elisa, Lustrino, Department of Biochemical Sciences “A. Rossi Fanelli” - Sapienza University of Rome. Dario, Marzella, Medical BioSciences Department - Radboud University Medical Center. Arne, Elofsson, Department of Biochemistry and Biophysics and Science for Life Laboratory - Stockholm University. Massimo, Panella, Department of Information Engineering, Electronics and Telecommunications - Sapienza University of Rome. Alessandro, Paiardini, Department of Biochemical Sciences “A. Rossi Fanelli” - Sapienza University of Rome.
G⁴REP: A deep learning framework for the prediction of human RNA G-quadruplex-binding proteins
Serena, Rosignoli, Centre for Regenerative Medicine “Stefano Ferrari” Department of Life Sciences - University of Modena and Reggio Emilia. Sophie, Taraglio, Department of Biochemical Sciences “A. Rossi Fanelli” - Sapienza University of Rome. Francesco, Di Luzio, Department of Information Engineering, Electronics and Telecommunications - Sapienza University of Rome. * Elisa, Lustrino, Department of Biochemical Sciences “A. Rossi Fanelli” - Sapienza University of Rome. Dario, Marzella, Medical BioSciences Department - Radboud University Medical Center. Arne, Elofsson, Department of Biochemistry and Biophysics and Science for Life Laboratory - Stockholm University. Massimo, Panella, Department of Information Engineering, Electronics and Telecommunications - Sapienza University of Rome. Alessandro, Paiardini, Department of Biochemical Sciences “A. Rossi Fanelli” - Sapienza University of Rome.
- Keywords
- RNA G-quadruplexesG-quadruplex-binding proteinsdeep learningLSTMstress granules
- Abstract
- Objective RNA G-quadruplex-binding proteins (RG4BPs) are key regulators of RNA metabolism, gene expression, and cellular stress responses, but their experimental identification remains technically demanding and resource-intensive. This study aimed to develop a scalable deep learning framework for the prediction of human RG4BPs from primary protein sequences and to make this approach accessible through a web-based platform. Methods Known human RG4BPs were collected from the literature and QUADRatlas, while negative examples were selected from proteins unlikely to bind G-quadruplexes based on functional annotation, localization, and structural features. After duplicate removal, homology reduction, length filtering, and dataset balancing, proteins were split into training, validation, and test sets. Protein sequences were encoded using one-hot encoding and ESM-2 protein language model embeddings. Five neural network architectures combining LSTM, CNN, attention, and fully connected layers were evaluated. Model performance was assessed using standard binary classification metrics, including accuracy and AUROC. The best model was then applied to the human proteome and to a curated set of stress granule-associated proteins. A residue-level scoring strategy was also implemented to highlight candidate RG4-binding regions, integrating intrinsic disorder propensity and enrichment of residues associated with G4 recognition. Results The best-performing model, named G⁴REP, combined ESM-2 embeddings with an LSTM architecture. It achieved strong predictive performance, with approximately 84–86% accuracy in distinguishing RG4BPs from non-binding proteins and an AUROC of 0.917 on the test set. Application to the human proteome identified more than 2,000 candidate RG4BPs, including 552 high-confidence candidates with prediction scores above 0.95. These high-scoring proteins were significantly enriched in intrinsically disordered regions. Analysis of stress granule-associated proteins further revealed a strong enrichment of predicted RG4BPs, supporting a link between RG4 recognition, intrinsic disorder, and cellular stress response. Structural modeling of selected candidates, such as FAM98B, suggested plausible G4-recognition modes involving RGG motifs and aromatic residues. Conclusion G⁴REP provides an effective and accessible deep learning framework for predicting human RNA G-quadruplex-binding proteins and prioritizing candidate binding regions. The results expand the known RG4BP landscape and suggest that RG4 recognition may be closely connected to stress granule biology and stress-responsive RNA regulation. The accompanying G⁴REP web server enables broad use of the model for RG4BP prediction and analysis.
REV-AGE and REV-AGE 2.0: From Systematic Evidence Mapping to AI-Guided Mechanistic Discovery in Aging Research
Alexandra Muntiu, Fondazione EBRI Rita Levi-Montalcini, Rome, Italy Veronica De Paolis, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Alessia Formato, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Alessia Gambadoro, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Ramona Lattao, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Sara Lazzari, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Silvia Mandillo, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Francesca Matarazzo, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Leiron Ferrarese, European Molecular Biology Laboratory (EMBL Rome), Monterotondo, Italy Chiara Parisi, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Miriam Pasquini, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Sara Vincenti, Fondazione Policlinico Universitario Agostino Gemelli IRCCS, Rome, Italy Francesco Chiani*, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy
REV-AGE and REV-AGE 2.0: From Systematic Evidence Mapping to AI-Guided Mechanistic Discovery in Aging Research
Alexandra Muntiu, Fondazione EBRI Rita Levi-Montalcini, Rome, Italy Veronica De Paolis, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Alessia Formato, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Alessia Gambadoro, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Ramona Lattao, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Sara Lazzari, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Silvia Mandillo, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Francesca Matarazzo, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Leiron Ferrarese, European Molecular Biology Laboratory (EMBL Rome), Monterotondo, Italy Chiara Parisi, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Miriam Pasquini, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy Sara Vincenti, Fondazione Policlinico Universitario Agostino Gemelli IRCCS, Rome, Italy Francesco Chiani*, Institute of Biochemistry and Cell Biology (IBBC), National Research Council (CNR), Monterotondo, Italy
- Keywords
- AgingArtificial IntelligenceMechanistic DiscoverySystematic ReviewAdverse Outcome Pathways
- Abstract
- Objective Aging is a multifactorial biological process involving interconnected molecular, cellular, metabolic, and neurodegenerative mechanisms. The increasing volume and complexity of biomedical literature make systematic interpretation and mechanistic integration increasingly challenging. We developed REV-AGE and REV-AGE 2.0, two complementary frameworks designed to combine systematic evidence mapping with AI-assisted mechanistic discovery in aging research. Methods REV-AGE was structured as a collaborative systematic-review platform aimed at collecting, organizing, and analyzing literature related to compounds and interventions potentially associated with aging modulation. Literature retrieval strategies were standardized across PubMed, Web of Science, and Scopus, followed by deduplication and active-learning assisted screening using ASReview. REV-AGE 2.0 extends this approach through a human-in-the-loop AI pipeline capable of extracting mechanistic biological relationships from scientific title and abstract datasets. The system integrates: • automated sentence classification, • large language model (LLM)-based extraction of biological triples, • post-filtering and normalization procedures, • semantic clustering, • evidence aggregation and support scoring. Extracted relationships are progressively organized into approximate AOP-like (Adverse Outcome Pathway-like) structures connecting molecular targets, pathways, cellular processes, and phenotypic outcomes. The pipeline was implemented using a FastAPI backend, local LLM inference infrastructure, and collaborative remote-access architecture. Results Preliminary analyses demonstrated the feasibility of combining systematic-review methodologies with AI-guided mechanistic extraction. The pipeline successfully identified biologically relevant interactions involving signaling pathways, neuroimmune mechanisms, inflammatory processes, neuroplasticity-related pathways, and oxidative stress modulation from curated aging-related literature datasets. The integration of semantic clustering and evidence-weighted relationship extraction enabled the generation of interpretable mechanistic graphs from heterogeneous literature sources. Initial testing suggested that post-filtering and support-based aggregation substantially improved signal-to-noise ratio by reducing generic or non-mechanistic relationships while preserving biologically meaningful interactions. The modular architecture also enabled collaborative multiuser workflows and scalable analysis pipelines applicable to compounds, stressors, and disease-oriented contexts. Conclusion REV-AGE and REV-AGE 2.0 represent an integrated framework combining systematic evidence synthesis with AI-assisted mechanistic discovery for aging research. Rather than replacing expert interpretation, the platform is designed to support researchers through scalable literature analysis, hypothesis generation, and mechanistic evidence organization. This approach may facilitate the identification of emerging biological pathways, improve literature interpretability, and contribute to the development of more structured mechanistic models in aging and age-related diseases
Closing the Design-Validation Loop for AI-Driven mRNA Therapeutics via Structured Knowledge Accumulation
Eric, Audemard, D2R-HeDS McGill University (Canada) *Sebastian, Ballesteros Ramirez, D2R-HeDS McGill University (Canada) Senthilkumar, Kailasam, D2R-HeDS McGill University (Canada) Maryam, Youssef, D2R-HeDS McGill University (Canada) Guillaume, Bourque, D2R-HeDS and Department of Human Genetics McGill University (Canada) Mathieu, Bourgey, D2R-HeDS McGill University (Canada)
Closing the Design-Validation Loop for AI-Driven mRNA Therapeutics via Structured Knowledge Accumulation
Eric, Audemard, D2R-HeDS McGill University (Canada) *Sebastian, Ballesteros Ramirez, D2R-HeDS McGill University (Canada) Senthilkumar, Kailasam, D2R-HeDS McGill University (Canada) Maryam, Youssef, D2R-HeDS McGill University (Canada) Guillaume, Bourque, D2R-HeDS and Department of Human Genetics McGill University (Canada) Mathieu, Bourgey, D2R-HeDS McGill University (Canada)
- Keywords
- mRNA TherapeuticsAI-readyData ManagementCodon OptimizationReproducibility
- Abstract
- Objective A central bottleneck in AI-driven mRNA therapeutic design is the lack of structured datasets pairing computational predictions with experimental validation. In-silico tools generate simulated metrics, but experimental results are collected in heterogeneous, unstructured formats, preventing systematic learning and model refinement. We present a two-platform framework designed to close this design-validation loop, positioning reproducible knowledge accumulation as the foundation of AI-driven mRNA therapeutic research. Methods DOTS-RNA formalizes the design side through a three-tier abstraction comprising sequences, constructs, and projects. Sequences carry content addressable identifiers enabling global deduplication and accumulate typed features (Kozak scores, secondary structure predictions, codon usage metrics) recorded across versioned, integrity-verified optimization runs. Multi-region constructs, including linear mRNA, self-amplifying RNA, circular RNA, and CAR-T modalities, are modeled as ordered, typed compositions preserving spatial context across 5′UTR, CDS, 3′UTR, and poly(A) regions. Independent optimization components operate per region, supporting a divide-and-conquer design and future integration with LLM-driven pipelines. BioDASH structures the validation side by tracking experimental samples across mRNA production and analysis workflows. It captures measured outputs (expression levels, capping efficiency, RNA integrity, encapsulation efficiency) from standard techniques including flow cytometry, ddPCR, RT-PCR, HPLC, and fragment analysis. Results are linked to samples via versioned workflow definitions and validated metric schemas, reducing data silos across teams. The two platforms are connected through an alignment layer: sequences selected in DOTS‑RNA are instantiated as BioDASH samples using shared identifiers, and experimental measurements are mapped back to their originating computational features, creating paired prediction–outcome records. Results As a proof of concept, we applied the framework to a published synonymous codon library of GFP variants in which the first 10 codons were systematically altered, and expression measured by flow cytometry. Analysis focused on mean fluorescence, consistent with prior observations that codon usage predominantly affects mean expression. Codon Adaptation Index (CAI) computed using the 2022 Codon Statistics Database explained more variance (R² = 0.16) than CAI using the Sharp (1987) reference (R² = 0.10), demonstrating that reference choice materially impacts predictive power. Additionally, the Codon Health Index (CHI) was incorporated; however, its narrow dynamic range in this dataset (0.63–0.99) prevented a statistically significant effect, instead confirming the framework's adaptability to new measurements. Conclusion This proof of concept demonstrates that heterogeneous predictions and experimental measurements can be consistently aligned and re-analyzed within a single structured system. Even on a small dataset, the integration produced insights, such as the sensitivity of CAI correlations to reference codon tables, while identifying when emerging metrics cannot be assessed due to constrained experimental ranges. By prioritizing data integration over isolated tool performance, this framework facilitates reproducible learning, comparative benchmarking, and iterative model refinement, offering a scalable foundation for AI-driven mRNA therapeutic development.
Artificial Intelligence in Oncology: Trends and Insights from 2 Systematic Evidence
Riccardo Carbonetti, Antonella Cammarota, Alex Martino Cinnera, Angela Mastronuzzi, Giovanni Morone, Francesco Negrini, Federica D’Antonio
Artificial Intelligence in Oncology: Trends and Insights from 2 Systematic Evidence
Riccardo Carbonetti, Antonella Cammarota, Alex Martino Cinnera, Angela Mastronuzzi, Giovanni Morone, Francesco Negrini, Federica D’Antonio
- Keywords
- Machine Learning; Deep Learning; Cancer; Diagnosis; Prognosis
- Abstract
- The use of artificial intelligence (AI) in medicine has surged. Our overview of systematic reviews (SRr) aims to provide a broad summary of the explored applications of AI in oncological medicine. Information on qualitative, methodological, and AI metrics were ex tracted from included SRs. AI metrics were categorized based on inputs, models, data training, and performance metrics. The quality of reporting of included SRs was evaluated using the AMSTAR 2. After screening 1,923 items from 5 databases were found, and 32 systematic reviews about the use of AI in oncology were included. AI was mainly utilized for diagnosis and prognosis. Most of the applications regarded nervous systems, gastro- intestinal, and prostate tumors. Images were the most common input of the AI models and were analyzed principally via support vector machine (SVM) and convolution neural network (CNN) models. Model performance was assessed using the area under the curve or via contingency tables. Indeed, inputs, models, and performance of AI were generally reported but training information was often lacking. The quality of reporting was generally unsatisfactory especially in the critical domains (44.2%) of the AMSTAR 2. The pre-sent overview highlights the increasing establishment of AI models for tumor diagnosis, particularly those utilizing images as input and employing algorithms like SVM and CNN. However, quality of the reporting must be improved to provide clear guidelines for the implementation of AI in oncological care.
VISMA: Vector Integration Site Mutation Analysis
*Francesco, Gazzo, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy; Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan Italy. Elena, Buscaroli, Department of Mathematics, Informatics and Geosciences, University of Trieste, Trieste, Italy. Yasmin Natalia, Serina Secanechia, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Pierangela, Gallina, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Laura, Rudilosso, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Giulio, Spinozzi, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Sara, Degl'Innocenti, GLP Test Facility, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Francesca, Sanvito, GLP Test Facility, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy; Pathology Unit, IRCCS San Raffaele Scientific Institute, Milan, Italy. Giulio, Caravagna, Department of Mathematics, Informatics and Geosciences, University of Trieste, Trieste, Italy. Marco, Masseroli, Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan Italy. Andrea, Calabria, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Daniela, Cesana, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Eugenio, Montini, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy.
VISMA: Vector Integration Site Mutation Analysis
*Francesco, Gazzo, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy; Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan Italy. Elena, Buscaroli, Department of Mathematics, Informatics and Geosciences, University of Trieste, Trieste, Italy. Yasmin Natalia, Serina Secanechia, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Pierangela, Gallina, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Laura, Rudilosso, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Giulio, Spinozzi, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Sara, Degl'Innocenti, GLP Test Facility, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Francesca, Sanvito, GLP Test Facility, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy; Pathology Unit, IRCCS San Raffaele Scientific Institute, Milan, Italy. Giulio, Caravagna, Department of Mathematics, Informatics and Geosciences, University of Trieste, Trieste, Italy. Marco, Masseroli, Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan Italy. Andrea, Calabria, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Daniela, Cesana, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy. Eugenio, Montini, San Raffaele Telethon Institute for Gene Therapy (SR-Tiget), IRCCS San Raffaele Scientific Institute, Milan, Italy.
- Keywords
- Integration site analysisSomatic variant callingLentiviral vector genotoxicityClonal dynamicsHematopoietic stem cell gene therapy
- Abstract
- Objective Hematopoietic stem cell gene therapy (HSC-GT) uses lentiviral vectors (LVs) to deliver therapeutic transgenes, but semi-random integration can perturb host genomic elements and promote clonal expansion. We developed VISMA, Vector Integration Site Mutation Analysis, a computational workflow to quantify somatic mutation burden in genomic regions flanking individual vector integration sites (IS), linking clonal dynamics with local mutagenesis in vivo. Methods VISMA extends VISPA2 IS calls by processing sequencing reads mapped around each IS to detect single-nucleotide variants (SNVs) and small insertions/deletions (indels). The pipeline removes optical duplicates, trims terminal nucleotides to reduce PCR and sequencing artifacts, assigns reads to their corresponding IS, and merges reads belonging to the same IS across cell lineages and longitudinal timepoints. Merged BAM files are converted to mpileup format for variant calling. A backtracing procedure reconstructs the temporal occurrence and variant allele frequency (VAF) of each mutation across samples. Variants with VAF >0.8 are removed as putative germline events, and variants in error-prone mono- or di-nucleotide stretches are filtered. Somatic burden is quantified through a Mutation Rate, normalized by covered bases, and a Mutation Index (MI), which further incorporates the number of IS and count-per-million abundance to provide an overall clone-level measure of mutational accumulation. Results We applied VISMA to a mouse HSC-GT model combining wild-type or Cdkn2a-/- HSCs with either genotoxic or GT-like non-genotoxic LVs, followed by longitudinal peripheral-blood IS profiling. Tumor incidence analysis showed that the genotoxic LV significantly accelerated tumor onset in Cdkn2a-/- recipients, log-rank p<0.0001. Overall, VISMA analyzed more than 200,000 unique IS and over 9 Gb of flanking sequence. Genotoxic LV groups showed significantly higher MI than non-genotoxic controls, Mann-Whitney p<0.001, with the strongest increase in Cdkn2a-/- mice receiving genotoxic LV, indicating synergy between vector genotoxicity and defective oncogene surveillance. Whole-genome sequencing confirmed higher mutation rates in genotoxic LV groups, supporting the absence of major bias in the flanking-region analysis. Conclusion VISMA provides a reproducible computational framework for quantifying somatic mutagenesis at the single-clone level from integration-site sequencing data. By integrating IS tracking, longitudinal variant backtracing, artifact-aware filtering and normalized mutation indices, VISMA enables systematic assessment of the relationship between LV integration, clonal selection and mutation dynamics. This approach can support bioinformatic evaluation of vector safety and genotoxic risk in HSC-GT.
Analytical variability across computational workflows in assessing tumor mutational burden and microsatellite instability from NGS data in male breast cancer
*Giovanni, Guglielmelli, Department of Molecular Medicine, Sapienza University of Rome Virginia,Valentini, Department of Molecular Medicine, Sapienza University of Rome Agostino, Bucalo, Department of Molecular Medicine, Sapienza University of Rome Virginia, Porzio, Department of Molecular Medicine, Sapienza University of Rome Annalisa, Platania, Department of Molecular Medicine, Sapienza University of Rome Davide, Alfieri, Department of Molecular Medicine, Sapienza University of Rome Ludovica, Perticone, Department of Molecular Medicine, Sapienza University of Rome Valentina, Silvestri, Department of Molecular Medicine, Sapienza University of Rome Laura, Ottini, Department of Molecular Medicine, Sapienza University of Rome
Analytical variability across computational workflows in assessing tumor mutational burden and microsatellite instability from NGS data in male breast cancer
*Giovanni, Guglielmelli, Department of Molecular Medicine, Sapienza University of Rome Virginia,Valentini, Department of Molecular Medicine, Sapienza University of Rome Agostino, Bucalo, Department of Molecular Medicine, Sapienza University of Rome Virginia, Porzio, Department of Molecular Medicine, Sapienza University of Rome Annalisa, Platania, Department of Molecular Medicine, Sapienza University of Rome Davide, Alfieri, Department of Molecular Medicine, Sapienza University of Rome Ludovica, Perticone, Department of Molecular Medicine, Sapienza University of Rome Valentina, Silvestri, Department of Molecular Medicine, Sapienza University of Rome Laura, Ottini, Department of Molecular Medicine, Sapienza University of Rome
- Keywords
- Male breast cancerTumor mutational burdenComputational workflows NGSMicrosatellite instability
- Abstract
- Objective Tumor mutational burden (TMB) and microsatellite instability (MSI) represent well-established biomarkers guiding the use of immune checkpoint inhibitors in solid tumors. However, standardized procedures for TMB and MSI detection from next-generation sequencing (NGS) data are still lacking. Currently, whole exome sequencing (WES) represents the gold standard for this analysis. On the other hand, TMB and MSI evaluation using gene panels is becoming increasingly widespread in clinical practice, but may be influenced by panel design and computational workflow. In this study, we assessed TMB and MSI in male breast cancer (MBC), a rare tumor in which these biomarkers remain poorly investigated, by comparing gene panel and WES data. Methods We analyzed 106 formalin-fixed paraffin-embedded MBC sequenced by Illumina TruSight Oncology 500 (TSO500) gene panel. Five cases had matched WES data for cross-platform comparison. TMB values generated by the TSO500 app were compared with those obtained using in-house pipelines applying different variant inclusion criteria, with attention to synonymous versus non-synonymous variants and coverage-based panel size correction. A cut-off of 10 mutations/Mb was used to define TMB-High status. MSI was assessed using the TSO500 app and the computational tools MSIsensor-pro and MSIsensor2, applying respective cut-offs of 10% and 20% for MSI-High classification. The distribution of mono-, di-, tri- and tetranucleotide unstable loci was also compared across tools and sequencing approaches. Results High TMB was identified in 11.3% of MBCs using the TSO500 app. In-house TMB estimates were strongly influenced by synonymous variant inclusion and coverage-based size correction, both altering absolute values and TMB-High/Low assignment. In the five matched WES/TSO500 cases, the definition based on non-synonymous variants without coverage-based correction showed the highest concordance. Using this optimized definition, 10.4% of MBCs were classified as TMB-High, with only partial concordance with the TSO500 app. A major methodological discrepancy emerged for MSI assessment. The TSO500 app classified 10.4% of cases as MSI-High, but none of these calls was confirmed by MSIsensor2. In contrast, MSIsensor-pro classified the majority of cases as MSI-High. Compared with WES, TSO500 showed an enrichment of trinucleotide sites detected by MSIsensor2, suggesting that MSI score distributions may be influenced by the algorithm used and by the microsatellite loci interrogated by the gene panel. Conclusion Our data indicate approximately 10% of MBCs may display TMB-High status, whereas MSI-High status is unlikely to be frequent. More importantly, MSI calls generated from NGS gene panels should be interpreted with caution. Our results emphasize the need for tailored approaches and harmonized computational pipelines to ensure reproducibility and reliable clinical interpretation of TMB and MSI values across platforms and datasets. Study supported by AIRC (IG28775) to LO GG VP AP are recipients of the Ph.D. programme of Molecular Medicine of Sapienza, University of Rome
Gene expression models for Alzheimer's Disease vs Mild Cognitive Impairment classification
Tommaso, Giorgini, Sapienza University of Rome. Eleonora, Grassucci, Sapienza University of Rome. Danilo, Comminiello, Sapienza University of Rome. Giulia, Fiscon, San Raffaele University of Rome.
Gene expression models for Alzheimer's Disease vs Mild Cognitive Impairment classification
Tommaso, Giorgini, Sapienza University of Rome. Eleonora, Grassucci, Sapienza University of Rome. Danilo, Comminiello, Sapienza University of Rome. Giulia, Fiscon, San Raffaele University of Rome.
- Keywords
- Alzheimer's diseaseMild Cognitive Impairmentblood transcriptomicsself-supervised learningtransformer modelsgene expression classification
- Abstract
- Objective. Blood transcriptomics offers a minimally invasive source of molecular information for Alzheimer's disease (AD), but distinguishing AD from Mild Cognitive Impairment (MCI) remains a difficult task. Unlike AD versus healthy-control classification, AD versus MCI involves clinically closer and biologically more heterogeneous conditions, where the discriminative signal is expected to be weaker. This work studies AD/MCI classification from gene expression data and investigates whether attention-based models and transcriptomic self-supervised pretraining can improve robustness in a high-dimensional, small-sample regime. Methods. We used a preprocessed, merged and batch-corrected expression matrix derived from the AddNeuroMed-related GEO cohorts GSE63060 and GSE63061, generated on Illumina HumanHT-12 platforms. The full reference matrix includes 711 samples and 19,460 genes; the AD/MCI task focuses on 473 subjects, including 284 AD and 189 MCI samples. We evaluated two families of supervised inputs: a biologically informed set of 788 differentially expressed genes from an AD versus MCI contrast, and top-variance gene subsets selected within the training fold. Transformer-style models for gene expression, including T-GEM and TxT, were compared with classical baselines such as logistic regression, support vector machines, random forest, k-nearest neighbours and Gaussian naive Bayes. Evaluation used repeated stratified train/validation/test splits and macro-F1, with attention to performance stability across seeds. In parallel, we developed a masked gene-expression pretraining pipeline on an expanded GEO transcriptomic corpus of 3,672 samples and 9,748 common genes, designed to initialize TxT before supervised fine-tuning. Results. Preliminary supervised experiments show that DEG-based inputs are more reliable than larger variance-based gene sets in this setting. T-GEM and TxT achieved competitive validation macro-F1 values, around 0.73 on DEG inputs, but their test performance was more variable across splits. Classical baselines, particularly RBF SVM, remained highly competitive, reaching an average test macro-F1 around 0.68 and showing lower variance. Increasing the number of input genes did not consistently improve generalization, suggesting that dimensionality and attention memory cost remain important limitations for small-cohort transcriptomic learning. Conclusion. These results indicate that AD versus MCI classification from blood gene expression remains challenging despite careful feature selection and modern neural architectures. Transformer-based models are promising but require strategies to reduce split sensitivity and overfitting. Self-supervised transcriptomic pretraining is therefore a central direction for learning more stable gene-expression representations and improving downstream AD/MCI classification.
geneslator: a R package for comprehensive gene identifier mapping and annotation
*Grete Francesca Privitera, Department of Clinical and Experimental Medicine, Bioinformatics Unit, University of Catania, Catania, Italy Giulia Cavallaro, Istituto Oncologico del Mediterraneo (IOM), Via Penninazzo 7, 95029 Viagrande, Italy Giovanni Micale, Department of Clinical and Experimental Medicine, Bioinformatics Unit, University of Catania, Catania, Italy Stefano Forte, Istituto Oncologico del Mediterraneo (IOM), Via Penninazzo 7, 95029 Viagrande, Italy Alfredo Pulvirenti, Department of Clinical and Experimental Medicine, Bioinformatics Unit, University of Catania, Catania, Italy Salvatore Alaimo, Department of Clinical and Experimental Medicine, Bioinformatics Unit, University of Catania, Catania, Italy
geneslator: a R package for comprehensive gene identifier mapping and annotation
*Grete Francesca Privitera, Department of Clinical and Experimental Medicine, Bioinformatics Unit, University of Catania, Catania, Italy Giulia Cavallaro, Istituto Oncologico del Mediterraneo (IOM), Via Penninazzo 7, 95029 Viagrande, Italy Giovanni Micale, Department of Clinical and Experimental Medicine, Bioinformatics Unit, University of Catania, Catania, Italy Stefano Forte, Istituto Oncologico del Mediterraneo (IOM), Via Penninazzo 7, 95029 Viagrande, Italy Alfredo Pulvirenti, Department of Clinical and Experimental Medicine, Bioinformatics Unit, University of Catania, Catania, Italy Salvatore Alaimo, Department of Clinical and Experimental Medicine, Bioinformatics Unit, University of Catania, Catania, Italy
- Keywords
- Data integrationGeneID conversionOrthologs mappingPathways mapping
- Abstract
- Objective: High-throughput sequencing generates large gene lists, making data interpretation challenging. Accurate gene annotation and reliable conversion between identifiers (e.g., Gene Symbol, Ensembl GeneIDs, Entrez GeneIDs) are essential for integrating datasets, functional analyses, and cross-species comparisons. Existing tools and databases facilitate annotation but often suffer from inconsistencies, missing mappings, and fragmented workflows, limiting reproducibility and interpretability. Methods: To address these limitations, we developed geneslator. Its annotation databases have been built for 8 model organisms. These databases integrate data collected from several sources: (i) information about genes (symbol, aliases, full name, and genetype) were extracted from NCBI Gene and Ensembl; (ii) Gene identifiers were taken from NCBI and Ensembl; (iii) protein IDs were taken from Uniprot; (iv) specific identifiers were obtained from the most popular species-specific genome database, such as HGNC for Human, MGI for Mouse, RGD for Rat, SGD for Yeast, WormBase for Worm, FlyBase for Fly, ZFIN for Zebrafish and TAIR for Arabidopsis. Our databases integrate old discontinued and replaced gene identifiers from NCBI gene and Ensembl. Gene orthologs have been taken from NCBI, Ensembl, AllianceGenome and HCOP. Functional annotation data include pathways collected from KEGG, Reactome and Wikipathways and gene ontologies taken from GO. Finally, annotation databases have been built as SQLite objects using the AnnotationForge R package. The databases are updated on a monthly basis. Results: geneslator, is a R package that unifies gene identifier conversion, orthologs mapping and pathway annotation across eight model organisms (H.sapiens, M.musculus, R.norvegicus, D.melanogaster, D.rerio, S.cerevisiae, C.elegans, A.thaliana). It provides an up-to-date, precise, and coherent framework preserving data integrity, enabling cross-species analyses and facilitating robust interpretation of gene function and regulation. Experiments show that geneslator is able to uniquely map a higher number of gene identifiers than the existing annotation tools with a very low rate of unmapped identifiers. In H.sapiens, Ensembl GeneID to Gene symbol conversion achieved a one-to-one mapping rate of 98.92%, outperforming biomaRt (73.38%), org.Hs.eg.db (61.70%), mygene (76.82%) and gprofiler2 (72.9%). Notably, the proportion of unmapped identifiers for geneslator was limited to 0.05%, whereas other tools exhibited missing rates ranging from 23.14% to 37.29%. Fisher’s exact test confirmed that the differences in mapping performance between geneslator and other packages were statistically significant (p<0.001). Moreover, the mapping from Gene symbols to Entrez GeneID increased gene coverage in the 17.09% of pathways (p < 0.001), with an average gain of 4.13 genes per pathway. Conclusion: Experimental results are coherent in all the tested organisms and show tangible consequences for downstream analyses. Even minor changes in annotation completeness could strongly affect biological conclusions. Finally, by minimizing the information loss at the annotation stage, geneslator preserves data complexity and clearly improves the robustness and reproducibility of downstream analyses. geneslator is available at https://github.com/knowmics-lab/geneslator.
MErlin – Methylation-driven Expression & Regulation Linkage in Interacting Nuclear-domains
Iacopo Passeri*, Università degli Studi di Firenze Solène Pety, Université Paris-Saclay (INRAE) Alessio Mengoni, Università degli Studi di Firenze Elena Perrin, Università degli Studi di Firenze
MErlin – Methylation-driven Expression & Regulation Linkage in Interacting Nuclear-domains
Iacopo Passeri*, Università degli Studi di Firenze Solène Pety, Université Paris-Saclay (INRAE) Alessio Mengoni, Università degli Studi di Firenze Elena Perrin, Università degli Studi di Firenze
- Keywords
- EpigeneticsMultiomicsMLDGEA
- Abstract
- Objective The role of epigenetic markers (i.e., m4C, m5C, hm5C, and m6A) in regulating gene expression and chromatin structure remains one of the major open problems in bacterial genomics. Although these omics have been individually examined, their combined effect remains poorly understood, especially in bacterial genomes with multipartite organization. In this work, we introduce MErlin, a tool for uncovering relationships between methylation, gene expression, and 3D spatial organization of the chromosomes, aiming at providing a reproducible tool for harmonizing and integrating multi-omics data, allowing biological insights through machine learning modeling. Methods MErlin is implemented as a user-friendly modular Snakemake pipeline that processes three primary data types: methylation basecalling (ONT or PacBio), RNA-seq quantification, and Hi-C contact matrices. Input datasets are standardized and aggregated at consistent genomic resolutions (e.g., gene-level or fixed bins), enabling direct comparison across omics layers. The pipeline performs feature engineering to generate biologically meaningful predictors, including normalized methylation counts, gene expression levels, and spatial interaction frequencies. A supervised machine learning model is then trained to predict expression patterns or regulatory states from combined features. To ensure interpretability, MErlin integrates a SHAP (SHapley Additive exPlanations) framework, allowing quantification of the contribution of each omics feature to model predictions across different experimental conditions. Report statistics, tables and plots are stored as output. Results Application of MErlin to bacterial multipartite genomes revealed biologically interesting associations between methylation, transcriptional activity, and chromosomal architecture. Preliminary analyses demonstrate that specific genomic regions exhibiting high spatial contact frequencies also show coordinated methylation patterns and gene expression levels. SHAP-based interpretations suggest that methylation signals in specific DNA motifs and coding regions contribute to predicting transcriptional output, while long-range chromosomal interactions provide additional explanatory power. In fact, the integration of Hi-C data improves model performance, highlighting the importance of 3D genome organization. These findings support the hypothesis that epigenetic regulation is influenced not only by local sequence context but also by higher-order chromosomal structure. Conclusion MErlin provides a complete and understandable framework for data integration, allowing for a deeper analysis of intricate relationships among DNA methylation, gene expression, and genome structure. This has been achieved through a combination of a powerful Snakemake workflow and SHAP-based explainability, thus providing a link between data integration and biological interpretability. The utility and importance of MErlin are demonstrated through its application to bacterial genomes, especially those with a multipartite structure, in which spatial organization is a critical aspect, and new regulatory processes can be revealed.
The voice as a diagnostic tool: Machine Learning-based speech model in discriminating patients with acute mental disorders
Alessandro, Comandatore, Sapienza University of Rome* Isabella, Getuli, Department of Human Neurosciences Policlinico Umberto I Sapienza University of Rome Riccardo, Serra, Department of Human Neurosciences Policlinico Umberto I Sapienza University of Rome Lorenzo, Tarsitani, Department of Human Neurosciences Policlinico Umberto I Sapienza University of Rome
The voice as a diagnostic tool: Machine Learning-based speech model in discriminating patients with acute mental disorders
Alessandro, Comandatore, Sapienza University of Rome* Isabella, Getuli, Department of Human Neurosciences Policlinico Umberto I Sapienza University of Rome Riccardo, Serra, Department of Human Neurosciences Policlinico Umberto I Sapienza University of Rome Lorenzo, Tarsitani, Department of Human Neurosciences Policlinico Umberto I Sapienza University of Rome
- Keywords
- Machine LearningVoice AnalysisPsychiatryAcute Mental DisordersDifferential diagnosis
- Abstract
- OBJECTIVES This study aims to analyse vocal speech features to identify mental disorders in the acute phase, using a Machine Learning (ML) model. Specific diagnostic categories were classified and compared pairwise and against non-clinical subjects. METHODS Patients with acute phase psychiatric disorders — Bipolar Disorder (BD), Depressive/Personality disorders (DP), or Schizophrenia/other psychotic disorders (SKZ) — admitted to the Psychiatric Intensive Care Unit of Policlinico Umberto I Hospital in Rome were consecutively enrolled. Non-Clinical Subjects (NCS) were recruited among university students, residents, and nurses. To date, 26 psychiatric patients (5 BD, 12 DP, 9 SKZ) and 11 NCS have been enrolled. Each participant was recorded performing two vocal tasks: a free speech task (responses to neutral examiner-administered prompts) and a picture description task (description of the Cookie Theft Picture). Vocal features were extracted via a signal-processing pipeline and grouped into seven macro-categories: Speech Tempo, Speech Pause, Prosodic Intonation, Prosodic Stress, Speech Spectrum, Vocal Quality, and Articulation. Data from standalone and combined tasks were used to train a Support Vector Machine (SVM) binary classifier. Model performance (AUC) was assessed primarily through leave-one-subject-out cross-validation and, in selected comparisons, five-fold cross-validation. RESULTS Key preliminary findings are as follows. The combined vocal task achieved good discriminatory performance between the patient and NCS groups (AUC = 0.86). Pairwise diagnostic comparisons revealed the strongest performance for BD vs. DP, particularly using the free speech task alone (AUC = 0.90). BD vs. SKZ showed moderate discriminability across both tasks (AUC = 0.71). SKZ vs. DP yielded the weakest performance under leave-one-subject-out cross-validation (AUC = 0.59); however, five-fold cross-validation improved this result (AUC = 0.72). Overall, discriminatory ability of the model across all comparisons ranged from moderate to excellent (AUC = 0.59-0.90) representing a promising preliminary result given the small sample size. CONCLUSIONS These preliminary findings suggest that BD, DP, and SKZ exhibit distinct vocal profiles and that an ML-based speech model may serve as a discriminative tool across psychiatric diagnoses in the acute phase. Validation through larger samples remains essential to clarify the relationships between vocal features, psychopathological dimensions, and pharmacological treatment. Future directions include feature importance analysis to identify the most diagnostically relevant vocal markers and finer stratification of diagnostic subgroups (e.g., SKZ positive vs. negative symptoms) to optimize model performance. Speech analysis is not intended to replace clinical expertise, but rather to provide an objective, reproducible tool capable of reducing subjectivity and rater-dependent variability in psychiatric assessment.
Beyond the Diagnosis: Semi-Automated Risk Analysis of Inappropriate Antibiotic Prescribing in Mexico and the Potential for Real-Time Alerts
María Cecilia Ishida Gutiérrez, UACH Kevin, Johnson Molina, UACH Josué Eleazar, Batres Martínez, UACH* Vianey Iraís, Jiménez Ramírez, UACH René, Núñez Bautista, UACH Roxana, Chávez Ávila, UTEP*
Beyond the Diagnosis: Semi-Automated Risk Analysis of Inappropriate Antibiotic Prescribing in Mexico and the Potential for Real-Time Alerts
María Cecilia Ishida Gutiérrez, UACH Kevin, Johnson Molina, UACH Josué Eleazar, Batres Martínez, UACH* Vianey Iraís, Jiménez Ramírez, UACH René, Núñez Bautista, UACH Roxana, Chávez Ávila, UTEP*
- Keywords
- Medical InformaticsDrug UtilizationPractice GuidelinesClinical Decision Support Systems
- Abstract
- Objective: Semi-automatically analyze 1-year records from a government health database to identify prescriptions for upper respiratory infections (URIs) and quantify the risk of inappropriate antibiotic prescribing for diagnoses that do not clinically require this therapy, according to international clinical guidelines. Methods: From a database containing 2,224,741 prescriptions, 100,437 diagnoses of upper respiratory infections (URIs) were automatically extracted and classified as requiring or not requiring antibiotics according to clinical practice guidelines. A pipeline was designed to analyze prescriptions using Python and Pandas libraries automatically: Diagnoses of upper respiratory infections were extracted using ICD-10 codes, and a library of systemic antibiotics was compiled from all drugs in the dataset. To determine whether antibiotic use was justified, medical experts classified antibiotic use according to international guidelines for URIs. A secondary automatic analysis identified the most frequently inappropriately prescribed antibiotics. Results: Antibiotics were prescribed without guideline indication in 12.92% of cases reviewed. Furthermore, in cases where antibiotic therapy was clinically indicated, the antibiotic selected was the recommended active substance by the corresponding clinical guideline in only 1.5% of the cases. The patients with diagnoses that do not require antibiotic therapy are 9.5 times more likely to receive an antibiotic prescription than those with diagnoses that do (IC 95%: 7.54–12.02). Stratified analysis revealed a statistically significant association between specific clinical diagnoses and inappropriate antibiotic prescription: acute pharyngitis due to other specified organisms (OR=22.23, IC 95%: 17.24–28.68) and acute tonsillitis (OR of 22.21, IC 95%: 17.50–28.19) stand out due to their high magnitude of association. Other diagnoses, such as influenza with ICD-10 codes J110 and J118, showed significant associations, though with wider confidence intervals (OR of 16.04, IC 95%: 1.77-145 and OR of 12.83, IC 95%: 1.48-111.22, respectively). Conclusion: There is strong statistical evidence of antibiotic overprescription in the years of study. The risk of receiving an unnecessary antibiotic prescription is between 3.8 and 22.2 times higher in patients with viral or non-bacterial diagnoses compared to correct clinical practice. The semiautomatic analysis can be upscaled and adapted to analyze other large health data databases and to generate real-time alerts focused on prescriptions with a high risk of inappropriate antibiotic use before a prescription is issued to the patient.
An LLM-RAG System for Interactive Multi-Omic Exploration of 3D Genome Organization and Gene Expression during Mouse Cortical Neurogenesis
Marco, Cosulich, Dipartimento di informatica, bioingegneria, robotica e ingegneria dei sistemi - DIBRIS, Genoa, Italy. *Francesca Anna, Cupaioli, National Research Council of Italy - Institute for Biomedical Technologies, Segrate, Italy. *Ivan, Merelli, National Research Council of Italy - Institute for Biomedical Technologies, Segrate, Italy. Daniele, D’Agostino, Dipartimento di informatica, bioingegneria, robotica e ingegneria dei sistemi - DIBRIS, Genoa, Italy.
An LLM-RAG System for Interactive Multi-Omic Exploration of 3D Genome Organization and Gene Expression during Mouse Cortical Neurogenesis
Marco, Cosulich, Dipartimento di informatica, bioingegneria, robotica e ingegneria dei sistemi - DIBRIS, Genoa, Italy. *Francesca Anna, Cupaioli, National Research Council of Italy - Institute for Biomedical Technologies, Segrate, Italy. *Ivan, Merelli, National Research Council of Italy - Institute for Biomedical Technologies, Segrate, Italy. Daniele, D’Agostino, Dipartimento di informatica, bioingegneria, robotica e ingegneria dei sistemi - DIBRIS, Genoa, Italy.
- Keywords
- Hi-CLLMRAGgenomic data integration
- Abstract
- Objective. High-throughput multi-omics technologies have produced datasets of unprecedented scale, but pulling them together and exploring them interactively remains a bottleneck. Combining Hi-C with RNA-seq is particularly informative, revealing directly how three-dimensional (3D) genome organisation and transcriptional activity change together across cell types and developmental stages. Standard bioinformatic pipelines were not designed for data of this size and heterogeneity, and tend to fall short when several layers must be cross-examined at once. The problem is especially acute in the developing nervous system, where gene regulatory programs act on structural, transcriptional and temporal scales simultaneously. Methods. We present GrapHiC, a Retrieval-Augmented Generation (RAG) system that couples large language models (LLMs) with a graph-structured Neo4j database to support natural language querying of multi-omics data from Hi-C and RNA-seq experiments. The RAG design addresses the two best-known weaknesses of standalone LLMs, namely the static knowledge cutoff and the tendency to produce plausible but inaccurate answers, by anchoring every response in evidence retrieved from a domain-specific database rather than in the model’s parametric memory. GrapHiC was applied to dataset GSE96107, which provides high-resolution Hi-C maps and matched RNA-seq profiles from mouse cortical neurogenesis. Hi-C contacts from embryonic stem cells (ESC) and neural progenitor cells (NPC) were represented as a weighted undirected graph containing approximately 20,000 gene nodes connected by around 450,000 chromatin contact edges. RNA-seq expression profiles were attached to the same nodes as additional attributes, so that transcriptional and structural information could be queried together. The pipeline is organised in three sequential modules. A Rephraser first breaks down complex natural language questions into simpler atomic sub-queries. A Parser then maps each sub-question onto one of eight predefined graph query functions and produces a structured JSON representation. Finally, a DB Interviewer executes the corresponding Cypher query against Neo4j and returns an answer in natural language, grounded in the retrieved evidence. When a question does not match any predefined function, the system generates Cypher code directly from the database schema. Results. On a benchmark of 78 graph-oriented questions involving connectivity analysis and graph metrics, GrapHiC reached 96% accuracy with an average response time of 30 seconds. Integrating RNA-seq enabled biologically meaningful multi-layer analyses that neither dataset would have supported alone. In particular, we could identify genes undergoing coordinated structural and transcriptional remodelling during the ESC-to-NPC transition, and detect structurally active but transcriptionally silent loci, a class of regions invisible when Hi-C or RNA-seq are analysed in isolation. Conclusion. GrapHiC provides a practical proof-of-concept that LLM-RAG systems can serve as an effective interface to complex multi-omics datasets, making large structural and transcriptional resources easier to query in a hypothesis-driven way. The framework is particularly well suited to investigating the regulatory mechanisms underlying neurogenesis and neurodevelopmental disease.
Advancing Facial Recognition for Next-Generation Phenotyping in Africa
*Japhet Dienda, Center for Human Genetics, Faculty of Medicine, University of Kinshasa, Kinshasa, DR Congo; African Rare Diseases Initiative, Reference Center for Rare and Undiagnosed Diseases, University of Kinshasa, Kinshasa, DR Congo; Computational Biology Division, IDM, Faculty of Health Sciences, University of Cape Town, Cape Town, South Africa Nicola Mulder, Computational Biology Division, IDM, Faculty of Health Sciences, University of Cape Town, Cape Town, South Africa Christian D. Bope, African Rare Diseases Initiative, Reference Center for Rare and Undiagnosed Diseases, University of Kinshasa, Kinshasa, DR Congo; Department of Mathematics and Computer Sciences, Faculty of Sciences, University of Kinshasa, DR Congo Hocine Bendou, Computational Biology Division, IDM, Faculty of Health Sciences, University of Cape Town, Cape Town, South Africa Aimé Lumaka, Center for Human Genetics, Faculty of Medicine, University of Kinshasa, Kinshasa, DR Congo; African Rare Diseases Initiative, Reference Center for Rare and Undiagnosed Diseases, University of Kinshasa, Kinshasa, DR Congo; Department of Paediatrics, Faculty of Medicine, University of Kinshasa, Kinshasa, DR Congo
Advancing Facial Recognition for Next-Generation Phenotyping in Africa
*Japhet Dienda, Center for Human Genetics, Faculty of Medicine, University of Kinshasa, Kinshasa, DR Congo; African Rare Diseases Initiative, Reference Center for Rare and Undiagnosed Diseases, University of Kinshasa, Kinshasa, DR Congo; Computational Biology Division, IDM, Faculty of Health Sciences, University of Cape Town, Cape Town, South Africa Nicola Mulder, Computational Biology Division, IDM, Faculty of Health Sciences, University of Cape Town, Cape Town, South Africa Christian D. Bope, African Rare Diseases Initiative, Reference Center for Rare and Undiagnosed Diseases, University of Kinshasa, Kinshasa, DR Congo; Department of Mathematics and Computer Sciences, Faculty of Sciences, University of Kinshasa, DR Congo Hocine Bendou, Computational Biology Division, IDM, Faculty of Health Sciences, University of Cape Town, Cape Town, South Africa Aimé Lumaka, Center for Human Genetics, Faculty of Medicine, University of Kinshasa, Kinshasa, DR Congo; African Rare Diseases Initiative, Reference Center for Rare and Undiagnosed Diseases, University of Kinshasa, Kinshasa, DR Congo; Department of Paediatrics, Faculty of Medicine, University of Kinshasa, Kinshasa, DR Congo
- Keywords
- Down syndrome; facial phenotyping; deep learning; convolutional neural network; ancestry-aware modelling; population diversity; sub-Saharan Africa; Grad-CAM; computer-assisted diagnosis; genetic syndromes
- Abstract
- Objective: Artificial intelligence (AI)-based facial phenotyping tools have shown strong potential for supporting the diagnosis of genetic syndromes such as Down syndrome (Hsieh et al., 2022; Gurovich et al., 2019). However, most existing systems are trained predominantly on European populations, limiting their performance in underrepresented groups, particularly African populations (Lumaka et al., 2017; Lesmann et al., 2024). Because facial morphology is strongly influenced by ancestry, ancestry-aware approaches are needed to reduce diagnostic bias and improve equitable genetic diagnosis. This study aimed to develop and evaluate an ancestry-aware and explainable deep learning framework for Down syndrome facial phenotyping using two-dimensional facial images from African and comparative global cohorts. Methods: Facial images of individuals with Down syndrome and controls were collected from cohorts in the Democratic Republic of Congo, Rwanda, Guadeloupe, and a public dataset, resulting in 6,956 images. Images were standardised through face alignment, resizing, normalisation, and augmentation (Soekarta & Ku-Mahamud, 2025). Multiple deep learning architectures, including ResNet50, DenseNet121, MobileNetV2, VGG16, ConvNeXt-Tiny, and a Vision Transformer, were trained using transfer learning (Cai et al., 2020). Model performance was evaluated using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUROC) (Hicks et al., 2022). Explainability was assessed using Grad-CAM visualisations to identify facial regions influencing predictions (Selvaraju et al., 2020), while cohort-level analyses examined robustness and ancestry-related bias. Results: Cohort-aware training improved model robustness and reduced performance variability across populations. ConvNeXt-Tiny achieved the best overall performance, reaching 92.5% accuracy and an AUROC of 0.979, while MobileNetV2 achieved the highest recall (0.971), reducing the risk of missed diagnoses. These results are comparable to previously reported AI-based facial phenotyping systems (Gurovich et al., 2019; Reiter et al., 2024). Strong classification performance was observed across the African cohorts from the Democratic Republic of Congo and Rwanda, as well as the Guadeloupe cohort. Mean facial representations highlighted the continued underrepresentation of African populations in public facial datasets (Lesmann et al., 2024). Grad-CAM analyses consistently showed attention to clinically meaningful facial regions associated with Down syndrome, particularly the periocular region and nasal bridge (Kruszka et al., 2017). No major systematic bias disadvantaging any cohort was observed. Conclusion: Ancestry-aware and explainable AI can achieve strong and equitable facial phenotyping performance across diverse populations, including underrepresented African cohorts. Incorporating cohort-aware training and inclusive datasets improves diagnostic reliability while reducing ancestry-related bias, supporting the future development of fair and clinically trustworthy AI tools for genetic diagnosis in Africa and other diverse populations.
Mixture-based Nonparametric Estimation of Spatial Covariance Functions with Applications to HIV Key Population Size Estimation across Sub-Saharan Africa
*Manushi, Siriwardana, Department of Statistics, Pennsylvania State University, University Park, PA, USA Hyebin, Song, Department of Statistics, Pennsylvania State University, University Park, PA, USA Le, Bao, Department of Statistics, Pennsylvania State University, University Park, PA, USA Stephen, Berg, Department of Statistics, Pennsylvania State University, University Park, PA, USA
Mixture-based Nonparametric Estimation of Spatial Covariance Functions with Applications to HIV Key Population Size Estimation across Sub-Saharan Africa
*Manushi, Siriwardana, Department of Statistics, Pennsylvania State University, University Park, PA, USA Hyebin, Song, Department of Statistics, Pennsylvania State University, University Park, PA, USA Le, Bao, Department of Statistics, Pennsylvania State University, University Park, PA, USA Stephen, Berg, Department of Statistics, Pennsylvania State University, University Park, PA, USA
- Keywords
- Key Population Size EstimationSpatial Covariance EstimationStationary Isotropic ProcessesGaussian Scale MixturesNonparametric Maximum Likelihood
- Abstract
- Objective Consistent data on the sizes of key populations, such as female sex workers (FSWs), are often scarce, particularly at the sub-national level. Accurate size estimates are critical to effectively allocate resources and achieve HIV targets. Since FSW population sizes may be spatially correlated across areas, models that account for spatial dependence can improve estimation. An important component of such models is the covariance function, which characterizes the spatial dependence structure of the underlying process. In this work, we study spatial covariance functions to estimate FSW population sizes in Sub-Saharan Africa (SSA). Methods Many spatial models rely on parametric covariance functions. However, parametric estimation can suffer from model misspecification, potentially leading to inefficient or biased predictions. We therefore develop a robust nonparametric approach for estimating the covariance function of a stationary isotropic process in R^d. We focus on a class of covariance functions that are valid in all dimensions, which includes popular kernels such as the exponential and Matern kernels. Leveraging the fact that such covariance functions can be represented as infinite mixtures of scaled Gaussian kernels, we propose two estimation methods: weighted least squares and nonparametric maximum likelihood (NPML) estimation to estimate the mixing measure of scaled Gaussian kernels. We also develop computationally efficient methods to solve these optimization problems using non-negative least squares and second-order descent updates. We evaluate the proposed methods through simulations and apply them to estimate the FSW population sizes at the sub-national level in SSA. Results We evaluate the estimation and predictive performance of the proposed nonparametric method using integrated squared error (ISE) and the mean squared prediction error (MSPE). Results from simulation studies show that the proposed method performs well across varying simulation models, in cases with and without nugget effects, and as the strength of the nugget effect increases. We also examine the approximation error in using finite approximations of infinite mixtures to estimate the covariance function relative to the true covariances and the approximation appears to estimate the covariance function well except for very coarse grids of the mixing measure. Furthermore, the proposed method yields the lowest MSPE, outperforming the conditional autoregressive integrated nested Laplace approximation (CAR-INLA) model and other parametric and nonparametric covariance estimation methods in estimating FSW populations sizes in SSA. Conclusion The proposed mixture-based nonparametric method based on NPML estimation accounts for local and long-range spatial relationships in the data while incorporating covariate information. Thus, the method provides a useful approach for producing FSW population size estimates (PSEs) in SSA. The method can also be applied to produce PSEs of other key populations in the future.
Torch-eCpG v2: A Scalable and Interpretable Framework for eQTM Mapping and Multi-Omic Network Analysis
*Kord, Kober, Helen Diller Family Comprehensive Cancer Center, University of California San Francisco, United States Anika, Rau, Helen Diller Family Comprehensive Cancer Center, University of California San Francisco, United States Adam, Olshen, Helen Diller Family Comprehensive Cancer Center, University of California San Francisco, United States
Torch-eCpG v2: A Scalable and Interpretable Framework for eQTM Mapping and Multi-Omic Network Analysis
*Kord, Kober, Helen Diller Family Comprehensive Cancer Center, University of California San Francisco, United States Anika, Rau, Helen Diller Family Comprehensive Cancer Center, University of California San Francisco, United States Adam, Olshen, Helen Diller Family Comprehensive Cancer Center, University of California San Francisco, United States
- Keywords
- Torch-eCpGExpression Quantitative Trait Methylation (eQTM)DNA MethylationGene ExpressionGPU Acceleration
- Abstract
- Objective Understanding the functional contribution of epigenetic variation to gene expression remains a central challenge in genomics. This is commonly addressed by identifying expression quantitative trait methylation (eQTM) loci, which capture associations between DNA methylation and transcriptional activity. Interest in integrating these data modalities is growing, yet scalable and interpretable tools for systematic eQTM analysis remain limited. Torch-eCpG was developed as the first freely available, open-source framework for large-scale multi-omic eQTM mapping. It has been successfully deployed in clinical studies, enabling the discovery of eQTM signatures linked to environmental exposures and cancer fatigue phenotypes, demonstrating its utility for uncovering clinically relevant epigenetic signals. Here, we present Torch-eCpG v2, which introduces key improvements in statistical modeling, empirical validation, interpretability, and hardware-aware scalability. Methods Torch-eCpG v2 implements a highly optimized least-squares multivariate linear regression framework for efficiently computing eQTM associations across high-dimensional datasets. Statistical confidence is strengthened through configurable bootstrap resampling, enabling empirical estimation of robustness for top-ranked associations. Model interpretability is enhanced with Explainable AI (XAI), using Integrated Gradients to quantify the contribution of individual covariates to gene expression predictions. To support large-scale data processing, the framework incorporates native support for compressed columnar storage formats (Parquet) and a host-aware configuration system that automatically adapts chunk sizes, threading, and execution parameters to available CPU or GPU resources. A bisection-based auto-chunking strategy, combined with dynamic memory estimation, prevents out-of-memory failures and ensures stable execution across heterogeneous environments. Additionally, a fully asynchronous compute pipeline, including chunk prefetching with parallelized I/O, overlaps data movement and computation to maximize throughput. Results The updated framework enables efficient and scalable identification of eQTM associations in large, high-dimensional cohorts. Bootstrap-derived confidence estimates provide a principled basis for prioritizing biologically meaningful associations, while XAI-based attribution improves interpretability by identifying key regulatory drivers within multivariate models. For downstream characterization, Torch-eCpG v2 includes an expanded network analysis and visualization suite. This includes Circos plots for global mapping of eQTM architectures and bipartite network modeling to identify co-regulatory modules. Extended graph analytics support projection to gene- or CpG-level networks and multiple edge-weighting strategies, enabling interrogation of epigenetic regulatory structure. Outputs are exported in standardized formats compatible with external tools (e.g., Cytoscape), facilitating downstream analysis. Performance profiling demonstrates stable memory utilization and computational throughput across hardware configurations ranging from consumer-grade systems to high-performance computing environments. Conclusion Torch-eCpG v2 provides a scalable, interpretable, and extensible framework for integrative eQTM analysis. By combining efficient regression modeling, empirical validation, explainable AI, advanced network analysis, and systems-level performance optimization, the platform enables more rigorous and biologically meaningful exploration of epigenetic regulatory mechanisms. These advances support the analysis of increasingly large and complex multi-omic datasets, facilitating the characterization of gene regulation across diverse biological contexts and population-scale studies.
Autoresearch Discovery of Interpretable Filter Rules for Antibody Binder Classification
Mikel Landajuela*, Lawrence Livermore National Laboratory, USA
Autoresearch Discovery of Interpretable Filter Rules for Antibody Binder Classification
Mikel Landajuela*, Lawrence Livermore National Laboratory, USA
- Keywords
- antibody designbinder classificationautoresearchinterpretable filtersAlphaFold3
- Abstract
- **Objective**: Antibody design campaigns increasingly generate many candidates before only a small subset can be tested experimentally, making candidate filtering a central bottleneck. We investigate whether an autoresearch loop can discover better training-free filters for antibody binder classification by systematically exploring rule-based variants and using experimental results to guide iterative refinement. **Methods**: We implemented an autoresearch loop that iteratively proposes filter rule variants, evaluates them under a fixed Leave-One-System-Out (LOSO) cross-validation protocol across seven antibody-antigen systems from the IgDesign benchmark, records each experiment in version control, and uses the accumulated results to inform the next iteration. The system explores combinations of AlphaFold3 confidence metrics (ipTM, pTM, PAE) and structural metrics (RMSD) as training-free classification rules. Each filter variant is committed to version control with descriptive metadata, enabling systematic tracking of the exploration trajectory. The LOSO protocol tests generalization to unseen antibody-antigen pairs by training on six systems and evaluating on one held-out system, rotating through all seven systems. We compare the discovered filters against supervised machine learning baselines (logistic regression, feature-selected balanced logistic regression) and prompted large language model baselines (GPT-4o and GPT-5 tabular few-shot prompting) evaluated on identical data splits. **Results**: Across 75 unique logged filter variants, the autoresearch loop improved average ROC-AUC from 0.6371 for the initial baseline to 0.8060 for the final discovered rule, an absolute gain of 0.1689 and a relative improvement of 26.5%. This final filter, which we call the RMSD-Tuned Triad rule, uses only three structural-confidence signals in an interpretable Boolean combination. The discovered filter exceeds logistic regression (0.7144 ROC-AUC), feature-selected balanced logistic regression (0.7536), and GPT-4o tabular few-shot prompting (0.7640). It approaches within 0.0044 ROC-AUC of the strongest GPT-5 tabular few-shot baseline (0.8104), while requiring no prompted examples and no LLM inference once numeric features are computed. **Conclusion**: Systematic autoresearch can turn simple structural-confidence signals into compact, interpretable filters competitive with supervised learning and prompted LLMs for antibody binder classification. The discovered training-free rule is particularly valuable when target-specific labeled data are scarce, as it generalizes across diverse antibody-antigen systems without requiring system-specific training. These results demonstrate that automated scientific exploration guided by version-controlled experiments can discover effective solutions in computational biology applications.
QFARM - hierarchical quantitative association rule mining
*Ledio, Deda, Biolabs, JetBrains Research, Munich, Germany; TUM School of Computation, Information and Technology (CIT), Technical University of Munich, Munich, Germany Konstantin, Zaitsev, Biolabs, JetBrains Research, Berlin, Germany
QFARM - hierarchical quantitative association rule mining
*Ledio, Deda, Biolabs, JetBrains Research, Munich, Germany; TUM School of Computation, Information and Technology (CIT), Technical University of Munich, Munich, Germany Konstantin, Zaitsev, Biolabs, JetBrains Research, Berlin, Germany
- Keywords
- quantitative association rule mininghierarchical rule miningmulti-objective evolutionary optimizationhigh-dimensional biomedical data analysiscomplex pattern recognition
- Abstract
- Objective Association rule mining (ARM) is a widely used approach for discovering interpretable relationships in data, especially in biomedical applications. Classical ARM methods are well suited for categorical data but are not directly applicable to numerical data. Existing quantitative ARM (QARM) methods offer limited support for analyzing the impact of added attributes on rule performance. Discretization-based QARM approaches may lose information and increase dimensionality, whereas evolutionary methods often overuse highly predictive attributes and underexplore large search spaces. Inspired by hierarchical ARM (FARM [1]) and evolutionary QARM methods, we present QFARM (Quantitative FARM), which combines hierarchical rule construction with evolutionary optimization of attribute ranges. Methods QFARM identifies rules of the form A ∈ [a1, a2] ∧ B ∈ [b1, b2] ∧ … → T ∈ [t1, t2], where the target attribute T and its range are user-defined, and the left-hand side consists of optimized ranges over numerical attributes. The output is organized as a hierarchical refinement tree: each edge adds one attribute and each root-to-node path defines the attribute set of that node. For each node, QFARM applies the multi-objective evolutionary algorithm NSGA-II [2] to optimize numerical ranges for the attribute set. Instead of selecting a single best rule, each node stores a Pareto front of non-dominated rules, preserving support-confidence trade-offs. The front is converted into sample-level scores to generate a ROC curve for each node. Only children with significant ROC/AUC improvement over their parent (DeLong test [3]) are retained in the tree. For interpretation, Pareto fronts are summarized using percentile-based histograms of attribute ranges, while individual rules and metrics remain available in human- and machine-readable formats for downstream analysis by researchers or AI agents. QFARM also validates whether rule structures discovered in one dataset remain meaningful in another. Ranges for previously learned attribute combinations are re-optimized using an independent dataset, and the resulting rules are assessed for consistency and predictive significance. Results We evaluated QFARM on the synthetic Friedman#1 dataset and public NHANES survival data. On Friedman#1, QFARM identified rules involving the known relevant variables and their combinations. Against strong rule-mining baselines, including RuleFit, PRIM bump hunting, subgroup discovery, and discretized association rules, QFARM achieved competitive performance, matching or outperforming baselines on the selected metrics. Beyond predictive performance, QFARM provides hierarchical rule organization, enabling interpretation of the impact of attribute additions. On NHANES survival data, the method identified biochemical variables associated with all-cause mortality, similar to variables identified by traditional Cox proportional hazards models. Conclusion QFARM provides a practical framework for structured rule discovery on numerical data with an integrated validation mechanism. By combining hierarchical rule construction and evolutionary optimization, it enables interpretable and robust rule discovery. This approach is particularly suitable for exploratory analysis in domains where both interpretability and generalization are critical. REFERENCES [1] Tsurinov, Petr & Shpynov, Oleg & Lukashina, Nina & Likholetova, Daria & Artyomov, Maxim. (2021). FARM: hierarchical association rule mining and visualization method. 1-1. 10.1145/3459930.3469499. [2] K. Deb, A. Pratap, S. Agarwal and T. Meyarivan, "A fast and elitist multiobjective genetic algorithm: NSGA-II," in IEEE Transactions on Evolutionary Computation, vol. 6, no. 2, pp. 182-197, April 2002, doi: 10.1109/4235.996017. [3] DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach. Biometrics. 1988 Sep;44(3):837-45. PMID: 3203132.
Doski-nf: an integrated computational pipeline for consensus variant calling and comprehensive cancer genomic profiling
*Ludovica, Perticone, Department of Molecular Medicine, Sapienza University of Rome Davide, Alfieri, Department of Molecular Medicine, Sapienza University of Rome Giovanni, Guglielmelli, Department of Molecular Medicine, Sapienza University of Rome Laura, Ottini, Department of Molecular Medicine, Sapienza University of Rome Valentina, Silvestri, Department of Molecular Medicine, Sapienza University of Rome
Doski-nf: an integrated computational pipeline for consensus variant calling and comprehensive cancer genomic profiling
*Ludovica, Perticone, Department of Molecular Medicine, Sapienza University of Rome Davide, Alfieri, Department of Molecular Medicine, Sapienza University of Rome Giovanni, Guglielmelli, Department of Molecular Medicine, Sapienza University of Rome Laura, Ottini, Department of Molecular Medicine, Sapienza University of Rome Valentina, Silvestri, Department of Molecular Medicine, Sapienza University of Rome
- Keywords
- cancer genomicscomputational pipelinevariant callingFFPE samplesNextflow
- Abstract
- Objective The widespread use of next-generation sequencing in cancer genomics has generated numerous bioinformatics tools with variable performance, potentially affecting downstream analyses and clinical decision-making. We developed a computational pipeline for matched germline and somatic cancer samples analysis, aimed at (i) building and validating a robust framework, optionally reproducing The Cancer Genome Atlas (TCGA)-like analyses, while integrating complementary analyses for comprehensive genomic characterization, including microsatellite instability, homologous recombination deficiency, tumor mutational burden, copy number alterations; and (ii) evaluating 20 different variant callers to optimize the workflow and fine-tuning optional features for specific sample types, i.e., Formalin-Fixed Paraffin-Embedded (FFPE) cancer samples. Methods We implemented a Nextflow workflow to process whole-exome sequencing short-reads data from FASTQ or BAM files and generate user-friendly output tables. When FASTQ files are available, reads are filtered, aligned to the reference genome and processed, depending on tool requirements. Germline and somatic variant calling are then performed in parallel, and the resulting VCFs are integrated through a consensus strategy retaining only variants supported by at least two different callers. A preliminary benchmarking was performed on two breast cancer datasets, one from fresh tissue and one from FFPE tissue, by comparing PASS-filtered VCFs from the 20 callers against ground truth sets, using precision and recall to calculate F1-scores. Results The TCGA-like pipeline integrating complementary genomic analyses was successfully implemented and is currently being employed for the comprehensive molecular characterization of in-house cancer samples. Within this framework, the performance of 20 different variant callers and their consensus combinations is being systematically evaluated to optimize variant detection across different sample types. Variant calling performance varied substantially across the evaluated tools and depended on the dataset analyzed. For the fresh tissue dataset, among individual callers, Mutect2 achieved the highest F1-score, while, for FFPE dataset both FreeBayes-som and ClairS performed best for SNVs and Strelka2 for INDELs. The consensus-based approach across multiple variant callers consistently improved performance by balancing precision and recall. In the fresh tissue dataset, systematic evaluation among all tested combinations showed that the best combination was ClairS-TO paired with DeepSomatic, reaching a peak F1-score=0.83. In contrast, performance in the FFPE dataset was generally lower and at least four callers are required. Among consensus setups, the combination of ClairS, Mutect2, Shimmer, and SomaticSniper performed best overall, achieving an F1-score=0.64. Conclusion This workflow may improve the robustness and reproducibility of variant calling while enabling comprehensive genomic characterization, improving the clinical translation of cancer genomic results. Benchmarking is currently being expanded to additional datasets derived from different tissues and experimental methodologies, and further filtering strategies will be implemented to reduce variant calling errors, specifically in FFPE cancer samples. Study supported by AIRC (MFAG30294) to V.S.
Machine learning approaches for predicting functional lncRNAs and key regulatory interactions in plant gene networks
Luiza Oliveira Romão*, Bioinformatics and Systems Biology Lab, University of Florida/ State University of Campinas Kevin Begcy, University of Florida, Microbiology and Cell Science Department, USA Renato Vicentini, Bioinformatics and Systems Biology Lab, State University of Campinas
Machine learning approaches for predicting functional lncRNAs and key regulatory interactions in plant gene networks
Luiza Oliveira Romão*, Bioinformatics and Systems Biology Lab, University of Florida/ State University of Campinas Kevin Begcy, University of Florida, Microbiology and Cell Science Department, USA Renato Vicentini, Bioinformatics and Systems Biology Lab, State University of Campinas
- Keywords
- lncRNAsplantsmachine learningdevelopmental transition
- Abstract
- OBJECTIVE Long non-coding RNAs (lncRNAs) are important regulatory components of gene expression in plants. However, their roles in controlling developmental transitions remain poorly characterized, particularly in crops with complex and polyploid genomes. Current computational approaches for lncRNA analyses are largely sequence-centered and predominantly developed for model plant species, limiting their ability to capture regulatory interactions in non-model crops. Similarly, gene regulatory networks (GRNs) provide a systems-level framework to investigate coordinated gene regulation, yet remain underexplored for lncRNA-mediated regulation. The main goal of this project focuses on investigating the involvement of lncRNAs and other regulatory RNAs in the transition from vegetative to reproductive development applying comparative genomics and machine learning based GRN inference. METHODS Genomic and transcriptomic datasets collected from several plant species including sugarcane (Saccharum spp.), Arabidopsis thaliana, Oryza sativa, Sorghum bicolor and Zea mays will be processed with Trimmomatic, quantified with Salmon and analysed for differential expression with DESeq2. Orthologous genes will be inferred through a pipeline combining OrthoFinder with targeted BLAST+, MAFFT and PhyML phylogenetic validation. Co-expression and regulatory networks will be inferred from expression matrices by integrating WGCNA, GENIE3 and ARACNe; supervised refinement will use SIRENE or GENIE3, complemented by unsupervised MRNET and contrastive learning via DeepMCL. Topological descriptors, structural features of lncRNAs derived from ViennaRNA, DMfold and RhoFold+, evolutionary conservation scores and co-expression patterns will be assembled into composite feature vectors. A network deep learning strategy inspired by the DLNet architecture will integrate expression profiles and network topology to discriminate vegetative from reproductive regulatory configurations, with k-fold cross validation assessing robustness and generalisation. RESULTS The integrative framework is expected to deliver a benchmarked computational pipeline for the systematic prediction of functional lncRNAs and the inference of lncRNA aware GRNs across species. Anticipated outputs include the prioritisation of conserved and species specific lncRNAs occupying central or connector positions within the inferred networks, supported by structural, evolutionary and co-expression evidence. Comparative analyses are expected to reveal conserved regulatory modules and species-specific rewiring associated with the vegetative to reproductive transition, alongside candidate lncRNAs repeatedly emerging as influential features in the deep learning classifier. The benchmarking of supervised, unsupervised and contrastive inference methods will quantify the added value of network aware deep learning for non-coding regulator discovery in polyploid crops. CONCLUSION By integrating open source tools into a benchmarked and reproducible workflow that mines public biomolecular databases, the framework will support knowledge discovery on the regulation exerted by lncRNAs in polyploid plant genomes. The systematic benchmarking of supervised, unsupervised, contrastive and topology aware deep learning strategies will yield reusable evidence on the performance of GRN inference methods for noncoding regulators, supporting testable hypotheses on the transition from vegetative to reproductive growth in sugarcane and the optimisation of breeding synchronisation programmes.
Impact of Non-Informative Censoring on the Performance of Propensity Scores Methods for Estimating Absolute Treatment Effects on Survival Outcome; Simulation study
1. Mahin Tatari*, Statistical Sciences Department, Sapienza University of Rome, Rome, Italy. And Istituto Superiore di Sanità, Rome, Italy. 2. Stefano Rosato, Istituto Superiore di Sanità, Rome, Italy. 3. Paola D’Errigo, Istituto Superiore di Sanità, Rome, Italy. 4. Giovanna Jona Lasinio, Statistical Sciences Department, Sapienza University of Rome, Rome, Italy.
Impact of Non-Informative Censoring on the Performance of Propensity Scores Methods for Estimating Absolute Treatment Effects on Survival Outcome; Simulation study
1. Mahin Tatari*, Statistical Sciences Department, Sapienza University of Rome, Rome, Italy. And Istituto Superiore di Sanità, Rome, Italy. 2. Stefano Rosato, Istituto Superiore di Sanità, Rome, Italy. 3. Paola D’Errigo, Istituto Superiore di Sanità, Rome, Italy. 4. Giovanna Jona Lasinio, Statistical Sciences Department, Sapienza University of Rome, Rome, Italy.
- Keywords
- Absolute measurecensoringperformancepropensity scoreSimulation study
- Abstract
- Introduction: Propensity Score (PS) methods are increasingly being used to reduce confounding in observational studies. While their performance for relative measure has been extensively studied, less is known about their accuracy in estimating absolute treatment effects when censoring is present. This study evaluates the performance of PS methods in estimating absolute treatment effects on survival outcomes under non-informative censoring. Methods and Materials: Monte Carlo simulations were conducted with 10 baseline covariates. Treatment assignment was generated via a logistic model. Mean and median survival times were estimated from Kaplan-Meier (KM) curves, with the mean calculated as the area under the curve. Four PS methods—matching, stratification, stabilized IPTW, and covariate adjustment—were applied to estimate absolute treatment effects in both the overall population and treated group (ATE, ATT). Performance was assessed across 54 scenarios varying by censoring rates, and hazard ratio (HR), using bias and mean squared error (MSE) as evaluation metrics in R version 4.4.2. Result: Stabilized IPTW consistently achieved the lowest bias and MSE, when censoring exceeds 50%. PS stratification produced unbiased estimates only under low censoring, while bias and MSE increased as treatment effects strengthened. PS adjustment exhibited the greatest sensitivity to censoring and large HR. For the ATT, both matching and stratification yielded accurate and stable estimates under low to moderate censoring, whereas stabilized IPTW maintained superior robustness under heavier censoring. Conclusion: Under non-informative censoring, Stabilized IPTW was the most robust for estimating absolute treatment effects on mean and median survival times in the presence of higher censoring.
A Multiplex Graph Framework for Higher-Order Gene Interactions in Lung Cancer Pathway Analysis
* Francesca Possenti, Sapienza Università di Roma Lorenzo Farina, Sapienza Università di Roma Paolo Tieri, CNR e Sapienza Università di Roma Manuela Petti, Sapienza Università di Roma
A Multiplex Graph Framework for Higher-Order Gene Interactions in Lung Cancer Pathway Analysis
* Francesca Possenti, Sapienza Università di Roma Lorenzo Farina, Sapienza Università di Roma Paolo Tieri, CNR e Sapienza Università di Roma Manuela Petti, Sapienza Università di Roma
- Keywords
- Hypergraphmultilayer-networkstranscriptomicsoncologynetwork-medicine
- Abstract
- Objective Lung adenocarcinoma (LUAD) is driven by coordinated alterations across major oncogenic pathways, including PI3K, cell cycle, p53, and DDR. Building on the core genes from 11 pathways – extensively studied and associated to the oncologic lung pathologies – the work aims to develop a multiplex (multi-layer) hypergraph structure to compare the cancer and normal conditions by studying how the behaviour emerging from their higher order multigene interactions differs. Methods Our approach develops a structured pipeline to identify and validate higher-order interactions in LUAD, contrasting normal and cancer samples. Candidate gene multiplets are selected using a hybrid strategy that integrates data-driven pairwise interaction structure with prior pathway knowledge, to reduce the search space while preserving interpretability. The information-theoretic Omega-information measure is used to quantify the candidates’ synergy (emergent cooperative effects) and redundancy (shared or overlapping information). Statistical validation is performed through a three-stage inferential framework that separates null-model significance testing from robustness assessment via bootstrap confidence intervals, ensuring that retained higher-order interactions are both statistically significant and stable, and not reducible to lower-order effects. The resulting framework is applied across the two conditions, yielding synergistic and redundant multiplex hypergraphs in which layers correspond to normal and cancer states. This enables direct comparison of higher-order informational organization between healthy and cancer lung tissue. Results Higher-order interaction analysis reveals a strong reorganization of pathway structure in LUAD compared to normal tissue. In normal samples, both synergistic and redundant interactions are widespread, forming a diverse landscape dominated by genome integrity modules involving cell cycle, DDR, and p53, with additional contributions from MYC, WNT/β-catenin, and NRF2_Oxidative_Stress signaling. In contrast, the cancer state shows an almost complete collapse of synergistic structure, with only a few patterns remaining, indicating a loss of emergent multi-gene cooperation. Redundant interactions are also markedly reduced and become highly concentrated in a restricted set of genes spanning core pathways, including Cell_cycle, DDR, MYC and NRF2_Oxidative_Stress response signaling – reflecting strong compression of pathway combinations. Conclusions The collapse of synergistic organization and the contraction of redundancy indicate a shift from a robust and interconnected regulatory architecture in normal tissue to a more constrained and specialized network in cancer. LUAD appears to reorganize higher-order interactions around a compact stress-adaptation core involving DNA damage control, cell-cycle regulation, redox homeostasis, and growth signaling pathways. Redundancy becomes concentrated within this restricted functional module rather than distributed across multiple processes, suggesting a reduction in global regulatory flexibility, alongside increased robustness of key mechanisms that sustain proliferation and survival.
Investigating Immunotherapy Response in Non–Small Cell Lung Cancer Using Machine Learning and Differential Networks
Salim, Sikder, MSc Program in Data Science Sapienza University of Rome Manuela, Petti, Sapienza University of Rome* Giulia, Fiscon, Università di Roma 'San Raffaele'
Investigating Immunotherapy Response in Non–Small Cell Lung Cancer Using Machine Learning and Differential Networks
Salim, Sikder, MSc Program in Data Science Sapienza University of Rome Manuela, Petti, Sapienza University of Rome* Giulia, Fiscon, Università di Roma 'San Raffaele'
- Keywords
- Differential NetworksMachine LearningImmunotherapy ResponseNon–small cell lung cancer
- Abstract
- Objective Non–small cell lung cancer (NSCLC) accounts for approximately 85% of all lung cancers and remains a leading cause of cancer-related mortality worldwide. Immune checkpoint inhibitors (ICIs), particularly PD-1/PD-L1 blockade, have transformed the therapeutic landscape, providing durable tumor control for a subset of patients. However, only a minority achieves long-term benefit, and the factors that distinguish patients who experience durable clinical benefit (DCB) from those who do not are still not fully understood. This motivates a central question: which clinical and genomic features distinguish patients who experience DCB from those who do not, and how do these groups differ in their systems-level feature interactions? Methods This work addresses this question through an integrated computational framework that combines supervised machine learning with differential phenotypic network analysis. Using a curated cohort of 240 NSCLC patients from the MSK-IMPACT study, we constructed a harmonized feature matrix encompassing demographic variables, treatment history, tumor mutational burden (TMB), genomic instability metrics, and smoking status. Preprocessing steps, including categorical encoding, redundancy pruning, SMOTE oversampling for class balancing and variance standardization, ensured compatibility across analytical pipelines. Four supervised models (Logistic Regression, Support Vector Machine, Decision Tree, and Random Forest) were trained using stratified 5-fold cross-validation. Moreover, to investigate multivariate structure beyond individual predictors, we constructed separate correlation networks for responder (DCB = YES) and non-responder (DCB = NO) groups and performed a 10,000-iteration permutation-based differential network analysis. Results Random Forest achieved the highest discriminative performance (AUPRC = 0.859, AUROC = 0.912), while Logistic Regression provided interpretable effect estimates. Across models, Tumor Mutational Burden, Fraction Genome Altered, and Lines of Treatment emerged as influential predictors. These patterns align with known immunobiological mechanisms, including the role of neoantigen load and the negative impact of genomic instability. Responders showed stronger coupling between TMB, genomic instability, treatment modality, and age, whereas non-responders exhibited distinct associations involving genomic instability, smoking status, and treatment type. The differential network highlighted several edges with meaningful shifts between groups, indicating that feature–feature coordination differs across clinical outcomes even when marginal effects appear similar. Conclusion Together, these findings demonstrate that durable immunotherapy benefit in NSCLC reflects not only the contribution of individual biomarkers but also coordinated interactions among mutational processes, genomic stability, treatment exposure, and clinical history. By integrating predictive modeling with systems-level network inference, this work provides a comprehensive and biologically grounded characterization of the determinants of durable clinical benefit and establishes a robust methodological framework for future biomarker discovery in immuno-oncology.
SOPHYSM: More Steps towards Digital Twins of Solid Tumours
* Diego, Cividini, Dipartimento di Informatica, Sistemistica e Comunicazione, Università degli Studi di Milano-Bicocca Marco Antoniotti, Dipartimento di Informatica, Sistemistica e Comunicazione, Università degli Studi di Milano-Bicocca; Biinformatics, Biostatistics, Bioimaging Centre (B4), Università degli Studi di Milano-Bicocca
SOPHYSM: More Steps towards Digital Twins of Solid Tumours
* Diego, Cividini, Dipartimento di Informatica, Sistemistica e Comunicazione, Università degli Studi di Milano-Bicocca Marco Antoniotti, Dipartimento di Informatica, Sistemistica e Comunicazione, Università degli Studi di Milano-Bicocca; Biinformatics, Biostatistics, Bioimaging Centre (B4), Università degli Studi di Milano-Bicocca
- Keywords
- SimulationCancerCell-lineagesDigital TwinImaging
- Abstract
- Objective: The notion of "digital twin" is well established in several engineering disciplines, but, due to the complexity of the field, it is still in the making when it comes to the Life Sciences. We are addressing this issue, while realizing that many different pieces must come together to progress toward the goal. The SOPHYSM system is a Julia tool-set for simulating the evolution of solid tumours, while tracking genotyping information at the Single-cell level. Methods: To be more useful for Life Science practitioners, SOPHYSM provides a histopathology image handling facility. The images are analysed to pinpoint cells in a tissue. The cells positions and inferred types are then fed into a spatial simulator as a basis to compute possible solid tumour evolutions. We describe the useful (and necessary) porting to Julia of one of the latest image analysis "algorithms'' for cell identification: the Python-based Cellpose. This allowed us to simplify our development pipeline for identifying (i.e., segment) cells' positions from histopathology images. Results: This effort required unpacking and reverse engineering what is a sophisticated AI tool, Cellpose, to port published weights to our Julia implementation, by leveraging the ONNX libraries. Conclusions: Our tests yield competitive measurements. It should be noted that this is an example of the importance of providing the weights of a trained AI model, to be able to reuse them in different settings, like, in our case, SOPHYSM.
FFPE-Open-ST: high-resolution, unbiased spatial transcriptomics profiling of archival samples
Martina, Brunetti*, Department of Experimental Medicine, Sapienza University, Rome, Italy Elena, Splendiani, Department of Experimental Medicine, Sapienza University, Rome, Italy Tanja Milena, Autilio, Department of Experimental Medicine, Sapienza University, Rome, Italy Salah, Ayoub, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Anastasiya, Boltengagen, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Daniel, León-Perinán, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Giorgia, Gugliuzza, Department of Experimental Medicine, Sapienza University, Rome, Italy Marie, Schott, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Cledi Alicia, Cerda-Jara, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Francesca, Gianno, Department of Radiological, Oncological and Anatomic Pathology, Sapienza University of Rome, Rome, Italy Marta, Moretti, Department of Experimental Medicine, Sapienza University, Rome, Italy Gwendolin, Thomas, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin, Germany Giuseppe, Macino, Fondazione per la ricerca Genomica ed Epigenomica ETS-FORGE, Udine, Italy Nikos, Karaiskos, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin, Germany Nikolaus, Rajewsky, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin, Germany Elisabetta, Ferretti, Department of Experimental Medicine, Sapienza University, Rome, Italy
FFPE-Open-ST: high-resolution, unbiased spatial transcriptomics profiling of archival samples
Martina, Brunetti*, Department of Experimental Medicine, Sapienza University, Rome, Italy Elena, Splendiani, Department of Experimental Medicine, Sapienza University, Rome, Italy Tanja Milena, Autilio, Department of Experimental Medicine, Sapienza University, Rome, Italy Salah, Ayoub, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Anastasiya, Boltengagen, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Daniel, León-Perinán, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Giorgia, Gugliuzza, Department of Experimental Medicine, Sapienza University, Rome, Italy Marie, Schott, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Cledi Alicia, Cerda-Jara, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin Francesca, Gianno, Department of Radiological, Oncological and Anatomic Pathology, Sapienza University of Rome, Rome, Italy Marta, Moretti, Department of Experimental Medicine, Sapienza University, Rome, Italy Gwendolin, Thomas, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin, Germany Giuseppe, Macino, Fondazione per la ricerca Genomica ed Epigenomica ETS-FORGE, Udine, Italy Nikos, Karaiskos, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin, Germany Nikolaus, Rajewsky, Laboratory for Systems Biology of Regulatory Elements, Berlin Institute for Medical Systems Biology (BIMSB), Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Hannoversche Str. 28, 10115 Berlin, Germany Elisabetta, Ferretti, Department of Experimental Medicine, Sapienza University, Rome, Italy
- Keywords
- spatial transcriptomicsFFPEcancerRNAbiomarkers
- Abstract
- Objective Formalin-fixed paraffin-embedded (FFPE) tissue blocks represent the gold-standard methods for preservation of tissue samples in clinical practice. Applying spatial transcriptomics (ST) to FFPE specimens provides unprecedented access to a vast repository of biological samples that can be useful for medical research including retrospective studies, given the long-term follow-up information, and documented therapeutic outcomes. However, transcriptome-wide ST profiling of FFPE tissues remains challenging due to formalin-induced crosslinking, which severely compromises RNA integrity and accessibility. FFPE-compatible ST technologies remain limited, and commercially available options rely mostly on probe-based approaches, which confine analyses to predefined gene panels. To overcome these limitations, we developed FFPE-Open-ST, a high-resolution whole-transcriptome platform for FFPE tissues with an integrated computational pipeline. Methods To compare FFPE and paired FF ST datasets, we first assessed the correlation between the protein-coding genes captured. Then, spatial domains were identified using Squidpy in Python and annotated based on marker genes. Pearson correlation was computed between shared cell types across FFPE and FF samples. To assess robustness and applicability to long-term archived specimens, we applied the method to FFPE mouse brain samples preserved for up to 6.5 years. Results FFPE Open-ST generates high-quality spatial transcriptomic data comparable to standard Open-ST on FF samples. Spatial expression patterns of canonical marker genes were highly consistent between FF and FFPE samples, both in terms of expression levels and spatial localization. The same cell types were identified in the FFPE and FF samples. Furthermore, in mouse samples stored for up to 6.5 years, we successfully identified the expected major cell populations and cell type-specific lncRNAs, with spatial patterns that remained stable over time. Conclusion FFPE-Open-ST enables high-resolution, unbiased spatial transcriptomics profiling of archival samples. This study lays the groundwork for broader studies on archived clinical samples with the aim to identify biomarkers and investigate mechanisms of therapy resistance.
Automation of Prescription Analysis in Older Adults: A Big Data Approach in Mexico
María Cecilia Ishida Gutiérrez*1 Vianey Iraís, Jiménez Ramírez1 René, Núñez Bautista1 Xóchitl Duque Alarcón1 Omar Fierro Fierro1 José López Loya1 Roxana Chávez Avila1,2 Fernanda Aimee Ríos Ruiz1
Automation of Prescription Analysis in Older Adults: A Big Data Approach in Mexico
María Cecilia Ishida Gutiérrez*1 Vianey Iraís, Jiménez Ramírez1 René, Núñez Bautista1 Xóchitl Duque Alarcón1 Omar Fierro Fierro1 José López Loya1 Roxana Chávez Avila1,2 Fernanda Aimee Ríos Ruiz1
- Keywords
- Drug UtilizationDecision Support Systems-ClinicalOlder adultAnti-Inflammatory Agents-Non-SteroidalPolypharmacy
- Abstract
- Objective: To evaluate the relationship between polypharmacy and the Potentially Inappropriate Prescribing (PIP) of non-steroidal anti-inflammatory drugs (NSAIDs) among adults aged 60 and older treated at a first-level medical institution in Chihuahua, Mexico. A key focus was the implementation and validation of an automated procedure, using specialised software, to systematically identify these pharmacological alerts at scale within electronic medical record databases. Methods: The study employed an analytical, cross-sectional, and retrospective design, processing a massive dataset of 21,634 medical consultation records. The Big Data methodology involved developing RStudio scripts, with artificial intelligence tools assisting in the logical structuring of the code. This automated system linked clinical diagnoses—standardised via International Classification of Diseases-10—with institutional medication catalogs to detect specific STOPP/START criteria, focusing on drug-drug and drug-disease interactions. Reliability was ensured through a technical validation process that compared the algorithm’s findings against a manual review, utilizing Cohen’s Kappa coefficient. Results: The prevalence of polypharmacy was 21.6%. The automated analysis revealed that consultations with polypharmacy had a significantly higher burden of at least one STOPP/START alert than those without (56.8% vs. 26.4%; p < 0.001). Among the most frequent issues, the algorithm identified STOPP A3 (therapeutic duplication) at 24.98% and STOPP H2 (NSAIDs in severe hypertension) at 30.37% within the polypharmacy group. Notably, START F3 (omission of gastroprotection) was the most prevalent alert overall, occurring in 43.04% of polypharmacy consultations compared to 25.14% in the non-polypharmacy group (p < 0.001). Logistic regression confirmed that polypharmacy is an independent risk factor, nearly tripling the likelihood of a PIP alert (Adjusted OR = 2.980). The algorithm achieved perfect concordance (κ = 1.000) and 100% sensitivity and specificity. Conclusion: The integration of automated data analysis in clinical settings transforms massive electronic health records into powerful surveillance tools. This study demonstrates that RStudio® is a valid and scalable alternative to manual review for identifying complex pharmacological risks. These systems optimize patient safety by identifying both medication errors (PIMs) and critical therapeutic omissions (PPOs), such as the lack of gastroprotection, while facilitating evidence-based health resource management in primary care.
Beyond Differential Expression: Network Centrality measures reveal Tissue-Specific Signatures in Tumor-Educated Platelets
Mattia*, Manna*, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma Manuela, Petti, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma Lorenzo, Farina, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma Alessandro, Taraborelli, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma Stefano, Rinaldi, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma
Beyond Differential Expression: Network Centrality measures reveal Tissue-Specific Signatures in Tumor-Educated Platelets
Mattia*, Manna*, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma Manuela, Petti, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma Lorenzo, Farina, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma Alessandro, Taraborelli, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma Stefano, Rinaldi, Dipartimento di Ingegneria Informatica - Automatica e Gestionale Antonio Ruberti (DIAG) - Sapienza Università di Roma
- Keywords
- RNA-SeqTumor educated plateletsNetwork oncologyNode based graph metricsDifferential Networks
- Abstract
- Objective Tumor-educated platelets (TEPs) are circulating blood components that play a central role in both systemic and local responses to tumor growth, leading to alterations in their RNA profiles. Recent studies have demonstrated that the TEP transcriptome can be leveraged for minimally invasive cancer diagnostics. In a previous analysis based on TEPs transcriptomic data, we highlighted the diagnostic potential of a specific set of key genes for glioblastoma multiforme (GBM). In that study, we applied network-based approaches and identified a subset of 42 candidates. These genes were subsequently validated through enrichment analysis, and a binary classifier used to distinguish between healthy individuals and cancer patients achieved high performance. In the present work, we apply this analytical framework to four additional cancer types—breast cancer, colorectal cancer, non-small cell lung cancer and pancreatic adenocarcinoma to investigate whether the key genes identified in GBM are tissue-specific or shared across different tumor types. Our objective is to demonstrate that the key genes identified for GBM differ from those associated with other cancer types, thereby suggesting that a network-based approach captures biologically meaningful, tissue-specific signals rather than noise, even when the data are derived from peripheral liquid biopsies, thus reinforcing the validity of TEPs based approaches. Methods To identify key genes, we employed a combination of differential expression analysis (DEA) and differential co-expression (DCE) network analysis. Initially, DEA was performed to detect differentially expressed genes (DEGs). These genes were then used to construct differential co-expression networks, which were subsequently analyzed to identify key nodes based on centrality metrics. This procedure was applied independently to five different cancer types, resulting in a distinct set of key genes for each condition. Then we performed a comparative analysis across these gene sets to determine which genes were tissue-specific. Results The analysis of key genes across all conditions and centrality metrics provides evidence that the network based approach identifies tissue specific genes. In contrast, the DEGs based approach does not recover tissue-specific signals and identifies largely overlapping gene sets across different cancer types. More specifically, for each condition, more than 70% of the key genes identified using betweenness, closeness, and degree centrality were tissue-specific. In contrast, for DEGs based selection, more than 70% of the genes were shared across conditions. Conclusion The results suggest that the network approach is effective in identifying tissue specific gene signatures, indicating its potential to uncover biologically relevant markers in cancer transcriptomic data, also the identification of tissue specific genes through a peripheral liquid biopsy (TEPs based) further supports the method.
Network-Based Comparison of Solid Tissue and Tumor-Educated Platelet Transcriptomes in Glioblastoma
Paolo, Meli, Sapienza Università di Roma. Stefano, Rinaldi ,Sapienza Università di Roma. Alessandro, Taraborelli, Sapienza Università di Roma. Mattia, Manna, Sapienza Università di Roma Lorenzo, Farina, Sapienza Università di Roma Manuela, Petti, Sapienza Università di Roma
Network-Based Comparison of Solid Tissue and Tumor-Educated Platelet Transcriptomes in Glioblastoma
Paolo, Meli, Sapienza Università di Roma. Stefano, Rinaldi ,Sapienza Università di Roma. Alessandro, Taraborelli, Sapienza Università di Roma. Mattia, Manna, Sapienza Università di Roma Lorenzo, Farina, Sapienza Università di Roma Manuela, Petti, Sapienza Università di Roma
- Keywords
- Tumor-Educated PlateletsNetwork BiologyDifferential Co-Expression NetworksNetwork ModelCommunity Detection
- Abstract
- Glioblastoma (GBM) is an aggressive brain tumor diagnosed using invasive techniques such as tissue biopsy and imaging. Liquid biopsy approaches have recently attracted interest, and several studies have demonstrated the potential of analysing tumor-educated platelets (TEPs) transcriptome in cancer diagnosis. However, tumor related signals in the blood are weaker and more diluted. Objective The aim of this work is to develop a network-based framework to compare co-expression patterns characterizing transcriptomic data from solid and liquid biopsy in glioblastoma, while providing a coherent interpretation of changes in gene relationships and their role in higher-order network organization. Methods We downloaded publicly available RNA-seq datasets from GBM patients and healthy controls (HC), including both solid tissue and platelet-derived samples. We filtered out the low-expression genes and selected the subset of differentially expressed genes in the tissue. We used Pearson correlation to assess the level of co-expression between genes in each condition: GBM tissue, HC tissue, GBM platelets and HC platelets. Two differential co-expression networks (tissue, platelets) were constructed by comparing Pearson correlation between GBM and healthy conditions through Fisher’s z-transformation and z-test. Given the 3 possible co-expression links, i.e. positive correlation (+), negative correlation (-) and no significant correlation (0), we introduced a Co-Expression Transition Network Model (CTNM) classifying edges of the differential networks into four transition types: appearance (0+) and disappearance (+0) of positive correlation, and appearance (0−) and disappearance (−0) of negative correlation in GBM compared to HC. Then, we studied the topology of the differential networks through a clustering strategy exploiting the proposed CTNM. Finally, Reactome enrichment analysis was performed to functionally compare communities between tissue and platelets networks. Results The application of the CTNM and the analysis of higher-order structures in the differential networks revealed that large cliques are composed exclusively of 0+ or +0 edges, and never both, suggesting cooperative but mutually exclusive behaviour. In contrast, 0− and −0 transitions behave antagonistically, not contributing to dense structures. Based on these observations, the clustering algorithm assigned each cluster a dominant cooperative transition and optimized intra-cluster coherence while favouring antagonistic interactions between clusters. In tissue network, the resulting structure identified three main communities functionally associated with immune response, neuronal system, and cellular organization. In platelet-derived data, only the presence of the functional clusters related to the immune response and cellular organization is confirmed. Conclusion This work introduces a transition-based framework for the analysis of differential co-expression networks. The proposed CTNM provides a coherent interpretation of differential interactions distinguishing between appearance and disappearance of positive and negative correlations, overcoming ambiguities associated with z-scores interpretation. Applied to glioblastoma transcriptomic data, the framework was able to detect biologically meaningful functional organization in solid tissue and partially preserved in platelet-derived data.
Quantum Optimization for Protein-Protein Interaction Network Alignment
* Merle, Stahl, Institute for Computational Systems Biomedicine, University of Hamburg, 22761 Hamburg, Germany Robert J., Banks, Parity Quantum Computing Germany GmbH, 20095 Hamburg, Germany Matthias, Traube, Parity Quantum Computing Germany GmbH, 20095 Hamburg, Germany Josua, Unger, Parity Quantum Computing GmbH, A-6020 Innsbruck, Austria Jan, Baumbach, Institute for Computational Systems Biomedicine, University of Hamburg, 22761 Hamburg, Germany; Department of Mathematics and Computer Science, University of Southern Denmark, 5230 Odense, Denmark Mhaned, Oubounyt, Institute for Computational Systems Biomedicine, University of Hamburg, 22761 Hamburg, Germany
Quantum Optimization for Protein-Protein Interaction Network Alignment
* Merle, Stahl, Institute for Computational Systems Biomedicine, University of Hamburg, 22761 Hamburg, Germany Robert J., Banks, Parity Quantum Computing Germany GmbH, 20095 Hamburg, Germany Matthias, Traube, Parity Quantum Computing Germany GmbH, 20095 Hamburg, Germany Josua, Unger, Parity Quantum Computing GmbH, A-6020 Innsbruck, Austria Jan, Baumbach, Institute for Computational Systems Biomedicine, University of Hamburg, 22761 Hamburg, Germany; Department of Mathematics and Computer Science, University of Southern Denmark, 5230 Odense, Denmark Mhaned, Oubounyt, Institute for Computational Systems Biomedicine, University of Hamburg, 22761 Hamburg, Germany
- Keywords
- Protein-protein interaction network alignmentMaximum Common SubgraphQuantum Approximate Optimization AlgorithmHybrid quantum-classical algorithms
- Abstract
- Objective: Protein-protein interaction (PPI) network alignment combines structural and biological information to detect conserved molecular systems across species, supporting functional annotation and the study of disease mechanisms. Global PPI alignment remains challenging: heuristic methods cannot guarantee optimal solutions, while exact methods are prohibitive at scale. Quantum optimization offers a promising route to navigating these combinatorial search spaces more exactly, potentially yielding alignments inaccessible to classical methods. Methods: We model PPI alignment as a Maximum Common Subgraph (MCS) problem on a modular product graph, where nodes represent candidate pairs weighted by sequence similarity and edges encode topological consistency. The MCS problem is reduced to a minimum weighted vertex cover, yielding a quantum-amenable binary formulation. We solve this with the Quantum Approximate Optimization Algorithm (QAOA), designing five variants spanning penalty-based to mixer-constrained formulations: (i) a baseline transverse-field method, (ii) a matching-based two-qubit mixer covering each matched edge, (iii) a clique-cover formulation with auxiliary qubits enabling two-body mixers, (iv) a clique-cover variant with multi-controlled mixers, and (v) a fully constrained formulation restricted to valid vertex covers. The four penalty-based formulations allow tunable constraint strength, ranging from strict exact alignment to softer matching tolerant of missing or noisy interactions. To address current hardware limitations, we embed these variants in a branch-and-bound framework: classical preprocessing and branching reduce problem size, while QAOA solves the resulting subproblems where quantum advantage is most tractable, preserving exactness throughout. Results: We evaluate the framework on synthetic networks from NAPAbench2 and real-world human-mouse PPI subnetworks from KEGG pathways. Alignment quality is measured on the mapped subgraphs using the symmetric substructure score for topology and functional conservation metrics tailored to each benchmark, alongside node coverage and quantum-resource metrics including runtime, ground-state probability, circuit depth, qubit count, and communication overhead. The QAOA branch-and-bound variants achieve exact topological consistency, while maintaining functional conservation comparable to leading classical aligners, at the cost of lower node coverage. The five variants trade off differently across quantum-resource metrics, with penalty-based formulations achieving the lowest qubit count and circuit depth, at the cost of weaker feasibility guarantees in the mixer. Performance is strongest on closely related networks, while more divergent networks benefit from the penalty-based formulations. Across selected KEGG pathways, aligned subnetworks retained most disease-associated genes, indicating preserved biological relevance despite reduced coverage. Conclusion: We present a hybrid quantum-classical framework that solves PPI alignment as a MCS problem, combining five problem-tailored QAOA variants with branch-and-bound to guarantee solution optimality. The framework matches the biological conservation of established aligners while enforcing perfect topological consistency. These results demonstrate the potential of quantum optimization to deliver exact, biologically accurate network alignment, with performance expected to scale as quantum hardware continues to mature.
An atlas of gene expression and alternative splicing across stress conditions in microalgal species: a focus on CO2 capture.
Meryam, Carrus, University of Tuscia* Jessica, Di Martino, University of Tuscia Manuel, Arcieri, Sapienza University of Rome Lorenzo, Arcioni, Sapienza University of Rome Paolo, Bottoni, Sapienza University of Rome Claudia, Cafaro, CNR-Institute of Atmospheric Pollution Research, Division of Rome Loretta, De Giorgi, CNR-Institute of Atmospheric Pollution Research, Division of Rome Antonio, Fardelli, CNR-Institute of Atmospheric Pollution Research, Division of Rome Marcella, Pasqualetti, University of Tuscia Daniele, Canestrelli, University of Tuscia Tiziana, Castrignanò, University of Tuscia
An atlas of gene expression and alternative splicing across stress conditions in microalgal species: a focus on CO2 capture.
Meryam, Carrus, University of Tuscia* Jessica, Di Martino, University of Tuscia Manuel, Arcieri, Sapienza University of Rome Lorenzo, Arcioni, Sapienza University of Rome Paolo, Bottoni, Sapienza University of Rome Claudia, Cafaro, CNR-Institute of Atmospheric Pollution Research, Division of Rome Loretta, De Giorgi, CNR-Institute of Atmospheric Pollution Research, Division of Rome Antonio, Fardelli, CNR-Institute of Atmospheric Pollution Research, Division of Rome Marcella, Pasqualetti, University of Tuscia Daniele, Canestrelli, University of Tuscia Tiziana, Castrignanò, University of Tuscia
- Keywords
- MicroalgaeComparative transcriptomicsCO₂ captureAlternative splicingCo-expression networks
- Abstract
- Objective Microalgae are highly efficient biological systems for carbon dioxide (CO₂) assimilation due to their remarkable photosynthetic efficiency, metabolic flexibility, and ability to adapt to environmental stress. However, the molecular and regulatory mechanisms underlying carbon fixation and stress acclimation remain only partially understood, particularly across evolutionarily distant lineages. This study aimed to characterize conserved and lineage-specific transcriptional responses associated with CO₂ capture in representative Chlorophyta and Ochrophyta species through an integrated comparative transcriptomic approach combining differential gene expression, alternative splicing, and gene co-expression network analyses. Methods Public RNA-seq datasets were retrieved from the NCBI Sequence Read Archive (SRA) and selected according to experimental relevance, biological replication, and reference genome availability. Six BioProjects comprising 163 RNA-seq experiments (~1.03 TB) from multiple Chlorophyta and Ochrophyta species were reanalyzed using a standardized bioinformatics workflow. Raw reads underwent quality assessment with FastQC, trimming with Trimmomatic, alignment with HISAT2, and transcript reconstruction and quantification with StringTie. Differential gene expression analysis was performed using DESeq2, while alternative splicing events were investigated using MAJIQ and VOILA through the identification of Local Splicing Variations (LSVs). Gene co-expression networks were reconstructed using WGCNA to identify co-regulated modules and hub genes associated with photosynthesis, carbon metabolism, and environmental adaptation. Results Comparative transcriptomic analyses revealed extensive but species-specific transcriptional remodeling under high CO₂, light-related, and diel stress conditions. Thousands of differentially expressed genes (DEGs) were identified across experimental comparisons, with particularly strong responses observed in Seminavis robusta and Skeletonema marinoi. Conserved up-regulated genes included key components of carbon concentrating mechanisms and photosynthetic pathways, such as carbonic anhydrase (CA), phosphoenolpyruvate carboxylase (PEPC), Calvin cycle enzymes, and photoprotective proteins. Distinct transcriptional configurations emerged between Chlorophyta and Ochrophyta, suggesting lineage-specific adaptive strategies. Alternative splicing analyses identified condition-dependent splicing variations affecting genes involved in photosynthesis, RuBisCO activation, ATP-citrate metabolism, and plastidial regulation, supporting a role for post-transcriptional regulation in stress adaptation. Co-expression network analysis identified modules enriched in genes associated with light harvesting, redox balance, and carbon fixation. Several hub genes, including light-harvesting complex proteins and carbon metabolism enzymes, emerged as central regulatory nodes potentially linked to enhanced photosynthetic efficiency and environmental resilience. Conclusion This integrative study demonstrates that efficient CO₂ capture in microalgae results from a multilayered regulatory organization involving coordinated transcriptional regulation, alternative splicing, and co-expression network dynamics. Despite conserved functional axes related to photosynthesis and carbon metabolism, Chlorophyta and Ochrophyta display distinct regulatory architectures underlying environmental adaptation. These findings provide new insights into microalgal resilience mechanisms and identify candidate molecular targets for future biotechnological applications aimed at improving sustainable carbon sequestration systems.
Evaluating Genome Language Model Embeddings as Transferable Priors for Metagenomic Binning
Noga, Aharony, Columbia University Mohammed, AlQuraishi, Columbia University
Evaluating Genome Language Model Embeddings as Transferable Priors for Metagenomic Binning
Noga, Aharony, Columbia University Mohammed, AlQuraishi, Columbia University
- Keywords
- Metagenomic binninggenome language modelstransfer learningmicrobial communitiesrepresentation learning
- Abstract
- Objective. Metagenomic binning, the assignment of assembled contigs to their genomes of origin, is central to characterizing unculturable microbial diversity. Current de novo binners rely heavily on tetranucleotide frequency (TNF) and read coverage, often with sample-specific optimization or retraining. This limits reuse across environments and remains challenging for short contigs, where composition statistics are noisy. We investigate whether pretrained genome language models (gLMs) can provide transferable contig representations for metagenomic binning, with the long-term goal of moving toward models that generalize across microbial communities. Methods. We evaluate gLM-derived representations on CAMI benchmark communities, simulated metagenomes with known genome-of-origin labels, using AMBER metrics including completeness, purity, and genome recovery. We compare unsupervised embedding-based binning, sample-specific representation learning, and cross-community transfer experiments against composition-based baselines. Results. Raw pretrained embeddings alone do not yet match established binning tools, suggesting that species-level binning is not directly recovered from off-the-shelf gLM representations. Sample-specific learning with gLM features also remains below strong TNF-based baselines. However, transfer experiments provide preliminary evidence that gLM embeddings retain reusable information across microbial communities, even when absolute binning performance remains below mature state-of-the-art methods. Conclusion. These results do not yet establish a competitive universal binner; existing tools reflect years of optimization around TNF, coverage, and per-sample training. Instead, they provide an early proof of concept for a property needed by universal binning models: cross-community transfer. Future work will test whether this signal can be scaled and optimized into a single model that generalizes to new metagenomes without per-sample retraining.
High-Resolution Spatial Transcriptomics Using Open-ST Reveals Alterations across Midbrain Subregions in an Alzheimer’s Disease Mouse Model
Pasquale, Sibilio*, Department of Wellbeing, Health and Environmental Sustainability - BESSA, Sapienza University of Rome, 02100 Rieti, Italy Antonio Angeloni, Department of Wellbeing, Health and Environmental Sustainability - BESSA, Sapienza University of Rome, 02100 Rieti, Italy Elena, Splendiani, Department of Experimental Medicine, Sapienza University of Rome, 00161 Rome, Italy Martina, Brunetti, Department of Experimental Medicine, Sapienza University of Rome, 00161 Rome, Italy Livia, Barbera, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation, 00143 Rome, Italy; Department of Medicine and Surgery, Università Campus Bio-Medico Di Roma, 00128 Rome, Italy Cristian, Fiorucci, Department of Sciences, University of Roma Tre, 00146 Rome, Italy Giorgia, Gugliuzza, Department of Experimental Medicine, Sapienza University of Rome, 00161 Rome, Italy Tanja Milena, Autilio, Department of Experimental Medicine, Sapienza University, Rome, 00161 Rome Italy Paraskevi, Krashia, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation , 00143 Rome, Italy; Department of Sciences and Technologies for Sustainable; Development and One Health, Università Campus Bio-Medico Di Roma, Via Alvaro del Portillo, 21–00128, Rome, Italy Annalisa, Nobili, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation , 00143 Rome, Italy; Department of Medicine and Surgery, Università Campus Bio-Medico Di Roma, 00128 Rome, Italy Luisa, Lo Iacono, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation , 00143 Rome, Italy Valerio, Fulci, Department of Molecular Medicine, Sapienza University of Rome, Rome, Italy Giuseppe, Macino, Fondazione per la ricerca Genomica ed Epigenomica ETS-FORGE, Udine, Italy Marcello, D'Amelio, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation , 00143 Rome, Italy; Department of Medicine and Surgery, Università Campus Bio-Medico Di Roma, 00128 Rome, Italy Elisabetta, Ferretti, Department of Experimental Medicine, Sapienza University of Rome, 00161 Rome, Italy
High-Resolution Spatial Transcriptomics Using Open-ST Reveals Alterations across Midbrain Subregions in an Alzheimer’s Disease Mouse Model
Pasquale, Sibilio*, Department of Wellbeing, Health and Environmental Sustainability - BESSA, Sapienza University of Rome, 02100 Rieti, Italy Antonio Angeloni, Department of Wellbeing, Health and Environmental Sustainability - BESSA, Sapienza University of Rome, 02100 Rieti, Italy Elena, Splendiani, Department of Experimental Medicine, Sapienza University of Rome, 00161 Rome, Italy Martina, Brunetti, Department of Experimental Medicine, Sapienza University of Rome, 00161 Rome, Italy Livia, Barbera, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation, 00143 Rome, Italy; Department of Medicine and Surgery, Università Campus Bio-Medico Di Roma, 00128 Rome, Italy Cristian, Fiorucci, Department of Sciences, University of Roma Tre, 00146 Rome, Italy Giorgia, Gugliuzza, Department of Experimental Medicine, Sapienza University of Rome, 00161 Rome, Italy Tanja Milena, Autilio, Department of Experimental Medicine, Sapienza University, Rome, 00161 Rome Italy Paraskevi, Krashia, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation , 00143 Rome, Italy; Department of Sciences and Technologies for Sustainable; Development and One Health, Università Campus Bio-Medico Di Roma, Via Alvaro del Portillo, 21–00128, Rome, Italy Annalisa, Nobili, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation , 00143 Rome, Italy; Department of Medicine and Surgery, Università Campus Bio-Medico Di Roma, 00128 Rome, Italy Luisa, Lo Iacono, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation , 00143 Rome, Italy Valerio, Fulci, Department of Molecular Medicine, Sapienza University of Rome, Rome, Italy Giuseppe, Macino, Fondazione per la ricerca Genomica ed Epigenomica ETS-FORGE, Udine, Italy Marcello, D'Amelio, Department of Experimental Neurosciences, IRCCS Santa Lucia Foundation , 00143 Rome, Italy; Department of Medicine and Surgery, Università Campus Bio-Medico Di Roma, 00128 Rome, Italy Elisabetta, Ferretti, Department of Experimental Medicine, Sapienza University of Rome, 00161 Rome, Italy
- Keywords
- BioinformaticsSpatial TranscriptomicsBiostatisticsAlzheimerPrecision Medicine
- Abstract
- Objective Early dysfunction of the dopaminergic circuits has emerged as a relevant feature of Alzheimer’s disease-related pathology, with evidence suggesting a selective vulnerability of ventral tegmental area (VTA) neurons compared with other midbrain dopaminergic populations. Here, we aimed to apply Open-ST, a high-resolution spatial transcriptomics approach, to investigate the region-specific molecular signatures of neuronal vulnerability and apoptosis in the dopaminergic midbrain neurons. Methods Coronal midbrain sections were profiled from four experimental groups defined by genotype and age at sacrifice: Heterozygous male Tg2576 mice (Tg) overexpressing a mutated human amyloid precursor protein (APP) and their wild-type (Wt) littermates sacrificed at 1- and 3-months developmental stages. Whole-transcriptome mRNA capture was performed using Open-ST, enabling spatially resolved transcriptomic profiling at subcellular resolution (0.6 μm spacing between capture spots). Spatial domains were identified through spatial clustering and multi-sample spatial subclustering using BANKSY in R, enabling the delineation of anatomically and transcriptionally distinct midbrain regions. Dopaminergic neuronal domains were further resolved into VTA and SNc compartments, and spatial clusters were annotated through marker-based cell-type assignment. Finally, differential pathway activation analysis was performed to compare apoptosis-related transcriptional programs between genotypes, ages, and anatomical regions. Results Open-ST resolved the cellular and anatomical organization of the coronal midbrain section, allowing marker-based annotation of major spatial domains. Dopaminergic neuronal territories were successfully subclustered into VTA and SNc, while additional midbrain structures, including substantia nigra pars reticulata and GABAergic interpeduncular nucleus domains, were also identified. Differential pathway activation analysis revealed a significant increase in apoptosis pathway activation in the VTA of 3-month-old Tg mice compared with age-matched Wt. This genotype-dependent activation was regionally selective, as a comparable trend was not observed in the SNc. At 1 month, APP transgenic mice already showed a similar direction of apoptosis pathway activation in the VTA compared with wild-type mice, although this effect did not reach statistical significance. Conclusion Our findings demonstrate that Open-ST can capture region-specific molecular alterations within anatomically complex midbrain structures at near-cellular/subcellular resolution. By combining high-resolution spatial transcriptomics with spatial clustering, cell-type annotation, and pathway-level analysis, we identified a selective activation of apoptosis-related programs in the VTA of Tg mice, preceding or accompanying early dopaminergic vulnerability. This approach provides a powerful framework to spatially map neurodegenerative processes and to distinguish region-specific disease mechanisms.
Pathway-Mediated Networks for Drug Repurposing in Breast Cancer
* Pedro, Duarte, University of Bologna Francesco, Durazzi, University of Bologna Daniel, Remondini, University of Bologna
Pathway-Mediated Networks for Drug Repurposing in Breast Cancer
* Pedro, Duarte, University of Bologna Francesco, Durazzi, University of Bologna Daniel, Remondini, University of Bologna
- Keywords
- drug repurposingmultilayer networksbreast cancerpathway mediationnetwork medicine
- Abstract
- Objective Many biological systems can be represented as multilevel networks, where direct associations between nodes in one layer are also shaped by other layers. However, most network analyses focus only on pairwise similarities, overlooking how higher-level organization alters node importance. Here, we apply a multilayer mediation framework to drug repurposing in breast cancer. We quantify how pathway structure redefines drug relevance, beyond simple target overlap. Methods The system is represented by a tripartite network composed of drugs, proteins, and biological pathways. The network contains drug-protein links obtained from drug–target associations, and protein-pathway links from curated pathway annotations. Drug–protein associations are extracted from the ChEMBL database, disease–protein associations from Open Targets, and protein–pathway mappings from Reactome. We restrict the analysis to the overlap of proteins that are affected by breast cancer, targeted by drugs and annotated in human pathways, yielding 846 relevant proteins. This formulation identifies relevant drugs using only shared protein targets, the direct representation, and examines how drug relevance changes when the pathway structure is included, which we call the mediated representation. For the direct representation, we contract the drug-protein bipartite network into a drug-drug network linking drugs with shared protein targets. For the mediated representation, we contract the protein-pathway bipartite network into a protein-protein network. Proteins are linked if they share a pathway, so that proximity reflects functional context. Drugs are mapped onto overlapping protein subsets, shortest-path distances between subsets are computed and average distances between subsets are converted into similarity weights, producing the pathway-mediated representation of the drug-drug network. We compute centralities C for both direct and mediated representations, and quantify the impact of pathway-mediation in the metric ΔC=C_mediated-C_direct. Results The mediated representation reorders drug rankings into distinct separated regimes. Drugs with high positive ΔC only become important when pathway mediation is considered. In this group, we find drug classes strongly associated with breast cancer biology, such as kinase inhibitors and signaling regulators. The framework identifies well-established breast cancer drugs, like lapatinib and talazoparib, as well as literature-supported repurposing candidates, like copanlisib. Drugs with strongly negative ΔC lose relevance when the functional context is considered and are dominated by promiscuous drugs, such as ion-channel blockers (like carbamazepine) and CNS-active agents (like citalopram), often deprioritized in network-based repurposing studies. ΔC≈0 identifies drugs whose importance remains stable across the direct and mediated representations. Conclusion Overall, the results show that the pathway-mediated structure amplifies biologically coherent signals, while suppressing artifacts arising from target overlap. These results suggest that multilayer structural mediation provides a biologically grounded and interpretable framework to identify functionally embedded drug candidates, using relational biological data, instead of large-scale training data. This may help prioritize novel treatments in less-studied diseases, where validated repurposing knowledge is limited.
Early Warning of Extremely Rare Events from Pediatric Cardiac Intensive Care Electronic Health Records
Davide, Passaro, Department of Statistical Sciences Sapienza - University of Rome Luca, Tardella, Department of Statistical Sciences Sapienza - University of Rome Tiziana, Fragasso, Pediatric Cardiac Intensive Care Unit Bambino Gesù Children’s Hospital Valeria, Raggi, Pediatric Cardiac Intensive Care Unit Bambino Gesù Children’s Hospital
Early Warning of Extremely Rare Events from Pediatric Cardiac Intensive Care Electronic Health Records
Davide, Passaro, Department of Statistical Sciences Sapienza - University of Rome Luca, Tardella, Department of Statistical Sciences Sapienza - University of Rome Tiziana, Fragasso, Pediatric Cardiac Intensive Care Unit Bambino Gesù Children’s Hospital Valeria, Raggi, Pediatric Cardiac Intensive Care Unit Bambino Gesù Children’s Hospital
- Keywords
- Rare Event PredictionMachine learningPediatric cardiac intensive care unitEarly Warning Systems
- Abstract
- Objective: The prediction of extremely rare events is becoming increasingly important in modern medicine. In intensive care units (ICUs), particularly severe events may occur that are often difficult to predict even for experienced clinicians. This study focuses on a specific ICU setting: the Pediatric Cardiac Intensive Care Unit (PICU). Due to the unique characteristics of pediatric cardiac patients, evidence and predictive models developed for adult ICUs are not directly transferable, making the development of dedicated algorithms necessary. The aim of this work is to investigate this problem using data from the Pediatric Cardiac Intensive Care Unit of the Bambino Gesù Children’s Hospital in Rome, focusing on the prediction of CLABSI (Central Line-Associated Bloodstream Infection), a severe bloodstream infection associated with central venous catheters. In particular, we developed an early warning system designed to activate when a CLABSI event is expected to occur within the following 12 hours. Methods: The early warning system was developed using data from 3,229 PICU admissions. Classification was performed using the eXtreme Gradient Boosting (XGBoost algorithm), which is considered one of the most effective approaches for tabular clinical data. A stratified group split at the patient level was applied to the dataset to ensure that the same patients did not appear in both training and test sets, while preserving a similar proportion of CLABSI events across the splits. Since XGBoost does not explicitly model temporal dependencies, temporal variability was incorporated by providing the algorithm not only with data at time t, but also with sequences of previous observations. This allowed the model to account for temporal changes in clinical variables despite using a tabular machine learning approach. Model performance was evaluated using the Area Under the Precision–Recall Curve (AUPRC). Results: The proposed approach achieved a substantial improvement compared with the baseline prevalence. SHAP analysis identified clinically meaningful variables among the most influential predictors, including procalcitonin, which clinicians commonly associate with CLABSI onset. Conclusion: This preliminary study demonstrates the potential of early warning algorithms for predicting rare but clinically significant events in pediatric cardiac intensive care settings.
Mutational signature in Mycobacterium tuberculosis under pretomanid and nitric oxide exposure
Rajender Kumar*1, Blanka Andersson1, Thomas Schön1,2,3 1Department of Biomedical and Clinical Sciences, Division of Infection and Inflammation, Linköping University, Linköping, Sweden 2Department of Infectious Diseases, Kalmar County Hospital, Linköping University, Kalmar, Sweden 3Department of Infectious Diseases in Östergötland, Linköping University, Linköping, Sweden.
Mutational signature in Mycobacterium tuberculosis under pretomanid and nitric oxide exposure
Rajender Kumar*1, Blanka Andersson1, Thomas Schön1,2,3 1Department of Biomedical and Clinical Sciences, Division of Infection and Inflammation, Linköping University, Linköping, Sweden 2Department of Infectious Diseases, Kalmar County Hospital, Linköping University, Kalmar, Sweden 3Department of Infectious Diseases in Östergötland, Linköping University, Linköping, Sweden.
- Keywords
- Mycobacterium tuberculosisNitric oxidePretomanidkatG mutationNGS
- Abstract
- WHO estimated that there are 400,000 cases of multidrug-resistant tuberculosis (TB) with a mortality close to 40%. During infection, Mycobacterium tuberculosis (Mtb) faces several stresses, one of them are reactive nitrogen intermediates like nitric oxide (NO), produced by host immune cells. Additionally, NO is produced when pretomanid (PRT) a nitroimidazole drug used for MDR TB, is activated inside macrophages. The katG (S315T) mutation is described as the first step towards MDR TB and is associated with reduced antioxidant capacity. Our hypothesis is that strains with a lack of anti-oxidative capacity due to a katG mutation are more prone to mutagenic changes from oxidative stress, including NO, compared to katG WT and a laboratory control strain. Methods: Clinical isolates of Mtb with and without a katG mutation derived from the same patient before and after INH drug resistance development, and as an internal control, the Mtb H37Rv were included. The strains were separately exposed to PRT (0.03 µg/ml) and NO (0.16, 0.3 mM for H37Rv) for 1X (one) and 3X (three times) exposure. Genomic DNA was extracted, and short-read genome sequencing was performed, followed by de novo assembly with the SPAdes program. Average nucleotide identities (ANI) were calculated using the Pyani program. Mutational signature analysis was performed to detect SNPs, classify different SBS models, and mutation profiles. Results: A total of 21 genomes were sequenced (H37Rv: 7, 17332: 7 and 18129: 7), including one control for each, 1X and 3X exposure for both NO and PRT. ANI showed ~0.2% differences between exposed and non-exposed clinical strains under NO and PRT, with clinical isolates clustering together and separately from control Mtb H37Rv. Mutational analysis showed H37Rv exhibited only 11 and 18 mutations under NO and PRT after 3X exposure, which was ~10x less compared to clinical isolates. The KatG mutant 17332 showed 2.9-fold and 1.2-fold higher mutation burden than the KatG wt 18129 under the same conditions. Both clinical isolates exhibited stronger G>A and C>T transitions (Ts/Tv: ~1.66) after 1X NO exposure, while 3X led to increase in transversions (Ts/Tv: ~0.69) that indicate DNA damage, probably associated with oxidative stress or chemical exposure. In the case of PRT, katG wt showed a shift from transitions to transversions with increasing exposure, whereas Ts/Tv decreased from 1.39 to 0.73 after 1X to 3X exposure. katG mutant showed only a minor change in Ts/Tv (0.84 to 0.98) after 1X to 3X exposure. Furthermore, protein-level analysis revealed that the PE and PPE protein families were dominant, but some other key proteins involved in metabolism, cellular biosynthesis, and drug tolerance (e.g., aftC, qcrB, glpK) were also detected under both exposure conditions. Conclusion: Clinical isolates exhibited a higher mutation rate, particularly the KatG mutant strain, and showed strain-specific mutation patterns after exposure to PRT and NO.
Integration of computational and functional screening reveals novel R-loop–inducing drugs for triple-negative breast cancer treatment
Roberto Dinami1*, Ludovica Bonanni1, Eleonora Petti1, Pasquale Sibilio1,2, Serena Di Vito1, Manuela Porru1, Pasquale Zizza1, Sara Iachettini1, Angela Rizzo1, Carmen Maresca1, Carmen D'Angelo1, Annamaria Biroccio1 1 Translational Oncology Research Unit, IRCCS-Regina Elena National Cancer Institute, Rome, Italy. 2 Current address: Department of Wellbeing, Health and Environmental Sustainability - BESSA, Sapienza University of Rome, 02100 Rieti, Italy
Integration of computational and functional screening reveals novel R-loop–inducing drugs for triple-negative breast cancer treatment
Roberto Dinami1*, Ludovica Bonanni1, Eleonora Petti1, Pasquale Sibilio1,2, Serena Di Vito1, Manuela Porru1, Pasquale Zizza1, Sara Iachettini1, Angela Rizzo1, Carmen Maresca1, Carmen D'Angelo1, Annamaria Biroccio1 1 Translational Oncology Research Unit, IRCCS-Regina Elena National Cancer Institute, Rome, Italy. 2 Current address: Department of Wellbeing, Health and Environmental Sustainability - BESSA, Sapienza University of Rome, 02100 Rieti, Italy
- Keywords
- Triple-Negative Breast CancerR-loopsRNA-DNA hybridsTargeted Therapydrugs repurposing
- Abstract
- Background Breast cancer (BC) is the most common cancer in women. Among all BCs, triple-negative breast cancer (TNBC) subtype still represents an unmet clinical need due to its high rate of metastasis and very poor prognosis. TNBC is a heterogeneous disease characterized by high genomic instability (GI), where the formation of R-loop structures represents an additional source of GI in BC. R-loops are transient three stranded nucleic acid structures formed when the newly transcribed nascent RNA anneals back to its template DNA forming RNA-DNA hybrid structures. Maintenance of R-loop balance is allowed by different classes of proteins including Topoisomerases and RNA-DNA helicases. Although they are involved in various physiological processes, disruption of R-loop homeostasis can lead to their aberrant accumulation causing Transcription-Replication Conflicts (TRCs), DNA damage and consequent cell death. Despite the growing interest in R-loop biology, no FDA-approved drugs specifically developed to induce toxic R-loop accumulation in cancer cells are currently available, although some anticancer agents, such as PARP inhibitors, have been shown to modulate R-loop levels. Objective: Identification of FDA-approved drugs capable of inducing aberrant R-loop accumulation and cell death in tumor cells characterized by high transcription/replication stress and GI. Methods: SaveRUNNER, a recently developed network-based algorithm for drug repurposing, has been used to predict R-loops to drugs association, by quantifying the interactions of R-loops-associated genes and drug targets in the human interactome. The list of 1837 FDA approved drugs’ targets were obtained from DrugBank, whereas the list of R-loops associated genes were retrived from R-loopBase database. Candidate drugs were initially screened for antiproliferative activity using an automated live-cell imaging (Incucyte). Selected drugs were subsequently evaluated by immunofluorescence using anti-RPA32pSer33 and S9.6 antibodies, as well as by dot blot analysis. Results: Here by in-silico screening we identified several candidate drugs from distinct pharmacological classes. Of note, we found PARP and DNA topoisomerase inhibitors, two categories already reported to modulate R-loops, supporting the specificity of our approach. Interestingly, Functional screening revealed that antimicrobacterial agents were among the most effective drugs in inhibiting proliferation of TNBC cell models. By immunofluorescence staining and dot blot we found that candidate drugs, including Moxifloxacine, vitamin B1 (Tiamina) and anti-reumatic drug (Auranofin), were able to increase the aberrant levels of R-loops upon treatment. To validate the direct effects on drug-induced cell growth inhibition and R-loop accumulation, rescue experiments based on RNaseH1 overexpression are currently ongoing. In addition, molecular mechanisms induced by the most interesting drugs will be studied. Conclusion: Using an integrated bioinformatic and experimental drug-repurposing strategy, we identified novel FDA-approved compounds capable of inducing aberrant R-loop accumulation and promoting TRCs, leading to cell death in aggressive and therapy-resistant TNBC cells. These findings highlight R-loop dysregulation as a promising therapeutic vulnerability for TNBC treatment.
Liquid biopsy-based multi-omics data integration for triple-negative breast cancer
Susan, Costantini*, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Palmina, Bagnara, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Alessandra, Calabrese, Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Carolina, Manzo, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Carlo, Vitagliano, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Carmela, Dello Russo, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Chiara, Argenziano, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Federica, Renza, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Tania, Moccia, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Raffaella, Di Monda, Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Claudia, Von Arx, Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Patrizia, Vici, UOSD Sperimentazioni di fase IV, IRCCS Istituto Nazionale Tumori Regina Elena, Roma, Italy Michelino, De Laurentiis, Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Alfredo, Budillon, Scientific Directorate, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Elena, Di Gennaro, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy
Liquid biopsy-based multi-omics data integration for triple-negative breast cancer
Susan, Costantini*, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Palmina, Bagnara, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Alessandra, Calabrese, Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Carolina, Manzo, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Carlo, Vitagliano, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Carmela, Dello Russo, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Chiara, Argenziano, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Federica, Renza, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Tania, Moccia, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Raffaella, Di Monda, Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Claudia, Von Arx, Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Patrizia, Vici, UOSD Sperimentazioni di fase IV, IRCCS Istituto Nazionale Tumori Regina Elena, Roma, Italy Michelino, De Laurentiis, Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Alfredo, Budillon, Scientific Directorate, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy Elena, Di Gennaro, Experimental Pharmacology Unit, Istituto Nazionale Tumori-IRCCS-Fondazione G. Pascale, Napoli, Italy
- Keywords
- data integrationmetabolomemiRnome
- Abstract
- Objective Plasma, as a "liquid biopsy" source, represents a minimally invasive sample for analyzing systemic and localized changes occurring within the body. Unlike tissue biopsies, which are limited by spatial heterogeneity, plasma contains a rich reservoir of circulating biomarkers, including miRNAs, proteins, and metabolites, which reflect the real-time status of the tumor and its microenvironment. Despite the enormous amount of data available, the real challenge lies in data integration. Analyzing each individual omic level provides only a fragmented picture. Therefore, multi-omic data integration has emerged as a revolutionary approach, synthesizing information from the genome, miRnome, proteome, and metabolome to provide a holistic view of biological systems. This integration is crucial to the advancement of precision medicine, enabling the development of personalized therapeutic strategies tailored to the specific molecular signature of each patient. Methods We are performing on plasma samples from triple negative breast cancer (TNBC) patients prior to any systemic therapy enrolled in a multicenter prospective observational study (NCT05817227) both metabolomics analysis by Nuclear Magnetic Resonance (NMR) (Bruker AVANCETM 600 MHz), and evaluations of protein and miRNA expression by ELISA assays and real time PCR QUANTSTUDIO 7 PRO, respectively. The associations between the significant metabolite, protein and miRNA levels and some parameters such as lifestyle (smoking and BMI), and comorbidity were analysed by logistic regression model, also adjusted by age and gender as categorical covariates. Moreover, hierarchical Clustering Heatmaps were performed to highlight correlations between significant metabolites, protein and miRNAs. Results The evaluation of the metabolomic profile and the expression levels of VCP/p97-related proteins, identified metabolites and proteins differentially expressed between TNBC patients and the healthy controls. Specifically, data analysis revealed higher levels of expression of VCP/p97, GPX1, GPX4, and TXNRD1 proteins, and lower levels of SELENOP in TNBC patients compared to the healthy group. The metabolomic profile was also evaluated on the same samples, revealing higher levels of 3-hydroxybutyrate, arginine, ATP, cysteine, citrulline, phosphocholine, glutamate, hydroxyproline, proline, and trimethylamine in TNBC patients, and lower levels of glucose, isoleucine, histidine, and leucine. We are currently also evaluating the expression of a panel of miRNAs related to TNBC in the same plasma samples. By integrating these data with clinical information, we found significant associations between VCP/p97 and SELENOP levels and Ki-67, as well as notable correlations between specific metabolite levels and lifestyle-related factors such as BMI and smoking. Conclusions Integrative approaches combining metabolomic profiling with the investigation of key regulators of stress responses like VCP/p97-associated pathways may provide new insights into TNBC biology. In the coming months, miRnome profiles will be assessed, and the integrated multi-omics data will be leveraged to identify potential signatures predictive of disease progression and patient outcome.
Single-Molecule Tissue-of-Origin Profiling for Prenatal and Oncology cfDNA Applications
Casper, Lumby, aitiologic GmbH Lukas, Lüftinger, aitiologic GmbH Olga, Frank, aitiologic GmbH Andreas, Posch, aitiologic GmbH Stephan, Beisken, aitiologic GmbH*
Single-Molecule Tissue-of-Origin Profiling for Prenatal and Oncology cfDNA Applications
Casper, Lumby, aitiologic GmbH Lukas, Lüftinger, aitiologic GmbH Olga, Frank, aitiologic GmbH Andreas, Posch, aitiologic GmbH Stephan, Beisken, aitiologic GmbH*
- Keywords
- cfDNAliquid biopsybayes classifier
- Abstract
- Background Circulating cfDNA is a heterogeneous mixture derived from multiple tissues, limiting the ability to pinpoint the biological source of observed molecular alterations. Tissue-of-origin (TOO) resolution of individual cfDNA fragments is therefore critical to enables more sensitive and interpretable molecular readouts than aggregate analyses, e.g., for prenatal testing by distinguishing fetal-derived signals from maternal contributions. Methods A single-molecule, tunable, Bayesian TOO classification framework was developed on 70 paired buffy coat and placenta samples, integrating epigenetic and fragmentomic modalities to identify placental cfDNA at defined, application-specific probability cut-offs. We validated the framework on 125 plasma samples with several million fetal- and maternal-specific fragments identified via SNP genotyping in matched tissue samples. Results The classifier’s overall ROC AUC is 0.75 on the total validation set and has been shown to increase for more targeted cfDNA fragment subgroups, e.g., with increasing lengths or CpG density. We demonstrate that enrichment of fetal fraction in cfDNA can be achieved in relation to the operator threshold, enabling us to adapt the classifier to scenarios spanning low-coverage, high-purity to high-coverage, low-purity cfDNA samples. These may include prenatal testing or tissue-specific mutation profiling in cancer.
Federated Real World Data, Zero Exchange Framework for Transportability
Soyean, Kim, Providence Health Care* , Jeong Eun Min, Providence Health Research Institute
Federated Real World Data, Zero Exchange Framework for Transportability
Soyean, Kim, Providence Health Care* , Jeong Eun Min, Providence Health Research Institute
- Keywords
- TransportabilityFederated learningclinical trialReal World Evidence (RWE)
- Abstract
- A decentralized framework for synthesizing regulatory grade real-world evidence from heterogeneous geographies must reconcile the need for cross-population transportability with the strict demands of statistical validity and multi-layered data security. To enable a global federated framework, our approach reconciles statistical techniques, like high-dimensional propensity matching, with robust federated learning system architecture. Specifically, we utilize an audit-ready gRPC protocol to facilitate the secure, real-time exchange of the model gradients and aggregates required for cross-border collaboration. Our approach ensures that the 'borrowing' of real-world data is governed by hardened security protocols that prevent data leakage while maintaining the mathematical accuracy required for regulatory evidence. Our system enables cross-border participant matching by calculating propensity scores from high-dimensional EHR data using highly structured medical codes such as ICD-10. By leveraging federated analytics using cross border real-world evidence, we can identify clinically similar patients in different geographies while keeping sensitive health records within their home jurisdictions. The core stages of secure computation include Local Processing (Input Party) where the raw Input Data remains exclusively within the perimeter of the data holder, Updating Transmission where TensorFlow Federated (TFF) transmits only non-personal data such as model updates (e.g., model gradients) and aggregate statistics, and Secure Aggregation where the central aggregator (the Output Party) receives these protected updates and combines them to refine a global model. By employing a high-dimensional propensity score (hdPS) for transportability, the system identifies thousands of pre-exposure variables to proxy for unobserved confounding. Because the algorithm operates without requiring access to patient-level data, it significantly streamlines the compliance process while maintaining high standards of data privacy. This method leverages external data (e.g., Korean EHR data for a Canadian trial) to improve statistical power and precision, offering a pathway to streamline clinical trial designs and lower the necessary patient enrollment.
AEGLOS V7: Hyperbolic Digital Pathology for Morphology-to-Genomics Prediction in Advanced Non-Small Cell Lung Cancer
Lorenzo, Nicolè, Pathology Unit, Angelo’s Hospital, Mestre, Venice, Italy Flavio, Sartori, AI and Computational Biomedicine Unit, Department of Medical Sciences, University of Torino, Torino, Italy Ivan, Rossi, Biodec SRL, Casalecchio di Reno, Bologna, Italy Tiziana, Sanavia, AI and Computational Biomedicine Unit, Department of Medical Sciences, University of Torino, Torino, Italy
AEGLOS V7: Hyperbolic Digital Pathology for Morphology-to-Genomics Prediction in Advanced Non-Small Cell Lung Cancer
Lorenzo, Nicolè, Pathology Unit, Angelo’s Hospital, Mestre, Venice, Italy Flavio, Sartori, AI and Computational Biomedicine Unit, Department of Medical Sciences, University of Torino, Torino, Italy Ivan, Rossi, Biodec SRL, Casalecchio di Reno, Bologna, Italy Tiziana, Sanavia, AI and Computational Biomedicine Unit, Department of Medical Sciences, University of Torino, Torino, Italy
- Keywords
- Digital PathologyHyperbolic LearningFoundation ModelsWhole-Slide ImagingPrecision Oncology
- Abstract
- Objective Morphological patterns associated with clinically actionable genomic alterations have been reported in haematoxylin-and-eosin whole-slide images from advanced non-small cell lung cancer (NSCLC), but their reproducibility across real-world cohorts remains unclear. We developed AEGLOS V7, a hyperbolic digital pathology framework designed to evaluate morphology-to-genomics transferability across independent cohorts at both pathway and actionable-target levels. The aim of this study was to assess the reproducibility of morphology-associated genomic signals across independent cohorts. Methods Diagnostic whole-slide images were processed using a multi-resolution computational pathology pipeline and encoded using multiple pathology foundation-model backbones, including Dinov2, Virchow2 and Gigapath. AEGLOS V7 projects slide-level representations into a hyperbolic latent space and jointly models two genomic layers: biological pathways and actionable therapeutic groups. The internal development cohort included 409 patients from The Cancer Genome Atlas with valid feature tensors and was evaluated using stratified 5-fold cross-validation. Actionable targets included EGFR sensitizing mutations, EGFR exon 20 insertions, KRAS G12C, BRAF V600E, ALK, ROS1, RET, NTRK and MET exon 14 alterations. Pathway targets included RTK/RAS/MAPK, PI3K/AKT/mTOR, P53/cell-cycle, WNT/TGF-beta, DNA-repair/telomere, immuno-metabolism, epigenetic/IDH and hormone/GPCR axes. External validation was performed on an independent real-world hospital cohort comprising 420 slide-derived tensors. Performance was evaluated using area under the precision-recall curve (AUPRC) and lift over target prevalence. Results Internal stratified 5-fold cross-validation on TCGA identified the strongest overall performance among Gigapath-based configurations, with mean pathway AUPRCs of 0.40 ± 0.02 and actionable-target AUPRCs of 0.12 ± 0.02. External validation confirmed heterogeneous but biologically interpretable transferability across cohorts. The strongest actionable enrichment was observed for ALK fusions, achieving an AUPRC of 0.14 and a 6.30-fold lift over prevalence. The most enriched pathway-level signals included WNT/TGF-beta status (AUPRC 0.11; 3.73-fold lift) and hormone/GPCR pathway status (AUPRC 0.10; 2.49-fold lift). Broader pathway alterations, including P53/cell-cycle and RTK/RAS/MAPK, achieved higher absolute AUPRCs but lower enrichment due to higher baseline prevalence. Conclusions AEGLOS V7 introduces a hyperbolic digital pathology framework for comparative morphology-genomics analysis with uncertainty-aware model selection and external validation across heterogeneous datasets. External validation demonstrated that selected actionable alterations and biological pathways may generate transferable morphology-associated signals detectable across independent real-world NSCLC cohorts using routine histology alone. These preliminary findings provide a baseline for future development of robust morphology-genomics approaches in advanced NSCLC, with ongoing work focused on calibration, additional hospital cohorts and improved characterization of target-specific transferability.
Learning Read-Level joint genomic and epigenomic status representations from Nanopore sequencing data using Jesica-Fetcher and NTv3-like architecture
*Tommaso Ducci Institute of Informatics and Telematics, National Research Council, Pisa, Department of Medical Biotechnologies, University of Siena; Monia Taranta Institute of Clinical Physiology, National Research Council, Siena Mario Chiariello Institute of Clinical Physiology, National Research Council, Siena CRL-ISPRO, Siena; Elia Giuseppe Ceroni Institute of Informatics and Telematics, National Research Council, Pisa; Romina D’Aurizio Institute of Informatics and Telematics, National Research Council, Pisa;
Learning Read-Level joint genomic and epigenomic status representations from Nanopore sequencing data using Jesica-Fetcher and NTv3-like architecture
*Tommaso Ducci Institute of Informatics and Telematics, National Research Council, Pisa, Department of Medical Biotechnologies, University of Siena; Monia Taranta Institute of Clinical Physiology, National Research Council, Siena Mario Chiariello Institute of Clinical Physiology, National Research Council, Siena CRL-ISPRO, Siena; Elia Giuseppe Ceroni Institute of Informatics and Telematics, National Research Council, Pisa; Romina D’Aurizio Institute of Informatics and Telematics, National Research Council, Pisa;
- Keywords
- Representation LearningNanopore sequencingDNA MethylationCancer
- Abstract
- Introduction: In recent years, DNA language models have emerged as a powerful approach for learning meaningful representations of genomic sequences. Existing models usually rely on curated multi-species datasets of DNA sequences, still excluding the exploration of learning capability in the context of experimental data from different biological conditions. Furthermore, few deep learning (DL) models learn joint representations of genomic and epigenomic status, with the exception of MethylBERT, which limits the investigation of fundamental mechanisms, particularly in cancer. We aimed to develop an efficient data-fetching framework for training DL models directly from experimental BAM files, and used it to train a model to learn representations of ONT reads. Methods: We sequenced the whole-genome of 4 glioblastoma cell lines (GSC11,TS576,RG01,12o89) by ONT (~32X). Bins of read were uniformly sampled from the BAMs with splitting policy: 0.81 train, 0.09 validation, 0.1 test. We designed a lightweight augmented BAM-index for lazy reads data loading from BAM files. The index stores compact read-level metadata along with the information of reads binning in intervals of the required size. The BAM together with its corresponding augmented index are used to build an iterable dataset that leverages targeted “pysam” fetch operations. A 30M-parameter model based on the NTv3 architecture was developed and trained using this read fetching method. The model used a 10-token vocabulary composed of 5 tokens representing both modified and unmodified bases, one token encoding strand orientation, and 4 additional special tokens. Training was performed using BERT-style masked language modeling and cross-entropy loss for masked token prediction for 100k steps on 8 NVIDIA A100 GPUs, with batches of 512 sequences of up to 4096 bases, for a total of nearly 210B processed tokens and 12 hours of training time. Results: A lightweight BAM-native read loader was developed which enables online streaming of reads directly from compressed BAM files with negligible RAM usage and minimal storage overhead, as the augmented index occupies roughly one tenth of the original BAM file size, without training time overhead. Under this framework, a customized NTv3 architecture was trained and tested on 4 GBM cell lines measuring its capability to learn informative joint genetic and epigenetic representations from ONT reads. Cross-entropy loss reached 0.88 on training, 0.89 on validation, and 0.91 on test sets, for masked tokens prediction. Conclusions: Here we present Jesica-Fetcher, the first data-fetching framework capable of directly streaming compressed raw ONT sequencing reads from BAM files to downstream DL models. Through it, we trained a methylation-aware model, reaching stable convergence. Given the pivotal role of coordinated genotype–methylation programs underlying cellular behavior, this new approach and the learned representations could provide a strong foundation for a wide range of downstream applications.
Benchmarking LLMs for automated systems biology model replication
Danilo, Tomasoni*, Fondazione The Microsoft Research – University of Trento Centre for Computational and Systems Biology (COSBI), Rovereto, Italy Marco, Zardo, Department of Mathematics, University of Trento, Povo, Italy Mario, Lauria, Department of Mathematics, University of Trento, Povo, Italy
Benchmarking LLMs for automated systems biology model replication
Danilo, Tomasoni*, Fondazione The Microsoft Research – University of Trento Centre for Computational and Systems Biology (COSBI), Rovereto, Italy Marco, Zardo, Department of Mathematics, University of Trento, Povo, Italy Mario, Lauria, Department of Mathematics, University of Trento, Povo, Italy
- Keywords
- AIPrompt EngineeringAgentic PipelineReflection PatternSystematic EvaluationSystems Biology
- Abstract
- Objective: Model reproducibility is a relevant issue in systems biology, mainly due to incomplete or missing information in the manuscript. Despite some attempts to address the problem with guidelines, reproducibility remains a major challenge that prevents re-using and extending published models. For this reason, it will be beneficial to devise automated methods to generate executable models from manuscripts and evaluate their compliance with reported outputs. Methods: We implemented a tool that automatically generates a formal mathematical model from a manuscript describing it. It is composed of an agentic reflection pipeline where a translation agent generates a mathematical model that is then simulated. Simulation errors are then handled by a critique agent, that reviews and corrects the model based on simulation errors, resulting in a final model that is compared against an expert-curated gold standard. We systematically evaluate the generated model quality across 80 articles whose reproducibility was manually assessed by experts of the field and for which an implementation in SBML is available. The SBML model is then automatically converted to the equivalent Antimony language, that was used to compare the performance of 6 mainstream LLMs: DeepSeek, Gemini 2.5 and 3.0 pro, ChatGPT 5.0, 5.2 and 5.2-pro. Results: Automated LLM reproducibility score was defined as the percentage of reproduced models for which agreement with published results was achieved, where agreement is determined by at least one simulated time series attaining an Absolute Average Fold Error (AAFE) < 2. Each manuscript was evaluated three times with identical LLM models and parameters to assess LLM generation stochasticity. We found that: 1) the top-performing gemini-3.0-pro generated an executable model 97% of the times (min=95%, max=100%) but reproduced accurately (AAFE<2) only 34.5% of the models (min=32%, max=37%) 2) models from the same provider improved in newer versions (OpenAI=+4%, Google=+9.3%), 3) lowering the temperature and top-p values of LLMs consistently improved performance (+1.8%) and reduced stochasticity (−0.7%) across tested LLMs, but the effect was not statistically significant (exact permutation test, p > 0.25), 4) supplementary materials contain crucial information for reproducing published results. We observed a relevant average decrease in reproducibility when supplementary materials were not accessible: from 23.3% to 8.7% across all LLMs (min=6.6% on DeepSeek, max=20% on gemini-2.5-pro). Finally, our evaluation dataset together with a set of key performance metrics can be used as a baseline benchmark against which future methods can be evaluated. Conclusion: To the best of our knowledge, this is the first systematic evaluation of LLM performance in the context of automated systems biology model replication, providing guidance on LLM model choice and parameter tuning as well as an evaluation framework that can be used to rigorously evaluate future agentic AI systems in the field.
HIC-Sutra: Inferring 3D genome organization from condition-specific transcriptomic signals
Umesh*, Bhati, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Academy of Scientific and Innovative Research (AcSIR), Ghaziabad-201002, India Ronit, Roy, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Academy of Scientific and Innovative Research (AcSIR), Ghaziabad-201002, India Ravi, Shankar, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Academy of Scientific and Innovative Research (AcSIR), Ghaziabad-201002, India
HIC-Sutra: Inferring 3D genome organization from condition-specific transcriptomic signals
Umesh*, Bhati, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Academy of Scientific and Innovative Research (AcSIR), Ghaziabad-201002, India Ronit, Roy, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Academy of Scientific and Innovative Research (AcSIR), Ghaziabad-201002, India Ravi, Shankar, CSIR-Institute of Himalayan Bioresource Technology (CSIR-IHBT), Academy of Scientific and Innovative Research (AcSIR), Ghaziabad-201002, India
- Keywords
- Chromatin interactionsCis-regulatory elementsDeep LearningTransformerTAD boundaries
- Abstract
- Objective Three-dimensional chromatin architecture is a primary determinant of cell-type identity, gene regulation, and disease aetiology. Despite the central importance of DNA–DNA interactions (DDIs) and topologically associating domain (TAD) boundaries in orchestrating transcriptional programmes, no existing deep learning framework simultaneously predicts Hi-C interactions in a cell-type-specific and condition-specific manner, generalises to previously unseen cell types, and provides statistically rigorous confidence estimates for differential interaction analysis. Current state-of-the-art models such as Akita, EPCOT, C.Origami, and ChromaFold are constrained by the requirement for Hi-C training data from every queried cell type and lack principled uncertainty quantification limiting their translational utility in developmental and disease contexts. Methods We developed HIC-Sutra, a two-stage attention-based deep ensemble architecture for cell-type-specific chromatin interaction prediction from RNA-seq, ATAC-seq, and DNA methylation data alone. In Stage I, a convolutional neural network (CNN) extracts transcription factor binding features for CTCF, YY1, and SP1 from one-hot-encoded DNA sequence combined with ATAC-seq chromatin accessibility signals and CpG methylation profiles, enabling epigenomic context-aware feature representation. In Stage II, a Transformer-based encoder integrates these learned embeddings to predict Hi-C interaction frequencies between accessible genomic loci up to 2 Mb apart, capturing long-range regulatory contacts with sequence-level resolution. Uncertainty was decomposed into aleatoric (data-inherent noise) and epistemic (model uncertainty) components through stochastic training and a deep ensemble of 10 independently trained models. HIC-Sutra was trained and rigorously evaluated on bulk ATAC-seq/Hi-C datasets from multiple mouse cell lines and independently validated on human cell lines, with additional adaptation and benchmarking against pseudo-bulk single-cell ATAC-seq (scATAC-seq) data to demonstrate single-cell scalability. Results HIC-Sutra achieved Spearman correlation coefficients exceeding 0.90 across all held-out test cell lines in bulk data settings and faithfully recapitulated distance-dependent interaction decay patterns characteristic of in situ Hi-C. The model consistently outperformed Akita, EPCOT, C.Origami, and ChromaFold across the majority of benchmarking metrics, particularly for unseen cell types on which existing methods exhibit substantial performance degradation. Critically, HIC-Sutra accurately resolved biologically interpretable differential interactions, including: disruption of TAD boundary insulation upon CTCF motif inversion; allele-specific chromatin contact rewiring at the colorectal cancer susceptibility variant rs1800734; dynamic chromatin reorganisation during LPS-mediated macrophage activation; and differential 3D genome architecture between differentiated and undifferentiated esophageal adenocarcinoma cell states each validated against orthogonal experimental evidence. Conclusion HIC-Sutra establishes a lightweight, generalizable, and uncertainty-aware paradigm for in silico prediction of cell-type-specific and differential chromatin interactions using only multi-omic profiling data, eliminating the experimental burden of generating Hi-C libraries for every biological condition. By quantifying both aleatoric and epistemic uncertainty, HIC-Sutra enables statistically principled identification of high-confidence interaction changes across developmental trajectories, disease transitions, and genetic perturbations. This framework offers a powerful computational resource for probing chromatin architecture in contexts where experimental Hi-C is cost-prohibitive, opening avenues for large-scale regulatory genomics, variant interpretation, and cancer epigenome research.
Large-scale analysis and clustering of RNA 3D structures
*Wojciech, Rączka, Poznan University of Technology Tomasz, Żok, Poznan University of Technology Jan, Pielesiak, Poznan University of Technology Maciej, Antczak, Poznan University of Technology Marta, Szachniuk, Poznan University of Technology
Large-scale analysis and clustering of RNA 3D structures
*Wojciech, Rączka, Poznan University of Technology Tomasz, Żok, Poznan University of Technology Jan, Pielesiak, Poznan University of Technology Maciej, Antczak, Poznan University of Technology Marta, Szachniuk, Poznan University of Technology
- Keywords
- RNA 3D structuresClusteringProcessing pipelineMachine Learning
- Abstract
- Multiloop is a motif inside an RNA structure where three or more helices connect. They present a problem for bioinformatics due to their rarity and simultaneously high impact on the overall RNA structural topology. Objective: Our team aims to further study Multiloop motifs, allowing researchers in turn to better understand RNA molecules and predict their 3D structures more effectively. To achieve this, we are working on tools creating datasets dedicated to RNA structures with multiloops motifs. With these datasets, we are using machine learning in order to cluster RNA’s secondary structures in hopes of gaining a better understanding of this subject. Methods: Two main objectives of this project are to develop a pipeline allowing fast processing of mmCIF (Macromolecular Crystallographic Information File) files through RNA analysis tools and to create a secondary structure clustering algorithm. To achieve this, we store RNA datasets obtained from EMBL-EBI, extract entities important to our research from preferred assemblies and then process them through secondary structure analysis tools. This allows us to have reliable access to the results and detect potential discrepancies between them with different tools used. Currently, we are working on deciding which clustering method would be best suited for detecting similarities between observed motifs found in the RNA secondary structures. To do this, we are clustering the dot-bracket representations of the secondary structures divided in groups by their length. This prior classification is necessary in order to remove the dot-bracket length from impacting clustering results. The metric used in clustering methods is distance (differences) between secondary structures. Distance is calculated using Tree Edit Distance (TED) algorithm between dot-bracket representations. Results: Currently we have clustering results for smaller scale dot-bracket samples as well as a working processing pipeline. Conclusion: Using this clustering algorithm we aim to obtain information about redundancy, the exclusion of which is necessary to reliably train machine learning models. In future we aim to train models allowing us to easily detect coaxial stacking in 3D structures and predict coaxial stacking from sequences and 2D structures. This work is supported by the National Science Centre, Poland (grant no. 2023/51/D/ST6/01207).
Proteogenomic insights into complex human diseases
Yuan Hu *, European Institute for the Biology of Aging, University Medical Center Groningen, Groningen, the Netherlands Yanick Paco Hagemeijer, Department of Analytical Biochemistry, University of Groningen, Groningen, the Netherlands Peter Horvatovich, Department of Analytical Biochemistry, University of Groningen, Groningen, the Netherlands Victor Guryev, European Institute for the Biology of Aging, University Medical Center Groningen, Groningen, the Netherlands Gyorgy Halmos, Department of Otolaryngology, Head and neck Surgery, University Medical Center Groningen, Groningen, the Netherlands
Proteogenomic insights into complex human diseases
Yuan Hu *, European Institute for the Biology of Aging, University Medical Center Groningen, Groningen, the Netherlands Yanick Paco Hagemeijer, Department of Analytical Biochemistry, University of Groningen, Groningen, the Netherlands Peter Horvatovich, Department of Analytical Biochemistry, University of Groningen, Groningen, the Netherlands Victor Guryev, European Institute for the Biology of Aging, University Medical Center Groningen, Groningen, the Netherlands Gyorgy Halmos, Department of Otolaryngology, Head and neck Surgery, University Medical Center Groningen, Groningen, the Netherlands
- Keywords
- Proteogenomics; transcriptomics; proteomics; laryngeal squamous cell carcinoma; aging
- Abstract
- Objective Although standardized treatment protocols exist for laryngeal squamous cell carcinoma (LSCC), current clinical guidelines do not explicitly account for age-related differences in disease manifestation. This study aims to characterize the proteogenomic landscape of LSCC across different age groups and to explore the molecular interplay between aging and tumorigenesis. Methods A proteogenomic analysis was conducted using paired tumor and adjacent control samples from younger (<65 years) and older (>75 years) LSCC patients. Multi-omics data integration was performed by combining mass spectrometry–based proteomics with transcriptomic information. Differential expression analysis, variant detection, and identification of non-reference peptides were carried out, followed by functional annotation and pathway enrichment analyses to elucidate biological processes associated with disease and aging. Results LSCC in older patients shared core molecular characteristics with younger cases, while also exhibiting distinct age-associated alterations. A total of seventy-one genes were concordantly dysregulated in both groups, primarily associated with extracellular matrix (ECM) remodeling and ethanol metabolism. In addition, rare variants and non-reference peptides specific to LSCC were identified across samples, revealing previously uncharacterized molecular features. Conversely, thirty-five genes demonstrated age-dependent expression patterns, showing either synergistic or antagonistic interactions between aging and oncogenesis. Conclusion This study provides a comprehensive proteogenomic characterization of LSCC across age groups and highlights the complex interaction between aging and cancer biology. The findings offer a molecular basis for understanding age-related heterogeneity in LSCC and support the development of age-adapted therapeutic strategies.
MoonLitDB: An LLM-powered database annotating moonlighting proteins from scientific literature
Davide Gotta*, Department of Biochemical Sciences "A. Rossi Fanelli", Sapienza University of Rome, Rome, Italy Allegra Via, Department of Biochemical Sciences "A. Rossi Fanelli", Sapienza University of Rome, Rome, Italy Teresa Colombo, Istituto di Biologia e Patologia Molecolari, Consiglio Nazionale delle Ricerche, Rome, Italy
MoonLitDB: An LLM-powered database annotating moonlighting proteins from scientific literature
Davide Gotta*, Department of Biochemical Sciences "A. Rossi Fanelli", Sapienza University of Rome, Rome, Italy Allegra Via, Department of Biochemical Sciences "A. Rossi Fanelli", Sapienza University of Rome, Rome, Italy Teresa Colombo, Istituto di Biologia e Patologia Molecolari, Consiglio Nazionale delle Ricerche, Rome, Italy
- Keywords
- Moonlighting proteinsliterature mininglarge language models
- Abstract
- **Objective** Moonlighting proteins are single polypeptides that perform multiple, mechanistically distinct functions, challenging the classical "one gene, one protein, one function" paradigm. Their multifunctionality offers insight into context-dependent regulation and the evolution of new protein functions, while their dysregulation has been implicated in cancer, infection, and other pathologies. However, current knowledge remains scattered across the primary literature, and existing resources compiling moonlighting proteins are sparse and rarely document the mechanisms underlying functional switching. This work aimed to build a comprehensive, transparent resource that systematically captures moonlighting functions together with their supporting evidence and switching mechanisms. **Methods** We developed MoonLitDB, a resource that mines full-text literature for moonlighting functions using local large language models with structured output, linking every extracted function to a supporting source sentence for transparency and verifiability. The pipeline separates canonical from putative moonlighting roles, maps functions to Gene Ontology (GO) slim terms, normalizes subcellular localizations using controlled vocabularies, and resolves gene identifiers using established entity-normalization tools. Mechanisms of functional switching are explicitly recorded for each annotation if available. The extracted data are organized into a web interface enabling both gene-level and function-level exploration. **Results** MoonLitDB hosts data extracted from a large corpus of publications, encompassing a substantial number of protein-function annotations that span both moonlighting and canonical roles across a wide range of unique genes. By systematizing this information, MoonLitDB recovers the great majority of previously known moonlighting examples while surfacing numerous novel candidates that are absent from current databases. Users can query the resource by single gene, gene list, or advanced Boolean search, and can move fluidly between gene-centric and function-centric cards. All results are downloadable in CSV, JSON, or spreadsheet formats to support downstream analysis. **Conclusion** MoonLitDB offers a systematic, transparent, and continuously updated platform for exploring moonlighting protein functions and the mechanisms driving functional switching. By coupling automated literature mining with rigorous normalization and sentence-level evidence linking, it fills a significant gap in the field and provides a scalable foundation for further discovery. The underlying mining pipeline will be progressively enhanced—for instance, through semantic search—to broaden coverage and accelerate future research into protein multifunctionality. MoonLitDB is freely available at https://bioinfo.bio.uniroma1.it/moonlitdb.
Hypothesis-driven integrative transcriptomics reveals distinct Z-disc gene signatures in ischemic and non-ischemic cardiomyopathy
Steve, Solun, Faculty of Engineering, Department of Digital Medical Technologies, Holon Institute of Technology (HIT), Holon, Israel David, Luria, Heart Institute, Hadassah-Hebrew University Medical Center, Jerusalem, Israel Yulia, Einav *, Faculty of Engineering, Department of Digital Medical Technologies, Holon Institute of Technology (HIT), Holon, Israel
Hypothesis-driven integrative transcriptomics reveals distinct Z-disc gene signatures in ischemic and non-ischemic cardiomyopathy
Steve, Solun, Faculty of Engineering, Department of Digital Medical Technologies, Holon Institute of Technology (HIT), Holon, Israel David, Luria, Heart Institute, Hadassah-Hebrew University Medical Center, Jerusalem, Israel Yulia, Einav *, Faculty of Engineering, Department of Digital Medical Technologies, Holon Institute of Technology (HIT), Holon, Israel
- Keywords
- integrative transcriptomicscardiomyopathyZ-disc genesmeta-analysismachine learning
- Abstract
- Objective Publicly archived cardiac transcriptomic datasets are individually underpowered and generated across heterogeneous microarray platforms, making it difficult to detect subtle, reproducible expression changes in functionally defined gene sets. Conventional genome-wide differential expression workflows compound this problem by distributing statistical power across tens of thousands of genes. We aimed to develop and demonstrate a hypothesis-driven "reverse" integrative framework that inverts this paradigm — beginning with a biologically curated gene panel rather than an agnostic screen — and to apply it to characterise the transcriptional landscape of ten Z-disc-associated genes across ischemic (ICM) and non-ischemic cardiomyopathy (NICM). Methods Six publicly archived human left-ventricular microarray datasets (GEO accessions GSE1145, GSE1869, GSE9128, GSE3585, GSE3586, GSE42955; n = 142: 45 controls, 25 ICM, 72 NICM; three Affymetrix platforms) were integrated via ComBat batch correction into a unified expression matrix of 3,878 common genes. Targeted differential expression analysis was performed with limma (empirical Bayes moderation; Benjamini–Hochberg FDR correction). Cross-study consistency was quantified by random-effects meta-analysis (Hedges' g, REML estimation; metafor). Subtype discrimination was assessed using an elastic-net logistic regression classifier evaluated in a leave-one-dataset-out (LODO) cross-validation framework (AUC via bootstrap resampling, 2,000 iterations). Differential transcriptional variability between subtypes was assessed with the Brown–Forsythe test, and principal component analysis was used to visualise subtype separation at both global and panel-focused resolution. Results In ICM, filamin C (FLNC) was the sole gene satisfying both significance and fold-change thresholds (log₂FC = +0.503, p = 0.005), while six additional genes reached nominal significance. In NICM, PDLIM3 was significantly downregulated (log₂FC = −0.450, FDR = 0.002) in a subtype-specific manner (p = 0.649 in ICM). Random-effects meta-analysis confirmed FHOD3 as the ischemic-specific gene with consistent cross-study effects (Hedges' g = 0.793, I² = 0%), and identified MYOZ2, PDLIM3, and TTN as significant in NICM. The elastic-net classifier achieved a pooled AUC of 0.649 (95% CI: 0.501–0.776), with MYOZ2 selected in 83% of LODO folds. Strikingly, NICM displayed marked transcriptional homogeneity across Z-disc genes (IQR range: 0.13–0.36), while ICM exhibited pronounced heterogeneity (IQR range: 7.87–10.12), a contrast confirmed by Brown–Forsythe differential variability analysis. Conclusion The hypothesis-driven reverse integrative workflow successfully extracted reproducible, subtype-specific molecular signals from small, cross-platform datasets where genome-wide approaches would lack power. The framework is biologically interpretable by design, computationally lightweight, and directly transferable to any functionally coherent gene panel in other diseases. These findings position FLNC and FHOD3 as candidate ischemic Z-disc markers and PDLIM3 as a NICM-specific target, while the ICM/NICM variability asymmetry suggests fundamentally different transcriptional regulation between cardiomyopathy subtypes.
Whole-Genome Sequencing Reveals Novel De Novo and Familial Variants Underlying Syndromic and Isolated Congenital Heart Defects in Tunisian Patients
Rim Khelifi 1 2 3* : 1 Unit of Common Services of Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Tunisia. 2 Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia. 3 University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany Andreas Rump3 : University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany. Amira Benzarti 12 : 1 Unit of Common Services of Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Tunisia. 2 Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia Amilcar Perez Riverol 3 : University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany. Rim Kooli 12 : Unit of Common Services of Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Tunisia. 2 Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia Rafiga Masmaliyeva 3 : University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany Enrique Audin Martinez 3 : University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany Ajmi Houda 4 : Paediatric Department University Hospital Sahloul Sousse, Tunisia, Faten Yahia 5 : Cardiology Department, University Hospital Sahloul Sousse, Tunisia. Sahbi Ghanem 6 : Neonatology Department, Tahar Sfar Hospital, Mahdia, Tunisia, Essia Sboui 7 : Paediatric Department, Regional Hospital Ibn El Jazzar, Kairouan, Tunisia Elies Neffati 5 :Cardiology Department, University Hospital Sahloul Sousse, Tunisia. Moez Gribaa 2 : Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia Marc Phillip Hitz 3: University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany Soumaya Mougu-Zerelli 1 2: 1 Unit of Common Services of Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Tunisia. 2 Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia
Whole-Genome Sequencing Reveals Novel De Novo and Familial Variants Underlying Syndromic and Isolated Congenital Heart Defects in Tunisian Patients
Rim Khelifi 1 2 3* : 1 Unit of Common Services of Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Tunisia. 2 Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia. 3 University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany Andreas Rump3 : University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany. Amira Benzarti 12 : 1 Unit of Common Services of Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Tunisia. 2 Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia Amilcar Perez Riverol 3 : University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany. Rim Kooli 12 : Unit of Common Services of Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Tunisia. 2 Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia Rafiga Masmaliyeva 3 : University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany Enrique Audin Martinez 3 : University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany Ajmi Houda 4 : Paediatric Department University Hospital Sahloul Sousse, Tunisia, Faten Yahia 5 : Cardiology Department, University Hospital Sahloul Sousse, Tunisia. Sahbi Ghanem 6 : Neonatology Department, Tahar Sfar Hospital, Mahdia, Tunisia, Essia Sboui 7 : Paediatric Department, Regional Hospital Ibn El Jazzar, Kairouan, Tunisia Elies Neffati 5 :Cardiology Department, University Hospital Sahloul Sousse, Tunisia. Moez Gribaa 2 : Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia Marc Phillip Hitz 3: University Institute for Medical Genetics Carl Von Ossietzky University Oldenburg, Germany Soumaya Mougu-Zerelli 1 2: 1 Unit of Common Services of Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Tunisia. 2 Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, University Hospital Farhat Hached Sousse, Tunisia
- Keywords
- Corpus callosum malformationsRare diseasesGenomic diagnosticsCopy number variantsBioinformatics
- Abstract
- Background: Congenital heart defects (CHDs) are the most common congenital anomalies worldwide and a leading cause of infant morbidity and mortality. Despite advances in genomic medicine, the genetic architecture of CHDs remains insufficiently characterized in North African populations, where high consanguinity rates may contribute to unique patterns of genetic variation. This study employed whole-genome sequencing (WGS) to investigate the molecular basis of syndromic and familial CHDs in Tunisian patients, aiming to identify pathogenic variants, assess diagnostic yield, and provide population-specific genomic insights. Materials and Methods: High-coverage WGS (30×, NovaSeq 6000) was performed on 66 individuals, including 25 syndromic and non-syndromic CHD index cases analyzed through trio/duo-based approaches and eight members of a consanguineous family presenting an isolated atrial septal defect (ASD). A comprehensive bioinformatics workflow was applied for variant calling, annotation, filtering, and prioritization. Variants were interpreted according to ACMG guidelines, while inheritance patterns and penetrance were evaluated using in silico analyses. Results: Three unrelated patients presenting classical CHARGE syndrome features harbored distinct pathogenic variants in CHD7, including a frameshift mutation (c.921_922delAG; p.Gly308fs) and two novel nonsense variants (c.1480C>T; p.Arg494* and c.2149G>T; p.Gly717*). All occurred de novo, representing a 19% prevalence of CHD7-associated CHARGE syndrome within the cohort. A patient diagnosed with Coffin–Siris syndrome carried a novel nonsense variant in ARID1B (c.3693C>G; p.Tyr1231*), while a patient with tetralogy of Fallot and neurodevelopmental impairment harbored a novel missense variant in CTCF (c.1414T>G; p.Cys472Gly). These findings expanded the mutational spectrum of syndromic CHDs and yielded a molecular diagnosis in 20% of cases, exclusively attributable to de novo pathogenic variants. In the familial ASD pedigree, affected siblings shared maternally inherited cis variants in MYH6 (c.1576G>A; p.Glu526Lys) and MYH7 (c.5779A>T; p.Ile1927Phe), whereas the mother remained asymptomatic, suggesting incomplete penetrance and the involvement of additional genetic modifiers. Conclusions: WGS demonstrated substantial diagnostic value for CHD patients from an underrepresented North African population, identifying novel de novo and familial variants while expanding the known genetic landscape of CHDs. However, approximately 80% of cases remained unresolved, highlighting the complexity of CHD genetics. Future integration of transcriptomic profiling and DRAGEN-based genomic analyses may facilitate the detection of regulatory, non-coding, and splicing alterations, improving variant interpretation and supporting precision medicine strategies for congenital heart disease.
Integrated genomic analysis of corpus callosum malformations in a Tunisian cohort reveals extensive genetic heterogeneity and improves genotype–phenotype characterization
Bochra Khadija*, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Higher Institute of Biotechnology of Monastir, University of Monastir, Monastir, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia Hamza Hadj Abdallah, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Wafa Slimani, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Ayda Bennour, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia Sarra Dimassi, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia Najla Soyah, Pediatrics Department, CHU Farhat Hached, Sousse, Tunisia Amira Benzarti, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Khouloud Rjiba, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Wafa Dahleb, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Molka Kammoun, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Hanen Hannechi, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Hela Ben Khelifa, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Neziha Gouider Khouja, Neuropediatrics Department, National Institute of Neurology of Tunis, Tunis, Tunisia Lamia Boughamoura, Pediatrics Department, CHU Farhat Hached, Sousse, Tunisia Ichrak Kraoua, Neuropediatrics Department, National Institute of Neurology of Tunis, Tunis, Tunisia Chahnaz Triki, Neuropediatrics Department, Hedi Chaker University Hospital, Sfax, Tunisia Amel Tej, Pediatrics Department, CHU Farhat Hached, Sousse, Tunisia Jihen Mathlouthi, Neonatology Department, CHU Farhat Hached, Sousse, Tunisia Aida Guith, Neonatology Department, CHU Farhat Hached, Sousse, Tunisia Saoussen Abroug, Pediatrics Department, Sahloul University Hospital, Sousse, Tunisia Raoudha Kebaili, Pediatrics Department, CHU Farhat Hached, Sousse, Tunisia Mohamed Ali Bouaziz, Specialist in Pediatrics, Tunis, Tunisia Habib Soua, Neonatology Department, Tahar Sfar Hospital, Mahdia, Tunisia Essia Sboui, Pediatrics Department, Ibn El Jazzar Hospital, Kairouan, Tunisia Samir Hadded, Pediatrics Department, Bechir Hamza Children's Hospital, Tunis, Tunisia Randa Ziadi, Pediatrics Department, CHU Fattouma Bourguiba, Monastir, Tunisia Kamel Monastiri, Neonatal Medicine Department, Monastir University Hospital, Monastir, Tunisia Hayet Ben Hamida, Neonatology Department, Farhat Hached University Hospital, Sousse, Tunisia Monji Ghanmi, Neonatology Department, Tahar Sfar Hospital, Mahdia, Tunisia Sayda Hassayoun, Specialist in Pediatrics and Neonatology, Clinique Les Oliviers, Sousse, Tunisia Mohamed Tahar Sfar, Pediatrics Department, Tahar Sfar Hospital, Mahdia, Tunisia Ali Saad, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia Christel Depienne, Institute of Human Genetics, University Hospital Essen, University of Duisburg-Essen, Essen, Germany Soumaya Mougou-Zerelli, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia
Integrated genomic analysis of corpus callosum malformations in a Tunisian cohort reveals extensive genetic heterogeneity and improves genotype–phenotype characterization
Bochra Khadija*, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Higher Institute of Biotechnology of Monastir, University of Monastir, Monastir, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia Hamza Hadj Abdallah, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Wafa Slimani, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Ayda Bennour, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia Sarra Dimassi, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia Najla Soyah, Pediatrics Department, CHU Farhat Hached, Sousse, Tunisia Amira Benzarti, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Khouloud Rjiba, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Wafa Dahleb, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Molka Kammoun, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Hanen Hannechi, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Hela Ben Khelifa, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia Neziha Gouider Khouja, Neuropediatrics Department, National Institute of Neurology of Tunis, Tunis, Tunisia Lamia Boughamoura, Pediatrics Department, CHU Farhat Hached, Sousse, Tunisia Ichrak Kraoua, Neuropediatrics Department, National Institute of Neurology of Tunis, Tunis, Tunisia Chahnaz Triki, Neuropediatrics Department, Hedi Chaker University Hospital, Sfax, Tunisia Amel Tej, Pediatrics Department, CHU Farhat Hached, Sousse, Tunisia Jihen Mathlouthi, Neonatology Department, CHU Farhat Hached, Sousse, Tunisia Aida Guith, Neonatology Department, CHU Farhat Hached, Sousse, Tunisia Saoussen Abroug, Pediatrics Department, Sahloul University Hospital, Sousse, Tunisia Raoudha Kebaili, Pediatrics Department, CHU Farhat Hached, Sousse, Tunisia Mohamed Ali Bouaziz, Specialist in Pediatrics, Tunis, Tunisia Habib Soua, Neonatology Department, Tahar Sfar Hospital, Mahdia, Tunisia Essia Sboui, Pediatrics Department, Ibn El Jazzar Hospital, Kairouan, Tunisia Samir Hadded, Pediatrics Department, Bechir Hamza Children's Hospital, Tunis, Tunisia Randa Ziadi, Pediatrics Department, CHU Fattouma Bourguiba, Monastir, Tunisia Kamel Monastiri, Neonatal Medicine Department, Monastir University Hospital, Monastir, Tunisia Hayet Ben Hamida, Neonatology Department, Farhat Hached University Hospital, Sousse, Tunisia Monji Ghanmi, Neonatology Department, Tahar Sfar Hospital, Mahdia, Tunisia Sayda Hassayoun, Specialist in Pediatrics and Neonatology, Clinique Les Oliviers, Sousse, Tunisia Mohamed Tahar Sfar, Pediatrics Department, Tahar Sfar Hospital, Mahdia, Tunisia Ali Saad, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia Christel Depienne, Institute of Human Genetics, University Hospital Essen, University of Duisburg-Essen, Essen, Germany Soumaya Mougou-Zerelli, Laboratory of Cytogenetics, Molecular Genetics and Human Reproductive Biology, CHU Farhat Hached, Sousse, Tunisia; Common Service Unit for Research in Genetics, Faculty of Medicine of Sousse, University of Sousse, Sousse, Tunisia
- Keywords
- Moonlighting proteinsliterature mininglarge language models
- Abstract
- Objective Corpus callosum malformations (CCM) are among the most common congenital brain abnormalities and represent a clinically and genetically heterogeneous group of neurodevelopmental disorders. Although numerous genetic causes have been reported, the molecular basis of CCM remains incompletely understood, particularly in underrepresented populations. This study aimed to characterize the genomic landscape of CCM in a large Tunisian cohort and evaluate the contribution of an integrated genomic strategy to diagnostic yield and genotype–phenotype correlation. Methods We investigated 107 Tunisian patients with CCM referred for genetic evaluation between 2010 and 2025. Clinical data, neuroimaging findings, and associated cerebral and extracerebral anomalies were systematically reviewed. A stepwise genomic workflow combining conventional cytogenetics (karyotyping), molecular cytogenetic techniques (FISH, MLPA, array-CGH), and next-generation sequencing approaches (whole-exome sequencing and whole-genome sequencing) was implemented. Detected variants were interpreted according to current clinical guidelines, and genotype–phenotype relationships were assessed through integrated analysis of copy number variants (CNVs) and sequence variants. Results A molecular diagnosis was achieved in 49.53% of patients, demonstrating the effectiveness of combining CNV detection and sequencing-based approaches. Pathogenic CNVs were identified in recurrent genomic regions, including deletions involving 14q12, 4p, 1q43q44, 1p32, 5p14.3, 5q35, 16q, and 18p, as well as duplications involving 4q, 2p23.3, and isochromosome 18p. Several individuals exhibited complex genomic architectures characterized by the coexistence of pathogenic CNVs and additional variants of uncertain significance. Sequencing analyses identified pathogenic or likely pathogenic variants in established CCM-associated genes, including RTTN, SMPD4, MAPK8IP3, AMPD2, EPG5, and TUBB3. Furthermore, potentially relevant variants were detected in candidate genes such as VPS39, TBC1D2, MSH3, PPP1R21, CWC27, MOCS1, and RNASEH2C, expanding the spectrum of genes potentially implicated in CCM pathogenesis. Integrated genomic analysis improved diagnostic resolution and facilitated more precise genotype–phenotype interpretation. Conclusion This study provides one of the largest genomic investigations of CCM in a North African population and highlights the substantial genetic heterogeneity underlying these disorders. The results demonstrate the complementarity of CNV analysis and high-throughput sequencing for maximizing diagnostic yield and refining genotype–phenotype correlations. Our findings support a complex genetic architecture for CCM and underscore the importance of integrative bioinformatics-driven genomic approaches for rare disease diagnosis, genetic counseling, and future functional investigations in underrepresented populations.
