Menu

Interactive Program

Interactive Program

Search the programme

Room

Track

Day 1

Wednesday 2 September

Opening Remarks

Aula Magna
08:45 - 09:15

Opening Scientific Introduction

Aula Magna
09:15 - 09:30

BIOINF Main Track 1

Aula MagnaChair: Davide Chicco
09:30 - 10:30
PaperBIOINF

An Enhanced Pipeline for Multi-Omic Integration Based on Topological Data Analysis

Veronica Paparozzi, Anna Plaksienko, Marco Pedicini, Christine Nardini

Abstract

The advent of high-throughput sequencing techniques have made it essential to employ ad- vanced tools for data integration and interpretation which represent an active area of research. In this work, we propose to build on two existing efforts to enlarge their scope and usabil- ity in multi-omic integration. We exploit Topological Data Analysis (TDA), capitalizing on its graph-based framework, to refine complex interactions characteristic of biological systems; on its intrinsic explainability to enable interpretation of model outputs in terms of biologically meaningful features; and on its robustness to small input perturbations, to ensure stable results in the presence of noisy data, such as omics. Specifically, we utilize a recent tool in TDA, called Harmonic Persistent Homology (HPH), and propose two relevant advances. First, we adopt an integration and clustering tool (i.e. iNETgrate) as an informed means to reduce data size and integrate multiple omic layers by aggregating methylation beta values from loci to the gene level, to be combined with gene expression values. This enables the application of HPH, otherwise limited by the computational burden. Second, as a consequence, we apply HPH to multi-omic molecular data, and not patients, enlarging HPH scope. Finally, we follow up on recent efforts toward performance standardization, and validate our results against biomarkers from a well-known breast cancer benchmark dataset from TGCA.

PaperBIOINF

Analysis of biological networks using Krylov subspace trajectories

H. Robert Frost

Abstract

We describe an approach for analyzing biological networks using rows of the Krylov subspace of the adjacency matrix. Specifically, we explore the scenario where the Krylov sub-space matrix is computed via power iteration using a non-random and potentially non-uniform initial vector that captures a specific biological state or perturbation. In this case, the rows the Krylov subspace matrix (i.e., Krylov trajectories) carry important functional information about the network nodes in the biological context represented by the initial vector. We demonstrate theu tility of this approach for community detection and perturbation analysis using the C. Elegans neural network.

PaperBIOINF

Evaluating SMT-Based Formal Inference for GRN Synthesis: From Data Density to Structural Rules

Ofri Caspi, Michal Greenberg, Eitan Tannenbaum, Hillel Kugler

Abstract

Synthesizing Gene Regulatory Networks (GRNs) from dynamic observations is a fundamental systems biology challenge, essential for deciphering the regulatory mechanisms governing cellular behavior. This work explores a formal inference strategy based on Satisfiability Modulo Theories (SMT) and the Reasoning Engine for Interaction Networks (RE:IN) to systematically identify all networks consistent with biological constraints and trajectory data. The research investigates two dimensions. First, we evaluate the capability of model reconstruction and, separately, examine the impact of observation volume and resolution on the number of consistent solutions (RQ1). Evaluation across 20 models from the Biodivine Boolean Models (BBM) repository shows a 100% recovery rate with five experimental observations, decreasing to 85% with ten observations. This decrease is due to an expressiveness limitation: RE:IN’s regulation condition templates cannot represent the complex or non-monotonic logic of certain biological update functions. Analysis of a subset of models also shows that increasing experimental volume and timestep resolution effectively reduces the solution space. Second, we characterize the formal identification of structural interactions (RQ2) to determine which are logically necessary to explain observed dynamics. We derived and validated six deterministic rules that characterize, within single-activator motifs, the conditions under which interactions are classified as structurally required or disallowed based on topology and trajectory alone. In conclusion, this work evaluates the reliability and exposes the limitations of SMT-based formal inference for genetic network synthesis. It provides a framework for predicting required or disallowed interactions under single-activator motifs, thereby bridging the gap between complex logical inference and biological understanding.

PaperBIOINF

A Pipeline for Predicting Variant Associations with Complex Disorders via Functional Annotations

Francesco Gualdi, Zeno Darani, Daniele Malpetti, Marco Scutari, Francesca Mangili

Abstract

The study of complex disorders is limited by the lack of structured data sets that integrate genetic association signals with rich functional annotations. While genome-wide association studies (GWAS) provide large collections of associated variants, these are heterogeneous, studydependent, and not directly suitable for machine learning applications. Although frameworks such as Combined Annotation Dependent Depletion (CADD) offer comprehensive variant-level features, there is no standardised procedure to construct curated, balanced, and reproducible data sets for the systematic study of complex traits. We introduce a scalable and reproducible pipeline for generating CADD-annotated variant data sets associated with complex disorders defined via Experimental Factor Ontology (EFO) terms. The pipeline integrates GWAS summary statistics, genome mapping, chromosomestratified negative sampling, and large-scale annotation, while addressing class imbalance and feature redundancy. The resulting data sets provide curated, balanced, and annotation-rich representations of variant–trait associations, suitable for statistical learning, systematic model evaluation and comparison, and analysis of functional annotation contributions and interactions.

Coffee Break

Terrazza del RettoratoChair: Antonello Maruotti
10:30 - 11:00

A Dependent Dirichlet Process Approach to Population Size Estimation

Lucia Gallucci

Abstract

Estimating the size of partially observed populations is a central problem in capture--recapture studies, where only a subset of individuals is detected in the available data. Classical approaches rely on strong assumptions such as population closure and homogeneous capture probabilities, which are often unrealistic in applications and limit their practical usefulness. We propose a flexible Bayesian nonparametric framework for population size estimation under heterogeneous detection. Building on the model of Guindani et al. (2014), we extend the standard capture-recapture formulation by incorporating individual-level covariates through a Dependent Dirichlet Process (DDP) construction. This allows capture probabilities to vary smoothly with observed characteristics, while retaining the flexibility of nonparametric mixture models. Moreover, our approach explicitly incorporates the covariates probability distribution into the DDP model, allowing both cluster allocation and capture intensity to enable a richer representation of heterogeneity. Moreover, the joint structure of the model allows covariate information to be coherently propagated to unobserved individuals, supporting principled inference on their latent covariate profiles and improving estimation of the total population size (n). We develop an efficient Markov chain Monte Carlo algorithm that combines a collapsed Chinese Restaurant Process representation with reversible-jump Metropolis--Hastings updates for the unknown population size. Despite the increased model flexibility, conjugacy is preserved wherever possible, leading to efficient posterior computation and scalable inference. We illustrate the proposed methodology in the context of estimating the prevalence of Autism Spectrum Disorder (ASD) using hospital discharge records in Italy, an inherently incomplete data source, with the aim of estimating the true prevalence.

Parallel sessions

11:00 - 12:00
Aula Gini
Poster session

SS1: AI and Computational Methods for Medical Informatics - Part I

Museo dell'Arte Classica

PaperMEDINF

Diagnosing Rejection Collapse via Uncertainty Decomposition

Walter Endrizzi, ...

PaperMEDINF

ClinRAG-BiLSTM: An Agentic Cloud AI Framework for Explainable and Traceable Breast Cancer Decision Support

Adaleta Gicic, ...

PaperMEDINF

Towards Federated and Explainable Quantum Machine Learning for Possible Epileptic Seizure Detection

Francesco Mercaldo, ...

Aula Magna
PaperBIOINF

Toward Simulation-Informed Lipid Nanoparticle Design with Martini Coarse-Grained Models

Mariana Valério, ...

PaperBIOINF

GPU-accelerated single-cell and spatial transcriptomics on NVIDIA DGX H100: a systematic benchmark of speed, scalability, and biological concordance

Luca Vedovelli, ...

PaperBIOINF

CudaMon: An R Package to Monitor NVIDIA GPUs, Showcased by Monitoring a GPU-accelerated Single-cell Analysis Workflow in R

Mohammad Amin Zadenoori, ...

PaperBIOINF

DualBioGraph: Patient-Held-Out Tumor--Normal Classification Using Dual Graph Neural Networks

Al Hamna Asif, ...

Aula III
PaperBIOINF

A topological object-proposal pipeline for cellular fiber segmentation

Riccardo Ceccaroni, ...

PaperBIOSTAT

Quantum Machine Learning for Missense Variant Pathogenicity Prediction in the KCNQ Gene Family

Markel Garcia, ...

PaperBIOSTAT

Zero-shot phenotype prediction and screening of engineered immune receptor variants using contrastive learning-based workflow

Jie Shi, ...

PaperBIOSTAT

A Landmarked Penalised Weibull Framework for Time-Updated Survival Modelling of Heterogeneous Clinical Outcomes in Multiple Sclerosis

Francisco José Aparicio Serrano, ...

Parallel sessions

12:00 - 13:15
Aula III
PaperBIOINF

CellChat Hotspot: A Focused Lens on Tumor Microenvironment Communication

Dario Monaco, ...

PaperBIOSTAT

Evaluating in-context learning with prompting regimes for transformer-based synthetic health tabular data generation

Amanda Bertgren, ...

PaperBIOSTAT

Multivariate Model-Based Landmarking Approach in Mixture Cure Models to Dynamic Prediction

Bianca Ferraro, ...

PaperBIOSTAT

Functional Fuzzy Clustering of Longitudinal biological marker's trajectories

Silvia D'Elia, ...

PaperBIOSTAT

Reducing the Size of Breastfeeding Reference Arms in Infant Nutrition Studies Using Prognostic Covariate Adjustment

Luca Lavalle

Aula Gini
PaperMEDINF

Predicting Alzheimer’s Progression Over Time Using Sheaf Neural Networks and Brain Graphs

Annamaria Defilippo, ...

PaperMEDINF

Ontology-Aware Candidate Reranking for SNOMED CT Entity Linking in Clinical Notes

Luka Blašković, ...

PaperMEDINF

Evaluating Locally-Deployable Large Language Models on Free-Text Data in an Italian Obstetric Context

Pierluigi Reali, ...

PaperMEDINF

Decoding HIV-1 Antibody Escape with Interpretable Protein Language Models

Fahsai Nakarin, ...

PaperMEDINF

AutoEncoder based approach for the identification of genomic regions responsible for poorly described oligo-patient disorders

Joanna Szyda

Aula Magna
PaperMEDINF

Prediction of treatment response in Multiple Sclerosis with Machine Learning models

Ariadna Masot-Llima, ...

PaperBIOINF

Secondary structure design for efficient translation initiation

Tobias von der Haar, ...

PaperBIOINF

Manufacturability of mRNA Therapeutics: Native vs. N-1-Methyl-Pseudouridine

Roland Huber, ...

PaperBIOINF

Optimizing RNA yield using deep neural networks coupled to massively parallel screening

Adrien Villain

PaperBIOINF

MAP-Net: An Interpretable Deep Learning Framework for Universal mRNA Manufacturability

Giorgio Ciano

Lunch

Terrazza del Rettorato
13:15 - 14:15

Computational metagenomics to unravel person-to-person microbiome transmission

Francesco Asnicar

BIOINF Main Track 2

Aula MagnaChair: Graziano Pesole
14:15 - 16:00
PaperBIOINF

Sequence-Derived Physicochemical Features and Stacked Learning for SUMOylation-Site Prediction

Junyan Li, Qiaobin Yao, Danda Rawat, Jiang Li, Shaolei Teng

Abstract

SUMOylation regulates protein stability, localization, and interaction networks, but computational prediction of SUMO-acceptor lysines remains difficult because many true sites do not strictly follow the canonical Ψ-K-x-E motif. We developed a sequence-based prediction framework that systematically evaluates AAindex-derived physicochemical descriptors for human SUMOylation-site prediction. The dataset comprised 5,888 experimentally verified SUMOylated lysines and 415,918 unlabeled lysines from 176 human proteins collected from dbPTM and UniProt. For each candidate lysine in sumoylated proteins, we generated lysine-centered windows of multiple lengths and encoded each window with 566 AAindex properties. We first ranked individual AAindex features using single-feature XGBoost performance on the full dataset, then benchmarked several classifiers (logistic regression, SVM, random forest, and XGBoost) across window sizes. Based on these results, we trained a XGBoost model and a stacked ensemble (XGBoost + random forest + logistic regression with a logistic regression meta-learner) using a balanced training subset. The best setting used an 11 (±5 from center lysines)-residue window and ~350 top-ranked features. Under this configuration, the stacked ensemble achieved an AUC-ROC of 0.9426 on the test set, substantially outperforming the best single model. Motif-level analysis confirmed the presence of the canonical Ψ-K-x-E consensus, but the strongest predictive signals arose from flexible, coil-promoting sequence regions flanking the lysine. Overall, our results show that rich physicochemical encoding combined with stacking can markedly improve SUMOylation-site prediction and helps highlight property patterns that are most informative around sumoylated lysines.

PaperBIOINF

LLM-Guided Integer Linear Programming for Antibody Library Design

Mikel Landajuela Larma

Abstract

Designing a library of antibody variants requires choosing which mutations to introduce while balancing competing goals such as binding-related scores and human-likeness. We study whether a large language model (LLM) can help a combinatorial solver make better choices. In the proposed pipeline, the LLM reads a compact summary of the antibody--antigen structure (derived from a Protein Data Bank entry) together with per-mutation scores, and returns strictly structured hints: mutations to avoid, mutations to suggest, avoid-combinations, and additive preference scores. These hints are mechanically translated into linear constraints and objective coefficients of an integer linear program (ILP), which then returns a diverse batch of designs. The design decouples reasoning (handled by the LLM, which is good at high-level priors but bad at satisfying many interacting constraints at once) from search (handled by the ILP, which guarantees feasibility and can be solved to global optimality for each scalarization). We evaluate on two antibodies, trastuzumab and D44.1, and report multi-objective diversity and quality metrics; for trastuzumab we additionally report the success rate of a learned oracle classifier. LLM-guided ILP improves the strongest previous baseline on every reported metric while reliably delivering a full library of 1{,}000 unique sequences; direct LLM generation, in contrast, under-produces and collapses in diversity. Improvements are consistent across both antibody targets and across LLM backbones (gpt-4o, gpt-5), showing that the structural priors surfaced by the LLM translate directly into better designs. A key practical benefit of the approach is that those priors are emitted as strictly structured JSON, so domain experts can inspect, edit, or veto individual constraints before the ILP is re-run --- a level of auditability that end-to-end neural generators do not offer.

PaperBIOINF

Cell Senescence Identification Using Single-Cell Transcriptomics Foundation Model

Yongqi Zhao, Fan Tong, Xiangwen Zheng, Yuanyuan Ma, Huajian Mao, Yuting Zhou, Peixiang Yang, Dongsheng Zhao

Abstract

Background: In recent years, single-cell foundation models have achieved good results in tasks such as cell type annotation, and how to use single-cell foundation models to identify senescent cells is a scientific question worth investigating. Data and Methods: This study selects 4 outstanding single-cell foundation models including scGPT, CellPLM, tGPT and UCE, combining multilayer perceptron as simple downstream networks, and uses 10 consensus senescence datasets e.g. GSE119807, GSE102090, and GSE94980 to train (80%) and test (20%) their performance on the task of cell senescence identification, preliminarily exploring the performance on their AUC evaluation. At the same time, using approaches such as decrease of training dataset, dimensionality reduction clustering analysis, and simplification of downstream models, the ability of the foundation model's cell embeddings to distinguish cellular ageing is evaluated. Result: The experimental results show that, except for a very few datasets, the constructed model generally achieved an AUC value of 99%, higher than the traditional domain models - SenCID (mean 95%) and hUSI (mean 92%), and comparable to DeepScence. When the training data is reduced to 60% or the downstream model is simplified to logistic regression, the AUC only slightly decreases (the mean AUC still exceeds 98%), and dimensionality reduction clustering analysis of embedding representations indicates that the foundation model has a strong ability to distinguish cell condition. Conclusion: The single-cell foundation model gains intrinsic knowledge related to cellular ageing through pre-training on massive sequencing data without expert feature engineering, providing a competitive new approach for research in cell senescence identification.

PaperBIOINF

OmicsGPT: Multi-Agent Bioinformatics Orchestration for Evidence-Grounded RNA-seq Interpretation

Patrick Roney, James Li

Abstract

Large language models can summarize oncology evidence, but unsupported stan- dalone use is vulnerable to hallucinated druggability, acronym collisions, and failure to respect tissue and assay context. We present OmicsGPT, a LangGraph-based multi-agent prototype that converts RNA-seq counts, metadata, optional context files, and a disease-specific prompt into an auditable oncology interpretation report. The workflow applies PyDESeq2/DESeq2-style differential-expression analysis and preranked GSEA, then routes selected targets through a planner-executor architecture querying PubMed, Open Targets, OncoKB, STRING, and GTEx- style tissue context. Role-specific agents flag lineage-mismatch artifacts, distinguish mutation- level from expression-level actionability, and filter acronym collisions before a report-writer agent synthesizes a source-traced report with evidence-bounded follow-up. As a public software- architecture test, we used GEO series GSE164641, an RNA-seq dataset profiling normal breast tissue from women at high or average risk for breast cancer. In a representative breast-risk run, the target roster included OLAH and BTN1A1; OmicsGPT treated these RNA-level find- ings as biomarker-hypothesis signals rather than direct therapeutic recommendations. Com- pared with a standalone LLM baseline, OmicsGPT preserved count-model statistics, exposed database provenance, separated RNA expression from DNA-level actionability, and retained weak or negative evidence when no direct claim was supported. The current system is intended for research interpretation rather than direct clinical decision support. Its main contribution is an interpretable glass-box architecture for evidence-grounded bioinformatics analysis and on- cology hypothesis generation.

PaperBIOINF

Identifying functional drivers of Hepatoblastoma outcomes via agent-based modeling and transcriptomics

Alessandro Ravoni, Yuanhua Liu, Stefano Cairo, Filippo Castiglione, Christine Nardini

Abstract

Hepatoblastoma (HB) is the most common pediatric liver cancer and represents a major clinical challenge, due to the lack of effective therapies for advanced stages and disease relapse. In this work, we use the results of a previously HB-tailored agent-based model of the immune system to investigate whether model-derived variables can be of use in the prediction of patients’ outcomes. To this aim, we apply factor analysis to the results of a simulated cohort of HB patients, to identify combinations of key immunological variables able to discriminate disease outcomes in the simulator, and we then assess the coherence of such predictions with independent results of differential expression and enrichment analyses on HB transcriptomics. Our analysis proposes that the ability of immune cells, particularly natural killer and CD8+ cytotoxic T cells, to recognize tumor-associated antigens and exert cytotoxic activity is essential for disease control following treatment.

Keynote — Francesco Asnicar

Aula Magna
16:00 - 16:45

Poster Session A / Coffee Break

Museo dell'Arte ClassicaChair: Davide Chicco
16:45 - 19:00

ELIXIR-IT: the Italian Research Infrastructure for Big Data in Life Sciences

Graziano Pesole

Welcome Cocktail + MUSA Jazz Concert

Terrazza del RettoratoChair: Annamaria Carissimo
19:00 - 20:00

AI-assisted Single-Cell transcriptomics reveals persistent malignant CD4+ activity in Sézary Syndrome patient

Domenico Palumbo, Viola Melone, Luigi Palo, Carlo Ferravante, Giulia Salvatore, Cristina Cristofoletti, Maria Grazia Narducci, Roberta Tarallo

Abstract

Sézary syndrome (SS) is a rare and aggressive leukemic variant of cutaneous T-cell lymphoma characterized by erythroderma, lymphadenopathy, and circulating malignant CD4+ T cells. Despite several therapeutic advances such as extracorporeal photopheresis (ECP), the disease remains difficult to treat. Here, we applied single-cell RNA sequencing to investigate transcriptional changes in CD4+ T cells from an SS patient undergoing ECP at three timepoints (T1: 2 cycles, T2: 20 cycles, T3: 74 cycles). Exploiting the Illumina Single Cell 3’ RNA Prep Kit, we analyzed output data by using a custom pipeline with an AI-based annotation tool. Across all timepoints, CD4+ T cells were predominant, with an increasing CD4+/CD8+ ratio, indicating persistent disease spread. The AI sub-clustering identified a single active CD4+ population, which was selected and analyzed for differential gene expression. We identified 42, 82, and 151 differentially expressed genes in T1-T2, T2-T3, and T1-T3 comparisons respectively, indicating progressive transcriptional changes. Pathway analysis revealed a good activation of the inflammatory response, even under ECP therapy. These results suggest that, in this patient, ECPs were not enough to prevent transcriptional evolution of malignant CD4+ T cells in SS, showing an ongoing immune activation and disease progression. For this reason, a perturbation analysis was performed to highlight possible drugs to be used in combination with ECP to maximize its effect. This study demonstrates the utility of single-cell transcriptomics and AI-based annotation for monitoring therapeutic response in CD4+ cell population, causative of the pathological state, and understanding the SS pathophysiology.