Pietro H. Guzzi

dblp:60/161 · also Pietro Hiram Guzzi · DBLP profile ↗
← Back
93ranked-venue papers
28as first author
35since 2021 · last 2026
0000-0001-5542-2997ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 82 · 26 first-author · 32 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 11 · 3 first-authorSystems, architecture and hardware · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 The next paradigm in bioinformatics: a review of multi-agent systems and foundational models for end-to-end scientific discovery
abstract
Bioinformatics is entering a new phase characterized by the integration of universal biological models and multi-agent systems to enable end-to-end scientific discoveries. This review argues that the next paradigm shift will go beyond traditional predictive models and generative artificial intelligence (AI) toward agentic AI: systems capable of planning, acting through tools, reflecting on results, and iterating until a goal is achieved. We first examine recent foundational models that produce transferable representations across omic modalities, such as scGPT, Nicheformer, and EpiAgent, and discuss their architectural choices, training regimes, and interpretability constraints. We then analyze biomedical agent frameworks through their main components (planning, action, reflection, and memory), highlighting representative systems such as ClinicalAgent and Biomni that operationalize these ideas in controlled environments. Next, we focus on hypothesis validation mechanisms, including retrieval-augmented generation for evidence grounding, sequential statistical testing, and benchmarking methodologies designed to quantify robustness and reproducibility. Finally, we summarize emerging applications in drug discovery and personalized medicine, from molecular literature analysis and protocol automation to drug repurposing for rare diseases and closed-loop synthesis. We conclude by outlining the main challenges ahead, namely hallucinations, interpretability, systemic biases, integration with clinical infrastructures, and regulatory and ethical requirements, and propose a roadmap for the development of scientific agents that are not only high-performing but also reliable, verifiable, and implementable in real biomedical contexts.
Francesco Branda, Mohamed Mustaf Ahmed, Massimo Ciccozzi, Pietro H. Guzzi, Fabio Scarpa
Briefings Bioinform.4
2026 Bidirectional Mamba-2 boosts EEG super-resolution via regression and diffusion
abstract
MOTIVATIONS: Electroencephalography (EEG) is a non-invasive method that records brain electrical activity from scalp electrodes, offering millisecond temporal resolution but limited spatial detail due to sparse sensor layouts. RESULTS: We present DiBiMa-EEGSR, a bidirectional Mamba-2 diffusion framework for spatio-temporal EEG super-resolution that reconstructs high-resolution signals from standard low-density recordings without additional hardware. The method formulates super-resolution as conditional generative inference and integrates a diffusion process with a bidirectional state-space backbone to model long-range temporal dependencies with linear complexity. Conditioning on low-resolution inputs, electrode positions and task labels enables anatomically coherent and context-aware reconstruction. A one-step sampling strategy substantially reduces inference time while preserving fidelity. Across two public benchmarks, the approach improves reconstruction accuracy, spatial coherence and spectral preservation over convolutional, transformer-based and prior diffusion models in both spatial and temporal upsampling tasks, providing a scalable pathway toward high-resolution electrophysiological imaging. AVAILABILITY AND IMPLEMENTATION: Code to reproduce ablation experiments, training and evaluation of the proposed BiMa and DiBiMa EEGSR models are available at https://github.com/UgoLomoio/DiBiMa-EEGSR.git. Model weights are available at https://huggingface.co/Ugo96/DiBiMa-EEGSR while an interactive demo for EEG spatial super-resolution using our models can be found at https://huggingface.co/spaces/Ugo96/DiBiMa-EEGSR-Demo.
Ugo Lomoio, Pietro Liò, Pietro H. Guzzi, Pierangelo Veltri
Bioinform.3
2026 Unsupervised synchronization of molecular dynamics trajectories via graph embedding and time warping
abstract
MOTIVATION: Molecular dynamics (MD) simulations provide detailed atomistic insights into biomolecular processes, but comparing independent trajectories remains challenging due to stochastic divergence. Misaligned simulations can obscure shared mechanisms or exaggerate differences, limiting reproducibility and mechanistic interpretation. A generalizable, unsupervised method for synchronizing and comparing MD trajectories across systems and conditions is, therefore, needed. RESULTS: We introduce NetMD, a computational framework that synchronizes and analyzes MD trajectories by integrating graph-based representations with dynamic time warping. Trajectory frames are converted into residue-contact graphs, entropy-filtered to retain variable interactions, and embedded as low-dimensional vectors. NetMD aligns these vectorized trajectories through time-warping barycenter averaging, generating a consensus trajectory while pruning outlier simulations. Applied to transporters, demethylases, and large protein complexes relevant to neurological disease pathways and cancer, NetMD revealed shared multiphase dynamics and identified mutation- or ligand-specific deviations. This unsupervised, time-resolved approach enables direct comparison of MD ensembles across heterogeneous conditions. NetMD is robust and broadly applicable, providing a tool for uncovering conserved patterns and critical divergences in biomolecular dynamics. AVAILABILITY AND IMPLEMENTATION: NetMD is freely available at https://github.com/mazzalab/NetMD.
Manuel Mangoni, Salvatore Daniele Bianco, Francesco Petrizzelli, Michele Pieroni, Pietro H. Guzzi, Viviana Caputo, Tommaso Biagini, Tommaso Mazza
Bioinform.5
2026 Segmentation of temporal graphs
Raffaele Giancotti, Francesco Gullo, Pietro H. Guzzi, Edoardo Serra, Pierangelo Veltri
Inf. Sci.3
2025 GraphNet: A Novel Method Based on Graph Neural Networks for Emergency Healthcare Management
Annamaria Defilippo, Pietro H. Guzzi, Pierangelo Veltri, Pietro Liò
AIME (2)2
2025 Comparative Analysis of a Custom Lightweight LLM Versus General-Purpose LLMs for Medical Query Handling
abstract
Large language models (LLMs) have demonstrated remarkable versatility in natural language understanding and generation, yet their reliability in specialized medical domains remains uncertain. We focus on the use of lightwight and specialized LLMs instances on cardiological medical data corpora. We present a comparative evaluation of four general-purpose LLMs, i.e., ChatGPT, Gemini, Claude AI, and PerplexityAI trained on cardiology clinical data. We assess response quality on set of clinically relevant queries with a score in the [0,1] interval. Score were grouped in three classes: good (scores greater than 0.75), sufficient (scores in the interval 0.51-0.74), insufficient (scores less than 0.5). Results show that the lightweight model achieved the highest proportion of sufficient answers evaluated as sufficient (54%) and a greater share of answers evaluated as good (23%) than general-purpose counterparts, while maintaining a relatively low proportion of insufficient responses$(23 \%)$. In contrast, general-purpose models exhibited greater variability, alternating between highly accurate and critically lacking outputs, with PerplexityAI performing weakest overall. These findings suggest that targeted domain adaptation, even with lightweight architectures, can yield more stable and clinically reliable outputs than large general-purpose systems. The study underscores the potential of lightweight, specialized LLMs as trustworthy components in medical decision support frameworks, where consistency and factual grounding are paramount.
Pietro H. Guzzi, Valentina Carbonari, Giovanni Canino, Giorgia Caronzolo, Fabiola Boccuto, Salvatore De Rosa, Daniele Torella, Pierangelo Veltri
BIBM1
2025 Transformer-Based Analysis for Detecting Pulmonary Nodules in CT Scans: Preliminary Results
abstract
Using artificial intelligence (AI) offers opportunities to analyze medical images and to support early cancer detection. For instance, neural networks, in different implementation, can be used to analyze data image parts (i.e., voxels), defining a trained network useful for lung cancer nodules detection. We present our experience in designing and testing a transformerbased deep learning architecture, aiming to detect pulmonary cancer nodule candidates using 3D Computed Tomography (CT) images. The module also includes a preprocessing pipeline based on dynamic sampling of voxels extracted from images, to support data filtering and results explainability. The proposed architecture has been implemented, trained, and tested using the LUNA16 publicly available dataset. Experimental results proved both high effectiveness and competitive performance metrics across standard evaluations. Trained module can be used on a large CT dataset aiming to support clinicians in lung cancer early detection as well as to support in followup for lung cancer patients treatments. This work represent, indeed, results for preliminary applications in a research project (Advancing Lung Cancer Screening: Artificial Intelligence, Multimodal Imaging and Cutting-Edge Technologies for Early Detection and Characterization), conducted in collaboration with San Raffaele Hospital (Italy), Campus Biomedico University (Italy) and University Hospital of Salerno.
Martina De Salazar, Fatih Aksu, Raffaele Giancotti, Fabrizia Gelardi, Patrizia Vizza, Pietro H. Guzzi, Paolo Soda, Giuseppe Tradigo, Arturo Chiti, Pierangelo Veltri
BIBM6
2025 Design and use of a Denoising Convolutional Autoencoder for reconstructing electrocardiogram signals at super resolution
abstract
Electrocardiogram signals play a pivotal role in cardiovascular diagnostics, providing essential information on electrical hearth activity. However, inherent noise and limited resolution can hinder an accurate interpretation of the recordings. In this paper an advanced Denoising Convolutional Autoencoder designed to process electrocardiogram signals, generating super-resolution reconstructions is proposed; this is followed by in-depth analysis of the enhanced signals. The autoencoder receives a signal window (of 5 s) sampled at 50 Hz (low resolution) as input and reconstructs a denoised super-resolution signal at 500 Hz. The proposed autoencoder is applied to publicly available datasets, demonstrating optimal performance in reconstructing high-resolution signals from very low-resolution inputs sampled at 50 Hz. The results were then compared with current state-of-the-art for electrocardiogram super-resolution, demonstrating the effectiveness of the proposed method. The method achieves a signal-to-noise ratio of 12.20 dB, a mean squared error of 0.0044, and a root mean squared error of 4.86%, which significantly outperforms current state-of-the-art alternatives. This framework can effectively enhance hidden information within signals, aiding in the detection of heart-related diseases. • We defined a novel architecture based on autocencoders which is able to denoise and reconstruct high resolution copies of input low resolution ECG signals. • This unique approach that has not been previously applied to ECG signals. • We also present a deep validation of our approach against traditional and contemporary methods in terms of signal-to-noise ratio, mean squared error, and root mean squared error with those of other widely used ECG signal processing techniques. • The results consistently showed superior performance, further validating the effectiveness of our approach. • Given the increasing reliance on effective and efficient diagnostic techniques in medical practice, especially in cardiology, the findings of our study have significant practical implications.
Ugo Lomoio, Pierangelo Veltri, Pietro H. Guzzi, Pietro Liò
Artif. Intell. Medicine3
2025 Differential causal networks highlight sex-based differences in human tissues
abstract
Sex differences appear in healthy and pathological conditions and may influence sex-specific therapeutic responses. Understanding such differences is a key activity for developing precision medicine strategies. This study investigates sex differences in gene expression across 40 human tissues by applying a Differential Causal Network (DCN) analysis using data from the Genotype-Tissue Expression project. We identified sex-based DCNs that highlight distinct molecular mechanisms influencing both health and disease in men and women. For example, in pancreas tissue, genes associated with immune system show significant differences in their regulatory patterns between sexes, demonstrating a possible different response to diseases such as diabetes mellitus and cancer. Our findings provide valuable information on the biological underpinnings of sex differences, offering potential pathways for the development of precision medicine strategies.
Annamaria Defilippo, Kimberly Glass, Federico Manuel Giorgi, Tamer Kahveci, Pierangelo Veltri, Pietro H. Guzzi
Briefings Bioinform.6
2024 Anomaly Detection in Individual Specific Networks through Explainable Generative Adversarial Attributed Networks
abstract
Recently, the availability of many omics data source has given the rise of modelling biological networks for each individual or patient. Such networks are able to represent individual-specific characteristics, providing insights into the condition of each person. Given a set of networks of individuals, a network representing a particular condition (e.g., an individual with a specific disease) may be seen as an anomaly network. Consequently, the use of Graph Anomaly Detection techniques may support such analysis. Among the others, Generative Adversarial Networks present optimal performances in anomaly detection. This paper presents ADIN (Anomaly Detection in Individual Networks), a framework based on Generative Adversarial Attributed Networks (GAANs) for anomaly detection in convergence/divergence patients attributed networks. Preliminary results on networks generated from computational biology gene expression data demonstrate the effectiveness of our approach in detecting and explaining bladder cancer patients.
Pietro H. Guzzi, Ugo Lomoio, Tommaso Mazza, Pierangelo Veltri
BIBM1
2024 Exploring Network Curvature Differences in Gene Expression Networks
abstract
Networks and their properties have been used to study complex biological systems. Recently, network curvature measures have demonstrated the ability to capture relevant network properties. This study employs network curvature measures to analyse gene expression correlations in various human tissues for identifying unique topological features that differentiate these groups. Preliminary findings suggest that curvature measures offer novel insights that could enhance our understanding of the biological systems.
Pietro H. Guzzi, Marianna Milano
BIBM1
2024 SHELLEY: Exploring Learning-Based Network Alignment on Biological Data
abstract
Global network alignment is the computational problem of determining the similarity between nodes of different networks to establish a one-to-one correspondence between them. It has important applications in the biological field, particularly for discovering similar roles between the elements of different systems or for transferring knowledge from a well-studied system to another. In this paper, we present SHELLEY, a tool that facilitates the development, testing, and combination of learning-based network alignment algorithms by providing a set of modules that allow for the recreation and combination of both representation learning methods (RLMs) and deep matching methods (DMMs). We then present a case study in which we apply this tool to a protein-protein interaction network (PPI), demonstrating how the representation phase of RLMs is crucial for model robustness against noise.The code of SHELLEY is available at: https://github.com/rickydeluca/shelley
Riccardo De Luca, Manuela Petti, Pietro H. Guzzi, Paolo Tieri
BIBM3
2024 Studying Cardiac infection by tracking clinical data flow: experiences using a REDCap instance
abstract
Studying infection-related diseases, such as those associated with cardiological surgical interventions, often requires the acquisition and analysis of heterogeneous data, including bioimages, microbiological data, and blood analytes. Multidisciplinary collaboration among clinicians and specialists is also essential for effective data management and to develop strategies for the prevention and treatment of infections. Acquiring and analyzing data for clinical studies is a vital approach to preventing infectious diseases. In this context, the REDCap platform provides a comprehensive solution for collecting, managing, and analyzing clinical and hospitalization information. REDCap enables data entry validation, automated reporting, and data integration, which facilitates data management and ensures data quality.This contribution describes a REDCap-based method to support predictive clinical studies on infective endocarditis. The application aims to advance clinical research on this disease, improving understanding and fostering the development of effective treatment strategies, while assisting clinicians in defining cardiac infection prevention methods.
Giuseppe Pozzi, Maria Ghita Cassano, Francesca Giovannenze, Eleonora Taddei, Giancarlo Scoppettuolo, Pietro H. Guzzi, Carlo Torti, Pierangelo Veltri
BIBM6
2024 Application of Generative Graph Models in Biological Network Regeneration: A Selective Review and Qualitative Analysis
abstract
Biological networks are essential for understanding the complex cellular mechanisms of living organisms. Simulating biological networks allows researchers to model and understand complex cellular processes without the need for extensive and costly experiments. The ability to generate synthetic graphs that closely resemble these real-world, complex biological processes is a vital research area. Graph generation in the context of biological networks is particularly significant because it can lead to insights into cellular functions, disease mechanisms, and therapeutic targets. Recent advances in deep learning, particularly in graph generative models, have opened new avenues for applications in the biological domain. These advancements have the potential to revolutionize our understanding of biological systems. However, despite the development of numerous effective graph generation models, there has been limited work assessing the qualitative aspects of the generated graphs within the context of biological networks. Addressing this gap is crucial for ensuring that synthetic graphs are not only structurally accurate but also biologically meaningful.Although various graph generation models are available, their application to biological network recreation has been limited. In this paper, we focus on graph generation, specifically edge and node-independent models for biological networks. We assess four candidate models across four gene expression networks. Our systematic assessment examines the models’ qualitative aspects, including graph structural properties, generation diversity, and computational efficiency. Our findings highlight the strengths and limitations of current models, offering insights to guide the development of more robust graph generation techniques that accurately replicate biological network characteristics.
Binon Teji, Swarup Roy, Pietro H. Guzzi, Dinabandhu Bhandari
BIBM3
2024 An architecture for Deep Learning based automatic bioimages segmentation for sarcopenia evaluation
abstract
Sarcopenia is a clinical condition marked by loss of muscle mass and strength, leading to reduced mobility and quality of life. Accurate identification and quantification of muscle mass are essential for the timely diagnosis and treatment of sarcopenia. To calculate muscle volumes and sarcopenia indexes, segmentation techniques are required on CT images, helping to adjust treatments for chronic diseases. However, evaluating muscle volumes and thus determining sarcopenia indexes currently is highly dependent on manual image segmentation by human operators. We propose a deep learning architecture for automatic muscle mass segmentation, integrating DeepLabv3+ and U-Net3+ models for 2D and 3D segmentation, respectively. These models have been tested on available datasets and combined using an ensemble learning approach to enhance predictive accuracy. This proposed architecture can be integrated into clinical workflows for the assessment of sarcopenia, increasing the reliability and efficiency of image-based diagnoses.
Giuseppe Timpano, Patrizia Vizza, Francesco Manti, Cascini Lucio Giuseppe, Pietro H. Guzzi, Pierangelo Veltri
BIBM5
2024 Non parametric differential network analysis: a tool for unveiling specific molecular signatures
abstract
BACKGROUND: The rewiring of molecular interactions in various conditions leads to distinct phenotypic outcomes. Differential network analysis (DINA) is dedicated to exploring these rewirings within gene and protein networks. Leveraging statistical learning and graph theory, DINA algorithms scrutinize alterations in interaction patterns derived from experimental data. RESULTS: Introducing a novel approach to differential network analysis, we incorporate differential gene expression based on sex and gender attributes. We hypothesize that gene expression can be accurately represented through non-Gaussian processes. Our methodology involves quantifying changes in non-parametric correlations among gene pairs and expression levels of individual genes. CONCLUSIONS: Applying our method to public expression datasets concerning diabetes mellitus and atherosclerosis in liver tissue, we identify gender-specific differential networks. Results underscore the biological relevance of our approach in uncovering meaningful molecular distinctions.
Pietro H. Guzzi, Arkaprava Roy, Marianna Milano, Pierangelo Veltri
BMC Bioinform.1
2023 Annotating omics Data with sex and age of samples: Enabling powerful omics studies
abstract
There is increasing evidence that many molecular processes exhibit differences with age and sex. Such differences produce also differences in the insurgence and progression of many complex diseases. For instance, demographic data on the insurgence of comorbidities of mellitus diabetes, on the lethality of COVID-19, and on some cancers shows differences between sex and age groups. Therefore, the growing interest in such areas requires the management of related data as well as the development of algorithms and tools for the analysis. The availability of omics data annotated with metadata related to age and sex is mandatory for building the analysis pipeline. The number of databases containing data related to age and sex is henceforth growing. We here show some databases and tools storing such data. Finally, future research directions are highlighted.
Pietro H. Guzzi, Mattia Cannistrà, Raffaele Giancotti, Ugo Lomoio, Barbara Puccio, Patrizia Vizza, Giuseppe Tradigo, Pierangelo Veltri
BIBM1
2023 A novel Network Science Algorithm for Improving Triage of Patients
abstract
Patient triage plays a crucial role in healthcare, ensuring timely and appropriate care based on the urgency of patient conditions. Traditional triage methods heavily rely on human judgment, which can be subjective and prone to errors. Recently, a growing interest has been in leveraging artificial intelligence (AI) to develop algorithms for triaging patients. This paper presents the development of a novel algorithm for triaging patients. It is based on the analysis of patient data to produce decisions regarding their prioritization. The algorithm was trained on a comprehensive data set containing relevant patient information, such as vital signs, symptoms, and medical history. The algorithm was designed to accurately classify patients into triage categories through rigorous preprocessing and feature engineering. Experimental results demonstrate that our algorithm achieved high accuracy and performance, outperforming traditional triage methods. By incorporating computer science into the triage process, healthcare professionals can benefit from improved efficiency, accuracy, and consistency, prioritizing patients effectively and optimizing resource allocation. Although further research is needed to address challenges such as biases in training data and model interpretability, the development of AI-based algorithms for triaging patients shows great promise in enhancing healthcare delivery and patient outcomes.
Pietro H. Guzzi, Annamaria Defilippo, Pierangelo Veltri
BIBM1
2023 An Artificial Intelligence-Based Framework for Supporting Management of Patients Affected by Dementia
abstract
Dementia is a major issue for healthcare systems worldwide, necessitating the development of creative and effective strategies for its management. This paper examines the use of Artificial Intelligence (AI) technologies in caring for and managing people with dementia. AI-driven solutions can improve diagnosis, personalize care, optimize medication management, and reduce the burden on caregivers. This paper discusses the implementation of an AI-based framework for creating videos related to people’s memories to support train-therapy or travel-therapy, a non-pharmacological intervention for Alzheimer’s disease patients.
Pietro H. Guzzi, Pierangelo Veltri
BIBM1
2023 An innovative platform to manage the access to social and health services for vulnerable people
abstract
Healthcare access (HA) is a multi-dimensional concept that includes health services availability and accessibility for the populations. These services should be determined by population healthcare needs, especially for vulnerable populations. Digital healthcare service became more important to facilitate the access to medical care by vulnerable people and by citizens in general. Digital health and technologies have provided many online e-services to address social and healthcare services.In this contribution, we propose the implementation of an innovative platform to support and manage the access to health and social services for vulnerable people. Two different use cases have been proposed to demonstrate the application of this platform to different healthcare contexts. The results shows the benefits of using the platform in terms of request management times and reduction of hospitalization.
Patrizia Vizza, Giuseppe Tradigo, Massimiliano Perri, Antonino Posterino, Raffaele Giancotti, Pietro H. Guzzi, Pierangelo Veltri
BIBM6
2023 Tracking and Predicting Productions in Agricultural Processes: Applications and Experiences
abstract
The quality and traceability of agricoltural and food products (indicated as agri-food) represents an important task for industries to focus on environments and wellness targets. In the context of milk and vegetable production processes, it is no possible to monitor and control animals behaviour, environmental conditions, and overall quality affecting these productions. Accurate and explainable predictions of quantities, as well as food properties qualities, is relevant for marketing and planning action in agri-food companies thus to in obtaining more efficient higher-quality productions and contribute to citizens wellness.We here report examples and experiences of machine learning algorithms application to evaluate and predict quantity and frequency of production in an large south of Italy farm. Data are extracted from a tracking system storing all production phases, i.e.: (i) from fruits plants to storage, cold maintaining and transportation, and (ii) cows management, fresh milk analysis and packaging. The here proposed experience contributes to evaluate and predict quantity and frequency of production, aiming to support farms in product planning and production phases.
Patrizia Vizza, Giuseppe Timpano, Francesco Vescio, Gianmichele Caligiuri, Fulvia Michela Caligiuri, Pasquale Lambardi, Pierangelo Veltri, Pietro H. Guzzi, Giuseppe Tradigo
IEEE Big Data8
2023 GTExVisualizer: a web platform for supporting ageing studies
abstract
MOTIVATION: Studying ageing effects on molecules is an important new topic for life science. To perform such studies, the need for data, models, algorithms, and tools arises to elucidate molecular mechanisms. GTEx (standing for Genotype-Tissue Expression) portal is a web-based data source allowing to retrieve patients' transcriptomics data annotated with tissues, gender, and age information. It represents the more complete data sources for ageing effects studies. Nevertheless, it lacks functionalities to query data at the sex/age level, as well as tools for protein interaction studies, thereby limiting ageing studies. As a result, users need to download query results to proceed to further analysis, such as retrieving the expression of a given gene on different age (or sex) classes in many tissues. RESULTS: We present the GTExVisualizer, a platform to query and analyse GTEx data. This tool contains a web interface able to: (i) graphically represent and study query results; (ii) analyse genes using sex/age expression patterns, also integrated with network-based modules; and (iii) report results as plot-based representation as well as (gene) networks. Finally, it allows the user to obtain basic statistics which evidence differences in gene expression among sex/age groups. CONCLUSION: The GTExVisualizer novelty consists in providing a tool for studying ageing/sex-related effects on molecular processes. AVAILABILITY AND IMPLEMENTATION: GTExVisualizer is available at: http://gtexvisualizer.herokuapp.com. The source code and data are available at: https://github.com/UgoLomoio/gtex_visualizer.
Pietro H. Guzzi, Ugo Lomoio, Pierangelo Veltri
Bioinform.1
2022 A novel framework based on network embedding for the simulation and analysis of disease progression
abstract
Modelling infectious disease spreading is crucial for planning effective containment measures, as shown in the COVID-19 pandemic. The effectiveness of planned measures can also be measured regarding saved lives and economic resources. Therefore, introducing methods able to model the evolution and the impact of measures, as well as planning tailored and updated measures, is a crucial step. Existing models for spreading modelling belong to two main classes: (i) compartmental models based on ordinary differential equations and (ii) contact-based models based on a contact structure using an underlining layer to simulate diffusion. Nevertheless, none of these methods can leverage the high computational power of artificial intelligence and deep learning. We propose a novel framework for simulating and analysing disease progression for these methods. The framework is based on the multiscale simulation of the spreading based on using a multiscale contact model built on top of a diffusion model customised by the user. The evolution of the spreading, modelled as a graph with attributed nodes, is then mapped into a latent space through graph embedding. Finally, deep learning models are used in the latent space to analyse and forecast methods without running expensive computational simulations of the contact-based model.
Francesco Chiodo, Mario Torchia, Enza Messina, Elisabetta Fersini, Tommaso Mazza, Pietro H. Guzzi
BIBM6
2022 A machine-learning based tool for bioimages managing and annotation
abstract
Magnetic Resonance Images (MRI) allow to extract meaningful structural information. Machine learning and neural network based algorithms are used to analyze such images, to extract features and to identify anomalies related to diseases. To perform anomaly detection tasks in MR images of the human brain, we propose the use of the Variational AutoEncoder (VAE) method. A VAE is a deep-learning method able to compress and reconstruct the original image through well-defined functions aiming to extract only significant features that are used to identify abnormal pattern. In this contribution, we present a tool based on VAE method for the identification and annotation of brain lesions in MRI aiming to support physicians in the detection of anomalies. Moreover, a MongoDB database is also used to store the data and manage the annotations.
Raffaele Giancotti, Ugo Lomoio, Pierangelo Veltri, Pietro H. Guzzi, Patrizia Vizza
BIBM4
2022 A network-based analysis of genes related to comorbidities in diabetes
abstract
Network medicine helps to shed light insight many chronic diseases by offering useful information about mechanistic information from omic data sets. Type 2 diabetes mellitus (T2DM) is one of the major challenges in medical research. it has been demonstrated that the odds of comorbidities is different considering age and sex of patients. Therefore the use of a framework such as a system and network approach may be useful to shed light into uncertainties related to sex, age effects and comorbidity. We first selected from T2Dico database the list of genes related to comorbidities. Then we first extracted networks of proteins connecting them. In parallel we analysed the pattern of expression of them considering both age and sex as factors, stored into the GTEx database. Preliminary results showed the action of few genes and the biological validation is currently carried out.
Pietro H. Guzzi, Francesca Cortese, Gaia Chiara Mannino, Elisabetta Pedace, Francesco Andreozzi, Pierangelo Veltri
BIBM1
2022 NOMA-DB: a framework for management and analysis of ageing-related gene-expression data
abstract
Recently there is a growing interest for the study of the molecular basis of ageing processes and on the differences among genders. These studies require many data, models and tools for inferring molecular mechanisms. Among the others, the Genotype-Tissue Expression (GTEx) database is one of the prominent resources for the analysis of expression data related to tissues, sex and age. The current version of the database has a lot of querying interfaces that enable many analysis centred on the expression of genes on tissues. Despite this, the database lacks on the analysis at sex/age level, thus the researcher has to download data and then write queries by hand (e.g. for retrieving the expression of a given gene on different age-class in many tissues). It also lacks on the integration with existing protein interaction data. Therefore, the need for the introduction of tools enabling easy access and powerful analysis capabilities (i.e. state of the art network based analysis and integration), arises. We here present NOMA-DB, a framework for ageing studies based on an extension of the GTEx database that enable easy querying at sex/age level, network based analysis. The framework is based on wrapping the GTEx database and on building an application logic level on top of existing data. The current version enables the analysis of genes by tissue, gene and age, thus it may be used in potentially future directions of analysis towards better comprehension of aging/sex-related molecular processes based on the analysis of expression data.
Pietro H. Guzzi, Ugo Lomoio, Rocco Scicchitano, Pierangelo Veltri
BIBM1
2022 Glucose Metabolism Evaluation by using cardiac PET images
abstract
Quantitative analysis of PET images is a clinical common practice. It is used to estimate the input function of 18F-FDG tracer in order to study a physiological process and to evaluate the response to a treatment. It allows the evaluation of coronary artery pathologies, as well as metabolic syndrome associated to cardiovascular diseases. We propose a method for analyzing the dynamic PET cardiac images aiming to assess the progress of glucose metabolism on large vessels as the aorta one. The aim is to study the relation among drug dosage with metabolic syndrome. Indeed, the aim is to correlate the glucose metabolism values (specifically MRGlu - Glucose Metabolic Rate) quantified in the aorta and in the left ventricle, by using PET dynamic images. The measures are presented and proposed for clinical drug validations.
Patrizia Vizza, Giuseppe Tradigo, Pietro H. Guzzi, Elena Succurro, Giuseppe Lucio Cascini, Pierangelo Veltri
BIBM3
2022 Disease spreading modeling and analysis: a survey
abstract
MOTIVATION: The control of the diffusion of diseases is a critical subject of a broad research area, which involves both clinical and political aspects. It makes wide use of computational tools, such as ordinary differential equations, stochastic simulation frameworks and graph theory, and interaction data, from molecular to social granularity levels, to model the ways diseases arise and spread. The coronavirus disease 2019 (COVID-19) is a perfect testbench example to show how these models may help avoid severe lockdown by suggesting, for instance, the best strategies of vaccine prioritization. RESULTS: Here, we focus on and discuss some graph-based epidemiological models and show how their use may significantly improve the disease spreading control. We offer some examples related to the recent COVID-19 pandemic and discuss how to generalize them to other diseases.
Pietro H. Guzzi, Francesco Petrizzelli, Tommaso Mazza
Briefings Bioinform.1
2022 Detection of pan-cancer surface protein biomarkers via a network-based approach on transcriptomics data
abstract
Cell surface proteins have been used as diagnostic and prognostic markers in cancer research and as targets for the development of anticancer agents. Many of these proteins lie at the top of signaling cascades regulating cell responses and gene expression, therefore acting as 'signaling hubs'. It has been previously demonstrated that the integrated network analysis on transcriptomic data is able to infer cell surface protein activity in breast cancer. Such an approach has been implemented in a publicly available method called 'SURFACER'. SURFACER implements a network-based analysis of transcriptomic data focusing on the overall activity of curated surface proteins, with the final aim to identify those proteins driving major phenotypic changes at a network level, named surface signaling hubs. Here, we show the ability of SURFACER to discover relevant knowledge within and across cancer datasets. We also show how different cancers can be stratified in surface-activity-specific groups. Our strategy may identify cancer-wide markers to design targeted therapies and biomarker-based diagnostic approaches.
Daniele Mercatelli, Chiara Cabrelle, Pierangelo Veltri, Federico Manuel Giorgi, Pietro H. Guzzi
Briefings Bioinform.5
2022 Modeling multi-scale data via a network of networks
abstract
MOTIVATION: Prediction of node and graph labels are prominent network science tasks. Data analyzed in these tasks are sometimes related: entities represented by nodes in a higher-level (higher scale) network can themselves be modeled as networks at a lower level. We argue that systems involving such entities should be integrated with a 'network of networks' (NoNs) representation. Then, we ask whether entity label prediction using multi-level NoN data via our proposed approaches is more accurate than using each of single-level node and graph data alone, i.e. than traditional node label prediction on the higher-level network and graph label prediction on the lower-level networks. To obtain data, we develop the first synthetic NoN generator and construct a real biological NoN. We evaluate accuracy of considered approaches when predicting artificial labels from the synthetic NoNs and proteins' functions from the biological NoN. RESULTS: For the synthetic NoNs, our NoN approaches outperform or are as good as node- and network-level ones depending on the NoN properties. For the biological NoN, our NoN approaches outperform the single-level approaches for just under half of the protein functions, and for 30% of the functions, only our NoN approaches make meaningful predictions, while node- and network-level ones achieve random accuracy. So, NoN-based data integration is important. AVAILABILITY AND IMPLEMENTATION: The software and data are available at https://nd.edu/~cone/NoNs. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shawn Gu, Meng Jiang 0001, Pietro H. Guzzi, Tijana Milenkovic
Bioinform.3
2022 PCN-Miner: an open-source extensible tool for the analysis of Protein Contact Networks
abstract
MOTIVATION: Protein Contact Network (PCN) is a powerful method for analysing the structure and function of proteins, with a specific focus on disclosing the molecular features of allosteric regulation through the discovery of modular substructures. The importance of PCN analysis has been shown in many contexts, such as the analysis of SARS-CoV-2 Spike protein and its complexes with the Angiotensin Converting Enzyme 2 (ACE2) human receptors. Even if there exist many software tools implementing such methods, there is a growing need for the introduction of tools integrating existing approaches. RESULTS: We present PCN-Miner, a software tool implemented in the Python programming language, able to (i) import protein structures from the Protein Data Bank; (ii) generate the corresponding PCN; (iii) model, analyse and visualize PCNs and related protein structures by using a set of known algorithms and metrics. The PCN-Miner can cover a large set of applications: from clustering to embedding and subsequent analysis. AVAILABILITY AND IMPLEMENTATION: The PCN-Miner tool is freely available at the following GitHub repository: https://github.com/hguzzi/ProteinContactNetworks. It is also available in the Python Package Index (PyPI) repository.
Pietro H. Guzzi, Luisa Di Paola, Alessandro Giuliani 0002, Pierangelo Veltri
Bioinform.1
2022 Editorial Deep Learning and Graph Embeddings for Network Biology
Pietro H. Guzzi, Marinka Zitnik
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 Data science in unveiling COVID-19 pathogenesis and diagnosis: evolutionary origin to drug repurposing
abstract
MOTIVATION: The outbreak of novel severe acute respiratory syndrome coronavirus (SARS-CoV-2, also known as COVID-19) in Wuhan has attracted worldwide attention. SARS-CoV-2 causes severe inflammation, which can be fatal. Consequently, there has been a massive and rapid growth in research aimed at throwing light on the mechanisms of infection and the progression of the disease. With regard to this data science is playing a pivotal role in in silico analysis to gain insights into SARS-CoV-2 and the outbreak of COVID-19 in order to forecast, diagnose and come up with a drug to tackle the virus. The availability of large multiomics, radiological, bio-molecular and medical datasets requires the development of novel exploratory and predictive models, or the customisation of existing ones in order to fit the current problem. The high number of approaches generates the need for surveys to guide data scientists and medical practitioners in selecting the right tools to manage their clinical data. RESULTS: Focusing on data science methodologies, we conduct a detailed study on the state-of-the-art of works tackling the current pandemic scenario. We consider various current COVID-19 data analytic domains such as phylogenetic analysis, SARS-CoV-2 genome identification, protein structure prediction, host-viral protein interactomics, clinical imaging, epidemiological research and drug discovery. We highlight data types and instances, their generation pipelines and the data science models currently in use. The current study should give a detailed sketch of the road map towards handling COVID-19 like situations by leveraging data science experts in choosing the right tools. We also summarise our review focusing on prime challenges and possible future research directions. CONTACT: [email protected], [email protected].
Jayanta Kumar Das, Giuseppe Tradigo, Pierangelo Veltri, Pietro H. Guzzi, Swarup Roy
Briefings Bioinform.4
2021 Using dual-network-analyser for communities detecting in dual networks
abstract
BACKGROUND: Representations of the relationships among data using networks are widely used in several research fields such as computational biology, medical informatics and social network mining. Recently, complex networks have been introduced to better capture the insights of the modelled scenarios. Among others, dual networks (DNs) consist of mapping information as pairs of networks containing the same set of nodes but with different edges: one, called physical network, has unweighted edges, while the other, called conceptual network, has weighted edges. RESULTS: We focus on DNs and we propose a tool to find common subgraphs (aka communities) in DNs with particular properties. The tool, called Dual-Network-Analyser, is based on the identification of communities that induce optimal modular subgraphs in the conceptual network and connected subgraphs in the physical one. It includes the Louvain algorithm applied to the considered case. The Dual-Network-Analyser can be used to study DNs, to find common modular communities. We report results on using the tool to identify communities on synthetic DNs as well as real cases in social networks and biological data. CONCLUSION: The proposed method has been tested by using synthetic and biological networks. Results demonstrate that it is well able to detect meaningful information from DNs.
Pietro H. Guzzi, Giuseppe Tradigo, Pierangelo Veltri
BMC Bioinform.1
2021 Parallel and distributed association rule mining in life science: A novel parallel algorithm to mine genomics data
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
Inf. Sci.2
2020 Evaluation of the Topological Agreement of Network Alignments
abstract
Aligning protein interaction networks (PPI) of two or more organisms consists of finding a mapping of the nodes (proteins) of the networks that captures important structural and functional associations (similarity). It is a well studied but difficult problem. It is provably NP-hard in some instances thus computationally very demanding. The problem comes in several versions: global versus local alignment; pairwise versus multiple alignment; one-to-one versus many-to-many alignment. Heuristics to address the various instances of the problem abound and they achieve some degree of success when their performance is measured in terms of node and/or edges conservation. However, as the evolutionary distance between the organisms being considered increases the results tend to degrade. Moreover, poor performance is achieved when the considered networks have remarkably different sizes in the number of nodes and/or edges. Here we address the challenge of analyzing and comparing different approaches to global network alignment, when a one-to-one mapping is sought. We consider and propose various measures to evaluate the agreement between alignments obtained by existing approaches. We show that some such measures indicate an agreement that is often about the same than what would be obtained by chance. That tends to occur even when the mappings exhibit a good performance based on standard measures.
Concettina Guerra, Pietro H. Guzzi
BIBM2
2020 A Framework for Patient Data Management and Analysis in Randomised Clinical Trials
abstract
The efficient management and analysis of patient data enrolled in clinical studies is a critical factor for both supporting data management and knowledge discovery from data. Recent trends in literature present many approaches that demonstrate that the integration of multiple data sources (e.g. biochemical parameters, geographical data as well as the behaviour of patients into social networks) may improve the quality of findings. Moreover, the collection of such data may enable the development of a tailored intervention for precision medicine. All these aspects rely on the design and development of novel solutions for data management, storing and consequently, analysis. We here report the design and development of a prototype for data management and sharing introduced during a collaboration of Bioinformatics Laboratory, the Fisiopatology Unit and the University Hospital of Catanzaro. Our findings are currently under the validation of the clinicians.
Pietro H. Guzzi, Tiziana Larussa, Rosarina Vallelunga, Ludovico Abenavoli, Giuseppe Tradigo, Francesco Luzza, Pierangelo Veltri
BIBM1
2020 A method to assess COVID-19 infected numbers in Italy during peak pandemic period
abstract
COVID-19 (SARS-CoV-2) is a pandemic disease diffused throughout the world. COVID-19 is usually identified by applying Reverse transcriptase-polymerase chain reaction (RT-PCR) analysis on swab tests. The high rate of diffusion of the disease caused many problems related to the managing part of limited healthcare resources such as Intensive Care Units (ICUs) services. Assessing the real number of infected as well as early identification of the more infected zones have been defined as a relevant issue to treat pandemic. COVID-19 infected citizens are identified by swab test applied on suspected cases as well as people that have been in touch with affected ones. For these reasons, recognised numbers of COVID-19 affected patients are significantly lower than real ones. We investigate the number of COVID-19 infections and the number of deaths, through Italian regions by comparing these data with respect to diseases caused by similar viruses. We assess several infections having a higher rate of dissemination than the ones currently measured. We focus on the characterisation of the pandemic diffusion by estimating the infected number of patients versus the number of death. We believe that our model can support the healthcare system to react as COVID-19 infection rate increases.
Giuseppe Tradigo, Pietro H. Guzzi, Tamer Kahveci, Pierangelo Veltri
BIBM2
2020 A programmable device to guide rehabilitation patients: design, testing and data collection
abstract
Physical therapy and rehabilitation therapy aim to support patients in dealing with the consequences of their and physical impairments in daily activities. The recent developments in biomedical sensors combined with the wireless network infrastructure will deeply transform healthcare systems and help physicians in designing better and more precise therapies with faster feedback from patients in terms of health-related measured data. Furthermore, these new systems will enable distributed healthcare services for remote patients who may live far away from health structures or who may have movement impairment. Efficiently monitoring or acquiring data from a large number of patients will cause improvements during rehabilitation and help in early diagnoses together with reducing the costs in the healthcare system with more effective prevention. We present a programmable rehabilitation device which can be useful to evaluate patients' performances in a set of physiotherapy exercises designed to evaluate subjects by neurophysiological impairments which slow down some types of movements. The tool is able to support the definition of rehabilitation exercises involving upper limbs and hands. The presented device assesses the responsiveness and movement capacity of patients undergoing physiotherapy aiming to test and measure the mobility, strength and functional ability of the hand during prone supination exercises.
Giuseppe Tradigo, Patrizia Vizza, Pietro H. Guzzi, Gionata Fragomeni, Antonio Ammendolia, Pierangelo Veltri
BIBM3
2020 On the use of clinical based infection data for pandemic case studies
abstract
Epidemiological models are relevant to study and analyze clinical as well as environmental and behavioural data, useful to support health studies. The target is to perform epidemiological analysis producing fast and reliable data access useful to guide prevention and curing processes. This is currently true in pandemic emergency as the current Covid-19 context. Epidemiological models should support in the early identification of pandemic phenomena and in making available data set for studying more accurate drug-based strategy for vaccines or virus containment.In this contribution we present an epidemiology database which integrates different types of clinical data to support research, follow-up and patient monitoring. The idea starts from an hospital databases cooperation integration where virus available data have been integrated to support statistical based studies. Starting from an available database containing 5 years data of infection related viruses (such as HPC, hepatitis) and patient anonymous data, the proposed system provide an integrated data access able to (i) extracting data filtered by means of clinical hypothesis based on patient profiles, environment and drugs and (ii) allowing to build large scale geographical data mappings in order to study correlations among chronic infection diseases and their relations with upcoming pandemic phenomena. Even if the application is in its infancy, the application is relevant with high very important applications.
Giuseppe Tradigo, Patrizia Vizza, Gabriel Gabriele, Maria Mazzitelli, Carlo Torti, Mattia Prosperi, Pietro H. Guzzi, Pierangelo Veltri
BIBM7
2020 An efficient and scalable SPARK preprocessing methodology for Genome Wide Association Studies
abstract
The importance of the use of high-performance software frameworks to analyze omics data obtained by using High-Throughput (HT) essays is widely recognized. HT methodologies comprise microarrays, Genome-Wide Association Studies (GWAS), and Next Generation Sequencing (NGS), which provide a vast amount of data per a single experiment. Each HT vendor provides to the users only the software frameworks and the proprietary libraries for the annotation, and summarization of raw data. Consequently, the needs of algorithms for the preprocessing and analysis of omics data arise. GWAS aims to highlight the association between genetic variants and diseases by examining single nucleotide polymorphisms (SNPs), which differ in a statistically significant way between cases and controls. The effectiveness of GWAS analysis increases with the number of analyzed samples per single experiment. GWAS data analyzed through the use of statistical methods can detect associations among a single allelic variant and the clinical conditions of samples. To overcome these limitations, and to make it possible to discover multiple associations among allelic variants, it is possible to use Association Rules mining. Consequently, the need for the introduction of scalable Association Rule Mining (ARM) algorithms able to analyze GWAS data arises. Hence, the use of high-performance data analytics framework is needed. For this purpose, we propose a software framework called GARMS (GWAS Association Rule Mining in Spark) built on top of Apache Spark for the preprocessing, and mining of association rules from GWAS data sets. GARMS comprises a two steps analysis methodology: (i) in the first step, the GWAS data are preprocessed, along with the identification of the frequent itemsets; (ii) in the second step, frequent itemsets are employed to mine association rules without scanning the input data. We implemented our algorithm, and we tested it on some synthetic GWAS data sets. Preliminary results confirm that our method may extract relevant association rules from GWAS data reducing the computational time.
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
PDP2
2020 BioPAX-Parser: parsing and enrichment analysis of BioPAX pathways
abstract
SUMMARY: Biological pathways are fundamental for learning about healthy and disease states. Many existing formats support automatic software analysis of biological pathways, e.g. BioPAX (Biological Pathway Exchange). Although some algorithms are available as web application or stand-alone tools, no general graphical application for the parsing of BioPAX pathway data exists. Also, very few tools can perform pathway enrichment analysis (PEA) using pathway encoded in the BioPAX format. To fill this gap, we introduce BiP (BioPAX-Parser), an automatic and graphical software tool aimed at performing the parsing and accessing of BioPAX pathway data, along with PEA by using information coming from pathways encoded in BioPAX. AVAILABILITY AND IMPLEMENTATION: BiP is freely available for academic and non-profit organizations at https://gitlab.com/giuseppeagapito/bip under the LGPL 2.1, the GNU Lesser General Public License. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Giuseppe Agapito, Chiara Pastrello, Pietro H. Guzzi, Igor Jurisica, Mario Cannataro
Bioinform.3
2019 Association Rule Mining from large datasets of clinical invoices document
abstract
The concept of massive data generation nowadays affects several domains such as marketing including electronic invoices of large retailers, web access log files, healthcare, life sciences and so on. All these web activities introduced a new way to pay through the concept of electronic invoices (eInvoice), replacing the paper invoices. For these reasons, eInvoicing can be thought of as an innovative digital infrastructure for the issue, transmission, and storage of invoices. The availability of large volumes of eInvoices allows the discovery of new knowledge through data mining in these domains. Thus, users by using data mining can extract knowledge from large invoices documents. In this paper, we present a software tool for mining association rules from invoices produced in healthcare centers. In particular, the tool adopt a novel preprocessing methodology that provides merging, cleaning, formatting and summarization of eInvocies. The methodology can improve the quality of a huge amount of clinical invoices reducing the quantity of irrelevant data, making the remaining data suitable to mine information in form of association rules. The core of the tool allows to extract association rules from eInvoices; as a case study, we discuss the mined rules, highlighting the relationships among the purchased goods.
Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Sabrina Graziano, Mario Cannataro
BIBM3
2019 Pathway Analysis for SNP microarray data
abstract
Pathway Analysis (PA) is a powerful method for data analysis in genomics, most often applied to gene expression analysis, but little used to analyze variants such as Single Nucleotide Polymorphisms (SNPs). PA could allow the interpretation of variants concerning the biological processes in which the affected genes and proteins are involved. Currently, the available PA software tools are not able to automatically perform pathway analysis using SNPs data. PA software tools cannot deal natively with SNPs data, hence several software tools have to be used to put SNPs data in the proper format for the analysis. To overcome these limitations, we present SNP Microarray Pathway Analysis (MPA), a software tool able to discriminate relevant genes from SNP microarrays to use in PA analysis. MPA automatically identifies relevant SNPs using the well known Fisher's test, with which to perform PA. Pathway analysis in MPA is obtained employing the Hypergeometric function. As a result, MPA provides to the user the list of enriched pathways from the identified SNPs. MPA software tool along with the user guide and datasets, are available for download at https://gitlab.com/giuseppeagapito/mpa under the GPL v3.0 license.
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
BIBM2
2019 Mining Association Rules From Disease Ontology
abstract
The Disease Ontology (DO) is standardized, controlled vocabulary that contains information about inherited, developmental and acquired human diseases. Each DO term is associated with disease concepts through an annotation process. The relevance and the specificity of DO terms are often evaluated by its Information Content (IC). An important research area focus on the analysis of annotated data with the goal to extract knowledge. For example, the analysis of annotated data using Association Rules (AR) may supply meaningful knowledge, discovering relevant associations. Classical association rules methods consider all annotation equally, do not taking into account that the DO terms have different Information Content, i.e. different relevance. This implies the generation of association rules with low IC. In this paper we presents WARDO (Weighted Association Rule mining from Disease Ontology), a methodology based on the extraction od Weighted Association Rules from the DO Ontology considering the IC of terms. To assess our methodology, we tested WARDO on DO annotation datasets. WARDO is publicly available at https://gitlab.com/giuseppeagapito/wardo.
Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro
BIBM3
2019 A geographical patients based health information system
abstract
Prevention is essential to counteract the onset of cancer. The analysis of clinical and environmental data as well as their integration are basic topics to prevent chronic diseases, especially neoplasms, and to identify the correlations between cancer and environmental factors. The proposed contribution aims to acquire, analyze and integrate clinical and geographical data to evaluate their possible correlations. A GIS application is here reported to correlate TSH (Thyroid-Stimulating Hormone) values with environmental data as case study.
Giuseppe Tradigo, Patrizia Vizza, Giuseppe Brescia, Pietro H. Guzzi, Pierangelo Veltri
BIBM4
2019 SISTABENE: an information system for the traceability of agricultural food production
abstract
Wellness can be related to the prevention of diseases by means of ensuring the quality of products and is an important challenge in food industry. To this end, food traceability has become a priority in the industry in order let the final users to verify which production phases the food went through. Furthermore it gives domain experts the opportunity to trace defect or issues in the production workflow when problems arise. The proposed information system aims to track the production process of milk and vegetable products. This system is useful both for producers than consumers, giving them a complete tool for food traceability. It provides data management and processing in order to check each production step, making traceability a simpler and more efficient process. Information about raw materials, nutritional facts and activities is readily available and guarantees a transparent and secure supply chain.
Giuseppe Tradigo, Patrizia Vizza, Pierangelo Veltri, Pasquale Lambardi, Fulvia Michela Caligiuri, Gianmichele Caligiuri, Pietro H. Guzzi
BIBM7
2019 GLAlign: A Novel Algorithm for Local Network Alignment
abstract
Networks are successfully used as a modelling framework in many application domains. For instance, Protein-Protein Interaction Networks (PPINs) model the set of interactions among proteins in a cell. A critical application of network analysis is the comparison among PPINs of different organisms to reveal similarities among the underlying biological processes. Algorithms for comparing networks (also referred to as network aligners) fall into two main classes: global aligners, which aim to compare two networks on a global scale, and local aligners that aim to evidence single sub-regions of similarity among networks. The possibility to improve the performance of the aligners by mixing the two approaches is a growing research area. In our previous work, we started to explore the possibility to use global alignment to improve the local one. We here explore further this possibility by using topological information extracted from global alignment to guide the steps of the local alignment. Therefore, we present Global Local Aligner (GLAlign), a methodology that improves the performances of local network aligners by exploiting a preliminary global alignment. Furthermore, we provide implementation of GLAlign. As a proof-of-principle, we evaluated the performance of the GLAlign prototype on the PPINs of five species. Results show that GLAlign methodology outperforms the state-of-the-arts local alignment algorithms. GLAlign is publicly available for academic use and can be downloaded here: https://sites.google.com/site/globallocalalignment/.
Marianna Milano, Pietro H. Guzzi, Mario Cannataro
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 S4S: RESTful Services to Collect, Integrate and Analyze SNPs and Clinical Data on the Web
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
BIBM2
2018 INTEGRO: an algorithm for data-integration and disease-gene association
Pietro Cinaglia, Pietro H. Guzzi, Pierangelo Veltri
BIBM2
2018 Chemical Characterization of Interacting Genes in Few Subnetworks of Alzheimer's Disease
Antara Sengupta, Pabitra Pal Choudhury, Hazel N. Manners, Pietro H. Guzzi, Swarup Roy
BIBM4
2018 Tracking agricultural products for wellness care
Patrizia Vizza, Giuseppe Tradigo, Pierangelo Veltri, Pasquale Lambardi, Claudia Garofalo, Fulvia Michela Caligiuri, Gianmichele Caligiuri, Pietro H. Guzzi
BIBM8
2018 Survey of local and global biological network alignment: the need to reconcile the two sides of the same coin
abstract
Analogous to genomic sequence alignment that allows for across-species transfer of biological knowledge between conserved sequence regions, biological network alignment can be used to guide the knowledge transfer between conserved regions of molecular networks of different species. Hence, biological network alignment can be used to redefine the traditional notion of a sequence-based homology to a new notion of network-based homology. Analogous to genomic sequence alignment, there exist local and global biological network alignments. Here, we survey prominent and recent computational approaches of each network alignment type and discuss their (dis)advantages. Then, as it was recently shown that the two approach types are complementary, in the sense that they capture different slices of cellular functioning, we discuss the need to reconcile the two network alignment types and present a recent first step in this direction. We conclude with some open research problems on this topic and comment on the usefulness of network alignment in other domains besides computational biology.
Pietro H. Guzzi, Tijana Milenkovic
Briefings Bioinform.1
2017 Network based algorithms for module extraction from RNASeq data: A quantitative assessment
abstract
Genes participating in a common module may cause clinically similar diseases and shares the common genetic origin of their associated disease phenotypes. Identifying such modules may be helpful in system level understanding of biological and cellular processes or their disruption caused in associated diseases. The choose dofthe appropriate method for gene selection is a difficult task. In this work we discuss and compare selective module finding methods.
Monica Jha, Pierangelo Veltri, Pietro H. Guzzi, Swarup Roy
BIBM3
2017 Performing local network alignment by ensembling global aligners
abstract
Interactions among proteins are important mechanisms in living cells. The whole set of interactions is often referred to as a protein-protein interaction network (PIN). Comparison among such networks may discover conserved (or disrupted) patterns of interactions among species. Such comparison is performed using network alignment algorithms. They help analyse PPI networks for a better understanding of biological processes such as finding conserved regions between species, giving us insight into their evolution. However, there is no best aligner or standard evaluation measure to assess the quality of alignments. In this work, we use several aligners to produce an ensembled result, which can further improve individual aligners' alignment quality. Two basic ensemble approaches are used: One by finding majority node mappings from aligners and another by combining their results into one final alignment. These alignments are then evaluated based on three scoring schemes: Gene Ontology Consistency (GOC), Node Coverage (NCV) and Generalised Symmetric Substructure Score (GS3) using IsoBase PPI networks. Results show that the majority voting based ensemble scheme performs well in GS3while the ensemble by the union of the decision by different aligners produces satisfactory outcomes in comparison in GOC and NCV scores.
Hazel N. Manners, Ahed Elmsallati, Pietro H. Guzzi, Swarup Roy, Jugal K. Kalita
BIBM3
2017 Parallel and Cloud-Based Analysis of Omics Data: Modelling and Simulation in Medicine
abstract
High throughput experimental platforms and diagnostic equipments available in clinical settings and in research laboratories, such as magnetic resonance imaging, microarray, mass spectrometry and next-generation sequencing, are producing an increasing volume of clinical and omics data. Moreover, Electronic Patients Records (EPRs), eHealth systems, personal mobile sensors and Social Networks are collecting an overwhelming volume of health and life style data that may be integrated with clinical data and more and more is used for the real-time monitoring of patient's health. This poses new issues in terms of secure data storage, effective models for data integration, efficient algorithms for data analysis, new models for health monitoring, that may be addressed, among the others, using high performance computing solutions. Parallel computing and Cloud Computing may offer efficient and scalable solutions in an orthogonal way. In fact, parallel, bioinformatics software, that exploit off-the-shelf high performance computers, may be used to preprocess and analyze omics data at a lower layer, for instance to highlight genetic variation associated with complex diseases. On the other hand, Cloud Computing offers large scale data storage, data sharing services, on-demand anytime and anywhere access to resources and applications, for the realization of elastic and scalable applications and services. Motivated by the increasing use of parallel computing and cloud computing in life sciences, in this paper we survey both parallel bioinformatics algorithms for the parallel preprocessing and statistical and data mining analysis of omics data, as well as Cloud-based healthcare and biomedicine services and systems for large scale applications. Moreover, the paper underlines main issues and problems related to the use of such platforms for the storage and analysis of health data, with special focus to the security and privacy of patients data, that are particularly important in fields such as personalized medicine. Finally, the paper presents some case studies about the parallel and distributed modelling and simulation in medicine and biology.
Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Gionata Fragomeni, Giuseppe Tradigo, Pierangelo Veltri, Mario Cannataro
PDP3
2017 An extensive assessment of network alignment algorithms for comparison of brain connectomes
abstract
BACKGROUND: Recently the study of the complex system of connections in neural systems, i.e. the connectome, has gained a central role in neurosciences. The modeling and analysis of connectomes are therefore a growing area. Here we focus on the representation of connectomes by using graph theory formalisms. Macroscopic human brain connectomes are usually derived from neuroimages; the analyzed brains are co-registered in the image domain and brought to a common anatomical space. An atlas is then applied in order to define anatomically meaningful regions that will serve as the nodes of the network - this process is referred to as parcellation. The atlas-based parcellations present some known limitations in cases of early brain development and abnormal anatomy. Consequently, it has been recently proposed to perform atlas-free random brain parcellation into nodes and align brains in the network space instead of the anatomical image space, as a way to deal with the unknown correspondences of the parcels. Such process requires modeling of the brain using graph theory and the subsequent comparison of the structure of graphs. The latter step may be modeled as a network alignment (NA) problem. RESULTS: In this work, we first define the problem formally, then we test six existing state of the art of network aligners on diffusion MRI-derived brain networks. We compare the performances of algorithms by assessing six topological measures. We also evaluated the robustness of algorithms to alterations of the dataset. CONCLUSION: The results confirm that NA algorithms may be applied in cases of atlas-free parcellation for a fully network-driven comparison of connectomes. The analysis shows MAGNA++ is the best global alignment algorithm. The paper presented a new analysis methodology that uses network alignment for validating atlas-free parcellation brain connectomes. The methodology has been experimented on several brain datasets.
Marianna Milano, Pietro H. Guzzi, Olga Tymofiyeva, Duan Xu, Christopher Paul Hess, Pierangelo Veltri, Mario Cannataro
BMC Bioinform.2
2017 On the Analysis of Diseases and Their Related Geographical Data
abstract
Electronic medical records (EMRs) store data related to patients information enrolled during their stay in health structures. Data stored into EMRs span from data crawled from biological laboratories to textual description of diseases and diagnostic device results (e.g., biomedical images). Each EMR is related to a diagnosis related group (DRG) record. A DRG record is a record associated with a citizen that has been cured in a hospital. It contains a code, called major diagnostic category (MDC), which summarizes the treated disease and allows to reimburse costs related to patient treatments during his staying in health structures. DRGs are used for administrative process (e.g., costs and reimbursement management) as well as disease monitoring. Associating diagnostic codes with external information (such as environmental and geographical data) and with information filtered from EMRs (e.g., biological results or analytes values) can be useful to monitor citizens wellness status. We propose a methodology to analyze such data based on a multistep process. First, we cross reference data by using a semantics-based clustering procedure, extract information from EMRs, and then, cluster them by looking for similar patterns of diseases. Then, biological records in each disease cluster are analyzed to evaluate intracluster similarity by selecting analytes typologies and values. Finally, biological data is related to diagnosis codes and geometrically projected in areas of interest in order to map calculated outlier patients. We applied the methodology on two case studies: 1) diagnosis codes and biochemical analytes of 20 000 biological analyses about hospitalized patients during one observation year and 2) the correlation between cardiovascular diseases and water quality in a southern Italian region. Preliminary findings show the effectiveness of our method.
Giovanni Canino, Pietro H. Guzzi, Giuseppe Tradigo, Aidong Zhang 0001, Pierangelo Veltri
IEEE J. Biomed. Health Informatics2
2016 GLAlign: Using global graph alignment to improve local graph alignment
abstract
During the last years, the graph alignment has been used as a possible way to compare biological networks in system biology. The techniques for the alignment of biological networks fall into two categories: global alignment, that aims to identify large common subnetworks optimizing a topological alignment quality, and local alignment that aims to evidence single sub-regions optimizing functional alignment quality. In this work, we presented GLAlign (Global Local Aligner), a novel algorithm that integrates global and local alignment, starting from the possibility that the topological information gathered by results of global alignment can be used to improve the local alignment building. Initially, the algorithm enables the calculation of global alignment, then it uses this one to guide the building of the local alignment. GLAlign is based on two global and local algorithms widely used in literature, MAGNA++ and AlignMCL. We tested GLAlign as proof-of-principle using the Protein Interaction Networks (PINs) of three species: fly, yeast and worm. GLAlign is publicly available for academic use at https://sites.google.com/site/globallocalalignment/.
Marianna Milano, Mario Cannataro, Pietro H. Guzzi
BIBM3
2016 GIDAC: A prototype for bioimages annotation and clinical data integration
abstract
The analysis of bioimages and their correlated clinical patient information allows to investigate specific diseases and define the corresponding medical protocols. To perform a correct diagnosis and apply a precise therapy, bioimages must be collected and studied together with others relevant data as well as laboratory results, medical annotations and patient history. Today, the management of these data is performed by single systems inside hospital departments that often do not provide dedicated data integration platforms among different departments as well as different health structures to exchange of relevant clinical information. Also, images cannot be annotated or enriched by physicians to trace temporal studies for patients or even among patients with similar diseases. In this contribution, we report the results of a research project called GIDAC (standing for Gestione Integrata DAti Clinici) that aims to define a general purpose framework for the bioimages management and annotations as well as clinical data view and integration in a simple-to-use information system. The proposed framework does not substitute any existing clinical information system but is able in gathering and integrating data by using a XML-based module. The novelty also consists in allowing annotations on DICOM images by means of simple user-interface to take trace of changes intra images as well as comparisons among patients. This system supports oncologists in the management of DICOM images from different devices (e.g., ecograph or PACS) to extract relevant information necessary to query (annotate) images and study similar clinical cases.
Patrizia Vizza, Pietro H. Guzzi, Pierangelo Veltri, Giuseppe Lucio Cascini, Rosario Curia, Loredana Sisca
BIBM2
2016 Experiences on quantitative cardiac PET analysis
abstract
Quantitative analysis of PET images is a useful as well as essential practice to perform an objective measurement of a physiological process. It allows to study diseases, evaluating treatment response and comparing patients data by quantify images. The analysis consists in estimating the quantity of radionuclide tracer uptaken by tissues. We focus on quantitative analysis of dynamic PET studies to evaluate the diseases of coronary artery and myocardium perfusion. We report experiences on quantitative cardiac PET analysis by using a commercial and largely used software to evaluate viable myocardium through Patlak method. We report also results obtained on PET images provided by clinical departments of the Magna Graecia University Medical School of Catanzaro.
Patrizia Vizza, Pietro H. Guzzi, Pierangelo Veltri, Annalisa Papa, Giuseppe Lucio Cascini, Giorgio Sesti, Elena Succurro
BIBM2
2016 DIETOS: A recommender system for adaptive diet monitoring and personalized food suggestion
abstract
Nowadays there is a widespread diffusion of mobile applications for weight and diet management. Even though, the most popular apps are not usually experimented in clinical contexts, as well as apps are not supported by medical evidence. Further research is necessary to assess the effectiveness of apps for weight and diet management. Moreover, there are few examples of food recommender systems that provide to the users nutritional facts about suitable food choices and take into account individual physiological status and environmental situations. We propose DIETOS (DIET Organizer System), a recommender system for the adaptive delivery of nutrition contents to improve the quality of life of both healthy people and individuals affected by chronic diet-related diseases. The proposed system is able to build a user's health profile, and provides individualized nutritional recommendation according to the health profile. The profile is created through the use of dynamic real-time questionnaires prepared by medical doctors and compiled by the users. The health profile includes information about health status and eventual chronic diseases. The first prototype of the system (available online at http://www.easyanalysis.it/dietos), includes a catalogue of typical Calabrian foods compiled by nutrition specialists (Calabria is a region of the southern Italy). DIETOS can suggest not only the use of specific foods compatible with the health status, but also it may give dietary indications related to some specific pathologies or health conditions.
Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Mario Cannataro, Mariadelina Simeoni, Ilaria Care, Theodora Lamprinoudi, Giorgio Fuiano, Arturo Pujia
WiMob3
2016 Methodologies and experimental platforms for generating and analysing microarray and mass spectrometry-based omics data to support P4 medicine
abstract
Predictive, preventive, personalized and participatory (P4) medicine is an emerging medical model that is based on the customization of all medical aspects (i.e. practices, drugs, decisions) of the individual patient. P4 medicine presupposes the elucidation of the so-called omic world, under the assumption that this knowledge may explain differences of patients with respect to disease prevention, diagnosis and therapies. Here, we elucidate the role of some selected omics sciences for different aspects of disease management, such as early diagnosis of diseases, prevention of diseases, selection of personalized appropriate and optimal therapies based on molecular profiling of patients. After introducing basic concepts of P4 medicine and omics sciences, we review some computational tools and approaches for analysing selected omics data, with a special focus on microarray and mass spectrometry data, which may be used to support P4 medicine. Some applications of biomarker discovery and pharmacogenomics and some experiences on the study of drug reactions are also described.
Pietro H. Guzzi, Giuseppe Agapito, Marianna Milano, Mario Cannataro
Briefings Bioinform.1
2016 Extracting Cross-Ontology Weighted Association Rules from Gene Ontology Annotations
abstract
Gene Ontology (GO) is a structured repository of concepts (GO Terms) that are associated to one or more gene products through a process referred to as annotation. The analysis of annotated data is an important opportunity for bioinformatics. There are different approaches of analysis, among those, the use of association rules (AR) which provides useful knowledge, discovering biologically relevant associations between terms of GO, not previously known. In a previous work, we introduced GO-WAR (Gene Ontology-based Weighted Association Rules), a methodology for extracting weighted association rules from ontology-based annotated datasets. We here adapt the GO-WAR algorithm to mine cross-ontology association rules, i.e., rules that involve GO terms present in the three sub-ontologies of GO. We conduct a deep performance evaluation of GO-WAR by mining publicly available GO annotated datasets, showing how GO-WAR outperforms current state of the art approaches.
Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 Guest Editorial for Special Section on Semantic-Based Approaches for Analysis of Biological Data
abstract
The papers in this special section present recent developments in semantic-based approaches for biological data analysis. The systematic integration of biological data with biological knowledge is a recent trend in bioinformatics. Current biological information is spread among multiple sources and encoded in different ontologies. Biological information is associated to biological concepts in a process known as annotation. The annotation of biological data with this additional information enable the use (and the development) of algorithms that use biological ontologies as a framework to mine annotated data based on the use of semantics. The use of such annotations for the analysis of protein data is a relatively novel research area that is becoming more and more important in the field. Indeed, as shown in literature, there is a positive trend in the use of biological information in the analysis of protein data.
Pietro H. Guzzi, Marco Mina
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 Overall Survival Analyzer: A software tool to analyze genotyping and clinical data enriched with temporal events
abstract
The estimation of survival distributions of patients is an important current problem in clinical oncology. The current trend is to integrate molecular data (such as genomic data) with clinical data (e.g. cancer type, stage of the disease, etc) and then to link survival distributions to molecular profile of patients. Recently, the Affymetrix DMET (Drug Metabolizing Enzymes and Transporters) microarray technology has enabled the possibility to determine the allelic variants of a patient and to relate them to phenotype (e.g. drug toxicity). Therefore, the analysis of survival distribution of patients starting from their profile obtained using DMET data may reveal important knowledge to clinicians. In order to provide support to this analysis we propose Overall Survival Analyzer (OS-Analyzer), a software tool able to compute the Overall Survival and Progression-Free Survival (PFS). The tool is able to perform an automatic analysis of data avoiding wasting time on the manual analysis. OS-Analyzer is available to download at the follows web address: https://sites.google.com/site/overallsurvivalanalyzer/.
Giuseppe Agapito, Pietro H. Guzzi, C. Botta, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro
BIBM2
2015 MODULA: A network module based local protein interaction network alignment method
abstract
Biological networks are usually used to model interactions among biological macromolecules in a cells. For instance protein-protein interaction networks (PIN) are used to model and analyse the set of interactions among proteins. The comparison of networks may result in the identification of conserved patterns of interactions corresponding to biological relevant entities such as protein complexes and pathways. Several algorithms, known as network alignment algorithms, have been proposed to unravel relations between different species at the interactome level. Algorithms may be categorized in two main classes: merge and mine and mine and merge. Algorithms belonging to the first class initially merge input network into a single integrated and then mine such networks. Conversely algorithms belonging to the second class initially analyze separately two input networks then integrate such results. In this paper we present MODULA (Network Module based PPI Aligner), a novel approach for local network alignment that belong to the second class. The algorithm at first identifies compact modules from input networks. Modules of both networks are then matched using functional knowledge. Then it uses high scoring pairs of modules as seeds to build a bigger alignment. In order to asses MODULA we compared it to the state of the art local alignment algorithms over a rather extensive and updated dataset.
Pietro H. Guzzi, Pierangelo Veltri, Swarup Roy, Jugal K. Kalita
BIBM1
2015 ICT Solutions for Health Education Model
abstract
Health promotion represents the process to empower the citizens to improve their health lifestyle and to achieve higher levels of wellness. The health models focus on helping people to prevent illnesses through their behavior, and on looking at ways in which a person can pursue better health or ideal health. We report on a project aiming to propose a new model for wellness improvement, consisting in actions to be performed to encourage individuals to become aware of their wellness and develop healthier habits.
Domenico Mirarchi, Patrizia Vizza, Mario Cannataro, Pietro H. Guzzi, Giuseppe Tradigo, Pierangelo Veltri
CBMS4
2015 DMET-Miner: Efficient discovery of association rules from pharmacogenomic data
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
J. Biomed. Informatics2
2014 Improving annotation quality in gene ontology by mining cross-ontology weighted association rules
abstract
The Gene Ontology (GO) is the major resource of annotations for genes and proteins. Despite the presence of large efforts to avoid errors and inconsistencies, some unreliabilities are still present. In particular electronically inferred annotations are more unreliable than manual ones and their number is growing. Thus, the need for an accurate evaluation of annotations in an automatic way arises. In the past, some approaches for improving annotation consistencies have been proposed using association rule mining to discover hidden relationships among GO terms. However such approaches consider all the GO terms equally, while GO terms have different Information Content, i.e. different relevance. Consequently we designed a novel algorithm, (GO-WAR), Mining Weighted Association Rules from GO, that is based on the extraction of weighted association rules considering the IC of terms. We evaluated our algorithm considering seven different species and all the GO ontologies. In all the experiments GO-WAR outperformed state of the art approaches.
Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro
BIBM3
2014 DMET-miner: Efficient learning of association rules from genotyping data for personalized medicine
abstract
Recent developments of microarray technology enable the investigation of allelic variants that may be correlated to phenotypes. In particular the Affymetrix DMET (Drug Metabolism Enzymes and Transporters) platform enables the simultaneous investigation of all the genes that are related to drug absorption, distribution, metabolism and excretion (ADME) and it has been used in clinical studies. In a previous work we developed DMET-Analyzer, a platform able to automatize the study of allelic variants, that has been validated in clinical studies. DMET-Analyzer is able to correlate a single variant for each probe (related to a portion of a gene) through the use of the Fisher test, on the other hand it is unable to discover multiple associations among allelic variants. To overcome those limitations, here we propose DMET-Miner, that is able to correlate the presence of a set of allelic variants by employing an Apriori-like discovery strategy. Preliminary experiments on a synthetic DMET dataset.
Pietro H. Guzzi, Giuseppe Agapito, Maria Teresa Di Martino, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro
BIBM1
2014 Biases in information content measurement of gene ontology terms
abstract
The Gene Ontology (GO) is used to achieve information about gene and protein functions by using a structured vocabulary of terms (GO Terms). GO Terms are related to biological concepts such as proteins or genes through the annotation process. There exist many different annotation processes identified by different evidence codes (EC). Annotated data are stored in public databases such as the Gene Ontology Annotation (GOA) database. Each term has a different specificity also referred to as Information Content (IC) of terms. Both the structure of GO and the corpora of annotation are continuously subject to change due to novel experimental findings. This process is often referred to as ontology evolution. This work focuses on how changes of annotations affect the IC of terms. The study confirms that statistically significant difference among many whole GOA versions exists on each species. Furthermore, there is also a statistically significant difference considering MF taxonomy for human, yeast, worm and fly. These results convey that annotation corpora changes have a high impact on IC.
Marianna Milano, Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
BIBM3
2014 coreSNP: Parallel Processing of Microarray Data
abstract
The availability of high-throughput technologies, such as next generation sequencing and microarray, and the diffusion of genomics studies to large populations are producing an increasing amount of experimental data. In particular, pharmacogenomics studies the impact of genetic variation on drug response in patients and correlates gene expression or single nucleotide polymorphisms (SNPs) with the toxicity or efficacy of a drug, with the aim to improve drug therapy with respect to the patients’ genotype ensuring maximum efficacy with minimal adverse effects. However, the storage, preprocessing, and analysis of experimental data are becoming a main bottleneck in the pharmacogenomics analysis pipeline, due to the increasing number of genes and patients investigated. This paper presents a new parallel software tool named coreSNP for the parallel preprocessing and statistical analysis of DMET (Drug Metabolism Enzymes and Transporters) SNP microarray data produced by Affymetrix for pharmacogenomics studies. The scalable multi-threaded implementation of coreSNP allows to handle the huge volumes of experimental pharmacogenomics data in a very efficient way, while its easy to use graphical user interface and its ability to annotate significant SNPs allow biologists to interpret the results easily. Performance evaluation conducted using real datasets shows good speed-up and scalability and effective response times.
Pietro H. Guzzi, Giuseppe Agapito, Mario Cannataro
IEEE Trans. Computers1
2014 Improving the Robustness of Local Network Alignment: Design and Extensive Assessmentof a Markov Clustering-Based Approach
abstract
The analysis of protein behavior at the network level had been applied to elucidate the mechanisms of protein interaction that are similar in different species. Published network alignment algorithms proved to be able to recapitulate known conserved modules and protein complexes, and infer new conserved interactions confirmed by wet lab experiments. In the meantime, however, a plethora of continuously evolving protein-protein interaction (PPI) data sets have been developed, each featuring different levels of completeness and reliability. For instance, algorithms performance may vary significantly when changing the data set used in their assessment. Moreover, existing papers did not deeply investigate the robustness of alignment algorithms. For instance, some algorithms performances vary significantly when changing the data set used in their assessment. In this work, we design an extensive assessment of current algorithms discussing the robustness of the results on the basis of input networks. We also present AlignMCL, a local network alignment algorithm based on an improved model of alignment graph and Markov Clustering. AlignMCL performs better than other state-of-the-art local alignment algorithms over different updated data sets. In addition, AlignMCL features high levels of robustness, producing similar results regardless the selected data set.
Marco Mina, Pietro H. Guzzi
IEEE ACM Trans. Comput. Biol. Bioinform.2
2013 Using open data in health care and tourism
abstract
Open Data refers to the possibility of freely sharing data among users and organization. Similarly Open Government Initiatives refer to the sharing of documents and data among public governments and citizens. Here we focus on an open initiative held by the Italian Ministry of Health who is making available through Internet a set of Open Data about drug stores, health centers, and other health-related data. In particular, we propose a Cloud-based software tool able to gather and integrate different datasets made available by the Italian Ministry of Health. The proposed Cloud-based tool, called Open Health Data for Tourist (OHT), is able to offer to the tourist information about nearest health care providers (drug stores, public emergency room, hospitals and medical doctors) in Italy through an application accessible from mobile devices.
Mario Cannataro, Pietro H. Guzzi, Pierangelo Veltri
BIBM2
2013 Application of different classification techniques on brain morphological data
abstract
The increasing number of people affected by Neurodegenerative diseases and the improvement of brain imaging diagnostic techniques are bringing to a massive production of brain images that need demanding preprocessing and analysis algorithms. We analyzed volumetric measures of critical brain areas by using different Data Mining methods. Structural magnetic resonance images, generated in our university, were preprocessed using a fully automated segmentation method and the extracted volumetric information was then analyzed by using different binary classifiers. We performed three binary classification experiments considering different data mining algorithms and neurological diseases. Naïve Bayes outperformed all the others classifiers in two experiments, obtaining respectively 93.75% and 95.00% accuracy, while in the third experiment the best classifier was SVM but with a lower accuracy (58,56%). Afterwards, using the Stacking technique we combined the predictions from the best detected three models to build a meta-learner. Meta-learner classification results suggest that the application of the Stacking technique needs more experimentation and the test of additional stackers.
Alessia Sarica, Claudia Critelli, Pietro H. Guzzi, Antonio Cerasa, Aldo Quattrone, Mario Cannataro
CBMS3
2013 Visualization of protein interaction networks: problems and solutions
abstract
BACKGROUND: Visualization concerns the representation of data visually and is an important task in scientific research. Protein-protein interactions (PPI) are discovered using either wet lab techniques, such mass spectrometry, or in silico predictions tools, resulting in large collections of interactions stored in specialized databases. The set of all interactions of an organism forms a protein-protein interaction network (PIN) and is an important tool for studying the behaviour of the cell machinery. Since graphic representation of PINs may highlight important substructures, e.g. protein complexes, visualization is more and more used to study the underlying graph structure of PINs. Although graphs are well known data structures, there are different open problems regarding PINs visualization: the high number of nodes and connections, the heterogeneity of nodes (proteins) and edges (interactions), the possibility to annotate proteins and interactions with biological information extracted by ontologies (e.g. Gene Ontology) that enriches the PINs with semantic information, but complicates their visualization. METHODS: In these last years many software tools for the visualization of PINs have been developed. Initially thought for visualization only, some of them have been successively enriched with new functions for PPI data management and PIN analysis. The paper analyzes the main software tools for PINs visualization considering four main criteria: (i) technology, i.e. availability/license of the software and supported OS (Operating System) platforms; (ii) interoperability, i.e. ability to import/export networks in various formats, ability to export data in a graphic format, extensibility of the system, e.g. through plug-ins; (iii) visualization, i.e. supported layout and rendering algorithms and availability of parallel implementation; (iv) analysis, i.e. availability of network analysis functions, such as clustering or mining of the graph, and the possibility to interact with external databases. RESULTS: Currently, many tools are available and it is not easy for the users choosing one of them. Some tools offer sophisticated 2D and 3D network visualization making available many layout algorithms, others tools are more data-oriented and support integration of interaction data coming from different sources and data annotation. Finally, some specialistic tools are dedicated to the analysis of pathways and cellular processes and are oriented toward systems biology studies, where the dynamic aspects of the processes being studied are central. CONCLUSION: A current trend is the deployment of open, extensible visualization tools (e.g. Cytoscape), that may be incrementally enriched by the interactomics community with novel and more powerful functions for PIN analysis, through the development of plug-ins. On the other hand, another emerging trend regards the efficient and parallel implementation of the visualization engine that may provide high interactivity and near real-time response time, as in NAViGaTOR. From a technological point of view, open-source, free and extensible tools, like Cytoscape, guarantee a long term sustainability due to the largeness of the developers and users communities, and provide a great flexibility since new functions are continuously added by the developer community through new plug-ins, but the emerging parallel, often closed-source tools like NAViGaTOR, can offer near real-time response time also in the analysis of very huge PINs.
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
BMC Bioinform.2
2012 CytoMCL: A Cytoscape plugin for fast clustering of protein interaction networks
abstract
The analysis of the whole set of molecular interactions in an organism, often referred to as interaction networks, is becoming an important research area. A main approach for such analysis resides on the application of clustering techniques to such networks. The meaning of discovered clusters, (i.e. highly interconnected regions), is strictly related to the type of networks. For instance in protein-protein interaction networks clusters may represent protein complexes. The Markov Clustering Algorithm (MCL) is a wellknown algorithm for clustering graphs. It does not provide a graphical user interface and cannot be used in the Cytoscape platform. We present CytoMCL a Cytoscape plugin that finds clusters in a graph by using MCL. It is based on an intuitive interface it is able to load a network from Cytoscape, to analyze it and to visualize resulting clusters into Cytoscape.
Pietro H. Guzzi, Mario Cannataro
CBMS1
2012 SySQ: A Web-based system for survey and questionnaire management in medicine
abstract
A questionnaire is a method for collecting data that can come from many different sources: from observations, telephone interviews or documentary sources. Whatever the source of data is, the questionnaire provides a framework of questions that facilitate researcher's work. A manual approach for collecting data using questionnaire, presents some limitations and introduces several sources of errors. An Informatic System was implemented to reduce these errors and to support researchers in epidemiological studies. For the experimentation of the prototype, a paper format questionnaire has been digitalized from the original sheet provided by the Chair of Hygiene of the Magna Graecia University of Catanzaro. The implemented system allows researchers to create questionnaires, adding sections and structured questions. The administrator of the system can visualize a preview of the questionnaire and decide which group of users can compile it. The system provides a control panel to analyze the collected data and that permits the exportation of saved data into statistical software compatible formats.
Alessia Sarica, Pietro H. Guzzi, Domenico Flotta, Carmelo G. A. Nobile, Mario Cannataro
CBMS2
2012 audioEPR: A specialized electronic patient records for the semi-automatic management of clinical data in audiology
abstract
General-purpose Electronic Patient Records (EPR) may lack in support of specialistic clinical data that characterizes each clinical domain. In many clinical settings often such data remain inside the computer systems associated with specialistic instruments and are not readily available to clinicians for conducting large studies for research purposes. In this paper we present audioEPR, a specialized EPR for the semi-automatic management and querying of audiological and otoneurological clinical data. The goal of the tool is to support the day-by-day clinical activity as well as the simple selection and extraction of clinical data for research activity. The realized prototype is able to support the semi-automatic storage of data related to different Audiology tests and the generation of related diagnostic reports. To encourage its use in a clinical setting its interface resembles the form of the paper-based documents currently used by operators, while maintaining the formal correctness of the produced documentation. On the other hand, its database allows an easy selection and extraction of set of clinical data for research purposes.
Salvatore Scaramuzzino, Pietro H. Guzzi, Claudio Petrolo, Giuseppe Chiarella, Pierangelo Veltri, Mario Cannataro
CBMS2
2012 Semantic similarity analysis of protein data: assessment with biological features and issues
abstract
The integration of proteomics data with biological knowledge is a recent trend in bioinformatics. A lot of biological information is available and is spread on different sources and encoded in different ontologies (e.g. Gene Ontology). Annotating existing protein data with biological information may enable the use (and the development) of algorithms that use biological ontologies as framework to mine annotated data. Recently many methodologies and algorithms that use ontologies to extract knowledge from data, as well as to analyse ontologies themselves have been proposed and applied to other fields. Conversely, the use of such annotations for the analysis of protein data is a relatively novel research area that is currently becoming more and more central in research. Existing approaches span from the definition of the similarity among genes and proteins on the basis of the annotating terms, to the definition of novel algorithms that use such similarities for mining protein data on a proteome-wide scale. This work, after the definition of main concept of such analysis, presents a systematic discussion and comparison of main approaches. Finally, remaining challenges, as well as possible future directions of research are presented.
Pietro H. Guzzi, Marco Mina, Concettina Guerra, Mario Cannataro
Briefings Bioinform.1
2012 DMET-Analyzer: automatic analysis of Affymetrix DMET Data
abstract
BACKGROUND: Clinical Bioinformatics is currently growing and is based on the integration of clinical and omics data aiming at the development of personalized medicine. Thus the introduction of novel technologies able to investigate the relationship among clinical states and biological machineries may help the development of this field. For instance the Affymetrix DMET platform (drug metabolism enzymes and transporters) is able to study the relationship among the variation of the genome of patients and drug metabolism, detecting SNPs (Single Nucleotide Polymorphism) on genes related to drug metabolism. This may allow for instance to find genetic variants in patients which present different drug responses, in pharmacogenomics and clinical studies. Despite this, there is currently a lack in the development of open-source algorithms and tools for the analysis of DMET data. Existing software tools for DMET data generally allow only the preprocessing of binary data (e.g. the DMET-Console provided by Affymetrix) and simple data analysis operations, but do not allow to test the association of the presence of SNPs with the response to drugs. RESULTS: We developed DMET-Analyzer a tool for the automatic association analysis among the variation of the patient genomes and the clinical conditions of patients, i.e. the different response to drugs. The proposed system allows: (i) to automatize the workflow of analysis of DMET-SNP data avoiding the use of multiple tools; (ii) the automatic annotation of DMET-SNP data and the search in existing databases of SNPs (e.g. dbSNP), (iii) the association of SNP with pathway through the search in PharmaGKB, a major knowledge base for pharmacogenomic studies. DMET-Analyzer has a simple graphical user interface that allows users (doctors/biologists) to upload and analyse DMET files produced by Affymetrix DMET-Console in an interactive way. The effectiveness and easy use of DMET Analyzer is demonstrated through different case studies regarding the analysis of clinical datasets produced in the University Hospital of Catanzaro, Italy. CONCLUSION: DMET Analyzer is a novel tool able to automatically analyse data produced by the DMET-platform in case-control association studies. Using such tool user may avoid wasting time in the manual execution of multiple statistical tests avoiding possible errors and reducing the amount of time needed for a whole experiment. Moreover annotations and the direct link to external databases may increase the biological knowledge extracted. The system is freely available for academic purposes at: https://sourceforge.net/projects/dmetanalyzer/files/
Pietro H. Guzzi, Giuseppe Agapito, Maria Teresa Di Martino, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro
BMC Bioinform.1
2011 Challenges in microarray data management and analysis
abstract
Microarray is a key technology in genomics and is increasingly used in molecular biology as well as in molecular medicine and clinical applications. The availability of different microarray types and vendors and the increasing number of samples forming microarray studies pose new challenges in the management and analysis of microarray data. The paper recalls main microarray types and goals and discusses the most important challenges in microarray management and analysis, including the following issues. Heterogeneity in microarray data format requires the application of different preprocessing and annotation tools and the use of different, vendor-specific, preprocessing libraries. Multiplicity of microarray types (e.g. gene expression, SNPs and miRNA arrays, to cite a few) requires the application of different analysis tools. The application of various preprocessing steps produces a number of different files for each study, such as raw, preprocessed, annotated and eventually filtered data, that must be managed and stored properly. Finally, the increasing volume of microarray data due to the use of large sets of samples poses further challenges for the storage of such data. Finally, some emerging directions to face such challenges are also described.
Pietro H. Guzzi, Mario Cannataro
CBMS1
2011 Automatic summarisation and annotation of microarray data
Pietro H. Guzzi, Maria Teresa Di Martino, Giuseppe Tradigo, Pierangelo Veltri, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro
Soft Comput.1
2010 mu-CS: An extension of the TM4 platform to manage Affymetrix binary data
abstract
BACKGROUND: A main goal in understanding cell mechanisms is to explain the relationship among genes and related molecular processes through the combined use of technological platforms and bioinformatics analysis. High throughput platforms, such as microarrays, enable the investigation of the whole genome in a single experiment. There exist different kind of microarray platforms, that produce different types of binary data (images and raw data). Moreover, also considering a single vendor, different chips are available. The analysis of microarray data requires an initial preprocessing phase (i.e. normalization and summarization) of raw data that makes them suitable for use on existing platforms, such as the TIGR M4 Suite. Nevertheless, the annotations of data with additional information such as gene function, is needed to perform more powerful analysis. Raw data preprocessing and annotation is often performed in a manual and error prone way. Moreover, many available preprocessing tools do not support annotation. Thus novel, platform independent, and possibly open source tools enabling the semi-automatic preprocessing and annotation of microarray data are needed. RESULTS: The paper presents mu-CS (Microarray Cel file Summarizer), a cross-platform tool for the automatic normalization, summarization and annotation of Affymetrix binary data. mu-CS is based on a client-server architecture. The mu-CS client is provided both as a plug-in of the TIGR M4 platform and as a Java standalone tool and enables users to read, preprocess and analyse binary microarray data, avoiding the manual invocation of external tools (e.g. the Affymetrix Power Tools), the manual loading of preprocessing libraries, and the management of intermediate files. The mu-CS server automatically updates the references to the summarization and annotation libraries that are provided to the mu-CS client before the preprocessing. The mu-CS server is based on the web services technology and can be easily extended to support more microarray vendors (e.g. Illumina). CONCLUSIONS: Thus mu-CS users can directly manage binary data without worrying about locating and invoking the proper preprocessing tools and chip-specific libraries. Moreover, users of the mu-CS plugin for TM4 can manage Affymetrix binary files without using external tools, such as APT (Affymetrix Power Tools) and related libraries. Consequently, mu-CS offers four main advantages: (i) it avoids to waste time for searching the correct libraries, (ii) it reduces possible errors in the preprocessing and further analysis phases, e.g. due to the incorrect choice of parameters or the use of old libraries, (iii) it implements the annotation of preprocessed data, and finally, (iv) it may enhance the quality of further analysis since it provides the most updated annotation libraries. The mu-CS client is freely available as a plugin of the TM4 platform as well as a standalone application at the project web site (http://bioingegneria.unicz.it/M-CS).
Pietro H. Guzzi, Mario Cannataro
BMC Bioinform.1
2010 IMPRECO: Distributed prediction of protein complexes
Mario Cannataro, Pietro H. Guzzi, Pierangelo Veltri
Future Gener. Comput. Syst.2
2009 GridSnake: A Grid-based implementation of the Snake segmentation algorithm
abstract
Medical imaging is becoming a key technique to visualize the internal structure of the body. Magnetic resonance imaging (MRI) is currently used to take different spatial images of organs, such as the heart. The output of such an analysis is a set of images representing different views of the body or of an organ. There exist many algorithms to pre-process and analyze medical images such as the well known Snake segmentation algorithm. An issue in medical imaging is the large size of images, that require large and efficient data stores, and the high computational power needed to process them. For these reasons, the Grid is being more and more used as an ideal environment for medical image processing. This paper presents a first experience in porting the snake algorithm on a Globus-based Grid.
Mario Cannataro, Pietro H. Guzzi, Marcelo Lobosco, Rodrigo Weber dos Santos
CBMS2
2009 Using ontologies for annotating and retrieving protein-protein interactions data
abstract
Protein-protein interaction (PPI) databases store the whole set of protein interactions in organism. In spite of the availability of much biological information spread on different sources (e.g. Gene Ontology), neither proteins nor interactions are generally annotated in PPI databases. This results in very poor querying capabilities of PPI databases that enable very simple queries. The annotation of proteins and interactions stored in PPI databases may allow the implementation of more powerful querying interfaces. The paper presents a software architecture for the annotation of existing PPI databases with information extracted from Gene Ontology. A simple extension of the query interface of an existent PPI database is discussed.
Mario Cannataro, Pietro H. Guzzi, Pierangelo Veltri
CBMS2
2008 myMCL: A Web Portal for Protein Complexes Prediction
abstract
Interactomics is the study of the Interactome, i.e. the whole set of macromolecular interactions within a cell. Proteins interact among them and different interactions are represented as graphs named Protein to Protein Interaction (PPI) networks. The interest in analyzing PPI networks is related to the possibility of predicting PPI properties on the basis of global properties of the graph (e.g. verify if homology among species involves PPI similarity), or to find set of protein interactions that has a biological meaning. The prediction of protein complexes has been faced in the last years by using different clustering algorithms. The Markov Clustering algorithm (MCL) is a method that presents one of the best performance but is currently available only as a stand alone application with a simple command-line interface available only on Linux platforms. Following a trend in bioinformatics, we provide a web portal (myMCL) allowing remote users to access MCL functions through the Internet. myMCL enables user to submit a job and stores results in a local database for further processing.
Mario Cannataro, Pietro H. Guzzi, Pierangelo Veltri
CBMS2
2008 A Tool for the Semiautomatic Acquisition of the Morphological Data of Blood Vessel Networks
abstract
The simulation of the dynamics of the blood flow in the venous system of the lower limb is an important tool for supporting clinical research and for suggesting possible treatments for many diseases, e.g. for enhancing the surgical treatment of chronic venous insufficiency (CVI). Nevertheless the accuracy of the simulation of the blood flow is strictly related to the morphological data characterizing the investigated venous system. Although some of these data can be extracted from the observation of the real blood flow of a patient, e.g. through the acquisition of a set of images, the extraction of such values is often performed in a manual way, so the need for the automatic induction of parameters arises. The paper presents a software module that allows the semiautomatic acquisition of the morphological data of the venous system of a patient. The tool, developed as a plugin of the ImageJ imaging platform, receives in input a DICOM file containing the computerized tomography (CT) of the vessels network of the lower limb, and produces in a semi-automatic way a weighted graph of the network. This model can be used as the input for a subsequent simulation of the system.
Mario Cannataro, Pietro H. Guzzi, Giuseppe Tradigo, Pierangelo Veltri
ISPA2
2007 Using ontologies for preprocessing and mining spectra data on the Grid
Mario Cannataro, Pietro H. Guzzi, Tommaso Mazza, Giuseppe Tradigo, Pierangelo Veltri
Future Gener. Comput. Syst.2
2006 Analysis and Classification of Proteomics Data, a Case Study
abstract
This paper presents a methodology for analyzing and classifying proteins identified in biological samples. In particular, such methodology consists in normalizing and classifying quantity and quality of proteins identified by using tandem mass spectrometry. A case study is considered and a classification experiment for protein discriminant is also reported
Pietro H. Guzzi, Mario Cannataro, Marco Gaspari, Tommaso Mazza, Barbara Quaresima, Pierangelo Veltri, Francesco Saverio Costanzo
CBMS1
2005 Preprocessing of Mass Spectrometry Proteomics Data on the Grid
abstract
The combined use of mass spectrometry and data mining is a novel approach in proteomic pattern analysis for discovering novel biomarkers or identifying patterns and associations in proteomic profiles. Data produced by mass spectrometers are affected by errors and noise due to sample preparation and instrument approximation, so different preprocessing techniques need to be applied before analysis is conducted. We survey different techniques for spectra preprocessing, and we present a first design of a software tool that allows the preprocessing, management and analysis of mass spectrometry data on the Grid.
Mario Cannataro, Pietro H. Guzzi, Tommaso Mazza, Giuseppe Tradigo, Pierangelo Veltri
CBMS2