Giuseppe Agapito

dblp:65/9707 · DBLP profile ↗
← Back
39ranked-venue papers
27as first author
18since 2021 · last 2025
0000-0003-2868-7732ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 19 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Comparative Analysis of Algorithms and Computational Architectures for Efficient Biological Data Processing
abstract
This paper presents a comparative analysis of algorithms and computational architectures for the efficient processing of biological data, particularly focusing on sparse matrix representations for biological pathways. Biological networks, such as protein-protein interaction (PPI) networks and metabolic pathways, are represented as graphs with nodes corresponding to molecular entities and edges representing biochemical interactions. These graphs are typically sparse, leading to challenges in data storage and computational efficiency. The study examines diverse sparse matrix formats, including Coordinate (COO), Compressed Sparse Row (CSR), and Compressed Sparse Column (CSC), emphasizing their application in biological data analysis. It further explores the potential of modern high-performance computing architectures, such as Graphics Processing Units (GPUs), to accelerate the execution of graph-based algorithms. Algorithms such as Dijkstra’s, Breadth-First Search (BFS), and PageRank are adapted to leverage the efficiencies of sparse matrix representations, optimizing the analysis of large-scale biological networks for improved performance and scalability. Performance evaluations show that GPUs significantly outperform Central Processing Units (CPUs) in processing largescale biological networks, reducing execution time and energy consumption while enhancing scalability. This research demonstrates how the use of an appropriate sparse matrix format and computational architecture can optimize the analysis of complex biological networks, providing insights into biological processes and therapeutic targets.
Giuseppe Agapito, Gaetano Guardasole, Mario Cannataro
PDP1
2025 Ten practical tips and tricks to improve the effectiveness of biological network alignment
abstract
Network alignment (NA) is a computational methodology employed to compare biological networks across different species or conditions. By identifying conserved structures, functions, and interactions, NA provides invaluable insights into shared biological processes, evolutionary relationships, and system-level behaviors. This manuscript presents a comprehensive overview of NA methodologies, including the importance of preprocessing network data, selecting suitable input formats, and understanding diverse network types such as attributed, temporal, and multilayer networks. Additionally, it explores key challenges such as seed nodes selection, algorithm configuration, and cross-species alignment, emphasizing the necessity of integrating functional annotations, sequence similarity, and network topology for biologically meaningful results. Various NA strategies, including Local and Global Network Alignment, are discussed alongside their respective advantages and limitations. Practical recommendations for effectively documenting and visualizing NA experiments are also provided, ensuring reproducibility and clarity in research. By leveraging diverse alignment tools and adopting best practices, researchers can unlock the potential of NA to advance our understanding of complex biological systems.
Giuseppe Agapito, Mario Cannataro, Pietro Cinaglia, Marianna Milano
PLoS Comput. Biol.1
2024 Multi Weighted Graphs as Magnifier to Discover Hidden Cross-Interactions among Biological Pathways
abstract
Scientists and researchers need robust and flexible methods to identify interactions among multiple biological pathways, which unveil crucial insights otherwise obscured when pathways are analyzed in isolation. Understanding these complex interrelations is vital for comprehensively mapping the processes relevant to their research.Biological pathways are categorized into three main classes: signaling, metabolic, and regulatory.The pathways representation can be significantly enhanced using multi weighted graphs models, a concept from graph theory where multiple edges can connect nodes. This allows for a more nuanced representation of the complex interactions and multiple relationships between pathway elements. In this model, each node represents a pathway element, and the multiple edges between nodes depict various types of interactions, capturing the dynamic and multifaceted nature of cellular processes.To address the needs of researchers for more sophisticated and accurate analysis, we have developed a new methodology to convert and represent biological pathways as multi weighted graphs. In this way, it is possible to recall the graph theory metrics to evaluate the elements within the graphs. We show that the use of multi weighted graphs to connect several pathways, allowing to identify hidden cross-pathways links, e.g., genes shared among pathways, through the use of PageRank and Kats index. The discovered genes allow to improve of some order the magnitude of pathway enrichment analysis results, as conveyed from the reached p-values.
Giuseppe Agapito, Mario Cannataro, Gaetano Guardasole
BIBM1
2024 Modeling UGT2B7 and NR1I3 genes through multilayer network to highlight hidden link with taxane neurotoxicity
abstract
Breast cancer (BC) remains a leading global malignancy, with taxane-based treatment (TBT), including paclitaxel and docetaxel, significantly enhancing outcomes across early-stage, locally advanced, and metastatic BC. Despite its efficacy, TBT can cause severe taxane-related peripheral neurotoxicity (TrPN), which limits dosing and affects patient quality of life. TrPN is cumulative, primarily sensory, and unpredictable, with genetic predisposition playing a key role in its development. Previous pharmacogenomic (PGx) study identified five single nucleotide polymorphisms (SNPs) in the NR1I3 and UGT2B7 genes, linked to protection against severe TrPN in BC patients. ROC analysis validated these findings in an independent BC dataset. Neuroprotective effects were in patients homozygous for allele variants 2 with ultrametabolizer phenotype, which accelerates taxane inactivation and reduces treatment efficacy, potentially worsening prognosis. NR1I3 and UGT2B7 genes could serve as predictive biomarkers for TrPN and taxane bioavailability. Further network and pathway enrichment analyses (PEA) provided insights into the molecular pathways involved in TrPN and BC progression. Network analysis, a powerful tool in computational biology, integrates genomic and proteomic data to uncover drug-response mechanisms. This approach advances personalized medicine by identifying key genes and biomarkers for predicting treatment response and adverse effects, offering new therapeutic opportunities.
Giuseppe Agapito, Marianna Milano, Francesca Scionti, Nicoletta Staropoli, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro, Mariamena Arbitrio
BIBM1
2024 A Graph Neural Network based fMRI classification
abstract
In this paper, we aim to describe a novel deep learning model (a machine learning subclass) to classify fMRI. We used a publicly available dataset to train and test the model which involved patients affected by depression.Our model is based on a specific deep neural network, the Graph Attention Network (GAT) which has proven its strength in dealing with graph data representation. The novelty of our approach is that it is based on the extraction from the original fMRIs of graph representation then passed to the deep learning model.We performed this crucial phase by using a Matlab based toolbox, CONN, which helped in data preparation and graph representation extraction. We then used the extracted fMRI representations to feed, train, and finally test our deep learning model.While classification results were encouraging, achieving approximately 73% accuracy, another aspect that we investigated was focused on the comparison of three architectural solutions, focusing on power consumption. We used an Apple Silicon platform compared to a NVIDIA based laptop and an edge device of the NVIDIA Jetson family.
Luca Barillaro, Marianna Milano, Giuseppe Agapito, Mario Cannataro
BIBM3
2024 Efficient GPU Processing Method to Analyze Large GWAS Data Sets
abstract
Genome-wide association studies (GWAS) have ex-panded rapidly, generating large genetic data sets that require advanced computational strategies for efficient analysis. This study introduces a novel method that uses Graphics Processing Units (GPUs) to analyze large scale GWAS data sets, enhancing computational efficiency and reducing processing time. Our approach leverages GPUs' inherent parallel processing capabilities to significantly accelerate computation-intensive tasks commonly encountered in GWAS, such a Genetic Risk Score (GRS). We explain how our GPU-GRS computational approach out-performs traditional CPU (Central Processing Unit) methods, outlining the architectural optimizations that enable parallel processing to handle large genomic data sets more effectively. We also demonstrate the scalability of our approach by ana-lyzing increasingly large synthetic G WAS data sets, showcasing its ability to manage the growing size and complexity of genetic data efficiently. Our findings suggest that adopting GPU-based methods in G WAS analysis can play a pivotal role in the era of big genomic data, offering a path towards more time-efficient and scalable solutions.
Giuseppe Agapito, Gaetano Guardasole, Mario Cannataro
PDP1
2023 Use Predictive Learning Model to Tackle Data Breaches in Healthcare Domain
abstract
Healthcare data breaches are a growing problem that seriously threatens patient privacy, the reputation and trustworthiness of both public and private healthcare organizations. The aim of this paper is to elucidate the severity of healthcare data breaches, their potential impact on patient privacy and healthcare organizations, providing a predictive data breaches transformer conceptual architecture that can monitoring users actions that can result in possible system security violation and consequently in data breaches. At this regard, we introduce the description of a concept architecture for implementing a predictive data breaches transformer highlighting weakness and strengthens, and in the same time how the adoption of a predictive system can significantly limit the risk of the onset of possible data breaches by making the operator more aware in carrying out his activity in the processing of each type of data including personal health data.
Giuseppe Agapito, Mario Cannataro, Pietro Cinaglia, Gaetano Guardasole, Marianna Milano
BIBM1
2023 Artificial Intelligence to Analyze Huge Amounts of Juridical Documents via Edge Computing
Gessica Fulciniti, Abdellah Kabli, Mario Cannataro, Gaetano Guardasole, Giuseppe Agapito
EWSN5
2023 Using Edge-based Deep Learning Model for Early Detection of Cancer
abstract
Cancer is one of the most frequent causes of death in the world. Usually, cancer can be easily diagnosed if characteristic symptoms occur. However, many people who are suffering from cancer have no symptoms. Early diagnosis of tumors is essential to contrast their progression, helping to define more effective treatments to provide long-term survival. Early cancer detection is effective if sensible data can be investigated through high-performance technologies like edge computing. Edge computing is a new paradigm for analyzing data as close to the source as possible, avoiding exporting them outside. Hence, edge-based deep learning models can be applied to improve early cancer detection. This paper provides an use case of a classification task on tumor-related data based on the famous UCI machine learning data sets repository using a deep learning approach based on edge computing. In addition, the manuscript provides an overview of the edge computing paradigm, highlighting its advantages and usability. We also described a small experiment with real tumor data to characterize performance considerations. Moreover, the presented model can be used with different data types, such as images, EGC, and ECC signals.
Luca Barillaro, Giuseppe Agapito, Mario Cannataro
PDP2
2023 High performance deep learning libraries for biomedical applications
abstract
Deep learning approaches are a topic of growing interest since they can achieve high precision in machine learning tasks and may be useful in several scenarios, while high performance computing (HPC) is one of the driving factors for deep learning applications since they require massive computational power. One of these scenarios is related to biomedical context since the massive growth of data generated by several medical procedures. Deep learning techniques, applied on these data may be useful both for medical procedures and for further knowledge discovery in specific field (in example gene interaction related to some diseases). Therefore the importance to have a deep learning library tailored for these task is evident. This paper aims to discuss about some libraries specifically designed to provide convenient high performance computing oriented deep learning support to biomedical applications. We describe two libraries developed inside a European project, named the Deep Health Project, to support both deep learning basic operations and computer vision tasks, oriented to a distributed computing fashion and with some special features for managing biomedical data. In addition we highlight some differences and comparisons with popular environments like Keras and Tensorflow by describing a simple use case.
Luca Barillaro, Giuseppe Agapito, Mario Cannataro
PDP2
2022 Edge-based Deep Learning in Medicine: Classification of ECG signals
abstract
This paper aims to discuss the use of edge computing in medicine with a focus on the analysis of ECG biosignals. Edge computing is a novel paradigm which aims to perform computations (or at least most of them) near the data source achieving some advantages over classical centralized or distributed approaches (e.g. on cloud). After introducing edge computing and the novel NVIDIA Jetson device, the paper presents a use case regarding the classification of ECG biosignals on such a device. Several experiments were conducted to show main differences between traditional and edge-based data analysis approaches. Performance evaluation showed little differences between a traditional approach given the power constrained scenario of edge device. Main results of the paper include an overview of the edge computing paradigm, and a first performance evaluation of deep learning applications on a NVIDIA Jetson device.
Luca Barillaro, Giuseppe Agapito, Mario Cannataro
BIBM2
2022 A parallel software pipeline to select relevant genes for pathway enrichment
abstract
The continuous technological development of experimental omics technologies such as microarrays, allows to perform large scale genomics studies. After the initial enthusiasm, it became pretty clear that even the results provided by microarrays in form of lists of differential expressed genes (DEGs), were mainly as enigmatic as the first sequence of the genome, because these lists of DEGs are detached from the influenced biological mechanisms. Pathway enrichment analysis (PEA) supports researchers to provide the clues necessary to link DEGs to the influenced biological pathways and consequently to the underlying biological mechanisms and processes. Putting DEGs data sets in a suitable format for the PEA can be a tedious error-prone and laborious process even for bioinformaticians, who needs to perform it manually before to be ready for the PEA. To fill this lack, we present a parallel software pipeline which uploads a list of DEGs and automatically provides as results the enriched pathways.The parallel software pipeline is implemented in Python and provides the following automated actions: i) parallel splitting of DEGs in groups; ii) parallel building of the similarity matrices related to the DEGs groups; iii) parallel mapping of similarity matrices in networks; iv) parallel pathway enrichment analysis for each group of identified DEGs.Preliminary results shown that the pipeline can help to analyze DEGs and easily generate in a few minutes a list of pathway enrichment results that otherwise would require numerous hours of manual work and several different scripts.The parallel software pipeline provides a two-fold benefits: first, it contributes to speed up the computation of pathway enrichment, automating several steps currently performed manually. Second, it provides a more peculiar list of DEGs to calculate pathway enrichment, contributing to improve the relevance and significance of the enriched pathways.
Giuseppe Agapito, Mario Cannataro
PDP1
2022 Pathway integration and annotation: building a puzzle with non-matching pieces and no reference picture
abstract
Biological pathways are a broadly used formalism for representing and interpreting the cascade of biochemical reactions underlying cellular and biological mechanisms. Pathway representation provides an ontological link among biomolecules such as RNA, DNA, small molecules, proteins, protein complexes, hormones and genes. Frequently, pathway annotations are used to identify mechanisms linked to genes within affected biological contexts. This important role and the simplicity and elegance in representing complex interactions led to an explosion of pathway representations and databases. Unfortunately, the lack of overlap across databases results in inconsistent enrichment analysis results, unless databases are integrated. However, due to absence of consensus, guidelines or gold standards in pathway definition and representation, integration of data across pathway databases is not straightforward. Despite multiple attempts to provide consolidated pathways, highly related, redundant, poorly overlapping or ambiguous pathways continue to render pathways analysis inconsistent and hard to interpret. Ontology-based integration will promote unbiased, comprehensive yet streamlined analysis of experiments, and will reduce the number of enriched pathways when performing pathway enrichment analysis. Moreover, appropriate and consolidated pathways provide better training data for pathway prediction algorithms. In this manuscript, we describe the current methods for pathway consolidation, their strengths and pitfalls, and highlight directions for future improvements to this research area.
Giuseppe Agapito, Chiara Pastrello, Yun Niu, Igor Jurisica
Briefings Bioinform.1
2022 A statistical network pre-processing method to improve relevance and significance of gene lists in microarray gene expression studies
abstract
BACKGROUND: Microarrays can perform large scale studies of differential expressed gene (DEGs) and even single nucleotide polymorphisms (SNPs), thereby screening thousands of genes for single experiment simultaneously. However, DEGs and SNPs are still just as enigmatic as the first sequence of the genome. Because they are independent from the affected biological context. Pathway enrichment analysis (PEA) can overcome this obstacle by linking both DEGs and SNPs to the affected biological pathways and consequently to the underlying biological functions and processes. RESULTS: To improve the enrichment analysis results, we present a new statistical network pre-processing method by mapping DEGs and SNPs on a biological network that can improve the relevance and significance of the DEGs or SNPs of interest to incorporate pathway topology information into the PEA. The proposed methodology improves the statistical significance of the PEA analysis in terms of computed p value for each enriched pathways and limit the number of enriched pathways. This helps reduce the number of relevant biological pathways with respect to a non-specific list of genes. CONCLUSION: The proposed method provides two-fold enhancements. Network analysis reveals fewer DEGs, by selecting only relevant DEGs and the detected DEGs improve the enriched pathways' statistical significance, rather than simply using a general list of genes.
Giuseppe Agapito, Marianna Milano, Mario Cannataro
BMC Bioinform.1
2022 Nine quick tips for pathway enrichment analysis
abstract
Pathway enrichment analysis (PEA) is a computational biology method that identifies biological functions that are overrepresented in a group of genes more than would be expected by chance and ranks these functions by relevance. The relative abundance of genes pertinent to specific pathways is measured through statistical methods, and associated functional pathways are retrieved from online bioinformatics databases. In the last decade, along with the spread of the internet, higher availability of computational resources made PEA software tools easy to access and to use for bioinformatics practitioners worldwide. Although it became easier to use these tools, it also became easier to make mistakes that could generate inflated or misleading results, especially for beginners and inexperienced computational biologists. With this article, we propose nine quick tips to avoid common mistakes and to out a complete, sound, thorough PEA, which can produce relevant and robust results. We describe our nine guidelines in a simple way, so that they can be understood and used by anyone, including students and beginners. Some tips explain what to do before starting a PEA, others are suggestions of how to correctly generate meaningful results, and some final guidelines indicate some useful steps to properly interpret PEA results. Our nine tips can help users perform better pathway enrichment analyses and eventually contribute to a better understanding of current biology.
Davide Chicco, Giuseppe Agapito
PLoS Comput. Biol.2
2021 Comprehensive pathway enrichment analysis workflows: COVID-19 case study
abstract
Abstract The coronavirus disease 2019 (COVID-19) outbreak due to the novel coronavirus named severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has been classified as a pandemic disease by the World Health Organization on the 12th March 2020. This world-wide crisis created an urgent need to identify effective countermeasures against SARS-CoV-2. In silico methods, artificial intelligence and bioinformatics analysis pipelines provide effective and useful infrastructure for comprehensive interrogation and interpretation of available data, helping to find biomarkers, explainable models and eventually cures. One class of such tools, pathway enrichment analysis (PEA) methods, helps researchers to find possible key targets present in biological pathways of host cells that are targeted by SARS-CoV-2. Since many software tools are available, it is not easy for non-computational users to choose the best one for their needs. In this paper, we highlight how to choose the most suitable PEA method based on the type of COVID-19 data to analyze. We aim to provide a comprehensive overview of PEA techniques and the tools that implement them.
Giuseppe Agapito, Chiara Pastrello, Igor Jurisica
Briefings Bioinform.1
2021 Using BioPAX-Parser (BiP) to enrich lists of genes or proteins with pathway data
abstract
BACKGROUND: Pathway enrichment analysis (PEA) is a well-established methodology for interpreting a list of genes and proteins of interest related to a condition under investigation. This paper aims to extend our previous work in which we introduced a preliminary comparative analysis of pathway enrichment analysis tools. We extended the earlier work by providing more case studies, comparing BiP enrichment performance with other well-known PEA software tools. METHODS: PEA uses pathway information to discover connections between a list of genes and proteins as well as biological mechanisms, helping researchers to overcome the problem of explaining biological entity lists of interest disconnected from the biological context. RESULTS: We compared the results of BiP with some existing pathway enrichment analysis tools comprising Centrality-based Pathway Enrichment, pathDIP, and Signaling Pathway Impact Analysis, considering three cancer types (colorectal, endometrial, and thyroid), for a total of six datasets (that is, two datasets per cancer type) obtained from the The Cancer Genome Atlas and Gene Expression Omnibus databases. We measured the similarities between the overlap of the enrichment results obtained using each couple of cancer datasets related to the same cancer. CONCLUSION: As a result, BiP identified some well-known pathways related to the investigated cancer type, validated by the available literature. We also used the Jaccard and meet-min indices to evaluate the stability and the similarity between the enrichment results obtained from each couple of cancer datasets. The obtained results show that BiP provides more stable enrichment results than other tools.
Giuseppe Agapito, Mario Cannataro
BMC Bioinform.1
2021 Parallel and distributed association rule mining in life science: A novel parallel algorithm to mine genomics data
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
Inf. Sci.1
2020 An efficient and scalable SPARK preprocessing methodology for Genome Wide Association Studies
abstract
The importance of the use of high-performance software frameworks to analyze omics data obtained by using High-Throughput (HT) essays is widely recognized. HT methodologies comprise microarrays, Genome-Wide Association Studies (GWAS), and Next Generation Sequencing (NGS), which provide a vast amount of data per a single experiment. Each HT vendor provides to the users only the software frameworks and the proprietary libraries for the annotation, and summarization of raw data. Consequently, the needs of algorithms for the preprocessing and analysis of omics data arise. GWAS aims to highlight the association between genetic variants and diseases by examining single nucleotide polymorphisms (SNPs), which differ in a statistically significant way between cases and controls. The effectiveness of GWAS analysis increases with the number of analyzed samples per single experiment. GWAS data analyzed through the use of statistical methods can detect associations among a single allelic variant and the clinical conditions of samples. To overcome these limitations, and to make it possible to discover multiple associations among allelic variants, it is possible to use Association Rules mining. Consequently, the need for the introduction of scalable Association Rule Mining (ARM) algorithms able to analyze GWAS data arises. Hence, the use of high-performance data analytics framework is needed. For this purpose, we propose a software framework called GARMS (GWAS Association Rule Mining in Spark) built on top of Apache Spark for the preprocessing, and mining of association rules from GWAS data sets. GARMS comprises a two steps analysis methodology: (i) in the first step, the GWAS data are preprocessed, along with the identification of the frequent itemsets; (ii) in the second step, frequent itemsets are employed to mine association rules without scanning the input data. We implemented our algorithm, and we tested it on some synthetic GWAS data sets. Preliminary results confirm that our method may extract relevant association rules from GWAS data reducing the computational time.
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
PDP1
2020 BioPAX-Parser: parsing and enrichment analysis of BioPAX pathways
abstract
SUMMARY: Biological pathways are fundamental for learning about healthy and disease states. Many existing formats support automatic software analysis of biological pathways, e.g. BioPAX (Biological Pathway Exchange). Although some algorithms are available as web application or stand-alone tools, no general graphical application for the parsing of BioPAX pathway data exists. Also, very few tools can perform pathway enrichment analysis (PEA) using pathway encoded in the BioPAX format. To fill this gap, we introduce BiP (BioPAX-Parser), an automatic and graphical software tool aimed at performing the parsing and accessing of BioPAX pathway data, along with PEA by using information coming from pathways encoded in BioPAX. AVAILABILITY AND IMPLEMENTATION: BiP is freely available for academic and non-profit organizations at https://gitlab.com/giuseppeagapito/bip under the LGPL 2.1, the GNU Lesser General Public License. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Giuseppe Agapito, Chiara Pastrello, Pietro H. Guzzi, Igor Jurisica, Mario Cannataro
Bioinform.1
2020 cPEA: a parallel method to perform pathway enrichment analysis using multiple pathways databases
Giuseppe Agapito, Mario Cannataro
Soft Comput.1
2019 Association Rule Mining from large datasets of clinical invoices document
abstract
The concept of massive data generation nowadays affects several domains such as marketing including electronic invoices of large retailers, web access log files, healthcare, life sciences and so on. All these web activities introduced a new way to pay through the concept of electronic invoices (eInvoice), replacing the paper invoices. For these reasons, eInvoicing can be thought of as an innovative digital infrastructure for the issue, transmission, and storage of invoices. The availability of large volumes of eInvoices allows the discovery of new knowledge through data mining in these domains. Thus, users by using data mining can extract knowledge from large invoices documents. In this paper, we present a software tool for mining association rules from invoices produced in healthcare centers. In particular, the tool adopt a novel preprocessing methodology that provides merging, cleaning, formatting and summarization of eInvocies. The methodology can improve the quality of a huge amount of clinical invoices reducing the quantity of irrelevant data, making the remaining data suitable to mine information in form of association rules. The core of the tool allows to extract association rules from eInvoices; as a case study, we discuss the mined rules, highlighting the relationships among the purchased goods.
Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Sabrina Graziano, Mario Cannataro
BIBM1
2019 Pathway Analysis for SNP microarray data
abstract
Pathway Analysis (PA) is a powerful method for data analysis in genomics, most often applied to gene expression analysis, but little used to analyze variants such as Single Nucleotide Polymorphisms (SNPs). PA could allow the interpretation of variants concerning the biological processes in which the affected genes and proteins are involved. Currently, the available PA software tools are not able to automatically perform pathway analysis using SNPs data. PA software tools cannot deal natively with SNPs data, hence several software tools have to be used to put SNPs data in the proper format for the analysis. To overcome these limitations, we present SNP Microarray Pathway Analysis (MPA), a software tool able to discriminate relevant genes from SNP microarrays to use in PA analysis. MPA automatically identifies relevant SNPs using the well known Fisher's test, with which to perform PA. Pathway analysis in MPA is obtained employing the Hypergeometric function. As a result, MPA provides to the user the list of enriched pathways from the identified SNPs. MPA software tool along with the user guide and datasets, are available for download at https://gitlab.com/giuseppeagapito/mpa under the GPL v3.0 license.
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
BIBM1
2019 Mining Association Rules From Disease Ontology
abstract
The Disease Ontology (DO) is standardized, controlled vocabulary that contains information about inherited, developmental and acquired human diseases. Each DO term is associated with disease concepts through an annotation process. The relevance and the specificity of DO terms are often evaluated by its Information Content (IC). An important research area focus on the analysis of annotated data with the goal to extract knowledge. For example, the analysis of annotated data using Association Rules (AR) may supply meaningful knowledge, discovering relevant associations. Classical association rules methods consider all annotation equally, do not taking into account that the DO terms have different Information Content, i.e. different relevance. This implies the generation of association rules with low IC. In this paper we presents WARDO (Weighted Association Rule mining from Disease Ontology), a methodology based on the extraction od Weighted Association Rules from the DO Ontology considering the IC of terms. To assess our methodology, we tested WARDO on DO annotation datasets. WARDO is publicly available at https://gitlab.com/giuseppeagapito/wardo.
Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro
BIBM1
2018 S4S: RESTful Services to Collect, Integrate and Analyze SNPs and Clinical Data on the Web
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
BIBM1
2017 A software pipeline for multiple microarray data analysis
abstract
Microarray platforms such as Gene expression, oligonucleotide (GeneChip), and cDNA (complementary DNA) can play a critical role in the understanding of genome sequencing, by providing information for hundreds or thousands genes in a single assay, information not available by using other methodologies of investigation. Microarray experiments aim to survey patterns of gene expression by assaying the expression levels of thousands of genes in parallel in a single assay. Thus, microarray are known as high-throughput technologies, allowing the investigation of genetic variations underlying the inter-individual variability in drug pharmacokinetics/pharmacodynamics, providing a complete understanding of gene function, regulation, and interactions. Thus, to exploit all the power of this massive amount of data in the short possible time (before that data becomes obsolete), the necessity to develop databases and software tools for efficient data collection and analysis arises. The establishing of efficient and scalable software tools avoids that researcher will be overwhelming from this ever-growing flow of data. At this reason, a preliminary design of a platform named microPipe, for the multiple analysis of microarray data sets is proposed. Specifically, the paper outlines the main issues and challenges relative to the design of such a platform.
Giuseppe Agapito, Mario Cannataro
BIBM1
2017 Parallel and Cloud-Based Analysis of Omics Data: Modelling and Simulation in Medicine
abstract
High throughput experimental platforms and diagnostic equipments available in clinical settings and in research laboratories, such as magnetic resonance imaging, microarray, mass spectrometry and next-generation sequencing, are producing an increasing volume of clinical and omics data. Moreover, Electronic Patients Records (EPRs), eHealth systems, personal mobile sensors and Social Networks are collecting an overwhelming volume of health and life style data that may be integrated with clinical data and more and more is used for the real-time monitoring of patient's health. This poses new issues in terms of secure data storage, effective models for data integration, efficient algorithms for data analysis, new models for health monitoring, that may be addressed, among the others, using high performance computing solutions. Parallel computing and Cloud Computing may offer efficient and scalable solutions in an orthogonal way. In fact, parallel, bioinformatics software, that exploit off-the-shelf high performance computers, may be used to preprocess and analyze omics data at a lower layer, for instance to highlight genetic variation associated with complex diseases. On the other hand, Cloud Computing offers large scale data storage, data sharing services, on-demand anytime and anywhere access to resources and applications, for the realization of elastic and scalable applications and services. Motivated by the increasing use of parallel computing and cloud computing in life sciences, in this paper we survey both parallel bioinformatics algorithms for the parallel preprocessing and statistical and data mining analysis of omics data, as well as Cloud-based healthcare and biomedicine services and systems for large scale applications. Moreover, the paper underlines main issues and problems related to the use of such platforms for the storage and analysis of health data, with special focus to the security and privacy of patients data, that are particularly important in fields such as personalized medicine. Finally, the paper presents some case studies about the parallel and distributed modelling and simulation in medicine and biology.
Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Gionata Fragomeni, Giuseppe Tradigo, Pierangelo Veltri, Mario Cannataro
PDP1
2016 DIETOS: A recommender system for adaptive diet monitoring and personalized food suggestion
abstract
Nowadays there is a widespread diffusion of mobile applications for weight and diet management. Even though, the most popular apps are not usually experimented in clinical contexts, as well as apps are not supported by medical evidence. Further research is necessary to assess the effectiveness of apps for weight and diet management. Moreover, there are few examples of food recommender systems that provide to the users nutritional facts about suitable food choices and take into account individual physiological status and environmental situations. We propose DIETOS (DIET Organizer System), a recommender system for the adaptive delivery of nutrition contents to improve the quality of life of both healthy people and individuals affected by chronic diet-related diseases. The proposed system is able to build a user's health profile, and provides individualized nutritional recommendation according to the health profile. The profile is created through the use of dynamic real-time questionnaires prepared by medical doctors and compiled by the users. The health profile includes information about health status and eventual chronic diseases. The first prototype of the system (available online at http://www.easyanalysis.it/dietos), includes a catalogue of typical Calabrian foods compiled by nutrition specialists (Calabria is a region of the southern Italy). DIETOS can suggest not only the use of specific foods compatible with the health status, but also it may give dietary indications related to some specific pathologies or health conditions.
Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Mario Cannataro, Mariadelina Simeoni, Ilaria Care, Theodora Lamprinoudi, Giorgio Fuiano, Arturo Pujia
WiMob1
2016 Methodologies and experimental platforms for generating and analysing microarray and mass spectrometry-based omics data to support P4 medicine
abstract
Predictive, preventive, personalized and participatory (P4) medicine is an emerging medical model that is based on the customization of all medical aspects (i.e. practices, drugs, decisions) of the individual patient. P4 medicine presupposes the elucidation of the so-called omic world, under the assumption that this knowledge may explain differences of patients with respect to disease prevention, diagnosis and therapies. Here, we elucidate the role of some selected omics sciences for different aspects of disease management, such as early diagnosis of diseases, prevention of diseases, selection of personalized appropriate and optimal therapies based on molecular profiling of patients. After introducing basic concepts of P4 medicine and omics sciences, we review some computational tools and approaches for analysing selected omics data, with a special focus on microarray and mass spectrometry data, which may be used to support P4 medicine. Some applications of biomarker discovery and pharmacogenomics and some experiences on the study of drug reactions are also described.
Pietro H. Guzzi, Giuseppe Agapito, Marianna Milano, Mario Cannataro
Briefings Bioinform.2
2016 Extracting Cross-Ontology Weighted Association Rules from Gene Ontology Annotations
abstract
Gene Ontology (GO) is a structured repository of concepts (GO Terms) that are associated to one or more gene products through a process referred to as annotation. The analysis of annotated data is an important opportunity for bioinformatics. There are different approaches of analysis, among those, the use of association rules (AR) which provides useful knowledge, discovering biologically relevant associations between terms of GO, not previously known. In a previous work, we introduced GO-WAR (Gene Ontology-based Weighted Association Rules), a methodology for extracting weighted association rules from ontology-based annotated datasets. We here adapt the GO-WAR algorithm to mine cross-ontology association rules, i.e., rules that involve GO terms present in the three sub-ontologies of GO. We conduct a deep performance evaluation of GO-WAR by mining publicly available GO annotated datasets, showing how GO-WAR outperforms current state of the art approaches.
Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 Overall Survival Analyzer: A software tool to analyze genotyping and clinical data enriched with temporal events
abstract
The estimation of survival distributions of patients is an important current problem in clinical oncology. The current trend is to integrate molecular data (such as genomic data) with clinical data (e.g. cancer type, stage of the disease, etc) and then to link survival distributions to molecular profile of patients. Recently, the Affymetrix DMET (Drug Metabolizing Enzymes and Transporters) microarray technology has enabled the possibility to determine the allelic variants of a patient and to relate them to phenotype (e.g. drug toxicity). Therefore, the analysis of survival distribution of patients starting from their profile obtained using DMET data may reveal important knowledge to clinicians. In order to provide support to this analysis we propose Overall Survival Analyzer (OS-Analyzer), a software tool able to compute the Overall Survival and Progression-Free Survival (PFS). The tool is able to perform an automatic analysis of data avoiding wasting time on the manual analysis. OS-Analyzer is available to download at the follows web address: https://sites.google.com/site/overallsurvivalanalyzer/.
Giuseppe Agapito, Pietro H. Guzzi, C. Botta, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro
BIBM1
2015 DMET-Miner: Efficient discovery of association rules from pharmacogenomic data
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
J. Biomed. Informatics1
2014 Improving annotation quality in gene ontology by mining cross-ontology weighted association rules
abstract
The Gene Ontology (GO) is the major resource of annotations for genes and proteins. Despite the presence of large efforts to avoid errors and inconsistencies, some unreliabilities are still present. In particular electronically inferred annotations are more unreliable than manual ones and their number is growing. Thus, the need for an accurate evaluation of annotations in an automatic way arises. In the past, some approaches for improving annotation consistencies have been proposed using association rule mining to discover hidden relationships among GO terms. However such approaches consider all the GO terms equally, while GO terms have different Information Content, i.e. different relevance. Consequently we designed a novel algorithm, (GO-WAR), Mining Weighted Association Rules from GO, that is based on the extraction of weighted association rules considering the IC of terms. We evaluated our algorithm considering seven different species and all the GO ontologies. In all the experiments GO-WAR outperformed state of the art approaches.
Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro
BIBM1
2014 DMET-miner: Efficient learning of association rules from genotyping data for personalized medicine
abstract
Recent developments of microarray technology enable the investigation of allelic variants that may be correlated to phenotypes. In particular the Affymetrix DMET (Drug Metabolism Enzymes and Transporters) platform enables the simultaneous investigation of all the genes that are related to drug absorption, distribution, metabolism and excretion (ADME) and it has been used in clinical studies. In a previous work we developed DMET-Analyzer, a platform able to automatize the study of allelic variants, that has been validated in clinical studies. DMET-Analyzer is able to correlate a single variant for each probe (related to a portion of a gene) through the use of the Fisher test, on the other hand it is unable to discover multiple associations among allelic variants. To overcome those limitations, here we propose DMET-Miner, that is able to correlate the presence of a set of allelic variants by employing an Apriori-like discovery strategy. Preliminary experiments on a synthetic DMET dataset.
Pietro H. Guzzi, Giuseppe Agapito, Maria Teresa Di Martino, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro
BIBM2
2014 Biases in information content measurement of gene ontology terms
abstract
The Gene Ontology (GO) is used to achieve information about gene and protein functions by using a structured vocabulary of terms (GO Terms). GO Terms are related to biological concepts such as proteins or genes through the annotation process. There exist many different annotation processes identified by different evidence codes (EC). Annotated data are stored in public databases such as the Gene Ontology Annotation (GOA) database. Each term has a different specificity also referred to as Information Content (IC) of terms. Both the structure of GO and the corpora of annotation are continuously subject to change due to novel experimental findings. This process is often referred to as ontology evolution. This work focuses on how changes of annotations affect the IC of terms. The study confirms that statistically significant difference among many whole GOA versions exists on each species. Furthermore, there is also a statistically significant difference considering MF taxonomy for human, yeast, worm and fly. These results convey that annotation corpora changes have a high impact on IC.
Marianna Milano, Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
BIBM2
2014 coreSNP: Parallel Processing of Microarray Data
abstract
The availability of high-throughput technologies, such as next generation sequencing and microarray, and the diffusion of genomics studies to large populations are producing an increasing amount of experimental data. In particular, pharmacogenomics studies the impact of genetic variation on drug response in patients and correlates gene expression or single nucleotide polymorphisms (SNPs) with the toxicity or efficacy of a drug, with the aim to improve drug therapy with respect to the patients’ genotype ensuring maximum efficacy with minimal adverse effects. However, the storage, preprocessing, and analysis of experimental data are becoming a main bottleneck in the pharmacogenomics analysis pipeline, due to the increasing number of genes and patients investigated. This paper presents a new parallel software tool named coreSNP for the parallel preprocessing and statistical analysis of DMET (Drug Metabolism Enzymes and Transporters) SNP microarray data produced by Affymetrix for pharmacogenomics studies. The scalable multi-threaded implementation of coreSNP allows to handle the huge volumes of experimental pharmacogenomics data in a very efficient way, while its easy to use graphical user interface and its ability to annotate significant SNPs allow biologists to interpret the results easily. Performance evaluation conducted using real datasets shows good speed-up and scalability and effective response times.
Pietro H. Guzzi, Giuseppe Agapito, Mario Cannataro
IEEE Trans. Computers2
2013 Visualization of protein interaction networks: problems and solutions
abstract
BACKGROUND: Visualization concerns the representation of data visually and is an important task in scientific research. Protein-protein interactions (PPI) are discovered using either wet lab techniques, such mass spectrometry, or in silico predictions tools, resulting in large collections of interactions stored in specialized databases. The set of all interactions of an organism forms a protein-protein interaction network (PIN) and is an important tool for studying the behaviour of the cell machinery. Since graphic representation of PINs may highlight important substructures, e.g. protein complexes, visualization is more and more used to study the underlying graph structure of PINs. Although graphs are well known data structures, there are different open problems regarding PINs visualization: the high number of nodes and connections, the heterogeneity of nodes (proteins) and edges (interactions), the possibility to annotate proteins and interactions with biological information extracted by ontologies (e.g. Gene Ontology) that enriches the PINs with semantic information, but complicates their visualization. METHODS: In these last years many software tools for the visualization of PINs have been developed. Initially thought for visualization only, some of them have been successively enriched with new functions for PPI data management and PIN analysis. The paper analyzes the main software tools for PINs visualization considering four main criteria: (i) technology, i.e. availability/license of the software and supported OS (Operating System) platforms; (ii) interoperability, i.e. ability to import/export networks in various formats, ability to export data in a graphic format, extensibility of the system, e.g. through plug-ins; (iii) visualization, i.e. supported layout and rendering algorithms and availability of parallel implementation; (iv) analysis, i.e. availability of network analysis functions, such as clustering or mining of the graph, and the possibility to interact with external databases. RESULTS: Currently, many tools are available and it is not easy for the users choosing one of them. Some tools offer sophisticated 2D and 3D network visualization making available many layout algorithms, others tools are more data-oriented and support integration of interaction data coming from different sources and data annotation. Finally, some specialistic tools are dedicated to the analysis of pathways and cellular processes and are oriented toward systems biology studies, where the dynamic aspects of the processes being studied are central. CONCLUSION: A current trend is the deployment of open, extensible visualization tools (e.g. Cytoscape), that may be incrementally enriched by the interactomics community with novel and more powerful functions for PIN analysis, through the development of plug-ins. On the other hand, another emerging trend regards the efficient and parallel implementation of the visualization engine that may provide high interactivity and near real-time response time, as in NAViGaTOR. From a technological point of view, open-source, free and extensible tools, like Cytoscape, guarantee a long term sustainability due to the largeness of the developers and users communities, and provide a great flexibility since new functions are continuously added by the developer community through new plug-ins, but the emerging parallel, often closed-source tools like NAViGaTOR, can offer near real-time response time also in the analysis of very huge PINs.
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro
BMC Bioinform.1
2013 Visual Data Mining of Biological Networks: One Size Does Not Fit All
abstract
High-throughput technologies produce massive amounts of data. However, individual methods yield data specific to the technique used and biological setup. The integration of such diverse data is necessary for the qualitative analysis of information relevant to hypotheses or discoveries. It is often useful to integrate these datasets using pathways and protein interaction networks to get a broader view of the experiment. The resulting network needs to be able to focus on either the large-scale picture or on the more detailed small-scale subsets, depending on the research question and goals. In this tutorial, we illustrate a workflow useful to integrate, analyze, and visualize data from different sources, and highlight important features of tools to support such analyses.
Chiara Pastrello, David Otasek, Kristen Fortney, Giuseppe Agapito, Mario Cannataro, Elize Shirdel, Igor Jurisica
PLoS Comput. Biol.4
2012 DMET-Analyzer: automatic analysis of Affymetrix DMET Data
abstract
BACKGROUND: Clinical Bioinformatics is currently growing and is based on the integration of clinical and omics data aiming at the development of personalized medicine. Thus the introduction of novel technologies able to investigate the relationship among clinical states and biological machineries may help the development of this field. For instance the Affymetrix DMET platform (drug metabolism enzymes and transporters) is able to study the relationship among the variation of the genome of patients and drug metabolism, detecting SNPs (Single Nucleotide Polymorphism) on genes related to drug metabolism. This may allow for instance to find genetic variants in patients which present different drug responses, in pharmacogenomics and clinical studies. Despite this, there is currently a lack in the development of open-source algorithms and tools for the analysis of DMET data. Existing software tools for DMET data generally allow only the preprocessing of binary data (e.g. the DMET-Console provided by Affymetrix) and simple data analysis operations, but do not allow to test the association of the presence of SNPs with the response to drugs. RESULTS: We developed DMET-Analyzer a tool for the automatic association analysis among the variation of the patient genomes and the clinical conditions of patients, i.e. the different response to drugs. The proposed system allows: (i) to automatize the workflow of analysis of DMET-SNP data avoiding the use of multiple tools; (ii) the automatic annotation of DMET-SNP data and the search in existing databases of SNPs (e.g. dbSNP), (iii) the association of SNP with pathway through the search in PharmaGKB, a major knowledge base for pharmacogenomic studies. DMET-Analyzer has a simple graphical user interface that allows users (doctors/biologists) to upload and analyse DMET files produced by Affymetrix DMET-Console in an interactive way. The effectiveness and easy use of DMET Analyzer is demonstrated through different case studies regarding the analysis of clinical datasets produced in the University Hospital of Catanzaro, Italy. CONCLUSION: DMET Analyzer is a novel tool able to automatically analyse data produced by the DMET-platform in case-control association studies. Using such tool user may avoid wasting time in the manual execution of multiple statistical tests avoiding possible errors and reducing the amount of time needed for a whole experiment. Moreover annotations and the direct link to external databases may increase the biological knowledge extracted. The system is freely available for academic purposes at: https://sourceforge.net/projects/dmetanalyzer/files/
Pietro H. Guzzi, Giuseppe Agapito, Maria Teresa Di Martino, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro
BMC Bioinform.2