VLDB 2026 Research / reviewers in the wild / expert
Mario Cannataro
dblp:c/MarioCannataro
· DBLP profile ↗
103ranked-venue papers
25as first author
36since 2021 · last 2025
0000-0003-1502-2387ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 70 · 9 first-author · 27 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 1 since 2021Systems, architecture and hardware · 15 · 13 first-authorHuman-computer interaction and ubiquitous computing · 15 · 6 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multiview Learning Pipeline for Contrast-Enhanced Mammography: Comparative Evaluation of Different Fusion StrategiesabstractAccurate characterization of breast lesions remains a challenge in oncological imaging. Contrast-Enhanced Spectral Mammography (CESM) has recently emerged as a cost-effective alternative to magnetic resonance imaging (MRI), providing both low-energy (LE) and subtracted (SUB) images. While histopathology remains the diagnostic reference, growing evidence suggests that morphological and functional information extracted from CESM can support the discrimination between malignant and benign lesions. In this paper, we present a reproducible multiview learning pipeline for CESM, designed to evaluate three fusion strategies -early, late, and hybrid- across both LE and SUB images. Experiments conducted on the public CDD-CESM dataset of 326 patients show that single-view models on SUB images achieve the best performance (Acc$=84.5 \%$, AUC$=88.7 \%$), whereas multiview fusion does not consistently improve over the strongest single-view baseline. These results confirm the high discriminative power of SUB images and suggest that naive fusion may add limited value. Our contribution is to provide the first transparent benchmark of CESM fusion strategies on a public dataset, enabling like-for-like comparisons across future studies. From a clinical perspective, this reproducible framework may inform the design of decision-support tools, with the potential to reduce unnecessary biopsies and assist radiologists in challenging diagnostic settings. Valeria Popello, Chiara Zucco, Marianna Milano, Mario Cannataro |
BIBM | 4 |
| 2025 | Comparative Analysis of Algorithms and Computational Architectures for Efficient Biological Data ProcessingabstractThis paper presents a comparative analysis of algorithms and computational architectures for the efficient processing of biological data, particularly focusing on sparse matrix representations for biological pathways. Biological networks, such as protein-protein interaction (PPI) networks and metabolic pathways, are represented as graphs with nodes corresponding to molecular entities and edges representing biochemical interactions. These graphs are typically sparse, leading to challenges in data storage and computational efficiency. The study examines diverse sparse matrix formats, including Coordinate (COO), Compressed Sparse Row (CSR), and Compressed Sparse Column (CSC), emphasizing their application in biological data analysis. It further explores the potential of modern high-performance computing architectures, such as Graphics Processing Units (GPUs), to accelerate the execution of graph-based algorithms. Algorithms such as Dijkstra’s, Breadth-First Search (BFS), and PageRank are adapted to leverage the efficiencies of sparse matrix representations, optimizing the analysis of large-scale biological networks for improved performance and scalability. Performance evaluations show that GPUs significantly outperform Central Processing Units (CPUs) in processing largescale biological networks, reducing execution time and energy consumption while enhancing scalability. This research demonstrates how the use of an appropriate sparse matrix format and computational architecture can optimize the analysis of complex biological networks, providing insights into biological processes and therapeutic targets. Giuseppe Agapito, Gaetano Guardasole, Mario Cannataro |
PDP | 3 |
| 2025 | fDESI: An open-source web application for visual bioinformatics pipeline designabstractIn bioinformatics, data analysis often involves complex workflows that require processing vast amounts of biological data. Pipelines play a crucial role in automating and organizing these sequences of computational tasks, ensuring reproducibility, efficiency, and scalability. In this paper, we presented F LENP DESI gner ( fDESI ), an open-source and user-friendly web application for visual bioinformatics pipeline design. It allows intuitive pipeline creation through a drag-and-drop interface, generating pipeline descriptions in a JSON-based meta-language. Furthermore, it provides an all-in-one environment to overcome the need for external tools, beyond the third-party software within the pipeline itself. Our contribution includes a flexible framework for pipeline modelling, built-in functionalities requiring no programming expertise, an integrated execution engine, and an open-source graphical interface for streamlined bioinformatics workflow design. Pietro Cinaglia, Mario Cannataro |
Neurocomputing | 2 |
| 2025 | Ten practical tips and tricks to improve the effectiveness of biological network alignmentabstractNetwork alignment (NA) is a computational methodology employed to compare biological networks across different species or conditions. By identifying conserved structures, functions, and interactions, NA provides invaluable insights into shared biological processes, evolutionary relationships, and system-level behaviors. This manuscript presents a comprehensive overview of NA methodologies, including the importance of preprocessing network data, selecting suitable input formats, and understanding diverse network types such as attributed, temporal, and multilayer networks. Additionally, it explores key challenges such as seed nodes selection, algorithm configuration, and cross-species alignment, emphasizing the necessity of integrating functional annotations, sequence similarity, and network topology for biologically meaningful results. Various NA strategies, including Local and Global Network Alignment, are discussed alongside their respective advantages and limitations. Practical recommendations for effectively documenting and visualizing NA experiments are also provided, ensuring reproducibility and clarity in research. By leveraging diverse alignment tools and adopting best practices, researchers can unlock the potential of NA to advance our understanding of complex biological systems. Giuseppe Agapito, Mario Cannataro, Pietro Cinaglia, Marianna Milano |
PLoS Comput. Biol. | 2 |
| 2024 | Multi Weighted Graphs as Magnifier to Discover Hidden Cross-Interactions among Biological PathwaysabstractScientists and researchers need robust and flexible methods to identify interactions among multiple biological pathways, which unveil crucial insights otherwise obscured when pathways are analyzed in isolation. Understanding these complex interrelations is vital for comprehensively mapping the processes relevant to their research.Biological pathways are categorized into three main classes: signaling, metabolic, and regulatory.The pathways representation can be significantly enhanced using multi weighted graphs models, a concept from graph theory where multiple edges can connect nodes. This allows for a more nuanced representation of the complex interactions and multiple relationships between pathway elements. In this model, each node represents a pathway element, and the multiple edges between nodes depict various types of interactions, capturing the dynamic and multifaceted nature of cellular processes.To address the needs of researchers for more sophisticated and accurate analysis, we have developed a new methodology to convert and represent biological pathways as multi weighted graphs. In this way, it is possible to recall the graph theory metrics to evaluate the elements within the graphs. We show that the use of multi weighted graphs to connect several pathways, allowing to identify hidden cross-pathways links, e.g., genes shared among pathways, through the use of PageRank and Kats index. The discovered genes allow to improve of some order the magnitude of pathway enrichment analysis results, as conveyed from the reached p-values. Giuseppe Agapito, Mario Cannataro, Gaetano Guardasole |
BIBM | 2 |
| 2024 | Modeling UGT2B7 and NR1I3 genes through multilayer network to highlight hidden link with taxane neurotoxicityabstractBreast cancer (BC) remains a leading global malignancy, with taxane-based treatment (TBT), including paclitaxel and docetaxel, significantly enhancing outcomes across early-stage, locally advanced, and metastatic BC. Despite its efficacy, TBT can cause severe taxane-related peripheral neurotoxicity (TrPN), which limits dosing and affects patient quality of life. TrPN is cumulative, primarily sensory, and unpredictable, with genetic predisposition playing a key role in its development. Previous pharmacogenomic (PGx) study identified five single nucleotide polymorphisms (SNPs) in the NR1I3 and UGT2B7 genes, linked to protection against severe TrPN in BC patients. ROC analysis validated these findings in an independent BC dataset. Neuroprotective effects were in patients homozygous for allele variants 2 with ultrametabolizer phenotype, which accelerates taxane inactivation and reduces treatment efficacy, potentially worsening prognosis. NR1I3 and UGT2B7 genes could serve as predictive biomarkers for TrPN and taxane bioavailability. Further network and pathway enrichment analyses (PEA) provided insights into the molecular pathways involved in TrPN and BC progression. Network analysis, a powerful tool in computational biology, integrates genomic and proteomic data to uncover drug-response mechanisms. This approach advances personalized medicine by identifying key genes and biomarkers for predicting treatment response and adverse effects, offering new therapeutic opportunities. Giuseppe Agapito, Marianna Milano, Francesca Scionti, Nicoletta Staropoli, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro, Mariamena Arbitrio |
BIBM | 7 |
| 2024 | A Graph Neural Network based fMRI classificationabstractIn this paper, we aim to describe a novel deep learning model (a machine learning subclass) to classify fMRI. We used a publicly available dataset to train and test the model which involved patients affected by depression.Our model is based on a specific deep neural network, the Graph Attention Network (GAT) which has proven its strength in dealing with graph data representation. The novelty of our approach is that it is based on the extraction from the original fMRIs of graph representation then passed to the deep learning model.We performed this crucial phase by using a Matlab based toolbox, CONN, which helped in data preparation and graph representation extraction. We then used the extracted fMRI representations to feed, train, and finally test our deep learning model.While classification results were encouraging, achieving approximately 73% accuracy, another aspect that we investigated was focused on the comparison of three architectural solutions, focusing on power consumption. We used an Apple Silicon platform compared to a NVIDIA based laptop and an edge device of the NVIDIA Jetson family. Luca Barillaro, Marianna Milano, Giuseppe Agapito, Mario Cannataro |
BIBM | 4 |
| 2024 | An information system for cataloging and annotating images of scoliosisabstractAdolescent idiopathic scoliosis (AIS) is a spinal deformity that tends to get worse as children grow and requires constant monitoring. Current AI models are trained on datasets consisting mainly of X-ray images of scoliosis patients, typically used to predict the Cobb angle. We aim to create a valuable and alternative dataset to traditional radiographic images that can be used to train AI models for both the detection and assessment of scoliosis, starting from images of scoliosis cases captured on smartphone and annotated. Our study may make up for limitations of invasive methods and address the challenge of lack of data for developing effective AI models. Lorella Bottino, Pietro Cinaglia, Mario Cannataro |
BIBM | 3 |
| 2024 | A GPU-based method for network alignmentabstractNetwork graph models are a powerful tool for handling objects, in terms of interactions and relationships. For instance, in bioinformatics networks are applied for analysing complex biological systems, topologically, as well as for investigating their own homology. In this context, the pairwise network alignment can be performed for producing a set of node-to-node mappings from a source network to a target one. For instance, network alignment can be applied for porting knowledge from a simpler to a more complex system (e.g., a simpler biological organism to a complex one). In this paper, we presented a GPU-based method for the pairwise Local Network Alignment, implemented by using a greedy strategy. It leverages GPU parallelism to accelerate the large-scale computation of a node similarity matrix of interest. Our experimentation showed a relevant improvement in terms of runtime when processing is executed on GPU, in comparison to the traditional CPU computing. Briefly, our method has proven to be an effective solution for the alignment of large networks, especially with dense similarity matrices, thus proving to be an ideal candidate for use in real-world study cases. Pietro Cinaglia, Mario Cannataro |
BIBM | 2 |
| 2024 | Ethics of Artificial Intelligence: challenges, opportunities and future prospectsabstractArtificial Intelligence (AI) has rapidly transformed numerous sectors, including healthcare, justice, and commerce, providing substantial benefits while also raising complex ethical questions: this article explores the main ethical challenges associated with AI, focusing on issues such as algorithmic bias, data privacy and security, transparency, and accountability. The importance of Explainable Artificial Intelligence (XAI) in enhancing the interpretability of algorithmic decisions is emphasized, particularly in healthcare, where model opacity can have a direct impact on patient outcomes. The paper further examines regulatory frameworks and ethical guidelines, including the European Union’s AI Act, advocating for a multidisciplinary approach that combines innovation and accountability to develop AI systems that respect human rights and foster user trust. In conclusion, the article underscores the need for interdisciplinary collaboration and adaptable regulations to ensure the ethical development of AI, promoting fairness and transparency. Fedra Rosita Falvo, Mario Cannataro |
BIBM | 2 |
| 2024 | Explainability techniques for Artificial Intelligence models in medical diagnosticabstractThe integration of artificial intelligence (AI) techniques into clinical settings presents critical challenges due to the opacity of machine learning models, often referred to as "black boxes": this study explores the application of explainability techniques, specifically Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP), in the context of medical diagnostics. Using the "Diabetes Health Indicators Dataset", we applied Logistic Regression as a predictive model to identify key risk factors for diabetes and evaluate the ability of explainability techniques to improve transparency and interpretability. The results demonstrate that SHAP provides a detailed global and local understanding of feature importance, offering clinicians insights into key predictors such as HighBP, CholCheck, and GenHlth; LIME complements this by delivering intuitive explanations for individual predictions, enabling rapid and accessible interpretation. The combination of these techniques enhances trust in AI systems by providing both comprehensive insights and actionable explanations. Challenges related to computational complexity, scalability, and the integration of these methods into clinical workflows are also discussed, along with recommendations for future research aimed at developing scalable, interpretable AI models for ethical and responsible medical use. Fedra Rosita Falvo, Mario Cannataro |
BIBM | 2 |
| 2024 | A Multilayer Network-Based Method for Brain Connectivity Analysis from EEG DataabstractMultilayer networks (MLNs) have emerged as a critical tool in the field of medicine, particularly in neuroscience, owing to their capacity to model the complex interactions between brain regions across multiple dimensions. Unlike traditional single-layer network approaches, which typically focus on functional or structural connectivity within a single frequency band, MLN provides a richer, more comprehensive framework that captures the dynamic and multi-frequency nature of brain activity. In this work, we propose the development of a pipeline for the design and analysis of multilayer brain networks based on electroencephalogram (EEG) data. The primary object is to explore how MLNs can be utilized to analyze brain activity by capturing both intra- and inter-frequency interactions, that coordinate the different neural processes. The EEG data used in this study come from a cohort of 75 patients, including 25 healthy subjects, 25 with psychogenic non-epileptic seizures (PNES), and 25 with epilepsy. The results revealed significant differences in both the structure of the graphical network representation and the multilayer network analysis across the groups studied. These findings underscore the potential of multilayer networks (MLNs) to offer valuable insights into the distinct network patterns associated with various neurological conditions, providing a promising framework for advancing research into the complexity of brain network interactions. Ilaria Lazzaro, Marianna Milano, Chiara Zucco, Miriam Sturniolo, Franco Pucci, Antonio Gambardella, Mario Cannataro |
BIBM | 7 |
| 2024 | Evolution of medical reports over time: an analysis using Dynamic Topic ModelingabstractIn recent years, the availability of data has increased across various fields, with a greater focus on the medical sector. Medical records, clinical charts, and diagnostic reports provide a large amount of information. The complexity of this data requires the use of advanced analytical methods to extract reliable and important information. This study aims to test Text Mining techniques, particularly Dynamic Topic Modeling (DTM), to monitor the evolution of medical practices within surgical records related to the descriptions of surgeries over time. By applying DTM to a large dataset of medical reports from a urology clinic, we were able to identify emerging topics, highlighting changes in topics over time and whether one or more surgeries are more or less frequent in the specific years 2022 and 2023. This method not only extracted emerging topics from the descriptions of surgical interventions but also provided a precise overview of how clinical practices and patient outcomes have evolved. Through this analysis, we can gain a more detailed understanding of therapeutic strategies and how they have evolved during the specified years, contributing to future clinical decision-making and improving patient care. The findings of this research provide detailed information on the importance of using advanced techniques to better understand medical data and its implications for healthcare practices. Maria Chiara Martinis, Chiara Zucco, Antonio Amodeo, Vincenzo Facente, Francesco Greco, Mario Cannataro |
BIBM | 6 |
| 2024 | Visualization of Comorbidities in Inflammatory Bowel Diseases through NetworksabstractInflammatory Bowel Diseases (IBD), including Crohn’s Disease (CD) and Ulcerative Colitis (UC), are chronic conditions characterized by a complex network of comorbidities that significantly impact patients’ quality of life.In this study, we employed a Network Visualization approach to explore the comorbidities associated with IBD, using data extracted from the clinical records of patients at the Digestive Pathophysiology Unit of the University Hospital “Renato Dulbecco”.By constructing networks, we mapped and compared the comorbidities of Crohn’s Disease and Ulcerative Colitis, considering gender differences as well. The graphs produced offer a clear view clearly shows the major comorbidities and their distribution, highlighting significant differences between the two diseases and across genders.This visual approach facilitates a better understanding of the complexity of clinical interactions in IBD, suggesting new perspectives for personalized and multidisciplinary therapeutic approaches in managing these diseases. Rosarina Vallelunga, Marianna Milano, Rocco Spagnuolo, Lidia Giubilei, Evelina Suraci, Mario Cannataro, Francesco Luzza |
BIBM | 6 |
| 2024 | Efficient GPU Processing Method to Analyze Large GWAS Data SetsabstractGenome-wide association studies (GWAS) have ex-panded rapidly, generating large genetic data sets that require advanced computational strategies for efficient analysis. This study introduces a novel method that uses Graphics Processing Units (GPUs) to analyze large scale GWAS data sets, enhancing computational efficiency and reducing processing time. Our approach leverages GPUs' inherent parallel processing capabilities to significantly accelerate computation-intensive tasks commonly encountered in GWAS, such a Genetic Risk Score (GRS). We explain how our GPU-GRS computational approach out-performs traditional CPU (Central Processing Unit) methods, outlining the architectural optimizations that enable parallel processing to handle large genomic data sets more effectively. We also demonstrate the scalability of our approach by ana-lyzing increasingly large synthetic G WAS data sets, showcasing its ability to manage the growing size and complexity of genetic data efficiently. Our findings suggest that adopting GPU-based methods in G WAS analysis can play a pivotal role in the era of big genomic data, offering a path towards more time-efficient and scalable solutions. Giuseppe Agapito, Gaetano Guardasole, Mario Cannataro |
PDP | 3 |
| 2023 | Use Predictive Learning Model to Tackle Data Breaches in Healthcare DomainabstractHealthcare data breaches are a growing problem that seriously threatens patient privacy, the reputation and trustworthiness of both public and private healthcare organizations. The aim of this paper is to elucidate the severity of healthcare data breaches, their potential impact on patient privacy and healthcare organizations, providing a predictive data breaches transformer conceptual architecture that can monitoring users actions that can result in possible system security violation and consequently in data breaches. At this regard, we introduce the description of a concept architecture for implementing a predictive data breaches transformer highlighting weakness and strengthens, and in the same time how the adoption of a predictive system can significantly limit the risk of the onset of possible data breaches by making the operator more aware in carrying out his activity in the processing of each type of data including personal health data. Giuseppe Agapito, Mario Cannataro, Pietro Cinaglia, Gaetano Guardasole, Marianna Milano |
BIBM | 2 |
| 2023 | Explanation of machine learning models for predicting obesity level using Shapley valuesabstractObesity is a multifactorial disease that includes genetic, biological and behavioral factors. Training machine learning algorithms on some of variables related to these factors can help healthcare professionals equip themselves with tools capable of predicting obesity. But for these tools to gain trust they must be understandable and explainable. SHAP and LIME are two methodologies that allow you to achieve this objective. Lorella Bottino, Mario Cannataro |
BIBM | 2 |
| 2023 | A method for modelling and executing customized pipelines in serverless computingabstractServerless computing is an emerging cloud service for executing distributed applications on cloud architecture. The possibility of performing functions without the need to manage any type of infrastructure has made this methodology particularly adopted in several fields, e.g., data processing and above all in parallel computing. The processing of large-scale genomic data needs many computational resources, resulting highly time-consuming. Therefore, the need of higher computing capabilities has translated into the increasing use of this technology. In this paper, we present a method for modelling and executing customized pipelines in serverless computing. We applied this one to the transcript-level expression analysis of samples from RNA sequencing (RNA-seq), by focusing on the most computationally expensive step: the mapping of reads to a reference genome. Our method has been implemented as an Amazon Web Services (AWS) Lambda function, that is deployed within our own serverless architecture. The parallel instances invoked in AWS Lambda are with negligible latencies, being managed by the provider, therefore, the average computational time are similar among experiments on similar samples. We denoted a relevant advantage in running time, by measuring an improvement up to 79.84% and 90.10% on the concurrent analysis of 10 samples, compared to the local environments having the following specifications: CPU 3.8 GHz 8 vcores and CPU 3.8 GHz 16 vcores, respectively. Pietro Cinaglia, Mario Cannataro |
BIBM | 2 |
| 2023 | Artificial Intelligence to Analyze Huge Amounts of Juridical Documents via Edge Computing
Gessica Fulciniti, Abdellah Kabli, Mario Cannataro, Gaetano Guardasole, Giuseppe Agapito |
EWSN | 3 |
| 2023 | Using Edge-based Deep Learning Model for Early Detection of CancerabstractCancer is one of the most frequent causes of death in the world. Usually, cancer can be easily diagnosed if characteristic symptoms occur. However, many people who are suffering from cancer have no symptoms. Early diagnosis of tumors is essential to contrast their progression, helping to define more effective treatments to provide long-term survival. Early cancer detection is effective if sensible data can be investigated through high-performance technologies like edge computing. Edge computing is a new paradigm for analyzing data as close to the source as possible, avoiding exporting them outside. Hence, edge-based deep learning models can be applied to improve early cancer detection. This paper provides an use case of a classification task on tumor-related data based on the famous UCI machine learning data sets repository using a deep learning approach based on edge computing. In addition, the manuscript provides an overview of the edge computing paradigm, highlighting its advantages and usability. We also described a small experiment with real tumor data to characterize performance considerations. Moreover, the presented model can be used with different data types, such as images, EGC, and ECC signals. Luca Barillaro, Giuseppe Agapito, Mario Cannataro |
PDP | 3 |
| 2023 | High performance deep learning libraries for biomedical applicationsabstractDeep learning approaches are a topic of growing interest since they can achieve high precision in machine learning tasks and may be useful in several scenarios, while high performance computing (HPC) is one of the driving factors for deep learning applications since they require massive computational power. One of these scenarios is related to biomedical context since the massive growth of data generated by several medical procedures. Deep learning techniques, applied on these data may be useful both for medical procedures and for further knowledge discovery in specific field (in example gene interaction related to some diseases). Therefore the importance to have a deep learning library tailored for these task is evident. This paper aims to discuss about some libraries specifically designed to provide convenient high performance computing oriented deep learning support to biomedical applications. We describe two libraries developed inside a European project, named the Deep Health Project, to support both deep learning basic operations and computer vision tasks, oriented to a distributed computing fashion and with some special features for managing biomedical data. In addition we highlight some differences and comparisons with popular environments like Keras and Tensorflow by describing a simple use case. Luca Barillaro, Giuseppe Agapito, Mario Cannataro |
PDP | 3 |
| 2023 | Distributed ICT solutions for scoliosis managementabstractScoliosis is a abnormal curvature of the spine often found in adolescents. Commonly the management of patients with scoliosis is done through manual methods. The use of smartphone applications with integrated sensors allows the remote distributed management of scoliosis. Patients and doctors can communicate and collaborate each other through two different methods: web-based methods and app-based methods. Scoliosis management moves from a centralized system to a decentralized system with obvious benefits for both the doctor and the patient. In particular, the applications allow the doctor to monitor the patients scoliosis curve remotely, saving time and money. Furthermore, they allow the patient to easily check the progress of the scoliotic curvature at home and receive immediate feedback about the correctness and effectiveness of the physical exercises prescribed. We report a brief survey of main apps for scoliosis management and depicts a possible distributed software architecture for scoliosis management. Lorella Bottino, Marzia Settino, Mario Cannataro |
PDP | 3 |
| 2023 | Network models in bioinformatics: modeling and analysis for complex diseasesabstractNetworks are present in different aspects of our life: communication networks, World Wide Web, Social Networks, and can be used to conveniently describe biological and clinical data, such as the interactions of proteins in an organism or the connections of neurons in the brain. Therefore, network science, focusing on the network representations of physical, biological and social phenomena and leading to predictive models of these phenomena, currently represents a vast field of application and research for many scientific and social disciplines. The mathematical background for the study and analysis of networks has its roots in the theory of graphs that allows studying real phenomena in a quantitative way. According to the formalism coming from graph theory, nodes of the graph represent entities, whereas edges represent the associations among them. Currently, in bioinformatics and systems biology, there is a growing interest in analyzing associations among biological molecules at a network level. Since the study of associations in a system-level scale has shown great potential, the use of networks has become the de facto standard for representing such associations, and its application fields span from molecular biology to brain connectome analysis [1]. Molecules of different types, e.g. genes, proteins, ribonucleic acids and metabolites, have fundamental roles in the mechanisms of the cellular processes. The study of their structure and interactions is crucial for different reasons, comprising the development of new drugs and the discovery of disease pathways. Thus, the modeling of the complete set of interactions and associations among biological molecules as a graph is convenient for a variety of reasons. Networks provide a simple and intuitive representation of heterogeneous and complex biological processes. Moreover, they facilitate modeling and understanding of complicated molecular mechanisms combining graph theory, machine learning and deep learning techniques. While proteomics and genomics data, represented as data streams or data tables, are mainly used to screen large populations in case–control studies (e.g. for early detection of diseases), interactomics data are represented as graphs and they add a new dimension of analysis, allowing, for instance, the graph-based comparison of organism’s properties. In general, complex biological systems represented as networks provide an integrated way to look into the dynamic behavior of the cellular system through the interactions of components. For instance, biological networks also referred to as Protein–Protein Interaction Networks, model biochemical interactions among proteins. Nodes represent the proteins from a given organism, and the edges represent the protein–protein interactions [2]. Also, gene regulatory network (GRN) is a collection of genes in a cell, which interact each other and with other substances in the cell, such as proteins or metabolites, thereby governing the rates at which genes in the network are transcribed into mRNA. Similarly, the graph-based modeling of the whole system of the brain elements and their relations, so-called brain connectome, is based on the representation of the regions of interest as nodes, and the representation of functional or anatomical connections as edges [3]. Furthermore, recent discoveries in biology have elucidated that the interplay of molecules of different types (e.g. genes, proteins and ribonucleic acids) is a constitutive block of mechanisms inside cells. Consequently, models describing the interplay should be able to consider the presence of multiple different agents and associations, i.e. multiple different types of nodes and edges, that yield to the so-called heterogeneous networks [4]. Networks and network analysis methods are a keystone in computational biology and bioinformatics and are increasingly used to study biological and clinical data in an integrated way. Network analysis consists of a collection of techniques with a shared methodological perspective, which allows to depict relations among entities and to analyze the structures that emerge from the recurrence of these relations. The basic assumption is that better explanations of different phenomena are yielded by the analysis of the relations among entities. Network analysis can be performed on networks built starting from omics data with the goal of extracting topological properties of the graph. These properties are then used to infer knowledge. For example, in interactomics, the identification of small subgraphs that are statistically overrepresented can be used to identify functionally relevant modules. Similarly, network analysis is used to highlight highly connected regions, assuming they can encode protein complexes. In addition, the systematic study of complex interactions among molecular components (i.e. DNA, RNA, microRNA, proteins and small molecules) is a new paradigm for discovering functional pathways on a global scale [5]. Finally, the network analysis can be conducted on clinical and biomedical data for investigating diseases. This Special Issue aims to collect relevant scientific contributions on fundamental network analysis methods and their applications in computational biology, bioinformatics and medicine. In particular, the Special Issue comprises contributions focusing both on networks modeling and analysis of omics data, as well as on modeling and analysis of clinical data. In Detection of pan-cancer surface protein biomarkers via a network-based approach on transcriptomics data, Daniele Mercatelli et al. [6] present a new network-based analysis protocol for transcriptomic data, in particular, focusing on the overall activity of curated surface proteins, with the final aim to identify those proteins driving major phenotype changes at a network level. The authors have demonstrated that their protocol is able to extract relevant knowledge within and across cancer data sets, by allowing to identify cancer-wide markers to design targeted therapies and biomarker-based diagnostic approaches. In Pathway integration and annotation: building a puzzle with non-matching pieces and no reference picture, Agapito et al. [7] present a review on the current methods for pathway consolidation. The authors start from the consideration that the absence of gold standards for pathway definition and representation as networks has led to the lack of overlap across databases and the lack of data integration across pathway databases. So the authors tackle the strengths and pitfalls of the current pathway consolidation methodologies and highlight directions for future improvements to this research area. In A generic parallel framework for inferring large-scale gene regulatory networks from expression profiles: application to Alzheimer’s disease network, Sebastian, et al. [8] present a framework that incorporates state-of-the-art methods as a black box, to infer GRN from expression profiles. The authors present a case study on the application of the framework to infer an Alzheimer’s disease-affected network from large expression profiles. On the other hand, the last article of the Special Issue focuses on networks modeling and analysis of clinical data. In Intersection of network medicine and machine learning towards investigating the key biomarkers and pathways underlying amyotrophic lateral sclerosis: a systematic review, Das et al. [9] present a review of the main network medicine approaches and implementations of network-based machine learning algorithms in amyotrophic lateral sclerosis, with the aim to identify critical pathways and biomarkers and therapeutic targets for personalized treatment. In summary, although investigating diseases through biological and biomedical data is continuously evolving, we hope this Special Issue will represent an authoritative and valuable resource for researchers. The editors are grateful to both the editor-in-chief and the publisher for having sustained this project, for their timely help and for having supported them in the day-to-day needs. A special thank is addressed to all the authors and reviewers, whose competence and effort allowed the realization of this Special Issue. Marianna Milano is an assistant professor and a senior research scientist in the field of omics data and biological networks analysis at the University ‘Magna Græcia’ of Catanzaro, Italy. Her research interests are focused on: the development of innovative algorithms for the analysis of clinical and omics data through the application of biological knowledge formalized in ontologies; the extraction of knowledge from biological and biomedical data; the use of formal knowledge representation tools in the field of computational biology; the development of algorithms for the analysis of biological and biomedical networks through the application of graph theory. She published one book and more than 60 papers in international journals and conference proceedings. She is a member of Bioinformatics Italian Society (BITS). Mario Cannataro is a full professor of computer engineering and the director of the Data Analytics research center at the University ‘Magna Græcia’ of Catanzaro, Italy. His current research interests include bioinformatics, health informatics, artificial intelligence, data mining, parallel computing. He has published six books and more than 300 papers in international journals and conference proceedings. Mario Cannataro is editor-in-chief of the Encyclopedia of Bioinformatics and Computational Biology, 2nd edn and associate editor of the Briefings in Bioinformatics and IEEE/ACM Transactions on Computational Biology and Bioinformatics journals. He is a senior member of ACM, ACM SIGBio, IEEE, IEEE Computer Society, Bioinformatics Italian Society (BITS) and Italian Society of Biomedical Informatics (SIBIM). He regularly co-organizes international workshops on bioinformatics and high-performance computing in primary conferences such as ACM-BCB, IEEE-BIBM and ICCS. Marianna Milano, Mario Cannataro |
Briefings Bioinform. | 2 |
| 2023 | Multilayer network alignment based on topological assessment via embeddingsabstractBACKGROUND: Network graphs allow modelling the real world objects in terms of interactions. In a multilayer network, the interactions are distributed over layers (i.e., intralayer and interlayer edges). Network alignment (NA) is a methodology that allows mapping nodes between two or multiple given networks, by preserving topologically similar regions. For instance, NA can be applied to transfer knowledge from one biological species to another. In this paper, we present DANTEml, a software tool for the Pairwise Global NA (PGNA) of multilayer networks, based on topological assessment. It builds its own similarity matrix by processing the node embeddings computed from two multilayer networks of interest, to evaluate their topological similarities. The proposed solution can be used via a user-friendly command line interface, also having a built-in guided mode (step-by-step) for defining input parameters. RESULTS: We investigated the performance of DANTEml based on (i) performance evaluation on synthetic multilayer networks, (ii) statistical assessment of the resulting alignments, and (iii) alignment of real multilayer networks. DANTEml over performed a method that does not consider the distribution of nodes and edges over multiple layers by 1193.62%, and a method for temporal NA by 25.88%; we also performed the statistical assessment, which corroborates the significance of its own node mappings. In addition, we tested the proposed solution by using a real multilayer network in presence of several levels of noise, in accordance with the same outcome pursued for the NA on our dataset of synthetic networks. In this case, the improvement is even more evident: +4008.75% and +111.72%, compared to a method that does not consider the distribution of nodes and edges over multiple layers and a method for temporal NA, respectively. CONCLUSIONS: DANTEml is a software tool for the PGNA of multilayer networks based on topological assessment, that is able to provide effective alignments both on synthetic and real multi layer networks, of which node mappings can be validated statistically. Our experimentation reported a high degree of reliability and effectiveness for the proposed solution. Pietro Cinaglia, Marianna Milano, Mario Cannataro |
BMC Bioinform. | 3 |
| 2023 | Advances and challenges in Bioinformatics and Biomedical Engineering: IWBBIO 2020abstractThis Supplement issue, presents five research articles which are distributed, mainly due to the subject they address, from the 8th International Work-Conference on Bioinformatics and Biomedical Engineering (IWBBIO 2020), which was held on line, during September, 30th-2nd October, 2020. These contributions have been chosen because of their quality and the importance of their findings. Those contributions were then invited to participate in this supplement for the following journals of BMC: BMC Bioinformatics and BMC Genomics. In the present Editorial in BMC journal, we summarize the contributions that provide a clear overview of the thematic areas covered by the IWBBIO conference, ranging from theoretical/review aspects to real-world applications of bioinformatic and biomedical engineering. Olga Valenzuela, Mario Cannataro, Irena Rusur, Jianxin Wang 0001, Zhongming Zhao, Ignacio Rojas |
BMC Bioinform. | 2 |
| 2023 | An overview of bioinformatics courses delivered at the academic level in Italy: Reflections and recommendations from BITSabstractIn Italian universities, bioinformatics courses are increasingly being incorporated into different study paths. However, the content of bioinformatics courses is usually selected by the professor teaching the course, in the absence of national guidelines that identify the minimum indispensable knowledge in bioinformatics that undergraduate students from different scientific fields should achieve. The Training&Teaching group of the Bioinformatics Italian Society (BITS) proposed to university professors a survey aimed at portraying the current situation of bioinformatics courses within undergraduate curricula in Italy (i.e., bioinformatics courses activated within both bachelor's and master's degrees). Furthermore, the Training&Teaching group took a cue from the survey outcomes to develop recommendations for the design and the inclusion of bioinformatics courses in academic curricula. Here, we present the outcomes of the survey, as well as the BITS recommendations, with the hope that they may support BITS members in identifying learning outcomes and selecting content for their bioinformatics courses. As we share our effort with the broader international community involved in teaching bioinformatics at academic level, we seek feedback and thoughts on our proposal and hope to start a fruitful debate on the topic, including how to better fulfill the real bioinformatics knowledge needs of the research and the labor market at both the national and international level. Roberto Marangoni, Vitoantonio Bevilacqua, Mario Cannataro, Bruno Hay Mele, Giancarlo Mauri, Anna Marabotti |
PLoS Comput. Biol. | 3 |
| 2022 | Edge-based Deep Learning in Medicine: Classification of ECG signalsabstractThis paper aims to discuss the use of edge computing in medicine with a focus on the analysis of ECG biosignals. Edge computing is a novel paradigm which aims to perform computations (or at least most of them) near the data source achieving some advantages over classical centralized or distributed approaches (e.g. on cloud). After introducing edge computing and the novel NVIDIA Jetson device, the paper presents a use case regarding the classification of ECG biosignals on such a device. Several experiments were conducted to show main differences between traditional and edge-based data analysis approaches. Performance evaluation showed little differences between a traditional approach given the power constrained scenario of edge device. Main results of the paper include an overview of the edge computing paradigm, and a first performance evaluation of deep learning applications on a NVIDIA Jetson device. Luca Barillaro, Giuseppe Agapito, Mario Cannataro |
BIBM | 3 |
| 2022 | Alignment of Dynamic Networks based on Temporal EmbeddingsabstractIn the real-world systems, the interactions between objects (e.g., molecules) are generally represented through the dynamic networks, based on a graph-model that evolve over the time. Network alignment allows evaluating different networks in terms of homology and topology. It is a method to map nodes from different networks, in order to match the same entities. For instance, we can assume that the topological similarity between the regions of two given networks corresponds to their functionality, e.g., in terms of biological process. Therefore, the alignment between the biological networks of two species may be useful to transfer the knowledge from the simplest (e.g., mouse) to the more complex (e.g., human).In this paper, we present DANTE (DynAmic Networks alignment based on Temporal Embeddings), a solution for the alignment of dynamic networks, based on temporal embeddings. DANTE takes into account the similarity between the nodes based on their evolution and relationships between one time point and the subsequent, for all ones in the networks. Our temporal embeddings concern a set of vectors representing the topological similarity between the nodes of two given networks, by representing each node as a word on the Skip-Gram model. In addition, we extend the embeddings process to the dynamic networks, in order to evaluate the similarity between two dynamic networks. DANTE applies an iterative process for maximizing globally the match score between the pair of nodes.DANTE is freely available on https://github.com/pietrocinaglia/dante. Pietro Cinaglia, Mario Cannataro |
BIBM | 2 |
| 2022 | Serverless computing for RNA-Seq data analysisabstractServerless is a computing model where the infrastructure orchestration is managed by the service provider. Amazon Web Services (AWS) offers a serverless computing service named Lambda (or AWS Lambda). It allows deploying a function on AWS by executing this one on a serverless environment. The analysis of RNA-Seq data needs a large amount of computational resources, and it is really time-consuming. Therefore, the solutions based on serverless computing could offer a tangible benefit, being well-prepared to the parallelization. In this paper, we proposed a serverless solution for RNA-Seq data analysis, by focusing the attention on the mapping/alignment of sequencing reads to a reference genome, in that it is the more computationally expensive step within an RNA-Seq pipeline. We deployed our solution on an in-house serverless environment, useful to overcome the limits of a traditional AWS Lambda configuration. The tests have been conducted on 16 Paired-end sequencing samples, by using the Ensembl’s GRCh38 (release 84, Homo sapiens) as reference genome. Furthermore, we also estimated the probable time needed to map 50, 100, 250, 500, and 1000 samples based on our 16 ones. The proposed solution has been compared with local analysis performed by using two different workstations. We demonstrate an obvious advantage for our use-case. Results showed that a serverless solution is only suitable for highly parallel computations or for a lot of number of independent atomic operations, while it is not recommended for a single atomic operation that requires a lot amount of CPU and/or memory. Pietro Cinaglia, José Luis Vázquez-Poletti, Mario Cannataro |
BIBM | 3 |
| 2022 | A parallel software pipeline to select relevant genes for pathway enrichmentabstractThe continuous technological development of experimental omics technologies such as microarrays, allows to perform large scale genomics studies. After the initial enthusiasm, it became pretty clear that even the results provided by microarrays in form of lists of differential expressed genes (DEGs), were mainly as enigmatic as the first sequence of the genome, because these lists of DEGs are detached from the influenced biological mechanisms. Pathway enrichment analysis (PEA) supports researchers to provide the clues necessary to link DEGs to the influenced biological pathways and consequently to the underlying biological mechanisms and processes. Putting DEGs data sets in a suitable format for the PEA can be a tedious error-prone and laborious process even for bioinformaticians, who needs to perform it manually before to be ready for the PEA. To fill this lack, we present a parallel software pipeline which uploads a list of DEGs and automatically provides as results the enriched pathways.The parallel software pipeline is implemented in Python and provides the following automated actions: i) parallel splitting of DEGs in groups; ii) parallel building of the similarity matrices related to the DEGs groups; iii) parallel mapping of similarity matrices in networks; iv) parallel pathway enrichment analysis for each group of identified DEGs.Preliminary results shown that the pipeline can help to analyze DEGs and easily generate in a few minutes a list of pathway enrichment results that otherwise would require numerous hours of manual work and several different scripts.The parallel software pipeline provides a two-fold benefits: first, it contributes to speed up the computation of pathway enrichment, automating several steps currently performed manually. Second, it provides a more peculiar list of DEGs to calculate pathway enrichment, contributing to improve the relevance and significance of the enriched pathways. Giuseppe Agapito, Mario Cannataro |
PDP | 2 |
| 2022 | A statistical network pre-processing method to improve relevance and significance of gene lists in microarray gene expression studiesabstractBACKGROUND: Microarrays can perform large scale studies of differential expressed gene (DEGs) and even single nucleotide polymorphisms (SNPs), thereby screening thousands of genes for single experiment simultaneously. However, DEGs and SNPs are still just as enigmatic as the first sequence of the genome. Because they are independent from the affected biological context. Pathway enrichment analysis (PEA) can overcome this obstacle by linking both DEGs and SNPs to the affected biological pathways and consequently to the underlying biological functions and processes. RESULTS: To improve the enrichment analysis results, we present a new statistical network pre-processing method by mapping DEGs and SNPs on a biological network that can improve the relevance and significance of the DEGs or SNPs of interest to incorporate pathway topology information into the PEA. The proposed methodology improves the statistical significance of the PEA analysis in terms of computed p value for each enriched pathways and limit the number of enriched pathways. This helps reduce the number of relevant biological pathways with respect to a non-specific list of genes. CONCLUSION: The proposed method provides two-fold enhancements. Network analysis reveals fewer DEGs, by selecting only relevant DEGs and the detected DEGs improve the enriched pathways' statistical significance, rather than simply using a general list of genes. Giuseppe Agapito, Marianna Milano, Mario Cannataro |
BMC Bioinform. | 3 |
| 2021 | Characterization of Long COVID using text mining on narrative medicine textsabstractThrough an adequate survey of the history of the disease, Narrative Medicine (NM) aims to allow the definition and implementation of an effective, appropriate, and shared treatment path. In the present study, standard text mining (TM) techniques are applied, as a Latent Dirichlet Allocation (LDA) model for topic modeling is used to characterize narrative medicine texts written on COVID-19. In particular, the focus was mainly on the writings of patients with Post-acute Sequelae of COVID-19, i.e., PASC, as opposed to writings by health professionals and general reflections on COVID-19. The results suggest that the testimonies of PASC patients can be used for identifying shared issues to focus on to be followed and supported appropriately, even from a psychological point of view. Ileana Scarpino, Chiara Zucco, Mario Cannataro |
BIBM | 3 |
| 2021 | Bioinformatics helping to mitigate the impact of COVID-19 - EditorialabstractThe coronavirus disease 2019 (COVID-19) pandemic caused by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) became known to the world at the end of 2019 [1]. The severity of the pandemic and its worldwide spread provoked an unprecedented effort of the scientific community and a lot of new research was conducted, especially by the medicine, biology, public health, bioinformatics and computer science researchers, that led to the rapid development of several novel vaccines [2]. At the biological level, SARS-CoV-2 and COVID-19 research involves several themes, including high-throughput technologies such as Next-Generation Sequencing for detecting the genome of SARS-CoV-2, databases storing SARS-CoV-2 genomes and variants, bioinformatics software tools and databases for analyzing and storing host–virus interactions [3]. At the medical level and in particular when considering the search for therapeutic strategies, the identification of COVID-19 biomarkers, the discovery of therapeutic targets for drugs and the bioinformatics approaches for drug repurposing, i.e. the use of already available drugs for the COVID-19 disease, are main research themes. At the epidemiological and public-health level, main research themes regard: the systematic collection and sharing of data about the spread of the infection, such as the number of cases, hospitalized, ICU and deceased patients, that may be helpful to manage the pandemic [4]; the biological tests for testing, and the computational methods for tracing and tracking infected people; the exploitation of the vast clinical data stored into the Electronic Health Records of COVID-19 patients [5]; the analysis of the impact of lockdown measures in various contexts, e.g. at socioeconomic level, that may benefit from sentiment analysis methods; and finally measures to help quarantined people, such as local healthcare service, robotics and virtual assistants. Finally, those unprecedented research efforts yield an overwhelming volume of scientific publications that require new methods and tools to improve learning from SARS-CoV-2 and COVID-19 literature, such as novel text mining and natural language processing techniques to distill relevant information [6]. This Special Issue aims to collect relevant scientific contributions on methods and applications of bioinformatics and informatics in themes related to COVID-19 and SARS-CoV-2. In particular, the special issue is organized in two main strands: one on Bioinformatics helping to mitigate the impact of Covid-19 and another one on Informatics helping to mitigate the impact of Covid-19. Here, we present the first-strand Bioinformatics helping to mitigate the impact of Covid-19 that comprises more than 60 manuscripts, each dealing with one of the following central key issues, as detailed below. Next-generation sequencing is the central technology for detecting genomes of SARS-CoV-2 that provides the basic data about the virus. Bioinformatics pipelines, biological and host–virus interaction databases, are key tools for computing such data and advancing knowledge on SARS-CoV-2. In Next-generation sequencing of SARS-CoV-2 genomes: challenges, applications and opportunities, Chiara, D’Erchia, Gissi, Manzari, Parisi, Resta, Zambelli, Picardi, Pavesi, Horner and Pesole discuss next-generation sequencing (NGS), a fundamental technology and method for tracing origins and understanding the evolution of infectious agents, and in particular to reconstruct the genomic sequence of SARS-CoV-2. Authors briefly introduce available platforms and approaches for the sequencing of SARS-CoV-2 genomes and outline current databases for SARS-CoV-2 genomic data. As a result, they provide some useful guidelines for the sharing and deposition of SARS-CoV-2 data and metadata, suggesting the use of efficient and standardized approaches for the production, handling and integration of SARS-CoV-2 sequencing data. In Bioinformatics resources for SARS-CoV-2 discovery and surveillance, Hu, J. Li, Zhou, C. Li, Holmes and Shi discuss the role of next-generation sequencing and available bioinformatics pipelines for the worldwide genomic surveillance of SARS-CoV-2, focusing on the tracking of COVID-19 spread and the analysis of evolution and patterns of SARS-CoV-2 variation on a global scale. The authors review the main bioinformatics resources available for the discovery and surveillance of SARS-CoV-2 and discuss their advantages and disadvantages, highlighting areas needing urgent technical improvements. In Computational strategies to combat COVID-19: useful tools to accelerate SARS-CoV-2 and coronavirus research, Franziska Hufsky et al. present bioinformatics tools that have been explicitly developed for SARS-CoV-2 with the aim to provide key tools for the detection, understanding and treatment of COVID-19. The reviewed tools include detection of SARS-CoV-2, analysis of sequencing data, tracking and containment of the COVID-19 pandemic, study of coronavirus evolution, discovery of potential drug targets and related therapeutic strategies. All analyzed tools are available online and free to use and for each tool the authors describe a use case and discuss the contribution to the SARS-CoV-2 research. In A review on viral data sources and search systems for perspective mitigation of COVID-19, Bernasconi, Canakoglu, Masseroli, Pinoli and Ceri discuss the data integration activities needed for accessing and searching SARS-CoV-2 genome sequences and metadata stored in main viral sequences databases. The authors review some host-pathogen integrated datasets and underline possible integrative surveillance mechanisms, e.g. based on the time-space distribution of common virus variants. They observe that while organizations already managing virus databases are offering novel specific SARS-CoV-2 data and services, novel specific approaches and resources to face COVID-19 are appearing, providing better accessibility of viral sequence data, integration with clinical data and with the genotype of the human host. The role of pathway enrichment analysis (PEA) in finding possible targets present in biological pathways of host cells that are targeted by SARS-CoV-2 is discussed in Comprehensive pathway enrichment analysis workflows: COVID-19 case study. To guide bioinformaticians in the choice of the many available PEA methods and software tools, Agapito, Pastrello and Jurisica highlight how to choose the most suitable PEA methods based on the type of SARS-CoV-2/COVID-19 data to analyze. In Web tools to fight pandemics: the COVID-19 experience, Mercatelli, Holding and Giorgi focus on the state of the art of COVID-19 online resources and review the most popular web tools for the analysis of COVID-19 data, focusing on the epidemiology, genomics, interactomics and pharmacology fields. The identification of COVID-19 biomarkers, the discovery of therapeutic targets for drugs and the bioinformatics approaches for drug repurposing are key research topics to address for facing the COVID-19 disease. The research in these fields is mainly driven by the SARS-CoV-2 proteins structure, protein dynamics produced by computer simulations, variants and mutations of the virus. In A review of COVID-19 biomarkers and drug targets: resources and tools, Caruso, Scala, Cerulo and Ceccarelli present a review of tools and resources to identify biomarkers and drug targets in COVID-19, through the automatic analysis of a consolidated corpus of 27 570 papers. Using latent Dirichlet analysis, authors extracted topics associated with computational methods for biomarker identification and drug repurposing, which include machine learning and artificial intelligence for disease characterization, vaccine development and therapeutic target identification. In Bioinformatics resources facilitate understanding and harnessing clinical research of SARS-CoV-2, Ahsan, Liu, Feng, Zhou, Ma, Bai and Chen review some bioinformatics resources, the status of drug development and various resources for enabling research toward effective treatment of COVID-19, including phylogenetic characteristics, genomic conservation and interaction data. The authors review several SARS-CoV-2-related tools and databases, focusing on bioinformatics approaches for target prioritization and drug repurposing. They present a web-portal named OverCOVID that provides a detailed description of SARS-CoV-2 basics and shares a collection of bioinformatics resources and information that may contribute to better understanding of SARS-CoV-2 and to therapeutic advances. In A review on drug repurposing applicable to COVID-19, Dotolo, Marabotti, Facchiano and Tagliaferri present a review of different drug repurposing strategies useful to face COVID-19 pandemic, i.e. strategies for discovering new applications of existing drugs to COVID-19, that may reduce costs and provide shorter time application. Authors categorize computational drug repurposing approaches into network, structure and artificial intelligence approaches. Network-based approaches, further categorized into clustering and propagation approaches, allow the identification of proteins that are functionally associated with COVID-19, evidencing novel drug–disease or drug–target relationships useful for new therapies. Structure-based approaches study how chemical compounds can interact with the macro molecular targets, finding new possible applications for existing drugs. Finally, artificial intelligence approaches are evaluated less relevant at the moment, due to the scarcity of data to learn models. In The impact of structural bioinformatics tools and resources on SARS-CoV-2 research and therapeutic strategies, Waman, Sen, Varadi, Daina, Wodak, Zoete, Velankar and Orengo review recent structural bioinformatics tools and discuss the impact of structure-based studies on SARS-CoV-2 research, with focus on the differences between SARS-CoV-2 and SARS-CoV, the SARS-CoV-2 residues involved in receptor–antibody recognition, the variants in host proteins that affect susceptibility to infection, and the computational analyses enabling structure-based drug and vaccine development. In SARS-CoV-2 3D database: understanding the coronavirus proteome and evaluating possible drug targets, Alsulami, Thomas, Jamasb, Beaudoin, Moghul, Bannerman, Copoiu, Vedithi, Torres and Blundell propose a new database containing 3D models of the SARS-CoV-2 proteome, including models of protomers and oligomers, protein-ligand docking, interactions of SARS-CoV-2 proteins with human proteins, impacts of mutations and experimental structures. The resulting SARS-CoV-2 3D database provides information for drug discovery, useful to evaluate targets and design new possible therapeutics. The unprecedented rate of SARS-CoV-2 and COVID-19 publications strongly accelerated the development of text mining and natural language processing techniques to analyze scientific literature. In Text mining approaches for dealing with the rapidly expanding literature on COVID-19, Wang and Lo tackle the problem of extracting recent knowledge from the overwhelming COVID-19 literature using text mining applications and discuss the corpora, models and systems that have been introduced for COVID-19. They analyzed 39 systems that support search, discovery, visualization and summarization of the COVID-19 literature, and categorized them through qualitative description, performance assessment and user interface. The authors note that some systems, in addition to standard functions such as search and discovery, provide new functions such as summary of multiple documents or connections between scientific articles and clinical trials. In How do we share data in COVID-19 research? A systematic review of COVID-19 datasets in PubMed Central Articles, Zuo, Chen, Ohno-Machado and Xu review more than 100 datasets about COVID-19 that were reported into several scientific articles available from PubMed Central. Starting from 12 324 COVID-19 full-text articles published until 31 May 2020, the authors extracted the links to 128 datasets that were manually reviewed using 10 variables. Although the analysis was performed in an initial stage of the pandemic, the authors found 128 unique dataset links. The largest portion (53.9%) are epidemiological datasets and most datasets (84.4%) were available for immediate download. The study found that GitHub was the most used repository and evidenced a great heterogeneity in the way the datasets are mentioned, shared, and updated. In addition to bioinformatics research, COVID-19 pandemic has given an impulse to the development and adaption of several informatics techniques, including computational methods for tracing and tracking infected people; collaborative data infrastructures for COVID-19 research; sentiment analysis methods for monitoring the impact of lockdown measures; artificial intelligence methods and robotics applications to support remote patients assistance (e.g. quarantined people). In Health informatics and EHR to support clinical research in the COVID-19 pandemic: an overview, Dagliati, Malovini, Tibollo and Bellazzi discuss the role of Electronic Health Records (EHR) that are primarily used to support day-by-day clinical activities, to enable global scale research on COVID-19. Authors review collaborative data infrastructures to support COVID-19 research, including studies on effectiveness of drugs and therapeutic strategies, and discuss the data sharing and governance issues emerged with the COVID-19 pandemic, that may prevent a full exploitation of EHR data, especially when considering international collaborations. The authors underline the data management, interoperability and governance issues, the modelling of healthcare processes and the management of data privacy regulations as primary aspects to boost collaborative research. In Robots as intelligent assistants to face COVID-19 pandemic, Seidita, Lanza, Pipitone and Chella discuss the role that an emerging technology such as robotics may have in the management and fight against the COVID-19 pandemic. Authors analyzed scientific articles and industrial initiatives underlining how robotics was used to face the pandemic, its level of readiness, what are the expectations from robots and what remains to do. Authors reviewed what is offered by research groups in terms of robot support for therapies and for prevention actions and discussed the maturity of robotics in dealing with situations like COVID-19. In HVIDB: a comprehensive database for human-virus protein-protein interaction, X. Yang, Lian, Fu, Wuchty, S. Yang and Zhang present HVIDB (Human-Virus Interaction DataBase), an annotated human-virus protein–protein interaction (PPI) database that contains experimentally verified human-virus PPIs about 35 virus families, experimentally verified 3D complex structures of human-virus PPIs, and integrates machine learning models to predict interactions between human host and viral proteins. Although research on SARS-CoV-2 and COVID-19 is continuously evolving, we hope this special issue will represent an authoritative and valuable resource for researchers. The Editors are grateful to both the Editor-in-Chief and the Publisher for having sustained this project, for their timely help, and for having supported them in the day-to-day needs. A special thank is addressed to all the authors and reviewers, whose competence and effort allowed the realization of this special issue. Mario Cannataro is a Full Professor of computer engineering at the University ‘Magna Graecia’ of Catanzaro, Italy, and the Director of the Data Analytics Research Center. His current research interests include bioinformatics, health informatics, artificial intelligence, data mining and parallel computing. He published three books and more than 300 papers in international journals and conference proceedings. Mario Cannataro is a Member of the Board of Directors of ACM SIGBio, and a Senior Member of ACM, IEEE, BITS (Bioinformatics Italian Society) and SIBIM (Italian Society of Biomedical Informatics). Andrew Harrison is a Senior Lecturer in Data Science at the University of Essex, UK. His current research interests include the analysis of Hi-C experiments as well as large-scale meta-analysis of gene expression experiments. He has published over 50 papers on a broad range of topics from data science, applied statistics, bioinformatics, analysis of high-throughput biological experiments and astrophysics. Mario Cannataro |
Briefings Bioinform. | 1 |
| 2021 | MMRFBiolinks: an R-package for integrating and analyzing MMRF-CoMMpass dataabstractIn order to understand the mechanisms underlying the onset and the drug responses in multiple myeloma (MM), the second most frequent hematological cancer, the use of appropriate bioinformatic tools for integrative analysis of publicly available genomic data is required. We present MMRFBiolinks, a new R package for integrating and analyzing datasets from the Multiple Myeloma Research Foundation (MMRF) CoMMpass (Clinical Outcomes in MM to Personal Assessment of Genetic Profile) study, available at MMRF Researcher Gateway (MMRF-RG), and from the National Cancer Institute Genomic Data Commons (NCI-GDC) Data Portal. The package provides several methods for integrative analysis (array-array intensity correlation, Kaplan-Meier survival analysis) and visualization (response to treatments plot) of MMRF data, for performing an easily comprehensible analysis workflow. MMRFBiolinks extends the TCGABiolinks package by providing 13 new functions to analyze MMRF-CoMMpass data: six dealing with MMRF-RG data and seven with NCI-GDC data. As validation of the tool, we present two cases studies for searching, downloading and analyzing MMRF data. The former presents a workflow for identifying genes involved in survival depending on treatment. The latter presents an analysis workflow for analyzing the Best Overall (BO) response through correlation plots between the BO Response with respect to treatments, time, duration of treatment and annotated variants, as well as through Kaplan-Meier survival curves. The case studies demonstrate how MMRFBiolinks is able of overcoming the limitations of the analysis tools available at NCI-GDC and MMRF-RG, facilitating and making more comprehensive the retrieval, downloading and analysis of MMRF data. Marzia Settino, Mario Cannataro |
Briefings Bioinform. | 2 |
| 2021 | Using BioPAX-Parser (BiP) to enrich lists of genes or proteins with pathway dataabstractBACKGROUND: Pathway enrichment analysis (PEA) is a well-established methodology for interpreting a list of genes and proteins of interest related to a condition under investigation. This paper aims to extend our previous work in which we introduced a preliminary comparative analysis of pathway enrichment analysis tools. We extended the earlier work by providing more case studies, comparing BiP enrichment performance with other well-known PEA software tools. METHODS: PEA uses pathway information to discover connections between a list of genes and proteins as well as biological mechanisms, helping researchers to overcome the problem of explaining biological entity lists of interest disconnected from the biological context. RESULTS: We compared the results of BiP with some existing pathway enrichment analysis tools comprising Centrality-based Pathway Enrichment, pathDIP, and Signaling Pathway Impact Analysis, considering three cancer types (colorectal, endometrial, and thyroid), for a total of six datasets (that is, two datasets per cancer type) obtained from the The Cancer Genome Atlas and Gene Expression Omnibus databases. We measured the similarities between the overlap of the enrichment results obtained using each couple of cancer datasets related to the same cancer. CONCLUSION: As a result, BiP identified some well-known pathways related to the investigated cancer type, validated by the available literature. We also used the Jaccard and meet-min indices to evaluate the stability and the similarity between the enrichment results obtained from each couple of cancer datasets. The obtained results show that BiP provides more stable enrichment results than other tools. Giuseppe Agapito, Mario Cannataro |
BMC Bioinform. | 2 |
| 2021 | Parallel and distributed association rule mining in life science: A novel parallel algorithm to mine genomics data
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro |
Inf. Sci. | 3 |
| 2020 | An efficient and scalable SPARK preprocessing methodology for Genome Wide Association StudiesabstractThe importance of the use of high-performance software frameworks to analyze omics data obtained by using High-Throughput (HT) essays is widely recognized. HT methodologies comprise microarrays, Genome-Wide Association Studies (GWAS), and Next Generation Sequencing (NGS), which provide a vast amount of data per a single experiment. Each HT vendor provides to the users only the software frameworks and the proprietary libraries for the annotation, and summarization of raw data. Consequently, the needs of algorithms for the preprocessing and analysis of omics data arise. GWAS aims to highlight the association between genetic variants and diseases by examining single nucleotide polymorphisms (SNPs), which differ in a statistically significant way between cases and controls. The effectiveness of GWAS analysis increases with the number of analyzed samples per single experiment. GWAS data analyzed through the use of statistical methods can detect associations among a single allelic variant and the clinical conditions of samples. To overcome these limitations, and to make it possible to discover multiple associations among allelic variants, it is possible to use Association Rules mining. Consequently, the need for the introduction of scalable Association Rule Mining (ARM) algorithms able to analyze GWAS data arises. Hence, the use of high-performance data analytics framework is needed. For this purpose, we propose a software framework called GARMS (GWAS Association Rule Mining in Spark) built on top of Apache Spark for the preprocessing, and mining of association rules from GWAS data sets. GARMS comprises a two steps analysis methodology: (i) in the first step, the GWAS data are preprocessed, along with the identification of the frequent itemsets; (ii) in the second step, frequent itemsets are employed to mine association rules without scanning the input data. We implemented our algorithm, and we tested it on some synthetic GWAS data sets. Preliminary results confirm that our method may extract relevant association rules from GWAS data reducing the computational time. Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro |
PDP | 3 |
| 2020 | BioPAX-Parser: parsing and enrichment analysis of BioPAX pathwaysabstractSUMMARY: Biological pathways are fundamental for learning about healthy and disease states. Many existing formats support automatic software analysis of biological pathways, e.g. BioPAX (Biological Pathway Exchange). Although some algorithms are available as web application or stand-alone tools, no general graphical application for the parsing of BioPAX pathway data exists. Also, very few tools can perform pathway enrichment analysis (PEA) using pathway encoded in the BioPAX format. To fill this gap, we introduce BiP (BioPAX-Parser), an automatic and graphical software tool aimed at performing the parsing and accessing of BioPAX pathway data, along with PEA by using information coming from pathways encoded in BioPAX. AVAILABILITY AND IMPLEMENTATION: BiP is freely available for academic and non-profit organizations at https://gitlab.com/giuseppeagapito/bip under the LGPL 2.1, the GNU Lesser General Public License. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Giuseppe Agapito, Chiara Pastrello, Pietro H. Guzzi, Igor Jurisica, Mario Cannataro |
Bioinform. | 5 |
| 2020 | cPEA: a parallel method to perform pathway enrichment analysis using multiple pathways databases
Giuseppe Agapito, Mario Cannataro |
Soft Comput. | 2 |
| 2019 | Association Rule Mining from large datasets of clinical invoices documentabstractThe concept of massive data generation nowadays affects several domains such as marketing including electronic invoices of large retailers, web access log files, healthcare, life sciences and so on. All these web activities introduced a new way to pay through the concept of electronic invoices (eInvoice), replacing the paper invoices. For these reasons, eInvoicing can be thought of as an innovative digital infrastructure for the issue, transmission, and storage of invoices. The availability of large volumes of eInvoices allows the discovery of new knowledge through data mining in these domains. Thus, users by using data mining can extract knowledge from large invoices documents. In this paper, we present a software tool for mining association rules from invoices produced in healthcare centers. In particular, the tool adopt a novel preprocessing methodology that provides merging, cleaning, formatting and summarization of eInvocies. The methodology can improve the quality of a huge amount of clinical invoices reducing the quantity of irrelevant data, making the remaining data suitable to mine information in form of association rules. The core of the tool allows to extract association rules from eInvoices; as a case study, we discuss the mined rules, highlighting the relationships among the purchased goods. Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Sabrina Graziano, Mario Cannataro |
BIBM | 5 |
| 2019 | Pathway Analysis for SNP microarray dataabstractPathway Analysis (PA) is a powerful method for data analysis in genomics, most often applied to gene expression analysis, but little used to analyze variants such as Single Nucleotide Polymorphisms (SNPs). PA could allow the interpretation of variants concerning the biological processes in which the affected genes and proteins are involved. Currently, the available PA software tools are not able to automatically perform pathway analysis using SNPs data. PA software tools cannot deal natively with SNPs data, hence several software tools have to be used to put SNPs data in the proper format for the analysis. To overcome these limitations, we present SNP Microarray Pathway Analysis (MPA), a software tool able to discriminate relevant genes from SNP microarrays to use in PA analysis. MPA automatically identifies relevant SNPs using the well known Fisher's test, with which to perform PA. Pathway analysis in MPA is obtained employing the Hypergeometric function. As a result, MPA provides to the user the list of enriched pathways from the identified SNPs. MPA software tool along with the user guide and datasets, are available for download at https://gitlab.com/giuseppeagapito/mpa under the GPL v3.0 license. Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro |
BIBM | 3 |
| 2019 | Mining Association Rules From Disease OntologyabstractThe Disease Ontology (DO) is standardized, controlled vocabulary that contains information about inherited, developmental and acquired human diseases. Each DO term is associated with disease concepts through an annotation process. The relevance and the specificity of DO terms are often evaluated by its Information Content (IC). An important research area focus on the analysis of annotated data with the goal to extract knowledge. For example, the analysis of annotated data using Association Rules (AR) may supply meaningful knowledge, discovering relevant associations. Classical association rules methods consider all annotation equally, do not taking into account that the DO terms have different Information Content, i.e. different relevance. This implies the generation of association rules with low IC. In this paper we presents WARDO (Weighted Association Rule mining from Disease Ontology), a methodology based on the extraction od Weighted Association Rules from the DO Ontology considering the IC of terms. To assess our methodology, we tested WARDO on DO annotation datasets. WARDO is publicly available at https://gitlab.com/giuseppeagapito/wardo. Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro |
BIBM | 4 |
| 2019 | GLAlign: A Novel Algorithm for Local Network AlignmentabstractNetworks are successfully used as a modelling framework in many application domains. For instance, Protein-Protein Interaction Networks (PPINs) model the set of interactions among proteins in a cell. A critical application of network analysis is the comparison among PPINs of different organisms to reveal similarities among the underlying biological processes. Algorithms for comparing networks (also referred to as network aligners) fall into two main classes: global aligners, which aim to compare two networks on a global scale, and local aligners that aim to evidence single sub-regions of similarity among networks. The possibility to improve the performance of the aligners by mixing the two approaches is a growing research area. In our previous work, we started to explore the possibility to use global alignment to improve the local one. We here explore further this possibility by using topological information extracted from global alignment to guide the steps of the local alignment. Therefore, we present Global Local Aligner (GLAlign), a methodology that improves the performances of local network aligners by exploiting a preliminary global alignment. Furthermore, we provide implementation of GLAlign. As a proof-of-principle, we evaluated the performance of the GLAlign prototype on the PPINs of five species. Results show that GLAlign methodology outperforms the state-of-the-arts local alignment algorithms. GLAlign is publicly available for academic use and can be downloaded here: https://sites.google.com/site/globallocalalignment/. Marianna Milano, Pietro H. Guzzi, Mario Cannataro |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | S4S: RESTful Services to Collect, Integrate and Analyze SNPs and Clinical Data on the Web
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro |
BIBM | 3 |
| 2018 | Survey of main tools for querying and analyzing TCGA Data
Marzia Settino, Mario Cannataro |
BIBM | 2 |
| 2018 | Predicting Abandonment in Telehomecare Programs Using Sentiment Analysis: A System Proposal
Chiara Zucco, Sergio Bella, Clarissa Paglia, Paola Tabarini, Mario Cannataro |
BIBM | 5 |
| 2018 | Explainable Sentiment Analysis with Applications in Medicine
Chiara Zucco, Huizhi Liang 0001, Giuseppe Di Fatta, Mario Cannataro |
BIBM | 4 |
| 2017 | A software pipeline for multiple microarray data analysisabstractMicroarray platforms such as Gene expression, oligonucleotide (GeneChip), and cDNA (complementary DNA) can play a critical role in the understanding of genome sequencing, by providing information for hundreds or thousands genes in a single assay, information not available by using other methodologies of investigation. Microarray experiments aim to survey patterns of gene expression by assaying the expression levels of thousands of genes in parallel in a single assay. Thus, microarray are known as high-throughput technologies, allowing the investigation of genetic variations underlying the inter-individual variability in drug pharmacokinetics/pharmacodynamics, providing a complete understanding of gene function, regulation, and interactions. Thus, to exploit all the power of this massive amount of data in the short possible time (before that data becomes obsolete), the necessity to develop databases and software tools for efficient data collection and analysis arises. The establishing of efficient and scalable software tools avoids that researcher will be overwhelming from this ever-growing flow of data. At this reason, a preliminary design of a platform named microPipe, for the multiple analysis of microarray data sets is proposed. Specifically, the paper outlines the main issues and challenges relative to the design of such a platform. Giuseppe Agapito, Mario Cannataro |
BIBM | 2 |
| 2017 | Sentiment analysis and affective computing for depression monitoringabstractDepression is one of the most common and disabling mental disorders that has a relevant impact on society. Semiautomatic and/or automatic health monitoring systems could be crucial and important to improve depression detection and follow-up. Sentiment Analysis refers to the use of natural language processing and text mining methodologies aiming to identify opinion or sentiment. Affective Computing is the study and development of systems and devices that can recognize, interpret, process, and simulate human affects. Sentiment Analysis and Affective Computing methodologies could provide effective tools and systems for an objective assessment and monitoring of psychological disorders and, in particular, of depression. In this paper, the application of sentiment analysis and affective computing methodologies to depression detection and monitoring are presented and discussed. Moreover, a preliminary design of an integrated multimodal system for depression monitoring, that includes sentiment analysis and affective computing techniques, is proposed. Specifically, the paper outlines the main issues and challenges relative to the design of such a system. Chiara Zucco, Barbara Calabrese, Mario Cannataro |
BIBM | 3 |
| 2017 | Parallel and Cloud-Based Analysis of Omics Data: Modelling and Simulation in MedicineabstractHigh throughput experimental platforms and diagnostic equipments available in clinical settings and in research laboratories, such as magnetic resonance imaging, microarray, mass spectrometry and next-generation sequencing, are producing an increasing volume of clinical and omics data. Moreover, Electronic Patients Records (EPRs), eHealth systems, personal mobile sensors and Social Networks are collecting an overwhelming volume of health and life style data that may be integrated with clinical data and more and more is used for the real-time monitoring of patient's health. This poses new issues in terms of secure data storage, effective models for data integration, efficient algorithms for data analysis, new models for health monitoring, that may be addressed, among the others, using high performance computing solutions. Parallel computing and Cloud Computing may offer efficient and scalable solutions in an orthogonal way. In fact, parallel, bioinformatics software, that exploit off-the-shelf high performance computers, may be used to preprocess and analyze omics data at a lower layer, for instance to highlight genetic variation associated with complex diseases. On the other hand, Cloud Computing offers large scale data storage, data sharing services, on-demand anytime and anywhere access to resources and applications, for the realization of elastic and scalable applications and services. Motivated by the increasing use of parallel computing and cloud computing in life sciences, in this paper we survey both parallel bioinformatics algorithms for the parallel preprocessing and statistical and data mining analysis of omics data, as well as Cloud-based healthcare and biomedicine services and systems for large scale applications. Moreover, the paper underlines main issues and problems related to the use of such platforms for the storage and analysis of health data, with special focus to the security and privacy of patients data, that are particularly important in fields such as personalized medicine. Finally, the paper presents some case studies about the parallel and distributed modelling and simulation in medicine and biology. Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Gionata Fragomeni, Giuseppe Tradigo, Pierangelo Veltri, Mario Cannataro |
PDP | 7 |
| 2017 | An extensive assessment of network alignment algorithms for comparison of brain connectomesabstractBACKGROUND: Recently the study of the complex system of connections in neural systems, i.e. the connectome, has gained a central role in neurosciences. The modeling and analysis of connectomes are therefore a growing area. Here we focus on the representation of connectomes by using graph theory formalisms. Macroscopic human brain connectomes are usually derived from neuroimages; the analyzed brains are co-registered in the image domain and brought to a common anatomical space. An atlas is then applied in order to define anatomically meaningful regions that will serve as the nodes of the network - this process is referred to as parcellation. The atlas-based parcellations present some known limitations in cases of early brain development and abnormal anatomy. Consequently, it has been recently proposed to perform atlas-free random brain parcellation into nodes and align brains in the network space instead of the anatomical image space, as a way to deal with the unknown correspondences of the parcels. Such process requires modeling of the brain using graph theory and the subsequent comparison of the structure of graphs. The latter step may be modeled as a network alignment (NA) problem. RESULTS: In this work, we first define the problem formally, then we test six existing state of the art of network aligners on diffusion MRI-derived brain networks. We compare the performances of algorithms by assessing six topological measures. We also evaluated the robustness of algorithms to alterations of the dataset. CONCLUSION: The results confirm that NA algorithms may be applied in cases of atlas-free parcellation for a fully network-driven comparison of connectomes. The analysis shows MAGNA++ is the best global alignment algorithm. The paper presented a new analysis methodology that uses network alignment for validating atlas-free parcellation brain connectomes. The methodology has been experimented on several brain datasets. Marianna Milano, Pietro H. Guzzi, Olga Tymofiyeva, Duan Xu, Christopher Paul Hess, Pierangelo Veltri, Mario Cannataro |
BMC Bioinform. | 7 |
| 2016 | GLAlign: Using global graph alignment to improve local graph alignmentabstractDuring the last years, the graph alignment has been used as a possible way to compare biological networks in system biology. The techniques for the alignment of biological networks fall into two categories: global alignment, that aims to identify large common subnetworks optimizing a topological alignment quality, and local alignment that aims to evidence single sub-regions optimizing functional alignment quality. In this work, we presented GLAlign (Global Local Aligner), a novel algorithm that integrates global and local alignment, starting from the possibility that the topological information gathered by results of global alignment can be used to improve the local alignment building. Initially, the algorithm enables the calculation of global alignment, then it uses this one to guide the building of the local alignment. GLAlign is based on two global and local algorithms widely used in literature, MAGNA++ and AlignMCL. We tested GLAlign as proof-of-principle using the Protein Interaction Networks (PINs) of three species: fly, yeast and worm. GLAlign is publicly available for academic use at https://sites.google.com/site/globallocalalignment/. Marianna Milano, Mario Cannataro, Pietro H. Guzzi |
BIBM | 2 |
| 2016 | DIETOS: A recommender system for adaptive diet monitoring and personalized food suggestionabstractNowadays there is a widespread diffusion of mobile applications for weight and diet management. Even though, the most popular apps are not usually experimented in clinical contexts, as well as apps are not supported by medical evidence. Further research is necessary to assess the effectiveness of apps for weight and diet management. Moreover, there are few examples of food recommender systems that provide to the users nutritional facts about suitable food choices and take into account individual physiological status and environmental situations. We propose DIETOS (DIET Organizer System), a recommender system for the adaptive delivery of nutrition contents to improve the quality of life of both healthy people and individuals affected by chronic diet-related diseases. The proposed system is able to build a user's health profile, and provides individualized nutritional recommendation according to the health profile. The profile is created through the use of dynamic real-time questionnaires prepared by medical doctors and compiled by the users. The health profile includes information about health status and eventual chronic diseases. The first prototype of the system (available online at http://www.easyanalysis.it/dietos), includes a catalogue of typical Calabrian foods compiled by nutrition specialists (Calabria is a region of the southern Italy). DIETOS can suggest not only the use of specific foods compatible with the health status, but also it may give dietary indications related to some specific pathologies or health conditions. Giuseppe Agapito, Barbara Calabrese, Pietro H. Guzzi, Mario Cannataro, Mariadelina Simeoni, Ilaria Care, Theodora Lamprinoudi, Giorgio Fuiano, Arturo Pujia |
WiMob | 4 |
| 2016 | Methodologies and experimental platforms for generating and analysing microarray and mass spectrometry-based omics data to support P4 medicineabstractPredictive, preventive, personalized and participatory (P4) medicine is an emerging medical model that is based on the customization of all medical aspects (i.e. practices, drugs, decisions) of the individual patient. P4 medicine presupposes the elucidation of the so-called omic world, under the assumption that this knowledge may explain differences of patients with respect to disease prevention, diagnosis and therapies. Here, we elucidate the role of some selected omics sciences for different aspects of disease management, such as early diagnosis of diseases, prevention of diseases, selection of personalized appropriate and optimal therapies based on molecular profiling of patients. After introducing basic concepts of P4 medicine and omics sciences, we review some computational tools and approaches for analysing selected omics data, with a special focus on microarray and mass spectrometry data, which may be used to support P4 medicine. Some applications of biomarker discovery and pharmacogenomics and some experiences on the study of drug reactions are also described. Pietro H. Guzzi, Giuseppe Agapito, Marianna Milano, Mario Cannataro |
Briefings Bioinform. | 4 |
| 2016 | Extracting Cross-Ontology Weighted Association Rules from Gene Ontology AnnotationsabstractGene Ontology (GO) is a structured repository of concepts (GO Terms) that are associated to one or more gene products through a process referred to as annotation. The analysis of annotated data is an important opportunity for bioinformatics. There are different approaches of analysis, among those, the use of association rules (AR) which provides useful knowledge, discovering biologically relevant associations between terms of GO, not previously known. In a previous work, we introduced GO-WAR (Gene Ontology-based Weighted Association Rules), a methodology for extracting weighted association rules from ontology-based annotated datasets. We here adapt the GO-WAR algorithm to mine cross-ontology association rules, i.e., rules that involve GO terms present in the three sub-ontologies of GO. We conduct a deep performance evaluation of GO-WAR by mining publicly available GO annotated datasets, showing how GO-WAR outperforms current state of the art approaches. Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2015 | Overall Survival Analyzer: A software tool to analyze genotyping and clinical data enriched with temporal eventsabstractThe estimation of survival distributions of patients is an important current problem in clinical oncology. The current trend is to integrate molecular data (such as genomic data) with clinical data (e.g. cancer type, stage of the disease, etc) and then to link survival distributions to molecular profile of patients. Recently, the Affymetrix DMET (Drug Metabolizing Enzymes and Transporters) microarray technology has enabled the possibility to determine the allelic variants of a patient and to relate them to phenotype (e.g. drug toxicity). Therefore, the analysis of survival distribution of patients starting from their profile obtained using DMET data may reveal important knowledge to clinicians. In order to provide support to this analysis we propose Overall Survival Analyzer (OS-Analyzer), a software tool able to compute the Overall Survival and Progression-Free Survival (PFS). The tool is able to perform an automatic analysis of data avoiding wasting time on the manual analysis. OS-Analyzer is available to download at the follows web address: https://sites.google.com/site/overallsurvivalanalyzer/. Giuseppe Agapito, Pietro H. Guzzi, C. Botta, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro |
BIBM | 7 |
| 2015 | A Genetic Algorithm for the selection of structural MRI features for classification of Mild Cognitive Impairment and Alzheimer's DiseaseabstractThis work investigates the problem of feature selection in neuroimaging features from structural MRI brain images for the classification of subjects as healthy controls, suffering from Mild Cognitive Impairment or Alzheimer's Disease. A Genetic Algorithm wrapper method for feature selection is adopted in conjunction with a Support Vector Machine classifier. In very large feature sets, feature selection is found to be redundant as the accuracy is often worsened when compared to an Support Vector Machine with no feature selection. However, when just the hippocampal subfields are used, feature selection shows a significant improvement of the classification accuracy. Three-class Support Vector Machines and two-class Support Vector Machines combined with weighted voting are also compared with the former and found more useful. The highest accuracy achieved at classifying the test data was 65.5% using a genetic algorithm for feature selection with a three-class Support Vector Machine classifier. Alexander Luke Spedding, Giuseppe Di Fatta, Mario Cannataro |
BIBM | 3 |
| 2015 | ICT Solutions for Health Education ModelabstractHealth promotion represents the process to empower the citizens to improve their health lifestyle and to achieve higher levels of wellness. The health models focus on helping people to prevent illnesses through their behavior, and on looking at ways in which a person can pursue better health or ideal health. We report on a project aiming to propose a new model for wellness improvement, consisting in actions to be performed to encourage individuals to become aware of their wellness and develop healthier habits. Domenico Mirarchi, Patrizia Vizza, Mario Cannataro, Pietro H. Guzzi, Giuseppe Tradigo, Pierangelo Veltri |
CBMS | 3 |
| 2015 | DMET-Miner: Efficient discovery of association rules from pharmacogenomic data
Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro |
J. Biomed. Informatics | 3 |
| 2014 | Improving annotation quality in gene ontology by mining cross-ontology weighted association rulesabstractThe Gene Ontology (GO) is the major resource of annotations for genes and proteins. Despite the presence of large efforts to avoid errors and inconsistencies, some unreliabilities are still present. In particular electronically inferred annotations are more unreliable than manual ones and their number is growing. Thus, the need for an accurate evaluation of annotations in an automatic way arises. In the past, some approaches for improving annotation consistencies have been proposed using association rule mining to discover hidden relationships among GO terms. However such approaches consider all the GO terms equally, while GO terms have different Information Content, i.e. different relevance. Consequently we designed a novel algorithm, (GO-WAR), Mining Weighted Association Rules from GO, that is based on the extraction of weighted association rules considering the IC of terms. We evaluated our algorithm considering seven different species and all the GO ontologies. In all the experiments GO-WAR outperformed state of the art approaches. Giuseppe Agapito, Marianna Milano, Pietro H. Guzzi, Mario Cannataro |
BIBM | 4 |
| 2014 | DMET-miner: Efficient learning of association rules from genotyping data for personalized medicineabstractRecent developments of microarray technology enable the investigation of allelic variants that may be correlated to phenotypes. In particular the Affymetrix DMET (Drug Metabolism Enzymes and Transporters) platform enables the simultaneous investigation of all the genes that are related to drug absorption, distribution, metabolism and excretion (ADME) and it has been used in clinical studies. In a previous work we developed DMET-Analyzer, a platform able to automatize the study of allelic variants, that has been validated in clinical studies. DMET-Analyzer is able to correlate a single variant for each probe (related to a portion of a gene) through the use of the Fisher test, on the other hand it is unable to discover multiple associations among allelic variants. To overcome those limitations, here we propose DMET-Miner, that is able to correlate the presence of a set of allelic variants by employing an Apriori-like discovery strategy. Preliminary experiments on a synthetic DMET dataset. Pietro H. Guzzi, Giuseppe Agapito, Maria Teresa Di Martino, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro |
BIBM | 7 |
| 2014 | Biases in information content measurement of gene ontology termsabstractThe Gene Ontology (GO) is used to achieve information about gene and protein functions by using a structured vocabulary of terms (GO Terms). GO Terms are related to biological concepts such as proteins or genes through the annotation process. There exist many different annotation processes identified by different evidence codes (EC). Annotated data are stored in public databases such as the Gene Ontology Annotation (GOA) database. Each term has a different specificity also referred to as Information Content (IC) of terms. Both the structure of GO and the corpora of annotation are continuously subject to change due to novel experimental findings. This process is often referred to as ontology evolution. This work focuses on how changes of annotations affect the IC of terms. The study confirms that statistically significant difference among many whole GOA versions exists on each species. Furthermore, there is also a statistically significant difference considering MF taxonomy for human, yeast, worm and fly. These results convey that annotation corpora changes have a high impact on IC. Marianna Milano, Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro |
BIBM | 4 |
| 2014 | EMG-Miner: Automatic Acquisition and Processing of Electromyographic Signals: First Experimentation in a Clinical Context for Gait Disorders EvaluationabstractIn the context of physical medicine and rehabilitation, gait analysis is the "gold standard" for an effective assessment of any problems in the locomotor patterns. Surface electromyography is one of the exams within the protocol of the gait analysis, allowing an assessment of functional limitations in the walking. Considering the Physical Medicine and Rehabilitation Unit of the University of Catanzaro, physicians are limited to a visual analysis of the electromyographic signals coming from the muscles of the lower limbs, to extract useful information for diagnosis and monitoring of treatment. The objective of this work is to provide to specialists a simple and flexible system that allows the extraction of quantitative synthetic parameters in time and frequency domain from EMG signals, in particular we propose a novel EMG data acquisition and processing system, referred as EMG-Miner, that allows the automated acquisition and analysis of EMG signals along the different stages of the rehabilitation process (follow-up) of a patient. Nicola Ielpo, Barbara Calabrese, Mario Cannataro, Arrigo Palumbo, S. Ciliberti, C. Grillo, Maurizio Iocco |
CBMS | 3 |
| 2014 | coreSNP: Parallel Processing of Microarray DataabstractThe availability of high-throughput technologies, such as next generation sequencing and microarray, and the diffusion of genomics studies to large populations are producing an increasing amount of experimental data. In particular, pharmacogenomics studies the impact of genetic variation on drug response in patients and correlates gene expression or single nucleotide polymorphisms (SNPs) with the toxicity or efficacy of a drug, with the aim to improve drug therapy with respect to the patients’ genotype ensuring maximum efficacy with minimal adverse effects. However, the storage, preprocessing, and analysis of experimental data are becoming a main bottleneck in the pharmacogenomics analysis pipeline, due to the increasing number of genes and patients investigated. This paper presents a new parallel software tool named coreSNP for the parallel preprocessing and statistical analysis of DMET (Drug Metabolism Enzymes and Transporters) SNP microarray data produced by Affymetrix for pharmacogenomics studies. The scalable multi-threaded implementation of coreSNP allows to handle the huge volumes of experimental pharmacogenomics data in a very efficient way, while its easy to use graphical user interface and its ability to annotate significant SNPs allow biologists to interpret the results easily. Performance evaluation conducted using real datasets shows good speed-up and scalability and effective response times. Pietro H. Guzzi, Giuseppe Agapito, Mario Cannataro |
IEEE Trans. Computers | 3 |
| 2013 | Using open data in health care and tourismabstractOpen Data refers to the possibility of freely sharing data among users and organization. Similarly Open Government Initiatives refer to the sharing of documents and data among public governments and citizens. Here we focus on an open initiative held by the Italian Ministry of Health who is making available through Internet a set of Open Data about drug stores, health centers, and other health-related data. In particular, we propose a Cloud-based software tool able to gather and integrate different datasets made available by the Italian Ministry of Health. The proposed Cloud-based tool, called Open Health Data for Tourist (OHT), is able to offer to the tourist information about nearest health care providers (drug stores, public emergency room, hospitals and medical doctors) in Italy through an application accessible from mobile devices. Mario Cannataro, Pietro H. Guzzi, Pierangelo Veltri |
BIBM | 1 |
| 2013 | Application of different classification techniques on brain morphological dataabstractThe increasing number of people affected by Neurodegenerative diseases and the improvement of brain imaging diagnostic techniques are bringing to a massive production of brain images that need demanding preprocessing and analysis algorithms. We analyzed volumetric measures of critical brain areas by using different Data Mining methods. Structural magnetic resonance images, generated in our university, were preprocessed using a fully automated segmentation method and the extracted volumetric information was then analyzed by using different binary classifiers. We performed three binary classification experiments considering different data mining algorithms and neurological diseases. Naïve Bayes outperformed all the others classifiers in two experiments, obtaining respectively 93.75% and 95.00% accuracy, while in the third experiment the best classifier was SVM but with a lower accuracy (58,56%). Afterwards, using the Stacking technique we combined the predictions from the best detected three models to build a meta-learner. Meta-learner classification results suggest that the application of the Stacking technique needs more experimentation and the test of additional stackers. Alessia Sarica, Claudia Critelli, Pietro H. Guzzi, Antonio Cerasa, Aldo Quattrone, Mario Cannataro |
CBMS | 6 |
| 2013 | Visualization of protein interaction networks: problems and solutionsabstractBACKGROUND: Visualization concerns the representation of data visually and is an important task in scientific research. Protein-protein interactions (PPI) are discovered using either wet lab techniques, such mass spectrometry, or in silico predictions tools, resulting in large collections of interactions stored in specialized databases. The set of all interactions of an organism forms a protein-protein interaction network (PIN) and is an important tool for studying the behaviour of the cell machinery. Since graphic representation of PINs may highlight important substructures, e.g. protein complexes, visualization is more and more used to study the underlying graph structure of PINs. Although graphs are well known data structures, there are different open problems regarding PINs visualization: the high number of nodes and connections, the heterogeneity of nodes (proteins) and edges (interactions), the possibility to annotate proteins and interactions with biological information extracted by ontologies (e.g. Gene Ontology) that enriches the PINs with semantic information, but complicates their visualization. METHODS: In these last years many software tools for the visualization of PINs have been developed. Initially thought for visualization only, some of them have been successively enriched with new functions for PPI data management and PIN analysis. The paper analyzes the main software tools for PINs visualization considering four main criteria: (i) technology, i.e. availability/license of the software and supported OS (Operating System) platforms; (ii) interoperability, i.e. ability to import/export networks in various formats, ability to export data in a graphic format, extensibility of the system, e.g. through plug-ins; (iii) visualization, i.e. supported layout and rendering algorithms and availability of parallel implementation; (iv) analysis, i.e. availability of network analysis functions, such as clustering or mining of the graph, and the possibility to interact with external databases. RESULTS: Currently, many tools are available and it is not easy for the users choosing one of them. Some tools offer sophisticated 2D and 3D network visualization making available many layout algorithms, others tools are more data-oriented and support integration of interaction data coming from different sources and data annotation. Finally, some specialistic tools are dedicated to the analysis of pathways and cellular processes and are oriented toward systems biology studies, where the dynamic aspects of the processes being studied are central. CONCLUSION: A current trend is the deployment of open, extensible visualization tools (e.g. Cytoscape), that may be incrementally enriched by the interactomics community with novel and more powerful functions for PIN analysis, through the development of plug-ins. On the other hand, another emerging trend regards the efficient and parallel implementation of the visualization engine that may provide high interactivity and near real-time response time, as in NAViGaTOR. From a technological point of view, open-source, free and extensible tools, like Cytoscape, guarantee a long term sustainability due to the largeness of the developers and users communities, and provide a great flexibility since new functions are continuously added by the developer community through new plug-ins, but the emerging parallel, often closed-source tools like NAViGaTOR, can offer near real-time response time also in the analysis of very huge PINs. Giuseppe Agapito, Pietro H. Guzzi, Mario Cannataro |
BMC Bioinform. | 3 |
| 2013 | Visual Data Mining of Biological Networks: One Size Does Not Fit AllabstractHigh-throughput technologies produce massive amounts of data. However, individual methods yield data specific to the technique used and biological setup. The integration of such diverse data is necessary for the qualitative analysis of information relevant to hypotheses or discoveries. It is often useful to integrate these datasets using pathways and protein interaction networks to get a broader view of the experiment. The resulting network needs to be able to focus on either the large-scale picture or on the more detailed small-scale subsets, depending on the research question and goals. In this tutorial, we illustrate a workflow useful to integrate, analyze, and visualize data from different sources, and highlight important features of tools to support such analyses. Chiara Pastrello, David Otasek, Kristen Fortney, Giuseppe Agapito, Mario Cannataro, Elize Shirdel, Igor Jurisica |
PLoS Comput. Biol. | 5 |
| 2012 | Knowledge-based compilation of magnetic resonance diagnosis reports in neuroradiologyabstractThe compilation of the neuroradiology diagnosis reports based on magnetic resonance (MR) exams comprises a deep analysis of images and related numerical values, usually done with specialized image processing tools, and the compilation of the different parts forming the reports, following well-defined schemes which depend on the kind of exam and pathology. Although the diagnosis report is a semi-structured document comprising different well-defined parts, usually it is compiled using simple text editors that may loose its structure. The drawback of this approach is twofold, first of all the specialist has to repeat the writing of some texts for each report, yielding to a time consuming process, and second and most importantly, the precious data, annotations and comments written in the report are not easily available for further analysis. In fact, when the information contained in the diagnosis reports is stored into unstructured documents such as texts, it is very difficult to query and extract useful information needed for conducting studies on large populations of patients. In this paper we propose a novel software tool able: (i) to store neuroradiology diagnosis reports and their schemes in a structured knowledge-base; and (ii) to support the specialist in the compilation of new diagnosis reports on the basis of the schemes and contents already stored in the knowledge-base. Mario Cannataro, Orlando Alfieri, Francesco Fera |
CBMS | 1 |
| 2012 | CytoMCL: A Cytoscape plugin for fast clustering of protein interaction networksabstractThe analysis of the whole set of molecular interactions in an organism, often referred to as interaction networks, is becoming an important research area. A main approach for such analysis resides on the application of clustering techniques to such networks. The meaning of discovered clusters, (i.e. highly interconnected regions), is strictly related to the type of networks. For instance in protein-protein interaction networks clusters may represent protein complexes. The Markov Clustering Algorithm (MCL) is a wellknown algorithm for clustering graphs. It does not provide a graphical user interface and cannot be used in the Cytoscape platform. We present CytoMCL a Cytoscape plugin that finds clusters in a graph by using MCL. It is based on an intuitive interface it is able to load a network from Cytoscape, to analyze it and to visualize resulting clusters into Cytoscape. Pietro H. Guzzi, Mario Cannataro |
CBMS | 2 |
| 2012 | SySQ: A Web-based system for survey and questionnaire management in medicineabstractA questionnaire is a method for collecting data that can come from many different sources: from observations, telephone interviews or documentary sources. Whatever the source of data is, the questionnaire provides a framework of questions that facilitate researcher's work. A manual approach for collecting data using questionnaire, presents some limitations and introduces several sources of errors. An Informatic System was implemented to reduce these errors and to support researchers in epidemiological studies. For the experimentation of the prototype, a paper format questionnaire has been digitalized from the original sheet provided by the Chair of Hygiene of the Magna Graecia University of Catanzaro. The implemented system allows researchers to create questionnaires, adding sections and structured questions. The administrator of the system can visualize a preview of the questionnaire and decide which group of users can compile it. The system provides a control panel to analyze the collected data and that permits the exportation of saved data into statistical software compatible formats. Alessia Sarica, Pietro H. Guzzi, Domenico Flotta, Carmelo G. A. Nobile, Mario Cannataro |
CBMS | 5 |
| 2012 | audioEPR: A specialized electronic patient records for the semi-automatic management of clinical data in audiologyabstractGeneral-purpose Electronic Patient Records (EPR) may lack in support of specialistic clinical data that characterizes each clinical domain. In many clinical settings often such data remain inside the computer systems associated with specialistic instruments and are not readily available to clinicians for conducting large studies for research purposes. In this paper we present audioEPR, a specialized EPR for the semi-automatic management and querying of audiological and otoneurological clinical data. The goal of the tool is to support the day-by-day clinical activity as well as the simple selection and extraction of clinical data for research activity. The realized prototype is able to support the semi-automatic storage of data related to different Audiology tests and the generation of related diagnostic reports. To encourage its use in a clinical setting its interface resembles the form of the paper-based documents currently used by operators, while maintaining the formal correctness of the produced documentation. On the other hand, its database allows an easy selection and extraction of set of clinical data for research purposes. Salvatore Scaramuzzino, Pietro H. Guzzi, Claudio Petrolo, Giuseppe Chiarella, Pierangelo Veltri, Mario Cannataro |
CBMS | 6 |
| 2012 | Semantic similarity analysis of protein data: assessment with biological features and issuesabstractThe integration of proteomics data with biological knowledge is a recent trend in bioinformatics. A lot of biological information is available and is spread on different sources and encoded in different ontologies (e.g. Gene Ontology). Annotating existing protein data with biological information may enable the use (and the development) of algorithms that use biological ontologies as framework to mine annotated data. Recently many methodologies and algorithms that use ontologies to extract knowledge from data, as well as to analyse ontologies themselves have been proposed and applied to other fields. Conversely, the use of such annotations for the analysis of protein data is a relatively novel research area that is currently becoming more and more central in research. Existing approaches span from the definition of the similarity among genes and proteins on the basis of the annotating terms, to the definition of novel algorithms that use such similarities for mining protein data on a proteome-wide scale. This work, after the definition of main concept of such analysis, presents a systematic discussion and comparison of main approaches. Finally, remaining challenges, as well as possible future directions of research are presented. Pietro H. Guzzi, Marco Mina, Concettina Guerra, Mario Cannataro |
Briefings Bioinform. | 4 |
| 2012 | DMET-Analyzer: automatic analysis of Affymetrix DMET DataabstractBACKGROUND: Clinical Bioinformatics is currently growing and is based on the integration of clinical and omics data aiming at the development of personalized medicine. Thus the introduction of novel technologies able to investigate the relationship among clinical states and biological machineries may help the development of this field. For instance the Affymetrix DMET platform (drug metabolism enzymes and transporters) is able to study the relationship among the variation of the genome of patients and drug metabolism, detecting SNPs (Single Nucleotide Polymorphism) on genes related to drug metabolism. This may allow for instance to find genetic variants in patients which present different drug responses, in pharmacogenomics and clinical studies. Despite this, there is currently a lack in the development of open-source algorithms and tools for the analysis of DMET data. Existing software tools for DMET data generally allow only the preprocessing of binary data (e.g. the DMET-Console provided by Affymetrix) and simple data analysis operations, but do not allow to test the association of the presence of SNPs with the response to drugs. RESULTS: We developed DMET-Analyzer a tool for the automatic association analysis among the variation of the patient genomes and the clinical conditions of patients, i.e. the different response to drugs. The proposed system allows: (i) to automatize the workflow of analysis of DMET-SNP data avoiding the use of multiple tools; (ii) the automatic annotation of DMET-SNP data and the search in existing databases of SNPs (e.g. dbSNP), (iii) the association of SNP with pathway through the search in PharmaGKB, a major knowledge base for pharmacogenomic studies. DMET-Analyzer has a simple graphical user interface that allows users (doctors/biologists) to upload and analyse DMET files produced by Affymetrix DMET-Console in an interactive way. The effectiveness and easy use of DMET Analyzer is demonstrated through different case studies regarding the analysis of clinical datasets produced in the University Hospital of Catanzaro, Italy. CONCLUSION: DMET Analyzer is a novel tool able to automatically analyse data produced by the DMET-platform in case-control association studies. Using such tool user may avoid wasting time in the manual execution of multiple statistical tests avoiding possible errors and reducing the amount of time needed for a whole experiment. Moreover annotations and the direct link to external databases may increase the biological knowledge extracted. The system is freely available for academic purposes at: https://sourceforge.net/projects/dmetanalyzer/files/ Pietro H. Guzzi, Giuseppe Agapito, Maria Teresa Di Martino, Mariamena Arbitrio, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro |
BMC Bioinform. | 7 |
| 2011 | Challenges in microarray data management and analysisabstractMicroarray is a key technology in genomics and is increasingly used in molecular biology as well as in molecular medicine and clinical applications. The availability of different microarray types and vendors and the increasing number of samples forming microarray studies pose new challenges in the management and analysis of microarray data. The paper recalls main microarray types and goals and discusses the most important challenges in microarray management and analysis, including the following issues. Heterogeneity in microarray data format requires the application of different preprocessing and annotation tools and the use of different, vendor-specific, preprocessing libraries. Multiplicity of microarray types (e.g. gene expression, SNPs and miRNA arrays, to cite a few) requires the application of different analysis tools. The application of various preprocessing steps produces a number of different files for each study, such as raw, preprocessed, annotated and eventually filtered data, that must be managed and stored properly. Finally, the increasing volume of microarray data due to the use of large sets of samples poses further challenges for the storage of such data. Finally, some emerging directions to face such challenges are also described. Pietro H. Guzzi, Mario Cannataro |
CBMS | 2 |
| 2011 | Automatic summarisation and annotation of microarray data
Pietro H. Guzzi, Maria Teresa Di Martino, Giuseppe Tradigo, Pierangelo Veltri, Pierfrancesco Tassone, Pierosandro Tagliaferri, Mario Cannataro |
Soft Comput. | 7 |
| 2010 | mu-CS: An extension of the TM4 platform to manage Affymetrix binary dataabstractBACKGROUND: A main goal in understanding cell mechanisms is to explain the relationship among genes and related molecular processes through the combined use of technological platforms and bioinformatics analysis. High throughput platforms, such as microarrays, enable the investigation of the whole genome in a single experiment. There exist different kind of microarray platforms, that produce different types of binary data (images and raw data). Moreover, also considering a single vendor, different chips are available. The analysis of microarray data requires an initial preprocessing phase (i.e. normalization and summarization) of raw data that makes them suitable for use on existing platforms, such as the TIGR M4 Suite. Nevertheless, the annotations of data with additional information such as gene function, is needed to perform more powerful analysis. Raw data preprocessing and annotation is often performed in a manual and error prone way. Moreover, many available preprocessing tools do not support annotation. Thus novel, platform independent, and possibly open source tools enabling the semi-automatic preprocessing and annotation of microarray data are needed. RESULTS: The paper presents mu-CS (Microarray Cel file Summarizer), a cross-platform tool for the automatic normalization, summarization and annotation of Affymetrix binary data. mu-CS is based on a client-server architecture. The mu-CS client is provided both as a plug-in of the TIGR M4 platform and as a Java standalone tool and enables users to read, preprocess and analyse binary microarray data, avoiding the manual invocation of external tools (e.g. the Affymetrix Power Tools), the manual loading of preprocessing libraries, and the management of intermediate files. The mu-CS server automatically updates the references to the summarization and annotation libraries that are provided to the mu-CS client before the preprocessing. The mu-CS server is based on the web services technology and can be easily extended to support more microarray vendors (e.g. Illumina). CONCLUSIONS: Thus mu-CS users can directly manage binary data without worrying about locating and invoking the proper preprocessing tools and chip-specific libraries. Moreover, users of the mu-CS plugin for TM4 can manage Affymetrix binary files without using external tools, such as APT (Affymetrix Power Tools) and related libraries. Consequently, mu-CS offers four main advantages: (i) it avoids to waste time for searching the correct libraries, (ii) it reduces possible errors in the preprocessing and further analysis phases, e.g. due to the incorrect choice of parameters or the use of old libraries, (iii) it implements the annotation of preprocessed data, and finally, (iv) it may enhance the quality of further analysis since it provides the most updated annotation libraries. The mu-CS client is freely available as a plugin of the TM4 platform as well as a standalone application at the project web site (http://bioingegneria.unicz.it/M-CS). Pietro H. Guzzi, Mario Cannataro |
BMC Bioinform. | 2 |
| 2010 | IMPRECO: Distributed prediction of protein complexes
Mario Cannataro, Pietro H. Guzzi, Pierangelo Veltri |
Future Gener. Comput. Syst. | 1 |
| 2010 | Special section: Biomedical and bioinformatics challenges to computer science
Mario Cannataro, Mathilde Romberg, Joakim Sundnes, Rodrigo Weber dos Santos |
Future Gener. Comput. Syst. | 1 |
| 2009 | GridSnake: A Grid-based implementation of the Snake segmentation algorithmabstractMedical imaging is becoming a key technique to visualize the internal structure of the body. Magnetic resonance imaging (MRI) is currently used to take different spatial images of organs, such as the heart. The output of such an analysis is a set of images representing different views of the body or of an organ. There exist many algorithms to pre-process and analyze medical images such as the well known Snake segmentation algorithm. An issue in medical imaging is the large size of images, that require large and efficient data stores, and the high computational power needed to process them. For these reasons, the Grid is being more and more used as an ideal environment for medical image processing. This paper presents a first experience in porting the snake algorithm on a Globus-based Grid. Mario Cannataro, Pietro H. Guzzi, Marcelo Lobosco, Rodrigo Weber dos Santos |
CBMS | 1 |
| 2009 | Using ontologies for annotating and retrieving protein-protein interactions dataabstractProtein-protein interaction (PPI) databases store the whole set of protein interactions in organism. In spite of the availability of much biological information spread on different sources (e.g. Gene Ontology), neither proteins nor interactions are generally annotated in PPI databases. This results in very poor querying capabilities of PPI databases that enable very simple queries. The annotation of proteins and interactions stored in PPI databases may allow the implementation of more powerful querying interfaces. The paper presents a software architecture for the annotation of existing PPI databases with information extracted from Gene Ontology. A simple extension of the query interface of an existent PPI database is discussed. Mario Cannataro, Pietro H. Guzzi, Pierangelo Veltri |
CBMS | 1 |
| 2009 | StiMaRe: A software tool supporting visual stimuli definition and analysis in magnetic resonanceabstractAnalyzing physiological brain responses to external stimuli helps neuroscientists to elucidate human behaviour and, more generally improves knowledge of neurological patients profile. It is well known that functional magnetic resonance imaging (fMRI) can provide important information when stimulating sensorimotor or cognitive functions in humans. In this paper we present a software tool supporting analysis of fMRI datasets. The tool allows medical operators to build sequences of stimuli that are presented to subjects within the MRI scanner and to define critical task parameters and timings. Patients feedbacks are recorded through a fiber-optic computer-controlled MR compatible system during the fMRI acquisition phase. The proposed tool, called StiMaRe (for Stimuli definition and analysis in Magnetic Resonance), includes a database layer allowing to store the defined pattern with the patient feedbacks. An XML based framework allows to distribute the pattern stimuli and results to remote sites, allowing the reissuing of the experiment on different samples. Giuseppe Tradigo, Pierangelo Veltri, Mario Cannataro, Francesco Fera |
CBMS | 3 |
| 2008 | myMCL: A Web Portal for Protein Complexes PredictionabstractInteractomics is the study of the Interactome, i.e. the whole set of macromolecular interactions within a cell. Proteins interact among them and different interactions are represented as graphs named Protein to Protein Interaction (PPI) networks. The interest in analyzing PPI networks is related to the possibility of predicting PPI properties on the basis of global properties of the graph (e.g. verify if homology among species involves PPI similarity), or to find set of protein interactions that has a biological meaning. The prediction of protein complexes has been faced in the last years by using different clustering algorithms. The Markov Clustering algorithm (MCL) is a method that presents one of the best performance but is currently available only as a stand alone application with a simple command-line interface available only on Linux platforms. Following a trend in bioinformatics, we provide a web portal (myMCL) allowing remote users to access MCL functions through the Internet. myMCL enables user to submit a job and stores results in a local database for further processing. Mario Cannataro, Pietro H. Guzzi, Pierangelo Veltri |
CBMS | 1 |
| 2008 | A Tool for the Semiautomatic Acquisition of the Morphological Data of Blood Vessel NetworksabstractThe simulation of the dynamics of the blood flow in the venous system of the lower limb is an important tool for supporting clinical research and for suggesting possible treatments for many diseases, e.g. for enhancing the surgical treatment of chronic venous insufficiency (CVI). Nevertheless the accuracy of the simulation of the blood flow is strictly related to the morphological data characterizing the investigated venous system. Although some of these data can be extracted from the observation of the real blood flow of a patient, e.g. through the acquisition of a set of images, the extraction of such values is often performed in a manual way, so the need for the automatic induction of parameters arises. The paper presents a software module that allows the semiautomatic acquisition of the morphological data of the venous system of a patient. The tool, developed as a plugin of the ImageJ imaging platform, receives in input a DICOM file containing the computerized tomography (CT) of the vessels network of the lower limb, and produces in a semi-automatic way a weighted graph of the network. This model can be used as the input for a subsequent simulation of the system. Mario Cannataro, Pietro H. Guzzi, Giuseppe Tradigo, Pierangelo Veltri |
ISPA | 1 |
| 2008 | Computational proteomics: management and analysis of proteomics dataabstractProteomics is about the study of the proteins expressed in a cell, organism, or tissue. This includes protein identification and quantification (or quantitation), protein–protein interactions, protein complexes prediction, protein modifications and protein localization in the cell. Mass Spectrometry (MS) is one of the main technologies in proteomics and is more and more used for its increasing precision and for the possibility to automate the proteomics analysis pipeline, yielding to large-scale high-throughput experiments. Since proteins play a central role in the life of an organism, proteomics is instrumental in many biomedical applications, such as biomarker discovery and drug treatment evaluation, as well as for investigating the dynamics of cells in Systems Biology. Computational Proteomics is about the computational methods, algorithms, databases and methodologies used to process, manage, analyze and interpret the data produced in proteomics experiments. The broad application of proteomics in different biological and medical fields, as well as the diffusion of high-throughput platforms, leads to increasing volumes of available proteomics data requiring efficient algorithms, new data management capabilities and novel analysis, inference and visualization techniques. Moreover, high-throughput production and collection of data pose new challenges in data handling and reusability as well as in tools interoperability and interconnection. On the other hand, the increasing availability of data and tools opens new research directions and opportunities (e.g. annotated spectral libraries) that can be exploited only through the rigorous application of computer science, machine learning, knowledge discovery, statistics and signal processing techniques. As in many scientific disciplines, the final goal of Computational Proteomics is to infer knowledge models (e.g. verify a hypothesis or identify proteins involved in a disease) from the inspection of biological samples. Such an activity involves different steps, happening both in wet and dry lab, and is the result of combination of many instruments, methods, tools, algorithms, databases, according to established or emerging working methodologies and standards. This special issue explores the current state-of-the-art research taking place in different areas of Computational Proteomics, with special emphasis on methodologies and tools for data handling and analysis, machine learning, knowledge discovery, biomarker discovery, data standardization and information quality, in MS-based proteomics. The discussion of all experimental techniques and computational methods taking place in proteomics, would be too large to be hosted in one journal issue, thus, this special issue is complemented by the March 2008 issue of Briefings in Functional Genomics and Proteomics, that discusses, among others, MS-based techniques for improving the study of protein structure and protein folding, as well as some applications and tools in Interactomics and Systems Biology. MS permits, with high accuracy, the determination of molecular weight of chemical compounds, ranging from small molecules to large, polar biopolymers [1]. The mass spectrometer separates gas phase ions according to their mass to charge ratio values. The output of the spectrometer (the spectrum) is a large sequence of value pairs. Each pair contains a measured intensity and a mass to charge ratio (m/z), which depend, respectively, on the quantity and molecular mass of the detected molecule. Although MS is only able to identify molecular masses, the adoption of advanced sample preparation techniques, such as chromatographic separation or labeling techniques, the use of protein database information, e.g. the theoretical spectrum associated to a protein sequence, and the application of powerful machine learning algorithms, make this technique a basic cornerstone for proteomics applications. In a proteomics experiment different steps can be identified, each one having one or more parameters of choice: (i) sample preparation, including separation and labeling; (ii) MS experiment, including mass spectrometer choice and configuration; (iii) spectra preprocessing, including spectra signal deconvolution, often happening inside the spectrometer software, baseline subtraction, noise removal, dimension reduction, peaks extraction, outlier detection; (iv) peptide/protein identification, including database searching, eventually coupled with de novo sequencing or sequence tagging; (v) peptide/protein quantitation, either performed through stable isotope labeling or through intensity profiling and (vi) knowledge discovery, including biomarker discovery in clinical applications, systems biology modeling and biological behavior explanation. In this step, a major problem to be faced by data mining algorithms is the extremely high dimensionality of proteomics data (i.e. the number of features or variables) that largely exceeds the size of sample set. While recent papers concentrate especially on protein/peptide identification and quantitation [2, 3], this special issue focuses on the overall knowledge discovery process behind computational proteomics, with special emphasis on machine learning methods, spectra data handling, biomarker discovery, standard-based and quality-aware management of proteomics experiments. The first group of papers expounds the biomarker discovery problem, presenting general methodologies for dimensionality reduction (DR) in proteomics data, machine learning methods and way to interconnect them for conducting biomarker discovery studies and two novel classification algorithms with special features for proteomics studies. The second group of papers reviews basic algorithms for spectra analysis and management and describes issues and opportunities when managing large datasets. Peptide/protein identification and quantitation are the basic building blocks for the comparative analysis of quantitative information (i.e. protein expression level) about identified proteins. Methods for protein quantitation and in particular for label-free comparative quantitation are reviewed and compared. The third group of papers presents two important and timely issues: standardization of proteomics data and methodologies and tools to ensure information quality in proteomics experiments. The main goal of biomarker discovery is to find the most discriminating features in a classification of samples (e.g. peaks or peptides/proteins discriminating healthy versus diseased subjects). An important problem is the huge number of features (e.g. peaks or peptides/proteins) that may represent potential biomarkers in a mass spectrum. The low number of cases/controls with respect to the number of features is another important problem. Moreover, collected data are sparse. The combination of such characteristics yields the so-called high-dimensional small-sample problem that may lead to selection bias and information leakage when analyzing data. Hilario and Kalousis describe in a comprehensive way the computational methods involved in biomarker discovery studies. They first focus on the high-dimensional small-sample problem, underlying significant data-analytical issues raised by biomarker discovery. Then, they describe the phases of the knowledge discovery process—see step (vi) explained earlier—in biomarker discovery, focusing on DR methods, learning methods and the way they can be coupled. The article provides taxonomy of DR methods that includes feature transformation and feature selection methods. The former transform or combine old features generating a smaller set of new features, while the latter eliminate directly irrelevant or redundant features. Considering how DR is coupled to the learning process, the authors identify filter methods that perform DR as a preprocessing step before the learning method, and wrapper methods that wrap DR around a specific learning process. The article then discusses the so-called DR method selection problem, i.e. the criteria and an objective evaluation methodology, which may be used to select a DR method. Barla et al. propose a Design Analysis Protocol (DAP) for organizing and controlling the different steps involved in biomarker discovery. The aim of the proposed methodology, inspired by similar initiatives in the microarray community, is to provide a set of guidelines for the development and validation of predictive models or classifiers. In order to ensure reproducibility of experiments and reliability of produced results, the authors focus on two critical problems in MS-based proteomics: information leakage and selection bias, also considering the above cited high-dimensional small-sample problem. After defining main blocks of a general predictive proteomics pipeline and illustrating key characteristics of elementary functions (e.g. spectra preprocessing, feature selection and learning models) they argue that reproducibility strongly depends on the methods used for selecting training and test datasets from the original dataset. Experimental studies based on publicly available datasets and on synthetic data are provided, including an example showing the effects of varying the preprocessing module in the pipeline. Villmann et al. present two recently developed classification algorithms for analysis of spectra that face the high-dimensional small-sample problem: the supervised neural gas and the fuzzy labeled self-organizing map algorithms. From a mathematical point of view, a consequence of the high-dimensional small-sample problem is that the data space to be explored is sparsely filled. Data separation during classification requires detecting the underlying data regularities. The presented algorithms are inherently self-regularizing and thus are able to deal with high-dimensional, sparse and noisy data. Moreover, the fuzzy labeled self-organizing map algorithm allows the processing of uncertain class information and returns a fuzzy classification scheme that allows deep class analysis. Moreover, encoding similar class information with similar colors allows class dependent data visualization. The way in which the proteomics steps (i)−(vi) are applied makes computational proteomics a multidimensional activity, leading to possibly diverse intermediate results on the basis of the choices made at each step. The second and third groups of papers deal with these aspects. For instance, the different combination of separation techniques [e.g. two-dimensional gel electrophoresis (2-DE) or liquid chromatography (LC)], ionization techniques [e.g. surface enhanced laser desorption ionization (SELDI) or matrix-assisted laser desorption ionization (MALDI)] and mass spectrometer type and configuration (e.g. MS, tandem MS/MS, etc.), as well as any eventual labeling techniques [e.g. isotope-coded affinity tags (ICAT) or stable isotope labeling with amino acids in cell culture (SILAC)], leads to different kind of result spectra, that need to be interpreted accordingly. Thus, steps (i)–(ii) determine the kind of produced spectra and the kind of proteomics analysis that can be performed. The preprocessing step (iii) contains other elements of choice and represents a tradeoff between the need to produce compact data to be analyzed and the need to avoid information leakage and selection bias in further knowledge discovery processes. Veltri, in his paper, discusses characteristics of spectra data, describing the data management issues posed when large datasets have to be managed, and underlying how the results obtained by the database, XML and information retrieval communities, can be beneficial to Computational Proteomics. As an example, an improvement of the peptide/protein identification process based on a library of peptides and advanced database technologies is presented. Moreover, an application of time series analysis to biomarker discovery is also discussed. Stable isotope labeling is a common technique for comparing protein levels between samples, but it requires a careful sample preparation that can be costly for large numbers of samples. The comparative quantification of label-free LC-MS/MS data is an emerging alternative to isotope labeling. The main idea is that under well-controlled conditions identical peptides across different LC-MS/MS datasets can be compared directly without the use of isotope labeling. Wong et al. review two main computational methods for comparative quantitative analysis of LC-MS proteomics data: extraction of peptide ion intensities and spectral counting. They survey available algorithms and tools and discuss statistical tests for evaluating the significance of the comparative results. They argue that the computational process behind intensity-based quantification is significantly more complex compared with spectral counting. On the other hand, spectral counting is more likely to be influenced by the acquisition program of the mass spectrometer and high abundance proteins can mask low abundance ones. Thus, although spectral counting represents the simplest method, both methods could be used simultaneously to improve confidence. Continuing along the computational proteomics pipeline, it is possible to note as the binding between the kind of spectra produced and the type of performed analysis is done by manual association or is implicit when using the software tools provided by the mass spectrometer vendor. While portable, open and vendor-neutral representation of structured data like spectra can be provided by the XML and RDF World Wide Web languages, semantic associations between the type, meaning and format of the generated data and the applicable preprocessing and analysis activities could be managed through domain ontologies. The use of ontologies to model the overall analysis pipeline and the relationship and interconnection between the steps of proteomics analysis is in its beginning [4]. Orchard and Hermjakob describe the HUPO (Human Proteome Organization) Proteomics Standards Initiative, which aims to provide the data standards and interchange formats for the storage and sharing of proteomics data. They present the current state of these standards and describe some standards compliant proteomics data repositories. It should be noted that such emerging initiative includes controlled vocabularies for guiding metadata generation and data annotation and domain ontologies (e.g. the HUPO PSI Mass Spectrometry mzOntology) to model relationship between objects of the domain. Currently available proteomics pipelines, although offering a set of tools for preprocessing and analysis in a comprehensive suite, do not offer the possibility of varying the internal preprocessing or analysis tools. While interoperability at the data layer is starting to be implemented through the exchange of data in standard formats (e.g. the HUPO PSI mzML standard discussed in Orchard and Hermjakob), interoperability at the application level is far from being completed. Web services and workflows are key technologies for enabling application-to-application interoperability, overcoming the limits of current web-based bioinformatics tools that are accessed through human intervention [5, 6]. Different bioinformatics tools are beginning to be offered as web services that can be composed through workflows, but agreed semantics and standard result formats are needed. Orchard and Hermjakob also report on the HUPO PSI common syntax for peptide/protein identification and for protein modification description, useful for capturing results from MS search engines and for giving them as input parameters to further analysis tools. Because the proteomics pipeline presents many steps and different options for each of them, the presence of inaccuracies throughout the processing pipeline may prevent reuse of data for comparative analysis and for experiment validation. Information quality of proteomics experiments and proteomics data repositories is thus an important issue that needs to be managed in an explicit way throughout the entire proteomics pipeline. Stead et al. describe factors that impact on the quality of experimental data and review current approaches for information quality management in proteomics. Data quality issues are considered along the entire pipeline of a proteomics experiment, from experiment design and technique selection, through data analysis, to archiving and sharing. To do this they define a set of quality parameters (e.g. completeness, reproducibility, accuracy and uniformity of information), and attempt to identify how they are affected, focusing on 2-DE and MS technologies and peptide/protein identification. They describe minimum information guidelines, good practice guidelines and standard data formats for achieving and managing information quality. Finally, they introduce an information quality management system with the capability to measure, filter and flag proteomics data based on its quality, and able to access quality-checking web services. In summary, high-throughput MS facilities must be coupled to flexible and accurate computational proteomics software platforms to result in an effective data analysis pipeline. Basic components of such platforms are specialized spectra services as well as advanced machine learning tools, surrounded by orthogonal services for data validation and quality management. Interoperability of basic software tools can be obtained through web services and workflows, while proteomics data repositories and ontologies can allow knowledge sharing and integration, opening new research directions. This special issue has been originated by the Special Track on Computational Proteomics held at the IEEE Symposium on Computer-Based Medical Systems since 2006. It would not have been possible without the encouragement of Tim Clark. A special thanks goes to my friends Gianni Cuda and Marco Gaspari for having introduced me in the world of proteomics. Mario Cannataro |
Briefings Bioinform. | 1 |
| 2008 | SIGMCC: A system for sharing meta patient records in a Peer-to-Peer environment
Mario Cannataro, Domenico Talia, Giuseppe Tradigo, Paolo Trunfio, Pierangelo Veltri |
Future Gener. Comput. Syst. | 1 |
| 2007 | The EIPeptiDi tool: enhancing peptide discovery in ICAT-based LC MS/MS experimentsabstractBACKGROUND: Isotope-coded affinity tags (ICAT) is a method for quantitative proteomics based on differential isotopic labeling, sample digestion and mass spectrometry (MS). The method allows the identification and relative quantification of proteins present in two samples and consists of the following phases. First, cysteine residues are either labeled using the ICAT Light or ICAT Heavy reagent (having identical chemical properties but different masses). Then, after whole sample digestion, the labeled peptides are captured selectively using the biotin tag contained in both ICAT reagents. Finally, the simplified peptide mixture is analyzed by nanoscale liquid chromatography-tandem mass spectrometry (LC-MS/MS). Nevertheless, the ICAT LC-MS/MS method still suffers from insufficient sample-to-sample reproducibility on peptide identification. In particular, the number and the type of peptides identified in different experiments can vary considerably and, thus, the statistical (comparative) analysis of sample sets is very challenging. Low information overlap at the peptide and, consequently, at the protein level, is very detrimental in situations where the number of samples to be analyzed is high. RESULTS: We designed a method for improving the data processing and peptide identification in sample sets subjected to ICAT labeling and LC-MS/MS analysis, based on cross validating MS/MS results. Such a method has been implemented in a tool, called EIPeptiDi, which boosts the ICAT data analysis software improving peptide identification throughout the input data set. Heavy/Light (H/L) pairs quantified but not identified by the MS/MS routine, are assigned to peptide sequences identified in other samples, by using similarity criteria based on chromatographic retention time and Heavy/Light mass attributes. EIPeptiDi significantly improves the number of identified peptides per sample, proving that the proposed method has a considerable impact on the protein identification process and, consequently, on the amount of potentially critical information in clinical studies. The EIPeptiDi tool is available at http://bioingegneria.unicz.it/~veltri/projects/eipeptidi/ with a demo data set. CONCLUSION: EIPeptiDi significantly increases the number of peptides identified and quantified in analyzed samples, thus reducing the number of unassigned H/L pairs and allowing a better comparative analysis of sample data sets. Mario Cannataro, Giovanni Cuda, Marco Gaspari, Sergio Greco, Giuseppe Tradigo, Pierangelo Veltri |
BMC Bioinform. | 1 |
| 2007 | MS-Analyzer: preprocessing and data mining services for proteomics applications on the GridabstractAbstract Mass spectrometry proteomics data contain much information about cell functions and disease conditions. The discovery of such information is enabled by the combined use of novel bioinformatics tools and data mining techniques requiring the integration of huge data sources and the composition of different software tools. The main phases of such emerging applications comprise the loading, management, preprocessing, mining, and visualization of spectra, as well as the analysis of discovered knowledge models. The collection, storage, and analysis of spectra produced in different laboratories can make use of the services of computational Grids, which offer efficient data transfer primitives, effective management of large data stores, and large computing power. In this paper we present MS‐Analyzer, a Grid‐based software platform for the integrated management and analysis of spectra data. MS‐Analyzer provides efficient spectra management through a specialized spectra database, and supports the semantic composition of spectra preprocessing services and data mining services to analyze spectra on the Grid. Copyright © 2006 John Wiley & Sons, Ltd. Mario Cannataro, Pierangelo Veltri |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Using ontologies for preprocessing and mining spectra data on the Grid
Mario Cannataro, Pietro H. Guzzi, Tommaso Mazza, Giuseppe Tradigo, Pierangelo Veltri |
Future Gener. Comput. Syst. | 1 |
| 2007 | Sharing mass spectrometry data in a grid-based distributed proteomics laboratory
Pierangelo Veltri, Mario Cannataro, Giuseppe Tradigo |
Inf. Process. Manag. | 2 |
| 2006 | Analysis and Classification of Proteomics Data, a Case StudyabstractThis paper presents a methodology for analyzing and classifying proteins identified in biological samples. In particular, such methodology consists in normalizing and classifying quantity and quality of proteins identified by using tandem mass spectrometry. A case study is considered and a classification experiment for protein discriminant is also reported Pietro H. Guzzi, Mario Cannataro, Marco Gaspari, Tommaso Mazza, Barbara Quaresima, Pierangelo Veltri, Francesco Saverio Costanzo |
CBMS | 2 |
| 2006 | Next-generation Grids: requirements and knowledge-based servicesabstractAbstract To be effectively adopted in different application domains, next‐generation Grids need to address different issues such as: an increasing complexity and distribution of applications; different goals, skills and habits of Grid users; availability of different programming and deployment models; heterogeneous capabilities and performances of access networks and devices. Moreover, scientific and commercial applications, as well as Grid middleware, will increasingly produce an overwhelming quantity of application and usage data. Although the ongoing convergence between Grids, Web Services, and the Semantic Web constitutes a milestone towards a service‐oriented Grid architecture, which has the potential to face important issues such as application programming and business modelling, many other issues need research and development efforts. The great availability of data and information at the different layers of Grids, the maturity of data exploration techniques able to extract and synthesize knowledge, such as data mining, text summarization, semantic modelling, and knowledge management, and the demand for intelligent services in different phases of application life cycle are the driving forces towards novel knowledge‐based Grid services. Guided by those considerations, the paper first introduces main requirements of next‐generation Grids and then describes some representative knowledge‐based Grid services for both applications support and system management. Simple cases study showing how such services could be employed are discussed. Copyright © 2005 John Wiley & Sons, Ltd. Mario Cannataro |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | Preprocessing of Mass Spectrometry Proteomics Data on the GridabstractThe combined use of mass spectrometry and data mining is a novel approach in proteomic pattern analysis for discovering novel biomarkers or identifying patterns and associations in proteomic profiles. Data produced by mass spectrometers are affected by errors and noise due to sample preparation and instrument approximation, so different preprocessing techniques need to be applied before analysis is conducted. We survey different techniques for spectra preprocessing, and we present a first design of a software tool that allows the preprocessing, management and analysis of mass spectrometry data on the Grid. Mario Cannataro, Pietro H. Guzzi, Tommaso Mazza, Giuseppe Tradigo, Pierangelo Veltri |
CBMS | 1 |
| 2004 | Distributed data mining on grids: services, tools, and applicationsabstractData mining algorithms are widely used today for the analysis of large corporate and scientific datasets stored in databases and data archives. Industry, science, and commerce fields often need to analyze very large datasets maintained over geographically distributed sites by using the computational power of distributed and parallel systems. The grid can play a significant role in providing an effective computational support for distributed knowledge discovery applications. For the development of data mining applications on grids we designed a system called Knowledge Grid. This paper describes the Knowledge Grid framework and presents the toolset provided by the Knowledge Grid for implementing distributed knowledge discovery. The paper discusses how to design and implement data mining applications by using the Knowledge Grid tools starting from searching grid resources, composing software and data components, and executing the resulting data mining process on a grid. Some performance results are also discussed. Mario Cannataro, Antonio Congiusta, Andrea Pugliese 0001, Domenico Talia, Paolo Trunfio |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2002 | XAHM: an adaptive hypermedia model based on XMLabstractThis paper presents XAHM, a model for Adaptive Hypermedia Systems based on XML. We introduce a multidimensional approach to model different aspects of the process, which is based on three different adaptivity dimensions: user's behavior (preferences and browsing activity), technology (network and user's terminal) and external environment (time, location, language, socio-political issues, etc.). An Adaptive Hypermedia is modeled with respect to such dimensions, and a view over it corresponds to each potential position of the user in the adaptation space. Finally, we propose a probabilistic algorithm for the classification of users. Mario Cannataro, Alfredo Cuzzocrea, Andrea Pugliese 0001 |
SEKE | 1 |
| 2002 | Distributed data mining on the grid
Mario Cannataro, Domenico Talia, Paolo Trunfio |
Future Gener. Comput. Syst. | 1 |
| 2002 | Parallel data intensive computing in scientific and commercial applications
Mario Cannataro, Domenico Talia, Pradip K. Srimani |
Parallel Comput. | 1 |
| 2000 | An XML-Based Architecture for Adaptive Web Hypermedia Systems Using a Probabilistic User ModelabstractWeb based hypermedia systems are becoming increasingly popular as tools for user driven access to information and services. The paper presents an architecture for the development of Web based adaptive hypermedia systems. The architecture uses weighted graphs of XML documents to describe the application domain and a probabilistic model to adapt the Web site content generation and presentation to the user's behaviour. The user's behaviour is modelled using a probabilistic model and the most promising profile, that is a "view" over the application domain, is dynamically assigned to the user, using a discrete probability density function. Mario Cannataro, Andrea Pugliese 0001 |
IDEAS | 1 |
| 1995 | A Parallel Cellular Automata Environment on Multicomputers for Computational Science
Mario Cannataro, Salvatore Di Gregorio, Rocco Rongo, William Spataro, Giandomenico Spezzano, Domenico Talia |
Parallel Comput. | 1 |
| 1992 | Design, implementation and evaluation of a deadlock-free routing algorithm for concurrent computersabstractAbstract This paper describes the design, the implementation, and the performance results of a routing algorithm which provides deadlock‐free communication in a tightly coupled message‐passing concurrent computer. The algorithm is adaptive, isolated and uses the store‐and‐forward technique. It allows message communication between two processes regardless of where they are physically located on the network. The routing algorithm has many positive characteristics including provable deadlock freedom, guaranteed message arrival, and automatic local congestion reduction. It can be used as a basis for the design of high‐level communication primitives. An Occam implementation on a network of inmos Transputers is discussed. The experimental results show that the routing algorithm is effective to support process to process communication on a concurrent computer. Mario Cannataro, Giandomenico Spezzano, Domenico Talia, E. Gallizzi |
Concurr. Pract. Exp. | 1 |
| 1992 | High level communication mechanisms for distributed parallel computers from an adaptive message routing
Mario Cannataro, Giandomenico Spezzano, Domenico Talia |
Future Gener. Comput. Syst. | 1 |
| 1992 | A model of efficient asynchronous parallel algorithms on multicomputer systems
Domenico Conforti, Lucio Grandinetti, Roberto Musmanno, Mario Cannataro, Giandomenico Spezzano, Domenico Talia |
Parallel Comput. | 4 |
| 1991 | A parallel logic system on a multicomputer architecture
Mario Cannataro, Giandomenico Spezzano, Domenico Talia |
Future Gener. Comput. Syst. | 1 |