VLDB 2026 Research / reviewers in the wild / expert
Gianvito Pio
dblp:118/2606
· DBLP profile ↗
35ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0003-2520-3616ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Similarity-Guided Forecasting of Multiple Sequences from Mechanical Testing Data of Additively Manufactured Components
Antonio Pellicani, Gianvito Pio, Donato Malerba, Michelangelo Ceci |
ISMIS | 2 |
| 2026 | Dynamic instance weighting for online learning in multi-cryptocurrency price and trend forecastingabstractAbstract The cryptocurrency market represents a significant innovation in the financial ecosystem, built upon cryptographic principles to ensure secure and transparent transactions. Cryptocurrencies experienced a global adoption, driven by their decentralized nature that enables borderless transactions without third-party intermediaries. The price of cryptocurrencies is characterized by a significant volatility, that introduces both opportunities and challenges. In this context, the development of accurate methods for the forecasting of price variation, able to work in real-time on data streams, has become vital for various stakeholders. In this paper, we propose a novel approach, called LEMON, for the online prediction of the price variation of cryptocurrencies, that leverages possible temporal correlations among them. Our approach stems from the empirical evidence that cryptocurrencies tend to form groups characterized by similar trends, a behavior often attributed to shared market dynamics and common external factors. Through the analysis of temporal correlations, LEMON dynamically identifies these groups, that are then exploited to learn multiple multi-target tree-based models, specifically designed for processing continuous data streams. LEMON also introduces a novel adaptive non-parametric weighting scheme, that automatically adjusts the importance of each instance based on the observed data distribution in real-time, improving the forecasting of the price variation. Our experiments, performed on 16 datasets related to 16 cryptocurrencies, demonstrate that LEMON outperforms state-of-the-art approaches in two distinct prediction tasks: forecasting the closing price variation (regression) and predicting the market trend direction (classification), making it an effective tool to support stakeholders requiring accurate real-time predictions. Antonio Pellicani, Gianvito Pio, Saso Dzeroski, Michelangelo Ceci |
Data Min. Knowl. Discov. | 2 |
| 2026 | Handling complex backgrounds and light perturbations for enhancing learning tasks from images of vegetablesabstractAbstract The quality assessment of fruits and vegetables is crucial in the agroalimentary supply chain, as it directly affects consumer satisfaction, market value and overall food security. Traditional approaches rely on visual inspections or destructive techniques, which are labor-intensive and time-consuming. On the contrary, non-destructive techniques emerged as promising alternatives, offering solutions that can be adopted in real environments. Previous studies emphasized that the color distribution over images plays a significant role in the quality evaluation of food. In this paper, we propose a solution that leverages an autoencoder architecture to extract groups of relevant colors from the complete histogram of colors. To enhance the analysis of real-world images with complex backgrounds, we employ a pre-trained U2-Net architecture for background removal. Moreover, we propose a novel procedure based on outlier detection to identify and remove parts of the background that are not fully eliminated, especially along the edges of the product. After this preprocessing, we extract a complete color histogram which is fed to an autoencoder architecture, to extract high-level features representing color groups at different levels of granularity. The goal is to make the learned models less sensitive to light and color perturbations. Our experiments, conducted on two real-world datasets related to two different learning tasks, demonstrated the effectiveness of the proposed solution, that outperformed several baseline and state-of-the-art approaches, also based on complex neural network architectures. Stefano Polimena, Gianvito Pio, Giovanni Attolico, Michelangelo Ceci |
J. Intell. Inf. Syst. | 2 |
| 2025 | A Novel AI Approach for the Diagnosis of Alzheimer's Disease from Multi-modal Incomplete Data
Veronica Buttaro, Giuseppe Lamanna, Donato Massaro, Claudio Basilio Caporusso, Gianvito Pio, Michelangelo Ceci |
DS | 5 |
| 2025 | CARROT: Simultaneous prediction of anomalies from groups of correlated cryptocurrency trendsabstractCryptocurrencies are virtual currencies that exploit cryptography to perform secure financial transactions. They gained widespread popularity in recent years due to their decentralized nature, (pseudo-)anonymity, and ability to facilitate cross-border transactions without the need for intermediaries. However, their price on the market exhibits a huge volatility, that makes them prone to market anomalies. Therefore, predicting anomalies in cryptocurrency time series can be considered an important task for financial institutions, traders, and investors, to maximize their profit or minimize losses. In this paper, we propose a novel approach for predicting anomalies in cryptocurrency time series by exploiting temporal correlations among different cryptocurrencies. Our approach, called CARROT, is based on the idea that groups of cryptocurrencies exhibit similar trends, possibly due to common influencing factors. CARROT analyzes the temporal correlation between different cryptocurrencies, and identifies clusters showing similar patterns that can be useful for gaining insights into future anomalies. Subsequently, CARROT exploits multiple (i.e., one for each cluster) multi-target LSTM models to predict anomalies. Our experiments, performed on a dataset of 17 cryptocurrencies, proved that CARROT outperforms single-target LSTM models of up to 20%, as well as other approaches based on neural networks, i.e., MLP and CNN, in terms of macro F1-score. Therefore, the proposed approach can be considered as a promising tool for predicting anomalies in cryptocurrency time series data and can potentially be used to improve risk management and trading strategies in the cryptocurrency market. • Analysis of cryptocurrency trends. • Clustering-based multi-target prediction of anomalies in time series. • Consistent improvements achieved over the single-target counterpart. Antonio Pellicani, Gianvito Pio, Michelangelo Ceci |
Expert Syst. Appl. | 2 |
| 2025 | Modeling spatio-temporal locality in multi-step forecasting of geo-referenced time seriesabstractAbstract Forecasting future measurements from geographically distributed sensors is essential across many application domains. However, the spatial distribution of these sensors raises multiple challenges, primarily due to spatial autocorrelation phenomena, that introduce inter-dependencies among nearby locations, that cannot therefore be treated independently by learning algorithms. While some existing approaches can capture such phenomena, they generally model the spatial dimension globally across all locations. On the other hand, the method we propose in this paper, called SPALT, focuses on capturing spatial relationships specifically among time series with similar trends, even if these trends occur at different times, thus modeling the spatio-temporal locality. SPALT leverages linear model trees, which allow us to naturally consider the spatial autocorrelation in a local manner: during the tree-building process, the adopted heuristics aim to group time series exhibiting similar trends into the same node, on which additional features considering the spatial dimension are selectively injected. Additionally, we propose a new pruning strategy, based on Reduced Error Pruning (REP), that also considers the spatio-temporal locality during the tree simplification. Designed for a multi-step setting, SPALT provides forecasts for multiple future time steps across multiple sensors simultaneously. The characteristics exhibited by SPALT can provide significant benefits in different domains, where measurements come from geographically distributed sensors. In this paper, we focus on data produced by sensors located in multiple renewable power plants measuring their energy production at regular, short intervals. Experiments on three real-world datasets demonstrate the effectiveness of SPALT in forecasting the production of energy at different time horizons, and its superior performance in comparison with tree-based models and state-of-the-art neural networks that incorporate both temporal and spatial dimensions. Annunziata D'Aversa, Gianvito Pio, Michelangelo Ceci |
Mach. Learn. | 2 |
| 2024 | Leveraging Spatio-Temporal Locality in Linear Model Trees for Multi-Step Time Series ForecastingabstractIn the era of Big Data, forecasting the future measurements of geo-distributed sensors can be considered among the most fundamental tasks for several application domains. However, the spatial distribution of such sensors introduces several challenges for forecasting methods. The most important one is due to the spatial dimension which introduces autocorrelation phenomena, according to which nearby locations are not independent and should not be treated as such by the learning algorithms. Some existing approaches are able to capture the spatial autocorrelation, but they tend to model spatial information globally across all locations. On the contrary, our method, called SPLiT, focuses on capturing spatial information only among time series with similar trends, also at different timing, modeling the so-called spatio-temporal locality. The proposed method is based on linear model trees, which allow us to naturally model any form of discontinuities in the spatial autocorrelation during the tree growing phase, guided by heuristics that identify time series with similar trends. SPLiT works in the multi-step predictive setting in order to simultaneously provide forecasts for multiple time steps for multiple sensors in the future.Our experiments, conducted on two real-world datasets, show the effectiveness of the proposed method in forecasting the energy production of geographically distributed renewable power plants. The comparison against several tree-based models and state-of-the-art neural networks that consider both temporal and spatial dimensions shows the superiority of the proposed method. Annunziata D'Aversa, Gianvito Pio, Michelangelo Ceci |
IEEE Big Data | 2 |
| 2024 | Improving the Robustness to Color Perturbations of Classification and Regression Models in the Visual Evaluation of Fruits and Vegetables
Stefano Polimena, Gianvito Pio, Giovanni Attolico, Michelangelo Ceci |
ISMIS | 2 |
| 2024 | Exploiting microRNA Expression Data for the Diagnosis of Disease Conditions and the Discovery of Novel Biomarkers
Daniele Rosa, Antonio Pellicani, Gianvito Pio, Domenica D'Elia, Michelangelo Ceci |
ISMIS | 3 |
| 2024 | Multi-class boosting for the analysis of multiple incomplete views on microbiome dataabstractBACKGROUND: Microbiome dysbiosis has recently been associated with different diseases and disorders. In this context, machine learning (ML) approaches can be useful either to identify new patterns or learn predictive models. However, data to be fed to ML methods can be subject to different sampling, sequencing and preprocessing techniques. Each different choice in the pipeline can lead to a different view (i.e., feature set) of the same individuals, that classical (single-view) ML approaches may fail to simultaneously consider. Moreover, some views may be incomplete, i.e., some individuals may be missing in some views, possibly due to the absence of some measurements or to the fact that some features are not available/applicable for all the individuals. Multi-view learning methods can represent a possible solution to consider multiple feature sets for the same individuals, but most existing multi-view learning methods are limited to binary classification tasks or cannot work with incomplete views. RESULTS: We propose irBoost.SH, an extension of the multi-view boosting algorithm rBoost.SH, based on multi-armed bandits. irBoost.SH solves multi-class classification tasks and can analyze incomplete views. At each iteration, it identifies one winning view using adversarial multi-armed bandits and uses its predictions to update a shared instance weight distribution in a learning process based on boosting. In our experiments, performed on 5 multi-view microbiome datasets, the model learned by irBoost.SH always outperforms the best model learned from a single view, its closest competitor rBoost.SH, and the model learned by a multi-view approach based on feature concatenation, reaching an improvement of 11.8% of the F1-score in the prediction of the Autism Spectrum disorder and of 114% in the prediction of the Colorectal Cancer disease. CONCLUSIONS: The proposed method irBoost.SH exhibited outstanding performances in our experiments, also compared to competitor approaches. The obtained results confirm that irBoost.SH can fruitfully be adopted for the analysis of microbiome data, due to its capability to simultaneously exploit multiple feature sets obtained through different sequencing and preprocessing pipelines. Andrea Simeon, Milos Radovanovic 0001, Tatjana Loncar-Turukalo, Michelangelo Ceci, Sanja Brdar, Gianvito Pio |
BMC Bioinform. | 6 |
| 2023 | On the exploitation of the blockchain technology in the healthcare sector: A systematic review
Valeria Merlo, Gianvito Pio, Francesco Giusto, Massimo Bilancia |
Expert Syst. Appl. | 2 |
| 2023 | Multi-view overlapping clustering for the identification of the subject matter of legal judgments
Graziella De Martino, Gianvito Pio, Michelangelo Ceci |
Inf. Sci. | 2 |
| 2023 | HURI: Hybrid user risk identification in social networksabstractAbstract The massive adoption of social networks increased the need to analyze users’ data and interactions to detect and block the spread of propaganda and harassment behaviors, as well as to prevent actions influencing people towards illegal or immoral activities. In this paper, we propose HURI, a method for social network analysis that accurately classifies users assafeorrisky, according to their behavior in the social network. Specifically, the proposed hybrid approach leverages both the topology of the network of interactions and the semantics of the content shared by users, leading to an accurate classification also in the presence of noisy data, such as users who may appear to be risky due to the topic of their posts, but are actually safe according to their relationships. The strength of the proposed approach relies on the full and simultaneous exploitation of both aspects, giving each of them equal consideration during the combination phase. This characteristic makes HURI different from other approaches that fully consider only a single aspect and graft partial or superficial elements of the other into the first. The achieved performance in the analysis of a real-world Twitter dataset shows that the proposed method offers competitive performance with respect to eight state-of-the-art approaches. Roberto Corizzo, Gianvito Pio, Emanuele Pio Barracchia, Antonio Pellicani, Nathalie Japkowicz, Michelangelo Ceci |
World Wide Web (WWW) | 2 |
| 2022 | Distributed Heterogeneous Transfer Learning for Link Prediction in the Positive Unlabeled SettingabstractTransfer learning focuses on enhancing predictive models for a target domain, by exploiting the knowledge coming from a related source domain. However, most existing transfer learning methods assume that source and target domains are described with the same feature spaces. Heterogeneous transfer learning approaches aim to overcome this limitation, but they usually introduce strong assumptions (e.g., on the number of features), cannot distribute the workload to handle large volumes of data, or cannot work in challenging settings like the Positive-Unlabeled (PU) setting, where only positive and unlabelled examples are available. In this paper, we present a novel heterogeneous distributed transfer learning method that can work also in PU learning setting and overcomes all such limitations.The experimental evaluation was conducted in the context of a link prediction task in the biological domain. The results showed the effectiveness of the proposed method, that outperformed three state-of-the-art heterogeneous transfer learning approaches. Paolo Mignone, Gianvito Pio, Michelangelo Ceci |
IEEE Big Data | 2 |
| 2022 | Leveraging Spatio-Temporal Autocorrelation to Improve the Forecasting of the Energy Consumption in Smart Grids
Annunziata D'Aversa, Stefano Polimena, Gianvito Pio, Michelangelo Ceci |
DS | 3 |
| 2022 | Identification of Paragraph Regularities in Legal Judgements Through Clustering and Textual Embedding
Graziella De Martino, Gianvito Pio |
ISMIS | 2 |
| 2022 | Integrating genome-scale metabolic modelling and transfer learning for human gene regulatory network reconstructionabstractMOTIVATION: Gene regulation is responsible for controlling numerous physiological functions and dynamically responding to environmental fluctuations. Reconstructing the human network of gene regulatory interactions is thus paramount to understanding the cell functional organization across cell types, as well as to elucidating pathogenic processes and identifying molecular drug targets. Although significant effort has been devoted towards this direction, existing computational methods mainly rely on gene expression levels, possibly ignoring the information conveyed by mechanistic biochemical knowledge. Moreover, except for a few recent attempts, most of the existing approaches only consider the information of the organism under analysis, without exploiting the information of related model organisms. RESULTS: We propose a novel method for the reconstruction of the human gene regulatory network, based on a transfer learning strategy that synergically exploits information from human and mouse, conveyed by gene-related metabolic features generated in silico from gene expression data. Specifically, we learn a predictive model from metabolic activity inferred via tissue-specific metabolic modelling of artificial gene knockouts. Our experiments show that the combination of our transfer learning approach with the constructed metabolic features provides a significant advantage in terms of reconstruction accuracy, as well as additional clues on the contribution of each constructed metabolic feature. AVAILABILITY AND IMPLEMENTATION: The method, the datasets and all the results obtained in this study are available at: https://doi.org/10.6084/m9.figshare.c.5237687. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gianvito Pio, Paolo Mignone, Giuseppe Magazzù, Guido Zampieri, Michelangelo Ceci, Claudio Angione |
Bioinform. | 1 |
| 2022 | LP-ROBIN: Link prediction in dynamic networks exploiting incremental node embedding
Emanuele Pio Barracchia, Gianvito Pio, Albert Bifet, Heitor Murilo Gomes, Bernhard Pfahringer, Michelangelo Ceci |
Inf. Sci. | 2 |
| 2022 | Relational tree ensembles and feature rankingsabstractAs the complexity of data increases, so does the importance of powerful representations, such as relational and logical representations, as well as the need for machine learning methods that can learn predictive models in such representations. A characteristic of these representations is that they give rise to a huge number of features to be considered, thus drastically increasing the difficulty of learning in terms of computational complexity and the curse of dimensionality. Despite this, methods for ranking features in this context, i.e., estimating their importance are practically non-existent. Among the most well-known methods for feature ranking are those based on ensembles, and in particular tree ensembles. To develop methods for feature ranking in a relational context, we adopt the relational tree ensemble approach. We thus first develop methods for learning ensembles of relational trees, extending a wide spectrum of tree-based ensemble methods from the propositional to the relational context, resulting in methods for bagging and random forests of relational trees, as well as gradient boosted ensembles thereof. Complex relational features are considered in our ensembles: by using complex aggregates, we extend the standard collection of features that correspond to existential queries, such as ‘Does this person have any children?’, to more complex features that correspond to aggregation queries, such as ‘What is the average age of this person’s children?’. We also calculate feature importance scores and rankings from the different kinds of relational tree ensembles learned, with different kinds of relational features. The rankings provide insight into and explain the ensemble models, which would be otherwise difficult to understand. We compare the methods for learning single trees and different tree ensembles, using only existential qualifiers and using the whole set of relational features, against 10 state-of-the-art methods on a collection of benchmark relational datasets, deriving also the corresponding feature rankings. Overall, the bagging ensembles perform the best, with gradient boosted ensembles following closely. The use of aggregates is beneficial and in some datasets drastically improves performance: In these cases, aggregate-based features clearly stand out in the feature rankings derived from the ensembles. Matej Petkovic, Michelangelo Ceci, Gianvito Pio, Blaz Skrlj, Kristian Kersting, Saso Dzeroski |
Knowl. Based Syst. | 3 |
| 2021 | Spatially-Aware Autoencoders for Detecting Contextual Anomalies in Geo-Distributed Data
Roberto Corizzo, Michelangelo Ceci, Gianvito Pio, Paolo Mignone, Nathalie Japkowicz |
DS | 3 |
| 2021 | BROCCOLI: overlapping and outlier-robust biclustering through proximal stochastic gradient descentabstractAbstract Matrix tri-factorization subject to binary constraints is a versatile and powerful framework for the simultaneous clustering of observations and features, also known as biclustering. Applications for biclustering encompass the clustering of high-dimensional data and explorative data mining, where the selection of the most important features is relevant. Unfortunately, due to the lack of suitable methods for the optimization subject to binary constraints, the powerful framework of biclustering is typically constrained to clusterings which partition the set of observations or features. As a result, overlap between clusters cannot be modelled and every item, even outliers in the data, have to be assigned to exactly one cluster. In this paper we propose Broccoli , an optimization scheme for matrix factorization subject to binary constraints, which is based on the theoretically well-founded optimization scheme of proximal stochastic gradient descent. Thereby, we do not impose any restrictions on the obtained clusters. Our experimental evaluation, performed on both synthetic and real-world data, and against 6 competitor algorithms, show reliable and competitive performance, even in presence of a high amount of noise in the data. Moreover, a qualitative analysis of the identified clusters shows that Broccoli may provide meaningful and interpretable clustering structures. Sibylle Hess, Gianvito Pio, Michiel E. Hochstenbach, Michelangelo Ceci |
Data Min. Knowl. Discov. | 2 |
| 2020 | Exploiting transfer learning for the reconstruction of the human gene regulatory networkabstractMOTIVATION: The reconstruction of gene regulatory networks (GRNs) from gene expression data has received increasing attention in recent years, due to its usefulness in the understanding of regulatory mechanisms involved in human diseases. Most of the existing methods reconstruct the network through machine learning approaches, by analyzing known examples of interactions. However, (i) they often produce poor results when the amount of labeled examples is limited, or when no negative example is available and (ii) they are not able to exploit information extracted from GRNs of other (better studied) related organisms, when this information is available. RESULTS: In this paper, we propose a novel machine learning method that overcomes these limitations, by exploiting the knowledge about the GRN of a source organism for the reconstruction of the GRN of the target organism, by means of a novel transfer learning technique. Moreover, the proposed method is natively able to work in the positive-unlabeled setting, where no negative example is available, by fruitfully exploiting a (possibly large) set of unlabeled examples. In our experiments, we reconstructed the human GRN, by exploiting the knowledge of the GRN of Mus musculus. Results showed that the proposed method outperforms state-of-the-art approaches and identifies previously unknown functional relationships among the analyzed genes. AVAILABILITY AND IMPLEMENTATION: http://www.di.uniba.it/∼mignone/systems/biosfer/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Paolo Mignone, Gianvito Pio, Domenica D'Elia, Michelangelo Ceci |
Bioinform. | 2 |
| 2020 | Prediction of new associations between ncRNAs and diseases exploiting multi-type hierarchical clusteringabstractBACKGROUND: The study of functional associations between ncRNAs and human diseases is a pivotal task of modern research to develop new and more effective therapeutic approaches. Nevertheless, it is not a trivial task since it involves entities of different types, such as microRNAs, lncRNAs or target genes whose expression also depends on endogenous or exogenous factors. Such a complexity can be faced by representing the involved biological entities and their relationships as a network and by exploiting network-based computational approaches able to identify new associations. However, existing methods are limited to homogeneous networks (i.e., consisting of only one type of objects and relationships) or can exploit only a small subset of the features of biological entities, such as the presence of a particular binding domain, enzymatic properties or their involvement in specific diseases. RESULTS: To overcome the limitations of existing approaches, we propose the system LP-HCLUS, which exploits a multi-type hierarchical clustering method to predict possibly unknown ncRNA-disease relationships. In particular, LP-HCLUS analyzes heterogeneous networks consisting of several types of objects and relationships, each possibly described by a set of features, and extracts multi-type clusters that are subsequently exploited to predict new ncRNA-disease associations. The extracted clusters are overlapping, hierarchically organized, involve entities of different types, and allow LP-HCLUS to catch multiple roles of ncRNAs in diseases at different levels of granularity. Our experimental evaluation, performed on heterogeneous attributed networks consisting of microRNAs, lncRNAs, diseases, genes and their known relationships, shows that LP-HCLUS is able to obtain better results with respect to existing approaches. The biological relevance of the obtained results was evaluated according to both quantitative (i.e., TPR@k, Areas Under the TPR@k, ROC and Precision-Recall curves) and qualitative (i.e., according to the consultation of the existing literature) criteria. CONCLUSIONS: The obtained results prove the utility of LP-HCLUS to conduct robust predictive studies on the biological role of ncRNAs in human diseases. The produced predictions can therefore be reliably considered as new, previously unknown, relationships among ncRNAs and diseases. Emanuele Pio Barracchia, Gianvito Pio, Domenica D'Elia, Michelangelo Ceci |
BMC Bioinform. | 2 |
| 2020 | Exploiting causality in gene network reconstruction based on graph embedding
Gianvito Pio, Michelangelo Ceci, Francesca Prisciandaro, Donato Malerba |
Mach. Learn. | 1 |
| 2018 | Positive Unlabeled Link Prediction via Transfer Learning for Gene Network Reconstruction
Paolo Mignone, Gianvito Pio |
ISMIS | 2 |
| 2018 | Multi-type clustering and classification from heterogeneous networks
Gianvito Pio, Francesco Serafino 0002, Donato Malerba, Michelangelo Ceci |
Inf. Sci. | 1 |
| 2018 | Ensemble Learning for Multi-Type Classification in Heterogeneous NetworksabstractHeterogeneous networks are networks consisting of different types of objects and links. They can be found in several fields, ranging from the Internet to social sciences, biology, epidemiology, geography, finance, and many others. In the literature, several methods have been proposed for the analysis of network data, but they usually focus on homogeneous networks, where all the objects are of the same type, and links among them describe a single type of relationship. More recently, the complexity of real scenarios has impelled researchers to design methods for the analysis of heterogeneous networks, especially focused on classification and clustering tasks. However, they often make assumptions on the structure of the network that are too restrictive or do not fully exploit different forms of network correlation and autocorrelation. Moreover, when nodes which are the main subject of the classification task are linked to several nodes of the network having missing values, standard methods can lead to either building incomplete classification models or to discarding possibly relevant dependencies (correlation or autocorrelation). In this paper, we propose an ensemble learning approach for multi-type classification. We adopt the system Mr-SBC, which is originally able to analyze heterogeneous networks of arbitrary structure, within an ensemble learning approach. The ensemble allows us to improve the classification accuracy of Mr-SBC by exploiting i) the possible presence of correlation and autocorrelation phenomena and ii) the classification of instances (which contain missing values) of other node types in the network. As a beneficial side effect, we have also that the models are more stable in terms of standard deviation of the accuracy, over different samples used for training. Experiments performed on real-world datasets show that the proposed method is able to significantly outperform the standard implementation of Mr-SBC. Moreover, it gives Mr-SBC the advantage of outperforming four other well-known algorithms for the classification of data organized in a network. Francesco Serafino 0002, Gianvito Pio, Michelangelo Ceci |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | LOCANDA: Exploiting Causality in the Reconstruction of Gene Regulatory Networks
Gianvito Pio, Michelangelo Ceci, Francesca Prisciandaro, Donato Malerba |
DS | 1 |
| 2015 | Hierarchical Multidimensional Classification of Web Documents with MultiWebClass
Francesco Serafino 0002, Gianvito Pio, Michelangelo Ceci, Donato Malerba |
Discovery Science | 2 |
| 2015 | ComiRNet: a web-based system for the analysis of miRNA-gene regulatory networksabstractBACKGROUND: The understanding of mechanisms and functions of microRNAs (miRNAs) is fundamental for the study of many biological processes and for the elucidation of the pathogenesis of many human diseases. Technological advances represented by high-throughput technologies, such as microarray and next-generation sequencing, have significantly aided miRNA research in the last decade. Nevertheless, the identification of true miRNA targets and the complete elucidation of the rules governing their functional targeting remain nebulous. Computational tools have been proven to be fundamental for guiding experimental validations for the discovery of new miRNAs, for the identification of their targets and for the elucidation of their regulatory mechanisms. DESCRIPTION: ComiRNet (Co-clustered miRNA Regulatory Networks) is a web-based database specifically designed to provide biologists and clinicians with user-friendly and effective tools for the study of miRNA-gene target interaction data and for the discovery of miRNA functions and mechanisms. Data in ComiRNet are produced by a combined computational approach based on: 1) a semi-supervised ensemble-based classifier, which learns to combine miRNA-gene target interactions (MTIs) from several prediction algorithms, and 2) the biclustering algorithm HOCCLUS2, which exploits the large set of produced predictions, with the associated probabilities, to identify overlapping and hierarchically organized biclusters that represent miRNA-gene regulatory networks (MGRNs). CONCLUSIONS: ComiRNet represents a valuable resource for elucidating the miRNAs' role in complex biological processes by exploiting data on their putative function in the context of MGRNs. ComiRnet currently stores about 5 million predicted MTIs between 934 human miRNAs and 30,875 mRNAs, as well as 15 bicluster hierarchies, each of which represents MGRNs at different levels of granularity. The database can be freely accessed at: http://comirnet.di.uniba.it. Gianvito Pio, Michelangelo Ceci, Donato Malerba, Domenica D'Elia |
BMC Bioinform. | 1 |
| 2015 | Non-negative Matrix Tri-Factorization for co-clustering: An analysis of the block matrix
Nicoletta Del Buono, Gianvito Pio |
Inf. Sci. | 2 |
| 2014 | Mining Temporal Evolution of Entities in a Stream of Textual Documents
Gianvito Pio, Pasqua Fabiana Lanotte, Michelangelo Ceci, Donato Malerba |
ISMIS | 1 |
| 2014 | Network Reconstruction for the Identification of miRNA: mRNA Interaction Networks
Gianvito Pio, Michelangelo Ceci, Domenica D'Elia, Donato Malerba |
ECML/PKDD (3) | 1 |
| 2014 | Integrating microRNA target predictions for the discovery of gene regulatory networks: a semi-supervised ensemble learning approachabstractBACKGROUND: MicroRNAs (miRNAs) are small non-coding RNAs which play a key role in the post-transcriptional regulation of many genes. Elucidating miRNA-regulated gene networks is crucial for the understanding of mechanisms and functions of miRNAs in many biological processes, such as cell proliferation, development, differentiation and cell homeostasis, as well as in many types of human tumors. To this aim, we have recently presented the biclustering method HOCCLUS2, for the discovery of miRNA regulatory networks. Experiments on predicted interactions revealed that the statistical and biological consistency of the obtained networks is negatively affected by the poor reliability of the output of miRNA target prediction algorithms. Recently, some learning approaches have been proposed to learn to combine the outputs of distinct prediction algorithms and improve their accuracy. However, the application of classical supervised learning algorithms presents two challenges: i) the presence of only positive examples in datasets of experimentally verified interactions and ii) unbalanced number of labeled and unlabeled examples. RESULTS: We present a learning algorithm that learns to combine the score returned by several prediction algorithms, by exploiting information conveyed by (only positively labeled/) validated and unlabeled examples of interactions. To face the two related challenges, we resort to a semi-supervised ensemble learning setting. Results obtained using miRTarBase as the set of labeled (positive) interactions and mirDIP as the set of unlabeled interactions show a significant improvement, over competitive approaches, in the quality of the predictions. This solution also improves the effectiveness of HOCCLUS2 in discovering biologically realistic miRNA:mRNA regulatory networks from large-scale prediction data. Using the miR-17-92 gene cluster family as a reference system and comparing results with previous experiments, we find a large increase in the number of significantly enriched biclusters in pathways, consistent with miR-17-92 functions. CONCLUSION: The proposed approach proves to be fundamental for the computational discovery of miRNA regulatory networks from large-scale predictions. This paves the way to the systematic application of HOCCLUS2 for a comprehensive reconstruction of all the possible multiple interactions established by miRNAs in regulating the expression of gene networks, which would be otherwise impossible to reconstruct by considering only experimentally validated interactions. Gianvito Pio, Donato Malerba, Domenica D'Elia, Michelangelo Ceci |
BMC Bioinform. | 1 |
| 2013 | A Novel Biclustering Algorithm for the Discovery of Meaningful Biological Correlations between microRNAs and their Target GenesabstractBACKGROUND: microRNAs (miRNAs) are a class of small non-coding RNAs which have been recognized as ubiquitous post-transcriptional regulators. The analysis of interactions between different miRNAs and their target genes is necessary for the understanding of miRNAs' role in the control of cell life and death. In this paper we propose a novel data mining algorithm, called HOCCLUS2, specifically designed to bicluster miRNAs and target messenger RNAs (mRNAs) on the basis of their experimentally-verified and/or predicted interactions. Indeed, existing biclustering approaches, typically used to analyze gene expression data, fail when applied to miRNA:mRNA interactions since they usually do not extract possibly overlapping biclusters (miRNAs and their target genes may have multiple roles), extract a huge amount of biclusters (difficult to browse and rank on the basis of their importance) and work on similarities of feature values (do not limit the analysis to reliable interactions). RESULTS: To overcome these limitations, HOCCLUS2 i) extracts possibly overlapping biclusters, to catch multiple roles of both miRNAs and their target genes; ii) extracts hierarchically organized biclusters, to facilitate bicluster browsing and to distinguish between universe and pathway-specific miRNAs; iii) extracts highly cohesive biclusters, to consider only reliable interactions; iv) ranks biclusters according to the functional similarities, computed on the basis of Gene Ontology, to facilitate bicluster analysis. CONCLUSIONS: Our results show that HOCCLUS2 is a valid tool to support biologists in the identification of context-specific miRNAs regulatory modules and in the detection of possibly unknown miRNAs target genes. Indeed, results prove that HOCCLUS2 is able to extract cohesiveness-preserving biclusters, when compared with competitive approaches, and statistically confirm (at a confidence level of 99%) that mRNAs which belong to the same biclusters are, on average, more functionally similar than mRNAs which belong to different biclusters. Finally, the hierarchy of biclusters provides useful insights to understand the intrinsic hierarchical organization of miRNAs and their potential multiple interactions on target genes. Gianvito Pio, Michelangelo Ceci, Domenica D'Elia, Corrado Loglisci, Donato Malerba |
BMC Bioinform. | 1 |