Gianvito Pio

dblp:118/2606 · DBLP profile ↗
← Back
11ranked-venue papers in the field
2as first author
7since 2021 · last 2026
0000-0003-2520-3616ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 5 (1 first)Data Mining & Knowledge Discovery · 3 (1 first)Big Data, Cloud & Distributed Data Systems · 2Database Systems & Data Management · 1
YearPublicationVenuePosition
2026 Dynamic instance weighting for online learning in multi-cryptocurrency price and trend forecasting
abstract
Abstract The cryptocurrency market represents a significant innovation in the financial ecosystem, built upon cryptographic principles to ensure secure and transparent transactions. Cryptocurrencies experienced a global adoption, driven by their decentralized nature that enables borderless transactions without third-party intermediaries. The price of cryptocurrencies is characterized by a significant volatility, that introduces both opportunities and challenges. In this context, the development of accurate methods for the forecasting of price variation, able to work in real-time on data streams, has become vital for various stakeholders. In this paper, we propose a novel approach, called LEMON, for the online prediction of the price variation of cryptocurrencies, that leverages possible temporal correlations among them. Our approach stems from the empirical evidence that cryptocurrencies tend to form groups characterized by similar trends, a behavior often attributed to shared market dynamics and common external factors. Through the analysis of temporal correlations, LEMON dynamically identifies these groups, that are then exploited to learn multiple multi-target tree-based models, specifically designed for processing continuous data streams. LEMON also introduces a novel adaptive non-parametric weighting scheme, that automatically adjusts the importance of each instance based on the observed data distribution in real-time, improving the forecasting of the price variation. Our experiments, performed on 16 datasets related to 16 cryptocurrencies, demonstrate that LEMON outperforms state-of-the-art approaches in two distinct prediction tasks: forecasting the closing price variation (regression) and predicting the market trend direction (classification), making it an effective tool to support stakeholders requiring accurate real-time predictions.
Antonio Pellicani, Gianvito Pio, Saso Dzeroski, Michelangelo Ceci
Data Min. Knowl. Discov.2
2026 Handling complex backgrounds and light perturbations for enhancing learning tasks from images of vegetables
abstract
Abstract The quality assessment of fruits and vegetables is crucial in the agroalimentary supply chain, as it directly affects consumer satisfaction, market value and overall food security. Traditional approaches rely on visual inspections or destructive techniques, which are labor-intensive and time-consuming. On the contrary, non-destructive techniques emerged as promising alternatives, offering solutions that can be adopted in real environments. Previous studies emphasized that the color distribution over images plays a significant role in the quality evaluation of food. In this paper, we propose a solution that leverages an autoencoder architecture to extract groups of relevant colors from the complete histogram of colors. To enhance the analysis of real-world images with complex backgrounds, we employ a pre-trained U2-Net architecture for background removal. Moreover, we propose a novel procedure based on outlier detection to identify and remove parts of the background that are not fully eliminated, especially along the edges of the product. After this preprocessing, we extract a complete color histogram which is fed to an autoencoder architecture, to extract high-level features representing color groups at different levels of granularity. The goal is to make the learned models less sensitive to light and color perturbations. Our experiments, conducted on two real-world datasets related to two different learning tasks, demonstrated the effectiveness of the proposed solution, that outperformed several baseline and state-of-the-art approaches, also based on complex neural network architectures.
Stefano Polimena, Gianvito Pio, Giovanni Attolico, Michelangelo Ceci
J. Intell. Inf. Syst.2
2024 Leveraging Spatio-Temporal Locality in Linear Model Trees for Multi-Step Time Series Forecasting
abstract
In the era of Big Data, forecasting the future measurements of geo-distributed sensors can be considered among the most fundamental tasks for several application domains. However, the spatial distribution of such sensors introduces several challenges for forecasting methods. The most important one is due to the spatial dimension which introduces autocorrelation phenomena, according to which nearby locations are not independent and should not be treated as such by the learning algorithms. Some existing approaches are able to capture the spatial autocorrelation, but they tend to model spatial information globally across all locations. On the contrary, our method, called SPLiT, focuses on capturing spatial information only among time series with similar trends, also at different timing, modeling the so-called spatio-temporal locality. The proposed method is based on linear model trees, which allow us to naturally model any form of discontinuities in the spatial autocorrelation during the tree growing phase, guided by heuristics that identify time series with similar trends. SPLiT works in the multi-step predictive setting in order to simultaneously provide forecasts for multiple time steps for multiple sensors in the future.Our experiments, conducted on two real-world datasets, show the effectiveness of the proposed method in forecasting the energy production of geographically distributed renewable power plants. The comparison against several tree-based models and state-of-the-art neural networks that consider both temporal and spatial dimensions shows the superiority of the proposed method.
Annunziata D'Aversa, Gianvito Pio, Michelangelo Ceci
IEEE Big Data2
2023 Multi-view overlapping clustering for the identification of the subject matter of legal judgments
Graziella De Martino, Gianvito Pio, Michelangelo Ceci
Inf. Sci.2
2022 Distributed Heterogeneous Transfer Learning for Link Prediction in the Positive Unlabeled Setting
abstract
Transfer learning focuses on enhancing predictive models for a target domain, by exploiting the knowledge coming from a related source domain. However, most existing transfer learning methods assume that source and target domains are described with the same feature spaces. Heterogeneous transfer learning approaches aim to overcome this limitation, but they usually introduce strong assumptions (e.g., on the number of features), cannot distribute the workload to handle large volumes of data, or cannot work in challenging settings like the Positive-Unlabeled (PU) setting, where only positive and unlabelled examples are available. In this paper, we present a novel heterogeneous distributed transfer learning method that can work also in PU learning setting and overcomes all such limitations.The experimental evaluation was conducted in the context of a link prediction task in the biological domain. The results showed the effectiveness of the proposed method, that outperformed three state-of-the-art heterogeneous transfer learning approaches.
Paolo Mignone, Gianvito Pio, Michelangelo Ceci
IEEE Big Data2
2022 LP-ROBIN: Link prediction in dynamic networks exploiting incremental node embedding
Emanuele Pio Barracchia, Gianvito Pio, Albert Bifet, Heitor Murilo Gomes, Bernhard Pfahringer, Michelangelo Ceci
Inf. Sci.2
2021 BROCCOLI: overlapping and outlier-robust biclustering through proximal stochastic gradient descent
abstract
Abstract Matrix tri-factorization subject to binary constraints is a versatile and powerful framework for the simultaneous clustering of observations and features, also known as biclustering. Applications for biclustering encompass the clustering of high-dimensional data and explorative data mining, where the selection of the most important features is relevant. Unfortunately, due to the lack of suitable methods for the optimization subject to binary constraints, the powerful framework of biclustering is typically constrained to clusterings which partition the set of observations or features. As a result, overlap between clusters cannot be modelled and every item, even outliers in the data, have to be assigned to exactly one cluster. In this paper we propose Broccoli , an optimization scheme for matrix factorization subject to binary constraints, which is based on the theoretically well-founded optimization scheme of proximal stochastic gradient descent. Thereby, we do not impose any restrictions on the obtained clusters. Our experimental evaluation, performed on both synthetic and real-world data, and against 6 competitor algorithms, show reliable and competitive performance, even in presence of a high amount of noise in the data. Moreover, a qualitative analysis of the identified clusters shows that Broccoli may provide meaningful and interpretable clustering structures.
Sibylle Hess, Gianvito Pio, Michiel E. Hochstenbach, Michelangelo Ceci
Data Min. Knowl. Discov.2
2018 Multi-type clustering and classification from heterogeneous networks
Gianvito Pio, Francesco Serafino 0002, Donato Malerba, Michelangelo Ceci
Inf. Sci.1
2018 Ensemble Learning for Multi-Type Classification in Heterogeneous Networks
abstract
Heterogeneous networks are networks consisting of different types of objects and links. They can be found in several fields, ranging from the Internet to social sciences, biology, epidemiology, geography, finance, and many others. In the literature, several methods have been proposed for the analysis of network data, but they usually focus on homogeneous networks, where all the objects are of the same type, and links among them describe a single type of relationship. More recently, the complexity of real scenarios has impelled researchers to design methods for the analysis of heterogeneous networks, especially focused on classification and clustering tasks. However, they often make assumptions on the structure of the network that are too restrictive or do not fully exploit different forms of network correlation and autocorrelation. Moreover, when nodes which are the main subject of the classification task are linked to several nodes of the network having missing values, standard methods can lead to either building incomplete classification models or to discarding possibly relevant dependencies (correlation or autocorrelation). In this paper, we propose an ensemble learning approach for multi-type classification. We adopt the system Mr-SBC, which is originally able to analyze heterogeneous networks of arbitrary structure, within an ensemble learning approach. The ensemble allows us to improve the classification accuracy of Mr-SBC by exploiting i) the possible presence of correlation and autocorrelation phenomena and ii) the classification of instances (which contain missing values) of other node types in the network. As a beneficial side effect, we have also that the models are more stable in terms of standard deviation of the accuracy, over different samples used for training. Experiments performed on real-world datasets show that the proposed method is able to significantly outperform the standard implementation of Mr-SBC. Moreover, it gives Mr-SBC the advantage of outperforming four other well-known algorithms for the classification of data organized in a network.
Francesco Serafino 0002, Gianvito Pio, Michelangelo Ceci
IEEE Trans. Knowl. Data Eng.2
2015 Non-negative Matrix Tri-Factorization for co-clustering: An analysis of the block matrix
Nicoletta Del Buono, Gianvito Pio
Inf. Sci.2
2014 Network Reconstruction for the Identification of miRNA: mRNA Interaction Networks
Gianvito Pio, Michelangelo Ceci, Domenica D'Elia, Donato Malerba
ECML/PKDD (3)1