VLDB 2026 Research / reviewers in the wild / expert
Luca Colomba
dblp:289/2692
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-2911-4522ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discovering SpatioTemporally Invariant Event Patterns From Mobility DataabstractThe discovery of sequential patterns from spatiotemporal data is known to be a very complex data mining task. The relevance of spatiotemporal patterns to study event correlations in mobility data is established. Prior works addressed either the separate analysis of spatial and temporal dependencies among data, such as the co-location of events, or the study of the joint spatiotemporal properties of the trajectories observed over a region of interest. The aim of this paper is instead to overcome existing approaches by extracting sequences of discrete events showing spatiotemporally invariant properties. For example, ifan arbitrary bike sharing station becomes full (all its docks are used)thenwe will observe an increase in the occupancy level of the bike sharing stations in the surrounding area within ten minutes. We denote such a new pattern as a SpatioTemporally Invariant (STInv) event pattern because we observe several instances in the source data differing just in spatiotemporal shifts. We also propose a new algorithm to mine STInvs based on a prefix-projected sequential pattern growth approach and different quality metrics to quantify the contribution of the spatial invariance. The proposed approach is empirically evaluated on two mobility datasets related to a bike sharing system and traffic data. The results confirm the usability of the proposed solution in real-world scenarios. Luca Colomba, Luca Cagliero, Paolo Garza |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | ITALIC: An Italian Intent Classification DatasetabstractRecent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects.We introduce ITALIC, the first largescale speech dataset designed for intent classification in Italian.The dataset comprises 16,521 crowdsourced audio samples recorded by 70 speakers from various Italian regions and annotated with intent labels and additional metadata.We explore the versatility of ITALIC by evaluating current state-of-the-art speech and text models.Results on intent classification suggest that increasing scale and running language adaptation yield better speech models, monolingual text models outscore multilingual ones, and that speech recognition on ITALIC is more challenging than on existing Italian benchmarks.We release both the dataset and the annotation scheme to streamline the development of new Italian SLU models and language-specific datasets. Alkis Koudounas, Moreno La Quatra, Lorenzo Vaiani, Luca Colomba, Giuseppe Attanasio, Eliana Pastor, Luca Cagliero, Elena Baralis |
INTERSPEECH | 4 |
| 2023 | Density-Based Clustering by Means of Bridge Point IdentificationabstractDensity-based clustering focuses on defining clusters consisting of contiguous regions characterized by similar densities of points. Traditional approaches identify core points first, whereas more recent ones initially identify the cluster borders and then propagate cluster labels within the delimited regions. Both strategies encounter issues in presence of multi-density regions or when clusters are characterized by noisy borders. To overcome the above issues, we present a new clustering algorithm that relies on the concept of bridge point. A bridge point is a point whose neighborhood includes points of different clusters. The key idea is to use bridge points, rather than border points, to partition points into clusters. We have proved that a correct bridge point identification yields a cluster separation consistent with the expectation. To correctly identify bridge points in absence of a priori cluster information we leverage an established unsupervised outlier detection algorithm. Specifically, we empirically show that, in most cases, the detected outliers are actually a superset of the bridge point set. Therefore, to define clusters we spread cluster labels like a wildfire until an outlier, acting as a candidate bridge point, is reached. The proposed algorithm performs statistically better than state-of-the-art methods on a large set of benchmark datasets and is particularly robust to the presence of intra-cluster multiple densities and noisy borders. Luca Colomba, Luca Cagliero, Paolo Garza |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | A Dataset for Burned Area Delineation and Severity Estimation from Satellite ImageryabstractThe ability to correctly identify areas damaged by forest wildfires is essential to plan and monitor the restoration process and estimate the environmental damages after such catastrophic events. The wide availability of satellite data, combined with the recent development of machine learning and deep learning methodologies applied to the computer vision field, makes it extremely interesting to apply the aforementioned techniques to the field of automatic burned area detection. One of the main issues in such a context is the limited amount of labeled data, especially in the context of semantic segmentation. In this paper, we introduce a publicly available dataset for the burned area detection problem for semantic segmentation. The dataset contains 73 satellite images of different forests damaged by wildfires across Europe with a resolution of up to 10m per pixel. Data were collected from the Sentinel-2 L2A satellite mission and the target labels were generated from the Copernicus Emergency Management Service (EMS) annotations, with five different severity levels, ranging from undamaged to completely destroyed. Finally, we report the benchmark values obtained by applying a Convolutional Neural Network on the proposed dataset to address the burned area identification problem. Luca Colomba, Alessandro Farasin, Simone Monaco, Salvatore Greco, Paolo Garza, Daniele Apiletti, Elena Baralis, Tania Cerquitelli |
CIKM | 1 |
| 2022 | Mining spatiotemporally invariant patternsabstractDiscovering patterns that represent key spatial or temporal dependencies among data is a well-known exploratory data mining task. However, prior works either separately analyze spatial and temporal dependencies or discover joint spatiotemporal properties of specific trajectories observed over a region of interest. With the goal of generalizing the information provided by spatiotemporal patterns, in this paper we extract sequences of discrete events showing spatiotemporally invariant properties. We seek patterns whose corresponding instances in the source data differ only due to an invariant spatiotemporal transformation. We denote such a new type of patterns as SpatioTemporally Invariant. We also propose an efficient algorithm to mine STInvs and validate its efficiency and effectiveness on real data. Luca Colomba, Luca Cagliero, Paolo Garza |
SIGSPATIAL/GIS | 1 |
| 2020 | Improving Wildfire Severity Classification of Deep Learning U-Nets from Satellite ImagesabstractUncontrolled wildfires are dangerous events capable of harming people safety. To contrast their increasing impact in recent years, a key task is an accurate detection of the affected areas and their damage assessment from satellite images. Current state-of-the-art solutions address such problem through a double convolutional neural network able to automatically detect wildfires in satellite acquisitions and associate a damage index from a defined scale. However, such deep-learning model performance is strongly dependent on many factors. In this work, we specifically focus on a key parameter, i.e., the loss function, exploited in the underlying neural networks. Besides the state-of-the-art solutions based on the Dice-MSE, among the many loss functions proposed in literature, we focus on the Binary Cross-Entropy (BCE) and the Intersection over Union (IoU), as two representatives of the distribution-based and region-based categories, respectively. Experiments show that the BCE loss function coupled with a double-step U-Net architecture provides better results than current state-of-the-art solutions on a public labeled dataset of European wildfires. Simone Monaco, Andrea Pasini, Daniele Apiletti, Luca Colomba, Paolo Garza, Elena Baralis |
IEEE BigData | 4 |