VLDB 2026 Research / reviewers in the wild / expert
Bruno Vento
dblp:320/6012
· DBLP profile ↗
16ranked-venue papers
0as first author
16since 2021 · last 2026
0009-0006-7687-4929ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Simultaneous person attribute recognition using task-specific attention network on embedded devicesabstractPedestrian attribute recognition has become an important task in computer vision, particularly for retail and marketing and critical applications in security and surveillance. Despite its potential, achieving real-time performance on embedded devices while maintaining accuracy has been a significant challenge. In this paper, we propose a novel multi-task method for simultaneous person attribute recognition using task-specific attention network, which shares low-level representations across related tasks, reducing computational and memory requirements without compromising accuracy. In particular, we employ a spatial-channel attention mechanism to selectively focus on relevant regions without increasing the computational complexity of the backbone. Furthermore, we use a knowledge distillation technique to deal with missing labels and gradient normalization for dealing with task imbalances, since varying task difficulties lead to disproportionate gradient magnitudes during training. The experimental results demonstrate the effectiveness of our approach, achieving a mean accuracy of 0.889 while maintaining real-time performance at 114 frames per second on an embedded board with limited resources. These results highlight the practical viability and novelty of our system as a robust and scalable solution for pedestrian attribute recognition on embedded devices in real-world scenarios. George Azzopardi, Antonio Greco 0001, Alessia Saggese, Bruno Vento |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Joint underwater image enhancement and multi-scale fish detectionabstractAutomatic fish detection in images plays a crucial role in marine biodiversity monitoring and environmental conservation. The advent of deep learning has significantly improved the results achievable in fish detection; however, underwater challenges such as scale variability, optical distortions, and color inconsistencies reveal some limitations of conventional deep learning-based detectors. To address these issues, we propose an improved fish detection model designed for underwater environments trained to perform joint underwater image enhancement and multi-scale fish detection. While incorporating structural enhancements for multi-scale robustness, such as a dedicated tiny object detection head and attention mechanisms, the core of our approach resides in the joint training of an underwater image enhancement module with the detector. We evaluate our method on two challenging test sets. The experimental results demonstrate that our model significantly outperforms state-of-the-art methods, achieving an average precision between 0.84 and 0.89. Despite its superior accuracy, the proposed solution maintains a lightweight architecture with only 24 millions of parameters and ensures real-time processing capabilities (13 frames per second on edge hardware), highlighting its potential for effective deployment in marine research and autonomous fisheries monitoring. • YOLO-JUICE boosts small-object detection with an extra tiny-object head. • YOLO-JUICE preserves details without extra cost with a space to depth module. • YOLO-JUICE refines spatial and channel features with a spatial-channel attention. • YOLO-JUICE is jointly trained for image classification and enhancement. Vincenzo Carletti, Antonio Greco 0001, Andrea Vincenzo Ricciardi, Alessia Saggese, Bruno Vento |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Multi-task road surface condition recognition through a self-conditioned multi-label approachabstractAutonomous vehicle applications are gaining growing attention for their increasing reliability and capability to reduce the ecological impact of transportation. However, most prototypes fail to perceive the conditions of the road surface, which are crucial to adapt the driving controllers and avoid risks, e.g., due to hydroplaning. We propose a method to identify the major causes of grip reduction, namely friction, unevenness, and type of material, using a sequential architecture based on an LSTM module. As the multiple combination of these conditions alters the appearance of the ground surface, our model is trained to identify the most evident conditions first, then to use the predicted information to infer the less evident conditions. After the sequential prediction, a final classification stage allows for surpassing the limits of the sequential approach and recovering errors that occurred in early iterations, thus reducing their impact on subsequent ones. Experimental results show that the method outperforms previous ones on a challenging 27-class and 3-task dataset, by about 2.5 and 5.0 percentage points in terms of overall and balanced accuracy, respectively, while marginally increasing the inference time. Moreover, the conducted thorough error analysis demonstrates that our approach can implicitly learn the application constraints, like the absence of the material label if the surface is covered by snow, and provide useful information to interpret the sequence of predictions produced at each iteration of the sequential architecture. Diego Gragnaniello, Antonio Greco 0001, Mattia Marseglia, Carlo Sansone, Bruno Vento |
Expert Syst. Appl. | 5 |
| 2026 | Benchmarking illegal waste dumping detection: A public video dataset and a reference baselineabstractIllegal waste dumping represents one of the most pervasive threats to environmental health and the well-being of terrestrial ecosystems. As widely documented in the literature, this phenomenon is deeply rooted in cultural factors and civic behavior, making early detection a crucial component in effective prevention strategies. Recent advances in artificial intelligence offer promising avenues for automatically identifying illicit dumping actions; however, progress in this direction is severely hindered by the scarcity of publicly available video datasets tailored to the Illegal Waste Dumping Detection (IWDD) task. To address this critical gap, we introduce Mivia-IWDD-500 , a novel, fully balanced dataset comprising 500 videos: 250 positive samples and 250 negative samples. The positive class is further evenly divided into 125 videos depicting static disposal events, intentional and spatially localized acts such as depositing garbage bags or bulky items, and 125 videos capturing dynamic disposal behaviors, which involve brief, spontaneous, and spatially dispersed actions such as discarding waste while walking or from a moving vehicle. Alongside the dataset, we present a baseline model designed to serve as a reference point for future research. The proposed baseline model achieves an F 1 -score of 0.79, demonstrating the viability of the dataset and establishing a solid foundation for subsequent advancements. Antonio Greco 0001, Andrea Vincenzo Ricciardi, Carlo Sansone, Bruno Vento |
Image Vis. Comput. | 4 |
| 2026 | Parvelous: pedestrian attribute recognition using a vision encoder for multi-task learning on unbalanced sample distributionsabstractAbstract Pedestrian attribute recognition (PAR) is a critical task for real-time video surveillance and person re-identification in-the-wild. While modern vision–language models pre-trained on billions of image–text pairs have recently achieved outstanding accuracy, their substantial latency and high memory requirements make them impractical for real-world deployments. To overcome these limitations, we present Parvelous , an efficient and versatile multi-task framework built on an optimized vision encoder pre-trained for image–text matching and specifically adapted to the PAR domain. Through targeted architectural refinements and a selective layer-wise fine-tuning strategy, our framework ensures both efficiency and strong task specialization. Modular task-specific branches equipped with channel-wise attention are tailored in depth to address the varying complexity of binary and multi-class attribute recognition tasks. Additionally, a multi-task loss based on asymmetric loss functions mitigates the severe class imbalance inherent in standard PAR datasets, fostering robust learning across diverse attributes. Extensive evaluations on the public MIVIA PAR KD benchmark demonstrate that Parvelous achieves a 0.957 accuracy rate, setting a new state of the art, while delivering up to an 80-fold inference speedup compared to competing vision language models. This combination of accuracy and computational efficiency positions Parvelous as a practical and deployable solution for real-world PAR applications. Antonio Greco 0001, Andrea Vincenzo Ricciardi, Bruno Vento, Antonio Vitale |
Neural Comput. Appl. | 3 |
| 2026 | Guest editorial: special issue "from bench to the wild: recent advances in computer vision methods (WILD-VISION)"
George Azzopardi, Laura Fernández-Robles, Antonio Greco 0001, Bruno Vento |
Pattern Recognit. | 4 |
| 2026 | FAN-TAST-IC: Fast Alarm Notification with Task-Aware Spatio-Temporal Image ClassificationabstractVideo-based alarm notification systems play a critical role in safety-critical applications such as fire detection and pedestrian monitoring, because they allow for prompt intervention in the event of dangerous situations. However, existing approaches often struggle to balance high sensitivity and specificity with real-time performance, particularly in complex scenes. To address these limitations, we propose a novel method enabling Fast Alarm Notification with Task-Aware Spatio-Temporal Image Classification (FAN-TAST-IC). It is an innovative and efficient framework that combines a lightweight task-specific object detector with a pre-trained Vision-Language Model (VLM) encoder and a binary classifier. Unlike end-to-end multimodal systems, FAN-TAST-IC leverages the VLM solely as a frozen visual feature extractor, preserving its rich semantic knowledge while ensuring computational efficiency. The object detector performs a rough but real-time filtering of the temporal frames and spatial positions where objects can be. This filtering is then refined in two distinct ways, depending on the time constraints of the specific application, either to discard temporally incoherent detections or to confirm those with high confidence. The selection of such candidates drastically reduces the input space for the VLM and classifier to dubious detections only. Thus, the latter is trained on detector-guided positive and negative samples, enabling precise alarm validation, improving specificity, and preserving sensitivity without requiring extensive labeled datasets or fine-tuning the VLM. Experiments on fire and pedestrian detection tasks demonstrate the effectiveness of our method since FAN-TAST-IC consistently outperforms all the compared approaches, achieving superior F-scores and precision on challenging benchmarks while maintaining real-time capabilities. Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Leveraging Vision-Language Models for Improving Detection of Obstacles on Railway Tracks
Vincenzo Carletti, Antonio Greco 0001, Alessia Saggese, Camilla Spingola, Bruno Vento |
CAIP (2) | 5 |
| 2025 | An Extended Dataset and a Baseline for Pedestrian Attribute Recognition with Advanced Neural Networks
Antonio Greco 0001, Bruno Vento |
CAIP (1) | 2 |
| 2025 | Real-time joint recognition of weather and ground surface conditions by a multi-task deep networkabstractClimate change and the occurrence of intense and unexpected weather events highlighted the need for real-time weather warning systems, especially in smart roads and isolated scenarios like rural areas. In this work, we propose to jointly recognize the weather and the ground surface conditions using existing video surveillance systems. Previous works separately tackled these two tasks even if they are correlated to each other. We propose a convolutional neural network with shared weights in the lower layers and two separate classification branches on top to exploit the correlation between the tasks and, at the same time, learn diverse high-level features for each task. Moreover, the network architecture implements attention mechanisms allowing the classification branches to focus on diverse image regions. The method is versatile and allows us to train the network on partially labeled data. The experimental analysis on real data demonstrate the effectiveness of the proposed method on both tasks, confirmed by the accuracy comparison with existing methods for the recognition of weather and ground surface conditions. The multi-task solution improves the inference speed (50 frames per second) and reduces the required memory (less than 1 GB) with respect to a system with two different single-task approaches; these results confirm that the proposed solution is ready for video surveillance applications to support smart cities. • A novel multi-task neural network for real-time recognition of weather and ground surface conditions is proposed. • A task-specific attention mechanism focuses the network’s receptive field on relevant image parts. • A masked asymmetric loss is proposed to deal with unbalanced and partially labeled datasets • We adaptively balance the backpropagated gradients of the two tasks. Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | FOCUS: Improving fire detection on videos by scenario adaptationabstractFire detection from video is effective for most video surveillance applications. The algorithm that processes the video acquired by cameras in real time has a twofold goal: detect as many fires as possible and keep the number of false alarms low. While existing approaches obtain the first goal, they often produce many false alarms due to their inability to account for the specific characteristics of diverse application environments. This paper introduces Fire Observation and Control Using Scenarios (FOCUS), a novel configurable fire detection method designed to bridge the gap between the literature methods and the application needs by exploiting scenario-specific knowledge. FOCUS leverages scalable and configurable modules for robust fire detection, incorporating three key steps: (1) fire detection, which identifies potential fire regions using visual cues; (2) fire candidate filtering through motion analysis, to eliminate false positives by analyzing the dynamic behavior of the identified fire candidates; (3) a vision-language model, which evaluates and confirms fire alarms by correlating visual evidence with contextual knowledge. By tailoring the configuration to the scenario and integrating the advanced filtering mechanisms according to the complexity of the environment, FOCUS improves performance in all the considered application scenarios. The analysis of the results shows that the proposed approach outperforms existing methods demonstrating higher resilience on real data, which enables its usage in real-world applications. Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento |
Image Vis. Comput. | 4 |
| 2025 | FLAME: fire detection in videos combining a deep neural network with a model-based motion analysisabstractAbstract Among the catastrophic natural events posing hazards to human lives and infrastructures, fire is the phenomenon causing more frequent damages. Thanks to the spread of smart cameras, video fire detection is gaining more attention as a solution to monitor wide outdoor areas where no specific sensors for smoke detection are available. However, state-of-the-art fire detectors assure a satisfactory Recall but exhibit a high false-positive rate that renders the application practically unusable. In this paper, we propose FLAME, an efficient and adaptive classification framework to address fire detection from videos. The framework integrates a state-of-the-art deep neural network for frame-wise object detection, in an automatic video analysis tool. The advantages of our approach are twofold. On the one side, we exploit advances in image detector technology to ensure a high Recall. On the other side, we design a model-based motion analysis that improves the system’s Precision by filtering out fire candidates occurring in the scene’s background or whose movements differ from those of the fire. The proposed technique, able to be executed in real-time on embedded systems, has proven to surpass the methods considered for comparison on a recent literature dataset representing several scenarios. The code and the dataset used for designing the system have been made publicly available by the authors at ( https://mivia.unisa.it/large-fire-dataset-with-negative-samples-lfdn/ ). Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento |
Neural Comput. Appl. | 4 |
| 2025 | Guest editorial: special issue on pedestrian attribute recognition and person re-identification
Antonio Greco 0001, Modesto Castrillón-Santana, Bruno Vento |
Pattern Anal. Appl. | 3 |
| 2025 | Video Fire Recognition Using Zero-Shot Vision-Language Models Guided by a Task-Aware Object DetectorabstractFire detection from images or videos has gained a growing interest in recent years due to the criticality of the application. Both reliable real-time detectors and efficient retrieval techniques, able to process large databases acquired by sensor networks, are needed. Even if the reliability of artificial vision methods improved in the last years, some issues are still open problems. In particular, literature methods often reveal a low generalization capability when employed in scenarios different from the training ones in terms of framing distance, surrounding environment, or weather conditions. This can be addressed by considering contextual information and, more specifically, using vision-language models capable of interpreting and describing the framed scene. In this work, we propose FIRE-TASTIC: Fire Recognition with Task-Aware Spatio-Temporal Image Captioning, a novel framework to use object detectors in conjunction with vision-language models for fire detection and information retrieval. The localization capability of the former makes it able to detect even tiny fire traces but expose the system to false alarms. These are strongly reduced by the impressive zero-shot generalization capability of the latter, which can recognize and describe fire-like objects without prior fine-tuning. We also present a variant of the FIRE-TASTIC framework based on visual question answering instead of image captioning, which allows one to customize the retrieved information with personalized questions. To integrate the high-level information provided by both neural networks, we propose a novel method to query the vision-language models using the temporal and spatial localization information provided by the object detector. The proposal can improve the retrieval performance, as evidenced by the experiments conducted on two recent fire detection datasets, showing the effectiveness and the generalization capabilities of FIRE-TASTIC, which surpasses the state of the art. Moreover, the vision-language model, which is unsuitable for video processing due to its high computational load, is executed only on suspicious frames, allowing for real-time processing. This makes FIRE-TASTIC suitable for both real-time processing and information retrieval on large datasets. Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Fire and smoke detection from videos: A literature review under a novel taxonomyabstractThe recent development of deep learning based fire detection techniques and the availability of smart cameras able to execute these algorithms on the edge paved the way for sophisticated and efficient video-based firefighting systems. However, the limited available data to train these algorithms cast shadows on their robustness and generalization capability. In this survey, we review 153 papers published in the literature and 17 publicly available fire detection datasets with the aim of identifying application scenarios that better describe real-world fire detection challenges. In the proposed taxonomy, these are characterized by two features: i) the fire size in the framed scene that depends on several parameters, foremost the distance from the fire but also the camera optic; ii) the background activity, due to the presence of moving objects that may mislead the detector. On this basis, we analyzed the existing methods under a common scheme according to this new taxonomy and matched the solutions with the needs of specific application scenarios. Similarly, for 9 interesting video datasets acquired from cameras, we labeled 536 videos according to the proposed taxonomy and shared these annotations with the community. The aim of this fire detection review is two-fold: on one hand, we classify the existing scientific works according to the real application scenarios, determining the features that are promising in specific operative conditions; on the other hand, we provide a detailed analysis and annotation of available datasets to promote the development of more reliable validation protocols and the collection of data from missing scenarios. Diego Gragnaniello, Antonio Greco 0001, Carlo Sansone, Bruno Vento |
Expert Syst. Appl. | 4 |
| 2023 | PAR Contest 2023: Pedestrian Attributes Recognition with Multi-task Learning
Antonio Greco 0001, Bruno Vento |
CAIP (1) | 2 |