Pasquale Coscia

dblp:179/2391 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-4726-3409ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
YearPublicationVenuePosition
2026 Diffusion-augmented direct classification: A few-shot learning framework for Synthetic Aperture Radar image automatic target recognition
Zilu Ying, Wenyu Ke, Yikui Zhai, Xinglin Liu, Pasquale Coscia, Angelo Genovese
Eng. Appl. Artif. Intell.6
2026 On the relevance of patch-based extraction methods for monocular depth estimation
abstract
Scene geometry estimation from images plays a key role in robotics, augmented reality, and autonomous systems. In particular, Monocular Depth Estimation (MDE) focuses on predicting depth using a single RGB image, avoiding the need for expensive sensors. State-of-the-art approaches use deep learning models for MDE while processing images as a whole, sub-optimally exploiting their spatial information. A recent research direction focuses on smaller image patches, as depth information varies across different regions of an image. This approach reduces model complexity and improves performance by capturing finer spatial details. From this perspective, we propose a novel warp patch-based extraction method which corrects perspective camera distortions, and employ it in tailored training and inference pipelines. Our experimental results show that our patch-based approach outperforms its full-image-trained counterpart and the classical crop patch-based extraction. With our technique, we obtain a general performance enhancements over recent state-of-the-art models. Code is available at https://github.com/AntonioFusillo/PatchMDE . • We propose a novel patch-based approach for monocular depth estimation. • Our method extracts patches from wide-aspect images, preserving camera parameters. • The designed patch-based inference outperforms full-image models in depth accuracy. • The proposed warp-based patch extraction is superior to patch cropping. • Our approach can wrap existing models, improving their performance.
Pasquale Coscia, Antonio Fusillo, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
Image Vis. Comput.1
2026 DLGCNet: Multimodal remote sensing semantic segmentation via dual diagonal low-rank adaptation and graph convolutional feature fusion
Jun-Ying Zeng, Xudong Jia 0001, Bin Deng 0003, Yikui Zhai, Chuanbo Qin, Pasquale Coscia, Angelo Genovese
Knowl. Based Syst.7
2026 DUR-Net+: Semi-Supervised Abdominal CT Pheochromocytoma Segmentation via Dynamic Uncertainty Rectified and Prior Knowledge From SAM-Med3D
abstract
Pheochromocytoma is a rare urological adrenal tumor disease. Automated segmentation of pheochromocytomas from computed tomography (CT) is essential for diagnosis and treatment. However, this task is a challenging one due to issues such as blurred boundaries, irregular shapes, variations in location and size, and the lack of annotated images for training. To address these issues, we propose a semi-supervised framework for pheochromocytoma segmentation that primarily consists of a dynamic uncertainty rectification mechanism and a supervised strategy based on SAM-Med3D prior knowledge. First, we design a semi-supervised segmentation model comprising a shared encoder and multiple independent decoders that dynamically select pseudo labels from the different decoder outputs. To mitigate the risk of unreliable predictions caused by sparse annotations during training, we introduce uncertainty estimation to prioritize reliable outputs. Additionally, an Attentional Convolution Block (ACB) is designed in the encoding stage to fully utilize both global and local features, improving tumor recognition in segmentation. Furthermore, SAM-Med3D prior knowledge is incorporated into the framework as supplementary supervisory information, aiding the model in learning from limited labeled data. To eliminate the labor-intensive requirement for manual prompts in SAM-Med3D, we leverage pseudo labels to generate high-quality mask prompts, thus transforming the clinical workflow. Experiments on two pheochromocytoma datasets from different centers demonstrate that our proposed method achieves competitive performance.
Chuanbo Qin, Zhuyuan Chen, Dong Wang 0083, Jun-Ying Zeng, Xudong Jia 0001, Maoqing Hu, Yikui Zhai, Pasquale Coscia, Angelo Genovese
IEEE J. Biomed. Health Informatics11
2026 Bidirectional Interactive Multi-Scale Aggregation Network for Vehicle Detection in Urban Traffic
abstract
Existing UAV vehicle-detection datasets, typically captured under static and uniform illumination, fail to adequately represent the variable lighting conditions, dense traffic, and frequent occlusions observed in real-world transportation hubs. To bridge this gap, a new dataset, UAV-HubSurveillance, is introduced to capture complex vehicle interactions across urban transportation nodes under diverse environmental scenarios. Although UAV-HubSurveillance provides rich and multidimensional interaction data, it still suffers from severe occlusions and adverse weather conditions that hinder detection and identification accuracy. To address these limitations, a novel vehicle detection framework, termed bidirectional interactive multi-scale aggregation-yolo (BIMSA-YOLO), is proposed, which integrates bidirectional feature interaction with adaptive multi-scale aggregation to enhance detection robustness. First, the bidirectional shallow fusion module (BSFM) facilitates cross-resolution information exchange through a lightweight gating strategy, preserving fine-grained details of small objects. Second, the interactive deep fusion module (IDFM) reinforces contextual coherence via attention-guided cross-level semantic fusion. Third, the multi-scale adaptive aggregation module (MSAAM) dynamically aligns and integrates multi-scale features to improve robustness against scale variation. Extensive experiments conducted on the UAV-HubSurveillance dataset demonstrate that BIMSA-YOLO significantly enhances detection performance under dynamic, occluded, and adverse-weather conditions. Specifically, the proposed model achieves an mAP0.5of 63.2%, surpassing the baseline by 5.3 percentage points. Furthermore, BIMSA-YOLO also exhibits strong generalization capabilities on VisDrone and CARPK datasets. Our code and dataset are available athttps://github.com/yikuizhai/BIMSA-YOLO
Chaojun Dong, Wenkang Qiu, Ye Li 0002, Yikui Zhai, Xiankun Liu, Chaoyun Mai, Hufei Zhu, Pasquale Coscia, Angelo Genovese, C. L. Philip Chen
IEEE Trans. Intell. Transp. Syst.9
2026 Useg-PanoDepth:Unified $360^{\circ }$ Depth Estimation for Indoor and Outdoor Scenes With Semantic Assistance
abstract
In complex$360^{\circ }$scenes, depth estimation is challenging for small objects and the depth of object boundaries, which cannot be effectively solved with existing works.$360^{\circ }$depth estimation is unable to produce uniform depth estimate findings in both indoor and outdoor settings due to the datasets. In this paper, the Useg-PanoDepth and PanoDepth dataset is proposed to improve the above problems effectively. The Diagonal-aware Attention Module (DAM) effectively estimates small objects in complex scenes. Enhanced Boundary Module (EBM), for enhancing boundary information,can also effectively solve the problem of depth unification of indoor and outdoor scenes. Extensive experiments on our constructed PanoDepth dataset, Useg-PanoDepth achieves SOTA results. The Relative accuracy (deltahttps://github.com/xjh6/Useg-PanoDepth.
Qingling Chang, Jingheng Xu, Yan Cui 0011, Yikui Zhai, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Multim.5
2025 Synthetic and (Un)Secure: Evaluating Generalized Membership Inference Attacks on Image Data
Pasquale Coscia, Stefano Ferrari, Vincenzo Piuri, Ayse Salman
SECRYPT1
2025 FELACS: Federated learning with adaptive client selection for IoT DDoS attack detection
abstract
Distributed denial-of-service (DDoS) attacks pose a significant threat to network security by overwhelming systems with malicious traffic, leading to service disruptions and potential data breaches. The traditional centralized machine learning (ML) methods for detecting DDoS attacks in Internet of Things (IoT) environments raise privacy and security concerns due to their collection and distribution of data to a central entity that may not be trusted to perform model training. Federated learning (FL) offers a privacy-preserving solution that enables distributed collaboration by training a model only on local clients, without data exchanges, where the central entity only performs global model aggregation. However, the current practice of random client selection, combined with the statistical heterogeneity of client data and the device heterogeneity encountered in IoT environments, requires many training rounds to reach optimal accuracy, increasing the imposed computational overhead. To address these challenges, we propose a multiobjective optimization-based FL with adaptive client selection (FELACS) approach that maximizes client importance scores while satisfying resource, performance, and data diversity constraints. Experiments are carried out on the CIC-IDS2018, CIC-DDoS2019, BoT-IoT, and CIC-IoT2023 datasets, demonstrating that FELACS improves upon the accuracy of the existing approaches while exhibiting increased convergence speed when training a model in an FL scenario, hence reducing the number of communication rounds required to achieve the target accuracy, making it highly effective for performing IoT-based DDoS attack detection in FL scenarios.
Mulualem Bitew Anley, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri
Comput. Secur.2
2025 OneN: Guided attention for natively-explainable anomaly detection
abstract
In industrial computer vision applications, anomaly detection (AD) is a critical task for ensuring product quality and system reliability. However, many existing AD systems follow a modular design that decouples classification from detection and localization tasks. Although this separation simplifies model development, it often limits generalizability and reduces practical effectiveness in real-world scenarios. Deep neural networks offer strong potential for unified solutions. Nonetheless, most current approaches still treat detection, localization and classification as separate components, hindering the development of more integrated and efficient AD pipelines. To bridge this gap, we propose OneN (One Network), a unified architecture that performs detection, localization, and classification within a single framework. Our approach distills knowledge from a high-capacity convolutional neural network (CNN) into an attention-based architecture trained under varying levels of supervision. The resulting attention maps act as interpretable pseudo-segmentation masks, enabling accurate localization of anomalous regions. To further enhance localization quality, we introduce a progressive focal loss that guides attention maps at each layer to focus on critical features. We validate our method through extensive experiments on both standardized and custom-defined industrial benchmarks. Even under weak supervision, it improves performance, reduces annotation effort, and facilitates scalable deployment in industrial environments.
Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
Image Vis. Comput.1
2025 PBSD-Net: Prismatic Battery Surface Defect Detection via Sliding Slice Amplification and Shunted Dynamic Snake Convolution
abstract
Automatically detecting surface defects in prismatic battery is crucial for ensuring quality meets established standards. Traditional methods face challenges in accurately identifying these defects due to their minute and varied shapes and high density of distribution. To address these issues, we propose an innovative network for prismatic battery surface defect (PBSD-Net), which employs shunted dynamic snake convolution and focal modulation to detect surface defects in prismatic battery. This network is integrated into the 2D-AOI system. Firstly, we introduce sliding slice amplification (SSA) as a training strategy to enhance the network’s ability to recognize densely clustered tiny defects. Secondly, we develop a novel method using the shunted dynamic snake convolution (SDSC) module and focal modulation (FM) to improve the extraction of deformation features, thereby addressing complex and sporadically scattered surface defects. By integrating the SDSC module and FM mechanism, the receptive field of the defect feature extraction network is expanded, enabling the acquisition of comprehensive defect edge features. Additionally, we introduce the quality focal loss (QFL) function to effectively tackle the issue of imbalanced sample types. Experimental results on the PBSD-RGB dataset demonstrate that our method achieves a mAP@50 of 85.8%, representing an improvement of approximately 7.7% over the baseline network. We have applied the PBSD-Net to an automatic defect detection system in a well-known battery production company. This enhancement significantly boosts the accuracy of surface defect detection in prismatic battery. The relevant code is at the https://github.com/yikuizhai/PBSD-Net.
Ying Xu 0005, Bo Li 0165, Yikui Zhai, Feng Ke, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans Autom. Sci. Eng.6
2025 GMTNet: Dense Object Detection via Global Dynamically Matching Transformer Network
abstract
In recent years, object detection models have been extensively applied across various industries, leveraging learned samples to recognize and locate objects. However, industrial environments present unique challenges, including complex backgrounds, dense object distributions, object stacking, and occlusion. To address these challenges, we propose the Global Dynamic Matching Transformer Network (GMTNet). GMTNet partitions images into blocks and employs a sliding window approach to capture information from each block and their interrelationships, mitigating background interference while acquiring global information for dense object recognition. By reweighting key-value pairs in multi-scale feature maps, GMTNet enhances global information relevance and effectively handles occlusion and overlap between objects. Furthermore, we introduce a dynamic sample matching method to tackle the issue of excessive candidate boxes in dense detection tasks. This method adaptively adjusts the number of matched positive samples according to the specific detection task, enabling the model to reduce the learning of irrelevant features and simplify post-processing. Experimental results demonstrate that GMTNet excels in dense detection tasks and outperforms current mainstream algorithms. The code will be available athttp://github.com/yikuizhai/GMTNet.
Chaojun Dong, Chengxuan Wang, Yikui Zhai, Ye Li 0002, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Circuits Syst. Video Technol.6
2025 AEGL-Net: Adaptive Multiscale Global-Local Feature Fusion Network for Remote Sensing Change Detection
abstract
With the rapid advancements in deep learning technology, the field of remote sensing change detection (RSCD) has witnessed significant improvements and innovations. In this context, bitemporal image processing, using features directly extracted by the backbone for subsequent fusion operations, may be obstructed by external environmental factors, potentially limiting the effective capture of complex feature variations. Moreover, overlooking local features during the fusion of bitemporal features can significantly affect the final detection results. As a result, achieving accurate change detection (CD) still encounters various challenges. To tackle these issues, this paper proposes a CD network (AEGL-Net) with Adaptive Multiscale Enhancement (AME) and Global-Local Feature Fusion (GLFF) modules. First, AME enhances features at each stage of backbone extraction through an adaptive strategy, balancing the enhancement of semantic information and texture details. Then, GLFF is used to fuse the bitemporal image features, which enhances the modeling of global dependencies while also fusing shared and context-aware weights to enhance the local features. Finally, the merged features are fed into the decoder to generate precise change maps. Experiments conducted with four open RSCD datasets (LEVIR-CD, S2Looking, SYSU-CD, and UAV-CD) demonstrate that our proposed AEGL-Net outperforms ten state-of-the-art models in the RSCD field. Our code is available at https://github.com/yikuizhai/AEGL-Net.
Zilu Ying, Yikui Zhai, Hufei Zhu, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.6
2025 Spatial Reconstruction and Joint Training in Transformer Network for Cross-Domain Remote Sensing Images Semantic Segmentation
abstract
Recently, Unsupervised Domain Adaptation (UDA) methods have attracted considerable attention in Remote Sensing Images (RSI) semantic segmentation. However, cross-domain RSI exhibit diverse scales, imbalanced distributions within domains, and significant inter-domain variations. In response to these challenges, we combine Spatial reconstruction and Joint training with the Transformer Network (SJT-Net). This framework introduces a spatial reconstruction method to address the issue of inconsistent ground sampling distances in cross domain RSI, which is rarely considered in existing approaches. Transferring domain knowledge at a similar spatial scale improves the spatial representation ability of UDA models. Unlike traditional adversarial training using ResNet for feature extraction, the SJT-Net employs Segformer, which enhances the model’s ability to capture in-class features across domains and improves global dependency modeling. Transmitting these refined features to the discriminator allows for more precise feature-level domain alignment. To enhance feature decoding, an interactive global-local decoder is constructed to efficiently capture both global relationships and local details of landform objects. Our framework leverages adversarial training to generate highly confident model weights and pseudo-labels for self-training in the target domain. Through iterative updates, the model’s generalization capability is gradually improved, eventually achieving optimal segmentation performance. Experimental results demonstrate that SJT-Net outperforms current UDA approaches and accomplishes state-of-the-art (SOTA) segmentation accuracy. The repository can be accessed at https://github.com/AnsonD0820/SJT-Net.
Jun-Ying Zeng, Senyao Deng, Yikui Zhai, Xudong Jia 0001, Chuanbo Qin, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.6
2025 CLIP-Vision Guided Few-Shot Metal Surface Defect Recognition
abstract
Metal surface defect recognition (MSDR) based on deep learning encounters the challenge of few-shot expert-labeled data. In this study, we proposed a CLIP-vision guided self supervised learning (CVGSSL) framework for representation learning of unlabeled data, completing MSDR using few-shot labeled data. This framework initially generates rich and diverse representation information through multiple CLIP-Vs to ensure effective SSL pretraining, followed by the design of an MLP-adapter to distill knowledge and adapt these representations to recognition tasks. In addition, we constructed a self-constrained loss to address the inherent problem of intraclass and interclass distance ambiguity that causes the representation to fall into an equivocal decision margin. Following label-free pretraining of CVGSSL, the downstream model adapts to one-shot to four-shot defect recognition tasks through fine-tuning. Experimental results demonstrate that CVGSSL outperforms state-of-the-art SSL methods across three public metal surface defect datasets, with the efficacy of the approach validated through extensive ablation experiments.
Tianlei Wang, Zeliang Li, Ying Xu 0005, Yikui Zhai, Xiaofen Xing, Kailing Guo, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Ind. Informatics7
2024 Features Disentanglement For Explainable Convolutional Neural Networks
abstract
Explainable methods for understanding deep neural networks are currently being employed for many visual tasks and provide valuable insights about their decisions. While post-hoc visual explanations offer easily understandable human cues behind neural networks’ decision-making processes, comparing their outcomes still remains challenging. Furthermore, balancing the performance-explainability trade-off could be a time-consuming process and require a deep domain knowledge. In this regard, we propose a novel auxiliary module, built upon convolutional-based encoders, which acts on the final layers of convolutional neural networks (CNNs) to learn orthogonal feature maps with a more discriminative and explainable power. This module is trained via a disentangle loss which specifically aims to decouple the object from the background in the input image. To quantitatively assess its impact on standard CNNs, and compare the quality of the resulting visual explanations, we employ metrics specifically designed for semantic segmentation tasks. These metrics rely on bounding-box annotations that may accompany image classification (or recognition) datasets, allowing us to compare both ground-truth and predicted regions. Finally, we explore the impact of various self-supervised pre-training strategies, due to their positive influence on vision tasks, and assess their effectiveness on our considered metrics.
Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri
ICIP1
2024 DS-HyFA-Net: A Deeply Supervised Hybrid Feature Aggregation Network With Multiencoders for Change Detection in High-Resolution Imagery
abstract
With the advancement of deep learning (DL) technologies, remarkable progress has been achieved in change detection (CD). Existing DL-based methods primarily focus on the discrepancy in bitemporal images, while overlooking the commonality in bitemporal images. However, one of the reasons hindering the improvement of CD performance is the inadequate utilization of image information. To address the above issue, we propose a Deeply Supervised Hybrid Feature Aggregation Network (DS-HyFA-Net). This network predicts changes by integrating the distinctness and the commonality in bitemporal images. Specifically, the DS-HyFA-Net primarily consists of a set of encoders and a Hybrid Feature Aggregation (HyFA) module. It uses a Siamese encoder (or Encoder I) and a specialized encoder (or Encoder II) to extract distinct and common features (CFs) in bitemporal images, respectively. The HyFA module efficiently aggregates distinct and common features (or hybrid features) and generates a change map using a predictor. In addition, a common feature learning strategy (CFLS) is introduced, based on deeply supervised (DS) techniques, to guide Encoder II in learning CFs. Experimental results on three well-recognized datasets demonstrate the effectiveness of the innovative DS-HyFA-Net, achieving F1-Scores of 93.33% on WHU-CD, 90.98% on LEVIR-CD, and 81.14% on SYSU-CD. Our code is available athttps://github.com/yikuizhai/DS-HyFA-Net.
Zilu Ying, Tingfeng Xian, Yikui Zhai, Xudong Jia 0001, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.7
2024 Efficient Adjacent Feature Harmonizer Network With UAV-CD+ Dataset for Remote Sensing Change Detection
abstract
Remote sensing change detection (RSCD) aims to identify changes within bi-temporal registered images. However, existing deep learning (DL)-based RSCD networks often suffer from large numbers of parameters, high computational complexity, and low inference speed, making it challenging to achieve efficient inference in real-world deployments. In addition, current models lack robust feature-fitting capabilities, necessitating the development of an efficient and powerful RSCD model to address this issue. Therefore, we propose a novel RSCD network named efficient adjacent feature harmonizer network (EAFH-Net) with fast computational speed and lightweight design. It is based on MobileNetV2, considering that change maps of different sizes contain temporal information of bitemporal features and spatial information at various scales, we introduce a multiscale feature neighbor fusion module (MFNFM) to address the lack of interaction between sophisticated-level and elementary-level features, and spatial and channel feature harmonizer module (SCFHM) to harmonize the spatiotemporal information of the change maps. Moreover, data-driven DL algorithms face another challenge due to insufficient granularity and the need for more practical datasets. Therefore, we present unmanned aerial vehicle (UAV)-CD+, a dataset comprising 2002 pairs of bi-temporal UAV low-altitude images, each sized at$1024\times 1024$. We performed experiments on three publicly accessible datasets in conjunction with UAV-CD+, comparing the results with other state-of-the-art (SOTA) methods. EAFH-Net attains the utmost precision, obtaining 91.74% on LEVIR-CD, 84.28% on SYSU-CD, 95.07% on WHU-CD, 79.12% on CLCD, and 70.12% on UAV-CD+. We have our model code available at the following link:https://github.com/yikuizhai/UCSFH-Net.
Yikui Zhai, Hongsheng Zhang 0001, Tingfeng Xian, Ying Xu 0005, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.6
2023 Adversarial Defect Synthesis for Industrial Products in Low Data Regime
abstract
Synthetic defect generation is an important aid for advanced manufacturing and production processes. Industrial scenarios rely on automated image-based quality control methods to avoid time-consuming manual inspections and promptly identify products not complying with specific quality standards. However, these methods show poor performance in the case of ill-posed low-data training regimes, and the lack of defective samples, due to operational costs or privacy policies, strongly limits their large-scale applicability.To overcome these limitations, we propose an innovative architecture based on an unpaired image-to-image (I2I) translation model to guide a transformation from a defect-free to a defective domain for common industrial products and propose simultaneously localizing their synthesized defects through a segmentation mask. As a performance evaluation, we measure image similarity and variability using standard metrics employed for generative models. Finally, we demonstrate that inspection networks, trained on synthesized samples, improve their accuracy in spotting real defective products.
Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri
ICIP1
2022 How many Observations are Enough? Knowledge Distillation for Trajectory Forecasting
abstract
Accurate prediction of future human positions is an essential task for modern video-surveillance systems. Current state-of-the-art models usually rely on a “history” of past tracked locations (e.g., 3 to 5 seconds) to predict a plausible sequence of future locations (e.g., up to the next 5 seconds). We feel that this common schema neglects critical traits of realistic applications: as the collection of input trajectories involves machine perception (i.e., detection and tracking), incorrect detection and fragmentation errors may accumulate in crowded scenes, leading to tracking drifts. On this account, the model would be fed with corrupted and noisy input data, thus fatally affecting its prediction performance. In this regard, we focus on delivering accurate predictions when only few input observations are used, thus potentially lowering the risks associated with automatic perception. To this end, we conceive a novel distillation strategy that allows a knowledge transfer from a teacher network to a student one, the latter fed with fewer observations (just two ones). We show that a properly defined teacher super-vision allows a student network to perform comparably to state-of-the-art approaches that demand more observations. Besides, extensive experiments on common trajectory forecasting datasets highlight that our student network better generalizes to unseen scenarios.
Alessio Monti, Angelo Porrello, Simone Calderara, Pasquale Coscia, Lamberto Ballan, Rita Cucchiara
CVPR4
2022 Early Pedestrian Intent Prediction via Features Estimation
abstract
Anticipating human motion is an essential requirement for autonomous vehicles and robots in order to primary guarantee people’s safety. In urban scenarios, they interact with humans, the surrounding environment, and other vehicles relying on several cues to forecast crossing or not crossing intentions. For these reasons, this challenging task is often tackled using both visual and non-visual features to anticipate future actions from 2 s to 1 s earlier the event. Our work primarily aims to revise this standard evaluation protocol to forecast crossing events as early as possible. To this end, we conceive a solution upon an extensively used model for egocentric action anticipation (RU-LSTM), proposing to envision future features, or modalities, that can better infer human intentions using a properly attention-based fusion mechanism. We validate our model against JAAD and PIE datasets and demonstrate that an intent prediction model can benefit from these additional clues for anticipating pedestrians crossing events.
Nada Osman, Enrico Cancelli, Guglielmo Camporese, Pasquale Coscia, Lamberto Ballan
ICIP4
2021 AC-VRNN: Attentive Conditional-VRNN for multi-future trajectory prediction
abstract
Anticipating human motion in crowded scenarios is essential for developing intelligent transportation systems, social-aware robots and advanced video surveillance applications. A key component of this task is represented by the inherently multi-modal nature of human paths which makes socially acceptable multiple futures when human interactions are involved. To this end, we propose a generative architecture for multi-future trajectory predictions based on Conditional Variational Recurrent Neural Networks (C-VRNNs). Conditioning mainly relies on prior belief maps, representing most likely moving directions and forcing the model to consider past observed dynamics in generating future positions. Human interactions are modelled with a graph-based attention mechanism enabling an online attentive hidden state refinement of the recurrent estimation. To corroborate our model, we perform extensive experiments on publicly-available datasets (e.g., ETH/UCY, Stanford Drone Dataset, STATS SportVU NBA, Intersection Drone Dataset and TrajNet++) and demonstrate its effectiveness in crowded scenes compared to several state-of-the-art methods.
Alessia Bertugli, Simone Calderara, Pasquale Coscia, Lamberto Ballan, Rita Cucchiara
Comput. Vis. Image Underst.3
2020 Knowledge Distillation for Action Anticipation via Label Smoothing
abstract
Human capability to anticipate near future from visual observations and non-verbal cues is essential for developing intelligent systems that need to interact with people. Several research areas, such as human-robot interaction (HRI), assisted living or autonomous driving need to foresee future events to avoid crashes or help people. Egocentric scenarios are classic examples where action anticipation is applied due to their numerous applications. Such challenging task demands to capture and model domain's hidden structure to reduce prediction uncertainty. Since multiple actions may equally occur in the future, we treat action anticipation as a multi-label problem with missing labels extending the concept of label smoothing. This idea resembles the knowledge distillation process since useful information is injected into the model during training. We implement a multi-modal framework based on long short-term memory (LSTM) networks to summarize past observations and make predictions at different time steps. We perform extensive experiments on EPIC-Kitchens and EGTEA Gaze+ datasets including more than 2500 and 100 action classes, respectively. The experiments show that label smoothing systematically improves performance of state-of-the-art models for action anticipation.
Guglielmo Camporese, Pasquale Coscia, Antonino Furnari, Giovanni Maria Farinella, Lamberto Ballan
ICPR2
2020 A CNN-RNN Framework for Image Annotation from Visual Cues and Social Network Metadata
abstract
Images represent a commonly used form of visual communication among people. Nevertheless, image classification may be a challenging task when dealing with unclear or non-common images needing more context to be correctly annotated. Metadata accompanying images on social-media represent an ideal source of additional information for retrieving proper neighborhoods easing image annotation task. To this end, we blend visual features extracted from neighbors and their metadata to jointly leverage context and visual cues. Our models use multiple semantic embeddings to achieve the dual objective of being robust to vocabulary changes between train and test sets and decoupling the architecture from the low-level metadata representation. Convolutional and recurrent neural networks (CNNs-RNNs) are jointly adopted to infer similarity among neighbors and query images. We perform comprehensive experiments on the NUS-WIDE dataset showing that our models outperform state-of-the-art architectures based on images and metadata, and decrease both sensory and semantic gaps to better annotate images.
Tobia Tesan, Pasquale Coscia, Lamberto Ballan
ICPR2
2018 Unsupervised Maritime Traffic Graph Learning with Mean-Reverting Stochastic Processes
abstract
Inspired by the fair regularity of the motion of ships, we present a method to derive a representation of the commercial maritime traffic in the form of a graph, whose nodes represent way-point areas, or regions of likely direction changes, and whose edges represent navigational legs with constant cruise velocity. The proposed method is based on the representation of a ship's velocity with an Ornstein-Uhlenbeck process and on the detection of changes of its long-run mean to identify navigational way-points. In order to assess the graph representativeness of the traffic, two performance metrics are introduced, leading to distinct graph construction criteria. Finally, the proposed method is validated against real-world Automatic Identification System data collected in a large area.
Pasquale Coscia, Francesco Palmieri 0001, Paolo Braca, Leonardo Maria Millefiori, Peter Willett 0001
FUSION1
2018 Long-term path prediction in urban scenarios using circular distributions
Pasquale Coscia, Francesco Castaldo, Francesco Palmieri 0001, Alexandre Alahi, Silvio Savarese, Lamberto Ballan
Image Vis. Comput.1
2016 Point-based path prediction from polar histograms
Pasquale Coscia, Francesco Castaldo, Francesco Palmieri 0001, Lamberto Ballan, Alexandre Alahi, Silvio Savarese
FUSION1