Angelo Genovese

dblp:89/10370 · DBLP profile ↗
← Back
38ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0002-3683-4723ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 12 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Security and privacy · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021
YearPublicationVenuePosition
2026 Diffusion-augmented direct classification: A few-shot learning framework for Synthetic Aperture Radar image automatic target recognition
Zilu Ying, Wenyu Ke, Yikui Zhai, Xinglin Liu, Pasquale Coscia, Angelo Genovese
Eng. Appl. Artif. Intell.7
2026 On the relevance of patch-based extraction methods for monocular depth estimation
abstract
Scene geometry estimation from images plays a key role in robotics, augmented reality, and autonomous systems. In particular, Monocular Depth Estimation (MDE) focuses on predicting depth using a single RGB image, avoiding the need for expensive sensors. State-of-the-art approaches use deep learning models for MDE while processing images as a whole, sub-optimally exploiting their spatial information. A recent research direction focuses on smaller image patches, as depth information varies across different regions of an image. This approach reduces model complexity and improves performance by capturing finer spatial details. From this perspective, we propose a novel warp patch-based extraction method which corrects perspective camera distortions, and employ it in tailored training and inference pipelines. Our experimental results show that our patch-based approach outperforms its full-image-trained counterpart and the classical crop patch-based extraction. With our technique, we obtain a general performance enhancements over recent state-of-the-art models. Code is available at https://github.com/AntonioFusillo/PatchMDE . • We propose a novel patch-based approach for monocular depth estimation. • Our method extracts patches from wide-aspect images, preserving camera parameters. • The designed patch-based inference outperforms full-image models in depth accuracy. • The proposed warp-based patch extraction is superior to patch cropping. • Our approach can wrap existing models, improving their performance.
Pasquale Coscia, Antonio Fusillo, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
Image Vis. Comput.3
2026 DLGCNet: Multimodal remote sensing semantic segmentation via dual diagonal low-rank adaptation and graph convolutional feature fusion
Jun-Ying Zeng, Xudong Jia 0001, Bin Deng 0003, Yikui Zhai, Chuanbo Qin, Pasquale Coscia, Angelo Genovese
Knowl. Based Syst.8
2026 DUR-Net+: Semi-Supervised Abdominal CT Pheochromocytoma Segmentation via Dynamic Uncertainty Rectified and Prior Knowledge From SAM-Med3D
abstract
Pheochromocytoma is a rare urological adrenal tumor disease. Automated segmentation of pheochromocytomas from computed tomography (CT) is essential for diagnosis and treatment. However, this task is a challenging one due to issues such as blurred boundaries, irregular shapes, variations in location and size, and the lack of annotated images for training. To address these issues, we propose a semi-supervised framework for pheochromocytoma segmentation that primarily consists of a dynamic uncertainty rectification mechanism and a supervised strategy based on SAM-Med3D prior knowledge. First, we design a semi-supervised segmentation model comprising a shared encoder and multiple independent decoders that dynamically select pseudo labels from the different decoder outputs. To mitigate the risk of unreliable predictions caused by sparse annotations during training, we introduce uncertainty estimation to prioritize reliable outputs. Additionally, an Attentional Convolution Block (ACB) is designed in the encoding stage to fully utilize both global and local features, improving tumor recognition in segmentation. Furthermore, SAM-Med3D prior knowledge is incorporated into the framework as supplementary supervisory information, aiding the model in learning from limited labeled data. To eliminate the labor-intensive requirement for manual prompts in SAM-Med3D, we leverage pseudo labels to generate high-quality mask prompts, thus transforming the clinical workflow. Experiments on two pheochromocytoma datasets from different centers demonstrate that our proposed method achieves competitive performance.
Chuanbo Qin, Zhuyuan Chen, Dong Wang 0083, Jun-Ying Zeng, Xudong Jia 0001, Maoqing Hu, Yikui Zhai, Pasquale Coscia, Angelo Genovese
IEEE J. Biomed. Health Informatics12
2026 Bidirectional Interactive Multi-Scale Aggregation Network for Vehicle Detection in Urban Traffic
abstract
Existing UAV vehicle-detection datasets, typically captured under static and uniform illumination, fail to adequately represent the variable lighting conditions, dense traffic, and frequent occlusions observed in real-world transportation hubs. To bridge this gap, a new dataset, UAV-HubSurveillance, is introduced to capture complex vehicle interactions across urban transportation nodes under diverse environmental scenarios. Although UAV-HubSurveillance provides rich and multidimensional interaction data, it still suffers from severe occlusions and adverse weather conditions that hinder detection and identification accuracy. To address these limitations, a novel vehicle detection framework, termed bidirectional interactive multi-scale aggregation-yolo (BIMSA-YOLO), is proposed, which integrates bidirectional feature interaction with adaptive multi-scale aggregation to enhance detection robustness. First, the bidirectional shallow fusion module (BSFM) facilitates cross-resolution information exchange through a lightweight gating strategy, preserving fine-grained details of small objects. Second, the interactive deep fusion module (IDFM) reinforces contextual coherence via attention-guided cross-level semantic fusion. Third, the multi-scale adaptive aggregation module (MSAAM) dynamically aligns and integrates multi-scale features to improve robustness against scale variation. Extensive experiments conducted on the UAV-HubSurveillance dataset demonstrate that BIMSA-YOLO significantly enhances detection performance under dynamic, occluded, and adverse-weather conditions. Specifically, the proposed model achieves an mAP0.5of 63.2%, surpassing the baseline by 5.3 percentage points. Furthermore, BIMSA-YOLO also exhibits strong generalization capabilities on VisDrone and CARPK datasets. Our code and dataset are available athttps://github.com/yikuizhai/BIMSA-YOLO
Chaojun Dong, Wenkang Qiu, Ye Li 0002, Yikui Zhai, Xiankun Liu, Chaoyun Mai, Hufei Zhu, Pasquale Coscia, Angelo Genovese, C. L. Philip Chen
IEEE Trans. Intell. Transp. Syst.10
2026 Useg-PanoDepth:Unified $360^{\circ }$ Depth Estimation for Indoor and Outdoor Scenes With Semantic Assistance
abstract
In complex$360^{\circ }$scenes, depth estimation is challenging for small objects and the depth of object boundaries, which cannot be effectively solved with existing works.$360^{\circ }$depth estimation is unable to produce uniform depth estimate findings in both indoor and outdoor settings due to the datasets. In this paper, the Useg-PanoDepth and PanoDepth dataset is proposed to improve the above problems effectively. The Diagonal-aware Attention Module (DAM) effectively estimates small objects in complex scenes. Enhanced Boundary Module (EBM), for enhancing boundary information,can also effectively solve the problem of depth unification of indoor and outdoor scenes. Extensive experiments on our constructed PanoDepth dataset, Useg-PanoDepth achieves SOTA results. The Relative accuracy (deltahttps://github.com/xjh6/Useg-PanoDepth.
Qingling Chang, Jingheng Xu, Yan Cui 0011, Yikui Zhai, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Multim.6
2025 FLIFRA: Hybrid Data Poisoning Attack Detection in Federated Learning for IoT Security
abstract
The rapid expansion of IoT devices has transformed numerous industries by enabling extensive data collection and real-time analytics. Federated Learning (FL) offers a decentralized model training paradigm that ensures data privacy, making it particularly suitable for IoT environments. Yet, it remains vulnerable to poisoning attacks that can severely compromise model integrity, wherein malicious clients compromise the global model by injecting poisoned updates. Existing defenses, which focus primarily on global model performance, often fail to effectively integrate local anomaly detection with global weighting mechanisms, thus limiting their efficacy against such threats. Addressing this research gap, we propose FLIFRA (Federated Learning Isolation Forest with Robust Aggregation), a hybrid defense framework that combines client-side anomaly detection using Isolation Forest (iForest) with dynamic reputation-based robust aggregation at the server. This dual-layer approach filters out malicious updates before aggregation and adjusts client reputations to mitigate adversarial influence. Our evaluation of three cybersecurity datasets (CIC-IDS2018, BoT-IoT, and UNSW-NB15) under various intensities of poisoning (10%, 20%, 30%, and 40%) demonstrates that the proposed method outperforms the traditional aggregation schemes of FedAvg, Krum, Trimmed Mean, DRRA, and WeiDetect in the literature. In particular, our framework achieves higher detection accuracy, faster convergence, and improved stability, even in highly heterogeneous data environments.
Mulualem Bitew Anley, Angelo Genovese, Tibebe Beshah Tesema, Vincenzo Piuri
SMC2
2025 FELACS: Federated learning with adaptive client selection for IoT DDoS attack detection
abstract
Distributed denial-of-service (DDoS) attacks pose a significant threat to network security by overwhelming systems with malicious traffic, leading to service disruptions and potential data breaches. The traditional centralized machine learning (ML) methods for detecting DDoS attacks in Internet of Things (IoT) environments raise privacy and security concerns due to their collection and distribution of data to a central entity that may not be trusted to perform model training. Federated learning (FL) offers a privacy-preserving solution that enables distributed collaboration by training a model only on local clients, without data exchanges, where the central entity only performs global model aggregation. However, the current practice of random client selection, combined with the statistical heterogeneity of client data and the device heterogeneity encountered in IoT environments, requires many training rounds to reach optimal accuracy, increasing the imposed computational overhead. To address these challenges, we propose a multiobjective optimization-based FL with adaptive client selection (FELACS) approach that maximizes client importance scores while satisfying resource, performance, and data diversity constraints. Experiments are carried out on the CIC-IDS2018, CIC-DDoS2019, BoT-IoT, and CIC-IoT2023 datasets, demonstrating that FELACS improves upon the accuracy of the existing approaches while exhibiting increased convergence speed when training a model in an FL scenario, hence reducing the number of communication rounds required to achieve the target accuracy, making it highly effective for performing IoT-based DDoS attack detection in FL scenarios.
Mulualem Bitew Anley, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri
Comput. Secur.3
2025 OneN: Guided attention for natively-explainable anomaly detection
abstract
In industrial computer vision applications, anomaly detection (AD) is a critical task for ensuring product quality and system reliability. However, many existing AD systems follow a modular design that decouples classification from detection and localization tasks. Although this separation simplifies model development, it often limits generalizability and reduces practical effectiveness in real-world scenarios. Deep neural networks offer strong potential for unified solutions. Nonetheless, most current approaches still treat detection, localization and classification as separate components, hindering the development of more integrated and efficient AD pipelines. To bridge this gap, we propose OneN (One Network), a unified architecture that performs detection, localization, and classification within a single framework. Our approach distills knowledge from a high-capacity convolutional neural network (CNN) into an attention-based architecture trained under varying levels of supervision. The resulting attention maps act as interpretable pseudo-segmentation masks, enabling accurate localization of anomalous regions. To further enhance localization quality, we introduce a progressive focal loss that guides attention maps at each layer to focus on critical features. We validate our method through extensive experiments on both standardized and custom-defined industrial benchmarks. Even under weak supervision, it improves performance, reduces annotation effort, and facilitates scalable deployment in industrial environments.
Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
Image Vis. Comput.2
2025 PBSD-Net: Prismatic Battery Surface Defect Detection via Sliding Slice Amplification and Shunted Dynamic Snake Convolution
abstract
Automatically detecting surface defects in prismatic battery is crucial for ensuring quality meets established standards. Traditional methods face challenges in accurately identifying these defects due to their minute and varied shapes and high density of distribution. To address these issues, we propose an innovative network for prismatic battery surface defect (PBSD-Net), which employs shunted dynamic snake convolution and focal modulation to detect surface defects in prismatic battery. This network is integrated into the 2D-AOI system. Firstly, we introduce sliding slice amplification (SSA) as a training strategy to enhance the network’s ability to recognize densely clustered tiny defects. Secondly, we develop a novel method using the shunted dynamic snake convolution (SDSC) module and focal modulation (FM) to improve the extraction of deformation features, thereby addressing complex and sporadically scattered surface defects. By integrating the SDSC module and FM mechanism, the receptive field of the defect feature extraction network is expanded, enabling the acquisition of comprehensive defect edge features. Additionally, we introduce the quality focal loss (QFL) function to effectively tackle the issue of imbalanced sample types. Experimental results on the PBSD-RGB dataset demonstrate that our method achieves a mAP@50 of 85.8%, representing an improvement of approximately 7.7% over the baseline network. We have applied the PBSD-Net to an automatic defect detection system in a well-known battery production company. This enhancement significantly boosts the accuracy of surface defect detection in prismatic battery. The relevant code is at the https://github.com/yikuizhai/PBSD-Net.
Ying Xu 0005, Bo Li 0165, Yikui Zhai, Feng Ke, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans Autom. Sci. Eng.7
2025 GMTNet: Dense Object Detection via Global Dynamically Matching Transformer Network
abstract
In recent years, object detection models have been extensively applied across various industries, leveraging learned samples to recognize and locate objects. However, industrial environments present unique challenges, including complex backgrounds, dense object distributions, object stacking, and occlusion. To address these challenges, we propose the Global Dynamic Matching Transformer Network (GMTNet). GMTNet partitions images into blocks and employs a sliding window approach to capture information from each block and their interrelationships, mitigating background interference while acquiring global information for dense object recognition. By reweighting key-value pairs in multi-scale feature maps, GMTNet enhances global information relevance and effectively handles occlusion and overlap between objects. Furthermore, we introduce a dynamic sample matching method to tackle the issue of excessive candidate boxes in dense detection tasks. This method adaptively adjusts the number of matched positive samples according to the specific detection task, enabling the model to reduce the learning of irrelevant features and simplify post-processing. Experimental results demonstrate that GMTNet excels in dense detection tasks and outperforms current mainstream algorithms. The code will be available athttp://github.com/yikuizhai/GMTNet.
Chaojun Dong, Chengxuan Wang, Yikui Zhai, Ye Li 0002, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Circuits Syst. Video Technol.7
2025 AEGL-Net: Adaptive Multiscale Global-Local Feature Fusion Network for Remote Sensing Change Detection
abstract
With the rapid advancements in deep learning technology, the field of remote sensing change detection (RSCD) has witnessed significant improvements and innovations. In this context, bitemporal image processing, using features directly extracted by the backbone for subsequent fusion operations, may be obstructed by external environmental factors, potentially limiting the effective capture of complex feature variations. Moreover, overlooking local features during the fusion of bitemporal features can significantly affect the final detection results. As a result, achieving accurate change detection (CD) still encounters various challenges. To tackle these issues, this paper proposes a CD network (AEGL-Net) with Adaptive Multiscale Enhancement (AME) and Global-Local Feature Fusion (GLFF) modules. First, AME enhances features at each stage of backbone extraction through an adaptive strategy, balancing the enhancement of semantic information and texture details. Then, GLFF is used to fuse the bitemporal image features, which enhances the modeling of global dependencies while also fusing shared and context-aware weights to enhance the local features. Finally, the merged features are fed into the decoder to generate precise change maps. Experiments conducted with four open RSCD datasets (LEVIR-CD, S2Looking, SYSU-CD, and UAV-CD) demonstrate that our proposed AEGL-Net outperforms ten state-of-the-art models in the RSCD field. Our code is available at https://github.com/yikuizhai/AEGL-Net.
Zilu Ying, Yikui Zhai, Hufei Zhu, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.7
2025 Spatial Reconstruction and Joint Training in Transformer Network for Cross-Domain Remote Sensing Images Semantic Segmentation
abstract
Recently, Unsupervised Domain Adaptation (UDA) methods have attracted considerable attention in Remote Sensing Images (RSI) semantic segmentation. However, cross-domain RSI exhibit diverse scales, imbalanced distributions within domains, and significant inter-domain variations. In response to these challenges, we combine Spatial reconstruction and Joint training with the Transformer Network (SJT-Net). This framework introduces a spatial reconstruction method to address the issue of inconsistent ground sampling distances in cross domain RSI, which is rarely considered in existing approaches. Transferring domain knowledge at a similar spatial scale improves the spatial representation ability of UDA models. Unlike traditional adversarial training using ResNet for feature extraction, the SJT-Net employs Segformer, which enhances the model’s ability to capture in-class features across domains and improves global dependency modeling. Transmitting these refined features to the discriminator allows for more precise feature-level domain alignment. To enhance feature decoding, an interactive global-local decoder is constructed to efficiently capture both global relationships and local details of landform objects. Our framework leverages adversarial training to generate highly confident model weights and pseudo-labels for self-training in the target domain. Through iterative updates, the model’s generalization capability is gradually improved, eventually achieving optimal segmentation performance. Experimental results demonstrate that SJT-Net outperforms current UDA approaches and accomplishes state-of-the-art (SOTA) segmentation accuracy. The repository can be accessed at https://github.com/AnsonD0820/SJT-Net.
Jun-Ying Zeng, Senyao Deng, Yikui Zhai, Xudong Jia 0001, Chuanbo Qin, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.7
2025 CLIP-Vision Guided Few-Shot Metal Surface Defect Recognition
abstract
Metal surface defect recognition (MSDR) based on deep learning encounters the challenge of few-shot expert-labeled data. In this study, we proposed a CLIP-vision guided self supervised learning (CVGSSL) framework for representation learning of unlabeled data, completing MSDR using few-shot labeled data. This framework initially generates rich and diverse representation information through multiple CLIP-Vs to ensure effective SSL pretraining, followed by the design of an MLP-adapter to distill knowledge and adapt these representations to recognition tasks. In addition, we constructed a self-constrained loss to address the inherent problem of intraclass and interclass distance ambiguity that causes the representation to fall into an equivocal decision margin. Following label-free pretraining of CVGSSL, the downstream model adapts to one-shot to four-shot defect recognition tasks through fine-tuning. Experimental results demonstrate that CVGSSL outperforms state-of-the-art SSL methods across three public metal surface defect datasets, with the efficacy of the approach validated through extensive ablation experiments.
Tianlei Wang, Zeliang Li, Ying Xu 0005, Yikui Zhai, Xiaofen Xing, Kailing Guo, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Ind. Informatics8
2024 Features Disentanglement For Explainable Convolutional Neural Networks
abstract
Explainable methods for understanding deep neural networks are currently being employed for many visual tasks and provide valuable insights about their decisions. While post-hoc visual explanations offer easily understandable human cues behind neural networks’ decision-making processes, comparing their outcomes still remains challenging. Furthermore, balancing the performance-explainability trade-off could be a time-consuming process and require a deep domain knowledge. In this regard, we propose a novel auxiliary module, built upon convolutional-based encoders, which acts on the final layers of convolutional neural networks (CNNs) to learn orthogonal feature maps with a more discriminative and explainable power. This module is trained via a disentangle loss which specifically aims to decouple the object from the background in the input image. To quantitatively assess its impact on standard CNNs, and compare the quality of the resulting visual explanations, we employ metrics specifically designed for semantic segmentation tasks. These metrics rely on bounding-box annotations that may accompany image classification (or recognition) datasets, allowing us to compare both ground-truth and predicted regions. Finally, we explore the impact of various self-supervised pre-training strategies, due to their positive influence on vision tasks, and assess their effectiveness on our considered metrics.
Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri
ICIP2
2024 Robust DDoS attack detection with adaptive transfer learning
abstract
In the evolving cybersecurity landscape, the rising frequency of Distributed Denial of Service (DDoS) attacks requires robust defense mechanisms to safeguard network infrastructure availability and integrity. Deep Learning (DL) models have emerged as a promising approach for DDoS attack detection and mitigation due to their capability of automatically learning feature representations and distinguishing complex patterns within network traffic data. However, the effectiveness of DL models in protecting against evolving attacks depends also on the design of adaptive architectures, through the combination of appropriate models, quality data, and thorough hyperparameter optimizations, which are scarcely performed in the literature. Also, within adaptive architectures for DDoS detection, no method has yet addressed how to transfer knowledge between different datasets to improve classification accuracy. In this paper, we propose an innovative approach for DDoS detection by leveraging Convolutional Neural Networks (CNN), adaptive architectures, and transfer learning techniques. Experimental results on publicly available datasets show that the proposed adaptive transfer learning method effectively identifies benign and malicious activities and specific attack categories.
Mulualem Bitew Anley, Angelo Genovese, Davide Agostinello, Vincenzo Piuri
Comput. Secur.2
2024 A decision support system for acute lymphoblastic leukemia detection based on explainable artificial intelligence
Angelo Genovese, Vincenzo Piuri, Fabio Scotti
Image Vis. Comput.1
2024 DGMA2-Net: A Difference-Guided Multiscale Aggregation Attention Network for Remote Sensing Change Detection
abstract
Remote sensing change detection (RSCD) focuses on identifying regions that have undergone changes between two remote sensing images captured at different times. Recently, convolutional neural networks (CNNs) have shown promising results in the challenging task of RSCD. However, these methods do not efficiently fuse bitemporal features and extract useful information that is beneficial to subsequent RSCD tasks. In addition, they did not consider multilevel feature interactions in feature aggregation and ignore relationships between difference features and bitemporal features, which thus affects the RSCD results. To address the above problems, a difference-guided multiscale aggregation attention network, DGMA2-Net, is developed. Bitemporal features at different levels are extracted through a Siamese convolutional network and a multiscale difference fusion module (MDFM) is then created to fuse bitemporal features and extract, in a multiscale manner, difference features containing rich contextual information. After the MDFM treatment, two difference aggregation modules (DAMs) are used to aggregate difference features at different levels for multilevel feature interactions. The features through DAMs are sent to the difference-enhanced attention modules (DEAMs) to strengthen the connections between bitemporal features and difference features and further refine change features. Finally, refined change features are superimposed from deep to shallow and a change map is produced. In validating the effectiveness of DGMA2-Net, a series of experiments are conducted on three public RSCD benchmark datasets (LEVIR-CD, BCDD, and SYSU-CD). The experimental results demonstrate that DGMA2-Net surpasses the current eight state-of-the-art methods in RSCD. Our code is released at https://github.com/yikuizhai/DGMA2-Net.
Zilu Ying, Zijun Tan, Yikui Zhai, Xudong Jia 0001, Wenba Li, Jun-Ying Zeng, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.7
2024 DS-HyFA-Net: A Deeply Supervised Hybrid Feature Aggregation Network With Multiencoders for Change Detection in High-Resolution Imagery
abstract
With the advancement of deep learning (DL) technologies, remarkable progress has been achieved in change detection (CD). Existing DL-based methods primarily focus on the discrepancy in bitemporal images, while overlooking the commonality in bitemporal images. However, one of the reasons hindering the improvement of CD performance is the inadequate utilization of image information. To address the above issue, we propose a Deeply Supervised Hybrid Feature Aggregation Network (DS-HyFA-Net). This network predicts changes by integrating the distinctness and the commonality in bitemporal images. Specifically, the DS-HyFA-Net primarily consists of a set of encoders and a Hybrid Feature Aggregation (HyFA) module. It uses a Siamese encoder (or Encoder I) and a specialized encoder (or Encoder II) to extract distinct and common features (CFs) in bitemporal images, respectively. The HyFA module efficiently aggregates distinct and common features (or hybrid features) and generates a change map using a predictor. In addition, a common feature learning strategy (CFLS) is introduced, based on deeply supervised (DS) techniques, to guide Encoder II in learning CFs. Experimental results on three well-recognized datasets demonstrate the effectiveness of the innovative DS-HyFA-Net, achieving F1-Scores of 93.33% on WHU-CD, 90.98% on LEVIR-CD, and 81.14% on SYSU-CD. Our code is available athttps://github.com/yikuizhai/DS-HyFA-Net.
Zilu Ying, Tingfeng Xian, Yikui Zhai, Xudong Jia 0001, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.8
2024 Efficient Adjacent Feature Harmonizer Network With UAV-CD+ Dataset for Remote Sensing Change Detection
abstract
Remote sensing change detection (RSCD) aims to identify changes within bi-temporal registered images. However, existing deep learning (DL)-based RSCD networks often suffer from large numbers of parameters, high computational complexity, and low inference speed, making it challenging to achieve efficient inference in real-world deployments. In addition, current models lack robust feature-fitting capabilities, necessitating the development of an efficient and powerful RSCD model to address this issue. Therefore, we propose a novel RSCD network named efficient adjacent feature harmonizer network (EAFH-Net) with fast computational speed and lightweight design. It is based on MobileNetV2, considering that change maps of different sizes contain temporal information of bitemporal features and spatial information at various scales, we introduce a multiscale feature neighbor fusion module (MFNFM) to address the lack of interaction between sophisticated-level and elementary-level features, and spatial and channel feature harmonizer module (SCFHM) to harmonize the spatiotemporal information of the change maps. Moreover, data-driven DL algorithms face another challenge due to insufficient granularity and the need for more practical datasets. Therefore, we present unmanned aerial vehicle (UAV)-CD+, a dataset comprising 2002 pairs of bi-temporal UAV low-altitude images, each sized at$1024\times 1024$. We performed experiments on three publicly accessible datasets in conjunction with UAV-CD+, comparing the results with other state-of-the-art (SOTA) methods. EAFH-Net attains the utmost precision, obtaining 91.74% on LEVIR-CD, 84.28% on SYSU-CD, 95.07% on WHU-CD, 79.12% on CLCD, and 70.12% on UAV-CD+. We have our model code available at the following link:https://github.com/yikuizhai/UCSFH-Net.
Yikui Zhai, Hongsheng Zhang 0001, Tingfeng Xian, Ying Xu 0005, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.7
2024 Learning to Count Arbitrary Industrial Manufacturing Workpieces
abstract
Man-made workpiece counting is a routine job for manufactory workers; however, this is an error-prone task. In this article, we are interested in detecting and counting arbitrary workpieces in industrial manufacturing. Therefore, we construct a comprehensive and large-scale open-world public benchmark dataset for workpiece counting, called workpiece counting dataset, which includes 121 475 instances of workpieces from 351 different categories. We also propose a novel method for workpiece detection and counting, named two-stage workpiece counting network. The first stage of the network is to develop a class-agnostic detector to localize each workpiece instance, followed by the second stage to employ an unsupervised deep clustering strategy with the backbone network pretrained in a workpiece convolutional autoencoder for decision boundary prediction, achieving workpiece clustering under unknownKvalues. Finally, our experiments show that the proposed method outperforms current mainstream methods, greatly enhancing the efficiency of factory operations.
Yikui Zhai, Feng Ke, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Ind. Informatics5
2024 Large-Scale High-Altitude UAV-Based Vehicle Detection via Pyramid Dual Pooling Attention Path Aggregation Network
abstract
UAVs can collect vehicle data in high-altitude scenes, playing a significant role in intelligent urban management due to their wide of view. Nevertheless, the current datasets for UAV-based vehicle detection are acquired at altitude below 150 meters. This contrasts with the data perspective obtained from high-altitude scenes, potentially leading to incongruities in data distribution. Consequently, it is challenging to apply these datasets effectively in high-altitude scenes, and there is an ongoing obstacle. To resolve this challenge, we developed a comprehensive vehicle dataset named LH-UAV-Vehicle, specifically collected at flight altitudes ranging from 250 to 400 meters. Collecting data at higher flight altitudes offers a broader perspective, but it concurrently introduces complexity and diversity in the background, which consequently impacts vehicle localization and recognition accuracy. In response, we proposed the pyramid dual pooling attention path aggregation network (PDPA-PAN), an innovative framework that improves detection performance in high-altitude scenes by combining spatial and semantic information. Object attention integration in both spatial and channel dimensions is aimed by the pyramid dual pooling attention module (PDPAM), which is achieved through the parallel integration of two distinct attention mechanisms. Furthermore, we have individually developed the pyramid pooling attention module (PPAM) and the dual pooling attention module (DPAM). The PPAM emphasizes channel attention, while the DPAM prioritizes spatial attention. This design aims to enhance vehicle information and suppress background interference more effectively. Extensive experiments conducted on the LH-UAV-Vehicle conclusively demonstrate the efficacy of the proposed vehicle detection method. Our code and dataset can be found at https://github.com/yikuizhai/PDPA-PAN.
Zilu Ying, Yikui Zhai, Hao Quan 0002, Wenba Li, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Intell. Transp. Syst.6
2023 Adversarial Defect Synthesis for Industrial Products in Low Data Regime
abstract
Synthetic defect generation is an important aid for advanced manufacturing and production processes. Industrial scenarios rely on automated image-based quality control methods to avoid time-consuming manual inspections and promptly identify products not complying with specific quality standards. However, these methods show poor performance in the case of ill-posed low-data training regimes, and the lack of defective samples, due to operational costs or privacy policies, strongly limits their large-scale applicability.To overcome these limitations, we propose an innovative architecture based on an unpaired image-to-image (I2I) translation model to guide a transformation from a defect-free to a defective domain for common industrial products and propose simultaneously localizing their synthesized defects through a segmentation mask. As a performance evaluation, we measure image similarity and variability using standard metrics employed for generative models. Finally, we demonstrate that inspection networks, trained on synthesized samples, improve their accuracy in spotting real defective products.
Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri
ICIP2
2023 Anomaly-Based Intrusion Detection System for DDoS Attack with Deep Learning Techniques
abstract
The increasing number of connected devices is fostering a rising frequency of cyber attacks, with Distributed Denial of Service (DDoS) attacks among the most common.To counteract DDoS, companies and large organizations are increasingly deploying anomaly-based Intrusion Detection Systems (IDS), which detect attack patterns by analyzing differences in malicious network traffic against a baseline of legitimate traffic.To differentiate malicious and normal traffic, methods based on artificial intelligence and, in particular, Deep Learning (DL) are being increasingly considered, due to their ability to automatically learn feature representations for the different traffic types, without need of explicit programming or handcrafted feature extraction.In this paper, we propose a novel methodology for simulating an anomaly-based IDS based on adaptive DL by designing multiple DL models working with both binary and multi-label classification on multiple datasets with different degrees of complexity.To make the DL models adaptable to different conditions, we consider adaptive architectures obtained by automatically tuning the number of neurons for each situation.Results on publicly-available datasets confirm the validity of our proposed methodology, with DL models adapting to the different conditions by increasing the number of neurons on more complex datasets and achieving the highest accuracy in the binary classification configuration.anomaly-based IDSs work by establishing a baseline of "normal" network traffic and detecting "malicious" traffic and hence possible attacks when significant differences from the baseline are detected.Recent
Davide Agostinello, Angelo Genovese, Vincenzo Piuri
SECRYPT2
2022 Histokt: Cross Knowledge Transfer in Computational Pathology
abstract
The lack of well-annotated datasets in computational pathology (CPath) obstructs the application of deep learning techniques for classifying medical images. Many CPath workflows involve transferring learned knowledge between various image domains through transfer learning. Currently, most transfer learning research follows a model-centric approach, tuning network parameters to improve transfer results over few datasets. In this paper, we take a data-centric approach to the transfer learning problem and examine the existence of generalizable knowledge between histopathological datasets. First, we create a standardization workflow for aggregating existing histopathological data. We then measure inter-domain knowledge by training ResNet18 models across multiple histopathological datasets, and cross-transferring between them to determine the quantity and quality of innate shared knowledge. Additionally, we use weight distillation to share knowledge between models without additional training. We find that hard to learn, multi-class datasets benefit most from pretraining, and a two stage learning framework incorporating a large source domain such as ImageNet allows for better utilization of smaller datasets. Furthermore, we find that weight distillation enables models trained on purely histopathological features to outperform models using external natural image data.
Ryan Zhang, Jiadai Zhu, Mahdi S. Hosseini, Angelo Genovese, Lina Chen, Corwyn Rowsell, Savvas Damaskinos, Sonal Varma, Konstantinos N. Plataniotis
ICASSP5
2021 Acute Lymphoblastic Leukemia Detection Based on Adaptive Unsharpening and Deep Learning
abstract
Computer Aided Diagnosis (CAD) systems are increasingly utilizing image analysis and Deep Learning (DL) techniques, due to their high accuracy in several medical imaging fields, including the detection of Acute Lymphoblastic (or Lymphocytic) Leukemia (ALL) from peripheral blood samples. However, no method in the literature has specifically analyzed the focus quality of ALL images or proposed a technique for sharpening the samples in an adaptive way for the purpose of classification. To address this issue, in this paper we propose the first machine learning-based approach able to enhance blood sample images by an adaptive unsharpening method. The method uses image processing techniques and DL to normalize the radius of the cell, estimate the focus quality, adaptively improve the sharpness of the images, and then perform the classification. We evaluated the methodology on a public database of ALL images, considering several state-of-the-art CNNs to perform the classification, with results showing the validity of the proposed approach. For a complete reproducibility of the work, the source code is available at: http://iebil.di.unimi.it/cnnALL/index.htm.
Angelo Genovese, Mahdi S. Hosseini, Vincenzo Piuri, Konstantinos N. Plataniotis, Fabio Scotti
ICASSP1
2021 I-SOCIAL-DB: A labeled database of images collected from websites and social media for Iris recognition
Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, Sarvesh Vishwakarma
Image Vis. Comput.2
2020 Unsupervised Learning From Limited Available Data by β-NMF and Dual Autoencoder
abstract
Unsupervised Learning (UL) models are a class of Machine Learning (ML) which concerns with reducing dimensionality, data factorization, disentangling and learning the representations among the data. The UL models gain their popularity due to their abilities to learn without any predefined label, and they are able to reduce the noise and redundancy among the data samples. However, generalizing the UL models for different applications including image generation, compression, encoding, and recognition faces different challenges due to limited available data for learning, diversity, and complex dimensions. To overcome such challenges, we propose a partial learning procedure by utilizing the β-Non Negative Matrix Factorization (β-NMF), which maps the data into two complementary subspaces constituting generalized driven priors among the data. Moreover, we employ a dual-shallow Autoencoder (AE) to learn the subspaces separately or jointly for image reconstruction and visualization tasks, where our model performance shows superior results to the literary works when learning the model with a small amount of data and generalizing it for large-scale unseen data.
Mohanad Abukmeil, Stefano Ferrari, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
ICIP3
2020 A Decision Support System for Wind Power Production
abstract
Renewable energy production is constantly growing worldwide, and some countries produce a relevant percentage of their daily electricity consumption through wind energy. Therefore, decision support systems that can make accurate predictions of wind-based power production are of paramount importance for the traders operating in the energy market and for the managers in charge of planning the nonrenewable energy production. In this paper, we present a decision support system that can predict electric power production, estimate a variability index for the prediction, and analyze the wind farm (WF) production characteristics. The main contribution of this paper is a novel system for long-term electric power prediction based solely on the weather forecasts; thus, it is suitable for the WFs that cannot collect or manage the real-time data acquired by the sensors. Our system is based on neural networks and on novel techniques for calibrating and thresholding the weather forecasts based on the distinctive characteristics of the WF orography. We tuned and evaluated the proposed system using the data collected from two WFs over a two-year period and achieved satisfactory results. We studied different feature sets, training strategies, and system configurations before implementing this system for a player in the energy market. This company evaluated the power production prediction performance and the impact of our system at ten different WFs under real-world conditions and achieved a significant improvement with respect to their previous approach.
Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, Gianluca Sforza
IEEE Trans. Syst. Man Cybern. Syst.2
2019 Towards Explainable Face Aging with Generative Adversarial Networks
abstract
Generative Adversarial Networks (GAN) are being increasingly used to perform face aging due to their capabilities of automatically generating highly-realistic synthetic images by using an adversarial model often based on Convolutional Neural Networks (CNN). However, GANs currently represent black box models since it is not known how the CNNs store and process the information learned from data. In this paper, we propose the first method that deals with explaining GANs, by introducing a novel qualitative and quantitative analysis of the inner structure of the model. Similarly to analyzing the common genes in two DNA sequences, we analyze the common filters in two CNNs. We show that the GANs for face aging partially share their parameters with GANs trained for heterogeneous applications and that the aging transformation can be learned using general purpose image databases and a fine-tuning step. Results on public databases confirm the validity of our approach, also enabling future studies on similar models.
Angelo Genovese, Vincenzo Piuri, Fabio Scotti
ICIP1
2019 PalmNet: Gabor-PCA Convolutional Networks for Touchless Palmprint Recognition
abstract
Touchless palmprint recognition systems enable high-accuracy recognition of individuals through less-constrained and highly usable procedures that do not require the contact of the palm with a surface. To perform this recognition, methods based on local texture descriptors and convolutional neural networks (CNNs) are currently used to extract highly discriminative features while compensating for variations in scale, rotation, and illumination in biometric samples. In particular, the main advantage of CNN-based methods is their ability to adapt to biometric samples captured with heterogeneous devices. However, the current methods rely on either supervised training algorithms, which require class labels (e.g., the identities of the individuals) during the training phase, or filters pretrained on general-purpose databases, which may not be specifically suitable for palmprint data. To achieve a high-recognition accuracy with touchless palmprint samples captured using different devices while neither requiring class labels for training nor using pretrained filters, we introduce PalmNet, which is a novel CNN that uses a newly developed method to tune palmprint-specific filters through an unsupervised procedure based on Gabor responses and principal component analysis (PCA), not requiring class labels during training. PalmNet is a new method of applying Gabor filters in a CNN and is designed to extract highly discriminative palmprint-specific descriptors and to adapt to heterogeneous databases. We validated the innovative PalmNet on several palmprint databases captured using different touchless acquisition procedures and heterogeneous devices, and in all cases, a recognition accuracy greater than that of the current methods in this paper was obtained.
Angelo Genovese, Vincenzo Piuri, Konstantinos N. Plataniotis, Fabio Scotti
IEEE Trans. Inf. Forensics Secur.1
2019 3-D Granulometry Using Image Processing
abstract
Image-based methods for estimating the particle size distribution (granulometry) usually analyze two-dimensional (2-D) samples of particles disposed on a conveyor belt. Such approaches have to deal with occlusions and cannot evaluate the thickness of each particle. Three-dimensional (3-D) vision systems can reduce the acquisition constraints and speed up the quality control process. This paper proposes a novel 3-D vision system for analyzing the granulometry of falling particles. The system is designed to work in real time and to compute a partial 3-D reconstruction of the particle from a single pair of two-view images, which is then enhanced by using a neural-based technique. The validation of the proposed approach has been performed by considering three application scenarios for which the system achieved satisfactory accuracy and robustness.
Ruggero Donida Labati, Angelo Genovese, Enrique Muñoz Ballester, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Ind. Informatics2
2018 A novel pore extraction method for heterogeneous fingerprint images using Convolutional Neural Networks
Ruggero Donida Labati, Angelo Genovese, Enrique Muñoz Ballester, Vincenzo Piuri, Fabio Scotti
Pattern Recognit. Lett.2
2016 Towards touchless pore fingerprint biometrics: A neural approach
abstract
Touchless fingerprint recognition systems are being increasingly used for a fast, hygienic, and distortion-free recognition. However, due to the greater complexity of the algorithms required for processing touchless fingerprint samples, currently only Level 1 and Level 2 features are being used for recognition, and Level 3 features are used only in touch-based optical devices with about 1000 ppi resolution. In this paper, we propose the first innovative method in the literature able to extract Level 3 features, in particular sweat pores, from fingerprint images captured with a touchless acquisition using a commercial off-the-shelf camera. The method uses image processing algorithms to extract a set of candidate sweat pores. Then, computational intelligence techniques based on neural networks are used to learn the local features of the real pores, and select only the actual sweat pores from the set of candidate points. The results show the validity of the proposed methodology, with the majority of the pores correctly extracted, indicating that a touchless fingerprint recognition using Level 3 features is feasible.
Angelo Genovese, Enrique Muñoz Ballester, Vincenzo Piuri, Fabio Scotti, Gianluca Sforza
CEC1
2016 Toward Unconstrained Fingerprint Recognition: A Fully Touchless 3-D System Based on Two Views on the Move
abstract
Touchless fingerprint recognition systems do not require contact of the finger with any acquisition surface and thus provide an increased level of hygiene, usability, and user acceptability of fingerprint-based biometric technologies. The most accurate touchless approaches compute 3-D models of the fingertip. However, a relevant drawback of these systems is that they usually require constrained and highly cooperative acquisition methods. We present a novel, fully touchless fingerprint recognition system based on the computation of 3-D models. It adopts an innovative and less-constrained acquisition setup compared with other previously reported 3-D systems, does not require contact with any surface or a finger placement guide, and simultaneously captures multiple images while the finger is moving. To compensate for possible differences in finger placement, we propose novel algorithms for computing 3-D models of the shape of a finger. Moreover, we present a new matching strategy based on the computation of multiple touch-compatible images. We evaluated different aspects of the biometric system: acceptability, usability, recognition performance, robustness to environmental conditions and finger misplacements, and compatibility and interoperability with touch-based technologies. The proposed system proved to be more acceptable and usable than touch-based techniques. Moreover, the system displayed satisfactory accuracy, achieving an equal error rate of 0.06% on a dataset of 2368 samples acquired in a single session and 0.22% on a dataset of 2368 samples acquired over the course of one year. The system was also robust to environmental conditions and to a wide range of finger rotations. The compatibility and interoperability with touch-based technologies was greater or comparable to those reported in public tests using commercial touchless devices.
Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Syst. Man Cybern. Syst.2
2013 Wildfire Smoke Detection Using Computational Intelligence Techniques Enhanced With Synthetic Smoke Plume Generation
abstract
An early wildfire detection is essential in order to assess an effective response to emergencies and damages. In this paper, we propose a low-cost approach based on image processing and computational intelligence techniques, capable to adapt and identify wildfire smoke from heterogeneous sequences taken from a long distance. Since the collection of frame sequences can be difficult and expensive, we propose a virtual environment, based on a cellular model, for the computation of synthetic wildfire smoke sequences. The proposed detection method is tested on both real and simulated frame sequences. The results show that the proposed approach obtains accurate results.
Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Syst. Man Cybern. Syst.2
2012 Low-cost volume estimation by two-view acquisitions: A computational intelligence approach
abstract
The estimation of the volume occupied by an object is an important task in the fields of granulometry, quality control, and archaeology. An accurate and well know technique for the volume measurement is based on the Archimedes' principle. However, in many applications it is not possible to use this technique and faster contact-less techniques based on image processing or laser scanning should be adopted. In this work, we propose a low-cost approach for the volume estimation of different kinds of objects by using a two-view vision approach. The method first computes a reduced three-dimensional model from a single couple of images, then extracts a series of features from the obtained model. Lastly, the features are processed using a computational intelligence approach, which is able to learn the relation between the features and the volume of the captured object, in order to estimate the volume independently of its position and angle, and without computing a full three-dimensional model. Results show that the approach is feasible and can obtain an accurate volume estimation. Compared to the direct computation of the volume from the three-dimensional models, the approach is more accurate and also less dependent to the position and angle of the measured objects with respect to the cameras.
Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IJCNN2
2012 Quality measurement of unwrapped three-dimensional fingerprints: A neural networks approach
abstract
Traditional biometric systems based on the fingerprint characteristics acquire the biometric samples using touch-based sensors. Some recent researches are focused on the design of touch-less fingerprint recognition systems based on CCD cameras. Most of these systems compute three-dimensional fingertip models and then apply unwrapping techniques in order to obtain images compatible with biometric methods designed for images captured by touch-based sensors. Unwrapped images can present different problems with respect to the traditional fingerprint images. The most important of them is the presence of deformations of the ridge pattern caused by spikes or badly reconstructed regions in the corresponding three-dimensional models. In this paper, we present a neural-based approach for the quality estimation of images obtained from the unwrapping of three-dimensional fingertip models. The paper also presents different sets of features that can be used to evaluate the quality of fingerprint images. Experimental results show that the proposed quality estimation method has an adequate accuracy for the quality classification. The performances of the proposed method are also evaluated in a complete biometric system and compared with the ones obtained by a well-known algorithm in the literature, obtaining satisfactory results.
Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IJCNN2