VLDB 2026 Research / reviewers in the wild / expert
Fabio Scotti
dblp:91/5031
· DBLP profile ↗
49ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0002-4277-3701ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 10 since 2021Human-computer interaction and ubiquitous computing · 4Security and privacy · 3Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the relevance of patch-based extraction methods for monocular depth estimationabstractScene geometry estimation from images plays a key role in robotics, augmented reality, and autonomous systems. In particular, Monocular Depth Estimation (MDE) focuses on predicting depth using a single RGB image, avoiding the need for expensive sensors. State-of-the-art approaches use deep learning models for MDE while processing images as a whole, sub-optimally exploiting their spatial information. A recent research direction focuses on smaller image patches, as depth information varies across different regions of an image. This approach reduces model complexity and improves performance by capturing finer spatial details. From this perspective, we propose a novel warp patch-based extraction method which corrects perspective camera distortions, and employ it in tailored training and inference pipelines. Our experimental results show that our patch-based approach outperforms its full-image-trained counterpart and the classical crop patch-based extraction. With our technique, we obtain a general performance enhancements over recent state-of-the-art models. Code is available at https://github.com/AntonioFusillo/PatchMDE . • We propose a novel patch-based approach for monocular depth estimation. • Our method extracts patches from wide-aspect images, preserving camera parameters. • The designed patch-based inference outperforms full-image models in depth accuracy. • The proposed warp-based patch extraction is superior to patch cropping. • Our approach can wrap existing models, improving their performance. Pasquale Coscia, Antonio Fusillo, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
Image Vis. Comput. | 5 |
| 2026 | Useg-PanoDepth:Unified $360^{\circ }$ Depth Estimation for Indoor and Outdoor Scenes With Semantic AssistanceabstractIn complex$360^{\circ }$scenes, depth estimation is challenging for small objects and the depth of object boundaries, which cannot be effectively solved with existing works.$360^{\circ }$depth estimation is unable to produce uniform depth estimate findings in both indoor and outdoor settings due to the datasets. In this paper, the Useg-PanoDepth and PanoDepth dataset is proposed to improve the above problems effectively. The Diagonal-aware Attention Module (DAM) effectively estimates small objects in complex scenes. Enhanced Boundary Module (EBM), for enhancing boundary information,can also effectively solve the problem of depth unification of indoor and outdoor scenes. Extensive experiments on our constructed PanoDepth dataset, Useg-PanoDepth achieves SOTA results. The Relative accuracy (deltahttps://github.com/xjh6/Useg-PanoDepth. Qingling Chang, Jingheng Xu, Yan Cui 0011, Yikui Zhai, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Multim. | 8 |
| 2025 | OneN: Guided attention for natively-explainable anomaly detectionabstractIn industrial computer vision applications, anomaly detection (AD) is a critical task for ensuring product quality and system reliability. However, many existing AD systems follow a modular design that decouples classification from detection and localization tasks. Although this separation simplifies model development, it often limits generalizability and reduces practical effectiveness in real-world scenarios. Deep neural networks offer strong potential for unified solutions. Nonetheless, most current approaches still treat detection, localization and classification as separate components, hindering the development of more integrated and efficient AD pipelines. To bridge this gap, we propose OneN (One Network), a unified architecture that performs detection, localization, and classification within a single framework. Our approach distills knowledge from a high-capacity convolutional neural network (CNN) into an attention-based architecture trained under varying levels of supervision. The resulting attention maps act as interpretable pseudo-segmentation masks, enabling accurate localization of anomalous regions. To further enhance localization quality, we introduce a progressive focal loss that guides attention maps at each layer to focus on critical features. We validate our method through extensive experiments on both standardized and custom-defined industrial benchmarks. Even under weak supervision, it improves performance, reduces annotation effort, and facilitates scalable deployment in industrial environments. Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
Image Vis. Comput. | 4 |
| 2025 | PBSD-Net: Prismatic Battery Surface Defect Detection via Sliding Slice Amplification and Shunted Dynamic Snake ConvolutionabstractAutomatically detecting surface defects in prismatic battery is crucial for ensuring quality meets established standards. Traditional methods face challenges in accurately identifying these defects due to their minute and varied shapes and high density of distribution. To address these issues, we propose an innovative network for prismatic battery surface defect (PBSD-Net), which employs shunted dynamic snake convolution and focal modulation to detect surface defects in prismatic battery. This network is integrated into the 2D-AOI system. Firstly, we introduce sliding slice amplification (SSA) as a training strategy to enhance the network’s ability to recognize densely clustered tiny defects. Secondly, we develop a novel method using the shunted dynamic snake convolution (SDSC) module and focal modulation (FM) to improve the extraction of deformation features, thereby addressing complex and sporadically scattered surface defects. By integrating the SDSC module and FM mechanism, the receptive field of the defect feature extraction network is expanded, enabling the acquisition of comprehensive defect edge features. Additionally, we introduce the quality focal loss (QFL) function to effectively tackle the issue of imbalanced sample types. Experimental results on the PBSD-RGB dataset demonstrate that our method achieves a mAP@50 of 85.8%, representing an improvement of approximately 7.7% over the baseline network. We have applied the PBSD-Net to an automatic defect detection system in a well-known battery production company. This enhancement significantly boosts the accuracy of surface defect detection in prismatic battery. The relevant code is at the https://github.com/yikuizhai/PBSD-Net. Ying Xu 0005, Bo Li 0165, Yikui Zhai, Feng Ke, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans Autom. Sci. Eng. | 9 |
| 2025 | GMTNet: Dense Object Detection via Global Dynamically Matching Transformer NetworkabstractIn recent years, object detection models have been extensively applied across various industries, leveraging learned samples to recognize and locate objects. However, industrial environments present unique challenges, including complex backgrounds, dense object distributions, object stacking, and occlusion. To address these challenges, we propose the Global Dynamic Matching Transformer Network (GMTNet). GMTNet partitions images into blocks and employs a sliding window approach to capture information from each block and their interrelationships, mitigating background interference while acquiring global information for dense object recognition. By reweighting key-value pairs in multi-scale feature maps, GMTNet enhances global information relevance and effectively handles occlusion and overlap between objects. Furthermore, we introduce a dynamic sample matching method to tackle the issue of excessive candidate boxes in dense detection tasks. This method adaptively adjusts the number of matched positive samples according to the specific detection task, enabling the model to reduce the learning of irrelevant features and simplify post-processing. Experimental results demonstrate that GMTNet excels in dense detection tasks and outperforms current mainstream algorithms. The code will be available athttp://github.com/yikuizhai/GMTNet. Chaojun Dong, Chengxuan Wang, Yikui Zhai, Ye Li 0002, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | AEGL-Net: Adaptive Multiscale Global-Local Feature Fusion Network for Remote Sensing Change DetectionabstractWith the rapid advancements in deep learning technology, the field of remote sensing change detection (RSCD) has witnessed significant improvements and innovations. In this context, bitemporal image processing, using features directly extracted by the backbone for subsequent fusion operations, may be obstructed by external environmental factors, potentially limiting the effective capture of complex feature variations. Moreover, overlooking local features during the fusion of bitemporal features can significantly affect the final detection results. As a result, achieving accurate change detection (CD) still encounters various challenges. To tackle these issues, this paper proposes a CD network (AEGL-Net) with Adaptive Multiscale Enhancement (AME) and Global-Local Feature Fusion (GLFF) modules. First, AME enhances features at each stage of backbone extraction through an adaptive strategy, balancing the enhancement of semantic information and texture details. Then, GLFF is used to fuse the bitemporal image features, which enhances the modeling of global dependencies while also fusing shared and context-aware weights to enhance the local features. Finally, the merged features are fed into the decoder to generate precise change maps. Experiments conducted with four open RSCD datasets (LEVIR-CD, S2Looking, SYSU-CD, and UAV-CD) demonstrate that our proposed AEGL-Net outperforms ten state-of-the-art models in the RSCD field. Our code is available at https://github.com/yikuizhai/AEGL-Net. Zilu Ying, Yikui Zhai, Hufei Zhu, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Spatial Reconstruction and Joint Training in Transformer Network for Cross-Domain Remote Sensing Images Semantic SegmentationabstractRecently, Unsupervised Domain Adaptation (UDA) methods have attracted considerable attention in Remote Sensing Images (RSI) semantic segmentation. However, cross-domain RSI exhibit diverse scales, imbalanced distributions within domains, and significant inter-domain variations. In response to these challenges, we combine Spatial reconstruction and Joint training with the Transformer Network (SJT-Net). This framework introduces a spatial reconstruction method to address the issue of inconsistent ground sampling distances in cross domain RSI, which is rarely considered in existing approaches. Transferring domain knowledge at a similar spatial scale improves the spatial representation ability of UDA models. Unlike traditional adversarial training using ResNet for feature extraction, the SJT-Net employs Segformer, which enhances the model’s ability to capture in-class features across domains and improves global dependency modeling. Transmitting these refined features to the discriminator allows for more precise feature-level domain alignment. To enhance feature decoding, an interactive global-local decoder is constructed to efficiently capture both global relationships and local details of landform objects. Our framework leverages adversarial training to generate highly confident model weights and pseudo-labels for self-training in the target domain. Through iterative updates, the model’s generalization capability is gradually improved, eventually achieving optimal segmentation performance. Experimental results demonstrate that SJT-Net outperforms current UDA approaches and accomplishes state-of-the-art (SOTA) segmentation accuracy. The repository can be accessed at https://github.com/AnsonD0820/SJT-Net. Jun-Ying Zeng, Senyao Deng, Yikui Zhai, Xudong Jia 0001, Chuanbo Qin, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | CLIP-Vision Guided Few-Shot Metal Surface Defect RecognitionabstractMetal surface defect recognition (MSDR) based on deep learning encounters the challenge of few-shot expert-labeled data. In this study, we proposed a CLIP-vision guided self supervised learning (CVGSSL) framework for representation learning of unlabeled data, completing MSDR using few-shot labeled data. This framework initially generates rich and diverse representation information through multiple CLIP-Vs to ensure effective SSL pretraining, followed by the design of an MLP-adapter to distill knowledge and adapt these representations to recognition tasks. In addition, we constructed a self-constrained loss to address the inherent problem of intraclass and interclass distance ambiguity that causes the representation to fall into an equivocal decision margin. Following label-free pretraining of CVGSSL, the downstream model adapts to one-shot to four-shot defect recognition tasks through fine-tuning. Experimental results demonstrate that CVGSSL outperforms state-of-the-art SSL methods across three public metal surface defect datasets, with the efficacy of the approach validated through extensive ablation experiments. Tianlei Wang, Zeliang Li, Ying Xu 0005, Yikui Zhai, Xiaofen Xing, Kailing Guo, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Ind. Informatics | 10 |
| 2024 | Features Disentanglement For Explainable Convolutional Neural NetworksabstractExplainable methods for understanding deep neural networks are currently being employed for many visual tasks and provide valuable insights about their decisions. While post-hoc visual explanations offer easily understandable human cues behind neural networks’ decision-making processes, comparing their outcomes still remains challenging. Furthermore, balancing the performance-explainability trade-off could be a time-consuming process and require a deep domain knowledge. In this regard, we propose a novel auxiliary module, built upon convolutional-based encoders, which acts on the final layers of convolutional neural networks (CNNs) to learn orthogonal feature maps with a more discriminative and explainable power. This module is trained via a disentangle loss which specifically aims to decouple the object from the background in the input image. To quantitatively assess its impact on standard CNNs, and compare the quality of the resulting visual explanations, we employ metrics specifically designed for semantic segmentation tasks. These metrics rely on bounding-box annotations that may accompany image classification (or recognition) datasets, allowing us to compare both ground-truth and predicted regions. Finally, we explore the impact of various self-supervised pre-training strategies, due to their positive influence on vision tasks, and assess their effectiveness on our considered metrics. Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri |
ICIP | 3 |
| 2024 | A decision support system for acute lymphoblastic leukemia detection based on explainable artificial intelligence
Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
Image Vis. Comput. | 3 |
| 2024 | DGMA2-Net: A Difference-Guided Multiscale Aggregation Attention Network for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) focuses on identifying regions that have undergone changes between two remote sensing images captured at different times. Recently, convolutional neural networks (CNNs) have shown promising results in the challenging task of RSCD. However, these methods do not efficiently fuse bitemporal features and extract useful information that is beneficial to subsequent RSCD tasks. In addition, they did not consider multilevel feature interactions in feature aggregation and ignore relationships between difference features and bitemporal features, which thus affects the RSCD results. To address the above problems, a difference-guided multiscale aggregation attention network, DGMA2-Net, is developed. Bitemporal features at different levels are extracted through a Siamese convolutional network and a multiscale difference fusion module (MDFM) is then created to fuse bitemporal features and extract, in a multiscale manner, difference features containing rich contextual information. After the MDFM treatment, two difference aggregation modules (DAMs) are used to aggregate difference features at different levels for multilevel feature interactions. The features through DAMs are sent to the difference-enhanced attention modules (DEAMs) to strengthen the connections between bitemporal features and difference features and further refine change features. Finally, refined change features are superimposed from deep to shallow and a change map is produced. In validating the effectiveness of DGMA2-Net, a series of experiments are conducted on three public RSCD benchmark datasets (LEVIR-CD, BCDD, and SYSU-CD). The experimental results demonstrate that DGMA2-Net surpasses the current eight state-of-the-art methods in RSCD. Our code is released at https://github.com/yikuizhai/DGMA2-Net. Zilu Ying, Zijun Tan, Yikui Zhai, Xudong Jia 0001, Wenba Li, Jun-Ying Zeng, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2024 | DS-HyFA-Net: A Deeply Supervised Hybrid Feature Aggregation Network With Multiencoders for Change Detection in High-Resolution ImageryabstractWith the advancement of deep learning (DL) technologies, remarkable progress has been achieved in change detection (CD). Existing DL-based methods primarily focus on the discrepancy in bitemporal images, while overlooking the commonality in bitemporal images. However, one of the reasons hindering the improvement of CD performance is the inadequate utilization of image information. To address the above issue, we propose a Deeply Supervised Hybrid Feature Aggregation Network (DS-HyFA-Net). This network predicts changes by integrating the distinctness and the commonality in bitemporal images. Specifically, the DS-HyFA-Net primarily consists of a set of encoders and a Hybrid Feature Aggregation (HyFA) module. It uses a Siamese encoder (or Encoder I) and a specialized encoder (or Encoder II) to extract distinct and common features (CFs) in bitemporal images, respectively. The HyFA module efficiently aggregates distinct and common features (or hybrid features) and generates a change map using a predictor. In addition, a common feature learning strategy (CFLS) is introduced, based on deeply supervised (DS) techniques, to guide Encoder II in learning CFs. Experimental results on three well-recognized datasets demonstrate the effectiveness of the innovative DS-HyFA-Net, achieving F1-Scores of 93.33% on WHU-CD, 90.98% on LEVIR-CD, and 81.14% on SYSU-CD. Our code is available athttps://github.com/yikuizhai/DS-HyFA-Net. Zilu Ying, Tingfeng Xian, Yikui Zhai, Xudong Jia 0001, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2024 | Efficient Adjacent Feature Harmonizer Network With UAV-CD+ Dataset for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) aims to identify changes within bi-temporal registered images. However, existing deep learning (DL)-based RSCD networks often suffer from large numbers of parameters, high computational complexity, and low inference speed, making it challenging to achieve efficient inference in real-world deployments. In addition, current models lack robust feature-fitting capabilities, necessitating the development of an efficient and powerful RSCD model to address this issue. Therefore, we propose a novel RSCD network named efficient adjacent feature harmonizer network (EAFH-Net) with fast computational speed and lightweight design. It is based on MobileNetV2, considering that change maps of different sizes contain temporal information of bitemporal features and spatial information at various scales, we introduce a multiscale feature neighbor fusion module (MFNFM) to address the lack of interaction between sophisticated-level and elementary-level features, and spatial and channel feature harmonizer module (SCFHM) to harmonize the spatiotemporal information of the change maps. Moreover, data-driven DL algorithms face another challenge due to insufficient granularity and the need for more practical datasets. Therefore, we present unmanned aerial vehicle (UAV)-CD+, a dataset comprising 2002 pairs of bi-temporal UAV low-altitude images, each sized at$1024\times 1024$. We performed experiments on three publicly accessible datasets in conjunction with UAV-CD+, comparing the results with other state-of-the-art (SOTA) methods. EAFH-Net attains the utmost precision, obtaining 91.74% on LEVIR-CD, 84.28% on SYSU-CD, 95.07% on WHU-CD, 79.12% on CLCD, and 70.12% on UAV-CD+. We have our model code available at the following link:https://github.com/yikuizhai/UCSFH-Net. Yikui Zhai, Hongsheng Zhang 0001, Tingfeng Xian, Ying Xu 0005, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2024 | Learning to Count Arbitrary Industrial Manufacturing WorkpiecesabstractMan-made workpiece counting is a routine job for manufactory workers; however, this is an error-prone task. In this article, we are interested in detecting and counting arbitrary workpieces in industrial manufacturing. Therefore, we construct a comprehensive and large-scale open-world public benchmark dataset for workpiece counting, called workpiece counting dataset, which includes 121 475 instances of workpieces from 351 different categories. We also propose a novel method for workpiece detection and counting, named two-stage workpiece counting network. The first stage of the network is to develop a class-agnostic detector to localize each workpiece instance, followed by the second stage to employ an unsupervised deep clustering strategy with the backbone network pretrained in a workpiece convolutional autoencoder for decision boundary prediction, achieving workpiece clustering under unknownKvalues. Finally, our experiments show that the proposed method outperforms current mainstream methods, greatly enhancing the efficiency of factory operations. Yikui Zhai, Feng Ke, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | Large-Scale High-Altitude UAV-Based Vehicle Detection via Pyramid Dual Pooling Attention Path Aggregation NetworkabstractUAVs can collect vehicle data in high-altitude scenes, playing a significant role in intelligent urban management due to their wide of view. Nevertheless, the current datasets for UAV-based vehicle detection are acquired at altitude below 150 meters. This contrasts with the data perspective obtained from high-altitude scenes, potentially leading to incongruities in data distribution. Consequently, it is challenging to apply these datasets effectively in high-altitude scenes, and there is an ongoing obstacle. To resolve this challenge, we developed a comprehensive vehicle dataset named LH-UAV-Vehicle, specifically collected at flight altitudes ranging from 250 to 400 meters. Collecting data at higher flight altitudes offers a broader perspective, but it concurrently introduces complexity and diversity in the background, which consequently impacts vehicle localization and recognition accuracy. In response, we proposed the pyramid dual pooling attention path aggregation network (PDPA-PAN), an innovative framework that improves detection performance in high-altitude scenes by combining spatial and semantic information. Object attention integration in both spatial and channel dimensions is aimed by the pyramid dual pooling attention module (PDPAM), which is achieved through the parallel integration of two distinct attention mechanisms. Furthermore, we have individually developed the pyramid pooling attention module (PPAM) and the dual pooling attention module (DPAM). The PPAM emphasizes channel attention, while the DPAM prioritizes spatial attention. This design aims to enhance vehicle information and suppress background interference more effectively. Extensive experiments conducted on the LH-UAV-Vehicle conclusively demonstrate the efficacy of the proposed vehicle detection method. Our code and dataset can be found at https://github.com/yikuizhai/PDPA-PAN. Zilu Ying, Yikui Zhai, Hao Quan 0002, Wenba Li, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2023 | Adversarial Defect Synthesis for Industrial Products in Low Data RegimeabstractSynthetic defect generation is an important aid for advanced manufacturing and production processes. Industrial scenarios rely on automated image-based quality control methods to avoid time-consuming manual inspections and promptly identify products not complying with specific quality standards. However, these methods show poor performance in the case of ill-posed low-data training regimes, and the lack of defective samples, due to operational costs or privacy policies, strongly limits their large-scale applicability.To overcome these limitations, we propose an innovative architecture based on an unpaired image-to-image (I2I) translation model to guide a transformation from a defect-free to a defective domain for common industrial products and propose simultaneously localizing their synthesized defects through a segmentation mask. As a performance evaluation, we measure image similarity and variability using standard metrics employed for generative models. Finally, we demonstrate that inspection networks, trained on synthesized samples, improve their accuracy in spotting real defective products. Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri |
ICIP | 3 |
| 2023 | Efficient IoT Big Data Streaming With Deep-Learning-Enabled DynamicsabstractInternet of Medical Things (IoMT) is igniting many emerging smart health applications, by continuously streaming the big data for data-driven innovations. One critical obstacle in IoMT big data is the power hungriness of long-term data transmission. Targeting this challenge, we propose a novel framework called, IoMT big-data Bayesian-backward deep-encoder learning (IBBD), which mines deep autoencoder (AE) configurations for data sparsification and determines optimal tradeoffs between information loss and power overhead. More specifically, the IBBD framework leverages an additional external Bayesian-backward loop that recommends AE configurations, on top of a traditional deep learning loop that executes and evaluate the AE quality. The IBBD recommendation is based on confidence to further minimize the regularized metrics that quantify the quality of AE configurations, and it further leverages regularization techniques to allow adjusting error–power tradeoffs in the mining process. We have conducted thorough experiments on a cardiac data streaming application and demonstrated the superiority of IBBD over the common practices such as discrete wavelet transform, and we have further generalized IBBD through validating the optimal AE configurations determined on one user to other users. This study is expected to greatly advance IoMT big data streaming practices toward precision medicine. Junhua Wong, Vincenzo Piuri, Fabio Scotti, Qingxue Zhang |
IEEE Internet Things J. | 3 |
| 2023 | MultiCardioNet: Interoperability between ECG and PPG biometricsabstractCompared to other well-known biometric technologies based on physiological traits (e.g., fingerprint, iris, and face), heart biometrics are more robust to presentation attacks and are particularly suitable for continuous/periodic recognition.Most studies on heart biometrics concern electrocardiogram (ECG) and photoplethysmogram (PPG).While the reported results are encouraging, to the best of our knowledge, no studies have been conducted on the interoperability between ECG and PPG biometrics.We present a novel method that is capable of performing single-domain and multiple-domain identity verifications for ECG and PPG signals, providing interoperability between the heterogeneous cardiac signals.Our method does not require the computation of any reference/fiducial point and uses a compact representation of the given signals.We propose MultiCardioNet, a novel Siamese neural network trained by using an ad hoc learning algorithm.MultiCardioNet computes a similarity score between two spectrogram-based representations of cardiac signals.Our learning algorithm iteratively computes a balanced subset of genuine and impostor pairs during the training epochs.We performed experiments on a dataset containing 1,008 pairs of ECG and PPG samples, obtaining accuracy comparable to that of the state-of-the-art methods for single-domain scenarios and demonstrating only a relatively small performance decrease in the multiple-domain scenario. Ruggero Donida Labati, Vincenzo Piuri, Francesco Rundo, Fabio Scotti |
Pattern Recognit. Lett. | 4 |
| 2022 | Photoplethysmographic biometrics: A comprehensive survey
Ruggero Donida Labati, Vincenzo Piuri, Francesco Rundo, Fabio Scotti |
Pattern Recognit. Lett. | 4 |
| 2022 | Weakly Contrastive Learning via Batch Instance Discrimination and Feature Clustering for Small Sample SAR ATRabstractIn recent years, impressive performance of deep learning technology has been recognized in synthetic aperture radar (SAR) automatic target recognition (ATR). Since a large amount of annotated data are required in this technique, it poses a trenchant challenge to the issue of obtaining a high recognition rate through less labeled data. To overcome this problem, inspired by the contrastive learning, we proposed a novel framework named batch instance discrimination and feature clustering (BIDFC). In this framework, different from that of the objective of general contrastive learning methods, embedding distance between samples should be moderate because of the high similarity between samples in the SAR images. Consequently, our flexible framework is equipped with adjustable distance between embedding, which we term as weakly contrastive learning. Technically, instance labels are assigned to the unlabeled data in per batch, and random augmentation and training are performedfewtimes on these augmented data. Meanwhile, a novel dynamic-weighted variance loss (DWV loss) function is also posed to cluster the embedding of enhanced versions for each sample. The experimental results on the moving and stationary target acquisition and recognition (MSTAR) database indicate a 91.25% classification accuracy of our method fine-tuned on only 3.13% training data. Even though a linear evaluation is performed on the same training data, the accuracy can still reach 90.13%. We also verified the effectiveness of BIDFC in OpenSarShip database, indicating that our method can be generalized to other data sets. Our code is available at:https://github.com/Wenlve-Zhou/BIDFC-master. Yikui Zhai, Wenlve Zhou, Bing Sun 0002, Jingwen Li 0003, Qirui Ke, Zilu Ying, Junying Gan, Chaoyun Mai, Ruggero Donida Labati, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 11 |
| 2021 | Acute Lymphoblastic Leukemia Detection Based on Adaptive Unsharpening and Deep LearningabstractComputer Aided Diagnosis (CAD) systems are increasingly utilizing image analysis and Deep Learning (DL) techniques, due to their high accuracy in several medical imaging fields, including the detection of Acute Lymphoblastic (or Lymphocytic) Leukemia (ALL) from peripheral blood samples. However, no method in the literature has specifically analyzed the focus quality of ALL images or proposed a technique for sharpening the samples in an adaptive way for the purpose of classification. To address this issue, in this paper we propose the first machine learning-based approach able to enhance blood sample images by an adaptive unsharpening method. The method uses image processing techniques and DL to normalize the radius of the cell, estimate the focus quality, adaptively improve the sharpness of the images, and then perform the classification. We evaluated the methodology on a public database of ALL images, considering several state-of-the-art CNNs to perform the classification, with results showing the validity of the proposed approach. For a complete reproducibility of the work, the source code is available at: http://iebil.di.unimi.it/cnnALL/index.htm. Angelo Genovese, Mahdi S. Hosseini, Vincenzo Piuri, Konstantinos N. Plataniotis, Fabio Scotti |
ICASSP | 5 |
| 2021 | I-SOCIAL-DB: A labeled database of images collected from websites and social media for Iris recognition
Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, Sarvesh Vishwakarma |
Image Vis. Comput. | 4 |
| 2020 | Unsupervised Learning From Limited Available Data by β-NMF and Dual AutoencoderabstractUnsupervised Learning (UL) models are a class of Machine Learning (ML) which concerns with reducing dimensionality, data factorization, disentangling and learning the representations among the data. The UL models gain their popularity due to their abilities to learn without any predefined label, and they are able to reduce the noise and redundancy among the data samples. However, generalizing the UL models for different applications including image generation, compression, encoding, and recognition faces different challenges due to limited available data for learning, diversity, and complex dimensions. To overcome such challenges, we propose a partial learning procedure by utilizing the β-Non Negative Matrix Factorization (β-NMF), which maps the data into two complementary subspaces constituting generalized driven priors among the data. Moreover, we employ a dual-shallow Autoencoder (AE) to learn the subspaces separately or jointly for image reconstruction and visualization tasks, where our model performance shows superior results to the literary works when learning the model with a small amount of data and generalizing it for large-scale unseen data. Mohanad Abukmeil, Stefano Ferrari, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
ICIP | 5 |
| 2020 | Weakly supervised facial expression recognition via transferred DAL-CNN and active incremental learning
Ying Xu 0005, Yikui Zhai, Junying Gan, Jun-Ying Zeng, He Cao, Fabio Scotti, Vincenzo Piuri, Ruggero Donida Labati |
Soft Comput. | 7 |
| 2020 | A Decision Support System for Wind Power ProductionabstractRenewable energy production is constantly growing worldwide, and some countries produce a relevant percentage of their daily electricity consumption through wind energy. Therefore, decision support systems that can make accurate predictions of wind-based power production are of paramount importance for the traders operating in the energy market and for the managers in charge of planning the nonrenewable energy production. In this paper, we present a decision support system that can predict electric power production, estimate a variability index for the prediction, and analyze the wind farm (WF) production characteristics. The main contribution of this paper is a novel system for long-term electric power prediction based solely on the weather forecasts; thus, it is suitable for the WFs that cannot collect or manage the real-time data acquired by the sensors. Our system is based on neural networks and on novel techniques for calibrating and thresholding the weather forecasts based on the distinctive characteristics of the WF orography. We tuned and evaluated the proposed system using the data collected from two WFs over a two-year period and achieved satisfactory results. We studied different feature sets, training strategies, and system configurations before implementing this system for a player in the energy market. This company evaluated the power production prediction performance and the impact of our system at ten different WFs under real-world conditions and achieved a significant improvement with respect to their previous approach. Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, Gianluca Sforza |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | Towards Explainable Face Aging with Generative Adversarial NetworksabstractGenerative Adversarial Networks (GAN) are being increasingly used to perform face aging due to their capabilities of automatically generating highly-realistic synthetic images by using an adversarial model often based on Convolutional Neural Networks (CNN). However, GANs currently represent black box models since it is not known how the CNNs store and process the information learned from data. In this paper, we propose the first method that deals with explaining GANs, by introducing a novel qualitative and quantitative analysis of the inner structure of the model. Similarly to analyzing the common genes in two DNA sequences, we analyze the common filters in two CNNs. We show that the GANs for face aging partially share their parameters with GANs trained for heterogeneous applications and that the aging transformation can be learned using general purpose image databases and a fine-tuning step. Results on public databases confirm the validity of our approach, also enabling future studies on similar models. Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
ICIP | 3 |
| 2019 | Non-ideal iris segmentation using Polar Spline RANSAC and illumination compensation
Ruggero Donida Labati, Enrique Muñoz Ballester, Vincenzo Piuri, Arun Ross, Fabio Scotti |
Comput. Vis. Image Underst. | 5 |
| 2019 | Deep-ECG: Convolutional Neural Networks for ECG biometric recognition
Ruggero Donida Labati, Enrique Muñoz Ballester, Vincenzo Piuri, Roberto Sassi, Fabio Scotti |
Pattern Recognit. Lett. | 5 |
| 2019 | PalmNet: Gabor-PCA Convolutional Networks for Touchless Palmprint RecognitionabstractTouchless palmprint recognition systems enable high-accuracy recognition of individuals through less-constrained and highly usable procedures that do not require the contact of the palm with a surface. To perform this recognition, methods based on local texture descriptors and convolutional neural networks (CNNs) are currently used to extract highly discriminative features while compensating for variations in scale, rotation, and illumination in biometric samples. In particular, the main advantage of CNN-based methods is their ability to adapt to biometric samples captured with heterogeneous devices. However, the current methods rely on either supervised training algorithms, which require class labels (e.g., the identities of the individuals) during the training phase, or filters pretrained on general-purpose databases, which may not be specifically suitable for palmprint data. To achieve a high-recognition accuracy with touchless palmprint samples captured using different devices while neither requiring class labels for training nor using pretrained filters, we introduce PalmNet, which is a novel CNN that uses a newly developed method to tune palmprint-specific filters through an unsupervised procedure based on Gabor responses and principal component analysis (PCA), not requiring class labels during training. PalmNet is a new method of applying Gabor filters in a CNN and is designed to extract highly discriminative palmprint-specific descriptors and to adapt to heterogeneous databases. We validated the innovative PalmNet on several palmprint databases captured using different touchless acquisition procedures and heterogeneous devices, and in all cases, a recognition accuracy greater than that of the current methods in this paper was obtained. Angelo Genovese, Vincenzo Piuri, Konstantinos N. Plataniotis, Fabio Scotti |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | 3-D Granulometry Using Image ProcessingabstractImage-based methods for estimating the particle size distribution (granulometry) usually analyze two-dimensional (2-D) samples of particles disposed on a conveyor belt. Such approaches have to deal with occlusions and cannot evaluate the thickness of each particle. Three-dimensional (3-D) vision systems can reduce the acquisition constraints and speed up the quality control process. This paper proposes a novel 3-D vision system for analyzing the granulometry of falling particles. The system is designed to work in real time and to compute a partial 3-D reconstruction of the particle from a single pair of two-view images, which is then enhanced by using a neural-based technique. The validation of the proposed approach has been performed by considering three application scenarios for which the system achieved satisfactory accuracy and robustness. Ruggero Donida Labati, Angelo Genovese, Enrique Muñoz Ballester, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Ind. Informatics | 5 |
| 2018 | A novel pore extraction method for heterogeneous fingerprint images using Convolutional Neural Networks
Ruggero Donida Labati, Angelo Genovese, Enrique Muñoz Ballester, Vincenzo Piuri, Fabio Scotti |
Pattern Recognit. Lett. | 5 |
| 2017 | SAR Automatic Target Recognition Based on Deep Convolutional Neural Network
Ying Xu 0005, Kaipin Liu, Zilu Ying, Lijuan Shang, Yikui Zhai, Vincenzo Piuri, Fabio Scotti |
ICIG (3) | 8 |
| 2017 | Deep Convolutional Neural Network for Facial Expression Recognition
Yikui Zhai, Jun-Ying Zeng, Vincenzo Piuri, Fabio Scotti, Zilu Ying, Ying Xu 0005, Junying Gan |
ICIG (1) | 5 |
| 2016 | Towards touchless pore fingerprint biometrics: A neural approachabstractTouchless fingerprint recognition systems are being increasingly used for a fast, hygienic, and distortion-free recognition. However, due to the greater complexity of the algorithms required for processing touchless fingerprint samples, currently only Level 1 and Level 2 features are being used for recognition, and Level 3 features are used only in touch-based optical devices with about 1000 ppi resolution. In this paper, we propose the first innovative method in the literature able to extract Level 3 features, in particular sweat pores, from fingerprint images captured with a touchless acquisition using a commercial off-the-shelf camera. The method uses image processing algorithms to extract a set of candidate sweat pores. Then, computational intelligence techniques based on neural networks are used to learn the local features of the real pores, and select only the actual sweat pores from the set of candidate points. The results show the validity of the proposed methodology, with the majority of the pores correctly extracted, indicating that a touchless fingerprint recognition using Level 3 features is feasible. Angelo Genovese, Enrique Muñoz Ballester, Vincenzo Piuri, Fabio Scotti, Gianluca Sforza |
CEC | 4 |
| 2016 | Toward Unconstrained Fingerprint Recognition: A Fully Touchless 3-D System Based on Two Views on the MoveabstractTouchless fingerprint recognition systems do not require contact of the finger with any acquisition surface and thus provide an increased level of hygiene, usability, and user acceptability of fingerprint-based biometric technologies. The most accurate touchless approaches compute 3-D models of the fingertip. However, a relevant drawback of these systems is that they usually require constrained and highly cooperative acquisition methods. We present a novel, fully touchless fingerprint recognition system based on the computation of 3-D models. It adopts an innovative and less-constrained acquisition setup compared with other previously reported 3-D systems, does not require contact with any surface or a finger placement guide, and simultaneously captures multiple images while the finger is moving. To compensate for possible differences in finger placement, we propose novel algorithms for computing 3-D models of the shape of a finger. Moreover, we present a new matching strategy based on the computation of multiple touch-compatible images. We evaluated different aspects of the biometric system: acceptability, usability, recognition performance, robustness to environmental conditions and finger misplacements, and compatibility and interoperability with touch-based technologies. The proposed system proved to be more acceptable and usable than touch-based techniques. Moreover, the system displayed satisfactory accuracy, achieving an equal error rate of 0.06% on a dataset of 2368 samples acquired in a single session and 0.22% on a dataset of 2368 samples acquired over the course of one year. The system was also robust to environmental conditions and to a wide range of finger rotations. The compatibility and interoperability with touch-based technologies was greater or comparable to those reported in public tests using commercial touchless devices. Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2013 | Wildfire Smoke Detection Using Computational Intelligence Techniques Enhanced With Synthetic Smoke Plume GenerationabstractAn early wildfire detection is essential in order to assess an effective response to emergencies and damages. In this paper, we propose a low-cost approach based on image processing and computational intelligence techniques, capable to adapt and identify wildfire smoke from heterogeneous sequences taken from a long distance. Since the collection of frame sequences can be difficult and expensive, we propose a virtual environment, based on a cellular model, for the computation of synthetic wildfire smoke sequences. The proposed detection method is tested on both real and simulated frame sequences. The results show that the proposed approach obtains accurate results. Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2012 | Low-cost volume estimation by two-view acquisitions: A computational intelligence approachabstractThe estimation of the volume occupied by an object is an important task in the fields of granulometry, quality control, and archaeology. An accurate and well know technique for the volume measurement is based on the Archimedes' principle. However, in many applications it is not possible to use this technique and faster contact-less techniques based on image processing or laser scanning should be adopted. In this work, we propose a low-cost approach for the volume estimation of different kinds of objects by using a two-view vision approach. The method first computes a reduced three-dimensional model from a single couple of images, then extracts a series of features from the obtained model. Lastly, the features are processed using a computational intelligence approach, which is able to learn the relation between the features and the volume of the captured object, in order to estimate the volume independently of its position and angle, and without computing a full three-dimensional model. Results show that the approach is feasible and can obtain an accurate volume estimation. Compared to the direct computation of the volume from the three-dimensional models, the approach is more accurate and also less dependent to the position and angle of the measured objects with respect to the cameras. Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IJCNN | 4 |
| 2012 | Quality measurement of unwrapped three-dimensional fingerprints: A neural networks approachabstractTraditional biometric systems based on the fingerprint characteristics acquire the biometric samples using touch-based sensors. Some recent researches are focused on the design of touch-less fingerprint recognition systems based on CCD cameras. Most of these systems compute three-dimensional fingertip models and then apply unwrapping techniques in order to obtain images compatible with biometric methods designed for images captured by touch-based sensors. Unwrapped images can present different problems with respect to the traditional fingerprint images. The most important of them is the presence of deformations of the ridge pattern caused by spikes or badly reconstructed regions in the corresponding three-dimensional models. In this paper, we present a neural-based approach for the quality estimation of images obtained from the unwrapping of three-dimensional fingertip models. The paper also presents different sets of features that can be used to evaluate the quality of fingerprint images. Experimental results show that the proposed quality estimation method has an adequate accuracy for the quality classification. The performances of the proposed method are also evaluated in a complete biometric system and compared with the ones obtained by a well-known algorithm in the literature, obtaining satisfactory results. Ruggero Donida Labati, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IJCNN | 4 |
| 2011 | All-IDB: The acute lymphoblastic leukemia image database for image processingabstractThe visual analysis of peripheral blood samples is an important test in the procedures for the diagnosis of leukemia. Automated systems based on artificial vision methods can speed up this operation and increase the accuracy and homogeneity of the response also in telemedicine applications. Unfortunately, there are not available public image datasets to test and compare such algorithms. In this paper, we propose a new public dataset of blood samples, specifically designed for the evaluation and the comparison of algorithms for segmentation and classification. For each image in the dataset, the classification of the cells is given, as well as a specific set of figures of merits to fairly compare the performances of different algorithms. This initiative aims to offer a new test tool to the image processing and pattern matching communities, direct to stimulating new studies in this important field of research. Ruggero Donida Labati, Vincenzo Piuri, Fabio Scotti |
ICIP | 3 |
| 2011 | Biometrics Privacy - Technologies and Applications
Vincenzo Piuri, Fabio Scotti |
SECRYPT | 2 |
| 2010 | Neural-based quality measurement of fingerprint images in contactless biometric systemsabstractTraditional fingerprint biometric systems capture the user fingerprint images by a contact-based sensor. Differently, contactless systems aim to capture the fingerprint images by an approach based on a vision system without the need of any contact of the user with the sensor. The user finger is placed in front of a special CCD-based system that captures the pattern of ridges and valleys of the fingertips. This approach is less constrained by the point of view of the user, but it requires much more capability of the system to deal with the focus of the moving target, the illumination problems and the complexity of the background in the captured image. During the acquisition procedure, the quality of each frame must be carefully evaluated in order to extract only the correct frames with valuable biometric information from the sequence. In this paper, we present a neural-based approach for the quality estimation of the contactless fingertips images. The application of the neural classification models allowed for a relevant reduction of the computational complexity permitting the application in real-time. Experimental results show that the proposed method has an adequate accuracy, and it can capture fingerprints at a distance up to 0.2 meters. Ruggero Donida Labati, Vincenzo Piuri, Fabio Scotti |
IJCNN | 3 |
| 2010 | Noisy iris segmentation with boundary regularization and reflections removal
Ruggero Donida Labati, Fabio Scotti |
Image Vis. Comput. | 2 |
| 2010 | Design of an Automatic Wood Types Classification System by Using Fluorescence SpectraabstractThe classification of wood types is needed in many industrial sectors, since it can provide relevant information concerning the features and characteristics of the final product (appearance, cost, mechanical properties, etc.). This analysis is typical in the furniture industries and the wood panel production. Usually, the analysis is performed by human experts, is not rapid, and has a nonuniform accuracy related mainly to the operator's experience and attention. This paper presents a methodology to effectively cope with the design of an automatic wood types classification system based on the analysis of the fluorescence spectra suitable for real-time applications. This paper presents an experimental set up based on a laser source, a spectrometer, and a processing system, and then, it discusses a set of techniques suitable to extract features from the spectra and how to exploit the extracted feature to train an inductive classification system capable to properly classify the wood types. Obtained experimental results show that the proposed approach can achieve a good accuracy in the classification and requires a limited computational power, hence allowing for the application in real-time industrial processes. Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2009 | Neural-based iterative approach for iris detection in iris recognition systemsabstractThe detection of the iris boundaries is considered in the literature as one of the most critical steps in the identification task of the iris recognition systems. In this paper we present an iterative approach to the detection of the iris center and boundaries by using neural networks. The proposed algorithm starts by an initial random point in the input image, then it processes a set of local image properties in a circular region of interest searching for the peculiar transition patterns of the iris boundaries. A trained neural network processes the parameters associated to the extracted boundaries and it estimates the offsets in the vertical and horizontal axis with respect to the estimated center. The coordinates of the starting point are then updated with the processed offsets. The steps are then iterated for a fixed number of epochs, producing an iterative refinements of the coordinates of the pupils center and its boundaries. Experiments showed that the method is feasible and it can be exploited even in non-ideal operative condition of iris recognition biometric systems. Ruggero Donida Labati, Vincenzo Piuri, Fabio Scotti |
CISDA | 3 |
| 2008 | Privacy-Aware Biometrics: Design and Implementation of a Multimodal Verification SystemabstractA serious concern in the design and use of biometric authentication systems is the privacy protection of the information derived from human biometric traits, especially since such traits cannot be replaced. Combining cryptography and biometrics, several recent works proposed to build the protection in the biometric templates themselves. While these solutions can increase the confidence in biometric systems when biometric information is stored for verification, they have been shown difficult to apply to real biometrics. In this work we present a biometric authentication technique that exploits multiple biometric traits. It is privacy-aware as it ensures privacy protection and allows the extraction of secure identifiers by means of cryptographic primitives. We also discuss the implementation of our approach by considering, as a significant example, the combination of iris and fingerprint biometrics and present experimental results obtained from real data. The implementation shows the feasibility of the scheme in practical applications. Stelvio Cimato, Marco Gamassi, Vincenzo Piuri, Roberto Sassi, Fabio Scotti |
ACSAC | 5 |
| 2006 | Exploiting application locality to design low-complexity, highly performing, and power-aware embedded classifiersabstractTemporal and spatial locality of the inputs, i.e., the property allowing a classifier to receive the same samples over time--or samples belonging to a neighborhood--with high probability, can be translated into the design of embedded classifiers. The outcome is a computational complexity and power aware design particularly suitable for implementation. A classifier based on the gated-parallel family has been found particularly suitable for exploiting locality properties: Subclassifiers are generally small, independent each other, and controlled by a master-enabling module granting that only a subclassifier is active at a time, the others being switched off. By exploiting locality properties we obtain classifiers with accuracy comparable with the ones designed without integrating locality but gaining a significant reduction in computational complexity and power consumption. Cesare Alippi, Fabio Scotti |
IEEE Trans. Neural Networks | 2 |
| 2005 | Fingerprint local analysis for high-performance minutiae extractionabstractThe paper presents a novel approach to identify the minutiae present in a fingerprint image based on the analysis of the local properties. The typical patterns of minutiae called ridge termination and bifurcation are identified by studying the intensity along squared paths in the image. The presented algorithm works both on grey-level image and binarized image and, despite its simplicity, it achieves good accuracy and it can be a good candidate to be implemented in hardware or executed on simple hardware architectures, for example in biometric systems embedded in portable applications such as cellular phones and smart cards. Fabio Scotti, Marco Gamassi, Vincenzo Piuri |
ICIP (3) | 1 |
| 2005 | Visual inspection of particle boards for quality assessmentabstractThe automatic visual inspection (AVI) systems are nowadays used with good results in a wide broad of applications. In particular we considered the problem of the defect detection of particle boards by means of a visual inspection of the printed surface. We propose an innovative defect detection approach capable to automatically extract the repetitive patterns that are generally present in the printed-matters. The extracted patterns are then used in the proper defect detection phase that can thus achieve high defect detection performances. Real wood patterns prove that our defect detection system achieves very high classification capability. Vincenzo Piuri, Fabio Scotti, Manuel Roveri |
ICIP (3) | 2 |
| 2004 | A training-time analysis of robustness in feed-forward neural networksabstractThe paper addresses the analysis of robustness over training time issue. Robustness is evaluated in the large, without assuming the small perturbation hypothesis, by means of randomised algorithms. We discovered that robustness is a strict property of the model -as it is accuracy- and, hence, it depends on the particular neural network family, application, training algorithm and training starting point. Complex neural networks are hence not necessarily more robust than less complex topologies. An early stopping algorithm is finally suggested which extends the one based on the test set inspection with robustness aspects. Cesare Alippi, Daniele Sana, Fabio Scotti |
IJCNN | 3 |