VLDB 2026 Research / reviewers in the wild / expert
Yi Li 0047
dblp:59/871-47
· DBLP profile ↗
12ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0003-2304-5802ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 9 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Complex-Cycle-Consistent Diffusion Model for Monaural Speech EnhancementabstractIn this paper, we present a novel diffusion model-based monaural speech enhancement method. Our approach incorporates the separate estimation of speech spectra's magnitude and phase in two diffusion networks. Throughout the diffusion process, noise clips from real-world noise interferences are added gradually to the clean speech spectra and a noise-aware reverse process is proposed to learn how to generate both clean speech spectra and noise spectra. Furthermore, to fully leverage the intrinsic relationship between magnitude and phase, we introduce a complex-cycle-consistent (CCC) mechanism that uses the estimated magnitude to map the phase, and vice versa. We implement this algorithm within a phase-aware speech enhancement diffusion model (SEDM). We conduct extensive experiments on public datasets to demonstrate the effectiveness of our method, highlighting the significant benefits of exploiting the intrinsic relationship between phase and magnitude information to enhance speech. The comparison to conventional diffusion models demonstrates the superiority of SEDM. Yi Li 0047, Plamen Angelov 0001 |
AAAI | 1 |
| 2025 | Detecting Cross-domain Deepfake Videos with Contrastive Prototype LearningabstractDeepfake videos are synthetic media generated using advanced deep learning techniques that manipulate or replace the visual and audio content of an original recording, enabling the creation of highly realistic yet entirely fabricated audiovisual content. The proliferation of such manipulated media poses significant societal risks, including potential misinformation, reputation damage, psychological manipulation, and erosion of trust in digital visual communication. Recent deep learning methods for deepfake detection have emerged, leveraging sophisticated machine learning models that analyze multi-modal cues, including facial inconsistencies, unnatural temporal dynamics, and visual misalignments to distinguish between authentic and synthetic content. However, these state-of-the-art detection approaches often struggle with the domain-shift challenge, where models trained on specific deepfake datasets fail to generalize effectively when confronted with unseen generation techniques or evolving synthesis technologies. To address this critical limitation, we propose a self-supervised contrastive learning framework called CPDD, introducing contrast between features and prototypes of original data to alleviate domain-specific distractions (i.e., deepfake generative models or datasets). We calculate the cosine similarity between two features or prototypes to scale the original distance, clustering the features around closely related prototypes. This process encodes the semantic structures discovered through clustering into the learned embedding space. The extensive experiments show that, compared to various benchmark deepfake detection models and domain generalization techniques, the proposed model achieves state-of-the-art performance on the cross-domain deepfake detection task across a wide range of scenarios. Yi Li 0047, Plamen Angelov 0001 |
IJCNN | 1 |
| 2025 | Vision-Based Landing Guidance Through Tracking and Orientation EstimationabstractFixed-wing aerial vehicles are equipped with functionalities such as ILS (instrument landing system), PAR (precision approach radar) and, DGPS (differential global positioning system), enabling fully automated landings. However, these systems impose significant costs on airport operations due to high installation and maintenance requirements. Moreover, since these navigation parameters come from ground or satellite signals, they are vulnerable to interference. A more cost-effective and independent alternative for guiding landing is a vision-based system that detects the runway and aligns the aircraft, reducing the pilot’s cognitive load. This paper proposes a novel framework that addresses three key challenges in developing autonomous vision-based landing systems. Firstly, to overcome the lack of aerial front-view video data, we created high-quality videos simulating landing approaches through the generator code available in the LARD (landing approach runway detection dataset) repository. Secondly, in contrast to former studies focusing on object detection for finding the runway, we chose the state-of-the-art model LoRAT to track runways within bounding boxes in each video frame. Thirdly, to align the aircraft with the designated landing runway, we extract runway keypoints from the resulting LoRAT frames and estimate the camera relative pose via the Perspective-n-Point algorithm. Our experimental results over a dataset of generated videos and original images from the LARD dataset consistently demonstrate the proposed framework’s highly accurate tracking and alignment capabilities. Our approach source code and the LoRAT model pre-trained with LARD videos are available at https:// github.com/ jpklock2/ visionbased-landing-guidance João P. K. Ferreira, João P. L. Pinto, Júlia S. Moura, Yi Li 0047, Cristiano Leite Castro, Plamen Angelov 0001 |
WACV | 4 |
| 2024 | Self-supervised Representation Learning for Adversarial Attack Detection
Yi Li 0047, Plamen Angelov 0001, Neeraj Suri |
ECCV (60) | 1 |
| 2024 | CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Ground Image Synthesis
Yuankun Chen, Dazhong Rong, Yi Li 0047 |
ICANN (3) | 3 |
| 2024 | Federated Adversarial Learning for Robust Autonomous Landing Runway Detection
Yi Li 0047, Plamen Angelov 0001, Zhengxin Yu, Alvaro Lopez Pellicer, Neeraj Suri |
ICANN (6) | 1 |
| 2024 | Rethinking Self-supervised Learning for Cross-domain Adversarial Sample RecoveryabstractAdversarial attacks can cause misclassification in machine learning pipelines, posing a significant safety risk in critical applications such as autonomous systems or medical applications. Supervised learning-based methods for adversarial sample recovery rely heavily on large volumes of labeled data, which often results in substantial performance degradation when applying the trained model to new domains. In this paper, differing from conventional self-supervised learning techniques such as data augmentation, we present a novel two-stage self-supervised representation learning framework for the task of adversarial sample recovery, aimed at overcoming these limitations. In the first stage, we employ a clean image autoencoder (CAE) to learn representations of clean images. Subsequently, the second stage utilizes an adversarial image autoencoder (AAE) to learn a shared latent space that captures the relationships between the representations acquired by CAE and AAE. It is noteworthy that the input clean images in the first stage and adversarial images in the second stage are cross-domain and not paired. To the best of our knowledge, this marks the first instance of self-supervised adversarial sample recovery work that operates without the need for labeled data. Our experimental evaluations, spanning a diverse range of images, consistently demonstrate the superior performance of the proposed method compared to conventional adversarial sample recovery methods. Yi Li 0047, Plamen Angelov 0001, Neeraj Suri |
IJCNN | 1 |
| 2024 | UNICAD: A Unified Approach for Attack Detection, Noise Reduction and Novel Class IdentificationabstractAs the use of Deep Neural Networks (DNNs) becomes pervasive, their vulnerability to adversarial attacks and limitations in handling unseen classes poses significant challenges. The state-of-the-art offers discrete solutions aimed to tackle individual issues covering specific adversarial attack scenarios, classification or evolving learning. However, real-world systems need to be able to detect and recover from a wide range of adversarial attacks without sacrificing classification accuracy and to flexibly act in unseen scenarios. In this paper, UNICAD, is proposed as a novel framework that integrates a variety of techniques to provide an adaptive solution.For the targeted image classification, UNICAD achieves accurate image classification, detects unseen classes, and recovers from adversarial attacks using Prototype and Similarity-based DNNs with denoising autoencoders. Our experiments performed on the CIFAR-10 dataset highlight UNICAD’s effectiveness in adversarial mitigation and unseen class classification, outperforming traditional models. Alvaro Lopez Pellicer, Kittipos Giatgong, Yi Li 0047, Neeraj Suri, Plamen Angelov 0001 |
IJCNN | 3 |
| 2024 | Adversarial Attack Detection via Fuzzy PredictionsabstractImage processing using neural networks act as a tool to speed up predictions for users, specifically on large-scale image samples. To guarantee the clean data for training accuracy, various deep learning-based adversarial attack detection techniques have been proposed. These crisp set-based detection methods directly determine whether an image is clean or attacked, while, calculating the loss is nondifferentiable and hinders training through normal back-propagation. Motivated by the recent success in fuzzy systems, in this work, we present an attack detection method to further improve detection performance, which is suitable for any pretrained neural network classifier. Subsequently, the fuzzification network is used to obtain feature maps to produce fuzzy sets of difference degree between clean and attacked images. The fuzzy rules control the intelligence that determines the detection boundaries. Different from previous fuzzy systems, we propose a fuzzy mean-intelligence mechanism with new support and confidence functions to improve fuzzy rule's quality. In the defuzzification layer, the fuzzy prediction from the intelligence is mapped back into the crisp model predictions for images. The loss between the prediction and label controls the rules to train the fuzzy detector. We show that the fuzzy rule-based network learns rich feature information than binary outputs and offer to obtain an overall performance gain. Experiment results show that compared to various benchmark fuzzy systems and adversarial attack detection methods, our fuzzy detector achieves better detection performance over a wide range of images. Yi Li 0047, Plamen Angelov 0001, Neeraj Suri |
IEEE Trans. Fuzzy Syst. | 1 |
| 2023 | Domain Generalization and Feature Fusion for Cross-domain Imperceptible Adversarial Attack DetectionabstractDeep learning-based imperceptible adversarial attack detection methods have recently seen significant progress. However, the accuracy, latency, and computational cost of previous methods remain insufficient. Particularly, trained attack detection models can potentially be applied in previously unseen conditions, such as new datasets or attacks for real-world applications. Therefore, to improve domain generalization performance, we propose a new method for cross-domain imperceptible adversarial attack detection by leveraging domain generalization, where we train the model's feature extractor or detector with a partner well-tuned for different domains. Different from conventional domain generalization methods, we use the global loss and local loss to train each feature extractor or detector. Moreover, to efficiently re-use high-resolution feature maps from the feature extractor, we propose a feature fusion network, which exploits feature maps from images that are attacked with different error rates and helps extract rich features to further improve the attack detection accuracy. Extensive experiments on four public datasets are used to demonstrate the efficacy of the proposed method. The source code of the proposed method is available at https://github.com/Yukino-3/DTAD. Yi Li 0047, Plamen Angelov 0001, Neeraj Suri |
IJCNN | 1 |
| 2023 | U-Shaped Transformer With Frequency-Band Aware Attention for Speech Enhancement
Yi Li 0047, Yang Sun 0003, Wenwu Wang 0001, Syed M. Naqvi |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Single-channel dereverberation and denoising based on lower band trained SA-LSTMsabstractThe supervised single‐channel speech enhancement presents one mixture recording at the input of the neural network and updates network parameters in order to generate an output as the reconstructed speech signal. However, current neural networks‐based single‐channel speech enhancement methods are not able to fully utilise pertinence with the specific frequency range of speech signals with limited computational complexity. In this study, the authors studied the power spectral density of mixtures with human speech and noise interferences. Based on the theory that the speech signal distributes at the lower band, they proposed a method to train signal approximation (SA) based neural networks with the lower frequency band of the speech mixture to improve the performance. To realise the lower band approach for single‐channel speech enhancement, the method uses a long short‐term memory (LSTM) block to exploit short‐time Fourier transform of the desired frequency range. Furthermore, in order to improve the speech enhancement performance within reverberant room environments, the dereverberation mask and the enhanced ratio mask are exploited as the training targets of two LSTM blocks, respectively. The detailed evaluations confirm that the proposed method outperforms the state‐of‐the‐art methods. Yi Li 0047, Yang Sun 0003, Syed M. Naqvi |
IET Signal Process. | 1 |