Yangjie Wei

dblp:135/1316 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-6615-5484ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HazFormer: Physics-Guided Hierarchical Transformer for Spatially Non-Uniform Image Dehazing
Cheng Jiayang, Yuxi Li 0002, Dong Ji, Yangjie Wei
ICIC (17)4
2026 ADP-SDM: An adaptive dynamic programming modeling framework for sequential decision-making in automatic speech recognition
Yangjie Wei, Jiayue Sun, Huaguang Zhang, Zhongyang Ming
Neurocomputing2
2026 Accurate target speaker extraction method with adaptive information interaction and updating in complex scenarios
Yangjie Wei, Ben Niu 0011, Yuqiao Wang
J. Supercomput.1
2025 An Abnormal Audio Generation Method for Fault Diagnosis of Power Transformers
abstract
Existing deep learning-based models can achieve a prompt diagnosis of operational anomalies by analyzing the audios emitted from power transformers. However, the practical abnormal data are insufficient for model training, resulting in limited diagnostic performance. To address this problem, we propose an abnormal audio generation method based on the improved cycle generative adversarial networks (ImCycleGAN) to augment the limited training dataset. In the ImCycleGAN, the generator and discriminator are redesigned and adversarially trained to generate the realistic-like audios of six abnormal statuses. Moreover, we combine adversarial, cycle consistency and identity mapping losses to optimize the training process of ImCycleGAN and enhance its ability to capture nonlinear features of audios. Finally, the generated data is evaluated in terms of similarity and fault classification. Experimental results show that our method can generate abnormal audios with high similarity to the real ones, and significantly improve the classification accuracy of existing fault diagnosis models.
Ben Niu 0011, Yangjie Wei, Zhuoran Yu
ICASSP2
2025 Multi-Level Speaker Representation for Target Speaker Extraction
abstract
Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In this work, we propose a multi-level speaker representation approach, from raw features to neural embeddings, to serve as the speaker reference cue. We generate a spectral-level representation from the enrollment magnitude spectrogram as a raw, low-level feature, which significantly improves the model’s generalization capability. Additionally, we propose a contextual embedding feature based on cross-attention mechanisms that integrate frame-level embeddings from a pre-trained speaker encoder. By incorporating speaker features across multiple levels, we significantly enhance the performance of the TSE model. Our approach achieves a 2.74 dB improvement and a 4.94% increase in extraction accuracy on Libri2mix test set over the baseline.
Shuai Wang 0016, Yangjie Wei, Yannan Wang, Haizhou Li 0001
ICASSP4
2025 VISE: Velocity-Guided Interpolant Diffusion for Efficient Speech Enhancement
Yangjie Wei, Ben Niu 0011, Yuqiao Wang, Shengling Yu
ICONIP (1)2
2025 StarGAN-Aug: A Cross-domain Fault Audio Generation Method for High-performance Fault Diagnosis of Power Transformers
Ben Niu 0011, Yangjie Wei, Yuqiao Wang, Shengling Yu
INTERSPEECH2
2025 Acoustic signal augmentation for fault diagnosis of power transformers based on improved cycle generative adversarial networks
Ben Niu 0011, Yangjie Wei, Zhuoran Yu, Yuqiao Wang
Expert Syst. Appl.2
2025 Improved Few-Shot Object Detection Method Based on Faster R-CNN
abstract
ABSTRACT Uneven distribution of object features and insufficient feature learning significantly affect the accuracy and generalizability of existing detection methods. This paper proposes an improved two‐stage few‐shot object detection method that builds upon the faster region‐based convolutional neural network framework to enhance its performance in detecting objects with limited training data. First, a modified data augmentation method for optical images is introduced, and a Gaussian optimization module of sample feature distribution is constructed to enhance the model's generalizability. Second, a parameter‐less 3D space attention module without additional parameters, is added to enhance the space features of a sample, where a neuron linear separability measurement and feature optimization module based on mathematical operations are used to adjust the feature distribution and reduce data distribution bias. Finally, a class feature vector extractor based on meta‐learning is provided to reconstruct the feature map by overlaying a class feature vector from the target domain onto the query image. This process improves accuracy and generalization performance, and multiple experiments on the PASCAL VOC dataset show that the proposed method has higher detection accuracy and stronger generalizability than other methods. Especially, the experiment using practical images under complicated environments indicates its potential effectiveness in real‐world scenarios.
Yangjie Wei, Shangwei Long
IET Image Process.1
2025 Improved AED with multi-stage feature extraction and fusion based on RFAConv and PSA
Yangjie Wei, Zekang Qi
Speech Commun.2
2023 Speaker Extraction with Detection of Presence and Absence of Target Speakers
Marvin Borsdorf, Zexu Pan, Haizhou Li 0001, Yangjie Wei
INTERSPEECH5
2016 Dynamic blind source separation based on source-direction prediction
Yangjie Wei, Wang Yi 0001
Neurocomputing1