Yuhong Yang 0001

dblp:52/5811-1 · DBLP profile ↗
← Back
61ranked-venue papers
4as first author
43since 2021 · last 2026
0000-0003-3001-7957ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 4 first-author · 32 since 2021Artificial intelligence and machine learning · 26 · 1 first-author · 25 since 2021Systems, architecture and hardware · 2Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 GLoMOT: Efficient Online GNN-based Low-Frame-Rate Multi-Object Tracker
abstract
Low-frame-rate (LFR) Multi-Object Tracking (MOT) is crucial for efficient tracking on edge devices, as it significantly reduces computational and storage demands. However, existing trackers struggle in LFR settings due to large temporal gaps, extreme appearance changes, and motion non-linearity. While Graph Neural Network (GNN)-based trackers are effective at associating objects across these gaps, most operate offline, which prevents their use for online tracking. To address these limitations, we propose GLoMOT, a novel online GNN-based Low-Frame-Rate Multi-Object Tracker designed for robust performance in LFR videos. To bridge the large temporal gaps, we introduce a Dynamic Node Buffer Pool. This acts as a long-term memory, caching the states of absent objects to enable their robust re-association. To tackle extreme motion uncertainty, we propose an adaptive context-aware module that dynamically adjusts the weights of positional and appearance features, generating more robust features for predicting node connections. Furthermore, we propose a pseudo-depth feature calculation method. This provides the GNN with critical geometric context, which helps resolve spatial ambiguity arising from occlusions. Extensive experiments on several public MOT benchmarks, including DanceTrack, MOT17, and VisDrone, demonstrate GLoMOT's effectiveness and superiority, particularly in challenging Low-Frame-Rate conditions.
Yaxuan Hu 0001, Jie Hua 0005, Gang Wu 0010, Yuhong Yang 0001, Atsushi Suzuki 0002, Zhongyuan Wang 0001
AAAI4
2026 DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery Localization
abstract
Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments in video and audio, offering strong interpretability for security and forensics. While recent State Space Models (SSMs) show promise in precise temporal reasoning, their use in TFL is hindered by ambiguous boundaries, sparse forgeries, and limited long-range modeling. We propose DeformTrace, which enhances SSMs with deformable dynamics and relay mechanisms to address these challenges. Specifically, Deformable Self-SSM (DS-SSM) introduces dynamic receptive fields into SSMs for precise temporal localization. To further enhance its capacity for temporal reasoning and mitigate long-range decay, a Relay Token Mechanism is integrated into DS-SSM. Besides, Deformable Cross-SSM (DC-SSM) partitions the global state space into query-specific subspaces, reducing non-forgery information accumulation and boosting sensitivity to sparse forgeries. These components are integrated into a hybrid architecture that combines the global modeling of Transformers with the efficiency of SSMs. Extensive experiments show that DeformTrace achieves state-of-the-art performance with fewer parameters, faster inference, and stronger robustness.
Suting Wang, Yuanming Zheng, Junqi Yang, Yangxu Liao, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001
AAAI6
2026 AudioJailbreak: Jailbreak Attacks Against End-to-End Large Audio-Language Models
abstract
Jailbreak attacks to Large audio-language models (LALMs) are studied recently, but they exclusively focused on the attack scenario where the adversary can fully manipulate user prompts (named strong adversary) and limited in effectiveness, applicability, and practicability. In this work, we first conduct an extensive evaluation showing that advanced text jailbreak attacks cannot be easily ported to end-to-end LALMs via text-to-speech (TTS) techniques. We then propose AUDIOJAILBREAK, a novel audio jailbreak attack, featuring (1) asynchrony: the jailbreak audios do not need to align with user prompts in the time axis by crafting suffixal jailbreak audios; (2) universality: a single jailbreak perturbation is effective for different prompts by incorporating multiple prompts into the perturbation generation; (3) stealthiness: the malicious intent of jailbreak audios is concealed by proposing various intent concealment strategies; and (4) over-the-air robustness: the jailbreak audios remain effective when being played over the air by incorporating reverberation into the perturbation generation. In contrast, all prior audio jailbreak attacks cannot offer asynchrony, universality, stealthiness, and/or over-the-air robustness. Moreover, AUDIOJAILBREAK is also applicable to a more practical and broader attack scenario where the adversary cannot fully manipulate user prompts (named weak adversary). Extensive experiments with thus far the most LALMs demonstrate the high effectiveness of AUDIOJAILBREAK, in particular, it can jailbreak openAI's GPT-4o-Audio and bypass Meta's Llama-Guard-3 safeguard, in the weak adversary scenario. We highlight that our work peeks into the security implications of audio jailbreak attacks against LALMs, and realistically fosters improving their robustness, especially for the newly proposed weak adversary.
Guangke Chen, Fu Song, Zhe Zhao 0007, Xiaojun Jia, Yang Liu 0003, Yanchen Qiao, Weizhe Zhang, Weiping Tu, Yuhong Yang 0001, Bo Du 0001
IEEE Trans. Dependable Secur. Comput.9
2025 FreqSense: Universal and Low-Latency Adversarial Example Detection for Speaker Recognition with Interpretability in Frequency Domain
abstract
Speaker recognition (SR) systems are particularly vulnerable to adversarial example (AE) attacks. To mitigate these attacks, AE detection systems are typically integrated into SR systems. To overcome the limitations of low detection accuracy, poor generalization, and high latency in existing schemes, this paper proposes FreqSense, an AE detection scheme based on frequency distribution features. FreqSense detects a variety of unknown AE attacks with low latency, and provides interpretability in its detection process. The basic idea of FreqSense is that AE typically introduce carefully designed noise in specific frequency bands that are associated with highly distinctive speaker identities. Therefore, leveraging the distributional variations in these frequency bands can effectively distinguish between AE and benign audio. FreqSense models frequency distribution features by integrating time-frequency transformation technology with a self-attention mechanism and employs a neural network-based classifier to distinguish between AE and benign audio. Experimental results show that FreqSense achieves an overall detection accuracy of 99.2%, surpassing state-of-the-art (SOTA) schemes by 27.2%. When confronting unknown AE attacks, FreqSense achieves a detection accuracy of 98.3% with a latency of just 0.0014 seconds.
Yihuan Huang, Yanzhen Ren, Weiping Tu, Yuhong Yang 0001
ICASSP5
2025 Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
abstract
Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same SSM to all sub-bands limits its inference capability. Additionally, when processing each time frame of the time-frequency representation, the SSM may forget certain high-frequency information of low energy, making the restoration of structure in the high-frequency bands challenging. For this reason, we propose Cross- and Sub-band Mamba (CSMamba). To assist the SSM in handling different sub-band features flexibly, we propose a band split block that splits the full-band into four sub-bands with different widths based on their information similarity. We then allocate independent weights to each sub-band, thereby reducing the inference burden on the SSM. Furthermore, to mitigate the forgetting of low-energy information in the high-frequency bands by the SSM, we introduce a spectrum restoration block that enhances the representation of the cross-band features from multiple perspectives. Experimental results on the DNS Challenge 2021 dataset demonstrate that CSMamba outperforms several state-of-the-art (SOTA) speech enhancement methods in three objective evaluation metrics with fewer parameters.
Jizhen Li, Weiping Tu, Yuhong Yang 0001, Xinmeng Xu, Yanzhen Ren
ICASSP3
2025 Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio Codecs
abstract
Existing end-to-end neural codecs have made great progress in preserving audio quality. Despite their success, they still face challenges in achieving accurate and efficient quantization. Specifically, these codecs often overlook which features have a greater impact on perceptual audio quality during quantization, leading to a quantization error distribution that fails to reflect the actual importance of latent features. They are also sensitive to unusual data points (outliers) because they use Mean Squared Error (MSE) to measure quantization errors, which can increase quantization noise or spectral artifacts. To address these limitations, we propose AW-CEQCodec which integrates an Attention Weighting (AW) module and a Conditional Entropy-driven Quantization (CEQ) loss. The AW enhances key regions of latent features before quantization, enabling more accurate quantizing critical features and reducing their quantization errors. After quantization, it restores global details from dequantized features, improving overall reconstruction. Moreover, the CEQ minimizes the uncertainty between latent and quantized features, effectively reflecting the distortion introduced by the quantization module. Experimental results on the CodecSuperb-STL dataset demonstrate that our method consistently outperforms baseline approaches, achieving superior audio quality at bitrates as low as 0.5 kbps, confirming its effectiveness in minimizing distortion and preserving perceptual quality. The reconstruction audio samples can be find at https://huazhi1024.github.io/first-page.
Weiping Tu, Yuhong Yang 0001, Xinmeng Xu, Yanzhen Ren
ICASSP3
2025 HAPG-SAQAM: Human Auditory Perception Guided Spatial Audio Quality Assessment Metric
abstract
Spatial audio quality evaluation is essential for applications like virtual and augmented reality, where accurate sound reproduction enhances user immersion. While subjective listening tests are the gold standard, they are costly and time-consuming. To address this, we propose HAPG-SAQAM, an objective metric for assessing timbre quality, spatial quality, and overall quality of binaural audio, guided by human auditory perception. Our contributions include: (1) the Multi-scale Auditory Guided Feature Extraction (MAGFE) module, incorporating gammatone frequency cepstral coefficients for better alignment with human perception; (2) Perceptual Weighted Loss (PWL), optimizing the weighting of timbre quality (TQ) and spatial quality (SQ) loss based on subjective test data; and (3) data augmentation techniques to enhance robustness by amplifying perceptual distortions. Experimental results show HAPG-SAQAM improves correlation with subjective scores by 10%, with ablation studies confirming the contributions of its components to enhanced spatial and overall audio quality.
Yuanming Zheng, Jiaxuan Yao, Xiangyu Deng, Yuhong Yang 0001, Ruiqi Liao, Weiping Tu, Cedar Lin
ICASSP4
2025 FIRING-Net: A filtered feature recycling network for speech enhancement
abstract
Current deep neural networks for speech enhancement (SE) aim to minimize the distance between the output signal and the clean target by filtering out noise features from input features. However, when noise and speech components are highly similar, SE models struggle to learn effective discrimination patterns. To address this challenge, we propose a Filter-Recycle-Interguide framework termed Filter-Recycle-INterGuide NETwork (FIRING-Net) for SE, which filters the input features to extract target features and recycles the filtered-out features as non-target features. These two feature sets then guide each other to refine the features, leading to the aggregation of speech information within the target features and noise information within the non-target features. The proposed FIRING-Net mainly consists of a Local Module (LM) and a Global Module (GM). The LM uses outputs of the speech extraction network as target features and the residual between input and output as non-target features. The GM leverages the energy distribution of self-attention map to extract target and non-target features guided by highest and lowest energy regions. Both LM and GM include interaction modules to leverage the two feature sets in an inter-guided manner for collecting speech from non-target features and filtering out noise from target features. Experiments confirm the effectiveness of the Filter-Recycle-Interguide framework, with FIRING-Net achieving a strong balance between SE performance and computational efficiency, surpassing comparable models across various SNR levels and noise environments.
Xinmeng Xu, Jizhen Li, Yuhong Yang 0001, Yong Luo 0002, Weiping Tu
ICLR4
2025 PGD-N2L: A Parameter-Guided Disentanglement Approach for Normal-To-Lombard Speech Conversion
abstract
The Normal-To-Lombard (N2L) speech conversion can effectively improve speech intelligibility in noisy communication scenarios and serve as a data augmentation tool for various speech-related algorithms. However, existing N2L methods did not aim to disentangle the Lombard effect from other speech attributes, leading to incomplete conversions. In this paper, we propose a Parameter-Guided Disentanglement approach for N2L speech conversion (PGD-N2L) which decomposes speech into linguistic content, speaker identity, and Lombard effect. To extract disentangled linguistic content, we propose a DeLomb-Based content encoder. To extract disentangled speaker identity and Lombard effect, we propose a style encoder that combines a fine-tuned speaker encoder and a learnable Lombard encoder to form a personalized style embedding. Furthermore, an En-Lomb-Based injection module is designed to accurately integrate the target Lombard effect and speaker identity into the linguistic content based on personalized style embedding, ensuring complete Lombard conversion. Experimental results demonstrate that our proposed method outperforms existing N2L models in speech intelligibility, acoustic similarity, and speech quality. Ablation studies confirm that the fine-tuned speaker encoder and the De-Lomb block effectively improve speech intelligibility and acoustic similarity, while the En-Lomb block enables the converted speech to more closely match the target Lombard speech.
Hongyang Chen 0004, Yuhong Yang 0001, Xinmeng Xu, Weiping Tu, Zhongyuan Wang 0001, Cedar Lin
ICME2
2025 Band-SCNet: A Causal, Lightweight Model for High-Performance Real-Time Music Source Separation
Junqi Yang, Yuhong Yang 0001, Weiping Tu, Cedar Lin
INTERSPEECH2
2025 FreeCodec: A Disentangled Neural Speech Codec with Fewer Tokens
Youqiang Zheng, Weiping Tu, Yueteng Kang, Li Xiao 0007, Yuhong Yang 0001
INTERSPEECH7
2025 Query-Based Audio-Visual Temporal Forgery Localization with Register-Enhanced Representation Learning
abstract
Temporal forgery in multimedia-where audio or video streams are subtly manipulated-poses critical challenges for content authenticity verification. While video-level detection has advanced, Temporal Forgery Localization (TFL) remains underexplored, often limited by weak audio-visual modeling and reliance on non-learnable post-processing. To address these challenges, we propose RegQAV, a Register-enhanced Query-based Audio-Visual framework for TFL. RegQAV exploits pretrained foundation models to capture fine-grained audio-visual correspondences and learnable registers are introduced to mitigate the model's tendency to overly focus on a limited set of temporal features. A query-based localization strategy enables end-to-end optimization without post-processing. We also introduce a Modality Fusion Adapter (MFA) for effective multi-scale integration of audio-visual data, a Deepfake Queries Generation (DQG) module for efficient query initialization, and a Poisson Count-Based Approach to dynamically predict the number of forgeries. Experiments on LAV-DF and AV-Deepfake1M show that RegQAV achieves state-of-the-art performance with fewer parameters, faster inference, and stronger generalization. This work offers significant potential for real-time deepfake detection and other multimedia verification applications. The code is available at https://github.com/zxd3099/RegQAV.
Suting Wang, Junqi Yang, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001
ACM Multimedia4
2025 Lombard-VLD: Voice Liveness Detection Based on Human Auditory Feedback
abstract
Voice Liveness Detection (VLD) aims to protect speaker authentication from speech spoofing by determining whether speeches come from live speakers or loudspeakers. Previous methods mainly focus on their differences at the signal level. In this paper, we propose the first VLD that uses the human auditory feedback mechanism (i.e., the Lombard effect), called Lombard-VLD. The key idea is that live speakers can physiologically and involuntarily adjust their speaking patterns in a noisy background but loudspeakers cannot. Moreover, we design a reference-based dual input mode and a differential SE-ResBlock to model the acoustic differences caused by the Lombard effect. Experimental results show that Lombard-VLD achieves 0% and 0.24% EER in two datasets, outperforming the state-of-the-art methods. It is robust to various environmental factors, including different distances, postures of the speaker, and environmental noise, with an average accuracy of over 98.51%. It also has a good generalization to unseen speakers, genders, and datasets, with EER lower than 2.68%, 3.44%, and 7.32%, respectively. This work shows the advantages of the Lombard effect in VLD, which has fewer user limitations and better detection performance.
Hongcheng Zhu, Zongkun Sun, Yanzhen Ren, Kun He 0008, Yongpeng Yan, Wuyang Liu, Yuhong Yang 0001, Weiping Tu
SP8
2025 Spatial information aided speech and noise feature discrimination for Monaural speech enhancement
Xinmeng Xu, Jizhen Li, Weiping Tu, Yuhong Yang 0001
Expert Syst. Appl.5
2025 Lightweight real-time speech enhancement: State-space models and multi-spectral scanning techniques
Junqi Yang, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001
Neural Networks3
2025 Luminance decomposition and reconstruction for high dynamic range Video Quality Assessment
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Jing Xiao 0004, Zixiang Xiong
Pattern Recognit.5
2024 SnoreOxiNet: Non-contact Diagnosis of Nocturnal Hypoxemia Using Cross-Domain Acoustic Features
Weiyan Yi, Xiuping Yang, Li Xiao 0007, Weiping Tu, Yuhong Yang 0001
ICANN (8)6
2024 EMALG: An Enhanced Mandarin Lombard Grid Corpus with Meaningful Sentences
abstract
This study investigates the Lombard effect, where individuals adapt their speech in noisy environments. We introduce an enhanced Mandarin Lombard grid (EMALG) corpus with meaningful sentences, enhancing the Mandarin Lombard grid (MALG) corpus. EMALG features 34 speakers and improves recording setups, addressing challenges faced by MALG with nonsense sentences. Our findings reveal that in Mandarin, meaningful sentences are more effective in enhancing the Lombard effect. Additionally, we uncover that female exhibit a more pronounced Lombard effect than male when uttering meaningful sentences. Moreover, our results reaffirm the consistency in the Lombard effect comparison between English and Mandarin found in previous research.
Baifeng Li, Qingmu Liu, Yuhong Yang 0001, Hongyang Chen 0004, Weiping Tu
ICASSP3
2024 Curricular Contrastive Regularization for Speech Enhancement with Self-Supervised Representations
abstract
Existing deep learning-based speech enhancement methods only adopt clean speech as positive samples to guide the training of speech enhancement networks while negative samples, i.e., noisy speech, are unexploited. In this paper, we adopt contrastive regularization (CR) built upon contrastive learning to exploit both the information of noisy and clean speech as negative and positive samples, respectively. Particularly, CR minimizes the distance between clean and enhanced speech and maximizes the distance between noisy and enhanced speech in the representation space of the self-supervised learning model. However, the contrastive samples are non-consensual, as the negatives are usually represented distantly from the clean speech, leaving the solution space still under-constricted. To tackle this issue, we provide the negative samples assembled from (1) the noisy speech, and (2) the corresponding enhanced speech without using CR, and we customize a curriculum learning strategy to define the importance of these negative samples to balance the learning difficulty caused by different similarities between the embeddings of the positive and negative samples. Experiments show that our proposal improves SE performance effectively without introducing additional computation/parameters.
Xinmeng Xu, Chang Han, Weiping Tu, Yuhong Yang 0001
ICASSP5
2024 An Efficient and Interpre Table Speech Enhancement Network Via Deep Dictionary Learning
abstract
Speech enhancement is a vital and highly ill-posed problem for many speech downstream tasks. While currently existing deep learning based speech enhancement methods have held state-of-the-art results, they still possess apparent shortcomings in that most of the deep learning based models lack interpretability. This deficiency results in unsatisfied speech enhancement performance in many sophisticated scenarios. To tackle this problem, we integrate dictionary learning and sparse coding into deep learning networks for speech enhancement and present a deep dictionary learning based speech enhancement network (DicLSENet). Specifically, the proposed DicLSENet strictly follows the principle of dictionary learning, learns the priors for both representation coefficients and dictionaries, and adaptively adjusts the dictionary for each input. Experimental results show that the proposed model outperforms state-of-the-art fully deep learning based methods with attractive computational costs.
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
ICASSP4
2024 LungAdapter: Efficient Adapting Audio Spectrogram Transformer for Lung Sound Classification
Li Xiao 0007, Lucheng Fang, Yuhong Yang 0001, Weiping Tu
INTERSPEECH3
2024 Exploring Sentence Type Effects on the Lombard Effect and Intelligibility Enhancement: A Comparative Study of Natural and Grid Sentences
Hongyang Chen 0004, Yuhong Yang 0001, Zhongyuan Wang 0001, Weiping Tu, Haojun Ai, Cedar Lin
INTERSPEECH2
2024 Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
Jizhen Li, Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
INTERSPEECH4
2024 SimuSOE: A Simulated Snoring Dataset for Obstructive Sleep Apnea-Hypopnea Syndrome Evaluation during Wakefulness
Xiuping Yang, Li Xiao 0007, Weiyan Yi, Yuhong Yang 0001, Weiping Tu
INTERSPEECH6
2024 Adaptive selection of local and non-local attention mechanisms for speech enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
Neural Networks3
2024 Joint Distortion Restoration and Quality Feature Learning for No-reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) methods, inspired by the free energy principle, improve the accuracy of image quality prediction by simulating the human brain’s repair process for distorted images. However, existing methods use separate optimization schemes for distortion restoration and quality prediction, which undermines the accurate mapping of feature representations to quality scores. To address this issue, we propose a joint restoration and quality feature learning NR-IQA (RQFL-IQA) method to jointly tackle distortion image restoration and quality prediction within a unified framework. To accurately establish the quality reconstruction relationship between distorted and restored images, a hybrid loss function based on pixel-wise and structure-wise representations is used to improve the restoration capability of the image restoration network. The proposed RQFL-IQA exploits rich labels, including restored images and quality scores, to enable the model to learn more discriminative features and establish a more accurate mapping from feature representation to quality scores. In addition, to avoid the impact of poor restoration on quality prediction, we propose a module with a cleaning function to reweight the fusion of restored and primitive features to achieve more perceptual consistency in feature fusion. Experimental results on public IQA datasets show that the proposed RQFL-IQA is superior over existing methods.
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Zixiang Xiong
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Auxiliary Information Guided Self-attention for Image Quality Assessment
abstract
Image quality assessment (IQA) is an important problem in computer vision with many applications. We propose a transformer-based multi-task learning framework for the IQA task. Two subtasks: constructing an auxiliary information error map and completing image quality prediction, are jointly optimized using a shared feature extractor. We use visual transformers (ViT) as a feature extractor for feature extraction and guide ViT to focus on image quality-related features by building auxiliary information error map subtask. In particular, we propose a fusion network that includes a channel focus module. Unlike the fusion methods commonly used in previous IQA methods, we use the fusion network, including the channel attention module, to fuse the auxiliary information error map features with the image features, which facilitates the model to mine the image quality features for more accurate image quality assessment. And by jointly optimizing the two subtasks, ViT focuses more on extracting image quality features and building a more precise mapping from feature representation to quality score. With slight adjustments to the model, our approach can be used in both no-reference (NR) and full-reference (FR) IQA environments. We evaluate the proposed method in multiple IQA databases, showing better performance than state-of-the-art FR and NR IQA methods.
Jifan Yang, Zhongyuan Wang 0001, Guangcheng Wang, Baojin Huang, Yuhong Yang 0001, Weiping Tu
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech Enhancement
abstract
Attention mechanisms, such as local and non-local attention, play a fundamental role in recent deep learning based speech enhancement (SE) systems. However, a natural speech contains many fast-changing and relatively briefly acoustic events, therefore, capturing the most informative speech features by indiscriminately using local and non-local attention is challenged. We observe that the noise type and speech feature vary within a sequence of speech and the local and non-local can respectively process different types of corrupted speech regions. To leverage this, we propose Selector-Enhancer, a dual-attention based convolution neural network (CNN) with a feature-filter that can dynamically select regions from low-resolution speech features and feed them to local or non-local attention operations. In particular, the proposed feature-filter is trained by using reinforcement learning (RL) with a developed difficulty-regulated reward that related to network performance, model complexity and “the difficulty of the SE task”. The results show that our method achieves comparable or superior performance to existing approaches. In particular, Selector-Enhancer is effective for real-world denoising, where the number and types of noise are varies on a single noisy mixture.
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
AAAI3
2023 Improving Acoustic Echo Cancellation by Mixing Speech Local and Global Features with Transformer
abstract
We propose MiT-Net, a novel mix-transformer neural network with a pyramid encoder operating in the time domain, for the task of acoustic echo cancellation. The MiT-Net formulates acoustic echo cancellation as a supervised speech separation problem, in which near-end speech is separated from a single microphone recording and sent to the far end, and consists of two key components. First, we apply a pyramid encoder, which adopts the coarse-to-fine structure, to extract the latent correlations between double-end signals and to fuse them in a multiscale manner. Second, we propose a mix-transformer, a combination of local and global attention in a parallel way, to leverage local and global speech information for separation. Experimental results show that the proposed method outperforms recent AEC methods in terms of objective evaluation metrics. In addition, exploring the correlation between speech local and global features by using the mix-transformer significantly improves the system performance and shows more robustness than the conventional transformer.
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001, Li Xiao 0007
ICASSP4
2023 Learning From Single-Expert Annotated Labels for Automatic Sleep Staging
abstract
Existing automatic sleep staging algorithms rely on accurately labeled data. However, due to the subjectivity of sleep experts, accurate labels must be obtained through joint labeling by multiple experts, which results in high time and labor costs. In this work, we treat labels mislabeled by a single expert as noisy labels and first propose SE-ASS, an automatic sleep staging learning framework based on single-expert annotated data. Since multiple models tend to produce inconsistent predictions for instances with incorrect labels during training, we use two networks with the same structure but different initializations and regularize them with a prediction consistency loss to prevent overfitting to noisy labels. Furthermore, we use a contrastive loss between models to enhance the exploration of feature representations without relying on potentially noisy labels. Our results on two publicly available datasets show that SE-ASS can effectively improve the performance of automatic sleep staging models trained on single-expert annotated datasets.
Zhiheng Luan, Yanzhen Ren, Xiuping Yang, Weiping Tu, Yuhong Yang 0001
ICASSP7
2023 PMMSD: Development of the Matrix Sentence Intelligibility Dataset for Mandarin with Lombard Effect
abstract
This paper presents a Paired Mandarin Matrix Sentence Dataset (PMMSD), which will be available after publication. PMMSD is the first Mandarin matrix sentence intelligibility dataset containing both plain and Lombard speech for scientific research. The results verify that different Lombard styles would affect word intelligibility to different degrees and the Lombard effect helps maintain homogeneous intelligibility against contextual interference. All of the discoveries indicate that the Lombard effect should be considered when building intelligibility datasets with noise in the future.
Hanchen Pei, Yuhong Yang 0001, Xufeng Chen, Qingmu Liu, Hongyang Chen 0004, Weiping Tu
ICASSP2
2023 ONEI: Unveiling Route and Phase of Breathing from Snoring Sounds
Baoai Han, Li Xiao 0007, Xiuping Yang, Weiping Tu, Weiyan Yi, Yuhong Yang 0001, Yanzhen Ren
ICONIP (9)9
2023 Leveraging Sound Local and Global Features for Language-Queried Target Sound Extraction
Xinmeng Xu, Yuhong Yang 0001, Weiping Tu
ICONIP (4)3
2023 Exploring the Interactions Between Target Positive and Negative Information for Acoustic Echo Cancellation
Chang Han, Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
INTERSPEECH4
2023 A Snoring Sound Dataset for Body Position Recognition: Collection, Annotation, and Analysis
Li Xiao 0007, Xiuping Yang, Weiping Tu, Weiyan Yi, Yuhong Yang 0001, Yanzhen Ren
INTERSPEECH8
2023 PCNN: A Lightweight Parallel Conformer Neural Network for Efficient Monaural Speech Enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
INTERSPEECH3
2023 CQNV: A Combination of Coarsely Quantized Bitstream and Neural Vocoder for Low Rate Speech Coding
Youqiang Zheng, Li Xiao 0007, Weiping Tu, Yuhong Yang 0001, Xinmeng Xu
INTERSPEECH4
2023 CASE-Net: Integrating local and non-local attention operations for speech enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
Speech Commun.3
2022 CS-CTCSCONV1D: Small footprint speaker verification with channel split time-channel-time separable 1-dimensional convolution
Linjun Cai, Yuhong Yang 0001, Xufeng Chen, Weiping Tu, Hongyang Chen 0004
INTERSPEECH2
2022 Speaker- and Phone-aware Convolutional Transformer Network for Acoustic Echo Cancellation
Chang Han, Weiping Tu, Yuhong Yang 0001
INTERSPEECH3
2022 Mandarin Lombard Grid: a Lombard-grid-like corpus of Standard Chinese
Yuhong Yang 0001, Xufeng Chen, Qingmu Liu, Weiping Tu, Hongyang Chen 0004, Linjun Cai
INTERSPEECH1
2022 Error model and simulation for multisource fusion indoor positioning
abstract
Seamless positioning services are of a critical concern in building smart cities. In a multisource fusion indoor positioning system, providing the guidance information for the deployment of positioning sources is a key technology, which can optimize the infrastructure resources to provide higher positioning accuracy. The error models of single-source positioning such as the received signal strength (RSS) fingerprint and the pedestrian dead reckoning (PDR) should be extended to meet the requirement of multisource indoor positioning for positioning error estimation. This paper proposes a model that combines the RSS fingerprint and PDR positioning error models for fusion positioning error simulation, which weights the PDR and RSS fingerprint positioning results and calculates the mean square error for the fusion positioning according to their positioning variances. This model is also used to establish an indoor positioning simulation system. To validate the proposed model, an experiment is performed which compared the actual positioning errors using the fusion positioning with the errors of the simulate model. The results show that the actual positioning error curves and the error curve predicted by the model are consistent. As a result, the proposed error model provides a solution for optimizing the deployment of positioning sources.
Haojun Ai, Jingjie Tao, Shan Ai, Tianshui Xu, Ning Li 0050, Kaifeng Tang, Yuhong Yang 0001, Shengchen Li
Int. J. Intell. Syst.8
2021 When Face Recognition Meets Occlusion: A New Benchmark
abstract
The existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models trained on existing datasets are almost ineffective for heavy occlusion. To this end, we pioneer a simulated occlusion face recognition dataset. In particular, we first collect a variety of glasses and masks as occlusion, and randomly combine the occlusion attributes (occlusion objects, textures,and colors) to achieve a large number of more realistic occlusion types. We then cover them in the proper position of the face image with the normal occlusion habit. Furthermore, we reasonably combine original normal face images and occluded face images to form our final dataset, termed as Webface-OCC. It covers 804,704 face images of 10,575 subjects, with diverse occlusion types to ensure its diversity and stability. Extensive experiments on public datasets show that the ArcFace retrained by our dataset significantly outperforms the state-of-the-arts. Webface-OCC is available at https://github.com/Baojin-Huang/Webface-OCC.
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Kangli Zeng, Zhen Han 0002, Xin Tian 0006, Yuhong Yang 0001
ICASSP8
2020 Attention-Guided Deraining Network Via Stage-Wise Learning
abstract
Due to diverse rain shapes, directions, densities as well as different distances to cameras, rain streaks in the air are interweaved and overlapped. However, most existing deraining methods are inherently oblivious this phenomenon and tend to learn a single rain streak layer to simulate this complex distribution, consequently failing to restore high-quality rain-free images. To solve this problem, along with the stage-wise learning, we propose a novel attention-guided deraining network (ADN) for rain streak removal. Specially, we decompose the rain streaks into multiple rain streak layers, and individually model them along the stages of the network to match the increasing abstracts. Moreover, the attention mechanism is utilized to guide the fusion of these rain streak layers by handling the overlaps between them. Extensive experiments on several benchmark datasets and real-world scenarios show substantial improvements both on quantitative indicators and visual effects over the current top-performing methods.
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Yuhong Yang 0001, Xin Tian 0006, Junjun Jiang
ICASSP5
2020 Constrained Ratio Mask for Speech Enhancement Using DNN
Hongjiang Yu, Wei-Ping Zhu 0001, Yuhong Yang 0001
INTERSPEECH3
2019 Kullback-Leibler Divergence Frequency Warping Scale for Acoustic Scene Classification Using Convolutional Neural Network
abstract
Most of current best performing Acoustic Scene Classification (ASC) systems utilize Mel scale spectrograms with Convolutional Neural Networks (CNNs). Mel scale is a common way to suit frequency warping of human ears, with strict decreasing frequency resolution on low to high frequency range. However, we find that significant frequency bins are located at mid to high frequency range for some acoustic scenes, such as travelling by bus, tram or train. In this paper, we show that a better frequency warping scale for ASC can be automatically learned from raw spectrograms, using Kullback-Leibler (KL) divergence scale. Our KL scale spectrograms with CNN method is evaluated on two public ASC datasets. The results show that we outperform the Mel scale method on both datasets. In addition, we also employ a Conditional Generative Adversarial Nets (Conditional-GAN) model for data augmentation, to prevent overfitting problem and allow further improvements on ASC.
Yuhong Yang 0001, Weiping Tu, Haojun Ai, Linjun Cai, Ruimin Hu
ICASSP1
2019 Towards a real-time production of immersive spatial audio of high individuality with an RBF neural network
Weiping Tu, Yuhong Yang 0001, Bo Du 0001, Jiaxi Zheng, Shuangxing Zhai
J. Parallel Distributed Comput.2
2018 The BLE Fingerprint Map Fast Construction Method for Indoor Localization
Haojun Ai, Yuhong Yang 0001
ICA3PP (4)3
2018 An RNN-Based Speech-Music Discrimination Used for Hybrid Audio Coder
Wanzhao Yang, Weiping Tu, Jiaxi Zheng, Yuhong Yang 0001, Yucheng Song
MMM (1)5
2017 Sound physical property matching between non central listening point and central listening point for NHK 22.2 system reproduction
abstract
NHK has proposed a famous 3D audio system: 22.2 multi-channel system, but its loudspeakers are too many and are troublesome to put in home. Ando and Wang has proposed two simplification methods to reduce its channel number, but only 3D sound field at the central listening point can be recovered well by NHK 22.2 system and its simplified systems, the listening experience at a non central listening point is worse than that at the central listening point. In real life, listeners may stay at arbitrary listening point: central or non central point. Conventional pressure matching and particle matching method could be used for non central zone sound field reproduction, but they have some theoretical shortcomings. To address these problems, this paper propose a universal non central listening point sound field reproduction method by matching sound physical property between a non central listening point and the central listening point. Subjective and objective experiments show the effectiveness of the proposed method.
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
ICASSP6
2017 3D Sound Field Reproduction at Non Central Point for NHK 22.2 System
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
MMM (1)5
2016 Level Ratio Based Inter and Intra Channel Prediction with Application to Stereo Audio Frame Loss Concealment
Yuhong Yang 0001, Yanye Wang, Ruimin Hu, Hongjiang Yu, Song Wang 0011
MMM (1)1
2016 Spatial Constrained Fine-Grained Color Name for Person Re-identification
Yang Yang 0062, Yuhong Yang 0001, Mang Ye, Wenxin Huang, Zheng Wang 0007, Chao Liang 0001, Chunjie Zhang 0001
MMM (1)2
2015 A down-mixing method for 22.2 multichannel system reproduction
abstract
This paper proposes a general multichannel system reproduction method. Firstly, relative to original multichannel system, a general global model is build up by guaranteeing sound pressure and the direction of particle velocity at the receiving point constant, and making the square error of particle velocity magnitude at the receiving point as little as possible. Then the model is equivalent to a least squares problems with non-negative constraints, it can be worked out by existing mature algorithms, and the global optimal solution of simplifying multichannel system are obtained. The proposed method can be used to simplify 22.2 multichannel system to 10.2 and 8.2 multichannel system, objective and subjective experimental results demonstrate that it performs better than traditional method.
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
ICASSP5
2015 Azimuthal Perceptual Resolution Model Based Adaptive 3D Spatial Parameter Coding
Ruimin Hu, Yuhong Yang 0001, Weiping Tu, Tingzhao Wu
MMM (1)3
2015 Signal-Aware Parametric Quality Model for Audio and Speech over IP Networks
Songbo Xie, Yuhong Yang 0001, Ruimin Hu, Yanye Wang, Hongjiang Yu, ShaoLong Dong
MMM (1)2
2014 Joint speech/audio coding based scalable perceptual audio coding
abstract
With the technical evolution of global mobile communications, various heterogeneous communication environments, frequently fluctuant bandwidth and multiform signals put new challenges to coding technology of multimedia signals. Scalable Audio Coding (SAC) can provide smooth transition between different coding qualities, which is an optimal choice for coding audio signals of different types and can produce more reliable and consistent service quality in multimedia communications. A scalable audio coding system based on joint speech/audio coding method and an auditory perceptual importance model based on bit-plane are proposed here. Both the audio content and the network bandwidth fluctuation will be considered in the system to obtain stable service qualities in mobile multimedia services. Experimental results indicate that with the same bit rates the subjective quality of proposed method is slightly better than G.729.1 and the SNR is improved by 0.3dB.
Ruimin Hu, Yuhong Yang 0001
ICIS3
2014 A spatial priority based scalable audio coding
abstract
A spatial priority scheme for scalable audio coding is presented in this paper. To improve the coding quality of important sounds with high attention, especially the moving sound, spatial information is introduced to assign the priorities of frequency subbands. Spatial cues and distance features are extracted in frequency subbands to represent the sound with fast changing direction and distance. Coding priorities are assigned to different frequency subbands according to the energy and spatial information. With trivial added side information and complexity, experimental results show that the perceptual quality is improved especially for the sound with high attention, especially the moving sound in scalable audio coding.
Ruimin Hu, Yuhong Yang 0001
ICASSP3
2014 Auditory attention based mobile audio quality assessment
abstract
Mobile audio services are growing with rising popularity of smart mobile devices using WiFi or cellular networks. A major issue facing mobile audio quality assessment is occasional background noises due to the prospect of sound recording at anytime and anywhere with smart mobile devices. Psychological study reveals that people pay selective attention to their interested sound in complex auditory input. In this paper, we model the mobile audio objective quality assessment based on auditory attention mechanism, with attention based horizontal azimuth parameters and timbre distortion parameters as additional Model Output Variables (MOVs). The results show that the prediction accuracy can be obtained by using such a method.
Yuhong Yang 0001, Hongjiang Yu, Ruimin Hu, Song Wang 0011, Qing Zhai, Songbo Xie
ICASSP1
2014 A 3D audio coding technique based on extracting the distance parameter
abstract
This paper presents a compression technique to improve the quality of three-dimensional (3D) audio produced by multiple loudspeaker channels or by headphone. The approach is based on extracting the side information of spatial sound sources within the three-dimensional space when capturing the sound sources. Different from other compression technique, the distances of sound sources are included in the side information. The separated signals of different sound sources are downmixed into one mono or stereo audio signal with the side information. The resulting downmixed signal is then compressed with traditional audio coder, resulting in a better perceptual quality of 3D audio by adding the distance parameter in the side information, and maintaining a low bit rates comparable with directional audio coding (DirAC).
Ruimin Hu, Liuyue Su, Weiping Tu, Xiaochen Wang 0001, Yuhong Yang 0001, Shi Dong 0004, Song Wang 0011, Maosheng Zhang, Furong Lei, Shiqing Li
ICME6
2013 Sound intensity and particle velocity based three-dimensional panning methods by five loudspeakers
abstract
In this paper, we present two new 3D panning methods. One method guarantees that time-average sound intensity and sound pressure of a virtual sound source at the receiving point are the same as time-averaged sound intensity and sound pressure of five loudspeakers at the receiving point. Another method maintains the direction of particle velocity and sound pressure. These methods can realize using five loudspeakers to replace a virtual sound source. These approaches relies on the assumption that the virtual sound source and five loudspeakers are on the same sphere and the virtual sound source needs to be in the area of a spherical pentagon that consists of the five loudspeakers. In the situation of five loudspeakers replacing a virtual source, these new methods do not need to undertake loudspeakers grouping. Compared with traditional 3D panning methods, these new methods are more convenient.
Song Wang 0011, Ruimin Hu, Yuhong Yang 0001
ICME4