Weiping Tu

dblp:119/0299 · DBLP profile ↗
← Back
81ranked-venue papers
1as first author
58since 2021 · last 2026
0000-0002-6933-3298ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 36 since 2021Artificial intelligence and machine learning · 34 · 31 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Security and privacy · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery Localization
abstract
Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments in video and audio, offering strong interpretability for security and forensics. While recent State Space Models (SSMs) show promise in precise temporal reasoning, their use in TFL is hindered by ambiguous boundaries, sparse forgeries, and limited long-range modeling. We propose DeformTrace, which enhances SSMs with deformable dynamics and relay mechanisms to address these challenges. Specifically, Deformable Self-SSM (DS-SSM) introduces dynamic receptive fields into SSMs for precise temporal localization. To further enhance its capacity for temporal reasoning and mitigate long-range decay, a Relay Token Mechanism is integrated into DS-SSM. Besides, Deformable Cross-SSM (DC-SSM) partitions the global state space into query-specific subspaces, reducing non-forgery information accumulation and boosting sensitivity to sparse forgeries. These components are integrated into a hybrid architecture that combines the global modeling of Transformers with the efficiency of SSMs. Extensive experiments show that DeformTrace achieves state-of-the-art performance with fewer parameters, faster inference, and stronger robustness.
Suting Wang, Yuanming Zheng, Junqi Yang, Yangxu Liao, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001
AAAI7
2026 Same Last-Item Confusion Unveiled: A Unified Mitigation Framework for Graph Learning in Session-Based Recommendation
abstract
Session-based recommendation (SBR), which focuses on next-item prediction for anonymous users based on short-term interaction sequences, has garnered increasing attention from researchers. While graph neural networks (GNNs) have become predominant in modeling complex item transition patterns, our empirical study reveals two critical limitations in existing GNN-based SBR methods. On the one hand, they struggle to differentiate between sessions sharing the same last item, resulting in indistinguishable session representations. On the other hand, the inherent popularity bias in session data leads to the over-recommendation of popular items. Inspired by contrastive learning techniques, this paper presents a unified mitigation framework for Same lAst-item confusion in Graph lEarning (SAGE) for SBR. In SAGE, we first obtain normalized session embeddings on constructed session graphs. We then build positive and negative samples of sessions through dual forward propagations and a novel negative sample selection strategy, followed by calculating contrastive loss. Finally, the enhanced session embeddings are utilized for prediction. Extensive experiments on two real-world datasets demonstrate that integrating SAGE with various state-of-the-art GNN-based SBR methods significantly improves their original performances.
Jinpeng Chen 0001, Jianxiang He, Yuan Cao 0003, Huan Li 0003, Zhenye Yang, Kaimin Wei, Xiongnan Jin, Senzhang Wang, Weiping Tu
WWW9
2026 AudioJailbreak: Jailbreak Attacks Against End-to-End Large Audio-Language Models
abstract
Jailbreak attacks to Large audio-language models (LALMs) are studied recently, but they exclusively focused on the attack scenario where the adversary can fully manipulate user prompts (named strong adversary) and limited in effectiveness, applicability, and practicability. In this work, we first conduct an extensive evaluation showing that advanced text jailbreak attacks cannot be easily ported to end-to-end LALMs via text-to-speech (TTS) techniques. We then propose AUDIOJAILBREAK, a novel audio jailbreak attack, featuring (1) asynchrony: the jailbreak audios do not need to align with user prompts in the time axis by crafting suffixal jailbreak audios; (2) universality: a single jailbreak perturbation is effective for different prompts by incorporating multiple prompts into the perturbation generation; (3) stealthiness: the malicious intent of jailbreak audios is concealed by proposing various intent concealment strategies; and (4) over-the-air robustness: the jailbreak audios remain effective when being played over the air by incorporating reverberation into the perturbation generation. In contrast, all prior audio jailbreak attacks cannot offer asynchrony, universality, stealthiness, and/or over-the-air robustness. Moreover, AUDIOJAILBREAK is also applicable to a more practical and broader attack scenario where the adversary cannot fully manipulate user prompts (named weak adversary). Extensive experiments with thus far the most LALMs demonstrate the high effectiveness of AUDIOJAILBREAK, in particular, it can jailbreak openAI's GPT-4o-Audio and bypass Meta's Llama-Guard-3 safeguard, in the weak adversary scenario. We highlight that our work peeks into the security implications of audio jailbreak attacks against LALMs, and realistically fosters improving their robustness, especially for the newly proposed weak adversary.
Guangke Chen, Fu Song, Zhe Zhao 0007, Xiaojun Jia, Yang Liu 0003, Yanchen Qiao, Weizhe Zhang, Weiping Tu, Yuhong Yang 0001, Bo Du 0001
IEEE Trans. Dependable Secur. Comput.8
2025 DREAM-OSA: Dual-Modal Transformer Framework for Early Warning of Obstructive Sleep Apnea via Transitional States Detection
abstract
Accurate early detection of obstructive sleep apnea (OSA) is critical for enabling timely auto-adjusting positive airway pressure (APAP) interventions. However, existing methods largely rely on binary classification (normal vs. apnea), failing to capture the transitional state preceding OSA onset and inducing therapy delays-hampered by open issues of ambiguous biomarkers and signal temporal misalignment. To address this, this study redefines sleep physiology into three distinct states: normal breathing, pre-apnea transitional (30 s pre-onset), and apnea. This study further proposes DREAM-OSA, a dual-modal transformer framework that specifically targets the transitional states, providing APAP with a sufficient advance response window (up to 10 s) to enable true early prediction of OSA events. It synergizes the complementary electroencephalogram (EEG) and respiratory signals through: 1) Modality-Specific Tokenization: EEG (decomposed into$\delta, \theta, \alpha, \beta, \gamma$bands) and respiratory signals are segmented into 1 s patches, encoded via dedicated 1DCNNs while preserving temporal-spectral structural information through learnable embeddings; and 2) Hierarchical Attention: Intra-modal self-attention captures temporal-spectral dynamics within each modality, while inter-modal cross-attention models bidirectional EEG-respiratory interactions. Evaluated on the MASS-SS1 dataset vs. the state-of-the-art methods towards real-time OSA early warning, DREAM-OSA achieves: overall accuracy up to 95.0%, and per-class F1-scores reaching 91.9% (normal), 94.9% (transitional), and 97.0% (apnea), demonstrating significantly more reliable detection of the transitional states, whereas its counterparts face performance bottleneck.
Qiyuan Yang, Dan Chen 0001, Feng Leng, Yiping Zuo, Weiping Tu, Xiaoli Li 0002
BIBM6
2025 FreqSense: Universal and Low-Latency Adversarial Example Detection for Speaker Recognition with Interpretability in Frequency Domain
abstract
Speaker recognition (SR) systems are particularly vulnerable to adversarial example (AE) attacks. To mitigate these attacks, AE detection systems are typically integrated into SR systems. To overcome the limitations of low detection accuracy, poor generalization, and high latency in existing schemes, this paper proposes FreqSense, an AE detection scheme based on frequency distribution features. FreqSense detects a variety of unknown AE attacks with low latency, and provides interpretability in its detection process. The basic idea of FreqSense is that AE typically introduce carefully designed noise in specific frequency bands that are associated with highly distinctive speaker identities. Therefore, leveraging the distributional variations in these frequency bands can effectively distinguish between AE and benign audio. FreqSense models frequency distribution features by integrating time-frequency transformation technology with a self-attention mechanism and employs a neural network-based classifier to distinguish between AE and benign audio. Experimental results show that FreqSense achieves an overall detection accuracy of 99.2%, surpassing state-of-the-art (SOTA) schemes by 27.2%. When confronting unknown AE attacks, FreqSense achieves a detection accuracy of 98.3% with a latency of just 0.0014 seconds.
Yihuan Huang, Yanzhen Ren, Weiping Tu, Yuhong Yang 0001
ICASSP4
2025 Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
abstract
Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same SSM to all sub-bands limits its inference capability. Additionally, when processing each time frame of the time-frequency representation, the SSM may forget certain high-frequency information of low energy, making the restoration of structure in the high-frequency bands challenging. For this reason, we propose Cross- and Sub-band Mamba (CSMamba). To assist the SSM in handling different sub-band features flexibly, we propose a band split block that splits the full-band into four sub-bands with different widths based on their information similarity. We then allocate independent weights to each sub-band, thereby reducing the inference burden on the SSM. Furthermore, to mitigate the forgetting of low-energy information in the high-frequency bands by the SSM, we introduce a spectrum restoration block that enhances the representation of the cross-band features from multiple perspectives. Experimental results on the DNS Challenge 2021 dataset demonstrate that CSMamba outperforms several state-of-the-art (SOTA) speech enhancement methods in three objective evaluation metrics with fewer parameters.
Jizhen Li, Weiping Tu, Yuhong Yang 0001, Xinmeng Xu, Yanzhen Ren
ICASSP2
2025 Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio Codecs
abstract
Existing end-to-end neural codecs have made great progress in preserving audio quality. Despite their success, they still face challenges in achieving accurate and efficient quantization. Specifically, these codecs often overlook which features have a greater impact on perceptual audio quality during quantization, leading to a quantization error distribution that fails to reflect the actual importance of latent features. They are also sensitive to unusual data points (outliers) because they use Mean Squared Error (MSE) to measure quantization errors, which can increase quantization noise or spectral artifacts. To address these limitations, we propose AW-CEQCodec which integrates an Attention Weighting (AW) module and a Conditional Entropy-driven Quantization (CEQ) loss. The AW enhances key regions of latent features before quantization, enabling more accurate quantizing critical features and reducing their quantization errors. After quantization, it restores global details from dequantized features, improving overall reconstruction. Moreover, the CEQ minimizes the uncertainty between latent and quantized features, effectively reflecting the distortion introduced by the quantization module. Experimental results on the CodecSuperb-STL dataset demonstrate that our method consistently outperforms baseline approaches, achieving superior audio quality at bitrates as low as 0.5 kbps, confirming its effectiveness in minimizing distortion and preserving perceptual quality. The reconstruction audio samples can be find at https://huazhi1024.github.io/first-page.
Weiping Tu, Yuhong Yang 0001, Xinmeng Xu, Yanzhen Ren
ICASSP2
2025 HAPG-SAQAM: Human Auditory Perception Guided Spatial Audio Quality Assessment Metric
abstract
Spatial audio quality evaluation is essential for applications like virtual and augmented reality, where accurate sound reproduction enhances user immersion. While subjective listening tests are the gold standard, they are costly and time-consuming. To address this, we propose HAPG-SAQAM, an objective metric for assessing timbre quality, spatial quality, and overall quality of binaural audio, guided by human auditory perception. Our contributions include: (1) the Multi-scale Auditory Guided Feature Extraction (MAGFE) module, incorporating gammatone frequency cepstral coefficients for better alignment with human perception; (2) Perceptual Weighted Loss (PWL), optimizing the weighting of timbre quality (TQ) and spatial quality (SQ) loss based on subjective test data; and (3) data augmentation techniques to enhance robustness by amplifying perceptual distortions. Experimental results show HAPG-SAQAM improves correlation with subjective scores by 10%, with ablation studies confirming the contributions of its components to enhanced spatial and overall audio quality.
Yuanming Zheng, Jiaxuan Yao, Xiangyu Deng, Yuhong Yang 0001, Ruiqi Liao, Weiping Tu, Cedar Lin
ICASSP6
2025 FIRING-Net: A filtered feature recycling network for speech enhancement
abstract
Current deep neural networks for speech enhancement (SE) aim to minimize the distance between the output signal and the clean target by filtering out noise features from input features. However, when noise and speech components are highly similar, SE models struggle to learn effective discrimination patterns. To address this challenge, we propose a Filter-Recycle-Interguide framework termed Filter-Recycle-INterGuide NETwork (FIRING-Net) for SE, which filters the input features to extract target features and recycles the filtered-out features as non-target features. These two feature sets then guide each other to refine the features, leading to the aggregation of speech information within the target features and noise information within the non-target features. The proposed FIRING-Net mainly consists of a Local Module (LM) and a Global Module (GM). The LM uses outputs of the speech extraction network as target features and the residual between input and output as non-target features. The GM leverages the energy distribution of self-attention map to extract target and non-target features guided by highest and lowest energy regions. Both LM and GM include interaction modules to leverage the two feature sets in an inter-guided manner for collecting speech from non-target features and filtering out noise from target features. Experiments confirm the effectiveness of the Filter-Recycle-Interguide framework, with FIRING-Net achieving a strong balance between SE performance and computational efficiency, surpassing comparable models across various SNR levels and noise environments.
Xinmeng Xu, Jizhen Li, Yuhong Yang 0001, Yong Luo 0002, Weiping Tu
ICLR6
2025 PGD-N2L: A Parameter-Guided Disentanglement Approach for Normal-To-Lombard Speech Conversion
abstract
The Normal-To-Lombard (N2L) speech conversion can effectively improve speech intelligibility in noisy communication scenarios and serve as a data augmentation tool for various speech-related algorithms. However, existing N2L methods did not aim to disentangle the Lombard effect from other speech attributes, leading to incomplete conversions. In this paper, we propose a Parameter-Guided Disentanglement approach for N2L speech conversion (PGD-N2L) which decomposes speech into linguistic content, speaker identity, and Lombard effect. To extract disentangled linguistic content, we propose a DeLomb-Based content encoder. To extract disentangled speaker identity and Lombard effect, we propose a style encoder that combines a fine-tuned speaker encoder and a learnable Lombard encoder to form a personalized style embedding. Furthermore, an En-Lomb-Based injection module is designed to accurately integrate the target Lombard effect and speaker identity into the linguistic content based on personalized style embedding, ensuring complete Lombard conversion. Experimental results demonstrate that our proposed method outperforms existing N2L models in speech intelligibility, acoustic similarity, and speech quality. Ablation studies confirm that the fine-tuned speaker encoder and the De-Lomb block effectively improve speech intelligibility and acoustic similarity, while the En-Lomb block enables the converted speech to more closely match the target Lombard speech.
Hongyang Chen 0004, Yuhong Yang 0001, Xinmeng Xu, Weiping Tu, Zhongyuan Wang 0001, Cedar Lin
ICME5
2025 Band-SCNet: A Causal, Lightweight Model for High-Performance Real-Time Music Source Separation
Junqi Yang, Yuhong Yang 0001, Weiping Tu, Cedar Lin
INTERSPEECH3
2025 FreeCodec: A Disentangled Neural Speech Codec with Fewer Tokens
Youqiang Zheng, Weiping Tu, Yueteng Kang, Li Xiao 0007, Yuhong Yang 0001
INTERSPEECH2
2025 Query-Based Audio-Visual Temporal Forgery Localization with Register-Enhanced Representation Learning
abstract
Temporal forgery in multimedia-where audio or video streams are subtly manipulated-poses critical challenges for content authenticity verification. While video-level detection has advanced, Temporal Forgery Localization (TFL) remains underexplored, often limited by weak audio-visual modeling and reliance on non-learnable post-processing. To address these challenges, we propose RegQAV, a Register-enhanced Query-based Audio-Visual framework for TFL. RegQAV exploits pretrained foundation models to capture fine-grained audio-visual correspondences and learnable registers are introduced to mitigate the model's tendency to overly focus on a limited set of temporal features. A query-based localization strategy enables end-to-end optimization without post-processing. We also introduce a Modality Fusion Adapter (MFA) for effective multi-scale integration of audio-visual data, a Deepfake Queries Generation (DQG) module for efficient query initialization, and a Poisson Count-Based Approach to dynamically predict the number of forgeries. Experiments on LAV-DF and AV-Deepfake1M show that RegQAV achieves state-of-the-art performance with fewer parameters, faster inference, and stronger generalization. This work offers significant potential for real-time deepfake detection and other multimedia verification applications. The code is available at https://github.com/zxd3099/RegQAV.
Suting Wang, Junqi Yang, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001
ACM Multimedia5
2025 Lombard-VLD: Voice Liveness Detection Based on Human Auditory Feedback
abstract
Voice Liveness Detection (VLD) aims to protect speaker authentication from speech spoofing by determining whether speeches come from live speakers or loudspeakers. Previous methods mainly focus on their differences at the signal level. In this paper, we propose the first VLD that uses the human auditory feedback mechanism (i.e., the Lombard effect), called Lombard-VLD. The key idea is that live speakers can physiologically and involuntarily adjust their speaking patterns in a noisy background but loudspeakers cannot. Moreover, we design a reference-based dual input mode and a differential SE-ResBlock to model the acoustic differences caused by the Lombard effect. Experimental results show that Lombard-VLD achieves 0% and 0.24% EER in two datasets, outperforming the state-of-the-art methods. It is robust to various environmental factors, including different distances, postures of the speaker, and environmental noise, with an average accuracy of over 98.51%. It also has a good generalization to unseen speakers, genders, and datasets, with EER lower than 2.68%, 3.44%, and 7.32%, respectively. This work shows the advantages of the Lombard effect in VLD, which has fewer user limitations and better detection performance.
Hongcheng Zhu, Zongkun Sun, Yanzhen Ren, Kun He 0008, Yongpeng Yan, Wuyang Liu, Yuhong Yang 0001, Weiping Tu
SP9
2025 Spatial information aided speech and noise feature discrimination for Monaural speech enhancement
Xinmeng Xu, Jizhen Li, Weiping Tu, Yuhong Yang 0001
Expert Syst. Appl.4
2025 From hippocampal neurons to broad spiking neural networks
Yiping Zuo, Dan Chen 0001, Weiping Tu, Albert Y. Zomaya, Xiaoli Li 0002
Neurocomputing4
2025 Self-training EEG discrimination model with weakly supervised sample construction: An age-based perspective on ASD evaluation
Tengfei Gao, Dan Chen 0001, Meiqi Zhou, Yiping Zuo, Weiping Tu, Xiaoli Li 0002, Jingying Chen 0001
Neural Networks6
2025 Lightweight real-time speech enhancement: State-space models and multi-spectral scanning techniques
Junqi Yang, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001
Neural Networks4
2024 SnoreOxiNet: Non-contact Diagnosis of Nocturnal Hypoxemia Using Cross-Domain Acoustic Features
Weiyan Yi, Xiuping Yang, Li Xiao 0007, Weiping Tu, Yuhong Yang 0001
ICANN (8)4
2024 EMALG: An Enhanced Mandarin Lombard Grid Corpus with Meaningful Sentences
abstract
This study investigates the Lombard effect, where individuals adapt their speech in noisy environments. We introduce an enhanced Mandarin Lombard grid (EMALG) corpus with meaningful sentences, enhancing the Mandarin Lombard grid (MALG) corpus. EMALG features 34 speakers and improves recording setups, addressing challenges faced by MALG with nonsense sentences. Our findings reveal that in Mandarin, meaningful sentences are more effective in enhancing the Lombard effect. Additionally, we uncover that female exhibit a more pronounced Lombard effect than male when uttering meaningful sentences. Moreover, our results reaffirm the consistency in the Lombard effect comparison between English and Mandarin found in previous research.
Baifeng Li, Qingmu Liu, Yuhong Yang 0001, Hongyang Chen 0004, Weiping Tu
ICASSP5
2024 Curricular Contrastive Regularization for Speech Enhancement with Self-Supervised Representations
abstract
Existing deep learning-based speech enhancement methods only adopt clean speech as positive samples to guide the training of speech enhancement networks while negative samples, i.e., noisy speech, are unexploited. In this paper, we adopt contrastive regularization (CR) built upon contrastive learning to exploit both the information of noisy and clean speech as negative and positive samples, respectively. Particularly, CR minimizes the distance between clean and enhanced speech and maximizes the distance between noisy and enhanced speech in the representation space of the self-supervised learning model. However, the contrastive samples are non-consensual, as the negatives are usually represented distantly from the clean speech, leaving the solution space still under-constricted. To tackle this issue, we provide the negative samples assembled from (1) the noisy speech, and (2) the corresponding enhanced speech without using CR, and we customize a curriculum learning strategy to define the importance of these negative samples to balance the learning difficulty caused by different similarities between the embeddings of the positive and negative samples. Experiments show that our proposal improves SE performance effectively without introducing additional computation/parameters.
Xinmeng Xu, Chang Han, Weiping Tu, Yuhong Yang 0001
ICASSP4
2024 An Efficient and Interpre Table Speech Enhancement Network Via Deep Dictionary Learning
abstract
Speech enhancement is a vital and highly ill-posed problem for many speech downstream tasks. While currently existing deep learning based speech enhancement methods have held state-of-the-art results, they still possess apparent shortcomings in that most of the deep learning based models lack interpretability. This deficiency results in unsatisfied speech enhancement performance in many sophisticated scenarios. To tackle this problem, we integrate dictionary learning and sparse coding into deep learning networks for speech enhancement and present a deep dictionary learning based speech enhancement network (DicLSENet). Specifically, the proposed DicLSENet strictly follows the principle of dictionary learning, learns the priors for both representation coefficients and dictionaries, and adaptively adjusts the dictionary for each input. Experimental results show that the proposed model outperforms state-of-the-art fully deep learning based methods with attractive computational costs.
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
ICASSP3
2024 Improving Acoustic Echo Cancellation by Exploring Speech and Echo Affinity with Multi-Head Attention
abstract
Deep learning-based approaches formulate acoustic echo cancellation (AEC) as a supervised speech separation task, where the mixture signal and the far-end signal are combined directly before or after the encoding stage. However, the mixture signal and the far-end signal are not integrated sufficiently due to the lack of interpretability for the affinity between speech and echo in a noisy mixture. In this paper, we propose DCA-Net, a dual-branch cross-attention neural network, to improve AEC performance by exploring the affinities between speech and echo in the representation space. In particular, the two branches predict speech and echo, respectively, and an interaction module is designed at several intermediate feature domains between the two branches to learn the correlations between these features of the two branches. Such an interaction can leverage features learned from one branch to restore missing information or counteract undesired information of the other by calculating the similarity between these features of two branches using multi-head cross attention. Evaluation results show that the proposed DCA-Net effectively suppresses acoustic echo and noise while preserving good speech quality.
Xinmeng Xu, Weiping Tu
ICASSP3
2024 Srcodec: Split-Residual Vector Quantization for Neural Speech Codec
abstract
End-to-end neural speech coding achieves state-of-the-art performance by using residual vector quantization. However, it is a challenge to quantize the latent variables with as few bits as possible. In this paper, we propose SRCodec, a neural speech codec that relies on a fully convolutional encoder/decoder network with specifically proposed split-residual vector quantization. In particular, it divides the latent representation into two parts with the same dimensions. We utilize two different quantizers to quantize the low-dimensional features and the residual between the low- and high-dimensional features. Meanwhile, we propose a dual attention module in split-residual vector quantization to improve information sharing along both dimensions. Both subjective and objective evaluations demonstrate that the effectiveness of our proposed method can achieve a higher quality of reconstructed speech at 0.95 kbps than Lyra-v1 at 3 kbps and Encodec at 3 kbps.
Youqiang Zheng, Weiping Tu, Li Xiao 0007, Xinmeng Xu
ICASSP2
2024 SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
abstract
Neural speech coding is a rapidly developing topic, where state-of-the-art approaches now exhibit superior compression performance than conventional methods. Despite significant progress, existing methods still have limitations in preserving and reconstructing fine details for optimal reconstruction, especially at low bitrates. In this study, we introduce SuperCodec, a neural speech codec that achieves state-of-the-art performance at low bitrates. It employs a novel back projection method with selective feature fusion for augmented representation. Specifically, we propose to use Selective Up-sampling Back Projection (SUBP) and Selective Down-sampling Back Projection (SDBP) modules to replace the standard up- and down-sampling layers at the encoder and decoder, respectively. Experimental results show that our method outperforms the existing neural speech codecs operating at various bitrates. Specifically, our proposed method can achieve higher quality reconstructed speech at 1 kbps than Lyra V2 at 3.2 kbps and Encodec at 6 kbps.
Youqiang Zheng, Weiping Tu, Li Xiao 0007, Xinmeng Xu
ICASSP2
2024 LungAdapter: Efficient Adapting Audio Spectrogram Transformer for Lung Sound Classification
Li Xiao 0007, Lucheng Fang, Yuhong Yang 0001, Weiping Tu
INTERSPEECH4
2024 Exploring Sentence Type Effects on the Lombard Effect and Intelligibility Enhancement: A Comparative Study of Natural and Grid Sentences
Hongyang Chen 0004, Yuhong Yang 0001, Zhongyuan Wang 0001, Weiping Tu, Haojun Ai, Cedar Lin
INTERSPEECH4
2024 Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
Jizhen Li, Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
INTERSPEECH3
2024 SimuSOE: A Simulated Snoring Dataset for Obstructive Sleep Apnea-Hypopnea Syndrome Evaluation during Wakefulness
Xiuping Yang, Li Xiao 0007, Weiyan Yi, Yuhong Yang 0001, Weiping Tu
INTERSPEECH7
2024 V2IED: Dual-view learning framework for detecting events of interictal epileptiform discharges
Zhekai Ming, Dan Chen 0001, Tengfei Gao, Yunbo Tang, Weiping Tu, Jingying Chen 0001
Neural Networks5
2024 Adaptive selection of local and non-local attention mechanisms for speech enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
Neural Networks2
2024 Deep Hashing Network With Hybrid Attention and Adaptive Weighting for Image Retrieval
abstract
Due to the low computational cost of Hamming distance, hashing-based image retrieval has been universally acknowledged. Therefore, it is becoming increasingly important to quickly generate high-precision hash codes (also hash features) from images. However, the existing deep hashing methods are vulnerable to image content variations; that is, it is difficult to generate stable and consistent hash codes for similar images. In addition, generating hash codes of different lengths requires retraining the model, which is expensive in training time. To address these problems, this paper proposes a deep hashing network (DHN) with a hybrid attention mechanism and adaptive weighting (HAAW) learning. It mainly consists of a feature extraction module, feature refinement module, classification layer, hash layer and an adaptive weight layer. In particular, the hybrid attention mechanism combines bottom-up pixel saliency and top-down semantic constraints, in which the former is achieved through channel and spatial attention (CSA) and the latter is supervised by classification labels. In this way, it encourages the network to focus on dominant semantic features without being disturbed by irrelevant objects so that semantically similar images can be mapped to approximate hash codes. We further propose an adaptive weighting learning algorithm to generate weights for each bit of the hash code generated by the deep network. Then, we directly generate shorter hash codes from the available long hash code according to the importance of bits represented by the weights. This avoids retraining the network for learning hash codes of different lengths. Extensive experiments on public CIFAR-10, NUS_WIDE and ImageNet datasets show that our method has achieved substantial improvements over the counterparts in terms of precision and speed.
Yingjiao Pei, Zhongyuan Wang 0001, Heling Chen, Baojin Huang, Weiping Tu
IEEE Trans. Multim.6
2024 Auxiliary Information Guided Self-attention for Image Quality Assessment
abstract
Image quality assessment (IQA) is an important problem in computer vision with many applications. We propose a transformer-based multi-task learning framework for the IQA task. Two subtasks: constructing an auxiliary information error map and completing image quality prediction, are jointly optimized using a shared feature extractor. We use visual transformers (ViT) as a feature extractor for feature extraction and guide ViT to focus on image quality-related features by building auxiliary information error map subtask. In particular, we propose a fusion network that includes a channel focus module. Unlike the fusion methods commonly used in previous IQA methods, we use the fusion network, including the channel attention module, to fuse the auxiliary information error map features with the image features, which facilitates the model to mine the image quality features for more accurate image quality assessment. And by jointly optimizing the two subtasks, ViT focuses more on extracting image quality features and building a more precise mapping from feature representation to quality score. With slight adjustments to the model, our approach can be used in both no-reference (NR) and full-reference (FR) IQA environments. We evaluate the proposed method in multiple IQA databases, showing better performance than state-of-the-art FR and NR IQA methods.
Jifan Yang, Zhongyuan Wang 0001, Guangcheng Wang, Baojin Huang, Yuhong Yang 0001, Weiping Tu
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech Enhancement
abstract
Attention mechanisms, such as local and non-local attention, play a fundamental role in recent deep learning based speech enhancement (SE) systems. However, a natural speech contains many fast-changing and relatively briefly acoustic events, therefore, capturing the most informative speech features by indiscriminately using local and non-local attention is challenged. We observe that the noise type and speech feature vary within a sequence of speech and the local and non-local can respectively process different types of corrupted speech regions. To leverage this, we propose Selector-Enhancer, a dual-attention based convolution neural network (CNN) with a feature-filter that can dynamically select regions from low-resolution speech features and feed them to local or non-local attention operations. In particular, the proposed feature-filter is trained by using reinforcement learning (RL) with a developed difficulty-regulated reward that related to network performance, model complexity and “the difficulty of the SE task”. The results show that our method achieves comparable or superior performance to existing approaches. In particular, Selector-Enhancer is effective for real-world denoising, where the number and types of noise are varies on a single noisy mixture.
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
AAAI2
2023 MBMS-GAN: Multi-Band Multi-Scale Adversarial Learning for Enhancement of Coded Speech at Very Low Rate
Weiping Tu, Yong Luo 0002, Xin Zhou 0003, Li Xiao 0007, Youqiang Zheng
ICANN (7)2
2023 Freevc: Towards High-Quality Text-Free One-Shot Voice Conversion
abstract
Voice conversion (VC) can be achieved by first extracting source content information and target speaker information, and then reconstructing waveform with these information. However, current approaches normally either extract dirty content information with speaker information leaked in, or demand a large amount of annotated data for training. Besides, the quality of reconstructed waveform can be degraded by the mismatch between conversion model and vocoder. In this paper, we adopt the end-to-end framework of VITS for high-quality waveform reconstruction, and propose strategies for clean content information extraction without text annotation. We disentangle content information by imposing an information bottleneck to WavLM features, and propose the spectrogram-resize based data augmentation to improve the purity of extracted content information. Experimental results show that the proposed method outperforms the latest VC models trained with annotated data and has greater robustness.
Weiping Tu, Li Xiao 0007
ICASSP2
2023 Improving Acoustic Echo Cancellation by Mixing Speech Local and Global Features with Transformer
abstract
We propose MiT-Net, a novel mix-transformer neural network with a pyramid encoder operating in the time domain, for the task of acoustic echo cancellation. The MiT-Net formulates acoustic echo cancellation as a supervised speech separation problem, in which near-end speech is separated from a single microphone recording and sent to the far end, and consists of two key components. First, we apply a pyramid encoder, which adopts the coarse-to-fine structure, to extract the latent correlations between double-end signals and to fuse them in a multiscale manner. Second, we propose a mix-transformer, a combination of local and global attention in a parallel way, to leverage local and global speech information for separation. Experimental results show that the proposed method outperforms recent AEC methods in terms of objective evaluation metrics. In addition, exploring the correlation between speech local and global features by using the mix-transformer significantly improves the system performance and shows more robustness than the conventional transformer.
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001, Li Xiao 0007
ICASSP3
2023 Learning From Single-Expert Annotated Labels for Automatic Sleep Staging
abstract
Existing automatic sleep staging algorithms rely on accurately labeled data. However, due to the subjectivity of sleep experts, accurate labels must be obtained through joint labeling by multiple experts, which results in high time and labor costs. In this work, we treat labels mislabeled by a single expert as noisy labels and first propose SE-ASS, an automatic sleep staging learning framework based on single-expert annotated data. Since multiple models tend to produce inconsistent predictions for instances with incorrect labels during training, we use two networks with the same structure but different initializations and regularize them with a prediction consistency loss to prevent overfitting to noisy labels. Furthermore, we use a contrastive loss between models to enhance the exploration of feature representations without relying on potentially noisy labels. Our results on two publicly available datasets show that SE-ASS can effectively improve the performance of automatic sleep staging models trained on single-expert annotated datasets.
Zhiheng Luan, Yanzhen Ren, Xiuping Yang, Weiping Tu, Yuhong Yang 0001
ICASSP6
2023 PMMSD: Development of the Matrix Sentence Intelligibility Dataset for Mandarin with Lombard Effect
abstract
This paper presents a Paired Mandarin Matrix Sentence Dataset (PMMSD), which will be available after publication. PMMSD is the first Mandarin matrix sentence intelligibility dataset containing both plain and Lombard speech for scientific research. The results verify that different Lombard styles would affect word intelligibility to different degrees and the Lombard effect helps maintain homogeneous intelligibility against contextual interference. All of the discoveries indicate that the Lombard effect should be considered when building intelligibility datasets with noise in the future.
Hanchen Pei, Yuhong Yang 0001, Xufeng Chen, Qingmu Liu, Hongyang Chen 0004, Weiping Tu
ICASSP6
2023 ONEI: Unveiling Route and Phase of Breathing from Snoring Sounds
Baoai Han, Li Xiao 0007, Xiuping Yang, Weiping Tu, Weiyan Yi, Yuhong Yang 0001, Yanzhen Ren
ICONIP (9)5
2023 Leveraging Sound Local and Global Features for Language-Queried Target Sound Extraction
Xinmeng Xu, Yuhong Yang 0001, Weiping Tu
ICONIP (4)4
2023 Exploring the Interactions Between Target Positive and Negative Information for Acoustic Echo Cancellation
Chang Han, Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
INTERSPEECH3
2023 A Snoring Sound Dataset for Body Position Recognition: Collection, Annotation, and Analysis
Li Xiao 0007, Xiuping Yang, Weiping Tu, Weiyan Yi, Yuhong Yang 0001, Yanzhen Ren
INTERSPEECH4
2023 PCNN: A Lightweight Parallel Conformer Neural Network for Efficient Monaural Speech Enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
INTERSPEECH2
2023 CQNV: A Combination of Coarsely Quantized Bitstream and Neural Vocoder for Low Rate Speech Coding
Youqiang Zheng, Li Xiao 0007, Weiping Tu, Yuhong Yang 0001, Xinmeng Xu
INTERSPEECH3
2023 Functional connectivity learning via Siamese-based SPD matrix representation of brain imaging data
Yunbo Tang, Dan Chen 0001, Jia Wu 0001, Weiping Tu, Jessica Monaghan, Paul F. Sowman, David McAlpine
Neural Networks4
2023 Corrigendum to "Functional Connectivity Learning via Siamese-based SPD Matrix Representation of Brain Imaging Data" [Neural Networks 163 (2023) 272-285]
Yunbo Tang, Dan Chen 0001, Jia Wu 0001, Weiping Tu, Jessica Monaghan, Paul F. Sowman, David McAlpine
Neural Networks4
2023 CASE-Net: Integrating local and non-local attention operations for speech enhancement
Xinmeng Xu, Weiping Tu, Yuhong Yang 0001
Speech Commun.2
2022 CS-CTCSCONV1D: Small footprint speaker verification with channel split time-channel-time separable 1-dimensional convolution
Linjun Cai, Yuhong Yang 0001, Xufeng Chen, Weiping Tu, Hongyang Chen 0004
INTERSPEECH4
2022 Speaker- and Phone-aware Convolutional Transformer Network for Acoustic Echo Cancellation
Chang Han, Weiping Tu, Yuhong Yang 0001
INTERSPEECH2
2022 Mandarin Lombard Grid: a Lombard-grid-like corpus of Standard Chinese
Yuhong Yang 0001, Xufeng Chen, Qingmu Liu, Weiping Tu, Hongyang Chen 0004, Linjun Cai
INTERSPEECH4
2022 Parallel discriminative subspace for city target detection from high dimension images
Yipeng Zhang 0001, Yiming Zhang 0027, Bo Du 0001, Weiping Tu
GeoInformatica6
2022 VISEL: A visual and magnetic fusion-based large-scale indoor localization system with improved high-precision semantic maps
abstract
Multisource fusion localization is a mainstream scheme for acquiring accurate locations in complex indoor scenes. To overcome the interference of indoor structures on radio and illumination variation on visual features, the semantic maps provide an effective way for multisource fusion localization. However, due to the lack of visual depth information, solutions of indoor semantic maps suffer from large semantic segmentation errors for similar objects, which leads to the unstable performance of localization systems. To overcome the issue in semantic and fusion localization, we develop a localization system to demonstrate the use of restudy semantic map and self-adapting fusion localization would achieve centimeter-level positioning accuracy, termed VISEL. VISEL uses the proposed spatial attention-aware semantic model to enhance the discrimination of semantic features for capturing accurate semantic maps. On the basis of high-precision semantic maps, VISEL completes an enhanced particle filter fusion localization module with adaptive reassign weight to different localization modules, which successfully improves accuracy through complementary advantages between different signals while overcoming the drawbacks of each signal and interference of complex environment. The extensive experimental results show that VISEL outperforms current state-of-the-art positioning systems and achieves an average positioning accuracy of 0.4 m. VISEL utilizes semantic maps with depth features and enhanced particle filter to reduce the fusion localization error by 38%, which suggests the high-precision semantic maps with depth features could provide a robust solution for the fusion localization system for indoor complex scenes.
Ning Li 0050, Weiping Tu, Haojun Ai, Huimin Deng, Jingjie Tao, Tan Hu, Xu Sun 0010
Int. J. Intell. Syst.2
2021 Adaptive Texture Distillation Network for Image Hybrid Super-Resolution
abstract
To save the transmission bandwidth of high-resolution (HR) images, we can send down-sampled low-resolution (LR) images and reconstruct them using super-resolution (SR) technology at the receiving end. However, image down-sampling by a large factor results in the loss of many spatial details. Instead, we use a combination of spatial down-sampling by a small factor and gray-level quantization to obtain the low hybrid-resolution images. Although the small down-sampling factor makes images retain more spatial details and real textures, the gray-level quantization introduces fake textures. Obviously, the real textures should be enhanced, and the fake textures should be eliminated. To address this issue, we propose a lightweight Adaptive Texture Distillation Network (ATDN) for image hybrid super-resolution. Our model uses the texture enhancement block (TEB) and the texture smoothing block (TSB) to handle real and fake textures in different ways. Considering that the mixing proportions of two kinds of textures in low hybrid-resolution images vary with regions, we specifically use a cascaded weight branch to adaptively adjust the weights of real and fake textures. Experiments reveal that our model can effectively deal with the mixing problem of real and fake textures, and our method can achieve superior performance to other lightweight methods.
Chunlei Liu 0006, Zhen Han 0002, Jiaxing Wen, Zhongyuan Wang 0001, Weiping Tu
IJCNN6
2021 Metric Learning via Penalized Optimization
abstract
Metric learning aims to project original data into a new space, where data points can be classified more accurately using kNN or similar types of classification algorithms. To avoid trivial learning results such as indistinguishably projecting the data onto a line, many existing approaches formulate metric learning as a constrained optimization problem, like finding a metric that minimizes the distance between data points from the same class, with a constraint of ensuring a certain separation for data points from different classes, and then they approximate the optimal solution to the constrained optimization in an iterative way. In order to improve the classification accuracy as much as possible, we try to find a metric that is able to minimize the intra-class distance and maximize the inter-class distance simultaneously. Towards this, we formulate metric learning as a penalized optimization problem, and provide design guideline, paradigms with a general formula, as well as two representative instantiations for the penalty term. In addition, we provide an analytical solution for the penalized optimization, with which costly computation can be avoid, and more importantly, there is no need to worry about the convergence rates or approximation ratios any more. Extensive experiments on real-world data sets are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Yanan Peng, Ting Gan, Weiping Tu, Ruiting Zhou, Sai Wu
KDD4
2021 Optimization of sound fields reproduction based Higher-Order Ambisonics (HOA) using the Generative Adversarial Network (GAN)
Lingkun Zhang, Xiaochen Wang 0001, Ruimin Hu, Dengshi Li, Weiping Tu
Multim. Tools Appl.5
2021 Estimation of spherical harmonic coefficients in sound field recording using feed-forward neural networks
Lingkun Zhang, Xiaochen Wang 0001, Ruimin Hu, Dengshi Li, Weiping Tu
Multim. Tools Appl.5
2021 Incorporating Distribution Matching into Uncertainty for Multiple Kernel Active Learning
abstract
Due to the lack of the labeled data and the complex structures of various data, it is very hard to learn the uncertainty and representativeness accurately in active learning. In this paper, we propose a multiple kernel active learning framework that incorporates a group regularizer of distribution information into the estimation of uncertainty. The proposed method takes the advantage of multiple kernel learning to learn the kernel space in which the complex structures can be well captured by kernel weights. Meanwhile, we have developed an efficient optimization algorithm to solve the proposed method. Experimental results on twelve UCI benchmark data sets and eight subsets of ImageNet show that the proposed method outperforms several state-of-the-art active learning methods. Moreover, we also have applied the proposed method to multiple feature scenario on Caltech101, and the promising results are also obtained compared with single feature scenario.
Zengmao Wang, Bo Du 0001, Weiping Tu, Lefei Zhang, Dacheng Tao
IEEE Trans. Knowl. Data Eng.3
2020 Story segmentation for news broadcast based on primary caption
abstract
In the information explosion era, people only want to access the news information that they are interested in. News broadcast story segmentation is strongly needed, which is an essential basis for personalized delivery and short video. The existing advanced story boundary segmentation methods utilize semantic similarity of subtitles, thus entailing complex semantic computation. The title texts of news broadcast programs include headline (or primary) captions, dialogue captions and the channel logo, while the same story clips only render one primary caption in most news broadcast. Inspired by this fact, we propose a simple method for story segmentation based on the primary caption, which combines YOLOv3 based primary caption extraction and preliminary location of boundaries. In particular, we introduce mean hash to achieve the fast and reliable comparison for detected small-size primary caption blocks. We further incorporate scene recognition to exact the preliminary boundaries, because the primary captions always appear later than the story boundary. Experimental results on two Chinese news broadcast datasets show that our method enjoys high accuracy in terms of R, P and F1-measures.
Heling Chen, Zhongyuan Wang 0001, Yingjiao Pei, Baojin Huang, Weiping Tu
MMAsia5
2020 Video scene detection based on link prediction using graph convolution network
abstract
With the development of the Internet, multimedia data grows by an exponential level. The demand for video organization, summarization and retrieval has been increasing where scene detection plays an essential role. Existing shot clustering algorithms for scene detection usually treat temporal shot sequence as unconstrained data. The graph based scene detection methods can locate the scene boundaries by taking the temporal relation among shots into account, while most of them only rely on low-level features to determine whether the connected shot pairs are similar or not. The optimized algorithms considering temporal sequence of shots or combining multi-modal features will bring parameter trouble and computational burden. In this paper, we propose a novel temporal clustering method based on graph convolution network and the link transitivity of shot nodes, without involving complicated steps and prior parameter setting such as the number of clusters. In particular, the graph convolution network is used to predict the link possibility of node pairs that are close in temporal sequence. The shots are then clustered into scene segments by merging all possible links. Experimental results on BBC and OVSD datasets show that our approach is more robust and effective than the comparison methods in terms of F1-score.
Yingjiao Pei, Zhongyuan Wang 0001, Heling Chen, Baojin Huang, Weiping Tu
MMAsia5
2020 Loudspeaker triplet selection based on low distortion within head for multichannel conversion of smart 3D home theater
abstract
Summary In recent years, with the vigorous development of 3D film industry, the demand for 3D Smart Home Theater, based on Internet of Things (IoT), continues to grow. Theaters are populated with a large number of loudspeakers for more realistic 3D sound effects. However the number of loudspeakers in home is limited. Therefore, multichannel conversion is required to achieve theater 3D sound effects in home. Traditionally, the replaced loudspeaker signal of the original system is assigned to a “loudspeaker triplet” of the converted system. A large amount of subjective evaluations is necessary to judge the consistency of the replaced loudspeaker position with the “phantom source” positions reconstructed by loudspeaker triplets. In this study, after calculating the least‐squares errors of the reproduced sound field within a given region, we explore the constraint between the low distortion of the reproduced sound field and the loudspeaker triplet positions. Using this constraint rule, we present a new loudspeaker triplet selection criteria that can greatly reduce the number and time of subjective evaluations for selecting the optimal loudspeaker triplet. Simulation and subjective evaluation experiments indicate that the proposed selection method outperforms the traditional method, and that the proposed method can be successfully applied to multichannel conversion.
Dengshi Li, Ruimin Hu, Xiaochen Wang 0001, Weiping Tu
Concurr. Comput. Pract. Exp.4
2020 Local manifold sparse model for image classification
Fulin Luo, Yajuan Huang, Weiping Tu
Neurocomputing3
2020 Bi-adapting kernel learning for unsupervised domain adaptation
Zengmao Wang, Weiping Tu, Bo Du 0001, Yanxiang Cheng
Neurocomputing3
2020 Homologous Component Analysis for Domain Adaptation
abstract
Covariate shift assumption based domain adaptation approaches usually utilize only one common transformation to align marginal distributions and make conditional distributions preserved. However, one common transformation may cause loss of useful information, such as variances and neighborhood relationship in both source and target domain. To address this problem, we propose a novel method called homologous component analysis (HCA) where we try to find two totally different but homologous transformations to align distributions with side information and make conditional distributions preserved. As it is hard to find a closed form solution to the corresponding optimization problem, we solve them by means of the alternating direction minimizing method (ADMM) in the context of Stiefel manifolds. We also provide a generalization error bound for domain adaptation in semi-supervised case and two transformations can help to decrease this upper bound more than only one common transformation does. Extensive experiments on synthetic and real data show the effectiveness of the proposed method by comparing its classification accuracy with the state-of-the-art methods and numerical evidence on chordal distance and Frobenius distance shows that resulting optimal transformations are different.
Youfa Liu, Weiping Tu, Bo Du 0001, Lefei Zhang, Dacheng Tao
IEEE Trans. Image Process.2
2020 LogDet Metric-Based Domain Adaptation
abstract
Domain adaptation has proven to be successful in dealing with the case where training and test samples are drawn from two kinds of distributions, respectively. Recently, the second-order statistics alignment has gained significant attention in the field of domain adaptation due to its superior simplicity and effectiveness. However, researchers have encountered major difficulties with optimization, as it is difficult to find an explicit expression for the gradient. Moreover, the used transformation employed here does not perform dimensionality reduction. Accordingly, in this article, we prove that there exits some scaled LogDet metric that is more effective for the second-order statistics alignment than the Frobenius norm, and hence, we consider it for second-order statistics alignment. First, we introduce the two homologous transformations, which can help to reduce dimensionality and excavate transferable knowledge from the relevant domain. Second, we provide an explicit gradient expression, which is an important ingredient for optimization. We further extend the LogDet model from single-source domain setting to multisource domain setting by applying the weighted Karcher mean to the LogDet metric. Experiments on both synthetic and realistic domain adaptation tasks demonstrate that the proposed approaches are effective when compared with state-of-the-art ones.
Youfa Liu, Bo Du 0001, Weiping Tu, Mingming Gong, Yuhong Guo, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.3
2019 Kullback-Leibler Divergence Frequency Warping Scale for Acoustic Scene Classification Using Convolutional Neural Network
abstract
Most of current best performing Acoustic Scene Classification (ASC) systems utilize Mel scale spectrograms with Convolutional Neural Networks (CNNs). Mel scale is a common way to suit frequency warping of human ears, with strict decreasing frequency resolution on low to high frequency range. However, we find that significant frequency bins are located at mid to high frequency range for some acoustic scenes, such as travelling by bus, tram or train. In this paper, we show that a better frequency warping scale for ASC can be automatically learned from raw spectrograms, using Kullback-Leibler (KL) divergence scale. Our KL scale spectrograms with CNN method is evaluated on two public ASC datasets. The results show that we outperform the Mel scale method on both datasets. In addition, we also employ a Conditional Generative Adversarial Nets (Conditional-GAN) model for data augmentation, to prevent overfitting problem and allow further improvements on ASC.
Yuhong Yang 0001, Weiping Tu, Haojun Ai, Linjun Cai, Ruimin Hu
ICASSP3
2019 Towards a real-time production of immersive spatial audio of high individuality with an RBF neural network
Weiping Tu, Yuhong Yang 0001, Bo Du 0001, Jiaxi Zheng, Shuangxing Zhai
J. Parallel Distributed Comput.1
2019 Domain Adaptation With Discriminative Distribution and Manifold Embedding for Hyperspectral Image Classification
abstract
Hyperspectral remote sensing image classification has drawn a great attention in recent years due to the development of remote sensing technology. To build a high confident classifier, the large number of labeled data is very important, e.g., the success of deep learning technique. Indeed, the acquisition of labeled data is usually very expensive, especially for the remote sensing images, which usually needs to survey outside. To address this problem, in this letter, we propose a domain adaptation method by learning the manifold embedding and matching the discriminative distribution in source domain with neural networks for hyperspectral image classification. Specifically, we use the discriminative information of source image to train the classifier for the source and target images. To make the classifier can work well on both domains, we minimize the distribution shift between the two domains in an embedding space with prior class distribution in the source domain. Meanwhile, to avoid the distortion mapping of the target domain in the embedding space, we try to keep the manifold relation of the samples in the embedding space. Then, we learn the embedding on source domain and target domain by minimizing the three criteria simultaneously based on a neural network. The experimental results on two hyperspectral remote sensing images have shown that our proposed method can outperform several baseline methods.
Zengmao Wang, Bo Du 0001, Qian Shi 0001, Weiping Tu
IEEE Geosci. Remote. Sens. Lett.4
2019 A Secure AMR Fixed Codebook Steganographic Scheme Based on Pulse Distribution Model
abstract
Adaptive multi-rate (AMR), a popular audio compression standard, is widely used in mobile communication and mobile Internet applications and has become a novel carrier for hiding information. To improve the statistical security, this paper presents a steganographic scheme in the AMR fixed codebook (FCB) domain based on the pulse distribution model (PDM-AFS), which is obtained from the distribution characteristics of the FCB value in the cover audio. The pulse positions in stego audio are controlled by message encoding and random masking to make the statistical distribution of the FCB parameters close to that of the cover audio. The experimental results show that the statistical security of the proposed scheme is better than that of the existing schemes. Furthermore, the hiding capacity is maintained compared with the existing schemes. The average hiding capacity can reach 2.06 kbps at an audio compression rate of 12.2 kbps, and the auditory concealment is good. To the best of our knowledge, this is the first secure AMR FCB steganographic scheme that improves the statistical security based on the distribution model of the cover audio. This scheme can be extended to other audio compression codecs under the principle of algebraic code excited linear prediction (ACELP), such as G.723.1 and G.729.
Yanzhen Ren, Hanyi Yang, Hongxia Wu, Weiping Tu, Lina Wang 0001
IEEE Trans. Inf. Forensics Secur.4
2018 An RNN-Based Speech-Music Discrimination Used for Hybrid Audio Coder
Wanzhao Yang, Weiping Tu, Jiaxi Zheng, Yuhong Yang 0001, Yucheng Song
MMM (1)2
2017 Sound physical property matching between non central listening point and central listening point for NHK 22.2 system reproduction
abstract
NHK has proposed a famous 3D audio system: 22.2 multi-channel system, but its loudspeakers are too many and are troublesome to put in home. Ando and Wang has proposed two simplification methods to reduce its channel number, but only 3D sound field at the central listening point can be recovered well by NHK 22.2 system and its simplified systems, the listening experience at a non central listening point is worse than that at the central listening point. In real life, listeners may stay at arbitrary listening point: central or non central point. Conventional pressure matching and particle matching method could be used for non central zone sound field reproduction, but they have some theoretical shortcomings. To address these problems, this paper propose a universal non central listening point sound field reproduction method by matching sound physical property between a non central listening point and the central listening point. Subjective and objective experiments show the effectiveness of the proposed method.
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
ICASSP7
2017 Frame-Independent and Parallel Method for 3D Audio Real-Time Rendering on Mobile Devices
Yucheng Song, Xiaochen Wang 0001, Wei Chen 0143, Weiping Tu
MMM (2)6
2017 3D Sound Field Reproduction at Non Central Point for NHK 22.2 System
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
MMM (1)6
2016 Multichannel reduction based on sound field within two ears
abstract
People hope to use a small number of loudspeakers to get the experience of the film 3D sound at home. Considering that people use two ears to listen, this paper provides a method which reproduce the sound field within the region of two ears. We develop the fundamental performance limits for the truncated spherical harmonic function expansions of the sound field within the region of ears. Based on this, the low distortion of reproduced sound field within two ears is maintained in the processing of reducing loudspeakers from Q to Q-1. The 22.2 multichannel sound system without two low-frequency effect channels can be simplified to 6 channels automatically and the total of loudspeaker arrangements is ten. The subjective evaluation of the proposed method is better than that of the previous multichannel reduction method with the decrease of the number of loudspeakers.
Dengshi Li, Ruimin Hu, Xiaochen Wang 0001, Guo Wu, Weiping Tu
ICME6
2016 Adaptive Multichannel Reduction Using Convex Polyhedral Loudspeaker Array
Lingkun Zhang, Ruimin Hu, Dengshi Li, Xiaochen Wang 0001, Weiping Tu
MMM (1)5
2015 A down-mixing method for 22.2 multichannel system reproduction
abstract
This paper proposes a general multichannel system reproduction method. Firstly, relative to original multichannel system, a general global model is build up by guaranteeing sound pressure and the direction of particle velocity at the receiving point constant, and making the square error of particle velocity magnitude at the receiving point as little as possible. Then the model is equivalent to a least squares problems with non-negative constraints, it can be worked out by existing mature algorithms, and the global optimal solution of simplifying multichannel system are obtained. The proposed method can be used to simplify 22.2 multichannel system to 10.2 and 8.2 multichannel system, objective and subjective experimental results demonstrate that it performs better than traditional method.
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
ICASSP6
2015 Azimuthal Perceptual Resolution Model Based Adaptive 3D Spatial Parameter Coding
Ruimin Hu, Yuhong Yang 0001, Weiping Tu, Tingzhao Wu
MMM (1)5
2014 A 3D audio coding technique based on extracting the distance parameter
abstract
This paper presents a compression technique to improve the quality of three-dimensional (3D) audio produced by multiple loudspeaker channels or by headphone. The approach is based on extracting the side information of spatial sound sources within the three-dimensional space when capturing the sound sources. Different from other compression technique, the distances of sound sources are included in the side information. The separated signals of different sound sources are downmixed into one mono or stereo audio signal with the side information. The resulting downmixed signal is then compressed with traditional audio coder, resulting in a better perceptual quality of 3D audio by adding the distance parameter in the side information, and maintaining a low bit rates comparable with directional audio coding (DirAC).
Ruimin Hu, Liuyue Su, Weiping Tu, Xiaochen Wang 0001, Yuhong Yang 0001, Shi Dong 0004, Song Wang 0011, Maosheng Zhang, Furong Lei, Shiqing Li
ICME4
2014 The Perceptual Characteristics of 3D Orientation
Ruimin Hu, Weiping Tu, Xiaochen Wang 0001
MMM (2)4
2013 An expanded Mid/Side coding for 3D audio signal compression
abstract
Three dimensional (3D) audio technologies are booming with the success of 3D video technology. The sharply increased audio channels make its huge data unacceptable for transmitting bandwidth and storage media. This paper investigates the conventional Mid/Side (M/S) coding method, and expands it to a Three-channel Dependent M/S coding (3D-M/S) method. 3D-M/S perform sum and difference coding based on three channels instead of conventional two channels, and corresponding transform matrixes are presented. Furthermore, a framework is proposed to enable 3D-M/S compress any number of audio channels. Experiment shows proposed method obtains 25.4% objective quality improvement comparing with independent channel coding, and only increases 11.3% complexity comparing with the 29.9% of PCA method.
Shi Dong 0004, Ruimin Hu, Xiaochen Wang 0001, Weiping Tu
ICASSP4
2012 Enhanced Principal Component Using Polar Coordinate PCA for Stereo Audio Coding
abstract
High efficiency audio compression is the basic technology in audio involved multimedia application. Down mixing and parametric coding are efficient coding scheme with widely applications in some up to date audio codecs such as PS in EAAC+ and MPEG-Surround, and PCA stereo coding followed this idea to map two channels to one channel with maximum energy and parameterize the secondary channel. This paper investigates the conventional PCA method performance under general stereo model with multiple sound sources and different directions, and then proposes a Polar Coordinate based PCA (PC-PCA) stereo coding method. It has been proved that when multiple sound sources exist with different directions, proposed method is better than the conventional PCA method in certain conditions. A stereo codec based on PC-PCA has also been proposed to validate the performance improvement of proposed method.
Shi Dong 0004, Ruimin Hu, Weiping Tu, Junjun Jiang, Song Wang 0011
ICME3