EDBT 2026 Demo / reviewers in the wild / expert
Jie Zhang 0042
dblp:84/6889-42
· DBLP profile ↗
61ranked-venue papers
13as first author
51since 2021 · last 2025
0000-0003-1124-0854ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 3 first-author · 37 since 2021Artificial intelligence and machine learning · 33 · 10 first-author · 25 since 2021Systems, architecture and hardware · 3Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational AutoencoderabstractSinging Voice Synthesis (SVS) aims to generate singing voices of high fidelity and expressiveness. Conventional SVS systems usually utilize an acoustic model to transform a music score into acoustic features, followed by a vocoder to reconstruct the singing voice. It was recently shown that end-to-end modeling is effective in the fields of SVS and Text to Speech (TTS). In this work, we thus present a fully end-to-end SVS method together with a chunkwise streaming inference to address the latency issue for practical usages. Note that this is the first attempt to fully implement end-to-end streaming audio synthesis using latent representations in VAE. We have made specific improvements to enhance the performance of streaming SVS using latent representations. Experimental results demonstrate that the proposed method achieves synthesized audio with high expressiveness and pitch accuracy in both streaming SVS and TTS tasks. Jianwei Cui 0003, Shihao Chen, Jie Zhang 0042, Li-Rong Dai 0001 |
AAAI | 4 |
| 2025 | Sinba: Singing-To-Accompaniment Generation With Pitch Guidance Via Mamba-Based Language ModelabstractIn this paper, we propose Sinba, a system that can directly generate corresponding background accompaniment music from vocal input, allowing users to create complete songs using only sung vocals. Sinba adopts a decoder-only backbone network architecture. We utilize the Mamba model, which is a linear-time sequence modeling method with selective state spaces and has been proven to achieve more advanced performance than Transformers as a foundation model in long-sequence modeling tasks. However, the Mamba was initially applied to audio tasks by pre-training directly on raw audio waveform samples as the backbone model. In this paper, we convert both the training targets and inputs into discretized tokens for direct training. We also extract pitch information from the vocal input as an additional feature for the model. The proposed model is trained using source-separated data pairs. Subjective and objective experimental results demonstrate that the proposed model can generate high-quality accompaniment that matches the style and rhythm of the vocal input, outperforming the Transformerbased baseline. Synthesized audio samples are available at: https://sounddemos.github.io/sinba. Jianwei Cui 0003, Shihao Chen, Jie Zhang 0042, Chengxing Li, Shan Yang 0001, Li-Rong Dai 0001 |
ASRU | 4 |
| 2025 | Leveraging Boolean Directivity Embedding for Binaural Target Speaker ExtractionabstractDirection-based target speaker extraction (TSE) attracts a constant attention due to the convenience of direction acquisition over assistive video or enrollment audio. The direction clue heavily affects the TSE performance, which might be more seriously in the case of binaural setups due to the small-sized and irregular microphone array. In this paper, we propose Boolean Directivity Embedding (BDE) as a new direction feature in order to precisely lock onto the target speaker independent on microphone array configurations for binaural TSE (BiTSE). We design an encoder that accurately aligns the BDE with the mixed audio signals for feature fusion. Considering that the Boolean representation may contain insufficient spatial and temporal information, we enhance the BDE by incorporating the previously-proposed spatiotemporal features, showing a compatibility and stronger capacity for BiTSE. The proposed BiTSE model, which is based on the narrow-band Conformer as the backbone, can adapt to the cases of target speaker switching and moving by simply modifying the frame-wise BDE. Experimental results demonstrate the efficacy of our method in both stationary and dynamic scenarios. The proposed method is open-sourced in https://github.com/ichi131/Direction-based-BiTSE. Yichi Wang 0001, Jie Zhang 0042, Chengqian Jiang, Weitai Zhang, Zhongyi Ye, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2025 | Learning-Based Utility Estimation with Application to Speech Enhancement of a Moving SpeakerabstractWireless acoustic sensor network (WASN) has become a useful platform for monitoring acoustic scenes and sound acquisition. It is likely that many acoustic devices have a marginal impact on performance, which facilitates a necessity of optimizing the tradeoff between performance and computational load by microphone subset selection (MSS). It was shown that microphone utility can measure the node-specific importance on sound acquisition, which can potentially guide downstream speech processing. However, existing utility estimation methods were primarily designed in cases with a single static speaker. In this paper, we propose a learning-based utility estimation model for moving sources, allowing for the time-varying informative MSS by simply choosing microphones with larger utilities. The proposed model takes the speech representation extracted by the pre-trained wav2vec2.0 and the DFT magnitudes as raw features and employs an encoder for feature compression. The microphone utility is estimated by a remote decoder using the received feature vectors. Experimental results show that the proposed method can obtain a more accurate utility estimate in terms of Pearson correlation coefficient (PCC) and ordering the estimated utilities turns out a more informative microphone subset for the enhancement of a moving speaker. Jie Zhang 0042, Chengqian Jiang, Yichi Wang 0001, Haoyin Yan |
ICASSP | 1 |
| 2025 | Improved Feature Extraction Network for Neuro-Oriented Target Speaker ExtractionabstractThe recent rapid development of auditory attention decoding (AAD) offers the possibility of using electroencephalography (EEG) as auxiliary information for target speaker extraction. However, effectively modeling long sequences of speech and resolving the identity of the target speaker from EEG signals remains a major challenge. In this paper, an improved feature extraction network (IFENet) is proposed for neuro-oriented target speaker extraction, which mainly consists of a speech encoder with dual-path Mamba and an EEG encoder with Kolmogorov-Arnold Networks (KAN). We propose SpeechBiMamba, which makes use of dual-path Mamba in modeling local and global speech sequences to extract speech features. In addition, we propose EEGKAN to effectively extract EEG features that are closely related to the auditory stimuli and locate the target speaker through the subject’s attention information. Experiments on the KUL and AVED datasets show that IFENet outperforms the state-of-the-art model, achieving 36% and 29% relative improvements in terms of scale-invariant signal-to-distortion ratio (SI-SDR) under an open evaluation condition. Cunhang Fan, Youdian Gao, Zexu Pan, Jie Zhang 0042, Zhao Lv |
ICASSP | 6 |
| 2025 | Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker RepresentationsabstractIn this paper, we propose a noise-invariant speaker representation learning (SRL) approach by aligning noisy-clean speech pairs at both the feature and embedding levels for model training. Specifically, we first construct noisy-clean pairs using data augmentation during training. The noisy features are then processed by a Conformer-based enhancement module. The feature-level alignment is achieved by minimizing the mean squared error between the enhanced and original clean data. At the embedding level, we introduce a supervised contrastive learning loss with noise-adaptive margin to simultaneously enhance the intra-speaker compactness and the inter-speaker separability and better adapt different noise levels, in combination with the Barlow Twins self-supervised loss to align the noisy-clean data pairs and reduce noise redundancy in the embedding space. Finally, these loss components are integrated with conventional classification loss to train the SRL network. Experimental results on various VoxCeleb1 test sets synthesized with noise sources demonstrate the effectiveness of the proposed method. Zuoliang Li, Yang Ai, Jie Zhang 0042, Shengyu Peng, Bin Gu 0004, Wu Guo |
ICASSP | 3 |
| 2025 | A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker VerificationabstractIn this paper, a multi-scale feature fusion paradigm is proposed to fully exploit the power of the pre-trained models for text-independent speaker verification. It contains a front-end feature extractor and an enhanced ECAPA-TDNN backend in a cascade manner. The feature extractor incorporates local representations of the CNN layers as well as the global clues of the Transformer layers of the pre-trained models, which are combined to construct the multi-scale discriminative features. The outputs of the feature extractor are then fed into the back-end model (tailored from ECAPA-TDNN) to obtain the final speaker embedding. Results on VoxCeleb datasets validate the superiority of the proposed method with equal error rates of 0.633% and 0.457% on the official trials of Vox1-O using the base and large pre-trained models, respectively. Shengyu Peng, Wu Guo, Jie Zhang 0042, Zuoliang Li, Bin Gu 0004, Yang Ai |
ICASSP | 3 |
| 2025 | LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech EnhancementabstractSpeech enhancement (SE) aims to extract the clean waveform from noise-contaminated measurements to improve the speech quality and intelligibility. Although learning-based methods can perform much better than traditional counterparts, the large computational complexity and model size heavily limit the deployment on latency-sensitive and low-resource edge devices. In this work, we propose a lightweight SE network (LiSenNet) for real-time applications. We design sub-band downsampling and upsampling blocks and a dual-path recurrent module to capture band-aware features and time-frequency patterns, respectively. A noise detector is developed to detect noisy regions in order to perform SE adaptively and save computational costs. Compared to recent higher-resource-dependent baseline models, the proposed LiSenNet can achieve a competitive performance with only 37k parameters (half of the state-of-the-art model) and 56M multiply-accumulate (MAC) operations per second. Haoyin Yan, Jie Zhang 0042, Cunhang Fan, Yeping Zhou, Peiqi Liu |
ICASSP | 2 |
| 2025 | Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech EnhancementabstractBrain-assisted speech enhancement (BASE) aims to extract the target speaker in complex multi-talker scenarios using electroencephalogram (EEG) signals as an assistive modality, as the auditory attention of the listener can be decoded from electroneurographic signals of the brain. This facilitates a potential integration of EEG electrodes with listening devices to improve the speech intelligibility of hearing-impaired listeners, which was shown by the recently-proposed BASEN model. As in general the multichannel EEG signals are highly correlated and some are even irrelevant to listening, blindly incorporating all EEG channels would lead to a high economic and computational cost. In this work, we therefore propose a geometry-constrained EEG channel selection approach for BASE. We design a new weighted multi-dilation temporal convolutional network (WD-TCN) as the backbone to replace the Conv-TasNet in BASEN. Given a raw channel set that is defined by the electrode geometry for feasible integration, we then propose a geometry-constrained convolutional regularization selection (GC-ConvRS) module for WD-TCN to find an informative EEG subset. Experimental results on a public dataset show the superiority of the proposed WD-TCN over BASEN. The GC-ConvRS can further refine the useful EEG subset subject to the geometry constraint, resulting in a better trade-off between performance and integration cost. Keying Zuo, Qingtian Xu, Jie Zhang 0042, Zhen-Hua Ling |
ICASSP | 3 |
| 2025 | Parameter-Efficient Fine-tuning with Instance-Aware Prompt and Parallel Adapters for Speaker Verification
Shengyu Peng, Wu Guo, Jie Zhang 0042, Lipeng Dai, Zuoliang Li |
INTERSPEECH | 3 |
| 2025 | DOA or Speaker Embedding: Which is Better for Multi-Microphone Target Speaker ExtractionabstractTarget speaker extraction (TSE) is a useful front-end to improve the speech quality and intelligibility for speech applications, whereas direction-of-arrival (DOA) and speaker embedding are two of the most often-used assistive clues to identify the target speaker in audio-only multi-microphone systems. Both can significantly improve the TSE performance compared to blind TSE models, which however have not yet been comprehensively compared in literature. In order to show their pros and cons, in this work we therefore build a unified framework for a fair comparison that allows for both DOA and speaker embedding as the assistive clue. The DOA is used to calculate multichannel spatiotemporal speech features and a speaker encoder is designed to extract the speaker embedding, either of which is then fused with the noisy speech features for TSE. We can then evaluate their respective strengths in diverse acoustic conditions, e.g., varying noise level, microphone number, speaker location. Results show that given true DOA angles, the DOA-based TSE model always outperforms the speaker embedding based counterpart regardless of noise/microphone/location conditions, meaning the stronger discriminativity of DOA in terms of speaker identity. This superiority becomes smaller if the DOA mis-match increases, and the latter can do better in the large DOA mismatch case. Jie Zhang 0042, Yichi Wang 0001, Haoyin Yan |
IEEE Signal Process. Lett. | 2 |
| 2024 | Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech RepresentationabstractSelf-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the scarcity of labeled multichannel data and complex ambient noises. The efficacy of self-supervised learning for far-field multichannel and multi-modal speech processing has not been well explored. Considering that visual information helps to improve speech recognition performance in noisy scenes, in this work we propose the multichannel multi-modal speech self-supervised learning framework AV-wav2vec2, which utilizes video and multichannel audio data as inputs. First, we propose a multi-path structure to process multi-channel audio streams and a visual stream in parallel, with intra-, and inter-channel contrastive as training targets to fully exploit the rich information in multi-channel speech data. Second, based on contrastive learning, we use additional single-channel audio data, which is trained jointly to improve the performance of multichannel multi-modal representation. Finally, we use a Chinese multichannel multi-modal dataset in real scenarios to validate the effectiveness of the proposed method on audio-visual speech recognition (AVSR), automatic speech recognition (ASR), visual speech recognition (VSR) and audio-visual speaker diarization (AVSD) tasks. Qiushi Zhu, Jie Zhang 0042, Li-Rong Dai 0001 |
AAAI | 2 |
| 2024 | Adversarial Speech for Voice Privacy Protection from Personalized Speech GenerationabstractThe rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human listeners, resulting in an urgent demand in protecting speakers' voices from malicious misuse. In this regard, we propose a speaker protection method based on adversarial attacks. The proposed method perturbs speech signals by minimally altering the original speech while rendering downstream speech generation models unable to accurately generate the voice of the target speaker. For validation, we employ the open-source pre-trained YourTTS model for speech generation and protect the target speaker's speech in the white-box scenario. Automatic speaker verification (ASV) evaluations were carried out on the generated speech as the assessment of the voice protection capability. Our experimental results show that we successfully perturbed the speaker encoder of the YourTTS model using the gradient-based I-FGSM adversarial perturbation method. Furthermore, the adversarial perturbation is effective in preventing the YourTTS model from generating the speech of the target speaker. Audio samples can be found in https://voiceprivacy.github.io/Adeversarial-Speech-with-YourTTS. Shihao Chen, Jie Zhang 0042, Kong-Aik Lee, Zhen-Hua Ling, Li-Rong Dai 0001 |
ICASSP | 3 |
| 2024 | Sifisinger: A High-Fidelity End-to-End Singing Voice Synthesizer Based on Source-Filter ModelabstractThis paper presents an advanced end-to-end singing voice synthesis (SVS) system based on the source-filter mechanism that directly translates lyrical and melodic cues into expressive and high-fidelity human-like singing. Similarly to VISinger 2, the proposed system also utilizes training paradigms evolved from VITS and incorporates elements like the fundamental pitch (F0) predictor and waveform generation decoder. To address the issue that the coupling of mel-spectrogram features with F0 information may introduce errors during F0 prediction, we consider two strategies. Firstly, we leverage mel-cepstrum (mcep) features to decouple the intertwined mel-spectrogram and F0 characteristics. Secondly, inspired by the neural source-filter models, we introduce source excitation signals as the representation of F0 in the SVS system, aiming to capture pitch nuances more accurately. Meanwhile, differentiable mcep and F0 losses are employed as the waveform decoder supervision to fortify the prediction accuracy of speech envelope and pitch in the generated speech. Experiments on the Opencpop dataset demonstrate efficacy of the proposed model in synthesis quality and intonation accuracy. Synthesized audio samples are available at: https://sounddemos.github.io/sifisinger. Jianwei Cui 0003, Chao Weng, Jie Zhang 0042, Li-Rong Dai 0001 |
ICASSP | 4 |
| 2024 | A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker ExtractionabstractTarget speaker extraction (TSE) based on direction of arrival (DOA) has a wide range of applications in e.g., remote conferencing, hearing aids, in-car speech interaction. Due to the inherent phase uncertainty, existing TSE methods usually suffer from speaker confusion within specific frequency bands. Imprecise DOA measurements caused by e.g., the calibration of the microphone array and ambient noises, can also deteriorate the TSE performance. In order to improve the robustness of TSE, in this work we propose several new multichannel spatiotemporal features to represent the discriminability of the target speaker. The narrow-band Conformer model is applied in combination with the proposed features to facilitate the extraction of the target speaker. In addition, we consider knowledge distillation for improving the model robustness, particularly in the presence of DOA mis-match. Experimental results on a public dataset verify the efficacy of the proposed method. Yichi Wang 0001, Jie Zhang 0042, Shihao Chen, Weitai Zhang, Zhongyi Ye, Xinyuan Zhou, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2024 | An End-to-End EEG Channel Selection Method with Residual Gumbel Softmax for Brain-Assisted Speech EnhancementabstractBrain-assisted speech enhancement (SE) has gained an increasing attention recently, as electroencephalogram (EEG) measurements somehow reflect auditory attention clues. The design of an EEG cap with sparse channel distributions can save the hardware cost, setup time as well as algorithmic complexity, which can be done by EEG channel selection, as it was shown that the multichannel EEG signals are highly correlated and redundant. In this paper, we thus propose an end-to-end EEG channel selection method based on a weighted residual structure, called Residual Gumbel Selection (ResGS), for the neuro-steered SE task. The use of residual connections can lead to a more efficient and stable training procedure. The proposed ResGS consists of the weighted residual training and fine-tuning steps. Experimental results on a public dataset validate the efficacy of the proposed method in channel selection and show that a small subset of channels is enough to achieve a near-optimal performance. Qing-Tian Xu, Jie Zhang 0042, Zhen-Hua Ling |
ICASSP | 2 |
| 2024 | LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
Shihao Chen, Jie Zhang 0042, Rilin Chen, Li-Rong Dai 0001 |
INTERSPEECH | 3 |
| 2024 | Refining Self-supervised Learnt Speech Representation using Brain Activations
Kangdi Mei, Zhaoci Liu, Yang Ai, Jie Zhang 0042, Zhen-Hua Ling |
INTERSPEECH | 6 |
| 2024 | FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial LandmarksabstractDepression is a prevalent mental health disorder that significantly impacts individuals' lives and well-being. Early detection and intervention are crucial for effective treatment and management of depression. Recently, there are many end-to-end deep learning methods leveraging the facial expression features for automatic depression detection. However, most current methods overlook the temporal dynamics of facial expressions. Although very recent 3DCNN methods remedy this gap, they introduce more computational cost due to the selection of CNN-based backbones and redundant facial features. To address the above limitations, by considering the timing correlation of facial expressions, we propose a novel framework called FacialPulse, which recognizes depression with high accuracy and speed. By harnessing the bidirectional nature and proficiently addressing long-term dependencies, the Facial Motion Modeling Module (FMMM) is designed in FacialPulse to fully capture temporal features. Since the proposed FMMM has parallel processing capabilities and has the gate mechanism to mitigate gradient vanishing, this module can also significantly boost the training speed. Besides, to effectively use facial landmarks to replace original images to decrease information redundancy, a Facial Landmark Calibration Module (FLCM) is designed to eliminate facial landmark errors to further improve recognition accuracy. Extensive experiments on the AVEC2014 dataset and MMDA dataset (a depression dataset) demonstrate the superiority of FacialPulse on recognition accuracy and speed, with the average MAE (Mean Absolute Error) decreased by 21% compared to baselines, and the recognition speed increased by 100% compared to state-of-the-art methods. Codes are released at https://github.com/volatileee/FacialPulse. Jinyang Huang, Jie Zhang 0042, Xin Liu 0104, Xiang Zhang 0011, Zhi Liu 0002, Peng Zhao 0024, Sigui Chen, Xiao Sun 0003 |
ACM Multimedia | 3 |
| 2024 | VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation LearningabstractAlthough speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework to integrate different modal information and leverage different resources (e.g., visual-audio pairs, audio-text pairs, unlabeled speech, and unlabeled text) to facilitate speech representation learning was not well explored. In this paper, we propose a unified cross-modal representation learning frameworkVatLM(Visual-Audio-Text Language Model). The proposedVatLMemploys a unified backbone network to model the modality-independent information and utilizes three simple modality-dependent modules to preprocess visual, speech, and text inputs. In order to integrate these three modalities into one shared semantic space,VatLMis optimized with a masked prediction task of unified tokens, given by our proposed unified tokenizer. We evaluate the pre-trainedVatLMon audio-visual related downstream tasks, including audio-visual speech recognition (AVSR), and visual speech recognition (VSR) tasks. Results show that the proposedVatLMoutperforms previous state-of-the-art models, such as the audio-visual pre-trained AV-HuBERT model, and analysis also demonstrates thatVatLMis capable of aligning different modalities into the same space. To facilitate future research, we release the code and pre-trained models athttps://aka.ms/vatlm. Qiushi Zhu, Shujie Liu 0001, Binxing Jiao, Jie Zhang 0042, Li-Rong Dai 0001, Daxin Jiang, Jinyu Li 0001, Furu Wei |
IEEE Trans. Multim. | 6 |
| 2023 | A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech EnhancementabstractAudio-visual speech enhancement (AVSE) was shown to be superior over conventional audio-only counterpart for improving the speech quality. However, most existing AVSE models are heavyweight in the sense of parameter count, which is inappropriate for the deployment and practical applications. In this paper, we therefore present a lightweight AVSE approach (called M3Net) by incorporating several multi-modality, multi-scale and multi-branch strategies. Three multi-scale techniques are designed for the visual and audio streams, including multi-scale average pooling (MSAP), multi-scale ResNet (MSResNet) and multi-scale short time Fourier transform (MSSTFT). It is shown that each multi-scale module positively contributes to the performance. Also, we consider four skip connections for the audio-visual feature aggregation, which have a great complementary effect on the designed multi-scale techniques. Experimental results show that these techniques are flexible in combination with existing approaches, and more importantly obtain a comparable performance with a smaller model size compared to the heavyweight networks. Liangfa Wei, Jie Zhang 0042, Jianming Yang, Yannan Wang, Tian Gao 0005, Li-Rong Dai 0001 |
ICASSP | 3 |
| 2023 | The NERCSLIP-USTC System for the L3DAS23 Challenge Task2: 3D Sound Event Localization and Detection (SELD)abstractSound event localization and detection (SELD) aims at identifying the temporal activities of a known set of sound event classes and estimating their locations. It remains challenging especially when there are overlapped acoustic events. In this work, a robust network architecture with data augmentation techniques is proposed to improve SELD performance, where ResNet and Conformer blocks are combined to model both local and global patterns. To address the data sparsity issue in SELD, SpecAugment, mixup and audio channel swapping (ACS) techniques are adopted. Our proposed system is evaluated in the Task2 of the L3DAS23 challenge and ranks the second place, achieving significant improvements over the baseline. Haoyin Yan, Qing Wang 0008, Jie Zhang 0042 |
ICASSP | 4 |
| 2023 | Robust Data2VEC: Noise-Robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive LearningabstractSelf-supervised pre-training methods based on contrastive learning or regression tasks can utilize more unlabeled data to improve the performance of automatic speech recognition (ASR). However, the robustness impact of combining the two pre-training tasks and constructing different negative samples for contrastive learning still remains unclear. In this paper, we propose a noise-robust data2vec for self-supervised speech representation learning by jointly optimizing the contrastive learning and regression tasks in the pre-training stage. Furthermore, we present two improved methods to facilitate contrastive learning. More specifically, we first propose to construct patch-based non-semantic negative samples to boost the noise robustness of the pre-training model, which is achieved by dividing the features into patches at different sizes (i.e., so-called negative samples). Second, by analyzing the distribution of positive and negative samples, we propose to remove the easily distinguishable negative samples to improve the discriminative capacity for pre-training models. Experimental results on the CHiME-4 dataset show that our method is able to improve the performance of the pre-trained model in noisy scenarios. We find that joint training of the contrastive learning and regression tasks can avoid the model collapse to some extent compared to only training the regression task. Qiushi Zhu, Jie Zhang 0042, Shujie Liu 0001, Yu-Chen Hu, Li-Rong Dai 0001 |
ICASSP | 3 |
| 2023 | BASEN: Time-Domain Brain-Assisted Speech Enhancement Network with Convolutional Cross Attention in Multi-talker Conditions
Jie Zhang 0042, Qing-Tian Xu, Qiushi Zhu, Zhen-Hua Ling |
INTERSPEECH | 1 |
| 2023 | CASA-ASR: Context-Aware Speaker-Attributed ASR
Mohan Shi, Zhihao Du, Qian Chen 0003, Fan Yu 0002, Yangze Li, Shiliang Zhang, Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 7 |
| 2023 | Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction
Mohan Shi, Yuchun Shu, Lingyun Zuo, Qian Chen 0003, Shiliang Zhang, Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 6 |
| 2023 | Real-Time Causal Spectro-Temporal Voice Activity Detection Based on Convolutional Encoding and Residual Decoding
Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 2 |
| 2023 | Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023abstractIn this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three different structures based on attention-guided feature gathering (AFG) are designed for deep feature fusion. Then, we introduce a joint decoding structure for emotion classification and valence regression in the decoding stage. A multi-task loss based on uncertainty is also designed to optimize the whole process. Finally, by combining three different structures on the posterior probability level, we obtain the final predictions of discrete and dimensional emotions. When tested on the dataset of multimodal emotion recognition challenge (MER 2023), the proposed framework yields consistent improvements in both emotion classification and valence regression. Our final system achieves state-of-the-art performance and ranks third on the leaderboard on MER-MULTI sub-challenge. Yuxuan Xi, Hang Chen 0001, Jun Du 0002, Yan Song 0001, Qing Wang 0008, Hengshun Zhou, Jiefeng Ma, Pengfei Hu 0006, Ya Jiang, Shi Cheng 0001, Jie Zhang 0042, Yuzhe Weng |
ACM Multimedia | 13 |
| 2023 | A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech RecognitionabstractBoth unpaired speech and text have shown to be beneficial for low-resource automatic speech recognition (ASR), which, however were either separately used for pre-training, self-training and language model (LM) training, or jointly used for designing hybrid models in literature. In this work, we leverage both unpaired speech and text to train a general ASR model, which are used in the form of data pairs by generating the missing parts in prior to model training. We propose to train a model alternatively using the prepared speech-PseudoLabel and SynthesizedAudio-text pairs and reveal the complementary property in both acoustic and linguistic features. The proposed method is thus called complementary joint training (CJT). Based on the basic CJT, label masking for pseudo-labels and parallel layers for synthesized audio are then proposed for re-training to further cope with the deviations from real data, termed as CJT++. In addition, the proposed CJT is extended to the scenario with zero paired data by considering an iterative CJT for the training of seed ASR model. Experimental results on Libri-light show the efficacy of joint training as well as two second-round training strategies, and the superiority over recent models is validated, particularly in extreme low-resource cases. Ye-Qian Du, Jie Zhang 0042, Ming-Hui Wu, Zhouwang Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Memory Storable Network Based Feature Aggregation for Speaker Representation LearningabstractLearning fixed-dimensional speaker representation using deep neural networks is a key step in speaker verification. In this work, we propose an auxiliary memory storable network (MSN) to assist a backbone network for learning discriminative features, which are sequentially aggregated from lower to deeper layers of the backbone. The proposed MSN has a similar architecture to the ResNet and contains a set of cascaded feature aggregation (FA) blocks. Each FA block first aggregates the multi-level features from the previous block and the features from the corresponding backbone layer. The output features of each intermediate layer within the backbone are then refined by the multi-level features of the corresponding FA block through masking and biasing operations. Finally, the features from the last layers of both MSN and the backbone are concatenated to form more discriminative speaker representations. Experimental results on five public datasets show significant and consistent improvements over conventional approaches. The effectiveness of the proposed method is also validated using ablation studies, showing a robust generalization capacity in combination with different backbone networks. Bin Gu 0004, Wu Guo, Jie Zhang 0042 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | A Dynamic Convolution Framework for Session-Independent Speaker Embedding LearningabstractSpeaker verification (SV) has suffered from session variability in complex acoustic scenarios, and learning session independent speaker representations remains a challenging problem. To tackle this, we propose a dynamic convolution framework for SV in this article, which dynamically adapts the model parameters to each input feature during inference, such that the model can flexibly extract robust speaker characteristics under different acoustic conditions. Specifically, we combine an adaptive context vector extraction (ACVE) module and a sub-kernel scaling (SKS) module in a cascaded manner. The ACVE uses a moving weighted sum and a sub-band self-attention in parallel to extract global-local information, which is then passed through the subsequent SKS module for efficient dynamic kernel generation. The proposed method can be easily implemented on the backbone network by replacing the conventional static counterparts. Experimental results on five public SV datasets show significant and consistent improvements over comparison approaches, and ablation studies and visual analysis further demonstrate the effectiveness of the proposed method. Bin Gu 0004, Jie Zhang 0042, Wu Guo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Energy-Efficient Sparsity-Driven Speech Enhancement in Wireless Acoustic Sensor NetworksabstractWireless acoustic sensor network (WASN) has shown a superiority over conventional microphone arrays in many aspects. There exists an important tradeoff between the performance and power consumption, as usually the sensors are power driven with a limited amount of battery resource. Given a prescribed performance bound, in literature sensor selection (SS) and rate allocation (RA) methods can be leveraged to optimize the energy efficiency. In this work, we propose a joint rate allocation and sensor selection (RASS) approach to simultaneously optimize the sensor subset and rate distribution, which is formulated by minimizing the total transmission power in terms of selection and bit-rate variables and constraining the residual noise power. It can be shown that under a set of linear constraints on beamforming, the linearly-constrained minimum variance (LCMV) beamformer is the optimal noise reduction filter. Based on this, the RASS reduces to a mixed semi-definite and bilinear programming problem, which is then solved using a two-step algorithm. As the selection and bit-rate unknowns are bilinear, we first consider to optimize their product, resulting in an upper bound of RASS. Then, we use McCormick envelopes to relax the bilinear constraint, resulting in a linear program. The final selection and bit-rate solutions are obtained by posterior randomized rounding. It can be shown that SS and RA are special cases of the proposed RASS. Numerical results using simulated WASNs validate the power efficiency of the proposed method as well as the robustness against dynamic factors. Jie Zhang 0042, Jun Du 0002, Li-Rong Dai 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | SDW-SWF: Speech Distortion Weighted Single-Channel Wiener Filter for Noise ReductionabstractSpeech enhancement shows an important necessity in many audio applications, particularly in noisy environments, where the speech quality needs to be improved. In this work, we consider the single-channel noise reduction (NR) problem from the conventional signal processing perspective. As conventional single-channel NR filters suffer from a serious speech distortion (SD) problem, we propose an SD weighted single-channel Wiener filter (SDW-SWF) in the short-time Fourier transform domain, which is obtained by minimizing the mean-square error (MSE) of the clean speech plus a$\mu$-weighted residual noise variance. Based on the generalized eigenvalue decomposition (GEVD) and rank-$r$approximation of the speech correlation matrix, the SDW-SWF can be written as a linear combination of eigenpairs, from which some special cases reduce to existing single-channel NR filters. As such, the proposed SDW-SWF has two parameters (i.e.,$\mu$and$r$) to tradeoff the MSE and SD. Then we theoretically analyze the impacts of the tradeoff parameters on the NR performance in SD, residual noise variance and the output signal-to-noise ratio (SNR). In addition, it is shown that the STFT-domain SDW-SWF can be further extended to the time domain, where the derived theorems still hold. Numerical results from several perspectives validate the effectiveness of the proposed method. Jie Zhang 0042, Jun Du 0002, Li-Rong Dai 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | A Joint Speech Enhancement and Self-Supervised Representation Learning Framework for Noise-Robust Speech RecognitionabstractThough speech enhancement (SE) can be used to improve speech quality in noisy environments, it may also cause distortions that degrade the performance of automatic speech recognition (ASR) models. Self-supervised pre-training, on the other hand, has been shown to improve the noise robustness of ASR models. However, the potential of the (optimal) integration of SE and self-supervised pre-training still remains unclear. In this paper, we propose a novel self-supervised pre-training framework that incorporates SE to improve ASR performance in noisy environments. First, in the pre-training phase the original noisy waveform or the waveform obtained by SE is fed into the self-supervised model to learn the contextual representation, where the quantized clean speech acts as the target. Second, we propose a dual-attention fusion method to fuse the features of noisy and enhanced speech, which can compensate for the information loss caused by separately using individual modules. Due to the flexible exploitation of clean/noisy/enhanced branches, the proposed method turns out to be a generalization of some existing noise-robust ASR models, e.g., enhanced wav2vec2.0. Finally, experimental results on both synthetic and real noisy datasets show that the proposed joint training approach can improve the ASR performance under various noisy settings, leading to a stronger noise robustness. Qiushi Zhu, Jie Zhang 0042, Li-Rong Dai 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Adaptive Video Streaming With Automatic Quality-of-Experience OptimizationabstractVideo streaming has grown tremendously in recent years and it is now one of the main applications on the Internet. Due to the networks' inherent bandwidth fluctuations, various rate-adaptive streaming algorithms have been developed to compensate for such fluctuations to improve Quality-of-Experience (QoE). However, in practice, the preference for QoE typically differs significantly across different viewers and there is no systematic way so far to comprehensively incorporate different sets of conflicting QoE objectives into the algorithm design. Thus, it is not surprising that the QoE performance achieved by the existing algorithms is in fact far from optimal. This work aims at attacking the heart of the problem by developing a novel framework called Post Streaming Quality Analysis (PSQA) that can maximize the QoE under any preference through automatically tuning the adaptation logic of the streaming algorithms. Evaluation results show that the QoE achieved by PSQA is substantially better than the existing approaches and in some scenarios even close to optimal. Moreover, PSQA can be readily implemented into real streaming platforms, offering a practical and reliable solution for high-performance streaming services. Jie Zhang 0042, Yan Liu 0047, Haibo Hu 0001, Jack Y. B. Lee, Vaneet Aggarwal |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | DUASVS: A Mobile Data Saving Strategy in Short-Form Video StreamingabstractFueled by the emerging short video applications (e.g., TikTok), streaming short-form videos nowadays is ubiquitous among mobile users. During the viewing, one common action is to scroll the screen to switch videos, which is a handy operation for the viewers to quickly search for content of interest. However, our empirical measurements reveal that frequent video switching can result in nearly half of the mobile data quota being used for transferring the video data that is never watched. This problem is called data loss in this work. Given the immense cost of the network infrastructure, such a high proportion of data loss is financially tremendous to both mobile users and streaming vendors. To tackle the problem, this study proposes a novel system called Data Usage Aware Short Video Streaming (DUASVS), where a new Integrated Learning is used to capture the characters of past network conditions and then trains intelligent adaptation models to reduce data loss and save data usage. Extensive evaluations show that DUASVS is able to save 70.7%∼83.2% of mobile data usage without incurring any QoE degradation. Moreover, the system exhibits strong robustness, performing consistently over a wide range of network environments as well as video streaming sessions. Jie Zhang 0042, Ke Liu 0004, Jack Y. B. Lee, Haibo Hu 0001, Vaneet Aggarwal |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Reference Microphone Selection and Low-Rank Approximation Based Multichannel Wiener Filter with Application to Speech RecognitionabstractFor multichannel speech recognition systems, it is necessary to use a speech enhancement module to suppress ambient noises. Given second-order statistics, the multichannel Wiener filter (MWF) can be designed for noise reduction. It was shown that the MWF noise reduction performance depends on the selection of reference microphone and the rank of the speech correlation matrix. It is questionable how the reference microphone and rank would affect the subsequent recognition accuracy. In this paper, we present an experimental study on the low-rank approximation and reference microphone selection based MWF with application to noisy speech recognition. Further, we propose to maximize the input signal-to-noise ratio (SNR) for reference selection in the sense of signal quality. Experimental results show that the output SNR of rank-1 MWF is independent of the reference, while the speech intelligibility is always related to both the rank and reference microphone. The word error rate is positively affected by the rank, and the proposed reference selection method can improve the performance in terms of both speech intelligibility and speech recognition. Xing-Yu Chen, Jie Zhang 0042, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2022 | Supervised and Self-Supervised Pretraining Based Covid-19 Detection Using Acoustic Breathing/Cough/Speech SignalsabstractA rapid-accurate detection method for COVID-19 is rather important for avoiding its pandemic. In this work, we propose a bi-directional long short-term memory (BiLSTM) network based COVID-19 detection method using breath/speech/cough signals. Three kinds of acoustic signals are taken to train the network and individual models for three tasks are built, respectively, whose parameters are averaged to obtain an average model, which is then used as the initialization for the BiLSTM model training of each task. It is shown that such an initialization method can significantly improve the detection performance on three tasks. This is called supervised pre-training based detection. Besides, we utilize an existing pre-trained wav2vec2.0 model and pre-train it using the DiCOVA dataset, which is utilized to extract a high-level representation as the model input to replace conventional mel-frequency cepstral coefficients (MFCC) features. This is called self-supervised pre-training based detection. To reduce the information redundancy contained in the recorded sounds, silent segment removal, amplitude normalization and time-frequency masking are also considered. The proposed detection model is evaluated on the DiCOVA dataset and results show that our method achieves an area under curve (AUC) score of 88.44% on blind test in the fusion track. It is shown that using high-level features together with MFCC features is helpful for diagnosing accuracy. Xing-Yu Chen, Qiushi Zhu, Jie Zhang 0042, Li-Rong Dai 0001 |
ICASSP | 3 |
| 2022 | A Noise-Robust Self-Supervised Pre-Training Model Based Speech Representation Learning for Automatic Speech RecognitionabstractWav2vec2.0 is a popular self-supervised pre-training framework for learning speech representations in the context of automatic speech recognition (ASR). It was shown that wav2vec2.0 has a good robustness against the domain shift, while the noise robustness is still unclear. In this work, we therefore first analyze the noise robustness of wav2vec2.0 via experiments. We observe that wav2vec2.0 pre-trained on noisy data can obtain good representations and thus improve the ASR performance on the noisy test set, which however brings a performance degradation on the clean test set. To avoid this issue, in this work we propose an enhanced wav2vec2.0 model. Specifically, the noisy speech and the corresponding clean version are fed into the same feature encoder, where the clean speech provides training targets for the model. Experimental results reveal that the proposed method can not only improve the ASR performance on the noisy test set which surpasses the original wav2vec2.0, but also ensure a tiny performance decrease on the clean test set. In addition, the effectiveness of the proposed method is demonstrated under different types of noise conditions. Qiushi Zhu, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2022 | Learning Contextually Fused Audio-Visual Representations For Audio-Visual Speech RecognitionabstractWith the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR) performance, as the multi-modal inputs contain more fruitful information in principle. In this paper, based on existing self-supervised representation learning methods for audio modality, we therefore propose an audio-visual representation learning approach. The proposed approach explores both the complementarity of audio-visual modalities and long-term context dependency using a transformer-based fusion module and a flexible masking strategy. After pre-training, the model is able to extract fused representations required by AVSR. Without loss of generality, it can be applied to single-modal tasks, e.g., audio/visual speech recognition by simply masking out one modality in the fusion module. The proposed pre-trained model is evaluated on speech recognition and lipreading tasks using one or two modalities, where the superiority is revealed. Jie Zhang 0042, Jianshu Zhang 0001, Ming-Hui Wu, Li-Rong Dai 0001 |
ICIP | 2 |
| 2022 | An Experimental Comparison between Low-Resource Semi-Supervised and High-Resource Supervised Automatic Speech Recognition ModelsabstractAutomatic speech recognition (ASR) is an important module in many multimedia applications. Recently, semi-supervised ASR has attracted increasing attention, which can be classified into self-supervised learning and self-training. It was shown that the combination of the two methods is beneficial for ASR on open datasets, e.g., LibriSpeech. However, it is still not completely clear how the semi-supervised model behaves on more challenging industrial datasets compared with supervised approaches using high-resource labeled data. In this paper, we therefore present an experimental study on the combination of self-supervised learning (e.g., wav2vec 2.0) and self-training (e.g., noisy student) in a low-resource industrial setting. The in-domain pre-trained and fine-tuned wav2vec 2.0 model is utilized to teach a baseline VGG-Transformer by pseudo-labeling in an industrial setting. Results reveal that without extra language models the combined semi-supervised acoustic model using much less labeled data can perform as well as the supervised counterpart, which would be rather beneficial for the application of semi-supervised ASR models in low-resource scenarios. Ao-Ran Gan, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001 |
ICME | 2 |
| 2022 | A Complementary Joint Training Approach Using Unpaired Speech and Text A Complementary Joint Training Approach Using Unpaired Speech and Text
Ye-Qian Du, Jie Zhang 0042, Qiushi Zhu, Li-Rong Dai 0001, Ming-Hui Wu, Zhouwang Yang |
INTERSPEECH | 2 |
| 2022 | Differential Time-frequency Log-mel Spectrogram Features for Vision Transformer Based Infant Cry Recognition
Hai-tao Xu, Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 2 |
| 2022 | External Text Based Data Augmentation for Low-Resource Speech Recognition in the Constrained Condition of OpenASR21 Challenge
Guolong Zhong, Hongyu Song, Ruoyu Wang 0029, Lei Sun 0010, Diyuan Liu, Jun Du 0002, Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 9 |
| 2022 | A Parametric Unconstrained Beamformer Based Binaural Noise Reduction for Assistive HearingabstractFor hearing-impaired listeners, it is required not only to enhance the target speech by suppressing ambient noises, but also to preserve the binaural cues of important directional sources, such that a complete spatial awareness of the acoustic scene is obtained. It was shown that the binaural minimum variance distortionless response (BMVDR) filter enables a joint enhancement and binaural cue preservation of the target source. By adding more linear constraints associated with the acoustic transfer functions of noise sources to BMVDR, the binaural linearly constrained minimum variance (BLCMV) filter is obtained, which can further preserve the spatial cues of interfering sources at the cost of sacrificing the noise reduction performance. Both BMVDR and BLCMV beamformers enforce ahard decisionon the spatial awareness. In this paper, we therefore propose a generalization for BMVDR and BLCMV, namely parametric unconstrained binaural (PUB) beamformer, which can achieve a controllablesoft decisionon noise reduction and binaural cues preservation. The involved parameters are chosen to control the distortion level of the target source and the preservation of interfering sources, respectively. Theoretical analysis on the performance comparison shows that the conventional BMVDR and BLCMV beamformers are special cases of the PUB beamformer. Experiments using both synthetic data and real audio measurements validate the superiority of the proposed method. Jie Zhang 0042 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Frequency-Invariant Sensor Selection for MVDR Beamforming in Wireless Acoustic Sensor NetworksabstractWireless acoustic sensor network (WASN) has a wide range of applications in internet of things, where signal estimation is one of the network design objectives. Due to the existence of ambient noises, the recorded audio signals are inevitably corrupted, resulting in a low signal-to-noise ratio (SNR), which triggers the necessity of signal enhancement. As using all sensor measurements brings a large amount of data transmissions and computational cost, the narrowband sensor selection was proposed to choose an informative subset of sensors to perform noise reduction in the audio context. However, the resulting frequency-dependent selection status has to be switched across frequencies. In order to avoid the complicated switching operations, we consider frequency-invariant sensor selection in this work. We propose to minimize the total power consumption over the WASN by constraining the broadband SNR, which can be solved using broadband semi-definite optimization (BroadOpt) or narrowband voting (NaVo) approaches. In order to further reduce the time complexity, we propose two near-optimal greedy methods, including gradient removal (GradR) and weighted input SNR removal (SnrR). As comparison, we also show a broadband energy removal (EnergyR) method. The greedy methods remove one sensor at each iteration from the complete network until the performance constraint is not satisfied. Numerical results using a simulated large-scale WASN show that the greedy methods can achieve a comparable performance compared to the optimization based counterparts, while the corresponding time complexity is much lower. In general, the sensors around the target source and the fusion center are more likely to be selected. Jie Zhang 0042, Li-Rong Dai 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2021 | An Improved Wav2Vec 2.0 Pre-Training Approach Using Enhanced Local Dependency Modeling for Speech Recognition
Qiushi Zhu, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001 |
Interspeech | 2 |
| 2021 | Multi-Granularity Sequence Alignment Mapping for Encoder-Decoder Based End-to-End ASRabstractEncoder-decoder based automatic speech recognition (ASR) methods are increasingly popular due to their simplified processing stages and low reliance on prior knowledge. Conventional encoder-decoder based approaches usually learn a sequence-to-sequence mapping function from the source speech to target units (e.g., subwords, characters) in an end-to-end manner. However, it is still unclear how to choose the optimal target unit, or granularity of multiple units. In general, as increasing the information available for learning sequence-to-sequence mapping functions can improve modeling effectiveness, we therefore propose a multi-granularity sequence alignment (MGSA) approach. This aims to enhance cross-sequence interactions between different granularity units for both modeling and inference stages in the encoder-decoder based ASR. Specifically, a decoder module is designed to generate multi-granularity sequence predictions. We then exploit the latent alignment mapping among units having different levels of granularity, by utilizing the decoded multi-level sequences as input for model prediction. The cross-sequence interaction can also be employed to re-calibrate output probabilities in the proposed post-inference algorithm. Experimental results on both WSJ-80 hrs and Switchboard-300 hrs datasets show the superiority of the proposed method compared to traditional multi-task methods as well as to single granularity baseline systems. Jie Zhang 0042, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | A Study on Reference Microphone Selection for Multi-Microphone Speech EnhancementabstractMulti-microphone speech enhancement methods typically require a reference position with respect to which the target signal is estimated. Often, this reference position is arbitrarily chosen as one of the reference microphones. However, it has been shown that the choice of the reference microphone can have a significant impact on the final noise reduction performance. In this paper, we therefore theoretically analyze the impact of selecting a reference on the noise reduction performance with near-end noise being taken into account. Following the generalized eigenvalue decomposition (GEVD) based optimal variable span filtering framework, we find that for any linear beamformer, the output signal-to-noise ratio (SNR) taking both the near-end and far-end noise into account is reference dependent. Only when the near-end noise is neglected, the output SNR of rank-1 beamformers does not depend on the reference position. However, in general for rank-r beamformers with r > 1 (e.g., the multichannel Wiener filter) the performance does depend on the reference position. Based on these, we propose an optimal algorithm for microphone reference selection that maximizes the output SNR. In addition, we propose a lower-complexity algorithm that is still optimal for rank-1 beamformers, but sub-optimal for the general r > 1 rank beamformers. Experiments using a simulated microphone array validate the effectiveness of both proposed methods and show that in terms of quality, several dB can be gained by selecting the proper reference microphone. Jie Zhang 0042, Li-Rong Dai 0001, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Sensor Selection for Relative Acoustic Transfer Function Steered Linearly-Constrained BeamformersabstractFor multi-microphone speech enhancement, different microphones might have different contributions, assome are even marginal. This is more likely to happen in wireless acoustic sensor networks (WASNs), where somesensors might be distant. In this work, we therefore consider sensor selection for linearly-constrained beamformers. Theproposed sensor selection approach is formulated by minimizing the total output noise power and constraining thenumber of selected sensors. As the considered sensor selection problem requires the relative acoustic transfer function(RTF), the covariance whitening based RTF estimation or a direct-path RTF approximation is exploited. For a singletarget source, we can thus substitute the estimated RTF or the assumed RTF to the original problem formulation in orderto design a minimum variance distortionless response (MVDR) beamformer. Alternatively, we can integrate the two RTFsto design a linearly constrained minimum variance (LCMV) beamformer in order to alleviate the effects of RTFestimation/approximation errors. By leveraging the superiority of LCMV beamformers, the proposed approach can beapplied to the multi-source case. An evaluation using a simulated large-scale WASN demonstrates that the integration ofRTFs for the sensor selection based LCMV beamformer can be beneficial as opposed to relying on either of theindividual RTF steered sensor selection based MVDR beamformers. We conclude that the sensors that are close to thetarget source(s) and also some around the coherent interferers are more informative. Jie Zhang 0042, Jun Du 0002, Li-Rong Dai 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Quantization-Aware Binaural MWF Based Noise Reduction Incorporating External Wireless DevicesabstractFor hearing-impaired listeners, both ambient noise suppression and directional source binaural cues preservation are required, such that a complete spatial awareness of the acoustic scene can be obtained. It was shown that the binaural multichannel Wiener filter (MWF) with partial noise preservation can achieve joint noise reduction and binaural cues preservation and incorporating an external microphone signal improves the performance of binaural MWFs. Motivated by this, we propose a binaural MWF incorporating external wireless devices in this paper. First, we theoretically analyze the performance of the MWF in terms of output signal-to-noise ratio (SNR) and binaural cues preservation errors. As in practice the external devices are power driven with a limited amount of battery resource and the power consumption heavily depends on the transmission rate, given an expected noise reduction performance we then optimize the bit-rate for a single external microphone. Further, we consider to minimize the total power consumption over multiple external devices under a constraint on the output SNR, which turns out to be a rate distribution problem. The proposed rate-distributed binaural MWF is evaluated using a hearing-aid setup with various dynamics. It is shown that the proposed method can obtain a desired SNR at a much lower bit-rate, and an expected trade-off between SNR gain and binaural cues preservation accuracy can be obtained by optimizing the bit-rates. Increasing the bit-rates improves both instrumental speech quality and speech intelligibility. Jie Zhang 0042, Changheng Li |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Distributed Rate-Constrained LCMV BeamformingabstractIn this letter, we propose a decentralized framework for rate-distributed linearly constrained minimum variance (LCMV) beamforming in wireless acoustic sensor networks. To save the energy usage within the network, we propose to minimize the transmission cost and put a constraint on the noise reduction performance. Subsequently, we decentralize the obtained LCMV filter structure by exploiting an imposed block diagonal form of the noise correlation matrix. As a result, the beamformer weights are calculated in a decentralized fashion and each node can determine its quantization rate locally. Finally, numerical results validate the proposed method. Jie Zhang 0042, Andreas I. Koutrouvelis, Richard Heusdens, Richard C. Hendriks |
IEEE Signal Process. Lett. | 1 |
| 2019 | Relative Acoustic Transfer Function Estimation in Wireless Acoustic Sensor NetworksabstractIn this paper, we present an algorithm to estimate the relative acoustic transfer function (RTF) of a target source in wireless acoustic sensor networks (WASNs). Two well-known methods to estimate the RTF are the covariance subtraction (CS) method and the covariance whitening (CW) approach, the latter based on the generalized eigenvalue decomposition. Both methods depend on the use of the noisy correlation matrix, which, in practice, has to be estimated using limited and (in WASNs) quantized data. The bit rate and the fact that we use limited data records therefore directly affect the accuracy of the estimated RTFs. Therefore, we first theoretically analyze the estimation performance of the two approaches in terms of bit rate. Second, we propose a rate-distribution method by minimizing the power usage and constraining the expected estimation error for both RTF estimators. The optimal rate distributions are found by using convex optimization techniques. The model-based methods, however, are impractical due to the dependence on the true RTFs. We therefore further develop two greedy rate-distribution methods for both approaches. Finally, numerical simulations on synthetic data and real audio recordings show the superiority of the proposed approaches in power usage compared to uniform rate allocation. We find that in order to satisfy the same RTF estimation accuracy, the rate-distributed CW methods consume much less transmission energy than the CS-based methods. Jie Zhang 0042, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Microphone Subset Selection for MVDR Beamformer Based Noise ReductionabstractIn large-scale wireless acoustic sensor networks (WASNs), many of the sensors will only have a marginal contribution to a certain estimation task. Involving all sensors increases the energy budget unnecessarily and decreases the lifetime of the WASN. Using microphone subset selection, also termed as sensor selection, the most informative sensors can be chosen from a set of candidate sensors to achieve a prescribed inference performance. In this paper, we consider microphone subset selection for minimum variance distortionless response (MVDR) beamformer based noise reduction. The best subset of sensors is determined by minimizing the transmission cost while constraining the output noise power (or signal-to-noise ratio). Assuming the statistical information on correlation matrices of the sensor measurements is available, the sensor selection problem for this model-driven scheme is first solved by utilizing convex optimization techniques. In addition, to avoid estimating the statistics related to all the candidate sensors beforehand, we also propose a data-driven approach to select the best subset using a greedy strategy. The performance of the greedy algorithm converges to that of the model-driven method, while it displays advantages in dynamic scenarios as well as on computational complexity. Compared to a sparse MVDR or radius-based beamformer, experiments show that the proposed methods can guarantee the desired performance with significantly less transmission costs. Jie Zhang 0042, Sundeep Prabhakar Chepuri, Richard C. Hendriks, Richard Heusdens |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Rate-Distributed Spatial Filtering Based Noise Reduction in Wireless Acoustic Sensor NetworksabstractIn wireless acoustic sensor networks (WASNs), sensors typically have a limited energy budget as they are often battery-driven. Energy efficiency is, therefore, essential for the design of algorithms in WASNs. One way to reduce energy costs is to select only the sensors that are most informative, a problem known as sensor selection . In this way, only sensors that significantly contribute to the task at hand will be involved. In this paper, we consider a more general approach, which is based on rate-distributed spatial filtering. Depending on the distance over which a transmission takes place, the bit rate directly influences the energy consumption. We try to minimize the battery usage due to transmission, while constraining the noise reduction performance. This results in an efficient rate allocation strategy, which depends on the underlying signal statistics, as well as the distance from sensors to a fusion center (FC). Through the utilization of a linearly constrained minimum variance beamformer, the problem is derived as a semidefinite program. Furthermore, we show that rate allocation is more general than sensor selection, and sensor selection can be seen as a special case of the presented rate-allocation solution, e.g., the best microphone subset can be determined by thresholding the rates. Finally, numerical simulations for estimating several target sources in a WASN demonstrate that the proposed method outperforms the sensor-selection-based approaches in terms of energy usage, and we find that the sensors close to the FC and point sources are allocated with higher rates. Jie Zhang 0042, Richard Heusdens, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Binaural Sound Localization Based on Reverberation Weighting and Generalized Parametric MappingabstractBinaural sound source localization is an important technique for speech enhancement, video conferencing, and human-robot interaction, etc. However, in realistic scenarios, the reverberation and environmental noise would degrade the precision of sound direction estimation. Therefore, reliable sound localization is essential to practical applications. To deal with these disturbances, this paper presents a novel binaural sound source localization approach based on reverberation weighting and generalized parametric mapping. First, the reverberation weighting as a preprocessing stage, is used to separately suppress the early and late reverberation, while preserving interaural cues. Then, two binaural cues, i.e., interaural time and intensity differences, are extracted from the frequency-domain representations of dereverberated binaural signals for the online localization. Their corresponding templates are established using the training data. Furthermore, the generalized parametric mapping is proposed to build a generalized parametric model for describing relationships between azimuth and binaural cues analytically. Finally, a two-step sound localization process is introduced to refine azimuth estimation based on the generalized parametric model and template matching. Experiments in both simulated and real scenarios validate that the proposed method can achieve better localization performance compared to state-of-the-art methods. Hong Liu 0008, Jie Zhang 0042, Xiaofei Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Probabilistic binaural multiple sources localization based on time-delay compensation estimator and clustering analysisabstractSound source localization (SSL) is an essential technique in many applications, such as robot audition, human-robot interaction and speech capturing. However, SSL from a binaural input is still a challenging problem, particularly when multiple sources are active simultaneously. In this work, we propose a multi-sources localization framework based on the time-delay compensation (TDC) estimator and clustering analysis. The TDC estimator is a simultaneous operator to estimate binaural cues, which breaks the limitation of independent processors for binaural cues extraction. The multi-sources decision is realized by clustering analysis for the binaural cues of multiple signal frames. In experiments, we demonstrate that the localization performance is improved compared to the methods that assume the number of spatial stationary sources to be known. Results with both simulated and recorded impulse responses show that robust performance can be achieved with limited prior training, and our method is also adaptive to different sound activities. Hong Liu 0008, Mengdi Yue, Jie Zhang 0042 |
IROS | 3 |
| 2015 | Binaural sound source localization based on generalized parametric model and two-layer matching strategy in complex environmentsabstractBinaural sound source localization is an important technique involving Human-Robot Interaction (HRI), video conference, speech enhancement, etc. In many real application scenarios, especially for closed environments, the affect of reverberation and noise would degrade the precision of position estimations. Therefore, a new binaural sound source localization method based on generalized parametric model and two-layer matching strategy is proposed in this paper for complex environments. Firstly, cepstral prefiltering is utilized for dereverberation of binaural signals. Then, two binaural cues computed from a dual-channel frequency representation, are combined to estimate the azimuths of sources. Additionally, the generalized parametric model is presented to describe the relationship between the azimuth and binaural cues through finding the optimal scaling factors from training data. At last, a two-layer matching strategy based on Bayesian rule is used to make the final decision, which can effectively decrease the computation complexity. Experiments have validated the proposed approach and show that it achieves favorably better results compared with several available methods without extra spacial burden. Hong Liu 0008, Jie Zhang 0042 |
ICRA | 3 |
| 2015 | Direction of arrival estimation based on reverberation weighting and noise error estimator
Jie Zhang 0042, Hong Liu 0008 |
INTERSPEECH | 2 |
| 2014 | A binaural sound source localization model based on time-delay compensation and interaural coherenceabstractBinaural sound source localization is an important technique involving speech capture and enhancement. However, the simple array structure makes it hard to localize sources in complex noisy conditions. This paper presents a novel algorithm based on time-delay compensation (TDC) and inter-aural coherence for binaural sound localization. Firstly, the TDC of binaural signals is used to estimate interaural time-delay (ITD) and interaural intensity difference (IID) instead of generalized cross correlation and logarithmic energy ratio. Then the interaural coherence is utilized to select reliable frames and reduce the variance of ITDs. Finally, a hierarchical framework, which successfully reduces computation complexity, is applied to make a decision of location based on Bayesian rule. Our innovation lies in that both ITD and IID are foremost yielded by TDC. Compared with other popular algorithms, experiments show that the most extrusive superiority of this method is complexity for both time and storage. Hong Liu 0008, Jie Zhang 0042 |
ICASSP | 2 |
| 2014 | A new hierarchical binaural sound source localization method based on Interaural Matching FilterabstractBinaural sound source localization is an important technique in friendly Human-Robot Interaction (HRI) for its easy-implementation with only two microphones. This paper develops a robust method based on a hierarchical probabilistic model. Reliable frequency sub-bands are used to PHAT — ργ method in the first layer to obtain a priori crude Interaural Time-Delay (ITD) and the probabilistic distribution of candidate azimuths. The second layer utilizes Interaural Intensity Difference (IID) to reduce matching time and refine candidate azimuths as well as elevations. A novel feature named Interaural Matching Filter (IMF), which can eliminate the difference between ITDs and IIDs, is proposed in the third layer. The probability of sound source location is acquired by computing the similarity between the IMF of received binaural signal and the IMFs in templates. Finally, combined with the former ITDs and IIDs, the similarity matrix is used to make a decision of sound source location based on a Bayes rule. Our innovation lies in adding selecting reliable frequency components into time-delay estimation and foremost taking IMF as a feature of sound source. Compared with several state-of-the-art algorithms, experimental results show our approach has a better performance even in noisy environments without increasing storages, and also has less time complexity. Hong Liu 0008, Jie Zhang 0042, Zhuo Fu |
ICRA | 2 |