Qiquan Zhang

dblp:211/6562 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0001-5089-6317ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement
abstract
Time-frequency (T-F) domain methods for monaural speech enhancement have benefited from the success of deep learning. Recently, focus has been put on designing two-stream network models to predict amplitude mask and phase separately, or, coupling the amplitude and phase into Cartesian coordinates and constructing real and imaginary pairs. However, most methods suffer from the alignment modeling of amplitude and phase (real and imaginary pairs) in a two-stream network framework, which inevitably incurs performance restrictions. In this paper, we introduce a graph Fourier transform defined with the singular value decomposition (GFT-SVD), resulting in real-valued time-graph representation for neural speech enhancement. This real-valued representation-based GFT-SVD provides an ability to align the modeling of amplitude and phase, leading to avoiding recovering the target speech phase information. Our findings demonstrate the effects of real-valued time-graph representation based on GFT-SVD for neutral speech enhancement. The extensive speech enhancement experiments establish that the combination of GFT-SVD and DNN outperforms the combination of GFT with the eigenvector decomposition (GFT-EVD) and magnitude estimation UNet, and outperforms the short-time Fourier transform (STFT) and DNN, regarding objective intelligibility and perceptual quality. We release our source code at: https://github.com/Wangfighting0015/GFTproject.
Tianrui Wang, Meng Ge, Qiquan Zhang, Zirui Ge, Zhen Yang 0001
ICASSP4
2025 Dual-stream Noise and Speech Information Perception based Speech Enhancement
Longbiao Wang, Qiquan Zhang, Jianwu Dang 0001
Expert Syst. Appl.3
2024 Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
abstract
Xiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, Leibny Paola Garcia Perera, EngSiong Chng, Lina Yao. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Xiangyu Zhang 0005, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, L. Paola García-Perera, Chng Eng Siong
EMNLP4
2024 When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
abstract
Depression is a critical concern in global mental health, prompting extensive research into AIbased detection methods.Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications.However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities.Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped.In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection.We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks.By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts.This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals.Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines.In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals.
Xiangyu Zhang 0005, Hexin Liu, Kaishuai Xu, Qiquan Zhang, Daijiao Liu, Beena Ahmed, Julien Epps
EMNLP4
2024 GLMB 3D Speaker Tracking with Video-Assisted Multi-Channel Audio Optimization Functions
abstract
Speaker tracking plays a significant role in numerous real-world human robot interaction (HRI) applications. In recent years, there has been a growing interest in utilizing multi-sensory information, such as complementary audio and visual signals, to address the challenges of speaker tracking. Despite the promising results, existing approaches still encounter difficulties in accurately determining the speaker’s true location, particularly in adverse conditions such as speech pauses, reverberation, or visual occlusions, leading to missed detections or spurious estimates. In this paper, we propose a novel speaker tracking method based on the Generalized Labelled Multi-Bernoulli (GLMB) filter. Our method operates in 3D space using audio information captured by a microphone array and video streams obtained from a monocular camera. The GLMB-based tracker effectively handles outliers in location estimates and maintains tracking during periods of missed detections. Experiments conducted on the publicly available AV16.3 dataset show that our proposal surpasses other competitive methods with improved results.
Xinyuan Qian 0001, Zexu Pan, Qiquan Zhang, Kainan Chen, Shoufeng Lin
ICASSP3
2024 An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech Enhancement
abstract
Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encoding exactly impacts speech enhancement based on Transformer architectures. In this paper, we perform a comprehensive empirical study evaluating five positional encoding methods, i.e., Sinusoidal and learned absolute position embedding (APE), T5-RPE, KERPLE, as well as the Transformer without positional encoding (No-Pos), across both causal and noncausal configurations. We conduct extensive speech enhancement experiments, involving spectral mapping and masking methods. Our findings establish that positional encoding is not quite helpful for the models in a causal configuration, which indicates that causal attention may implicitly incorporate position information. In a noncausal configuration, the models significantly benefit from the use of positional encoding. In addition, we find that among the four position embeddings, relative position embeddings outperform APEs.
Qiquan Zhang, Meng Ge, Hongxu Zhu, Eliathamby Ambikairajah, Zhaoheng Ni, Haizhou Li 0001
ICASSP1
2024 Binaural Selective Attention Model for Target Speaker Extraction
Hanyu Meng, Qiquan Zhang, Xiangyu Zhang 0005, Vidhyasaharan Sethu, Eliathamby Ambikairajah
INTERSPEECH2
2024 An Exploration of Length Generalization in Transformer-Based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Eliathamby Ambikairajah, Haizhou Li 0001
INTERSPEECH1
2024 Deep Cross-Modal Retrieval Between Spatial Image and Acoustic Speech
abstract
Cross-modal Retrieval (CMR) is formulated for the scenarios where the queries and retrieval results are of different modalities. Existing Cross-modal Retrieval (CMR) studies mainly focus on the common contextualized information between text transcripts and images, and the synchronized event information in audio-visual recordings. Unlike all previous works, in this article, we investigate the geometric correspondence between images and speech recordings captured in the same space and formulate a novel CMR task, called Spatial Image-Acoustic Retrieval (SIAR). To this end, we first design a novel speech encoder that consists of convolution neural networks and transformer layers, to learn space-aware speech representations. Then, to eliminate the cross-modal inherent discrepancy, we propose the Contrastive Speech Image Retrieval (CSIR) method which uses supervised contrastive learning to attract the same-space cross-modal features while repelling the ones from different spaces. Finally, image and speech features are directly compared and we predict the SIAR result with the maximum similarity. Extensive experiments demonstrate that our proposed speech encoder can recognize space from human speeches with superior performance over the other prevailing networks. It also sets our penultimate goal of speech-to-speech retrieval. Furthermore, our CSIR proposal can successfully perform bi-directional SIAR between spatial images and reverberant speeches with promising results. Code and data will be available.
Xinyuan Qian 0001, Wei Xue 0002, Qiquan Zhang, Ruijie Tao, Haizhou Li 0001
IEEE Trans. Multim.3
2023 Ripple Sparse Self-Attention for Monaural Speech Enhancement
abstract
The use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally prohibited for long speech recordings. Moreover, it allows each time frame to attend to all time frames, neglecting the strong local correlations of speech signals. This study presents a simple yet effective sparse self-attention for speech enhancement, called ripple attention, which simultaneously performs fine- and coarse-grained modeling for local and global dependencies, respectively. Specifically, we employ local band attention to enable each frame to attend to its closest neighbor frames in a window at fine granularity, while employing dilated attention outside the window to model the global dependencies at a coarse granularity. We evaluate the efficacy of our ripple attention for speech enhancement on two commonly used training objectives. Extensive experimental results consistently confirm the superior performance of the ripple attention design over standard full self-attention, blockwise attention, and dual-path attention (Sep-Former) in terms of speech quality and intelligibility.
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Zhaoheng Ni, Haizhou Li 0001
ICASSP1
2023 PoE: A Panel of Experts for Generalized Automatic Dialogue Assessment
abstract
Chatbots are expected to be knowledgeable across multiple domains, e.g. for daily chit-chat, exchange of information, and grounding in emotional situations. To effectively measure the quality of such conversational agents, a model-based automatic dialogue evaluation metric (ADEM) is expected to perform well across multiple domains. Despite significant progress, existing ADEMs tend to perform well only on data that are similar to its training data (overfit to its training domain). This calls for a domain-generalized metric that can assess dialogues of different characteristics. To this end, we propose aPanel of Experts(PoE), a multitask network that consists of a shared transformer encoder and a collection of lightweight adapters. The shared encoder captures the general knowledge of dialogues across domains, while each adapter specializes in one specific domain and serves as a domain expert. To validate the idea, we construct a high-quality multi-domain dialogue dataset leveraging data augmentation and pseudo-labeling. The PoE network is comprehensively assessed on 16 dialogue evaluation datasets spanning a wide range of dialogue domains. It achieves state-of-the-art performance in terms of mean Spearman correlation over all the evaluation datasets. It exhibits better zero-shot generalization than existing state-of-the-art ADEMs and the ability to easily adapt to new domains with few-shot transfer learning.
Chen Zhang 0055, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 A Time-Frequency Attention Module for Neural Speech Enhancement
abstract
Speech enhancement plays an essential role in a wide range of speech processing applications. Recent studies on speech enhancement tend to investigate how to effectively capture the long-term contextual dependencies of speech signals to boost performance. However, these studies generally neglect the time-frequency (T-F) distribution information of speech spectral components, which is equally important for speech enhancement. In this paper, we propose a simple yet very effective network module, which we term the T-F attention (TFA) module, that uses two parallel attention branches, i.e., time-frame attention and frequency-channel attention, to explicitly exploit position information to generate a 2-D attention map to characterise the salient T-F speech distribution. We validate our TFA module as part of two widely used backbone networks (residual temporal convolution network and Transformer) and conduct speech enhancement with four most popular training objectives. Our extensive experiments demonstrate that our proposed TFA module consistently leads to substantial enhancement performance improvements in terms of the five most widely used objective metrics, with negligible parameter overheads. In addition, we further evaluate the efficacy of speech enhancement as a front-end for a downstream speech recognition task. Our evaluation results show that the TFA module significantly improves the robustness of the system to noisy conditions.
Qiquan Zhang, Xinyuan Qian 0001, Zhaoheng Ni, Aaron Nicolson, Eliathamby Ambikairajah, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Speech-Oriented Sparse Attention Denoising for Voice User Interface Toward Industry 5.0
abstract
The adoption of voice user interface (VUI) will promote network automation with enhanced efficiency with reduced simplicity and operating expense in Industry 5.0. Given the noisy environments, speech denoising is indispensable for the VUI in Internet of Things (IoT) or Industrial IoT (IIoT). Despite Transformer's recent success in speech denoising, the adopted full self-attention suffers from quadratic complexity, which challenges the computational power of the IoT/IIoT components. Considering the strong local correlations of speech signals, a speech-oriented sparse attention denoising scheme is developed to keep the meaningful local and global dependencies while mitigating the redundant attentions, resulting in a significant reduction in computational complexity. With the full self-attention as the baseline, experimental results revealed that the proposed scheme achieves a better denoising performance and yields a lower computational cost, indicating the strong potential for various VUI application scenarios in IoT and IIoT toward Industry 5.0.
Hongxu Zhu, Qiquan Zhang, Peng Gao 0005, Xinyuan Qian 0001
IEEE Trans. Ind. Informatics2
2022 FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation
abstract
Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment 1 .However, they either perform turn-level evaluation or look at a single dialogue quality dimension.One would expect a good evaluation metric to assess multiple quality dimensions at the dialogue level.To this end, we are motivated to propose a multi-dimensional dialogue-level metric, which consists of three sub-metrics with each targeting a specific dimension.The submetrics are trained with novel self-supervised objectives and exhibit strong correlations with human judgment for their respective dimensions.Moreover, we explore two approaches to combine the sub-metrics: metric ensemble and multitask learning.Both approaches yield a holistic metric that significantly outperforms individual sub-metrics.Compared to the existing state-of-the-art metric, the combined metrics achieve around 16% relative improvement on average across three high-quality dialoguelevel evaluation benchmarks.
Chen Zhang 0055, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs, Haizhou Li 0001
EMNLP3
2022 Time-Frequency Attention for Monaural Speech Enhancement
abstract
Most studies on speech enhancement generally don’t explicitly consider the energy distribution of speech in time-frequency (T-F) representation, which is important for accurate prediction of mask or spectra. In this paper, we present a simple yet effective T-F attention (TFA) module, where a 2-D attention map is produced to provide differentiated weights to the spectral components of T-F representation. To validate the effectiveness of our proposed TFA module, we use the residual temporal convolution network (ResTCN) as the backbone network and conduct extensive experiments on two commonly used training targets. Our experiments demonstrate that applying our TFA module significantly improves the performance in terms of five objective evaluation metrics with negligible parameter overhead. The evaluation results show that the proposed ResTCN with the TFA module (ResTCN+TFA) consistently outperforms other baselines by a large margin.
Qiquan Zhang, Zhaoheng Ni, Aaron Nicolson, Haizhou Li 0001
ICASSP1
2022 Deep Audio-Visual Beamforming for Speaker Localization
abstract
Generalized Cross Correlation (GCC) is the most popular localization technique over the past decades and can be extended with the beamforming method e.g. Steered Response Power (SRP) when multiple microphone pairs exist. Considering the promising results of Deep Learning (DL) strategies over classical approaches, in this work, instead of directly using Generalized Cross Correlation (GCC), SRP is derived with the DL-learnt ideal correlation functions for each pair of a microphone array. To deploy visual information, we explore the Conditional Variational Auto-Encoder (CVAE) framework in which the audio generative process is conditioned on the visual features encoded by face detections. The vision-derived auxiliary correlation function eventually contributes to the back-end beamformer for improved localization performance. To the best of our knowledge, this is the first deep-generative audiovisual method for speaker localization. Experimental results demonstrate our superior performance over other competitive methods, especially when the speech signal is corrupted by noise.
Xinyuan Qian 0001, Qiquan Zhang, Guohui Guan 0001, Wei Xue 0002
IEEE Signal Process. Lett.2
2021 Temporal Convolutional Network with Frequency Dimension Adaptive Attention for Speech Enhancement
abstract
Despite much progress, most temporal convolutional networks (TCN) based speech enhancement models are mainly focused on modeling the long-term temporal contextual dependencies of speech frames, without taking into account the distribution information of speech signal in frequency dimension. In this study, we propose a frequency dimension adaptive attention (FAA) mechanism to improve TCNs, which guides the model selectively emphasize the frequency-wise features with important speech information and also improves the representation capability of network. Our extensive experimental investigation demonstrates that the proposed FAA mechanism is able to consistently provide significant improvements in terms of speech quality (PESQ), intelligibility (STOI) and three other composite metrics. More promisingly, it has better generalization ability to real-world noisy environment.
Qiquan Zhang, Aaron Nicolson, Haizhou Li 0001
Interspeech1
2021 PhaseDCN: A Phase-Enhanced Dual-Path Dilated Convolutional Network for Single-Channel Speech Enhancement
abstract
Recent deep neural network (DNN) based single-channel speech enhancement methods have achieved remarkable results in the time-frequency (TF) magnitude domain. To further improve the quality and intelligibility of enhanced speech, the attention to phase enhancement is also increasing. In this paper, we propose a novel dilated convolutional network (DCN) model to simultaneously enhance the magnitude and phase of noisy speech. Unlike the direct complex spectral mapping methods, we take the complex spectrum of the signal as the main target and the ideal ratio mask (IRM) as the auxiliary target in a multi-target learning framework to achieve their complementary advantages. Firstly, a feature extraction module is introduced to achieve the fusion of local and long-term features. Two different targets are learned separately, but share the common feature extraction module, which is helpful to extract more general and suitable features. During the joint learning, the intermediate estimation of the IRM target in the auxiliary path, contributing as the attention gating factors, helps to distinguish the speech or non-speech components of the complex-valued signals in the main path. To leverage more fine-grained long-term contextual information, we introduce a multi-scale dilated convolution approach for feature encoding. Moreover, the proposed model is a causal system, which can fully meet the low latency requirements of real-time speech products. Experimental results show that, compared with other advanced systems, the proposed model not only has better speech denoising performance and phase estimation accuracy, but also generalizes better in the speaker, noise, and channel mismatch cases.
Lu Zhang 0055, Mingjiang Wang, Qiquan Zhang
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Probability decision-driven speech enhancement algorithm based on human acoustic perception
abstract
In this study, a novel human acoustic perception motivated Wiener filter speech enhancement system is presented to cope with real‐world interfering background noises. Guiding by the speech presence probability, two alternative methods are proposed to reduce the noise by adopting the audible sound pressure level (SPL) and the masking characteristic of the human auditory system to achieve better listening comfort level. More specifically, when the probability of speech presence in the noisy signal is less than the decision threshold, a new SPL compressed method effectively reduces the noise. When the speech presence probability is more than the decision threshold, an improved acoustical mask threshold constrained Wiener filter approach enhances the noisy speech. Moreover, in order to evaluate the performance of the new system, the proposed algorithm is compared with the classic prior signal‐to‐noise ratio‐based Wiener filter and three acoustic perception related algorithms. The experimental results show that the proposed algorithm significantly outperforms the four comparing algorithms in terms of speech quality and intelligibility either in stationary or moderate non‐stationary noisy environments. Thus, the intended approach can be employed as the front‐end module for various speech‐related applications.
Lu Zhang 0055, Mingjiang Wang, Qiquan Zhang
IET Signal Process.4
2020 Learning reinforced attentional representation for end-to-end visual tracking
Peng Gao 0005, Qiquan Zhang, Fei Wang 0036, Liyi Xiao, Hamido Fujita, Yan Zhang 0066
Inf. Sci.2
2020 DeepMMSE: A Deep Learning Approach to MMSE-Based Noise Power Spectral Density Estimation
abstract
An accurate noise power spectral density (PSD) tracker is an indispensable component of a single-channel speech enhancement system. Bayesian-motivated minimum mean-square error (MMSE)-based noise PSD estimators have been the most prominent in recent time. However, they lack the ability to track highly non-stationary noise sources due to current methods of a priori signal-to-noise (SNR) estimation. This is caused by the underlying assumption that the noise signal changes at a slower rate than the speech signal. As a result, MMSE-based noise PSD trackers exhibit a large tracking delay and produce noise PSD estimates that require bias compensation. Motivated by this, we propose an MMSE-based noise PSD tracker that employs a temporal convolutional network (TCN) a priori SNR estimator. The proposed noise PSD tracker, called DeepMMSE makes no assumptions about the characteristics of the noise or the speech, exhibits no tracking delay, and produces an accurate estimate that requires no bias correction. Our extensive experimental investigation shows that the proposed DeepMMSE method outperforms state-of-the-art noise PSD trackers and demonstrates the ability to track abrupt changes in the noise level. Furthermore, when employed in a speech enhancement framework, the proposed DeepMMSE method is able to outperform state-of-the-art noise PSD trackers, as well as multiple deep learning approaches to speech enhancement. Availability: DeepMMSE is available at: https://github.com/anicolson/DeepXi.
Qiquan Zhang, Aaron Nicolson, Mingjiang Wang, Kuldip K. Paliwal, Chenxu Wang 0002
IEEE ACM Trans. Audio Speech Lang. Process.1