Hao Zhang 0112

dblp:55/2270-112 · DBLP profile ↗
← Back
20ranked-venue papers
13as first author
17since 2021 · last 2026
0000-0002-6749-8320ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 12 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning
abstract
Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded substantial benefits for audio question answering, and effectively leveraging deep reasoning remains an open challenge, with LALMs still falling short of achieving human-level auditory-language reasoning. To address these limitations, we propose Audio-Thinker, a reinforcement learning framework designed to enhance the reasoning capabilities of LALMs through improved adaptability, consistency, and effectiveness. Our approach introduces an adaptive think accuracy reward, enabling the model to adjust its reasoning strategies based on task complexity. Furthermore, we incorporate an external reward model to evaluate the overall consistency and quality of the reasoning process, complemented by think-based rewards that assist the model in distinguishing between valid and flawed reasoning paths during training. Experimental results demonstrate that Audio-Thinker models outperform existing reasoning-oriented LALMs across various benchmark tasks, exhibiting superior reasoning and generalization capabilities.
Chenxing Li, Wenfu Wang, Hao Zhang 0112, Hualei Wang, Meng Yu 0003, Dong Yu 0001
AAAI4
2026 Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
abstract
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored.Crucially, there is no consensus on whether audio-language models can build effective general-purpose audio encoders, nor a systematic understanding of how pretraining objectives behave across diverse tasks and scales.We identify three key barriers: limited scale of audio-text corpora, limited coverage of audio attributes in existing caption corpora, and lack of systematic exploration and evaluation.To fill this gap, we present the first principled empirical study of ALP.We first introduce Cap-tionStew, a 10.7M caption dataset aggregating open-source audio-text corpora across multiple domains and captioning focuses.We then conduct the first comprehensive evaluation comparing contrastive and captioning objectives for learning audio representation across speech, music, and environmental sound tasks.Our results not only demonstrate that ALP yields competitive, transferable representations, but reveal critical trade-offs: contrastive learning offers superior data efficiency, while captioning exhibits better scalability.Furthermore, we find that the benefits of supervised initialization often diminish at larger scales, challenging common practices.By grounding these claims in empirical evidence, we establish a viable pathway toward general-purpose audio representation learning, guiding future research.
Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao, Hao Zhang 0112, Dong Yu 0001
ACL (1)5
2025 Continual Pre-training for Codec-Based Speech LLMs: Balancing Understanding and Generation
abstract
Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially with codec-based representations. We propose a continual pre-training (CPT) framework that adapts a textual LLM to handle codec-discretized speech, mitigating modality mismatch and preserving linguistic reasoning. Our unified model supports both understanding and generation, achieving strong results across ASR, TTS, S2T-Trans, and S2S-Trans. Notably, we present the first end-to-end, single-pass S2S-Trans system using only neural codec tokens, without intermediate transcriptions, translations, or semantic tokens. CPT proves essential for crossmodal alignment and task generalization, making it a powerful tool for building robust, unified speech LLMs.
Jiatong Shi, Jinchuan Tian, Junrui Ni, Hao Zhang 0112, Shinji Watanabe 0001, Dong Yu 0001
ASRU5
2025 Preference Alignment Improves Language Model-Based TTS
abstract
Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences of reward models, enhancing the desirability of the generated content. This study presents a thorough empirical evaluation of how preference alignment algorithms, particularly Direct Preference Optimization (DPO), enhance LM-based TTS. With a 1.15B parameter LM-based TTS model, we demonstrate that preference alignment consistently improves intelligibility, speaker similarity, and proxy subjective evaluation scores, with the latter two metrics surpassing even human speech in certain evaluations. We also show preference alignment is applicable to low-resource scenarios and effectively generalized to out-of-domain applications.
Jinchuan Tian, Jiatong Shi, Hao Zhang 0112, Jianwei Yu 0001, Shinji Watanabe 0001, Dong Yu 0001
ICASSP4
2025 EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Jiarui Hai, Yong Xu 0004, Hao Zhang 0112, Chenxing Li, Helin Wang, Mounya Elhilali, Dong Yu 0001
INTERSPEECH3
2024 SMRU: Split-And-Merge Recurrent-Based UNet For Acoustic Echo Cancellation And Noise Suppression
abstract
The proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which yield promising performance. Nevertheless, they rarely consider the deployment generality in different processing scenarios, such as edge devices, and cloud processing. To this end, this paper proposes a general model, termed SMRU, to cover different application scenarios. The novelty lies in two-fold. First, a multi-scale band split layer and band merge layer are proposed to effectively fuse local frequency bands for lower complexity modeling. Besides, by simulating the multi-resolution feature modeling characteristic of the classical UNet structure, a novel recurrent-dominated UNet is devised. It consists of multiple variable frame rate blocks, each of which involves the causal time down-/upsampling layer with varying compression ratios and the dualpath structure for inter- and intra-band modeling. The model is configured from $50 \mathrm{M} / \mathrm{s}$ to $6.8 \mathrm{G} / \mathrm{s}$ in terms of MACs, and the experimental results show that the proposed approach yields competitive or even better performance over existing baselines, and has the full potential to adapt to more general scenarios with varying complexity requirements.
Zhihang Sun, Andong Li, Rilin Chen, Hao Zhang 0112, Meng Yu 0003, Yi Zhou 0014, Dong Yu 0001
SLT4
2024 Enhanced Acoustic Howling Suppression via Hybrid Kalman Filter and Deep Learning Models
abstract
This paper presents a comprehensive study addressing the challenging problem of acoustic howling suppression (AHS) through the fusion of Kalman filter and deep learning techniques. We introduce two integration approaches: HybridAHS, which concatenates Kalman and neural networks (NN), and NeuralKalmanAHS, where NN modules are embedded inside the Kalman filter for signal and parameter estimation. In HybridAHS, we explore two implementation methods. One is trained offline using pre-processed signals with a light training burden, while the other employs a recursive training strategy with training signals generated adaptively. The offline model serves as an initialization for recursively training the other model. With NeuralKalmanAHS, we harness the power of NN modules to refine the reference signal and improve covariance matrices estimation in the Kalman filter, resulting in enhanced feedback suppression. Our methods capitalize on the strengths of traditional and deep learning-based AHS techniques. We have explored different variants of combining Kalman filter and NN and systematically compared their howling suppression performance, providing users with versatile solutions for addressing AHS. Furthermore, by employing the proposed recursive training, we effectively mitigate the mismatch issues that plagued previous NN-based AHS methods. Extensive experimental results show the superiority of our approach over baseline techniques.
Hao Zhang 0112, Yixuan Zhang 0005, Meng Yu 0003, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Neuralkalman: A Learnable Kalman Filter for Acoustic Echo Cancellation
abstract
The robustness of the Kalman filter to double talk and its rapid convergence make it a popular approach for addressing acoustic echo cancellation (AEC) challenges. However, the inability to model nonlinearity and the need to tune control parameters cast limitations on such adaptive filtering algorithms. In this paper, we integrate the frequency domain Kalman filter (FDKF) and deep neural networks (DNNs) into a hybrid method, called NeuralKalman, to leverage the advantages of deep learning and adaptive filtering algorithms. Specifically, we employ a DNN to estimate nonlinearly distorted far-end signals, a transition factor, and the nonlinear transition function in the state equation of the FDKF algorithm. Experimental results show that the proposed NeuralKalman improves the performance of FDKF significantly and outperforms strong baseline methods.
Yixuan Zhang 0005, Meng Yu 0003, Hao Zhang 0112, Dong Yu 0001, DeLiang Wang
ASRU3
2023 Hybrid AHS: A Hybrid of Kalman Filter and Deep Learning for Acoustic Howling Suppression
Hao Zhang 0112, Meng Yu 0003, Yuzhong Wu, Dong Yu 0001
INTERSPEECH1
2023 Deep MCANC: A deep learning approach to multi-channel active noise control
Hao Zhang 0112, DeLiang Wang
Neural Networks1
2023 Low-Latency Active Noise Control Using Attentive Recurrent Network
abstract
Processing latency is a critical issue for active noise control (ANC) due to the causality constraint of ANC systems. This paper addresses low-latency ANC in the context of deep learning (i.e. deep ANC). A time-domain method using an attentive recurrent network (ARN) is employed to perform deep ANC with smaller frame sizes, thus reducing algorithmic latency of deep ANC. In addition, we introduce a delay-compensated training to perform ANC using predicted noise for several milliseconds. Moreover, a revised overlap-add method is utilized during signal resynthesis to avoid the latency introduced due to overlaps between neighboring time frames. Experimental results show the effectiveness of the proposed strategies for achieving low-latency deep ANC. Combining the proposed strategies is capable of yielding zero, even negative, algorithmic latency without affecting ANC performance much, thus alleviating the causality constraint in ANC design.
Hao Zhang 0112, Ashutosh Pandey 0004, DeLiang Wang
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Neural Cascade Architecture for Joint Acoustic Echo and Noise Suppression
abstract
In this paper, we propose a neural cascade architecture for joint acoustic echo and noise suppression. The proposed cascade architecture consists of two modules. A convolutional recurrent network (CRN) is employed in the first module for complex spectral mapping. The output is then fed as an additional input to the second module, where a long short-term memory network (LSTM) is utilized for magnitude mask estimation. The entire architecture is trained in an end-to-end manner with the two modules optimized jointly using a single loss function. The final output is generated using the enhanced phase and magnitude obtained from the first and the second module, respectively. The cascade architecture enables the proposed method to obtain robust magnitude estimation as well as phase enhancement. Evaluation results show that the proposed method effectively suppresses acoustic echo and noise while preserving good speech quality, and significantly outperforms related methods.
Hao Zhang 0112, DeLiang Wang
ICASSP1
2022 Attentive Recurrent Network for Low-Latency Active Noise Control
abstract
Processing latency is a critical issue for active noise control (ANC) due to the causality constraint of ANC systems. This paper addresses low-latency ANC in the deep learning framework (i.e. deep ANC). A time-domain method using an attentive recurrent network is employed to perform deep ANC with smaller frame sizes, thus reducing algorithmic latency of deep ANC. In addition, a delay-compensated training strategy is introduced to perform ANC using predicted noise for several milliseconds. Moreover, we utilize a revised overlap-add method during signal resynthesis to avoid the latency introduced due to overlaps between neighboring time frames. Experimental results show that the proposed strategies are effective for achieving low-latency deep ANC. Combining the proposed strategies is capable of yielding zero, even negative, algorithmic latency without significantly affecting ANC performance.
Hao Zhang 0112, Ashutosh Pandey 0004, DeLiang Wang
INTERSPEECH1
2022 Neural Cascade Architecture for Multi-Channel Acoustic Echo Suppression
abstract
Traditional acoustic echo cancellation (AEC) works by identifying an acoustic impulse response using adaptive algorithms. This paper proposes a neural cascade architecture for joint acoustic echo and noise suppression to address both single-channel and multi-channel AEC (MCAEC) problems. The proposed cascade architecture consists of two modules. A convolutional recurrent network (CRN) is employed in the first module for complex spectral mapping. Its output is fed as an additional input to the second module, where a long short-term memory network (LSTM) is utilized for magnitude mask estimation. The entire architecture is trained in an end-to-end manner with the two modules optimized jointly using a single loss function. The final output is generated using the enhanced phase and magnitude obtained from the first and the second module, respectively. The cascade architecture enables the proposed method to obtain robust magnitude estimation as well as phase enhancement. The proposed method is investigated under different AEC setups. We find that the deep learning based approach avoids the no-uniqueness problem in traditional MCAEC. For MCAEC setups with multiple microphones, combining deep MCAEC with supervised beamforming further improves the system performance. Evaluation results show that the proposed approach effectively suppresses acoustic echo and noise while preserving speech quality, and consistently outperforms related methods under different setups.
Hao Zhang 0112, DeLiang Wang
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 A Deep Learning Method to Multi-Channel Active Noise Control
Hao Zhang 0112, DeLiang Wang
Interspeech1
2021 A Deep Learning Approach to Multi-Channel and Multi-Microphone Acoustic Echo Cancellation
Hao Zhang 0112, DeLiang Wang
Interspeech1
2021 Deep ANC: A deep learning approach to active noise control
Hao Zhang 0112, DeLiang Wang
Neural Networks1
2020 A Deep Learning Approach to Active Noise Control
Hao Zhang 0112, DeLiang Wang
INTERSPEECH1
2019 Deep Learning for Joint Acoustic Echo and Noise Cancellation with Nonlinear Distortions
Hao Zhang 0112, Ke Tan 0001, DeLiang Wang
INTERSPEECH1
2018 Deep Learning for Acoustic Echo Cancellation in Noisy and Double-Talk Scenarios
Hao Zhang 0112, DeLiang Wang
INTERSPEECH1