VLDB 2026 Research / reviewers in the wild / expert
Chuanzeng Huang
dblp:298/1082
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FAF-Filt: Frequency-aware Fourier Filter for Sound Event DetectionabstractCapturing time-frequency patterns along the frequency axis, is crucial for the precision of sound event detection systems. Frequency dynamic convolution (FDY) and a series of its variants which incorporate frequency-adaptive kernels in standard 2D convolutions, have demonstrated remarkable performance, yet also suffered from high computational costs. To address the issue, we propose an efficient and light-weighted frequency-aware Fourier filter (FAF-Filt), which performs a 2D Fourier transform on features to the frequency domain and employs a learnable frequency-aware filter to process the transformed features, thereby integrating global information more effectively to extract decisive frequency components. In addition, frequency-adaptive convolution (FA-Conv) is adopted to further strengthen the representative ability of convolution, which incorporates the frequency-aware attention mechanism into the inputs and outputs of the convolutions. Experimental results exhibit superiority of the proposed method, achieving comparable performance with FDY-CRNN in terms of polyphonic sound event scores (PSDS) with a significantly 56% reduction in parameters. Xiaohuai Le, Zhuangqi Chen, Xianjun Xia, Chuanzeng Huang |
ICASSP | 5 |
| 2025 | AF-Vocoder: Artifact-Free Neural Vocoder with Global Artifact Filter
Zhuangqi Chen, Xianjun Xia, Xiaohuai Le, Chuanzeng Huang |
INTERSPEECH | 5 |
| 2025 | Multistage Universal Speech Enhancement System for URGENT Challenge
Xiaohuai Le, Zhuangqi Chen, Xianjun Xia, Chuanzeng Huang |
INTERSPEECH | 5 |
| 2025 | CBA-Whisper: Curriculum Learning-Based AdaLoRA Fine-Tuning on Whisper for Low-Resource Dysarthric Speech Recognition
Tianyi Tan, Xiaohuai Le, Wenzhi Fan, Xianjun Xia, Chuanzeng Huang |
INTERSPEECH | 6 |
| 2024 | RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
Mingshuai Liu, Zhuangqi Chen, Xiaopeng Yan, Yuanjun Lv, Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001 |
INTERSPEECH | 6 |
| 2024 | BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
Xianjun Xia, Chuanzeng Huang, Yijian Xiao, Lei Xie 0001 |
INTERSPEECH | 3 |
| 2023 | Pretraining Conformer with ASR for Speaker VerificationabstractThis paper proposes to pretrain Conformer with automatic speech recognition (ASR) task for speaker verification. Conformer combines convolution neural network (CNN) and Transformer model for modeling local and global features, respectively. Recently, multi-scale feature aggregation Conformer (MFA-Conformer) has been proposed for automatic speaker verification. MFA-Conformer concatenates frame-level outputs from all Conformer blocks for further pooling. However, our experiments show that Conformer can be easily overfitted with limited speaker recognition training data. To avoid overfitting, we propose to transfer the knowledge learned from ASR to speaker verification. Specifically, an ASR pretrained Conformer is used to initialize the training of MFA-Conformer for speaker verification. Our experiments show that pretraining Conformer with ASR leads to significant performance gains across model sizes. The best model achieves 0.48%, 0.71% and 1.54% EER on Voxceleb1-O, Voxceleb1-E, and Voxceleb1-H, respectively. Danwei Cai, Weiqing Wang 0004, Ming Li 0026, Chuanzeng Huang |
ICASSP | 5 |
| 2023 | Exploring Universal Singing Speech Language Identification Using Self-Supervised Learning Based Front-End FeaturesabstractDespite the great performance of language identification (LID), there is a lack of large-scale singing LID databases to support the research of singing language identification (SLID). This paper proposed a over 3200 hours dataset used for singing language identification, called Slingua. As the baseline, we explore two self-supervised learning (SSL) models, WavLM and Wav2vec2, as the feature extractors for both SLID and universal singing speech language identification (ULID), compared with the traditional handcraft feature. Moreover, by training with speech language corpus, we compare the performance difference of the universal singing speech language identification. The final results show that the SSL-based features exhibit more robust generalization, especially for low-resource and open-set scenarios. The database can be downloaded following this repository: https://github.com/Doctor-Do/Slingua. Xingming Wang, Chuanzeng Huang, Ming Li 0026 |
ICASSP | 4 |
| 2023 | Memory Augmented Lookup Dictionary Based Language Modeling for Automatic Speech RecognitionabstractRecent studies have shown that using an external Language Model (LM) benefits the end-to-end Automatic Speech Recognition (ASR).However, predicting tokens that appear less frequently in the training set is still quite challenging.The longtail prediction problems have been widely studied in many applications, but only been addressed by a few studies for ASR and LMs.In this paper, we propose a new memory augmented lookup dictionary based Transformer architecture for LM.The newly introduced lookup dictionary incorporates rich contextual information in training set, which is vital to correctly predict long-tail tokens.With intensive experiments on Chinese and English data sets, our proposed method is proved to outperform the baseline Transformer LM by a great margin on both word/character error rate and tail tokens error rate.This is achieved without impact on the decoding efficiency.Overall, we demonstrate the effectiveness of our proposed method in boosting the ASR decoding performance, especially for longtail tokens. Yukun Feng, Chuanzeng Huang, Yuxuan Wang 0002 |
INTERSPEECH | 4 |
| 2023 | Language-universal Phonetic Encoder for Low-resource Speech Recognition
Chuanzeng Huang, Yuxuan Wang 0002 |
INTERSPEECH | 4 |
| 2023 | Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition
Chuanzeng Huang |
INTERSPEECH | 4 |
| 2022 | VoiceFixer: A Unified Framework for High-Fidelity Speech RestorationabstractSpeech restoration aims to remove distortions in speech signals. Prior methods mainly focus on a single type of distortion, such as speech denoising or dereverberation. However, speech signals can be degraded by several different distortions simultaneously in the real world. It is thus important to extend speech restoration models to deal with multiple distortions. In this paper, we introduce VoiceFixer, a unified framework for high-fidelity speech restoration. VoiceFixer restores speech from multiple distortions (e.g., noise, reverberation, and clipping) and can expand degraded speech (e.g., noisy speech) with a low bandwidth to 44.1 kHz full-bandwidth high-fidelity speech. We design VoiceFixer based on (1) an analysis stage that predicts intermediate-level features from the degraded speech, and (2) a synthesis stage that generates waveform using a neural vocoder. Both objective and subjective evaluations show that VoiceFixer is effective on severely degraded speech, such as real-world historical speech recordings. Samples of VoiceFixer are available at https://haoheliu.github.io/voicefixer. Haohe Liu, Xubo Liu 0001, Qiuqiang Kong, Qiao Tian 0001, Yan Zhao 0010, DeLiang Wang, Chuanzeng Huang, Yuxuan Wang 0002 |
INTERSPEECH | 7 |
| 2022 | Non-intrusive Speech Quality Assessment with a Multi-Task Learning based Subband Adaptive Attention Temporal Convolutional Neural Network
Xiaofeng Shu, Chuxiang Shang, Yan Zhao 0010, Chengshuai Zhao, Yehang Zhu, Chuanzeng Huang, Yuxuan Wang 0002 |
INTERSPEECH | 7 |