VLDB 2026 Research / reviewers in the wild / expert
Hangting Chen
dblp:226/2015
· DBLP profile ↗
31ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0002-4085-4364ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 21 · 5 first-author · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language ModelsabstractExtending pre-trained text Large Language Models (LLMs)’s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention in the speech research community. However, building a unified speech understanding and generation model still faces the following challenges: (1) Due to the huge modality gap between speech and text tokens, extending text LLMs to unified speech LLMs relies on large-scale paired data for fine-tuning, and (2) Generation and understanding tasks prefer information at different levels, e.g., generation benefits from detailed acoustic features, while understanding favors high-level semantics. This divergence leads to difficult performance optimization in one unified model. To solve these challenges, in this paper, we present two key insights in speech tokenization and speech language modeling. Specifically, we first propose an Understanding-driven Speech Tokenizer (USTokenizer), which extracts high-level semantic information essential for accomplishing understanding tasks using text LLMs. In this way, USToken enjoys better modality commonality with text, which reduces the difficulty of modality alignment in adapting text LLMs to speech LLMs. Secondly, we present DualSpeechLM, a dual-token modeling framework that concurrently models USToken as input and acoustic token as output within a unified, end-to-end framework, seamlessly integrating speech understanding and generation capabilities. Furthermore, we propose a novel semantic supervision loss and a Chain-of-Condition (CoC) strategy to stabilize model training and enhance speech generation performance. Experimental results demonstrate that our proposed approach effectively fosters a complementary relationship between understanding and generation tasks, highlighting the promising strategy of mutually enhancing both tasks in one unified model. Dongchao Yang, Yiwen Shao, Hangting Chen, Jiankun Zhao, Zhiyong Wu 0001, Helen M. Meng, Xixin Wu |
AAAI | 4 |
| 2025 | SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task EditorabstractThe emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about partial adjustments or editing of existing songs is still underexplored, which allows for more flexible and effective production. In this paper, we present SongEditor, the first song editing paradigm that introduces the editing capabilities into language-modeling song generation approaches, facilitating both segment-wise and track-wise modifications. SongEditor offers the flexibility to adjust lyrics, vocals, and accompaniments, as well as synthesizing songs from scratch. The core components of SongEditor include a music tokenizer, an autoregressive language model, and a diffusion generator, enabling generating an entire section, masked lyrics, or even separated vocals and background music. Extensive experiments demonstrate that the proposed SongEditor achieves exceptional performance in end-to-end song editing, as evidenced by both objective and subjective metrics. Shuai Wang 0016, Hangting Chen, Jianwei Yu 0001, Wei Tan 0011, Rongzhi Gu, Yaoxun Xu, Yizhi Zhou, Haina Zhu, Haizhou Li 0001 |
AAAI | 3 |
| 2025 | AudioComposer: Towards Fine-grained Audio Generation with Natural Language DescriptionsabstractCurrent Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control networks. However, this usually leads to complex system design and difficulties due to the requirement for reference frame-level conditions. To address these challenges, we propose AudioComposer, a novel TTA generation framework that relies solely on natural language descriptions (NLDs) to provide both content specification and style control information. To further enhance audio generative modeling, we employ flow-based diffusion transformers with the cross-attention mechanism to incorporate text descriptions effectively into audio generation processes, which can not only simultaneously consider the content and style information in the text inputs, but also accelerate generation compared to other architectures. Furthermore, we propose a novel and comprehensive automatic data simulation pipeline to construct data with fine-grained text descriptions, which significantly alleviates the problem of data scarcity in the area. Experiments demonstrate the effectiveness of our framework using solely NLDs as inputs for content specification and style control. The generation quality and controllability surpass stateof-the-art TTA models, even with a smaller model size.1 Hangting Chen, Dongchao Yang, Zhiyong Wu 0001, Xixin Wu |
ICASSP | 2 |
| 2025 | UniSep: Universal Target Audio Separation with Language Models at ScaleabstractWe propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited source numbers. We formulate the separation task as a sequence-to-sequence problem, and a large language model (LLM) is used to model the audio sequence in the discrete latent space, leveraging the power of LLM in handling complex mixture audios with large-scale data. Moreover, a novel pre-training strategy is proposed to utilize audio-only data, which reduces the efforts of large-scale data simulation and enhances the ability of LLMs to understand the consistency and correlation of information within audio sequences. We also demonstrate the effectiveness of scaling datasets in an audio separation task: we use large-scale data (36.5k hours), including speech, music, and sound, to train a universal target audio separation model that is not limited to a specific domain. Experiments show that UniSep achieves competitive subjective and objective evaluation results compared with single-task models. Hangting Chen, Dongchao Yang, Guangzhi Li, Shan Yang 0001, Zhiyong Wu 0001, Helen M. Meng, Xixin Wu |
ICME | 2 |
| 2025 | TSDT-Net: Ultra-Low-Complexity Two-Stage Model Combining Dual-Path-Transformer and Transform-Average-Concatenate Network for Speech Enhancement
Hangting Chen, Qingshan Yang, Jingcong Chen |
INTERSPEECH | 2 |
| 2025 | WAKE: Watermarking Audio with Key Enrichment
Yaoxun Xu, Jianwei Yu 0001, Hangting Chen, Zhiyong Wu 0001, Xixin Wu, Dong Yu 0001, Rongzhi Gu, Yi Luo 0004 |
INTERSPEECH | 3 |
| 2025 | TVC-MusicGen: Time-Varying Structure Control for Background Music Generation via Self-Supervised Training
Hangting Chen, Shuai Wang 0016, Haina Zhu, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2025 | MuCodec: Ultra Low-Bitrate Music Codec for Music GenerationabstractMusic generation is pivotal in multimedia, aiding creation and lowering the creative threshold. It focuses on generating music with clear vocals and harmonious accompaniment based on lyrics, combining high artistic creativity with technical challenges. The music codec is an important bridging component in large language model-based music generation, connecting language models with the generated music. However, existing neural codecs typically require token rates exceeding 50 Hz to achieve acceptable music quality, resulting in a context length that surpasses 12,000 tokens for a 4-minute song-a scale that is computationally demanding. This highlights the need for high-compression, high-fidelity music codecs that can reconstruct both vocals and accompaniment with high quality at low frame rates and bitrates, thereby better assisting music generation. To address this, we introduce MuCodec, designed for high-quality music reconstruction at ultra-low bitrates, facilitating more efficient music generation. MuCodec employs a two-stage training method, enabling its encoder, MuEncoder, to extract semantic and acoustic features in a unified representation. These features are discretized using residual vector quantization and converted into Mel-VAE features through flow matching, with reconstruction quality improved by representation alignment during training. The Mel-VAE features are then reconstructed into music using a pretrained Mel-VAE decoder and HiFi-GAN. To the best of our knowledge, MuCodec is the first codec capable of reconstructing 48kHz stereo music at an ultra-low bitrate of 0.35 kbps (25 Hz), achieving state-of-the-art performance in both subjective and objective evaluations, and can more effectively support music generation. Code and Demo: https://mucodec.github.io/Mucodec/. Yaoxun Xu, Hangting Chen, Jianwei Yu 0001, Wei Tan 0011, Shun Lei, Rongzhi Gu, Zhiyong Wu 0001 |
ACM Multimedia | 2 |
| 2025 | LeVo: High-Quality Song Generation with Multi-Preference AlignmentabstractRecent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation.
However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limitations in audio quality, musicality, instruction following, and vocal-instrument harmony.
To address these challenges, we introduce LeVo, a language model based framework consisting of LeLM and Music Codec.
LeLM is capable of parallel modeling of two types of tokens: mixed tokens, which represent the combined audio of vocals and accompaniment to achieve better vocal-instrument harmony, and dual-track tokens, which separately encode vocals and accompaniment for high-quality song generation.
It employs two decoder-only transformers and a modular extension training strategy to prevent interference between different token types.
To further enhance musicality and instruction following ability, we introduce a multi-preference alignment method based on Direct Preference Optimization (DPO).
This method handles diverse human preferences through a semi-automatic data construction process and post-training.
Experimental results demonstrate that LeVo significantly outperforms existing open-source methods in both objective and subjective metrics, while performing competitively with industry systems.
Ablation studies further justify the effectiveness of our designs.
Audio examples and source code are available at https://levo-demo.github.io and https://github.com/tencent-ailab/songgeneration. Shun Lei, Yaoxun Xu, Huaicheng Zhang, Wei Tan 0011, Hangting Chen, Yixuan Zhang 0005, Haina Zhu, Shuai Wang 0016, Zhiyong Wu 0001, Dong Yu 0001 |
NeurIPS | 6 |
| 2025 | SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion RefinementabstractGenerating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with local fidelity, resulting in outputs that lack musicality or suffer from incoherent progression and mismatched lyrics. This paper introduces SongBloom, a novel framework for full-length song generation that leverages an interleaved paradigm of autoregressive sketching and diffusion-based refinement. SongBloom employs an autoregressive diffusion model that combines the high fidelity of diffusion models with the scalability of language models.
Specifically, it gradually extends a musical sketch from short to long and refines the details from coarse to fine-grained. The interleaved generation paradigm effectively integrates prior semantic and acoustic context to guide the generation process.
Experimental results demonstrate that SongBloom outperforms existing methods across both subjective and objective metrics and achieves performance comparable to the state-of-the-art commercial music generation platforms.
Audio samples are available on our demo page:
https://cypress-yang.github.io/SongBloom_demo. Shuai Wang 0016, Hangting Chen, Wei Tan 0011, Jianwei Yu 0001, Haizhou Li 0001 |
NeurIPS | 3 |
| 2024 | SECap: Speech Emotion Captioning with Large Language ModelabstractSpeech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SECap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests. Yaoxun Xu, Hangting Chen, Jianwei Yu 0001, Qiaochu Huang, Zhiyong Wu 0001, Shixiong Zhang 0001, Guangzhi Li, Yi Luo 0004, Rongzhi Gu |
AAAI | 2 |
| 2024 | Complexity Scaling for Speech DenoisingabstractComputational complexity is critical when deploying deep learning-based speech denoising models for on-device applications. Most prior research focused on optimizing model architectures to meet specific computational cost constraints, often creating distinct neural network architectures for different complexity limitations. This study conducts complexity scaling for speech denoising tasks, aiming to consolidate models with various complexities into a unified architecture. We present a Multi-Path Transform-based (MPT) architecture to handle both low- and high-complexity scenarios. A series of MPT networks present high performance covering a wide range of computational complexities on the DNS challenge dataset. Moreover, inspired by the scaling experiments in natural language processing, we explore the empirical relationship between model performance and computational cost on the denoising task. As the complexity number of multiply-accumulate operations (MACs) is scaled from 50M/s to 15G/s on MPT networks, we observe a linear increase in the values of PESQ-WB and SI-SNR, proportional to the logarithm of MACs, which might contribute to the understanding and application of complexity scaling in speech denoising tasks.1 Hangting Chen, Chao Weng |
ICASSP | 1 |
| 2024 | Consistent and Relevant: Rethink the Query Embedding in General Sound SeparationabstractThe query-based audio separation usually employs specific queries to extract target sources from a mixture of audio signals. Currently, most query-based separation models need additional networks to obtain query embedding. In this way, separation model is optimized to be adapted to the distribution of query embedding. However, query embedding may exhibit mismatches with separation models due to inconsistent structures and independent information. In this paper, we present CaRE-SEP, a consistent and relevant embedding network for general sound separation to encourage a comprehensive reconsideration of query usage in audio separation. CaRE-SEP alleviates the potential mismatch between queries and separation in two aspects, including sharing network structure and sharing feature information. First, a Swin-Unet model with a shared encoder is conducted to unify query encoding and sound separation into one model, eliminating the network architecture difference and generating consistent distribution of query and separation features. Second, by initializing CaRE-SEP with a pretrained classification network and allowing gradient backpropagation, the query embedding is optimized to be relevant to the separation feature, further alleviating the feature mismatch problem. Experimental results indicate the proposed CaRE-SEP model substantially improves the performance of separation tasks. Moreover, visualizations validate the potential mismatch and how CaRE-SEP solves it. Hangting Chen, Dongchao Yang, Jianwei Yu 0001, Chao Weng, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 2 |
| 2024 | AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech DataabstractRecently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained. One reason for this limitation is that a considerable amount of the publicly available speech data is compromised by background noise, speech overlapping, lack of speech segmentation information, missing speaker labels, and incomplete transcriptions, which can largely hinder their usefulness. On the other hand, human annotation of speech data is both time-consuming and costly. To address this issue, we introduce an automatic in-the-wild speech data preprocessing framework (AutoPrep) in this paper, which is designed to enhance speech quality, generate speaker labels, and produce transcriptions automatically. The proposed AutoPrep framework comprises six components: speech enhancement, speech segmentation, speaker clustering, target speech extraction, quality filtering and automatic speech recognition. Experiments conducted on the open-sourced WenetSpeech and our self-collected AutoPrepWild corpora demonstrate that the proposed AutoPrep framework can generate preprocessed data with similar DNSMOS and PDNSMOS scores compared to several open-sourced TTS datasets. The corresponding TTS system can achieve up to 0.68 in-domain speaker similarity.1 Jianwei Yu 0001, Hangting Chen, Yanyao Bian, Yi Luo 0004, Jinchuan Tian, Mengyang Liu, Jiayi Jiang, Shuai Wang 0016 |
ICASSP | 2 |
| 2024 | Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex RecordingsabstractTarget speaker extraction (TSE) aims to extract the target speaker’s voice from the input mixture. Previous studies have concentrated on high-overlapping scenarios. However, real-world applications usually meet more complex scenarios like variable speaker overlapping and target speaker absence. In this paper, we introduces a framework to perform continuous TSE (C-TSE), comprising a target speaker voice activation detection (TSVAD) and a TSE model. This framework significantly improves TSE performance on similar speakers and enhances personalization, which is lacking in traditional diarization methods. In detail, unlike conventional TSVAD deployed to refine the diarization results, the proposed Attention-target speaker voice activation detection (A-TSVAD) to directly generate timestamps of the target speaker. We also explore some different integration methods of A-TSVAD and TSE by comparing the cascaded and parallel methods. The framework’s effectiveness is assessed using a range of metrics, including diarization and enhancement metrics. Our experiments demonstrate that A-TSVAD outperforms conventional methods in reducing diarization errors, when integrating A-TSVAD and TSE in a sequential cascaded manner further enhances extraction accuracy. Audio demos are available on our demo page1. Hangting Chen, Jianwei Yu 0001, Yuehai Wang |
IJCNN | 2 |
| 2023 | TSpeech-AI System Description to the 5th Deep Noise Suppression (DNS) ChallengeabstractThis report presents the development of Tencent AI Lab’s personalized speech enhancement system for the 2023 ICASSP Signal Processing Grand Challenge – deep noise suppression (DNS) challenge1, which includes the use of a modified band-split recurrent neural network (BSRNN) and a multi-resolution spectrogram discriminator to improve perceptual quality metrics. The proposed system outperforms the baseline system by 0.047 and 0.036 final scores in the headset track and speaker phone track, respectively, and ranks among the top-32in both tracks. Jianwei Yu 0001, Hangting Chen, Yi Luo 0004, Rongzhi Gu, Chao Weng |
ICASSP | 2 |
| 2023 | Ultra Dual-Path Compression For Joint Echo Cancellation And Noise SuppressionabstractEcho cancellation and noise reduction are essential for full-duplex communication, yet most existing neural networks have high computational costs and are inflexible in tuning model complexity. In this paper, we introduce time-frequency dual-path compression to achieve a wide range of compression ratios on computational cost. Specifically, for frequency compression, trainable filters are used to replace manually designed filters for dimension reduction. For time compression, only using frame skipped prediction causes large performance degradation, which can be alleviated by a post-processing network with full sequence modeling. We have found that under fixed compression ratios, dual-path compression combining both the time and frequency methods will give further performance improvement, covering compression ratios from 4x to 32x with little model size change. Moreover, the proposed models show competitive performance compared with fast FullSubNet and DeepFilterNet. Hangting Chen, Jianwei Yu 0001, Yi Luo 0004, Rongzhi Gu, Zhuocheng Lu, Chao Weng |
INTERSPEECH | 1 |
| 2023 | Bayes Risk Transducer: Transducer with Controllable Alignment Prediction
Jinchuan Tian, Jianwei Yu 0001, Hangting Chen, Brian Yan, Chao Weng, Dong Yu 0001, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2023 | High Fidelity Speech Enhancement with Band-split RNN
Jianwei Yu 0001, Hangting Chen, Yi Luo 0004, Rongzhi Gu, Chao Weng |
INTERSPEECH | 2 |
| 2023 | How to make embeddings suitable for PLDA
Zhuo Li 0020, Runqiu Xiao, Hangting Chen, Zhenduo Zhao, Pengyuan Zhang |
Comput. Speech Lang. | 3 |
| 2023 | First coarse, fine afterward: A lightweight two-stage complex approach for monaural speech enhancement
Feng Dang, Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
Speech Commun. | 2 |
| 2022 | DPT-FSNet: Dual-Path Transformer Based Full-Band and Sub-Band Fusion Network for Speech EnhancementabstractSub-band models have achieved promising results due to their ability to model local patterns in the spectrogram. Some studies further improve the performance by fusing sub-band and full-band information. However, the structure for the full-band and sub-band fusion model was not fully explored. This paper proposes a dual-path transformer-based full-band and sub-band fusion network (DPT-FSNet) for speech enhancement in the frequency domain. The intra and inter parts of the dual-path transformer model sub-band and full-band information, respectively. The features utilized by our proposed method are more interpretable than those utilized by the time-domain dual-path transformer. We conducted experiments on the Voice Bank + DEMAND and Interspeech 2020 Deep Noise Suppression (DNS) datasets to evaluate the proposed method. Experimental results show that the proposed method outperforms the current state-of-the-art. Feng Dang, Hangting Chen, Pengyuan Zhang |
ICASSP | 2 |
| 2022 | Beam-Guided TasNet: An Iterative Speech Separation Framework with Multi-Channel OutputabstractTime-domain audio separation network (TasNet) has achieved remarkable performance in blind source separation (BSS).Classic multi-channel speech processing framework employs signal estimation and beamforming.For example, Beam-TasNet links multi-channel convolutional TasNet (MC-Conv-TasNet) with minimum variance distortionless response (MVDR) beamforming, which leverages the strong modeling ability of data-driven network and boosts the performance of beamforming with an accurate estimation of speech statistics.Such integration can be viewed as a directed acyclic graph by accepting multi-channel input and generating multi-source output.In this paper, we design a "multi-channel input, multi-channel multi-source output" (MIMMO) speech separation system entitled "Beam-Guided TasNet", where MC-Conv-TasNet and MVDR can interact and promote each other more compactly under a directed cyclic flow.Specifically, the first stage uses Beam-TasNet to generate estimated single-speaker signals, which favors the separation in the second stage.The proposed framework facilitates iterative signal refinement with the guide of beamforming and seeks to reach the upper bound of the MVDR-based methods.Experimental results on the spatialized WSJ0-2MIX demonstrate that the Beam-Guided TasNet has achieved an SDR of 21.5 dB, exceeding the baseline Beam-TasNet by 4.1 dB under the same model size and narrowing the gap with the oracle signal-based MVDR to 2 dB. Hangting Chen, Yi Yang 0057, Feng Dang, Pengyuan Zhang |
INTERSPEECH | 1 |
| 2022 | The HCCL System for the NIST SRE21
Zhuo Li 0020, Runqiu Xiao, Hangting Chen, Zhenduo Zhao |
INTERSPEECH | 3 |
| 2021 | Power Pooling: An Adaptive Pooling Function for Weakly Labelled Sound Event DetectionabstractAccess to large corpora with strongly labelled sound events is expensive and difficult in engineering applications. Many researches turn to address the problem of how to detect both the types and the timestamps of sound events with weak labels that only specify the types. This task can be treated as a multiple instance learning (MIL) problem, and a key to it in the sound event detection (SED) task is the design of a pooling function. The linear softmax pooling function achieves state-of-the-art performance since it can vary both the signs and the magnitudes of gradients. However, linear softmax pooling cannot flexibly deal with sound events of different time scales. In this paper, we propose a power pooling function which can automatically adapt to various sound events. By adding a trainable parameter to each event, power pooling can provide more accurate gradients for frames in a clip than other pooling functions. On both weakly supervised and semi-supervised SED datasets, the proposed power pooling function outperforms linear softmax pooling on both coarse-grained and fine-grained metrics. Specifically, it improves the event-based F1 score by 11.4% and 10.2% relatively on the two datasets. While this paper focuses on SED applications, the proposed method can be applied to MIL tasks in other domains. Yuzhuo Liu, Hangting Chen, Pengyuan Zhang |
IJCNN | 2 |
| 2021 | Improved Speech Enhancement Using a Complex-Domain GAN with Fused Time-Domain and Time-Frequency Domain Constraints
Feng Dang, Pengyuan Zhang, Hangting Chen |
Interspeech | 3 |
| 2021 | A dual-stream deep attractor network with multi-domain learning for speech dereverberation and separation
Hangting Chen, Pengyuan Zhang |
Neural Networks | 1 |
| 2020 | Improved Guided Source Separation Integrated with a Strong Back-End for the CHiME-6 Dinner Party Scenario
Hangting Chen, Pengyuan Zhang, Qian Shi 0001, Zuozhen Liu |
INTERSPEECH | 1 |
| 2019 | An Audio Scene Classification Framework with Embedded Filters and a DCT-based Temporal ModuleabstractDeep convolutional neural network (DCNN) has recently improved the performance of acoustic scene classification. However, the input features of the network are usually based on predefined hand-tailored filters, which may not apply to the specific tasks. To overcome this, we propose a hybrid framework that jointly trains the front-end filters and the back-end DCNN. Also, a novel temporal module based on the discrete cosine transform (DCT) is inserted after the high-level feature map of the network, thus enabling us to utilize time information without a reduction of training samples. Our single system, composed of the fine-tuned wavelet front-end and the DCNN back-end, with the integrated DCT-based temporal module, has achieved an accuracy of 79.20% in the evaluation set in DCASE17, gaining around 3% and 8% accuracy improvement compared with scalogram-DCNN and FBank-DCNN systems, respectively. Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 1 |
| 2019 | Speaker-Invariant Feature-Mapping for Distant Speech Recognition via Adversarial Teacher-Student Learning
Long Wu, Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2018 | Deep Convolutional Neural Network with Scalogram for Audio Scene Modeling
Hangting Chen, Pengyuan Zhang, Haichuan Bai, Qingsheng Yuan, Xiuguo Bao, Yonghong Yan 0002 |
INTERSPEECH | 1 |