EDBT 2026 Demo / reviewers in the wild / expert
Zhisheng Zheng
dblp:311/6002
· DBLP profile ↗
17ranked-venue papers
3as first author
17since 2021 · last 2025
0000-0001-7761-9790ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Rich Style-Prompted Text-to-Speech DatasetsabstractWe introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions.While rich abstract tags (e.g.guttural, nasal, pained) have been explored in small-scale human-annotated datasets, existing large-scale datasets only cover basic tags (e.g.low-pitched, slow, loud).We combine off-the-shelf text and speech embedders, classifiers and an audio language model to automatically scale rich tag annotations for the first time.ParaSpeechCaps covers a total of 59 style tags, including both speaker-level intrinsic tags and utterance-level situational tags.It consists of 282 hours of human-labelled data (PSC-Base) and 2427 hours of automatically annotated data (PSC-Scaled).We finetune Parler-TTS, an open-source style-prompted TTS model, on ParaSpeechCaps, and achieve improved style consistency (+7.9% Consistency MOS) and speech quality (+15.5% Naturalness MOS) over the best performing baseline that combines existing rich style tag datasets.We ablate several of our dataset design choices to lay the foundation for future work in this space.Our dataset, models and code are released at https://github. com/ajd12342/paraspeechcaps. Anuj Diwan, Zhisheng Zheng, David F. Harwath, Eunsol Choi |
EMNLP | 2 |
| 2025 | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech EditingabstractZhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Zhu Liu, Vimal Bhat, David Harwath. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Vimal Bhat, David F. Harwath |
EMNLP | 1 |
| 2025 | SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMsabstractAutomated Audio Captioning (AAC) aims to generate natural textual descriptions for input audio signals. Recent progress in audio pre-trained models and large language models (LLMs) has significantly enhanced audio understanding and textual reasoning capabilities, making improvements in AAC possible. In this paper, we propose SLAM-AAC to further enhance AAC with paraphrasing augmentation and CLAP-Refine through LLMs. Our approach uses the self-supervised EAT model to extract fine-grained audio representations, which are then aligned with textual embeddings via lightweight linear layers. The caption generation LLM is efficiently fine-tuned using the LoRA adapter. Drawing inspiration from the back-translation method in machine translation, we implement paraphrasing augmentation to expand the Clotho dataset during pre-training. This strategy helps alleviate the limitation of scarce audio-text pairs and generates more diverse captions from a small set of audio clips. During inference, we introduce the plug-and-play CLAP-Refine strategy to fully exploit multiple decoding outputs, akin to the n-best rescoring strategy in speech recognition. Using the CLAP model for audio-text similarity calculation, we could select the textual descriptions generated by multiple searching beams that best match the input audio. Experimental results show that SLAM-AAC achieves state-of-the-art performance on Clotho V2 and AudioCaps, surpassing previous mainstream models. Ziyang Ma 0001, Xiquan Li, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Kai Yu 0004, Xie Chen 0001 |
ICASSP | 6 |
| 2025 | DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio CaptioningabstractWhile automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when transferring across domains. To overcome these limitations, we present DRCap, a data-efficient and flexible zero-shot audio captioning system that requires text-only data for training and can quickly adapt to new domains without additional fine-tuning. DRCap integrates a contrastive language-audio pre-training (CLAP) model and a large language model (LLM) as its backbone. During training, the model predicts the ground-truth caption with a fixed text encoder from CLAP, whereas, during inference, the text encoder is replaced with the audio encoder to generate captions for audio clips in a zero-shot manner. To mitigate the modality gap of the CLAP model, we use both the projection strategy from the encoder side and the retrieval-augmented generation strategy from the decoder side. Specifically, audio embeddings are first projected onto a text embedding support to absorb extensive semantic information within the joint multi-modal space of CLAP. At the same time, similar captions retrieved from a datastore are fed as prompts to instruct the LLM, incorporating external knowledge to take full advantage of its strong generative capability. Conditioned on both the projected CLAP embedding and the retrieved similar captions, the model is able to produce a more accurate and semantically rich textual description. By tailoring the text embedding support and the caption datastore to the target domain, DRCap acquires a robust ability to adapt to new domains in a training-free manner. Experimental results demonstrate that DRCap outperforms all other zero-shot models in in-domain scenarios and achieves state-of-the-art performance in cross-domain scenarios. Xiquan Li, Ziyang Ma 0001, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Qiuqiang Kong, Xie Chen 0001 |
ICASSP | 6 |
| 2025 | Compositional Neural Distance Field with Latent Code Embedding for Dynamic ObjectsabstractIn recent years, implicit neural representation of 3D scenes has evolved significantly and has been rapidly extended to multiple application scenarios. However, using this type of approach for highly dynamic urban scenes remains a challenging problem. Typical images of objects captured by cameras on board autonomous vehicles from one trajectory contains limited number of views, leading unsatisfactory reconstruction quality. To address this problem, we propose a novel hybrid network structure that decomposes the urban scene into a static background and multiple dynamic objects. The background model outputs a signed distance field with its accuracy enhanced by surface constraints. To deal with the issues of insufficient camera view we use a pre-training strategy to leverage shape and appearance prior to the object model from external datasets. We design an autoencoder architecture to encode a category of objects and employ an attention-based fusion module to better extract features from multiple object images. Furthermore, we introduce a “symmetric completion” approach to leverage the inherent symmetry property of normal cars. During experiments with data from Carla Platform, we find that our model can reconstruct scenes with high-fidelity, generate novel-view successfully and edit 3D scene freely. Zhisheng Zheng, Alois C. Knoll |
ICTAI | 2 |
| 2025 | MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their MixabstractWe introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. These findings underscore the urgent need for greater research attention in audio-language reasoning, including both data and algorithm innovation. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area. Ziyang Ma 0001, Yinghao Ma, Yanqiao Zhu 0003, Yi-Wen Chao, Yuanzhe Chen, Zhuo Chen 0006, Jian Cong, Keliang Li, Siyou Li, Xinfeng Li, Xiquan Li, Zheng Lian 0004, Yuzhe Liang, Minghao Liu 0003, Zhikang Niu, Tianrui Wang, Yuping Wang 0005, Yuxuan Wang 0002, Guanrou Yang, Jianwei Yu 0001, Ruibin Yuan, Zhisheng Zheng, Ziya Zhou, Haina Zhu, Wei Xue 0002, Emmanouil Benetos, Kai Yu 0004, Chng Eng Siong, Xie Chen 0001 |
NeurIPS | 27 |
| 2024 | Leveraging Speech PTM, Text LLM, And Emotional TTS For Speech Emotion RecognitionabstractIn this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervised pre-trained models, and we found that data2vec has a good representation ability on the SER task. Second, we employed a powerful large language model (LLM), GPT-4, and emotional text-to-speech (TTS) model, Azure TTS, to generate emotionally congruent text and speech. We carefully designed the text prompt and dataset construction, to obtain the synthetic emotional speech data with high quality. Third, we studied different ways of data augmentation to promote the SER task with synthetic speech, including random mixing, adversarial training, transfer learning, and curriculum learning. Experiments and ablation studies on the IEMOCAP dataset demonstrate the effectiveness of our method, compared with other data augmentation methods, and data augmentation with other synthetic data. Ziyang Ma 0001, Wen Wu 0007, Zhisheng Zheng, Qian Chen 0003, Shiliang Zhang, Xie Chen 0001 |
ICASSP | 3 |
| 2024 | BAT: Learning to Reason about Spatial Sounds with Large Language ModelsabstractSpatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT’s superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments. Zhisheng Zheng, Puyuan Peng, Ziyang Ma 0001, Xie Chen 0001, Eunsol Choi, David F. Harwath |
ICML | 1 |
| 2024 | EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
Yuzhe Liang, Ziyang Ma 0001, Zhisheng Zheng, Xie Chen 0001 |
IJCAI | 4 |
| 2024 | EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
Ziyang Ma 0001, Hezhao Zhang, Zhisheng Zheng, Xiquan Li, Jiaxin Ye, Xie Chen 0001, Thomas Hain |
INTERSPEECH | 4 |
| 2023 | Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech RecognitionabstractSelf-supervised learning (SSL) has achieved great success in speech processing, but always with a large model size to increase the modeling capacity. This may limit its potential applications due to the expensive computation and memory costs introduced by the oversize model. Compression for SSL models has become an important research direction of practical value. To this end, we explore the effective distillation of HuBERT-based SSL models for automatic speech recognition. First, a comprehensive study of different student model structures is conducted. On top of this, as a supplement to the regression loss widely adopted in previous works, a discriminative loss is introduced for HuBERT to enhance the distillation performance, especially in low-resource scenarios. In addition, we design a simple and effective algorithm to distill the front-end input from waveform to Fbank feature, resulting in 17% parameter reduction and doubling inference speed, at marginal performance degradation. Changli Tang, Ziyang Ma 0001, Zhisheng Zheng, Xie Chen 0001, Weiqiang Zhang 0001 |
ASRU | 4 |
| 2023 | Fast-Hubert: an Efficient Training Framework for Self-Supervised Speech Representation LearningabstractRecent years have witnessed significant advancements in self-supervised learning (SSL) methods for speech-processing tasks. Various speech-based SSL models have been developed and present promising performance on a range of downstream tasks including speech recognition. However, existing speech-based SSL models face a common dilemma in terms of computational cost, which might hinder their potential application and in-depth academic research. To address this issue, we first analyze the computational cost of different modules during HuBERT pre-training and then introduce a stack of efficiency optimizations, which is named Fast-HuBERT in this paper. The proposed Fast-HuBERT can be trained in 1.1 days with 8 V100 GPUs on the Librispeech 960 h benchmark, without performance degradation, resulting in a 5.2x speedup, compared to the original implementation. Moreover, we explore two well-studied techniques in the Fast-HuBERT and demonstrate consistent improvements as reported in previous work.11The code for Fast-HuBERT training is available at https://github.com/yanghaha0908/FastHuBERT Guanrou Yang, Ziyang Ma 0001, Zhisheng Zheng, Yakun Song, Zhikang Niu, Xie Chen 0001 |
ASRU | 3 |
| 2023 | Front-End Adapter: Adapting Front-End Input of Speech Based Self-Supervised Learning for Speech RecognitionabstractRecent years have witnessed a boom in self-supervised learning (SSL) in various areas including speech processing. Speech based SSL models present promising performance in a range of speech related tasks. However, the training of SSL models is computationally expensive and a common practice is to fine-tune a released SSL model on the specific task. It is essential to use consistent front-end input during pre-training and fine-tuning. This consistency may introduce potential issues when the optimal front-end is not the same as that used in pre-training. In this paper, we propose a simple but effective front-end adapter to address this front-end discrepancy. By minimizing the distance between the outputs of different front-ends, the filterbank feature (Fbank) can be compatible with SSL models which are pre-trained with waveform. The experiment results demonstrate the effectiveness of our proposed front-end adapter on several popular SSL models for the speech recognition task. Xie Chen 0001, Ziyang Ma 0001, Changli Tang, Zhisheng Zheng |
ICASSP | 5 |
| 2023 | MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets
Ziyang Ma 0001, Zhisheng Zheng, Changli Tang, Xie Chen 0001 |
INTERSPEECH | 2 |
| 2023 | Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation
Ziyang Ma 0001, Zhisheng Zheng, Guanrou Yang, Yu Wang 0027, Chao Zhang 0031, Xie Chen 0001 |
INTERSPEECH | 2 |
| 2023 | Unsupervised Active Learning: Optimizing Labeling Cost-Effectiveness for Automatic Speech Recognition
Zhisheng Zheng, Ziyang Ma 0001, Yu Wang 0027, Xie Chen 0001 |
INTERSPEECH | 1 |
| 2022 | Noniterative f -x-y Streaming Prediction Filtering for Random Noise Attenuation on Seismic DataabstractRandom noise is unavoidable in seismic exploration, especially under complex-surface conditions and in deep-exploration environments. The current problems in random noise attenuation include preserving the nonstationary characteristics of the signal and reducing computational cost of broadband, wide-azimuth, and high-density data acquisition. To obtain high-quality images, traditional prediction filters (PFs) have proved effective for random noise attenuation, but these methods typically assume that the signal is stationary. Most nonstationary PFs use an iterative strategy to calculate the coefficients, which leads to high computational costs. In this study, we extended the streaming prediction theory to the frequency domain and proposed the$f$-$x$-$y$streaming prediction filter (SPF) to attenuate random noise. Instead of using the iterative optimization algorithm, we formulated a constraint least-squares problem to calculate the SPF and derived an analytical solution to this problem. The multidimensional streaming constraints are used to increase the accuracy of the SPF. We also modified the recursive algorithm to update the SPF with the snaky processing path, which takes full advantage of the streaming structure to improve the effectiveness of the SPF in high dimensions. In comparison with 2-D$f$-$x$SPF and 3-D$f$-$x$-$y$regularized nonstationary autoregression (RNA), we tested the practicality of the proposed method in attenuating random noise. Numerical experiments show that the 3-D$f$-$x$-$y$SPF is suitable for large-scale seismic data with the advantages of low computational cost, reasonable nonstationary signal protection, and effective random noise attenuation. Yang Liu 0354, Zhisheng Zheng |
IEEE Trans. Geosci. Remote. Sens. | 2 |