VLDB 2026 Research / reviewers in the wild / expert
Peiwen Sun
dblp:249/4025
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language ModelabstractRecent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focused on enhancing the architecture and scale of audio language models, as well as leveraging larger datasets, and generally, acoustic codecs, such as EnCodec, are used for audio tokenization. However, these codecs were originally designed for audio compression, which may lead to suboptimal performance in the context of audio LLM. Our research aims to address the shortcomings of current audio LLM codecs, particularly their challenges in maintaining semantic integrity in generated audio. For instance, existing methods like VALL-E, which condition acoustic token generation on text transcriptions, often suffer from content inaccuracies and elevated word error rates (WER) due to semantic misinterpretations of acoustic tokens, resulting in word skipping and errors. To overcome these issues, we propose a straightforward yet effective approach called X-Codec. X-Codec incorporates semantic features from a pre-trained semantic encoder before the Residual Vector Quantization (RVQ) stage and introduces a semantic reconstruction loss after RVQ. By enhancing the semantic ability of the codec, X-Codec significantly reduces WER in speech synthesis tasks and extends these benefits to non-speech applications, including music and sound generation. Our experiments in text-to-speech, music continuation, and text-to-sound tasks demonstrate that integrating semantic information substantially improves the overall performance of language models in audio generation. Zhen Ye 0006, Peiwen Sun, Jiahe Lei, Hongzhan Lin 0001, Xu Tan 0003, Zheqi Dai, Qiuqiang Kong, Jianyi Chen, Yike Guo, Wei Xue 0002 |
AAAI | 2 |
| 2025 | Both Ears Wide Open: Towards Language-Driven Spatial Audio GenerationabstractRecently, diffusion models have achieved great success in mono-channel audio generation.
However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions.
Controlling stereo audio with spatial contexts remains challenging due to high data costs and unstable generative models.
To the best of our knowledge, this work represents the first attempt to address these issues.
We first construct a large-scale, simulation-based, and GPT-assisted dataset, BEWO-1M, with abundant soundscapes and descriptions even including moving and multiple sources.
Beyond text modality, we have also acquired a set of images and rationally paired stereo audios through retrieval to advance multimodal generation.
Existing audio generation models tend to generate rather random and indistinct spatial audio.
To provide accurate guidance for Latent Diffusion Models, we introduce the SpatialSonic model utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance.
By leveraging spatial guidance, our model not only achieves the objective of generating immersive and controllable spatial audio from text but also extends to other modalities as the pioneer attempt.
Finally, under fair settings, we conduct subjective and objective evaluations on simulated and real-world data to compare our approach with prevailing methods.
The results demonstrate the effectiveness of our method, highlighting its capability to generate spatial audio that adheres to physical rules. Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 0006, Huadai Liu, Honggang Zhang 0002, Wei Xue 0002, Yike Guo |
ICLR | 1 |
| 2025 | OmniAudio: Generating Spatial Audio from 360-Degree VideoabstractTraditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and perspective video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets are available at https://github.com/liuhuadai/OmniAudio. The project website is available at https://OmniAudio-360V2SA.github.io. Huadai Liu, Tianyi Luo, Kaicheng Luo, Qikai Jiang, Peiwen Sun, Rongjie Huang 0001, Qian Chen 0003, Wen Wang 0001, Xiangtai Li, Shiliang Zhang, Zhijie Yan, Zhou Zhao 0001, Wei Xue 0002 |
ICML | 5 |
| 2024 | Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation
Juncheng Ma, Peiwen Sun, Yaoting Wang, Di Hu 0001 |
ECCV (73) | 2 |
| 2024 | Can Textual Semantics Mitigate Sounding Object Segmentation Preference?
Yaoting Wang, Peiwen Sun, Yuanchao Li, Honggang Zhang 0002, Di Hu 0001 |
ECCV (74) | 2 |
| 2024 | Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li 0001, Honggang Zhang 0002, Di Hu 0001 |
ECCV (74) | 2 |
| 2024 | Enhancing Few-shot Classification through Token Selection for Balanced LearningabstractIn recent years, patch-based approaches have shown promise in few-shot learning, with further improvements observed through the use of self-supervised learning. However, we observe that the mainstream object-oriented approach focuses mainly on the salient part of the subject and ignores the non-annotated part of the image. Based on the assumption that any patch of the image is beneficial to learning, we present an end-to-end learning framework, which reconsiders the whole image from a multi-level perspective. The learning of annotated subjects involves Direct Patch Learning (DPL) to promote balanced learning of different features and Gaussian Mixup (GMIX) to provide extra mixed patch-level labels. As for the non-annotated part, we utilize a cascading token selection strategy along with self-supervised learning to better utilize knowledge in the background in the current context by learning the consistent representation of different views from the same image. Finally, in inductive few-shot learning, our method outperforms many previous methods and achieves new state-of-the-art performance. Furthermore, it provides an insight that non-annotated parts are also favorable for few-shot learning. As an ablation study, the effectiveness of each designed component is verified. Wangding Zeng, Peiwen Sun, Honggang Zhang 0002 |
IJCNN | 2 |
| 2024 | Unveiling and Mitigating Bias in Audio Visual SegmentationabstractCommunity researchers have developed a range of advanced audio-visual segmentation models aimed at improving the quality of sounding objects' masks. While masks created by these models may initially appear plausible, they occasionally exhibit anomalies with incorrect grounding logic. We attribute this to real-world inherent preferences and distributions as a simpler signal for learning than the complex audio-visual grounding, which leads to the disregard of important modality information. Generally, the anomalous phenomena are often complex and cannot be directly observed systematically. In this study, we made a pioneering effort with the proper synthetic data to categorize and analyze phenomena as two types "audio priming bias" and "visual prior" according to the source of anomalies. For audio priming bias, to enhance audio sensitivity to different intensities and semantics, a perception module specifically for audio perceives the latent semantic information and incorporates information into a limited set of queries, namely active queries. Moreover, the interaction mechanism related to such active queries in the transformer decoder is customized to adapt to the need for interaction regulating among audio semantics. For visual prior, multiple contrastive training strategies are explored to optimize the model by incorporating a biased branch, without even changing the structure of the model. During experiments, observation demonstrates the presence and the impact that has been produced by the biases of the existing model. Finally, through experimental evaluation of AVS benchmarks, we demonstrate the effectiveness of our methods in handling both types of biases, achieving competitive performance across all three subsets. Peiwen Sun, Honggang Zhang 0002, Di Hu 0001 |
ACM Multimedia | 1 |
| 2024 | FlashSpeech: Efficient Zero-Shot Speech SynthesisabstractRecent progress in large-scale zero-shot speech synthesis has been significantly advanced by language models and diffusion models. However, the generation process of both methods is slow and computationally intensive. Efficient speech synthesis using a lower computing budget to achieve quality on par with previous work remains a significant challenge. In this paper, we present FlashSpeech, a large-scale zero-shot speech synthesis system with approximately 5% of the inference time compared with previous work. FlashSpeech is built on the latent consistency model and applies a novel adversarial consistency training approach that can train from scratch without the need for a pre-trained diffusion model as the teacher. Furthermore, a new prosody generator module enhances the diversity of prosody, making the rhythm of the speech sound more natural. The generation processes of FlashSpeech can be achieved efficiently with one or two sampling steps while maintaining high audio quality and high similarity to the audio prompt for zero-shot speech generation. Our experimental results demonstrate the superior performance of FlashSpeech. Notably, FlashSpeech can be about 20 times faster than other zero-shot speech synthesis systems while maintaining comparable performance in terms of voice quality and similarity. Furthermore, FlashSpeech demonstrates its versatility by efficiently performing tasks like voice conversion, speech editing, and diverse speech sampling. Audio samples can be found in https://flashspeech.github.io/ Zhen Ye 0006, Zeqian Ju, Haohe Liu, Xu Tan 0003, Jianyi Chen, Peiwen Sun, Weizhen Bian, Shulin He, Wei Xue 0002, Yike Guo |
ACM Multimedia | 7 |
| 2023 | A Method of Audio-Visual Person Verification by Mining Connections between Time Series
Peiwen Sun, Zishan Liu, Yougen Yuan, Taotao Zhang, Honggang Zhang 0002, Pengfei Hu 0004 |
INTERSPEECH | 1 |