VLDB 2026 Research / reviewers in the wild / expert
Xize Cheng
dblp:334/2167
· DBLP profile ↗
45ranked-venue papers
5as first author
45since 2021 · last 2026
0000-0001-9708-3225ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 5 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and ColloquialnessabstractJingyu Lu, Yuhan Wang, Fan Zhuo, Xize Cheng, Changhao Pan, Xueyi Pu, Yifu Chen, Chenyuhao Wen, Tianle Liang, Zhou Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jingyu Lu 0001, Fan Zhuo, Xize Cheng, Changhao Pan, Xueyi Pu, Chenyuhao Wen, Tianle Liang, Zhou Zhao 0001 |
ACL (1) | 4 |
| 2026 | Emphasizing Domain Differences Through Interactive-Augmented Prompts in Continual Audio-Visual Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) has been studied for a long time in the literature. By leveraging the complementary information from both acoustic and visual modalities, this approach offers a promising solution for robust speech transcription. While recent AVSR models have achieved impressive performance on large-scale, uniformly distributed datasets, they often overlook the challenges posed by real-world scenarios-where data is collected across multiple sessions and environments, leading to significant domain shifts and heterogeneous distributions. Such heterogeneity can result in catastrophic forgetting and hinder the generalization ability of the conventional models. To bridge this gap, we introduce the Continual Audio-Visual Speech Recognition (CL-AVSR) problem, which formulates AVSR as a continual learning task. We establish a dedicated benchmark for CL-AVSR by designing three experimental scenarios that reflect real-world challenges: introducing varying background noise for the audio stream, degrading video quality for the visual stream, and dividing tasks by speaker characteristics to jointly affect both modalities. These scenarios systematically evaluate the model's ability to adapt and retain knowledge across dynamic and non-stationary data streams. To address the unique challenges of CL-AVSR, we propose the Interaction-enhanced Multimodal Prompt learning (IMP) framework. IMP builds upon a pre-trained AV-HuBERT backbone and integrates task-relevant soft prompts with cross-modal and cross-task interactions, enabling efficient knowledge transfer from high-quality source domains to typical low-quality target domains with minimal parameter overhead. The interactive prompts facilitate fine-grained alignment and adaptation between modalities and tasks, while contrastive regularization further mitigates catastrophic forgetting. Furthermore, we devise a multi-modal prompt selection strategy that leverages clustering-based feature analysis, empowering the model to dynamically select optimal prompts for unseen data distributions during inference. Extensive experiments on the LRS2 dataset demonstrate that IMP achieves substantial improvements over strong baselines, setting new state-of-the-art performance in all CL-AVSR scenarios. Our results highlight the effectiveness of IMP in enhancing continual learning capabilities for AVSR, paving the way for more robust and adaptable multi-modal speech recognition systems in real-world applications. Xize Cheng, Jingyuan Chen 0003, Tao Jin 0004, Zhongfei Zhang |
IEEE Trans. Image Process. | 2 |
| 2025 | A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal AdapterabstractEfficient transfer learning methods such as adapter-based methods have shown great success in unimodal models and vision-language models. However, existing methods have two main challenges in fine-tuning multimodal models. Firstly, they are designed for vision-language tasks and fail to extend to situations where there are more than two modalities. Secondly, they exhibit limited exploitation of interactions between modalities and lack efficiency. To address these issues, in this paper, we propose the loW-rank sequence multimodal adapter (Wander). We first use the outer product to fuse the information from different modalities in an element-wise way effectively. For efficiency, we use CP decomposition to factorize tensors into rank-one components and achieve substantial parameter reduction. Furthermore, we implement a token-level low-rank decomposition to extract more fine-grained features and sequence relationships between modalities. With these designs, Wander enables token-level interactions between sequences of different modalities in a parameter-efficient way. We conduct extensive experiments on datasets with different numbers of modalities, where Wander outperforms state-of-the-art efficient transfer learning methods consistently. The results fully demonstrate the effectiveness, efficiency and universality of Wander. Zirun Guo, Xize Cheng, Tao Jin 0004 |
AAAI | 2 |
| 2025 | T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI FeedbackabstractZehan Wang, Ke Lei, Chen Zhu, Jiawei Huang, Sashuai Zhou, Luping Liu, Xize Cheng, Shengpeng Ji, Zhenhui Ye, Tao Jin, Zhou Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zehan Wang 0001, Ke Lei, Jiawei Huang 0008, Sashuai Zhou, Luping Liu, Xize Cheng, Shengpeng Ji, Zhenhui Ye, Tao Jin 0004, Zhou Zhao 0001 |
ACL (1) | 7 |
| 2025 | CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic ModelingabstractMinghui Fang, Shengpeng Ji, Jialong Zuo, Hai Huang, Yan Xia, Jieming Zhu, Xize Cheng, Xiaoda Yang, Wenrui Liu, Gang Wang, Zhenhua Dong, Zhou Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Minghui Fang 0002, Shengpeng Ji, Jialong Zuo, Hai Huang 0013, Yan Xia 0006, Jieming Zhu, Xize Cheng, Xiaoda Yang, Wenrui Liu 0003, Zhenhua Dong, Zhou Zhao 0001 |
ACL (1) | 7 |
| 2025 | ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style ControlabstractIn this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker’s voice and enabling arbitrary control and adjustment of speaking style. Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS models cannot perform speaker-specific voice generation. Therefore, ControlSpeech focuses on a more challenging task—a TTS system with controllable timbre, content, and style at the same time. ControlSpeech takes speech prompts, content prompts, and style prompts as inputs and utilizes bidirectional attention and mask-based parallel decoding to capture codec representations corresponding to timbre, content, and style in a discrete decoupling codec space. Moreover, we analyze the many-to-many issue in textual style control and propose the Style Mixture Semantic Density (SMSD) module, which is based on Gaussian mixture density networks, to resolve this problem. To facilitate empirical validations, we make available a new style controllable dataset called VccmDataset. Our experimental results demonstrate that ControlSpeech exhibits comparable or state-of-the-art (SOTA) performance in terms of controllability, timbre similarity, audio quality, robustness, and generalizability. Codes are available at https://github.com/jishengpeng/ControlSpeech. Shengpeng Ji, Qian Chen 0003, Wen Wang 0001, Jialong Zuo, Minghui Fang 0002, Ziyue Jiang 0004, Hai Huang 0013, Zehan Wang 0001, Xize Cheng, Zhou Zhao 0001 |
ACL (1) | 9 |
| 2025 | Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow MatchingabstractJialong Zuo, Shengpeng Ji, Minghui Fang, Mingze Li, Ziyue Jiang, Xize Cheng, Xiaoda Yang, Chen Feiyang, Xinyu Duan, Zhou Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jialong Zuo, Shengpeng Ji, Minghui Fang 0002, Ziyue Jiang 0001, Xize Cheng, Xiaoda Yang, Feiyang Chen 0001, Xinyu Duan, Zhou Zhao 0001 |
ACL (1) | 6 |
| 2025 | VoxpopuliTTS: a large-scale multilingual TTS corpus for zero-shot speech generationabstractIn recent years, speech generation fields have achieved significant advancements, primarily due to improvements in large TTS (text-to-speech) systems and scalable TTS datasets. However, there is still a lack of large-scale multilingual TTS datasets, which limits the development of cross-language and multilingual TTS systems. Hence, we refine Voxpopuli dataset and propose VoxpopuliTTS dataset. This dataset comprises 30,000 hours of high-quality speech data, across 3 languages with multiple speakers and styles, suitable for various speech tasks such as TTS and ASR. To enhance the quality of speech data from Voxpopuli, we improve the existing processing pipeline by: 1) filtering out low-quality speech-text pairs based on ASR confidence scores, and 2) concatenating short transcripts by checking semantic information completeness to generate the long transcript. Experimental results demonstrate the effectiveness of the VoxpopuliTTS dataset and the proposed processing pipeline. Wenrui Liu 0003, Jionghao Bai, Xize Cheng, Jialong Zuo, Ziyue Jiang 0001, Shengpeng Ji, Minghui Fang 0002, Xiaoda Yang, Qian Yang 0006, Zhou Zhao 0001 |
COLING | 3 |
| 2025 | SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative LanguageabstractContrastive Language-Image Pre-training (CLIP) learns robust visual models through language supervision, making it a crucial visual encoding technique for various applications. However, CLIP struggles with comprehending spatial concepts in images, potentially restricting the spatial intelligence of CLIP-based AI systems. In this work, we propose SpatialCLIP, an enhanced version of CLIP with better spatial understanding capabilities. To capture the intricate 3D spatial relationships in images, we improve both "visual model" and "language supervision" of CLIP. Specifically, we design 3D-inspired ViT to replace the standard ViT in CLIP. By lifting 2D image tokens into 3D space and incorporating design insights from point cloud networks, our visual model gains greater potential for spatial perception. Meanwhile, captions with accurate and detailed spatial information are very rare. To explore better language supervision for spatial understanding, we re-caption images and perturb their spatial phrases as negative descriptions, which compels the visual model to seek spatial cues to distinguish these hard negative captions. With the enhanced visual model, we introduce SpatialLLaVA, following the same LLaVA-1.5 training protocol, to investigate the importance of visual representations for MLLM’s spatial intelligence. Furthermore, we create SpatialBench, a benchmark specifically designed to evaluate CLIP and MLLM in spatial reasoning. Spatial-CLIP and SpatialLLaVA achieve substantial performance improvements, demonstrating stronger capabilities in spatial perception and reasoning, while maintaining comparable results on general-purpose benchmarks. Zehan Wang 0001, Sashuai Zhou, Shaoxuan He, Haifeng Huang 0001, Lihe Yang, Xize Cheng, Shengpeng Ji, Tao Jin 0004, Hengshuang Zhao, Zhou Zhao 0001 |
CVPR | 7 |
| 2025 | PACHAT: Persona-Aware Speech Assistant for Multi-party DialogueabstractExtensive research on LLM-based spoken dialogue systems has significantly advanced the development of intelligent voice assistants.However, the integration of role information within speech remains an underexplored area, limiting its application in real-world scenarios, particularly in multi-party dialogue settings.With the growing demand for personalization, voice assistants that can recognize and remember users establish a deeper connection with them.We focus on enabling LLMs with speaker-awareness capabilities and enhancing their understanding of character settings through synthetic data to generate contextually appropriate responses.We introduce Persona-Dialogue, the first large-scale multi-party spoken dialogue dataset that incorporates speaker profiles.Based on this dataset, we propose PAChat, an architecture that simultaneously models both linguistic content and speaker features, allowing LLMs to map character settings to speaker identities in speech.Through extensive experiments, we demonstrate that PAChat successfully achieves speaker-specific responses, character understanding, and the generation of targeted replies in multi-party dialogue scenarios, surpassing existing spoken dialogue systems.For more details, please visit our demo page at https Xize Cheng, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin 0004 |
EMNLP | 2 |
| 2025 | Curriculum Learning aided Audio-Visual Speech Recognition with Arbitrary Speaker NumberabstractRecently, audio-visual speech recognition has attracted increasing attention. However, most existing works only focused on scenarios with two speakers. In this work, we study the effect of speaker number in AVSR task and propose an end-to-end audio-visual speech recognition framework under a more realistic condition where the speaker number is arbitrary. Specifically, we adopted curriculum learning to train models from easy scenarios to hard ones and introduce a new training strategy named Challenge-based Curriculum Learning (CBCL) that forces the model to focus on hard, challenging data instead of easy ones during training. Further, to avoid scenario bias from unbalanced sampling during curriculum learning, we propose a Speaker-number Aware Mixture-of-Expert (SA-MoE) mechanism to explicitly model the characteristic difference in scenarios with different speaker numbers. Yuxiao Lin, Tao Jin 0004, Xize Cheng, Zhou Zhao 0001, Fei Wu 0001 |
ICASSP | 3 |
| 2025 | Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching ModelabstractThis paper introduces PFlow-VC, a conditional flow matching voice conversion model that leverages fine-grained discrete pitch tokens and target speaker prompt information for expressive voice conversion (VC). Previous VC works primarily focus on speaker conversion, with further exploration needed in enhancing expressiveness (such as prosody and emotion) for timbre conversion. Unlike previous methods, we adopt a simple and efficient approach to enhance the style expressiveness of voice conversion models. Specifically, we pretrain a self-supervised pitch VQVAE model to discretize speaker-irrelevant pitch information and leverage a masked pitch-conditioned flow matching model for Mel-spectrogram synthesis, which provides in-context pitch modeling capabilities for the speaker conversion model, effectively improving the voice style transfer capacity. Additionally, we improve timbre similarity by combining global timbre embeddings with time-varying timbre tokens. Experiments on unseen LibriTTS test-clean and emotional speech dataset ESD show the superiority of the PFlow-VC model in both timbre conversion and style transfer. Audio samples are available on the demo page https://speechai-demo.github.io/PFlow-VC/. Jialong Zuo, Shengpeng Ji, Minghui Fang 0002, Ziyue Jiang 0001, Xize Cheng, Qian Yang 0006, Wenrui Liu 0003, Guangyan Zhang, Zehai Tu, Yiwen Guo, Zhou Zhao 0001 |
ICASSP | 5 |
| 2025 | OmniBind: Large-scale Omni Multimodal Representation via Binding SpacesabstractRecently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Meanwhile, multimodal representation models have emerged as the foundation for these versatile multimodal understanding and generation pipeline. Models like CLIP, CLAP and ImageBind can map their specialized modalities into respective joint spaces. To construct a high-quality omni representation space that can be shared and expert in any modality, we propose to merge these advanced models into a unified space in scale. With this insight, we present \textbf{OmniBind}, advanced multimodal joint representation models via fusing knowledge of 14 pre-trained spaces, which support 3D, audio, image, video and language inputs. To alleviate the interference between different knowledge sources in integrated space, we dynamically assign weights to different spaces by learning routers with two objectives: cross-modal overall alignment and language representation decoupling. Notably, since binding and routing spaces only require lightweight networks, OmniBind is extremely training-efficient. Extensive experiments demonstrate the versatility and superiority of OmniBind as an omni representation model, highlighting its great potential for diverse applications, such as any-query and composable multimodal understanding. Zehan Wang 0001, Minjie Hong, Luping Liu, Rongjie Huang 0001, Xize Cheng, Shengpeng Ji, Tao Jin 0004, Hengshuang Zhao, Zhou Zhao 0001 |
ICLR | 7 |
| 2025 | VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?abstractWith the rapid advancement of large models, voice assistants are gradually acquiring the ability to engage in open-ended daily conversations with humans. However, current spoken dialogue systems often overlook multi-modal information in audio beyond text, such as speech rate, volume, emphasis, and background sounds. Relying solely on Automatic Speech Recognition (ASR) can lead to the loss of valuable auditory cues, thereby weakening the system’s ability to generate contextually appropriate responses. To address this limitation, we propose \textbf{VoxDialogue}, a comprehensive benchmark for evaluating the ability of spoken dialogue systems to understand multi-modal information beyond text. Specifically, we have identified 12 attributes highly correlated with acoustic information beyond words and have meticulously designed corresponding spoken dialogue test sets for each attribute, encompassing a total of 4.5K multi-turn spoken dialogue samples. Finally, we evaluated several existing spoken dialogue models, analyzing their performance on the 12 attribute subsets of VoxDialogue. Experiments have shown that in spoken dialogue scenarios, many acoustic cues cannot be conveyed through textual information and must be directly interpreted from the audio input. In contrast, while direct spoken dialogue systems excel at processing acoustic signals, they still face limitations in handling complex dialogue tasks due to their restricted context understanding capabilities. All data and code will be open source at \url{https://voxdialogue.github.io/}. Xize Cheng, Ruofan Hu 0002, Xiaoda Yang, Jingyu Lu 0001, Zehan Wang 0001, Shengpeng Ji, Rongjie Huang 0001, Tao Jin 0004, Zhou Zhao 0001 |
ICLR | 1 |
| 2025 | OmniSep: Unified Omni-Modality Sound Separation with Query-MixupabstractQuery-based sound separation (QSS) effectively isolate sound signals that match the content of a given query, enhancing the understanding of audio data. However, most existing QSS methods rely on a single modality for separation, lacking the ability to fully leverage homologous but heterogeneous information across multiple modalities for the same sound signal. To address this limitation, we introduce Omni-modal Sound Separation (**OmniSep**), a novel framework capable of isolating clean soundtracks based on omni-modal queries, encompassing both single-modal and multi-modal composed queries. Specifically, we introduce the **Query-Mixup** strategy, which blends query features from different modalities during training. This enables OmniSep to optimize multiple modalities concurrently, effectively bringing all modalities under a unified framework for sound separation. We further enhance this flexibility by allowing queries to influence sound separation positively or negatively, facilitating the retention or removal of specific sounds as desired. Finally, OmniSep employs a retrieval-augmented approach known as **Query-Aug**, which enables open-vocabulary sound separation. Experimental evaluations on MUSIC, VGGSOUND-CLEAN+, and MUSIC-CLEAN+ datasets demonstrate effectiveness of OmniSep, achieving state-of-the-art performance in text-, image-, and audio-queried sound separation tasks. For samples and further information, please visit the demo page at \url{https://omnisep.github.io/}. Xize Cheng, Zehan Wang 0001, Minghui Fang 0002, Rongjie Huang 0001, Shengpeng Ji, Jialong Zuo, Tao Jin 0004, Zhou Zhao 0001 |
ICLR | 1 |
| 2025 | WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingabstractLanguage models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1) extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2) improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The code is available at https://github.com/jishengpeng/WavTokenizer. Shengpeng Ji, Ziyue Jiang 0001, Wen Wang 0001, Minghui Fang 0002, Jialong Zuo, Qian Yang 0006, Xize Cheng, Zehan Wang 0001, Ruiqi Li 0002, Xiaoda Yang, Rongjie Huang 0001, Yidi Jiang, Qian Chen 0003, Zhou Zhao 0001 |
ICLR | 8 |
| 2025 | GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer
Minghui Fang 0002, Shengpeng Ji, Jialong Zuo, Xize Cheng, Wenrui Liu 0003, Xiaoda Yang, Ruofan Hu 0002, Jieming Zhu, Zhou Zhao 0001 |
INTERSPEECH | 4 |
| 2025 | Multimodal Conditional Retrieval with High ControllabilityabstractSearching for images using text has limitations because language has difficulties in expressing certain abstract intentions, e.g. artistic styles are difficult to describe for non-experts. As for the image search image model, images can convey abstract intentions, but cannot express the specific purpose, so many of the current graph search works only have a single function, such as content search and style search. Our work aims to combine the strengths of both, merging the ability of text to express specific ideas with the ability of images to convey abstract concepts, thus achieving a better capture of the user's intentions. To this end, we propose CCSR, a multimodal conditional content-style joint retrieval model. Our model is the first to apply contrastive learning to conditional retrieval and introduces a novel Mixture-of-Expert models (MOE) system to enable collaboration between multiple expert systems. We adopt a novel prompt learning strategy that allows the model to adaptively select specific prompts, thereby enhancing its focus on the current task. In addition, to evaluate the joint content-style retrieval capability of our model, we present a new dataset, StyleCoco, containing rich content categories and style categories. The experimental results indicate that CCSR has achieved state-of-the-art performance in conditional style retrieval, content retrieval, and style-content retrieval. The dataset and code will be publicly available on https://mccsr.github.io/. Xiaoda Yang, Xize Cheng, Minghui Fang 0002, Hongshun Qiu, Jiaqi Duan, Sihang Cai, Zehan Wang 0001, Ruofan Hu 0002, Zhou Zhao 0001, Tao Jin 0004 |
KDD (2) | 2 |
| 2025 | AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language ModelsabstractHallucinations present a significant challenge in the development and evaluation of large language models (LLMs), directly affecting their reliability and accuracy. While notable advancements have been made in research on textual and visual hallucinations, there is still a lack of a comprehensive benchmark for evaluating auditory hallucinations in large audio language models (LALMs). To fill this gap, we introduce AHa-Bench, a systematic and comprehensive benchmark for audio hallucinations. Audio data, in particular, uniquely combines the multi-attribute complexity of visual data with the semantic richness of textual data, leading to auditory hallucinations that share characteristics with both visual and textual hallucinations. Based on the source of these hallucinations, AHa-Bench categorizes them into semantic hallucinations, acoustic hallucinations, and semantic-acoustic confusion hallucinations. In addition, we systematically evaluate seven open-source local perception language models (LALMs), demonstrating the challenges these models face in audio understanding, especially when it comes to jointly understanding semantic and acoustic information. Through the development of a comprehensive evaluation framework, AHa-Bench aims to enhance the robustness and stability of LALMs, fostering more reliable and nuanced audio understanding in LALMs. The benchmark dataset is available at \url{https://huggingface.co/datasets/ahabench/AHa-Bench}. Xize Cheng, Chenyuhao Wen, Shannon Yu, Zehan Wang 0001, Shengpeng Ji, Siddhant Arora, Tao Jin 0004, Shinji Watanabe 0001, Zhou Zhao 0001 |
NeurIPS | 1 |
| 2025 | Multi-talker audio-visual speech recognition towards diverse scenariosabstractRecently, audio–visual speech recognition (AVSR) has attracted increasing attention. However, most existing works simplify the complex challenges in real-world applications and only focus on scenarios with two speakers and perfectly aligned audio-video clips. In this work, we study the effect of speaker number and modal misalignment in the AVSR task, and propose an end-to-end AVSR framework under a more realistic condition. Specifically, we propose a speaker-number-aware mixture-of-experts (SA-MoE) mechanism to explicitly model the characteristic difference in scenarios with different speaker numbers, and a cross-modal realignment (CMR) module for robust handling of asynchronous inputs. We also use the underlying difficulty difference and introduce a new training strategy named challenge-based curriculum learning (CBCL), which forces the model to focus on difficult, challenging data instead of simple data to improve efficiency. Yuxiao Lin, Tao Jin 0004, Xize Cheng, Zhou Zhao 0001, Fei Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2024 | Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and AccompanimentabstractZhiqing Hong, Rongjie Huang, Xize Cheng, Yongqi Wang, Ruiqi Li, Fuming You, Zhou Zhao, Zhimeng Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhiqing Hong, Rongjie Huang 0001, Xize Cheng, Ruiqi Li 0002, Fuming You, Zhou Zhao 0001 |
ACL (1) | 3 |
| 2024 | Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results AlignmentabstractTransformer-based methods have gone mainstream in multimodal sequential learning.The intra and inter modality interactions are captured by the query-key associations of multihead attention.In this way, the calculated multimodal contexts (attentional results) are expected to be relevant to the query modality.However, in existing literature, the alignment degree between different calculated attentional results of the same query are under-explored.Based on this concern, we propose a new constrained scheme called Multimodal Contextual Contrast (MCC), which could align the multiple attentional results from both local and global perspectives, making the information capture more efficient.Concretely, the calculated attentional results of different modalities are mapped into a common feature space, those attentional vectors with the same query are considered as a positive group and the remaining sets are negative.From local perspective, we sample the negative groups for a positive group by randomly changing the sequential step of one specific context and keeping the other stay the same.From coarse global perspective, we divide all the contextual groups into two sets (i.e., aligned and unaligned), making the total score of aligned group relatively large.We extend the vectorial inner product operation for more input and calculate the aligned score for each multimodal group.Considering that the computational complexity scales exponentially to the number of modalities, we adopt stochastic expectation approximation (SEA) for the real process.The extensive experimental results on several tasks reveal the effectiveness of our contributions. Tao Jin 0004, Ye Wang 0018, Linjun Li, Xize Cheng, Zhou Zhao 0001 |
ACL (1) | 5 |
| 2024 | Uni-Dubbing: Zero-Shot Speech Synthesis from Visual ArticulationabstractSongju Lei, Xize Cheng, Mengjiao Lyu, Jianqiao Hu, Jintao Tan, Runlin Liu, Lingyu Xiong, Tao Jin, Xiandong Li, Zhou Zhao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Songju Lei, Xize Cheng, Mengjiao Lyu, Jianqiao Hu, Jintao Tan, Runlin Liu, Lingyu Xiong, Tao Jin 0004, Xiandong Li, Zhou Zhao 0001 |
ACL (1) | 2 |
| 2024 | AudioVSR: Enhancing Video Speech Recognition with Audio DataabstractXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu, Minjie Hong, Minghui Fang, Shengpeng Ji, Jialong Zuo, Zhiqing Hong, Zhimeng Zhang, Tao Jin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu, Minjie Hong, Minghui Fang 0002, Shengpeng Ji, Jialong Zuo, Zhiqing Hong, Tao Jin 0004 |
EMNLP | 2 |
| 2024 | Landmark-guided Diffusion Model for High-fidelity and Temporally Coherent Talking Head GenerationabstractAudio-driven talking head generation is a significant and challenging task applicable to various fields such as virtual avatars, film production, and online conferences. However, the existing GAN-based models emphasize generating well-synchronized lip shapes but overlook the visual quality of generated frames, while diffusion-based models prioritize generating high-quality frames but neglect lip shape matching, resulting in jittery mouth movements. To address the aforementioned problems, we introduce a two-stage diffusion-based model. The first stage involves generating synchronized facial landmarks based on the given speech. In the second stage, these generated landmarks serve as a condition in the denoising process, aiming to optimize mouth jitter issues and generate high-fidelity, well-synchronized, and temporally coherent talking head videos. Extensive experiments demonstrate that our model yields the best performance. Jintao Tan, Xize Cheng, Lingyu Xiong, Xiandong Li, Xianjia Wu |
ICME | 2 |
| 2024 | FreeBind: Free Lunch in Unified Multimodal Space via Knowledge FusionabstractUnified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained unified spaces. In this work, we propose FreeBind, an idea that treats multimodal representation spaces as basic units, and freely augments pre-trained unified space by integrating knowledge from extra expert spaces via “space bonds". Specifically, we introduce two kinds of basic space bonds: 1) Space Displacement Bond and 2) Space Combination Bond. Based on these basic bonds, we design Complex Sequential & Parallel Bonds to effectively integrate multiple spaces simultaneously. Benefiting from the modularization concept, we further propose a coarse-to-fine customized inference strategy to flexibly adjust the enhanced unified space for different purposes. Experimentally, we bind ImageBind with extra image-text and audio-text expert spaces, resulting in three main variants: ImageBind++, InternVL_IB, and InternVL_IB++. These resulting spaces outperform ImageBind on 5 audio-image-text downstream tasks across 9 datasets. Moreover, via customized inference, it even surpasses the advanced audio-text and image-text expert spaces. Our code and checkpoints are released at https://github.com/zehanwang01/FreeBind Zehan Wang 0001, Xize Cheng, Rongjie Huang 0001, Luping Liu, Zhenhui Ye, Haifeng Huang 0001, Yang Zhao 0022, Tao Jin 0004, Peng Gao 0007, Zhou Zhao 0001 |
ICML | 3 |
| 2024 | InstructSpeech: Following Speech Editing Instructions via Large Language ModelsabstractInstruction-guided speech editing aims to follow the user’s natural language instruction to manipulate the semantic and acoustic attributes of a speech. In this work, we construct triplet paired data (instruction, input speech, output speech) to alleviate data scarcity and train a multi-task large language model named InstructSpeech. To mitigate the challenges of accurately executing user’s instructions, we 1) introduce the learned task embeddings with a fine-tuned Flan-T5-XL to guide the generation process towards the correct generative task; 2) include an extensive and diverse set of speech editing and processing tasks to enhance model capabilities; 3) investigate chain-of-thought reasoning for free-form semantic content editing; and 4) propose a hierarchical adapter that effectively updates a small portion of parameters for generalization to new tasks. To assess instruction speech editing in greater depth, we introduce a benchmark evaluation with contrastive instruction-speech pre-training (CISP) to test the speech quality and instruction-speech alignment faithfulness. Experimental results demonstrate that InstructSpeech achieves state-of-the-art results in eleven tasks, for the first time unlocking the ability to edit speech’s acoustic and semantic attributes following a user’s instruction. Audio samples are available at https://InstructSpeech.github.io Rongjie Huang 0001, Ruofan Hu 0002, Zehan Wang 0001, Xize Cheng, Ziyue Jiang 0001, Zhenhui Ye, Dongchao Yang, Luping Liu, Peng Gao 0007, Zhou Zhao 0001 |
ICML | 5 |
| 2024 | Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented PromptsabstractIn the burgeoning field of Audio-Visual Speech Recognition (AVSR), extant research has predominantly concentrated on the training paradigms tailored for high-quality resources. However, owing to the challenges inherent in real-world data collection, audio-visual data are frequently affected by modality-distortion, which encompasses audio-visual asynchrony, video noise and audio noise. The recognition accuracy of existing AVSR method is significantly compromised when multiple modality-distortion coexist in low-resource data. In light of the above challenges, we propose PCD: cluster-Prompt with Contrastive Decomposition, a robust framework for modality-distortion speech recognition, specifically devised to transpose the pre-trained knowledge from high-resource domain to the targeted domain by leveraging contrast-augmented prompts. In contrast to previous studies, we take into consideration the possibility of various types of distortion in both the audio and visual modalities. Concretely, we design bespoke prompts to delineate each modality-distortion, guiding the model to achieve speech recognition applicable to various distortion scenarios with quite few learnable parameters. To materialize the prompt mechanism, we employ multiple cluster-based strategies that better suits the pre-trained audio-visual model. Additionally, we design a contrastive decomposition mechanism to restrict the explicit relationships among various modality conditions, given their shared task knowledge and disparate modality priors. Extensive results on LRS2 dataset demonstrate that PCD achieves state-of-the-art performance for audio-visual speech recognition under the constraints of distorted resources. Code is available at https://github.com/ballooncatt/PCD. Xize Cheng, Xiaoda Yang, Hanting Wang, Zhou Zhao 0001, Tao Jin 0004 |
ACM Multimedia | 2 |
| 2024 | VoiceTuner: Self-Supervised Pre-training and Efficient Fine-tuning For Voice GenerationabstractVoice large language models (LLMs) cast voice synthesis as a language modeling task in a discrete space, and have demonstrated significant progress to date. Despite the recent success, the current development of voice LLMs in low-resource applications is hampered by data scarcity and high computational cost. In this work, we propose VoiceTuner, with a self-supervised pre-training and efficient fine-tuning approach for low-resource voice generation. Specifically, 1) to mitigate data scarcity, we leverage large-scale unlabeled dataset and pre-train VoiceTuner-SSL without pre-defined applications, which can be fine-tuned in downstream tasks; 2) to further reduce the high training cost in complete fine-tuning, we introduce a multiscale transformer adapter to effectively update only around 1% parameters as a plug-and-play module. Experimental results demonstrate that VoiceTuner-SSL presents strong acoustic continuations, and VoiceTuner achieves state-of-the-art results in rich-resource TTS evaluation compared with competitive baseline models. Low-resource (1h, 10h, 30h) downstream applications including zero-shot TTS, instruction TTS, and singing voice synthesis present VoiceTuner's superior audio quality and style similarity with reduced data requirement and computational cost. Audio samples are available at https://VoiceTuner.github.io Rongjie Huang 0001, Ruofan Hu 0002, Xiaoshan Xu, Zhiqing Hong, Dongchao Yang, Xize Cheng, Zehan Wang 0001, Ziyue Jiang 0001, Zhenhui Ye, Luping Liu, Zhou Zhao 0001 |
ACM Multimedia | 7 |
| 2024 | AudioLCM: Efficient and High-Quality Text-to-Audio Generation with Minimal Inference Steps
Huadai Liu, Rongjie Huang 0001, Yang Liu 0278, Hengyuan Cao, Xize Cheng, Zhou Zhao 0001 |
ACM Multimedia | 6 |
| 2024 | SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local EditingabstractAudio-driven talking face generation aims to synthesize video with lip movements synchronized to input audio. However, current generative techniques face challenges in preserving intricate regional textures (skin, teeth). To address the aforementioned challenges, we propose a novel framework called SegTalker to decouple lip movements and image textures by introducing segmentation as intermediate representation. Specifically, given the mask of image employed by a parsing network, we first leverage the speech to drive the mask and generate talking segmentation. Then we disentangle semantic regions of image into style codes using a mask-guided encoder. Ultimately, we inject the previously generated talking segmentation and style codes into a mask-guided StyleGAN to synthesize video frame. In this way, most of textures are fully preserved. Moreover, our approach can inherently achieve background separation and facilitate mask-guided facial local editing. In particular, by editing the mask and swapping the region textures from a given reference image (e.g. hair, lip, eyebrows), our approach enables facial editing seamlessly when generating talking face video. Experiments demonstrate that our proposed approach can effectively preserve texture details and generate temporally consistent video while remaining competitive in lip synchronization. Quantitative and qualitative results on the HDTF and MEAD datasets illustrate the superior performance of our method over existing methods. Lingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu, Xiandong Li, Fei Ma 0006, Minglei Li 0001, Huang Xu 0003 |
ACM Multimedia | 2 |
| 2024 | SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task LearningabstractTalking Face Generation (TFG) reconstructs facial motions concerning lips given speech input, which aims to generate highquality, synchronized, and lip-readable videos. Previous efforts have achieved success in generating quality and synchronization, and recently, there has been an increasing focus on the importance of intelligibility. Despite these efforts, there remains a challenge in achieving a balance among quality, synchronization, and intelligibility, often resulting in trade-offs that compromise one aspect in favor of another. In light of this, we propose SyncTalklip, a novel dual-tower framework designed to overcome the challenges of synchronization while improving lip-reading performance. To enhance the performance of SyncTalklip in both synchronization and intelligibility, we design AV-SyncNet, a pre-trained multi-task model, aiming to achieve a dual-focus on synchronization and intelligibility. Moreover, we propose a novel cross-modal contrastive learning bringing audio and video closer to enhance synchronization. Experimental results demonstrate that SyncTalklip achieves state-of-the-art performance in quality, intelligibility, and synchronization. Furthermore, extensive experiments have demonstrated our model's generalizability across domains. The code and demo is available at https://sync-talklip.github.io. Xiaoda Yang, Xize Cheng, Minghui Fang 0002, Jialong Zuo, Shengpeng Ji, Zhou Zhao 0001, Tao Jin 0004 |
ACM Multimedia | 2 |
| 2024 | Chat-Scene: Bridging 3D Scene and Large Language Models with Object IdentifiersabstractRecent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in general referencing and grounding capabilities for intricate scene comprehension. In this paper, we introduce the use of object identifiers and object-centric representations to interact with scenes at the object level. Specifically, we decompose the input 3D scene into a set of object proposals, each assigned a unique identifier token, which enables efficient object referencing and grounding during user-assistant interactions. Given the scarcity of scene-language data, we model the scene embeddings as a sequence of explicit object-level embeddings, derived from semantic-rich 2D or 3D representations. By employing object identifiers, we transform diverse 3D scene-language tasks into a unified question-answering format, facilitating joint training without the need for additional task-specific heads. With minimal fine-tuning on all downstream tasks, our model significantly outperforms existing methods on benchmarks including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D. Haifeng Huang 0001, Zehan Wang 0001, Rongjie Huang 0001, Runsen Xu, Luping Liu, Xize Cheng, Yang Zhao 0022, Jiangmiao Pang, Zhou Zhao 0001 |
NeurIPS | 8 |
| 2024 | MimicTalk: Mimicking a personalized and expressive 3D talking face in minutesabstractTalking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically solve this problem by learning an individual neural radiance field (NeRF) for each identity to implicitly store its static and dynamic information, we find it inefficient and non-generalized due to the per-identity-per-training framework and the limited training data. To this end, we propose MimicTalk, the first attempt that exploits the rich knowledge from a NeRF-based person-agnostic generic model for improving the efficiency and robustness of personalized TFG. To be specific, (1) we first come up with a person-agnostic 3D TFG model as the base model and propose to adapt it into a specific identity; (2) we propose a static-dynamic-hybrid adaptation pipeline to help the model learn the personalized static appearance and facial dynamic features; (3) To generate the facial motion of the personalized talking style, we propose an in-context stylized audio-to-motion model that mimics the implicit talking style provided in the reference video without information loss by an explicit style representation. The adaptation process to an unseen identity can be performed in 15 minutes, which is 47 times faster than previous person-dependent methods. Experiments show that our MimicTalk surpasses previous baselines regarding video quality, efficiency, and expressiveness. Video samples are available at https://mimictalk.github.io . Zhenhui Ye, Tianyun Zhong, Yi Ren 0006, Ziyue Jiang 0001, Jiawei Huang 0008, Rongjie Huang 0001, Jinglin Liu, Jinzheng He, Chen Zhang 0020, Zehan Wang 0001, Xize Cheng, Xiang Yin 0006, Zhou Zhao 0001 |
NeurIPS | 11 |
| 2024 | Extending Multi-modal Contrastive RepresentationsabstractMulti-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Inspired by recent C-MCR, this paper proposes $\textbf{Ex}$tending $\textbf{M}$ultimodal $\textbf{C}$ontrastive $\textbf{R}$epresentation (Ex-MCR), a training-efficient and paired-data-free method to build unified contrastive representation for many modalities. Since C-MCR is designed to learn a new latent space for the two non-overlapping modalities and projects them onto this space, a significant amount of information from their original spaces is lost in the projection process. To address this issue, Ex-MCR proposes to extend one modality's space into the other's, rather than mapping both modalities onto a completely new space. This method effectively preserves semantic alignment in the original space. Experimentally, we extend pre-trained audio-text and 3D-image representations to the existing vision-text space. Without using paired data, Ex-MCR achieves comparable performance to advanced methods on a series of audio-image-text and 3D-image-text tasks and achieves superior performance when used in parallel with data-driven methods. Moreover, semantic alignment also emerges between the extended modalities (e.g., audio and 3D). Zehan Wang 0001, Luping Liu, Rongjie Huang 0001, Xize Cheng, Zhenhui Ye, Huadai Liu, Haifeng Huang 0001, Yang Zhao 0022, Tao Jin 0004, Zhou Zhao 0001 |
NeurIPS | 5 |
| 2023 | OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality AlignmentabstractSpeech Recognition builds a bridge between the multimedia streaming (audio-only, visualonly or audio-visual) and the corresponding text transcription.However, when training the specific model of new domain, it often gets stuck in the lack of new-domain utterances, especially the labeled visual utterances.To break through this restriction, we attempt to achieve zero-shot modality transfer by maintaining the multi-modality alignment in phoneme space learned with unlabeled multimedia utterances in the high resource domain during the pretraining (Shi et al., 2022), and propose a training system Open-modality Speech Recognition (OpenSR) that enables the models trained on a single modality (e.g., audio-only) applicable to more modalities (e.g., visual-only and audio-visual).Furthermore, we employ a cluster-based prompt tuning strategy to handle the domain shift for the scenarios with only common words in the new domain utterances.We demonstrate that OpenSR enables modality transfer from one to any in three different settings (zero-, few-and fullshot), and achieves highly competitive zeroshot performance compared to the existing fewshot and full-shot lip-reading methods.To the best of our knowledge, OpenSR achieves the state-of-the-art performance of word error rate in LRS2 on audio-visual speech recognition and lip-reading with 2.7% and 25.0%, respectively.The code and demo are available at https://github.com/Exgc/OpenSR. Xize Cheng, Tao Jin 0004, Linjun Li, Xinyu Duan, Zhou Zhao 0001 |
ACL (1) | 1 |
| 2023 | AV-TranSpeech: Audio-Visual Robust Speech-to-Speech TranslationabstractRongjie Huang, Huadai Liu, Xize Cheng, Yi Ren, Linjun Li, Zhenhui Ye, Jinzheng He, Lichao Zhang, Jinglin Liu, Xiang Yin, Zhou Zhao. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Rongjie Huang 0001, Huadai Liu, Xize Cheng, Yi Ren 0006, Linjun Li, Zhenhui Ye, Jinzheng He, Jinglin Liu, Xiang Yin 0006, Zhou Zhao 0001 |
ACL (1) | 3 |
| 2023 | TAVT: Towards Transferable Audio-Visual Text GenerationabstractAudio-visual text generation aims to understand multi-modality contents and translate them into texts.Although various transfer learning techniques of text generation have been proposed, they focused on uni-modal analysis (e.g., text-to-text, visual-to-text) and lack consideration of multi-modal content and cross-modal relation.Motivated by the fact that humans can recognize the timbre of the same low-level concepts (e.g., footstep, rainfall, and laughing), even in different visual conditions, we aim to mitigate the domain discrepancies by audiovisual correlation.In this paper, we propose a novel Transferable Audio-Visual Text Generation framework, named TAVT, which consists of two key components: Audio-Visual Meta-Mapper (AVMM) and Dual Counterfactual Contrastive Learning (DCCL).(1) AVMM first introduces a universal auditory semantic space and drifts the domain-invariant low-level concepts into visual prefixes.Then the reconstructbased learning encourages the AVMM to learn "which pixels belong to the same sound" and achieve audio-enhanced visual prefix.The welltrained AVMM can be further applied to unimodal setting.(2) Furthermore, DCCL leverages the destructive counterfactual transformations to provide cross-modal constraints for AVMM from the perspective of feature distribution and text generation.(3) The experimental results show that TAVT outperforms the stateof-the-art methods across multiple domains (cross-datasets, cross-categories) and various modal settings (uni-modal, multi-modal). Tao Jin 0004, Wenwen Pan 0003, Linjun Li, Xize Cheng, Ye Wang 0018, Zhou Zhao 0001 |
ACL (1) | 5 |
| 2023 | Weakly-Supervised Spoken Video Grounding via Semantic Interaction LearningabstractThe task of spoken video grounding aims to localize moments in videos that are relevant to descriptive spoken queries.However, extracting semantic information from speech and modeling the cross-modal correlation pose two critical challenges.Previous studies solve them by representing spoken queries based on the matched video frames, which require tremendous effort for frame-level labeling.In this work, we investigate weakly-supervised spoken video grounding, i.e., learning to localize moments without expensive temporal annotations.To effectively represent the cross-modal semantics, we propose Semantic Interaction Learning (SIL), a novel framework consisting of the acoustic-semantic pre-training (ASP) and acoustic-visual contrastive learning (AVCL).In ASP, we pre-train an effective encoder for the grounding task with three comprehensive tasks, where the robustness task enhances stability by explicitly capturing the invariance between time-and frequency-domain features, the conciseness task avoids over-smooth attention by compressing long sequence into segments, and the semantic task improves spoken language understanding by modeling the precise semantics.In AVCL, we mine pseudo labels with discriminative sampling strategies and directly strengthen the interaction between speech and video by maximizing their mutual information.Extensive experiments demonstrate the effectiveness and superiority of our method.1 Ye Wang 0018, Shengyu Zhang 0001, Tao Jin 0004, Linjun Li, Xize Cheng, Zhou Zhao 0001 |
ACL (1) | 6 |
| 2023 | 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Groundingabstract3D visual grounding aims to localize the target object in a 3D point cloud by a free-form language description.Typically, the sentences describing the target object tend to provide information about its relative relation between other objects and its position within the whole scene.In this work, we propose a relation-aware onestage framework, named 3D Relative Positionaware Network (3DRP-Net), which can effectively capture the relative spatial relationships between objects and enhance object attributes.Specifically, 1) we propose a 3D Relative Position Multi-head Attention (3DRP-MA) module to analyze relative relations from different directions in the context of object pairs, which helps the model to focus on the specific object relations mentioned in the sentence.2) We designed a soft-labeling strategy to alleviate the spatial ambiguity caused by redundant points, which further stabilizes and enhances the learning process through a constant and discriminative distribution.Extensive experiments conducted on three benchmarks (i.e., ScanRefer and Nr3D/Sr3D) demonstrate that our method outperforms all the state-of-the-art methods in general. Zehan Wang 0001, Haifeng Huang 0001, Yang Zhao 0022, Linjun Li, Xize Cheng, Aoxiong Yin, Zhou Zhao 0001 |
EMNLP | 5 |
| 2023 | MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and RecognitionabstractMulti-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on visual speech. This lack of research is mainly due to the absence of datasets containing visual speech and translated text pairs. In this paper, we present AVMuST-TED, the first dataset for Audio-Visual Multilingual Speech Translation, derived from TED talks. Nonetheless, visual speech is not as distinguishable as audio speech, making it difficult to develop a mapping from source speech phonemes to the target language text. To address this issue, we propose MixSpeech, a cross-modality self-learning framework that utilizes audio speech to regularize the training of visual speech tasks. To further minimize the cross-modality gap and its impact on knowledge transfer, we suggest adopting mixed speech, which is created by interpolating audio and visual streams, along with a curriculum learning strategy to adjust the mixing ratio as needed. MixSpeech enhances speech translation in noisy environments, improving BLEU scores for four languages on AVMuST-TED by +1.4 to +4.2. Moreover, it achieves state-of-the-art performance in lip reading on CMLR (11.1%), LRS2 (25.5%), and LRS3 (28.0%). Xize Cheng, Tao Jin 0004, Rongjie Huang 0001, Linjun Li, Zehan Wang 0001, Ye Wang 0018, Huadai Liu, Aoxiong Yin, Zhou Zhao 0001 |
ICCV | 1 |
| 2023 | Exploring Group Video Captioning with Efficient Relational ApproximationabstractCurrent video captioning efforts most focus on describing a single video while the need for captioning videos in groups has increased considerably. In this study, we propose a new task, group video captioning, which aims to infer the desired content among a group of target videos and describe it with another group of related reference videos. This task requires the model to effectively summarize the target videos and accurately describe the distinguishing content compared to the reference videos, and it becomes more difficult as the video length increases. To solve this problem, 1) First, we propose an efficient relational approximation (ERA) to identify the shared content among videos while the complexity is linearly related to the number of videos. 2) Then, we introduce a contextual feature refinery with intra-group self-supervision to capture the contextual information and further refine the common properties. 3) In addition, we construct two group video captioning datasets derived from the YouCook2 and the ActivityNet Captions. The experimental results demonstrate the effectiveness of our method on this new task. Tao Jin 0004, Ye Wang 0018, Wenwen Pan 0003, Linjun Li, Xize Cheng, Zhou Zhao 0001 |
ICCV | 6 |
| 2023 | Distilling Coarse-to-Fine Semantic Matching Knowledge for Weakly Supervised 3D Visual Groundingabstract3D visual grounding involves finding a target object in a 3D scene that corresponds to a given sentence query. Although many approaches have been proposed and achieved impressive performance, they all require dense object-sentence pair annotations in 3D point clouds, which are both time-consuming and expensive. To address the problem that fine-grained annotated data is difficult to obtain, we propose to leverage weakly supervised annotations to learn the 3D visual grounding model, i.e., only coarse scene-sentence correspondences are used to learn object-sentence links. To accomplish this, we design a novel semantic matching model that analyzes the semantic similarity between object proposals and sentences in a coarse-to-fine manner. Specifically, we first extract object proposals and coarsely select the top-K candidates based on feature and class similarity matrices. Next, we reconstruct the masked keywords of the sentence using each candidate one by one, and the reconstructed accuracy finely reflects the semantic similarity of each candidate to the query. Additionally, we distill the coarse-to-fine semantic matching knowledge into a typical two-stage 3D visual grounding model, which reduces inference costs and improves performance by taking full advantage of the well-studied structure of the existing architectures. We conduct extensive experiments on ScanRefer, Nr3D, and Sr3D, which demonstrate the effectiveness of our proposed method. Zehan Wang 0001, Haifeng Huang 0001, Yang Zhao 0022, Linjun Li, Xize Cheng, Aoxiong Yin, Zhou Zhao 0001 |
ICCV | 5 |
| 2023 | Rethinking Missing Modality Learning from a Decoding PerspectiveabstractConventional pipeline of multimodal learning consists of three stages, including encoding, fusion, and decoding. Most existing methods under missing modality condition focus on the first stage and aim to learn the modality invariant representation or reconstruct missing features. However, these methods rely on strong assumptions (i.e., all the pre-defined modalities are available for each input sample during training and the number of modalities is fixed). To solve this problem, we propose a simple yet effective method called Interaction Augmented Prototype Decomposition (IPD) for a more general setting, where the number of modalities is arbitrary and there are various incomplete modality conditions happening in both training and inference phases, even there are unseen testing conditions. Different from the previous methods, we improve the decoding stage. Concretely, IPD jointly learns the common and modality-specific task prototypes. Considering that the number of missing modality conditions scales exponentially with the number of modalities O(2n) and different conditions may have implicit interaction, the low-rank partial prototype decomposition with enough theoretical analysis is employed for modality-specific components to reduce the complexity. The decomposition also can promote unseen generalization with the modality factors of existing conditions. To simulate the low-rank setup, we further constrain the explicit interaction of specific modality conditions by employing disentangled contrastive constraints. Extensive results on the newly-created benchmarks of multiple tasks illustrate the effectiveness of our proposed model. Tao Jin 0004, Xize Cheng, Linjun Li, Ye Wang 0018, Zhou Zhao 0001 |
ACM Multimedia | 2 |
| 2023 | Connecting Multi-modal Contrastive RepresentationsabstractMulti-modal Contrastive Representation (MCR) learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across various modalities. However, the reliance on massive high-quality data pairs limits its further development on more modalities. This paper proposes a novel training-efficient method for learning MCR without paired data called Connecting Multi-modal Contrastive Representations (C-MCR). Specifically, given two existing MCRs pre-trained on $(\mathcal{A}$, $\mathcal{B})$ and $(\mathcal{B}$, $\mathcal{C})$ modality pairs, we project them to a new space and use the data from the overlapping modality $\mathcal{B}$ to aligning the two MCRs in the new space. Meanwhile, since the modality pairs $(\mathcal{A}$, $\mathcal{B})$ and $(\mathcal{B}$, $\mathcal{C})$ are already aligned within each MCR, the connection learned by overlapping modality can also be transferred to non-overlapping modality pair $(\mathcal{A}$, $\mathcal{C})$. To unleash the potential of C-MCR, we further introduce a semantic-enhanced inter- and intra-MCR connection method. We first enhance the semantic consistency and completion of embeddings across different modalities for more robust alignment. Then we utilize the inter-MCR alignment to establish the connection, and employ the intra-MCR alignment to better maintain the connection for inputs from non-overlapping modalities. To demonstrate the effectiveness of C-MCR, we take the field of audio-visual and 3D-language learning as examples. Specifically, we connect CLIP and CLAP via texts to derive audio-visual representations, and integrate CLIP and ULIP via images for 3D-language representations. Remarkably, without using any paired data, C-MCR for audio-visual achieves state-of-the-art performance on audio-image retrieval, audio-visual source localization, and counterfactual audio-image recognition tasks. Furthermore, C-MCR for 3D-language also attains advanced zero-shot 3D point cloud classification accuracy on ModelNet40. Our project page is available at \url{https://c-mcr.github.io/C-MCR/} Zehan Wang 0001, Yang Zhao 0022, Xize Cheng, Haifeng Huang 0001, Jiageng Liu, Aoxiong Yin, Linjun Li, Zhou Zhao 0001 |
NeurIPS | 3 |