Yifei Xin

dblp:271/3101 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021
YearPublicationVenuePosition
2025 ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
abstract
The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets for embodied video question answering lack comprehensive and systematic evaluation frameworks. Critical embodied cognitive issues, such as robotic self-cognition, dynamic scene perception, and hallucination, are rarely addressed. To tackle these challenges, we propose ECBench, a high-quality benchmark designed to systematically evaluate the embodied cognitive abilities of LVLMs. ECBench features a diverse range of scene video sources, open and varied question formats, and 30 dimensions of embodied cognition. To ensure quality, balance, and high visual dependence, ECBench uses class-independent meticulous human annotation and multi-round question screening strategies. Additionally, we introduce ECEval, a comprehensive evaluation system that ensures the fairness and rationality of the indicators. Utilizing ECBench, we conduct extensive evaluations of proprietary, open-source, and task-specific LVLMs. ECBench is pivotal in advancing the embodied cognitive capabilities of LVLMs, laying a solid foundation for developing reliable core models for embodied agents. All data and code is available at https://github.com/RhDang/ECBench.
Ronghao Dang, Yuqian Yuan, Wenqi Zhang 0001, Yifei Xin, Boqiang Zhang, Liuyi Wang, Qinyang Zeng, Xin Li 0056, Lidong Bing
CVPR4
2025 Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
abstract
Vector Quantization (VQ) is essential for discretizing continuous representations in unsupervised learning but suffers from representation collapse, causing low codebook utilization and limiting scalability. Existing solutions often rely on complex optimizations or reduce latent dimensionality, which compromises model capacity and fails to fully solve the problem. We identify the root cause as disjoint codebook optimization, where only a few code vectors are updated via gradient descent. To fix this, we propose \textbf{Sim}ple\textbf{VQ}, which reparameterizes code vectors through a learnable linear transformation layer over a latent basis, optimizing the \textit{entire linear space} rather than nearest \textit{individual code vectors}. Although the multiplication of two linear matrices is equivalent to applying a single linear layer, this simple approach effectively prevents collapse. Extensive experiments on image and audio tasks demonstrate that SimVQ improves codebook usage, is easy to implement, and generalizes well across modalities and architectures. The code is available at https://github.com/youngsheen/SimVQ.
Yongxin Zhu 0003, Bocheng Li, Yifei Xin, Zhihua Xia, Linli Xu 0002
ICCV3
2024 Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup
abstract
Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information, which has received widespread attention recently.It has been verified that visual information brings greater performance gains when the textual information is limited.However, most previous works ignore to take advantage of the complete textual inputs and the limited textual inputs at the same time, which limits the overall performance.To solve this issue, we propose a mixup method termed Soul-Mix to enhance MMT by using visual information more effectively.We mix the predicted translations of complete textual input and the limited textual inputs.Experimental results on the Multi30K dataset of three translation directions show that our Soul-Mix significantly outperforms existing approaches and achieves new state-of-the-art performance with fewer parameters than some previous models.Besides, the strength of Soul-Mix is more obvious on more challenging MSCOCO dataset which includes more out-of-domain instances with lots of ambiguous verbs.
Xuxin Cheng, Ziyu Yao 0001, Yifei Xin, Hongxiang Li 0004, Yaowei Li 0001, Yuexian Zou
ACL (1)3
2024 DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
Yifei Xin, Xuxin Cheng, Zhihong Zhu 0001, Xusheng Yang, Yuexian Zou
INTERSPEECH1
2024 Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
Yifei Xin, Zhihong Zhu 0001, Xuxin Cheng, Xusheng Yang, Yuexian Zou
INTERSPEECH1
2024 MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
Yifei Xin, Zhesong Yu, Bilei Zhu, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH2
2023 Improving Speech Enhancement via Event-Based Query
abstract
Existing deep learning based speech enhancement (SE) methods either use blind end-to-end training or explicitly incorporate speaker embedding or phonetic information into the SE network to enhance speech quality. In this paper, we perceive speech and noises as different types of sound events and propose an event-based query method for SE. Specifically, speech embeddings that can discriminate speech from noises are first pre-trained with the sound event detection (SED) task. The embeddings are then clustered into fixed golden speech queries, i.e., general but representative speech embeddings, on a diverse clean speech dataset to assist the SE network. The golden speech queries can be obtained offline and generalizable to different SE datasets and networks. Therefore, little extra complexity is introduced and no enrollment is needed for each speaker. Experimental results show that the proposed method yields significant gains compared with baselines and the golden queries are well generalized to different datasets.
Yifei Xin, Xiulian Peng, Yan Lu 0001
ICASSP1
2023 Improving Weakly Supervised Sound Event Detection with Causal Intervention
abstract
Existing weakly supervised sound event detection (WSSED) work has not explored both types of co-occurrences simultaneously, i.e., some sound events often co-occur, and their occurrences are usually accompanied by specific background sounds, so they would be inevitably entangled, causing misclassification and biased localization results with only clip-level supervision. To tackle this issue, we first establish a structural causal model (SCM) to reveal that the context is the main cause of co-occurrence confounders that mislead the model to learn spurious correlations between frames and clip-level labels. Based on the causal analysis, we propose a causal intervention (CI) method for WSSED to remove the negative impact of co-occurrence confounders by iteratively accumulating every possible context of each class and then re-projecting the contexts to the frame-level features for making the event boundary clearer. Experiments show that our method effectively improves the performance on multiple datasets and can generalize to various baseline models.
Yifei Xin, Dongchao Yang, Fan Cui, Yuexian Zou
ICASSP1
2023 Improving Text-Audio Retrieval by Text-Aware Attention Pooling and Prior Matrix Revised Loss
abstract
In text-audio retrieval (TAR) tasks, due to the heterogeneity of contents between text and audio, the semantic information contained in the text is only similar to certain frames within the audio. Yet, existing works aggregate the entire audio without considering the text, such as mean-pooling over the frames, which is likely to encode misleading audio information not described in the given text. In this paper, we present a text-aware attention pooling (TAP) module for TAR, which is essentially a scaled dot product attention for a text to attend to its most semantically similar frames. Furthermore, previous methods only conduct the softmax for every single-side retrieval, ignoring the potential cross-retrieval information. By exploring the intrinsic prior of each text-audio pair, we introduce a prior matrix revised (PMR) loss to filter the hard case with high (or low) text-to-audio but low (or high) audio-to-text similarity scores, thus achieving the dual optimal match. Experiments show that our TAP significantly outperforms various text-agnostic pooling functions. Moreover, our PMR loss also shows stable performance gains on multiple datasets.
Yifei Xin, Dongchao Yang, Yuexian Zou
ICASSP1
2023 Masked Audio Modeling with CLAP and Multi-Objective Learning
Yifei Xin, Xiulian Peng, Yan Lu 0001
INTERSPEECH1
2023 Background-aware Modeling for Weakly Supervised Sound Event Detection
Yifei Xin, Dongchao Yang, Yuexian Zou
INTERSPEECH1
2023 Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
Yifei Xin, Yuexian Zou
INTERSPEECH1
2023 Cooperative Game Modeling With Weighted Token-Level Alignment for Audio-Text Retrieval
abstract
Previous audio-text retrieval (ATR) methods primarily concentrate on constructing contrastive pairs between entire audio clips and full caption sentences, while neglecting fine-grained cross-modal relationships. In this letter, we first introduce a weighted token-level alignment (WTA) module for ATR to learn fine-grained semantic interactions. Besides, due to the unavailability of manually labeling the fine-grained sequential correspondence between audio-text pairs, we attempt to model ATR as a cooperative game process to flexibly handle the uncertainty during audio-text semantic interactions. Specifically, we treat audio frames and text words as players and present a game theoretic interaction (GTI) method to assess potential correspondence between audio frames and text words, which can also be seen as an additional learning signal to improve the pure audio-text contrastive learning. Furthermore, to implement multi-level WTA and GTI, we develop a token cluster module to cluster the frames/words and calculate the interaction scores between the clustered tokens. Experiments show that our WTA significantly improves the ATR performance on multiple datasets. By combining our GTI, the retrieval performance is further boosted by a large margin.
Yifei Xin, Baojun Wang, Lifeng Shang
IEEE Signal Process. Lett.1
2022 Audio Pyramid Transformer with Domain Adaption for Weakly Supervised Sound Event Detection and Audio Classification
Yifei Xin, Dongchao Yang, Yuexian Zou
INTERSPEECH1