Yexin Liu

dblp:255/8704 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Vision and language · 21% Efficient and distributed learning · 15% Segmentation and scene understanding · 10%
Computer graphics and multimedia
1 paper
Computational photography and imaging · 100%

Topics — the 25 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.722025
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding · NeurIPS 2025
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly · CVPR 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.622025
Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse Distillation · ICCV 2025
GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-Aware Panoramic Semantic Segmentation · CVPR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
1.422024
GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-Aware Panoramic Semantic Segmentation · CVPR 2024
Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic Segmentation · CVPR 2023
Machine learning › Trustworthy machine learning › interpretability
attention analysis
0.912025
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly · CVPR 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
cross-modal distillation
0.912025
Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse Distillation · ICCV 2025
Machine learning › Generative modeling
diffusion model
0.912025
Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models · ICCV 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding · NeurIPS 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models · ICCV 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly · CVPR 2025
Natural language and speech › Language models and text generation › LLM agents
long-term memory
0.912025
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation · ACL (1) 2025
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.912025
Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse Distillation · ICCV 2025
Computer vision › Vision and language
multimodal dialogue
0.912025
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation · ACL (1) 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
0.912025
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue
0.912025
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation · ACL (1) 2025
Computer vision › Video understanding and tracking
object tracking
0.912025
Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment · NeurIPS 2025
Computer vision › Image recognition and object detection
scene text spotting
0.912025
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models · ICCV 2025
Computer vision › Video understanding and tracking
video instance segmentation
0.912025
Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment · NeurIPS 2025
Computer vision › Vision and language
visual question answering
0.912025
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly · CVPR 2025
Machine learning › Transfer learning and domain adaptation
cross-domain transfer
0.812024
GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-Aware Panoramic Semantic Segmentation · CVPR 2024
Computer vision › Segmentation and scene understanding › semantic segmentation › geometry-aware semantic segmentation
panoramic semantic segmentation
0.812024
GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-Aware Panoramic Semantic Segmentation · CVPR 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.712023
Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic Segmentation · CVPR 2023
Computer vision › 3D vision › depth estimation › deep depth estimation
depth foundation model
0.312025
Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse Distillation · ICCV 2025
Computer vision › Vision and language › multimodal understanding
scene text understanding
0.312025
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding · NeurIPS 2025
Computational photography and imaging
panoramic imaging
0.212023
Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic Segmentation · CVPR 2023

Methods — techniques the papers use, named apart from their topics

visual attention refinement · 0.9sparsity-aware feature mixture · 0.9profiling-based feature reuse · 0.9paired positive and negative data construction · 0.9note-taking strategy · 0.9feature distillation · 0.9content guided refinement · 0.9consistency regularization · 0.9benchmark construction · 0.9attention analysis · 0.9cross-projection training · 0.7contrastive learning · 0.7adversarial training · 0.7
YearPublicationVenuePosition
2025 MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation
abstract
Recent multimodal large language models (MLLMs) have demonstrated significant potential in open-ended conversation, generating more accurate and personalized responses. However, their abilities to memorize, recall, and reason in sustained interactions within real-world scenarios remain underexplored. This paper introduces MMRC, a Multi-Modal Real-world Conversation benchmark for evaluating six core open-ended abilities of MLLMs: information extraction, multi-turn reasoning, information update, image management, memory recall, and answer refusal. With data collected from real-world scenarios, MMRC comprises 5,120 conversations and 28,720 corresponding manually labeled questions, posing a significant challenge to existing MLLMs. Evaluations on 20 MLLMs in MMRC indicate an accuracy drop during open-ended interactions. We identify four common failure patterns: long-term memory degradation, inadequacies in updating factual knowledge, accumulated assumption of error propagation, and reluctance to “say no.” To mitigate these issues, we propose a simple yet effective NOTE-TAKING strategy, which can record key information from the conversation and remind the model during its responses, enhancing conversational capabilities. Experiments across six MLLMs demonstrate significant performance improvements.
Haochen Xue, Yexin Liu, Qidong Huang, Yulong Li 0002, Zhongxing Xu, Chong Zhang 0006, Yutong Xie 0001, Muhammad Imran Razzak, ZongYuan Ge, Jionglong Su, Junjun He, Yu Qiao 0001
ACL (1)4
2025 Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
abstract
Multimodal Large Language Models (MLLMs) have displayed remarkable performance in multi-modal tasks, particularly in visual comprehension. However, we reveal that MLLMs often generate incorrect answers even when they understand the visual content. To this end, we manually construct a benchmark with 12 categories and design evaluation metrics that assess the degree of error in MLLM responses even when the visual content is seemingly understood. Based on this benchmark, we test 15 leading MLLMs and analyze the distribution of attention maps and logits of some MLLMs. Our investigation identifies two primary issues: 1) most instruction tuning datasets predominantly feature questions that "directly" relate to the visual content, leading to a bias in MLLMs’ responses to other indirect questions, and 2) MLLMs’ attention to visual tokens is notably lower than to system and question tokens. We further observe that attention scores between questions and visual tokens as well as the model’s confidence in the answers are lower in response to misleading questions than to straightforward ones. To address the first challenge, we introduce a paired positive and negative data construction pipeline to diversify the dataset. For the second challenge, we propose to enhance the model’s focus on visual content during decoding by refining the text and visual prompt. For the text prompt, we propose a content guided refinement strategy that performs preliminary visual content analysis to generate structured information before answering the question. Additionally, we employ a visual attention refinement strategy that highlights question-relevant visual tokens to increase the model’s attention to visual content that aligns with the question. Extensive experiments demonstrate that these challenges can be significantly mitigated with our proposed dataset and techniques.
Yexin Liu, Zhengyang Liang, Yueze Wang, Xianfeng Wu, Muyang He, Jian Li 0062, Zheng Liu 0011, Harry Yang, Ser-Nam Lim, Bo Zhao 0015
CVPR1
2025 Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models
Xuran Ma, Yexin Liu, Yaofu Liu, Xianfeng Wu, Mingzhe Zheng, Ser-Nam Lim, Harry Yang
ICCV2
2025 Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse Distillation
abstract
With the superior sensitivity of event cameras to high-speed motion and extreme lighting conditions, event-based monocular depth estimation has gained popularity to predict structural information about surrounding scenes in challenging environments. However, the scarcity of labeled event data constrains prior supervised learning methods. Unleashing the promising potential of the existing RGB-based depth foundation model, DAM, we propose Depth Any Event stream (EventDAM) to achieve high-performance event based monocular depth estimation in an annotation-free manner. EventDAM effectively combines paired dense RGB images with sparse event data by incorporating three key cross-modality components: Sparsity-aware Feature Mixture (SFM), Sparsity-aware Feature Distillation (SFD), and Sparsity-invariant Consistency Module (SCM). With the proposed sparsity metric, SFM mixes features from RGB images and event data to generate auxiliary depth predictions, while SFD facilitates adaptive feature distillation. Furthermore, SCM ensures output consistency across varying sparsity levels in event data, thereby endowing EventDAM with zero shot capabilities across diverse scenes. Extensive experiments across a variety of benchmark datasets, compared to approaches using diverse input modalities, robustly substantiate the generalization and zero-shot capabilities of EventDAM.
Jinjing Zhu, Tianbo Pan, Zidong Cao, Yexin Liu, James T. Kwok, Hui Xiong 0001
ICCV4
2025 When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
abstract
Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visually incorrect answers, which we refer to as semantic hallucination. In this work, we investigate the underlying causes of semantic hallucination and identify a key finding: Transformer layers in LLM with stronger attention focus on scene text regions are less prone to producing semantic hallucinations. Thus, we propose a training-free semantic hallucination mitigation framework comprising two key components: (1) ZoomText, a coarse-to-fine strategy that identifies potential text regions without external detectors; and (2) Grounded Layer Correction, which adaptively leverages the internal representations from layers less prone to hallucination to guide decoding, correcting hallucinated outputs for non-semantic samples while preserving the semantics of meaningful ones. To enable rigorous evaluation, we introduce TextHalu-Bench, a benchmark of 1,740 samples spanning both semantic and non-semantic cases, with manually curated question–answer pairs designed to probe model hallucinations. Extensive experiments demonstrate that our method not only effectively mitigates semantic hallucination but also achieves strong performance on public benchmarks for scene text spotting and understanding.
Hangui Lin, Yexin Liu, Gangyan Zeng, Yu Zhou 0015, Ser-Nam Lim, Harry Yang, Nicu Sebe
NeurIPS3
2025 Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment
abstract
360 video captures the complete surrounding scenes with the ultra-large field of view of 360x180. This makes 360 scene understanding tasks, e.g., segmentation and tracking, crucial for appications, such as autonomous driving, robotics. With the recent emergence of foundation models, the community is, however, impeded by the lack of large-scale, labelled real-world datasets. This is caused by the inherent spherical properties, e.g., severe distortion in polar regions, and content discontinuities, rendering the annotation costly yet complex. This paper introduces Leader360V, the first large-scale (10K+), labeled real-world 360 video datasets for instance segmentation and tracking. Our datasets enjoy high scene diversity, ranging from indoor and urban settings to natural and dynamic outdoor scenes. To automate annotation, we design an automatic labeling pipeline, which subtly coordinates pre-trained 2D segmentors and large language models (LLMs) to facilitate the labeling. The pipeline operates in three novel stages. Specifically, in the Initial Annotation Phase, we introduce a Semantic- and Distortion-aware Refinement (SDR) module, which combines object mask proposals from multiple 2D segmentors with LLM-verified semantic labels. These are then converted into mask prompts to guide SAM2 in generating distortion-aware masks for subsequent frames. In the Auto-Refine Annotation Phase, missing or incomplete regions are corrected either by applying the SDR again or resolving the discontinuities near the horizontal borders. The Manual Revision Phase finally incorporates LLMs and human annotators to further refine and validate the annotations. Extensive user studies and evaluations demonstrate the effectiveness of our labeling pipeline. Meanwhile, experiments confirm that Leader360V significantly enhances model performance for 360 video segmentation and tracking, paving the way for more scalable 360 scene understanding. We release our dataset and code at {https://leader360v.github.io/Leader360V_HomePage/} for better understanding.
Dingwen Xiao, Aobotao Dai, Yexin Liu, Tianbo Pan, Shiqi Wen, Lei Chen 0002, Lin Wang 0040
NeurIPS4
2025 Unsupervised Visible-Infrared ReID via Pseudo-Label Correction and Modality-Level Alignment
abstract
Unsupervised visible-infrared person reidentification (UVI-ReID) has recently gained great attention due to its potential for enhancing human detection in diverse environments without labeling. Previous methods utilize intramodality clustering and cross-modality feature matching to achieve UVI-ReID. However, there exist two challenges: 1) noisy pseudo-labels might be generated in the clustering process and 2) the cross-modality feature alignment via matching the marginal distribution of visible and infrared modalities may misalign the different identities from the two modalities. In this article, we first conduct a theoretical analysis where an interpretable generalization upper bound is introduced. Based on the analysis, we then propose a novel unsupervised cross-modality person reidentification framework (PRAISE). Specifically, to address the first challenge, we propose a pseudo-label correction (PLC) strategy that utilizes a beta mixture model (BMM) to predict the probability of misclustering-based network's memory effect and rectifies the correspondence by adding a perceptual term to contrastive learning. Next, we introduce a modality-level alignment (MLA) strategy that generates paired visible-infrared latent features and reduces the modality gap by aligning the labeling function of visible and infrared features to learn identity-discriminative and modality-invariant features. Experimental results on two benchmark datasets demonstrate that our method achieves a state-of-the-art (SOTA) performance than the unsupervised visible-ReID methods.
Yexin Liu, Weiming Zhang 0001, Athanasios V. Vasilakos, Lin Wang 0025
IEEE Trans. Neural Networks Learn. Syst.1
2024 GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-Aware Panoramic Semantic Segmentation
abstract
This paper tackles a novel yet challenging problem: how to transfer knowledge from the emerging Segment Anything Model (SAM) - which reveals impressive zero-shot instance segmentation capacity - to learn a compact panoramic semantic segmentation model, i.e., student, without requiring any labeled data. This poses considerable challenges due to SAM's inability to provide semantic labels and the large capacity gap between SAM and the student. To this end, we propose a novel framework, called GoodSAM, that introduces a teacher assistant (TA) to provide semantic information, integrated with SAM to generate ensemble logits to achieve knowledge transfer. Specifically, we propose a Distortion-Aware Rectification (DAR) module that first addresses the distortion problem of panoramic images by imposing prediction-level consistency and boundary enhancement. This subtly enhances TA's prediction capacity on panoramic images. DAR then incorporates a cross-task complementary fusion block to adaptively merge the predictions of SAM and TA to obtain more reliable ensemble logits. Moreover, we introduce a Multi-level Knowledge Adaptation (MKA) module to efficiently transfer the multi-level feature knowledge from TA and ensemble logits to learn a compact student model. Extensive experiments on two benchmarks show that our GoodSAM achieves a remarkable +3.75% mIoU improvement over the state-of-the-art (SOTA) domain adaptation methods, e.g., [41]. Also, our most lightweight model achieves comparable performance to the SOTA methods with only 3.7M parameters.
Yexin Liu, Lin Wang 0025
CVPR2
2023 Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic Segmentation
abstract
The ability of scene understanding has sparked active research for panoramic image semantic segmentation. However, the performance is hampered by distortion of the equirectangular projection (ERP) and a lack of pixel-wise annotations. For this reason, some works treat the ERP and pinhole images equally and transfer knowledge from the pinhole to ERP images via unsupervised domain adaptation (UDA). However, they fail to handle the domain gaps caused by: 1) the inherent differences between camera sensors and captured scenes; 2) the distinct image formats (e.g., ERP and pinhole images). In this paper, we propose a novel yet flexible dual-path UDA framework, DPPASS, taking ERP and tangent projection (TP) images as inputs. To reduce the domain gaps, we propose cross-projection and intra-projection training. The cross-projection training includes tangent-wise feature contrastive training and prediction consistency training. That is, the former formulates the features with the same projection locations as positive examples and vice versa, for the models' awareness of distortion, while the latter ensures the consistency of cross-model predictions between the ERP and TP. Moreover, adversarial intra-projection training is proposed to reduce the inherent gap, between the features of the pinhole images and those of the ERP and TP images, respectively. Importantly, the TP path can be freely removed after training, leading to no additional inference cost. Extensive experiments on two benchmarks show that our DPPASS achieves + 1.06% mIoU increment than the state-of-the-art approaches. https://vlis2022.github.io/cvpr23/DPPASS
Xu Zheng 0002, Jinjing Zhu, Yexin Liu, Zidong Cao, Chong Fu 0001, Lin Wang 0025
CVPR3
2022 FCP-Net: A Feature-Compression-Pyramid Network Guided by Game-Theoretic Interactions for Medical Image Segmentation
abstract
Medical image segmentation is a crucial step in diagnosis and analysis of diseases for clinical applications. Deep convolutional neural network methods such as DeepLabv3+ have successfully been applied for medical image segmentation, but multi-level features are seldom integrated seamlessly into different attention mechanisms, and few studies have fully explored the interactions between medical image segmentation and classification tasks. Herein, we propose a feature-compression-pyramid network (FCP-Net) guided by game-theoretic interactions with a hybrid loss function (HLF) for the medical image segmentation. The proposed approach consists of segmentation branch, classification branch and interaction branch. In the encoding stage, a new strategy is developed for the segmentation branch by applying three modules, e.g., embedded feature ensemble, dilated spatial mapping and channel attention (DSMCA), and branch layer fusion. These modules allow effective extraction of spatial information, efficient identification of spatial correlation among various features, and fully integration of multi-receptive field features from different branches. In the decoding stage, a DSMCA module and a multi-scale feature fusion module are used to establish multiple skip connections for enhancing fusion features. Classification and interaction branches are introduced to explore the potential benefits of the classification information task to the segmentation task. We further explore the interactions of segmentation and classification branches from a game theoretic view, and design an HLF. Based on this HLF, the segmentation, classification and interaction branches can collaboratively learn and teach each other throughout the training process, thus applying the conjoint information between the segmentation and classification tasks and improving the generalization performance. The proposed model has been evaluated using several datasets, including ISIC2017, ISIC2018, REFUGE, Kvasir-SEG, BUSI, and PH2, and the results prove its competitiveness compared with other state-of-the-art techniques.
Yexin Liu, Lizhu Liu, Zhengjia Zhan, Yueqiang Hu, Huigao Duan
IEEE Trans. Medical Imaging1