Yuanhan Zhang

dblp:257/7120 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-9063-7886ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Video-MMMU: Evaluating Knowledge Acquisition from Multidisciplinary Professional Videos
abstract
Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Xiang Yue, Bo Li, Yuanhan Zhang, Ziwei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Xiang Yue, Bo Li 0080, Yuanhan Zhang, Ziwei Liu 0002
ACL (1)7
2025 EgoLife: Towards Egocentric Life Assistant
abstract
We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one week, continuously recording their daily activities—including discussions, shopping, cooking, social-izing, and entertainment—using AI glasses for multimodal person-view video references. This effort resulted in EgoLife Dataset, a comprehensive 300-hour egocentric, terpersonal, multiview, and multimodal daily life with intensive annotation. Leveraging this dataset, we troduce EgoLifeQA, a suite of long-context, life-oriented question-answering tasks designed to provide meaningful sistance in daily life by addressing practical questions as recalling past relevant events, monitoring health and offering personalized recommendations.To address the key technical challenges of 1) developing robust visual-audio models for egocentric data, 2) enabling identity recognition, and 3) facilitating long-context question answering over extensive temporal information, we introduce EgoBulter, an integrated system comprising EgoGPT and EgoRAG. EgoGPT is an omni-modal model trained on egocentric datasets, achieving state-of-the-art performance on egocentric video understanding. EgoRAG is a retrieval-based component that supports answering ultra-long-context questions. Our experimental studies verify their working mechanisms and reveal critical factors and bottlenecks, guiding future improvements. By releasing our datasets, models, and benchmarks, we aim to stimulate further research in egocentric AI assistants.
Shuai Liu 0002, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li 0080, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Jörg Widmer, Francesco Gringoli, Lei Yang 0059, Ziwei Liu 0002
CVPR16
2025 Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
abstract
Human intelligence requires correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintains consistent performance in challenging conditions. Despite advances in video large language models (video LLMs), existing benchmarks inadequately reflect the gap between these models and human intelligence in maintaining correctness and robustness in video interpretation. We introduce the Video Thinking Test (Video-TT), to assess if video LLMs can interpret real-world videos as effectively as humans. Video-TT reflects genuine gaps in understanding complex visual narratives, and evaluates robustness against natural adversarial questions. Video-TT comprises 1,000 YouTube Shorts videos, each with one open-ended question and four adversarial questions that probe visual and narrative complexity. Our evaluation shows a significant gap between video LLMs and human performance.
Yuanhan Zhang, Yunice Chew, Yuhao Dong, Aria Leo
ICCV1
2025 LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
abstract
Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains less explored. Additionally, prior LMM research separately tackles different scenarios, leaving it impossible to generalize cross scenarios with new emerging capabilities. To this end, we introduce LLaVA-Interleave, which simultaneously tackles Multi-image, Multi-frame (video), Multi-view (3D), and Multi-patch (single-image) scenarios in LMMs. To enable these capabilities, we regard the interleaved data format as a general template and compile the M4-Instruct dataset with 1,177.6k samples, spanning 4 primary domains with 14 tasks and 41 datasets. We also curate the LLaVA-Interleave Bench to comprehensively evaluate the multi-image performance of LMMs. Through extensive experiments, LLaVA-Interleave achieves leading results in multi-image, video, and 3D benchmarks, while maintaining the performance of single-image tasks. Besides, our model also exhibits several emerging capabilities, e.g., transferring tasks across different settings and modalities.
Feng Li 0040, Renrui Zhang, Hao Zhang 0097, Yuanhan Zhang, Bo Li 0080, Wei Li 0119, Zejun Ma 0001, Chunyuan Li
ICLR4
2025 Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
abstract
Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander G Hauptmann, Yonatan Bisk, Yiming Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alex Hauptmann 0001, Yonatan Bisk, Yiming Yang 0002
NAACL (Long Papers)6
2025 Bamboo: Building Mega-Scale Vision Dataset Continually with Human-Machine Synergy
Yuanhan Zhang, Qinghong Sun, Yichun Zhou, Zexin He, Zhenfei Yin, Kun Wang 0056, Lu Sheng, Yu Qiao 0001, Ziwei Liu 0002
Int. J. Comput. Vis.1
2025 Otter: A Multi-Modal Model With In-Context Instruction Tuning
abstract
Recent advances in Large Multimodal Models (LMMs) have unveiled great potential as visual assistants. However, most existing works focus on responding to individual instructions or using previous dialogues for contextual understanding. There is little discussion on employing both images and text as in-context examples to enhance the instruction following capability. To bridge this gap, we introduce the Otter model to leverage both textual and visual in-context examples for instruction tuning. Specifically, Otter builds upon Flamingo with Perceiver architecture, and has been instruction tuned for general purpose multi-modal assistant. Otter seamlessly processes multi-modal inputs, supporting modalities including text, multiple images, and dynamic video content. To support the training of Otter, we present the MIMIC-IT (MultI-Modal In-Context Instruction Tuning) dataset, which encompasses over 3 million multi-modal instruction-response pairs, including approximately 2.2 million unique instructions across a broad spectrum of images and videos. MIMIC-IT has been carefully curated to feature a diverse array of in-context examples for each entry. Comprehensive evaluations suggest that instruction tuning with these in-context examples substantially enhances model convergence and generalization capabilities. Notably, the extensive scenario coverage provided by the MIMIC-IT dataset empowers the Otter model to excel in tasks involving complex video and multi-image understanding.
Bo Li 0080, Yuanhan Zhang, Liangyu Chen 0005, Fanyi Pu, Joshua Adrian Cahyono, Chunyuan Li, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Neural Prompt Search
abstract
The size of vision models has grown exponentially over the last few years, especially after the emergence of Vision Transformer. This has motivated the development of parameter-efficient tuning methods, such as learning adapter layers or visual prompt tokens, which allow a tiny portion of model parameters to be trained whereas the vast majority obtained from pre-training are frozen. However, designing a proper tuning method is non-trivial: one might need to try out a lengthy list of design choices, not to mention that each downstream dataset often requires custom designs. In this paper, we view the existing parameter-efficient tuning methods as “prompt modules” and propose Neural prOmpt seArcH (NOAH), a novel approach that learns, for large vision models, the optimal design of prompt modules through a neural architecture search algorithm, specifically for each downstream dataset. By conducting extensive experiments on over 20 vision datasets, we demonstrate that NOAH (i) is superior to individual prompt modules, (ii) has good few-shot learning ability, and (iii) is domain-generalizable.
Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Learning Without Forgetting for Vision-Language Models
abstract
Class-Incremental Learning (CIL) or continual learning is a desired capability in the real world, which requires a learning system to adapt to new tasks without forgetting former ones. While traditional CIL methods focus on visual information to grasp core features, recent advances in Vision-Language Models (VLM) have shown promising capabilities in learning generalizable representations with the aid of textual information. However, when continually trained with new classes, VLMs often suffer from catastrophic forgetting of former knowledge. Applying VLMs to CIL poses two major challenges: 1) how to adapt the model without forgetting and 2) how to make full use of the multi-modal information. To this end, we propose PROjectiOn Fusion (Proof) that enables VLMs to learn without forgetting. To handle the first challenge, we propose training task-specific projections based on the frozen image/text encoders. When facing new tasks, new projections are expanded, and former projections are fixed, alleviating the forgetting of old concepts. For the second challenge, we propose the fusion module to better utilize the cross-modality information. By jointly adjusting visual and textual features, the model can capture better task-specific semantic information that facilitates recognition. Extensive experiments on nine benchmark datasets with various continual learning scenarios and various VLMs validate that Proof achieves state-of-the-art performance.
Da-Wei Zhou 0001, Yuanhan Zhang, Jingyi Ning, Han-Jia Ye, De-Chuan Zhan, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Robust face anti-spoofing with Dual Probabilistic Modeling
Yuanhan Zhang, Yichao Wu, Zhenfei Yin, Ziwei Liu 0002
Pattern Recognit.1
2024 VBench: Comprehensive Benchmark Suite for Video Generative Models
abstract
Video generation has witnessed significant advance-ments, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal eval-uation system should provide insights to inform future de-velopments of video generation. To this end, we present VBench, a comprehensive benchmark suite that dissects “video generation quality” into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench has three appealing proper-ties: 1) Comprehensive Dimensions: VBench comprises 16 dimensions in video generation (e.g., subject identity in-consistency, motion smoothness, temporal flickering, and spatial relationship, etc.). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investi-gate the gaps between video and image generation models. We will open-source VBench, including all prompts, evaluation methods, generated videos, and human preference an-notations, and also include more video generation models in VBench to drive forward the field of video generation.
Yinan He, Jiashuo Yu, Fan Zhang 0045, Chenyang Si, Yuming Jiang 0003, Yuanhan Zhang, Tianxing Wu 0002, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang 0001, Limin Wang 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002
CVPR7
2024 MMBench: Is Your Multi-modal Model an All-Around Player?
Yuan Liu 0025, Haodong Duan, Yuanhan Zhang, Bo Li 0080, Songyang Zhang 0001, Wangbo Zhao, Yike Yuan, Jiaqi Wang 0003, Conghui He, Ziwei Liu 0002, Kai Chen 0026, Dahua Lin
ECCV (6)3
2024 🐱 FunQA: Towards Surprising Video Comprehension
Binzhu Xie, Zitang Zhou, Bo Li 0080, Yuanhan Zhang, Jack Hessel, Ziwei Liu 0002
ECCV (1)5
2024 Octopus: Embodied Vision-Language Programmer from Environmental Feedback
Yuhao Dong, Shuai Liu 0002, Bo Li 0080, Haoran Tan, Chencheng Jiang, Jiamu Kang, Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu 0002
ECCV (1)9
2024 3D Point Cloud Pre-Training with Knowledge Distilled from 2D Images
abstract
The success of pre-trained 2D vision models can largely be attributed to their ability to learn from large-scale datasets. However, compared with 2D image datasets, current pre-training data for 3D point clouds are limited. In this paper, we propose a knowledge distillation method for pre-training 3D point cloud models by directly acquiring knowledge from a 2D representation learning model, specifically the image encoder of CLIP. To close the significant domain gap between 2D images and 3D point clouds, we propose to align the features from the two domains at the concept level. Our method utilizes a cross-attention mechanism to extract concept features from 3D point clouds and compares them with the corresponding information from 2D images. This approach bridges the two domains and allows point cloud models to learn directly from the rich information contained in 2D teacher models. Extensive experiments show that our proposed knowledge distillation scheme achieves higher accuracy than the state-of-the-art 3D pre-training methods for synthetic and real-world datasets on various downstream tasks, including object classification, object detection, semantic segmentation and part segmentation.
Yuanhan Zhang, Zhenfei Yin, Jiebo Luo 0001, Wanli Ouyang, Xiaoshui Huang
ICME2
2023 What Makes Good Examples for Visual In-Context Learning?
abstract
Large vision models with billions of parameters and trained on broad data have great potential in numerous downstream applications. However, these models are typically difficult to adapt due to their large parameter size and sometimes lack of accesss to their weights---entities able to develop large vision models often provide APIs only. In this paper, we study how to better utilize large vision models through the lens of in-context learning, a concept that has been well-known in natural language processing but has only been studied very recently in computer vision. In-context learning refers to the ability to perform inference on tasks never seen during training by simply conditioning on in-context examples (i.e., input-output pairs) without updating any internal model parameters. To demystify in-context learning in computer vision, we conduct an extensive research and identify a critical problem: downstream performance is highly sensitivie to the choice of visual in-context examples. To address this problem, we propose a prompt retrieval framework specifically for large vision models, allowing the selection of in-context examples to be fully automated. Concretely, we provide two implementations: (i) an unsupervised prompt retrieval method based on nearest example search using an off-the-shelf model, and (ii) a supervised prompt retrieval method, which trains a neural network to choose examples that directly maximize in-context learning performance. Both methods do not require access to the internal weights of large vision models. Our results demonstrate that our methods can bring non-trivial improvements to visual in-context learning in comparison to the commonly-used random selection. Code and models will be released.
Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu 0002
NeurIPS1
2022 Benchmarking Omni-Vision Representation Through the Lens of Visual Realms
abstract
Though impressive performance has been achieved in specific visual realms (e.g. faces, dogs, and places), an omni-vision representation generalizing to many natural visual domains is highly desirable. But, existing benchmarks are biased and inefficient to evaluate the omni-vision representation—these benchmarks either only include several specific realms, or cover most realms at the expense of subsuming numerous datasets that have extensive realm overlapping. In this paper, we propose Omni-Realm Benchmark (OmniBenchmark). It includes 21 realm-wise datasets with 7,372 concepts and 1,074,346 images. Without semantic overlapping, these datasets cover most visual realms comprehensively and meanwhile efficiently. In addition, we propose a new supervised contrastive learning framework, namely Relational Contrastive learning (ReCo), for a better omni-vision representation. Beyond pulling two instances from the same concept closer—the typical supervised contrastive learning framework—ReCo also pulls two instances from the same semantic realm closer, encoding the semantic relation between concepts, facilitating omni-vision representation learning. We benchmark ReCo and other advances in omni-vision representation studies that are different in architectures (from CNNs to transformers) and in learning paradigms (from supervised learning to self-supervised learning) on OmniBenchmark. We illustrate the superior of ReCo to other supervised contrastive learning methods, and reveal multiple practical observations to facilitate future research. The code and models are available at https://zhangyuanhan-ai.github.io/OmniBenchmark .
Yuanhan Zhang, Zhenfei Yin, Ziwei Liu 0002
ECCV (7)1
2020 CelebA-Spoof: Large-Scale Face Anti-spoofing Dataset with Rich Annotations
Yuanhan Zhang, Zhenfei Yin, Yidong Li, Guojun Yin, Ziwei Liu 0002
ECCV (12)1