VLDB 2026 Research / reviewers in the wild / expert
Yunhao Fang
dblp:348/9991
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Vision and language · 26% Generative modeling · 16% Efficient and distributed learning · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 21 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-language model |
1.7 | 2 | 2025 | LongVILA: Scaling Long-Context Visual Language Models for Long Videos · ICLR 2025 NVILA: Efficient Frontier Visual Language Models · CVPR 2025 |
Machine learning › Generative modeling
autoregressive model |
0.9 | 1 | 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation · ICLR 2025 |
Robotics › Robot manipulation
dexterous manipulation |
0.9 | 1 | 2025 | Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model · ICLR 2025 |
Machine learning › Efficient and distributed learning
distributed training |
0.9 | 1 | 2025 | LongVILA: Scaling Long-Context Visual Language Models for Long Videos · ICLR 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.9 | 1 | 2025 | Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model · ICLR 2025 |
Computer vision › Video understanding and tracking
long video understanding |
0.9 | 1 | 2025 | LongVILA: Scaling Long-Context Visual Language Models for Long Videos · ICLR 2025 |
Machine learning › Reinforcement learning › policy optimization
residual policy learning |
0.9 | 1 | 2025 | Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model · ICLR 2025 |
Machine learning › Efficient and distributed learning › distributed training › model parallelism
sequence parallelism |
0.9 | 1 | 2025 | LongVILA: Scaling Long-Context Visual Language Models for Long Videos · ICLR 2025 |
Machine learning › Generative modeling › image generation
token-based image generation |
0.9 | 1 | 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation · ICLR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
unified image understanding and generation |
0.9 | 1 | 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation · ICLR 2025 |
Computer vision › Vision and language › vision-language model › vision-language model architecture
unified vision-language model |
0.9 | 1 | 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation · ICLR 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | WorldModelBench: Judging Video Generation Models As World Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.7 | 1 | 2023 | Deductive Verification of Chain-of-Thought Reasoning · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.7 | 1 | 2023 | Distilling Large Vision-Language Model with Out-of-Distribution Generalizability · ICCV 2023 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.7 | 1 | 2023 | Distilling Large Vision-Language Model with Out-of-Distribution Generalizability · ICCV 2023 |
Machine learning › Trustworthy machine learning › verification
self-verification |
0.7 | 1 | 2023 | Deductive Verification of Chain-of-Thought Reasoning · NeurIPS 2023 |
Computer vision › Vision and language
vision-language model distillation |
0.7 | 1 | 2023 | Distilling Large Vision-Language Model with Out-of-Distribution Generalizability · ICCV 2023 |
Robotics › Robot manipulation
grasping |
0.3 | 1 | 2025 | Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model · ICLR 2025 |
Natural language and speech › Language models and text generation
instruction following |
0.3 | 1 | 2025 | WorldModelBench: Judging Video Generation Models As World Models · NeurIPS 2025 |
Parallel and multicore computing › parallelization strategies
model parallelism |
0.3 | 1 | 2025 | LongVILA: Scaling Long-Context Visual Language Models for Long Videos · ICLR 2025 |
Automated reasoning and model checking › program verification
deductive verification |
0.2 | 1 | 2023 | Deductive Verification of Chain-of-Thought Reasoning · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.7multi-modal sequence parallelism · 1.7long context extension · 1.7visual token compression · 0.9scale-then-compress · 0.9discrete visual tokenization · 0.9diffusion policy · 0.9controlled exploration · 0.9behavior transformer · 0.9autoregressive next-token prediction · 0.9natural program · 0.7chain-of-thought prompting · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NVILA: Efficient Frontier Visual Language ModelsabstractVisual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention. This paper introduces NVILA, a family of open VLMs designed to optimize both efficiency and accuracy. Building on top of VILA, we improve its model architecture by first scaling up the spatial and temporal resolutions, and then compressing visual tokens. This "scale-then-compress" approach enables NVILA to efficiently process high-resolution images and long videos. We also conduct a systematic investigation to enhance the efficiency of NVILA throughout its entire lifecycle, from training to deployment. NVILA matches or surpasses the accuracy of many leading open and proprietary VLMs across a wide range of image and video benchmarks. At the same time, it reduces training costs by 1.9-5.1×, prefilling latency by 1.6-2.2×, and decoding latency by 1.2-2.8×. Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Haotian Tang, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Jinyi Hu, Sifei Liu, Ranjay Krishna, Pavlo Molchanov 0001, Jan Kautz, Hongxu Yin, Song Han 0003, Yao Lu 0006 |
CVPR | 13 |
| 2025 | LongVILA: Scaling Long-Context Visual Language Models for Long VideosabstractLong-context capability is critical for multi-modal foundation models, especially for long video understanding. We introduce LongVILA, a full-stack solution for long-context visual-language models by co-designing the algorithm and system. For model training, we upgrade existing VLMs to support long video understanding by incorporating two additional stages, i.e., long context extension and long video supervised fine-tuning. However, training on long video is computationally and memory intensive. We introduce the long-context Multi-Modal Sequence Parallelism (MM-SP) system that efficiently parallelizes long video training and inference, enabling 2M context length training on 256 GPUs without any gradient checkpointing. LongVILA efficiently extends the number of video frames of VILA from 8 to 2048, achieving 99.8% accuracy in 6,000-frame (more than 1 million tokens) video needle-in-a-haystack. LongVILA-7B demonstrates strong accuracy on 9 popular video benchmarks, e.g., 65.1% VideoMME with subtitle. Besides, MM-SP is 2.1x - 5.7x faster than ring style sequence parallelism and 1.1x - 1.4x faster than Megatron with a hybrid context and tensor parallelism. Moreover, it seamlessly integrates with Hugging Face Transformers. Yukang Chen, Fuzhao Xue, Dacheng Li, Qinghao Hu 0004, Ligeng Zhu, Xiuyu Li, Yunhao Fang, Haotian Tang, Shang Yang, Yihui He, Hongxu Yin, Pavlo Molchanov 0001, Jan Kautz, Linxi Fan, Yuke Zhu, Yao Lu 0006, Song Han 0003 |
ICLR | 7 |
| 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and GenerationabstractVILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework. Yecheng Wu, Zhuoyang Zhang, Junyu Chen 0003, Haotian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi 0001, Song Han 0003, Yao Lu 0006 |
ICLR | 6 |
| 2025 | Policy Decorator: Model-Agnostic Online Refinement for Large Policy ModelabstractRecent advancements in robot learning have used imitation learning with large models and extensive demonstrations to develop effective policies. However, these models are often limited by the quantity quality, and diversity of demonstrations. This paper explores improving offline-trained imitation learning models through online interactions with the environment. We introduce Policy Decorator, which uses a model-agnostic residual policy to refine large imitation learning models during online interactions. By implementing controlled exploration strategies, Policy Decorator enables stable, sample-efficient online learning. Our evaluation spans eight tasks across two benchmarks—ManiSkill and Adroit—and involves two state-of-the-art imitation learning models (Behavior Transformer and Diffusion Policy). The results show Policy Decorator effectively improves the offline-trained policies and preserves the smooth motion of imitation learning models, avoiding the erratic behaviors of pure RL policies. See our [project page](https://policydecorator.github.io/) for videos. Xiu Yuan, Tongzhou Mu, Stone Tao, Yunhao Fang, Mengke Zhang, Hao Su 0001 |
ICLR | 4 |
| 2025 | WorldModelBench: Judging Video Generation Models As World ModelsabstractVideo generation models have rapidly progressed, positioning themselves as video world models capable of supporting decision-making applications like robotics and autonomous driving. However, current benchmarks fail to rigorously evaluate these claims, focusing only on general video quality, ignoring important factors to world models such as physics adherence.To bridge this gap, we propose WorldModelBench, a benchmark designed to evaluate the world modeling capabilities of video generation models in application-driven domains. WorldModelBench offers two key advantages: (1) Against to nuanced world modeling violations: By incorporating instruction-following and physics-adherence dimensions, WorldModelBench detects subtle violations, such as irregular changes in object size that breach the mass conservation law—issues overlooked by prior benchmarks. (2) Aligned with large-scale human preferences: We crowd-source 67K human labels to accurately measure 14 frontier models. Using our high-quality human labels, we further fine-tune an accurate judger to automate the evaluation procedure, achieving 9.9% lower error in predicting world modeling violations than GPT-4o with 2B parameters. In addition, we demonstrate that training to align human annotations by maximizing the rewards from the judger noticeably improve the world modeling capability. The dataset is hosted in HuggingFace at https://huggingface.co/datasets/Efficient-Large-Model/worldmodelbench. The code to run evaluation is available at https://github.com/WorldModelBench-Team/WorldModelBench. Dacheng Li, Yunhao Fang, Yukang Chen, Shuo Yang 0011, Shiyi Cao, Justin Wong, Michael Luo, Xiaolong Wang 0004, Hongxu Yin, Joseph Gonzalez 0001, Ion Stoica, Song Han 0003, Yao Lu 0006 |
NeurIPS | 2 |
| 2023 | Distilling Large Vision-Language Model with Out-of-Distribution GeneralizabilityabstractLarge vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distillation, the process of creating smaller, faster models that maintain the performance of larger models, is a promising direction towards the solution. This paper investigates the distillation of visual representations in large teacher vision-language models into lightweight student models using a small- or mid-scale dataset. Notably, this study focuses on open-vocabulary out-of-distribution (OOD) generalization, a challenging problem that has been overlooked in previous model distillation literature. We propose two principles from vision and language modality perspectives to enhance student’s OOD generalization: (1) by better imitating teacher’s visual representation space, and carefully promoting better coherence in vision-language alignment with the teacher; (2) by enriching the teacher’s language representations with informative and fine-grained semantic attributes to effectively distinguish between different labels. We propose several metrics and conduct extensive experiments to investigate their techniques. The results demonstrate significant improvements in zero-shot and few-shot student performance on open-vocabulary out-of-distribution classification, highlighting the effectiveness of our proposed approaches. Code released at this link. Yunhao Fang, Minghua Liu, Zhan Ling, Zhuowen Tu, Hao Su 0001 |
ICCV | 2 |
| 2023 | Deductive Verification of Chain-of-Thought ReasoningabstractLarge Language Models (LLMs) significantly benefit from Chain-of-thought (CoT) prompting in performing various reasoning tasks. While CoT allows models to produce more comprehensive reasoning processes, its emphasis on intermediate reasoning steps can inadvertently introduce hallucinations and accumulated errors, thereby limiting models’ ability to solve complex reasoning tasks. Inspired by how humans engage in careful and meticulous deductive logical reasoning processes to solve tasks, we seek to enable language models to perform explicit and rigorous deductive reasoning, and also ensure the trustworthiness of their reasoning process through self-verification. However, directly verifying the validity of an entire deductive reasoning process is challenging, even with advanced models like ChatGPT. In light of this, we propose to decompose a reasoning verification process into a series of step-by-step subprocesses, each only receiving their necessary context and premises. To facilitate this procedure, we propose Natural Program, a natural language-based deductive reasoning format. Our approach enables models to generate precise reasoning steps where subsequent steps are more rigorously grounded on prior steps. It also empowers language models to carry out reasoning self-verification in a step-by-step manner. By integrating this verification process into each deductive reasoning stage, we significantly enhance the rigor and trustfulness of generated reasoning steps. Along this process, we also improve the answer correctness on complex reasoning tasks. Zhan Ling, Yunhao Fang, Zhiao Huang, Mingu Lee, Roland Memisevic, Hao Su 0001 |
NeurIPS | 2 |