Hao Tian 0006

dblp:29/1374-6 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2025
0009-0003-0941-3629ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Vision and language · 52% 3D vision · 15% Autonomous driving · 8%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
3.652025
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text · ICLR 2025
PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation · ICCV 2025
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models · CVPR 2025
Visual content generation and editing
image editing
1.722025
GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing · NeurIPS 2025
PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation · ICCV 2025
Visual content generation and editing › image generation
text-to-image generation
1.722025
GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing · NeurIPS 2025
PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation · ICCV 2025
Robotics › Autonomous driving › perception › 3d perception
bird's-eye-view perception
1.422024
Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe · IEEE Trans. Pattern Anal. Mach. Intell. 2024
BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision · CVPR 2023
Computer vision › Vision and language
multimodal in-context learning
0.912025
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text · ICLR 2025
Computer vision › Vision and language
multimodal understanding
0.912025
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models · ICLR 2025
Computer vision › Vision and language
video-language model
0.912025
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models · CVPR 2025
Computer vision › Vision and language
vision-language dataset
0.912025
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text · ICLR 2025
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.912025
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models · ICLR 2025
Machine learning › Efficient and distributed learning › model compression › token compression
visual token compression
0.912025
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models · CVPR 2025
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing
0.912025
GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing · NeurIPS 2025
Natural language and speech › Language models and text generation
instruction tuning
0.812024
MMInstruct: a high-quality multi-modal instruction tuning dataset with extensive diversity · Sci. China Inf. Sci. 2024
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal instruction tuning
0.812024
MMInstruct: a high-quality multi-modal instruction tuning dataset with extensive diversity · Sci. China Inf. Sci. 2024
Robotics › Robot navigation and mapping
sensor fusion
0.812024
Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › 3D vision
view transformation
0.812024
Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › 3D vision › 3d object detection
bird's-eye-view detection
0.712023
BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision · CVPR 2023
Computer vision › 3D vision › point cloud analysis
LiDAR point cloud understanding
0.512021
Unsupervised Object Detection With LIDAR Clues · CVPR 2021
Computer vision › Image recognition and object detection › object detection
unsupervised object detection
0.512021
Unsupervised Object Detection With LIDAR Clues · CVPR 2021
Natural language and speech › Language models and text generation
large language model
0.312025
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability
visual explanation
0.312025
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models · CVPR 2025
Computer vision › Image recognition and object detection
object detection
0.212023
BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision · CVPR 2023
Computer vision › Image recognition and object detection › object detection
region-based detection
0.212023
BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision · CVPR 2023

Methods — techniques the papers use, named apart from their topics

multimodal pretraining · 1.7instruction tuning · 1.7diffusion model · 1.7chain-of-thought reasoning · 1.7token compression · 0.9temporal redundancy exploitation · 0.9semantic-spatial guidance · 0.9multiple-choice evaluation · 0.9data filtering · 0.9data engine · 0.9benchmark construction · 0.9
YearPublicationVenuePosition
2025 PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
abstract
Large Vision-Language Models (VLMs) have been extended to understand both images and videos. Visual token compression is leveraged to reduce the considerable token length of visual inputs. To meet the needs of different tasks, existing high-performance models usually process images and videos separately with different token compression strategies, limiting the capabilities of combining images and videos. To this end, we extend each image into a "static" video and introduce a unified token compression strategy called Progressive Visual Token Compression (PVC), where the tokens of each frame are progressively encoded and adaptively compressed to supplement the information not extracted from previous frames. Video tokens are efficiently compressed with exploiting the inherent temporal redundancy. Images are repeated as static videos, and the spatial details can be gradually supplemented in multiple frames. PVC unifies the token compressing of images and videos. With a limited number of tokens per frame (64 tokens by default), spatial details and temporal changes can still be preserved. Experiments show that our model achieves state-of-the-art performance across various video understanding benchmarks, including long video tasks and fine-grained short video tasks. Meanwhile, our unified token compression strategy incurs no performance loss on image benchmarks, particularly in detail-sensitive tasks. Code is released at https://github.com/OpenGVLab/PVC.
Xizhou Zhu, Weijie Su 0002, Jiahao Wang 0005, Hao Tian 0006, Zhe Chen 0017, Wenhai Wang, Lewei Lu, Jifeng Dai
CVPR6
2025 PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation
abstract
Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for visual content generation. However, existing works have insufficiently addressed the varying granularity demands of different image generation tasks within a unified MLLM paradigm - from the diversity required in text-to-image generation to the precise controllability needed in image manipulation. In this work, we propose PUMA, emPowering Unified MLLM with Multi-grAnular visual generation. PUMA unifies multi-granular visual features as both inputs and outputs of MLLMs, elegantly addressing the different granularity requirements of various image generation tasks within a unified MLLM framework. Following multimodal pretraining and task-specific instruction tuning, PUMA demonstrates proficiency in a wide range of multimodal tasks. This work represents a significant step towards a truly unified MLLM capable of adapting to the granularity demands of various visual tasks. The code and model will be released in https://github.com/rongyaofang/PUMA.
Rongyao Fang, Chengqi Duan, Kun Wang 0056, Hao Li 0069, Linjiang Huang, Hao Tian 0006, Xingyu Zeng, Rui Zhao 0001, Jifeng Dai, Hongsheng Li 0001, Xihui Liu
ICCV6
2025 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
abstract
Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains the capabilities of large language models during multimodal fine-tuning. However, the limited scale and diversity of current image-text interleaved data restrict the development of multimodal large language models. In this paper, we introduce OmniCorpus, a 10 billion-scale image-text interleaved dataset. Using an efficient data engine, we filter and extract large-scale high-quality documents, which contain 8.6 billion images and 1,696 billion text tokens. Compared to counterparts (e.g., MMC4, OBELICS), our dataset 1) has 15 times larger scales while maintaining good data quality; 2) features more diverse sources, including both English and non-English websites as well as video-centric websites; 3) is more flexible, easily degradable from an image-text interleaved format to pure text corpus and image-text pairs. Through comprehensive analysis and experiments, we validate the quality, usability, and effectiveness of the proposed dataset. We hope this could provide a solid data foundation for future multimodal model research.
Qingyun Li, Zhe Chen 0017, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen 0004, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian 0006, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai
ICLR12
2025 MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
abstract
The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their evaluation has not kept pace with their development. To fill this gap, we introduce the Multimodal Multi-image Understanding (MMIU) benchmark, a comprehensive evaluation suite designed to assess LVLMs across a wide range of multi-image tasks. MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of nearly 30 popular LVLMs, including both open-source and proprietary models, reveals significant challenges in multi-image comprehension, particularly in tasks involving spatial understanding. Even the most advanced models, such as GPT-4o, achieve only 55.7\% accuracy on MMIU. Through multi-faceted analytical experiments, we identify key performance gaps and limitations, providing valuable insights for future model and data improvements. We aim for MMIU to advance the frontier of LVLM research and development. We release the data and code at https://github.com/MMIUBenchmark/MMIU.
Fanqing Meng, Chuanhao Li 0001, Quanfeng Lu, Hao Tian 0006, Tianshuo Yang, Jiaqi Liao, Xizhou Zhu, Jifeng Dai, Yu Qiao 0001, Ping Luo 0002, Kaipeng Zhang, Wenqi Shao
ICLR5
2025 GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing
abstract
Current image generation and editing methods primarily process textual prompts as direct inputs without explicit reasoning about visual composition or operational steps. We present Generation Chain-of-Thought (GoT), a novel paradigm that empowers a Multimodal Large Language Model (MLLM) to first generate an explicit, structured reasoning chain in natural language—detailing semantic relationships, object attributes, and, crucially, precise spatial coordinates—before any image synthesis occurs. This intermediate reasoning output directly guides the subsequent visual generation or editing process. This approach transforms conventional text-to-image generation and editing into a reasoning-guided framework that analyzes semantic relationships and spatial arrangements. We define the formulation of GoT and construct large-scale GoT datasets containing over \textbf{9M} samples with detailed reasoning chains capturing semantic-spatial relationships. To leverage the advantages of GoT, we implement a unified framework that integrates Qwen2.5-VL for reasoning chain generation with an end-to-end diffusion model enhanced by our novel Semantic-Spatial Guidance Module. Experiments show our GoT framework achieves excellent performance on both generation and editing tasks, with significant improvements over baselines. Additionally, our approach enables interactive visual generation, allowing users to explicitly modify reasoning steps for precise image adjustments. GoT pioneers a new direction for reasoning-driven visual generation and editing, producing images that better align with human intent. We will release our datasets and models to facilitate future research.
Rongyao Fang, Chengqi Duan, Kun Wang 0056, Linjiang Huang, Hao Li 0069, Hao Tian 0006, Shilin Yan, Weihao Yu 0005, Xingyu Zeng, Jifeng Dai, Xihui Liu, Hongsheng Li 0001
NeurIPS6
2024 How far are we to GPT-4V? Closing the gap to commercial multimodal models with open-source suites
Zhe Chen 0017, Weiyun Wang, Hao Tian 0006, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma 0012, Jiaqi Wang 0003, Xiaoyi Dong, Hang Yan 0001, Hewei Guo, Conghui He, Botian Shi, Zhenjiang Jin, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai, Licheng Wen, Xiangchao Yan, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu 0002, Dahua Lin, Yu Qiao 0001, Jifeng Dai, Wenhai Wang
Sci. China Inf. Sci.3
2024 MMInstruct: a high-quality multi-modal instruction tuning dataset with extensive diversity
Yangzhou Liu, Zhangwei Gao, Weiyun Wang, Zhe Chen 0017, Wenhai Wang, Hao Tian 0006, Lewei Lu, Xizhou Zhu, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
Sci. China Inf. Sci.7
2024 Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe
abstract
Learning powerful representations in bird's-eye-view (BEV) for perception tasks is trending and drawing extensive attention both from industry and academia. Conventional approaches for most autonomous driving algorithms perform detection, segmentation, tracking, etc., in a front or perspective view. As sensor configurations get more complex, integrating multi-source information from different sensors and representing features in a unified view come of vital importance. BEV perception inherits several advantages, as representing surrounding scenes in BEV is intuitive and fusion-friendly; and representing objects in BEV is most desirable for subsequent modules as in planning and/or control. The core problems for BEV perception lie in (a) how to reconstruct the lost 3D information via view transformation from perspective view to BEV; (b) how to acquire ground truth annotations in BEV grid; (c) how to formulate the pipeline to incorporate features from different sources and views; and (d) how to adapt and generalize algorithms as sensor configurations vary across different scenarios. In this survey, we review the most recent works on BEV perception and provide an in-depth analysis of different solutions. Moreover, several systematic designs of BEV approach from the industry are depicted as well. Furthermore, we introduce a full suite of practical guidebook to improve the performance of BEV perception tasks, including camera, LiDAR and fusion inputs. At last, we point out the future research directions in this area. We hope this report will shed some light on the community and encourage more research effort on BEV perception.
Hongyang Li 0001, Chonghao Sima, Jifeng Dai, Wenhai Wang, Lewei Lu, Huijie Wang, Jiazhi Yang, Hanming Deng, Hao Tian 0006, Enze Xie, Jiangwei Xie, Li Chen 0008, Tianyu Li 0004, Yang Li 0189, Yulu Gao, Xiaosong Jia, Si Liu 0001, Jianping Shi, Dahua Lin, Yu Qiao 0001
IEEE Trans. Pattern Anal. Mach. Intell.11
2023 BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision
abstract
We present a novel bird's-eye-view (BEV) detector with perspective supervision, which converges faster and bet-suits modern image backbones. Existing state-of-the-art BEV detectors are often tied to certain depth pretrained backbones like Vo Vn et, hindering the synergy between booming image backbones and BEV detectors. To address this limitation, we prioritize easing the optimization of BEV detectors by introducing perspective view supervision. To this end, we propose a two-stage BEV detector; where proposals from the perspective head are fed into the bird’ s-eye-view head for final predictions. To evaluate the effectiveness of our model, we conduct extensive ablation studies focusing on the form of supervision and the gener-ality of the proposed detector. The proposed method is ver-ified with a wide spectrum of traditional and modern image backbones and achieves new SoTA results on the large-scale nuScenes dataset. The code shall be released soon.
Yuntao Chen, Hao Tian 0006, Chenxin Tao, Xizhou Zhu, Zhaoxiang Zhang 0001, Gao Huang 0001, Hongyang Li 0001, Yu Qiao 0001, Lewei Lu, Jie Zhou 0001, Jifeng Dai
CVPR3
2021 Unsupervised Object Detection With LIDAR Clues
abstract
Despite the importance of unsupervised object detection, to the best of our knowledge, there is no previous work addressing this problem. One main issue, widely known to the community, is that object boundaries derived only from 2D image appearance are ambiguous and unreliable. To address this, we exploit LiDAR clues to aid unsupervised object detection. By exploiting the 3D scene structure, the issue of localization can be considerably mitigated. We further identify another major issue, seldom noticed by the community, that the long-tailed and open-ended (sub-)category distribution should be accommodated. In this paper, we present the first practical method for unsupervised object detection with the aid of LiDAR clues. In our approach, candidate object segments based on 3D point clouds are firstly generated. Then, an iterative segment labeling process is conducted to assign segment labels and to train a segment labeling network, which is based on features from both 2D images and 3D point clouds. The labeling process is carefully designed so as to mitigate the issue of long-tailed and open-ended distribution. The final segment labels are set as pseudo annotations for object detection network training. Extensive experiments on the large-scale Waymo Open dataset suggest that the derived unsupervised object detection method achieves reasonable accuracy compared with that of strong supervision within the LiDAR visible range.
Hao Tian 0006, Yuntao Chen, Jifeng Dai, Zhaoxiang Zhang 0001, Xizhou Zhu
CVPR1