Chunrui Han

dblp:227/5014 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0001-9725-280XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Vision and language · 43% Generative modeling · 14% 3D vision · 11%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
3.142025
Perception-R1: Pioneering Perception Policy with Reinforcement Learning · NeurIPS 2025
ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning · IJCAI 2024
DreamLLM: Synergistic Multimodal Comprehension and Creation · ICLR 2024
Computer vision › Face, body and person analysis
face recognition
0.922022
Personalized Convolution for Face Recognition · Int. J. Comput. Vis. 2022
Face Recognition with Contrastive Convolution · ECCV (9) 2018
Machine learning › Trustworthy machine learning › AI safety
human-aligned evaluation
0.912025
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation · ICLR 2025
Computer vision › Vision and language
multimodal evaluation
0.912025
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation · ICLR 2025
Machine learning › Generative modeling › diffusion model
personalized image generation
0.912025
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation · ICLR 2025
Machine learning › Reinforcement learning › knowledge-based reinforcement learning
rule-based reinforcement learning
0.912025
Perception-R1: Pioneering Perception Policy with Reinforcement Learning · NeurIPS 2025
Computer vision › 3D vision › 3d shape analysis
3d shape understanding
0.812024
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction · ECCV (43) 2024
Computer vision › 3D vision › 3d scene understanding
multimodal 3d understanding
0.812024
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction · ECCV (43) 2024
Machine learning › Generative modeling
multimodal generation
0.812024
DreamLLM: Synergistic Multimodal Comprehension and Creation · ICLR 2024
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal instruction tuning
0.812024
ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning · IJCAI 2024
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal understanding and generation
0.812024
DreamLLM: Synergistic Multimodal Comprehension and Creation · ICLR 2024
Computer vision › Vision and language › vision-language model › multimodal large language model
chart understanding
0.212024
OneChart: Purify the Chart Structural Extraction via One Auxiliary Token · ACM Multimedia 2024
Machine learning › Generative modeling
image tokenization
0.212024
Vary: Scaling up the Vision Vocabulary for Large Vision-Language Model · ECCV (4) 2024
Natural language and speech › Language models and text generation
instruction tuning
0.212024
ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning · IJCAI 2024
Machine learning › Deep learning architectures and training
convolutional neural network
0.112018
Face Recognition with Contrastive Convolution · ECCV (9) 2018

Methods — techniques the papers use, named apart from their topics

large language model · 1.5task reinforcement · 0.9reward design · 0.9multimodal GPT · 0.9GRPO · 0.9vision vocabulary · 0.8referring instruction tuning · 0.8large vision-language model · 0.8diffusion model · 0.8auxiliary token · 0.8
YearPublicationVenuePosition
2025 DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
abstract
Personalized image generation holds great promise in assisting humans in everyday work and life due to its impressive function in creatively generating personalized content. However, current evaluations either are automated but misalign with humans or require human evaluations that are time-consuming and expensive. In this work, we present DreamBench++, a human-aligned benchmark that advanced multimodal GPT models automate. Specifically, we systematically design the prompts to let GPT be both human-aligned and self-aligned, empowered with task reinforcement. Further, we construct a comprehensive dataset comprising diverse images and prompts. By benchmarking 7 modern generative models, we demonstrate that \dreambench results in significantly more human-aligned evaluation, helping boost the community with innovative findings.
Yuang Peng, Haomiao Tang, Zekun Qi, Runpei Dong, Chunrui Han, Zheng Ge, Xiangyu Zhang 0005, Shutao Xia
ICLR7
2025 Perception-R1: Pioneering Perception Policy with Reinforcement Learning
abstract
Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, our initial experiments reveal that incorporating a thinking process through RL does not consistently lead to performance gains across all visual perception tasks. This leads us to delve into the essential role of RL in the context of visual perception. In this work, we return to the fundamentals and explore the effects of RL on different perception tasks. We observe that the perceptual perplexity is a major factor in determining the effectiveness of RL. We also observe that reward design plays a crucial role in further approaching the upper limit of model perception. To leverage these findings, we propose Perception-R1, a scalable RL framework using GRPO during MLLM post-training. With a standard Qwen2-VL-2B-Instruct, Perception-R1 achieves +4.2% on RefCOCO+, +17.9% on PixMo-Count, +4.2% on PageOCR, and notably, 31.9% AP on COCO2017 val for the first time, establishing a strong baseline for perception policy learning.
En Yu, Kangheng Lin, Jisheng Yin, Yana Wei, Yuang Peng, Jianjian Sun, Chunrui Han, Zheng Ge, Xiangyu Zhang 0005, Daxin Jiang, Jingyu Wang 0001, Wenbing Tao
NeurIPS9
2024 ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
Zekun Qi, Runpei Dong, Shaochen Zhang, Chunrui Han, Zheng Ge, Li Yi 0001, Kaisheng Ma
ECCV (43)5
2024 Vary: Scaling up the Vision Vocabulary for Large Vision-Language Model
Lingyu Kong, Jinyue Chen, Zheng Ge, Jianjian Sun, Chunrui Han, Xiangyu Zhang 0005
ECCV (4)8
2024 DreamLLM: Synergistic Multimodal Comprehension and Creation
abstract
This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation. DreamLLM operates on two fundamental principles. The first focuses on the generative modeling of both language and image posteriors by direct sampling in the raw multimodal space. This approach circumvents the limitations and information loss inherent to external feature extractors like CLIP, and a more thorough multimodal understanding is obtained. Second, DreamLLM fosters the generation of raw, interleaved documents, modeling both text and image contents, along with unstructured layouts. This allows DreamLLM to learn all conditional, marginal, and joint multimodal distributions effectively. As a result, DreamLLM is the first MLLM capable of generating free-form interleaved content. Comprehensive experiments highlight DreamLLM's superior performance as a zero-shot multimodal generalist, reaping from the enhanced learning synergy. Project page: https://dreamllm.github.io.
Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jianjian Sun, Xiangwen Kong, Xiangyu Zhang 0005, Kaisheng Ma, Li Yi 0001
ICLR2
2024 ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning
En Yu, Zheng Ge, Jianjian Sun, Yuang Peng, Runpei Dong, Chunrui Han, Xiangyu Zhang 0005
IJCAI10
2024 OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
Jinyue Chen, Lingyu Kong, Zheng Ge, Jianjian Sun, Chunrui Han, Xiangyu Zhang 0005
ACM Multimedia8
2022 Personalized Convolution for Face Recognition
Chunrui Han, Shiguang Shan, Meina Kan, Shuzhe Wu, Xilin Chen 0001
Int. J. Comput. Vis.1
2021 Learning to Learn Adaptive Classifier-Predictor for Few-Shot Learning
abstract
Few-shot learning aims to learn a well-performing model from a few labeled examples. Recently, quite a few works propose to learn a predictor to directly generate model parameter weights with episodic training strategy of meta-learning and achieve fairly promising performance. However, the predictor in these works is task-agnostic, which means that the predictor cannot adjust to novel tasks in the testing phase. In this article, we propose a novel meta-learning method to learn how to learn task-adaptive classifier-predictor to generate classifier weights for few-shot classification. Specifically, a meta classifier-predictor module, (MPM) is introduced to learn how to adaptively update a task-agnostic classifier-predictor to a task-specialized one on a novel task with a newly proposed center-uniqueness loss function. Compared with previous works, our task-adaptive classifier-predictor can better capture characteristics of each category in a novel task and thus generate a more accurate and effective classifier. Our method is evaluated on two commonly used benchmarks for few-shot classification, i.e., miniImageNet and tieredImageNet. Ablation study verifies the necessity of learning task-adaptive classifier-predictor and the effectiveness of our newly proposed center-uniqueness loss. Moreover, our method achieves the state-of-the-art performance on both benchmarks, thus demonstrating its superiority.
Nan Lai, Meina Kan, Chunrui Han, Xingguang Song, Shiguang Shan
IEEE Trans. Neural Networks Learn. Syst.3
2021 Corrections to "Learning to Learn Adaptive Classifier-Predictor for Few-Shot Learning"
abstract
In the above article [1], the results of "Fully-supervised (Upper bound)" in Tables III and IV were inadvertently set to intermediate records that were used as placeholders. This error has no effect on any of the interpretations and conclusions. Tables I and II of this amendment show the corrected results (highlighted in italics) of the original Tables III and IV.
Nan Lai, Meina Kan, Chunrui Han, Xingguang Song, Shiguang Shan
IEEE Trans. Neural Networks Learn. Syst.3
2018 Face Recognition with Contrastive Convolution
Chunrui Han, Shiguang Shan, Meina Kan, Shuzhe Wu, Xilin Chen 0001
ECCV (9)1