Hongjin Lu

dblp:367/3174 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 39% Language models and text generation · 14% Reinforcement learning · 14%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
data-efficient learning
0.912025
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model inference
inference-time computation
0.912025
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension · ICCV 2025
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
0.912025
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
sample selection
0.912025
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement · NeurIPS 2025
Computer vision › Vision and language
visual reasoning
0.912025
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences · ACL (1) 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.312025
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
temporal reasoning
0.212024
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

vision value model · 0.9reinforcement fine-tuning · 0.9monte carlo tree search · 0.9inference-time search · 0.9multimodal evaluation · 0.8
YearPublicationVenuePosition
2025 Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Zhengyuan Yang, Hongjin Lu, Yuancheng Xu, Chung-Ching Lin, Furong Huang
ICCV4
2025 SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
abstract
We introduce ThinkLite-VL, a family of visual reasoning models that achieve state-of-the-art (SoTA) performance using an order of magnitude fewer training samples, relying purely on reinforcement fine-tuning (RFT) self-improvement without any knowledge distillation. Our central insight is that sample difficulty critically influences RFT effectiveness: appropriately challenging examples can drive substantial reasoning improvements, even in low-data regimes. However, quantifying sample difficulty in a reliable and scalable manner remains non-trivial. To address this, we repurpose Monte Carlo Tree Search (MCTS) to measure sample difficulty via the number of reasoning iterations a vision-language model (VLM) requires to solve each instance. This MCTS-based selection procedure identifies samples that induce deeper reasoning while remaining solvable, allowing us to filter a high-quality subset from 70k open-source examples spanning math, natural image understanding, and chart comprehension. Using this approach, we select just 11k challenging samples for RFT on Qwen2.5-VL-7B-Instruct and 7.5k samples for Qwen2.5-VL-72B-Instruct. The resulting models, ThinkLite-VL-7B and ThinkLite-VL-72B, significantly outperform their respective base models across eight visual reasoning benchmarks. In particular, ThinkLite-VL-7B improves the average performance of Qwen2.5-VL-7B-Instruct by 7\% and surpasses all existing 7B-level models, as well as much larger models such as GPT-4o, O1 and Qwen2.5-VL-72B, achieving a new SoTA score of 75.1 on MathVista. ThinkLite-VL-72B further advances the SoTA frontier, achieving an accuracy of 79.7 on MathVista and an average benchmark improvement of 4.42 over the open-source SOTA. These results demonstrate that MCTS-guided difficulty filtering provides a scalable and effective path toward data-efficient self-improvement in multimodal reasoning.
Zhengyuan Yang, Hongjin Lu, Chung-Ching Lin, Furong Huang
NeurIPS4
2024 Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
abstract
Xiyao Wang, Yuhang Zhou, Xiaoyu Liu, Hongjin Lu, Yuancheng Xu, Feihong He, Jaehong Yoon, Taixi Lu, Fuxiao Liu, Gedas Bertasius, Mohit Bansal, Huaxiu Yao, Furong Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiaoyu Liu 0003, Hongjin Lu, Yuancheng Xu, Feihong He, Jaehong Yoon, Taixi Lu, Fuxiao Liu, Gedas Bertasius, Mohit Bansal, Huaxiu Yao, Furong Huang
ACL (1)4