Minghe Gao

dblp:223/7088 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
16since 2021 · last 2026
0009-0000-9705-5398ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
abstract
Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiheng Xi, Dingwen Yang, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang 0001, Zhonghang Lu, Jiazheng Zhang, Dingwei Zhu, Junzhe Wang 0001, Zhihao Zhang 0002, Yuming Yang 0001, Junjie Ye 0005, Minghe Gao, Dongrui Liu, Jiaming Ji, Tao Gui, Xuanjing Huang 0001
ACL (1)19
2026 Momentor++: Advancing Video Large Language Models With Fine-Grained Long Video Reasoning
abstract
Large Language Models (LLMs) exhibit remarkable proficiency in understanding and managing text-based tasks.Many works try to transfer these capabilities to the video domain, which are referred to as Video-LLMs. However, current Video-LLMs can only grasp the coarse-grained semantics and are unable to efficiently handle tasks involving the comprehension or localization of specific video segments. To address these challenges, we propose Momentor, a Video-LLM designed to perform fine-grained temporal understanding tasks. To facilitate the training of Momentor, we develop an automatic data generation engine to build Moment-10M, a large-scale video instruction dataset with segment-level instruction data. Building upon the foundation of the previously published Momentor and the Moment-10M dataset, we further extend this work by introducing a Spatio-Temporal Token Consolidation (STTC) method, which can merge redundant visual tokens spatio-temporally in a parameter-free manner, thereby significantly promoting computational efficiency while preserving fine-grained visual details. We integrate STTC with Momentor to develop Momentor++ and validate its performance on various benchmarks. Momentor demonstrates robust capabilities in fine-grained temporal understanding and localization. Further, Momentor++ excels in efficiently processing and analyzing extended videos with complex events, showcasing marked advancements in handling extensive temporal contexts.
Juncheng Li 0006, Minghe Gao, Xiangnan He 0001, Siliang Tang, Wei-Shi Zheng 0001, Jun Xiao 0001, Meng Wang 0001, Tat-Seng Chua, Yueting Zhuang
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Structure-Induced Gradient Regulation for Generalizable Vision-Language Models
abstract
Prompt tuning, a recently emerging paradigm, adapts vision-language pre-trained models to new tasks efficiently by learning "soft prompts" for frozen models. However, in few-shot scenarios, its effectiveness is limited by sensitivity to the initialization and the time-consuming search for optimal initialization, hindering rapid adaptation. Additionally, prompt tuning risks reducing the models' generalizability due to overfitting on scarce training samples. To overcome these challenges, we introduce a novel Gradient-RegulAted Meta-prompt learning (GRAM) framework that jointly meta-learns an efficient soft prompt initialization for better adaptation and a lightweight gradient regulating function for strong cross-domain generalizability in a meta-learning paradigm using only the weakly labeled image-text pre-training data. This is achieved through a Cross-Modal Hierarchical Clustering algorithm that organizes extensive image-text data into a structured hierarchy, facilitating robust meta-learning across diverse domains. Rather than designing a specific prompt tuning method, our GRAM can be easily incorporated into various prompt tuning methods in a model-agnostic way and bring about consistent improvement for them. Further, we consider a more practical but challenging setting: test-time prompt tuning with only unlabeled test samples and propose an improved structure-induced gradient regulating function to leverage the structured semantics of the meta-learning data for zero-shot generalization. This novel approach exploits the hierarchically clustered meta-learning data to model relationships between test-time data and meta-learning prototypes, facilitating the transfer of invariant knowledge without explicit annotations. Meanwhile, we introduce a structure complexity-informed strategy for adaptively constructing meta-training tasks and generating prototypes, which fully considers the diverse semantics within hierarchical clusters of different complexities. Comprehensive experiments demonstrate the state-of-the-art few- and zero-shot generalizability of our method.
Juncheng Li 0006, Minghe Gao, Siliang Tang, Longhui Wei, Jun Xiao 0001, Fei Wu 0001, Richang Hong, Meng Wang 0001, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
abstract
Video Large Language Models (Video-LLMs) have recently shown strong performance in basic video understanding tasks, such as captioning and coarse-grained question answering, but struggle with compositional reasoning that requires multi-step spatio-temporal inference across object relations, interactions, and events. The hurdles to enhancing this capability include extensive manual labor, the lack of spatio-temporal compositionality in existing training data and the absence of explicit reasoning supervision. In this paper, we propose STEP, a novel graph-guided self-training method that enables Video-LLMs to generate reasoning-rich fine-tuning data from any raw videos to improve itself. Specifically, we first induce Spatio-Temporal Scene Graph (STSG) representation of diverse videos to capture fine-grained, multi-granular video semantics. Then, the STSGs guide the derivation of multi-step reasoning Question-Answer (QA) data with Chain-of-Thought (CoT) rationales. Both answers and rationales are integrated as training objective, aiming to enhance model’s reasoning abilities by supervision over explicit reasoning steps. Experimental results demonstrate the effectiveness of STEP across models of varying scales, with a significant 21.3% improvement in tasks requiring three or more reasoning steps. Furthermore, it achieves superior performance with a minimal amount of self-generated rationale-enriched training samples in both compositional reasoning and comprehensive understanding benchmarks, highlighting the broad applicability and vast potential.
Haiyi Qiu, Minghe Gao, Kaihang Pan, Juncheng Li 0006, Wenjie Wang 0007, Siliang Tang, Yueting Zhuang, Tat-Seng Chua
CVPR2
2025 Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
abstract
Recent advancements in reward signal usage for Large Language Models (LLMs) are remarkable. However, significant challenges exist when transitioning reward signal to the multimodal domain, including labor-intensive annotations, over-reliance on one-step rewards, and inadequate evaluation. To address these issues, we propose SVIP, a novel approach to train a step-level multi-dimensional Chain-of-Thought (CoT) reward model automatically. It generates code for solving visual tasks and transforms the analysis of code blocks into the evaluation of CoT step as training samples. Then, we train SVIP-Reward model using a multi-head attention mechanism called TriAtt-CoT. The advantages of SVIP-Reward are evident throughout the entire process of MLLM. We also introduce a benchmark for CoT reward model training and testing. Experimental results demonstrate that SVIP-Reward improves MLLM performance across training and inference-time scaling, yielding better results on benchmarks while reducing hallucinations and enhancing reasoning ability.
Minghe Gao, Xuqi Liu, Zhongqi Yue, Juncheng Li 0006, Siliang Tang, Fei Wu 0001, Tat-Seng Chua, Yueting Zhuang
ICCV1
2025 Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
abstract
Digital agents are increasingly employed to automate tasks in interactive digital environments such as web pages, software applications, and operating systems. While text-based agents built on Large Language Models (LLMs) often require frequent updates due to platform-specific APIs, visual agents leveraging Multimodal Large Language Models (MLLMs) offer enhanced adaptability by interacting directly with Graphical User Interfaces (GUIs). However, these agents face significant challenges in visual perception, particularly when handling high-resolution, visually complex digital environments. This paper introduces Iris, a foundational visual agent that addresses these challenges through two key innovations: Information-Sensitive Cropping (ISC) and Self-Refining Dual Learning (SRDL). ISC dynamically identifies and prioritizes visually dense regions using a edge detection algorithm, enabling efficient processing by allocating more computational resources to areas with higher information density. SRDL enhances the agent's ability to handle complex tasks by leveraging a dual-learning loop, where improvements in referring (describing UI elements) reinforce grounding (locating elements) and vice versa, all without requiring additional annotated data. Empirical evaluations demonstrate that Iris achieves state-of-the-art performance across multiple benchmarks with only 850K GUI annotations, outperforming methods using 10x more training data. These improvements further translate to significant gains in both web and OS agent downstream tasks.
Zhiqi Ge, Juncheng Li 0006, Xinglei Pang, Minghe Gao, Kaihang Pan, Hao Fei 0001, Wenqiao Zhang, Siliang Tang, Yueting Zhuang
ICCV4
2025 On Path to Multimodal Generalist: General-Level and General-Bench
abstract
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of language-based LLMs. Unlike their specialist predecessors, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple modalities, these models have advanced to not only comprehend but also generate across modalities. Their capabilities have expanded from coarse-grained to fine-grained multimodal understanding and from supporting singular modalities to accommodating a wide array of or even arbitrary modalities. To assess the capabilities of various MLLMs, a diverse array of benchmark test sets has been proposed. This leads to a critical question: Can we simply assume that higher performance across tasks indicates a stronger MLLM capability, bringing us closer to human-level AI? We argue that the answer is not as straightforward as it seems. In this project, we introduce an evaluation framework to delineate the capabilities and behaviors of current multimodal generalists. This framework, named General-Level, establishes 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI (Artificial General Intelligence). Central to our framework is the use of Synergy as the evaluative criterion, categorizing capabilities based on whether MLLMs preserve synergy across comprehension and generation, as well as across multimodal interactions. To evaluate the comprehensive abilities of various generalists, we present a massive multimodal benchmark, General-Bench, which encompasses a broader spectrum of skills, modalities, formats, and capabilities, including over 700 tasks and 325,800 instances. The evaluation results that involve over 100 existing state-of-the-art MLLMs uncover the capability rankings of generalists, highlighting the challenges in reaching genuine AI. We expect this project to pave the way for future research on next-generation multimodal foundation models, providing a robust infrastructure to accelerate the realization of AGI. Project Page: https://generalist.top/, Leaderboard: https://generalist.top/leaderboard/, Benchmark: https://huggingface.co/General-Level/.
Hao Fei 0001, Yuan Zhou 0016, Juncheng Li 0006, Xiangtai Li, Qingshan Xu 0001, Bobo Li 0001, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang 0042, Weiming Wu, Tianjie Ju, Zixiang Meng, Shilin Xu 0001, Liyu Jia, Meng Luo 0010, Jiebo Luo 0001, Tat-Seng Chua, Shuicheng Yan, Hanwang Zhang
ICML14
2025 What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
abstract
As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual annotation, and a lack of multidimensional evaluation. In response to these challenges, we introduce OmniBench, a self-generating, graph-based benchmark with an automated pipeline for synthesizing tasks of controllable complexity through subtask composition. To evaluate the diverse capabilities of virtual agents on the graph, we further present OmniEval, a multidimensional evaluation framework that includes subtask-level evaluation, graph-based metrics, and comprehensive tests across 10 capabilities. Our synthesized dataset contains 36k graph-structured tasks across 20 scenarios, achieving a 91% human acceptance rate. Training on our graph-structured data shows that it improves generalization across environments. We conduct multidimensional evaluations for virtual agents, revealing their performance across various capabilities and paving the way for future advancements. Our project is available at https://omni-bench.github.io.
Wendong Bu, Minghe Gao, Bingchen Miao, Zhenkui Zhang, Kaihang Pan, Liyunfei, Mengze Li 0001, Wei Ji 0008, Juncheng Li 0006, Siliang Tang, Yueting Zhuang
ICML4
2025 Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
abstract
The development of Generalist Virtual Agents (GVAs) has shown significant promise in autonomous task execution. However, current training paradigms face critical limitations, including reliance on outcome supervision and labor-intensive human annotations. To address these challenges, we propose Similar, a step-wise multi-dimensional generalist reward model, which offers fine-grained signals for agent training and can choose better actions for inference-time scaling. Specifically, we begin by systematically defining five dimensions for evaluating agent actions. Building on this framework, we design an MCTS-P algorithm to automatically collect and annotate step-wise, five-dimensional agent execution data. Using this data, we train Similar with our crafted Triple-M strategy. Furthermore, we introduce the first benchmark in the virtual agent domain for step-wise, multi-dimensional reward model training and evaluation, named SRM. This benchmark consists of two components: SRMTrain, which serves as the training set for Similar, and SRMEval, a manually selected test set for evaluating the reward model. Experimental results demonstrate that Similar, through its step-wise, multi-dimensional assessment and synergistic gain, provides GVAs with effective intermediate signals during both training and inference-time scaling. The code is available at https://github.com/antgroup/Similar.
Bingchen Miao, Minghe Gao, Wendong Bu, Wenqiao Zhang, Siliang Tang, Tat-Seng Chua, Juncheng Li 0006
ICML3
2025 Counterfactual Evolution of Multimodal Datasets via Visual Programming
abstract
The rapid development of Multimodal Large Language Models (MLLMs) poses increasing demands on the diversity and complexity of multimodal datasets. Yet manual annotation pipelines can no longer keep pace. Existing augmentation methods often follow fixed rules and lack verifiable control over sample diversity and reasoning complexity. To address this, we introduce Scalable COunterfactual Program Evolution (SCOPE), a framework that uses symbolic Visual Programming to guide program evolution via counterfactual reasoning. SCOPE performs the three steps of counterfactual inference: (1) Abduction, by generating verifiable programs to model reasoning associations; (2) Action, by intervening on program structure along three axes—reasoning path, visual context, and cross-instance composition; and (3) Prediction, by categorizing evolved instances by difficulty, structure, and input multiplicity. Based on this process, we build SCOPE-Train and SCOPE-Test, evolving benchmarks with expert validation. To support training, we propose MAP, a curriculum learning strategy that aligns model capacity with sample difficulty. Experiments show that SCOPE improves reasoning performance, exposes model blind spots, and enhances visual dialog capabilities.
Minghe Gao, Zhongqi Yue, Wei Ji 0008, Siliang Tang, Jun Xiao 0001, Tat-Seng Chua, Yueting Zhuang, Juncheng Li 0006
NeurIPS1
2024 An Extended Few-Shot Learning-Based Approach for Histopathological Image Classification of Pan-Cancer in the Digestive System
Md Mamunur Rahaman, Hongzan Sun, Jinzhu Yang, Minghe Gao, Marcin Grzegorzek, Tao Jiang 0014, Xinyu Huang 0003, Chen Li 0022
ADMA (4)6
2024 Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
abstract
Recent advancements in Multimodal Large Language Models (MLLMs) have been utilizing Visual Prompt Generators (VPGs) to convert visual features into tokens that LLMs can recognize. This is achieved by training the VPGs on millions of image-caption pairs, where the VPG-generated tokens of images are fed into a frozen LLM to generate the corresponding captions. However, this image-captioning based training objective inherently biases the VPG to concentrate solely on the primary visual contents sufficient for caption generation, often neglecting other visual details. This shortcoming results in MLLMs’ underperformance in comprehending demonstrative instructions consisting of multiple, interleaved, and multimodal instructions that demonstrate the required context to complete a task. To address this issue, we introduce a generic and lightweight Visual Prompt Generator Complete module (VPG-C), which can infer and complete the missing details essential for comprehending demonstrative instructions. Further, we propose a synthetic discriminative training strategy to fine-tune VPG-C, eliminating the need for supervised demonstrative instructions. As for evaluation, we build DEMON, a comprehensive benchmark for demonstrative instruction understanding. Synthetically trained with the proposed strategy, VPG-C achieves significantly stronger zero-shot performance across all tasks of DEMON. Further evaluation on the MME and OwlEval benchmarks also demonstrate the superiority of VPG-C. The code and models are available at https://github.com/DCDmllm/Cheetah.
Juncheng Li 0006, Kaihang Pan, Zhiqi Ge, Minghe Gao, Wei Ji 0008, Wenqiao Zhang, Tat-Seng Chua, Siliang Tang, Hanwang Zhang, Yueting Zhuang
ICLR4
2024 De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
abstract
Visual programming, a modular paradigm, integrates different modules and Python operators to solve various vision-language tasks. Unlike end-to-end models that need task-specific data, it performs visual processing and inference in an unsupervised manner. Current visual programming methods generate programs in a single pass where the ability to evaluate and optimize based on feedback, unfortunately, is lacking, which consequentially limits their effectiveness for complex, multi-step problems. Inspired by benders decomposition, we introduce De-fine, a training-free framework that automatically decomposes complex tasks into simpler subtasks and refines programs through auto-feedback. This model-agnostic approach can improve logical reasoning performance by integrating the strengths of multiple models. Our experiments across various visual tasks show that De-fine creates more robust programs. Moreover, viewing each feedback module as an independent agent will yield fresh prospects for the field of agent research.
Minghe Gao, Juncheng Li 0006, Hao Fei 0001, Liang Pang 0001, Wei Ji 0008, Guoming Wang, Zheqi Lv, Wenqiao Zhang, Siliang Tang, Yueting Zhuang
ACM Multimedia1
2024 Fact : Teaching MLLMs with Faithful, Concise and Transferable Rationales
abstract
The remarkable performance of Multimodal Large Language Models (MLLMs) has demonstrated their proficient understanding capabilities in handling various visual tasks. Nevertheless, the opaque nature of black-box reasoning processes persists as an enigma, rendering them uninterpretable and struggling with hallucination. Their ability to execute intricate reasoning tasks is also constrained, culminating in stagnation of progression. In this work, we introduce Fact, a novel paradigm designed to generate multimodal rationales that are faithful, concise, and transferable for teaching MLLMs. This paradigm utilizes verifiable visual programming to generate executable code guaranteeing faithfulness. Through a series of operations including pruning, merging, and bridging, the rationale enhances its conciseness. Furthermore, we filter rationales that can be transferred to end-to-end paradigms from programming paradigms to guarantee transferability. Empirical evidence from experiments demonstrates the superiority of Fact across models of varying parameter sizes, significantly enhancing their compositional reasoning and generalization ability and reducing hallucinations owing to its high correlation between images and text.
Minghe Gao, Liang Pang 0001, Yuan Yao 0013, Jisheng Dang, Wenqiao Zhang, Juncheng Li 0006, Siliang Tang, Yueting Zhuang, Tat-Seng Chua
ACM Multimedia1
2023 Predicting PD-L1 status of esophageal cancer from H&E images based on FusedNet model
abstract
For esophageal cancer immunotherapy, Programmed Death Ligand-l (PD-LI) is considered a predictive biomarker. However, immunehistochemistry (IHC) methods used to quantify PD-LI are challenged by high cost, time and variability. In contrast, hematoxylin and eosin (H&E) staining is a reliable method commonly used in cancer diagnosis. By employing advanced deep learning techniques, this study demonstrates the feasibility of predicting PD-LI expression from H&E stained images. With the help of pathologists, a dataset is constructed to evaluate the validity of PD-LI prediction in esophageal cancer by H&E using the FusedNet model. In 227 patients, PD-LI status is systematically predicted. Consistent prediction performance is demonstrated through validation of the validation set, proving that the system can be used as a decision support and quality assurance system in clinical practice.
Minghe Gao, Chen Li 0022, Hechen Yang, Liyu Shi, Yujie Jing, Shuaiyi Tian, Hongzan Sun, Marcin Grzegorzek
IEEE Big Data1
2023 Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language Models
abstract
Prompt tuning, a recently emerging paradigm, enables the powerful vision-language pre-training models to adapt to downstream tasks in a parameter- and data- efficient way, by learning the "soft prompts" to condition frozen pretraining models. Though effective, it is particularly problematic in the few-shot scenario, where prompt tuning performance is sensitive to the initialization and requires a time-consuming process to find a good initialization, thus restricting the fast adaptation ability of the pre-training models. In addition, prompt tuning could undermine the generalizability of the pre-training models, because the learnable prompt tokens are easy to overfit to the limited training samples. To address these issues, we introduce a novel Gradient-RegulAted Meta-prompt learning (GRAM) framework that jointly meta-learns an efficient soft prompt initialization for better adaptation and a lightweight gradient regulating function for strong cross-domain generalizability in a meta-learning paradigm using only the unlabeled image-text pre-training data. Rather than designing a specific prompt tuning method, our GRAM can be easily incorporated into various prompt tuning methods in a model-agnostic way, and comprehensive experiments show that GRAM brings about consistent improvement for them in several settings (i.e., few-shot learning, cross-domain generalization, cross-dataset generalization, etc.) over 11 datasets. Further, experiments show that GRAM enables the orthogonal methods of textual and visual prompt tuning to work in a mutually-enhanced way, offering better generalizability beyond the uni-modal prompt tuning methods.
Juncheng Li 0006, Minghe Gao, Longhui Wei, Siliang Tang, Wenqiao Zhang, Mengze Li 0001, Wei Ji 0008, Qi Tian 0001, Tat-Seng Chua, Yueting Zhuang
ICCV2
2019 Extracting geographic features from the Internet: A geographic information mining framework
Ying Zhang 0010, Qunfei Ma, Yao-Yi Chiang, Craig A. Knoblock, Puhai Yang, Minghe Gao
Knowl. Based Syst.7
2018 Integrated Route and Charging Planning for Electric Vehicles Considering Nonlinear Charging Functions
abstract
Considering the carbon emission reduction and energy saving, the electric vehicles (EVs) have been identified as a promising alternative to classical fossil fuel-driven vehicles. Consequently, the Electric Vehicle Route Planning with Recharging (EVRC) problem extends traditional routing problems to take the limited capacity of the vehicle's battery into consideration. In order to avoid the depletion of battery, planned detours should be introduced to recharge the battery. As the previous research works mainly focused on the three assumptions: (1) EVs will be fully charged when they get a CS; (2) The charging time rises linearly as the battery charge level ascends;(3) Serve time is not considered in the whole process. However, in the actual situations, the location of the used charging stations and the amount of charge are decision variables. In addition, the charging function is not linear. In this work, we propose an integrated route and charging planning for EVs, and take partial charging, nonlinear charging functions as well as serve time into consideration. The experimental result shows that it leads to infeasible or excessively costly solutions when the serve time and the features of partial and nonlinear charging process are ignored.
Qunfei Ma, Minghe Gao, Puhai Yang, Baltabay Aliya
CSCWD4