Mingqian Feng

dblp:335/2516 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 34% Video understanding and tracking · 19% Trustworthy machine learning · 15%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
2.632025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos? · CVPR 2025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Machine learning › Trustworthy machine learning
robustness
1.022025
Discover and Mitigate Multiple Biased Subgroups in Image Classifiers · CVPR 2024
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
Computer vision › Vision and language
video captioning
1.012026
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026
Computer vision › Vision and language › vision-language model › prompt learning
vision-language model prompting
1.012026
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026
Computer vision › Video understanding and tracking › multimodal video understanding › audio-visual video understanding
audio-visual event localization
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Machine learning › Deep learning architectures and training › neural network training › local learning
forward learning
0.912025
FLOPS: Forward Learning with OPtimal Sampling · ICLR 2025
Machine learning › Optimization for machine learning
gradient estimation
0.912025
FLOPS: Forward Learning with OPtimal Sampling · ICLR 2025
Computer vision › 3D vision › multi-view geometry › camera geometry
perspective geometry
0.912025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
0.912025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
Computer vision › Video understanding and tracking › temporal understanding
temporal video understanding
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Machine learning › Optimization for machine learning
variance reduction
0.912025
FLOPS: Forward Learning with OPtimal Sampling · ICLR 2025
Computer vision › Vision and language
video-language model
0.912025
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding · AAAI 2025
Machine learning › Trustworthy machine learning › fairness
bias mitigation
0.812024
Discover and Mitigate Multiple Biased Subgroups in Image Classifiers · CVPR 2024
Machine learning › Trustworthy machine learning
fairness
0.812024
Discover and Mitigate Multiple Biased Subgroups in Image Classifiers · CVPR 2024
Computer vision › Video understanding and tracking
video object segmentation
0.312026
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting · AAAI 2026
Computer vision › Vision and language
cross-modal alignment
0.312025
FLOPS: Forward Learning with OPtimal Sampling · ICLR 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.312025
FLOPS: Forward Learning with OPtimal Sampling · ICLR 2025
Natural language and speech › Language models and text generation
prompt tuning
0.312025
FLOPS: Forward Learning with OPtimal Sampling · ICLR 2025
Computer vision › 3D vision
spatial consistency
0.312025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
Computer vision › Image recognition and object detection
image classification
0.212024
Discover and Mitigate Multiple Biased Subgroups in Image Classifiers · CVPR 2024

Methods — techniques the papers use, named apart from their topics

chain-of-thought prompting · 1.9multimodal LLM fine-tuning · 1.7event-based video clustering · 1.7temporal analysis · 1.0segment anything · 1.0multimodal prompting · 1.0reparameterization · 0.9monte carlo sampling · 0.9benchmark evaluation · 0.9benchmark construction · 0.9
YearPublicationVenuePosition
2026 Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
abstract
In this work, we introduce CAT-V (Caption Anything in Video), a training-free framework for fine-grained object-centric video captioning of user-selected instances. CAT-V combines (i) a SAMURAI-based Segmenter for precise object masks across frames, (ii) a TRACE-Uni Temporal Analyzer for event boundary detection and coarse event descriptions, and (iii) an InternVL-2.5 Captioner that, conditioned on spatiotemporal visual prompts and chain-of-thought (CoT) guidance, produces detailed, temporally coherent captions about object attributes, actions, states, interactions, and context. The system supports point, box, and region prompts and maintains temporal sensitivity by tracking object states across segments. In contrast to vanilla video captioning that is overly abstract and dense video captioning that is often terse, CAT-V enables object-level specificity with spatial accuracy and temporal coherence, without additional training data.
Yunlong Tang 0002, Jing Bi 0002, Chao Huang 0033, Susan Liang, Daiki Shimada, Hang Hua, Yunzhong Xiao, Pinxin Liu, Mingqian Feng, Junjia Guo, Luchuan Song, Ali Vosoughi, Jinxi He, Zeliang Zhang 0001, Jiebo Luo 0001, Chenliang Xu
AAAI10
2026 Video Understanding With Large Language Models: A Survey
abstract
With the rapid growth of online video platforms and the escalating volume of video content, the need for proficient video understanding tools has increased significantly. Given the remarkable capabilities of large language models (LLMs) in language and multimodal tasks, this survey provides a detailed overview of recent advances in video understanding that harness the power of LLMs (Vid-LLMs). The emergent capabilities of Vid-LLMs are surprisingly advanced, particularly their ability for open-ended multi-granularity (abstract, temporal, and spatiotemporal) reasoning combined with common-sense knowledge, suggesting a promising path for future video understanding. We examine the unique characteristics and capabilities of Vid-LLMs, categorizing the approaches into three main types:Video Analyzer × LLM, Video Embedder × LLM, and (Analyzer + Embedder) × LLM. We identify five subtypes based on the functions of LLMs in Vid-LLMs:LLMas Summarizer,LLMas Manager,LLMas Text Decoder,LLMas Regressor, andLLMas Hidden Layer. This survey also presents a comprehensive study of the tasks, datasets, benchmarks, and evaluation methods for Vid-LLMs. Additionally, it explores the extensive applications of Vid-LLMs in various domains, highlighting their remarkable scalability and versatility in real-world video understanding challenges. Additionally, it summarizes the limitations of existing Vid-LLMs and outlines directions for future research. For more information, readers are encouraged to visit the repository at https://github.com/yunlong10/Awesome-LLMs-for-Video-Understanding.
Yunlong Tang 0002, Jing Bi 0002, Siting Xu, Luchuan Song, Susan Liang, Teng Wang 0007, Daoan Zhang, Jie An 0002, Rongyi Zhu, Ali Vosoughi, Chao Huang 0033, Zeliang Zhang 0001, Pinxin Liu, Mingqian Feng, Feng Zheng 0001, Jianguo Zhang 0001, Ping Luo 0002, Jiebo Luo 0001, Chenliang Xu
IEEE Trans. Circuits Syst. Video Technol.15
2025 Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in natural language and multimodal domains. By fine-tuning multimodal LLMs with temporal annotations from well-annotated datasets, e.g., dense video captioning datasets, their temporal understanding capacity in video-language tasks can be obtained. However, there is a notable lack of untrimmed audio-visual video datasets with precise temporal annotations for events. This deficiency hinders LLMs from learning the alignment between time, audio-visual events, and text tokens, thus impairing their ability to localize audio-visual events in videos temporally. To address this gap, we introduce PU-VALOR, a comprehensive audio-visual dataset comprising over 114,081 pseudo-untrimmed videos with detailed temporal annotations. PU-VALOR is derived from the large-scale but coarse-annotated audio-visual dataset VALOR, through a subtle method involving event-based video clustering, random temporal scaling, and permutation. By fine-tuning a multimodal LLM on PU-VALOR, we developed AVicuna, a model capable of aligning audio-visual events with temporal intervals and corresponding text tokens. AVicuna excels in temporal localization and time-aware dialogue capabilities. Our experiments demonstrate that AVicuna effectively handles temporal understanding in audio-visual videos and achieves state-of-the-art performance on open-ended video QA, audio-visual QA, and audio-visual event dense localization tasks.
Yunlong Tang 0002, Daiki Shimada, Jing Bi 0002, Mingqian Feng, Hang Hua, Chenliang Xu
AAAI4
2025 VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
abstract
The advancement of Multimodal Large Language Models (MLLMs) has enabled significant progress in multi-modal understanding, expanding their capacity to analyze video content. However, existing evaluation benchmarks for MLLMs primarily focus on abstract video comprehension, lacking a detailed assessment of their ability to understand video compositions, the nuanced interpretation of how visual elements combine and interact within highly compiled video contexts. We introduce VidComposition, a new benchmark specifically designed to evaluate the video composition understanding capabilities of MLLMs using carefully curated compiled videos and cinematic-level annotations. VidComposition includes 982 videos with 1706 multiple-choice questions, covering various compositional aspects such as camera movement, angle, shot size, narrative structure, character actions and emotions, etc. Our comprehensive evaluation of 33 open-source and proprietary MLLMs reveals a significant performance gap between human and model capabilities. This highlights the limitations of current MLLMs in understanding complex, compiled video compositions and offers insights into areas for further improvement. Our benchmark is publicly available at https://yunlong10.github.io/VidComposition/.
Yunlong Tang 0002, Junjia Guo, Hang Hua, Susan Liang, Mingqian Feng, Rui Mao 0017, Chao Huang 0033, Jing Bi 0002, Zeliang Zhang 0001, Pooyan Fazli, Chenliang Xu
CVPR5
2025 FLOPS: Forward Learning with OPtimal Sampling
abstract
Given the limitations of backpropagation, perturbation-based gradient computation methods have recently gained focus for learning with only forward passes, also referred to as queries. Conventional forward learning consumes enormous queries on each data point for accurate gradient estimation through Monte Carlo sampling, which hinders the scalability of those algorithms. However, not all data points deserve equal queries for gradient estimation. In this paper, we study the problem of improving the forward learning efficiency from a novel perspective: how to reduce the gradient estimation variance with minimum cost? For this, we allocate the optimal number of queries within a set budget during training to balance estimation accuracy and computational efficiency. Specifically, with a simplified proxy objective and a reparameterization technique, we derive a novel plug-and-play query allocator with minimal parameters. Theoretical results are carried out to verify its optimality. We conduct extensive experiments for fine-tuning Vision Transformers on various datasets and further deploy the allocator to two black-box applications: prompt tuning and multimodal alignment for foundation models. All findings demonstrate that our proposed allocator significantly enhances the scalability of forward-learning algorithms, paving the way for real-world applications. The implementation is available at https://github.com/RTkenny/FLOPS-Forward-Learning-with-OPtimal-Sampling.
Tao Ren 0006, Zishi Zhang, Jinyang Jiang 0001, Zeliang Zhang 0001, Mingqian Feng, Yijie Peng
ICLR6
2025 MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
abstract
Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark specifically designed to systematically evaluate MLLMs' understanding of perspective through 10 carefully crafted tasks across three complementary dimensions: Perspective Perception, Reasoning, and Robustness. Our benchmark comprises 2,711 real-world and synthetic image instances with 5,083 question-answer pairs that probe key capabilities, such as vanishing point perception and counting, perspective type reasoning, line relationship understanding in 3D space, invariance to perspective-preserving transformations, etc. Through a comprehensive evaluation of 43 state-of-the-art MLLMs, we uncover significant limitations: while models demonstrate competence on surface-level perceptual tasks, they struggle with compositional reasoning and maintaining spatial consistency under perturbations. Our analysis further reveals intriguing patterns between model architecture, scale, and perspective capabilities, highlighting both robustness bottlenecks and the benefits of chain-of-thought prompting. MMPerspective establishes a valuable testbed for diagnosing and advancing spatial understanding in vision-language systems. Resources are available at https://yunlong10.github.io/MMPerspective/
Yunlong Tang 0002, Pinxin Liu, Mingqian Feng, Zhangyun Tan, Rui Mao 0017, Chao Huang 0033, Jing Bi 0002, Yunzhong Xiao, Susan Liang, Hang Hua, Ali Vosoughi, Luchuan Song, Zeliang Zhang 0001, Chenliang Xu
NeurIPS3
2024 Discover and Mitigate Multiple Biased Subgroups in Image Classifiers
abstract
Machine learning models can perform well on indistribution data but often fail on biased subgroups that are underrepresented in the training data, hindering the robustness of models for reliable applications. Such subgroups are typically unknown due to the absence of subgroup labels. Discovering biased subgroups is the key to understanding models' failure modes and further improving models' robustness. Most previous works of subgroup discovery make an implicit assumption that models only underperform on a single biased subgroup, which does not hold on in-the-wild data where multiple biased subgroups exist. In this work, we propose Decomposition, Interpretation, and Mitigation (DIM), a novel method to address a more challenging but also more practical problem of discovering multiple biased subgroups in image classifiers. Our approach decomposes the image features into multiple components that represent multiple subgroups. This decomposition is achieved via a bilinear dimension reduction method, Partial Least Square (PLS), guided by useful supervision from the image classifier. We further interpret the semantic meaning of each subgroup component by generating natural language descriptions using vision-language foundation models. Finally, DIM mitigates multiple biased subgroups simultaneously via two strategies, including the data and model-centric strategies. Extensive experiments on CIFAR-100 and Breeds datasets demonstrate the effectiveness of DIM in discovering and mitigating multiple biased subgroups. Furthermore, DIM uncovers the failure modes of the classifier on Hard ImageNet, showcasing its broader applicability to understanding model bias in image classifiers. The code is available at https://github.com/ZhangAIPI/DIM.
Zeliang Zhang 0001, Mingqian Feng, Zhiheng Li 0002, Chenliang Xu
CVPR2