VLDB 2026 Research / reviewers in the wild / expert
Mengxue Qu
dblp:325/4683
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-9432-0205ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Vision and language · 26% Segmentation and scene understanding · 19% Learning paradigms · 15% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding › 3d segmentation
3d referring expression segmentation |
1.0 | 1 | 2026 | 3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation · IEEE Trans. Multim. 2026 |
Machine learning › Learning paradigms
semi-supervised learning |
1.0 | 1 | 2026 | 3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation · IEEE Trans. Multim. 2026 |
Computer vision › 3D vision
3d object detection |
0.9 | 1 | 2025 | Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention · ICLR 2025 |
Computer vision › 3D vision › 3d scene understanding
3d visual grounding |
0.9 | 1 | 2025 | Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention · ICLR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | ReCot: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language Models · ICCV 2025 |
Computer vision › Vision and language › video grounding
spatio-temporal video grounding |
0.9 | 1 | 2025 | Single-Frame Supervision for Spatio-Temporal Video Grounding · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Learning paradigms
weakly supervised learning |
0.9 | 1 | 2025 | Single-Frame Supervision for Spatio-Temporal Video Grounding · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent reasoning
intention inference |
0.7 | 1 | 2023 | RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments · NeurIPS 2023 |
Computer vision › Vision and language
multimodal reasoning |
0.7 | 1 | 2023 | RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments · NeurIPS 2023 |
Computer vision › Image recognition and object detection
object detection |
0.7 | 1 | 2023 | RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments · NeurIPS 2023 |
Computer vision › Segmentation and scene understanding › interactive segmentation
point-based segmentation |
0.7 | 1 | 2023 | Learning to Segment Every Referring Object Point by Point · CVPR 2023 |
Computer vision › Segmentation and scene understanding
referring image segmentation |
0.7 | 1 | 2023 | Learning to Segment Every Referring Object Point by Point · CVPR 2023 |
Machine learning › Learning theory › online learning
sequence prediction |
0.7 | 1 | 2023 | Learning to Segment Every Referring Object Point by Point · CVPR 2023 |
Computer vision › Vision and language
visual grounding |
0.6 | 1 | 2022 | SiRi: A Simple Selective Retraining Mechanism for Transformer-Based Visual Grounding · ECCV (35) 2022 |
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection |
0.2 | 1 | 2023 | RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
teacher-student learning · 1.0pseudo-label filtering · 1.0dynamic weighting · 1.0reflective self-correction · 0.9multiple instance learning · 0.9language-based 3d object detection · 0.9fine-tuning · 0.9curriculum learning · 0.9cross-attention · 0.9cascaded adaptive learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentationabstract3D Referring Expression Segmentation (3D-RES) typically requires extensive instance-level annotations for fully supervised learning, a process that is both time-consuming and costly. Semi-supervised learning (SSL) can address this by using a small amount of labeled data and a large amount of unlabeled data, improving performance while reducing the annotation costs. SSL adopts a teacher-student learning paradigm, where the teacher produces pseudo-labels to guide the student, often using high-confidence threshold filtering to improve pseudo-label quality. However, in the context of 3D-RES, where each label corresponds to a single mask and labeled data is scarce, existing SSL methods treat high-quality pseudo-labels merely as auxiliary supervision, which hinders their ability to fully boost the model's learning potential. The reliance on high-confidence thresholds for filtering often results in potentially valuable pseudo-labels being discarded, restricting the model's ability to leverage the abundant unlabeled data. Therefore, we identify two critical challenges in semi-supervised 3D-RES, namely, inefficient utilization of highquality pseudo-labels and wastage of useful information from lowquality pseudo-labels. In this paper, we introduce the first semisupervised learning framework for 3D-RES, presenting a robust baseline method named 3DResT. To address these challenges, we propose two novel designs called Teacher-Student Consistency- Based Sampling (TSCS) and Quality-Driven Dynamic Weighting (QDW). TSCS aids in the selection of high-quality pseudo-labels, integrating them into the labeled dataset to strengthen the labeled supervision signals. QDW preserves low-quality pseudo-labels by dynamically assigning them lower weights, allowing for the effective extraction of useful information rather than discarding them. Extensive experiments conducted on the widely used benchmark demonstrate the effectiveness of our method. Notably, with only 1 points compared to the fully supervised method. Code will be available at:https://github.com/Wind010321/3DResT Mengxue Qu, Weitai Kang, Yan Yan 0002, Yao Zhao 0001, Yunchao Wei |
IEEE Trans. Multim. | 2 |
| 2025 | ReCot: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language Models
Mengxue Qu, Kunyang Han, Yunchao Wei, Yao Zhao 0001 |
ICCV | 1 |
| 2025 | Intent3D: 3D Object Detection in RGB-D Scans Based on Human IntentionabstractIn real-life scenarios, humans seek out objects in the 3D world to fulfill their daily needs or intentions. This inspires us to introduce 3D intention grounding, a new task in 3D object detection employing RGB-D, based on human intention, such as "I want something to support my back." Closely related, 3D visual grounding focuses on understanding human reference. To achieve detection based on human intention, it relies on humans to observe the scene, reason out the target that aligns with their intention ("pillow" in this case), and finally provide a reference to the AI system, such as "A pillow on the couch". Instead, 3D intention grounding challenges AI agents to automatically observe, reason and detect the desired target solely based on human intention. To tackle this challenge, we introduce the new Intent3D dataset, consisting of 44,990 intention texts associated with 209 fine-grained classes from 1,042 scenes of the ScanNet dataset. We also establish several baselines based on different language-based 3D object detection models on our benchmark. Finally, we propose IntentNet, our unique approach, designed to tackle this intention-based detection problem. It focuses on three key aspects: intention understanding, reasoning to identify object candidates, and cascaded adaptive learning that leverages the intrinsic priority logic of different losses for multiple objective optimization. Weitai Kang, Mengxue Qu, Jyoti Kini, Yunchao Wei, Mubarak Shah, Yan Yan 0002 |
ICLR | 2 |
| 2025 | A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly AnalysisabstractMost video-anomaly research stops at frame-wise detection, offering little insight into why an event is abnormal, typically outputting only frame-wise anomaly scores without spatial or semantic context. Recent video anomaly localization and video anomaly understanding methods improve explainability but remain data-dependent and task-specific. We propose a unified reasoning framework that bridges the gap between temporal detection, spatial localization, and textual explanation. Our approach is built upon a chained test-time reasoning process that sequentially connects these tasks, enabling holistic zero-shot anomaly analysis without any additional training. Specifically, our approach leverages intra-task reasoning to refine temporal detections and inter-task chaining for spatial and semantic understanding, yielding improved interpretability and generalization in a fully zero-shot manner. Without any additional data or gradients, our method achieves state-of-the-art zero-shot performance across multiple video anomaly detection, localization, and explanation benchmarks. The results demonstrate that careful prompt design with task-wise chaining can unlock the reasoning power of foundation models, enabling practical, interpretable video anomaly analysis in a fully zero-shot manner. Project Page: https://rathgrith.github.io/Unified_Frame_VAA/. Dongheng Lin, Mengxue Qu, Kunyang Han, Jianbo Jiao, Xiaojie Jin 0004, Yunchao Wei |
NeurIPS | 2 |
| 2025 | RelFormer: Advancing contextual relations for transformer-based dense captioning
Weiqi Jin, Mengxue Qu, Caijuan Shi, Yao Zhao 0001, Yunchao Wei |
Comput. Vis. Image Underst. | 2 |
| 2025 | Single-Frame Supervision for Spatio-Temporal Video GroundingabstractSpatio-Temporal Video Grounding (STVG) aims at localizing the spatio-temporal tube of a specific object in an untrimmed video given a free-form natural language query. As the annotation of tubes is labor intensive, researchers are motivated to explore weakly supervised approaches in recent works, which usually results in significant performance degradation. To achieve a less expensive STVG method with acceptable accuracy, this work investigates the "single-frame supervision" paradigm that requires a single frame labeled with a bounding box within the temporal boundary of the fully supervised counterpart as the supervisory signal. Based on the characteristics of the STVG problem, we propose a Two-Stage Multiple Instance Learning (T-SMILE) method, which creates pseudo labels by expanding the annotated frame to its contextual frames, thereby establishing a fully-supervised problem to facilitate further model training. The innovations of the proposed method are three-folded, including 1) utilizing multiple instance learning to dynamically select instances in positive bags for the recognition of starting and ending timestamps, 2) learning highly discriminative query features by incorporating spatial prior constraints in cross-attention, and 3) designing a curriculum learning-based strategy that iterative assigns dynamic weights to spatial and temporal branches, thereby gradually adapting to the learning branch with larger difficulty. To facilitate future research on this task, we also contribute a large-scale benchmark containing 12,469 videos on complex scenes with single-frame annotation. The extensive experiments on two benchmarks demonstrate that T-SMILE significantly outperforms all weakly-supervised methods. Remarkably, it also performs better than some fully-supervised methods associated with much more annotation labor costs. Kun Liu 0016, Mengxue Qu, Yang Liu 0235, Yunchao Wei, Wenming Zhe, Yao Zhao 0001, Wu Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | AdGPT: Explore Meaningful Advertising with ChatGPTabstractAdvertising is pervasive in everyday life. Some advertisements are not as readily comprehensible, as they convey a deeper message or purpose, which is referred to as “meaningful advertising.” These ads often aim to create an emotional connection with the audience or promote a social cause. Developing a method for automatically understanding meaningful advertising would be advantageous for the dissemination and creation of such ads. However, current models of ad understanding primarily focus on the superficial aspects of images. In this article, we introduce AdGPT, a model that leverages visual expert analysis to guide Large Language Models (LLMs) in generating adaptive reasoning chains. Informed by these chains of thought, the model can intelligently comprehend meaningful ads regarding category, content, and sentiment. To assess the effectiveness of our approach, we extract a subset of meaningful ads from the widely used Pitt’s ad images for analysis. Beyond employing traditional ad understanding metrics to evaluate the LLMs’ comprehensive ad comprehension, we also develop a novel generative metric that aligns with user study evaluations for consistent performance assessment. Experiments show that our methods outperform existing state-of-the-art (SOTA) approaches directly linking visual expert models and LLMs and large-scale visual-language models. Code is available at https://github.com/Rbrq03/AdGPT . Jiannan Huang 0002, Mengxue Qu, Yunchao Wei |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Learning to Segment Every Referring Object Point by PointabstractReferring Expression Segmentation (RES) can facilitate pixel-level semantic alignment between vision and language. Most of the existing RES approaches require massive pixel-level annotations, which are expensive and exhaustive. In this paper, we propose a new partially supervised training paradigm for RES, i.e., training using abundant referring bounding boxes and only a few (e.g., 1%) pixel-level referring masks. To maximize the transferability from the REC model, we construct our model based on the point-based sequence prediction model. We propose the co-content teacher-forcing to make the model explicitly associate the point coordinates (scale values) with the referred spatial features, which alleviates the exposure bias caused by the limited segmentation masks. To make the most of referring bounding box annotations, we further propose the resampling pseudo points strategy to select more accurate pseudo-points as supervision. Extensive experiments show that our model achieves 52.06% in terms of accuracy (versus 58.93% in fully supervised setting) on Re-fCOCO+@testA, when only using 1% of the mask annotations. Code is available at https://github.com/qumengxue/Partial-RES.git. Mengxue Qu, Yu Wu 0011, Yunchao Wei, Wu Liu 0005, Xiaodan Liang, Yao Zhao 0001 |
CVPR | 1 |
| 2023 | RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open EnvironmentsabstractIntention-oriented object detection aims to detect desired objects based on specific intentions or requirements. For instance, when we desire to "lie down and rest", we instinctively seek out a suitable option such as a "bed" or a "sofa" that can fulfill our needs. Previous work in this area is limited either by the number of intention descriptions or by the affordance vocabulary available for intention objects. These limitations make it challenging to handle intentions in open environments effectively. To facilitate this research, we construct a comprehensive dataset called Reasoning Intention-Oriented Objects (RIO). In particular, RIO is specifically designed to incorporate diverse real-world scenarios and a wide range of object categories. It offers the following key features: 1) intention descriptions in RIO are represented as natural sentences rather than a mere word or verb phrase, making them more practical and meaningful; 2) the intention descriptions are contextually relevant to the scene, enabling a broader range of potential functionalities associated with the objects; 3) the dataset comprises a total of 40,214 images and 130,585 intention-object pairs. With the proposed RIO, we evaluate the ability of some existing models to reason intention-oriented objects in open environments. Mengxue Qu, Yu Wu 0011, Wu Liu 0005, Xiaodan Liang, Jingkuan Song, Yao Zhao 0001, Yunchao Wei |
NeurIPS | 1 |
| 2022 | SiRi: A Simple Selective Retraining Mechanism for Transformer-Based Visual Grounding
Mengxue Qu, Yu Wu 0011, Wu Liu 0005, Qiqi Gong, Xiaodan Liang, Olga Russakovsky, Yao Zhao 0001, Yunchao Wei |
ECCV (35) | 1 |