Mengnan Liu 0001

dblp:261/3036-1 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0002-3700-3785ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Video understanding and tracking · 40% Segmentation and scene understanding · 40% Vision and language · 21%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › action detection
temporal action localization
1.622025
Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal Transport · CVPR 2025
Stepwise Multi-grained Boundary Detector for Point-Supervised Temporal Action Localization · ECCV (7) 2024
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label generation
0.912025
Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal Transport · CVPR 2025
Computer vision › Vision and language
temporal grounding
0.912025
Moment Quantization for Video Temporal Grounding · ICCV 2025
Computer vision › Segmentation and scene understanding
boundary detection
0.812024
Stepwise Multi-grained Boundary Detector for Point-Supervised Temporal Action Localization · ECCV (7) 2024

Methods — techniques the papers use, named apart from their topics

vector quantization · 0.9optimal transport · 0.9hungarian algorithm · 0.9codebook · 0.9clustering · 0.9DETR · 0.9point supervision · 0.8
YearPublicationVenuePosition
2025 Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal Transport
abstract
Point-supervised Temporal Action Localization poses significant challenges due to the difficulty of identifying complete actions with a single-point annotation per action. Existing methods typically employ Multiple Instance Learning, which struggles to capture global temporal context and requires heuristic post-processing. In research on fully-supervised tasks, DETR-based structures have effectively addressed these limitations. However, it is nontrivial to merely adapt DETR to this task, encountering two major bottlenecks. (1) How to integrate point label information into the model and (2) How to select optimal decoder proposals for training in the absence of complete action segment annotations. To address this issue, we introduce an end-to-end framework by integrating Query Reformation and Optimal Transport (QROT). Specifically, we encode point labels through a set of semantic consensus queries, enabling effective focus on action-relevant snippets. Furthermore, we integrate an optimal transport mechanism to generate high-quality pseudo labels. These pseudo-labels facilitate precise proposals selection based on the Hungarian algorithm, significantly enhancing localization accuracy in point-supervised settings. Extensive experiments on the THUMOS14 and ActivityNet-v1.3 datasets demonstrate that our method outperforms existing MIL-based approaches, offering more stable and accurate temporal action localization in point-level supervision.
Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Xiaolong Sun, Gang Hua 0001
CVPR1
2025 Moment Quantization for Video Temporal Grounding
abstract
Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused on learning continuous features exhibit weak differentiation between foreground and background features. In this paper, we propose a novel Moment-Quantization based Video Temporal Grounding method (MQVTG), which quantizes the input video into various discrete vectors to enhance the discrimination between relevant and irrelevant moments. Specifically, MQVTG maintains a learnable moment codebook, where each video moment matches a codeword. Considering the visual diversity, i.e., various visual expressions for the same moment, MQVTG treats moment-codeword matching as a clustering process without using discrete vectors, avoiding the loss of useful information from direct hard quantization. Additionally, we employ effective prior-initialization and joint-projection strategies to enhance the maintained moment codebook. With its simple implementation, the proposed method can be integrated into existing temporal grounding models as a plug-and-play component. Extensive experiments on six popular benchmarks demonstrate the effectiveness and generalizability of MQVTG, significantly outperforming state-of-the-art methods. Further qualitative analysis shows that our method effectively groups relevant features and separates irrelevant ones, aligning with our goal of enhancing discrimination.
Xiaolong Sun, Le Wang 0003, Sanping Zhou, Liushuai Shi, Mengnan Liu 0001, Gang Hua 0001
ICCV6
2024 Stepwise Multi-grained Boundary Detector for Point-Supervised Temporal Action Localization
Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Qilin Zhang 0004, Gang Hua 0001
ECCV (7)1