Minghang Zheng

dblp:279/3301 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-1612-975XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Large-Scale Pre-Trained Models Empowering Phrase Generalization in Temporal Sentence Localization
Yang Liu 0105, Minghang Zheng, Qingchao Chen, Shaogang Gong, Yuxin Peng 0001
Int. J. Comput. Vis.2
2026 Cross-Modal Retrieval from Coarse-Grained to Fine-Grained Perspectives: A Survey
Yu-Xin Peng, Minghang Zheng
J. Comput. Sci. Technol.2
2025 Hierarchical Event Memory for Accurate and Low-Latency Online Video Temporal Grounding
Minghang Zheng, Yuxin Peng 0001, Benyuan Sun
ICCV1
2025 InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
Xinhao Cai, Minghang Zheng, Xin Jin 0015, Yang Liu 0105
ACM Multimedia2
2024 Diff-BGM: A Diffusion Model for Video Background Music Generation
abstract
When editing a video, a piece of attractive background music is indispensable. However, video background music generation tasks face several challenges, for example, the lack of suitable training datasets, and the difficulties in flexibly controlling the music generation process and sequentially aligning the video and music. In this work, we first propose a high-quality music-video dataset BGM909 with detailed annotation and shot detection to provide multimodal information about the video and music. We then present evaluation metrics to assess music quality, including music diversity and alignment between music and video with retrieval precision metrics. Finally, we propose the Diff-BGM framework to automatically generate the background music for a given video, which uses different signals to control different aspects of the music during the generation process, i.e., uses dynamic video features to control music rhythm and semantic features to control the melody and atmosphere. We propose to align the video and music sequentially by introducing a segment-aware crossattention layer. Experiments verify the effectiveness of our proposed method. The code and models are available at https://github.com/sizhelee/Diff-BGM.
Minghang Zheng
CVPR3
2024 Training-Free Video Temporal Grounding Using Large-Scale Pre-trained Models
Minghang Zheng, Xinhao Cai, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105
ECCV (82)1
2024 ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
abstract
Visual grounding aims to localize the object referred to in an image based on a natural language query. Although progress has been made recently, accurately localizing target objects within multiple-instance distractions (multiple objects of the same category as the target) remains a significant challenge. Existing methods demonstrate a significant performance drop when there are multiple distractions in an image, indicating an insufficient understanding of the fine-grained semantics and spatial relationships between objects. In this paper, we propose a novel approach, the Relation and Semantic-sensitive Visual Grounding (ResVG) model, to address this issue. Firstly, we enhance the model's understanding of fine-grained semantics by injecting semantic prior information derived from text queries into the model. This is achieved by leveraging text-to-image generation models to produce images representing the semantic attributes of target objects described in queries. Secondly, we tackle the lack of training samples with multiple distractions by introducing a relation-sensitive data augmentation method. This method generates additional training data by synthesizing images containing multiple objects of the same category and pseudo queries based on their spatial relationships. The proposed ReSVG model significantly improves the model's ability to comprehend both object semantics and spatial relations, leading to enhanced performance in visual grounding tasks, particularly in scenarios with multiple-instance distractions. We conduct extensive experiments to validate the effectiveness of our methods on five datasets. Code is available at https://github.com/minghangz/ResVG.
Minghang Zheng, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105
ACM Multimedia1
2023 Phrase-Level Temporal Relationship Mining for Temporal Sentence Localization
abstract
In this paper, we address the problem of video temporal sentence localization, which aims to localize a target moment from videos according to a given language query. We observe that existing models suffer from a sheer performance drop when dealing with simple phrases contained in the sentence. It reveals the limitation that existing models only capture the annotation bias of the datasets but lack sufficient understanding of the semantic phrases in the query. To address this problem, we propose a phrase-level Temporal Relationship Mining (TRM) framework employing the temporal relationship relevant to the phrase and the whole sentence to have a better understanding of each semantic entity in the sentence. Specifically, we use phrase-level predictions to refine the sentence-level prediction, and use Multiple Instance Learning to improve the quality of phrase-level predictions. We also exploit the consistency and exclusiveness constraints of phrase-level and sentence-level predictions to regularize the training process, thus alleviating the ambiguity of each phrase prediction. The proposed approach sheds light on how machines can understand detailed phrases in a sentence and their compositions in their generality rather than learning the annotation biases. Experiments on the ActivityNet Captions and Charades-STA datasets show the effectiveness of our method on both phrase and sentence temporal localization and enable better model interpretability and generalization when dealing with unseen compositions of seen concepts. Code can be found at https://github.com/minghangz/TRM.
Minghang Zheng, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105
AAAI1
2023 Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence Localization
abstract
Video sentence localization aims to locate moments in an unstructured video according to a given natural language query.A main challenge is the expensive annotation costs and the annotation bias.In this work, we study video sentence localization in a zero-shot setting, which learns with only video data without any annotation.Existing zero-shot pipelines usually generate event proposals and then generate a pseudo query for each event proposal.However, their event proposals are obtained via visual feature clustering, which is query-independent and inaccurate; and the pseudo-queries are short or less interpretable.Moreover, existing approaches ignores the risk of pseudo-label noise when leveraging them in training.To address the above problems, we propose a Structurebased Pseudo Label generation (SPL), which first generate free-form interpretable pseudo queries before constructing query-dependent event proposals by modeling the event temporal structure.To mitigate the effect of pseudolabel noise, we propose a noise-resistant iterative method that repeatedly re-weight the training sample based on noise estimation to train a grounding model and correct pseudo labels.Experiments on the ActivityNet Captions and Charades-STA datasets demonstrate the advantages of our approach.Code can be found at https://github.com/minghangz/SPL.
Minghang Zheng, Shaogang Gong, Hailin Jin, Yuxin Peng 0001, Yang Liu 0105
ACL (1)1
2023 Recent Advances in Class-Incremental Learning
Dejie Yang, Minghang Zheng, Weishuai Wang, Yang Liu 0105
ICIG (2)2
2022 Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining
abstract
Video moment localization aims at localizing the video segments which are most related to the given free-form natural language query. The weakly supervised setting, where only video level description is available during training, is getting more and more attention due to its lower annotation cost. Prior weakly supervised methods mainly use sliding windows to generate temporal proposals, which are independent of video content and low quality, and train the model to distinguish matched video-query pairs and unmatched ones collected from different videos, while neglecting what the model needs is to distinguish the unaligned segments within the video. In this work, we propose a novel weakly supervised solution by introducing Contrastive Negative sample Mining (CNM). Specifically, we use a learnable Gaussian mask to generate positive samples, highlighting the video frames most related to the query, and consider other frames of the video and the whole video as easy and hard negative samples respectively. We then train our network with the Intra-Video Contrastive loss to make our positive and negative samples more discriminative. Our method has two advantages: (1) Our proposal generation process with a learnable Gaussian mask is more efficient and makes our positive sample higher quality. (2) The more difficult intra-video negative samples enable our model to distinguish highly confusing scenes. Experiments on two datasets show the effectiveness of our method. Code can be found at https://github.com/minghangz/cnm.
Minghang Zheng, Yanjie Huang, Qingchao Chen, Yang Liu 0105
AAAI1
2022 Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal Learning
abstract
Temporal sentence grounding aims to detect the most salient moment corresponding to the natural language query from untrimmed videos. As labeling the temporal boundaries is labor-intensive and subjective, the weakly- supervised methods have recently received increasing attention. Most of the existing weakly-supervised methods gen-erate the proposals by sliding windows, which are content- independent and of low quality. Moreover, they train their model to distinguish positive visual-language pairs from negative ones randomly collected from other videos, ignoring the highly confusing video segments within the same video. In this paper, we propose Contrastive Proposal Learning(CPL) to overcome the above limitations. Specifi-cally, we use multiple learnable Gaussian functions to gen-erate both positive and negative proposals within the same video that can characterize the multiple events in a long video. Then, we propose a controllable easy to hard neg-ative proposal mining strategy to collect negative samples within the same video, which can ease the model opti-mization and enables CPL to distinguish highly confusing scenes. The experiments show that our method achieves state-of-the-art performance on Charades-STA and Activi-tyNet Captions datasets. The code and models are available at https://github.com/minghangz/cpl.
Minghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105
CVPR1
2022 Phrase-level Prediction for Video Temporal Localization
abstract
Video temporal localization aims to locate a period that semantically matches a natural language query in a given untrimmed video. We empirically observe that although existing approaches gain steady progress on sentence localization, the performance of phrase localization is far from satisfactory. In principle, the phrase should be easier to localize as fewer combinations of visual concepts need to be considered; such incapability indicates that the existing models only capture the sentence annotation bias in the benchmark but lack sufficient understanding of the intrinsic relationship between simple visual and language concepts, thus the model generalization and interpretability is questioned. This paper proposes a unified framework that can deal with both sentence and phrase-level localization, namely Phrase Level Prediction Net (PLPNet). Specifically, based on the hypothesis that similar phrases tend to focus on similar video cues, while dissimilar ones should not, we build a contrastive mechanism to restrain phrase-level localization without fine-grained phrase boundary annotation required in training. Moreover, considering the sentence's flexibility and wide discrepancy among phrases, we propose a clustering-based batch sampler to ensure that contrastive learning can be conducted efficiently. Extensive experiments demonstrate that our method surpasses state-of-the-art methods of phrase-level temporal localization while maintaining high performance in sentence localization and boosting the model's interpretability and generalization capability. Our code is available at https://github.com/sizhelee/PLPNet.
Minghang Zheng, Yang Liu 0105
ICMR3
2021 End-to-End Object Detection with Adaptive Clustering Transformer
Minghang Zheng, Peng Gao 0007, Renrui Zhang, Kunchang Li 0002, Hongsheng Li 0001, Hao Dong 0003
BMVC1
2021 Fast Convergence of DETR with Spatially Modulated Co-Attention
abstract
The recently proposed Detection Transformer (DETR) model successfully applies Transformer to objects detection and achieves comparable performance with two-stage object detection frameworks, such as Faster-RCNN. However, DETR suffers from its slow convergence. Training DETR [4] from scratch needs 500 epochs to achieve a high accuracy. To accelerate its convergence, we propose a simple yet effective scheme for improving the DETR framework, namely Spatially Modulated Co-Attention (SMCA) mechanism. The core idea of SMCA is to conduct location-aware co-attention in DETR by constraining co-attention responses to be high near initially estimated bounding box locations. Our proposed SMCA increases DETR’s convergence speed by replacing the original co-attention mechanism in the decoder while keeping other operations in DETR unchanged. Furthermore, by integrating multi-head and scale-selection attention designs into SMCA, our fully-fledged SMCA can achieve better performance compared to DETR with a dilated convolution-based backbone (45.6 mAP at 108 epochs vs. 43.3 mAP at 500 epochs). We perform extensive ablation studies on COCO dataset to validate SMCA. Code is released at https://github.com/gaopengcuhk/SMCA-DETR.
Peng Gao 0007, Minghang Zheng, Xiaogang Wang 0001, Jifeng Dai, Hongsheng Li 0001
ICCV2