Mingyao Zhou

dblp:82/9542 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric Network
abstract
With the exponential growth of video content, aiming at localizing relevant video moments based on natural language queries, video moment retrieval (VMR) has gained significant attention. Existing weakly supervised VMR methods focus on designing various feature modeling and modal interaction modules to alleviate the reliance on precise temporal annotations. However, these methods have poor generalization capabilities on compositional queries with novel syntactic structures or vocabulary in real-world scenarios. To this end, we propose a new task: weakly supervised compositional moment retrieval (WSCMR). This task trains models using only video-query pairs without precise temporal annotations, while enabling generalization to complex compositional queries. Furthermore, a proposal-centric network (PC-Net) is proposed to tackle this challenging task. First, video and query features are extracted through frozen feature extractors, followed by modality interaction to obtain multimodal features. Second, to handle compositional queries with explicit temporal associations, a dual-granularity proposal generator decodes multimodal global and frame-level features to obtain query-relevant proposal boundaries with fine-grained temporal perception. Third, to improve the discrimination of proposal features, a proposal feature aggregator is constructed to conduct semantic alignment of frames and queries, and employ a learnable peak-aware Gaussian distributor to fit the frame weights within the proposals to derive proposal features from the video frame features. Finally, the proposal quality is assessed based on the results of reconstructing the masked query using the obtained proposal features. To further enhance the model's ability to capture semantic associations between proposals and queries, a quality margin regularizer is constructed to dynamically stratify proposals into high and low query-relevance subsets and enhance the association between queries and common elements within proposals, and suppress spurious correlations via inter-subset contrastive learning. Notably, PC-Net achieves superior performance with 54\% fewer parameters than prior works by parameter-efficient design. Experiments on Charades-CG and ActivityNet-CG demonstrate PC-Net’s ability to generalize across diverse compositional queries. Code is available at https://github.com/mingyao1120/PC-Net.
Mingyao Zhou, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Chengji Wang, Mang Ye
NeurIPS1
2025 Two-layer tensor decomposition for temporal kge graph completion
Lupeng Yue, Kaisheng Zeng, Mingyao Zhou, Jian Wan 0001
Expert Syst. Appl.6
2024 TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
abstract
Video moment retrieval (MR) and highlight detection (HD) based on natural language queries are two highly related tasks, which aim to obtain relevant moments within videos and highlight scores of each video clip. Recently, several methods have been devoted to building DETR-based networks to solve both MR and HD jointly. These methods simply add two separate task heads after multi-modal feature extraction and feature interaction, achieving good performance. Nevertheless, these approaches underutilize the reciprocal relationship between two tasks. In this paper, we propose a task-reciprocal transformer based on DETR (TR-DETR) that focuses on exploring the inherent reciprocity between MR and HD. Specifically, a local-global multi-modal alignment module is first built to align features from diverse modalities into a shared latent space. Subsequently, a visual feature refinement is designed to eliminate query-irrelevant information from visual features for modal interaction. Finally, a task cooperation module is constructed to refine the retrieval pipeline and the highlight score prediction process by utilizing the reciprocity between MR and HD. Comprehensive experiments on QVHighlights, Charades-STA and TVSum datasets demonstrate that TR-DETR outperforms existing state-of-the-art methods. Codes are available at https://github.com/mingyao1120/TR-DETR.
Hao Sun 0014, Mingyao Zhou, Wenjing Chen 0003, Wei Xie 0008
AAAI2
2024 Cross-Modal Multiscale Difference-Aware Network for Joint Moment Retrieval and Highlight Detection
abstract
Since the goals of both Moment Retrieval (MR) and Highlight Detection (HD) are to quickly obtain the required content from the video according to user needs, several works have attempted to take advantage of the commonality between both tasks to design transformer-based networks for joint MR and HD. Although these methods achieve impressive performance, they still face some problems: a) Semantic gaps across different modalities. b) Various durations of different query-relevant moments and highlights. c) Smooth transitions among diverse events. To this end, we propose a Cross-modal Multiscale Difference-aware Network, named CMDNet. First, a clip-text alignment module is constructed to narrow semantic gaps between different modalities. Second, a multiscale difference perception module is utilized to mine the differential information between adjacent clips and perform multiscale modeling to obtain discriminative representations. Finally, these representations are fed into the MR and HD task heads to retrieve relevant moments and estimate highlight scores precisely. Extensive experiments on three popular datasets demonstrate that CMDNet achieves state-of-the-art performance.
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008
ICASSP1
2024 A novel device placement approach based on position-aware subgraph neural networks
Meiting Xue, Mingyao Zhou
Neurocomputing6
2024 Complex expressional characterizations learning based on block decomposition for temporal knowledge graph completion
Lupeng Yue, Kaisheng Zeng, Jian Wan 0001, Mingyao Zhou
Knowl. Based Syst.7
2024 Query-aware multi-scale proposal network for weakly supervised temporal sentence grounding in videos
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Xiaoqiang Lu
Knowl. Based Syst.1
2023 FedECCR: Federated Learning Method with Encoding Comparison and Classification Rectification
Xin Wang 0157, Beibei Zhang 0007, Mingyao Zhou
CollaborateCom (3)5
2023 Analysis of Performance and Optimization in MindSpore on Ascend NPUs
abstract
With the rapid advancement of artificial intelligence, the complexity and depth of deep neural networks continues to grow, placing higher demands on computational power. To meet these requirements, various manufacturers have developed specialized computing processors for the training process of deep learning, such as Huawei’s Ascend Neural Processing Unit (NPU). In order to fully leverage the capabilities of the Ascend NPU, Huawei has introduced the MindSpore deep learning framework. The computational performance of deep learning frameworks plays a critical role for developers. However, there is a lack of comprehensive research on the analysis of performance and optimization in MindSpore framework on the Ascend NPU, leading to a scarcity of relevant references for deep learning development utilizing MindSpore on the Ascend NPU. To address this gap, this study examined the performance and optimization of MindSpore on Ascend NPUs through detailed experiments involving diverse workloads and multiple analysis metrics. The analysis is conducted at three levels: operations, models, and techniques, investigating the operations configurations, performance bottleneck, along with techniques selections, within the training process of DNNs. Consequently, this study offers significant insights and guidance for DL researchers and practitioners using MindSpore on Ascend NPUs.
Bangchuan Wang, Chuying Yang, Mingyao Zhou, Nenggan Zheng
ICPADS5
2023 An Auto-Parallel Method for Deep Learning Models Based on Genetic Algorithm
abstract
As the size of datasets and neural network models increases, automatic parallelization methods for models have become a research hotspot in recent years. The existing auto-parallel methods based on machine learning or graph algorithms still have issues with search efficiency and applicability. This paper proposes an automatic parallel method based on a dual-population genetic algorithm, TGA, which transforms model partitioning and placement into an integer linear programming problem and constructs a cost model to evaluate the solution. The solution space is built using the neural network’s dataflow graph and device cluster’s topology, and the dual-population genetic algorithm is used to search for the optimal model parallel strategy. Experiments with various models show that the proposed method can improve single-step execution time by up to 42% compared to the Baechi method and up to 37.7% compared to the Hierarchical method.
Chengchuang Huang, Yijie Ni, Chunbao Zhou, Jue Wang 0013, Mingyao Zhou, Meiting Xue, Yunquan Zhang
ICPADS7
2023 Performance Evaluation of MindSpore and PyTorch Based on Ascend NPU
abstract
The development of deep learning depends on the support of deep learning processing units and frameworks. HUAWEI's Ascend Neural Processing Unit (Ascend NPU), as a chip designed specifically for neural network computation acceleration, not only provides support for its self-developed framework MindSpore, but also provides adaptation for PyTorch. However, there is a lack of comparative evaluation research on MindSpore and other frameworks on Ascend NPU, making it difficult to understand their actual performance. Therefore, comprehensive evaluation experiments in both model level and operator level on MindSpore and PyTorch based on Ascend NPU are conducted in this paper, analyzing and comparing the performance of the two frameworks. With the conclusions analyzed on the model evaluation experiments and operator evaluation experiments, we provide references and suggestions for deep learning researchers and practitioners in terms of framework selection and optimization strategies for Ascend NPU.
Zeling Zhu, Bangchuan Wang, Chuying Yang, Mingyao Zhou, Nenggan Zheng
ICPADS5
2021 Computing infrastructure construction and optimization for high-performance computing and artificial intelligence
Jipeng Zhou, Jiangyong Ying, Mingyao Zhou
CCF Trans. High Perform. Comput.4