Ming Li 0069

dblp:181/2821-69 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-1341-5585ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Enhancing Multi-View Omnidirectional Depth Estimation With Semantic-Aware Cost Aggregation and Spatial Propagation
abstract
Omnidirectional depth estimation predicts 360-degree depth information using multiple fisheye cameras arranged in a surround-view configuration. However, due to the lack of reference panorama and differences between the predicted depth viewpoint and input cameras, it is challenging to construct and utilize semantic information to improve depth accuracy, resulting in limited accurate in complex regions such as non-overlapping, weak textures, object boundaries and occlusions. This paper proposes a novel model architecture that effectively extracts and leverages semantic information to enhance the accuracy of omnidirectional depth estimation. Specifically, the proposed algorithm combines the variance and mean of multi-view image features to construct the fused matching cost and utilize both geometry and semantic constraints. The model extracts 360-degree semantic context during matching cost aggregation, and predict the corresponding panoramas jointly with omnidirectional depth maps. A semantic-aware spatial propagation module is then employed to further refine the depth estimation. We leverage a multi-scale multi-task learning strategy to supervise the prediction of omnidirectional depth maps and panoramas jointly. The proposed approach achieves state-of-the-art performance on public datasets, and also demonstrates high-precision results on real-world data. The experiments with varying camera configurations validate the generalization ability and flexibility of the algorithm.
Ming Li 0069, Xuejiao Hu, Zihang Gao, Sidan Du, Yang Li 0063
IEEE Trans. Circuits Syst. Video Technol.1
2025 Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
abstract
The target of video moment retrieval (VMR) is predicting temporal spans within a video that semantically match a given linguistic query. Existing VMR methods based on multimodal large language models (MLLMs) overly rely on expensive high-quality datasets and time-consuming fine-tuning. Although some recent studies introduce a zero-shot setting to avoid fine-tuning, they overlook inherent language bias in the query, leading to erroneous localization. To tackle the aforementioned challenges, this paper proposes Moment-GPT, a tuning-free pipeline for zero-shot VMR utilizing frozen MLLMs. Specifically, we first employ LLaMA-3 to correct and rephrase the query to mitigate language bias. Subsequently, we design a span generator combined with MiniGPT-v2 to produce candidate spans adaptively. Finally, to leverage the video comprehension capabilities of MLLMs, we apply Video-ChatGPT and span scorer to select the most appropriate spans. Our proposed method substantially outperforms the state-of-the-art MLLM-based and zero-shot models on several public datasets, including QVHighlights, ActivityNet-Captions, and Charades-STA.
Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li 0069, Wenxin Liang, Yang Li 0063, Sidan Du
AAAI4
2025 HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
abstract
Recent advancements in diffusion-based technologies have made significant strides, particularly in identity-preserved portrait generation (IPG). However, when using multiple reference images from the same ID, existing methods typically produce lower-fidelity portraits and struggle to customize face attributes precisely. To address these issues, this paper presents HiFi-Portrait, a high-fidelity method for zero-shot portrait generation. Specifically, we first introduce the face refiner and landmark generator to obtain fine-grained multi-face features and 3D-aware face landmarks. The landmarks include the reference ID and the target attributes. Then, we design HiFi-Net to fuse multi-face features and align them with landmarks, which improves ID fidelity and face control. In addition, we devise an automated pipeline to construct an ID-based dataset for training HiFi-Portrait. Extensive experimental results demonstrate that our method surpasses the SOTA approaches in face similarity and controllability. Furthermore, our method is also compatible with previous SDXL-based works.
Yifang Xu, Benxiang Zhai, Yunzhuo Sun, Ming Li 0069, Yang Li 0063, Sidan Du
CVPR4
2025 Robust and Flexible Omnidirectional Depth Estimation With Multiple 360-Degree Cameras
Ming Li 0069, Xueqian Jin, Xuejiao Hu, Jinghao Cao, Sidan Du, Yang Li 0063
IET Image Process.1
2025 An efficient action proposal processing approach for temporal action detection
Xuejiao Hu, Jingzhao Dai, Ming Li 0069, Yang Li 0063, Sidan Du
Neurocomputing3
2024 CASSC: Context-aware method for depth guided semantic scene completion
abstract
Abstract Semantic scene completion is a crucial end‐to‐end 3D perception task, and the 3D information perception subjects is vital for autonomous driving. This paper presents CASSC, a novel adaptive context‐aware method based on Transformer networks, aimed at realizing camera‐based semantic scene completion algorithms. The key idea is to leverage rich context information from images to obtain pixel‐level label proposals, followed by designing a multiscale fusion mechanism to merge this information and match it with voxel space. A weakly supervised training strategy is proposed to obtain semantic label distribution features from images and introduce an adaptive multiscale fusion module to fuse and adaptively match these features with voxel space. Here, CASSC achieves state‐of‐the‐art performance on the SemanticKITTI dataset and demonstrates excellent performance on the SSC‐Bench dataset. Ablation experiments validate the rationality and effectiveness of our design, and the model and code of CASSC will be open‐sourced on https://github.com/dogooooo/CASSC .
Jinghao Cao, Ming Li 0069, Sheng Liu 0013, Yang Li 0063, Sidan Du
IET Image Process.2
2024 Time-attentive fusion network: An efficient model for online detection of action start
abstract
Abstract Online detection of action start is a significant and challenging task that requires prompt identification of action start positions and corresponding categories within streaming videos. This task presents challenges due to data imbalance, similarity in boundary content, and real‐time detection requirements. Here, a novel Time‐Attentive Fusion Network is introduced to address the requirements of improved action detection accuracy and operational efficiency. The time‐attentive fusion module is proposed, which consists of long‐term memory attention and the fusion feature learning mechanism, to improve spatial‐temporal feature learning. The temporal memory attention mechanism captures more effective temporal dependencies by employing weighted linear attention. The fusion feature learning mechanism facilitates the incorporation of current moment action information with historical data, thus enhancing the representation. The proposed method exhibits linear complexity and parallelism, enabling rapid training and inference speed. This method is evaluated on two challenging datasets: THUMOS’14 and ActivityNet v1.3. The experimental results demonstrate that the proposed method significantly outperforms existing state‐of‐the‐art methods in terms of both detection accuracy and inference speed.
Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du
IET Image Process.3
2024 Distribution-Aware Activity Boundary Representation for Online Detection of Action Start in Untrimmed Videos
abstract
The Online Detection of Action Start (ODAS) has attracted the attention of researchers because of its practical applications in areas such as security and emergency response. However, online detection of activity boundaries remains a challenging task due to the inherent ambiguity of boundary definition and the significant imbalance in the number of boundaries and nonboundary points. To address this issue, this study proposes a novel Distribution-aware Activity Boundary Representation (DABR) method that utilizes a continuous probability density function to smooth the probability of moments near activity boundaries. The proposed DABR reduces the penalty for detecting moments near ground-truth boundary points, while increasing the number of samples related to boundary points. Additionally, we introduce a two-stage framework that incorporates class-informed information in temporal localization for more efficient activity boundary localization. Extensive experiments demonstrate that our method achieves state-of-the-art results on two standard datasets, particularly exhibiting a significant improvement of 11.5% at average p-mAP on the THUMOS'14 dataset.
Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du
IEEE Signal Process. Lett.3
2023 The multi-learning for food analyses in computer vision: a survey
Jingzhao Dai, Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du
Multim. Tools Appl.3
2022 MODE: Multi-view Omnidirectional Depth Estimation with 360$^\circ $ Cameras
Ming Li 0069, Xueqian Jin, Xuejiao Hu, Jingzhao Dai, Sidan Du, Yang Li 0063
ECCV (33)1
2022 Online human action detection and anticipation in videos: A survey
Xuejiao Hu, Jingzhao Dai, Ming Li 0069, Chenglei Peng, Yang Li 0063, Sidan Du
Neurocomputing3
2022 The study of stereo matching optimization based on multi-baseline trinocular model
Chenglei Peng, Ming Li 0069, Yang Li 0063, Sidan Du
Multim. Tools Appl.3
2021 Pyramid Feature Attention Network for Monocular Depth Prediction
abstract
Deep convolutional neural networks (DCNNs) have achieved great success in monocular depth estimation (MDE). However, few existing works take the contributions for MDE of different levels feature maps into account, leading to inaccurate spatial layout, ambiguous boundaries and discontinuous object surface in the prediction. To better tackle these problems, we propose a Pyramid Feature Attention Network (PFANet) to improve the high-level context features and lowlevel spatial features. In the proposed PFANet, we design a Dual-scale Channel Attention Module (DCAM) to employ channel attention in different scales, which aggregate global context and local information from the high-level feature maps. To exploit the spatial relationship of visual features, we design a Spatial Pyramid Attention Module (SPAM) which can guide the network attention to multi-scale detailed information in the low-level feature maps. Finally, we introduce scale-invariant gradient loss to increase the penalty on errors in depth-wise discontinuous regions. Experimental results show that our method outperforms state-of-the-art methods on the KITTI dataset.
Yifang Xu, Chenglei Peng, Ming Li 0069, Yang Li 0063, Sidan Du
ICME3
2021 Omnidirectional stereo depth estimation based on spherical deep network
Ming Li 0069, Xuejiao Hu, Jingzhao Dai, Yang Li 0063, Sidan Du
Image Vis. Comput.1