EDBT 2026 Demo / reviewers in the wild / expert
Yang Li 0063
dblp:37/4190-63
· DBLP profile ↗
23ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0001-6769-4076ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | URPose: The model with unbiased rectified projection and reconstruction error for monocular unsupervised 3D human pose estimation
Sheng Liu 0013, Yang Li 0063, Sidan Du |
Neurocomputing | 2 |
| 2026 | Enhancing Multi-View Omnidirectional Depth Estimation With Semantic-Aware Cost Aggregation and Spatial PropagationabstractOmnidirectional depth estimation predicts 360-degree depth information using multiple fisheye cameras arranged in a surround-view configuration. However, due to the lack of reference panorama and differences between the predicted depth viewpoint and input cameras, it is challenging to construct and utilize semantic information to improve depth accuracy, resulting in limited accurate in complex regions such as non-overlapping, weak textures, object boundaries and occlusions. This paper proposes a novel model architecture that effectively extracts and leverages semantic information to enhance the accuracy of omnidirectional depth estimation. Specifically, the proposed algorithm combines the variance and mean of multi-view image features to construct the fused matching cost and utilize both geometry and semantic constraints. The model extracts 360-degree semantic context during matching cost aggregation, and predict the corresponding panoramas jointly with omnidirectional depth maps. A semantic-aware spatial propagation module is then employed to further refine the depth estimation. We leverage a multi-scale multi-task learning strategy to supervise the prediction of omnidirectional depth maps and panoramas jointly. The proposed approach achieves state-of-the-art performance on public datasets, and also demonstrates high-precision results on real-world data. The experiments with varying camera configurations validate the generalization ability and flexibility of the algorithm. Ming Li 0069, Xuejiao Hu, Zihang Gao, Sidan Du, Yang Li 0063 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language ModelsabstractThe target of video moment retrieval (VMR) is predicting temporal spans within a video that semantically match a given linguistic query. Existing VMR methods based on multimodal large language models (MLLMs) overly rely on expensive high-quality datasets and time-consuming fine-tuning. Although some recent studies introduce a zero-shot setting to avoid fine-tuning, they overlook inherent language bias in the query, leading to erroneous localization. To tackle the aforementioned challenges, this paper proposes Moment-GPT, a tuning-free pipeline for zero-shot VMR utilizing frozen MLLMs. Specifically, we first employ LLaMA-3 to correct and rephrase the query to mitigate language bias. Subsequently, we design a span generator combined with MiniGPT-v2 to produce candidate spans adaptively. Finally, to leverage the video comprehension capabilities of MLLMs, we apply Video-ChatGPT and span scorer to select the most appropriate spans. Our proposed method substantially outperforms the state-of-the-art MLLM-based and zero-shot models on several public datasets, including QVHighlights, ActivityNet-Captions, and Charades-STA. Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li 0069, Wenxin Liang, Yang Li 0063, Sidan Du |
AAAI | 6 |
| 2025 | HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face FusionabstractRecent advancements in diffusion-based technologies have made significant strides, particularly in identity-preserved portrait generation (IPG). However, when using multiple reference images from the same ID, existing methods typically produce lower-fidelity portraits and struggle to customize face attributes precisely. To address these issues, this paper presents HiFi-Portrait, a high-fidelity method for zero-shot portrait generation. Specifically, we first introduce the face refiner and landmark generator to obtain fine-grained multi-face features and 3D-aware face landmarks. The landmarks include the reference ID and the target attributes. Then, we design HiFi-Net to fuse multi-face features and align them with landmarks, which improves ID fidelity and face control. In addition, we devise an automated pipeline to construct an ID-based dataset for training HiFi-Portrait. Extensive experimental results demonstrate that our method surpasses the SOTA approaches in face similarity and controllability. Furthermore, our method is also compatible with previous SDXL-based works. Yifang Xu, Benxiang Zhai, Yunzhuo Sun, Ming Li 0069, Yang Li 0063, Sidan Du |
CVPR | 5 |
| 2025 | FaceSnap: Enhanced ID-Fidelity Network for Tuning-Free Portrait Customization
Benxiang Zhai, Yifang Xu, Guofeng Zhang 0026, Yang Li 0063, Sidan Du |
ICANN (2) | 4 |
| 2025 | Robust and Flexible Omnidirectional Depth Estimation With Multiple 360-Degree Cameras
Ming Li 0069, Xueqian Jin, Xuejiao Hu, Jinghao Cao, Sidan Du, Yang Li 0063 |
IET Image Process. | 6 |
| 2025 | An efficient action proposal processing approach for temporal action detection
Xuejiao Hu, Jingzhao Dai, Ming Li 0069, Yang Li 0063, Sidan Du |
Neurocomputing | 4 |
| 2025 | ATHENA - Autonomous Vehicle Trajectory Planning Considered Human Action AwarenessabstractLarge language models have brought revolutionary changes to autonomous driving algorithms, ushering them into the era of multimodality. However, existing vehicle trajectory planning methods primarily focus on obstacle avoidance in autonomous driving scenarios, overlooking interactions with entities within the scene, such as humans. In this letter, we propose a new research direction: vehicle trajectory planning that takes into account human actions. We establish ATHENA, the first autonomous driving dataset that integrates multimodal human actions, comprising 33,855 scenarios. Each scenario contains status information about ego vehicle, as well as pedestrian actions that actively or passively interact with the vehicle, such as signaling the vehicle to proceed by waving and the unexpected falls by pedestrians that force the vehicle to stop. Based on each type of interaction, ATHENA also provides the corresponding driving suggestions and the reasons behind them. Moreover, we present an LLM-based baseline that consists of two agents: the Action Understanding Agent and the Vehicle Control Agent. Our baseline implements the generation of driving recommendations and vehicle control functions, which are guided by pedestrian actions. Experiments demonstrate the effectiveness and strong performance of our method. Our dataset and code will be publicly available at https://github.com/dogooooo/ATHENA. Jinghao Cao, Sheng Liu 0013, Chaofan Wu, Yang Li 0063, Sidan Du |
IEEE Signal Process. Lett. | 4 |
| 2024 | A Stereo Matching Method for Specular Objects via Cascaded Network and Joint Supervision
Yongkang Feng, Jianghai Shuai, Pinzhi Wang, Yang Li 0063, Sidan Du |
PRCV (3) | 4 |
| 2024 | CASSC: Context-aware method for depth guided semantic scene completionabstractAbstract Semantic scene completion is a crucial end‐to‐end 3D perception task, and the 3D information perception subjects is vital for autonomous driving. This paper presents CASSC, a novel adaptive context‐aware method based on Transformer networks, aimed at realizing camera‐based semantic scene completion algorithms. The key idea is to leverage rich context information from images to obtain pixel‐level label proposals, followed by designing a multiscale fusion mechanism to merge this information and match it with voxel space. A weakly supervised training strategy is proposed to obtain semantic label distribution features from images and introduce an adaptive multiscale fusion module to fuse and adaptively match these features with voxel space. Here, CASSC achieves state‐of‐the‐art performance on the SemanticKITTI dataset and demonstrates excellent performance on the SSC‐Bench dataset. Ablation experiments validate the rationality and effectiveness of our design, and the model and code of CASSC will be open‐sourced on https://github.com/dogooooo/CASSC . Jinghao Cao, Ming Li 0069, Sheng Liu 0013, Yang Li 0063, Sidan Du |
IET Image Process. | 4 |
| 2024 | Time-attentive fusion network: An efficient model for online detection of action startabstractAbstract Online detection of action start is a significant and challenging task that requires prompt identification of action start positions and corresponding categories within streaming videos. This task presents challenges due to data imbalance, similarity in boundary content, and real‐time detection requirements. Here, a novel Time‐Attentive Fusion Network is introduced to address the requirements of improved action detection accuracy and operational efficiency. The time‐attentive fusion module is proposed, which consists of long‐term memory attention and the fusion feature learning mechanism, to improve spatial‐temporal feature learning. The temporal memory attention mechanism captures more effective temporal dependencies by employing weighted linear attention. The fusion feature learning mechanism facilitates the incorporation of current moment action information with historical data, thus enhancing the representation. The proposed method exhibits linear complexity and parallelism, enabling rapid training and inference speed. This method is evaluated on two challenging datasets: THUMOS’14 and ActivityNet v1.3. The experimental results demonstrate that the proposed method significantly outperforms existing state‐of‐the‐art methods in terms of both detection accuracy and inference speed. Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du |
IET Image Process. | 4 |
| 2024 | ARES: Text-Driven Automatic Realistic Simulator for Autonomous TrafficabstractThe large-scale generation of real-world scenario datasets is a pivotal task in the field of autonomous driving. Existing methods emphasize solely on single-frame rendering, which need complex inputs for continuous scenario rendering. In this letter, ARES: a text-driven automatic realistic simulator is proposed, which can generate extensive realistic datasets with just a single text input. Its core idea is to generate vehicle trajectories based on the textual description, and then render the scenario by vehicle attributes associated with these trajectories. For learning trajectories generating, supervisory signal temporal logic is proposed to assist conditional diffusion model, which incorporates prior physical information. We annotate textual descriptions for KITTI-MOT dataset and establish an objective quantitative evaluation system. The superiority of our method is demonstrated by its high performance, which is reflected in a matching score of 3.54 and an FID of 8.93in the trajectory reconstruction task, along with a speed accuracy of 0.99 and a direction accuracy of 0.93in the trajectory editing task. The scenarios rendered by the proposed method exhibit high quality and realism, which indicates its great potential in testing of autonomous driving algorithms with vehicle-in-the-loop simulations. Jinghao Cao, Sheng Liu 0013, Yang Li 0063, Sidan Du |
IEEE Signal Process. Lett. | 4 |
| 2024 | Distribution-Aware Activity Boundary Representation for Online Detection of Action Start in Untrimmed VideosabstractThe Online Detection of Action Start (ODAS) has attracted the attention of researchers because of its practical applications in areas such as security and emergency response. However, online detection of activity boundaries remains a challenging task due to the inherent ambiguity of boundary definition and the significant imbalance in the number of boundaries and nonboundary points. To address this issue, this study proposes a novel Distribution-aware Activity Boundary Representation (DABR) method that utilizes a continuous probability density function to smooth the probability of moments near activity boundaries. The proposed DABR reduces the penalty for detecting moments near ground-truth boundary points, while increasing the number of samples related to boundary points. Additionally, we introduce a two-stage framework that incorporates class-informed information in temporal localization for more efficient activity boundary localization. Extensive experiments demonstrate that our method achieves state-of-the-art results on two standard datasets, particularly exhibiting a significant improvement of 11.5% at average p-mAP on the THUMOS'14 dataset. Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du |
IEEE Signal Process. Lett. | 4 |
| 2023 | MMDA: Multi-person marginal distribution awareness for monocular 3D pose estimationabstractAbstract Most existing 3D pose representations cannot completely decouple the overlapping two or more human joints of the same type. In this paper, the authors propose a novel 2.5 D representation of the human pose by projecting human joints in 3D space onto the three orthogonal planes. The authors apply for the first time the permutation module to a multi‐person 3D human pose estimation task and use Geometric Constraints Loss (GCL) to guide the learning of the model. The authors overcome the negative effects of the inductive bias of convolutional neural networks (CNNs) by aligning the intermediate feature space with the output feature space. The effectiveness of the authors’ approach is validated on the carnegie mellon university (CMU) panoptic dataset and MuPoTS‐3D dataset. The authors’ proposed representations can effectively decouple the human joints in their selected data from overlapping human joints. Sheng Liu 0013, Jianghai Shuai, Yang Li 0063, Sidan Du |
IET Image Process. | 3 |
| 2023 | The multi-learning for food analyses in computer vision: a survey
Jingzhao Dai, Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du |
Multim. Tools Appl. | 4 |
| 2022 | MODE: Multi-view Omnidirectional Depth Estimation with 360$^\circ $ Cameras
Ming Li 0069, Xueqian Jin, Xuejiao Hu, Jingzhao Dai, Sidan Du, Yang Li 0063 |
ECCV (33) | 6 |
| 2022 | Online human action detection and anticipation in videos: A survey
Xuejiao Hu, Jingzhao Dai, Ming Li 0069, Chenglei Peng, Yang Li 0063, Sidan Du |
Neurocomputing | 5 |
| 2022 | The study of stereo matching optimization based on multi-baseline trinocular model
Chenglei Peng, Ming Li 0069, Yang Li 0063, Sidan Du |
Multim. Tools Appl. | 4 |
| 2021 | A Study of General Data Improvement for Large-Angle Head Pose Estimation
Jue Bai, Chenglei Peng, Zhaoxu Li, Sidan Du, Yang Li 0063 |
CAIP (2) | 5 |
| 2021 | Pyramid Feature Attention Network for Monocular Depth PredictionabstractDeep convolutional neural networks (DCNNs) have achieved great success in monocular depth estimation (MDE). However, few existing works take the contributions for MDE of different levels feature maps into account, leading to inaccurate spatial layout, ambiguous boundaries and discontinuous object surface in the prediction. To better tackle these problems, we propose a Pyramid Feature Attention Network (PFANet) to improve the high-level context features and lowlevel spatial features. In the proposed PFANet, we design a Dual-scale Channel Attention Module (DCAM) to employ channel attention in different scales, which aggregate global context and local information from the high-level feature maps. To exploit the spatial relationship of visual features, we design a Spatial Pyramid Attention Module (SPAM) which can guide the network attention to multi-scale detailed information in the low-level feature maps. Finally, we introduce scale-invariant gradient loss to increase the penalty on errors in depth-wise discontinuous regions. Experimental results show that our method outperforms state-of-the-art methods on the KITTI dataset. Yifang Xu, Chenglei Peng, Ming Li 0069, Yang Li 0063, Sidan Du |
ICME | 4 |
| 2021 | Omnidirectional stereo depth estimation based on spherical deep network
Ming Li 0069, Xuejiao Hu, Jingzhao Dai, Yang Li 0063, Sidan Du |
Image Vis. Comput. | 4 |
| 2017 | Hearing Loss Detection in Medical Multimedia Data by Discrete Wavelet Packet Entropy and Single-Hidden Layer Neural Network Trained by Adaptive Learning-Rate Back Propagation
Shuihua Wang, Sidan Du, Yang Li 0063, Huimin Lu 0001, Ming Yang 0011, Bin Liu 0043, Yudong Zhang 0001 |
ISNN (2) | 3 |
| 2012 | A low complexity fast lattice reduction algorithm for MIMO detectionabstractBased on the well known Lenstra Lenstra Lovász (LLL) algorithm, we propose a possible swap LLL algorithm (P-SLLL) for lattice reduction aided (LRA) MIMO detection in this paper. The reduction process of the new algorithm is modified by searching for the next column swap through the whole basis, instead of the sequential implementation in the original LLL algorithm. Two different searching criteria are proposed, i.e. the random selection criterion and the optimal swap selection criterion. Comparing to the LLL algorithm, the PSLLL algorithm enjoys fast termination property and lower computational complexity, which can benefit practical hardware implementation. Simulation results prove our analysis and show that PSLLL aided linear MIMO detectors achieve the same performance as the LLL aided methods. Kanglian Zhao, Yang Li 0063, Sidan Du |
PIMRC | 2 |