EDBT 2026 Demo / reviewers in the wild / expert
Zehai Niu
dblp:304/1151
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0003-1732-8627ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GS2Physics: Semantic-Region-Aware Gaussian Splatting for Physical Property PredictionabstractPredicting the physical properties of reconstructed 3D assets is essential for virtual reality interactions. However, current systems often depend on manually assigning properties such as stiffness and density, which can be inefficient and prone to errors. To address this issue, we present GS2Physics, a novel framework based on 3D Gaussian Splatting. This framework is designed to predict physical properties accurately while maintaining improved consistency in semantic segmentation. Unlike existing approaches, which either struggle with region inconsistency or misalign semantic 3D features, GS2Physics embeds semantic-region-aware features directly into the Gaussian Splatting representation. This allows for region-consistent and accurate physical property prediction, achieving state-of-the-art performance on the ABO-500 mass prediction benchmark. To further evaluate our segmentation capabilities, we introduce PhysSeg-15, a subset dataset of ABO-500 featuring physical property segmentation masks for 15 different 3D objects captured from five viewpoints. Our method significantly outperforms existing approaches in segmentation accuracy. Qualitative results demonstrate more consistent material predictions across different object regions and improved accuracy in physical property prediction. In addition, we showcase the effectiveness of GS2Physics in 3D interaction tasks, where our predicted physical properties result in more realistic object motion. Our dataset and results are available at https://github.com/momaiyc/GS2Physics. Bin Huang 0016, Jiayi Lyu, Zehai Niu, LinLin Shen, Jinbao Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Text-to-Any-Skeleton Motion Generation Without Retargeting
Qingyuan Liu 0001, Ke Lu 0002, Kun Dong 0001, Jian Xue 0002, Zehai Niu, Jinbao Wang 0001 |
ICCV | 5 |
| 2025 | FoodSAM: Any Food SegmentationabstractIn this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generated masks, we propose a novel framework, calledFoodSAM. This innovative approach integrates the coarse semantic mask with SAM-generated masks to enhance semantic segmentation quality. Besides, we recognize that the ingredients in food can be supposed as independent individuals, which motivated us to perform instance segmentation on food images. Furthermore, FoodSAM extends its zero-shot capability to encompass panoptic segmentation by incorporating an object detector, which renders FoodSAM to effectively capture non-food object information. Drawing inspiration from the recent success of promptable segmentation, we also extend FoodSAM to promptable segmentation, supporting various prompt variants. Consequently, FoodSAM emerges as an all-encompassing solution capable of segmenting food items at multiple levels of granularity. Remarkably, this pioneering framework stands as the first-ever work to achieve instance, panoptic, and promptable segmentation on food images. Extensive experiments demonstrate the feasibility and impressing performance of FoodSAM, validating SAM's potential as a prominent and influential tool within the domain of food image segmentation. Xing Lan, Jiayi Lyu, Hanyu Jiang 0004, Kun Dong 0001, Zehai Niu, Yi Zhang 0162, Jian Xue 0002 |
IEEE Trans. Multim. | 5 |
| 2024 | VS3D: A Vote-Based Semi-Supervised 3D Object Detection Framework for Point CloudsabstractIn recent years, the 3D object detection method has undergone rapid evolution, heavily relying on substantial amounts of high-quality labeled data. However, the process of annotating 3D data is both time-consuming and costly. In response to this challenge, we propose a vote-based semi-supervised 3D object detection framework called VS3D. First, a data augmentation technique named Random Grid Deleting (RGD) is proposed to detect occluded objects and small objects more robustly. Then, an auxiliary branch with Voting Consistency Learning (VCL) is added to predict object centers more accurately. Additionally, a Teacher-Student Matching (TSM) module with stricter consistency constraints is designed to accelerate network convergence and improve detection performance. Our method can integrate any vote-based fully supervised network seamlessly. Extensive experiments on SUN RGB-D and ScanNet V2 datasets demonstrate that the proposed method outperforms the state-of-the-art fully supervised model when using only 70% labeled data. Ke Lu 0002, Yang Zhao 0028, Hengsheng Lun, Zehai Niu, Jian Xue 0002 |
ICME | 5 |
| 2024 | Realistic Full-Body Motion Generation from Sparse Tracking with State Space ModelabstractIn the domain of generative multimedia and interactive experiences, generating realistic and accurate full-body poses from sparse tracking is crucial for many real-world applications, while achieving sequence modeling and efficient motion generation remains challenging. Recently, state space models (SSMs) with efficient hardware-aware designs (i.e., Mamba) have shown great potential for sequence modeling, particularly in temporal contexts. However, processing motion data is still challenging for SSMs. Specifically, the sparsity of input conditions makes motion generation an ill-posed problem. Moreover, the complex structure of the human body further complicates this task. To address these issues, we present Motion Mamba Diffusion (MMD), a novel conditional diffusion model, which effectively utilizes the sequence modeling capability of SSMs and the robust generation ability of diffusion models to track full-body poses accurately. In particular, we design a bidirectional Temporal Mamba Module (TMM) to model motion sequence. Additionally, a Spatial Mamba Module (SMM) is further proposed for feature enhancement within a single frame. Extensive experiments on the large motion capture dataset (AMASS) demonstrate that our proposed approach outperforms the latest methods in terms of accuracy and smoothness, thus providing a crucial advancement for creating realistic virtual avatars in various applications. Kun Dong 0001, Jian Xue 0002, Zehai Niu, Xing Lan, Ke Lu 0002, Qingyuan Liu 0001, Xiaoyu Qin 0001 |
ACM Multimedia | 3 |
| 2024 | Skeleton Cluster Tracking for robust multi-view multi-person 3D human pose estimation
Zehai Niu, Ke Lu 0002, Jian Xue 0002, Jinbao Wang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2024 | Dynamic spatial-temporal topology graph network for skeleton-based action recognition
Lian Chen, Ke Lu 0002, Zehai Niu, Runchen Wei, Jian Xue 0002 |
Multim. Syst. | 3 |
| 2024 | From Methods to Applications: A Review of Deep 3D Human Motion CaptureabstractMotion capture technology is crucial in various applications like animation, virtual reality and sports analysis. With the development of deep learning methods, significant progress has been experienced in this field, producing cost-effective and user-friendly solutions for various applications. This paper provides a comprehensive review of deep learning-based human motion capture techniques. Our review aims to bridge the gap between academic research and practical applications, providing valuable insights and guidance for researchers and practitioners in deep learning-based human motion capture. Our study puts forth a new application-oriented taxonomy that comprehensively summarises five fundamental routes of motion capture technology. In addition to that, we also delve into the research priorities linked with each route, following the structure of “hardware requirements - technical routes - datasets - evaluation metrics” and extending the necessary criteria for transferring traditional motion capture systems to deep learning-based ones. Meanwhile, for the motion capture technology, the current state of the art is reviewed, the challenges are identified, and the future directions of the research are outlined. Zehai Niu, Ke Lu 0002, Jian Xue 0002, Xiaoyu Qin 0001, Jinbao Wang 0001, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Multi-view 3D Smooth Human Pose Estimation based on Heatmap Filtering and Spatio-temporal InformationabstractThe estimation of 3D human poses from time-synchronized, calibrated multi-view video usually consists of two steps: (1) a 2D detector to locate the 2D coordinate point position of the joint via heatmaps for each frame and (2) a post-processing method such as the recursive pictorial structure model or robust triangulation to obtain 3D coordinate points. However, most existing methods are based on a single frame only. They do not take advantage of the temporal characteristics of the video sequence itself, and must rely on post-processing algorithms. They are also susceptible to human self-occlusion, and the generated sequences suffer from jitter. Therefore, we propose a network model incorporating spatial and temporal features. Using a coarse-to-fine approach, the proposed heatmap temporal network (HTN) generates temporal heatmap information, with an occlusion heatmap filter used to filter low-quality heatmaps before they are sent to the HTN. The heatmap fusion and the triangulation weights are dynamically adjusted, and intermediate supervision is employed to enable better integration of temporal and spatial information. Our network is also end-to-end differentiable. This overcomes the long-standing problem of skeleton jitter being generated and ensures that the sequence is smooth and stable. Zehai Niu, Ke Lu 0002, Jian Xue 0002, Runchen Wei |
ACM Multimedia | 1 |