EDBT 2026 Demo / reviewers in the wild / expert
Sen Wang 0003
dblp:69/6403-3
· DBLP profile ↗
24ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0002-1808-5239ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Test-time Adaptation for 3D Human Pose and Shape Estimation
Zhixiang Chi, Sen Wang 0003, Yang Wang 0003, Xinxin Zuo |
FG | 3 |
| 2025 | Highly Efficient 3D Human Pose Tracking From Events With Spiking Spatiotemporal TransformerabstractEvent camera, as an asynchronous vision sensor capturing scene dynamics, presents new opportunities for highly efficient 3D human pose tracking. Existing approaches typically adopt modern-day Artificial Neural Networks (ANNs), such as CNNs or Transformer, where sparse events are converted into dense images or paired with additional gray-scale images as input. Such practices, however, ignore the inherent sparsity of events, resulting in redundant computations, increased energy consumption, and potentially degraded performance. Motivated by these observations, we introduce the first sparse Spiking Neural Networks (SNNs) framework for 3D human pose tracking based solely on events. Our approach eliminates the need to convert sparse data to dense formats or incorporate additional images, thereby fully exploiting the innate sparsity of input events. Central to our framework is a novel Spiking Spatio-temporal Transformer, which enables bi-directional spatio-temporal fusion of spike pose features and provides a guaranteed similarity measurement between binary spike features in spiking attention. Moreover, we have constructed a largescale synthetic dataset, SynEventHPD, that features a broad and diverse set of 3D human motions, as well as much longer hours of event streams. Empirical experiments demonstrate the superiority of our approach over existing state-of-the-art (SOTA) ANN-based methods, requiring only 19.1% FLOPs and 3.6% energy cost. Furthermore, our approach outperforms existing SNN-based benchmarks in this task, highlighting the effectiveness of our proposed SNN framework. The dataset will be released upon acceptance, and code can be found at https://github.com/JimmyZou/HumanPoseTracking_SNN. Shihao Zou, Yuxuan Mu, Wei Ji 0011, Zi-An Wang, Xinxin Zuo, Sen Wang 0003, Weixin Si, Li Cheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | MoMask: Generative Masked Modeling of 3D Human MotionsabstractWe introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In Mo-Mask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity details. Starting at the base layer, with a sequence of motion tokens obtained by vector quan-tization, the residual tokens of increasing orders are de-rived and stored at the subsequent layers of the hierar-chy. This is consequently followed by two distinct bidirectional transformers. For the base-layer motion tokens, a Masked Transformer is designated to predict randomly masked motion tokens conditioned on text input at training stage. During generation (i. e. inference) stage, starting from an empty sequence, our Masked Transformer iteratively fills up the missing tokens; Subsequently, a Residual Transformer learns to progressively predict the next-layer tokens based on the results from current layer. Extensive experiments demonstrate that MoMask outperforms the state-of-art methods on the text-to-motion generation task, with an FID of 0.045 (vs e.g. 0.141 of T2M-GPT) on the HumanML3D dataset, and 0.228 (vs 0.514) on KIT-ML, respectively. MoMask can also be seamlessly applied in related tasks without further model fine-tuning, such as text-guided temporal inpainting. Chuan Guo 0002, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang 0003, Li Cheng 0001 |
CVPR | 4 |
| 2023 | Human Pose and Shape Estimation From Single Polarization ImagesabstractThis paper focuses on a new problem of estimating human pose and shape from single polarization images. Polarization camera is known to be able to capture the polarization of reflected lights that preserves rich geometric cues of an object surface. Inspired by the recent applications in surface normal reconstruction from polarization images, in this paper, we attempt to estimate human pose and shape from single polarization images by leveraging the polarization-induced geometric cues. A dedicated two-stage pipeline is proposed: given a single polarization image, stage one (Polar2Normal) focuses on the fine detailed human body surface normal estimation; stage two (Polar2Shape) then reconstructs clothed human shape from the polarization image and the estimated surface normal. To empirically validate our approach, a dedicated dataset (PHSPD) is constructed, consisting of over 500 K frames with accurate pose and parametric shape annotations. Empirical evaluations on this real-world dataset as well as a synthetic dataset, SURREAL, demonstrate the effectiveness of our approach. It suggests polarization camera as a promising alternative to the more conventional RGB camera for human pose and shape estimation. Shihao Zou, Xinxin Zuo, Sen Wang 0003, Yiming Qian, Chuan Guo 0002, Li Cheng 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Generating Diverse and Natural 3D Human Motions from TextabstractAutomated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem with a two-stage approach: text2length sampling and text2motion generation. Text2length involves sampling from the learned distribution function of motion lengths conditioned on the input text. This is followed by our text2motion module using temporal variational autoen-coder to synthesize a diverse set of human motions of the sampled lengths. Instead of directly engaging with pose sequences, we propose motion snippet code as our internal motion representation, which captures local semantic motion contexts and is empirically shown to facilitate the generation of plausible motions faithful to the input text. Moreover, a large-scale dataset of scripted 3D Human motions, HumanML3D, is constructed, consisting of 14,616 motion clips and 44,970 text descriptions. Chuan Guo 0002, Shihao Zou, Xinxin Zuo, Sen Wang 0003, Wei Ji 0011, Li Cheng 0001 |
CVPR | 4 |
| 2022 | TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Li Cheng 0001 |
ECCV (35) | 3 |
| 2022 | Object Wake-Up: 3D Object Rigging from a Single Image
Xinxin Zuo, Sen Wang 0003, Zhenbo Yu, Bingbing Ni, Minglun Gong, Li Cheng 0001 |
ECCV (2) | 3 |
| 2022 | Action2video: Generating Videos of Human 3D Actions
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Xinshuang Liu, Shihao Zou, Minglun Gong, Li Cheng 0001 |
Int. J. Comput. Vis. | 3 |
| 2022 | Part-Level Car Parsing and Reconstruction in Single Street View ImagesabstractPart information has been proven to be resistant to occlusions and viewpoint changes, which are main difficulties in car parsing and reconstruction. However, in the absence of datasets and approaches incorporating car parts, there are limited works that benefit from it. In this paper, we propose the first part-aware approach for joint part-level car parsing and reconstruction in single street view images. Without labor-intensive part annotations on real images, our approach simultaneously estimates pose, shape, and semantic parts of cars. There are two contributions in this paper. First, our network introduces dense part information to facilitate pose and shape estimation, which is further optimized with a novel 3D loss. To obtain part information in real images, a class-consistent method is introduced to implicitly transfer part knowledge from synthesized images. Second, we construct the first high-quality dataset containing 348 car models with physical dimensions and part annotations. Given these models, 60K synthesized images with randomized configurations are generated. Experimental results demonstrate that part knowledge can be effectively transferred with our class-consistent method, which significantly improves part segmentation performance on real street views. By fusing dense part information, our pose and shape estimation results achieve the state-of-the-art performance on the ApolloCar3D and outperform previous approaches by large margins in terms of both A3DP-Abs and A3DP-Rel. Qichuan Geng, Hong Zhang 0009, Feixiang Lu, Xinyu Huang 0001, Sen Wang 0003, Zhong Zhou, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Detailed Avatar Recovery From Single ImageabstractThis paper presents a novel framework to recover detailed avatar from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, texture, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric-based template that lacks the surface details. As such resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of the parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. Our method can restore detailed human body shapes with complete textures beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | 3D pose estimation and future motion prediction from 2D images
Youdong Ma, Xinxin Zuo, Sen Wang 0003, Minglun Gong, Li Cheng 0001 |
Pattern Recognit. | 4 |
| 2021 | LiDAR-Aug: A General Rendering-Based Augmentation Framework for 3D Object DetectionabstractAnnotating the LiDAR point cloud is crucial for deep learning-based 3D object detection tasks. Due to expensive labeling costs, data augmentation has been taken as a necessary module and plays an important role in training the neural network. "Copy" and "paste" (i.e., GT-Aug) is the most commonly used data augmentation strategy, however, the occlusion between objects has not been taken into consideration. To handle the above limitation, we propose a rendering-based LiDAR augmentation frame-work (i.e., LiDAR-Aug) to enrich the training data and boost the performance of LiDAR-based 3D object detectors. The proposed LiDAR-Aug is a plug-and-play module that can be easily integrated into different types of 3D object detection frameworks. Compared to the traditional object augmentation methods, LiDAR-Aug is more realistic and effective. Finally, we verify the proposed framework on the public KITTI dataset with different 3D object detectors. The experimental results show the superiority of our method compared to other data augmentation strategies. We plan to make our data and code public to help other researchers reproduce our results. Xinxin Zuo, Dingfu Zhou, Shengze Jin, Sen Wang 0003, Liangjun Zhang |
CVPR | 5 |
| 2021 | EventHPE: Event-based 3D Human Pose and Shape EstimationabstractEvent camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges: rather than capturing static body postures, the event signals are best at capturing local motions. This leads us to propose a two-stage deep learning approach, called EventHPE. The first-stage, FlowNet, is trained by unsupervised learning to infer optical flow from events. Both events and optical flow are closely related to human body dynamics, which are fed as input to the ShapeNet in the second stage, to estimate 3D human shapes. To mitigate the discrepancy between image-based flow (optical flow) and shape-based flow (vertices movement of human body shape), a novel flow coherence loss is introduced by exploiting the fact that both flows are originated from the identical human motion. An in-house event-based 3D human dataset is curated that comes with 3D pose and shape annotations, which is by far the largest one to our knowledge. Empirical evaluations on DHP19 dataset and our in-house dataset demonstrate the effectiveness of our approach. Shihao Zou, Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Pengyu Wang 0007, Xiaoqin Hu, Shoushun Chen, Minglun Gong, Li Cheng 0001 |
ICCV | 4 |
| 2021 | SparseFusion: Dynamic Human Avatar Modeling From Sparse RGBD ImagesabstractIn this paper, we propose a novel approach to reconstruct 3D human body shapes based on a sparse set of RGBD frames using a single RGBD camera. We specifically focus on the realistic settings where human subjects move freely during the capture. The main challenge is how to robustly fuse these sparse frames into a canonical 3D model, under pose changes and surface occlusions. This is addressed by our new framework consisting of the following steps. First, based on a generative human template, for every two frames having sufficient overlap, an initial pairwise alignment is performed; It is followed by a global non-rigid registration procedure, in which partial results from RGBD frames are collected into a unified 3D shape, under the guidance of correspondences from the pairwise alignment; Finally, the texture map of the reconstructed human model is optimized to deliver a clear and spatially consistent texture. Empirical evaluations on synthetic and real datasets demonstrate both quantitatively and qualitatively the superior performance of our framework in reconstructing complete 3D human models with high fidelity. It is worth noting that our framework is flexible, with potential applications going beyond shape reconstruction. As an example, we showcase its use in reshaping and reposing to a new avatar. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Minglun Gong, Ruigang Yang, Li Cheng 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | 3D Human Shape Reconstruction from a Polarization Image
Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang 0003, Chi Xu 0002, Minglun Gong, Li Cheng 0001 |
ECCV (14) | 4 |
| 2020 | Angus Cattle Recognition Using Deep LearningabstractAngus cattle have significant economical values. Individualized management is expected to improve the efficiency and prevent financial loss in the farming industry. However Angus cattle, being all black, are a challenging case for visual recognition. We present a system for image segmentation and identification on Angus cattle using deep learning methods. Two databases of cattle were first collected and annotated, one is frontal face only, captured in a lab setting with controlled lighting and pose in the same day. The second was captured in a farm with natural light and background at three different days. The full body of cattle is captured from different angles. Using three popular neutral networks: PrimNet, VGG16 and ResNet50, we have evaluated a number of design choices for cattle identifications, including face only, face + body, and with/without background segmentation. The best result is obtained using face + body image without background, achieving 85.45% accuracy with the VGG16 net. If we use images captured under different days as training and testing datasets, the accuracy drops dramatically below 10%. It remains as a challenging open problem to be resolved. Shunnan Chen, Sen Wang 0003, Xinxin Zuo, Ruigang Yang |
ICPR | 2 |
| 2020 | Action2Motion: Conditioned Generation of 3D Human MotionsabstractAction recognition is a relatively established task, where given an input sequence of human motion, the goal is to predict its action category. This paper, on the other hand, considers a relatively new problem, which could be thought of as an inverse of action recognition: given a prescribed action type, we aim to generate plausible human motion sequences in 3D. Importantly, the set of generated motions are expected to maintain its diversity to be able to explore the entire action-conditioned motion space; meanwhile, each sampled sequence faithfully resembles a natural human body articulation dynamics. Motivated by these objectives, we follow the physics law of human kinematics by adopting the Lie Algebra theory to represent the natural human motions; we also propose a temporal Variational Auto-Encoder (VAE) that encourages a diverse sampling of the motion space. A new 3D human motion dataset, HumanAct12, is also constructed. Empirical experiments over three distinct human motion datasets (including ours) demonstrate the effectiveness of our approach. Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, Li Cheng 0001 |
ACM Multimedia | 3 |
| 2020 | Detailed Surface Geometry and Albedo Recovery from RGB-D Video under Natural IlluminationabstractThis article presents a novel approach for depth map enhancement from an RGB-D video sequence. The basic idea is to exploit the photometric information in the color sequence to resolve the inherent ambiguity of shape from shading problem. Instead of making any assumption about surface albedo or controlled object motion and lighting, we use the lighting variations introduced by casual object movement. We are effectively calculating photometric stereo from a moving object under natural illuminations. One of the key technical challenges is to establish correspondences over the entire image set. We, therefore, develop a lighting insensitive robust pixel matching technique that out-performs optical flow method in presence of lighting variations. An adaptive reference frame selection procedure is introduced to get more robust to imperfect lambertian reflections. In addition, we present an expectation-maximization framework to recover the surface normal and albedo simultaneously, without any regularization term. We have validated our method on both synthetic and real datasets to show its superior performance on both surface details recovery and intrinsic decomposition. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh DeformationabstractThis paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template that lacks the surface details. As such the resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. We are able to restore detailed human body shapes beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. The code is available in https://github.com/zhuhao-nju/hmd.git. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
CVPR | 3 |
| 2017 | Detailed Surface Geometry and Albedo Recovery from RGB-D Video under Natural IlluminationabstractIn this paper we present a novel approach for depth map enhancement from an RGB-D video sequence. The basic idea is to exploit the photometric information in the color sequence. Instead of making any assumption about surface albedo or controlled object motion and lighting, we use the lighting variations introduced by casual object movement. We are effectively calculating photometric stereo from a moving object under natural illuminations. The key technical challenge is to establish correspondences over the entire image set. We therefore develop a lighting insensitive robust pixel matching technique that out-performs optical flow method in presence of lighting variations. In addition we present an expectation-maximization framework to recover the surface normal and albedo simultaneously, without any regularization term. We have validated our method on both synthetic and real datasets to show its superior performance on both surface details recovery and intrinsic decomposition. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ICCV | 2 |
| 2017 | A generative human-robot motion retargeting approach using a single depth sensorabstractThe goal of human-robot motion retargeting is to let a robot follow the movements performed by a human subject. This is traditionally achieved by applying the estimated poses from a human pose tracking system to a robot via explicit joint mapping strategies. In this paper, we present a novel approach that combine the human pose estimation and the motion retarget procedure in a unified generative framework. A 3D parametric human-robot model is proposed that has the specific joint and stability configurations as a robot while its shape resembles a human subject. Using a single depth camera to monitor human pose, we use its raw depth map as input and drive the human-robot model to fit the input 3D point cloud. The calculated joint angles of the fitted model can be applied onto the robots for retargeting. The robot's joint angles, instead of fitted individually, are fitted globally so that the transformed surface shape is as consistent as possible to the input point cloud. The robot configurations including its skeleton proportion, joint limitation, and DoF are enforced implicitly in the formulation. No explicit and pre-defined joints mapping strategies are needed. This framework is tested with both simulations and real robots that have different skeleton proportion and DoFs compared with human to show its effectiveness for motion retargeting. Sen Wang 0003, Xinxin Zuo, Runxiao Wang, Fuhua (Frank) Cheng, Ruigang Yang |
ICRA | 1 |
| 2016 | High-speed Depth Stream Generation from a Hybrid CameraabstractHigh-speed video has been commonly adopted in consumer-grade cameras, augmenting these videos with a corresponding depth stream will enable new multimedia applications, such as 3D slow-motion video. In this paper, we present a hybrid camera system that combines a high-speed color camera with a depth sensor, e.g. Kinect depth sensor, to generate a depth stream that can produce both high-speed and high-resolution RGB+depth stream. Simply interpolating the low-speed depth frames is not satisfactory, where interpolation artifacts and lose in surface details are often visible. We have developed a novel framework that utilizes both shading constraints within each frame and optical flow constraints between neighboring frames. More specifically we present (a) an effective method to find the intrinsics images to allow more accurate normal estimation; and (b) an optimization-based framework to estimate the high-resolution/high-speed depth stream, taking into consideration temporal smoothness and shading/depth consistency. We evaluated our holistic framework with both synthetic and real sequences, it showed superior performance than previous state-of-the-art. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ACM Multimedia | 2 |
| 2015 | Interactive Visual Hull Refinement for Specular and Transparent Object Surface ReconstructionabstractIn this paper we present a method of using standard multi-view images for 3D surface reconstruction of non-Lambertian objects. We extend the original visual hull concept to incorporate 3D cues presented by internal occluding contours, i.e., occluding contours that are inside the object's silhouettes. We discovered that these internal contours, which are results of convex parts on an object's surface, can lead to a tighter fit than the original visual hull. We formulated a new visual hull refinement scheme -- Locally Convex Carving that can completely reconstruct concavity caused by two or more intersecting convex surfaces. In addition we develop a novel approach for contour tracking given labeled contours in sparse key frames. It is designed specifically for highly specular or transparent objects, for which assumptions made in traditional contour detection/tracking methods, such as highest gradient and stationary texture edges, are no longer valid. It is formulated as an energy minimization function where several novel terms are developed to increase robustness. Based on the two core algorithms, we have developed an interactive system for 3D modeling. We have validated our system, both quantitatively and qualitatively, with four datasets of different object materials. Results show that we are able to generate visually pleasing models for very challenging cases. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ICCV | 3 |
| 2015 | Multiple Depth Maps Integration for 3D Reconstruction Using Geodesic Graph CutsabstractDepth images, in particular depth maps estimated from stereo vision, may have a substantial amount of outliers and result in inaccurate 3D modelling and reconstruction. To address this challenging issue, in this paper, a graph-cut based multiple depth maps integration approach is proposed to obtain smooth and watertight surfaces. First, confidence maps for the depth images are estimated to suppress noise, based on which reliable patches covering the object surface are determined. These patches are then exploited to estimate the path weight for 3D geodesic distance computation, where an adaptive regional term is introduced to deal with the "shorter-cuts" problem caused by the effect of the minimal surface bias. Finally, the adaptive regional term and the boundary term constructed using patches are combined in the graph-cut framework for more accurate and smoother 3D modelling. We demonstrate the superior performance of our algorithm on the well-known Middlebury multi-view database and additionally on real-world multiple depth images captured by Kinect. The experimental results have shown that our method is able to preserve the object protrusions and details while maintaining surface smoothness. Jiangbin Zheng 0001, Xinxin Zuo, Jinchang Ren, Sen Wang 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 4 |