VLDB 2026 Research / reviewers in the wild / expert
Xinxin Zuo
dblp:167/3181
· DBLP profile ↗
40ranked-venue papers
5as first author
29since 2021 · last 2026
0000-0002-7116-9634ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 3 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 20 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Test-time Adaptation for 3D Human Pose and Shape Estimation
Zhixiang Chi, Sen Wang 0003, Yang Wang 0003, Xinxin Zuo |
FG | 5 |
| 2026 | An Empirical Study of Self-supervised Pretraining in X-Ray Security Screening
Niloofar Akbari, Yang Wang 0003, Xinxin Zuo |
ICPR (8) | 3 |
| 2026 | Controllable Diffusion-Based Data Augmentation for X-Ray Object Detection
Jacob Kingi, Yang Wang 0003, Xinxin Zuo |
ICPR (8) | 3 |
| 2026 | ProtoAug: Prototype-Guided Uncertainty-Aware Augmentation for Long-Tail Motion Prediction
Ziheng Lu, Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yang Wang 0003, Xinxin Zuo |
IEEE Internet Things J. | 6 |
| 2025 | IMFine: 3D Inpainting via Geometry-guided Multi-view RefinementabstractCurrent 3D inpainting and object removal methods are largely limited to front-facing scenes, facing substantial challenges when applied to diverse, "unconstrained" scenes where the camera orientation and trajectory are unrestricted. To bridge this gap, we introduce a novel approach that produces inpainted 3D scenes with consistent visual quality and coherent underlying geometry across both front-facing and unconstrained scenes. Specifically, we propose a robust 3D inpainting pipeline that incorporates geometric priors and a multi-view refinement network trained via test-time adaptation, building on a pre-trained image inpainting model. Additionally, we develop a novel inpainting mask detection technique to derive targeted inpainting masks from object masks, boosting the performance in handling unconstrained scenes. To validate the efficacy of our approach, we create a challenging and diverse benchmark that spans a wide range of scenes. Comprehensive experiments demonstrate that our proposed method substantially outperforms existing state-of-the-art approaches. Zhihao Shi, Dong Huo, Yuhongze Zhou, Yan Min, Juwei Lu, Xinxin Zuo |
CVPR | 6 |
| 2025 | BOOTPLACE: Bootstrapped Object Placement with Detection TransformersabstractIn this paper, we tackle the copy-paste image-to-image composition problem with a focus on object placement learning. Prior methods have leveraged generative models to reduce the reliance for dense supervision. However, this often limits their capacity to model complex data distributions. Alternatively, transformer networks with a sparse contrastive loss have been explored, but their over-relaxed regularization often leads to imprecise object placement. We introduce BootPlace, a novel paradigm that formulates object placement as a placement-by-detection problem. Our approach begins by identifying suitable regions of interest for object placement. This is achieved by training a specialized detection transformer on object-subtracted backgrounds, enhanced with multi-object supervisions. It then semantically associates each target compositing object with detected regions based on their complementary characteristics. Through a boostrapped training approach applied to randomly object-subtracted images, our model enforces meaningful placements through extensive paired data augmentation. Experimental results on established benchmarks demonstrate BootPlace’s superior performance in object repositioning, markedly surpassing state-of-the-art baselines on Cityscapes and OPA datasets with notable improvements in IOU scores. Additional ablation studies further showcase the compositionality and generalizability of our approach, supported by user study evaluations. Code is available at https://github.com/RyanHangZhou/BootPlace Hang Zhou 0007, Xinxin Zuo, Li Cheng 0001 |
CVPR | 2 |
| 2025 | Tera: Rethinking Text-Guided Realistic 3D Avatar GenerationabstractIn this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two-stage training strategy for learning a native 3D avatar generative model. Initially, we distill a decoder to derive a structured latent space from a large human reconstruction model. Subsequently, a text-controlled latent diffusion model is trained to generate photorealistic 3D human avatars within this latent space. TeRA enhances the model performance by eliminating slow iterative optimization and enables text-based partial customization through a structured 3D human representation. Experiments have proven our approach's superiority over previous text-to-avatar generative models in subjective and objective evaluation. Yiyu Zhuang, Yifei Zeng, Xun Cao, Xinxin Zuo, Hao Zhu 0004 |
ICCV | 7 |
| 2025 | MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked TransformerabstractGenerative masked transformer have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the animation domain, large datasets are not always available. Applying generative masked modeling to generate diverse instances from a single MoCap reference may lead to overfitting, a challenge that remains unexplored. In this work, we present MotionDreamer, a localized masked modeling paradigm designed to learn motion internal patterns from a given motion with arbitrary topology and duration. By embedding the given motion into quantized tokens with a novel distribution regularization method, MotionDreamer constructs a robust and informative codebook for local motion patterns. Moreover, a sliding window local attention is introduced in our masked transformer, enabling the generation of natural yet diverse animations that closely resemble the reference motion patterns. As demonstrated through comprehensive experiments, MotionDreamer outperforms the state-of-the-art methods that are typically GAN or Diffusion-based in both faithfulness and diversity. Thanks to the consistency and robustness of quantization-based approach, MotionDreamer can also effectively perform downstream tasks such as temporal motion editing, crowd motion synthesis, and beat-aligned dance generation, all using a single reference motion. Our implementation, learned models and results are to be made publicly available upon paper acceptance. Chuan Guo 0002, Yuxuan Mu, Muhammad Gohar Javed, Xinxin Zuo, Juwei Lu, Hai Jiang 0001, Li Cheng 0001 |
ICLR | 5 |
| 2025 | Collaborative Cloud-edge Generalized Category DiscoveryabstractGeneralized category discovery (GCD) aims to group unlabeled samples from known and unknown classes when only part of the labeled data in the known classes is given. It allows the model to adapt to dynamic environments by discovering novel categories. However, when we applied the GCD approach to the decentralized open world, we still encountered the following challenges: (1) none of labeled data easily obtained in the open world, (2) heterogeneous label spaces across different environments, (3)representation degradation caused by fine-tuning models with limited data in specific environments. To address the above challenges, we introduce a new and practical task, namely Cloud-edge GCD (CE-GCD). Different from semi-supervised GCD, CE-GCD assumes that we only have a base model trained on common public categories, and aims to perform personalized unsupervised novel category discovery in multiple environments with heterogeneous label spaces. Data from different environments or clients cannot be shared, only model parameters can be transferred. To tackle this problem, we propose a novel GCD framework based on energy-guided known class discrimination and multi-level contrastive learning. In each client, we first use the classifier of the base model to distinguish between known and unknown classes, and then perform unsupervised learning on the unknown classes. Each client transfers category information through prototypes to assist learning. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach. Yingbing Liu, Fei Ma 0001, Xinxin Zuo, Fan Zhang 0007, Yang Wang 0003 |
ACM Multimedia | 4 |
| 2025 | PointMAC: Meta-Learned Adaptation for Robust Test-Time Point Cloud CompletionabstractPoint cloud completion is essential for robust 3D perception in safety-critical applications such as robotics and augmented reality. However, existing models perform static inference and rely heavily on inductive biases learned during training, limiting their ability to adapt to novel structural patterns and sensor-induced distortions at test time.
To address this limitation, we propose PointMAC, a meta-learned framework for robust test-time adaptation in point cloud completion. It enables sample-specific refinement without requiring additional supervision.
Our method optimizes the completion model under two self-supervised auxiliary objectives that simulate structural and sensor-level incompleteness.
A meta-auxiliary learning strategy based on Model-Agnostic Meta-Learning (MAML) ensures that adaptation driven by auxiliary objectives is consistently aligned with the primary completion task.
During inference, we adapt the shared encoder on-the-fly by optimizing auxiliary losses, with the decoder kept fixed. To further stabilize adaptation, we introduce Adaptive $\lambda$-Calibration, a meta-learned mechanism for balancing gradients between primary and auxiliary objectives.
Extensive experiments on synthetic, simulated, and real-world datasets demonstrate that PointMAC achieves state-of-the-art results by refining each sample individually to produce high-quality completions. To the best of our knowledge, this is the first work to apply meta-auxiliary test-time adaptation to point cloud completion. Linlian Jiang, Li Gu, Ziqiang Wang 0003, Xinxin Zuo, Yang Wang 0003 |
NeurIPS | 5 |
| 2025 | Sketch2PoseNet: Efficient and Generalized Sketch to 3D Human Pose Predictionabstract3D human pose estimation from sketches has broad applications in computer animation and film production. Unlike traditional human pose estimation, this task presents unique challenges due to the abstract and disproportionate nature of sketches. Previous sketch-to-pose methods, constrained by the lack of large-scale sketch-3D pose annotations, primarily relied on optimization with heuristic rules—an approach that is both time-consuming and limited in generalizability. To address these challenges, we propose a novel approach leveraging a "learn from synthesis" strategy. Firstly, a diffusion model is learned to synthesize sketch images from 2D poses projected from 3D human poses, mimicking disproportionate human structures in sketches. This process enables the creation of a synthetic dataset, SKEP-120K, consisting of 120k accurate sketch-3D pose annotation pairs across various sketch styles. Building on this synthetic dataset, we introduce an end-to-end data-driven framework for estimating human poses and shapes from diverse sketch styles. Our framework combines existing 2D pose detectors and generative diffusion priors for sketch feature extraction with a feed-forward neural network for efficient 2D pose estimation. Multiple heuristic loss functions have been incorporated to guarantee geometric coherence between the derived 3D poses and the detected 2D poses while preserving accurate self-contacts. Qualitative, quantitative, and subjective evaluations collectively affirm that our proposed model substantially surpasses previous ones in both estimation accuracy and speed for sketch-to-pose tasks. Yiyu Zhuang, Xun Cao, Chuan Guo 0002, Xinxin Zuo, Hao Zhu 0004 |
SIGGRAPH Asia | 6 |
| 2025 | Towards 4D human video stylizationabstractWe present a first step towards 4D (3D space and time) human video stylization, which addresses style transfer, novel view synthesis, and human animation within a unified framework. While numerous video stylization methods have been developed, they are typically restricted to rendering images in specific viewpoints of the input video, lacking the capability to generalize to novel views and novel poses in dynamic scenes. To overcome these limitations, we leverage Neural Radiance Fields (NeRFs) to represent and stylize videos within a single framework. Our method involves simultaneously representing the human subject and the surrounding scene using two NeRFs. This dual representation facilitates the animation of human subjects across various poses and novel viewpoints. A key innovation is our introduction of a geometry-guided tri-plane representation, which significantly boosts the efficiency and robustness of the feature representation compared to direct tri-plane optimization. Stylization is performed within the NeRF rendered feature space, which can reduce the computational burden compared to applying style transformation to the feature vector of sampled points. Extensive experiments demonstrate that the proposed method strikes a superior balance between stylized textures and temporal coherence, surpassing existing approaches. Furthermore, our framework uniquely extends its capabilities to accommodate novel poses and viewpoints, making it a versatile tool for creative human video stylization. The source code and results will be available at this github site . The stylized videos are available in this Youtube video . • We present a 3D video stylization framework for novel views and human poses. • A geometric prior is introduced to enhance tri-plane feature learning efficiency. • Our method balances stylized textures and temporal coherence for dynamic scenes. Xinxin Zuo, Fangzhou Mu, Jian Wang 0100, Ming-Hsuan Yang 0001 |
Comput. Vis. Image Underst. | 2 |
| 2025 | Highly Efficient 3D Human Pose Tracking From Events With Spiking Spatiotemporal TransformerabstractEvent camera, as an asynchronous vision sensor capturing scene dynamics, presents new opportunities for highly efficient 3D human pose tracking. Existing approaches typically adopt modern-day Artificial Neural Networks (ANNs), such as CNNs or Transformer, where sparse events are converted into dense images or paired with additional gray-scale images as input. Such practices, however, ignore the inherent sparsity of events, resulting in redundant computations, increased energy consumption, and potentially degraded performance. Motivated by these observations, we introduce the first sparse Spiking Neural Networks (SNNs) framework for 3D human pose tracking based solely on events. Our approach eliminates the need to convert sparse data to dense formats or incorporate additional images, thereby fully exploiting the innate sparsity of input events. Central to our framework is a novel Spiking Spatio-temporal Transformer, which enables bi-directional spatio-temporal fusion of spike pose features and provides a guaranteed similarity measurement between binary spike features in spiking attention. Moreover, we have constructed a largescale synthetic dataset, SynEventHPD, that features a broad and diverse set of 3D human motions, as well as much longer hours of event streams. Empirical experiments demonstrate the superiority of our approach over existing state-of-the-art (SOTA) ANN-based methods, requiring only 19.1% FLOPs and 3.6% energy cost. Furthermore, our approach outperforms existing SNN-based benchmarks in this task, highlighting the effectiveness of our proposed SNN framework. The dataset will be released upon acceptance, and code can be found at https://github.com/JimmyZou/HumanPoseTracking_SNN. Shihao Zou, Yuxuan Mu, Wei Ji 0011, Zi-An Wang, Xinxin Zuo, Sen Wang 0003, Weixin Si, Li Cheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling
Dong Huo, Zixin Guo, Xinxin Zuo, Zhihao Shi, Juwei Lu, Peng Dai 0002, Songcen Xu, Li Cheng 0001, Yee-Hong Yang |
ECCV (38) | 3 |
| 2024 | GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction
Yuxuan Mu, Xinxin Zuo, Chuan Guo 0002, Juwei Lu, Songcen Xu, Peng Dai 0002, Youliang Yan, Li Cheng 0001 |
ECCV (79) | 2 |
| 2024 | Generative Human Motion Stylization in Latent SpaceabstractHuman motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the \textit{latent space} of pretrained autoencoders as a more expressive and robust representation for motion extraction and infusion. Building upon this, we present a novel \textit{generative} model that produces diverse stylization results of a single motion (latent) code. During training, a motion code is decomposed into two coding components: a deterministic content code, and a probabilistic style code adhering to a prior distribution; then a generator massages the random combination of content and style codes to reconstruct the corresponding motion codes. Our approach is versatile, allowing the learning of probabilistic style space from either style labeled or unlabeled motions, providing notable flexibility in stylization as well. In inference, users can opt to stylize a motion using style cues from a reference motion or a label. Even in the absence of explicit style input, our model facilitates novel re-stylization by sampling from the unconditional style prior distribution. Experimental results show that our proposed stylization models, despite their lightweight design, outperform the state-of-the-arts in style reeanactment, content preservation, and generalization across various applications and settings. Chuan Guo 0002, Yuxuan Mu, Xinxin Zuo, Peng Dai 0002, Youliang Yan, Juwei Lu, Li Cheng 0001 |
ICLR | 3 |
| 2023 | TM2D: Bimodality Driven 3D Dance Generation via Music-Text IntegrationabstractWe propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce richer dance movements guided by the instructive information provided by the text. However, the lack of paired motion data with both music and text modalities limits the ability to generate dance movements that integrate both. To alleviate this challenge, we propose to utilize a 3D human motion VQ-VAE to project the motions of the two datasets into a latent space consisting of quantized vectors, which effectively mix the motion tokens from the two datasets with different distributions for training. Additionally, we propose a cross-modal transformer to integrate text instructions into motion generation architecture for generating 3D dance movements without degrading the performance of music-conditioned dance generation. To better evaluate the quality of the generated motion, we introduce two novel metrics, namely Motion Prediction Distance (MPD) and Freezing Score (FS), to measure the coherence and freezing percentage of the generated motion. Extensive experiments show that our approach can generate realistic and coherent dance movements conditioned on both text and music while maintaining comparable performance with the two single modalities. Code is available at https://garfield-kh.github.io/TM2D/. Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo 0002, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, Xinchao Wang |
ICCV | 6 |
| 2023 | SiamCLIM: Text-Based Pedestrian Search Via Multi-Modal Siamese Contrastive LearningabstractText-based pedestrian search (TBPS) aims at retrieving target persons from the image gallery through descriptive text queries. Despite remarkable progress in recent state-of-the-art approaches, previous works still struggle to efficiently extract discriminative features from multi-modal data. To address the problem of cross-modal fine-grained text-to-image, we proposed a novel Siamese Contrastive Language-Image Model (SiamCLIM). The model implements textual description and target-person interaction through deep bilateral projection, and siamese network structure to capture the relationship between text and image. Experiments show that our model significantly outperforms the state-of-the-art methods on cross-modal fine-grained matching tasks. We conduct the downstream task experiments on the benchmark dataset CUHK-PEDES and the experimental results demonstrate that our model is state-of-the-art and outperforms the current methods by 11.55%, 11.02%, and 7.76% in terms of top-1, top-5, and top-10 accuracy, respectively. Runlin Huang, Shuyang Wu, Leiping Jie, Xinxin Zuo |
ICIP | 4 |
| 2023 | Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the WildabstractThis paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependent texture, and a diffuse UV texture. Subsequently, users can either manually edit the UV or provide a prompt for the automatic generation of a new 3D-consistent texture. To achieve high-quality 3D texture generation, we propose a structure-aware score distillation sampling method to optimize a neural UV texture based on user-defined text and empower an image diffusion model with 3D-consistent generation capability. Furthermore, we introduce a few-view resampling training method and utilize a super-resolution model to obtain refined high-resolution UV textures (2048$\times$2048) for 3D texturing. Extensive experiments collectively validate the superior performance of Decorate3D in retexturing real-world 3D objects. Project page: https://decorate3d.github.io/Decorate3D/. Xinxin Zuo, Peng Dai 0002, Juwei Lu, Li Cheng 0001, Youliang Yan, Songcen Xu |
NeurIPS | 2 |
| 2023 | Human Pose and Shape Estimation From Single Polarization ImagesabstractThis paper focuses on a new problem of estimating human pose and shape from single polarization images. Polarization camera is known to be able to capture the polarization of reflected lights that preserves rich geometric cues of an object surface. Inspired by the recent applications in surface normal reconstruction from polarization images, in this paper, we attempt to estimate human pose and shape from single polarization images by leveraging the polarization-induced geometric cues. A dedicated two-stage pipeline is proposed: given a single polarization image, stage one (Polar2Normal) focuses on the fine detailed human body surface normal estimation; stage two (Polar2Shape) then reconstructs clothed human shape from the polarization image and the estimated surface normal. To empirically validate our approach, a dedicated dataset (PHSPD) is constructed, consisting of over 500 K frames with accurate pose and parametric shape annotations. Empirical evaluations on this real-world dataset as well as a synthetic dataset, SURREAL, demonstrate the effectiveness of our approach. It suggests polarization camera as a promising alternative to the more conventional RGB camera for human pose and shape estimation. Shihao Zou, Xinxin Zuo, Sen Wang 0003, Yiming Qian, Chuan Guo 0002, Li Cheng 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Generating Diverse and Natural 3D Human Motions from TextabstractAutomated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem with a two-stage approach: text2length sampling and text2motion generation. Text2length involves sampling from the learned distribution function of motion lengths conditioned on the input text. This is followed by our text2motion module using temporal variational autoen-coder to synthesize a diverse set of human motions of the sampled lengths. Instead of directly engaging with pose sequences, we propose motion snippet code as our internal motion representation, which captures local semantic motion contexts and is empirically shown to facilitate the generation of plausible motions faithful to the input text. Moreover, a large-scale dataset of scripted 3D Human motions, HumanML3D, is constructed, consisting of 14,616 motion clips and 44,970 text descriptions. Chuan Guo 0002, Shihao Zou, Xinxin Zuo, Sen Wang 0003, Wei Ji 0011, Li Cheng 0001 |
CVPR | 3 |
| 2022 | TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Li Cheng 0001 |
ECCV (35) | 2 |
| 2022 | Object Wake-Up: 3D Object Rigging from a Single Image
Xinxin Zuo, Sen Wang 0003, Zhenbo Yu, Bingbing Ni, Minglun Gong, Li Cheng 0001 |
ECCV (2) | 2 |
| 2022 | Action2video: Generating Videos of Human 3D Actions
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Xinshuang Liu, Shihao Zou, Minglun Gong, Li Cheng 0001 |
Int. J. Comput. Vis. | 2 |
| 2022 | Detailed Avatar Recovery From Single ImageabstractThis paper presents a novel framework to recover detailed avatar from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, texture, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric-based template that lacks the surface details. As such resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of the parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. Our method can restore detailed human body shapes with complete textures beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | 3D pose estimation and future motion prediction from 2D images
Youdong Ma, Xinxin Zuo, Sen Wang 0003, Minglun Gong, Li Cheng 0001 |
Pattern Recognit. | 3 |
| 2021 | LiDAR-Aug: A General Rendering-Based Augmentation Framework for 3D Object DetectionabstractAnnotating the LiDAR point cloud is crucial for deep learning-based 3D object detection tasks. Due to expensive labeling costs, data augmentation has been taken as a necessary module and plays an important role in training the neural network. "Copy" and "paste" (i.e., GT-Aug) is the most commonly used data augmentation strategy, however, the occlusion between objects has not been taken into consideration. To handle the above limitation, we propose a rendering-based LiDAR augmentation frame-work (i.e., LiDAR-Aug) to enrich the training data and boost the performance of LiDAR-based 3D object detectors. The proposed LiDAR-Aug is a plug-and-play module that can be easily integrated into different types of 3D object detection frameworks. Compared to the traditional object augmentation methods, LiDAR-Aug is more realistic and effective. Finally, we verify the proposed framework on the public KITTI dataset with different 3D object detectors. The experimental results show the superiority of our method compared to other data augmentation strategies. We plan to make our data and code public to help other researchers reproduce our results. Xinxin Zuo, Dingfu Zhou, Shengze Jin, Sen Wang 0003, Liangjun Zhang |
CVPR | 2 |
| 2021 | EventHPE: Event-based 3D Human Pose and Shape EstimationabstractEvent camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges: rather than capturing static body postures, the event signals are best at capturing local motions. This leads us to propose a two-stage deep learning approach, called EventHPE. The first-stage, FlowNet, is trained by unsupervised learning to infer optical flow from events. Both events and optical flow are closely related to human body dynamics, which are fed as input to the ShapeNet in the second stage, to estimate 3D human shapes. To mitigate the discrepancy between image-based flow (optical flow) and shape-based flow (vertices movement of human body shape), a novel flow coherence loss is introduced by exploiting the fact that both flows are originated from the identical human motion. An in-house event-based 3D human dataset is curated that comes with 3D pose and shape annotations, which is by far the largest one to our knowledge. Empirical evaluations on DHP19 dataset and our in-house dataset demonstrate the effectiveness of our approach. Shihao Zou, Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Pengyu Wang 0007, Xiaoqin Hu, Shoushun Chen, Minglun Gong, Li Cheng 0001 |
ICCV | 3 |
| 2021 | SparseFusion: Dynamic Human Avatar Modeling From Sparse RGBD ImagesabstractIn this paper, we propose a novel approach to reconstruct 3D human body shapes based on a sparse set of RGBD frames using a single RGBD camera. We specifically focus on the realistic settings where human subjects move freely during the capture. The main challenge is how to robustly fuse these sparse frames into a canonical 3D model, under pose changes and surface occlusions. This is addressed by our new framework consisting of the following steps. First, based on a generative human template, for every two frames having sufficient overlap, an initial pairwise alignment is performed; It is followed by a global non-rigid registration procedure, in which partial results from RGBD frames are collected into a unified 3D shape, under the guidance of correspondences from the pairwise alignment; Finally, the texture map of the reconstructed human model is optimized to deliver a clear and spatially consistent texture. Empirical evaluations on synthetic and real datasets demonstrate both quantitatively and qualitatively the superior performance of our framework in reconstructing complete 3D human models with high fidelity. It is worth noting that our framework is flexible, with potential applications going beyond shape reconstruction. As an example, we showcase its use in reshaping and reposing to a new avatar. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Minglun Gong, Ruigang Yang, Li Cheng 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Speech2Video Synthesis with 3D Skeleton Regularization and Expressive Body Poses
Miao Liao, Peng Wang 0001, Hao Zhu 0004, Xinxin Zuo, Ruigang Yang |
ACCV (5) | 5 |
| 2020 | 3D Human Shape Reconstruction from a Polarization Image
Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang 0003, Chi Xu 0002, Minglun Gong, Li Cheng 0001 |
ECCV (14) | 2 |
| 2020 | Angus Cattle Recognition Using Deep LearningabstractAngus cattle have significant economical values. Individualized management is expected to improve the efficiency and prevent financial loss in the farming industry. However Angus cattle, being all black, are a challenging case for visual recognition. We present a system for image segmentation and identification on Angus cattle using deep learning methods. Two databases of cattle were first collected and annotated, one is frontal face only, captured in a lab setting with controlled lighting and pose in the same day. The second was captured in a farm with natural light and background at three different days. The full body of cattle is captured from different angles. Using three popular neutral networks: PrimNet, VGG16 and ResNet50, we have evaluated a number of design choices for cattle identifications, including face only, face + body, and with/without background segmentation. The best result is obtained using face + body image without background, achieving 85.45% accuracy with the VGG16 net. If we use images captured under different days as training and testing datasets, the accuracy drops dramatically below 10%. It remains as a challenging open problem to be resolved. Shunnan Chen, Sen Wang 0003, Xinxin Zuo, Ruigang Yang |
ICPR | 3 |
| 2020 | Action2Motion: Conditioned Generation of 3D Human MotionsabstractAction recognition is a relatively established task, where given an input sequence of human motion, the goal is to predict its action category. This paper, on the other hand, considers a relatively new problem, which could be thought of as an inverse of action recognition: given a prescribed action type, we aim to generate plausible human motion sequences in 3D. Importantly, the set of generated motions are expected to maintain its diversity to be able to explore the entire action-conditioned motion space; meanwhile, each sampled sequence faithfully resembles a natural human body articulation dynamics. Motivated by these objectives, we follow the physics law of human kinematics by adopting the Lie Algebra theory to represent the natural human motions; we also propose a temporal Variational Auto-Encoder (VAE) that encourages a diverse sampling of the motion space. A new 3D human motion dataset, HumanAct12, is also constructed. Empirical experiments over three distinct human motion datasets (including ours) demonstrate the effectiveness of our approach. Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, Li Cheng 0001 |
ACM Multimedia | 2 |
| 2020 | Detailed Surface Geometry and Albedo Recovery from RGB-D Video under Natural IlluminationabstractThis article presents a novel approach for depth map enhancement from an RGB-D video sequence. The basic idea is to exploit the photometric information in the color sequence to resolve the inherent ambiguity of shape from shading problem. Instead of making any assumption about surface albedo or controlled object motion and lighting, we use the lighting variations introduced by casual object movement. We are effectively calculating photometric stereo from a moving object under natural illuminations. One of the key technical challenges is to establish correspondences over the entire image set. We, therefore, develop a lighting insensitive robust pixel matching technique that out-performs optical flow method in presence of lighting variations. An adaptive reference frame selection procedure is introduced to get more robust to imperfect lambertian reflections. In addition, we present an expectation-maximization framework to recover the surface normal and albedo simultaneously, without any regularization term. We have validated our method on both synthetic and real datasets to show its superior performance on both surface details recovery and intrinsic decomposition. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh DeformationabstractThis paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template that lacks the surface details. As such the resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. We are able to restore detailed human body shapes beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. The code is available in https://github.com/zhuhao-nju/hmd.git. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
CVPR | 2 |
| 2017 | Detailed Surface Geometry and Albedo Recovery from RGB-D Video under Natural IlluminationabstractIn this paper we present a novel approach for depth map enhancement from an RGB-D video sequence. The basic idea is to exploit the photometric information in the color sequence. Instead of making any assumption about surface albedo or controlled object motion and lighting, we use the lighting variations introduced by casual object movement. We are effectively calculating photometric stereo from a moving object under natural illuminations. The key technical challenge is to establish correspondences over the entire image set. We therefore develop a lighting insensitive robust pixel matching technique that out-performs optical flow method in presence of lighting variations. In addition we present an expectation-maximization framework to recover the surface normal and albedo simultaneously, without any regularization term. We have validated our method on both synthetic and real datasets to show its superior performance on both surface details recovery and intrinsic decomposition. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ICCV | 1 |
| 2017 | A generative human-robot motion retargeting approach using a single depth sensorabstractThe goal of human-robot motion retargeting is to let a robot follow the movements performed by a human subject. This is traditionally achieved by applying the estimated poses from a human pose tracking system to a robot via explicit joint mapping strategies. In this paper, we present a novel approach that combine the human pose estimation and the motion retarget procedure in a unified generative framework. A 3D parametric human-robot model is proposed that has the specific joint and stability configurations as a robot while its shape resembles a human subject. Using a single depth camera to monitor human pose, we use its raw depth map as input and drive the human-robot model to fit the input 3D point cloud. The calculated joint angles of the fitted model can be applied onto the robots for retargeting. The robot's joint angles, instead of fitted individually, are fitted globally so that the transformed surface shape is as consistent as possible to the input point cloud. The robot configurations including its skeleton proportion, joint limitation, and DoF are enforced implicitly in the formulation. No explicit and pre-defined joints mapping strategies are needed. This framework is tested with both simulations and real robots that have different skeleton proportion and DoFs compared with human to show its effectiveness for motion retargeting. Sen Wang 0003, Xinxin Zuo, Runxiao Wang, Fuhua (Frank) Cheng, Ruigang Yang |
ICRA | 2 |
| 2016 | High-speed Depth Stream Generation from a Hybrid CameraabstractHigh-speed video has been commonly adopted in consumer-grade cameras, augmenting these videos with a corresponding depth stream will enable new multimedia applications, such as 3D slow-motion video. In this paper, we present a hybrid camera system that combines a high-speed color camera with a depth sensor, e.g. Kinect depth sensor, to generate a depth stream that can produce both high-speed and high-resolution RGB+depth stream. Simply interpolating the low-speed depth frames is not satisfactory, where interpolation artifacts and lose in surface details are often visible. We have developed a novel framework that utilizes both shading constraints within each frame and optical flow constraints between neighboring frames. More specifically we present (a) an effective method to find the intrinsics images to allow more accurate normal estimation; and (b) an optimization-based framework to estimate the high-resolution/high-speed depth stream, taking into consideration temporal smoothness and shading/depth consistency. We evaluated our holistic framework with both synthetic and real sequences, it showed superior performance than previous state-of-the-art. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ACM Multimedia | 1 |
| 2015 | Interactive Visual Hull Refinement for Specular and Transparent Object Surface ReconstructionabstractIn this paper we present a method of using standard multi-view images for 3D surface reconstruction of non-Lambertian objects. We extend the original visual hull concept to incorporate 3D cues presented by internal occluding contours, i.e., occluding contours that are inside the object's silhouettes. We discovered that these internal contours, which are results of convex parts on an object's surface, can lead to a tighter fit than the original visual hull. We formulated a new visual hull refinement scheme -- Locally Convex Carving that can completely reconstruct concavity caused by two or more intersecting convex surfaces. In addition we develop a novel approach for contour tracking given labeled contours in sparse key frames. It is designed specifically for highly specular or transparent objects, for which assumptions made in traditional contour detection/tracking methods, such as highest gradient and stationary texture edges, are no longer valid. It is formulated as an energy minimization function where several novel terms are developed to increase robustness. Based on the two core algorithms, we have developed an interactive system for 3D modeling. We have validated our system, both quantitatively and qualitatively, with four datasets of different object materials. Results show that we are able to generate visually pleasing models for very challenging cases. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ICCV | 1 |
| 2015 | Multiple Depth Maps Integration for 3D Reconstruction Using Geodesic Graph CutsabstractDepth images, in particular depth maps estimated from stereo vision, may have a substantial amount of outliers and result in inaccurate 3D modelling and reconstruction. To address this challenging issue, in this paper, a graph-cut based multiple depth maps integration approach is proposed to obtain smooth and watertight surfaces. First, confidence maps for the depth images are estimated to suppress noise, based on which reliable patches covering the object surface are determined. These patches are then exploited to estimate the path weight for 3D geodesic distance computation, where an adaptive regional term is introduced to deal with the "shorter-cuts" problem caused by the effect of the minimal surface bias. Finally, the adaptive regional term and the boundary term constructed using patches are combined in the graph-cut framework for more accurate and smoother 3D modelling. We demonstrate the superior performance of our algorithm on the well-known Middlebury multi-view database and additionally on real-world multiple depth images captured by Kinect. The experimental results have shown that our method is able to preserve the object protrusions and details while maintaining surface smoothness. Jiangbin Zheng 0001, Xinxin Zuo, Jinchang Ren, Sen Wang 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 2 |