Zhenzhen Weng

dblp:248/3274 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
6since 2021 · last 2024
0009-0004-1108-4155ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
abstract
Recent text-to-image generative models have exhibited remarkable abilities in generating high-fidelity and photorealistic images. However, despite the visually impressive results, these models often struggle to preserve plausible human structure in the generations. Due to this reason, while generative models have shown promising results in aiding downstream image recognition tasks by generating large volumes of synthetic data, they are not suitable for improving downstream human pose perception and understanding. In this work, we propose a Diffusion model with Human Pose Correction (Diffusion-HPC), a text-conditioned method that generates photo-realistic images with plausible posed humans by injecting prior knowledge about human body structure. Our generated images are accompanied by 3D meshes that serve as ground truths for improving Human Mesh Recovery tasks, where a short-age of 3D training data has long been an issue. Furthermore, we show that Diffusion-HPC effectively improves the realism of human generations under varying conditioning strategies.11Code: https://github.com/ZZWENG/Diffusion_HPC
Zhenzhen Weng, Laura Bravo-Sánchez, Serena Yeung-Levy
3DV1
2023 NeMo: 3D Neural Motion Fields from Multiple Video Instances of the Same Action
abstract
The task of reconstructing 3D human motion has wide-ranging applications. The gold standard Motion capture (MoCap) systems are accurate but inaccessible to the general public due to their cost, hardware, and space constraints. In contrast, monocular human mesh recovery (HMR) methods are much more accessible than MoCap as they take single-view videos as inputs. Replacing the multi-view MoCap systems with a monocular HMR method would break the current barriers to collecting accurate 3D motion thus making exciting applications like motion analysis and motion-driven animation accessible to the general public. However, the performance of existing HMR methods degrades when the video contains challenging and dynamic motion that is not in existing MoCap datasets used for training. This reduces its appeal as dynamic motion is frequently the target in 3D motion recovery in the aforementioned applications. Our study aims to bridge the gap between monocular HMR and multi-view MoCap systems by leveraging information shared across multiple video instances of the same action. We introduce the Neural Motion (NeMo) field. It is optimized to represent the underlying 3D motions across a set of videos of the same action. Empirically, we show that NeMo can recover 3D motion in sports using videos from the Penn Action dataset, where NeMo outperforms existing HMR methods in terms of 2D keypoint detection. To further validate NeMo using 3D metrics, we collected a small MoCap dataset mimicking actions in Penn Action, and show that NeMo achieves better 3D reconstruction compared to various baselines.
Kuan-Chieh Wang, Zhenzhen Weng, Maria Xenochristou, João Pedro Araújo 0001, Jeffrey Gu, C. Karen Liu, Serena Yeung-Levy
CVPR2
2023 3D Human Keypoints Estimation from Point Clouds in the Wild without Human Labels
abstract
Training a 3D human keypoint detector from point clouds in a supervised manner requires large volumes of high quality labels. While it is relatively easy to capture large amounts of human point clouds, annotating 3D key-points is expensive, subjective, error prone and especially difficult for long-tail cases (pedestrians with rare poses, scooterists, etc.). In this work, we propose GC-KPL - Geometry Consistency inspired Key Point Leaning, an approach for learning 3D human joint locations from point clouds without human labels. We achieve this by our novel unsupervised loss formulations that account for the structure and movement of the human body. We show that by training on a large training set from Waymo Open Dataset [21] without any human annotated keypoints, we are able to achieve reasonable performance as compared to the fully supervised approach. Further, the backbone benefits from the unsupervised training and is useful in downstream few-shot learning of keypoints, where fine-tuning on only 10 percent of the labeled training data gives comparable performance to fine-tuning on the entire set. We demonstrated that GC-KPL outperforms by a large margin over SoTA when trained on entire dataset and efficiently leverages large volumes of unlabeled data.
Zhenzhen Weng, Alexander S. Gorban, Jingwei Ji, Mahyar Najibi, Dragomir Anguelov
CVPR1
2022 Domain Adaptive 3D Pose Augmentation for In-the-Wild Human Mesh Recovery
abstract
The ability to perceive 3D human bodies from a single image has a multitude of applications ranging from entertainment and robotics to neuroscience and healthcare. A fundamental challenge in human mesh recovery is in collecting the ground truth 3D mesh targets required for training, which requires burdensome motion capturing systems and is often limited to indoor laboratories. As a result, while progress is made on benchmark datasets collected in these restrictive settings, models fail to generalize to real-world "in-the-wild" scenarios due to distribution shifts. We propose Domain Adaptive 3D Pose Augmentation (DAPA), a data augmentation method that enhances the model's generalization ability in in-the-wild scenarios. DAPA combines the strength of methods based on synthetic datasets by getting direct supervision from the synthesized meshes, and domain adaptation methods by using ground truth 2D keypoints from the target dataset. We show quantitatively that finetuning with DAPA effectively improves results on benchmarks 3DPW [38]and AGORA[32]. We further demonstrate the utility of DAPA on a challenging dataset curated from videos of real-world parent-child interaction.
Zhenzhen Weng, Kuan-Chieh Wang, Angjoo Kanazawa, Serena Yeung-Levy
3DV1
2021 Unsupervised Discovery of the Long-Tail in Instance Segmentation Using Hierarchical Self-Supervision
abstract
Instance segmentation is an active topic in computer vision that is usually solved by using supervised learning approaches over very large datasets composed of object level masks. Obtaining such a dataset for any new domain can be very expensive and time-consuming. In addition, models trained on certain annotated categories do not generalize well to unseen objects. The goal of this paper is to propose a method that can perform unsupervised discovery of long-tail categories in instance segmentation, through learning instance embeddings of masked regions. Leveraging rich relationship and hierarchical structure between objects in the images, we propose self-supervised losses for learning mask embeddings. Trained on COCO [34] dataset without additional annotations of the long-tail objects, our model is able to discover novel and more fine-grained objects than the common categories in COCO. We show that the model achieves competitive quantitative results on LVIS [17] as compared to the supervised and partially supervised methods.
Zhenzhen Weng, Mehmet Giray Ogut, Shai Limonchik, Serena Yeung-Levy
CVPR1
2021 Holistic 3D Human and Scene Mesh Estimation From Single View Images
abstract
The 3D world limits the human body pose and the human body pose conveys information about the surrounding objects. Indeed, from a single image of a person placed in an indoor scene, we as humans are adept at resolving ambiguities of the human pose and room layout through our knowledge of the physical laws and prior perception of the plausible object and human poses. However, few computer vision models fully leverage this fact. In this work, we pro-pose a holistically trainable model that perceives the 3D scene from a single RGB image, estimates the camera pose and the room layout, and reconstructs both human body and object meshes. By imposing a set of comprehensive and sophisticated losses on all aspects of the estimations, we show that our model outperforms existing human body mesh methods and indoor scene reconstruction methods. To the best of our knowledge, this is the first model that outputs both object and human predictions at the mesh level, and performs joint optimization on the scene and human poses.
Zhenzhen Weng, Serena Yeung-Levy
CVPR1
2019 Utilizing Weak Supervision to Infer Complex Objects and Situations in Autonomous Driving Data
abstract
While the detection and classification of simple objects encountered during autonomous driving sessions has been widely researched, the detection of complex objects and situations based on the combinations of objects in a scene remains relatively overlooked. This is especially difficult due to the cost of gathering labels for each complex scenario of interest before training a specialized model. To address this bottleneck of training data, we explore the applicability of weak supervision, or relying on higher level, noisier forms of supervision to label training data. Specifically, we use data programming, a paradigm that can learn the accuracy and dependency structure of these sources without using any ground truth labels and assign training labels accordingly. We focus on an example task of cyclist detection by comparing weak supervision, which relies on a set of user-defined rules over the outputs of detectors that identify people and bikes separately, to CyDet [1], which detects the cyclist as a complete object. We find that the weak supervision method can achieve a performance of 96.8 F1 points, 4.6 F1 higher than CyDet, without relying on any ground truth labels on the newly released Specialized Cyclist Dataset. We then discuss how heuristics can detect complex objects such as cyclists and by extension, situations, based on the output of existing object detection algorithms.
Zhenzhen Weng, Paroma Varma, Alexander Masalov, Jeffrey M. Ota, Christopher Ré
IV1