Zhong Li 0007

dblp:70/3488-7 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
21since 2021 · last 2025
0000-0002-7416-1216ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion
abstract
We present DualMat, a novel dual-path diffusion framework for estimating Physically Based Rendering (PBR) materials from single images under complex lighting conditions. Our approach operates in two distinct latent spaces: an albedo-optimized path leveraging pretrained visual knowledge through RGB latent space, and a material-specialized path operating in a compact latent space designed for precise metallic and roughness estimation. To ensure coherent predictions between the albedo-optimized and material-specialized paths, we introduce feature distillation during training. We employ rectified flow to enhance efficiency by reducing inference steps while maintaining quality. Our framework extends to high-resolution and multi-view inputs through patch-based estimation and cross-view attention, enabling seamless integration into image-to-3D pipelines. DualMat achieves state-of-the-art performance on both Objaverse and real-world data, significantly outperforming existing methods with up to 28% improvement in albedo estimation and 39% reduction in metallic-roughness prediction errors. Our project can be found at yifehuang97.github.io/DualMatProjPage/.
Yi Xu 0002, Minh Hoai, Zhong Li 0007
ACM Multimedia5
2025 Scalable High-Fidelity 3D Hand Shape Reconstruction via Graph-Image Frequency Mapping and Graph Frequency Decomposition
abstract
Despite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high-fidelity hand modeling is required, e.g., personalized hand modeling. To address this problem, we design a frequency split network to generate 3D hand meshes using different frequency bands in a coarse-to-fine manner. To capture high-frequency personalized details, we transform the 3D mesh into the frequency domain, and proposed a novel frequency decomposition loss to supervise each frequency component. By leveraging such a coarse-to-fine scheme, hand details that correspond to the higher frequency domain can be preserved. In addition, the proposed network is scalable, and can stop the inference at any resolution level to accommodate different hardware with varying computational powers. To feed the scalable frequency network with frequency split image features, we proposed an image-graph ring feature mapping strategy. To train our network with per-vertex supervision, we use a bidirectional registration strategy to generate a topology-fixed ground-truth. To quantitatively evaluate the performance of our method in terms of recovering personalized shape details, we introduce a new evaluation metric named Mean-frequency Signal-to-Noise Ratio (MSNR) to measure the mean signal-to-noise ratio of mesh signal on each frequency component. Extensive experiments demonstrate that our approach generates fine-grained details for high-fidelity 3D hand reconstruction, and our evaluation metric is more effective than traditional metrics for measuring mesh details.
Tianyu Luan, Yuanhao Zhai 0001, Jingjing Meng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Spectrum AUC Difference (SAUCD): Human-Aligned 3D Shape Evaluation
abstract
Existing 3D mesh shape evaluation metrics mainly focus on the overall shape but are usually less sensitive to local details. This makes them inconsistent with human evaluation, as human perception cares about both overall and detailed shape. In this paper, we propose an analytic metric named Spectrum Area Under the Curve Difference (SAUCD) that demonstrates better consistency with human evaluation. To compare the difference between two shapes, we first transform the 3D mesh to the spectrum domain using the discrete Laplace-Beltrami operator and Fourier transform. Then, we calculate the Area Under the Curve (AUC) difference between the two spectrums, so that each frequency band that captures either the overall or detailed shape is equitably considered. Taking human sensitivity across frequency bands into account, we further extend our metric by learning suitable weights for each frequency band which better aligns with human perception. To measure the performance of SAUCD, we build a 3D mesh evaluation dataset called Shape Grading, along with manual annotations from more than 800 subjects. By measuring the correlation between our metric and human evaluation, we demonstrate that SAUCD is well aligned with human evaluation, and outperforms previous 3D mesh metrics. Our project page: https://bit.ly/saucd.
Tianyu Luan, Zhong Li 0007, Lichang Chen, Yi Xu 0002, Junsong Yuan 0001
CVPR2
2024 PanoFree: Tuning-Free Holistic Multi-view Image Generation with Cross-View Self-guidance
Aoming Liu, Zhong Li 0007, Nannan Li 0004, Yi Xu 0002, Bryan A. Plummer
ECCV (27)2
2024 NeuSmoke: Efficient Smoke Reconstruction and View Synthesis with Neural Transportation Fields
Jiaxiong Qiu, Ruihong Cen, Zhong Li 0007, Han Yan 0012, Ming-Ming Cheng, Bo Ren 0003
SIGGRAPH Asia3
2024 NeuralTO: Neural Reconstruction and View Synthesis of Translucent Objects
abstract
Learning from multi-view images using neural implicit signed distance functions shows impressive performance on 3D Reconstruction of opaque objects. However, existing methods struggle to reconstruct accurate geometry when applied to translucent objects due to the non-negligible bias in their rendering function. To address the inaccuracies in the existing model, we have reparameterized the density function of the neural radiance field by incorporating an estimated constant extinction coefficient. This modification forms the basis of our innovative framework, which is geared towards highfidelity surface reconstruction and the novel-view synthesis of translucent objects. Our framework contains two stages. In the reconstruction stage, we introduce a novel weight function to achieve accurate surface geometry reconstruction. Following the recovery of geometry, the second phase involves learning the distinct scattering properties of the participating media to enhance rendering. A comprehensive dataset, comprising both synthetic and real translucent objects, has been built for conducting extensive experiments. Experiments reveal that our method outperforms existing approaches in terms of reconstruction and novel-view synthesis.
Jiaxiong Qiu, Zhong Li 0007, Bo Ren 0003
ACM Trans. Graph.3
2023 High Fidelity 3D Hand Shape Reconstruction via Scalable Graph Frequency Decomposition
abstract
Despite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high-fidelity hand modeling is required, e.g., personalized hand modeling. To address this problem, we design a frequency split network to generate 3D hand mesh using different frequency bands in a coarse-to-fine manner. To capture high-frequency personalized details, we transform the 3D mesh into the frequency domain, and propose a novel frequency decomposition loss to supervise each frequency component. By leveraging such a coarse-to-fine scheme, hand details that correspond to the higher frequency domain can be preserved. In addition, the proposed network is scalable, and can stop the inference at any resolution level to accommodate different hardware with varying computational powers. To quantitatively evaluate the performance of our method in terms of recovering personalized shape details, we introduce a new evaluation metric named Mean Signal-to-Noise Ratio (MSNR) to measure the signal-to-noise ratio of each mesh frequency component. Extensive experiments demonstrate that our approach generates fine-grained details for high-fidelity 3D hand reconstruction, and our evaluation metric is more effective for measuring mesh details compared with traditional metrics. The code is available at https://github.com/tyluann/FreqHand.
Tianyu Luan, Yuanhao Zhai 0001, Jingjing Meng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001
CVPR4
2023 3D-aware Facial Landmark Detection via Multi-view Consistent Training on Synthetic Data
abstract
Accurate facial landmark detection on wild images plays an essential role in human-computer interaction, entertainment, and medical applications. Existing approaches have limitations in enforcing 3D consistency while detecting 3D/2D facial landmarks due to the lack of multi-view in-the-wild training data. Fortunately, with the recent advances in generative visual models and neural rendering, we have witnessed rapid progress towards high quality 3D image synthesis. In this work, we leverage such approaches to construct a synthetic dataset and propose a novel multi-view consistent learning strategy to improve 3D facial landmark detection accuracy on in-the-wild images. The proposed 3D-aware module can be plugged into any learning-based landmark detection algorithm to enhance its accuracy. We demonstrate the superiority of the proposed plug-in module with extensive comparison against state-of-the-art methods on several real and synthetic datasets.
Libing Zeng, Wentao Bao, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001, Nima Khademi Kalantari
CVPR4
2023 Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting
abstract
Hand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space which is inadequate for 3D real-world applications. In this paper, we set up an egocentric 3D hand trajectory forecasting task that aims to predict hand trajectories in a 3D space from early observed RGB videos in a first-person view. To fulfill this goal, we propose an uncertainty-aware state space Transformer (USST) that takes the merits of the attention mechanism and aleatoric uncertainty within the framework of the classical state-space model. The model can be further enhanced by the velocity constraint and visual prompt tuning (VPT) on large vision transformers. Moreover, we develop an annotation workflow to collect 3D hand trajectories with high quality. Experimental results on H2O and EgoPAT3D datasets demonstrate the superiority of USST for both 2D and 3D trajectory forecasting. The code and datasets are publicly released: https://actionlab-cv.github.io/EgoHandTrajPred.
Wentao Bao, Libing Zeng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001, Yu Kong 0001
ICCV4
2023 NeuRBF: A Neural Fields Representation with Adaptive Radial Basis Functions
abstract
We present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for storing local neural features and N-dimensional linear kernels for interpolating features at continuous query points. The spatial positions of their neural features are fixed on grid nodes and cannot well adapt to target signals. Our method instead builds upon general radial bases with flexible kernel position and shape, which have higher spatial adaptivity and can more closely fit target signals. To further improve the channel-wise capacity of radial basis functions, we propose to compose them with multi-frequency sinusoid functions. This technique extends a radial basis to multiple Fourier radial bases of different frequency bands without requiring extra parameters, facilitating the representation of details. Moreover, by marrying adaptive radial bases with grid-based ones, our hybrid combination inherits both adaptivity and interpolation smoothness. We carefully designed weighting schemes to let radial bases adapt to different types of signals effectively. Our experiments on 2D image and 3D signed distance field representation demonstrate the higher accuracy and compactness of our method than prior arts. When applied to neural radiance field reconstruction, our method achieves state-of-the-art rendering quality, with small model size and comparable training speed.
Zhong Li 0007, Liangchen Song, Jingyi Yu 0001, Junsong Yuan 0001, Yi Xu 0002
ICCV2
2023 Relit-NeuLF: Efficient Relighting and Novel View Synthesis via Neural 4D Light Field
abstract
In this paper, we address the problem of simultaneous relighting and novel view synthesis of a complex scene from multi-view images with a limited number of light sources. We propose an analysis-synthesis approach called Relit-NeuLF. Following the recent neural 4D light field network (NeuLF)[22], Relit-NeuLF first leverages a two-plane light field representation to parameterize each ray in a 4D coordinate system, enabling efficient learning and inference. Then, we recover the spatially-varying bidirectional reflectance distribution function (SVBRDF) of a 3D scene in a self-supervised manner. A DecomposeNet learns to map each ray to its SVBRDF components: albedo, normal, and roughness. Based on the decomposed BRDF components and conditioning light directions, a RenderNet learns to synthesize the color of the ray. To self-supervise the SVBRDF decomposition, we encourage the predicted ray color to be close to the physically-based rendering result using the microfacet model. Comprehensive experiments demonstrate that the proposed method is efficient and effective on both synthetic data and real-world human face data, and outperforms the state-of-the-art results.
Zhong Li 0007, Liangchen Song, Xiangyu Du, Junsong Yuan 0001, Yi Xu 0002
ACM Multimedia1
2023 OpenIllumination: A Multi-Illumination Dataset for Inverse Rendering Evaluation on Real Objects
abstract
We introduce OpenIllumination, a real-world dataset containing over 108K images of 64 objects with diverse materials, captured under 72 camera views and a large number of different illuminations. For each image in the dataset, we provide accurate camera parameters, illumination ground truth, and foreground segmentation masks. Our dataset enables the quantitative evaluation of most inverse rendering and material decomposition methods for real objects. We examine several state-of-the-art inverse rendering methods on our dataset and compare their performances. The dataset and code can be found on the project page: https://oppo-us-research.github.io/OpenIllumination.
Isabella Liu, Ziyang Fu, Liwen Wu, Haian Jin, Zhong Li 0007, Chin Ming Ryan Wong, Yi Xu 0002, Ravi Ramamoorthi, Zexiang Xu, Hao Su 0001
NeurIPS6
2023 Full-Volume 3D Fluid Flow Reconstruction With Light Field PIV
abstract
Particle Imaging Velocimetry (PIV) is a classical method that estimates fluid flow by analyzing the motion of injected particles. To reconstruct and track the swirling particles is a difficult computer vision problem, as the particles are dense in the fluid volume and have similar appearances. Further, tracking a large number of particles is particularly challenging due to heavy occlusion. Here we present a low-cost PIV solution that uses compact lenslet-based light field cameras as imaging device. We develop novel optimization algorithms for dense particle 3D reconstruction and tracking. As a single light field camera has limited capacity in resolving depth (z-dimension measurement), the resolution of 3D reconstruction on the x-y plane is much higher than along the z-axis. To compensate for the imbalanced resolution in 3D, we use two light field cameras positioned at an orthogonal angle to capture particle images. In this way, we can achieve high-resolution 3D particle reconstruction in the full fluid volume. For each time frame, we first estimate particle depths under a single viewpoint by exploiting the focal stack symmetry of light field. We then fuse the recovered 3D particles in two views by solving a linear assignment problem (LAP). Specifically, we propose an anisotropic point-to-ray distance as matching cost to handle the resolution mismatch. Finally, given a sequence of 3D particle reconstructions over time, we recover the full-volume 3D fluid flow with a physically-constrained optical flow, which enforces local motion rigidity and fluid incompressibility. We perform comprehensive experiments on synthetic and real data for ablation and evaluation. We show that our method recovers full-volume 3D fluid flows of various types. Two-view reconstruction results achieves higher accuracy than those with one view only.
Yuqi Ding, Zhong Li 0007, Yu Ji 0001, Jingyi Yu 0001, Jinwei Ye
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 RobustFusion: Robust Volumetric Performance Reconstruction Under Human-Object Interactions From Monocular RGBD Stream
abstract
High-quality 4D reconstruction of human performance with complex interactions to various objects is essential in real-world scenarios, which enables numerous immersive VR/AR applications. However, recent advances still fail to provide reliable performance reconstruction, suffering from challenging interaction patterns and severe occlusions, especially for the monocular setting. To fill this gap, in this paper, we propose RobustFusion, a robust volumetric performance reconstruction system for human-object interaction scenarios using only a single RGBD sensor, which combines various data-driven visual and interaction cues to handle the complex interaction patterns and severe occlusions. We propose a semantic-aware scene decoupling scheme to model the occlusions explicitly, with a segmentation refinement and robust object tracking to prevent disentanglement uncertainty and maintain temporal consistency. We further introduce a robust performance capture scheme with the aid of various data-driven cues, which not only enables re-initialization ability, but also models the complex human-object interaction patterns in a data-driven manner. To this end, we introduce a spatial relation prior to prevent implausible intersections, as well as data-driven interaction cues to maintain natural motions, especially for those regions under severe human-object occlusions. We also adopt an adaptive fusion scheme for temporally coherent human-object reconstruction with occlusion analysis and human parsing cue. Extensive experiments demonstrate the effectiveness of our approach to achieve high-quality 4D human performance reconstruction under complex human-object interactions whilst still maintaining the lightweight monocular setting.
Zhuo Su 0006, Lan Xu 0003, Dawei Zhong, Zhong Li 0007, Fan Deng 0005, Shuxue Quan, Lu Fang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Real-Time Lighting Estimation for Augmented Reality via Differentiable Screen-Space Rendering
abstract
Augmented Reality (AR) applications aim to provide realistic blending between the real-world and virtual objects. One of the important factors for realistic AR is the correct lighting estimation. In this article, we present a method that estimates the real-world lighting condition from a single image in real time, using information from an optional support plane provided by advanced AR frameworks (e.g., ARCore, ARKit, etc.). By analyzing the visual appearance of the real scene, our algorithm can predict the lighting condition from the input RGB photo. In the first stage, we use a deep neural network to decompose the scene into several components: lighting, normal, and Bidirectional Reflectance Distribution Function (BRDF). Then we introduce differentiable screen-space rendering, a novel approach to providing the supervisory signal for regressing lighting, normal, and BRDF jointly. We recover the most plausible real-world lighting condition using Spherical Harmonics and the main directional lighting. Through a variety of experimental results, we demonstrate that our method can provide improved results than prior works quantitatively and qualitatively, and it can enhance the real-time AR experiences.
Celong Liu, Zhong Li 0007, Shuxue Quan, Yi Xu 0002
IEEE Trans. Vis. Comput. Graph.3
2023 NeRFPlayer: A Streamable Dynamic Scene Representation with Decomposed Neural Radiance Fields
abstract
Visually exploring in a real-world 4D spatiotemporal space freely in VR has been a long-term quest. The task is especially appealing when only a few or even single RGB cameras are used for capturing the dynamic scene. To this end, we present an efficient framework capable of fast reconstruction, compact modeling, and streamable rendering. First, we propose to decompose the 4D spatiotemporal space according to temporal characteristics. Points in the 4D space are associated with probabilities of belonging to three categories: static, deforming, and new areas. Each area is represented and regularized by a separate neural field. Second, we propose a hybrid representations based feature streaming scheme for efficiently modeling the neural fields. Our approach, coined NeRFPlayer, is evaluated on dynamic scenes captured by single hand-held cameras and multi-camera arrays, achieving comparable or superior rendering performance in terms of quality and speed comparable to recent state-of-the-art methods, achieving reconstruction in 10 seconds per frame and interactive rendering. Project website: https://bit.ly/nerfplayer.
Liangchen Song, Anpei Chen, Zhong Li 0007, Junsong Yuan 0001, Yi Xu 0002, Andreas Geiger 0001
IEEE Trans. Vis. Comput. Graph.3
2022 NeuLF: Efficient Novel View Synthesis with Neural 4D Light Field
abstract
In this paper, we present an efficient and robust deep learning solution for novel view synthesis of complex scenes. In our approach, a 3D scene is represented as a light field, i.e., a set of rays, each of which has a corresponding color when reaching the image plane. For efficient novel view rendering, we adopt a two-plane parameterization of the light field, where each ray is characterized by a 4D parameter. We then formulate the light field as a function that indexes rays to corresponding color values. We train a deep fully connected network to optimize this implicit function and memorize the 3D scene. Then, the scene-specific model is used to synthesize novel views. Different from previous light field approaches which require dense view sampling to reliably render novel views, our method can render novel views by sampling rays and querying the color for each ray from the network directly, thus enabling high-quality light field rendering with a sparser set of training images. Per-ray depth can be optionally predicted by the network, thus enabling applications such as auto refocus. Our novel view synthesis results are comparable to the state-of-the-arts, and even superior in some challenging scenes with refraction and reflection. We achieve this while maintaining an interactive frame rate and a small memory footprint.
Zhong Li 0007, Liangchen Song, Celong Liu, Junsong Yuan 0001, Yi Xu 0002
EGSR (ST)1
2022 PoP-Net: Pose over Parts Network for Multi-Person 3D Pose Estimation from a Depth Image
abstract
In this paper, a real-time method called PoP-Net is proposed to predict multi-person 3D poses from a depth image. PoP-Net learns to predict bottom-up part representations and top-down global poses in a single shot. Specifically, a new part-level representation, called Truncated Part Displacement Field (TPDF), is introduced which enables an explicit fusion process to unify the advantages of bottom-up part detection and global pose detection. Meanwhile, an effective mode selection scheme is introduced to automatically resolve the conflicting cases between global pose and part detections. Finally, due to the lack of high-quality depth datasets for developing multi-person 3D pose estimation, we introduce Multi-Person 3D Human Pose Dataset (MP-3DHP) as a new benchmark. MP-3DHP is designed to enable effective multi-person and background data augmentation in model training, and to evaluate 3D human pose estimators under uncontrolled multi-person scenarios. We show that PoP-Net achieves the state-of-the-art results both on MP-3DHP and on the widely used ITOP dataset, and has significant advantages in efficiency for multi-person processing. MP-3DHP Dataset and the evaluation code have been made available at: https://github.com/oppo-us-research/PoP-Net.
Yuliang Guo, Zhong Li 0007, Zekun Li 0011, Xiangyu Du, Shuxue Quan, Yi Xu 0002
WACV2
2021 Learning Kinematic Formulas from Multiple View Videos
abstract
Given a set of multiple view videos, which records the motion trajectory of an object, we propose to find out the objects' kinematic formulas with neural rendering techniques. For example, if the input multiple view videos record the free fall motion of an object with different initial speed v, the network aims to learn its kinematics: Δ=vt-1over 2 gt2, where Δ, g and t are displacement, gravitational acceleration and time. To achieve this goal, we design a novel framework consisting of a motion network and a differentiable renderer. For the differentiable renderer, we employ Neural Radiance Field (NeRF) since the geometry is implicitly modeled by querying coordinates in the space. The motion network is composed of a series of blending functions and linear weights, enabling us to analytically derive the kinematic formulas after training. The proposed framework is trained end to end and only requires knowledge of cameras' intrinsic and extrinsic parameters. To validate the proposed framework, we design three experiments to demonstrate its effectiveness and extensibility. The first experiment is the video of free fall and the framework can be easily combined with the principle of parsimony, resulting in the correct free fall kinematics. The second experiment is on the large angle pendulum which does not have analytical kinematics. We use the differential equation controlling pendulum dynamics as a physical prior in the framework and demonstrate that the convergence speed becomes much faster. Finally, we study the explosion animation and demonstrate that our framework can well handle such black-box-generated motions.
Liangchen Song, Sheng Liu 0017, Celong Liu, Zhong Li 0007, Yuqi Ding, Yi Xu 0002, Junsong Yuan 0001
ACM Multimedia4
2021 Animated 3D human avatars from a single image with GAN-based texture inference
Zhong Li 0007, Celong Liu, Fuyao Zhang, Zekun Li 0011, Yuanzhou Ha, Chenliang Xu, Shuxue Quan, Yi Xu 0002
Comput. Graph.1
2021 Structure From Motion on XSlit Cameras
abstract
We present a structure-from-motion (SfM) framework based on a special type of multi-perspective camera called the cross-slit or XSlit camera. Traditional perspective camera based SfM suffers from the scale ambiguity which is inherent to the pinhole camera geometry. In contrast, an XSlit camera captures rays passing through two oblique lines in 3D space and we show such ray geometry directly resolves the scale ambiguity when employed for SfM. To accommodate the XSlit cameras, we develop tailored feature matching, camera pose estimation, triangulation, and bundle adjustment techniques. Specifically, we devise a SIFT feature variant using non-uniform Gaussian kernels to handle the distortions in XSlit images for reliable feature matching. Moreover, we demonstrate that the XSlit camera exhibits ambiguities in pose estimation process which can not be handled by existing work. Consequently, we propose a 14 point algorithm to properly handle the XSlit degeneracy and estimate the relative pose between XSlit cameras from feature correspondences. We further exploit the unique depth-dependent aspect ratio (DDAR) property to improve the bundle adjustment for the XSlit camera. Synthetic and real experiments demonstrate that the proposed XSlit SfM can conduct reliable and high fidelity 3D reconstruction at an absolute scale.
Wei Yang 0034, Yingliang Zhang, Jinwei Ye, Yu Ji 0001, Zhong Li 0007, Mingyuan Zhou, Jingyi Yu 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Talking-Head Generation with Rhythmic Head Motion
Guofeng Cui, Celong Liu, Zhong Li 0007, Ziyi Kou, Yi Xu 0002, Chenliang Xu
ECCV (9)4
2020 3D Fluid Flow Reconstruction Using Compact Light Field PIV
Zhong Li 0007, Yu Ji 0001, Jingyi Yu 0001, Jinwei Ye
ECCV (16)1
2019 Pose2Body: Pose-Guided Human Parts Segmentation
abstract
Reliable human parts segmentation on 2D images plays an important role in many human-centric computer vision tasks. While significant achievements have been made on human pose estimation, the performance on human parts segmentation remains low. In this paper, we present a novel technique that we call Pose2Body that robustly conducts human parts segmentation based on the pose estimation results. We partition an image into superpixels and set out to assign a segment label to each superpixel most consistent with the pose. We design special feature vectors for every superpixel-label assignment as well as superpixel-superpixel pairs and model optimal labeling as to solve for a conditional random field (CRF). Comprehensive experiments show that our technique achieves substantial improvements over the state-of-the-art solutions.
Zhong Li 0007, Xin Chen 0040, Wangyiteng Zhou, Yingliang Zhang, Jingyi Yu 0001
ICME1
2018 4D Human Body Correspondences From Panoramic Depth Maps
abstract
The availability of affordable 3D full body reconstruction systems has given rise to free-viewpoint video (FVV) of human shapes. Most existing solutions produce temporally uncorrelated point clouds or meshes with unknown point/vertex correspondences. Individually compressing each frame is ineffective and still yields to ultra-large data sizes. We present an end-to-end deep learning scheme to establish dense shape correspondences and subsequently compress the data. Our approach uses sparse set of "panoramic" depth maps or PDMs, each emulating an inward-viewing concentric mosaics (CM) [45]. We then develop a learning-based technique to learn pixel-wise feature descriptors on PDMs. The results are fed into an autoencoder-based network for compression. Comprehensive experiments demonstrate our solution is robust and effective on both public and our newly captured datasets.
Zhong Li 0007, Minye Wu, Wangyiteng Zhou, Jingyi Yu 0001
CVPR1
2017 Robust 3D Human Motion Reconstruction via Dynamic Template Construction
abstract
In multi-view human body capture systems, the recovered 3D geometry or even the acquired imagery data can be heavily corrupted due to occlusions, noise, limited fieldof- view, etc. Direct estimation of 3D pose, body shape or motion on these low-quality data has been traditionally challenging.In this paper, we present a graph-based non-rigid shape registration framework that can simultaneously recover 3D human body geometry and estimate pose/motion at high fidelity.Our approach first generates a global full-body template by registering all poses in the acquired motion sequence.We then construct a deformable graph by utilizing the rigid components in the global template. We directly warp the global template graph back to each motion frame in order to fill in missing geometry. Specifically,we combine local rigidity and temporal coherence constraints to maintain geometry and motion consistencies. Comprehensive experiments on various scenes show that our method is accurate and robust even in the presence of drastic motions.
Zhong Li 0007, Yu Ji 0001, Wei Yang 0034, Jinwei Ye, Jingyi Yu 0001
3DV1
2017 The light field 3D scanner
abstract
We present a novel light field structure-from-motion (SfM) framework for reliable 3D object reconstruction. Specifically, we use the light field (LF) camera such as Lytro and Raytrix as a virtual 3D scanner. We move an LF camera around the object and register between multiple LF shots. We show that applying conventional SfM on sub-aperture images is not only expensive but also unreliable due to ultra-small baseline and low image resolution. Instead, our LF-SfM scheme maps ray manifolds across LFs. Specifically, we show how rays passing through a common 3D point transform between two LFs and we develop reliable technique for extracting extrinsic parameters from this ray transform. Next, we apply a new edge-preserving stereo matching technique on individual LFs and conduct LF bundle adjustment to jointly optimize pose and geometry. Comprehensive experiments show our solution outperforms many state-of-the-art passive and even active techniques especially on topologically complex objects.
Yingliang Zhang, Zhong Li 0007, Wei Yang 0034, Peihong Yu, Haiting Lin, Jingyi Yu 0001
ICCP2