EDBT 2026 Demo / reviewers in the wild / expert
Feng Xu 0005
dblp:03/2611-5
· DBLP profile ↗
92ranked-venue papers
7as first author
60since 2021 · last 2026
0000-0002-0953-1057ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 76 · 5 first-author · 45 since 2021Artificial intelligence and machine learning · 28 · 2 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the WildabstractReconstructing topologically consistent facial geometry is crucial for the digital avatar creation pipelines. Existing methods either require tedious manual efforts, lack generalization to in-the-wild data, or are constrained by the limited expressiveness of 3D Morphable Models. To address these limitations, we propose VGGTFace, an automatic approach that innovatively applies the 3D foundation model, i.e. VGGT, for topologically consistent facial geometry reconstruction from in-the-wild multi-view images captured by everyday users. Our key insight is that, by leveraging VGGT, our method naturally inherits strong generalization ability and expressive power from its large-scale training and point map representation. However, it is unclear how to reconstruct a topologically consistent mesh from VGGT, as the topology information is missing in its prediction. To this end, we augment VGGT with Pixel3DMM for injecting topology information via pixel-aligned UV values. In this manner, we convert the pixel-aligned point map of VGGT to a point cloud with topology. Tailored to this point cloud with known topology, we propose a novel Topology-Aware Bundle Adjustment strategy to fuse them, where we construct a Laplacian energy for the Bundle Adjustment objective. Our method achieves high-quality reconstruction in 10 seconds for 16 views on a single NVIDIA RTX 4090. Experiments demonstrate state-of-the-art results on benchmarks and impressive generalization to in-the-wild data. Xin Ming, Feng Xu 0005 |
AAAI | 4 |
| 2026 | High-Quality Facial Geometry and Appearance Capture at Home
Feng Xu 0005, Junfeng Lyu |
Int. J. Comput. Vis. | 1 |
| 2026 | MAT: Mixing Attention Transfer From Multiple Transformers for Medical TasksabstractTransformer has been widely used for image analysis tasks, but in medicine, it suffers from limited data availability. To overcome this challenge, we propose a novel approach specially designed for transformers to transfer knowledge from multiple sources to target medical tasks with limited data, named Mixing Attention Transfer (MAT). MAT aims to harness and merge knowledge from multiple source transformers at the token and layer level to improve the performance of target medical tasks. The core component of MAT is the Mixing Attention layer, which encompasses: 1) token-level Routing and Fusion modules that allocate input images to adequate source modules; 2) sequence-level Aligned-Attention module that adaptively aligns outputs produced by different source modules. To the best of our knowledge, this is the first multi-source transfer learning approach specifically designed for transformers. Through extensive evaluations, we demonstrate the effectiveness of MAT on three medical scenarios: noisy-labeled, class-imbalanced, and fine-grained tasks. Zihao Bo, Lishan Ye, Feng Xu 0005 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Learning a Delighting Prior for Facial Appearance Capture in the WildabstractHigh-quality facial appearance capture has traditionally required costly studio recording. Recent works consider an in-the-wild smartphone-based setup; however, their model-based inverse rendering paradigm struggles with the complex disentanglement of reflectance from unknown illumination. To bridge this gap, we propose to shift the paradigm into training a powerful delighting network as a prior to constrain the optimization. We leverage the OLAT dataset and the rendered Light Stage scans for training, and propose Dataset Latent Modulation (DLM) to seamlessly integrate these heterogeneous data sources. Specifically, by conditioning the core network on learnable source-aware tokens, we decouple dataset-specific styles from physical delighting principles, enabling the emergence of a delighting prior that outperforms existing proprietary models. This powerful delighting prior enables a simple and automatic appearance capture pipeline that achieves high-quality reflectance estimation from casual video inputs, outperforming prior arts by a large margin. Furthermore, we leverage our appearance capture method to transform the multi-view NeRSemble dataset into NeRSemble-Scan, a large-scale collection of 4K-resolution relightable scans. By open-sourcing our model and the NeRSemble-Scan dataset, we democratize high-end facial capture and provide a new foundation for the research community to build photorealistic digital humans. Xin Ming, Zhuofan Shen, Qixuan Zhang, Lan Xu 0003, Feng Xu 0005 |
ACM Trans. Graph. | 7 |
| 2026 | Sample Matching for Joint Extinction Gradient Estimation in Differentiable Volume RenderingabstractDifferentiable volume rendering enables gradient-based optimization of volumetric scenes, but unbiased estimators suffer from high gradient variance. We observe that the extinction gradients split into two components on structurally different integration domains: a scattering term evaluated at a single path vertex, and a transmittance term integrated along the ray segment. Because the domains are mismatched, existing estimators sample the two components at different locations, leaving the negative correlation between their opposite-signed contributions unexploited. We expose this overlooked correlation and exploit it through a principle we call sample matching : evaluate both components at shared sample locations. To enable this, we derive the first reformulation of the differential path integral that couples the two contributions within a single integrand, yielding an unbiased Monte Carlo estimator that ties them together by construction. For efficiency, the estimator reuses partially sampled light paths and amortizes in-scattering cost by evaluating gradients at multiple probe points per segment. On voxel-grid reconstruction, our estimator reduces gradient variance by up to 80% over differential ratio tracking (DRT), yielding faster convergence and higher reconstruction quality. Ruihan Yu, Jingwang Ling, Feng Xu 0005 |
ACM Trans. Graph. | 4 |
| 2026 | Effective Gaussian Management for High-Fidelity Scene ReconstructionabstractThis paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent Gaussian Splatting (GS) pipelines that treat all primitives uniformly during optimization, our framework explicitly manages the attribute activation, representation and pruning of Gaussian. Specifically, our framework first introduces GauSep, a novel densification strategy that selectively activates Gaussian color or normal attributes to alleviate destructive gradient conflicts arising from dual supervision. We further propose GauRep, an adaptive Gaussian representation that dynamically adjusts spherical harmonics (SHs) orders and performs task-decoupled pruning to reduce redundancy at both the individual and global levels. To provide reliable geometric supervision for above mangement process, we additionally introduce CoRe, an regularized surface reconstruction module that distills robust normal fields from an SDF branch to the Gaussian representation through a confidence mechanism. Notably, the proposed Gaussian management is compatible with various reconstruction architectures and can be seamlessly integrated to improve performance while reducing size of the model. Extensive experiments demonstrate that our approach achieves superior or comparable performance in appearance and geometry reconstruction compared with state-of-the-art methods, while using significantly fewer parameters. Jiateng Liu, Hao Gao 0005, Jiucheng Xie, Chi-Man Pun, Jian Xiong 0005, Haolun Li 0001, Junxin Chen 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2026 | Errata to "DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular Camera"abstractIn the originally published version of this article, the Acknowledgements section was inadvertently omitted. The correct Acknowledgements are provided below. Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | 2DGH: 2D Gaussian-Hermite Splatting for High-Quality Rendering and Better Geometry Featuresabstract2D Gaussian Splatting has recently emerged as a significant method in 3D reconstruction, enabling novel view synthesis and geometry reconstruction simultaneously. While the well-known Gaussian kernel is broadly used, its lack of anisotropy and deformation ability leads to dim and vague edges at object silhouettes, limiting the reconstruction quality of current Gaussian splatting methods. To enhance the representation power, we draw inspiration from quantum physics and propose to use the Gaussian-Hermite kernel as the new primitive in Gaussian splatting. The new kernel takes a unified mathematical form and extends the Gaussian function, which serves as the zero-rank special case in the updated general formulation. Our experiments demonstrate that the proposed Gaussian-Hermite kernel achieves improved performance over traditional Gaussian Splatting kernels on both geometry reconstruction and novel-view synthesis tasks. Specifically, on the DTU dataset, our method yields more accurate geometry reconstruction, while on datasets such as MipNeRF360 and our customized Detail dataset, it achieves better results in novel-view synthesis. These results highlight the potential of the Gaussian-Hermite kernel for high-quality 3D reconstruction and rendering. Ruihan Yu, Jingwang Ling, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Refined Mamba-Based Lower Limbs Motor Estimator for Parkinson's Disease DiagnosticabstractBradykinesia, a key Parkinson's disease (PD) symptom, requires accurate lower limbs assessment, yet current clinical assessments are subjective and biased, while computer vision methods lack precision in skeleton extraction and PD-specific movement analysis. Furthermore, clothing-induced foot occlusion further aggravates keypoint localization errors. To address these gaps, we propose a vision-assisted diagnostic framework for PD lower limbs assessment. Our approach incorporates the CSDPose model, which employs CNN for local feature extraction, SSM (State Space Model) for global feature capture, and DCT (Discrete Cosine Transform) for frequency domain analysis, to enhance 2D pose estimation accuracy. These keypoints are used to compute objective PD motor indicators that we proposed, which are subsequently analyzed by a classification model to grade lower limbs dysfunction severity. Experiments demonstrate the algorithm achieves over$90\%$accuracy in classifying lower limbs motions. Validated clinical trials confirm that the automated severity ratings and motor indicators effectively support diagnostic decision-making. Xinyuan Dong, Hao Gao 0005, Yikang He, Yiqin Yao, Chi-Man Pun, Haolun Li 0001, Feng Xu 0005 |
BIBM | 8 |
| 2025 | A Noise-Resistant 3D Hand Motion Estimator Framework for Parkinson's Tremor AssessmentabstractParkinson's disease (PD) is a progressive neurodegenerative disorder, with tremor being one of its representative motor symptoms. Current clinical evaluations primarily rely on subjective scales such as the MDS-UPDRS, which often introduce significant inter-rater variability. Although vision-based evaluation offers objective motion analysis, existing pose tracking frameworks struggle with accurate tremor quantification due to inter-frame jitter and limited precision, failing to capture fine-grained spatiotemporal dynamics. To address this, we propose Motion-aware Hierarchical Grouping Mamba Network (MHG-Mamba), a non-contact video-based framework for automated evaluation of the 'finger-to-nose' task. Our approach first employs feature extraction to decouple high-degree-of-freedom finger joint movements from stable palm joint motions, enabling accurate finger pose estimation. Second, by incorporating with the hierarchical spatiotemporal scanning mechanism in Mamba's SSM, the model captures global motion features while preserving anatomical constraints, resulting in a temporally smooth and plausible skeletal sequence. Finally, based on the predicted skeletal sequences, we introduce several objective metrics to quantify motion features and apply a classifier for precise objective severity rating. Experimental results demonstrate that MHG-Mamba significantly improves the accuracy of 3D hand pose estimation and reduces noise in the motion sequences. The system achieved a classification accuracy of 93.2 % on the 'finger-to-nose' task. Moreover, clinicians using our system exhibited reduced variability in their assessments, highlighting its high clinical value. Yixing Ye, Hao Gao 0005, Yikang He, Haolun Li 0001, Chi-Man Pun, Feng Xu 0005 |
BIBM | 7 |
| 2025 | MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic DisturbancesabstractThis paper proposes a novel method called MagShield, designed to address the issue of magnetic interference in sparse inertial motion capture (MoCap) systems. Existing Inertial Measurement Unit (IMU) systems are prone to orientation estimation errors in magnetically disturbed environments, limiting their practical application in real-world scenarios. To address this problem, MagShield employs a "detect-then-correct" strategy, first detecting magnetic disturbances through multi-IMU joint analysis, and then correcting orientation errors using human motion priors. MagShield can be integrated with most existing sparse inertial MoCap systems, improving their performance in magnetically disturbed environments. Experimental results demonstrate that MagShield significantly enhances the accuracy of motion capture under magnetic interference and exhibits good compatibility across different sparse inertial MoCap systems. Yunzhe Shao, Xinyu Yi, Shihui Guo, Jun-Hai Yong, Feng Xu 0005 |
ICCV | 6 |
| 2025 | Teeth Reconstruction and Performance Capture Using a Phone Camera
Weixi Zheng, Jingwang Ling, Zhibo Wang 0003, Feng Xu 0005 |
ICCV | 5 |
| 2025 | MotionRefineNet: Fine-Grained Pose Sequence Smoothing and RefinementabstractCapturing human motion with existing monocular estimators often results in large errors when dealing with rare poses, occlusions, truncations, and frame blurring, leading to jitter and long-term drift. Although previous methods have introduced post-processing networks for pose refinement, they struggle to balance global smoothing and fine-grained correction. In this work, we propose MotionRefineNet, which leverages the synergy and complementarity between long- and short-term features in the temporal domain and high- and low-frequency features in the frequency domain to address these challenges. The temporal branch is designed as a hierarchical motion structure to learn multi-time scale features, where long-term features learn motion smoothness, and short-term features capture local rapid changes. The frequency branch employs different frequency band learning strategies based on the degrees of freedom (DoF) of body parts. For body parts with low DoF, the focus is on low-frequency features that represent overall motion trends and regular actions. For body parts with high DoF, we design a filter to adaptively extract useful information from all frequency bands, including subtle motion changes in the high-frequency bands. Extensive experiments on multiple datasets and estimators demonstrate that MotionRefineNet outperforms existing methods in refining 2D, 3D, and SMPL poses, achieving superior pose smoothing and deviation correction. Our code is available at: https://github.com/Wheels319/MotionRefineNet. Haolun Li 0001, Weihuang Liu, Jiateng Liu, Zhenhua Tang 0001, Chi-Man Pun, Qiguang Miao, Feng Xu 0005, Hao Gao 0005 |
ACM Multimedia | 7 |
| 2025 | BaroPoser: Real-time Human Motion Tracking from IMUs and Barometers in Everyday Devices
Xinyu Yi, Feng Xu 0005 |
UIST | 3 |
| 2025 | Facial Appearance Capture at Home with Patch-Level Reflectance PriorabstractExisting facial appearance capture methods can reconstruct plausible facial reflectance from smartphone-recorded videos. However, the reconstruction quality is still far behind the ones based on studio recordings. This paper fills the gap by developing a novel daily-used solution with a co-located smartphone and flashlight video capture setting in a dim room. To enhance the quality, our key observation is to solve facial reflectance maps within the data distribution of studio-scanned ones. Specifically, we first learn a diffusion prior over the Light Stage scans and then steer it to produce the reflectance map that best matches the captured images. We propose to train the diffusion prior at the patch level to improve generalization ability and training stability, as current Light Stage datasets are in ultra-high resolution but limited in data size. Tailored to this prior, we propose a patch-level posterior sampling technique to sample seamless full-resolution reflectance maps from this patch-level diffusion model. Experiments demonstrate our method closes the quality gap between low-cost and studio recordings by a large margin, opening the door for everyday users to clone themselves to the digital world. Junfeng Lyu, Kuan Sheng, Minghao Que, Qixuan Zhang, Lan Xu 0003, Feng Xu 0005 |
ACM Trans. Graph. | 7 |
| 2025 | Improving Global Motion Estimation in Sparse IMU-based Motion Capture with PhysicsabstractBy learning human motion priors, motion capture can be achieved by 6 inertial measurement units (IMUs) in recent years with the development of deep learning techniques, even though the sensor inputs are sparse and noisy. However, human global motions are still challenging to be reconstructed by IMUs. This paper aims to solve this problem by involving physics. It proposes a physical optimization scheme based on multiple contacts to enable physically plausible translation estimation in the full 3D space where the z-directional motion is usually challenging for previous works. It also considers gravity in local pose estimation which well constrains human global orientations and refines local pose estimation in a joint estimation manner. Experiments demonstrate that our method achieves more accurate motion capture for both local poses and global motions. Furthermore, by deeply integrating physics, we can also estimate 3D contact, contact forces, joint torques, and interacting proxy surfaces. Code is available at https://xinyu-yi.github.io/GlobalPose/. Xinyu Yi, Shaohua Pan 0002, Feng Xu 0005 |
ACM Trans. Graph. | 3 |
| 2025 | Shape-aware Inertial Poser: Motion Tracking for Humans with Diverse Shapes Using Sparse Inertial SensorsabstractHuman motion capture with sparse inertial sensors has gained significant attention recently. However, existing methods almost exclusively rely on a template adult body shape to model the training data, which poses challenges when generalizing to individuals with largely different body shapes (such as a child). This is primarily due to the variation in IMU-measured acceleration caused by changes in body shape. To fill this gap, we propose Shape-aware Inertial Poser (SAIP), the first solution considering body shape differences in sparse inertial-based motion capture. Specifically, we decompose the sensor measurements related to shape and pose in order to effectively model their joint correlations. Firstly, we train a regression model to transfer the IMU-measured accelerations of a real body to match the template adult body model, compensating for the shape-related sensor measurements. Then, we can easily follow the state-of-the-art methods to estimate the full body motions of the template-shaped body. Finally, we utilize a second regression model to map the joint velocities back to the real body, combined with a shape-aware physical optimization strategy to calculate global motions on the subject. Furthermore, our method relies on body shape awareness, introducing the first inertial shape estimation scheme. This is accomplished by modeling the shape-conditioned IMU-pose correlation using an MLP-based network. To validate the effectiveness of SAIP, we also present the first IMU motion capture dataset containing individuals of different body sizes. This dataset features 10 children and 10 adults, with heights ranging from 110 cm to 190 cm, and a total of 400 minutes of paired IMU-Motion samples. Extensive experimental results demonstrate that SAIP can effectively handle motion capture tasks for diverse body shapes. The code and dataset are available at https://github.com/yinlu5942/SAIP . Ziying Shi, Yinghao Wu, Xinyu Yi, Feng Xu 0005, Shihui Guo |
ACM Trans. Graph. | 5 |
| 2025 | Transformer IMU Calibrator: Dynamic On-body IMU Calibration for Inertial Motion CaptureabstractIn this paper, we propose a novel dynamic calibration method for sparse inertial motion capture systems, which is the first to break the restrictive absolute static assumption in IMU calibration, i.e., the coordinate drift R G′ G and measurement offset R BS remain constant during the entire motion, thereby significantly expanding their application scenarios. Specifically, we achieve real-time estimation of R G′ G and R BS under two relaxed assumptions: i) the matrices change negligibly in a short time window; ii) the human movements/IMU readings are diverse in such a time window. Intuitively, the first assumption reduces the number of candidate matrices, and the second assumption provides diverse constraints, which greatly reduces the solution space and allows for accurate estimation of R G′ G and R BS from a short history of IMU readings in real time. To achieve this, we created synthetic datasets of paired R G′ G , R BS matrices and IMU readings, and learned their mappings using a Transformer-based model. We also designed a calibration trigger based on the diversity of IMU readings to ensure that assumption ii) is met before applying our method. To our knowledge, we are the first to achieve implicit IMU calibration (i.e., seamlessly putting IMUs into use without the need for an explicit calibration process), as well as the first to enable long-term and accurate motion capture using sparse IMUs. The code and dataset are available at https://github.com/ZuoCX1996/TIC. Chengxu Zuo, Xiangren Shi, Xinyu Yi, Feng Xu 0005, Shihui Guo, Yipeng Qin |
ACM Trans. Graph. | 8 |
| 2025 | DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular CameraabstractCombining sparse IMUs and a monocular camera is a new promising setting to perform real-time human motion capture. This paper proposes a diffusion-based solution to learn human motion priors and fuse the two modalities of signals together seamlessly in a unified framework. By delicately considering the characteristics of the two signals, the sequential visual information is considered as a whole and transformed into a condition embedding, while the inertial measurement is concatenated with the noisy body pose frame by frame to construct a sequential input for the diffusion model. Firstly, we observe that the visual information may be unavailable in some frames due to occlusions or subjects moving out of the camera view. Thus incorporating the sequential visual features as a whole to get a single feature embedding is robust to the occasional degenerations of visual information in those frames. On the other hand, the IMU measurements are robust to occlusions and always stable when signal transmission has no problem. So incorporating them frame-wisely could better explore the temporal information for the system. Experiments have demonstrated the effectiveness of the system design and its state-of-the-art performance in pose estimation compared with the previous works. The code will be released. Shaohua Pan 0002, Xinyu Yi, Yan Zhou 0003, Weihua Jian, Yuan Zhang 0020, Pengfei Wan 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | GaussianHead: High-Fidelity Head Avatars With Learnable Gaussian DerivationabstractCreating lifelike 3D head avatars and generating compelling animations for diverse subjects remain challenging in computer vision. This paper presents GaussianHead, which models the active head based on anisotropic 3D Gaussians. Our method integrates a motion deformation field and a single-resolution tri-plane to capture the head's intricate dynamics and detailed texture. Notably, we introduce a customized derivation scheme for each 3D Gaussian, facilitating the generation of multiple "doppelgangers" through learnable parameters for precise position transformation. This approach enables efficient representation of diverse Gaussian attributes and ensures their precision. Additionally, we propose an inherited derivation strategy for newly added Gaussians to expedite training. Extensive experiments demonstrate GaussianHead's efficacy, achieving high-fidelity visual results with a remarkably compact model size ($\approx 12$≈12 MB). Our method outperforms state-of-the-art alternatives in tasks such as reconstruction, cross-identity reenactment, and novel view synthesis. Jie Wang 0137, Jiucheng Xie, Xianyan Li, Feng Xu 0005, Chi-Man Pun, Hao Gao 0005 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Relightable and Animatable Neural Avatars from VideosabstractLightweight creation of 3D digital avatars is a highly desirable but challenging task. With only sparse videos of a person under unknown illumination, we propose a method to create relightable and animatable neural avatars, which can be used to synthesize photorealistic images of humans under novel viewpoints, body poses, and lighting. The key challenge here is to disentangle the geometry, material of the clothed body, and lighting, which becomes more difficult due to the complex geometry and shadow changes caused by body motions. To solve this ill-posed problem, we propose novel techniques to better model the geometry and shadow changes. For geometry change modeling, we propose an invertible deformation field, which helps to solve the inverse skinning problem and leads to better geometry quality. To model the spatial and temporal varying shading cues, we propose a pose-aware part-wise light visibility network to estimate light occlusion. Extensive experiments on synthetic and real datasets show that our approach reconstructs high-quality geometry and generates realistic shadows under different body poses. Code and data are available at https://wenbin-lin.github.io/RelightableAvatar-page. Chengwei Zheng, Jun-Hai Yong, Feng Xu 0005 |
AAAI | 4 |
| 2024 | An Automatic Assessment of Parkinson's Disease in Arising from Chair Task via Refined Diffusion-based Pose EstimatorabstractParkinson’s disease (PD) is a progressively common neurodegenerative disorder characterized by a decline in motor function. The diagnosis of PD typically relies on the Movement Disorder Society-Unified Parkinson’s Disease Rating Scale (MDS-UPDRS), which involves subjective scoring through observation of targeted movements. However, this objective method heavily depends on professional experience and has relatively high misdiagnosis rates. In this paper, we introduce a novel vision-based architecture for automated assessment of the ‘arising from chair’ task, which is one of the key MDS-UPDRS components. First, a diffusion-based 2D pose estimator is proposed to enhance keypoint accuracy by iteratively learning the distribution of ground-truth data and then denoising noisy poses. Second, a keypoint trajectory refinement network is introduced to eliminate the jitter error by considering motion information such as position, velocity, acceleration, and jerk. Finally, based on the predicted skeleton keypoint trajectories, we propose several objective indicators to assess the movement characteristics and perform the final rating using the classifier. The experiment substantiates the proposed algorithm, achieving a precision of 98.7% and an accuracy of 95.8% in classifying the ‘arising from chair’ task. Furthermore, the classification results and the proposed objective indicators have been validated as effective aids for neurologists to provide more precise diagnoses. Chi-Man Pun, Haolun Li 0001, Mingliang Zhai, Feng Xu 0005, Hao Gao 0005 |
BIBM | 5 |
| 2024 | High-Quality Facial Geometry and Appearance Capture at HomeabstractFacial geometry and appearance capture have demonstrated tremendous success in 3D scanning real humans in studios. Recent works propose to democratize this technique while keeping the results high quality. However, they are still inconvenient for daily usage. In addition, they focus on an easier problem of only capturing facial skin. This paper proposes a novel method for high-quality face capture, featuring an easy-to-use system and the capability to model the complete face with skin, mouth interior, hair, and eyes. We reconstruct facial geometry and appearance from a single co-located smartphone flashlight sequence captured in a dim room where the flashlight is the dominant light source (e.g. rooms with curtains or at night). To model the complete face, we propose a novel hybrid representation to effectively model both eyes and other facial regions, along with novel techniques to learn it from images. We apply a combined lighting model to compactly represent real illuminations and exploit a morphable face albedo model as a reflectance prior to disentangle diffuse and specular. Experiments show that our method can capture high-quality 3D relightable scans. Our code will be released. Junfeng Lyu, Feng Xu 0005 |
CVPR | 3 |
| 2024 | Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear JacketabstractExisting wearable motion capture methods typically demand tight on-body fixation (often using straps) for reliable sensing, limiting their application in everyday life. In this paper, we introduce Loose Inertial Poser, a novel motion capture solution with high wearing comfortableness, by integrating four Inertial Measurement Units (IMUs) into a loose-wear jacket. Specifically, we address the challenge of scarce loose-wear IMU training data by proposing a Secondary Motion AutoEncoder (SeMo-AE) that learns to model and synthesize the effects of secondary motion between the skin and loose clothing on IMU data. SeMo-AE is leveraged to generate a diverse synthetic dataset of loose-wear IMU data to augment training for the pose estimation network and significantly improve its accuracy. For validation, we collected a dataset with various subjects and 2 wearing styles (zipped and unzipped). Experimental results demonstrate that our approach maintains high-quality real-time posture estimation even in loose-wear scenarios. Our dataset and code are available at: https://github.com/ZuoCX1966/Loose-Inertial-Poser Chengxu Zuo, Lishuang Zhan, Shihui Guo, Xinyu Yi, Feng Xu 0005, Yipeng Qin |
CVPR | 6 |
| 2024 | High-Quality Mesh Blendshape Generation from Face Videos via Neural Inverse Rendering
Xin Ming, Jingwang Ling, Feng Xu 0005 |
ECCV (70) | 5 |
| 2024 | High-quality Animatable Eyelid Shapes from Lightweight CapturesabstractHigh-quality eyelid reconstruction and animation are challenging for the subtle details and complicated deformations. Previous works usually suffer from the trade-off between the capture costs and the quality of details. In this paper, we propose a novel method that can achieve detailed eyelid reconstruction and animation by only using an RGB video captured by a mobile phone. Our method utilizes both static and dynamic information of eyeballs (e.g., positions and rotations) to assist the eyelid reconstruction, cooperating with an automatic eyeball calibration method to get the required eyeball parameters. Furthermore, we develop a neural eyelid control module to achieve the semantic animation control of eyelids. To the best of our knowledge, we present the first method for high-quality eyelid reconstruction and animation from lightweight captures. Extensive experiments on both synthetic and real data show that our method can provide more detailed and realistic results compared with previous methods based on the same-level capture setups. The code is available at https://github.com/StoryMY/AniEyelid. Junfeng Lyu, Feng Xu 0005 |
SIGGRAPH Asia | 2 |
| 2024 | Ethics-aware face recognition aided by synthetic face imagesabstractCurrent face recognition models trained on large-scale face datasets have achieved promising performance. However, using face images to train a face recognition model without consent would lead to severe privacy and ethical issues. Moreover, existing face recognition models also exhibit uneven performance on different races, thus perplexing vulnerable populations. To address the aforementioned two issues, this work investigates an ethics-aware face recognition method and examines whether we can leverage synthesized faces to achieve a high-accuracy racial balanced recognition model. In a nutshell, we introduce a race-controllable and identity-innumerable face synthesis approach to generate synthetic face images, and then employ the synthesized images to improve face recognition accuracy and mitigate recognition imbalance among different races despite the scarcity of consenting images (less than 100 individuals). More importantly, the synthetic data enable us to analyze the potential impacts of races on face recognition models quantitatively and facilitate the eradication of racial imbalance in face recognition. Extensive experiments demonstrate that employing our synthetic face data improves face recognition accuracy by a large margin while mitigating the recognition imbalance across different race groups. Xiaobiao Du, Xin Yu 0002, Beifen Dai, Feng Xu 0005 |
Neurocomputing | 5 |
| 2024 | EditableNeRF: Editing Topologically Varying Neural Radiance Fields by Key PointsabstractNeural radiance fields (NeRF) achieve highly photo-realistic novel-view synthesis, but it's a challenging problem to edit the scenes modeled by NeRF-based methods, especially for dynamic scenes. We propose editable neural radiance fields that enable end-users to easily edit dynamic scenes and support topological changes. Input with an image sequence from a single camera, our network is trained automatically and models topologically varying dynamics using our picked-out surface key points. Then end-users can edit the scene by easily dragging the key points to desired new positions. To achieve this, we propose a scene analysis method to detect and initialize key points by considering the dynamics in the scene, and a weighted key points strategy to model topologically varying dynamics by joint key points and weights optimization. Our method supports intuitive multi-dimensional (up to 3D) editing and can generate novel scenes that are unseen in the input sequence. Experiments demonstrate that our method achieves high-quality editing on various dynamic scenes and outperforms the state-of-the-art. Chengwei Zheng, Feng Xu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | A Parkinson's Auxiliary Diagnosis Algorithm Based on a Hyperparameter Optimization Method of Deep LearningabstractParkinson's disease is a common mental disease in the world, especially in the middle-aged and elderly groups. Today, clinical diagnosis is the main diagnostic method of Parkinson's disease, but the diagnosis results are not ideal, especially in the early stage of the disease. In this paper, a Parkinson's auxiliary diagnosis algorithm based on a hyperparameter optimization method of deep learning is proposed for the Parkinson's diagnosis. The diagnosis system uses ResNet50 to achieve feature extraction and Parkinson's classification, mainly including speech signal processing part, algorithm improvement part based on Artificial Bee Colony algorithm (ABC) and optimizing the hyperparameters of ResNet50 part. The improved algorithm is called Gbest Dimension Artificial Bee Colony algorithm (GDABC), proposing "Range pruning strategy" which aims at narrowing the scope of search and "Dimension adjustment strategy" which is to adjust gbest dimension by dimension. The accuracy of the diagnosis system in the verification set of Mobile Device Voice Recordings at King's College London (MDVR-CKL) dataset can reach more than 96%. Compared with current Parkinson's sound diagnosis methods and other optimization algorithms, our auxiliary diagnosis system shows better classification performance on the dataset within limited time and resources. Shujuan Li, Chi-Man Pun, Yijing Guo, Feng Xu 0005, Hao Gao 0005, Huimin Lu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | UniFRD: A Unified Method for Facial Image Restoration Based on Diffusion Probabilistic ModelabstractThis paper presents a Unified Facial image and video Restoration method based on the Diffusion probabilistic model (UniFRD), designed to effectively address both single- and multi-type image degradation. The noise predictor in UniFRD consists of a ViT-based encoder and a novel Separation Fusion Decoding Module (SFDM). The flexible feature optimization strategy allows for decoding complex conditional noise without being limited by degradation patterns. Specifically, SFDM adjusts and refines the channel correlation and expressive power of high-dimensional features step by step, enabling the network to more accurately perceive and enhance the interaction between posterior probabilities and conditional inputs. This process is crucial for improving the visual quality and stability of the restoration results. Extensive experiments demonstrate that even when facial images suffer from both pixel-level and image-level degradation, UniFRD can still guarantee the restoration of rich details and maintain attribute consistency. In summary, compared to existing methods, the solution proposed in this study for facial restoration tasks offers greater generality and adaptability. Moreover, it has high practical value for applications involving faces in complex and unconstrained outdoor scenarios. Muwei Jian, Rui Wang 0199, Feng Xu 0005, Hui Yu 0001, Kin-Man Lam 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Multi-Grained Radiology Report Generation With Sentence-Level Image-Language Contrastive LearningabstractThe automatic generation of accurate radiology reports is of great clinical importance and has drawn growing research interest. However, it is still a challenging task due to the imbalance between normal and abnormal descriptions and the multi-sentence and multi-topic nature of radiology reports. These features result in significant challenges to generating accurate descriptions for medical images, especially the important abnormal findings. Previous methods to tackle these problems rely heavily on extra manual annotations, which are expensive to acquire. We propose a multi-grained report generation framework incorporating sentence-level image-sentence contrastive learning, which does not require any extra labeling but effectively learns knowledge from the image-report pairs. We first introduce contrastive learning as an auxiliary task for image feature learning. Different from previous contrastive methods, we exploit the multi-topic nature of imaging reports and perform fine-grained contrastive learning by extracting sentence topics and contents and contrasting between sentence contents and refined image contents guided by sentence topics. This forces the model to learn distinct abnormal image features for each specific topic. During generation, we use two decoders to first generate coarse sentence topics and then the fine-grained text of each sentence. We directly supervise the intermediate topics using sentence topics learned by our contrastive objective. This strengthens the generation constraint and enables independent fine-tuning of the decoders using reinforcement learning, which further boosts model performance. Experiments on two large-scale datasets MIMIC-CXR and IU-Xray demonstrate that our approach outperforms existing state-of-the-art methods, evaluated by both language generation metrics and clinical accuracy. Aohan Liu, Jun-Hai Yong, Feng Xu 0005 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | An Optimized-Skeleton-Based Parkinsonian Gait Auxiliary Diagnosis Method with Both Monitoring Indicators and Assisted RatingsabstractAbnormal gait is one of the indispensable diagnostic sources of Parkinson’s disease (PD) diagnosis, typically presenting as small shuffling steps and gait bradykinesia. However, its diagnostic accuracy is lower due to the subjective judgments of doctors. To assist in improving the accuracy and reducing the subjectivity of the doctors, we propose an optimized-skeleton-based Parkinsonian gait auxiliary diagnosis method with both monitoring indicators and assisted ratings. By inputting a patient gait video captured from the side, our PD symptom-applicable pose trajectory model will extract a more precise and stable 2D skeleton sequence of patients. Next, the sequence will be used to calculate our proposed five monitoring indicators: gait frequency, ankle speed, whole speed, ankle angle speed, ankle acceleration, and previous work indicators: arm swing angle, leg angle, two feet x-axis distance to record the patient’s gait details at every moment. The extracted gait frequency will then be input into a random forest model to obtain the gait rating. Lastly, doctors can make more accurate judgments by referring to our objective monitoring indicators and assisted ratings. Experimental results show that our monitored indicators improve the doctors’ diagnosis accuracy by 16%, the skeleton speed and acceleration error of our optimized-skeleton extraction method achieve 4.18 cm/s and 5.71 cm/s2, and our random forest model has reached a classification accuracy of 95.8%. Gaoqi Li, Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Feng Xu 0005, Hao Gao 0005 |
BIBM | 5 |
| 2023 | Multi-modal Contrastive-Generative Pre-training for Fine-grained Skin Disease DiagnosisabstractVision-language pre-training (VLP) leverages easily accessible image-text pairs instead of high-cost expert-annotated labels for pre-training, which has achieved promising performance and attracted considerable attention. There are many works on coarse-grained VLP in natural and medical radiology applications. However, as a common problem, the fine-grained setting is still unexplored, especially in medical applications like skin disease diagnosis. In fine-grained cases, the visual appearance of different objects is highly similar, and the language information may be sparse and noisy, both remarkably increasing the difficulty of learning effective features by VLP. In this paper, we address these difficulties and propose a novel Multilevel Multi-modal Contrastive-Generative (M2CG) pre-training method. M2CG has 1) a feature-level multi-modal contrastive module to learn fine-grained features via semantic knowledge, and 2) a semantic-level cross-modal generation module to enforce the model to capture key and discriminative features. We construct a multi-modal skin disease dataset, containing user-taken lesion photos, chief complaints, and consultation dialogues, to perform VLP with M2CG and evaluate the performance on three public skin disease benchmarks and a fine-grained dataset with 64 categories collected from real-world applications. M2CG outperforms the state-of-the-art VLP methods by up to 11.11% in diagnosis accuracy, yielding consistent and significant promotion and facilitating skin disease diagnosis. To our knowledge, this is the first VLP study presented for fine-grained skin disease diagnosis. We believe that the success of M2CG will inspire more innovations in fine-grained VLP for medical practice. Liangdi Ma, Feng Xu 0005 |
BIBM | 5 |
| 2023 | Learning a 3D Morphable Face Reflectance Model from Low-Cost DataabstractModeling non-Lambertian effects such as facial specularity leads to a more realistic 3D Morphable Face Model. Existing works build parametric models for diffuse and specular albedo using Light Stage data. However, only diffuse and specular albedo cannot determine the full BRDF. In addition, the requirement of Light Stage data is hard to fulfill for the research communities. This paper proposes the first 3D morphable face reflectance model with spatially varying BRDF using only low-cost publicly-available data. We apply linear shiness weighting into parametric modeling to represent spatially varying specular intensity and shiness. Then an inverse rendering algorithm is developed to reconstruct the reflectance parameters from non-Light Stage data, which are used to train an initial morphable reflectance model. To enhance the model's generalization capability and expressive power, we further propose an update-by-reconstruction strategy to finetune it on an in-the-wild dataset. Experimental results show that our method obtains decent rendering results with plausible facial specularities. Our code is released here. Zhibo Wang 0003, Feng Xu 0005 |
CVPR | 3 |
| 2023 | ShadowNeuS: Neural SDF Reconstruction by Shadow Ray SupervisionabstractBy supervising camera rays between a scene and multi-view image planes, NeRF reconstructs a neural scene representation for the task of novel view synthesis. On the other hand, shadow rays between the light source and the scene have yet to be considered. Therefore, we propose a novel shadow ray supervision scheme that optimizes both the samples along the ray and the ray location. By supervising shadow rays, we successfully reconstruct a neural SDF of the scene from single-view images under multiple lighting conditions. Given single-view binary shadows, we train a neural network to reconstruct a complete scene not limited by the camera's line of sight. By further modeling the correlation between the image colors and the shadow rays, our technique can also be effectively extended to RGB inputs. We compare our method with previous works on challenging tasks of shape reconstruction from single-view binary shadow or RGB images and observe significant improvements. The code and data are available at https://github.com/gerwang/ShadowNeuS. Jingwang Ling, Zhibo Wang 0003, Feng Xu 0005 |
CVPR | 3 |
| 2023 | EditableNeRF: Editing Topologically Varying Neural Radiance Fields by Key PointsabstractNeural radiance fields (NeRF) achieve highly photorealistic novel-view synthesis, but it's a challenging problem to edit the scenes modeled by NeRF-based methods, especially for dynamic scenes. We propose editable neural radiance fields that enable end-users to easily edit dynamic scenes and even support topological changes. Input with an image sequence from a single camera, our network is trained fully automatically and models topologically varying dynamics using our picked-out surface key points. Then end-users can edit the scene by easily dragging the key points to desired new positions. To achieve this, we propose a scene analysis method to detect and initialize key points by considering the dynamics in the scene, and a weighted key points strategy to model topologically varying dynamics by joint key points and weights optimization. Our method supports intuitive multi-dimensional (up to 3D) editing and can generate novel scenes that are unseen in the input sequence. Experiments demonstrate that our method achieves high-quality editing on various dynamic scenes and outperforms the state-of-the-art. Our code and captured data are available at https://chengwei-zheng.github.io/EditableNeRF/ Chengwei Zheng, Feng Xu 0005 |
CVPR | 3 |
| 2023 | Towards Eyeglasses Refraction in Appearance-based Gaze EstimationabstractFor myopia and hyperopia subjects, eyeglasses would change the position of objects in their views, leading to different eyeball rotations for the same gaze target (Fig. 1). Existing appearance-based gaze estimation methods ignore this effect, while this paper investigates it and proposes an effective method to consider it in gaze estimation, achieving noticeable improvements. Specifically, we discover that the appearance-gaze mapping differs for spectacled and unspectacled conditions, and the deviations are nearly consistent with the physical laws of the ideal lens. Based on this discovery, we propose a novel multi-task training strategy that encourages networks to regress gaze and classify the wearing conditions simultaneously. We apply the proposed strategy to some popular methods, including supervised and unsupervised ones, and evaluate them on different datasets with various backbones. The results show that the multi-task training strategy could be used on the existing methods to improve the performance of gaze estimation. To the best of our knowledge, we are the first to clearly reveal and explicitly consider eyeglasses refraction in appearance-based gaze estimation. Data and code are available at https://github.com/StoryMY/RefractionGaze. Junfeng Lyu, Feng Xu 0005 |
ISMAR | 2 |
| 2023 | Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion CaptureabstractEither RGB images or inertial signals have been used for the task of motion capture (mocap), but combining them together is a new and interesting topic. We believe that the combination is complementary and able to solve the inherent difficulties of using one modality input, including occlusions, extreme lighting/texture, and out-of-view for visual mocap and global drifts for inertial mocap. To this end, we propose a method that fuses monocular images and sparse IMUs for real-time human motion capture. Our method contains a dual coordinate strategy to fully explore the IMU signals with different goals in motion capture. To be specific, besides one branch transforming the IMU signals to the camera coordinate system to combine with the image information, there is another branch to learn from the IMU signals in the body root coordinate system to better estimate body poses. Furthermore, a hidden state feedback mechanism is proposed for both two branches to compensate for their own drawbacks in extreme input cases. Thus our method can easily switch between the two kinds of signals or combine them in different cases to achieve a robust mocap. Quantitative and qualitative results demonstrate that by delicately designing the fusion method, our technique significantly outperforms the state-of-the-art vision, IMU, and combined methods on both global orientation and local pose estimation. Our codes are available for research at https://shaohua-pan.github.io/robustcap-page/. Shaohua Pan 0002, Xinyu Yi, Xingkang Zhou, Jijunnan Li, Feng Xu 0005 |
SIGGRAPH Asia | 8 |
| 2023 | EgoLocate: Real-time Motion Capture, Localization, and Mapping with Sparse Body-mounted SensorsabstractHuman and environment sensing are two important topics in Computer Vision and Graphics. Human motion is often captured by inertial sensors, while the environment is mostly reconstructed using cameras. We integrate the two techniques together in EgoLocate, a system that simultaneously performs human motion capture (mocap), localization, and mapping in real time from sparse body-mounted sensors, including 6 inertial measurement units (IMUs) and a monocular phone camera. On one hand, inertial mocap suffers from large translation drift due to the lack of the global positioning signal. EgoLo-cate leverages image-based simultaneous localization and mapping (SLAM) techniquesto locate the human in the reconstructed scene. Onthe other hand, SLAM often fails when the visual feature is poor. EgoLocate involves inertial mocap to provide a strong prior for the camera motion. Experiments show that localization, a key challenge for both two fields, is largely improved by our technique, compared with the state of the art of the two fields. Our codes are available for research at https://xinyu-yi.github.io/EgoLocate/. Xinyu Yi, Yuxiao Zhou 0001, Marc Habermann, Vladislav Golyanik, Shaohua Pan 0002, Christian Theobalt, Feng Xu 0005 |
ACM Trans. Graph. | 7 |
| 2023 | Semantically Disentangled Variational Autoencoder for Modeling 3D Facial DetailsabstractParametric face models, such as morphable and blendshape models, have shown great potential in face representation, reconstruction, and animation. However, all these models focus on large-scale facial geometry. Facial details such as wrinkles are not parameterized in these models, impeding accuracy and realism. In this article, we propose a method to learn a Semantically Disentangled Variational Autoencoder (SDVAE) to parameterize facial details and support independent detail manipulation as an extension of an off-the-shelf large-scale face model. Our method utilizes the non-linear capability of Deep Neural Networks for detail modeling, achieving better accuracy and greater representation power compared with linear models. In order to disentangle the semantic factors of identity, expression and age, we propose to eliminate the correlation between different factors in an adversarial manner. Therefore, wrinkle-level details of various identities, expressions, and ages can be generated and independently controlled by changing latent vectors of our SDVAE. We further leverage our model to reconstruct 3D faces via fitting to facial scans and images. Benefiting from our parametric model, we achieve accurate and robust reconstruction, and the reconstructed details can be easily animated and manipulated. We evaluate our method on practical applications, including scan fitting, image fitting, video tracking, model manipulation, and expression and age animation. Extensive experiments demonstrate that the proposed method can robustly model facial details and achieve better results than alternative methods. Jingwang Ling, Zhibo Wang 0003, Ming Lu 0002, Chen Qian 0006, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | A Self-Occlusion Aware Lighting Model for Real-Time Dynamic ReconstructionabstractIn real-time dynamic reconstruction, geometry and motion are the major focuses while appearance is not fully explored, leading to the low-quality appearance of the reconstructed surfaces. In this article, we propose a lightweight lighting model that considers spatially varying lighting conditions caused by self-occlusion. This model estimates per-vertex masks on top of a single Spherical Harmonic (SH) lighting to represent spatially varying lighting conditions without adding too much computation cost. The mask is estimated based on the local geometry of a vertex to model the self-occlusion effect, which is the major reason leading to the spatial variation of lighting. Furthermore, to use this model in dynamic reconstruction, we also improve the motion estimation quality by adding a real-time per-vertex displacement estimation step. Experiments demonstrate that both the reconstructed appearance and the motion are largely improved compared with the current state-of-the-art techniques. Chengwei Zheng, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | OcclusionFusion: Occlusion-aware Motion Estimation for Real-time Dynamic 3D ReconstructionabstractRGBD-based real-time dynamic 3D reconstruction suffers from inaccurate inter-frame motion estimation as errors may accumulate with online tracking. This problem is even more severe for single-view-based systems due to strong occlusions. Based on these observations, we propose OcclusionFusion, a novel method to calculate occlusion-aware 3D motion to guide the reconstruction. In our technique, the motion of visible regions is first estimated and combined with temporal information to infer the motion of the occluded regions through an LSTM-involved graph neural network. Furthermore, our method computes the confidence of the estimated motion by modeling the network output with a probabilistic model, which alleviates untrust-worthy motions and enables robust tracking. Experimental results on public datasets and our own recorded data show that our technique outperforms existing single-view-based real-time methods by a large margin. With the reduction of the motion errors, the proposed technique can handle long and challenging motion sequences. Please check out the project page for sequence results: https://wenbinlin.github.io/OcclusionFusion. Chengwei Zheng, Jun-Hai Yong, Feng Xu 0005 |
CVPR | 4 |
| 2022 | Portrait Eyeglasses and Shadow Removal by Leveraging 3D Synthetic DataabstractIn portraits, eyeglasses may occlude facial regions and generate cast shadows on faces, which degrades the performance of many techniques like face verification and expression recognition. Portrait eyeglasses removal is critical in handling these problems. However, completely removing the eyeglasses is challenging because the lighting effects (e.g., cast shadows) caused by them are often complex. In this paper, we propose a novel framework to remove eyeglasses as well as their cast shadows from face images. The method works in a detect-then-remove manner, in which eyeglasses and cast shadows are both detected and then removed from images. Due to the lack of paired data for supervised training, we present a new synthetic portrait dataset with both intermediate and final supervisions for both the detection and removal tasks. Furthermore, we apply a cross-domain technique to fill the gap between the synthetic and real data. To the best of our knowledge, the proposed technique is the first to remove eyeglasses and their cast shadows simultaneously. The code and synthetic dataset are available at https://gethub.com/StoryMY/take-off-eyeglasses. Junfeng Lyu, Zhibo Wang 0003, Feng Xu 0005 |
CVPR | 3 |
| 2022 | Physical Inertial Poser (PIP): Physics-aware Real-time Human Motion Tracking from Sparse Inertial SensorsabstractMotion capture from sparse inertial sensors has shown great potential compared to image-based approaches since occlusions do not lead to a reduced tracking quality and the recording space is not restricted to be within the viewing frustum of the camera. However, capturing the motion and global position only from a sparse set of inertial sensors is inherently ambiguous and challenging. In consequence, recent state-of-the-art methods can barely handle very long period motions, and unrealistic artifacts are common due to the unawareness of physical constraints. To this end, we present the first method which combines a neural kinematics estimator and a physics-aware motion optimizer to track body motions with only 6 inertial sensors. The kinematics module first regresses the motion status as a reference, and then the physics module refines the motion to satisfy the physical constraints. Experiments demonstrate a clear improvement over the state of the art in terms of capture accuracy, temporal stability, and physical correctness. Xinyu Yi, Yuxiao Zhou 0001, Marc Habermann, Soshi Shimada, Vladislav Golyanik, Christian Theobalt, Feng Xu 0005 |
CVPR | 7 |
| 2022 | Structure-Aware Editable Morphable Model for 3D Facial Detail Animation and Manipulation
Jingwang Ling, Zhibo Wang 0003, Ming Lu 0002, Chen Qian 0006, Feng Xu 0005 |
ECCV (3) | 6 |
| 2022 | Physical Interaction: Reconstructing Hand-object Interactions with PhysicsabstractSingle view-based reconstruction of hand-object interaction is challenging due to the severe observation missing caused by occlusions. This paper proposes a physics-based method to better solve the ambiguities in the reconstruction. It first proposes a force-based dynamic model of the in-hand object, which not only recovers the unobserved contacts but also solves for plausible contact forces. Next, a confidence-based slide prevention scheme is proposed, which combines both the kinematic confidences and the contact forces to jointly model static and sliding contact motion. Qualitative and quantitative experiments show that the proposed technique reconstructs both physically plausible and more accurate hand-object interaction and estimates plausible contact forces in real-time with a single RGBD sensor. Xinyu Yi, Hao Zhang 0042, Jun-Hai Yong, Feng Xu 0005 |
SIGGRAPH Asia | 5 |
| 2022 | Action Recognition Framework in Traffic Scene for Autonomous Driving SystemabstractFor the autonomous driving system, accurately recognizing the actions of different roles in the traffic scene is the prerequisite for realizing this kind of human-vehicle information interaction. In this paper, we propose a complete framework based on 3D human pose estimation to recognize the actions of different roles on the road. The main objects recognized include traffic police, cyclists, and some passersby in need. We perform action recognition based on a dynamic adaptive graph convolutional network, which can realize the action recognition of objects based on 3D human pose. In addition to the action recognition module, we have optimized both the object detection module and the human pose estimation module in the framework so that the framework can handle multiple objects at the same time, which can be closer to the real traffic scene. To realize complex and changeable human action recognition, we built a multi-view camera system to collect responsible 3D human pose datasets containing traffic police gestures, cyclist gestures, and pedestrians’ body movements. In the experiments, compared to other state-of-the-art researches, the proposed framework can achieve comparable results with the same dataset. Satisfactory performance has also been obtained on the real data we collected, which can handle a variety of different action recognition tasks at the same time. Feiyi Xu, Feng Xu 0005, Jiucheng Xie, Chi-Man Pun, Huimin Lu 0001, Hao Gao 0005 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Virtual Reality Aided High-Quality 3D Reconstruction by Remote DronesabstractArtificial intelligence including deep learning and 3D reconstruction methods is changing the daily life of people. Now, an unmanned aerial vehicle that can move freely in the air and avoid harsh ground conditions has been commonly adopted as a suitable tool for 3D reconstruction. The traditional 3D reconstruction mission based on drones usually consists of two steps: image collection and offline post-processing. But there are two problems: one is the uncertainty of whether all parts of the target object are covered, and another is the tedious post-processing time. Inspired by modern deep learning methods, we build a telexistence drone system with an onboard deep learning computation module and a wireless data transmission module that perform incremental real-time dense reconstruction of urban cities by itself. Two technical contributions are proposed to solve the preceding issues. First, based on the popular depth fusion surface reconstruction framework, we combine it with a visual-inertial odometry estimator that integrates the inertial measurement unit and allows for robust camera tracking as well as high-accuracy online 3D scan. Second, the capability of real-time 3D reconstruction enables a new rendering technique that can visualize the reconstructed geometry of the target as navigation guidance in the HMD. Therefore, it turns the traditional path-planning-based modeling process into an interactive one, leading to a higher level of scan completeness. The experiments in the simulation system and our real prototype demonstrate an improved quality of the 3D model using our artificial intelligence leveraged drone system. Feng Xu 0005, Chi-Man Pun, Yang Yang 0002, Rushi Lan, Yujie Li 0001, Hao Gao 0005 |
ACM Trans. Internet Techn. | 2 |
| 2022 | PlaneFusion: Real-Time Indoor Scene Reconstruction With Planar PriorabstractReal-time dense SLAM techniques aim to reconstruct the dense three-dimensional geometry of a scene in real time with an RGB or RGB-D sensor. An indoor scene is an important type of working environment for these techniques. The planar prior can be used in this scenario to improve the reconstruction quality, especially for large low-texture regions that commonly occur in an indoor scene. This article fully explores the planar prior in a dense SLAM pipeline. First, we propose a novel plane detection and segmentation method that runs at 200 Hz on a modern graphics processing unit. Our algorithm for constructing global plane constraints is very efficient; hence, we use it in the process of each input frame for the camera pose estimation while maintaining the real-time performance. Second, we propose herein a plane-based map representation that greatly reduces the memory footprint of plane regions while keeping the geometric details on planes. The experiments reveal that our system yields superior reconstruction results with planar information running at more than 30 fps. Aside from speed and storage improvements, our technique also handles the low-texture problem in plane regions. Bingjian Gong, Zunjie Zhu, Chenggang Yan 0001, Zhiguo Shi 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Emotion-Preserving Blendshape Update With Real-Time Face TrackingabstractBlendshape representations are widely used in facial animation. Consistent semantics must be maintained for all the blendshapes to build the blendshapes of one character. However, this is difficult for real characters because the face shape of the same semantics varies significantly across identities. Previous studies have handled this issue by asking users to perform a set of predefined expressions with specified semantics. We observe that facial emotions can be used to define semantics. Herein, we propose a real-time technique that directly updates blendshapes without predefined expressions. Its aim is to preserve semantics based on the emotion information extracted from an arbitrary facial motion sequence. In addition, we have designed corresponding algorithms to efficiently update blendshapes with large- and middle-scale face shapes and fine-scale facial details, such as wrinkles, in a real-time face tracking system. The experimental results indicate that using a commodity RGBD sensor, we can achieve real-time online blendshape updates with well-preserved semantics and user-specific facial features and details. Zhibo Wang 0003, Jingwang Ling, Chengzeng Feng, Ming Lu 0002, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | DTexFusion: Dynamic Texture Fusion Using a Consumer RGBD SensorabstractIn addition to 3D geometry, accurate representation of texture is important when digitizing real objects in virtual worlds. Based on a single consumer RGBD sensor, accurate texture representation for static objects can be realized by fusing multi-frame information; however, extending the process to dynamic objects, which typically have time-varying textures, is difficult. Thus, to address this problem, we propose a compact keyframe-based representation that decouples a dynamic texture into a basic static texture and a set of multiplicative changing maps. With this representation, the proposed method first aligns textures recorded from multiple keyframes with the reconstructed dynamic geometry of the object. Errors in the alignment and geometry are then compensated in an innovative iterative linear optimization framework. With the reconstructed texture, we then employ a scheme to synthesize the dynamic object from arbitrary viewpoints. By considering temporal and local pose similarities jointly, dynamic textures in all keyframes are fused to guarantee high-quality image generation. Experimental results demonstrate that the proposed method handles various dynamic objects, including faces, bodies, cloth, and toys. In addition, qualitative and quantitative comparisons demonstrate that the proposed method outperforms state-of-the-art solutions. Chengwei Zheng, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Audio-Driven Emotional Video PortraitsabstractDespite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important features on natural human faces, is always neglected in their methods. In this work, we present Emotional Video Portraits (EVP), a system for synthesizing high-quality video portraits with vivid emotional dynamics driven by audios. Specifically, we propose the Cross-Reconstructed Emotion Disentanglement technique to decompose speech into two decoupled spaces, i.e., a duration-independent emotion space and a duration- dependent content space. With the disentangled features, dynamic 2D emotional facial landmarks can be deduced. Then we propose the Target-Adaptive Face Synthesis technique to generate the final high-quality video portraits, by bridging the gap between the deduced landmarks and the natural head poses of target videos. Extensive experiments demonstrate the effectiveness of our method both qualitatively and quantitatively.1 Xinya Ji, Hang Zhou 0009, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, Feng Xu 0005 |
CVPR | 7 |
| 2021 | Monocular Real-Time Full Body Capture With Inter-Part CorrelationsabstractWe present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network architecture that exploits correlations between body and hands at high computational efficiency. Unlike previous works, our approach is jointly trained on multiple datasets focusing on hand, body or face separately, without requiring data where all the parts are annotated at the same time, which is much more difficult to create at sufficient variety. The possibility of such multi-dataset training enables superior generalization ability. In contrast to earlier monocular full body methods, our approach captures more expressive 3D face geometry and color by estimating the shape, expression, albedo and illumination parameters of a statistical face model. Our method achieves competitive accuracy on public benchmarks, while being significantly faster and providing more complete face reconstructions. Yuxiao Zhou 0001, Marc Habermann, Ikhsanul Habibie, Ayush Tewari, Christian Theobalt, Feng Xu 0005 |
CVPR | 6 |
| 2021 | A Rate-based Drone Control with Adaptive Origin Update in TelexistenceabstractA new form of telexistence is achieved by recording videos with a camera on an Uncrewed aerial vehicle (UAV) and playing the videos to a user via a head-mounted display (HMD). One key problem here is how to let the user freely and naturally control the UAV and thus the viewpoint. In this paper, we develop an HMD-based telexistence technique that achieves full 6- DOF control of the viewpoint. The core of our technique is an improved rate-based control technique with our adaptive origin update (AOU), in which the origin of the coordinate system of the user changes adaptively. This makes the user naturally perceive the origin and thus easily perform the control motion to get his/her desired viewpoint changing. As a consequence, without the aid of any auxiliary equipment, the AOU scheme handles the well known self-centering problem in the rate-based control methods. A real prototype is also built to evaluate this feature of our technique. To explore the advantage of our telexistence technique, we further use it as an interactive tool to perform the task of 3D scene reconstruction. User studies demonstrate that comparing with other telexistence solutions and the widely used joystick-based solutions, our solution largely reduces the workload and saves time and moving distance for the user. Chi-Man Pun, Yang Yang 0002, Hao Gao 0005, Feng Xu 0005 |
VR | 5 |
| 2021 | Enhancement of ridge-valley features in point cloud based on position and normal guidance
Jianhui Nie, Zhaochen Zhang, Ye Liu 0005, Hao Gao 0005, Feng Xu 0005, Wenkai Shi |
Comput. Graph. | 5 |
| 2021 | Bas-relief generation from point clouds based on normal space compression with real-time adjustment on CPU
Jianhui Nie, Wenkai Shi, Ye Liu 0005, Hao Gao 0005, Feng Xu 0005, Zhaochen Zhang |
Graph. Model. | 5 |
| 2021 | TransPose: real-time 3D human translation and pose estimation with six inertial sensorsabstractMotion capture is facing some new possibilities brought by the inertial sensing technologies which do not suffer from occlusion or wide-range recordings as vision-based solutions do. However, as the recorded signals are sparse and quite noisy, online performance and global translation estimation turn out to be two key difficulties. In this paper, we present TransPose, a DNN-based approach to perform full motion capture (with both global translations and body poses) from only 6 Inertial Measurement Units (IMUs) at over 90 fps. For body pose estimation, we propose a multi-stage network that estimates leaf-to-full joint positions as intermediate results. This design makes the pose estimation much easier, and thus achieves both better accuracy and lower computation cost. For global translation estimation, we propose a supporting-foot-based method and an RNN-based method to robustly solve for the global translations with a confidence-based fusion technique. Quantitative and qualitative comparisons show that our method outperforms the state-of-the-art learning- and optimization-based methods with a large margin in both accuracy and efficiency. As a purely inertial sensor-based approach, our method is not limited by environmental settings (e.g., fixed cameras), making the capture free from common difficulties such as wide-range motion space and strong occlusion. Xinyu Yi, Yuxiao Zhou 0001, Feng Xu 0005 |
ACM Trans. Graph. | 3 |
| 2021 | Single Depth View Based Real-Time Reconstruction of Hand-Object InteractionsabstractReconstructing hand-object interactions is a challenging task due to strong occlusions and complex motions. This article proposes a real-time system that uses a single depth stream to simultaneously reconstruct hand poses, object shape, and rigid/non-rigid motions. To achieve this, we first train a joint learning network to segment the hand and object in a depth image, and to predict the 3D keypoints of the hand. With most layers shared by the two tasks, computation cost is saved for the real-time performance. A hybrid dataset is constructed here to train the network with real data (to learn real-world distributions) and synthetic data (to cover variations of objects, motions, and viewpoints). Next, the depth of the two targets and the keypoints are used in a uniform optimization to reconstruct the interacting motions. Benefitting from a novel tangential contact constraint, the system not only solves the remaining ambiguities but also keeps the real-time performance. Experiments show that our system handles different hand and object shapes, various interactive motions, and moving cameras. Hao Zhang 0042, Yuxiao Zhou 0001, Yifei Tian, Jun-Hai Yong, Feng Xu 0005 |
ACM Trans. Graph. | 5 |
| 2021 | A Hybrid Feature Selection Algorithm Based on a Discrete Artificial Bee Colony for Parkinson's DiagnosisabstractParkinson's disease is a neurodegenerative disease that affects millions of people around the world and cannot be cured fundamentally. Automatic identification of early Parkinson's disease on feature data sets is one of the most challenging medical tasks today. Many features in these datasets are useless or suffering from problems like noise, which affect the learning process and increase the computational burden. To ensure the optimal classification performance, this article proposes a hybrid feature selection algorithm based on an improved discrete artificial bee colony algorithm to improve the efficiency of feature selection. The algorithm combines the advantages of filters and wrappers to eliminate most of the uncorrelated or noisy features and determine the optimal subset of features. In the filter, three different variable ranking methods are employed to pre-rank the candidate features, then the population of artificial bee colony is initialized based on the significance degree of the re-rank features. In the wrapper part, the artificial bee colony algorithm evaluates individuals (feature subsets) based on the classification accuracy of the classifier to achieve the optimal feature subset. In addition, for the first time, we introduce a strategy that can automatically select the best classifier in the search framework more quickly. By comparing with several publicly available datasets, the proposed method achieves better performance than other state-of-the-art algorithms and can extract fewer effective features. Haolun Li 0001, Chi-Man Pun, Feng Xu 0005, Longsheng Pan, Rui Zong, Hao Gao 0005, Huimin Lu 0001 |
ACM Trans. Internet Techn. | 3 |
| 2021 | Data-Driven 3D Neck Modeling and AnimationabstractIn this article, we present a data-driven approach for modeling and animation of 3D necks. Our method is based on a new neck animation model that decomposes the neck animation into local deformation caused by larynx motion and global deformation driven by head poses, facial expressions, and speech. A skinning model is introduced for modeling local deformation and underlying larynx motions, while the global neck deformation caused by each factor is modeled by its corrective blendshape set, respectively. Based on this neck model, we introduce a regression method to drive the larynx motion and neck deformation from speech. Both the neck model and the speech regressor are learned from a dataset of 3D neck animation sequences captured from different identities. Our neck model significantly improves the realism of facial animation and allows users to easily create plausible neck animations from speech and facial expressions. We verify our neck model and demonstrate its advantages in 3D neck tracking and animation. Chengwei Zheng, Feng Xu 0005, Xin Tong 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Monocular Real-Time Hand Shape and Motion Capture Using Multi-Modal DataabstractWe present a novel method for monocular hand shape and pose estimation at unprecedented runtime performance of 100fps and at state-of-the-art accuracy. This is enabled by a new learning based architecture designed such that it can make use of all the sources of available hand training data: image data with either 2D or 3D annotations, as well as stand-alone 3D animations without corresponding image data. It features a 3D hand joint detection module and an inverse kinematics module which regresses not only 3D joint positions but also maps them to joint rotations in a single feed-forward pass. This output makes the method more directly usable for applications in computer vision and graphics compared to only regressing 3D joint positions. We demonstrate that our architectural design leads to a significant quantitative and qualitative improvement over the state of the art on several challenging benchmarks. We will make our code publicly available for future research. Yuxiao Zhou 0001, Marc Habermann, Weipeng Xu, Ikhsanul Habibie, Christian Theobalt, Feng Xu 0005 |
CVPR | 6 |
| 2020 | Eyeglasses 3D Shape Reconstruction from a Single Face Image
Feng Xu 0005 |
ECCV (25) | 3 |
| 2020 | New multi-view human motion capture frameworkabstractEstimating human pose and shape without markers is a challenging problem. This study proposes a multiple‐view markerless human motion capture framework. Firstly, a multi‐view camera system is built for capturing real‐time images of moving humans on multiple views. Secondly, by employing the OpenPose method, the authors calculate robust 3D key points from 2D key points of the human body, which are estimated from the multi‐view images. And dense 3D point cloud is reconstructed from images. Thirdly, they propose a novel SMPL‐based method to represent human motion by fitting the SMPL model to 3D key points and 3D point clouds. In order to achieve a more accurate human pose, a penalty term is utilised to solve the problem of error accumulation in the process of human motion capture. In addition, they present a dense mesh template‐based SMPL that can be deformed to point cloud to recover a real human body shape. Finally, they map multi‐view colour images onto the human mesh model to acquire rendered mesh. The experimental results show that the proposed method improves the accuracy of human pose and realises the 3D human body model more realistic. Feiyi Xu, Chi-Man Pun, Wenqi Xiao, Jianhui Nie, Jian Xiong 0005, Hao Gao 0005, Feng Xu 0005 |
IET Image Process. | 8 |
| 2020 | 3D Room Layout Estimation From a Single RGB Imageabstract3D layout is crucial for scene understanding and reconstruction, and very useful in applications like real estate and furniture design. In this paper, we propose a fully automatic solution to estimate 3D layout of an indoor scene from a single 2D image. Our technique contains two key components. Firstly, we train a neural network that directly estimates room structure lines from the input image. Secondly, we propose a novel technique to automatically identify the layout topology of an input image, followed by a nonlinear optimization with equality constraints to estimate the final 3D layout of a scene. Based on our knowledge, this is the first fully automatic technique to achieve single image-based 3D layout estimation of an indoor scene. We evaluate our method on the public datasets LSUN, Hedau and 3DGP and the results show that the proposed method achieves accurate 3D layout reconstruction on various images with different layout topologies. Chenggang Yan 0001, Biyao Shao, Hao Zhao 0002, Ruixin Ning, Yongdong Zhang 0001, Feng Xu 0005 |
IEEE Trans. Multim. | 6 |
| 2020 | Single image portrait relighting via explicit multiple reflectance channel modelingabstractPortrait relighting aims to render a face image under different lighting conditions. Existing methods do not explicitly consider some challenging lighting effects such as specular and shadow, and thus may fail in handling extreme lighting conditions. In this paper, we propose a novel framework that explicitly models multiple reflectance channels for single image portrait relighting, including the facial albedo, geometry as well as two lighting effects, i.e. , specular and shadow. These channels are finally composed to generate the relit results via deep neural networks. Current datasets do not support learning such multiple reflectance channel modeling. Therefore, we present a large-scale dataset with the ground-truths of the channels, enabling us to train the deep neural networks in a supervised manner. Furthermore, we develop a novel module named Lighting guided Feature Modulation (LFM). In contrast to existing methods which simply incorporate the given lighting in the bottleneck of a network, LFM fuses the lighting by layer-wise feature modulation to deliver more convincing results. Extensive experiments demonstrate that our proposed method achieves better results and is able to generate challenging lighting effects. Zhibo Wang 0003, Xin Yu 0002, Ming Lu 0002, Chen Qian 0006, Feng Xu 0005 |
ACM Trans. Graph. | 6 |
| 2019 | A Closed-Form Solution to Universal Style TransferabstractUniversal style transfer tries to explicitly minimize the losses in feature space, thus it does not require training on any pre-defined styles. It usually uses different layers of VGG network as the encoders and trains several decoders to invert the features into images. Therefore, the effect of style transfer is achieved by feature transform. Although plenty of methods have been proposed, a theoretical analysis of feature transform is still missing. In this paper, we first propose a novel interpretation by treating it as the optimal transport problem. Then, we demonstrate the relations of our formulation with former works like Adaptive Instance Normalization (AdaIN) and Whitening and Coloring Transform (WCT). Finally, we derive a closed-form solution named Optimal Style Transfer (OST) under our formulation by additionally considering the content loss of Gatys. Comparatively, our solution can preserve better structure and achieve visually pleasing results. It is simple yet effective and we demonstrate its advantages both quantitatively and qualitatively. Besides, we hope our theoretical analysis can inspire future works in neural style transfer. Ming Lu 0002, Hao Zhao 0002, Anbang Yao, Yurong Chen 0001, Feng Xu 0005, Li Zhang 0023 |
ICCV | 5 |
| 2019 | Real-time Indoor Scene Reconstruction with RGBD and Inertial InputabstractCamera motion estimation is a key technique for 3D scene reconstruction. Previous works usually assume slow camera motions, which limit the usage in many real cases. We propose an end-to-end 3D reconstruction system which combines color, depth and inertial measurements to achieve robust reconstruction with fast sensor motions. Our framework utilizes extended Kalman filter to fuse the three kinds of information and involve an iterative method to jointly optimize feature correspondences, camera poses and scene geometry. We also propose a novel geometry-aware patch deformation technique to adapt the feature appearance in image domain, leading to a more accurate feature matching under fast camera motions. Experiments show that our patch deformation method improves the accuracy of feature tracking, and our 3D reconstruction framework outperforms the state-of-the-art solutions under fast camera motions. Zunjie Zhu, Feng Xu 0005, Chenggang Yan 0001, Xinhong Hao, Xiangyang Ji, Yongdong Zhang 0001, Qionghai Dai |
ICME | 2 |
| 2019 | A 6-DOF Telexistence Drone Controlled by a Head Mounted DisplayabstractRecently, a new form of telexistence is achieved by recording images with cameras on an unmanned aerial vehicle (UAV) and displaying them to the user via a head mounted display (HMD). A key problem here is how to provide a free and natural mechanism for the user to control the viewpoint and watch a scene. To this end, we propose an improved rate-control method with an adaptive origin update (AOU) scheme. Without the aid of any auxiliary equipment, our scheme handles the self-centering problem. In addition, we present a full 6-DOF viewpoint control method to manipulate the motion of a stereo camera, and we build a real prototype to realize this by utilizing a pan-tilt-zoom (PTZ) which not only provides 2-DOF to the camera but also compensates the jittering motion of the UAV to record more stable image streams. Xingyu Xia, Chi-Man Pun, Yang Yang 0002, Huimin Lu 0001, Hao Gao 0005, Feng Xu 0005 |
VR | 7 |
| 2019 | Image generation from bounding box-represented semantic labels
Congying Liu, Zexi Yang, Feng Xu 0005, Jun-Hai Yong |
Comput. Graph. | 3 |
| 2019 | Real-time indoor scene reconstruction with Manhattan assumption
Zunjie Zhu, Feng Xu 0005, Chenggang Yan 0001, Bingjian Gong, Yongdong Zhang 0001, Qionghai Dai |
Multim. Tools Appl. | 2 |
| 2019 | InteractionFusion: real-time reconstruction of hand poses and deformable objects in hand-object interactionsabstractHand-object interaction is challenging to reconstruct but important for many applications like HCI, robotics and so on. Previous works focus on either the hand or the object while we jointly track the hand poses, fuse the 3D object model and reconstruct its rigid and nonrigid motions, and perform all these tasks in real time. To achieve this, we first use a DNN to segment the hand and object in the two input depth streams and predict the current hand pose based on the previous poses by a pre-trained LSTM network. With this information, a unified optimization framework is proposed to jointly track the hand poses and object motions. The optimization integrates the segmented depth maps, the predicted motion, a spatial-temporal varying rigidity regularizer and a real-time contact constraint. A nonrigid fusion technique is further involved to reconstruct the object model. Experiments demonstrate that our method can solve the ambiguity caused by heavy occlusions between hand and object, and generate accurate results for various objects and interacting motions. Hao Zhang 0042, Zihao Bo, Jun-Hai Yong, Feng Xu 0005 |
ACM Trans. Graph. | 4 |
| 2018 | DDRNet: Depth Map Denoising and Refinement for Consumer Depth Cameras Using Cascaded CNNs
Shi Yan 0007, Chenglei Wu, Lizhen Wang 0002, Feng Xu 0005, Liang An 0001, Yebin Liu |
ECCV (10) | 4 |
| 2018 | Generating VR Live Videos with Tripod Panoramic RigabstractRecent breakthrough in consumer-level virtual reality (VR) devices brings an increasing demand of VR live content. As converting real life content into VR need complex computations, current techniques can not synthesize 360° 3D VR content with high performance, not to mention real time. We propose an end-to-end system that records a scene using a tripod panoramic rig and broadcasts 360° stereo panorama videos in real time. The system performs a panorama stitching technique which pre-compute 3 stitching seam candidates for dynamic seam switching in the live broadcasting. This technique achieves high frame rates (>30fps) with minimum foreground cutoff and temporal jittering artifacts. Stereo vision quality is also better preserved by a proposed weighting-based image alignment scheme. We demonstrate the effectiveness of our approach on a variety of videos delivering live events. And our system has been successfully used in broadcasting live shows to mobile phone users on a professional live broadcasting platform with about 390 million user visits per month. Feng Xu 0005, Bicheng Luo, Qionghai Dai |
VR | 1 |
| 2018 | Robust Non-Rigid Motion Tracking and Surface Reconstruction Using L0 RegularizationabstractWe present a new motion tracking technique to robustly reconstruct non-rigid geometries and motions from a single view depth input recorded by a consumer depth sensor. The idea is based on the observation that most non-rigid motions (especially human-related motions) are intrinsically involved in articulate motion subspace. To take this advantage, we propose a novel based motion regularizer with an iterative solver that implicitly constrains local deformations with articulate structures, leading to reduced solution space and physical plausible deformations. The strategy is integrated into the available non-rigid motion tracking pipeline, and gradually extracts articulate joints information online with the tracking, which corrects the tracking errors in the results. The information of the articulate joints is used in the following tracking procedure to further improve the tracking accuracy and prevent tracking failures. Extensive experiments over complex human body motions with occlusions, facial and hand motions demonstrate that our approach substantially improves the robustness and accuracy in motion tracking. Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | Errata to "Robust Non-Rigid Motion Tracking and Surface Reconstruction Using L0 Regularization"abstractPresents corrections to grant number information from the paper, “Robust non-rigid motion tracking and surface reconstruction using L0 regularization,” (Guo, K., et al), IEEE Trans. Vis. Comput. Graph., vol. 24, no. 5, pp. 1770–1783, May 2018. Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | Parallax360: Stereoscopic 360° Scene Representation for Head-Motion ParallaxabstractWe propose a novel 360° scene representation for converting real scenes into stereoscopic 3D virtual reality content with head-motion parallax. Our image-based scene representation enables efficient synthesis of novel views with six degrees-of-freedom (6-DoF) by fusing motion fields at two scales: (1) disparity motion fields carry implicit depth information and are robustly estimated from multiple laterally displaced auxiliary viewpoints, and (2) pairwise motion fields enable real-time flow-based blending, which improves the visual fidelity of results by minimizing ghosting and view transition artifacts. Based on our scene representation, we present an end-to-end system that captures real scenes with a robotic camera arm, processes the recorded data, and finally renders the scene in a head-mounted display in real time (more than 40 Hz). Our approach is the first to support head-motion parallax when viewing real 360° scenes. We demonstrate compelling results that illustrate the enhanced visual experience - and hence sense of immersion-achieved with our approach compared to widely-used stereoscopic panoramas. Bicheng Luo, Feng Xu 0005, Christian Richardt, Jun-Hai Yong |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | MixedFusion: Real-Time Reconstruction of an Indoor Scene with Dynamic ObjectsabstractReal-time indoor scene reconstruction aims to recover the 3D geometry of an indoor scene in real time with a sensor scanning the scene. Previous works of this topic consider pure static scenes, but in this paper, we focus on more challenging cases that the scene contains dynamic objects, for example, moving people and floating curtains, which are quite common in reality and thus are eagerly required to be handled. We develop an end-to-end system using a depth sensor to scan a scene on the fly. By proposing a Sigmoid-based Iterative Closest Point (S-ICP) method, we decouple the camera motion and the scene motion from the input sequence and segment the scene into static and dynamic parts accordingly. The static part is used to estimate the camera rigid motion, while for the dynamic part, graph node-based motion representation and model-to-depth fitting are applied to reconstruct the scene motions. With the camera and scene motions reconstructed, we further propose a novel mixed voxel allocation scheme to handle static and dynamic scene parts with different mechanisms, which helps to gradually fuse a large scene with both static and dynamic objects. Experiments show that our technique successfully fuses the geometry of both the static and dynamic objects in a scene in real time, which extends the usage of the current techniques for indoor scene reconstruction. Hao Zhang 0042, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Decoder Network over Lightweight Reconstructed Feature for Fast Semantic Style TransferabstractRecently, the community of style transfer is trying to incorporate semantic information into traditional system. This practice achieves better perceptual results by transferring the style between semantically-corresponding regions. Yet, few efforts are invested to address the computation bottleneck of back-propagation. In this paper, we propose a new framework for fast semantic style transfer. Our method decomposes the semantic style transfer problem into feature reconstruction part and feature decoder part. The reconstruction part tactfully solves the optimization problem of content loss and style loss in feature space by particularly reconstructed feature. This significantly reduces the computation of propagating the loss through the whole network. The decoder part transforms the reconstructed feature into the stylized image. Through a careful bridging of the two modules, the proposed approach not only achieves competitive results as backward optimization methods but also is about two orders of magnitude faster. Ming Lu 0002, Hao Zhao 0002, Anbang Yao, Feng Xu 0005, Yurong Chen 0001, Li Zhang 0023 |
ICCV | 4 |
| 2017 | BodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth CameraabstractWe propose BodyFusion, a novel real-time geometry fusion method that can track and reconstruct non-rigid surface motion of a human performance using a single consumer-grade depth camera. To reduce the ambiguities of the non-rigid deformation parameterization on the surface graph nodes, we take advantage of the internal articulated motion prior for human performance and contribute a skeleton-embedded surface fusion (SSF) method. The key feature of our method is that it jointly solves for both the skeleton and graph-node deformations based on information of the attachments between the skeleton and the graph nodes. The attachments are also updated frame by frame based on the fused surface geometry and the computed deformations. Overall, our method enables increasingly denoised, detailed, and complete surface reconstruction as well as the updating of the skeleton and attachments as the temporal depth frames are fused. Experimental results show that our method exhibits substantially improved nonrigid motion fusion performance and tracking robustness compared with previous state-of-the-art fusion methods. We also contribute a dataset for the quantitative evaluation of fusion-based dynamic scene reconstruction algorithms using a single depth camera. Tao Yu 0007, Feng Xu 0005, Zhaoqi Su, Jianhui Zhao 0002, Qionghai Dai, Yebin Liu |
ICCV | 3 |
| 2017 | Real-time 3D eyelids tracking from semantic edgesabstractState-of-the-art real-time face tracking systems still lack the ability to realistically portray subtle details of various aspects of the face, particularly the region surrounding the eyes. To improve this situation, we propose a technique to reconstruct the 3D shape and motion of eyelids in real time. By combining these results with the full facial expression and gaze direction, our system generates complete face tracking sequences with more detailed eye regions than existing solutions in real-time. To achieve this goal, we propose a generative eyelid model which decomposes eyelid variation into two low-dimensional linear spaces which efficiently represent the shape and motion of eyelids. Then, we modify a holistically-nested DNN model to jointly perform semantic eyelid edge detection and identification on images. Next, we correspond vertices of the eyelid model to 2D image edges, and employ polynomial curve fitting and a search scheme to handle incorrect and partial edge detections. Finally, we use the correspondences in a 3D-to-2D edge fitting scheme to reconstruct eyelid shape and pose. By integrating our fast fitting method into a face tracking system, the estimated eyelid results are seamlessly fused with the face and eyeball results in real time. Experiments show that our technique applies to different human races, eyelid shapes, and eyelid motions, and is robust to changes in head pose, expression and gaze direction. Quan Wen 0002, Feng Xu 0005, Ming Lu 0002, Jun-Hai Yong |
ACM Trans. Graph. | 2 |
| 2017 | Real-Time Geometry, Albedo, and Motion Reconstruction Using a Single RGB-D CameraabstractThis article proposes a real-time method that uses a single-view RGB-D input (a depth sensor integrated with a color camera) to simultaneously reconstruct a casual scene with a detailed geometry model, surface albedo, per-frame non-rigid motion, and per-frame low-frequency lighting, without requiring any template or motion priors. The key observation is that accurate scene motion can be used to integrate temporal information to recover the precise appearance, whereas the intrinsic appearance can help to establish true correspondence in the temporal domain to recover motion. Based on this observation, we first propose a shading-based scheme to leverage appearance information for motion estimation. Then, using the reconstructed motion, a volumetric albedo fusing scheme is proposed to complete and refine the intrinsic appearance of the scene by incorporating information from multiple frames. Since the two schemes are iteratively applied during recording, the reconstructed appearance and motion become increasingly more accurate. In addition to the reconstruction results, our experiments also show that additional applications can be achieved, such as relighting, albedo editing, and free-viewpoint rendering of a dynamic scene, since geometry, appearance, and motion are all reconstructed by our technique. Feng Xu 0005, Tao Yu 0007, Qionghai Dai, Yebin Liu |
ACM Trans. Graph. | 2 |
| 2017 | Real-Time 3D Eye Performance Reconstruction for RGBD CamerasabstractThis paper proposes a real-time method for 3D eye performance reconstruction using a single RGBD sensor. Combined with facial surface tracking, our method generates more pleasing facial performance with vivid eye motions. In our method, a novel scheme is proposed to estimate eyeball motions by minimizing the differences between a rendered eyeball and the recorded image. Our method considers and handles different appearances of human irises, lighting variations and highlights on images via the proposed eyeball model and the -based optimization. Robustness and real-time optimization are achieved through the novel 3D Taylor expansion-based linearization. Furthermore, we propose an online bidirectional regression method to handle occlusions and other tracking failures on either of the two eyes from the information of the opposite eye. Experiments demonstrate that our technique achieves robust and accurate eye performance reconstruction for different iris appearances, with various head/face/eye motions, and under different lighting conditions. Quan Wen 0002, Feng Xu 0005, Jun-Hai Yong |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Robust Non-rigid Motion Tracking and Surface Reconstruction Using L0 RegularizationabstractWe present a new motion tracking method to robustly reconstruct non-rigid geometries and motions from single view depth inputs captured by a consumer depth sensor. The idea comes from the observation of the existence of intrinsic articulated subspace in most of non-rigid motions. To take advantage of this characteristic, we propose a novel L0based motion regularizer with an iterative optimization solver that can implicitly constrain local deformation only on joints with articulated motions, leading to reduced solution space and physical plausible deformations. The L0strategy is integrated into the available non-rigid motion tracking pipeline, forming the proposed L0-L2non-rigid motion tracking method that can adaptively stop the tracking error propagation. Extensive experiments over complex human body motions with occlusions, face and hand motions demonstrate that our approach substantially improves tracking robustness and surface reconstruction accuracy. Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai |
ICCV | 2 |
| 2015 | Video-audio driven real-time facial animationabstractWe present a real-time facial tracking and animation system based on a Kinect sensor with video and audio input. Our method requires no user-specific training and is robust to occlusions, large head rotations, and background noise. Given the color, depth and speech audio frames captured from an actor, our system first reconstructs 3D facial expressions and 3D mouth shapes from color and depth input with a multi-linear model. Concurrently a speaker-independent DNN acoustic model is applied to extract phoneme state posterior probabilities (PSPP) from the audio frames. After that, a lip motion regressor refines the 3D mouth shape based on both PSPP and expression weights of the 3D mouth shapes, as well as their confidences. Finally, the refined 3D mouth shape is combined with other parts of the 3D face to generate the final result. The whole process is fully automatic and executed in real time. The key component of our system is a data-driven regresor for modeling the correlation between speech data and mouth shapes. Based on a precaptured database of accurate 3D mouth shapes and associated speech audio from one speaker, the regressor jointly uses the input speech and visual features to refine the mouth shape of a new actor. We also present an improved DNN acoustic model. It not only preserves accuracy but also achieves real-time performance. Our method efficiently fuses visual and acoustic information for 3D facial performance capture. It generates more accurate 3D mouth motions than other approaches that are based on audio or video input only. It also supports video or audio only input for real-time facial animation. We evaluate the performance of our system with speech and facial expressions captured from different actors. Results demonstrate the efficiency and robustness of our method. Feng Xu 0005, Jinxiang Chai, Xin Tong 0001, Qiang Huo |
ACM Trans. Graph. | 2 |
| 2014 | A Data-Driven Approach for Facial Expression Retargeting in VideoabstractThis paper presents a data-driven approach for facial expression retargeting in video, i.e., synthesizing a face video of a target subject that mimics the expressions of a source subject in the input video. Our approach takes advantage of a pre-existing facial expression database of the target subject to achieve realistic synthesis. First, for each frame of the input video, a new facial expression similarity metric is proposed for querying the expression database of the target person to select multiple candidate images that are most similar to the input. The similarity metric is developed using a metric learning approach to reliably handle appearance difference between different subjects. Secondly, we employ an optimization approach to choose the best candidate image for each frame, resulting in a retrieved sequence that is temporally coherent. Finally, a spatio-temporal expression mapping method is employed to further improve the synthesized sequence. Experimental results show that our system is capable of generating high quality facial expression videos that match well with the input sequences, even when the source and target subjects have big identity difference. In addition, extensive evaluations demonstrate the high accuracy of the learned expression similarity metric and the effectiveness of our retrieval strategy. Kai Li 0016, Qionghai Dai, Ruiping Wang 0001, Yebin Liu, Feng Xu 0005, Jue Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2014 | Efficient Patch-Wise Non-Uniform Deblurring for a Single ImageabstractIn this paper, we address the problem of estimating a latent sharp image from a single spatially variant blurred image. Non-uniform deblurring methods based on projective motion path models formulate the blur as a linear combination of homographic projections of a clear image. But they are computationally expensive and require large memory due to the calculation and storage of a large number of the projections. Patch-wise non-uniform deblurring algorithms have been proposed to estimate each kernel locally by a uniform deblurring algorithm, which does not require to calculate and store the projections. The key issues of these methods are the accuracy of kernel estimation and the identification of erroneous kernels. To perform accurate kernel estimation, we employ the total variation (TV) regularization to recover a latent image, in which the edges are better enhanced and the ringing artifacts are reduced, rather than Tikhonov regularization that previous algorithms adopt. Thus blur kernels can be estimated more accurately from the latent image and estimated in a closed form while previous methods cannot estimate kernels in closed forms. To identify the erroneous kernels, we develop a novel metric, which is able to measure the similarity between the neighboring kernels. After replacing the erroneous kernels with the well-estimated ones, a clear image is obtained. The experiments show that our approach can achieve better results on the real-world blurry images while using less computation and memory. Xin Yu 0002, Feng Xu 0005, Shunli Zhang 0005, Li Zhang 0023 |
IEEE Trans. Multim. | 2 |
| 2014 | Controllable high-fidelity facial performance transferabstractRecent technological advances in facial capture have made it possible to acquire high-fidelity 3D facial performance data with stunningly high spatial-temporal resolution. Current methods for facial expression transfer, however, are often limited to large-scale facial deformation. This paper introduces a novel facial expression transfer and editing technique for high-fidelity facial performance data. The key idea of our approach is to decompose high-fidelity facial performances into high-level facial feature lines, large-scale facial deformation and fine-scale motion details and transfer them appropriately to reconstruct the retargeted facial animation in an efficient optimization framework. The system also allows the user to quickly modify and control the retargeted facial sequences in the spatial-temporal domain. We demonstrate the power of our approach by transferring and editing high-fidelity facial animation data from high-resolution source models to a wide range of target models, including both human faces and non-human faces such as "monster" and "dog". Feng Xu 0005, Jinxiang Chai, Xin Tong 0001 |
ACM Trans. Graph. | 1 |
| 2013 | Video-based hand manipulation capture through composite motion controlabstractThis paper describes a new method for acquiring physically realistic hand manipulation data from multiple video streams. The key idea of our approach is to introduce a composite motion control to simultaneously model hand articulation, object movement, and subtle interaction between the hand and object. We formulate video-based hand manipulation capture in an optimization framework by maximizing the consistency between the simulated motion and the observed image data. We search an optimal motion control that drives the simulation to best match the observed image data. We demonstrate the effectiveness of our approach by capturing a wide range of high-fidelity dexterous manipulation data. We show the power of our recovered motion controllers by adapting the captured motion data to new objects with different properties. The system achieves superior performance against alternative methods such as marker-based motion capture and kinematic hand motion tracking. Yangang Wang 0001, Jianyuan Min, Jianjie Zhang, Yebin Liu, Feng Xu 0005, Qionghai Dai, Jinxiang Chai |
ACM Trans. Graph. | 5 |
| 2012 | A data-driven approach for facial expression synthesis in videoabstractThis paper presents a method to synthesize a realistic facial animation of a target person, driven by a facial performance video of another person. Different from traditional facial animation approaches, our system takes advantage of an existing facial performance database of the target person, and generates the final video by retrieving frames from the database that have similar expressions to the input ones. To achieve this we develop an expression similarity metric for accurately measuring the expression difference between two video frames. To enforce temporal coherence, our system employs a shortest path algorithm to choose the optimal image for each frame from a set of candidate frames determined by the similarity metric. Finally, our system adopts an expression mapping method to further minimize the expression difference between the input and retrieved frames. Experimental results show that our system can generate high quality facial animation using the proposed data-driven approach. Kai Li 0016, Feng Xu 0005, Jue Wang 0001, Qionghai Dai, Yebin Liu |
CVPR | 2 |
| 2011 | Video-object segmentation and 3D-trajectory estimation for monocular video sequences
Feng Xu 0005, Kin-Man Lam 0001, Qionghai Dai |
Image Vis. Comput. | 1 |
| 2011 | Occlusion-Aware Motion Layer Extraction Under Large Interframe MotionsabstractExtracting motion layers from videos is an important task for video representation, analysis, and compression. For videos with large interframe motions, motion layer extraction is challenging in two respects: the estimation of large disparity motions and the awareness of large occluded regions. In this paper, we propose an effective method for motion layer extraction under large disparity motions. To robustly estimate large displacement motions, we have developed an efficient voting-based method that estimates planar homographies from sparse feature matches. To handle occlusions, we first integrate color and motion consistency into a Markov random field framework to achieve per-pixel assignment with occlusion detection. Then, we perform motion-color segmentation and an earth mover's distance-based comparison to determine motion labels for occluded pixels. Experimental results show that our proposed method achieves good performance in automatically extracting multiple moving objects under large disparity motions while maintaining a low computational cost. Feng Xu 0005, Qionghai Dai |
IEEE Trans. Image Process. | 1 |
| 2011 | Video-based characters: creating new human performances from a multi-view video databaseabstractWe present a method to synthesize plausible video sequences of humans according to user-defined body motions and viewpoints. We first capture a small database of multi-view video sequences of an actor performing various basic motions. This database needs to be captured only once and serves as the input to our synthesis algorithm. We then apply a marker-less model-based performance capture approach to the entire database to obtain pose and geometry of the actor in each database frame. To create novel video sequences of the actor from the database, a user animates a 3D human skeleton with novel motion and viewpoints. Our technique then synthesizes a realistic video sequence of the actor performing the specified motion based only on the initial database. The first key component of our approach is a new efficient retrieval strategy to find appropriate spatio-temporally coherent database frames from which to synthesize target video frames. The second key component is a warping-based texture synthesis approach that uses the retrieved most-similar database frames to synthesize spatio-temporally coherent target video frames. For instance, this enables us to easily create video sequences of actors performing dangerous stunts without them being placed in harm's way. We show through a variety of result videos and a user study that we can synthesize realistic videos of people, even if the target motions and camera views are different from the database content. Feng Xu 0005, Yebin Liu, Carsten Stoll, James Tompkin 0001, Gaurav Bharaj, Qionghai Dai, Hans-Peter Seidel, Jan Kautz, Christian Theobalt |
ACM Trans. Graph. | 1 |