VLDB 2026 Research / reviewers in the wild / expert
Chenglei Wu
dblp:49/8116
· DBLP profile ↗
57ranked-venue papers
13as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 11 first-author · 14 since 2021Computer networks · 13 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Practical Congestion Control Algorithm for Low-Latency Interactive Video StreamingabstractCongestion control (CC) plays a pivotal role in low-latency interactive video streaming such as cloud gaming. However, existing end-to-end CC methods often cause self-induced network queuing. As a result, they may largely delay video frame transmission and undermine the user’s quality of experience. In this paper, we present a new, practical CC algorithm namedPudicathat strives to achieve near-zero queuing delay and high link utilization while respecting cross-flow fairness. Pudica introduces several judicious approaches to utilize the paced frame to probe the bandwidth utilization ratio (BUR) instead of bandwidth itself. By leveraging BUR estimations, Pudica designs a holistic bitrate adjustment policy to balance low queuing, efficiency, and fairness. We conducted thorough and comprehensive evaluations in real production networks. In comparison to the state-of-the-art methods, Pudica reduces the average and tailed frame delay by 3.1$\times$and 5.1$\times$, respectively. Meanwhile, it increases the frame bitrate by 12.1%. Pudica has been deployed in a large-scale cloud gaming platform, currently serving millions of players. Shibo Wang 0002, Jianjun Xiao 0003, Chenglei Wu, Shusen Yang, Cong Zhao 0001, Chenren Xu, Hong Xu 0001, Jing Wang 0077 |
IEEE Trans. Netw. | 4 |
| 2024 | Diffusion Shape Prior for Wrinkle-Accurate Cloth RegistrationabstractRegistering clothes from 4D scans with vertex-accurate correspondence is challenging, yet important for dynamic appearance modeling and physics parameter estimation from real-world data. However, previous methods either rely on texture information, which is not always reliable, or achieve only coarse-level alignment. In this work, we present a novel approach to enabling accurate surface registration of texture-less clothes with large deformation. Our key idea is to effectively leverage a shape prior learned from pre-captured clothing using diffusion models. We also propose a multi-stage guidance scheme based on learned functional maps, which stabilizes registration for large-scale deformation even when they vary significantly from training data. Using high-fidelity real captured clothes, our experiments show that the proposed approach based on diffusion models generalizes better than surface registration with VAE or PCA-based priors, outperforming both optimization-based and learning-based non-rigid registration methods for both interpolation and extrapolation tests. Jingfan Guo, Fabian Prada, Donglai Xiang, Javier Romero 0002, Chenglei Wu, Hyun Soo Park, Takaaki Shiratori, Shunsuke Saito |
3DV | 5 |
| 2024 | Authentic Hand Avatar from a Phone Scan via Universal Hand ModelabstractThe authentic 3D hand avatar with every identifiable information, such as hand shapes and textures, is necessary for immersive experiences in AR/VR. In this paper, we present a universal hand model (UHM), which 1) can universally represent high-fidelity 3D hand meshes of arbitrary identities (IDs) and 2) can be adapted to each person with a short phone scan for the authentic hand avatar. For effective universal hand modeling, we perform tracking and modeling at the same time, while previous 3D hand models perform them separately. The conventional separate pipeline suffers from the accumulated errors from the tracking stage, which cannot be recovered in the modeling stage. On the other hand, ours does not suffer from the accumulated errors while having a much more concise overall pipeline. We additionally introduce a novel image matching loss function to address a skin sliding during the tracking and modeling, while existing works have not focused on it much. Finally, using learned priors from our UHM, we effectively adapt our UHM to each person's short phone scan for the authentic hand avatar. Gyeongsik Moon, Weipeng Xu, Rohan Joshi, Chenglei Wu, Takaaki Shiratori |
CVPR | 4 |
| 2024 | Pudica: Toward Near-Zero Queuing Delay in Congestion Control for Cloud Gaming
Shibo Wang 0002, Shusen Yang, Chenglei Wu, Longwei Jiang, Chenren Xu, Cong Zhao 0001, Xuesong Yang, Jianjun Xiao 0003, Changxi Zheng, Jing Wang 0077 |
NSDI | 4 |
| 2024 | AUGUR: Practical Mobile Multipath Transport Service for Low Tail Latency in Real-Time Streaming
Tingfeng Wang, Liying Wang 0011, Nian Wen, Jing Wang 0077, Chenglei Wu, Jiafeng Chen, Longwei Jiang, Shibo Wang 0002, Chenren Xu |
NSDI | 7 |
| 2024 | FabricDiffusion: High-Fidelity Texture Transfer for 3D Garments Generation from In-The-Wild Images
Cheng Zhang 0014, Yuanhao Wang 0011, Francisco Vicente 0001, Chenglei Wu, Thabo Beeler, Fernando De la Torre |
SIGGRAPH Asia | 4 |
| 2024 | Learning to Stabilize FacesabstractAbstract Nowadays, it is possible to scan faces and automatically register them with high quality. However, the resulting face meshes often need further processing: we need tostabilizethem to remove unwanted head movement. Stabilization is important for tasks like game development or movie making which require facial expressions to be cleanly separated from rigid head motion. Since manual stabilization is labor‐intensive, there have been attempts to automate it. However, previous methods remain impractical: they either still require some manual input, produce imprecise alignments, rely on dubious heuristics and slow optimization, or assume a temporally ordered input. Instead, we present a new learning‐based approach that is simple and fully automatic. We treat stabilization as a regression problem: given two face meshes, our network directly predicts the rigid transform between them that brings their skulls into alignment. We generate synthetic training data using a 3D Morphable Model (3DMM), exploiting the fact that 3DMM parameters separate skull motion from facial skin motion. Through extensive experiments we show that our approach outperforms the state‐of‐the‐art both quantitatively and qualitatively on the tasks of stabilizing discrete sets of facial expressions as well as dynamic facial performances. Furthermore, we provide an ablation study detailing the design choices and best practices to help others adopt our approach for their own uses. Jan Bednarík, Erroll Wood, Vasileios Choutas, Timo Bolkart, Daoye Wang, Chenglei Wu, Thabo Beeler |
Comput. Graph. Forum | 6 |
| 2023 | Buffer Awareness Neural Adaptive Video Streaming for Avoiding Extra Buffer Consumption
Tianchi Huang, Chao Zhou 0003, Rui-Xiao Zhang, Chenglei Wu, Lifeng Sun |
INFOCOM | 4 |
| 2023 | Who is the Rising Star? Demystifying the Promising Streamers in Crowdsourced Live StreamingabstractStreamers are the core competency of the crowd-sourced live streaming (CLS) platform. However, little work has explored how different factors relate to their popularity evolution patterns. In this paper, we will investigate a critical problem, i.e., how to discover the promising streamers in their early stage? To tackle this problem, we first conduct large-scale measurement on a real-world CLS dataset. We find that streamers can indeed be clustered into two evolution types (i.e., rising type and normal type), and these two types of streamers will show differences in some inherent properties. Traditional time-sequential models cannot handle this problem, because they are unable to capture the complicated interactivity and extensive heterogeneity in CLS scenarios. To address their shortcomings, we further propose Niffler, a novel heterogeneous attention temporal graph framework (HATG) for predicting the evolution types of CLS streamers. Specifically, through the graph neural network (GNN) and gated-recurrent-unit (GRU) structure, Niffler can capture both the interactive features and the evolutionary dynamics. Moreover, by integrating the attention mechanism in the model design, Niffler can intelligently preserve the heterogeneity when learning different levels of node representations. We systematically compare Niffler against multiple baselines from different categories, and the experimental results show that our proposed model can achieve the best prediction performance. Rui-Xiao Zhang, Tianchi Huang, Chenglei Wu, Lifeng Sun |
INFOCOM | 3 |
| 2023 | Owl: A Pre-and Post-processing Framework for Video Analytics in Low-light SurroundingsabstractThe low-light environment is an integral surrounding in real-world video analytic applications. Conventional wisdom claims that in order to adapt to the extensive computation requirement of the analytics model and achieve high inference accuracy, the overall pipeline should leverage a client-to-cloud framework that designs a cloud-based inference with on-demand video streaming. However, we show that due to the amplified noise, directly streaming the video in low-light scenarios can introduce significant bandwidth inefficiency.In this paper, we propose Owl, an intelligent framework to optimize the bandwidth utilization and inference accuracy for the low-light video analytic pipeline. The core idea of Owl is two-fold: on the one hand, we will deploy a light-weighted pre-processing module before transmission, through which we will get the denoised video and significantly reduce the transmitted data; on the other hand, we recover the information from the denoised video via an enhancement module in the server-side. Specifically, through well-designed training mechanism and content representation technique, Owl can dynamically select the best configuration for time-varying videos. Experiments with a variety of datasets and tasks show that Owl achieves significant bandwidth benefits, while consistently optimizing the inference accuracy. Rui-Xiao Zhang, Chaoyang Li 0002, Chenglei Wu, Tianchi Huang, Lifeng Sun |
INFOCOM | 3 |
| 2023 | Optimizing Adaptive Video Streaming with Human FeedbackabstractQuality of Experience (QoE)-driven adaptive bitrate (ABR) algorithms are typically optimized using QoE models that are based on the mean opinion score (MOS), while such principles may not account for user heterogeneity on rating scales, resulting in unexpected behaviors. In this paper, we propose Jade, which leverages reinforcement learning with human feedback(RLHF) technologies to better align the users' opinion scores. Jade's rank-based QoE model considers relative values of user ratings to interpret the subjective perception of video sessions. We implement linear-based and Deep Neural Network (DNN)-based architectures for satisfying both accuracy and generalization ability. We further propose entropy-aware reinforced mechanisms for training policies with the integration of the proposed QoE models. Experimental results demonstrate that Jade performs favorably on conventional metrics, such as quality and stall ratio, and improves QoE by 8.09%-38.13% in different network conditions, emphasizing the importance of user heterogeneity in QoE modeling and the potential of combining linear-based and DNN-based models for performance improvement. Tianchi Huang, Rui-Xiao Zhang, Chenglei Wu, Lifeng Sun |
ACM Multimedia | 3 |
| 2023 | Drivable Avatar Clothing: Faithful Full-Body Telepresence with Dynamic Clothing Driven by Sparse RGB-D InputabstractClothing is an important part of human appearance but challenging to model in photorealistic avatars. In this work we present avatars with dynamically moving loose clothing that can be faithfully driven by sparse RGB-D inputs as well as body and face motion. We propose a Neural Iterative Closest Point (N-ICP) algorithm that can efficiently track the coarse garment shape given sparse depth input. Given the coarse tracking results, the input RGB-D images are then remapped to texel-aligned features, which are fed into the drivable avatar models to faithfully reconstruct appearance details. We evaluate our method against recent image-driven synthesis baselines, and conduct a comprehensive analysis of the N-ICP algorithm. We demonstrate that our method can generalize to a novel testing environment, while preserving the ability to produce high-fidelity and faithful clothing dynamics and appearance. Donglai Xiang, Fabian Prada, Zhe Cao 0003, Chenglei Wu, Jessica K. Hodgins, Timur M. Bagautdinov |
SIGGRAPH Asia | 5 |
| 2023 | CT2Hair: High-Fidelity 3D Hair Modeling using Computed TomographyabstractWe introduce CT2Hair, a fully automatic framework for creating high-fidelity 3D hair models that are suitable for use in downstream graphics applications. Our approach utilizes real-world hair wigs as input, and is able to reconstruct hair strands for a wide range of hair styles. Our method leverages computed tomography (CT) to create density volumes of the hair regions, allowing us to see through the hair unlike image-based approaches which are limited to reconstructing the visible surface. To address the noise and limited resolution of the input density volumes, we employ a coarse-to-fine approach. This process first recovers guide strands with estimated 3D orientation fields, and then populates dense strands through a novel neural interpolation of the guide strands. The generated strands are then refined to conform to the input density volumes. We demonstrate the robustness of our approach by presenting results on a wide variety of hair styles and conducting thorough evaluations on both real-world and synthetic datasets. Code and data for this paper are at github.com/facebookresearch/CT2Hair. Yuefan Shen, Shunsuke Saito, Olivier Maury, Chenglei Wu, Jessica K. Hodgins, Youyi Zheng, Giljoo Nam |
ACM Trans. Graph. | 5 |
| 2023 | Practical Cloud-Edge Scheduling for Large-Scale Crowdsourced Live StreamingabstractEven though conventional wisdom claims that in order to improve viewer engagement, the cloud-edge providers should serve the viewers with the nearest edge nodes, however, we show that doing this for crowdsourced live streaming (CLS) services can introduce significant costs inefficiency. In this paper, we first carry out large-scale measurement analysis by using the real-world service data from Huawei Cloud, a representative cloud-edge provider in China. We observe that the massive number of channels has proposed great burdens to the operating expenditure of the cloud-edge providers, and most importantly, unbalanced viewer distribution makes the edge nodes suffer significant costs inefficiency. To tackle the above concerns, we proposeAggCast, a novel CLS scheduling framework to optimize the edge node utilization for the cloud-edge provider. The core idea ofAggCastis to aggregate some viewers that are initially scattered on different regions, and assign them to fewer pre-selected nodes, thereby reducing bandwidth costs. In particular, by integrating the useful insights obtained from our large-scale measurement,AggCastcan not only ensure that quality of experience (QoS) does not suffer degradation, but also satisfy the systematic requirements of CLS services.AggCasthas been A/B tested and fully deployed. The online and trace-driven experiments show that, compared to the most prevalent method,AggCastsaves over 16.3%back-to-source(BTS) bandwidth costs while significantly improving QoS (startup latency, stall frequency and stall time are reduced over 12.3%, 4.57% and 3.91%, respectively). Rui-Xiao Zhang, Changpeng Yang, Xiaochan Wang, Tianchi Huang, Chenglei Wu, Jiangchuan Liu, Lifeng Sun |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Neural Strands: Learning Hair Geometry and Appearance from Multi-view Images
Radu Alexandru Rosu, Shunsuke Saito, Chenglei Wu, Sven Behnke, Giljoo Nam |
ECCV (33) | 4 |
| 2022 | AggCast: Practical Cost-effective Scheduling for Large-scale Cloud-edge Crowdsourced Live StreamingabstractConventional wisdom claims that in order to improve viewer engagement, the cloud-edge providers should serve the viewers with the nearest edge nodes, however, we show that doing this for crowdsourced live streaming (CLS) services can introduce significant costs inefficiency. We observe that the massive number of channels has greatly burdened the operating expenditure of the cloud-edge providers, and most importantly, unbalanced viewer distribution makes the edge nodes suffer significant costs inefficiency. To tackle the above concerns, we propose AggCast, a novel CLS scheduling framework to optimize the edge node utilization for the cloud-edge provider. The core idea of AggCast is to aggregate some viewers who are initially scattered on different regions, and assign them to fewer pre-selected nodes, thereby reducing bandwidth costs. In particular, by leveraging the insights obtained from our large-scale measurement, AggCast can not only ensure quality of experience (QoS), but also satisfy the systematic requirements of CLS services. AggCast has been A/B tested and fully deployed in a top cloud-edge provider in China for over eight months. The online and trace-driven experiments show that, compared to the common practice, AggCast can save over 15% back-to-source (BTS) bandwidth costs while having no negative impacts on QoS. Rui-Xiao Zhang, Changpeng Yang, Xiaochan Wang, Tianchi Huang, Chenglei Wu, Jiangchuan Liu, Lifeng Sun |
ACM Multimedia | 5 |
| 2022 | Learning Tailored Adaptive Bitrate Algorithms to Heterogeneous Network Conditions: A Domain-Specific Priors and Meta-Reinforcement Learning ApproachabstractInternet adaptive video streaming is a typical form of video delivery that leverages adaptive bitrate (ABR) algorithms to provide video services with high quality of experience (QoE) for various users in diverse and unique network conditions. Such heterogeneous network environments, which can be viewed as exogenous input processes, often lead to the unstable performance of ABR algorithms. Unfortunately, learning-based ABR algorithm which generated by state-of-the-art reinforcement learning (RL) technologies achievesgood average performancebut fails to perform well in all kinds of network conditions. In this work, considering the video playback process as the Input-driven Markov Decision Process (IMDP), we propose$\text{A}^{2}$BR (Adaptation of ABR), a novel meta-RL ABR approach.$\text{A}^{2}$BR is mainly composed of an online stage and an offline stage. It leverages meta-RL to learn an initial meta-policy with various network conditions at the offline stage and makes decisions in personalized network conditions at the online stage. At the same time, we continually optimize the meta-policy to the tailor-made ABR policy for varying the current network environment within few shots. Moreover, in order to improve the learning efficiency, we fully utilize domain knowledge for implementing a virtual player to replay the previously experienced network. Using trace-driven experiments on various scenarios including different vehicles, users, network types, and heterogeneous user-preferences, we show that$\text{A}^{2}$BR outperforming recent ABR approaches with rapidly adapting to the personalized QoE metrics and specific network conditions. Testbed experimental results also illustrate the superiority of$\text{A}^{2}$BR in adapting to the unseen environments. Tianchi Huang, Chao Zhou 0003, Rui-Xiao Zhang, Chenglei Wu, Lifeng Sun |
IEEE J. Sel. Areas Commun. | 4 |
| 2022 | Pattern-Based Cloth Registration and Sparse-View AnimationabstractWe propose a novel multi-view camera pipeline for the reconstruction and registration of dynamic clothing. Our proposed method relies on a specifically designed pattern that allows for precise video tracking in each camera view. We triangulate the tracked points and register the cloth surface in a fine-grained geometric resolution and low localization error. Compared to state-of-the-art methods, our registration exhibits stable correspondence, tracking the same points on the deforming cloth surface along the temporal sequence. As an application, we demonstrate how the use of our registration pipeline greatly improves state-of-the-art pose-based drivable cloth models. Furthermore, we propose a novel model, Garment Avatar , for driving cloth from a dense tracking signal which is obtained from two opposing camera views. The method produces realistic reconstructions which are faithful to the actual geometry of the deforming cloth. In this setting, the user wears a garment with our custom pattern which enables our driving model to reconstruct the geometry. Our code and data are available at https://github.com/HalimiOshri/Pattern-Based-Cloth-Registration-and-Sparse-View-Animation. The released data includes our pattern and registered mesh sequences containing four different subjects and 15k frames in total. Oshri Halimi, Tuur Stuyck, Donglai Xiang, Timur M. Bagautdinov, He Wen 0001, Ron Kimmel, Takaaki Shiratori, Chenglei Wu, Yaser Sheikh, Fabian Prada |
ACM Trans. Graph. | 8 |
| 2022 | Dressing Avatars: Deep Photorealistic Appearance for Physically Simulated ClothingabstractDespite recent progress in developing animatable full-body avatars, realistic modeling of clothing - one of the core aspects of human self-expression - remains an open challenge. State-of-the-art physical simulation methods can generate realistically behaving clothing geometry at interactive rates. Modeling photorealistic appearance, however, usually requires physically-based rendering which is too expensive for interactive applications. On the other hand, data-driven deep appearance models are capable of efficiently producing realistic appearance, but struggle at synthesizing geometry of highly dynamic clothing and handling challenging body-clothing configurations. To this end, we introduce pose-driven avatars with explicit modeling of clothing that exhibit both photorealistic appearance learned from real-world data and realistic clothing dynamics. The key idea is to introduce a neural clothing appearance model that operates on top of explicit geometry: at training time we use high-fidelity tracking, whereas at animation time we rely on physically simulated geometry. Our core contribution is a physically-inspired appearance network, capable of generating photorealistic appearance with view-dependent and dynamic shadowing effects even for unseen body-clothing configurations. We conduct a thorough evaluation of our model and demonstrate diverse animation results on several subjects and different types of clothing. Unlike previous work on photorealistic full-body avatars, our approach can produce much richer dynamics and more realistic deformations even for many examples of loose clothing. We also demonstrate that our formulation naturally allows clothing to be used with avatars of different people while staying fully animatable, thus enabling, for the first time, photorealistic avatars with novel clothing. Donglai Xiang, Timur M. Bagautdinov, Tuur Stuyck, Fabian Prada, Javier Romero 0002, Weipeng Xu, Shunsuke Saito, Jingfan Guo, Breannan Smith, Takaaki Shiratori, Yaser Sheikh, Jessica K. Hodgins, Chenglei Wu |
ACM Trans. Graph. | 13 |
| 2021 | A Spherical Mixture Model Approach for 360 Video Virtual Cinematographyabstract360 video virtual cinematography attempts to direct a virtual camera and capture the most salient regions of 360 videos. In this paper, we propose a data-drive solution to achieve high-quality and diversified 360 cinematography based on crowd-sourced viewing histories. Specifically, we try to address two problems: 1) how to locate the semantically important regions of interest (RoI) from raw data, 2) how to generate virtual camera paths that follow chronological narratives. We first design a dynamic spherical mixture model based algorithm to locate variable number of RoIs on each video frame. We then model the camera transition and chronological orders with a Bayesian network and conditional probabilities. With the above two designs, we can generate “optimal” cinematography paths based on a dynamic programming algorithm. By modeling the RoIs as spherical mixture model, we are also able to provide diversified cinematography results. We show its effectiveness through extensive experiments. Chenglei Wu, Zhi Wang 0001, Lifeng Sun |
ICIP | 1 |
| 2021 | PAAS: a preference-aware deep reinforcement learning approach for 360° video streamingabstractConventional tile-based 360° video streaming methods, including deep reinforcement learning (DRL) based, ignore the interactive nature of 360° video streaming and download tiles following fixed sequential orders, thus failing to respond to the user's head motion changes. We show that these existing solutions suffer from either the prefetch accuracy or the playback stability drop. Furthermore, these methods are constrained to serve only one fixed streaming preference, causing extra training overhead and the lack of generalization on unseen preferences. In this paper, we propose a dual-queue streaming framework, with accuracy and stability purposes respectively, to enable the DRL agent to determine and change the tile download order without incurring overhead. We also design a preference-aware DRL algorithm to incentivize the agent to learn preference-dependent ABR decisions efficiently. Compared with state-of-the-art DRL baselines, our method not only significantly improves the streaming quality, e.g., increasing the average streaming quality by 13.6% on a public dataset, but also demonstrates better performance and generalization under dynamic preferences, e.g., an average quality improvement of 19.9% on unseen preferences. Chenglei Wu, Zhi Wang 0001, Lifeng Sun |
NOSSDAV | 1 |
| 2021 | Driving-signal aware full-body avatarsabstractWe present a learning-based method for building driving-signal aware full-body avatars. Our model is a conditional variational autoencoder that can be animated with incomplete driving signals, such as human pose and facial keypoints, and produces a high-quality representation of human geometry and view-dependent appearance. The core intuition behind our method is that better drivability and generalization can be achieved by disentangling the driving signals and remaining generative factors, which are not available during animation. To this end, we explicitly account for information deficiency in the driving signal by introducing a latent space that exclusively captures the remaining information, thus enabling the imputation of the missing factors required during full-body animation, while remaining faithful to the driving signal. We also propose a learnable localized compression for the driving signal which promotes better generalization, and helps minimize the influence of global chance-correlations often found in real datasets. For a given driving signal, the resulting variational model produces a compact space of uncertainty for missing factors that allows for an imputation strategy best suited to a particular application. We demonstrate the efficacy of our approach on the challenging problem of full-body animation for virtual telepresence with driving signals acquired from minimal sensors placed in the environment and mounted on a VR-headset. Timur M. Bagautdinov, Chenglei Wu, Tomas Simon, Fabian Prada, Takaaki Shiratori, Shih-En Wei, Weipeng Xu, Yaser Sheikh, Jason M. Saragih |
ACM Trans. Graph. | 2 |
| 2021 | Modeling clothing as a separate layer for an animatable human avatarabstractWe have recently seen great progress in building photorealistic animatable full-body codec avatars, but generating high-fidelity animation of clothing is still difficult. To address these difficulties, we propose a method to build an animatable clothed body avatar with an explicit representation of the clothing on the upper body from multi-view captured videos. We use a two-layer mesh representation to register each 3D scan separately with the body and clothing templates. In order to improve the photometric correspondence across different frames, texture alignment is then performed through inverse rendering of the clothing geometry and texture predicted by a variational autoencoder. We then train a new two-layer codec avatar with separate modeling of the upper clothing and the inner body layer. To learn the interaction between the body dynamics and clothing states, we use a temporal convolution network to predict the clothing latent code based on a sequence of input skeletal poses. We show photorealistic animation output for three different actors, and demonstrate the advantage of our clothed-body avatars over the single-layer avatars used in previous work. We also show the benefit of an explicit clothing model that allows the clothing texture to be edited in the animation output. Donglai Xiang, Fabian Prada, Timur M. Bagautdinov, Weipeng Xu, He Wen 0001, Jessica K. Hodgins, Chenglei Wu |
ACM Trans. Graph. | 8 |
| 2020 | MonoClothCap: Towards Temporally Coherent Clothing Capture from Monocular RGB VideoabstractWe present a method to capture temporally coherent dynamic clothing deformation from a monocular RGB video input. In contrast to the existing literature, our method does not require a pre-scanned personalized mesh template, and thus can be applied to in-the-wild videos. To constrain the output to a valid deformation space, we build statistical deformation models for three types of clothing: T- shirt, short pants and long pants. A differentiable renderer is utilized to align our captured shapes to the input frames by minimizing the difference in both silhouette, segmentation, and texture. We develop a UV texture growing method which expands the visible texture region of the clothing sequentially in order to minimize drift in deformation tracking. We also extract fine-grained wrinkle detail from the input videos by fitting the clothed surface to the normal maps estimated by a convolutional neural network. Our method produces temporally coherent reconstruction of body and clothing from monocular video. We demonstrate successful clothing capture results from a variety of challenging videos. Extensive quantitative experiments demonstrate the effectiveness of our method on metrics including body pose error and surface reconstruction error of the clothing. Donglai Xiang, Fabian Prada, Chenglei Wu, Jessica K. Hodgins |
3DV | 3 |
| 2020 | Stick: A Harmonious Fusion of Buffer-based and Learning-based Approach for Adaptive StreamingabstractOff-the-shelf buffer-based approaches leverage a simple yet effective buffer-bound to control the adaptive bitrate (ABR) streaming system. Nevertheless, such approaches in standard parameters fail to always provide high quality of experience (QoE) video streaming services under all considered network conditions. Meanwhile, state-of-the-art learning-based ABR approach Pensieve outperforms existing schemes but is impractical to deploy. Therefore, how to harmoniously fuse the buffer-based and learning-based approach has become a key challenge for further enhancing ABR methods. In this paper, we propose Stick, an ABR algorithm that fuses the deep learning method and traditional buffer-based method. Stick utilizes the deep reinforcement learning (DRL) method to train the neural network, which outputs the buffer-bound to control the buffer-based approach for maximizing the QoE metric with different parameters. Trace-driven emulation illustrates that Stick betters Pensieve by 3.5% - 9.41% with an overhead reduction of 88%. Moreover, aiming to further reduce the computational costs while preserving the performances, we propose Trigger, a light-weighted neural network that determines whether the buffer-bound should be adjusted. Experimental results show that Stick+Trigger rivals or outperforms existing schemes in average QoE by 1.7%-28%, and significantly reduces the Stick's computational overhead by 24%-61%. Meanwhile, we show that Trigger also helps other ABR schemes mitigate the overhead. Extensive results on real-world evaluation demonstrate the superiority of Stick over existing state-of-the-art approaches. Tianchi Huang, Chao Zhou 0003, Rui-Xiao Zhang, Chenglei Wu, Xin Yao 0003, Lifeng Sun |
INFOCOM | 4 |
| 2020 | Fully Convolutional Mesh Autoencoder using Efficient Spatially Varying KernelsabstractLearning latent representations of registered meshes is useful for many 3D tasks. Techniques have recently shifted to neural mesh autoencoders. Although they demonstrate higher precision than traditional methods, they remain unable to capture fine-grained deformations. Furthermore, these methods can only be applied to a template-specific surface mesh, and is not applicable to more general meshes, like tetrahedrons and non-manifold meshes. While more general graph convolution methods can be employed, they lack performance in reconstruction precision and require higher memory usage. In this paper, we propose a non-template-specific fully convolutional mesh autoencoder for arbitrary registered mesh data. It is enabled by our novel convolution and (un)pooling operators learned with globally shared weights and locally varying coefficients which can efficiently capture the spatially varying contents presented by irregular mesh connections. Our model outperforms state-of-the-art methods on reconstruction accuracy. In addition, the latent codes of our network are fully localized thanks to the fully convolutional structure, and thus have much higher interpolation capability than many traditional 3D mesh generation models. Yi Zhou 0023, Chenglei Wu, Zimo Li, Chen Cao 0001, Yuting Ye, Jason M. Saragih, Hao Li 0015, Yaser Sheikh |
NeurIPS | 2 |
| 2020 | Quality-Aware Neural Adaptive Video Streaming With Lifelong Imitation LearningabstractExisting Adaptive Bitrate (ABR) algorithms pick future video chunks' bitrates via fixed rules or offline trained models to ensure good quality of experience (QoE) for Internet video. Nevertheless, data analysis demonstrates that a good ABR algorithm is required to continually and fast update for adapting itself to time-varying network conditions. Therefore, we propose Comyco, a video quality-aware learning-based ABR approach that enormously improves recent schemes by i) picking the chunk with higher perceptual video qualities rather than video bitrates; ii) training the policy via imitating expert trajectories given by the expert strategy; iii) employing the lifelong learning method to continually train the model w.r.t the fresh trace collected by the users. To achieve this, we develop a complete quality-aware lifelong imitation learning-based ABR system, construct quality-based neural network architecture, collect a quality-driven video dataset, and estimate QoE metrics with video quality features. Using trace-driven and real-world experiments, we demonstrate Comyco reaches 1700-fold improvements in the number of samples required and 16-fold speedup in the training time compared with the prior work. Meanwhile, Comyco outperforms existing methods, with the improvements on average QoE of 7.5%-16.79%. Moreover, experimental results on continual training also illustrate that lifelong learning helps Comyco further improve the average QoE of 1.07%-9.81% in comparison to the offline trained model. Tianchi Huang, Chao Zhou 0003, Xin Yao 0003, Rui-Xiao Zhang, Chenglei Wu, Lifeng Sun |
IEEE J. Sel. Areas Commun. | 5 |
| 2020 | Constraining dense hand surface tracking with elasticityabstractMany of the actions that we take with our hands involve self-contact and occlusion: shaking hands, making a fist, or interlacing our fingers while thinking. This use of of our hands illustrates the importance of tracking hands through self-contact and occlusion for many applications in computer vision and graphics, but existing methods for tracking hands and faces are not designed to treat the extreme amounts of self-contact and self-occlusion exhibited by common hand gestures. By extending recent advances in vision-based tracking and physically based animation, we present the first algorithm capable of tracking high-fidelity hand deformations through highly self-contacting and self-occluding hand gestures, for both single hands and two hands. By constraining a vision-based tracking algorithm with a physically based deformable model, we obtain an algorithm that is robust to the ubiquitous self-interactions and massive self-occlusions exhibited by common hand gestures, allowing us to track two hand interactions and some of the most difficult possible configurations of a human hand. Breannan Smith, Chenglei Wu, He Wen 0001, Patrick Peluse, Yaser Sheikh, Jessica K. Hodgins, Takaaki Shiratori |
ACM Trans. Graph. | 2 |
| 2020 | A Practical Learning-based Approach for Viewer Scheduling in the Crowdsourced Live StreamingabstractScheduling viewers effectively among different Content Delivery Network (CDN) providers is challenging owing to the extreme diversity in the crowdsourced live streaming (CLS) scenarios. Abundant algorithms have been proposed in recent years, which, however, suffer from a critical limitation: Due to their inaccurate feature engineering or naive rules, they cannot optimally schedule viewers. To address this concern, we put forward LTS (Learn to Schedule), a novel scheduling algorithm that can adapt to the dynamics from both viewer traffics and CDN performance. In detail, we first propose LTS-RL, an approach that schedules CLS viewers based on deep reinforcement learning (DRL). Since LTS-RL is trained in an end-to-end way, it can automatically learn scheduling algorithms without any pre-programmed models or assumptions about the environment dynamics. At the same time, to practically deploy LTS-RL, we then use the decision tree and imitation learning to convert LTS-RL into a more light-weighted and interpretable model, which is denoted as Fast-LTS. After the extensive evaluation of the real data from a leading CLS platform in China, we demonstrate that our proposed model (both LTS-RL and Fast-LTS) can improve the average quality of experience (QoE) over state-of-the-art approaches by 8.71--15.63%. At the same time, we also demonstrate that Fast-LTS can faithfully convert the complicated LTS-RL with slight performance degradation (< 2%), while significantly reducing the decision time (×7--10). Rui-Xiao Zhang, Tianchi Huang, Haitian Pang, Xin Yao 0003, Chenglei Wu, Lifeng Sun |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2019 | Strand-Accurate Multi-View Hair CaptureabstractHair is one of the most challenging objects to reconstruct due to its micro-scale structure and a large number of repeated strands with heavy occlusions. In this paper, we present the first method to capture high-fidelity hair geometry with strand-level accuracy. Our method takes three stages to achieve this. In the first stage, a new multi-view stereo method with a slanted support line is proposed to solve the hair correspondences between different views. In detail, we contribute a novel cost function consisting of both photo-consistency term and geometric term that reconstructs each hair pixel as a 3D line. By merging all the depth maps, a point cloud, as well as local line directions for each point, is obtained. Thus, in the second stage, we feature a novel strand reconstruction method with the mean-shift to convert the noisy point data to a set of strands. Lastly, we grow the hair strands with multi-view geometric constraints to elongate the short strands and recover the missing strands, thus significantly increasing the reconstruction completeness. We evaluate our method on both synthetic data and real captured data, showing that our method can reconstruct hair strands with sub-millimeter accuracy. Giljoo Nam, Chenglei Wu, Min H. Kim 0001, Yaser Sheikh |
CVPR | 2 |
| 2019 | Towards Faster and Better Federated Learning: A Feature Fusion ApproachabstractFederated learning enables on-device training over distributed networks consisting of a massive amount of modern smart devices, such as smartphones and IoT devices. However, the leading optimization algorithm in such settings, i.e., federated averaging, suffers from heavy communication cost and inevitable performance drop, especially when the local data is distributed in a Non-IID way. In this paper, we propose a feature fusion method to address this problem. By aggregating the features from both the local and global models, we achieve a higher accuracy at less communication cost. Furthermore, the feature fusion modules offer better initialization for newly incoming clients and thus speed up the process of convergence. Experiments in popular federated learning scenarios show that our federated learning algorithm with feature fusion mechanism outperforms baselines in both accuracy and generalization ability while reducing the number of communication rounds by more than 60%. Xin Yao 0003, Tianchi Huang, Chenglei Wu, Rui-Xiao Zhang, Lifeng Sun |
ICIP | 3 |
| 2019 | Tiyuntsong: A Self-Play Reinforcement Learning Approach for ABR Video StreamingabstractExisting reinforcement learning (RL)-based adaptive bitrate (ABR) approaches outperform the previous fixed control rules based methods by improving the Quality of Experience (QoE) score, as the QoE metric can hardly provide clear guidance for optimization, finally resulting in the unexpected strategies. In this paper, we propose Tiyuntsong, a self-play reinforcement learning approach with generative adversarial network (GAN)-based method for ABR video streaming. Tiyuntsong learns strategies automatically by training two agents who are competing against each other. Note that the competition results are determined by a set of rules rather than a numerical QoE score that allows clearer optimization objectives. Meanwhile, we propose GAN Enhancement Module to extract hidden features from the past status for preserving the information without the limitations of sequence lengths. Using testbed experiments, we show that the utilization of GAN significantly improves the Tiyuntsong's performance. By comparing the performance of ABRs, we observe that Tiyuntsong also betters existing ABR algorithms in the underlying metrics. Tianchi Huang, Xin Yao 0003, Chenglei Wu, Rui-Xiao Zhang, Zhengyuan Pang, Lifeng Sun |
ICME | 3 |
| 2019 | Comyco: Quality-Aware Adaptive Video Streaming via Imitation LearningabstractLearning-based Adaptive Bit Rate~(ABR) method, aiming to learn outstanding strategies without any presumptions, has become one of the research hotspots for adaptive streaming. However, it is still suffering from several issues, i.e., low sample efficiency and lack of awareness of the video quality information. In this paper, we propose Comyco, a video quality-aware ABR approach that enormously improves the learning-based methods by tackling the above issues. Comyco trains the policy via imitating expert trajectories given by the instant solver, which can not only avoid redundant exploration but also make better use of the collected samples. Meanwhile, Comyco attempts to pick the chunk with higher perceptual video qualities rather than video bitrates. To achieve this, we construct Comyco's neural network architecture, video datasets and QoE metrics with video quality features. Using trace-driven and real world experiments, we demonstrate significant improvements of Comyco's sample efficiency in comparison to prior work, with 1700x improvements in terms of the number of samples required and 16x improvements on training time required. Moreover, results illustrate that Comyco outperforms previously proposed methods, with the improvements on average QoE of 7.5% - 16.79%. Especially, Comyco also surpasses state-of-the-art approach Pensieve by 7.37% on average video quality under the same rebuffering time. Tianchi Huang, Chao Zhou 0003, Rui-Xiao Zhang, Chenglei Wu, Xin Yao 0003, Lifeng Sun |
ACM Multimedia | 4 |
| 2019 | Livesmart: A QoS-Guaranteed Cost-Minimum Framework of Viewer Scheduling for Crowdsourced Live StreamingabstractViewer scheduling among different CDN providers in crowdsourced live streaming (CLS) service is especially challenging due to the large-scale dynamic viewers as well as the time-variant performance of the content delivery network. A practical scheduling method should tackle the following challenges: 1) accurate modeling of viewer patterns and CDN performance; 2) intelligent workload offloading to save costs while guaranteeing the quality of service (QoS); 3) and ease of integration with practical CDN infrastructure in CLS platforms. Rui-Xiao Zhang, Tianchi Huang, Haitian Pang, Xin Yao 0003, Chenglei Wu, Jiangchuan Liu, Lifeng Sun |
ACM Multimedia | 6 |
| 2019 | Generalizing Rate Control Strategies for Realtime Video Streaming via Learning from Deep LearningabstractThe leading learning-based rate control method, i.e., QARC, achieves state-of-the-art performances but fails to interpret the fundamental principles, and thus lacks the abilities to further improve itself efficiently. In this paper, we propose EQARC (Explainable QARC) via reconstructing QARC's modules, aiming to demystify how QARC works. In details, we first utilize a novel hybrid attention-based CNN+GRU model to re-characterize the original quality prediction network and reasonably replace the QARC's 1D-CNN layers with 2D-CNN layers. Using trace-driven experiment, we demonstrate the superiority of EQARC over existing state-of-the-art approaches. Next, we collect several useful information from each interpretable modules and learn the insight of EQARC. Following this step, we further propose AQARC (Advanced QARC), which is the light-weighted version of QARC. Experimental results show that AQARC achieves the same performances as the QARC with an overhead reduction of 90%. In short, through learning from deep learning, we generalize a rate control method which can both reach high performance and reduce computation cost. Tianchi Huang, Rui-Xiao Zhang, Chenglei Wu, Xin Yao 0003, Chao Zhou 0003, Lifeng Sun |
MMAsia | 3 |
| 2019 | Enhancing the crowdsourced live streaming: a deep reinforcement learning approachabstractWith the growing demand for crowdsourced live streaming (CLS), how to schedule the large-scale dynamic viewers effectively among different Content Delivery Network (CDN) providers has become one of the most significant challenges for CLS platforms. Although abundant algorithms have been proposed in recent years, they suffer from a critical limitation: due to their inaccurate feature engineering or naive rules, they cannot optimally schedule viewers. To address this concern, we propose LTS (Learn to schedule), a deep reinforcement learning (DRL) based scheduling approach that can dynamically adapt to the variation of both viewer traffics and CDN performance. After the extensive evaluation the real data from a leading CLS platform in China, we demonstrate that LTS improves the average quality of experience (QoE) over state-of-the-art approach by 8.71%-15.63%. Rui-Xiao Zhang, Tianchi Huang, Haitian Pang, Xin Yao 0003, Chenglei Wu, Lifeng Sun |
NOSSDAV | 6 |
| 2019 | Adversarial Feature Alignment: Avoid Catastrophic Forgetting in Incremental Task Lifelong Learningabstract, is one of the major roadblocks that prevent deep neural networks from achieving human-level artificial intelligence. Several research efforts (e.g., lifelong or continual learning algorithms) have proposed to tackle this problem. However, they either suffer from an accumulating drop in performance as the task sequence grows longer, or require storing an excessive number of model parameters for historical memory, or cannot obtain competitive performance on the new tasks. In this letter, we focus on the incremental multitask image classification scenario. Inspired by the learning process of students, who usually decompose complex tasks into easier goals, we propose an adversarial feature alignment method to avoid catastrophic forgetting. In our design, both the low-level visual features and high-level semantic features serve as soft targets and guide the training process in multiple stages, which provide sufficient supervised information of the old tasks and help to reduce forgetting. Due to the knowledge distillation and regularization phenomena, the proposed method gains even better performance than fine-tuning on the new tasks, which makes it stand out from other methods. Extensive experiments in several typical lifelong learning scenarios demonstrate that our method outperforms the state-of-the-art methods in both accuracy on new tasks and performance preservation on old tasks. Xin Yao 0003, Tianchi Huang, Chenglei Wu, Rui-Xiao Zhang, Lifeng Sun |
Neural Comput. | 3 |
| 2018 | Modeling Facial Geometry Using Compositional VAEsabstractWe propose a method for learning non-linear face geometry representations using deep generative models. Our model is a variational autoencoder with multiple levels of hidden variables where lower layers capture global geometry and higher ones encode more local deformations. Based on that, we propose a new parameterization of facial geometry that naturally decomposes the structure of the human face into a set of semantically meaningful levels of detail. This parameterization enables us to do model fitting while capturing varying level of detail under different types of geometrical constraints. Timur M. Bagautdinov, Chenglei Wu, Jason M. Saragih, Pascal Fua, Yaser Sheikh |
CVPR | 2 |
| 2018 | Learning Patch Reconstructability for Accelerating Multi-View StereoabstractWe present an approach to accelerate multi-view stereo (MVS) by prioritizing computation on image patches that are likely to produce accurate 3D surface reconstructions. Our key insight is that the accuracy of the surface reconstruction from a given image patch can be predicted significantly faster than performing the actual stereo matching. The intuition is that non-specular, fronto-parallel, in-focus patches are more likely to produce accurate surface reconstructions than highly specular, slanted, blurry patches - and that these properties can be reliably predicted from the image itself. By prioritizing stereo matching on a subset of patches that are highly reconstructable and also cover the 3D surface, we are able to accelerate MVS with minimal reduction in accuracy and completeness. To predict the reconstructability score of an image patch from a single view, we train an image-to-reconstructability neural network: the I2RNet. This reconstructability score enables us to efficiently identify image patches that are likely to provide the most accurate surface estimates before performing stereo matching. We demonstrate that the I2RNet, when trained on the ScanNet dataset, generalizes to the DTU and Tanks & Temples MVS datasets. By using our I2RNet with an existing MVS implementation, we show that our method can achieve more than a 30× speed-up over the baseline with only an minimal loss in completeness. Alex Poms, Chenglei Wu, Shoou-I Yu, Yaser Sheikh |
CVPR | 2 |
| 2018 | DDRNet: Depth Map Denoising and Refinement for Consumer Depth Cameras Using Cascaded CNNs
Shi Yan 0007, Chenglei Wu, Lizhen Wang 0002, Feng Xu 0005, Liang An 0001, Yebin Liu |
ECCV (10) | 2 |
| 2018 | Deep incremental learning for efficient high-fidelity face trackingabstractIn this paper, we present an incremental learning framework for efficient and accurate facial performance tracking. Our approach is to alternate the modeling step, which takes tracked meshes and texture maps to train our deep learning-based statistical model, and the tracking step, which takes predictions of geometry and texture our model infers from measured images and optimize the predicted geometry by minimizing image, geometry and facial landmark errors. Our Geo-Tex VAE model extends the convolutional variational autoencoder for face tracking, and jointly learns and represents deformations and variations in geometry and texture from tracked meshes and texture maps. To accurately model variations in facial geometry and texture, we introduce the decomposition layer in the Geo-Tex VAE architecture which decomposes the facial deformation into global and local components. We train the global deformation with a fully-connected network and the local deformations with convolutional layers. Despite running this model on each frame independently - thereby enabling a high amount of parallelization - we validate that our framework achieves sub-millimeter accuracy on synthetic data and outperforms existing methods. We also qualitatively demonstrate high-fidelity, long-duration facial performance tracking on several actors. Chenglei Wu, Takaaki Shiratori, Yaser Sheikh |
ACM Trans. Graph. | 1 |
| 2017 | A Dataset for Exploring User Behaviors in VR Spherical Video StreamingabstractWith Virtual Reality (VR) devices and content getting increasingly popular, understanding user behaviors in virtual environment is important for not only VR product design but also user experience improvement. In VR applications, the head movement is one of the most important user behaviors, which can reflect a user's visual attention, preference, and even unique motion pattern. However, to the best of our knowledge, no dataset containing this information is publicly available. In this paper, we present a head tracking dataset composed of 48 users (24 males and 24 females) watching 18 sphere videos from 5 categories. We carefully record how users watch the videos, how their heads move in each session, what directions they focus, and what content they can remember after each session. Based on this dataset, we show that people share certain common patterns in VR spherical video streaming, which are different from conventional video streaming. We believe the dataset can serve good resource for exploring user behavior patterns in VR applications. Chenglei Wu, Zhihao Tan, Zhi Wang 0001, Shiqiang Yang |
MMSys | 1 |
| 2016 | Crowdsourced Live Streaming over Aggregated Edge NetworksabstractRecent years have witnessed a dramatic increase of user-generated video services. In such user-generated video services, crowdsourced live streaming (e.g., Periscope, Twitch) has significantly challenged today's content delivery infrastructure: today's edge networks (e.g., 4G, Wi-Fi) have limited uplink capacity support, making high-bitrate live streaming over such links fundamentally impossible. In this paper, we propose to let broadcasters (i.e., users who generate the video) upload crowdsourced video streams using aggregated network resources from multiple edge networks. There are several challenges in the proposal: First, how to design a framework that aggregates bandwidth from multiple edge networks? Second, how to make this framework transparent to today's crowdsourced live stream- ing services? Third, how to maximize the streaming quality for the whole system? We design a multi-objective and deployable bandwidth aggregation system BASS to address these challenges: (1) We propose an aggregation framework transparent to today's crowdsourced live streaming services, using an edge proxy box and aggregation cloud paradigm; (2) We dynamically allocate geo- distributed cloud aggregation servers to enable MPTCP (i.e., multi- path TCP), according to location and network characteristics of both broadcasters and the original streaming servers; (3) We maximize the overall performance gain for the whole system, by matching streams with the best aggregation paths. Chenglei Wu, Zhi Wang 0001, Jiangchuan Liu, Shiqiang Yang |
GLOBECOM | 1 |
| 2016 | Corrective 3D reconstruction of lips from monocular videoabstractIn facial animation, the accurate shape and motion of the lips of virtual humans is of paramount importance, since subtle nuances in mouth expression strongly influence the interpretation of speech and the conveyed emotion. Unfortunately, passive photometric reconstruction of expressive lip motions, such as a kiss or rolling lips, is fundamentally hard even with multi-view methods in controlled studios. To alleviate this problem, we present a novel approach for fully automatic reconstruction of detailed and expressive lip shapes along with the dense geometry of the entire face, from just monocular RGB video. To this end, we learn the difference between inaccurate lip shapes found by a state-of-the-art monocular facial performance capture approach, and the true 3D lip shapes reconstructed using a high-quality multi-view system in combination with applied lip tattoos that are easy to track. A robust gradient domain regressor is trained to infer accurate lip shapes from coarse monocular reconstructions, with the additional help of automatically extracted inner and outer 2D lip contours. We quantitatively and qualitatively show that our monocular approach reconstructs higher quality lip shapes, even for complex shapes like a kiss or lip rolling, than previous monocular approaches. Furthermore, we compare the performance of person-specific and multi-person generic regression strategies and show that our approach generalizes to new individuals and general scenes, enabling high-fidelity reconstruction even from commodity video footage. Pablo Garrido 0001, Michael Zollhöfer, Chenglei Wu, Derek Bradley, Patrick Pérez, Thabo Beeler, Christian Theobalt |
ACM Trans. Graph. | 3 |
| 2016 | Model-based teeth reconstructionabstractIn recent years, sophisticated image-based reconstruction methods for the human face have been developed. These methods capture highly detailed static and dynamic geometry of the whole face, or specific models of face regions, such as hair, eyes or eye lids. Unfortunately, image-based methods to capture the mouth cavity in general, and the teeth in particular, have received very little attention. The accurate rendering of teeth, however, is crucial for the realistic display of facial expressions, and currently high quality face animations resort to tooth row models created by tedious manual work. In dentistry, special intra-oral scanners for teeth were developed, but they are invasive, expensive, cumbersome to use, and not readily available. In this paper, we therefore present the first approach for non-invasive reconstruction of an entire person-specific tooth row from just a sparse set of photographs of the mouth region. The basis of our approach is a new parametric tooth row prior learned from high quality dental scans. A new model-based reconstruction approach fits teeth to the photographs such that visible teeth are accurately matched and occluded teeth plausibly synthesized. Our approach seamlessly integrates into photogrammetric multi-camera reconstruction setups for entire faces, but also enables high quality teeth modeling from normal uncalibrated photographs and even short videos captured with a mobile phone. Chenglei Wu, Derek Bradley, Pablo Garrido 0001, Michael Zollhöfer, Christian Theobalt, Markus Gross 0001, Thabo Beeler |
ACM Trans. Graph. | 1 |
| 2016 | An anatomically-constrained local deformation model for monocular face captureabstractWe present a new anatomically-constrained local face model and fitting approach for tracking 3D faces from 2D motion data in very high quality. In contrast to traditional global face models, often built from a large set of blendshapes, we propose a local deformation model composed of many small subspaces spatially distributed over the face. Our local model offers far more flexibility and expressiveness than global blendshape models, even with a much smaller model size. This flexibility would typically come at the cost of reduced robustness, in particular during the under-constrained task of monocular reconstruction. However, a key contribution of this work is that we consider the face anatomy and introduce subspace skin thickness constraints into our model, which constrain the face to only valid expressions and helps counteract depth ambiguities in monocular tracking. Given our new model, we present a novel fitting optimization that allows 3D facial performance reconstruction from a single view at extremely high quality, far beyond previous fitting approaches. Our model is flexible, and can be applied also when only sparse motion data is available, for example with marker-based motion capture or even face posing from artistic sketches. Furthermore, by incorporating anatomical constraints we can automatically estimate the rigid motion of the skull, obtaining a rigid stabilization of the performance for free. We demonstrate our model and single-view fitting method on a number of examples, including, for the first time, extreme local skin deformation caused by external forces such as wind, captured from a single high-speed camera. Chenglei Wu, Derek Bradley, Markus Gross 0001, Thabo Beeler |
ACM Trans. Graph. | 1 |
| 2015 | Shading-based refinement on volumetric signed distance functionsabstractWe present a novel method to obtain fine-scale detail in 3D reconstructions generated with low-budget RGB-D cameras or other commodity scanning devices. As the depth data of these sensors is noisy, truncated signed distance fields are typically used to regularize out the noise, which unfortunately leads to over-smoothed results. In our approach, we leverage RGB data to refine these reconstructions through shading cues, as color input is typically of much higher resolution than the depth data. As a result, we obtain reconstructions with high geometric detail, far beyond the depth resolution of the camera itself. Our core contribution is shading-based refinement directly on the implicit surface representation, which is generated from globally-aligned RGB-D images. We formulate the inverse shading problem on the volumetric distance field, and present a novel objective function which jointly optimizes for fine-scale surface geometry and spatially-varying surface reflectance. In order to enable the efficient reconstruction of sub-millimeter detail, we store and process our surface using a sparse voxel hashing scheme which we augment by introducing a grid hierarchy. A tailored GPU-based Gauss-Newton solver enables us to refine large shape models to previously unseen resolution within only a few seconds. Michael Zollhöfer, Angela Dai, Matthias Innmann, Chenglei Wu, Marc Stamminger, Christian Theobalt, Matthias Nießner |
ACM Trans. Graph. | 4 |
| 2014 | Real-time non-rigid reconstruction using an RGB-D cameraabstractWe present a combined hardware and software solution for markerless reconstruction of non-rigidly deforming physical objects with arbitrary shape in real-time . Our system uses a single self-contained stereo camera unit built from off-the-shelf components and consumer graphics hardware to generate spatio-temporally coherent 3D models at 30 Hz. A new stereo matching algorithm estimates real-time RGB-D data. We start by scanning a smooth template model of the subject as they move rigidly. This geometric surface prior avoids strong scene assumptions, such as a kinematic human skeleton or a parametric shape model. Next, a novel GPU pipeline performs non-rigid registration of live RGB-D data to the smooth template using an extended non-linear as-rigid-as-possible (ARAP) framework. High-frequency details are fused onto the final mesh using a linear deformation model. The system is an order of magnitude faster than state-of-the-art methods, while matching the quality and robustness of many offline algorithms. We show precise real-time reconstructions of diverse scenes, including: large deformations of users' heads, hands, and upper bodies; fine-scale wrinkles and folds of skin and clothing; and non-rigid interactions performed by users on flexible objects such as toys. We demonstrate how acquired models can be used for many interactive scenarios, including re-texturing, online performance capture and preview, and real-time shape and motion re-targeting. Michael Zollhöfer, Matthias Nießner, Shahram Izadi, Christoph Rhemann, Christopher Zach, Matthew Fisher, Chenglei Wu, Andrew W. Fitzgibbon, Charles T. Loop, Christian Theobalt, Marc Stamminger |
ACM Trans. Graph. | 7 |
| 2013 | Capturing Relightable Human Performances under General Uncontrolled IlluminationabstractAbstract We present a novel approach to create relightable free‐viewpoint human performances from multi‐view video recorded under general uncontrolled and uncalibated illumination. We first capture a multi‐view sequence of an actor wearing arbitrary apparel and reconstruct a spatio‐temporal coherent coarse 3D model of the performance using a marker‐less tracking approach. Using these coarse reconstructions, we estimate the low‐frequency component of the illumination in a spherical harmonics (SH) basis as well as the diffuse reflectance, and then utilize them to estimate the dynamic geometry detail of human actors based on shading cues. Given the high‐quality time‐varying geometry, the estimated illumination is extended to the all‐frequency domain by re‐estimating it in the wavelet basis. Finally, the high‐quality all‐frequency illumination is utilized to reconstruct the spatially‐varying BRDF of the surface. The recovered time‐varying surface geometry and spatially‐varying non‐Lambertian reflectance allow us to generate high‐quality model‐based free view‐point videos of the actor under novel illumination conditions. Our method enables plausible reconstruction of relightable dynamic scene models without a complex controlled lighting apparatus, and opens up a path towards relightable performance capture in less constrained environments and using less complex acquisition setups. Chenglei Wu, Carsten Stoll, Yebin Liu, Kiran Varanasi, Qionghai Dai, Christian Theobalt |
Comput. Graph. Forum | 2 |
| 2013 | Reconstructing detailed dynamic face geometry from monocular videoabstractDetailed facial performance geometry can be reconstructed using dense camera and light setups in controlled studios. However, a wide range of important applications cannot employ these approaches, including all movie productions shot from a single principal camera. For post-production, these require dynamic monocular face capture for appearance modification. We present a new method for capturing face geometry from monocular video. Our approach captures detailed, dynamic, spatio-temporally coherent 3D face geometry without the need for markers. It works under uncontrolled lighting, and it successfully reconstructs expressive motion including high-frequency face detail such as folds and laugh lines. After simple manual initialization, the capturing process is fully automatic, which makes it versatile, lightweight and easy-to-deploy. Our approach tracks accurate sparse 2D features between automatically selected key frames to animate a parametric blend shape model, which is further refined in pose, expression and shape by temporally coherent optical flow and photometric stereo. We demonstrate performance capture results for long and complex face sequences captured indoors and outdoors, and we exemplify the relevance of our approach as an enabling technology for model-based face editing in movies and video, such as adding new facial textures, as well as a step towards enabling everyone to do facial performance capture with a single affordable camera. Pablo Garrido 0001, Levi Valgaerts, Chenglei Wu, Christian Theobalt |
ACM Trans. Graph. | 3 |
| 2013 | On-set performance capture of multiple actors with a stereo cameraabstractState-of-the-art marker-less performance capture algorithms reconstruct detailed human skeletal motion and space-time coherent surface geometry. Despite being a big improvement over marker-based motion capture methods, they are still rarely applied in practical VFX productions as they require ten or more cameras and a studio with controlled lighting or a green screen background. If one was able to capture performances directly on a general set using only the primary stereo camera used for principal photography, many possibilities would open up in virtual production and previsualization, the creation of virtual actors, and video editing during post-production. We describe a new algorithm which works towards this goal. It is able to track skeletal motion and detailed surface geometry of one or more actors from footage recorded with a stereo rig that is allowed to move. It succeeds in general sets with uncontrolled background and uncontrolled illumination, and scenes in which actors strike non-frontal poses. It is one of the first performance capture methods to exploit detailed BRDF information and scene illumination for accurate pose tracking and surface refinement in general scenes. It also relies on a new foreground segmentation approach that combines appearance, stereo, and pose tracking results to segment out actors from the background. Appearance, segmentation, and motion cues are combined in a new pose optimization framework that is robust under uncontrolled lighting, uncontrolled background and very sparse camera views. Chenglei Wu, Carsten Stoll, Levi Valgaerts, Christian Theobalt |
ACM Trans. Graph. | 1 |
| 2012 | Full Body Performance Capture under Uncontrolled and Varying Illumination: A Shading-Based Approach
Chenglei Wu, Kiran Varanasi, Christian Theobalt |
ECCV (4) | 1 |
| 2012 | Lightweight binocular facial performance capture under uncontrolled lightingabstractRecent progress in passive facial performance capture has shown impressively detailed results on highly articulated motion. However, most methods rely on complex multi-camera set-ups, controlled lighting or fiducial markers. This prevents them from being used in general environments, outdoor scenes, during live action on a film set, or by freelance animators and everyday users who want to capture their digital selves. In this paper, we therefore propose a lightweight passive facial performance capture approach that is able to reconstruct high-quality dynamic facial geometry from only a single pair of stereo cameras. Our method succeeds under uncontrolled and time-varying lighting, and also in outdoor scenes. Our approach builds upon and extends recent image-based scene flow computation, lighting estimation and shading-based refinement algorithms. It integrates them into a pipeline that is specifically tailored towards facial performance reconstruction from challenging binocular footage under uncontrolled lighting. In an experimental evaluation, the strong capabilities of our method become explicit: We achieve detailed and spatio-temporally coherent results for expressive facial motion in both indoor and outdoor scenes -- even from low quality input images recorded with a hand-held consumer stereo camera. We believe that our approach is the first to capture facial performances of such high quality from a single stereo rig and we demonstrate that it brings facial performance capture out of the studio, into the wild, and within the reach of everybody. Levi Valgaerts, Chenglei Wu, Andrés Bruhn, Hans-Peter Seidel, Christian Theobalt |
ACM Trans. Graph. | 2 |
| 2011 | High-quality shape from multi-view stereo and shading under general illuminationabstractMulti-view stereo methods reconstruct 3D geometry from images well for sufficiently textured scenes, but often fail to recover high-frequency surface detail, particularly for smoothly shaded surfaces. On the other hand, shape-from-shading methods can recover fine detail from shading variations. Unfortunately, it is non-trivial to apply shape-from-shading alone to multi-view data, and most shading-based estimation methods only succeed under very restricted or controlled illumination. We present a new algorithm that combines multi-view stereo and shading-based refinement for high-quality reconstruction of 3D geometry models from images taken under constant but otherwise arbitrary illumination. We have tested our algorithm on several scenes that were captured under several general and unknown lighting conditions, and we show that our final reconstructions rival laser range scans. Chenglei Wu, Bennett Wilburn, Yasuyuki Matsushita, Christian Theobalt |
CVPR | 1 |
| 2011 | Shading-based dynamic shape refinement from multi-view video under general illuminationabstractWe present an approach to add true fine-scale spatio-temporal shape detail to dynamic scene geometry captured from multi-view video footage. Our approach exploits shading information to recover the millimeter-scale surface structure, but in contrast to related approaches succeeds under general unconstrained lighting conditions. Our method starts off from a set of multi-view video frames and an initial series of reconstructed coarse 3D meshes that lack any surface detail. In a spatio-temporal maximum a posteriori probability (MAP) inference framework, our approach first estimates the incident illumination and the spatially-varying albedo map on the mesh surface for every time instant. Thereafter, albedo and illumination are used to estimate the true geometric detail visible in the images and add it to the coarse reconstructions. The MAP framework uses weak temporal priors on lighting, albedo and geometry which improve reconstruction quality yet allow for temporal variations in the data. Chenglei Wu, Kiran Varanasi, Yebin Liu, Hans-Peter Seidel, Christian Theobalt |
ICCV | 1 |
| 2011 | Fusing Multiview and Photometric Stereo for 3D Reconstruction under Uncalibrated IlluminationabstractWe propose a method to obtain a complete and accurate 3D model from multiview images captured under a variety of unknown illuminations. Based on recent results showing that for Lambertian objects, general illumination can be approximated well using low-order spherical harmonics, we develop a robust alternating approach to recover surface normals. Surface normals are initialized using a multi-illumination multiview stereo algorithm, then refined using a robust alternating optimization method based on the l(1) metric. Erroneous normal estimates are detected using a shape prior. Finally, the computed normals are used to improve the preliminary 3D model. The reconstruction system achieves watertight and robust 3D reconstruction while neither requiring manual interactions nor imposing any constraints on the illumination. Experimental results on both real world and synthetic data show that the technique can acquire accurate 3D models for Lambertian surfaces, and even tolerates small violations of the Lambertian assumption. Chenglei Wu, Yebin Liu, Qionghai Dai, Bennett Wilburn |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2009 | Multi-view reconstruction under varying illumination conditionsabstractThis paper addresses the problem of complete and detailed 3D model reconstruction of objects filmed by multiple cameras under varying illumination. Firstly, initial normal maps are obtained to enhance the correspondence mapping. Then, the depth for every pixel is estimated by combining photometric constraint with occlusion robust photo-consistency. Finally, after filtering the point cloud, a Poisson surface reconstruction is applied to obtain a watertight mesh. In contrast with traditional photometric stereo techniques, the proposed algorithm does not directly calculate the photometric normal but integrates the photometric constraint into the depth estimation. Furthermore, different from classic multi-view stereo(MVS), we consider the counterpart under changing light conditions. The algorithm has been implemented based on our multi-camera and multi-light acquisition system. We validate the method by complete reconstruction of challenging real objects and show experimentally that this technique can greatly improve on correspondence-based MVS results. Chenglei Wu, Yebin Liu, Xiangyang Ji, Qionghai Dai |
ICME | 1 |