VLDB 2026 Research / reviewers in the wild / expert
Lingyu Wei
dblp:165/9987
· DBLP profile ↗
9ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0001-7278-4228ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
7 papers |
Visual content generation and editing · 26% Geometric modeling and processing · 17% Virtual and augmented reality · 17% | |
| Artificial intelligence
7 papers |
Generative modeling · 52% 3D vision · 44% Face, body and person analysis · 4% |
Topics — the 20 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d face reconstruction |
0.8 | 2 | 2021 | Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement · CVPR 2021 Photorealistic Facial Texture Inference Using Deep Neural Networks · CVPR 2017 |
Machine learning › Generative modeling
generative adversarial network |
0.6 | 2 | 2021 | Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement · CVPR 2021 Real-Time Hair Rendering Using Sequential Adversarial Networks · ECCV (4) 2018 |
Machine learning › Generative modeling
avatar generation |
0.5 | 1 | 2021 | Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement · CVPR 2021 |
Machine learning › Generative modeling › face synthesis
StyleGAN-based face generation |
0.5 | 1 | 2021 | Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement · CVPR 2021 |
Visual content generation and editing › face editing
facial expression editing |
0.4 | 1 | 2019 | Deep face normalization · ACM Trans. Graph. 2019 |
Rendering › appearance modeling
hair rendering |
0.3 | 1 | 2018 | Real-Time Hair Rendering Using Sequential Adversarial Networks · ECCV (4) 2018 |
Geometric modeling and processing › shape modeling
3d hair modeling |
0.3 | 1 | 2017 | Avatar digitization from a single image for real-time rendering · ACM Trans. Graph. 2017 |
Rendering
appearance modeling |
0.3 | 1 | 2017 | Photorealistic Facial Texture Inference Using Deep Neural Networks · CVPR 2017 |
Visual content generation and editing
texture synthesis |
0.3 | 1 | 2017 | Photorealistic Facial Texture Inference Using Deep Neural Networks · CVPR 2017 |
Computer vision › 3D vision
shape matching |
0.2 | 1 | 2016 | Dense Human Body Correspondences Using Convolutional Networks · CVPR 2016 |
Computer animation and physical simulation
facial animation |
0.2 | 1 | 2015 | Facial performance sensing head-mounted display · ACM Trans. Graph. 2015 |
Computer animation and physical simulation › performance capture
facial performance capture |
0.2 | 1 | 2015 | Facial performance sensing head-mounted display · ACM Trans. Graph. 2015 |
Virtual and augmented reality › immersive display
head-mounted display |
0.2 | 1 | 2015 | Facial performance sensing head-mounted display · ACM Trans. Graph. 2015 |
Computer vision › 3D vision › 3d face modeling
3d morphable model |
0.1 | 1 | 2021 | Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement · CVPR 2021 |
Computer vision › Face, body and person analysis
face recognition |
0.1 | 1 | 2019 | Deep face normalization · ACM Trans. Graph. 2019 |
Machine learning › Generative modeling › generative adversarial network
conditional GAN |
0.1 | 1 | 2018 | paGAN: real-time avatars using dynamic textures · ACM Trans. Graph. 2018 |
Computational photography and imaging › image-based modeling
3d reconstruction from images |
0.1 | 1 | 2017 | Photorealistic Facial Texture Inference Using Deep Neural Networks · CVPR 2017 |
Virtual and augmented reality › avatar rendering
real-time avatar rendering |
0.1 | 1 | 2017 | Avatar digitization from a single image for real-time rendering · ACM Trans. Graph. 2017 |
Geometric modeling and processing
surface reconstruction |
0.1 | 1 | 2016 | Capturing Dynamic Textured Surfaces of Moving Targets · ECCV (7) 2016 |
Wearable and physiological sensing
strain sensing |
0.1 | 1 | 2015 | Facial performance sensing head-mounted display · ACM Trans. Graph. 2015 |
Methods — techniques the papers use, named apart from their topics
regression network · 0.8generative adversarial network · 0.8dense flow field prediction · 0.8sequential adversarial networks · 0.7conditional generative adversarial network · 0.7blendshape blending · 0.7UV texture mapping · 0.7deep convolutional neural network · 0.6photogrammetry · 0.5perceptual refinement · 0.5StyleGAN2 · 0.5feature correlation · 0.3convex optimization · 0.3gaussian mixture model · 0.2flexible electronic materials · 0.2RGB-D camera · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Normalized Avatar Synthesis Using StyleGAN and Perceptual RefinementabstractWe introduce a highly robust GAN-based framework for digitizing a normalized 3D avatar of a person from a single unconstrained photo. While the input image can be of a smiling person or taken in extreme lighting conditions, our method can reliably produce a high-quality textured model of a person’s face in neutral expression and skin textures under diffuse lighting condition. Cutting-edge 3D face reconstruction methods use non-linear morphable face models combined with GAN-based decoders to capture the likeness and details of a person but fail to produce neutral head models with unshaded albedo textures which is critical for creating relightable and animation-friendly avatars for integration in virtual environments. The key challenges for existing methods to work is the lack of training and ground truth data containing normalized 3D faces. We propose a two-stage approach to address this problem. First, we adopt a highly robust normalized 3D face generator by embedding a non-linear morphable face model into a StyleGAN2 network. This allows us to generate detailed but normalized facial assets. This inference is then followed by a perceptual refinement step that uses the generated assets as regularization to cope with the limited available training samples of normalized faces. We further introduce a Normalized Face Dataset, which consists of a combination photogrammetry scans, carefully selected photographs, and generated fake people with neutral expressions in diffuse lighting conditions. While our prepared dataset contains two orders of magnitude less subjects than cutting edge GAN-based 3D facial reconstruction methods, we show that it is possible to produce high-quality normalized face models for very challenging unconstrained input images, and demonstrate superior performance to the current state-of-the-art. Huiwen Luo, Koki Nagano, Han-Wei Kung, Qingguo Xu, Zejian Wang, Lingyu Wei, Liwen Hu 0001, Hao Li 0015 |
CVPR | 6 |
| 2019 | Deep face normalizationabstractFrom angling smiles to duck faces, all kinds of facial expressions can be seen in selfies, portraits, and Internet pictures. These photos are taken from various camera types, and under a vast range of angles and lighting conditions. We present a deep learning framework that can fully normalize unconstrained face images, i.e., remove perspective distortions, relight to an evenly lit environment, and predict a frontal and neutral face. Our method can produce a high resolution image while preserving important facial details and the likeness of the subject, along with the original background. We divide this ill-posed problem into three consecutive normalization steps, each using a different generative adversarial network that acts as an image generator. Perspective distortion removal is performed using a dense flow field predictor. A uniformly illuminated face is obtained using a lighting translation network, and the facial expression is neutralized using a generalized facial expression synthesis framework combined with a regression network based on deep features for facial recognition. We introduce new data representations for conditional inference, as well as training methods for supervised learning to ensure that different expressions of the same person can yield to not only a plausible but also a similar neutral face. We demonstrate our results on a wide range of challenging images collected in the wild. Key applications of our method range from robust image-based 3D avatar creation, portrait manipulation, to facial enhancement and reconstruction tasks for crime investigation. We also found through an extensive user study, that our normalization results can be hardly distinguished from ground truth ones if the person is not familiar. Koki Nagano, Huiwen Luo, Zejian Wang, Jaewoo Seo, Jun Xing, Liwen Hu 0001, Lingyu Wei, Hao Li 0015 |
ACM Trans. Graph. | 7 |
| 2018 | Real-Time Hair Rendering Using Sequential Adversarial Networks
Lingyu Wei, Liwen Hu 0001, Vladimir G. Kim, Ersin Yumer, Hao Li 0015 |
ECCV (4) | 1 |
| 2018 | paGAN: real-time avatars using dynamic texturesabstractWith the rising interest in personalized VR and gaming experiences comes the need to create high quality 3D avatars that are both low-cost and variegated. Due to this, building dynamic avatars from a single unconstrained input image is becoming a popular application. While previous techniques that attempt this require multiple input images or rely on transferring dynamic facial appearance from a source actor, we are able to do so using only one 2D input image without any form of transfer from a source image. We achieve this using a new conditional Generative Adversarial Network design that allows fine-scale manipulation of any facial input image into a new expression while preserving its identity. Our photoreal avatar GAN (paGAN) can also synthesize the unseen mouth interior and control the eye-gaze direction of the output, as well as produce the final image from a novel viewpoint. The method is even capable of generating fully-controllable temporally stable video sequences, despite not using temporal information during training. After training, we can use our network to produce dynamic image-based avatars that are controllable on mobile devices in real time. To do this, we compute a fixed set of output images that correspond to key blendshapes, from which we extract textures in UV space. Using a subject's expression blendshapes at run-time, we can linearly blend these key textures together to achieve the desired appearance. Furthermore, we can use the mouth interior and eye textures produced by our network to synthesize on-the-fly avatar animations for those regions. Our work produces state-of-the-art quality image and video synthesis, and is the first to our knowledge that is able to generate a dynamically textured avatar with a mouth interior, all from a single image. Koki Nagano, Jaewoo Seo, Jun Xing, Lingyu Wei, Zimo Li, Shunsuke Saito, Aviral Agarwal, Jens Fursund, Hao Li 0015 |
ACM Trans. Graph. | 4 |
| 2017 | Photorealistic Facial Texture Inference Using Deep Neural NetworksabstractWe present a data-driven inference method that can synthesize a photorealistic texture map of a complete 3D face model given a partial 2D view of a person in the wild. After an initial estimation of shape and low-frequency albedo, we compute a high-frequency partial texture map, without the shading component, of the visible face area. To extract the fine appearance details from this incomplete input, we introduce a multi-scale detail analysis technique based on mid-layer feature correlations extracted from a deep convolutional neural network. We demonstrate that fitting a convex combination of feature correlations from a high-resolution face database can yield a semantically plausible facial detail description of the entire face. A complete and photorealistic texture map can then be synthesized by iteratively optimizing for the reconstructed feature correlations. Using these high-resolution textures and a commercial rendering framework, we can produce high-fidelity 3D renderings that are visually comparable to those obtained with state-of-the-art multi-view face capture systems. We demonstrate successful face reconstructions from a wide range of low resolution input images, including those of historical figures. In addition to extensive evaluations, we validate the realism of our results using a crowdsourced user study. Shunsuke Saito, Lingyu Wei, Liwen Hu 0001, Koki Nagano, Hao Li 0015 |
CVPR | 2 |
| 2017 | Avatar digitization from a single image for real-time renderingabstractWe present a fully automatic framework that digitizes a complete 3D head with hair from a single unconstrained image. Our system offers a practical and consumer-friendly end-to-end solution for avatar personalization in gaming and social VR applications. The reconstructed models include secondary components (eyes, teeth, tongue, and gums) and provide animation-friendly blendshapes and joint-based rigs. While the generated face is a high-quality textured mesh, we propose a versatile and efficient polygonal strips (polystrips) representation for the hair. Polystrips are suitable for an extremely wide range of hairstyles and textures and are compatible with existing game engines for real-time rendering. In addition to integrating state-of-the-art advances in facial shape modeling and appearance inference, we propose a novel single-view hair generation pipeline, based on 3D-model and texture retrieval, shape refinement, and polystrip patching optimization. The performance of our hairstyle retrieval is enhanced using a deep convolutional neural network for semantic hair attribute classification. Our generated models are visually comparable to state-of-the-art game characters designed by professional artists. For real-time settings, we demonstrate the flexibility of polystrips in handling hairstyle variations, as opposed to conventional strand-based representations. We further show the effectiveness of our approach on a large number of images taken in the wild, and how compelling avatars can be easily created by anyone. Liwen Hu 0001, Shunsuke Saito, Lingyu Wei, Koki Nagano, Jaewoo Seo, Jens Fursund, Iman Sadeghi, Carrie Sun, Hao Li 0015 |
ACM Trans. Graph. | 3 |
| 2016 | Dense Human Body Correspondences Using Convolutional NetworksabstractWe propose a deep learning approach for finding dense correspondences between 3D scans of people. Our method requires only partial geometric information in the form of two depth maps or partial reconstructed surfaces, works for humans in arbitrary poses and wearing any clothing, does not require the two people to be scanned from similar view-points, and runs in real time. We use a deep convolutional neural network to train a feature descriptor on depth map pixels, but crucially, rather than training the network to solve the shape correspondence problem directly, we train it to solve a body region classification problem, modified to increase the smoothness of the learned descriptors near region boundaries. This approach ensures that nearby points on the human body are nearby in feature space, and vice versa, rendering the feature descriptor suitable for computing dense correspondences between the scans. We validate our method on real and synthetic data for both clothed and unclothed humans, and show that our correspondences are more robust than is possible with state-of-the-art unsupervised methods, and more accurate than those found using methods that require full watertight 3D geometry. Lingyu Wei, Qixing Huang, Duygu Ceylan, Etienne Vouga, Hao Li 0015 |
CVPR | 1 |
| 2016 | Capturing Dynamic Textured Surfaces of Moving Targets
Ruizhe Wang 0002, Lingyu Wei, Etienne Vouga, Qixing Huang, Duygu Ceylan, Gérard G. Medioni, Hao Li 0015 |
ECCV (7) | 2 |
| 2015 | Facial performance sensing head-mounted displayabstractThere are currently no solutions for enabling direct face-to-face interaction between virtual reality (VR) users wearing head-mounted displays (HMDs). The main challenge is that the headset obstructs a significant portion of a user's face, preventing effective facial capture with traditional techniques. To advance virtual reality as a next-generation communication platform, we develop a novel HMD that enables 3D facial performance-driven animation in real-time. Our wearable system uses ultra-thin flexible electronic materials that are mounted on the foam liner of the headset to measure surface strain signals corresponding to upper face expressions. These strain signals are combined with a head-mounted RGB-D camera to enhance the tracking in the mouth region and to account for inaccurate HMD placement. To map the input signals to a 3D face model, we perform a single-instance offline training session for each person. For reusable and accurate online operation, we propose a short calibration step to readjust the Gaussian mixture distribution of the mapping before each use. The resulting animations are visually on par with cutting-edge depth sensor-driven facial performance capture systems and hence, are suitable for social interactions in virtual worlds. Hao Li 0015, Laura C. Trutoiu, Kyle Olszewski, Lingyu Wei, Tristan Trutna, Pei-Lun Hsieh, Aaron Nicholls, Chongyang Ma |
ACM Trans. Graph. | 4 |