Yanlin Weng

dblp:03/3634 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0001-5223-4253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 8 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 RGBAvatar: Reduced Gaussian Blendshapes for Online Modeling of Head Avatars
abstract
We present Reduced Gaussian Blendshapes Avatar (RGBAvatar), a method for reconstructing photorealistic, animatable head avatars at speeds sufficient for on-the-fly reconstruction. Unlike prior approaches that utilize linear bases from 3D morphable models (3DMM) to model Gaussian blendshapes, our method maps tracked 3DMM parameters into reduced blendshape weights with an MLP, leading to a compact set of blendshape bases. The learned compact base composition effectively captures essential facial details for specific individuals, and does not rely on the fixed base composition weights of 3DMM, leading to enhanced reconstruction quality and higher efficiency. To further expedite the reconstruction process, we develop a novel color initialization estimation method and a batch-parallel Gaussian rasterization process, achieving state-of-the-art quality with training throughput of about 630 images per second. Moreover, we propose a local-global sampling strategy that enables direct on-the-fly reconstruction, immediately reconstructing the model as video streams in real time while achieving quality comparable to offline settings. Our source code is available at https://github.com/gapszju/RGBAvatar.
Linzhou Li, Yanlin Weng, Youyi Zheng, Kun Zhou 0001
CVPR3
2025 Relightable Gaussian blendshapes for head avatar animation
Shengjie Ma, Youyi Zheng, Yanlin Weng, Kun Zhou 0001
Comput. Graph.3
2025 Animatable 3D Gaussians for modeling dynamic humans
Yukun Xu, Keyang Ye, Tianjia Shao, Yanlin Weng
Frontiers Comput. Sci.4
2025 A Semantic Talking Style Space for Speech-Driven Facial Animation
abstract
We present a latent talking style space with semantic meanings for speech-driven 3D facial animation. The style space is learned from 3D speech facial animations via a self-supervision paradigm without any style labeling, leading to an automatic separation of high-level attributes, i.e., different channels of the latent style code possess different semantic meanings, such as a wide/slightly open mouth, a grinning/round mouth, and frowning/raising eyebrows. The style space enables intuitive and flexible control of talking styles in speech-driven facial animation through manipulating the channels of style code. To effectively learn such a style space, we propose a two-stage approach, involving two deep neural networks, to disentangle the person identity, speech content, and talking style contained in 3D speech facial animations. The training is performed on a novel dataset of 3D talking faces of various styles, constructed from over ten hours of videos of 200 subjects collected from the Internet.
Yujin Chai, Yanlin Weng, Tianjia Shao, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.2
2025 Single-View 3D Hair Modeling With Clumping Optimization
abstract
Deep learning advancements have enabled the generation of visually plausible hair geometry from a single image, but the results still do not meet the realism required for further applications (e.g., high quality hair rendering and simulation). One of the essential element that is missing in previous single-view hair reconstruction methods is the clumping effect of hair, which is influenced by scalp secretions and oils, and is a key ingredient for high-quality hair rendering and simulation. Inspired by common practices in industrial production which simulates realistic hair clumping by allowing artists to adjust clumping parameters, we aim to integrate these clumping effects into single-view hair reconstruction. We introduce a hierarchical hair representation that incorporates a clumping modifier into the guide hair and skinning-based hair expressions. This representation utilizes guide strands and skinning weights to express the basic geometric structure of the hair. The clumping modifier allows for the expression of more detailed and realistic clumping effects. Based on this representation, We design a fully differentiable framework integrating a neural measurement of clumping and a line-based rasterization renderer to iteratively solve guide strands positions and clumping parameters. Our method demonstrates superior performance both qualitatively and quantitatively compared to state-of-the-art techniques.
Zhongsi Tang, Jiahao Geng, Yanlin Weng, Youyi Zheng, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.3
2025 As-Rigid-As-Possible Deformation of Gaussian Radiance Fields
abstract
3D Gaussian Splatting (3DGS) models radiance fields as sparsely distributed 3D Gaussians, providing a compelling solution to novel view synthesis at high resolutions and real-time frame rates. However, deforming objects represented by 3D Gaussians remains a challenging task. Existing methods deform a 3DGS object by editing Gaussians geometrically. These approaches ignore the fact that it is the radiance field that rasterizes and renders the final image. The inconsistency between the deformed 3D Gaussians and the desired radiance field inevitably leads to artifacts in the final results. In this paper, we propose an interactive method for as-rigid-as-possible (ARAP) deformation of the Gaussian radiance fields. Specifically, after performing geometric edits on the Gaussians, we further optimize Gaussians to ensure its rasterization yields a similar result as the deformed radiance field. To facilitate this objective, we design radial features to mathematically describe the radial difference before and after the deformation, which are densely sampled across the radiance field. Additionally, we propose an adaptive anisotropic spatial low-pass filter to prevent aliasing issues during sampling and to preserve the field with the varying non-uniform sampling intervals. Users can interactively employ this tool to achieve large-scale ARAP deformations of the radiance field. Since our method maintains the consistency of the Gaussian radiance field before and after deformation, it avoids artifacts that are common in existing 3DGS deformation frameworks. Meanwhile, our method keeps the high quality and efficiency of 3DGS in rendering.
Xinhao Tong, Tianjia Shao, Yanlin Weng, Yin Yang 0002, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Personalized Audio-Driven 3D Facial Animation via Style-Content Disentanglement
abstract
We present a learning-based approach for generating 3D facial animations with the motion style of a specific subject from arbitrary audio inputs. The subject style is learned from a video clip (1-2 minutes) either downloaded from the Internet or captured through an ordinary camera. Traditional methods often require many hours of the subject's video to learn a robust audio-driven model and are thus unsuitable for this task. Recent research efforts aim to train a model from video collections of a few subjects but ignore the discrimination between the subject style and underlying speech content within facial motions, leading to inaccurate style or articulation. To solve the problem, we propose a novel framework that disentangles subject-specific style and speech content from facial motions. The disentanglement is enabled by two novel training mechanisms. One is two-pass style swapping between two random subjects, and the other is joint training of the decomposition network and audio-to-motion network with a shared decoder. After training, the disentangled style is combined with arbitrary audio inputs to generate stylized audio-driven 3D facial animations. Compared with start-of-the-art methods, our approach achieves better results qualitatively and quantitatively, especially in difficult cases like bilabial plosive and bilabial nasal phonemes.
Yujin Chai, Tianjia Shao, Yanlin Weng, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.3
2023 Unsupervised image translation with distributional semantics awareness
abstract
Unsupervised image translation (UIT) studies the mapping between two image domains. Since such mappings are under-constrained, existing research has pursued various desirable properties such as distributional matching or two-way consistency. In this paper, we re-examine UIT from a new perspective: distributional semantics consistency, based on the observation that data variations contain semantics, e.g., shoes varying in colors. Further, the semantics can be multi-dimensional, e.g., shoes also varying in style, functionality, etc. Given two image domains, matching these semantic dimensions during UIT will produce mappings with explicable correspondences, which has not been investigated previously. We propose distributional semantics mapping (DSM), the first UIT method which explicitly matches semantics between two domains. We show that distributional semantics has been rarely considered within and beyond UIT, even though it is a common problem in deep learning. We evaluate DSM on several benchmark datasets, demonstrating its general ability to capture distributional semantics. Extensive comparisons show that DSM not only produces explicable mappings, but also improves image quality in general.
Zhexi Peng, He Wang 0002, Yanlin Weng, Yin Yang 0002, Tianjia Shao
Comput. Vis. Media3
2022 Speech-driven facial animation with spectral gathering and temporal attention
Yujin Chai, Yanlin Weng, Lvdi Wang, Kun Zhou 0001
Frontiers Comput. Sci.2
2021 Single-view facial reflectance inference with a differentiable renderer
Jiahao Geng, Yanlin Weng, Lvdi Wang, Kun Zhou 0001
Sci. China Inf. Sci.2
2020 H-CNN: Spatial Hashing Based CNN for 3D Shape Analysis
abstract
We present a novel spatial hashing based data structure to facilitate 3D shape analysis using convolutional neural networks (CNNs). Our method builds hierarchical hash tables for an input model under different resolutions that leverage the sparse occupancy of 3D shape boundary. Based on this data structure, we design two efficient GPU algorithms namely hash2col and col2hash so that the CNN operations like convolution and pooling can be efficiently parallelized. The perfect spatial hashing is employed as our spatial hashing scheme, which is not only free of hash collision but also nearly minimal so that our data structure is almost of the same size as the raw input. Compared with existing 3D CNN methods, our data structure significantly reduces the memory footprint during the CNN training. As the input geometry features are more compactly packed, CNN operations also run faster with our data structure. The experiment shows that, under the same network structure, our method yields comparable or better benchmark results compared with the state-of-the-art while it has only one-third memory consumption when under high resolutions (i.e., 2563).
Tianjia Shao, Yin Yang 0002, Yanlin Weng, Qiming Hou, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.3
2018 Automatic Mechanism Modeling from a Single Image with CNNs
abstract
Abstract This paper presents a novel system that enables a fully automatic modeling of both 3D geometry and functionality of a mechanism assembly from a single RGB image. The resulting 3D mechanism model highly resembles the one in the input image with the geometry, mechanical attributes, connectivity, and functionality of all the mechanical parts prescribed in a physically valid way. This challenging task is realized by combining various deep convolutional neural networks to provide high‐quality and automatic part detection, segmentation, camera pose estimation and mechanical attributes retrieval for each individual part component. On the top of this, we use a local/global optimization algorithm to establish geometric interdependencies among all the parts while retaining their desired spatial arrangement. We use an interaction graph to abstract the inter‐part connection in the resulting mechanism system. If an isolated component is identified in the graph, our system enumerates all the possible solutions to restore the graph connectivity, and outputs the one with the smallest residual error. We have extensively tested our system with a wide range of classic mechanism photos, and experimental results show that the proposed system is able to build high‐quality 3D mechanism models without user guidance.
Minmin Lin, Tianjia Shao, Youyi Zheng, Zhong Ren 0001, Yanlin Weng, Yin Yang 0002
Comput. Graph. Forum5
2018 Warp-guided GANs for single-photo facial animation
abstract
This paper introduces a novel method for realtime portrait animation in a single photo. Our method requires only a single portrait photo and a set of facial landmarks derived from a driving source (e.g., a photo or a video sequence), and generates an animated image with rich facial details. The core of our method is a warp-guided generative model that instantly fuses various fine facial details (e.g., creases and wrinkles), which are necessary to generate a high-fidelity facial expression, onto a pre-warped image. Our method factorizes out the nonlinear geometric transformations exhibited in facial expressions by lightweight 2D warps and leaves the appearance detail synthesis to conditional generative neural networks for high-fidelity facial animation generation. We show such a factorization of geometric transformation and appearance synthesis largely helps the network better learn the high nonlinearity of the facial expression functions and also facilitates the design of the network architecture. Through extensive experiments on various portrait photos from the Internet, we show the significant efficacy of our method compared with prior arts.
Jiahao Geng, Tianjia Shao, Youyi Zheng, Yanlin Weng, Kun Zhou 0001
ACM Trans. Graph.4
2018 Modeling hair from an RGB-D camera
abstract
Creating realistic 3D hairs that closely match the real-world inputs remains challenging. With the increasing popularity of lightweight depth cameras featured in devices such as iPhone X, Intel RealSense and DJI drones, depth cues can be very helpful in consumer applications, for example, the Animated Emoji. In this paper, we introduce a fully automatic, data-driven approach to model the hair geometry and compute a complete strand-level 3D hair model that closely resembles the input from a single RGB-D camera. Our method heavily exploits the geometric cues contained in the depth channel and leverages exemplars in a 3D hair database for high-fidelity hair synthesis. The core of our method is a local-similarity based search and synthesis algorithm that simultaneously reasons about the hair geometry, strands connectivity, strand orientation, and hair structural plausibility. We demonstrate the efficacy of our method using a variety of complex hairstyles and compare our method with prior arts.
Hongzhi Wu, Yanlin Weng, Youyi Zheng, Kun Zhou 0001
ACM Trans. Graph.4
2016 Real-time facial animation with image-based dynamic avatars
abstract
We present a novel image-based representation for dynamic 3D avatars, which allows effective handling of various hairstyles and headwear, and can generate expressive facial animations with fine-scale details in real-time. We develop algorithms for creating an image-based avatar from a set of sparsely captured images of a user, using an off-the-shelf web camera at home. An optimization method is proposed to construct a topologically consistent morphable model that approximates the dynamic hair geometry in the captured images. We also design a real-time algorithm for synthesizing novel views of an image-based avatar, so that the avatar follows the facial motions of an arbitrary actor. Compelling results from our pipeline are demonstrated on a variety of cases.
Hongzhi Wu, Yanlin Weng, Tianjia Shao, Kun Zhou 0001
ACM Trans. Graph.3
2016 AutoHair: fully automatic hair modeling from a single image
abstract
We introduceAutoHair, the first fully automatic method for 3D hair modeling from a single portrait image, with no user interaction or parameter tuning. Our method efficiently generates complete and high-quality hair geometries, which are comparable to those generated by the state-of-the-art methods, where user interaction is required. The core components of our method are: a novel hierarchical deep neural network for automatic hair segmentation and hair growth direction estimation, trained over an annotated hair image database; and an efficient and automatic data-driven hair matching and modeling algorithm, based on a large set of 3D hair exemplars. We demonstrate the efficacy and robustness of our method on Internet photos, resulting in a database of around 50K 3D hair models and a corresponding hairstyle space that covers a wide variety of real-world hairstyles. We also show novel applications enabled by our method, including 3D hairstyle space navigation and hair-aware image retrieval.
Menglei Chai, Tianjia Shao, Hongzhi Wu, Yanlin Weng, Kun Zhou 0001
ACM Trans. Graph.4
2014 Real-time facial animation on mobile devices
Yanlin Weng, Qiming Hou, Kun Zhou 0001
Graph. Model.1
2014 FaceWarehouse: A 3D Facial Expression Database for Visual Computing
abstract
We present FaceWarehouse, a database of 3D facial expressions for visual computing applications. We use Kinect, an off-the-shelf RGBD camera, to capture 150 individuals aged 7-80 from various ethnic backgrounds. For each person, we captured the RGBD data of her different expressions, including the neutral expression and 19 other expressions such as mouth-opening, smile, kiss, etc. For every RGBD raw data record, a set of facial feature points on the color image such as eye corners, mouth contour, and the nose tip are automatically localized, and manually adjusted if better accuracy is required. We then deform a template facial mesh to fit the depth data as closely as possible while matching the feature points on the color image to their corresponding points on the mesh. Starting from these fitted face meshes, we construct a set of individual-specific expression blendshapes for each person. These meshes with consistent topology are assembled as a rank-3 tensor to build a bilinear face model with two attributes: identity and expression. Compared with previous 3D facial databases, for every person in our database, there is a much richer matching collection of expressions, enabling depiction of most human facial actions. We demonstrate the potential of FaceWarehouse for visual computing with four applications: facial image manipulation, face component transfer, real-time performance-based facial image animation, and facial animation retargeting from video to image.
Yanlin Weng, Yiying Tong, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.2
2013 As-Rigid-AsPossible Distance Field Metamorphosis
abstract
Abstract Widely used for morphing between objects with arbitrary topology, distance field interpolation (DFI) handles topological transition naturally without the need for correspondence or remeshing, unlike surface‐based interpolation approaches. However, lack of correspondence in DFI also leads to ineffective control over the morphing process. In particular, unless the user specifies a dense set of landmarks, it is not even possible to measure the distortion of intermediate shapes during interpolation, let alone control it. To remedy such issues, we introduce an approach for establishing correspondence between the interior of two arbitrary objects, formulated as an optimal mass transport problem with a sparse set of landmarks. This correspondence enables us to compute non‐rigid warping functions that better align the source and target objects as well as to incorporate local rigidity constraints to perform as‐rigid‐aspossible DFI. We demonstrate how our approach helps achieve flexible morphing results with a small number of landmarks.
Yanlin Weng, Menglei Chai, Weiwei Xu 0003, Yiying Tong, Kun Zhou 0001
Comput. Graph. Forum1
2013 Hair Interpolation for Portrait Morphing
abstract
Abstract In this paper we study the problem of hair interpolation: given two 3D hair models, we want to generate a sequence of intermediate hair models that transform from one input to another both smoothly and aesthetically pleasing. We propose an automatic method that efficiently calculates a many‐to‐many strand correspondence between two or more given hair models, taking into account the multi‐scale clustering structure of hair. Experiments demonstrate that hair interpolation can be used for producing more vivid portrait morphing effects and enabling a novel example‐based hair styling methodology, where a user can interactively create new hairstyles by continuously exploring a “style space” spanning multiple input hair models.
Yanlin Weng, Lvdi Wang, Menglei Chai, Kun Zhou 0001
Comput. Graph. Forum1
2013 Compact combinatorial maps: A volume mesh data structure
Yuanzhen Wang, Yanlin Weng, Yiying Tong
Graph. Model.3
2013 Orientation Field Guided Texture Synthesis
Yanlin Weng, Jian-Nan Wang, Yiying Tong
J. Comput. Sci. Technol.2
2013 3D shape regression for real-time facial animation
abstract
We present a real-time performance-driven facial animation system based on 3D shape regression. In this system, the 3D positions of facial landmark points are inferred by a regressor from 2D video frames of an ordinary web camera. From these 3D points, the pose and expressions of the face are recovered by fitting a user-specific blendshape model to them. The main technical contribution of this work is the 3D regression algorithm that learns an accurate, user-specific face alignment model from an easily acquired set of training data, generated from images of the user performing a sequence of predefined facial poses and expressions. Experiments show that our system can accurately recover 3D face shapes even for fast motions, non-frontal faces, and exaggerated expressions. In addition, some capacity to handle partial occlusions and changing lighting conditions is demonstrated.
Yanlin Weng, Stephen Lin 0001, Kun Zhou 0001
ACM Trans. Graph.2
2013 Dynamic hair manipulation in images and videos
abstract
This paper presents a single-view hair modeling technique for generating visually and physically plausible 3D hair models with modest user interaction. By solving an unambiguous 3D vector field explicitly from the image and adopting an iterative hair generation algorithm, we can create hair models that not only visually match the original input very well but also possess physical plausibility (e.g., having strand roots fixed on the scalp and preserving the length and continuity of real strands in the image as much as possible). The latter property enables us to manipulate hair in many new ways that were previously very difficult with a single image, such as dynamic simulation or interactive hair shape editing. We further extend the modeling approach to handle simple video input, and generate dynamic 3D hair models. This allows users to manipulate hair in a video or transfer styles from images to videos.
Menglei Chai, Lvdi Wang, Yanlin Weng, Xiaogang Jin 0001, Kun Zhou 0001
ACM Trans. Graph.3
2013 Texture mapping subdivision surfaces with hard constraints
Yanlin Weng, Dongping Li, Yiying Tong
Vis. Comput.1
2012 Compact Combinatorial Maps in 3D
Yuanzhen Wang, Yanlin Weng, Yiying Tong
CVM3
2012 Constrained Texture Mapping on Subdivision Surfaces
Yanlin Weng, Dongping Li, Yiying Tong
CVM1
2012 Single-view hair modeling for portrait manipulation
abstract
Human hair is known to be very difficult to model or reconstruct. In this paper, we focus on applications related to portrait manipulation and take an application-driven approach to hair modeling. To enable an average user to achieve interesting portrait manipulation results, we develop a single-view hair modeling technique with modest user interaction to meet the unique requirements set by portrait manipulation. Our method relies on heuristics to generate a plausible high-resolution strand-based 3D hair model. This is made possible by an effective high-precision 2D strand tracing algorithm, which explicitly models uncertainty and local layering during tracing. The depth of the traced strands is solved through an optimization, which simultaneously considers depth constraints, layering constraints as well as regularization terms. Our single-view hair modeling enables a number of interesting applications that were previously challenging, including transferring the hairstyle of one subject to another in a potentially different pose, rendering the original portrait in a novel view and image-space hair editing.
Menglei Chai, Lvdi Wang, Yanlin Weng, Yizhou Yu, Baining Guo, Kun Zhou 0001
ACM Trans. Graph.3
2011 New Technique: Sketch-based rotation editing
abstract
We present a sketch-based rotation editing system for enriching rotational motion in keyframe animations. Given a set of keyframe orientations of a rigid object, the user first edits its angular velocity trajectory by sketching curves, and then the system computes the altered rotational motion by solving a variational curve fitting problem. The solved rotational motion not only satisfies the orientation constraints at the keyframes, but also fits well the user-specified angular velocity trajectory. Our system is simple and easy to use. We demonstrate its usefulness by adding interesting and realistic rotational details to several keyframe animations.
Weiwei Xu 0003, Yizhou Yu, Yanlin Weng
J. Zhejiang Univ. Sci. C4
2008 Keyframe Based Video Object Deformation
abstract
This paper presents a novel deformation system for video objects. The system is designed to minimize the amount of user interaction, while providing flexible and precise user control. It has a keyframe-based user interface. The user only needs to manipulate the video object at several keyframes. Our algorithm smoothly propagate the editing result from the keyframes to the remaining frames and automatically generate the new video object. The algorithm is able to preserve the temporal coherence as well as the shape features of the video object in the original video clips. We demonstrate the potential of our system with a variety of examples.
Yanlin Weng, Shichao Hu, Baining Guo
CW1
2008 Sketching MLS Image Deformations On the GPU
abstract
Abstract In this paper, we present an image editing tool that allows the user to deform images using a sketch‐based interface. The user simply sketches a set of source curves in the input image, and also some target curves that the source curves should be deformed to. Then the moving least squares (MLS) deformation technique [ SMW06 ] is adapted to produce realistic deformations while satisfying the curves' positional constraints. We also propose a scheme to reduce image fold‐overs in MLS deformations. Our system has a very intuitive user interface, generates physically plausible deformations, and can be easily implemented on the GPU for real‐time performance.
Yanlin Weng, Hujun Bao
Comput. Graph. Forum1
2006 2D shape deformation using nonlinear least squares optimization
Yanlin Weng, Weiwei Xu 0003, Yanchen Wu, Kun Zhou 0001, Baining Guo
Vis. Comput.1