EDBT 2026 Demo / reviewers in the wild / expert
Juyong Zhang
dblp:38/5125
· DBLP profile ↗
116ranked-venue papers
11as first author
66since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 100 · 10 first-author · 55 since 2021Artificial intelligence and machine learning · 40 · 2 first-author · 25 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GSwap: Realistic Head Swapping With Dynamic Neural Gaussian FieldabstractWe present GSwap, a novel consistent and realistic video head-swapping system empowered by dynamic neural Gaussian portrait priors, which significantly advances the state of the art in face and head replacement. Unlike previous methods that rely primarily on 2D generative models or 3D Morphable Face Models (3DMM), our approach overcomes their inherent limitations, including poor 3D consistency, unnatural facial expressions, and restricted synthesis quality. Moreover, existing techniques struggle with full head-swapping tasks due to insufficient holistic head modeling and ineffective background blending, often resulting in visible artifacts and misalignments. To address these challenges, GSwap introduces an intrinsic 3D Gaussian feature field embedded within a full-body SMPL-X surface, effectively elevating 2D portrait videos into a dynamic neural Gaussian field. This innovation ensures high-fidelity, 3D-consistent portrait rendering while preserving natural head-torso relationships and seamless motion dynamics. To facilitate training, we adapt a pretrained 2D portrait generative model to the source head domain using only a few reference images, enabling efficient domain adaptation. Furthermore, we propose a neural re-rendering strategy that harmoniously integrates the synthesized foreground with the original background, eliminating blending artifacts and enhancing realism. Extensive experiments demonstrate that GSwap surpasses existing methods in multiple aspects, including visual quality, temporal coherence, identity preservation, and 3D consistency. Xuan Gao 0003, Dongyu Liu, Junhui Hou, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | Computational Caustic Design for Surface Light SourceabstractDesigning freeform surfaces to control light based on real-world illumination patterns is challenging, as existing caustic lens designs often assume oversimplified point or parallel light sources. We propose representing surface light sources using an optimized set of point sources, whose parameters are fitted to the real light source's illumination using a novel differentiable rendering framework. Our physically-based rendering approach simulates light transmission using flux, without requiring prior knowledge of the light source's intensity distribution. To efficiently explore the light source parameter space during optimization, we apply a contraction mapping that converts the constrained problem into an unconstrained one. Using the optimized light source model, we then design the freeform lens shape considering flux consistency and normal integrability. Simulations and physical experiments show our method more accurately represents real surface light sources compared to point-source approximations, yielding caustic lenses that produce images closely matching the target light distributions. Sizhuo Zhou, Yuou Sun, Bailin Deng, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | FGO-SLAM++: Real-Time Geometry-Aware Gaussian SLAM With Continuous Opacity FieldabstractWe present FGO-SLAM++, a real-time geometry-aware Gaussian SLAM system capable of processing video streams from arbitrary modalities. The system performs multi-view consistent map reconstruction by maintaining a continuous opacity field. Existing Gaussian SLAM systems face challenges in achieving efficient tracking and mapping while supporting multiple input modalities, which often necessitates a trade-off between rendering quality and geometric accuracy. This study demonstrates that it is possible to satisfy all these requirements simultaneously using video streams of arbitrary modalities. The core of the proposed method involves explicit geometric feature extraction to capture the underlying scene structure and estimate camera poses, followed by a Gaussian-based ray tracing strategy to construct and optimize an opacity field. Upon the detection of loop closures, the system performs global adjustment to enhance map consistency. Furthermore, the surface is directly extracted using the marching tetrahedra method and refined through a geometric constraint field. Extensive experiments demonstrate that the proposed method achieves superior performance in tracking accuracy, rendering quality, geometric reconstruction, and real-time efficiency. Peichen Liu, Hui Zhu 0010, Chunmao Jiang, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Oblique-MERF: Revisiting and Improving MERF for Oblique PhotographyabstractNeural radiance fields (NeRF) have established a new paradigm for 3D scene reconstruction, with subsequent work achieving high-quality real-time rendering. However, reconstructing large-scale scenes from oblique aerial photography presents unique challenges, such as varying spatial scale distributions and a constrained range of tilt angles, often resulting in high memory consumption and reduced rendering quality at extrapolated viewpoints. To address these issues, we propose a novel approach named Oblique-MERF to accommodate the distinctive characteristics of oblique photography datasets and support real-time rendering on various common devices. Firstly, an innovative adaptive occupancy plane is proposed to constrain the sampling space. Additionally, we propose a smoothness regularization loss for view-dependent color to enhance the MLP's ability to generalize to untrained viewpoints. Experimental results demonstrate that Oblique-MERF reduces VRAM usage by approximately 40% while maintaining competitive rendering quality compared to baseline methods, and achieves higher frame rates with more realistic rendering even at untrained extrapolated viewpoints. Project page: https://ustc3dv.github.ioIOblique-MERFI Xiaoyi Zeng, Kaiwen Song, Leyuan Yang, Bailin Deng, Juyong Zhang |
3DV | 5 |
| 2025 | HERA: Hybrid Explicit Representation for Ultra-Realistic Head AvatarsabstractWe introduce a novel approach to creating ultra-realistic head avatars and rendering them in real time (≥ 30 fps at 2048 × 1334 resolution). First, we propose a hybrid explicit representation that combines the advantages of two primitive based efficient rendering techniques. UV-mapped 3D mesh is utilized to capture sharp and rich textures on smooth surfaces, while 3D Gaussian Splatting is employed to represent complex geometric structures. In the pipeline of modeling an avatar, after tracking parametric models based on captured multi-view RGB videos, our goal is to simultaneously optimize the texture and opacity map of mesh, as well as a set of 3D Gaussian splats localized and rigged onto the mesh facets. Specifically, we perform α-blending on the color and opacity values based on the merged and reordered z-buffer from the rasterization results of mesh and 3DGS. This process involves the mesh and 3DGS adaptively fitting the captured visual information to outline a high-fidelity digital avatar. To avoid artifacts caused by Gaussian splats crossing the mesh facets, we design a stable hybrid depth sorting strategy. Experiments illustrate that our modeled results exceed those of state-of-the-art approaches. Hongrui Cai, Xuan Wang 0009, Jiafei Li, Yanbo Fan, Shenghua Gao, Juyong Zhang |
CVPR | 8 |
| 2025 | D^3-Human: Dynamic Disentangled Digital Human from Monocular VideoabstractWe introduce D3-Human, a method for reconstructing Dynamic Disentangled Digital Human geometry from monocular videos. Past monocular video human reconstruction primarily focuses on reconstructing undecoupled clothed human bodies or only reconstructing clothing, making it difficult to apply directly in applications such as animation production. The challenge in reconstructing decoupled clothing and body lies in the occlusion caused by clothing over the body. To this end, the details of the visible area and the plausibility of the invisible area must be ensured during the reconstruction process. Our proposed method combines explicit and implicit representations to model the decoupled clothed human body, leveraging the robustness of explicit representations and the flexibility of implicit representations. Specifically, we reconstruct the visible region as SDF and propose a novel human manifold signed distance field (hmSDF) to segment the visible clothing and visible body, and then merge the visible and invisible body. Extensive experimental results demonstrate that, compared with existing reconstruction schemes, D3-Human can achieve high-quality decoupled reconstruction of the human body wearing different clothing, and can be directly applied to clothing transfer and animation production. Code is available at https://ustc3dv.github.io/D3Human/. Honghu Chen, Yunfan Tao, Juyong Zhang |
CVPR | 4 |
| 2025 | Expressive Talking Human from Single-Image with Imperfect Priors
Leipeng Hu, Boyang Guo, Yancheng Yuan, Juyong Zhang |
ICCV | 6 |
| 2025 | Constructing Diffusion Avatar with Learnable EmbeddingsabstractRecent advances in diffusion models have made significant progress in digital human generation. However, most existing models still struggle to maintain 3D consistency, temporal coherence, and motion accuracy. These limitations primarily stem from two key factors: the limited representation ability of commonly used control signals (e.g., landmarks, depth maps), and the lack of diversity in identity and pose variations within publicly available datasets. In this paper, we construct a powerful head model from both aspects by constructing learnable control signals and enabling the model to adaptively leverage synthetic data. Firstly, we introduce a novel control signal representation that is learnable, dense, expressive, and 3D consistent. Our method embeds learnable Gaussians onto a parametric head surface, which significantly enhances the consistency and expressiveness of diffusion-based head models. Secondly, in terms of data, we synthesize a large-scale dataset covering diverse poses and identities. To reduce the negative impact of artifacts in synthetic data, we introduce real/synthetic embeddings that allow the model to distinguish between real and synthetic samples and learn to utilize them adaptively. Extensive experiments show that our model outperforms existing methods in terms of realism, expressiveness, and 3D consistency. Our code, synthetic datasets, and pre-trained models will be released at https://ustc3dv.github.io/Learn2Control. Xuan Gao 0003, Dongyu Liu, Yuqi Zhou 0004, Juyong Zhang |
SIGGRAPH Asia | 5 |
| 2025 | Foreword to chinagraph 2024 special section
Song-Hai Zhang, Juyong Zhang |
Comput. Graph. | 3 |
| 2025 | Joint Deblurring and 3D Reconstruction for MacrophotographyabstractAbstract Macro lens has the advantages of high resolution and large magnification, and 3D modeling of small and detailed objects can provide richer information. However, defocus blur in macrophotography is a long‐standing problem that heavily hinders the clear imaging of the captured objects and high‐quality 3D reconstruction of them. Traditional image deblurring methods require a large number of images and annotations, and there is currently no multi‐view 3D reconstruction method for macrophotography. In this work, we propose a joint deblurring and 3D reconstruction method for macrophotography. Starting from multi‐view blurry images captured, we jointly optimize the clear 3D model of the object and the defocus blur kernel of each pixel. The entire framework adopts a differentiable rendering method to self‐supervise the optimization of the 3D model and the defocus blur kernel. Extensive experiments show that from a small number of multi‐view images, our proposed method can not only achieve high‐quality image deblurring but also recover high‐fidelity 3D appearance. Liangchen Li, Yuqi Zhou 0004, Kai Wang 0012, Juyong Zhang |
Comput. Graph. Forum | 6 |
| 2025 | SPARE: Symmetrized Point-to-Plane Distance for Robust Non-Rigid 3D RegistrationabstractExisting optimization-based methods for non-rigid registration typically minimize an alignment error metric based on the point-to-point or point-to-plane distance between corresponding point pairs on the source surface and target surface. However, these metrics can result in slow convergence or a loss of detail. In this paper, we propose SPARE, a novel formulation that utilizes a symmetrized point-to-plane distance for robust non-rigid registration. The symmetrized point-to-plane distance relies on both the positions and normals of the corresponding points, resulting in a more accurate approximation of the underlying geometry and can achieve higher accuracy than existing methods. To solve this optimization problem efficiently, we introduce an as-rigid-as-possible regulation term to estimate the deformed normals and propose an alternating minimization solver using a majorization-minimization strategy. Moreover, for effective initialization of the solver, we incorporate a deformation graph-based coarse alignment that improves registration quality and efficiency. Extensive experiments show that the proposed method greatly improves the accuracy of non-rigid registration problems and maintains relatively high solution efficiency. Yuxin Yao 0001, Bailin Deng, Junhui Hou, Juyong Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | End-to-end Surface Optimization for Light ControlabstractDesigning a freeform surface to reflect or refract light to achieve a target distribution is a challenging inverse problem. In this article, we propose an end-to-end optimization strategy for an optical surface mesh. Our formulation leverages a novel differentiable rendering model, and is directly driven by the difference between the resulting light distribution and the target distribution. We also enforce geometric constraints related to fabrication requirements, to facilitate CNC milling and polishing of the designed surface. To address the issue of local minima, we formulate a face-based optimal transport problem between the current mesh and the target distribution, which makes effective large changes to the surface shape. The combination of our optimal transport update and rendering-guided optimization produces an optical surface design with a resulting image closely resembling the target, while the geometric constraints in our optimization help to ensure consistency between the rendering model and the final physical results. The effectiveness of our algorithm is demonstrated on a variety of target images using both simulated rendering and physical prototypes. Yuou Sun, Bailin Deng, Juyong Zhang |
ACM Trans. Graph. | 3 |
| 2025 | Double Reference Guided Interactive 2D and 3D Caricature GenerationabstractIn this article, we propose the first geometry and texture (double) referenced interactive two-dimensional (2D) and 3D caricature generating and editing method. The main challenge of caricature generation lies in the fact that it not only exaggerates the facial geometry but also refreshes the facial texture. We address this challenge by utilizing the semantic segmentation maps as an intermediary domain, removing the influence of photo texture while preserving the person-specific geometry features. Specifically, our proposed method consists of two main components: 3D-CariNet and CariMaskGAN. 3D-CariNet uses sketches or caricatures to exaggerate the input photo into several types of 3D caricatures. To generate a CariMask, we geometrically exaggerate the photos using the projection of exaggerated 3D landmarks, after which CariMask is converted into a caricature by CariMaskGAN. In this step, users can edit and adjust the geometry of caricatures freely. Moreover, we propose a semantic detail preprocessing approach that considerably increases the details of generated caricatures and allows modification of hair strands, wrinkles, and beards. By rendering high-quality 2D caricatures as textures, we produce 3D caricatures with a variety of texture styles. Extensive experimental results have demonstrated that our method can produce higher-quality caricatures as well as support interactive modification with ease. Hongrui Cai, Juyong Zhang, Feng Tian 0006, Jinyuan Jia 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Neural-ABC: Neural Parametric Models for Articulated Body With ClothesabstractIn this article, we introduce Neural-ABC, a novel parametric model based on neural implicit functions that can represent clothed human bodies with disentangled latent spaces for identity, clothing, shape, and pose. Traditional mesh-based representations struggle to represent articulated bodies with clothes due to the diversity of human body shapes and clothing styles, as well as the complexity of poses. Our proposed model provides a unified framework for parametric modeling, which can represent the identity, clothing, shape and pose of the clothed human body. Our proposed approach utilizes the power of neural implicit functions as the underlying representation and integrates well-designed structures to meet the necessary requirements. Specifically, we represent the underlying body as a signed distance function and clothing as an unsigned distance function, and they can be uniformly represented as unsigned distance fields. Different types of clothing do not require predefined topological structures or classifications, and can follow changes in the underlying body to fit the body. Additionally, we construct poses using a controllable articulated structure. The model is trained on both open and newly constructed datasets, and our decoupling strategy is carefully designed to ensure optimal performance. Our model excels at disentangling clothing and identity in different shape and poses while preserving the style of the clothing. We demonstrate that Neural-ABC fits new observations of different types of clothing. Compared to other state-of-the-art parametric models, Neural-ABC demonstrates powerful advantages in the reconstruction of clothed human bodies, as evidenced by fitting raw scans, depth maps and images. We show that the attributes of the fitted results can be further edited by adjusting their identities, clothing, shape and pose codes. Honghu Chen, Yuxin Yao 0001, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | RGAvatar: Relightable 4D Gaussian Avatar From Monocular VideosabstractRelightable 4D avatar reconstruction which enables high fidelity and real-time rendering continues to be a crucial but challenging problem, especially from monocular videos. Previous NeRF-based 4D avatars enable photo-realistic relighting but are too slow for rendering, while point-based or mesh-based 4D avatars are efficient but have limited rendering quality. The recent success of 3D Gaussian Splatting, i.e., 3DGS, has inspired a series of impressive 4D Gaussian avatars, however, most of which only focus on faithful appearance reconstruction but are not relightable. To address such issues, this article proposes a new Relightable 4D Gaussian Avatar, i.e., RGAvatar, tailored for high fidelity relightable rendering from monocular videos. Our key idea is to introduce a new relightable 4D Gaussian representation, based on which we can directly perform high fidelity Physically Based Rendering, and an effective joint learning mechanism for compact 4D Gaussian reconstruction with SDF regulation and accurate materials and lighting decomposition. By comparing with previous state-of-the-art approaches, RGAvatar can significantly outperform previous approaches in relightable rendering quality and speed. To our best knowledge, RGAvatar contributes a new state-of-the-art 4D Gaussian avatar from monocular videos, which enables high fidelity relightable rendering in a quite efficient manner. Zhe Fan, Shi-Sheng Huang, Dachao Shang, Juyong Zhang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | PICA: Physics-Integrated Clothed AvatarabstractWe introduce PICA, a novel representation for high-fidelity animatable clothed human avatars with physics-plausible dynamics, even for loose clothing. Previous neural rendering-based representations of animatable clothed humans typically employ a single model to represent both the clothing and the underlying body. While efficient, these approaches often fail to represent complex garment dynamics, leading to incorrect deformations and noticeable rendering artifacts, especially for sliding or loose garments. Furthermore, most previous works represent garment dynamics as pose-dependent deformations and facilitate novel pose animations in a data-driven manner. This often results in outcomes that do not faithfully represent the mechanics of motion and are prone to generating artifacts in out-of-distribution poses. To address these issues, we employ two individual 2D Gaussian Splatting (2DGS) models with different deformation characteristics, modeling the human body and clothing separately. This distinction allows for better handling of their respective motion characteristics. With this representation, we integrate a graph neural network (GNN)-based clothing physics simulation module to ensure a better representation of clothing dynamics. Our method, through its carefully designed features, achieves high-fidelity rendering of clothed human bodies in complex and novel driving poses, outperforming previous methods under the same settings. Bo Peng 0020, Yunfan Tao, Haoyu Zhan, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | EGAvatar: Efficient GAN Inversion for Generalizable Head Avatar From Few-Shot ImagesabstractControllable head avatar reconstruction via the inversion of few-shot images using 3D generative models has demonstrated significant potential for efficient avatar creation. However, under limited input conditions, existing one-shot inversion methods often fail to produce high-fidelity results, frequently leading to shape distortions, expression deviations, and identity inconsistencies. To address these limitations, we propose EGAvatar, a novel and efficient 3DGAN inversion framework designed to generate high-fidelity, generalizable head avatars from few-shot images. The core principle of EGAvatar is a decoupling-by-inverting strategy, built upon an animatable 3DGAN prior. Specifically, we introduce an effective animatable 3DGAN model that synthesizes high-quality 3D avatars by integrating a coarse 3D triplane representation (derived from a latent 3DGAN) with an offset 3D triplane (learned via a triplane 3DGAN). Leveraging this architecture, we design a 3DGAN-based inversion approach to reconstruct 3D avatars efficiently. Additionally, we incorporate an expression-view disentanglement mechanism to maintain consistent appearance across varying expressions and viewpoints, thereby enhancing the generalizability of avatar reconstruction from limited input images. Extensive experiments conducted on two publicly available benchmarks and a private dataset demonstrate that EGAvatar outperforms existing state-of-the-art methods in both qualitative and quantitative evaluations. Notably, EGAvatar achieves superior performance while requiring significantly fewer input images and offering more efficient training and inference. Hao Pan Ren, Wan Yu Li, Shi-Sheng Huang, Juyong Zhang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Scalable and High-Quality Neural Implicit Representation for 3D ReconstructionabstractVarious SDF-based neural implicit surface reconstruction methods have been proposed recently, and have demonstrated remarkable modeling capabilities. However, due to the global nature and limited representation ability of a single network, existing methods still suffer from many drawbacks, such as limited accuracy and scale of the reconstruction. In this paper, we propose a versatile, scalable and high-quality neural implicit representation to address these issues. We integrate a divide-and-conquer approach into the neural SDF-based reconstruction. Specifically, we model the object or scene as a fusion of multiple independent local neural SDFs with overlapping regions. The construction of our representation involves three key steps: (1) constructing the distribution and overlap relationship of the local radiance fields based on object structure or data distribution, (2) relative pose registration for adjacent local SDFs, and (3) SDF blending. Thanks to the independent representation of each local region, our approach can not only achieve high-fidelity surface reconstruction, but also enable scalable scene reconstruction. Extensive experimental results demonstrate the effectiveness and practicality of our proposed method. Leyuan Yang, Bailin Deng, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Zero-Shot Real Facial Attribute Separation and Transfer at Novel Views
Dingyun Zhang, Heyuan Li, Juyong Zhang |
CVM (2) | 3 |
| 2024 | L0-Sampler: An L0Model Guided Volume Sampling for NeRFabstractSince its proposal, Neural Radiance Fields (NeRF) has achieved great success in related tasks, mainly adopting the hierarchical volume sampling (HVS) strategy for volume rendering. However, the HVS of NeRF approximates distributions using piecewise constant functions, which provides a relatively rough estimation. Based on the obser-vation that a well-trained weight function$w(t)$and the$L_{0}$distance between points and the surface have very high sim-ilarity, we propose$L_{0}$-Sampler by incorporating the$L_{0}$model into$w(t)$to guide the sampling process. Specif-ically, we propose using piecewise exponential functions rather than piecewise constant functions for interpolation, which can not only approximate quasi-$L_{0}$weight distri-butions along rays quite well but can be easily imple-mented with a few lines of code change without additional computational burden. Stable performance improvements can be achieved by applying$L_{0}$-Sampler to NeRF and re-lated tasks like 3D reconstruction. Code is available at https://ustc3dv.github.io/L0-Sampler/. Liangchen Li, Juyong Zhang |
CVPR | 2 |
| 2024 | FlashAvatar: High-Fidelity Head Avatar with Efficient Gaussian EmbeddingabstractWe propose FlashAvatar, a novel and lightweight 3D animatable avatar representation that could reconstruct a digital avatar from a short monocular video sequence in minutes and render high-fidelity photo-realistic images at 300FPS on a consumer-grade GPU. To achieve this, we maintain a uniform 3D Gaussian field embedded in the surface of a parametric face model and learn extra spatial offset to model non-surface regions and subtle facial details. While full use of geometric priors can capture high-frequency facial details and preserve exaggerated expressions, proper initialization can help reduce the number of Gaussians, thus enabling super-fast rendering speed. Extensive experimental results demonstrate that FlashAvatar outperforms existing works regarding visual quality and personalized details and is almost an order of magnitude faster in rendering speed. Project page: https://ustc3dv.github.io/FlashAvatar/ Xuan Gao 0003, Juyong Zhang |
CVPR | 4 |
| 2024 | City-on-Web: Real-Time Neural Rendering of Large-Scale Scenes on the Web
Kaiwen Song, Xiaoyi Zeng, Chenqu Ren, Juyong Zhang |
ECCV (47) | 4 |
| 2024 | DynoSurf: Neural Deformation-Based Temporally Consistent Dynamic Surface Reconstruction
Yuxin Yao 0001, Junhui Hou, Juyong Zhang, Wenping Wang 0001 |
ECCV (33) | 5 |
| 2024 | Deformable NeRF using Recursively Subdivided Tetrahedra
Zherui Qiu, Chenqu Ren, Kaiwen Song, Xiaoyi Zeng, Leyuan Yang, Juyong Zhang |
ACM Multimedia | 6 |
| 2024 | Portrait Video Editing Empowered by Multimodal Generative Priors
Xuan Gao 0003, Haiyao Xiao, Chenglai Zhong, Shimin Hu 0004, Juyong Zhang |
SIGGRAPH Asia | 6 |
| 2024 | iShapEditing: Intelligent Shape Editing with Diffusion ModelsabstractAbstract Recent advancements in generative models have enabled image editing very effective with impressive results. By extending this progress to 3D geometry models, we introduce iShapEditing, a novel framework for 3D shape editing which is applicable to both generated and real shapes. Users manipulate shapes by dragging handle points to corresponding targets, offering an intuitive and intelligent editing interface. Leveraging the Triplane Diffusion model and robust intermediate feature correspondence, our framework utilizes classifier guidance to adjust noise representations during sampling process, ensuring alignment with user expectations while preserving plausibility. For real shapes, we employ shape predictions at each time step alongside a DDPM‐based inversion algorithm to derive their latent codes, facilitating seamless editing. iShapEditing provides effective and intelligent control over shapes without the need for additional model training or fine‐tuning. Experimental examples demonstrate the effectiveness and superiority of our method in terms of editing accuracy and plausibility. Jing Li 0113, Juyong Zhang, Falai Chen |
Comput. Graph. Forum | 2 |
| 2024 | Multi-scale hash encoding based neural geometry representationabstractRecently, neural implicit function-based representation has attracted more and more attention, and has been widely used to represent surfaces using differentiable neural networks. However, surface reconstruction from point clouds or multi-view images using existing neural geometry representations still suffer from slow computation and poor accuracy. To alleviate these issues, we propose a multi-scale hash encoding-based neural geometry representation which effectively and efficiently represents the surface as a signed distance field. Our novel neural network structure carefully combines low-frequency Fourier position encoding with multi-scale hash encoding. The initialization of the geometry network and geometry features of the rendering module are accordingly redesigned. Our experiments demonstrate that the proposed representation is at least 10 times faster for reconstructing point clouds with millions of points. It also significantly improves speed and accuracy of multi-view reconstruction. Our code and models are available at https://github.com/Dengzhi-USTC/Neural-Geometry-Reconstruction . Haoyao Xiao, Yining Lang, Juyong Zhang |
Comput. Vis. Media | 5 |
| 2024 | A Closer Look at the Reflection Formulation in Single Image Reflection RemovalabstractHow to model the effect of reflection is crucial for single image reflection removal (SIRR) task. Modern SIRR methods usually simplify the reflection formulation with the assumption of linear combination of a transmission layer and a reflection layer. However, the large variations in image content and the real-world picture-taking conditions often result in far more complex reflection. In this paper, we introduce a new screen-blur combination based on two important factors, namely the intensity and the blurriness of reflection, to better characterize the reflection formulation in SIRR. Specifically, we present Screen-blur Reflection Networks (SRNet), which executes the screen-blur formulation in its network design and adapts to the complex reflection on real scenes. Technically, SRNet consists of three components: a blended image generator, a reflection estimator and a reflection removal module. The image generator exploits the screen-blur combination to synthesize the training blended images. The reflection estimator learns the reflection layer and a blur degree that measures the level of blurriness for reflection. The reflection removal module further uses the blended image, blur degree and reflection layer to filter out the transmission layer in a cascaded manner. Superior results on three different SIRR methods are reported when generating the training data on the principle of the screen-blur combination. Moreover, extensive experiments on six datasets quantitatively and qualitatively demonstrate the efficacy of SRNet over the state-of-the-art methods. Fuchen Long, Zhaofan Qiu, Juyong Zhang, Zhengjun Zha, Ting Yao 0003, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | IntrinsicNGP: Intrinsic Coordinate Based Hash Encoding for Human NeRFabstractRecently, many works have been proposed to use the neural radiance field for novel view synthesis of human performers. However, most of these methods require hours of training, making them difficult for practical use. To address this challenging problem, we propose IntrinsicNGP, which can be trained from scratch and achieve high-fidelity results in a few minutes with videos of a human performer. To achieve this goal, we introduce a continuous and optimizable intrinsic coordinate instead of the original explicit euclidean coordinate in the hash encoding module of InstantNGP. With this novel intrinsic coordinate, IntrinsicNGP can aggregate interframe information for dynamic objects using proxy geometry shapes. Moreover, the results trained with the given rough geometry shapes can be further refined with an optimizable offset field based on the intrinsic coordinate. Extensive experimental results on several datasets demonstrate the effectiveness and efficiency of IntrinsicNGP. We also illustrate the ability of our approach to edit the shape of reconstructed objects. Bo Peng 0020, Xuan Gao 0003, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Facial optical flow estimation via neural non-rigid registrationabstractOptical flow estimation in human facial video, which provides 2D correspondences between adjacent frames, is a fundamental pre-processing step for many applications, like facial expression capture and recognition. However, it is quite challenging as human facial images contain large areas of similar textures, rich expressions, and large rotations. These characteristics also result in the scarcity of large, annotated real-world datasets. We propose a robust and accurate method to learn facial optical flow in a self-supervised manner. Specifically, we utilize various shape priors, including face depth, landmarks, and parsing, to guide the self-supervised learning task via a differentiable nonrigid registration framework. Extensive experiments demonstrate that our method achieves remarkable improvements for facial optical flow estimation in the presence of significant expressions and large rotations. Zhuang Peng, Boyi Jiang, Haofei Xu, Wanquan Feng, Juyong Zhang |
Comput. Vis. Media | 5 |
| 2023 | Fast and Robust Non-Rigid Registration Using Accelerated Majorization-MinimizationabstractNon-rigid 3D registration, which deforms a source 3D shape in a non-rigid way to align with a target 3D shape, is a classical problem in computer vision. Such problems can be challenging because of imperfect data (noise, outliers and partial overlap) and high degrees of freedom. Existing methods typically adopt the$\ell _{p}$type robust norm to measure the alignment error and regularize the smoothness of deformation, and use a proximal algorithm to solve the resulting non-smooth optimization problem. However, the slow convergence of such algorithms limits their wide applications. In this paper, we propose a formulation for robust non-rigid registration based on a globally smooth robust norm for alignment and regularization, which can effectively handle outliers and partial overlaps. The problem is solved using the majorization-minimization algorithm, which reduces each iteration to a convex quadratic problem with a closed-form solution. We further apply Anderson acceleration to speed up the convergence of the solver, enabling the solver to run efficiently on devices with limited compute capability. Extensive experiments demonstrate the effectiveness of our method for non-rigid alignment between two shapes with outliers and partial overlaps, with quantitative evaluation showing that it outperforms state-of-the-art methods in terms of registration accuracy and computational speed. The source code is available athttps://github.com/yaoyx689/AMM_NRR. Yuxin Yao 0001, Bailin Deng, Weiwei Xu 0003, Juyong Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Flattening-Net: Deep Regular 2D Representation for 3D Point Cloud AnalysisabstractPoint clouds are characterized by irregularity and unstructuredness, which pose challenges in efficient data exploitation and discriminative feature extraction. In this paper, we present an unsupervised deep neural architecture called Flattening-Net to represent irregular 3D point clouds of arbitrary geometry and topology as a completely regular 2D point geometry image (PGI) structure, in which coordinates of spatial points are captured in colors of image pixels. Intuitively, Flattening-Net implicitly approximates a locally smooth 3D-to-2D surface flattening process while effectively preserving neighborhood consistency. As a generic representation modality, PGI inherently encodes the intrinsic property of the underlying manifold structure and facilitates surface-style point feature aggregation. To demonstrate its potential, we construct a unified learning framework directly operating on PGIs to achieve diverse types of high-level and low-level downstream applications driven by specific task networks, including classification, segmentation, reconstruction, and upsampling. Extensive experiments demonstrate that our methods perform favorably against the current state-of-the-art competitors. Qijian Zhang, Junhui Hou, Yiming Zeng 0002, Juyong Zhang, Ying He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Contextualized Graph Attention Network for Recommendation With Item Knowledge GraphabstractGraph neural networks (GNN) have recently been applied to exploit knowledge graph (KG) for recommendation. Existing GNN-based methods explicitly model the dependency between an entity and its local graph context in KG (i.e., the set of its first-order neighbors), but may not be effective in capturing its non-local graph context (i.e., the set of most related high-order neighbors). In this paper, we propose a novel recommendation framework, named Contextualized Graph Attention Network (CGAT), which can explicitly exploit both local and non-local graph context information of an entity in KG. More specifically, CGAT captures the local context information by a user-specific graph attention mechanism, considering a user's personalized preferences on entities. In addition, CGAT employs a biased random walk sampling process to extract the non-local context of an entity, and utilizes a Recurrent Neural Network (RNN) to model the dependency between the entity and its non-local contextual entities. To capture the user's personalized preferences on items, an item-specific attention mechanism is also developed to model the dependency between a target item and the contextual items extracted from the user's historical behaviors. We compared CGAT with state-of-the-art KG-based recommendation methods on real datasets, and the experimental results demonstrate the effectiveness of CGAT. Yong Liu 0020, Susen Yang, Chunyan Miao, Min Wu 0008, Juyong Zhang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Learning Hierarchical Review Graph Representations for RecommendationabstractThe user review data have been demonstrated to be effective in solving different recommendation problems. Previous review-based recommendation methods usually employ sophisticated compositional models, such as Recurrent Neural Networks (RNN) and Convolutional Neural Networks (CNN), to learn semantic representations from the review data for recommendation. However, these methods mainly capture the local dependency between neighboring words in a word window, and they treat each review equally. Therefore, they may not be effective in capturing the global dependency between words and tend to be easily biased by noise review information. In this paper, we propose a novel review-based recommendation model, named Review Graph Neural Network (RGNN). Specifically, RGNN builds a specific review graph for each individual user/item, which provides a global view about the user/item properties to help weaken the biases caused by noise review information. A type-aware graph attention mechanism is developed to learn semantic embeddings of words. Moreover, a personalized graph pooling operator is proposed to learn hierarchical representations of the review graph to form the semantic representation for each user/item. We compared RGNN with state-of-the-art review-based recommendation approaches on two real-world datasets. The experimental results indicate that RGNN consistently outperforms baseline methods, in terms of Mean Square Error (MSE). Yong Liu 0020, Susen Yang, Yinan Zhang 0002, Chunyan Miao, Zaiqing Nie, Juyong Zhang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Audio-Driven Talking Face Video Generation With Dynamic Convolution KernelsabstractIn this paper, we present a dynamic convolution kernel (DCK) strategy for convolutional neural networks. Using a fully convolutional network with the proposed DCKs, high-quality talking-face video can be generated from multi-modal sources (i.e., unmatched audio and video) in real time, and our trained model is robust to different identities, head postures, and input audios. Our proposed DCKs are specially designed for audio-driven talking face video generation, leading to a simple yet effective end-to-end system. We also provide a theoretical analysis to interpret why DCKs work. Experimental results show that our method can generate high-quality talking-face video with background at 60 fps. Comparison and evaluation between our method and the state-of-the-art methods demonstrate the superiority of our method. Zipeng Ye, Mengfei Xia, Ran Yi 0002, Juyong Zhang, Yukun Lai, Xuwei Huang, Guo-Xin Zhang, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Predicting Personalized Head Movement From Short Video and Speech SignalabstractAudio-driven talking face video generation has attracted much attention recently. However, few existing works pay attention to machine learning of talking head movement, especially based on the phonetic study. Observing that real-world talking faces often accompany natural head movement, in this paper, we model the relation between speech signal and talking head movement, which is a typical one-to-many mapping problem. To solve this problem, we propose a novel two-step mapping strategy: (1) in the first step, we train an encoder that predicts a head motion behavior pattern (modeled as a feature vector) from the head motion sequence of a short video of 10–15 seconds, and (2) in the second step, we train a decoder that predict a unique head motion sequence from both the motion behavior pattern and the auditory features of an arbitrary speech signal. Based on the proposed mapping strategy, we build a deep neural network model that takes a speech signal of a source person and a short video of a target person as input, and outputs a synthesized high-fidelity talking face video with personalized head pose. Extensive experiments and a user study show that our method can generate high-quality personalized head movement in synthesized talking face videos, and meanwhile, has comparable facial animation quality (e.g., lip synchronization and expression) with the state-of-the-art methods. Ran Yi 0002, Zipeng Ye, Zhiyao Sun, Juyong Zhang, Guo-Xin Zhang, Pengfei Wan 0001, Hujun Bao, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Robust Mesh Representation Learning via Efficient Local Structure-Aware Anisotropic ConvolutionabstractMesh is a type of data structure commonly used for 3-D shapes. Representation learning for 3-D meshes is essential in many computer vision and graphics applications. The recent success of convolutional neural networks (CNNs) for structured data (e.g., images) suggests the value of adapting insights from CNN for 3-D shapes. However, 3-D shape data are irregular since each node's neighbors are unordered. Various graph neural networks for 3-D shapes have been developed with isotropic filters or predefined local coordinate systems to overcome the node inconsistency on graphs. However, isotropic filters or predefined local coordinate systems limit the representation power. In this article, we propose a local structure-aware anisotropic convolutional operation (LSA-Conv) that learns adaptive weighting matrices for each template's node according to its neighboring structure and performs shared anisotropic filters. In fact, the learnable weighting matrix is similar to the attention matrix in the random synthesizer-a new Transformer model for natural language processing (NLP). Since the learnable weighting matrices require large amounts of parameters for high-resolution 3-D shapes, we introduce a matrix factorization technique to notably reduce the parameter size, denoted as LSA-small. Furthermore, a residual connection with a linear transformation is introduced to improve the performance of our LSA-Conv. Comprehensive experiments demonstrate that our model produces significant improvement in 3-D shape reconstruction compared to state-of-the-art methods. Zhongpai Gao, Junchi Yan, Guangtao Zhai, Juyong Zhang, Xiaokang Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Sketch2PQ: Freeform Planar Quadrilateral Mesh Design via a Single SketchabstractThe freeform architectural modeling process often involves two important stages: concept design and digital modeling. In the first stage, architects usually sketch the overall 3D shape and the panel layout on a physical or digital paper briefly. In the second stage, a digital 3D model is created using the sketch as a reference. The digital model needs to incorporate geometric requirements for its components, such as the planarity of panels due to consideration of construction costs, which can make the modeling process more challenging. In this work, we present a novel sketch-based system to bridge the concept design and digital modeling of freeform roof-like shapes represented as planar quadrilateral (PQ) meshes. Our system allows the user to sketch the surface boundary and contour lines under axonometric projection and supports the sketching of occluded regions. In addition, the user can sketch feature lines to provide directional guidance to the PQ mesh layout. Given the 2D sketch input, we propose a deep neural network to infer in real-time the underlying surface shape along with a dense conjugate direction field, both of which are used to extract the final PQ mesh. To train and validate our network, we generate a large synthetic dataset that mimics architect sketching of freeform quadrilateral patches. The effectiveness and usability of our system are demonstrated with quantitative and qualitative evaluation as well as user studies. Yang Liu 0014, Hao Pan 0001, Wassim Jabi, Juyong Zhang, Bailin Deng |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | 3D-CariGAN: An End-to-End Solution to 3D Caricature Generation From Normal Face PhotosabstractCaricature is a type of artistic style of human faces that attracts considerable attention in the entertainment industry. So far a few 3D caricature generation methods exist and all of them require some caricature information (e.g., a caricature sketch or 2D caricature) as input. This kind of input, however, is difficult to provide by non-professional users. In this paper, we propose an end-to-end deep neural network model that generates high-quality 3D caricatures directly from a normal 2D face photo. The most challenging issue for our system is that the source domain of face photos (characterized by normal 2D faces) is significantly different from the target domain of 3D caricatures (characterized by 3D exaggerated face shapes and textures). To address this challenge, we: (1) build a large dataset of 5,343 3D caricature meshes and use it to establish a PCA model in the 3D caricature shape space; (2) reconstruct a normal full 3D head from the input face photo and use its PCA representation in the 3D caricature shape space to establish correspondences between the input photo and 3D caricature shape; and (3) propose a novel character loss and a novel caricature loss based on previous psychological studies on caricatures. Experiments including a novel two-level user study show that our system can generate high-quality 3D caricatures directly from normal face photos. Zipeng Ye, Mengfei Xia, Yanan Sun 0006, Ran Yi 0002, Minjing Yu, Juyong Zhang, Yukun Lai, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | HeadNeRF: A Realtime NeRF-based Parametric Head ModelabstractIn this paper, we propose HeadNeRF, a novel NeRF-based parametric head model that integrates the neural radiance field to the parametric representation of the human head. It can render high fidelity head images in real-time on modern GPUs, and supports directly controlling the generated images' rendering pose and various semantic attributes. Different from existing related parametric models, we use the neural radiance fields as a novel 3D proxy instead of the traditional 3D textured mesh, which makes that HeadNeRF is able to generate high fidelity images. However, the computationally expensive rendering process of the original NeRF hinders the construction of the parametric NeRF model. To address this issue, we adopt the strategy of integrating 2D neural rendering to the rendering process of NeRF and design novel loss terms. As a result, the rendering speed of HeadNeRF can be significantly accelerated, and the rendering time of one frame is reduced from 5s to 25ms. The well designed loss terms also improve the rendering accuracy, and the fine-level details of the human head, such as the gaps between teeth, wrinkles, and beards, can be represented and synthesized by HeadNeRF. Extensive experimental results and several applications demonstrate its effectiveness. The trained parametric model is available at https://github.com/CrisHY1995/headnerf. Yang Hong 0003, Haiyao Xiao, Ligang Liu 0001, Juyong Zhang |
CVPR | 5 |
| 2022 | Neural Points: Point Cloud Representation with Neural Fields for Arbitrary UpsamplingabstractIn this paper, we propose Neural Points, a novel point cloud representation and apply it to the arbitrary-factored upsampling task. Different from traditional point cloud representation where each point only represents a position or a local plane in the 3D space, each point in Neural Points represents a local continuous geometric shape via neural fields. Therefore, Neural Points contain more shape information and thus have a stronger representation ability. Neural Points is trained with surface containing rich geometric details, such that the trained model has enough expression ability for various shapes. Specifically, we extract deep local features on the points and construct neural fields through the local isomorphism between the 2D parametric domain and the 3D local patch. In the final, local neural fields are integrated together to form the global surface. Experimental results show that Neural Points has powerful representation ability and demonstrate excellent robustness and generalization ability. With Neural Points, we can resample point cloud with arbitrary resolutions, and it outperforms the state-of-the-art point cloud upsampling methods. Code is available at https://github.com/WanquanF/NeuralPoints. Wanquan Feng, Hongrui Cai, Juyong Zhang |
CVPR | 5 |
| 2022 | SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoabstractWe propose SelfRecon, a clothed human body reconstruction method that combines implicit and explicit repre-sentations to recover space-time coherent geometries from a monocular self-rotating human video. Explicit methods require a predefined template mesh for a given sequence, while the template is hard to acquire for a specific subject. Meanwhile, the fixed topology limits the reconstruction accuracy and clothing types. Implicit representation supports arbitrary topology and can represent high-fidelity geometry shapes due to its continuous nature. However, it is difficult to integrate multi-frame information to produce a consistent registration sequence for downstream applications. We propose to combine the advantages of both representations. We utilize differential mask loss of the explicit mesh to obtain the coherent overall shape, while the details on the implicit surface are refined with the differentiable neural rendering. Meanwhile, the explicit mesh is updated periodically to adjust its topology changes, and a consistency loss is designed to match both representations. Compared with existing methods, SelfRecon can produce high-fidelity surfaces for arbitrary clothed humans with self-supervised optimization. Extensive experimental results demonstrate its effectiveness on real captured monocular videos. The source code is available at https://github.com/jby1993/SelfReconCode. Boyi Jiang, Yang Hong 0003, Hujun Bao, Juyong Zhang |
CVPR | 4 |
| 2022 | TalkingFlow: Talking Facial Landmark Generation with Multi-Scale Normalizing Flow NetworkabstractDeterministic models dominate the field of talking facial land-mark generation by directly mapping speech signals to a certain lip-sync facial landmark sequence, which often suffer from regression to the mean face. In contrast, probability generative models are more beneficial to handle complex data space and generate diverse samples. In this work, we pro-pose a flow-based probabilistic network named TalkingFlow to generate natural talking facial landmark with head movements from speech data. It is implemented by a weighted multi-scale architecture to improve model representation capability and a conditional Temporal Convolutional Network module to fuse speech data. Extensive experiments results show that it can effectively generate diverse and natural facial landmark from speech data. All code will be made publicly available online. Sen Liang, Zhize Zhou, Juyong Zhang, Hujun Bao |
ICASSP | 4 |
| 2022 | CariPainter: Sketch Guided Interactive Caricature GenerationabstractIn this paper, we propose CariPainter, the first interactive caricature generating and editing method. The main challenge of caricature generation lies in the fact that it not only exaggerates the facial geometry but also refreshes the facial texture. We solve this challenging problem by utilizing the semantic segmentation maps as an intermediary domain, removing the influence of photo texture while preserving the person-specific geometry features. Specifically, our proposed method consists of two main components: CariSketchNet and CariMaskGAN. CariSketchNet exaggerates the photo segmentation map to construct CariMask. Then, CariMask is converted into a caricature by CariMaskGAN. In this step, users can edit and adjust the geometry of caricatures freely. Additionally, we propose a semantic detail pre-processing approach, which considerably increases details of generated images and allows modification of hair strands, wrinkles, and beards. Extensive experimental results show that our method produces higher-quality caricatures as well as supports easily used interactive modification. Hongrui Cai, Juyong Zhang, Jinyuan Jia 0002 |
ACM Multimedia | 4 |
| 2022 | Neural Surface Reconstruction of Dynamic Scenes with Monocular RGB-D CameraabstractWe propose Neural-DynamicReconstruction (NDR), a template-free method to recover high-fidelity geometry and motions of a dynamic scene from a monocular RGB-D camera. In NDR, we adopt the neural implicit function for surface representation and rendering such that the captured color and depth can be fully utilized to jointly optimize the surface and deformations. To represent and constrain the non-rigid deformations, we propose a novel neural invertible deforming network such that the cycle consistency between arbitrary two frames is automatically satisfied. Considering that the surface topology of dynamic scene might change over time, we employ a topology-aware strategy to construct the topology-variant correspondence for the fused frames. NDR also further refines the camera poses in a global optimization manner. Experiments on public datasets and our collected dataset demonstrate that NDR outperforms existing monocular dynamic reconstruction methods. Hongrui Cai, Wanquan Feng, Xuetao Feng, Juyong Zhang |
NeurIPS | 5 |
| 2022 | A Survey of Non-Rigid 3D RegistrationabstractAbstract Non‐rigid registration computes an alignment between a source surface with a target surface in a non‐rigid manner. In the past decade, with the advances in 3D sensing technologies that can measure time‐varying surfaces, non‐rigid registration has been applied for the acquisition of deformable shapes and has a wide range of applications. This survey presents a comprehensive review of non‐rigid registration methods for 3D shapes, focusing on techniques related to dynamic shape acquisition and reconstruction. In particular, we review different approaches for representing the deformation field, and the methods for computing the desired deformation. Both optimization‐based and learning‐based methods are covered. We also review benchmarks and datasets for evaluating non‐rigid registration methods, and discuss potential future research directions. Bailin Deng, Yuxin Yao 0001, Roberto M. Dyke, Juyong Zhang |
Comput. Graph. Forum | 4 |
| 2022 | RegGeoNet: Learning Regular Representations for Large-Scale 3D Point Clouds
Qijian Zhang, Junhui Hou, Antoni B. Chan, Juyong Zhang, Ying He 0001 |
Int. J. Comput. Vis. | 5 |
| 2022 | Robust 3D face modeling and tracking from RGB-D images
Changwei Luo, Juyong Zhang, Changcun Bao, Yali Li 0001, Shengjin Wang |
Multim. Syst. | 2 |
| 2022 | Fast and Robust Iterative Closest PointabstractThe iterative closest point (ICP) algorithm and its variants are a fundamental technique for rigid registration between two point sets, with wide applications in different areas from robotics to 3D reconstruction. The main drawbacks for ICP are its slow convergence, as well as its sensitivity to outliers, missing data, and partial overlaps. Recent work such as Sparse ICP achieves robustness via sparsity optimization at the cost of computational speed. In this paper, we propose a new method for robust registration with fast convergence. First, we show that the classical point-to-point ICP can be treated as a majorization-minimization (MM) algorithm, and propose an Anderson acceleration approach to speed up its convergence. In addition, we introduce a robust error metric based on the Welsch's function, which is minimized efficiently using the MM algorithm with Anderson acceleration. On challenging datasets with noises and partial overlaps, we achieve similar or better accuracy than Sparse ICP while being at least an order of magnitude faster. Finally, we extend the robust formulation to point-to-plane ICP, and solve the resulting problem using a similar Anderson-accelerated MM strategy. Our robust ICP methods improve the registration accuracy on benchmark datasets while being competitive in computational time. Juyong Zhang, Yuxin Yao 0001, Bailin Deng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Prior-Guided Multi-View 3D Head ReconstructionabstractRecovery of a 3D head model including the complete face and hair regions is still a challenging problem in computer vision and graphics. In this paper, we consider this problem using only a few multi-view portrait images as input. Previous multi-view stereo methods that have been based, either on optimization strategies or deep learning techniques, suffer from low-frequency geometric structures such as unclear head structures and inaccurate reconstruction in hair regions. To tackle this problem, we propose a prior-guided implicit neural rendering network. Specifically, we model the head geometry with a learnable signed distance field (SDF) and optimize it via an implicit differentiable renderer with the guidance of some human head priors, including the facial prior knowledge, head semantic segmentation information and 2D hair orientation maps. The utilization of these priors can improve the reconstruction accuracy and robustness, leading to a high-quality integrated 3D head model. Extensive ablation studies and comparisons with state-of-the-art methods demonstrate that our method can generate high-fidelity 3D head geometries with the guidance of these priors. Zhongqi Yang, Juyong Zhang |
IEEE Trans. Multim. | 4 |
| 2022 | Reconstructing Personalized Semantic Facial NeRF Models from Monocular VideoabstractWe present a novel semantic model for human head defined with neural radiance field. The 3D-consistent head model consist of a set of disentangled and interpretable bases, and can be driven by low-dimensional expression coefficients. Thanks to the powerful representation ability of neural radiance field, the constructed model can represent complex facial attributes including hair, wearings, which can not be represented by traditional mesh blendshape. To construct the personalized semantic facial model, we propose to define the bases as several multi-level voxel fields. With a short monocular RGB video as input, our method can construct the subject's semantic facial NeRF model with only ten to twenty minutes, and can render a photorealistic human head image in tens of miliseconds with a given expression coefficient and view direction. With this novel representation, we apply it to many tasks like facial retargeting and expression editing. Experimental results demonstrate its strong representation ability and training/inference speed. Demo videos and released code are provided in our project page: https://ustc3dv.github.io/NeRFBlendShape/ Xuan Gao 0003, Chenglai Zhong, Yang Hong 0003, Juyong Zhang |
ACM Trans. Graph. | 6 |
| 2022 | GDR-Net: A Geometric Detail Recovering Network for 3D Scanned ObjectsabstractThis article addresses the problem of mesh super-resolution such that the geometry details which are not well represented in the low-resolution models can be recovered and well represented in the generated high-quality models. The main challenges of this problem are the nonregularity of 3D mesh representation and the high complexity of 3D shapes. We propose a deep neural network called GDR-Net to solve this ill-posed problem, which resolves the two challenges simultaneously. First, to overcome the nonregularity, we regress a displacement in radial basis function parameter space instead of the vertex-wise coordinates in the euclidean space. Second, to overcome the high complexity, we apply the detail recovery process to small surface patches extracted from the input surface and obtain the overall high-quality mesh by fusing the refined surface patches. To train the network, we constructed a dataset composed of both real-world and synthetic scanned models, including high/low-quality pairs. Our experimental results demonstrate that GDR-Net works well for general models and outperforms previous methods for recovering geometric details. Wanquan Feng, Juyong Zhang, Yuanfeng Zhou, Shi-Qing Xin |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | GeodesicEmbedding (GE): A High-Dimensional Embedding Approach for Fast Geodesic Distance QueriesabstractIn this article, we develop a novel method for fast geodesic distance queries. The key idea is to embed the mesh into a high-dimensional space, such that the euclidean distance in the high-dimensional space can induce the geodesic distance in the original manifold surface. However, directly solving the high-dimensional embedding problem is not feasible due to the large number of variables and the fact that the embedding problem is highly nonlinear. We overcome the challenges with two novel ideas. First, instead of taking all vertices as variables, we embed only the saddle vertices, which greatly reduces the problem complexity. We then compute a local embedding for each non-saddle vertex. Second, to reduce the large approximation error resulting from the purely euclidean embedding, we propose a cascaded optimization approach that repeatedly introduces additional embedding coordinates with a non-euclidean function to reduce the approximation residual. Using the precomputation data, our approach can determine the geodesic distance between any two vertices in near-constant time. Computational testing results show that our method is more desirable than previous geodesic distance queries methods. Qianwei Xia, Juyong Zhang, Zheng Fang 0008, Bailin Deng, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Facial Expression Retargeting From Human to Avatar Made EasyabstractFacial expression retargeting from humans to virtual characters is a useful technique in computer graphics and animation. Traditional methods use markers or blendshapes to construct a mapping between the human and avatar faces. However, these approaches require a tedious 3D modeling process, and the performance relies on the modelers' experience. In this article, we propose a brand-new solution to this cross-domain expression transfer problem via nonlinear expression embedding and expression domain translation. We first build low-dimensional latent spaces for the human and avatar facial expressions with variational autoencoder. Then we construct correspondences between the two latent spaces guided by geometric and perceptual constraints. Specifically, we design geometric correspondences to reflect geometric matching and utilize a triplet data structure to express users' perceptual preference of avatar expressions. A user-friendly method is proposed to automatically generate triplets for a system allowing users to easily and efficiently annotate the correspondences. Using both geometric and perceptual correspondences, we trained a network for expression domain translation from human to avatar. Extensive experimental results and user studies demonstrate that even nonprofessional users can apply our method to generate high-quality facial expression retargeting results with less time and effort. Juyong Zhang, Jianmin Zheng |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | Learning Local Neighboring Structure for Robust 3D Shape RepresentationabstractMesh is a powerful data structure for 3D shapes. Representation learning for 3D meshes is important in many computer vision and graphics applications. The recent success of convolutional neural networks (CNNs) for structured data (e.g., images) suggests the value of adapting insight from CNN for 3D shapes. However, 3D shape data are irregular since each node's neighbors are unordered. Various graph neural networks for 3D shapes have been developed with isotropic filters or predefined local coordinate systems to overcome the node inconsistency on graphs. However, isotropic filters or predefined local coordinate systems limit the representation power. In this paper, we propose a local structure-aware anisotropic convolutional operation (LSA-Conv) that learns adaptive weighting matrices for each node according to the local neighboring structure and performs shared anisotropic filters. In fact, the learnable weighting matrix is similar to the attention matrix in random synthesizer -- a new Transformer model for natural language processing (NLP). Comprehensive experiments demonstrate that our model produces significant improvement in 3D shape reconstruction compared to state-of-the-art methods. Zhongpai Gao, Junchi Yan, Guangtao Zhai, Juyong Zhang, Yiyan Yang, Xiaokang Yang 0001 |
AAAI | 4 |
| 2021 | StereoPIFu: Depth Aware Clothed Human Digitization via Stereo VisionabstractIn this paper, we propose StereoPIFu, which integrates the geometric constraints of stereo vision with implicit function representation of PIFu, to recover the 3D shape of the clothed human from a pair of low-cost rectified images. First, we introduce the effective voxel-aligned features from a stereo vision-based network to enable depth-aware reconstruction. Moreover, the novel relative z-offset is employed to associate predicted high-fidelity human depth and occupancy inference, which helps restore fine-level surface de-tails. Second, a network structure that fully utilizes the geometry information from the stereo images is designed to improve the human body reconstruction quality. Consequently, our StereoPIFu can naturally infer the human body’s spatial location in camera space and maintain the correct relative position of different parts of the human body, which enables our method to capture human performance. Compared with previous works, our StereoPIFu significantly improves the robustness, completeness, and accuracy of the clothed human reconstruction, which is demonstrated by extensive experimental results. Yang Hong 0003, Juyong Zhang, Boyi Jiang, Ligang Liu 0001, Hujun Bao |
CVPR | 2 |
| 2021 | Recurrent Multi-View Alignment Network for Unsupervised Surface RegistrationabstractLearning non-rigid registration in an end-to-end manner is challenging due to the inherent high degrees of freedom and the lack of labeled training data. In this paper, we resolve these two challenges simultaneously. First, we propose to represent the non-rigid transformation with a point-wise combination of several rigid transformations. This representation not only makes the solution space well-constrained but also enables our method to be solved iteratively with a recurrent framework, which greatly reduces the difficulty of learning. Second, we introduce a differentiable loss function that measures the 3D shape similarity on the projected multi-view 2D depth images so that our full framework can be trained end-to-end without ground truth supervision. Extensive experiments on several different datasets demonstrate that our proposed method outperforms the previous state-of-the-art by a large margin. Wanquan Feng, Juyong Zhang, Hongrui Cai, Haofei Xu, Junhui Hou, Hujun Bao |
CVPR | 2 |
| 2021 | A Robust Loss for Point Cloud Registration
Yuxin Yao 0001, Bailin Deng, Juyong Zhang |
ICCV | 4 |
| 2021 | AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisabstractGenerating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation networks. Our method is completely different from existing methods that rely on intermediate representations like 2D landmarks or 3D face models to bridge the gap between audio input and video output. Specifically, the feature of input audio signal is directly fed into a conditional implicit function to generate a dynamic neural radiance field, from which a high-fidelity talking-head video corresponding to the audio signal is synthesized using volume rendering. Another advantage of our framework is that not only the head (with hair) region is synthesized as previous methods did, but also the upper body is generated via two individual neural radiance fields. Experimental results demonstrate that our novel framework can (1) produce high-fidelity and natural results, and (2) support free adjustment of audio signals, viewing directions, and background images. Code is available at https://github.com/YudongGuo/AD-NeRF. Sen Liang, Yong-Jin Liu 0001, Hujun Bao, Juyong Zhang |
ICCV | 6 |
| 2021 | High-Resolution Optical Flow from 1D Attention and CorrelationabstractOptical flow is inherently a 2D search problem, and thus the computational complexity grows quadratically with respect to the search window, making large displacements matching infeasible for high-resolution images. In this paper, we take inspiration from Transformers and propose a new method for high-resolution optical flow estimation with significantly less computation. Specifically, a 1D attention operation is first applied in the vertical direction of the target image, and then a simple 1D correlation in the horizontal direction of the attended image is able to achieve 2D correspondence modeling effect. The directions of attention and correlation can also be exchanged, resulting in two 3D cost volumes that are concatenated for optical flow estimation. The novel 1D formulation empowers our method to scale to very high-resolution input images while maintaining competitive performance. Extensive experiments on Sintel, KITTI and real-world 4K (2160 × 3840) resolution images demonstrated the effectiveness and superiority of our proposed method. Code and models are available at https://github.com/haofeixu/flow1d. Haofei Xu, Jiaolong Yang, Jianfei Cai 0001, Juyong Zhang, Xin Tong 0001 |
ICCV | 4 |
| 2021 | Pre-training Graph Transformer with Multimodal Side Information for RecommendationabstractSide information of items, e.g., images and text description, has shown to be effective in contributing to accurate recommendations. Inspired by the recent success of pre-training models on natural language and images, we propose a pre-training strategy to learn item representations by considering both item side information and their relationships. We relate items by common user activities, e.g., co-purchase, and construct a homogeneous item graph. This graph provides a unified view of item relations and their associated side information in multimodality. We develop a novel sampling algorithm named MCNSampling to select contextual neighbors for each item. The proposed Pre-trained Multimodal Graph Transformer (PMGT) learns item representations with two objectives: 1) graph structure reconstruction, and 2) masked node feature reconstruction. Experimental results on real datasets demonstrate that the proposed PMGT model effectively exploits the multimodality side information to achieve better accuracies in downstream tasks including item recommendation and click-through ratio prediction. In addition, we also report a case study of testing PMGT in an online setting with 600 thousand users. Yong Liu 0020, Susen Yang, Chenyi Lei, Guoxin Wang 0002, Haihong Tang, Juyong Zhang, Aixin Sun, Chunyan Miao |
ACM Multimedia | 6 |
| 2021 | Landmark Detection and 3D Face Reconstruction for Caricature using a Nonlinear Parametric Model
Hongrui Cai, Zhuang Peng, Juyong Zhang |
Graph. Model. | 4 |
| 2021 | Real-time face view correction for front-facing camerasabstractFace views are particularly important in person-to-person communication. Differenes between the camera location and the face orientation can result in undesirable facial appearances of the participants during video conferencing. This phenomenon is particularly noticeable when using devices where the front-facing camera is placed in unconventional locations such as below the display or within the keyboard. In this paper, we take a video stream from a single RGB camera as input, and generate a video stream that emulates the view from a virtual camera at a designated location. The most challenging issue in this problem is that the corrected view often needs out-of-plane head rotations. To address this challenge, we reconstruct the 3D face shape and re-render it into synthesized frames according to the virtual camera location. To output the corrected video stream with natural appearance in real time, we propose several novel techniques including accurate eyebrow reconstruction, high-quality blending between the corrected face image and background, and template-based 3D reconstruction of glasses. Our system works well for different lighting conditions and skin tones, and can handle users wearing glasses. Extensive experiments and user studies demonstrate that our method provides high-quality results. Juyong Zhang, Hongrui Cai, Zhangjin Huang, Bailin Deng |
Comput. Vis. Media | 2 |
| 2021 | Deformation representation based convolutional mesh autoencoder for 3D hand generation
Xinqian Zheng, Boyi Jiang, Juyong Zhang |
Neurocomputing | 3 |
| 2021 | Parallel and Scalable Heat Methods for Geodesic Distance ComputationabstractIn this paper, we propose a parallel and scalable approach for geodesic distance computation on triangle meshes. Our key observation is that the recovery of geodesic distance with the heat method [1] can be reformulated as optimization of its gradients subject to integrability, which can be solved using an efficient first-order method that requires no linear system solving and converges quickly. Afterward, the geodesic distance is efficiently recovered by parallel integration of the optimized gradients in breadth-first order. Moreover, we employ a similar breadth-first strategy to derive a parallel Gauss-Seidel solver for the diffusion step in the heat method. To further lower the memory consumption from gradient optimization on faces, we also propose a formulation that optimizes the projected gradients on edges, which reduces the memory footprint by about 50 percent. Our approach is trivially parallelizable, with a low memory footprint that grows linearly with respect to the model size. This makes it particularly suitable for handling large models. Experimental results show that it can efficiently compute geodesic distance on meshes with more than 200 million vertices on a desktop PC with 128 GB RAM, outperforming the original heat method and other state-of-the-art geodesic distance solvers. Jiong Tao, Juyong Zhang, Bailin Deng, Zheng Fang 0008, Ying He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | 3D Face From X: Learning Face Shape From Diverse SourcesabstractWe present a novel method to jointly learn a 3D face parametric model and 3D face reconstruction from diverse sources. Previous methods usually learn 3D face modeling from one kind of source, such as scanned data or in-the-wild images. Although 3D scanned data contain accurate geometric information of face shapes, the capture system is expensive and such datasets usually contain a small number of subjects. On the other hand, in-the-wild face images are easily obtained and there are a large number of facial images. However, facial images do not contain explicit geometric information. In this paper, we propose a method to learn a unified face model from diverse sources. Besides scanned face data and face images, we also utilize a large number of RGB-D images captured with an iPhone X to bridge the gap between the two sources. Experimental results demonstrate that with training data from more sources, we can learn a more powerful face model. Lin Cai 0005, Juyong Zhang |
IEEE Trans. Image Process. | 3 |
| 2020 | Diversified Interactive Recommendation with Implicit FeedbackabstractInteractive recommender systems that enable the interactions between users and the recommender system have attracted increasing research attention. Previous methods mainly focus on optimizing recommendation accuracy. However, they usually ignore the diversity of the recommendation results, thus usually results in unsatisfying user experiences. In this paper, we propose a novel diversified recommendation model, named Diversified Contextual Combinatorial Bandit (DC2B), for interactive recommendation with users' implicit feedback. Specifically, DC2B employs determinantal point process in the recommendation procedure to promote diversity of the recommendation results. To learn the model parameters, a Thompson sampling-type algorithm based on variational Bayesian inference is proposed. In addition, theoretical regret analysis is also provided to guarantee the performance of DC2B. Extensive experiments on real datasets are performed to demonstrate the effectiveness of the proposed method in balancing the recommendation accuracy and diversity. Yong Liu 0020, Yingtai Xiao, Qiong Wu 0001, Chunyan Miao, Juyong Zhang, Binqiang Zhao, Haihong Tang |
AAAI | 5 |
| 2020 | Lightweight Photometric Stereo for Facial Details RecoveryabstractRecently, 3D face reconstruction from a single image has achieved great success with the help of deep learning and shape prior knowledge, but they often fail to produce accurate geometry details. On the other hand, photometric stereo methods can recover reliable geometry details, but require dense inputs and need to solve a complex optimization problem. In this paper, we present a lightweight strategy that only requires sparse inputs or even a single image to recover high-fidelity face shapes with images captured under near-field lights. To this end, we construct a dataset containing 84 different subjects with 29 expressions under 3 different lights. Data augmentation is applied to enrich the data in terms of diversity in identity, lighting, expression, etc. With this constructed dataset, we propose a novel neural network specially designed for photometric stereo based 3D face reconstruction. Extensive experiments and comparisons demonstrate that our method can generate high-quality reconstruction results with one to three facial images captured under near-field lights. Our full framework is available at https://github.com/Juyong/FacePSNet. Bailin Deng, Juyong Zhang |
CVPR | 4 |
| 2020 | AANet: Adaptive Aggregation Network for Efficient Stereo MatchingabstractDespite the remarkable progress made by learning based stereo matching algorithms, one key challenge remains unsolved. Current state-of-the-art stereo models are mostly based on costly 3D convolutions, the cubic computational complexity and high memory consumption make it quite expensive to deploy in real-world applications. In this paper, we aim at completely replacing the commonly used 3D convolutions to achieve fast inference speed while maintaining comparable accuracy. To this end, we first propose a sparse points based intra-scale cost aggregation method to alleviate the well-known edge-fattening issue at disparity discontinuities. Further, we approximate traditional cross-scale cost aggregation algorithm with neural network layers to handle large textureless regions. Both modules are simple, lightweight, and complementary, leading to an effective and efficient architecture for cost aggregation. With these two modules, we can not only significantly speed up existing top-performing models (e.g., 41x than GC-Net, 4x than PSMNet and 38x than GA-Net), but also improve the performance of fast stereo models (e.g., StereoNet). We also achieve competitive results on Scene Flow and KITTI datasets while running at 62ms, demonstrating the versatility and high efficiency of the proposed method. Our full framework is available at https://github.com/haofeixu/aanet. Haofei Xu, Juyong Zhang |
CVPR | 2 |
| 2020 | Quasi-Newton Solver for Robust Non-Rigid RegistrationabstractImperfect data (noise, outliers and partial overlap) and high degrees of freedom make non-rigid registration a classical challenging problem in computer vision. Existing methods typically adopt the l_p type robust estimator to regularize the fitting and smoothness, and the proximal operator is used to solve the resulting non-smooth problem. However, the slow convergence of these algorithms limits its wide applications. In this paper, we propose a formulation for robust non-rigid registration based on a globally smooth robust estimator for data fitting and regularization, which can handle outliers and partial overlaps. We apply the majorization-minimization algorithm to the problem, which reduces each iteration to solving a simple least-squares problem with L-BFGS. Extensive experiments demonstrate the effectiveness of our method for non-rigid alignment between two shapes with outliers and partial overlap. with quantitative evaluation showing that it outperforms state-of-the-art methods in terms of registration accuracy and computational speed. The source code is available at https://github.com/Juyong/Fast_RNRR. Yuxin Yao 0001, Bailin Deng, Weiwei Xu 0003, Juyong Zhang |
CVPR | 4 |
| 2020 | BCNet: Learning Body and Cloth Shape from a Single Image
Boyi Jiang, Juyong Zhang, Yang Hong 0003, Jinhao Luo, Ligang Liu 0001, Hujun Bao |
ECCV (20) | 2 |
| 2020 | Modeling Caricature Expressions by 3D Blendshape and Dynamic TextureabstractThe problem of deforming an artist-drawn caricature according to a given normal face expression is of interest in applications such as social media, animation and entertainment. This paper presents a solution to the problem, with an emphasis on enhancing the ability to create desired expressions and meanwhile preserve the identity exaggeration style of the caricature, which imposes challenges due to the complicated nature of caricatures. The key of our solution is a novel method to model caricature expression, which extends traditional 3DMM representation to caricature domain. The method consists of shape modelling and texture generation for caricatures. Geometric optimization is developed to create identity-preserving blendshapes for reconstructing accurate and stable geometric shape, and a conditional generative adversarial network (cGAN) is designed for generating dynamic textures under target expressions. The combination of both shape and texture components makes the non-trivial expressions of a caricature be effectively defined by the extension of the popular 3DMM representation and a caricature can thus be flexibly deformed into arbitrary expressions with good results visually in both shape and color spaces. The experiments demonstrate the effectiveness of the proposed method. Jianmin Zheng, Jianfei Cai 0001, Juyong Zhang |
ACM Multimedia | 4 |
| 2020 | Anderson Acceleration for Nonconvex ADMM Based on Douglas-Rachford SplittingabstractAbstract The alternating direction multiplier method (ADMM) is widely used in computer graphics for solving optimization problems that can be nonsmooth and nonconvex. It converges quickly to an approximate solution, but can take a long time to converge to a solution of high‐accuracy. Previously, Anderson acceleration has been applied to ADMM, by treating it as a fixed‐point iteration for the concatenation of the dual variables and a subset of the primal variables. In this paper, we note that the equivalence between ADMM and Douglas‐Rachford splitting reveals that ADMM is in fact a fixed‐point iteration in a lower‐dimensional space. By applying Anderson acceleration to such lower‐dimensional fixed‐point iteration, we obtain a more effective approach for accelerating ADMM. We analyze the convergence of the proposed acceleration method on nonconvex problems, and verify its effectiveness on a variety of computer graphics including geometry processing and physical simulation. Wenqing Ouyang, Yuxin Yao 0001, Juyong Zhang, Bailin Deng |
Comput. Graph. Forum | 4 |
| 2020 | Robust RGB-D Face Recognition Using Attribute-Aware LossabstractExisting convolutional neural network (CNN) based face recognition algorithms typically learn a discriminative feature mapping, using a loss function that enforces separation of features from different classes and/or aggregation of features within the same class. However, they may suffer from bias in the training data such as uneven sampling density, because they optimize the adjacency relationship of the learned features without considering the proximity of the underlying faces. Moreover, since they only use facial images for training, the learned feature mapping may not correctly indicate the relationship of other attributes such as gender and ethnicity, which can be important for some face recognition applications. In this paper, we propose a new CNN-based face recognition approach that incorporates such attributes into the training process. Using an attribute-aware loss function that regularizes the feature mapping using attribute proximity, our approach learns more discriminative features that are correlated with the attributes. We train our face recognition model on a large-scale RGB-D data set with over 100K identities captured under real application conditions. By comparing our approach with other methods on a variety of experiments, we demonstrate that depth channel and attribute-aware loss greatly improve the accuracy and robustness of face recognition. Luo Jiang, Juyong Zhang, Bailin Deng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Disentangled Human Body Embedding Based on Deep Hierarchical Neural NetworkabstractHuman bodies exhibit various shapes for different identities or poses, but the body shape has certain similarities in structure and thus can be embedded in a low-dimensional space. This article presents an autoencoder-like network architecture to learn disentangled shape and pose embedding specifically for the 3D human body. This is inspired by recent progress of deformation-based latent representation learning. To improve the reconstruction accuracy, we propose a hierarchical reconstruction pipeline for the disentangling process and construct a large dataset of human body models with consistent connectivity for the learning of the neural network. Our learned embedding can not only achieve superior reconstruction accuracy but also provide great flexibility in 3D human body generation via interpolation, bilinear interpolation, and latent space sampling. The results from extensive experiments demonstrate the powerfulness of our learned 3D human body embedding in various applications. Boyi Jiang, Juyong Zhang, Jianfei Cai 0001, Jianmin Zheng |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Disentangled Representation Learning for 3D Face ShapeabstractIn this paper, we present a novel strategy to design disentangled 3D face shape representation. Specifically, a given 3D face shape is decomposed into identity part and expression part, which are both encoded and decoded in a nonlinear way. To solve this problem, we propose an attribute decomposition framework for 3D face mesh. To better represent face shapes which are usually nonlinear deformed between each other, the face shapes are represented by a vertex based deformation representation rather than Euclidean coordinates. The experimental results demonstrate that our method has better performance than existing methods on decomposing the identity and expression parts. Moreover, more natural expression transfer results can be achieved with our method than existing methods. Zihang Jiang, Qianyi Wu, Juyong Zhang |
CVPR | 4 |
| 2019 | Region Deformer Networks for Unsupervised Depth Estimation from Unconstrained Monocular VideosabstractWhile learning based depth estimation from images/videos has achieved substantial progress, there still exist intrinsic limitations. Supervised methods are limited by a small amount of ground truth or labeled data and unsupervised methods for monocular videos are mostly based on the static scene assumption, not performing well on real world scenarios with the presence of dynamic objects. In this paper, we propose a new learning based method consisting of DepthNet, PoseNet and Region Deformer Networks (RDN) to estimate depth from unconstrained monocular videos without ground truth supervision. The core contribution lies in RDN for proper handling of rigid and non-rigid motions of various objects such as rigidly moving cars and deformable humans. In particular, a deformation based motion representation is proposed to model individual object motion on 2D images. This representation enables our method to be applicable to diverse unconstrained monocular videos. Our method can not only achieve the state-of-the-art results on standard benchmarks KITTI and Cityscapes, but also show promising results on a crowded pedestrian tracking dataset, which demonstrates the effectiveness of the deformation based motion representation. Code and trained models are available at https://github.com/haofeixu/rdn4depth. Haofei Xu, Jianmin Zheng, Jianfei Cai 0001, Juyong Zhang |
IJCAI | 4 |
| 2019 | Vectorization Based Color Transfer for Portrait Images
Ying He 0001, Fei Hou 0001, Juyong Zhang, Anxiang Zeng, Yong-Jin Liu 0001 |
Comput. Aided Des. | 4 |
| 2019 | A fast numerical solver for local barycentric coordinates
Jiong Tao, Bailin Deng, Juyong Zhang |
Comput. Aided Geom. Des. | 3 |
| 2019 | Conditional adversarial synthesis of 3D facial action units
Zhilei Liu, Guoxian Song, Jianfei Cai 0001, Tat-Jen Cham, Juyong Zhang |
Neurocomputing | 5 |
| 2019 | CNN-Based Real-Time Dense Face Reconstruction with Inverse-Rendered Photo-Realistic Face ImagesabstractWith the powerfulness of convolution neural networks (CNN), CNN based face reconstruction has recently shown promising performance in reconstructing detailed face shape from 2D face images. The success of CNN-based methods relies on a large number of labeled data. The state-of-the-art synthesizes such data using a coarse morphable face model, which however has difficulty to generate detailed photo-realistic images of faces (with wrinkles). This paper presents a novel face data generation method. Specifically, we render a large number of photo-realistic face images with different attributes based on inverse rendering. Furthermore, we construct a fine-detailed face image dataset by transferring different scales of details from one image to another. We also construct a large number of video-type adjacent frame pairs by simulating the distribution of real video data.11.All these coarse-scale and fine-scale photo-realistic face image datasets can be downloaded from https://github.com/Juyong/3DFace. With these nicely constructed datasets, we propose a coarse-to-fine learning framework consisting of three convolutional networks. The networks are trained for real-time detailed 3D face reconstruction from monocular video as well as from a single image. Extensive experimental results demonstrate that our framework can produce high-quality reconstruction but with much less computation time compared to the state-of-the-art. Moreover, our method is robust to pose, expression and lighting due to the diversity of data. Juyong Zhang, Jianfei Cai 0001, Boyi Jiang, Jianmin Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Real-Time Head Pose Estimation and Face Modeling From a Depth ImageabstractWe address the issues of 3-D head pose estimation and face modeling from a depth image. Given a depth image, random forests are effective for estimating the location and orientation of a person's head. However, the accuracy of the estimation is not high enough. We propose using corrected regression votes. The corrected votes are obtained by considering the cooperation of all trees, leading to significant improvement of head pose estimation accuracy. Based on the head pose estimator, we present a face modeling system. In our system, the face model is generated by aligning a deformable face model to a depth image using an iterative closest point (ICP) algorithm. The novelty of our approach is that an optimal weight for each vertex is incorporated into the ICP algorithm with point to plane constraints. Experiments show that our system can automatically estimate the head pose and generate a realistic face model from a single depth image. We also provide a detailed evaluation that shows the benefits of our approach. Changwei Luo, Juyong Zhang, Jun Yu 0001, Chang Wen Chen, Shengjin Wang |
IEEE Trans. Multim. | 2 |
| 2019 | Accelerating ADMM for efficient simulation and optimizationabstractThe alternating direction method of multipliers (ADMM) is a popular approach for solving optimization problems that are potentially non-smooth and with hard constraints. It has been applied to various computer graphics applications, including physical simulation, geometry processing, and image processing. However, ADMM can take a long time to converge to a solution of high accuracy. Moreover, many computer graphics tasks involve non-convex optimization, and there is often no convergence guarantee for ADMM on such problems since it was originally designed for convex optimization. In this paper, we propose a method to speed up ADMM using Anderson acceleration, an established technique for accelerating fixed-point iterations. We show that in the general case, ADMM is a fixed-point iteration of the second primal variable and the dual variable, and Anderson acceleration can be directly applied. Additionally, when the problem has a separable target function and satisfies certain conditions, ADMM becomes a fixed-point iteration of only one variable, which further reduces the computational overhead of Anderson acceleration. Moreover, we analyze a particular non-convex problem structure that is common in computer graphics, and prove the convergence of ADMM on such problems under mild assumptions. We apply our acceleration technique on a variety of optimization problems in computer graphics, with notable improvement on their convergence speed. Juyong Zhang, Wenqing Ouyang, Bailin Deng |
ACM Trans. Graph. | 1 |
| 2019 | Static/Dynamic Filtering for Mesh GeometryabstractThe joint bilateral filter, which enables feature-preserving signal smoothing according to the structural information from a guidance, has been applied for various tasks in geometry processing. Existing methods either rely on a static guidance that may be inconsistent with the input and lead to unsatisfactory results, or a dynamic guidance that is automatically updated but sensitive to noises and outliers. Inspired by recent advances in image filtering, we propose a new geometry filtering technique called static/dynamic filter, which utilizes both static and dynamic guidances to achieve state-of-the-art results. The proposed filter is based on a nonlinear optimization that enforces smoothness of the signal while preserving variations that correspond to features of certain scales. We develop an efficient iterative solver for the problem, which unifies existing filters that are based on static or dynamic guidances. The filter can be applied to mesh face normals followed by vertex position update, to achieve scale-aware and feature-preserving filtering of mesh geometry. It also works well for other types of signals defined on mesh surfaces, such as texture colors. Extensive experimental results demonstrate the effectiveness of the proposed filter for various geometry processing applications such as mesh denoising, geometry feature enhancement, and texture color filtering. Juyong Zhang, Bailin Deng, Yang Hong 0003, Wenjie Qin, Ligang Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | Alive Caricature From 2D to 3DabstractCaricature is an art form that expresses subjects in abstract, simple and exaggerated views. While many caricatures are 2D images, this paper presents an algorithm for creating expressive 3D caricatures from 2D caricature images with minimum user interaction. The key idea of our approach is to introduce an intrinsic deformation representation that has the capability of extrapolation, enabling us to create a deformation space from standard face datasets, which maintains face constraints and meanwhile is sufficiently large for producing exaggerated face models. Built upon the proposed deformation representation, an optimization model is formulated to find the 3D caricature that captures the style of the 2D caricature image automatically. The experiments show that our approach has better capability in expressing caricatures than those fitting approaches directly using classical parametric face models such as 3DMM and FaceWareHouse. Moreover, our approach is based on standard face datasets and avoids constructing complicated 3D caricature training sets, which provides great flexibility in real applications. Qianyi Wu, Juyong Zhang, Yukun Lai, Jianmin Zheng, Jianfei Cai 0001 |
CVPR | 2 |
| 2018 | Real-time 3D Face-Eye Performance Capture of a Person Wearing VR HeadsetabstractTeleconference or telepresence based on virtual reality (VR) head-mount display (HMD) device is a very interesting and promising application since HMD can provide immersive feelings for users. However, in order to facilitate face-to-face communications for HMD users, real-time 3D facial performance capture of a person wearing HMD is needed, which is a very challenging task due to the large occlusion caused by HMD. The existing limited solutions are very complex either in setting or in approach as well as lacking the performance capture of 3D eye gaze movement. In this paper, we propose a convolutional neural network (CNN) based solution for real-time 3D face-eye performance capture of HMD users without complex modification to devices. To address the issue of lacking training data, we generate massive pairs of HMD face-label dataset by data synthesis as well as collecting VR-IR eye dataset from multiple subjects. Then, we train a dense-fitting network for facial region and an eye gaze network to regress 3D eye model parameters. Extensive experimental results demonstrate that our system can efficiently and effectively produce in real time a vivid personalized 3D avatar with the correct identity, pose, expression and eye motion corresponding to the HMD user. Guoxian Song, Jianfei Cai 0001, Tat-Jen Cham, Jianmin Zheng, Juyong Zhang, Henry Fuchs |
ACM Multimedia | 5 |
| 2018 | Shading-Based Surface Detail Recovery Under General Unknown IlluminationabstractReconstructing the shape of a 3D object from multi-view images under unknown, general illumination is a fundamental problem in computer vision. High quality reconstruction is usually challenging especially when fine detail is needed and the albedo of the object is non-uniform. This paper introduces vertex overall illumination vectors to model the illumination effect and presents a total variation (TV) based approach for recovering surface details using shading and multi-view stereo (MVS). Behind the approach are the two important observations: (1) the illumination over the surface of an object often appears to be piecewise smooth and (2) the recovery of surface orientation is not sufficient for reconstructing the surface, which was often overlooked previously. Thus we propose to use TV to regularize the overall illumination vectors and use visual hull to constrain partial vertices. The reconstruction is formulated as a constrained TV-minimization problem that simultaneously treats the shape and illumination vectors as unknowns. An augmented Lagrangian method is proposed to quickly solve the TV-minimization problem. As a result, our approach is robust, stable and is able to efficiently recover high-quality surface details even when starting with a coarse model obtained using MVS. These advantages are demonstrated by extensive experiments on the state-of-the-art MVS database, which includes challenging objects with varying albedo. Di Xu 0012, Qi Duan, Jianmin Zheng, Juyong Zhang, Jianfei Cai 0001, Tat-Jen Cham |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | 3D Face Reconstruction With Geometry Details From a Single Imageabstract3D face reconstruction from a single image is a classical and challenging problem, with wide applications in many areas. Inspired by recent works in face animation from RGBD or monocular video inputs, we develop a novel method for reconstructing 3D faces from unconstrained 2D images, using a coarse-to-fine optimization strategy. First, a smooth coarse 3D face is generated from an example-based bilinear face model, by aligning the projection of 3D face landmarks with 2D landmarks detected from the input image. Afterwards, using local corrective deformation fields, the coarse 3D face is refined using photometric consistency constraints, resulting in a medium face shape. Finally, a shape-from-shading method is applied on the medium face to recover fine geometric details. Our method outperforms stateof- the-art approaches in terms of accuracy and detail recovery, which is demonstrated in extensive experiments using real world models and publicly available datasets. Luo Jiang, Juyong Zhang, Bailin Deng, Ligang Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Anderson acceleration for geometry optimization and physics simulationabstractMany computer graphics problems require computing geometric shapes subject to certain constraints. This often results in non-linear and non-convex optimization problems with globally coupled variables, which pose great challenge for interactive applications. Local-global solvers developed in recent years can quickly compute an approximate solution to such problems, making them an attractive choice for applications that prioritize efficiency over accuracy. However, these solvers suffer from lower convergence rate, and may take a long time to compute an accurate result. In this paper, we propose a simple and effective technique to accelerate the convergence of such solvers. By treating each local-global step as a fixed-point iteration, we apply Anderson acceleration, a well-established technique for fixed-point solvers, to speed up the convergence of a local-global solver. To address the stability issue of classical Anderson acceleration, we propose a simple strategy to guarantee the decrease of target energy and ensure its global convergence. In addition, we analyze the connection between Anderson acceleration and quasi-Newton methods, and show that the canonical choice of its mixing parameter is suitable for accelerating local-global solvers. Moreover, our technique is effective beyond classical local-global solvers, and can be applied to iterative methods with a common structure. We evaluate the performance of our technique on a variety of geometry optimization and physics simulation problems. Our approach significantly reduces the number of iterations required to compute an accurate result, with only a slight increase of computational cost per iteration. Its simplicity and effectiveness makes it a promising tool for accelerating existing algorithms as well as designing efficient new algorithms. Bailin Deng, Juyong Zhang, Fanyu Geng, Wenjie Qin, Ligang Liu 0001 |
ACM Trans. Graph. | 3 |
| 2017 | Pixel-level guided face editing with fully convolution networksabstractFace editing has a variety of applications, especially with the increasing popularity of photography using mobile devices. In this work, we argue that the performance of face image editing can be further improved by using semantic segmentation which marks each pixel with a label that indicates its corresponding facial part. To this end, we propose a deep learning based method for automatic pixel-level labeling on face images. Our approach achieves state-of-the-art labeling accuracy on publicly available datasets, at a significantly higher speed than existing labeling methods. Then we show how the label information can be applied to various face image editing applications, such as face smoothing, face cloning and face blending. Extensive experimental results demonstrate the effectiveness of our method in editing face images with convincing visual quality. Zhenxi Li, Juyong Zhang |
ICME | 2 |
| 2017 | Joint head pose and facial landmark regression from depth imagesabstractThis paper presents a joint head pose and facial landmark regression method with input from depth images for real-time application. Our main contributions are: firstly, a joint optimization method to estimate head pose and facial landmarks, i.e., the pose regression result provides supervised initialization for cascaded facial landmark regression, while the regression result for the facial landmarks can also help to further refine the head pose at each stage. Secondly, we classify the head pose space into 9 sub-spaces, and then use a cascaded random forest with a global shape constraint for training facial landmarks in each specific space. This classification-guided method can effectively handle the problem of large pose changes and occlusion. Lastly, we have built a 3D face database containing 73 subjects, each with 14 expressions in various head poses. Experiments on challenging databases show our method achieves state-of-the-art performance on both head pose estimation and facial landmark regression. Juyong Zhang, Changwei Luo, Falai Chen |
Comput. Vis. Media | 2 |
| 2016 | Hand Pose Regression via a Classification-Guided Approach
Juyong Zhang |
ACCV (3) | 2 |
| 2016 | Photometric stereo using mesh face based optimizationabstractThe state-of-the-art photometric stereo (PS) methods typically apply shading cues on each vertex and represent a vertex normal as a non-linear function of its neighboring vertices. Such vertex-based representation leads to huge computational cost and limits it from processing dense meshes. In this paper, we propose a PS based surface reconstruction using mesh face based representation. In particular, we propose to apply the shading cue on each mesh face instead of each vertex and optimize face normals instead of vertex normals. We develop a two-step approach to solve the surface recovery problem, where we first optimize face normals using shading cues and then update vertices by the optimized face normals. Experimental results show that, compared with the state-of-the-art, such a two-step approach is able to reduce the runtime significantly as well as handling much denser meshes. Di Xu 0012, Jianfei Cai 0001, Jianmin Zheng, Juyong Zhang |
VCIP | 4 |
| 2016 | Upright orientation of 3D shapes with Convolutional Networks
Zishun Liu 0003, Juyong Zhang, Ligang Liu 0001 |
Graph. Model. | 2 |
| 2016 | FrameFab: robotic fabrication of frame shapesabstractFrame shapes, which are made of struts, have been widely used in many fields, such as art, sculpture, architecture, and geometric modeling, etc. An interest in robotic fabrication of frame shapes via spatial thermoplastic extrusion has been increasingly growing in recent years. In this paper, we present a novel algorithm to generate a feasible fabrication sequence for general frame shapes. To solve this non-trivial combinatorial problem, we develop a divide-and-conquer strategy that first decomposes the input frame shape into stable layers via a constrained sparse optimization model. Then we search a feasible sequence for each layer via a local optimization method together with a backtracking strategy. The generated sequence guarantees that the already-printed part is in a stable equilibrium state at all stages of fabrication, and that the 3D printing extrusion head does not collide with the printed part during the fabrication. Our algorithm has been validated by a built prototype robotic fabrication system made by a 6-axis KUKA robotic arm with a customized extrusion head. Experimental results demonstrate the feasibility and applicability of our algorithm. Yijiang Huang, Juyong Zhang, Xin Hu 0005, Guoxian Song, Zhongyuan Liu, Ligang Liu 0001 |
ACM Trans. Graph. | 2 |
| 2015 | l1-Regression based subdivision schemes for noisy data
Ghulam Mustafa 0001, Juyong Zhang, Jiansong Deng |
Comput. Aided Des. | 3 |
| 2015 | Guided Mesh Normal FilteringabstractThe joint bilateral filter is a variant of the standard bilateral filter, where the range kernel is evaluated using a guidance signal instead of the original signal. It has been successfully applied to various image processing problems, where it provides more flexibility than the standard bilateral filter to achieve high quality results. On the other hand, its success is heavily dependent on the guidance signal, which should ideally provide a robust estimation about the features of the output signal. Such a guidance signal is not always easy to construct. In this paper, we propose a novel mesh normal filtering framework based on the joint bilateral filter, with applications in mesh denoising. Our framework is designed as a two-stage process: first, we apply joint bilateral filtering to the face normals, using a properly constructed normal field as the guidance; afterwards, the vertex positions are updated according to the filtered face normals. We compute the guidance normal on a face using a neighboring patch with the most consistent normal orientations, which provides a reliable estimation of the true normal even with a high-level of noise. The effectiveness of our approach is validated by extensive experimental results. Wangyu Zhang, Bailin Deng, Juyong Zhang, Sofien Bouaziz, Ligang Liu 0001 |
Comput. Graph. Forum | 3 |
| 2015 | Survey on sparsity in geometric modeling and processing
Linlin Xu, Juyong Zhang, Zhouwang Yang, Jiansong Deng, Falai Chen, Ligang Liu 0001 |
Graph. Model. | 3 |
| 2015 | Locating Facial Landmarks Using Probabilistic Random ForestabstractRandom forest is a useful tool for face alignment/tracking. The method of regressing local binary features learned from random forest has achieved state-of-the-art performance both in fitting accuracy and speed. Despite the great success of this method, it has certain weaknesses: the number of available local binary features is rather limited and is not optimal for face alignment; the binary features inevitably lead to serious jitter when tracking a video sequence. To address these problems, we propose learning probability features from probabilistic random forest (PRF). The proposed PRF is the same as standard random forest except that it models the probability of a sample belonging to the nodes of a tree. By using the probability features, our method significantly outperforms the state-of-the-art in terms of accuracy. It also achieves about 60 fps for locating a few facial landmarks. In addition, our method shows excellent stability in face tracking. Changwei Luo, Zengfu Wang, Shaobiao Wang, Juyong Zhang, Jun Yu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2015 | Variational Mesh Denoising Using Total Variation and Piecewise Constant Function SpaceabstractMesh surface denoising is a fundamental problem in geometry processing. The main challenge is to remove noise while preserving sharp features (such as edges and corners) and preventing generating false edges. We propose in this paper to combine total variation (TV) and piecewise constant function space for variational mesh denoising. We first give definitions of piecewise constant function spaces and associated operators. A variational mesh denoising method will then be presented by combining TV and piecewise constant function space. It is proved that, the solution of the variational problem (the key part of the method) is in some sense continuously dependent on its parameter, indicating that the solution is robust to small perturbations of this parameter. To solve the variational problem, we propose an efficient iterative algorithm (with an additional algorithmic parameter) based on variable splitting and augmented Lagrangian method, each step of which has closed form solution. Our denoising method is discussed and compared to several typical existing methods in various aspects. Experimental results show that our method outperforms all the compared methods for both CAD and non-CAD meshes at reasonable costs. It can preserve different levels of features well, and prevent generating false edges in most cases, even with the parameters evaluated by our estimation formulae. Huayan Zhang, Juyong Zhang, Jiansong Deng |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | Recovering Surface Details under General Unknown Illumination Using Shading and Coarse Multi-view StereoabstractSummary form only given. Reconstructing the shape of a 3D object from multi-view images under unknown, general illumination is a fundamental problem in computer vision and high quality reconstruction is usually challenging especially when high detail is needed. This paper presents a total variation (TV) based approach for recovering surface details using shading and multi-view stereo (MVS). Behind the approach are our two important observations: (1) the illumination over the surface of an object tends to be piecewise smooth and (2) the recovery of surface orientation is not sufficient for reconstructing geometry, which were previously overlooked. Thus we introduce TV to regularize the lighting and use visual hull to constrain partial vertices. The reconstruction is formulated as a constrained TV minimization problem that treats the shape and lighting as unknowns simultaneously. An augmented Lagrangian method is proposed to quickly solve the TV-minimization problem. As a result, our approach is robust, stable and is able to efficiently recover high quality of surface details even starting with a coarse MVS. These advantages are demonstrated by the experiments with synthetic and real world examples. Di Xu 0012, Qi Duan, Jianming Zheng, Juyong Zhang, Jianfei Cai 0001, Tat-Jen Cham |
CVPR | 4 |
| 2014 | Robust interactive image segmentation via iterative refinementabstractImage segmentation with user inputs gets more and more popular in recent years and always performs better compared with automatic methods. However, existing interactive image segmentation methods still might fail if the image contains messy textures, or the user inputs are sparse or at inappropriate locations. In this paper, we propose a novel iterative refinement framework which leads to robust segmentation performance even with sparse and improper input strokes. Specifically, a geodesic distance based energy is introduced and combined with convex active contour model, and an iterative seeds refinement technique is put forward to handle the sparse input problem. Extensive experiments using real world images, and segmentation benchmark dataset show that our proposed method has superior performance compared with representative state-of-the-art methods. Juyong Zhang, Yancheng Yuan, Shuyuan Zhu |
ICIP | 2 |
| 2014 | Analytical model for camera distance related 3D virtual view distortion estimationabstractWe propose an analytical model to estimate the depth-error-induced synthesis distortion in 3D video, taking into account the configuration of the cameras. In particular, the model mathematically relates the Distance between camera positions (reference view and virtual view) to the Virtual View Distortion (VVD), thus it is denoted as DVVD model. Specifically, the DVVD model accounts for two modules: distribution of disparity errors and shift-induced distortion. The former one is derived under a Laplacian distribution assumption of depth errors, and the latter one is estimated under a Quadratic model. We further propose a linear Steady-State model by performing Taylor series approximation of the DVVD model over a region of practical interest. Experiment results demonstrate that both the DVVD and Steady-State models are capable of estimating the relationship between VVD and the distance between virtual/reference view. Therefore, our model can effectively inform camera setup for capturing, in particular, the setup of the cameras in situation where depth information will be compressed subsequently. Yijian Xiang, Ngai-Man Cheung, Juyong Zhang, Lu Fang 0001 |
ICIP | 3 |
| 2014 | Iso-level tool path planning for free-form surfaces
Qiang Zou 0007, Juyong Zhang, Bailin Deng |
Comput. Aided Des. | 2 |
| 2014 | Robust surface reconstruction via dictionary learningabstractSurface reconstruction from point cloud is of great practical importance in computer graphics. Existing methods often realize reconstruction via a few phases with respective goals, whose integration may not give an optimal solution. In this paper, to avoid the inherent limitations of multi-phase processing in the prior art, we propose a unified framework that treats geometry and connectivity construction as one joint optimization problem. The framework is based on dictionary learning in which the dictionary consists of the vertices of the reconstructed triangular mesh and the sparse coding matrix encodes the connectivity of the mesh. The dictionary learning is formulated as a constrained ℓ 2,q -optimization (0 < q < 1), aiming to find the vertex position and triangulation that minimize an energy function composed of point-to-mesh metric and regularization. Our formulation takes many factors into account within the same framework, including distance metric, noise/outlier resilience, sharp feature preservation, no need to estimate normal, etc., thus providing a global and robust algorithm that is able to efficiently recover a piecewise smooth surface from dense data points with imperfections. Extensive experiments using synthetic models, real world models, and publicly available benchmark show that our method outperforms the state-of-the-art in terms of accuracy, robustness to noise and outliers, geometric feature and detail preservation, and mesh connectivity. Shiyao Xiong, Juyong Zhang, Jianmin Zheng, Jianfei Cai 0001, Ligang Liu 0001 |
ACM Trans. Graph. | 2 |
| 2014 | Local barycentric coordinatesabstractBarycentric coordinates yield a powerful and yet simple paradigm to interpolate data values on polyhedral domains. They represent interior points of the domain as an affine combination of a set of control points, defining an interpolation scheme for any function defined on a set of control points. Numerous barycentric coordinate schemes have been proposed satisfying a large variety of properties. However, they typically define interpolation as a combination ofallcontrol points. Thus alocalchange in the value at a single control point will create aglobalchange by propagation into the whole domain. In this context, we present a family oflocal barycentric coordinates(LBC), which select for each interior point a small set of control points and satisfy common requirements on barycentric coordinates, such as linearity, non-negativity, and smoothness. LBC are achieved through a convex optimization based on total variation, and provide a compact representation that reduces memory footprint and allows for fast deformations. Our experiments show that LBC provide more local and finer control on shape deformation than previous approaches, and lead to more intuitive deformation results. Juyong Zhang, Bailin Deng, Zishun Liu 0003, Giuseppe Patanè 0001, Sofien Bouaziz, Kai Hormann, Ligang Liu 0001 |
ACM Trans. Graph. | 1 |
| 2013 | Exploring Local Modifications for Constrained MeshesabstractAbstract Mesh editing under constraints is a challenging task with numerous applications in geometric modeling, industrial design, and architectural form finding. Recent methods support constraint‐based exploration of meshes with fixed connectivity, but commonly lack local control. Because constraints are often globally coupled, a local modification by the user can have global effects on the surface, making iterative design exploration and refinement difficult. Simply fixing a local region of interest a priori is problematic, as it is not clear in advance which parts of the mesh need to be modified to obtain an aesthetically pleasing solution that satisfies all constraints. We propose a novel framework for exploring local modifications of constrained meshes. Our solution consists of three steps. First, a user specifies target positions for one or more vertices. Our algorithm computes a sparse set of displacement vectors that satisfies the constraints and yields a smooth deformation. Then we build a linear subspace to allow realtime exploration of local variations that satisfy the constraints approximately. Finally, after interactive exploration, the result is optimized to fully satisfy the set of constraints. We evaluate our framework on meshes where each face is constrained to be planar. Bailin Deng, Sofien Bouaziz, Mario Deuss, Juyong Zhang, Yuliy Schwartzburg, Mark Pauly |
Comput. Graph. Forum | 4 |
| 2012 | Constrained active contours for boundary refinement in interactive image segmentationabstractThe state-of-the-art interactive image segmentation algorithms are often not able to produce accurate segmentation results with one-shot user input, and they frequently rely on laborious user editing to refine the segmentation boundary. In this paper, we propose a constrained active contour method for boundary refinement, which can be used to improve the segmentation results of many existing region-based interactive segmentation algorithms. Our constrained active contour model exhibits many desired properties for a good boundary refinement tool, including the robustness to user inputs, the ability to produce a smooth and accurate boundary contour, and the ability to handle topology changes. Experimental results show that the proposed refinement tool is highly effective and can significantly improve initial segmentation results without additional user inputs. Thi Nhat Anh Nguyen, Jianfei Cai 0001, Juyong Zhang, Jianmin Zheng |
ISCAS | 3 |
| 2012 | Robust Interactive Image Segmentation Using Convex Active ContoursabstractThe state-of-the-art interactive image segmentation algorithms are sensitive to the user inputs and often unable to produce an accurate boundary with a small amount of user interaction. They frequently rely on laborious user editing to refine the segmentation boundary. In this paper, we propose a robust and accurate interactive method based on the recently developed continuous-domain convex active contour model. The proposed method exhibits many desirable properties of an effective interactive image segmentation algorithm, including robustness to user inputs and different initializations, the ability to produce a smooth and accurate boundary contour, and the ability to handle topology changes. Experimental results on a benchmark data set show that the proposed tool is highly effective and outperforms the state-of-the-art interactive image segmentation algorithms. Thi Nhat Anh Nguyen, Jianfei Cai 0001, Juyong Zhang, Jianmin Zheng |
IEEE Trans. Image Process. | 3 |
| 2012 | Variational mesh decompositionabstractThe problem of decomposing a 3D mesh into meaningful segments (or parts) is of great practical importance in computer graphics. This article presents a variational mesh decomposition algorithm that can efficiently partition a mesh into a prescribed number of segments. The algorithm extends the Mumford-Shah model to 3D meshes that contains a data term measuring the variation within a segment using eigenvectors of a dual Laplacian matrix whose weights are related to the dihedral angle between adjacent triangles and a regularization term measuring the length of the boundary between segments. Such a formulation simultaneously handles segmentation and boundary smoothing, which are usually two separate processes in most previous work. The efficiency is achieved by solving the Mumford-Shah model through a saddle-point problem that is solved by a fast primal-dual method. A preprocess step is also proposed to determine the number of segments that the mesh should be decomposed into. By incorporating this preprocessing step, the proposed algorithm can automatically segment a mesh into meaningful parts. Furthermore, user interaction is allowed by incorporating the user's inputs into the variational model to reflect the user's special intention. Experimental results show that the proposed algorithm outperforms competitive segmentation methods when evaluated on the Princeton Segmentation Benchmark. Juyong Zhang, Jianmin Zheng, Jianfei Cai 0001 |
ACM Trans. Graph. | 1 |
| 2011 | Fast optimization for multichannel total variation minimization with non-quadratic fidelity
Juyong Zhang |
Signal Process. | 1 |
| 2011 | Interactive Mesh Cutting Using Constrained Random WalksabstractThis paper considers the problem of interactively finding the cutting contour to extract components from an existing mesh. First, we propose a constrained random walks algorithm that can add constraints to the random walks procedure and thus allows for a variety of intuitive user inputs. Second, we design an optimization process that uses the shortest graph path to derive a nice cut contour. Then a new mesh cutting algorithm is developed based on the constrained random walks plus the optimization process. Within the same computational framework, the new algorithm provides a novel user interface for interactive mesh cutting that supports three typical user inputs and also their combinations: 1) foreground/background seed inputs: the user draws strokes specifying seeds for “foreground” (i.e., the part to be cut out) and “background” (i.e., the rest); 2) soft constraint inputs: the user draws strokes on the mesh indicating the region which the cuts should be made nearby; and 3) hard constraint inputs: the marks which the cutting contour must pass. The algorithm uses feature sensitive metrics that are based on surface geometric properties and cognitive theory. The integration of the constrained random walks algorithm, the optimization process, the feature sensitive metrics, and the varieties of user inputs makes the algorithm intuitive, flexible, and effective as well. The experimental examples show that the proposed cutting method is fast, reliable, and capable of producing good results reflecting user intention and geometric attributes. Juyong Zhang, Jianmin Zheng, Jianfei Cai 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2010 | A diffusion approach to seeded image segmentationabstractSeeded image segmentation is a popular type of supervised image segmentation in computer vision and image processing. Previous methods of seeded image segmentation treat the image as a weighted graph and minimize an energy function on the graph to produce a segmentation. In this paper, we propose to conduct the seeded image segmentation according to the result of a heat diffusion process in which the seeded pixels are considered to be the heat sources and the heat diffuses on the image starting from the sources. After the diffusion reaches a stable state, the image is segmented based on the pixel temperatures. It is also shown that our proposed framework includes the RandomWalk algorithm for image segmentation as a special case which diffuses only along the two coordinate axes. To better control diffusion, we propose to incorporate the attributes (such as the geometric structure) of the image into the diffusion process, yielding an anisotropic diffusion method for image segmentation. The experiments show that the proposed anisotropic diffusion method usually produces better segmentation results. In particular, when the method is tested using the groundtruth dataset of Microsoft Research Cambridge (MSRC), an error rate of 4.42% can be achieved, which is lower than the reported error rates of other state-of-the-art algorithms. Juyong Zhang, Jianmin Zheng, Jianfei Cai 0001 |
CVPR | 1 |
| 2010 | Mesh Snapping: Robust Interactive Mesh Cutting Using Fast Geodesic Curvature FlowabstractAbstract This paper considers the problem of interactively finding the cutting contour to extract components from a given mesh. Some existing methods support cuts of arbitrary shape but require careful and tedious input from the user. Others need little user input however they are sensitive to user input and need a postprocessing step to smooth the generated jaggy cutting contours. The popular geometric snake can be used to optimize the cutting contour, but it cannot deal with the topology change. In this paper, we propose a geodesic curvature flow based framework to overcome all these problems. Since in many cases the meaningful cutting contour on a 3D mesh is locally shortest in the sense of some weighted curve length, the geodesic curvature flow is an ideal tool for our problem. It evolves the cutting contour to the nearby local minimum. We should mention that the previous numerical scheme, discretized geodesic curvature flow (dGCF) is too slow and has not been applied to mesh segmentation. With a careful observation to dGCF, we devise here a fast computation scheme called fast geodesic curvature flow (FGCF), which only needs to solve a smaller and easier problem. The initial cutting contour is generated by a variant of random walks algorithm, which is very fast and gives reasonable cutting result with little user input. Experiment results on the benchmark mesh segmentation data set show that our proposed framework is robust to user input and capable of producing good results reflecting geometric features and human shape perception. Juyong Zhang, Jianfei Cai 0001, Jianmin Zheng, Xue-Cheng Tai |
Comput. Graph. Forum | 1 |
| 2010 | Progressive Coding and Illumination and View Dependent Transmission of 3-D Meshes Using R-D OptimizationabstractFor transmitting complex 3-D models over bandwidth-limited networks, efficient mesh coding and transmission are indispensable. The state-of-the-art 3-D mesh transmission system employs a wavelet-based progressive mesh coder, which converts an irregular mesh into a semi-regular mesh and directly applies the zerotree-like image coders to compress the wavelet vectors, and view-dependent transmission, which saves the transmission bandwidth through only delivering the visible portions of a mesh model. We propose methods to improve both progressive mesh coding and transmission based on thorough rate-distortion analysis. In particular, by noticing that the dependency among the wavelet coefficients generated in remeshing is not being considered in the existing approaches, we propose to introduce a preprocessing step to scale up the wavelets so that the inherent dependency of wavelets can be truly understood by the zerotree-like image compression algorithms. The weights used in the scaling process are carefully designed through thoroughly analyzing the distortions of wavelets at different refinement levels. For the transmission part, we propose to incorporate the illumination effects into the existing view-depend progressive mesh transmission system to further improve the performance. We develop a novel distortion model that considers both illumination distortion and geometry distortion. Based on our proposed distortion model, given the viewing and lighting parameters, we are able to optimally allocate bits among different segments in real time. Simulation results show significant improvements in both progressive compression and transmission. Jianfei Cai 0001, Juyong Zhang, Jianmin Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Re-examination of applying wavelet based progressive image coder for 3D semi-regular mesh compressionabstractThe latest wavelet based 3D mesh coding schemes convert an irregular mesh into a semi-regular mesh and directly apply the zerotree-like image coders to compress the wavelet vectors generated in the remeshing process. The major problem of such type of approaches is that the particular properties of semi-regular meshes are not being considered in the zerotree-like image coders. In this paper, we propose an improved wavelet based 3D mesh coder. The basic idea is to introduce a preprocessing step to scale up the vector wavelets generated in remeshing so that the inherent dependency of wavelets can be truly understood by the zerotree-like image compression algorithms. The weights used in the scaling process are carefully designed through thoroughly analyzing the distortions of wavelets at different refinement levels. Experimental results show that our proposed mesh coder significantly outperforms the state-of-the-art wavelet based 3D mesh compression scheme. Juyong Zhang, Jianfei Cai 0001, Jianmin Zheng, Susanto Rahardja |
ICME | 1 |