VLDB 2026 Research / reviewers in the wild / expert
Hao Zhu 0004
dblp:10/3520-4
· DBLP profile ↗
44ranked-venue papers
7as first author
35since 2021 · last 2026
0000-0003-1596-4366ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 3 first-author · 28 since 2021Artificial intelligence and machine learning · 29 · 4 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Split-Layer: Enhancing Implicit Neural Representation by Maximizing the Dimensionality of Feature SpaceabstractImplicit neural representation (INR) models signals as continuous functions using neural networks, offering efficient and differentiable optimization for inverse problems across diverse disciplines. However, the representational capacity of INR—defined by the range of functions the neural network can characterize—is inherently limited by the low-dimensional feature space in conventional multilayer perceptron (MLP) architectures. While widening the MLP can linearly increase feature space dimensionality, it also leads to a quadratic growth in computational and memory costs. To address this limitation, we propose the split-layer, a novel reformulation of MLP construction. The split-layer divides each layer into multiple parallel branches and integrates their outputs via Hadamard product, effectively constructing a high-degree polynomial space. This approach significantly enhances INR’s representational capacity by expanding the feature space dimensionality without incurring prohibitive computational overhead. Extensive experiments demonstrate that the split-layer substantially improves INR performance, surpassing existing methods across multiple tasks, including 2D image fitting, 2D CT reconstruction, 3D shape representation, and 5D novel view synthesis. Zhicheng Cai, Hao Zhu 0004, Linsen Chen, Qiu Shen, Xun Cao |
AAAI | 2 |
| 2026 | RHINO: regularizing the hash-based implicit neural representation
Hao Zhu 0004, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
Sci. China Inf. Sci. | 1 |
| 2026 | X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video RenderingabstractWe present X2Video, the first diffusion model for rendering photorealistic videos guided by a sequence of intrinsic channels including albedo, normal, roughness, metallicity, and irradiance, while supporting intuitive multi-modal controls with reference images and text prompts for both global and local regions. The intrinsic guidance allows accurate manipulation of color, material, geometry, and lighting, while reference images and text prompts provide intuitive adjustments in the absence of intrinsic information. To enable these functionalities, we extend the intrinsic-guided image generation model XRGB to video generation by employing a novel and efficient Hybrid Self-Attention, which ensures temporal consistency across video frames and also enhances fidelity to reference images. We further develop a Masked Cross-Attention to disentangle global and local text prompts, applying them effectively onto respective local and global regions. For generating long videos, our novel Recursive Sampling method incorporates progressive frame sampling, combining keyframe prediction and frame interpolation to maintain long-range temporal consistency while preventing error accumulation. To support the training of X2Video, we assembled a video dataset named InteriorVideo, featuring 1,154 rooms from 295 interior scenes, complete with reliable ground-truth intrinsic channel sequences and smooth camera trajectories. Both qualitative and quantitative evaluations demonstrate that X2Video can produce long, temporally consistent, and photorealistic videos guided by intrinsic conditions. Additionally, X2Video effectively accommodates multi-modal controls with reference images, global and local text prompts, and simultaneously supports editing on color, material, geometry, and lighting through parametric tuning. Upon acceptance, we will publicly release our model and dataset. Zhitong Huang, Mohan Zhang, Renhan Wang, Rui Tang 0015, Hao Zhu 0004, Jing Liao 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid PriorabstractAudio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, facial motion, head pose generation, and video quality. However, no model has yet led or tied on all these metrics due to the one-to-many mapping between audio and motion. In this paper, we propose VividTalk, a two-stage generic framework that supports generating high-visual quality talking head videos with all the above properties. Specifically, in the first stage, we map the audio to mesh by learning two motions, including non-rigid facial motion and rigid head motion. For facial motion, both blendshape and vertex are adopted as the intermediate representation to maximize the representation ability of the model. For head motion, a novel learnable head pose codebook with a two-phase training mechanism is proposed. In the second stage, we proposed a dual branch motion-vae and a generator to transform the meshes into dense motion and synthesize high-quality video frame-by-frame. Extensive experiments show that the proposed VividTalk can generate high-visual quality talking head videos with lip-sync and realistic enhanced by a large margin, and outperforms previous state-of-the-art works in objective and subjective comparisons. The code will be publicly released upon publication. Xusen Sun, Longhao Zhang, Hao Zhu 0004, Peng Zhang 0080, Bang Zhang, Xinya Ji, Kangneng Zhou, Daiheng Gao, Liefeng Bo, Xun Cao |
3DV | 3 |
| 2025 | From 2D CAD Drawings to 3D Parametric Models: A Vision-Language ApproachabstractIn this paper, we present CAD2Program, a new method for reconstructing 3D parametric models from 2D CAD drawings. Our proposed method is inspired by recent successes in vision-language models (VLMs), and departs from traditional methods which rely on task-specific data representations and/or algorithms. Specifically, on the input side, we simply treat the 2D CAD drawing as a raster image, regardless of its original format, and encode the image with a standard ViT model. We show that such an encoding scheme achieves competitive performance against existing methods that operate on vector-graphics inputs, while imposing substantially fewer restrictions on the 2D drawings. On the output side, our method auto-regressively predicts a general-purpose language describing 3D parametric models in text form. Compared to other sequence modeling methods for CAD which use domain-specific sequence representations with fixed-size slots, our text-based representation is more flexible, and can be easily extended to arbitrary geometric entities and semantic or functional properties. Experimental results on a large-scale dataset of cabinet models demonstrate the effectiveness of our method. Xilin Wang, Jia Zheng 0002, Yuanchao Hu, Hao Zhu 0004, Zihan Zhou 0001 |
AAAI | 4 |
| 2025 | FATE: Full-head Gaussian Avatar with Textural Editing from Monocular VideoabstractReconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities, notable challenges remain, including incomplete reconstruction and inefficient Gaussian representation. To address these challenges, we introduce FATE — a novel method for reconstructing an editable full-head avatar from a single monocular video. FATE integrates a sampling-based densification strategy to ensure optimal positional distribution of points, improving rendering efficiency. A neural baking technique is introduced to convert discrete Gaussian representations into continuous attribute maps, facilitating intuitive appearance editing. Furthermore, we propose a universal completion framework to recover non-frontal appearance, culminating in a 360° -renderable 3D head avatar. FATE outperforms previous approaches in both qualitative and quantitative evaluations, achieving state-of-the-art performance. To the best of our knowledge, FATE is the first animatable and 360° full-head monocular reconstruction method for a 3D head avatar. Project page and code are available at this link. Zhiyang Liang 0002, Dongfang Hu, Yao Yao 0008, Xun Cao, Hao Zhu 0004 |
CVPR | 8 |
| 2025 | Mitigating Ambiguities in 3D Classification with Gaussian Splattingabstract3D classification with point cloud input is a fundamental problem in 3D vision. However, due to the discrete nature and the insufficient material description of point cloud representations, there are ambiguities in distinguishing wire-like and flat surfaces, as well as transparent or reflective objects. To address these issues, we propose Gaussian Splatting (GS) point cloud-based 3D classification. We find that the scale and rotation coefficients in the GS point cloud help characterize surface types. Specifically, wire-like surfaces consist of multiple slender Gaussian ellipsoids, while flat surfaces are composed of a few flat Gaussian ellipsoids. Additionally, the opacity in the GS point cloud represents the transparency characteristics of objects. As a result, ambiguities in point cloud-based 3D classification can be mitigated utilizing GS point cloud as input. To verify the effectiveness of GS point cloud input, we construct the first real-world GS point cloud dataset in the community, which includes 20 categories with 200 objects in each category. Experiments not only validate the superiority of GS point cloud input, especially in distinguishing ambiguous objects, but also demonstrate the generalization ability across different classification methods. Our project page: https://ruiqi-nju.github.io/MACGS. Hao Zhu 0004, Qi Zhang 0029, Xun Cao, Zhan Ma 0001 |
CVPR | 2 |
| 2025 | IDOL: Instant Photorealistic 3D Human Creation from a Single ImageabstractCreating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the perspectives of dataset, model, and representation. First, we introduce a large-scale HUman-centric GEnerated dataset, HuGe100K, consisting of 100K diverse, photorealistic sets of human images. Each set contains 24-view frames in specific human poses, generated using a pose-controllable image-to-multi-view model. Next, leveraging the diversity in views, poses, and appearances within HuGe100K, we develop a scalable feed-forward transformer model to predict a 3D human Gaussian representation in a uniform space from a given human image. This model is trained to disentangle human pose, body shape, clothing geometry, and texture. The estimated Gaussians can be animated without post-processing. We conduct comprehensive experiments to validate the effectiveness of the proposed dataset and method. Our model demonstrates the ability to efficiently reconstruct photorealistic humans at 1K resolution from a single input image using a single GPU instantly. Additionally, it seamlessly supports various applications, as well as shape and texture editing tasks. Yiyu Zhuang, Jiaxi Lv, Hao Wen 0005, Qing Shuai, Ailing Zeng, Hao Zhu 0004, Shifeng Chen, Yujiu Yang 0001, Xun Cao, Wei Liu 0005 |
CVPR | 6 |
| 2025 | Tera: Rethinking Text-Guided Realistic 3D Avatar GenerationabstractIn this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two-stage training strategy for learning a native 3D avatar generative model. Initially, we distill a decoder to derive a structured latent space from a large human reconstruction model. Subsequently, a text-controlled latent diffusion model is trained to generate photorealistic 3D human avatars within this latent space. TeRA enhances the model performance by eliminating slow iterative optimization and enables text-based partial customization through a structured 3D human representation. Experiments have proven our approach's superiority over previous text-to-avatar generative models in subjective and objective evaluation. Yiyu Zhuang, Yifei Zeng, Xun Cao, Xinxin Zuo, Hao Zhu 0004 |
ICCV | 8 |
| 2025 | WIPES: Wavelet-based Visual PrimitivesabstractPursuing a continuous visual representation that offers flexible frequency modulation and fast rendering speed has recently garnered increasing attention in the fields of 3D vision and graphics. However, existing representations often rely on frequency guidance or complex neural network decoding, leading to spectrum loss or slow rendering. To address these limitations, we propose WIPES, a universal Wavelet-based vIsual PrimitivES for representing multi-dimensional visual signals. Building on the spatial-frequency localization advantages of wavelets, WIPES effectively captures both the low-frequency "forest" and the high-frequency "trees." Additionally, we develop a wavelet-based differentiable rasterizer to achieve fast visual rendering. Experimental results on various visual tasks, including 2D image representation, 5D static and 6D dynamic novel view synthesis, demonstrate that WIPES, as a visual primitive, offers higher rendering quality and faster inference than INR-based methods, and outperforms Gaussian-based representations in rendering quality. Hao Zhu 0004, Delong Wu, Linchao Bao, Xun Cao, Zhan Ma 0001 |
ICCV | 2 |
| 2025 | Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image AnimationabstractRecent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities.First, we extend the method to produce long-duration videos. To address substantial challenges such as appearance drift and temporal artifacts, we investigate augmentation strategies within the image space of conditional motion frames. Specifically, we introduce a patch-drop technique augmented with Gaussian noise to enhance visual consistency and temporal coherence over long duration.Second, we achieve 4K resolution portrait video generation. To accomplish this, we implement vector quantization of latent codes and apply temporal alignment techniques to maintain coherence across the temporal dimension. By integrating a high-quality decoder, we realize visual synthesis at 4K resolution.Third, we incorporate adjustable semantic textual labels for portrait expressions as conditional inputs. This extends beyond traditional audio cues to improve controllability and increase the diversity of the generated content. To the best of our knowledge, Hallo2, proposed in this paper, is the first method to achieve 4K resolution and generate hour-long, audio-driven portrait image animations enhanced with textual prompts. We have conducted extensive experiments to evaluate our method on publicly available datasets, including HDTF, CelebV, and our introduced ''Wild'' dataset. The experimental results demonstrate that our approach achieves state-of-the-art performance in long-duration portrait video animation, successfully generating rich and controllable content at 4K resolution for duration extending up to tens of minutes. Jiahao Cui 0003, Yao Yao 0008, Hao Zhu 0004, Hanlin Shang, Kaihui Cheng, Hang Zhou 0009, Siyu Zhu 0001, Jingdong Wang 0001 |
ICLR | 4 |
| 2025 | SpatialLM: Training Large Language Models for Structured Indoor ModelingabstractSpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs include architectural elements like walls, doors, windows, and oriented object boxes with their semantic categories. Unlike previous methods which exploit task-specific network designs, our model adheres to the standard multimodal LLM architecture and is fine-tuned directly from open-source LLMs.
To train SpatialLM, we collect a large-scale, high-quality synthetic dataset consisting of the point clouds of 12,328 indoor scenes (54,778 rooms) with ground-truth 3D annotations, and conduct a careful study on various modeling and training decisions. On public benchmarks, our model gives state-of-the-art performance in layout estimation and competitive results in 3D object detection. With that, we show a feasible path for enhancing the spatial understanding capabilities of modern LLMs for applications in augmented reality, embodied robotics, and more. Yongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng 0002, Rui Tang 0015, Hao Zhu 0004, Ping Tan 0002, Zihan Zhou 0001 |
NeurIPS | 6 |
| 2025 | NeLiF: Neural Lighting Function Generation for Real-Time Indoor RenderingabstractRecent advances in neural rendering have mainly focused on modeling radiance fields with neural representations, often overlooking the underlying mechanisms for producing various lighting effects, and consequently leading to the limited adaptability to dynamic scenes. Lighting effects, such as highlights, shadows, and indirect illuminations, are typically computed using physically-based rendering methods like path tracing, which can be computationally intensive for complex indoor luminaires. Although several recent studies have aimed to model global illumination effects with neural representations, they commonly suffer from long training times or poor generalizability to new scenes. Addressing these challenges, this work presents a novel neural lighting function generation model capable of synthesizing diverse lighting effects in real time for unseen dynamic scenes and complex indoor luminaires, achieving results comparable to state-of-the-art rendering pipelines. Our model operates in two stages. First, multi-view observation images of the luminaire are captured to encode a compact, scene-independent 3D neural lighting field. Subsequently, light information is sampled from this neural lighting field and integrated with G-buffers and shadow clues to produce the shading results. In parallel, we employ a state-of-the-art generative model together with our training-free Inverse HDR Splatting module to generate HDR 3D Gaussians representing the luminaire. This strategy capitalizes on the powerful generalization capabilities of advanced generative models, enabling efficient and accurate appearance reconstruction for a diverse range of complex luminaires. In our experiments, the model trained on a dataset of 10,000 modern indoor scenes and thousands of illuminations demonstrates strong generalizability, high efficiency, and visually convincing results across a wide range of test scenes, highlighting its potential as a practical and flexible solution for high-fidelity, real-time neural indoor rendering. Hongtao Sheng, Yuchi Huo, Chuankun Zheng, Guangzhi Han, Yifan Peng 0001, Bin Zang, Hao Zhu 0004, Rui Tang 0015, Rui Wang 0004, Hujun Bao |
SIGGRAPH Asia | 8 |
| 2025 | Sketch2PoseNet: Efficient and Generalized Sketch to 3D Human Pose Predictionabstract3D human pose estimation from sketches has broad applications in computer animation and film production. Unlike traditional human pose estimation, this task presents unique challenges due to the abstract and disproportionate nature of sketches. Previous sketch-to-pose methods, constrained by the lack of large-scale sketch-3D pose annotations, primarily relied on optimization with heuristic rules—an approach that is both time-consuming and limited in generalizability. To address these challenges, we propose a novel approach leveraging a "learn from synthesis" strategy. Firstly, a diffusion model is learned to synthesize sketch images from 2D poses projected from 3D human poses, mimicking disproportionate human structures in sketches. This process enables the creation of a synthetic dataset, SKEP-120K, consisting of 120k accurate sketch-3D pose annotation pairs across various sketch styles. Building on this synthetic dataset, we introduce an end-to-end data-driven framework for estimating human poses and shapes from diverse sketch styles. Our framework combines existing 2D pose detectors and generative diffusion priors for sketch feature extraction with a feed-forward neural network for efficient 2D pose estimation. Multiple heuristic loss functions have been incorporated to guarantee geometric coherence between the derived 3D poses and the detected 2D poses while preserving accurate self-contacts. Qualitative, quantitative, and subjective evaluations collectively affirm that our proposed model substantially surpasses previous ones in both estimation accuracy and speed for sketch-to-pose tasks. Yiyu Zhuang, Xun Cao, Chuan Guo 0002, Xinxin Zuo, Hao Zhu 0004 |
SIGGRAPH Asia | 7 |
| 2024 | A Pre-convolved Representation for Plug-and-Play Neural Illumination FieldsabstractRecent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light modeling without a solid basis, can lead to an undesirable decomposition between lighting and material. To address this, we propose a fully differentiable framework named Neural Illumination Fields (NeIF) that uses radiance fields as a lighting model to handle complex lighting in a physically based way. Together with integral lobe encoding for roughness-adaptive specular lobe and leveraging the pre-convolved background for accurate decomposition, the proposed method represents a significant step towards integrating physically based rendering into the NeRF representation. The experiments demonstrate the superior performance of novel-view rendering compared to previous works, and the capability to re-render objects under arbitrary NeRF-style environments opens up exciting possibilities for bridging the gap between virtual and real-world scenes. Yiyu Zhuang, Qi Zhang 0029, Xuan Wang 0009, Hao Zhu 0004, Xiaoyu Li 0002, Ying Shan, Xun Cao |
AAAI | 4 |
| 2024 | Batch Normalization Alleviates the Spectral Bias in Coordinate NetworksabstractRepresenting signals using coordinate networks domi-nates the area of inverse problems recently, and is widely applied in various scientific computing tasks. Still, there exists an issue of spectral bias in coordinate networks, lim-iting the capacity to learn high-frequency components. This problem is caused by the pathological distribution of the neural tangent kernel's (NTK's) eigenvalues of coordinate networks. We find that, this pathological distribution could be improved using the classical batch normalization (BN), which is a common deep learning technique but rarely used in coordinate networks. BN greatly reduces the maximum and variance of NTK's eigenvalues while slightly modifies the mean value, considering the max eigenvalue is much larger than the most, this variance change results in a shift of eigenvalues' distribution from a lower one to a higher one, therefore the spectral bias could be alleviated (see Fig. 1). This observation is substantiated by the significant improvements of applying BN-based coordinate networks to various tasks, including the image compression, computed tomography reconstruction, shape representation, magnetic resonance imaging and novel view synthesis. Zhicheng Cai, Hao Zhu 0004, Qiu Shen, Xun Cao |
CVPR | 2 |
| 2024 | Relightable 3D Gaussians: Realistic Point Cloud Relighting with BRDF Decomposition and Ray Tracing
Jian Gao 0009, Chun Gu, Youtian Lin, Zhihao Li 0002, Hao Zhu 0004, Xun Cao, Li Zhang 0040, Yao Yao 0008 |
ECCV (45) | 5 |
| 2024 | EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head
Qianyun He, Xinya Ji, Yuanxun Lu, Zhengyu Diao, Linjia Huang, Yao Yao 0008, Siyu Zhu 0001, Zhan Ma 0001, Songcen Xu, Zixiao Zhang, Xun Cao, Hao Zhu 0004 |
ECCV (57) | 14 |
| 2024 | Head360: Learning a Parametric 3D Full-Head for Free-View Synthesis in 360$^\circ $
Yuxiao He, Yiyu Zhuang, Yao Yao 0008, Siyu Zhu 0001, Xiaoyu Li 0002, Qi Zhang 0029, Xun Cao, Hao Zhu 0004 |
ECCV (56) | 9 |
| 2024 | STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians
Yifei Zeng, Yanqin Jiang, Siyu Zhu 0001, Yuanxun Lu, Youtian Lin, Hao Zhu 0004, Weiming Hu 0004, Xun Cao, Yao Yao 0008 |
ECCV (36) | 6 |
| 2024 | Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Zilong Dong, Xun Cao, Yao Yao 0008, Hao Zhu 0004, Siyu Zhu 0001 |
ECCV (55) | 8 |
| 2024 | Neural Global Illumination via Superposed Deformable Feature Fields
Chuankun Zheng, Yuchi Huo, Hongxiang Huang, Hongtao Sheng, Junrong Huang, Rui Tang 0015, Hao Zhu 0004, Rui Wang 0004, Hujun Bao |
SIGGRAPH Asia | 7 |
| 2024 | Compressing 3D Gaussian Splatting via a Generalizable Neural CoderabstractAs a promising technique for 3D representation, 3D Gaussian Splatting (3DGS) offers fast rendering speed and high fidelity while generating large data volumes. This challenges storage and transmission, so an efficient compression solution is required. Existing implicit methods require pre-scene optimization (online), leading to a long optimization time. By contrast, this paper regards the 3DGS as a point cloud and pre-trains a generalizable (offline) neural coder for compression. After obtaining the 3DGS representation, we focus on the data compression process, which is friendly to applications already equipped with a PCC codec. The neural coder employed is extended from a typical AIbased point cloud compression method, which uses a multiscale and multistage framework to exploit spatial correlations across scales and stages for conditional coding. Experimental results show that our method significantly outperforms existing 3DGS representations without compromising fidelity, achieving more than 39× and 6.8× compression ratio compared to the original 3DGS and SOTA Scaffold-GS, respectively. More importantly, our approach does not require additional time to optimize the compression model. Junteng Zhang, Tong Chen 0004, Hao Zhu 0004, Dandan Ding, Zhan Ma 0001 |
VCIP | 3 |
| 2023 | RAFaRe: Learning Robust and Accurate Non-parametric 3D Face Reconstruction from Pseudo 2D&3D PairsabstractWe propose a robust and accurate non-parametric method for single-view 3D face reconstruction (SVFR). While tremendous efforts have been devoted to parametric SVFR, a visible gap still lies between the result 3D shape and the ground truth. We believe there are two major obstacles: 1) the representation of the parametric model is limited to a certain face database; 2) 2D images and 3D shapes in the fitted datasets are distinctly misaligned. To resolve these issues, a large-scale pseudo 2D&3D dataset is created by first rendering the detailed 3D faces, then swapping the face in the wild images with the rendered face. These pseudo 2D&3D pairs are created from publicly available datasets which eliminate the gaps between 2D and 3D data while covering diverse appearances, poses, scenes, and illumination. We further propose a non-parametric scheme to learn a well-generalized SVFR model from the created dataset, and the proposed hierarchical signed distance function turns out to be effective in predicting middle-scale and small-scale 3D facial geometry. Our model outperforms previous methods on FaceScape-wild/lab and MICC benchmarks and is well generalized to various appearances, poses, expressions, and in-the-wild environments. The code is released at https://github.com/zhuhao-nju/rafare. Longwei Guo, Hao Zhu 0004, Yuanxun Lu, Menghua Wu, Xun Cao |
AAAI | 2 |
| 2023 | High-fidelity 3D Face Generation from Natural Language DescriptionsabstractSynthesizing high-quality 3D face models from natural language descriptions is very valuable for many applications, including avatar creation, virtual reality, and telepresence. However, little research ever tapped into this task. We argue the major obstacle lies in 1) the lack of high-quality 3D face data with descriptive text annotation, and 2) the complex mapping relationship between descriptive language space and shape/appearance space. To solve these problems, we build Describe3D dataset, the first large-scale dataset with fine-grained text descriptions for text-to-3D face generation task. Then we propose a two-stage framework to first generate a 3D face that matches the concrete descriptions, then optimize the parameters in the 3D shape and texture space with abstract description to refine the 3D face model. Extensive experimental results show that our method can produce a faithful 3D face that conforms to the input descriptions with higher accuracy and quality than previous methods. The code and Describe3D dataset are released at https://github.com/zhuhao-nju/describe3D. Menghua Wu, Hao Zhu 0004, Linjia Huang, Yiyu Zhuang, Yuanxun Lu, Xun Cao |
CVPR | 2 |
| 2023 | DINER: Disorder-Invariant Implicit Neural RepresentationabstractImplicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the capacity of INR is limited by the spectral bias in the network training. In this paper, we find that such a frequency-related problem could be largely solved by re-arranging the coordinates of the input signal, for which we propose the disorder-invariant implicit neural representation (DINER) by augmenting a hash-table to a traditional INR backbone. Given discrete signals sharing the same histogram of attributes and different arrangement orders, the hash-table could project the coordinates into the same distribution for which the mapped signal can be better modeled using the subsequent INR network, leading to significantly alleviated spectral bias. Experiments not only reveal the generalization of the DINER for different INR backbones (MLP vs. SIREN) and various tasks (image/video representation, phase retrieval, and refractive index recovery) but also show the superiority over the state-of-the-art algorithms both in quality and speed. Project page: https://ezio77.github.io/DINER-website/ Shaowen Xie, Hao Zhu 0004, Zhen Liu 0031, Qi Zhang 0029, Xun Cao, Zhan Ma 0001 |
CVPR | 2 |
| 2023 | Efficient FPGA-Based Accelerator of the L-BFGS Algorithm for IoT ApplicationsabstractThe Internet of Things (IoT)-centric applications, such as augmented reality and self-driven cars, require real-time task processing, large bandwidth, and low data transmission latency. FPGA-based edge computing is considered an effective solution to tackle these challenges. As an excellent tool in these applications, nonlinear optimization methods involve computation-intensive and data-dependency operations leading to limited real-time applications. The limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm ranks among the most efficient algorithms for large-scale optimization problems. In this paper, we propose, for the first time, a high-parallel FPGA-based architecture for the two key parts of the L-BFGS algorithm: the search direction computation and line searching. Compared with the implementation on CPU, the search direction computation and line searching implementation on FPGA achieve$\mathbf{39.73}\times$and$\mathbf{5.50}\times$speedups, respectively. Compared with the straightforward implementation on GPU, the search direction computation on FPGA obtains a speedup of$\mathbf{31.03}\times$. Huiyang Xiong, Bohang Xiong, Jing Tian 0004, Hao Zhu 0004, Zhongfeng Wang 0001 |
ISCAS | 5 |
| 2023 | Anti-Aliased Neural Implicit Surfaces with Encoding Level of DetailabstractWe present LoD-NeuS, an efficient neural representation for high-frequency geometry detail recovery and anti-aliased novel view rendering. Drawing inspiration from voxel-based representations with the level of detail (LoD), we introduce a multi-scale tri-plane-based scene representation that is capable of capturing the LoD of the signed distance function (SDF) and the space radiance. Our representation aggregates space features from a multi-convolved featurization within a conical frustum along a ray and optimizes the LoD feature volume through differentiable rendering. Additionally, we propose an error-guided sampling strategy to guide the growth of the SDF during the optimization. Both qualitative and quantitative evaluations demonstrate that our method achieves superior surface reconstruction and photorealistic view synthesis compared to state-of-the-art approaches. Yiyu Zhuang, Qi Zhang 0029, Hao Zhu 0004, Yao Yao 0008, Xiaoyu Li 0002, Yan-Pei Cao 0001, Ying Shan, Xun Cao |
SIGGRAPH Asia | 4 |
| 2023 | Pyramid NeRF: Frequency Guided Fast Radiance Field Optimization
Junyu Zhu, Hao Zhu 0004, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
Int. J. Comput. Vis. | 2 |
| 2023 | FaceScape: 3D Facial Dataset and Benchmark for Single-View 3D Face ReconstructionabstractIn this article, we present a large-scale detailed 3D face dataset, FaceScape, and the corresponding benchmark to evaluate single-view facial 3D reconstruction. By training on FaceScape data, a novel algorithm is proposed to predict elaborate riggable 3D face models from a single image input. FaceScape dataset releases 16,940 textured 3D faces, captured from 847 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniform. These fine 3D facial models can be represented as a 3D morphable model for coarse shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different from most previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. We also use FaceScape data to generate the in-the-wild and in-the-lab benchmark to evaluate recent methods of single-view face reconstruction. The accuracy is reported and analyzed on the dimensions of camera pose and focal length, which provides a faithful and comprehensive evaluation and reveals new challenges. The unprecedented dataset, benchmark, and code have been released to the public for research purpose. Hao Zhu 0004, Longwei Guo, Mingkai Huang, Menghua Wu, Qiu Shen, Ruigang Yang, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Detailed Facial Geometry Recovery from Multi-View Images by Learning an Implicit FunctionabstractRecovering detailed facial geometry from a set of calibrated multi-view images is valuable for its wide range of applications. Traditional multi-view stereo (MVS) methods adopt an optimization-based scheme to regularize the matching cost. Recently, learning-based methods integrate all these into an end-to-end neural network and show superiority of efficiency. In this paper, we propose a novel architecture to recover extremely detailed 3D faces within dozens of seconds. Unlike previous learning-based methods that regularize the cost volume via 3D CNN, we propose to learn an implicit function for regressing the matching cost. By fitting a 3D morphable model from multi-view images, the features of multiple images are extracted and aggregated in the mesh-attached UV space, which makes the implicit function more effective in recovering detailed facial shape. Our method outperforms SOTA learning-based MVS in accuracy by a large margin on the FaceScape dataset. The code and data are released in https://github.com/zhuhao-nju/mvfr. Yunze Xiao, Hao Zhu 0004, Zhengyu Diao, Xiangju Lu, Xun Cao |
AAAI | 2 |
| 2022 | MoFaNeRF: Morphable Facial Neural Radiance Field
Yiyu Zhuang, Hao Zhu 0004, Xusen Sun, Xun Cao |
ECCV (3) | 2 |
| 2022 | Migrating Face Swap to Mobile Devices: A Lightweight Framework and a Supervised Training SolutionabstractExisting face swap methods rely heavily on large-scale networks for adequate capacity to generate visually plausible results, which inhibits its applications on resource-constraint platforms. In this work, we propose MobileFSGAN, a novel lightweight GAN for face swap that can run on mobile devices with much fewer parameters while achieving competitive performance. A lightweight encoder-decoder structure is designed especially for image synthesis tasks, which is only 10.2MB and can run on mobile devices at a real-time speed. To tackle the unstability of training such a small network, we construct the FSTriplets dataset utilizing facial attribute editing techniques. FSTriplets provides source-target-result training triplets, yielding pixel-level labels thus for the first time making the training process supervised. We also designed multi-scale gradient losses for efficient back-propagation, resulting in faster and better convergence. Experimental results show that our model reaches comparable performance towards state-of-the-art methods, while significantly reducing the number of network parameters. Codes and the dataset have been released11https://githuh.com/HoiM/MobileFSGAN. Haiming Yu, Hao Zhu 0004, Xiangju Lu |
ICME | 2 |
| 2022 | Detailed Avatar Recovery From Single ImageabstractThis paper presents a novel framework to recover detailed avatar from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, texture, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric-based template that lacks the surface details. As such resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of the parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. Our method can restore detailed human body shapes with complete textures beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Lossy Point Cloud Geometry Compression via End-to-End LearningabstractThis paper presents a novel end-to-endLearned Point Cloud Geometry Compression(a.k.a., Learned-PCGC) system, leveraging stacked Deep Neural Networks (DNN) based Variational AutoEncoder (VAE) to efficiently compress the Point Cloud Geometry (PCG). In this systematic exploration, PCG is first voxelized, and partitioned into non-overlapped 3D cubes, which are then fed into stacked 3D convolutions for compact latent feature and hyperprior generation. Hyperpriors are used to improve the conditional probability modeling of entropy-coded latent features. A Weighted Binary Cross-Entropy (WBCE) loss is applied in training while an adaptive thresholding is used in inference to remove false voxels and reduce the distortion. Objectively, our method exceeds the Geometry-based Point Cloud Compression (G-PCC) algorithm standardized by the Moving Picture Experts Group (MPEG) with a significant performance margin, e.g., at least 60% BD-Rate (Bjöntegaard Delta Rate) savings, using common test datasets, and other public datasets. Subjectively, our method has presented better visual quality with smoother surface reconstruction and appealing details, in comparison to all existing MPEG standard compliant PCC methods. Our method requires about 2.5 MB parameters in total, which is a fairly small size for practical implementation, even on embedded platform. Additional ablation studies analyze a variety of aspects (e.g., thresholding, kernels, etc) to examine the generalization, and application capacity of our Learned-PCGC. We would like to make all materials publicly accessible athttps://njuvision.github.io/PCGCv1/for reproducible research. Jianqiang Wang 0006, Hao Zhu 0004, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Speech2Video Synthesis with 3D Skeleton Regularization and Expressive Body Poses
Miao Liao, Peng Wang 0001, Hao Zhu 0004, Xinxin Zuo, Ruigang Yang |
ACCV (5) | 4 |
| 2020 | Self-Supervised Human Depth Estimation From Monocular VideosabstractPrevious methods on estimating detailed human depth often require supervised training with ‘ground truth’ depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes training data collection simple and improves the generalization of the learned network. The self-supervised learning is achieved by minimizing a photo-consistency loss, which is evaluated between a video frame and its neighboring frames warped according to the estimated depth and the 3D non-rigid motion of the human body. To solve this non-rigid motion, we first estimate a rough SMPL model at each video frame and compute the non-rigid body motion accordingly, which enables self-supervised learning on estimating the shape details. Experiments demonstrate that our method enjoys better generalization, and performs much better on data in the wild. Feitong Tan, Hao Zhu 0004, Zhaopeng Cui, Siyu Zhu 0001, Marc Pollefeys, Ping Tan 0002 |
CVPR | 2 |
| 2020 | FaceScape: A Large-Scale High Quality 3D Face Dataset and Detailed Riggable 3D Face PredictionabstractIn this paper, we present a large-scale detailed 3D face dataset, FaceScape, and propose a novel algorithm that is able to predict elaborate riggable 3D face models from a single image input. FaceScape dataset provides 18,760 textured 3D faces, captured from 938 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniformed. These fine 3D facial models can be represented as a 3D morphable model for rough shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different than the previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. The unprecedented dataset and code will be released to public for research purpose. Hao Zhu 0004, Mingkai Huang, Qiu Shen, Ruigang Yang, Xun Cao |
CVPR | 2 |
| 2020 | Interactive free-viewpoint video generationabstractFree-viewpoint video (FVV) is processed video content in which viewers can freely select the viewing position and angle. FVV delivers an improved visual experience and can also help synthesize special effects and virtual reality content. In this paper, a complete FVV system is proposed to interactively control the viewpoints of video relay programs through multimedia terminals such as computers and tablets. The hardware of the FVV generation system is a set of synchronously controlled cameras, and the software generates videos in novel viewpoints from the captured video using view interpolation. The interactive interface is designed to visualize the generated video in novel viewpoints and enable the viewpoint to be changed interactively. Experiments show that our system can synthesize plausible videos in intermediate viewpoints with a view range of up to 180°. Hao Zhu 0004, Wei Li 0111, Xun Cao, Ruigang Yang |
Virtual Real. Intell. Hardw. | 3 |
| 2019 | Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh DeformationabstractThis paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template that lacks the surface details. As such the resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. We are able to restore detailed human body shapes beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. The code is available in https://github.com/zhuhao-nju/hmd.git. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
CVPR | 1 |
| 2018 | View Extrapolation of Human Body From a Single ImageabstractWe study how to synthesize novel views of human body from a single image. Though recent deep learning based methods work well for rigid objects, they often fail on objects with large articulation, like human bodies. The core step of existing methods is to fit a map from the observable views to novel views by CNNs; however, the rich articulation modes of human body make it rather challenging for CNNs to memorize and interpolate the data well. To address the problem, we propose a novel deep learning based pipeline that explicitly estimates and leverages the geometry of the underlying human body. Our new pipeline is a composition of a shape estimation network and an image generation network, and at the interface a perspective transformation is applied to generate a forward flow for pixel value transportation. Our design is able to factor out the space of data variation and makes learning at each step much easier. Empirically, we show that the performance for pose-varying objects can be improved dramatically. Our method can also be applied on real data captured by 3D sensors, and the flow generated by our methods can be used for generating high quality results in higher resolution. Hao Zhu 0004, Peng Wang 0001, Xun Cao, Ruigang Yang |
CVPR | 1 |
| 2017 | The role of prior in image based 3D modeling: a survey
Hao Zhu 0004, Yongming Nie, Tao Yue 0003, Xun Cao |
Frontiers Comput. Sci. | 1 |
| 2017 | Robust multi-view stereo synthesized by various parameters model
Yongming Nie, Tao Yue 0003, Hao Zhu 0004, Sidan Du, Xun Cao |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Video-Based Outdoor Human ReconstructionabstractA human body scanning system of great practical convenience, which can be used in an outdoor environment, is proposed. The system uses only a single conventional video camera without the aid of special sensors or controlled illuminations. We leverage the structure from motion calibration results directly and improve the available video-based dense 3D reconstruction by integrating the surface smoothness constraints. The point cloud reinforcement is proposed to detect and adjust the conflict point data for the slender and shaky body parts. Combined with the silhouette adaptation, the proposed point cloud reinforcement achieves reasonable and plausible mesh reconstruction on these challenging parts. We further introduce the close-shot frames to refine the prereconstructed mesh model, leading to a colored watertight model. The overall system is approximate to automatic since only one or two times of painting brush interaction are required for robust and high-quality multiview image segmentation. The experiment results on various test sequences demonstrate the effectiveness and the robustness of the proposed method, even under very challenging scenarios when shaking body, varying illumination, and textureless regions occur. Hao Zhu 0004, Yebin Liu, Jingtao Fan, Qionghai Dai, Xun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |