VLDB 2026 Research / reviewers in the wild / expert
Zhuoman Liu
dblp:284/0962
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-2991-5242ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Monocular and Generalizable Gaussian Talking Head AnimationabstractIn this work, we introduce Monocular and Generalizable Gaussian Talking Head Animation (MGGTalk), which requires monocular datasets and generalizes to unseen identities without personalized re-training. Compared with previous 3D Gaussian Splatting (3DGS) methods that requires elusive multi-view datasets or tedious personalized learning/inference, MGGtalk enables more practical and broader applications. However, in the absence of multi-view and personalized training data, the incompleteness of geometric and appearance information poses a significant challenge. To address these challenges, MGGTalk explores depth information to enhance geometric and facial symmetry characteristics to supplement both geometric and appearance features. Initially, based on the pixel-wise geometric information obtained from depth estimation, we incorporate symmetry operations and point cloud filtering techniques to ensure a complete and precise position parameter for 3DGS. Subsequently, we adopt a two-stage strategy with symmetric priors for predicting the remaining 3DGS parameters. We begin by predicting Gaussian parameters for the visible facial regions of the source image. These parameters are subsequently utilized to improve the prediction of Gaussian parameters for the non-visible regions. Extensive experiments demonstrate that MGGTalk surpasses previous state-of-the-art methods, achieving superior performance across various metrics. Project page: https://scut-mmpr.github.io/MGGTalk-Homepage/. Shengjie Gong, Jiapeng Tang, Dongming Hu, Shuangping Huang, Tianshui Chen, Zhuoman Liu |
CVPR | 8 |
| 2025 | Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene SimulationabstractRealistic simulation of dynamic scenes requires accurately capturing diverse material properties and modeling complex object interactions grounded in physical principles. However, existing methods are constrained to basic material types with limited predictable parameters, making them insufficient to represent the complexity of real-world materials. We introduce PhysFlow, a novel approach that leverages multi-modal foundation models and video diffusion to achieve enhanced 4D dynamic scene simulation. Our method utilizes multi-modal models to identify material types and initialize material parameters through image queries, while simultaneously inferring 3D Gaussian splats for detailed scene representation. We further refine these material parameters using video diffusion with a differentiable Material Point Method (MPM) and optical flow guidance rather than render loss or Score Distillation Sampling (SDS) loss. This integrated framework enables accurate prediction and realistic simulation of dynamic interactions in real-world scenarios, advancing both accuracy and flexibility in physics-based simulations. Our code and data are available at https://zhuomanliu.github.io/PhysFlow Zhuoman Liu, Weicai Ye, Yan Luximon, Pengfei Wan 0001, Di Zhang 0026 |
CVPR | 1 |
| 2025 | SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lightingabstract3D inverse rendering in indoor scenes with strong light sources presents a significant challenge, primarily due to the substantial ambiguity in material recovery caused by the complex interaction between lighting and shadows. To address this, we propose a novel approach that integrates an implicit-explicit shadow predictor with a three-stage material estimation process. Our method enhances shadow realism by accurately predicting light interactions, while our material estimation process improves SVBRDF quality under challenging lighting conditions. Extensive experiments demonstrate the effectiveness of our method in both quantitative and qualitative metrics, enabling realistic object insertion and material replacement with proper shadow rendering under strong indoor light sources. Xiaokang Wei, Zhuoman Liu, Ping Li 0016, Yan Luximon |
ICME | 2 |
| 2025 | DDF-ISM: Internal Structure Modeling of Human Head Using Probabilistic Directed Distance FieldabstractThe increasing interest surrounding 3D human heads for digital avatars and simulations has highlighted the need for accurate internal modeling rather than solely focusing on external approximations. Existing approaches rely on traditional optimization techniques applied to explicit 3D representations like point clouds and meshes, leading to computational inefficiencies and challenges in capturing local geometric features. To tackle these problems, we propose a novel modeling method called DDF-ISM. It leverages a probabilistic Directed Distance Field for Internal Structure Modeling, facilitating efficient and anatomically accurate deformation of different parts of the human head. DDF-ISM comprises two key components: 1) a probabilistic DDF network for implicit representation of the target model to provide crucial local geometric information, and 2) a conditioned deformation network guided by the local geometry. Additionally, we introduce a large-scale dataset of human heads with internal structures derived from high-quality Computed Tomography (CT) scans, along with well-designed template models encompassing skull, mandible, brain, and head surface. Evaluation on this dataset showcases the superiority of our approach over existing methods, exhibiting superior performance in both modeling quality and efficiency. Zhuoman Liu, Yan Luximon, Wei Lin Ng, Eric Chung |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Disentangling Writer and Character Styles for Handwriting GenerationabstractTraining machines to synthesize diverse handwritings is an intriguing task. Recently, RNN-based methods have been proposed to generate stylized online Chinese characters. However, these methods mainly focus on capturing a person's overall writing style, neglecting subtle style inconsistencies between characters written by the same person. For example, while a person's handwriting typically exhibits general uniformity (e.g., glyph slant and aspect ratios), there are still small style variations in finer details (e.g., stroke length and curvature) of characters. In light of this, we propose to disentangle the style representations at both writer and character levels from individual handwritings to synthesize realistic stylized online handwritten characters. Specifically, we present the style-disentangled Transformer (SDT), which employs two complementary contrastive objectives to extract the style commonalities of reference samples and capture the detailed style patterns of each sample, respectively. Extensive experiments on various language scripts demonstrate the effectiveness of SDT. Notably, our empirical findings reveal that the two learned style representations provide information at different frequency magnitudes, underscoring the importance of separate style extraction. Our source code is public at: https://github.com/dailenson/SDT. Gang Dai 0002, Yifan Zhang 0004, Zhu Liang Yu, Zhuoman Liu, Shuangping Huang |
CVPR | 6 |
| 2023 | RayDF: Neural Ray-surface Distance Fields with Multi-view ConsistencyabstractIn this paper, we study the problem of continuous 3D shape representations. The majority of existing successful methods are coordinate-based implicit neural representations. However, they are inefficient to render novel views or recover explicit surface points. A few works start to formulate 3D shapes as ray-based neural functions, but the learned structures are inferior due to the lack of multi-view geometry consistency. To tackle these challenges, we propose a new framework called RayDF. It consists of three major components: 1) the simple ray-surface distance field, 2) the novel dual-ray visibility classifier, and 3) a multi-view consistency optimization module to drive the learned ray-surface distances to be multi-view geometry consistent. We extensively evaluate our method on three public datasets, demonstrating remarkable performance in 3D surface point reconstruction on both synthetic and challenging real-world 3D scenes, clearly surpassing existing coordinate-based and ray-based baselines. Most notably, our method achieves a 1000x faster speed than coordinate-based methods to render an 800x800 depth image, showing the superiority of our method for 3D shape representation. Our code and data are available at https://github.com/vLAR-group/RayDF Zhuoman Liu, Bo Yang 0027, Yan Luximon |
NeurIPS | 1 |
| 2022 | Deep View Synthesis via Self-Consistent Generative NetworkabstractView synthesis aims to produce unseen views from a set of views captured by two or more cameras at different positions. This task is non-trivial since it is hard to conduct pixel-level matching among different views. To address this issue, most existing methods seek to exploit the geometric information to match pixels. However, when the distinct cameras have a large baseline (i. e., far away from each other), severe geometry distortion issues would occur and the geometric information may fail to provide useful guidance, resulting in very blurry synthesized images. To address the above issues, in this paper, we propose a novel deep generative model, called Self-Consistent Generative Network (SCGN), which synthesizes novel views from the given input views without explicitly exploiting the geometric information. The proposed SCGN model consists of two main components, i. e., a View Synthesis Network (VSN) and a View Decomposition Network (VDN), both employing an Encoder-Decoder structure. Here, the VDN seeks to reconstruct input views from the synthesized novel view to preserve the consistency of view synthesis. Thanks to VDN, SCGN is able to synthesize novel views without using any geometric rectification before encoding, making it easier for both training and applications. Finally, adversarial loss is introduced to improve the photo-realism of novel views. Both qualitative and quantitative comparisons against several state-of-the-art methods on two benchmark tasks demonstrated the superiority of our approach. Zhuoman Liu, Ming Yang 0039, Peiyao Luo, Mingkui Tan |
IEEE Trans. Multim. | 1 |