VLDB 2026 Research / reviewers in the wild / expert
Haozhe Jia
dblp:201/7583
· DBLP profile ↗
14ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0002-1224-354XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
5 papers |
Rendering · 40% Computer animation and physical simulation · 25% Image and video coding · 13% | |
| Artificial intelligence
3 papers |
Generative modeling · 54% 3D vision · 41% Representation and self-supervised learning · 5% | |
| Computer networks
1 paper |
Wireless sensing and localization · 77% Cellular and mobile networks · 23% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.9 | 2 | 2026 | Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priors · IEEE Trans. Pattern Anal. Mach. Intell. 2026 DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space · ICML 2025 |
Rendering
gaussian splatting |
1.8 | 2 | 2026 | Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priors · IEEE Trans. Pattern Anal. Mach. Intell. 2026 Topo4D: Topology-Preserving Gaussian Splatting for High-fidelity 4D Head Capture · ECCV (34) 2024 |
Computer vision › 3D vision
3d reconstruction |
1.0 | 1 | 2026 | Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priors · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › 3D vision › 3d reconstruction
surface reconstruction |
1.0 | 1 | 2026 | Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priors · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Rendering
inverse rendering |
1.0 | 1 | 2026 | Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priors · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer animation and physical simulation › motion synthesis › human motion synthesis
diffusion-based motion generation |
0.9 | 1 | 2025 | ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model · ACM Multimedia 2025 |
Image and video coding › transform coding
discrete cosine transform |
0.9 | 1 | 2025 | DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space · ICML 2025 |
Computer animation and physical simulation › motion synthesis › human motion synthesis
text-to-motion generation |
0.9 | 1 | 2025 | ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model · ACM Multimedia 2025 |
Wireless sensing and localization › radio map
radio map estimation |
0.9 | 1 | 2025 | Physics-Informed Representation Alignment for Sparse Radio-Map Reconstruction · ACM Multimedia 2025 |
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion autoencoder |
0.8 | 1 | 2024 | DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing · ACM Multimedia 2024 |
Geometric modeling and processing
3d reconstruction |
0.8 | 1 | 2024 | Topo4D: Topology-Preserving Gaussian Splatting for High-fidelity 4D Head Capture · ECCV (34) 2024 |
Visual content generation and editing
face editing |
0.8 | 1 | 2024 | DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing · ACM Multimedia 2024 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.2 | 1 | 2024 | DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
physically-based rendering · 2.0blendshape expression priors · 2.03d gaussian splatting · 2.0dit · 1.7diffusion sampling · 1.7UViT · 1.7diffusion autoencoder · 1.5u-net · 0.9spectral analysis · 0.9representation alignment · 0.9physics-informed neural network · 0.9diffusion model · 0.9classifier-free guidance scheduling · 0.9topology-preserving optimization · 0.8disentangled control · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priorsabstract4D head capture aims to generate dynamic facial meshes in the same topology with corresponding UV maps, which requires temporal correspondence between 3D head models. Existing pipelines either involve manual processing of artists or employ constraints such as landmark tracking and optical flow, failing to achieve a trade-off between accuracy and efficiency. To enhance this process, we propose Topo4D++, a novel framework for automatic geometry and texture reconstruction that optimizes densely aligned 4D heads and 8 K BRDF maps directly from calibrated multi-view videos. Our key insight is to represent facial models as a set of dynamic 3D Gaussians with fixed topology, where the Gaussian centers are bound to the mesh vertices. This enables tracking all vertices rather than sparse vertices on the face accurately by leveraging the inverse rendering capabilities of 3D Gaussian Splatting (3DGS), while also enabling ultra-high-resolution texture generation. To maintain face structure during dynamic 3DGS optimization, we propose to optimize geometry and texture alternatively under physical and topological constraints frame-by-frame and employ blendshape-based expression priors to address extreme expressions. Then, we propose to extract dynamic facial meshes in a regular wiring arrangement and high-fidelity textures with pore-level details from the learned Gaussians. Finally, we train a diffusion-based model to generate BRDF texture maps to achieve physically based rendering. Given the absence of a universal benchmark, we construct JHead, a novel benchmark for the comprehensive evaluation of 4D head capture methods. Extensive experiments on different datasets demonstrate that our method is generalized to different capture systems, identities, and expressions, outperforming current state-of-the-art head reconstruction methods in both mesh and texture qualitatively and quantitatively. Yuhao Cheng, Xuanchen Li, Xingyu Ren, Haozhe Jia, Di Xu 0012, Wenhan Zhu, Bingbing Ni, Yichao Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT SpaceabstractThis paper explores image modeling from the frequency space and introduces DCTdiff, an end-to-end diffusion generative paradigm that efficiently models images in the discrete cosine transform (DCT) space. We investigate the design space of DCTdiff and reveal the key design factors. Experiments on different frameworks (UViT, DiT), generation tasks, and various diffusion samplers demonstrate that DCTdiff outperforms pixel-based diffusion models regarding generative quality and training efficiency. Remarkably, DCTdiff can seamlessly scale up to 512$\times$512 resolution without using the latent diffusion paradigm and beats latent diffusion (using SD-VAE) with only 1/4 training cost. Finally, we illustrate several intriguing properties of DCT image modeling. For example, we provide a theoretical proof of why `image diffusion can be seen as spectral autoregression', bridging the gap between diffusion and autoregressive models. The effectiveness of DCTdiff and the introduced properties suggest a promising direction for image modeling in the frequency space. The code is at https://github.com/forever208/DCTdiff. Mang Ning, Mingxiao Li 0002, Jianlin Su, Haozhe Jia, Lanmiao Liu, Martin Benes 0001, Wenshuo Chen, Albert Ali Salah, Itir Önal |
ICML | 4 |
| 2025 | ANT: Adaptive Neural Temporal-Aware Text-to-Motion ModelabstractWhile diffusion models advance text-to-motion generation, their static semantic conditioning ignores temporal-frequency demands: early denoising requires structural semantics for motion foundations while later stages need localized details for text alignment. This mismatch mirrors biological morphogenesis where developmental phases demand distinct genetic programs. Inspired by epigenetic regulation governing morphological specialization, we propose **(ANT)**, an **A**daptive **N**eural **T**emporal-Aware architecture. ANT orchestrates semantic granularity through: **(i) Semantic Temporally Adaptive (STA) Module:** Automatically partitions denoising into low-frequency structural planning and high-frequency refinement via spectral analysis. **(ii) Dynamic Classifier-Free Guidance scheduling (DCFG):** Adaptively adjusts conditional to unconditional ratio enhancing efficiency while maintaining fidelity. Extensive experiments show that ANT can be applied to various baselines, significantly improving model performance, and achieving state-of-the-art semantic alignment on StableMoFusion. Wenshuo Chen, Kuimou Yu, Haozhe Jia, Kaishen Yuan, Zexu Huang, Songning Lai, Hongru Xiao, Erhang Zhang, Lei Wang 0108, Yutao Yue |
ACM Multimedia | 3 |
| 2025 | Physics-Informed Representation Alignment for Sparse Radio-Map ReconstructionabstractWith the rapid development of wireless communication technology, the efficient utilization of spectrum resources, optimization of communication quality, and intelligent communication have become critical. Radio map reconstruction is essential for enabling advanced applications, yet challenges such as complex signal propagation and sparse observational data hinder accurate reconstruction in practical scenarios. Existing methods often fail to align physical constraints with data-driven features, particularly under sparse measurement conditions. To address these issues, we propose Physics-Aligned Radio Map Diffusion Model (PhyRMDM), a novel framework that establishes cross-domain representation alignment between physical principles and neural network features through dual learning pathways. The proposed model integrates Physics-Informed Neural Networks (PINNs) with a representation alignment mechanism that explicitly enforces consistency between Helmholtz equation constraints and environmental propagation patterns. Our architecture employs two synergistic U-Nets: the first ensures physical consistency by minimizing PDE residuals and boundary conditions through latent space alignment, while the second refines predictions via diffusion-based denoising with attention-guided feature fusion. This dual alignment strategy enables simultaneous satisfaction of wave propagation laws and data distribution characteristics. Experimental results demonstrate significant improvements over state-of-the-art methods, achieving NMSE of 0.0031 and RMSE of 0.0125 under Static Radio Map (SRM) conditions, and NMSE of 0.0047 with RMSE of 0.0146 in Dynamic Radio Map (DRM) scenarios. The proposed representation alignment paradigm provides 37.2% accuracy enhancement in ultra-sparse cases (1% sampling rate), confirming its effectiveness in bridging physics-based modeling and deep learning for radio map reconstruction. These advancements establish a new framework for sparse signal environment characterization, with direct applications in 5G/6G network optimization and intelligent spectrum management. The code can be found on the website: https://github.com/Hxxxz0/RMDM Haozhe Jia, Wenshuo Chen, Lei Wang 0108, Hongru Xiao, Nanqian Jia, Keming Wu, Songning Lai, Yutao Yue |
ACM Multimedia | 1 |
| 2024 | Topo4D: Topology-Preserving Gaussian Splatting for High-fidelity 4D Head Capture
Xuanchen Li, Yuhao Cheng, Xingyu Ren, Haozhe Jia, Di Xu 0012, Wenhan Zhu, Yichao Yan |
ECCV (34) | 4 |
| 2024 | DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing
Haozhe Jia, Yan Li 0129, Hengfei Cui, Di Xu 0012, Yuwang Wang, Tao Yu 0007 |
ACM Multimedia | 1 |
| 2022 | Learning multi-scale synergic discriminative features for prostate image segmentation
Haozhe Jia, Tom Weidong Cai, Heng Huang 0001, Yong Xia 0001 |
Pattern Recognit. | 1 |
| 2020 | Learning High-Resolution and Efficient Non-local Features for Brain Glioma Segmentation in MR Images
Haozhe Jia, Yong Xia 0001, Tom Weidong Cai, Heng Huang 0001 |
MICCAI (4) | 1 |
| 2020 | Thorax-Net: An Attention Regularized Deep Neural Network for Classification of Thoracic Diseases on Chest RadiographyabstractDeep learning techniques have been increasingly used to provide more accurate and more accessible diagnosis of thorax diseases on chest radiographs. However, due to the lack of dense annotation of large-scale chest radiograph data, this computer-aided diagnosis task is intrinsically a weakly supervised learning problem and remains challenging. In this paper, we propose a novel deep convolutional neural network called Thorax-Net to diagnose 14 thorax diseases using chest radiography. Thorax-Net consists of a classification branch and an attention branch. The classification branch serves as a uniform feature extraction-classification network to free users from the troublesome hand-crafted feature extraction and classifier construction. The attention branch exploits the correlation between class labels and the locations of pathological abnormalities via analyzing the feature maps learned by the classification branch. Feeding a chest radiograph to the trained Thorax-Net, a diagnosis is obtained by averaging and binarizing the outputs of two branches. The proposed Thorax-Net model has been evaluated against three state-of-the-art deep learning models using the patientwise official split of the ChestX-ray14 dataset and against other five deep learning models using the imagewise random data split. Our results show that Thorax-Net achieves an average per-class area under the receiver operating characteristic curve (AUC) of 0.7876 and 0.896 in both experiments, respectively, which are higher than the AUC values obtained by other deep models when they were all trained with no external data. Hongyu Wang 0011, Haozhe Jia, Le Lu 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | 3D APA-Net: 3D Adversarial Pyramid Anisotropic Convolutional Network for Prostate Segmentation in MR ImagesabstractAccurate and reliable segmentation of the prostate gland using magnetic resonance (MR) imaging has critical importance for the diagnosis and treatment of prostate diseases, especially prostate cancer. Although many automated segmentation approaches, including those based on deep learning have been proposed, the segmentation performance still has room for improvement due to the large variability in image appearance, imaging interference, and anisotropic spatial resolution. In this paper, we propose the 3D adversarial pyramid anisotropic convolutional deep neural network (3D APA-Net) for prostate segmentation in MR images. This model is composed of a generator (i.e., 3D PA-Net) that performs image segmentation and a discriminator (i.e., a six-layer convolutional neural network) that differentiates between a segmentation result and its corresponding ground truth. The 3D PA-Net has an encoder-decoder architecture, which consists of a 3D ResNet encoder, an anisotropic convolutional decoder, and multi-level pyramid convolutional skip connections. The anisotropic convolutional blocks can exploit the 3D context information of the MR images with anisotropic resolution, the pyramid convolutional blocks address both voxel classification and gland localization issues, and the adversarial training regularizes 3D PA-Net and thus enables it to generate spatially consistent and continuous segmentation results. We evaluated the proposed 3D APA-Net against several state-of-the-art deep learning-based segmentation approaches on two public databases and the hybrid of the two. Our results suggest that the proposed model outperforms the compared approaches on three databases and could be used in a routine clinical workflow. Haozhe Jia, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Heng Huang 0001, Yanning Zhang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 1 |
| 2019 | HD-Net: Hybrid Discriminative Network for Prostate Segmentation in MR Images
Haozhe Jia, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai, Yong Xia 0001 |
MICCAI (2) | 1 |
| 2018 | Densely Connected Large Kernel Convolutional Network for Semantic Membrane Segmentation in Microscopy ImagesabstractStructural analysis of neurons can provide valuable insights of brain function. Semantic segmentation of neurons thus becomes an important technique in bioinformatics. Deep learning approaches have shown promising performance in various semantic segmentation problems. However, segmentation of neurons in Electron Microscopy (EM) images has some differences compared with typical segmentation tasks due to the image noise and the disturbance of the intracellular structures. In our work, we propose a network with a ResNet encoder and densely connected decoder with large kernels, and then refinement with simple morphological post-possessing. Two main advantages of our method are: 1) the network can prevent the loss of high-resolution information and enlarge the reception field; 2) the post-processing method is simple and can be directly applied to the probability map from the network to enhance the unconfident area. Evaluated on the ISBI2012 EM membrane segmentation challenge, the proposed method achieves competitive performance. Dongnan Liu, Donghao Zhang 0004, Siqi Liu 0001, Yang Song 0001, Haozhe Jia, David Dagan Feng, Yong Xia 0001, Tom Weidong Cai |
ICIP | 5 |
| 2018 | Panoptic Segmentation with an End-to-End Cell R-CNN for Pathology Image Analysis
Donghao Zhang 0004, Yang Song 0001, Dongnan Liu, Haozhe Jia, Siqi Liu 0001, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai |
MICCAI (2) | 4 |
| 2018 | Atlas registration and ensemble deep convolutional neural network-based prostate segmentation using magnetic resonance imaging
Haozhe Jia, Yong Xia 0001, Yang Song 0001, Tom Weidong Cai, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 1 |