VLDB 2026 Research / reviewers in the wild / expert
Hao Zhu 0005
dblp:10/3520-5
· DBLP profile ↗
20ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-6756-9571ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward the Spectral Bias Alleviation by Normalizations in Coordinate NetworksabstractRepresenting signals using coordinate networks dominates the area of inverse problems recently, and is widely applied in various scientific computing tasks. Still, there exists an issue of spectral bias in coordinate networks, limiting the capacity to learn high-frequency components. This problem is caused by the pathological distribution of the neural tangent kernel's (NTK's) eigenvalues of coordinate networks. We find that, this pathological distribution could be improved using classical normalization techniques (batch normalization and layer normalization), which are commonly used in convolutional neural networks but rarely used in coordinate networks. We prove that normalization techniques greatly reduces the maximum and variance of NTK's eigenvalues while slightly modifies the mean value, considering the max eigenvalue is much larger than the most, this variance change results in a shift of eigenvalues' distribution from a lower one to a higher one, therefore the spectral bias could be alleviated (see Fig. 1). Furthermore, we propose two new normalization techniques by combining these two techniques in different ways. The efficacy of these normalization techniques is substantiated by the significant improvements and new state-of-the-arts achieved by applying normalization-based coordinate networks to various tasks, including the image compression, computed tomography reconstruction, shape representation, magnetic resonance imaging, novel view synthesis and multi-view stereo reconstruction. Zhicheng Cai, Hao Zhu 0005, Qiu Shen, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Light Field Reconstruction Using Multi-orientation Epipolar Plane ImagesabstractLight field reconstruction is one of the most important techniques for future glass-free 3D media production. However, current techniques suffer from low view-consistency and fixed patterns of view-trajectory. This article presents the 3D Multi-orientation Epipolar Plane Image (MOEPI) representation for high-quality light field reconstruction with both the inter-view and extra-view settings. Each layer in MOEPI is composed of EPI lines with a fixed orientation. To infer the MOEPI, a new Multi-reference Focal Stack (MRFS) intermediate representation is proposed. The optimization of MOEPI could be regarded as the problem of the most-focused content extraction from the MRFS. This optimization is implemented with a 3D U-shaped network. We also propose the LPIPS-EPI metric for evaluating the view-consistency. Experiments on light fields with both high and low signal-to-noise ratios demonstrate that the proposed MOEPI representation could synthesize high-quality light fields especially in occlusion or un-captured areas. Hao Zhu 0005, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2026 | Hierarchical Bayesian Guided Spatial-, Angular- and Temporal-Consistent View SynthesisabstractNeural Radiance Fields (NeRF) have gained significant attention due to their precise reconstruction and rapid inference capabilities, making them highly promising for applications in virtual reality and gaming. However, extending NeRF's capabilities to dynamic scenes remains underexplored, particularly in ensuring consistent and coherent reconstructions across space, time, and viewing angles. To address this challenge, we propose Scale-NeRF, a novel approach that organizes the training of dynamic NeRFs as a progressive, scale-based refinement process, grounded in hierarchical Bayesian theory. Scale-NeRF begins by reconstructing the radiance fields using coarse, large-scale frames and iteratively refines them with progressively smaller-scale frames. This hierarchical strategy, combined with a corresponding sampling approach and a newly introduced structural loss, ensures consistency and integrity throughout the reconstruction process. Experiments on public datasets validate the superiority of Scale-NeRF over traditional methods, especially in terms of the proposed metrics evaluating spatial, angular, and temporal consistency. Furthermore, Scale-NeRF demonstrates excellent dynamic reconstruction capabilities with real-time rendering, offering a significant advancement for applications demanding both high fidelity and real-time performance. Junyu Zhu, Hao Zhu 0005, Zhan Ma 0001, Xun Cao |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | FINER: Flexible Spectral-Bias Tuning in Implicit NEural Representation by Variableperiodic Activation FunctionsabstractImplicit Neural Representation (INR), which utilizes a neural network to map coordinate inputs to corresponding attributes, is causing a revolution in the field of signal processing. However, current INR techniques suffer from a re-stricted capability to tune their supported frequency set, re-sulting in imperfect performance when representing complex signals with multiple frequencies. We have identified that this frequency-related problem can be greatly alleviated by introducing variableperiodic activation functions, for which we propose FINER. By initializing the bias of the neural network within different ranges, sub-functions with various frequencies in the variableperiodic function are selected for activation. Consequently, the supported frequency set of FINER can be flexibly tuned, leading to improved performance in signal representation. We demon-strate the capabilities of FINER in the contexts of2D image fitting, 3D signed distance field representation, and 5D neural radiance fields optimization, and we show that it outper-forms existing INRs. Zhen Liu 0031, Hao Zhu 0005, Qi Zhang 0029, Jingde Fu, Weibing Deng, Zhan Ma 0001, Yanwen Guo 0001, Xun Cao |
CVPR | 2 |
| 2024 | Neural Poisson Solver: A Universal and Continuous Framework for Natural Signal Blending
Delong Wu, Hao Zhu 0005, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
ECCV (80) | 2 |
| 2024 | Sheared Epipolar Focus Spectrum for Dense Light Field ReconstructionabstractThis paper presents a novel technique for the dense reconstruction of light fields (LFs) from sparse input views. Our approach leverages the Epipolar Focus Spectrum (EFS) representation, which models the LF in the transformed spatial-focus domain, avoiding the dependence on the scene depth and providing a high-quality basis for dense LF reconstruction. Previous EFS-based LF reconstruction methods learn the cross-view, occlusion, depth and shearing terms simultaneously, which makes the training difficult due to stability and convergence problems and further results in limited reconstruction performance for challenging scenarios. To address this issue, we conduct a theoretical study on the transformation between the EFSs derived from one LF with sparse and dense angular samplings, and propose that a dense EFS can be decomposed into a linear combination of the EFS of the sparse input, the sheared EFS, and a high-order occlusion term explicitly. The devised learning-based framework with the input of the under-sampled EFS and its sheared version provides high-quality reconstruction results, especially in large disparity areas. Comprehensive experimental evaluations show that our approach outperforms state-of-the-art methods, especially achieves at most dB advantages in reconstructing scenes containing thin structures. Xue Wang 0006, Guoqing Zhou 0003, Hao Zhu 0005, Qing Wang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Disorder-Invariant Implicit Neural RepresentationabstractImplicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the expressive power of INR is limited by the spectral bias in the network training. In this paper, we find that such a frequency-related problem could be greatly solved by re-arranging the coordinates of the input signal, for which we propose the disorder-invariant implicit neural representation (DINER) by augmenting a hash-table to a traditional INR backbone. Given discrete signals sharing the same histogram of attributes and different arrangement orders, the hash-table could project the coordinates into the same distribution for which the mapped signal can be better modeled using the subsequent INR network, leading to significantly alleviated spectral bias. Furthermore, the expressive power of the DINER is determined by the width of the hash-table. Different width corresponds to different geometrical elements in the attribute space, e.g., 1D curve, 2D curved-plane and 3D curved-volume when the width is set as 1, 2 and 3, respectively. More covered areas of the geometrical elements result in stronger expressive power. Experiments not only reveal the generalization of the DINER for different INR backbones (MLP versus SIREN) and various tasks (image/video representation, phase retrieval, refractive index recovery, and neural radiance field optimization) but also show the superiority over the state-of-the-art algorithms both in quality and speed. Hao Zhu 0005, Shaowen Xie, Zhen Liu 0031, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Continuous 3D Myocardial Motion Tracking via EchocardiographyabstractMyocardial motion tracking stands as an essential clinical tool in the prevention and detection of cardiovascular diseases (CVDs), the foremost cause of death globally. However, current techniques suffer from incomplete and inaccurate motion estimation of the myocardium in both spatial and temporal dimensions, hindering the early identification of myocardial dysfunction. To address these challenges, this paper introduces the Neural Cardiac Motion Field (NeuralCMF). NeuralCMF leverages implicit neural representation (INR) to model the 3D structure and the comprehensive 6D forward/backward motion of the heart. This method surpasses pixel-wise limitations by offering the capability to continuously query the precise shape and motion of the myocardium at any specific point throughout the cardiac cycle, enhancing the detailed analysis of cardiac dynamics beyond traditional speckle tracking. Notably, NeuralCMF operates without the need for paired datasets, and its optimization is self-supervised through the physics knowledge priors in both space and time dimensions, ensuring compatibility with both 2D and 3D echocardiogram video inputs. Experimental validations across three representative datasets support the robustness and innovative nature of the NeuralCMF, marking significant advantages over existing state-of-the-art methods in cardiac imaging and motion tracking. Code is available at: https://njuvision.github.io/NeuralCMF. Chengkang Shen, Hao Zhu 0005, Si Yi, Weipeng Zhao, David J. Brady, Xun Cao, Zhan Ma 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Dense light field reconstruction based on epipolar focus spectrumabstractExisting light field (LF) representations, such as epipolar plane image (EPI) and sub-aperture images, do not consider the structural characteristics across the views, so they usually require additional disparity and spatial structure cues for follow-up tasks. Besides, they have difficulties dealing with occlusions or large disparity scenes. To this end, this paper proposes a novel Epipolar Focus Spectrum (EFS) representation by rearranging the EPI spectrum. Different from the classical EPI representation where an EPI line corresponds to a specific depth, there is a one-to-one mapping from the EFS line to the view. By exploring the EFS sampling task, the analytical function is derived for constructing a non-aliasing EFS. To demonstrate its effectiveness, we develop a trainable EFS-based pipeline for light field reconstruction, where a dense light field can be reconstructed by compensating the missing EFS lines given a sparse light field, yielding promising results with cross-view consistency, especially in the presence of severe occlusion and large disparity. Experimental results on both synthetic and real-world datasets demonstrate the validity and superiority of the proposed method over SOTA methods. Xue Wang 0006, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006 |
Pattern Recognit. | 3 |
| 2023 | Learning Reliable Gradients From Undersampled Circular Light Field for 3D ReconstructionabstractThe paper presents a 3D reconstruction algorithm from an undersampled circular light field (LF). With an ultra-dense angular sampling rate, every scene point captured by a circular LF corresponds to a smooth trajectory in the circular epipolar plane volume (CEPV). Thus per-pixel disparities can be calculated by retrieving the local gradients of the CEPV-trajectories. However, the continuous curve will be broken up into discrete segments in an undersampled circular LF, which leads to a noticeable deterioration of the 3D reconstruction accuracy. We observe that the coherent structure is still embedded in the discrete segments. With less noise and ambiguity, the scene points can be reconstructed using gradients from reliable epipolar plane image (EPI) regions. By analyzing the geometric characteristics of the coherent structure in the CEPV, both the trajectory itself and its gradients could be modeled as 3D predictable series. Thus a mask-guided CNN+LSTM network is proposed to learn the mapping from the CEPV with a lower angular sampling rate to the gradients under a higher angular sampling rate. To segment the reliable regions, the reliable-mask-based loss that assesses the difference between learned gradients and ground truth gradients is added to the loss function. We construct a synthetic circular LF dataset with ground truth for depth and foreground/background segmentation to train the network. Moreover, a real-scene circular LF dataset is collected for performance evaluation. Experimental results on both public and self-constructed datasets demonstrate the superiority of the proposed method over existing state-of-the-art methods. Zhengxi Song, Xue Wang 0006, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Revisiting Spatio-Angular Trade-off in Light Field Cameras and Extended Applications in Super-ResolutionabstractLight field cameras (LFCs) have received increasing attention due to their wide-spread applications. However, current LFCs suffer from the well-known spatio-angular trade-off, which is considered an inherent and fundamental limit for LFC designs. In this article, by doing a detailed optical analysis of the sampling process in an LFC, we show that the effective resolution is generally higher than the number of micro-lenses. This contribution makes it theoretically possible to super-resolve a light field. Further optical analysis proves the "2D predictable series" nature of the 4D light field, which provides new insights for analyzing light field using series processing techniques. To model this nature, a specifically designed epipolar plane image (EPI) based CNN-LSTM network is proposed to super-resolve a light field in the spatial and angular dimensions simultaneously. Rather than leveraging semantic information, our network focuses on extracting geometric continuity in the EPI domain. This gives our method an improved generalization ability and makes it applicable to a wide range of previously unseen scenes. Experiments on both synthetic and real light fields demonstrate the improvements over state-of-the-arts, especially in large disparity areas. Hao Zhu 0005, Mantang Guo, Hongdong Li, Qing Wang 0006, Antonio Robles-Kelly |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Accurate 3D Reconstruction from Circular Light Field Using CNN-LSTMabstractA light field is formed by densely capturing images on a regular sub-aperture grid. Geometry information endowed in the epipolar plane images(EPI) can only lead to a 2. 5D reconstruction. In order to obtain a full 360°view of an object, we focus on light fields captured by a circularly moving camera, resulting in circular light fields (or Cir-LFs in short). Compared with traditional EPIs, Circular EPIs(CEPIs) provide unique advantages, such as that corresponding points forming a 3D sinusoid like curve instead of a 2D straight line and geometry information encoded sequentially in multiple adjacent views along the curve. However, current reconstruction methods only focus on the 2D projection of 3D curve, leading to distortions in the reconstructed upper and lower surfaces. We propose to analyze 3D features contained in the 3D CEPI volume and we develop a deep CNN-LSTM network to model the gradient map in the CEPI volume. Additionally, a large scale Cir-LF dataset is constructed for research purpose. Experiments on both synthetic and real scenes demonstrate the effectiveness and generaliability of the proposed method. Zhengxi Song, Hao Zhu 0005, Xue Wang 0006, Hongdong Li, Qing Wang 0006 |
ICME | 2 |
| 2020 | 4D Light Field Superpixel and SegmentationabstractSuperpixel segmentation of 2D images has been widely used in many computer vision tasks. Previous algorithms model the color, position, or higher spectral information for segmenting a 2D image. However, limited to the Gaussian imaging principle in a traditional camera, where each pixel is formed by summing lots of light rays from different angles, there is not a thorough segmentation solution to eliminate the ambiguity in defocus and occlusion boundary areas. In this paper, we consider the essential element of image pixel, i.e., rays in light space, and propose light field superpixel (LFSP) to eliminate the ambiguity. The LFSP is first defined mathematically and then two evaluation metrics, named LFSP self-similarity and effective label ratio, are proposed to evaluate the refocus-invariant and full-sliced properties of segmentation. By building a clique system containing 80 neighbors in light field, a robust refocus-invariant LFSP segmentation algorithm is developed. Experimental results on both synthetic and real light field datasets demonstrate the advantages over the current state of the art in terms of traditional evaluation metrics. Additionally, the LFSP self-similarity evaluations under different light field refocus levels show the refocus-invariance of the proposed algorithm. The full-sliced property of the proposed LFSP algorithm is verified by comparing it with the classical supervoxel algorithms. Finally, an LFSP-based application is demonstrated to show the effectiveness of LFSP in light field editing. Hao Zhu 0005, Qi Zhang 0029, Qing Wang 0006, Hongdong Li |
IEEE Trans. Image Process. | 1 |
| 2018 | Dense Light Field Reconstruction from Sparse Sampling Using Residual Network
Mantang Guo, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006 |
ACCV (6) | 2 |
| 2017 | 4D Light Field Superpixel and SegmentationabstractSuperpixel segmentation of 2D image has been widely used in many computer vision tasks. However, limited to the Gaussian imaging principle, there is not a thorough segmentation solution to the ambiguity in defocus and occlusion boundary areas. In this paper, we consider the essential element of image pixel, i.e., rays in the light space and propose light field superpixel (LFSP) segmentation to eliminate the ambiguity. The LFSP is first defined mathematically and then a refocus-invariant metric named LFSP self-similarity is proposed to evaluate the segmentation performance. By building a clique system containing 80 neighbors in light field, a robust refocus-invariant LFSP segmentation algorithm is developed. Experimental results on both synthetic and real light field datasets demonstrate the advantages over the state-of-the-arts in terms of traditional evaluation metrics. Additionally the LFSP self-similarity evaluation under different light field refocus levels shows the refocus-invariance of the proposed algorithm. Hao Zhu 0005, Qi Zhang 0029, Qing Wang 0006 |
CVPR | 1 |
| 2017 | High angular resolution light field reconstruction with coded-aperture maskabstractIn the past decade, light field imaging has greatly extended the imaging capabilities of traditional photography. However, the applications of light field imaging are limited by the aliasing artifacts due to the plenoptic sampling trade-off between angular and spatial domains. We propose to use a coded aperture light field camera instead of the traditional one, which can get more angular information without losing spatial resolution. To that end, we exploit a theoretical model to explain the relationship between light field and the raw data captured by the sensor. Then, we design a mask to code the rays using compressive sensing. Last, the sparse characteristic of light field in gradient domain and the corresponding optimization methods are utilized to reconstruct the high angular resolution light field. Experimental results on synthetic data and real data demonstrate that our system can obtain high angular resolution light field by producing a low-aliasing refocused image and high PSNR multi-view images. Wanxin Qu, Guoqing Zhou 0003, Hao Zhu 0005, Zhaolin Xiao, Qing Wang 0006, René Vidal |
ICIP | 3 |
| 2017 | Extending the FOV from disparity and color consistencies in multiview light fieldsabstractLight field, which is captured by a plenoptic camera, is always limited in its narrow field of view (FOV) by the physical size of the aperture. To break through the restriction, we propose to extend the FOV using multiview light fields. A series of light fields are acquired by translating the camera at isometric spatial positions. In contrast to previous methods, our algorithm is the first that achieves light field registration and rendering based on epipolar plane image (EPI) properties, including disparity and color consistencies. Furthermore, the aliasing caused by the under-sampling in the angular space is eliminated by synthesizing novel views in the EPI space. Experimental results on the real scene data have demonstrated the effectiveness of our algorithm. Zhao Ren, Qi Zhang 0029, Hao Zhu 0005, Qing Wang 0006 |
ICIP | 3 |
| 2017 | Light field imaging: models, calibrations, reconstructions, and applicationsabstractLight field imaging is an emerging technology in computational photography areas. Based on innovative designs of the imaging model and the optical path, light field cameras not only record the spatial intensity of threedimensional (3D) objects, but also capture the angular information of the physical world, which provides new ways to address various problems in computer vision, such as 3D reconstruction, saliency detection, and object recognition. In this paper, three key aspects of light field cameras, i.e., model, calibration, and reconstruction, are reviewed extensively. Furthermore, light field based applications on informatics, physics, medicine, and biology are exhibited. Finally, open issues in light field imaging and long-term application prospects in other natural sciences are discussed. Hao Zhu 0005, Qing Wang 0006, Jingyi Yu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2016 | LFHOG: A discriminative descriptor for live face detection from light field imageabstractHow to avoid the invading of the attack in the biometric system, such as 2D printed photos, gradually becomes an important research hotspot. In this paper, we present a novel descriptor in light field to tackle the issue. Based on the angular and spatial information in light field, the proposed light field histogram of gradient (LFHoG) descriptor is derived from three directions, including vertical, horizontal and depth. Different with traditional HoG in 2D image, the gradient in depth direction is distinctive in light field. To validate the effectiveness of the proposed LFHoG descriptor, experiments have been carried out on light field datasets taken by a Lytro camera. The descriptor can achieve 99.75% accuracy on the user collected dataset, which proves the correctness and effectiveness of the LFHoG descriptor. Hao Zhu 0005, Qing Wang 0006 |
ICIP | 2 |
| 2016 | Accurate disparity estimation in light field using ground control pointsabstractThe recent development of light field cameras has received growing interest, as their rich angular information has potential benefits for many computer vision tasks. In this paper, we introduce a novel method to obtain a dense disparity map by use of ground control points (GCPs) in the light field. Previous work optimizes the disparity map by local estimation which includes both reliable points and unreliable points. To reduce the negative effect of the unreliable points, we predict the disparity at non-GCPs from GCPs. Our method performs more robustly in shadow areas than previous methods based on GCP work, since we combine color information and local disparity. Experiments and comparisons on a public dataset demonstrate the effectiveness of our proposed method. Hao Zhu 0005, Qing Wang 0006 |
Comput. Vis. Media | 1 |