EDBT 2026 Demo / reviewers in the wild / expert
Wenbo Zhao 0004
dblp:30/5943-4
· DBLP profile ↗
11ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0002-5517-1304ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spatial Annealing for Efficient Few-shot Neural RenderingabstractNeural Radiance Fields (NeRF) with hybrid representations have shown impressive capabilities for novel view synthesis, delivering high efficiency. Nonetheless, their performance significantly drops with sparse input views. Various regularization strategies have been devised to address these challenges. However, these strategies either require additional rendering costs or involve complex pipeline designs, leading to a loss of training efficiency. Although FreeNeRF has introduced an efficient frequency annealing strategy, its operation on frequency positional encoding is incompatible with the efficient hybrid representations. In this paper, we introduce an accurate and efficient few-shot neural rendering method named Spatial Annealing regularized NeRF (SANeRF), which adopts the pre-filtering design of a hybrid representation. We initially establish the analytical formulation of the frequency band limit for a hybrid architecture by deducing its filtering process. Based on this analysis, we propose a universal form of frequency annealing in the spatial domain, which can be implemented by modulating the sampling kernel to exponentially shrink from an initial one with a narrow grid tangent kernel spectrum. This methodology is crucial for stabilizing the early stages of the training phase and significantly contributes to enhancing the subsequent process of detail refinement. Our extensive experiments reveal that, by adding merely one line of code, SANeRF delivers superior rendering quality and much faster reconstruction speed compared to current few-shot neural rendering methods. Notably, SANeRF outperforms FreeNeRF on the Blender dataset, achieving 700X faster reconstruction speed. Yuru Xiao, Deming Zhai, Wenbo Zhao 0004, Kui Jiang, Junjun Jiang, Xianming Liu 0005 |
AAAI | 3 |
| 2025 | FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust FusionabstractImage-event joint depth estimation methods leverage complementary modalities for robust perception, yet face challenges in generalizability stemming from two factors: 1) limited annotated image-event-depth datasets causing insufficient cross-modal supervision, and 2) inherent frequency mismatches between static images and dynamic event streams with distinct spatiotemporal patterns, leading to ineffective feature fusion. To address this dual challenge, we propose Frequency-decoupled Unified Self-supervised Encoder (FUSE) with two synergistic components: The Parameter-efficient Self-supervised Transfer (PST) leverages image foundation models for cross-modal knowledge transfer, effectively mitigating data scarcity by enabling joint encoding without depth ground truth. Complementing this, the Frequency-Decoupled Fusion module (FreDFuse) resolves modality-specific frequency mismatches by decoupling features into high- and low-frequency bands and then performing a guided cross-attention fusion, where the modality dominant in each band steers the integration. This combined approach enables FUSE to construct a universal image-event encoder that only requires lightweight decoder adaptation for target datasets. Extensive experiments demonstrate state-of-the-art performance with 14% and 24.9% improvements in Abs.Rel on MVSEC and DENSE datasets. The framework exhibits remarkable robustness and generalization in challenging scenarios, including extreme lighting and motion blur, significantly advancing its real-world deployment capabilities. The source code for our method is publicly available at: https://github.com/sunpihai-up/FUSE. Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen, Wenbo Zhao 0004, Kui Jiang, Xianming Liu 0005 |
IROS | 5 |
| 2025 | A Wavelet-based Image Coding Framework for Data Storage on DNAabstractIn the face of the exponential growth of digital data, DNA is expected to become a new storage medium. Image data makes up a large proportion of digital data. However, existing DNA data storage models are mainly designed for general files. To address this issue, we propose a novel image encoding method for DNA data storage. We employ discrete wavelet transform to decompose the image and utilize an improved exponent-mantissa representation for numerical data. Subsequently, we achieve enhanced compression performance through context-adaptive arithmetic coding. Additionally, we construct a dictionary between ternary sequences and oligonucleotides to generate nucleotide sequences that meet the specified constraints. Experiments show that our method outperforms JPEG-DNA and BioCoder in compression performance and generates higher-quality nucleotide sequences. Chen Qin, Yuanchao Bai, Wenbo Zhao 0004, Xianming Liu 0005 |
VCIP | 3 |
| 2025 | LOD-PCAC: Level-of-Detail-Based Deep Lossless Point Cloud Attribute CompressionabstractPoint cloud attribute compression is a challenging issue in efficiently compressing large volumes of attributes. Despite notable advancements in lossy point cloud compression using deep learning, progress in lossless compression remains limited. Some methods have employed octree- or voxel-based partitioning techniques derived from geometric compression, achieving success on dense point clouds. However, these voxel-based approaches struggle with sparse or unevenly distributed point clouds, leading to performance degradation. In this work, we introduce a novel framework for learning-based lossless point cloud attribute compression, named LOD-PCAC, which leverages a Level-of-Detail (LOD) structure to ensure density-robust compression. Specifically, the input point cloud is divided into multiple detail levels, and vertices from these levels are selected to construct a Reference Set as context, which effectively captures multi-level information. Then we propose the Bit-level Residual Coder for efficient attribute compression. Instead of directly compressing attributes, our method first predicts attribute values and organizes the residual bits into a Bit Matrix as another context, simplifying predictions and fully exploiting channel correlations. Finally, a neural network with specialized encoders processes the context to estimate the probability of each residual bit. Experimental results demonstrate that the proposed method outperforms both traditional and learning-based approaches across various point clouds, exhibiting strong generalization across datasets and robustness to varying densities. Wenbo Zhao 0004, Wei Gao 0003, Dingquan Li, Jing Wang 0115 |
IEEE Trans. Image Process. | 1 |
| 2024 | Context-Adaptive Entropy Model With Adapters For Lossless Point Cloud Geometry CompressionabstractLearning-based point cloud compression has achieved tremendous progress in recent years. However, existing methods often train an optimal occupancy distribution predictor for the entire train dataset in an amortization sense, which struggles to handle point clouds with unique characteristics. In this work, we focus on the lossless point cloud compression, and propose a novel context-adaptive entropy model to achieve adaptive occupancy prediction. Specifically, given a baseline entropy model and a point cloud, we firstly integrate adapters into diverse feature extraction modules. These adapters are then trained to be specifically attuned to the input cloud. Finally, the trained adapter parameters are encoded and transmitted along with the point cloud bitstream, which allow us to recover the integrated model in decoder. The experimental results demonstrate that our method can enhance the performance of the entropy model, especially improving the compression performance of data that performs poorly in conventional methods. Wenbo Zhao 0004, Daxin Li, Junjun Jiang, Xianming Liu 0005 |
ICIP | 2 |
| 2023 | Self-Supervised Arbitrary-Scale Implicit Point Clouds UpsamplingabstractPoint clouds upsampling (PCU), which aims to generate dense and uniform point clouds from the captured sparse input of 3D sensor such as LiDAR, is a practical yet challenging task. It has potential applications in many real-world scenarios, such as autonomous driving, robotics, AR/VR, etc. Deep neural network based methods achieve remarkable success in PCU. However, most existing deep PCU methods either take the end-to-end supervised training, where large amounts of pairs of sparse input and dense ground-truth are required to serve as the supervision; or treat up-scaling of different factors as independent tasks, where multiple networks are required for different scaling factors, leading to significantly increased model complexity and training time. In this article, we propose a novel method that achieves self-supervised and magnification-flexible PCU simultaneously. No longer explicitly learning the mapping between sparse and dense point clouds, we formulate PCU as the task of seeking nearest projection points on the implicit surface for seed points. We then define two implicit neural functions to estimate projection direction and distance respectively, which can be trained by the pretext learning tasks. Moreover, the projection rectification strategy is tailored to remove outliers so as to keep the shape of object clear and sharp. Experimental results demonstrate that our self-supervised learning based scheme achieves competitive or even better performance than state-of-the-art supervised methods. Wenbo Zhao 0004, Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Local Surface Descriptor for Geometry and Feature Preserved Mesh Denoisingabstract3D meshes are widely employed to represent geometry structure of 3D shapes. Due to limitation of scanning sensor precision and other issues, meshes are inevitably affected by noise, which hampers the subsequent applications. Convolultional neural networks (CNNs) achieve great success in image processing tasks, including 2D image denoising, and have been proven to own the capacity of modeling complex features at different scales, which is also particularly useful for mesh denoising. However, due to the nature of irregular structure, CNNs-based denosing strategies cannot be trivially applied for meshes. To circumvent this limitation, in the paper, we propose the local surface descriptor (LSD), which is able to transform the local deformable surface around a face into 2D grid representation and thus facilitates the deployment of CNNs to generate denoised face normals. To verify the superiority of LSD, we directly feed LSD into the classical Resnet without any complicated network design. The extensive experimental results show that, compared to the state-of-the-arts, our method achieves encouraging performance with respect to both objective and subjective evaluations. Wenbo Zhao 0004, Xianming Liu 0005, Junjun Jiang, Debin Zhao, Ge Li 0002, Xiangyang Ji |
AAAI | 1 |
| 2022 | Self-Supervised Arbitrary-Scale Point Clouds Upsampling via Implicit Neural RepresentationabstractPoint clouds upsampling is a challenging issue to gener-ate dense and uniform point clouds from the given sparse input. Most existing methods either take the end-to-end su-pervised learning based manner, where large amounts of pairs of sparse input and dense ground-truth are exploited as supervision information; or treat up-scaling of different scale factors as independent tasks, and have to build multiple networks to handle upsampling with varying factors. In this paper, we propose a novel approach that achieves self-supervised and magnification-flexible point clouds upsampling simultaneously. We formulate point clouds upsampling as the task of seeking nearest projection points on the implicit surface for seed points. To this end, we define two implicit neural functions to estimate projection direction and distance respectively, which can be trained by two pretext learning tasks. Experimental results demonstrate that our self-supervised learning based scheme achieves competitive or even better performance than supervised learning based state-of-the-art methods. The source code is publicly available at https://github.com/xnowbzhaolsapcu. Wenbo Zhao 0004, Xianming Liu 0005, Zhiwei Zhong 0001, Junjun Jiang, Wei Gao 0003, Ge Li 0002, Xiangyang Ji |
CVPR | 1 |
| 2021 | NormalNet: Learning-Based Mesh Normal Denoising via Local Partition NormalizationabstractMesh denoising is a critical technology in geometry processing that aims to recover high-fidelity 3D mesh models of objects from noise-corrupted versions. In this work, we propose a learning-based mesh normal denoising scheme, calledNormalNet, which employs deep networks to find the correlation between the volumetric representation and denoised face normal. Overall,NormalNetfollows the iterative framework of filtering-based mesh denoising. During each iteration, firstly, a local partition normalization strategy is applied to split the local structure around each face into dense voxels, in which both the structure and face normal information can be preserved during this transformation. Benefiting from the thorough information preservation, we can use simple residual networks, which employ the volumetric representation as the input and produce the learned denoised face normal, to achieve satisfactory results. Finally, the vertex positions are updated according to the denoised normals. Besides introducing normalization into mesh denoising, our main contributions include a classification-based training faces selection strategy for balancing the training set and a mismatched-faces rejection strategy for removing the mismatched faces between noisy mesh and ground truth. Compared to state-of-the-art works,NormalNetcan effectively remove noise while preserving the original features and avoiding pseudo-features. Wenbo Zhao 0004, Xianming Liu 0005, Yongsen Zhao, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Graph-Based Feature-Preserving Mesh Normal FilteringabstractDistinguishing between geometric features and noise is of paramount importance for mesh denoising. In this paper, a graph-based feature-preserving mesh normal filtering scheme is proposed, which includes two stages: graph-based feature detection and feature-aware guided normal filtering. In the first stage, faces in the input noisy mesh are represented by patches, which are then modelled as weighted graphs. In this way, feature detection can be cast as a graph-cut problem. Subsequently, an iterative normalized cut algorithm is applied on each patch to separate the patch into smooth regions according to the detected features. In the second stage, a feature-aware guidance normal is constructed for each face, and guided normal filtering is applied to achieve robust feature-preserving mesh denoising. The results of experiments on synthetic and real scanned models indicate that the proposed scheme outperforms state-of-the-art mesh denoising works in terms of both objective and subjective evaluations. Wenbo Zhao 0004, Xianming Liu 0005, Shiqi Wang 0001, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | Region-of-interest based coding scheme for synthesized videoabstractIn many multimedia applications, such as online speech, video chat and online conference, multiple source videos are synthesized in a single scene for explicit presentation and the synthesized video is compressed for transmission. The source video with important contents deserves more compression resources for quality preservation under the bandwidth constraint. To address this problem, a region-of-interest (ROI) based coding scheme for synthesized video is proposed in this paper aiming at achieve better and consistent quality for ROI source videos with the bitrate meeting the constraint bandwidth. In the proposed coding scheme, ROI based rate-distortion (R-D) models are established, in which different R-D models are built for different source video. Then an objective function is defined with respect to the video quality and the consistency of video quality. By minimizing the objective function, the optimal quantization parameters for the ROI and non-ROI source videos are obtained. The experimental results show that the proposed coding scheme achieves better and consistent quality for ROI source videos. Wenbo Zhao 0004, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Debin Zhao |
VCIP | 1 |