VLDB 2026 Research / reviewers in the wild / expert
Zunjie Zhu
dblp:216/2729
· DBLP profile ↗
23ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0001-6107-4538ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CMTNet: A hybrid CNN-Mamba-Transformer network for point cloud salient object detection
Langtao Gan, Hongfa Wen, Haicheng Tu, Zunjie Zhu |
Neural Networks | 5 |
| 2026 | LLFeat: Noise-Aware Feature Matching Under Various Low-Light Conditions
Longjian Zeng, Zunjie Zhu, Ming Lu 0002, Bolun Zheng, Rongfeng Lu, Tingyu Wang 0002, Zhongtian Zheng, Yaoqi Sun, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Hybrid Debiasing Transformer With Adaptive Regularization for Video Moment LocalizationabstractVideo Moment Localization (VML) is a task that seeks to pinpoint the most pertinent segment within an untrimmed video using a linguistic query. Previous works expose the severe data bias issues in VML and note that models avoid understanding visual-textual content by adapting the timestamp distribution. The work investigates data biases from both intrinsic and extrinsic perspectives: The former arises primarily from moment boundary ambiguity and the inputoutput information imbalance. The latter is attributed to the longtail distribution and the semantic bias resulting from the limited tail samples. To reduce the issues, we develop a hybrid multimodal debiasing network with a temporal consistency constraint for VML. Firstly, we propose a multi-temporal Transformer to alleviate boundary ambiguity by merging frame-wise features into segment-wise representations and dynamically aligning with moment boundaries. Subsequently, we implement a temporal consistency constraint to accentuate action information from complex moment context and mitigate the intrinsic bias caused by information imbalance. Moreover, we develop a hybrid linguistic activation module to mitigate the long-tail bias, which offers prior guidance to emphasize distinguishing clues from tail samples. Additionally, we introduce the prior-guided Transformer to alleviate the semantic bias by learning the global semantics of sentences, thereby circumventing the tail-sample overfitting issue. Comprehensive experiments demonstrate the efficacy of our proposed method across three datasets. Jiong Yin, Liang Li 0003, Chenggang Yan 0001, Hongkui Wang, Yaoqi Sun, Zunjie Zhu |
IEEE Trans. Multim. | 7 |
| 2025 | 4D Gaussian Splatting SLAMabstractSimultaneously localizing camera poses and constructing Gaussian radiance fields in dynamic scenes establish a crucial bridge between 2D images and the 4D real world. Instead of removing dynamic objects as distractors and reconstructing only static environments, this paper proposes an efficient architecture that incrementally tracks camera poses and establishes the 4D Gaussian radiance fields in unknown scenarios by using a sequence of RGB-D images. First, by generating motion masks, we obtain static and dynamic priors for each pixel. To eliminate the influence of static scenes and improve the efficiency on learning the motion of dynamic objects, we classify the Gaussian primitives into static and dynamic Gaussian sets, while the sparse control points along with an MLP is utilized to model the transformation fields of the dynamic Gaussians. To more accurately learn the motion of dynamic Gaussians, a novel 2D optical flow map reconstruction algorithm is designed to render optical flows of dynamic objects between neighbor images, which are further used to supervise the 4D Gaussian radiance fields along with traditional photometric and geometric constraints. In experiments, qualitative and quantitative evaluation results show that the proposed method achieves robust tracking and high-quality view synthesis performance in real-world environments. Yanyan Li 0001, Youxu Fang, Zunjie Zhu, Kunyi Li, Federico Tombari |
ICCV | 3 |
| 2025 | ThermalGaussian: Thermal 3D Gaussian SplattingabstractThermography is especially valuable for the military and other users of surveillance cameras. Some recent methods based on Neural Radiance Fields (NeRF) are proposed to reconstruct the thermal scenes in 3D from a set of thermal and RGB images. However, unlike NeRF, 3D Gaussian splatting (3DGS) prevails due to its rapid training and real-time rendering. In this work, we propose ThermalGaussian, the first thermal 3DGS approach capable of rendering high-quality images in RGB and thermal modalities. We first calibrate the RGB camera and the thermal camera to ensure that both modalities are accurately aligned. Subsequently, we use the registered images to learn the multimodal 3D Gaussians. To prevent the overfitting of any single modality, we introduce several multimodal regularization constraints. We also develop smoothing constraints tailored to the physical characteristics of the thermal modality.
Besides, we contribute a real-world dataset named RGBT-Scenes, captured by a hand-hold thermal-infrared camera, facilitating future research on thermal scene reconstruction. We conduct comprehensive experiments to show that ThermalGaussian achieves photorealistic rendering of thermal images and improves the rendering quality of RGB images. With the proposed multimodal regularization constraints, we also reduced the model's storage cost by 90\%. Our project page is at https://thermalgaussian.github.io/. Rongfeng Lu, Zunjie Zhu, Yuhang Qin, Ming Lu 0002, Chenggang Yan 0001, Anke Xue |
ICLR | 3 |
| 2025 | K-Buffers: A Plug-in Method for Enhancing Neural Fields with Multiple BuffersabstractNeural fields are now the central focus of research in 3D vision and computer graphics. Existing methods mainly focus on various scene representations, such as neural points and 3D Gaussians. However, few works have studied the rendering process to enhance the neural fields. In this work, we propose a plug-in method named K-Buffers that leverages multiple buffers to improve the rendering performance. Our method first renders K buffers from scene representations and constructs K pixel-wise feature maps. Then, We introduce a K-Feature Fusion Network (KFN) to merge the K pixel-wise feature maps. Finally, we adopt a feature decoder to generate the rendering image. We also introduce an acceleration strategy to improve rendering speed and quality. We apply our method to well-known radiance field baselines, including neural point fields and 3D Gaussian Splatting (3DGS). Extensive experiments demonstrate that our method effectively enhances the rendering performance of neural point fields and 3DGS. Haofan Ren, Zunjie Zhu, Xiang Chen 0015, Ming Lu 0002, Rongfeng Lu, Chenggang Yan 0001 |
IJCAI | 2 |
| 2025 | DepthDark: Robust Monocular Depth Estimation for Low-Light EnvironmentsabstractIn recent years, foundation models for monocular depth estimation have received increasing attention. Current methods mainly address typical daylight conditions, but their effectiveness notably decreases in low-light environments. There is a lack of robust foundational models for monocular depth estimation specifically designed for low-light scenarios. This largely stems from the absence of large-scale, high-quality paired depth datasets for low-light conditions and the effective parameter-efficient fine-tuning (PEFT) strategy. To address these challenges, we propose DepthDark, a robust foundation model for low-light monocular depth estimation. We first introduce a flare-simulation module and a noise-simulation module to accurately simulate the imaging process under nighttime conditions, producing high-quality paired depth datasets for low-light conditions. Additionally, we present an effective low-light PEFT strategy that utilizes illumination guidance and multiscale feature fusion to enhance the model's capability in low-light environments. Our method achieves state-of-the-art depth estimation performance on the challenging nuScenes-Night and RobotCar-Night datasets, validating its effectiveness using limited training data and computing resources. Longjian Zeng, Zunjie Zhu, Rongfeng Lu, Ming Lu 0002, Bolun Zheng, Chenggang Yan 0001, Anke Xue |
ACM Multimedia | 2 |
| 2025 | Consistency perception network for 360° omnidirectional salient object detection
Hongfa Wen, Zunjie Zhu, Xiaofei Zhou 0003, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 2 |
| 2025 | Hierarchical Frequency-Based Upsampling and Refining for HEVC Compressed Video EnhancementabstractVideo compression artifacts arise from quantization applied in the frequency domain. Video quality enhancement aims to reduce such compression artifacts and reconstruct a visually pleasant result. While existing methods effectively reduce artifacts in the spatial domain, they often overlook the rich frequency domain information, especially in addressing multi-scale compression artifacts. This work introduces a frequency-domain upsampling strategy within a multi-scale framework, specifically designed to focus on high-frequency details rather than simply blending neighboring pixels during the upsampling process. Our proposed hierarchical frequency-based upsampling and refinement neural network (HFUR) consists of two modules: implicit frequency upsampling (ImpFreqUp) and hierarchical and iterative refinement (HIR). ImpFreqUp exploits the DCT-domain prior derived through an implicit DCT transform, and accurately reconstructs the DCT-domain signal via a coarse-to-fine transfer. Additionally, HIR is introduced to facilitate cross-collaboration and information compensation between the scales, further refining the feature maps and promoting the visual quality of the final output. We demonstrate the effectiveness of the proposed modules via ablation experiments and visualized results. Experimental results demonstrate that HFUR outperforms the state-of-the-art methods up to 0.13dB/0.17dB on both constant bit rate and constant QP modes. The code is available athttps://github.com/zqqqyu/HFUR. Qianyu Zhang 0002, Bolun Zheng, Xingying Chen, Zunjie Zhu, Canjin Wang, Zongpeng Li, Xu Jia 0012, Chengang Yan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Rethinking Boundary Discontinuity Problem for Oriented Object DetectionabstractOriented object detection has been developed rapidly in the past few years, where rotation equivariance is crucial for detectors to predict rotated boxes. It is expected that the prediction can maintain the corresponding rotation when objects rotate, but severe mutation in angular prediction is sometimes observed when objects rotate near the boundary angle, which is well-known boundary discontinuity problem. The problem has been long believed to be caused by the sharp loss increase at the angular boundary, and widely used joint-optim IoU-like methods deal with this problem by loss-smoothing. However, we experimentally find that even state-of-the-art IoU-like methods actually fail to solve the problem. On further analysis, we find that the key to solution lies in encoding mode of the smoothing function rather than in joint or independent optimization. In existing IoU-like methods, the model essentially attempts to fit the angular relationship between box and object, where the break point at angular boundary makes the predictions highly unstable. To deal with this issue, we propose a dual-optimization paradigm for angles. We decouple reversibility and joint-optim from single smoothing function into two distinct entities, which for the first time achieves the objectives of both correcting angular boundary and blending angle with other parameters. Extensive experiments on multiple datasets show that boundary discontinuity problem is well-addressed. More-over, typical IoU-like methods are improved to the same level without obvious performance gap. The code is available at https://github.com/hangxu-cv/cvpr24acm. Xinyuan Liu 0003, Yike Ma, Zunjie Zhu, Chenggang Yan 0001 |
CVPR | 5 |
| 2024 | Superpixel-based Efficient Sampling for Learning Neural Fields from Large InputabstractIn recent years, neural field-based methods for synthesizing novel views have gained popularity due to their exceptional rendering quality and fast training speed. However, the computational cost of volumetric rendering has significantly increased with the advancement of camera technology and the subsequent rise in average camera resolution. Despite extensive efforts to accelerate the training process, the training duration remains unacceptable for high-resolution inputs. Therefore, it's crucial to develop efficient sampling methods to optimize the learning process of neural fields from large inputs. In this paper, we present a new technique called Superpixel-based Efficient Sampling (SES) to improve the learning efficiency of neural fields. Our approach optimizes pixel-level ray sampling by segmenting the error map into multiple superpixels and dynamically updating their errors during training to increase ray sampling in superpixel areas with higher rendering errors. Compared with other methods, our approach leverages the flexibility of superpixels, effectively reducing redundant sampling while considering local information. Our method not only speeds up the learning process but also enhances the rendering quality learned from large inputs. We conduct extensive experiments to evaluate the effectiveness of our method across several baselines and datasets. The code will be released. Zhongwei Xuan, Zunjie Zhu, Shuai Wang 0003, Haibing Yin, Hongkui Wang, Ming Lu 0002 |
ACM Multimedia | 2 |
| 2024 | GoLDFormer: A global-local deformable window transformer for efficient image restoration
Bolun Zheng, Chenggang Yan 0001, Zunjie Zhu, Tingyu Wang 0002, Gregory Slabaugh, Shanxin Yuan |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | GINet:Graph interactive network with semantic-guided spatial refinement for salient object detection in optical remote sensing images
Chenwei Zhu, Xiaofei Zhou 0003, Liuxin Bao, Hongkui Wang, Shuai Wang 0003, Zunjie Zhu, Chenggang Yan 0001, Jiyong Zhang 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2024 | Non-local degradation modeling for spatially adaptive single image super-resolution
Qianyu Zhang 0002, Bolun Zheng, Zongpeng Li, Yu Liu 0005, Zunjie Zhu, Gregory Slabaugh, Shanxin Yuan |
Neural Networks | 5 |
| 2024 | TMNet: Triple-modal interaction encoder and multi-scale fusion decoder network for V-D-T salient object detection
Bin Wan, Chengtao Lv, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Hongkui Wang, Chenggang Yan 0001 |
Pattern Recognit. | 5 |
| 2024 | Learning Cross-View Geo-Localization Embeddings via Dynamic Weighted Decorrelation RegularizationabstractIn the domain of cross-view geo-localization, the challenge lies in accurately matching images captured from distinct perspectives, such as aerial drone imagery and satellite imagery of the same geographical location. Existing methods predominantly concentrate on minimizing distances between feature embeddings in the representational space, inadvertently overlooking the significance of reducing embedding redundancy. This oversight potentially hampers the extraction of diverse and distinctive visual patterns critical for precise localization. This work argues that minimizing embedding redundancy is a pivotal factor in enhancing a model’s ability to discriminate diverse scene characteristics. To support this claim, we introduce a straightforward yet effective regularization technique, termed dynamic weighted decorrelation regularization (DWDR). DWDR serves to actively promote the learning of orthogonal feature channels within neural networks. By dynamically adjusting weights, DWDR targets the minimization of interchannel correlations, guiding the correlation matrix toward diagonality, indicative of independence among channels. The dynamic weighting mechanism adaptively prioritizes the decorrelation of channels that remain highly correlated throughout training. Additionally, we devise a symmetrical sampling strategy for cross-view scenarios to ensure that the training examples are balanced across different imaging platforms in a batch. Despite its simplicity, the integration of DWDR and the proposed sampling scheme yields remarkable performance across four extensive benchmark datasets: University-1652, CVUSA, CVACT, and VIGOR. Notably, in stringent conditions, such as when constrained to exceedingly compact feature dimensions of 64, our methodology significantly outperforms conventional baselines, thereby affirming its efficacy and robustness under challenging constraints. Tingyu Wang 0002, Zhedong Zheng, Zunjie Zhu, Yaoqi Sun, Chenggang Yan 0001, Yi Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Reducing Intrinsic and Extrinsic Data Biases for Moment Localization with Natural LanguageabstractMoment Localization with Natural Language (MLNL) aims to locate the target moment from an untrimmed video by a linguistic query. Recent works reveal the severe data bias problem in MLNL and point out that the multi-modal content may not be understood by fitting the timestamp distribution. In this paper, we study the data biases on the intrinsic and extrinsic aspects: the former is mainly caused by the ambiguity of the moment boundary and the information imbalance between input and output; The latter results from the long-tail distribution of moments in MLNL datasets. To alleviate this, we propose a hybrid multi-modal debiasing network with temporal consistency constraint for MLNL. Specifically, we first design the multi-temporal Transformer to mitigate the ambiguity of boundary by integrating frame-wise features into segment-wise and dynamically matching with moment boundaries. Then, we introduce the temporal consistency constraint that highlights the action information in complex moment content to overcome the intrinsic bias from information imbalance.Furthermore, we design the hybrid linguistic activating module with external knowledge to relieve the extrinsic bias, which introduces a prior guidance to focus the discriminative information from the tail samples. Extensive experiments on three public datasets demonstrate that our model outperforms the existing methods. Jiong Yin, Liang Li 0003, Chenggang Yan 0001, Lei Zhang 0119, Zunjie Zhu |
ACM Multimedia | 6 |
| 2023 | GFNet: gated fusion network for video saliency prediction
Songhe Wu, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001 |
Appl. Intell. | 5 |
| 2023 | SMINet: Semantics-aware multi-level feature interaction network for surface defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Haibing Yin, Ji Hu 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Aggregating transformers and CNNs for salient object detection in optical remote sensing images
Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Haibing Yin, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 5 |
| 2022 | PlaneFusion: Real-Time Indoor Scene Reconstruction With Planar PriorabstractReal-time dense SLAM techniques aim to reconstruct the dense three-dimensional geometry of a scene in real time with an RGB or RGB-D sensor. An indoor scene is an important type of working environment for these techniques. The planar prior can be used in this scenario to improve the reconstruction quality, especially for large low-texture regions that commonly occur in an indoor scene. This article fully explores the planar prior in a dense SLAM pipeline. First, we propose a novel plane detection and segmentation method that runs at 200 Hz on a modern graphics processing unit. Our algorithm for constructing global plane constraints is very efficient; hence, we use it in the process of each input frame for the camera pose estimation while maintaining the real-time performance. Second, we propose herein a plane-based map representation that greatly reduces the memory footprint of plane regions while keeping the geometric details on planes. The experiments reveal that our system yields superior reconstruction results with planar information running at more than 30 fps. Aside from speed and storage improvements, our technique also handles the low-texture problem in plane regions. Bingjian Gong, Zunjie Zhu, Chenggang Yan 0001, Zhiguo Shi 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Real-time Indoor Scene Reconstruction with RGBD and Inertial InputabstractCamera motion estimation is a key technique for 3D scene reconstruction. Previous works usually assume slow camera motions, which limit the usage in many real cases. We propose an end-to-end 3D reconstruction system which combines color, depth and inertial measurements to achieve robust reconstruction with fast sensor motions. Our framework utilizes extended Kalman filter to fuse the three kinds of information and involve an iterative method to jointly optimize feature correspondences, camera poses and scene geometry. We also propose a novel geometry-aware patch deformation technique to adapt the feature appearance in image domain, leading to a more accurate feature matching under fast camera motions. Experiments show that our patch deformation method improves the accuracy of feature tracking, and our 3D reconstruction framework outperforms the state-of-the-art solutions under fast camera motions. Zunjie Zhu, Feng Xu 0005, Chenggang Yan 0001, Xinhong Hao, Xiangyang Ji, Yongdong Zhang 0001, Qionghai Dai |
ICME | 1 |
| 2019 | Real-time indoor scene reconstruction with Manhattan assumption
Zunjie Zhu, Feng Xu 0005, Chenggang Yan 0001, Bingjian Gong, Yongdong Zhang 0001, Qionghai Dai |
Multim. Tools Appl. | 1 |