Zhongkai Zhou

dblp:251/2193 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 R-FGDepth: Towards foundation models for recurrent depth learning with frequency-Guided initialization and refinement
Zhaoxin Fan, Zhongkai Zhou
Pattern Recognit.3
2025 Uncertainty-Aware Multimodal MRI Fusion for HIV-Associated Asymptomatic Neurocognitive Impairment Prediction
Zige Chen, Wei Wang 0411, Zhongkai Zhou, Yuqi Fang, Caifeng Shan
MICCAI (15)4
2024 CrackYOLO: towards efficient dam crack detection for underwater scenes
Shen Shao, Xinnan Fan, Yuanxue Xin, Zhongkai Zhou, Sisi Zhu
Pattern Anal. Appl.5
2024 Recurrent Multiscale Feature Modulation for Geometry Consistent Depth Learning
abstract
The U-Net-like coarse-to-fine network design is currently the dominant choice for dense prediction tasks. Although this design can often achieve competitive performance, it suffers from some inherent limitations, such as training error propagation from low to high resolution and the dependency on the deeper and heavier backbones. To design an effective network that performs better, we instead propose Recurrent Multiscale Feature Modulation (R-MSFM), a new lightweight network design for self-supervised monocular depth estimation. R-MSFM extracts per-pixel features, builds a multiscale feature modulation module, and performs recurrent depth refinement through a parameter-shared decoder at a fixed resolution. This network design enables our R-MSFM to maintain a more lightweight architecture and fundamentally avoid error propagation caused by the coarse-to-fine design. Furthermore, we introduce the mask geometry consistency loss to facilitate our R-MSFM for geometry consistent depth learning. This loss penalizes the inconsistency of the estimated depths between adjacent views within the nonoccluded and nonstationary regions. Experimental results demonstrate the superiority of our proposed R-MSFM both at model size and inference speed, and show state-of-the-art results on two datasets: KITTI and Make3D.
Zhongkai Zhou, Xinnan Fan, Yuanxue Xin, Dongliang Duan, Liuqing Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 R-LKDepth: Recurrent Depth Learning With Larger Kernel
abstract
Monocular depth estimation is a critical task with significant potential for real-time applications. However, current networks for monocular depth estimation exhibit a trade-off between computational overhead and accuracy. Current top-performing coarse-to-fine networks excel in building multi-scale features and thus get a satisfactory result, but suffer from an excessive count of model parameters of its deeper backbone. In contrast, recurrent networks significantly save the use of model parameters, but their performance is often restricted by the smaller receptive field, leading to insufficient local evidence to handle complex objects and homogeneous regions. In this letter, we aim to alleviate the limitation and thus propose an efficient recurrent network R-LKDepth. Specifically, we first introduce the larger kernel scheme with a careful design to effectively extract the image features with larger receptive fields. Furthermore, we propose a group-wise multi-scale self-attention based refine-feed (GMSA-RF) module, which efficiently aggregates multi-scale information for recurrent refinement procedure. The resulting network, R-LKDepth, significantly improves the depth estimation performance over complex objects and homogeneous regions and achieves state-of-the-art performance on standard benchmarks with superior cross-dataset generalization capability.
Zhongkai Zhou, Xinnan Fan, Yuanxue Xin
IEEE Signal Process. Lett.1
2023 LCIF-Net: Local criss-cross attention based optical flow method using multi-scale image features and feature pyramid
Zige Wang, Zhen Chen 0004, Congxuan Zhang, Zhongkai Zhou, Hao Chen 0060
Signal Process. Image Commun.4
2022 RAFM: Recurrent Atrous Feature Modulation for Accurate Monocular Depth Estimating
abstract
Current top-performing coarse-to-fine monocular depth estimation systems mainly depend on deeper backbones, such as a full ResNet50. These systems benefit from the powerful multiscale feature representations but suffer from the high computational costs and memory overheads. Conversely, the recurrent depth refinement systems fulfill the lightweight architecture, while the local multiscale context often limits their accuracies. To handle this problem, we propose an atrous spatial (AS) module by utilizing atrous convolutions with different rates, which efficiently captures the multiscale context. We further assemble our AS module to the current recurrent depth refinement system R-MSFM, constructing a powerful monocular depth estimation system called RAFM. To evaluate its effectiveness, we conduct the comprehensive experiments and show that our RAFM outperforms the previous SOTA methods in all metrics on KITTI while keeping the real-time speed. For accuracy, our RAFM achieves a square relative error (Sq Rel) of 0.702 on KITTI, a 12.6% error reduction from the previous SOTA method (0.802). For model size and inference speed, our RAFM gets a frame rate of 33fps on a GPU with 7.5 M parameters, which is faster and more lightweight than the recent SOTA method (5fps and 128 M). The code is available athttps://github.com/jsczzzk/RAFM.
Xinnan Fan, Zhongkai Zhou, Yuanxue Xin
IEEE Signal Process. Lett.2
2022 Self-Attention-Based Multiscale Feature Learning Optical Flow With Occlusion Feature Map Prediction
abstract
Even though optical flow approaches based on convolutional neural networks have achieved remarkable performance with respect to both accuracy and efficiency, large displacements and motion occlusions remain challenges for most existing learning-based models. To address the abovementioned issues, we propose in this paper a self-attention-based multiscale feature learning optical flow computation method with occlusion feature map prediction. First, we exploit a self-attention mechanism-based multiscale feature learning module to compensate for large displacement optical flows, and the presented module is able to capture long-range dependencies from the input frames. Second, we design a simple but effective self-learning module to acquire an occlusion feature map, in which the predicted occlusion map is utilized to correct the optical flow estimation in occluded areas. Third, we explore a hybrid loss function that integrates the photometric and smoothness losses into the classical endpoint error (EPE)-based loss to ensure the accuracy and robustness of the presented network. Finally, we compare the proposed method with some state-of-the-art approaches using the MPI-Sintel and KITTI test databases. The experimental results demonstrate that the proposed method achieved competitive performance with respect to both accuracy and robustness, and it produced the better results compared to other methods under large displacements and motion occlusions.
Congxuan Zhang, Zhongkai Zhou, Zhen Chen 0004, Weiming Hu 0004, Ming Li 0056, Shaofeng Jiang
IEEE Trans. Multim.2
2021 R-MSFM: Recurrent Multi-Scale Feature Modulation for Monocular Depth Estimating
abstract
In this paper, we propose Recurrent Multi-Scale Feature Modulation (R-MSFM), a new deep network architecture for self-supervised monocular depth estimation. R-MSFM extracts per-pixel features, builds a multi-scale feature modulation module, and iteratively updates an inverse depth through a parameter-shared decoder at the fixed resolution. This architecture enables our R-MSFM to maintain semantically richer while spatially more precise representations and avoid the error propagation caused by the traditional U-Net-like coarse-to-fine architecture widely used in this domain, resulting in strong generalization and efficient parameter count. Experimental results demonstrate the superiority of our proposed R-MSFM both at model size and inference speed, and show the state-of-the-art results on the KITTI benchmark. Code is available at https://github.com/jsczzzk/R-MSFM
Zhongkai Zhou, Xinnan Fan, Yuanxue Xin
ICCV1