EDBT 2026 Demo / reviewers in the wild / expert
Zhibo Rao
dblp:239/8521
· DBLP profile ↗
21ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0001-7832-2913ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FACT: Frequency adaptive consistency tuning for efficient zero-shot unified image restoration
Zhidong Zhu, Shuzhen Yu, Jinhao Zhu, Zhibo Rao, Qiaofeng Ou, Xing Li 0040 |
Neurocomputing | 5 |
| 2026 | Adaptive control for 3D Gaussian splatting: a systematic regularization framework
Wenxuan Xiong, Fusheng Wang 0015, Xing Li 0040, Zhidong Zhu, Zhibo Rao |
Vis. Comput. | 7 |
| 2025 | WCG-Net: Warping Consistency Compensation Guided Multi-Feature Fusion For Stereo MatchingabstractDespite the significant progress achieved by iterative optimization-based stereo matching methods, a critical challenge persists: these state-of-the-art models continue to face difficulties when handling ill-posed regions. This stems from lighting variations and viewpoint differences, which may cause the feature distributions extracted from the left and right images to differ, resulting in unreliable cost volume construction in ill-posed regions. To remedy this issue, we propose a novel network for stereo matching, named WCG-Net. In WCG-Net, we develop a warping consistency compensation module (WCCM) that employs consistency attention to identify feature differences and generate cross-view features, which are then used to construct the warping correlation volume. By introducing warping correlation volume into warping-guided fusion recurrent unit (WFRU), our method refines the disparity map iteratively, focusing on correcting errors in ill-posed regions. Extensive experimental evaluation demonstrates that WCG-Net achieves competitive results on the KITTI and ETH3D datasets, with particularly superior performance compared to other state-of-the-art algorithms on the KITTI dataset. Zhibo Rao, Zhen Chen 0004, Congxuan Zhang |
ICME | 3 |
| 2025 | Constraining multimodal distribution for domain adaptation in stereo matching
Zhelun Shen, Chenming Wu, Zhibo Rao, Lina Liu 0010, Yuchao Dai, Liangjun Zhang |
Pattern Recognit. | 4 |
| 2025 | FRCL-MNER: A Finer Grained Rank-Based Contrastive Learning Framework for Multimodal NERabstractMultimodal named entity recognition (MNER) is an emerging field that aims to automatically detect named entities and classify their categories, utilizing input text and auxiliary resources such as images. While previous studies have leveraged object detectors to preprocess images and fuse textual semantics with corresponding image features, these methods often overlook the potential finer grained information within each modality and may exacerbate error propagation due to predetection. To address these issues, we propose a finer grained rank-based contrastive learning (FRCL) framework for MNER. This framework employs a global-level contrastive learning to align multimodal semantic features and a Top-K rank-based mask strategy to construct positive-negative pairs, thereby learning a finer grained multimodal interaction representation. Experimental results from three well-known social media datasets reveal that our approach surpasses existing strong baselines, and achieves up to a 1.54% improvement on the Twitter2015 dataset. Extensive discussions further confirm the effectiveness of our approach. We will release the source code on https://github.com/augusyan/FRCL. Tianwei Yan 0001, Shan Zhao 0002, Wentao Ma 0003, Shezheng Song, Chengyu Wang 0008, Zhibo Rao, Shizhao Chen, Zhigang Luo, Xinwang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | GPDF-Net: geometric prior-guided stereo matching with disparity fusion refinement
Congxuan Zhang, Zhibo Rao, Zhen Chen 0004, Zige Wang, Ke Lu 0002 |
Vis. Comput. | 3 |
| 2025 | WaveUIR: wavelet-based guided transformer model for efficient universal image restoration
Zhidong Zhu, Zhibo Rao, Jinhao Zhu, Qiaofeng Ou, Xing Li 0040 |
Vis. Comput. | 3 |
| 2024 | MaskRecon: High-quality human reconstruction via masked autoencoders using a single RGB-D image
Xing Li 0040, Yangyu Fan, Zhibo Rao, Yu Duan 0001, Shiya Liu |
Neurocomputing | 4 |
| 2023 | Masked Representation Learning for Domain Generalized Stereo MatchingabstractRecently, many deep stereo matching methods have begun to focus on cross-domain performance, achieving impressive achievements. However, these methods did not deal with the significant volatility of generalization performance among different training epochs. Inspired by masked representation learning and multi-task learning, this paper designs a simple and effective masked representation for domain generalized stereo matching. First, we feed the masked left and complete right images as input into the models. Then, we add a lightweight and simple decoder following the feature extraction module to recover the original left image. Finally, we train the models with two tasks (stereo matching and image reconstruction) as a pseudo-multi-task learning framework, promoting models to learn structure information and to improve generalization performance. We implement our method on two well-known architectures (CFNet and LacGwcNet) to demonstrate its effectiveness. Experimental results on multi-datasets show that: (1) our method can be easily plugged into the current various stereo matching models to improve generalization performance; (2) our method can reduce the significant volatility of generalization performance among different training epochs; (3) we find that the current methods prefer to choose the best results among different training epochs as generalization performance, but it is impossible to select the best performance by ground truth in practice. Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen, Xing Li 0040 |
CVPR | 1 |
| 2023 | Digging Into Uncertainty-Based Pseudo-Label for Robust Stereo MatchingabstractDue to the domain differences and unbalanced disparity distribution across multiple datasets, current stereo matching approaches are commonly limited to a specific dataset and generalize poorly to others. Such domain shift issue is usually addressed by substantial adaptation on costly target-domain ground-truth data, which cannot be easily obtained in practical settings. In this paper, we propose to dig into uncertainty estimation for robust stereo matching. Specifically, to balance the disparity distribution, we employ a pixel-level uncertainty estimation to adaptively adjust the next stage disparity searching space, in this way driving the network progressively prune out the space of unlikely correspondences. Then, to solve the limited ground truth data, an uncertainty-based pseudo-label is proposed to adapt the pre-trained model to the new domain, where pixel-level and area-level uncertainty estimation are proposed to filter out the high-uncertainty pixels of predicted disparity maps and generate sparse while reliable pseudo-labels to align the domain gap. Experimentally, our method shows strong cross-domain, adapt, and joint generalization and obtains 1st place on the stereo task of Robust Vision Challenge 2020. Additionally, our uncertainty-based pseudo-labels can be extended to train monocular depth estimation networks in an unsupervised way and even achieves comparable performance with the supervised methods. Zhelun Shen, Xibin Song, Yuchao Dai, Dingfu Zhou, Zhibo Rao, Liangjun Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Rethinking Training Strategy in Stereo MatchingabstractIn stereo matching, various learning-based approaches have shown impressive performance in solving traditional difficulties on multiple datasets. While most progress is obtained on a specific dataset with a dataset-specific network design, the performance on the single dataset and cross dataset affected by training strategy is often ignored. In this article, we analyze the relationship between different training strategies and performance by retraining some representative state-of-the-art methods (e.g., geometry and context network (GC-Net), pyramid stereo matching network (PSM-Net), and guided aggregation network (GA-Net), etc.). According to our research, it is surprising that the performance of networks on single or cross datasets is significantly improved by pre-training and data augmentation without any particular structure acquirement. Based on this discovery, we improve our previous non-local context attention network (NLCA-Net) to NLCA-Net v2 and train it with the novel strategy and rethink the training strategy of stereo matching concurrently. The quantitative experiments demonstrate that: 1) our model is capable of reaching top performance on both the single dataset and the multiple datasets with the same parameters in this study, which also won the 2nd place in the stereo task of the ECCV Robust vision Challenge 2020 (RVC 2020); and 2) on small datasets (e.g., KITTI, ETH3D, and Middlebury), the model's generalization and robustness are significantly affected by pre-training and data augmentation, even exceeding the network structure's influence in some cases. These observations present a challenge to the conventional wisdom of network architectures in this stage. We expect these discoveries to encourage researchers to rethink the current paradigm of "excessive attention on the performance of a single small dataset" in stereo matching. Zhibo Rao, Yuchao Dai, Zhelun Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | PCW-Net: Pyramid Combination and Warping Cost Volume for Stereo Matching
Zhelun Shen, Yuchao Dai, Xibin Song, Zhibo Rao, Dingfu Zhou, Liangjun Zhang |
ECCV (32) | 4 |
| 2022 | Sliding space-disparity transformer for stereo matching
Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen |
Neural Comput. Appl. | 1 |
| 2022 | Improving Stereo Matching Generalization via Fourier-Based Amplitude TransformabstractStereo matching CNNs suffer from performance deteriorate when evaluated under different distributions from training data. Previous domain adaptation/generalization methods are hard to maintain a robust performance in different baselines and usually require difficult adversarial optimization or intricate network structure. To solve this problem, we propose Fourier-based amplitude transform (FAT), mapping the source image to the target style without altering semantic content, which requires no training to perform the domain alignment. Specifically, we leverage the Fourier transform and its inverse to swap the low-frequency amplitude component of the source data with the target data. To effectively map style and relieve the artifacts, we introduce two factors to control the replacing area: the distance of HSV distribution between source and target images; and the difference between the source left image and its warped left image. Experiments testify FAT can significantly bridge domain gaps, making source data distribution closer to target data. Furthermore, when only training on synthetic datasets, FAT can also help different baselines achieve competitive cross-domain generalization capabilities on real datasets. Xing Li 0040, Yangyu Fan, Zhibo Rao, Guoyun Lv |
IEEE Signal Process. Lett. | 3 |
| 2022 | Synthetic-to-Real Domain Adaptation Joint Spatial Feature Transform for Stereo MatchingabstractMost deep learning-based state-of-the-art stereo matching methods significantly depend on large-scale datasets. However, it is implausible to collect sufficient real-world samples with dense and clear ground-truth disparity maps in practice. Although synthetic datasets’ appearance has alleviated the demand for extensive real data, there is a domain shift between synthetic and real sets. To tackle this problem, we propose an individually trained synthetic-to-real domain adaptation (SDA) network that maps synthetic images into the real domain. Specifically, our approach translates the data style from synthetic domain to real domain while maintaining the content and the spatial information. First, edge cues are leveraged to guide domain adaptation in preserving the spatial consistency between input and the generated image. Second, we combine the spatial feature transform (SFT) layer to effectively fuse features from the edge map and the source image. Extensive experiments demonstrate that: 1) when only trained on synthetic data and generalized to real data, our model evidently outperforms many state-of-the-art domain adaptation methods; 2) our translated synthetic datasets (TSD) help to improve the generalization capability of any stereo matching CNNs. Codes and data will be available athttps://github.com/Archaic-Atom/SDA_network. Xing Li 0040, Yangyu Fan, Zhibo Rao, Guoyun Lv, Shiya Liu |
IEEE Signal Process. Lett. | 3 |
| 2022 | Patch attention network with generative adversarial model for semi-supervised binocular disparity prediction
Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen |
Vis. Comput. | 1 |
| 2021 | CFNet: Cascade and Fused Cost Volume for Robust Stereo MatchingabstractRecently, the ever-increasing capacity of large-scale annotated datasets has led to profound progress in stereo matching. However, most of these successes are limited to a specific dataset and cannot generalize well to other datasets. The main difficulties lie in the large domain differences and unbalanced disparity distribution across a variety of datasets, which greatly limit the real-world applicability of current deep stereo matching models. In this paper, we propose CFNet, a Cascade and Fused cost volume based network to improve the robustness of the stereo matching network. First, we propose a fused cost volume representation to deal with the large domain difference. By fusing multiple low-resolution dense cost volumes to enlarge the receptive field, we can extract robust structural representations for initial disparity estimation. Second, we propose a cascade cost volume representation to alleviate the unbalanced disparity distribution. Specifically, we employ a variance-based uncertainty estimation to adaptively adjust the next stage disparity search space, in this way driving the network progressively prune out the space of unlikely correspondences. By iteratively narrowing down the disparity search space and improving the cost volume resolution, the disparity estimation is gradually refined in a coarse-to-fine manner. When trained on the same training images and evaluated on KITTI, ETH3D, and Middlebury datasets with the fixed model parameters and hyperparameters, our proposed method achieves the state-of-the-art overall performance and obtains the 1st place on the stereo task of Robust Vision Challenge 2020. The code will be available at https://github.com/gallenszl/CFNet. Zhelun Shen, Yuchao Dai, Zhibo Rao |
CVPR | 3 |
| 2021 | Him-Net: A New Neural Network Approach for SAR and Optical Image Template MatchingabstractSAR and optical images provide highly complementary information about observed scenes. The integrated use of these two data is desired in many data fusion tasks. However, traditional similarity methods cannot correctly match SAR and optical images due to the significant non-linear radio-metric difference between them. This paper proposed a template matching neural network based on stereo matching for SAR and optical image matching. Unlike the classical template matching methods doing feature extraction and similarity calculating separately, our network is a complete end-to-end approach, which allows optimizing the matching between the SAR and optical images through training. Moreover, a heatmap loss function is designed for image template matching, and better result is obtained. Our experiments confirmed our proposed network advantage over the state-of-the-art similarity approaches (such as NCC, CARMI, DeepMatch, and QATM) and superior matching performance. Mingyi He, Zhibo Rao, Wenyao Li 0002 |
ICIP | 3 |
| 2021 | Bidirectional Guided Attention Network for 3-D Semantic Detection of Remote Sensing ImagesabstractSemantic segmentation and disparity estimation are in the research frontier of the computer vision and remote sensing (RS) fields. However, existing methods mostly deal with these two problems separately or use a combination of multiple models to solve these two tasks. Due to a lack of sufficient information sharing and fusion, they still have difficulties in coping with seasonal appearance differences in 3-D RS problems. In this article, we propose a novel multitask learning architecture that considers the bottom–up and up–bottom visual attention mechanism for 3-D semantic detection, named bidirectional guided attention network (BGA-Net). BGA-Net consists of five modules: unified backbone module (UBM), bidirectional guided attention module (BGAM), semantic segmentation module (SSM), feature matching module (FMM), and bidirectional fusion module (BFM). First, in UBM, we use a shared backbone to extract unified features and share them with three branches/modules (BGAM, SSM, and FMM). Then, SSM and FMM branches are applied to estimate segmentation and disparity maps, whereas the third branch/module (BGAM) shares the global features to guide the task-specific learning via attention mechanism. Finally, we fuse the results of the two tasks by BFM to improve the final performance. Extensive experiments demonstrate that: 1) our BGA-Net can handle the two tasks simultaneously and can be trained in an end-to-end way; 2) these modules fully take advantage of the two tasks’ information to share features and enhance the scene understanding ability, effectively against seasons change of RS images; and 3) BGA-Net has notable superiority and greater flexibility and also sets a new state of the art on the urban semantic 3-D (US3D) benchmark. Moreover, BGA-Net also provides insights into the intelligent interpretation of RS data images. Zhibo Rao, Mingyi He, Zhidong Zhu, Yuchao Dai |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | MVS2: Deep Unsupervised Multi-View Stereo with Multi-View SymmetryabstractThe success of existing deep-learning based multi-view stereo (MVS) approaches greatly depends on the availability of large-scale supervision in the form of dense depth maps. Such supervision, while not always possible, tends to hinder the generalization ability of the learned models in never-seen-before scenarios. In this paper, we propose the first unsupervised learning based MVS network, which learns the multi-view depth maps from the input multi-view images and does not need ground-truth 3D training data. Our network is symmetric in predicting depth maps for all views simultaneously, where we enforce cross-view consistency of multi-view depth maps during both training and testing stages. Thus, the learned multi-view depth maps naturally comply with the underlying 3D scene geometry. Besides, our network also learns the multi-view occlusion maps, which further improves the robustness of our network in handling real-world occlusions. Experimental results on multiple benchmarking datasets demonstrate the effectiveness of our network and the excellent generalization ability. Yuchao Dai, Zhidong Zhu, Zhibo Rao, Bo Li 0090 |
3DV | 3 |
| 2019 | Input-Perturbation-Sensitivity for Performance Analysis of CNNS on Image RecognitionabstractPerformance assessment is critical to learning systems, but it is tough to explain the relationship between data, model, and performance. In this paper, the Input-Perturbation-Sensitivity (IPS) is proposed to investigate this problem in a class of Convolutional Neural Networks (CNNs) for image recognition and try to explain their relationships. First, IPS is defined and the CNNs parameters are divided into groups according to the layers of the model. Second, the output perturbations of the CNNs caused by input perturbations are analyzed with a group of local sensitivities (LS). Third, global sensitivity (GS) is obtained over all local IPS. Finally, experiments are carried out on a few CNNs with different hyper-parameters on the image recognition datasets. The analytic and experimental results show that the proposed method correlates well with data, model, and performance. Moreover, the IPS can provide a reasonable explanation of the different networks for image classification, showing the potential to evaluate other learning systems. Zhibo Rao, Mingyi He, Zhidong Zhu |
ICIP | 1 |