Chaoqiang Zhao

dblp:255/5327 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-3651-2177ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Boundary-Based Active Domain Adaptation for Semantic Segmentation Under Adverse Conditions
abstract
Existing domain adaptation semantic segmentation (DASS) methods under adverse conditions often depend on pseudo-labels for network training. However, these pseudo-labels are frequently plagued by noise and bias toward high-confidence predictions, thereby impeding the enhancement of segmentation performance. This article tackles the above challenge by proposing a novel boundary-based active domain adaptation (ADA) framework, which efficiently selects both informative low-confidence samples and high-confident but misclassified samples to be labeled while maximizing the segmentation performance under a limited annotation budget. For the evaluation of sample confidence and informativeness, we first propose ranking weighted feature space impurity (RWFSI) metric to quantify category distribution among a sample's nearest neighbors within the feature space and consider the samples with higher RWFSI values as low-confidence samples around the decision boundary, which can also alleviate the category imbalance of active labels. Subsequently, we apply Gaussian mixture models (GMMs) to model the distribution across source and target domains. Using the spatial arrangement of each GMM component, we define the intraclass domain shift score (ICDSS), which identifies samples with high ICDSS values as those more likely to be high-confidence but misclassified, aiding in refining sample selection. Extensive experiments demonstrate that our method is superior to the existing state-of-the-art domain adaptation and active learning (AL) methods and comparable with those of full supervision. The code will be released at https://github.com/1061018609/BADA.
Gary G. Yen, Chaoqiang Zhao, Qiyu Sun, Wenqi Ren, Lu Sheng, Yang Tang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Self-Supervised Monocular Depth Estimation in the Dark: Towards Data Distribution Compensation
Chaoqiang Zhao, Lu Sheng, Yang Tang 0001
IJCAI2
2024 Learn to Adapt for Self-Supervised Monocular Depth Estimation
abstract
Monocular depth estimation is one of the fundamental tasks in environmental perception and has achieved tremendous progress by virtue of deep learning. However, the performance of trained models tends to degrade or deteriorate when employed on other new datasets due to the gap between different datasets. Though some methods utilize domain adaptation technologies to jointly train different domains and narrow the gap between them, the trained models cannot generalize to new domains that are not involved in training. To boost the transferability of self-supervised monocular depth estimation models and mitigate the issue of meta-overfitting, we train the model in the pipeline of meta-learning and propose an adversarial depth estimation task. We adopt model-agnostic meta-learning (MAML) to obtain universal initial parameters for further adaptation and train the network in an adversarial manner to extract domain-invariant representations for easing meta-overfitting. In addition, we propose a constraint to impose upon cross-task depth consistency to compel the depth estimation to be identical in different adversarial tasks, which improves the performance of our method and smoothens the training process. Experiments on four new datasets demonstrate that our method adapts quite fast to new domains. Our method trained after 0.5 epoch achieves comparable results with the state-of-the-art methods trained at least 20 epochs.
Qiyu Sun, Gary G. Yen, Yang Tang 0001, Chaoqiang Zhao
IEEE Trans. Neural Networks Learn. Syst.4
2023 CMDA: Cross-Modality Domain Adaptation for Nighttime Semantic Segmentation
abstract
Most nighttime semantic segmentation studies are based on domain adaptation approaches and image input. However, limited by the low dynamic range of conventional cameras, images fail to capture structural details and boundary information in low-light conditions. Event cameras, as a new form of vision sensors, are complementary to conventional cameras with their high dynamic range. To this end, we propose a novel unsupervised Cross-Modality Domain Adaptation (CMDA) framework to leverage multi-modality (Images and Events) information for nighttime semantic segmentation, with only labels on daytime images. In CMDA, we design the Image Motion-Extractor to extract motion information and the Image Content-Extractor to extract content information from images, in order to bridge the gap between different modalities (Images ⇌ Events) and domains (Day ⇌ Night). Besides, we introduce the first image-event nighttime semantic segmentation dataset. Extensive experiments on both the public image dataset and the proposed image-event dataset demonstrate the effectiveness of our proposed approach. We open-source our code, models, and dataset at https://github.com/XiaRho/CMDA.
Ruihao Xia, Chaoqiang Zhao, Meng Zheng 0002, Ziyan Wu 0001, Qiyu Sun, Yang Tang 0001
ICCV2
2023 GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor Scenes
abstract
This paper tackles the challenges of self-supervised monocular depth estimation in indoor scenes caused by large rotation between frames and low texture. We ease the learning process by obtaining coarse camera poses from monocular sequences through multi-view geometry to deal with the former. However, we found that limited by the scale ambiguity across different scenes in the training dataset, a naïve introduction of geometric coarse poses cannot play a positive role in performance improvement, which is counter-intuitive. To address this problem, we propose to refine those poses during training through rotation and translation/scale optimization. To soften the effect of the low texture, we combine the global reasoning of vision transformers with an overfitting-aware, iterative self-distillation mechanism, providing more accurate depth guidance coming from the network itself. Experiments on NYUv2, ScanNet, 7scenes, and KITTI datasets support the effectiveness of each component in our framework, which sets a new state-of-the-art for indoor self-supervised monocular depth estimation, as well as outstanding generalization ability. Code and models are available at https://github.com/zxcqlf/GasMono
Chaoqiang Zhao, Matteo Poggi, Fabio Tosi, Qiyu Sun, Yang Tang 0001, Stefano Mattoccia
ICCV1
2023 Rethinking Unsupervised Domain Adaptation for Nighttime Tracking
Qiyu Sun, Chaoqiang Zhao, Wenqi Ren, Yang Tang 0001
ICONIP (14)3
2023 Multi-Dimensional Deformable Object Manipulation Using Equivariant Models
abstract
Manipulating deformable objects, such as ropes (1D), fabrics (2D), and bags (3D), poses a significant challenge in robotics research due to their high degree of freedom in physical state and nonlinear dynamics. Compared with single-dimensional deformable objects, multi-dimensional object manipulation suffers from the difficulty in recognizing the characteristics of the object correctly and making an accurate action decision on the deformable object of various dimensions. Some methods are proposed to use neural networks to rearrange deformable objects in all dimensions, but their approaches are not accurate in predicting the motion of the robot as they just consider the equivariance in the picking objects. To address this problem, we present a novel Transporter Network encoded and decoded with equivariance to generalize to different picking and placing positions. Additionally, we propose an equivariant goal-conditioned model to enable the robot to manipulate deformable objects into flexible configurations without relying on artificially marked visual anchors for the target position. Finally, experiments conducted in both Deformable-Ravens and the real world demonstrate that our equivariant models are more sample efficient than the traditional Transporter Network. The video is available at https://youtu.be/5_q5ff9c9FU.
Tianyu Fu 0007, Yang Tang 0001, Xiaowu Xia, Jianrui Wang, Chaoqiang Zhao
IROS6
2023 Self-supervised depth super-resolution with contrastive multiview pre-training
Chenyang Ge, Chaoqiang Zhao, Fabio Tosi, Matteo Poggi, Stefano Mattoccia
Neural Networks3
2023 Perception and Navigation in Autonomous Systems in the Era of Learning: A Survey
abstract
Autonomous systems possess the features of inferring their own state, understanding their surroundings, and performing autonomous navigation. With the applications of learning systems, like deep learning and reinforcement learning, the visual-based self-state estimation, environment perception, and navigation capabilities of autonomous systems have been efficiently addressed, and many new learning-based algorithms have surfaced with respect to autonomous visual perception and navigation. In this review, we focus on the applications of learning-based monocular approaches in ego-motion perception, environment perception, and navigation in autonomous systems, which is different from previous reviews that discussed traditional methods. First, we delineate the shortcomings of existing classical visual simultaneous localization and mapping (vSLAM) solutions, which demonstrate the necessity to integrate deep learning techniques. Second, we review the visual-based environmental perception and understanding methods based on deep learning, including deep learning-based monocular depth estimation, monocular ego-motion prediction, image enhancement, object detection, semantic segmentation, and their combinations with traditional vSLAM frameworks. Then, we focus on the visual navigation based on learning systems, mainly including reinforcement learning and deep reinforcement learning. Finally, we examine several challenges and promising directions discussed and concluded in related research of learning systems in the era of computer science and robotics.
Yang Tang 0001, Chaoqiang Zhao, Jianrui Wang, Chongzhen Zhang, Qiyu Sun, Wei Xing Zheng 0001, Wenli Du, Feng Qian 0004, Jürgen Kurths
IEEE Trans. Neural Networks Learn. Syst.2
2022 MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer
abstract
Self-supervised monocular depth estimation is an attractive solution that does not require hard-to-source depth la-bels for training. Convolutional neural networks (CNNs) have recently achieved great success in this task. However, their limited receptive field constrains existing network architectures to reason only locally, dampening the effectiveness of the self-supervised paradigm. In the light of the recent successes achieved by Vision Transformers (ViTs), we propose MonoViT, a brand-new framework combining the global reasoning enabled by ViT models with the flexibility of self-supervised monocular depth estimation. By combining plain convolutions with Transformer blocks, our model can reason locally and globally, yielding depth prediction at a higher level of detail and accuracy, allowing MonoViT to achieve state-of-the-art performance on the established KITTI dataset. Moreover, MonoViT proves its superior generalization capacities on other datasets such as Make3D and DrivingStereo. Source code available at https://github.com/zxcqlf/MonoViT
Chaoqiang Zhao, Youmin Zhang 0008, Matteo Poggi, Fabio Tosi, Xianda Guo, Guan Huang 0003, Yang Tang 0001, Stefano Mattoccia
3DV1
2022 Deep Direct Visual Odometry
abstract
Traditional monocular direct visual odometry (DVO) is one of the most famous methods to estimate the ego-motion of robots and map environments from images simultaneously. However, DVO heavily relies on high-quality images and accurate initial pose estimation during tracking. With the outstanding performance of deep learning, previous works have shown that deep neural networks can effectively learn 6-DoF (Degree of Freedom) poses between frames from monocular image sequences in the unsupervised manner. However, these unsupervised deep learning-based frameworks cannot accurately generate the full trajectory of a long monocular video because of the scale-inconsistency between each pose. To address this problem, we use several geometric constraints to improve the scale-consistency of the pose network, including improving the previous loss function and proposing a novel scale-to-trajectory constraint for unsupervised training. We call the pose network trained by the proposed novel constraint as TrajNet. In addition, a new DVO architecture, called deep direct sparse odometry (DDSO), is proposed to overcome the drawbacks of the previous direct sparse odometry (DSO) framework by embedding TrajNet. Extensive experiments on the KITTI dataset show that the proposed constraints can effectively improve the scale-consistency of TrajNet when compared with previous unsupervised monocular methods, and integration with TrajNet makes the initialization and tracking of DSO more robust and accurate.
Chaoqiang Zhao, Yang Tang 0001, Qiyu Sun, Athanasios V. Vasilakos
IEEE Trans. Intell. Transp. Syst.1
2022 Unsupervised Estimation of Monocular Depth and VO in Dynamic Environments via Hybrid Masks
abstract
Deep learning-based methods mymargin have achieved remarkable performance in 3-D sensing since they perceive environments in a biologically inspired manner. Nevertheless, the existing approaches trained by monocular sequences are still prone to fail in dynamic environments. In this work, we mitigate the negative influence of dynamic environments on the joint estimation of depth and visual odometry (VO) through hybrid masks. Since both the VO estimation and view reconstruction process in the joint estimation framework is vulnerable to dynamic environments, we propose the cover mask and the filter mask to alleviate the adverse effects, respectively. As the depth and VO estimation are tightly coupled during training, the improved VO estimation promotes depth estimation as well. Besides, a depth-pose consistency loss is proposed to overcome the scale inconsistency between different training samples of monocular sequences. Experimental results show that both our depth prediction and globally consistent VO estimation are state of the art when evaluated on the KITTI benchmark. We evaluate our depth prediction model on the Make3D dataset to prove the transferability of our method as well.
Qiyu Sun, Yang Tang 0001, Chongzhen Zhang, Chaoqiang Zhao, Feng Qian 0004, Jürgen Kurths
IEEE Trans. Neural Networks Learn. Syst.4
2021 Multitask GANs for Semantic Segmentation and Depth Completion With Cycle Consistency
abstract
Semantic segmentation and depth completion are two challenging tasks in scene understanding, and they are widely used in robotics and autonomous driving. Although several studies have been proposed to jointly train these two tasks using some small modifications, such as changing the last layer, the result of one task is not utilized to improve the performance of the other one despite that there are some similarities between these two tasks. In this article, we propose multitask generative adversarial networks (Multitask GANs), which are not only competent in semantic segmentation and depth completion but also improve the accuracy of depth completion through generated semantic images. In addition, we improve the details of generated semantic images based on CycleGAN by introducing multiscale spatial pooling blocks and the structural similarity reconstruction loss. Furthermore, considering the inner consistency between semantic and geometric structures, we develop a semantic-guided smoothness loss to improve depth completion results. Extensive experiments on the Cityscapes data set and the KITTI depth completion benchmark show that the Multitask GANs are capable of achieving competitive performance for both semantic segmentation and depth completion tasks.
Chongzhen Zhang, Yang Tang 0001, Chaoqiang Zhao, Qiyu Sun, Zhencheng Ye, Jürgen Kurths
IEEE Trans. Neural Networks Learn. Syst.3
2021 Masked GAN for Unsupervised Depth and Pose Prediction With Scale Consistency
abstract
Previous work has shown that adversarial learning can be used for unsupervised monocular depth and visual odometry (VO) estimation, in which the adversarial loss and the geometric image reconstruction loss are utilized as the mainly supervisory signals to train the whole unsupervised framework. However, the performance of the adversarial framework and image reconstruction is usually limited by occlusions and the visual field changes between the frames. This article proposes a masked generative adversarial network (GAN) for unsupervised monocular depth and ego-motion estimations. The MaskNet and Boolean mask scheme are designed in this framework to eliminate the effects of occlusions and impacts of visual field changes on the reconstruction loss and adversarial loss, respectively. Furthermore, we also consider the scale consistency of our pose network by utilizing a new scale-consistency loss, and therefore, our pose network is capable of providing the full camera trajectory over a long monocular sequence. Extensive experiments on the KITTI data set show that each component proposed in this article contributes to the performance, and both our depth and trajectory predictions achieve competitive performance on the KITTI and Make3D data sets.
Chaoqiang Zhao, Gary G. Yen, Qiyu Sun, Chongzhen Zhang, Yang Tang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Trajectory Planning for Unmanned Aircraft Vehicle via Set-Valued Filter
abstract
Trajectory planning in complex environments with different kinds of obstacles, like static obstacles, dynamic obstacles and noncooperative agents, is of great pratical importance for Unmanned Aircraft Vehicle (UAV). Although various algorithms are proposed to solve the obstacle avoidance problem, these methods only consider one or two kinds of obstacles. In this paper, a novel algorithm is proposed to plan a trajectory for UAV in such an environment which simultaneously includes static obstacles, dynamic obstacles and especially noncooperative agents. Besides, we consider different types of approach modes of noncooperative agents and propose corresponding strategies to avoid collisions with them. A set-valued filter-based method is proposed to predict the position of dynamic obstacles and noncooperative agents whose motion model is not available to UAV. Meanwhile, the measurement noise is considered in the set-valued filter to improve the safety of UAV. Some simulations are implemented to verify the effectiveness of the proposed method and they confirm that the proposed method can plan a safe trajectory for UAV in complex environments.
Hailong Qian, Weimin Zhong, Chaoqiang Zhao, Wenle Zhang, Yang Tang 0001
IECON3