Yijun Cao

dblp:195/3749 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Personalized federated learning via multifaceted feature matching and element-wise classifier fusion
Yijun Cao, Hongjiao Li, Botao Zhang 0006, Ning Xue, Hongliang Yin
Multim. Syst.1
2025 Interpretable localization of false data injection attacks in smart grids: a multi-head graph convolutional attention network approach
Hongjiao Li, Botao Zhang 0006, Ning Xue, Hongliang Yin, Yijun Cao
J. Supercomput.6
2024 LVP-net: A deep network of learning visual pathway for edge detection
Chuan Lin 0003, Fuzhang Li, Yijun Cao, Yongjie Li 0001
Image Vis. Comput.4
2024 Nighttime Thermal Infrared Image Colorization With Feedback-Based Object Appearance Learning
abstract
Stable imaging in adverse environments (e.g., total darkness) makes thermal infrared (TIR) cameras a prevalent option for night scene perception. However, the low contrast and lack of chromaticity of TIR images are detrimental to human interpretation and subsequent deployment of RGB-based vision algorithms. Therefore, it makes sense to colorize the nighttime TIR images by translating them into the corresponding daytime color images (NTIR2DC). Despite the impressive progress made in the NTIR2DC task, how to improve the translation performance of small object classes is under-explored. To address this problem, we propose a generative adversarial network incorporating feedback-based object appearance learning (FoalGAN). Specifically, an occlusion-aware mixup module and corresponding appearance consistency loss are proposed to reduce the context dependence of object translation. As a representative example of small objects in nighttime street scenes, we illustrate how to enhance the realism of traffic light by designing a traffic light appearance loss. To further improve the appearance learning of small objects, we devise a dual feedback learning strategy to selectively adjust the learning frequency of different samples. In addition, we provide pixel-level annotation for a subset of the Brno dataset, which can facilitate the research of NTIR image understanding under multiple weather conditions. Extensive experiments illustrate that the proposed FoalGAN is not only effective for appearance learning of small objects, but also outperforms other image translation methods in terms of semantic preservation and edge consistency for the NTIR2DC task. Compared with the state-of-the-art NTIR2DC approach, FoalGAN achieves at least 5.4% improvement in semantic consistency and at least 2% lead in edge consistency.
Fuya Luo, Yijun Cao, Kaifu Yang, Chang-Yong Xie, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Memory-Guided Collaborative Attention for Nighttime Thermal Infrared Image Colorization of Traffic Scenes
abstract
Robust imaging under challenging conditions, such as starlit nights, has broadened the adoption of thermal infrared (TIR) cameras for nighttime driving scenes. Given that TIR images are monochromatic, which makes them difficult to interpret by humans and limits the applicability of RGB-based algorithms, it is reasonable to perform colorization of nighttime TIR (NTIR) images by converting them into corresponding daytime color images (NTIR2DC). Despite the impressive results achieved by previous NTIR2DC methods, how to improve the colorization performance of small-sample categories without semantic annotation is under-explored. To address this issue, we propose a novel learning framework called Memory-guided cOllaboRative atteNtion Generative Adversarial Network (MornGAN), which is inspired by the analogical reasoning mechanisms of humans. Specifically, we first propose an online semantic distillation module to mine and refine the semantic cues of NTIR images. Then, a memory-guided sample selection strategy and adaptive collaborative attention loss are devised to enhance the semantic preservation of small-sample categories. Further, a new conditional gradient repair loss is introduced for reducing edge distortion during translation. Extensive experiments on the NTIR2DC task show that the proposed MornGAN significantly outperforms other image-to-image translation methods in terms of semantic preservation and edge consistency, which helps improve the object detection accuracy remarkably.
Fuya Luo, Yijun Cao, Kaifu Yang, Gang Wang 0031, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.2
2023 Toward Better SSIM Loss for Unsupervised Monocular Depth Estimation
Yijun Cao, Fuya Luo, Yongjie Li 0001
ICIG (1)1
2023 Learning generalized visual odometry using position-aware optical flow and geometric bundle adjustment
abstract
Recent visual odometry (VO) methods incorporating geometric algorithm into deep-learning architecture have shown outstanding performance on the challenging monocular VO task. Despite encouraging results are shown, previous methods ignore the requirement of generalization capability under noisy environment and various scenes. To address this challenging issue, this work first proposes a novel optical flow network (PANet). Compared with previous methods that predict optical flow as a direct regression task , our PANet computes optical flow by predicting it into the discrete position space with optical flow probability volume, and then converting it to optical flow. Next, we improve the bundle adjustment module to fit the self-supervised training pipeline by introducing multiple sampling, ego-motion initialization, dynamic damping factor adjustment, and Jacobi matrix weighting. In addition, a novel normalized photometric loss function is advanced to improve the depth estimation accuracy. The experiments show that the proposed system not only achieves comparable performance with other state-of-the-art self-supervised learning-based methods on the KITTI dataset, but also significantly improves the generalization capability compared with geometry-based, learning-based and hybrid VO systems on the noisy KITTI and the challenging outdoor (KAIST) scenes.
Yijun Cao, Fuya Luo, Chuan Lin 0003, Kaifu Yang, Yongjie Li 0001
Pattern Recognit.1
2023 Unsupervised Visual Odometry and Action Integration for PointGoal Navigation in Indoor Environment
abstract
PointGoal navigation in indoor environment is a fundamental task for personal robots to navigate to a specified point. Recent studies solved this PointGoal navigation task with near-perfect success rate in photo-realistically simulated environments, under the assumptions with noiseless actuation and most importantly, perfect localization with GPS and compass sensors. However, accurate GPS signalis difficult to be obtained in real indoor environment. To improve the PointGoal navigation accuracy without GPS signal, we use visual odometry (VO) and propose a novel action integration module (AIM) trained in unsupervised manner. Sepecifically, unsupervised VO computes the relative pose of the agent from the re-projection error of two adjacent frames, and then replaces the accurate GPS signal with the path integration. The pseudo position estimated by VO is used to train action integration which assists agent to update their internal perception of location and helps improve the success rate of navigation. The training and inference process only use RGB, depth, collision as well as self-action information. The experiments show that the proposed system achieves satisfactory results and outperforms the partially supervised learning algorithms on the popular Gibson dataset.
Yijun Cao, Fuya Luo, Chuan Lin 0003, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 A bio-inspired contour detection model using multiple cues inhibition in primary visual cortex
Chuan Lin 0003, Ze-Qi Wen, Gui-Li Xu, Yijun Cao, Yongcai Pan 0001
Multim. Tools Appl.4
2021 Nighttime Thermal Infrared Image Colorization with Dynamic Label Mining
Fuya Luo, Yijun Cao, Yongjie Li 0001
ICIG (3)2
2021 Learning Crisp Boundaries Using Deep Refinement Network and Adaptive Weighting Loss
abstract
Significant progress has been made in boundary detection with the help of convolutional neural networks. Recent boundary detection models not only focus on real object boundary detection but also “crisp” boundaries (precisely localized along the object's contour). There are two methods to evaluate crisp boundary performance. One uses more strict tolerance to measure the distance between the ground truth and the detected contour. The other focuses on evaluating the contour map without any postprocessing. In this study, we analyze both methods and conclude that both methods are two aspects of crisp contour evaluation. Accordingly, we propose a novel network named deep refinement network (DRNet) that stacks multiple refinement modules to achieve richer feature representation and a novel loss function, which combines cross-entropy and dice loss through effective adaptive fusion. Experimental results demonstrated that we achieve state-of-the-art performance for several available datasets.
Yijun Cao, Chuan Lin 0003, Yongjie Li 0001
IEEE Trans. Multim.1
2020 Lateral refinement network for contour detection
Chuan Lin 0003, Linhao Cui, Fuzhang Li, Yijun Cao
Neurocomputing4
2019 Bio-inspired contour detection model based on multi-bandwidth fusion and logarithmic texture inhibition
abstract
Relevant physiological studies have revealed that the response of the classical receptive field (CRF) to visual stimuli could be suppressed by non‐CRF (nCRF) inhibition of the kernel in the primary visual cortex (V1). Based on this mechanism, many bio‐inspired contour detection models have been proposed, which are mainly achieved through CRF responses and nCRF surround inhibition calculation. In fact, the dynamic characteristics of neurons play an important role in contour detection in biological vision. Inspired by these visual mechanisms, the authors propose a contour detection model that emulates these dynamic characteristics. By introducing a multi‐bandwidth Gabor filter, according to the target image, they can effectively adjust the weight ratios of the filter to protect the contours and filter the background textures in the calculation of CRF responses. Additionally, they logarithmically modulate the nCRF inhibition kernel to make texture suppression more flexible and effective, thus improving the accuracy of detection algorithm as a whole. Compared with existing bio‐inspired contour detection models, the proposed model is more effective at contour detection, which will aid engineering applications that utilise pattern recognition in machine vision.
Chuan Lin 0003, Fuzhang Li, Yijun Cao, Haojun Zhao
IET Image Process.3
2019 Application of the center-surround mechanism to contour detection
Yijun Cao, Chuan Lin 0003, Yi-Jian Pan, Hao-Jun Zhao
Multim. Tools Appl.1
2018 Contour detection model based on neuron behaviour in primary visual cortex
abstract
In the mammalian primary visual cortex, the response of the classical receptive field (CRF) to visual stimuli can be suppressed by inhibition of non‐CRF (nCRF) neurons. Although many biologically plausible models based on these centre–surround interaction properties have been proposed, most of these models have failed to account for two important behaviours of neurons in the primary visual cortex (V1). First, saturation properties of neuron response. Second, the properties of fixational eye movements (FEyeMs). In the present study, the authors proposed a biologically motivated counter detection approach based on these properties. The authors’ work is significant in that they utilised a simple threshold method to ensure that CRF responses were observed within a meaningful range, and multichannel filter bank was proposed to simulate the influence of FEyeMs on nCRF. Both methods effectively preserved object contours and inhibition isolated textures. Extensive experiments indicated that the authors’ model can preserve more object contours and suppress more textures than previous biologically based models.
Chuan Lin 0003, Guili Xu, Yijun Cao
IET Comput. Vis.3
2018 Contour detection model using linear and non-linear modulation based on non-CRF suppression
abstract
Psychophysical and neurophysiological investigations on the human visual system show that most neurons in the primary visual cortex (V1) possess a non‐classical receptive field (nCRF) region in addition to the CRF region. The nCRF has a modulatory, normally inhibitory, effect on the responses to visual stimuli generated within the CRF. In computational terms, this mechanism suppresses the response to edges in the presence of similar edges in the surroundings. Many computational techniques have been proposed to address the surround suppression mechanism. These methods introduce an inhibition term that is required to suppress the textures and protect the contours. Several studies have found that the spatial summation properties over the receptive fields of retinal X cells are approximately linear, while they are non‐linear for Y cells. Inspired by the visual information processing in the X–Y channel and spatial summation properties of X and Y cells, the authors propose a contour detector using linear and non‐linear modulations based on nCRF suppression. Extensive experimental evaluations demonstrate that their contour detector significantly outperforms other algorithms. The methods proposed in this study are expected to facilitate the development of efficient computational models in the field of machine vision.
Chuan Lin 0003, Guili Xu, Yijun Cao
IET Image Process.3
2017 Optimizing ZNCC calculation in binocular stereo matching
Chuan Lin 0003, Guili Xu, Yijun Cao
Signal Process. Image Commun.4