Liguang Zhou

dblp:216/8298 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-0237-1377ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Peer Learning Approach to Unbiased Scene Graph Generation for Traffic Scene Understanding
abstract
The biased scene graph generation problem arises from the inherent long-tailed distributions of predicates, which are challenging to handle effectively with a single network. In this paper, we introduce a novel framework called peer learning, designed to address the issue of unbiased scene graph generation (USGG) through a divide-and-vote approach. To address the long-tailed problem, our framework operates in three steps. Firstly, we partition the heavily long-tailed distribution into subsets of more balanced sub-distribution groups, including head, body, and tail classes with a predicate sampling module. Next, we establish a peer network consisting of multiple peers, where each peer receives a combination of sub-distributions. This division enables peers to focus on different aspects of the scene graph generation task. Then, a novel peer learning loss function is introduced to cultivate the learning process among peer networks. Lastly, we employ the voting strategies for making final predictions within the peer network, boosting the influence of the majority’s opinion while downplaying the minority’s perspective. To illustrate the applicability of the proposed framework in intelligent transportation systems (ITSs), we further conduct qualitative evaluations on traffic scene understanding tasks. The results demonstrate that peer learning markedly enhances the reliability of interpreting complex traffic scenarios. Experimental results on the Visual Genome and Open Images V6 datasets further verify the effectiveness of our proposed model. These results highlight that the peer learning framework is well-suited for addressing the challenges of unbiased scene graph generation, offering practical benefits for ITS applications such as traffic analysis and monitoring. The code is available at: PL.
Liguang Zhou, Junjie Hu 0003, Yuhongze Zhou, Tin Lun Lam, Yangsheng Xu
IEEE Trans. Intell. Transp. Syst.1
2025 Class Relevance Learning for Out-of-Distribution Detection
abstract
Image classification plays a pivotal role across diverse robotic applications, yet challenges persist when models are deployed in real-world scenarios. These models often fail to detect out-of-distribution (OOD) samples, classes not included in their training. This makes OOD detection a significant challenge for safe and effective real-world use. While existing techniques, like max logits, aim to leverage logits for OOD identification, they often disregard the intricate interclass relationships that underlie effective detection. This paper presents an innovative class relevance learning (CRL) method tailored for OOD detection. Our method establishes a comprehensive class relevance learning framework, strategically harnessing interclass relationships within the OOD pipeline. This framework significantly augments OOD detection capabilities. Extensive experimentation on diverse datasets, encompassing generic image classification datasets (Near OOD and Far OOD datasets), demonstrates the superiority of our method over state-of-the-art alternatives for OOD detection. The code is available on GitHub at: CRL.
Liguang Zhou, Butian Xiong, Tin Lun Lam, Yangsheng Xu
ICASSP1
2025 Lifelong-MonoDepth: Lifelong Learning for Multidomain Monocular Metric Depth Estimation
abstract
With the rapid advancements in autonomous driving and robot navigation, there is a growing demand for lifelong learning (LL) models capable of estimating metric (absolute) depth. LL approaches potentially offer significant cost savings in terms of model training, data storage, and collection. However, the quality of RGB images and depth maps is sensor-dependent, and depth maps in the real world exhibit domain-specific characteristics, leading to variations in depth ranges. These challenges limit existing methods to LL scenarios with small domain gaps and relative depth map estimation. To facilitate lifelong metric depth learning, we identify three crucial technical challenges that require attention: 1) developing a model capable of addressing the depth scale variation through scale-aware depth learning; 2) devising an effective learning strategy to handle significant domain gaps; and 3) creating an automated solution for domain-aware depth inference in practical applications. Based on the aforementioned considerations, in this article, we present 1) a lightweight multihead framework that effectively tackles the depth scale imbalance; 2) an uncertainty-aware LL solution that adeptly handles significant domain gaps; and 3) an online domain-specific predictor selection method for real-time inference. Through extensive numerical studies, we show that the proposed method can achieve good efficiency, stability, and plasticity, leading the benchmarks by 8%-15%. The code is available at https://github.com/FreeformRobotics/Lifelong-MonoDepth.
Junjie Hu 0003, Chenyou Fan, Liguang Zhou, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam
IEEE Trans. Neural Networks Learn. Syst.3
2023 Affinity Learning With Blind-Spot Self-Supervision for Image Denoising
abstract
In this paper, we extend the blind-spot based self-supervised denoising by using affinity learning to remove noise from affected pixels. Inspired by inpainting, we introduce a novel Mask Guided Residual Convolution (MGRConv) to learn a neighboring image pixel affinity map that gradually removes noise and refines blind-spot denoising process. We show that mask convolution plays an important role in blind-spot denoising since it is theoretically aligned with $\mathcal{J} - invariance$, which blind-spot based self-supervised denoising frameworks are built upon. The theoretical analysis further shows the motivation behind using more adaptive mask convolutions. Our MGRConv not only enables dynamic mask learning without external trainable parameters, but also preserves appropriate mask constraints by sigmoid activation and residual summation. Our MGRConv is a balance between partial convolution and learnable attention maps, and boosts denoising performance better than other inpainting convolutions with similar or even less parameters, memory, and training/inference time. Extensive experiments show that our proposed plug-and-play MGRConv can assist blind-spot based denoising networks to reach promising results on both existing single-image based and dataset based benchmarks.
Yuhongze Zhou, Liguang Zhou, Issam H. Laradji, Tin Lun Lam, Yangsheng Xu
ICASSP2
2023 Sampling Propagation Attention With Trimap Generation Network for Natural Image Matting
abstract
Natural image matting aims to precisely separate foreground objects from backgrounds using alpha mattes. Fully automatic natural image matting without external annotations is challenging. Well-performed matting methods usually require accurate labor-intensive handcrafted trimap as an extra input while the performance of automatic trimap generation method, e.g., erosion/dilation manipulation on foreground segmentation, fluctuates with segmentation quality. Therefore, we argue that how to produce a high-quality trimap using coarse segmentation is a major issue in automatic matting. In this paper, we present a two-stage trimap-free natural image matting pipeline that does not need trimap and background as input. Specifically, guided by a coarse segmentation, Trimap Generation Network (TGN) estimates a trimap where the coarse segmentation can be produced by segmentation/salient object detection/matting approaches, which enables more flexibility for matting to adapt into different scenarios. Then, with an estimated trimap as guidance, our Sampling Propagation Attention Matting Network (SPAMattNet) estimates an alpha matte. Different from previous propagation-based matting networks, inspired by traditional sampling/propagation matting approaches, we propose Sampling Propagation Attention (SPA) for matting network to incorporate sampling and propagation procedures in deep learning based manner for network explainability and performance improvement. It explicitly investigates local spatial and global semantic relationships to reconstruct alpha features. To better harvest sampling/propagation and local/global information, a Cross-Fusion Contextual Module (CFC) is introduced to aggregate features from different sources. Extensive experiments are conducted to show that our matting approach is competitive compared to other state-of-the-art methods in both trimap-free and trimap-needed aspects on several challenging matting benchmarks.
Yuhongze Zhou, Liguang Zhou, Tin Lun Lam, Yangsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2022 OSM: An Open Set Matting Framework with OOD Detection and Few-Shot Learning
Yuhongze Zhou, Issam H. Laradji, Liguang Zhou, Derek Nowrouzezahrai
BMVC3
2022 Toward Better Accuracy-Efficiency Trade-Offs: Divide and Co-Training
abstract
The width of a neural network matters since increasing the width will necessarily increase the model capacity. However, the performance of a network does not improve linearly with the width and soon gets saturated. In this case, we argue that increasing the number of networks (ensemble) can achieve better accuracy-efficiency trade-offs than purely increasing the width. To prove it, one large network is divided into several small ones regarding its parameters and regularization components. Each of these small networks has a fraction of the original one's parameters. We then train these small networks together and make them see various views of the same data to increase their diversity. During this co-training process, networks can also learn from each other. As a result, small networks can achieve better ensemble performance than the large one with few or no extra parameters or FLOPs, i. e., achieving better accuracy-efficiency trade-offs. Small networks can also achieve faster inference speed than the large one by concurrent running. All of the above shows that the number of networks is a new dimension of model scaling. We validate our argument with 8 different neural architectures on common benchmarks through extensive experiments.
Shuai Zhao 0006, Liguang Zhou, Wenxiao Wang 0001, Deng Cai 0001, Tin Lun Lam, Yangsheng Xu
IEEE Trans. Image Process.2
2021 Long-Range Hand Gesture Recognition via Attention-based SSD Network
abstract
Hand gesture recognition plays an essential role in the human-robot interaction (HRI) field. Most previous research only studies hand gesture recognition in a short distance, which cannot be applied for interaction with mobile robots like unmanned aerial vehicles (UAVs) at a longer and safer distance. Therefore, we investigate the challenging long-range hand gesture recognition problem for the interaction between humans and UAVs. To this end, we propose a novel attention-based single shot multibox detector (SSD) model that incorporates both spatial and channel attention for hand gesture recognition. We notably extend the recognition distance from 1 meter to 7 meters through the proposed model without sacrificing speed. Besides, we present a long-range hand gesture (LRHG) dataset collected by the USB camera mounted on mobile robots. The hand gestures are collected at discrete distance levels from 1 meter to 7 meters, where most of the hand gestures are small and at low resolution. Experiments with the self-built LRHG dataset show our methods reach the surprising performance-boosting over the state-of-the-art method like the SSD network on both short-range (1 meter) and long-range (up to 7 meters) hand gesture recognition tasks.
Liguang Zhou, Chenping Du, Zhenglong Sun 0001, Tin Lun Lam, Yangsheng Xu
ICRA1
2021 Object-to-Scene: Learning to Transfer Object Knowledge to Indoor Scene Recognition
abstract
Accurate perception of the surrounding scene is helpful for robots to make reasonable judgments and behaviours. Therefore, developing effective scene representation and recognition methods are of significant importance in robotics. Currently, a large body of research focuses on developing novel auxiliary features and networks to improve indoor scene recognition ability. However, few of them focus on directly constructing object features and relations for indoor scene recognition. In this paper, we analyze the weaknesses of current methods and propose an Object-to-Scene (OTS) method, which extracts object features and learns object relations to recognize indoor scenes. The proposed OTS first extracts object features based on the segmentation network and the proposed object feature aggregation module (OFAM). Afterwards, the object relations are calculated and the scene representation is constructed based on the proposed object attention module (OAM) and global relation aggregation module (GRAM). The final results in this work show that OTS successfully extracts object features and learns object relations from the segmentation network. Moreover, OTS outperforms the state-of-the-art methods by more than 2% on indoor scene recognition without using any additional streams. Code is publicly available at: https://github.com/FreeformRobotics/OTS.
Bo Miao, Liguang Zhou, Ajmal Mian, Tin Lun Lam, Yangsheng Xu
IROS2
2021 Design of an SSVEP-based BCI Stimuli System for Attention-based Robot Navigation in Robotic Telepresence
abstract
Brain-computer interface (BCI)-based robotic telepresence provides an opportunity for people with disabilities to control robots remotely without any actual physical movement. However, traditional BCI systems usually require the user to select the navigation direction from visual stimuli in a fixed background, which makes it difficult to control the robot in a dynamic environment during the locomotion. In this paper, a novel SSVEP-based BCI stimuli system is proposed for robotic telepresence. The novel system utilized the live video streamed from the robot onboard camera as the input. By altering and flickering the detected objects in the scene with different frequencies predefined based on their relative positions on the screen, the robot can be navigated based on the user’s attention in a dynamic manner. In order to better differentiate multiple objects (more than the number of frequencies predefined), the task-related component analysis (TRCA) model was trained with a priori offline experimental data to select the front objects with priority. Experiments were conducted to validate the proposed system. Using the system, four human subjects are able to control a humanoid robot to navigate through multiple objects to reach the desired goal. The success rate reaches 87.5% in average.
Xingchao Wang, Xiaopeng Huang, Liguang Zhou, Zhenglong Sun 0001, Yangsheng Xu
IROS4
2021 BORM: Bayesian Object Relation Model for Indoor Scene Recognition
abstract
Scene recognition is a fundamental task in robotic perception. For human beings, scene recognition is reasonable because they have abundant object knowledge of the real world. The idea of transferring prior object knowledge from humans to scene recognition is significant but still less exploited. In this paper, we propose to utilize meaningful object representations for indoor scene representation. First, we utilize an improved object model (IOM) as a baseline that enriches the object knowledge by introducing a scene parsing algorithm pretrained on the ADE20K dataset with rich object categories related to the indoor scene. To analyze the object co-occurrences and pairwise object relations, we formulate the IOM from a Bayesian perspective as the Bayesian object relation model (BORM). Meanwhile, we incorporate the proposed BORM with the PlacesCNN model as the combined Bayesian object relation model (CBORM) for scene recognition and significantly outperforms the state-of-the-art methods on the reduced Places365 dataset, and SUN RGB-D dataset without retraining, showing the excellent generalization ability of the proposed method. Code can be found at https://github.com/FreeformRobotics/BORM.
Liguang Zhou, Jun Cen, Xingchao Wang, Zhenglong Sun 0001, Tin Lun Lam, Yangsheng Xu
IROS1