Guanglei Yang

dblp:06/10218 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-5324-3642ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Parameter-efficient multimodal adaptation for adverse condition depth estimation
Guanglei Yang, Yongqiang Zhang 0007, Zhun Zhong, Wangmeng Zuo
Expert Syst. Appl.1
2026 ConSept: Continual semantic segmentation via adapter-based vision transformer
Bowen Dong 0001, Guanglei Yang, Lei Zhang 0006, Wangmeng Zuo
Pattern Recognit. Lett.2
2025 ReMP-AD: Retrieval-Enhanced Multi-Modal Prompt Fusion for Few-Shot Industrial Visual Anomaly Detection
Hongchi Ma, Guanglei Yang, Debin Zhao, Yanli Ji, Wangmeng Zuo
ICCV2
2025 Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
abstract
With the emergence of transformer-based architectures and large language models (LLMs), the accuracy of road scene perception has substantially advanced. Nonetheless, current road scene segmentation approaches are predominantly trained on closed-set data, resulting in insufficient detection capabilities for out-of-distribution (OOD) objects. To overcome this limitation, road anomaly detection methods have been proposed. However, existing methods primarily depend on image inpainting and OOD distribution detection techniques, facing two critical issues: (1) inadequate consideration of the objectiveness attributes of anomalous regions, causing incomplete segmentation when anomalous objects share similarities with known classes, and (2) insufficient attention to environmental constraints, leading to the detection of anomalies irrelevant to autonomous driving tasks. In this paper, we propose a novel framework termed Segmenting Objectiveness and Task-Awareness (SOTA) for autonomous driving scenes. Specifically, SOTA enhances the segmentation of objectiveness through a Semantic Fusion Block (SFB) and filters anomalies irrelevant to road navigation tasks using a Scene-understanding Guided Prompt-Context Adaptor (SG-PCA). Extensive empirical evaluations on multiple benchmark datasets, including Fishyscapes Lost and Found, Segment-Me-If-You-Can, and RoadAnomaly, demonstrate that the proposed SOTA consistently improves OOD detection performance across diverse detectors, achieving robust and accurate segmentation outcomes.
Mi Zheng, Guanglei Yang, Zitong Huang, Zhenhua Guo 0001, Kevin Han, Wangmeng Zuo
ACM Multimedia2
2025 MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
abstract
Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to adequately distinguish between perception-induced hallucinations and reasoning-induced hallucinations. This failure constitutes a significant issue and hinders the diagnosis of multimodal reasoning failures within MLLMs. To address this, we propose the MIRAGE benchmark, which isolates reasoning hallucinations by constructing questions where input images are correctly perceived by MLLMs yet reasoning errors persist. MIRAGE introduces multi-granular evaluation metrics: accuracy, factuality, and LLMs hallucination score for hallucination quantification. Our analysis reveals strong correlations between question types and specific hallucination patterns, particularly systematic failures of MLLMs in spatial reasoning involving complex relationships (\emph{e.g.}, complex geometric patterns across images). This highlights a critical limitation in the reasoning capabilities of current MLLMs and provides targeted insights for hallucination mitigation on specific types. To address these challenges, we propose Logos, a method that combines curriculum reinforcement fine-tuning to encourage models to generate logic-consistent reasoning chains by stepwise reducing learning difficulty, and collaborative hint inference to reduce reasoning complexity. Logos establishes a baseline on MIRAGE, and reduces the logical hallucinations in original base models. Link: \url{https://bit.ly/25mirage}.
Bowen Dong 0001, Minheng Ni, Zitong Huang, Guanglei Yang, Wangmeng Zuo, Lei Zhang 0006
NeurIPS4
2025 DualAug: Exploiting additional heavy augmentation with OOD data rejection
Yiwen Guo, Qizhang Li, Guanglei Yang, Wangmeng Zuo
Neurocomputing4
2025 FILP-3D: Enhancing 3D few-shot class-incremental learning with pre-trained vision-language models
Wan Xu, Tianyuan Qu, Guanglei Yang, Yiwen Guo, Wangmeng Zuo
Pattern Recognit.4
2025 Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multitask Learning Perspective
abstract
Human beings can leverage knowledge from relative tasks to improve learning on a primary task. Similarly, multitask learning (MTL) methods suggest using auxiliary tasks to enhance a neural network's performance on a specific primary task. However, previous methods often select auxiliary tasks carefully but treat them as secondary during training. The weights assigned to auxiliary losses are typically smaller than the primary loss weight, leading to insufficient training on auxiliary tasks and ultimately failing to support the main task effectively. To address this issue, we propose an uncertainty-based impartial learning method that ensures balanced training across all tasks. In addition, we consider both gradients and uncertainty information during backpropagation to further improve performance on the primary task. Extensive experiments show that our method achieves performance comparable to or better than state-of-the-art approaches. Moreover, our weighting strategy is effective and robust in enhancing the performance of the primary task regardless of the noise auxiliary tasks' pseudolabels.
Yuanze Li, Chun-Mei Feng 0001, Qilong Wang 0001, Guanglei Yang, Wangmeng Zuo
IEEE Trans. Neural Networks Learn. Syst.4
2024 UniM2AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving
Jian Zou 0005, Guanglei Yang, Zhenhua Guo 0001, Tao Luo 0014, Chun-Mei Feng 0001, Wangmeng Zuo
ECCV (22)3
2024 DPDFormer: A Coarse-to-Fine Model for Monocular Depth Estimation
abstract
Monocular depth estimation attracts great attention from computer vision researchers for its convenience in acquiring environment depth information. Recently classification-based MDE methods show its promising performance and begin to act as an essential role in many multi-view applications such as reconstruction and 3D object detection. However, existed classification-based MDE models usually apply fixed depth range discretization strategy across a whole scene. This fixed depth range discretization leads to the imbalance of discretization scale among different depth ranges, resulting in the inexact depth range localization. In this article, to alleviate the imbalanced depth range discretization problem in classification-based monocular depth estimation (MDE) method we follow the coarse-to-fine principle and propose a novel depth range discretization method called depth post-discretization (DPD). Based on a coarse depth anchor roughly indicating the depth range, the DPD generates the depth range discretization adaptively for every position. The depth range discretization with DPD is more fine-grained around the actual depth, which is beneficial for locating the depth range more precisely for each scene position. Besides, to better manage the prediction of the coarse depth anchor and depth probability distribution for calculating the final depth, we design a dual-decoder transformer-based network, i.e., DPDFormer, which is more compatible with our proposed DPD method. We evaluate DPDFormer on popular depth datasets NYU Depth V2 and KITTI. The experimental results prove the superior performance of our proposed method.
Chunpu Liu, Guanglei Yang, Wangmeng Zuo, Tianyi Zang
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Relative order constraint for monocular depth estimation
Chunpu Liu, Wangmeng Zuo, Guanglei Yang, Wanlong Li, Hongbo Zhang 0004, Tianyi Zang
Appl. Intell.3
2023 Self-training transformer for source-free domain adaptation
Guanglei Yang, Zhun Zhong, Mingli Ding, Nicu Sebe, Elisa Ricci 0001
Appl. Intell.1
2023 Uncertainty-Aware Contrastive Distillation for Incremental Semantic Segmentation
abstract
A fundamental and challenging problem in deep learning is catastrophic forgetting, i.e., the tendency of neural networks to fail to preserve the knowledge acquired from old tasks when learning new tasks. This problem has been widely investigated in the research community and several Incremental Learning (IL) approaches have been proposed in the past years. While earlier works in computer vision have mostly focused on image classification and object detection, more recently some IL approaches for semantic segmentation have been introduced. These previous works showed that, despite its simplicity, knowledge distillation can be effectively employed to alleviate catastrophic forgetting. In this paper, we follow this research direction and, inspired by recent literature on contrastive learning, we propose a novel distillation framework, Uncertainty-aware Contrastive Distillation (UCD). In a nutshell, UCDis operated by introducing a novel distillation loss that takes into account all the images in a mini-batch, enforcing similarity between features associated to all the pixels from the same classes, and pulling apart those corresponding to pixels from different classes. In order to mitigate catastrophic forgetting, we contrast features of the new model with features extracted by a frozen model learned at the previous incremental step. Our experimental results demonstrate the advantage of the proposed distillation technique, which can be used in synergy with previous IL approaches, and leads to state-of-art performance on three commonly adopted benchmarks for incremental semantic segmentation.
Guanglei Yang, Enrico Fini, Dan Xu 0002, Paolo Rota, Mingli Ding, Moin Nabi, Xavier Alameda-Pineda, Elisa Ricci 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Continual Attentive Fusion for Incremental Learning in Semantic Segmentation
abstract
International audience
Guanglei Yang, Enrico Fini, Dan Xu 0002, Paolo Rota, Mingli Ding, Hao Tang 0005, Xavier Alameda-Pineda, Elisa Ricci 0001
IEEE Trans. Multim.1
2022 Bi-directional class-wise adversaries for unsupervised domain adaptation
Guanglei Yang, Mingli Ding, Yongqiang Zhang 0007
Appl. Intell.1
2021 Transformer-Based Attention Networks for Continuous Pixel-Wise Prediction
abstract
While convolutional neural networks have shown a tremendous impact on various computer vision tasks, they generally demonstrate limitations in explicitly modeling long-range dependencies due to the intrinsic locality of the convolution operation. Initially designed for natural language processing tasks, Transformers have emerged as alternative architectures with innate global self-attention mechanisms to capture long-range dependencies. In this paper, we propose TransDepth, an architecture that benefits from both convolutional neural networks and transformers. To avoid the network losing its ability to capture locallevel details due to the adoption of transformers, we propose a novel decoder that employs attention mechanisms based on gates. Notably, this is the first paper that applies transformers to pixel-wise prediction problems involving continuous labels (i.e., monocular depth prediction and surface normal estimation). Extensive experiments demonstrate that the proposed TransDepth achieves state-of-theart performance on three challenging datasets. Our code is available at: https://github.com/ygjwd12345/TransDepth.
Guanglei Yang, Hao Tang 0005, Mingli Ding, Nicu Sebe, Elisa Ricci 0001
ICCV1
2020 Bi-Directional Generation for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation facilitates the unlabeled target domain relying on well-established source domain information. The conventional methods forcefully reducing the domain discrepancy in the latent space will result in the destruction of intrinsic data structure. To balance the mitigation of domain gap and the preservation of the inherent structure, we propose a Bi-Directional Generation domain adaptation model with consistent classifiers interpolating two intermediate domains to bridge source and target domains. Specifically, two cross-domain generators are employed to synthesize one domain conditioned on the other. The performance of our proposed method can be further enhanced by the consistent classifiers and the cross-domain alignment constraints. We also design two classifiers which are jointly optimized to maximize the consistency on target sample prediction. Extensive experiments verify that our proposed model outperforms the state-of-the-art on standard cross domain visual benchmarks.
Guanglei Yang, Haifeng Xia, Mingli Ding, Zhengming Ding
AAAI1
2015 High-order information for robust iris recognition under less controlled conditions
abstract
Iris recognition has achieved great progress in cooperative environments in the past decades. However, in less controlled conditions it is still an open and challenging problem because of severe noisy factors induced by non-cooperative subjects. For handling this challenging problem, we propose a method called ordinal measure of outer product tensor (O2PT) which leverages the high-order information of image features. O2PT consists of two components. First we compute outer product tensors of raw features (e.g. SIFT) which are vectorized and locally aggregated, characterizing the second-order statistics of raw features. And then we compute the ordinal measure of the aggregated outer product tensors to model the order relation of iris texture, which makes the representation more compact and robust to noise and illumination changes. Furthermore, we combine two modalities to improve the matching performance, namely, O2PT for iris image matching and Fisher Vector (FV), which also exploits the high-order information, for eye image matching. We have achieved competitive matching performance on the challenging UBIRIS.v2 and CASIA-Iris-Thousand databases.
Guanglei Yang, Hui Zeng 0001, Peihua Li, Lei Zhang 0006
ICIP1