Guodong Wang 0006

dblp:77/168-6 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0003-8642-396XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Label-informed knowledge integration: Advancing visual prompt for VLMs adaptation
Yunhong Wang 0001, Guodong Wang 0006, Yingjie Gao 0001, Xiuguo Bao, Di Huang 0001
Comput. Vis. Image Underst.3
2025 Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
abstract
Training-free open-vocabulary semantic segmentation aims to explore the potential of frozen vision-language models (VLM) for segmentation tasks. Recent works reform the inference process of CLIP and utilize the features from the final layer to reconstruct dense representations for segmentation, demonstrating promising performance. However, the final layer tends to prioritize global components over local representations, leading to suboptimal robustness and effectiveness of existing methods. In this paper, we propose CLIPSeg, a novel training-free framework that fully exploits the diverse knowledge across layers in CLIP for dense predictions. Our study unveils two key discoveries: Firstly, the features in the middle layers exhibit high locality awareness and feature coherence compared to the final layer, based on which we propose the coherence enhanced residual attention module that generates semantic-aware attention. Secondly, despite not being directly aligned with the text, the deep layers capture valid local semantics that complement those in the final layer. Leveraging this insight, we introduce the deep semantic integration module to boost the patch semantics in the final block. Experiments conducted on 9 segmentation benchmarks with various CLIP models demonstrate that CLIPSeg consistently outperforms all training-free methods by substantial margins, e.g., a 7.8 % improvement in average mIoU for CLIP with a ViT-L backbone, and competes with learning-based counterparts in generalizing to novel concepts in an efficient way.
Guodong Wang 0006, Qingjie Liu 0001, Di Huang 0001
AAAI2
2025 Towards Training-free Anomaly Detection with Vision and Language Foundation Models
abstract
Anomaly detection is valuable for real-world applications, such as industrial quality inspection. However, most approaches focus on detecting local structural anomalies while neglecting compositional anomalies incorporating logical constraints. In this paper, we introduce LogSAD, a novel multi-modal framework that requires no training for both Logical and Structural Anomaly Detection. First, we propose a match-of-thought architecture that employs advanced large multi-modal models (i.e. GPT-4V) to generate matching proposals, formulating interests and compositional rules of thought for anomaly detection. Second, we elaborate on multi-granularity anomaly detection, consisting of patch tokens, sets of interests, and composition matching with vision and language foundation models. Subsequently, we present a calibration module to align anomaly scores from different detectors, followed by integration strategies for the final decision. Consequently, our approach addresses both logical and structural anomaly detection within a unified framework and achieves state-of-the-art results without the need for training, even when compared to supervised approaches, highlighting its robustness and effectiveness. Code is available at https://github.com/zhang0jhon/LogSAD.
Guodong Wang 0006, Yizhou Jin, Di Huang 0001
CVPR2
2025 Anomaly-aware self-supervised feature learning for weakly supervised video anomaly detection
Zhen Yang 0037, Guodong Wang 0006, Yuanfang Guo, Xiuguo Bao, Di Huang 0001
Comput. Vis. Image Underst.2
2025 Multi-Grained Contrastive Learning for Text-Supervised Open-Vocabulary Semantic Segmentation
abstract
Learning open-vocabulary semantic segmentation (OVSS) from text supervision has recently received increasing attention for its promising potential in real-world applications. However, only with image-level supervision, it struggles to achieve dense and robust cross-modal alignment and thus limits pixel-level predictions. In this article, we present a novel approach to this task with M ulti- G rained C ross-modal C ontrastive L earning, named MGCCL. Specifically, unlike current solutions restricted by coarse image/object-text alignment, MGCCL constructs pseudo multi-granular semantic correspondences at the object-, part-, and pixel-level and collaborates with hard sampling strategies to conduct cross-modal contrastive learning, significantly facilitating fine-grained alignment. Further, we develop an adaptive semantic unit which flexibly harnesses the learned multi-grained cross-modal alignment capabilities to effectively mitigate the under- and over-segmentation issues arising from the per-group and per-pixel units. Extensive experiments over a broad suite of eight segmentation benchmarks show that our approach delivers significant advancements over state-of-the-art counterparts, demonstrating its effectiveness.
Pu Ge, Guodong Wang 0006, Qingjie Liu 0001, Di Huang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
Zhi Cai, Guodong Wang 0006, Zheng Ge, Xiangyu Zhang 0005, Di Huang 0001
BMVC3
2024 Rotation Has Two Sides: Evaluating Data Augmentation for Deep One-class Classification
abstract
One-class classification (OCC) involves predicting whether a new data is normal or anomalous based solely on the data from a single class during training. Various attempts have been made to learn suitable representations for OCC within a self-supervised framework. Notably, discriminative methods that use geometric visual transformations, such as rotation, to generate pseudo-anomaly samples have exhibited impressive detection performance. Although rotation is commonly viewed as a distribution-shifting transformation and is widely used in the literature, the cause of its effectiveness remains a mystery. In this study, we are the first to make a surprising observation: there exists a strong linear relationship (Pearson's Correlation, $r > 0.9$) between the accuracy of rotation prediction and the performance of OCC. This suggests that a classifier that effectively distinguishes different rotations is more likely to excel in OCC, and vice versa. The root cause of this phenomenon can be attributed to the transformation bias in the dataset, where representations learned from transformations already present in the dataset tend to be less effective, making it essential to accurately estimate the transformation distribution before utilizing pretext tasks involving these transformations for reliable self-supervised representation learning. To the end, we propose a novel two-stage method to estimate the transformation distribution within the dataset. In the first stage, we learn general representations through standard contrastive pre-training. In the second stage, we select potentially semantics-preserving samples from the entire augmented dataset, which includes all rotations, by employing density matching with the provided reference distribution. By sorting samples based on semantics-preserving versus shifting transformations, we achieve improved performance on OCC benchmarks.
Guodong Wang 0006, Yunhong Wang 0001, Xiuguo Bao, Di Huang 0001
ICLR1
2023 Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation for Anomaly Detection
abstract
Anomaly detection (AD), aiming to find samples that deviate from the training distribution, is essential in safety-critical applications. Though recent self-supervised learning based attempts achieve promising results by creating virtual outliers, their training objectives are less faithful to AD which requires a concentrated inlier distribution as well as a dispersive outlier distribution. In this paper, we propose Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation (UniCon-HA), taking into account both the requirements above. Specifically, we explicitly encourage the concentration of inliers and the dispersion of virtual outliers via supervised and unsupervised contrastive losses, respectively. Considering that standard contrastive data augmentation for generating positive views may induce outliers, we additionally introduce a soft mechanism to re-weight each augmented inlier according to its deviation from the inlier distribution, to ensure a purified concentration. Moreover, to prompt a higher concentration, inspired by curriculum learning, we adopt an easy-to-hard hierarchical augmentation strategy and perform contrastive aggregation at different depths of the network based on the strengths of data augmentation. Our method is evaluated under three AD settings including unlabeled one-class, unlabeled multi-class, and labeled multi-class, demonstrating its consistent superiority over other competitors.
Guodong Wang 0006, Yunhong Wang 0001, Jie Qin 0004, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
ICCV1
2022 Video Anomaly Detection by Solving Decoupled Spatio-Temporal Jigsaw Puzzles
Guodong Wang 0006, Yunhong Wang 0001, Jie Qin 0004, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
ECCV (10)1
2021 Student-Teacher Feature Pyramid Matching for Anomaly Detection
Guodong Wang 0006, Shumin Han, Errui Ding, Di Huang 0001
BMVC1