VLDB 2026 Research / reviewers in the wild / expert
Haoyan Guan
dblp:239/7368
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-1936-2442ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Depth perception in virtual reality: The impact of spatial interaction and color-based cues
Yusi Sun, Haoyan Guan, Leith Kin Yip Chan |
Comput. Graph. | 2 |
| 2025 | Few-Shot Anomaly Detection via Category-Agnostic Registration LearningabstractMost existing anomaly detection (AD) methods require a dedicated model for each category. Such a paradigm, despite its promising results, is computationally expensive and inefficient, thereby failing to meet the requirements for real-world applications. Inspired by how humans detect anomalies, by comparing a query image to known normal ones, this article proposes a novel few-shot AD (FSAD) framework. Using a training set of normal images from various categories, registration, aiming to align normal images of the same categories, is leveraged as the proxy task for self-supervised category-agnostic representation learning. At test time, an image and its corresponding support set, consisting of a few normal images from the same category, are supplied, and anomalies are identified by comparing the registered features of the test image to its corresponding support image features. Such a setup enables the model to generalize to novel test categories. It is, to our best knowledge, the first FSAD method that requires no model fine-tuning for novel categories: enabling a single model to be applied to all categories. Extensive experiments demonstrate the effectiveness of the proposed method. Particularly, it improves the current state-of-the-art (SOTA) for FSAD by 11.3% and 8.3% on the MVTec and MPDD benchmarks, respectively. The source code is available at https://github.com/Haoyan-Guan/CAReg. Chaoqin Huang, Haoyan Guan, Aofan Jiang, Ya Zhang 0002, Michael W. Spratling, Xinchao Wang, Yanfeng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | One Prompt Word is Enough to Boost Adversarial Robustness for Pre-Trained Vision-Language ModelsabstractLarge pre-trained Vision-Language Models (VLMs) like CLIP, despite having remarkable generalization ability, are highly vulnerable to adversarial examples. This work studies the adversarial robustness of VLMs from the novel perspective of the text prompt instead of the extensively studied model weights (frozen in this work). We first show that the effectiveness of both adversarial attack and defense are sensitive to the used text prompt. Inspired by this, we propose a method to improve resilience to adversarial attacks by learning a robust text prompt for VLMs. The proposed method, named Adversarial Prompt Tuning (APT), is effective while being both computationally and data efficient. Extensive experiments are conducted across 15 datasets and 4 data sparsity schemes (from 1-shot to full training data settings) to show APT's superiority over hand-engineered prompts and other state-of-the-art adaption methods. APT demonstrated excellent abilities in terms of the in-distribution performance and the generalization under input distribution shift and across datasets. Surprisingly, by simply adding one learned word to the prompts, APT can significantly boost the accuracy and robustness ($\epsilon=4/255$) over the hand-engineered prompts by +13% and +8.5% on average respectively. The improvement further increases, in our most effective setting, to +26.4% for accuracy and +16.7% for robustness. Code is available at https://github.com/TreeLLi/APT. Lin Li 0070, Haoyan Guan, Jianing Qiu, Michael W. Spratling |
CVPR | 2 |
| 2024 | Query semantic reconstruction for background in few-shot segmentationabstractAbstract Few-shot segmentation (FSS) aims to segment unseen classes using a few annotated samples. Typically, a prototype representing the foreground class is extracted from annotated support image(s) and is matched to features representing each pixel in the query image. However, models learnt in this way are insufficiently discriminatory, and often produce false positives: misclassifying background pixels as foreground. Some FSS methods try to address this issue by using the background in the support image(s) to help identify the background in the query image. However, the backgrounds of these images are often quite distinct, and hence, the support image background information is uninformative. This article proposes a method, QSR, that extracts the background from the query image itself, and as a result is better able to discriminate between foreground and background features in the query image. This is achieved by modifying the training process to associate prototypes with class labels including known classes from the training data and latent classes representing unknown background objects. This class information is then used to extract a background prototype from the query image. To successfully associate prototypes with class labels and extract a background prototype that is capable of predicting a mask for the background regions of the image, the machinery for extracting and using foreground prototypes is induced to become more discriminative between different classes. Experiments achieves state-of-the-art results for both 1-shot and 5-shot FSS on the PASCAL- $$5^{i}$$ 5 i and COCO- $$20^{i}$$ 20 i dataset. As QSR operates only during training, results are produced with no extra computational complexity during testing. Haoyan Guan, Michael W. Spratling |
Vis. Comput. | 1 |
| 2022 | Registration Based Few-Shot Anomaly Detection
Chaoqin Huang, Haoyan Guan, Aofan Jiang, Ya Zhang 0002, Michael W. Spratling, Yanfeng Wang 0001 |
ECCV (24) | 2 |
| 2022 | CobNet: Cross Attention on Object and Background for Few-Shot SegmentationabstractFew-shot segmentation aims to segment images containing objects from previously unseen classes using only a few annotated samples. Most current methods focus on using object information extracted, with the aid of human annotations, from support images to identify the same objects in new query images. However, background information can also be useful to distinguish objects from their surroundings. Hence, some previous methods also extract background information from the support images. In this paper, we argue that such information is of limited utility, as the background in different images can vary widely. To overcome this issue, we propose CobNet which utilises information about the background that is extracted from the query images without annotations of those images. Experiments show that our method achieves a mean Intersection-over-Union score of 61.4% and 37.8% for 1-shot segmentation on PASCAL-5iand COCO-20irespectively, outperforming previous methods. It is also shown to produce state-of-the-art performances of 53.7% for weakly-supervised few-shot segmentation, where no annotations are provided for the support images. Haoyan Guan, Michael W. Spratling |
ICPR | 1 |
| 2018 | Deep Dual-view Network with Smooth Loss for Spinal Metastases ClassificationabstractSpinal metastases have a high incidence among cancer patients and may later develop to metastatic spinal cord compression (MSCC). Early detection of spinal metastases is critical for optimal treatment. The diagnosis is usually facilitated with computed tomography (CT) scans, which requires considerable efforts from well-trained radiologist. In this paper, we explore automatic spinal metastases classification based on CT images. Considering the unique characteristics of spinal CT images, a novel Deep Dual-view Network is proposed which contains two branches: a X-Y Conv Branch to extract the features for each individual cross-sectional image slice, and a Z Conv Branch to capture z-direction features from neighboring images. The features from the two branches are then fused to generate the final prediction, which imitates the doctor's way of combining cross sections with sagittal or coronal sections. Considering a tumor usually presents in multiple consecutive image slices, a smooth loss is introduced to maintain the label consistency of adjacent images. To validate the proposed approach, we collect a dataset of 316 patients with spinal metastases. Experimental results on this data set have demonstrated the effectiveness and the robustness of the proposed Deep Dual-view network. Haoyan Guan, Guangyu Yao, Yexun Zhang, Yujun Gu, Ya Zhang 0002, Xiao Gu 0001 |
VCIP | 1 |