VLDB 2026 Research / reviewers in the wild / expert
Peiyan Guan
dblp:312/3524
· DBLP profile ↗
6ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0003-1613-2756ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Progressive Self-Supervised Pretraining for Hyperspectral Image ClassificationabstractSelf-supervised learning has demonstrated considerable success in hyperspectral image (HSI) classification when limited labeled data is available. However, inherent dissimilarities among HSIs require self-supervised pre-training from scratch for each HSI dataset. Pre-training on a large amount of unlabeled data can be time consuming. Additionally, the poor quality of some HSIs can limit the performance of self-supervised learning algorithms. To address these issues, we propose to enhance self-supervised pre-training on HSIs with transfer learning. We introduce a progressive self-supervised pre-training framework that acquires strong initialization for the final pre-training on the target HSI dataset by sequentially performing self-supervised pre-training on datasets that are increasingly similar to the target HSI, specifically, first on a large general vision dataset and then on a related HSI dataset. This sequential strategy enables the model to progressively learn from domain-general vision knowledge to target-specific hyperspectral knowledge. To mitigate the catastrophic forgetting in sequential training, we develop a regularization method, called self-supervised elastic weight consolidation, to impose adaptive constraints on the changes to model parameters. Thorough classification experiments on various HSI datasets demonstrate that our framework significantly and consistently improves the self-supervised pre-training on HSIs in terms of both convergence speed and representation quality. Furthermore, our framework exhibits high generalizability and can be applied to various self-supervised learning algorithms. Transfer learning continues to prove its usefulness in self-supervised settings. Peiyan Guan, Edmund Y. Lam |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video RetrievalabstractText-video retrieval is a fundamental task with high practical value in multi-modal research. Inspired by the great success of pre-trained image-text models with large-scale data, such as CLIP, many methods are proposed to transfer the strong representation learning capability of CLIP to text-video retrieval. However, due to the modality difference between videos and images, how to effectively adapt CLIP to the video domain is still underexplored. In this paper, we investigate this problem from two aspects. First, we enhance the transferred image encoder of CLIP for fine-grained video understanding in a seamless fashion. Second, we conduct fine-grained contrast between videos and texts from both model improvement and loss design. Particularly, we propose a fine-grained contrastive model equipped with parallel isomeric attention and dynamic routing, namely PIDRo, for text-video retrieval. The parallel isomeric attention module is used as the video encoder, which consists of two parallel branches modeling the spatial-temporal information of videos from both patch and frame levels. The dynamic routing module is constructed to enhance the text encoder of CLIP, generating informative word representations by distributing the fine-grained information to the related word tokens within a sentence. Such model design provides us with informative patch, frame and word representations. We then conduct token-wise interaction upon them. With the enhanced encoders and the token-wise loss, we are able to achieve finer-grained text-video alignment and more accurate retrieval. PIDRo obtains state-of-the-art performance over various text-video retrieval benchmarks, including MSR-VTT, MSVD, LSMDC, DiDeMo and ActivityNet. Peiyan Guan, Renjing Pei, Jianzhuang Liu, Weimian Li, Jiaxi Gu, Hang Xu 0004, Songcen Xu, Youliang Yan, Edmund Y. Lam |
ICCV | 1 |
| 2022 | Three-Branch Multilevel Attentive Fusion Network for Hyperspectral PansharpeningabstractIn this paper, we propose a three-branch multilevel attentive fusion network (TMA-Net) for hyperspectral pansharpening, which aims to merge low-resolution hyperspectral images (LR-HSIs) and high-resolution panchromatic images (HR-PANs) to obtain HSIs with high resolution. We construct three branches to extract rich features of the two images and the correlation between them, which enables us to capture abundant useful information for pansharpening. We merge the multilevel features extracted by each branch in multiple steps to fully fuse the useful information. An attentive fusion module (AFM) is designed to guide the fusion procedure. It explores the relation between different features and employs attention mechanism to refine them adaptively. The experimental results illustrate the superiority of the TMA-Net. Peiyan Guan, Edmund Y. Lam |
IGARSS | 1 |
| 2022 | Spatial-Spectral Contrastive Learning for Hyperspectral Image ClassificationabstractIn spite of being widely used in hyperspectral image (HSI) classification, most deep learning algorithms require plenty of labeled samples to achieve satisfactory performance. However, manual labeling is very time-consuming and laborious in practice. To solve such problem, we propose a method, called spatial-spectral contrastive learning (SSCL), to learn the representations of HSIs suitable for classification in an unsuper-vised manner. We demonstrate that the useful contents (e.g., semantics) are invariant under the spatial and spectral domains while the uninformative ones are usually not. Thus, we learn powerful representations that model domain-invariant information by defining a contrastive prediction task. Specifically, two signals are constructed for an HSI sample to include information of the two domains, and the representations of these two signals are then optimized to be similar, such that the domain-invariant contents are extracted. We conduct classification experiments on the learned representation with very few labels, the results of which verify the superiority of our method over the state-of-the-art techniques. Peiyan Guan, Edmund Y. Lam |
IGARSS | 1 |
| 2022 | Multistage Dual-Attention Guided Fusion Network for Hyperspectral PansharpeningabstractDeep learning, especially the convolutional neural network, has been widely applied to solve the hyperspectral pansharpening problem. However, most do not explore the intraimage characteristics and the interimage correlation concurrently due to the limited representation ability of the networks, which may lead to insufficient fusion of valuable information encoded in the high-resolution panchromatic images (HR-PANs) and low-resolution hyperspectral images (LR-HSIs). To cope with this problem, we develop a hyperspectral pansharpening method called multistage dual-attention guided fusion network (MDA-Net) to fully extract the important information and accurately fuse them. It employs a three-stream structure, which enables the network to incorporate the intrinsic characteristics of each input and correlation among them simultaneously. In order to combine as much information as possible, we merge the features extracted from three streams in multiple stages, where a dual-attention guided fusion block (DAFB) with spectral and spatial attention mechanisms is utilized to fuse the features efficiently. It identifies the useful components in both spatial and spectral domains, which are beneficial to improving the fusion accuracy. Moreover, we design a multiscale residual dense block (MRDB) to extract dense and hierarchical features, which improves the representation power of the network. Experiments are conducted on both real and simulated datasets. The evaluation results validate the superiority of the MDA-Net. Peiyan Guan, Edmund Y. Lam |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Cross-Domain Contrastive Learning for Hyperspectral Image ClassificationabstractDespite the success of deep learning algorithms in hyperspectral image (HSI) classification, most deep learning models require a large amount of labeled data to optimize the numerous parameters. However, it is very expensive and time-consuming to collect a lot of labeled HSI samples. To cope with this problem, we propose a cross-domain contrastive learning (XDCL) framework to learn representations of HSIs in an unsupervised manner. We demonstrate that the features that are valuable for category identification are shared across the spectral and spatial domains, while the less useful contents tend to be independent. The XDCL extracts such domain-invariant information with a cross-domain discrimination task, i.e., predicting which two representations of different domains are matched. With this insight, our method learns semantically meaningful HSI representations. We develop a simple method to construct effective signals representing the two domains, respectively. Moreover, we randomly mask the signals to improve their semantic level and encourage the representations to dig out more useful abstract factors. In order to evaluate the representation quality, we use the learned representations to train a linear classifier on three hyperspectral datasets with limited labeled samples. Experimental results demonstrate that our method surpasses the state-of-the-art methods by a large margin. Peiyan Guan, Edmund Y. Lam |
IEEE Trans. Geosci. Remote. Sens. | 1 |