Lin Qi 0004

dblp:82/3116-4 · DBLP profile ↗
← Back
55ranked-venue papers
8as first author
29since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 6 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 4Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 MSF-Net: Multi-stage Feature Extraction and Fusion for Robust Photometric Stereo
Shiyu Qin, Lin Qi 0004, Junyu Dong
PRCV (10)4
2025 Weakly supervised camouflaged object detection based on the SAM model and mask guidance
Lin Qi 0004, Junyu Dong
Image Vis. Comput.3
2025 Dynamic Cross-Modal Feature Interaction Network for Hyperspectral and LiDAR Data Classification
abstract
Hyperspectral image (HSI) and light detection and ranging (LiDAR) data joint classification is a challenging task. Existing multisource remote sensing data classification methods often rely on human-designed frameworks for feature extraction, which heavily depend on expert knowledge. To address these limitations, we propose a novel dynamic cross-modal feature interaction network (DCMNet), the first framework leveraging a dynamic routing mechanism for HSI and LiDAR classification. Specifically, our approach introduces three feature interaction blocks: bilinear spatial attention block (BSAB), bilinear channel attention block (BCAB), and integration convolutional block (ICB). These blocks are designed to effectively enhance spatial, spectral, and discriminative feature interactions. A multilayer routing space with routing gates is designed to determine optimal computational paths, enabling data-dependent feature fusion. Additionally, bilinear attention mechanisms are employed to enhance feature interactions in spatial and channel representations. Extensive experiments on three public HSI and LiDAR datasets demonstrate the superiority of DCMNet over the state-of-the-art methods. Our codes are available athttps://github.com/oucailab/DCMNet.
Junyan Lin, Feng Gao 0005, Lin Qi 0004, Junyu Dong, Qian Du 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Sparse Focus Network for Multi-Source Remote Sensing Data Classification
abstract
Multi-source remote sensing data classification has emerged as a prominent research topic with the advancement of various sensors. Existing multi-source data classification methods are susceptible to irrelevant information interference during multi-source feature extraction and fusion. To solve this issue, we propose a sparse focus network for multi-source data classification. Sparse attention is employed in Transformer block for HSI and SAR/LiDAR feature extraction, thereby the most useful self-attention values are maintained for better feature aggregation. Furthermore, cross-attention is used to enhance multi-source feature interactions, and further improves the efficiency of cross-modal feature fusion. Experimental results on the Berlin and Houston2018 datasets highlight the effectiveness of SF-Net, outperforming existing state-of-the-art methods.
Xuepeng Jin, Junyan Lin, Feng Gao 0005, Lin Qi 0004
IGARSS4
2024 Dual-Stream Attention Network for Hyperspectral Image Unmixing
abstract
Hyperspectral image (HSI) contains abundant spatial and spectral information, making it highly valuable for unmixing. In this paper, we propose a Dual-Stream Attention Network (DSANet) for HSI unmixing. The endmembers and abundance of a pixel in HSI have high correlations with its adjacent pixels. Therefore, we adopt a “many to one” strategy to estimate the abundance of the central pixel. In addition, we adopt multiview spectral method, dividing spectral bands into multiple partitions with low correlations to estimate abundances. To aggregate the estimated abundances for complementary from the two branches, we design a cross-fusion attention network to enhance valuable information. Extensive experiments have been conducted on two real datasets, which demonstrate the effectiveness of our DSANet.
Yufang Wang, Wenmin Wu, Lin Qi 0004, Feng Gao 0005
IGARSS3
2024 Exploring Cross-Domain Few-Shot Classification via Frequency-Aware Prompting
Tiange Zhang, Feng Gao 0005, Lin Qi 0004, Junyu Dong
IJCAI4
2024 Superpixel Cost Volume Excitation for Stereo Matching
Shanglong Liu, Lin Qi 0004, Junyu Dong, Wen-Xiang Gu, Liyi Xu
PRCV (6)2
2024 Image Gradient-Aided Photometric Stereo Network
Lin Qi 0004, Shiyu Qin, Yakun Ju, Junyu Dong
PRICAI (3)2
2024 A Sketch-texture Retrieval Framework using Perceptual Similarity
Ying Gao 0005, Hafiza Sadia Nawaz, Lin Qi 0004, Junyu Dong
Knowl. Based Syst.4
2024 Hierarchical Attention and Parallel Filter Fusion Network for Multisource Data Classification
abstract
Hyperspectral image (HSI) and synthetic aperture radar (SAR) data joint classification is a crucial and yet challenging task in the field of remote sensing image interpretation. However, feature modeling in the existing methods is deficient to exploit the abundant global, spectral, and local features simultaneously, leading to suboptimal classification performance. To solve the problem, we propose a hierarchical attention and parallel filter fusion network for multisource data classification. Concretely, we design a hierarchical attention module (HAM) for hyperspectral feature extraction. This module integrates global, spectral, and local features simultaneously to provide more comprehensive feature representation. In addition, we develop parallel filter fusion module (PFFM), which enhances cross-modal feature interactions among different spatial locations in the frequency domain. Extensive experiments on two multisource remote sensing data classification datasets verify the superiority of our proposed method over current state-of-the-art classification approaches. Specifically, our proposed method achieves 91.44% and 80.51% of overall accuracy (OA) on the respective datasets, highlighting its superior performance.
Feng Gao 0005, Junyu Dong, Lin Qi 0004
IEEE Geosci. Remote. Sens. Lett.4
2024 Deep Attention-Guided Spatial-Spectral Network for Hyperspectral Image Unmixing
abstract
Deep-learning (DL)-based methods have been increasingly used in hyperspectral unmixing (HU), especially the recent trend of unsupervised autoencoder (AE) networks, which have achieved excellent performances. Although some existing unmixing methods take spatial information into account, the utilization of spatial structure is not sufficient and effective. In this letter, we present a deep attention-guided spatial–spectral network for hyperspectral image (HSI) unmixing called DASS-Net, which adopts a parallel dual-stream structure. We design a neighborhood spatial attention (NSA) module, where the abundance features of the central pixel are dynamically weighted by the coarse-grained features of the neighborhood pixels. In addition, a dual-gated mechanism is introduced to further integrate and express the spatial and spectral information. Experimental results show that the proposed DASS-Net performs particularly well in endmember extraction and outperforms all compared methods.
Lin Qi 0004, Mengyi Yue, Feng Gao 0005, Bing Cao 0002, Junyu Dong, Xinbo Gao 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 Camouflaged Object Detection with Feature Grafting and Distractor Aware
abstract
The task of Camouflaged Object Detection (COD) aims to accurately segment camouflaged objects that integrated into the environment, which is more challenging than ordinary detection as the texture between the target and background is visually indistinguishable. In this paper, we proposed a novel Feature Grafting and Distractor Aware network (FDNet) to handle the COD task. Specifically, we use CNN and Transformer to encode multi-scale images in parallel. In order to better explore the advantages of the two encoders, we design a cross-attention-based Feature Grafting Module to graft features extracted from Transformer branch into CNN branch, after which the features are aggregated in the Feature Fusion Module. A Distractor Aware Module is designed to explicitly model the two possible distractor in the COD task to refine the coarse camouflage map. We also proposed the largest artificial camouflaged object dataset which contains 2000 images with annotations, named ACOD2K. We conducted extensive experiments on four widely used benchmark datasets and the ACOD2K dataset. The results show that our method significantly outperforms other state-of-the-art methods. The code and the ACOD2K will be available at https://github.com/syxvision/FDNet.
Lin Qi 0004
ICME3
2023 Edge-Aware Mirror Network for Camouflaged Object Detection
abstract
Existing edge-aware camouflaged object detection (COD) methods normally output the edge prediction in the early stage. However, edges are important and fundamental factors in the following segmentation task. Due to the high visual similarity between camouflaged targets and the surroundings, edge prior predicted in early stage usually introduces erroneous foreground-background and contaminates features for segmentation. To tackle this problem, we propose a novel Edge-aware Mirror Network (EAMNet), which models edge detection and camouflaged object segmentation as a cross refinement process. More specifically, EAMNet has a two-branch architecture, where a segmentation-induced edge aggregation module and an edge- induced integrity aggregation module are designed to cross-guide the segmentation branch and edge detection branch. A guided-residual channel attention module which leverages the residual connection and gated convolution finally better extracts structural details from low-level features. Quantitative and qualitative experiment results show that EAMNet outperforms existing cutting-edge baselines on three widely used COD datasets. Codes are available at https://github.com/sdy1999/EAMNet.
Dongyue Sun, Shiyao Jiang, Lin Qi 0004
ICME3
2023 Hyperspectral and SAR Image Classification via Recursive Feature Interactive Fusion Network
abstract
Most of existing mutli-source remote sensing data classification methods are based on convolutional neural networks. Recently, the emergence of Vision Transformer greatly challenges the dominance of CNN-based methods. The self-attention mechanism in Transformer and other dynamic networks imply that high-order feature interactions are beneficial to improve the feature representation and fusion. To explore the high-order feature interactions in multi-source image fusion, in this paper, we proposed a novel recursive feature interactive fusion network. It is composed of cross-shaped window self-attention encoder, and recursive feature interactive fusion. We use gated convolution recursively to mix multi-modal features and exploit their spatial relations. Experimental results on two datasets show that the proposed method achieves better performance than closely related methods.
Junyan Lin, Feng Gao 0005, Lin Qi 0004, Junyu Dong
IGARSS4
2023 Multiview Siamese Collaborative Network for Hyperspectral Image Unmixing
abstract
Existing hyperspectral image unmixing methods acquire information from single-view, and therefore can hardly make full use of diverse spectral information, and the learning of space and spectrum is often relatively independent and cannot be combined effectively. In order to solve the problem of insufficient feature representation caused by the single-view spectral information and the lack of close relationship between spatial and spectral learning, we introduce the idea of multiview data construction, which divides the spectral bands into different views for multiview learning. In addition, we proposed a spatial-spectral siamese network. Deep collaborative learning is used to construct an unmixing model by combining multiview representation and the siamese network. Experimental results on the Jasper Ridge and Urban datasets demonstrate the effectiveness of the proposed method.
Zimo Yang, Lin Qi 0004, Feng Gao 0005, Junyan Lin
IGARSS2
2023 CLIP-Hand3D: Exploiting 3D Hand Pose Estimation via Context-Aware Prompting
abstract
Contrastive Language-Image Pre-training (CLIP) starts to emerge in many computer vision tasks and has achieved promising performance. However, it remains underexplored whether CLIP can be generalized to 3D hand pose estimation, as bridging text prompts with pose-aware features presents significant challenges due to the discrete nature of joint positions in 3D space. In this paper, we make one of the first attempts to propose a novel 3D hand pose estimator from monocular images, dubbed as CLIP-Hand3D, which successfully bridges the gap between text prompts and irregular detailed pose distribution. In particular, the distribution order of hand joints in various 3D space directions is derived from pose labels, forming corresponding text prompts that are subsequently encoded into text representations. Simultaneously, 21 hand joints in the 3D space are retrieved, and their spatial distribution (in x, y, and z axes) is encoded to form pose-aware features. Subsequently, we maximize semantic consistency for a pair of pose-text features following a CLIP-based contrastive learning paradigm. Furthermore, a coarse-to-fine mesh regressor is designed, which is capable of effectively querying joint-aware cues from the feature pyramid. Extensive experiments on several public hand benchmarks show that the proposed model attains a significantly faster inference speed while achieving state-of-the-art performance compared to methods utilizing the similar scale backbone. Code is available at: https://github.com/ShaoXiang23/CLIP_Hand_Demo.
Shaoxiang Guo, Lin Qi 0004, Junyu Dong
ACM Multimedia3
2023 SAWU-Net: Spatial Attention Weighted Unmixing Network for Hyperspectral Images
abstract
Hyperspectral unmixing is a critical yet challenging task in hyperspectral image interpretation.Recently, great efforts have been made to solve the hyperspectral unmixing task via deep autoencoders.However, existing networks mainly focus on extracting spectral features from mixed pixels, and the employment of spatial feature prior knowledge is still insufficient.To this end, we put forward a spatial attention weighted unmixing network, dubbed as SAWU-Net, which learns a spatial attention network and a weighted unmixing network in an end-to-end manner for better spatial feature exploitation.In particular, we design a spatial attention module, which consists of a pixel attention block and a window attention block to efficiently model pixelbased spectral information and patch-based spatial information, respectively.While in the weighted unmixing framework, the central pixel abundance is dynamically weighted by the coarsegrained abundances of surrounding pixels.In addition, SAWU-Net generates dynamically adaptive spatial weights through the spatial attention mechanism, so as to dynamically integrate surrounding pixels more effectively.Experimental results on real and synthetic datasets demonstrate the better accuracy and superiority of SAWU-Net, which reflects the effectiveness of the proposed spatial attention mechanism.
Lin Qi 0004, Xuewen Qin, Feng Gao 0005, Junyu Dong, Xinbo Gao 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 Multiview Spatial-Spectral Two-Stream Network for Hyperspectral Image Unmixing
abstract
Linear spectral unmixing is an important technique in the analysis of mixed pixels in hyperspectral images. In recent years, deep learning-based methods have been garnering increasing attention in hyperspectral unmixing; especially, unsupervised autoencoder (AE) networks that have achieved excellent unmixing performance are a recent trend. While most approaches use spatial information, it is well known that hyperspectral data are characterized by a large number of narrow spectral bands. In order to take full advantage of the hyperspectral bands in unmixing and the spatial information, in this article, we explore multiview spectral and spatial information in an AE-based unmixing framework. We introduce multiview spectral information through spectral partitioning and propose a multiview spatial–spectral two-stream network, called MSSS-Net, which simultaneously learns a spatial stream network and a multiview spectral stream network in an end-to-end fashion for more efficient unmixing. The MSSS-Net is a two-stream deep unmixing network sharing a decoder, where its two AE networks employ recurrent neural networks (RNNs) to collaboratively utilize multiview spectral and spatial information. The spatial stream network branch extracts the spatial features of pixels and its neighbors, while the multiview spectral stream network branch exploits the multiview spectral bands of a pixel. Meanwhile, we design a cascaded bidirectional and unidirectional RNNs’ encoder structure for multiview spatial–spectral information to learn more discriminative deep patch-pixel features. Extensive ablation studies and experiments on both synthetic and real datasets demonstrate the superiority of the MSSS-Net over state-of-the-art unmixing methods.
Lin Qi 0004, Feng Gao 0005, Junyu Dong, Xinbo Gao 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Hyperspectral Image Denoising Based on Parallel Cross-Fusion Network
abstract
Hyperspectral images (HSIs) are widely used in agriculture and environmental monitoring. However, due to the climate and weather conditions, the acquired images commonly contain noise or detail loss, which greatly affects the image interpretation. To solve the problem, we propose a HSI denoising network based on the combination of Transformer and convolutional neural network (CNN), which can effectively improve the complex noise while ensuring image quality by using the powerful local analysis ability of CNN and the effective global feature interaction of Transformer. Extensive experiments on HSI dataset show that the proposed method outperforms four closely related methods.
Zhuoran Gong, Feng Gao 0005, Junyu Dong, Lin Qi 0004
IGARSS4
2022 Change Detection in Sar Images Based on A Multi-Scale Attention Convolution Network
abstract
Self-attention mechanism has been widely used in deep neural networks to improve the feature representation. In this paper, the self-attention mechanism is employed to solve the SAR change detection task. Specifically, we propose a Multi-scale Attention Convolution Network, MACNet for short, which exploits the multiscale features by a linear attention weight module. To enrich the feature representation space, we extract the spatial information of feature maps in a multi-branch way. Furthermore, we design a linear attention weight module, which emphasizes the important channels adaptively and fuses the contextual information from different scales. Experiments are conducted on three SAR datasets, and the results demonstrate the superior performance of the proposed method.
Feng Gao 0005, Junyu Dong, Lin Qi 0004
IGARSS4
2022 Multi-Scale Feature Fusion for Hyperspectral and Lidar Data Joint Classification
abstract
To exploit the multi-scale information to improve the feature representation, we propose a multi-scale feature fusion network for hyperspectral image (HSI) and LiDAR data joint classification. The model is comprised of three parts: feature extraction module, feature fusion module, and MSF(Multi-Scale Fusion) module. The feature fusion module integrates the multi-modal information into the two inputs by fusing attention. MSF module integrates richer semantic information by integrating multi-scale information to improve performance. Experimental results show that the proposed method is effective in multi-modal data joint classification.
Maqun Zhang, Feng Gao 0005, Junyu Dong, Lin Qi 0004
IGARSS4
2022 NormAttention-PSN: A High-frequency Region Enhanced Photometric Stereo Network with Normalized Attention
Yakun Ju, Boxin Shi, Muwei Jian, Lin Qi 0004, Junyu Dong, Kin-Man Lam 0001
Int. J. Comput. Vis.4
2022 Near-field photometric stereo using a ring-light imaging device
Hao Fan 0004, Yuan Rao 0001, Eric Rigall, Lin Qi 0004, Zhile Wang, Junyu Dong
Signal Process. Image Commun.4
2022 Wallpaper Texture Generation and Style Transfer Based on Multi-Label Semantics
abstract
Textures contain a wealth of image information and are widely used in various fields such as computer graphics and computer vision. With the development of machine learning, the texture synthesis and generation have been greatly improved. As a very common element in everyday life, wallpapers contain a wealth of texture information, making it difficult to annotate with a simple single label. Moreover, wallpaper designers spend significant time to create different styles of wallpaper. For this purpose, this paper proposes to describe wallpaper texture images by using multi-label semantics. Based on these labels and generative adversarial networks, we present a framework for perception driven wallpaper texture generation and style transfer. In this framework, a perceptual model is trained to recognize whether the wallpapers produced by the generator network are sufficiently realistic and have the attribute designated by given perceptual description; these multi-label semantic attributes are treated as condition variables to generate wallpaper images. The generated wallpaper images can be converted to those with well-known artist styles using CycleGAN. Finally, using the aesthetic evaluation method, the generated wallpaper images are quantitatively measured. The experimental results demonstrate that the proposed method can generate wallpaper textures conforming to human aesthetics and have artistic characteristics.
Ying Gao 0005, Xiaohan Feng, Tiange Zhang, Eric Rigall, Huiyu Zhou 0001, Lin Qi 0004, Junyu Dong
IEEE Trans. Circuits Syst. Video Technol.6
2022 Multimodal Gait Recognition for Neurodegenerative Diseases
abstract
In recent years, single modality-based gait recognition has been extensively explored in the analysis of medical images or other sensory data, and it is recognized that each of the established approaches has different strengths and weaknesses. As an important motor symptom, gait disturbance is usually used for diagnosis and evaluation of diseases; moreover, the use of multimodality analysis of the patient's walking pattern compensates for the one-sidedness of single modality gait recognition methods that only learn gait changes in a single measurement dimension. The fusion of multiple measurement resources has demonstrated promising performance in the identification of gait patterns associated with individual diseases. In this article, as a useful tool, we propose a novel hybrid model to learn the gait differences between three neurodegenerative diseases, between patients with different severity levels of Parkinson's disease, and between healthy individuals and patients, by fusing and aggregating data from multiple sensors. A spatial feature extractor (SFE) is applied to generating representative features of images or signals. In order to capture temporal information from the two modality data, a new correlative memory neural network (CorrMNN) architecture is designed for extracting temporal features. Afterward, we embed a multiswitch discriminator to associate the observations with individual state estimations. Compared with several state-of-the-art techniques, our proposed framework shows more accurate classification results.
Aite Zhao, Junyu Dong, Lin Qi 0004, Qianni Zhang, Ning Li 0012, Xin Wang 0068, Huiyu Zhou 0001
IEEE Trans. Cybern.4
2022 SSCU-Net: Spatial-Spectral Collaborative Unmixing Network for Hyperspectral Images
abstract
Linear spectral unmixing is an essential technique in hyperspectral image (HSI) processing and interpretation. In recent years, deep learning-based approaches have shown great promise in hyperspectral unmixing (HU), in particular, unsupervised unmixing methods based on autoencoder (AE) networks are a recent trend. The AE model, which automatically learns low-dimensional representations (abundances) and reconstructs data with their corresponding bases (endmembers), has achieved superior performance in HU. In this article, we explore the effective utilization of spatial and spectral information in AE-based unmixing networks. Important findings on the use of spatial and spectral information in the AE framework are discussed. Inspired by these findings, we propose a spatial–spectral collaborative unmixing network, called SSCU-Net, which learns a spatial AE network and a spectral AE network in an end-to-end manner to more effectively improve the unmixing performance. SSCU-Net is a two-stream deep network and shares an alternating architecture, where the two AE networks are efficiently trained in a collaborative way for estimation of endmembers and abundances. Meanwhile, we propose a new spatial AE network by introducing a superpixel segmentation method based on abundance information, which greatly facilitates the employment of spatial information and improves the accuracy of unmixing network. Moreover, extensive ablation studies are carried out to investigate the performance gain of SSCU-Net. Experimental results on both synthetic and real hyperspectral datasets illustrate the effectiveness and competitiveness of the proposed SSCU-Net compared with several state-of-the-art HU methods.
Lin Qi 0004, Feng Gao 0005, Junyu Dong, Xinbo Gao 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Associated Spatio-Temporal Capsule Network for Gait Recognition
abstract
It is a challenging task to identify a person based on her/his gait patterns. State-of-the-art approaches rely on the analysis of temporal or spatial characteristics of gait, and gait recognition is usually performed on single modality data (such as images, skeleton joint coordinates, or force signals). Evidence has shown that using multi-modality data is more conducive to gait research. Therefore, we here establish an automated learning system, with an associated spatio-temporal capsule network (ASTCapsNet) trained on multi-sensor datasets, to analyze multimodal information for gait recognition. Specifically, we first design a low-level feature extractor and a high-level feature extractor for spatio-temporal feature extraction of gait with a novel recurrent memory unit and a relationship layer. Subsequently, a Bayesian model is employed for the decision-making of class labels. Extensive experiments on several public datasets (normal and abnormal gait) validate the effectiveness of the proposed ASTCapsNet, compared against several state-of-the-art methods.
Aite Zhao, Junyu Dong, Lin Qi 0004, Huiyu Zhou 0001
IEEE Trans. Multim.4
2021 Remote Sensing Imagery Scene Classification Based on Spiking Neural Network
abstract
In order to overcome the large computational cost of deep neural networks (DNNs), spiking neural networks (SNNs) have been proposed, which are more biologically reasonable. It has the potential to achieve energy efficiency while maintaining performance comparable to DNNs. Although SNNs have achieved good results on the MNIST and CIFAR10 data sets, its potential in remote sensing has not been studied and explored. This paper adopts the idea of converting a trained DNN into an SNN, and proposes a multi-bit-based SNN, and introduces channel normalization (channel-norm) to replace the previous layer normalization to achieve remote sensing imagery scene classification tasks. By achieving multi-bit spiking and channel-norm in the way of DNN conversion to SNN, the SNN proposed in this paper achieves lossless conversion on the UC Merced data set and WHU-RS data set.
Saifei Wu, Jie Li 0001, Lin Qi 0004, Xinbo Gao 0001
IGARSS3
2021 Graph-Based CNNs With Self-Supervised Module for 3D Hand Pose Estimation From Monocular RGB
abstract
Hand pose estimation in 3D space from a single RGB image is a highly challenging problem due to self-geometric ambiguities, diverse texture, viewpoints, and self-occlusions. Existing work proves that a network structure with multi-scale resolution subnets, fused in parallel can more effectively shows the spatial accuracy of 2D pose estimation. Nevertheless, the features extracted by traditional convolutional neural networks cannot efficiently express the unique topological structure of hand key points based on discrete and correlated properties. Some applications of hand pose estimation based on traditional convolutional neural networks have demonstrated that the structural similarity between the graph and hand key points can improve the accuracy of the 3D hand pose regression. In this paper, we design and implement an end-to-end network for predicting 3D hand pose from a single RGB image. We first extract multiple feature maps from different resolutions and make parallel feature fusion, and then model a graph-based convolutional neural network module to predict the initial 3D hand key points. Next, we use 2D spatial relationships and 3D geometric knowledge to build a self-supervised module to eliminate domain gaps between 2D and 3D space. Finally, the final 3D hand pose is calculated by averaging the 3D hand poses from the GCN output and the self-supervised module output. We evaluate the proposed method on two challenging benchmark datasets for 3D hand pose estimation. Experimental results show the effectiveness of our proposed method that achieves state-of-the-art performance on the benchmark datasets.
Shaoxiang Guo, Eric Rigall, Lin Qi 0004, Xinghui Dong, Junyu Dong
IEEE Trans. Circuits Syst. Video Technol.3
2020 Arbitrary Style Transfer with Parallel Self-Attention
abstract
Neural style transfer aims to create artistic images by synthesizing patterns from a given style image. Recently, the Adaptive Instance Normalization (AdaIN) layer is proposed to achieve real-time arbitrary style transfer. However, we observed that if crucial features based on AdaIN can be further emphasized during transfer, both content and style information will be better reflected in stylized images. Furthermore, it is always essential to preserve more details and reduce unexpected artifacts in order to generate appealing results. In this paper, we introduce an improved arbitrary style transfer method based on the self-attention mechanism. A self-attention module is designed to learn what and where to emphasize in the input image. In addition, an extra Laplacian loss is applied to preserve structure details of the content while eliminating artifacts. Experimental results demonstrate that the proposed method outperforms AdaIN and can generate more appealing results.
Tiange Zhang, Ying Gao 0005, Feng Gao 0005, Lin Qi 0004, Junyu Dong
ICPR4
2020 Spatial-Spectral Autoencoder Networks for Hyperspectral Unmixing
abstract
We present a spatial-spectral autoencoder (SSAE) for hyperspectral unmixing, including a net for endmember extraction (EENet) and a net for abundance estimation (AENet). The EENet exploits the spatial information in hyperspectral image by a “many to one” strategy, i.e., the abundance of a pixel is combined by the abundances of its adjacent pixels. The idea is based on the assumption: once an endmember is mixed in a pixel, it is mixed in the surrounding pixels with high probability. The strategy promotes a continuous and smooth spatial distribution of abundances, and it is more effective than the other methods for endmember extraction. Besides, to make full use of the rich spectral information and obtain more accurate abundances, we design an AENet, which applies the deep convolutional neural network to estimate the abundances with the endmembers acquired from the EENet. The experiments are conducted on two real datasets, which show the SSAE outperforms the state-of-the-art methods.
Yongfa Huang, Jie Li 0001, Lin Qi 0004, Ying Wang 0007, Xinbo Gao 0001
IGARSS3
2020 Hyperspectral Unmixing via Recurrent Neural Network With Chain Classifier
abstract
Recently, the development of deep learning brings new opportunities for hyperspectral unmixing. In this paper, we propose a new architecture based on Recurrent Neural Networks and Bidirectional-LSTM (BiLSTM) that consists of two blocks: the feature extraction stage based on BiLSTM and the abundance estimation stage via Chain Classifier. BiLSTM can capture the long-distance relation among spectral bands better and Chain Classifier is more suitable to deal with multi-label task that exists in abundance estimation process. We evaluate the proposed method on two real HSI data sets including Jasper Ridge and Urban compared with three state-of-the-art approaches, and the results show significant improvements in accuracy of unmixing.
Mingyu Lei, Jie Li 0001, Lin Qi 0004, Ying Wang 0007, Xinbo Gao 0001
IGARSS3
2020 Pay Attention to Devils: A Photometric Stereo Network for Better Details
abstract
We present an attention-weighted loss in a photometric stereo neural network to improve 3D surface recovery accuracy in complex-structured areas, such as edges and crinkles, where existing learning-based methods often failed. Instead of using a uniform penalty for all pixels, our method employs the attention-weighted loss learned in a self-supervise manner for each pixel, avoiding blurry reconstruction result in such difficult regions. The network first estimates a surface normal map and an adaptive attention map, and then the latter is used to calculate a pixel-wise attention-weighted loss that focuses on complex regions. In these regions, the attention-weighted loss applies higher weights of the detail-preserving gradient loss to produce clear surface reconstructions. Experiments on real datasets show that our approach significantly outperforms traditional photometric stereo algorithms and state-of-the-art learning-based methods.
Yakun Ju, Kin-Man Lam 0001, Yang Chen 0036, Lin Qi 0004, Junyu Dong
IJCAI4
2020 MPS-Net: Learning to recover surface normal for multispectral photometric stereo
Yakun Ju, Lin Qi 0004, Jichao He, Xinghui Dong, Feng Gao 0005, Junyu Dong
Neurocomputing2
2020 A joint guidance-enhanced perceptual encoder and atrous separable pyramid-convolutions for image inpainting
Yongle Zhang 0001, Yingyu Wang, Junyu Dong, Lin Qi 0004, Hao Fan 0004, Xinghui Dong, Muwei Jian, Hui Yu 0001
Neurocomputing4
2020 Hybrid regression and isophote curvature for accurate eye center localization
abstract
Abstract The eye center localization is a crucial requirement for various human-computer interaction applications such as eye gaze estimation and eye tracking. However, although significant progress has been made in the field of eye center localization in recent years, it is still very challenging for tasks under the significant variability situations caused by different illumination, shape, color and viewing angles. In this paper, we propose a hybrid regression and isophote curvature for accurate eye center localization under low resolution. The proposed method first applies the regression method, which is called Supervised Descent Method (SDM), to obtain the rough location of eye region and eye centers. SDM is robust against the appearance variations in the eye region. To make the center points more accurate, isophote curvature method is employed on the obtained eye region to obtain several candidate points of eye center. Finally, the proposed method selects several estimated eye center locations from the isophote curvature method and SDM as our candidates and a SDM-based means of gradient method further refine the candidate points. Therefore, we combine regression and isophote curvature method to achieve robustness and accuracy. In the experiment, we have extensively evaluated the proposed method on the two public databases which are very challenging and realistic for eye center localization and compared our method with existing state-of-the-art methods. The results of the experiment confirm that the proposed method outperforms the state-of-the-art methods with a significant improvement in accuracy and robustness and has less computational complexity.
Jianwen Lou, Junyu Dong, Lin Qi 0004, Gongfa Li, Hui Yu 0001
Multim. Tools Appl.4
2020 A dual-cue network for multispectral photometric stereo
Yakun Ju, Xinghui Dong, Yingyu Wang, Lin Qi 0004, Junyu Dong
Pattern Recognit.4
2020 Deep spectral convolution network for hyperspectral image unmixing with spectral library
Lin Qi 0004, Jie Li 0001, Ying Wang 0007, Mingyu Lei, Xinbo Gao 0001
Signal Process.1
2020 A Perception-Inspired Deep Learning Framework for Predicting Perceptual Texture Similarity
abstract
Similarity learning plays a fundamental role in the fields of multimedia retrieval and pattern recognition. Prediction of perceptual similarity is a challenging task as in most cases we lack human labeled ground-truth data and robust models to mimic human visual perception. Although in the literature, some studies have been dedicated to similarity learning, they mainly focus on the evaluation of whether or not two images are similar, rather than prediction of perceptual similarity which is consistent with human perception. Inspired by the human visual perception mechanism, we here propose a novel framework in order to predict perceptual similarity between two texture images. Our proposed framework is built on the top of Convolutional Neural Networks (CNNs). The proposed framework considers both powerful features and perceptual characteristics of contours extracted from the images. The similarity value is computed by aggregating resemblances between the corresponding convolutional layer activations of the two texture maps. Experimental results show that the predicted similarity values are consistent with the human-perceived similarity data.
Ying Gao 0005, Yanhai Gan, Lin Qi 0004, Huiyu Zhou 0001, Xinghui Dong, Junyu Dong
IEEE Trans. Circuits Syst. Video Technol.3
2020 Spectral-Spatial-Weighted Multiview Collaborative Sparse Unmixing for Hyperspectral Images
abstract
Spectral unmixing is an important task in hyperspectral image (HSI) analysis and processing. Sparse representation has become a promising semisupervised method for remotely sensed hyperspectral unmixing and incorporating the spectral or spatial information to improve the spectral unmixing results under a weighted sparse unmixing framework is a recent trend. While most methods focus on analyzing HSI by exploring the spatial information, it is known that hyperspectral data are characterized by its large contiguous set of wavelengths. This information can be naturally used to improve the representation of pixels in HSI. In order to take the advantage of the hyper spectral information as well as the spatial information for hyperspectral unmixing, in this article, we explore and introduce a multiview data processing approach through spectral partitioning to benefit from the abundant spectral information in HSI. Some important findings on the application of multiview data set in sparse unmixing are discussed. Meanwhile, we develop a new spectral–spatial-weighted multiview collaborative sparse unmixing (MCSU) model to tackle such a multiview data set. The MCSU uses a weighted sparse regularizer, which includes both multiview spectral and spatial weighting factors to further impose sparsity on the fractional abundances. The weights are adaptively updated associated with the abundances, and the proposed MCSU can be solved by the alternating direction method of multipliers efficiently. The experimental results on both the simulated and real hyperspectral data sets demonstrate the effectiveness of the proposed MCSU, which can significantly improve the abundance estimation results.
Lin Qi 0004, Jie Li 0001, Ying Wang 0007, Yongfa Huang, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2019 A novel joint dictionary framework for sparse hyperspectral unmixing incorporating spectral library
Lin Qi 0004, Jie Li 0001, Xinbo Gao 0001, Ying Wang 0007, Chongyue Zhao, Yu Zheng 0006
Neurocomputing1
2019 A procedural texture generation framework based on semantic descriptions
Junyu Dong, Jun Liu 0055, Ying Gao 0005, Lin Qi 0004, Xin Sun 0003
Knowl. Based Syst.5
2019 Region-Based Multiview Sparse Hyperspectral Unmixing Incorporating Spectral Library
abstract
Hyperspectral image (HSI) is characterized by its huge contiguous set of wavelengths. It is possible and needed to benefit from the “hyper” spectral information as well as the spatial information. For this purpose, we propose a new multiview data generation approach that takes full advantage of the rich spectral and spatial information in HSI, by dividing the original HSI into several spatially homogeneous regions with different band margins. Then, a new sparse unmixing algorithm, called region-based multiview sparse unmixing (RMSU), is presented to tackle such a multiview data model in this letter. The RMSU algorithm combines the multiview learning anda prioriinformation to improve the performance of sparse unmixing by incorporating the multiview information and spectral library into the dictionary learning framework. We also show that RMSU can serve as a dictionary pruning algorithm, which provides a possibility that unmixing algorithms could have higher accuracy and efficiency. Experimental results on both simulated and real hyperspectral data demonstrate the effectiveness of the proposed RMSU algorithm both visually and quantitatively.
Lin Qi 0004, Jie Li 0001, Ying Wang 0007, Xinbo Gao 0001
IEEE Geosci. Remote. Sens. Lett.1
2018 Predicting and Generating Wallpaper Texture with Semantic Properties
abstract
Humans naturally use semantic descriptions to express their visual perception of textures; this is also the fact for perception and description of wallpaper texture. Classification of wallpaper's style is mainly based on understanding of visual information. However, the complexity of real-world wallpaper images is difficult to be captured by existing datasets. Inspired by a publicly available Procedural Textures Dataset, a number of wallpaper images was collected and assembled into a wallpaper dataset. A series of psychophysical experiments was performed to further collect semantic descriptions for this dataset. Each wallpaper was labeled with 5-10 semantic descriptions. More importantly, our dataset contains complex wallpaper images with rich annotations. To our best knowledge, our dataset is the first public wallpaper dataset with semantic descriptions. We use label distribution to analysis semantic descriptions and texture characteristics. Furthermore, a texture generation method based on GAN was tested using our wallpaper dataset, which produced state-of-the-art results.
Xiaohan Feng, Lin Qi 0004, Yanhai Gan, Ying Gao 0005, Hui Yu 0001, Junyu Dong
HSI2
2018 Dynamic 3D Surface Reconstruction Using a Hand-Held Camera
abstract
This paper proposes a dynamic 3D reconstruction method for recovering a surface shape from a set of images that are captured by a hand-held camera. A light source is attached to the camera as a photometric constraint. Thus, we can effectively calculate photometric stereo using the relative moving camera. The key contributions of our work are a robust pixel matching method to build effective correspondences between images for normal estimation, and an optimization method to correct the deviation in the recovered surface shape that is caused by the nonideal illumination in a close-range lighting condition. Specially we correct the recovered shape by adding an interpolation surface that is estimated using sparse control points from the structure from motion. The effectiveness of our method is verified on real datasets with a digital camera and a smart phone.
Hao Fan 0004, Lin Qi 0004, Junyu Dong, Gongfa Li, Hui Yu 0001
IECON2
2018 Perception-driven procedural texture generation from examples
Jun Liu 0055, Yanhai Gan, Junyu Dong, Lin Qi 0004, Xin Sun 0003, Muwei Jian, Hui Yu 0001
Neurocomputing4
2018 A hybrid spatio-temporal model for detection and severity rating of Parkinson's disease from gait data
Aite Zhao, Lin Qi 0004, Jie Li 0001, Junyu Dong, Hui Yu 0001
Neurocomputing2
2018 Dual channel LSTM based multi-feature extraction in gait for diagnosis of Neurodegenerative diseases
Aite Zhao, Lin Qi 0004, Junyu Dong, Hui Yu 0001
Knowl. Based Syst.2
2016 Robust Photometric Stereo in a scattering medium via Low-Rank Matrix Completion and Recovery
abstract
Photometric Stereo is a popular method for 3D reconstruction from images due to its high level of details handling. However, when it is used in a scattering medium such as lakes and oceans, the recovery result will be negatively impacted by the light absorption, light scattering and the impurities in the water. In this paper, we present a new method to solve the problem of better 3D reconstruction via Low-Rank Matrix Completion and Recovery. First, we use the dark points, like shadows and darkness in the water to fit the scattering effect distribution and then remove the scattering from the image. Next, we use the Robust Principal Component Analysis method (RPCA) to recover the image by removing the sparse noise including shadows, impurities and some corrupted points caused by backscatter compensation. Finally, we combine the RPCA results and the least-squares (LS) results to get the surface normal and accomplish the 3D reconstruction. Extensive experimental results demonstrate that our method achieves more accurate estimates of surface normal and 3D reconstruction than previous techniques.
Hao Fan 0004, Yisong Luo, Lin Qi 0004, Junyu Dong, Hui Yu 0001
HSI3
2016 Learning perceptual texture similarity and relative attributes from computational features
abstract
Previous work has shown that perceptual texture similarity and relative attributes cannot be well described by computational features. In this paper, we propose to predict human's visual perception of texture images by learning a non-linear mapping from computational feature space to perceptual space. Hand-crafted features and deep features, which were successfully applied in texture classification tasks, were extracted and used to train Random Forest and rankSVM models against perceptual data from psychophysical experiments. Three texture datasets were used to test our proposed method and the experiments show that the predictions of such learnt models are in high correlation with human's results.
Jianwen Lou, Lin Qi 0004, Junyu Dong, Hui Yu 0001, Guoqiang Zhong 0001
IJCNN2
2015 Toward a psychophysical-based procedural texture generation system for interactive design
abstract
Procedural textures have been widely used as they can be easily generated from various mathematical models. However, the model parameters are not perceptually meaningful or uniform for non-expert users. In this paper, we proposed a system that can generate procedural textures interactively along certain perceptual dimensions. We built a procedural texture dataset and measured twelve perceptual properties of a small subset through psychophysical experiments. The perceived magnitude of the rest textures was estimated by Support Vector Machines using computational features from a cascaded PCA network. For a given texture displayed on a touch screen, the user makes finger gestures which were then transferred to magnitude changes in perceptual space. The texture in the database that matches the new perceptual scale and with nearest distance in computational feature space will be chosen and displayed. We reported our experiment results for two particular perceptual properties: surface roughness and directionality. Other properties can be manipulated similarly.
Xiaoxu Cai, Jun Liu 0055, Lin Qi 0004, Junyu Dong, Ying Gao 0005, Hui Yu 0001
HSI3
2015 Combining Kinect and PnP for camera pose estimation
abstract
This paper presents a novel method to conduct camera pose estimation though combining Kinect and Perspective-n-points algorithms. Most existing camera pose estimation methods suffer from the errors caused by inevitable outliers between 2D-3D correspondences. To this end, we propose to use a random down sampling process to deal with outliers in this paper. The proposed method is divided into two main steps, which are 2D-3D correspondences generation and pose estimation. The method has been tested in a real project, and the experiment has shown encouraging results compared to the ground truth.
Shu Zhang 0002, Hui Yu 0001, Junyu Dong, Ting Wang 0018, Lin Qi 0004, Honghai Liu 0001
HSI5
2015 Fast 3D face reconstruction based on uncalibrated photometric stereo
Yujuan Sun, Junyu Dong, Muwei Jian, Lin Qi 0004
Multim. Tools Appl.4
2008 Texture synthesis by Support Vector Machines
abstract
We introduce a simple texture synthesis method based on Support Vector Machines (SVM). Although SVM has been effectively used for various pattern recognition tasks, there is no report available on directly applying SVM for texture synthesis. The advantage of using SVM is that the sample can be simply modeled by a linear model and is not required during the synthesis stage. In addition, the method can be further extended to synthesize 3D surface texture or Bidirectional Texture Functions. Our experimental results show that the method can successfully model and synthesize semi or highly structured textures, which can be difficult subjects for previous texture synthesis methods based on parametric models.
Junyu Dong, Yuanxu Duan, Guimei Sun, Lin Qi 0004
ICPR4
2008 Directionality measurement and illumination estimation of 3D surface textures by using mojette transform
abstract
This paper presents a new approach to measure texture directions and estimate illumination tilt angle of 3D surface textures by using mojette transform. Feature vectors are generated from variances of 72 mojette transform projections with different projection angles. The measured texture directions are compared with human perceptual judgement. Furthermore, we estimate illumination tilt angles by minimizing the Euclidean distance of the feature vector between the test image and the training sets. Experimental results show the effectiveness and accuracy of our proposed approach.
Junyu Dong, Lin Qi 0004, Florent Autrusseau
ICPR3