EDBT 2026 Demo / reviewers in the wild / expert
Yanshan Li
dblp:64/599
· DBLP profile ↗
41ranked-venue papers
20as first author
27since 2021 · last 2027
0000-0002-8814-4628ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 11 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Mgrad-CAM: Class activation map based on multi-label gradient feedback
Yanshan Li, Xinhua Mo, Hongfang Zheng, Jianlin Xiang, Linhui Dai |
Expert Syst. Appl. | 1 |
| 2026 | Physics informed Dual-Layer Bidirectional Gated Recurrent Unit for Nuclear-Grade Electric Gate Valves Fault Prognostics
Chenwei Tang, Jiancheng Lv 0001, Yanping Huang, Yanshan Li |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Fast and Effective Video Inpainting via Implicit Motion-Guided Propagation and Sparse AttentionabstractVideo inpainting aims to reconstruct missing or corrupted regions in video frames, with applications in video editing, restoration, and special effects. Current deep video inpainting methods rely on optical flow to guide the propagation of effective features and spatiotemporal attention mechanisms to model relationships between frames. However, as an explicit motion representation, the optical flow extracted offline in preceding steps often suffers from instability and errors during estimation. These errors accumulate during subsequent content hallucination, resulting in artifacts and blurring. Meanwhile, although traditional spatiotemporal attention effectively captures frame relationships, its dense computational nature introduces redundant information, disrupting inpainting tasks and reducing efficiency. To address these issues, we propose an implicit motion-guided approach for efficient video inpainting. Instead of relying on optical flow, our method uses implicit motion in the latent feature space to guide the dual-domain propagation of images and features end-to-end, avoiding error accumulation from the independent optical flow estimation process. Additionally, we introduce a self-correcting module that enables feedback between image and feature propagation, reducing errors during propagation. Furthermore, we design an adaptive sparse video attention mechanism to focus on highly relevant regions, minimizing the impact of irrelevant information. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches both qualitatively and quantitatively, while also delivering superior efficiency. Yuanman Li, Bin Li 0011, Yanshan Li, Jiantao Zhou 0001, Xia Li 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | $\ell _{0}$-RASC-NN: An Efficient Spatiotemporal Data Completion Method for Edge DevicesabstractWith the widespread deployment of large models on edge devices, more and more SpatioTemporal (ST) data processing tasks need to be performed locally on these devices. However, edge devices often face challenges, such as limited computational resources, stringent real-time and robustness requirements, scarcity of labeled data, and difficulties with manual intervention. To address these challenges, we propose a lightweight, fast, and robust method, named$\ell _{0}$-norm Rank-Adaptive Spatiotemporal data Completion Neural Network ($\ell _{0}$-RASC-NN). We first formulate the target as an optimization model and employ Block Coordinate Descent (BCD) combined with alternating optimization for iterative solving. This iterative process is then transformed into an artificial neural network using a deep unrolling algorithm. This approach not only reduces computational cost and improves real-time performance but also better addresses the challenge of limited labeled data compared to conventional deep learning methods. To enhance robustness, the model incorporates an$\ell _{0}$-norm regularization term and a dynamic threshold adjustment strategy to handle anomalies. Furthermore, an adaptive rank selection mechanism and hyperparameter reparameterization are introduced to minimize the need for manual intervention. Experiments on five real-world ST datasets demonstrate that our method outperforms other state-of-the-art approaches in both ST data recovery and anomaly detection. Notably, under the same conditions,$\ell _{0}$-RASC-NN reduces computational time by one to two orders of magnitude compared to existing methods, while maintaining or even enhancing recovery accuracy. The code of our proposed method is provided athttps://tinyurl.com/36crh4v9. Hao Wang 0075, Chenyu Guan, Linfang Yu, Lishuai Li, Yanshan Li, Lei Gong 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | MARSNet: Scalable Deep Coding of LiDAR Point Clouds via Multimodal and Residual Learning
Yanji Huang, Runnan Huang, Jianlong Zhou, Yingqi Zhuo, Yanshan Li, Miaohui Wang |
ICIG (1) | 5 |
| 2025 | OMA-SSR: Optical-guided multi-kernel attention based SAR image super-resolution reconstruction networkabstractAbstract Synthetic aperture radar (SAR) has been widely studied and applied in many fields. Although image super‐resolution technology has been successfully applied to SAR imaging in recent years, there is less research on large‐scale factor SAR image super‐resolution methods. A more effective method is to obtain comprehensive information to guide the reconstruction of SAR images. In fact, the co‐registered characteristics of high‐resolution optical images have been successfully applied to improve the quality of SAR images. Inspired by this, an optical‐guided multi‐kernel attention based SAR image super‐resolution reconstruction network (OMA‐SSR) is proposed. The proposed multi‐modal mutual attention (MMA) module in this network can effectively establish the dependency between SAR image features and optical image features. This network also designs a deep feature extraction module for SAR images, which includes a channel‐splitted multi‐kernel attention (CSMA) module and residual connections. CSMA module splits SAR image channels, extracts features in different ranges through multi‐kernel convolution, and finally fuses the extracted features between different channels. Experimental results on the Sen1‐2 and QXS datasets show that the proposed OMA‐SSR performs well in evaluation indicators and visual effects of SAR image super‐resolution reconstruction. Yanshan Li |
IET Image Process. | 1 |
| 2025 | Small object detection network based on progressive enhanced multi-level feature fusion
Yanshan Li, Fuxing Liu, Yusong Qin, Linhui Dai, Weixin Xie |
Neurocomputing | 1 |
| 2025 | STD-Explain: Generalizing explanations for spatio-temporal graph convolutional networks based on spatio-temporal decoupled perturbation
Yanshan Li, Suixuan He, Rui Yu 0004, Weixin Xie |
Neurocomputing | 1 |
| 2025 | WB-LRP: Layer-wise relevance propagation with weight-dependent baseline
Yanshan Li, Huajie Liang, Lirong Zheng 0003 |
Pattern Recognit. | 1 |
| 2025 | suLPCC: A Novel LiDAR Point Cloud Compression Framework for Scene Understanding TasksabstractLight detection and ranging (LiDAR) point cloud compression (LPCC) plays an important role in managing the storage, transmission, and perception of the rapidly expanding volume of LiDAR point cloud (LPC) data. However, there has been a noticeable lack of comprehensive investigation into LPCC methods specifically designed for environmental perception and understanding. To address this gap, we propose a new LPCC framework aimed at meeting the unique requirements of various scene understanding tasks, enhancing the adaptability of LPCCs in real-world scenarios. Specifically, we divide the input LPCs into an object and a scene component through a distinction module, design a new point completion-based method to encode object LPCs, and develop novel structure-aware intracoding and motion-optimized intercoding schemes to compress scene LPCs. Experimental results on three benchmark datasets demonstrate the effectiveness of our proposed method on the localization, mapping, and detection tasks. We believe that the findings presented in this article will contribute to a deeper understanding of LPCCs as well as promote further development of LiDAR sensor-based systems. Miaohui Wang, Runnan Huang, Ye Liu 0005, Yanshan Li, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Visual Redundancy Removal for Composite Images: A Benchmark Dataset and a Multi-Visual-Effects Driven Incremental MethodabstractComposite images (CIs) typically combine various elements from different scenes, views, and styles, which are a very important information carrier in the era of mixed media such as virtual reality, mixed reality, metaverse, etc. However, the complexity of CI content presents a significant challenge for subsequent visual perception modeling and compression. In addition, the lack of benchmark CI databases also hinders the use of recent advanced data-driven methods. To address these challenges, we first establish one of the earliest visual redundancy prediction (VRP) databases for CIs. Moreover, we propose a multi-visual effect (MVE)-driven incremental learning method that combines the strengths of hand-crafted and data-driven approaches to achieve more accurate VRP modeling. Specifically, we design special incremental rules to learn the visual knowledge flow of MVE. To effectively capture the associated features of MVE, we further develop a three-stage incremental learning approach for VRP based on an encoder-decoder network. Extensive experimental results validate the superiority of the proposed method in terms of subjective, objective, and compression experiments. Miaohui Wang, Lirong Huang, Yanshan Li |
AAAI | 4 |
| 2024 | A Fast Blind Deblurring Algorithm Using Local Gradient Product PriorabstractBlind image deblurring is the restoration of latent clear images from blurred images without knowing the blur kernel. Recently, a large number of priors have been proposed to effectively address the ill-posed nature of blind deblurring. However, most methods approximate the proposed priors and only use the first-order term of the blurred image. We observe the gradient inner product of clear images is significantly larger than that of blurry ones. In this paper, we propose a novel Local Gradient Product (LGP) prior based on this observation. This prior not only offers a more accurate approximation but also uses the relationship between image pixels involving quadratic terms. The employment of the LGP prior avoids the inversion of large matrices, thereby improving solution efficiency. Extensive experiments demonstrate that our algorithm achieves state-of-the-art performance and competitive running speed on benchmark datasets. Our MATLAB code and experimental results are available at github.com/JixuanLiang/deblur-LGP. Jixuan Liang, Yanshan Li |
ICASSP | 2 |
| 2024 | Spatiotemporal adaptive hybrid dynamic graph convolutional network for traffic flow predictionabstractTraffic flow prediction is a crucial research area that has been extensively studied using graph-based prediction methods. However, existing approaches often rely on static or dynamic graphs to model spatial dependencies, which may not capture diverse spatial correlations and dependencies arising from intricate traffic patterns. In this paper, we propose a novel Spatiotemporal Adaptive Hybrid Dynamic Graph Convolutional Network (STAHDGCN) to enhance traffic prediction accuracy. Specifically, we propose a hybrid spatial graph learning module designed to capture diverse spatial stability and contingency in the road network at different times. This module incorporates both a static adaptive learning module and a dynamic learning module. Following this, a spatial gate fusion module is employed to conduct feature fusion, effectively simulating the complex spatiotemporal dependence within road networks. Finally, a proposed adaptive spatiotemporal module utilizes an attention mechanism to effectively capture potential dependence patterns in both time and space, addressing the impact of spatial heterogeneity. Experimental evaluations on two public datasets, METRLA and PEMS-BAY, demonstrate the superior performance of our model. Yamin Wen, Bin Ren 0006, Yanshan Li, Yuming Huang 0008, Lianghong Wu |
IJCNN | 3 |
| 2024 | GeoExplainer: Interpreting Graph Convolutional Networks with geometric masking
Rui Yu 0004, Yanshan Li, Huajie Liang |
Neurocomputing | 2 |
| 2024 | Aircraft type recognition in 3D-view optical image with contour segmentation
Zhixiang Liang, Yanshan Li, Rui Yu 0004, Kaihao Zhang |
Multim. Tools Appl. | 2 |
| 2024 | GT-CAM: Game Theory Based Class Activation Map for GCNabstractGraph Convolutional Networks (GCN) have shown outstanding performance in skeleton-based behavior recognition. However, their opacity hampers further development. Researches on the explainability of deep learning have provided solutions to this issue, with Class Activation Map (CAM) algorithms being a class of explainable methods. However, existing CAM algorithms applies to GCN often independently compute the contribution of individual nodes, overlooking the interactions between nodes in the skeleton. Therefore, we propose a game theory based class activation map for GCN (GT-CAM). First, GT-CAM integrates Shapley values with gradient weights to calculate node importance, producing an activation map that highlights the critical role of nodes in decision-making. It also reveals the cooperative dynamics between nodes or local subgraphs for a more comprehensive explanation. Second, to reduce the computational burden of Shapley values, we propose a method for calculating Shapley values of node coalitions. Lastly, to evaluate the rationality of coalition partitioning, we propose a rationality evaluation method based on bipartite game interaction and cooperative game theory. Additionally, we introduce an efficient calculation method for the coalition rationality coefficient based on the Monte Carlo method. Experimental results demonstrate that GT-CAM outperforms other competitive interpretation methods in visualization and quantitative analysis. Yanshan Li, Weixin Xie |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | CR-CAM: Generating explanations for deep neural networks by contrasting and ranking features
Yanshan Li, Huajie Liang, Hongfang Zheng, Rui Yu 0004 |
Pattern Recognit. | 1 |
| 2024 | A Ladder Water Level Prediction Model for the Yangtze River Based on Transfer Learning and TransformerabstractWater level prediction is of great importance in alleviating the increasing water scarcity and preventing frequent floods. However, current water level prediction models do not consider the spatial and temporal features in water levels at monitoring stations. This study proposes a ladder water level prediction model for the Yangtze River based on transfer learning and Transformer to obtain more accurate predictions of water levels under tidal interactions. Our model utilizes the attention mechanism of the Transformer and incorporates spatial features and correlations of water level variations at the tidal limit of rivers. In addition, transfer learning is employed to explore and analyze the temporal characteristics of the wet season and dry season. Numerical experiments conducted at monitoring stations in the lower reach of the Yangtze River in China validate the effectiveness of our model. In 24-h water level prediction, our model achieves an average reduction of 52.14% in mean absolute error (MAE), 52.99% in root mean square error (RMSE), and an average increase of 5.70% in the pass rate within ±0.3 m compared to existing studies. Yanshan Li, Ya Zhang 0001, Xiang Liu 0019, Mingyan Xia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | BI-CAM: Generating Explanations for Deep Neural Networks Using Bipolar InformationabstractThe higher requirements for deep neural networks are driving researchers to have a deeper understanding of the internals of neural networks. The class activation map (CAM) based methods can provide a convincing interpretation of the features extracted by the neural network from both visual and quantitative perspectives. However, the existing CAM methods do not take into account that the non-target region also contains target-related activation, which results in the generated saliency map containing noise from unrelated regions. In addition, the soft mask with continuous value not only contains more non-target regions for gradient-free CAM, but also causes the characteristics and distribution of the target region to be disturbed. This paper proposed a novel CAM method named Bipolar Information CAM (BI-CAM) to interpret convolutional neural networks (CNNs) and graph convolutional networks (GCNs). Firstly, dual-stream information is proposed to precisely quantify the relationship between the target region and the non-target region for an image/graph. Secondly, binary reformation is also proposed to generate a hard mask that can retain the original features and regions. Finally, we propose to use concise and effective Point-wise Mutual Information (PMI) to measure the quantitative relationship between the image and the local region with respect to the label. The results of the experiment show that the proposed BI-CAM achieves significantly better performance in the faithfulness evaluation from the perspectives of visualization and quantitative analysis than other competitive interpretation methods. Yanshan Li, Huajie Liang, Rui Yu 0004 |
IEEE Trans. Multim. | 1 |
| 2023 | Time-sequential hesitant fuzzy entropy, cross-entropy and correlation coefficient and their application to decision making
Lingyu Meng, Weixin Xie, Yanshan Li |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Hesitant hierarchical T-S fuzzy system with fuzzily weighted recursive least square
Lingyu Meng, Weixin Xie, Yanshan Li |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Dual attention based spatial-temporal inference network for volleyball group activity recognition
Yanshan Li, Rui Yu 0004, Hailin Zong, Weixin Xie |
Multim. Tools Appl. | 1 |
| 2023 | T-Net: Deep Stacked Scale-Iteration Network for Image DehazingabstractHaze reduces the visibility of image content and leads to failure in handling subsequent computer vision tasks. In this paper, we address the problem of single image dehazing by proposing a dehazing network named T-Net, which consists of a backbone network based on the U-Net architecture and a dual attention module. Multi-scale feature fusion can be achieved by using skip connections with a new fusion strategy. Furthermore, by repeatedly unfolding the plain T-Net, Stack T-Net is proposed to take advantage of the dependence of deep features across stages via a recursive strategy. To reduce network parameters, the intra-stage recursive computation of ResNet is adopted in our Stack T-Net. We take both the stage-wise result and the original hazy image as input to each T-Net and finally output the prediction of the clean image. Experimental results on both synthetic and real-world images demonstrate that our plain T-Net and the advanced Stack T-Net perform favorably against state-of-the-art dehazing algorithms and show that our Stack T-Net could further improve the dehazing effect, demonstrating the effectiveness of the recursive strategy. Lirong Zheng 0003, Yanshan Li, Kaihao Zhang, Wenhan Luo |
IEEE Trans. Multim. | 2 |
| 2022 | A Multiscale Spatial-Spectral Prototypical Network for Hyperspectral Image Few-Shot ClassificationabstractDue to the complex environment of hyperspectral image (HSI) gathering area, it is difficult to obtain a large number of labeled samples for HSI. Therefore, how to effectively achieve the HSI few-shot classification is a hot spot of current research. Prototypical network (PN) is one of the most classical few-shot learning algorithms, which has been widely employed for few-shot image classification and few-shot object detection. However, existing PN-based algorithms for HSI only utilize the single-scale spatial-spectral feature extracted from the last layer, ignoring the semantic information with different scales contained in the other layers. To solve this problem, a novel multi-scale spatial-spectral prototypical network (MSSPN) is proposed in this letter. The contribution of this letter is threefold. Firstly, a multi-scale spatial-spectral feature extraction algorithm based on ladder structure is proposed to effectively achieve the integration of spatial-spectral features with different scales. Secondly, with the theory of ladder-structure-based extraction algorithm, we design a multi-scale spatial-spectral prototype representation, which is suggested to be more robust and effective in the multi-scale spatial-spectral metric space. Finally, our proposed MSSPN has the advantage of expandability, and can be easily applied for the other PN-based few-shot learning methods. The experimental results on HSI few-shot classification indicate that our proposed MSSPN algorithm can achieve higher accuracy than the representative HSI classifiers and the existing PN-based algorithms. Haojin Tang, Zhiquan Huang, Yanshan Li, Weixin Xie |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Multidimensional Local Binary Pattern for Hyperspectral Image ClassificationabstractFor the large amount of spatial and spectral information contained in hyperspectral image (HSI), feature description of HSI has attracted widespread concern in recent years. Existing deep learning-based HSI feature description algorithms require a large number of training samples and have poor interpretability. Therefore, it is necessary to develop an efficient HSI features description algorithm with interpretability based on machine learning. Local binary pattern (LBP) is a classical descriptor used to extract the local spatial texture features of images, which has been widely applied to image feature description and matching. However, the existing LBP algorithms for HSI are based on the single-dimensional description, which leads to the limitations on the expression of spatial–spectral information. Therefore, a multidimensional LBP (MDLBP) based on Clifford algebra for HSI is proposed in this article, which is able to extract spatial–spectral feature from multiple dimensions. First, with the theory of the Clifford algebra, a new representation of HSI including spatial and spectral information is built. Second, the geometric relationship between the local geometry of HSI in Clifford algebra space is calculated to realize the local multidimensional description of the local spatial–spectral information. Finally, a novel LBP coding algorithm for HSI is implemented based on the local multidimensional description to calculate the feature descriptor of HSI. The experimental results on HSI classification show that our proposed MDLBP algorithm can achieve higher accuracy than the representative spatial–spectral features and the existing LBP algorithms, especially in the scenery of small-scale training samples. Yanshan Li, Haojin Tang, Weixin Xie, Wenhan Luo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | LAGA-Net: Local-and-Global Attention Network for Skeleton Based Action RecognitionabstractSkeleton-based action recognition has attracted significant attention and obtained widespread applications due to the robustness of 3D skeleton data. One of the key challenges is how to extract discriminative and robust spatio-temporal features from sparse skeleton data to describe actions and improve recognition accuracy. To address this issue, this paper combines convolutions with attention mechanisms and proposes a deep network for skeleton-based action recognition, termed as local-and-global attention network (LAGA-Net). First, we encode skeleton sequences into joint feature evolution maps to compactly describe the spatial and temporal characteristics of skeleton sequences. Then, a motion guided channel attention module (MGCAM) is proposed to model the interdependencies between feature channels by calculating temporal frame-level motion and enhance motion-salient features in a channel-wise way. Further, a spatio-temporal attention module (STAM) is proposed to model spatio-temporal context-aware collaboration at sequence level and extract spatio-temporal attention features that involve long-range dependencies. Together, MGCAM and STAM are combined to form LAGA-Net, which extracts discriminative features integrating both local and global representations of skeleton sequences. Moreover, a two-stream architecture is proposed to learn complementary features from joint and bone aspects. We conduct extensive experiments to verify the effectiveness and superiority of our proposed method over state-of-the-art approaches on several benchmarks (e.g., NTU RGB+D, Northwestern-UCLA, UTD-MHAD and NTU RGB+D 120). Rongjie Xia, Yanshan Li, Wenhan Luo |
IEEE Trans. Multim. | 2 |
| 2021 | Adaptive multi-view graph convolutional networks for skeleton-based action recognition
Yanshan Li, Rongjie Xia |
Neurocomputing | 2 |
| 2020 | Relative view based holistic-separate representations for two-person interaction recognition using multiple graph convolutional networks
Yanshan Li, Tianyu Guo 0002, Rongjie Xia |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | A Spatial-Spectral Prototypical Network for Hyperspectral Remote Sensing ImageabstractHyperspectral remote sensing image (HRSI) can provide additional spectral information of objects and have been widely used in many fields. However, due to the complex environment of the HRSI gathering area, collecting the labeled samples of HRSI is time-consuming and labor-intensive. The scarcity of labeled samples is one of the major difficulties for HRSI analysis and processing. In this letter, a spatial-spectral prototypical network (SSPN) for HRSI is proposed for solving the problem of lack of labeled samples. The contribution of this letter is threefold. First, we design a novel local pattern coding algorithm to combine the spatial and spectral information of HRSI pixels based on spatial neighborhood correlation. Then, a spatial-spectral feature extraction algorithm based on 1-D convolutional neural network (1-D-CNN) is suggested to learn the spatial-spectral metric space where HRSI pixels can be correctly classified with only a few labeled samples. Finally, a novel prototype representation for HRSI in spatial-spectral metric space is proposed to better classify the mixed pixels existing in HRSI. The experimental results on three popular HRSI data sets demonstrate that the proposed SSPN is significantly better than the traditional algorithms. Haojin Tang, Yanshan Li, Qinghua Huang, Weixin Xie |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Learning shape and motion representations for view invariant skeleton-based action recognition
Yanshan Li, Rongjie Xia |
Pattern Recognit. | 1 |
| 2019 | Learning Shape-Motion Representations from Geometric Algebra Spatio-Temporal Model for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition has been widely applied in intelligent video surveillance and human behavior analysis. Previous works have successfully applied Convolutional Neural Networks (CNN) to learn spatio-temporal characteristics of the skeleton sequence. However, they merely focus on the coordinates of isolated joints, which ignore the spatial relationships between joints and only implicitly learn the motion representations. To solve these problems, we propose an effective method to learn comprehensive representations from skeleton sequences by using Geometric Algebra. Firstly, a frontal orientation based spatio-temporal model is constructed to represent the spatial configuration and temporal dynamics of skeleton sequences, which owns the robustness against view variations. Then the shape-motion representations which mutually compensate are learned to describe skeleton actions comprehensively. Finally, a multi-stream CNN model is applied to extract and fuse deep features from the complementary shape-motion representations. Experimental results on NTU RGB+D and Northwestern-UCLA datasets consistently verify the superiority of our method. Yanshan Li, Rongjie Xia, Qinghua Huang |
ICME | 1 |
| 2019 | Robust multi-view representation for spatial-spectral domain in application of hyperspectral image classificationabstractSpatial–spectral representation plays an important role in hyperspectral images (HSIs) classification. However, many of the existing local feature algorithms for HSIs are based on the two‐dimensional image and do not take full advantage of the information hidden in HSI, such as spatial–spectral locality correlation information, thereby reducing the robustness of these algorithms. In response to these problems, this study presents a robust multi‐view spatial–spectral representation method with the characteristics of HSIs. There are two key techniques in this representation method, called spatial–spectral locality constrained linear coding (SSLLC) and spatial–spectral pyramid matching model (SSPM). Firstly, SSLLC applies the locality information of the feature points and visual words and uses the discriminant information provided by the nearest‐neighbouring spatial–spectral feature points in HSIs. Secondly, SSPM works by partitioning the image into increasingly fine sub‐cubes and uses the cubes to match the local features of the HSIs. The multi‐view representation is tolerant to illumination change, image rotation, affine distortion etc. To assess the validity of authors' algorithm, the authors compared their results with several existing approaches, including a deep learning method. The experimental results show that this representation method can effectively improve the accuracy of HSIs classification. Yanshan Li, Xianchen Wang, Qinghua Huang, Weixin Xie |
IET Comput. Vis. | 1 |
| 2019 | A spatial-spectral SIFT for hyperspectral image matching and classification
Yanshan Li, Qingteng Li, Weixin Xie |
Pattern Recognit. Lett. | 1 |
| 2018 | Two-stage local constrained sparse coding for fine-grained visual categorization
Lihua Guo, Chenggang Guo, Qinghua Huang, Yanshan Li, Xuelong Li 0001 |
Sci. China Inf. Sci. | 5 |
| 2018 | Extreme-constrained spatial-spectral corner detector for image-level hyperspectral image classification
Yanshan Li, Jianjie Xu, Rongjie Xia, Qinghua Huang, Weixin Xie, Xuelong Li 0001 |
Pattern Recognit. Lett. | 1 |
| 2018 | Discovery of trading points based on Bayesian modeling of trading rules
Qinghua Huang, Zhoufan Kong, Yanshan Li, Jie Yang 0002, Xuelong Li 0001 |
World Wide Web | 3 |
| 2016 | Energy efficient design for multiuser downlink energy and uplink information transfer in 5G
Chunguo Li, Yanshan Li, Luxi Yang |
Sci. China Inf. Sci. | 2 |
| 2016 | Traffic anomaly detection based on image descriptor in videos
Yanshan Li, Weiming Liu 0003, Qinghua Huang |
Multim. Tools Appl. | 1 |
| 2016 | Fuzzy bag of words for social image description
Yanshan Li, Weiming Liu 0003, Qinghua Huang, Xuelong Li 0001 |
Multim. Tools Appl. | 1 |
| 2015 | A novel visual codebook model based on fuzzy geometry for large-scale image classification
Yanshan Li, Qinghua Huang, Weixin Xie, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2014 | GA-SIFT: A new scale invariant feature transform for multispectral image using geometric algebra
Yanshan Li, Weiming Liu 0003, Xiaotang Li, Qinghua Huang, Xuelong Li 0001 |
Inf. Sci. | 1 |