Leiquan Wang

dblp:150/8440 · DBLP profile ↗
← Back
48ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0003-4314-0030ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 5 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 ZS2Net: Frequency-aware semantic segmentation for zooplankton microscopic image
Dekun Yuan, Leiquan Wang, Yanping Qi, Zheng Qiao, Jie Zhang 0019
Expert Syst. Appl.3
2026 AGPL-KEM : Attribute-guided prompt learning with knowledge experts mixture for few-shot remote sensing image classification
Chunlei Wu, Congzheng Zhu, Qinfu Xu, Yongzhen Zhang, Leiquan Wang, Jie Wu 0033
Knowl. Based Syst.6
2026 A memory-tree driven network for multi-view fusion anomaly detection
Chunlei Wu, Huan Zhang 0015, Leiquan Wang
Pattern Recognit.4
2026 Reconstruction error-aware collaborative memory network for unsupervised anomaly detection
Chunlei Wu, Huan Zhang 0015, Leiquan Wang
Pattern Recognit.4
2025 Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution Constraints
abstract
Sentiment analysis is rapidly advancing by utilizing various data modalities (e.g., text, video, and audio). However, most existing techniques only learn the atomic-level features that reflect strong correlations, while ignoring more complex compositions in multimodal data. Moreover, they also neglected the incongruity in semantic distribution among modalities. In light of this, we introduce a novel Hierarchical Correlation Modeling Network (HCMNet), which enhances the multimodal sentiment analysis by exploring both the atomic-level correlations based on dynamic attention reasoning and the composition-level correlations through topological graph reasoning. In addition, we also alleviate the impact of distributional inconsistencies between modalities from both atomic-level and composition-level perspectives. Specifically, we first design an atomic-level contrastive loss that constrains the semantic distribution across modalities to mitigate the atomic-level inconsistency. Then, we design a graph optimal transport module that integrates transport flows with different graphs to constrain the composition-level semantic distribution, thus reducing the inconsistency of compositional nodes. Experiments on three public benchmark datasets have demonstrated the superiority of the proposed model over the state-of-the-art methods.
Qinfu Xu, Chunlei Wu, Leiquan Wang, Shaozu Yuan, Jie Wu 0033, Jing Lu 0013, Hengyang Zhou
AAAI4
2025 Multiple Feature Refining Network for Visual Emotion Distribution Learning
abstract
The significance of visual emotion distribution learning (VEDL) has surged, particularly with the growing inclination to convey emotions through images. The key of VEDL lies in capturing both low- and high-level features within the same visual content, thus promoting the model for salient and subtle emotion awareness. To learn the distribution of emotions involved in images, most previous works learn coarse semantic knowledge with unbiased filtering. Consequently, they focus on the entire scene and suffer from the redundancy of semantic-irrelevant information, which diminishes the affective coherence, impeding the comprehension of emotional attributes within the treated features. In light of this, we reanalyze from the perspective of information filtering and propose a novel method called Multiple Feature Refining Network (MFRN). To minimize low-level feature redundancy, we design a wavelet-based separated frequency modeling, named Spectral Mixer, to learn invariant representations and enhance emotion saliency in low-level image features. At the higher semantic level, we design a Semantic Graph Prompt Learning for emotional semantic filtering, ensuring the purity of emotional information and providing the model with richer content semantics. Experiments conducted on three commonly used datasets have demonstrated the superiority of our MFRN model over cutting-edge methods.
Qinfu Xu, Shaozu Yuan, Jie Wu 0033, Leiquan Wang, Chunlei Wu
AAAI5
2025 Recursive bidirectional cross-modal reasoning network for vision-and-language navigation
Jie Wu 0033, Chunlei Wu, Xiuxuan Shen, Fengjiang Wu, Leiquan Wang
Expert Syst. Appl.5
2025 Adaptive Cross-Modal Experts Network with Uncertainty-Driven Fusion for Vision-Language Navigation
Jie Wu 0033, Chunlei Wu, Xiuxuan Shen, Leiquan Wang
Knowl. Based Syst.4
2025 Instance-Wise Domain Generalization for Cross-Scene Wetland Classification With Hyperspectral and LiDAR Data
abstract
Wetland is one of the three ecosystems in the world, and collaborative monitoring using hyperspectral images (HSIs) and light detection and ranging (LiDAR) has been important for wetland ecological protection. However, because of the domain shift of different images, cross-scene wetland classification of HSIs and LiDAR is a practical challenge, necessitating the development of models trained solely on the source domain (SD) and directly transferred to the target domain (TD) without retraining. To address this issue, an instance-wise domain generalization network (IDGnet) is proposed for HSI and LiDAR cross-scene wetland classification. An instance-wise random domain expansion module (IWR-DEM) is developed to simulate the domain shift, establishing the extended domain (ED). Specifically, the original HSI and LiDAR data are separated as semantic and background information in the frequency domain, a random background shift is applied to the HSI, and a semantic random shift is deployed to LiDAR. The HSI and LiDAR fusion features are extracted from the SD and ED by a weight-shared network. Multiple condition constraints are proposed for domain and class alignment, learning the domain-invariant and class-specific information and improving model generalization. Experiments conducted on two wetland datasets demonstrate the superiority of the proposed IDGnet for cross-scene wetland classification with HSI and LiDAR data. The codes will be available from the website:https://github.com/bigshot-g/IEEE_TGRS_IDGnet.
Fangming Guo, Guangbo Ren, Leiquan Wang, Jie Zhang 0019, Jianbu Wang, Yabin Hu
IEEE Trans. Geosci. Remote. Sens.4
2024 Memory Self-Calibrated Network for Visual Grounding
abstract
Visual Grounding (VG) aims to locate the most relevant object or region in an image according to a natural language query. Existing methods in VG utilize fixed image and text representations to capture cross-modal semantic consistency, which limits the flexibility in adjusting image representations according to diverse textual information and hinders performance. To handle this limitation, we propose a novel Memory Self-Calibrated Network (MSCN) by dynamically refining image representations based on the query, thereby improving the semantic consistency between texts and images for visual grounding. Specifically, we introduce two modules: Semantic Relevance Filtering Module (SRFM) and Adaptive Memory Fusion Module (AMFM), to explicitly model the relationship between image and text. SRFM focuses on filtering out image information that is irrelevant to the query, while AMFM adaptively fuses text-related representations with initial image features to enhance the understanding ability of the MSCN model. Comprehensive experiments on three datasets demonstrate the superiority of our method compared to existing approaches.
Jie Wu 0033, Chunlei Wu, Xiuxuan Shen, Leiquan Wang
ICASSP5
2024 Improving visual grounding with multi-scale discrepancy information and centralized-transformer
Jie Wu 0033, Chunlei Wu, Fuyan Wang, Leiquan Wang
Expert Syst. Appl.4
2024 Vertical-horizontal latent space with iterative memory review network for multi-class anomaly detection
Chunlei Wu, Jie Wu 0033, Huan Zhang 0015, Leiquan Wang
Knowl. Based Syst.5
2024 Towards visual emotion analysis via Multi-Perspective Prompt Learning with Residual-Enhanced Adapter
Chunlei Wu, Qinfu Xu, Shaozu Yuan, Jie Wu 0033, Leiquan Wang
Knowl. Based Syst.6
2024 Multistage Synergistic Aggregation Network for Remote Sensing Visual Grounding
abstract
Visual Grounding has a broad application prospect in the field of remote sensing. Current state-of-the-art methods predominantly are based on the transformer architecture, utilizing multi-head self-attention in multi-modal encoders to integrate visual and textual features. However, they typically rely on a single fusion approach, which may limit the model’s capacity to learn intricate correlations between textual semantics and visual information. Moreover, they did not establish a direct dependency between features and bounding box representations, thereby restricting the fusion features to conventional object detection paradigm. Consequently, the interactions between regression results and encoded features are constrained. To address these limitations, a generative paradigm is harnessed to directly generate discrete coordinates sequence in an auto-regressive manner, which explores the interaction between direct regression features and encoded multi-modal features. Meanwhile, a novel multi-stage synergistic aggregation module is proposed to facilitate the acquisition of multi-modal features at multiple scales by effectively aggregating visual and textual contexts, enhancing the overall performance. In this work, we validate our framework on the DIOR-RSVG dataset and conduct a comparative analysis with existing methods, achieving a noteworthy improvement in accuracy. The proposed approach presents a promising direction for advancing visual grounding techniques in the context of remote sensing applications. The related code and weights are available at https://github.com/waynamigo/MSAM.
Fuyan Wang, Chunlei Wu, Jie Wu 0033, Leiquan Wang, Canwei Li
IEEE Geosci. Remote. Sens. Lett.4
2024 Learning Depth-Density Priors for Fourier-Based Unpaired Image Restoration
abstract
Deep learning-based image restoration methods trained on synthetic datasets have witnessed notable progress, but suffer from significant performance drops on real-world images due to huge domain shifts. To alleviate this issue, some recent methods strive to improve the generalization ability of models with unpaired training. However, these solutions typically handle each problem individually and ignore the shared physical properties of different harsh scenarios, i.e., heavy rain, hazy and low-light images degrade more densely with increasing scene depth. Such limitations make them generalize poorly to real-world images. In this paper, we propose a novel Physically Oriented Generative Adversarial Network (POGAN) for unpaired image restoration with depth-density priors. Specifically, our POGAN consists of two core designs: Physical Restoration Network (PRNet) and Degradation Rendering Network (DRNet). The former focuses on estimating the physical components related to the depth and density distribution for restoration, while the latter re-renders degradation effects guided by the estimated depth information. To further facilitate learning the above physical prior, we design a Spatial-Frequency Interaction Residual block (SFIR), which efficiently learns global frequency information and local spatial features in an interactive manner. Extensive experiments on synthetic and real-world datasets demonstrate the superiority of our method in heavy rain, haze, and low-light scenarios.
Yuanjian Qiao 0001, Ming-Wen Shao, Leiquan Wang, Wangmeng Zuo
IEEE Trans. Circuits Syst. Video Technol.3
2024 Multisource Feature Embedding and Interaction Fusion Network for Coastal Wetland Classification With Hyperspectral and LiDAR Data
abstract
With the development of earth observation technology, hyperspectral image (HSI) and light detection and ranging (LiDAR) data collaborative monitoring has shown great potential in the ecological protection and restoration of coastal wetlands. However, due to the different working principle adopted by the HSI sensor and LiDAR sensor, the data obtained by them has different distribution characteristics. The distribution difference limits the fusion of HSI and LiDAR data, bringing a great challenge for coastal wetland classification. To tackle this problem, a multi-source feature embedding and interaction fusion network is proposed for coastal wetland classification, named MsFE-IFN. First, the HSI and LiDAR data are embedded in the same feature space, where the feature distribution of multi-source remote sensing are aligned to alleviate data distribution differences. Second, the aligned HSI and LiDAR features interact information in channels and pixels, which is able to establish the relationship of spectral, elevation and geospatial. Third, the HSI and LiDAR feature are sent into the feature fusion network, in which the low-frequency residual is retained to enrich intra-class features. Finally, the fused feature is applied for final class prediction. Experiments conducted on three coastal wetland HSI-LiDAR datasets created by ourselves demonstrate the superiority of the proposed MsFE-IFN for coastal wetland classification. The codes will be available from the website:https://github.com/bigshot-g/IEEE_TGRS_MsFE-IFN.
Fangming Guo, Guangbo Ren, Leiquan Wang, Jie Zhang 0019, Rongyu Xin, Yabin Hu
IEEE Trans. Geosci. Remote. Sens.5
2024 Cycle Self-Training With Joint Adversarial for Cross-Scene Hyperspectral Image Classification
abstract
Cross-scene hyperspectral image classification (HSIC) leverages existing knowledge to categorize unknown scenes, aligning with the practical applications of remote sensing monitoring. However, spectral shifts across domains pose substantial challenges for this classification, aggravated by the insufficient number of labeled samples. Most existing methods predominantly address domain alignment from a singular perspective, rendering them inadequate to sustain robust classification performance in the presence of significant domain shifts. Additionally, although self-training can mitigate the labeling deficiency by leveraging unlabeled data, existing methods often fail to ensure the effective and accurate utilization of such data. Consequently, this article proposes a hyperspectral image (HSI) cross-scene classification architecture based on cycle self-training with joint adversarial (CSJA), which mitigates the impact of spectral shifts on cross-scene classification. Specifically, the proposed approach incorporates domain adversarial modules to reconcile domain distributions at varying granularities, coupled with a class adversarial module for joint adversarial alignment. Moreover, the cycle self-training (CST) module is devised to explicitly enforce pseudo-label generalization, thereby fully harnessing the informative content of the target domain. To effectively exploit both spatial and spectral information in HSIs and extract discriminative features, a convolutionally enhanced Transformer feature extraction network is proposed to generate feature-rich representations for both domains. Experimental evaluations on two publicly available cross-scene hyperspectral datasets and two self-made UAV hyperspectral datasets validate the superiority of the proposed algorithm.
Yajie Yang, Leiquan Wang, Mingming Xu 0001, Ziqi Xin, Yuewen Wang
IEEE Trans. Geosci. Remote. Sens.3
2024 Summator-Subtractor Network: Modeling Spatial and Channel Differences for Change Detection
abstract
The field of remote sensing (RS) image change detection (CD) has made significant progress, largely due to the powerful feature representation abilities of deep learning. However, traditional methods have not fully exploited the valuable information in differences. These methods often treat deep models as tools to extract features from individual images, which limits their ability to effectively describe differences. Additionally, many approaches tend to focus on spatial differences, while neglecting variations in the channel dimension. In this study, we introduce a novel Summator–Subtractor network for CD (${S}^{2}$CD), which adeptly captures subtle differences within both the spatial and channel aspects of bi-temporal images. The initial spatial and channel differences are derived through summation and subtraction operations on the bi-temporal images. The summator computes initial channel variations, while the subtractor captures initial spatial disparities. Transformers are then used to pull out meaningful differences in both spatial and channel patterns, allowing for a more nuanced understanding than methods relying solely on features from individual images. Finally, a heterogeneous modulation block integrates channel and spatial difference features, thus amplifying overall differences. Through extensive experimentation on four widely acknowledged CD benchmark datasets, our proposed${S}^{2}$CD method outperforms existing techniques, showcasing its superior performance and promising potential. The codes of this work will be available for the sake of reproducibility at:https://github.com/qianday/SSCD-CD.
Leiquan Wang, Ye Fang, Chunlei Wu, Mingming Xu 0001, Ming-Wen Shao
IEEE Trans. Geosci. Remote. Sens.1
2024 TDWCNet: Triple UNet With Dual-Window Convolution for Hyperspectral Anomaly Detection
abstract
In recent years, deep learning technology has emerged as the primary research focus in the field of hyperspectral anomaly detection (HAD) and has demonstrated satisfactory detection performance. Existing deep learning-based methods mainly utilize reconstruction errors as criteria for anomaly detection. However, they lack effective suppression of anomaly information in the background reconstruction process, and encounter challenges in addressing scenarios involving the coexistence of multiscale anomalies, which limits the performance of HAD. In order to reconstruct clean and reliable background images, this article proposes a Triple-UNet with dual-window convolution called TDWCNet for HAD. Specifically, we introduce a dual-window convolution that shields pixels within the inner window and only utilizes pixels between the inner and outer windows to reconstruct the central pixel. Based on the dual-window convolution, we construct the DWCBlock module, which serves as the core component for background reconstruction. To address the coexistence of multiscale anomalies, the Triple-UNet structure is designed, which organically combines three DWCBlock modules to gradually eliminate abnormal pixels that are inadvertently reconstructed due to inappropriate convolution kernels. Furthermore, adaptive mean-squared error (mse) and structural similarity index (SSIM) losses are employed to suppress anomaly reconstruction. Extensive experiments conducted on four publicly available datasets demonstrate that TDWCNet achieves satisfactory detection performance. The codes of this work will be available for the sake of reproducibility at:https://github.com/szc2277/TDWCNet-HAD.
Leiquan Wang, Zhicheng Sun 0005, Chunlei Wu, Mingming Xu 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Multilevel Class Token Transformer With Cross TokenMixer for Hyperspectral Images Classification
abstract
The transformer has become a prominent technique for hyperspectral image (HSI) classification, attributed to its capability to model global dependencies between features. Nevertheless, the predominant transformer-based methods rely on a direct information flow with a fixed number of tokens, causing the sequential transformer encoders to lack crucial interaction. This deficiency results in an inappropriate granularity of discriminative features and the loss of subtle patterns. In response to this limitation, we introduce a novel approach named Multi-level Class Token Transformer with Cross TokenMixer (MCTT) for HSI classification. Specifically, we explore a CNN stem network that incorporates 3D, 2D, and pointwise convolutions to encode local spatial-spectral information. The spectral-spatial features undergo transformation into semantic tokens using a semantic tokenizer. These tokens are then input into the transformer encoder to capture global interactions between different pixels. To create a hierarchical semantic representation, we propose a cross tokenmixer that integrates different levels of class tokens and patch tokens, enabling a multi-grained representation. The cross tokenmixers, with their varied number of tokens, facilitate the learning of distinct discriminative spectral-spatial representations and enable a comprehensive understanding of the HSI through a voting mechanism. Extensive experiments and ablation studies are conducted on three public HSI datasets to evaluate the performance of our proposed method. The results demonstrate the effectiveness and superior performance of our approach in HSI classification.
Leiquan Wang, Neeraj Kumar 0001, Fangming Guo, Peiying Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Nested Attention Network with Graph Filtering for Visual Question and Answering
abstract
Recently, Visual Question Answering(VQA), which is required to generate the answer by understanding both visual and textual content, has attracted considerable research interest. Most existing works extract visual features with the CNN network and learn its feature embedding with an attention mechanism. However, this mechanism may ignore the interaction between entities in the image, which has a fuzzy impact on the answer generation. To better explore the relationship between different entities in the image, a novel Nested Attention Network with Graph Filtering (NANGF) is proposed. It composes of two novel designed modules: a graph filtering mechanism to mine more precise visual semantics and avoid understanding deviation and nested attention to effectively guide the integration of visual features and question features. Extensive experiments conducted on the VQA2.0 datasets demonstrate the effectiveness of the proposed method.
Jing Lu 0013, Chunlei Wu, Leiquan Wang, Shaozu Yuan, Jie Wu 0033
ICASSP3
2023 Multi-view inter-modality representation with progressive fusion for image-text matching
Jie Wu 0033, Leiquan Wang, Chenglizhao Chen, Jing Lu 0013, Chunlei Wu
Neurocomputing2
2023 SOR-TC: Self-attentive octave ResNet with temporal consistency for compressed video action recognition
Junsan Zhang, Yao Wan 0001, Leiquan Wang, Jian Wang 0010, Philip S. Yu
Neurocomputing4
2023 Hyperspectral Image Few-Shot Classification Network With Brownian Distance Covariance
abstract
At present, how to achieve high precision hyperspectral image classification (HSIC) under the condition of few samples is a hot research issue. Metric-based meta-learning methods have proved to be very successful in this field. However, in terms of quantifying the dependencies between embedded features of hyperspectral samples, previous methods either only model marginal distribution and ignore joint distribution, limiting expressive capability of feature representation, or bring large computational cost though considering joint distribution. In this paper we propose a novel few shot learning (FSL) method based on Brownian distance covariance (BDC) for HSIC, which learns hyperspectral images’ representations by measuring the discrepancy between joint characteristic functions of embedded features and product of the marginals. In addition, a lightweight feature extraction network based on tied block convolution is proposed to better model cross channel correlation and aggregate global spectral-spatial features across channels. Extensive evaluations on several datasets show the effectiveness of the proposed method.
Ziqi Xin, Leiquan Wang, Mingming Xu 0001
IEEE Geosci. Remote. Sens. Lett.2
2023 A multimodal dialogue system for improving user satisfaction via knowledge-enriched response and image recommendation
Jiangnan Wang, Hai-Sheng Li 0002, Leiquan Wang, Chunlei Wu
Neural Comput. Appl.3
2023 Adversarial MixUp with implicit semantic preservation for semi-supervised hyperspectral image classification
Leiquan Wang, Jialiang Zhou, Chunlei Wu, Mingming Xu 0001
Signal Process.1
2023 Dynamic Pruning of Regions for Image-Sentence Matching
abstract
Image–sentence matching is becoming increasingly essential in the integrated understanding of vision and language. Prior approaches apply a pre-trained detection model to extract region features and explore fine-grained relationships between image and sentence by aggregating the similarities of all region–word pairs. However, all images are represented by the same number of regions, regardless of their respective semantic complexity, which results in a large number of redundant regions interfering with semantic inference and bringing additional computational burden. To address the lack of flexibility in image representation and information redundancy, a novel method named Dynamic Pruning of Regions for Image–Sentence Matching (DPRM) is proposed to efficiently capture relationships between text and image. In particular, a dynamic region pruning module is presented to dynamically select the appropriate number of regions according to the semantic complexity of each image, thus pruning redundant regions and reducing superfluous computations. Moreover, an inter-modality refinement module is designed to refine the fine-grained relationships of region–word pairs by retaining meaningful interaction features and suppressing interference from redundant alignments, which learns the more accurate semantic correspondences. Extensive experiments on MSCOCO and Flickr30K datasets prove the superiority of DPRM compared with previous approaches.
Jie Wu 0033, Weifeng Liu 0001, Leiquan Wang, Xiuxuan Shen, Chunlei Wu
Signal Process. Image Commun.3
2023 Eliminating Spatial Correlations of Anomaly: Corner-Visible Network for Unsupervised Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) is crucial for identifying and analyzing abnormal objects in various domains. While existing methods have shown promising results by designing detection methods tailored to specific anomaly characteristics, there is a need for a highly versatile approach that can effectively handle anomalies, particularly those with large spatial sizes. In this article, we propose an end-to-end corner-visible network (CVNet) for unsupervised HAD. Specifically, we introduce a corner-visible convolution that leverages the statistical dependencies of the background within the receptive field while eliminating spatial correlations with potential anomalies for background generation. To address the grid effect caused by the corner-visible convolution, a background smoothing module is employed by using the conventional convolution. Furthermore, adaptive mean-squared error (MSE) and structural similarity index (SSIM) losses are employed to suppress anomaly reconstruction, resulting in a reliable reconstructed background map. Anomalies are identified through the residual of the original hyperspectral image (HSI) and the reconstructed background. Extensive experiments conducted on three publicly datasets demonstrate the effectiveness of our proposed method in handling different types of anomalies. The state-of-the-art performance showcases the versatility and applicability of CVNet in HAD. The codes of this work will be available for the sake of reproducibility athttps://github.com/Cloudynewbee/CVNet-HAD.
Leiquan Wang, Chunlei Wu, Mingming Xu 0001, Ming-Wen Shao
IEEE Trans. Geosci. Remote. Sens.1
2023 Attentive-Adaptive Network for Hyperspectral Images Classification With Noisy Labels
abstract
With the development of deep neural networks, hyperpsectral image (HSI) classification systems have achieved a significant improvement. These systems require numerous and accurate labeled hyperspectral data to be adequately trained. However, noisy labels are inherent in real-world hyperspectral systems, resulting in unreliable decisions. To handle noisy labels in hyperpsectral classification, an end-to-end attentive-adaptive network (AAN) is proposed for robust HSI classification training. The goal is to build a classifier with strong generalization capabilities that can be applied to both clean and noisy training sets without explicit noise label pre-treatment. Specifically, a spectral stem network with non-adjacent shortcut is exploited initially to re-distribute the sensitive layers for noisy labels to achieve robust spectral representation. Then, a group-shuffle attention module is proposed to capture the discriminative and robust spatial-spectral features in the presence of noisy labels. Finally, an adaptive noise-robust loss function is developed to fight against noisy labels by learning a parameter to balance the normalized cross entropy (NCE) and reverse cross entropy (RCE). Experimental results on three HSI benchmark datasets with simulated noisy labels demonstrate the effectiveness of AAN on HSI classification.
Leiquan Wang, Tongchuan Zhu, Neeraj Kumar 0001, Chunlei Wu, Peiying Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Dual Graph Convolution Joint Dense Networks for Hyperspectral and LiDAR Data Classification
abstract
With the increasing demand of observation, multi-source remote sensing data has been widely used. Hyperspectral Images (HSI) and Light Detection and Ranging (LiDAR) data have shown the great potential in land cover classification. However, the redundant information of multi-source data influences the effectiveness of heterogeneous data features, which reduces the accuracy of joint classification. To tackle this problem, a dual graph convolution joint dense networks is proposed for HSI and LiDAR classification. In this method, a dual graph convolution network (GCN)is extracted the spectral feature from euclidean graph and cosine graph, which contains the spectrum absolute and relative differences. A dense network is employed to acquire spatial feature from LiDAR data. Finally, a fully connected network fuses the spectral and spatial feature for classification. Experiments conducted on the Huston dataset demonstrate the effectiveness of the proposed method on joint classification.
Fangming Guo, Leiquan Wang, Jie Zhang 0019
IGARSS4
2022 Convolution Enhanced Spatial-Spectral Unified Transformer Network for Hyperspectral Image Classification
abstract
Convolutional neural network has achieved great success in hyperspectral image classification for its excellent local context modeling capabilities. However, the convolution operation with fixed-size local receptive fields is difficult to establish long-distance dependence in hyperspectral im-age. To address this problem, we propose a spatial-spectral unified transformer network, which utilizes self-attention mechanisms to extract global spatial and spectral features. In addition, in order to introduce local spatial and spectral information, the convolution operation is integrated into the network. Specifically, spatial and spectral convolutional embedding layers are designed to generate embeddings of spatial patches and spectral bands. Besides, depthwise convolution is exploited in the locally-enhanced feedforward layer to bring locality into transformer. Experimental results on two datasets demonstrate that our proposed network has greatly improved compared with other state-of-the-art methods.
Ziqi Xin, Mingming Xu 0001, Leiquan Wang
IGARSS4
2022 Anti-jamming heart rate estimation using a spatial-temporal fusion network
Chunlei Wu, Ziyu Yuan, Shaohua Wan 0001, Leiquan Wang, Weishan Zhang
Comput. Vis. Image Underst.4
2022 Generating diverse chinese poetry from images via unsupervised method
Jiangnan Wang, Hai-Sheng Li 0002, Chunlei Wu, Faming Gong, Leiquan Wang
Neurocomputing5
2022 Region Reinforcement Network With Topic Constraint for Image-Text Matching
abstract
Image and sentence matching has attracted increasing attention since it is associated with two important modalities of vision and language. Previous methods aim to find the latent correspondences between image regions and words by aggregating the similarities of the region-word pairs. However, these approaches consider little about the relationships of diverse regions in the image and treat the similarities of all region-word pairs equally. Moreover, focusing on fine-grained alignment overly, the true meaning of the original image will be likely distorted. In this paper, a novel Region Reinforcement Network with Topic Constraint (RRTC) is proposed to explore the correspondences between images and texts. Specifically, the region reinforcement network is built to infer fine-grained correspondence by considering the relationships of regions and re-assigning region-word similarities. Meanwhile, the topic constraint module is presented to summarize the central theme of images, which constrains the original image deviation. Extensive experimental results on MSCOCO and Flickr30k datasets verify the effectiveness of our proposed RRTC.
Jie Wu 0033, Chunlei Wu, Jing Lu 0013, Leiquan Wang, Xue-rong Cui
IEEE Trans. Circuits Syst. Video Technol.4
2022 Multiscale Contrastive Learning Networks for Automatic Denoising of Geological Sedimentary Model Images
abstract
The well-described river courses in geological sedimentary models are essential for identifying oil and gas reservoirs. However, numerous noises are generated around the river course, including irregular noises and regions due to errors in the geological data, as well as traces left over from the printed model. The cluttered distributions of noises make the complete river course unobtainable, which interferes with reservoir prediction. To the best of our knowledge, deep learning-based methods for automatic noise detection and removal have not been explored in the field of processing geological sedimentary model images. In this paper, we present Multi-scale Contrastive Learning Networks (CLGAN) for detecting noise in geological sedimentary model images. A multi-scale contrastive training strategy is proposed to capture noise locations and river discontinuities by comparing the distributions with and without noise in image space and feature space. A paired dataset that consists of images with noise and images without noise is constructed for contrastive training. Moreover, cyclic denoising is proposed to completely denoise by constantly modifying the pixel values at the noise, which prevents missed detections by calling the model multiple times. Subsequently, pixel-level denoising is designed to initially remove a portion of the noise by contouring non-river regions, which reduces the pressure on the cyclic denoising method. The detection results demonstrate the effectiveness of CLGAN. The denoising results with excellent river connectivity and integrity highlight the superiority of the proposed denoising methods.
Chunlei Wu, Huan Zhang 0015, Leiquan Wang
IEEE Trans. Geosci. Remote. Sens.5
2021 Dual-View Semantic Inference Network for image-text matching
Chunlei Wu, Jie Wu 0033, Haiwen Cao, Leiquan Wang
Neurocomputing5
2021 Generate classical Chinese poems with theme-style from images
Chunlei Wu, Jiangnan Wang, Shaozu Yuan, Leiquan Wang, Weishan Zhang
Pattern Recognit. Lett.4
2020 Multi-Attention Generative Adversarial Network for image captioning
Leiquan Wang, Haiwen Cao, Ming-Wen Shao, Chunlei Wu
Neurocomputing2
2019 Video-level Multi-model Fusion for Action Recognition
abstract
The approaches based on spatio-temporal features for video action recognition have emerged such as two-stream based methods and 3D convolution based methods. However, current methods suffer from the problems caused by partial observation, or restricted to single information modeling, and so on. Segment-level recognition results obtained from dense sampling can not represent the entire video and, therefore lead to partial observation. And a single model is hard to capture the complementary information on spacial, temporal and spatio-temporal information from video at the same time. Therefore, the challenge is to build the video-level representation and capture multiple information. In this paper, a video-level multi-model fusion action recognition method is proposed to solve these problems. Firstly, an efficient video-level 3D convolution model is proposed to get the global information in the video which assembling segment-level 3D convolution models. Secondly, a multi-model fusion architecture is proposed for video action recognition to capture multiple information. The spatial, temporal and spatio-temporal information are aggregate with SVM classifier. Experimental results show that this method achieves the state-of-the-art performance on the datasets of UCF-101(97.6%) without pre-training on Kinetics.
Junsan Zhang, Leiquan Wang, Philip S. Yu, Hai-Sheng Li 0002
CIKM3
2019 Jointly Predicting Future Sequence and Steering Angles for Dynamic Driving Scenes
abstract
Generative Adversarial Network (GAN) has attracted rising attention for video future sequence prediction in driving scenes. However, the images generated by GAN often miss the target for lack of any constraints for its generated target. In this paper, an encoder-decoder based multi-task video prediction network - SegVAE is proposed by simultaneously accomplishing the predictions (generations) of both future sequence and steering angles for egocentric driving videos at pixel-level. Specifically, the encoder is constructed based on Varitional Auto-Encoder (VAE) to learn the complex latent distribution of real driving scenes. The decoder is exploited with a multi-task manner to jointly predict the future sequence and steering angles of dynamic driving scenes, where an enhanced generation mechanism is also proposed. Varitional Auto-Encoder (VAE) and Long Short Term Memory Networks (LSTM) are introduced to optimize the learning of SegVAE. The experimental results on public KITTI and NVIDIA driving datasets indicate that the proposed Seg-VAE can effectively mimic humans prediction mechanism, and outperform standard VAE and CNN-based generative adversarial network.
Zhicheng Zhao 0001, Leiquan Wang, Chaohong An
ICASSP4
2018 Improving deep neural networks with multi-layer maxout networks and a novel initialization method
Weichen Sun, Leiquan Wang
Neurocomputing3
2018 Hierarchical attention-based multimodal fusion for video captioning
Chunlei Wu, Xiaoliang Chu, Weichen Sun, Leiquan Wang
Neurocomputing6
2018 Modeling visual and word-conditional semantic attention for image captioning
Chunlei Wu, Xiaoliang Chu, Leiquan Wang
Signal Process. Image Commun.5
2017 Modeling intra- and inter-pair correlation via heterogeneous high-order preserving for cross-modal retrieval
Leiquan Wang, Weichen Sun, Zhicheng Zhao 0001
Signal Process.1
2016 Deep canonical correlation analysis with progressive and hypergraph learning for cross-modal retrieval
Jie Shao 0014, Leiquan Wang, Zhicheng Zhao 0001, Anni Cai
Neurocomputing2
2016 Efficient multi-modal hypergraph learning for social image classification with complex label correlations
Leiquan Wang, Zhicheng Zhao 0001
Neurocomputing1
2014 Improving deep neural networks with multilayer maxout networks
abstract
For the purpose of enhancing discriminability of convolutional neural networks (CNNs) and facilitating optimization, a multilayer structured variant of the maxout unit (named Multilayer Maxout Network, MMN) is proposed in this paper. CNNs with maxout units employ linear convolution filters followed by maxout units to abstract representations from less abstract ones. Our model instead applies MMNs as activation functions of CNNs to abstract representations, which inherits advantages of both maxout units and deep neural networks, and is a more general nonlinear function approximator as well. Experimental results show that our proposed model yields better performance on three image classification benchmark datasets (CIFAR-10, CIFAR-100 and MNIST) than some state-of-the-art methods. Furthermore, the influence of MMN in different hidden layers is analyzed, and a trade-off scheme between the accuracy and computing resources is given.
Weichen Sun, Leiquan Wang
VCIP3
2014 Tag-based social image search with hyperedges correlation
abstract
In social image search, most existing hypergraph methods use the visual and textual features in isolation by treating each feature term as a hyperedge. Nevertheless, they neglect the correlations of visual and textual hyperedges, which are more robust to represent the high-order relationship among vertices. In this paper, we propose a hypergraph with correlated hyperedges (CHH), which introduces high-order relationship of hyperedges into hypergraph learning. Based on CHH, a pairwise visual-textual correlation hypergraph (VTCH) model is used for tag-based social image search. To overcome the large number of newly generated hybrid hyperedges, a bagging-based method is adopted to balance the accuracy and speed. Finally, adaptive hyperedges learning method is used to obtain the relevance score for social image search. The experiments conducted on MIR Flickr show the effectiveness of our proposed method.
Leiquan Wang, Zhicheng Zhao 0001
VCIP1