Wenjing Chen 0003

dblp:74/8490-3 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-3878-7681ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Spectral-spatial fusion Mamba for spectral super-resolution from single RGB image
Wenjing Chen 0003, Mingtao Zhong, Zhiwei Ye
Knowl. Based Syst.1
2026 Reconstruction-Contrast Coupling Learning for Open-Set Semi-Supervised Hyperspectral Image Classification
abstract
Although numerous semi-supervised learning methods have been elaborately designed for hyperspectral image (HSI) classification, most existing semi-supervised learning paradigms still rely on a closed-set assumption. These methods implicitly assume that the category spaces of labeled and unlabeled samples are completely aligned, that is, all unlabeled samples must belong to a pre-defined known category set. However, the closed-set assumption is particularly problematic in practical remote sensing scenarios because partial unlabeled data inevitably belong to unknown categories. To address this challenge, this paper proposes a reconstruction-contrast coupling learning (ReCo2L) method for open-set semi-supervised HSI classification, fully leveraging the complementarity between masked feature reconstruction learning and contrastive learning to enhance the encoder’s local detail sensitivity and global discriminative ability. Specifically, we first apply a masked feature reconstruction learning with an adaptive masking strategy to enhance the encoder’s ability to capture local details by high-quality spectral-spatial feature reconstruction. Then, we employ contrastive learning to strengthen the encoder’s capability to extract global characteristics by pulling semantically similar samples closer and pushing dissimilar ones farther apart in the feature space. Finally, a pixel-prototype deviation loss is proposed to further improve both inter-category distinguishability and intra-category compactness by reducing the distances between labeled sample features and their corresponding class anchors. Extensive experiments on three benchmark datasets demonstrate that our proposed ReCo2L achieves superior classification performance in both known and unknown categories and significantly surpasses 10 state-of-the-art HSI classification methods. The code will be available at https://github.com/repository-AI-chen/ReCo2L.
Hao Sun 0014, Renyi Chen, Yong Chen 0024, Wenjing Chen 0003, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Image Process.4
2025 Relation-Aware Multiprototype Learning for Semi-Supervised Hyperspectral Image Classification
Wenjing Chen 0003, Renyi Chen, Zhiwei Ye, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.1
2025 Learning Positive-Negative Prompts for Open-Set Remote Sensing Scene Classification
Hao Sun 0014, Hanlizi Chen, Wenjing Chen 0003, Chengji Wang, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2024 TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
abstract
Video moment retrieval (MR) and highlight detection (HD) based on natural language queries are two highly related tasks, which aim to obtain relevant moments within videos and highlight scores of each video clip. Recently, several methods have been devoted to building DETR-based networks to solve both MR and HD jointly. These methods simply add two separate task heads after multi-modal feature extraction and feature interaction, achieving good performance. Nevertheless, these approaches underutilize the reciprocal relationship between two tasks. In this paper, we propose a task-reciprocal transformer based on DETR (TR-DETR) that focuses on exploring the inherent reciprocity between MR and HD. Specifically, a local-global multi-modal alignment module is first built to align features from diverse modalities into a shared latent space. Subsequently, a visual feature refinement is designed to eliminate query-irrelevant information from visual features for modal interaction. Finally, a task cooperation module is constructed to refine the retrieval pipeline and the highlight score prediction process by utilizing the reciprocity between MR and HD. Comprehensive experiments on QVHighlights, Charades-STA and TVSum datasets demonstrate that TR-DETR outperforms existing state-of-the-art methods. Codes are available at https://github.com/mingyao1120/TR-DETR.
Hao Sun 0014, Mingyao Zhou, Wenjing Chen 0003, Wei Xie 0008
AAAI3
2024 Spatial Formation-Guided Network for Group Activity Recognition
abstract
Effectively modeling the interactions among actors is critical and challenging for Group Activity Recognition (GAR). Previous methods usually divide actors into subgroups based on the similarity of appearance features for modeling multilevel interactions among actors. However, the appearance feature-based grouping scheme does not fully consider the spatial relations of actors, which can provide a discriminative clue for GAR. In this paper, we propose a Spatial Formation-Guided Network (SFGN) to capture effective interactions under the guidance of spatial formations. We first design a spatial formation extractor to excavate latent spatial relations among actors for extracting spatial formation features. Then, a formation-guided interaction module is built to utilize the spatial formation features to guide the interactions among actors. Finally, a cross-formation interaction module is further designed to explore the complementarity among diverse spatial formations. Extensive experiments on the volleyball dataset and the collective activity dataset demonstrate that SFGN outperforms the state-of-the-art methods.
Dunbo Ning, Wenjing Chen 0003, Wei Xie 0008, Hao Sun 0014
ICASSP2
2024 Cross-Modal Multiscale Difference-Aware Network for Joint Moment Retrieval and Highlight Detection
abstract
Since the goals of both Moment Retrieval (MR) and Highlight Detection (HD) are to quickly obtain the required content from the video according to user needs, several works have attempted to take advantage of the commonality between both tasks to design transformer-based networks for joint MR and HD. Although these methods achieve impressive performance, they still face some problems: a) Semantic gaps across different modalities. b) Various durations of different query-relevant moments and highlights. c) Smooth transitions among diverse events. To this end, we propose a Cross-modal Multiscale Difference-aware Network, named CMDNet. First, a clip-text alignment module is constructed to narrow semantic gaps between different modalities. Second, a multiscale difference perception module is utilized to mine the differential information between adjacent clips and perform multiscale modeling to obtain discriminative representations. Finally, these representations are fed into the MR and HD task heads to retrieve relevant moments and estimate highlight scores precisely. Extensive experiments on three popular datasets demonstrate that CMDNet achieves state-of-the-art performance.
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008
ICASSP2
2024 Spatial Dual Context Learning for Weakly-supervised Group Activity Recognition in Still-images
abstract
This paper investigates a new task, Weakly- supervised Group Activity Recognition in Still-images (WGARS), which aims to extend the applicability of Group Activity Recognition (GAR) to broader scenarios, such as low-latency domains. To tackle this challenge, we propose a Spatial Dual Context Transformer (SDCT), comprising a Dual Context Encoder (DCE) and a Dual Context Decoder (DCD). The DCE module individually encodes holistic context with integral relations of overall actors, and encodes partial context with individual features in still images. Subsequently, the DCD module explores the complementarity between holistic and partial contexts, and alternatively updates these encoded contexts to enhance the interaction of actors. Additionally, auxiliary supervised contrastive learning is incorporated to mitigate activity confusion. The proposed SDCT attains state-of-the-art performance on Volleyball and NBA datasets in WGARS. Notably, SDCT even outperforms recent methods when extended to the weakly-supervised GAR in videos task on Volleyball dataset.
Dunbo Ning, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004
ICME3
2024 Query-aware multi-scale proposal network for weakly supervised temporal sentence grounding in videos
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Xiaoqiang Lu
Knowl. Based Syst.2
2024 Prototype-Based Pseudo-Label Refinement for Semi-Supervised Hyperspectral Image Classification
abstract
Pseudo-label learning-based methods usually regard class confidence above a certain threshold for unlabeled samples as pseudo-labels, which may result in pseudo-labels still containing wrong labels. In this letter, we propose a prototype-based pseudo-label refinement (PPLR) for semi-supervised hyperspectral image classification. The proposed PPLR filters wrong labels from pseudo-labels using class prototypes, which can improve the discrimination of the network. First, PPLR uses multi-head attentions to extract the spectral-spatial features, and designs an adaptive threshold that can be dynamically adjusted to generate high-confidence pseudo-labels. Then, PPLR constructs class prototypes for different categories using labeled sample features and unlabeled sample features with refined pseudo-labels to improve the quality of pseudo-labels by filtering wrong labels. Finally, PPLR further assigns reliable weights to these pseudo-labels in calculating their supervised loss, and introduces a center loss to improve the discrimination of features. When 10 labeled samples per category are utilized for training, PPLR achieves the overall accuracies of 82.11%, 86.70% and 92.50% on the Indian Pines, Houston2013 and Salinas datasets, respectively.
Renyi Chen, Huaxiong Yao, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.3
2024 Orientational Clustering Learning for Open-Set Hyperspectral Image Classification
abstract
Recently, some literature has begun to pay attention to the open-set problem in remote sensing application scenarios and studied various open-set hyperspectral image classification (OSHIC) methods. These OSHIC methods are usually based on deep neural networks, using the nondirectional Euclidean distance losses to constrain latent sample representations of known classes to be compact. Nonetheless, the potential effect of the spatial distribution of sample representations is ignored, resulting in degraded classification performance in OSHIC. In this letter, we propose an orientational clustering learning (OCL) method for OSHIC. First, in the feature space generated by the convolutional neural network, a class anchor strategy is employed to bring features of the same class closer while keeping features of different classes distant. Then, we utilize the orientational learning to further tighten the intraclass feature space. OCL directionally optimizes the spatial distribution of hyperspectral sample representations to improve the ability to identify known classes and distinguish unknown classes. Experiments show that the OCL achieves overall accuracies of 94.43%, 92.27%, and 76.94% on the Pavia University, Salinas, and Indian Pines datasets, respectively.
Wenjing Chen 0003, Hailong Ning, Hao Sun 0014, Wei Xie 0008
IEEE Geosci. Remote. Sens. Lett.2
2024 Cross-Modal Feature Fusion-Based Knowledge Transfer for Text-Based Person Search
abstract
Text-based person search aims to retrieve corresponding images of person from a large gallery based on text descriptions. Existing methods strive to bridge the modality gap between images and texts and have made promising progress. However, these approaches disregard the knowledge imbalance between images and texts caused by the reporting bias. To resolve this issue, we present a cross-modal feature fusion-based knowledge transfer network to balance identity information between images and texts. First, we design an identity information emphasis module to enhance person-relevant information and suppress person-irrelevant information. Second, we design an intermediate modal-guided knowledge transfer module to balance the knowledge between images and texts. Experimental results on CUHK-PEDES, ICFG-PEDE, and RSTPReid datasets demonstrate that our method achieves state-of-the-art performance.
Kaiyang You, Wenjing Chen 0003, Chengji Wang, Hao Sun 0014, Wei Xie 0008
IEEE Signal Process. Lett.2
2023 Deep Feature Reconstruction Learning for Open-Set Classification of Remote-Sensing Imagery
abstract
Existing remote sensing scene image (RSSI) classification methods usually rely on static closed-set assumption that testing samples do not belong to unknown classes. However, practical applications are usually the open-set classification problem, which means that RSSIs from unknown classes will appear in the testing set. Most existing methods are prone to forcibly misclassify RSSIs of unknown classes into known classes, resulting in poor practical performance. In this letter, a deep feature reconstruction learning (DFRL) framework is proposed for open-set classification of RSSIs. The proposed DFRL unifies discriminative feature learning and feature reconstruction into an end-to-end network. Firstly, a feature extraction module is utilized to project raw input data from the image space to the feature space to extract deep features. Then, the deep features are fed to a deep feature reconstruction module for distinguishing known and unknown classes based on feature-level reconstruction errors. The feature-level reconstruction can effectively suppress the interference of complex backgrounds. In addition, a sparse regularization is introduced to improve the discrimination of image representation. Experiments on three RSSI datasets demonstrate the effectiveness of DFRL for open-set classification of RSSIs.
Hao Sun 0014, Jie Yu 0003, Dongbo Zhou, Wenjing Chen 0003, Xiangtao Zheng, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.5
2023 Pseudolabel-Based Unreliable Sample Learning for Semi-Supervised Hyperspectral Image Classification
abstract
Recently, pseudo-label-based deep learning methods have shown excellent performance in semi-supervised hyperspectral image (HSI) classification. These methods usually select high-confidence unlabeled samples to help optimize backbone classification networks. However, a large number of remaining low-confidence unlabeled samples, which contain rich land-covers information, are underutilized. In this paper, we propose a pseudo-label-based unreliable sample learning (PUSL) method to fully exploit low-confidence unlabeled samples for semi-supervised HSI classification. Firstly, to avoid overfitting the spatial distribution of labeled samples, we build a position-free transformer (PFT) as the backbone classification network. Secondly, PFT is initially trained with labeled samples in a supervised learning manner to obtain an initial classifier, which is then used to split unlabeled samples into reliable and unreliable unlabeled samples based on the predicted confidence. Thirdly, reliable unlabeled samples participate in training along with labeled samples. Finally, unreliable unlabeled samples are treated as negative samples for corresponding categories to improve the discrimination of PFT in a contrastive learning paradigm. Extensive experiments on three HSI datasets demonstrate that PUSL outperforms compared methods.
Huaxiong Yao, Renyi Chen, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2022 Semisupervised Spectral Degradation Constrained Network for Spectral Super-Resolution
abstract
Recently, various deep learning-based methods have been designed to improve the spectral resolution of the multispectral image (MSI) to obtain the hyperspectral image (HSI). These methods usually rely on sufficient MSI/HSI pairs for supervised training. However, collecting plentiful HSIs is time-consuming. In this letter, a semisupervised spectral degradation constrained network (SSDCN) is proposed to improve the spectral resolution of MSI. SSDCN is an autoencoder-like network that is composed of an encoder subnetwork for estimating HSI from input MSI and a decoder subnetwork for reconstructing MSI from the estimated HSI. A semisupervised training method is proposed to explore both MSI/HSI pairs and MSIs without ground-truth HSIs to optimize SSDCN. Simulated and two real databases are employed to demonstrate the effectiveness of SSDCN.
Wenjing Chen 0003, Xiangtao Zheng, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.1
2022 Spectral Super-Resolution of Multispectral Images Using Spatial-Spectral Residual Attention Network
abstract
The spectral super-resolution of multispectral image (MSI) refers to improving the spectral resolution of the MSI to obtain the hyperspectral image (HSI). Most recent works are based on the sparse representation to unfold the MSI into the 2-D matrix in advance for subsequent operations, which results in that the spatial information of MSI cannot be fully explored. In this article, a spatial–spectral residual attention network (SSRAN) is proposed to simultaneously explore the spatial and spectral information of MSI for reconstructing the HSI. The proposed SSRAN is composed of the feature extraction part, the nonlinear mapping part, and the reconstruction part. Firstly, the multispectral features of the input MSI are extracted in the feature extraction part. Second, in the nonlinear mapping part, the spatial–spectral residual blocks are proposed to explore spatial and spectral information of MSI for mapping the multispectral features to the hyperspectral features. Finally, in the reconstruction part, a 2-D convolution is used to reconstruct the HSI from the hyperspectral features. Also, a neighboring spectral attention module is specially designed to explicitly constrain the reconstructed HSI to maintain the correlation among neighboring spectral bands. The proposed SSRAN outperforms the state-of-the-art methods on both simulated and real databases.
Xiangtao Zheng, Wenjing Chen 0003, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.2
2020 Unregistered Hyperspectral and Multispectral Image Fusion with Synchronous Nonnegative Matrix Factorization
Wenjing Chen 0003, Xiaoqiang Lu
PRCV (1)1