Wei Xie 0008

dblp:87/1010-8 · DBLP profile ↗
← Back
36ranked-venue papers
0as first author
27since 2021 · last 2026
0000-0001-7734-1274ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Reconstruction-Contrast Coupling Learning for Open-Set Semi-Supervised Hyperspectral Image Classification
abstract
Although numerous semi-supervised learning methods have been elaborately designed for hyperspectral image (HSI) classification, most existing semi-supervised learning paradigms still rely on a closed-set assumption. These methods implicitly assume that the category spaces of labeled and unlabeled samples are completely aligned, that is, all unlabeled samples must belong to a pre-defined known category set. However, the closed-set assumption is particularly problematic in practical remote sensing scenarios because partial unlabeled data inevitably belong to unknown categories. To address this challenge, this paper proposes a reconstruction-contrast coupling learning (ReCo2L) method for open-set semi-supervised HSI classification, fully leveraging the complementarity between masked feature reconstruction learning and contrastive learning to enhance the encoder’s local detail sensitivity and global discriminative ability. Specifically, we first apply a masked feature reconstruction learning with an adaptive masking strategy to enhance the encoder’s ability to capture local details by high-quality spectral-spatial feature reconstruction. Then, we employ contrastive learning to strengthen the encoder’s capability to extract global characteristics by pulling semantically similar samples closer and pushing dissimilar ones farther apart in the feature space. Finally, a pixel-prototype deviation loss is proposed to further improve both inter-category distinguishability and intra-category compactness by reducing the distances between labeled sample features and their corresponding class anchors. Extensive experiments on three benchmark datasets demonstrate that our proposed ReCo2L achieves superior classification performance in both known and unknown categories and significantly surpasses 10 state-of-the-art HSI classification methods. The code will be available at https://github.com/repository-AI-chen/ReCo2L.
Hao Sun 0014, Renyi Chen, Yong Chen 0024, Wenjing Chen 0003, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Image Process.5
2025 PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric Network
abstract
With the exponential growth of video content, aiming at localizing relevant video moments based on natural language queries, video moment retrieval (VMR) has gained significant attention. Existing weakly supervised VMR methods focus on designing various feature modeling and modal interaction modules to alleviate the reliance on precise temporal annotations. However, these methods have poor generalization capabilities on compositional queries with novel syntactic structures or vocabulary in real-world scenarios. To this end, we propose a new task: weakly supervised compositional moment retrieval (WSCMR). This task trains models using only video-query pairs without precise temporal annotations, while enabling generalization to complex compositional queries. Furthermore, a proposal-centric network (PC-Net) is proposed to tackle this challenging task. First, video and query features are extracted through frozen feature extractors, followed by modality interaction to obtain multimodal features. Second, to handle compositional queries with explicit temporal associations, a dual-granularity proposal generator decodes multimodal global and frame-level features to obtain query-relevant proposal boundaries with fine-grained temporal perception. Third, to improve the discrimination of proposal features, a proposal feature aggregator is constructed to conduct semantic alignment of frames and queries, and employ a learnable peak-aware Gaussian distributor to fit the frame weights within the proposals to derive proposal features from the video frame features. Finally, the proposal quality is assessed based on the results of reconstructing the masked query using the obtained proposal features. To further enhance the model's ability to capture semantic associations between proposals and queries, a quality margin regularizer is constructed to dynamically stratify proposals into high and low query-relevance subsets and enhance the association between queries and common elements within proposals, and suppress spurious correlations via inter-subset contrastive learning. Notably, PC-Net achieves superior performance with 54\% fewer parameters than prior works by parameter-efficient design. Experiments on Charades-CG and ActivityNet-CG demonstrate PC-Net’s ability to generalize across diverse compositional queries. Code is available at https://github.com/mingyao1120/PC-Net.
Mingyao Zhou, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Chengji Wang, Mang Ye
NeurIPS3
2025 Correlation-based switching mean teacher for semi-supervised medical image segmentation
Guiyuhan Deng, Hao Sun 0014, Wei Xie 0008
Neurocomputing3
2025 Learning Positive-Negative Prompts for Open-Set Remote Sensing Scene Classification
Hao Sun 0014, Hanlizi Chen, Wenjing Chen 0003, Chengji Wang, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2025 Class-Aware Consistency Learning for Open-Set Semi-Supervised Hyperspectral Image Classification
abstract
Semi-supervised hyperspectral image (HSI) classification methods focus on exploring the spectral and spatial information of unlabeled samples. However, existing methods generally follow the closed-set setting, assuming that unlabeled samples do not contain novel classes, which is hard to hold in practical applications. This paper aims to study semi-supervised HSI classification in the open-set setting, i.e., unlabeled samples fall into novel classes, and proposes a class-aware consistency learning (CACL) method. First, to explore discriminative spectral-spatial features, a position-aware transformer is developed, which effectively models spatial position priors between the center pixel and its neighboring pixels via a symmetric position-aware encoding. Then, to reduce the interference from novel class samples on the model’s discrimination, a prototype-driven consistency learning is proposed, which accurately selects unlabeled samples belonging to known classes via a known class sampler, and efficiently utilizes their spectral-spatial information by modeling consistent predictions across different views. Finally, to further improve the distinguishability between known classes, a prototype contrastive optimization is proposed to decrease the distance between samples from the same class and increase the distances between those from different classes in the feature domain. Furthermore, an adaptive segmentation threshold is designed to accurately predict known classes and reject novel classes. Extensive experiments verify that our CACL outperforms the state-of-the-art methods, achieving the overall accuracy of 81.65%, 88.61%, and 92.88% on the Indian Pines, Salinas, and Pavia University datasets, with 10 labeled samples in each known class. The code is available at https://github.com/rock-in/CACL-main.
Hao Sun 0014, Renyi Chen, Huaxiong Yao, Yaxiong Chen, Wei Xie 0008, Guirong Feng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.6
2024 TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
abstract
Video moment retrieval (MR) and highlight detection (HD) based on natural language queries are two highly related tasks, which aim to obtain relevant moments within videos and highlight scores of each video clip. Recently, several methods have been devoted to building DETR-based networks to solve both MR and HD jointly. These methods simply add two separate task heads after multi-modal feature extraction and feature interaction, achieving good performance. Nevertheless, these approaches underutilize the reciprocal relationship between two tasks. In this paper, we propose a task-reciprocal transformer based on DETR (TR-DETR) that focuses on exploring the inherent reciprocity between MR and HD. Specifically, a local-global multi-modal alignment module is first built to align features from diverse modalities into a shared latent space. Subsequently, a visual feature refinement is designed to eliminate query-irrelevant information from visual features for modal interaction. Finally, a task cooperation module is constructed to refine the retrieval pipeline and the highlight score prediction process by utilizing the reciprocity between MR and HD. Comprehensive experiments on QVHighlights, Charades-STA and TVSum datasets demonstrate that TR-DETR outperforms existing state-of-the-art methods. Codes are available at https://github.com/mingyao1120/TR-DETR.
Hao Sun 0014, Mingyao Zhou, Wenjing Chen 0003, Wei Xie 0008
AAAI4
2024 Segmentation Foundation Model-Aided Medical Image Segmentation
abstract
Accurate medical image segmentation is significant for reliable clinical diagnoses and pathology research. Deep learning methods rely on large high-quality dataset for supervised training to achieve satisfactory performance, but manually annotating large-scale datasets is time-consuming and costly. A recent breakthrough in the segmentation foundation model SAM shows promise for aiding annotation. However, when using SAM for annotation, errors are inevitably introduced. To utilize SAM for effective labeling, we adopt a dual-stream collaborative learning framework. Initially, a subset from the dataset is divided and then SAM is used to generate noisy labels. The first branch incorporates a dynamic weight fusion module, which adaptively fuses complementary features from model predictions and noisy labels, reconstructing more informative labels. The second branch aims to improve segmentation accuracy by training the model on an accurate subset and sharing parameters with the other branch. Our method outperforms state-of-the-art segmentation methods on the ISIC2018 and BUSI datasets.
Shiqi Hua, Dunbo Ning, Wei Xie 0008, Hao Sun 0014
BIBM3
2024 Spatial Formation-Guided Network for Group Activity Recognition
abstract
Effectively modeling the interactions among actors is critical and challenging for Group Activity Recognition (GAR). Previous methods usually divide actors into subgroups based on the similarity of appearance features for modeling multilevel interactions among actors. However, the appearance feature-based grouping scheme does not fully consider the spatial relations of actors, which can provide a discriminative clue for GAR. In this paper, we propose a Spatial Formation-Guided Network (SFGN) to capture effective interactions under the guidance of spatial formations. We first design a spatial formation extractor to excavate latent spatial relations among actors for extracting spatial formation features. Then, a formation-guided interaction module is built to utilize the spatial formation features to guide the interactions among actors. Finally, a cross-formation interaction module is further designed to explore the complementarity among diverse spatial formations. Extensive experiments on the volleyball dataset and the collective activity dataset demonstrate that SFGN outperforms the state-of-the-art methods.
Dunbo Ning, Wenjing Chen 0003, Wei Xie 0008, Hao Sun 0014
ICASSP3
2024 Cross-Modal Multiscale Difference-Aware Network for Joint Moment Retrieval and Highlight Detection
abstract
Since the goals of both Moment Retrieval (MR) and Highlight Detection (HD) are to quickly obtain the required content from the video according to user needs, several works have attempted to take advantage of the commonality between both tasks to design transformer-based networks for joint MR and HD. Although these methods achieve impressive performance, they still face some problems: a) Semantic gaps across different modalities. b) Various durations of different query-relevant moments and highlights. c) Smooth transitions among diverse events. To this end, we propose a Cross-modal Multiscale Difference-aware Network, named CMDNet. First, a clip-text alignment module is constructed to narrow semantic gaps between different modalities. Second, a multiscale difference perception module is utilized to mine the differential information between adjacent clips and perform multiscale modeling to obtain discriminative representations. Finally, these representations are fed into the MR and HD task heads to retrieve relevant moments and estimate highlight scores precisely. Extensive experiments on three popular datasets demonstrate that CMDNet achieves state-of-the-art performance.
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008
ICASSP4
2024 Spatial Dual Context Learning for Weakly-supervised Group Activity Recognition in Still-images
abstract
This paper investigates a new task, Weakly- supervised Group Activity Recognition in Still-images (WGARS), which aims to extend the applicability of Group Activity Recognition (GAR) to broader scenarios, such as low-latency domains. To tackle this challenge, we propose a Spatial Dual Context Transformer (SDCT), comprising a Dual Context Encoder (DCE) and a Dual Context Decoder (DCD). The DCE module individually encodes holistic context with integral relations of overall actors, and encodes partial context with individual features in still images. Subsequently, the DCD module explores the complementarity between holistic and partial contexts, and alternatively updates these encoded contexts to enhance the interaction of actors. Additionally, auxiliary supervised contrastive learning is incorporated to mitigate activity confusion. The proposed SDCT attains state-of-the-art performance on Volleyball and NBA datasets in WGARS. Notably, SDCT even outperforms recent methods when extended to the weakly-supervised GAR in videos task on Volleyball dataset.
Dunbo Ning, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004
ICME5
2024 Image-Centered Pseudo Label Generation for Weakly Supervised Text-Based Person Re-Identification
Weizhi Nie, Chengji Wang, Hao Sun 0014, Wei Xie 0008
PRCV (12)4
2024 An improved brain storm optimization with chunking-grouping method
abstract
Summary Brain storm optimization (BSO) is a population‐based intelligence algorithm for optimization problems, which has attracted researchers' growing attention due to its simplicity and efficiency. An improved BSO, called CIBSO, is presented in this article. First of all, a new grouping method, in which the population is partitioned into chunks according to the fitness and recombined to groups, is developed to balance each group with same quality‐level. Afterwards, a new mutation strategy is designed in CIBSO and a learning mechanism is used to adaptively select appropriate strategy. Experiments on the CEC2014 test suite indicate that CIBSO is better or at least competitive performance against the compared BSO variants.
Jinglei Guo, Shouyong Jiang, Wei Xie 0008, Zhijian Wu
Concurr. Comput. Pract. Exp.4
2024 Query-aware multi-scale proposal network for weakly supervised temporal sentence grounding in videos
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Xiaoqiang Lu
Knowl. Based Syst.4
2024 Prototype-Based Pseudo-Label Refinement for Semi-Supervised Hyperspectral Image Classification
abstract
Pseudo-label learning-based methods usually regard class confidence above a certain threshold for unlabeled samples as pseudo-labels, which may result in pseudo-labels still containing wrong labels. In this letter, we propose a prototype-based pseudo-label refinement (PPLR) for semi-supervised hyperspectral image classification. The proposed PPLR filters wrong labels from pseudo-labels using class prototypes, which can improve the discrimination of the network. First, PPLR uses multi-head attentions to extract the spectral-spatial features, and designs an adaptive threshold that can be dynamically adjusted to generate high-confidence pseudo-labels. Then, PPLR constructs class prototypes for different categories using labeled sample features and unlabeled sample features with refined pseudo-labels to improve the quality of pseudo-labels by filtering wrong labels. Finally, PPLR further assigns reliable weights to these pseudo-labels in calculating their supervised loss, and introduces a center loss to improve the discrimination of features. When 10 labeled samples per category are utilized for training, PPLR achieves the overall accuracies of 82.11%, 86.70% and 92.50% on the Indian Pines, Houston2013 and Salinas datasets, respectively.
Renyi Chen, Huaxiong Yao, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.5
2024 Orientational Clustering Learning for Open-Set Hyperspectral Image Classification
abstract
Recently, some literature has begun to pay attention to the open-set problem in remote sensing application scenarios and studied various open-set hyperspectral image classification (OSHIC) methods. These OSHIC methods are usually based on deep neural networks, using the nondirectional Euclidean distance losses to constrain latent sample representations of known classes to be compact. Nonetheless, the potential effect of the spatial distribution of sample representations is ignored, resulting in degraded classification performance in OSHIC. In this letter, we propose an orientational clustering learning (OCL) method for OSHIC. First, in the feature space generated by the convolutional neural network, a class anchor strategy is employed to bring features of the same class closer while keeping features of different classes distant. Then, we utilize the orientational learning to further tighten the intraclass feature space. OCL directionally optimizes the spatial distribution of hyperspectral sample representations to improve the ability to identify known classes and distinguish unknown classes. Experiments show that the OCL achieves overall accuracies of 94.43%, 92.27%, and 76.94% on the Pavia University, Salinas, and Indian Pines datasets, respectively.
Wenjing Chen 0003, Hailong Ning, Hao Sun 0014, Wei Xie 0008
IEEE Geosci. Remote. Sens. Lett.6
2024 Cross-Modal Feature Fusion-Based Knowledge Transfer for Text-Based Person Search
abstract
Text-based person search aims to retrieve corresponding images of person from a large gallery based on text descriptions. Existing methods strive to bridge the modality gap between images and texts and have made promising progress. However, these approaches disregard the knowledge imbalance between images and texts caused by the reporting bias. To resolve this issue, we present a cross-modal feature fusion-based knowledge transfer network to balance identity information between images and texts. First, we design an identity information emphasis module to enhance person-relevant information and suppress person-irrelevant information. Second, we design an intermediate modal-guided knowledge transfer module to balance the knowledge between images and texts. Experimental results on CUHK-PEDES, ICFG-PEDE, and RSTPReid datasets demonstrate that our method achieves state-of-the-art performance.
Kaiyang You, Wenjing Chen 0003, Chengji Wang, Hao Sun 0014, Wei Xie 0008
IEEE Signal Process. Lett.5
2023 Semi-Supervised Facial Expression Recognition by Exploring False Pseudo-Labels
abstract
Pseudo-labels are popular in semi-supervised facial expression recognition. Recent methods usually exploit the confidence as the criterion for pseudo-label generation, and utilize the high-confidence pseudo-labels as the ground-truth for training. However, high confidence cannot guarantee the correctness of pseudo-labels. False pseudo-labels can weaken the feature discrimination and degrade recognition performance. In this paper, we propose a Critical Feature Refinement Network (CFRN) to alleviate the interference of false pseudo-labels on the model performance. Specially, a feature dropout module and a feature emphasis module are proposed to improve the feature discrimination of CFRN. Then, a mean-absolute error loss is further exploited to improve the robustness against false pseudo-labels. Experimental results on three challenging datasets RAF-DB, SFEW and Affectnet demonstrate that the proposed CFRN outperforms the state-of-the-art methods.
Hao Sun 0014, Chenchen Pi, Wei Xie 0008
ICME3
2023 Efficient Dynamic Multi-key FHE Scheme from LWE for Untrusted Cloud Environments
abstract
Fully Homomorphic Encryption (FHE) provides a good solution to directly operate on the ciphertext, and the decryption result is equivalent to the corresponding operation on the plaintext. As a technique suitable for distributed environments, multi-key Fully Homomorphic Encryption (MKFHE) scheme is the most common variant of the FHE scheme since it allows encrypted data to be computed under different keys. Unfortunately, the existing dynamic MKFHE schemes based on learning with errors (LWE) still suffer from the inefficiency of long public keys, which typically grow cube in size along the lattice dimension. Moreover, there current constructions fail to provide reliable and fast algorithms to simultaneously expand ciphertexts with multiple additional keys. In order to solve the above problems, a new faster dynamic MKFHE scheme with shorter public key in asymmetric key setting from LWE is proposed in this paper, in which the size of the public key is further reduced from $\tilde O\left( {{n^3}{{(K + L)}^2}} \right)$ to $\tilde O\left( {{n^2}{{(K + L)}^2}} \right)$. In addition, our scheme cleverly adopts the dual-user cooperation method in distributed system to realize the ciphertext expansion locally, thereby reducing the computing overhead of the cloud server. More interestingly, we design a flexible parallel ciphertext expansion algorithm for the first time based on the basic algorithm. This algorithm realizes the ciphertext expansion when multiple keys are added at the same time, thus significantly improving the computational efficiency of ciphertext expansion in the dynamic MKFHE scheme. Finally, the CPA-secure of our scheme based on standard LWE assumptions is proven.
Shuchang Zeng, Jianqun Cui, Wei Xie 0008, Qihang Hou
ICPADS4
2023 Pseudolabel-Based Unreliable Sample Learning for Semi-Supervised Hyperspectral Image Classification
abstract
Recently, pseudo-label-based deep learning methods have shown excellent performance in semi-supervised hyperspectral image (HSI) classification. These methods usually select high-confidence unlabeled samples to help optimize backbone classification networks. However, a large number of remaining low-confidence unlabeled samples, which contain rich land-covers information, are underutilized. In this paper, we propose a pseudo-label-based unreliable sample learning (PUSL) method to fully exploit low-confidence unlabeled samples for semi-supervised HSI classification. Firstly, to avoid overfitting the spatial distribution of labeled samples, we build a position-free transformer (PFT) as the backbone classification network. Secondly, PFT is initially trained with labeled samples in a supervised learning manner to obtain an initial classifier, which is then used to split unlabeled samples into reliable and unreliable unlabeled samples based on the predicted confidence. Thirdly, reliable unlabeled samples participate in training along with labeled samples. Finally, unreliable unlabeled samples are treated as negative samples for corresponding categories to improve the discrimination of PFT in a contrastive learning paradigm. Extensive experiments on three HSI datasets demonstrate that PUSL outperforms compared methods.
Huaxiong Yao, Renyi Chen, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2023 Graph-aware transformer for skeleton-based action recognition
Wei Xie 0008, Chao Wang 0010, Ruide Tu, Zhigang Tu 0001
Vis. Comput.2
2022 Deformable Attention U-Shaped Network with Progressively Supervised Learning for Subarachnoid Hemorrhage Image Segmentation
abstract
Subarachnoid hemorrhage (SAH) is a common acute disease, which belongs to a subtype of intracranial hemorrhage. In this paper, a deformable attention u-shaped network (DAUN) is specially designed for SAH image segmentation. Firstly, a deformable attention module is embedded at the end of each encoding layer in Res-UNet to adaptively adjust the attention domain for alleviating the introduction of irrelevant information. Then, to improve the segmentation accuracy on irregular edges and small lesions, a region-boundary-aware loss is utilized to optimize the model. Finally, a progressively supervised learning strategy is proposed to train the proposed DAUN, which enables DAUN to find a balance between the focus on semantic information and position information of each pixel. A novel SAHCT dataset is constructed to demonstrate the performance of DAUN. In addition, the Monuseg dataset is utilized to evaluate the generalization ability of DAUN.
Hao Sun 0014, Lianghao Jin, Wei Xie 0008
BIBM3
2022 Multi-Hyperedge Hypergraph for Group Activity Recognition
abstract
Group activity recognition aims to identify group activities from the videos. Most of the previous methods focus on modeling between individuals (one-to-one), which ignores the fact that a single individual's behavior may be jointly determined by multiple individual behaviors (many-to-one). For this reason, we propose a Multi-Hyperedge Hypergraph (MHH) to capture high-order relationships between multiple people. Specifically, we build three different types of hyperedges on the hypergraph structure. Each hyperedge can accommodate the characteristics of multiple nodes to capture different types of high-order relationships between nodes. Then, we use the late fusion method to fuse the three features to further enhance the overall behavioral representation. Finally, we perform a series of experiments on two of the most widely used benchmarks in group activity recognition, which have proved the effectiveness of MHH. More importantly, as far as we know, this is the first case of using a hypergraph structure for group activity recognition.
Wanxin Li, Wei Xie 0008, Zhigang Tu 0001, Lianghao Jin
IJCNN2
2022 Multi-Part Adaptive Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
In skeleton-based action recognition task, graph convolutional network has attracted widespread attention and achieved remarkable results. However, most of the current methods are performing graph convolution on the entire skeleton graph, ignoring the fact that people are composed of different body parts. In addition, previous work ignores the temporal and spatial independence and relevance of different parts. Thus, to solve these issues, we optimize the representation of the skeleton graph, graph convolution and temporal convolution respectively. In this work, we propose multi-part adaptive graph convolution (MPA-GC) to adaptively learn the topology of each part of the body and dynamically aggregate the relevance between them. Meanwhile, we add a multi-scale temporal convolution module to better obtain temporal dimension features. Ultimately, we develop a powerful graph convolutional network named MPA-GCN, and extensive experiments on two public large-scale datasets NTU-RGB+D and NTU-RGB+D120 demonstrate the effectiveness of our module, which outperforms state-of-the-art methods.
Wei Xie 0008, Zhigang Tu 0001, Wanxin Li, Lianghao Jin
IJCNN2
2022 Video anomaly detection with spatio-temporal dissociation
Yunpeng Chang, Zhigang Tu 0001, Wei Xie 0008, Bin Luo 0005, Shifu Zhang, Haigang Sui, Junsong Yuan 0001
Pattern Recognit.3
2022 Zoom Transformer for Skeleton-Based Group Activity Recognition
abstract
Skeleton-based human action recognition has attracted increasing attention and many methods have been proposed to boost the performance. However, these methods still confront three main limitations: 1) Focusing on single-person action recognition while neglecting the group activity of multiple people (more than 5 people). In practice, multi-person group activity recognition via skeleton data is also a meaningful problem. 2) Unable to mine high-level semantic information from the skeleton data, such as interactions among multiple people and their positional relationships. 3) Existing datasets used for multi-person group activity recognition are all RGB videos involved, which cannot be directly applied to skeleton-based group activity analysis. To address these issues, we propose a novel Zoom Transformer to exploit both the low-level single-person motion information and the high-level multi-person interaction information in a uniform model structure with carefully designed Relation-aware Maps. Besides, we estimate the multi-person skeletons from the existing real-world video datasets i.e. Kinetics and Volleyball-Activity, and release two new benchmarks to verify the effectiveness of our Zoom Transfromer. Extensive experiments demonstrate that our model can effectively cope with the skeleton-based multi-person group activity. Additionally, experiments on the large-scale NTU-RGB+D dataset validate that our model also achieves remarkable performance for single-person action recognition. The code and the skeleton data are publicly available athttps://github.com/Kebii/Zoom-Transformer
Yifan Jia 0007, Wei Xie 0008, Zhigang Tu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Rotation-Invariant Attention Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification refers to identifying land-cover categories of pixels based on spectral signatures and spatial information of HSIs. In recent deep learning-based methods, to explore the spatial information of HSIs, the HSI patch is usually cropped from original HSI as the input. And 3 ×3 convolution is utilized as a key component to capture spatial features for HSI classification. However, the 3 ×3 convolution is sensitive to the spatial rotation of inputs, which results in that recent methods perform worse in rotated HSIs. To alleviate this problem, a rotation-invariant attention network (RIAN) is proposed for HSI classification. First, a center spectral attention (CSpeA) module is designed to avoid the influence of other categories of pixels to suppress redundant spectral bands. Then, a rectified spatial attention (RSpaA) module is proposed to replace 3 ×3 convolution for extracting rotation-invariant spectral-spatial features from HSI patches. The CSpeA module, the 1 ×1 convolution and the RSpaA module are utilized to build the proposed RIAN for HSI classification. Experimental results demonstrate that RIAN is invariant to the spatial rotation of HSIs and has superior performance, e.g., achieving an overall accuracy of 86.53% (1.04% improvement) on the Houston database. The codes of this work are available at https://github.com/spectralpublic/RIAN.
Xiangtao Zheng, Hao Sun 0014, Xiaoqiang Lu, Wei Xie 0008
IEEE Trans. Image Process.4
2021 Multi-Attribute Enhancement Network for Person Search
abstract
Person Search is designed to jointly solve the problems of Person Detection and Person Re-identification (Re-ID), in which the target person will be located in a large number of uncut images. Over the past few years, Person Search based on deep learning has made great progress. Visual character attributes play a key role in retrieving the query person, which has been explored in Re-ID but has been ignored in Person Search. So, we introduce attribute learning into the model, allowing the use of attribute features for retrieval task. Specifically, we propose a simple and effective model called Multi-Attribute Enhancement (MAE) which introduces attribute tags to learn local features. In addition to learning the global representation of pedestrians, it also learns the local representation, and combines the two aspects to learn robust features to promote the search performance. Additionally, we verify the effectiveness of our module on the existing benchmark dataset, CUHK-SYSU and PRW. Ultimately, our model achieves state-of-the-art among end-to-end methods, especially reaching 91.8% of mAP and 93.0% of rank-1 on CUHK-SYSU. Codes and models are available at https:// github. com/chenlq123/ MAE.
Lequan Chen, Wei Xie 0008, Zhigang Tu 0001, Jinglei Guo, Yaping Tao
IJCNN2
2020 Session-Based Recommendation Model Based on Multiple Neural Networks Hybrid Extraction Feature
abstract
The problem of session-based recommendation model aims to predict user actions based on anonymous sessions. Although, previous models achieved promising results, there are still some problems, for example, we are unable to take into account the effects of session sequences of different lengths. Generally speaking, the effect of long sequence is not as good as that of short sequence in the same model. The reason of above is that the characteristics of different length session will vary greatly. Generally, the shorter the session, the tighter the relationship between items, and the longer the session, the more likely there are items that have no relationship with each other. So, we propose a model named session-based recommendation model based on multiple neural networks hybrid extraction feature. This model uses different feature extractor to deal with the features of long sessions and short sessions respectively. In SR-MNN, we use Graph Convolutional Network to extract the features of long session and use Recurrent Neural Network to extract the features of short session. Each session is then represented as the composition of the global preference, the initial interest of that session, and the current interest of that session using an attention network. Experiments on two real datasets show that SR-MNN evidently outperforms the state-of-the-art session-based recommendation methods consistently.
Huaxiong Yao, Jiabei Hu, Wenqi Xie, Wei Xie 0008
IEEE BigData5
2020 Clustering Driven Deep Autoencoder for Video Anomaly Detection
Yunpeng Chang, Zhigang Tu 0001, Wei Xie 0008, Junsong Yuan 0001
ECCV (15)3
2020 Triangular Gaussian mutation to differential evolution
Jinglei Guo, Wei Xie 0008, Shouyong Jiang
Soft Comput.3
2019 A survey of variational and CNN-based optical flow techniques
Zhigang Tu 0001, Wei Xie 0008, Dejun Zhang, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Signal Process. Image Commun.2
2019 Semantic Cues Enhanced Multimodality Multistream CNN for Action Recognition
abstract
This paper addresses the issue of video-based action recognition by exploiting an advanced multistream convolutional neural network (CNN) to fully use semantics-derived multiple modalities in both spatial (appearance) and temporal (motion) domains, since the performance of the CNN-based action recognition methods heavily relates to two factors: semantic visual cues and the network architecture. Our work consists of two major parts. First, to extract useful human-related semantics accurately, we propose a novel spatiotemporal saliency-based video object segmentation (STS) model. By fusing different distinctive saliency maps, which are computed according to object signatures of complementary object detection approaches, a refined STS maps can be obtained. In this way, various challenges in the realistic video can be handled jointly. Based on the estimated saliency maps, an energy function is constructed to segment two semantic cues: the actor and one distinctive acting part of the actor. Second, we modify the architecture of the two-stream network (TS-Net) to design a multistream network that consists of three TS-Nets with respect to the extracted semantics, which is able to use deeper abstract visual features of multimodalities in multi-scale spatiotemporally. Importantly, the performance of action recognition is significantly boosted when integrating the captured human-related semantics into our framework. Experiments on four public benchmarks-JHMDB, HMDB51, UCF-Sports, and UCF101-demonstrate that the proposed method outperforms the state-of-the-art algorithms.
Zhigang Tu 0001, Wei Xie 0008, Justin Dauwels, Baoxin Li, Junsong Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Multi-stream CNN: Learning representations based on human-related regions for action recognition
Zhigang Tu 0001, Wei Xie 0008, Qianqing Qin, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Pattern Recognit.2
2017 Variational method for joint optical flow estimation and edge-aware image restoration
Zhigang Tu 0001, Wei Xie 0008, Coert Van Gemeren, Ronald Poppe, Remco C. Veltkamp
Pattern Recognit.2
2017 Fusing disparate object signatures for salient object detection in video
Zhigang Tu 0001, Zuwei Guo, Wei Xie 0008, Mengjia Yan 0003, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Pattern Recognit.3
2015 Dissipative differential evolution with self-adaptive control parameters
abstract
Differential evolution (DE) is one of the most powerful and effective evolutionary algorithms for the global optimization problems. However, the performance of DE highly depends on control parameters. To solve this problem, dissipative differential evolution with self-adaptive control parameters (DSDE) is proposed in this paper. In DSDE approach, the values of control parameters are adjusted by the fitness information between the target vector and trial vector. Because the population diversity is a key to avoid falling into the local optima, DSDE develops dissipative scheme to make the population far away equilibrium state. Experimental studies on comprehensive set of benchmark functions show DSDE achieves better results for the majority of test cases.
Jinglei Guo, Wei Xie 0008
CEC3