EDBT 2026 Demo / reviewers in the wild / expert
Xin Wang 0068
dblp:10/5630-68
· DBLP profile ↗
14ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0003-0203-9964ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Texture, Shape and Order Matter: A New Transformer Design for Sequential DeepFake DetectionabstractSequential DeepFake detection is an emerging task that predicts the manipulation sequence in order. Existing methods typically formulate it as an image-to-sequence problem, employing conventional Transformer architectures. However, these methods lack dedicated design and consequently result in limited performance. As such, this paper describes a new Transformer design, called TSOM, by exploring three perspectives: Texture, Shape, and Order of Manipulations. Our method features four major improvements: we describe a new texture-aware branch that effectively captures subtle manipulation traces with a Diversiform Pixel Difference Attention module. Then we introduce a Multi-source Cross-attention module to seek deep correlations among spatial and sequential features, enabling effective modeling of complex manipulation traces. To further enhance the cross-attention, we describe a Shape-guided Gaussian mapping strategy, providing initial priors of the manipulation shape. Finally, observing that the subsequent manipulation in a sequence may influence traces left in the preceding one, we intriguingly invert the prediction order from forward to backward, leading to notable gains as expected. Extensive experimental results demonstrate that our method outperforms others by a large margin, highlighting the superiority of our method. Yuezun Li, Xin Wang 0068, Baoyuan Wu, Jiaran Zhou, Junyu Dong |
WACV | 3 |
| 2025 | ORENet: Oriented Rotation Equivariant Network for Remote Sensing Object DetectionabstractCurrent state-of-the-art methods use oriented bounding boxes instead of horizontal bounding boxes for oriented object detection in remote sensing images (RSIs). However, this introduces several challenges, such as imprecise rotated feature extraction, inefficient oriented region proposal generation, unstable rotated bounding box representation, and inconsistency between loss functions and evaluation metrics. To address these challenges, we propose an oriented rotation equivariant network called ORENet. First, to achieve accurate feature extraction, an equivariant steerable module (ESM) is designed to guide feature learning for arbitrarily oriented objects. Second, an oriented efficient region proposal network (OERPN) is developed to improve the efficiency of proposal generation. Third, a center-offset six-parameter representation is adopted for robust bounding box modeling. Finally, a Gaussian distance loss module (GDLM) is introduced to tackle the inconsistency issue. Detailed ablation studies validate the effectiveness of the proposed components. Compared with existing state-of-the-art approaches, our method achieves substantially reduced computational complexity while maintaining superior performance, establishing it as a highly efficient lightweight framework. The source code is publicly available at: https://github.com/WangXin81/ORENet2.0/. Xin Wang 0068, Zhilu Zhang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | CT-PromptSAM: Collaborative Prompting With Hybrid CNN-Transformer for Remote Sensing Semantic SegmentationabstractRemote sensing (RS) semantic segmentation (SS) faces critical challenges in small-object detection and boundary precision due to scale variation and complex scenes. This letter proposes CT-PromptSAM, a specialist-generalist framework integrating a CNN-Transformer specialist network and a SAM-adaptive generalist network. The specialist network employs a tandem-encoder for local-global feature fusion, coupled with three decoder modules: Up-sampling Feature Refinement Module (UFRM), Multi-scale Feature Interaction Module (MFIM), and Mutual-guided Semantic-Spatial Attention Module (MSSAM), to generate task-adaptive masks for small-class enhancement. The generalist network adapts Segment Anything Model (SAM) to RS via multi-modal prompts, forming a “local-global-boundary” collaborative loop for boundary refinement. On Vaihingen and Potsdam datasets, CT-PromptSAM achieves state-of-the-art small-class IoU and mean Boundary F1-score, outperforming competitors with comparable parameter efficiency. By fusing RS-specific semantic priors (from the specialist) and SAM’s geometric precision (from the generalist), the framework balances small-object and boundary tasks, validating its efficacy in complex RS scenes. Our code is publicly available at: https://github.com/WangXin81/CT-PromptSAM/. Xin Wang 0068 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | TBRNet: Two-Branch Reinforcement Network for Few-Shot Semantic Segmentation in Remote Sensing ImagesabstractTraditional semantic segmentation methods for remote sensing images (RSIs) require abundant labeled data yet falter when samples are scarce. Few-shot semantic segmentation (FSS) innovatively resolves this data scarcity bottleneck. However, existing FSS models face unique challenges in RSIs: natural-to-remote-sensing domain gaps, intraclass variance, and multicategory coexistence-induced generalization collapse. These problems may lead to the loss of discriminative query features in model performance. In this article, we present a novel two-branch reinforcement network called TBRNet to tackle these challenges. Specifically, we first propose a prototype reinforcement module (PRM) to generate enhanced context-adaptive prototypes by dynamically weighting query-support feature contexts, which effectively mitigates intraclass variance via strengthening support images’ perception of discriminative query features. In addition, to deal with the challenge of the coexistence of multiple target categories, we develop a multilevel guidance reinforcement module (MGRM), which provides multilevel guidance maps across resolutions to model cross-level semantic dependencies and emphasize discriminative subregions. Extensive experiments conducted on the iSAID-$5^{i}$and DLRSD-$5^{i}$remote sensing (RS) datasets have shown the superiority of our proposed TBRNet compared with several state-of-the-art approaches. Ablation studies have also verified the effectiveness of the proposed modules. Our source code is made publicly available athttps://github.com/WangXin81/TBRNet Xin Wang 0068, Huiyu Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Attention-Aware Three-Branch Network for Salient Object Detection in Remote Sensing ImagesabstractAlthough remarkable advances have been achieved on salient object detection (SOD) for natural scene images (NSIs), SOD for optical remote sensing images (RSIs) still remains a big challenge due to the unique imaging conditions and various scene patterns. To enable effective SOD for RSIs, this letter proposes a novel end-to-end network, called attention-aware three-branch network (AATBNet). First, an attention feature encoding branch is constructed for learning more discriminative features. Then, a hierarchical feature decoding branch, equipped with three streams, i.e., a decoding stream, a dilated reverse attention stream, and a fusion dense up-sampling convolution stream, is proposed to effectively and robustly compute saliency maps and salient edge maps. Third, a two losses computation branch is designed to further boost SOD performance. Comprehensive evaluations on two well-known RSIs benchmarks, as well as comparisons with 20 state-of-the-art technologies validate the superiority of our AATBNet. The code of our method is publicly available at: https://github.com/WangXin81/AATBNet. Xin Wang 0068, Zhilu Zhang 0003, Shihan Jing, Huiyu Zhou 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Multilevel Feature Fusion Networks With Adaptive Channel Dimensionality Reduction for Remote Sensing Scene ClassificationabstractScene classification in very high-resolution (VHR) remote sensing (RS) images is a challenging task due to the complex and diverse content of the images. Recently, convolution neural networks (CNNs) have been utilized to tackle this task. However, CNNs cannot fully meet the needs of scene classification due to clutters and small objects in VHR images. To handle these challenges, this letter presents a novel multilevel feature fusion (MLFF) network with adaptive channel dimensionality reduction for RS scene classification. Specifically, an adaptive method is designed for channel dimensionality reduction of high-dimensional features. Then, an MLFF module is introduced to fuse the features in an efficient way. Experiments on three widely used data sets show that our model outperforms several state-of-the-art methods in terms of both accuracy and stability. Xin Wang 0068, Lin Duan, Aiye Shi, Huiyu Zhou 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Dropout-Based Adversarial Training Networks for Remote Sensing Scene ClassificationabstractScene classification in remote sensing (RS) images is a challenging task due to the lack of well-labeled data. Recently, deep transfer learning (DTL) has been proposed to handle this task. However, the intraclass variations and interclass similarities remain challenges. To handle these challenges, this letter presents a novel dropout-based adversarial training network (DATN) for RS scene classification. Specifically, a dropout-based label classifier (DLC) module is designed to reduce the selection of ambiguous features on class boundaries. Then, a dropout-based domain discriminator (DDD) module is constructed to capture multimodal structures of RS images so as to achieve fine-grained alignment between cross-domain distributions. Third, a joint distribution of features and labels is built to further enhance the performance. Experiments on seven public RS datasets show that our model outperforms several states of the art (SOTAs) under different conditions. The code of our method is publicly available athttps://github.com/WangXin81/DATN-Submitted-to-IEEE-GRSL. Xin Wang 0068, Zhipeng Mao, Aiye Shi, Zhilu Zhang 0003, Huiyu Zhou 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | SWIPENET: Object detection in noisy underwater scenesabstractDeep learning based object detection methods have achieved promising performance in controlled environments. However, these methods lack sufficient capabilities to handle underwater object detection due to these challenges: (1) images in the underwater datasets and real applications are blurry whilst accompanying severe noise that confuses the detectors and (2) objects in real applications are usually small. In this paper, we propose a Sample-WeIghted hyPEr Network (SWIPENET), and a novel training paradigm named Curriculum Multi-Class Adaboost (CMA), to address these two problems at the same time. Firstly, the backbone of SWIPENET produces multiple high resolution and semantic-rich Hyper Feature Maps, which significantly improve small object detection. Secondly, inspired by the human education process that drives the learning from easy to hard concepts, we propose the noise-robust CMA training paradigm that learns the clean data first and then move on to learns the diverse noisy data. Experiments on four underwater object detection datasets show that the proposed SWIPENET+CMA framework achieves better or competitive accuracy in object detection against several state-of-the-art approaches. Long Chen 0019, Feixiang Zhou, Shengke Wang, Junyu Dong, Ning Li 0012, Haiping Ma, Xin Wang 0068, Huiyu Zhou 0001 |
Pattern Recognit. | 7 |
| 2022 | Multimodal Gait Recognition for Neurodegenerative DiseasesabstractIn recent years, single modality-based gait recognition has been extensively explored in the analysis of medical images or other sensory data, and it is recognized that each of the established approaches has different strengths and weaknesses. As an important motor symptom, gait disturbance is usually used for diagnosis and evaluation of diseases; moreover, the use of multimodality analysis of the patient's walking pattern compensates for the one-sidedness of single modality gait recognition methods that only learn gait changes in a single measurement dimension. The fusion of multiple measurement resources has demonstrated promising performance in the identification of gait patterns associated with individual diseases. In this article, as a useful tool, we propose a novel hybrid model to learn the gait differences between three neurodegenerative diseases, between patients with different severity levels of Parkinson's disease, and between healthy individuals and patients, by fusing and aggregating data from multiple sensors. A spatial feature extractor (SFE) is applied to generating representative features of images or signals. In order to capture temporal information from the two modality data, a new correlative memory neural network (CorrMNN) architecture is designed for extracting temporal features. Afterward, we embed a multiswitch discriminator to associate the observations with individual state estimations. Compared with several state-of-the-art techniques, our proposed framework shows more accurate classification results. Aite Zhao, Junyu Dong, Lin Qi 0004, Qianni Zhang, Ning Li 0012, Xin Wang 0068, Huiyu Zhou 0001 |
IEEE Trans. Cybern. | 7 |
| 2022 | Recurrent Attention and Semantic Gate for Remote Sensing Image CaptioningabstractThe remote sensing image captioning has attracted wide spread attention in remote sensing field due to its application potentiality. However, most existing approaches model limited interactions between image content and sentence and fail to exploit special characteristics of the remote sensing images. We introduce a novel recurrent attention and semantic gate (RASG) framework to facilitate the remote sensing image captioning in this article, which integrates competitive visual features and a recurrent attention mechanism to generate a better context vector for the images every time as well as enhances the representations of the current word state. Specifically, we first project each image into competitive visual features by taking the advantage of both static visual features and multiscale features. Then, a novel recurrent attention mechanism is developed to extract the high-level attentive maps from encoded features and nonvisual features, which can help the decoder recognize and focus on the effective information for understanding the complex content of the remote sensing images. Finally, the hidden states from the long short-term memory (LSTM) and other semantic references are incorporated into a semantic gate, which contributes to more comprehensive and precise semantic understanding. Comprehensive experiments on three widely used datasets, Sydney-Captions, UCM-Captions, and Remote Sensing Image Captioning Dataset, have demonstrated the superiority of the proposed RASG over a series of attentive models based on image captioning methods. Yunpeng Li 0010, Xiangrong Zhang, Chen Li 0011, Xin Wang 0068, Xu Tang 0004, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | An Object Recognition Approach for Synthetic Aperture Radar Images
Chen Ning, Wenbo Liu 0001, Gong Zhang 0002, Xin Wang 0068 |
Mob. Networks Appl. | 4 |
| 2021 | An Effective Algorithm for Single Image Fog Removal
Xin Wang 0068, Hangcheng Zhu, Chen Ning |
Mob. Networks Appl. | 1 |
| 2021 | Enhanced Feature Pyramid Network With Deep Semantic Embedding for Remote Sensing Scene ClassificationabstractRecent progress on remote sensing (RS) scene classification is substantial, benefiting mostly from the explosive development of convolutional neural networks (CNNs). However, different from the natural images in which the objects occupy most of the space, objects in RS images are usually small and separated. Therefore, there is still a large room for improvement of the vanilla CNNs that extract global image-level features for RS scene classification, ignoring local object-level features. In this article, we propose a novel RS scene classification method via enhanced feature pyramid network (EFPN) with deep semantic embedding (DSE). Our proposed framework extracts multiscale multilevel features using an EFPN. Then, to leverage the complementary advantages of the multilevel and multiscale features, we design a DSE module to generate discriminative features. Third, a feature fusion module, called two-branch deep feature fusion (TDFF), is introduced to aggregate the features at different levels in an effective way. Our method produces state-of-the-art results on two widely used RS scene classification benchmarks, with better effectiveness and accuracy than the existing algorithms. Beyond that, we conduct an exhaustive analysis on the role of each module in the proposed architecture, and the experimental results further verify the merits of the proposed method. Xin Wang 0068, Chen Ning, Huiyu Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Scene Attention Mechanism for Remote Sensing Image Caption GenerationabstractRemote sensing images play an important role in various applications. To make it easier for humans to understand remote sensing images, the task of remote sensing image captioning attracts more and more researchers' attention. Inspired from the way human receives visual information, attention mechanism has been widely used in remote sensing image understanding. To catch more scene information and improve the stability of the generated sentences, a new attention mechanism called scene attention is proposed. Except for the current attention via the current hidden state of the long shortterm memory network (LSTM), our proposed method simultaneously explores the global visual information from the mean feature of all convolutional features. The effectiveness of the proposed method is evaluated on UCM-captions, Sydney-captions and RSICD datasets. The results of our experiment show that comparing with some other captioning methods, our method is more stable and obtains a better performance. Shiqi Wu, Xiangrong Zhang, Xin Wang 0068, Chen Li 0011, Licheng Jiao |
IJCNN | 3 |