VLDB 2026 Research / reviewers in the wild / expert
Xianbin Wen
dblp:94/5397
· DBLP profile ↗
36ranked-venue papers
1as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 13 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EFIN: A Novel Enhanced Feature Interaction Network for Temporal Sentence Grounding in VideosabstractTemporal sentence grounding in videos (TSGV) is a challenging task that aims to match text queries with semantically relevant segments in untrimmed videos. However, existing methods face limitations in modeling modality features, which constrains the expressive power of candidate moment features. To address this challenge, we propose a novel Enhanced Feature Interaction Network (EFIN) that effectively captures semantic information within each modality and aligns relationships between modalities. Additionally, EFIN enhances the fusion of information between candidate moments and modality features. Specifically, our model begins by extracting modality features to generate candidate moments as priors. Building upon these modality features, we introduce an enhanced feature encoder to extract semantic information within each modality, thereby improving intra-modality feature representation. Simultaneously, the encoder captures alignment relationships between modalities to optimize cross-modality feature representation, enhancing the overall modeling capacity of modality features. Moreover, we design an information fusion module to enrich the comprehension of modality information for candidate moments. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed EFIN model. Notably, EFIN achieves a maximum performance improvement of approximately 1.67% and 1.91% across different evaluation metrics on TACoS dataset. Chongxu Hu, Xianbin Wen, Yibo Zhao 0001, Chunjie Ma, Weili Guan, Riwei Wang, Zan Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | ACDR-CRAFF Net: A Multi-Scale Network Based on Adaptive Channel and Coordinate Relational Attention Network for Remote Sensing Scene ClassificationabstractABSTRACT Accurate classification of remote sensing scene images is crucial for diverse applications, from environmental monitoring to urban planning. While convolutional neural networks (CNNs) have dramatically improved classification accuracy, challenges remain due to the complex distribution of small objects, varied spatial configurations, and intra‐class multimodality in remote sensing images. In this work, we make three key contributions to address these challenges. (1) We propose the adaptive channel and coordinate relational attention network (ACDR‐CRAFF), a novel multi‐scale feature fusion framework designed to enhance feature representation across scales. (2) We introduce two innovative modules: the adaptive channel dimensionality reduction (ACDR) module, which dynamically adjusts channel representations to retain essential low‐dimensional features, and the coordinate relational attention multi‐scale feature fusion (CRAFF) module, which effectively captures and transfers spatial information between feature levels. (3) By integrating ACDR and CRAFF, our model achieves a progressive fusion of local to global features, ensuring robust feature expressiveness at multiple scales. Experimental results on four widely used benchmark datasets demonstrate that ACDR‐CRAFF consistently outperforms several state‐of‐the‐art methods, achieving significant improvements in classification accuracy and setting a new benchmark for complex remote sensing scene classification tasks. These results underscore the effectiveness of our approach in addressing the limitations of existing methods and advancing the state of the art in remote sensing image analysis. Haixia Xu 0003, Furong Shi, Xianbin Wen |
IET Image Process. | 6 |
| 2025 | PLDF-S3: Pseudo-Label-Driven Framework for Offshore to Inshore Unsupervised SAR Image Ship SegmentationabstractRecently, unsupervised ship segmentation methods for Synthetic Aperture Radar (SAR) images have achieved promising results in offshore scenes. However, these methods generate a large number of false alarms in inshore scenes. To address this issue, we propose the pseudo-label driven framework for offshore to inshore SAR image ship segmentation (PLDF-S3), which leverages ship segmentation results from offshore scenes to assist inshore ship segmentation. In particular, to account for the anisotropy of ships, which are characterized by a dominant long-axis direction, we design a directional feature enhancement module (DFEM) in PLDF-S3to extract ship features with varying orientations. Additionally, due to the diverse size variations of ships in SAR images, we propose a hierarchical context enhancement module (HCEM) to capture ship features at different scales. Experimental results show that the proposed unsupervised PLDF-S3achieves comparable segmentation performance than several supervised methods under challenging inshore scenarios. Furong Shi, Xianbin Wen, Jiao Liu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Hfffap-net: unsupervised fundus image enhancement with high-frequency feature fusion and artifact processing
Xiaoming Yang, Xianbin Wen |
Multim. Syst. | 3 |
| 2025 | CACP: Covariance-Aware Cross-Domain Prototypes for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to reduce domain shifts / discrepancies between source and target domains, improving the source domain model's generalization ability to the target domain. Recently, prototypical methods, which primarily use single-source or single-target domain prototypes as category centers to aggregate features from both domains, have achieved competitive performance in this task. However, due to large domain shifts, single-source domain prototypes have finite generalization ability and not all source domain knowledge is conducive to model generalization. Single-target domain prototypes are noisy because they are prematurely initialized with all features filtered by pseudo labels, which causes error accumulation in the prototypes. To address these issues, we propose a covariance-aware cross-domain prototypes method (CACP) to achieve robust domain adaptation. We propose to use both domain prototypes to dynamically rectify pseudo labels in the target domain, effectively reducing the recognition difficulty of hard target domain samples and narrowing the gap between features of the same category in both domains. In addition, to further generalize the model to the target domain, we propose two modules based on covariance correlation, FSPC (Features Selection by Prototypes Covariances) and WSPC (Weighting Source by Prototypes Coefficients), to learn discriminative characteristics. FSPC selects highly correlated features to update target domain prototypes online, denoising and enhancing discriminativeness between categories. WSPC utilizes the correlation coefficients between target domain prototypes and source domain features to weight each point in the source domain, eliminating the information interference from the source domain. In particular, CACP achieves excellent performance on the GTA5$\to$Cityscapes and SYNTHIA$\to$Cityscapes tasks with minimal computational resources and time. Yanbing Xue, Feifei Zhang 0001, Xianbin Wen, Zan Gao 0002, Shengyong Chen |
IEEE Trans. Multim. | 4 |
| 2024 | Visual Tracking based on deformable Transformer and spatiotemporal information
Ruixu Wu, Xianbin Wen, Yanli Liu 0005 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | LPNet: A remote sensing scene classification method based on large kernel convolution and parameter fusionabstractAbstract Remote sensing scene images contain numerous feature targets with unrelated semantic information, so how to extract to the local key information and semantic features of the image becomes the key to achieving accurate classification. Existing Convolutional Neural Networks (CNNs) mostly concentrate on the global representation of an image and lose the shallow features. To overcome these issues, this paper proposes LPNet for remote sensing scene image classification. First, LPNet employs LKConv to extract the semantic features in the image, while using standard convolution to extract local key information in the image. Additionally, the LPNet applies a shortcut residual concatenation branch to reuse features. Then, parameter fusion combines parameters from previous branches, improving the capacity of the model to obtain a more comprehensive and rich feature representation of the image. Finally, considering the relationship between the classification ability of the model and the depth of feature extraction, the Feature Mixture (FM) Block is used to deepen the model for feature extraction. Comparative experiments on four publicly available datasets show that LPNet provides comparable results to other state‐of‐the‐art methods. The effectiveness of LPNet is further demonstrated by visualizing the effective receptive fields (ERFs). Furong Shi, Haixia Xu 0003, Xianbin Wen |
IET Image Process. | 6 |
| 2024 | X-CDNet: A real-time crosswalk detector based on YOLOX
Xingyuan Lu, Yanbing Xue, Xianbin Wen |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | Multiscale Attention-Based Subdomain Dynamic Adaptation for Cross-Domain Scene ClassificationabstractCross-domain scene classification is proposed to address the problems of difficult labeling of remote-sensing (RS) image datasets and poor generalization ability of supervised models, aiming to better utilize existing knowledge. To minimize the difference between the source- and target-domain distributions, many deep-domain adaptation methods have been generated, but most of them are based on the difference metric function, which only globally aligns the edge distributions without taking into account the effect of each sample on the network in different and the relationship between related subdomains in different domains of the same category, resulting in a large amount of fine-grained information being lost. In addition, existing domain adaptation methods do not adaptively balance the weights of the marginal and conditional distributions well. To overcome the above difficulties, a multiscale attention-based subdomain dynamic adaptation (SAMRA) method is proposed. The relative importance of the two is balanced by calculating the dynamic weights of each sample in the different domains to globally adjust the marginal distribution and by obtaining finer-grained key information from the subdomains to locally adjust the conditional distribution. Moreover, multiscale feature generation and attention mechanisms support the extraction of more robust features and more complete information, as well as more purposeful transfer. The results of cross-domain experiments show that the improvement of SAMRA over state-of-the-art deep-domain adaptation methods is significant. Furong Shi, Xianbin Wen |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | A lightweight and stochastic depth residual attention network for remote sensing scene classificationabstractAbstract Due to the rapid development of satellite technology, high‐spatial‐resolution remote sensing (HRRS) images have highly complex spatial distributions and multiscale features, making the classification of such images a challenging task. The key to scene classification is to accurately understand the main semantic information contained in images. Convolutional neural networks (CNNs) have outstanding advantages in this field. Deep CNNs (D‐CNNs) with better performance tend to have more parameters and higher complexity. However, shallow CNNs have difficulty extracting the key features of complex remote sensing images. In this paper, we propose a lightweight network with a random depth strategy for remote sensing scene classification (LRSCM). We construct a convolutional feature extraction module, DCAB, which incorporates depthwise separable convolutional and inverted residual structures, effectively reducing the numbers of required parameters and computations, and retains and utilizes low‐level features. In addition, coordinate attention (CA) is integrated into the module, thereby further improving the network's ability to extract key local information. To further reduce the complexity of model training, the residual module adopts a stochastic depth strategy, providing the network with a random depth. Comparative experiments on five public datasets show that the LRSCM network can achieve results comparable to those of other state‐of‐the‐art methods. Haixia Xu 0003, Xianbin Wen |
IET Image Process. | 4 |
| 2023 | Pose-guided feature region-based fusion network for occluded person re-identification
Gengsheng Xie, Xianbin Wen, Jianchen Wang, Changlun Guo, Yansong Jia |
Multim. Syst. | 2 |
| 2023 | Generalized attention-based deep multi-instance learning
Kun Hao, Xianbin Wen |
Multim. Syst. | 4 |
| 2023 | RMHNet: A Relation-Aware Multi-granularity Hierarchical Network for Person Re-identification
Gengsheng Xie, Xianbin Wen |
Neural Process. Lett. | 2 |
| 2023 | The Impact of Arabic Diacritization on Word EmbeddingsabstractWord embedding is used to represent words for text analysis. It plays an essential role in many Natural Language Processing (NLP) studies and has hugely contributed to the extraordinary developments in the field in the last few years. In Arabic, diacritic marks are a vital feature for the readability and understandability of the language. Current Arabic word embeddings are non-diacritized. In this article, we aim to develop and compare word embedding models based on diacritized and non-diacritized corpora to study the impact of Arabic diacritization on word embeddings. We propose evaluating the models in four different ways: clustering of the nearest words; morphological semantic analysis; part-of-speech tagging; and semantic analysis. For a better evaluation, we took the challenge to create three new datasets from scratch for the three downstream tasks. We conducted the downstream tasks with eight machine learning algorithms and two deep learning algorithms. Experimental results show that the diacritized model exhibits a better ability to capture syntactic and semantic relations and in clustering words of similar categories. Overall, the diacritized model outperforms the non-diacritized model. We obtained some more interesting findings. For example, from the morphological semantics analysis, we found that with the increase in the number of target words, the advantages of the diacritized model are also more obvious, and the diacritic marks have more significance in POS tagging than in other tasks. Mohamed Abbache, Ahmed Abbache, Farid Meziane, Xianbin Wen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2022 | Deep Multi-Instance Learning with Adaptive Recurrent Pooling for Medical Image ClassificationabstractRecently, deep multi-instance neural networks have been successfully applied for medical image classification, where only image-level labels rather than fine-grained patch-level labels are available for use. One key issue for these multi-instance neural networks is how to aggregate all patch (instance) features into an entire image (bag) representation, referred to as multi-instance pooling, e.g., max, mean, and attention based pooling. Nevertheless, these multi-instance pooling operations do not take the structural information within an image into account. This is obviously inappropriate for medical image classification since there often exist certain dependencies among regional patches/lesions. We propose an adaptive recurrent pooling based deep multi-instance neural network in this paper. In this network, we first extract multi-view global structural features from every bag using the self-attention mechanism, and then aggregate these multi-view features into a whole bag representation based on the adaptive recurrent pooling operation in order to further capture the contextual information within the bag. Moreover, we introduce the cross-normalization operation used in the Unit Force Operated Vision Transformer into the self-attention module to reduce its computational complexity. We have experimentally evaluated the performance of the proposed network on three medical image datasets, namely UCSB breast, Messidor, and Colon cancer. The results demonstrate the advantage of our network over current state-of-the-art deep multi-instance networks in terms of classification accuracy and interpretability. Xianbin Wen |
BIBM | 4 |
| 2022 | Multi-View Representation Learning for Multi-Instance Learning with Applications to Medical Image ClassificationabstractMulti-Instance Learning (MIL) is a weakly supervised learning paradigm, in which every training example is a labeled bag of unlabeled instances. In typical MIL applications, instances are often used for describing the features of regions/parts in a whole object, e.g., regional patches/lesions in an eye-fundus image. However, for a (semantically) complex part the standard MIL formulation puts a heavy burden on the representation ability of the corresponding instance. To alleviate this pressure, we still adopt a bag-of-instances as an example in this paper, but extract from each instance a set of representations using $1 \times1$ convolutions. The advantages of this tactic are two-fold: i) This set of representations can be regarded as multi-view representations for an instance; ii) Compared to building multi-view representations directly from scratch, extracting them automatically using $1 \times1$ convolutions is more economical, and may be more effective since $1 \times1$ convolutions can be embedded into the whole network. Furthermore, we apply two consecutive multi-instance pooling operations on the reconstituted bag that has actually become a bag of sets of multi-view representations. We have conducted extensive experiments on several canonical MIL data sets from different application domains. The experimental results show that the proposed framework outperforms the standard MIL formulation in terms of classification performance and has good interpretability. Zhenliang Li, Xianbin Wen |
BIBM | 4 |
| 2022 | Global Correlative Network for Person re-identification
Gengsheng Xie, Xianbin Wen, Haixia Xu 0003, Zhanlu Liu |
Neurocomputing | 2 |
| 2022 | DASFTOT: Dual attention spatiotemporal fused transformer for object tracking
Ruixu Wu, Xianbin Wen, Haixia Xu 0003 |
Knowl. Based Syst. | 2 |
| 2022 | STASiamRPN: visual tracking based on spatiotemporal and attention
Ruixu Wu, Xianbin Wen, Zhanlu Liu, Haixia Xu 0003 |
Multim. Syst. | 2 |
| 2022 | Correction: STASiamRPN: visual tracking based on spatiotemporal and attention
Ruixu Wu, Xianbin Wen, Zhanlu Liu, Haixia Xu 0003 |
Multim. Syst. | 2 |
| 2022 | Multiview clustering via consistent and specific nonnegative matrix factorization with graph regularization
Haixia Xu 0003, Limin Gong, Hai-Zhen Xuan, Xusheng Zheng, Zan Gao 0001, Xianbin Wen |
Multim. Syst. | 6 |
| 2022 | Multiple Regularization and Analysis of Deep Capsule Network
Xianbin Wen |
Pattern Anal. Appl. | 4 |
| 2021 | Multi-Scale Attention Network Based on Multi-Feature Fusion for Person Re-IdentificationabstractAs a challenging and high practical research topic in public safety, person re-identification (Re-ID) technology has attracted increasing attention in the field of computer vision. Due to the success of deep learning, Convolutional Neural Networks (CNNs) have become the main techniques to extract discriminative features for person Re-ID++. However, there are many problems in the real scene of pedestrian images, such as pedestrian posture changes, inconsistent shooting perspective and object occlusion, etc. The global features extracted from images by CNNs are easily disturbed by these problems, resulting in the lack of robustness and discrimination, which further leads to low recognition accuracy. To solve these issues, we propose a multi-scale attention network based on multi-feature fusion (MSAN), which adopts a multi-branch deep network structure consisting of a global feature learning branch, two local feature learning branches and a shallow-level feature learning branch. It can sample the features of different depth of the network and get discriminative feature embedding by combining global and local cues, and then the sampled features are fused to predict the pedestrians. We also use the attention mechanism to make the network to focus on the key information of different scale feature maps thus enhance the learning of key parts of the human body and alleviate the interference caused by image changes. Experimental results on three mainstream benchmark datasets Market-1501, DukeMTMC-reID and CUHK03 show that our method can significantly improve the performances and outperform most mainstream methods. Xianbin Wen, Jianchen Wang, Gengsheng Xie, Yansong Jia |
IJCNN | 3 |
| 2021 | Pairwise attention network for cross-domain image recognition
Zan Gao 0002, Guangping Xu, Xianbin Wen |
Neurocomputing | 4 |
| 2021 | Channel-exchanged feature representations for person re-identification
Jianchen Wang, Haixia Xu 0003, Gengsheng Xie, Xianbin Wen |
Inf. Sci. | 5 |
| 2021 | Non-full multi-layer feature representations for person re-identification
Jianchen Wang, Jianguang Zhang, Xianbin Wen |
Multim. Tools Appl. | 3 |
| 2021 | Dense capsule networks with fewer parameters
Xianbin Wen, Haixia Xu 0003 |
Soft Comput. | 2 |
| 2020 | Learnable Bag Similarity Based Deep Multi-Instance Network for Breast Cancer DiagnosisabstractComputer-aided diagnosis for breast cancer is an important and challenging research problem in medical image analysis. The main difficulty lies in that only image-level labels rather than fine-grained patch-level labels can be gotten for medical images in general. This situation fits well with the settings of Multi-Instance Learning (MIL). Following this line of research, Bag Similarity Network (BSN) uses the inter-bag similarities to learn the relationship between bags (images) and achieves a good performance in automatic diagnosis of breast cancer. Nevertheless, the inter-bag similarities are pre-defined rather than learnable. In this paper, we propose a Learnable Bag Similarity Network for deep MIL, called LBSN, to aid breast cancer diagnosis. To implement automatic similarity learning, LBSN first extracts fixed numbers of global representations of each bag using the attention mechanism, and then employs channel and spatial attention on similarity matrices for learning the relation of bags and that of instances, respectively. The experimental results on a publicly available breast cancer data set demonstrate that the proposed LBSN outperforms BSN by a large margin in terms of classification accuracy. Rui Cheng 0008, Haixia Xu 0003, Zhenliang Li, Xianbin Wen |
BIBM | 5 |
| 2020 | Deep Multi-Instance Learning with Induced Self-Attention for Medical Image ClassificationabstractExisting Multi-Instance learning (MIL) methods for medical image classification typically segment an image (bag) into small patches (instances) and learn a classifier to predict the label of an unknown bag. Most of such methods assume that instances within a bag are independently and identically distributed. However, instances in the same bag often interact with each other. In this paper, we propose an Induced SelfAttention based deep MIL method that uses the self-attention mechanism for learning the global structure information within a bag. To alleviate the computational complexity of the naive implementation of self-attention, we introduce an inducing point based scheme into the self-attention block. We show empirically that the proposed method is superior to other deep MIL methods in terms of performance and interpretability on three medical image data sets. We also employ a synthetic MIL data set to provide an intensive analysis of the effectiveness of our method. The experimental results reveal that the induced self-attention mechanism can learn very discriminative and different features for target and non-target instances within a bag, and thus fits more generalized MIL problems. Zhenliang Li, Haixia Xu 0003, Rui Cheng 0008, Xianbin Wen |
BIBM | 5 |
| 2020 | Local-Variance-Based Attention For Visual TrackingabstractThe RoIAlign module incorporated into the deep tracking-by-detection framework, which can thus receive the entire image as the input to the convolutional layer, alleviating high computational complexity induced by multiple proposals. Nevertheless, this would also produce an ambiguous feature discriminative boundary between the target and background in the feature map, which makes the following target identification and localization very difficult. To solve this problem, we apply a novel local-variance-based regularization for optimizing the convolutional layer, the local variance calculated from the attention map, i.e., the average pooling of the convolutional feature map. Therefore, the binary classification loss function integrated with local-variance-based regularization item can explicitly make the response of the target and background very distinguishable, specifically strengthening the response of target and weakening that of background. Extensive experiments on large-scale benchmark data sets demonstrate that the proposed algorithm is highly comparable to other state-of-the-art methods. Changlun Guo, Xianbin Wen, Haixia Xu 0003 |
ICME | 2 |
| 2020 | Integrating aspect-aware interactive attention and emotional position-aware for multi-aspect sentiment analysisabstractAspect-level Sentiment Analysis is a fine-grained sentiment analysis task, which aims to infer the corresponding sentiment polarity with different aspects in an opinion sentence. Attention-based neural networks have proven to be effective in extracting aspect terms, but the prior models are based on context-dependent. Moreover, the prior works only attend aspect terms to detect the sentiment word and cannot consider the sentiment words that might be influenced by domain-specific knowledge. In this work, we proposed a novel integrating Aspect-aware Interactive Attention and Emotional Position-aware module for multi-aspect sentiment analysis (abbreviated to AIAEP) where the aspect-aware interactive attention is utilized to extract aspect terms, and it fuses the domain-specific information of an aspect and context and learns their relationship representations by global context and local context attention mechanisms. Specifically, in the sentiment lexicon, the syntactic parse is used to increase the prior domain knowledge. Then we propose a novel position-aware fusion scheme to compose aspect-sentiment pairs. It combines absolute distance and relative distance from aspect terms and sentiment words, which can improve the accuracy of polarity classification. Extensive experimental results on SemEval2014 task4 restaurant and AIChallenge2018 datasets demonstrate that AIAEP can outperform state-of-the-art approaches, and it is very effective for aspect-level sentiment analysis. Xiaowen Zhou, Xianbin Wen, Hongyun Ning |
MMAsia | 5 |
| 2018 | Multiple- Instance Learning with Empirical Estimation Guided Instance SelectionabstractThe embedding based framework handles the multiple-instance learning (MIL) via the instance selection and embedding. It is how to select instance prototypes that becomes the main difference between various algorithms. Most current studies depend on single criteria for selecting instance prototypes. In this paper, we adopt two kinds of instance-selection criteria from two different views. For the combination of the two-view criteria, we also present an empirical estimator under which the two criteria compete for the instance selection. Experimental results validate the effectiveness of the proposed empirical estimator based instance-selection method for MIL. Xianbin Wen, Haixia Xu 0003 |
ICPR | 2 |
| 2018 | An Iterative Instance Selection Based Framework for Multiple-Instance LearningabstractThe instance selection based model is an effective multiple-instance learning (MIL) framework, which solves the MIL problems by embedding examples (bags of instances) into a new feature space formed by some concepts (represented by some selected instances). Most previous studies use single-point concepts for the instance selection, where every possible concept is represented by only a single instance. In this paper, we apply multiple-point concepts for choosing instances, in which each possible concept is jointly represented by a group of similar instances. Furthermore, we establish an iterative instance selection based MIL framework based on multiple-point concepts, which is guaranteed to automatically converge to the needed number of concepts for a given problem. The experimental results demonstrate that the proposed framework can better handle not only common MIL problems but also hybrid ones compared to state-of-the-art MIL algorithms. Xianbin Wen, Haixia Xu 0003 |
ICTAI | 2 |
| 2017 | A robust approach of watermarking in contourlet domain based on probabilistic neural network
Jiaxing Liu 0001, Xianbin Wen, Haixia Xu 0003 |
Multim. Tools Appl. | 2 |
| 2005 | Classification of SAR Imagery Using Multiscale Self-organizing Network
Xianbin Wen |
ISNN (2) | 1 |
| 2004 | The Cook Projection Index Estimation Using the Wavelet Kernel Function
Wei Lin 0006, Zheng Tian 0001, Xianbin Wen |
ISNN (1) | 4 |