Xin Xiong 0016

dblp:57/1151-16 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-2998-6494ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CLIP-based partial-wise prompt learning for unsupervised vehicle re-identification
Qi Wang 0061, Xin Xiong 0016
Expert Syst. Appl.6
2026 Dual-Student Adversarial Framework With Discriminator and Consistency-Driven Learning for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised medical image segmentation is essential for alleviating the cost of manual annotation in clinical applications. However, existing methods often suffer from unreliable pseudo-labels and confirmation bias in consistency-based training, which can lead to unstable optimization and degraded performance. To address these issues, a novel method named dual-Student adversarial framework with discriminator and consistency-driven learning for semi-supervised medical image segmentation is proposed. Specifically, an adversarial learning-based segmentation refinement (ALSR) module is designed to encourage prediction diversity between two student networks and leverage a shared discriminator for adversarial refinement of pseudo-labels. To further stabilize the consistency process, a residual exponential moving average (R-EMA) is applied in the uncertainty estimation with inter-instance consistency measurement (UIM) module to construct a robust teacher model, while noisy voxel predictions are selectively filtered based on uncertainty estimation. In addition, a Contrastive Representation Stabilization (CRS) module is developed to enhance voxel-level semantic alignment by performing contrastive learning only on confident regions, improving feature discriminability and structural consistency. Extensive experiments on benchmark datasets demonstrate that our method consistently outperforms prior state-of-the-art approaches.
Haifan Wu, Yuhan Geng, Di Gai, Jieying Tu, Xin Xiong 0016, Qi Wang 0061
IEEE J. Biomed. Health Informatics5
2025 Vehiclemae: View-Asymmetry Mutual Learning for Vehicle Re-Identification Pre-Training Via Masked Autoencoders
Qi Wang 0061, Dong Wang 0080, Di Gai, Xin Xiong 0016, Jiyang Xu, Ruihua Zhou
ICCV5
2024 Vision-language constraint graph representation learning for unsupervised vehicle re-identification
Dong Wang 0080, Qi Wang 0061, Zhiwei Tu, Weidong Min, Xin Xiong 0016, Yuling Zhong, Di Gai
Expert Syst. Appl.5
2024 Feature ensemble network for medical image segmentation with multi-scale atrous transformer
abstract
Abstract Recent years have witnessed notable advancements in medical image segmentation through deep convolutional neural networks. However, a notable limitation lies in the local operation of convolution, which hinders the ability to fully exploit global semantic information. To overcome the challenges prevalent in medical image segmentation, the feature ensemble network with multi‐scale atrous transformer is proposed. At the core of the approach lies the multi‐scale contextual integration module, which is based on the multi‐scale atrous transformer and facilitates contextual integration of multi‐level features. To extract discriminative fine‐grained features of the target region, a hybrid attention mechanism that synergistically combines spatial and channel attention, thereby sharpening the model's focus on crucial target information within high‐level features, is incorporated. Additionally, the channel‐aware feature reconstruction module is introduced as an innovative component engineered to tackle feature similarity issues across different categories. This module performs feature reconstruction based on channel perception, effectively widening the feature gap between categories and enhancing the segmentation capability. It is worth mentioning that our approach surpasses the state‐of‐the‐art method using three benchmark datasets in medical image segmentation.
Di Gai, Yuhan Geng, Xin Xiong 0016, Ruihua Zhou, Qi Wang 0061
IET Image Process.5
2024 Semi-supervised contextual cognitive augmentation-based cross-teaching network for multiclass medical image segmentation
abstract
Abstract The application of medical image segmentation technology enables accurate localization of human tissues, providing doctors with a reliable foundation for diagnosis. While deep learning methods have proven effective in this task, most current approaches rely on a single prediction framework, which overlooks Edge semantic features and results in flawed texture features. Moreover, existing supervised methods face challenges due to limited availability of high‐quality annotations in the field of medical imaging. In this article, a Semi‐supervised Contextual Cognitive Augmentation‐based Cross‐teaching Network is proposed. A Contextual Cognitive Enhancement Module is introduced consisting of two components: data augmentation and information extraction. The data augmentation component provides multi‐level data distribution by incorporating diverse perturbation strategies such as Discrete Cosine Transform and Gaussian noise. The information extraction component employs the Comprehensive Information Extraction module, which consists of Global Perception Information Extraction module and Multi‐channel Information Extraction module to extract perceptual information from images and enhance interaction between image channels, respectively. Additionally, a cross‐teaching strategy is adopted and a hybrid loss function is utilized to encourage knowledge sharing among the networks, leveraging the advantages of dual networks for improved performance. Experimental results demonstrate significant enhancements in multiclass medical image segmentation compared to several state‐of‐the‐art single‐framework networks.
Di Gai, Yusong Xiao, Yuhan Geng, Xin Xiong 0016, An-qi Zhong
IET Image Process.6
2024 Semi-supervised medical image classification based on class prototype matching for soft pseudo labels with consistent regularization
Di Gai, Ruonan Xiong, Weidong Min, Qi Wang 0061, Xin Xiong 0016, Chunjiang Peng
Multim. Tools Appl.6
2023 SAR ship localization method with denoising and feature refinement
abstract
Synthetic Aperture Radar (SAR) ship detection is greatly important to marine transportation monitoring and fishery resource management. To improve the detection accuracy of small ships, an SAR ship localization method with Denoising and Feature Refinement (DFR) is proposed in this paper. It consists of three parts. The first part is the denoising module, which uses non-local mean to suppress the speckle noise of the SAR image . The second part is Hierarchical Feature Fusion (HFF) module. It can integrate more low-level features by adding skip connections. This prevents the low-level spatial position information of the fused features from being diluted by high-level semantic information, therefore it is beneficial to the detection of small ships. The third part is a center-based ship predictor with Feature Refinement (FR). The FR module is proposed to refine the features and reduce the background interference, which is conducive to locate ships more accurately. Extensive experiments are conducted. The experimental results show that after adding the denoising and FR modules, the value of AP 0.5 is increased by 1.7% and 2.3%, respectively, which proves the effectiveness of these two modules. In inshore and offshore scenarios, the AP 0.5 values of DFR are 0.884 and 0.966, respectively, achieving the best results. The proposed method can also be generalized to mark lesion locations in medical images and detect offshore oil production platforms .
Cheng Zha, Weidong Min, Wei Li 0151, Xin Xiong 0016, Qi Wang 0061
Eng. Appl. Artif. Intell.5
2023 Human Skeleton Feature Optimizer and Adaptive Structure Enhancement Graph Convolution Network for Action Recognition
abstract
Human action recognition based on the graph convolution network (GCN) is a hot topic in computer vision. Existing GCN-based methods fail to capture internal implicit information when extracting action features, thereby leading to over-smoothing in the training stage. These issues result in poor performance and inaccurate extraction of action features. To address these problems, a new GCN is constructed. In this paper, a human skeleton feature optimizer (SFO) and adaptive structure enhancement graph convolution network (ASE-GCN) for action recognition are proposed in an end-to-end manner. To obtain discriminative features, the SFO is proposed to construct a new skeleton representation for action recognition through the connection criterion, which extracts the internal implicit information of action. The action feature of the joint coordinates is extracted by graph structure mask (GSM), directed graph mapping (DGM), and adaptive pooling operation (APO) in the proposed ASE-GCN network. The GSM acts as the regularizer of skeleton structure information to strengthen the representation of the graph structure. The DGM correlates the directed graph with human motion information through kinematic principle, and the APO strengthens the global high-frequency features to alleviate over-smoothing. The proposed method achieves comparable or superior results over state-of-the-art methods when used in experiments on two large public-scale datasets, NTU-RGB+D and Kinetics.
Xin Xiong 0016, Weidong Min, Qi Wang 0061, Cheng Zha
IEEE Trans. Circuits Syst. Video Technol.1
2022 Multiple Granularity Spatiotemporal Network for Sea Surface Temperature Prediction
abstract
Sea surface temperature (SST) prediction has an important practical value in marine disaster prevention and mitigation. Most current methods only use the temporal correlation of SST during prediction, but the spatial correlation is not considered, resulting in low prediction accuracies. In addition, the changing trend of SST as reflected by the single granularity feature is unreliable, and the degrees of dependence between historical SST and future SST tend to vary. In order to overcome these issues, the multiple granularity spatiotemporal network (MGSN) is proposed for SST prediction. The proposed method consists of three parts. First, a multibranch network structure is constructed to extract different temporal features of different granularities. Second, a temporal dependence representation module is developed to represent the different degrees of dependence between historical SST and predicted SST in the temporal dimension. Third, the spatiotemporal fusion prediction module is used to achieve a spatiotemporal prediction of the SST and fuse the prediction results of different granular features. Comparative experiments have been conducted. The experimental results show that the root-mean-square error (RMSE) of the proposed method is reduced by 0.1360, 0.1608, and 0.1448 compared with the RMSE of convolutional LSTM (ConvLSTM), when predicting SST for the next one day, three days, and seven days, respectively. Our method has strong spatiotemporal feature modeling capabilities and is suitable for regional SST prediction.
Cheng Zha, Weidong Min, Xin Xiong 0016, Qi Wang 0061
IEEE Geosci. Remote. Sens. Lett.4
2022 Multi-view face generation via unpaired images
Yanni Zou, Weidong Min, Xin Xiong 0016
Vis. Comput.5
2021 Multimodal graph inference network for scene graph generation
Jingwen Duan, Weidong Min, Deyu Lin, Xin Xiong 0016
Appl. Intell.5
2021 Viewpoint adaptation learning with cross-view distance metric for robust vehicle re-identification
Qi Wang 0061, Weidong Min, Ziyuan Yang 0001, Xin Xiong 0016
Inf. Sci.5
2021 Driver Yawning Detection Based on Subtle Facial Action Recognition
abstract
Various investigations have shown that driver fatigue is the main cause of traffic accidents. Research on the use of computer vision techniques to detect signs of fatigue from facial actions, such as yawning, has demonstrated good potential. However, accurate and robust detection of yawning is difficult because of the complicated facial actions and expressions of drivers in the real driving environment. Several facial actions and expressions have the same mouth deformation as yawning. Thus, a novel approach to detecting yawning based on subtle facial action recognition is proposed in this study to alleviate the abovementioned problems. A 3D deep learning network with a low time sampling characteristic is proposed for subtle facial action recognition. This network uses 3D convolutional and bidirectional long short-term memory networks for spatiotemporal feature extraction and adopts SoftMax for classification. A keyframe selection algorithm is designed to select the most representative frame sequence from subtle facial actions. This algorithm rapidly eliminates redundant frames using image histograms with low computation cost and detects outliers by median absolute deviation. A series of experiments are also conducted on YawDD benchmark and self-collected datasets. Compared with several state-of-the-art methods, the proposed method has high yawning detection rates and can effectively distinguish yawning from similar facial actions.
Hao Yang 0027, Li Liu 0010, Weidong Min, Xiaosong Yang, Xin Xiong 0016
IEEE Trans. Multim.5
2020 S3D-CNN: skeleton-based 3D consecutive-low-pooling neural network for fall detection
Xin Xiong 0016, Weidong Min, Wei-Shi Zheng 0001, Pin Liao, Hao Yang 0027
Appl. Intell.1