VLDB 2026 Research / reviewers in the wild / expert
Xin Xiong 0016
dblp:57/1151-16
· DBLP profile ↗
15ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-2998-6494ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIP-based partial-wise prompt learning for unsupervised vehicle re-identification
Qi Wang 0061, Xin Xiong 0016 |
Expert Syst. Appl. | 6 |
| 2026 | Dual-Student Adversarial Framework With Discriminator and Consistency-Driven Learning for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation is essential for alleviating the cost of manual annotation in clinical applications. However, existing methods often suffer from unreliable pseudo-labels and confirmation bias in consistency-based training, which can lead to unstable optimization and degraded performance. To address these issues, a novel method named dual-Student adversarial framework with discriminator and consistency-driven learning for semi-supervised medical image segmentation is proposed. Specifically, an adversarial learning-based segmentation refinement (ALSR) module is designed to encourage prediction diversity between two student networks and leverage a shared discriminator for adversarial refinement of pseudo-labels. To further stabilize the consistency process, a residual exponential moving average (R-EMA) is applied in the uncertainty estimation with inter-instance consistency measurement (UIM) module to construct a robust teacher model, while noisy voxel predictions are selectively filtered based on uncertainty estimation. In addition, a Contrastive Representation Stabilization (CRS) module is developed to enhance voxel-level semantic alignment by performing contrastive learning only on confident regions, improving feature discriminability and structural consistency. Extensive experiments on benchmark datasets demonstrate that our method consistently outperforms prior state-of-the-art approaches. Haifan Wu, Yuhan Geng, Di Gai, Jieying Tu, Xin Xiong 0016, Qi Wang 0061 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Vehiclemae: View-Asymmetry Mutual Learning for Vehicle Re-Identification Pre-Training Via Masked Autoencoders
Qi Wang 0061, Dong Wang 0080, Di Gai, Xin Xiong 0016, Jiyang Xu, Ruihua Zhou |
ICCV | 5 |
| 2024 | Vision-language constraint graph representation learning for unsupervised vehicle re-identification
Dong Wang 0080, Qi Wang 0061, Zhiwei Tu, Weidong Min, Xin Xiong 0016, Yuling Zhong, Di Gai |
Expert Syst. Appl. | 5 |
| 2024 | Feature ensemble network for medical image segmentation with multi-scale atrous transformerabstractAbstract Recent years have witnessed notable advancements in medical image segmentation through deep convolutional neural networks. However, a notable limitation lies in the local operation of convolution, which hinders the ability to fully exploit global semantic information. To overcome the challenges prevalent in medical image segmentation, the feature ensemble network with multi‐scale atrous transformer is proposed. At the core of the approach lies the multi‐scale contextual integration module, which is based on the multi‐scale atrous transformer and facilitates contextual integration of multi‐level features. To extract discriminative fine‐grained features of the target region, a hybrid attention mechanism that synergistically combines spatial and channel attention, thereby sharpening the model's focus on crucial target information within high‐level features, is incorporated. Additionally, the channel‐aware feature reconstruction module is introduced as an innovative component engineered to tackle feature similarity issues across different categories. This module performs feature reconstruction based on channel perception, effectively widening the feature gap between categories and enhancing the segmentation capability. It is worth mentioning that our approach surpasses the state‐of‐the‐art method using three benchmark datasets in medical image segmentation. Di Gai, Yuhan Geng, Xin Xiong 0016, Ruihua Zhou, Qi Wang 0061 |
IET Image Process. | 5 |
| 2024 | Semi-supervised contextual cognitive augmentation-based cross-teaching network for multiclass medical image segmentationabstractAbstract The application of medical image segmentation technology enables accurate localization of human tissues, providing doctors with a reliable foundation for diagnosis. While deep learning methods have proven effective in this task, most current approaches rely on a single prediction framework, which overlooks Edge semantic features and results in flawed texture features. Moreover, existing supervised methods face challenges due to limited availability of high‐quality annotations in the field of medical imaging. In this article, a Semi‐supervised Contextual Cognitive Augmentation‐based Cross‐teaching Network is proposed. A Contextual Cognitive Enhancement Module is introduced consisting of two components: data augmentation and information extraction. The data augmentation component provides multi‐level data distribution by incorporating diverse perturbation strategies such as Discrete Cosine Transform and Gaussian noise. The information extraction component employs the Comprehensive Information Extraction module, which consists of Global Perception Information Extraction module and Multi‐channel Information Extraction module to extract perceptual information from images and enhance interaction between image channels, respectively. Additionally, a cross‐teaching strategy is adopted and a hybrid loss function is utilized to encourage knowledge sharing among the networks, leveraging the advantages of dual networks for improved performance. Experimental results demonstrate significant enhancements in multiclass medical image segmentation compared to several state‐of‐the‐art single‐framework networks. Di Gai, Yusong Xiao, Yuhan Geng, Xin Xiong 0016, An-qi Zhong |
IET Image Process. | 6 |
| 2024 | Semi-supervised medical image classification based on class prototype matching for soft pseudo labels with consistent regularization
Di Gai, Ruonan Xiong, Weidong Min, Qi Wang 0061, Xin Xiong 0016, Chunjiang Peng |
Multim. Tools Appl. | 6 |
| 2023 | SAR ship localization method with denoising and feature refinementabstractSynthetic Aperture Radar (SAR) ship detection is greatly important to marine transportation monitoring and fishery resource management. To improve the detection accuracy of small ships, an SAR ship localization method with Denoising and Feature Refinement (DFR) is proposed in this paper. It consists of three parts. The first part is the denoising module, which uses non-local mean to suppress the speckle noise of the SAR image . The second part is Hierarchical Feature Fusion (HFF) module. It can integrate more low-level features by adding skip connections. This prevents the low-level spatial position information of the fused features from being diluted by high-level semantic information, therefore it is beneficial to the detection of small ships. The third part is a center-based ship predictor with Feature Refinement (FR). The FR module is proposed to refine the features and reduce the background interference, which is conducive to locate ships more accurately. Extensive experiments are conducted. The experimental results show that after adding the denoising and FR modules, the value of AP 0.5 is increased by 1.7% and 2.3%, respectively, which proves the effectiveness of these two modules. In inshore and offshore scenarios, the AP 0.5 values of DFR are 0.884 and 0.966, respectively, achieving the best results. The proposed method can also be generalized to mark lesion locations in medical images and detect offshore oil production platforms . Cheng Zha, Weidong Min, Wei Li 0151, Xin Xiong 0016, Qi Wang 0061 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Human Skeleton Feature Optimizer and Adaptive Structure Enhancement Graph Convolution Network for Action RecognitionabstractHuman action recognition based on the graph convolution network (GCN) is a hot topic in computer vision. Existing GCN-based methods fail to capture internal implicit information when extracting action features, thereby leading to over-smoothing in the training stage. These issues result in poor performance and inaccurate extraction of action features. To address these problems, a new GCN is constructed. In this paper, a human skeleton feature optimizer (SFO) and adaptive structure enhancement graph convolution network (ASE-GCN) for action recognition are proposed in an end-to-end manner. To obtain discriminative features, the SFO is proposed to construct a new skeleton representation for action recognition through the connection criterion, which extracts the internal implicit information of action. The action feature of the joint coordinates is extracted by graph structure mask (GSM), directed graph mapping (DGM), and adaptive pooling operation (APO) in the proposed ASE-GCN network. The GSM acts as the regularizer of skeleton structure information to strengthen the representation of the graph structure. The DGM correlates the directed graph with human motion information through kinematic principle, and the APO strengthens the global high-frequency features to alleviate over-smoothing. The proposed method achieves comparable or superior results over state-of-the-art methods when used in experiments on two large public-scale datasets, NTU-RGB+D and Kinetics. Xin Xiong 0016, Weidong Min, Qi Wang 0061, Cheng Zha |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Multiple Granularity Spatiotemporal Network for Sea Surface Temperature PredictionabstractSea surface temperature (SST) prediction has an important practical value in marine disaster prevention and mitigation. Most current methods only use the temporal correlation of SST during prediction, but the spatial correlation is not considered, resulting in low prediction accuracies. In addition, the changing trend of SST as reflected by the single granularity feature is unreliable, and the degrees of dependence between historical SST and future SST tend to vary. In order to overcome these issues, the multiple granularity spatiotemporal network (MGSN) is proposed for SST prediction. The proposed method consists of three parts. First, a multibranch network structure is constructed to extract different temporal features of different granularities. Second, a temporal dependence representation module is developed to represent the different degrees of dependence between historical SST and predicted SST in the temporal dimension. Third, the spatiotemporal fusion prediction module is used to achieve a spatiotemporal prediction of the SST and fuse the prediction results of different granular features. Comparative experiments have been conducted. The experimental results show that the root-mean-square error (RMSE) of the proposed method is reduced by 0.1360, 0.1608, and 0.1448 compared with the RMSE of convolutional LSTM (ConvLSTM), when predicting SST for the next one day, three days, and seven days, respectively. Our method has strong spatiotemporal feature modeling capabilities and is suitable for regional SST prediction. Cheng Zha, Weidong Min, Xin Xiong 0016, Qi Wang 0061 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Multi-view face generation via unpaired images
Yanni Zou, Weidong Min, Xin Xiong 0016 |
Vis. Comput. | 5 |
| 2021 | Multimodal graph inference network for scene graph generation
Jingwen Duan, Weidong Min, Deyu Lin, Xin Xiong 0016 |
Appl. Intell. | 5 |
| 2021 | Viewpoint adaptation learning with cross-view distance metric for robust vehicle re-identification
Qi Wang 0061, Weidong Min, Ziyuan Yang 0001, Xin Xiong 0016 |
Inf. Sci. | 5 |
| 2021 | Driver Yawning Detection Based on Subtle Facial Action RecognitionabstractVarious investigations have shown that driver fatigue is the main cause of traffic accidents. Research on the use of computer vision techniques to detect signs of fatigue from facial actions, such as yawning, has demonstrated good potential. However, accurate and robust detection of yawning is difficult because of the complicated facial actions and expressions of drivers in the real driving environment. Several facial actions and expressions have the same mouth deformation as yawning. Thus, a novel approach to detecting yawning based on subtle facial action recognition is proposed in this study to alleviate the abovementioned problems. A 3D deep learning network with a low time sampling characteristic is proposed for subtle facial action recognition. This network uses 3D convolutional and bidirectional long short-term memory networks for spatiotemporal feature extraction and adopts SoftMax for classification. A keyframe selection algorithm is designed to select the most representative frame sequence from subtle facial actions. This algorithm rapidly eliminates redundant frames using image histograms with low computation cost and detects outliers by median absolute deviation. A series of experiments are also conducted on YawDD benchmark and self-collected datasets. Compared with several state-of-the-art methods, the proposed method has high yawning detection rates and can effectively distinguish yawning from similar facial actions. Hao Yang 0027, Li Liu 0010, Weidong Min, Xiaosong Yang, Xin Xiong 0016 |
IEEE Trans. Multim. | 5 |
| 2020 | S3D-CNN: skeleton-based 3D consecutive-low-pooling neural network for fall detection
Xin Xiong 0016, Weidong Min, Wei-Shi Zheng 0001, Pin Liao, Hao Yang 0027 |
Appl. Intell. | 1 |