VLDB 2026 Research / reviewers in the wild / expert
Xun Gong 0002
dblp:58/5901-2
· DBLP profile ↗
31ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0002-1494-0955ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prompt-driven multi-instance CLIP: Aligning heterogeneous modalities with missing data tolerance for multi-modal medical analysis
Yafei Ou, Cenyang Zheng, Xun Gong 0002 |
Expert Syst. Appl. | 4 |
| 2026 | Dual-branch decoupling framework for medical one-class classification
Xun Gong 0002, Lingsong Huang, Yafei Ou |
Pattern Recognit. | 2 |
| 2025 | Efficient Visual Representation Learning with Heat Conduction EquationabstractFoundation models, such as CNNs and ViTs, have powered the development of image representation learning. However, general guidance to model architecture design is still missing. Inspired by the connection between image representation learning and heat conduction, we model images by the heat conduction equation, where the essential idea is to conceptualize image features as temperatures and model their information interaction as the diffusion of thermal energy. Based on this idea, we find that many modern model architectures, such as residual structures, SE block, and feed-forward networks, can be interpreted from the perspective of the heat conduction equation. Therefore, we leverage the heat equation to design new and more interpretable models. As an example, we propose the Heat Conduction Layer and the Refinement Approximation Layer inspired by solving the heat conduction equation using Finite Difference Method and Fourier series, respectively. The main goal of this paper is to integrate the overall architectural design of neural networks into the theoretical framework of heat conduction. Nevertheless, our Heat Conduction Network (HcNet) still shows competitive performance, e.g., HcNet-T achieves 83.0% top-1 accuracy on ImageNet-1K while only requiring 28M parameters and 4.1G MACs. The code is publicly available at: https://github.com/ZheminZhang1/HcNet. Zhemin Zhang, Xun Gong 0002 |
IJCAI | 2 |
| 2025 | GDKansformer: A group-wise dynamic Kolmogorov-Arnold transformer with multi-view gated attention for pathological image diagnosis
Xun Gong 0002, Yuanyuan Chen 0004 |
Expert Syst. Appl. | 2 |
| 2025 | AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
Wei Li 0249, Xun Gong 0002, Xiaobin Sun |
Knowl. Based Syst. | 2 |
| 2025 | Cycle-VQA: A Cycle-Consistent Framework for Robust Medical Visual Question Answering
Xun Gong 0002, Cenyang Zheng, Xuli Tan, Yafei Ou |
Pattern Recognit. | 2 |
| 2025 | Anti-noise face: A resilient model for face recognition with labeled noise data
Lei Wang 0276, Xun Gong 0002 |
Signal Process. Image Commun. | 2 |
| 2025 | Cross-Level Adaptive Pyramid Pooling Network for Cardiac Structure SegmentationabstractThe segmentation of cardiac structures in two-dimensional echocardiographic video is crucial for the diagnosis of various cardiac diseases. However, existing video datasets typically contain only sparse annotations for one or two frames. Furthermore, due to the complex variations caused by cardiac motion, the pooling mechanisms employed in current methods result in the loss of global context for the segmentation targets. To address these issues, we propose a Cross-Level Adaptive Pyramid Pooling Network (CAPP) for cardiac image segmentation under sparse annotations. Specifically, the proposed dynamic adaptive network adjusts its architecture based on the size and spacing of input patches from images of different scales, enhances the network's focus on beneficial features through a cross-channel interaction mechanism, and employs data augmentation to increase data complexity, thereby achieving better training outcomes and addressing the segmentation challenges posed by sparse annotations. Additionally, we propose the Cross-Level Pyramid Pooling Module, which captures contextual information between different levels in the network to extract relevant features of cardiac structures. Our method achieves state-of-the-art performance in cardiac structure segmentation on two publicly available echocardiographic datasets and the corresponding code will be made available on GitHub. Biyu Yan, Jinrong Lv, Xun Gong 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Generating Multi-Center Classifier via Conditional Gaussian DistributionabstractIn real-world data, one class can contain several local clusters,e.g., birds of different poses, which makes it difficult to represent the feature distribution of each class using only a single center. Existing intra-class multimodal representation methods employ sub-centers to capture intra-class variations. However, these sub-center methods have some limitations, i.e., they ignore the relationship between sub-centers and do not ensure the diversity of sub-centers. To address these limitations, we propose a novel multi-center classifier. Different from the vanilla multi-center classifier, our proposal is established on the assumption that the deep features of the training set follow a Gaussian Mixture distribution. Specifically, we create a conditional Gaussian distribution for each class and then sample multiple sub-centers from that distribution to extend the linear classifier. This approach allows the model to capture intra-class local structures more efficiently. In addition, we propose a novel label assignment strategy, the Multi-Center Class Label, to ensure that each sub-center is effectively involved in the training. Extensive experiments on various recognition benchmarks like ImageNet, CIFAR, and Mini-ImageNet demonstrate the effectiveness of our proposal. Zhemin Zhang, Xun Gong 0002 |
IEEE Signal Process. Lett. | 2 |
| 2025 | CLIP-Based Camera-Agnostic Feature Learning for Intra-Camera Supervised Person Re-IdentificationabstractContrastive Language-Image Pre-Training (CLIP) model excels in traditional person re-identification (ReID) tasks due to its inherent advantage in generating textual descriptions for pedestrian images. However, applying CLIP directly to intra-camera supervised person re-identification (ICS ReID) presents challenges. ICS ReID requires independent identity labeling within each camera, without associations across cameras. This limits the effectiveness of text-based enhancements. To address this, we propose a novel framework called CLIP-based Camera-Agnostic Feature Learning (CCAFL) for ICS ReID. Accordingly, two custom modules are designed to guide the model to actively learn camera-agnostic pedestrian features: Intra-Camera Discriminative Learning (ICDL) and Inter-Camera Adversarial Learning (ICAL). Specifically, we first establish learnable textual prompts for intra-camera pedestrian images to obtain crucial semantic supervision signals for subsequent intra- and inter-camera learning. Then, we design ICDL to increase inter-class variation by considering the hard positive and hard negative samples within each camera, thereby learning intra-camera finer-grained pedestrian features. Additionally, we propose ICAL to reduce inter-camera pedestrian feature discrepancies by penalizing the model’s ability to predict the camera from which a pedestrian image originates, thus enhancing the model’s capability to recognize pedestrians from different viewpoints. Extensive experiments on popular ReID datasets demonstrate the effectiveness of our approach. Especially, on the challenging MSMT17 dataset, we arrive at 58.9% in terms of mAP accuracy, surpassing state-of-the-art methods by 7.6%. Code is available athttps://gitee.com/swjtugx/classmate/tree/master/OurGroup/CCAFL. Xun Gong 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | RandomViG: Random Vision Graph Neural Network for Image ClassificationabstractVision Graph Neural Network (ViG) is the first graph neural network model capable of directly processing image data. The community primarily focuses on the model structures to improve ViG's performance but lacks attention to its graph construction method. To avoid quadratic computational complexity, ViG uses clustering algorithms (K-nearest neighbor) to construct graph structures. Nevertheless, clustering algorithms introduce biases, which limit ViG's ability to obtain global information. To address this problem, we propose RandomViG, which abandons clustering algorithms and uses a random manner to obtain relationships between nodes. Our RandomViG is sparse in computation and can approximate a complete graph, enabling ViG to gain global interaction capability. In order to obtain the local dependence, we design a local feature extraction module for RandomViG. In addition, to alleviate the over-smoothing problem, we propose a novel method called MRN (maintaining relationships among nodes). Considering that the increased feature diversity does not necessarily lead to better performance, MRN does not aim to maximize the feature diversity of the model but instead strives to maintain consistency between the feature similarity and the inherent similarity of the original image. We validate our proposal in three major computer visual tasks, including image classification, object detection, and instance segmentation. Without extra data, RandomViG-Ti achieves 79.4% ImageNet-1 K top-1 accuracy, outperforming the baseline (ViG) by 1.2%. Under the same model scale, our RandomViG performs better with fewer FLOPs compared with existing state-of-the-art models. Xun Gong 0002, Daisong Yan, Zhemin Zhang |
IEEE Trans. Multim. | 1 |
| 2024 | Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute AnalysisabstractMed-VQA, intersecting medicine and VQA, is a challenging field offering benefits like patient engagement and clinical collaboration. However, current joint embedding methods lack explanation of reasoning, undermining VQA credibility. In this paper, motivated by causal effect, we propose a novel Triangular Reasoning VQA (Tri-VQA) framework, which constructs reverse causal questions from the perspective of "Why this answer?" to elucidate the source of the answer and stimulate more reasonable forward reasoning processes. We evaluate our method on EUS multi-attribute annotated datasets from five centers and medical VQA datasets, demonstrating superior performance. Our codes and pre-trained models are available at https://github.com/hahaha111111/Tri-VQA. Xun Gong 0002, Cenyang Zheng, Yafei Ou |
BIBM | 2 |
| 2024 | Bilinear Fine-grained Classification of Ultrasound Images Integrated with Interpretable Radiomics
Chenzhong Wang, Xun Gong 0002, Weiji Kong |
PRCV (15) | 2 |
| 2024 | Video-based person re-identification with scene and person attributes
Xun Gong 0002 |
Multim. Tools Appl. | 1 |
| 2024 | Contrastive Mean Teacher for Intra-Camera Supervised Person Re-IdentificationabstractIntra-camera supervision (ICS) person reidentification (Re-ID) assumes that a person’s identity labels are independently annotated within each camera, lacking inter-camera association for person identities. Recently, several ICS methods have achieved significant results by using two stages: intra-camera learning and inter-camera learning for model training. However, in the intra-camera learning stage, these methods only focus on pedestrian features within each camera, which increases the variance of the same person across different cameras. In the inter-camera learning stage, due to lighting variations and background shifts, the generated pseudo-labels from feature similarity contain significant noise, and the unassociated outlier samples are not fully utilized. To address these issues, we propose a Contrastive Mean Teacher (CMT) framework combiningMean-teacherparadigm and contrastive learning. Specifically, by conducting both intra-camera and inter-camera learning simultaneously, we can fully leverage predefined intra-camera labels and inter-camera-associated labels. This method can effectively learn pedestrian features under various cameras. Moreover, the teacher model provides more stable predictions, which helps to establish a better inter-camera association and improves the model’s generalization capabilities. Finally, we design a background filtering module that employs attention mechanisms to guide instance normalization, further reducing variations in identity features caused by lighting and background changes. We validate our method on three large-scale person re-identification datasets, and the results show that our approach outperforms all existing ICS methods. Specifically, our approach achieves a state-of-the-art accuracy 88.9% mAP and 95.8% Rank-1 on the challenging Market1501 benchmarked with ResNet-50, even surpassing the performance of state-of-the-art fully supervised methods. Xun Gong 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Positional Label for Self-Supervised Vision TransformerabstractPositional encoding is important for vision transformer (ViT) to capture the spatial structure of the input image. General effectiveness has been proven in ViT. In our work we propose to train ViT to recognize the positional label of patches of the input image, this apparently simple task actually yields a meaningful self-supervisory task. Based on previous work on ViT positional encoding, we propose two positional labels dedicated to 2D images including absolute position and relative position. Our positional labels can be easily plugged into various current ViT variants. It can work in two ways: (a) As an auxiliary training target for vanilla ViT for better performance. (b) Combine the self-supervised ViT to provide a more powerful self-supervised signal for semantic feature learning. Experiments demonstrate that with the proposed self-supervised methods, ViT-B and Swin-B gain improvements of 1.20% (top-1 Acc) and 0.74% (top-1 Acc) on ImageNet, respectively, and 6.15% and 1.14% improvement on Mini-ImageNet. The code is publicly available at: https://github.com/zhangzhemin/PositionalLabel. Zhemin Zhang, Xun Gong 0002 |
AAAI | 2 |
| 2023 | DSF-net: occluded person re-identification based on dual structure features
Yueqiao Fan, Xun Gong 0002, Yuning He |
Neural Comput. Appl. | 2 |
| 2023 | Multi-Grained Temporal Segmentation Attention Modeling for Skeleton-Based Action RecognitionabstractThe Transformer network has been widely studied for skeleton-based action recognition, and existing methods have made significant progress. However, these methods still suffer from severe overfitting issues and are less effective at capturing local relationships compared to methods based on graph convolutional neural networks (GCN). Building on previous research on attention, we propose MTS-Former, a novel multi-granularity temporal segmentation attention-based modeling method, to address these challenges. MTS-Former dynamically learns the topology of the skeleton by position-sensitive axis-attention, eliminating the constraints of manually crafted adjacency matrices. Furthermore, it effectively reduces overfitting by combining attention-guided regularization. It models the global-local relationship of skeleton sequences through segmental sampling and a multi-granularity aggregation strategy, thus enabling more robust local motion feature extraction. MTS-Former outperforms state-of-the-art Transformer networks on four benchmark datasets, DHG, SHREC, NUT RGB+D 60, and NTU RGB+D 120. Jinrong Lv, Xun Gong 0002 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Rotated and Masked Image Modeling: A Superior Self-Supervised Method for ClassificationabstractMask image modeling (MIM) has performed excellently as a transformer-based self-supervised method via random masking and reconstruction. However, since the unmasked image patches are non-participation in the loss computation, MIM cannot effectively utilize the data and waste much computation. This drawback usually limits the learning ability of the pre-training model when pre-training on small-scale datasets. To solve this problem, we propose a novel self-supervised learning method for small-scale datasets called RotMIM. Unlike MIM, RotMIM has a different pretext task: recognizing the rotation angle that is applied to the unmasked patches. RotMIM can fully utilize data and provide a stronger self-supervised signal. Moreover, to fit RotMIM, we propose a data augmentation method called FeaMix. Our proposal ensures that the mixing area with RotMIM understands that each basic unit of semantic information in an image has the same size. This consistency guarantees clean tokenization during fine-tuning after pre-training. Our proposals outperform state-of-the-art self-supervised methods on three popular datasets, Mini-ImageNet, Caltech256, and Cifar100. Daisong Yan, Xun Gong 0002, Zhemin Zhang |
IEEE Signal Process. Lett. | 2 |
| 2022 | BoundaryFace: A Mining Framework with Noise Label Self-correction for Face Recognition
Xun Gong 0002 |
ECCV (13) | 2 |
| 2022 | Bounding box regression with balance for harmonious object detection
Chenzhong Wang, Xun Gong 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Dynamic Channel-Aware Subgraph Interactive Networks for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition has obtained remarkable success due to the rapid development of Graph Neural Networks (GNNs). Existing skeleton-based methods primarily focus on the spatial aspect of skeleton graphs, but rarely mine long-range temporal relationships to learn the intrinsic dependencies across different frames. The latter is crucial for extracting discriminative motion patterns. To acquire a more precise spatial-temporal representation for the human skeleton data, we develop a lightweight yet practical method termed Dynamic Channel-Aware Subgraph Interactive Network (DCA-SGIN), using the interactive motion to adaptively capture long-range temporal nuance among the skeleton sequences. Moreover, we unify a novel paradigm to address the inherent structural limitations of GCNs. Specifically, the proposed graph interactive learners utilize the collaborative channel-aware topology of multiple subgraphs to model spatial relations. Unlike previous methods, the global and local features are simultaneously considered in spatial modeling without calculating any adjacency matrix, making it more efficient. Each proposed individual component in DCA-SGIN can be treated as a plug-in module that can be easily applied to other GNNs. Extensive experiments on the three challenging datasets show that the performance of DCA-SGIN outperforms the state-of-the-art methods with fewer FLOPs. Kunlun Wu, Xun Gong 0002 |
IEEE Signal Process. Lett. | 2 |
| 2022 | LAG-Net: Multi-Granularity Network for Person Re-Identification via Local Attention SystemabstractPerson re-identification (Re-ID) is a challenging research topic which aims to retrieve the pedestrian images of the same person that captured by non-overlapping cameras. Existing methods either assume the body parts of the same person are well-aligned, or use attention selection mechanisms to constrain the effective region of feature learning. But these methods concentrate only on coarse feature representation and cannot model complex real scenes effectively. We propose a novel Local Attention Guided Network (LAG-Net) to not only exploit the most salient area among different people, but also extract important local detail through a Local Attention System (LAS). LAS is an attention selection unit that could extract approximate semantic local features of human body parts without extra supervision. To learn discriminative attention feature representation, we explore an attention feature regularization scheme to enhance the relevance of body part features that belong to same personal identity. Considering the effectiveness of feature augmentation in the Re-ID task and the defect of the existing methods, we propose a Batch Attention DropBlock (BA-DropBlock) to further improve DropBlock by combining the attention selection mechanism. Results on mainstream datasets demonstrate the superiority of our model over the state-of-the-art. Especially, our approach exceeds the current best method by a large margin of 4.6${\%}$on the most challenging dataset CUHK03. Xun Gong 0002, Zu Yao, Xin Li 0003, Yueqiao Fan, Jianfeng Fan, Boji Lao |
IEEE Trans. Multim. | 1 |
| 2021 | Face recognition based on adaptive margin and diversity regularization constraintsabstractAbstract In recent years, a more robust facial feature can be learned by convolutional neural networks once introducing margins into loss functions. Those methods set a margin for each class manually to squeeze the intra‐class variations within each class equally. However, the internal feature distributions of different persons in the real world are highly unbalanced, and the distance between different identities is not uniform either. As a result, applying the same margin on all classes might not lead to higher inter‐class differences. To address this problem, this paper proposes an adaptive margin based on feature distribution to squeeze the feature interior spaces of different classes. Simultaneously, because the inter‐class margin can adequately represent the distribution of different classes in the feature space, this paper proposes a novel diversity regularization method. The regularization weights of each class are dynamically set depending on their margins. This method proposed in this paper is intuitively interpretable and can be easily applied to other classification scenarios. Experiments on current existing benchmarks have demonstrated the superiority of our method over state‐of‐the‐art competitors. Zhemin Zhang, Xun Gong 0002, Junzhou Chen 0001 |
IET Image Process. | 2 |
| 2021 | A directional margin paradigm for noise suppression in face recognition
Yang Zhou 0057, Xun Gong 0002, Peng Yang 0007 |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | A Cross-Dimension Annotations Method for 3D Structural Facial Landmark ExtractionabstractAbstract Recent methods for 2D facial landmark localization perform well on close‐to‐frontal faces, but 2D landmarks are insufficient to represent 3D structure of a facial shape. For applications that require better accuracy, such as facial motion capture and 3D shape recovery, 3DA‐2D (2D Projections of 3D Facial Annotations) is preferred. Inferring the 3D structure from a single image is an ill‐posed problem whose accuracy and robustness are not always guaranteed. This paper aims to solve accurate 2D facial landmark localization and the transformation between 2D and 3DA‐2D landmarks. One way to increase the accuracy is to input more precisely annotated facial images. The traditional cascaded regressions cannot effectively handle large or noisy training data sets. In this paper, we propose a Mini‐Batch Cascaded Regressions (MBCR) method that can iteratively train a robust model from a large data set. Benefiting from the incremental learning strategy and a small learning rate, MBCR is robust to noise in training data. We also propose a new Cross‐Dimension Annotations Conversion (CDAC) method to map facial landmarks from 2D to 3DA‐2D coordinates and vice versa. The experimental results showed that CDAC combined with MBCR outperforms the‐state‐of‐the‐art methods in 3DA‐2D facial landmark localization. Moreover, CDAC can run efficiently at up to 110 fps on a 3.4 GHz‐CPU workstation. Thus, CDAC provides a solution to transform existing 2D alignment methods into 3DA‐2D ones without slowing down the speed. Training and testing code as well as the data set can be downloaded from https://github.com/SWJTU‐3DVision/CDAC. Xun Gong 0002, Zhemin Zhang, Yue Xiang, Xin Li 0003 |
Comput. Graph. Forum | 1 |
| 2020 | A Discriminative Multi-Channel Facial Shape (MCFS) Representation and Feature Extraction for 3D Human FacesabstractAbstract Building an effective representation for 3D face geometry is essential for face analysis tasks, that is, landmark detection, face recognition and reconstruction. This paper proposes to use a Multi‐Channel Facial Shape (MCFS) representation that consists of depth, hand‐engineered feature and attention maps to construct a 3D facial descriptor. And, a multi‐channel adjustment mechanism, named filtered squeeze and reversed excitation (FSRE), is proposed to re‐organize MCFS data. To assign a suitable weight for each channel, FSRE is able to learn the importance of each layer automatically in the training phase. MCFS and FSRE blocks collaborate together effectively to build a robust 3D facial shape representation, which has an excellent discriminative ability. Extensive experimental results, testing on both high‐resolution and low‐resolution face datasets, show that facial features extracted by our framework outperform existing methods. This representation is stable against occlusions, data corruptions, expressions and pose variations. Also, unlike traditional methods of 3D face feature extraction, which always take minutes to create 3D features, our system can run in real time. Xun Gong 0002, Xin Li 0003, Tianrui Li 0001, Yongqing Liang 0001 |
Comput. Graph. Forum | 1 |
| 2020 | A review on crowd simulation and modeling
Shanwen Yang, Tianrui Li 0001, Xun Gong 0002, Bo Peng 0006, Jie Hu 0007 |
Graph. Model. | 3 |
| 2019 | An LSTM based Encoder-Decoder Model for MultiStep Traffic Flow PredictionabstractTraffic flow prediction has been regarded as a key research problem in the intelligent transportation system. In this paper, we propose an encoder-decoder model with temporal attention mechanism for multi-step forward traffic flow prediction task, which uses LSTM as the encoder and decoder to learn the long dependencies features and nonlinear characteristics of multivariate traffic flow related time series data, and also introduces a temporal attention mechanism for more accurately traffic flow prediction. Through the real traffic flow dataset experiments, it has shown that the proposed model has better prediction ability than classic shallow learning and baseline deep learning models. And the predicted traffic flow value can be well matched with the ground truth value not only under short step forward prediction condition but also under longer step forward prediction condition, which validates that the proposed model is a good option for dealing with the realtime and forward-looking problems of traffic flow prediction task. Shengdong Du, Tianrui Li 0001, Yan Yang 0001, Xun Gong 0002, Shi-Jinn Horng |
IJCNN | 4 |
| 2013 | Using Sorted Switching Median Filter to remove high-density impulse noises
Shi-Jinn Horng, Ling-Yuan Hsu, Tianrui Li 0001, Shaojie Qiao, Xun Gong 0002, Hsien-Hsin Chou, Muhammad Khurram Khan |
J. Vis. Commun. Image Represent. | 5 |
| 2009 | Single 2D Image-based 3D Face Reconstruction and Its Application in Pose EstimationabstractHuman beings are born with a natural capacity of recovering shape from merely one image. However, it is still a challenging mission for current techniques to make a computer have such an ability. To simulate the modeling procedure of human visual system, a Ternary Deformation Framework (TDF) is proposed to reconstruct a realistic 3D face from one 2D frontal facial image, with prior knowledge regarding facial shape learnt from a 3D face data set. Based upon the reconstructed 3D face, a novelmethod via linear regression is then proposed to estimate that person's pose on another image with pose variations. Simulation results show that TDF outperforms the conventional methods with respect to the modeling precision and that reconstructions on real photographs have achieved favorable visual effects. Moreover, the comparison results validated the effectiveness of using the 3D face in the proposed pose estimation method. Xun Gong 0002, Guoyin Wang 0001, Lili Xiong |
Fundam. Informaticae | 1 |