Di Gai

dblp:244/8532 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Receptive field weighted representation and context enhancement for SAR ship detection
Cheng Zha, Weidong Min, Qi Wang 0061, Di Gai, Hongyue Xiang
Expert Syst. Appl.5
2026 Residual Mamba-Driven Multiscale Attentive Network With Boundary Enhancement for IoT-Enabled Medical Image Segmentation
abstract
In IoT-enabled intelligent healthcare systems, medical images are frequently acquired in real-time from heterogeneous imaging sensors such as dermoscopic devices and MRI scanners. In resource-limited or edge-deployed settings, achieving precise and rapid image segmentation plays a crucial role in facilitating early diagnosis and supporting clinical decision-making. This paper proposes a medical image segmentation method based on a residual Mamba backbone network, combining a multi-scale gated attention (MGA) module with a boundary enhancement (BE) module to effectively enhance the model’s feature representation and boundary localization capabilities. Specifically, the R-Mamba backbone network combines the advantages of statespace modeling and convolutional feature extraction, achieving efficient fusion of global context and local details. The MGA module dynamically captures multi-scale semantic information through dilated convolutions and gating mechanisms, enhancing the model’s adaptability to targets of different scales and shapes. The BE module significantly strengthens boundary representation and fine-grained structural segmentation through multi-scale convolutions and channel-spatial dual attention mechanisms. Additionally, this paper designs a multi-loss function joint optimization strategy to comprehensively constrain region overlap, pixel classification, and structural consistency. Experimental validation on ISIC skin lesion and LGG brain tumor datasets shows competitive performance compared to several mainstream models under the tested conditions.
Guoqiang Ren, Qi Wang 0061, Jieying Tu, Pengxiang Su, Hengrui Liu, Di Gai, Peng Luo 0005, Shuxiao Li
IEEE Internet Things J.9
2026 Decoupled dual-branch prototype network with fine-grained context mining for self-supervised few-shot abdominal image segmentation
Di Gai, Yuxuan Zou, Jieying Tu, Yuhan Geng, Dengyun Xu, Pengxiang Su
Image Vis. Comput.1
2026 Dual-Student Adversarial Framework With Discriminator and Consistency-Driven Learning for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised medical image segmentation is essential for alleviating the cost of manual annotation in clinical applications. However, existing methods often suffer from unreliable pseudo-labels and confirmation bias in consistency-based training, which can lead to unstable optimization and degraded performance. To address these issues, a novel method named dual-Student adversarial framework with discriminator and consistency-driven learning for semi-supervised medical image segmentation is proposed. Specifically, an adversarial learning-based segmentation refinement (ALSR) module is designed to encourage prediction diversity between two student networks and leverage a shared discriminator for adversarial refinement of pseudo-labels. To further stabilize the consistency process, a residual exponential moving average (R-EMA) is applied in the uncertainty estimation with inter-instance consistency measurement (UIM) module to construct a robust teacher model, while noisy voxel predictions are selectively filtered based on uncertainty estimation. In addition, a Contrastive Representation Stabilization (CRS) module is developed to enhance voxel-level semantic alignment by performing contrastive learning only on confident regions, improving feature discriminability and structural consistency. Extensive experiments on benchmark datasets demonstrate that our method consistently outperforms prior state-of-the-art approaches.
Haifan Wu, Yuhan Geng, Di Gai, Jieying Tu, Xin Xiong 0016, Qi Wang 0061
IEEE J. Biomed. Health Informatics3
2025 CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identification
abstract
With the emergence of vision-language pre-trained models, such as CLIP, some textual prompts have been gradually introduced recently into re-identification (Re-ID) tasks to obtain considerably robust multimodal information. However, most textual descriptions based on vehicle Re-ID tasks only contain identity index words without specific words to describe vehicle view information, thereby resulting in difficulty to be widely applied in vehicle Re-ID tasks with view variations. This case inspires us to propose a CLIP-driven view-aware prompt learning framework for unsupervised vehicle Re-ID. We first design a learnable textual prompt template called view-aware context optimization (ViewCoOp) based on dynamic multi-view word embeddings, which can fully obtain the proportion and position encoding of each view in the whole vehicle body region. Subsequently, a cross-modal mutual graph is constructed to explore the connections between inter-modal and intra-modal. Each sample is treated as a graph node, which extracts textual features based on ViewCoOp and the visual features of images. Moreover, leveraging the inter-cluster and intra-cluster correlation in the bimodal clustering results in the determination of connectivity between graph node pairs. Lastly, the proposed cross-modal mutual graph method utilizes supervised information from the bimodal gap to directly fine-tune the image encoder of CLIP for downstream unsupervised vehicle Re-ID tasks. Extensive experiments verify that the proposed method is capable of effectively obtaining cross-modal description ability from multiple views.
Jiyang Xu, Di Gai, Ruihua Zhou
AAAI4
2025 Vehiclemae: View-Asymmetry Mutual Learning for Vehicle Re-Identification Pre-Training Via Masked Autoencoders
Qi Wang 0061, Dong Wang 0080, Di Gai, Xin Xiong 0016, Jiyang Xu, Ruihua Zhou
ICCV4
2025 Threefold Encoder Interaction: Hierarchical Multi-Grained Semantic Alignment for Cross-Modal Food Retrieval
abstract
Current cross-modal food retrieval approaches focus mainly on the global visual appearance of food without explicitly considering multi-grained information. Additionally, direct calculation of the global similarity of image-recipe pairs is not particularly effective in terms of latent alignment, which suffers from mismatch during the mutual image-recipe retrieval process. This paper proposes a threefold encoder interaction (TEI) cross-modal food retrieval framework to maintain the multi-granularity of food images and the multi-levels of textual recipes to address the aforementioned challenges. The TEI framework comprises an image encoder, a recipe encoder, and a multi-grained interaction encoder. We simultaneously propose a multi-grained relation-aware attention (MRA) embedded in the multi-grained interaction encoder to capture multi-grained food visual features. The multi-grained interaction similarity scores are calculated to better establish the multi-grained correlation between recipe and image entities based on the extracted hierarchical textual and multi-grained visual features. Finally, a hierarchical multi-grained semantic alignment loss is designed to supervise the whole process of cross-modal training using the multi-grained interaction similarity scores. Extensive qualitative and quantitative experiments on the Recipe1M dataset have demonstrated that the proposed TEI framework achieves multi-grained semantic alignment between image and text modalities and is superior to other state-of-the-art methods in cross-modal food retrieval tasks.
Qi Wang 0061, Dong Wang 0080, Weidong Min, Di Gai, Cheng Zha, Yuling Zhong
IEEE Trans. Multim.4
2024 SAM-driven MAE pre-training and background-aware meta-learning for unsupervised vehicle re-identification
abstract
Distinguishing identity-unrelated background information from discriminative identity information poses a challenge in unsupervised vehicle re-identification (Re-ID). Re-ID models suffer from varying degrees of background interference caused by continuous scene variations. The recently proposed segment anything model (SAM) has demonstrated exceptional performance in zero-shot segmentation tasks. The combination of SAM and vehicle Re-ID models can achieve efficient separation of vehicle identity and background information. This paper proposes a method that combines SAM-driven mask autoencoder (MAE) pre-training and background-aware meta-learning for unsupervised vehicle Re-ID. The method consists of three sub-modules. First, the segmentation capacity of SAM is utilized to separate the vehicle identity region from the background. SAM cannot be robustly employed in exceptional situations, such as those with ambiguity or occlusion. Thus, in vehicle Re-ID downstream tasks, a spatially-constrained vehicle background segmentation method is presented to obtain accurate background segmentation results. Second, SAM-driven MAE pre-training utilizes the aforementioned segmentation results to select patches belonging to the vehicle and to mask other patches, allowing MAE to learn identity-sensitive features in a self-supervised manner. Finally, we present a background-aware meta-learning method to fit varying degrees of background interference in different scenarios by combining different background region ratios. Our experiments demonstrate that the proposed method has state-of-the-art performance in reducing background interference variations.
Dong Wang 0080, Qi Wang 0061, Weidong Min, Di Gai, Yuhan Geng
Comput. Vis. Media4
2024 Vision-language constraint graph representation learning for unsupervised vehicle re-identification
Dong Wang 0080, Qi Wang 0061, Zhiwei Tu, Weidong Min, Xin Xiong 0016, Yuling Zhong, Di Gai
Expert Syst. Appl.7
2024 Feature ensemble network for medical image segmentation with multi-scale atrous transformer
abstract
Abstract Recent years have witnessed notable advancements in medical image segmentation through deep convolutional neural networks. However, a notable limitation lies in the local operation of convolution, which hinders the ability to fully exploit global semantic information. To overcome the challenges prevalent in medical image segmentation, the feature ensemble network with multi‐scale atrous transformer is proposed. At the core of the approach lies the multi‐scale contextual integration module, which is based on the multi‐scale atrous transformer and facilitates contextual integration of multi‐level features. To extract discriminative fine‐grained features of the target region, a hybrid attention mechanism that synergistically combines spatial and channel attention, thereby sharpening the model's focus on crucial target information within high‐level features, is incorporated. Additionally, the channel‐aware feature reconstruction module is introduced as an innovative component engineered to tackle feature similarity issues across different categories. This module performs feature reconstruction based on channel perception, effectively widening the feature gap between categories and enhancing the segmentation capability. It is worth mentioning that our approach surpasses the state‐of‐the‐art method using three benchmark datasets in medical image segmentation.
Di Gai, Yuhan Geng, Xin Xiong 0016, Ruihua Zhou, Qi Wang 0061
IET Image Process.1
2024 Semi-supervised contextual cognitive augmentation-based cross-teaching network for multiclass medical image segmentation
abstract
Abstract The application of medical image segmentation technology enables accurate localization of human tissues, providing doctors with a reliable foundation for diagnosis. While deep learning methods have proven effective in this task, most current approaches rely on a single prediction framework, which overlooks Edge semantic features and results in flawed texture features. Moreover, existing supervised methods face challenges due to limited availability of high‐quality annotations in the field of medical imaging. In this article, a Semi‐supervised Contextual Cognitive Augmentation‐based Cross‐teaching Network is proposed. A Contextual Cognitive Enhancement Module is introduced consisting of two components: data augmentation and information extraction. The data augmentation component provides multi‐level data distribution by incorporating diverse perturbation strategies such as Discrete Cosine Transform and Gaussian noise. The information extraction component employs the Comprehensive Information Extraction module, which consists of Global Perception Information Extraction module and Multi‐channel Information Extraction module to extract perceptual information from images and enhance interaction between image channels, respectively. Additionally, a cross‐teaching strategy is adopted and a hybrid loss function is utilized to encourage knowledge sharing among the networks, leveraging the advantages of dual networks for improved performance. Experimental results demonstrate significant enhancements in multiclass medical image segmentation compared to several state‐of‐the‐art single‐framework networks.
Di Gai, Yusong Xiao, Yuhan Geng, Xin Xiong 0016, An-qi Zhong
IET Image Process.1
2024 Semi-supervised medical image classification based on class prototype matching for soft pseudo labels with consistent regularization
Di Gai, Ruonan Xiong, Weidong Min, Qi Wang 0061, Xin Xiong 0016, Chunjiang Peng
Multim. Tools Appl.1
2023 M-AResNet: a novel multi-scale attention residual network for melting curve image classification
Pengxiang Su, Xuanjing Shen, Haipeng Chen 0002, Di Gai, Yu Liu 0004
Multim. Tools Appl.4
2023 Dual similarity pre-training and domain difference encouragement learning for vehicle re-identification in the wild
Qi Wang 0061, Yuling Zhong, Weidong Min, Di Gai
Pattern Recognit.5
2023 Spatiotemporal Learning Transformer for Video-Based Human Pose Estimation
abstract
Multi-frame human pose estimation has long been an appealing and fundamental issue in visual perception. Owing to the frequent rapid motion and pose occlusion in videos, this task is extremely challenging. Current state-of-the-art methods seek to model spatiotemporal features by equally fusing each frame in the local sequence, which weakens the target frame information. In addition, existing approaches usually emphasize more on deep features while ignoring the detailed information implied in the shallow feature maps, resulting in the dropping of crucial features. To address the above problems, we propose an effective framework, namely spatiotemporal learning transformer for video-based human pose estimation (SLT-Pose), which consists of a Personalized Feature Extraction Module (PFEM), Self-feature Refinement Module (SRM), Cross-frame Temporal Learning Module (CTLM) and Disentangled Keypoint Detector (DKD). To be specific, we propose PFEM which extracts and modulates the individual frame features to adapt to the varying human shape, and integrates single-frame features to obtain the spatiotemporal features. We further present SRM to establish global correlation spatial cues on the target frame to attain the refinement feature. Then, a CTLM is designed to search for the information most closely related to the target frame from the spatiotemporal features to intensify the interaction between the target frame and the local sequence, using both the shallow detailed and the deep semantic representations. Finally, we employ DKD to extract the disentangled characteristics of each joint and encode the articulated joint pairs in the human body, promoting the model to reasonably and accurately predict the keypoint heatmaps. Extensive experiments on three huamn motion benchmarks, including PoseTrack2017, PoseTrack2018, and Sub-JHMDB dataset, demonstrate that SLT-Pose plays favorably against state-of-the-art approaches in terms of both objective evaluation and subjective visual performance.
Di Gai, Runyang Feng, Weidong Min, Xiaosong Yang, Pengxiang Su, Qi Wang 0061
IEEE Trans. Circuits Syst. Video Technol.1
2023 Trade-off background joint learning for unsupervised vehicle re-identification
Qi Wang 0061, Weidong Min, Di Gai, Haowen Luo
Vis. Comput.5
2020 Multi-focus noisy image fusion based on gradient regularized convolutional sparse representatione
abstract
The method proposes a multi-focus noisy image fusion algorithm combining gradient regularized convolutional sparse representatione and spatial frequency. Firstly, the source image is decomposed into a base layer and a detail layer through two-scale image decomposition. The detail layer uses the Alternating Direction Method of Multipliers (ADMM) to solve the convolutional sparse coefficients with gradient penalties to complete the fusion of detail layer coefficients. Then, The base layer uses the spatial frequency to judge the focus area, the spatial frequency and the "choose-max" strategy are applied to achieved the multi-focus fusion result of base layer. Finally, the fused image is calculated as a superposition of the base layer and the detail layer. Experimental results show that compared with other algorithms, this algorithm provides excellent subjective visual perception and objective evaluation metrics.
Xuanjing Shen, Haipeng Chen 0002, Di Gai
MMAsia4
2020 Medical image fusion using the PCNN based on IQPSO in NSST domain
abstract
In this study, an improved quantum‐behaved particle swarm optimisation based pulse‐coupled neural network (IQPSO‐PCNN) is proposed in the non‐subsampled shearlet transform (NSST) domain for medical image fusion. First, NSST tool is used to decompose the source image into low‐frequency and high‐frequency subbands. Then, for low‐frequency subbands, the fusion rules of two different functions are presented, which simultaneously addresses two key issues of energy preservation and detail extraction. For high‐frequency subbands, unlike conventional PCNN‐based methods, parameters are manually set based on experience, and the decomposed high‐frequency subbands share a set of parameters. The IQPSO‐PCNN model can obtain the optimal parameters for each high‐frequency subband adaptively according to its own information. Finally, the fused low‐frequency subband and high‐frequency subbands are inversely transformed by NSST to acquire the final fused image. The proposed algorithm uses >90 pairs of images with four different modalities. In addition, fusion experiments are performed on different sequences of the three modes. The experimental results demonstrate that the proposed method is superior to existing state‐of‐art methods in subjective visual performance and objective evaluation.
Di Gai, Xuanjing Shen, Haipeng Chen 0002, Zeyu Xie, Pengxiang Su
IET Image Process.1
2020 Multi-focus image fusion method based on two stage of convolutional neural network
Di Gai, Xuanjing Shen, Haipeng Chen 0002, Pengxiang Su
Signal Process.1