EDBT 2026 Demo / reviewers in the wild / expert
Xudong Jia 0001
dblp:152/9995-1
· DBLP profile ↗
14ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0001-7911-8869ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DLGCNet: Multimodal remote sensing semantic segmentation via dual diagonal low-rank adaptation and graph convolutional feature fusion
Jun-Ying Zeng, Xudong Jia 0001, Bin Deng 0003, Yikui Zhai, Chuanbo Qin, Pasquale Coscia, Angelo Genovese |
Knowl. Based Syst. | 3 |
| 2026 | Activation Trained Layer Based Controls Continual Knowledge Editing Approach
Yimin Wen, Jun-Ying Zeng, Xudong Jia 0001, Tao Chen 0026 |
Mach. Learn. | 4 |
| 2026 | DUR-Net+: Semi-Supervised Abdominal CT Pheochromocytoma Segmentation via Dynamic Uncertainty Rectified and Prior Knowledge From SAM-Med3DabstractPheochromocytoma is a rare urological adrenal tumor disease. Automated segmentation of pheochromocytomas from computed tomography (CT) is essential for diagnosis and treatment. However, this task is a challenging one due to issues such as blurred boundaries, irregular shapes, variations in location and size, and the lack of annotated images for training. To address these issues, we propose a semi-supervised framework for pheochromocytoma segmentation that primarily consists of a dynamic uncertainty rectification mechanism and a supervised strategy based on SAM-Med3D prior knowledge. First, we design a semi-supervised segmentation model comprising a shared encoder and multiple independent decoders that dynamically select pseudo labels from the different decoder outputs. To mitigate the risk of unreliable predictions caused by sparse annotations during training, we introduce uncertainty estimation to prioritize reliable outputs. Additionally, an Attentional Convolution Block (ACB) is designed in the encoding stage to fully utilize both global and local features, improving tumor recognition in segmentation. Furthermore, SAM-Med3D prior knowledge is incorporated into the framework as supplementary supervisory information, aiding the model in learning from limited labeled data. To eliminate the labor-intensive requirement for manual prompts in SAM-Med3D, we leverage pseudo labels to generate high-quality mask prompts, thus transforming the clinical workflow. Experiments on two pheochromocytoma datasets from different centers demonstrate that our proposed method achieves competitive performance. Chuanbo Qin, Zhuyuan Chen, Dong Wang 0083, Jun-Ying Zeng, Xudong Jia 0001, Maoqing Hu, Yikui Zhai, Pasquale Coscia, Angelo Genovese |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Multi-label classification of tongue images using label semantic embedding and dual-branch network
Yue Feng 0003, Xudong Jia 0001, Tao Chen 0026 |
Multim. Syst. | 3 |
| 2025 | Spatial Reconstruction and Joint Training in Transformer Network for Cross-Domain Remote Sensing Images Semantic SegmentationabstractRecently, Unsupervised Domain Adaptation (UDA) methods have attracted considerable attention in Remote Sensing Images (RSI) semantic segmentation. However, cross-domain RSI exhibit diverse scales, imbalanced distributions within domains, and significant inter-domain variations. In response to these challenges, we combine Spatial reconstruction and Joint training with the Transformer Network (SJT-Net). This framework introduces a spatial reconstruction method to address the issue of inconsistent ground sampling distances in cross domain RSI, which is rarely considered in existing approaches. Transferring domain knowledge at a similar spatial scale improves the spatial representation ability of UDA models. Unlike traditional adversarial training using ResNet for feature extraction, the SJT-Net employs Segformer, which enhances the model’s ability to capture in-class features across domains and improves global dependency modeling. Transmitting these refined features to the discriminator allows for more precise feature-level domain alignment. To enhance feature decoding, an interactive global-local decoder is constructed to efficiently capture both global relationships and local details of landform objects. Our framework leverages adversarial training to generate highly confident model weights and pseudo-labels for self-training in the target domain. Through iterative updates, the model’s generalization capability is gradually improved, eventually achieving optimal segmentation performance. Experimental results demonstrate that SJT-Net outperforms current UDA approaches and accomplishes state-of-the-art (SOTA) segmentation accuracy. The repository can be accessed at https://github.com/AnsonD0820/SJT-Net. Jun-Ying Zeng, Senyao Deng, Yikui Zhai, Xudong Jia 0001, Chuanbo Qin, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Multimodal Feature Fusion Network With Text Difference Enhancement for Remote Sensing Change DetectionabstractAlthough deep learning has advanced remote sensing change detection (RSCD), most methods rely solely on image modality, limiting feature representation, change pattern modeling, and generalization—especially under illumination and noise disturbances. To address this, we propose MMChange, a multimodal RSCD method that combines image and text modalities to enhance accuracy and robustness. An Image Feature Refinement (IFR) module is introduced to highlight key regions and suppress environmental noise. To overcome the semantic limitations of image features, we employ a vision-language model (VLM) to generate semantic descriptions of bi-temporal images. A Textual Difference Enhancement (TDE) module then captures fine-grained semantic shifts, guiding the model toward meaningful changes. To bridge the heterogeneity between modalities, we design an Image-Text Feature Fusion (ITFF) module that enables deep cross-modal integration. Extensive experiments on LEVIR-CD, WHU-CD, and SYSU-CD demonstrate that MMChange consistently surpasses state-of-the-art methods across multiple metrics, validating its effectiveness for multimodal RSCD. Code is available at: https://github.com/yikuizhai/MMChange. Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou 0001, Xudong Jia 0001, Hongsheng Zhang 0001, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | 3D Vehicle Detection in Roadside Traffic Flow Using Complex-YOLOabstractVehicle detection and classification are important for intelligent transportation planning, directly impacting traffic flow and management efficiency. There have been a lot of research on vehicle detection and classification for autonomous driving, but research on roadside traffic flow has not been fully investigated. This research explores various methods for vehicle detection and classification that could be applied to 3D roadside traffic flow analysis. Using data collected from a Velodyne 32C Lidar sensor 3D camera and a stereo-based depth 2D camera, this project implemented a Complex-YOLO projection-based model, by converting 3D point cloud data into a 2D bird-eye-view projection and applying the YOLO model for vehicle detection and classification. Our experimental results demonstrated that using transfer learning as a training technique, along with normalization of the rotation angle, enhanced the model performance for vehicle detection and classification. Without normalization of the rotation angle, the highest precision achieved was 89.71 %, with a recall of 85.15 %. With normalization of the rotation angle, the best model achieved a precision of 92.20 % with a recall of 84.90 %, demonstrating the effectiveness of both transfer learning and normalization in handling and aligning the dataset for more effective learning. Jonathan Cordova, Xunfei Jiang, Xudong Jia 0001 |
ICMLA | 3 |
| 2024 | MSFA-Net : Multiple Spatial-Channel Feature Aggregation Network for Change Detection and a UAV-CD DatasetabstractChange detection in remote sensing images is pivotal for monitoring and comprehending dynamic environmental phenomena. Nonetheless, conventional change detection models have grappled with correlating information between channel and spatial dimensions due to inherent feature extraction limitations. Hence, this paper proposes an innovative change detection framework and a high resolution UAV change detection dataset named UAV-CD dataset. The introduced network embraces a Siamese network, amalgamating a feature extraction backbone network along with spatial and channel reconstruction convolution (ScConv) and Ghost modules. A primary contribution of this paper is the incorporation of ScConv into the change detection network, facilitating the reconstruction of information in both spatial and channel dimensions. Additionally, the Ghost module is employed to fortify information across distinct channel dimensions within the feature maps. Compared with current state-of-the-art methods, it is indicated that the proposed approach achieves superior performance on the LEVIR-CD, SYSU-CD, and our proposed dataset UAV-CD. Yikui Zhai, Haolin Lv, Tingfeng Xian, Zilu Ying, Hao Quan 0002, Xudong Jia 0001 |
IGARSS | 7 |
| 2024 | A Scale-Temporal Interaction Network For Remote Sensing Image Change Detection And A UAV-CD DatasetabstractRemote sensing (RS) image change detection (CD) is a challenging visual task due to its rich and complex image information. Nowadays, CD has yielded fruitful results. However, insufficient feature interaction hinders further improvement of CD performance. In this paper, we introduce a scale-temporal interaction network (STI-Net). It extracts multi-scale bitemporal features using a depth-separable convolution-based Siamese encoder, followed by both Cross-Scale Feature Interaction (CSFI) and Cross-Temporal Feature Interaction (CTFI). Finally, we employ a straightforward decoder to generate the change map. Additionally, to enrich the CD data, we introduced a new dataset based on UAV optical image, named UAV-CD. This dataset comprises 2660 pairs of images sized at 768×768 pixels, focusing primarily on building and land changes. Experiments demonstrate that our method outperforms existing state-of-the-art methods on two public CD datasets as well as UAV-CD, showcasing excellent performance. Tingfeng Xian, Zilu Ying, Haolin Lv, Yikui Zhai, Hao Quan 0002, Xudong Jia 0001 |
IGARSS | 7 |
| 2024 | Semi-supervised segmentation for primary nasopharyngeal carcinoma tumors using local-region constraint and mixed feature-level consistency
Jun-Ying Zeng, Xiuping Zhang, Xudong Jia 0001, Chuanbo Qin |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | PheoSeg: A 3D transfer learning framework for accurate abdominal CT pheochromocytoma segmentation and surgical grade prediction
Dong Wang 0083, Jun-Ying Zeng, Guolin Huang, Dong Xu 0023, Xudong Jia 0001, Chuanbo Qin |
Knowl. Based Syst. | 5 |
| 2024 | DGMA2-Net: A Difference-Guided Multiscale Aggregation Attention Network for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) focuses on identifying regions that have undergone changes between two remote sensing images captured at different times. Recently, convolutional neural networks (CNNs) have shown promising results in the challenging task of RSCD. However, these methods do not efficiently fuse bitemporal features and extract useful information that is beneficial to subsequent RSCD tasks. In addition, they did not consider multilevel feature interactions in feature aggregation and ignore relationships between difference features and bitemporal features, which thus affects the RSCD results. To address the above problems, a difference-guided multiscale aggregation attention network, DGMA2-Net, is developed. Bitemporal features at different levels are extracted through a Siamese convolutional network and a multiscale difference fusion module (MDFM) is then created to fuse bitemporal features and extract, in a multiscale manner, difference features containing rich contextual information. After the MDFM treatment, two difference aggregation modules (DAMs) are used to aggregate difference features at different levels for multilevel feature interactions. The features through DAMs are sent to the difference-enhanced attention modules (DEAMs) to strengthen the connections between bitemporal features and difference features and further refine change features. Finally, refined change features are superimposed from deep to shallow and a change map is produced. In validating the effectiveness of DGMA2-Net, a series of experiments are conducted on three public RSCD benchmark datasets (LEVIR-CD, BCDD, and SYSU-CD). The experimental results demonstrate that DGMA2-Net surpasses the current eight state-of-the-art methods in RSCD. Our code is released at https://github.com/yikuizhai/DGMA2-Net. Zilu Ying, Zijun Tan, Yikui Zhai, Xudong Jia 0001, Wenba Li, Jun-Ying Zeng, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DS-HyFA-Net: A Deeply Supervised Hybrid Feature Aggregation Network With Multiencoders for Change Detection in High-Resolution ImageryabstractWith the advancement of deep learning (DL) technologies, remarkable progress has been achieved in change detection (CD). Existing DL-based methods primarily focus on the discrepancy in bitemporal images, while overlooking the commonality in bitemporal images. However, one of the reasons hindering the improvement of CD performance is the inadequate utilization of image information. To address the above issue, we propose a Deeply Supervised Hybrid Feature Aggregation Network (DS-HyFA-Net). This network predicts changes by integrating the distinctness and the commonality in bitemporal images. Specifically, the DS-HyFA-Net primarily consists of a set of encoders and a Hybrid Feature Aggregation (HyFA) module. It uses a Siamese encoder (or Encoder I) and a specialized encoder (or Encoder II) to extract distinct and common features (CFs) in bitemporal images, respectively. The HyFA module efficiently aggregates distinct and common features (or hybrid features) and generates a change map using a predictor. In addition, a common feature learning strategy (CFLS) is introduced, based on deeply supervised (DS) techniques, to guide Encoder II in learning CFs. Experimental results on three well-recognized datasets demonstrate the effectiveness of the innovative DS-HyFA-Net, achieving F1-Scores of 93.33% on WHU-CD, 90.98% on LEVIR-CD, and 81.14% on SYSU-CD. Our code is available athttps://github.com/yikuizhai/DS-HyFA-Net. Zilu Ying, Tingfeng Xian, Yikui Zhai, Xudong Jia 0001, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | CAS-Net: Comparison-Based Attention Siamese Network for Change Detection With an Open High-Resolution UAV Image DatasetabstractChange detection (CD) is a process of extracting changes on the Earth’s surface from bitemporal images. Current CD methods that use high-resolution remote sensing images require extensive computational resources and are vulnerable to the presence of irrelevant noises in the images. In addressing these challenges, a comparison-based attention Siamese network (CAS-Net) is proposed. The network utilizes contrastive attention modules (CAMs) for feature fusion and employs a classifier to determine similarities and differences of bitemporal image patches. It simplifies pixel-level CDs by comparing image patches. As such, the influences of image background noises on change predictions are reduced. Along with the CAS-Net, an unmanned aerial vehicle (UAV) similarity detection (UAV-SD) dataset is built using high-resolution remote sensing images. This dataset, serving as a benchmark for CD, comprises 10000 pairs of UAV images with a size of$256 \times 256$. Experiments of the CAS-Net on the UAV-SD dataset demonstrate that the CAS-Net is superior to other baseline CD networks. The CAS-Net detection accuracy is 93.1% on the UAV-SD dataset. The code and the dataset can be found athttps://github.com/WenbaLi/CAS-Net. Yikui Zhai, Wenba Li, Tingfeng Xian, Xudong Jia 0001, Hongsheng Zhang 0001, Zijun Tan, Jun-Ying Zeng, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |