EDBT 2026 Demo / reviewers in the wild / expert
Xukun Zhang
dblp:257/7646
· DBLP profile ↗
19ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-2869-9434ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nested resolution mesh-graph CNN for automated extraction of liver surface anatomical landmarksabstractThe anatomical landmarks on the liver (mesh) surface, including the falciform ligament and liver ridge, are composed of triangular meshes of varying shapes, sizes, and positions, making them highly complex. Extracting and segmenting these landmarks is critical for augmented reality-based intraoperative navigation and monitoring. The key to this task lies in comprehensively understanding the overall geometric shape and local topological information of the liver mesh. However, due to the liver's variations in shape and appearance, coupled with limited data, deep learning methods often struggle with automatic liver landmark segmentation. To address this, we propose a two-stage automatic framework combining mesh-CNN and graph-CNN. In the first stage, dynamic graph convolution (DGCNN) is employed on low-resolution meshes to achieve rapid global understanding, generating initial landmark proposals at two levels, "dilation" and "erosion", and mapping them onto the original high-resolution surface. Subsequently, a refinement network based on mesh convolution fuses these landmark proposals from edge features along the local topology of the high-resolution mesh surface, producing refined segmentation results. Additionally, we incorporate an anatomy-aware Dice loss to address resolution imbalance and better handle sparse anatomical regions. Extensive experiments on two liver datasets, both in-distribution and out-of-distribution, demonstrate that our method accurately processes liver meshes of different resolutions, outperforming state-of-the-art methods. The reconstructed liver mesh dataset and the source code are available at https://github.com/xukun-zhang/MeshGraphCNN. Xukun Zhang, Jinghui Feng, Peng Liu 0074, Minghao Han, Yanlan Kang, Sharib Ali, Lihua Zhang 0002 |
Medical Image Anal. | 1 |
| 2026 | Towards unified molecule-enhanced pathology image representation learning via integrating spatial transcriptomics
Minghao Han, Dingkang Yang, Jiabei Cheng, Xukun Zhang, Zizhi Chen, Haopeng Kuang, Lihua Zhang 0002 |
Pattern Recognit. | 4 |
| 2025 | Heterogeneous Knowledge and Global-Local Structure Integration Framework for Drug RepositioningabstractDrug repositioning offers a cost-effective strategy for identifying new therapeutic uses for existing drugs. However, existing methods struggle to fully utilize domain knowledge and structural information to capture complex network and entity features. In this work, we propose a Heterogeneous Knowledge and Global-Local Structure Integration Framework (HKGLS) for drug repositioning. HKGLS leverages graph neural networks to represent multi-source knowledge from proteins, diseases, and drugs, then utilizes a heterogeneous graph structure for knowledge fusion. Specifically, we constructed protein structure graphs based on binding pocket residues to achieve fine-grained protein representation. To comprehensively capture complex network structures, HKGLS also integrates global semantic and pharmacological information as well as local structural features through graph neural networks. Extensive experiments on three public drug repositioning datasets show that HKGLS outperforms eight state-of-the-art baseline methods across various evaluation metrics, validating its effectiveness and superiority. Xukun Zhang, Hongzhi Liu 0001, Zhonghai Wu |
BIBM | 1 |
| 2025 | Collaborative Adaptive Metric Learning for Personalized Recommendation
Xukun Zhang, Zhaoyu Zhou, Hongzhi Liu 0001 |
ICIC (8) | 1 |
| 2025 | VGAT: A Cancer Survival Analysis Framework Transitioning from Generative Visual Question Answering to Genomic ReconstructionabstractMultimodal learning combining pathology images and genomic sequences enhances cancer survival analysis but faces clinical implementation barriers due to limited access to genomic sequencing in under-resourced regions. To enable survival prediction using only whole-slide images (WSI), we propose the Visual-Genomic Answering-Guided Transformer (VGAT), a framework integrating Visual Question Answering (VQA) techniques for genomic modality reconstruction. By adapting VQA’s text feature extraction approach, we derive stable genomic representations that circumvent dimensionality challenges in raw genomic data. Simultaneously, a cluster-based visual prompt module selectively enhances discriminative WSI patches, addressing noise from unfiltered image regions. Evaluated across five TCGA datasets, VGAT outperforms existing WSI-only methods, demonstrating the viability of genomic-informed inference without sequencing. This approach bridges multimodal research and clinical feasibility in resource-constrained settings. The code link is https://github.com/CZZZZZZZZZZZZZZZZZ/VGAT. Zizhi Chen, Minghao Han, Xukun Zhang, Shuwei Ma, Tao Liu 0050, Lihua Zhang 0002 |
ICME | 3 |
| 2025 | VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual StainingabstractIn histopathology, tissue sections are typically stained using common H&E staining or special stains (MAS, PAS, PASM, etc. ) to clearly visualize specific tissue structures. The rapid advancement of deep learning offers an effective solution for generating virtually stained images, significantly reducing the time and labor costs associated with traditional histochemical staining. However, a new challenge arises in separating the fundamental visual characteristics of tissue sections from the visual differences induced by staining agents. Additionally, virtual staining often overlooks essential pathological knowledge and the physical properties of staining, resulting in only style-level transfer. To address these issues, we introduce, for the first time in virtual staining tasks, a pathological vision-language large model (VLM) as an auxiliary tool. We integrate contrastive learnable prompts, foundational concept anchors for tissue sections, and staining-specific concept anchors to leverage the extensive knowledge of the pathological VLM. This approach is designed to describe, frame, and enhance the direction of virtual staining. Furthermore, we have developed a data augmentation method based on the constraints of the VLM. This method utilizes the VLM's powerful image interpretation capabilities to further integrate image style and structural information, proving beneficial in high-precision pathological diagnostics. Extensive evaluations on publicly available multi-domain unpaired staining datasets demonstrate that our method can generate highly realistic images and enhance the accuracy of downstream tasks, such as glomerular detection and segmentation. Our code. https://github.com/CZZZZZZZZZZZZZZZZZ/VPGAN-HARBOR is available. Zizhi Chen, Minghao Han, Yizhou Liu 0002, Ziyun Qian, Xukun Zhang, Jingwei Wei, Lihua Zhang 0002 |
ACM Multimedia | 7 |
| 2025 | MDFGNN-SMMA: prediction of potential small molecule-miRNA associations based on multi-source data fusion and graph neural networksabstractBACKGROUND: MicroRNAs (miRNAs) are pivotal in the initiation and progression of complex human diseases and have been identified as targets for small molecule (SM) drugs. However, the expensive and time-intensive characteristics of conventional experimental techniques for identifying SM-miRNA associations highlight the necessity for efficient computational methodologies in this field. RESULTS: In this study, we proposed a deep learning method called Multi-source Data Fusion and Graph Neural Networks for Small Molecule-MiRNA Association (MDFGNN-SMMA) to predict potential SM-miRNA associations. Firstly, MDFGNN-SMMA extracted features of Atom Pairs fingerprints and Molecular ACCess System fingerprints to derive fusion feature vectors for small molecules (SMs). The K-mer features were employed to generate the initial feature vectors for miRNAs. Secondly, cosine similarity measures were computed to construct the adjacency matrices for SMs and miRNAs, respectively. Thirdly, these feature vectors and adjacency matrices were input into a model comprising GAT and GraphSAGE, which were utilized to generate the final feature vectors for SMs and miRNAs. Finally, the averaged final feature vectors were utilized as input for a multilayer perceptron to predict the associations between SMs and miRNAs. CONCLUSIONS: The performance of MDFGNN-SMMA was assessed using 10-fold cross-validation, demonstrating superior compared to the four state-of-the-art models in terms of both AUC and AUPR. Moreover, the experimental results of an independent test set confirmed the model's generalization capability. Additionally, the efficacy of MDFGNN-SMMA was substantiated through three case studies. The findings indicated that among the top 50 predicted miRNAs associated with Cisplatin, 5-Fluorouracil, and Doxorubicin, 42, 36, and 36 miRNAs, respectively, were corroborated by existing literature and the RNAInter database. Xukun Zhang |
BMC Bioinform. | 2 |
| 2025 | An objective comparison of methods for augmented reality in laparoscopic liver resection by preoperative-to-intraoperative image fusion from the MICCAI2022 challengeabstractAugmented reality for laparoscopic liver resection is a visualisation mode that allows a surgeon to localise tumours and vessels embedded within the liver by projecting them on top of a laparoscopic image. Preoperative 3D models extracted from Computed Tomography (CT) or Magnetic Resonance (MR) imaging data are registered to the intraoperative laparoscopic images during this process. Regarding 3D-2D fusion, most algorithms use anatomical landmarks to guide registration, such as the liver's inferior ridge, the falciform ligament, and the occluding contours. These are usually marked by hand in both the laparoscopic image and the 3D model, which is time-consuming and prone to error. Therefore, there is a need to automate this process so that augmented reality can be used effectively in the operating room. We present the Preoperative-to-Intraoperative Laparoscopic Fusion challenge (P2ILF), held during the Medical Image Computing and Computer Assisted Intervention (MICCAI 2022) conference, which investigates the possibilities of detecting these landmarks automatically and using them in registration. The challenge was divided into two tasks: (1) A 2D and 3D landmark segmentation task and (2) a 3D-2D registration task. The teams were provided with training data consisting of 167 laparoscopic images and 9 preoperative 3D models from 9 patients, with the corresponding 2D and 3D landmark annotations. A total of 6 teams from 4 countries participated in the challenge, whose results were assessed for each task independently. All the teams proposed deep learning-based methods for the 2D and 3D landmark segmentation tasks and differentiable rendering-based methods for the registration task. The proposed methods were evaluated on 16 test images and 2 preoperative 3D models from 2 patients. In Task 1, the teams were able to segment most of the 2D landmarks, while the 3D landmarks showed to be more challenging to segment. In Task 2, only one team obtained acceptable qualitative and quantitative registration results. Based on the experimental outcomes, we propose three key hypotheses that determine current limitations and future directions for research in this domain. Sharib Ali, Yamid Espinel, Yueming Jin, Peng Liu 0074, Bianca Güttner, Xukun Zhang, Lihua Zhang 0002, Thomas Dowrick, Matthew J. Clarkson, Shiting Xiao, Yifan Wu 0021, Lei Zhu 0003, Dai Sun, Micha Pfeiffer, Shahid Farid, Lena Maier-Hein, Emmanuel Buc, Adrien Bartoli |
Medical Image Anal. | 6 |
| 2025 | Enhanced multi-modal abdominal image registration via structural awareness and region-specific optimization
Xukun Zhang, Lihua Zhang 0002 |
Pattern Recognit. Lett. | 2 |
| 2025 | Denoising Transformer for BEV 3D Object Detection via Multiview Multiscale Cross-Attention
Xukun Zhang, Xiaoyi Wei, Zhi Xu 0010, Youxing Wang, Peng Zhai, Lihua Zhang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | MSCPT: Few-Shot Whole Slide Image Classification With Multi-Scale and Context-Focused Prompt TuningabstractMultiple instance learning (MIL) has become a standard paradigm for the weakly supervised classification of whole slide images (WSIs). However, this paradigm relies on using a large number of labeled WSIs for training. The lack of training data and the presence of rare diseases pose significant challenges for these methods. Prompt tuning combined with pre-trained Vision-Language models (VLMs) is an effective solution to the Few-shot Weakly Supervised WSI Classification (FSWC) task. Nevertheless, applying prompt tuning methods designed for natural images to WSIs presents three significant challenges: 1) These methods fail to fully leverage the prior knowledge from the VLM's text modality; 2) They overlook the essential multi-scale and contextual information in WSIs, leading to suboptimal results; and 3) They lack exploration of instance aggregation methods. To address these problems, we propose a Multi-Scale and Context-focused Prompt Tuning (MSCPT) method for FSWC task. Specifically, MSCPT employs the frozen large language model to generate pathological visual language prior knowledge at multiple scales, guiding hierarchical prompt tuning. Additionally, we design a graph prompt tuning module to learn essential contextual information within WSI, and finally, a non-parametric cross-guided instance aggregation module has been introduced to derive the WSI-level features. Extensive experiments, visualizations, and interpretability analyses were conducted on five datasets and three downstream tasks using three VLMs, demonstrating the strong performance of our MSCPT. All codes have been made publicly accessible at https://github.com/Hanminghao/MSCPT. Minghao Han, Linhao Qu, Dingkang Yang, Xukun Zhang, Lihua Zhang 0002 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Multi-Scale Heterogeneity-Aware Hypergraph Representation for Histopathology Whole Slide ImagesabstractSurvival prediction is a complex ordinal regression task that aims to predict the survival coefficient ranking among a cohort of patients, typically achieved by analyzing patients’ whole slide images. Existing deep learning approaches mainly adopt multiple instance learning or graph neural networks under weak supervision. Most of them are unable to uncover the diverse interactions between different types of biological entities(e.g., cell cluster and tissue block) across multiple scales, while such interactions are crucial for patient survival prediction. In light of this, we propose a novel multi-scale heterogeneity-aware hypergraph representation framework. Specifically, our framework first constructs a multi-scale heterogeneity-aware hypergraph and assigns each node with its biological entity type. It then mines diverse interactions between nodes on the graph structure to obtain a global representation. Experimental results demonstrate that our method outperforms state-of-the-art approaches on three benchmark datasets. Code is publicly available at https://github.com/Hanminghao/H2GT. Minghao Han, Xukun Zhang, Dingkang Yang, Tao Liu 0050, Haopeng Kuang, Jinghui Feng, Lihua Zhang 0002 |
ICME | 2 |
| 2024 | Bio-Inspired Feature Selection via an Improved Binary Golden Jackal Optimization Algorithm
Jinghui Feng, Xukun Zhang, Lihua Zhang 0002 |
KSEM (2) | 2 |
| 2024 | MaskBEV: Towards A Unified Framework for BEV Detection and Map SegmentationabstractAccurate and robust multimodal multi-task perception is crucial for modern autonomous driving systems. However, current multimodal perception research follows independent paradigms designed for specific perception tasks, leading to a lack of complementary learning among tasks and decreased performance in multi-task learning (MTL) due to joint training. In this paper, we propose MaskBEV, a masked attention-based MTL paradigm that unifies 3D object detection and bird's eye view (BEV) map segmentation. MaskBEV introduces a task-agnostic Transformer decoder to process these diverse tasks, enabling MTL to be completed in a unified decoder without requiring additional design of specific task heads. To fully exploit the complementary information between BEV map segmentation and 3D object detection tasks in BEV space, we propose spatial modulation and scene-level context aggregation strategies. These strategies consider the inherent dependencies between BEV segmentation and 3D detection, naturally boosting MTL performance. Extensive experiments on nuScenes dataset show that compared with previous state-of-the-art MTL methods, MaskBEV achieves 1.3 NDS improvement in 3D object detection and 2.7 mIoU improvement in BEV map segmentation, while also demonstrating slightly leading inference speed. Xukun Zhang, Dingkang Yang, Mingcheng Li, Shunli Wang 0001, Lihua Zhang 0002 |
ACM Multimedia | 2 |
| 2024 | 3DLaneFormer: End-to-End 3D Lane Detection with Voxel Descriptors
Qiangbin Xie, Xukun Zhang, Shunli Wang 0001, Lihua Zhang 0002 |
PRCV (4) | 3 |
| 2024 | Dual knowledge-guided two-stage model for precise small organ segmentation in abdominal CT imagesabstractAbstract Multi‐organ segmentation from abdominal CT scans is crucial for various medical examinations and diagnoses. Despite the remarkable achievements of existing deep‐learning‐based methods, accurately segmenting small organs remains challenging due to their small size and low contrast. This article introduces a novel knowledge‐guided cascaded framework that utilizes two types of knowledge—image intrinsic (anatomy) and clinical expertise (radiology)—to improve the segmentation accuracy of small abdominal organs. Specifically, based on the anatomical similarities in abdominal CT scans, the approach employs entropy‐based registration techniques to map high‐quality segmentation results onto inaccurate results from the first stage, thereby guiding precise localization of small organs. Additionally, inspired by the practice of annotating images from multiple perspectives by radiologists, novel Multi‐View Fusion Convolution (MVFC) operator is developed, which can extract and adaptively fuse features from various directions of CT images to refine segmentation of small organs effectively. Simultaneously, the MVFC operator offers a seamless alternative to conventional convolutions within diverse model architectures. Extensive experiments on the Abdominal Multi‐Organ Segmentation (AMOS) dataset demonstrate the superiority of the method, setting a new benchmark in the segmentation of small organs. Tao Liu 0050, Xukun Zhang, Zhongwei Yang, Minghao Han, Haopeng Kuang, Shuwei Ma, Lihua Zhang 0002 |
IET Image Process. | 2 |
| 2023 | D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object DetectionabstractAlthough CNN-based and Transformer-based detectors have made impressive improvements in 3D object detection, these two network paradigms suffer from the interference of insufficient receptive field and local detail weakening, which significantly limits the feature extraction performance of the backbone. In this paper, we propose to fuse convolution and transformer, and simultaneously considering the different contributions of non-empty voxels at different positions in 3D space to object detection, it is not consistent with applying standard convolution and transformer directly on voxels. Specifically, we design a novel deformable sparse transformer to perform long-range information interaction on fine-grained local detail semantics aggregated by focal sparse convolution, termed D-Conformer. D-Conformer learns valuable voxels with position-wise in sparse space and can be applied to most voxel-based detectors as a backbone. Extensive experiments demonstrate that our method achieves satisfactory detection results and outperforms state-of-the-art 3D detection methods by a large margin. Liuzhen Su, Xukun Zhang, Dingkang Yang, Shunli Wang 0001, Peng Zhai, Lihua Zhang 0002 |
ICASSP | 3 |
| 2023 | Anatomical-Aware Point-Voxel Network for Couinaud Segmentation in Liver CT
Xukun Zhang, Yang Liu 0007, Sharib Ali, Minghao Han, Tao Liu 0050, Peng Zhai, Zhiming Cui 0001, Peixuan Zhang, Lihua Zhang 0002 |
MICCAI (3) | 1 |
| 2021 | Confidence-Aware Cascaded Network for Fetal Brain Segmentation on MR Images
Xukun Zhang, Zhiming Cui 0001, Changan Chen, Jingjiao Lou, Wenxin Hu, He Zhang 0023, Tao Zhou 0002, Feng Shi 0001, Dinggang Shen |
MICCAI (3) | 1 |