VLDB 2026 Research / reviewers in the wild / expert
Xiaorou Zheng
dblp:263/8334
· DBLP profile ↗
17ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0002-4941-5907ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning What to Ignore: Mitigating Negative Transfer in Medical Knowledge Fusion via Clinical Task-Adaptive SelectionabstractIntegrating external medical knowledge into longitudinal electronic health record modeling is a prevailing paradigm to mitigate clinical data sparsity.However, existing approaches face a reliability-timeliness dilemma, struggling to balance the structural authority of static ontologies with the reasoning flexibility of large language models.Furthermore, most frameworks overlook the risk of relative negative transfer, where indiscriminately fusing task-irrelevant knowledge can introduce noise or even cause conflicts that weakens patient-specific signals.In this paper, we propose TrustKE, a Trustworthy Knowledge Enhancement framework.First, we construct a dual-layer knowledge graph that anchors dynamic, evidence-based chain-of-thought reasoning from medical literature within the stable structure of medical knowledge graph.Second, we introduce a task-adaptive knowledge selection mechanism that dynamically optimizes the graph, retaining only task-specific signals.Extensive experiments on MIMIC-III and MIMIC-IV across four clinical tasks show that TrustKE outperforms state-of-the-art baselines.Our analysis confirms that TrustKE effectively mitigates negative transfer while offering transparent reasoning for clinical decision-making. Shoubin Dong, Xiaorou Zheng |
ACL (1) | 3 |
| 2026 | Implicit-Explicit Segmentation Synergy: A Dual-Guided Fusion Network for Joint Lesion Localization and Disease Classification
Xiaorou Zheng, Shoubin Dong |
ICPR (4) | 3 |
| 2026 | Dynamic patient similarity modeling with multi-source fused clinical knowledge for enhanced disease prediction
Xiaorou Zheng, Shoubin Dong |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | GRFusion: Graph reconstruction-aware fusion of incomplete multi-modal data for cancer diagnosis and prognosis
Ziye Zhang 0004, Yuying Huang, Xiaorou Zheng, Shoubin Dong |
Inf. Sci. | 4 |
| 2026 | Causality-Guided Diffusion and Fusion of incomplete multi-modal data for robust survival prognosis
Yuying Huang, Xiaorou Zheng, Shoubin Dong |
Medical Image Anal. | 2 |
| 2026 | A knowledge enhanced framework for interpretable medical visual question and answering via large foundation model
Yinxin Xu, Xiaorou Zheng, Shoubin Dong |
Multim. Syst. | 3 |
| 2026 | Unsupervised SAM-guided mixture-of-multimodal-experts fusion network for medical image diagnosis
Jing Li 0174, Xiaorou Zheng, Shoubin Dong |
Neural Networks | 3 |
| 2026 | Pretraining-Based Relevance-Aware Visit Similarity Network for Drug RecommendationabstractDrug recommendation based on electronic health record (EHR) is fundamental to effective disease treatment. Similar to commercial sequence-based recommendation systems, the accuracy of drug recommendation largely depends on precise patient modeling. However, patient modeling is more complex, as it not only requires sequence modeling of patient's disease course, but also needs to refer to the information of patients with similar medical medication. In EHR data, many patients have only one visit record, and the similarity between patients is often vague and unclear, which may cause noise and ambiguity. This leads to significant challenges for the drug recommendation field, especially when patient records are sparse or when patient similarity is vague. To address the above challenges, we propose RaVSNet (Relevance aware Visit Similarity Network), which improves drug recommendation by leveraging both longitudinal and transversal visit similarity and integrating medical relevance knowledge. RaVSNet utilizes multi-dimensional visit information similar to the patient's current visit as a reference, and employs a relevance-aware network to explicitly model the matching relationships between medical conditions and medications. Additionally, RaVSNet designs a general pretraining framework specifically for drug recommendation, including two tasks, Medication Sequence Reconstruction (MSR) and Causal Effect Inference (CEI), to discover the deep connections between medical information and medications. Experimental results on two public EHR datasets, MIMIC-III and MIMIC-IV demonstrate that the proposed algorithm outperforms state-of-the-art methods, yielding more accurate drug recommendation combinations, and the proposed general pretraining framework can be seamlessly integrated into most drug recommendation methods to achieve performance improvements. Shoubin Dong, Xiaorou Zheng, Jinlong Hu 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Confidence-Aware Adaptive Fusion Leaning of Imbalance Multi-Modal Data for Cancer Diagnosis and PrognosisabstractThe effective fusion of pathological images and molecular omics holds significant potential for precision medicine. However, pathological and molecular data are highly heterogeneous, and large-scale multi-modal cancer data often suffer from incomplete information. Predicting clinical tasks from such imbalanced multi-modal data presents a major challenge. Therefore, we propose a confidence-aware adaptive fusion framework CAFusion. The framework adopts a modular design, providing independent and flexible modal feature learning modules to capture high-quality features. To address issues of modal imbalance caused by heterogeneous and incomplete modal, we design a confidence-aware method that evaluates the features of each modal and automatically adjusts their weights. To effectively fuse pathological and molecular modals, we propose an adaptive deep network, which features a flexible, non-fixed layer structure that effectively extracts hidden joint information from multi-modal features, ensuring high generalizability. Experiment results demonstrate that the performance of the CAFusion framework outperforms other state-of-the-art methods, both on complete and incomplete datasets. Moreover, the CAFusion framework offers reasonable medical interpretability. Ziye Zhang 0004, Yuying Huang, Xiaorou Zheng, Shoubin Dong |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Multi-Vector Biomedical Dense Retrieval with Knowledge-Enhanced Entity-Type ClusteringabstractSingle-vector dense retrieval models, which are foundational to modern Retrieval-Augmented Generation (RAG) systems, struggle to represent the multifaceted semantics of complex documents, particularly in specialized fields like biomedicine. This semantic bottleneck limits their ability to provide comprehensive context for generation tasks. To address this, we propose ELK-Multi, a novel multi-vector retrieval framework that constructs fine-grained document representations through knowledge-enhanced entity-type clustering. By leveraging a knowledge-aware encoder, ELK-Multi first identifies and groups entities by their type, generating a distinct vector for each semantic cluster. These targeted representations are then combined with a global document vector using principled aggregation strategies to balance fine-grained detail with holistic context. Extensive experiments on the TREC-COVID and NFCorpus datasets validate our approach, where ELK-Multi establishes new state-of-the-art results in NDCG and Recall. This is complemented by a detailed efficiency analysis demonstrating that our model achieves this performance while remaining within the efficient dual-encoder paradigm, alongside a qualitative analysis with case studies and visualizations. Jiajie Tan, Xiaorou Zheng, Shoubin Dong |
ACM Trans. Knowl. Discov. Data | 3 |
| 2025 | Adaptive disentangled target representation for unsupervised domain adaptation in remote sensing segmentation
Runuo Lu, Shoubin Dong, Jianxin Jia, Jinsong Chen 0001, Shanxin Guo, Xiaorou Zheng |
Eng. Appl. Artif. Intell. | 8 |
| 2025 | Reconstruction of Thermal Infrared Sea Surface Temperature Using a Diffusion ModelabstractSea surface temperature (SST) is a vital parameter in oceanography and climate science, influencing various fields. While remote sensing provides daily SST data, thermal infrared (TIR)-based sensors offer higher spatial resolution but struggle with cloud penetration, often resulting in data gaps and inaccuracies. This study proposes an SST reconstruction framework based on a diffusion model that integrates spatiotemporal information using Himawari TIR and OSTIA SST data, yielding a fully covered SST dataset with a spatial resolution of 0.02°. The model demonstrates good performance in reconstructing SST in the South China Sea (SCS), achieving a coefficient of determination ($R^{2}$) of 0.92, a bias of 0.06 °C, a root-mean-square error (RMSE) of 0.39 °C, and a peak signal-to-noise ratio (PSNR) of 57.93. The transferability of the model is, furthermore, confirmed through accurate SST predictions in the Indian Ocean after training on SCS data, indicating its applicability across different regions with limited Himawari data and its potential for broader geographical applications. Compared to the original Himawari data, the reconstructed SST reduces the RMSE from 1.01 °C to 0.29 °C, increases the$R^{2}$from 0.71 to 0.86, and adjusts the bias from -0.54 °C to 0.02 °C, thereby enhancing accuracy. By integrating temporal information, the proposed approach captures both spatial and temporal characteristics of SST, effectively representing seasonal variations, small-scale pulsations, rapid coastal changes, and stable offshore fluctuations. This study lays the groundwork for applying the proposed framework to other regions, highlighting its potential for broader applications in generating high-resolution, all-weather SST data. Yawei Wang 0001, Rongjin Guo, Xiaonning Song, Tinghui Zhang, Yueli Chen, Xiaorou Zheng |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2025 | Car Damage Detection Based on Multi-View Fusion and Alignment: Dataset and MethodabstractTraffic accidents remain a significant concern due to their potential severity and impact on society. The rapid and accurate detection of car damage is increasingly crucial. Manual assessment of car damage usually relies on multi-view car images taken at the scene, which can provide richer information for damage assessment. However, most car damage algorithms are based on single-view datasets, and it is hard to fully leverage the complementary information and alignment information between distant view and close-up images. In this paper, we propose the Multi-View Car Damage Detection model (MVA-CDD), comprising three key modules: Feature Split (FS), Feature Fusion (FF), and Image Alignment (IA). The FS module extracts global and detailed information from distant-view and close-up images separately, which are then combined by the FF module. The IA module effectively aligns car damage information in distant-view and close-up images to correct errors and biases. Meanwhile, we created the new Car Damage Detection Multi-view dataset (CDDM), which has a significant advantage in both image quantity and diversity across categories, addressing the shortcomings of existing multi-view datasets. Our proposed MVA-CDD outperforms the state-of-the-art single-view and multi-view models with the dataset. Results from ablation studies further confirm the efficiency of MVA-CDD. This study contributes to optimizing the car damage detection and claims adjudication process, leading to significant labor and material cost savings. CDDM dataset is available athttps://github.com/SCUT-CCNL/CDDM. Jinbo Peng, Shoubin Dong, Xiaorou Zheng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Supervoxels-based Self-supervised Few-shot 3D Medical Image Segmentation via Multiple Features TransferabstractIn recent years, medical image segmentation technology has made great progress. However, the annotated data of 3D medical images is relatively small, and although fewshot segmentation can solve this problem, there are still many challenges. Too few samples of support images may lead to the fact that it is difficult to fully represent 3D medical images, especially the important 3D spatial information, and the global correlations between support and query images are not fully utilized. In this paper, we propose a novel few-shot 3D medical image segmentation pipeline framework, SMFT-Net, which can efficiently accomplish the 3D medical image segmentation task using only one labeled sample. Specifically, we proposed pretrained feature transfer module (PFTM) and bidirectional feature transfer module (BFTM) for multiple feature transfer of 3D medical image. PFTM can be used for 3D feature transfer to ensure that the 3D spatial information of medical images is preserved. And BFTM can perform bi-directional feature transfer between the query image and the support image to eliminate extraneous information from the surrounding pixels. Extensive experiments on four medical image datasets demonstrate that our method outperforms the state-of-the-art methods. Jing Li 0174, Xiaorou Zheng, Shoubin Dong |
BIBM | 3 |
| 2024 | Explicit High-Level Semantic Network for Domain Generalization in Hyperspectral Image ClassificationabstractWhen applied across different scenes, hyperspectral image (HSI) classification models often struggle to generalize due to the data distribution disparities and labels’ scarcity, leading to domain shift (DS) problems. Recently, the high-level semantics from text has demonstrated the potential to address the DS problem, by improving the generalization capability of image encoders through aligning image-text pairs. However, the main challenge still lies in crafting appropriate texts that accurately represent the intricate interrelationships and the fragmented nature of land cover in HSIs and effectively extracting spectral-spatial features from HSI data. This article proposes a domain generalization (DG) method, EHSnet, to address these issues by leveraging multilayered explicit high-level semantic (EHS) information from different types of texts to provide precisely relevant semantic information for the image encoder. A multilayered EHS information paradigm is well-defined, aiming to extract the HSI’s intricate interrelationships and the fragmented land-cover features, and a dual-residual encoder connected by a 2-D convolution is designed, which combines CNNs with residual structure and Vision Transformers (ViTs) with short-range cross-layer connections to explore the spectral-spatial features of HSIs. By aligning text features with image features in the semantic space, EHSnet improves the representation capability of the image encoder and is endowed with zero-shot generalization ability for cross-scene tasks. Extensive experiments conducted on three hyperspectral datasets, including Houston, Pavia, and XS datasets, validate the effectiveness and superiority of EHSnet, with the Kappa coefficient improved by 8.17%, 3.22%, and 3.62% across three datasets compared to the state-of-the-art (SOTA) methods. The code is available athttps://github.com/SCUT-CCNL/EHSnet. Shoubin Dong, Xiaorou Zheng, Runuo Lu, Jianxin Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Removing Stripe Noise Based on Improved Statistics for Hyperspectral ImagesabstractStripe noise still affects full-spectrum airborne hyperspectral imager (FAHI) images after laboratory radiometric calibration, which seriously affects the subsequent applications of the imager. Therefore, two state-of-the-art methods, median linear correction (MLC) and Fourier transform filtering (FTF), were proposed to restore FAHI images, and the residual stripes were removed in most cases. However, these methods have their own limitations. For instance, the restored image has a slight “shadow” in cases where the high-response digital numbers (DNs) of the detector are aligned with the flight direction. This letter proposes a new method based on improved statistics to restore FAHI images. In this method, the hyperspectral image data from the adjacent flight paths is used to obtain the uniform response DNs for nearly identical low and high irradiances. Subsequently, a statistics-based MLC method is used to eliminate the stripe noise. To quantitatively evaluate the restoration results, we compared results with MLC and FTF methods. The change in mean value and mean relative deviation of the proposed method for the high-response DNs area of the image is 0.25% and 0.97%, respectively, better than that of the other two methods. The experimental results demonstrate that the proposed method is effective for removing stripe noise and preserving accurate image information of push-broom hyperspectral imagery. Jianxin Jia, Xiaorou Zheng, Shanxin Guo, Yueming Wang 0002, Jinsong Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Tradeoffs in the Spatial and Spectral Resolution of Airborne Hyperspectral Imaging Systems: A Crop Identification Case StudyabstractAirborne hyperspectral images are used for crop identification with a high classification accuracy because of their high spectral resolution, spatial resolution, and signal-to-noise ratio (SNR). However, the tradeoffs between the three core parameters of a hyperspectral imager (SNR, spatial resolution, and spectral resolution) should be considered for designing an efficient imaging system. Only a few reported studies on the analysis of the impact of SNR on identification accuracy are available. Further, the tradeoffs and mutual interactions among these parameters are rarely considered. In this empirical study, our aim was to understand the relationship among the core parameters and their effects on crop identification accuracy by analyzing the tradeoffs and mutual interactions among these parameters. We analyzed the hyperspectral images of a typical plain agricultural area in Xiongan, China, acquired by the newly developed sensor airborne multimodular imaging spectrometer (AMMIS). The fundamental images were transformed to form datasets with different ranges of spectral resolution, spatial resolution, and SNR using data reconstruction methods. We adopted the classification and regression tree (CART), random forest (RF), and k-nearest neighbor (kNN) classifiers, and observed the overall accuracy (OA) across the degraded hyperspectral datasets. The experimental results indicated that the OA decreased with a decreasing SNR. As the spectral resolution became coarser, the OA first increased, plateaued, and then decreased. However, the OA increased with decreasing spatial resolution. This study was performed with the goal of bridging the knowledge gap between the back-end hyperspectral sensor designing and its front-end applications. Jianxin Jia, Jinsong Chen 0001, Xiaorou Zheng, Yueming Wang 0002, Shanxin Guo, Haibin Sun 0002, Changhui Jiang, Mika Karjalainen, Kirsi Karila, Zhiyong Duan, Tinghuai Wang, Juha Hyyppä, Yuwei Chen 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |