VLDB 2026 Research / reviewers in the wild / expert
Kun Yu 0002
dblp:51/4486-2
· DBLP profile ↗
14ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-9973-3321ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health RecordsabstractLarge Language Models (LLMs) hold significant promise for improving clinical decision support and reducing physician burnout by synthesizing complex, longitudinal cancer Electronic Health Records (EHRs). However, their implementation in this critical field faces three primary challenges: the inability to effectively process the extensive length and fragmented nature of patient records for accurate temporal analysis; a heightened risk of clinical hallucination, as conventional grounding techniques such as Retrieval-Augmented Generation (RAG) do not adequately incorporate process-oriented clinical guidelines; and unreliable evaluation metrics that hinder the validation of AI systems in oncology. To address these issues, we propose CliCARE, a framework for Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records. The framework operates by transforming unstructured, longitudinal EHRs into patient-specific Temporal Knowledge Graphs (TKGs) to capture long-range dependencies, and then grounding the decision support process by aligning these real-world patient trajectories with a normative guideline knowledge graph. This approach provides oncologists with evidence-grounded decision support by generating a high-fidelity clinical summary and an actionable recommendation. We validated our framework using large-scale, longitudinal data from a private Chinese cancer dataset and the public English MIMIC-IV dataset. In these settings, CliCARE significantly outperforms baselines, including leading long-context LLMs and Knowledge Graph-enhanced RAG methods. The clinical validity of our results is supported by a robust evaluation protocol, which demonstrates a high correlation with assessments made by oncologists. Jitao Liang, Wei Li 0117, Longbing Cao, Kun Yu 0002 |
AAAI | 6 |
| 2026 | Synergistic graph-aware multi-objective differential evolution for high-dimensional feature selection
Weidong Xie, Zhengwei Yuan, Kun Yu 0002, Wei Li 0117 |
Expert Syst. Appl. | 4 |
| 2026 | TP-LReID: Lifelong person re-identification using text prompts
Zhaoshuo Liu, Chaolu Feng, Wei Li 0117, Kun Yu 0002, Jun Hu 0020, Jinzhu Yang |
Pattern Recognit. | 5 |
| 2025 | FactVAE: a factorized variational autoencoder for single-cell multi-omics data integration analysisabstractSingle-cell multi-omics technologies have revolutionized the study of cell states and functions by simultaneously profiling multiple molecular layers within individual cells. However, existing methods for integrating these data struggle to preserve critical feature information and fail to exploit known regulatory knowledge, which is essential for understanding cell functions. This limitation hinders their ability to provide comprehensive and accurate insights into cells. Here, we propose FactVAE, an innovative factorized variational autoencoder designed for the robust and accurate understanding of single-cell multi-omics data. FactVAE integrates the factorization principle into the variational autoencoder framework, ensuring the preservation of feature information while leveraging the non-linear capture of sample information by neural networks. Additionally, known regulatory knowledge is incorporated during model training, and a knowledge transfer strategy is employed for cell embedding optimization and data augmentation. Comparative analyses of single-cell multi-omics datasets from different protocols and the spatial multi-omics dataset demonstrate that FactVAE not only outperforms benchmark methods in clustering performance but also generates augmented data that reveals the clearest cell-type-specific motif expression. Moreover, the feature embeddings captured by FactVAE enable the inference of potential and reliable gene regulatory relationships. Overall, FactVAE's superior performance and strong scalability make it a promising new solution for single-cell multi-omics data analysis. Huixia Zhang, Weidong Xie, Kun Yu 0002, Wei Li 0117, Dazhe Zhao |
Briefings Bioinform. | 5 |
| 2025 | Domain diversity based meta learning for continual person re-identification
Zhaoshuo Liu, Chaolu Feng, Kun Yu 0002, Jiangdian Song, Wei Li 0117 |
Pattern Anal. Appl. | 3 |
| 2024 | nsDCC: dual-level contrastive clustering with nonuniform sampling for scRNA-seq data analysisabstractDimensionality reduction and clustering are crucial tasks in single-cell RNA sequencing (scRNA-seq) data analysis, treated independently in the current process, hindering their mutual benefits. The latest methods jointly optimize these tasks through deep clustering. However, contrastive learning, with powerful representation capability, can bridge the gap that common deep clustering methods face, which requires pre-defined cluster centers. Therefore, a dual-level contrastive clustering method with nonuniform sampling (nsDCC) is proposed for scRNA-seq data analysis. Dual-level contrastive clustering, which combines instance-level contrast and cluster-level contrast, jointly optimizes dimensionality reduction and clustering. Multi-positive contrastive learning and unit matrix constraint are introduced in instance- and cluster-level contrast, respectively. Furthermore, the attention mechanism is introduced to capture inter-cellular information, which is beneficial for clustering. The nsDCC focuses on important samples at category boundaries and in minority categories by the proposed nearest boundary sparsest density weight assignment algorithm, making it capable of capturing comprehensive characteristics against imbalanced datasets. Experimental results show that nsDCC outperforms the six other state-of-the-art methods on both real and simulated scRNA-seq data, validating its performance on dimensionality reduction and clustering of scRNA-seq data, especially for imbalanced data. Simulation experiments demonstrate that nsDCC is insensitive to "dropout events" in scRNA-seq. Finally, cluster differential expressed gene analysis confirms the meaningfulness of results from nsDCC. In summary, nsDCC is a new way of analyzing and understanding scRNA-seq data. Wei Li 0117, Fanghui Zhou, Kun Yu 0002, Chaolu Feng, Dazhe Zhao |
Briefings Bioinform. | 4 |
| 2024 | GCReID: Generalized continual person re-identification via meta learning and knowledge accumulation
Zhaoshuo Liu, Chaolu Feng, Kun Yu 0002, Jun Hu 0020, Jinzhu Yang |
Neural Networks | 3 |
| 2023 | A two-stage hybrid biomarker selection method based on ensemble filter and binary differential evolution incorporating binary African vultures optimizationabstractBACKGROUND: In the field of genomics and personalized medicine, it is a key issue to find biomarkers directly related to the diagnosis of specific diseases from high-throughput gene microarray data. Feature selection technology can discover biomarkers with disease classification information. RESULTS: We use support vector machines as classifiers and use the five-fold cross-validation average classification accuracy, recall, precision and F1 score as evaluation metrics to evaluate the identified biomarkers. Experimental results show classification accuracy above 0.93, recall above 0.92, precision above 0.91, and F1 score above 0.94 on eight microarray datasets. METHOD: This paper proposes a two-stage hybrid biomarker selection method based on ensemble filter and binary differential evolution incorporating binary African vultures optimization (EF-BDBA), which can effectively reduce the dimension of microarray data and obtain optimal biomarkers. In the first stage, we propose an ensemble filter feature selection method. The method combines an improved fast correlation-based filter algorithm with Fisher score. obviously redundant and irrelevant features can be filtered out to initially reduce the dimensionality of the microarray data. In the second stage, the optimal feature subset is selected using an improved binary differential evolution incorporating an improved binary African vultures optimization algorithm. The African vultures optimization algorithm has excellent global optimization ability. It has not been systematically applied to feature selection problems, especially for gene microarray data. We combine it with a differential evolution algorithm to improve population diversity. CONCLUSION: Compared with traditional feature selection methods and advanced hybrid methods, the proposed method achieves higher classification accuracy and identifies excellent biomarkers while retaining fewer features. The experimental results demonstrate the effectiveness and advancement of our proposed algorithmic model. Wei Li 0117, Yuhuan Chi, Kun Yu 0002, Weidong Xie |
BMC Bioinform. | 3 |
| 2023 | An abnormal surgical record recognition model with keywords combination patterns based on TextRank for medical insurance fraud detection
Wei Li 0117, Panpan Ye, Kun Yu 0002, Xin Min, Weidong Xie |
Multim. Tools Appl. | 3 |
| 2022 | A Hybrid Feature Selection Method Based on Binary Differential Evolution and Feature Subset Correlation for Microarray Data*abstractObtaining essential genes from microarray data that can diagnose diseases can be very useful for researchers to understand diseases and develop drugs. However, the high computational cost due to the “curse of dimensionality” and the high redundancy among features limit the application of evolutionary algorithms to the feature selection problem for high-dimensional data. This paper proposes a two-stage hybrid feature selection method to address this problem. In the first stage, a simple and efficient filtering me thod is used to initially filter redundant features, reduce the feature dimensionality, and reduce the search space of the evolutionary algorithm in the second stage. In the second stage, we propose an improved differential evolution algorithm. We redesign the binary quantization of the differential evolution algorithm for the characteristics of microarray data and improve the algorithm’s variation process to balance the algorithm’s search efficiency and accuracy. In addition, we define the redundancy of feature subsets and add it to the fitness function to reduce the redundancy of the final feature subsets. The proposed method is compared with classical feature selection methods and advanced hybrid feature selection methods on eight publicly available microarray data, and the effectiveness and advancement of the proposed method are demonstrated. Weidong Xie, Wei Li 0117, Yushan Fang, Yuhuan Chi, Kun Yu 0002 |
BIBM | 5 |
| 2021 | MMBDE: A Two-stage Hybrid Feature Selection Method From Microarray DataabstractThe discovery of diagnostically significant genes from microarray data is essential for disease diagnosis and drug research. However, the difficulty of analyzing microarray data comes from its high dimensionality and small sample size. Feature selection can effectively remove irrelevant and redundant features, reduce data dimensionality, and improve the accuracy of classifiers. This paper proposes a two-stage hybrid feature selection method MMBDE based on the improved min-Redundancy and Max-Relevance (mRMR) and the improved Binary Differential Evolution (BDE) algorithm. The improved mRMR is used to reduce the feature dimensionality at a coarse-scale significantly. In contrast, the improved BDE is used to refine the feature dimensionality at fine-scale further and select the best features. The experimental results show that MMBDE successfully reduces the dimensionality of microarray gene expression data, obtains high classification accuracy, and extracts effective features closely related to diseases from microarray gene expression data. The relevant datasets and codes can be obtained from https://github.com/xwdshiwo/MMBDE. Weidong Xie, Yuhuan Chi, Kun Yu 0002, Wei Li 0117 |
BIBM | 4 |
| 2021 | ILRC: a hybrid biomarker discovery algorithm based on improved L1 regularization and clustering in microarray dataabstractBACKGROUND: Finding significant genes or proteins from gene chip data for disease diagnosis and drug development is an important task. However, the challenge comes from the curse of the data dimension. It is of great significance to use machine learning methods to find important features from the data and build an accurate classification model. RESULTS: The proposed method has proved superior to the published advanced hybrid feature selection method and traditional feature selection method on different public microarray data sets. In addition, the biomarkers selected using our method show a match to those provided by the cooperative hospital in a set of clinical cleft lip and palate data. METHOD: In this paper, a feature selection algorithm ILRC based on clustering and improved L1 regularization is proposed. The features are firstly clustered, and the redundant features in the sub-clusters are deleted. Then all the remaining features are iteratively evaluated using ILR. The final result is given according to the cumulative weight reordering. CONCLUSION: The proposed method can effectively remove redundant features. The algorithm's output has high stability and classification accuracy, which can potentially select potential biomarkers. Kun Yu 0002, Weidong Xie, Wei Li 0117 |
BMC Bioinform. | 1 |
| 2020 | SP-MIOV: A novel framework of shadow proxy based medical image online visualization in computing and storage resource restrained environments
Wei Li 0117, Kun Yu 0002, Chaolu Feng, Dazhe Zhao |
Future Gener. Comput. Syst. | 2 |
| 2020 | BCEFCM_S: Bias correction embedded fuzzy c-means with spatial constraint to segment multiple spectral images with intensity inhomogeneities and noises
Chaolu Feng, Wei Li 0117, Jun Hu 0020, Kun Yu 0002, Dazhe Zhao |
Signal Process. | 4 |