VLDB 2026 Research / reviewers in the wild / expert
Wei Li 0117
dblp:64/6025-117
· DBLP profile ↗
26ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0001-5422-2678ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health RecordsabstractLarge Language Models (LLMs) hold significant promise for improving clinical decision support and reducing physician burnout by synthesizing complex, longitudinal cancer Electronic Health Records (EHRs). However, their implementation in this critical field faces three primary challenges: the inability to effectively process the extensive length and fragmented nature of patient records for accurate temporal analysis; a heightened risk of clinical hallucination, as conventional grounding techniques such as Retrieval-Augmented Generation (RAG) do not adequately incorporate process-oriented clinical guidelines; and unreliable evaluation metrics that hinder the validation of AI systems in oncology. To address these issues, we propose CliCARE, a framework for Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records. The framework operates by transforming unstructured, longitudinal EHRs into patient-specific Temporal Knowledge Graphs (TKGs) to capture long-range dependencies, and then grounding the decision support process by aligning these real-world patient trajectories with a normative guideline knowledge graph. This approach provides oncologists with evidence-grounded decision support by generating a high-fidelity clinical summary and an actionable recommendation. We validated our framework using large-scale, longitudinal data from a private Chinese cancer dataset and the public English MIMIC-IV dataset. In these settings, CliCARE significantly outperforms baselines, including leading long-context LLMs and Knowledge Graph-enhanced RAG methods. The clinical validity of our results is supported by a robust evaluation protocol, which demonstrates a high correlation with assessments made by oncologists. Jitao Liang, Wei Li 0117, Longbing Cao, Kun Yu 0002 |
AAAI | 3 |
| 2026 | Synergistic graph-aware multi-objective differential evolution for high-dimensional feature selection
Weidong Xie, Zhengwei Yuan, Kun Yu 0002, Wei Li 0117 |
Expert Syst. Appl. | 5 |
| 2026 | TP-LReID: Lifelong person re-identification using text prompts
Zhaoshuo Liu, Chaolu Feng, Wei Li 0117, Kun Yu 0002, Jun Hu 0020, Jinzhu Yang |
Pattern Recognit. | 4 |
| 2026 | Distance self-adaptive fuzzy c-means and its application to image segmentation
Shuaizheng Chen, Chaolu Feng, Dongxiu Li, Zijian Bian, Wei Li 0117, Dazhe Zhao |
Signal Process. Image Commun. | 5 |
| 2025 | Multi-Context Modeling with Spatial Adaptive Enhancement for Domain IdentificationabstractSpatial transcriptomics (ST) provides groundbreaking opportunities to study biological processes and disease mechanisms by measuring gene expression profiles and spatial coordinates. Accurate spatial domain identification via effective data representation is crucial for biological discovery. We propose a multi-context modeling framework with spatial adaptive enhancement (SAE) for ST data analysis. The SAE algorithm leverages spatial information to denoise and enhance expression data, eliminating the adverse impacts of high noise and sparsity inherent in ST data. Given the enhanced data, we incorporate diverse sample context relationships, including spatial, expression, and random non-neighbor contexts, via masked attention. This multi-context strategy overcomes the reliance of current methods on local neighbors, further enhancing the expressiveness of sample embeddings. A graph Laplacian loss preserves sample adjacency in latent space, ensuring spatial coherence of the identified domains. Experiments on five public datasets demonstrate that our method outperforms seven state-of-the-art benchmarks in domain identification and robustness. As a plug-and-play algorithm, SAE significantly improves benchmarks' performance, highlighting its broad application potential. Weidong Xie, Huixia Zhang, Dazhe Zhao, Wei Li 0117 |
BIBM | 5 |
| 2025 | Consistency-Constrained Contrastive Learning with Hard-Negative for Spatial Domain IdentificationabstractSpatial transcriptomics can simultaneously measure gene expression levels and locate their location information, thus assisting in the study of tissue structure. Spatial domain identification plays an important role in the study of organizational structure information. Recent research shows that contrastive learning has been applied to spatial domain identification, achieving promising results. However, existing contrastive learning methods typically generate negative samples by applying global random disturbances to the samples, which introduces significant noise and limits the model's learning ability. Therefore, we propose a novel hierarchical hard-negative sample generation strategy, which customizes negative samples based on the original data for contrastive learning. Meanwhile, we use consistency constraints to guide contrastive learning, which effectively solves the potential local optima and enables the model to learn a more meaningful low-dimensional potential representation. The experimental results show that the model has good performance in terms of spatial domain identification accuracy on benchmark datasets and diversified organizational data across multiple technology platforms. This model shows unique advantages in analyzing the fine spatial domain in complex organizational structures, and provides a powerful tool for in-depth study of tissue heterogeneity. Fanghui Zhou, Huixia Zhang, Wei Li 0117 |
BIBM | 4 |
| 2025 | OFIA: An Object-centric Fine-grained Alignment Enhancement for Video-Text RetrievalabstractText-video alignment is crucial for text-video retrieval. Empirical studies have suggested that coarse-grained alignment overlooks the rich cross-modal details and therefore performs worse than fine-grained alignment. In the current transfer learning paradigm, although researchers primarily utilize patch-level or frame-level embedding as the fine-grained video representation for alignment, the frame embeddings miss the informative visual information, while a single patch only captures limited local details. These defects hinder the potential improvement in video-text retrieval. To address the defects of the existing fine-grained alignment approach, this paper proposes the Object-centric Fine-grained Alignment Enhancement for Video-Text Retrieval, namely OFIA, which consists of a text-guided object-text alignment module and a similarity-wise frame aggregation module to enhance video-text alignment. The text-guided object-text alignment module leverages the textual descriptions to detect and extract more relevant objects, enabling more precise local similarities within video frames. However, not all frames contribute equally to effective alignments. The similarity-wise frame aggregation module assigns greater importance to informative frames, overcoming the challenge of insignificant ones and optimizing the similarity of matched video-text pairs. The empirical evaluation on benchmark datasets including MSRVTT, MSVD, DiDeMo, and ActivityNet demonstrate state-of-the-art performance of the proposed method. Zhengqi Huang, Wei Li 0117, Chuang Dong |
CIKM | 2 |
| 2025 | FactVAE: a factorized variational autoencoder for single-cell multi-omics data integration analysisabstractSingle-cell multi-omics technologies have revolutionized the study of cell states and functions by simultaneously profiling multiple molecular layers within individual cells. However, existing methods for integrating these data struggle to preserve critical feature information and fail to exploit known regulatory knowledge, which is essential for understanding cell functions. This limitation hinders their ability to provide comprehensive and accurate insights into cells. Here, we propose FactVAE, an innovative factorized variational autoencoder designed for the robust and accurate understanding of single-cell multi-omics data. FactVAE integrates the factorization principle into the variational autoencoder framework, ensuring the preservation of feature information while leveraging the non-linear capture of sample information by neural networks. Additionally, known regulatory knowledge is incorporated during model training, and a knowledge transfer strategy is employed for cell embedding optimization and data augmentation. Comparative analyses of single-cell multi-omics datasets from different protocols and the spatial multi-omics dataset demonstrate that FactVAE not only outperforms benchmark methods in clustering performance but also generates augmented data that reveals the clearest cell-type-specific motif expression. Moreover, the feature embeddings captured by FactVAE enable the inference of potential and reliable gene regulatory relationships. Overall, FactVAE's superior performance and strong scalability make it a promising new solution for single-cell multi-omics data analysis. Huixia Zhang, Weidong Xie, Kun Yu 0002, Wei Li 0117, Dazhe Zhao |
Briefings Bioinform. | 6 |
| 2025 | Domain diversity based meta learning for continual person re-identification
Zhaoshuo Liu, Chaolu Feng, Kun Yu 0002, Jiangdian Song, Wei Li 0117 |
Pattern Anal. Appl. | 5 |
| 2024 | BECNN: Bias Field Estimation CNN Trained with a Dual Route Implicit Supervised Learning StrategyabstractBias fields adversely affect various automatic analysis technologies. Therefore, bias field correction is essential. However, deep learning based methods encounter challenges in obtaining ground truth. Although existing methods attempt to address this problem by using training-free techniques or constructing datasets with approximations of the ground truth, the lack of task-oriented guidance, training instability, and inappropriate use of approximations still impact performance. Additionally, different approaches for bias field correction, i.e., estimating bias fields versus directly restoring clean images, exhibit different performances. However, there is no consensus on the best way, resulting in limited performance in some cases. To address these problems, we propose a bias field generation method to construct the dataset and provide task-oriented information. We then propose the concept of Equivalent of Residual Mapping (ERM) to analyze the advantages of estimating bias fields. According to ERM, we propose the Bias field Estimation Convolutional Neural Network (BECNN). Finally, we propose the Dual Route Implicit Supervised Learning (DRISL) strategy to balance the guidance derived from approximations of the ground truth with over-dependence on them. The proposed method is compared qualitatively and quantitatively with the correlated methods. Experiment results demonstrate that the proposed method performs effectively both on bias field estimation and correction. Shuaizheng Chen, Chaolu Feng, Wei Li 0117, Jinzhu Yang, Dazhe Zhao |
BIBM | 3 |
| 2024 | nsDCC: dual-level contrastive clustering with nonuniform sampling for scRNA-seq data analysisabstractDimensionality reduction and clustering are crucial tasks in single-cell RNA sequencing (scRNA-seq) data analysis, treated independently in the current process, hindering their mutual benefits. The latest methods jointly optimize these tasks through deep clustering. However, contrastive learning, with powerful representation capability, can bridge the gap that common deep clustering methods face, which requires pre-defined cluster centers. Therefore, a dual-level contrastive clustering method with nonuniform sampling (nsDCC) is proposed for scRNA-seq data analysis. Dual-level contrastive clustering, which combines instance-level contrast and cluster-level contrast, jointly optimizes dimensionality reduction and clustering. Multi-positive contrastive learning and unit matrix constraint are introduced in instance- and cluster-level contrast, respectively. Furthermore, the attention mechanism is introduced to capture inter-cellular information, which is beneficial for clustering. The nsDCC focuses on important samples at category boundaries and in minority categories by the proposed nearest boundary sparsest density weight assignment algorithm, making it capable of capturing comprehensive characteristics against imbalanced datasets. Experimental results show that nsDCC outperforms the six other state-of-the-art methods on both real and simulated scRNA-seq data, validating its performance on dimensionality reduction and clustering of scRNA-seq data, especially for imbalanced data. Simulation experiments demonstrate that nsDCC is insensitive to "dropout events" in scRNA-seq. Finally, cluster differential expressed gene analysis confirms the meaningfulness of results from nsDCC. In summary, nsDCC is a new way of analyzing and understanding scRNA-seq data. Wei Li 0117, Fanghui Zhou, Kun Yu 0002, Chaolu Feng, Dazhe Zhao |
Briefings Bioinform. | 2 |
| 2024 | Graph neural collaborative filtering with medical content-aware pre-training for treatment pattern recommendation
Xin Min, Wei Li 0117, Ruiqi Han, Tianlong Ji, Weidong Xie |
Pattern Recognit. Lett. | 2 |
| 2023 | A two-stage hybrid biomarker selection method based on ensemble filter and binary differential evolution incorporating binary African vultures optimizationabstractBACKGROUND: In the field of genomics and personalized medicine, it is a key issue to find biomarkers directly related to the diagnosis of specific diseases from high-throughput gene microarray data. Feature selection technology can discover biomarkers with disease classification information. RESULTS: We use support vector machines as classifiers and use the five-fold cross-validation average classification accuracy, recall, precision and F1 score as evaluation metrics to evaluate the identified biomarkers. Experimental results show classification accuracy above 0.93, recall above 0.92, precision above 0.91, and F1 score above 0.94 on eight microarray datasets. METHOD: This paper proposes a two-stage hybrid biomarker selection method based on ensemble filter and binary differential evolution incorporating binary African vultures optimization (EF-BDBA), which can effectively reduce the dimension of microarray data and obtain optimal biomarkers. In the first stage, we propose an ensemble filter feature selection method. The method combines an improved fast correlation-based filter algorithm with Fisher score. obviously redundant and irrelevant features can be filtered out to initially reduce the dimensionality of the microarray data. In the second stage, the optimal feature subset is selected using an improved binary differential evolution incorporating an improved binary African vultures optimization algorithm. The African vultures optimization algorithm has excellent global optimization ability. It has not been systematically applied to feature selection problems, especially for gene microarray data. We combine it with a differential evolution algorithm to improve population diversity. CONCLUSION: Compared with traditional feature selection methods and advanced hybrid methods, the proposed method achieves higher classification accuracy and identifies excellent biomarkers while retaining fewer features. The experimental results demonstrate the effectiveness and advancement of our proposed algorithmic model. Wei Li 0117, Yuhuan Chi, Kun Yu 0002, Weidong Xie |
BMC Bioinform. | 1 |
| 2023 | Multi-channel hypergraph topic neural network for clinical treatment pattern mining
Xin Min, Wei Li 0117, Panpan Ye, Tianlong Ji, Weidong Xie |
Inf. Process. Manag. | 2 |
| 2023 | An abnormal surgical record recognition model with keywords combination patterns based on TextRank for medical insurance fraud detection
Wei Li 0117, Panpan Ye, Kun Yu 0002, Xin Min, Weidong Xie |
Multim. Tools Appl. | 1 |
| 2022 | A Hybrid Feature Selection Method Based on Binary Differential Evolution and Feature Subset Correlation for Microarray Data*abstractObtaining essential genes from microarray data that can diagnose diseases can be very useful for researchers to understand diseases and develop drugs. However, the high computational cost due to the “curse of dimensionality” and the high redundancy among features limit the application of evolutionary algorithms to the feature selection problem for high-dimensional data. This paper proposes a two-stage hybrid feature selection method to address this problem. In the first stage, a simple and efficient filtering me thod is used to initially filter redundant features, reduce the feature dimensionality, and reduce the search space of the evolutionary algorithm in the second stage. In the second stage, we propose an improved differential evolution algorithm. We redesign the binary quantization of the differential evolution algorithm for the characteristics of microarray data and improve the algorithm’s variation process to balance the algorithm’s search efficiency and accuracy. In addition, we define the redundancy of feature subsets and add it to the fitness function to reduce the redundancy of the final feature subsets. The proposed method is compared with classical feature selection methods and advanced hybrid feature selection methods on eight publicly available microarray data, and the effectiveness and advancement of the proposed method are demonstrated. Weidong Xie, Wei Li 0117, Yushan Fang, Yuhuan Chi, Kun Yu 0002 |
BIBM | 2 |
| 2022 | Feature Selection for Microarray Data via Community Detection Fusing Multiple Gene Relation Networks InformationabstractIn recent decades, the rapid development of gene sequencing and computer technology has increased the growth of high-dimensional microarray data. Some machine learning methods have been successfully applied to it to help classify cancer. In most cases, high dimensionality and the small sample size of microarray data restricted the performance of cancer classification. This problem usually issolved bysome feature selection methods. However, most of them neglect the exploitation of relations among genes. This paper proposes a novel feature selection method by fusing multiple gene relation network information based on community detection (MGRCD). The proposed method divides all genes into different communities. Then, the genes most associated with cancer classification are selected from each community. The proposed method satisfies both maximum relevances gene with cancer and minimum redundancy among genes for the selected optimal feature subset. The experiment results show that the proposed gene selection method can effectively improve classification performance. Shoujia Zhang, Wei Li 0117, Weidong Xie |
BIBM | 2 |
| 2022 | TCAM-Resnet: A convolutional neural network for screening DR and AMD based on OCT imagesabstractDiabetic retinopathy (DR) and age-related macular degeneration (AMD) are important causes of blindness and visual loss. Optical coherence tomography (OCT) is a non-invasive optical imaging method that can capture retinal vascular information and even pathological information. In order to improve the screening rate and accuracy of these two diseases, we propose a new network structure named TCAM-Resnet, which uses OCT three-dimensional images to screen and classify AMD and DR. TCAM-Resnet is based on the Resnet network. A three-dimensional convolution attention module (TCAM) is added. The attention module can extract the weight features of blood vessels from 3D images and uses the residual-like structure when interacting with the Resnet network, which makes the original data retain more information during attention. Experimental results on the OCTA-500 dataset show that a three-dimensional convolution network is superior to a two-dimensional convolution network in lesion feature extraction. With the addition of the new module TCAM, Resnet3D has achieved higher accuracy in disease classification t asks, with the accuracy of AMD, DR, and NORMAL reaching 83.3%, and the accuracy of AMD and NORMAL reaching 98%. Deshui Yu, Yuhui Cai, Bojun Li, Wei Li 0117 |
BIBM | 6 |
| 2022 | A novel biomarker selection method combining graph neural network and gene relationships applied to microarray dataabstractBACKGROUND: The discovery of critical biomarkers is significant for clinical diagnosis, drug research and development. Researchers usually obtain biomarkers from microarray data, which comes from the dimensional curse. Feature selection in machine learning is usually used to solve this problem. However, most methods do not fully consider feature dependence, especially the real pathway relationship of genes. RESULTS: Experimental results show that the proposed method is superior to classical algorithms and advanced methods in feature number and accuracy, and the selected features have more significance. METHOD: This paper proposes a feature selection method based on a graph neural network. The proposed method uses the actual dependencies between features and the Pearson correlation coefficient to construct graph-structured data. The information dissemination and aggregation operations based on graph neural network are applied to fuse node information on graph structured data. The redundant features are clustered by the spectral clustering method. Then, the feature ranking aggregation model using eight feature evaluation methods acts on each clustering sub-cluster for different feature selection. CONCLUSION: The proposed method can effectively remove redundant features. The algorithm's output has high stability and classification accuracy, which can potentially select potential biomarkers. Weidong Xie, Wei Li 0117, Shoujia Zhang, Jinzhu Yang, Dazhe Zhao |
BMC Bioinform. | 2 |
| 2022 | Dual-level diagnostic feature learning with recurrent neural networks for treatment sequence recommendation
Xin Min, Wei Li 0117, Jinzhao Yang, Weidong Xie, Dazhe Zhao |
J. Biomed. Informatics | 2 |
| 2021 | MMBDE: A Two-stage Hybrid Feature Selection Method From Microarray DataabstractThe discovery of diagnostically significant genes from microarray data is essential for disease diagnosis and drug research. However, the difficulty of analyzing microarray data comes from its high dimensionality and small sample size. Feature selection can effectively remove irrelevant and redundant features, reduce data dimensionality, and improve the accuracy of classifiers. This paper proposes a two-stage hybrid feature selection method MMBDE based on the improved min-Redundancy and Max-Relevance (mRMR) and the improved Binary Differential Evolution (BDE) algorithm. The improved mRMR is used to reduce the feature dimensionality at a coarse-scale significantly. In contrast, the improved BDE is used to refine the feature dimensionality at fine-scale further and select the best features. The experimental results show that MMBDE successfully reduces the dimensionality of microarray gene expression data, obtains high classification accuracy, and extracts effective features closely related to diseases from microarray gene expression data. The relevant datasets and codes can be obtained from https://github.com/xwdshiwo/MMBDE. Weidong Xie, Yuhuan Chi, Kun Yu 0002, Wei Li 0117 |
BIBM | 5 |
| 2021 | ILRC: a hybrid biomarker discovery algorithm based on improved L1 regularization and clustering in microarray dataabstractBACKGROUND: Finding significant genes or proteins from gene chip data for disease diagnosis and drug development is an important task. However, the challenge comes from the curse of the data dimension. It is of great significance to use machine learning methods to find important features from the data and build an accurate classification model. RESULTS: The proposed method has proved superior to the published advanced hybrid feature selection method and traditional feature selection method on different public microarray data sets. In addition, the biomarkers selected using our method show a match to those provided by the cooperative hospital in a set of clinical cleft lip and palate data. METHOD: In this paper, a feature selection algorithm ILRC based on clustering and improved L1 regularization is proposed. The features are firstly clustered, and the redundant features in the sub-clusters are deleted. Then all the remaining features are iteratively evaluated using ILR. The final result is given according to the cumulative weight reordering. CONCLUSION: The proposed method can effectively remove redundant features. The algorithm's output has high stability and classification accuracy, which can potentially select potential biomarkers. Kun Yu 0002, Weidong Xie, Wei Li 0117 |
BMC Bioinform. | 4 |
| 2020 | SP-MIOV: A novel framework of shadow proxy based medical image online visualization in computing and storage resource restrained environments
Wei Li 0117, Kun Yu 0002, Chaolu Feng, Dazhe Zhao |
Future Gener. Comput. Syst. | 1 |
| 2020 | BCEFCM_S: Bias correction embedded fuzzy c-means with spatial constraint to segment multiple spectral images with intensity inhomogeneities and noises
Chaolu Feng, Wei Li 0117, Jun Hu 0020, Kun Yu 0002, Dazhe Zhao |
Signal Process. | 2 |
| 2017 | A multi-kernel based framework for heterogeneous feature selection and over-sampling for computer-aided detection of pulmonary nodules
Peng Cao 0001, Xiaoli Liu 0001, Jinzhu Yang, Dazhe Zhao, Wei Li 0117, Min Huang 0001, Osmar R. Zaïane |
Pattern Recognit. | 5 |
| 2013 | CRNN: Integrating classification rules into neural networkabstractAssociation classification has been an important type of the rule-based classification. A variety of approaches have been proposed to build a classifier based on classification rules. In the prediction stage of the extant approaches, most of the existing association classifiers use the ensemble quality measurement of each rule in a subset of rules to predict the class label of the new data. This method still suffers the following two problems. The classification rules are used individually thus the coupling relations between rules [1] are ignored in the prediction. However, in real-world rule set, rules are often inter-related and a new data object may partially satisfy many rules. Furthermore, the classification rule based prediction model lacks a general expression of the decision methodology. This paper proposes a classification method that integrating classification rules into neural network (CRNN, for short), which presents a general form of the rule based decision methodology by rule-based network. In comparison with the extant rule-based classifiers, such as C4.5, CBA, CMAR and CPAR, our approach has two advantages. First, CRNN takes the coupling relations between rules from the training data into account in the prediction step. Second, CRNN automatically obtains higher performance on the structure and parameter learning than traditional neural network. CRNN uses the linear computing algorithm in neural network instead of the costly iterative learning algorithm. Two ways of the classification rule set generation are conducted in this paper for the CRNN evaluation, and CRNN achieves the satisfactory performance. Wei Li 0117, Longbing Cao, Dazhe Zhao, Xia Cui 0002, Jinzhu Yang |
IJCNN | 1 |