VLDB 2026 Research / reviewers in the wild / expert
Jie Li 0055
dblp:17/2703-55
· DBLP profile ↗
19ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0001-9359-3586ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Pathway-Informed Deep Learning Models in Cancer Research: A SurveyabstractBiological pathways play an important role in complex disease research, drug development, and precision medicine. Recently, tens of pathway-informed deep learning models in cancer research have been proposed due to their interpretability and high performance. However, there is still a lack of reviews that have a specific focus on the application strategy of pathway information in the pathway-informed deep learning models. Hence, we surveyed the pathway-informed deep learning models in cancer research. In this survey, the pathway information used in pathway-informed models was summarized into four categories; pathway-informed deep learning models were divided into three major categories and seven subcategories based on the pathway information utilized and the strategies employed for its application. For each subcategory, the application strategy of pathway information is illustrated, and the advantages and disadvantages of the pathway-informed models are summarized. Besides, the commonly used interpretability methods and pathway databases are provided to assist in the design of more effective models. Finally, the challenges faced in developing better pathway-informed deep learning models are presented. Dechen Xu, Jiahuan Jin, Zhengnan Zhao, Jie Li 0055 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | A Deep Transfer Learning Method Based on Multi-Perspective Feature Alignment for Cancer Drug Response PredictionabstractTumor cell subtypes exhibit significant heterogeneity in drug sensitivity, limiting the efficacy of conventional therapies and leading to recurrence or metastasis. Predicting single-cell drug responses is thus critical for precision medicine, but the scarcity of such data hinders model training. Leveraging abundant cancer cell line data, we propose DTLCDR, a deep transfer learning method that predicts singlecell drug responses by transferring knowledge from cell lines. DTLCDR integrates supervised and unsupervised approaches to identify key gene features, followed by deep learning-based cross-domain alignment. Its innovation lies in multi-perspective feature alignment, including: (1) graph attention for semantic propagation, (2) multi-level domain discriminators for feature alignment, and (3) centroid alignment loss for class-specific consistency. Evaluated on six single-cell datasets, DTLCDR outperformed traditional and existing deep transfer learning methods across nine metrics. Ablation studies and visualizations further validated its robustness. DTLCDR provides a reliable tool for single-cell drug response prediction, advancing personalized cancer therapy. Dechen Xu, Jie Li 0055 |
BIBM | 2 |
| 2025 | M-NET: Transforming Single Nucleotide Variations Into Patient Feature Images for the Prediction of Prostate Cancer Metastasis and Identification of Significant PathwaysabstractHigh-performance prediction of prostate cancer metastasis based on single nucleotide variations remains a challenge. Therefore, we developed a novel biologically informed deep learning framework, named M-NET, for the prediction of prostate cancer metastasis. Within the framework, we transformed single nucleotide variations into patient feature images that are optimal for fitting convolutional neural networks. Moreover, we identified significant pathways associated with the metastatic status. The experimental results showed that M-NET significantly outperformed other comparison methods based on single nucleotide variations, achieving improvements in accuracy, precision, recall, F1-score, area under the receiver operating characteristics curve, and area under the precision-recall curve by 6.3%, 8.4%, 5.1%, 0.070, 0.041, and 0.026, respectively. Furthermore, M-NET identified some important pathways associated with the metastatic status, such as signaling by the hedgehog pathway. In summary, compared with other comparative methods, M-NET exhibited a better performance in the prediction of prostate cancer metastasis. Jie Li 0055, Weilong Tan |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | DVA: predicting the functional impact of single nucleotide missense variantsabstractBACKGROUND: In the past decade, single nucleotide variants (SNVs) have been identified as having a significant relationship with the development and treatment of diseases. Among them, prioritizing missense variants for further functional impact investigation is an essential challenge in the study of common disease and cancer. Although several computational methods have been developed to predict the functional impacts of variants, the predictive ability of these methods is still insufficient in the Mendelian and cancer missense variants. RESULTS: We present a novel prediction method called the disease-related variant annotation (DVA) method that predicts the effect of missense variants based on a comprehensive feature set of variants, notably, the allele frequency and protein-protein interaction network feature based on graph embedding. Benchmarked against datasets of single nucleotide missense variants, the DVA method outperforms the state-of-the-art methods by up to 0.473 in the area under receiver operating characteristic curve. The results demonstrate that the proposed method can accurately predict the functional impact of single nucleotide missense variants and substantially outperforms existing methods. CONCLUSIONS: DVA is an effective framework for identifying the functional impact of disease missense variants based on a comprehensive feature set. Based on different datasets, DVA shows its generalization ability and robustness, and it also provides innovative ideas for the study of the functional mechanism and impact of SNVs. Dong Wang 0066, Jie Li 0055, Edwin Wang, Yadong Wang 0001 |
BMC Bioinform. | 2 |
| 2024 | CLUE: Contrastive language-guided learning for referring video object segmentation
Wanjun Zhong, Jie Li 0055, Tiejun Zhao |
Pattern Recognit. Lett. | 3 |
| 2023 | Leveraging Visual Prompts To Guide Language Modeling for Referring Video Object SegmentationabstractReferring Video Object Segmentation (R-VOS) aims to segment object masks in a target video given a language query describing the object. It is a challenging task that requires modeling the semantics of a natural language query and its correspondence to the target video. Previous works directly use visual-agnostic language features from uni-modal language models, and only interact with visual features in late decoding stages. We propose to encode visual-enriched language features by using visual prompts as guidance in the early encoding stage. The proposed visual prompt is constructed by modulating visual features of key frames with alignment scores to text inputs. The alignment score is computed with a pre-trained visual-language contrastive model. We concatenate visual prompts with text inputs to encode visual-enriched language features, which serve as queries for target object segmentation in a Transformer-based decoder. Our method outperforms the previous state-of-the-art method (+2.3) on Refer-Youtube-VOS benchmark. Wanjun Zhong, Jie Li 0055, Tiejun Zhao |
ICIP | 3 |
| 2023 | MTGDC: A Multi-Scale Tensor Graph Diffusion Clustering for Single-Cell RNA Sequencing DataabstractSingle-cell RNA sequencing (scRNA-seq) is a new technology that focuses on the expression levels for each cell to study cell heterogeneity. Thus, new computational methods matching scRNA-seq are designed to detect cell types among various cell groups. Herein, we propose a Multi-scale Tensor Graph Diffusion Clustering (MTGDC) for single-cell RNA sequencing data. It has the following mechanisms: 1) To mine potential similarity distributions among cells, we design a multi-scale affinity learning method to construct a fully connected graph between cells; 2) For each affinity matrix, we propose an efficient tensor graph diffusion learning framework to learn high-order information among multi-scale affinity matrices. First, the tensor graph is explicitly introduced to measure cell-cell edges with local high-order relationship information. To further preserve more global topology structure information in the tensor graph, MTGDC implicitly considers the propagation of information via a data diffusion process by designing a simple and efficient tensor graph diffusion update algorithm. 3) Finally, we mix together the multi-scale tensor graphs to obtain the fusion high-order affinity matrix and apply it to spectral clustering. Experiments and case studies showed that MTGDC had obvious advantages over the state-of-art algorithms in robustness, accuracy, visualization, and speed. Qiaoming Liu, Dong Wang 0066, Jie Li 0055, Guohua Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | MicroRNA Promoter Identification in Human With a Three-level Prediction MethodabstractThe accurate annotation of miRNA promoters is critical for the mechanistic understanding of miRNA gene regulation. Various computational methods have been developed for the prediction of miRNA promoters solely employing a single classifier. Most of these computational methods extract either sequence features or one-sided signal features, and the accuracy and reliability of predictions need to be improved. To address these issues, we present miPTP, a three-level prediction method that combines SVM, RF, and correlation coefficients. It is capable of identifying miRNA promoters based on both DNA sequence and ChIP-Seq data (RPol II). By sequentially integrating these two types of information sources with the three methods selected, miPTP can identify miRNA promoters with higher accuracy and sensitivity compared to specific existing methods. Finally, the reliability of miPTP is validated by examining the conservation, CpG content, and activating histone marks in the identified miRNA promoters. Xin Wang 0124, Jie Li 0055, Guohua Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Predicting Drug-Disease Associations Through Similarity Network Fusion and Multi-View Feature Projection RepresentationabstractPredicting drug-disease associations (DDAs) through computational methods has become a prevalent trend in drug development because of their high efficiency and low cost. Existing methods usually focus on constructing heterogeneous networks by collecting multiple data resources to improve prediction ability. However, potential association possibilities of numerous unconfirmed drug-related or disease-related pairs are not sufficiently considered. In this article, we propose a novel computational model to predict new DDAs. First, a heterogeneous network is constructed, including four types of nodes (drugs, targets, cell lines, diseases) and three types of edges (associations, association scores, similarities). Second, an updating and merging-based similarity network fusion method, termed UM-SF, is presented to fuse various similarity networks with diverse weights. Finally, an intermediate layer-mediated multi-view feature projection representation method, termed IM-FP, is proposed to calculate the predicted DDA scores. This method uses multiple association scores to construct multi-view drug features, then projects them into disease space through the intermediate layer, where an intermediate layer similarity constraint is designed to learn the projection matrices. Results of comparative experiments reveal the effectiveness of our innovations. Comparisons with other state-of-the-art models by the 10-fold cross-validation experiment indicate our model's advantage on AUROC and AUPR metrics. Moreover, our proposed model successfully predicted 107 novel high-ranked DDAs. Jie Li 0055, Dong Wang 0066, Dechen Xu, Jiahuan Jin, Yadong Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Real-time Image Enhancement with Attention AggregationabstractImage enhancement has stimulated significant research works over the past years for its great application potential in video conferencing scenarios. Nevertheless, most existing image enhancement approaches are still struggling to find a good tradeoff that reduces the computational cost as much as possible while maintaining plausible result quality. Recently, curve-based mapping methods are proposed and have shown great potential for real-time and high-quality image enhancement of arbitrary resolutions. In this article, we take advantage of the curve-based mapping representation and focus on further improving the enhancement quality and robustness, while minimizing additional computational costs. Specifically, we (1) carefully re-formulate the curve function to improve learning stability, and (2) aggregate different semantic attention into the curve regression process, which can overcome the major problems of curve-based methods that generate moderate results with low contrast. The semantic attention is jointly learned with the supervision from class activation mapping of pre-trained feature extractors, thus reducing the manual annotation cost of semantic labels. Experiments have shown that our proposed method significantly improves curve-based methods both qualitatively and quantitatively, achieving visually plausible results compared with other deep neural network-based enhancement methods, and maintains a very low computational cost, i.e., taking 18.7 ms for a 360p image on a single P40 GPU. Extensive experiments demonstrate that our method is also capable of video enhancement tasks. Jie Li 0055, Tiejun Zhao, Yadong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | scESI: evolutionary sparse imputation for single-cell transcriptomes from nearest neighbor cellsabstractThe ubiquitous dropout problem in single-cell RNA sequencing technology causes a large amount of data noise in the gene expression profile. For this reason, we propose an evolutionary sparse imputation (ESI) algorithm for single-cell transcriptomes, which constructs a sparse representation model based on gene regulation relationships between cells. To solve this model, we design an optimization framework based on nondominated sorting genetics. This framework takes into account the topological relationship between cells and the variety of gene expression to iteratively search the global optimal solution, thereby learning the Pareto optimal cell-cell affinity matrix. Finally, we use the learned sparse relationship model between cells to improve data quality and reduce data noise. In simulated datasets, scESI performed significantly better than benchmark methods with various metrics. By applying scESI to real scRNA-seq datasets, we discovered scESI can not only further classify the cell types and separate cells in visualization successfully but also improve the performance in reconstructing trajectories differentiation and identifying differentially expressed genes. In addition, scESI successfully recovered the expression trends of marker genes in stem cell differentiation and can discover new cell types and putative pathways regulating biological processes. Qiaoming Liu, Ximei Luo, Jie Li 0055, Guohua Wang 0001 |
Briefings Bioinform. | 3 |
| 2022 | StackCirRNAPred: computational classification of long circRNA from other lncRNA based on stacking strategyabstractBACKGROUND: CircRNAs are essential for the regulation of post-transcriptional gene expression, including as miRNA sponges, and play an important role in disease development. Some computational tools have been proposed recently to predict circRNA, since only one classifier is used, there is still much that can be done to improve the performance. RESULTS: StackCirRNAPred was proposed, the computational classification of long circRNA from other lncRNA based on stacking strategy. In order to cope with the potential problem that a single feature might not be able to distinguish circRNA well from other lncRNA, we first extracted features from different sources, including nucleic acid composition, sequence spatial features and physicochemical properties, Alu and tandem repeats. We innovatively apply the stacking strategy to integrate the more advantageous classifiers of RF, LightGBM, XGBoost. This allows the model to incorporate these features more flexibly. StackCirRNAPred was found to be significantly better than other tools, with precision, accuracy, F1, recall and MCC of 0.843, 0.833, 0.831, 0.819 and 0.666 respectively. We tested it directly on the mouse dataset. StackCirRNAPred was still significantly better than other methods, with precision, accuracy, F1, recall and MCC of 0.837, 0.839, 0.839, 0.841, 0.677. CONCLUSIONS: We proposed StackCirRNAPred based on stacking strategy to distinguish long circRNAs from other lncRNAs. With the test results demonstrating the validity and robustness of StackCirRNAPred, we hope StackCirRNAPred will complement existing circRNA prediction methods and is helpful in down-stream research. Xin Wang 0124, Yadong Liu 0001, Jie Li 0055, Guohua Wang 0001 |
BMC Bioinform. | 3 |
| 2022 | M2PP: a novel computational model for predicting drug-targeted pathogenic proteinsabstractBACKGROUND: Detecting pathogenic proteins is the origin way to understand the mechanism and resist the invasion of diseases, making pathogenic protein prediction develop into an urgent problem to be solved. Prediction for genome-wide proteins may be not necessarily conducive to rapidly cure diseases as developing new drugs specifically for the predicted pathogenic protein always need major expenditures on time and cost. In order to facilitate disease treatment, computational method to predict pathogenic proteins which are targeted by existing drugs should be exploited. RESULTS: In this study, we proposed a novel computational model to predict drug-targeted pathogenic proteins, named as M2PP. Three types of features were presented on our constructed heterogeneous network (including target proteins, diseases and drugs), which were based on the neighborhood similarity information, drug-inferred information and path information. Then, a random forest regression model was trained to score unconfirmed target-disease pairs. Five-fold cross-validation experiment was implemented to evaluate model's prediction performance, where M2PP achieved advantageous results compared with other state-of-the-art methods. In addition, M2PP accurately predicted high ranked pathogenic proteins for common diseases with public biomedical literature as supporting evidence, indicating its excellent ability. CONCLUSIONS: M2PP is an effective and accurate model to predict drug-targeted pathogenic proteins, which could provide convenience for the future biological researches. Jie Li 0055, Yadong Wang 0001 |
BMC Bioinform. | 2 |
| 2022 | A Neighborhood-Based Global Network Model to Predict Drug-Target InteractionsabstractThe detection of drug-target interactions (DTIs) plays an important role in drug discovery and development, making DTI prediction urgent to be solved. Existing computational methods usually utilize drug similarity, target similarity and DTI information to make prediction, providing the convenience of fast time and low cost. However, they usually learn features for drugs and targets separately, lacking of a global consideration. In this study, we proposed a novel neighborhood-based global network model, named as NGN, to accurately predict DTIs from the global perspective. We designed a distance constraint for features of all entities (drugs and targets) in the latent space to ensure the close distance between adjacent entities, and defined a global probability matrix to compute the predicted DTI scores on our constructed neighborhood-based global network. Results showed that NGN obtained advantageous performance compared with other state-of-the-art methods, especially surpassing them by 4.2-9.1 percent on AUPR values in the biggest dataset. Furthermore, several novel high-ranked DTIs were successfully predicted with confirmations by public sources, demonstrating the effectiveness of our method. Jie Li 0055, Yadong Wang 0001, Liran Juan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | WMMDCA: Prediction of Drug Responses by Weight-Based Modular Mapping in Cancer Cell LinesabstractDue to the high consumption of cost and time for experimental verification in clinical trials, drug response prediction by computational models have become important challenges. The existing drug response data in diverse cell lines enable prediction of potential sensitive associations. Here, we propose a weight-based modular mapping method, named as WMMDCA, to predict drug-cell line associations. The method fully considers the effects of drugs' chemical structural feature, and adds modular information into the network projection. Leave-one-out cross-validation was used to evaluate the predictive ability of WMMDCA, which showed the best performance among several state-of-the-art methods in not only the whole dataset but also the major tissue types of cell lines. Literature support of highly ranked potential associations was found manually, demonstrating the effectiveness of WMMDCA on drug response prediction. Jie Li 0055, Yadong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | Y-SPCR: A new dimensionality reduction method for gene expression data classificationabstractWith the increase of the scale and complexity of massive data, data dimensionality reduction technologies, such as principal component analysis, have developed rapidly. The performance of dimension reduction technologies still needs to be further improved. In the paper we proposed a new dimensionality reduction method (Y-SPCR) based Supervised Principal Component Regression (SPCR) and Y-aware Principal Component Regression (Y-aware PCR). Experimental results on four gene expression data sets show that Y-SPCR effectively overcomes the shortcomings of SPCR and Y-aware PCR and improves the accuracy and stability of on gene expression data classification. Jie Li 0055, Zhun Zhao, Yadong Wang 0001 |
BIBM | 1 |
| 2016 | A network-based pathway-expanding approach for pathway analysisabstractBACKGROUND: Pathway analysis combining multiple types of high-throughput data, such as genomics and proteomics, has become the first choice to gain insights into the pathogenesis of complex diseases. Currently, several pathway analysis methods have been developed to study complex diseases. However, these methods did not take into account the interaction between internal and external genes of the pathway and between pathways. Hence, these approaches still face some challenges. Here, we propose a network-based pathway-expanding approach that takes the topological structures of biological networks into account. RESULTS: First, two weighted gene-gene interaction networks (tumor and normal) are constructed integrating protein-protein interaction(PPI) information, gene expression data and pathway databases. Then, they are used to identify significant pathways through testing the difference of topological structures of expanded pathways in the two weighted networks. The proposed method is employed to analyze two breast cancer data. As a result, the top 15 pathways identified using the proposed method are supported by biological knowledge from the published literatures and other methods. In addition, the proposed method is also compared with other methods, such as GSEA and SPIA, and estimated using the classification performance of the top 15 expanded pathways. CONCLUSIONS: A novel network-based pathway-expanding approach is proposed to avoid the limitations of existing pathway analysis approaches. Experimental results indicate that the proposed method can accurately and reliably identify significant pathways which are related to the corresponding disease. Jie Li 0055, Haozhe Xie, Hanqing Xue, Yadong Wang 0001 |
BMC Bioinform. | 2 |
| 2015 | Using Semantic Association to Extend and Infer Literature-Oriented Relativity Between TermsabstractRelative terms often appear together in the literature. Methods have been presented for weighting relativity of pairwise terms by their co-occurring literature and inferring new relationship. Terms in the literature are also in the directed acyclic graph of ontologies, such as Gene Ontology and Disease Ontology. Therefore, semantic association between terms may help for establishing relativities between terms in literature. However, current methods do not use these associations. In this paper, an adjusted R-scaled score (ARSS) based on information content (ARSSIC) method is introduced to infer new relationship between terms. First, set inclusion relationship between terms of ontology was exploited to extend relationships between these terms and literature. Next, the ARSS method was presented to measure relativity between terms across ontologies according to these extensional relationships. Then, the ARSSIC method using ratios of information shared of term's ancestors was designed to infer new relationship between terms across ontologies. The result of the experiment shows that ARSS identified more pairs of statistically significant terms based on corresponding gene sets than other methods. And the high average area under the receiver operating characteristic curve (0.9293) shows that ARSSIC achieved a high true positive rate and a low false positive rate. Data is available at http://mlg.hit.edu.cn/ARSSIC/. Liang Cheng 0006, Jie Li 0055, Yang Hu 0008, Yongzhuang Liu, Yan-Shuo Chu, Yadong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2007 | A new framework for identifying differentially expressed genes
Jie Li 0055, Xianglong Tang, Wei Zhao 0008 |
Pattern Recognit. | 1 |