EDBT 2026 Demo / reviewers in the wild / expert
Jianing Xi
dblp:184/2479
· DBLP profile ↗
18ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0001-6785-5618ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Channel transformer based multi field-of-view model to detect tumor spread through air space in histopathological imagesabstractAccurate detection of Tumor Spread Through Air Spaces (STAS) is pivotal for patient prognosis assessment and therapeutic strategies. Current mainstream histopathology images detection methodologies rely on single Field-of-View (FoV), which overlook global context and are susceptible to misclassifying STAS. In addition, despite good progress in utilizing multi FoV for histopathology images detection, the semantic gap that exists between different FoV tasks reduces model performance. To address these issues, we introduce the Multi FoV Channel-wise Transformer Detection Model (MFCT), which harnesses cross attention mechanisms to fuse multi FoV features. MFCT enhances the capacity of small FoV to capture contextual information, thereby augmenting model precision. Moreover, In order to enhance the application scenarios of the model, the model employs a detection head and a segmentation head respectively. To evaluate MFCT’s performance, experiments were carried out using both single and multi FoV setups, employing the publicly available Ocelot dataset as well as our custom STAS dataset. MFCT achieved a F1 score of 73.4% on the Ocelot dataset and 85.2% on our STAS dataset. Specifically, F1 scores improved by 5% compared to the single FoV model and by 3% compared to the multi FoV baseline model on the Ocelot, and the results indicate that MFCT outperforms the other method compared. The empirical evidence suggests that our model furnishes a detection framework that facilitates and accelerates the exploitation of multi FoV information in histopathological image detection. • Leveraging multi Field of View information for accurate pathology image detection. • Building a channel-based Transformer to leverage large Field of View information. • Model performance exceeds current methods using large Field of View information. • Detection model uses large Field of View info without extra annotations. Haotian Gong, Jianing Xi, Sisi Chen, Shuanlong Che, Ling Qi, Guiying Zhang |
Expert Syst. Appl. | 2 |
| 2025 | Explainable Reasoning Path Inference of Anti-Cancer Drug Sensitivity on Genomic Knowledge Graph via Macro-Micro Agent Collaborative Reinforcement LearningabstractArtificial intelligence (AI) based anticancer drug recommendation systems have emerged as powerful tools for precision dosing. Although existing methods have advanced in terms of predictive accuracy, they encounter three significant obstacles, including the "black-box" problem resulting in unexplainable reasoning, the computational difficulty for graph-based structures, and the combinatorial explosion during multi-step reasoning. To tackle these issues, we introduce a novel Macro-Micro agent Drug sensitivity inference (MarMirDrug). Specifically, our methodology enhances interpretability via knowledge graphs (KG). To reduce computational overhead, our MarMirDrug also transforms the graph structures of KG into low-dimensional embeddings. To manage the combinatorial explosion of graph paths that occurs with increasing inference steps, we incorporate a macro-micro dual-agent reinforcement learning algorithm, and combines reinforcement learning with KG-based reasoning to infer path reasoning. In computational experimental outcomes, the efficiency of our embedding model surpasses that of the baseline model, achieving an average improvement of 117.65% for hit score. Also, our dual-agent framework exhibits superior performance and interpretability in drug response prediction, with AUC values outperforming conventional baseline methods. In summary, our approach integrates interpretability, computational efficiency, and predictive accuracy, providing novel contributions to precision medicine. Minhua Feng, Juntao Liang, Zhimin Zheng, Ranran Guo, Wen Shi 0009, Jianing Xi |
IEEE Trans. Comput. Biol. Bioinform. | 9 |
| 2023 | LUR: An Online Learning Model for EEG Emotion RecognitionabstractEmotion recognition based on EEG (Electroen-cephalogram) has been widely used in may scenarios, such as Brain-computer interface, medical health and entertainment, etc. However, large differences exist in subjects due to individual characteristics, and EEG data of different time periods usually distribute inconsistently, which hinder the further development of EEG-based research. In this paper, we propose a LUR model to partially address these challenges. In the first stage, we extracted candidate features based on neuroscience research and learned them using SVM or Naive Bayes algorithm. If the performance during the validation phase is satisfactory, we keep the features. Otherwise we replace the candidate features and use an online learning method named LUR(Learn, Unlearn, and Relearn), which continuously prunes the irrelevant connections of the current data and retains important connections to constructing a model more suitable for the subject. Because the proposed model is based on streaming data, it can continuously relearn and solve the problem that the input data is not i.i.d. to a certain extent. Experiments were conducted on two public datasets, DEAP and DREAMER, and competitive results were obtained. Liying Yang 0001, Qian Zhang 0074, Jianing Xi, Chengchuang Tang |
BIBM | 4 |
| 2023 | Pathological Tissue-level Contour Genomic Profile Interpretation of Lung Adenocarcinoma via Spatial and Morphological Features Co-action Graph Neural NetworkabstractThe affections from genomics to morphology can prompt the genomic profile interpretation from inexpensive pathological image data rather than highly cost genomic sequencing data. Due to the extremely large size of Whole Slide Image (WSI), directly processing the complete Lung Adenocarcinoma (LUAD) WSI with traditional deep learning methods, will lead to memory overflow. In comparison to complete WSI, the smaller size patches can be easier processed by the existing deep learning based methods. Nevertheless, the split patches severely break the potential relationships between genomic abnormalities and morphological features, and the traditional deep learning methods may be difficult to capture the information of the broken relationship. Fortunately, a recent study has shown that the graph-structure representation can feasibly demonstrate the relationships among both local and remote regions. In consideration of the obstacles of both the break of remote area relationships and the lack of tissue-level contour for genomics-to-morphology associations, we propose Spatial and Morphological Features Co-action Graph Neural Network model (SMCGNN) to achieve the pathological tissue-level contour genomic profile interpretation of LUAD. Our SMCGNN achieves better performance on the genomic profile interpretation task than those of previous researches, yielding a relative performance increment of 9.3%. To the best of our knowledge, our SMCGNN is the first model to interpret the biological tissue-level contour of the genomic abnormality-related morphological regions. In summary, our method can provide pathologists with more fine-grained hints to molecular profile. The interpreted tissue-level regions can be accessed via the link: https://github.com/xianyvxxx/tissue-level-contour-regions-associated-with-genomic-profile. Wen Shi 0009, Guoxi Xie, Jianing Xi |
BIBM | 4 |
| 2022 | A parametric model for clustering single-cell mutation dataabstractClustering tumor single-cell mutation data has formed an important paradigm for deciphering tumor subclones and evolutionary history. This type of data may often be heavily complicated by incompleteness, false positives and false negatives errors. Despite to the fact that several computational methods have been developed for clustering binary mutation data, their applications still suffer from degraded accuracy on large datasets or datasets with high sparsity. Therefore, more effective methods are sorely required. Here, we propose a novel method called CBM for reliably Clustering Binary Mutation data. CBM formulates the binary mutation data under a probabilistic framework through parameterizing false positive errors, false negative errors, presence probability distribution of subclones and their binary mutation profiles. To cope with the difficulty of optimizing discrete parameters, Gibbs sampling for mixtures is employed to iteratively sample cell-to-cluster assignments and cluster centers from the posterior. Extensive evaluations on simulated and real datasets demonstrate CBM outperforms the state-of-the-art tools in different performance metrics such as ARI for clustering and accuracy for genotyping. CBM can be integrated into the pipeline of reconstructing tumor evolutionary tree, and detecting subclones using CBM can be employed as a pre-text task of tumor subclonal tree inference, which will significantly improve computational efficiency of phylogenetic analysis especially on large datasets. CBM software is freely available at https://github.com/zhyu-lab/cbm. Jiaqian Yan, Jianing Xi, Zhenhua Yu 0002 |
BIBM | 2 |
| 2022 | Knowledge tensor embedding framework with association enhancement for breast ultrasound diagnosis of limited labeled samples
Jianing Xi, Zhaoji Miao, Longzhong Liu, Xuebing Yang, Wensheng Zhang 0002, Qinghua Huang, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2022 | Structure-aware siamese graph neural networks for encounter-level patient similarity learning
Xuebing Yang, Lei Tian 0007, Jicheng Lv, Jianing Xi, Guilan Kong, Wensheng Zhang 0002 |
J. Biomed. Informatics | 8 |
| 2021 | Tolerating Data Missing in Breast Cancer Diagnosis from Clinical Ultrasound Reports via Knowledge Graph InferenceabstractMedical diagnosis through artificial intelligence has been drawing increasing attention currently. For breast lesions, the clinical ultrasound reports are the most commonly used data in the diagnosis of breast cancer. Nevertheless, the input reports always encounter the inevitable issue of data missing. Unfortunately, despite the efforts made in previous approaches that made progress on tackling data imprecision, nearly all of these approaches cannot accept inputs with data missing. A common way to alleviate the data missing issue is to fill the missing values with artificial data. However, the data filling strategy actually brings in additional noises that do not exist in the raw data. Inspired by the advantage of open world assumption, we regard the missing data in clinical ultrasound reports as non-observed terms of facts, and propose a Knowledge Graph embedding based model KGSeD with the capability of tolerating data missing, which can successfully circumvent the pollution caused by data filling. Our KGSeD is designed via an encoder-decoder framework, where the encoder incorporates structural information of the graph via embedding, and the decoder diagnose patients by inferring their links to clinical outcomes. Comparative experiments show that KGSeD achieves noticeable diagnosis performances. When data missing occurred, KGSeD yields the most stable performance over those of existing approaches, showing better tolerance to data missing. Jianing Xi, Liping Ye, Qinghua Huang, Xuelong Li 0001 |
KDD | 1 |
| 2021 | A Local Outlier Factor-Based Detection of Copy Number Variations From NGS DataabstractCopy number variation (CNV) is a major type of genomic structural variations that play an important role in human disorders. Next generation sequencing (NGS) has fueled the advancement in algorithm design to detect CNVs at base-pair resolution. However, accurate detection of CNVs of low amplitudes remains a challenging task. This paper proposes a new computational method, CNV-LOF, to identify CNVs of full-range amplitudes from NGS data. CNV-LOF is distinctly different from traditional methods, which mainly consider aberrations from a global perspective and rely on some assumed distribution of NGS read depths. In contrast, CNV-LOF takes a local view on the read depths and assigns an outlier factor to each genome segment. With the outlier factor profile, CNV-LOF uses a boxplot procedure to declare CNVs without the reliance of any distribution assumptions. Simulation experiments indicate that CNV-LOF outperforms five existing methods with respect to F1-measure, sensitivity, and precision. CNV-LOF is further validated on real sequencing samples, yielding highly consistent results with peer methods. CNV-LOF is able to detect CNVs of low and moderate amplitudes where the other existing methods fail, and it is expected to become a routine approach for the discovery of novel CNVs on whole sequencing genome. Xiguo Yuan, Junping Li, Jianing Xi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | STIC: Predicting Single Nucleotide Variants and Tumor Purity in Cancer GenomeabstractSingle nucleotide variant (SNV) plays an important role in cellular proliferation and tumorigenesis in various types of human cancer. Next-generation sequencing (NGS) has provided high-throughput data at an unprecedented resolution to predict SNVs. Currently, there exist many computational methods for either germline or somatic SNV discovery from NGS data, but very few of them are versatile enough to adapt to any situations. In the absence of matched normal samples, the prediction of somatic SNVs from single-tumor samples becomes considerably challenging, especially when the tumor purity is unknown. Here, we propose a new approach, STIC, to predict somatic SNVs and estimate tumor purity from NGS data without matched normal samples. The main features of STIC include: (1) extracting a set of SNV-relevant features on each site and training the BP neural network algorithm on the features to predict SNVs; (2) creating an iterative process to distinguish somatic SNVs from germline ones by disturbing allele frequency; and (3) establishing a reasonable relationship between tumor purity and allele frequencies of somatic SNVs to accurately estimate the purity. We quantitatively evaluate the performance of STIC on both simulation and real sequencing datasets, the results of which indicate that STIC outperforms competing methods. Xiguo Yuan, Haiyong Zhao, Liying Yang 0001, Shuzhen Wang, Jianing Xi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | CNV_IFTV: An Isolation Forest and Total Variation-Based Detection of CNVs from Short-Read Sequencing DataabstractAccurate detection of copy number variations (CNVs) from short-read sequencing data is challenging due to the uneven distribution of reads and the unbalanced amplitudes of gains and losses. The direct use of read depths to measure CNVs tends to limit performance. Thus, robust computational approaches equipped with appropriate statistics are required to detect CNV regions and boundaries. This study proposes a new method called CNV_IFTV to address this need. CNV_IFTV assigns an anomaly score to each genome bin through a collection of isolation trees. The trees are trained based on isolation forest algorithm through conducting subsampling from measured read depths. With the anomaly scores, CNV_IFTV uses a total variation model to smooth adjacent bins, leading to a denoised score profile. Finally, a statistical model is established to test the denoised scores for calling CNVs. CNV_IFTV is tested on both simulated and real data in comparison to several peer methods. The results indicate that the proposed method outperforms the peer methods. CNV_IFTV is a reliable tool for detecting CNVs from short-read sequencing data even for low-level coverage and tumor purity. The detection results on tumor samples can aid to evaluate known cancer genes and to predict target drugs for disease diagnosis. Xiguo Yuan, Jianing Xi, Liying Yang 0001, Junliang Shang, Junbo Duan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Inferring subgroup-specific driver genes from heterogeneous cancer samples via subspace learning with subgroup indicationabstractMOTIVATION: Detecting driver genes from gene mutation data is a fundamental task for tumorigenesis research. Due to the fact that cancer is a heterogeneous disease with various subgroups, subgroup-specific driver genes are the key factors in the development of precision medicine for heterogeneous cancer. However, the existing driver gene detection methods are not designed to identify subgroup specificities of their detected driver genes, and therefore cannot indicate which group of patients is associated with the detected driver genes, which is difficult to provide specifically clinical guidance for individual patients. RESULTS: By incorporating the subspace learning framework, we propose a novel bioinformatics method called DriverSub, which can efficiently predict subgroup-specific driver genes in the situation where the subgroup annotations are not available. When evaluated by simulation datasets with known ground truth and compared with existing methods, DriverSub yields the best prediction of driver genes and the inference of their related subgroups. When we apply DriverSub on the mutation data of real heterogeneous cancers, we can observe that the predicted results of DriverSub are highly enriched for experimentally validated known driver genes. Moreover, the subgroups inferred by DriverSub are significantly associated with the annotated molecular subgroups, indicating its capability of predicting subgroup-specific driver genes. AVAILABILITY AND IMPLEMENTATION: The source code is publicly available at https://github.com/JianingXi/DriverSub. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianing Xi, Xiguo Yuan, Ao Li 0001, Xuelong Li 0001, Qinghua Huang |
Bioinform. | 1 |
| 2020 | Dual-Layer Strengthened Collaborative Topic Regression Modeling for Predicting Drug SensitivityabstractAn effective way to facilitate the development of modern oncology precision medicine is the systematical analysis of the known drug sensitivities that have emerged in recent years. Meanwhile, the screening of drug response in cancer cell lines provides an estimable genomic and pharmacological data towards high accuracy prediction. Existing works primarily utilize genomic or functional genomic features to classify or regress the drug response. Here in this work, by the migration and extension of the conventional merchandise recommendation methods, we introduce an innovation model on accurate drug sensitivity prediction by using dual-layer strengthened collaborative topic regression (DS-CTR), which incorporates not only the graphic model to jointly learn drugs and cell lines feature from pharmacogenomics data but also drug and cell line similarity network model to strengthen the correlation of the prediction results. Using Genomics of Drug Sensitivity in Cancer project (GDSC) as benchmark datasets, the 5-fold cross-validation experiment demonstrates that DS-CTR model significantly improves drug response prediction performance compared with four categories of state-of-the-art algorithms as for both Receiver Operator Curve (ROC) and the Area Under Receiver Operator Curve (AUC). By uncovering the unknown cell-drug associations with advanced literature evidences, our novel model DS-CTR is validated and supported. The model also provides the possibility to make the discovery of new anti-cancer therapeutics in the preclinical trials cheaper and faster. Jianing Xi, Ao Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | HetRCNA: A Novel Method to Identify Recurrent Copy Number Alternations from Heterogeneous Tumor Samples Based on Matrix Decomposition FrameworkabstractA common strategy to discovering cancer associated copy number aberrations (CNAs) from a cohort of cancer samples is to detect recurrent CNAs (RCNAs). Although the previous methods can successfully identify communal RCNAs shared by nearly all tumor samples, detecting subgroup-specific RCNAs and their related subgroup samples from cancer samples with heterogeneity is still invalid for these existing approaches. In this paper, we introduce a novel integrated method called HetRCNA, which can identify statistically significant subgroup-specific RCNAs and their related subgroup samples. Based on matrix decomposition framework with weight constraint, HetRCNA can successfully measure the subgroup samples by coefficients of left vectors with weight constraint and subgroup-specific RCNAs by coefficients of the right vectors and significance test. When we evaluate HetRCNA on simulated dataset, the results show that HetRCNA gives the best performances among the competing methods and is robust to the noise factors of the simulated data. When HetRCNA is applied on a real breast cancer dataset, our approach successfully identifies a bunch of RCNA regions and the result is highly correlated with the results of the other two investigated approaches. Notably, the genomic regions identified by HetRCNA harbor many breast cancer related genes reported by previous researches. Jianing Xi, Ao Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | LPGNMF: Predicting Long Non-Coding RNA and Protein Interaction Using Graph Regularized Nonnegative Matrix FactorizationabstractLong non-coding RNAs (lncRNA) play crucial roles in a variety of biological processes and complex diseases. Massive studies have indicated that lncRNAs interact with related proteins to exert regulation of cellular biological processes. Because it is time-consuming and expensive to determine lncRNA-protein interaction by experiment, more accurate predictions of interaction by computational methods are imperative. We propose a novel computational approach, predicting lncRNA-protein interaction using graph regularized nonnegative matrix factorization (LPGNMF), to discover unobserved lncRNA-protein association. First, we calculate lncRNA similarity and protein similarity by integrating the lncRNA expression information and gene ontology information. Subsequently, we utilize graph regularized nonnegative matrix factorization framework to predict potential interactions for all lncRNA simultaneously. In the cross validation test, LPGNMF achieves an AUC of 85.2 percent, higher than those of other compared methods. In addition, novel lncRNA-protein interactions detected by LPGNMF are validated by literatures or database. The results indicate that our method is effective to discover potential lncRNA-protein interaction. Jianing Xi, Ao Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Discovering mutated driver genes through a robust and sparse co-regularized matrix factorization framework with prior information from mRNA expression patterns and interaction networkabstractBACKGROUND: Discovery of mutated driver genes is one of the primary objective for studying tumorigenesis. To discover some relatively low frequently mutated driver genes from somatic mutation data, many existing methods incorporate interaction network as prior information. However, the prior information of mRNA expression patterns are not exploited by these existing network-based methods, which is also proven to be highly informative of cancer progressions. RESULTS: To incorporate prior information from both interaction network and mRNA expressions, we propose a robust and sparse co-regularized nonnegative matrix factorization to discover driver genes from mutation data. Furthermore, our framework also conducts Frobenius norm regularization to overcome overfitting issue. Sparsity-inducing penalty is employed to obtain sparse scores in gene representations, of which the top scored genes are selected as driver candidates. Evaluation experiments by known benchmarking genes indicate that the performance of our method benefits from the two type of prior information. Our method also outperforms the existing network-based methods, and detect some driver genes that are not predicted by the competing methods. CONCLUSIONS: In summary, our proposed method can improve the performance of driver gene discovery by effectively incorporating prior information from interaction network and mRNA expression patterns into a robust and sparse co-regularized matrix factorization framework. Jianing Xi, Ao Li 0001 |
BMC Bioinform. | 1 |
| 2018 | A novel unsupervised learning model for detecting driver genes from pan-cancer data through matrix tri-factorization framework with pairwise similarities constraints
Jianing Xi, Ao Li 0001 |
Neurocomputing | 1 |
| 2016 | Discovering Recurrent Copy Number Aberrations in Complex Patterns via Non-Negative Sparse Singular Value DecompositionabstractRecurrent copy number aberrations (RCNAs) in multiple cancer samples are strongly associated with tumorigenesis, and RCNA discovery is helpful to cancer research and treatment. Despite the emergence of numerous RCNA discovering methods, most of them are unable to detect RCNAs in complex patterns that are influenced by complicating factors including aberration in partial samples, co-existing of gains and losses and normal-like tumor samples. Here, we propose a novel computational method, called non-negative sparse singular value decomposition (NN-SSVD), to address the RCNA discovering problem in complex patterns. In NN-SSVD, the measurement of RCNA is based on the aberration frequency in a part of samples rather than all samples, which can circumvent the complexity of different RCNA patterns. We evaluate NN-SSVD on synthetic dataset by comparison on detection scores and Receiver Operating Characteristics curves, and the results show that NN-SSVD outperforms existing methods in RCNA discovery and demonstrate more robustness to RCNA complicating factors. Applying our approach on a breast cancer dataset, we successfully identify a number of genomic regions that are strongly correlated with previous studies, which harbor a bunch of known breast cancer associated genes. Jianing Xi, Ao Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |