EDBT 2026 Demo / reviewers in the wild / expert
Xuegong Zhang
dblp:36/881
· DBLP profile ↗
68ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-9684-5643ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 48 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-dataset annotation harmonization for cell-type hierarchy constructionabstractMOTIVATION: Single-cell transcriptomic datasets annotate cell types with diverse schemes and varying resolution. This poses challenges in building unified hierarchical cell-type structures and hinders integration of large-scale datasets. To address this, several computational methods have been developed to harmonize cell type annotations across datasets and build data-driven hierarchies of cell types. RESULTS: Here, we benchmarked three state-of-the-art methods: scHPL, treeArches, and CellHint. We evaluated these methods across five simulated scenarios and five real-world scenarios across cell types and organs. To assess harmonization results, we designed three metrics, Annotation Harmonization F1-score (AH-F1), Tree Edit Distance Similarity and Parent-Children Branches Similarity, comparing the constructed cell-type hierarchies and the knowledge-based ones. Based on the benchmarking results, we found that methods performed well in simulated scenarios but still have room for improvement in complex real-world data. Thus, we developed OTHarmonizer, a tool based on partial optimal transport (OT) for cell-type harmonization and hierarchy construction. OTHarmonizer excels in accurately capturing equivalent and hierarchical relationships between cell types, offering a more effective approach for the cell-type hierarchy construction across datasets. AVAILABILITY AND IMPLEMENTATION: The simulated and real-world datasets in the benchmark are available on https://figshare.com/articles/dataset/OTHarmonizer/28243205. The source codes for the benchmark and OTHarmonizer are available online on GitHub at https://github.com/Duck-Boss/OTHarmonizer. Tianhong Zhou, Yingtao Zhu, Jinmeng Jia, Xuegong Zhang, Lei Wei 0009 |
Bioinform. | 5 |
| 2025 | Multi-Modal Follow-Up Data-Guided Aggregated Representation for Predicting Gout Recurrence RiskabstractGout recurrence is common in real-world settings. While traditional machine learning methods are applicable, their performance is often limited by a lack of diverse data modalities, insufficient understanding of inter-modality interactions, and poor model generalizability. To address these challenges, this work proposes ARL-GRP, a novel framework for forecasting the risk of gout recurrence. This framework is built upon three essential modules: continuous learning utilising real-world multimodel follow-up data, feature representation aggregation employing a pretrained large encoder, and predicting recurrent gout risk using a multilayer perceptron. The experimental comparison demonstrates that our proposed approach generally outperforms conventional machine learning techniques. ARL-GRP can effectively combine structured clinical data and unstructured medical narratives into unified patient representations, significantly outperforming traditional machine learning methods (Accuracy: 0.931, AUC: 0.969). Our method demonstrates strong predictive capability, enabling precise risk assessment and personalised clinical decision-making. Furthermore, the effectiveness of our method is also consolidated through additional analysis using ROC curves and a heatmap. Baisong Li, Ruohan Liu, Xuegong Zhang, Hairong Lv |
BIBM | 5 |
| 2025 | DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?abstractVision–language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks do not systematically evaluate whether these models truly reason like human clinicians or merely imitate superficial patterns. To address this gap, we propose DrVD-Bench, the first multimodal benchmark for clinical visual reasoning. DrVD-Bench consists of three modules: Visual Evidence Comprehension, Reasoning Trajectory Assessment, and Report Generation Evaluation, comprising a total of 7,789 image–question pairs. Our benchmark covers 20 task types, 17 diagnostic categories, and five imaging modalities—CT, MRI, ultrasound, radiography, and pathology. DrVD-Bench is explicitly structured to reflect the clinical reasoning workflow from modality recognition to lesion identification and diagnosis. We benchmark 19 VLMs, including general-purpose and medical-specific, open-source and proprietary models, and observe that performance drops sharply as reasoning complexity increases. While some models begin to exhibit traces of human-like reasoning, they often still rely on shortcut correlations rather than grounded visual understanding. DrVD-Bench offers a rigorous and structured evaluation framework to guide the development of clinically trustworthy VLMs. Tianhong Zhou, Yingtao Zhu, Chuxi Xiao, Haiyang Bian, Lei Wei 0009, Xuegong Zhang |
NeurIPS | 7 |
| 2025 | Computational methods and data resources for predicting tumor neoantigensabstractNeoantigens are tumor-specific antigens presented exclusively by cancer cells. These antigens are recognized as nonself by the host immune system, thereby eliciting an antitumor T-cell response. This response is significantly enhanced through neoantigen-based immunotherapies, such as personalized cancer vaccines. The repertoire of neoantigens is unique to each cancer patient, necessitating neoantigen prediction for designing patient-specific immunotherapies. This review presents the computational methods and data resources used for neoantigen prediction, as well as the prediction-associated challenges. Neoantigen prediction typically uses human leukocyte antigen typing, RNA-seq transcript quantification, somatic variant calling, peptide-major histocompatibility complex (pMHC) presentation prediction, and pMHC recognition prediction as the main computational steps. The immunoinformatics tools used for these steps and for the overall prediction of neoantigens are systematically summarized and detailed in this review. Xiaofei Zhao 0005, Lei Wei 0009, Xuegong Zhang |
Briefings Bioinform. | 3 |
| 2025 | uHAF: a unified hierarchical annotation framework for cell type standardization and harmonizationabstractSUMMARY: In single-cell transcriptomics, inconsistent cell type annotations due to varied naming conventions and hierarchical granularity impede data integration, machine learning applications, and meaningful evaluations. To address this challenge, we developed the unified Hierarchical Annotation Framework (uHAF), which includes organ-specific hierarchical cell type trees (uHAF-T) and a mapping tool (uHAF-Agent) based on large language models. uHAF-T provides standardized hierarchical references for 38 organs, allowing for consistent label unification and analysis at different levels of granularity. uHAF-Agent leverages GPT-4 to accurately map diverse and informal cell type labels onto uHAF-T nodes, streamlining the harmonization process. By simplifying label unification, uHAF enhances data integration, supports machine learning applications, and enables biologically meaningful evaluations of annotation methods. Our framework serves as an essential resource for standardizing cell type annotations and fostering collaborative refinement in the single-cell research community. AVAILABILITY AND IMPLEMENTATION: uHAF is publicly available at: https://uhaf.unifiedcellatlas.org and https://github.com/SuperBianC/uhaf. Haiyang Bian, Yinxin Chen, Lei Wei 0009, Xuegong Zhang |
Bioinform. | 4 |
| 2025 | Label Informed Contrastive Pretraining for Node Importance Estimation on Knowledge GraphsabstractNode importance estimation (NIE) is the task of inferring the importance scores of the nodes in a graph. Due to the availability of richer data and knowledge, recent research interests of NIE have been dedicated to knowledge graphs (KGs) for predicting future or missing node importance scores. Existing state-of-the-art NIE methods train the model by available labels, and they consider every interested node equally before training. However, the nodes with higher importance often require or receive more attention in real-world scenarios, e.g., people may care more about the movies or webpages with higher importance. To this end, we introduce Label Informed ContrAstive Pretraining (LICAP) to the NIE problem for being better aware of the nodes with high importance scores. Specifically, LICAP is a novel type of contrastive learning (CL) framework that aims to fully utilize continuous labels to generate contrastive samples for pretraining embeddings. Considering the NIE problem, LICAP adopts a novel sampling strategy called top nodes preferred hierarchical sampling to first group all interested nodes into a top bin and a nontop bin based on node importance scores, and then divide the nodes within the top bin into several finer bins also based on the scores. The contrastive samples are generated from those bins and are then used to pretrain node embeddings of KGs via a newly proposed predicate-aware graph attention networks (PreGATs), so as to better separate the top nodes from nontop nodes, and distinguish the top nodes within the top bin by keeping the relative order among finer bins. Extensive experiments demonstrate that the LICAP pretrained embeddings can further boost the performance of existing NIE methods and achieve new state-of-the-art performance regarding both regression and ranking metrics. The source code for reproducibility is available at https://github.com/zhangtia16/LICAP. Chengbin Hou, Rui Jiang 0001, Xuegong Zhang, Chenghu Zhou, Ke Tang 0001, Hairong Lv |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Molecular Graph Representation Learning Integrating Large Language Models with Domain-specific Small ModelsabstractMolecular property prediction is a crucial foundation for drug discovery. In recent years, pre-trained deep learning models have been widely applied to this task. Some approaches that incorporate prior biological domain knowledge into the pre-training framework have achieved impressive results. However, these methods heavily rely on biochemical experts, and retrieving and summarizing vast amounts of domain knowledge literature is both time-consuming and expensive. Large Language Models (LLMs) have demonstrated remarkable performance in understanding and efficiently providing general knowledge. Nevertheless, they occasionally exhibit hallucinations and lack precision in generating domain-specific knowledge. Conversely, Domain-specific Small Models (DSMs) possess rich domain knowledge and can accurately calculate molecular domain-related metrics. However, due to their limited model size and singular functionality, they lack the breadth of knowledge necessary for comprehensive representation learning. To leverage the advantages of both approaches in molecular property prediction, we propose a novel Molecular Graph representation learning framework that integrates Large language models and Domain-specific small models (MolGraph-LarDo). Technically, we design a two-stage prompt strategy where DSMs are introduced to calibrate the knowledge provided by LLMs, enhancing the accuracy of domain-specific information and thus enabling LLMs to generate more precise textual descriptions for molecular samples. Subsequently, we employ a multi-modal alignment method to coordinate various modalities, including molecular graphs and their corresponding descriptive texts, to guide the pre-training of molecular representations. Extensive experiments demonstrate the effectiveness of the proposed method. Yuxiang Ren, Chengbin Hou, Hairong Lv, Xuegong Zhang |
BIBM | 5 |
| 2024 | scMulan: A Multitask Generative Pre-Trained Language Model for Single-Cell Analysis
Haiyang Bian, Xiaomin Dong, Chen Li 0001, Minsheng Hao, Jinyi Hu, Maosong Sun 0001, Lei Wei 0009, Xuegong Zhang |
RECOMB | 10 |
| 2024 | scDecouple: decoupling cellular response from infected proportion bias in scCRISPR-seqabstractSingle-cell clustered regularly interspaced short palindromic repeats-sequencing (scCRISPR-seq) is an emerging high-throughput CRISPR screening technology where the true cellular response to perturbation is coupled with infected proportion bias of guide RNAs (gRNAs) across different cell clusters. The mixing of these effects introduces noise into scCRISPR-seq data analysis and thus obstacles to relevant studies. We developed scDecouple to decouple true cellular response of perturbation from the influence of infected proportion bias. scDecouple first models the distribution of gene expression profiles in perturbed cells and then iteratively finds the maximum likelihood of cell cluster proportions as well as the cellular response for each gRNA. We demonstrated its performance in a series of simulation experiments. By applying scDecouple to real scCRISPR-seq data, we found that scDecouple enhances the identification of biologically perturbation-related genes. scDecouple can benefit scCRISPR-seq data analysis, especially in the case of heterogeneous samples or complex gRNA libraries. Qiuchen Meng, Lei Wei 0009, Joshua W. K. Ho, Yinqing Li, Xuegong Zhang |
Briefings Bioinform. | 8 |
| 2024 | Benchmarking multi-omics integration algorithms across single-cell RNA and ATAC dataabstractRecent advancements in single-cell sequencing technologies have generated extensive omics data in various modalities and revolutionized cell research, especially in the single-cell RNA and ATAC data. The joint analysis across scRNA-seq data and scATAC-seq data has paved the way to comprehending the cellular heterogeneity and complex cellular regulatory networks. Multi-omics integration is gaining attention as an important step in joint analysis, and the number of computational tools in this field is growing rapidly. In this paper, we benchmarked 12 multi-omics integration methods on three integration tasks via qualitative visualization and quantitative metrics, considering six main aspects that matter in multi-omics data analysis. Overall, we found that different methods have their own advantages on different aspects, while some methods outperformed other methods in most aspects. We therefore provided guidelines for selecting appropriate methods for specific scenarios and tasks to help obtain meaningful insights from multi-omics data integration. Chuxi Xiao, Qiuchen Meng, Lei Wei 0009, Xuegong Zhang |
Briefings Bioinform. | 5 |
| 2024 | scDiffusion: conditional generation of high-quality single-cell data using diffusion modelabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) data are important for studying the laws of life at single-cell level. However, it is still challenging to obtain enough high-quality scRNA-seq data. To mitigate the limited availability of data, generative models have been proposed to computationally generate synthetic scRNA-seq data. Nevertheless, the data generated with current models are not very realistic yet, especially when we need to generate data with controlled conditions. In the meantime, diffusion models have shown their power in generating data with high fidelity, providing a new opportunity for scRNA-seq generation. RESULTS: In this study, we developed scDiffusion, a generative model combining the diffusion model and foundation model to generate high-quality scRNA-seq data with controlled conditions. We designed multiple classifiers to guide the diffusion process simultaneously, enabling scDiffusion to generate data under multiple condition combinations. We also proposed a new control strategy called Gradient Interpolation. This strategy allows the model to generate continuous trajectories of cell development from a given cell state. Experiments showed that scDiffusion could generate single-cell gene expression data closely resembling real scRNA-seq data. Also, scDiffusion can conditionally produce data on specific cell types including rare cell types. Furthermore, we could use the multiple-condition generation of scDiffusion to generate cell type that was out of the training data. Leveraging the Gradient Interpolation strategy, we generated a continuous developmental trajectory of mouse embryonic cells. These experiments demonstrate that scDiffusion is a powerful tool for augmenting the real scRNA-seq data and can provide insights into cell fate research. AVAILABILITY AND IMPLEMENTATION: scDiffusion is openly available at the GitHub repository https://github.com/EperLuo/scDiffusion or Zenodo https://zenodo.org/doi/10.5281/zenodo.13268742. Erpai Luo, Minsheng Hao, Lei Wei 0009, Xuegong Zhang |
Bioinform. | 4 |
| 2024 | Weakly Supervised Causal Discovery Based on Fuzzy Knowledge and Complex Data ComplementarityabstractCausal discovery based on observational data is important for deciphering the causal mechanism behind complex systems. However, the effectiveness of existing causal discovery methods is limited due to inferior prior knowledge, domain inconsistencies, and the challenges of high-dimensional datasets with small sample sizes. To address this gap, we propose a novel weakly supervised fuzzy knowledge and data co-driven causal discovery method named KEEL. KEEL introduces a fuzzy causal knowledge schema to encapsulate diverse types of fuzzy knowledge, and forms corresponding weakened constraints. This schema not only lessens the dependency on expertise but also allows various types of limited and error-prone fuzzy knowledge to guide causal discovery. It can enhance the generalization and robustness of causal discovery, especially in high-dimensional and small-sample scenarios. In addition, we integrate the extended linear causal model into KEEL for dealing with the multi-distribution and incomplete data. Extensive experiments with different datasets demonstrate the superiority of KEEL over several state-of-the-art methods in accuracy, robustness and efficiency. The effectiveness of KEEL is also verified in limited real protein signal transduction process data, with the better performance than benchmark methods. In summary, KEEL is effective to tackle the causal discovery tasks with higher accuracy while alleviating the requirement for extensive domain expertise. Wei Zhang 0241, Qinghao Zhang, Xuegong Zhang, Xiaowo Wang |
IEEE Trans. Fuzzy Syst. | 4 |
| 2023 | xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq DataabstractAdvances in high-throughput sequencing technology have led to significant progress in measuring gene expressions at the single-cell level. The amount of publicly available single-cell RNA-seq (scRNA-seq) data is already surpassing 50M records for humans with each record measuring 20,000 genes. This highlights the need for unsupervised representation learning to fully ingest these data, yet classical transformer architectures are prohibitive to train on such data in terms of both computation and memory. To address this challenge, we propose a novel asymmetric encoder-decoder transformer for scRNA-seq data, called xTrimoGene$^\alpha$ (or xTrimoGene for short), which leverages the sparse characteristic of the data to scale up the pre-training. This scalable design of xTrimoGene reduces FLOPs by one to two orders of magnitude compared to classical transformers while maintaining high accuracy, enabling us to train the largest transformer models over the largest scRNA-seq dataset today. Our experiments also show that the performance of xTrimoGene improves as we scale up the model sizes, and it also leads to SOTA performance over various downstream tasks, such as cell type annotation, perturb-seq effect prediction, and drug combination prediction.
xTrimoGene model is now available for use as a service via the following link: https://api.biomap.com/xTrimoGene/apply. Minsheng Hao, Xingyi Cheng, Chiming Liu, Jianzhu Ma, Xuegong Zhang, Taifeng Wang |
NeurIPS | 7 |
| 2023 | Decoding functional cell-cell communication events by multi-view graph learning on spatial transcriptomicsabstractCell-cell communication events (CEs) are mediated by multiple ligand-receptor (LR) pairs. Usually only a particular subset of CEs directly works for a specific downstream response in a particular microenvironment. We name them as functional communication events (FCEs) of the target responses. Decoding FCE-target gene relations is: important for understanding the mechanisms of many biological processes, but has been intractable due to the mixing of multiple factors and the lack of direct observations. We developed a method HoloNet for decoding FCEs using spatial transcriptomic data by integrating LR pairs, cell-type spatial distribution and downstream gene expression into a deep learning model. We modeled CEs as a multi-view network, developed an attention-based graph learning method to train the model for generating target gene expression with the CE networks, and decoded the FCEs for specific downstream genes by interpreting trained models. We applied HoloNet on three Visium datasets of breast cancer and liver cancer. The results detangled the multiple factors of FCEs by revealing how LR signals and cell types affect specific biological processes, and specified FCE-induced effects in each single cell. We conducted simulation experiments and showed that HoloNet is more reliable on LR prioritization in comparison with existing methods. HoloNet is a powerful tool to illustrate cell-cell communication landscapes and reveal vital FCEs that shape cellular phenotypes. HoloNet is available as a Python package at https://github.com/lhc17/HoloNet. Haochen Li 0003, Tianxing Ma, Minsheng Hao, Wenbo Guo 0010, Jin Gu, Xuegong Zhang, Lei Wei 0009 |
Briefings Bioinform. | 6 |
| 2023 | simCAS: an embedding-based method for simulating single-cell chromatin accessibility sequencing dataabstractMOTIVATION: Single-cell chromatin accessibility sequencing (scCAS) technology provides an epigenomic perspective to characterize gene regulatory mechanisms at single-cell resolution. With an increasing number of computational methods proposed for analyzing scCAS data, a powerful simulation framework is desirable for evaluation and validation of these methods. However, existing simulators generate synthetic data by sampling reads from real data or mimicking existing cell states, which is inadequate to provide credible ground-truth labels for method evaluation. RESULTS: We present simCAS, an embedding-based simulator, for generating high-fidelity scCAS data from both cell- and peak-wise embeddings. We demonstrate simCAS outperforms existing simulators in resembling real data and show that simCAS can generate cells of different states with user-defined cell populations and differentiation trajectories. Additionally, simCAS can simulate data from different batches and encode user-specified interactions of chromatin regions in the synthetic data, which provides ground-truth labels more than cell states. We systematically demonstrate that simCAS facilitates the benchmarking of four core tasks in downstream analysis: cell clustering, trajectory inference, data integration, and cis-regulatory interaction inference. We anticipate simCAS will be a reliable and flexible simulator for evaluating the ongoing computational methods applied on scCAS data. AVAILABILITY AND IMPLEMENTATION: simCAS is freely available at https://github.com/Chen-Li-17/simCAS. Xiaoyang Chen 0007, Shengquan Chen, Rui Jiang 0001, Xuegong Zhang |
Bioinform. | 5 |
| 2022 | ARIC: accurate and robust inference of cell type proportions from bulk gene expression or DNA methylation dataabstractQuantifying cell proportions, especially for rare cell types in some scenarios, is of great value in tracking signals associated with certain phenotypes or diseases. Although some methods have been proposed to infer cell proportions from multicomponent bulk data, they are substantially less effective for estimating the proportions of rare cell types which are highly sensitive to feature outliers and collinearity. Here we proposed a new deconvolution algorithm named ARIC to estimate cell type proportions from gene expression or DNA methylation data. ARIC employs a novel two-step marker selection strategy, including collinear feature elimination based on the component-wise condition number and adaptive removal of outlier markers. This strategy can systematically obtain effective markers for weighted $\upsilon$-support vector regression to ensure a robust and precise rare proportion prediction. We showed that ARIC can accurately estimate fractions in both DNA methylation and gene expression data from different experiments. We further applied ARIC to the survival prediction of ovarian cancer and the condition monitoring of chronic kidney disease, and the results demonstrate the high accuracy and robustness as well as clinical potentials of ARIC. Taken together, ARIC is a promising tool to solve the deconvolution problem of bulk data where rare components are of vital importance. Wei Zhang 0241, Rong Qiao, Bixi Zhong, Xianglin Zhang, Jin Gu, Xuegong Zhang, Lei Wei 0009, Xiaowo Wang |
Briefings Bioinform. | 7 |
| 2022 | scGraph: a graph neural network-based approach to automatically identify cell typesabstractMOTIVATION: Single-cell technologies play a crucial role in revolutionizing biological research over the past decade, which strengthens our understanding in cell differentiation, development and regulation from a single-cell level perspective. Single-cell RNA sequencing (scRNA-seq) is one of the most common single cell technologies, which enables probing transcriptional states in thousands of cells in one experiment. Identification of cell types from scRNA-seq measurements is a fundamental and crucial question to answer. Most previous studies directly take gene expression as input while ignoring the comprehensive gene-gene interactions. RESULTS: We propose scGraph, an automatic cell identification algorithm leveraging gene interaction relationships to enhance the performance of the cell-type identification. scGraph is based on a graph neural network to aggregate the information of interacting genes. In a series of experiments, we demonstrate that scGraph is accurate and outperforms eight comparison methods in the task of cell-type identification. Moreover, scGraph automatically learns the gene interaction relationships from biological data and the pathway enrichment analysis shows consistent findings with previous analysis, providing insights on the analysis of regulatory mechanism. AVAILABILITY AND IMPLEMENTATION: scGraph is freely available at https://github.com/QijinYin/scGraph and https://figshare.com/articles/software/scGraph/17157743. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qijin Yin, Qiao Liu 0008, Zhuoran Fu, Wanwen Zeng, Boheng Zhang, Xuegong Zhang, Rui Jiang 0001, Hairong Lv |
Bioinform. | 6 |
| 2022 | DualGCN: a dual graph convolutional network model to predict cancer drug responseabstractAbstract Background Drug resistance is a critical obstacle in cancer therapy. Discovering cancer drug response is important to improve anti-cancer drug treatment and guide anti-cancer drug design. Abundant genomic and drug response resources of cancer cell lines provide unprecedented opportunities for such study. However, cancer cell lines cannot fully reflect heterogeneous tumor microenvironments. Transferring knowledge studied from in vitro cell lines to single-cell and clinical data will be a promising direction to better understand drug resistance. Most current studies include single nucleotide variants (SNV) as features and focus on improving predictive ability of cancer drug response on cell lines. However, obtaining accurate SNVs from clinical tumor samples and single-cell data is not reliable. This makes it difficult to generalize such SNV-based models to clinical tumor data or single-cell level studies in the future. Results We present a new method, DualGCN, a unified Dual Graph Convolutional Network model to predict cancer drug response. DualGCN encodes both chemical structures of drugs and omics data of biological samples using graph convolutional networks. Then the two embeddings are fed into a multilayer perceptron to predict drug response. DualGCN incorporates prior knowledge on cancer-related genes and protein–protein interactions, and outperforms most state-of-the-art methods while avoiding using large-scale SNV data. Conclusions The proposed method outperforms most state-of-the-art methods in predicting cancer drug response without the use of large-scale SNV data. These favorable results indicate its potential to be extended to clinical and single-cell tumor samples and advancements in precision medicine. Tianxing Ma, Qiao Liu 0008, Haochen Li 0003, Mu Zhou, Rui Jiang 0001, Xuegong Zhang |
BMC Bioinform. | 6 |
| 2022 | CS-CO: A Hybrid Self-Supervised Visual Representation Learning Method for H&E-stained Histopathological Images
Pengshuai Yang, Xiaoxu Yin, Haiming Lu, Zhongliang Hu, Xuegong Zhang, Rui Jiang 0001, Hairong Lv |
Medical Image Anal. | 5 |
| 2022 | AggEnhance: Aggregation Enhancement by Class Interior Points in Federated Learning with Non-IID DataabstractFederated learning (FL) is a privacy-preserving paradigm for multi-institutional collaborations, where the aggregation is an essential procedure after training on the local datasets. Conventional aggregation algorithms often apply a weighted averaging of the updates generated from distributed machines to update the global model. However, while the data distributions are non-IID, the large discrepancy between the local updates might lead to a poor averaged result and a lower convergence speed, i.e., more iterations required to achieve a certain performance. To solve this problem, this article proposes a novel method named AggEnhance for enhancing the aggregation, where we synthesize a group of reliable samples from the local models and tune the aggregated result on them. These samples, named class interior points (CIPs) in this work, bound the relevant decision boundaries that ensure the performance of aggregated result. To the best of our knowledge, this is the first work to explicitly design an enhancing method for the aggregation in prevailing FL pipelines. A series of experiments on real data demonstrate that our method has noticeable improvements of the convergence in non-IID scenarios. In particular, our approach reduces the iterations by 31.87% on average for the CIFAR10 dataset and 43.90% for the PASCAL VOC dataset. Since our method does not modify other procedures of FL pipelines, it is easy to apply to most existing FL frameworks. Furthermore, it does not require additional data transmitted from the local clients to the global server, thus holding the same security level as the original FL algorithms. Jinxiang Ou, Yunheng Shen, Feng Wang 0047, Qiao Liu 0008, Xuegong Zhang, Hairong Lv |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2021 | stPlus: a reference-based method for the accurate enhancement of spatial transcriptomicsabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) techniques have revolutionized the investigation of transcriptomic landscape in individual cells. Recent advancements in spatial transcriptomic technologies further enable gene expression profiling and spatial organization mapping of cells simultaneously. Among the technologies, imaging-based methods can offer higher spatial resolutions, while they are limited by either the small number of genes imaged or the low gene detection sensitivity. Although several methods have been proposed for enhancing spatially resolved transcriptomics, inadequate accuracy of gene expression prediction and insufficient ability of cell-population identification still impede the applications of these methods. RESULTS: We propose stPlus, a reference-based method that leverages information in scRNA-seq data to enhance spatial transcriptomics. Based on an auto-encoder with a carefully tailored loss function, stPlus performs joint embedding and predicts spatial gene expression via a weighted k-nearest-neighbor. stPlus outperforms baseline methods with higher gene-wise and cell-wise Spearman correlation coefficients. We also introduce a clustering-based approach to assess the enhancement performance systematically. Using the data enhanced by stPlus, cell populations can be better identified than using the measured data. The predicted expression of genes unique to scRNA-seq data can also well characterize spatial cell heterogeneity. Besides, stPlus is robust and scalable to datasets of diverse gene detection sensitivity levels, sample sizes and number of spatially measured genes. We anticipate stPlus will facilitate the analysis of spatial transcriptomics. AVAILABILITY AND IMPLEMENTATION: stPlus with detailed documents is freely accessible at http://health.tsinghua.edu.cn/software/stPlus/ and the source code is openly available on https://github.com/xy-chen16/stPlus. Shengquan Chen, Boheng Zhang, Xiaoyang Chen 0007, Xuegong Zhang, Rui Jiang 0001 |
Bioinform. | 4 |
| 2021 | SOMDE: a scalable method for identifying spatially variable genes with self-organizing mapabstractMOTIVATION: Recent developments of spatial transcriptomic sequencing technologies provide powerful tools for understanding cells in the physical context of tissue microenvironments. A fundamental task in spatial gene expression analysis is to identify genes with spatially variable expression patterns, or spatially variable genes (SVgenes). Several computational methods have been developed for this task. Their high computational complexity limited their scalability to the latest and future large-scale spatial expression data. RESULTS: We present SOMDE, an efficient method for identifying SVgenes in large-scale spatial expression data. SOMDE uses self-organizing map to cluster neighboring cells into nodes, and then uses a Gaussian process to fit the node-level spatial gene expression to identify SVgenes. Experiments show that SOMDE is about 5-50 times faster than existing methods with comparable results. The adjustable resolution of SOMDE makes it the only method that can give results in ∼5 min in large datasets of more than 20 000 sequencing sites. SOMDE is available as a python package on PyPI at https://pypi.org/project/somde free for academic use. AVAILABILITY AND IMPLEMENTATION: SOMDE is available for download from PyPI, and the source code is openly available from the Github repository https://github.com/XuegongLab/somde. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Minsheng Hao, Kui Hua, Xuegong Zhang |
Bioinform. | 3 |
| 2021 | CellTracker: an automated toolbox for single-cell segmentation and tracking of time-lapse microscopy imagesabstractSUMMARY: Recent advances of long-term time-lapse microscopy have made it easy for researchers to quantify cell behavior and molecular dynamics at single-cell resolution. However, the lack of easy-to-use software tools optimized for customized research is still a major challenge for quantitatively understanding biological processes through microscopy images. Here, we present CellTracker, a highly integrated graphical user interface software, for automated cell segmentation and tracking of time-lapse microscopy images. It covers essential steps in image analysis including project management, image pre-processing, cell segmentation, cell tracking, manually correction and statistical analysis such as the quantification of cell size and fluorescence intensity, etc. Furthermore, CellTracker provides an annotation tool and supports model training from scratch, thus proposing a flexible and scalable solution for customized dataset analysis. AVAILABILITY AND IMPLEMENTATION: CellTracker is an open-source software under the GPL-3.0 license. It is implemented in Python and provides an easy-to-use graphical user interface. The source code, instruction manual and demos can be found at https://github.com/WangLabTHU/CellTracker. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tao Hu 0020, Shixiong Xu, Lei Wei 0009, Xuegong Zhang, Xiaowo Wang |
Bioinform. | 4 |
| 2021 | HGC: fast hierarchical clustering for large-scale single-cell dataabstractSUMMARY: Clustering is a key step in revealing heterogeneities in single-cell data. Most existing single-cell clustering methods output a fixed number of clusters without the hierarchical information. Classical hierarchical clustering (HC) provides dendrograms of cells, but cannot scale to large datasets due to high computational complexity. We present HGC, a fast Hierarchical Graph-based Clustering tool to address both problems. It combines the advantages of graph-based clustering and HC. On the shared nearest-neighbor graph of cells, HGC constructs the hierarchical tree with linear time complexity. Experiments showed that HGC enables multiresolution exploration of the biological hierarchy underlying the data, achieves state-of-the-art accuracy on benchmark data and can scale to large datasets. AVAILABILITY AND IMPLEMENTATION: The R package of HGC is available at https://bioconductor.org/packages/HGC/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ziheng Zou, Kui Hua, Xuegong Zhang |
Bioinform. | 3 |
| 2021 | A Method for Generating Synthetic Electronic Medical Record TextabstractMachine learning (ML) and Natural Language Processing (NLP) have achieved remarkable success in many fields and have brought new opportunities and high expectation in the analyses of medical data, of which the most common type is the massive free-text electronic medical records (EMR). However, the free EMR texts are lacking consistent standards, rich of private information, and limited in availability. Also, it is often hard to have a balanced number of samples for the types of diseases under study. These problems hinder the development of ML and NLP methods for EMR data analysis. To tackle these problems, we developed a model called Medical Text Generative Adversarial Network or mtGAN, to generate synthetic EMR text. It is based on the GAN framework and is trained by the REINFORCE algorithm. It takes disease tags as inputs and generates synthetic texts as EMRs for the corresponding diseases. We evaluate the model from micro-level, macro-level and application-level on a Chinese EMR text dataset. The results show that the method has a good capacity to fit real data and can generate realistic and diverse EMR samples. This provides a novel way to avoid potential leakage of patient privacy while still supply sufficient well-controlled cohort data for developing downstream ML and NLP methods. Jiaqi Guan, Runzhe Li, Sheng Yu 0002, Xuegong Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | SCeQTL: an R package for identifying eQTL from single-cell parallel sequencing dataabstractBACKGROUND: With the rapid development of single-cell genomics, technologies for parallel sequencing of the transcriptome and genome in each single cell is being explored in several labs and is becoming available. This brings us the opportunity to uncover association between genotypes and gene expression phenotypes at single-cell level by eQTL analysis on single-cell data. New method is needed for such tasks due to special characteristics of single-cell sequencing data. RESULTS: We developed an R package SCeQTL that uses zero-inflated negative binomial regression to do eQTL analysis on single-cell data. It can distinguish two type of gene-expression differences among different genotype groups. It can also be used for finding gene expression variations associated with other grouping factors like cell lineages or cell types. CONCLUSIONS: The SCeQTL method is capable for eQTL analysis on single-cell data as well as detecting associations of gene expression with other grouping factors. The R package of the method is available at https://github.com/XuegongLab/SCeQTL/. Xuegong Zhang |
BMC Bioinform. | 4 |
| 2019 | Pixel-Level Clustering Reveals Intra-Tumor Heterogeneity in Non-Small Cell Lung CancerabstractIntra-tumor heterogeneity in non-small cell lung cancer (NSCLC) is a current topic of many genomic studies, but there is no effective technology for characterizing tumor heterogeneity before surgery or biopsy. Since computed tomography (CT) is a standard technology for the diagnosis of lung cancer, it will be of great scientific and clinical use if tumor heterogeneity information can be extracted from CT images. In this work, we proposed a strategy for studying intra-tumor heterogeneity using pixel-level multi-resolution clustering on radiomics features of CT images. On a set of pretreatment CT images of NSCLC patients, we observed three high-order image patterns “uni-core”, “multi-core” and “diffused” that can be associated with characteristics of intra-tumor heterogeneity. We checked the clinical records of the two groups of patterns: “uni-core” and “diffused” patterns that have enough samples, and found significant difference in their survival time. The results showed the potential of using radiomic analysis on CT images to characterize intra-tumor heterogeneity and to predict patient prognosis based on high-order image patterns. Haiming Lu, Xuegong Zhang |
BIBM | 5 |
| 2019 | A new statistic for efficient detection of repetitive sequencesabstractMOTIVATION: Detecting sequences containing repetitive regions is a basic bioinformatics task with many applications. Several methods have been developed for various types of repeat detection tasks. An efficient generic method for detecting most types of repetitive sequences is still desirable. Inspired by the excellent properties and successful applications of the D2 family of statistics in comparative analyses of genomic sequences, we developed a new statistic D2R that can efficiently discriminate sequences with or without repetitive regions. RESULTS: Using the statistic, we developed an algorithm of linear time and space complexity for detecting most types of repetitive sequences in multiple scenarios, including finding candidate clustered regularly interspaced short palindromic repeats regions from bacterial genomic or metagenomics sequences. Simulation and real data experiments show that the method works well on both assembled sequences and unassembled short reads. AVAILABILITY AND IMPLEMENTATION: The codes are available at https://github.com/XuegongLab/D2R_codes under GPL 3.0 license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fengzhu Sun, Michael S. Waterman, Xuegong Zhang |
Bioinform. | 5 |
| 2018 | Generation of Synthetic Electronic Medical Record Text
Jiaqi Guan, Runzhe Li, Sheng Yu 0002, Xuegong Zhang |
BIBM | 4 |
| 2018 | An Overview of Bioinformatics Challenges for Human Cell Atlas
Xuegong Zhang |
BIBM | 1 |
| 2018 | DEsingle for detecting three types of differential expression in single-cell RNA-seq dataabstractSummary: The excessive amount of zeros in single-cell RNA-seq (scRNA-seq) data includes 'real' zeros due to the on-off nature of gene transcription in single cells and 'dropout' zeros due to technical reasons. Existing differential expression (DE) analysis methods cannot distinguish these two types of zeros. We developed an R package DEsingle which employed Zero-Inflated Negative Binomial model to estimate the proportion of real and dropout zeros and to define and detect three types of DE genes in scRNA-seq data with higher accuracy. Availability and implementation: The R package DEsingle is freely available at Bioconductor (https://bioconductor.org/packages/DEsingle). Supplementary information: Supplementary data are available at Bioinformatics online. Zhun Miao, Xiaowo Wang, Xuegong Zhang |
Bioinform. | 4 |
| 2017 | Reading the Underlying Information From Massive Metagenomic Sequencing DataabstractMicroorganisms are everywhere. Recent studies showed that the mixture of microbes or the microbiome on the human body plays important roles in human physiology and diseases. Metagenomic sequencing is a key technology for studying microbiomes. It produces massive amounts of data in the form of short sequencing reads. A single metagenomic sample can contain 107to 108reads of about 100-nucleotide (nt) length each in a typical shotgun metagenomic sequencing study. They contain rich information about microbiomes and their functions, but reading out those information from the huge highly fragmented data has multiple challenges for mathematical models, bioinformatics methods, and computer algorithms. In this paper, we review the basic bioinformatics tasks and existing methods in processing and analyzing metagenomic data, and discuss remaining open challenges and practical observations. The aim of the paper is to provide readers a whole picture of metagenomic data processing and analysis, and a reference and perspective to start with for computational scientists who are interested in this exciting field. Xuegong Zhang, Shansong Liu, Hongfei Cui, Ting Chen 0006 |
Proc. IEEE | 1 |
| 2015 | dslice: an R package for nonparametric testing of associations with application in QTL and gene set analysisabstractUNLABELLED: Many statistical problems in bioinformatics and genetics can be formulated as the testing of associations between a categorical variable and a continuous variable. A dynamic slicing method was proposed for non-parametric dependence testing, which has been demonstrated to have higher powers compared with traditional methods such as Kolmogorov-Smirnov test. We introduce an R package dslice to facilitate the use of dynamic slicing method in bioinformatic applications such as quantitative trait loci study and gene set enrichment analysis. AVAILABILITY AND IMPLEMENTATION: dslice is implemented in Rcpp and available in the Comprehensive R Archive Network. The package is distributed under the GNU General Public License (version 2 or later). Bo Jiang 0005, Xuegong Zhang, Jun S. Liu |
Bioinform. | 3 |
| 2014 | RNAseqViewer: visualization tool for RNA-Seq dataabstractSUMMARY: With the advances of RNA sequencing technologies, scientists need new tools to analyze transcriptome data. We introduce RNAseqViewer, a new visualization tool dedicated to RNA-Seq data. The program offers innovative ways to represent transcriptome data for single or multiple samples. It is a handy tool for scientists who use RNA-Seq data to compare multiple transcriptomes, for example, to compare gene expression and alternative splicing of cancer samples or of different development stages. AVAILABILITY AND IMPLEMENTATION: RNAseqViewer is freely available for academic use at http://bioinfo.au.tsinghua.edu.cn/software/RNAseqViewer/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xavier Rogé, Xuegong Zhang |
Bioinform. | 2 |
| 2013 | NURD: an implementation of a new method to estimate isoform expression from non-uniform RNA-seq dataabstractBACKGROUND: RNA-Seq technology has been used widely in transcriptome study, and one of the most important applications is to estimate the expression level of genes and their alternative splicing isoforms. There have been several algorithms published to estimate the expression based on different models. Recently Wu et al. published a method that can accurately estimate isoform level expression by considering position-related sequencing biases using nonparametric models. The method has advantages in handling different read distributions, but there hasn't been an efficient program to implement this algorithm. RESULTS: We developed an efficient implementation of the algorithm in the program NURD. It uses a binary interval search algorithm. The program can correct both the global tendency of sequencing bias in the data and local sequencing bias specific to each gene. The correction makes the isoform expression estimation more reliable under various read distributions. And the implementation is computationally efficient in both the memory cost and running time and can be readily scaled up for huge datasets. CONCLUSION: NURD is an efficient and reliable tool for estimating the isoform expression level. Given the reads mapping result and gene annotation file, NURD will output the expression estimation result. The package is freely available for academic use at http://bioinfo.au.tsinghua.edu.cn/software/NURD/. Xinyun Ma, Xuegong Zhang |
BMC Bioinform. | 2 |
| 2013 | Network-based differential gene expression analysis suggests cell cycle related genes regulated by E2F1 underlie the molecular difference between smoker and non-smoker lung adenocarcinomaabstractBACKGROUND: Differential gene expression (DGE) analysis is commonly used to reveal the deregulated molecular mechanisms of complex diseases. However, traditional DGE analysis (e.g., the t test or the rank sum test) tests each gene independently without considering interactions between them. Top-ranked differentially regulated genes prioritized by the analysis may not directly relate to the coherent molecular changes underlying complex diseases. Joint analyses of co-expression and DGE have been applied to reveal the deregulated molecular modules underlying complex diseases. Most of these methods consist of separate steps: first to identify gene-gene relationships under the studied phenotype then to integrate them with gene expression changes for prioritizing signature genes, or vice versa. It is warrant a method that can simultaneously consider gene-gene co-expression strength and corresponding expression level changes so that both types of information can be leveraged optimally. RESULTS: In this paper, we develop a gene module based method for differential gene expression analysis, named network-based differential gene expression (nDGE) analysis, a one-step integrative process for prioritizing deregulated genes and grouping them into gene modules. We demonstrate that nDGE outperforms existing methods in prioritizing deregulated genes and discovering deregulated gene modules using simulated data sets. When tested on a series of smoker and non-smoker lung adenocarcinoma data sets, we show that top differentially regulated genes identified by the rank sum test in different sets are not consistent while top ranked genes defined by nDGE in different data sets significantly overlap. nDGE results suggest that a differentially regulated gene module, which is enriched for cell cycle related genes and E2F1 targeted genes, plays a role in the molecular differences between smoker and non-smoker lung adenocarcinoma. CONCLUSIONS: In this paper, we develop nDGE to prioritize deregulated genes and group them into gene modules by simultaneously considering gene expression level changes and gene-gene co-regulations. When applied to both simulated and empirical data, nDGE outperforms the traditional DGE method. More specifically, when applied to smoker and non-smoker lung cancer sets, nDGE results illustrate the molecular differences between smoker and non-smoker lung cancer. Xuegong Zhang |
BMC Bioinform. | 3 |
| 2013 | Improving cis-regulatory elements modeling by consensus scaffolded mixture models
Hongshan Jiang, Xuegong Zhang |
Sci. China Inf. Sci. | 5 |
| 2013 | Detecting DNA Modifications from SMRT Sequencing Data by Modeling Sequence Context Dependence of Polymerase KineticabstractDNA modifications such as methylation and DNA damage can play critical regulatory roles in biological systems. Single molecule, real time (SMRT) sequencing technology generates DNA sequences as well as DNA polymerase kinetic information that can be used for the direct detection of DNA modifications. We demonstrate that local sequence context has a strong impact on DNA polymerase kinetics in the neighborhood of the incorporation site during the DNA synthesis reaction, allowing for the possibility of estimating the expected kinetic rate of the enzyme at the incorporation site using kinetic rate information collected from existing SMRT sequencing data (historical data) covering the same local sequence contexts of interest. We develop an Empirical Bayesian hierarchical model for incorporating historical data. Our results show that the model could greatly increase DNA modification detection accuracy, and reduce requirement of control data coverage. For some DNA modifications that have a strong signal, a control sample is not even needed by using historical data as alternative to control. Thus, sequencing costs can be greatly reduced by using the model. We implemented the model in a R package named seqPatch, which is available at https://github.com/zhixingfeng/seqPatch. Zhixing Feng, Gang Fang 0004, Jonas Korlach, Tyson Clark, Khai Luong, Xuegong Zhang, Wing Wong, Eric E. Schadt |
PLoS Comput. Biol. | 6 |
| 2012 | Integrating gene expression and protein-protein interaction network to prioritize cancer-associated genesabstractBACKGROUND: To understand the roles they play in complex diseases, genes need to be investigated in the networks they are involved in. Integration of gene expression and network data is a promising approach to prioritize disease-associated genes. Some methods have been developed in this field, but the problem is still far from being solved. RESULTS: In this paper, we developed a method, Networked Gene Prioritizer (NGP), to prioritize cancer-associated genes. Applications on several breast cancer and lung cancer datasets demonstrated that NGP performs better than the existing methods. It provides stable top ranking genes between independent datasets. The top-ranked genes by NGP are enriched in the cancer-associated pathways. The top-ranked genes by NGP-PLK1, MCM2, MCM3, MCM7, MCM10 and SKP2 might coordinate to promote cell cycle related processes in cancer but not normal cells. CONCLUSIONS: In this paper, we have developed a method named NGP, to prioritize cancer-associated genes. Our results demonstrated that NGP performs better than the existing methods. Xuegong Zhang |
BMC Bioinform. | 3 |
| 2011 | Using non-uniform read distribution models to improve isoform expression inference in RNA-SeqabstractMOTIVATION: RNA-Seq technology based on next-generation sequencing provides the unprecedented ability of studying transcriptomes at high resolution and accuracy, and the potential of measuring expression of multiple isoforms from the same gene at high precision. Solved by maximum likelihood estimation, isoform expression can be inferred in RNA-Seq using statistical models based on the assumption that sequenced reads are distributed uniformly along transcripts. Modification of the model is needed when considering situations where RNA-Seq data do not follow uniform distribution. RESULTS: We proposed two curves, the global bias curve (GBC) and the local bias curves (LBCs), to describe the non-uniformity of read distributions for all genes in a transcriptome and for each gene, respectively. Incorporating the bias curves into the uniform read distribution (URD) model, we introduced non-URD (N-URD) models to infer isoform expression levels. On a series of systematic simulation studies, the proposed models outperform the original model in recovering major isoforms and the expression ratio of alternative isoforms. We also applied the new model to real RNA-Seq datasets and found that its inferences on expression ratios of alternative isoforms are more reasonable. The experiments indicate that incorporating N-URD information can improve the accuracy in modeling and inferring isoform expression in RNA-Seq. Zhengpeng Wu, Xi Wang 0002, Xuegong Zhang |
Bioinform. | 3 |
| 2011 | Network-based group variable selection for detecting expression quantitative trait loci (eQTL)abstractBACKGROUND: Analysis of expression quantitative trait loci (eQTL) aims to identify the genetic loci associated with the expression level of genes. Penalized regression with a proper penalty is suitable for the high-dimensional biological data. Its performance should be enhanced when we incorporate biological knowledge of gene expression network and linkage disequilibrium (LD) structure between loci in high-noise background. RESULTS: We propose a network-based group variable selection (NGVS) method for QTL detection. Our method simultaneously maps highly correlated expression traits sharing the same biological function to marker sets formed by LD. By grouping markers, complex joint activity of multiple SNPs can be considered and the dimensionality of eQTL problem is reduced dramatically. In order to demonstrate the power and flexibility of our method, we used it to analyze two simulations and a mouse obesity and diabetes dataset. We considered the gene co-expression network, grouped markers into marker sets and treated the additive and dominant effect of each locus as a group: as a consequence, we were able to replicate results previously obtained on the mouse linkage dataset. Furthermore, we observed several possible sex-dependent loci and interactions of multiple SNPs. CONCLUSIONS: The proposed NGVS method is appropriate for problems with high-dimensional data and high-noise background. On eQTL problem it outperforms the classical Lasso method, which does not consider biological knowledge. Introduction of proper gene expression and loci correlation information makes detecting causal markers more accurate. With reasonable model settings, NGVS can lead to novel biological findings. Xuegong Zhang |
BMC Bioinform. | 2 |
| 2010 | DEGseq: an R package for identifying differentially expressed genes from RNA-seq dataabstractAbstract Summary: High-throughput RNA sequencing (RNA-seq) is rapidly emerging as a major quantitative transcriptome profiling platform. Here, we present DEGseq, an R package to identify differentially expressed genes or isoforms for RNA-seq data from different samples. In this package, we integrated three existing methods, and introduced two novel methods based on MA-plot to detect and visualize gene expression difference. Availability: The R package and a quick-start vignette is available at http://bioinfo.au.tsinghua.edu.cn/software/degseq Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Likun Wang 0003, Zhixing Feng, Xi Wang 0002, Xiaowo Wang, Xuegong Zhang |
Bioinform. | 5 |
| 2009 | The Seventh Asia Pacific Bioinformatics Conference (APBC2009)abstractThe Asia Pacific Bioinformatics Conference (APBC) series, founded in 2003, is an annual international forum for exploring research, development and applications of Bioinformatics and Computational Biology.The Seventh Asia Pacific Bioinformatics Conference (APBC2009) was held at Tsinghua University, Michael Q. Zhang, Michael S. Waterman, Xuegong Zhang |
BMC Bioinform. | 3 |
| 2008 | Multiobjective fuzzy biclustering in microarray data: Method and a new performance measureabstractObjective of any biclustering algorithm in microarray data is to discover a subset of genes that are expressed similarly in a subset of conditions. The boundaries of biclusters usually overlap as genes and conditions may belong to different biclusters with different membership degrees. Hence the notion of fuzzy sets is useful for discovering such overlapping biclusters. In this article an attempt has been made to develop a multiobjective genetic algorithm based approach for probabilistic fuzzy biclustering that minimizes the residual and maximizes cluster size and expression profile variance. A novel variable string length encoding has been proposed in this regard that encodes multiple biclusters in a single string. Also a new performance measure that reflects how a bicluster is statistically distinguished from the background is proposed. Performance of the proposed algorithm has been compared with some well known biclustering algorithms. Ujjwal Maulik, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Michael Q. Zhang, Xuegong Zhang |
IEEE Congress on Evolutionary Computation | 5 |
| 2008 | Estimating the Confidence Interval for Prediction Errors of Support Vector Machine Classifiers
Bo Jiang 0005, Xuegong Zhang, Tianxi Cai |
J. Mach. Learn. Res. | 2 |
| 2007 | Prediction of Kinase-Specific Phosphorylation Sites by One-Class SVMsabstractProtein phosphorylation is one of the most important post-translational modifications. Detecting possible phosphorylation sites and their corresponding protein kinases is crucial for studying the function of proteins. This task has been formulated as a machine learning problem of discriminating phosphorylation sites from non-phosphorylation sites with features of amino acid sequences. However, as phosphorylation is a dynamic event, it is hard to collect a set of protein sequences which can be safely regarded as non- phosphorylatable. Here we present a new prediction system, called OCSPP, to predict kinase-specific phosphorylation sites according to peptide sequences using one-class support vector machines. This method only needs positive training samples. Experiments on the published datasets show that both the specificity and sensitivity of the method can reach higher than 90% at kinase-family level for some kinase families. Comparison results show the superiority of OCSPP over Scansite, both of which only use positive samples in prediction. OCSPP is available at http://bioinfo.au.tsinghua.edu.cn/OCSPP/. Xuegong Zhang |
BIBM | 3 |
| 2007 | Integrative data mining in systems biology: from text to network mining
Yonghong Peng, Xuegong Zhang |
Artif. Intell. Medicine | 2 |
| 2007 | OSCAR: One-class SVM for accurate recognition of cis-elementsabstractMOTIVATION: Traditional methods to identify potential binding sites of known transcription factors still suffer from large number of false predictions. They mostly use sequence information in a position-specific manner and neglect other types of information hidden in the proximal promoter regions. Recent biological and computational researches, however, suggest that there exist not only locational preferences of binding, but also correlations between transcription factors. RESULTS: In this article, we propose a novel approach, OSCAR, which utilizes one-class SVM algorithms, and incorporates multiple factors to aid the recognition of transcription factor binding sites. Using both synthetic and real data, we find that our method outperforms existing algorithms, especially in the high sensitivity region. The performance of our method can be further improved by taking into account locational preference of binding events. By testing on experimentally-verified binding sites of GATA and HNF transcription factor families, we show that our algorithm can infer the true co-occurring motif pairs accurately, and by considering the co-occurrences of correlated motifs, we not only filter out false predictions, but also increase the sensitivity. AVAILABILITY: An online server based on OSCAR is available at http://bioinfo.au.tsinghua.edu.cn/oscar. Bo Jiang 0005, Michael Q. Zhang, Xuegong Zhang |
Bioinform. | 3 |
| 2007 | Computing exact P-values for DNA motifsabstractMOTIVATION: Many heuristic algorithms have been designed to approximate P-values of DNA motifs described by position weight matrices, for evaluating their statistical significance. They often significantly deviate from the true P-value by orders of magnitude. Exact P-value computation is needed for ranking the motifs. Furthermore, surprisingly, the complexity of the problem is unknown. RESULTS: We show the problem to be NP-hard, and present MotifRank, software based on dynamic programming, to calculate exact P-values of motifs. We define the exact P-value on a general and more precise model. Asymptotically, MotifRank is faster than the best exact P-value computing algorithm, and is in fact practical. Our experiments clearly demonstrate that MotifRank significantly improves the accuracy of existing approximation algorithms. AVAILABILITY: MotifRank is available from http://bio.dlg.cn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jing Zhang 0012, Bo Jiang 0005, Ming Li 0001, John Tromp, Xuegong Zhang, Michael Q. Zhang |
Bioinform. | 5 |
| 2007 | Identifications of conserved 7-mers in 3'-UTRs and microRNAs in DrosophilaabstractBACKGROUND: MicroRNAs (miRNAs) are a class of endogenous regulatory small RNAs which play an important role in posttranscriptional regulations by targeting mRNAs for cleavage or translational repression. The base-pairing between the 5'-end of miRNA and the target mRNA 3'-UTRs is essential for the miRNA:mRNA recognition. Recent studies show that many seed matches in 3'-UTRs, which are fully complementary to miRNA 5'-ends, are highly conserved. Based on these features, a two-stage strategy can be implemented to achieve the de novo identification of miRNAs by requiring the complete base-pairing between the 5'-end of miRNA candidates and the potential seed matches in 3'-UTRs. RESULTS: We presented a new method, which combined multiple pairwise conservation information, to identify the frequently-occurred and conserved 7-mers in 3'-UTRs. A pairwise conservation score (PCS) was introduced to describe the conservation of all 7-mers in 3'-UTRs between any two Drosophila species. Using PCSs computed from 6 pairs of flies, we developed a support vector machine (SVM) classifier ensemble, named Cons-SVM and identified 689 conserved 7-mers including 63 seed matches covering 32 out of 38 known miRNA families in the reference dataset. In the second stage, we searched for 90 nt conserved stem-loop regions containing the complementary sequences to the identified 7-mers and used the previously published miRNA prediction software to analyze these stem-loops. We predicted 47 miRNA candidates in the genome-wide screen. CONCLUSION: Cons-SVM takes advantage of the independent evolutionary information from the 6 pairs of flies and shows high sensitivity in identifying seed matches in 3'-UTRs. Combining the multiple pairwise conservation information by the machine learning approach, we finally identified 47 miRNA candidates in D. melanogaster. Jin Gu, Hu Fu 0001, Xuegong Zhang, Yanda Li |
BMC Bioinform. | 3 |
| 2007 | Neighbor number, valley seeking and clustering
Xuegong Zhang, Michael Q. Zhang, Yanda Li |
Pattern Recognit. Lett. | 2 |
| 2006 | Embryonics: A Path to Artificial Life?abstractElectronic systems, no matter how clever and intelligent they are, cannot yet demonstrate the reliability that biological systems can. Perhaps we can learn from these processes, which have developed through millions of years of evolution, in our pursuit of highly reliable systems. This article discusses how such systems, inspired by biological principles, might be built using simple embryonic cells. We illustrate how they can monitor their own functional integrity in order to protect themselves from internal failure or from hostile environmental effects and how faults caused by DNA mutation or cell death can be repaired and thus full system functionality restored. Xuegong Zhang, Gabriel Dragffy, Anthony G. Pipe |
Artif. Life | 1 |
| 2006 | Predicting methylation status of CpG islands in the human brainabstractMOTIVATION: Over 50% of human genes contain CpG islands in their 5'-regions. Methylation patterns of CpG islands are involved in tissue-specific gene expression and regulation. Mis-epigenetic silencing associated with aberrant CpG island methylation is one mechanism leading to the loss of tumor suppressor functions in cancer cells. Large-scale experimental detection of DNA methylation is still both labor-intensive and time-consuming. Therefore, it is necessary to develop in silico approaches for predicting methylation status of CpG islands. RESULTS: Based on a recent genome-scale dataset of DNA methylation in human brain tissues, we developed a classifier called MethCGI for predicting methylation status of CpG islands using a support vector machine (SVM). Nucleotide sequence contents as well as transcription factor binding sites (TFBSs) are used as features for the classification. The method achieves specificity of 84.65% and sensitivity of 84.32% on the brain data, and can also correctly predict about two-third of the data from other tissues reported in the MethDB database. AVAILABILITY: An online predictor based on MethCGI is available at http://166.111.201.7/MethCGI.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data available at Bioinformatics online and http://166.111.201.7/help.html. Shicai Fan, Xuegong Zhang, Michael Q. Zhang |
Bioinform. | 3 |
| 2006 | Recursive SVM feature selection and sample classification for mass-spectrometry and microarray dataabstractBACKGROUND: Like microarray-based investigations, high-throughput proteomics techniques require machine learning algorithms to identify biomarkers that are informative for biological classification problems. Feature selection and classification algorithms need to be robust to noise and outliers in the data. RESULTS: We developed a recursive support vector machine (R-SVM) algorithm to select important genes/biomarkers for the classification of noisy data. We compared its performance to a similar, state-of-the-art method (SVM recursive feature elimination or SVM-RFE), paying special attention to the ability of recovering the true informative genes/biomarkers and the robustness to outliers in the data. Simulation experiments show that a 5%- approximately 20% improvement over SVM-RFE can be achieved regard to these properties. The SVM-based methods are also compared with a conventional univariate method and their respective strengths and weaknesses are discussed. R-SVM was applied to two sets of SELDI-TOF-MS proteomics data, one from a human breast cancer study and the other from a study on rat liver cirrhosis. Important biomarkers found by the algorithm were validated by follow-up biological experiments. CONCLUSION: The proposed R-SVM method is suitable for analyzing noisy high-throughput proteomics and microarray data and it outperforms SVM-RFE in the robustness to noise and in the ability to recover informative features. The multivariate SVM-based method outperforms the univariate method in the classification performance, but univariate methods can reveal more of the differentially expressed features especially when there are correlations between the features. Xuegong Zhang, Xiu-qin Xu, Hon-chiu E. Leung, Lyndsay N. Harris, James D. Iglehart, Alexander Miron, Jun S. Liu, Wing Hung Wong |
BMC Bioinform. | 1 |
| 2006 | Significance of Gene Ranking for Classification of Microarray SamplesabstractMany methods for classification and gene selection with microarray data have been developed. These methods usually give a ranking of genes. Evaluating the statistical significance of the gene ranking is important for understanding the results and for further biological investigations, but this question has not been well addressed for machine learning methods in existing works. Here, we address this problem by formulating it in the framework of hypothesis testing and propose a solution based on resampling. The proposed r-test methods convert gene ranking results into position p-values to evaluate the significance of genes. The methods are tested on three real microarray data sets and three simulation data sets with support vector machines as the method of classification and gene selection. The obtained position p-values help to determine the number of genes to be selected and enable scientists to analyze selection results by sophisticated multivariate methods under the same statistical inference paradigm as for simple hypothesis testing methods. Xuegong Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2005 | A Neural Network Based Method for Shape Measurement in Steel Plate Forming Robot
Peifa Jia, Xuegong Zhang |
ISNN (3) | 3 |
| 2005 | A Novel Chamber Scheduling Method in Etching Tools Using Adaptive Neural Networks
Peifa Jia, Xuegong Zhang |
ISNN (3) | 3 |
| 2005 | MicroRNA identification based on sequence and structure alignmentabstractMOTIVATION: MicroRNAs (miRNA) are approximately 22 nt long non-coding RNAs that are derived from larger hairpin RNA precursors and play important regulatory roles in both animals and plants. The short length of the miRNA sequences and relatively low conservation of pre-miRNA sequences restrict the conventional sequence-alignment-based methods to finding only relatively close homologs. On the other hand, it has been reported that miRNA genes are more conserved in the secondary structure rather than in primary sequences. Therefore, secondary structural features should be more fully exploited in the homologue search for new miRNA genes. RESULTS: In this paper, we present a novel genome-wide computational approach to detect miRNAs in animals based on both sequence and structure alignment. Experiments show this approach has higher sensitivity and comparable specificity than other reported homologue searching methods. We applied this method on Anopheles gambiae and detected 59 new miRNA genes. AVAILABILITY: This program is available at http://bioinfo.au.tsinghua.edu.cn/miralign. SUPPLEMENTARY INFORMATION: Supplementary information is available at http://bioinfo.au.tsinghua.edu.cn/miralign/supplementary.htm. Xiaowo Wang, Jing Zhang 0010, Jin Gu, Xuegong Zhang, Yanda Li |
Bioinform. | 6 |
| 2005 | htSNPer1.0: software for haplotype block partition and htSNPs selectionabstractBACKGROUND: There is recently great interest in haplotype block structure and haplotype tagging SNPs (htSNPs) in the human genome for its implication on htSNPs-based association mapping strategy for complex disease. Different definitions have been used to characterize the haplotype block structure in the human genome, and several different performance criteria and algorithms have been suggested on htSNPs selection. RESULTS: A heuristic algorithm, generalized branch-and-bound algorithm, is applied to the searching of minimal set of haplotype tagging SNPs (htSNPs) according to different htSNPs performance criteria. We develop a software htSNPer1.0 to implement the algorithm, and integrate three htSNPs performance criteria and four haplotype block definitions for haplotype block partitioning. It is a software with powerful Graphical User Interface (GUI), which can be used to characterize the haplotype block structure and select htSNPs in the candidate gene or interested genomic regions. It can find the global optimization with only a fraction of the computing time consumed by exhaustive searching algorithm. CONCLUSION: htSNPer1.0 allows molecular geneticists to perform haplotype block analysis and htSNPs selection using different definitions and performance criteria. The software is a powerful tool for those focusing on association mapping based on strategy of haplotype block and htSNPs. Keyue Ding, Jing Zhang 0010, Kaixin Zhou, Xuegong Zhang |
BMC Bioinform. | 5 |
| 2005 | Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machineabstractBACKGROUND: MicroRNAs (miRNAs) are a group of short (approximately 22 nt) non-coding RNAs that play important regulatory roles. MiRNA precursors (pre-miRNAs) are characterized by their hairpin structures. However, a large amount of similar hairpins can be folded in many genomes. Almost all current methods for computational prediction of miRNAs use comparative genomic approaches to identify putative pre-miRNAs from candidate hairpins. Ab initio method for distinguishing pre-miRNAs from sequence segments with pre-miRNA-like hairpin structures is lacking. Being able to classify real vs. pseudo pre-miRNAs is important both for understanding of the nature of miRNAs and for developing ab initio prediction methods that can discovery new miRNAs without known homology. RESULTS: A set of novel features of local contiguous structure-sequence information is proposed for distinguishing the hairpins of real pre-miRNAs and pseudo pre-miRNAs. Support vector machine (SVM) is applied on these features to classify real vs. pseudo pre-miRNAs, achieving about 90% accuracy on human data. Remarkably, the SVM classifier built on human data can correctly identify up to 90% of the pre-miRNAs from other species, including plants and virus, without utilizing any comparative genomics information. CONCLUSION: The local structure-sequence features reflect discriminative and conserved characteristics of miRNAs, and the successful ab initio classification of real and pseudo pre-miRNAs opens a new approach for discovering new miRNAs. Chenghai Xue, Guo-Ping Liu 0003, Yanda Li, Xuegong Zhang |
BMC Bioinform. | 6 |
| 2004 | A simple strategy for detecting outlier samples in microarray dataabstractMicroarrays can monitor expression levels of thousands of genes simultaneously. Many people have used the gene expression data obtained with microarrays to classify different groups of samples, such as different types or subtypes of cancers. In our experiments as well as those of some other investigators, it has been observed that in some microarray data sets, there might be outlier samples which are either caused by imperfectness in the experiments or by possible mislabeling at certain steps. The existence of such samples impacts classification accuracy and may even cause misleading conclusions. In this paper, we studied this problem with two simulated data sets of typical scenarios and formed a simple but powerful strategy for detecting such outlier or mislabeled samples, built upon cross validation of the basic SVM classifier. The strategy was applied to a public colon cancer data set and it successfully detected 6 outlier cases. This work suggests an effective scheme for detecting outlier samples in a data set and for evaluating the sample quality. Yanda Li, Xuegong Zhang |
ICARCV | 3 |
| 2004 | Discovering possible context dependences around SNP sites in human genes with Bayesian network learningabstractSingle nucleotide polymorphisms (SNPs) are loci on the genome where different alleles are observed in the population. It has been observed that there might be some patterns or context dependences in the sequence segments adjacent to SNPs sites. Discovering such dependences is very important for understanding possible origins of SNPs in evolution. We collected 519,767 bi-allelic SNPs of human in gene regions from HGBASE and separated them in 6 groups according to the types of alleles at the SNP loci. Bayesian network structure learning technique is applied to discovery of possible dependences in sequence segments around these sites as well as in reference sequences collected as comparison. Noticeable probabilistic correlations among some loci were detected in all the 6 SNP groups and nothing significant was found in the reference sequences. The dependence relations found with different SNP groups are different. These putative context dependences around SNP sites provide important hints for further analyzing SNP-related sequences patterns. The work also illustrates the powerfulness of the Bayesian network method as a tool for biological sequence analysis. Xi Ma, Wei Hu 0002, Yimin Zhang 0002, Yanda Li, Xuegong Zhang |
ICARCV | 6 |
| 2004 | Parallelization of Bayesian Network based SNPs Pattern Analysis and Performance Characterization on SMP/HT
Justin J. Song, Eric Q. Li, Wei Hu 0002, Steven Ge, Chunrong Lai, Yimin Zhang 0002, Xuegong Zhang |
ICPADS | 7 |
| 2004 | Kernels based on weighted Levenshtein distanceabstractIn some real world applications, the sample could be described as a string of symbols rather than a vector of real numbers. It is necessary to determine the similarity or dissimilarity of two strings in many training algorithms. The widely used notion of similarity of two strings with different lengths is the weighted Levenshtein distance (WLD), which implies the minimum total weights of single symbol insertions, deletions and substitutions required to transform one string into another. In order to incorporate prior knowledge of strings into kernels used in support vector machine and other kernel machines, we utilize variants of this distance to replace distance measure in the RBF and exponential kernels and inner product in polynomial and sigmoid kernels, and form a new class of string kernels: Levenshtein kernels in this paper. Combining our new kernels with support vector machine, the error rate and variance on UCI splice site recognition dataset over 20 run is 5.88/spl mnplus/0.53, which is better than the best result 9.5/spl mnplus/0.7 from other five training algorithms. Xuegong Zhang |
IJCNN | 2 |
| 2004 | A Learning Algorithm with Gaussian Regularizer for Kernel Neuron
Xuegong Zhang |
ISNN (1) | 2 |
| 2002 | Kernel Nearest Neighbor Algorithm
Xuegong Zhang |
Neural Process. Lett. | 3 |
| 2000 | Joint speech signal enhancement based on spectral subtraction and SVD filter
Wenkai Lu, Xuegong Zhang, Yanda Li, Li Qin Shen, Weibin Zhu |
INTERSPEECH | 2 |
| 2000 | The signal reconstruction of speech by KPCA
Xuegong Zhang, Yanda Li, Li Qin Shen, Weibin Zhu |
INTERSPEECH | 2 |