Zhandong Liu

dblp:123/7924 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-7608-0831ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Image inpainting based on non-causal state-space duality
Zhandong Liu, Ruixia Song, Shuping Chen, Xiangwei Qi
J. Vis. Commun. Image Represent.2
2025 Positional Frequency Chaos Game Representation for Machine Learning-Based Classification of Crop lncRNAs
abstract
Alignment-based methods are fundamental for sequence comparison but are often computationally prohibitive for large-scale genomic analyses. This limitation has spurred the development of quicker, alignment-free alternatives, such as k-mer analysis, which are crucial for studying long noncoding ribonucleic acids (lncRNAs) in plants. These lncRNAs play critical roles in regulating gene expression at both the epigenetic and transcriptomic levels. However, existing alignmentfree approaches typically lose positional information, which can be vital for achieving accurate classification. We propose positional frequency chaos game representation (PFCGR), a novel encoding that improves the traditional frequency chaos game representation (FCGR) by incorporating four statistical moments of k-mer positions: mean, standard deviation, skewness, and kurtosis. This creates a multi-channel image representation of genomic sequences, enabling machine learning models such as Logistic Regression, Random Forests, and Convolutional Neural Networks to classify plant lncRNAs directly from raw genomic sequences. Tested on seven major crop species, our PFCGRbased classifiers achieve classification accuracies comparable to or exceeding those of the computationally intensive DNABERTbased model [1], while requiring 80 % to 95 % less computational time. These results demonstrate PFCGR's potential as an efficient and accurate tool for plant IncRNA identification, as well as its ability to facilitate large-scale computational studies in genomics.
Athanasios Papastathopoulos-Katsaros, Zhandong Liu
BIBE2
2025 Improving Physics-Informed Neural Network Extrapolation via Transfer Learning and Adaptive Activation Functions
Athanasios Papastathopoulos-Katsaros, Alexandra Stavrianidi, Zhandong Liu
ICANN (4)3
2025 Adaptive graph neural network protection algorithm based on differential privacy
Yong Li 0019, Zhandong Liu, Qianren Yang
J. Syst. Softw.3
2024 CoRegNet: unraveling gene co-regulation networks from public RNA-Seq repositories using a beta-binomial statistical model
abstract
Millions of RNA sequencing samples have been deposited into public databases, providing a rich resource for biological research. These datasets encompass tens of thousands of experiments and offer comprehensive insights into human cellular regulation. However, a major challenge is how to integrate these experiments that acquired at different conditions. We propose a new statistical tool based on beta-binomial distributions that can construct robust gene co-regulation network (CoRegNet) across tens of thousands of experiments. Our analysis of over 12 000 experiments involving human tissues and cells shows that CoRegNet significantly outperforms existing gene co-expression-based methods. Although the majority of the genes are linearly co-regulated, we did discover an interesting set of genes that are non-linearly co-regulated; half of the time they change in the same direction and the other half they change in the opposite direction. Additionally, we identified a set of gene pairs that follows the Simpson's paradox. By utilizing public domain data, CoRegNet offers a powerful approach for identifying functionally related gene pairs, thereby revealing new biological insights.
Ying-Wooi Wan, Rami Al-Ouran, Meichen Huang, Zhandong Liu
Briefings Bioinform.5
2023 SPA-STOCSY: an automated tool for identifying annotated and non-annotated metabolites in high-throughput NMR spectra
abstract
MOTIVATION: Nuclear magnetic resonance spectroscopy (NMR) is widely used to analyze metabolites in biological samples, but the analysis requires specific expertise, it is time-consuming, and can be inaccurate. Here, we present a powerful automate tool, SPatial clustering Algorithm-Statistical TOtal Correlation SpectroscopY (SPA-STOCSY), which overcomes challenges faced when analyzing NMR data and identifies metabolites in a sample with high accuracy. RESULTS: As a data-driven method, SPA-STOCSY estimates all parameters from the input dataset. It first investigates the covariance pattern among datapoints and then calculates the optimal threshold with which to cluster datapoints belonging to the same structural unit, i.e. the metabolite. Generated clusters are then automatically linked to a metabolite library to identify candidates. To assess SPA-STOCSY's efficiency and accuracy, we applied it to synthesized spectra and spectra acquired on Drosophila melanogaster tissue and human embryonic stem cells. In the synthesized spectra, SPA outperformed Statistical Recoupling of Variables (SRV), an existing method for clustering spectral peaks, by capturing a higher percentage of the signal regions and the close-to-zero noise regions. In the biological data, SPA-STOCSY performed comparably to the operator-based Chenomx analysis while avoiding operator bias, and it required <7 min of total computation time. Overall, SPA-STOCSY is a fast, accurate, and unbiased tool for untargeted analysis of metabolites in the NMR spectra. It may thus accelerate the use of NMR for scientific discoveries, medical diagnostics, and patient-specific decision making. AVAILABILITY AND IMPLEMENTATION: The codes of SPA-STOCSY are available at https://github.com/LiuzLab/SPA-STOCSY.
Li-Hua Ma, Ismael Ai-Ramahi, Juan Botas, Kevin Mackenzie, Genevera I. Allen, Damian W. Young, Zhandong Liu, Mirjana Maletic-Savatic
Bioinform.9
2023 Unravelling spatial gene associations with SEAGAL: a Python package for spatial transcriptomics data analysis and visualization
abstract
SUMMARY: In the era where transcriptome profiling moves toward single-cell and spatial resolutions, the traditional co-expression analysis lacks the power to fully utilize such rich information to unravel spatial gene associations. Here, we present a Python package called Spatial Enrichment Analysis of Gene Associations using L-index (SEAGAL) to detect and visualize spatial gene correlations at both single-gene and gene-set levels. Our package takes spatial transcriptomics datasets with gene expression and the aligned spatial coordinates as input. It allows for analyzing and visualizing genes' spatial correlations and cell types' colocalization within the precise spatial context. The output could be visualized as volcano plots and heatmaps with a few lines of code, thus providing an easy-yet-comprehensive tool for mining spatial gene associations. AVAILABILITY AND IMPLEMENTATION: The Python package SEAGAL can be installed using pip: https://pypi.org/project/seagal/. The source code and step-by-step tutorials are available at: https://github.com/linhuawang/SEAGAL.
Linhua Wang, Chaozhong Liu, Xiang H.-F. Zhang, Zhandong Liu
Bioinform.5
2023 Single-cell multi-omics integration for unpaired data by a siamese network with graph-based contrastive loss
abstract
BACKGROUND: Single-cell omics technology is rapidly developing to measure the epigenome, genome, and transcriptome across a range of cell types. However, it is still challenging to integrate omics data from different modalities. Here, we propose a variation of the Siamese neural network framework called MinNet, which is trained to integrate multi-omics data on the single-cell resolution by using graph-based contrastive loss. RESULTS: By training the model and testing it on several benchmark datasets, we showed its accuracy and generalizability in integrating scRNA-seq with scATAC-seq, and scRNA-seq with epitope data. Further evaluation demonstrated our model's unique ability to remove the batch effect, a common problem in actual practice. To show how the integration impacts downstream analysis, we established model-based smoothing and cis-regulatory element-inferring method and validated it with external pcHi-C evidence. Finally, we applied the framework to a COVID-19 dataset to bolster the original work with integration-based analysis, showing its necessity in single-cell multi-omics research. CONCLUSIONS: MinNet is a novel deep-learning framework for single-cell multi-omics sequencing data integration. It ranked top among other methods in benchmarking and is especially suitable for integrating datasets with batch and biological variances. With the single-cell resolution integration results, analysis of the interplay between genome and transcriptome can be done to help researchers understand their data and question.
Chaozhong Liu, Linhua Wang, Zhandong Liu
BMC Bioinform.3
2023 Correction: Single-cell multi-omics integration for unpaired data by a siamese network with graph-based contrastive loss
Chaozhong Liu, Linhua Wang, Zhandong Liu
BMC Bioinform.3
2022 C/PAS-BERT: A Novel Attention-Based Deep Learning Approach to Identify and Characterize Cleavage and Polyadenylation Sites
Venkata S. Jonnakuti, Zhandong Liu, Hari Krishna Yalamanchili
AMIA2
2021 MFECN: Multi-level Feature Enhanced Cumulative Network for Scene Text Detection
abstract
Recently, many scene text detection algorithms have achieved impressive performance by using convolutional neural networks. However, most of them do not make full use of the context among the hierarchical multi-level features to improve the performance of scene text detection. In this article, we present an efficient multi-level features enhanced cumulative framework based on instance segmentation for scene text detection. At first, we adopt a Multi-Level Features Enhanced Cumulative ( MFEC ) module to capture features of cumulative enhancement of representational ability. Then, a Multi-Level Features Fusion ( MFF ) module is designed to fully integrate both high-level and low-level MFEC features, which can adaptively encode scene text information. To verify the effectiveness of the proposed method, we perform experiments on six public datasets (namely, CTW1500, Total-text, MSRA-TD500, ICDAR2013, ICDAR2015, and MLT2017), and make comparisons with other state-of-the-art methods. Experimental results demonstrate that the proposed Multi-Level Features Enhanced Cumulative Network (MFECN) detector can well handle scene text instances with irregular shapes (i.e., curved, oriented, and horizontal) and achieves better or comparable results.
Zhandong Liu, Wengang Zhou 0001, Houqiang Li
ACM Trans. Multim. Comput. Commun. Appl.1
2020 PRIME: a probabilistic imputation method to reduce dropout effects in single-cell RNA sequencing
abstract
SUMMARY: Single-cell RNA sequencing technology provides a novel means to analyze the transcriptomic profiles of individual cells. The technique is vulnerable, however, to a type of noise called dropout effects, which lead to zero-inflated distributions in the transcriptome profile and reduce the reliability of the results. Single-cell RNA sequencing data, therefore, need to be carefully processed before in-depth analysis. Here, we describe a novel imputation method that reduces dropout effects in single-cell sequencing. We construct a cell correspondence network and adjust gene expression estimates based on transcriptome profiles for the local subnetwork of cells of the same type. We comprehensively evaluated this method, called PRIME (PRobabilistic IMputation to reduce dropout effects in Expression profiles of single-cell sequencing), on synthetic and eight real single-cell sequencing datasets and verified that it improves the quality of visualization and accuracy of clustering analysis and can discover gene expression patterns hidden by noise. AVAILABILITY AND IMPLEMENTATION: The source code for the proposed method is freely available at https://github.com/hyundoo/PRIME. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hyundoo Jeong, Zhandong Liu
Bioinform.2
2020 GDASC: a GPU parallel-based web server for detecting hidden batch factors
abstract
SUMMARY: We developed GDASC, a web version of our former DASC algorithm implemented with GPU. It provides a user-friendly web interface for detecting batch factors. Based on the good performance of DASC algorithm, it is able to give the most accurate results. For two steps of DASC, data-adaptive shrinkage and semi-non-negative matrix factorization, we designed parallelization strategies facing convex clustering solution and decomposition process. It runs more than 50 times faster than the original version on the representative RNA sequencing quality control dataset. With its accuracy and high speed, this server will be a useful tool for batch effects analysis. AVAILABILITY AND IMPLEMENTATION: http://bioinfo.nankai.edu.cn/gdasc.php. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiao Wang 0099, Haidong Yi, Jia Wang 0051, Zhandong Liu, Yanbin Yin, Han Zhang 0017
Bioinform.4
2020 AB-LSTM: Attention-based Bidirectional LSTM Model for Scene Text Detection
abstract
Detection of scene text in arbitrary shapes is a challenging task in the field of computer vision. Most existing scene text detection methods exploit the rectangle/quadrangular bounding box to denote the detected text, which fails to accurately fit text with arbitrary shapes, such as curved text. In addition, recent progress on scene text detection has benefited from Fully Convolutional Network. Text cues contained in multi-level convolutional features are complementary for detecting scene text objects. How to explore these multi-level features is still an open problem. To tackle the above issues, we propose an Attention-based Bidirectional Long Short-Term Memory (AB-LSTM) model for scene text detection. First, word stroke regions (WSRs) and text center blocks (TCBs) are extracted by two AB-LSTM models, respectively. Then, the union of WSRs and TCBs are used to represent text objects. To verify the effectiveness of the proposed method, we perform experiments on four public benchmarks: CTW1500, Total-text, ICDAR2013, and MSRA-TD500, and compare it with existing state-of-the-art methods. Experiment results demonstrate that the proposed method can achieve competitive results, and well handle scene text objects with arbitrary shapes (i.e., curved, oriented, and horizontal forms).
Zhandong Liu, Wengang Zhou 0001, Houqiang Li
ACM Trans. Multim. Comput. Commun. Appl.1
2019 Scene text detection with fully convolutional neural networks
Zhandong Liu, Wengang Zhou 0001, Houqiang Li
Multim. Tools Appl.1
2018 Detecting hidden batch factors through data-adaptive adjustment for biological effects
abstract
Motivation: Batch effects are one of the major source of technical variations that affect the measurements in high-throughput studies such as RNA sequencing. It has been well established that batch effects can be caused by different experimental platforms, laboratory conditions, different sources of samples and personnel differences. These differences can confound the outcomes of interest and lead to spurious results. A critical input for batch correction algorithms is the knowledge of batch factors, which in many cases are unknown or inaccurate. Hence, the primary motivation of our paper is to detect hidden batch factors that can be used in standard techniques to accurately capture the relationship between gene expression and other modeled variables of interest. Results: We introduce a new algorithm based on data-adaptive shrinkage and semi-Non-negative Matrix Factorization for the detection of unknown batch effects. We test our algorithm on three different datasets: (i) Sequencing Quality Control, (ii) Topotecan RNA-Seq and (iii) Single-cell RNA sequencing (scRNA-Seq) on Glioblastoma Multiforme. We have demonstrated a superior performance in identifying hidden batch effects as compared to existing algorithms for batch detection in all three datasets. In the Topotecan study, we were able to identify a new batch factor that has been missed by the original study, leading to under-representation of differentially expressed genes. For scRNA-Seq, we demonstrated the power of our method in detecting subtle batch effects. Availability and implementation: DASC R package is available via Bioconductor or at https://github.com/zhanglabNKU/DASC. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Haidong Yi, Ayush T. Raman, Han Zhang 0017, Genevera I. Allen, Zhandong Liu
Bioinform.5
2017 CRISPRcloud: a secure cloud-based pipeline for CRISPR pooled screen deconvolution
abstract
SUMMARY: We present a user-friendly, cloud-based, data analysis pipeline for the deconvolution of pooled screening data. This tool, CRISPRcloud, serves a dual purpose of extracting, clustering and analyzing raw next generation sequencing files derived from pooled screening experiments while at the same time presenting them in a user-friendly way on a secure web-based platform. Moreover, CRISPRcloud serves as a useful web-based analysis pipeline for reanalysis of pooled CRISPR screening datasets. Taken together, the framework described in this study is expected to accelerate development of web-based bioinformatics tool for handling all studies which include next generation sequencing data. AVAILABILITY AND IMPLEMENTATION: http://crispr.nrihub.org. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hyun-hwan Jeong, Seon Young Kim, Maxime W. C. Rousseaux, Huda Y. Zoghbi, Zhandong Liu
Bioinform.5
2017 Comprehensive evaluation of RNA-seq quantification methods for linearity
abstract
BACKGROUND: Deconvolution is a mathematical process of resolving an observed function into its constituent elements. In the field of biomedical research, deconvolution analysis is applied to obtain single cell-type or tissue specific signatures from a mixed signal and most of them follow the linearity assumption. Although recent development of next generation sequencing technology suggests RNA-seq as a fast and accurate method for obtaining transcriptomic profiles, few studies have been conducted to investigate best RNA-seq quantification methods that yield the optimum linear space for deconvolution analysis. RESULTS: Using a benchmark RNA-seq dataset, we investigated the linearity of abundance estimated from seven most popular RNA-seq quantification methods both at the gene and isoform levels. Linearity is evaluated through parameter estimation, concordance analysis and residual analysis based on a multiple linear regression model. Results show that count data gives poor parameter estimations, large intercepts and high inter-sample variability; while TPM value from Kallisto and Salmon shows high linearity in all analyses. CONCLUSIONS: Salmon and Kallisto TPM data gives the best fit to the linear model studied. This suggests that TPM values estimated from Salmon and Kallisto are the ideal RNA-seq measurements for deconvolution studies.
Haijing Jin, Ying-Wooi Wan, Zhandong Liu
BMC Bioinform.3
2017 The International Conference on Intelligent Biology and Medicine (ICIBM) 2016: from big data to big analytical tools
abstract
The 2016 International Conference on Intelligent Biology and Medicine (ICIBM 2016) was held on December 8-10, 2016 in Houston, Texas, USA. ICIBM included eight scientific sessions, four tutorials, one poster session, four highlighted talks and four keynotes that covered topics on 3D genomics structural analysis, next generation sequencing (NGS) analysis, computational drug discovery, medical informatics, cancer genomics, and systems biology. Here, we present a summary of the nine research articles selected from ICIBM 2016 program for publishing in BMC Bioinformatics.
Zhandong Liu, W. Jim Zheng, Genevera I. Allen, Jianhua Ruan, Zhongming Zhao
BMC Bioinform.1
2017 Method for unconstrained text detection in natural scene image
abstract
Text detection in natural scene images is an important prerequisite for many content‐based multimedia understanding applications. The authors present a simple and effective text detection method in natural scene image. Firstly, MSERs are extracted by the V‐MSER algorithm from channels of G , H , S , O 1 , and O 2 , as component candidates. Since text is composed of character candidates, the authors design an MRF model to exploit the relationship between characters. Secondly, in order to filter out non‐text components, they design a set of two‐layers filtering scheme: most of the non‐text components can be filtered by the first layer of the filtering scheme; the second layer filtering scheme is an AdaBoost classifier, which is trained by the features of compactness, horizontal variance and vertical variance, and aspect ratio. Then, only four simple features are adopted to generate component pairs. Finally, according to the orientation similarity of the component pairs, component pairs which have roughly the same orientation are merged into text lines. The proposed method is evaluated on two public datasets: ICDAR 2011 and MSRA‐TD500. It achieves 82.94 and 75% F ‐measure, respectively. Especially, the experimental results, on their URMQ_LHASA‐TD220 dataset which contains 220 images for multi‐orientation and multi‐language text lines evaluation, show that the proposed method is general for detecting scene text lines in different languages.
Zhandong Liu, Xiangwei Qi, Mei Nian, Reziwanguli Xiamixiding
IET Comput. Vis.1
2016 Chinese sign language recognition based on trajectory and hand shape features
abstract
Sign language recognition(SLR) is a challenging task due to the diversity of the signs. To tackle the problem, this paper utilize both trajectory features and hand shape features. Since the trajectory features and hand shape features are not in the same domain, it is unreasonable to concatenate them naively or model them with a unified model. To deal with the issue, we adopt Support Vector Machine(SVM) and validation Hidden Markov Models(VHMM), respectively. To depict the direction of the trajectory, we first employ histogram of oriented displacement(HOD) with SVM to SLR. We propose the relative distance features(RDF) by using VHMM to consider the relationship between hands and the other body parts. As for hand shape feature, we explore histogram of oriented gradient(HOG) in local hand regions with VHMM, too. To facilitate late fusion, we normalize the probabilities of different features to the same range and fuse them for the final classification. To demonstrate the effectiveness of our proposed method, we conduct the experiments both in ChaLearn dataset and our self-build Kinect-based Chinese sign language dataset. The results show that our method outperforms the classical methods and some state-of-the-art methods.
Zhandong Liu
VCIP2
2016 TCGA2STAT: simple TCGA data access for integrated statistical analysis in R
abstract
MOTIVATION: Massive amounts of high-throughput genomics data profiled from tumor samples were made publicly available by the Cancer Genome Atlas (TCGA). RESULTS: We have developed an open source software package, TCGA2STAT, to obtain the TCGA data, wrangle it, and pre-process it into a format ready for multivariate and integrated statistical analysis in the R environment. In a user-friendly format with one single function call, our package downloads and fully processes the desired TCGA data to be seamlessly integrated into a computational analysis pipeline. No further technical or biological knowledge is needed to utilize our software, thus making TCGA data easily accessible to data scientists without specific domain knowledge. AVAILABILITY AND IMPLEMENTATION: TCGA2STAT is available from the https://cran.r-project.org/web/packages/TCGA2STAT/index.html SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected].
Ying-Wooi Wan, Genevera I. Allen, Zhandong Liu
Bioinform.3
2015 Graphical models via univariate exponential family distributions
Eunho Yang, Pradeep Ravikumar, Genevera I. Allen, Zhandong Liu
J. Mach. Learn. Res.4
2014 Mixed Graphical Models via Exponential Families
abstract
Markov Random Fields, or undirected graphical models are widely used to model high-dimensional multivariate data. Classical instances of these models, such as Gaussian Graphical and Ising Models, as well as recent extensions to graphical models specified by univariate exponential families, assume all variables arise from the same distribution. Complex data from high-throughput genomics and social networking for example, often contain discrete, count, and continuous variables measured on the same set of samples. To model such heterogeneous data, we develop a \emphnovel class of mixed graphical models by specifying that each node-conditional distribution is a member of a possibly different univariate exponential family. We study several instances of our model, and propose scalable M-estimators for recovering the underlying network structure. Simulations as well as an application to learning mixed genomic networks from next generation sequencing and mutation data demonstrate the versatility of our methods.
Eunho Yang, Yulia Baker, Pradeep Ravikumar, Genevera I. Allen, Zhandong Liu
AISTATS5
2014 Combinatorial therapy discovery using mixed integer linear programming
abstract
MOTIVATION: Combinatorial therapies play increasingly important roles in combating complex diseases. Owing to the huge cost associated with experimental methods in identifying optimal drug combinations, computational approaches can provide a guide to limit the search space and reduce cost. However, few computational approaches have been developed for this purpose, and thus there is a great need of new algorithms for drug combination prediction. RESULTS: Here we proposed to formulate the optimal combinatorial therapy problem into two complementary mathematical algorithms, Balanced Target Set Cover (BTSC) and Minimum Off-Target Set Cover (MOTSC). Given a disease gene set, BTSC seeks a balanced solution that maximizes the coverage on the disease genes and minimizes the off-target hits at the same time. MOTSC seeks a full coverage on the disease gene set while minimizing the off-target set. Through simulation, both BTSC and MOTSC demonstrated a much faster running time over exhaustive search with the same accuracy. When applied to real disease gene sets, our algorithms not only identified known drug combinations, but also predicted novel drug combinations that are worth further testing. In addition, we developed a web-based tool to allow users to iteratively search for optimal drug combinations given a user-defined gene set. AVAILABILITY: Our tool is freely available for noncommercial use at http://www.drug.liuzlab.org/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kaifang Pang, Ying-Wooi Wan, William T. Choi, Lawrence A. Donehower, Jingchun Sun, Dhruv Pant, Zhandong Liu
Bioinform.7
2013 Conditional Random Fields via Univariate Exponential Families
abstract
Conditional random fields, which model the distribution of a multivariate response conditioned on a set of covariates using undirected graphs, are widely used in a variety of multivariate prediction applications. Popular instances of this class of models such as categorical-discrete CRFs, Ising CRFs, and conditional Gaussian based CRFs, are not however best suited to the varied types of response variables in many applications, including count-valued responses. We thus introduce a “novel subclass of CRFs”, derived by imposing node-wise conditional distributions of response variables conditioned on the rest of the responses and the covariates as arising from univariate exponential families. This allows us to derive novel multivariate CRFs given any univariate exponential distribution, including the Poisson, negative binomial, and exponential distributions. Also in particular, it addresses the common CRF problem of specifying feature'' functions determining the interactions between response variables and covariates. We develop a class of tractable penalized $M$-estimators to learn these CRF distributions from data, as well as a unified sparsistency analysis for this general class of CRFs showing exact structure recovery can be achieved with high probability."
Eunho Yang, Pradeep Ravikumar, Genevera I. Allen, Zhandong Liu
NIPS4
2013 On Poisson Graphical Models
abstract
Undirected graphical models, such as Gaussian graphical models, Ising, and multinomial/categorical graphical models, are widely used in a variety of applications for modeling distributions over a large number of variables. These standard instances, however, are ill-suited to modeling count data, which are increasingly ubiquitous in big-data settings such as genomic sequencing data, user-ratings data, spatial incidence data, climate studies, and site visits. Existing classes of Poisson graphical models, which arise as the joint distributions that correspond to Poisson distributed node-conditional distributions, have a major drawback: they can only model negative conditional dependencies for reasons of normalizability given its infinite domain. In this paper, our objective is to modify the Poisson graphical model distribution so that it can capture a rich dependence structure between count-valued variables. We begin by discussing two strategies for truncating the Poisson distribution and show that only one of these leads to a valid joint distribution; even this model, however, has limitations on the types of variables and dependencies that may be modeled. To address this, we propose two novel variants of the Poisson distribution and their corresponding joint graphical model distributions. These models provide a class of Poisson graphical models that can capture both positive and negative conditional dependencies between count-valued variables. One can learn the graph structure of our model via penalized neighborhood selection, and we demonstrate the performance of our methods by learning simulated networks as well as a network from microRNA-Sequencing data.
Eunho Yang, Pradeep Ravikumar, Genevera I. Allen, Zhandong Liu
NIPS4
2013 Digital sorting of complex tissues for cell type-specific gene expression profiles
abstract
BACKGROUND: Cellular heterogeneity is present in almost all gene expression profiles. However, transcriptome analysis of tissue specimens often ignores the cellular heterogeneity present in these samples. Standard deconvolution algorithms require prior knowledge of the cell type frequencies within a tissue or their in vitro expression profiles. Furthermore, these algorithms tend to report biased estimations. RESULTS: Here, we describe a Digital Sorting Algorithm (DSA) for extracting cell-type specific gene expression profiles from mixed tissue samples that is unbiased and does not require prior knowledge of cell type frequencies. CONCLUSIONS: The results suggest that DSA is a specific and sensitivity algorithm in gene expression profile deconvolution and will be useful in studying individual cell types of complex tissues.
Ying-Wooi Wan, Kaifang Pang, Lionel M. L. Chow, Zhandong Liu
BMC Bioinform.5
2012 A Log-Linear Graphical Model for inferring genetic networks from high-throughput sequencing data
abstract
Gaussian graphical models are often used to infer gene networks based on microarray expression data. Many scientists, however, have begun using high-throughput sequencing technologies to measure gene expression. As the resulting high-dimensional count data consists of counts of sequencing reads for each gene, Gaussian graphical models are not optimal for modeling gene networks based on this discrete data. We develop a novel method for estimating high-dimensional Poisson graphical models, the Log-Linear Graphical Model, allowing us to infer networks based on high-throughput sequencing data. Our model assumes a pair-wise Markov property: conditional on all other variables, each variable is Poisson. We estimate our model locally via neighborhood selection by fitting 1-norm penalized log-linear models. Additionally, we develop a fast parallel algorithm permitting us to fit our graphical model to high-dimensional genomic data sets. We illustrate the effectiveness of our methods for recovering network structure from count data through simulations and a case study on breast cancer microRNA networks.
Genevera I. Allen, Zhandong Liu
BIBM2
2012 Graphical Models via Generalized Linear Models
abstract
Undirected graphical models, or Markov networks, such as Gaussian graphical models and Ising models enjoy popularity in a variety of applications. In many settings, however, data may not follow a Gaussian or binomial distribution assumed by these models. We introduce a new class of graphical models based on generalized linear models (GLM) by assuming that node-wise conditional distributions arise from exponential families. Our models allow one to estimate networks for a wide class of exponential distributions, such as the Poisson, negative binomial, and exponential, by fitting penalized GLMs to select the neighborhood for each node. A major contribution of this paper is the rigorous statistical analysis showing that with high probability, the neighborhood of our graphical models can be recovered exactly. We provide examples of high-throughput genomic networks learned via our GLM graphical models for multinomial and Poisson distributed data.
Eunho Yang, Pradeep Ravikumar, Genevera I. Allen, Zhandong Liu
NIPS4