EDBT 2026 Demo / reviewers in the wild / expert
Shuangge Ma
dblp:07/3898
· DBLP profile ↗
46ranked-venue papers
17as first author
19since 2021 · last 2025
0000-0001-9001-4999ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 44 · 17 first-author · 17 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bilevel Network Learning via Hierarchically Structured SparsityabstractAccurate network estimation serves as the cornerstone for understanding complex systems across scientific domains, from decoding gene regulatory networks in systems biology to identifying social relationship patterns in computational sociology. Modern applications demand methods that simultaneously address two critical challenges: capturing nonlinear dependencies between variables and reconstructing inherent hierarchical structures where higher-level entities coordinate lower-level components (e.g., functional pathways organizing gene clusters). Traditional Gaussian graphical models fundamentally fail in these aspects due to their restrictive linear assumptions and flat network representations. We propose NNBLNet, a neural network-based learning framework for bi-level network inference. The core innovation lies in hierarchical selection layers that enforce structural consistency between high-level coordinator groups and their constituent low-level connections via adaptive sparsity constraints. This architecture is integrated with a compositional neural network architecture that learn cross-level association patterns through constrained nonlinear transformations, explicitly preserving hierarchical dependencies while overcoming the representational limitations of linear methods. Crucially, we establish formal theoretical guarantees for the consistent recovery of both high-level connections and their internal low-level structures under general statistical regimes. Extensive validation demonstrates NNBLNet's effectiveness across synthetic and real-world scenarios, achieving superior F1 scores compared to competitive methods and particularly beneficial for complex systems analysis through its interpretable bi-level structure discovery. Jingyuan Yang 0021, Shuangge Ma, Mengyun Wu |
NeurIPS | 3 |
| 2025 | GE-IA-NAM: gene-environment interaction analysis via imaging-assisted neural additive modelabstractMOTIVATION: Gene-environment (G-E) interaction analysis is crucial in cancer research, offering insights into how genetic and environmental factors jointly influence cancer outcomes. Most existing G-E interaction methods are regression-based, which may lack flexibility to capture complex data patterns. Recent advances have investigated deep neural network-based G-E models. However, these methods may be more vulnerable to information deficiency due to challenges such as limited sample size and high dimensionality. Apart from genetic and environmental data, pathological images have emerged as a widely accessible and informative resource for cancer modeling, presenting its potential to enhance G-E modeling. RESULTS: We propose the pathological imaging-assisted neural additive model for G-E analysis (GE-IA-NAM). The flexible and interpretable additive network architecture is adopted to account for individualized effects associated with genetic factors, environmental factors, and their interactions. To improve G-E modeling, an assisted-learning strategy is investigated, which adopts a joint analysis to integrate information from pathological images. Simulations and the analysis of lung and skin cancer datasets from The Cancer Genome Atlas demonstrate the competitive performance of the proposed method. AVAILABILITY AND IMPLEMENTATION: Python code implementing the proposed method is available at https://github.com/Mr-maoge/NAM-IA-GE. The data that support the findings in this article are openly available in TCGA (The Cancer Genome Atlas) at https://portal.gdc.cancer.gov/. Jingmao Li, Yaqing Xu, Shuangge Ma, Kuangnan Fang |
Bioinform. | 3 |
| 2025 | Joint modeling of mixed outcomes using a rank-based sparse neural network
Jiajing Xue, Yaqing Xu, Jingmao Li, Shuangge Ma, Kuangnan Fang |
J. Biomed. Informatics | 4 |
| 2024 | EditorialabstractIt has been slightly more than a year since I took over the position of Editor-in-Chief of Briefings in Bioinformatics (BIB). When I look back, the first thing that comes to mind is great appreciation. I would like to take this opportunity to thank all readers, authors, reviewers, and the editorial team—the journal could not have stayed this strong without your fully dedicated support. BIB remains one of the most highly ranked and competitive journals in bioinformatics and a top choice for numerous bioinformatics researchers worldwide. Similar to many other bioinformatics journals, we have witnessed a surge in submissions that involve AI, deep learning, and machine learning. BIB has become a leading venue for publishing new methods, applications, and reviews in AI and deep learning. There is no doubt that such techniques have and will continue to revolutionize science including bioinformatics. Along the way, there can be a few aspects worth some additional consideration. Many bioinformatics datasets are perceived as being challenged by small sample sizes and high-dimensional input. It is not fully clear whether/how deep neural networks can be immune from the curse of dimensionality, which may lead to a lack of stability, inferior prediction performance, etc. Some recent techniques have grown increasingly complicated, with more complex architectures and more layers/nodes. There is a strong and increasing demand for stability evaluation metrics and techniques as well as methods that can purposely improve stability. A naturally related question is: with a certain sample size (amount of information), what is the maximum complexity a deep neural network model can have? In classic analytics, there are some rules of thumb (for example, in regression analysis 101, it is recommended that the ratio of sample size/number of input variables is at least 5–10). However, such guidelines are missing in AI and deep/machine learning-based bioinformatics studies. It is recommended that more attention to complexity, stability, limitations in training data, etc. is paid in future studies. Different from some scientific domains, interpretability can be highly desirable in many bioinformatics studies. In classic bioinformatics, interpretability is enhanced by the relatively lucid model structures, differentiation/identification of signals from noises, quantification of effect sizes, and elucidation of causal paths. It is recognized that, with the foundational differences of AI and deep/machine learning, we need to rethink some of those perspectives. However, interpretability overall is still and may be more desirable—this view is shared by numerous research organizations, funding agencies, and researchers. We welcome more discussions/research on interpretability and the development of (more) interpretable AI/deep learning techniques. For most if not all bioinformatics problems and datasets, there are multiple available approaches. In the process of developing new approaches, it is critical to know how they compare against existing alternatives. In practical applications, it is necessary to know when an approach works (and, equally importantly, when it does not). Rigorous theoretical investigations remain somewhat rare in bioinformatics technology development. Well-designed, extensive, and fair simulations and comparisons with state-of-the-art existing techniques can serve this purpose to a certain extent. Recognizing the limitations of synthetic data, and the ultimate goal of analyzing practical data, we recommend carefully gauged comparisons based on extensive, unbiased, and high-quality data. Studies with selection bias in data and benchmarks will not be sufficiently appealing. “Plurality should not be posited without necessary”—Occam’s razor. Bioinformatics problems are growing more complicated and demand the development of more complex tools. On the other hand, it is also recognized that some practical problems can be equally resolved with existing tools that may be simpler, more robust, and lucid. As such, when a new (and likely more complicated) approach is developed, it is strongly recommended that it is benchmarked against existing alternatives comprehensively in terms of model fitting, computational cost, stability, and other aspects, which will better inform practitioners how to choose tools properly. To date, BIB does not uniformly require making software and data fully publicly available. Rather, this is handled on a case-by-case basis. The public availability of software and data ensures reproducibility and facilitates broad utilization. It will be a huge burden to society if users have to re-develop software for a published technique. For data, we encourage sharing when it is feasible (for example, allowed by funding agencies), though require access to data if needed for peer review. For software, especially when it is a new and nontrivial approach, we very strongly encourage depositing at stable and publicly accessible repositories. Such recommendations have been strongly advocated by many reviewers and readers. We welcome input from all readers and authors regarding reproducibility and publication of data and software. The field of bioinformatics has never been as dynamic as it is right now. With your support, I am fully confident of the even brighter future of BIB. Again, thank you. Shuangge Ma |
Briefings Bioinform. | 1 |
| 2024 | Heterogeneity-aware Clustered Distributed Learning for Multi-source Data AnalysisabstractIn diverse fields ranging from finance to omics, it is increasingly common that data is distributed with multiple individual sources (referred to as “clients” in some studies). Integrating raw data, although powerful, is often not feasible, for example, when there are considerations on privacy protection. Distributed learning techniques have been developed to integrate summary statistics as opposed to raw data. In many existing distributed learning studies, it is stringently assumed that all the clients have the same model. To accommodate data heterogeneity, some federated learning methods allow for client-specific models. In this article, we consider the scenario that clients form clusters, those in the same cluster have the same model, and different clusters have different models. Further considering the clustering structure can lead to a better understanding of the “interconnections” among clients and reduce the number of parameters. To this end, we develop a novel penalization approach. Specifically, group penalization is imposed for regularized estimation and selection of important variables, and fusion penalization is imposed to automatically cluster clients. An effective ADMM algorithm is developed, and the estimation, selection, and clustering consistency properties are established under mild conditions. Simulation and data analysis further demonstrate the practical utility and superiority of the proposed approach. Yuanxing Chen, Qingzhao Zhang 0002, Shuangge Ma, Kuangnan Fang |
J. Mach. Learn. Res. | 3 |
| 2023 | EditorialabstractWhen I was informed that I had been selected as the next editor in chief, I was thrilled. And then very quickly, the magnitude of the responsibility set in. Under the unparalleled leadership of Dr Martin Bishop, and with strong support from the publisher, Briefings in Bioinformatics (BIB) has grown into one of the best journals in bioinformatics. Although there will be challenges, it is the whole editorial team’s full intention to keep the leading position of the journal and take it one big step forward. I would like to take this opportunity to first congratulate Martin on his stellar editorial and career achievements and wish him a happy retirement, and also thank the publisher for trusting me with leading BIB. In his editorial, Martin highlighted the completion of the Human Genome Project in 2003. I was a mathematical statistics student then. I can still clearly recall now the excitement of that time, as it coincided with my own initial fascination with molecular research and bioinformatics. My first bioinformatics publication was in Bioinformatics, the sister journal of BIB. I published my first BIB paper (which was a review with Dr Jian Huang on penalized feature selection and classification in bioinformatics; https://academic.oup.com/bib/article/9/5/392/267473) in 2008 and have been a veteran author and editorial board member since then. I very much appreciate this opportunity to briefly share some of my personal views on the journal. As the old saying goes, ‘A man with one watch always knows what time it is. A man with two watches is never sure’. Over the years, there has been a significant accumulation of bioinformatics analysis methods, software packages, and data collection and analysis platforms, as partly reflected in BIB publications. For the most pressing problems, there have been many different tools, making different data assumptions, taking different strategies and building on different techniques. For practitioners, it is critical to have access to head-to-head comparisons in methodology and empirical performance, so as to make proper methodological choices. Many BIB review papers have successfully served this purpose and provided insightful guidance to practice, and their high level of citation testifies to this point. In the coming years, the journal will continue to welcome review and original research papers that provide comprehensive, thorough and fair evaluations and comparisons of bioinformatics tools for both old and new fields and topics, and, more importantly, critiques that can truly assist practice. As reflected in Martin’s final editorial (https://doi.org/10.1093/bib/bbad176), bioinformatics tools have evolved significantly and fast. Many of the early tools have been based on ‘classic’ statistical concepts/techniques such as hypothesis testing, regression and clustering. In recent years, we have witnessed a significant increase in studies that are built on deep learning and AI techniques—this may also be true for other bioinformatics journals. While we do not know when and what the next analytical or technical revolution will be, we are sure there will be one. I strongly encourage bioinformaticians to publish their leading-edge methods in the journal. On the other hand, I would also like to cautiously note that all new methods should be built on solid scientific ground with well-justified rationale. They should be gaged against well-tested benchmarks in a rigorous and fair way—newer is not necessarily better. And their weaknesses should be clearly identified along with strengths. There should be sufficient attention to ‘simple’ problems: the risk of overfitting may increase because of the significantly increased number of parameters; some methods that have superior performance on specific data sets may have narrow applicability and much worse performance on others; some deep learning models have weak biological interpretability; and they sometimes have high computational complexity and weak stability and replicability. The scientific problems bioinformaticians address have also evolved significantly. When I was a student, my early projects were microarray gene expression normalization and identification of differential genes. Twenty years later, my current projects include modeling heterogeneous gene networks and integrating multi-omics data for gene–environment analysis. In the near future, it is still expected that all BIB papers will have a strong molecular component. On the other hand, we can also foresee more and more ‘combining power’ with other types of data, for example, single-cell sequencing data with pathological imaging data, bulk sequencing data with radiological imaging data and disease-specific gene signatures with electronic health record data. Integration of multiple data modalities has brought unprecedentedly rich information and a strong demand for more sophisticated bioinformatics methods. Unique challenges to be addressed include but are not limited to error accumulation, higher computational cost, a higher risk of overfitting, significantly decreased stability/replicability, less lucid interpretation and the tradeoff between improvement in performance and increased complexity. We hope that BIB authors will take a leading role in addressing these challenges. In the past few years, BIB has experienced a steady increase in quality and impact and number of submissions/publications. Our whole editorial team will strive to support the journal’s ongoing success by maintaining the highest review standards and further improving the submission and reviewing experience. We will closely monitor developments in the field and adapt the journal. While we will continue to welcome submissions from previously published authors, we also invite new authors to help us expand coverage and readership of the journal. I very much look forward to working with all authors, readers, editorial board members and publishing staff. Shuangge Ma |
Briefings Bioinform. | 1 |
| 2023 | FunctanSNP: an R package for functional analysis of dense SNP data (with interactions)abstractSUMMARY: Densely measured SNP data are routinely analyzed but face challenges due to its high dimensionality, especially when gene-environment interactions are incorporated. In recent literature, a functional analysis strategy has been developed, which treats dense SNP measurements as a realization of a genetic function and can 'bypass' the dimensionality challenge. However, there is a lack of portable and friendly software, which hinders practical utilization of these functional methods. We fill this knowledge gap and develop the R package FunctanSNP. This comprehensive package encompasses estimation, identification, and visualization tools and has undergone extensive testing using both simulated and real data, confirming its reliability. FunctanSNP can serve as a convenient and reliable tool for analyzing SNP and other densely measured data. AVAILABILITY AND IMPLEMENTATION: The package is available at https://CRAN.R-project.org/package=FunctanSNP. Kuangnan Fang, Qingzhao Zhang 0002, Shuangge Ma |
Bioinform. | 4 |
| 2023 | Prior information-assisted integrative analysis of multiple datasetsabstractMOTIVATION: Analyzing genetic data to identify markers and construct predictive models is of great interest in biomedical research. However, limited by cost and sample availability, genetic studies often suffer from the "small sample size, high dimensionality" problem. To tackle this problem, an integrative analysis that collectively analyzes multiple datasets with compatible designs is often conducted. For regularizing estimation and selecting relevant variables, penalization and other regularization techniques are routinely adopted. "Blindly" searching over a vast number of variables may not be efficient. RESULTS: We propose incorporating prior information to assist integrative analysis of multiple genetic datasets. To obtain accurate prior information, we adopt a convolutional neural network with an active learning strategy to label textual information from previous studies. Then the extracted prior information is incorporated using a group LASSO-based technique. We conducted a series of simulation studies that demonstrated the satisfactory performance of the proposed method. Finally, data on skin cutaneous melanoma are analyzed to establish practical utility. AVAILABILITY AND IMPLEMENTATION: Code is available at https://github.com/ldz7/PAIA. The data that support the findings in this article are openly available in TCGA (The Cancer Genome Atlas) at https://portal.gdc.cancer.gov/. Dongzuo Liang, Yang Li 0070, Shuangge Ma |
Bioinform. | 4 |
| 2023 | Spatio-temporally smoothed deep survival neural network
Dongzuo Liang, Shuangge Ma, Chenjin Ma |
J. Biomed. Informatics | 3 |
| 2023 | Aligned deep neural network for integrative analysis with high-dimensional input
Shunqin Zhang, Sanguo Zhang, Huangdi Yi, Shuangge Ma |
J. Biomed. Informatics | 4 |
| 2022 | Analysis of cancer omics data: a selective review of statistical techniquesabstractCancer is an omics disease. The development in high-throughput profiling has fundamentally changed cancer research and clinical practice. Compared with clinical, demographic and environmental data, the analysis of omics data-which has higher dimensionality, weaker signals and more complex distributional properties-is much more challenging. Developments in the literature are often 'scattered', with individual studies focused on one or a few closely related methods. The goal of this review is to assist cancer researchers with limited statistical expertise in establishing the 'overall framework' of cancer omics data analysis. To facilitate understanding, we mainly focus on intuition, concepts and key steps, and refer readers to the original publications for mathematical details. This review broadly covers unsupervised and supervised analysis, as well as individual-gene-based, gene-set-based and gene-network-based analysis. We also briefly discuss 'special topics' including interaction analysis, multi-datasets analysis and multi-omics analysis. Chenjin Ma, Mengyun Wu, Shuangge Ma |
Briefings Bioinform. | 3 |
| 2022 | Replicability in cancer omics data analysis: measures and empirical explorationsabstractIn biomedical research, the replicability of findings across studies is highly desired. In this study, we focus on cancer omics data, for which the examination of replicability has been mostly focused on important omics variables identified in different studies. In published literature, although there have been extensive attention and ad hoc discussions, there is insufficient quantitative research looking into replicability measures and their properties. The goal of this study is to fill this important knowledge gap. In particular, we consider three sensible replicability measures, for which we examine distributional properties and develop a way of making inference. Applying them to three The Cancer Genome Atlas (TCGA) datasets reveals in general low replicability and significant across-data variations. To further comprehend such findings, we resort to simulation, which confirms the validity of the findings with the TCGA data and further informs the dependence of replicability on signal level (or equivalently sample size). Overall, this study can advance our understanding of replicability for cancer omics and other studies that have identification as a key goal. Hongmin Liang, Qingzhao Zhang 0002, Shuangge Ma |
Briefings Bioinform. | 4 |
| 2022 | iSFun: an R package for integrative dimension reduction analysisabstractSUMMARY: In the analysis of high-dimensional omics data, dimension reduction techniques-including principal component analysis (PCA), partial least squares (PLS) and canonical correlation analysis (CCA)-have been extensively used. When there are multiple datasets generated by independent studies with compatible designs, integrative analysis has been developed and shown to outperform meta-analysis, other multidatasets analysis, and individual-data analysis. To facilitate integrative dimension reduction analysis in daily practice, we develop the R package iSFun, which can comprehensively conduct integrative sparse PCA, PLS and CCA, as well as meta-analysis and stacked analysis. The package can conduct analysis under the homogeneity and heterogeneity models and with the magnitude- and sign-based contrasted penalties. As a 'byproduct', this article is the first to develop integrative analysis built on the CCA technique, further expanding the scope of integrative analysis. AVAILABILITY AND IMPLEMENTATION: The package is available at https://CRAN.R-project.org/package=iSFun. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. Kuangnan Fang, Qingzhao Zhang 0002, Shuangge Ma |
Bioinform. | 4 |
| 2022 | Network-based cancer heterogeneity analysis incorporating multi-view of prior informationabstractMOTIVATION: Cancer genetic heterogeneity analysis has critical implications for tumour classification, response to therapy and choice of biomarkers to guide personalized cancer medicine. However, existing heterogeneity analysis based solely on molecular profiling data usually suffers from a lack of information and has limited effectiveness. Many biomedical and life sciences databases have accumulated a substantial volume of meaningful biological information. They can provide additional information beyond molecular profiling data, yet pose challenges arising from potential noise and uncertainty. RESULTS: In this study, we aim to develop a more effective heterogeneity analysis method with the help of prior information. A network-based penalization technique is proposed to innovatively incorporate a multi-view of prior information from multiple databases, which accommodates heterogeneity attributed to both differential genes and gene relationships. To account for the fact that the prior information might not be fully credible, we propose a weighted strategy, where the weight is determined dependent on the data and can ensure that the present model is not excessively disturbed by incorrect information. Simulation and analysis of The Cancer Genome Atlas glioblastoma multiforme data demonstrate the practical applicability of the proposed method. AVAILABILITY AND IMPLEMENTATION: R code implementing the proposed method is available at https://github.com/mengyunwu2020/PECM. The data that support the findings in this paper are openly available in TCGA (The Cancer Genome Atlas) at https://portal.gdc.cancer.gov/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Li 0070, Shaodong Xu, Shuangge Ma, Mengyun Wu |
Bioinform. | 3 |
| 2022 | GEInfo: an R package for gene-environment interaction analysis incorporating prior informationabstractSUMMARY: Gene-environment (G-E) interactions have important implications for many complex diseases. With higher dimensionality and weaker signals, G-E interaction analysis is more challenged than the analysis of main G (and E) effects. The accumulation of published literature makes it possible to borrow strength from prior information and improve analysis. In a recent study, a 'quasi-likelihood + penalization' approach was developed to effectively incorporate prior information. Here, we first extend it to linear, logistic and Poisson regressions. Such models are much more popular in practice. More importantly, we develop the R package GEInfo, which realizes this approach in a user-friendly manner. To facilitate direct comparison and routine data analysis, the package also includes functions for alternative methods and visualization. AVAILABILITY AND IMPLEMENTATION: The package is available at https://CRAN.R-project.org/package=GEInfo. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. Hongduo Liu, Shuangge Ma |
Bioinform. | 3 |
| 2021 | Vertical integration methods for gene expression data analysisabstractGene expression data have played an essential role in many biomedical studies. When the number of genes is large and sample size is limited, there is a 'lack of information' problem, leading to low-quality findings. To tackle this problem, both horizontal and vertical data integrations have been developed, where vertical integration methods collectively analyze data on gene expressions as well as their regulators (such as mutations, DNA methylation and miRNAs). In this article, we conduct a selective review of vertical data integration methods for gene expression data. The reviewed methods cover both marginal and joint analysis and supervised and unsupervised analysis. The main goal is to provide a sketch of the vertical data integration paradigm without digging into too many technical details. We also briefly discuss potential pitfalls, directions for future developments and application notes. Mengyun Wu, Huangdi Yi, Shuangge Ma |
Briefings Bioinform. | 3 |
| 2021 | DReSS: a method to quantitatively describe the influence of structural perturbations on state spaces of genetic regulatory networksabstractStructures of genetic regulatory networks are not fixed. These structural perturbations can cause changes to the reachability of systems' state spaces. As system structures are related to genotypes and state spaces are related to phenotypes, it is important to study the relationship between structures and state spaces. However, there is still no method can quantitively describe the reachability differences of two state spaces caused by structural perturbations. Therefore, Difference in Reachability between State Spaces (DReSS) is proposed. DReSS index family can quantitively describe differences of reachability, attractor sets between two state spaces and can help find the key structure in a system, which may influence system's state space significantly. First, basic properties of DReSS including non-negativity, symmetry and subadditivity are proved. Then, typical examples are shown to explain the meaning of DReSS and the differences between DReSS and traditional graph distance. Finally, differences of DReSS distribution between real biological regulatory networks and random networks are compared. Results show most structural perturbations in biological networks tend to affect reachability inside and between attractor basins rather than to affect attractor set itself when compared with random networks, which illustrates that most genotype differences tend to influence the proportion of different phenotypes and only a few ones can create new phenotypes. DReSS can provide researchers with a new insight to study the relation between genotypes and phenotypes. Ziqiao Yin, Shuangge Ma, Zhilong Mi, Zhiming Zheng 0001 |
Briefings Bioinform. | 3 |
| 2021 | HeteroGGM: an R package for Gaussian graphical model-based heterogeneity analysisabstractSUMMARY: Heterogeneity is a hallmark of many complex human diseases, and unsupervised heterogeneity analysis has been extensively conducted using high-throughput molecular measurements and histopathological imaging features. 'Classic' heterogeneity analysis has been based on simple statistics such as mean, variance and correlation. Network-based analysis takes interconnections as well as individual variable properties into consideration and can be more informative. Several Gaussian graphical model (GGM)-based heterogeneity analysis techniques have been developed, but friendly and portable software is still lacking. To facilitate more extensive usage, we develop the R package HeteroGGM, which conducts GGM-based heterogeneity analysis using the advanced penaliztaion techniques, can provide informative summary and graphical presentation, and is efficient and friendly. AVAILABILITYAND IMPLEMENTATION: The package is available at https://CRAN.R-project.org/package=HeteroGGM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mingyang Ren, Sanguo Zhang, Qingzhao Zhang 0002, Shuangge Ma |
Bioinform. | 4 |
| 2021 | GEInter: an R package for robust gene-environment interaction analysisabstractSUMMARY: For understanding complex diseases, gene-environment (G-E) interactions have important implications beyond main G and E effects. Most of the existing analysis approaches and software packages cannot accommodate data contamination/long-tailed distribution. We develop GEInter, a comprehensive R package tailored to robust G-E interaction analysis. For both marginal and joint analysis, for data without and with missingness, for continuous and censored survival responses, it comprehensively conducts identification, estimation, visualization and prediction. It can fill an important gap in the existing literature and enjoy broad applicability. AVAILABILITY AND IMPLEMENTATION: TCGA data is analyzed as demonstrating examples. It is well known that such data is publicly available https://cran.r-project.org/web/packages/GEInter/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengyun Wu, Xing Qin, Shuangge Ma |
Bioinform. | 3 |
| 2020 | NCutYX: a package for clustering analysis of multilayer omics dataabstractSUMMARY: Multilayer omics profiling has become a major venue for understanding complex diseases. We develop NCutYX, an R package for clustering analysis of multilayer omics data. The package and methods jointly analyze multiple layers of omics measurements and effectively accommodate their regulations. They systematically conduct a series of analysis based on the normalized cut technique, including the clusterings of subjects and omics measurements and biclustering. The package can be valuable for its timely context, novel methods, and comprehensiveness. AVAILABILITY: https://cran.r-project.org/web/packages/NCutYX/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sebastian J. Teran Hidalgo, Mengyun Wu, Shuangge Ma |
Bioinform. | 3 |
| 2019 | Robust genetic interaction analysisabstractFor the risk, progression, and response to treatment of many complex diseases, it has been increasingly recognized that genetic interactions (including gene-gene and gene-environment interactions) play important roles beyond the main genetic and environmental effects. In practical genetic interaction analyses, model mis-specification and outliers/contaminations in response variables and covariates are not uncommon, and demand robust analysis methods. Compared with their nonrobust counterparts, robust genetic interaction analysis methods are significantly less popular but are gaining attention fast. In this article, we provide a comprehensive review of robust genetic interaction analysis methods, on their methodologies and applications, for both marginal and joint analysis, and for addressing model mis-specification as well as outliers/contaminations in response variables and covariates. Mengyun Wu, Shuangge Ma |
Briefings Bioinform. | 2 |
| 2016 | Group-combined P-values with applications to genetic association studiesabstractMOTIVATION: In large-scale genetic association studies with tens of hundreds of single nucleotide polymorphisms (SNPs) genotyped, the traditional statistical framework of logistic regression using maximum likelihood estimator (MLE) to infer the odds ratios of SNPs may not work appropriately. This is because a large number of odds ratios need to be estimated, and the MLEs may be not stable when some of the SNPs are in high linkage disequilibrium. Under this situation, the P-value combination procedures seem to provide good alternatives as they are constructed on the basis of single-marker analysis. RESULTS: The commonly used P-value combination methods (such as the Fisher's combined test, the truncated product method, the truncated tail strength and the adaptive rank truncated product) may lose power when the significance level varies across SNPs. To tackle this problem, a group combined P-value method (GCP) is proposed, where the P-values are divided into multiple groups and then are combined at the group level. With this strategy, the significance values are integrated at different levels, and the power is improved. Simulation shows that the GCP can effectively control the type I error rates and have additional power over the existing methods-the power increase can be as high as over 50% under some situations. The proposed GCP method is applied to data from the Genetic Analysis Workshop 16. Among all the methods, only the GCP and ARTP can give the significance to identify a genomic region covering gene DSC3 being associated with rheumatoid arthritis, but the GCP provides smaller P-value. AVAILABILITY AND IMPLEMENTATION: http://www.statsci.amss.ac.cn/yjscy/yjy/lqz/201510/t20151027_313273.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaonan Hu, Sanguo Zhang, Shuangge Ma, Qizhai Li |
Bioinform. | 4 |
| 2016 | EPS: an empirical Bayes approach to integrating pleiotropy and tissue-specific information for prioritizing risk genesabstractMOTIVATION: Researchers worldwide have generated a huge volume of genomic data, including thousands of genome-wide association studies (GWAS) and massive amounts of gene expression data from different tissues. How to perform a joint analysis of these data to gain new biological insights has become a critical step in understanding the etiology of complex diseases. Due to the polygenic architecture of complex diseases, the identification of risk genes remains challenging. Motivated by the shared risk genes found in complex diseases and tissue-specific gene expression patterns, we propose as an Empirical Bayes approach to integrating Pleiotropy and Tissue-Specific information (EPS) for prioritizing risk genes. RESULTS: As demonstrated by extensive simulation studies, EPS greatly improves the power of identification for disease-risk genes. EPS enables rigorous hypothesis testing of pleiotropy and tissue-specific risk gene expression patterns. All of the model parameters can be adaptively estimated from the developed expectation-maximization (EM) algorithm. We applied EPS to the bipolar disorder and schizophrenia GWAS from the Psychiatric Genomics Consortium, along with the gene expression data for multiple tissues from the Genotype-Tissue Expression project. The results of the real data analysis demonstrate many advantages of EPS. AVAILABILITY AND IMPLEMENTATION: The EPS software is available on https://sites.google.com/site/liujin810822 CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jin Liu 0011, Shuangge Ma, Can Yang 0002 |
Bioinform. | 3 |
| 2015 | Measures for the degree of overlap of gene signatures and applications to TCGAabstractFor cancer and many other complex diseases, a large number of gene signatures have been generated. In this study, we use cancer as an example and note that other diseases can be analyzed in a similar manner. For signatures generated in multiple independent studies on the same cancer type and outcome, and for signatures on different cancer types, it is of interest to evaluate their degree of overlap. Many of the existing studies simply count the number (or percentage) of overlapped genes shared by two signatures. Such an approach has serious limitations. In this study, as a demonstrating example, we consider cancer prognosis data under the Cox model. Lasso, which is representative of a large number of regularization methods, is adopted for generating gene signatures. We examine two families of measures for quantifying the degree of overlap. The first family is based on the Cox-Lasso estimates at the optimal tunings, and the second family is based on estimates across the whole solution paths. Within each family, multiple measures, which describe the overlap from different perspectives, are introduced. The analysis of TCGA (The Cancer Genome Atlas) data on five cancer types shows that the degree of overlap varies across measures, cancer types and types of (epi)genetic measurements. More investigations are needed to better describe and understand the overlaps among gene signatures. Xingjie Shi, Huangdi Yi, Shuangge Ma |
Briefings Bioinform. | 3 |
| 2015 | A selective review of robust variable selection with applications in bioinformaticsabstractA drastic amount of data have been and are being generated in bioinformatics studies. In the analysis of such data, the standard modeling approaches can be challenged by the heavy-tailed errors and outliers in response variables, the contamination in predictors (which may be caused by, for instance, technical problems in microarray gene expression studies), model mis-specification and others. Robust methods are needed to tackle these challenges. When there are a large number of predictors, variable selection can be as important as estimation. As a generic variable selection and regularization tool, penalization has been extensively adopted. In this article, we provide a selective review of robust penalized variable selection approaches especially designed for high-dimensional data from bioinformatics and biomedical studies. We discuss the robust loss functions, penalty functions and computational algorithms. The theoretical properties and implementation are also briefly examined. Application examples of the robust penalization approaches in representative bioinformatics and biomedical studies are also illustrated. Cen Wu, Shuangge Ma |
Briefings Bioinform. | 2 |
| 2015 | Combining multidimensional genomic measurements for predicting cancer prognosis: observations from TCGAabstractWith accumulating research on the interconnections among different types of genomic regulations, researchers have found that multidimensional genomic studies outperform one-dimensional studies in multiple aspects. Among many sources of multidimensional genomic data, The Cancer Genome Atlas (TCGA) provides the public with comprehensive profiling data on >30 cancer types, making it an ideal test bed for conducting and comparing different analyses. In this article, the analysis goal is to apply several existing methods and associate multidimensional genomic measurements with cancer outcomes in particular prognosis, with special focus on the predictive power of genomic signatures. We exploit clinical data and four types of genomic measurement including mRNA gene expression, DNA methylation, microRNA and copy number alterations for breast invasive carcinoma, glioblastoma multiforme, acute myeloid leukemia and lung squamous cell carcinoma collected by TCGA. To accommodate the high dimensionality, we extract important features using Principal Component Analysis, Partial Least Squares and Least Absolute Shrinkage and Selection Operator (Lasso), which are representative of dimension reduction and variable selection techniques and have been extensively adopted, and fit Cox survival models with combined important features. We calibrate the predictive power of each type of genomic measurement for the prognosis of four cancer types and find that the results vary across cancers. Our analysis also suggests that for most of the cancers in our study and the adopted methods, there is no substantial improvement in prediction when adding other genomic measurement after gene expression and clinical covariates have been included in the model. This is consistent with the findings that molecular features measured at the transcription level affect clinical outcomes more directly than those measured at the DNA/epigenetic level. Xingjie Shi, Jian Huang 0003, Ben-Chang Shia, Shuangge Ma |
Briefings Bioinform. | 6 |
| 2015 | Deciphering the associations between gene expression and copy number alteration using a sparse double Laplacian shrinkage approachabstractMOTIVATION: Both gene expression levels (GEs) and copy number alterations (CNAs) have important biological implications. GEs are partly regulated by CNAs, and much effort has been devoted to understanding their relations. The regulation analysis is challenging with one gene expression possibly regulated by multiple CNAs and one CNA potentially regulating the expressions of multiple genes. The correlations among GEs and among CNAs make the analysis even more complicated. The existing methods have limitations and cannot comprehensively describe the regulation. RESULTS: A sparse double Laplacian shrinkage method is developed. It jointly models the effects of multiple CNAs on multiple GEs. Penalization is adopted to achieve sparsity and identify the regulation relationships. Network adjacency is computed to describe the interconnections among GEs and among CNAs. Two Laplacian shrinkage penalties are imposed to accommodate the network adjacency measures. Simulation shows that the proposed method outperforms the competing alternatives with more accurate marker identification. The Cancer Genome Atlas data are analysed to further demonstrate advantages of the proposed method. AVAILABILITY AND IMPLEMENTATION: R code is available at http://works.bepress.com/shuangge/49/. Xingjie Shi, Jian Huang 0003, Shuangge Ma |
Bioinform. | 5 |
| 2014 | Similarity of markers identified from cancer gene expression studies: observations from GEOabstractGene expression profiling has been extensively conducted in cancer research. The analysis of multiple independent cancer gene expression datasets may provide additional information and complement single-dataset analysis. In this study, we conduct multi-dataset analysis and are interested in evaluating the similarity of cancer-associated genes identified from different datasets. The first objective of this study is to briefly review some statistical methods that can be used for such evaluation. Both marginal analysis and joint analysis methods are reviewed. The second objective is to apply those methods to 26 Gene Expression Omnibus (GEO) datasets on five types of cancers. Our analysis suggests that for the same cancer, the marker identification results may vary significantly across datasets, and different datasets share few common genes. In addition, datasets on different cancers share few common genes. The shared genetic basis of datasets on the same or different cancers, which has been suggested in the literature, is not observed in the analysis of GEO data. Xingjie Shi, Shihao Shen, Jin Liu 0011, Jian Huang 0003, Shuangge Ma |
Briefings Bioinform. | 6 |
| 2012 | Adjusting confounders in ranking biomarkers: a model-based ROC approachabstractHigh-throughput studies have been extensively conducted in the research of complex human diseases. As a representative example, consider gene-expression studies where thousands of genes are profiled at the same time. An important objective of such studies is to rank the diagnostic accuracy of biomarkers (e.g. gene expressions) for predicting outcome variables while properly adjusting for confounding effects from low-dimensional clinical risk factors and environmental exposures. Existing approaches are often fully based on parametric or semi-parametric models and target evaluating estimation significance as opposed to diagnostic accuracy. Receiver operating characteristic (ROC) approaches can be employed to tackle this problem. However, existing ROC ranking methods focus on biomarkers only and ignore effects of confounders. In this article, we propose a model-based approach which ranks the diagnostic accuracy of biomarkers using ROC measures with a proper adjustment of confounding effects. To this end, three different methods for constructing the underlying regression models are investigated. Simulation study shows that the proposed methods can accurately identify biomarkers with additional diagnostic power beyond confounders. Analysis of two cancer gene-expression studies demonstrates that adjusting for confounders can lead to substantially different rankings of genes. Jialiang Li 0001, Shuangge Ma |
Briefings Bioinform. | 3 |
| 2012 | Integrative prescreening in analysis of multiple cancer genomic studiesabstractBACKGROUND: In high throughput cancer genomic studies, results from the analysis of single datasets often suffer from a lack of reproducibility because of small sample sizes. Integrative analysis can effectively pool and analyze multiple datasets and provides a cost effective way to improve reproducibility. In integrative analysis, simultaneously analyzing all genes profiled may incur high computational cost. A computationally affordable remedy is prescreening, which fits marginal models, can be conducted in a parallel manner, and has low computational cost. RESULTS: An integrative prescreening approach is developed for the analysis of multiple cancer genomic datasets. Simulation shows that the proposed integrative prescreening has better performance than alternatives, particularly including prescreening with individual datasets, an intensity approach and meta-analysis. We also analyze multiple microarray gene profiling studies on liver and pancreatic cancers using the proposed approach. CONCLUSIONS: The proposed integrative prescreening provides an effective way to reduce the dimensionality in cancer genomic studies. It can be coupled with existing analysis methods to identify cancer markers. Rui Song 0006, Jian Huang 0003, Shuangge Ma |
BMC Bioinform. | 3 |
| 2011 | Principal component analysis based methods in bioinformatics studiesabstractIn analysis of bioinformatics data, a unique challenge arises from the high dimensionality of measurements. Without loss of generality, we use genomic study with gene expression measurements as a representative example but note that analysis techniques discussed in this article are also applicable to other types of bioinformatics studies. Principal component analysis (PCA) is a classic dimension reduction approach. It constructs linear combinations of gene expressions, called principal components (PCs). The PCs are orthogonal to each other, can effectively explain variation of gene expressions, and may have a much lower dimensionality. PCA is computationally simple and can be realized using many existing software packages. This article consists of the following parts. First, we review the standard PCA technique and their applications in bioinformatics data analysis. Second, we describe recent 'non-standard' applications of PCA, including accommodating interactions among genes, pathways and network modules and conducting PCA with estimating equations as opposed to gene expressions. Third, we introduce several recently proposed PCA-based techniques, including the supervised PCA, sparse PCA and functional PCA. The supervised PCA and sparse PCA have been shown to have better empirical performance than the standard PCA. The functional PCA can analyze time-course gene expression data. Last, we raise the awareness of several critical but unsolved problems related to PCA. The goal of this article is to make bioinformatics researchers aware of the PCA technique and more importantly its most recent development, so that this simple yet effective dimension reduction technique can be better employed in bioinformatics data analysis. Shuangge Ma |
Briefings Bioinform. | 1 |
| 2011 | Ranking prognosis markers in cancer genomic studiesabstractIn cancer research, high-throughput genomic studies have been extensively conducted, searching for markers associated with cancer diagnosis, prognosis and variation in response to treatment. In this article, we analyze cancer prognosis studies and investigate ranking markers based on their marginal prognosis power. To avoid ambiguity, we focus on microarray gene expression studies where genes are the markers, but note that the methodology and results are applicable to other high-throughput studies. The objectives of this study are 2-fold. First, we investigate ranking markers under three commonly adopted semiparametric models, namely the Cox, accelerated failure time and additive risk models. Data analysis shows that the ranking may vary significantly under different models. Second, we describe a nonparametric concordance measure, which has roots in the time-dependent ROC (receiver operating characteristic) framework and relies on much weaker assumptions than the semiparametric models. In simulation, it is shown that ranking using the concordance measure is not sensitive to model specification whereas ranking under the semiparametric models is. In data analysis, the concordance measure generates rankings significantly different from those under the semiparametric models. Shuangge Ma |
Briefings Bioinform. | 1 |
| 2010 | Semiparametric prognosis models in genomic studiesabstractDevelopment of high-throughput technologies makes it possible to survey the whole genome. Genomic studies have been extensively conducted, searching for markers with predictive power for prognosis of complex diseases such as cancer, diabetes and obesity. Most existing statistical analyses are focused on developing marker selection techniques, while little attention is paid to the underlying prognosis models. In this article, we review three commonly used prognosis models, namely the Cox, additive risk and accelerated failure time models. We conduct simulation and show that gene identification can be unsatisfactory under model misspecification. We analyze three cancer prognosis studies under the three models, and show that the gene identification results, prediction performance of all identified genes combined, and reproducibility of each identified gene are model-dependent. We suggest that in practical data analysis, more attention should be paid to the model assumption, and multiple models may need to be considered. Shuangge Ma, Jian Huang 0003, Mingyu Shi, Yang Li 0070, Ben-Chang Shia |
Briefings Bioinform. | 1 |
| 2010 | Identification of non-Hodgkin's lymphoma prognosis signatures using the CTGDR methodabstractMOTIVATION: Although NHL (non-Hodgkin's lymphoma) is the fifth leading cause of cancer incidence and mortality in the USA, it remains poorly understood and is largely incurable. Biomedical studies have shown that genomic variations, measured with SNPs (single nucleotide polymorphisms) in genes, may have independent predictive power for disease-free survival in NHL patients beyond clinical measurements. RESULTS: We apply the CTGDR (clustering threshold gradient directed regularization) method to genetic association studies using SNPs, analyze data from an association study of NHL and identify prognosis signatures to diffuse large B cell lymphoma (DLBCL) and follicular lymphoma (FL), the two most common subtypes of NHL. With the CTGDR method, we are able to account for the joint effects of multiple genes/SNPs, whereas most existing studies are single-marker based. In addition, we are able to account for the 'gene and SNP-within-gene' hierarchical structure and identify not only predictive genes but also predictive SNPs within identified genes. In contrast, existing studies are limited to either gene or SNP identification, but not both. We propose using resampling methods to evaluate the predictive power and reproducibility of identified genes and SNPs. Simulation study and data analysis suggest satisfactory performance of the CTGDR method. Shuangge Ma, Jian Huang 0003, Xuesong Han, Theodore Holford, Qing Lan, Nathaniel Rothman, Peter Boyle, Tongzhang Zheng |
Bioinform. | 1 |
| 2010 | Detection of gene pathways with predictive power for breast cancer prognosisabstractBACKGROUND: Prognosis is of critical interest in breast cancer research. Biomedical studies suggest that genomic measurements may have independent predictive power for prognosis. Gene profiling studies have been conducted to search for predictive genomic measurements. Genes have the inherent pathway structure, where pathways are composed of multiple genes with coordinated functions. The goal of this study is to identify gene pathways with predictive power for breast cancer prognosis. Since our goal is fundamentally different from that of existing studies, a new pathway analysis method is proposed. RESULTS: The new method advances beyond existing alternatives along the following aspects. First, it can assess the predictive power of gene pathways, whereas existing methods tend to focus on model fitting accuracy only. Second, it can account for the joint effects of multiple genes in a pathway, whereas existing methods tend to focus on the marginal effects of genes. Third, it can accommodate multiple heterogeneous datasets, whereas existing methods analyze a single dataset only. We analyze four breast cancer prognosis studies and identify 97 pathways with significant predictive power for prognosis. Important pathways missed by alternative methods are identified. CONCLUSIONS: The proposed method provides a useful alternative to existing pathway analysis methods. Identified pathways can provide further insights into breast cancer prognosis. Shuangge Ma, Michael R. Kosorok |
BMC Bioinform. | 1 |
| 2010 | Incorporating gene co-expression network in identification of cancer prognosis markersabstractBACKGROUND: Extensive biomedical studies have shown that clinical and environmental risk factors may not have sufficient predictive power for cancer prognosis. The development of high-throughput profiling technologies makes it possible to survey the whole genome and search for genomic markers with predictive power. Many existing studies assume the interchangeability of gene effects and ignore the coordination among them. RESULTS: We adopt the weighted co-expression network to describe the interplay among genes. Although there are several different ways of defining gene networks, the weighted co-expression network may be preferred because of its computational simplicity, satisfactory empirical performance, and because it does not demand additional biological experiments. For cancer prognosis studies with gene expression measurements, we propose a new marker selection method that can properly incorporate the network connectivity of genes. We analyze six prognosis studies on breast cancer and lymphoma. We find that the proposed approach can identify genes that are significantly different from those using alternatives. We search published literature and find that genes identified using the proposed approach are biologically meaningful. In addition, they have better prediction performance and reproducibility than genes identified using alternatives. CONCLUSIONS: The network contains important information on the functionality of genes. Incorporating the network structure can improve cancer marker identification. Shuangge Ma, Mingyu Shi, Yang Li 0070, Danhui Yi, Ben-Chang Shia |
BMC Bioinform. | 1 |
| 2009 | Enriching PubMed Related Article Search with Sentence Level Co-citations
Nam Tran, Pedro Alves, Shuangge Ma, Michael Krauthammer |
AMIA | 3 |
| 2009 | Identification of differential gene pathways with principal component analysisabstractMOTIVATION: Development of high-throughput technology makes it possible to measure expressions of thousands of genes simultaneously. Genes have the inherent pathway structure, where pathways are composed of multiple genes with coordinated biological functions. It is of great interest to identify differential gene pathways that are associated with the variations of phenotypes. RESULTS: We propose the following approach for detecting differential gene pathways. First, we construct gene pathways using databases such as KEGG or GO. Second, for each pathway, we extract a small number of representative features, which are linear combinations of gene expressions and/or their transformations. Specifically, we propose using (i) principal components (PCs) of gene expression sets, (ii) PCs of expanded gene expression sets and (iii) expanded sets of PCs of gene expressions, as the representative features. Third, we identify differential gene pathways as those with representative features significantly associated with the variations of phenotypes, particularly disease clinical outcomes, in regression models. The false discovery rate approach is used to adjust for multiple comparisons. Analysis of three gene expression datasets suggests that (i) the proposed approach can effectively identify differential gene pathways; (ii) PCs that explain only a small amount of variations of gene expressions may bear significant associations between gene pathways and phenotypes; (iii) including second-order terms of gene expressions may lead to identification of new differential gene pathways; (iv) the proposed approach is relatively insensitive to additional noises; and (v) the proposed approach can identify gene pathways missed by alternative approaches. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuangge Ma, Michael R. Kosorok |
Bioinform. | 1 |
| 2009 | Regularized gene selection in cancer microarray meta-analysisabstractBACKGROUND: In cancer studies, it is common that multiple microarray experiments are conducted to measure the same clinical outcome and expressions of the same set of genes. An important goal of such experiments is to identify a subset of genes that can potentially serve as predictive markers for cancer development and progression. Analyses of individual experiments may lead to unreliable gene selection results because of the small sample sizes. Meta analysis can be used to pool multiple experiments, increase statistical power, and achieve more reliable gene selection. The meta analysis of cancer microarray data is challenging because of the high dimensionality of gene expressions and the differences in experimental settings amongst different experiments. RESULTS: We propose a Meta Threshold Gradient Descent Regularization (MTGDR) approach for gene selection in the meta analysis of cancer microarray data. The MTGDR has many advantages over existing approaches. It allows different experiments to have different experimental settings. It can account for the joint effects of multiple genes on cancer, and it can select the same set of cancer-associated genes across multiple experiments. Simulation studies and analyses of multiple pancreatic and liver cancer experiments demonstrate the superior performance of the MTGDR. CONCLUSION: The MTGDR provides an effective way of analyzing multiple cancer microarray studies and selecting reliable cancer-associated genes. Shuangge Ma, Jian Huang 0003 |
BMC Bioinform. | 1 |
| 2008 | Penalized feature selection and classification in bioinformaticsabstractIn bioinformatics studies, supervised classification with high-dimensional input variables is frequently encountered. Examples routinely arise in genomic, epigenetic and proteomic studies. Feature selection can be employed along with classifier construction to avoid over-fitting, to generate more reliable classifier and to provide more insights into the underlying causal relationships. In this article, we provide a review of several recently developed penalized feature selection and classification techniques--which belong to the family of embedded feature selection methods--for bioinformatics studies with high-dimensional input. Classification objective functions, penalty functions and computational algorithms are discussed. Our goal is to make interested researchers aware of these feature selection and classification methods that are applicable to high-dimensional bioinformatics data. Shuangge Ma, Jian Huang 0003 |
Briefings Bioinform. | 1 |
| 2007 | Clustering threshold gradient descent regularization: with applications to microarray studiesabstractMOTIVATION: An important goal of microarray studies is to discover genes that are associated with clinical outcomes, such as disease status and patient survival. While a typical experiment surveys gene expressions on a global scale, there may be only a small number of genes that have significant influence on a clinical outcome. Moreover, expression data have cluster structures and the genes within a cluster have correlated expressions and coordinated functions, but the effects of individual genes in the same cluster may be different. Accordingly, we seek to build statistical models with the following properties. First, the model is sparse in the sense that only a subset of the parameter vector is non-zero. Second, the cluster structures of gene expressions are properly accounted for. RESULTS: For gene expression data without pathway information, we divide genes into clusters using commonly used methods, such as K-means or hierarchical approaches. The optimal number of clusters is determined using the Gap statistic. We propose a clustering threshold gradient descent regularization (CTGDR) method, for simultaneous cluster selection and within cluster gene selection. We apply this method to binary classification and censored survival analysis. Compared to the standard TGDR and other regularization methods, the CTGDR takes into account the cluster structure and carries out feature selection at both the cluster level and within-cluster gene level. We demonstrate the CTGDR on two studies of cancer classification and two studies correlating survival of lymphoma patients with microarray expressions. AVAILABILITY: R code is available upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuangge Ma, Jian Huang 0003 |
Bioinform. | 1 |
| 2007 | Additive risk survival model with microarray dataabstractBACKGROUND: Microarray techniques survey gene expressions on a global scale. Extensive biomedical studies have been designed to discover subsets of genes that are associated with survival risks for diseases such as lymphoma and construct predictive models using those selected genes. In this article, we investigate simultaneous estimation and gene selection with right censored survival data and high dimensional gene expression measurements. RESULTS: We model the survival time using the additive risk model, which provides a useful alternative to the proportional hazards model and is adopted when the absolute effects, instead of the relative effects, of multiple predictors on the hazard function are of interest. A Lasso (least absolute shrinkage and selection operator) type estimate is proposed for simultaneous estimation and gene selection. Tuning parameter is selected using the V-fold cross validation. We propose Leave-One-Out cross validation based methods for evaluating the relative stability of individual genes and overall prediction significance. CONCLUSION: We analyze the MCL and DLBCL data using the proposed approach. A small number of probes represented on the microarrays are identified, most of which have sound biological implications in lymphoma development. The selected probes are relatively stable and the proposed approach has overall satisfactory prediction power. Shuangge Ma, Jian Huang 0003 |
BMC Bioinform. | 1 |
| 2007 | Supervised group Lasso with applications to microarray data analysisabstractBACKGROUND: A tremendous amount of efforts have been devoted to identifying genes for diagnosis and prognosis of diseases using microarray gene expression data. It has been demonstrated that gene expression data have cluster structure, where the clusters consist of co-regulated genes which tend to have coordinated functions. However, most available statistical methods for gene selection do not take into consideration the cluster structure. RESULTS: We propose a supervised group Lasso approach that takes into account the cluster structure in gene expression data for gene selection and predictive model building. For gene expression data without biological cluster information, we first divide genes into clusters using the K-means approach and determine the optimal number of clusters using the Gap method. The supervised group Lasso consists of two steps. In the first step, we identify important genes within each cluster using the Lasso method. In the second step, we select important clusters using the group Lasso. Tuning parameters are determined using V-fold cross validation at both steps to allow for further flexibility. Prediction performance is evaluated using leave-one-out cross validation. We apply the proposed method to disease classification and survival analysis with microarray data. CONCLUSION: We analyze four microarray data sets using the proposed approach: two cancer data sets with binary cancer occurrence as outcomes and two lymphoma data sets with survival outcomes. The results show that the proposed approach is capable of identifying a small number of influential gene clusters and important genes within those clusters, and has better prediction performance than existing methods. Shuangge Ma, Jian Huang 0003 |
BMC Bioinform. | 1 |
| 2006 | Empirical study of supervised gene screeningabstractBACKGROUND: Microarray studies provide a way of linking variations of phenotypes with their genetic causations. Constructing predictive models using high dimensional microarray measurements usually consists of three steps: (1) unsupervised gene screening; (2) supervised gene screening; and (3) statistical model building. Supervised gene screening based on marginal gene ranking is commonly used to reduce the number of genes in the model building. Various simple statistics, such as t-statistic or signal to noise ratio, have been used to rank genes in the supervised screening. Despite of its extensive usage, statistical study of supervised gene screening remains scarce. Our study is partly motivated by the differences in gene discovery results caused by using different supervised gene screening methods. RESULTS: We investigate concordance and reproducibility of supervised gene screening based on eight commonly used marginal statistics. Concordance is assessed by the relative fractions of overlaps between top ranked genes screened using different marginal statistics. We propose a Bootstrap Reproducibility Index, which measures reproducibility of individual genes under the supervised screening. Empirical studies are based on four public microarray data. We consider the cases where the top 20%, 40% and 60% genes are screened. CONCLUSION: From a gene discovery point of view, the effect of supervised gene screening based on different marginal statistics cannot be ignored. Empirical studies show that (1) genes passed different supervised screenings may be considerably different; (2) concordance may vary, depending on the underlying data structure and percentage of selected genes; (3) evaluated with the Bootstrap Reproducibility Index, genes passed supervised screenings are only moderately reproducible; and (4) concordance cannot be improved by supervised screening based on reproducibility. Shuangge Ma |
BMC Bioinform. | 1 |
| 2006 | Regularized binormal ROC method in disease classificationusing microarray dataabstractBACKGROUND: An important application of microarrays is to discover genomic biomarkers, among tens of thousands of genes assayed, for disease diagnosis and prognosis. Thus it is of interest to develop efficient statistical methods that can simultaneously identify important biomarkers from such high-throughput genomic data and construct appropriate classification rules. It is also of interest to develop methods for evaluation of classification performance and ranking of identified biomarkers. RESULTS: The ROC (receiver operating characteristic) technique has been widely used in disease classification with low dimensional biomarkers. Compared with the empirical ROC approach, the binormal ROC is computationally more affordable and robust in small sample size cases. We propose using the binormal AUC (area under the ROC curve) as the objective function for two-sample classification, and the scaled threshold gradient directed regularization method for regularized estimation and biomarker selection. Tuning parameter selection is based on V-fold cross validation. We develop Monte Carlo based methods for evaluating the stability of individual biomarkers and overall prediction performance. Extensive simulation studies show that the proposed approach can generate parsimonious models with excellent classification and prediction performance, under most simulated scenarios including model mis-specification. Application of the method to two cancer studies shows that the identified genes are reasonably stable with satisfactory prediction performance and biologically sound implications. The overall classification performance is satisfactory, with small classification errors and large AUCs. CONCLUSION: In comparison to existing methods, the proposed approach is computationally more affordable without losing the optimality possessed by the standard ROC method. Shuangge Ma, Jian Huang 0003 |
BMC Bioinform. | 1 |
| 2005 | Regularized ROC method for disease classification and biomarker selection with microarray dataabstractMOTIVATION: An important application of microarrays is to discover genomic biomarkers, among tens of thousands of genes assayed, for disease classification. Thus there is a need for developing statistical methods that can efficiently use such high-throughput genomic data, select biomarkers with discriminant power and construct classification rules. The ROC (receiver operator characteristic) technique has been widely used in disease classification with low-dimensional biomarkers because (1) it does not assume a parametric form of the class probability as required for example in the logistic regression method; (2) it accommodates case-control designs and (3) it allows treating false positives and false negatives differently. However, due to computational difficulties, the ROC-based classification has not been used with microarray data. Moreover, the standard ROC technique does not incorporate built-in biomarker selection. RESULTS: We propose a novel method for biomarker selection and classification using the ROC technique for microarray data. The proposed method uses a sigmoid approximation to the area under the ROC curve as the objective function for classification and the threshold gradient descent regularization method for estimation and biomarker selection. Tuning parameter selection based on the V-fold cross validation and predictive performance evaluation are also investigated. The proposed approach is demonstrated with a simulation study, the Colon data and the Estrogen data. The proposed approach yields parsimonious models with excellent classification performance. Shuangge Ma, Jian Huang 0003 |
Bioinform. | 1 |