Taesung Park

dblp:55/4543 · DBLP profile ↗
← Back
88ranked-venue papers
6as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 67 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 VideoGigaGAN: Towards Detail-rich Video Super-Resolution
abstract
Video super-resolution (VSR) models achieve temporal consistency but often produce blurrier results than their image-based counterparts due to limited generative capacity. This prompts the question: can we adapt a generative image upsampler for VSR while preserving temporal consistency? We introduce VideoGigaGAN, a new generative VSR model that combines high-frequency detail with temporal stability, building on the large-scale GigaGAN image upsampler. Simple adaptations of GigaGAN for VSR led to flickering issues, so we propose techniques to enhance temporal consistency. We validate the effectiveness of VideoGigaGAN by comparing it with state-of-the-art VSR models on public datasets and showcasing video results with 8× upsampling.
Taesung Park, Richard Zhang 0001, Yang Zhou 0009, Eli Shechtman, Feng Liu 0015, Jia-Bin Huang 0001, Difan Liu
CVPR2
2025 Expressive Image Generation and Editing with Rich Text
Songwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin Huang 0001
Int. J. Comput. Vis.2
2024 Confidence Interval Estimation for Machine Learning Models in Forecasting Infectious Diseases
abstract
Forecasting models have been instrumental in managing the COVID-19 pandemic by facilitating efficient resource distribution and proper interventions. Although various forecasting models—such as mathematical models, statistical models, and machine learning models—are available, a few models account for the uncertainty in their predictions. The lack of uncertainty information constrains the reliability of forecasting results, making them less applicable for use in decision-making. Confidence Intervals (CIs) are widely used in statistical inference to provide uncertainty information. In this study, we introduce a framework that provides bootstrap-based CIs readily applicable to various forecasting models. The key concept of this framework is repeated training on bootstrap data based on resampled residuals from the forecasting model. After fitting the forecasting model, bootstrap datasets are generated considering time series characteristics to retrain the forecasting models. Then, CIs are calculated based on the distribution of prediction values. This framework was applied to forecasting COVID-19 outcomes—such as daily confirmed COVID-19 cases, deaths, and ICU patients—in South Korea. Our results demonstrate that bootstrap-based CIs can be successfully applied to provide additional information on the reliability of prediction. In conclusion, our framework can provide uncertainty information without requiring complex assumptions, regardless of the forecasting model or response variable type. Therefore, this approach can support public health policymakers in gaining a deeper understanding of epidemic trends, thereby offering valuable insights for decision-making.
Taewan Goo, Kyulhee Han, Hanbyul Song, Jooha Oh, Sayooj Aby Jose, Taesung Park
BIBM8
2024 Uncertainty Quantification and Statistical Inference for Biologically Informed Neural Networks
abstract
In precision medicine, deep learning models play an important role in identifying relevant biomarkers and making predictions for the early diagnosis and prognosis of diseases. While successful in developing accurate prediction models for prognosis and diagnosis, the ‘black box’ nature of deep learning poses significant challenges, which makes it difficult to have a meaningful biological interpretation. Recently, several biologically informed neural networks (BINNs), such as the Visible Neural Network and P-NET, have been proposed to incorporate biological knowledge for improved interpretation. These models assess node importance using quantitative measures such as SHAP and DeepLIFT. However, these approaches have limitations because they typically provide a single estimate of node importance without any information on the uncertainty of these estimates. This study aims to assess uncertainty of node importance through a novel statistical framework. Bootstrapping is used to calculate confidence interval of node importance. Our findings validate that the proposed statistical framework quantifies the uncertainty of nodes in BINNs. By providing robust confidence intervals of pathways and genes, this framework enhances the interpretability and reliability of BINNs.
Taewan Goo, Sangyeon Shin, Haeyoung Kim, Taesung Park
BIBM5
2024 One-Step Diffusion with Distribution Matching Distillation
abstract
Diffusion models generate high-quality images but require dozens of forward passes. We introduce Distribution Matching Distillation (DMD), a procedure to transform a diffusion model into a one-step image generator with minimal impact on image quality. We enforce the one-step image generator match the diffusion model at distribution level, by minimizing an approximate KL divergence whose gradient can be expressed as the difference between 2 score functions, one of the target distribution and the other of the synthetic distribution being produced by our one-step generator. The score functions are parameterized as two diffusion models trained separately on each distribution. Combined with a simple regression loss matching the large-scale structure of the multi-step diffusion outputs, our method outperforms all published few-step diffusion approaches, reaching 2.62 FID on ImageNet 64×64 and 11.49 FID on zero-shot COCO-30k, comparable to Stable Diffusion but orders of magnitude faster. Utilizing FP16 inference, our model can generate images at 20 FPS on modern hardware.
Tianwei Yin, Michaël Gharbi, Richard Zhang 0001, Eli Shechtman, Frédo Durand, William T. Freeman, Taesung Park
CVPR7
2024 Distilling Diffusion Models Into Conditional GANs
Minguk Kang, Richard Zhang 0001, Connelly Barnes, Sylvain Paris, Suha Kwak, Jaesik Park, Eli Shechtman, Jun-Yan Zhu, Taesung Park
ECCV (28)9
2024 Editable Image Elements for Controllable Synthesis
Jiteng Mu, Michaël Gharbi, Richard Zhang 0001, Eli Shechtman, Nuno Vasconcelos, Xiaolong Wang 0004, Taesung Park
ECCV (2)7
2024 Lazy Diffusion Transformer for Interactive Image Editing
Yotam Nitzan, Zongze Wu 0002, Richard Zhang 0001, Eli Shechtman, Daniel Cohen-Or, Taesung Park, Michaël Gharbi
ECCV (24)6
2024 Improved Distribution Matching Distillation for Fast Image Synthesis
abstract
Recent approaches have shown promises distilling expensive diffusion models into efficient one-step generators. Amongst them, Distribution Matching Distillation (DMD) produces one-step generators that match their teacher in distribution, i.e., the distillation process does not enforce a one-to-one correspondence with the sampling trajectories of their teachers. However, to ensure stable training in practice, DMD requires an additional regression loss computed using a large set of noise--image pairs, generated by the teacher with many steps of a deterministic sampler. This is not only computationally expensive for large-scale text-to-image synthesis, but it also limits the student's quality, tying it too closely to the teacher's original sampling paths. We introduce DMD2, a set of techniques that lift this limitation and improve DMD training. First, we eliminate the regression loss and the need for expensive dataset construction. We show that the resulting instability is due to the "fake" critic not estimating the distribution of generated samples with sufficient accuracy and propose a two time-scale update rule as a remedy. Second, we integrate a GAN loss into the distillation procedure, discriminating between generated samples and real images. This lets us train the student model on real data, thus mitigating the imperfect "real" score estimation from the teacher model, and thereby enhancing quality. Third, we introduce a new training procedure that enables multi-step sampling in the student, and addresses the training--inference input mismatch of previous work, by simulating inference-time generator samples during training. Taken together, our improvements set new benchmarks in one-step image generation, with FID scores of 1.28 on ImageNet-64×64 and 8.35 on zero-shot COCO 2014, surpassing the original teacher despite a 500X reduction in inference cost. Further, we show our approach can generate megapixel images by distilling SDXL, demonstrating exceptional visual quality among few-step methods, and surpassing the teacher. We release our code and pretrained models.
Tianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang 0001, Eli Shechtman, Frédo Durand, William T. Freeman
NeurIPS3
2024 Customizing Text-to-Image Diffusion with Object Viewpoint Control
Nupur Kumari, Grace Su, Richard Zhang 0001, Taesung Park, Eli Shechtman, Jun-Yan Zhu
SIGGRAPH Asia4
2023 Scaling up GANs for Text-to-Image Synthesis
abstract
The recent success of text-to-image synthesis has taken the world by storm and captured the general public's imagination. From a technical standpoint, it also marked a drastic change in the favored architecture to design generative image models. GANs used to be the de facto choice, with techniques like StyleGAN. With DALL.E 2, autoregressive and diffusion models became the new standard for large-scale generative models overnight. This rapid shift raises a fundamental question: can we scale up GANs to benefit from large datasets like LAION? We find that naïvely increasing the capacity of the StyleGan architecture quickly becomes unstable. We introduce GigaGAN, a new GAN architecture that far exceeds this limit, demonstrating GANs as a viable option for text-to-image synthesis. GigaGAN offers three major advantages. First, it is orders of magnitude faster at inference time, taking only 0.13 seconds to synthesize a 512px image. Second, it can synthesize high-resolution images, for example, 16-megapixel images in 3.66 seconds. Finally, GigaGAN supports various latent space editing applications such as latent interpolation, style mixing, and vector arithmetic operations.
Minguk Kang, Jun-Yan Zhu, Richard Zhang 0001, Jaesik Park, Eli Shechtman, Sylvain Paris, Taesung Park
CVPR7
2023 Domain Expansion of Image Generators
abstract
Can one inject new concepts into an already trained generative model, while respecting its existing structure and knowledge? We propose a new task - domain expansion - to address this. Given a pretrained generator and novel (but related) domains, we expand the generator to jointly model all domains, old and new, harmoniously. First, we note the generator contains a meaningful, pretrained latent space. Is it possible to minimally perturb this hard-earned representation, while maximally representing the new domains? Interestingly, we find that the latent space offers unused, “dormant” directions, which do not affect the output. This provides an opportunity: By “repurposing” these directions, we can represent new domains without perturbing the original representation. In fact, we find that pretrained generators have the capacity to add several- even hundreds - of new domains! Using our expansion method, one “expanded” model can supersede numerous domain-specific models, without expanding the model size. Additionally, a single expanded generator natively supports smooth transitions between domains, as well as composition of domains. Code and project page available here.
Yotam Nitzan, Michaël Gharbi, Richard Zhang 0001, Taesung Park, Jun-Yan Zhu, Daniel Cohen-Or, Eli Shechtman
CVPR4
2023 Expressive Text-to-Image Generation with Rich Text
abstract
Plain text has become a prevalent interface for text-to-image synthesis. However, its limited customization options hinder users from accurately describing desired outputs. For example, plain text makes it hard to specify continuous quantities, such as the precise RGB color value or importance of each word. Furthermore, creating detailed text prompts for complex scenes is tedious for humans to write and challenging for text encoders to interpret. To address these challenges, we propose using a rich-text editor supporting formats such as font style, size, color, and footnote. We extract each word’s attributes from rich text to enable local style control, explicit token reweighting, precise color rendering, and detailed region synthesis. We achieve these capabilities through a region-based diffusion process. We first obtain each word’s region based on attention maps of a diffusion process using plain text. For each region, we enforce its text attributes by creating region-specific detailed prompts and applying region-specific guidance, and maintain its fidelity against plain-text generation through region-based injections. We present various examples of image generation from rich text and demonstrate that our method outperforms strong baselines with quantitative evaluations.
Songwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin Huang 0001
ICCV2
2023 Holistic Evaluation of Text-to-Image Models
abstract
The stunning qualitative improvement of text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (HEIM). Whereas previous evaluations focus mostly on image-text alignment and image quality, we identify 12 aspects, including text-image alignment, image quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency. We curate 62 scenarios encompassing these aspects and evaluate 26 state-of-the-art text-to-image models on this benchmark. Our results reveal that no single model excels in all aspects, with different models demonstrating different strengths. We release the generated images and human evaluation results for full transparency at https://crfm.stanford.edu/heim/latest and the code at https://github.com/stanford-crfm/helm, which is integrated with the HELM codebase
Michihiro Yasunaga, Chenlin Meng, Yifan Mai 0001, Joon Sung Park 0001, Agrim Gupta, Deepak Narayanan, Hannah Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei 0001, Jiajun Wu 0001, Stefano Ermon, Percy Liang
NeurIPS12
2022 BlobGAN: Spatially Disentangled Scene Representations
Dave Epstein, Taesung Park, Richard Zhang 0001, Eli Shechtman, Alexei A. Efros
ECCV (15)2
2022 DeepHisCoM: deep learning pathway analysis using hierarchical structural component models
abstract
Many statistical methods for pathway analysis have been used to identify pathways associated with the disease along with biological factors such as genes and proteins. However, most pathway analysis methods neglect the complex nonlinear relationship between biological factors and pathways. In this study, we propose a Deep-learning pathway analysis using Hierarchical structured CoMponent models (DeepHisCoM) that utilize deep learning to consider a nonlinear complex contribution of biological factors to pathways by constructing a multilayered model which accounts for hierarchical biological structure. Through simulation studies, DeepHisCoM was shown to have a higher power in the nonlinear pathway effect and comparable power for the linear pathway effect when compared to the conventional pathway methods. Application to hepatocellular carcinoma (HCC) omics datasets, including metabolomic, transcriptomic and metagenomic datasets, demonstrated that DeepHisCoM successfully identified three well-known pathways that are highly associated with HCC, such as lysine degradation, valine, leucine and isoleucine biosynthesis and phenylalanine, tyrosine and tryptophan. Application to the coronavirus disease-2019 (COVID-19) single-nucleotide polymorphism (SNP) dataset also showed that DeepHisCoM identified four pathways that are highly associated with the severity of COVID-19, such as mitogen-activated protein kinase (MAPK) signaling pathway, gonadotropin-releasing hormone (GnRH) signaling pathway, hypertrophic cardiomyopathy and dilated cardiomyopathy. Codes are available at https://github.com/chanwoo-park-official/DeepHisCoM.
Chanwoo Park, Boram Kim, Taesung Park
Briefings Bioinform.3
2022 Kernel-based hierarchical structural component models for pathway analysis
abstract
MOTIVATION: Pathway analyses have led to more insight into the underlying biological functions related to the phenotype of interest in various types of omics data. Pathway-based statistical approaches have been actively developed, but most of them do not consider correlations among pathways. Because it is well known that there are quite a few biomarkers that overlap between pathways, these approaches may provide misleading results. In addition, most pathway-based approaches tend to assume that biomarkers within a pathway have linear associations with the phenotype of interest, even though the relationships are more complex. RESULTS: To model complex effects including non-linear effects, we propose a new approach, Hierarchical structural CoMponent analysis using Kernel (HisCoM-Kernel). The proposed method models non-linear associations between biomarkers and phenotype by extending the kernel machine regression and analyzes entire pathways simultaneously by using the biomarker-pathway hierarchical structure. HisCoM-Kernel is a flexible model that can be applied to various omics data. It was successfully applied to three omics datasets generated by different technologies. Our simulation studies showed that HisCoM-Kernel provided higher statistical power than other existing pathway-based methods in all datasets. The application of HisCoM-Kernel to three types of omics dataset showed its superior performance compared to existing methods in identifying more biologically meaningful pathways, including those reported in previous studies. AVAILABILITY AND IMPLEMENTATION: The HisCoM-Kernel software is freely available at http://statgen.snu.ac.kr/software/HisCom-Kernel/. The RNA-seq data underlying this article are available at https://xena.ucsc.edu/, and the others will be shared on reasonable request to the corresponding author. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Suhyun Hwangbo, Sungyoung Lee 0002, Seungyeoun Lee, Heungsun Hwang, Inyoung Kim, Taesung Park
Bioinform.6
2022 PATHOME-Drug: a subpathway-based polypharmacology drug-repositioning method
abstract
MOTIVATION: Drug repositioning reveals novel indications for existing drugs and in particular, diseases with no available drugs. Diverse computational drug repositioning methods have been proposed by measuring either drug-treated gene expression signatures or the proximity of drug targets and disease proteins found in prior networks. However, these methods do not explain which signaling subparts allow potential drugs to be selected, and do not consider polypharmacology, i.e. multiple targets of a known drug, in specific subparts. RESULTS: Here, to address the limitations, we developed a subpathway-based polypharmacology drug repositioning method, PATHOME-Drug, based on drug-associated transcriptomes. Specifically, this tool locates subparts of signaling cascading related to phenotype changes (e.g. disease status changes), and identifies existing approved drugs such that their multiple targets are enriched in the subparts. We show that our method demonstrated better performance for detecting signaling context and specific drugs/compounds, compared to WebGestalt and clusterProfiler, for both real biological and simulated datasets. We believe that our tool can successfully address the current shortage of targeted therapy agents. AVAILABILITY AND IMPLEMENTATION: The web-service is available at http://statgen.snu.ac.kr/software/pathome. The source codes and data are available at https://github.com/labnams/pathome-drug. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Seungyoon Nam, Sungyoung Lee 0002, Jinhyuk Lee, Aron Park, Yon Hui Kim, Taesung Park
Bioinform.7
2022 Hierarchical Structured Component Analysis for Microbiome Data Using Taxonomy Assignments
abstract
The recent advent of high-throughput sequencing technology has enabled us to study the associations between human microbiome and diseases. The DNA sequences of microbiome samples are clustered as operational taxonomic units (OTUs) according to their similarity. The OTU table containing counts of OTUs present in each sample is used to measure correlations between OTUs and disease status and find key microbes for prediction of the disease status. Various statistical methods have been proposed for such microbiome data analysis. However, none of these methods reflects the hierarchy of taxonomy information. In this paper, we propose a hierarchical structural component model for microbiome data (HisCoM-microb) using taxonomy information as well as OTU table data. The proposed HisCoM-microb consists of two layers: one for OTUs and the other for taxa at the higher taxonomy level. Then we calculate simultaneously coefficient estimates of OTUs and taxa of the two layers inserted in the hierarchical model. Through this analysis, we can infer the association between taxa or OTUs and disease status, considering the impact of taxonomic structure on disease status. Both simulation study and real microbiome data analysis show that HisCoM-microb can successfully reveal the relations between each taxon and disease status and identify the key OTUs of the disease at the same time.
Sun Ah Kim, Nayeon Kang, Taesung Park
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 ASSET: autoregressive semantic scene editing with transformers at high resolutions
abstract
We present ASSET, a neural architecture for automatically modifying an input high-resolution image according to a user's edits on its semantic segmentation map. Our architecture is based on a transformer with a novel attention mechanism. Our key idea is to sparsify the transformer's attention matrix at high resolutions, guided by dense attention extracted at lower image resolutions. While previous attention mechanisms are computationally too expensive for handling high-resolution images or are overly constrained within specific image regions hampering long-range interactions, our novel attention mechanism is both computationally efficient and effective. Our sparsified attention mechanism is able to capture long-range interactions and context, leading to synthesizing interesting phenomena in scenes, such as reflections of landscapes onto water or fora consistent with the rest of the landscape, that were not possible to generate reliably with previous convnets and transformer approaches. We present qualitative and quantitative results, along with user studies, demonstrating the effectiveness of our method. Our code and dataset are available at our project page: https://github.com/DifanLiu/ASSET
Difan Liu, Sandesh Shetty, Tobias Hinz, Matthew Fisher, Richard Zhang 0001, Taesung Park, Evangelos Kalogerakis
ACM Trans. Graph.6
2021 Spatial rank-based multifactor dimensionality reduction to detect gene-gene interactions for multivariate phenotypes
abstract
Abstract Background Identifying interaction effects between genes is one of the main tasks of genome-wide association studies aiming to shed light on the biological mechanisms underlying complex diseases. Multifactor dimensionality reduction (MDR) is a popular approach for detecting gene–gene interactions that has been extended in various forms to handle binary and continuous phenotypes. However, only few multivariate MDR methods are available for multiple related phenotypes. Current approaches use Hotelling’s T2statistic to evaluate interaction models, but it is well known that Hotelling’s T2statistic is highly sensitive to heavily skewed distributions and outliers. Results We propose a robust approach based on nonparametric statistics such as spatial signs and ranks. The new multivariate rank-based MDR (MR-MDR) is mainly suitable for analyzing multiple continuous phenotypes and is less sensitive to skewed distributions and outliers. MR-MDR utilizes fuzzy k-means clustering and classifies multi-locus genotypes into two groups. Then, MR-MDR calculates a spatial rank-sum statistic as an evaluation measure and selects the best interaction model with the largest statistic. Our novel idea lies in adopting nonparametric statistics as an evaluation measure for robust inference. We adopt tenfold cross-validation to avoid overfitting. Intensive simulation studies were conducted to compare the performance of MR-MDR with current methods. Application of MR-MDR to a real dataset from a Korean genome-wide association study demonstrated that it successfully identified genetic interactions associated with four phenotypes related to kidney function. The R code for conducting MR-MDR is available at https://github.com/statpark/MR-MDR . Conclusions Intensive simulation studies comparing MR-MDR with several current methods showed that the performance of MR-MDR was outstanding for skewed distributions. Additionally, for symmetric distributions, MR-MDR showed comparable power. Therefore, we conclude that MR-MDR is a useful multivariate non-parametric approach that can be used regardless of the phenotype distribution, the correlations between phenotypes, and sample size.
Mira Park 0002, Hoe-Bin Jeong, Taesung Park
BMC Bioinform.4
2020 Contrastive Learning for Unpaired Image-to-Image Translation
Taesung Park, Alexei A. Efros, Richard Zhang 0001, Jun-Yan Zhu
ECCV (9)1
2020 Swapping Autoencoder for Deep Image Manipulation
abstract
Deep generative models have become increasingly effective at producing realistic images from randomly sampled seeds, but using such models for controllable manipulation of existing images remains challenging. We propose the Swapping Autoencoder, a deep model designed specifically for image manipulation, rather than random sampling. The key idea is to encode an image into two independent components and enforce that any swapped combination maps to a realistic image. In particular, we encourage the components to represent structure and texture, by enforcing one component to encode co-occurrent patch statistics across different parts of the image. As our method is trained with an encoder, finding the latent codes for a new input image becomes trivial, rather than cumbersome. As a result, our method enables us to manipulate real input images in various ways, including texture swapping, local and global editing, and latent code vector arithmetic. Experiments on multiple datasets show that our model produces better results and is substantially more efficient compared to recent generative models.
Taesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu, Eli Shechtman, Alexei A. Efros, Richard Zhang 0001
NeurIPS1
2019 Two stage pattern clustering analysis in cross-over experimental design
abstract
In interventional studies, biomarkers such as metabolites, are usually measured across serial time points. And when the interest lies in comparing expression levels between different experimental conditions, summary measures such as area under curve (AUC), have been widely used. Although the summary measure based approaches have been successful in identifying novel biomarkers, they do not reveal anything about time-dependent changing patterns of biomarkers which can demonstrate the reactivity of biomarkers to various physiological conditions. To account for such patterns, all measurements across time points need to be used, and clustering analysis with the measurements can group together biomarkers having similar changing patterns. Some such popularly used clustering methods include hierarchical- and K-means clustering. While these may provide some well-clustered results, their patterns are quite dependent on input data sets, making it difficult to obtain consistent patterns across different interventional studies. In addition, it is problematic for these methods to discriminate biomarkers with weakly active patterns that need to be grouped as static, compared to those having strongly active patterns, when their patterns are highly similar. To address these issues, we propose a new clustering method for improving identification of changing patterns. Our approach is based on a two-stage process: the first is elimination of stable markers using Euclidean distances, while the second stage assigns the remaining biomarkers to predefined patterns using 1-correlation distance measure. By simulation studies, we showed that our proposed method had superior classification performances, compared to other unsupervised clustering methods. We expect that this approach can complement the existing summary measure based approaches.
Ik-Soo Huh, Sunghoon Choi, Youjin Kim, Soo-yeon Park, Oran Kwon, Taesung Park
BIBM6
2019 Identification of hyperparameters with high effects on performance of deep neural networks: application to clinicopathological data of ovarian cancer
abstract
Recent advances in deep learning have emerged as an effective approach for precision medicine. The applications of deep learning to medicine have been applied mainly to medical image data but not clinicopathological data. One of challenges of deep learning model to clinicopathological data is to optimize hyperparameters to get high predictive power. In this study, we identified hyperparameters of deep learning model that have large effects on power. Specifically, we focused on predicting platinum-based chemotherapy response for ovarian cancer patients. As a performance metric, we used the area under the curve. We optimized six hyperparameters: the number of hidden layers, number of hidden units, optimization algorithm, weight initialization, activation function, and dropout rate. We also identified significant interaction effects between hyperparameters. We successfully found the combination of hyperparameters having large effects on prediction. These optimal combinations are expected to increase the prediction accuracy for the response to chemotherapy for a variety of cancer patients.
Suhyun Hwangbo, Se Ik Kim, Untack Cho, Yongsang Song, Taesung Park
BIBM5
2019 Cancer Subtype Classification based on Superlayered Neural Network
abstract
Targeted treatment on different cancer subtype has been of clinical interest. However, accurate molecular identification of pathological cancer subtypes has been challenging. To meet such needs, a new subtype classification model based on deep neural network, Sparse Cross-modal Superlayered Neural Network is presented in this study. The model focuses on selecting biomarkers with considering integration of the high dimensional RNA sequencing data and DNA methylation data. For a real data application, the multi-omics data of lung adenocarcinoma and squamous cell carcinoma from The Cancer Genomic Atlas was fitted to the proposed model. Our model was compared to other existing methods such as principal component analysis, penalized logistic regression, and artificial neural network. With only a small number of biomarkers, the proposed model was able to effectively classify these lung cancer subtypes. Gene set analysis of selected biomarkers revealed significant difference in epidermis development and cornification pathway activation level between two lung cancer subtypes.
Prasoon Joshi, Seokho Jeong, Taesung Park
BIBM3
2019 Comparison of Statistical Models for Cross-over design
abstract
Cross-over designs have been widely used in clinical trials to investigate the efficacy of new treatments. In cross-over design, each subject is treated subsequently with different treatments. Many methods such as linear mixed models (LMMs) and generalized estimating equation (GEE) models have been used to analyze the repeated measurements from cross-over design. When we consider repeated measured response variables, estimation of random components for LMMs is not always easy. In this article, we applied the GEE method to cross-over design to overcome the limitation of LMMs. To apply the GEE model to the data from the cross-over designs, we need to switch the role of variables in LMM such a way that the independent variable in LMMs is considered as a response variable in GEE model and vice versa. The purpose of this study is to compare the performance of these GEE models and LMMs for cross-over designs. Through simulation studies, we checked the type I errors and compared power to evaluate the performance of the proposed GEE model and LMMs.
Yonggab Kim, Md. Kamruzzaman 0005, Yeni Lim, Oran Kwon, Taesung Park
BIBM5
2019 Semantic Image Synthesis With Spatially-Adaptive Normalization
abstract
We propose spatially-adaptive normalization, a simple but effective layer for synthesizing photorealistic images given an input semantic layout. Previous methods directly feed the semantic layout as input to the network, forcing the network to memorize the information throughout all the layers. Instead, we propose using the input layout for modulating the activations in normalization layers through a spatially-adaptive, learned affine transformation. Experiments on several challenging datasets demonstrate the superiority of our method compared to existing approaches, regarding both visual fidelity and alignment with input layouts. Finally, our model allows users to easily control the style and content of image synthesis results as well as create multi-modal results. Code is available upon publication.
Taesung Park, Ming-Yu Liu 0001, Ting-Chun Wang, Jun-Yan Zhu
CVPR1
2019 Detecting differential DNA methylation from sequencing of bisulfite converted DNA of diverse species
abstract
DNA methylation is one of the most extensively studied epigenetic modifications of genomic DNA. In recent years, sequencing of bisulfite-converted DNA, particularly via next-generation sequencing technologies, has become a widely popular method to study DNA methylation. This method can be readily applied to a variety of species, dramatically expanding the scope of DNA methylation studies beyond the traditionally studied human and mouse systems. In parallel to the increasing wealth of genomic methylation profiles, many statistical tools have been developed to detect differentially methylated loci (DMLs) or differentially methylated regions (DMRs) between biological conditions. We discuss and summarize several key properties of currently available tools to detect DMLs and DMRs from sequencing of bisulfite-converted DNA. However, the majority of the statistical tools developed for DML/DMR analyses have been validated using only mammalian data sets, and less priority has been placed on the analyses of invertebrate or plant DNA methylation data. We demonstrate that genomic methylation profiles of non-mammalian species are often highly distinct from those of mammalian species using examples of honey bees and humans. We then discuss how such differences in data properties may affect statistical analyses. Based on these differences, we provide three specific recommendations to improve the power and accuracy of DML and DMR analyses of invertebrate data when using currently available statistical tools. These considerations should facilitate systematic and robust analyses of DNA methylation from diverse species, thus advancing our understanding of DNA methylation.
Ik-Soo Huh, Taesung Park, Soojin V. Yi
Briefings Bioinform.3
2018 Detecting population structures by independent component analysis
Mira Park 0002, Eunbin Choi, Yongkang Kim, Taesung Park
BIBM4
2018 Deep Learning-based Identification of Cancer or Normal Tissue using Gene Expression Data
TaeJin Ahn, Taewan Goo, Sungmin Kim, Kyullhee Han, Sangick Park, Taesung Park
BIBM7
2018 Gene expression based prediction of prognostic outcome in ovarian cancer
TaeJin Ahn, Nayeon Kang, Yonggab Kim, Taesung Park
BIBM4
2018 Topological data analysis can extract subgroups with high incidence rates of Type 2 diabetes
Hyung Sun Kim, Changwoo Yi, Yongkang Kim, Uhnmee Park, Woong Kook, Hyuk Kim, Taesung Park
BIBM7
2018 An approximation method of extremely low p-values using permutation test
Sangseob Leem, Taesung Park
BIBM2
2018 CyCADA: Cycle-Consistent Adversarial Domain Adaptation
abstract
Domain adaptation is critical for success in new, unseen environments. Adversarial adaptation models have shown tremendous progress towards adapting to new environments by focusing either on discovering domain invariant representations or by mapping between unpaired image domains. While feature space methods are difficult to interpret and sometimes fail to capture pixel-level and low-level domain shifts, image space methods sometimes fail to incorporate high level semantic knowledge relevant for the end task. We propose a model which adapts between domains using both generative image space alignment and latent representation space alignment. Our approach, Cycle-Consistent Adversarial Domain Adaptation (CyCADA), guides transfer between domains according to a specific discriminatively trained task and avoids divergence by enforcing consistency of the relevant semantics before and after adaptation. We evaluate our method on a variety of visual recognition and prediction settings, including digit classification and semantic segmentation of road scenes, advancing state-of-the-art performance for unsupervised adaptation from synthetic to real world driving domains.
Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei A. Efros, Trevor Darrell
ICML3
2018 Drug response prediction model using a hierarchical structural component modeling method
abstract
BACKGROUND: Component-based structural equation modeling methods are now widely used in science, business, education, and other fields. This method uses unobservable variables, i.e., "latent" variables, and structural equation model relationships between observable variables. Here, we applied this structural equation modeling method to biologically structured data. To identify candidate drug-response biomarkers, we first used proteomic peptide-level data, as measured by multiple reaction monitoring mass spectrometry (MRM-MS), for liver cancer patients. MRM-MS is a highly sensitive and selective method for proteomic targeted quantitation of peptide abundances in complex biological samples. RESULTS: We developed a component-based drug response prediction model, having the advantage that it first combines collapsed peptide-level data into protein-level information, facilitating subsequent biological interpretation. Our model also uses an alternating least squares algorithm, to efficiently estimate both coefficients of peptides and proteins. This approach also considers correlations between variables, without constraint, by a multiple testing problem. Using estimated peptide and protein coefficients, we selected significant protein biomarkers by permutation testing, resulting in our model for predicting liver cancer response to the tyrosine kinase inhibitor sorafenib. CONCLUSIONS: Using data from a cohort of liver cancer patients, we then "fine-tuned" our model to successfully predict drug responses, as demonstrated by a high area under the curve (AUC) score. Such drug response prediction models may eventually find clinical translation in identifying individual patients likely to respond to specific therapies.
Sungtae Kim, Sungkyoung Choi, Jung-Hwan Yoon, Seungyeoun Lee, Taesung Park
BMC Bioinform.6
2018 Hierarchical structural component modeling of microRNA-mRNA integration analysis
abstract
BACKGROUND: Identification of multi-markers is one of the most challenging issues in personalized medicine era. Nowadays, many different types of omics data are generated from the same subject. Although many methods endeavor to identify candidate markers, for each type of omics data, few or none can facilitate such identification. RESULTS: It is well known that microRNAs affect phenotypes only indirectly, through regulating mRNA expression and/or protein translation. Toward addressing this issue, we suggest a hierarchical structured component analysis of microRNA-mRNA integration ("HisCoM-mimi") model that accounts for this biological relationship, to efficiently study and identify such integrated markers. In simulation studies, HisCoM-mimi showed the better performance than the other three methods. Also, in real data analysis, HisCoM-mimi successfully identified more gives more informative miRNA-mRNA integration sets relationships for pancreatic ductal adenocarcinoma (PDAC) diagnosis, compared to the other methods. CONCLUSION: As exemplified by an application to pancreatic cancer data, our proposed model effectively identified integrated miRNA/target mRNA pairs as markers for early diagnosis, providing a much broader biological interpretation.
Yongkang Kim, Sungyoung Lee 0002, Sungkyoung Choi, Jin-Young Jang, Taesung Park
BMC Bioinform.5
2018 Pathway-based approach using hierarchical components of rare variants to analyze multiple phenotypes
abstract
BACKGROUND: As one possible solution to the "missing heritability" problem, many methods have been proposed that apply pathway-based analyses, using rare variants that are detected by next generation sequencing technology. However, while a number of methods for pathway-based rare-variant analysis of multiple phenotypes have been proposed, no method considers a unified model that incorporate multiple pathways. RESULTS: Simulation studies successfully demonstrated advantages of multivariate analysis, compared to univariate analysis, and comparison studies showed the proposed approach to outperform existing methods. Moreover, real data analysis of six type 2 diabetes-related traits, using large-scale whole exome sequencing data, identified significant pathways that were not found by univariate analysis. Furthermore, strong relationships between the identified pathways, and their associated metabolic disorder risk factors, were found via literature search, and one of the identified pathway, was successfully replicated by an analysis with an independent dataset. CONCLUSIONS: Herein, we present a powerful, pathway-based approach to investigate associations between multiple pathways and multiple phenotypes. By reflecting the natural hierarchy of biological behavior, and considering correlation between pathways and phenotypes, the proposed method is capable of analyzing multiple phenotypes and multiple pathways simultaneously.
Sungyoung Lee 0002, Yongkang Kim, Sungkyoung Choi, Heungsun Hwang, Taesung Park
BMC Bioinform.5
2018 Functional conservation of sequence determinants at rapidly evolving regulatory regions across mammals
abstract
Recent advances in epigenomics have made it possible to map genome-wide regulatory regions using empirical methods. Subsequent comparative epigenomic studies have revealed that regulatory regions diverge rapidly between genome of different species, and that the divergence is more pronounced in enhancers than in promoters. To understand genomic changes underlying these patterns, we investigated if we can identify specific sequence fragments that are over-enriched in regulatory regions, thus potentially contributing to regulatory functions of such regions. Here we report numerous sequence fragments that are statistically over-enriched in enhancers and promoters of different mammals (which we refer to as 'sequence determinants'). Interestingly, the degree of statistical enrichment, which presumably is associated with the degree of regulatory impacts of the specific sequence determinant, was significantly higher for promoter sequence determinants than enhancer sequence determinants. We further used a machine learning method to construct prediction models using sequence determinants. Remarkably, prediction models constructed from one species could be used to predict regulatory regions of other species with high accuracy. This observation indicates that even though the precise locations of regulatory regions diverge rapidly during evolution, the functional potential of sequence determinants underlying regulatory sequences may be conserved between species.
Ik-Soo Huh, Isabel Mendizabal, Taesung Park, Soojin V. Yi
PLoS Comput. Biol.3
2018 Fuzzy heaping mechanism for heaped count data with imprecision
Hye-Young Jung 0001, Heawon Choi, Taesung Park
Soft Comput.3
2017 Risk prediction using common and rare genetic variants: Application to Type 2 diabetes
abstract
학위논문 (석사)-- 서울대학교 대학원 : 자연과학대학 협동과정 생물정보학전공, 2018. 2. 박태성.
Sunghwan Bae, Taesung Park
BIBM2
2017 Sample size calculation for comparing multiple groups in cross-over designs
abstract
In clinical research, computing sample size is important. A cross-over design (CD) is widely used for comparing treatment groups such as active treatment and control groups. In the analysis of multi-omics data such as metabolomics data, we often adopt the CD to identify biomarkers that have treatment effects. In this study, we propose a new method to compute the sample size of CDs with multiple treatment groups.
Taesung Park
BIBM2
2017 Cluster-based multifactor dimensionality reduction method to identify gene-gene interactions for quantitative traits in genome-wide studies
abstract
One of the main challenges in genome-wide association studies is finding gene-gene interactions for complex diseases. For binary traits, multifactor dimensionality reduction (MDR) has been proposed as a non-parametric approach for identifying gene-gene interactions. However, few statistical methods are currently available for determining the genetic interactions associated with quantitative traits such as blood pressure, body mass index, and patient survival times. To address this problem, we propose a new method termed clustering MDR (CL-MDR). In the CL-MDR method, we first apply a clustering technique as a dimensionality reduction step. By comparing the ratios of two classes in each multifactor cell to the ratio from the whole data, each genotype combination is classified as high-risk or low-risk. Evaluation processes are conducted as in the standard MDR method to identify the best combination for representing interactions between genes. The proposed method was investigated through simulation studies, which showed that CL-MDR successfully identified genetic interactions associated with a quantitative trait.
Yoojung Lee, Hyein Kim, Taesung Park, Mira Park 0002
BIBM3
2017 A comparison study of statistical methods for the analysis metagenome data
abstract
As many diseases are known to be related to microbes, interests in statistical methods for Microbiome-Wide Association Studies (MWAS) are also increasing. In this respect, we systematically investigate the properties of statistical methods for MWAS and compare their performances using simulation data generated from Human Microbiome Project data. We first assessed the type I error rates of eight commonly used methods over different levels of sparsity. Among these, ANCOM and metagenomeSeq methods yielded the well preserved type I error rates regardless of sparsity. In the power comparison study, metagenomeSeq showed higher power than ANCOM. In conclusion, we recommend using metagenomeSeq for the analysis of metagenome data.
Seungyeoun Lee, Taesung Park
BIBM3
2017 Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks
abstract
Image-to-image translation is a class of vision and graphics problems where the goal is to learn the mapping between an input image and an output image using a training set of aligned image pairs. However, for many tasks, paired training data will not be available. We present an approach for learning to translate an image from a source domain X to a target domain Y in the absence of paired examples. Our goal is to learn a mapping G : X → Y such that the distribution of images from G(X) is indistinguishable from the distribution Y using an adversarial loss. Because this mapping is highly under-constrained, we couple it with an inverse mapping F : Y → X and introduce a cycle consistency loss to push F(G(X)) ≈ X (and vice versa). Qualitative results are presented on several tasks where paired training data does not exist, including collection style transfer, object transfiguration, season transfer, photo enhancement, etc. Quantitative comparisons against several prior methods demonstrate the superiority of our approach.
Jun-Yan Zhu, Taesung Park, Phillip Isola, Alexei A. Efros
ICCV2
2016 Analysis for doubly repeated omics data from crossover design
abstract
Some crossover clinical trials produce doubly repeated omics data with two repeated factors. Linear mixed effect models (LMMs) are commonly applied to the data from the crossover design focusing on the analysis of repeatedly observed omics data themselves. Alternatively, the univariate analyses using the single summary measurements such as differences between time points and incremental area under curve (iAUC) are also widely used. In this study, we compare the performance of both methods for real doubly repeated omics data from a crossover study.
Sunghoon Choi, Soo-yeon Park, Hoejin Kim, Oran Kwon, Taesung Park
BIBM5
2016 Quality control plot for high dimensional omics data
abstract
Quality control (QC) becomes more important in pre-processing analysis of high dimensional omics data. Several routine QC processes became a standard process in omics data analysis. The standard QC analysis includes calculating quality-related measures, checking the consistency among samples, detecting outlying observations and so forth. QC analysis tends to be more important in the era of high dimensional omics data. Although several QC analysis tools providing simple graphical display have been developed by many researchers, they usually require a subjective decision on QC. Here, we propose high-dimensional data quality control (HidQC) plot which is a simple and efficient QC tool for handling high dimensional omics data. HidQC plot primarily focuses on identifying samples of poor quality by conducting a contrast analysis for the between/within group distances and the summary distances. HidQC plot checks the quality by investigating the consistency of samples for each group. Unlike other QC plots, HidQC plot provides the p-value of each sample based on the permutation test, which can be used as a more objective criterion to determine whether to use the sample or not. We applied HidQC plot to MicroArray Quality Control (MAQC) project 1 data to demonstrate its usefulness.
Gyu-Tae Kim, Yongkang Kim, Min-Seok Kwon, Taesung Park
BIBM4
2016 Multivariate approach to the analysis of correlated RNA-seq data
abstract
High-throughput RNA-seq technology has emerged as a powerful tool for understanding the molecular basis of phenotype variation in biology, including disease. Recently, some correlated RNA-seq datasets started to be generated. While there have been several approaches proposed for identifying the differentially expressed genes (DEGs), not many methods can analyze correlated RNA-seq data. We expect the simultaneous analysis of correlated RNA-seq data to increase of power of detecting DEGs. We propose a multivariate method to find DEGs on correlated RNA-seq data based on the Generalized Estimating Equations (GEE) approach. The advantage of the proposed method is to consider correlated RNA-seq data simultaneously while accounting for correlations. Through real data analysis and simulation studies, we show that our multivariate approach has higher power of detecting DEGs than the existing methods.
Hyunjin Park, Seungyeoun Lee, Ye Jin Kim, Myung-Sook Choi, Taesung Park
BIBM5
2016 Pathway-based approach using hierarchical components of collapsed rare variants
abstract
MOTIVATION: To address 'missing heritability' issue, many statistical methods for pathway-based analyses using rare variants have been proposed to analyze pathways individually. However, neglecting correlations between multiple pathways can result in misleading solutions, and pathway-based analyses of large-scale genetic datasets require massive computational burden. We propose a Pathway-based approach using HierArchical components of collapsed RAre variants Of High-throughput sequencing data (PHARAOH) for the analysis of rare variants by constructing a single hierarchical model that consists of collapsed gene-level summaries and pathways and analyzes entire pathways simultaneously by imposing ridge-type penalties on both gene and pathway coefficient estimates; hence our method considers the correlation of pathways without constraint by a multiple testing problem. RESULTS: Through simulation studies, the proposed method was shown to have higher statistical power than the existing pathway-based methods. In addition, our method was applied to the large-scale whole-exome sequencing data with levels of a liver enzyme using two well-known pathway databases Biocarta and KEGG. This application demonstrated that our method not only identified associated pathways but also successfully detected biologically plausible pathways for a phenotype of interest. These findings were successfully replicated by an independent large-scale exome chip study. AVAILABILITY AND IMPLEMENTATION: An implementation of PHARAOH is available at http://statgen.snu.ac.kr/software/pharaoh/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sungyoung Lee 0002, Sungkyoung Choi, Young Jin Kim 0008, Bong-Jo Kim, Heungsun Hwang, Taesung Park
Bioinform.6
2016 Gene-set association tests for next-generation sequencing data
abstract
MOTIVATION: Recently, many methods have been developed for conducting rare-variant association studies for sequencing data. These methods have primarily been based on gene-level associations but have not been proven to be as effective as expected. Gene-set-level tests have shown great advantages over gene-level tests in terms of power and robustness, because complex diseases are often caused by multiple genes that comprise of biological gene sets. RESULTS: Here, we propose several novel gene-set tests that employ rapid and efficient dimensionality reduction. The performance of these tests was investigated using extensive simulations and application to 1058 whole-exome sequences from a Korean population. We identified some known pathways and novel pathways whose rare or common variants are associated with elevated liver enzymes and replicated the results in an independent cohort. AVAILABILITY AND IMPLEMENTATION: Source R code for our algorithm is freely available at http://statgen.snu.ac.kr/software/QTest CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Young Jin Kim 0008, Bong-Jo Kim, Seungyeoun Lee, Taesung Park
Bioinform.6
2016 A unified model based multifactor dimensionality reduction framework for detecting gene-gene interactions
abstract
MOTIVATION: Gene-gene interaction (GGI) is one of the most popular approaches for finding and explaining the missing heritability of common complex traits in genome-wide association studies. The multifactor dimensionality reduction (MDR) method has been widely studied for detecting GGI effects. However, there are several disadvantages of the existing MDR-based approaches, such as the lack of an efficient way of evaluating the significance of multi-locus models and the high computational burden due to intensive permutation. Furthermore, the MDR method does not distinguish marginal effects from pure interaction effects. METHODS: We propose a two-step unified model based MDR approach (UM-MDR), in which, the significance of a multi-locus model, even a high-order model, can be easily obtained through a regression framework with a semi-parametric correction procedure for controlling Type I error rates. In comparison to the conventional permutation approach, the proposed semi-parametric correction procedure avoids heavy computation in order to achieve the significance of a multi-locus model. The proposed UM-MDR approach is flexible in the sense that it is able to incorporate different types of traits and evaluate significances of the existing MDR extensions. RESULTS: The simulation studies and the analysis of a real example are provided to demonstrate the utility of the proposed method. UM-MDR can achieve at least the same power as MDR for most scenarios, and it outperforms MDR especially when there are some single nucleotide polymorphisms that only have marginal effects, which masks the detection of causal epistasis for the existing MDR approaches. CONCLUSIONS: UM-MDR provides a very good supplement of existing MDR method due to its efficiency in achieving significance for every multi-locus model, its power and its flexibility of handling different types of traits. AVAILABILITY AND IMPLEMENTATION: A R package "umMDR" and other source codes are freely available at http://statgen.snu.ac.kr/software/umMDR/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wenbao Yu, Seungyeoun Lee, Taesung Park
Bioinform.3
2015 Genetic association tests for cigarattes per day
abstract
Cigarettes per day (CPD) is one of most commonly used phenotypes in nicotine dependence (ND) study. For example, a genetic association study of ND focuses on identifying single nucleotide polymorphisms (SNPs) that correlate with CPD. However, analysis of CPD data is always challenging, since the estimation of CPD distribution is difficult due to spikes at some specific values, say 10, 20 and so forth. Thus, standard maximum likelihood estimation is not appropriate. In this study, we focus on genetic association tests for identifying SNPs that correlate with CPD. We first reviewed previously proposed approaches applicable to CPD data, in which the CPD data is a dependent variable and the SNP is an ordinal independent variable. We then considered a calibration model in which the SNP is the ordinal dependent variable and the CPD is the independent variable. Unlike a standard modeling approach, this calibration modeling approach becomes sufficiently robust to accommodate distributional assumptions of CPD data. We applied our robust calibration modeling approach to CPD data from the Korean Association Resource project data of 4,183 male samples.
Haewon Choi, Hye-Young Jung 0001, Taesung Park
BIBM3
2015 Competitive pathway analysis using Structural Equation Models (CPA-SEM) for gene expression data
abstract
There is an increasing interest in the pathway analysis of multiple genes and complex traits in association studies. Recently, a number of methods of pathway analysis have been developed to detect the novel pathways associated with human complex traits. In this paper, we propose a novel statistical approach for competitive pathway analysis based on Structural Equation Modeling (CPA-SEM), taking advantage of prior knowledge on existing relationships between genes in a pathway. Our CPA-SEM identifies pathways associated with traits of interest. The CPA-SEM approach is different from the previous SEM-based approaches in that it considers all possible sub-pathways into account and performs permutation based robust analysis. We applied the proposed CPA-SEM method to gene expression data of gastric cancer (GSE27342), and found that mTOR signaling pathway was significantly associated with gastric cancer. This pathway has previously been reported to be associated with gastric cancer. In conclusion, our CPA-SEM analysis provides a better understanding of biological mechanism by identifying pathways associated with a trait of interest.
Sungkyoung Choi, Sungyoung Lee 0002, Ik-Soo Huh, Heungsun Hwang, Taesung Park
BIBM5
2015 A test for detecting differentially methylated regions
abstract
DNA methylation is one of the most important phenomena in epigenetics. Because the DNA methylation process regulates gene expression without altering the corresponding DNA sequences, it can explain many biological processes such as onset of cancer, aging, and differentiation of species. Therefore, detecting differentially methylated region (DMR) between groups has been considered one of important analyses in epigenetics. For the datasets generated by the methylated DNA immunoprecipitation (MeDIP) technique, DMR analysis has employed the previous methods developed for the microarray-based gene expression dataset. However, for the datasets generated by more recent bisulfite-sequencing (BS-seq) technologies, totally different approaches need to be used. Unfortunately, only a few methods are available for DMR analysis among which the Fisher's exact test is most widely used. We first propose a new DMR test for BS-seq data. Our approach employs Cochran-Mantel-Hantzel (CMH) test and extend it by incorporating correlation structures between adjacent CpG cytosine sites. Then we apply it to a real dataset of two subspecies of honeybee for detecting DMR regions.
Ik-Soo Huh, Soojin V. Yi, Taesung Park
BIBM3
2015 Developing cancer prediction model based on stepwise selection by AUC measure for proteomics data
abstract
Since most of the cancer markers that have been reported are obtained directly from cancer tissues, it is difficult to use them for early diagnosis of cancer without surgery. Thus, development of markers that can be detected by blood is crucial for making early diagnosis of cancer easier. One of the most feasible types of markers that can be detected by blood is a protein marker. Here, we focus on building prediction methods using the protein markers for early diagnosis of cancer. To develop a prediction model with high prediction ability, it is critical to choose appropriate markers first. Here, we consider a stepwise selection method using area under the receiver operating characteristic curve (Step-AUC) in order to construct a multi-protein prediction model. We showed that the performance of Step-AUC highly depends on the tuning parameter. We compared our proposed Step-AUC method to stepwise selection using information criteria and support vector machine recursive feature extraction (SVM-RFE). We observed that Step-AUC and stepwise selection using Bayesian information criteria (Step-BIC) perform better than other methods. The importance of each marker can be chosen using a new stepwise selection consistency (SSC) measure. The final models include the markers with high SSC measures. We applied our stepwise procedure to pancreatic cancer data and found two markers of interest.
Yongkang Kim, Seungyeoun Lee, Min-Seok Kwon, Ahrum Na, Yonghwan Choi, Sung-Gon Yi, Junghyun Namkung, Sangjo Han, Meejoo Kang, Sun Whe Kim, Jin-Young Jang, Yikwon Kim, Taesung Park
BIBM14
2015 VizEpis : A visualization and mapping tool for interpreting epistasis
abstract
An important issue in genetic association studies is gene-gene interaction (epistasis) underlying common complex diseases that are affected by multiple genetic factors. Although many methods have been proposed to analyze epistasis, the interpretation of the identified gene-gene interactions is not straightforward. In order to efficiently provide the statistical interpretation and biological evidences of gene-gene interactions, we developed the VizEpis, a tool for visualizing of gene-gene interactions in genetic association analysis and mapping of epistatic interaction to the biological evidence from public interaction databases. Using interaction network and circular plot, the VizEpis provides to explore the interaction network integrated with biological evidences in epigenetic regulation, splicing, transcription, translation and post-translation level. To aid statistical interaction in genotype level, the VizEpis provides pairwise checkerboard, forest, and funnel. VizEpis provides for the user-specified variants with statistical interaction to find biological evidences.
Min-Seok Kwon, Sungyoung Lee 0002, Yongkang Kim, Taesung Park
BIBM4
2015 Family-based association analysis: a fast and efficient method of multivariate association analysis with multiple variants
abstract
BACKGROUND: Many disease phenotypes are outcomes of the complicated interplay between multiple genes, and multiple phenotypes are affected by a single or multiple genotypes. Therefore, joint analysis of multiple phenotypes and multiple markers has been considered as an efficient strategy for genome-wide association analysis, and in this work we propose an omnibus family-based association test for the joint analysis of multiple genotypes and multiple phenotypes. RESULTS: The proposed test can be applied for both quantitative and dichotomous phenotypes, and it is robust under the presence of population substructure, as long as large-scale genomic data is available. Using simulated data, we showed that our method is statistically more efficient than the existing methods, and the practical relevance is illustrated by application of the approach to obesity-related phenotypes. CONCLUSIONS: The proposed method may be more statistically efficient than the existing methods. The application was developed in C++ and is available at the following URL: http://healthstat.snu.ac.kr/software/mfqls/ .
Sungho Won, Wonji Kim, Sungyoung Lee 0002, Joohon Sung, Taesung Park
BMC Bioinform.6
2014 Genome-wide association analysis with matched samples discloses additional novel risk loci
abstract
Genome-wide association studies have identified many candidate causal variants associated with common complex diseases and traits, but most of them have been drawn from nonrandomized case/control designs. In nonrandomized experiments, the results drawn from two different groups can be misleading because the units exposed to one group generally differ systematically from the units exposed to the other group. Propensity score is widely used to group case and control units for a more direct and significant comparison even with nonrandomized experiments. This propensity score matching can help with prioritizing additional uncovered variants on disease risk via sub-group analysis in genome-wide association studies. The aim of this work is to propose a post-hoc association test based on the subsets of samples. For that purpose, this paper presents a new paradigm for a post-hoc genome-wide association test when the sample size of controls are larger than that of cases: selecting control samples by equating the distribution of covariates in the case and control groups and re-performing association analysis upon these matched samples. We demonstrated the feasibility of this approach by applying it to 2752 type II diabetes patients in 8842 Korean population. Genome-wide association approach with matched samples is able to disclose 9 additional novel variants and 7 out of 9 have not identified from the association test of whole control samples. The process described here can successfully be combined with other types of case/control studies with large covariate information. This indicates that there a possibility of obtaining additional candidate causal variants responsible for common diseases through genome-wide association analysis with matched samples.
Jungsoo Gim, Sungkyoung Choi, Jongho Im, Taesung Park
BIBM5
2014 A normalization method for multiple reaction monitoring (MRM) data
abstract
High accuracy and reproducibility provided by selected or multiple reaction monitoring (SRM/MRM) experiments suggest that it is likely to become the platform of choice for identifying and verifying candidate biomarkers. Although several methods have been developed to quantify and discover the significant changes in SRM/MRM expression, we show normalization continues to be an essential step in the differential expression analysis. Here, we first investigate what makes normalization across the samples difficult in SRM assay. Then, we suggest a simple but effective normalization method to overcome the difficulty. Both real and simulated data sets are examined and improved results, compared to an existing method, are shown for inferring differential expression.
Jungsoo Gim, Taesung Park
BIBM2
2014 Multifactor dimendionality reduction analysis for gene-gene interaction of multiple binary traits
abstract
Recent advances in genotyping technology have facilitated the use of genome-wide association studies (GWAS) to successfully identify genetic variants that are associated with common complex traits. Following the successes in identification of single variants, joint identification including gene-gene interaction has been studied vigorously and produced many novel results. However, most genome-wide association studies have been conducted by focusing on one trait of interest for identifying genetic variants associated with common complex traits. Since many complex diseases having severe influences on the public health are pleiotropic, simple univariate analysis focusing on a single trait does not well detect full genetic architecture of complex diseases. For example, hyperlipidemia is diagnosed by four multiple traits: Total cholesterol (Tchl), High density lipoprotein (HDL) cholesterol, Low density lipoprotein (LDL), and cholesterol and Triglycerides (TG). Surprisingly, however, only few studies handle multiple traits simultaneously so far. Therefore, in order to improve power and reflect biological association more expansively, we investigate a multivariate approach which considers multiple traits simultaneously. Especially for the gene-gene interaction analysis for the multiple traits, we extend original multifactor dimensionality reduction (MDR) to handle multiple traits. We then demonstrate its superiority to univariate analysis through simulation studies. We confirm that the multivariate approach provides more stable and precise accuracy measures compared to univariate analysis. We applied the multivariate MDR approach to a GWA dataset of 8,842 Korean individuals and detected genetic variants associated with hypertension traits using systolic blood pressure (SBP) and Diastolic blood pressure (DBP).
Ik-Soo Huh, Taesung Park
BIBM2
2014 Identifying differentially expressed genes for ordinal phenotypes
abstract
A popular goal of microarray analysis is identification of differentially expressed genes (DEGs) between groups, which usually involves two-group comparisons. Many statistical methods have been developed toward this end, such as the t-test and the permutation test. In some cases, more than two groups of interest may be compared, for example, in identification of DEGs across three or four different stages of cancer or across different stages of the cell cycle. Several statistical approaches are also available for such multi-group analyses, including analysis of variance (ANOVA) models. We hypothesized that statistical methods developed for identifying DEGs for ordered groups would provide higher power for such ordered information. Although there are some methods available for ordered group comparisons, they have been rarely applied to the analysis of microarray data. In this paper, we consider various statistical tests for identifying DEGs in comparisons involving more than two groups with ordered information (i.e., cancer stage and cell cycle data). We first consider a constraint ANOVA (CANOVA) model by extending an ANOVA model without using order information, and then employ a proportional odds (PO) model by extending a general logit model. Finally, a simple correlation-based approach is considered. Through extensive simulation studies, we evaluated the performance of the CANOVA, PO, and correlation approaches by comparing the sizes and powers of these methods. The CANOVA, PO, and correlation approaches were applied to real microarray data of The Cancer Genome Atlas (TCGA). We specifically focused on the acute myeloid leukemia (AML) mRNA microarray data set and considered the results of cytogenetic analyses as group information of AML. To identify the genes related to these risk categories, we selected 25 good samples, 25 intermediate samples, and 25 poor samples in the TCGA data set.
Yongkang Kim, Taesung Park
BIBM2
2014 Biomarker development for pancreatic ductal adenocarcinoma using integrated analysis of mRNA and miRNA expression
abstract
Pancreatic ductal adenocarcinoma (PDAC) is the most common type of pancreatic cancer, which has dismal prognosis because of its silent early symptoms, high metastatic potential, and resistance to conventional therapies. Although a PDAC patient who is diagnosed at an early stage would have a substantial increase in chance of survival, the survival rate is poor because there is no efficient non-invasive diagnostic test in the early stage. In this study, we developed an efficient prediction models to detect PDAC in its early stages. Our prediction models use both mRNA and miRNA expression data from 104 PDAC tissues and 17 normal pancreatic tissues using microarray technology. After quality control, we built prediction models based on support vector machine (SVM) from mRNA and miRNA expressions for detecting early PDAC. To prevent over-fitting effect, we conducted leave-one-out cross validation (LOOCV) and 5-fold cross validation (CV). For independent validation of prediction models, we performed evaluation on independent datasets from Gene Expression Omnibus (GEO). After the validation, we identified 28 single markers and 231 combinations of markers with powerful prediction performance. In addition, the marker candidates are annotated with cancer pathways using gene ontology analysis. Our prediction models for PDAC may have potential for early diagnosis of PDAC.
Min-Seok Kwon, Yongkang Kim, Seungyeoun Lee, Junghyun Namkung, Taegyun Yun, Sung-Gon Yi, Sangjo Han, Meejoo Kang, Sun Whe Kim, Jin-Young Jang, Taesung Park
BIBM11
2014 Estimating cancer gene pathway proximity using network interaction
abstract
It is known that the gene level aberrations for a given cancer could vary across patients. As a result, a single therapy may not be suitable for every patient. However, these genetic aberrations may occur in similar pathways across patients. Therefore a study at pathway/subnetwork is more effective than at gene level. In this paper, we propose a method at this level to classify pathways (sub-networks) as functionally coupled and functionally independent. For this, we propose novel interaction measures. We show how these can be used to link and classify subnetworks using breast cancer as an example. Such methods will play an important role in patient stratification in order to develop personalized treatment options.
Rama Srikanth Mallavarapu, TaeJin Ahn, Subhankar Mukherjee 0001, Ajit S. Bopardikar, Garima Agarwal, Taesung Park
BIBM6
2014 Cancer survival classification using integrated data sets and intermediate information
Shinuk Kim, Taesung Park, Mark Kon
Artif. Intell. Medicine2
2014 Personalized identification of altered pathways in cancer using accumulated normal tissue data
abstract
MOTIVATION: Identifying altered pathways in an individual is important for understanding disease mechanisms and for the future application of custom therapeutic decisions. Existing pathway analysis techniques are mainly focused on discovering altered pathways between normal and cancer groups and are not suitable for identifying the pathway aberrance that may occur in an individual sample. A simple way to identify individual's pathway aberrance is to compare normal and tumor data from the same individual. However, the matched normal data from the same individual are often unavailable in clinical situation. Therefore, we suggest a new approach for the personalized identification of altered pathways, making special use of accumulated normal data in cases when a patient's matched normal data are unavailable. The philosophy behind our method is to quantify the aberrance of an individual sample's pathway by comparing it with accumulated normal samples. We propose and examine personalized extensions of pathway statistics, overrepresentation analysis and functional class scoring, to generate individualized pathway aberrance score. RESULTS: Collected microarray data of normal tissue of lung and colon mucosa are served as reference to investigate a number of cancer individuals of lung adenocarcinoma (LUAD) and colon cancer, respectively. Our method concurrently captures known facts of cancer survival pathways and identifies the pathway aberrances that represent cancer differentiation status and survival. It also provides more improved validation rate of survival-related pathways than when a single cancer sample is interpreted in the context of cancer-only cohort. In addition, our method is useful in classifying unknown samples into cancer or normal groups. Particularly, we identified 'amino acid synthesis and interconversion' pathway is a good indicator of LUAD (Area Under the Curve (AUC) 0.982 at independent validation). Clinical importance of the method is providing pathway interpretation of single cancer, even though its matched normal data are unavailable. AVAILABILITY AND IMPLEMENTATION: The method was implemented using the R software, available at our Web site: http://bibs.snu.ac.kr/ipas. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
TaeJin Ahn, Eunjin Lee, Nam Huh, Taesung Park
Bioinform.4
2014 FARVAT: a family-based rare variant association test
abstract
MOTIVATION: Individuals in each family are genetically more homogeneous than unrelated individuals, and family-based designs are often recommended for the analysis of rare variants. However, despite the importance of family-based samples analysis, few statistical methods for rare variant association analysis are available. RESULTS: In this report, we propose a FAmily-based Rare Variant Association Test (FARVAT). FARVAT is based on the quasi-likelihood of whole families, and is statistically and computationally efficient for the extended families. FARVAT assumed that families were ascertained with the disease status of family members, and incorporation of the estimated genetic relationship matrix to the proposed method provided robustness under the presence of the population substructure. Depending on the choice of working matrix, our method could be a burden test or a variance component test, and could be extended to the SKAT-O-type statistic. FARVAT was implemented in C++, and application of the proposed method to schizophrenia data and simulated data for GAW17 illustrated its practical importance. AVAILABILITY: The software calculates various statistics for the analysis of related samples, and it is freely downloadable from http://healthstats.snu.ac.kr/software/farvat. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: supplementary data are available at Bioinformatics online.
Sungkyoung Choi, Sungyoung Lee 0002, Sven Cichon, Markus M. Nöthen, Christoph Lange 0003, Taesung Park, Sungho Won
Bioinform.6
2013 Pathway based identification of rare mutation effect in cancer
abstract
Identifying driver mutation is important in understanding disease mechanism and future application of custom tailored therapeutic decision. Functional analysis of mutational impact usually focuses on the gene expression level of the mutated gene. However, complex regulatory network may cause differential gene expression among functional neighbors of the mutated gene. We suggest a new approach for discovering rare mutations that have real impact in the context of pathway; philosophy of our method is iteratively combining rare mutations until no more mutations can be added, under the condition that the combined mutational event can statistically discriminate pathway level mRNA expression between groups with and without mutation events. Our approach is shown to sensitively capture mutations that change pathway level gene expression at breast cancer data.
TaeJin Ahn, Taesung Park
BIBM2
2012 Gene-gene interaction analysis for the survival phenotype based on the Cox model
abstract
MOTIVATION: For the past few decades, many statistical methods in genome-wide association studies (GWAS) have been developed to identify SNP-SNP interactions for case-control studies. However, there has been less work for prospective cohort studies, involving the survival time. Recently, Gui et al. (2011) proposed a novel method, called Surv-MDR, for detecting gene-gene interactions associated with survival time. Surv-MDR is an extension of the multifactor dimensionality reduction (MDR) method to the survival phenotype by using the log-rank test for defining a binary attribute. However, the Surv-MDR method has some drawbacks in the sense that it needs more intensive computations and does not allow for a covariate adjustment. In this article, we propose a new approach, called Cox-MDR, which is an extension of the generalized multifactor dimensionality reduction (GMDR) to the survival phenotype by using a martingale residual as a score to classify multi-level genotypes as high- and low-risk groups. The advantages of Cox-MDR over Surv-MDR are to allow for the effects of discrete and quantitative covariates in the frame of Cox regression model and to require less computation than Surv-MDR. RESULTS: Through simulation studies, we compared the power of Cox-MDR with those of Surv-MDR and Cox regression model for various heritability and minor allele frequency combinations without and with adjusting for covariate. We found that Cox-MDR and Cox regression model perform better than Surv-MDR for low minor allele frequency of 0.2, but Surv-MDR has high power for minor allele frequency of 0.4. However, when the effect of covariate is adjusted for, Cox-MDR and Cox regression model perform much better than Surv-MDR. We also compared the performance of Cox-MDR and Surv-MDR for a real data of leukemia patients to detect the gene-gene interactions with the survival time. CONTACT: [email protected]; [email protected].
Seungyeoun Lee, Min-Seok Kwon, Jung Mi Oh, Taesung Park
Bioinform.4
2012 A novel method to identify high order gene-gene interactions in genome-wide association studies: Gene-based MDR
abstract
BACKGROUND: Because common complex diseases are affected by multiple genes and environmental factors, it is essential to investigate gene-gene and/or gene-environment interactions to understand genetic architecture of complex diseases. After the great success of large scale genome-wide association (GWA) studies using the high density single nucleotide polymorphism (SNP) chips, the study of gene-gene interaction becomes a next challenge. Multifactor dimensionality reduction (MDR) analysis has been widely used for the gene-gene interaction analysis. In practice, however, it is not easy to perform high order gene-gene interaction analyses via MDR in genome-wide level because it requires exploring a huge search space and suffers from a computational burden due to high dimensionality. RESULTS: We propose dimensional reduction analysis, Gene-MDR analysis for the fast and efficient high order gene-gene interaction analysis. The proposed Gene-MDR method is composed of two-step applications of MDR: within- and between-gene MDR analyses. First, within-gene MDR analysis summarizes each gene effect via MDR analysis by combining multiple SNPs from the same gene. Second, between-gene MDR analysis then performs interaction analysis using the summarized gene effects from within-gene MDR analysis. We apply the Gene-MDR method to bipolar disorder (BD) GWA data from Wellcome Trust Case Control Consortium (WTCCC). The results demonstrate that Gene-MDR is capable of detecting high order gene-gene interactions associated with BD. CONCLUSION: By reducing the dimension of genome-wide data from SNP level to gene level, Gene-MDR efficiently identifies high order gene-gene interactions. Therefore, Gene-MDR can provide the key to understand complex disease etiology.
Sohee Oh, Min-Seok Kwon, Bruce Weir, Kyooseob Ha, Taesung Park
BMC Bioinform.6
2011 Efficient and Fast Analysis for Detecting High Order Gene-by-Gene Interactions in a Genome-Wide Association Study
abstract
Most common complex traits are affected by multiple genes and/or environmental factors. To understand genetic architecture of complex traits, the investigation of gene-gene and gene-environment interactions can be essential. However, conducting gene-gene interaction using genome-wide data requires exploring a huge search space and suffers from a computation burden due to high dimensionality of genetic data. To identify gene-gene interaction more efficiently, we propose a gene-based reduction method which first summarizes the gene effect by combining multiple single nucleotide polymorphism (SNP) and then performs the gene-gene interaction via the summarized gene effect. By reducing the search space from SNPs to gene, our gene-based method becomes efficient and fast for identifying gene-gene interaction in genome wide association studies. The gene-based reduction method is illustrated by hypertension data from a Korean population.
Sohee Oh, Min-Seok Kwon, Kyunga Kim, Taesung Park
BIBM5
2011 Enhanced peptide quantification using spectral count clustering and cluster abundance
abstract
BACKGROUND: Quantification of protein expression by means of mass spectrometry (MS) has been introduced in various proteomics studies. In particular, two label-free quantification methods, such as spectral counting and spectra feature analysis have been extensively investigated in a wide variety of proteomic studies. The cornerstone of both methods is peptide identification based on a proteomic database search and subsequent estimation of peptide retention time. However, they often suffer from restrictive database search and inaccurate estimation of the liquid chromatography (LC) retention time. Furthermore, conventional peptide identification methods based on the spectral library search algorithms such as SEQUEST or SpectraST have been found to provide neither the best match nor high-scored matches. Lastly, these methods are limited in the sense that target peptides cannot be identified unless they have been previously generated and stored into the database or spectral libraries.To overcome these limitations, we propose a novel method, namely Quantification method based on Finding the Identical Spectral set for a Homogenous peptide (Q-FISH) to estimate the peptide's abundance from its tandem mass spectrometry (MS/MS) spectra through the direct comparison of experimental spectra. Intuitively, our Q-FISH method compares all possible pairs of experimental spectra in order to identify both known and novel proteins, significantly enhancing identification accuracy by grouping replicated spectra from the same peptide targets. RESULTS: We applied Q-FISH to Nano-LC-MS/MS data obtained from human hepatocellular carcinoma (HCC) and normal liver tissue samples to identify differentially expressed peptides between the normal and disease samples. For a total of 44,318 spectra obtained through MS/MS analysis, Q-FISH yielded 14,747 clusters. Among these, 5,777 clusters were identified only in the HCC sample, 6,648 clusters only in the normal tissue sample, and 2,323 clusters both in the HCC and normal tissue samples. While it will be interesting to investigate peptide clusters only found from one sample, further examined spectral clusters identified both in the HCC and normal samples since our goal is to identify and assess differentially expressed peptides quantitatively. The next step was to perform a beta-binomial test to isolate differentially expressed peptides between the HCC and normal tissue samples. This test resulted in 84 peptides with significantly differential spectral counts between the HCC and normal tissue samples. We independently identified 50 and 95 peptides by SEQUEST, of which 24 and 56 peptides, respectively, were found to be known biomarkers for the human liver cancer. Comparing Q-FISH and SEQUEST results, we found 22 of the differentially expressed 84 peptides by Q-FISH were also identified by SEQUEST. Remarkably, of these 22 peptides discovered both by Q-FISH and SEQUEST, 13 peptides are known for human liver cancer and the remaining 9 peptides are known to be associated with other cancers. CONCLUSIONS: We proposed a novel statistical method, Q-FISH, for accurately identifying protein species and simultaneously quantifying the expression levels of identified peptides from mass spectrometry data. Q-FISH analysis on human HCC and liver tissue samples identified many protein biomarkers that are highly relevant to HCC. Q-FISH can be a useful tool both for peptide identification and quantification on mass spectrometry data analysis. It may also prove to be more effective in discovering novel protein biomarkers than SEQUEST and other standard methods.
Seungmook Lee, Min-Seok Kwon, Hyoung-Joo Lee, Young-Ki Paik, Haixu Tang, Jae K. Lee, Taesung Park
BMC Bioinform.7
2011 Integrated analysis of the heterogeneous microarray data
abstract
BACKGROUND: As the magnitude of the experiment increases, it is common to combine various types of microarrays such as paired and non-paired microarrays from different laboratories or hospitals. Thus, it is important to analyze microarray data together to derive a combined conclusion after accounting for heterogeneity among data sets. One of the main objectives of the microarray experiment is to identify differentially expressed genes among the different experimental groups. We propose the linear mixed effect model for the integrated analysis of the heterogeneous microarray data sets. RESULTS: The proposed linear mixed effect model was illustrated using the data from 133 microarrays collected at three different hospitals. Though simulation studies, we compared the proposed linear mixed effect model approach with the meta-analysis and the ANOVA model approaches. The linear mixed effect model approach was shown to provide higher powers than the other approaches. CONCLUSIONS: The linear mixed effect model has advantages of allowing for various types of covariance structures over ANOVA model. Further, it can handle easily the correlated microarray data such as paired microarray data and repeated microarray data from the same subject.
Sung Yi, Taesung Park
BMC Bioinform.2
2010 Integrated analysis of the various types of microarray data using linear-mixed effects models
abstract
As the magnitude of the experiment increases, it is common to combine various types of microarrays such as paired and non-paired microarrays from different laboratories or hospitals. Thus, it is important to analyze microarray data together to derive a combined conclusion after accounting for heterogeneity among data sets. One of the main objectives of the microarray experiment is to identify differentially expressed genes among the different experimental groups. We propose the linear-mixed effect model for the integrated analysis of the heterogeneous microarray data sets. The proposed LMe model was illustrated using the data from 133 microarrays collected at three different hospitals. Though simulation studies, we compared the proposed LMe model approach with the meta-analysis and the ANOVA model approaches. The LMe model approach was shown to provide higher powers than the other approaches.
Sung-Gon Yi, Taesung Park
BIBM2
2010 Global analysis of microarray data reveals intrinsic properties in gene expression and tissue selectivity
abstract
MOTIVATION: It is expected that individual genes have intrinsically different variability in the global expressional trend among them. Thus, the consideration of gene-specific expressional properties will help us to distinguish target-selective gene expression over non-selective over-expression. RESULTS: The re-standardization and integration of heterogeneous microarray datasets, available from public databases, have enabled us to determine the global expression properties of individual genes across a wide variety of experimental conditions and samples. The global averages and SDs of expression for each gene in the integrated microarray datasets were found to be intrinsic properties, which were consistent among independent collections of datasets using different microarray platforms. Using the gene-specific intrinsic parameters to rescale the microarray data, we were able to distinguish novel selective gene expression [cartilage oligomeric matrix protein (COMP) and Collagen X] in breast cancer tissues from non-selective over-expression, a difference that has not been detectable by conventional methods. AVAILABILITY AND IMPLEMENTATION: The web-based tool for GS-LAGE is available at http://lage.sookmyung.ac.kr
Changsik Kim, Hyunjin Park, Yunsun Park, Jungsun Park, Taesung Park, Kwanghui Cho, Young Yang, Sukjoon Yoon
Bioinform.6
2009 Double error shrinkage method for identifying protein binding sites observed by tiling arrays with limited replication
abstract
MOTIVATION: ChIP-chip has been widely used for various genome-wide biological investigations. Given the small number of replicates (typically two to three) per biological sample, methods of analysis that control the variance are desirable but in short supply. We propose a double error shrinkage (DES) method by using moving average statistics based on local-pooled error estimates which effectively control both heterogeneous error variances and correlation structures of an extremely large number of individual probes on tiling arrays. RESULTS: Applying DES to ChIP-chip tiling array study for discovering genome-wide protein-binding sites, we identified 8400 target regions that include highly likely TFIID binding sites. About 33% of these were well matched with the known transcription starting sites on the DBTSS library, while many other newly identified sites have a high chance to be real binding sites based on a high positive predictive value of DES. We also showed the superior performance of DES compared with other commonly used methods for detecting actual protein binding sites.
Youngchul Kim, Stefan Bekiranov, Jae K. Lee, Taesung Park
Bioinform.4
2009 New evaluation measures for multifactor dimensionality reduction classifiers in gene-gene interaction analysis
abstract
MOTIVATION: Gene-gene interactions are important contributors to complex biological traits. Multifactor dimensionality reduction (MDR) is a method to analyze gene-gene interactions and has been applied to many genetics studies of complex diseases. In order to identify the best interaction model associated with disease susceptibility, MDR classifiers corresponding to interaction models has been constructed and evaluated as a predictor of disease status via a certain measure such as balanced accuracy (BA). It has been shown that the performance of MDR tends to depend on the choice of the evaluation measures. RESULTS: In this article, we introduce two types of new evaluation measures. First, we develop weighted BA (wBA) that utilizes the quantitative information on the effect size of each multi-locus genotype on a trait. Second, we employ ordinal association measures to assess the performance of MDR classifiers. Simulation studies were conducted to compare the proposed measures with BA, a current measure. Our results showed that the wBA and tau(b) improved the power of MDR in detecting gene-gene interactions. Noticeably, the power increment was higher when data contains the greater number of genetic markers. Finally, we applied the proposed evaluation measures to real data.
Junghyun Namkung, Kyunga Kim, Sung-Gon Yi, Wonil Chung, Min-Seok Kwon, Taesung Park
Bioinform.6
2009 Identification of differentially expressed subnetworks based on multivariate ANOVA
abstract
BACKGROUND: Since high-throughput protein-protein interaction (PPI) data has recently become available for humans, there has been a growing interest in combining PPI data with other genome-wide data. In particular, the identification of phenotype-related PPI subnetworks using gene expression data has been of great concern. Successful integration for the identification of significant subnetworks requires the use of a search algorithm with a proper scoring method. Here we propose a multivariate analysis of variance (MANOVA)-based scoring method with a greedy search for identifying differentially expressed PPI subnetworks. RESULTS: Given the MANOVA-based scoring method, we performed a greedy search to identify the subnetworks with the maximum scores in the PPI network. Our approach was successfully applied to human microarray datasets. Each identified subnetwork was annotated with the Gene Ontology (GO) term, resulting in the phenotype-related functional pathway or complex. We also compared these results with those of other scoring methods such as t statistic- and mutual information-based scoring methods. The MANOVA-based method produced subnetworks with a larger number of proteins than the other methods. Furthermore, the subnetworks identified by the MANOVA-based method tended to consist of highly correlated proteins. CONCLUSION: This article proposes a MANOVA-based scoring method to combine PPI data with expression data using a greedy search. This method is recommended for the highly sensitive detection of large subnetworks.
Taeyoung Hwang, Taesung Park
BMC Bioinform.2
2008 Evolutionary design principles of modules that control cellular differentiation: consequences for hysteresis and multistationarity
abstract
MOTIVATION: Gene regulatory networks (GRNs) govern cellular differentiation processes and enable construction of multicellular organisms from single cells. Although such networks are complex, there must be evolutionary design principles that shape the network to its present form, gaining complexity from simple modules. RESULTS: To isolate particular design principles, we have computationally evolved random regulatory networks with a preference to result either in hysteresis (switching threshold depending on current state), or in multistationarity (having multiple steady states), two commonly observed dynamical features of GRNs related to differentiation processes. We have analyzed the resulting evolved networks and compared their structures and characteristics with real GRNs reported from experiments. CONCLUSION: We found that the artificially evolved networks have particular topologies and it was notable that these topologies share important features and similarities with the real GRNs, particularly in contrasting properties of positive and negative feedback loops. We conclude that the structures of real GRNs are consistent with selection to favor one or other of the dynamical features of multistationarity or hysteresis. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Junil Kim, Tae-Geon Kim, Sung Hoon Jung, Jeong-Rae Kim, Taesung Park, Pat Heslop-Harrison, Kwang-Hyun Cho
Bioinform.5
2008 Response projected clustering for direct association with physiological and clinical response data
abstract
BACKGROUND: Microarray gene expression data are often analyzed together with corresponding physiological response and clinical metadata of biological subjects, e.g. patients' residual tumor sizes after chemotherapy or glucose levels at various stages of diabetic patients. Current clustering analysis cannot directly incorporate such quantitative metadata into the clustering heatmap of gene expression. It will be quite useful if these clinical response data can be effectively summarized in the high-dimensional clustering display so that important groups of genes can be intuitively discovered with different degrees of relevance to target disease phenotypes. RESULTS: We introduced a novel clustering analysis approach, response projected clustering (RPC), which uses a high-dimensional geometrical projection of response data to the gene expression space. The projected response vector, which becomes the origin in the projected space, is then clustered together with the projected gene vectors based on their different degrees of association with the response vector. A bootstrap-counting based RPC analysis is also performed to evaluate statistical tightness of identified gene clusters. Our RPC analysis was applied to the in vitro growth-inhibition and microarray profiling data on the NCI-60 cancer cell lines and the microarray gene expression study of macrophage differentiation in atherogenesis. These RPC applications enabled us to identify many known and novel gene factors and their potential pathway associations which are highly relevant to the drug's chemosensitivity activities and atherogenesis. CONCLUSION: We have shown that RPC can effectively discover gene networks with different degrees of association with clinical metadata. Performed on each gene's response projected vector based on its degree of association with the response data, RPC effectively summarizes individual genes' association with metadata as well as their own expression patterns. Thus, RPC greatly enhances the utility of clustering analysis on investigating high-dimensional microarray gene expression data with quantitative metadata.
Sung-Gon Yi, Taesung Park, Jae K. Lee
BMC Bioinform.2
2007 Odds ratio based multifactor-dimensionality reduction method for detecting gene-gene interactions
abstract
MOTIVATION: The identification and characterization of genes that increase the susceptibility to common complex multifactorial diseases is a challenging task in genetic association studies. The multifactor dimensionality reduction (MDR) method has been proposed and implemented by Ritchie et al. (2001) to identify the combinations of multilocus genotypes and discrete environmental factors that are associated with a particular disease. However, the original MDR method classifies the combination of multilocus genotypes into high-risk and low-risk groups in an ad hoc manner based on a simple comparison of the ratios of the number of cases and controls. Hence, the MDR approach is prone to false positive and negative errors when the ratio of the number of cases and controls in a combination of genotypes is similar to that in the entire data, or when both the number of cases and controls is small. Hence, we propose the odds ratio based multifactor dimensionality reduction (OR MDR) method that uses the odds ratio as a new quantitative measure of disease risk. RESULTS: While the original MDR method provides a simple binary measure of risk, the OR MDR method provides not only the odds ratio as a quantitative measure of risk but also the ordering of the multilocus combinations from the highest risk to lowest risk groups. Furthermore, the OR MDR method provides a confidence interval for the odds ratio for each multilocus combination, which is extremely informative in judging its importance as a risk factor. The proposed OR MDR method is illustrated using the dataset obtained from the CDC Chronic Fatigue Syndrome Research Group. AVAILABILITY: The program written in R is available.
Yujin Chung, Seung Yeoun Lee, Robert C. Elston, Taesung Park
Bioinform.4
2007 Log-linear model-based multifactor dimensionality reduction method to detect gene-gene interactions
abstract
MOTIVATION: The identification and characterization of susceptibility genes that influence the risk of common and complex diseases remains a statistical and computational challenge in genetic association studies. This is partly because the effect of any single genetic variant for a common and complex disease may be dependent on other genetic variants (gene-gene interaction) and environmental factors (gene-environment interaction). To address this problem, the multifactor dimensionality reduction (MDR) method has been proposed by Ritchie et al. to detect gene-gene interactions or gene-environment interactions. The MDR method identifies polymorphism combinations associated with the common and complex multifactorial diseases by collapsing high-dimensional genetic factors into a single dimension. That is, the MDR method classifies the combination of multilocus genotypes into high-risk and low-risk groups based on a comparison of the ratios of the numbers of cases and controls. When a high-order interaction model is considered with multi-dimensional factors, however, there may be many sparse or empty cells in the contingency tables. The MDR method cannot classify an empty cell as high risk or low risk and leaves it as undetermined. RESULTS: In this article, we propose the log-linear model-based multifactor dimensionality reduction (LM MDR) method to improve the MDR in classifying sparse or empty cells. The LM MDR method estimates frequencies for empty cells from a parsimonious log-linear model so that they can be assigned to high-and low-risk groups. In addition, LM MDR includes MDR as a special case when the saturated log-linear model is fitted. Simulation studies show that the LM MDR method has greater power and smaller error rates than the MDR method. The LM MDR method is also compared with the MDR method using as an example sporadic Alzheimer's disease.
Seung Yeoun Lee, Yujin Chung, Robert C. Elston, Youngchul Kim, Taesung Park
Bioinform.5
2007 Boolean networks using the chi-square test for inferring large-scale gene regulatory networks
abstract
BACKGROUND: Boolean network (BN) modeling is a commonly used method for constructing gene regulatory networks from time series microarray data. However, its major drawback is that its computation time is very high or often impractical to construct large-scale gene networks. We propose a variable selection method that are not only reduces BN computation times significantly but also obtains optimal network constructions by using chi-square statistics for testing the independence in contingency tables. RESULTS: Both the computation time and accuracy of the network structures estimated by the proposed method are compared with those of the original BN methods on simulated and real yeast cell cycle microarray gene expression data sets. Our results reveal that the proposed chi-square testing (CST)-based BN method significantly improves the computation time, while its ability to identify all the true network mechanisms was effectively the same as that of full-search BN methods. The proposed BN algorithm is approximately 70.8 and 7.6 times faster than the original BN algorithm when the error sizes of the Best-Fit Extension problem are 0 and 1, respectively. Further, the false positive error rate of the proposed CST-based BN algorithm tends to be less than that of the original BN. CONCLUSION: The CST-based BN method dramatically improves the computation time of the original BN algorithm. Therefore, it can efficiently infer large-scale gene regulatory network mechanisms.
Haseong Kim, Jae K. Lee, Taesung Park
BMC Bioinform.3
2007 Robust imputation method for missing values in microarray data
abstract
BACKGROUND: When analyzing microarray gene expression data, missing values are often encountered. Most multivariate statistical methods proposed for microarray data analysis cannot be applied when the data have missing values. Numerous imputation algorithms have been proposed to estimate the missing values. In this study, we develop a robust least squares estimation with principal components (RLSP) method by extending the local least square imputation (LLSimpute) method. The basic idea of our method is to employ quantile regression to estimate the missing values, using the estimated principal components of a selected set of similar genes. RESULTS: Using the normalized root mean squares error, the performance of the proposed method was evaluated and compared with other previously proposed imputation methods. The proposed RLSP method clearly outperformed the weighted k-nearest neighbors imputation (kNNimpute) method and LLSimpute method, and showed competitive results with Bayesian principal component analysis (BPCA) method. CONCLUSION: Adapting the principal components of the selected genes and employing the quantile regression model improved the robustness and accuracy of missing value imputation. Thus, the proposed RLSP method is, according to our empirical studies, more robust and accurate than the widely used kNNimpute and LLSimpute methods.
Dankyu Yoon, Taesung Park
BMC Bioinform.3
2006 arrayQCplot: software for checking the quality of microarray data
abstract
UNLABELLED: arrayQCplot is a software for the exploratory analysis of microarray data. This software focuses on quality control and generates newly developed plots for quality and reproducibility checks. It is developed using R and provides a user-friendly graphical interface for graphics and statistical analysis. Therefore, novice users will find arrayQCplot as an easy-to-use software for checking the quality of their data by a simple mouse click. AVAILABILITY: arrayQCplot software is available from Bioconductor at http://www.bioconductor.org. A more detailed manual is available at http://bibs.snu.ac.kr/software/arrayQCplot CONTACT: [email protected].
Sung-Gon Yi, Taesung Park
Bioinform.3
2006 Combining multiple microarrays in the presence of controlling variables
abstract
MOTIVATION: Microarray technology enables the monitoring of expression levels for thousands of genes simultaneously. When the magnitude of the experiment increases, it becomes common to use the same type of microarrays from different laboratories or hospitals. Thus, it is important to analyze microarray data together to derive a combined conclusion after accounting for the differences. One of the main objectives of the microarray experiment is to identify differentially expressed genes among the different experimental groups. The analysis of variance (ANOVA) model has been commonly used to detect differentially expressed genes after accounting for the sources of variation commonly observed in the microarray experiment. RESULTS: We extended the usual ANOVA model to account for an additional variability resulting from many confounding variables such as the effect of different hospitals. The proposed model is a two-stage ANOVA model. The first stage is the adjustment for the effects of no interests. The second stage is the detection of differentially expressed genes among the experimental groups using the residuals obtained from the first stage. Based on these residuals, we propose a permutation test to detect the differentially expressed genes. The proposed model is illustrated using the data from 133 microarrays collected at three different hospitals. The proposed approach is more flexible to use, and it is easier to accommodate the individual covariates in this model than using the meta-analysis approach. AVAILABILITY: A set of programs written in R will be electronically sent upon request.
Taesung Park, Sung-Gon Yi, Young Kee Shin, Seung Yeoun Lee
Bioinform.1
2004 Two-stage normalization using background intensities in cDNA microarray data
abstract
BACKGROUND: In the microarray experiment, many undesirable systematic variations are commonly observed. Normalization is the process of removing such variation that affects the measured gene expression levels. Normalization plays an important role in the earlier stage of microarray data analysis. The subsequent analysis results are highly dependent on normalization. One major source of variation is the background intensities. Recently, some methods have been employed for correcting the background intensities. However, all these methods focus on defining signal intensities appropriately from foreground and background intensities in the image analysis. Although a number of normalization methods have been proposed, no systematic methods have been proposed using the background intensities in the normalization process. RESULTS: In this paper, we propose a two-stage method adjusting for the effect of background intensities in the normalization process. The first stage fits a regression model to adjust for the effect of background intensities and the second stage applies the usual normalization method such as a nonlinear LOWESS method to the background-adjusted intensities. In order to carry out the two-stage normalization method, we consider nine different background measures and investigate their performances in normalization. The performance of two-stage normalization is compared to those of global median normalization as well as intensity dependent nonlinear LOWESS normalization. We use the variability among the replicated slides to compare performance of normalization methods. CONCLUSIONS: For the selected background measures, the proposed two-stage normalization method performs better than global or intensity dependent nonlinear LOWESS normalization method. Especially, when there is a strong relationship between the background intensity and the signal intensity, the proposed method performs much better. Regardless of background correction methods used in the image analysis, the proposed two-stage normalization method can be applicable as long as both signal intensity and background intensity are available.
Dankyu Yoon, Sung-Gon Yi, Ju-Han Kim, Taesung Park
BMC Bioinform.4
2003 Statistical tests for identifying differentially expressed genes in time-course microarray experiments
abstract
MOTIVATION: Microarray technology allows the monitoring of expression levels for thousands of genes simultaneously. In time-course experiments in which gene expression is monitored over time, we are interested in testing gene expression profiles for different experimental groups. However, no sophisticated analytic methods have yet been proposed to handle time-course experiment data. RESULTS: We propose a statistical test procedure based on the ANOVA model to identify genes that have different gene expression profiles among experimental groups in time-course experiments. Especially, we propose a permutation test which does not require the normality assumption. For this test, we use residuals from the ANOVA model only with time-effects. Using this test, we detect genes that have different gene expression profiles among experimental groups. The proposed model is illustrated using cDNA microarrays of 3840 genes obtained in an experiment to search for changes in gene expression profiles during neuronal differentiation of cortical stem cells.
Taesung Park, Sung-Gon Yi, Seungmook Lee, Seung Yeoun Lee, Dong-Hyun Yoo, Jun-Ik Ahn, Yong-Sung Lee
Bioinform.1
2003 Evaluation of normalization methods for microarray data
abstract
BACKGROUND: Microarray technology allows the monitoring of expression levels for thousands of genes simultaneously. This novel technique helps us to understand gene regulation as well as gene by gene interactions more systematically. In the microarray experiment, however, many undesirable systematic variations are observed. Even in replicated experiment, some variations are commonly observed. Normalization is the process of removing some sources of variation which affect the measured gene expression levels. Although a number of normalization methods have been proposed, it has been difficult to decide which methods perform best. Normalization plays an important role in the earlier stage of microarray data analysis. The subsequent analysis results are highly dependent on normalization. RESULTS: In this paper, we use the variability among the replicated slides to compare performance of normalization methods. We also compare normalization methods with regard to bias and mean square error using simulated data. CONCLUSIONS: Our results show that intensity-dependent normalization often performs better than global normalization methods, and that linear and nonlinear normalization methods perform similarly. These conclusions are based on analysis of 36 cDNA microarrays of 3,840 genes obtained in an experiment to search for changes in gene expression profiles during neuronal differentiation of cortical stem cells. Simulation studies confirm our findings.
Taesung Park, Sung-Gon Yi, Sung-Hyun Kang, Seung Yeoun Lee, Yong-Sung Lee, Richard Simon
BMC Bioinform.1