Jiangwen Sun

dblp:64/3134 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0000-8905-7553ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Task-Aware Benchmarking of Batch Correction Methods for scRNA Triple-Negative Breast Cancer Atlas Construction
Peter Scheible, Jing He 0002, Amy H. Tang, Jiangwen Sun
ISBRA (2)4
2025 Optimizing Deep Learning Models for DNA Methylation Prediction
abstract
Deep learning has become an essential technique for deciphering genomic sequences and predicting regulatory activities in cellular biology. However, training a deep neural network to achieve a model with optimal performance for prediction tasks in functional genomics remains a challenge due to the high complexity in the regulatory biology of DNA sequence. In this study, we explore the impact of multiple key network training parameters, including the optimization algorithm, training objective, input context length, and network architecture on model performance for DNA methylation prediction from DNA sequence. Our results show that deeper architectures improve predictive accuracy but require careful selection of optimization algorithms to prevent overfitting. We also find that increasing input context length enhances model performance up to a certain threshold, beyond which diminishing returns occur. Additionally, loss function choice significantly influences model stability and generalization, with certain formulations better capturing sequence-level dependencies. By systematically evaluating these factors across multiple model architectures, we identify approaches that enhance predictive accuracy and consistency. Our findings provide insights into the trade-offs between model depth, optimization techniques, loss function selection, and sequence length offering a framework for improving deep learning applications in genomics. This work contributes to the development of more effective computational tools for analyzing regulatory sequences and understanding their role in gene regulation and disease.
Bashar A. Fakhreddin, Sanjeeva Dodlapati, Jiangwen Sun
BIBM3
2025 Efficient Longitudinal Feature Selection Via Binarized Transformation: Theory and Case Studies
Jason Orender, Jiangwen Sun, Mohammed Zubair
IEEE Big Data2
2025 LTBoost: Boosting Recall Uniformity in Long-Tailed Learning
Morteza Mohammady Gharasuie, Fengjio Wang, Ravi Mukkamala, Jiangwen Sun
CAIP (1)4
2022 The Combined Focal Cross Entropy and Dice Loss Function for Segmentation of Protein Secondary Structures from Cryo-EM 3D Density maps
abstract
Although cryo-electron microscopy (cryo-EM) has been successfully used to derive atomic structures for many proteins, it is still challenging to derive atomic structure when the resolution of cryo-EM density maps is in the medium resolution range such as 5-10 A. Although multiple neural networks have been proposed for the problem of secondary structure detection from cryo-EM 3D images, loss functions used in the existing networks are primarily based on cross entropy loss (CE). In order to study the behavior of various loss functions in the secondary structure detection problem, we investigated five loss functions and compared their performances. Using a U-net architecture in DeepSSETracer and a test set of 65 protein chains of atomic structures and their corresponding cryo-EM density component maps, we found that the combined function with focal cross entropy loss (FCE) and Dice loss (DL) provides the best overall detection of secondary structures. In particular, the combined loss function has a significant enhancement of an overall F1score of 6.7% when compared to CE in detection of $\beta-$sheet voxels that are generally much harder to be detected accurately than for helix voxels. Our work shows the potential of designing effective loss functions to enhance the detection of hard cases in the segmentation of secondary structure problem.
Yongcheng Mu, Jiangwen Sun, Jing He 0002
BIBM2
2022 Highly Scalable Task Grouping for Deep Multi-Task Learning in Prediction of Epigenetic Events
abstract
DNNs trained for predicting cellular events from DNA sequence have become emerging tools to help elucidate biological mechanisms underlying associations identified in genome-wide association studies. To enhance the training, multi-task learning (MTL) has been commonly exploited in previous works where trained networks were needed for multiple profiles differing in either event modality or cell type. All existing works adopted a simple MTL framework where all tasks share a single feature extraction network. Such a strategy even though effective to a certain extent leads to substantial negative transfer, meaning the existence of a large portion of tasks for which models obtained through MTL perform worse than those by single-task learning. There have been methods developed to address such negative transfer in other domains, such as computer vision. However, these methods are generally with limited scalability. In this paper, we propose a highly scalable task grouping framework to address negative transfer by only jointly training tasks that are potentially beneficial to each other. The proposed method exploits the network weights associated with task-specific classification heads that can be cheaply obtained by one-time joint training of all tasks. Our results using a dataset consisting of 367 epigenetic profiles demonstrate the effectiveness of the proposed approach and its superiority over baseline methods.
Mohammad Shiri, Jiangwen Sun
BIBM2
2022 LASSO Logic Engine: harnessing the logic parsing capabilities of the LASSO algorithm for longitudinal feature learning
abstract
Longitudinal data, which is widely used in many disciplines to study cause and effect, poses significant computational challenges to both modeling and analysis. Longitudinal data is composed of readings on the same variable collected over time and is often high-dimensional with correlated features. The combinatorial search approach for identifying the optimal features is unrealistic for most applications. The alternative approaches, such as heuristics, greedy searches, and regularization techniques, including LASSO, can result in models that suffer from both low accuracy and unclear feature attribution. In this paper, we propose a binary transformation on the data before applying LASSO for feature learning. As demonstrated in the paper, the binary transformation enhances signal in the data, resulting in highly accurate feature attribution, including associated time lags. It avoids the typical shortcomings of the LASSO algorithm, including saturation of the feature space and arbitrary or inconsistent sparse feature selection. Both synthetic data and real-world data sets were used to demonstrate the value of the proposed transformation and in all cases substantial improvements in feature learning were seen. In addition, the scalable parallelism of the solution is superior to that of the standard LASSO since transformation itself occurs in linear time and computing the LASSO solution using the transformed data results in a speedup of almost double.
Jason Orender, Mohammad Zubair, Jiangwen Sun
IEEE Big Data3
2021 Multi-view spectral graph convolution with consistent edge attention for molecular modeling
Qinqing Liu, Jiangwen Sun, Minghu Song, Jinbo Bi
Neurocomputing4
2019 Multi-view cluster analysis with incomplete data to understand treatment effects
Guoqing Chao, Jiangwen Sun, Jin Lu 0001, An-Li Wang, Daniel D. Langleben, Chiang-shan Ray Li, Jinbo Bi
Inf. Sci.2
2017 Collaborative phenotype inference from comorbid substance use disorders and genotypes
abstract
Data in large-scale genetic studies of complex human diseases, such as substance use disorders, are often incomplete. Despite great progress in genotype imputation, e.g., the IMPUTE2 method, considerably less progress has been made in inferring phenotypes. We designed a novel approach to integrate individuals' comorbid conditions with their genotype data to infer missing (unreported) diagnostic criteria of a disorder. The premise of our approach derives from correlations among symptoms and the shared biological bases of concurrent disorders such as co-dependence on cocaine and opioids. We describe a matrix completion method to construct a bi-linear model based on the interactions of genotypes and known symptoms of related disorders to infer unknown values of another set of symptoms or phenotypes. An efficient stochastic and parallel algorithm based on the linearized alternating direction method of multipliers was developed to solve the proposed optimization problem. Empirical evaluation of the approach in comparison with other advanced data matrix completion methods via a case study shows that it both significantly improves imputation accuracy and provides greater computational efficiency.
Jin Lu 0001, Jiangwen Sun, Xinyu Wang 0055, Henry R. Kranzler, Joel Gelernter, Jinbo Bi
BIBM2
2017 VIGAN: Missing view imputation with generative adversarial networks
abstract
In an era when big data are becoming the norm, there is less concern with the quantity but more with the quality and completeness of the data. In many disciplines, data are collected from heterogeneous sources, resulting in multi-view or multi-modal datasets. The missing data problem has been challenging to address in multi-view data analysis. Especially, when certain samples miss an entire view of data, it creates the missing view problem. Classic multiple imputations or matrix completion methods are hardly effective here when no information can be based on in the specific view to impute data for such samples. The commonly-used simple method of removing samples with a missing view can dramatically reduce sample size, thus diminishing the statistical power of a subsequent analysis. In this paper, we propose a novel approach for view imputation via generative adversarial networks (GANs), which we name by VIGAN. This approach first treats each view as a separate domain and identifies domain-to-domain mappings via a GAN using randomly-sampled data from each view, and then employs a multi-modal denoising autoencoder (DAE) to reconstruct the missing view from the GAN outputs based on paired data across the views. Then, by optimizing the GAN and DAE jointly, our model enables the knowledge integration for domain mappings and view correspondences to effectively recover the missing view. Empirical results on benchmark datasets validate the VIGAN approach by comparing against the state of the art. The evaluation of VIGAN in a genetic study of substance use disorders further proves the effectiveness and usability of this approach in life science.
Aaron Palmer, Jiangwen Sun, Ko-Shin Chen, Jin Lu 0001, Jinbo Bi
IEEE BigData3
2016 A Sparse Interactive Model for Matrix Completion with Side Information
abstract
Matrix completion methods can benefit from side information besides the partially observed matrix. The use of side features describing the row and column entities of a matrix has been shown to reduce the sample complexity for completing the matrix. We propose a novel sparse formulation that explicitly models the interaction between the row and column side features to approximate the matrix entries. Unlike early methods, this model does not require the low-rank condition on the model parameter matrix. We prove that when the side features can span the latent feature space of the matrix to be recovered, the number of observed entries needed for an exact recovery is $O(\log N)$ where $N$ is the size of the matrix. When the side features are corrupted latent features of the matrix with a small perturbation, our method can achieve an $\epsilon$-recovery with $O(\log N)$ sample complexity, and maintains a $\O(N^{3/2})$ rate similar to classfic methods with no side information. An efficient linearized Lagrangian algorithm is developed with a strong guarantee of convergence. Empirical results show that our approach outperforms three state-of-the-art methods both in simulations and on real world datasets.
Jin Lu 0001, Guannan Liang, Jiangwen Sun, Jinbo Bi
NIPS3
2016 A cross-species bi-clustering approach to identifying conserved co-regulated genes
abstract
MOTIVATION: A growing number of studies have explored the process of pre-implantation embryonic development of multiple mammalian species. However, the conservation and variation among different species in their developmental programming are poorly defined due to the lack of effective computational methods for detecting co-regularized genes that are conserved across species. The most sophisticated method to date for identifying conserved co-regulated genes is a two-step approach. This approach first identifies gene clusters for each species by a cluster analysis of gene expression data, and subsequently computes the overlaps of clusters identified from different species to reveal common subgroups. This approach is ineffective to deal with the noise in the expression data introduced by the complicated procedures in quantifying gene expression. Furthermore, due to the sequential nature of the approach, the gene clusters identified in the first step may have little overlap among different species in the second step, thus difficult to detect conserved co-regulated genes. RESULTS: We propose a cross-species bi-clustering approach which first denoises the gene expression data of each species into a data matrix. The rows of the data matrices of different species represent the same set of genes that are characterized by their expression patterns over the developmental stages of each species as columns. A novel bi-clustering method is then developed to cluster genes into subgroups by a joint sparse rank-one factorization of all the data matrices. This method decomposes a data matrix into a product of a column vector and a row vector where the column vector is a consistent indicator across the matrices (species) to identify the same gene cluster and the row vector specifies for each species the developmental stages that the clustered genes co-regulate. Efficient optimization algorithm has been developed with convergence analysis. This approach was first validated on synthetic data and compared to the two-step method and several recent joint clustering methods. We then applied this approach to two real world datasets of gene expression during the pre-implantation embryonic development of the human and mouse. Co-regulated genes consistent between the human and mouse were identified, offering insights into conserved functions, as well as similarities and differences in genome activation timing between the human and mouse embryos. AVAILABILITY AND IMPLEMENTATION: The R package containing the implementation of the proposed method in C ++ is available at: https://github.com/JavonSun/mvbc.git and also at the R platform https://www.r-project.org/ CONTACT: [email protected].
Jiangwen Sun, Zongliang Jiang, Xiuchun Tian, Jinbo Bi
Bioinform.1
2016 Multiplicative Multitask Feature Learning
abstract
We investigate a general framework of multiplicative multitask feature learning which decomposes individual task's model parameters into a multiplication of two components. One of the components is used across all tasks and the other component is task-specific. Several previous methods can be proved to be special cases of our framework. We study the theoretical properties of this framework when different regularization conditions are applied to the two decomposed components. We prove that this framework is mathematically equivalent to the widely used multitask feature learning methods that are based on a joint regularization of all model parameters, but with a more general form of regularizers. Further, an analytical formula is derived for the across-task component as related to the task- specific component for all these regularizers, leading to a better understanding of the shrinkage effects of different regularizers. Study of this framework motivates new multitask learning algorithms. We propose two new learning formulations by varying the parameters in the proposed framework. An efficient blockwise coordinate descent algorithm is developed suitable for solving the entire family of formulations with rigorous convergence analysis. Simulation studies have identified the statistical properties of data that would be in favor of the new formulations. Extensive empirical studies on various classification and regression benchmark data sets have revealed the relative advantages of the two new formulations by comparing with the state of the art, which provides instructive insights into the feature learning problem with multiple tasks.
Xin Wang 0023, Jinbo Bi, Shipeng Yu, Jiangwen Sun, Minghu Song
J. Mach. Learn. Res.4
2015 Quantifying feed efficiency of dairy cattle for genome-wide association analysis
abstract
Improving feed efficiency in dairy production is an important endeavor, as it can reduce feed costs and negative impacts of production on the environment. Feed efficiency is a multivariate phenotype that is characterized by a variety of phenotypic variables, such as dry matter intake, body weight gain, and milk yield. Currently, there is no consensus method for quantifying the feed efficiency of lactating dairy cattle for the purpose of breeding selection. Residual feed intake, which is the difference between actual feed intake and predicted intake, has been one of the commonly used measures for feed efficiency. However, such a measure is heterogeneous showing substantial variation in the cow population and has relatively low heritability (0.01~0.38). Hence, its utility in breeding selection is limited. In particular, no prior study has utilized genetic data directly in the development of feed efficiency measures. In this paper, we aim to identify cattle clusters with homogeneous feed efficiency features that are ready to link to genetic variants, and thus can have greater utility in breading selection. In order to achieve this goal, we explore a new multi-view clustering method that jointly analyzes two views of data: phenotypic measures and genotypes, and identifies cattle clusters that are characterized by specific phenotypic features and also associated with genetic markers. Using a set of feed efficiency data collected by USDA, three cattle subgroups have been identified by our analysis, and they offer instructive insights into future feed efficiency studies.
Tingyang Xu, Jiangwen Sun, Erin E. Connor, Jinbo Bi
BIBM2
2015 Multi-view Sparse Co-clustering via Proximal Alternating Linearized Minimization
abstract
When multiple views of data are available for a set of subjects, co-clustering aims to identify subject clusters that agree across the different views. We explore the problem of co-clustering when the underlying clusters exist in different subspaces of each view. We propose a proximal alternating linearized minimization algorithm that simultaneously decomposes multiple data matrices into sparse row and columns vectors. This approach is able to group subjects consistently across the views and simultaneously identify the subset of features in each view that are associated with the clusters. The proposed algorithm can globally converge to a critical point of the problem. A simulation study validates that the proposed algorithm can identify the hypothesized clusters and their associated features. Comparison with several latest multi-view co-clustering methods on benchmark datasets demonstrates the superior performance of the proposed approach.
Jiangwen Sun, Jin Lu 0001, Tingyang Xu, Jinbo Bi
ICML1
2015 Longitudinal LASSO: Jointly Learning Features and Temporal Contingency for Outcome Prediction
abstract
Longitudinal analysis is important in many disciplines, such as the study of behavioral transitions in social science. Only very recently, feature selection has drawn adequate attention in the context of longitudinal modeling. Standard techniques, such as generalized estimating equations, have been modified to select features by imposing sparsity-inducing regularizers. However, they do not explicitly model how a dependent variable relies on features measured at proximal time points. Recent graphical Granger modeling can select features in lagged time points but ignores the temporal correlations within an individual's repeated measurements. We propose an approach to automatically and simultaneously determine both the relevant features and the relevant temporal points that impact the current outcome of the dependent variable. Meanwhile, the proposed model takes into account the non-i.i.d nature of the data by estimating the within-individual correlations. This approach decomposes model parameters into a summation of two components and imposes separate block-wise LASSO penalties to each component when building a linear model in terms of the past τ measurements of features. One component is used to select features whereas the other is used to select temporal contingent points. An accelerated gradient descent algorithm is developed to efficiently solve the related optimization problem with detailed convergence analysis and asymptotic analysis. Computational results on both synthetic and real world problems demonstrate the superior performance of the proposed approach over existing techniques.
Tingyang Xu, Jiangwen Sun, Jinbo Bi
KDD2
2014 A sparse integrative cluster analysis for understanding soybean phenotypes
abstract
Soybean is one of the most important crops for food, feed and bio-energy world-wide. The study of soybean phenotypic variation at different geographical locations can help the understanding of soybean domestication, population structure of soybean, and the conservation of soybean biodiversity. We investigate if soybean varieties can be identified that they differ from other varieties on multiple traits even when growing at different geographical locations. When a collection of traits are observed for the same soybean type at different locations (different views), joint analysis of the multiple-view data is required in order to identify the same soybean clusters based on data from different locations. We employ a new multi-view singular value decomposition approach that simultaneously decomposes the data matrix gathered at each location into sparse singular vectors. This approach is able to group soybean samples consistently across the different locations and simultaneously identify the phenotypes at each location on which the soybean samples within a cluster are the most similar. Comparison with several latest multi-view co-clustering methods demonstrates the superior performance of the proposed approach.
Jinbo Bi, Jiangwen Sun, Tingyang Xu, Jin Lu 0001, Yansong Ma, Lijuan Qiu
BIBM2
2014 Identifying heritable composite traits from multivariate phenotypes and genome-wide SNPs
abstract
An important approach to reducing missing heritability and enhancing success of genome-wide association studies (GWAS) for complex diseases is the identification of traits that are highly heritable and homogeneous in their etiology. Many approaches have been proposed to define such traits based on either cluster analysis or pedigree-based heritable component analysis. None of the existing methods, however, exploit the dense genome-wide genotypic data that are now readily available from GWAS, and with exome and whole genome sequencing more data will be available in the future. Moreover, because a phenotype can vary with respect to a covariate, such as age or race. The fixed effect due to the covariates may lead to a spuriously elevated estimate of heritability. Existing heritable component analysis methods have not considered covariate effects. We propose an optimization approach to identify composite traits with high heritability as a function of multiple phenotypic variables where heritability is estimated from genome-wide single neucleotide polymorphisms (SNPs). Our approach can model the covariate effects within heritability analysis. The proposed optimization problem can be efficiently solved by a sequential quadratic programming algorithm. A case study demonstrates the effectiveness of the proposed approach for finding composite traits with high SNP-based heritability.
Jiangwen Sun, Jinbo Bi, Henry R. Kranzler
BIBM1
2014 On Multiplicative Multitask Feature Learning
Xin Wang 0023, Jinbo Bi, Shipeng Yu, Jiangwen Sun
NIPS4
2014 Multiview Comodeling to Improve Subtyping and Genetic Association of Complex Diseases
abstract
Genetic association analysis of complex diseases has been limited by heterogeneity in their clinical manifestations and genetic etiology. Research has made it possible to differentiate homogeneous subtypes of the disease phenotype. Currently, the most sophisticated subtyping methods perform unsupervised cluster analysis using only clinical features of a disorder, resulting in subtypes for which genetic association may be limited. In this study, we seek to derive a novel multiview data analytic method that integrates two views of the data: the clinical features and the genetic markers of the same set of patients. Our method is based on multiobjective programming that is capable of clinically categorizing a disease phenotype so as to discover genetically different subtypes.We optimize two objectives jointly: 1) in cluster analysis, the derived clusters should differ significantly in clinical features; 2) these clusters can be well separated using genetic markers by constructed classifiers. Extensive computational experiments with two substance-use disorders using two populations show that the proposed algorithm is superior to existing subtyping methods.
Jiangwen Sun, Jinbo Bi, Henry R. Kranzler
IEEE J. Biomed. Health Informatics1
2013 Multi-view biclustering for genotype-phenotype association studies of complex diseases
abstract
Complex disorders exhibit great heterogeneity in both clinical manifestation and genetic etiology. This heterogeneity substantially limits the identification of geneotype-phenotype associations. Differentiating homogeneous subtypes of a complex phenotype will enable the detection of genetic variants contributing to the effect of subtypes that cannot be detected by the non-differentiated phenotype. However, the most sophisticated subtyping methods available so far perform unsupervised cluster analysis or latent class analysis on only phenotypic features. Without guidance from the genetic dimension, the resultant subtypes can be suboptimal and genetic associations may fail. We propose a multi-view biclustering approach that integrates phenotypic features and genetic markers to detect confirming evidence in the two views for a disease subtype. This approach groups subjects in clusters that are consistent between the phenotypic and genetic views, and simultaneously identifies the phenotypic features that are used to define a subtype and the genotypes that are associated with the subtype. Our simulation study validates this approach, and our extensive comparison with several biclustering and multi-view data analytics on real-life disease data demonstrates the superior performance of the proposed approach.
Jiangwen Sun, Jinbo Bi, Henry R. Kranzler
BIBM1
2013 Quadratic optimization to identify highly heritable quantitative traits from complex phenotypic features
abstract
Identifying genetic variation underlying a complex disease is important. Many complex diseases have heterogeneous phenotypes and are products of a variety of genetic and environmental factors acting in concert. Deriving highly heritable quantitative traits of a complex disease can improve the identification of genetic risk of the disease. The most sophisticated methods so far perform unsupervised cluster analysis on phenotypic features; and then a quantitative trait is derived based on each resultant cluster. Heritability is estimated to assess the validity of the derived quantitative traits. However, none of these methods explicitly maximize the heritability of the derived traits. We propose a quadratic optimization approach that directly utilizes heritability as an objective during the derivation of quantitative traits of a disease. This method maximizes an objective function that is formulated by decomposing the traditional maximum likelihood method for estimating heritability of a quantitative trait. We demonstrate the effectiveness of the proposed method on both synthetic data and real-world problems. We apply our algorithm to identify highly heritable traits of complex human-behavior disorders including opioid and cocaine use disorders, and highly heritable traits of dairy cattle that are economically important. Our approach outperforms standard cluster analysis and several previous methods.
Jiangwen Sun, Jinbo Bi, Henry R. Kranzler
KDD1
2013 A machine learning approach to college drinking prediction and risk factor identification
abstract
Alcohol misuse is one of the most serious public health problems facing adolescents and young adults in the United States. National statistics shows that nearly 90% of alcohol consumed by youth under 21 years of age involves binge drinking and 44% of college students engage in high-risk drinking activities. Conventional alcohol intervention programs, which aim at installing either an alcohol reduction norm or prohibition against underage drinking, have yielded little progress in controlling college binge drinking over the years. Existing alcohol studies are deductive where data are collected to investigate a psychological/behavioral hypothesis, and statistical analysis is applied to the data to confirm the hypothesis. Due to this confirmatory manner of analysis, the resulting statistical models are cohort-specific and typically fail to replicate on a different sample. This article presents two machine learning approaches for a secondary analysis of longitudinal data collected in college alcohol studies sponsored by the National Institute on Alcohol Abuse and Alcoholism. Our approach aims to discover knowledge, from multiwave cohort-sequential daily data, which may or may not align with the original hypothesis but quantifies predictive models with higher likelihood to generalize to new samples. We first propose a so-called temporally-correlated support vector machine to construct a classifier as a function of daily moods, stress, and drinking expectancies to distinguish days with nighttime binge drinking from days without for individual students. We then propose a combination of cluster analysis and feature selection, where cluster analysis is used to identify drinking patterns based on averaged daily drinking behavior and feature selection is used to identify risk factors associated with each pattern. We evaluate our methods on two cohorts of 530 total college students recruited during the Spring and Fall semesters, respectively. Cross validation on these two cohorts and further on 100 random partitions of the total students demonstrate that our methods improve the model generalizability in comparison with traditional multilevel logistic regression. The discovered risk factors and the interaction of these factors delineated in our models can set a potential basis and offer insights to a new design of more effective college alcohol interventions.
Jinbo Bi, Jiangwen Sun, Howard Tennen, Stephen Armeli
ACM Trans. Intell. Syst. Technol.2
2012 A multi-objective program for quantitative subtyping of clinically relevant phenotypes
abstract
Identifying genetic variations that underlie human disease is very important to advance our understanding of the disease's pathophysiology and promote its personalized treatment. However, many disease phenotypes have complex clinical manifestations and a complicated etiology. Gene finding efforts for complex diseases have had limited success to date. Research results suggest that one way to enhance these efforts is to differentiate subtypes of a complex multifactorial disease phenotype. Existing subtyping methods rely on cluster analysis using only clinical features of a disorder without guidance from genetic data, resulting in subtypes for which genotype association may be limited. In this work, we seek to derive a novel computational method based on multi-objective programming that is capable of clinically categorizing a disease phenotype so as to discover genetically different subtypes. Our approach optimizes two objectives: (1) the cluster-derived subtypes should differ significantly on clinical features; (2) these subtypes can be well separated using candidate genes. This work has been motivated by clinical studies of opioid dependence, a serious, prevalent disorder that is heterogeneous phenotypically. Analyses on a sample of 1,470 European American subjects aggregated from multiple genetic studies of opioid dependence show that the proposed algorithm is superior to existing subtyping methods.
Jiangwen Sun, Jinbo Bi, Henry R. Kranzler
BIBM1