VLDB 2026 Research / reviewers in the wild / expert
Philip M. Kim
dblp:77/177
· DBLP profile ↗
17ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-3683-152XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlowPacker: protein side-chain packing with torsional flow matchingabstractMOTIVATION: Accurate prediction of protein side-chain conformations is necessary to understand protein folding, protein-protein interactions and facilitate de novo protein design. RESULTS: Here, we apply torsional flow matching and equivariant graph attention to develop FlowPacker, a fast and performant model to predict protein side-chain conformations conditioned on the protein sequence and backbone. We show that FlowPacker outperforms previous state-of-the-art baselines across most metrics with improved runtime. We further show that FlowPacker can be used to inpaint missing side-chain coordinates and also for multimeric targets, and exhibits strong performance on a test set of antibody-antigen complexes. AVAILABILITY AND IMPLEMENTATION: Code is available at https://gitlab.com/mjslee0921/flowpacker. Jin Sub Lee, Philip M. Kim |
Bioinform. | 2 |
| 2024 | EuDockScore: Euclidean graph neural networks for scoring protein-protein interfacesabstractMOTIVATION: Protein-protein interactions are essential for a variety of biological phenomena including mediating biochemical reactions, cell signaling, and the immune response. Proteins seek to form interfaces which reduce overall system energy. Although determination of single polypeptide chain protein structures has been revolutionized by deep learning techniques, complex prediction has still not been perfected. Additionally, experimentally determining structures is incredibly resource and time expensive. An alternative is the technique of computational docking, which takes the solved individual structures of proteins to produce candidate interfaces (decoys). Decoys are then scored using a mathematical function that assess the quality of the system, known as scoring functions. Beyond docking, scoring functions are a critical component of assessing structures produced by many protein generative models. Scoring models are also used as a final filtering in many generative deep learning models including those that generate antibody binders, and those which perform docking. RESULTS: In this work, we present improved scoring functions for protein-protein interactions which utilizes cutting-edge Euclidean graph neural network architectures, to assess protein-protein interfaces. These Euclidean docking score models are known as EuDockScore, and EuDockScore-Ab with the latter being antibody-antigen dock specific. Finally, we provided EuDockScore-AFM a model trained on antibody-antigen outputs from AlphaFold-Multimer (AFM) which proves useful in reranking large numbers of AFM outputs. AVAILABILITY AND IMPLEMENTATION: The code for these models is available at https://gitlab.com/mcfeemat/eudockscore. Matthew McFee, Philip M. Kim |
Bioinform. | 3 |
| 2023 | HelixGAN a deep-learning methodology for conditional de novo design of α-helix structuresabstractMOTIVATION: Protein and peptide engineering has become an essential field in biomedicine with therapeutics, diagnostics and synthetic biology applications. Helices are both abundant structural feature in proteins and comprise a major portion of bioactive peptides. Precise design of helices for binding or biological activity is still a challenging problem. RESULTS: Here, we present HelixGAN, the first generative adversarial network method to generate de novo left-handed and right-handed alpha-helix structures from scratch at an atomic level. We developed a gradient-based search approach in latent space to optimize the generation of novel α-helical structures by matching the exact conformations of selected hotspot residues. The designed α-helical structures can bind specific targets or activate cellular receptors. There is a significant agreement between the helix structures generated with HelixGAN and PEP-FOLD, a well-known de novo approach for predicting peptide structures from amino acid sequences. HelixGAN outperformed RosettaDesign, and our previously developed structural similarity method to generate D-peptides matching a set of given hotspots in a known L-peptide. As proof of concept, we designed a novel D-GLP1_1 analog that matches the conformations of critical hotspots for the GLP1 function. MD simulations revealed a stable binding mode of the D-GLP1_1 analog coupled to the GLP1 receptor. This novel D-peptide analog is more stable than our previous D-GLP1 design along the MD simulations. We envision HelixGAN as a critical tool for designing novel bioactive peptides with specific properties in the early stages of drug discovery. AVAILABILITY AND IMPLEMENTATION: https://github.com/xxiexuezhi/helix_gan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xuezhi Xie, Pedro A. Valiente, Philip M. Kim |
Bioinform. | 3 |
| 2023 | Gate-based quantum computing for protein designabstractProtein design is a technique to engineer proteins by permuting amino acids in the sequence to obtain novel functionalities. However, exploring all possible combinations of amino acids is generally impossible due to the exponential growth of possibilities with the number of designable sites. The present work introduces circuits implementing a pure quantum approach, Grover's algorithm, to solve protein design problems. Our algorithms can adjust to implement any custom pair-wise energy tables and protein structure models. Moreover, the algorithm's oracle is designed to consist of only adder functions. Quantum computer simulators validate the practicality of our circuits, containing up to 234 qubits. However, a smaller circuit is implemented on real quantum devices. Our results show that using [Formula: see text] iterations, the circuits find the correct results among all N possibilities, providing the expected quadratic speed up of Grover's algorithm over classical methods (i.e., [Formula: see text]). Mohammad Hassan Khatami, Udson C. Mendes, Nathan Wiebe, Philip M. Kim |
PLoS Comput. Biol. | 4 |
| 2017 | Data driven flexible backbone protein designabstractProtein design remains an important problem in computational structural biology. Current computational protein design methods largely use physics-based methods, which make use of information from a single protein structure. This is despite the fact that multiple structures of many protein folds are now readily available in the PDB. While ensemble protein design methods can use multiple protein structures, they treat each structure independently. Here, we introduce a flexible backbone strategy, FlexiBaL-GP, which learns global protein backbone movements directly from multiple protein structures. FlexiBaL-GP uses the machine learning method of Gaussian Process Latent Variable Models to learn a lower dimensional representation of the protein coordinates that best represent backbone movements. These learned backbone movements are used to explore alternative protein backbones, while engineering a protein within a parallel tempered MCMC framework. Using the human ubiquitin-USP21 complex as a model we demonstrate that our design strategy outperforms current strategies for the interface design task of identifying tight binding ubiquitin variants for USP21. Mark G. F. Sun, Philip M. Kim |
PLoS Comput. Biol. | 2 |
| 2016 | JBASE: Joint Bayesian Analysis of Subphenotypes and EpistasisabstractMOTIVATION: Rapid advances in genotyping and genome-wide association studies have enabled the discovery of many new genotype-phenotype associations at the resolution of individual markers. However, these associations explain only a small proportion of theoretically estimated heritability of most diseases. In this work, we propose an integrative mixture model called JBASE: joint Bayesian analysis of subphenotypes and epistasis. JBASE explores two major reasons of missing heritability: interactions between genetic variants, a phenomenon known as epistasis and phenotypic heterogeneity, addressed via subphenotyping. RESULTS: Our extensive simulations in a wide range of scenarios repeatedly demonstrate that JBASE can identify true underlying subphenotypes, including their associated variants and their interactions, with high precision. In the presence of phenotypic heterogeneity, JBASE has higher Power and lower Type 1 Error than five state-of-the-art approaches. We applied our method to a sample of individuals from Mexico with Type 2 diabetes and discovered two novel epistatic modules, including two loci each, that define two subphenotypes characterized by differences in body mass index and waist-to-hip ratio. We successfully replicated these subphenotypes and epistatic modules in an independent dataset from Mexico genotyped with a different platform. AVAILABILITY AND IMPLEMENTATION: JBASE is implemented in C++, supported on Linux and is available at http://www.cs.toronto.edu/∼goldenberg/JBASE/jbase.tar.gz. The genotype data underlying this study are available upon approval by the ethics review board of the Medical Centre Siglo XXI. Please contact Dr Miguel Cruz at [email protected] for assistance with the application. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Recep Colak, TaeHyung Kim, Hilal Kazan, Yoomi Oh, Adan Valladares-Salgado, Jesus Peralta, Jorge Escobedo, Esteban J. Parra, Philip M. Kim, Anna Goldenberg |
Bioinform. | 10 |
| 2016 | ELASPIC web-server: proteome-wide structure-based prediction of mutation effects on protein stability and binding affinityabstractUNLABELLED: ELASPIC is a novel ensemble machine-learning approach that predicts the effects of mutations on protein folding and protein-protein interactions. Here, we present the ELASPIC webserver, which makes the ELASPIC pipeline available through a fast and intuitive interface. The webserver can be used to evaluate the effect of mutations on any protein in the Uniprot database, and allows all predicted results, including modeled wild-type and mutated structures, to be managed and viewed online and downloaded if needed. It is backed by a database which contains improved structural domain definitions, and a list of curated domain-domain interactions for all known proteins, as well as homology models of domains and domain-domain interactions for the human proteome. Homology models for proteins of other organisms are calculated on the fly, and mutations are evaluated within minutes once the homology model is available. AVAILABILITY AND IMPLEMENTATION: The ELASPIC webserver is available online at http://elaspic.kimlab.org CONTACT: [email protected] or [email protected] data: Supplementary data are available at Bioinformatics online. Daniel K. Witvliet, Alexey Strokach, Andrés Felipe Giraldo-Forero, Joan Teyra, Recep Colak, Philip M. Kim |
Bioinform. | 6 |
| 2016 | PAT: predictor for structured units and its application for the optimization of target molecules for the generation of synthetic antibodiesabstractBACKGROUND: The identification of structured units in a protein sequence is an important first step for most biochemical studies. Importantly for this study, the identification of stable structured region is a crucial first step to generate novel synthetic antibodies. While many approaches to find domains or predict structured regions exist, important limitations remain, such as the optimization of domain boundaries and the lack of identification of non-domain structured units. Moreover, no integrated tool exists to find and optimize structural domains within protein sequences. RESULTS: Here, we describe a new tool, PAT ( http://www.kimlab.org/software/pat ) that can efficiently identify both domains (with optimized boundaries) and non-domain putative structured units. PAT automatically analyzes various structural properties, evaluates the folding stability, and reports possible structural domains in a given protein sequence. For reliability evaluation of PAT, we applied PAT to identify antibody target molecules based on the notion that soluble and well-defined protein secondary and tertiary structures are appropriate target molecules for synthetic antibodies. CONCLUSION: PAT is an efficient and sensitive tool to identify structured units. A performance analysis shows that PAT can characterize structurally well-defined regions in a given sequence and outperforms other efforts to define reliable boundaries of domains. Specially, PAT successfully identifies experimentally confirmed target molecules for antibody generation. PAT also offers the pre-calculated results of 20,210 human proteins to accelerate common queries. PAT can therefore help to investigate large-scale structured domains and improve the success rate for synthetic antibody generation. Jouhyun Jeon, Roland Arnold, Fateh Singh, Joan Teyra, Tatjana Braun, Philip M. Kim |
BMC Bioinform. | 6 |
| 2013 | Distinct Types of Disorder in the Human Proteome: Functional Implications for Alternative SplicingabstractIntrinsically disordered regions have been associated with various cellular processes and are implicated in several human diseases, but their exact roles remain unclear. We previously defined two classes of conserved disordered regions in budding yeast, referred to as "flexible" and "constrained" conserved disorder. In flexible disorder, the property of disorder has been positionally conserved during evolution, whereas in constrained disorder, both the amino acid sequence and the property of disorder have been conserved. Here, we show that flexible and constrained disorder are widespread in the human proteome, and are particularly common in proteins with regulatory functions. Both classes of disordered sequences are highly enriched in regions of proteins that undergo tissue-specific (TS) alternative splicing (AS), but not in regions of proteins that undergo general (i.e., not tissue-regulated) AS. Flexible disorder is more highly enriched in TS alternative exons, whereas constrained disorder is more highly enriched in exons that flank TS alternative exons. These latter regions are also significantly more enriched in potential phosphosites and other short linear motifs associated with cell signaling. We further show that cancer driver mutations are significantly enriched in regions of proteins associated with TS and general AS. Collectively, our results point to distinct roles for TS alternative exons and flanking exons in the dynamic regulation of protein interaction networks in response to signaling activity, and they further suggest that alternatively spliced regions of proteins are often functionally altered by mutations responsible for cancer. Recep Colak, TaeHyung Kim, Magali Michaut, Mark G. F. Sun, Manuel Irimia, Jeremy Bellay, Chad L. Myers, Benjamin J. Blencowe, Philip M. Kim |
PLoS Comput. Biol. | 9 |
| 2012 | Network Evolution: Rewiring and Signatures of Conservation in SignalingabstractThe analysis of network evolution has been hampered by limited availability of protein interaction data for different organisms. In this study, we investigate evolutionary mechanisms in Src Homology 3 (SH3) domain and kinase interaction networks using high-resolution specificity profiles. We constructed and examined networks for 23 fungal species ranging from Saccharomyces cerevisiae to Schizosaccharomyces pombe. We quantify rates of different rewiring mechanisms and show that interaction change through binding site evolution is faster than through gene gain or loss. We found that SH3 interactions evolve swiftly, at rates similar to those found in phosphoregulation evolution. Importantly, we show that interaction changes are sufficiently rapid to exhibit saturation phenomena at the observed timescales. Finally, focusing on the SH3 interaction network, we observe extensive clustering of binding sites on target proteins by SH3 domains and a strong correlation between the number of domains that bind a target protein (target in-degree) and interaction conservation. The relationship between in-degree and interaction conservation is driven by two different effects, namely the number of clusters that correspond to interaction interfaces and the number of domains that bind to each cluster leads to sequence specific conservation, which in turn results in interaction conservation. In summary, we uncover several network evolution mechanisms likely to generalize across peptide recognition modules. Mark G. F. Sun, Martin Sikora, Michael Costanzo, Charles Boone, Philip M. Kim |
PLoS Comput. Biol. | 5 |
| 2011 | Measuring the Evolutionary Rewiring of Biological NetworksabstractWe have accumulated a large amount of biological network data and expect even more to come. Soon, we anticipate being able to compare many different biological networks as we commonly do for molecular sequences. It has long been believed that many of these networks change, or "rewire", at different rates. It is therefore important to develop a framework to quantify the differences between networks in a unified fashion. We developed such a formalism based on analogy to simple models of sequence evolution, and used it to conduct a systematic study of network rewiring on all the currently available biological networks. We found that, similar to sequences, biological networks show a decreased rate of change at large time divergences, because of saturation in potential substitutions. However, different types of biological networks consistently rewire at different rates. Using comparative genomics and proteomics data, we found a consistent ordering of the rewiring rates: transcription regulatory, phosphorylation regulatory, genetic interaction, miRNA regulatory, protein interaction, and metabolic pathway network, from fast to slow. This ordering was found in all comparisons we did of matched networks between organisms. To gain further intuition on network rewiring, we compared our observed rewirings with those obtained from simulation. We also investigated how readily our formalism could be mapped to other network contexts; in particular, we showed how it could be applied to analyze changes in a range of "commonplace" networks such as family trees, co-authorships and linux-kernel function dependencies. Chong Shou, Nitin Bhardwaj, Hugo Y. K. Lam, Koon-Kiu Yan, Philip M. Kim, Michael Snyder 0001, Mark Gerstein |
PLoS Comput. Biol. | 5 |
| 2010 | MOTIPS: Automated Motif Analysis for Predicting Targets of Modular Protein DomainsabstractBACKGROUND: Many protein interactions, especially those involved in signaling, involve short linear motifs consisting of 5-10 amino acid residues that interact with modular protein domains such as the SH3 binding domains and the kinase catalytic domains. One straightforward way of identifying these interactions is by scanning for matches to the motif against all the sequences in a target proteome. However, predicting domain targets by motif sequence alone without considering other genomic and structural information has been shown to be lacking in accuracy. RESULTS: We developed an efficient search algorithm to scan the target proteome for potential domain targets and to increase the accuracy of each hit by integrating a variety of pre-computed features, such as conservation, surface propensity, and disorder. The integration is performed using naïve Bayes and a training set of validated experiments. CONCLUSIONS: By integrating a variety of biologically relevant features to predict domain targets, we demonstrated a notably improved prediction of modular protein domain targets. Combined with emerging high-resolution data of domain specificities, we believe that our approach can assist in the reconstruction of many signaling pathways. Hugo Y. K. Lam, Philip M. Kim, Janine Mok, Raffi Tonikian, Sachdev S. Sidhu, Benjamin E. Turk, Michael Snyder 0001, Mark Gerstein |
BMC Bioinform. | 2 |
| 2009 | Multi-level learning: improving the prediction of protein, domain and residue interactions by allowing information flow between levelsabstractBACKGROUND: Proteins interact through specific binding interfaces that contain many residues in domains. Protein interactions thus occur on three different levels of a concept hierarchy: whole-proteins, domains, and residues. Each level offers a distinct and complementary set of features for computationally predicting interactions, including functional genomic features of whole proteins, evolutionary features of domain families and physical-chemical features of individual residues. The predictions at each level could benefit from using the features at all three levels. However, it is not trivial as the features are provided at different granularity. RESULTS: To link up the predictions at the three levels, we propose a multi-level machine-learning framework that allows for explicit information flow between the levels. We demonstrate, using representative yeast interaction networks, that our algorithm is able to utilize complementary feature sets to make more accurate predictions at the three levels than when the three problems are approached independently. To facilitate application of our multi-level learning framework, we discuss three key aspects of multi-level learning and the corresponding design choices that we have made in the implementation of a concrete learning algorithm. 1) Architecture of information flow: we show the greater flexibility of bidirectional flow over independent levels and unidirectional flow; 2) Coupling mechanism of the different levels: We show how this can be accomplished via augmenting the training sets at each level, and discuss the prevention of error propagation between different levels by means of soft coupling; 3) Sparseness of data: We show that the multi-level framework compounds data sparsity issues, and discuss how this can be dealt with by building local models in information-rich parts of the data. Our proof-of-concept learning algorithm demonstrates the advantage of combining levels, and opens up opportunities for further research. AVAILABILITY: The software and a readme file can be downloaded at http://networks.gersteinlab.org/mll. The programs are written in Java, and can be run on any platform with Java 1.4 or higher and Apache Ant 1.7.0 or higher installed. The software can be used without a license. Kevin Y. Yip, Philip M. Kim, Drew McDermott, Mark Gerstein |
BMC Bioinform. | 2 |
| 2008 | An integrated system for studying residue coevolution in proteinsabstractUNLABELLED: Residue coevolution has recently emerged as an important concept, especially in the context of protein structures. While a multitude of different functions for quantifying it have been proposed, not much is known about their relative strengths and weaknesses. Also, subtle algorithmic details have discouraged implementing and comparing them. We addressed this issue by developing an integrated online system that enables comparative analyses with a comprehensive set of commonly used scoring functions, including Statistical Coupling Analysis (SCA), Explicit Likelihood of Subset Variation (ELSC), mutual information and correlation-based methods. A set of data preprocessing options are provided for improving the sensitivity and specificity of coevolution signal detection, including sequence weighting, residue grouping and the filtering of sequences, sites and site pairs. A total of more than 100 scoring variations are available. The system also provides facilities for studying the relationship between coevolution scores and inter-residue distances from a crystal structure if provided, which may help in understanding protein structures. AVAILABILITY: The system is available at http://coevolution.gersteinlab.org. The source code and JavaDoc API can also be downloaded from the web site. Kevin Y. Yip, Prianka Patel, Philip M. Kim, Donald M. Engelman, Drew McDermott, Mark Gerstein |
Bioinform. | 3 |
| 2007 | The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networksabstractBioinformatics (2006) 22(23), 2968–2970 The authors would like to apologize for the omission of a reference from the below paper. N-Browse was used and referred to in the following article: Lall,S., Grun,D., Krek,A., Chen,K., Wang,Y.L., Dewey,C.N., Sood,P., Colombo,T., Bray,N., Macmenamin,P., Kao,H.L., Gunsalus,K.C., Pachter,L., Piano,F. and Rajewsky,N. (2006) A genome-wide map of conserved microRNA targets in C. elegans. Current Biol., 16, 460–471. Also, it should be noted that the URL of N-Browse given in our paper is a temporary redirect. The official URL should be http://www.gnetbrowse.org/. Kevin Y. Yip, Haiyuan Yu, Philip M. Kim, Martin H. Schultz, Mark Gerstein |
Bioinform. | 3 |
| 2007 | The Importance of Bottlenecks in Protein Networks: Correlation with Gene Essentiality and Expression DynamicsabstractIt has been a long-standing goal in systems biology to find relations between the topological properties and functional features of protein networks. However, most of the focus in network studies has been on highly connected proteins ("hubs"). As a complementary notion, it is possible to define bottlenecks as proteins with a high betweenness centrality (i.e., network nodes that have many "shortest paths" going through them, analogous to major bridges and tunnels on a highway map). Bottlenecks are, in fact, key connector proteins with surprising functional and dynamic properties. In particular, they are more likely to be essential proteins. In fact, in regulatory and other directed networks, betweenness (i.e., "bottleneck-ness") is a much more significant indicator of essentiality than degree (i.e., "hub-ness"). Furthermore, bottlenecks correspond to the dynamic components of the interaction network-they are significantly less well coexpressed with their neighbors than non-bottlenecks, implying that expression dynamics is wired into the network topology. Haiyuan Yu, Philip M. Kim, Emmett Sprecher, Valery Trifonov, Mark Gerstein |
PLoS Comput. Biol. | 2 |
| 2006 | The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networksabstractUNLABELLED: Biological processes involve complex networks of interactions between molecules. Various large-scale experiments and curation efforts have led to preliminary versions of complete cellular networks for a number of organisms. To grapple with these networks, we developed TopNet-like Yale Network Analyzer (tYNA), a Web system for managing, comparing and mining multiple networks, both directed and undirected. tYNA efficiently implements methods that have proven useful in network analysis, including identifying defective cliques, finding small network motifs (such as feed-forward loops), calculating global statistics (such as the clustering coefficient and eccentricity), and identifying hubs and bottlenecks. It also allows one to manage a large number of private and public networks using a flexible tagging system, to filter them based on a variety of criteria, and to visualize them through an interactive graphical interface. A number of commonly used biological datasets have been pre-loaded into tYNA, standardized and grouped into different categories. AVAILABILITY: The tYNA system can be accessed at http://networks.gersteinlab.org/tyna. The source code, JavaDoc API and WSDL can also be downloaded from the website. tYNA can also be accessed from the Cytoscape software using a plugin. Kevin Y. Yip, Haiyuan Yu, Philip M. Kim, Martin H. Schultz, Mark Gerstein |
Bioinform. | 3 |