EDBT 2026 Demo / reviewers in the wild / expert
Jens Meiler
dblp:80/1028
· DBLP profile ↗
28ranked-venue papers
0as first author
10since 2021 · last 2024
0000-0001-8945-193XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 8 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery BenchmarkingabstractWhile deep learning has revolutionized computer-aided drug discovery, the AI community has predominantly focused on model innovation and placed less emphasis on establishing best benchmarking practices. We posit that without a sound model evaluation framework, the AI community's efforts cannot reach their full potential, thereby slowing the progress and transfer of innovation into real-world drug discovery.Thus, in this paper, we seek to establish a new gold standard for small molecule drug discovery benchmarking, WelQrate. Specifically, our contributions are threefold: WelQrate dataset collection - we introduce a meticulously curated collection of 9 datasets spanning 5 therapeutic target classes. Our hierarchical curation pipelines, designed by drug discovery experts, go beyond the primary high-throughput screen by leveraging additional confirmatory and counter screens along with rigorous domain-driven preprocessing, such as Pan-Assay Interference Compounds (PAINS) filtering, to ensure the high-quality data in the datasets; WelQrate Evaluation Framework - we propose a standardized model evaluation framework considering high-quality datasets, featurization, 3D conformation generation, evaluation metrics, and data splits, which provides a reliable benchmarking for drug discovery experts conducting real-world virtual screening; Benchmarking - we evaluate model performance through various research questions using the WelQrate dataset collection, exploring the effects of different models, dataset quality, featurization methods, and data splitting strategies on the results.In summary, we recommend adopting our proposed WelQrate as the gold standard in small molecule drug discovery benchmarking. The WelQrate dataset collection, along with the curation codes, and experimental scripts are all publicly available at www.WelQrate.org. Yunchao Liu 0001, Ha Dong, Xin Wang 0061, Rocco Moretti, Yu Wang 0160, Zhaoqian Su, Jiawei Gu, Bobby Bodenheimer, Charles David Weaver, Jens Meiler, Tyler Derr |
NeurIPS | 10 |
| 2024 | Combining machine learning with structure-based protein design to predict and engineer post-translational modifications of proteinsabstractPost-translational modifications (PTMs) of proteins play a vital role in their function and stability. These modifications influence protein folding, signaling, protein-protein interactions, enzyme activity, binding affinity, aggregation, degradation, and much more. To date, over 400 types of PTMs have been described, representing chemical diversity well beyond the genetically encoded amino acids. Such modifications pose a challenge to the successful design of proteins, but also represent a major opportunity to diversify the protein engineering toolbox. To this end, we first trained artificial neural networks (ANNs) to predict eighteen of the most abundant PTMs, including protein glycosylation, phosphorylation, methylation, and deamidation. In a second step, these models were implemented inside the computational protein modeling suite Rosetta, which allows flexible combination with existing protocols to model the modified sites and understand their impact on protein stability as well as function. Lastly, we developed a new design protocol that either maximizes or minimizes the predicted probability of a particular site being modified. We find that this combination of ANN prediction and structure-based design can enable the modification of existing, as well as the introduction of novel, PTMs. The potential applications of our work include, but are not limited to, glycan masking of epitopes, strengthening protein-protein interactions through phosphorylation, as well as protecting proteins from deamidation liabilities. These applications are especially important for the design of new protein therapeutics where PTMs can drastically change the therapeutic properties of a protein. Our work adds novel tools to Rosetta's protein engineering toolbox that allow for the rational design of PTMs. Moritz Ertelt, Vikram Khipple Mulligan, Jack B. Maguire, Sergey Lyskov, Rocco Moretti, Torben Schiffner, Jens Meiler, Clara T. Schoeder |
PLoS Comput. Biol. | 7 |
| 2023 | Interpretable Chirality-Aware Graph Neural Network for Quantitative Structure Activity Relationship Modeling in Drug DiscoveryabstractIn computer-aided drug discovery, quantitative structure activity relation models are trained to predict biological activity from chemical structure. Despite the recent success of applying graph neural network to this task, important chemical information such as molecular chirality is ignored. To fill this crucial gap, we propose Molecular-Kernel Graph NeuralNetwork (MolKGNN) for molecular representation learning, which features SE(3)-/conformation invariance, chirality-awareness, and interpretability. For our MolKGNN, we first design a molecular graph convolution to capture the chemical pattern by comparing the atom's similarity with the learnable molecular kernels. Furthermore, we propagate the similarity score to capture the higher-order chemical pattern. To assess the method, we conduct a comprehensive evaluation with nine well-curated datasets spanning numerous important drug targets that feature realistic high class imbalance and it demonstrates the superiority of MolKGNN over other graph neural networks in computer-aided drug discovery. Meanwhile, the learned kernels identify patterns that agree with domain knowledge, confirming the pragmatic interpretability of this approach. Our code and supplementary material are publicly available at https://github.com/meilerlab/MolKGNN. Yunchao Liu 0001, Yu Wang 0160, Oanh Vu, Rocco Moretti, Bobby Bodenheimer, Jens Meiler, Tyler Derr |
AAAI | 6 |
| 2023 | Docking cholesterol to integral membrane proteins with RosettaabstractLipid molecules such as cholesterol interact with the surface of integral membrane proteins (IMP) in a mode different from drug-like molecules in a protein binding pocket. These differences are due to the lipid molecule's shape, the membrane's hydrophobic environment, and the lipid's orientation in the membrane. We can use the recent increase in experimental structures in complex with cholesterol to understand protein-cholesterol interactions. We developed the RosettaCholesterol protocol consisting of (1) a prediction phase using an energy grid to sample and score native-like binding poses and (2) a specificity filter to calculate the likelihood that a cholesterol interaction site may be specific. We used a multi-pronged benchmark (self-dock, flip-dock, cross-dock, and global-dock) of protein-cholesterol complexes to validate our method. RosettaCholesterol improved sampling and scoring of native poses over the standard RosettaLigand baseline method in 91% of cases and performs better regardless of benchmark complexity. On the β2AR, our method found one likely-specific site, which is described in the literature. The RosettaCholesterol protocol quantifies cholesterol binding site specificity. Our approach provides a starting point for high-throughput modeling and prediction of cholesterol binding sites for further experimental validation. Brennica Marlow, Georg Kuenze, Jens Meiler, Julia Koehler Leman |
PLoS Comput. Biol. | 3 |
| 2022 | Computational epitope mapping of class I fusion proteins using low complexity supervised learning methodsabstractAntibody epitope mapping of viral proteins plays a vital role in understanding immune system mechanisms of protection. In the case of class I viral fusion proteins, recent advances in cryo-electron microscopy and protein stabilization techniques have highlighted the importance of cryptic or 'alternative' conformations that expose epitopes targeted by potent neutralizing antibodies. Thorough epitope mapping of such metastable conformations is difficult but is critical for understanding sites of vulnerability in class I fusion proteins that occur as transient conformational states during viral attachment and fusion. We introduce a novel method Accelerated class I fusion protein Epitope Mapping (AxIEM) that accounts for fusion protein flexibility to improve out-of-sample prediction of discontinuous antibody epitopes. Harnessing data from previous experimental epitope mapping efforts of several class I fusion proteins, we demonstrate that accuracy of epitope prediction depends on residue environment and allows for the prediction of conformation-dependent antibody target residues. We also show that AxIEM can identify common epitopes and provide structural insights for the development and rational design of vaccines. Marion F. S. Fischer, James E. Crowe Jr., Jens Meiler |
PLoS Comput. Biol. | 3 |
| 2022 | Predicting the functional impact of KCNQ1 variants with artificial neural networksabstractRecent advances in experimental and computational protein structure determination have provided access to high-quality structures for most human proteins and mutants thereof. However, linking changes in structure in protein mutants to functional impact remains an active area of method development. If successful, such methods can ultimately assist physicians in taking appropriate treatment decisions. This work presents three artificial neural network (ANN)-based predictive models that classify four key functional parameters of KCNQ1 variants as normal or dysfunctional using PSSM-based evolutionary and/or biophysical descriptors. Recent advances in predicting protein structure and variant properties with artificial intelligence (AI) rely heavily on the availability of evolutionary features and thus fail to directly assess the biophysical underpinnings of a change in structure and/or function. The central goal of this work was to develop an ANN model based on structure and physiochemical properties of KCNQ1 potassium channels that performs comparably or better than algorithms using only on PSSM-based evolutionary features. These biophysical features highlight the structure-function relationships that govern protein stability, function, and regulation. The input sensitivity algorithm incorporates the roles of hydrophobicity, polarizability, and functional densities on key functional parameters of the KCNQ1 channel. Inclusion of the biophysical features outperforms exclusive use of PSSM-based evolutionary features in predicting activation voltage dependence and deactivation time. As AI is increasingly applied to problems in biology, biophysical understanding will be critical with respect to 'explainable AI', i.e., understanding the relation of sequence, structure, and function of proteins. Our model is available at www.kcnq1predict.org. Saksham Phul, Georg Kuenze, Carlos G. Vanoye, Charles R. Sanders, Alfred L. George Jr., Jens Meiler |
PLoS Comput. Biol. | 6 |
| 2021 | Methodology for rigorous modeling of protein conformational changes by Rosetta using DEER distance restraintsabstractWe describe an approach for integrating distance restraints from Double Electron-Electron Resonance (DEER) spectroscopy into Rosetta with the purpose of modeling alternative protein conformations from an initial experimental structure. Fundamental to this approach is a multilateration algorithm that harnesses sets of interconnected spin label pairs to identify optimal rotamer ensembles at each residue that fit the DEER decay in the time domain. Benchmarked relative to data analysis packages, the algorithm yields comparable distance distributions with the advantage that fitting the DEER decay and rotamer ensemble optimization are coupled. We demonstrate this approach by modeling the protonation-dependent transition of the multidrug transporter PfMATE to an inward facing conformation with a deviation to the experimental structure of less than 2Å Cα RMSD. By decreasing spin label rotamer entropy, this approach engenders more accurate Rosetta models that are also more closely clustered, thus setting the stage for more robust modeling of protein conformational changes. Diego del Alamo, Kevin L. Jagessar, Jens Meiler, Hassane S. Mchaourab |
PLoS Comput. Biol. | 3 |
| 2021 | Computational redesign of a fluorogen activating protein with RosettaabstractThe use of unnatural fluorogenic molecules widely expands the pallet of available genetically encoded fluorescent imaging tools through the design of fluorogen activating proteins (FAPs). While there is already a handful of such probes available, each of them went through laborious cycles of in vitro screening and selection. Computational modeling approaches are evolving incredibly fast right now and are demonstrating great results in many applications, including de novo protein design. It suggests that the easier task of fine-tuning the fluorogen-binding properties of an already functional protein in silico should be readily achievable. To test this hypothesis, we used Rosetta for computational ligand docking followed by protein binding pocket redesign to further improve the previously described FAP DiB1 that is capable of binding to a BODIPY-like dye M739. Despite an inaccurate initial docking of the chromophore, the incorporated mutations nevertheless improved multiple photophysical parameters as well as the overall performance of the tag. The designed protein, DiB-RM, shows higher brightness, localization precision, and apparent photostability in protein-PAINT super-resolution imaging compared to its parental variant DiB1. Moreover, DiB-RM can be cleaved to obtain an efficient split system with enhanced performance compared to a parental DiB-split system. The possible reasons for the inaccurate ligand binding pose prediction and its consequence on the outcome of the design experiment are further discussed. Nina G. Bozhanova, Joel M. Harp, Brian Joseph Bender, Alexey S. Gavrikov, Dmitry A. Gorbachev, Mikhail S. Baranov, Christina B. Mercado, Konstantin A. Lukyanov, Alexander S. Mishin, Jens Meiler |
PLoS Comput. Biol. | 11 |
| 2021 | Prediction of amphipathic helix - membrane interactions with RosettaabstractAmphipathic helices have hydrophobic and hydrophilic/charged residues situated on opposite faces of the helix. They can anchor peripheral membrane proteins to the membrane, be attached to integral membrane proteins, or exist as independent peptides. Despite the widespread presence of membrane-interacting amphipathic helices, there is no computational tool within Rosetta to model their interactions with membranes. In order to address this need, we developed the AmphiScan protocol with PyRosetta, which runs a grid search to find the most favorable position of an amphipathic helix with respect to the membrane. The performance of the algorithm was tested in benchmarks with the RosettaMembrane, ref2015_memb, and franklin2019 score functions on six engineered and 44 naturally-occurring amphipathic helices using membrane coordinates from the OPM and PDBTM databases, OREMPRO server, and MD simulations for comparison. The AmphiScan protocol predicted the coordinates of amphipathic helices within less than 3Å of the reference structures and identified membrane-embedded residues with a Matthews Correlation Constant (MCC) of up to 0.57. Overall, AmphiScan stands as fast, accurate, and highly-customizable protocol that can be pipelined with other Rosetta and Python applications. Alican Gulsevin, Jens Meiler |
PLoS Comput. Biol. | 2 |
| 2021 | Rosetta design with co-evolutionary information retains protein functionabstractComputational protein design has the ambitious goal of crafting novel proteins that address challenges in biology and medicine. To overcome these challenges, the computational protein modeling suite Rosetta has been tailored to address various protein design tasks. Recently, statistical methods have been developed that identify correlated mutations between residues in a multiple sequence alignment of homologous proteins. These subtle inter-dependencies in the occupancy of residue positions throughout evolution are crucial for protein function, but we found that three current Rosetta design approaches fail to recover these co-evolutionary couplings. Thus, we developed the Rosetta method ResCue (residue-coupling enhanced) that leverages co-evolutionary information to favor sequences which recapitulate correlated mutations, as observed in nature. To assess the protocols via recapitulation designs, we compiled a benchmark of ten proteins each represented by two, structurally diverse states. We could demonstrate that ResCue designed sequences with an average sequence recovery rate of 70%, whereas three other protocols reached not more than 50%, on average. Our approach had higher recovery rates also for functionally important residues, which were studied in detail. This improvement has only a minor negative effect on the fitness of the designed sequences as assessed by Rosetta energy. In conclusion, our findings support the idea that informing protocols with co-evolutionary signals helps to design stable and native-like proteins that are compatible with the different conformational states required for a complex function. Samuel Schmitz, Moritz Ertelt, Rainer Merkl, Jens Meiler |
PLoS Comput. Biol. | 4 |
| 2020 | 3D Deep Learning for Biological Function Prediction from Physical FieldsabstractPredicting the biological function of molecules, be it proteins or drug-like compounds, from their atomic structure is an important and long-standing problem. The electron density field and electrostatic potential field of a molecule contain the “raw fingerprint” of how this molecule can fit to binding partners. In this paper, we show that deep learning can predict biological function of molecules directly from their raw 3D approximated electron density and electrostatic potential fields. Protein function based on Enzyme Commission numbers is predicted from the approximated electron density field. In another experiment, the activity of small molecules is predicted with quality comparable to state-of-the-art descriptor-based methods. We propose several alternative computational models for the GPU with different memory and runtime requirements for different sizes of molecules and of databases. We also propose application-specific multi-channel data representations. Vladimir Golkov, Marcin J. Skwark, Atanas Mirchev, Georgi Dikov, Alexander R. Geanes, Jeffrey L. Mendenhall, Jens Meiler, Daniel Cremers |
3DV | 7 |
| 2020 | Foldit Drug Design Game Usability Study: Comparison of Citizen and Expert ScientistsabstractIn building a new drug design mode for the popular citizen scientist game Foldit, we focus on creating an easy-to-use and intuitive interface to confer complex scientific concepts to citizen scientist players. We hypothesize that to be efficient in the hands of citizen scientists such an interface will look different from well-established drug-design software used by experts. We used the relaxed think-aloud method to compare citizen and expert scientists working with our prototype interface for Foldit Drug Design Mode (FDDM). First, we tested if the two groups are providing different feedback when it comes to the usability of the prototype interface. Second, we investigated how the difference between the two groups might inform a new game design. As expected, the results confirm that experienced scientists differ from citizen scientists in engaging their background knowledge when interacting with the game. We then provided a prioritization list of background knowledge employed by the expert scientists to derive design suggestions for FDDM. Yunchao Liu 0001, Rocco Moretti, Bobby Bodenheimer, Jens Meiler |
MIG | 4 |
| 2020 | PyIR: a scalable wrapper for processing billions of immunoglobulin and T cell receptor sequences using IgBLASTabstractBACKGROUND: Recent advances in DNA sequencing technologies have enabled significant leaps in capacity to generate large volumes of DNA sequence data, which has spurred a rapid growth in the use of bioinformatics as a means of interrogating antibody variable gene repertoires. Common tools used for annotation of antibody sequences are often limited in functionality, modularity and usability. RESULTS: We have developed PyIR, a Python wrapper and library for IgBLAST, which offers a minimal setup CLI and API, FASTQ support, file chunking for large sequence files, JSON and Python dictionary output, and built-in sequence filtering. CONCLUSIONS: PyIR offers improved processing speed over multithreaded IgBLAST (version 1.14) when spawning more than 16 processes on a single computer system. Its customizable filtering and data encapsulation allow it to be adapted to a wide range of computing environments. The API allows for IgBLAST to be used in customized bioinformatics workflows. Cinque S. Soto, Jessica A. Finn, Jordan R. Willis, Samuel B. Day, Robert S. Sinkovits, Taylor Jones, Samuel Schmitz, Jens Meiler, Andre Branchizio, James E. Crowe Jr. |
BMC Bioinform. | 8 |
| 2020 | Improving homology modeling from low-sequence identity templates in Rosetta: A case study in GPCRsabstractAs sequencing methodologies continue to advance, the availability of protein sequences far outpaces the ability of structure determination. Homology modeling is used to bridge this gap but relies on high-identity templates for accurate model building. G-protein coupled receptors (GPCRs) represent a significant target class for pharmaceutical therapies in which homology modeling could fill the knowledge gap for structure-based drug design. To date, only about 17% of druggable GPCRs have had their structures characterized at atomic resolution. However, modeling of the remaining 83% is hindered by the low sequence identity between receptors. Here we test key inputs in the model building process using GPCRs as a focus to improve the pipeline in two critical ways: Firstly, we use a blended sequence- and structure-based alignment that accounts for structure conservation in loop regions. Secondly, by merging multiple template structures into one comparative model, the best possible template for every region of a target can be used expanding the conformational space sampled in a meaningful way. This optimization allows for accurate modeling of receptors using templates as low as 20% sequence identity, which accounts for nearly the entire druggable space of GPCRs. A model database of all non-odorant GPCRs is made available at www.rosettagpcr.org. Additionally, all protocols are made available with insights into modifications that may improve accuracy at new targets. Brian Joseph Bender, Brennica Marlow, Jens Meiler |
PLoS Comput. Biol. | 3 |
| 2020 | Better together: Elements of successful scientific software development in a distributed collaborative communityabstractMany scientific disciplines rely on computational methods for data analysis, model generation, and prediction. Implementing these methods is often accomplished by researchers with domain expertise but without formal training in software engineering or computer science. This arrangement has led to underappreciation of sustainability and maintainability of scientific software tools developed in academic environments. Some software tools have avoided this fate, including the scientific library Rosetta. We use this software and its community as a case study to show how modern software development can be accomplished successfully, irrespective of subject area. Rosetta is one of the largest software suites for macromolecular modeling, with 3.1 million lines of code and many state-of-the-art applications. Since the mid 1990s, the software has been developed collaboratively by the RosettaCommons, a community of academics from over 60 institutions worldwide with diverse backgrounds including chemistry, biology, physiology, physics, engineering, mathematics, and computer science. Developing this software suite has provided us with more than two decades of experience in how to effectively develop advanced scientific software in a global community with hundreds of contributors. Here we illustrate the functioning of this development community by addressing technical aspects (like version control, testing, and maintenance), community-building strategies, diversity efforts, software dissemination, and user support. We demonstrate how modern computational research can thrive in a distributed collaborative community. The practices described here are independent of subject area and can be readily adopted by other software development communities. Julia Koehler Leman, Brian D. Weitzner, P. Douglas Renfrew, Steven M. Lewis, Rocco Moretti, Andrew M. Watkins, Vikram Khipple Mulligan, Sergey Lyskov, Jared Adolf-Bryfogle, Jason W. Labonte, Justyna Krys, Christopher Bystroff, William R. Schief, Dominik Gront, Ora Schueler-Furman, David Baker 0001, Philip Bradley, Roland L. Dunbrack Jr., Tanja Kortemme, Andrew Leaver-Fay, Charlie E. M. Strauss, Jens Meiler, Brian Kuhlman, Jeffrey J. Gray, Richard Bonneau |
PLoS Comput. Biol. | 23 |
| 2020 | Multi-state design of flexible proteins predicts sequences optimal for conformational changeabstractComputational protein design of an ensemble of conformations for one protein-i.e., multi-state design-determines the side chain identity by optimizing the energetic contributions of that side chain in each of the backbone conformations. Sampling the resulting large sequence-structure search space limits the number of conformations and the size of proteins in multi-state design algorithms. Here, we demonstrated that the REstrained CONvergence (RECON) algorithm can simultaneously evaluate the sequence of large proteins that undergo substantial conformational changes. Simultaneous optimization of side chain conformations across all conformations increased sequence conservation when compared to single-state designs in all cases. More importantly, the sequence space sampled by RECON MSD resembled the evolutionary sequence space of flexible proteins, particularly when confined to predicting the mutational preferences of limited common ancestral descent, such as in the case of influenza type A hemagglutinin. Additionally, we found that sequence positions which require substantial changes in their local environment across an ensemble of conformations are more likely to be conserved. These increased conservation rates are better captured by RECON MSD over multiple conformations and thus multiple local residue environments during design. To quantify this rewiring of contacts at a certain position in sequence and structure, we introduced a new metric designated 'contact proximity deviation' that enumerates contact map changes. This measure allows mapping of global conformational changes into local side chain proximity adjustments, a property not captured by traditional global similarity metrics such as RMSD or local similarity metrics such as changes in φ and ψ angles. Marion F. S. Fischer, Alexander M. Sevy, James E. Crowe Jr., Jens Meiler |
PLoS Comput. Biol. | 4 |
| 2019 | Immune repertoire fingerprinting by principal component analysis reveals shared features in subject groups with common exposuresabstractBACKGROUND: Advances in next-generation sequencing (NGS) of antibody repertoires have led to an explosion in B cell receptor sequence data from donors with many different disease states. These data have the potential to detect patterns of immune response across populations. However, to this point it has been difficult to interpret such patterns of immune response between disease states in the absence of functional data. There is a need for a robust method that can be used to distinguish general patterns of immune responses at the antibody repertoire level. RESULTS: We developed a method for reducing the complexity of antibody repertoire datasets using principal component analysis (PCA) and refer to our method as "repertoire fingerprinting." We reduce the high dimensional space of an antibody repertoire to just two principal components that explain the majority of variation in those repertoires. We show that repertoires from individuals with a common experience or disease state can be clustered by their repertoire fingerprints to identify common antibody responses. CONCLUSIONS: Our repertoire fingerprinting method for distinguishing immune repertoires has implications for characterizing an individual disease state. Methods to distinguish disease states based on pattern recognition in the adaptive immune response could be used to develop biomarkers with diagnostic or prognostic utility in patient care. Extending our analysis to larger cohorts of patients in the future should permit us to define more precisely those characteristics of the immune response that result from natural infection or autoimmunity. Alexander M. Sevy, Cinque S. Soto, Robin G. Bombardi, Jens Meiler, James E. Crowe Jr. |
BMC Bioinform. | 4 |
| 2018 | Three-dimensional spatial analysis of missense variants in RTEL1 identifies pathogenic variants in patients with Familial Interstitial PneumoniaabstractBACKGROUND: Next-generation sequencing of individuals with genetic diseases often detects candidate rare variants in numerous genes, but determining which are causal remains challenging. We hypothesized that the spatial distribution of missense variants in protein structures contains information about function and pathogenicity that can help prioritize variants of unknown significance (VUS) and elucidate the structural mechanisms leading to disease. RESULTS: To illustrate this approach in a clinical application, we analyzed 13 candidate missense variants in regulator of telomere elongation helicase 1 (RTEL1) identified in patients with Familial Interstitial Pneumonia (FIP). We curated pathogenic and neutral RTEL1 variants from the literature and public databases. We then used homology modeling to construct a 3D structural model of RTEL1 and mapped known variants into this structure. We next developed a pathogenicity prediction algorithm based on proximity to known disease causing and neutral variants and evaluated its performance with leave-one-out cross-validation. We further validated our predictions with segregation analyses, telomere lengths, and mutagenesis data from the homologous XPD protein. Our algorithm for classifying RTEL1 VUS based on spatial proximity to pathogenic and neutral variation accurately distinguished 7 known pathogenic from 29 neutral variants (ROC AUC = 0.85) in the N-terminal domains of RTEL1. Pathogenic proximity scores were also significantly correlated with effects on ATPase activity (Pearson r = -0.65, p = 0.0004) in XPD, a related helicase. Applying the algorithm to 13 VUS identified from sequencing of RTEL1 from patients predicted five out of six disease-segregating VUS to be pathogenic. We provide structural hypotheses regarding how these mutations may disrupt RTEL1 ATPase and helicase function. CONCLUSIONS: Spatial analysis of missense variation accurately classified candidate VUS in RTEL1 and suggests how such variants cause disease. Incorporating spatial proximity analyses into other pathogenicity prediction tools may improve accuracy for other genes and genetic diseases. R. Michael Sivley, Jonathan H. Sheehan, Jonathan A. Kropski, Joy D. Cogan, Timothy S. Blackwell, John A. Phillips, William S. Bush, Jens Meiler, John A. Capra |
BMC Bioinform. | 8 |
| 2018 | Integrating linear optimization with structural modeling to increase HIV neutralization breadthabstractComputational protein design has been successful in modeling fixed backbone proteins in a single conformation. However, when modeling large ensembles of flexible proteins, current methods in protein design have been insufficient. Large barriers in the energy landscape are difficult to traverse while redesigning a protein sequence, and as a result current design methods only sample a fraction of available sequence space. We propose a new computational approach that combines traditional structure-based modeling using the Rosetta software suite with machine learning and integer linear programming to overcome limitations in the Rosetta sampling methods. We demonstrate the effectiveness of this method, which we call BROAD, by benchmarking the performance on increasing predicted breadth of anti-HIV antibodies. We use this novel method to increase predicted breadth of naturally-occurring antibody VRC23 against a panel of 180 divergent HIV viral strains and achieve 100% predicted binding against the panel. In addition, we compare the performance of this method to state-of-the-art multistate design in Rosetta and show that we can outperform the existing method significantly. We further demonstrate that sequences recovered by this method recover known binding motifs of broadly neutralizing anti-HIV antibodies. Finally, our approach is general and can be extended easily to other protein systems. Although our modeled antibodies were not tested in vitro, we predict that these variants would have greatly increased breadth compared to the wild-type antibody. Alexander M. Sevy, Swetasudha Panda, James E. Crowe Jr., Jens Meiler, Yevgeniy Vorobeychik |
PLoS Comput. Biol. | 4 |
| 2016 | Protein contact prediction from amino acid co-evolution using convolutional networks for graph-valued imagesabstractProteins are the "building blocks of life", the most abundant organic molecules, and the central focus of most areas of biomedicine. Protein structure is strongly related to protein function, thus structure prediction is a crucial task on the way to solve many biological questions. A contact map is a compact representation of the three-dimensional structure of a protein via the pairwise contacts between the amino acid constituting the protein. We use a convolutional network to calculate protein contact maps from inferred statistical coupling between positions in the protein sequence. The input to the network has an image-like structure amenable to convolutions, but every "pixel" instead of color channels contains a bipartite undirected edge-weighted graph. We propose several methods for treating such "graph-valued images" in a convolutional network. The proposed method outperforms state-of-the-art methods by a large margin. It also allows for a great flexibility with regard to the input data, which makes it useful for studying a wide range of problems. Vladimir Golkov, Marcin J. Skwark, Antonij Golkov, Alexey Dosovitskiy, Thomas Brox, Jens Meiler, Daniel Cremers |
NIPS | 6 |
| 2015 | Design of Protein Multi-specificity Using an Independent Sequence Search Reduces the Barrier to Low Energy SequencesabstractComputational protein design has found great success in engineering proteins for thermodynamic stability, binding specificity, or enzymatic activity in a 'single state' design (SSD) paradigm. Multi-specificity design (MSD), on the other hand, involves considering the stability of multiple protein states simultaneously. We have developed a novel MSD algorithm, which we refer to as REstrained CONvergence in multi-specificity design (RECON). The algorithm allows each state to adopt its own sequence throughout the design process rather than enforcing a single sequence on all states. Convergence to a single sequence is encouraged through an incrementally increasing convergence restraint for corresponding positions. Compared to MSD algorithms that enforce (constrain) an identical sequence on all states the energy landscape is simplified, which accelerates the search drastically. As a result, RECON can readily be used in simulations with a flexible protein backbone. We have benchmarked RECON on two design tasks. First, we designed antibodies derived from a common germline gene against their diverse targets to assess recovery of the germline, polyspecific sequence. Second, we design "promiscuous", polyspecific proteins against all binding partners and measure recovery of the native sequence. We show that RECON is able to efficiently recover native-like, biologically relevant sequences in this diverse set of protein complexes. Alexander M. Sevy, Tim M. Jacobs, James E. Crowe Jr., Jens Meiler |
PLoS Comput. Biol. | 4 |
| 2013 | Human Germline Antibody Gene Segments Encode Polyspecific AntibodiesabstractStructural flexibility in germline gene-encoded antibodies allows promiscuous binding to diverse antigens. The binding affinity and specificity for a particular epitope typically increase as antibody genes acquire somatic mutations in antigen-stimulated B cells. In this work, we investigated whether germline gene-encoded antibodies are optimal for polyspecificity by determining the basis for recognition of diverse antigens by antibodies encoded by three VH gene segments. Panels of somatically mutated antibodies encoded by a common VH gene, but each binding to a different antigen, were computationally redesigned to predict antibodies that could engage multiple antigens at once. The Rosetta multi-state design process predicted antibody sequences for the entire heavy chain variable region, including framework, CDR1, and CDR2 mutations. The predicted sequences matched the germline gene sequences to a remarkable degree, revealing by computational design the residues that are predicted to enable polyspecificity, i.e., binding of many unrelated antigens with a common sequence. The process thereby reverses antibody maturation in silico. In contrast, when designing antibodies to bind a single antigen, a sequence similar to that of the mature antibody sequence was returned, mimicking natural antibody maturation in silico. We demonstrated that the Rosetta computational design algorithm captures important aspects of antibody/antigen recognition. While the hypervariable region CDR3 often mediates much of the specificity of mature antibodies, we identified key positions in the VH gene encoding CDR1, CDR2, and the immunoglobulin framework that are critical contributors for polyspecificity in germline antibodies. Computational design of antibodies capable of binding multiple antigens may allow the rational design of antibodies that retain polyspecificity for diverse epitope binding. Jordan R. Willis, Bryan S. Briney, Samuel L. DeLuca, James E. Crowe Jr., Jens Meiler |
PLoS Comput. Biol. | 5 |
| 2012 | Bcl∷ChemInfo - Qualitative analysis of machine learning models for activation of HSD involved in Alzheimer's DiseaseabstractIn this case study, a ligand-based virtual high throughput screening suite, bcl::ChemInfo, was applied to screen for activation of the protein target 17-beta hydroxysteroid dehydrogenase type 10 (HSD) involved in Alzheimer's Disease. bcl::ChemInfo implements a diverse set of machine learning techniques such as artificial neural networks (ANN), support vector machines (SVM) with the extension for regression, kappa nearest neighbor (KNN), and decision trees (DT). Molecular structures were converted into a distinct collection of descriptor groups involving 2D- and 3D-autocorrelation, and radial distribution functions. A confirmatory high-throughput screening data set contained over 72,000 experimentally validated compounds, available through PubChem. Here, the systematical model development was achieved through optimization of feature sets and algorithmic parameters resulting in a theoretical enrichment of 11 (44% of maximal enrichment), and an area under the ROC curve (AUC) of 0.75 for the best performing machine learning technique on an independent data set. In addition, consensus combinations of all involved predictors were evaluated and achieved the best enrichment of 13 (50%), and AUC of 0.86. All models were computed in silico and represent a viable option in guiding the drug discovery process through virtual library screening and compound prioritization a priori to synthesis and biological testing. The best consensus predictor will be made accessible for the academic community at www.meilerlab.org. Mariusz Butkiewicz, Edward W. Lowe, Jens Meiler |
CIBCB | 3 |
| 2012 | GPU-accelerated machine learning techniques enable QSAR modeling of large HTS dataabstractQuantitative structure activity relationship (QSAR) modeling using high-throughput screening (HTS) data is a powerful technique which enables the construction of predictive models. These models are utilized for the in silico screening of libraries of molecules for which experimental screening methods are both cost- and time-expensive. Machine learning techniques excel in QSAR modeling where the relationship between structure and activity is often complex and non-linear. As these HTS data sets continue to increase in number of compounds screened, extensive feature selection and cross validation becomes computationally expensive. Leveraging massively parallel architectures such as graphics processing units (GPUs) to accelerate the training algorithms for these machine learning techniques is a cost-efficient manner in which to combat this problem. In this work, several machine learning techniques are ported in OpenCL for GPU-acceleration to enable construction of QSAR ensemble models using HTS data. We report computational performance numbers using several HTS data sets freely available from PubChem database. We also report results of a case study using HTS data for a target of pharmacological and pharmaceutical relevance, cytochrome P450 3A4, for which an enrichment of 94% of the theoretical maximum is achieved. Edward W. Lowe, Mariusz Butkiewicz, Nils Woetzel, Jens Meiler |
CIBCB | 4 |
| 2011 | Comparative analysis of machine learning techniques for the prediction of logPabstractSeveral machine learning techniques were evaluated for the prediction of logP. The algorithms used include artificial neural networks (ANN), support vector machines (SVM) with the extension for regression, and kappa nearest neighbor (k-NN). Molecules were described using optimized feature sets derived from a series of scalar, two- and three-dimensional descriptors including 2-D and 3-D autocorrelation, and radial distribution function. Feature optimization was performed as a sequential forward feature selection. The data set contained over 25,000 molecules with experimentally determined logP values collected from the Reaxys and MDDR databases, as well as data mining through SciFinder. LogP, the logarithm of the equilibrium octanol-water partition coefficient for a given substance is a metric of the hydrophobicity. This property is an important metric for drug absorption, distribution, metabolism, and excretion (ADME). In this work, models were built by systematically optimizing feature sets and algorithmic parameters that predict logP with a root mean square deviation (rmsd) of 0.86 for compounds in an independent test set. This result presents a substantial improvement over XlogP, an incremental system that achieves a rmsd of 1.41 over the same dataset. The final models were 5-fold cross-validated. These fully in silico models can be useful in guiding early stages of drug discovery, such as virtual library screening and analogue prioritization prior to synthesis and biological testing. These models are freely available for academic use. Edward W. Lowe, Mariusz Butkiewicz, Matthew Spellings, Albert Omlor, Jens Meiler |
CIBCB | 5 |
| 2009 | Application of machine learning approaches on quantitative structure activity relationshipsabstractMachine Learning techniques are successfully applied to establish quantitative relations between chemical structure and biological activity (QSAR), i.e. classify compounds as active or inactive with respect to a specific target biological system. This paper presents a comparison of artificial neural networks (ANN), support vector machines (SVM), and decision trees (DT) in an effort to identify potentiators of metabotropic glutamate receptor 5 (mGluR5), compounds that have potential as novel treatments against schizophrenia. When training and testing each of the three techniques on the same dataset enrichments of 61, 64, and 43 were obtained and an area under the curve (AUC) of 0.77, 0.78, and 0.63 was determined for ANNs, SVMs, and DTs, respectively. For the top percentile of predicted active compounds, the true positives for all three methods were highly similar, while the inactives were diverse offering the potential use of jury approaches to improve prediction accuracy. Mariusz Butkiewicz, Ralf Mueller, Danilo Selic, Eric Dawson, Jens Meiler |
CIBCB | 5 |
| 2009 | Improved prediction of trans-membrane spans in proteins using an artificial neural networkabstractTools for the identification of trans-membrane spans from the protein sequence are widely used in the experimental community. Computational structural biology seeks to increase the prediction accuracy of such methods since they represent a first step towards membrane protein tertiary structure prediction from the amino acid sequence. We introduce a predictor that is able to identify trans-membrane spans from the sequence of a protein. The novelty of the approach presented here is the simultaneous prediction of trans-membrane spanning alpha-helices and beta-strands within a single tool. An artificial neural network was trained on databases of 102 membrane proteins and 3499 soluble proteins. Prediction accuracies of up to 92% for soluble residues, 75% for residues in the interface, and 73% for TM residues are achieved. On average the algorithm predicts 79% of the residues correctly which is a substantial improvement from a previously published implementation which achieved 57% accuracy (Koehler et al., Proteins: Structure, Function, and Bioinformatics, 2008). The algorithm was applied to four membrane proteins to illustrate the applicability to both alpha-helical bundles and beta-barrels. Julia Koehler Leman, Ralf Mueller, Jens Meiler |
CIBCB | 3 |
| 2009 | A Correspondence Between Solution-State Dynamics of an Individual Protein and the Sequence and Conformational Diversity of its FamilyabstractConformational ensembles are increasingly recognized as a useful representation to describe fundamental relationships between protein structure, dynamics and function. Here we present an ensemble of ubiquitin in solution that is created by sampling conformational space without experimental information using "Backrub" motions inspired by alternative conformations observed in sub-Angstrom resolution crystal structures. Backrub-generated structures are then selected to produce an ensemble that optimizes agreement with nuclear magnetic resonance (NMR) Residual Dipolar Couplings (RDCs). Using this ensemble, we probe two proposed relationships between properties of protein ensembles: (i) a link between native-state dynamics and the conformational heterogeneity observed in crystal structures, and (ii) a relation between dynamics of an individual protein and the conformational variability explored by its natural family. We show that the Backrub motional mechanism can simultaneously explore protein native-state dynamics measured by RDCs, encompass the conformational variability present in ubiquitin complex structures and facilitate sampling of conformational and sequence variability matching those occurring in the ubiquitin protein family. Our results thus support an overall relation between protein dynamics and conformational changes enabling sequence changes in evolution. More practically, the presented method can be applied to improve protein design predictions by accounting for intrinsic native-state dynamics. Gregory D. Friedland, Nils-Alexander Lakomek, Christian Griesinger, Jens Meiler, Tanja Kortemme |
PLoS Comput. Biol. | 4 |