VLDB 2026 Research / reviewers in the wild / expert
Konrad Krawczyk
dblp:141/5208
· DBLP profile ↗
11ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-0697-5522ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RIOT - Rapid Immunoglobulin Overview Tool - annotation of nucleotide and amino acid immunoglobulin sequences using an open germline databaseabstractAntibodies are a cornerstone of the immune system, playing a pivotal role in identifying and neutralizing infections caused by bacteria, viruses, and other pathogens. Understanding their structure, and function, can provide insights into both the body's natural defenses and the principles behind many therapeutic interventions, including vaccines and antibody-based drugs. The analysis and annotation of antibody sequences, including the identification of variable, diversity, joining, and constant genes, as well as the delineation of framework regions and complementarity-determining regions, is essential for understanding their structure and function. Currently analyzing large volumes of antibody sequences is routine in antibody discovery, requiring fast and accurate tools. While there are existing tools designed for the annotation and numbering of antibody sequences, they often have limitations such as being restricted to either nucleotide or amino acid sequences; slow execution times; or reliance on germline databases that are closed, frequently changed, or have sparse coverage for some species. Here, we present the Rapid Immunoglobulin Overview Tool (RIOT), a novel open-source solution for antibody numbering that addresses these shortcomings. RIOT handles nucleotide and amino acid sequence processing, comes integrated with an Open Germline Receptor Database, and is computationally efficient. We hope that the tool will facilitate rapid annotation of antibody sequencing outputs for the benefit of understanding antibody biology and discovering novel therapeutics. Pawel Dudzic, Bartosz Janusz, Tadeusz Satlawa, Dawid Chomicz, Tomasz Gawlowski, Rafal Grabowski, Przemek Józwiak, Mateusz Tarkowski, Maciej Mycielski, Sonia Wróbel, Konrad Krawczyk |
Briefings Bioinform. | 11 |
| 2024 | LAP: Liability Antibody Profiler by sequence & structural mapping of natural and therapeutic antibodiesabstractAntibody-based therapeutics must not undergo chemical modifications that would impair their efficacy or hinder their developability. A commonly used technique to de-risk lead biotherapeutic candidates annotates chemical liability motifs on their sequence. By analyzing sequences from all major sources of data (therapeutics, patents, GenBank, literature, and next-generation sequencing outputs), we find that almost all antibodies contain an average of 3-4 such liability motifs in their paratopes, irrespective of the source dataset. This is in line with the common wisdom that liability motif annotation is over-predictive. Therefore, we have compiled three computational flags to prioritize liability motifs for removal from lead drug candidates: 1. germline, to reflect naturally occurring motifs, 2. therapeutic, reflecting chemical liability motifs found in therapeutic antibodies, and 3. surface, indicative of structural accessibility for chemical modification. We show that these flags annotate approximately 60% of liability motifs as benign, that is, the flagged liabilities have a smaller probability of undergoing degradation as benchmarked on two experimental datasets covering deamidation, isomerization, and oxidation. We combined the liability detection and flags into a tool called Liability Antibody Profiler (LAP), publicly available at lap.naturalantibody.com. We anticipate that LAP will save time and effort in de-risking therapeutic molecules. Tadeusz Satlawa, Mateusz Tarkowski, Sonia Wróbel, Pawel Dudzic, Tomasz Gawlowski, Tomasz Klaus, Marek Orlowski, Anna Kostyn, Andrew Buchanan, Konrad Krawczyk |
PLoS Comput. Biol. | 11 |
| 2023 | Comparison of Different Surrogate Models for the JADE Algorithm
Konrad Krawczyk, Jaroslaw Arabas |
IJCCI | 1 |
| 2022 | Machine-designed biotherapeutics: opportunities, feasibility and advantages of deep learning in computational antibody discoveryabstractAntibodies are versatile molecular binders with an established and growing role as therapeutics. Computational approaches to developing and designing these molecules are being increasingly used to complement traditional lab-based processes. Nowadays, in silico methods fill multiple elements of the discovery stage, such as characterizing antibody-antigen interactions and identifying developability liabilities. Recently, computational methods tackling such problems have begun to follow machine learning paradigms, in many cases deep learning specifically. This paradigm shift offers improvements in established areas such as structure or binding prediction and opens up new possibilities such as language-based modeling of antibody repertoires or machine-learning-based generation of novel sequences. In this review, we critically examine the recent developments in (deep) machine learning approaches to therapeutic antibody design with implications for fully computational antibody design. Wiktoria Wilman, Sonia Wróbel, Weronika Bielska, Piotr Deszynski, Pawel Dudzic, Igor Jaszczyszyn, Jedrzej Kaniewski, Jakub Mlokosiewicz, Anahita Rouyan, Tadeusz Satlawa, Victor Greiff, Konrad Krawczyk |
Briefings Bioinform. | 13 |
| 2022 | AbDiver: a tool to explore the natural antibody landscape to aid therapeutic designabstractMOTIVATION: Rational design of therapeutic antibodies can be improved by harnessing the natural sequence diversity of these molecules. Our understanding of the diversity of antibodies has recently been greatly facilitated through the deposition of hundreds of millions of human antibody sequences in next-generation sequencing (NGS) repositories. Contrasting a query therapeutic antibody sequence to naturally observed diversity in similar antibody sequences from NGS can provide a mutational roadmap for antibody engineers designing biotherapeutics. Because of the sheer scale of the antibody NGS datasets, performing queries across them is computationally challenging. RESULTS: To facilitate harnessing antibody NGS data, we developed AbDiver (http://naturalantibody.com/abdiver), a free portal allowing users to compare their query sequences to those observed in the natural repertoires. AbDiver offers three antibody-specific use-cases: (i) compare a query antibody to positional variability statistics precomputed from multiple independent studies, (ii) retrieve close full variable sequence matches to a query antibody and (iii) retrieve CDR3 or clonotype matches to a query antibody. We applied our system to a set of 742 therapeutic antibodies, demonstrating that for each use-case our system can retrieve relevant results for most sequences. AbDiver facilitates the navigation of vast antibody mutation space for the purpose of rational therapeutic antibody design. AVAILABILITY AND IMPLEMENTATION: AbDiver is freely accessible at http://naturalantibody.com/abdiver. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jakub Mlokosiewicz, Piotr Deszynski, Wiktoria Wilman, Igor Jaszczyszyn, Rajkumar Ganesan, Aleksandr Kovaltsuk, Jinwoo Leem, Jacob D. Galson, Konrad Krawczyk |
Bioinform. | 9 |
| 2022 | MS2AI: automated repurposing of public peptide LC-MS data for machine learning applicationsabstractMOTIVATION: Liquid-chromatography mass-spectrometry (LC-MS) is the established standard for analyzing the proteome in biological samples by identification and quantification of thousands of proteins. Machine learning (ML) promises to considerably improve the analysis of the resulting data, however, there is yet to be any tool that mediates the path from raw data to modern ML applications. More specifically, ML applications are currently hampered by three major limitations: (i) absence of balanced training data with large sample size; (ii) unclear definition of sufficiently information-rich data representations for e.g. peptide identification; (iii) lack of benchmarking of ML methods on specific LC-MS problems. RESULTS: We created the MS2AI pipeline that automates the process of gathering vast quantities of MS data for large-scale ML applications. The software retrieves raw data from either in-house sources or from the proteomics identifications database, PRIDE. Subsequently, the raw data are stored in a standardized format amenable for ML, encompassing MS1/MS2 spectra and peptide identifications. This tool bridges the gap between MS and AI, and to this effect we also present an ML application in the form of a convolutional neural network for the identification of oxidized peptides. AVAILABILITY AND IMPLEMENTATION: An open-source implementation of the software can be found at https://gitlab.com/roettgerlab/ms2ai. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tobias Greisager Rehfeldt, Konrad Krawczyk, Mathias Bøgebjerg, Veit Schwämmle, Richard Röttger |
Bioinform. | 2 |
| 2020 | Computational approaches to therapeutic antibody design: established methods and emerging trendsabstractAntibodies are proteins that recognize the molecular surfaces of potentially noxious molecules to mount an adaptive immune response or, in the case of autoimmune diseases, molecules that are part of healthy cells and tissues. Due to their binding versatility, antibodies are currently the largest class of biotherapeutics, with five monoclonal antibodies ranked in the top 10 blockbuster drugs. Computational advances in protein modelling and design can have a tangible impact on antibody-based therapeutic development. Antibody-specific computational protocols currently benefit from an increasing volume of data provided by next generation sequencing and application to related drug modalities based on traditional antibodies, such as nanobodies. Here we present a structured overview of available databases, methods and emerging trends in computational antibody analysis and contextualize them towards the engineering of candidate antibody therapeutics. Richard A. Norman, Francesco Ambrosetti, Alexandre M. J. J. Bonvin, Lucy J. Colwell, Sebastian Kelm, Konrad Krawczyk |
Briefings Bioinform. | 7 |
| 2020 | The evolution of contact prediction: evidence that contact selection in statistical contact prediction is changingabstractMOTIVATION: Over the last few years, the field of protein structure prediction has been transformed by increasingly accurate contact prediction software. These methods are based on the detection of coevolutionary relationships between residues from multiple sequence alignments (MSAs). However, despite speculation, there is little evidence of a link between contact prediction and the physico-chemical interactions which drive amino-acid coevolution. Furthermore, existing protocols predict only a fraction of all protein contacts and it is not clear why some contacts are favoured over others. Using a dataset of 863 protein domains, we assessed the physico-chemical interactions of contacts predicted by CCMpred, MetaPSICOV and DNCON2, as examples of direct coupling analysis, meta-prediction and deep learning. RESULTS: We considered correctly predicted contacts and compared their properties against the protein contacts that were not predicted. Predicted contacts tend to form more bonds than non-predicted contacts, which suggests these contacts may be more important than contacts that were not predicted. Comparing the contacts predicted by each method, we found that metaPSICOV and DNCON2 favour accuracy, whereas CCMPred detects contacts with more bonds. This suggests that the push for higher accuracy may lead to a loss of physico-chemically important contacts. These results underscore the connection between protein physico-chemistry and the coevolutionary couplings that can be derived from MSAs. This relationship is likely to be relevant to protein structure prediction and functional analysis of protein structure and may be key to understanding their utility for different problems in structural biology. AVAILABILITY AND IMPLEMENTATION: We use publicly available databases. Our code is available for download at https://opig.stats.ox.ac.uk/. SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online. Mark Chonofsky, Saulo Henrique Pires de Oliveira, Konrad Krawczyk, Charlotte M. Deane |
Bioinform. | 3 |
| 2018 | In silico structural modeling of multiple epigenetic marks on DNAabstractAbstract There are four known epigenetic cytosine modifications in mammals: methylation (5mC), hydroxymethylation (5hmC), formylation (5fC) and carboxylation (5caC). The biological effects of 5mC are well understood but the roles of the remaining modifications remain elusive. Experimental and computational studies suggest that a single epigenetic mark has little structural effect but six of them can radically change the structure of DNA to a new form, F-DNA. Investigating the collective effect of multiple epigenetic marks requires the ability to interrogate all possible combinations of epigenetic states (e.g. methylated/non-methylated) along a stretch of DNA. Experiments on such complex systems are only feasible on small, isolated examples and there currently exist no systematic computational solutions to this problem. We address this issue by extending the use of Natural Move Monte Carlo to simulate the conformations of epigenetic marks. We validate our protocol by reproducing in silico experimental observations from two recently published high-resolution crystal structures that contain epigenetic marks 5hmC and 5fC. We further demonstrate that our protocol correctly finds either the F-DNA or the B-DNA states more energetically favorable depending on the configuration of the epigenetic marks. We hope that the computational efficiency and ease of use of this novel simulation framework would form the basis for future protocols and facilitate our ability to rapidly interrogate diverse epigenetic systems. Availability and implementation The code together with examples and tutorials are available from http://www.cs.ox.ac.uk/mosaics Supplementary information Supplementary data are available at Bioinformatics online. Konrad Krawczyk, Samuel Demharter, Bernhard Knapp, Charlotte M. Deane, Peter Minary |
Bioinform. | 1 |
| 2016 | Progress and challenges in predicting protein interfacesabstractThe majority of biological processes are mediated via protein-protein interactions. Determination of residues participating in such interactions improves our understanding of molecular mechanisms and facilitates the development of therapeutics. Experimental approaches to identifying interacting residues, such as mutagenesis, are costly and time-consuming and thus, computational methods for this purpose could streamline conventional pipelines. Here we review the field of computational protein interface prediction. We make a distinction between methods which address proteins in general and those targeted at antibodies, owing to the radically different binding mechanism of antibodies. We organize the multitude of currently available methods hierarchically based on required input and prediction principles to provide an overview of the field. Reyhaneh Esmaielbeiki, Konrad Krawczyk, Bernhard Knapp, Jean-Christophe Nebel, Charlotte M. Deane |
Briefings Bioinform. | 2 |
| 2014 | Improving B-cell epitope prediction and its application to global antibody-antigen dockingabstractMOTIVATION: Antibodies are currently the most important class of biopharmaceuticals. Development of such antibody-based drugs depends on costly and time-consuming screening campaigns. Computational techniques such as antibody-antigen docking hold the potential to facilitate the screening process by rapidly providing a list of initial poses that approximate the native complex. RESULTS: We have developed a new method to identify the epitope region on the antigen, given the structures of the antibody and the antigen-EpiPred. The method combines conformational matching of the antibody-antigen structures and a specific antibody-antigen score. We have tested the method on both a large non-redundant set of antibody-antigen complexes and on homology models of the antibodies and/or the unbound antigen structure. On a non-redundant test set, our epitope prediction method achieves 44% recall at 14% precision against 23% recall at 14% precision for a background random distribution. We use our epitope predictions to rescore the global docking results of two rigid-body docking algorithms: ZDOCK and ClusPro. In both cases including our epitope, prediction increases the number of near-native poses found among the top decoys. AVAILABILITY AND IMPLEMENTATION: Our software is available from http://www.stats.ox.ac.uk/research/proteins/resources. Konrad Krawczyk, Terry Baker, Jiye Shi, Charlotte M. Deane |
Bioinform. | 1 |