VLDB 2026 Research / reviewers in the wild / expert
Olga V. Kalinina
dblp:24/1249
· DBLP profile ↗
10ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-9445-477XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ALPAR: automated learning pipeline for antimicrobial resistanceabstractSUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR). Alper Yurtseven, Roman Joeres, Olga V. Kalinina |
Bioinform. | 3 |
| 2026 | Interpretable prediction of DNA replication origins in S. cerevisiae using DNABERT and DNABERT-2abstractBACKGROUND: DNA replication is a biological process in which a single DNA molecule is duplicated, initiating from multiple genomic sites known as replication origins. Identifying replication origins and analyzing their underlying base sequence composition is crucial for understanding the mechanisms of DNA replication. Although there are various machine learning and deep learning approaches for origin prediction, many rely on labor intensive feature engineering or lack interpretability. We fine-tune two genome-based pretrained language models, DNABERT and DNABERT-2, to predict replication origins in budding yeast and unravel the DNA base composition behind them. The key contribution of this study is a systematic framework for analyzing genomic language models for replication origin prediction, combining controlled dataset design with model-specific explainability pipelines to examine how different tokenization strategies influence learned sequence features and whether such approaches can highlight biologically meaningful signals. RESULTS: We evaluate both models on the designed datasets to ensure robustness and support explainability. DNABERT demonstrates consistent performance, achieving an average accuracy of 0.72 for more challenging and 0.83 for the easier dataset. In comparison, DNABERT-2 achieved comparable scores of 0.72 and 0.81 on the same datasets. Our attention-based motif discovery pipeline enhances the interpretability of DNABERT, by identifying motifs from high-attention fragments that closely match known sequence patterns of replication origins. Perturbation-based explanation methods, including Shapley additive explanations, were applied to interpret DNABERT-2's learning mechanism. This analysis identified tokens with high attribution scores aligned with biologically relevant sequence composition. CONCLUSION: Our study demonstrates that both models identify replication origin sequences, albeit through different learning strategies. Tokenization appears to influence model learning and attention behavior in these models. The overlapping k-mer tokenization used in DNABERT yields more interpretable attention maps compared to the byte pair encoding tokenization employed in DNABERT-2. We show that despite sharing the same BERT-style architecture, DNABERT captures relevant short-range patterns and some sequence dependencies beyond just local context, as reflected in its attention maps. In contrast, DNABERT-2's alternative tokenization strategy biases its learning toward relevant short-range patterns by optimizing token weighting. Zohreh Piroozeh, Ildem Akerman, Olga V. Kalinina, Stefan Kesselheim, Alina Bazarova |
BMC Bioinform. | 3 |
| 2023 | MetaProFi: an ultrafast chunked Bloom filter for storing and querying protein and nucleotide sequence data for accurate identification of functionally relevant genetic variantsabstractMOTIVATION: Bloom filters are a popular data structure that allows rapid searches in large sequence datasets. So far, all tools work with nucleotide sequences; however, protein sequences are conserved over longer evolutionary distances, and only mutations on the protein level may have any functional significance. RESULTS: We present MetaProFi, a Bloom filter-based tool that, for the first time, offers the functionality to build indexes of amino acid sequences and query them with both amino acid and nucleotide sequences, thus bringing sequence comparison to the biologically relevant protein level. MetaProFi implements additional efficient engineering solutions, such as a shared memory system, chunked data storage and efficient compression. In addition to its conceptual novelty, MetaProFi demonstrates state-of-the-art performance and excellent memory consumption-to-speed ratio when applied to various large datasets. AVAILABILITY AND IMPLEMENTATION: Source code in Python is available at https://github.com/kalininalab/metaprofi. Sanjay Kumar Srikakulam, Fawaz Dabbaghie, Robert Bals, Olga V. Kalinina |
Bioinform. | 5 |
| 2022 | Phylogenetic inference of changes in amino acid propensities with single-position resolutionabstractFitness conferred by the same allele may differ between genotypes and environments, and these differences shape variation and evolution. Changes in amino acid propensities at protein sites over the course of evolution have been inferred from sequence alignments statistically, but the existing methods are data-intensive and aggregate multiple sites. Here, we develop an approach to detect individual amino acids that confer different fitness in different groups of species from combined sequence and phylogenetic data. Using the fact that the probability of a substitution to an amino acid depends on its fitness, our method looks for amino acids such that substitutions to them occur more frequently in one group of lineages than in another. We validate our method using simulated evolution of a protein site under different scenarios and show that it has high specificity for a wide range of assumptions regarding the underlying changes in selection, while its sensitivity differs between scenarios. We apply our method to the env gene of two HIV-1 subtypes, A and B, and to the HA gene of two influenza A subtypes, H1 and H3, and show that the inferred fitness changes are consistent with the fitness differences observed in deep mutational scanning experiments. We find that changes in relative fitness of different amino acid variants within a site do not always trigger episodes of positive selection and therefore may not result in an overall increase in the frequency of substitutions, but can still be detected from changes in relative frequencies of different substitutions. Galya V. Klink, Olga V. Kalinina, Georgii A. Bazykin |
PLoS Comput. Biol. | 2 |
| 2021 | An extended catalogue of tandem alternative splice sites in human tissue transcriptomesabstractTandem alternative splice sites (TASS) is a special class of alternative splicing events that are characterized by a close tandem arrangement of splice sites. Most TASS lack functional characterization and are believed to arise from splicing noise. Based on the RNA-seq data from the Genotype Tissue Expression project, we present an extended catalogue of TASS in healthy human tissues and analyze their tissue-specific expression. The expression of TASS is usually dominated by one major splice site (maSS), while the expression of minor splice sites (miSS) is at least an order of magnitude lower. Among 46k miSS with sufficient read support, 9k (20%) are significantly expressed above the expected noise level, and among them 2.5k are expressed tissue-specifically. We found significant correlations between tissue-specific expression of RNA-binding proteins (RBP), tissue-specific expression of miSS, and miSS response to RBP inactivation by shRNA. In combination with RBP profiling by eCLIP, this allowed prediction of novel cases of tissue-specific splicing regulation including a miSS in QKI mRNA that is likely regulated by PTBP1. The analysis of human primary cell transcriptomes suggested that both tissue-specific and cell-type-specific factors contribute to the regulation of miSS expression. More than 20% of tissue-specific miSS affect structured protein regions and may adjust protein-protein interactions or modify the stability of the protein core. The significantly expressed miSS evolve under the same selection pressure as maSS, while other miSS lack signatures of evolutionary selection and conservation. Using mixture models, we estimated that not more than 15% of maSS and not more than 54% of tissue-specific miSS are noisy, while the proportion of noisy splice sites among non-significantly expressed miSS is above 63%. Andrey A. Mironov, Stepan Denisov, Alexander Greß, Olga V. Kalinina, Dmitri D. Pervouchine |
PLoS Comput. Biol. | 4 |
| 2020 | SphereCon - a method for precise estimation of residue relative solvent accessible area from limited structural informationabstractMOTIVATION: In proteins, solvent accessibility of individual residues is a factor contributing to their importance for protein function and stability. Hence one might wish to calculate solvent accessibility in order to predict the impact of mutations, their pathogenicity and for other biomedical applications. A direct computation of solvent accessibility is only possible if all atoms of a protein three-dimensional structure are reliably resolved. RESULTS: We present SphereCon, a new precise measure that can estimate residue relative solvent accessibility (RSA) from limited data. The measure is based on calculating the volume of intersection of a sphere with a cone cut out in the direction opposite of the residue with surrounding atoms. We propose a method for estimating the position and volume of residue atoms in cases when they are not known from the structure, or when the structural data are unreliable or missing. We show that in cases of reliable input structures, SphereCon correlates almost perfectly with the directly computed RSA, and outperforms other previously suggested indirect methods. Moreover, SphereCon is the only measure that yields accurate results when the identities of amino acids are unknown. A significant novel feature of SphereCon is that it can estimate RSA from inter-residue distance and contact matrices, without any information about the actual atom coordinates. AVAILABILITY AND IMPLEMENTATION: https://github.com/kalininalab/spherecon. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander Greß, Olga V. Kalinina |
Bioinform. | 2 |
| 2016 | BALL-SNPgp - from genetic variants toward computational diagnosticsabstractUNLABELLED: In medical research, it is crucial to understand the functional consequences of genetic alterations, for example, non-synonymous single nucleotide variants (nsSNVs). NsSNVs are known to be causative for several human diseases. However, the genetic basis of complex disorders such as diabetes or cancer comprises multiple factors. Methods to analyze putative synergetic effects of multiple such factors, however, are limited. Here, we concentrate on nsSNVs and present BALL-SNPgp, a tool for structural and functional characterization of nsSNVs, which is aimed to improve pathogenicity assessment in computational diagnostics. Based on annotated SNV data, BALL-SNPgp creates a three-dimensional visualization of the encoded protein, collects available information from different resources concerning disease relevance and other functional annotations, performs cluster analysis, predicts putative binding pockets and provides data on known interaction sites. AVAILABILITY AND IMPLEMENTATION: BALL-SNPgp is based on the comprehensive C ++ framework Biochemical Algorithms Library (BALL) and its visualization front-end BALLView. Our tool is available at www.ccb.uni-saarland.de/BALL-SNPgp CONTACT: [email protected]. Sabine C. Mueller, Christina Backes, Alexander Greß, Nina Baumgarten, Olga V. Kalinina, Andreas Moll, Oliver Kohlbacher, Eckart Meese, Andreas Keller |
Bioinform. | 5 |
| 2016 | Patterns of amino acid conservation in human and animal immunodeficiency virusesabstractMOTIVATION: Due to their high genomic variability, RNA viruses and retroviruses present a unique opportunity for detailed study of molecular evolution. Lentiviruses, with HIV being a notable example, are one of the best studied viral groups: hundreds of thousands of sequences are available together with experimentally resolved three-dimensional structures for most viral proteins. In this work, we use these data to study specific patterns of evolution of the viral proteins, and their relationship to protein interactions and immunogenicity. RESULTS: We propose a method for identification of two types of surface residues clusters with abnormal conservation: extremely conserved and extremely variable clusters. We identify them on the surface of proteins from HIV and other animal immunodeficiency viruses. Both types of clusters are overrepresented on the interaction interfaces of viral proteins with other proteins, nucleic acids or low molecular-weight ligands, both in the viral particle and between the virus and its host. In the immunodeficiency viruses, the interaction interfaces are not more conserved than the corresponding proteins on an average, and we show that extremely conserved clusters coincide with protein-protein interaction hotspots, predicted as the residues with the largest energetic contribution to the interaction. Extremely variable clusters have been identified here for the first time. In the HIV-1 envelope protein gp120, they overlap with known antigenic sites. These antigenic sites also contain many residues from extremely conserved clusters, hence representing a unique interacting interface enriched both in extremely conserved and in extremely variable clusters of residues. This observation may have important implication for antiretroviral vaccine development. AVAILABILITY AND IMPLEMENTATION: A Python package is available at https://bioinf.mpi-inf.mpg.de/publications/viral-ppi-pred/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Olga S. Voitenko, Andi Dhroso, Anna Hake, Dmitry Korkin, Olga V. Kalinina |
Bioinform. | 5 |
| 2011 | Combinations of Protein-Chemical Complex Structures Reveal New Targets for Established DrugsabstractBiological networks are powerful tools for predicting undocumented relationships between molecules. The underlying principle is that existing interactions between molecules can be used to predict new interactions. Here we use this principle to suggest new protein-chemical interactions via the network derived from three-dimensional structures. For pairs of proteins sharing a common ligand, we use protein and chemical superimpositions combined with fast structural compatibility screens to predict whether additional compounds bound by one protein would bind the other. The method reproduces 84% of complexes in a benchmark, and we make many predictions that would not be possible using conventional modeling techniques. Within 19,578 novel predicted interactions are 7,793 involving 718 drugs, including filaminast, coumarin, alitretonin and erlotinib. The growth rate of confident predictions is twice that of experimental complexes, meaning that a complete structural drug-protein repertoire will be available at least ten years earlier than by X-ray and NMR techniques alone. Olga V. Kalinina, Oliver Wichmann, Gordana Apic, Robert B. Russell |
PLoS Comput. Biol. | 1 |
| 2009 | Combining specificity determining and conserved residues improves functional site predictionabstractBACKGROUND: Predicting the location of functionally important sites from protein sequence and/or structure is a long-standing problem in computational biology. Most current approaches make use of sequence conservation, assuming that amino acid residues conserved within a protein family are most likely to be functionally important. Most often these approaches do not consider many residues that act to define specific sub-functions within a family, or they make no distinction between residues important for function and those more relevant for maintaining structure (e.g. in the hydrophobic core). Many protein families bind and/or act on a variety of ligands, meaning that conserved residues often only bind a common ligand sub-structure or perform general catalytic activities. RESULTS: Here we present a novel method for functional site prediction based on identification of conserved positions, as well as those responsible for determining ligand specificity. We define Specificity-Determining Positions (SDPs), as those occupied by conserved residues within sub-groups of proteins in a family having a common specificity, but differ between groups, and are thus likely to account for specific recognition events. We benchmark the approach on enzyme families of known 3D structure with bound substrates, and find that in nearly all families residues predicted by SDPsite are in contact with the bound substrate, and that the addition of SDPs significantly improves functional site prediction accuracy. We apply SDPsite to various families of proteins containing known three-dimensional structures, but lacking clear functional annotations, and discusse several illustrative examples. CONCLUSION: The results suggest a better means to predict functional details for the thousands of protein structures determined prior to a clear understanding of molecular function. Olga V. Kalinina, Mikhail S. Gelfand, Robert B. Russell |
BMC Bioinform. | 1 |