EDBT 2026 Demo / reviewers in the wild / expert
Anne-Florence Bitbol
dblp:155/3475
· DBLP profile ↗
13ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-1020-494XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Systematic analysis of CDR contacts and pairing constraints between T cell receptor αβ chainsabstractMOTIVATION: The six complementarity determining regions (CDRs) of the T cell receptor (TCR) form multiple contacts with cognate peptide and major histocompatibility complex, thus determining antigen specificity. However, the contacts between the CDRs themselves are less understood. RESULTS: Our systematic study of all available TCR crystallographic structures identified consistent patterns of intra- and inter-chain CDR contacts in both free and antigen-bound TCRs. In addition, the protein sequences of TCRα and TCRβ from sets of TCRs which recognise a shared antigen shared mutual information and were not independent. As a result, sequence-based models can partially predict TCRα/TCRβ pairing de novo. The conserved patterns of CDR amino acid contacts, and the mutual sequence constraints between antigen-specific sets of TCR α and β chains represent an under-appreciated element of TCR structure, which may play an important role in T cell antigen recognition. AVAILABILITY AND IMPLEMENTATION: The code and data necessary to reproduce the analyses are available at https://github.com/mm523/TCRab-pairing. Martina Milighetti, Yuta Nagano, Uri Hershberg, Andreas Tiffeau-Mayer, Anne-Florence Bitbol, Benjamin Chain |
Bioinform. | 6 |
| 2025 | DiffPaSS - high-performance differentiable pairing of protein sequences using soft scoresabstractMOTIVATION: Identifying interacting partners from two sets of protein sequences has important applications in computational biology. Interacting partners share similarities across species due to their common evolutionary history, and feature correlations in amino acid usage due to the need to maintain complementary interaction interfaces. Thus, the problem of finding interacting pairs can be formulated as searching for a pairing of sequences that maximizes a sequence similarity or a coevolution score. Several methods have been developed to address this problem, applying different approximate optimization methods to different scores. RESULTS: We introduce Differentiable Pairing using Soft Scores (DiffPaSS), a differentiable framework for flexible, fast, and hyperparameter-free optimization for pairing interacting biological sequences, which can be applied to a wide variety of scores. We apply it to a benchmark prokaryotic dataset, using mutual information and neighbor graph alignment scores. DiffPaSS outperforms existing algorithms for optimizing the same scores. We demonstrate the usefulness of our paired alignments for the prediction of protein complex structure. DiffPaSS does not require sequences to be aligned, and we also apply it to nonaligned sequences from T-cell receptors. AVAILABILITY AND IMPLEMENTATION: A PyTorch implementation and installable Python package are available at https://github.com/Bitbol-Lab/DiffPaSS. Umberto Lupo, Damiano Sgarbossa, Martina Milighetti, Anne-Florence Bitbol |
Bioinform. | 4 |
| 2025 | ProtMamba: a homology-aware but alignment-free protein state space modelabstractMOTIVATION: Protein language models are enabling advances in elucidating the sequence-to-function mapping, and have important applications in protein design. Models based on multiple sequence alignments efficiently capture the evolutionary information in homologous protein sequences, but multiple sequence alignment construction is imperfect. RESULTS: We present ProtMamba, a homology-aware but alignment-free protein language model based on the Mamba architecture. In contrast with attention-based models, ProtMamba efficiently handles very long context, comprising hundreds of protein sequences. It is also computationally efficient. We train ProtMamba on a large dataset of concatenated homologous sequences, using two GPUs. We combine autoregressive modeling and masked language modeling through a fill-in-the-middle training objective. This makes the model adapted to various protein design applications. We demonstrate ProtMamba's usefulness for sequence generation, motif inpainting, fitness prediction, and modeling intrinsically disordered regions. For homolog-conditioned sequence generation, ProtMamba outperforms state-of-the-art models. ProtMamba's competitive performance, despite its relatively small size, sheds light on the importance of long-context conditioning. AVAILABILITY AND IMPLEMENTATION: A Python implementation of ProtMamba is freely available in our GitHub repository: https://github.com/Bitbol-Lab/ProtMamba-ssm and archived at https://doi.org/10.5281/zenodo.15584634. Damiano Sgarbossa, Cyril Malbranke, Anne-Florence Bitbol |
Bioinform. | 3 |
| 2025 | Spatial structure facilitates evolutionary rescue by drug resistanceabstractBacterial populations often have complex spatial structures, which can impact their evolution. Here, we study how spatial structure affects the evolution of antibiotic resistance in a bacterial population. We consider a minimal model of spatially structured populations where all demes (i.e., subpopulations) are identical and connected to each other by identical migration rates. We show that spatial structure can facilitate the survival of a bacterial population to antibiotic treatment, starting from a sensitive inoculum. Specifically, the bacterial population can be rescued if antibiotic resistant mutants appear and are present when drug is added, and spatial structure can impact the fate of these mutants and the probability that they are present. Indeed, the probability of fixation of neutral or deleterious mutations providing drug resistance is increased in smaller populations. This promotes local fixation of resistant mutants in the structured population, which facilitates evolutionary rescue by drug resistance in the rare mutation regime. Once the population is rescued by resistance, migrations allow resistant mutants to spread in all demes. Our main result that spatial structure facilitates evolutionary rescue by antibiotic resistance extends to more complex spatial structures, and to the case where there are resistant mutants in the inoculum. Cecilia Fruet, Ella Linxia Müller, Claude Loverdo, Anne-Florence Bitbol |
PLoS Comput. Biol. | 4 |
| 2024 | Mutant fate in spatially structured populations on graphs: Connecting models to experimentsabstractIn nature, most microbial populations have complex spatial structures that can affect their evolution. Evolutionary graph theory predicts that some spatial structures modelled by placing individuals on the nodes of a graph affect the probability that a mutant will fix. Evolution experiments are beginning to explicitly address the impact of graph structures on mutant fixation. However, the assumptions of evolutionary graph theory differ from the conditions of modern evolution experiments, making the comparison between theory and experiment challenging. Here, we aim to bridge this gap by using our new model of spatially structured populations. This model considers connected subpopulations that lie on the nodes of a graph, and allows asymmetric migrations. It can handle large populations, and explicitly models serial passage events with migrations, thus closely mimicking experimental conditions. We analyze recent experiments in light of this model. We suggest useful parameter regimes for future experiments, and we make quantitative predictions for these experiments. In particular, we propose experiments to directly test our recent prediction that the star graph with asymmetric migrations suppresses natural selection and can accelerate mutant fixation or extinction, compared to a well-mixed population. Alia Abbara, Lisa Pagani, Celia García-Pareja, Anne-Florence Bitbol |
PLoS Comput. Biol. | 4 |
| 2024 | Impact of phylogeny on the inference of functional sectors from protein sequence dataabstractStatistical analysis of multiple sequence alignments of homologous proteins has revealed groups of coevolving amino acids called sectors. These groups of amino-acid sites feature collective correlations in their amino-acid usage, and they are associated to functional properties. Modeling showed that nonlinear selection on an additive functional trait of a protein is generically expected to give rise to a functional sector. These modeling results motivated a principled method, called ICOD, which is designed to identify functional sectors, as well as mutational effects, from sequence data. However, a challenge for all methods aiming to identify sectors from multiple sequence alignments is that correlations in amino-acid usage can also arise from the mere fact that homologous sequences share common ancestry, i.e. from phylogeny. Here, we generate controlled synthetic data from a minimal model comprising both phylogeny and functional sectors. We use this data to dissect the impact of phylogeny on sector identification and on mutational effect inference by different methods. We find that ICOD is most robust to phylogeny, but that conservation is also quite robust. Next, we consider natural multiple sequence alignments of protein families for which deep mutational scan experimental data is available. We show that in this natural data, conservation and ICOD best identify sites with strong functional roles, in agreement with our results on synthetic data. Importantly, these two methods have different premises, since they respectively focus on conservation and on correlations. Thus, their joint use can reveal complementary information. Nicola Dietler, Alia Abbara, Subham Choudhury, Anne-Florence Bitbol |
PLoS Comput. Biol. | 4 |
| 2023 | Combining phylogeny and coevolution improves the inference of interaction partners among paralogous proteinsabstractPredicting protein-protein interactions from sequences is an important goal of computational biology. Various sources of information can be used to this end. Starting from the sequences of two interacting protein families, one can use phylogeny or residue coevolution to infer which paralogs are specific interaction partners within each species. We show that these two signals can be combined to improve the performance of the inference of interaction partners among paralogs. For this, we first align the sequence-similarity graphs of the two families through simulated annealing, yielding a robust partial pairing. We next use this partial pairing to seed a coevolution-based iterative pairing algorithm. This combined method improves performance over either separate method. The improvement obtained is striking in the difficult cases where the average number of paralogs per species is large or where the total number of sequences is modest. Carlos A. Gandarilla-Pérez, Sergio Pinilla, Anne-Florence Bitbol, Martin Weigt |
PLoS Comput. Biol. | 3 |
| 2022 | Correlations from structure and phylogeny combine constructively in the inference of protein partners from sequencesabstractInferring protein-protein interactions from sequences is an important task in computational biology. Recent methods based on Direct Coupling Analysis (DCA) or Mutual Information (MI) allow to find interaction partners among paralogs of two protein families. Does successful inference mainly rely on correlations from structural contacts or from phylogeny, or both? Do these two types of signal combine constructively or hinder each other? To address these questions, we generate and analyze synthetic data produced using a minimal model that allows us to control the amounts of structural constraints and phylogeny. We show that correlations from these two sources combine constructively to increase the performance of partner inference by DCA or MI. Furthermore, signal from phylogeny can rescue partner inference when signal from contacts becomes less informative, including in the realistic case where inter-protein contacts are restricted to a small subset of sites. We also demonstrate that DCA-inferred couplings between non-contact pairs of sites improve partner inference in the presence of strong phylogeny, while deteriorating it otherwise. Moreover, restricting to non-contact pairs of sites preserves inference performance in the presence of strong phylogeny. In a natural data set, as well as in realistic synthetic data based on it, we find that non-contact pairs of sites contribute positively to partner inference performance, and that restricting to them preserves performance, evidencing an important role of phylogeny. Andonis Gerardos, Nicola Dietler, Anne-Florence Bitbol |
PLoS Comput. Biol. | 3 |
| 2020 | Resist or perish: Fate of a microbial population subjected to a periodic presence of antimicrobialabstractThe evolution of antimicrobial resistance can be strongly affected by variations of antimicrobial concentration. Here, we study the impact of periodic alternations of absence and presence of antimicrobial on resistance evolution in a microbial population, using a stochastic model that includes variations of both population composition and size, and fully incorporates stochastic population extinctions. We show that fast alternations of presence and absence of antimicrobial are inefficient to eradicate the microbial population and strongly favor the establishment of resistance, unless the antimicrobial increases enough the death rate. We further demonstrate that if the period of alternations is longer than a threshold value, the microbial population goes extinct upon the first addition of antimicrobial, if it is not rescued by resistance. We express the probability that the population is eradicated upon the first addition of antimicrobial, assuming rare mutations. Rescue by resistance can happen either if resistant mutants preexist, or if they appear after antimicrobial is added to the environment. Importantly, the latter case is fully prevented by perfect biostatic antimicrobials that completely stop division of sensitive microorganisms. By contrast, we show that the parameter regime where treatment is efficient is larger for biocidal drugs than for biostatic drugs. This sheds light on the respective merits of different antimicrobial modes of action. Loïc Marrec, Anne-Florence Bitbol |
PLoS Comput. Biol. | 2 |
| 2019 | Phylogenetic correlations can suffice to infer protein partners from sequencesabstractDetermining which proteins interact together is crucial to a systems-level understanding of the cell. Recently, algorithms based on Direct Coupling Analysis (DCA) pairwise maximum-entropy models have allowed to identify interaction partners among paralogous proteins from sequence data. This success of DCA at predicting protein-protein interactions could be mainly based on its known ability to identify pairs of residues that are in contact in the three-dimensional structure of protein complexes and that coevolve to remain physicochemically complementary. However, interacting proteins possess similar evolutionary histories. What is the role of purely phylogenetic correlations in the performance of DCA-based methods to infer interaction partners? To address this question, we employ controlled synthetic data that only involve phylogeny and no interactions or contacts. We find that DCA accurately identifies the pairs of synthetic sequences that share evolutionary history. While phylogenetic correlations confound the identification of contacting residues by DCA, they are thus useful to predict interacting partners among paralogs. We find that DCA performs as well as phylogenetic methods to this end, and slightly better than them with large and accurate training sets. Employing DCA or phylogenetic methods within an Iterative Pairing Algorithm (IPA) allows to predict pairs of evolutionary partners without a training set. We further demonstrate the ability of these various methods to correctly predict pairings among real paralogous proteins with genome proximity but no known direct physical interaction, illustrating the importance of phylogenetic correlations in natural data. However, for physically interacting and strongly coevolving proteins, DCA and mutual information outperform phylogenetic methods. We finally discuss how to distinguish physically interacting proteins from proteins that only share a common evolutionary history. Guillaume Marmier, Martin Weigt, Anne-Florence Bitbol |
PLoS Comput. Biol. | 3 |
| 2019 | Revealing evolutionary constraints on proteins through sequence analysisabstractStatistical analysis of alignments of large numbers of protein sequences has revealed "sectors" of collectively coevolving amino acids in several protein families. Here, we show that selection acting on any functional property of a protein, represented by an additive trait, can give rise to such a sector. As an illustration of a selected trait, we consider the elastic energy of an important conformational change within an elastic network model, and we show that selection acting on this energy leads to correlations among residues. For this concrete example and more generally, we demonstrate that the main signature of functional sectors lies in the small-eigenvalue modes of the covariance matrix of the selected sequences. However, secondary signatures of these functional sectors also exist in the extensively-studied large-eigenvalue modes. Our simple, general model leads us to propose a principled method to identify functional sectors, along with the magnitudes of mutational effects, from sequence data. We further demonstrate the robustness of these functional sectors to various forms of selection, and the robustness of our approach to the identification of multiple selected traits. Shouwen Wang, Anne-Florence Bitbol, Ned S. Wingreen |
PLoS Comput. Biol. | 2 |
| 2018 | Inferring interaction partners from protein sequences using mutual informationabstractFunctional protein-protein interactions are crucial in most cellular processes. They enable multi-protein complexes to assemble and to remain stable, and they allow signal transduction in various pathways. Functional interactions between proteins result in coevolution between the interacting partners, and thus in correlations between their sequences. Pairwise maximum-entropy based models have enabled successful inference of pairs of amino-acid residues that are in contact in the three-dimensional structure of multi-protein complexes, starting from the correlations in the sequence data of known interaction partners. Recently, algorithms inspired by these methods have been developed to identify which proteins are functional interaction partners among the paralogous proteins of two families, starting from sequence data alone. Here, we demonstrate that a slightly higher performance for partner identification can be reached by an approximate maximization of the mutual information between the sequence alignments of the two protein families. Our mutual information-based method also provides signatures of the existence of interactions between protein families. These results stand in contrast with structure prediction of proteins and of multi-protein complexes from sequence data, where pairwise maximum-entropy based global statistical models substantially improve performance compared to mutual information. Our findings entail that the statistical dependences allowing interaction partner prediction from sequence data are not restricted to the residue pairs that are in direct contact at the interface between the partner proteins. Anne-Florence Bitbol |
PLoS Comput. Biol. | 1 |
| 2014 | Quantifying the Role of Population Subdivision in Evolution on Rugged Fitness LandscapesabstractNatural selection drives populations towards higher fitness, but crossing fitness valleys or plateaus may facilitate progress up a rugged fitness landscape involving epistasis. We investigate quantitatively the effect of subdividing an asexual population on the time it takes to cross a fitness valley or plateau. We focus on a generic and minimal model that includes only population subdivision into equivalent demes connected by global migration, and does not require significant size changes of the demes, environmental heterogeneity or specific geographic structure. We determine the optimal speedup of valley or plateau crossing that can be gained by subdivision, if the process is driven by the deme that crosses fastest. We show that isolated demes have to be in the sequential fixation regime for subdivision to significantly accelerate crossing. Using Markov chain theory, we obtain analytical expressions for the conditions under which optimal speedup is achieved: valley or plateau crossing by the subdivided population is then as fast as that of its fastest deme. We verify our analytical predictions through stochastic simulations. We demonstrate that subdivision can substantially accelerate the crossing of fitness valleys and plateaus in a wide range of parameters extending beyond the optimal window. We study the effect of varying the degree of subdivision of a population, and investigate the trade-off between the magnitude of the optimal speedup and the width of the parameter range over which it occurs. Our results, obtained for fitness valleys and plateaus, also hold for weakly beneficial intermediate mutations. Finally, we extend our work to the case of a population connected by migration to one or several smaller islands. Our results demonstrate that subdivision with migration alone can significantly accelerate the crossing of fitness valleys and plateaus, and shed light onto the quantitative conditions necessary for this to occur. Anne-Florence Bitbol, David J. Schwab |
PLoS Comput. Biol. | 1 |