Aleksandra M. Walczak

dblp:28/11026 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-2686-5702ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 5 since 2021
YearPublicationVenuePosition
2026 Paraplume: A fast and accurate antibody paratope prediction method provides insights into repertoire-scale binding dynamics
abstract
The specific region of an antibody responsible for binding to an antigen, known as the paratope, is essential for immune recognition. Accurate identification of this small yet critical region can accelerate the development of therapeutic antibodies. Determining paratope locations typically relies on modeling the antibody structure, which is computationally intensive and difficult to scale across large antibody repertoires. We introduce Paraplume, a sequence-based paratope prediction method that leverages embeddings from protein language models (PLMs), without the need for structural input and achieves superior performance across multiple benchmarks compared to current methods. In addition, reweighting PLM embeddings using Paraplume predictions yields more informative sequence representations, improving downstream tasks such as binder classification and epitope binning. Applied to large antibody repertoires, Paraplume reveals that antigen-specific somatic hypermutations are associated with larger paratopes, suggesting a potential mechanism for affinity enhancement. Our findings position PLM-based paratope prediction as a powerful, scalable alternative to structure-dependent approaches, opening new avenues for understanding antibody evolution.
Gabriel Athènes, Adam Woolfe, Thierry Mora, Aleksandra M. Walczak
PLoS Comput. Biol.4
2025 Learning predictive signatures of HLA type from T-cell repertoires
abstract
T cells recognize a wide range of pathogens using surface receptors that interact directly with peptides presented on major histocompatibility complexes (MHC) encoded by the HLA loci in humans. Understanding the association between T cell receptors (TCR) and HLA alleles is an important step towards predicting TCR-antigen specificity from sequences. Here we analyze the TCR alpha and beta repertoires of large cohorts of HLA-typed donors to systematically infer such associations, by looking for overrepresentation of TCRs in individuals with a common allele.TCRs, associated with a specific HLA allele, exhibit sequence similarities that suggest prior antigen exposure. Immune repertoire sequencing has produced large numbers of datasets, however the HLA type of the corresponding donors is rarely available. Using our TCR-HLA associations, we trained a computational model to predict the HLA type of individuals from their TCR repertoire alone. We propose an iterative procedure to refine this model by using data from large cohorts of untyped individuals, by recursively typing them using the model itself. The resulting model shows good predictive performance, even for relatively rare HLA alleles.
María Ruiz Ortega, Mikhail Pogorelyy, Anastasia A. Minervina, Paul G. Thomas, Thierry Mora, Aleksandra M. Walczak
PLoS Comput. Biol.6
2022 Learning the statistics and landscape of somatic mutation-induced insertions and deletions in antibodies
abstract
Affinity maturation is crucial for improving the binding affinity of antibodies to antigens. This process is mainly driven by point substitutions caused by somatic hypermutations of the immunoglobulin gene. It also includes deletions and insertions of genomic material known as indels. While the landscape of point substitutions has been extensively studied, a detailed statistical description of indels is still lacking. Here we present a probabilistic inference tool to learn the statistics of indels from repertoire sequencing data, which overcomes the pitfalls and biases of standard annotation methods. The model includes antibody-specific maturation ages to account for variable mutational loads in the repertoire. After validation on synthetic data, we applied our tool to a large dataset of human immunoglobulin heavy chains. The inferred model allows us to identify universal statistical features of indels in heavy chains. We report distinct insertion and deletion hotspots, and show that the distribution of lengths of indels follows a geometric distribution, which puts constraints on future mechanistic models of the hypermutation process.
Cosimo Lupo, Natanael Spisak, Aleksandra M. Walczak, Thierry Mora
PLoS Comput. Biol.3
2021 Probing T-cell response by sequence-based probabilistic modeling
abstract
With the increasing ability to use high-throughput next-generation sequencing to quantify the diversity of the human T cell receptor (TCR) repertoire, the ability to use TCR sequences to infer antigen-specificity could greatly aid potential diagnostics and therapeutics. Here, we use a machine-learning approach known as Restricted Boltzmann Machine to develop a sequence-based inference approach to identify antigen-specific TCRs. Our approach combines probabilistic models of TCR sequences with clone abundance information to extract TCR sequence motifs central to an antigen-specific response. We use this model to identify patient personalized TCR motifs that respond to individual tumor and infectious disease antigens, and to accurately discriminate specific from non-specific responses. Furthermore, the hidden structure of the model results in an interpretable representation space where TCRs responding to the same antigen cluster, correctly discriminating the response of TCR to different viral epitopes. The model can be used to identify condition specific responding TCRs. We focus on the examples of TCRs reactive to candidate neoantigens and selected epitopes in experiments of stimulated TCR clone expansion.
Barbara Bravi, Vinod P. Balachandran, Benjamin D. Greenbaum, Aleksandra M. Walczak, Thierry Mora, Rémi Monasson, Simona Cocco
PLoS Comput. Biol.4
2021 Optimal prediction with resource constraints using the information bottleneck
abstract
Responding to stimuli requires that organisms encode information about the external world. Not all parts of the input are important for behavior, and resource limitations demand that signals be compressed. Prediction of the future input is widely beneficial in many biological systems. We compute the trade-offs between representing the past faithfully and predicting the future using the information bottleneck approach, for input dynamics with different levels of complexity. For motion prediction, we show that, depending on the parameters in the input dynamics, velocity or position information is more useful for accurate prediction. We show which motion representations are easiest to re-use for accurate prediction in other motion contexts, and identify and quantify those with the highest transferability. For non-Markovian dynamics, we explore the role of long-term memory in shaping the internal representation. Lastly, we show that prediction in evolutionary population dynamics is linked to clustering allele frequencies into non-overlapping memories.
Vedant Sachdeva, Thierry Mora, Aleksandra M. Walczak, Stephanie E. Palmer
PLoS Comput. Biol.3
2020 SOS: online probability estimation and generation of T-and B-cell receptors
abstract
SUMMARY: Recent advances in modelling VDJ recombination and subsequent selection of T- and B-cell receptors provide useful tools to analyse and compare immune repertoires across time, individuals and tissues. A suite of tools-IGoR, OLGA and SONIA-have been publicly released to the community that allow for the inference of generative and selection models from high-throughput sequencing data. However, using these tools requires some scripting or command-line skills and familiarity with complex datasets. As a result, the application of the above models has not been available to a broad audience. In this application note, we fill this gap by presenting Simple OLGA & SONIA (SOS), a web-based interface where users with no coding skills can compute the generation and post-selection probabilities of their sequences, as well as generate batches of synthetic sequences. The application also functions on mobile phones. AVAILABILITY AND IMPLEMENTATION: SOS is freely available to use at sites.google.com/view/statbiophysens/sos with source code at github.com/statbiophys/sos.
Giulio Isacchini, Carlos Olivares, Armita Nourmohammad, Aleksandra M. Walczak, Thierry Mora
Bioinform.4
2020 Population variability in the generation and selection of T-cell repertoires
abstract
The diversity of T-cell receptor (TCR) repertoires is achieved by a combination of two intrinsically stochastic steps: random receptor generation by VDJ recombination, and selection based on the recognition of random self-peptides presented on the major histocompatibility complex. These processes lead to a large receptor variability within and between individuals. However, the characterization of the variability is hampered by the limited size of the sampled repertoires. We introduce a new software tool SONIA to facilitate inference of individual-specific computational models for the generation and selection of the TCR beta chain (TRB) from sequenced repertoires of 651 individuals, separating and quantifying the variability of the two processes of generation and selection in the population. We find not only that most of the variability is driven by the VDJ generation process, but there is a large degree of consistency between individuals with the inter-individual variance of repertoires being about ∼2% of the intra-individual variance. Known viral-specific TCRs follow the same generation and selection statistics as all TCRs.
Zachary Sethna, Giulio Isacchini, Thomas Dupic, Thierry Mora, Aleksandra M. Walczak, Yuval Elhanati
PLoS Comput. Biol.5
2020 Inferring the immune response from repertoire sequencing
abstract
High-throughput sequencing of B- and T-cell receptors makes it possible to track immune repertoires across time, in different tissues, and in acute and chronic diseases or in healthy individuals. However, quantitative comparison between repertoires is confounded by variability in the read count of each receptor clonotype due to sampling, library preparation, and expression noise. Here, we present a general Bayesian approach to disentangle repertoire variations from these stochastic effects. Using replicate experiments, we first show how to learn the natural variability of read counts by inferring the distributions of clone sizes as well as an explicit noise model relating true frequencies of clones to their read count. We then use that null model as a baseline to infer a model of clonal expansion from two repertoire time points taken before and after an immune challenge. Applying our approach to yellow fever vaccination as a model of acute infection in humans, we identify candidate clones participating in the response.
Maximilian Puelma Touzel, Aleksandra M. Walczak, Thierry Mora
PLoS Comput. Biol.2
2019 OLGA: fast computation of generation probabilities of B- and T-cell receptor amino acid sequences and motifs
abstract
MOTIVATION: High-throughput sequencing of large immune repertoires has enabled the development of methods to predict the probability of generation by V(D)J recombination of T- and B-cell receptors of any specific nucleotide sequence. These generation probabilities are very non-homogeneous, ranging over 20 orders of magnitude in real repertoires. Since the function of a receptor really depends on its protein sequence, it is important to be able to predict this probability of generation at the amino acid level. However, brute-force summation over all the nucleotide sequences with the correct amino acid translation is computationally intractable. The purpose of this paper is to present a solution to this problem. RESULTS: We use dynamic programming to construct an efficient and flexible algorithm, called OLGA (Optimized Likelihood estimate of immunoGlobulin Amino-acid sequences), for calculating the probability of generating a given CDR3 amino acid sequence or motif, with or without V/J restriction, as a result of V(D)J recombination in B or T cells. We apply it to databases of epitope-specific T-cell receptors to evaluate the probability that a typical human subject will possess T cells responsive to specific disease-associated epitopes. The model prediction shows an excellent agreement with published data. We suggest that OLGA may be a useful tool to guide vaccine design. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/zsethna/OLGA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zachary Sethna, Yuval Elhanati, Curtis G. Callan Jr., Aleksandra M. Walczak, Thierry Mora
Bioinform.4
2019 Genesis of the αβ T-cell receptor
abstract
The T-cell (TCR) repertoire relies on the diversity of receptors composed of two chains, called α and β, to recognize pathogens. Using results of high throughput sequencing and computational chain-pairing experiments of human TCR repertoires, we quantitively characterize the αβ generation process. We estimate the probabilities of a rescue recombination of the β chain on the second chromosome upon failure or success on the first chromosome. Unlike β chains, α chains recombine simultaneously on both chromosomes, resulting in correlated statistics of the two genes which we predict using a mechanistic model. We find that ∼35% of cells express both α chains. Altogether, our statistical analysis gives a complete quantitative mechanistic picture that results in the observed correlations in the generative process. We learn that the probability to generate any TCRαβ is lower than 10(-12) and estimate the generation diversity and sharing properties of the αβ TCR repertoire.
Thomas Dupic, Quentin Marcou, Aleksandra M. Walczak, Thierry Mora
PLoS Comput. Biol.3
2019 Size and structure of the sequence space of repeat proteins
abstract
The coding space of protein sequences is shaped by evolutionary constraints set by requirements of function and stability. We show that the coding space of a given protein family-the total number of sequences in that family-can be estimated using models of maximum entropy trained on multiple sequence alignments of naturally occuring amino acid sequences. We analyzed and calculated the size of three abundant repeat proteins families, whose members are large proteins made of many repetitions of conserved portions of ∼30 amino acids. While amino acid conservation at each position of the alignment explains most of the reduction of diversity relative to completely random sequences, we found that correlations between amino acid usage at different positions significantly impact that diversity. We quantified the impact of different types of correlations, functional and evolutionary, on sequence diversity. Analysis of the detailed structure of the coding space of the families revealed a rugged landscape, with many local energy minima of varying sizes with a hierarchical structure, reminiscent of fustrated energy landscapes of spin glass in physics. This clustered structure indicates a multiplicity of subtypes within each family, and suggests new strategies for protein design.
Jacopo Marchi, Ezequiel A. Galpern, Rocío Espada, Diego U. Ferreiro, Aleksandra M. Walczak, Thierry Mora
PLoS Comput. Biol.5
2018 Active degradation of MarA controls coordination of its downstream targets
abstract
Several key transcription factors have unusually short half-lives compared to other cellular proteins. Here, we explore the utility of active degradation in shaping how the multiple antibiotic resistance activator MarA coordinates its downstream targets. MarA controls a variety of stress response genes in Escherichia coli. We modify its half-life either by knocking down the protease that targets it via CRISPRi or by engineering MarA to protect it from degradation. Our experimental and analytical results indicate that active degradation can impact both the rate of coordination and the maximum coordination that downstream genes can achieve. In the context of multi-gene regulation, trade-offs between these properties show that perfect information fidelity and instantaneous coordination cannot coexist.
Nicholas A. Rossi, Thierry Mora, Aleksandra M. Walczak, Mary J. Dunlop
PLoS Comput. Biol.3
2018 Precision in a rush: Trade-offs between reproducibility and steepness of the hunchback expression pattern
abstract
Fly development amazes us by the precision and reproducibility of gene expression, especially since the initial expression patterns are established during very short nuclear cycles. Recent live imaging of hunchback promoter dynamics shows a stable steep binary expression pattern established within the three minute interphase of nuclear cycle 11. Considering expression models of different complexity, we explore the trade-off between the ability of a regulatory system to produce a steep boundary and minimize expression variability between different nuclei. We show how a limited readout time imposed by short developmental cycles affects the gene's ability to read positional information along the embryo's anterior posterior axis and express reliably. Comparing our theoretical results to real-time monitoring of the hunchback transcription dynamics in live flies, we discuss possible regulatory strategies, suggesting an important role for additional binding sites, gradients or non-equilibrium binding and modified transcription factor search strategies.
Jonathan Desponds, Carmina Perez Romero, Mathieu Coppey, Cecile Fradin, Nathalie Dostatni, Aleksandra M. Walczak
PLoS Comput. Biol.7
2017 Inferring repeat-protein energetics from evolutionary information
abstract
Natural protein sequences contain a record of their history. A common constraint in a given protein family is the ability to fold to specific structures, and it has been shown possible to infer the main native ensemble by analyzing covariations in extant sequences. Still, many natural proteins that fold into the same structural topology show different stabilization energies, and these are often related to their physiological behavior. We propose a description for the energetic variation given by sequence modifications in repeat proteins, systems for which the overall problem is simplified by their inherent symmetry. We explicitly account for single amino acid and pair-wise interactions and treat higher order correlations with a single term. We show that the resulting evolutionary field can be interpreted with structural detail. We trace the variations in the energetic scores of natural proteins and relate them to their experimental characterization. The resulting energetic evolutionary field allows the prediction of the folding free energy change for several mutants, and can be used to generate synthetic sequences that are statistically indistinguishable from the natural counterparts.
Rocío Espada, R. Gonzalo Parra, Thierry Mora, Aleksandra M. Walczak, Diego U. Ferreiro
PLoS Comput. Biol.4
2017 Variable habitat conditions drive species covariation in the human microbiota
abstract
Two species with similar resource requirements respond in a characteristic way to variations in their habitat-their abundances rise and fall in concert. We use this idea to learn how bacterial populations in the microbiota respond to habitat conditions that vary from person-to-person across the human population. Our mathematical framework shows that habitat fluctuations are sufficient for explaining intra-bodysite correlations in relative species abundances from the Human Microbiome Project. We explicitly show that the relative abundances of closely related species are positively correlated and can be predicted from taxonomic relationships. We identify a small set of functional pathways related to metabolism and maintenance of the cell wall that form the basis of a common resource sharing niche space of the human microbiota.
Charles K. Fisher, Thierry Mora, Aleksandra M. Walczak
PLoS Comput. Biol.3
2017 Persisting fetal clonotypes influence the structure and overlap of adult human T cell receptor repertoires
abstract
The diversity of T-cell receptors recognizing foreign pathogens is generated through a highly stochastic recombination process, making the independent production of the same sequence rare. Yet unrelated individuals do share receptors, which together constitute a "public" repertoire of abundant clonotypes. The TCR repertoire is initially formed prenatally, when the enzyme inserting random nucleotides is downregulated, producing a limited diversity subset. By statistically analyzing deep sequencing T-cell repertoire data from twins, unrelated individuals of various ages, and cord blood, we show that T-cell clones generated before birth persist and maintain high abundances in adult organisms for decades, slowly decaying with age. Our results suggest that large, low-diversity public clones are created during pre-natal life, and survive over long periods, providing the basis of the public repertoire.
Mikhail Pogorelyy, Yuval Elhanati, Quentin Marcou, Anastasiia L. Sycheva, Ekaterina Komech, Vadim Nazarov, Olga V. Britanova, Dmitry Chudakov, Ilgar Z. Mamedov, Yuri B. Lebedev, Thierry Mora, Aleksandra M. Walczak
PLoS Comput. Biol.12
2016 repgenHMM: a dynamic programming tool to infer the rules of immune receptor generation from sequence data
abstract
MOTIVATION: The diversity of the immune repertoire is initially generated by random rearrangements of the receptor gene during early T and B cell development. Rearrangement scenarios are composed of random events-choices of gene templates, base pair deletions and insertions-described by probability distributions. Not all scenarios are equally likely, and the same receptor sequence may be obtained in several different ways. Quantifying the distribution of these rearrangements is an essential baseline for studying the immune system diversity. Inferring the properties of the distributions from receptor sequences is a computationally hard problem, requiring enumerating every possible scenario for every sampled receptor sequence. RESULTS: We present a Hidden Markov model, which accounts for all plausible scenarios that can generate the receptor sequences. We developed and implemented a method based on the Baum-Welch algorithm that can efficiently infer the parameters for the different events of the rearrangement process. We tested our software tool on sequence data for both the alpha and beta chains of the T cell receptor. To test the validity of our algorithm, we also generated synthetic sequences produced by a known model, and confirmed that its parameters could be accurately inferred back from the sequences. The inferred model can be used to generate synthetic sequences, to calculate the probability of generation of any receptor sequence, as well as the theoretical diversity of the repertoire. We estimate this diversity to be [Formula: see text] for human T cells. The model gives a baseline to investigate the selection and dynamics of immune repertoires. AVAILABILITY AND IMPLEMENTATION: Source code and sample sequence files are available at https://bitbucket.org/yuvalel/repgenhmm/downloads CONTACT: [email protected] or [email protected] or [email protected].
Yuval Elhanati, Quentin Marcou, Thierry Mora, Aleksandra M. Walczak
Bioinform.4
2016 Precision of Readout at the hunchback Gene: Analyzing Short Transcription Time Traces in Living Fly Embryos
abstract
The simultaneous expression of the hunchback gene in the numerous nuclei of the developing fly embryo gives us a unique opportunity to study how transcription is regulated in living organisms. A recently developed MS2-MCP technique for imaging nascent messenger RNA in living Drosophila embryos allows us to quantify the dynamics of the developmental transcription process. The initial measurement of the morphogens by the hunchback promoter takes place during very short cell cycles, not only giving each nucleus little time for a precise readout, but also resulting in short time traces of transcription. Additionally, the relationship between the measured signal and the promoter state depends on the molecular design of the reporting probe. We develop an analysis approach based on tailor made autocorrelation functions that overcomes the short trace problems and quantifies the dynamics of transcription initiation. Based on live imaging data, we identify signatures of bursty transcription initiation from the hunchback promoter. We show that the precision of the expression of the hunchback gene to measure its position along the anterior-posterior axis is low both at the boundary and in the anterior even at cycle 13, suggesting additional post-transcriptional averaging mechanisms to provide the precision observed in fixed embryos.
Jonathan Desponds, Teresa Ferraro, Tanguy Lucas, Carmina Perez Romero, Aurelien Guillou, Cecile Fradin, Mathieu Coppey, Nathalie Dostatni, Aleksandra M. Walczak
PLoS Comput. Biol.10
2016 Noise Expands the Response Range of the Bacillus subtilis Competence Circuit
abstract
Gene regulatory circuits must contend with intrinsic noise that arises due to finite numbers of proteins. While some circuits act to reduce this noise, others appear to exploit it. A striking example is the competence circuit in Bacillus subtilis, which exhibits much larger noise in the duration of its competence events than a synthetically constructed analog that performs the same function. Here, using stochastic modeling and fluorescence microscopy, we show that this larger noise allows cells to exit terminal phenotypic states, which expands the range of stress levels to which cells are responsive and leads to phenotypic heterogeneity at the population level. This is an important example of how noise confers a functional benefit in a genetic decision-making circuit.
Andrew Mugler, Mark Kittisopikul, Luke Hayden, Chris Wiggins 0001, Gürol M. Süel, Aleksandra M. Walczak
PLoS Comput. Biol.7
2015 Capturing coevolutionary signals inrepeat proteins
abstract
BACKGROUND: The analysis of correlations of amino acid occurrences in globular domains has led to the development of statistical tools that can identify native contacts - portions of the chains that come to close distance in folded structural ensembles. Here we introduce a direct coupling analysis for repeat proteins - natural systems for which the identification of folding domains remains challenging. RESULTS: We show that the inherent translational symmetry of repeat protein sequences introduces a strong bias in the pair correlations at precisely the length scale of the repeat-unit. Equalizing for this bias in an objective way reveals true co-evolutionary signals from which local native contacts can be identified. Importantly, parameter values obtained for all other interactions are not significantly affected by the equalization. We quantify the robustness of the procedure and assign confidence levels to the interactions, identifying the minimum number of sequences needed to extract evolutionary information in several repeat protein families. CONCLUSIONS: The overall procedure can be used to reconstruct the interactions at distances larger than repeat-pairs, identifying the characteristics of the strongest couplings in each family, and can be applied to any system that appears translationally symmetric.
Rocío Espada, R. Gonzalo Parra, Thierry Mora, Aleksandra M. Walczak, Diego U. Ferreiro
BMC Bioinform.4
2008 The Energy Landscapes of Repeat-Containing Proteins: Topology, Cooperativity, and the Folding Funnels of One-Dimensional Architectures
abstract
Repeat-proteins are made up of near repetitions of 20- to 40-amino acid stretches. These polypeptides usually fold up into non-globular, elongated architectures that are stabilized by the interactions within each repeat and those between adjacent repeats, but that lack contacts between residues distant in sequence. The inherent symmetries both in primary sequence and three-dimensional structure are reflected in a folding landscape that may be analyzed as a quasi-one-dimensional problem. We present a general description of repeat-protein energy landscapes based on a formal Ising-like treatment of the elementary interaction energetics in and between foldons, whose collective ensemble are treated as spin variables. The overall folding properties of a complete "domain" (the stability and cooperativity of the repeating array) can be derived from this microscopic description. The one-dimensional nature of the model implies there are simple relations for the experimental observables: folding free-energy (DeltaG(water)) and the cooperativity of denaturation (m-value), which do not ordinarily apply for globular proteins. We show how the parameters for the "coarse-grained" description in terms of foldon spin variables can be extracted from more detailed folding simulations on perfectly funneled landscapes. To illustrate the ideas, we present a case-study of a family of tetratricopeptide (TPR) repeat proteins and quantitatively relate the results to the experimentally observed folding transitions. Based on the dramatic effect that single point mutations exert on the experimentally observed folding behavior, we speculate that natural repeat proteins are "poised" at particular ratios of inter- and intra-element interaction energetics that allow them to readily undergo structural transitions in physiologically relevant conditions, which may be intrinsically related to their biological functions.
Diego U. Ferreiro, Aleksandra M. Walczak, Elizabeth A. Komives, Peter G. Wolynes
PLoS Comput. Biol.2