EDBT 2026 Demo / reviewers in the wild / expert
Sarah A. Teichmann
dblp:90/4458
· DBLP profile ↗
14ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0002-6294-6366ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 96% Medical and health informatics · 4% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics |
1.6 | 2 | 2025 | Estimation of single-cell and tissue perturbation effect in spatial transcriptomics via Spatial Causal Disentanglement · ICLR 2025 Bin2cell reconstructs cells from high resolution Visium HD data · Bioinform. 2024 |
Bioinformatics and computational biology › gene regulation
gene regulatory network |
0.9 | 1 | 2025 | Estimation of single-cell and tissue perturbation effect in spatial transcriptomics via Spatial Causal Disentanglement · ICLR 2025 |
Bioinformatics and computational biology › functional genomics
perturbation effect estimation |
0.9 | 1 | 2025 | Estimation of single-cell and tissue perturbation effect in spatial transcriptomics via Spatial Causal Disentanglement · ICLR 2025 |
Bioinformatics and computational biology › omics data analysis
batch effect correction |
0.4 | 1 | 2020 | BBKNN: fast batch alignment of single cell transcriptomes · Bioinform. 2020 |
Bioinformatics and computational biology › single-cell analysis
single-cell transcriptomics |
0.4 | 1 | 2020 | BBKNN: fast batch alignment of single cell transcriptomes · Bioinform. 2020 |
Medical and health informatics › medical imaging › medical image analysis
medical image segmentation |
0.2 | 1 | 2024 | Bin2cell reconstructs cells from high resolution Visium HD data · Bioinform. 2024 |
Bioinformatics and computational biology
data integration |
0.1 | 1 | 2020 | BBKNN: fast batch alignment of single cell transcriptomes · Bioinform. 2020 |
Bioinformatics and computational biology
genomics |
0.1 | 2 | 2006 | FlyTF: a systematic review of site-specific transcription factors in the fruit fly Drosophila melanogaster · Bioinform. 2006 Genome sequences and protein structures · RECOMB 1999 |
Bioinformatics and computational biology › gene regulation › transcription factor analysis
transcription factor identification |
0.1 | 1 | 2006 | FlyTF: a systematic review of site-specific transcription factors in the fruit fly Drosophila melanogaster · Bioinform. 2006 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.1 | 1 | 2005 | Statistical analysis of domains in interacting protein pairs · Bioinform. 2005 |
Bioinformatics and computational biology
proteomics |
0.1 | 1 | 2005 | Statistical analysis of domains in interacting protein pairs · Bioinform. 2005 |
Bioinformatics and computational biology › sequence analysis
sequence similarity search |
0.1 | 2 | 2000 | Fast assignment of protein structures to sequences using the Intermediate Sequence Library PDB-ISL · Bioinform. 2000 PDB_ISL: an intermediate sequence library for protein structure assignment · RECOMB 2000 |
Bioinformatics and computational biology › genomics
structural genomics |
0.1 | 2 | 2000 | PDB_ISL: an intermediate sequence library for protein structure assignment · RECOMB 2000 Genome sequences and protein structures · RECOMB 1999 |
Bioinformatics and computational biology › protein structure analysis
protein structure annotation |
0.0 | 1 | 2000 | PDB_ISL: an intermediate sequence library for protein structure assignment · RECOMB 2000 |
Bioinformatics and computational biology
protein structure prediction |
0.0 | 1 | 2000 | Fast assignment of protein structures to sequences using the Intermediate Sequence Library PDB-ISL · Bioinform. 2000 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 1 | 2000 | PDB_ISL: an intermediate sequence library for protein structure assignment · RECOMB 2000 |
Bioinformatics and computational biology › sequence analysis
sequence comparison |
0.0 | 1 | 2000 | PDB_ISL: an intermediate sequence library for protein structure assignment · RECOMB 2000 |
Bioinformatics and computational biology › genomics
genome sequencing |
0.0 | 1 | 1999 | Genome sequences and protein structures · RECOMB 1999 |
Bioinformatics and computational biology › structural bioinformatics
protein structure |
0.0 | 1 | 1999 | Genome sequences and protein structures · RECOMB 1999 |
Bioinformatics and computational biology › protein function prediction › protein classification
protein family classification |
0.0 | 1 | 1998 | DIVCLUS: an automatic method in the GEANFAMMER package that finds homologous domains in single- and multi-domain proteins · Bioinform. 1998 |
Bioinformatics and computational biology › protein sequence analysis
protein sequence clustering |
0.0 | 1 | 1998 | DIVCLUS: an automatic method in the GEANFAMMER package that finds homologous domains in single- and multi-domain proteins · Bioinform. 1998 |
Bioinformatics and computational biology
structural bioinformatics |
0.0 | 1 | 2005 | Statistical analysis of domains in interacting protein pairs · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
graph neural network · 0.9generative model · 0.9causal inference · 0.9graph-based integration · 0.4structural DNA-binding domain assignment · 0.1manual annotation · 0.1gene ontology annotation · 0.1statistical significance testing · 0.1power-law fitting · 0.1PSI-BLAST · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Estimation of single-cell and tissue perturbation effect in spatial transcriptomics via Spatial Causal DisentanglementabstractModels of Virtual Cells and Virtual Tissues at single-cell resolution would allow us to test perturbations in silico and accelerate progress in tissue and cell engineering.
However, most such models are not rooted in causal inference and as a result, could mistake correlation for causation.
We introduce Celcomen, a novel generative graph neural network grounded in mathematical causality to disentangle intra- and inter-cellular gene regulation in spatial transcriptomics and single-cell data.
Celcomen can also be prompted by perturbations to generate spatial counterfactuals, thus offering insights into experimentally inaccessible states, with potential applications in human health.
We validate the model's disentanglement and identifiability through simulations, and demonstrate its counterfactual predictions in clinically relevant settings, including human glioblastoma and fetal spleen, recovering inflammation-related gene programs post immune system perturbation.
Moreover, it supports mechanistic interpretability, as its parameters can be reverse-engineered from observed behavior, making it an accessible model for understanding both neural networks and complex biological systems. Stathis Megas, Daniel G. Chen, Krzysztof Polanski, Moshe Eliasof, Carola-Bibiane Schönlieb, Sarah A. Teichmann |
ICLR | 6 |
| 2024 | Bin2cell reconstructs cells from high resolution Visium HD dataabstractSUMMARY: Visium HD by 10X Genomics is the first commercially available platform capable of capturing full scale transcriptomic data paired with a reference morphology image from archived FFPE blocks at sub-cellular resolution. However, aggregation of capture regions to single cells poses challenges. Bin2cell reconstructs cells from the highest resolution data (2 μm bins) by leveraging morphology image segmentation and gene expression information. It is compatible with established Python single cell and spatial transcriptomics software, and operates efficiently in a matter of minutes without requiring a GPU. We demonstrate improvements in downstream analysis when using the reconstructed cells over default 8 μm bins on mouse brain and human colorectal cancer data. AVAILABILITY AND IMPLEMENTATION: Bin2cell is available at https://github.com/Teichlab/bin2cell, along with documentation and usage examples, and can be installed from pip. Probe design functionality is available at https://github.com/Teichlab/gene2probe. Krzysztof Polanski, Raquel Bartolomé-Casado, Ioannis Sarropoulos, Nick England, Frode L. Jahnsen, Sarah A. Teichmann, Nadav Yayon |
Bioinform. | 7 |
| 2020 | BBKNN: fast batch alignment of single cell transcriptomesabstractMOTIVATION: Increasing numbers of large scale single cell RNA-Seq projects are leading to a data explosion, which can only be fully exploited through data integration. A number of methods have been developed to combine diverse datasets by removing technical batch effects, but most are computationally intensive. To overcome the challenge of enormous datasets, we have developed BBKNN, an extremely fast graph-based data integration algorithm. We illustrate the power of BBKNN on large scale mouse atlasing data, and favourably benchmark its run time against a number of competing methods. AVAILABILITY AND IMPLEMENTATION: BBKNN is available at https://github.com/Teichlab/bbknn, along with documentation and multiple example notebooks, and can be installed from pip. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Krzysztof Polanski, Matthew D. Young, Zhichao Miao, Kerstin B. Meyer, Sarah A. Teichmann, Jong-Eun Park |
Bioinform. | 5 |
| 2010 | Assessing Computational Methods of Cis-Regulatory Module PredictionabstractComputational methods attempting to identify instances of cis-regulatory modules (CRMs) in the genome face a challenging problem of searching for potentially interacting transcription factor binding sites while knowledge of the specific interactions involved remains limited. Without a comprehensive comparison of their performance, the reliability and accuracy of these tools remains unclear. Faced with a large number of different tools that address this problem, we summarized and categorized them based on search strategy and input data requirements. Twelve representative methods were chosen and applied to predict CRMs from the Drosophila CRM database REDfly, and across the human ENCODE regions. Our results show that the optimal choice of method varies depending on species and composition of the sequences in question. When discriminating CRMs from non-coding regions, those methods considering evolutionary conservation have a stronger predictive power than methods designed to be run on a single genome. Different CRM representations and search strategies rely on different CRM properties, and different methods can complement one another. For example, some favour homotypical clusters of binding sites, while others perform best on short CRMs. Furthermore, most methods appear to be sensitive to the composition and structure of the genome to which they are applied. We analyze the principal features that distinguish the methods that performed well, identify weaknesses leading to poor performance, and provide a guide for users. We also propose key considerations for the development and evaluation of future CRM-prediction methods. Sarah A. Teichmann, Thomas A. Down |
PLoS Comput. Biol. | 2 |
| 2009 | Protein domain organisation: adding orderabstractBACKGROUND: Domains are the building blocks of proteins. During evolution, they have been duplicated, fused and recombined, to produce proteins with novel structures and functions. Structural and genome-scale studies have shown that pairs or groups of domains observed together in a protein are almost always found in only one N to C terminal order and are the result of a single recombination event that has been propagated by duplication of the multi-domain unit. Previous studies of domain organisation have used graph theory to represent the co-occurrence of domains within proteins. We build on this approach by adding directionality to the graphs and connecting nodes based on their relative order in the protein. Most of the time, the linear order of domains is conserved. However, using the directed graph representation we have identified non-linear features of domain organization that are over-represented in genomes. Recognising these patterns and unravelling how they have arisen may allow us to understand the functional relationships between domains and understand how the protein repertoire has evolved. RESULTS: We identify groups of domains that are not linearly conserved, but instead have been shuffled during evolution so that they occur in multiple different orders. We consider 192 genomes across all three kingdoms of life and use domain and protein annotation to understand their functional significance. To identify these features and assess their statistical significance, we represent the linear order of domains in proteins as a directed graph and apply graph theoretical methods. We describe two higher-order patterns of domain organisation: clusters and bi-directionally associated domain pairs and explore their functional importance and phylogenetic conservation. CONCLUSION: Taking into account the order of domains, we have derived a novel picture of global protein organization. We found that all genomes have a higher than expected degree of clustering and more domain pairs in forward and reverse orientation in different proteins relative to random graphs with identical degree distributions. While these features were statistically over-represented, they are still fairly rare. Looking in detail at the proteins involved, we found strong functional relationships within each cluster. In addition, the domains tended to be involved in protein-protein interaction and are able to function as independent structural units. A particularly striking example was the human Jak-STAT signalling pathway which makes use of a set of domains in a range of orders and orientations to provide nuanced signaling functionality. This illustrated the importance of functional and structural constraints (or lack thereof) on domain organisation. Sarah K. Kummerfeld, Sarah A. Teichmann |
BMC Bioinform. | 2 |
| 2007 | The (In)dependence of Alternative Splicing and Gene DuplicationabstractAlternative splicing (AS) and gene duplication (GD) both are processes that diversify the protein repertoire. Recent examples have shown that sequence changes introduced by AS may be comparable to those introduced by GD. In addition, the two processes are inversely correlated at the genomic scale: large gene families are depleted in splice variants and vice versa. All together, these data strongly suggest that both phenomena result in interchangeability between their effects. Here, we tested the extent to which this applies with respect to various protein characteristics. The amounts of AS and GD per gene are anticorrelated even when accounting for different gene functions or degrees of sequence divergence. In contrast, the two processes appear to be independent in their influence on variation in mRNA expression. Further, we conducted a detailed comparison of the effect of sequence changes in both alternative splice variants and gene duplicates on protein structure, in particular the size, location, and types of sequence substitutions and insertions/deletions. We find that, in general, alternative splicing affects protein sequence and structure in a more drastic way than gene duplication and subsequent divergence. Our results reveal an interesting paradox between the anticorrelation of AS and GD at the genomic level, and their impact at the protein level, which shows little or no equivalence in terms of effects on protein sequence, structure, and function. We discuss possible explanations that relate to the order of appearance of AS and GD in a gene family, and to the selection pressure imposed by the environment. David Talavera, Christine Vogel, Modesto Orozco, Sarah A. Teichmann, Xavier de la Cruz |
PLoS Comput. Biol. | 4 |
| 2006 | FlyTF: a systematic review of site-specific transcription factors in the fruit fly Drosophila melanogasterabstractSUMMARY: We present a manually annotated catalogue of site-specific transcription factors (TFs) in the fruit fly Drosophila melanogaster. These were identified from a list of candidate proteins with transcription-related Gene Ontology (Go) annotation as well as structural DNA-binding domain assignments. For all 1052 candidate proteins, a defined set of rules was applied to classify information from the literature and computational data sources with respect to both DNA-binding and transcriptional regulatory properties. We propose a set of 753 TFs in the fruit fly, of which 23 are confident novel predictions of this function for previously uncharacterized proteins. Boris Adryan, Sarah A. Teichmann |
Bioinform. | 2 |
| 2006 | 3D Complex: A Structural Classification of Protein ComplexesabstractMost of the proteins in a cell assemble into complexes to carry out their function. It is therefore crucial to understand the physicochemical properties as well as the evolution of interactions between proteins. The Protein Data Bank represents an important source of information for such studies, because more than half of the structures are homo- or heteromeric protein complexes. Here we propose the first hierarchical classification of whole protein complexes of known 3-D structure, based on representing their fundamental structural features as a graph. This classification provides the first overview of all the complexes in the Protein Data Bank and allows nonredundant sets to be derived at different levels of detail. This reveals that between one-half and two-thirds of known structures are multimeric, depending on the level of redundancy accepted. We also analyse the structures in terms of the topological arrangement of their subunits and find that they form a small number of arrangements compared with all theoretically possible ones. This is because most complexes contain four subunits or less, and the large majority are homomeric. In addition, there is a strong tendency for symmetry in complexes, even for heteromeric complexes. Finally, through comparison of Biological Units in the Protein Data Bank with the Protein Quaternary Structure database, we identified many possible errors in quaternary structure assignments. Our classification, available as a database and Web server at http://www.3Dcomplex.org, will be a starting point for future work aimed at understanding the structure and evolution of protein complexes. Emmanuel D. Levy, José B. Pereira-Leal, Cyrus Chothia, Sarah A. Teichmann |
PLoS Comput. Biol. | 4 |
| 2005 | The properties of protein family space depend on experimental designabstractMOTIVATION: Databases of protein families often exhibit drastically different properties of the protein family space. RESULTS: We compared the properties of protein family space as reflected by exhaustive protein family databases and databases with predefined families. We used TRIBES, Protomap, ProDom and COGs as representatives of the exhaustive databases, and Pfam-A and Superfamily as databases that predefine families. We observe a power-law distribution of family sizes in all these databases, albeit in predefined databases the power-law line collapses before reaching smaller sized families. We discuss the future trends of this power-law distribution and suggest that saturation in the sampling of protein family space will result in a distortion of the power law in small family sizes. For larger genome sizes, predefined databases show logarithmic growth of the number of families per genome, whereas exhaustive databases exhibit a virtually linear relationship. All databases consistently differ in the proportion of protein families shared between taxa. Predefined databases have a larger number of protein families shared between the three domains of life, while exhaustive databases show a much more fragmented distribution. We argue that these discrepancies reflect alternative approaches to the trade-off issue of sensitivity versus specificity in the detection of homologous proteins. We conclude that these properties are complementary rather than contradictory, while describing the protein universe from different perspectives. Victor Kunin, Sarah A. Teichmann, Martijn A. Huynen, Christos A. Ouzounis |
Bioinform. | 2 |
| 2005 | Statistical analysis of domains in interacting protein pairsabstractMOTIVATION: Several methods have recently been developed to analyse large-scale sets of physical interactions between proteins in terms of physical contacts between the constituent domains, often with a view to predicting new pairwise interactions. Our aim is to combine genomic interaction data, in which domain-domain contacts are not explicitly reported, with the domain-level structure of individual proteins, in order to learn about the structure of interacting protein pairs. Our approach is driven by the need to assess the evidence for physical contacts between domains in a statistically rigorous way. RESULTS: We develop a statistical approach that assigns p-values to pairs of domain superfamilies, measuring the strength of evidence within a set of protein interactions that domains from these superfamilies form contacts. A set of p-values is calculated for SCOP superfamily pairs, based on a pooled data set of interactions from yeast. These p-values can be used to predict which domains come into contact in an interacting protein pair. This predictive scheme is tested against protein complexes in the Protein Quaternary Structure (PQS) database, and is used to predict domain-domain contacts within 705 interacting protein pairs taken from our pooled data set. Tom M. W. Nye, Carlo Berzuini, Walter R. Gilks, M. Madan Babu, Sarah A. Teichmann |
Bioinform. | 5 |
| 2000 | PDB_ISL: an intermediate sequence library for protein structure assignmentabstractFor large scale structural assignment to sequences, as in computational structural genomics, a fast yet sensitive homology search procedure is essential. A new approach using intermediate sequences was tested as a shortcut to iterative multiple sequence search methods such as PSI-BLAST and hidden Markov models. A library containing potential intermediate sequences for proteins of known structure (PDB_ISL) was constructed. The sequences in the library were collected from a large sequence database using the sequences of the domains of proteins of known structure as the query sequences and the program PSI-BLAST. Sequences of proteins of unknown structure can be matched to distantly related proteins of known structure by using any pairwise sequence comparison methods to find homologues in PDB_ISL. Searches of PDB_ISL were calibrated, and the number of correct matches found at a given error rate was the same as that found by PSI_BLAST. The advantage of this library is that it uses pairwise sequence comparison methods, such as FASTA or BLAST2, and can, therefore, be searched easily and, in many cases, much more quickly than an iterative multiple sequence comparison method. The procedure is roughly twenty times faster than PSI-BLAST for small genomes and several hundred times for large genomes such as C. elegans. Sarah A. Teichmann, Cyrus Chothia, George M. Church, Jong-Chan Park |
RECOMB | 1 |
| 2000 | Fast assignment of protein structures to sequences using the Intermediate Sequence Library PDB-ISLabstractMOTIVATION: For large-scale structural assignment to sequences, as in computational structural genomics, a fast yet sensitive sequence search procedure is essential. A new approach using intermediate sequences was tested as a shortcut to iterative multiple sequence search methods such as PSI-BLAST. RESULTS: A library containing potential intermediate sequences for proteins of known structure (PDB-ISL) was constructed. The sequences in the library were collected from a large sequence database using the sequences of the domains of proteins of known structure as the query sequences and the program PSI-BLAST. Sequences of proteins of unknown structure can be matched to distantly related proteins of known structure by using pairwise sequence comparison methods to find homologues in PDB-ISL. Searches of PDB-ISL were calibrated, and the number of correct matches found at a given error rate was the same as that found by PSI-BLAST. The advantage of this library is that it uses pairwise sequence comparison methods, such as FASTA or BLAST2, and can, therefore, be searched easily and, in many cases, much more quickly than an iterative multiple sequence comparison method. The procedure is roughly 20 times faster than PSI-BLAST for small genomes and several hundred times for large genomes. AVAILABILITY: Sequences can be submitted to the PDB-ISL servers at http://stash.mrc-lmb.cam.ac.uk/PDB_ISL/ or http://cyrah.ebi.ac.uk:1111/Serv/PDB_ISL/ and can be downloaded from ftp://ftp.ebi.ac.uk/pub/contrib/jong/PDB_+ ++ISL/ CONTACT: [email protected] and [email protected] Sarah A. Teichmann, Cyrus Chothia, George M. Church, Jong-Chan Park |
Bioinform. | 1 |
| 1999 | Genome sequences and protein structures
Sarah A. Teichmann, Jong-Chan Park, Cyrus Chothia |
RECOMB | 1 |
| 1998 | DIVCLUS: an automatic method in the GEANFAMMER package that finds homologous domains in single- and multi-domain proteinsabstractMOTIVATION: Large-scale determination of relationships between the proteins produced by genome sequences is now common. All protein sequences are matched and those that have high match scores are clustered into families. In cases where the proteins are built of several domains or duplication modules, this can lead to misleading results. Consider the very simple example of three proteins: 1, formed by duplication modules A and B; 2, formed by duplication modules B' and C; and 3, formed by duplication modules C' and D. Duplication modules B and B' are homologous, as are C and C'. Matching the sequences of 1, 2 and 3 followed by simple single-linkage clustering would put all three in the same family, even though proteins 1 and 3 are not related. This is because the different parts of 2 match 1 and 3. This paper describes a procedure, DIVCLUS, that divides such complex clusters of partially related sequences into simple clusters that contain only related duplication modules. In the example just given, it would produce two groups of sequences: the first with domains B of sequence 1 and B of sequence 2, and the second with domain C of sequence 2 and C of sequence 3. DIVCLUS is part of a package called GEANFAMMER, for GEnome ANalysis and protein FAMily MakER. The package automates the detection of families of duplication modules from a protein sequence database. RESULTS: DIVCLUS has been applied to the division of single-linkage clusters generated from the protein sequences of six completely sequenced bacterial genomes. Out of 12 013 genes in these six genomes, 4563 single- and multi-domain sequences formed 1071 complex clusters. Application of the DIVCLUS program resolved these clusters into 2113 clusters corresponding to single duplication modules. AVAILABILITY: The perl5 program and its documentation are available at the following address: http://www.mrc-lmb.cam.ac.uk/genomes/ and by anonymous ftp at ftp.mrc-lmb.cam.ac.uk in the directory /pub/genomes/Software/. CONTACT: [email protected]; jong@mrc-lmb. cam.ac.uk Jong-Chan Park, Sarah A. Teichmann |
Bioinform. | 2 |