Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ewa Szczurek

dblp:48/1715 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-1320-6695ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 5 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 30% Segmentation and scene understanding · 20% Generative modeling · 10%
Interdisciplinary, comprehensive, and emerging computing
7 papers
Bioinformatics and computational biology · 100%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
cancer genomics
0.942017
Epistasis in Genomic and Survival Data of Cancer Patients · RECOMB 2017
TiMEx: a waiting time model for mutually exclusive cancer alterations · Bioinform. 2016
Inferring the paths of somatic evolution in cancer · Bioinform. 2014
Computer vision › Segmentation and scene understanding › medical image segmentation
3d medical image segmentation
0.912025
Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation · NeurIPS 2025
Machine learning › Efficient and distributed learning
active learning
0.912025
ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods · NeurIPS 2025
Machine learning › Generative modeling
generative model
0.912025
ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods · NeurIPS 2025
Machine learning › Deep learning architectures and training › mixture of experts
hierarchical mixture of experts
0.912025
Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation · NeurIPS 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.912025
Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation · NeurIPS 2025
Computer vision › Segmentation and scene understanding
medical image segmentation
0.912025
Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation · NeurIPS 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo
0.912025
ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods · NeurIPS 2025
Machine learning › Deep learning architectures and training
state space model
0.912025
Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation · NeurIPS 2025
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
surrogate model
0.912025
ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods · NeurIPS 2025
Bioinformatics and computational biology
protein design
0.912025
ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods · NeurIPS 2025
Bioinformatics and computational biology › protein design
protein sequence design
0.912025
ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods · NeurIPS 2025
Bioinformatics and computational biology › cancer genomics
mutual exclusivity analysis
0.422016
TiMEx: a waiting time model for mutually exclusive cancer alterations · Bioinform. 2016
Modeling Mutual Exclusivity of Cancer Mutations · RECOMB 2014
Bioinformatics and computational biology › biological network › network biology
signaling network inference
0.412019
Learning signaling networks from combinatorial perturbations by exploiting siRNA off-target effects · Bioinform. 2019
Bioinformatics and computational biology › genomics
computational genomics
0.312017
Epistasis in Genomic and Survival Data of Cancer Patients · RECOMB 2017
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
epistasis analysis
0.312017
Epistasis in Genomic and Survival Data of Cancer Patients · RECOMB 2017
Bioinformatics and computational biology › functional genomics
perturbation effect estimation
0.212016
Linear effects models of signaling pathways from combinatorial perturbation data · Bioinform. 2016
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
signaling pathway analysis
0.212016
Linear effects models of signaling pathways from combinatorial perturbation data · Bioinform. 2016
Bioinformatics and computational biology
survival analysis
0.112017
Epistasis in Genomic and Survival Data of Cancer Patients · RECOMB 2017
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
epistasis
0.112014
Inferring the paths of somatic evolution in cancer · Bioinform. 2014

Methods — techniques the papers use, named apart from their topics

sequential monte carlo · 1.7pre-trained generative model · 1.7active learning · 1.7soft mixture-of-experts routing · 0.9mamba selective state space model · 0.9linear effects model · 0.6network structure learning · 0.4RNA interference · 0.4survival analysis · 0.3epistasis analysis · 0.3waiting time model · 0.2generative probabilistic model · 0.2
YearPublicationVenuePosition
2025 Factor Analysis with Correlated Topic Model for Multi-Modal Data
abstract
Integrating various data modalities brings valuable insights into underlying phenomena. Multimodal factor analysis (FA) uncovers shared axes of variation underlying different simple data modalities, where each sample is represented by a vector of features. However, FA is not suited for structured data modalities, such as text or single cell sequencing data, where multiple data points are measured per each sample and exhibit a clustering structure. To overcome this challenge, we introduce FACTM, a novel, multi-view and multi-structure Bayesian model that combines FA with correlated topic modeling and is optimized using variational inference. Additionally, we introduce a method for rotating latent factors to enhance interpretability with respect to binary features. On text and video benchmarks as well as real-world music and COVID-19 datasets, we demonstrate that FACTM outperforms other methods in identifying clusters in structured data, and integrating them with simple modalities via the inference of shared, interpretable factors.
Malgorzata Lazecka, Ewa Szczurek
AISTATS2
2025 ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods
abstract
Designing protein sequences of both high fitness and novelty is a challenging task in data-efficient protein engineering. Exploration beyond wild-type neighborhoods often leads to biologically implausible sequences or relies on surrogate models that lose fidelity in novel regions. Here, we propose ProSpero, an active learning framework in which a frozen pre-trained generative model is guided by a surrogate updated from oracle feedback. By integrating fitness-relevant residue selection with biologically-constrained Sequential Monte Carlo sampling, our approach enables exploration beyond wild-type neighborhoods while preserving biological plausibility. We show that our framework remains effective even when the surrogate is misspecified. ProSpero consistently outperforms or matches existing methods across diverse protein engineering tasks, retrieving sequences of both high fitness and novelty.
Michal Kmicikiewicz, Vincent Fortuin, Ewa Szczurek
NeurIPS3
2025 Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation
abstract
In recent years, artificial intelligence has significantly advanced medical image segmentation. Nonetheless, challenges remain, including efficient 3D medical image processing across diverse modalities and handling data variability. In this work, we introduce Hierarchical Soft Mixture-of-Experts (HoME), a two-level token-routing layer for efficient long-context modeling, specifically designed for 3D medical image segmentation. Built on the Mamba Selective State Space Model (SSM) backbone, HoME enhances sequential modeling through adaptive expert routing. In the first level, a Soft Mixture-of-Experts (SMoE) layer partitions input sequences into local groups, routing tokens to specialized per-group experts for localized feature extraction. The second level aggregates these outputs through a global SMoE layer, enabling cross-group information fusion and global context refinement. This hierarchical design, combining local expert routing with global expert refinement, enhances generalizability and segmentation performance, surpassing state-of-the-art results across datasets from the three most widely used 3D medical imaging modalities and varying data qualities. The code is publicly available at https://github.com/gmum/MambaHoME.
Szymon Plotka, Gizem Mert, Maciej Chrabaszcz, Ewa Szczurek, Arkadiusz Sitek
NeurIPS4
2020 A mathematical model of the metastatic bottleneck predicts patient outcome and response to cancer treatment
abstract
Metastases are the main reason for cancer-related deaths. Initiation of metastases, where newly seeded tumor cells expand into colonies, presents a tremendous bottleneck to metastasis formation. Despite its importance, a quantitative description of metastasis initiation and its clinical implications is lacking. Here, we set theoretical grounds for the metastatic bottleneck with a simple stochastic model. The model assumes that the proliferation-to-death rate ratio for the initiating metastatic cells increases when they are surrounded by more of their kind. For a total of 159,191 patients across 13 cancer types, we found that a single cell has an extremely low median probability of successful seeding of the order of 10-8. With increasing colony size, a sharp transition from very unlikely to very likely successful metastasis initiation occurs. The median metastatic bottleneck, defined as the critical colony size that marks this transition, was between 10 and 21 cells. We derived the probability of metastasis occurrence and patient outcome based on primary tumor size at diagnosis and tumor type. The model predicts that the efficacy of patient treatment depends on the primary tumor size but even more so on the severity of the metastatic bottleneck, which is estimated to largely vary between patients. We find that medical interventions aiming at tightening the bottleneck, such as immunotherapy, can be much more efficient than therapies that decrease overall tumor burden, such as chemotherapy.
Ewa Szczurek, Tyll Krüger, Barbara Klink, Niko Beerenwinkel
PLoS Comput. Biol.1
2019 Learning signaling networks from combinatorial perturbations by exploiting siRNA off-target effects
abstract
MOTIVATION: Perturbation experiments constitute the central means to study cellular networks. Several confounding factors complicate computational modeling of signaling networks from this data. First, the technique of RNA interference (RNAi), designed and commonly used to knock-down specific genes, suffers from off-target effects. As a result, each experiment is a combinatorial perturbation of multiple genes. Second, the perturbations propagate along unknown connections in the signaling network. Once the signal is blocked by perturbation, proteins downstream of the targeted proteins also become inactivated. Finally, all perturbed network members, either directly targeted by the experiment, or by propagation in the network, contribute to the observed effect, either in a positive or negative manner. One of the key questions of computational inference of signaling networks from such data are, how many and what combinations of perturbations are required to uniquely and accurately infer the model? RESULTS: Here, we introduce an enhanced version of linear effects models (LEMs), which extends the original by accounting for both negative and positive contributions of the perturbed network proteins to the observed phenotype. We prove that the enhanced LEMs are identified from data measured under perturbations of all single, pairs and triplets of network proteins. For small networks of up to five nodes, only perturbations of single and pairs of proteins are required for identifiability. Extensive simulations demonstrate that enhanced LEMs achieve excellent accuracy of parameter estimation and network structure learning, outperforming the previous version on realistic data. LEMs applied to Bartonella henselae infection RNAi screening data identified known interactions between eight nodes of the infection network, confirming high specificity of our model and suggested one new interaction. AVAILABILITY AND IMPLEMENTATION: https://github.com/EwaSzczurek/LEM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jerzy Tiuryn, Ewa Szczurek
Bioinform.2
2019 Correction: Epistasis in genomic and survival data of cancer patients
abstract
[This corrects the article DOI: 10.1371/journal.pcbi.1005626.].
Dariusz Matlak, Ewa Szczurek
PLoS Comput. Biol.2
2017 Epistasis in Genomic and Survival Data of Cancer Patients
Dariusz Matlak, Ewa Szczurek
RECOMB2
2017 Epistasis in genomic and survival data of cancer patients
abstract
Cancer aggressiveness and its effect on patient survival depends on mutations in the tumor genome. Epistatic interactions between the mutated genes may guide the choice of anticancer therapy and set predictive factors of its success. Inhibitors targeting synthetic lethal partners of genes mutated in tumors are already utilized for efficient and specific treatment in the clinic. The space of possible epistatic interactions, however, is overwhelming, and computational methods are needed to limit the experimental effort of validating the interactions for therapy and characterizing their biomarkers. Here, we introduce SurvLRT, a statistical likelihood ratio test for identifying epistatic gene pairs and triplets from cancer patient genomic and survival data. Compared to established approaches, SurvLRT performed favorable in predicting known, experimentally verified synthetic lethal partners of PARP1 from TCGA data. Our approach is the first to test for epistasis between triplets of genes to identify biomarkers of synthetic lethality-based therapy. SurvLRT proved successful in identifying the known gene TP53BP1 as the biomarker of success of the therapy targeting PARP in BRCA1 deficient tumors. Search for other biomarkers for the same interaction revealed a region whose deletion was a more significant biomarker than deletion of TP53BP1. With the ability to detect not only pairwise but twelve different types of triple epistasis, applicability of SurvLRT goes beyond cancer therapy, to the level of characterization of shapes of fitness landscapes.
Dariusz Matlak, Ewa Szczurek
PLoS Comput. Biol.2
2016 TiMEx: a waiting time model for mutually exclusive cancer alterations
abstract
MOTIVATION: Despite recent technological advances in genomic sciences, our understanding of cancer progression and its driving genetic alterations remains incomplete. RESULTS: We introduce TiMEx, a generative probabilistic model for detecting patterns of various degrees of mutual exclusivity across genetic alterations, which can indicate pathways involved in cancer progression. TiMEx explicitly accounts for the temporal interplay between the waiting times to alterations and the observation time. In simulation studies, we show that our model outperforms previous methods for detecting mutual exclusivity. On large-scale biological datasets, TiMEx identifies gene groups with strong functional biological relevance, while also proposing new candidates for biological validation. TiMEx possesses several advantages over previous methods, including a novel generative probabilistic model of tumorigenesis, direct estimation of the probability of mutual exclusivity interaction, computational efficiency and high sensitivity in detecting gene groups involving low-frequency alterations. AVAILABILITY AND IMPLEMENTATION: TiMEx is available as a Bioconductor R package at www.bsse.ethz.ch/cbg/software/TiMEx CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Simona Constantinescu, Ewa Szczurek, Pejman Mohammadi 0001, Jörg Rahnenführer, Niko Beerenwinkel
Bioinform.2
2016 Linear effects models of signaling pathways from combinatorial perturbation data
abstract
MOTIVATION: Perturbations constitute the central means to study signaling pathways. Interrupting components of the pathway and analyzing observed effects of those interruptions can give insight into unknown connections within the signaling pathway itself, as well as the link from the pathway to the effects. Different pathway components may have different individual contributions to the measured perturbation effects, such as gene expression changes. Those effects will be observed in combination when the pathway components are perturbed. Extant approaches focus either on the reconstruction of pathway structure or on resolving how the pathway components control the downstream effects. RESULTS: Here, we propose a linear effects model, which can be applied to solve both these problems from combinatorial perturbation data. We use simulated data to demonstrate the accuracy of learning the pathway structure as well as estimation of the individual contributions of pathway components to the perturbation effects. The practical utility of our approach is illustrated by an application to perturbations of the mitogen-activated protein kinase pathway in Saccharomyces cerevisiaeAvailability and Implementation: lem is available as a R package at http://www.mimuw.edu.pl/∼szczurek/lem CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ewa Szczurek, Niko Beerenwinkel
Bioinform.1
2014 Modeling Mutual Exclusivity of Cancer Mutations
Ewa Szczurek, Niko Beerenwinkel
RECOMB1
2014 Inferring the paths of somatic evolution in cancer
abstract
MOTIVATION: Cancer cell genomes acquire several genetic alterations during somatic evolution from a normal cell type. The relative order in which these mutations accumulate and contribute to cell fitness is affected by epistatic interactions. Inferring their evolutionary history is challenging because of the large number of mutations acquired by cancer cells as well as the presence of unknown epistatic interactions. RESULTS: We developed Bayesian Mutation Landscape (BML), a probabilistic approach for reconstructing ancestral genotypes from tumor samples for much larger sets of genes than previously feasible. BML infers the likely sequence of mutation accumulation for any set of genes that is recurrently mutated in tumor samples. When applied to tumor samples from colorectal, glioblastoma, lung and ovarian cancer patients, BML identifies the diverse evolutionary scenarios involved in tumor initiation and progression in greater detail, but broadly in agreement with prior results. AVAILABILITY AND IMPLEMENTATION: Source code and all datasets are freely available at bml.molgen.mpg.de. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Navodit Misra, Ewa Szczurek, Martin Vingron
Bioinform.2
2014 Modeling Mutual Exclusivity of Cancer Mutations
abstract
In large collections of tumor samples, it has been observed that sets of genes that are commonly involved in the same cancer pathways tend not to occur mutated together in the same patient.Such gene sets form mutually exclusive patterns of gene alterations in cancer genomic data.Computational approaches that detect mutually exclusive gene sets, rank and test candidate alteration patterns by rewarding the number of samples the pattern covers and by punishing its impurity, i.e., additional alterations that violate strict mutual exclusivity.However, the extant approaches do not account for possible observation errors.In practice, false negatives and especially false positives can severely bias evaluation and ranking of alteration patterns.To address these limitations, we develop a fully probabilistic, generative model of mutual exclusivity, explicitly taking coverage, impurity, as well as error rates into account, and devise efficient algorithms for parameter estimation and pattern ranking.Based on this model, we derive a statistical test of mutual exclusivity by comparing its likelihood to the null model that assumes independent gene alterations.Using extensive simulations, the new test is shown to be more powerful than a permutation test applied previously.When applied to detect mutual exclusivity patterns in glioblastoma and in pan-cancer data from twelve tumor types, we identify several significant patterns that are biologically relevant, most of which would not be detected by previous approaches.Our statistical modeling framework of mutual exclusivity provides increased flexibility and power to detect cancer pathways from genomic alteration data in the presence of noise.A summary of this paper appears in the proceedings of the RECOMB 2014 conference, April 2 5.
Ewa Szczurek, Niko Beerenwinkel
PLoS Comput. Biol.1
2011 Deregulation upon DNA damage revealed by joint analysis of context-specific perturbation data
abstract
BACKGROUND: Deregulation between two different cell populations manifests itself in changing gene expression patterns and changing regulatory interactions. Accumulating knowledge about biological networks creates an opportunity to study these changes in their cellular context. RESULTS: We analyze re-wiring of regulatory networks based on cell population-specific perturbation data and knowledge about signaling pathways and their target genes. We quantify deregulation by merging regulatory signal from the two cell populations into one score. This joint approach, called JODA, proves advantageous over separate analysis of the cell populations and analysis without incorporation of knowledge. JODA is implemented and freely available in a Bioconductor package 'joda'. CONCLUSIONS: Using JODA, we show wide-spread re-wiring of gene regulatory networks upon neocarzinostatin-induced DNA damage in Human cells. We recover 645 deregulated genes in thirteen functional clusters performing the rich program of response to damage. We find that the clusters contain many previously characterized neocarzinostatin target genes. We investigate connectivity between those genes, explaining their cooperation in performing the common functions. We review genes with the most extreme deregulation scores, reporting their involvement in response to DNA damage. Finally, we investigate the indirect impact of the ATM pathway on the deregulated genes, and build a hypothetical hierarchy of direct regulation. These results prove that JODA is a step forward to a systems level, mechanistic understanding of changes in gene regulation between different cell populations.
Ewa Szczurek, Florian Markowetz, Irit Gat-Viks, Przemyslaw Biecek, Jerzy Tiuryn, Martin Vingron
BMC Bioinform.1
2009 On Subset Seeds for Protein Alignment
abstract
We apply the concept of subset seeds to similarity search in protein sequences. The main question studied is the design of efficient seed alphabets to construct seeds with optimal sensitivity/selectivity trade-offs. We propose several different design methods and use them to construct several alphabets. We then perform a comparative analysis of seeds built over those alphabets and compare them with the standard Blastp seeding method, as well as with the family of vector seeds. While the formalism of subset seeds is less expressive (but less costly to implement) than the cumulative principle used in Blastp and vector seeds, our seeds show a similar or even better performance than Blastp on Bernoulli models of proteins compatible with the common BLOSUM62 matrix. Finally, we perform a large-scale benchmarking of our seeds against several main databases of protein alignments. Here again, the results show a comparable or better performance of our seeds versus Blastp.
Mikhail A. Roytberg, Anna Gambin, Laurent Noé, Slawomir Lasota 0001, Eugenia Furletova, Ewa Szczurek, Gregory Kucherov
IEEE ACM Trans. Comput. Biol. Bioinform.6