Helena U. Zacharias

dblp:197/8303 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-3633-1330ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021
YearPublicationVenuePosition
2025 HIDE: hierarchical cell-type deconvolution
abstract
MOTIVATION: Cell-type deconvolution is a computational approach to infer cellular distributions from bulk transcriptomics data. Several methods have been proposed, each with its own advantages and disadvantages. Reference based approaches make use of archetypic transcriptomic profiles representing individual cell types. Those reference profiles are ideally chosen such that the observed bulks can be reconstructed as a linear combination thereof. This strategy, however, ignores the fact that cellular populations arise through the process of cellular differentiation, which entails the gradual emergence of cell groups with diverse morphological and functional characteristics. RESULTS: Here, we propose Hierarchical cell-type Deconvolution (HIDE), a cell-type deconvolution approach which incorporates a cell hierarchy for improved performance and interpretability. This is achieved by a hierarchical procedure that preserves estimates of major cell populations while inferring their respective subpopulations. We show in simulation studies that this procedure produces more reliable and more consistent results than other state-of-the-art approaches. Finally, we provide an example application of HIDE to explore breast cancer specimens from TCGA. AVAILABILITY AND IMPLEMENTATION: A python implementation of HIDE is available at zenodo (doi: 10.5281/zenodo.14724906).
Dennis Völkl, Malte Mensching-Buhr, Thomas Sterr, Sarah Bolz, Andreas Schäfer 0005, Nicole Seifert, Jana Tauschke, Austin Rayford, Oddbjørn Straume, Helena U. Zacharias, Sushma Nagaraja Grellscheid, Tim Beißbarth, Michael Altenbuchinger, Franziska Görtler
Bioinform.10
2024 SpaCeNet: Spatial Cellular Networks from Omics Data
Stefan Schrod, Niklas Lück, Robert Lohmayer, Stefan Solbrig, Tina Wipfler, Katherine H. Shutta, Marouen Ben Guebila, Andreas Schäfer 0005, Tim Beißbarth, Helena U. Zacharias, Peter J. Oefner, John Quackenbush, Michael Altenbuchinger
RECOMB10
2024 Adaptive digital tissue deconvolution
abstract
MOTIVATION: The inference of cellular compositions from bulk and spatial transcriptomics data increasingly complements data analyses. Multiple computational approaches were suggested and recently, machine learning techniques were developed to systematically improve estimates. Such approaches allow to infer additional, less abundant cell types. However, they rely on training data which do not capture the full biological diversity encountered in transcriptomics analyses; data can contain cellular contributions not seen in the training data and as such, analyses can be biased or blurred. Thus, computational approaches have to deal with unknown, hidden contributions. Moreover, most methods are based on cellular archetypes which serve as a reference; e.g. a generic T-cell profile is used to infer the proportion of T-cells. It is well known that cells adapt their molecular phenotype to the environment and that pre-specified cell archetypes can distort the inference of cellular compositions. RESULTS: We propose Adaptive Digital Tissue Deconvolution (ADTD) to estimate cellular proportions of pre-selected cell types together with possibly unknown and hidden background contributions. Moreover, ADTD adapts prototypic reference profiles to the molecular environment of the cells, which further resolves cell-type specific gene regulation from bulk transcriptomics data. We verify this in simulation studies and demonstrate that ADTD improves existing approaches in estimating cellular compositions. In an application to bulk transcriptomics data from breast cancer patients, we demonstrate that ADTD provides insights into cell-type specific molecular differences between breast cancer subtypes. AVAILABILITY AND IMPLEMENTATION: A python implementation of ADTD and a tutorial are available at Gitlab and zenodo (doi:10.5281/zenodo.7548362).
Franziska Görtler, Malte Mensching-Buhr, Ørjan Skaar, Stefan Schrod, Thomas Sterr, Andreas Schäfer 0005, Tim Beißbarth, Anagha Joshi, Helena U. Zacharias, Sushma Nagaraja Grellscheid, Michael Altenbuchinger
Bioinform.9
2024 CODEX: COunterfactual Deep learning for the in silico EXploration of cancer cell line perturbations
abstract
MOTIVATION: High-throughput screens (HTS) provide a powerful tool to decipher the causal effects of chemical and genetic perturbations on cancer cell lines. Their ability to evaluate a wide spectrum of interventions, from single drugs to intricate drug combinations and CRISPR-interference, has established them as an invaluable resource for the development of novel therapeutic approaches. Nevertheless, the combinatorial complexity of potential interventions makes a comprehensive exploration intractable. Hence, prioritizing interventions for further experimental investigation becomes of utmost importance. RESULTS: We propose CODEX (COunterfactual Deep learning for the in silico EXploration of cancer cell line perturbations) as a general framework for the causal modeling of HTS data, linking perturbations to their downstream consequences. CODEX relies on a stringent causal modeling strategy based on counterfactual reasoning. As such, CODEX predicts drug-specific cellular responses, comprising cell survival and molecular alterations, and facilitates the in silico exploration of drug combinations. This is achieved for both bulk and single-cell HTS. We further show that CODEX provides a rationale to explore complex genetic modifications from CRISPR-interference in silico in single cells. AVAILABILITY AND IMPLEMENTATION: Our implementation of CODEX is publicly available at https://github.com/sschrod/CODEX. All data used in this article are publicly available.
Stefan Schrod, Helena U. Zacharias, Tim Beißbarth, Anne-Christin Hauschild, Michael Altenbuchinger
Bioinform.2
2023 Anomaly detection in mixed high-dimensional molecular data
abstract
MOTIVATION: Mixed molecular data combines continuous and categorical features of the same samples, such as OMICS profiles with genotypes, diagnoses, or patient sex. Like all high-dimensional molecular data, it is prone to incorrect values that can stem from various sources for example the technical limitations of the measurement devices, errors in the sample preparation, or contamination. Most anomaly detection algorithms identify complete samples as outliers or anomalies. However, in most cases, not all measurements of those samples are erroneous but only a few one-dimensional features within the samples are incorrect. These one-dimensional data errors are continuous measurements that are either located outside or inside the normal ranges of their features but in both cases show atypical values given all other continuous and categorical features in the sample. Additionally, categorical anomalies can occur for example when the genotype or diagnosis was submitted wrongly. RESULTS: We introduce ADMIRE (Anomaly Detection using MIxed gRaphical modEls), a novel approach for the detection and correction of anomalies in mixed high-dimensional data. Hereby, we focus on the detection of single (one-dimensional) data errors in the categorical and continuous features of a sample. For that the joint distribution of continuous and categorical features is learned by mixed graphical models, anomalies are detected by the difference between measured and model-based estimations and are corrected using imputation. We evaluated ADMIRE in simulation and by screening for anomalies in one of our own metabolic datasets. In simulation experiments, ADMIRE outperformed the state-of-the-art methods of Local Outlier Factor, stray, and Isolation Forest. AVAILABILITY AND IMPLEMENTATION: All data and code is available at https://github.com/spang-lab/adadmire. ADMIRE is implemented in a Python package called adadmire which can be found at https://pypi.org/project/adadmire.
Lena Buck, Maren Feist, Philipp Schwarzfischer, Dieter Kube, Peter J. Oefner, Helena U. Zacharias, Michael Altenbuchinger, Katja Dettmer, Wolfram Gronwald, Rainer Spang
Bioinform.7
2022 BITES: balanced individual treatment effect for survival data
abstract
MOTIVATION: Estimating the effects of interventions on patient outcome is one of the key aspects of personalized medicine. Their inference is often challenged by the fact that the training data comprises only the outcome for the administered treatment, and not for alternative treatments (the so-called counterfactual outcomes). Several methods were suggested for this scenario based on observational data, i.e. data where the intervention was not applied randomly, for both continuous and binary outcome variables. However, patient outcome is often recorded in terms of time-to-event data, comprising right-censored event times if an event does not occur within the observation period. Albeit their enormous importance, time-to-event data are rarely used for treatment optimization. We suggest an approach named BITES (Balanced Individual Treatment Effect for Survival data), which combines a treatment-specific semi-parametric Cox loss with a treatment-balanced deep neural network; i.e. we regularize differences between treated and non-treated patients using Integral Probability Metrics (IPM). RESULTS: We show in simulation studies that this approach outperforms the state of the art. Furthermore, we demonstrate in an application to a cohort of breast cancer patients that hormone treatment can be optimized based on six routine parameters. We successfully validated this finding in an independent cohort. AVAILABILITY AND IMPLEMENTATION: We provide BITES as an easy-to-use python implementation including scheduled hyper-parameter optimization (https://github.com/sschrod/BITES). The data underlying this article are available in the CRAN repository at https://rdrr.io/cran/survival/man/gbsg.html and https://rdrr.io/cran/survival/man/rotterdam.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Stefan Schrod, Andreas Schäfer 0005, Stefan Solbrig, Robert Lohmayer, Wolfram Gronwald, Peter J. Oefner, Tim Beißbarth, Rainer Spang, Helena U. Zacharias, Michael Altenbuchinger
Bioinform.9
2017 Reference point insensitive molecular data analysis
abstract
MOTIVATION: In biomedicine, every molecular measurement is relative to a reference point, like a fixed aliquot of RNA extracted from a tissue, a defined number of blood cells, or a defined volume of biofluid. Reference points are often chosen for practical reasons. For example, we might want to assess the metabolome of a diseased organ but can only measure metabolites in blood or urine. In this case, the observable data only indirectly reflects the disease state. The statistical implications of these discrepancies in reference points have not yet been discussed. RESULTS: Here, we show that reference point discrepancies compromise the performance of regression models like the LASSO. As an alternative, we suggest zero-sum regression for a reference point insensitive analysis. We show that zero-sum regression is superior to the LASSO in case of a poor choice of reference point both in simulations and in an application that integrates intestinal microbiome analysis with metabolomics. Moreover, we describe a novel coordinate descent based algorithm to fit zero-sum elastic nets. AVAILABILITY AND IMPLEMENTATION: The R-package "zeroSum" can be downloaded at https://github.com/rehbergT/zeroSum Moreover, we provide all R-scripts and data used to produce the results of this manuscript as Supplementary Material CONTACT: [email protected], [email protected] and [email protected] information: Supplementary material is available at Bioinformatics online.
Michael Altenbuchinger, Thorsten Rehberg, Helena U. Zacharias, Frank Stämmler, K. Dettmer, D. Weber, Andreas Hiergeist, A. Gessner, E. Holler, Peter J. Oefner, Rainer Spang
Bioinform.3