Andre Franke

dblp:10/7236 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-1530-5811ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 4 since 2021
YearPublicationVenuePosition
2025 Splicing-aware scRNA-Seq resolution reveals execution-ready programs in effector Tregs
abstract
Single-cell RNA sequencing (scRNA-Seq) provides valuable insights into cell biology. However, current scRNA-Seq analytic approaches do not distinguish between spliced and unspliced mRNA at the level of dimensionality reduction. RNA velocity paradigm suggests that the presence of unspliced mRNA reflects transitional cell states, informative for studies of dynamic processes such as embryogenesis or tissue regeneration. Alternatively, stable cell subsets may also maintain translationally repressed spliced mRNA (e.g., in P-bodies) and/or unspliced mRNA reservoirs for prompt initiation of transcription-independent expression. Thus, functional cell subsets may differ not only in the current levels of actively produced mRNAs, but also in which mRNAs and in what forms are stored in the nucleus and cytoplasm. To enable splicing-aware analysis of scRNA-Seq data, we developed a method called SANSARA (Splicing-Aware scrNa-Seq AppRoAch). We employed SANSARA to characterize peripheral blood regulatory T cell (Treg) subsets, revealing a complementary interplay between the FOXP3 and Helios master transcription factors and high levels of spliced IL10RA, LGALS3, FCRL3, CD38, ITGAL, and LEF1 mRNAs in effector Tregs. Among Th1 and cytotoxic CD4+ T cell subsets, SANSARA also revealed substantial splicing heterogeneity across subset-specific genes. SANSARA is straightforward to implement in current data analysis pipelines and opens new dimensions for scRNA-Seq-based discoveries.
Daniil K. Lukyanov, Evgeniy S. Egorov, Valeriia V. Kriukova, Denis Syrko, Victor V. Kotliar, Kristin Ladell, David A. Price, Andre Franke, Dmitry Chudakov
PLoS Comput. Biol.8
2023 Pathogen detection in RNA-seq data with Pathonoia
abstract
BACKGROUND: Bacterial and viral infections may cause or exacerbate various human diseases and to detect microbes in tissue, one method of choice is RNA sequencing. The detection of specific microbes using RNA sequencing offers good sensitivity and specificity, but untargeted approaches suffer from high false positive rates and a lack of sensitivity for lowly abundant organisms. RESULTS: We introduce Pathonoia, an algorithm that detects viruses and bacteria in RNA sequencing data with high precision and recall. Pathonoia first applies an established k-mer based method for species identification and then aggregates this evidence over all reads in a sample. In addition, we provide an easy-to-use analysis framework that highlights potential microbe-host interactions by correlating the microbial to the host gene expression. Pathonoia outperforms state-of-the-art methods in microbial detection specificity, both on in silico and real datasets. CONCLUSION: Two case studies in human liver and brain show how Pathonoia can support novel hypotheses on microbial infection exacerbating disease. The Python package for Pathonoia sample analysis and a guided analysis Jupyter notebook for bulk RNAseq datasets are available on GitHub.
Anna-Maria Liebhoff, Kevin Menden, Alena Laschtowitz, Andre Franke, Christoph Schramm, Stefan Bonn
BMC Bioinform.4
2022 <tt>MAGScoT</tt>: a fast, lightweight and accurate bin-refinement tool
abstract
MOTIVATION: Recovery of metagenome-assembled genomes (MAGs) from shotgun metagenomic data is an important task for the comprehensive analysis of microbial communities from variable sources. Single binning tools differ in their ability to leverage specific aspects in MAG reconstruction, the use of ensemble binning refinement tools is often time consuming and computational demand increases with community complexity. We introduce MAGScoT, a fast, lightweight and accurate implementation for the reconstruction of highest-quality MAGs from the output of multiple genome-binning tools. RESULTS: MAGScoT outperforms popular bin-refinement solutions in terms of quality and quantity of MAGs as well as computation time and resource consumption. AVAILABILITY AND IMPLEMENTATION: MAGScoT is available via GitHub (https://github.com/ikmb/MAGScoT) and as an easy-to-use Docker container (https://hub.docker.com/repository/docker/ikmb/magscot). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Malte Rühlemann, Eike Matthias Wacker, David Ellinghaus, Andre Franke
Bioinform.4
2021 Immunopeptidomics toolkit library (IPTK): a python-based modular toolbox for analyzing immunopeptidomics data
abstract
BACKGROUND: The human leukocyte antigen (HLA) proteins play a fundamental role in the adaptive immune system as they present peptides to T cells. Mass-spectrometry-based immunopeptidomics is a promising and powerful tool for characterizing the immunopeptidomic landscape of HLA proteins, that is the peptides presented on HLA proteins. Despite the growing interest in the technology, and the recent rise of immunopeptidomics-specific identification pipelines, there is still a gap in data-analysis and software tools that are specialized in analyzing and visualizing immunopeptidomics data. RESULTS: We present the IPTK library which is an open-source Python-based library for analyzing, visualizing, comparing, and integrating different omics layers with the identified peptides for an in-depth characterization of the immunopeptidome. Using different datasets, we illustrate the ability of the library to enrich the result of the identified peptidomes. Also, we demonstrate the utility of the library in developing other software and tools by developing an easy-to-use dashboard that can be used for the interactive analysis of the results. CONCLUSION: IPTK provides a modular and extendable framework for analyzing and integrating immunopeptidomes with different omics layers. The library is deployed into PyPI at https://pypi.org/project/IPTKL/ and into Bioconda at https://anaconda.org/bioconda/iptkl , while the source code of the library and the dashboard, along with the online tutorials are available at https://github.com/ikmb/iptoolkit .
Hesham ElAbd, Frauke Degenhardt, Tomás Koudelka, Ann-Kristin Kamps, Andreas Tholey, Petra Bacher, Tobias L. Lenz, Andre Franke, Mareike Wendorff
BMC Bioinform.8
2020 Amino acid encoding for deep learning applications
abstract
BACKGROUND: The number of applications of deep learning algorithms in bioinformatics is increasing as they usually achieve superior performance over classical approaches, especially, when bigger training datasets are available. In deep learning applications, discrete data, e.g. words or n-grams in language, or amino acids or nucleotides in bioinformatics, are generally represented as a continuous vector through an embedding matrix. Recently, learning this embedding matrix directly from the data as part of the continuous iteration of the model to optimize the target prediction - a process called 'end-to-end learning' - has led to state-of-the-art results in many fields. Although usage of embeddings is well described in the bioinformatics literature, the potential of end-to-end learning for single amino acids, as compared to more classical manually-curated encoding strategies, has not been systematically addressed. To this end, we compared classical encoding matrices, namely one-hot, VHSE8 and BLOSUM62, to end-to-end learning of amino acid embeddings for two different prediction tasks using three widely used architectures, namely recurrent neural networks (RNN), convolutional neural networks (CNN), and the hybrid CNN-RNN. RESULTS: By using different deep learning architectures, we show that end-to-end learning is on par with classical encodings for embeddings of the same dimension even when limited training data is available, and might allow for a reduction in the embedding dimension without performance loss, which is critical when deploying the models to devices with limited computational capacities. We found that the embedding dimension is a major factor in controlling the model performance. Surprisingly, we observed that deep learning models are capable of learning from random vectors of appropriate dimension. CONCLUSION: Our study shows that end-to-end learning is a flexible and powerful method for amino acid encoding. Further, due to the flexibility of deep learning systems, amino acid encoding schemes should be benchmarked against random vectors of the same dimension to disentangle the information content provided by the encoding scheme from the distinguishability effect provided by the scheme.
Hesham ElAbd, Yana Bromberg, Adrienne Hoarfrost, Tobias Lenz, Andre Franke, Mareike Wendorff
BMC Bioinform.5
2018 Comparing genome versus proteome-based identification of clinical bacterial isolates
abstract
Whole-genome sequencing (WGS) is gaining importance in the analysis of bacterial cultures derived from patients with infectious diseases. Existing computational tools for WGS-based identification have, however, been evaluated on previously defined data relying thereby unwarily on the available taxonomic information.Here, we newly sequenced 846 clinical gram-negative bacterial isolates representing multiple distinct genera and compared the performance of five tools (CLARK, Kaiju, Kraken, DIAMOND/MEGAN and TUIT). To establish a faithful 'gold standard', the expert-driven taxonomy was compared with identifications based on matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry (MS) analysis. Additionally, the tools were also evaluated using a data set of 200 Staphylococcus aureus isolates.CLARK and Kraken (with k =31) performed best with 626 (100%) and 193 (99.5%) correct species classifications for the gram-negative and S. aureus isolates, respectively. Moreover, CLARK and Kraken demonstrated highest mean F-measure values (85.5/87.9% and 94.4/94.7% for the two data sets, respectively) in comparison with DIAMOND/MEGAN (71 and 85.3%), Kaiju (41.8 and 18.9%) and TUIT (34.5 and 86.5%). Finally, CLARK, Kaiju and Kraken outperformed the other tools by a factor of 30 to 170 fold in terms of runtime.We conclude that the application of nucleotide-based tools using k-mers-e.g. CLARK or Kraken-allows for accurate and fast taxonomic characterization of bacterial isolates from WGS data. Hence, our results suggest WGS-based genotyping to be a promising alternative to the MS-based biotyping in clinical settings. Moreover, we suggest that complementary information should be used for the evaluation of taxonomic classification tools, as public databases may suffer from suboptimal annotations.
Valentina Galata, Christina Backes, Cedric Christian Laczny, Georg Hemmrich-Stanisak, Howard Li, Laura Smoot, Andreas E. Posch, Susanne Schmolke, Markus Bischoff, Lutz von Müller, Achim Plum, Andre Franke, Andreas Keller
Briefings Bioinform.12
2018 A high-resolution map of the human small non-coding transcriptome
abstract
Motivation: Although the amount of small non-coding RNA-sequencing data is continuously increasing, it is still unclear to which extent small RNAs are represented in the human genome. Results: In this study we analyzed 303 billion sequencing reads from nearly 25 000 datasets to answer this question. We determined that 0.8% of the human genome are reliably covered by 874 123 regions with an average length of 31 nt. On the basis of these regions, we found that among the known small non-coding RNA classes, microRNAs were the most prevalent. In subsequent steps, we characterized variations of miRNAs and performed a staged validation of 11 877 candidate miRNAs. Of these, many were actually expressed and significantly dysregulated in lung cancer. Selected candidates were finally validated by northern blots. Although isolated miRNAs could still be present in the human genome, our presented set likely contains the largest fraction of human miRNAs. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Tobias Fehlmann, Christina Backes, Julia Alles, Ulrike Fischer, Martin Hart, Fabian Kern, Hilde Langseth, Trine Rounge, Sinan U. Umu, Mustafa Kahraman, Thomas Laufer, Jan Haas, Cord Stähler, Nicole Ludwig 0001, Matthias Hübenthal, Benjamin Meder, Andre Franke, Hans-Peter Lenhof, Eckart Meese, Andreas Keller
Bioinform.17
2017 Interdisciplinary approach towards a systems medicine toolbox using the example of inflammatory diseases
abstract
Electronic access to multiple data types, from generic information on biological systems at different functional and cellular levels to high-throughput molecular data from human patients, is a prerequisite of successful systems medicine research. However, scientists often encounter technical and conceptual difficulties that forestall the efficient and effective use of these resources. We summarize and discuss some of these obstacles, and suggest ways to avoid or evade them.The methodological gap between data capturing and data analysis is huge in human medical research. Primary data producers often do not fully apprehend the scientific value of their data, whereas data analysts maybe ignorant of the circumstances under which the data were collected. Therefore, the provision of easy-to-use data access tools not only helps to improve data quality on the part of the data producers but also is likely to foster an informed dialogue with the data analysts.We propose a means to integrate phenotypic data, questionnaire data and microbiome data with a user-friendly Systems Medicine toolbox embedded into i2b2/tranSMART. Our approach is exemplified by the integration of a basic outlier detection tool and a more advanced microbiome analysis (alpha diversity) script. Continuous discussion with clinicians, data managers, biostatisticians and systems medicine experts should serve to enrich even further the functionality of toolboxes like ours, being geared to be used by 'informed non-experts' but at the same time attuned to existing, more sophisticated analysis tools.
Christian R. Bauer, Carolin Knecht, Christoph Fretter, Benjamin Baum, Sandra Jendrossek, Malte Rühlemann, Femke-Anouska Heinsen, Nadine Umbach, Bodo Grimbacher, Andre Franke, Wolfgang Lieb, Michael Krawczak, Marc-Thorsten Hütt, Ulrich Sax
Briefings Bioinform.10
2017 Boolean analysis reveals systematic interactions among low-abundance species in the human gut microbiome
abstract
The analysis of microbiome compositions in the human gut has gained increasing interest due to the broader availability of data and functional databases and substantial progress in data analysis methods, but also due to the high relevance of the microbiome in human health and disease. While most analyses infer interactions among highly abundant species, the large number of low-abundance species has received less attention. Here we present a novel analysis method based on Boolean operations applied to microbial co-occurrence patterns. We calibrate our approach with simulated data based on a dynamical Boolean network model from which we interpret the statistics of attractor states as a theoretical proxy for microbiome composition. We show that for given fractions of synergistic and competitive interactions in the model our Boolean abundance analysis can reliably detect these interactions. Analyzing a novel data set of 822 microbiome compositions of the human gut, we find a large number of highly significant synergistic interactions among these low-abundance species, forming a connected network, and a few isolated competitive interactions.
Jens Christian Claussen, Jurgita Skieceviciene, Jun Wang 0074, Philipp Rausch, Tom H. Karlsen, Wolfgang Lieb, John F. Baines, Andre Franke, Marc-Thorsten Hütt
PLoS Comput. Biol.8
2016 Haplotype synthesis analysis reveals functional variants underlying known genome-wide associated susceptibility loci
abstract
MOTIVATION: The functional mechanisms underlying disease association remain unknown for Genome-wide Association Studies (GWAS) susceptibility variants located outside coding regions. Synthesis of effects from multiple surrounding functional variants has been suggested as an explanation of hard-to-interpret findings. We define filter criteria based on linkage disequilibrium measures and allele frequencies which reflect expected properties of synthesizing variant sets. For eligible candidate sets, we search for haplotype markers that are highly correlated with associated variants. RESULTS: Via simulations we assess the performance of our approach and suggest parameter settings which guarantee 95% sensitivity at 20-fold reduced computational cost. We apply our method to 1000 Genomes data and confirmed Crohn's Disease (CD) and Type 2 Diabetes (T2D) variants. A proportion of 36.9% allowed explanation by three-variant-haplotypes carrying at least two functional variants, as compared to 16.4% for random variants ([Formula: see text]). Association could be explained by missense variants for MUC19, PER3 (CD) and HMG20A (T2D). In a CD GWAS-imputed using haplotype reference consortium data (64 976 haplotypes)-we could confirm the syntheses of MUC19 and PER3 and identified synthesis by missense variants for 6 further genes (ZGPAZ, GPR65, CLN3/NPIPB8, LOC102723878, rs2872507, GCKR). In all instances, the odds ratios of the synthesizing haplotypes were virtually identical to that of the index SNP. In summary, we demonstrate the potential of synthesis analysis to guide functional follow-up of GWAS findings. AVAILABILITY AND IMPLEMENTATION: All methods are implemented in the C/C ++ toolkit GetSynth, available at http://sourceforge.net/projects/getsynth/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
André Lacour, David Ellinghaus, Stefan Schreiber, Andre Franke, Tim Becker
Bioinform.4
2012 B-SOLANA: an approach for the analysis of two-base encoding bisulfite sequencing data
abstract
SUMMARY: Bisulfite sequencing, a combination of bisulfite treatment and high-throughput sequencing, has proved to be a valuable method for measuring DNA methylation at single base resolution. Here, we present B-SOLANA, an approach for the analysis of two-base encoding (colorspace) bisulfite sequencing data on the SOLiD platform of Life Technologies. It includes the alignment of bisulfite sequences and the determination of methylation levels in CpG as well as non-CpG sequence contexts. B-SOLANA enables a fast and accurate analysis of large raw sequence datasets. AVAILABILITY AND IMPLEMENTATION: The source code, released under the GNU GPLv3 licence, is freely available at http://code.google.com/p/bsolana/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Benjamin Kreck, George Marnellos, Julia Richter, Felix Krueger, Reiner Siebert, Andre Franke
Bioinform.6
2010 CNVineta: a data mining tool for large case-control copy number variation datasets
abstract
MOTIVATION: Copy number variation (CNV), a major contributor to human genetic variation, comprises >/= 1 kb genomic deletions and insertions. Yet, the identification of CNVs from microarray data is still hampered by high false negative and positive prediction rates due to the noisy nature of the raw data. Here, we present CNVineta, an R package for rapid data mining and visualization of CNVs in large case-control datasets genotyped with single nucleotide polymorphism oligonucleotide arrays. CNVineta is compatible with various established CNV prediction algorithms, can be used for genome-wide association analysis of rare and common CNVs and enables rapid and serial display of log(2) of raw data ratios as well as B-allele frequencies for visual quality inspection. In summary, CNVineta aides in the interpretation of large-scale CNV datasets and prioritization of target regions for follow-up experiments. AVAILABILITY AND IMPLEMENTATION: CNVineta is available as an R package and can be downloaded from http://www.ikmb.uni-kiel.de/CNVineta/; the package contains a tutorial outlining a typical workflow. The CNVineta compatible HapMap dataset can also be downloaded from the link above.
Michael Wittig, Ingo Helbig, Stefan Schreiber, Andre Franke
Bioinform.4
2010 SNPexp - A web tool for calculating and visualizing correlation between HapMap genotypes and gene expression levels
abstract
BACKGROUND: Expression levels for 47294 transcripts in lymphoblastoid cell lines from all 270 HapMap phase II individuals, and genotypes (both HapMap phase II and III) of 3.96 million single nucleotide polymorphisms (SNPs) in the same individuals are publicly available. We aimed to generate a user-friendly web based tool for visualization of the correlation between SNP genotypes within a specified genomic region and a gene of interest, which is also well-known as an expression quantitative trait locus (eQTL) analysis. RESULTS: SNPexp is implemented as a server-side script, and publicly available on this website: http://tinyurl.com/snpexp. Correlation between genotype and transcript expression levels are calculated by performing linear regression and the Wald test as implemented in PLINK and visualized using the UCSC Genome Browser. Validation of SNPexp using previously published eQTLs yielded comparable results. CONCLUSIONS: SNPexp provides a convenient and platform-independent way to calculate and visualize the correlation between HapMap genotypes within a specified genetic region anywhere in the genome and gene expression levels. This allows for investigation of both cis and trans effects. The web interface and utilization of publicly available and widely used software resources makes it an attractive supplement to more advanced bioinformatic tools. For the advanced user the program can be used on a local computer on custom datasets.
Kristian Holm, Espen Melum, Andre Franke, Tom H. Karlsen
BMC Bioinform.3
2009 GMFilter and SXTestPlate: software tools for improving the SNPlexTM genotyping system
abstract
BACKGROUND: Genotyping of single-nucleotide polymorphisms (SNPs) is a fundamental technology in modern genetics. The SNPlex mid-throughput genotyping system (Applied Biosystems, Foster City, CA, USA) enables the multiplexed genotyping of up to 48 SNPs simultaneously in a single DNA sample. The high level of automation and the large amount of data produced in a high-throughput laboratory require advanced software tools for quality control and workflow management. RESULTS: We have developed two programs, which address two main aspects of quality control in a SNPlex genotyping environment: GMFilter improves the analysis of SNPlex plates by removing wells with a low overall signal intensity. It enables scientists to automatically process the raw data in a standardized way before analyzing a plate with the proprietary GeneMapper software from Applied Biosystems. SXTestPlate examines the genotype concordance of a SNPlex test plate, which was typed with a control SNP set. This program allows for regular quality control checks of a SNPlex genotyping platform. It is compatible to other genotyping methods as well. CONCLUSION: GMFilter and SXTestPlate provide a valuable tool set for laboratories engaged in genotyping based on the SNPlex system. The programs enhance the analysis of SNPlex plates with the GeneMapper software and enable scientists to evaluate the performance of their genotyping platform.
Markus Teuber, Michael H. Wenz, Stefan Schreiber, Andre Franke
BMC Bioinform.4