Philipp Ehmele

dblp:290/1890 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0001-5945-7839ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 100%

Topics — the 3 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › comparative genomics › pangenomics
pangenome graph
0.812024
Cluster-efficient pangenome graph construction with nf-core/pangenome · Bioinform. 2024
High-performance computing › cluster computing
cluster deployment
0.212024
Cluster-efficient pangenome graph construction with nf-core/pangenome · Bioinform. 2024
High-performance computing
scientific workflow
0.212024
Cluster-efficient pangenome graph construction with nf-core/pangenome · Bioinform. 2024

Methods — techniques the papers use, named apart from their topics

nextflow · 1.5biocontainers · 1.5
YearPublicationVenuePosition
2024 Cluster-efficient pangenome graph construction with nf-core/pangenome
abstract
MOTIVATION: Pangenome graphs offer a comprehensive way of capturing genomic variability across multiple genomes. However, current construction methods often introduce biases, excluding complex sequences or relying on references. The PanGenome Graph Builder (PGGB) addresses these issues. To date, though, there is no state-of-the-art pipeline allowing for easy deployment, efficient and dynamic use of available resources, and scalable usage at the same time. RESULTS: To overcome these limitations, we present nf-core/pangenome, a reference-unbiased approach implemented in Nextflow following nf-core's best practices. Leveraging biocontainers ensures portability and seamless deployment in High-Performance Computing (HPC) environments. Unlike PGGB, nf-core/pangenome distributes alignments across cluster nodes, enabling scalability. Demonstrating its efficiency, we constructed pangenome graphs for 1000 human chromosome 19 haplotypes and 2146 Escherichia coli sequences, achieving a two to threefold speedup compared to PGGB without increasing greenhouse gas emissions. AVAILABILITY AND IMPLEMENTATION: nf-core/pangenome is released under the MIT open-source license, available on GitHub and Zenodo, with documentation accessible at https://nf-co.re/pangenome/docs/usage.
Simon Heumos, Michael L. Heuer, Friederike Hanssen, Lukas Heumos, Andrea Guarracino, Peter Heringer, Philipp Ehmele, Pjotr Prins, Erik Garrison, Sven Nahnsen
Bioinform.7
2023 mlf-core: a framework for deterministic machine learning
abstract
MOTIVATION: Machine learning has shown extensive growth in recent years and is now routinely applied to sensitive areas. To allow appropriate verification of predictive models before deployment, models must be deterministic. Solely fixing all random seeds is not sufficient for deterministic machine learning, as major machine learning libraries default to the usage of nondeterministic algorithms based on atomic operations. RESULTS: Various machine learning libraries released deterministic counterparts to the nondeterministic algorithms. We evaluated the effect of these algorithms on determinism and runtime. Based on these results, we formulated a set of requirements for deterministic machine learning and developed a new software solution, the mlf-core ecosystem, which aids machine learning projects to meet and keep these requirements. We applied mlf-core to develop deterministic models in various biomedical fields including a single-cell autoencoder with TensorFlow, a PyTorch-based U-Net model for liver-tumor segmentation in computed tomography scans, and a liver cancer classifier based on gene expression profiles with XGBoost. AVAILABILITY AND IMPLEMENTATION: The complete data together with the implementations of the mlf-core ecosystem and use case models are available at https://github.com/mlf-core.
Lukas Heumos, Philipp Ehmele, Luis Kuhn Cuellar, Kevin Menden, Edmund Miller, Steffen Lemke 0003, Gisela Gabernet, Sven Nahnsen
Bioinform.2