VLDB 2026 Research / reviewers in the wild / expert
Daniel J. Blankenberg
dblp:39/7105
· DBLP profile ↗
10ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-6833-9049ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancerabstractThe continuously decreasing cost of next-generation sequencing has recently led to a significant increase in the number of microbiome-related studies, providing invaluable information for understanding host-microbiome interactions and their relation to diseases. A common approach in metagenomics consists of determining the composition of samples in terms of the amount and types of microbial species that populate them, with the goal of identifying microbes whose profiles are able to differentiate samples under different conditions with advanced feature selection techniques. Here, we propose a novel backward variable selection method based on the hyperdimensional computing (HDC) paradigm, which takes inspiration from how the human brain works in the classification of concepts by encoding features into vectors in a high-dimensional space. We validated our method on public metagenomic samples collected from patients affected by colorectal cancer in a case/control scenario, by performing a comparative analysis with other state-of-the-art feature selection methods, obtaining promising results. AUTHOR SUMMARY: Characterizing the microbial composition of metagenomic samples is crucial for identifying potential biomarkers that can distinguish between healthy and diseased states. However, the high dimensionality and complexity of metagenomic data present significant challenges in the context of accurately selecting features. Our backward variable selection method, based on the HDC paradigm, offers a promising approach to overcoming these challenges. By effectively reducing the feature space while preserving essential information, this method enhances the ability to detect critical microbial signatures associated with diseases like colorectal cancer, leading to more precise diagnostic tools. Fabio Cumbo, Simone Truglia, Emanuel Weitschek, Daniel J. Blankenberg |
Briefings Bioinform. | 4 |
| 2023 | Evaluating homophily in networks via HONTO (HOmophily Network TOol): a case study of chromosomal interactions in human PPI networksabstractSUMMARY: It has been observed in different kinds of networks, such as social or biological ones, a typical behavior inspired by the general principle 'similarity breeds connections'. These networks are defined as homophilic as nodes belonging to the same class preferentially interact with each other. In this work, we present HONTO (HOmophily Network TOol), a user-friendly open-source Python3 package designed to evaluate and analyze homophily in complex networks. The tool takes in input from the network along with a partition of its nodes into classes and yields a matrix whose entries are the homophily/heterophily z-score values. To complement the analysis, the tool also provides z-score values of nodes that do not interact with any other node of the same class. Homophily/heterophily z-scores values are presented as a heatmap allowing a visual at-a-glance interpretation of results. AVAILABILITY AND IMPLEMENTATION: Tool's source code is available at https://github.com/cumbof/honto under the MIT license, installable as a package from PyPI (pip install honto) and conda-forge (conda install -c conda-forge honto), and has a wrapper for the Galaxy platform available on the official Galaxy ToolShed (Blankenberg et al., 2014) at https://toolshed.g2.bx.psu.edu/view/fabio/honto. Nicola Apollonio, Daniel J. Blankenberg, Fabio Cumbo, Paolo Giulio Franciosa, Daniele Santoni |
Bioinform. | 2 |
| 2023 | Galaxy Training: A powerful framework for teaching!abstractThere is an ongoing explosion of scientific datasets being generated, brought on by recent technological advances in many areas of the natural sciences. As a result, the life sciences have become increasingly computational in nature, and bioinformatics has taken on a central role in research studies. However, basic computational skills, data analysis, and stewardship are still rarely taught in life science educational programs, resulting in a skills gap in many of the researchers tasked with analysing these big datasets. In order to address this skills gap and empower researchers to perform their own data analyses, the Galaxy Training Network (GTN) has previously developed the Galaxy Training Platform (https://training.galaxyproject.org), an open access, community-driven framework for the collection of FAIR (Findable, Accessible, Interoperable, Reusable) training materials for data analysis utilizing the user-friendly Galaxy framework as its primary data analysis platform. Since its inception, this training platform has thrived, with the number of tutorials and contributors growing rapidly, and the range of topics extending beyond life sciences to include topics such as climatology, cheminformatics, and machine learning. While initially aimed at supporting researchers directly, the GTN framework has proven to be an invaluable resource for educators as well. We have focused our efforts in recent years on adding increased support for this growing community of instructors. New features have been added to facilitate the use of the materials in a classroom setting, simplifying the contribution flow for new materials, and have added a set of train-the-trainer lessons. Here, we present the latest developments in the GTN project, aimed at facilitating the use of the Galaxy Training materials by educators, and its usage in different learning environments. Saskia D. Hiltemann, Helena Rasche, Simon L. Gladman, Hans-Rudolf Hotz, Delphine Larivière, Daniel J. Blankenberg, Pratik D. Jagtap, Thomas Wollmann, Anthony Bretaudeau, Nadia Goué, Timothy J. Griffin, Coline Royaux, Yvan Le Bras, Subina P. Mehta, Anna Syme, Frederik Coppens, Bert Droesbeke, Nicola Soranzo, Wendi Bacon, Fotis E. Psomopoulos, Cristóbal Gallardo-Alba, Melanie Christine Föll, Matthias Fahrner, Maria A. Doyle, Beatriz Serrano-Solano, Anne Fouilloux, Peter van Heusden, Wolfgang Maier 0003, Dave Clements, Florian Heyl, Björn A. Grüning, Bérénice Batut |
PLoS Comput. Biol. | 6 |
| 2022 | PDAUG: a Galaxy based toolset for peptide library analysis, visualization, and machine learning modelingabstractBACKGROUND: Computational methods based on initial screening and prediction of peptides for desired functions have proven to be effective alternatives to lengthy and expensive biochemical experimental methods traditionally utilized in peptide research, thus saving time and effort. However, for many researchers, the lack of expertise in utilizing programming libraries, access to computational resources, and flexible pipelines are big hurdles to adopting these advanced methods. RESULTS: To address the above mentioned barriers, we have implemented the peptide design and analysis under Galaxy (PDAUG) package, a Galaxy-based Python powered collection of tools, workflows, and datasets for rapid in-silico peptide library analysis. In contrast to existing methods like standard programming libraries or rigid single-function web-based tools, PDAUG offers an integrated GUI-based toolset, providing flexibility to build and distribute reproducible pipelines and workflows without programming expertise. Finally, we demonstrate the usability of PDAUG in predicting anticancer properties of peptides using four different feature sets and assess the suitability of various ML algorithms. CONCLUSION: PDAUG offers tools for peptide library generation, data visualization, built-in and public database peptide sequence retrieval, peptide feature calculation, and machine learning (ML) modeling. Additionally, this toolset facilitates researchers to combine PDAUG with hundreds of compatible existing Galaxy tools for limitless analytic strategies. Jayadev Joshi, Daniel J. Blankenberg |
BMC Bioinform. | 2 |
| 2021 | SimText: a text mining framework for interactive analysis and visualization of similarities among biomedical entitiesabstractSUMMARY: Literature exploration in PubMed on a large number of biomedical entities (e.g. genes, diseases or experiments) can be time-consuming and challenging, especially when assessing associations between entities. Here, we describe SimText, a user-friendly toolset that provides customizable and systematic workflows for the analysis of similarities among a set of entities based on text. SimText can be used for (i) text collection from PubMed and extraction of words with different text mining approaches, and (ii) interactive analysis and visualization of data using unsupervised learning techniques in an interactive app. AVAILABILITY AND IMPLEMENTATION: We developed SimText as an open-source R software and integrated it into Galaxy (https://usegalaxy.eu), an online data analysis platform with supporting self-learning training material available at https://training.galaxyproject.org. A command-line version of the toolset is available for download from GitHub (https://github.com/dlal-group/simtext) or as Docker image (https://hub.docker.com/r/dlalgroup/simtext/tags.). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marie Macnee, Eduardo Pérez-Palma, Sarah Schumacher-Bass, Jarrod Dalton, Costin Leu, Daniel J. Blankenberg, Dennis Lal |
Bioinform. | 6 |
| 2020 | A crowdsourced set of curated structural variants for the human genomeabstractA high quality benchmark for small variants encompassing 88 to 90% of the reference genome has been developed for seven Genome in a Bottle (GIAB) reference samples. However a reliable benchmark for large indels and structural variants (SVs) is more challenging. In this study, we manually curated 1235 SVs, which can ultimately be used to evaluate SV callers or train machine learning models. We developed a crowdsourcing app-SVCurator-to help GIAB curators manually review large indels and SVs within the human genome, and report their genotype and size accuracy. SVCurator displays images from short, long, and linked read sequencing data from the GIAB Ashkenazi Jewish Trio son [NIST RM 8391/HG002]. We asked curators to assign labels describing SV type (deletion or insertion), size accuracy, and genotype for 1235 putative insertions and deletions sampled from different size bins between 20 and 892,149 bp. 'Expert' curators were 93% concordant with each other, and 37 of the 61 curators had at least 78% concordance with a set of 'expert' curators. The curators were least concordant for complex SVs and SVs that had inaccurate breakpoints or size predictions. After filtering events with low concordance among curators, we produced high confidence labels for 935 events. The SVCurator crowdsourced labels were 94.5% concordant with the heuristic-based draft benchmark SV callset from GIAB. We found that curators can successfully evaluate putative SVs when given evidence from multiple sequencing technologies. Lesley M. Chapman, Noah Spies, Patrick Pai, Chun Shen Lim, Andrew Carroll, Giuseppe Narzisi, Christopher M. Watson, Christos Proukakis, Wayne E. Clarke, Naoki Nariai, Eric T. Dawson, Garan Jones, Daniel J. Blankenberg, Christian Brueffer, Chunlin Xiao, Sree Rohit Raj Kolora, Noah Alexander, Paul Wolujewicz, Azza E. Ahmed, Saadlee Shehreen, Aaron M. Wenger, Marc L. Salit, Justin M. Zook |
PLoS Comput. Biol. | 13 |
| 2014 | Wrangling Galaxy's reference dataabstractUNLABELLED: The Galaxy platform has developed into a fully featured collaborative workbench, with goals of inherently capturing provenance to enable reproducible data analysis, and of making it straightforward to run one's own server. However, many Galaxy platform tools rely on the presence of reference data, such as alignment indexes, to function efficiently. Until now, the building of this cache of data for Galaxy has been an error-prone manual process lacking reproducibility and provenance. The Galaxy Data Manager framework is an enhancement that changes the management of Galaxy's built-in data cache from a manual procedure to an automated graphical user interface (GUI) driven process, which contains the same openness, reproducibility and provenance that is afforded to Galaxy's analysis tools. Data Manager tools allow the Galaxy administrator to download, create and install additional datasets for any type of reference data in real time. AVAILABILITY AND IMPLEMENTATION: The Galaxy Data Manager framework is implemented in Python and has been integrated as part of the core Galaxy platform. Individual Data Manager tools can be defined locally or installed from a ToolShed, allowing the Galaxy community to define additional Data Manager tools as needed, with full versioning and dependency support. Daniel J. Blankenberg, James E. Johnson, James Taylor 0001, Anton Nekrutenko |
Bioinform. | 1 |
| 2011 | Making whole genome multiple alignments usable for biologistsabstractSUMMARY: Here we describe a set of tools implemented within the Galaxy platform designed to make analysis of multiple genome alignments truly accessible for biologists. These tools are available through both a web-based graphical user interface and a command-line interface. AVAILABILITY AND IMPLEMENTATION: This open-source toolset was implemented in Python and has been integrated into the online data analysis platform Galaxy (public web access: http://usegalaxy.org; download: http://getgalaxy.org). Additional help is available as a live supplement from http://usegalaxy.org/u/dan/p/maf. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniel J. Blankenberg, James Taylor 0001, Anton Nekrutenko |
Bioinform. | 1 |
| 2010 | Manipulation of FASTQ data with GalaxyabstractSUMMARY: Here, we describe a tool suite that functions on all of the commonly known FASTQ format variants and provides a pipeline for manipulating next generation sequencing data taken from a sequencing machine all the way through the quality filtering steps. AVAILABILITY AND IMPLEMENTATION: This open-source toolset was implemented in Python and has been integrated into the online data analysis platform Galaxy (public web access: http://usegalaxy.org; download: http://getgalaxy.org). Two short movies that highlight the functionality of tools described in this manuscript as well as results from testing components of this tool suite against a set of previously published files are available at http://usegalaxy.org/u/dan/p/fastq Daniel J. Blankenberg, Assaf Gordon, Gregory Von Kuster, Nathan Coraor, James Taylor 0001, Anton Nekrutenko |
Bioinform. | 1 |
| 2007 | Quantitative sequence-function relationships in proteins based on gene ontologyabstractBACKGROUND: The relationship between divergence of amino-acid sequence and divergence of function among homologous proteins is complex. The assumption that homologs share function--the basis of transfer of annotations in databases--must therefore be regarded with caution. Here, we present a quantitative study of sequence and function divergence, based on the Gene Ontology classification of function. We determined the relationship between sequence divergence and function divergence in 6828 protein families from the PFAM database. Within families there is a broad range of sequence similarity from very closely related proteins--for instance, orthologs in different mammals--to very distantly-related proteins at the limit of reliable recognition of homology. RESULTS: We correlated the divergence in sequences determined from pairwise alignments, and the divergence in function determined by path lengths in the Gene Ontology graph, taking into account the fact that many proteins have multiple functions. Our results show that, among homologous proteins, the proportion of divergent functions decreases dramatically above a threshold of sequence similarity at about 50% residue identity. For proteins with more than 50% residue identity, transfer of annotation between homologs will lead to an erroneous attribution with a totally dissimilar function in fewer than 6% of cases. This means that for very similar proteins (about 50 % identical residues) the chance of completely incorrect annotation is low; however, because of the phenomenon of recruitment, it is still non-zero. CONCLUSION: Our results describe general features of the evolution of protein function, and serve as a guide to the reliability of annotation transfer, based on the closeness of the relationship between a new protein and its nearest annotated relative. Vineet Sangar, Daniel J. Blankenberg, Naomi Altman, Arthur M. Lesk |
BMC Bioinform. | 2 |