VLDB 2026 Research / reviewers in the wild / expert
Benjamin Schubert
dblp:121/1174
· DBLP profile ↗
20ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Multi-Epitope Vaccine Effectiveness Through State-of-the-Art Proteasomal Cleavage Prediction With Deep LearningabstractMulti-epitope vaccines (EVs) represent a versatile and promising approach for combating a wide range of diseases, including viral and bacterial infections, parasitic diseases, and cancer. A critical aspect of EV design is ensuring efficient proteasomal cleavage of the synthetic polypeptide to recover therapeutic epitopes for presentation to$\mathbf{T}$cells. This process is influenced by the selection and arrangement of epitopes, as well as the design of linkers joining them. Modern EV design frameworks leverage proteasomal cleavage predictors to optimize epitope recovery and vaccine efficacy. However, the predictive power of these tools remains a limiting factor, particularly for challenging cleavage sites such as N -terminals. In this work, we systematically review and benchmark recent advances in proteasomal cleavage prediction, focusing on deep learning-based methods. We evaluate a range of architectures, including MLPs, CNNs, LSTMs, and Transformers, alongside simpler baselines such as logistic regression. Our results demonstrate that while complex models achieve marginally higher predictive performance, simpler models remain competitive and offer significant computational efficiency. We also explore the impact of dataset size, window size, and training techniques, finding diminishing returns from increasingly larger datasets and more complex models. Notably, our benchmarking highlights substantial improvements in predicting N-terminal cleavage sites, which are often overlooked but critical for EV design. Our findings provide practical guidance for the development of nextgeneration proteasomal cleavage predictors and underscore the importance of considering the probabilistic nature of cleavage in EV design. By consolidating the state of the art and introducing efficient baselines, this work aims to advance the field of multiepitope vaccine design and support the development of more effective and personalized immunotherapies.11Code for all experiments is available at this repository: https://github.com/ziegleringo/cleavage_benchmark. Ingo Ziegler, Bolei Ma, Bernd Bischl, Benjamin Schubert, Emilio Dorigatti |
BIBM | 4 |
| 2025 | TransFactor - prediction of pro-viral SARS-CoV-2 host factors using a protein language modelabstractMOTIVATION: Recent pandemics have revealed significant gaps in our understanding of viral pathogenesis, exposing an urgent need for methods to identify and prioritize key host proteins (host factors) as potential targets for antiviral treatments. De novo generation of experimental datasets is limited by their heterogeneity, and for looming future pandemics, may not be feasible due to limitations of experimental approaches. RESULTS: Here, we present TransFactor, a computational framework for predicting and prioritizing candidate host factors using only protein sequence data. It leverages the pre-trained ESM-2 protein language model, fine-tuned on a limited set of experimentally determined host factors aggregated from 33 independent SARS-CoV-2 studies. TransFactor outperforms machine and deep learning baselines and its predictions align with Gene Ontology enrichments of known host factors, but also provide interpretability through a computational alanine scan, enabling the identification of pro-viral protein domains such as COMM, PX, and RRM, that may be used to direct experimental investigations of virus biology and guide rational design of antiviral therapies. Our findings demonstrate the potential of transformer-based models to advance host factor prediction, providing a framework extendable to orthogonal input modalities and other infectious diseases, enhancing our preparedness for current and future viral threats. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/marsico-lab/TransFactor. A full reproducibility package, including code, trained models, and data, is archived on Zenodo (https://doi.org/10.5281/zenodo.16793684). Valter Bergant, Samuele Firmani, Corinna Grünke, Batiste Bonnal, Alexander Henrici, Andreas Pichlmair, Benjamin Schubert, Annalisa Marsico |
Bioinform. | 8 |
| 2024 | Advancing Neonatal Care: A Deep Learning Approach for Non-Contact Heart Rate MonitoringabstractHeart rate is an important indicator of newborn health status. Conventional wired heart rate monitoring is affected by motion, can limit parental bonding and is prone to damage the fragile newborn skin. Video-based heart rate monitoring in adults has shown potential for the assessment of cardiac functions in optimal acquisition conditions. However, automated methods adapted for neonates, trained with limited sample sizes and capable of handling occlusions and variable illumination, still need to be explored. This work proposes a new deep-learning pipeline for video-based neonatal heart rate measurement, integrating color and infrared signals from mul-tiple neonatal body regions. We train deep learning models for pose estimation and heart rate detection in a pilot cohort of five neonates recorded in the clinic for up to 60 minutes. Our methods show generalization in a leave-one-out cross-validation scheme; the newborn pose estimation model presents high performance (average precision=0.85±0.07), while the heart rate detection model achieves a mean absolute error of 3.83±1.22 beats per minute compared to the electrocardiogram heart rate. Automated video-based heart rate measurement could provide a non-contact, low-cost complement to current HR monitoring technologies in the clinic or outpatient settings. Alex Grafton, Alejandra Castelblanco, Joana M. Warnecke, Lynn Thomson, Benjamin Schubert, Anne Hilgendorff, Julia A. Schnabel, Joan Lasenby, Kathryn Beardsall |
HealthCom | 5 |
| 2023 | Frequentist Uncertainty Quantification in Semi-Structured Neural NetworksabstractSemi-structured regression (SSR) models jointly learn the effect of structured (tabular) and unstructured (non-tabular) data through additive predictors and deep neural networks (DNNs), respectively. Inference in SSR models aims at deriving confidence intervals for the structured predictor, although current approaches ignore the variance of the DNN estimation of the unstructured effects. This results in an underestimation of the variance of the structured coefficients and, thus, an increase of Type-I error rates. To address this shortcoming, we present here a theoretical framework for structured inference in SSR models that incorporates the variance of the DNN estimate into confidence intervals for the structured predictor. By treating this estimate as a random offset with known variance, our formulation is agnostic to the specific deep uncertainty quantification method employed. Through numerical experiments and a practical application on a medical dataset, we show that our approach results in increased coverage of the true structured coefficients and thus a reduction in Type-I error rate compared to ignoring the variance of the neural network, naive ensembling of SSR models, and a variational inference baseline. Emilio Dorigatti, Benjamin Schubert, Bernd Bischl, David Rügamer |
AISTATS | 2 |
| 2023 | Multi-omics regulatory network inference in the presence of missing dataabstractA key problem in systems biology is the discovery of regulatory mechanisms that drive phenotypic behaviour of complex biological systems in the form of multi-level networks. Modern multi-omics profiling techniques probe these fundamental regulatory networks but are often hampered by experimental restrictions leading to missing data or partially measured omics types for subsets of individuals due to cost restrictions. In such scenarios, in which missing data is present, classical computational approaches to infer regulatory networks are limited. In recent years, approaches have been proposed to infer sparse regression models in the presence of missing information. Nevertheless, these methods have not been adopted for regulatory network inference yet. In this study, we integrated regression-based methods that can handle missingness into KiMONo, a Knowledge guided Multi-Omics Network inference approach, and benchmarked their performance on commonly encountered missing data scenarios in single- and multi-omics studies. Overall, two-step approaches that explicitly handle missingness performed best for a wide range of random- and block-missingness scenarios on imbalanced omics-layers dimensions, while methods implicitly handling missingness performed best on balanced omics-layers dimensions. Our results show that robust multi-omics network inference in the presence of missing data with KiMONo is feasible and thus allows users to leverage available multi-omics data to its full extent. Juan D. Henao, Michael Lauber, Manuel Azevedo, Anastasiia Grekova, Fabian J. Theis, Markus List, Christoph Ogris, Benjamin Schubert |
Briefings Bioinform. | 8 |
| 2020 | Joint epitope selection and spacer design for string-of-beads vaccinesabstractMOTIVATION: Conceptually, epitope-based vaccine design poses two distinct problems: (i) selecting the best epitopes to elicit the strongest possible immune response and (ii) arranging and linking them through short spacer sequences to string-of-beads vaccines, so that their recovery likelihood during antigen processing is maximized. Current state-of-the-art approaches solve this design problem sequentially. Consequently, such approaches are unable to capture the inter-dependencies between the two design steps, usually emphasizing theoretical immunogenicity over correct vaccine processing, thus resulting in vaccines with less effective immunogenicity in vivo. RESULTS: In this work, we present a computational approach based on linear programming, called JessEV, that solves both design steps simultaneously, allowing to weigh the selection of a set of epitopes that have great immunogenic potential against their assembly into a string-of-beads construct that provides a high chance of recovery. We conducted Monte Carlo cleavage simulations to show that a fixed set of epitopes often cannot be assembled adequately, whereas selecting epitopes to accommodate proper cleavage requirements substantially improves their recovery probability and thus the effective immunogenicity, pathogen and population coverage of the resulting vaccines by at least 2-fold. AVAILABILITY AND IMPLEMENTATION: The software and the data analyzed are available at https://github.com/SchubertLab/JessEV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Emilio Dorigatti, Benjamin Schubert |
Bioinform. | 2 |
| 2020 | Graph-theoretical formulation of the generalized epitope-based vaccine design problemabstractEpitope-based vaccines have revolutionized vaccine research in the last decades. Due to their complex nature, bioinformatics plays a pivotal role in their development. However, existing algorithms address only specific parts of the design process or are unable to provide formal guarantees on the quality of the solution. We present a unifying formalism of the general epitope vaccine design problem that tackles all phases of the design process simultaneously and combines all prevalent design principles. We then demonstrate how to formulate the developed formalism as an integer linear program, which guarantees optimality of the designs. This makes it possible to explore new regions of the vaccine design space, analyze the trade-offs between the design phases, and balance the many requirements of vaccines. Emilio Dorigatti, Benjamin Schubert |
PLoS Comput. Biol. | 2 |
| 2019 | The EVcouplings Python framework for coevolutionary sequence analysisabstractSUMMARY: Coevolutionary sequence analysis has become a commonly used technique for de novo prediction of the structure and function of proteins, RNA, and protein complexes. We present the EVcouplings framework, a fully integrated open-source application and Python package for coevolutionary analysis. The framework enables generation of sequence alignments, calculation and evaluation of evolutionary couplings (ECs), and de novo prediction of structure and mutation effects. The combination of an easy to use, flexible command line interface and an underlying modular Python package makes the full power of coevolutionary analyses available to entry-level and advanced users. AVAILABILITY AND IMPLEMENTATION: https://github.com/debbiemarkslab/evcouplings. Thomas A. Hopf, Anna G. Green, Benjamin Schubert, Sophia Mersmann, Charlotta Schärfe, John Ingraham, Ágnes Tóth-Petróczy, Kelly Brock, Adam J. Riesselman, Perry Palmedo, Chan Kang, Robert P. Sheridan, Eli J. Draizen, Christian Dallago, Chris Sander, Debora S. Marks |
Bioinform. | 3 |
| 2019 | ClinOmicsTrailbc: a visual analytics tool for breast cancer treatment stratificationabstractMOTIVATION: Breast cancer is the second leading cause of cancer death among women. Tumors, even of the same histopathological subtype, exhibit a high genotypic diversity that impedes therapy stratification and that hence must be accounted for in the treatment decision-making process. RESULTS: Here, we present ClinOmicsTrailbc, a comprehensive visual analytics tool for breast cancer decision support that provides a holistic assessment of standard-of-care targeted drugs, candidates for drug repositioning and immunotherapeutic approaches. To this end, our tool analyzes and visualizes clinical markers and (epi-)genomics and transcriptomics datasets to identify and evaluate the tumor's main driver mutations, the tumor mutational burden, activity patterns of core cancer-relevant pathways, drug-specific biomarkers, the status of molecular drug targets and pharmacogenomic influences. In order to demonstrate ClinOmicsTrailbc's rich functionality, we present three case studies highlighting various ways in which ClinOmicsTrailbc can support breast cancer precision medicine. ClinOmicsTrailbc is a powerful integrated visual analytics tool for breast cancer research in general and for therapy stratification in particular, assisting oncologists to find the best possible treatment options for their breast cancer patients based on actionable, evidence-based results. AVAILABILITY AND IMPLEMENTATION: ClinOmicsTrailbc can be freely accessed at https://clinomicstrail.bioinf.uni-sb.de. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lara Schneider, Tim Kehl, Kristina Thedinga, Nadja Liddy Grammes, Christina Backes, Christopher Mohr, Benjamin Schubert, Kerstin Lenhof, Nico Gerstner, Andreas Daniel Hartkopf, Markus Wallwiener, Oliver Kohlbacher, Andreas Keller, Eckart Meese, Norbert Graf 0001, Hans-Peter Lenhof |
Bioinform. | 7 |
| 2018 | Population-specific design of de-immunized protein biotherapeuticsabstractImmunogenicity is a major problem during the development of biotherapeutics since it can lead to rapid clearance of the drug and adverse reactions.The challenge for biotherapeutic design is therefore to identify mutants of the protein sequence that minimize immunogenicity in a target population whilst retaining pharmaceutical activity and protein function.Current approaches are moderately successful in designing sequences with reduced immunogenicity, but do not account for the varying frequencies of different human leucocyte antigen alleles in a specific population and in addition, since many designs are non-functional, require costly experimental post-screening.Here, we report a new method for de-immunization design using multi-objective combinatorial optimization.The method simultaneously optimizes the likelihood of a functional protein sequence at the same time as minimizing its immunogenicity tailored to a target population.We bypass the need for three-dimensional protein structure or molecular simulations to identify functional designs by automatically generating sequences using probabilistic models that have been used previously for mutation effect prediction and structure prediction.As proof-of-principle we designed sequences of the C2 domain of Factor VIII and tested them experimentally, resulting in a good correlation with the predicted immunogenicity of our model. Author summaryTherapeutic proteins have become an important area of pharmaceutical research and have been successfully applied to treat many diseases in the last decades.However, biotherapeutics suffer from the formation of anti-drug antibodies, which can reduce the efficacy of the drug or even result in severe adverse effects.A main contributor to the antibody formation is a T-cell mediated immune reaction caused by presentation of small immunogenic peptides derived from the biotherapeutic.Targeting these peptides via Benjamin Schubert, Charlotta Schärfe, Pierre Dönnes, Thomas A. Hopf, Debora S. Marks, Oliver Kohlbacher |
PLoS Comput. Biol. | 1 |
| 2017 | ImmunoNodes - graphical development of complex immunoinformatics workflowsabstractBACKGROUND: Immunoinformatics has become a crucial part in biomedical research. Yet many immunoinformatics tools have command line interfaces only and can be difficult to install. Web-based immunoinformatics tools, on the other hand, are difficult to integrate with other tools, which is typically required for the complex analysis and prediction pipelines required for advanced applications. RESULT: We present ImmunoNodes, an immunoinformatics toolbox that is fully integrated into the visual workflow environment KNIME. By dragging and dropping tools and connecting them to indicate the data flow through the pipeline, it is possible to construct very complex workflows without the need for coding. CONCLUSION: ImmunoNodes allows users to build complex workflows with an easy to use and intuitive interface with a few clicks on any desktop computer. Benjamin Schubert, Luis de la Garza, Christopher Mohr, Mathias Walzer, Oliver Kohlbacher |
BMC Bioinform. | 1 |
| 2016 | FRED 2: an immunoinformatics framework for PythonabstractUNLABELLED: Immunoinformatics approaches are widely used in a variety of applications from basic immunological to applied biomedical research. Complex data integration is inevitable in immunological research and usually requires comprehensive pipelines including multiple tools and data sources. Non-standard input and output formats of immunoinformatics tools make the development of such applications difficult. Here we present FRED 2, an open-source immunoinformatics framework offering easy and unified access to methods for epitope prediction and other immunoinformatics applications. FRED 2 is implemented in Python and designed to be extendable and flexible to allow rapid prototyping of complex applications. AVAILABILITY AND IMPLEMENTATION: FRED 2 is available at http://fred-2.github.io CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin Schubert, Mathias Walzer, Hans-Philipp Brachvogel, András Szolek, Christopher Mohr, Oliver Kohlbacher |
Bioinform. | 1 |
| 2015 | Failure Sketches: A Better Way to Debug
Baris Kasikci, Cristiano Pereira, Gilles Pokam, Benjamin Schubert, Madan Musuvathi, George Candea |
HotOS | 4 |
| 2015 | Improved error resilience for volte and VoIP with 3GPP EVS channel aware codingabstractA highly error resilient mode of the newly standardized 3GPP EVS speech codec is described. Compared to the AMR-WB codec and other conversational codecs, the EVS channel aware mode offers significantly improved error resilience in voice communication over packet-switched networks such as Voice-over-IP (VoIP) and Voice-over-LTE (VoLTE). The error resilience is achieved using a form of in-band forward error correction. Source-controlled coding techniques are used to identify candidate speech frames for bitrate reduction, leaving spare bits for transmission of partial copies of prior frames such that a constant bit rate is maintained. The self-contained partial copies are used to improve the error robustness in case the original primary frame is lost or discarded due to late arrival. Subjective evaluation results from ITU-T P.800 Mean Opinion Score (MOS) tests are provided, showing improved quality under channel impairments as well as negligible impact to clean channel performance. Venkatraman Atti, Daniel J. Sinder, Shaminda Subasingha, Vivek Rajendran, Duminda A. Dewasurendra, Venkata Chebiyyam, Imre Varga, Venkatesh Krishnan, Benjamin Schubert, Jérémie Lecomte, Xingtao Zhang, Lei Miao 0004 |
ICASSP | 9 |
| 2015 | Failure sketching: a technique for automated root cause diagnosis of in-production failuresabstractDevelopers spend a lot of time searching for the root causes of software failures. For this, they traditionally try to reproduce those failures, but unfortunately many failures are so hard to reproduce in a test environment that developers spend days or weeks as ad-hoc detectives. The shortcomings of many solutions proposed for this problem prevent their use in practice. Baris Kasikci, Benjamin Schubert, Cristiano Pereira, Gilles Pokam, George Candea |
SOSP | 2 |
| 2015 | EpiToolKit - a web-based workbench for vaccine designabstractUNLABELLED: EpiToolKit is a virtual workbench for immunological questions with a focus on vaccine design. It offers an array of immunoinformatics tools covering MHC genotyping, epitope and neo-epitope prediction, epitope selection for vaccine design, and epitope assembly. In its recently re-implemented version 2.0, EpiToolKit provides a range of new functionality and for the first time allows combining tools into complex workflows. For inexperienced users it offers simplified interfaces to guide the users through the analysis of complex immunological data sets. AVAILABILITY AND IMPLEMENTATION: http://www.epitoolkit.de CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin Schubert, Hans-Philipp Brachvogel, Christopher Jürges, Oliver Kohlbacher |
Bioinform. | 1 |
| 2014 | Sinusoidal substitution - An integrated parametric tool for enhancement of transform-based perceptual audio codersabstractTransform-based audio coders are the preferred technique for music data compression. However, at low bitrates, traditional coders based on Modified Discrete Cosine Transform are prone to strong warbling and roughness artifacts originating from sparsely coded tonal components. Parametric coders, in turn, suffer from an unpleasantly artificial sound and do not scale well up to perceptual transparency. Hybrid transform-based and parametric coding could potentially overcome the limits of the individual approaches. Yet, existing hybrid coders are hampered by the lack of integrative interplay between both techniques. We outline our ideas how to tightly integrate transform-based coding and parametric coding to obtain an enhanced perceptual quality and scalability. Also, we provide listening test results which demonstrate the benefits of our hybrid coder design. Sascha Disch, Benjamin Schubert |
ICASSP | 2 |
| 2014 | OptiType: precision HLA typing from next-generation sequencing dataabstractMOTIVATION: The human leukocyte antigen (HLA) gene cluster plays a crucial role in adaptive immunity and is thus relevant in many biomedical applications. While next-generation sequencing data are often available for a patient, deducing the HLA genotype is difficult because of substantial sequence similarity within the cluster and exceptionally high variability of the loci. Established approaches, therefore, rely on specific HLA enrichment and sequencing techniques, coming at an additional cost and extra turnaround time. RESULT: We present OptiType, a novel HLA genotyping algorithm based on integer linear programming, capable of producing accurate predictions from NGS data not specifically enriched for the HLA cluster. We also present a comprehensive benchmark dataset consisting of RNA, exome and whole-genome sequencing data. OptiType significantly outperformed previously published in silico approaches with an overall accuracy of 97% enabling its use in a broad range of applications. András Szolek, Benjamin Schubert, Christopher Mohr, Marc Sturm, Magdalena Feldhahn, Oliver Kohlbacher |
Bioinform. | 2 |
| 2013 | Cheap beeps - Efficient synthesis of sinusoids and sweeps in the MDCT domainabstractModern transform audio coders often employ parametric enhancements, like noise substitution or bandwidth extension. In addition to these well-known parametric tools, it might also be desirable to synthesize parametric sinusoidal tones in the decoder. Low computational complexity is an important criterion in codec development and essential for acceptance and deployment. Therefore, efficient ways of generating these tones are needed. Since contemporary codecs like AAC or USAC are based on an MDCT domain representation of audio, we propose to generate synthetic tones by patching tone patterns into the MDCT spectrum at the decoder. We demonstrate how appropriate spectral patterns can be derived and adapted to their target location in (and between) the MDCT time/frequency (t/f) grid to seamlessly synthesize high quality sinusoidal tones including sweeps. Sascha Disch, Benjamin Schubert, Bernd Edler |
ICASSP | 2 |
| 2006 | FIR Linear Relay Network with Frequency Selective ChannelsabstractIn this work, the concept of linear relay network code [1] for an arbitrary number of relay nodes is extended to scalar frequency selective channels. Furthermore, we present performance criteria derived from the pairwise error probability (PEP) of a maximum likelihood sequence estimator (MLSE). By neglecting the transient processes of finite frames, deeper insights into the performance of FIR linear relay networks can be gained. Finally, the performance of a linear relay network using an MLSE sphere decoder at the destination is illustrated by some Monte Carlo simulations. Tobias J. Oechtering, Benjamin Schubert, Holger Boche |
ICC | 2 |