Daniel Beck

dblp:55/4410 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 24% Optimization for machine learning · 19% Kernel, tree and ensemble methods · 19%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.522020
BOSS: Bayesian Optimization over String Spaces · NeurIPS 2020
Joint Emotion Analysis via Multi-task Gaussian Processes · EMNLP 2014
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.412020
BOSS: Bayesian Optimization over String Spaces · NeurIPS 2020
Machine learning › Kernel, tree and ensemble methods › kernel methods › structured kernel
string kernel
0.412020
BOSS: Bayesian Optimization over String Spaces · NeurIPS 2020
Information retrieval › query processing
dynamic pruning
0.412019
Accelerated Query Processing Via Similarity Score Prediction · SIGIR 2019
Information retrieval
query processing
0.412019
Accelerated Query Processing Via Similarity Score Prediction · SIGIR 2019
Information retrieval › similarity search
top-k retrieval
0.412019
Accelerated Query Processing Via Similarity Score Prediction · SIGIR 2019
Machine learning › Graph learning › graph autoencoder › graph encoder-decoder
graph-to-sequence learning
0.312018
Graph-to-Sequence Learning using Gated Graph Neural Networks · ACL (1) 2018
Natural language and speech › Machine translation
neural machine translation
0.312018
Graph-to-Sequence Learning using Gated Graph Neural Networks · ACL (1) 2018
Bioinformatics and computational biology › computational microbiology
microbiome analysis
0.212015
Seed: a user-friendly tool for exploring and visualizing microbial community data · Bioinform. 2015
Natural language and speech › Information extraction and text analysis
emotion recognition
0.212014
Joint Emotion Analysis via Multi-task Gaussian Processes · EMNLP 2014
Bioinformatics and computational biology › computational microbiology › microbiome analysis
microbial community analysis
0.112012
mcaGUI: microbial community analysis R-Graphical User Interface (GUI) · Bioinform. 2012
Bioinformatics and computational biology
metagenomics
0.112011
OTUbase: an R infrastructure package for operational taxonomic unit data · Bioinform. 2011
Visualization and visual analytics
biological data visualization
0.112015
Seed: a user-friendly tool for exploring and visualizing microbial community data · Bioinform. 2015
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
multi-output gaussian process
0.112014
Joint Emotion Analysis via Multi-task Gaussian Processes · EMNLP 2014
Bioinformatics and computational biology › sequence analysis
sequence classification
0.012011
OTUbase: an R infrastructure package for operational taxonomic unit data · Bioinform. 2011
Bioinformatics and computational biology › metagenomics
taxonomic classification
0.012011
OTUbase: an R infrastructure package for operational taxonomic unit data · Bioinform. 2011

Methods — techniques the papers use, named apart from their topics

principal coordinate analysis · 0.4hierarchical clustering · 0.4heatmap · 0.4genetic algorithm · 0.4context-free grammar · 0.4acquisition function maximization · 0.4term-dependent features · 0.4risk-sensitive prediction · 0.4gated graph neural network · 0.3multi-task gaussian process · 0.2low-rank coregionalisation · 0.2r · 0.1clustering · 0.1
YearPublicationVenuePosition
2026 Rubric-guided Prompting for Automatic Assessment Scoring
Samiha Haque, Danula Hettiachchi, Daniel Beck
L@S3
2025 Zero-Shot Performance Prediction for Probabilistic Scaling Laws
abstract
The prediction of learning curves for Natural Language Processing (NLP) models enables informed decision-making to meet specific performance objectives, while reducing computational overhead and lowering the costs associated with dataset acquisition and curation. In this work, we formulate the prediction task as a multitask learning problem, where each task’s data is modelled as being organized within a two-layer hierarchy. To model the shared information and dependencies across tasks and hierarchical levels, we employ latent variable multi-output Gaussian Processes, enabling to account for task correlations and supporting zero-shot prediction of learning curves (LCs). We demonstrate that this approach facilitates the development of probabilistic scaling laws at lower costs. Applying an active learning strategy, LCs can be queried to reduce predictive uncertainty and provide predictions close to ground truth scaling laws. We validate our framework on three small-scale NLP datasets with up to $30$ LCs. These are obtained from nanoGPT models, from bilingual translation using mBART and Transformer models, and from multilingual translation using M2M100 models of varying sizes.
Viktoria Schram, Markus Hiller, Daniel Beck, Trevor Cohn
NeurIPS3
2023 Performance Prediction via Bayesian Matrix Factorisation for Multilingual Natural Language Processing Tasks
abstract
Performance prediction for Natural Language Processing (NLP) seeks to reduce the experimental burden resulting from the myriad of different evaluation scenarios, e.g., the combination of languages used in multilingual transfer.In this work, we explore the framework of Bayesian matrix factorisation for performance prediction, as many experimental settings in NLP can be naturally represented in matrix format.Our approach outperforms the stateof-the-art in several NLP benchmarks, including machine translation and cross-lingual entity linking.Furthermore, it also avoids hyperparameter tuning and is able to provide uncertainty estimates over predictions.
Viktoria Schram, Daniel Beck, Trevor Cohn
EACL2
2023 Graph embedding-based link prediction for literature-based discovery in Alzheimer's Disease
abstract
OBJECTIVE: We explore the framing of literature-based discovery (LBD) as link prediction and graph embedding learning, with Alzheimer's Disease (AD) as our focus disease context. The key link prediction setting of prediction window length is specifically examined in the context of a time-sliced evaluation methodology. METHODS: We propose a four-stage approach to explore literature-based discovery for Alzheimer's Disease, creating and analyzing a knowledge graph tailored to the AD context, and predicting and evaluating new knowledge based on time-sliced link prediction. The first stage is to collect an AD-specific corpus. The second stage involves constructing an AD knowledge graph with identified AD-specific concepts and relations from the corpus. In the third stage, 20 pairs of training and testing datasets are constructed with the time-slicing methodology. Finally, we infer new knowledge with graph embedding-based link prediction methods. We compare different link prediction methods in this context. The impact of limiting prediction evaluation of LBD models in the context of short-term and longer-term knowledge evolution for Alzheimer's Disease is assessed. RESULTS: We constructed an AD corpus of over 16 k papers published in 1977-2021, and automatically annotated it with concepts and relations covering 11 AD-specific semantic entity types. The knowledge graph of Alzheimer's Disease derived from this resource consisted of ∼11 k nodes and ∼394 k edges, among which 34% were genotype-phenotype relationships, 57% were genotype-genotype relationships, and 9% were phenotype-phenotype relationships. A Structural Deep Network Embedding (SDNE) model consistently showed the best performance in terms of returning the most confident set of link predictions as time progresses over 20 years. A huge improvement in model performance was observed when changing the link prediction evaluation setting to consider a more distant future, reflecting the time required for knowledge accumulation. CONCLUSION: Neural network graph-embedding link prediction methods show promise for the literature-based discovery context, although the prediction setting is extremely challenging, with graph densities of less than 1%. Varying prediction window length on the time-sliced evaluation methodology leads to hugely different results and interpretations of LBD studies. Our approach can be generalized to enable knowledge discovery for other diseases. AVAILABILITY: Code, AD ontology, and data are available at https://github.com/READ-BioMed/readbiomed-lbd.
Yiyuan Pu, Daniel Beck, Karin Verspoor
J. Biomed. Informatics2
2023 Modelling Emotion Dynamics in Song Lyrics with State Space Models
abstract
Abstract Most previous work in music emotion recognition assumes a single or a few song-level labels for the whole song. While it is known that different emotions can vary in intensity within a song, annotated data for this setup is scarce and difficult to obtain. In this work, we propose a method to predict emotion dynamics in song lyrics without song-level supervision. We frame each song as a time series and employ a State Space Model (SSM), combining a sentence-level emotion predictor with an Expectation-Maximization (EM) procedure to generate the full emotion dynamics. Our experiments show that applying our method consistently improves the performance of sentence-level baselines without requiring any annotated songs, making it ideal for limited training data scenarios. Further analysis through case studies shows the benefits of our method while also indicating the limitations and pointing to future directions.
Yingjin Song, Daniel Beck
Trans. Assoc. Comput. Linguistics2
2023 On the Effectiveness of Images in Multi-modal Text Classification: An Annotation Study
abstract
Combining different input modalities beyond text is a key challenge for natural language processing. Previous work has been inconclusive as to the true utility of images as a supplementary information source for text classification tasks, motivating this large-scale human study of labelling performance given text-only, images-only, or both text and images. To this end, we create a new dataset accompanied with a novel annotation method—Japanese Entity Labeling with Dynamic Annotation—to deepen our understanding of the effectiveness of images for multi-modal text classification. By performing careful comparative analysis of human performance and the performance of state-of-the-art multi-modal text classification models, we gain valuable insights into differences between human and model performance, and the conditions under which images are beneficial for text classification.
Chunpeng Ma, Aili Shen, Hiyori Yoshikawa, Tomoya Iwakura, Daniel Beck, Timothy Baldwin
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2022 Uncertainty Estimation and Reduction of Pre-trained Models for Text Regression
abstract
Abstract State-of-the-art classification and regression models are often not well calibrated, and cannot reliably provide uncertainty estimates, limiting their utility in safety-critical applications such as clinical decision-making. While recent work has focused on calibration of classifiers, there is almost no work in NLP on calibration in a regression setting. In this paper, we quantify the calibration of pre- trained language models for text regression, both intrinsically and extrinsically. We further apply uncertainty estimates to augment training data in low-resource domains. Our experiments on three regression tasks in both self-training and active-learning settings show that uncertainty estimation can be used to increase overall performance and enhance model generalization.
Yuxia Wang 0003, Daniel Beck, Karin Verspoor, Timothy Baldwin
Trans. Assoc. Comput. Linguistics2
2021 On the (In)Effectiveness of Images for Text Classification
abstract
Chunpeng Ma, Aili Shen, Hiyori Yoshikawa, Tomoya Iwakura, Daniel Beck, Timothy Baldwin. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Chunpeng Ma, Aili Shen, Hiyori Yoshikawa, Tomoya Iwakura, Daniel Beck, Timothy Baldwin
EACL5
2021 Generating Diverse Descriptions from Semantic Graphs
abstract
Text generation from semantic graphs is traditionally performed with deterministic methods, which generate a unique description given an input graph. However, the generation problem admits a range of acceptable textual outputs, exhibiting lexical, syntactic and semantic variation. To address this disconnect, we present two main contributions. First, we propose a stochastic graph-to-text model, incorporating a latent variable in an encoder-decoder model, and its use in an ensemble. Second, to assess the diversity of the generated sentences, we propose a new automatic evaluation metric which jointly evaluates output diversity and quality in a multi-reference setting. We evaluate the models on WebNLG datasets in English and Russian, and show an ensemble of stochastic models produces diverse sets of generated sentences while, retaining similar quality to state-of-the-art models.
Jiuzhou Han, Daniel Beck, Trevor Cohn
INLG2
2021 Predicting environmentally responsive transgenerational differential DNA methylated regions (epimutations) in the genome using a hybrid deep-machine learning approach
abstract
BACKGROUND: Deep learning is an active bioinformatics artificial intelligence field that is useful in solving many biological problems, including predicting altered epigenetics such as DNA methylation regions. Deep learning (DL) can learn an informative representation that addresses the need for defining relevant features. However, deep learning models are computationally expensive, and they require large training datasets to achieve good classification performance. RESULTS: One approach to addressing these challenges is to use a less complex deep learning network for feature selection and Machine Learning (ML) for classification. In the current study, we introduce a hybrid DL-ML approach that uses a deep neural network for extracting molecular features and a non-DL classifier to predict environmentally responsive transgenerational differential DNA methylated regions (DMRs), termed epimutations, based on the extracted DL-based features. Various environmental toxicant induced epigenetic transgenerational inheritance sperm epimutations were used to train the model on the rat genome DNA sequence and use the model to predict transgenerational DMRs (epimutations) across the entire genome. CONCLUSION: The approach was also used to predict potential DMRs in the human genome. Experimental results show that the hybrid DL-ML approach outperforms deep learning and traditional machine learning methods.
Pegah Mavaie, Lawrence B. Holder, Daniel Beck, Michael K. Skinner
BMC Bioinform.3
2020 BOSS: Bayesian Optimization over String Spaces
abstract
This article develops a Bayesian optimization (BO) method which acts directly over raw strings, proposing the first uses of string kernels and genetic algorithms within BO loops. Recent applications of BO over strings have been hindered by the need to map inputs into a smooth and unconstrained latent space. Learning this projection is computationally and data-intensive. Our approach instead builds a powerful Gaussian process surrogate model based on string kernels, naturally supporting variable length inputs, and performs efficient acquisition function maximization for spaces with syntactic constraints. Experiments demonstrate considerably improved optimization over existing approaches across a broad range of constraints, including the popular setting where syntax is governed by a context-free grammar.
Henry B. Moss, David S. Leslie, Daniel Beck, Javier González 0002, Paul Rayson
NeurIPS3
2019 A Unified Neural Architecture for Instrumental Audio Tasks
abstract
Within Music Information Retrieval (MIR), prominent tasks - including pitch-tracking, source-separation, super-resolution, and synthesis - typically call for specialised methods, despite their similarities. Conditional Generative Adversarial Networks (cGANs) have been shown to be highly versatile in learning general image-to-image translations, but have not yet been adapted across MIR. In this work, we present an end-to-end supervisable architecture to perform all aforementioned audio tasks, consisting of a WaveNet synthesiser conditioned on the output of a jointly-trained cGAN spectrogram translator. In doing so, we demonstrate the potential of such flexible techniques to unify MIR tasks, promote efficient transfer learning, and converge research to the improvement of powerful, general methods. Finally, to the best of our knowledge, we present the first application of GANs to guided instrument synthesis.
Steven Spratley, Daniel Beck, Trevor Cohn
ICASSP2
2019 Accelerated Query Processing Via Similarity Score Prediction
abstract
Processing top-k bag-of-words queries is critical to many information retrieval applications, including web-scale search. In this work, we consider algorithmic properties associated with dynamic pruning mechanisms. Such algorithms maintain a score threshold (the k th highest similarity score identified so far) so that low-scoring documents can be bypassed, allowing fast top-k retrieval with no loss in effectiveness. In standard pruning algorithms the score threshold is initialized to the lowest possible value. To accelerate processing, we make use of term- and query-dependent features to predict the final value of that threshold, and then employ the predicted value right from the commencement of processing. Because of the asymmetry associated with prediction errors (if the estimated threshold is too high the query will need to be re-executed in order to assure the correct answer), the prediction process must be risk-sensitive. We explore techniques for balancing those factors, and provide detailed experimental results that show the practical usefulness of the new approach.
Matthias Petri, Alistair Moffat, Joel Mackenzie, J. Shane Culpepper, Daniel Beck
SIGIR5
2018 Graph-to-Sequence Learning using Gated Graph Neural Networks
abstract
Many NLP applications can be framed as a graph-to-sequence learning problem. Previous work proposing neural architectures on graph-to-sequence obtained promising results compared to grammar-based approaches but still rely on linearisation heuristics and/or standard recurrent networks to achieve the best performance. In this work propose a new model that encodes the full structural information contained in the graph. Our architecture couples the recently proposed Gated Graph Neural Networks with an input transformation that allows nodes and edges to have their own hidden representations, while tackling the parameter explosion problem present in previous work. Experimental results shows that our model outperforms strong baselines in generation from AMR graphs and syntax-based neural machine translation.
Daniel Beck, Gholamreza Haffari, Trevor Cohn
ACL (1)1
2016 Exploring Prediction Uncertainty in Machine Translation Quality Estimation
abstract
Machine Translation Quality Estimation is a notoriously difficult task, which lessens its usefulness in real-world translation environments.Such scenarios can be improved if quality predictions are accompanied by a measure of uncertainty.However, models in this task are traditionally evaluated only in terms of point estimate metrics, which do not take prediction uncertainty into account.We investigate probabilistic methods for Quality Estimation that can provide well-calibrated uncertainty estimates and evaluate them in terms of their full posterior predictive distributions.We also show how this posterior information can be useful in an asymmetric risk scenario, which aims to capture typical situations in translation workflows.
Daniel Beck, Lucia Specia, Trevor Cohn
CoNLL1
2016 Speed-Constrained Tuning for Statistical Machine Translation Using Bayesian Optimization
abstract
Daniel Beck, Adrià de Gispert, Gonzalo Iglesias, Aurelien Waite, Bill Byrne. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Daniel Beck, Adrià de Gispert, Gonzalo Iglesias, Aurelien Waite, William J. Byrne
HLT-NAACL1
2015 Seed: a user-friendly tool for exploring and visualizing microbial community data
abstract
SUMMARY: In this article we present Simple Exploration of Ecological Data (Seed), a data exploration tool for microbial communities. Seed is written in R using the Shiny library. This provides access to powerful R-based functions and libraries through a simple user interface. Seed allows users to explore ecological datasets using principal coordinate analyses, scatter plots, bar plots, hierarchal clustering and heatmaps. AVAILABILITY AND IMPLEMENTATION: Seed is open source and available at https://github.com/danlbek/Seed. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniel Beck, Christopher Dennis, James A. Foster
Bioinform.1
2015 Learning Structural Kernels for Natural Language Processing
abstract
Structural kernels are a flexible learning paradigm that has been widely used in Natural Language Processing. However, the problem of model selection in kernel-based methods is usually overlooked. Previous approaches mostly rely on setting default values for kernel hyperparameters or using grid search, which is slow and coarse-grained. In contrast, Bayesian methods allow efficient model selection by maximizing the evidence on the training data through gradient-based methods. In this paper we show how to perform this in the context of structural kernels by using Gaussian Processes. Experimental results on tree kernels show that this procedure results in better prediction performance compared to hyperparameter optimization via grid search. The framework proposed in this paper can be adapted to other structures besides trees, e.g., strings and graphs, thereby extending the utility of kernel-based methods.
Daniel Beck, Trevor Cohn, Christian Hardmeier, Lucia Specia
Trans. Assoc. Comput. Linguistics1
2014 Joint Emotion Analysis via Multi-task Gaussian Processes
abstract
We propose a model for jointly predicting multiple emotions in natural language sentences.Our model is based on a low-rank coregionalisation approach, which combines a vector-valued Gaussian Process with a rich parameterisation scheme.We show that our approach is able to learn correlations and anti-correlations between emotions on a news headlines dataset.The proposed model outperforms both singletask baselines and other multi-task approaches.
Daniel Beck, Trevor Cohn, Lucia Specia
EMNLP1
2014 GA-based selection of vaginal microbiome features associated with bacterial vaginosis
abstract
In this paper, we successfully apply GEFeS (Genetic & Evolutionary Feature Selection) to identify the key features in the human vaginal microbiome and in patient meta-data that are associated with bacterial vaginosis (BV). The vaginal microbiome is the community of bacteria found in a patient, and meta-data include behavioral practices and demographic information. Bacterial vaginosis is a disease that afflicts nearly one third of all women, but the current diagnostics are crude at best. We describe two types of classifies for BV diagnosis, and show that each is associated with one of two treatments. Our results show that the classifiers associated with the 'Treat Any Symptom' version have better performances that the classifier associated with the 'Treat Based on N-Score Value'. Our long term objective is to develop a more accurate and objective diagnosis and treatment of BV.
Joi Carter, Daniel Beck, Henry Williams, Gerry V. Dozier, James A. Foster
GECCO2
2012 mcaGUI: microbial community analysis R-Graphical User Interface (GUI)
abstract
UNLABELLED: Microbial communities have an important role in natural ecosystems and have an impact on animal and human health. Intuitive graphic and analytical tools that can facilitate the study of these communities are in short supply. This article introduces Microbial Community Analysis GUI, a graphical user interface (GUI) for the R-programming language (R Development Core Team, 2010). With this application, researchers can input aligned and clustered sequence data to create custom abundance tables and perform analyses specific to their needs. This GUI provides a flexible modular platform, expandable to include other statistical tools for microbial community analysis in the future. AVAILABILITY: The mcaGUI package and source are freely available as part of Bionconductor at http://www.bioconductor.org/packages/release/bioc/html/mcaGUI.html
Wade K. Copeland, Vandhana Krishnan, Daniel Beck, Matthew Settles, James A. Foster, Kyu-Chul Cho, Mitch Day, Roxana Hickey, Ursel M. E. Schütte, Christopher J. Williams, Larry J. Forney, Zaid Abdo
Bioinform.3
2011 OTUbase: an R infrastructure package for operational taxonomic unit data
abstract
SUMMARY: OTUbase is an R package designed to facilitate the analysis of operational taxonomic unit (OTU) data and sequence classification (taxonomic) data. Currently there are programs that will cluster sequence data into OTUs and/or classify sequence data into known taxonomies. However, there is a need for software that can take the summarized output of these programs and organize it into easily accessed and manipulated formats. OTUbase provides this structure and organization within R, to allow researchers to easily manipulate the data with the rich library of R packages currently available for additional analysis. AVAILABILITY: OTUbase is an R package available through Bioconductor. It can be found at http://www.bioconductor.org/packages/release/bioc/html/OTUbase.html.
Daniel Beck, Matthew Settles, James A. Foster
Bioinform.1
2009 Robust Collision Avoidance in Unknown Domestic Environments
Stefan Jacobs 0001, Alexander Ferrein, Stefan Schiffer 0002, Daniel Beck, Gerhard Lakemeyer
RoboCup4
2008 Landmark-Based Representations for Navigating Holonomic Soccer Robots
Daniel Beck, Alexander Ferrein, Gerhard Lakemeyer
RoboCup1
2007 A Simulation Environment for Middle-Size Robots with Multi-level Abstraction
Daniel Beck, Alexander Ferrein, Gerhard Lakemeyer
RoboCup1
2006 Automatic Testing and Evaluation of Multilingual Language Technology Resources and Components
Ulrich Schäfer 0001, Daniel Beck
LREC2