Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daniel Ward

dblp:211/4049 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-2571-3114ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 39% Language models and text generation · 26% Learning theory · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive estimation
0.912025
SoftCVI: Contrastive variational inference with self-generated soft labels · ICLR 2025
Natural language and speech › Language models and text generation
decoding
0.912025
The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation · ICLR 2025
Natural language and speech › Language models and text generation › text generation
open-ended text generation
0.912025
The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation · ICLR 2025
Machine learning › Learning theory
overfitting
0.912025
The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.912025
SoftCVI: Contrastive variational inference with self-generated soft labels · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
model misspecification
0.612022
Robust Neural Posterior Estimation and Statistical Model Criticism · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference › simulation-based inference
neural posterior estimation
0.612022
Robust Neural Posterior Estimation and Statistical Model Criticism · NeurIPS 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
Robust Neural Posterior Estimation and Statistical Model Criticism · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference
0.612022
Robust Neural Posterior Estimation and Statistical Model Criticism · NeurIPS 2022
Bioinformatics and computational biology › genomics
primer design
0.412019
PrimedRPA: primer design for recombinase polymerase amplification assays · Bioinform. 2019
Bioinformatics and computational biology › metagenomics
pathogen detection
0.112019
PrimedRPA: primer design for recombinase polymerase amplification assays · Bioinform. 2019

Methods — techniques the papers use, named apart from their topics

normalizing flow · 0.9markov chain monte carlo · 0.9greedy decoding · 0.9fine-tuning · 0.9contrastive estimation · 0.9robust neural posterior estimation · 0.6model criticism · 0.6sequence alignment · 0.4cross-reactivity filtering · 0.4
YearPublicationVenuePosition
2025 The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation
abstract
This paper introduces the counter-intuitive generalization results of overfitting pre-trained large language models (LLMs) on very small datasets. In the setting of open-ended text generation, it is well-documented that LLMs tend to generate repetitive and dull sequences, a phenomenon that is especially apparent when generating using greedy decoding. This issue persists even with state-of-the-art LLMs containing billions of parameters, trained via next-token prediction on large datasets. We find that by further fine-tuning these models to achieve a near-zero training loss on a small set of samples -- a process we refer to as hyperfitting -- the long-sequence generative capabilities are greatly enhanced. Greedy decoding with these Hyperfitted models even outperform Top-P sampling over long-sequences, both in terms of diversity and human preferences. This phenomenon extends to LLMs of various sizes, different domains, and even autoregressive image generation. We further find this phenomena to be distinctly different from that of Grokking and double descent. Surprisingly, our experiments indicate that hyperfitted models rarely fall into repeating sequences they were trained on, and even explicitly blocking these sequences results in high-quality output. All hyperfitted models produce extremely low-entropy predictions, often allocating nearly all probability to a single token.
Fredrik Carlsson, Fangyu Liu 0001, Daniel Ward, Murathan Kurfali, Joakim Nivre
ICLR3
2025 SoftCVI: Contrastive variational inference with self-generated soft labels
abstract
Estimating a distribution given access to its unnormalized density is pivotal in Bayesian inference, where the posterior is generally known only up to an unknown normalizing constant. Variational inference and Markov chain Monte Carlo methods are the predominant tools for this task; however, both are often challenging to apply reliably, particularly when the posterior has complex geometry. Here, we introduce Soft Contrastive Variational Inference (SoftCVI), which allows a family of variational objectives to be derived through a contrastive estimation framework. The approach parameterizes a classifier in terms of a variational distribution, reframing the inference task as a contrastive estimation problem aiming to identify a single true posterior sample among a set of samples. Despite this framing, we do not require positive or negative samples, but rather learn by sampling the variational distribution and computing ground truth soft classification labels from the unnormalized posterior itself. The objectives have zero variance gradient when the variational approximation is exact, without the need for specialized gradient estimators. We empirically investigate the performance on a variety of Bayesian inference tasks, using both simple (e.g. normal) and expressive (normalizing flow) variational distributions. We find that SoftCVI can be used to form objectives which are stable to train and mass-covering, frequently outperforming inference with other variational approaches.
Daniel Ward, Mark Beaumont, Matteo Fasiolo
ICLR1
2022 Robust Neural Posterior Estimation and Statistical Model Criticism
abstract
Computer simulations have proven a valuable tool for understanding complex phenomena across the sciences. However, the utility of simulators for modelling and forecasting purposes is often restricted by low data quality, as well as practical limits to model fidelity. In order to circumvent these difficulties, we argue that modellers must treat simulators as idealistic representations of the true data generating process, and consequently should thoughtfully consider the risk of model misspecification. In this work we revisit neural posterior estimation (NPE), a class of algorithms that enable black-box parameter inference in simulation models, and consider the implication of a simulation-to-reality gap. While recent works have demonstrated reliable performance of these methods, the analyses have been performed using synthetic data generated by the simulator model itself, and have therefore only addressed the well-specified case. In this paper, we find that the presence of misspecification, in contrast, leads to unreliable inference when NPE is used naïvely. As a remedy we argue that principled scientific inquiry with simulators should incorporate a model criticism component, to facilitate interpretable identification of misspecification and a robust inference component, to fit ‘wrong but useful’ models. We propose robust neural posterior estimation (RNPE), an extension of NPE to simultaneously achieve both these aims, through explicitly modelling the discrepancies between simulations and the observed data. We assess the approach on a range of artificially misspecified examples, and find RNPE performs well across the tasks, whereas naïvely using NPE leads to misleading and erratic posteriors.
Daniel Ward, Patrick Cannon, Mark Beaumont, Matteo Fasiolo, Sebastian M. Schmon
NeurIPS1
2022 COVID-profiler: a webserver for the analysis of SARS-CoV-2 sequencing data
abstract
BACKGROUND: SARS-CoV-2 virus sequencing has been applied to track the COVID-19 pandemic spread and assist the development of PCR-based diagnostics, serological assays, and vaccines. With sequencing becoming routine globally, bioinformatic tools are needed to assist in the robust processing of resulting genomic data. RESULTS: We developed a web-based bioinformatic pipeline ("COVID-Profiler") that inputs raw or assembled sequencing data, displays raw alignments for quality control, annotates mutations found and performs phylogenetic analysis. The pipeline software can be applied to other (re-) emerging pathogens. CONCLUSIONS: The webserver is available at http://genomics.lshtm.ac.uk/ . The source code is available at https://github.com/jodyphelan/covid-profiler .
Jody Phelan, Wouter Deelder, Daniel Ward, Susana G. Campino, Martin L. Hibberd, Taane G. Clark
BMC Bioinform.3
2020 Temporally Coherent Embeddings for Self-Supervised Video Representation Learning
abstract
This paper presents TCE: Temporally Coherent Embeddings for self-supervised video representation learning. The proposed method exploits inherent structure of unlabeled video data to explicitly enforce temporal coherency in the embedding space, rather than indirectly learning it through ranking or predictive proxy tasks. In the same way that high-level visual information in the world changes smoothly, we believe that nearby frames in learned representations will benefit from demonstrating similar properties. Using this assumption, we train our TCE model to encode videos such that adjacent frames exist close to each other and videos are separated from one another. Using TCE we learn robust representations from large quantities of unlabeled video data. We thoroughly analyse and evaluate our self-supervised learned TCE models on a downstream task of video action recognition using multiple challenging benchmarks (Kinetics400, UCF101, HMDB51). With a simple but effective 2D-CNN backbone and only RGB stream inputs, TCE pre-trained representations outperform all previous self-supervised 2D-CNN and 3D-CNN pre-trained on UCF101. The code and pre-trained models for this paper can be downloaded at: https://github.com/csiro-robotics/TCE.
Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, Peyman Moghadam
ICPR3
2020 Scalable learning for bridging the species gap in image-based plant phenotyping
Daniel Ward, Peyman Moghadam
Comput. Vis. Image Underst.1
2019 PrimedRPA: primer design for recombinase polymerase amplification assays
abstract
SUMMARY: Recombinase polymerase amplification (RPA), an isothermal nucleic acid amplification method, is enhancing our ability to detect a diverse array of pathogens, thereby assisting the diagnosis of infectious diseases and the detection of microorganisms in food and water. However, new bioinformatics tools are needed to automate and improve the design of the primers and probes sets to be used in RPA, particularly to account for the high genetic diversity of circulating pathogens and cross detection of genetically similar organisms. PrimedRPA is a python-based package that automates the creation and filtering of RPA primers and probe sets. It aligns several sequences to identify conserved targets, and filters regions that cross react with possible background organisms. AVAILABILITY AND IMPLEMENTATION: PrimedRPA was implemented in Python 3 and supported on Linux and MacOS and is freely available from http://pathogenseq.lshtm.ac.uk/PrimedRPA.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matthew Higgins, Matt Ravenhall, Daniel Ward, Jody Phelan, Amy Ibrahim, Matthew S. Forrest, Taane G. Clark, Susana G. Campino
Bioinform.3
2018 Deep Leaf Segmentation Using Synthetic Data
Daniel Ward, Peyman Moghadam, Nicolas Hudson
BMVC1