VLDB 2026 Research / reviewers in the wild / expert
Tom Diethe
dblp:33/1098 · also Thomas Robert Diethe
· DBLP profile ↗
25ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-0776-5407ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 25% Vision and language · 20% Generative modeling · 19% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 90% Computational complexity · 10% | |
| Network and information security
2 papers |
Privacy and data protection · 100% | |
| Human-computer interaction and pervasive computing
2 papers |
Ubiquitous computing and smart environments · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 70% Medical and health informatics · 30% |
Topics — the 30 heaviest of 43, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.6 | 2 | 2025 | Diffusion Instruction Tuning · ICML 2025 Tackling Structural Hallucination in Image Translation with Local Diffusion · ECCV (81) 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
1.0 | 2 | 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces · ICML 2025 Distribution calibration for regression · ICML 2019 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.9 | 1 | 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
cross-modal transfer |
0.9 | 1 | 2025 | Diffusion Instruction Tuning · ICML 2025 |
Natural language and speech › Language models and text generation › large language model inference
inference-time techniques |
0.9 | 1 | 2025 | Balancing Act: Diversity and Consistency in Large Language Model Ensembles · ICLR 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | Diffusion Instruction Tuning · ICML 2025 |
Computer vision › Vision and language › visual grounding
language-guided segmentation |
0.9 | 1 | 2025 | Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation · ICML 2025 |
Natural language and speech › Language models and text generation › large language model
large language model ensemble |
0.9 | 1 | 2025 | Balancing Act: Diversity and Consistency in Large Language Model Ensembles · ICLR 2025 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
latent space bayesian optimization |
0.9 | 1 | 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces · ICML 2025 |
Natural language and speech › Language models and text generation › LLM agents › LLM collaboration
mixture of agents |
0.9 | 1 | 2025 | Balancing Act: Diversity and Consistency in Large Language Model Ensembles · ICLR 2025 |
Computer vision › Segmentation and scene understanding › open-world segmentation
open-set segmentation |
0.9 | 1 | 2025 | Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation · ICML 2025 |
Natural language and speech › Language models and text generation › decoding
self-consistency decoding |
0.9 | 1 | 2025 | Balancing Act: Diversity and Consistency in Large Language Model Ensembles · ICLR 2025 |
Machine learning › Generative modeling
variational autoencoder |
0.9 | 1 | 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces · ICML 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Diffusion Instruction Tuning · ICML 2025 |
Privacy and data protection
differential privacy |
0.8 | 2 | 2020 | Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations · WSDM 2020 Leveraging Hierarchical Representations for Preserving Privacy and Utility in Text · ICDM 2019 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | Improving Antibody Humanness Prediction using Patent Data · ICML 2024 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.8 | 1 | 2024 | Tackling Structural Hallucination in Image Translation with Local Diffusion · ECCV (81) 2024 |
Machine learning › Generative modeling
image translation |
0.8 | 1 | 2024 | Tackling Structural Hallucination in Image Translation with Local Diffusion · ECCV (81) 2024 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.8 | 1 | 2024 | An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning · ICML 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning · ICML 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
weakly-supervised contrastive learning |
0.8 | 1 | 2024 | Improving Antibody Humanness Prediction using Patent Data · ICML 2024 |
Bioinformatics and computational biology › protein design
antibody design |
0.8 | 1 | 2024 | Improving Antibody Humanness Prediction using Patent Data · ICML 2024 |
Mathematical optimization
combinatorial optimization |
0.8 | 1 | 2024 | Measures of diversity and space-filling designs for categorical data · ICML 2024 |
Mathematical optimization
discrete optimization |
0.8 | 1 | 2024 | Measures of diversity and space-filling designs for categorical data · ICML 2024 |
Mathematical optimization › optimization
diversity maximization |
0.8 | 1 | 2024 | Measures of diversity and space-filling designs for categorical data · ICML 2024 |
Mathematical optimization › sparse learning
feature selection |
0.8 | 1 | 2024 | Measures of diversity and space-filling designs for categorical data · ICML 2024 |
Mathematical optimization
submodular optimization |
0.8 | 1 | 2024 | Measures of diversity and space-filling designs for categorical data · ICML 2024 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.4 | 1 | 2020 | Optimal Continual Learning has Perfect Memory and is NP-hard · ICML 2020 |
Machine learning › Learning paradigms
continual learning |
0.4 | 1 | 2020 | Optimal Continual Learning has Perfect Memory and is NP-hard · ICML 2020 |
Privacy and data protection › location privacy
geo-indistinguishability |
0.4 | 1 | 2020 | Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations · WSDM 2020 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 0.9stable diffusion · 0.9prompt regularization · 0.9mixture refinement · 0.9ensemble gating · 0.9diffusion model · 0.9cross-validation · 0.9cross-attention · 0.9bayesian update · 0.9attention alignment · 0.9submodular optimization · 0.8multi-stage training · 0.8multi-loss training · 0.8linear programming relaxation · 0.8cross-entropy loss · 0.8approximation algorithm · 0.8sensor data analysis · 0.7machine learning deployment · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Balancing Act: Diversity and Consistency in Large Language Model EnsemblesabstractEnsembling strategies for Large Language Models (LLMs) have demonstrated significant potential in improving performance across various tasks by combining the strengths of individual models. However, identifying the most effective ensembling method remains an open challenge, as neither maximizing output consistency through self-consistency decoding nor enhancing model diversity via frameworks like "Mixture of Agents" has proven universally optimal. Motivated by this, we propose a unified framework to examine the trade-offs between task performance, model diversity, and output consistency in ensembles. More specifically, we introduce a consistency score that defines a gating mechanism for mixtures of agents and an algorithm for mixture refinement to investigate these trade-offs at the semantic and model levels, respectively. We incorporate our insights into a novel inference-time LLM ensembling strategy called the Dynamic Mixture of Agents (DMoA) and demonstrate that it achieves a new state-of-the-art result in the challenging Big Bench Hard mixed evaluations benchmark. Our analysis reveals that cross-validation bias can enhance performance, contingent on the expertise of the constituent models. We further demonstrate that distinct reasoning tasks—such as arithmetic reasoning, commonsense reasoning, and instruction following—require different model capabilities, leading to inherent task-dependent trade-offs that DMoA balances effectively. Ahmed Abdulaal, Nina Montaña Brown, Aryo Pradipta Gema, Daniel C. Castro, Daniel C. Alexander, Philip Teare, Tom Diethe, Dino Oglic, Amrutha Saseendran |
ICLR | 8 |
| 2025 | Diffusion Instruction TuningabstractWe introduce Lavender, a simple supervised fine-tuning (SFT) method that boosts the performance of advanced vision-language models (VLMs) by leveraging state-of-the-art image generation models such as Stable Diffusion. Specifically, Lavender aligns the text-vision attention in the VLM transformer with the equivalent used by Stable Diffusion during SFT, instead of adapting separate encoders. This alignment enriches the model’s visual understanding and significantly boosts performance across in- and out-of-distribution tasks. Lavender requires just 0.13 million training examples—2.5% of typical large-scale SFT datasets—and fine-tunes on standard hardware (8 GPUs) in a single day. It consistently improves state-of-the-art open-source multimodal LLMs (e.g., Llama-3.2-11B, MiniCPM-Llama3-v2.5), achieving up to 30% gains and a 68% boost on challenging out-of-distribution medical QA tasks. By efficiently transferring the visual expertise of image generators with minimal supervision, Lavender offers a scalable solution for more accurate vision-language systems. Code, training data, and models are available on the project page. Ryutaro Tanno, Amrutha Saseendran, Tom Diethe, Philip Teare |
ICML | 4 |
| 2025 | Segment Anyword: Mask Prompt Inversion for Open-Set Grounded SegmentationabstractOpen-set image segmentation poses a significant challenge because existing methods often demand extensive training or fine-tuning and generally struggle to segment unified objects consistently across diverse text reference expressions. Motivated by this, we propose Segment Anyword, a novel training-free visual concept prompt learning approach for open-set language grounded segmentation that relies on token-level cross-attention maps from a frozen diffusion model to produce segmentation surrogates or *mask prompts*, which are then refined into targeted object masks. Initial prompts typically lack coherence and consistency as the complexity of the image-text increases, resulting in suboptimal mask fragments. To tackle this issue, we further introduce a novel linguistic-guided visual prompt regularization that binds and clusters visual prompts based on sentence dependency and syntactic structural information, enabling the extraction of robust, noise-tolerant mask prompts, and significant improvements in segmentation accuracy. The proposed approach is effective, generalizes across different open-set segmentation tasks, and achieves state-of-the-art results of 52.5 (+6.8 relative) mIoU on Pascal Context 59, 67.73 (+25.73 relative) cIoU on gRefCOCO, and 67.4 (+1.1 relative to fine-tuned methods) mIoU on GranDf, which is the most complex open-set grounded segmentation task in the field. Amrutha Saseendran, Xilin He, Fariba Yousefi, Nikolay Burlutskiy, Dino Oglic, Tom Diethe, Philip Teare, Huiyu Zhou 0001 |
ICML | 8 |
| 2025 | Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured SpacesabstractBayesian optimisation in the latent space of a VAE is a powerful framework for optimisation tasks over complex structured domains, such as the space of valid molecules. However, existing approaches tightly couple the surrogate and generative models, which can lead to suboptimal performance when the latent space is not tailored to specific tasks, which in turn has led to the proposal of increasingly sophisticated algorithms. In this work, we explore a new direction, instead proposing a decoupled approach that trains a generative model and a GP surrogate separately, then combines them via a simple yet principled Bayesian update rule. This separation allows each component to focus on its strengths— structure generation from the VAE and predictive modelling by the GP. We show that our decoupled approach improves our ability to identify high-potential candidates in molecular optimisation problems under constrained evaluation budgets. Henry B. Moss, Sebastian W. Ober, Tom Diethe |
ICML | 3 |
| 2024 | Tackling Structural Hallucination in Image Translation with Local Diffusion
Seunghoi Kim, Tom Diethe, Matteo Figini, Henry F. J. Tregidgo, Asher Mullokandov, Philip Teare, Daniel C. Alexander |
ECCV (81) | 3 |
| 2024 | An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt LearningabstractTextural Inversion, a prompt learning method, learns a singular text embedding for a new "word" to represent image style and appearance, allowing it to be integrated into natural language sentences to generate novel synthesised images. However, identifying multiple unknown object-level concepts within one scene remains a complex challenge. While recent methods have resorted to cropping or masking individual images to learn multiple concepts, these techniques often require prior knowledge of new concepts and are labour-intensive. To address this challenge, we introduce *Multi-Concept Prompt Learning (MCPL)*, where multiple unknown "words" are simultaneously learned from a single sentence-image pair, without any imagery annotations. To enhance the accuracy of word-concept correlation and refine attention mask boundaries, we propose three regularisation techniques: *Attention Masking*, *Prompts Contrastive Loss*, and *Bind Adjective*. Extensive quantitative comparisons with both real-world categories and biomedical images demonstrate that our method can learn new semantically disentangled concepts. Our approach emphasises learning solely from textual embeddings, using less than 10% of the storage space compared to others. The project page, code, and data are available at [https://astrazeneca.github.io/mcpl.github.io](https://astrazeneca.github.io/mcpl.github.io). Ryutaro Tanno, Amrutha Saseendran, Tom Diethe, Philip Teare |
ICML | 4 |
| 2024 | Measures of diversity and space-filling designs for categorical dataabstractSelecting a small subset of items that represent the diversity of a larger population lies at the heart of many data analysis and machine learning applications. However, when it comes to items described by discrete features, the lack of natural ordering and the combinatorial nature of the search space pose significant challenges to the current selection techniques and make existing methods ill-suited. In this paper, we propose to make a step in that direction by proposing novel methods to select subsets of diverse categorical data based on the advances in combinatorial optimization. First, we start to cast the subset selection problem through the lens of the optimization of three diversity metrics. We then provide novel bounds for this problem and present exact solvers that unfortunately come with a high computational cost. To overcome this bottleneck, we go on and show how to employ tools from linear programming and submodular optimization by introducing two computationally plausible methods that still present approximation guarantees about the diversity metrics. Finally, a numerical assessment is provided to illustrate the potential of the designs with respect to state-of-the-art methods. Cédric Malherbe, Emilio Domínguez-Sánchez, Merwan Barlier, Igor Colin, Haitham Bou-Ammar, Tom Diethe |
ICML | 6 |
| 2024 | Improving Antibody Humanness Prediction using Patent DataabstractWe investigate the potential of patent data for improving the antibody humanness prediction using a multi-stage, multi-loss training process. Humanness serves as a proxy for the immunogenic response to antibody therapeutics, one of the major causes of attrition in drug discovery and a challenging obstacle for their use in clinical settings. We pose the initial learning stage as a weakly-supervised contrastive-learning problem, where each antibody sequence is associated with possibly multiple identifiers of function and the objective is to learn an encoder that groups them according to their patented properties. We then freeze a part of the contrastive encoder and continue training it on the patent data using the cross-entropy loss to predict the humanness score of a given antibody sequence. We illustrate the utility of the patent data and our approach by performing inference on three different immunogenicity datasets, unseen during training. Our empirical results demonstrate that the learned model consistently outperforms the alternative baselines and establishes new state-of-the-art on five out of six inference tasks, irrespective of the used metric. Talip Ucar, Aubin Ramon, Dino Oglic, Rebecca Croasdale-Wood, Tom Diethe, Pietro Sormanni |
ICML | 5 |
| 2020 | Optimal Continual Learning has Perfect Memory and is NP-hardabstractContinual Learning (CL) algorithms incrementally learn a predictor or representation across multiple sequentially observed tasks. Designing CL algorithms that perform reliably and avoid so-called catastrophic forgetting has proven a persistent challenge. The current paper develops a theoretical approach that explains why. In particular, we derive the computational properties which CL algorithms would have to possess in order to avoid catastrophic forgetting. Our main finding is that such optimal CL algorithms generally solve an NP-hard problem and will require perfect memory to do so. The findings are of theoretical interest, but also explain the excellent performance of CL algorithms using experience replay, episodic memory and core sets relative to regularization-based approaches. Jeremias Knoblauch, Hisham Husain, Tom Diethe |
ICML | 3 |
| 2020 | Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate PerturbationsabstractAccurately learning from user data while providing quantifiable privacy guarantees provides an opportunity to build better ML models while maintaining user trust. This paper presents a formal approach to carrying out privacy preserving text perturbation using the notion of d_χ-privacy designed to achieve geo-indistinguishability in location data. Our approach applies carefully calibrated noise to vector representation of words in a high dimension space as defined by word embedding models. We present a privacy proof that satisfies d_χ-privacy where the privacy parameter $\varepsilon$ provides guarantees with respect to a distance metric defined by the word embedding space. We demonstrate how $\varepsilon$ can be selected by analyzing plausible deniability statistics backed up by large scale analysis on GloVe and fastText embeddings. We conduct privacy audit experiments against $2$ baseline models and utility experiments on 3 datasets to demonstrate the tradeoff between privacy and utility for varying values of varepsilon on different task types. Our results demonstrate practical utility (< 2% utility loss for training binary classifiers) while providing better privacy guarantees than baseline models. Oluwaseyi Feyisetan, Borja Balle, Thomas Drake, Tom Diethe |
WSDM | 4 |
| 2020 | Automatic Discovery of Privacy-Utility Pareto FrontsabstractAbstract Differential privacy is a mathematical framework for privacy-preserving data analysis. Changing the hyperparameters of a differentially private algorithm allows one to trade off privacy and utility in a principled way. Quantifying this trade-off in advance is essential to decision-makers tasked with deciding how much privacy can be provided in a particular application while maintaining acceptable utility. Analytical utility guarantees offer a rigorous tool to reason about this tradeoff, but are generally only available for relatively simple problems. For more complex tasks, such as training neural networks under differential privacy, the utility achieved by a given algorithm can only be measured empirically. This paper presents a Bayesian optimization methodology for efficiently characterizing the privacy– utility trade-off of any differentially private algorithm using only empirical measurements of its utility. The versatility of our method is illustrated on a number of machine learning tasks involving multiple models, optimizers, and datasets. Brendan Avent, Javier González 0002, Tom Diethe, Andrei Paleyes, Borja Balle |
Proc. Priv. Enhancing Technol. | 3 |
| 2019 | $β^3$-IRT: A New Item Response Model and its ApplicationsabstractItem Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the $\beta^3$-IRT model, which models continuous responses and can generate a much enriched family of Item Characteristic Curves. In experiments we applied the proposed model to data from an online exam platform, and show our model outperforms a more standard 2PL-ND model on all datasets. Furthermore, we show how to apply $\beta^3$-IRT to assess the ability of machine learning classifiers.This novel application results in a new metric for evaluating the quality of the classifier’s probability estimates, based on the inferred difficulty and discrimination of data instances. Telmo de Menezes e Silva Filho, Ricardo B. C. Prudêncio, Tom Diethe, Peter A. Flach |
AISTATS | 4 |
| 2019 | Leveraging Hierarchical Representations for Preserving Privacy and Utility in TextabstractGuaranteeing a certain level of user privacy in an arbitrary piece of text is a challenging issue. However, with this challenge comes the potential of unlocking access to vast data stores for training machine learning models and supporting data driven decisions. We address this problem through the lens of dx-privacy, a generalization of Differential Privacy to non Hamming distance metrics. In this work, we explore word representations in Hyperbolic space as a means of preserving privacy in text. We provide a proof satisfying dx-privacy, then we define a probability distribution in Hyperbolic space and describe a way to sample from it in high dimensions. Privacy is provided by perturbing vector representations of words in high dimensional Hyperbolic space to obtain a semantic generalization. We conduct a series of experiments to demonstrate the tradeoff between privacy and utility. Our privacy experiments illustrate protections against an authorship attribution algorithm while our utility experiments highlight the minimal impact of our perturbations on several downstream machine learning models. Compared to the Euclidean baseline, we observe > 20x greater guarantees on expected privacy against comparable worst case statistics. Oluwaseyi Feyisetan, Tom Diethe, Thomas Drake |
ICDM | 2 |
| 2019 | Distribution calibration for regressionabstractWe are concerned with obtaining well-calibrated output distributions from regression models. Such distributions allow us to quantify the uncertainty that the model has regarding the predicted target value. We introduce the novel concept of distribution calibration, and demonstrate its advantages over the existing definition of quantile calibration. We further propose a post-hoc approach to improving the predictions from previously trained regression models, using multi-output Gaussian Processes with a novel Beta link function. The proposed method is experimentally verified on a set of common regression models and shows improvements for both distribution-level and quantile-level calibration. Hao Song 0007, Tom Diethe, Meelis Kull, Peter A. Flach |
ICML | 2 |
| 2019 | An application of hierarchical Gaussian processes to the detection of anomalies in star light curves
Niall Twomey, Haoyan Chen, Tom Diethe, Peter A. Flach |
Neurocomputing | 3 |
| 2018 | Anomaly detection in star light curves using hierarchical Gaussian processes
Haoyan Chen, Tom Diethe, Niall Twomey, Peter A. Flach |
ESANN | 2 |
| 2018 | Releasing eHealth Analytics into the Wild: Lessons Learnt from the SPHERE ProjectabstractThe SPHERE project is devoted to advancing eHealth in a smart-home context, and supports full-scale sensing and data analysis to enable a generic healthcare service. We describe, from a data-science perspective, our experience of taking the system out of the laboratory into more than thirty homes in Bristol, UK. We describe the infrastructure and processes that had to be developed along the way, describe how we train and deploy Machine Learning systems in this context, and give a realistic appraisal of the state of the deployed systems. Tom Diethe, Mike Holmes, Meelis Kull, Miquel Perelló-Nieto, Kacper Sokol, Hao Song 0007, Emma Tonkin, Niall Twomey, Peter A. Flach |
KDD | 1 |
| 2017 | Unsupervised learning of sensor topologies for improving activity recognition in smart environments
Niall Twomey, Tom Diethe, Ian Craddock, Peter A. Flach |
Neurocomputing | 2 |
| 2016 | Active transfer learning for activity recognition
Tom Diethe, Niall Twomey, Peter A. Flach |
ESANN | 1 |
| 2016 | ADL™: A Topic Model for Discovery of Activities of Daily Living in a Smart Home
Tom Diethe, Peter A. Flach |
IJCAI | 2 |
| 2016 | On the need for structure modelling in sequence predictionabstractThere is no uniform approach in the literature for modelling sequential correlations in sequence classification problems. It is easy to find examples of unstructured models ( e.g. logistic regression) where correlations are not taken into account at all, but there are also many examples where the correlations are explicitly incorporated into a—potentially computationally expensive—structured classification model ( e.g. conditional random fields). In this paper we lay theoretical and empirical foundations for clarifying the types of problem which necessitate direct modelling of correlations in sequences, and the types of problem where unstructured models that capture sequential aspects solely through features are sufficient. The theoretical work in this paper shows that the rate of decay of auto-correlations within a sequence is related to the excess classification risk that is incurred by ignoring the structural aspect of the data. This is an intuitively appealing result, demonstrating the intimate link between the auto-correlations and excess classification risk. Drawing directly on this theory, we develop well-founded visual analytics tools that can be applied a priori on data sequences and we demonstrate how these tools can guide practitioners in specifying feature representations based on auto-correlation profiles. Empirical analysis is performed on three sequential datasets. With baseline feature templates, structured and unstructured models achieve similar performance, indicating no initial preference for either model. We then apply the visual analytics tools to the datasets, and show that classification performance in all cases is improved over baseline results when our tools are involved in defining feature representations. Niall Twomey, Tom Diethe, Peter A. Flach |
Mach. Learn. | 2 |
| 2015 | Bayesian Modelling of the Temporal Aspects of Smart Home Activity with Circular Statistics
Tom Diethe, Niall Twomey, Peter A. Flach |
ECML/PKDD (2) | 1 |
| 2013 | Online Learning with (Multiple) Kernels: A ReviewabstractThis review examines kernel methods for online learning, in particular, multiclass classification. We examine margin-based approaches, stemming from Rosenblatt's original perceptron algorithm, as well as nonparametric probabilistic approaches that are based on the popular gaussian process framework. We also examine approaches to online learning that use combinations of kernels--online multiple kernel learning. We present empirical validation of a wide range of methods on a protein fold recognition data set, where different biological feature types are available, and two object recognition data sets, Caltech101 and Caltech256, where multiple feature spaces are available in terms of different image feature extraction methods. Tom Diethe, Mark A. Girolami |
Neural Comput. | 1 |
| 2010 | Constructing Nonlinear Discriminants from Multiple Data Views
Tom Diethe, David R. Hardoon, John Shawe-Taylor |
ECML/PKDD (1) | 1 |
| 2009 | Kernel Polytope Faces Pursuit
Tom Diethe, Zakria Hussain |
ECML/PKDD (1) | 1 |