Niall Twomey

dblp:55/10941 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0002-3225-2654ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 80% Learning theory · 20%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
class-conditional noise
0.812024
Hypothesis Testing for Class-Conditional Noise Using Local Maximum Likelihood · AAAI 2024
Machine learning › Learning theory
hypothesis testing
0.812024
Hypothesis Testing for Class-Conditional Noise Using Local Maximum Likelihood · AAAI 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Inherently Interpretable Time Series Classification via Multiple Instance Learning · ICLR 2024
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.812024
Hypothesis Testing for Class-Conditional Noise Using Local Maximum Likelihood · AAAI 2024
Machine learning › Trustworthy machine learning › interpretability
multiple instance learning explanation
0.812024
Inherently Interpretable Time Series Classification via Multiple Instance Learning · ICLR 2024
Information retrieval
cross-language information retrieval
0.512021
Backretrieval: An Image-Pivoted Evaluation Metric for Cross-Lingual Text Representations Without Parallel Corpora · SIGIR 2021
Information retrieval
evaluation
0.512021
Backretrieval: An Image-Pivoted Evaluation Metric for Cross-Lingual Text Representations Without Parallel Corpora · SIGIR 2021
Ubiquitous computing and smart environments
smart home
0.112018
Releasing eHealth Analytics into the Wild: Lessons Learnt from the SPHERE Project · KDD 2018

Methods — techniques the papers use, named apart from their topics

multiple instance learning · 0.8local maximum likelihood estimation · 0.8hypothesis testing · 0.8deep learning · 0.8sensor data analysis · 0.7machine learning deployment · 0.7image-text pairing · 0.5human study · 0.5
YearPublicationVenuePosition
2025 Multi-Teacher Knowledge Distillation for Efficient Object Segmentation
abstract
Segment Anything Model 2 (SAM2) has demonstrated state-of-the-art performance in image/video object segmentation across many domains, but its large encoder makes it challenging for resource-constrained devices or real-time applications. One solution to this problem is to carry out knowledge distillation from the bulky encoder to a lightweight encoder, but this can result in degraded performance. In this work, we investigate multi-teacher distillation to mitigate performance degradation for distilled segmentation models. Using several foundation teacher models, our multi-teacher distilled models achieve 3.2 times speedup during end-to-end inference compared to SAM2 while achieving the best results of 74.4 and 71.1 (72.1 and 69.6 for single-teacher distillation) mIoU on the COCO and LVIS image segmentation datasets, as well as showing competitive results on video segmentation. Our results show that multi-teacher distillation offers a powerful solution for efficient image/video segmentation, while also maintaining compelling performance.
Simon Zeng, Kurt Cutajar, Hanting Xie, Massimo Camplani, Richard Tomsett, Niall Twomey, Jas Kandola, Gavin K. C. Cheung
ICIP6
2024 Hypothesis Testing for Class-Conditional Noise Using Local Maximum Likelihood
abstract
In supervised learning, automatically assessing the quality of the labels before any learning takes place remains an open research question. In certain particular cases, hypothesis testing procedures have been proposed to assess whether a given instance-label dataset is contaminated with class-conditional label noise, as opposed to uniform label noise. The existing theory builds on the asymptotic properties of the Maximum Likelihood Estimate for parametric logistic regression. However, the parametric assumptions on top of which these approaches are constructed are often too strong and unrealistic in practice. To alleviate this problem, in this paper we propose an alternative path by showing how similar procedures can be followed when the underlying model is a product of Local Maximum Likelihood Estimation that leads to more flexible nonparametric logistic regression models, which in turn are less susceptible to model misspecification. This different view allows for wider applicability of the tests by offering users access to a richer model class. Similarly to existing works, we assume we have access to anchor points which are provided by the users. We introduce the necessary ingredients for the adaptation of the hypothesis tests to the case of nonparametric logistic regression and empirically compare against the parametric approach presenting both synthetic and real-world case studies and discussing the advantages and limitations of the proposed approach.
Weisong Yang, Rafael Poyiadzi, Niall Twomey, Raúl Santos-Rodríguez
AAAI3
2024 Inherently Interpretable Time Series Classification via Multiple Instance Learning
abstract
Conventional Time Series Classification (TSC) methods are often black boxes that obscure inherent interpretation of their decision-making processes. In this work, we leverage Multiple Instance Learning (MIL) to overcome this issue, and propose a new framework called MILLET: Multiple Instance Learning for Locally Explainable Time series classification. We apply MILLET to existing deep learning TSC models and show how they become inherently interpretable without compromising (and in some cases, even improving) predictive performance. We evaluate MILLET on 85 UCR TSC datasets and also present a novel synthetic dataset that is specially designed to facilitate interpretability evaluation. On these datasets, we show MILLET produces sparse explanations quickly that are of higher quality than other well-known interpretability methods. To the best of our knowledge, our work with MILLET is the first to develop general MIL methods for TSC and apply them to an extensive variety of domains.
Joseph Early, Gavin K. C. Cheung, Kurt Cutajar, Hanting Xie, Jas Kandola, Niall Twomey
ICLR6
2022 Equitable Ability Estimation in Neurodivergent Student Populations with Zero-Inflated Learner Models
Niall Twomey, Sarah McMullan, Anat Elhalal, Rafael Poyiadzi, Luis Miguel Vaquero González
EDM1
2022 Hypothesis Testing for Class-Conditional Label Noise
Rafael Poyiadzi, Weisong Yang, Niall Twomey, Raúl Santos-Rodríguez
ECML/PKDD (3)3
2021 Backretrieval: An Image-Pivoted Evaluation Metric for Cross-Lingual Text Representations Without Parallel Corpora
abstract
Cross-lingual text representations have gained popularity lately and act as the backbone of many tasks such as unsupervised machine translation and cross-lingual information retrieval, to name a few. However, evaluation of such representations is difficult in the domains beyond standard benchmarks due to the necessity of obtaining domain-specific parallel language data across different pairs of languages. In this paper, we propose an automatic metric for evaluating the quality of cross-lingual textual representations using images as a proxy in a paired image-text evaluation dataset. Experimentally, Backretrieval is shown to highly correlate with ground truth metrics on annotated datasets, and our analysis shows statistically significant improvements over baselines. Our experiments conclude with a case study on a recipe dataset without parallel cross-lingual data. We illustrate how to judge cross-lingual embedding quality with Backretrieval, and validate the outcome with a small human study.
Mikhail Fain, Niall Twomey, Danushka Bollegala
SIGIR2
2020 Neural ODEs with Stochastic Vector Field Mixtures
abstract
It was recently shown that neural ordinary differential equation models cannot solve fundamental and seemingly straightforward tasks even with high-capacity vector field representations. This paper introduces two other fundamental tasks to the set that baseline methods cannot solve, and proposes mixtures of stochastic vector fields as a model class that is capable of solving these essential problems. Dynamic vector field selection is of critical importance for our model, and our approach is to propagate component uncertainty over the integration interval with a technique based on forward filtering. We also formalise several loss functions that encourage desirable properties on the trajectory paths, and of particular interest are those that directly encourage fewer expected function evaluations. Experimentally, we demonstrate that our model class is capable of capturing the natural dynamics of human behaviour; a notoriously volatile application area. Baseline approaches cannot model this problem.
Niall Twomey, Michal Kozlowski, Raúl Santos-Rodríguez
ECAI1
2020 Towards Multi-Language Recipe Personalisation and Recommendation
abstract
Multi-language recipe personalisation and recommendation is an under-explored field of information retrieval in academic and production systems. The existing gaps in our current understanding are numerous, even on fundamental questions such as whether consistent and high-quality recipe recommendation can be delivered across languages. Motivated by this need, we consider the multi-language recipe recommendation setting and present grounding results that will help to establish the potential and absolute value of future work in this area. Our work draws on several billion events from millions of recipes, with published recipes and users incorporating several languages, including Arabic, English, Indonesian, Russian, and Spanish. We represent recipes using a combination of normalised ingredients, standardised skills and image embeddings obtained without human intervention. In modelling, we take a classical approach based on optimising an embedded bi-linear user-item metric space towards the interactions that most strongly elicit cooking intent. For users without interaction histories, a bespoke content-based cold-start model that predicts context and recipe affinity is introduced. We show that our approach to personalisation is stable and scales well to new languages. A robust cross-validation campaign is employed and consistently rejects baseline models and representations, strongly favouring those we propose. Our results are presented in a language-oriented (as opposed to model-oriented) fashion to emphasise the language-based goals of this work. We believe that this is the first large-scale work that evaluates the value and potential of multi-language recipe recommendation and personalisation.
Niall Twomey, Mikhail Fain, Andrey Ponikar, Nadine Sarraf
RecSys1
2020 Energy-efficient activity recognition framework using wearable accelerometers
Atis Elsts, Niall Twomey, Ryan McConville, Ian Craddock
J. Netw. Comput. Appl.2
2019 Active Learning with Label Proportions
abstract
Active Learning (AL) refers to the setting where the learner has the ability to perform queries to an oracle to acquire the true label of an instance or, sometimes, a set of instances. Even though Active Learning has been studied extensively, the setting is usually restricted to assume that the oracle is trustworthy and will provide the actual label. We argue that, while common, this approach can be made more flexible to account for different forms of supervision. In this paper, we propose a new framework that allows the algorithm to request the label for a bag of samples at a time. Although this label will come in the form of proportions of class labels in the bags and therefore encode less information, we demonstrate that we can still learn effectively.
Rafael Poyiadzi, Raúl Santos-Rodríguez, Niall Twomey
ICASSP3
2019 An application of hierarchical Gaussian processes to the detection of anomalies in star light curves
Niall Twomey, Haoyan Chen, Tom Diethe, Peter A. Flach
Neurocomputing1
2018 Anomaly detection in star light curves using hierarchical Gaussian processes
Haoyan Chen, Tom Diethe, Niall Twomey, Peter A. Flach
ESANN3
2018 Person Identification and Discovery With Wrist Worn Accelerometer Data
Ryan McConville, Raúl Santos-Rodríguez, Niall Twomey
ESANN3
2018 Efficient approximate representations for computationally expensive features
Raúl Santos-Rodríguez, Niall Twomey
ESANN2
2018 On-Board Feature Extraction from Acceleration Data for Activity Recognition
Atis Elsts, Ryan McConville, Xenofon Fafoutis, Niall Twomey, Robert J. Piechocki, Raúl Santos-Rodríguez, Ian Craddock
EWSN4
2018 Releasing eHealth Analytics into the Wild: Lessons Learnt from the SPHERE Project
abstract
The SPHERE project is devoted to advancing eHealth in a smart-home context, and supports full-scale sensing and data analysis to enable a generic healthcare service. We describe, from a data-science perspective, our experience of taking the system out of the laboratory into more than thirty homes in Bristol, UK. We describe the infrastructure and processes that had to be developed along the way, describe how we train and deploy Machine Learning systems in this context, and give a realistic appraisal of the state of the deployed systems.
Tom Diethe, Mike Holmes, Meelis Kull, Miquel Perelló-Nieto, Kacper Sokol, Hao Song 0007, Emma Tonkin, Niall Twomey, Peter A. Flach
KDD8
2017 Unsupervised learning of sensor topologies for improving activity recognition in smart environments
Niall Twomey, Tom Diethe, Ian Craddock, Peter A. Flach
Neurocomputing1
2016 Active transfer learning for activity recognition
Tom Diethe, Niall Twomey, Peter A. Flach
ESANN2
2016 On the need for structure modelling in sequence prediction
abstract
There is no uniform approach in the literature for modelling sequential correlations in sequence classification problems. It is easy to find examples of unstructured models ( e.g. logistic regression) where correlations are not taken into account at all, but there are also many examples where the correlations are explicitly incorporated into a—potentially computationally expensive—structured classification model ( e.g. conditional random fields). In this paper we lay theoretical and empirical foundations for clarifying the types of problem which necessitate direct modelling of correlations in sequences, and the types of problem where unstructured models that capture sequential aspects solely through features are sufficient. The theoretical work in this paper shows that the rate of decay of auto-correlations within a sequence is related to the excess classification risk that is incurred by ignoring the structural aspect of the data. This is an intuitively appealing result, demonstrating the intimate link between the auto-correlations and excess classification risk. Drawing directly on this theory, we develop well-founded visual analytics tools that can be applied a priori on data sequences and we demonstrate how these tools can guide practitioners in specifying feature representations based on auto-correlation profiles. Empirical analysis is performed on three sequential datasets. With baseline feature templates, structured and unstructured models achieve similar performance, indicating no initial preference for either model. We then apply the visual analytics tools to the datasets, and show that classification performance in all cases is improved over baseline results when our tools are involved in defining feature representations.
Niall Twomey, Tom Diethe, Peter A. Flach
Mach. Learn.1
2015 Bayesian Modelling of the Temporal Aspects of Smart Home Activity with Circular Statistics
Tom Diethe, Niall Twomey, Peter A. Flach
ECML/PKDD (2)2
2014 A Machine Learning Approach to Objective Cardiac Event Detection
abstract
This paper presents an automated framework for the detection of the QRS complex from Electrocardiogram (ECG) signals. We introduce an artefact-tolerant pre-processing algorithm which emphasises a number of characteristics of the ECG that are representative of the QRS complex. With this processed ECG signal we train Logistic Regression and Support Vector Machine classification models. With our approach we obtain over 99.7% detection sensitivity and precision on the MIT-BIH database without using supplementary de-noising or pre-emphasis filters.
Niall Twomey, Peter A. Flach
CISIS1
2014 Automated Detection of Perturbed Cardiac Physiology During Oral Food Allergen Challenge in Children
abstract
This paper investigates the fully automated computer-based detection of allergic reaction in oral food challenges using pediatric ECG signals. Nonallergic background is modeled using a mixture of Gaussians during oral food challenges, and the model likelihoods are used to determine whether a subject is allergic to a food type. The system performance is assessed on the dataset of 24 children (15 allergic and 9 nonallergic) totaling 34 h of data. The proposed detector correctly classified all nonallergic subjects (100% specificity) and 12 allergic subjects (80% sensitivity) and is capable of detecting allergy on average 17 min earlier than trained clinicians during oral food challenges, the gold standard of allergy diagnosis. Inclusion of the developed allergy classification platform during oral food challenges recorded would result in a 30% reduction of doses administered to allergic subjects. The results of study introduce the possibility to halt challenges earlier which can safely advance the state of clinical art of allergy diagnosis by reducing the overall exposure to the allergens.
Niall Twomey, Andriy Temko, Jonathan Hourihane, William P. Marnane
IEEE J. Biomed. Health Informatics1