Benjamin Denham

dblp:221/5109 · DBLP profile ↗
← Back
7ranked-venue papers
7as first author
3since 2021 · last 2024
0000-0002-5104-2361ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2024 Dynamic Quantification With Constrained Error Under Unknown General Dataset Shift
abstract
Quantification research has sought to accurately estimate class distributions under dataset shift. While existing methods perform well under assumed conditions of shift, it is not always clear whether such assumptions will hold in a given application. This work extends the analysis and experimental evaluation of our Gain-Some-Lose-Some (GSLS) model for quantification under general dataset shift and incorporates it into a method for dynamically selecting the most appropriate quantification method. Selection by a Kolmogorov-Smirnov test for any shift followed by a newly proposed “Adjusted Kolmogorov-Smirnov” test for non-prior shift is found to best balance quantification and runtime performance. We also present a framework for constraining quantification prediction intervals to user-specified limits by requesting a smaller set of instance class labels from the user than required with confidence-based rejection.
Benjamin Denham, Edmund M.-K. Lai, Roopak Sinha, Muhammad Asif Naeem
IEEE Trans. Knowl. Data Eng.1
2022 Witan: Unsupervised Labelling Function Generation for Assisted Data Programming
abstract
Effective supervised training of modern machine learning models often requires large labelled training datasets, which could be prohibitively costly to acquire for many practical applications. Research addressing this problem has sought ways to leverage weak supervision sources, such as the user-defined heuristic labelling functions used in the data programming paradigm, which are cheaper and easier to acquire. Automatic generation of these functions can make data programming even more efficient and effective. However, existing approaches rely on initial supervision in the form of small labelled datasets or interactive user feedback. In this paper, we propose Witan, an algorithm for generating labelling functions without any initial supervision. This flexibility affords many interaction modes, including unsupervised dataset exploration before the user even defines a set of classes. Experiments in binary and multi-class classification demonstrate the efficiency and classification accuracy of Witan compared to alternative labelling approaches.
Benjamin Denham, Edmund M.-K. Lai, Roopak Sinha, Muhammad Asif Naeem
Proc. VLDB Endow.1
2021 Gain-Some-Lose-Some: Reliable Quantification Under General Dataset Shift
abstract
When applying supervised learning to estimate class distributions of unlabelled samples (so-called quantification), dataset shift is an expected yet challenging problem. Existing quantification methods make strong assumptions on the nature of dataset shift that often will not hold in practice. We propose a novel Gain-Some-Lose-Some (GSLS) model that accounts for more general conditions of dataset shift. We present a method for fitting the GSLS model without any labelled instances from the target sample, and experimentally demonstrate that GSLS can produce reliable quantification prediction intervals under broader conditions of shift than existing quantification methods.
Benjamin Denham, Edmund M.-K. Lai, Roopak Sinha, Muhammad Asif Naeem
ICDM1
2020 Null-Labelling: A Generic Approach for Learning in the Presence of Class Noise
abstract
Class noise in datasets presents a significant challenge to accurate classification, requiring classifiers that can refuse to classify noisy instances. We demonstrate the inability of the popular confidence-thresholding rejection method to learn from relationships between input features and not-at-random class noise. To take advantage of these relationships, we propose a novel null-labelling scheme based on iterative re-training with relabelled datasets that enables a classifier to learn to reject instances that are likely to be misclassified. We demonstrate the ability of null-labelling to achieve a significantly better tradeoff between classification error and coverage than confidence-thresholding. Models generated by the null-labelling scheme have the added advantage of interpretability, in that they are able to identify features correlated with class noise. We also unify prior theories for combining and evaluating sets of rejecting classifiers.
Benjamin Denham, Russel Pears, Muhammad Asif Naeem
ICDM1
2020 Enhancing random projection with independent and cumulative additive noise for privacy-preserving data stream mining
Benjamin Denham, Russel Pears, Muhammad Asif Naeem
Expert Syst. Appl.1
2020 HDSM: A distributed data mining approach to classifying vertically distributed data streams
Benjamin Denham, Russel Pears, Muhammad Asif Naeem
Knowl. Based Syst.1
2018 Evaluating the Quality of Drupal Software Modules
abstract
Evaluating software modules for inclusion in a Drupal website is a crucial and complex task that currently requires manual assessment of a number of module facets. This study applied data-mining techniques to identify quality-related metrics associated with highly popular and unpopular Drupal modules. The data-mining approach produced a set of important metrics and thresholds that highlight a strong relationship between the overall perceived reliability of a module and its popularity. Areas for future research into open-source software quality are presented, including a proposed module evaluation tool to aid developers in selecting high-quality modules.
Benjamin Denham, Russel Pears, Andy M. Connor
Int. J. Softw. Eng. Knowl. Eng.1