Niall M. Adams

dblp:76/6355 · DBLP profile ↗
← Back
42ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0003-0342-0513ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 2 since 2021Security and privacy · 7Applied, interdisciplinary, general and emerging computing · 6Graphics, computer vision, multimedia, augmented reality and games · 4Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 81% Information retrieval · 19%
Artificial intelligence
1 paper
Trustworthy machine learning · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
pattern mining
0.122003
An iterative hypothesis-testing strategy for pattern discovery · KDD 2003
Unsupervised Segmentation of Categorical Time Series into Episodes · ICDM 2002
Information retrieval › evaluation
statistical significance testing
0.012003
An iterative hypothesis-testing strategy for pattern discovery · KDD 2003
Data mining › temporal data mining
time series mining
0.012002
Unsupervised Segmentation of Categorical Time Series into Episodes · ICDM 2002
Data mining › time series analysis
time series segmentation
0.012002
Unsupervised Segmentation of Categorical Time Series into Episodes · ICDM 2002
Data mining › predictive modeling › classification
classifier evaluation
0.011999
The Impact of Changing Populations on Classifier Performance · KDD 1999
Data mining › anomaly detection
outlier detection
0.012003
An iterative hypothesis-testing strategy for pattern discovery · KDD 2003
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.011999
The Impact of Changing Populations on Classifier Performance · KDD 1999

Methods — techniques the papers use, named apart from their topics

classifier performance analysis · 0.0residual analysis · 0.0hypothesis testing · 0.0voting experts · 0.0boundary entropy · 0.0goal definition · 0.0
YearPublicationVenuePosition
2022 Estimation, Forecasting, and Anomaly Detection for Nonstationary Streams Using Adaptive Estimation
abstract
Streaming data provides substantial challenges for data analysis. From a computational standpoint, these challenges arise from constraints related to computer memory and processing speed. Statistically, the challenges relate to constructing procedures that can handle the so-called concept drift-the tendency of future data to have different underlying properties to current and historic data. The issue of handling structure, such as trend and periodicity, remains a difficult problem for streaming estimation. We propose the real-time adaptive component (RAC), a penalized-regression modeling framework that satisfies the computational constraints of streaming data, and provides the capability for dealing with concept drift. At the core of the estimation process are techniques from adaptive filtering. The RAC procedure adopts a specified basis to handle local structure, along with a least absolute shrinkage operator-like penalty procedure to handle over fitting. We enhance the RAC estimation procedure with a streaming anomaly detection capability. The experiments with simulated data suggest the procedure can be considered as a competitive tool for a variety of scenarios, and an illustration with real cyber-security data further demonstrates the promise of the method.
Henrique Hoeltgebaum, Niall M. Adams, Cristiano Fernandes
IEEE Trans. Cybern.2
2021 Streaming changepoint detection for transition matrices
abstract
Abstract Sequentially detecting multiple changepoints in a data stream is a challenging task. Difficulties relate to both computational and statistical aspects, and in the latter, specifying control parameters is a particular problem. Choosing control parameters typically relies on unrealistic assumptions, such as the distributions generating the data, and their parameters, being known. This is implausible in the streaming paradigm, where several changepoints will exist. Further, current literature is mostly concerned with streams of continuous-valued observations, and focuses on detecting a single changepoint. There is a dearth of literature dedicated to detecting multiple changepoints in transition matrices , which arise from a sequence of discrete states. This paper makes the following contributions: a complete framework is developed for adaptively and sequentially estimating a Markov transition matrix in the streaming data setting. A change detection method is then developed, using a novel moment matching technique, which can effectively monitor for multiple changepoints in a transition matrix. This adaptive detection and estimation procedure for transition matrices, referred to as ADEPT-M , is compared to several change detectors on synthetic data streams, and is implemented on two real-world data streams – one consisting of over nine million HTTP web requests, and the other being a well-studied electricity market data set.
Joshua Plasse, Henrique Hoeltgebaum, Niall M. Adams
Data Min. Knowl. Discov.3
2021 Correction to: Streaming changepoint detection for transition matrices
Joshua Plasse, Henrique Hoeltgebaum, Niall M. Adams
Data Min. Knowl. Discov.3
2021 Multi-type relational clustering for enterprise cyber-security networks
Elizabeth Riddle-Workman, Marina Evangelou, Niall M. Adams
Pattern Recognit. Lett.3
2020 An anomaly detection framework for cyber-security data
Marina Evangelou, Niall M. Adams
Comput. Secur.2
2018 A Study of Data Fusion for Predicting Novel Activity in Enterprise Cyber-Security
abstract
Modern computer networks allow for the collection of vast amounts of data. A wide variety of sources record data relating to different aspects of computer and network activity. This wealth of available data, coupled with the persistent rise in successful cyber-security breaches, motivates the need for data-driven approaches to complement existing cyber-defence systems. Although obtainable, most of this data remains unexploited due to issues of data collection and privacy concerns. The majority of research has therefore been constrained to utilise limited data sets, usually obtained from only one of the many available data sources. We use a recently assembled public domain data set, which associates data from multiple sources in a real-world enterprise computer network, to demonstrate the advantages of data and entity fusion for cyber-security. We formulate an anomaly detection task employing time-delayed labels, which enables the use of supervised learning as a means of predicting novel activity. Our results show that an appropriate fusion of data from multiple sources and entities improves predictive accuracy.
Jack Hogan, Niall M. Adams
ISI2
2018 Adaptive Anomaly Detection on Network Data Streams
abstract
As the number of cyber-attacks increases, there has been increasing emphasis on developing complementary methods of detection to the existing signature-based approaches. This work builds upon a previously discovered persistent structure within the Los Alamos National Laboratory network data sources, to develop a regression based streaming anomaly detection mechanism that can adapt to the network behaviour over time. The methodology has also been applied to a new data set of the same network to assess the extent of its pertinence in time.
Elizabeth Riddle-Workman, Marina Evangelou, Niall M. Adams
ISI3
2017 Clustering and monitoring edge behaviour in enterprise network traffic
abstract
This paper takes an unsupervised learning approach for monitoring edge activity within an enterprise computer network. Using NetFlow records, features are gathered across the active connections (edges) in 15-minute time windows. Then, edges are grouped into clusters using the k-means algorithm. This process is repeated over contiguous windows. A series of informative indicators are derived by examining the relationship of edges with the observed cluster structure. This leads to an intuitive method for monitoring network behaviour and a temporal description of edge behaviour at global and local levels.
Christopher Schon, Niall M. Adams, Marina Evangelou
ISI2
2016 Handling delayed labels in temporally evolving data streams
abstract
Streaming classification is well studied in the machine learning community. In real-world applications labels for previously observed feature vectors may only arrive after appreciable lag - that is, the labels are delayed. These delayed labels are an important aspect of streaming analysis, one that is not properly appreciated or addressed in the literature. This paper provides a taxonomy of delayed labeling and a framework for incorporating such labels into a streaming classifier. We provide a real-world demonstration of the utility of correctly handling delayed labels, in the context of a temporally adaptive linear classifier. This simple illustration shows that appropriately handling delayed labels can lead to an increase in performance, suggesting an opportunity for new research.
Joshua Plasse, Niall M. Adams
IEEE BigData2
2016 Predictability of NetFlow data
abstract
The behaviour of individual devices connected to an enterprise network can vary dramatically, as a device's activity depends on the user operating the device as well as on all behind the scenes operations between the device and the network. Being able to understand and predict a device's behaviour in a network can work as the foundation of an anomaly detection framework, as devices may show abnormal activity as part of a cyber attack. The aim of this work is the construction of a predictive regression model for a device's behaviour at normal state. The behaviour of a device is presented by a quantitative response and modelled to depend on historic data recorded by NetFlow.
Marina Evangelou, Niall M. Adams
ISI2
2016 Disassortativity of computer networks
abstract
Network data is ubiquitous in cyber-security applications. Accurately modelling such data allows discovery of anomalous edges, subgraphs or paths, and is key to many signature-free cyber-security analytics. We present a recurring property of graphs originating from cyber-security applications, often considered a `corner case' in the main literature on network data analysis, that greatly affects the performance of standard `off-the-shelf' techniques. This is the property that similarity, in terms of network behaviour, does not imply connectivity, and in fact the reverse is often true. We call this disassortivity. The phenomenon is illustrated using network flow data collected on an enterprise network, and we show how Big Data analytics designed to detect unusual connectivity patterns can be improved.
Patrick Rubin-Delanchy, Niall M. Adams, Nicholas A. Heard
ISI2
2016 Activity-based temporal anomaly detection in enterprise-cyber security
abstract
Statistical anomaly detection is emerging as an important complement to signature-based methods for enterprise network defence. In this paper, we isolate a persistent structure in two different enterprise network data sources. This structure provides the basis of a regression-based anomaly detection method. The procedure is demonstrated on a large public domain data set.
Mark Whitehouse, Marina Evangelou, Niall M. Adams
ISI3
2016 Improving clustering performance by incorporating uncertainty
Maha Bakoben, Anthony Bellotti, Niall M. Adams
Pattern Recognit. Lett.3
2013 When Does Active Learning Work?
Lewis Percival Gordon Evans, Niall M. Adams, Christoforos Anagnostopoulos
IDA2
2012 Exponentially weighted moving average charts for detecting concept drift
Gordon J. Ross, Niall M. Adams, Dimitris K. Tasoulis, David J. Hand
Pattern Recognit. Lett.2
2012 Erratum to "Exponentially weighted moving average charts for detecting concept drift" [Pattern Recognition Lett. 33(2) (2012) 191-198]
Gordon J. Ross, Niall M. Adams, Dimitris K. Tasoulis, David J. Hand
Pattern Recognit. Lett.2
2011 Recognizing players' activities and hidden state
abstract
This paper describes a machine learning approach to classifying the activities of players in games. Instances of activities generally are not identical because they play out in different contexts, so the challenge is to extract the "essences" of activities from instances. We show how this problem may be mapped to a sequence alignment problem, for which there are polynomial-time solutions. The method works well even when some features of activities are not observable (e.g., the emotional states of players). In fact, these features can in some conditions be inferred with high accuracy.
Wesley Kerr, Paul R. Cohen, Niall M. Adams
FDG3
2011 lambda-Perceptron: An adaptive classifier for data streams
Nicos G. Pavlidis, Dimitris K. Tasoulis, Niall M. Adams, David J. Hand
Pattern Recognit.3
2010 On-Line Adaptation of Exploration in the One-Armed Bandit with Covariates Problem
abstract
Many sequential decision making problems require an agent to balance exploration and exploitation to maximise long-term reward. Existing policies that address this tradeoff typically have parameters that are set a priori to control the amount of exploration. In finite-time problems, the optimal values of these parameters are highly dependent on the problem faced. In this paper, we propose adapting the amount of exploration performed on-line, as information is gathered by the agent. To this end we introduce a novel algorithm, e-ADAPT, which has no free parameters. The algorithm adapts as it plays and sequentially chooses whether to explore or exploit, driven by the amount of uncertainty in the system. We provide simulation results for the one armed bandit with covariates problem, which demonstrate the effectiveness of e-ADAPT to correctly control the amount of exploration in finite-time problems and yield rewards that are close to optimally tuned off-line policies. Furthermore, we show that e-ADAPT is robust to a high-dimensional covariate, as well as misspecified models. Finally, we describe how our methods could be extended to other sequential decision making problems, such as dynamic bandit problems with changing reward structures.
Adam M. Sykulski, Niall M. Adams, Nicholas R. Jennings
ICMLA2
2010 Changing the Focus of the IDA Symposium
Niall M. Adams, Paul R. Cohen, Michael R. Berthold
IDA1
2010 Streaming Covariance Selection with Applications to Adaptive Querying in Sensor Networks
abstract
Sensor networks can be naturally represented as graphical models, where the edge set encodes the presence of sparsity in the correlation structure between sensors. Such graphical representations can be valuable for information mining purposes as well as for optimizing bandwidth and battery usage with minimal loss of estimation accuracy. We use a computationally efficient technique for estimating sparse graphical models which fits a sparse linear regression locally at each node of the graph via the Lasso estimator. Using a recently suggested online, temporally adaptive implementation of the Lasso, we propose an algorithm for streaming graphical model selection over sensor networks. With battery consumption minimization applications in mind, we use this algorithm as the basis of an adaptive querying scheme. We discuss implementation issues in the context of environmental monitoring using sensor networks, where the objective is short-term forecasting of local wind direction. The algorithm is tested against real UK weather data and conclusions are drawn about certain tradeoffs inherent in decentralized sensor networks data analysis.
Christoforos Anagnostopoulos, Niall M. Adams, David J. Hand
Comput. J.2
2010 Prospects for Bandit Solutions in Sensor Management
abstract
Sensor management in information-rich and dynamic environments can be posed as a sequential action selection problem with side information. To study such problems we employ the dynamic multi-armed bandit with covariates framework. In this generalization of the multi-armed bandit, the expected rewards are time-varying linear functions of the covariate vector. The learning goal is to associate the covariate with the optimal action at each instance, essentially learning to partition the covariate space adaptively. Applications of sensor management are frequently in environments in which the precise nature of the dynamics is unknown. In such settings, the sensor manager tracks the evolving environment by observing only the covariates and the consequences of the selected actions. This creates difficulties not encountered in static problems, and changes the exploitation–exploration dilemma. We study the relationship between the different factors of the problem and provide interesting insights. The impact of the environment dynamics on the action selection problem is influenced by the covariate dimensionality. We present the surprising result that strategies that perform very little or no exploration perform surprisingly well in dynamic environments.
Nicos G. Pavlidis, Niall M. Adams, David Nicholson, David J. Hand
Comput. J.2
2009 Intelligent Data Analysis in the 21st Century
Paul R. Cohen, Niall M. Adams
IDA2
2009 Transaction aggregation as a strategy for credit card fraud detection
Christopher Whitrow, David J. Hand, Piotr Juszczak, David John Weston, Niall M. Adams
Data Min. Knowl. Discov.5
2008 Online optimization for variable selection in data streams
abstract
Variable selection for regression is a classical statistical problem, motivated by concerns that too many covariates invite overfitting. Existing approaches notably include a class of convex optimisation techniques, such as the Lasso algorithm. Such techniques are invariably reliant on assumptions that are unrealistic in streaming contexts, namely that the data is available off-line and the correlation structure is static. In this paper, we relax both these constraints, proposing for the first time an online implementation of the Lasso algorithm with exponential forgetting. We also optimise the model dimension and the speed of forgetting in an online manner, resulting in a fully automatic scheme. In simulations our scheme improves on recursive least squares in dynamic environments, while also featuring model discovery and changepoint detection capabilities.
Christoforos Anagnostopoulos, Dimitris K. Tasoulis, David J. Hand, Niall M. Adams
ECAI4
2008 The Design, Deployment and Evaluation of the AnimalWatch Intelligent Tutoring System
abstract
Europe and the U.S. both face the challenges of urban schools with low-achieving adolescent learners, many of whom are not proficient in the language of instruction. This paper describes the deployment and evaluation of the AnimalWatch intelligent tutoring system for mathematics in challenging classrooms. Previous studies demonstrated that AnimalWatch benefits 12–14 year-old students in relatively controlled conditions. The current study indicates that the system can help older, very low-achieving students in challenging secondary schools that serve diverse student populations.
Paul R. Cohen, Carole R. Beal, Niall M. Adams
ECAI3
2008 Dynamic Multi-Armed Bandit with Covariates
Nicos G. Pavlidis, Dimitris K. Tasoulis, Niall M. Adams, David J. Hand
ECAI3
2007 Visualising the Cluster Structure of Data Streams
Dimitris K. Tasoulis, Gordon J. Ross, Niall M. Adams
IDA3
2007 Voting experts: An unsupervised algorithm for segmenting sequences
Paul R. Cohen, Niall M. Adams, Brent Heeringa
Intell. Data Anal.2
2007 Quantitative Analysis of Cell Nucleus Organisation
abstract
There are almost 1,300 entries for higher eukaryotes in the Nuclear Protein Database. The proteins' subcellular distribution patterns within interphase nuclei can be complex, ranging from diffuse to punctate or microspeckled, yet they all work together in a coordinated and controlled manner within the three-dimensional confines of the nuclear volume. In this review we describe recent advances in the use of quantitative methods to understand nuclear spatial organisation and discuss some of the practical applications resulting from this work.
Carol Shiels, Niall M. Adams, Suhail A. Islam, David A. Stephens, Paul S. Freemont
PLoS Comput. Biol.2
2006 The Transcriptional Regulator CBP Has Defined Spatial Associations within Interphase Nuclei
abstract
It is becoming increasingly clear that nuclear macromolecules and macromolecular complexes are compartmentalized through binding interactions into an apparent three-dimensionally ordered structure. This ordering, however, does not appear to be deterministic to the extent that chromatin and nonchromatin structures maintain a strict 3-D arrangement. Rather, spatial ordering within the cell nucleus appears to conform to stochastic rather than deterministic spatial relationships. The stochastic nature of organization becomes particularly problematic when any attempt is made to describe the spatial relationship between proteins involved in the regulation of the genome. The CREB-binding protein (CBP) is one such transcriptional regulator that, when visualised by confocal microscopy, reveals a highly punctate staining pattern comprising several hundred individual foci distributed within the nuclear volume. Markers for euchromatic sequences have similar patterns. Surprisingly, in most cases, the predicted one-to-one relationship between transcription factor and chromatin sequence is not observed. Consequently, to understand whether spatial relationships that are not coincident are nonrandom and potentially biologically important, it is necessary to develop statistical approaches. In this study, we report on the development of such an approach and apply it to understanding the role of CBP in mediating chromatin modification and transcriptional regulation. We have used nearest-neighbor distance measurements and probability analyses to study the spatial relationship between CBP and other nuclear subcompartments enriched in transcription factors, chromatin, and splicing factors. Our results demonstrate that CBP has an order of spatial association with other nuclear subcompartments. We observe closer associations between CBP and RNA polymerase II-enriched foci and SC35 speckles than nascent RNA or specific acetylated histones. Furthermore, we find that CBP has a significantly higher probability of being close to its known in vivo substrate histone H4 lysine 5 compared with the closely related H4 lysine 12. This study demonstrates that complex relationships not described by colocalization exist in the interphase nucleus and can be characterized and quantified. The subnuclear distribution of CBP is difficult to reconcile with a model where chromatin organization is the sole determinant of the nuclear organization of proteins that regulate transcription but is consistent with a close link between spatial associations and nuclear functions.
Kirk J. McManus, David A. Stephens, Niall M. Adams, Suhail A. Islam, Paul S. Freemont, Michael J. Hendzel
PLoS Comput. Biol.3
2003 An iterative hypothesis-testing strategy for pattern discovery
abstract
Pattern discovery has emerged as a direct result of increased data storage and analytic capabilities available to the data analyst. Without a massive amount of data, we do not have the evidence to support the discovery of the local deterministic structures that we call patterns. As such, pattern discovery is one of the few areas of data mining that cannot be considered simply as a 'scaling-up' of current statistical methodology to analyze large data sets. However, the philosophies of hypothesis testing and modeling in traditional statistics do lend themselves to forming a framework for pattern discovery, and we can also draw from ideas relating to outlier discovery and residual analysis to discover patterns. We illustrate an iterative strategy in a statistical framework by way of its application to one simulated and two real data sets.
Richard J. Bolton, Niall M. Adams
KDD2
2002 Unsupervised Segmentation of Categorical Time Series into Episodes
abstract
This paper describes an unsupervised algorithm for segmenting categorical time series into episodes. The VOTING-EXPERTS algorithm first collects statistics about the frequency and boundary entropy of ngrams, then passes a window over the series and has two "expert methods" decide where in the window boundaries should be drawn. The algorithm successfully segments text into words in four languages. The algorithm also segments time series of robot sensor data into subsequences that represent episodes in the life of the robot. We claim that VOTING-EXPERTS finds meaningful episodes in categorical time series because it exploits two statistical characteristics of meaningful episodes.
Paul R. Cohen, Brent Heeringa, Niall M. Adams
ICDM3
2001 Robot Baby 2001
Paul R. Cohen, Tim Oates 0001, Niall M. Adams, Carole R. Beal
ALT3
2001 Robot Baby 2001
Paul R. Cohen, Tim Oates 0001, Niall M. Adams, Carole R. Beal
Discovery Science3
2001 An Algorithm for Segmenting Categorical Time Series into Meaningful Episodes
Paul R. Cohen, Niall M. Adams
IDA2
2001 The IDA'01 Robot Data Challenge
Paul R. Cohen, Niall M. Adams, David J. Hand
IDA2
2000 Improving the Practice of Classifier Performance Assessment
abstract
In this note we use examples from the literature to illustrate some poor practices in assessing the performance of supervised classification rules, and we suggest guidelines for better methodology. We also describe a new assessment criterion that is suitable for the needs of many practical problems.
Niall M. Adams, David J. Hand
Neural Comput.1
1999 Supervised Classification Problems: How to Be Both Judge and Jury
Mark G. Kelly, David J. Hand, Niall M. Adams
IDA3
1999 The Impact of Changing Populations on Classifier Performance
abstract
Article Free Access Share on The impact of changing populations on classifier performance Authors: Mark G. Kelly Department of Mathematics, Imperial College, 180 Queen's Gate, London, SW7 2BZ, UK Department of Mathematics, Imperial College, 180 Queen's Gate, London, SW7 2BZ, UKView Profile , David J. Hand Department of Mathematics, Imperial College, 180 Queen's Gate, London, SW7 2BZ, UK Department of Mathematics, Imperial College, 180 Queen's Gate, London, SW7 2BZ, UKView Profile , Niall M. Adams Department of Mathematics, Imperial College, 180 Queen's Gate, London, SW7 2BZ, UK Department of Mathematics, Imperial College, 180 Queen's Gate, London, SW7 2BZ, UKView Profile Authors Info & Claims KDD '99: Proceedings of the fifth ACM SIGKDD international conference on Knowledge discovery and data miningAugust 1999 Pages 367–371https://doi.org/10.1145/312129.312285Published:01 August 1999Publication History 125citation965DownloadsMetricsTotal Citations125Total Downloads965Last 12 Months110Last 6 weeks10 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Mark G. Kelly, David J. Hand, Niall M. Adams
KDD3
1999 Comparing classifiers when the misallocation costs are uncertain
Niall M. Adams, David J. Hand
Pattern Recognit.1
1998 Defining the Goals to Optimise Data Mining Performance
Mark G. Kelly, David J. Hand, Niall M. Adams
KDD3