Dan Hudson 0001

dblp:274/2353 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-2917-4659ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Theory of computation · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A comparison of graph construction techniques for applying graph signal processing to soil moisture networks
abstract
Abstract Sensor networks let a farmer keep their eye on multiple locations in an agricultural field simultaneously but can be expensive to install, maintain and analyse. Furthermore, sensors often suffer from gaps in the recording process which leads to missing data points or what are essentially ‘blind spots’ in the network structure. To cater for missing values, effective methods for data imputation are essential. In this paper, we use graphs to impute these missing values within sensor networks using a technique called graph signal processing (GSP) applied to soil moisture recordings. Using this method, we simulate network conditions involving missing sensors or inconsistently collected data. This enables farmers to reliably estimate the sensor readings that would have been obtained, thereby increasing the fault tolerance of their agricultural sensor networks. In this work, we are specifically interested in the relative accuracy of data imputation between several graph construction techniques within the GSP framework, both geometric, i.e., dependent on the geographical coordinates, and data-driven techniques, e.g., correlations between the sensor readings. We evaluated seven graph construction techniques, also comparing with a simple mean imputation baseline, for creating edges. By masking sensor values, we identify how accurately sensor values can be inferred. This is done by gradually masking sensors from the network with 1000 random sensor combinations per mask size and then imputing these “missing” sensors. For our experiments, we make use of the Cook Agronomy Farm (CAF) dataset for GSP imputation that contains soil moisture data recorded with 42 sensors. At almost at every timestamp not even once all moisture sensors recorded the data simultaneously, showcasing the value of correct data imputation in these sparse sensor networks. Our results indicate that data-driven graphs, that connect nodes (e.g., sensors) based on the underlying sensor recordings, tend to capture the relationships between sensors most accurately, where the data-driven Gaussian kernel graph (a signal similarity approach) consistently outperforms other graphs on average with 15% improvement across all experiments. Furthermore, compared to a simple baseline, error reduces between 50 and 70% depending on the underlying data. This suggests that the Gaussian kernel graph can function as a solid enhancement in applying GSP when sensors networks are either prone to faults or sparsely placed. Additional analysis showed that the interplay between graph density, signal smoothness and structural connectivity should be balanced for optimal performance.
Jurgen van den Hoogen, Dan Hudson 0001, Martin Atzmüller
Discov. Comput.2
2024 Graph Signal Processing Unearths the Best Locations for Soil Moisture Sensors
abstract
In this paper, we apply graph signal processing to optimise a soil moisture sensor network by identifying and re-moving redundant sensors. We evaluated which of seven proposed graph construction techniques best models the relationships between soil moisture measurements at different places in an agricultural field. Here, we consider a sensor location to be redundant if the moisture value can be imputed from information elsewhere in the graph. We gradually remove redundant sensors from the network in a top-down manner, imputing the masked sensors using Tikhonov minimisation - looking for the graph structure that gives us the most accurate imputed values. Our results indicate that the thresholded Gaussian kernel has the best performance in terms of error, while Delaunay triangulation, a parameter-free method, performs similarly. Furthermore, as expected, it seems that the edge sensors are most important while sensors close-by each other or in the centre of the field seem to be less relevant.
Jurgen van den Hoogen, Dan Hudson 0001, Martin Atzmüller
ICMLA2
2023 Hyperparameter Analysis of Wide-Kernel CNN Architectures in Industrial Fault Detection - An Exploratory Study
abstract
In recent years, industrial fault detection has become more data-driven due to advancements in automated data analysis using Deep Learning (DL). These techniques facilitate meaningful feature extraction, e.g., in time series data retrieved from sensors, which is typically of complex nature. This enables effective fault detection and prognostics, which increases efficiency and productivity of industrial equipment. However, the optimal settings for these DL architectures are generally use-case specific.
Jurgen van den Hoogen, Dan Hudson 0001, Stefan Bloemheuvel, Martin Atzmüller
DSAA2
2023 Enhanced Explanations for Knowledge-Augmented Clustering using Subgroup Discovery
abstract
Contemporary machine learning techniques are capable of extracting complex structure from data in a way that complements or exceeds manual examination, yet, as is welldocumented, many of these techniques suffer from a lack of interpretability. This paper extends previous work on explainable and interpretable machine learning, in particular on the ‘Knowledge-Augmented Clusters (KnAC)’ approach, allowing human users to benefit from uninterpretable ‘black box’ models to extract structure from datasets by clustering and to make this better understandable. One of the key functions of KnAC is to relate expert-annotated clusters to clusters that have been identified by a machine learning method, and then provide a comprehensible explanation, thus clarifying the relationships that KnAC discovered. Our novel contribution in this paper is to examine the usefulness of subgroup discovery as a way to generate comprehensible explanations within KnAC, and to compare this to the existing approach based on the XAI algorithm Anchors through a detailed evaluation. We find that the approach using subgroup discovery performs equally or better in our extensive experimentation testing this on six different datasets.
Maciej Szelazek, Dan Hudson 0001, Szymon Bobek, Grzegorz J. Nalepa, Martin Atzmüller
DSAA2
2021 Local Exceptionality Detection in Time Series Using Subgroup Discovery: An Approach Exemplified on Team Interaction Data
Dan Hudson 0001, Travis J. Wiltshire, Martin Atzmüller
DS1