EDBT 2026 Demo / reviewers in the wild / expert
Russel Pears
dblp:71/4680
· DBLP profile ↗
52ranked-venue papers
7as first author
7since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 23 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 8Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Novel method for optimizing performance in resource constrained distributed data streams
Rashi Bhalla, Russel Pears, Muhammad Asif Naeem, Farhaan Mirza |
Appl. Intell. | 2 |
| 2022 | Analyzing and repairing concept drift adaptation in data stream classification
Ben Halstead, Yun Sing Koh, Patricia J. Riddle, Russel Pears, Mykola Pechenizkiy, Albert Bifet, Gustavo Olivares, Guy Coulson |
Mach. Learn. | 4 |
| 2021 | Analyzing and Repairing Concept Drift Adaptation in Data Stream ClassificationabstractData collected over time often exhibit changes in distribution, or concept drift, caused by changes in hidden context relevant to the classification task, e.g. weather conditions. Adaptive learning methods are able to retain performance in changing conditions by explicitly detecting concept drift and changing the classifier used to make predictions. However, in realworld conditions, existing methods often select classifiers which poorly represent current data due to adaptation errors, where change in context is misidentified. We propose the AiRStream system, which uses a novel repair algorithm to identify and correct adaptation errors. We identify errors by periodically testing the performance of inactive classifiers. If an error is identified, a backtracking procedure repairs training done under the misidentified context. AiRStream achieves higher accuracy compared to baseline methods and selects classifiers which better match changes in context. A case study on a real-world air quality inference task shows that AiRStream is able to build a robust model of environmental conditions, allowing the adaptions made to concept drift to be analysed and related to changes in weather. Ben Halstead, Yun Sing Koh, Patricia J. Riddle, Russel Pears, Mykola Pechenizkiy, Albert Bifet, Gustavo Olivares, Guy Coulson |
DSAA | 4 |
| 2021 | Fingerprinting Concepts in Data Streams with Supervised and Unsupervised Meta-InformationabstractStreaming sources of data are becoming more common as the ability to collect data in real-time grows. A major concern in dealing with data streams is concept drift, a change in the distribution of data over time, for example, due to changes in environmental conditions. Representing concepts (stationary periods featuring similar behaviour) is a key idea in adapting to concept drift. By testing the similarity of a concept representation to a window of observations, we can detect concept drift to a new or previously seen recurring concept. Concept representations are constructed using meta-information features, values describing aspects of concept behaviour. We find that previously proposed concept representations rely on small numbers of meta-information features. These representations often cannot distinguish concepts, leaving systems vulnerable to concept drift. We propose FiCSUM, a general framework to represent both supervised and unsupervised behaviours of a concept in a fingerprint, a vector of many distinct meta-information features able to uniquely identify more concepts. Our dynamic weighting strategy learns which meta-information features describe concept drift in a given dataset, allowing a diverse set of meta-information features to be used at once. FiCSUM outperforms state-of-the-art methods over a range of 11 real world and synthetic datasets in both accuracy and modeling underlying concept drift. Ben Halstead, Yun Sing Koh, Patricia J. Riddle, Mykola Pechenizkiy, Albert Bifet, Russel Pears |
ICDE | 6 |
| 2021 | Recurring concept memory management in data streams: exploiting data stream concept evolution to improve performance and transparency
Ben Halstead, Yun Sing Koh, Patricia J. Riddle, Russel Pears, Mykola Pechenizkiy, Albert Bifet |
Data Min. Knowl. Discov. | 4 |
| 2021 | Topical affinity in short text microblogs
Herman Masindano Wandabwa, Muhammad Asif Naeem, Farhaan Mirza, Russel Pears |
Inf. Syst. | 4 |
| 2021 | Multi-interest semantic changes over time in short-text microblogs
Herman Masindano Wandabwa, Muhammad Asif Naeem, Farhaan Mirza, Russel Pears |
Knowl. Based Syst. | 4 |
| 2020 | Null-Labelling: A Generic Approach for Learning in the Presence of Class NoiseabstractClass noise in datasets presents a significant challenge to accurate classification, requiring classifiers that can refuse to classify noisy instances. We demonstrate the inability of the popular confidence-thresholding rejection method to learn from relationships between input features and not-at-random class noise. To take advantage of these relationships, we propose a novel null-labelling scheme based on iterative re-training with relabelled datasets that enables a classifier to learn to reject instances that are likely to be misclassified. We demonstrate the ability of null-labelling to achieve a significantly better tradeoff between classification error and coverage than confidence-thresholding. Models generated by the null-labelling scheme have the added advantage of interpretability, in that they are able to identify features correlated with class noise. We also unify prior theories for combining and evaluating sets of rejecting classifiers. Benjamin Denham, Russel Pears, Muhammad Asif Naeem |
ICDM | 2 |
| 2020 | Automating Inspection of Moveable Lane Barrier for Auckland Harbour Bridge Traffic Safety
Boris Bacic, Munish Rathee, Russel Pears |
ICONIP (1) | 3 |
| 2020 | Enhancing random projection with independent and cumulative additive noise for privacy-preserving data stream mining
Benjamin Denham, Russel Pears, Muhammad Asif Naeem |
Expert Syst. Appl. | 2 |
| 2020 | HDSM: A distributed data mining approach to classifying vertically distributed data streams
Benjamin Denham, Russel Pears, Muhammad Asif Naeem |
Knowl. Based Syst. | 2 |
| 2018 | The incremental Fourier classifier: Leveraging the discrete Fourier transform for classifying high speed data streams
Chamari I. Kithulgoda, Russel Pears, Muhammad Asif Naeem |
Expert Syst. Appl. | 2 |
| 2018 | Evaluating the Quality of Drupal Software ModulesabstractEvaluating software modules for inclusion in a Drupal website is a crucial and complex task that currently requires manual assessment of a number of module facets. This study applied data-mining techniques to identify quality-related metrics associated with highly popular and unpopular Drupal modules. The data-mining approach produced a set of important metrics and thresholds that highlight a strong relationship between the overall perceived reliability of a module and its popularity. Areas for future research into open-source software quality are presented, including a proposed module evaluation tool to aid developers in selecting high-quality modules. Benjamin Denham, Russel Pears, Andy M. Connor |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2017 | Detecting precursor patterns for frequency fluctuation in an electrical gridabstractPrecursor pattern identification addresses the problem of detecting warning signals in data that herald an impending event of extraordinary interest. In the context of electrical power systems, identifying precursors to fluctuations in power generation in advance would enable engineers to put in place measures that mitigate against the effects of such fluctuations. In this research we use the Morlet wavelet to transform a time series defined on electrical power generation frequency which was sampled at intervals of 30 seconds to identify potential precursor patterns. The power spectrum that results is then used to select high coefficient regions that capture a large faction of the energy in the spectrum. We then subjected the high coefficient regions together with a contrasting low coefficient region to a non-parametric ANOVA test and our results indicate that one high coefficient region dominates by predicting an overwhelming percentage of the variation that occurs during the subsequent fluctuation event. These results suggest that the wavelet is an effective mechanism to identify precursor activity in electricity time series data. Md. Shahidul Islam, Russel Pears, Boris Bacic |
SNPD | 2 |
| 2017 | A composite spatio-temporal modeling approach for age invariant face recognition
Fahad Bashir Alvi, Russel Pears |
Expert Syst. Appl. | 2 |
| 2016 | Staged Online Learning: A new approach to classification in high speed data streamsabstractIn this research we present a new framework and associated algorithms for mining high speed data streams that take advantage of concept recurrence. Different from previous work our approach detects volatility in a stream and then matches the learning paradigm to the degree of volatility. In high volatility stream segments a decision forest is used as the learning mechanism whereas in low volatility segments an approach driven by the use of stored Fourier spectra are used for learning. Our empirical results show significant processing and memory advantages over an approach that does not take advantage of volatility in the stream. Chamari I. Kithulgoda, Russel Pears |
IJCNN | 2 |
| 2016 | Capturing recurring concepts using discrete Fourier transformabstractSummary In this research, we address the problem of capturing recurring concepts in a data stream environment. Recurrence capture enables the reuse of previously learned classifiers without the need for relearning while providing for better accuracy during the concept recurrence interval. We capture concepts by applying the discrete Fourier transform to decision tree classifiers to obtain highly compressed versions of the trees at concept drift points in the stream and store such trees in a repository for future use. In addition, the impact of drift detector in enabling stable performance is also studied with the two drift detectors: ADWIN and SeqDrift2 in recurring concept capturing context. Our empirical results on real world and synthetic data exhibiting varying degrees of recurrence show that the Fourier compressed trees are more robust to noise and are able to capture recurring concepts with higher precision than a meta‐learning approach that chooses to reuse classifiers in their originally occurring form. A case study on a flight dataset that closely matches the target data stream environment where concepts recur in similar form in a time critical system is also conducted and the benefits of discrete Fourier transform application is evaluated. Copyright © 2016 John Wiley & Sons, Ltd. Sripirakas Sakthithasan, Russel Pears |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | HI-Tree: Mining High Influence Patterns Using External and Internal Utility Values
Yun Sing Koh, Russel Pears |
DaWaK | 2 |
| 2015 | Use of ensembles of Fourier spectra in capturing recurrent concepts in data streamsabstractIn this research, we apply ensembles of Fourier encoded spectra to capture and mine recurring concepts in a data stream environment. Previous research showed that compact versions of Decision Trees can be obtained by applying the Discrete Fourier Transform to accurately capture recurrent concepts in a data stream. However, in highly volatile environments where new concepts emerge often, the approach of encoding each concept in a separate spectrum is no longer viable due to memory overload and thus in this research we present an ensemble approach that addresses this problem. Our empirical results on real world data and synthetic data exhibiting varying degrees of recurrence reveal that the ensemble approach outperforms the single spectrum approach in terms of classification accuracy, memory and execution time. Sripirakas Sakthithasan, Russel Pears, Albert Bifet, Bernhard Pfahringer |
IJCNN | 2 |
| 2015 | A graph based approach to inferring item weights for pattern mining
Russel Pears, Songwut Pisalpanus, Yun Sing Koh |
Expert Syst. Appl. | 1 |
| 2015 | Goal-oriented dynamic test generation
TheAnh Do, Siau-Cheng Khoo, Russel Pears, Thanh Tho Quan |
Inf. Softw. Technol. | 4 |
| 2014 | Mining Recurrent Concepts in Data Streams Using the Discrete Fourier Transform
Sripirakas Sakthithasan, Russel Pears |
DaWaK | 2 |
| 2014 | Web Materialization Formulation: Modelling Feasible Solutions
Srdan Zagorac, Russel Pears |
DEXA (2) | 2 |
| 2014 | Detecting Volatility Shift in Data StreamsabstractCurrent drift detection techniques detect a change in distribution within a stream. However, there are no current techniques that analyze the change in the rate of these detected changes. We coin the term stream volatility, to describe the rate of changes in a stream. A stream has a high volatility if changes are detected frequently and has a low volatility if changes are detected infrequently. We are particularly interested in a volatility shift which is a change in the rate of change (e.g. From high volatility to low volatility). We introduce and define the concept of stream volatility, and propose a novel technique to detect volatility on data streams in the presence of concept drifts. In the experiments we show our algorithm to be both fast and efficient. We also propose a new algorithm for drift detection called SEED that is faster and more memory efficient than the existing state-of-the-art drift detection approach. A faster drift detection algorithm has a flow-on benefit to the subsequent volatility detection stage because both algorithms run concurrently on the data stream. David Tse Jung Huang, Yun Sing Koh, Gillian Dobbie, Russel Pears |
ICDM | 4 |
| 2014 | Extracting temporal knowledge from time series: A case study in ecological dataabstractThis research presents a generic framework and methods for mining temporal rules from multiple time-series data and its application to ecological data. The aphids dataset that tracks the trajectory of aphid infestations over time has been well researched in a number of studies. Those studies concentrated on predicting the scale of infestation over time. The focus of our research is to identify environmental factors that predict, in a temporal fashion, high incidence of aphid activity. This required the development of a novel framework for knowledge extraction from multiple time-series data and a method for discretization of numeric data as well-known methods such as SAX did not perform adequately due to the non-Gaussian nature of the data involved. Our experimentation yielded new insights into the environmental factors that may influence pest outbreak which are captured in the form of simple actionable rules that would be of interest to the farming community. Reggio N. Hartono, Russel Pears, Nikola K. Kasabov, Susan P. Worner |
IJCNN | 2 |
| 2014 | Detecting Changes in Rare Patterns from Data Streams
David Tse Jung Huang, Yun Sing Koh, Gillian Dobbie, Russel Pears |
PAKDD (2) | 4 |
| 2014 | Synthetic Minority Over-sampling TEchnique (SMOTE) for Predicting Software Build Outcomes
Jacqui Finlay, Russel Pears, Andy M. Connor |
SEKE | 2 |
| 2014 | Efficient negative association rule mining based on chance thresholdsabstractTypically association rule mining only considers positive frequent itemsets in rule generation, where rules involving only the presence of items are generated. In this paper we consider the complementary problem of negative association rule mining, w Yun Sing Koh, Russel Pears |
Intell. Data Anal. | 2 |
| 2014 | Data stream mining for predicting software build outcomes using source code metrics
Jacqui Finlay, Russel Pears, Andy M. Connor |
Inf. Softw. Technol. | 2 |
| 2014 | Detecting concept change in dynamic data streams - A sequential approach based on reservoir sampling
Russel Pears, Sripirakas Sakthithasan, Yun Sing Koh |
Mach. Learn. | 1 |
| 2013 | Tracking Drift Types in Changing Data Streams
David Tse Jung Huang, Yun Sing Koh, Gillian Dobbie, Russel Pears |
ADMA (1) | 4 |
| 2013 | One Pass Concept Change Detection for Data Streams
Sripirakas Sakthithasan, Russel Pears, Yun Sing Koh |
PAKDD (2) | 2 |
| 2013 | Discovering diverse association rules from multidimensional schema
Muhammad Usman 0005, Russel Pears |
Expert Syst. Appl. | 2 |
| 2013 | Weighted association rule mining via a graph based connectivity model
Russel Pears, Yun Sing Koh, Gillian Dobbie, Wai-Kiang Yeap |
Inf. Sci. | 1 |
| 2013 | A data mining approach to knowledge discovery from multidimensional cube structures
Muhammad Usman 0005, Russel Pears |
Knowl. Based Syst. | 2 |
| 2012 | Extrapolation Prefix Tree for Data Stream Mining Using a Landmark Model
Yun Sing Koh, Russel Pears, Gillian Dobbie |
DaWaK | 2 |
| 2012 | Precise Guidance to Dynamic Test Generation
TheAnh Do, Russel Pears |
ENASE | 3 |
| 2012 | WeightTransmitter: Weighted Association Rule Mining Using Landmark Weights
Yun Sing Koh, Russel Pears, Gillian Dobbie |
PAKDD (2) | 2 |
| 2011 | Discriminatory Confidence Analysis in Pattern Mining
Russel Pears, Yun Sing Koh, Gillian Dobbie |
ADMA (1) | 1 |
| 2011 | How Effective is Model Checking in Practice?
TheAnh Do, Russel Pears |
ENASE | 3 |
| 2011 | Automatic Assignment of Item Weights for Pattern Mining on Data Streams
Yun Sing Koh, Russel Pears, Gillian Dobbie |
PAKDD (1) | 2 |
| 2011 | Multiple Time-Series Prediction through Multiple Time-Series Relationships Profiling and Clustered Recurring Trends
Harya Widiputra, Russel Pears, Nikola K. Kasabov |
PAKDD (2) | 2 |
| 2011 | Mining Software Metrics from JazzabstractIn this paper, we describe the extraction of source code metrics from the Jazz repository and the application of data mining techniques to identify the most useful of those metrics for predicting the success or failure of an attempt to construct a working instance of the software product. We present results from a systematic study using the J48 classification method. The results indicate that only a relatively small number of the available software metrics that we considered have any significance for predicting the outcome of a build. These significant metrics are discussed and implication of the results discussed, particularly the relative difficulty of being able to predict failed build attempts. Jacqui Finlay, Andy M. Connor, Russel Pears |
SERA | 3 |
| 2011 | Dynamic Interaction Networks versus Local Trend Models for Multiple Time-Series PredictionabstractTime-series modeling and prediction have been very well researched by both the statistical and data mining communities. However, the multiple time-series problem of modeling and predicting simultaneous movements of a collection of time-sensitive variables that are related to each other has received much less attention. Strong relationships between variables suggest that trajectories of given variables involved in the relationships can be improved by including the nature and strength of these relationships in a prediction model. The key challenge is to capture the dynamics of the relationships to reflect changes that take place continuously over time. This research presents a global model to capture inclusive patterns of dynamic interactions between multiple time-series and a local trend model to extract localized profiles of relationships and recurring trends in multiple time-series. Our experimentation revealed that the global and local models specially developed for multiple time-series prediction outperformed methods such as multiple linear regression and the multilayer perceptron that were developed for predicting single time-series. Harya Widiputra, Russel Pears, Nikola K. Kasabov |
Cybern. Syst. | 2 |
| 2010 | EWGen: Automatic Generation of Item Weights for Weighted Association Rule Mining
Russel Pears, Yun Sing Koh, Gillian Dobbie |
ADMA (1) | 1 |
| 2010 | Valency Based Weighted Association Rule Mining
Yun Sing Koh, Russel Pears, Wai-Kiang Yeap |
PAKDD (1) | 2 |
| 2009 | A Novel Evolving Clustering Algorithm with Polynomial Regression for Chaotic Time-Series Prediction
Harya Widiputra, Henry Kho, Lukas, Russel Pears, Nikola K. Kasabov |
ICONIP (2) | 4 |
| 2009 | CBDT: A Concept Based Approach to Data Stream Mining
Stefan Hoeglinger, Russel Pears, Yun Sing Koh |
PAKDD | 2 |
| 2008 | Personalised Modelling for Multiple Time-Series Data Prediction: A Preliminary Investigation in Asia Pacific Stock Market Indexes Movement
Harya Widiputra, Russel Pears, Nikola K. Kasabov |
ICONIP (1) | 2 |
| 2008 | Transaction Clustering Using a Seeds Based Approach
Yun Sing Koh, Russel Pears |
PAKDD | 2 |
| 2007 | FGC: An Efficient Constraint Based Frequent Set MinerabstractDespite advances in algorithmic design, association rule mining remains problematic from a performance viewpoint when the size of the underlying transaction database is large. The well-known a priori approach, while reducing the computational effort involved still suffers from the problem of scalability due to its reliance on generating candidate itemsets. In this paper we present a novel approach that combines the power of preprocessing with the application of user-defined constraints to prune the itemset space prior to building a compact FP-tree. Experimentation shows that that our algorithm significantly outperforms the current state of the art algorithm, FP-bonsai. Russel Pears, Sangeetha Kutty |
AICCSA | 1 |
| 2007 | Optimization of Multidimensional Aggregates in Data WarehousesabstractThe computation of multidimensional aggregates is a common operation in OLAP applications. The major bottleneck is the large volume of data that needs to be processed which leads to prohibitively expensive query execution times. On the other hand, data analysts are primarily concerned with discerning trends in the data and thus a system that provides approximate answers in a timely fashion would suit their requirements better. In this article we present the prime factor scheme, a novel method for compressing data in a warehouse. Our data compression method is based on aggregating data on each dimension of the data warehouse. We used both real world and synthetic data to compare our scheme against the Haar wavelet and our experiments on range-sum queries show that it outperforms the latter scheme with respect to both decoding time and error rate, while maintaining comparable compression ratios. One encouraging feature is the stability of the error rate when compared to the Haar wavelet. Although wavelets have been shown to be effective at compressing data, the approximate answers they provide varies widely, even for identical types of queries on nearly identical values in distinct parts of the data. This problem has been attributed to the thresholding technique used to reduce the size of the encoded data and is an integral part of the wavelet compression scheme. In contrast the prime factor scheme does not rely on thresholding but keeps a smaller version of every data element from the original data and is thus able to achieve a much higher degree of error stability which is important from a Data Analysts point of view. Russel Pears, Bryan Houliston |
J. Database Manag. | 1 |