EDBT 2026 Demo / reviewers in the wild / expert
Sally I. McClean
dblp:m/SallyIMcClean
· DBLP profile ↗
128ranked-venue papers
23as first author
10since 2021 · last 2026
0000-0002-6871-3504ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 17 first-author · 3 since 2021Databases, data management, data science and information retrieval · 40 · 11 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 7 first-authorHuman-computer interaction and ubiquitous computing · 23 · 7 first-author · 1 since 2021Computer networks · 13Software engineering, systems software and programming languages · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Systems, architecture and hardware · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A zero-shot framework for cross-project vulnerability detection in source codeabstractAbstract The growing prevalence of software vulnerabilities has increased the need for effective detection methods, particularly in cross-project settings where domain differences create significant challenges. Existing vulnerability detection models often struggle to generalise across projects due to variations in coding styles, feature distributions, and the absence of labelled target data. This paper presents ZSVulD, a zero-shot, cross-project vulnerability detection framework designed to operate without target-domain labels. ZSVulD uses domain-agnostic CodeBERT embeddings to capture both syntactic and semantic features of source code, enabling knowledge transfer between projects. The framework applies an iterative pseudo-labelling process in which a neural network and XGBoost classifier collaboratively refine predictions for the target domain. Feature alignment is incorporated as a diagnostic technique to assess and visualise distributional differences between source and target datasets. Experiments on the Devign and REVEAL datasets show that ZSVulD achieves higher recall, F1, and F2 scores compared to existing methods, with an emphasis on reducing false negatives. These findings indicate that ZSVulD can support automated vulnerability detection pipelines, contributing to more reliable security assessments across different software projects. Radowanul Haque, Aftab Ali, Sally I. McClean |
Empir. Softw. Eng. | 3 |
| 2025 | Autonomous Management of IoT Devices in Smart Homes Using Semi-Markov Models
Dongwei Wang, Sally I. McClean, Ian R. McChesney, Zeeshan Tariq 0001 |
AINA (8) | 2 |
| 2025 | Automating mixture model fitting of task durations for process conformance checkingabstractAbstract Process task duration data often exhibit multiple peaks, indicating differences in, for example, customer ages and preferences, resource capabilities or the day/hour of a week. This heterogeneous data, which captures diverse customer patterns, should be represented using different models, resulting in an overall mixture model. This paper introduces gamma mixture models to represent various customer patterns in task duration data, with a focus on automating the fitting process. The approach involves a two-stage procedure: first, divide-and-conquer using peak-, equidistance- and cluster-based techniques to partition data, and automatically fit gamma distributions to each subset. The second stage then improves the fitted mixture model by directly searching the log-likelihood surface. The method is compared with the expectation–maximization (EM) algorithm and an open tool (HyperStar), using both artificially generated datasets and a publicly available hospital billing dataset, demonstrating its effectiveness and time efficiency in modelling heterogeneous process duration data. Furthermore, a case study on process conformance checking is conducted using the hospital billing dataset, highlighting a potential application area for the method in process mining. Lingkai Yang, Sally I. McClean, Malcolm J. Faddy, Mark P. Donnelly, Kashaf Khan, Kevin Burke |
Data Min. Knowl. Discov. | 2 |
| 2025 | Modelling process durations with gamma mixtures for right-censored data: Applications in customer clustering, pattern recognition, drift detection, and rationalisation
Lingkai Yang, Sally I. McClean, Kevin Burke, Mark P. Donnelly, Kashaf Khan |
Data Knowl. Eng. | 2 |
| 2024 | Detecting Process Duration Drift Using Gamma Mixture Models in a Left-Truncated and Right-Censored EnvironmentabstractWithin the realm of business context, process duration signifies time spent by customers between successive activities. This temporal perspective offers important insight to customer behavior, highlighting potential bottlenecks, and influencing business management decisions. The distribution of these process duration often changes over time due to factors such as seasonality, emerging legislation, changes to supply chains, and customer demand. Referred to as concept drift, these variations pose challenges for robust process modeling, understanding, and refinement. Subsequently, gamma mixture models are widely employed to model durations. These source data can, however, become left-truncated and right-censored within any specific observation window thereby necessitating a (well-known) modification to the likelihood function. The approach reported in this article leveraged this adapted likelihood across a series of observation windows, applying the likelihood ratio test to identify duration changes/concept drift. Due to its flexibility in modelling any duration distribution, the gamma mixture model was used with Nelder–Mead optimized likelihood for the left-truncated and right-censored data. The number of gamma components was determined by the Bayesian information criterion. The proposed framework underwent validation through simulated exponential samples, leading to recommendations for its practical application. Subsequently, we applied the methodology to three real-life event logs exhibiting diverse characteristics. Experimental results showcase the effectiveness of our approach in terms of data fitting, as compared to Kaplan–Meier curves, and in detecting instances of drift. This comprehensive validation underscores the practical utility and reliability of our framework for dynamic business scenarios. Lingkai Yang, Sally I. McClean, Mark P. Donnelly, Kashaf Khan, Kevin Burke |
ACM Trans. Knowl. Discov. Data | 2 |
| 2022 | A multi-components approach to monitoring process structure and customer behaviour concept drift
Lingkai Yang, Sally I. McClean, Mark P. Donnelly, Kevin Burke, Kashaf Khan |
Expert Syst. Appl. | 2 |
| 2022 | Towards automatic placement of media objects in a personalised TV experience
Brahim Allan, Ian Kegel, Sri Harish Kalidass, Andriy Kharechko, Michael Milliken, Sally I. McClean, Bryan W. Scotney, Shuai Zhang 0001 |
Multim. Syst. | 6 |
| 2022 | Modelling mobile-based technology adoption among people with dementiaabstractAbstract The work described in this paper builds upon our previous research on adoption modelling and aims to identify the best subset of features that could offer a better understanding of technology adoption. The current work is based on the analysis and fusion of two datasets that provide detailed information on background, psychosocial, and medical history of the subjects. In the process of modelling adoption, feature selection is carried out followed by empirical analysis to identify the best classification models. With a more detailed set of features including psychosocial and medical history information, the developed adoption model, using kNN algorithm, achieved a prediction accuracy of 99.41% when tested on 173 participants. The second-best algorithm built, using NN, achieved 94.08% accuracy. Both these results have improved accuracy in comparison to the best accuracy achieved (92.48%) in our previous work, based on psychosocial and self-reported health data for the same cohort. It has been found that psychosocial data is better than medical data for predicting technology adoption. However, for the best results, we should use a combination of psychosocial and medical data where it is preferable that the latter is provided from reliable medical sources, rather than self-reported. Priyanka Chaurasia, Sally I. McClean, Chris D. Nugent, Ian Cleland, Shuai Zhang 0001, Mark P. Donnelly, Bryan W. Scotney, Chelsea Sanders, Ken Smith, Maria C. Norton, JoAnn T. Tschanz |
Pers. Ubiquitous Comput. | 2 |
| 2021 | Discriminating features-based cost-sensitive approach for software defect predictionabstractAbstract Correlated quality metrics extracted from a source code repository can be utilized to design a model to automatically predict defects in a software system. It is obvious that the extracted metrics will result in a highly unbalanced data, since the number of defects in a good quality software system should be far less than the number of normal instances. It is also a fact that the selection of the best discriminating features significantly improves the robustness and accuracy of a prediction model. Therefore, the contribution of this paper is twofold, first it selects the best discriminating features that help in accurately predicting a defect in a software component. Secondly, a cost-sensitive logistic regression and decision tree ensemble-based prediction models are applied to the best discriminating features for precisely predicting a defect in a software component. The proposed models are compared with the most recent schemes in the literature in terms of accuracy, area under the curve, and recall. The models are evaluated using 11 datasets and it is evident from the results and analysis that the performance of the proposed prediction models outperforms the schemes in the literature. Aftab Ali, Mamun I. Abu-Tair, Joost Noppen, Sally I. McClean, Ian R. McChesney |
Autom. Softw. Eng. | 5 |
| 2021 | Dual contextual module for neural machine translationabstractAbstract Self-attention-based encoder-decoder frameworks have drawn increasing attention in recent years. The self-attention mechanism generates contextual representations by attending to all tokens in the sentence. Despite improvements in performance, recent research argues that the self-attention mechanism tends to concentrate more on the global context with less emphasis on the contextual information available within the local neighbourhood of tokens. This work presents the Dual Contextual (DC) module, an extension of the conventional self-attention unit, to effectively leverage both the local and global contextual information. The goal is to further improve the sentence representation ability of the encoder and decoder subnetworks, thus enhancing the overall performance of the translation model. Experimental results on WMT’14 English-German (En $$\rightarrow $$ → De) and eight IWSLT translation tasks show that the DC module can further improve the translation performance of the Transformer model. Isaac K. E. Ampomah, Sally I. McClean, Glenn I. Hawe |
Mach. Transl. | 2 |
| 2020 | Comparison of Analogue and Digital Fronthaul for 5G MIMO SignalsabstractThis paper investigates architectural and capacity issues associated with analogue and digital radio over fibre fronthaul for MIMO in 5G cellular systems. The capacity of both systems is evaluated in terms of a static system deployment and also with varying traffic load. The results show that in a leased line scenario, the analogue systems offer an opportunity to reduce the cost of fibre infrastructure by more than 93% due to the more efficient use of bandwidth. but since the ARoF equipment is currently more expensive than the DRoF equivalent, there is a trade-off between CAPEX and OPEX for a given deployment. Philip Perry, Colm Browning, Bryan W. Scotney, Amol Delmade, Sally I. McClean, Liam P. Barry, Adaranijo Peters, Philip J. Morrow |
ICC | 5 |
| 2020 | Novel Martingale Approaches for Change Point Detection
Jonathan Etumusei, Jorge Martínez Carracedo, Sally I. McClean |
ISDA | 3 |
| 2020 | Quantifying consensus of rankings based on q-support patterns
Zhengui Xue, Zhiwei Lin 0002, Hui Wang 0001, Sally I. McClean |
Inf. Sci. | 4 |
| 2020 | Performance of a Steady-State Visual Evoked Potential and Eye Gaze Hybrid Brain-Computer Interface on Participants With and Without a Brain InjuryabstractThe brain-computer interface (BCI) and the tracking of eye gaze provide modalities for human-machine communication and control. In this article, we provide the evaluation of a collaborative BCI and eye gaze approach, known as a hybrid BCI. The combined inputs interact with a virtual environment to provide actuation according to a four-way menu system. The following two approaches are evaluated: first, steady-state visual evoked potential (SSVEP) BCI with on-screen stimulation; second, hybrid BCI, which combined eye gaze and SSVEP for navigation and selection. A study comprises participants without known brain injury (non-BI, N = 30) and participants with known brain injury (BI, N = 14). A total of 29 out of 30 non-BI participants can successfully control the hybrid BCI, while nine out of the 14 BI participants are able to achieve control, as evidenced by task completion. The hybrid BCI provides a mean accuracy of 99.84% in the cohort of non-BI participants and 99.14% in the cohort of BI participants. Information transfer rates are 24.41 bpm in non-BI participants and 15.87 bpm in BI participants. The research goal is to quantify usage of SSVEP and ET approaches in cohorts of non-BI and BI participants. The hybrid is the preferred interaction modality for most participants for both cohorts. When compared to non-BI participants, it is encouraging that nine out of 14 participants with known BI can use the hBCI technology with equivalent accuracy and efficiency, albeit with slower transfer rates. Chris P. Brennan, Paul J. McCullagh, Gaye Lightbody, Leo Galway, Sally I. McClean, Piotr Stawicki, Felix Gembler, Ivan Volosyak, Elaine Armstrong, Eileen Thompson |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2019 | Using Phase-type Models to Monitor and Predict Process Target ComplianceabstractProcesses are ubiquitous, spanning diverse areas such as business, production, telecommunications and healthcare. \nThey have been studied and modelled for many years in an attempt to increase understanding, improve \nefficiency and predict future pathways, events and outcomes. More recently, process mining has emerged with \nthe intention of discovering, monitoring, and improving processes, typically using data extracted from event \nlogs. This may include discovering the tasks within the overall processes, predicting future trajectories, or \nidentifying anomalous tasks. We focus on using phase-type process modelling to measure compliance with \nknown targets and, inversely, determine suitable targets given a threshold percentage required for satisfactory \nperformance. We illustrate the ideas with an application to a stroke patient care process, where there are multiple \noutcomes for patients, namely discharge to normal residence, nursing home, or death. Various scenarios \nare explored, with a focus on determining compliance with given targets; such KPIs are commonly used in \nHealthcare as well as for Business and Industrial processes. We believe that this approach has considerable \npotential to be extended to include more detailed and explicit models that allow us to assess complex scenarios. \nPhase-type models have an important role in this work. Sally I. McClean, David A. Stanford, Lalit Garg |
ICORES | 1 |
| 2019 | Gated Task Interaction Framework for Multi-task Sequence TaggingabstractRecent studies have shown that neural models can achieve high performance on several sequence labelling/tagging problems without the explicit use of linguistic features such as part-of-speech (POS) tags. These models are trained only using the character-level and the word embedding vectors as inputs. Others have shown that linguistic features can improve the performance of neural models on tasks such as chunking and named entity recognition (NER). However, the change in performance depends on the degree of semantic relatedness between the linguistic features and the target task; in some instances, linguistic features can have a negative impact on performance. This paper presents an approach to jointly learn these linguistic features along with the target sequence labelling tasks with a new multi-task learning (MTL) framework called Gated Tasks Interaction (GTI) network for solving multiple sequence tagging tasks. The GTI network exploits the relations between the multiple tasks via neural gate modules. These gate modules control the flow of information between the different tasks. Experiments on benchmark datasets for chunking and NER show that our framework outperforms other competitive baselines trained with and without external training resources. Isaac K. E. Ampomah, Zhiwei Lin 0002, Sally I. McClean, Glenn I. Hawe |
IJCNN | 3 |
| 2019 | JASs: Joint Attention Strategies for Paraphrase Generation
Isaac K. E. Ampomah, Sally I. McClean, Zhiwei Lin 0002, Glenn I. Hawe |
NLDB | 2 |
| 2019 | Protection of records and data authentication based on secret shares and watermarking
Zulfiqar Ali 0001, Muhammad Imran 0001, Sally I. McClean, Muhammad Shoaib 0005 |
Future Gener. Comput. Syst. | 3 |
| 2018 | eZiGait: Toward an AI Gait Analysis And Sssistant System
Graham McCalmont, Philip J. Morrow, Huiru Zheng, Anas Samara, Sara Yasaei, Haiying Wang 0001, Sally I. McClean |
BIBM | 7 |
| 2017 | Quality of Service Scheme for Intra/Inter-Data Center CommunicationsabstractMost Information Technology (IT) services nowadays rely on one or multiple data centers. In this paper, we propose a Quality of Service (QoS) scheme for Intra/Inter data center communications systems. The proposed scheme is based on differentiation of traffic into different class categories. The differentiation is based on the source of traffic as well as the specific traffic's requirements. We conduct a comprehensive performance study for the proposed scheme. Our results demonstrate the benefits of deploying a QoS scheme in data centers especially if there is a considerable number of live virtual machine migrations among its servers. By completing such migration quickly, the Quality of Experience (QoE) will not degrade and the workload will be kept balanced among all the servers in the Data Center. Additionally, our proposed scheme considers the live virtual machine migrations among different Data Centers which is currently rapidly increasing. Mamun I. Abu-Tair, Md Israfil Biswas, Philip J. Morrow, Sally I. McClean, Bryan W. Scotney, Gerard P. Parr |
AINA | 4 |
| 2017 | Applications of phase type survival trees in HIV disease progression modellingabstractIt is important to model progression of a disease to understanding if the patient's condition is improving or getting worse. In the case of HIV disease, the change in the patient's CD4+ T cell count is used to calculate the progression of HIV disease i.e. if the CD4 count goes down it represent the progression of the patient's HIV disease. Due to the lack of an effective cure for HIV disease, it is crucial to monitor the disease progression to managing HIV disease effectively. Therefore, this study is aimed to model HIV disease progression by using phase type survival trees to cluster patients into homogenous groups based on their disease progression to understand the effect of different factors of prognostic significance and their interactions affecting the disease progression. The proposed methods are evaluated using an empirical data of 1,838 HIV-infected patients. The methods developed in this study can also be used for modelling the progression of other chronic conditions or diseases. Marija Gafa, Lalit Garg, Giovanni Masala, Sally I. McClean |
ICCSA (7) | 4 |
| 2017 | Sensor-Based Change Detection for Timely Solicitation of User EngagementabstractThe accurate detection of changes has the potential to form a fundamental component of systems which autonomously solicit user interaction based on transitions within an input stream, for example, electrocardiogram data or accelerometry obtained from a mobile device. This solicited interaction may be utilized for diverse scenarios such as responding to changes in a patient's vital signs within a medical domain or requesting user activity labels for generating real-world labelled datasets. Within this paper, we extend our previous work on the Multivariate Online Change detection Algorithm subsequently exploring the utility of incorporating the Benjamini Hochberg method of correcting for multiple comparisons. Furthermore, we evaluate our approach against similarly light-weight Multivariate Exponentially Weighted Moving Average and Cumulative Sum based techniques. Results are presented based on manually labelled change points in accelerometry data captured using 10 participants. Each participant performed nine distinct activities for a total period of 35 minutes. The results subsequently demonstrate the practical potential of our approach from both accuracy and computational perspectives. Timothy Patterson, Sally I. McClean, Chris D. Nugent, Shuai Zhang 0001, Ian Cleland |
IEEE Trans. Mob. Comput. | 3 |
| 2016 | Using Fitt's Law to Model Arm Motion Tracked in 3D by a Leap Motion Controller for Virtual Reality Upper Arm Stroke RehabilitationabstractEarly and intensive physical therapy can improve upper arm and hand functionality for stroke survivors. Virtual reality systems that utilize state-of-art natural user interface tracking sensors and adaptive user profiling software has the potential to supplement traditional physiotherapy and engage patients to sustain beneficial quantity and quality of rehabilitation. In this paper we provide an overview of the problem area and present results from an initial experiment with healthy users. The potential for Fitts's law to model user motion effectively in reach and touch tasks within 3D virtual environments is investigated. Results indicate that Fitts's law may be effective though we propose that it would best be used as part of a more complex model of user motion. Dominic E. Holmes, Darryl Charles, Philip J. Morrow, Sally I. McClean, Suzanne McDonough |
CBMS | 4 |
| 2016 | Using Genetic Algorithms for Optimal Change Point Detection in Activity MonitoringabstractActivity Monitoring is a key feature of health and well-being assessment that has received increased consideration from the research community over the last few decades. Body worn sensors and smart devices are widely used in Activity Monitoring in order to capture and classify large amounts of data over short periods of time, in a relatively un-obtrusive manner. Change point detection is a technique at the core of the data processing of the sensory data recorded used to identify the transition from one underlying time series generation model to another. The sudden change in mean, variance or both may represent change point in time series data. Accurate and automatic change point detection in data is not only used to identify events (transition from one activity to another), however, can also be used for labelling activities to generate real world annotated datasets. This paper proposes a genetic algorithm (GA) that identifies the optimal set of parameters for a Multivariate Exponentially Weighted Moving Average (MEWMA) approach to change point detection. The proposed technique optimizes different parameters of the MEWMA in an effort to find the maximum F-measure, which subsequently identifies the exact location of the change point from an existing activity to a new one. Results have been evaluated based on real and synthetic datasets collected from accelerometer data during a set of 8 different activities for two users with a high degree of accuracy form 99.4% to 99.8% and F-measure to 66.7%. Sally I. McClean, Shuai Zhang 0001, Chris D. Nugent |
CBMS | 2 |
| 2016 | Tree Similarity Measurement for Classifying Questions by Syntactic Structures
Zhiwei Lin 0002, Hui Wang 0001, Sally I. McClean |
ICIC (3) | 3 |
| 2016 | Modelling assistive technology adoption for people with dementia
Priyanka Chaurasia, Sally I. McClean, Chris D. Nugent, Ian Cleland, Shuai Zhang 0001, Mark P. Donnelly, Bryan W. Scotney, Chelsea Sanders, Ken Smith, Maria C. Norton, JoAnn T. Tschanz |
J. Biomed. Informatics | 2 |
| 2015 | Facilitating Delivery and Remote Monitoring of Behaviour Change Interventions to Reduce Risk of Developing Alzheimer's Disease: The Gray Matters Study
Phillip J. Hartin, Ian Cleland, Chris D. Nugent, Sally I. McClean, Timothy Patterson, JoAnn T. Tschanz, Christine Clark, Maria C. Norton |
ICOST | 4 |
| 2015 | Relaying for 5G: A novel low-error relaying protocolabstractFuture 5G networks have stringent end-user requirements on data rate and error performance. In order to satisfy these requirements, innovative wireless networking technologies and models need be researched. One particular example is the two-way relaying channel, which can have as much as 100% higher theoretical data rate than current systems where transmissions are arranged in an orthogonal manner. However, benefits of this model cannot be achieved without the application of proper relaying protocols. This paper proposes a novel protocol that directly addresses the problems of existing protocols of two-way relaying models, e.g. analogy network coding and physical network coding, and has improved performance. By combining direct and differential demodulation-forward schemes based on wireless channel qualities and signal to noise ratio, a new hybrid protocol is created. Theoretical analysis and numerical experiments show that the proposed solution has lower error rate than the existing ones, and can thus be applied to support future 5G networks. Chunbo Luo, Gerard P. Parr, Sally I. McClean, Cathryn Peoples, Xinheng Wang 0001, James Nightingale, Qi Wang 0001 |
ISCC | 3 |
| 2015 | Hybrid Demodulate-Forward Relay Protocol for Two-Way Relay ChannelsabstractTwo-Way Relay Channel (TWRC) plays an important role in relay networks, and efficient relaying protocols are particularly important for this model. However, existing protocols may not be able to realize the potential of TWRC if the two independent fading channels are not carefully handled. In this paper, a Hybrid DeModulate-Forward (HDMF) protocol is proposed to address such a problem. We first introduce the two basic components of HDMF - direct and differential DMF, and then propose the key decision criterion for HDMF based on the corresponding log-likelihood ratios. We further enhance the protocol so that it can be applied independently from the modulation schemes. Through extensive mathematical analysis, theoretical performance of the proposed protocol is investigated. By comparing with existing protocols, the proposed HDMF has lower error rate. A novel scheduling scheme for the proposed protocol is introduced, which has lower length than the benchmark method. The results also reveal the protocol's potential to improve spectrum efficiency of relay channels with unbalanced bilateral traffic. Chunbo Luo, Gerard P. Parr, Sally I. McClean, Cathryn Peoples, Xinheng Wang 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2014 | Timely autonomous identification of UAV safe landing zones
Timothy Patterson, Sally I. McClean, Philip J. Morrow, Gerard P. Parr, Chunbo Luo |
Image Vis. Comput. | 2 |
| 2014 | Development of a Technology Adoption and Usage Prediction Tool for Assistive Technology for People with DementiaabstractIn the current work, data gleaned from an assistive technology (reminding technology), which has been evaluated with people with Dementia over a period of several years was retrospectively studied to extract the factors that contributed to successful adoption. The aim was to develop a prediction model with the capability of prospectively assessing whether the assistive technology would be suitable for persons with Dementia (and their carer), based on user characteristics, needs and perceptions. Such a prediction tool has the ability to empower a formal carer to assess, through a very limited amount of questions, whether the technology will be adopted and used. Sonja O'Neill, Sally I. McClean, Mark P. Donnelly, Chris D. Nugent, Leo Galway, Ian Cleland, Shuai Zhang 0001, Terry Young, Bryan W. Scotney, Sarah C. Mason, David Craig |
Interact. Comput. | 2 |
| 2014 | Clustering semantically heterogeneous distributed aggregate databases
Shuai Zhang 0001, Sally I. McClean, Bryan W. Scotney |
Knowl. Inf. Syst. | 2 |
| 2014 | A Predictive Model for Assistive Technology Adoption for People With DementiaabstractAssistive technology has the potential to enhance the level of independence of people with dementia, thereby increasing the possibility of supporting home-based care. In general, people with dementia are reluctant to change; therefore, it is important that suitable assistive technologies are selected for them. Consequently, the development of predictive models that are able to determine a person's potential to adopt a particular technology is desirable. In this paper, a predictive adoption model for a mobile phone-based video streaming system, developed for people with dementia, is presented. Taking into consideration characteristics related to a person's ability, living arrangements, and preferences, this paper discusses the development of predictive models, which were based on a number of carefully selected data mining algorithms for classification. For each, the learning on different relevant features for technology adoption has been tested, in conjunction with handling the imbalance of available data for output classes. Given our focus on providing predictive tools that could be used and interpreted by healthcare professionals, models with ease-of-use, intuitive understanding, and clear decision making processes are preferred. Predictive models have, therefore, been evaluated on a multi-criterion basis: in terms of their prediction performance, robustness, bias with regard to two types of errors and usability. Overall, the model derived from incorporating a k-Nearest-Neighbour algorithm using seven features was found to be the optimal classifier of assistive technology adoption for people with dementia (prediction accuracy 0.84 ± 0.0242). Shuai Zhang 0001, Sally I. McClean, Chris D. Nugent, Mark P. Donnelly, Leo Galway, Bryan W. Scotney, Ian Cleland |
IEEE J. Biomed. Health Informatics | 2 |
| 2013 | Costing Mixed Coxian Phase-type Systems in a given time intervalabstractPreviously we have introduced a modelling framework to classify individuals in Mixed Coxian Phase-type Systems. We here add costs and obtain results for moments of total costs in (0, t], for an individual, and a cohort arriving at time zero. Based on data from the Belfast City Hospital Stroke Unit we use the overall modelling framework to obtain results for total cost in a given time interval to facilitate planners who have limited time horizons for budget planning. Sally I. McClean, Lalit Garg, Jennifer Gillespie, Ken Fullerton |
CBMS | 1 |
| 2013 | Activity recognition and resource optimization in mobile cloud through MapReduceabstractMobile cloud computing aims at improving user experience through enhancing the ability of mobile applications by doing intensive tasks in the cloud. In this paper we consider an environment similar to a hybrid cloud in which the mobile device works as a private cloud. Given that the mobile phone has both limited processing resources and battery time, the proposed mobile application architecture has been designed with the capability of sending specified data/parameters to the cloud. This data is subsequently used for further processing/mining and visualization to assist in inferring further information through mapreduce. This information gives details about resource and battery consumption which will help in optimizing the relationship between the mobile device and cloud. It will also be beneficial to the optimization of the mobile application through the trends visualized in the cloud. In this paper we created an activity recognition health application as an example and helped the user about his health along with giving an insight into abnormal behavior and lifestyle trends. Shujaat Hussain, Muhammad Bilal Amin, Jae Hun Bang, Manhyung Han, Sungyoung Lee 0001, Chris D. Nugent, Sally I. McClean, Bryan W. Scotney, Gerard P. Parr |
Healthcom | 7 |
| 2013 | Energy aware scheduling across 'green' cloud data centres
Cathryn Peoples, Gerard P. Parr, Sally I. McClean, Philip J. Morrow, Bryan W. Scotney |
IM | 3 |
| 2013 | Multiple-source multiple-destinations relay channels with network codingabstractThe essential broadcasting feature of radio propagation channels provides an opportunity for multiple nodes to exchange information and work cooperatively, where each node, for example, mobile sensor or robot, plays both the role of transmitter and receiver. Especially in such machine‐to‐machine communication scenarios, the network performance and redundancy of information from different providers highly affect the work efficiency of each individual node and the whole system. To address these problems, a relay assisted centralised network model with physical layer network coding implemented in the relay is proposed in this study. This structure has the advantage of flexible data exchange and the capability to reduce redundancy in information. Its theoretical performance is analysed by the diversity multiplexing tradeoff, which proves the proposed model is versatile in reliable and high spectral‐efficiency information exchange. Experiments of multiple nodes in a machine‐to‐machine scenario – unmanned aerial vehicles, further reveal its potential in improving efficiency of communication and cooperation. Chunbo Luo, Sally I. McClean, Gerard P. Parr, Peng Ren 0001 |
IET Commun. | 2 |
| 2013 | Assessing Gait Patterns of Healthy Adults Climbing Stairs Employing Machine Learning TechniquesabstractSo far, stair climbing has not been studied as extensively as gait has, although the significance of the prevention of falling on stairs has been well recognized. Based on acceleration data taken from 25 healthy subjects climbing up and down a set of 13 stairs with an accelerometer placed on the lumbo-sacral joint, this paper aims to assess gait patterns of younger and older adults climbing stairs using a machine learning approach. A total of 14 gait features were extracted and analyzed. The performance of six representative classification models: Multilayer Perceptron (MLP), KStar, Support Vector Machine (SVM), Naïve Bayesian (NB), C4.5 Decision Trees, and Random Forests were evaluated in terms of their ability to discriminate between younger and older adults climbing up- and downstairs. MLP was found to provide the highest accuracy for classification. Accuracy of 95.7% was found for classifying a subject walking either up or down the stairs and an accuracy of 80.6% for classifying whether the subject was younger or older. An evaluation of individual features showed poor performance of classification for younger and older subjects climbing up- and downstairs, and in most cases failed to distinguish between the two classes. To access which set of features derived from a triaxial accelerometer can better describe the performance differences between younger and older adults climbing up- and downstairs, two feature selection algorithms, sequential feature selection and correlation-based feature selection, were implemented. Results show that 10 features derived from correlation-based feature selection were able to produce a 96.8% accuracy for classification between subjects climbing up and down. A subset of seven features achieved a performance of 84.9% accuracy for classification between younger and older subjects. Herman Chan, Mingjing Yang 0001, Haiying Wang 0001, Huiru Zheng, Sally I. McClean, Roy Sterritt, Ruth E. Mayagoitia |
Int. J. Intell. Syst. | 5 |
| 2012 | Using phase type distributions for modelling HIV disease progressionabstractDisease progression models are useful tools for gaining a systems' understanding of the transitions to disease states, and characterizing the relationship between disease progress and factors affecting it such as patients' profile, treatment and the HIV diagnosis stage. Patients are classified into four states (based on CD4+ T-lymphocyte count) and all the transitions are allowed. Examinations to identify disease progression of the patient are carried out routinely throughout the follow-up period. Therefore, the times spent at the various HIV infection stages are interval censored or right censored. This makes difficult to use simple statistical methods such as regression to model the disease progression and its relationship with the diagnosis stage. We present a novel, more intuitive and realistic approach based on phase type distributions to model progression of HIV infection and the effects and prognostic significance of HIV diagnosis stage. The approach is illustrated using a real database of total 2,092 HIV infected patients enrolled in the Italian public structures from January 1996 to January 2008. The approach can also be used to examine the effect of other covariates such as patient's profile. Lalit Garg, Giovanni Masala, Sally I. McClean, Marco Micocci, Giuseppina Cannas |
CBMS | 3 |
| 2012 | Using phase-type models to cost a cohort of stroke patientsabstractStroke disease incurs long periods of hospital and community care, with associated costs. Also stroke is a highly complex disease with diverse outcomes and multiple strategies and care options for therapy and care. Previously we have developed a modeling framework which classifies patients with respect to their length of stay (LOS); phase-type models can then be used to describe patient flows for each class. Also multiple outcomes, such as discharge to normal residence, nursing home, or death can be included. We here add costs to this model of the total stroke care system and determine moments for total cost of a cohort of patients and also the separate costs in different phases and stages of care. Based on stroke patients' data from the Belfast City Hospital, various scenarios are explored with a focus on comparing the costs of thrombolysis, a clot-busting therapy under different regimes. Sally I. McClean, Jennifer Gillespie, Bryan W. Scotney, Ken Fullerton |
CBMS | 1 |
| 2012 | A Smart Garment for Older Walkers
William P. Burns, Chris D. Nugent, Paul J. McCullagh, Dewar D. Finlay, Ian Cleland, Sally I. McClean, Bryan W. Scotney, Jane McCann |
ICOST | 6 |
| 2012 | Stakeholder Involvement Guidelines to Improve the Design Process of Assistive Technology
Leo Galway, Sonja O'Neill, Mark P. Donnelly, Chris D. Nugent, Sally I. McClean, Bryan W. Scotney |
ICOST | 5 |
| 2012 | Performance analysis of Bayesian Networks-based distributed Call Admission Control for NGNabstractThe efficient management of networks and the provisioning of services with desired QoS guarantees is a challenge which needs to be addressed through autonomous mechanisms which are intelligent, lightweight and scalable. Recent focus on applying Machine Learning approaches to model the network and service behavioural patterns have proved to be quite effective in fulfilling the objectives of autonomous management. To this end, this paper advances on the idea of implementing a distributed management solution which harnesses the predictive capability of Bayesian Networks (BN). A multi-node distributed Call Admission Control solution (termed as BNDAC) is proposed and implemented to demonstrate the modelling and prediction power of BN. A thorough evaluation of BNDAC is presented in terms of its prediction accuracy, algorithmic complexity and decision-making speed. In an online setup, performance of BNDAC is evaluated and compared with a centralised scenario, to demonstrate its superior performance for Call Blocking Probability and QoS provisioning. Simulation results based on Opnet Modeler and Hugin Researcher show the feasibility and applicability of BNDAC solution for real-time operation and management of real world networks such as the NGN. Abul Bashar, Gerard P. Parr, Sally I. McClean, Bryan W. Scotney, Detlef D. Nauck |
NOMS | 3 |
| 2012 | Coupling edge and region-based information for boundary finding in biomedical imagery
Huaizhong Zhang, Philip J. Morrow, Sally I. McClean, Kurt Saetzler |
Pattern Recognit. | 3 |
| 2012 | Probabilistic Learning From Incomplete Data for Recognition of Activities of Daily Living in Smart HomesabstractLearning behavioral patterns for activities of daily living in a smart home environment can be challenged by the limited number of training data that may be available. This may be due to the infrequent repetition of routine activities (e.g., once daily), the expense of using observers to label activities, and the intrusion that would be caused by the presence of observers over long time periods. It is important, therefore, to make as much use of any labeled data that are collected, however, incomplete these data may be. In this paper, we propose an algorithm for learning behavioral patterns for multi-inhabitants living in a single smart home environment, by making full use of all limited labeled activities, including incomplete data resulting from unreliable low-level sensors in this environment. Through maximum-likelihood estimation, using Expectation-Maximization, we build a model that captures both environmental uncertainties from sensor readings and user uncertainties, including variations in how individuals carry out activities. Our algorithm outperforms models that cannot handle data incompleteness, with increasing performance gains as incompleteness increases. The approach also enables the impact of particular sensors to be assessed and can thus inform sensor maintenance and deployment. Shuai Zhang 0001, Sally I. McClean, Bryan W. Scotney |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2012 | A Multidimensional Sequence Approach to Measuring Tree SimilarityabstractTree is one of the most common and well-studied data structures in computer science. Measuring the similarity of such structures is key to analyzing this type of data. However, measuring tree similarity is not trivial due to the inherent complexity of trees and the ensuing large search space. Tree kernel, a state of the art similarity measurement of trees, represents trees as vectors in a feature space and measures similarity in this space. When different features are used, different algorithms are required. Tree edit distance is another widely used similarity measurement of trees. It measures similarity through edit operations needed to transform one tree to another. Without any restrictions on edit operations, the computation cost is too high to be applicable to large volume of data. To improve efficiency of tree edit distance, some approximations were introduced into tree edit distance. However, their effectiveness can be compromised. In this paper, a novel approach to measuring tree similarity is presented. Trees are represented as multidimensional sequences and their similarity is measured on the basis of their sequence representations. Multidimensional sequences have their sequential dimensions and spatial dimensions. We measure the sequential similarity by the all common subsequences sequence similarity measurement or the longest common subsequence measurement, and measure the spatial similarity by dynamic time warping. Then we combine them to give a measure of tree similarity. A brute force algorithm to calculate the similarity will have high computational cost. In the spirit of dynamic programming two efficient algorithms are designed for calculating the similarity, which have quadratic time complexity. The new measurements are evaluated in terms of classification accuracy in two popular classifiers (k-nearest neighbor and support vector machine) and in terms of search effectiveness and efficiency in k-nearest neighbor similarity search, using three different data sets from natural language processing and information retrieval. Experimental results show that the new measurements outperform the benchmark measures consistently and significantly. Zhiwei Lin 0002, Hui Wang 0001, Sally I. McClean |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Intelligent Patient Management and Resource Planning for Complex, Heterogeneous, and Stochastic Healthcare SystemsabstractEffective resource requirement forecasting is necessary to reduce the escalating cost of care by ensuring optimum utilization and availability of scarce health resources. Patient hospital length of stay (LOS) and thus resource requirements depend on many factors including covariates representing patient characteristics such as age, gender, and diagnosis. We therefore propose the use of such covariates for better hospital capacity planning. Likewise, estimation of the patient's expected destination after discharge will help in allocating scarce community resources. Also, probable discharge destination may well affect a patient's LOS in hospital. For instance, it might be required to delay the discharge of a patient so as to make appropriate care provision in the community. A number of deterministic models such as ratio-based methods have failed to address inherent variability in complex health processes. To address such complexity, various stochastic models have therefore been proposed. However, such models fail to consider inherent heterogeneity in patient behavior. Therefore, we here use a phase-type survival tree for groups of patients that are homogeneous with respect to LOS distribution, on the basis of covariates such as time of admission, gender, and disease diagnosed; these homogeneous groups of patients can then model patient flow through a care system following stochastic pathways that are characterized by the covariates. Our phase-type model is then extended by further growing the survival tree based on covariates representing outcome measures such as treatment outcome or discharge destinations. These extended phase-type survival trees are very effective in modeling interrelationship between a patient's LOS and such outcome measures and allow us to describe patient movements through an integrated care system including hospital, social, and community components. In this paper, we first propose a generalization of the Coxian phase-type distribution to a Markov process with more than one absorbing state; we call this the multi-absorbing state phase-type distribution. We then describe how the model can be used with the extended phase-type survival tree for forecasting hospital, social, and community care resource requirements, estimating cost of care, predicting patient demography at a given time in the future, and admission scheduling. We can, thus, provide a stochastic approach to capacity planning across complex heterogeneous care systems. The approach is illustrated using a five year retrospective data of patients admitted to the stroke unit of the Belfast City Hospital. Lalit Garg, Sally I. McClean, Maria Barton, Brian Meenan, Ken Fullerton |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2011 | Markovian Workload Characterization for QoS Prediction in the CloudabstractResource allocation in the cloud is usually driven by performance predictions, such as estimates of the future incoming load to the servers or of the quality-of-service(QoS) offered by applications to end users. In this context, characterizing web workload fluctuations in an accurate way is fundamental to understand how to provision cloud resources under time-varying traffic intensities. In this paper, we investigate the Markovian Arrival Processes (MAP) and the related MAP/MAP/1 queueing model as a tool for performance prediction of servers deployed in the cloud. MAPs are a special class of Markov models used as a compact description of the time-varying characteristics of workloads. In addition, MAPs can fit heavy-tail distributions, that are common in HTTP traffic, and can be easily integrated within analytical queueing models to efficiently predict system performance without simulating. By comparison with traced riven simulation, we observe that existing techniques for MAP parameterization from HTTP log files often lead to inaccurate performance predictions. We then define a maximum likelihood method for fitting MAP parameters based on data commonly available in Apache log files, and a new technique to cope with batch arrivals, which are notoriously difficult to model accurately. Numerical experiments demonstrate the accuracy of our approach for performance prediction of web systems. Sergio Pacheco-Sanchez, Giuliano Casale, Bryan W. Scotney, Sally I. McClean, Gerard P. Parr, Stephen Dawson |
IEEE CLOUD | 4 |
| 2011 | Using model-based clustering to discretise duration information for activity recognitionabstractActivity recognition is an important component of patient management in smart homes where high level activities can be learned from low level sensor data. Such activity recognition utilises sensor ID, task order and time of activation to learn about patient behavior, detect anomalies and provide prompts or other interventions. In this paper we use the sensor activation times to calculate durations and then investigate several model-based clustering approaches with a view to discretising the duration data and using such data to improve activity prediction. We explore several popular approaches to characterising such duration data, namely Coxian phase type distributions and Gaussian mixture distributions. We then show how we can utilise the learned clustering components for discretisation. Finally we use simulated data, based on a real smart kitchen deployment, to compare these approaches and evaluate the discretisation results with regard to activity prediction. Sally I. McClean, Lalit Garg, Priyanka Chaurasia, Bryan W. Scotney, Chris D. Nugent |
CBMS | 1 |
| 2011 | A framework for context-aware online physiological monitoringabstractWith the challenge of healthcare for the increasing number of elderly people and the prevalence of chronic disease, research has been carried out on the development of assistive technologies and devices. This paper proposes a framework of context-aware physiological analysis for remote and efficient healthcare. With the relationship between the physiological function and daily activities, the online detection of abnormal situation needs to be carried out given such rich context information. Two core modules in the framework are discussed in details by proposing hierarchical online activity recognition and dynamic Cumulative Sum Control Chart (CUSUM) methods for process control. Corresponding experiments have been set up to collect both ECG data and upper-body accelerations from two healthy participants. This framework also has great potential to be used for long term health drift detection by comparison of the physiological function patterns given the activity across different periods of time. Shuai Zhang 0001, Sally I. McClean, Bryan W. Scotney, Leo Galway, Chris D. Nugent |
CBMS | 2 |
| 2011 | Utilizing Wearable Sensors to Investigate the Impact of Everyday Activities on Heart Rate
Leo Galway, Shuai Zhang 0001, Chris D. Nugent, Sally I. McClean, Dewar D. Finlay, Bryan W. Scotney |
ICOST | 4 |
| 2011 | Evaluation of Video Reminding Technology for Persons with Dementia
Chris D. Nugent, Sonja O'Neill, Mark P. Donnelly, Guido Parente, Mark Beattie, Sally I. McClean, Bryan W. Scotney, Sarah C. Mason, David Craig |
ICOST | 6 |
| 2011 | Dynamic bit-rate adjustment based on traffic characteristics for metro and core networksabstractThis paper reports on an investigation into bit-rate adjustment on high capability metro and core routers. The bit-rate is controlled in response to traffic load using EWMA (Exponential Weighted Moving Average), with parameters set to minimize the performance impact of the adjustments. The purpose of the adjustment is to enable energy savings. Our calculations show that 22.5% power saving is achievable by bit rate adjustment without significant performance degradation. Huseyin Abaci, Gerard P. Parr, Sally I. McClean, Adrian Moore 0001, Louise Krug |
Integrated Network Management | 3 |
| 2011 | Novel distributed call admission control solution based on machine learning approachabstractThe advent of IP-based Next Generation Network (NGN) and its guaranteed QoS promise has attracted significant attention from both service providers and subscribers. However, to fulfil the said promise, there is a need to provide effective Call Admission Control (CAC) based QoS provisioning solutions which are autonomous, intelligent and scalable. Abul Bashar, Gerard P. Parr, Sally I. McClean, Bryan W. Scotney, Detlef D. Nauck |
Integrated Network Management | 3 |
| 2011 | Hybrid optical and wireless technology integrations for next generation broadband access networksabstractHybrid optical and wireless technology integrations have been considered as one of the most promising candidates for the next generation broadband access networks for quite some time. The integration scheme provides the bandwidth advantages of the optical networks and mobility features of the wireless networks for Subscriber Stations (SSs). It also brings economic efficiency to the network providers particularly in rural area where the existing wired telecommunication infrastructures such as Digital Subscriber Line (DSL), Cable Modem (CM), T-1/E-1 networks or fibre deployments are either costly or unreachable. For successful integration of the optical and wireless technologies there are some technical issues which need to be addressed efficiently in order to provide End-to-End (ETE) and diverse Quality of Service (QoS) for various service classes. This paper investigates the possible challenging issues for the integrated structure of the Time Division Multiplexing and Wavelength Division Multiplexing Ethernet Passive Optical Networks (TDM EPON and WDM EPON ) with the Worldwide Interoperability for Microwave Access and Wireless Fidelity (WiMAX and Wi-Fi) networks. To reduce the ETE delay and provide the QoS for diverse service classes, we have compared six existing upstream scheduling mechanisms in two levels which are distributed on Access Points (APs) from Wi-Fi domain and Base Stations (BSs) from WiMAX domain. Performance evaluations of the existing scheduling techniques for three popular service classes (Quad-play) have been studied which show the strong impact of using the efficient up-link scheduler in converged scenario. We have also proposed a dynamic scheduling algorithm for optical and wireless integration scheme, which is under the implementation and evaluation process. Naghmeh Moradpoor Sheykhkanloo, Gerard P. Parr, Sally I. McClean, Bryan W. Scotney, Gilbert Owusu |
Integrated Network Management | 3 |
| 2011 | Context-aware characterisation of energy consumption in data centresabstractCarbon emissions are receiving increased attention and scrutiny in all walks of life and the ICT sector is no exception. With the increase in on-demand applications and services together with on-demand compute/storage facilities in server farms or data centres there are self-evident increases in the power requirements to maintain such systems. Proponents of the impact of increased carbon emissions when powering electrical systems in general however, regularly impress negative side-effects such as influence on climate change. Action is subsequently being encouraged to halt further environmental damage. The problem is explored in this paper from the point of view of carbon emissions from data centre operations and the development of energy-aware management and energy-efficient networking solutions. Data centre energy consumption costs drive the evaluation process within a Data Centre Energy-Efficient Context-Aware Broker (DCe-CAB) algorithm designed as an original solution to this significant carbon-contributing network scenario. In this paper, performance requirements and objectives of the DCe-CAB are defined, along with case study demonstration of the way in which it optimises selection and operation of data centres using context-awareness. Cathryn Peoples, Gerard P. Parr, Sally I. McClean |
Integrated Network Management | 3 |
| 2011 | Modelling and evaluation of a policy-based resource management framework for converged next generation networksabstractAs fixed and wireless access networks converge towards Internet Protocol based transport in next generation networks, the requirement for effective and scalable control and management solutions to address the complexity introduced by their heterogeneity becomes ever more critical. Over the years, policy-based management has emerged as a viable tool to address this challenge. Within the IU-ATC project, we are developing a policy-based framework, (Converged Networks QoS Framework, CNQF) for end-to-end QoS control and resource management in converged next generation networks. CNQF is designed to support QoS control and context-driven resource management policies across all networks (access, metro, core) on the end-to-end transport layer of converged networks. This paper describes CNQF and develops an exemplary use case scenario for policy-based resource management based on our CNQF architecture. The paper also presents the development of stochastic-analytic models to characterise and evaluate the impact of policies on operational performance. The model is used to analyse the performance of dynamic CNQF context-driven policies for admission control case study. Suleiman Y. Yerima, Gerard P. Parr, Sally I. McClean, Philip J. Morrow |
Integrated Network Management | 3 |
| 2011 | Characterization, monitoring and evaluation of operational performance trends on server processor hardwareabstractEnterprise IT environments have seen a sharp growth in content use due to the popularity of on-demand data-intensive applications. In turn, the huge demand in content has spawned off major developments such as growth and distribution of computing nodes as well as the adoption of various implementation technologies. Given the complexity brought to the makeup of business computing environments in addressing the above-mentioned factors, the critical planning task of determining the appropriate infrastructure sizes for supporting firm Quality of Service (QoS) guarantees becomes a very challenging undertaking to fulfil. Benchmarking methods are widely employed in calibrating attainable performance in IT solutions, but these have the drawback of presenting output performance metrics as composite measurements that only give an end-to-end perspective. As an enhancement to benchmarking approaches, we explore the use of Performance Monitoring Counters (PMCs) in obtaining detailed operational performance of CPU and memory hardware. Performance Monitoring Counters (PMCs) are onchip registers found on most modern processor hardware. We use PMC-derived measurements to validate cache performance trends that have been derived analytically, and in the course of validations, PMC data is also used to investigate the nature and character of surges in cache miss events, which emerge as the memory load generated by runtime processes increases. Ernest Sithole, Sally I. McClean, Bryan W. Scotney, Gerard P. Parr, Adrian Moore 0001, David W. Bustard, Stephen Dawson |
ICPE | 2 |
| 2010 | Using mixed phase-type distributions to model patient pathwaysabstractModeling length of stay (LOS) in hospital is an important aspect of developing integrated models that describe and predict movements of patients. However patient pathways and LOS distributions are highly heterogeneous, particularly with regard to patient diagnosis, age, gender and outcome. We here use a mixed Coxian phase-type distribution (MC-PH distribution) to describe such heterogeneity in terms of covariates, where a Coxian phase-type survival tree is used to estimate parameters for the MC-PH distribution. Multiple absorbing states (such as discharge to home, discharge to private nursing home, or death) are considered, and, based on the MC-PH distribution, expressions presented for key performance indicators of interest. The approach is illustrated using data for stroke patients from the Belfast City Hospital. Sally I. McClean, Lalit Garg, Maria Barton, Ken Fullerton |
CBMS | 1 |
| 2010 | Weighting common syntactic structures for natural language based information retrievalabstractNatural Language Processing (NLP) techniques are believed to hold the potential to assist "bag-of-words" Information Retrieval (IR) in terms of retrieval accuracy. In this paper, we report a natural language based IR approach where the common syntactic structures between documents and the query is regarded to as a query-dependent feature for documents. Specifically, a "structural weight" is proposed for query terms, which can be seen as a weight to model the degree of term's involvement in the common syntactic structures. This structural weight is used together with the TF-IDF weighting scheme, which results in a new ranking function. The accumulation of this structural weight of all the query terms in the new ranking function will be seen as a measure of how much a document and a query share the common syntactic structures. The experimental results show that by using this ranking function, significant improvements in the retrieval performance are achieved. Hui Wang 0001, Sally I. McClean, Epaminondas Kapetanios, Denis Carroll |
CIKM | 3 |
| 2010 | Machine learning based Call Admission Control approaches: A comparative studyabstractThe importance of providing guaranteed Quality of Service (QoS) cannot be overemphasised, especially in the NGN environment which supports converged services on a common IP transport network. Call Admission Control (CAC) mechanisms do provide QoS to class-based services in a proactive manner. However, due to the factors of complexity, scale and dynamicity of NGN, Machine Learning techniques are favoured to analytical approaches for providing autonomous CAC. This paper is an effort to compare the performance of two such approaches - Neural Networks (NN) and Bayesian Networks (BN), to model the network behaviour and to estimate QoS metrics to be used in the CAC algorithm. It provides a way to find the optimum model training size for accurate predictions. Performance comparison is based on a wide range of experiments through a simulated network in Opnet. The outcome of this comparative study provides some interesting insights into the behaviour of NN and BN models and how they can be utilised for better CAC implementations. Abul Bashar, Gerard P. Parr, Sally I. McClean, Bryan W. Scotney, Detlef D. Nauck |
CNSM | 3 |
| 2010 | A framework for context-driven end-to-end QoS control in Converged NetworksabstractThis paper presents a framework for context-driven policy-based QoS control and end-to-end resource management in converged next generation networks. The Converged Networks QoS Framework (CNQF) is being developed within the IU-ATC project, and comprises distributed functional entities whose instances co-ordinate the converged network infrastructure to facilitate scalable and efficient end-to-end QoS management. The CNQF design leverages aspects of TISPAN, IETF and 3GPP policy-based management architectures whilst also introducing important innovative extensions to support context-aware QoS control in converged networks. The framework architecture is presented and its functionalities and operation in specific application scenarios are described. Suleiman Y. Yerima, Gerard P. Parr, Cathryn Peoples, Sally I. McClean, Philip J. Morrow |
CNSM | 4 |
| 2010 | Collaborative Filtering: The Aim of Recommender Systems and the Significance of User Ratings
Jennifer Louise Redpath, David H. Glass, Sally I. McClean, Liming Chen 0001 |
ECIR | 3 |
| 2010 | Knowledge Discovery Using Bayesian Network Framework for Intelligent Telecommunication Network Management
Abul Bashar, Gerard P. Parr, Sally I. McClean, Bryan W. Scotney, Detlef D. Nauck |
KSEM | 3 |
| 2010 | Incorporating Duration Information in Activity Recognition
Priyanka Chaurasia, Bryan W. Scotney, Sally I. McClean, Shuai Zhang 0001, Chris D. Nugent |
KSEM | 3 |
| 2010 | Measuring Tree Similarity for Natural Language Processing Based Information Retrieval
Zhiwei Lin 0002, Hui Wang 0001, Sally I. McClean |
NLDB | 3 |
| 2010 | A stability approach to convergence of curve evolution methods
Huaizhong Zhang, Philip J. Morrow, Sally I. McClean, Kurt Saetzler |
Pattern Recognit. Lett. | 3 |
| 2009 | MCMC-Based Algorithm to Adjust Scale Bias in Large Series of Electron Microscopical Ultrathin Sections
Huaizhong Zhang, E. Patricia Rodriguez, Philip J. Morrow, Sally I. McClean, Kurt Saetzler |
CAIP | 4 |
| 2009 | Clustering patient length of stay using mixtures of Gaussian models and phase type distributionsabstractGaussian mixture distributions and Coxian phase type distributions have been popular choices model based clustering of patients' length of stay data. This paper compares these models and presents an idea for a mixture distribution comprising of components of both of the above distributions. Also a mixed distribution survival tree is presented. A stroke dataset available from the English Hospital Episode Statistics database is used as a running example. Lalit Garg, Sally I. McClean, Brian Meenan, Elia El-Darzi, Peter H. Millard |
CBMS | 2 |
| 2009 | Neighborhood counting for financial time series forecastingabstractTime series data abound and analysis of such data is challenging and potentially rewarding. One example is financial time series analysis. Most of the intelligent data analysis methods can be applied in principle, but evolutionary computing is becoming increasingly popular and powerful. In this paper we focus on one task of financial time series analysis - stock price forecasting based on historical data. The premise of this task is that the current price of a stock is dependent on the price of the same stock in the past. Here we consider an additional assumption, i.e., time dependency relevance, that the price in the nearer past is more relevant to the current price than that in the more distant past. This assumption appears intuitively sound, but needs formally validated. In this paper we set to test this assumption by introducing time weighting into similarity measures, as similarity is one of the key notions in time series analysis methods including evolutionary computing. We consider the generic neighborhood counting similarity as it can be specialized for various forms of data by defining the notion of neighborhood in a way that satisfies different requirements. We do so with a view to capturing time weights in time series. This results in a novel time weighted similarity for time series. A formula is also discovered for the similarity so that it can be computed efficiently. Experiments show that this similarity outperforms the standard Euclidean distance and a time weighted variant of it. We conclude that the time dependency relevance assumption is sound. Zhiwei Lin 0002, Yu Huang 0009, Hui Wang 0001, Sally I. McClean |
IEEE Congress on Evolutionary Computation | 4 |
| 2009 | Evidential fusion of sensor data for activity recognition in smart homes
Chris D. Nugent, Maurice D. Mulvenna, Sally I. McClean, Bryan W. Scotney, Steven Devlin |
Pervasive Mob. Comput. | 4 |
| 2008 | Optimal Control of Patient Admissions to Satisfy Resource RestrictionsabstractAn effective admission policy is an essential part of successful management of a patient care system. This admission policy should consider the future availability of resources, e.g. number of beds or budgets available. All possible pathways through the whole patient care system need to be identified to better understand the process dynamics of a care system. In the current paper we show how such patient pathways can be used to predict the care resource requirements in the future. The patient pathways, developed using Markov chain modelling, are used to design an appropriate admission policy to meet future resource availability for the care system. We demonstrate an application of control theory to manage current admission rates to meet future scarce budgetary resource availability. This model can also be used to estimate the effect of an admission policy. Geriatric Patients' data from a London hospital are used to illustrate the approach. An optimization technique is also presented to reduce the implementation complexity and time. Lalit Garg, Sally I. McClean, Brian Meenan, Peter H. Millard |
CBMS | 2 |
| 2008 | Decision Support for Alzheimer's Patients in Smart HomesabstractAssistive technology in smart homes for elderly people with Alzheimer's disease is needed to support 'aging in place'. In this paper, we propose a probabilistic learning approach to characterise behavioural patterns for multi-inhabitants in smart homes. Decision support is then provided to monitor and assist patients to complete activities of daily living (ADL). Reasoning is based on the learned profiles and partially observed low-level sensors information. Data are stored in the proposed snow-flake schema based on homeML (an XML based schema for representation of information within smart homes). A laboratory has been developed for studying activities of 'making drinks' for multiple users. Evaluations of our learning and decision support approach are carried out on both real and simulated data. The potential of our approach to support assistive living and home-health monitoring of Alzheimer's patients is demonstrated. Shuai Zhang 0001, Sally I. McClean, Bryan W. Scotney, Chris D. Nugent, Maurice D. Mulvenna |
CBMS | 2 |
| 2008 | Assessment of the Impact of Sensor Failure in the Recognition of Activities of Daily Living
Chris D. Nugent, Maurice D. Mulvenna, Sally I. McClean, Bryan W. Scotney, Steven Devlin |
ICOST | 4 |
| 2008 | Integrating semantically heterogeneous aggregate views of distributed databases
Sally I. McClean, Bryan W. Scotney, Philip J. Morrow, Kieran Greer |
Distributed Parallel Databases | 1 |
| 2008 | Deriving Evidence Theoretical Functions in Multivariate Data Spaces: A Systematic ApproachabstractThe mathematical theory of evidence is a generalization of the Bayesian theory of probability. It is one of the primary tools for knowledge representation and uncertainty and probabilistic reasoning and has found many applications. Using this theory to solve a specific problem is critically dependent on the availability of a mass function (or basic belief assignment). In this paper, we consider the important problem of how to systematically derive mass functions from the common multivariate data spaces and also the ensuing problem of how to compute the various forms of belief function efficiently. We also consider how such a systematic approach can be used in practical pattern recognition problems. More specifically, we propose a novel method in which a mass function can be systematically derived from multivariate data and present new methods that exploit the algebraic structure of a multivariate data space to compute various belief functions including the belief, plausibility, and commonality functions in polynomial-time. We further consider the use of commonality as an equality check. We also develop a plausibility-based classifier. Experiments show that the equality checker and the classifier are comparable to state-of-the-art algorithms. Hui Wang 0001, Sally I. McClean |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2007 | Incorporating Knowledge into Unsupervised Model-Based Clustering for Satellite ImagesabstractThe identification and classification of landcover types from remotely sensed data is traditionally based on the assumption that pixels with similar spatial distribution patterns belong to the same spectral class. However, spectral data on its own has proven to be insufficient for classification. In addition, it is difficult to obtain enough accurate labelled samples from such data. Contextual data can be incorporated or fused' with spectral data to improve the estimation of class labels and therefore enhance the accuracy of the classification process as a whole when labelled data is not available. In this paper we use Dempster-Shafer theory of evidence to fuse the output of an unsupervised model-based clustering (MBC) technique and contextual data in the form of a digital elevation model. The final classification accuracy is shown to improve when using this approach. Bilal Al Momani, Sally I. McClean, Philip J. Morrow |
AICCSA | 2 |
| 2007 | Model-Based Segmentation of Multimodal Images
Sally I. McClean, Bryan W. Scotney, Philip J. Morrow |
CAIP | 2 |
| 2007 | Using Markov Models to Find Interesting Patient PathwaysabstractOver recent years the concept of Interestingness has come to underpin Data Mining, leading to the discovery of much new knowledge. In particular recognition of interesting patient pathways can lead to the discovery of important rules and patterns such as high probability pathways, groups of patients who incur exceptional high costs or pathways that are very long lasting. In the current paper we show how Markov models can be used to identify such patient pathways. Using Markov modelling we show how patient pathways may be extracted and describe an algorithm based on branch and bound that we have developed to efficiently extract a number of interesting pathways, subject to the number of pathways required, or some other criterion being specified. The approach is illustrated using data on geriatric patients from an administrative database of a London hospital, and we identify interesting pathways for geriatric patients. Such an approach might be used in association with healthcare process improvement technologies, such as Lean Thinking or Six Sigma. Sally I. McClean, Lalit Garg, Brian Meenan, Peter H. Millard |
CBMS | 1 |
| 2007 | A Lightweight, Scalable and Distributed Admission Control Algorithm for Voice TrafficabstractThe idea of carrying voice traffic on IP networks has been found to be very lucrative. However to achieve a quality similar to that offered by the existing telephone networks, the IP network should be in a position to honour the stringent quality of service (QoS) requirements of voice traffic. If traffic in excess of the network capacity is admitted into the network, QoS may be violated resulting in performance degradation. One way of providing sustained and consistent QoS is by regulating the number of voice calls that are admitted into the network such that the load on the network is less than or equal to the capacity of the network. This mechanism is known as call admission control. This paper proposes a scalable and distributed call admission control algorithm that operates at the network edge and factors the local passive measurements into the admission decision. Performance evaluation of the proposed algorithm through ns2 simulations reveals that it is successful in detecting rate mismatches (input rate greater than output rate) and subsequently rejecting admission requests (as long as input rate is greater than output rate) thereby delivering on the QoS guarantees demanded by voice applications. Moreover, it is simple and lightweight from an implementation perspective. Parag G. Kulkarni, Petre Dini, Sally I. McClean, Gerard P. Parr, Michaela M. Black |
ICC | 3 |
| 2007 | Applying statistical principles to data fusion in information retrievalabstractData fusion in information retrieval has been investigated by many researchers and quite a few data fusion methods have been proposed. However, their impact on effectiveness has not been well understood. In this paper, we apply statistical principles to data fusion and present a statistical data fusion model, which specifies the algorithm for fusion and conditions to be satisfied. The statistical model can be used as a guideline for data fusion methods. Based on this analysis, we compare CombSum and CombMNZ, which are the two best-known data fusion methods. We explain why sometimes CombMNZ does outperform Comb- Sum and what can be done to make CombSum more effective. Experimental results with TREC data are reported to support the conclusion that our enhancements to the algorithm improve effectiveness. Shengli Wu 0001, Yaxin Bi, Sally I. McClean |
SMC | 3 |
| 2007 | Result merging methods in distributed information retrieval with overlapping databases
Shengli Wu 0001, Sally I. McClean |
Inf. Retr. | 2 |
| 2006 | On Combining Multiple Classifiers Using an Evidential Approach
Yaxin Bi, Sally I. McClean, Terry J. Anderson |
AAAI | 2 |
| 2006 | Using Markov Models to Assess the Performance of a Health and Community Care SystemabstractMarkov chain modelling, in particular phase type modelling, has been previously used for hospital and community care systems, where we describe hospital care as a series of phases, such as acute, rehabilitation, or long-stay and likewise social care in the community may be modelled using phases such as dependent, convalescent, or nursing home. Such an approach allows us to adopt a holistic systems approach to health and community care modelling and management rather than focusing on the improvement of part of the system to the possible detriment of other components. We here extend this approach by showing how the general Markov framework can be exploited to extract various metrics of interest. In particular, we derive formulae for the mean and variance of the number of spells spent by a patient in hospital and in the community, and the expected total length of time in hospital and in community care, subsequent to first admission. Results are obtained for geriatric patients who have been admitted to hospital care. Finally we discuss how covariates, particularly time dependent covariates such as age, can be incorporated into the analysis Sally I. McClean, Malcolm J. Faddy, Peter H. Millard |
CBMS | 1 |
| 2006 | Evaluation of System Measures for Incomplete Relevance Judgment in IR
Shengli Wu 0001, Sally I. McClean |
FQAS | 2 |
| 2006 | Evidential Integration of Semantically Heterogeneous Aggregates in Distributed Databases with Imprecision
Sally I. McClean, Bryan W. Scotney, Philip J. Morrow |
IDEAL | 2 |
| 2006 | Performance prediction of data fusion for information retrieval
Shengli Wu 0001, Sally I. McClean |
Inf. Process. Manag. | 2 |
| 2006 | Improving high accuracy retrieval by eliminating the uneven correlation effect in data fusionabstractAbstract The aim of this research is twofold. On the one hand, high accuracy retrieval has been a concern of the information retrieval community for some time. We aim to investigate this issue via data fusion. On the other hand, the correlation among component results has been proven harmful to data fusion, but it has not been taken into account in data fusion algorithms. In the hope of achieving better performance, we propose a group of algorithms to eliminate the effect of uneven correlation among component results by assigning different weights to all component results or their combinations. Then the linear combination method or a variation is used for fusion. Extensive experimentation is carried out to evaluate the performances of these algorithms with six groups of component results, which are the top 10 systems submitted to Text REtrieval Conference (TREC) 6, 7, 8, 9, 2001, and 2002. The experimental results show that all eight data fusion methods involved outperform the best component system on average. Therefore, we demonstrate that the data fusion technique in general is effective with accurate retrieval results. The experimental results also demonstrate that all six methods presented in this article are effective for eliminating the effect of uneven correlation among component results. All of them outperform CombSum and five of them outperform CombMNZ on average. Shengli Wu 0001, Sally I. McClean |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2006 | Combining Wavelet Analysis and Bayesian Networks for the Classification of Auditory Brainstem ResponseabstractThe auditory brainstem response (ABR) has become a routine clinical tool for hearing and neurological assessment. In order to pick out the ABR from the background EEG activity that obscures it, stimulus-synchronized averaging of many repeated trials is necessary, typically requiring up to 2000 repetitions. This number of repetitions can be very difficult, time consuming and uncomfortable for some subjects. In this study, a method combining wavelet analysis and Bayesian networks is introduced to reduce the required number of repetitions, which could offer a great advantage in the clinical situation. 314 ABRs with 64 repetitions and 155 ABRs with 128 repetitions recorded from eight subjects are used here. A wavelet transform is applied to each of the ABRs, and the important features of the ABRs are extracted by thresholding and matching the wavelet coefficients. The significant wavelet coefficients that represent the extracted features of the ABRs are then used as the variables to build the Bayesian network for classification of the ABRs. In order to estimate the performance of this approach, stratified ten-fold cross-validation is used. Rui Zhang 0012, Gerry McAllister, Bryan W. Scotney, Sally I. McClean, Glen Houston |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2006 | Lightweight proactive queue managementabstractThe quest for better resource control has been the driving force behind Active Queue Management (AQM) research. Random Early Detection (RED), the defacto standard and its variants have been proposed as simple solutions to the AQM problem. These approaches, however, are known to suffer from problems like parameter sensitivity and inability to capture input traffic load fluctuations accurately, thereby resulting in instability. This paper presents a proactive queue management algorithm called PAQMAN that captures input traffic load fluctuations accurately and regulates the queue size around the desirable level. PAQMAN draws from the predictability in the underlying traffic by employing the Recursive Least Squares (RLS) algorithm to forecast the average queue size over the next prediction interval using the average queue size information of the past intervals. The packet drop probability is then computed as a function of this predicted average queue size. The performance of PAQMAN has been evaluated and compared against existing AQM schemes through ns-2 simulations that encompass varying network conditions for networks comprising of single as well as multiple bottleneck links. Simulation results demonstrate that PAQMAN maintains a relatively low queue size, while at the same time achieving high link utilization and low packet loss. Moreover, the computational overhead of PAQMAN is negligible (lightweight) which further justifies its use. Parag G. Kulkarni, Sally I. McClean, Gerard P. Parr, Michaela M. Black |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2005 | Markov Model-Based Clustering for Efficient Patient CareabstractPhase-type distributions were used to carry out model-based clustering of patients using the time spent by the patients in hospital, with maximum likelihood estimation of the model parameters. These parameters were allowed to vary with covariates so that the probability of cluster membership was dependent on these covariates. Expressions for the cluster membership probabilities and corresponding distributions of length of stay in care were found where the membership probabilities can be updated to take account of length of stay to date. The approach was applied to data on geriatric patients from an administrative database of a London hospital. The age of the patients at admission to care and the year of admission were included as covariates. Differential effects of these covariates on the various parameters of the fitted model were demonstrated, and interpretations of these effects made. The clusters here corresponded to patient pathways, with different length of stay distributions, varying care needs and different associated costs. By using the membership probabilities to assign patients to such clusters, care may thus be suited to their predicted pathway. Such an approach might be used in association with healthcare process improvement technologies, such as Lean Thinking or Six Sigma. Sally I. McClean, Malcolm J. Faddy, Peter H. Millard |
CBMS | 1 |
| 2005 | Classification of the Auditory Brainstem Response (ABR) Using Wavelet Analysis and Bayesian NetworkabstractThe auditory brainstem response (ABR) has become a routine clinical tool for hearing and neurological assessment. In order to pick out the ABR from the background EEG activity that obscures it, stimulus-synchronized averaging of many repeated trials is necessary and it typically requires up to 2000 repetitions. This number of repetitions can be very difficult, time consuming and uncomfortable for some subjects. In this study a method combining the wavelet analysis and the Bayesian network is introduced to reduce the required number of repetitions, which could offer a great advantage in the clinical situation. The important features of the ABR are extracted by thresholding and matching the wavelet coefficients. These extracted features are then used as the variables to build up the Bayesian network for classifying the ABR. 172 ABRs with 64 repetitions are applied in this study to learn the Bayesian network and estimate the conditional probability tables (CPTs). A further 142 ABRs with 64 repetitions are used to test the network. Moreover, this Bayesian network can also be applied to classify the ABRs with 128 repetitions. Rui Zhang 0012, Gerry McAllister, Bryan W. Scotney, Sally I. McClean, Glen Houston |
CBMS | 4 |
| 2005 | Data grid performance analysis through study of replication and storage infrastructure parametersabstractRunning data grid applications such as high energy nuclear physics (HENP) and weather modelling experiments involves working with huge data sets possibly of hundreds of Terabytes to Petabytes in size often kept over wide area networks. Data replication is a useful technique for reducing latency across communication networks over which the source data are accessed. As a starting point towards developing a multifaceted optimisation solution for data grids, this paper considers the effect of replication and storage parameter settings on data grid performance. The simulation results we obtained suggest that replication at local (Tier2) nodes has significant impact on data grid performance while cache settings at remote (Tier 1) node result in minimal performance improvement. Ernest Sithole, Gerard P. Parr, Sally I. McClean |
CCGRID | 3 |
| 2005 | Data Fusion with Correlation Weights
Shengli Wu 0001, Sally I. McClean |
ECIR | 2 |
| 2005 | Improving Classification Decisions by Multiple KnowledgeabstractAn important issue in data mining is how to make use of multiple discovered knowledge to improve future decisions. In this paper, we propose a new approach to combining multiple sets of rules for text categorization using Dempster's rule of combination. We develop a boosting-like technique for generating multiple sets of rules based on rough set theory and model classification decisions from multiple sets of rules as pieces of evidence which can be combined by Dempster's rule of combination. We apply these methods to 10 out of the 20-newsgroups - a benchmark data collection, individually and in combination. Our experimental results show that the performance of the best combination of the multiple sets of rules on the 10 groups of the benchmark data is statistically significantly better than that of the best single set of rules. The comparative analysis between the Dempster-Shafer and the majority voting methods along with an overfitting study confirm the advantage and the robustness of our approach. Yaxin Bi, Sally I. McClean, Terry J. Anderson |
ICTAI | 2 |
| 2005 | Knowledge discovery by probabilistic clustering of distributed databases
Sally I. McClean, Bryan W. Scotney, Philip J. Morrow, Kieran Greer |
Data Knowl. Eng. | 1 |
| 2004 | Combining Rules for Text Categorization Using Dempster's Rule of Combination
Yaxin Bi, Terry J. Anderson, Sally I. McClean |
IDEAL | 3 |
| 2004 | Using Domain Knowledge to Learn from Heterogeneous Distributed Databases
Sally I. McClean, Bryan W. Scotney, Mary Shapcott |
KES | 1 |
| 2004 | MISSION: An Agent-Based System for Semantic Integration of Heterogeneous Distributed Statistical Information Sources
Sally I. McClean, Bryan W. Scotney, Hans Rutjes, Jannes Hartkamp, Isambo Karali, Michael Hatzopoulos, Joanne Lamb, Defeng Ma |
SSDBM | 1 |
| 2004 | Knowledge Discovery from Databases on the Semantic Web
Bryan W. Scotney, Sally I. McClean |
SSDBM | 2 |
| 2003 | Can the Grid Help to Solve the Data Integration Problems in Molecular Biology?abstractMolecular biology is increasingly relying on globally distributed information repositories. The quality and performance of research and development in this field will depend on highly flexible, performant and intuitive ways of accessing and integrating these resources into local information processing environments. In contrast to many ongoing Grid developments, the molecular biology community requires high-informational as opposed to high-performance computing Grids. This paper outlines a general distributed informational computing scenario in the context of molecular biology. It focuses on information integration (data warehousing) and data mining as two crucial Grid methodologies of future molecular biology computing environments. The presented scenario raises a number of questions, as the basis for discussion, on the role of emerging Grid technologies. Brian Sturgeon, Damian McCourt, John Cowper, Fiona Palmer, Sally I. McClean, Werner Dubitzky |
CCGRID | 5 |
| 2003 | Metadata with a MISSION: Using Metadata to Query Distributed Statistical Meta-information Syste
Sally I. McClean, Bryan W. Scotney, Hans Rutjes |
Dublin Core Conference | 1 |
| 2003 | Evaluation of inherent performance of intelligent medical decision support systems: utilising neural networks as an example
Ann E. Smith, Chris D. Nugent, Sally I. McClean |
Artif. Intell. Medicine | 3 |
| 2003 | Database aggregation of imprecise and uncertain evidence
Bryan W. Scotney, Sally I. McClean |
Inf. Sci. | 2 |
| 2003 | A rough set model with ontologies for discovering maximal association rules in document collections
Yaxin Bi, Terry J. Anderson, Sally I. McClean |
Knowl. Based Syst. | 3 |
| 2003 | Learning temporal concepts from heterogeneous data sequences
Sally I. McClean, Bryan W. Scotney, Fiona Palmer |
Soft Comput. | 1 |
| 2003 | A Scalable Approach to Integrating Heterogeneous Aggregate Views of Distributed DatabasesabstractAggregate views are commonly used for summarizing information held in very large databases such as those encountered in data warehousing, large scale transaction management, and statistical databases. Such applications often involve distributed databases that have developed independently and therefore may exhibit incompatibility, heterogeneity, and data inconsistency. We are here concerned with the integration of aggregates that have heterogeneous classification schemes where local ontologies, in the form of such classification schemes, may be mapped onto a common ontology. In previous work, we have developed a method for the integration of such aggregates; the method previously developed is efficient, but cannot handle innate data inconsistencies that are likely to arise when a large number of databases are being integrated. In this paper, we develop an approach that can handle data inconsistencies and is thus inherently much more scalable. In our new approach, we first construct a dynamic shared ontology by analyzing the correspondence graph that relates the heterogeneous classification schemes; the aggregates are then derived by minimization of the Kullback-Leibler information divergence using the EM (Expectation-Maximization) algorithm. Thus, we may assess whether global queries on such aggregates are answerable, partially answerable, or unanswerable in advance of computing the aggregates themselves. Sally I. McClean, Bryan W. Scotney, Kieran Greer |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2002 | Conceptual Clustering of Heterogeneous Sequences via Schema Mapping
Sally I. McClean, Bryan W. Scotney, Fiona Palmer |
ISMIS | 1 |
| 2002 | A Negotiation Agent for Distributed Heterogeneous Statistical DatabasesabstractThe World-Wide Web provides an ever-increasing source of diverse information. We focus on query agents, in particular the matching and negotiation agents that are responsible for pre-integration where the matching agent decomposes the query into sub-queries, and then searches metadata to find datasets that match the query fragments. In the case of heterogeneous data, the matching agent utilises a negotiation agent to find datasets that match the query fragments, provides mappings from the data to the query, and constructs the appropriate (sub-)query re-writing rules. Such matching is done by generalising the data and testing if the (sub) query is matchable to the generalised (meta) data: we call this g-matchable; if it is then we can construct an operator stack to transform the data to match the (sub) query. Such an approach provides a capability of automating the process of executing queries on heterogeneous statistical databases that are distributed over the Internet. The novelty lies in the provision of automated methods for statistical aggregates, where the heterogeneity essentially resides in the classification schemes of categorical data, including both heterogeneity of nomenclature and heterogeneity of granularity. In addition, our solution permits queries to be specified in a goal-driven query-by-example format. Rather than impose an a priori global standard, the user can query through a unified interface where integration is done at run-time. Sally I. McClean, Rónán Páircéir, Bryan W. Scotney, Kieran Greer |
SSDBM | 1 |
| 2002 | Learning with Concept Hierarchies in Probabilistic Relational Data Mining
Mary Shapcott, Sally I. McClean, Kenneth Adamson |
WAIM | 3 |
| 2001 | A data mining approach to the prediction of corporate failure
Feng Yu Lin, Sally I. McClean |
Knowl. Based Syst. | 2 |
| 2001 | Aggregation of Imprecise and Uncertain Information in DatabasesabstractInformation stored in a database is often subject to uncertainty and imprecision. Probability theory provides a well-known and well understood way of representing uncertainty and may thus be used to provide a mechanism for storing uncertain information in a database. We consider the problem of aggregation using an imprecise probability data model that allows us to represent imprecision by partial probabilities and uncertainty using probability distributions. Most work to date has concentrated on providing functionality for extending the relational algebra with a view to executing traditional queries on uncertain or imprecise data. However, for imprecise and uncertain data, we often require aggregation operators that provide information on patterns in the data. Thus, while traditional query processing is tuple-driven, processing of uncertain data is often attribute-driven where we use aggregation operators to discover attribute properties. The aggregation operator that we define uses the Kullback-Leibler information divergence between the aggregated probability distribution and the individual tuple values to provide a probability distribution for the domain values of an attribute or group of attributes. The provision of such aggregation operators is a central requirement in furnishing a database with the capability to perform the operations necessary for knowledge discovery in databases. Sally I. McClean, Bryan W. Scotney, Mary Shapcott |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2000 | Discovery of multi-level rules and exceptions from a distributed databaseabstractArticle Discovery of multi-level rules and exceptions from a distributed database Share on Authors: Rónán Páircéir School of Information and Software Engineering, Faculty of Informatics, University of Ulster, Cromore Road, Coleraine, BT52 1SA, Northern Ireland School of Information and Software Engineering, Faculty of Informatics, University of Ulster, Cromore Road, Coleraine, BT52 1SA, Northern IrelandView Profile , Sally McClean School of Information and Software Engineering, Faculty of Informatics, University of Ulster, Cromore Road, Coleraine, BT52 1SA, Northern Ireland School of Information and Software Engineering, Faculty of Informatics, University of Ulster, Cromore Road, Coleraine, BT52 1SA, Northern IrelandView Profile , Bryan Scotney School of Information and Software Engineering, Faculty of Informatics, University of Ulster, Cromore Road, Coleraine, BT52 1SA, Northern Ireland School of Information and Software Engineering, Faculty of Informatics, University of Ulster, Cromore Road, Coleraine, BT52 1SA, Northern IrelandView Profile Authors Info & Claims KDD '00: Proceedings of the sixth ACM SIGKDD international conference on Knowledge discovery and data miningAugust 2000 Pages 523–532https://doi.org/10.1145/347090.347196Online:01 August 2000Publication History 4citation767DownloadsMetricsTotal Citations4Total Downloads767Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Rónán Páircéir, Sally I. McClean, Bryan W. Scotney |
KDD | 2 |
| 2000 | Learning Dynamic Bayesian Belief Networks Using Conditional Phase-Type Distributions
Adele H. Marshall, Sally I. McClean, Mary Shapcott, Peter H. Millard |
PKDD | 2 |
| 2000 | Exploring dynamic Bayesian belief networks for intelligent fault management systemsabstractSystems that are subject to uncertainty in their behaviour are often modelled by Bayesian belief networks (BBNs). These are probabilistic models of the system in which the independence relations between the variables of interest are represented explicitly. A directed graph is used, in which two nodes are connected by an edge if one is a 'direct cause' of the other. However the Bayesian paradigm does not provide any direct means for modelling dynamic systems. There has been a considerable amount of research effort in recent years to address this. We review these approaches and propose a new dynamic extension to the BBN. Our discussion then focuses on fault management of complex telecommunications and how the dynamic Bayesian models can assist in the prediction of faults. Roy Sterritt, Adele H. Marshall, Mary Shapcott, Sally I. McClean |
SMC | 4 |
| 2000 | Using Background Knowledge in the Aggregation of Imprecise Evidence in Databases
Sally I. McClean, Bryan W. Scotney, Mary Shapcott |
Data Knowl. Eng. | 1 |
| 2000 | Rule discovery for event histories
Sally I. McClean, Bryan W. Scotney, Mary Shapcott |
Intell. Data Anal. | 1 |
| 2000 | Knowledge discovery in distributed databases using evidence theoryabstractDistributed databases allow us to integrate data from different sources which have not previously been combined. The Dempster–Shafer theory of evidence and evidential reasoning are particularly suited to the integration of distributed databases. Evidential functions are suited to represent evidence from different sources. Evidential reasoning is carried out by the well-known orthogonal sum. Previous work has defined linguistic summaries to discover knowledge by using fuzzy set theory and using evidence theory to define summaries. In this paper we study linguistic summaries and their applications to knowledge discovery in distributed databases. © 2000 John Wiley & Sons, Inc. D. Cai, Michael F. McTear, Sally I. McClean |
Int. J. Intell. Syst. | 3 |
| 2000 | Incorporating domain knowledge into attribute-oriented data miningabstractIt is frequently the case that data mining is carried out in an environment which contains noisy and missing data. This is particularly likely to be true when the data were originally collected for different purposes, as is commonly the case in data warehousing. In this paper we discuss the use of domain knowledge, e.g., integrity constraints or a concept hierarchy, to re-engineer the database and allocate sets to which missing or unacceptable outlying data may belong. Attribute-oriented knowledge discovery has proved to be a powerful approach for mining multi-level data in large databases. Such methods are set-oriented in that attribute values are considered to belong to subsets of the domain. These subsets may be provided directly by the database or derived from a knowledge base using inductive logic programming to re-engineer the database. In this paper we develop an algorithm which allows us to aggregate imprecise data and use it for multi-level rule induction and knowledge discovery. ©2000 John Wiley & Sons, Inc. Sally I. McClean, Bryan W. Scotney, Mary Shapcott |
Int. J. Intell. Syst. | 1 |
| 1999 | Automated Discovery of Rules and Exeptions from Distributed Databases Using Aggregates
Rónán Páircéir, Sally I. McClean, Bryan W. Scotney |
PKDD | 2 |
| 1999 | Optimal and Efficient Integration of Heterogeneous Summary Tables in a Distributed Database
Bryan W. Scotney, Sally I. McClean, Máire Rodgers |
Data Knowl. Eng. | 2 |
| 1999 | Efficient knowledge discovery through the integration of heterogeneous data
Bryan W. Scotney, Sally I. McClean |
Inf. Softw. Technol. | 2 |
| 1998 | Aggregation of Imprecise and Uncertain Information for Knowledge Discovery in Databases
Sally I. McClean, Bryan W. Scotney, Mary Shapcott |
KDD | 1 |
| 1997 | Using evidence theory for the integration of distributed databasesabstractDistributed databases allow us to integrate data from different sources which have not previously been combined. In this article, we are concerned with the situation where the data sources are held in a distributed database. Integration of the data is then accomplished using the Dempster–Shafer representation of evidence. The weighted sum operator is developed and this operator is shown to provide an appropriate mechanism for the integration of such data. This representation is particularly suited to statistical samples which may include missing values and be held at different levels of aggregation. Missing values are incorporated into the representation to provide lower and upper probabilities for propositions of interest. The weighted sum operator facilitates combination of samples with different classification schemes. Such a capability is particularly useful for knowledge discovery when we are searching for rules within the concept hierarchy, defined in terms of probabilities or associations. By integrating information from different sources, we may thus be able to induce new rules or strengthen rules which have already been obtained. We develop a framework for describing such rules and show how we may then integrate rules at a high level without having to resort to the raw data, a useful facility for knowledge discovery where efficiency is of the essence. © 1997 John Wiley & Sons, Inc. Sally I. McClean, Bryan W. Scotney |
Int. J. Intell. Syst. | 1 |
| 1992 | Framework for query optimization in distributed statistical databases
Mohammad Hadi Sadreddini, David A. Bell, Sally I. McClean |
Inf. Softw. Technol. | 3 |
| 1990 | The use of simulated annealing for clustering data in databases
F. J. McErlean, David A. Bell, Sally I. McClean |
Inf. Syst. | 3 |
| 1990 | Application of simulated annealing to clustering tuples in databasesabstractTechniques for general purpose optimization have been derived from the Metropolis Monte Carlo method of simulating the behavior of particles in substances as they are slowly cooled to form crystals. Simulated Annealing is such a derivative and its value for placement problems (e.g., in circuit board layout design) suggests that it could be advantageously applied to clustering tuples in databases in order to enhance responsiveness to queries. In this article we investigate this issue and compare the performance of this technique with a Graph-Collapsing clustering method which is known to perform very well, in order to gain insights into which approach is better for incorporation in a performance-oriented database design tool. We judge that, whilst the new method does give superior results to the graph-based method in many cases, these improvements are gained at such a very considerable expense of algorithm run time as to rule the new technique out of our consideration as a real world general purpose design tool (but perhaps not for some special-purpose databases). © 1990 John Wiley & Sons, Inc. David A. Bell, F. J. McErlean, P. M. Stewart, Sally I. McClean |
J. Am. Soc. Inf. Sci. | 4 |
| 1989 | Pragmatic Estimation of Join Sizes and Attribute CorrelationsabstractA method is presented for modeling attribute value distributions in database relations for the purpose of obtaining accurate estimates of intermediate relation sizes during query evaluation. The basic idea is that instead of keeping a single (average) value to represent the number of occurrences of each attribute value, m (typically ten) parameters are kept, each representing the number of occurrences of attribute values in a piece, or partition, corresponding to a subrange of 1/mth of the original value range. The uniformity assumption, taken as an estimation technique rather than as an assumption, holds for each partition, hence the name piecewise uniform. The distribution method is extended to the modeling of important intrarelational attribute correlations. This and other enhancements to the technique such as application to semijoin operation are suggested. The technique is being used on two multidatabase management systems.> David A. Bell, Daniel Hiak Ong Ling, Sally I. McClean |
ICDE | 3 |