Yulia Hicks

dblp:h/YuliaHicks · also I. A. Karaulova, Y. A. Hicks, Yulia A. Hicks · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-7179-4587ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-authorHuman-computer interaction and ubiquitous computing · 5Applied, interdisciplinary, general and emerging computing · 5Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Deep Learning for Quality Assessment of Echocardiographic Images
abstract
The growing need for standardised and automated cardiac ultrasound (US) acquisition has driven the integration of deep learning into echocardiographic workflows. While existing deep learning (DL) models have shown promising results in tasks such as view classification and image quality assessment, most of these approaches focus either on differentiating among standard views or grading image quality within a standard view. However, these methods lack the capacity to model the sequential spatial transitions that occur during the acquisition process, limiting their applicability to real-time probe guidance and robotic control. To address this gap, we propose a classification framework designed for the process of acquiring the parasternal long-axis (PLAX) view under a fixed scanning protocol. Based on extensive probe movement experiments across multiple patients, we identified four representative echocardiographic views that appear during the search for the optimal PLAX position. These views correspond to distinct probe positions and orientations and reflect varying levels of image completeness. A dataset of 7,200 annotated images was used to train a ResNet50-based deep network for multi-class classification. The model achieved robust performance with accuracy, sensitivity, specificity, and F1 scores above 89%, and AUC exceeding 97% in patient-level cross-validation. It effectively captures spatially relevant features, distinguishes subtle view differences, and generalizes well to unseen data. The outputs provide interpretable feedback correlating image quality with probe position, enabling real-time scanning assessment. Furthermore, this work introduces a novel problem formulation and multi-class view classification under a fixed acquisition protocol. It provides a foundation for developing the next generation of intelligent US systems. By linking image classification to probe position and orientation, the proposed framework enables real-time feedback that can ultimately support autonomous scanning agents in locating diagnostically optimal cardiac views.
Shuping Kang, Yulia Hicks, Rossitza Setchi
KES2
2025 MOVEXor: Motion-Video Attention Explainer for Low Back Pain Classification
abstract
Accurate classification of Movement Impairment (MI) and Motor Control Impairment (MCI) in non-specific low back pain (NSLBP) is essential for targeted rehabilitation but remains challenging due to subjective assessments and subtle movement differences. We present MOVEXor, a lightweight and explainable multi-modal framework that integrates spinal curvature images and motion-derived features through a modality-aware attention gating mechanism. MOVEXor achieves high classification performance (up to 97.5% accuracy) while offering transparent decision-making via Grad-CAM and Integrated Gradients (IG). Our analysis shows that the model focuses on physiologically meaningful movement phases, particularly minimal flexion angle, and relies heavily on motion stability for classification. The fused attention-based design outperforms static fusion methods, especially when handling noisy inputs. With minimal hardware requirements and real-time explainability, MOVEXor holds strong potential as a clinical decision-support tool for both in-clinic and remote settings, enabling objective, interpretable, and personalised rehabilitation exercise of LBP subgroups.
Zebang Liu, Yulia Hicks, Liba Sheeran
KES2
2025 Enhancing Clinical Decision Support with LLMs: A Feasibility Study Integrating CoT, RAG, and QLoRA
abstract
The integration of artificial intelligence (AI) into healthcare holds significant promise for enhancing clinical decision-making and improving patient outcomes. However, general-purpose large language models (LLMs) frequently exhibit limitations such as hallucinations, lack of domain-specific accuracy, and opaque reasoning processes, posing risks in clinical applications. This study addressed these challenges by proposing and exploring an innovative integration of Chain-of-Thought (CoT) prompting, Retrieval-Augmented Generation (RAG), and parameter-efficient fine-tuning using Quantized Low-Rank Adaptation (QLoRA) specifically tailored for medical use. A distilled 14-billion-parameter variant of the DeepSeek R1 model was fine-tuned using a structured clinical dataset that emphasizes step-by-step reasoning. Additionally, external medical references were incorporated through RAG, employing embedding models for precise context retrieval. The combination of these techniques was systematically evaluated on a challenging set of open-ended medical questions, resulting in accuracy improvements—from a baseline accuracy of 55% to a final performance of 81%. Further qualitative evaluation involving three practising General Practitioners (GPs) and three fourth-year medical students from Cardiff University underscored the proposed system’s clinical utility and transparent reasoning capabilities while also identifying areas for improvement, such as conciseness and explicit adherence to national-specific clinical guidelines. This research demonstrated that integrating CoT, RAG, and QLoRA provides a practical pathway toward reliable, transparent, and clinically relevant AI support for healthcare professionals. Recommendations for future work include scaling models, incorporating comprehensive patient data, and enhancing customization for clinical application contexts.
George Matthew, Yulia Hicks
KES2
2025 Human intention recognition using context relationships in complex scenes
abstract
Recognizing human intentions is a key challenge in human-robot interaction research. Much of the current work in this area centers on identifying human intentions within specific activities, often relying on a limited set of features. In contrast, this paper introduces a more versatile framework for intention recognition and introduces a novel model: the Spatial-Temporal Graph Attention Informer Neural Network (STGAIN). To recognize intentions, this model leverages spatial relationships between humans and objects in different scenes, along with their temporal evolution. In addition, to address an existing research gap, this research developed a new dataset called Dynamic Scene Graph (DSG) with representative dynamic relationships, derived from 471 videos covering 20 categories of human intentions. This dataset represents people and objects in different scenes, and the relationships between them. The model was tested rigorously at different points in the videos to track how the scenes evolved and to assess prediction accuracy, comparing the results to a range of advanced algorithms. Our findings clearly demonstrate that STGAIN outperforms these models, showcasing its potential for advanced human intention recognition applications. This model represents a significant advance toward creating more human-centered robots, capable of understanding and adapting to human intentions in real-world situations.
Rossitza Setchi, Yulia Hicks
Expert Syst. Appl.3
2024 Simulation-based dataset acquisition for robotic cardiac ultrasound examinations
abstract
Over the past decade, automating ultrasound scanning has been the subject of intense research. However, training a robot to perform automated ultrasound examinations requires a substantial corpus of training data. Traditionally, researchers have sought to obtain such data either through publicly available datasets or by engaging professional sonographers in experiments aimed at dataset generation. The former approach often yields incomplete datasets insufficient for specialized research objectives, while the latter entails logistical challenges, including the necessity for frequent manual experimentation and ready access to medical professionals. Therefore, the acquisition of a comprehensive and suitable dataset remains an essential yet formidable challenge. Here, we propose a novel framework for achieving the automated acquisition of cardiac ultrasound datasets by controlling robot behaviour using a digital twin within a simulated environment. This framework consists of two modules: physical and virtual. Within the virtual simulation module, diverse body models of varying dimensions can be inputted, enabling the planning of robot arm path and ultrasound scanning manoeuvrers. Then the physical robot arm clones the actions of the robot in the simulation environment and updates its current state in the virtual module. The proposed framework was used to collect 43,000 cardiac ultrasound images from 8 patients with different pathologies and 1 healthy individual using a KUKA LBR Med robot and Intelligent Ultrasound Simulator. It is also expected to be feasible for a real-person dataset collection.
Shuping Kang, Thomas Daniels 0003, Rossitza Setchi, Yulia Hicks
KES4
2024 SpineSighter: An AI-Driven Approach for Automatic Classification of Spinal Function from Video
abstract
Low Back Pain (LBP) is a prevalent musculoskeletal disorder affecting over 80% of the population over their lifetime and is a leading cause of disability globally. The most frequent type, non-specific LBP (NSLBP) does not have a clearly identifiable pathology cause. Current clinical guidelines advocate for tailored management and self-care approaches for NSLBP. The effectiveness of these personalised management plans significantly depends on accurate and on-going assessment of the patient’s spinal function. This presents considerable challenges for both clinicians and patients. This study introduces “SpineSighter”, an artificial intelligence (AI) model developed to tailor management of NSLBP by categorising patients based on their spinal function either into High Function (HF) and Low Function (LF) subsets. Utilising standard video recordings and computer vision technology, SpineSighter analyses motion features such as angular displacement, velocity, and acceleration during repeated forward flexion tests. The model showed high accuracy in classifying spinal function, achieving an accuracy of 95.13%, sensitivity of 93.81%, specificity of 96.00%, and an F1 score of 0.9442. This innovative use of AI highlights the importance of velocity as a critical indicator of spinal functional differences, opening new avenues for personalised clinical management, self-care and recovery strategies of NSLBP.
Zebang Liu, Yulia Hicks, Liba Sheeran
KES2
2024 Indoor human activity recognition based on context relationships
abstract
Human activity recognition, as a significant branch of artificial intelligence, requires increasingly generalized and precise methodologies due to growing demands. Therefore, this paper proposes a context-based method for recognising indoor human activities, which interlinks indoor human activities with interactions with objects, making the contextual relationship between humans and objects particularly crucial. In addition, this research has developed a new dynamic graph dataset based on publicly available video datasets and their associated descriptive scripts, instantiating the relationships between humans and objects. A novel architecture for human activity recognition is developed in this research. This architecture utilizes graph neural networks and self-attention mechanisms to learn the significance of the interactions between humans and objects and capture the relationships between video frames on a temporal level. The results demonstrate that the classification accuracy reaches 0.86 and it also performs better than other current advanced algorithms STGAT and STGCN. It is noteworthy that the approach also effectively reduces ambiguity in activity recognition.
Rossitza Setchi, Yulia Hicks
KES3
2024 Guest Editorial: Special Issue on the British Machine Vision Conference 2022
Guang Yang 0006, Angelica I. Avilés-Rivero, Yingying Fang, Zhenhua Feng 0001, Gianluigi Ciocca, Yulia Hicks, Constantino Carlos Reyes-Aldasoro
Int. J. Comput. Vis.6
2022 XAI & I: Self-explanatory AI facilitating mutual understanding between AI and human experts
abstract
Traditionally, explainable artificial intelligence seeks to provide explanation and interpretability of high-performing black-box models such as deep neural networks. Interpretation of such models remains difficult, because of their high complexity. An alternative method is to instead force a deep-neural network to use human-intelligible features as the basis for its decisions. We tested this approach using the natural category domain of rock types. We compared the performance of a black-box implementation of transfer-learning using Resnet50 to that of a network first trained to predict expert-identified features and then forced to use these features to categorise rock images. The performance of this feature-constrained network was virtually identical to that of the unconstrained network. Further, a partially constrained network forced to condense down to a small number of features that was not trained with expert features did not result in these abstracted features being intelligible; nevertheless, an affine transformation of these features could be found that aligned well with expert-intelligible features. These findings show that making an AI intrinsically intelligible need not be at the cost of performance.
Jacques A. Grange, Henrijs Princis, Theodor R. W. Kozlowski, Aissa Amadou-Dioffo, Yulia Hicks, Mark K. Johansen
KES6
2022 Context change and triggers for human intention recognition
abstract
In human-robot interaction, understanding human intention is important to smooth interaction between humans and robots. Proactive human-robot interactions are the trend. They rely on recognising human intentions to complete tasks. The reasoning is accomplished based on the current human state, environment and context, and human intention recognition and prediction. Many factors may affect human intention, including clues which are difficult to recognise directly from the action but may be perceived from the change in the environment or context. The changes that affect human intention are the triggers and serve as strong evidence for identifying human intention. Therefore, detecting such changes and identifying such triggers are the promising approach to assist in human intention recognition. This paper discusses the current state of art in human intention recognition in human-computer interaction and illustrates the importance of context change and triggers for human intention recognition in a variety of examples.
Rossitza Setchi, Yulia Hicks
KES3
2021 Development of Recreation Game for Measurement of Eye Movement Using Tangram
abstract
The increasing number of dementia patients is one of the major social problems in Japan. Early detection and prevention of dementia is important. Many welfare facilities use check tests to measure the progression of dementia. However, some elderly people can be very nervous about assessment tests. In addition, the evaluation tests should be conducted regularly to assess the cognitive function of the patient over time, which can be a huge burden for medical and care providers. On the other hand, research papers have recently reported that it is possible to measure cognitive functions by focusing on brain functions and eye movements. In this paper, the authors aimed to develop a new dementia evaluation system to reduce the burden on medical staff and elderly persons. As the first step of this project, we focused on eye movement and employed a simple puzzle game to collect a patient’s eye movement. Because of COVID-19, we could not conduct experiments at care houses; instead, we conducted a preliminary experiment with healthy subjects and collected eye movement data during the puzzle game.
Rise Morimoto, Hiroharu Kawanaka, Yulia Hicks, Rossitza Setchi
KES3
2020 Automatic Low Back Pain Classification Using Inertial Measurement Units: A Preliminary Analysis
abstract
Low back pain (LBP) is a major health problem that has now become leading cause of disability worldwide. The majority of LBP has no specific pathological cause. Classification of non-specific LBP (NSLBP) into subgroups corresponding to the reported symptoms has been identified as an essential step towards the provision of personalised management and rehabilitation plans. Currently, clinicians classify low back pain patients into clinical subgroups based on clinical judgement and expertise, which is a time-consuming process open to human error. This paper introduces a novel approach for automatic classification of NSLBP patients into clinical subgroups on the basis of the MTw2 inertial measurement unit (MTw2 IMU tracker) motion data, which are portable units and thus desirable for clinical use. Four MTw2 IMU trackers tracking movement during a number of physical assessment tests were investigated in their ability to distinguish between clinically recognized NSLBP subgroups. Simple motion features such as the angular range of displacement were used in classification experiments to reflect how clinicians make decisions when classifying NSLBP. The achieved results were comparable to the state of art results in automatic NSLBP classification using optical motion capture data and demonstrated the feasibility of developing an automatic classification system on the basis of the MTw2 IMU tracker motion data obtained with an individual performing a battery of standard physical assessment tests. Further developments could address gaps in current medical and engineering literature and improve clinical outcomes.
Zoë Bacon, Yulia Hicks, Mohammad Al-Amri, Liba Sheeran
KES2
2020 Significant Features of Hand Motion for Dementia Evaluation in the Simple Recreation Game
abstract
Increasing the number of dementia patients is one of the big social problems in Japan. Early detection and prevention of dementia are essential. Medical specialists and care workers usually use some dementia evaluation methods. However, some elderly persons often become very nervous about the evaluation tests. Also, these evaluation tests should be conducted continually to assess a patient’s cognitive function. On the other hand, recently, research papers on the relationship between simple human actions and a patient’s cognitive function have been reported. The authors focused on hand dexterity and movement of both hands because these factors have some relationships to cognitive functions. We developed a new recreational system to evaluate hand dexterity and activity of both hands to assess a subject’s cognitive functions. Also, the authors discussed the relationship between the extracted features and a subject’s cognitive functions. As a result of experiments, we determined the feature descriptor(s) about hand motion to estimate and evaluate a patient’s cognitive function.
Kanta Umemura, Hiroharu Kawanaka, Yulia Hicks, Rossitza Setchi
KES3
2020 Bimodal Automated Carotid Ultrasound Segmentation Using Geometrically Constrained Deep Neural Networks
abstract
For asymptomatic patients suffering from carotid stenosis, the assessment of plaque morphology is an important clinical task which allows monitoring of the risk of plaque rupture and future incidents of stroke. Ultrasound Imaging provides a safe and non-invasive modality for this, and the segmentation of media-adventitia boundaries and lumen-intima boundaries of the Carotid artery form an essential part in this monitoring process. In this paper, we propose a novel Deep Neural Network as a fully automated segmentation tool, and its application in delineating both the media-adventitia boundary and the lumen-intima boundary. We develop a new geometrically constrained objective function as part of the Network's Stochastic Gradient Descent optimisation, thus tuning it to the problem at hand. Furthermore, we also apply a bimodal fusion of amplitude and phase congruency data proposed by us in previous work, as an input to the network, as the latter provides an intensity-invariant data source to the network. We finally report the segmentation performance of the network on transverse sections of the carotid. Tests are carried out on an augmented dataset of 81,000 images, and the results are compared to other studies by reporting the DICE coefficient of similarity, modified Hausdorff Distance, sensitivity and specificity. Our proposed modification is shown to yield improved results on the standard network over this larger dataset, with the advantage of it being fully automated. We conclude that Deep Neural Networks provide a reliable trained manner in which carotid ultrasound images may be automatically segmented, using amplitude data and intensity invariant phase congruency maps as a data source.
Carl Azzopardi, Kenneth P. Camilleri, Yulia Hicks
IEEE J. Biomed. Health Informatics3
2019 A Pilot Study on Detecting Violence in Videos Fusing Proxy Models
Marc Roig Vilamala, Liam Hiley, Yulia Hicks, Alun D. Preece, Federico Cerutti 0001
FUSION3
2018 Intelligent Signal Processing Mechanisms for Nuanced Anomaly Detection in Action Audio-Visual Data Streams
abstract
We consider the problem of anomaly detection in an audiovisual analysis system designed to interpret sequences of actions from visual and audio cues. The scene activity recognition is based on a generative framework, with a high-level inference model for contextual recognition of sequences of actions. The system is endowed with anomaly detection mechanisms, which facilitate differentiation of various types of anomalies. This is accomplished using intelligence provided by a classifier incongruence detector, classifier confidence module and data quality assessment system, in addition to the classical outlier detection module. The paper focuses on one of the mechanisms, the classifier incongruence detector, the purpose of which is to flag situations when the video and audio modalities disagree in action interpretation. We demonstrate the merit of using the Delta divergence measure for this purpose. We show that this measure significantly enhances the incongruence detection rate in the Human Action Manipulation complex activity recognition data set.
Josef Kittler, Ioannis Kaloskampis, Cemre Zor, Yulia Hicks, Wenwu Wang 0001
ICASSP5
2018 A 3D morphometric perspective for facial gender analysis and classification using geodesic path curvature features
abstract
The relationship between the shape and gender of a face, with particular application to automatic gender classification, has been the subject of significant research in recent years. Determining the gender of a face, especially when dealing with unseen examples, presents a major challenge. This is especially true for certain age groups, such as teenagers, due to their rapid development at this phase of life. This study proposes a new set of facial morphological descriptors, based on 3D geodesic path curvatures, and uses them for gender analysis. Their goal is to discern key facial areas related to gender, specifically suited to the task of gender classification. These new curvature-based features are extracted along the geodesic path between two biological landmarks located in key facial areas. Classification performance based on the new features is compared with that achieved using the Euclidean and geodesic distance measures traditionally used in gender analysis and classification. Five different experiments were conducted on a large teenage face database (4745 faces from the Avon Longitudinal Study of Parents and Children) to investigate and justify the use of the proposed curvature features. Our experiments show that the combination of the new features with geodesic distances provides a classification accuracy of 89%. They also show that nose-related traits provide the most discriminative facial feature for gender classification, with the most discriminative features lying along the 3D face profile curve.
Hawraa Abbas, Yulia Hicks, David Marshall 0001, Alexei I. Zhurov, Stephen Richmond
Comput. Vis. Media2
2018 Error sensitivity analysis of Delta divergence - a novel measure for classifier incongruence detection
abstract
The state of classifier incongruence in decision making systems incorporating multiple classifiers is often an indicator of anomaly caused by an unexpected observation or an unusual situation. Its assessment is important as one of the key mechanisms for domain anomaly detection. In this paper, we investigate the sensitivity of Delta divergence, a novel measure of classifier incongruence, to estimation errors. Statistical properties of Delta divergence are analysed both theoretically and experimentally. The results of the analysis provide guidelines on the selection of threshold for classifier incongruence detection based on this measure.
Josef Kittler, Cemre Zor, Ioannis Kaloskampis, Yulia Hicks, Wenwu Wang 0001
Pattern Recognit.4
2017 Clock Drawing Test Interpretation System
abstract
A clock drawing test (CDT) is a neurological test used for the assessment of cognitive impairment based on sketches of a clock completed by a patient. Usually, a medical expert assesses the sketches to discover any deficiencies in the cognitive processes of the patient. More recently, automatic tools for assessing such tests have been developed. However, the problem of automatic interpretation of clock drawings, especially those sketched by people with cognitive impairment, is not fully solved, and in more difficult cases, the automatic systems have to revert to the help of human assessors in labelling the sketched objects forming the clock drawing. Moreover, the labelling of the sketched objects could be more reliable if prior knowledge of the expected CDT sketch structure and human reasoning could be integrated into the drawing interpretation system. This paper proposes a novel CDT sketch interpretation system, which represents the prior knowledge of the CDT structure by using ontology and integrating human reasoning through a fuzzy inference engine. The combination of the above technologies fuses multiple sources of information concerning the sketch structure and the visual appearance of the sketched objects whilst dealing with the interpretation uncertainty inherent to CDT sketches. The proposed CDT interpretation system is evaluated using two CDT data sets. The first data set consists of 65 drawings made by healthy people, while the second set contains 100 drawings reproduced from the drawings of dementia patients to simulate the kind of challenging sketches the system may have to work with. The evaluation analysis shows an improved interpretation performance of the proposed system in comparison with the classical approach, which does not receive benefits from the prior knowledge of the CDT sketch structure or simulated human reasoning.
Zainab Harbi, Yulia Hicks, Rossitza Setchi
KES2
2016 Huntington's Disease Assessment Using Tri Axis Accelerometers
abstract
Huntington's disease (HD) is a progressive inherited neurodegenerative disorder, causing involuntary movement and cognitive problems, severely affecting the quality of life. Controlling upper limb function is a core feature of daily activity and can prove problematic for people with HD. The Money Box Test (MBT) has been developed with a purpose of quantifying the involuntary movement frequently seen in people with HD. In this research, wearable and highly sensitive accelerometers are used to collect the acceleration of the hands and chest during the performance of the MBT. Using this data, a new approach is proposed to automatically classify the participants into two classes, healthy and HD, on the basis of the time series accelerometer data. A set of 90 time domain features is extracted from the accelerometer data, a feature selection technique is used to analyse the feature significance and to reduce the dimensionality of the dataset, and finally an SVM classifier is used to classify subjects into healthy and HD classes. The data of seven healthy controls and 15 HD patients are used in this study. The highest accuracy with the most significant eight features is 86.36% with the sensitivity and the specificity values being 87.50%, and 83.33% respectively.
Mohamed Bennasar, Yulia Hicks, Susanne Clinch, Philippa Jones, Anne Rosser, Monica Busse, Catherine A. Holt
KES2
2016 Automatic Detection and Quantification of Abdominal Aortic Calcification in Dual Energy X-ray Absorptiometry
abstract
Cardiovascular disease (CVD) is a major cause of mortality and the main cause of morbidity worldwide. CVD may lead to heart attacks and strokes and most of these are caused by atherosclerosis; this is a medical condition in which the arteries become narrowed and hardened due to an excessive build-up of plaque on the inner artery wall. Arterial calcification and, in particular, abdominal aortic calcification (AAC) is a manifestation of atherosclerosis and a prognostic indicator of CVD. In this paper, a two-stage automatic method to detect and quantify the severity of AAC is described; it is based on the analysis of lateral vertebral fracture assessment (VFA) images. These images were obtained on a dual energy x-ray absorptiometry (DXA) scanner used in single energy mode. First, an active appearance model was used to segment the lumbar vertebrae L1-L4 and the aorta on VFA images; the segmentation of the aorta was based on its position with respect to the vertebrae. In the second stage, feature vectors representing calcified regions in the aorta were extracted to quantify the severity of AAC. The presence and severity of AAC was also determined using an established visual scoring system (AC24). The abdominal aorta was divided into four parts immediately anterior to each vertebra, and the severity of calcification in the anterior and posterior walls was graded separately for each part on a 0-3 scale. The results were summed to give a composite severity score ranging from 0 to 24. This severity score was classified as follows: mild AAC (score 0-4), moderate AAC (score 5-12) and severe AAC (score 12-24). Two classification algorithms (k-nearest neighbour and support vector machine) were trained and tested to assign the automatically extracted feature vectors into the three classes. There was good agreement between the automatic and visual AC24 methods and the accuracy of the automated technique relative to visual classification indicated that it is capable of identifying and quantifying AAC over a range of severity.
Karima Elmasri, Yulia Hicks, Xin Yang 0033, Xianfang Sun, Rebecca Pettit, William Evans
KES2
2016 Clock Drawing Test Digit Recognition Using Static and Dynamic Features
abstract
The clock drawing test (CDT) is a standard neurological test for detection of cognitive impairment. A computerised version of the test promises to improve the accessibility of the test in addition to obtaining more detailed data about the subject's performance. Automatic handwriting recognition is one of the first stages in the analysis of the computerised test, which produces a set of recognized digits and symbols together with their positions on the clock face. Subsequently, these are used in the test scoring. This is a challenging problem because the average CDT taker has a high likelihood of cognitive impairment, and writing is one of the first functional activities to be affected. Current handwritten digit recognition system perform less well on this kind of data due to its unintelligibility. In this paper, a new system for numeral handwriting recognition in the CDT is proposed. The system is based on two complementary sources of data, namely static and dynamic features extracted from handwritten data. The main novelty of this paper is the new handwriting digit recognition system, which combines two classifiers—fuzzy k-nearest neighbour for dynamic stroke-based features and convolutional neural network for static image- based features, which can take advantage of both static and dynamic data. The proposed digit recognition system is tested on two sets of data: first, Pendigits online handwriting digits; and second, digits from the actual CDTs. The latter data set came from 65 drawings made by healthy people and 100 drawings reproduced from the drawings by dementia patients. The test on both data sets shows that the proposed combination system can outperform each classifier individually in terms of recognition accuracy, especially when assessing the handwriting of people with dementia.
Zainab Harbi, Yulia Hicks, Rossitza Setchi
KES2
2016 Dementia Detection Using Weighted Direction Index Histograms and SVM for Clock Drawing Test
abstract
Increasing the number of elderly persons who have dementia is one of the severe social problems. In Japan, the Ministry of Health, Labor and Welfare expects that the number of dementia patients will be around 5 million in 2025. It is also easily estimated that they require various living supports. Therefore, early detection and prevention of dementia are important. The authors have been developing a new system for quantitative and accurate evaluation of dementia. The basic concept of our system is evaluating a patient's dementia types and progression without awareness. To realize this, we are now developing the system using daily conversations, drawings, facial expressions and so on. In this paper, we focused on Clock Drawing Test (CDT) and proposed a dementia evaluation method for CDT. In the proposed method, Weighted Direction Index Histogram Method was used to extract features from given images, and Support Vector Machine (SVM) detected dementia cases from them. As a result of evaluation experiments, the proposed method could detect 97.1% of dementia cases correctly.
Tomoaki Shigemori, Hiroharu Kawanaka, Yulia Hicks, Rossitza Setchi, Haruhiko Takase, Shinji Tsuruoka
KES3
2015 Automatic Classification of Facial Morphology for Medical Applications
abstract
Facial morphology measurement and classification play important role in the face anthropometry of many medical applications. This usually involves the investigation of medical abnormalities where specific facial features are studied by taking a number of measurements of the facial area under investigation. The measurements are often obtained from the three-dimensional (3D) scans of the faces; however, the measurements are often made manually, which is tedious and time consuming process. Moreover, in gene related studies thousands of measurements may be necessary in order to find statistically significant relationships between facial features and genes. Normative studies, from which typical populous models can be built, also require many measurements. Thus an automatic method to extract morphological measurements and interpret them is desirable. In this article, an automatic method for classification of facial morphology on the basis of a number of geometric measurements obtained automatically from 3D facial scans is presented. Among different facial features the philtrum, which is the vertical groove extending from the nose to the upper lip and the lip area, plays an important role in defining the interaction between the genes and craniofacial anomalies such as, for example, cleft lip and palate. In this paper, geometric features are analysed for their suitability to classify philtrum into three classes previously proposed by medical experts. Moreover, further analysis is conducted to assess the best number of classes to model the underlying data distribution from the point of view of classification accuracy. The obtained classification results are compared with the ground truth manual labelling of 3D face meshes provided by a medical expert. The dataset used for this research is taken from ALSPAC dataset and consists of 1000 3D face meshes. The proposed method achieves classification accuracy of 97% for this data set using the Mean, Minimum and Maximum curvature features in combination.
Hawraa Abbas, Yulia Hicks, David Marshall 0001
KES2
2015 Cognitive Network Framework for Heterogeneous Wireless Networks
abstract
The Internet is used by more than two billion customers around the world and is expected to serve as a global platform for interconnecting cyber-physical objects that form the Internet of Things (IoT). Within the next decade, traffic demands are expected to increase a thousand-fold. This challenge can be addressed by introducing and expanding heterogeneous wireless technologies, which provide higher network capacity, wider coverage and higher quality of service (QoS). However, the heterogeneity and complexity of these networks are a major challenge for traditional control and management systems. Therefore, there is a need for self-manageable and self-configurable networks that support the data produced by the different IoT devices and provide opportunities for data analytics. In this work, a cognitive network framework is proposed, in which the network protocol stack is integrated with a semantic system. The proposed framework provides the bases for building smart networks that observe data from different layers in the network protocol stack and represents it in a hierarchical structure in a knowledge base. The framework employs an ontology that provides an abstraction model for the different heterogeneous wireless devices. The ontology determines the relationships between technology-dependent parameters in the network protocol stack and enables, through the use of inferences, the utilization of the observed data from the network. The use of a cognitive network framework with the network protocol stack allows adding ontologies to describe the data, a solution which could solve the problem of analysing, searching or visualising data.
Ahmed Al-Saadi, Rossitza Setchi, Yulia Hicks
KES3
2015 Segmentation of Clock Drawings Based on Spatial and Temporal Features
abstract
The Clock Drawing Test (CDT) is an inexpensive and effective measure for early detection of cognitive impairment in the elderly, which is important for timely diagnosis and initiation of appropriate treatment. Currently, medical experts assess the drawings based on their judgement and a number of available scoring systems. An automatic system for assessment of CDT drawings would simultaneously decrease the waiting time for a specialist appointment and improve accessibility of the test to the patients. Published research has only started to address the problem of automatic assessment of CDT drawings and existing systems require user intervention during the segmentation of the CDT drawing into its composing parts, such as numbers and clock hands. In this paper, a new set of temporal and spatial features automatically extracted from the CDT data acquired using a graphics tablet is proposed. Consequently, a Support Vector Machine (SVM) classifier is employed to segment the CDT drawings into their elements, such as numbers and clock hands, on the basis of the extracted features. The proposed algorithm is tested on two data sets, the first set consisting of 65 drawings made by healthy people, and the second consisting of 100 drawings reproduced from actual drawings of dementia patients. The test on both data sets shows that the proposed method outperforms the current state-of-the-art method for CDT drawing segmentation.
Zainab Harbi, Yulia Hicks, Rossitza Setchi, Antony Bayer
KES2
2015 Ontology-based Framework for Risk Assessment in Road Scenes Using Videos
abstract
Recent advances in autonomous vehicle technology pose an important problem of automatic risk assessment in road scenes. This article addresses the problem by proposing a novel ontology tool for assessment of risk in unpredictable road traffic environment, as it does not assume that the road users always obey the traffic rules. A framework for video-based assessment of the risk in a road scene encompassing the above ontology is also presented in the paper. The framework uses as input the video from a monocular video camera only, avoiding the need for additional sometimes expensive sensors. The key entities in the road scene (vehicles, pedestrians, environment objects etc.) are organised into an ontology which encodes their hierarchy, relations and interactions. The ontology tool infers the degree of risk in a given scene using as knowledge video-based features, related to the key entities. The evaluation of the proposed framework focuses on scenarios in which risk results from pedestrian behaviour. A dataset consisting of real-world videos illustrating pedestrian movement is built. Features related to the key entities in the road scene are extracted and fed to the ontology, which evaluates the degree of risk in the scene. The experimental results indicate that the proposed framework is capable of assessing risk resulting from pedestrian behaviour in various road scenes accurately.
Mahmud Abdulla Mohammad, Ioannis Kaloskampis, Yulia Hicks, Rossitza Setchi
KES3
2015 Feature Extraction Method for Clock Drawing Test
abstract
Recently, the number of elderly persons with dementia has been increasing. In the past, we proposed a dementia evaluation system using daily conversations and developed the system with a conversational robot. However, the current system is not ready for practical use because it can only evaluate time/geographical orientation and short-term memory, and some methods to evaluate other orientations and functions is required as well. In this paper, we discuss a new dementia evaluation system using not only daily conversations but also drawing tests. The authors employed a Clock Drawing Test (CDT) as a new dementia evaluation test and implemented it in a tablet device. This paper discusses a feature extraction and recognition method to distinguish normal cases from dementia cases. After evaluation experiments, the proposed method could recognize 87.6% of the clock drawing images.
Tomoaki Shigemori, Zainab Harbi, Hiroharu Kawanaka, Yulia Hicks, Rossitza Setchi, Haruhiko Takase, Shinji Tsuruoka
KES4
2015 Feature selection using Joint Mutual Information Maximisation
abstract
Feature selection is used in many application areas relevant to expert and intelligent systems, such as data mining and machine learning, image processing, anomaly detection, bioinformatics and natural language processing. Feature selection based on information theory is a popular approach due its computational efficiency, scalability in terms of the dataset dimensionality, and independence from the classifier. Common drawbacks of this approach are the lack of information about the interaction between the features and the classifier, and the selection of redundant and irrelevant features. The latter is due to the limitations of the employed goal functions leading to overestimation of the feature significance. To address this problem, this article introduces two new nonlinear feature selection methods, namely Joint Mutual Information Maximisation (JMIM) and Normalised Joint Mutual Information Maximisation (NJMIM); both these methods use mutual information and the ‘maximum of the minimum’ criterion, which alleviates the problem of overestimation of the feature significance as demonstrated both theoretically and experimentally. The proposed methods are compared using eleven publically available datasets with five competing methods. The results demonstrate that the JMIM method outperforms the other methods on most tested public datasets, reducing the relative average classification error by almost 6% in comparison to the next best performing method. The statistical significance of the results is confirmed by the ANOVA test. Moreover, this method produces the best trade-off between accuracy and stability.
Mohamed Bennasar, Yulia Hicks, Rossitza Setchi
Expert Syst. Appl.2
2015 Conceptual Framework for Evaluating Intuitive Interaction Based on Image Schemas
abstract
Intuitive interaction is an important aspect of usability in interface design. This paper contributes to the research in this area by proposing a conceptual framework for evaluating intuitive interaction based on image schemas. The framework comprises four phases: goal identification, image schemas extraction, analysis and assessment. It quantifies intuitive interaction by comparing the image schemas envisaged by the designer of a product with those used by its users. The proposed framework is evaluated through a study involving 42 participants completing a set task with a product. The study identified the image schemas, which were correctly used in accordance with the designer's intent and those that were incorrectly used and contributed to the difficulties that many participants experienced. The inter-rater reliability and empirical validity were examined. The proposed framework provides a structured approach to usability testing by enabling both quantitative and qualitative evaluation of intuitive interaction.
Obokhai Kess Asikhia, Rossitza Setchi, Yulia Hicks, Andrew Walters
Interact. Comput.3
2014 Multi-rate medium access protocol based on reinforcement learning
abstract
Many wireless devices employ multi-rate techniques to improve network performance. However, despite the significant amount of research aimed at dynamically adjusting the transmission rate, the majority of this effort considers neither the competing nodes in wireless mesh networks nor the congestion in the nodes. This work employs distributed intelligent agents to observe the surrounding environment in order to dynamically adjust the individual node transmission rates. Reinforcement learning is employed to control the way each node updates its transmission rate based on the transmission rate of the adjacent node as well as the traffic load. This work is validated through extensive simulations that compare the proposed model with three of the most widely cited schemes. The results indicate significant improvement in system throughput.
Ahmed Al-Saadi, Rossitza Setchi, Yulia Hicks, Stuart M. Allen
SMC3
2014 Cascade classification for diagnosing dementia
abstract
Dementia is a syndrome caused by a chronic or progressive disease of the brain, which affects memory, orientation, thinking, calculation, learning ability and language. The Clock Drawing Test (CDT) and Mini Mental State Examination (MMSE) are well-known cognitive assessment tests. A known obstacle to the wider usage of the CDT assessments is the scoring and interpretation of the results. This paper introduces a novel cascade CDT classifier, which can help in the diagnosis of three stages of dementia. The data used in this research are 604 clock drawings produced by patients and healthy individuals. The study employs 47 visual features, which are selected following a comprehensive analysis of the available data and the most common CDT scoring systems reported in the medical literature. These features are used to build a new digitized dataset needed to train and validate the proposed classifier. The results show significant improvement of 6.8% in differentiating between three levels of dementia (normal/functional, mild cognitive impairment/mild dementia, and moderate/severe dementia) when compared to a single stage classifier. In particular, the results show classification accuracy of over 89% when discriminating between normal and abnormal conditions only.
Mohamed Bennasar, Rossitza Setchi, Yulia Hicks, Antony Bayer
SMC3
2014 Joint EEG-fMRI model for EEG source separation
abstract
Electroencephalography (EEG) offers a rich representation of human brain activity in the time domain. EEG would in many circumstances be the preferred technique for analysing brain activity, as it is less expensive and more practical to use than other modalities like functional Magnetic Resonance Imaging (fMRI), notably due to its size. However, its spatial resolution is limited, hampering its ability to characterise activity across spatially distributed brain networks. In comparison, functional Magnetic Resonance Imaging (fMRI) offers very good spatial resolution but the haemodynamic nature of the signal limits its temporal resolution to the order of seconds. A possible solution to this problem is to use both EEG and fMRI signals, but this approach would lead to the loss of convenience of EEG alone. We would like to bring in the advantages of fMRI signal into EEG assessment of brain state and brain responses without the necessity for the presence of the fMRI equipment on site. In this article, we propose a joint statistical model of fMRI/EEG signals and then exploit the learnt correlations to improve the results of signal processing of EEG on its own. We compare the performance of a Blind Source Separation (BSS) method on its own with one, which uses our joint EEG-fMRI model, and show the improvement in the precision of the source separation.
Yulia Hicks, Rossitza Setchi
SMC2
2013 Feature Selection based on Information Theory in the Clock Drawing Test
abstract
The Clock Drawing Test is one of the most widely used screening tools for cognitive impairment and dementia. Since its introduction, more than fifteen scoring systems have been developed to assess the clock drawings. However, very little research has been conducted to study the significance of the elements (features) of the clock drawings for the correct diagnosis of dementia. This paper employs a feature selection method called Feature Interaction Maximization (FIM) to identify the most significant visual features of the test, which can be associated with dementia. The proposed approach is tested with a dataset of 648 clock drawings produced by dementia patients and healthy individuals. The results are compared with other methods used by medical experts. Furthermore, the paper compares the FIM method with an alternative feature selection method based on Information Gain. The results show that the FIM method selects features with higher discriminative power which leads to a deeper understanding of the Clock Drawing Test.
Mohamed Bennasar, Rossitza Setchi, Antony Bayer, Yulia Hicks
KES4
2013 Feature Interaction Maximisation
Mohamed Bennasar, Rossitza Setchi, Yulia Hicks
Pattern Recognit. Lett.3
2012 Hybrid phoneme based clustering approach for audio driven facial animation
abstract
We consider the problem of producing accurate facial animation corresponding to a given input speech signal. A popular technique previously used for Audio Driven Facial Animation is to build a joint audio-visual model using Active Appearance Models (AAMs) to represent possible facial variations and Hidden Markov Models (HMMs) to select the correct appearance based on the input audio. However there are several questions that remained unanswered. In particular the choice of clustering technique and the choice of the number of clusters in the HMM may have significant influence over the quality of the produced videos. We have investigated a range of clustering techniques in order to improve the quality of the HMM produced, and proposed a new structure based on using Gaussian Mixture Models (GMMs) to model each phoneme separately. We compared our approach to several alternatives using a public dataset of 300 phonetically labeled sentences spoken by a single person and found that our approach produces more accurate animation. In addition, we use a hybrid approach where the training data is phonetically labeled thus producing a model with better separation of phonemes, but test audio data is not labeled, thus making our approach for generating facial animation less laborious and fully automatic.
Benjamin Havell, Paul L. Rosin, Saeid Sanei, Andrew J. Aubrey, David Marshall 0001, Yulia Hicks
ICASSP6
2012 Unsupervised Discretization Method based on Adjustable Intervals
abstract
Discretization is a process applied to transform continuous data into data with discrete attributes. It makes the learning step of many classification algorithms more accurate and faster. Although many efficient supervised discretization methods have been proposed, unsupervised methods such as Equal Width Discretization (EWD) and Equal Frequency Discretization (EFD) are still in use especially with datasets when classification is not available. Each of these algorithms has its drawbacks. To improve the classification accuracy of EWD, a new method based on adjustable intervals is proposed in this paper. The new method is tested using benchmarking datasets from the UCI repository of machine learning databases; the C4.5 classification algorithm is then used to test the classification accuracy. The experimental results show that the method improves the classification accuracy by about 5% compared to the conventional EWD and EFD methods, and is as good as the supervised Entropy Minimization Discretization (EMD) method.
Mohamed Bennasar, Rossitza Setchi, Yulia Hicks
KES3
2011 Reinforcing conceptual engineering design with a hybrid computer vision, machine learning and knowledge based system framework
abstract
We propose a novel system that aids engineers in the conceptual stage of design. Our system's goal is to support the engineer without limiting his creative role; thus, our proposed method does not produce ready study solutions but rather actively monitors the design procedure, verifying design stages and pointing out potential mistakes. This is achieved with a hybrid computer vision, machine learning and knowledge based system framework. Design stage identification is performed with a novel algorithm which comprises a classification stage based on Random Forests and examination of the temporal relationships between the engineer's actions with the aid of statistical graphical models. Experimental results captured in a complex, real life scenario demonstrate our system's ability to efficiently support the engineer's decisions during the conceptual stage of design.
Ioannis Kaloskampis, Yulia Hicks, David Marshall 0001
SMC2
2010 An evolving MoG for online image sequence segmentation
abstract
When segmenting image sequences, it is important to ensure the coherency of the produced segments across successive frames. In this paper, we present a method for evolving a Mixture of Gaussian (MoG) to produce such coherent segments. Using a MoG allows us to select the number of components automatically and in a principled way. The parameters of the evolving MoG can vary smoothly to track online the continuous evolution of the feature's distribution. In addition, the complexity of the MoG can vary to cope with incoming or disappearing objects in the sequence. The method is tested on several video sequences and the results are compared to another method, which shows the advantage of the ability to change the number of components automatically for tracking changes in the scene.
Cyril Charron, Yulia Hicks
ICIP2
2009 Incremental Learning of Dynamical Models of Faces
abstract
Active Appearance Models (AAM) are a useful and popular tool for modelling facial variations. They have been used in face tracking, recognition and synthesis applications. For modelling facial dynamics of speech, they have been used in conjunction with Hid-den Markov Models (HMM). However, the high dimensionality of the training data and of the resulting AAMs leads to long learning time of HMMs and thus imposes serious limitations on their joint use. Here, we propose a new method for learning HMMs of facial dynamics incremen-tally. Our algorithm is fully unsupervised and can be used for on-line learning as new data becomes available. Another important feature of our algorithm is the automatic choice of the number of states in the model. We show in experiments an improvement in learning speed of three orders of magnitude. Finally, we demonstrate the quality of the learned HMMs by generating video footage of a talking face. 1 Introduction and
Cyril Charron, Yulia Hicks, Peter Hall 0001, Darren Cosker
BMVC2
2007 A Geometrically Constrained Multimodal Approach for Convolutive Blind Source Separation
abstract
A novel constrained multimodal approach for convolutive blind source separation is presented which incorporates video information related to geometrical position of both the speakers and the microphones, and the directionality of the speakers into the separation algorithm. The separation is performed in the frequency domain and the constraints are incorporated through a penalty function-based formulation. The separation results show a considerable improvement over traditional frequency domain convolutive BSS systems such as that developed by Parra and Spence. Importantly, the inherent permutation problem in the frequency domain BSS is potentially solved.
Saeid Sanei, Syed M. Naqvi, Jonathon A. Chambers, Yulia Hicks
ICASSP (3)4
2006 A model of diatom shape and texture for analysis, synthesis and identification
Yulia Hicks, David Marshall 0001, Paul L. Rosin, Ralph R. Martin, David G. Mann, S. J. M. Droop
Mach. Vis. Appl.1
2005 Video assisted speech source separation
abstract
We investigate the problem of integrating the complementary audio and visual modalities for speech separation. Rather than using independence criteria suggested in most blind source separation (BSS) systems, we use visual features from a video signal as additional information to optimize the unmixing matrix. We achieve this by using a statistical model characterizing the nonlinear coherence between audio and visual features as a separation criterion for both instantaneous and convolutive mixtures. We acquire the model by applying the Bayesian framework to the fused feature observations based on a training corpus. We point out several key existing challenges to the success of the system. Experimental results verify the proposed approach, which outperforms the audio only separation system in a noisy environment, and also provides a solution to the permutation problem.
Wenwu Wang 0001, Darren Cosker, Yulia Hicks, Saeid Sanei, Jonathon A. Chambers
ICASSP (5)3
2003 A method to add Hidden Markov Models with application to learning articulated motion
abstract
In this paper we present a method for adding Hidden Markov Models. The main advantages of our method are that it does not require the data the models had been trained on, allows a change in the number of components, does not assume independence of the components to be added and is resistant to the order in which the training data arrives. We assessed the method in the experiments with synthetic data, which showed good accuracy. Finally, we present an application in computer vision. 1
Yulia Hicks, Peter Hall 0001, Dave Marshall
BMVC1
2002 Modelling life cycle related and individual shape variation in biological specimens
abstract
The main purpose of this research is to develop methods for automatic identification of biological specimens in digital photographs and drawings held in a database. Incorporation of taxonomic drawings into a visual indexing system has not been attempted to date. Diatoms are a single cell microscopic algae that provide a particularly suitable case study. Identification of diatoms is a challenging task due to the huge number of the species, blurred boundaries between species, and life cycle related shape changes. A novel model based on principal curves representing the life cycle related shape variation of a number of diatom species has been developed. Our model is suitable for reconstruction purposes, allowing us to produce drawings of a variety of diatom shapes, thus providing a link between the photographs and drawings. We present the classification results of photographed and drawn specimens based on the model and compare our results to another recent system for diatom identification. Finally, given a diatom specimen, we are able not only to identify the species it belongs to but also to pinpoint the stage in the life cycle it represents.
Yulia Hicks, David Marshall 0001, Ralph R. Martin, Paul L. Rosin, Micha Bayer, David G. Mann
BMVC1
2002 Automatic landmarking for building biological shape models
abstract
We present a new method for automatic landmark extraction from the contours of biological specimens. Our ultimate goal is to enable automatic identification of biological specimens in photographs and drawings held in a database. We propose to use active appearance models for visual indexing of both photographs and drawings. Automatic landmark extraction will assist us in building the models. We describe the results of using our method on drawings and photographs of examples of diatoms, and present an active shape model built using automatically extracted data.
Yulia Hicks, David Marshall 0001, Ralph R. Martin, Paul L. Rosin, Micha Bayer, David G. Mann
ICIP (2)1
2002 Tracking people in three dimensions using a hierarchical model of dynamics
Yulia Hicks, Peter Hall 0001, David Marshall 0001
Image Vis. Comput.1
2000 A Hierarchical Model of Dynamics for Tracking People with a Single Video Camera
abstract
We propose a novel hierarchical model of human dynamics for view independent tracking of the human body in monocular video sequences. The model is trained using real data from a collection of people. Kinematics are encoded using Hierarchical Principal Component Analysis, and dynamics are encoded using Hidden Markov Models. The top of the hierarchy contains information about the whole body. The lower levels of the hierarchy contain more detailed information about possible poses of some subpart of the body. When tracking, the lower levels of the hierarchy are shown to improve accuracy. In this article we describe our model and present experiments that show we can recover 3D skeletons from 2D images in a view independent manner, and also track people the system was not trained on.
Yulia Hicks, Peter Hall 0001, David Marshall 0001
BMVC1