David Marshall 0001

dblp:m/ADavidMarshall · also A. David Marshall · DBLP profile ↗
← Back
71ranked-venue papers
4as first author
12since 2021 · last 2024
0000-0003-2789-1395ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 39 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2024 Decision fusion based multimodal hierarchical method for speech emotion recognition from audio and text
abstract
Expressing emotions is essential in human interaction.Often, individuals convey emotions through neutral speech, while the underlying meaning carries emotional weight.Conversely, tone can also convey emotion despite neutral words.Most Speech Emotion Recognition research overlooks this.We address this gap with a multimodal emotion recognition system using hierarchical classifiers and a novel decision fusion method.Our approach analyses emotional cues from speech and text, measuring their impact on predicted classes, considering emotional or neutral contributions for each instance.Results on the IEMOCAP dataset show our method's effectiveness: 69.45% and 65.26% weighted accuracy in speaker-dependent and speaker-independent settings, respectively.
Nawal Alqurashi, Yuhua Li 0001, Kirill A. Sidorov, David Marshall 0001
ESANN4
2024 Fast Explanation of RBF-Kernel SVM Models Using Activation Patterns
abstract
Machine learning models have significantly enriched the toolbox in the field of neuroimaging analysis. Among them, Support Vector Machines (SVM) have been one of the most popular models for supervised learning, but their use primarily relies on linear SVM models due to their explainability. Kernel SVM models are capable classifiers but more opaque. Recent advances in eXplainable AI (XAI) have developed several feature importance methods to address the explainability problem. However, noise variables can affect these explanations, making irrelevant variables regarded as important variables. This problem also appears in explaining linear models, which the linear pattern can address. This paper proposes a fast method to explain RBF kernel SVM globally by adopting the notion of linear pattern in kernel space. Our method can generate global explanations with low computational cost and is less affected by noise variables. We successfully evaluate our method on simulated and real MEG/EEG datasets.
Matthias Treder, David Marshall 0001, Yuhua Li 0001
IJCNN3
2024 Improving High-Frequency Details in Cerebellum for Brain MRI Super-Resolution
abstract
Deep-learning-based single-image super-resolution models are typically trained using image patches, rather than the whole images, due to hardware limits. Since different brain regions have disparate structures and their size varies, such as the cerebrum and cerebellum, models trained using image patches can be dominated by the structures of the larger brain regions and ignore the fine-grained details in smaller areas. In this paper, we first evaluate several previously proposed models using more blurry low-resolution images than previous studies, as input. Then, we propose an effective approach for the conventional patch-based strategy by balancing the proportion of patches containing high-frequency details. This makes the model focus more on high-frequency information in tiny regions, especially for the cerebellum. Compared with the conventional patch-based strategy, the resultant super-resolved image from our approach achieves comparable image quality in the whole brain. In contrast, it improves significantly on the high-frequency details in the cerebellum.
Hanzhi Wang 0002, David Marshall 0001, Derek K. Jones, Yuhua Li 0001
ISCC2
2024 Explaining the predictions of kernel SVM models for neuroimaging data analysis
abstract
Machine learning methods have shown great performance in many areas, including neuroimaging data analysis. However, model performance is only one objective in neuroimaging analysis. Gaining insight from the data is also critical in this field, such as identifying regions where detected signals are relevant to cognitive and diagnostic tasks. To fulfil this need, enabling the explainability of a model’s decision-making process is critical. Predictions of complex machine learning models are notoriously difficult to explain. This limits the use of complex models like kernel support vector machines (SVM) in neuroimaging analysis. Recently, several permutation-based methods have been developed to explain these complex models. However, the explanation results are affected by class-irrelevant features like suppressor variables and high background noise variables. This problem may also happen when explaining linear models. One possible reason is that the permutation process will produce unrealistic data instances when features are not independent, e.g. correlated. These unrealistic data instances will influence the explanation results. In neuroimaging analysis, the activation pattern, the estimated weight of the assumed generative model corresponding to the current classifier, is used to deal with this problem for linear models. This method does not rely on a permutation process but rather on the available data information. In this paper, we propose a novel method of Explanation through Activation Pattern (EAP) to explain the SVM models with different types of kernels for neuroimaging data analysis. Our method can generate a global feature importance score by estimating the activation pattern of kernel SVM models. We evaluate our method against three popular methods on both simulation datasets and a publicly available EEG/MEG dataset on visual tasks. The experimental results demonstrate that the proposed EAP method can provide explanations with low computational cost and is less affected by class-irrelevant features than the other three methods. In the experiment using the MEG/EEG dataset of visual tasks, the proposed EAP method provides agreement results with the brain’s electrical activity patterns reported in the literature on the visual tasks EEG/MEG data and is significantly faster than the other explanation methods.
Matthias Treder, David Marshall 0001, Yuhua Li 0001
Expert Syst. Appl.3
2024 Predicting Radiologists' Gaze With Computational Saliency Models in Mammogram Reading
abstract
Previous studies have shown that there is a strong correlation between radiologists' diagnoses and their gaze when reading medical images. The extent to which gaze is attracted by content in a visual scene can be characterised as visual saliency. There is a potential for the use of visual saliency in computer-aided diagnosis in radiology. However, little is known about what methods are effective for diagnostic images, and how these methods could be adapted to address specific applications in diagnostic imaging. In this study, we investigate 20 state-of-the-art saliency models including 10 traditional models and 10 deep learning-based models in predicting radiologists' visual attention while reading 196 mammograms. We found that deep learning-based models represent the most effective type of methods for predicting radiologists' gaze in mammogram reading; and that the performance of these saliency models can be significantly improved by transfer learning. In particular, an enhanced model can be achieved by pre-training the model on a large-scale natural image saliency dataset and then fine-tuning it on the target medical image dataset. In addition, based on a systematic selection of backbone networks and network architectures, we proposed a parallel multi-stream encoded model which outperforms the state-of-the-art approaches for predicting saliency of mammograms.
Jianxun Lou, Hanhe Lin, Philippa Young, Zelei Yang, Susan Cheng Shelmerdine, David Marshall 0001, Emiliano Spezi, Marco Palombo, Hantao Liu
IEEE Trans. Multim.7
2023 A Skewed Loss Function for Correcting Predictive Bias in Brain Age Prediction
abstract
In neuroimaging, the difference between predicted brain age and chronological age, known as brain age delta, has shown its potential as a biomarker related to various pathological phenotypes. There is a frequently observed bias when estimating brain age delta using regression models. This bias manifests as an overestimation of brain age for young participants and an underestimation of brain age for older participants. Therefore, the brain age delta is negatively correlated with chronological age, which can be problematic when evaluating relationships between brain age delta and other age-associated variables. This paper proposes a novel bias correction method for regression models by introducing a skewed loss function to replace the normal symmetric loss function. The regression model then behaves differently depending on whether it makes overestimations or underestimations. Our approach works with any type of MR image and no specific preprocessing is required, as long as the image is sensitive to age-related changes. The proposed approach has been validated using three classic deep learning models, namely ResNet, VGG, and GoogleNet on publicly available neuroimaging aging datasets. It shows flexibility across different model architectures and different choices of hyperparameters. The corrected brain age delta from our approach then has no linear relationship with chronological age and achieves higher predictive accuracy than a commonly-used two-stage approach.
Hanzhi Wang 0002, Matthias Treder, David Marshall 0001, Derek K. Jones, Yuhua Li 0001
IEEE Trans. Medical Imaging3
2022 Predicting Radiologist Attention During Mammogram Reading with Deep and Shallow High-Resolution Encoding
abstract
Radiologists’ eye-movement during diagnostic image reading reflects their personal training and experience, which means that their diagnostic decisions are related to their perceptual processes. For training, monitoring, and performance evaluation of radiologists, it would be beneficial to be able to automatically predict the spatial distribution of the radiologist’s visual attention on the diagnostic images. The measurement of visual saliency is a well-studied area that allows for prediction of a person’s gaze attention. However, compared with the extensively studied natural image visual saliency (in free viewing tasks), the saliency for diagnostic images is less studied; there could be fundamental differences in eye-movement behaviours between these two domains. Most current saliency prediction models have been optimally developed for natural images, which could lead them to be less adept at predicting the visual attention of radiologists during the diagnosis. In this paper, we propose a method specifically for automatically capturing the visual attention of radiologists during mammogram reading. By adopting high-resolution image representations from both deep and shallow encoders, the proposed method avoids potential detail losses and achieves superior results on multiple evaluation metrics in a large mammogram eye-movement dataset.
Jianxun Lou, Hanhe Lin, David Marshall 0001, Young Yang, Susan Cheng Shelmerdine, Hantao Liu
ICIP3
2022 SWAG-V: Explanations for Video using Superpixels Weighted by Average Gradients
abstract
CNN architectures that take videos as an input are often overlooked when it comes to the development of explanation techniques. This is despite their use in often critical domains such as surveillance and healthcare. Explanation techniques developed for these networks must take into account the additional temporal domain if they are to be successful. In this paper we introduce SWAG-V, an extension of SWAG for use with networks that take video as an input. In addition we show how these explanations can be created in such a way that they are balanced between fine and coarse explanations. By creating superpixels that incorporate the frames of the input video we are able to create explanations that better locate regions of the input that are important to the networks prediction. We compare SWAG-V against a number of similar techniques using metrics such as insertion and deletion, and weak localisation. We compute these using Kinetics-400 with both the C3D and R(2+1)D network architectures and find that SWAG-V is able to outperform multiple techniques.
Thomas Hartley, Kirill A. Sidorov, Chris Willis 0001, David Marshall 0001
WACV4
2022 TranSalNet: Towards perceptually relevant visual saliency prediction
abstract
Convolutional neural networks (CNNs) have significantly advanced computational modelling for saliency prediction. However, accurately simulating the mechanisms of visual attention in the human cortex remains an academic challenge. It is critical to integrate properties of human vision into the design of CNN architectures, leading to perceptually more relevant saliency prediction. Due to the inherent inductive biases of CNN architectures, there is a lack of sufficient long-range contextual encoding capacity. This hinders CNN-based saliency models from capturing properties that emulate viewing behaviour of humans. Transformers have shown great potential in encoding long-range information by leveraging the self-attention mechanism. In this paper, we propose a novel saliency model that integrates transformer components to CNNs to capture the long-range contextual visual information. Experimental results show that the transformers provide added value to saliency prediction, enhancing its perceptual relevance in the performance. Our proposed saliency model using transformers has achieved superior results on public benchmarks and competitions for saliency prediction models. The source code of our proposed saliency model TranSalNet is available at: https://github.com/LJOVO/TranSalNet.
Jianxun Lou, Hanhe Lin, David Marshall 0001, Dietmar Saupe, Hantao Liu
Neurocomputing3
2021 Learning Precise Temporal Point Event Detection with Misaligned Labels
Julien Schroeter, Kirill A. Sidorov, David Marshall 0001
AAAI3
2021 Jitter-CAM: Improving the Spatial Resolution of CAM-Based Explanations
Thomas Hartley, Kirill A. Sidorov, Chris Willis 0001, David Marshall 0001
BMVC4
2021 SWAG: Superpixels Weighted by Average Gradients for Explanations of CNNs
abstract
Providing an explanation of the operation of CNNs that is both accurate and interpretable is becoming essential in fields like medical image analysis, surveillance, and autonomous driving. In these areas, it is important to have confidence that the CNN is working as expected and explanations from saliency maps provide an efficient way of doing this. In this paper, we propose a pair of complementary contributions that improve upon the state of the art for region-based explanations in both accuracy and utility. The first is SWAG, a method for generating accurate explanations quickly using superpixels for discriminative regions which is meant to be a more accurate, efficient, and tunable drop in replacement method for Grad-CAM, LIME, or other region-based methods. The second contribution is based on an investigation into how to best generate the superpixels used to represent the features found within the image. Using SWAG, we compare using superpixels created from the image, a combination of the image and backpropagated gradients, and the gradients themselves. To the best of our knowledge, this is the first method proposed to generate explanations using superpixels explicitly created to represent the discriminative features important to the network. To compare we use both ImageNet and challenging fine-grained datasets over a range of metrics. We demonstrate experimentally that our methods provide the best local and global accuracy compared to Grad-CAM, Grad-CAM++, LIME, XRAI, and RISE.
Thomas Hartley, Kirill A. Sidorov, Chris Willis 0001, David Marshall 0001
WACV4
2020 Learning Multi-instance Sub-pixel Point Localization
Julien Schroeter, Tinne Tuytelaars, Kirill A. Sidorov, David Marshall 0001
ACCV (5)4
2019 Weakly-Supervised Temporal Localization via Occurrence Count Learning
abstract
We propose a novel model for temporal detection and localization which allows the training of deep neural networks using only counts of event occurrences as training labels. This powerful weakly-supervised framework alleviates the burden of the imprecise and time consuming process of annotating event locations in temporal data. Unlike existing methods, in which localization is explicitly achieved by design, our model learns localization implicitly as a byproduct of learning to count instances. This unique feature is a direct consequence of the model’s theoretical properties. We validate the effectiveness of our approach in a number of experiments (drum hit and piano onset detection in audio, digit detection in images) and demonstrate performance comparable to that of fully-supervised state-of-the-art methods, despite much weaker training requirements.
Julien Schroeter, Kirill A. Sidorov, David Marshall 0001
ICML3
2018 Computational Paralinguistics: Automatic Assessment of Emotions, Mood and Behavioural State from Acoustics of Speech
Zafi Sherhan Syed, Julien Schroeter, Kirill A. Sidorov, David Marshall 0001
INTERSPEECH4
2018 A 3D morphometric perspective for facial gender analysis and classification using geodesic path curvature features
abstract
The relationship between the shape and gender of a face, with particular application to automatic gender classification, has been the subject of significant research in recent years. Determining the gender of a face, especially when dealing with unseen examples, presents a major challenge. This is especially true for certain age groups, such as teenagers, due to their rapid development at this phase of life. This study proposes a new set of facial morphological descriptors, based on 3D geodesic path curvatures, and uses them for gender analysis. Their goal is to discern key facial areas related to gender, specifically suited to the task of gender classification. These new curvature-based features are extracted along the geodesic path between two biological landmarks located in key facial areas. Classification performance based on the new features is compared with that achieved using the Euclidean and geodesic distance measures traditionally used in gender analysis and classification. Five different experiments were conducted on a large teenage face database (4745 faces from the Avon Longitudinal Study of Parents and Children) to investigate and justify the use of the proposed curvature features. Our experiments show that the combination of the new features with geodesic distances provides a classification accuracy of 89%. They also show that nose-related traits provide the most discriminative facial feature for gender classification, with the most discriminative features lying along the 3D face profile curve.
Hawraa Abbas, Yulia Hicks, David Marshall 0001, Alexei I. Zhurov, Stephen Richmond
Comput. Vis. Media3
2017 An open-data, agent-based model of alcohol related crime
abstract
The allocation of resources to challenge city centre violent crime traditionally relies on historical data to identify hot-spots. The usefulness of such data-driven approaches is limited when historical data is scarce or unavailable (e.g. planning of a new city) or insufficiently representative (e.g. does not account for novel events, such as Olympic Games). In some cities, crime data is not systematically accumulated at all. We present a graph-constrained agent based simulation model of alcohol-related violent crime that is capable of predicting areas of likely violent crime without requiring any historical data. The only inputs to our simulation are publicly available geographical data, which makes our method immediately applicable to a wide range of tasks, such as optimal city planning, police patrol optimisation, devising alcohol licensing policies. In experiments, we evaluate our model and demonstrate agreement of our model's predictions on where and when violence will occur with real-world violent crime data. Analyses indicate that our agent based model may be able to make a significant contribution to attempts to prevent violence through deterrence or by design.
Joseph Redfern, Kirill A. Sidorov, Paul L. Rosin, Simon C. Moore, Padraig Corcoran, David Marshall 0001
AVSS6
2017 4D Analysis of Facial Ageing Using Dynamic Features
abstract
Facial ageing analysis based on 4D data (3D plus time) is much more robust to pose changes and illumination variations than using 2D image and video. The purpose of this investigation was to measure the effects of age and gender related facial changes using dynamic 3D facial scans. Experiments were carried out on the subjects, who were divided into two groups by age (15-30 years and 31-60 years). Each group was further subdivided by gender. 3D scans of the subjects were processed to extract facial features which were tracked through the duration of the data capture. Subsequently, a set of dynamic features were computed from these facial features, as well as static features for comparison. Two-way multivariate analysis of variance (MANOVA) of these features demonstrated that statistically significant age and gender related differences could be detected. We show that 3D facial dynamics provide more useful information than static features for the characterisation of smiles.
Khtam Al-Meyah, David Marshall 0001, Paul L. Rosin
KES2
2017 Detecting violent and abnormal crowd activity using temporal analysis of grey level co-occurrence matrix (GLCM)-based texture measures
abstract
The severity of sustained injury resulting from assault-related violence can be minimised by reducing detection time. However, it has been shown that human operators perform poorly at detecting events found in video footage when presented with simultaneous feeds. We utilise computer vision techniques to develop an automated method of abnormal crowd detection that can aid a human operator in the detection of violent behaviour. We observed that behaviour in city centre environments often occurs in crowded areas, resulting in individual actions being occluded by other crowd members. We propose a real-time descriptor that models crowd dynamics by encoding changes in crowd texture using temporal summaries of grey level co-occurrence matrix features. We introduce a measure of inter-frame uniformity and demonstrate that the appearance of violent behaviour changes in a less uniform manner when compared to other types of crowd behaviour. Our proposed method is computationally cheap and offers real-time description. Evaluating our method using a privately held CCTV dataset and the publicly available Violent Flows, UCF Web Abnormality and UMN Abnormal Crowd datasets, we report a receiver operating characteristic score of 0.9782, 0.9403, 0.8218 and 0.9956, respectively.
Kaelon Lloyd, Paul L. Rosin, David Marshall 0001, Simon C. Moore
Mach. Vis. Appl.3
2017 High-speed video generation with an event camera
David Marshall 0001, Luping Shi, Shi-Min Hu 0001
Vis. Comput.3
2015 Towards 4D Coupled Models of Conversational Facial Expression Interactions
abstract
In this paper we introduce a novel approach for building 4D coupled statistical models of conversational facial expression interactions. To build these coupled models we use 3D AAMs for feature extraction, 4D polynomial fitting for sequence representation, and concatenated feature vectors of frontchannel-backchannel interactions (with offset values) for the coupled model. Using a coupled model of conversation smile interactions, we predicted each sequence’s backchannel signal. In a subsequent experiment, human observers rated predicted sequences as highly similar to the originals. Our results demonstrate the usefulness of coupled models as powerful tools to analyse and synthesise key aspects of conversational interactions, including conversation timings, backchannel responses to frontchannel signals, and the spatial and temporal dynamics of conversational facial expression interactions.
Jason Vandeventer, Lukas Gräser, Magdalena Rychlowska, Paul L. Rosin, David Marshall 0001
BMVC5
2015 Automatic Classification of Facial Morphology for Medical Applications
abstract
Facial morphology measurement and classification play important role in the face anthropometry of many medical applications. This usually involves the investigation of medical abnormalities where specific facial features are studied by taking a number of measurements of the facial area under investigation. The measurements are often obtained from the three-dimensional (3D) scans of the faces; however, the measurements are often made manually, which is tedious and time consuming process. Moreover, in gene related studies thousands of measurements may be necessary in order to find statistically significant relationships between facial features and genes. Normative studies, from which typical populous models can be built, also require many measurements. Thus an automatic method to extract morphological measurements and interpret them is desirable. In this article, an automatic method for classification of facial morphology on the basis of a number of geometric measurements obtained automatically from 3D facial scans is presented. Among different facial features the philtrum, which is the vertical groove extending from the nose to the upper lip and the lip area, plays an important role in defining the interaction between the genes and craniofacial anomalies such as, for example, cleft lip and palate. In this paper, geometric features are analysed for their suitability to classify philtrum into three classes previously proposed by medical experts. Moreover, further analysis is conducted to assess the best number of classes to model the underlying data distribution from the point of view of classification accuracy. The obtained classification results are compared with the ground truth manual labelling of 3D face meshes provided by a medical expert. The dataset used for this research is taken from ALSPAC dataset and consists of 1000 3D face meshes. The proposed method achieves classification accuracy of 97% for this data set using the Mean, Minimum and Maximum curvature features in combination.
Hawraa Abbas, Yulia Hicks, David Marshall 0001
KES3
2015 Feature Neighbourhood Mutual Information for multi-modal image registration: An application to eye fundus imaging
Philip A. Legg, Paul L. Rosin, David Marshall 0001, James E. Morgan
Pattern Recognit.3
2014 Learnt Real-time Meshless Simulation
abstract
Abstract We present a new real‐time approach to simulate deformable objects using a learnt statistical model to achieve a high degree of realism. Our approach improves upon state‐of‐the‐art interactive shape‐matching meshless simulation methods by not only capturing important nuances of an object's kinematics but also of its dynamic texture variation. We are able to achieve this in an automated pipeline from data capture to simulation. Our system allows for the capture of idiosyncratic characteristics of an object's dynamics which for many simulations (e.g. facial animation) is essential. We allow for the plausible simulation of mechanically complex objects without knowledge of their inner workings. The main idea of our approach is to use a flexible statistical model to achieve a geometrically‐driven simulation that allows for arbitrarily complex yet easily learned deformations while at the same time preserving the desirable properties (stability, speed and memory efficiency) of current shape‐matching simulation systems. The principal advantage of our approach is the ease with which a pseudo‐mechanical model can be learned from 3D scanner data to yield realistic animation. We present examples of non‐trivial biomechanical objects simulated on a desktop machine in real‐time, demonstrating superior realism over current geometrically motivated simulation techniques.
Kirill A. Sidorov, David Marshall 0001
Comput. Graph. Forum2
2014 Facial expression recognition in dynamic sequences: An integrated approach
Hui Fang 0003, Neil Mac Parthaláin, Andrew J. Aubrey, Gary K. L. Tam, Rita Borgo, Paul L. Rosin, Phil W. Grant, David Marshall 0001, Min Chen 0001
Pattern Recognit.8
2014 Virtual unrolling and information recovery from scanned scrolled historical documents
Oksana Samko, Yukun Lai, David Marshall 0001, Paul L. Rosin
Pattern Recognit.3
2013 Making bas-reliefs from photographs of human faces
Jing Wu 0004, Ralph R. Martin, Paul L. Rosin, Xianfang Sun, Frank C. Langbein, Yukun Lai, David Marshall 0001
Comput. Aided Des.7
2013 Visualizing Natural Image Statistics
abstract
Natural image statistics is an important area of research in cognitive sciences and computer vision. Visualization of statistical results can help identify clusters and anomalies as well as analyze deviation, distribution, and correlation. Furthermore, they can provide visual abstractions and symbolism for categorized data. In this paper, we begin our study of visualization of image statistics by considering visual representations of power spectra, which are commonly used to visualize different categories of images. We show that they convey a limited amount of statistical information about image categories and their support for analytical tasks is ineffective. We then introduce several new visual representations, which convey different or more information about image statistics. We apply ANOVA to the image statistics to help select statistically more meaningful measurements in our design process. A task-based user evaluation was carried out to compare the new visual representations with the conventional power spectra plots. Based on the results of the evaluation, we made further improvement of visualizations by introducing composite visual representations of image statistics.
Hui Fang 0003, Gary K. L. Tam, Rita Borgo, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, Christian Wallraven, Douglas W. Cunningham, David Marshall 0001, Min Chen 0001
IEEE Trans. Vis. Comput. Graph.9
2013 Water Surface Modeling from a Single Viewpoint Video
abstract
We introduce a video-based approach for producing water surface models. Recent advances in this field output high-quality results but require dedicated capturing devices and only work in limited conditions. In contrast, our method achieves a good tradeoff between the visual quality and the production cost: It automatically produces a visually plausible animation using a single viewpoint video as the input. Our approach is based on two discoveries: first, shape from shading (SFS) is adequate to capture the appearance and dynamic behavior of the example water; second, shallow water model can be used to estimate a velocity field that produces complex surface dynamics. We will provide qualitative evaluation of our method and demonstrate its good performance across a wide range of scenes.
Chuan Li 0001, David Pickup, Thomas Saunders, Darren Cosker, David Marshall 0001, Peter Hall 0001, Philip J. Willis
IEEE Trans. Vis. Comput. Graph.5
2013 Registration of 3D Point Clouds and Meshes: A Survey from Rigid to Nonrigid
abstract
Three-dimensional surface registration transforms multiple three-dimensional data sets into the same coordinate system so as to align overlapping components of these sets. Recent surveys have covered different aspects of either rigid or nonrigid registration, but seldom discuss them as a whole. Our study serves two purposes: 1) To give a comprehensive survey of both types of registration, focusing on three-dimensional point clouds and meshes and 2) to provide a better understanding of registration from the perspective of data fitting. Registration is closely related to data fitting in which it comprises three core interwoven components: model selection, correspondences and constraints, and optimization. Study of these components 1) provides a basis for comparison of the novelties of different techniques, 2) reveals the similarity of rigid and nonrigid registration in terms of problem representations, and 3) shows how overfitting arises in nonrigid registration and the reasons for increasing interest in intrinsic techniques. We further summarize some practical issues of registration which include initializations and evaluations, and discuss some of our own observations, insights and foreseeable research trends.
Gary K. L. Tam, Zhi-Quan Cheng, Yukun Lai, Frank C. Langbein, Yonghuai Liu, David Marshall 0001, Ralph R. Martin, Xianfang Sun, Paul L. Rosin
IEEE Trans. Vis. Comput. Graph.6
2012 Hybrid phoneme based clustering approach for audio driven facial animation
abstract
We consider the problem of producing accurate facial animation corresponding to a given input speech signal. A popular technique previously used for Audio Driven Facial Animation is to build a joint audio-visual model using Active Appearance Models (AAMs) to represent possible facial variations and Hidden Markov Models (HMMs) to select the correct appearance based on the input audio. However there are several questions that remained unanswered. In particular the choice of clustering technique and the choice of the number of clusters in the HMM may have significant influence over the quality of the produced videos. We have investigated a range of clustering techniques in order to improve the quality of the HMM produced, and proposed a new structure based on using Gaussian Mixture Models (GMMs) to model each phoneme separately. We compared our approach to several alternatives using a public dataset of 300 phonetically labeled sentences spoken by a single person and found that our approach produces more accurate animation. In addition, we use a hybrid approach where the training data is phonetically labeled thus producing a model with better separation of phonemes, but test audio data is not labeled, thus making our approach for generating facial animation less laborious and fully automatic.
Benjamin Havell, Paul L. Rosin, Saeid Sanei, Andrew J. Aubrey, David Marshall 0001, Yulia Hicks
ICASSP5
2011 Segmentation of Parchment Scrolls for Virtual Unrolling
abstract
In this paper we introduce a framework for the segmentation of scanned scrolled parchments, based on a novel graph cut based approach with an additional shape prior, in combination with anisotropic diffusion and geometry-constrained postprocessing. This problem has not been investigated by the computer vision community properly yet due to the parchment scanning technology novelty, and is extremely important for effective data recovery from historical scrolled documents whose content is inaccessible due to the deterioration of the parchment. To date, parchment segmentation has required user interaction, which is very time consuming for such data. We demonstrate with real examples how our algorithm is able to solve the major problem for scrolled parchment analysis, namely segment connected layers, and process the data without user interaction.
Oksana Samko, Yukun Lai, David Marshall 0001, Paul L. Rosin
BMVC3
2011 Efficient groupwise non-rigid registration of textured surfaces
abstract
Advances in 3D imaging have recently made 3D surface scanners, capable of capturing textured surfaces at video rate, affordable and common in computer vision. This is a relatively new source of data, the potential of which has not yet been fully exploited as the problem of non-rigid registration of surfaces is difficult. While registration based on shape alone has been an active research area for some time, the problem of registering surfaces based on texture information has not been addressed in a principled way. We propose a novel, efficient and reliable, fully automatic method for performing groupwise non-rigid registration of textured surfaces, such as those obtained with 3D scanners. We demonstrate the robustness of our approach on 3D scans of human faces, including the notoriously difficult case of inter-subject registration. We show how our method can be used to build high-quality 3D models of appearance fully automatically.
Kirill A. Sidorov, Stephen Richmond, David Marshall 0001
CVPR3
2011 Mapping and manipulating facial dynamics
abstract
This paper describes a novel approach to building models of temporal dynamics for facial animation with applications in performing perceptual testing of trustworthiness. A vital component of the system is a method to bring two image sequences into temporal alignment. Our approach is to project the two sequences into face space (built using shape models [1]) and apply dynamic time warping (DTW). However, the variability in the sequences causes the standard DTW algorithm to perform poorly on our data, and so we have overcome this by extending DTW in the following ways: 1) the signal magnitudes are augmented by incorporating derivatives [2], and a scheme for estimating weights in the cost function is proposed, 2) the set of sequences is used to build a graph, with nodes representing sequences and edges indicating the cost of applying the extended DTW to align pairs of sequences; better alignments between sequences can now be found by traversing the minimum cost path through the graph. Once all signals are aligned to a common temporal reference it is straightforward to map the temporal dynamics from one face to another. A remapped face is synthesised using the new trajectory in face space to drive an active appearance model [1]. Furthermore, the common temporal reference allows us to build a statistical model of the dynamics. This can be used to both identify dynamics of interest and also to manipulate the dynamics, e.g. to reduce or exaggerate facial dynamics.
Andrew J. Aubrey, Vedran Kajic, Ivana Cingovska, Paul L. Rosin, David Marshall 0001
FG5
2011 Intelligent filtering by semantic importance for single-view 3D reconstruction from Snooker video
abstract
In this paper we investigate the challenge of 3D reconstruction from Snooker video data. We propose a system pipeline for intelligent filtering based on semantic importance in Snooker. The system can be divided into table detection and correction, followed by ball detection, classification and tracking. It is apparent from previous work that there are several challenges presented here. Firstly, previous methods tend to use a fixed top-down camera mounted above the table. To capture a full table view from this is challenging due to space limitations above the table. Instead, we capture video data from a tripod and correct the viewpoint through processing. Secondly, previous methods tend to simply detect the balls without considering other interfering objects such as player and cue. This becomes even more apparent when the player strikes the cue ball. Our intelligent filtering avoids such issues to give accurate 3D table reconstruction.
Philip A. Legg, Matthew L. Parry, David H. S. Chung, Richard Jiang 0001, Adrian Morris, Iwan W. Griffiths, David Marshall 0001, Min Chen 0001
ICIP7
2011 Reinforcing conceptual engineering design with a hybrid computer vision, machine learning and knowledge based system framework
abstract
We propose a novel system that aids engineers in the conceptual stage of design. Our system's goal is to support the engineer without limiting his creative role; thus, our proposed method does not produce ready study solutions but rather actively monitors the design procedure, verifying design stages and pointing out potential mistakes. This is achieved with a hybrid computer vision, machine learning and knowledge based system framework. Design stage identification is performed with a novel algorithm which comprises a classification stage based on Random Forests and examination of the temporal relationships between the engineer's actions with the aid of statistical graphical models. Experimental results captured in a complex, real life scenario demonstrate our system's ability to efficiently support the engineer's decisions during the conceptual stage of design.
Ioannis Kaloskampis, Yulia Hicks, David Marshall 0001
SMC3
2011 Visualization of Time-Series Data in Parameter Space for Understanding Facial Dynamics
abstract
Abstract Over the past decade, computer scientists and psychologists have made great efforts to collect and analyze facial dynamics data that exhibit different expressions and emotions. Such data is commonly captured as videos and are transformed into feature‐based time‐series prior to any analysis. However, the analytical tasks, such as expression classification, have been hindered by the lack of understanding of the complex data space and the associated algorithm space. Conventional graph‐based time‐series visualization is also found inadequate to support such tasks. In this work, we adopt a visual analytics approach by visualizing the correlation between the algorithm space and our goal – classifying facial dynamics. We transform multiple feature‐based time‐series for each expression in measurement space to a multi‐dimensional representation in parameter space. This enables us to utilize parallel coordinates visualization to gain an understanding of the algorithm space, providing a fast and cost‐effective means to support the design of analytical algorithms.
Gary K. L. Tam, Hui Fang 0003, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, David Marshall 0001, Min Chen 0001
Comput. Graph. Forum6
2010 Empty set biasing issues in the Transferable Belief Model for fusing and decision making
Gavin Powell, David Marshall 0001
FUSION3
2010 Improving Joint Tracking and Classification with the Transferable Belief Model and terrain information
David Marshall 0001, Gavin Powell
FUSION2
2010 Assessing the Uniqueness and Permanence of Facial Actions for Use in Biometric Applications
abstract
Although the human face is commonly used as a physiological biometric, very little work has been done to exploit the idiosyncrasies of facial motions for person identification. In this paper, we investigate theuniquenessandpermanenceof facial actions to determine whether these can be used as a behavioral biometric. Experiments are carried out using 3-D video data of participants performing a set of very short verbal and nonverbal facial actions. The data have been collected over long time intervals to assess the variability of the subjects' emotional and physical conditions. Quantitative evaluations are performed for both the identification and the verification problems; the results indicate that emotional expressions (e.g., smile and disgust) are not sufficiently reliable for identity recognition in real-life situations, whereas speech-related facial movements show promising potential.
Lanthao Benedikt, Darren Cosker, Paul L. Rosin, David Marshall 0001
IEEE Trans. Syst. Man Cybern. Part A4
2009 An efficient stochastic approach to groupwise non-rigid image registration
abstract
The groupwise approach to non-rigid image registration, solving the dense correspondence problem, has recently been shown to be a useful tool in many applications, including medical imaging, automatic construction of statistical models of appearance and analysis of facial dynamics. Such an approach overcomes limitations of traditional pairwise methods but at a cost of having to search for the solution (optimal registration) in a space of much higher dimensionality which grows rapidly with the number of examples (images) being registered. Techniques to overcome this dimensionality problem have not been addressed sufficiently in the groupwise registration literature. In this paper, we propose a novel, fast and reliable, fully unsupervised stochastic algorithm to search for optimal groupwise dense correspondence in large sets of unmarked images. The efficiency of our approach stems from novel dimensionality reduction techniques specific to the problem of groupwise image registration and from comparative insensitivity of the adopted optimisation scheme (simultaneous perturbation stochastic approximation (SPSA)) to the high dimensionality of the search space. Additionally, our algorithm is formulated in way readily suited to implementation on graphics processing units (GPU). In evaluation of our method we show a high robustness and success rate, fast convergence on various types of test data, including facial images featuring large degrees of both inter- and intra-person variation, and show considerable improvement in terms of accuracy of solution and speed compared to traditional methods.
Kirill A. Sidorov, Stephen Richmond, David Marshall 0001
CVPR3
2009 A Robust Solution to Multi-modal Image Registration by Combining Mutual Information with Multi-scale Derivatives
Philip A. Legg, Paul L. Rosin, David Marshall 0001, James E. Morgan
MICCAI (1)3
2008 Facial Dynamics in Biometric Identification
abstract
This paper investigates the use of facial gestures for identity recognition. This is the first time that such a quantitative evaluation is conducted, comparing the analyses of 2D versus 3D dynamic data of verbal and nonverbal facial actions. Suitable data processing and feature extraction methods are examined, then a number of pattern matching techniques including the Fréchet distance, Correlation Coefficients, Hidden-Markov Models, Dynamic Time Warping and its derived forms are compared, in light of which an improved algorithm is proposed. Finally, a face recognition prototype using facial dynamics is built, achieving an Equal Error Rate EER=1.6%. 1
Lanthao Benedikt, Vedran Kajic, Darren Cosker, Paul L. Rosin, David Marshall 0001
BMVC5
2006 A model of diatom shape and texture for analysis, synthesis and identification
Yulia Hicks, David Marshall 0001, Paul L. Rosin, Ralph R. Martin, David G. Mann, S. J. M. Droop
Mach. Vis. Appl.2
2006 Selection of the optimal parameter value for the Isomap algorithm
Oksana Samko, David Marshall 0001, Paul L. Rosin
Pattern Recognit. Lett.2
2005 Toward Perceptually Realistic Talking Heads: Models, Methods, and McGurk
abstract
Motivated by the need for an informative, unbiased, and quantitative perceptual method for the evaluation of a talking head we are developing, we propose a new test based on the “McGurk Effect.” Our approach helps to identify strengths and weaknesses in visual--speech synthesis algorithms for talking heads and facial animations, in general, and uses this insight to guide further development. We also evaluate the behavioral quality of our facial animations in comparison to real-speaker footage and demonstrate our tests by applying them to our current speech-driven facial animation system.
Darren Cosker, David Marshall 0001, Paul L. Rosin, Susan Paddock, Simon K. Rushton
ACM Trans. Appl. Percept.2
2004 Local topological beautification of reverse engineered models
C. H. Gao, Frank C. Langbein, David Marshall 0001, Ralph R. Martin
Comput. Aided Des.3
2004 Choosing consistent constraints for beautification of reverse engineered geometric models
Frank C. Langbein, David Marshall 0001, Ralph R. Martin
Comput. Aided Des.2
2004 Proceedings of the 13th British Machine Vision Conference
Paul L. Rosin, David Marshall 0001
Image Vis. Comput.2
2003 Approximate Congruence Detection of Model Features for Reverse Engineering
abstract
Reverse engineering allows the geometric reconstruction of simple mechanical parts. However, the resulting models suffer from inaccuracies caused by errors in measurement and reconstruction so such models do not have the exact congruence, symmetries and other regularities the original designer intended. We wish to impose such regularities in a beautification process. The paper discusses the particular problem of detecting approximate congruence between parts (e.g. a pair of handles) of a reconstructed B-rep model, so that a subsequent step can enforce them exactly. A practical detection algorithm is given for models defined using planes, spheres, cylinders, cones and tori. Analysis of the algorithm and experimental results show that expected congruence are detected reasonably quickly.
C. H. Gao, Frank C. Langbein, David Marshall 0001, Ralph R. Martin
Shape Modeling International3
2002 Modelling life cycle related and individual shape variation in biological specimens
abstract
The main purpose of this research is to develop methods for automatic identification of biological specimens in digital photographs and drawings held in a database. Incorporation of taxonomic drawings into a visual indexing system has not been attempted to date. Diatoms are a single cell microscopic algae that provide a particularly suitable case study. Identification of diatoms is a challenging task due to the huge number of the species, blurred boundaries between species, and life cycle related shape changes. A novel model based on principal curves representing the life cycle related shape variation of a number of diatom species has been developed. Our model is suitable for reconstruction purposes, allowing us to produce drawings of a variety of diatom shapes, thus providing a link between the photographs and drawings. We present the classification results of photographed and drawn specimens based on the model and compare our results to another recent system for diatom identification. Finally, given a diatom specimen, we are able not only to identify the species it belongs to but also to pinpoint the stage in the life cycle it represents.
Yulia Hicks, David Marshall 0001, Ralph R. Martin, Paul L. Rosin, Micha Bayer, David G. Mann
BMVC2
2002 Numerical Methods for Beautification of Reverse Engineered Geometric Models
abstract
Boundary representation models reconstructed from 3D range data suffer from various inaccuracies caused by noise in the data and the model building software. The quality of such models can be improved in a beautification step, which finds geometric regularities approximately present in the model and tries to impose a consistent subset of these regularities on the model. A framework for beautification and numerical methods to select and solve a consistent set of constraints deduced from a set of regularities are presented. For the initial selection of consistent regularities likely to be part of the model's ideal design priorities, and rules indicating simple inconsistencies between the regularities are employed. By adding regularities consecutively to an equation system and trying to solve it by using quasi-Newton optimization methods, inconsistencies and redundancies are detected. The results of experiments are encouraging and show potential for an expansion of the methods based on degree of freedom analysis.
Frank C. Langbein, David Marshall 0001, Ralph R. Martin
GMP2
2002 Automatic landmarking for building biological shape models
abstract
We present a new method for automatic landmark extraction from the contours of biological specimens. Our ultimate goal is to enable automatic identification of biological specimens in photographs and drawings held in a database. We propose to use active appearance models for visual indexing of both photographs and drawings. Automatic landmark extraction will assist us in building the models. We describe the results of using our method on drawings and photographs of examples of diatoms, and present an active shape model built using automatically extracted data.
Yulia Hicks, David Marshall 0001, Ralph R. Martin, Paul L. Rosin, Micha Bayer, David G. Mann
ICIP (2)2
2002 Multimodal retinal imaging: new strategies for the detection of glaucoma
abstract
Glaucoma is a serious worldwide disease whose treatment can be improved by early detection. As part of a new clinical approach this paper introduces some preliminary studies in the computerised detection of the disease. In particular we consider the problems of registering 3D laser data with a digital image. We introduce a new method based on windowed mutual information and show that it performs better than the standard mutual information technique.
Paul L. Rosin, David Marshall 0001, James E. Morgan
ICIP (3)2
2002 A prototype store choice and location modelling system using Dempster-Shafer theory
abstract
This paper concerns the study of destination choice modelling, more specifically identifying within some area (e.g. a city) the region where a particular store is the most favourable to be visited by individuals. An influence measure is constructed for each individual, which incorporates the modern technique known as Dempster–Shafer theory. Based on the evidence of the shopping destinations of individuals, geographical regions are found for levels of largest belief and plausibility (within Dempster–Shafer theory) for specific stores being the most favourable to visit. Additionally, this method may be used to identify the possible position of new stores, based on regions of most uncertainty or conflict in store choice. A prototype choice modelling system is introduced to enable the series of associated results to be easily visualized and analysed.
Malcolm J. Beynon, Benjamin Griffiths, David Marshall 0001
Expert Syst. J. Knowl. Eng.3
2002 Adding and subtracting eigenspaces with eigenvalue decomposition and singular value decomposition
Peter Hall 0001, David Marshall 0001, Ralph R. Martin
Image Vis. Comput.2
2002 Tracking people in three dimensions using a hierarchical model of dynamics
Yulia Hicks, Peter Hall 0001, David Marshall 0001
Image Vis. Comput.3
2001 Recognizing Geometric Patterns for Beautification of Reconstructed Solid Models
abstract
Boundary representation models reconstructed from 3D range data suffer from various inaccuracies caused by noise in the data and the model building software. The quality of such models can be improved in a beautification step, which finds regular geometric patterns approximately present in the model and imposes a maximal consistent subset of constraints deduced from these patterns on the model. This paper presents analysis methods seeking geometric patterns defined by similarities. Their specific types are derived from a part survey estimating the frequencies of the patterns in simple mechanical components. The methods seek clusters of similar objects which describe properties of faces, loops, edges and vertices, try to find special values representing the clusters, and seek approximate symmetries of the model. Experiments show that the patterns detected appear to be suitable for the subsequent beautification steps.
Frank C. Langbein, Bruce I. Mills, David Marshall 0001, Ralph R. Martin
Shape Modeling International3
2001 An expert system for multi-criteria decision making using Dempster Shafer theory
Malcolm J. Beynon, Darren Cosker, David Marshall 0001
Expert Syst. Appl.3
2001 Robust Segmentation of Primitives from Range Data in the Presence of Geometric Degeneracy
abstract
This paper addresses a common problem in the segmentation of range images. We present methods for the least-squares fitting of spheres, cylinders, cones, and tori to 3D point data, and their application within a segmentation framework. Least-squares fitting of surfaces other than planes, even of simple geometric type, has rarely been studied. Our main application areas of this research are reverse engineering of solid models from depth-maps and automated 3D inspection where reliable extraction of these surfaces is essential. Our fitting method has the particular advantage of being robust in the presence of geometric degeneracy, i.e., as the principal curvatures of the surfaces being fitted decrease, the results returned naturally become closer and closer to those surfaces of "simpler type", i.e., planes, cylinders, cones, or spheres, which best describe the data. Many other methods diverge because, in such cases, various parameters or their combination become infinite.
David Marshall 0001, Gábor Lukács, Ralph R. Martin
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 A Hierarchical Model of Dynamics for Tracking People with a Single Video Camera
abstract
We propose a novel hierarchical model of human dynamics for view independent tracking of the human body in monocular video sequences. The model is trained using real data from a collection of people. Kinematics are encoded using Hierarchical Principal Component Analysis, and dynamics are encoded using Hidden Markov Models. The top of the hierarchy contains information about the whole body. The lower levels of the hierarchy contain more detailed information about possible poses of some subpart of the body. When tracking, the lower levels of the hierarchy are shown to improve accuracy. In this article we describe our model and present experiments that show we can recover 3D skeletons from 2D images in a view independent manner, and also track people the system was not trained on.
Yulia Hicks, Peter Hall 0001, David Marshall 0001
BMVC3
2000 Simulation of FLIR and LADAR Data Using Graphics Animation Software
abstract
Presents an implementation of forward-looking infrared (FLIR) and laser radar (LADAR) data simulation for use in developing a multi-sensor data-fusion automated target recognition (ATR) system. Through the use of commercial models and software, we can create a highly detailed scene model, which provides a rich data set for later processing. We embed our own modules within this software to extract data from the scene and then present it to the FLIR and LADAR sensor models. These models produce simulated range and temperature readings for objects within the scene. These data are used to create a variety of images to aid in the visualisation of the data and to test our ATR system. Frames from approach sequences show the LADAR and FLIR data in different formats. The simulated data and subsequent images are accurate and rapid to produce, and provide invaluable data resources for testing an ATR system.
Gavin Powell, Ralph R. Martin, David Marshall 0001, Keith Markham
PG3
2000 Merging and Splitting Eigenspace Models
abstract
We present new deterministic methods that, given two eigenspace models-each representing a set of n-dimensional observations-will: 1) merge the models to yield a representation of the union of the sets and 2) split one model from another to represent the difference between the sets. As this is done, we accurately keep track of the mean. Here, we give a theoretical derivation of the methods, empirical results relating to the efficiency and accuracy of the techniques, and three general applications, including the construction of Gaussian mixture models that are dynamically updateable.
Peter Hall 0001, David Marshall 0001, Ralph R. Martin
IEEE Trans. Pattern Anal. Mach. Intell.2
1999 Adding and Subtracting Eigenspaces
abstract
This paper provides two algorithms; one for adding eigenspaces, another for subtracting them, thus allowing for incremental updating and downdat-ing of data models. Importantly, and unlike previous work, we keep an ac-curate track of the mean of the data, which allows our methods to be used in classification applications. The result of adding eigenspaces, each made from a set of data, is an approximation to that which would obtain were the sets of data taken together. Subtracting eigenspaces yields a result approximating that which would obtain were a subset of data used. Using our algorithms it is possible to perform “arithmetic ” on eigenspaces without reference to the orig-inal data. We illustrate the use of our algorithms in three generic applications, including the dynamic construction of Gaussian mixture models. 1
Peter Hall 0001, David Marshall 0001, Ralph R. Martin
BMVC2
1998 Incremental Eigenanalysis for Classification
abstract
Eigenspace models are a convenientway to represent sets of observations with widespread applications, including classi#cation. In this paper we describe a new constructive method for incrementally adding observations to an eigenspace model. Our contribution is to explicitly account forachange in origin as well as a change in the number of eigenvectors needed in the basis set. No other method wehave seen considers change of origin, yet both are needed if an eigenspace model is to be used for classi#cation purposes. We empirically compare our incremental method with two alternatives from the literature and show our method is the more useful for classi#cation because it computes the smaller eigenspace model representing the observations. 1 Introduction The contribution of this paper is a method for incrementally computing eigenspace models in the context of using them for classi#cation. Eigenspace models are widely used in computer vision. Applications include: face recognition #8# where...
Peter Hall 0001, David Marshall 0001, Ralph R. Martin
BMVC2
1998 Viewpoint Selection for Complete Surface Coverage of Three Dimensional Objects
abstract
Many machine vision tasks, e.g. object recognition and object inspection, cannot be performed robustly from a single image. For certain tasks (e.g. 3D object recognition and automated inspection) the availability of multiple views of an object is a requirement. This paper presents a novel approach to selecting a minimised number of views that allow each object face to be adequately viewed according to specified constraints on viewpoints and other features. The planner is generic and can be employed for a wide range of multiple view acquisition systems, ranging from camera systems mounted on the end of a robot arm, i.e. an eye-in-hand camera setup, to a turntable and fixed stereo cameras to allow different views of an object to be obtained. The results (both simulated and real) given focus on planning with a fixed camera and turntable. 1 Introduction Considerable research in machine vision has been directed at object recognition and object inspection. For complete 3D recogn...
D. R. Roberts, David Marshall 0001
BMVC2
1998 Faithful Least-Squares Fitting of Spheres, Cylinders, Cones and Tori for Reliable Segmentation
Gábor Lukács, Ralph R. Martin, David Marshall 0001
ECCV (1)3
1996 Region Template Correlation for FLIR Target Tracking
abstract
This paper deals with the problem of tracking an object from a sequence of images captured from a camera moving towards the object. Many conventional image based trackers trace key reference points over an image sequence using simple correlation processes [1]. However, the presence of noise and the magnification of image features as the camera approaches the object may cause a tracked point to drift. This paper introduces an alternative technique based on tracking regions formed by image segmentation. A region template of the target is acquired in the first frame and correlated over segmented images in future frames. The new technique produces significant reductions in drift rate and results have been obtained for real forward looking infrared (FLIR) images. This paper then discusses the advantages of integrating the new technique with a conventional point correlation based tracker. 1 Introduction The tracking of features in an image sequence is important in the areas of ...
H. S. Parry, David Marshall 0001, Keith Markham
BMVC2
1996 Interactive hypermedia courseware for the World Wide Web
abstract
article Interactive hypermedia courseware for the World Wide Web Share on Authors: A. D. Marshall Department of Computing Science, University of Wales, Cardiff, CF2 3XF Department of Computing Science, University of Wales, Cardiff, CF2 3XFView Profile , S. Hurley Department of Computing Science, University of Wales, Cardiff, CF2 3XF Department of Computing Science, University of Wales, Cardiff, CF2 3XFView Profile Authors Info & Claims ACM SIGCUE OutlookVolume 24Issue 1-3Jan.-July, 1996 pp 1–5https://doi.org/10.1145/1013718.237478Online:01 January 1996Publication History 11citation793DownloadsMetricsTotal Citations11Total Downloads793Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
David Marshall 0001, Stephen Hurley
ITiCSE1
1996 The Design, Development and Evaluation of Hypermedia Courseware for the World Wide Web
David Marshall 0001, Stephen Hurley
Multim. Tools Appl.1
1991 Automatic inspection of mechanical parts using geometric models and laser range finder data
David Marshall 0001, Ralph R. Martin, David Hutber
Image Vis. Comput.1