David C. Hogg

dblp:h/DHogg · also David Crossland Hogg, David Hogg 0001 · DBLP profile ↗
← Back
104ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-6125-9564ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 87 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 66 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Score Before You Speak: Improving Persona Consistency in Dialogue Generation Using Response Quality Scores
abstract
Persona-based dialogue generation is an important milestone towards building conversational artificial intelligence. Despite the ever-improving capabilities of large language models (LLMs), effectively integrating persona fidelity in conversations remains challenging due to the limited diversity in existing dialogue data. We propose a novel framework SBS (Score-Before-Speaking), which outperforms previous methods and yields improvements for both million and billion-parameter models. Unlike previous methods, SBS unifies the learning of responses and their relative quality into a single step. The key innovation is to train a dialogue model to correlate augmented responses with a quality score during training and then leverage this knowledge at inference. We use noun-based substitution for augmentation and semantic similarity-based scores as a proxy for response quality. Through extensive experiments with benchmark datasets (PERSONA-CHAT and ConvAI2), we show that score-conditioned training allows existing models to better capture a spectrum of persona-consistent dialogues. Our ablation studies also demonstrate that including scores in the input prompt during training is superior to conventional training setups. Code and further details are available at https://arpita2512.github.io/score_before_you_speak.
Arpita Saggar, Jonathan C. Darling, Vania Dimitrova, Duygu Sarikaya, David C. Hogg
ECAI5
2025 Learning to sculpt neural cityscapes
abstract
Abstract We introduce a system that learns to sculpt 3D models of massive urban environments. The majority of humans live their lives in urban environments, using detailed virtual models for applications as diverse as virtual worlds, special effects, and urban planning. Generating such 3D models from exemplars manually is time-consuming, while 3D deep learning approaches have high memory costs. In this paper, we present a technique for training 2D neural networks to repeatedly sculpt a plane into a large-scale 3D urban environment. An initial coarse depth map is created by a GAN model, from which we refine 3D normal and depth using an image translation network regularized by a linear system. The networks are trained using real-world data to allow generative synthesis of meshes at scale. We exploit sculpting from multiple viewpoints to generate a highly detailed, concave, and water-tight 3D mesh. We show cityscapes at scales of $$100 \times 1600$$ 100 × 1600 meters with more than 2 million triangles, and demonstrate that our results are objectively and subjectively similar to our exemplars.
Jialin Zhu 0001, He Wang 0002, David C. Hogg
Vis. Comput.3
2024 Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
abstract
Artificial intelligence (AI) has made remarkable progress across various domains, with large language models like ChatGPT gaining substantial attention for their human-like text-generation capabilities. Despite these achievements, improving spatial reasoning remains a significant challenge for these models. Benchmarks like StepGame evaluate AI spatial reasoning, where ChatGPT has shown unsatisfactory performance. However, the presence of template errors in the benchmark has an impact on the evaluation results. Thus there is potential for ChatGPT to perform better if these template errors are addressed, leading to more accurate assessments of its spatial reasoning capabilities. In this study, we refine the StepGame benchmark, providing a more accurate dataset for model evaluation. We analyze GPT’s spatial reasoning performance on the rectified benchmark, identifying proficiency in mapping natural language text to spatial relations but limitations in multi-hop reasoning. We provide a flawless solution to the benchmark by combining template-to-relation mapping with logic-based reasoning. This combination demonstrates proficiency in performing qualitative reasoning on StepGame without encountering any errors. We then address the limitations of GPT models in spatial reasoning. To improve spatial reasoning, we deploy Chain-of-Thought and Tree-of-thoughts prompting strategies, offering insights into GPT’s cognitive process. Our investigation not only sheds light on model deficiencies but also proposes enhancements, contributing to the advancement of AI with more robust spatial reasoning capabilities.
Fangjun Li, David C. Hogg, Anthony G. Cohn 0001
AAAI2
2024 Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
Fangjun Li, David C. Hogg, Anthony G. Cohn 0001
IJCAI2
2024 Understanding the vulnerability of skeleton-based Human Activity Recognition via black-box attack
Yunfeng Diao, He Wang 0002, Tianjia Shao, Kun Zhou 0001, David C. Hogg, Meng Wang 0001
Pattern Recognit.6
2023 3D Shape Reconstruction of Semi-Transparent Worms
abstract
3D shape reconstruction typically requires identifying object features or textures in multiple images of a subject. This approach is not viable when the subject is semi-transparent and moving in and out of focus. Here we overcome these challenges by rendering a candidate shape with adaptive blurring and transparency for comparison with the images. We use the microscopic nematode Caenorhabditis elegans as a case study as it freely explores a 3D complex fluid with constantly changing optical properties. We model the slender worm as a 3D curve using an intrinsic parametrisation that naturally admits biologically-informed constraints and regularisation. To account for the changing optics we develop a novel differentiable renderer to construct images from 2D projections and compare against raw images to generate a pixel-wise error to jointly update the curve, camera and renderer parameters using gradient descent. The method is robust to interference such as bubbles and dirt trapped in the fluid, stays consistent through complex sequences of postures, recovers reliable estimates from blurry images and provides a significant improvement on previous attempts to track C. elegans in 3D. Our results demonstrate the potential of direct approaches to shape estimation in complex physical environments in the absence of ground-truth data.
Thomas P. Ilett, Omer Yuval, Thomas Ranner, Netta Cohen, David C. Hogg
CVPR5
2023 Of Mice and Pose: 2D Mouse Pose Estimation from Unlabelled Data and Synthetic Prior
Jose Sosa, Sharn Perry, Jane E. Alty, David C. Hogg
ICVS4
2022 Using Graph Representation Learning with Schema Encoders to Measure the Severity of Depressive Symptoms
Simin Hong, Anthony G. Cohn 0001, David C. Hogg
ICLR3
2022 Talking Head from Speech Audio using a Pre-trained Image Generator
abstract
We propose a novel method for generating high-resolution videos of talking-heads from speech audio and a single 'identity' image. Our method is based on a convolutional neural network model that incorporates a pre-trained StyleGAN generator. We model each frame as a point in the latent space of StyleGAN so that a video corresponds to a trajectory through the latent space. Training the network is in two stages. The first stage is to model trajectories in the latent space conditioned on speech utterances. To do this, we use an existing encoder to invert the generator, mapping from each video frame into the latent space. We train a recurrent neural network to map from speech utterances to displacements in the latent space of the image generator. These displacements are relative to the back-projection into the latent space of an identity image chosen from the individuals depicted in the training dataset. In the second stage, we improve the visual quality of the generated videos by tuning the image generator on a single image or a short video of any chosen identity. We evaluate our model on standard measures (PSNR, SSIM, FID and LMD) and show that it significantly outperforms recent state-of-the-art methods on one of two commonly used datasets and gives comparable performance on the other. Finally, we report on ablation experiments that validate the components of the model. The code and videos from experiments can be found at https://mohammedalghamdi.github.io/talking-heads-acm-mm/
Mohammed M. Alghamdi, He Wang 0002, Andrew J. Bulpitt, David C. Hogg
ACM Multimedia4
2022 Online perceptual learning and natural language acquisition for autonomous robots
abstract
In this work, the problem of bootstrapping knowledge in language and vision for autonomous robots is addressed through novel techniques in grammar induction and word grounding to the perceptual world. In particular, we demonstrate a system, called OLAV, which is able, for the first time, to (1) learn to form discrete concepts from sensory data; (2) ground language (n-grams) to these concepts; (3) induce a grammar for the language being used to describe the perceptual world; and moreover to do all this incrementally, without storing all previous data. The learning is achieved in a loosely-supervised manner from raw linguistic and visual data. Moreover, the learnt model is transparent, rather than a black-box model and is thus open to human inspection. The visual data is collected using three different robotic platforms deployed in real-world and simulated environments and equipped with different sensing modalities, while the linguistic data is collected using online crowdsourcing tools and volunteers. The analysis performed on these robots demonstrates the effectiveness of the framework in learning visual concepts, language groundings and grammatical structure in these three online settings.
Muhannad Al-Omari, Fangjun Li, David C. Hogg, Anthony G. Cohn 0001
Artif. Intell.3
2021 Understanding the Robustness of Skeleton-Based Action Recognition Under Adversarial Attack
abstract
Action recognition has been heavily employed in many applications such as autonomous vehicles, surveillance, etc, where its robustness is a primary concern. In this paper, we examine the robustness of state-of-the-art action recognizers against adversarial attack, which has been rarely investigated so far. To this end, we propose a new method to attack action recognizers which rely on the 3D skeletal motion. Our method involves an innovative perceptual loss which ensures the imperceptibility of the attack. Empirical studies demonstrate that our method is effective in both white-box and black-box scenarios. Its generalizability is evidenced on a variety of action recognizers and datasets. Its versatility is shown in different attacking strategies. Its deceitfulness is proven in extensive perceptual studies. Our method shows that adversarial attack on 3D skeletal motions, one type of time-series data, is significantly different from traditional adversarial attack problems. Its success raises serious concern on the robustness of action recognizers and provides insights on potential improvements.
He Wang 0002, Feixiang He, Zhexi Peng, Tianjia Shao, Kun Zhou 0001, David C. Hogg
CVPR7
2019 A lattice-based approach for navigating design configuration spaces
abstract
Design configurations, such as Bills of Materials (BoMs), are indispensable parts of any product development process and integral to the design descriptions stored in proprietary Computer Aided Design and Product Lifecycle Management systems. Engineers use BoMs and other design configurations as lenses to repurpose design descriptions for specific purposes. For this reason, multiple BoMs typically occur in any given product development process. For example, an engineering BoM may be used to define a configuration that best supports a design activity whereas a manufacturing BoM may be used to define the configuration of parts that best supports a manufacturing process. Current practice for the definition of BoMs involves the use of indented parts lists and dendograms that are prone to error because it is easy to create discrepancies across BoMs that, in essence, are defined through collections of part identifiers such as names and part numbers. Such errors have a significant detrimental effect on the performance of product development processes by creating the need for rework, adding costs and increasing time to market. This paper introduces a design description capability that ensures consistency across BoMs for a given design. A boolean hypercube lattice is used to define a design configuration space that includes all possible configurations for a given design description. Valid operations within the space are governed by the mathematics of hypercube lattices. The design description capability is demonstrated through an early engineering design configuration software tool that offers significant benefits by ensuring consistency across the BoMs for a given design. The software uses and generates design descriptions that are exported from and imported to commercially available design systems through a standard (ISO 10303-214) interface format. In this way, potential for early impact on industry practice is high.
Alison McKay, Hau Hing Chau, Christopher F. Earl, Amar Kumar Behera, Alan de Pennington, David C. Hogg
Adv. Eng. Informatics6
2019 Unsupervised human activity analysis for intelligent mobile robots
abstract
The success of intelligent mobile robots operating and collaborating with humans in daily living environments depends on their ability to generalise and learn human movements, and obtain a shared understanding of an observed scene. In this paper we aim to understand human activities being performed in real-world environments from long-term observation from an autonomous mobile robot. For our purposes, a human activity is defined to be a changing spatial configuration of a person's body interacting with key objects that provide some functionality within an environment. To alleviate the perceptual limitations of a mobile robot, restricted by its obscured and incomplete sensory modalities, potentially noisy visual observations are mapped into an abstract qualitative space in order to generalise patterns invariant to exact quantitative positions within the real world. A number of qualitative spatial-temporal representations are used to capture different aspects of the relations between the human subject and their environment. Analogously to information retrieval on text corpora, a generative probabilistic technique is used to recover latent, semantically-meaningful concepts in the encoded observations in an unsupervised manner. The small number of concepts discovered are considered as human activity classes, granting the robot a low-dimensional understanding of visually observed complex scenes. Finally, variational inference is used to facilitate incremental and continuous updating of such concepts that allows the mobile robot to efficiently learn and update its models of human activity over time resulting in efficient life-long learning.
Paul Duckworth, David C. Hogg, Anthony G. Cohn 0001
Artif. Intell.2
2018 Learning Hierarchical Models of Complex Daily Activities from Annotated Videos
abstract
Effective recognition of complex long-term activities is becoming an increasingly important task in artificial intelligence. In this paper, we propose a novel approach for building models of complex long-term activities. First, we automatically learn the hierarchical structure of activities by learning about the 'parent-child' relation of activity components from a video using the variability in annotations acquired using multiple annotators. This variability allows for extracting the inherent hierarchical structure of the activity in a video. We consolidate hierarchical structures of the same activity from different videos into a unified stochastic grammar describing the overall activity. We then describe an inference mechanism to interpret new instances of activities. We use three datasets, which have been annotated by multiple annotators, of daily activity videos to demonstrate the effectiveness of our system.
Jawad Tayyub, Majd Hawasly, David C. Hogg, Anthony G. Cohn 0001
WACV3
2017 Natural Language Acquisition and Grounding for Embodied Robotic Systems
abstract
We present a cognitively plausible novel framework capable of learning the grounding in visual semantics and the grammar of natural language commands given to a robot in a table top environment. The input to the system consists of video clips of a manually controlled robot arm, paired with natural language commands describing the action. No prior knowledge is assumed about the meaning of words, or the structure of the language, except that there are different classes of words (corresponding to observable actions, spatial relations, and objects and their observable properties). The learning process automatically clusters the continuous perceptual spaces into concepts corresponding to linguistic input. A novel relational graph representation is used to build connections between language and vision. As well as the grounding of language to perception, the system also induces a set of probabilistic grammar rules. The knowledge learned is used to parse new commands involving previously unseen objects.
Muhannad Al-Omari, Paul Duckworth, David C. Hogg, Anthony G. Cohn 0001
AAAI3
2017 Latent Dirichlet Allocation for Unsupervised Activity Analysis on an Autonomous Mobile Robot
abstract
For autonomous robots to collaborate on joint tasks with humans they require a shared understanding of an observed scene. We present a method for unsupervised learning of common human movements and activities on an autonomous mobile robot, which generalises and improves on recent results. Our framework encodes multiple qualitative abstractions of RGBD video from human observations and does not require external temporal segmentation. Analogously to information retrieval in text corpora, each human detection is modelled as a random mixture of latent topics. A generative probabilistic technique is used to recover topic distributions over an auto-generated vocabulary of discrete, qualitative spatio-temporal code words. We show that the emergent categories align well with human activities as interpreted by a human. This is a particularly challenging task on a mobile robot due to the varying camera viewpoints which lead to incomplete, partial and occluded human detections.
Paul Duckworth, Muhannad Al-Omari, James Charles, David C. Hogg, Anthony G. Cohn 0001
AAAI4
2017 Anomaly Detection using a Convolutional Winner-Take-All Autoencoder
Hanh Tran, David C. Hogg
BMVC2
2017 Grounding of Human Environments and Activities for Autonomous Robots
abstract
With the recent proliferation of human-oriented robotic applications in domestic and industrial scenarios, it is vital for robots to continually learn about their environments and about the humans they share their environments with. In this paper, we present a novel, online, incremental framework for unsupervised symbol grounding in real-world, human environments for autonomous robots. We demonstrate the flexibility of the framework by learning about colours, people names, usable objects and simple human activities, integrating state-of-the-art object segmentation, pose estimation, activity analysis along with a number of sensory input encodings into a continual learning framework. Natural language is grounded to the learned concepts, enabling the robot to communicate in a human-understandable way. We show, using a challenging real-world dataset of human activities as perceived by a mobile robot, that our framework is able to extract useful concepts, ground natural language descriptions to them, and, as a proof-of-concept, generate simple sentences from templates to describe people and the activities they are engaged in.
Muhannad Al-Omari, Paul Duckworth, Nils Bore, Majd Hawasly, David C. Hogg, Anthony G. Cohn 0001
IJCAI5
2016 Personalizing Human Video Pose Estimation
abstract
We propose a personalized ConvNet pose estimator that automatically adapts itself to the uniqueness of a person's appearance to improve pose estimation in long videos. We make the following contributions: (i) we show that given a few high-precision pose annotations, e.g. from a generic ConvNet pose estimator, additional annotations can be generated throughout the video using a combination of image-based matching for temporally distant frames, and dense optical flow for temporally local frames, (ii) we develop an occlusion aware self-evaluation model that is able to automatically select the high-quality and reject the erroneous additional annotations, and (iii) we demonstrate that these high-quality annotations can be used to fine-tune a ConvNet pose estimator and thereby personalize it to lock on to key discriminative features of the person's appearance. The outcome is a substantial improvement in the pose estimates for the target video using the personalized ConvNet compared to the original generic ConvNet. Our method outperforms the state of the art (including top ConvNet methods) by a large margin on three standard benchmarks, as well as on a new challenging YouTube video dataset. Furthermore, we show that training from the automatically generated annotations can be used to improve the performance of a generic ConvNet on other benchmarks.
James Charles, Tomas Pfister, Derek R. Magee, David C. Hogg, Andrew Zisserman
CVPR4
2016 Unsupervised Activity Recognition Using Latent Semantic Analysis on a Mobile Robot
abstract
We show that by using qualitative spatio-temporal abstraction methods, we can learn common human movements and activities from long term observation by a mobile robot. Our novel framework encodes multiple qualitative abstractions of RGBD video from detected activities performed by a human as encoded by a skeleton pose estimator. Analogously to informational retrieval in text corpora, we use Latent Semantic Analysis (LSA) to uncover latent, semantically meaningful, concepts in an unsupervised manner, where the vocabulary is occurrences of qualitative spatio-temporal features extracted from video clips, and the discovered concepts are regarded as activity classes. The limited field of view of a mobile robot represents a particular challenge, owing to the obscured, partial and noisy human detections and skeleton pose-estimates from its environment. We show that the abstraction into a qualitative space helps the robot to generalise and compare multiple noisy and partial observations in a real world dataset and that a vocabulary of latent activity classes (expressed using qualitative features) can be recovered.
Paul Duckworth, Muhannad Al-Omari, Yiannis Gatsoulis, David C. Hogg, Anthony G. Cohn 0001
ECAI4
2016 Feature Space Analysis for Human Activity Recognition in Smart Environments
abstract
Activity classification from smart environment data is typically done employing ad hoc solutions customised to the particular dataset at hand. In this work we introduce a general purpose collection of features for recognising human activities across datasets of different type, size and nature. The first experimental test of our feature collection achieves state of the art results on well known datasets, and we provide a feature importance analysis in order to compare the potential relevance of features for activity classification in different datasets.
Eris Chinellato, David C. Hogg, Anthony G. Cohn 0001
Intelligent Environments2
2016 Unsupervised Grounding of Textual Descriptions of Object Features and Actions in Video
Muhannad Al-Omari, Eris Chinellato, Yiannis Gatsoulis, David C. Hogg, Anthony G. Cohn 0001
KR4
2016 Adapting pedestrian detectors to new domains: A comprehensive review
Kyaw Kyaw Htike, David C. Hogg
Eng. Appl. Artif. Intell.2
2016 Weakly supervised activity analysis with spatio-temporal localisation
Feng Gu 0006, Muralikrishna Sridhar, Anthony G. Cohn 0001, David C. Hogg, Francisco Flórez-Revuelta, Dorothy Ndedi Monekosso, Paolo Remagnino
Neurocomputing4
2015 Joint Tracking and Event Analysis for Carried Object Detection
abstract
This paper proposes a novel method for jointly estimating the track of a moving object and the events in which it participates. The method is intended for dealing with generic objects that are hard to localise and track with the performance of current detection algorithms - our focus is on events involving carried objects. The tracks for other objects with which the target object interacts (e.g. the carrying person) are assumed to be given. The method is posed as maximisation of a posterior probability defined over event sequences and temporally-disjoint subsets of the tracklets from an earlier tracking process. The probability function is a Hidden Markov Model coupled with a term that penalises non-smooth tracks and large gaps in the observed data. We evaluate the method using tracklets output by three state of the art trackers on the new created MINDSEYE2015 dataset and demonstrate improved performance.
Aryana Tavanai, Muralikrishna Sridhar, Eris Chinellato, Anthony G. Cohn 0001, David C. Hogg
BMVC5
2015 Learning Relational Event Models from Video
abstract
Event models obtained automatically from video can be used in applications ranging from abnormal event detection to content based video retrieval. When multiple agents are involved in the events, characterizing events naturally suggests encoding interactions as relations. Learning event models from this kind of relational spatio-temporal data using relational learning techniques such as Inductive Logic Programming (ILP) hold promise, but have not been successfully applied to very large datasets which result from video data. In this paper, we present a novel framework REMIND (Relational Event Model INDuction) for supervised relational learning of event models from large video datasets using ILP. Efficiency is achieved through the learning from interpretations setting and using a typing system that exploits the type hierarchy of objects in a domain. The use of types also helps prevent over generalization. Furthermore, we also present a type-refining operator and prove that it is optimal. The learned models can be used for recognizing events from previously unseen videos. We also present an extension to the framework by integrating an abduction step that improves the learning performance when there is noise in the input data. The experimental results on several hours of video data from two challenging real world domains (an airport domain and a physical action verbs domain) suggest that the techniques are suitable to real world scenarios.
Krishna Sandeep Reddy Dubba, Anthony G. Cohn 0001, David C. Hogg, Mehul Bhatt, Frank Dylla
J. Artif. Intell. Res.3
2014 Qualitative and Quantitative Spatio-temporal Relations in Daily Living Activity Recognition
Jawad Tayyub, Aryana Tavanai, Yiannis Gatsoulis, Anthony G. Cohn 0001, David C. Hogg
ACCV (5)5
2014 Real-time Activity Recognition by Discerning Qualitative Relationships Between Randomly Chosen Visual Features
Ardhendu Behera, Anthony G. Cohn 0001, David C. Hogg
BMVC3
2014 Upper Body Pose Estimation with Temporal Sequential Forests
James Charles, Tomas Pfister, Derek R. Magee, David C. Hogg, Andrew Zisserman
BMVC4
2014 Weakly supervised pedestrian detector training by unsupervised prior learning and cue fusion in videos
abstract
The growth in the amount of collected video data in the past decade necessitates automated video analysis for which pedestrian detection plays a key role. Training a pedestrian detector using supervised machine learning requires tedious manual annotation of pedestrians in the form of precise bounding boxes. In this paper, we propose a novel weakly supervised algorithm to train a pedestrian detector that only requires annotations of estimated centers of pedestrians instead of bounding boxes. Our algorithm makes use of a pedestrian prior learnt in an unsupervised way from the video and this prior is fused with the given weak supervision information in a principled manner. We show on publicly available datasets that our weakly supervised algorithm reduces the cost of manual annotation by over 4 times while achieving similar performance to a pedestrian detector trained with bounding box annotations.
Kyaw Kyaw Htike, David C. Hogg
ICIP2
2014 Efficient Non-iterative Domain Adaptation of Pedestrian Detectors to Video Scenes
abstract
Pedestrian detection is an essential step in many important applications of Computer Vision. Most detectors require manually annotated ground-truth to train, the collection of which is labor intensive and time-consuming. Generally, this training data is from representative views of pedestrians captured from a variety of scenes. Unsurprisingly, the performance of a detector on a new scene can be improved by tailoring the detector to the specific viewpoint, background and imaging conditions of the scene. Unfortunately, for many applications it is not practical to acquire this scene-specific training data by hand. In this paper, we propose a novel algorithm to automatically adapt and tune a generic pedestrian detector to specific scenes which may possess different data distributions than the original dataset from which the detector was trained. Most state-of-the-art approaches can be inefficient, require manually set number of iterations to converge and some form of human intervention. Our algorithm is a step towards overcoming these problems and although simple to implement, our algorithm exceeds state-of-the-art performance.
Kyaw Kyaw Htike, David C. Hogg
ICPR2
2014 Context Aware Detection and Tracking
abstract
This paper presents a novel approach to incorporate multiple contextual factors into a tracking process, for the purpose of reducing false positive detections. While much previous work has focused on improving object detection on static images using context, these have not been integrated into the tracking process. Our hypothesis is that a significant improvement can result from the use of context in dynamically influencing the linking of object detections, during the tracking process. To verify this hypothesis, we augment a state of the art dynamic programming based tracker with contextual information by reformulating the maximum a posteriori (MAP) estimation formulation. This formulation introduces contextual factors that first of all augment detection strengths and secondly provides temporal context. We allow both these types of factors to contribute organically to the linking process by learning the relative contribution of each of these factors jointly during a gradient decent based optimisation process. Our experiments demonstrate that the proposed approach contributes to a significantly superior performance on a recent challenging video dataset, which captures complex scenes with a wide range of object types and diverse backgrounds.
Aryana Tavanai, Muralikrishna Sridhar, Feng Gu 0006, Anthony G. Cohn 0001, David C. Hogg
ICPR5
2013 Domain Adaptation for Upper Body Pose Tracking in Signed TV Broadcasts
abstract
The objective of this work is to estimate upper body pose for signers in TV broadcasts. Given suitable training data, the pose is estimated using a random forest body joint detector. However, obtaining such training data can be costly. The novelty of this paper is a method of transfer learning which is able to harness existing training data and use it for new domains. Our contributions are: (i) a method for adapting existing training data to generate new training data by synthesis for signers with different appearances, and (ii) a method for personalising training data. As a case study we show how the appearance of the arms for different clothing, specifically short and long sleeved clothes, can be modelled to obtain person-specific trackers. We demonstrate that the transfer learning and person specific trackers significantly improve pose estimation performance.
James Charles, Tomas Pfister, Derek R. Magee, David C. Hogg, Andrew Zisserman
BMVC4
2013 Automated Ground-Plane Estimation for Trajectory Rectification
Ian Hales, David C. Hogg, Kia Ng, Roger D. Boyle
CAIP (2)2
2013 Carried Object Detection and Tracking Using Geometric Shape Models and Spatio-temporal Consistency
Aryana Tavanai, Muralikrishna Sridhar, Feng Gu 0006, Anthony G. Cohn 0001, David C. Hogg
ICVS5
2013 Robust abandoned object detection integrating wide area visual surveillance and social context
James M. Ferryman, David C. Hogg, Jan Sochman, Ardhendu Behera, José A. Rodríguez-Serrano, Simon F. Worgan, Longzhen Li, Valerie Leung, Murray Evans, Philippe Cornic, Stéphane Herbin, Stefan Schlenger, Michael Dose
Pattern Recognit. Lett.2
2012 Egocentric Activity Monitoring and Recovery
Ardhendu Behera, David C. Hogg, Anthony G. Cohn 0001
ACCV (3)2
2012 Workflow Activity Monitoring Using Dynamics of Pair-Wise Qualitative Spatial Relations
Ardhendu Behera, Anthony G. Cohn 0001, David C. Hogg
MMM3
2012 Building semantic scene models from unconstrained video
Hannah M. Dee, Anthony G. Cohn 0001, David C. Hogg
Comput. Vis. Image Underst.3
2012 Explaining Activities as Consistent Groups of Events - A Bayesian Framework Using Attribute Multiset Grammars
Dima Damen, David C. Hogg
Int. J. Comput. Vis.2
2012 Detecting Carried Objects from Sequences of Walking Pedestrians
abstract
This paper proposes a method for detecting objects carried by pedestrians, such as backpacks and suitcases, from video sequences. In common with earlier work [14], [16] on the same problem, the method produces a representation of motion and shape (known as a temporal template) that has some immunity to noise in foreground segmentations and phase of the walking cycle. Our key novelty is for carried objects to be revealed by comparing the temporal templates against view-specific exemplars generated offline for unencumbered pedestrians. A likelihood map of protrusions, obtained from this match, is combined in a Markov random field for spatial continuity, from which we obtain a segmentation of carried objects using the MAP solution. We also compare the previously used method of periodicity analysis to distinguish carried objects from other protrusions with using prior probabilities for carried-object locations relative to the silhouette. We have reimplemented the earlier state-of-the-art method [14] and demonstrate a substantial improvement in performance for the new method on the PETS2006 data set. The carried-object detector is also tested on another outdoor data set. Although developed for a specific problem, the method could be applied to the detection of irregularities in appearance for other categories of object that move in a periodic fashion.
Dima Damen, David C. Hogg
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 In Memoriam: Mark Everingham
abstract
Recounts the career and contributions pf Mark Everingham.
Andrew Zisserman, John M. Winn, Andrew W. Fitzgibbon, Luc Van Gool, Josef Sivic, Christopher K. I. Williams, David C. Hogg
IEEE Trans. Pattern Anal. Mach. Intell.7
2011 Temporal Structure Models for Event Recognition
abstract
In many areas of visual surveillance, the observed activity follows re-occurring patterns.This paper demonstrates how such patterns can be exploited to improve the detection rate of independent event detectors.We present a temporal model based on pairwise correlations between event timings, which efficiently exploits limited training data.This is combined with the response from potentially heterogeneous independent event detectors to improve the robustness of detections over extended sequences.We demonstrate the efficacy of our system with rigorous testing on a large real-world dataset of aircraft servicing operations.We describe the implementation of a binary classifier based on local histograms of optical flow which is used as the independent event detector in our experiments.
John Greenall, David C. Hogg, Anthony G. Cohn 0001
BMVC2
2011 From Video to RCC8: Exploiting a Distance Based Semantics to Stabilise the Interpretation of Mereotopological Relations
Muralikrishna Sridhar, Anthony G. Cohn 0001, David C. Hogg
COSIT3
2011 Exploiting petri-net structure for activity classification and user instruction within an industrial setting
abstract
Live workflow monitoring and the resulting user interaction in industrial settings faces a number of challenges. A formal workflow may be unknown or implicit, data may be sparse and certain isolated actions may be undetectable given current visual feature extraction technology. This paper attempts to address these problems by inducing a structural workflow model from multiple expert demonstrations. When interacting with a naive user, this workflow is combined with spatial and temporal information, under a Bayesian framework, to give appropriate feedback and instruction. Structural information is captured by translating a Markov chain of actions into a simple place/transition petri-net. This novel petri-net structure maintains a continuous record of the current workbench configuration and allows multiple sub-sequences to be monitored without resorting to second order processes. This allows the user to switch between multiple sub-tasks, while still receiving informative feedback from the system. As this model captures the complete workflow, human inspection of safety critical processes and expert annotation of user instructions can be made. Activity classification and user instruction results show a significant on-line performance improvement when compared to the existing Hidden Markov Model or pLSA based state of the art. Further analysis reveals that the majority of our model's classification errors are caused by small de-synchronisation events rather than significant workflow deviations. We conclude with a discussion of the generalisability of the induced place/transition petri-net to other activity recognition tasks and summarise the developments of this model.
Simon F. Worgan, Ardhendu Behera, Anthony G. Cohn 0001, David C. Hogg
ICMI4
2011 Interleaved Inductive-Abductive Reasoning for Learning Complex Event Models
Krishna Sandeep Reddy Dubba, Mehul Bhatt, Frank Dylla, David C. Hogg, Anthony G. Cohn 0001
ILP4
2010 Unsupervised Learning of Event Classes from Video
abstract
We present a method for unsupervised learning of event classes from videos in which multiple actions might occur simultaneously. It is assumed that all such activities are produced from an underlying set of event class generators. The learning task is then to recover this generative process from visual data. A set of event classes is derived from the most likely decomposition of the tracks into a set of labelled events involving subsets of interacting tracks. Interactions between subsets of tracks are modelled as a relational graph structure that captures qualitative spatio-temporal relationships between these tracks. The posterior probability of candidate solutions favours decompositions in which events of the same class have a similar relational structure, together with other measures of well-formedness. A Markov Chain Monte Carlo (MCMC) procedure is used to efficiently search for the MAP solution. This search moves between possible decompositions of the tracks into sets of unlabelled events and at each move adds a close to optimal labelling (for this decomposition) using spectral clustering. Experiments on real data show that the discovered event classes are often semantically meaningful and correspond well with groundtruth event classes assigned by hand.
Muralikrishna Sridhar, Anthony G. Cohn 0001, David C. Hogg
AAAI3
2010 Segmentation using Deformable Spatial Priors with Application to Clothing
abstract
We present a method for segmenting the parts of multiple instances of a known object category exhibiting large variations in projected shape and colour. The method builds on an existing MRF formulation incorporating a prior shape model and colour distributions for the constituent parts. We propose a novel shape model consisting of a deformable spatial prior probability for the part-label at each pixel. We also make a simple extension to the MRF formulation to deal simultaneously with multiple objects within a global optimisation. Finally, we evaluate the method for the task of segmenting individual items of clothing in images depicting groups of people, and demonstrate improved performance against the state of the art for this task.
Basela Hasan, David C. Hogg
BMVC2
2010 Event Model Learning from Complex Videos using ILP
Krishna Sandeep Reddy Dubba, Anthony G. Cohn 0001, David C. Hogg
ECAI3
2010 Discovering an Event Taxonomy from Video using Qualitative Spatio-temporal Graphs
abstract
This work proposes a graph mining based approach to mine a taxonomy of events from activities for complex videos which are represented in terms of qualitative spatio-temporal relationships. A Hidden Markov Model to obtain stable qualitative spatial relations from noisy measurements is introduced. The effectiveness of the approach is demonstrated through experimental results for a complex aircraft turnaround apron scenario.
Muralikrishna Sridhar, Anthony G. Cohn 0001, David C. Hogg
ECAI3
2009 Attribute Multiset Grammars for Global Explanations of Activities
abstract
Recognizing multiple interleaved activities in a video requires implicitly partitioning the detections for each activity. Furthermore, constraints between activities are impor- tant in finding valid explanations for all detections. We use Attribute Multiset Gram- mars (AMGs) as a formal representation for a domain’s knowledge to encode intra- and inter-activity constraints. We show how AMGs can be used to parse all the observa- tions into ‘feasible’ global explanations. We also present an algorithm for building a Bayesian network (BN) given an AMG and a set of detections. The set of labellings of the BN corresponds to the set of all possible parse trees. Finding the best explanation then amounts to finding the maximum a posteriori labeling of the BN. The technique is successfully applied to two different problems including the challenging problem of associating pedestrians and carried objects entering and departing a building.
Dima Damen, David C. Hogg
BMVC2
2009 Scene Modelling and Classification Using Learned Spatial Relations
Hannah M. Dee, David C. Hogg, Anthony G. Cohn 0001
COSIT2
2009 Recognizing linked events: Searching the space of feasible explanations
abstract
The ambiguity inherent in a localized analysis of events from video can be resolved by exploiting constraints between events and examining only feasible global explanations. We show how jointly recognizing and linking events can be formulated as labeling of a Bayesian network. The framework can be extended to multiple linking layers, expressing explanations as compositional hierarchies. The best global explanation is the maximum a posteriori (MAP) solution over a set of feasible explanations. The search space is sampled using reversible jump Markov chain Monte Carlo (RJMCMC). We propose a set of general move types that is extensible to multiple layers of linkage, and use simulated annealing to find the MAP solution given all observations. We provide experimental results for a challenging two-layer linkage problem, demonstrating the ability to recognise and link drop and pick events of bicycles in a rack over five days.
Dima Damen, David C. Hogg
CVPR2
2009 Navigational strategies in behaviour modelling
Hannah M. Dee, David C. Hogg
Artif. Intell.2
2008 Building artificial personalities - expressive communication channels based on an interlingua for a human-robot dance
John Bryden, David C. Hogg, Sita Popat, Mick Wallis
ALIFE2
2008 Learning Functional Object-Categories from a Relational Spatio-Temporal Representation
abstract
We propose a framework that learns functional object-categories from spatio-temporal data sets such as those abstracted from video. The data is represented as one activity graph that encodes qualitative spatio-temporal patterns of interaction between objects. Event classes are induced by statistical generalization, the instances of which encode similar patterns of spatio-temporal relationships between objects. Equivalence classes of objects are discovered on the basis of their similar role in multiple event instantiations. Objects are represented in a multidimensional space that captures their role in all the events. Unsupervised learning in this space results in functional object-categories. Experiments in the domain of food preparation suggest that our techniques represent a significant step in unsupervised learning of functional object categories from spatio-temporal patterns of object interaction.
Muralikrishna Sridhar, Anthony G. Cohn 0001, David C. Hogg
ECAI3
2008 Detecting Carried Objects in Short Video Sequences
Dima Damen, David C. Hogg
ECCV (3)2
2008 Motion segmentation by consensus
abstract
We present a method for merging multiple partitions into a single partition, by minimising the ratio of pairwise agreements and contradictions between the equivalence relations corresponding to the partitions. The number of equivalence classes is determined automatically. This method is advantageous when merging segmentations obtained independently. We propose using this consensus approach to merge segmentations of features tracked on video. Each segmentation is obtained by clustering on the basis of mean velocity during a particular time interval.
Roberto Fraile, David C. Hogg, Anthony G. Cohn 0001
ICPR2
2008 Enhanced tracking and recognition of moving objects by reasoning about spatio-temporal continuity
Brandon Bennett, Derek R. Magee, Anthony G. Cohn 0001, David C. Hogg
Image Vis. Comput.4
2007 Associating People Dropping off and Picking up Objects
abstract
Several interesting monitoring applications concern people entering a prescribed area, where they deposit an object in their possession, or collect an object deposited earlier. One example arises in the use of bicycle racks. We propose a novel method for associating each person who deposits an object with the person who later collects it. Our main contribution is to deal with ambiguity in the visual data through the use of global constraints on what is possible. The method is evaluated on a set of practical experiments in a bicycle rack, and applied to online theft detection by comparing the colour profile of associated individuals. 1
Dima Damen, David C. Hogg
BMVC2
2005 On the feasibility of using a cognitive model to filter surveillance data
abstract
This paper describes a novel approach to the problem of automated visual surveillance. The authors have extended an existing algorithm which uses a cognitive model of navigation to explain behaviour in a surveillance setting. We then take this cognitive model and apply it to the problem of filtering surveillance data: typically, a surveillance or CCTV installation will have a limited number of operatives monitoring a large number of cameras. The proposed system filters upon inexplicability scores, on the grounds that those trajectories which we can explain in terms of simple goals are exactly those trajectories which are uninteresting: it is only those we cannot simply explain which are worth attending to. Initial results are promising, with over 50% of uninteresting trajectories being excluded.
Hannah M. Dee, David C. Hogg
AVSS2
2005 Protocols from perceptual observations
Chris J. Needham, Paulo E. Santos, Derek R. Magee, Vincent E. Devin, David C. Hogg, Anthony G. Cohn 0001
Artif. Intell.5
2004 Detecting inexplicable behaviour
abstract
This paper presents a novel approach to the detection of unusual or interesting events in videos involving certain types of intentional behaviour, such as pedestrian scenes. The approach is not based upon a statistical measure of typicality, but upon building an understanding of the way people navigate towards a goal. The activity of agents moving around within the scene is evaluated based upon whether the behaviour in question is consistent with a simple model of goal-directed behaviour and a model of those goals and obstacles known to be in the scene. The advantages of such an approach are multiple: it handles the presence of movable obstacles (for example, parked cars) with ease; trajectories which have never before been presented to the system can be classified as explicable; and the technique as a whole has a prima facie psychological plausibility. A system based upon these principles is demonstrated in two scenes: a car-park, and in a foyer scenario 1. 1
Hannah M. Dee, David C. Hogg
BMVC2
2004 Using Spatio-Temporal Continuity Constraints to Enhance Visual Tracking of Moving Objects
Brandon Bennett, Derek R. Magee, Anthony G. Cohn 0001, David C. Hogg
ECAI4
2004 Combining Multiple Answers for Learning Mathematical Structures from Visual Observation
Paulo E. Santos, Derek R. Magee, Anthony G. Cohn 0001, David C. Hogg
ECAI4
2003 Reactive memories: an interactive talking-head
Vincent E. Devin, David C. Hogg
Image Vis. Comput.2
2002 Modeling Interaction Using Learnt Qualitative Spatio-Temporal Relations and Variable Length Markov Models
Aphrodite Galata, Anthony G. Cohn 0001, Derek R. Magee, David C. Hogg
ECAI4
2002 Representation and synthesis of behaviour using Gaussian mixtures
Neil Johnson 0001, David C. Hogg
Image Vis. Comput.2
2001 Reactive Memories: An Interactive Talking-Head
abstract
We demonstrate a novel method for producing a synthetic talking head. The method is based on earlier work in which the behaviour of a synthetic individual is generated by reference to a probabilistic model of interactive behaviour within the visual domain - such models are learnt automatically from typical interactions. We extend this work into a combined visual and auditory domain and employ a state-of-the-art facial appearance model. The result is a synthetic talking head that responds appropriately and with correct timing to simple forms of greeting with variations in facial expression and intonation.
Vincent E. Devin, David C. Hogg
BMVC2
2001 Learning Variable-Length Markov Models of Behavior
Aphrodite Galata, Neil Johnson 0001, David C. Hogg
Comput. Vis. Image Underst.3
2000 Statistical Models of Object Interaction
Richard J. Morris, David C. Hogg
Int. J. Comput. Vis.2
2000 Constructing qualitative event models automatically from video input
Jonathan H. Fernyhough, Anthony G. Cohn 0001, David C. Hogg
Image Vis. Comput.3
1999 Learning Behaviour Models of Human Activities
abstract
In recent years there has been an increased interest in the modelling and recognition of human activities involving highly structured and semantically rich behaviour such as dance, aerobics, and sign language. A novel approach is presented for automatically acquiring stochastic models of the high-level structure of an activity without the assumption of any prior knowledge. The process involves temporal segmentation intoplausible atomic behaviour com-ponents and the use of variable length Markov models for the efficient rep-resentation of behaviours. Experimental results are presented which demon-strate the generation of realistic sample behaviours and evaluate the perfor-mance of models for long-term temporal prediction. 1
Aphrodite Galata, Neil Johnson 0001, David C. Hogg
BMVC3
1999 Hybrid Approach to the Construction of Triangulated 3D Models of Building Interiors
Erik Wolfart, Vítor Sequeira, Kia Ng, Stuart Butterfield, João G. M. Gonçalves, David C. Hogg
ICVS6
1998 The Acquisition and Use of Interaction Behavior Models
abstract
Providing a machine with the ability to learn and use models of natural interaction is a challenging and largely unaddressed problem. A framework is developed enabling both the acquisition of interaction behaviours from the observation of humans, and the use of the acquired behaviour models to simulate a plausible partner during interaction. Statistically based interaction behaviour models are acquired automatically from the observation of interacting humans. Interaction with a virtual human is achieved using the model together with a stochastic tracking algorithm. Experimental results demonstrate the generation and use of the model for a simple human interaction.
Neil Johnson 0001, Aphrodite Galata, David C. Hogg
CVPR3
1998 Building Qualitative Event Models Automatically from Visual Input
abstract
We describe an implemented technique for generating event models automatically based on qualitative reasoning and a statistical analysis of video input. Using an existing tracking program which generates labelled contours for objects in every frame, the view from a fixed camera is partitioned into semantically relevant regions based on the paths followed by moving objects. The paths are indexed with temporal information so objects moving along the same path at different speeds can be distinguished. Using a notion of proximity based on the speed of the moving objects and qualitative spatial reasoning techniques, event models describing the behaviour of pairs of objects can be built, again using statistical methods. The system has been tested on a traffic domain and learns various event models expressed in the qualitative calculus which represent human observable events. The system can then be used to recognise subsequent selected event occurrences or unusual behaviours.
Jonathan H. Fernyhough, Anthony G. Cohn 0001, David C. Hogg
ICCV3
1998 Wormholes in Shape Space: Tracking Through Discontinuous Changes in Shape
abstract
Existing object tracking algorithms generally use some form of local optimisation, assuming that an object's position and shape change smoothly over time. In some situations this assumption is not valid: the track able shape of an object may change discontinuously, for example if it is the 2D silhouette of a 3D object. In this paper we propose a novel method for modelling temporal shape discontinuities explicitly. Allowable shapes are represented as a union of (learned) bounded regions within a shape space. Discontinuous shape changes are described in terms of transitions between these regions. Transition probabilities are learned from training sequences and stored in a Markov model. In this way we can create 'wormholes' in shape space. Tracking with such models is via an adaptation, of the CONDENSATION algorithm.
Tony Heap, David C. Hogg
ICCV2
1997 Improving Specificity in PDMs using a Hierarchical Approach
Tony Heap, David C. Hogg
BMVC2
1997 An Integrated Traffic and Pedestrian Model-Based Vision System
Paolo Remagnino, Adam Baumberg, T. Grove, David C. Hogg, Tieniu Tan, Anthony D. Worrall, Keith D. Baker
BMVC4
1997 Neural networks in human motion tracking - An experimental study
Li-Qun Xu, David C. Hogg
Image Vis. Comput.2
1996 Neural Networks in Human Motion Tracking - An Experimental Study
abstract
A novel application of neural networks is proposed for tracking the motion of a walking pedestrian. First, the motion is summarized by trajectories consisting of sequences of state vectors, each vector defining a 2D shape contour as well as the position of the pedestrian in image coordinates. Next, the task of tracking the motion is conducted in the context of multivariate time series prediction on the motion trajectories, in which neural networks are employed to model and generalize the spatial-temporal variations underlying the motion. The experiments have shown that the proposed system is capable of tracking eight typical movements of a walking pedestrian, even though the motion sequences are highly noisy and each motion trajectory is very short in time.
Li-Qun Xu, David C. Hogg
BMVC2
1996 Generation of Semantic Regions from Image Sequences
Jonathan H. Fernyhough, Anthony G. Cohn 0001, David C. Hogg
ECCV (2)3
1996 Global Alignment of MR Images Using a Scale Based Hierarchical Model
S. Fletcher, Andrew J. Bulpitt, David C. Hogg
ECCV (2)3
1996 Towards 3D hand tracking using a deformable model
abstract
In this paper we first describe how we have constructed a 3D deformable Point Distribution Model of the human hand, capturing training data semi-automatically from volume images via a physically-based model. We then show how we have attempted to use this model in tracking an unmarked hand moving with 6 degrees of freedom (plus deformation) in real time using a single video camera. In the course of this we show how to improve on a weighted least-squares pose parameter approximation at little computational cost. We note the successes and shortcomings of our system and discuss how it might be improved.
Tony Heap, David C. Hogg
FG2
1996 Generating spatiotemporal models from examples
Adam Baumberg, David C. Hogg
Image Vis. Comput.2
1996 Extending the Point Distribution Model using polar coordinates
Tony Heap, David C. Hogg
Image Vis. Comput.2
1996 Learning the distribution of object trajectories for event recognition
Neil Johnson 0001, David C. Hogg
Image Vis. Comput.2
1995 An Adaptive Eigenshape Model
abstract
There has been a great deal of recent interest in statistical models of 2D landmark data for generating compact deformable models of a given object. This paper extends this work to a class of parametrised shapes where there are no landmarks available. A rigorous statistical framework for the eigenshape model is introduced, which is an extension to the conventional Linear Point Distribution Model. One of the problems associated with landmark free methods is that a large degree of variability in any shape descriptor may be due to the choice of parametrisation. An automated training method is described which utilises an iterative feedback method to overcome this problem. The result is an automatically generated compact linear shape model. The model has been successfully applied to a problem of tracking the outline of a walking pedestrian in real time. 1 Introduction Statistical Analysis of 2D landmark data has become a well established tool in computer vision (e.g. morpholog...
Adam Baumberg, David C. Hogg
BMVC2
1995 Generating Spatiotemporal Models from Examples
Adam Baumberg, David C. Hogg
BMVC2
1995 Automated Pivot Location for the Cartesian-Polar Hybrid Point Distribution Model
abstract
The Point Distribution Model (PDM) has already proved useful for many tasks involving the location or tracking of deformable objects. A principal limitation is that non-linear variations must be approximated by combining linear variations, which sometimes results in a non-optimal model producing implausible object shapes. The Cartesian-Polar Hybrid PDM helps to overcome this limitation; selective use of polar geometry allows bending or pivotal deformation to be modelled more accurately; model components which exhibit no such trend remain in the Cartesian domain. Use of the Hybrid PDM currently requires the identification of pivot points by hand. In this paper we present a method for automatically identifying rigid model parts and pivot points from the training data. Experimental results are given for real data from human hands and for synthetic data from a simple jointed object. Keywords: Deformable models, Point Distribution Model, polar coordinates. 1 Introduction Models are used wi...
Tony Heap, David C. Hogg
BMVC2
1995 Learning the Distribution of Object Trajectories for Event Recognition
Neil Johnson 0001, David C. Hogg
BMVC2
1995 Extending the Point Distribution Model Using Polar Coordinates
Tony Heap, David C. Hogg
CAIP2
1995 Generic 3-D Shape Model: Acquisitions and Applications
Xinquan Shen, David C. Hogg
CAIP2
1995 3D shape recovery using a deformable model
Xinquan Shen, David C. Hogg
Image Vis. Comput.2
1994 3-D Shape Recovery Using A Deformable Model
abstract
This paper describes a method for recovering the 3D shape of a moving object from a sequence of images. While following the motion of the object, a 3D surface model, initialized to be spherical, progressively deforms under the action of simulated external forces driving its profile towards the object profile extracted from the image. Internal forces coupled to the model encourage the surface to be smooth and to deform smoothly at each time step. The model is aligned with the object in 3D by rotating on the ground plane to orientate parallel to the back-projected motion trajectory and translating to minimize the distance between the model profile and the object profile. Experimental results are presented for the recovery of shape models of vehicles
Xinquan Shen, David C. Hogg
BMVC2
1994 Learning Flexible Models from Image Sequences
Adam Baumberg, David C. Hogg
ECCV (1)2
1994 Shape Models from Image Sequences
Xinquan Shen, David C. Hogg
ECCV (1)2
1993 Special issue: British Machine Vision Conference 1992
David C. Hogg
Image Vis. Comput.1
1993 Shape in machine vision
David C. Hogg
Image Vis. Comput.1
1992 Building a Model of a Road Junction Using Moving Vehicle Information
Li-Qun Xu, David S. Young, David C. Hogg
BMVC3
1988 The use of digital map data in the segmentation and classification of remotely-sensed images
abstract
Remotely-sensed data constitute a major potential source of input to geographical information systems (GIS)However, these data often have a relatively poor classification accuracy compared with that of the cartographic data from maps with which they may be combined in the course of GIS analysis. The possibility exists of using data sets (in the form of digital maps) resident within a GIS in order to improve this accuracy, before the classified image is incorporated into the GIS. Results are discussed from a British Alvey Information Technology project to develop a system for the knowledge-based segmentation and classification of remotely-sensed terrain images, in which the knowledge contained in digital map
David C. Mason, D. G. Corr, Alan Cross, David C. Hogg, D. H. Lawrence, Maria Petrou, Anita Tailor
Int. J. Geogr. Inf. Sci.4
1986 Knowledge-based interpretation of remotely sensed images
Anita Tailor, Alan Cross, David C. Hogg, David C. Mason
Image Vis. Comput.3
1983 Model-based vision: a program to see a walking person
David C. Hogg
Image Vis. Comput.1
1977 A Methodology for Real Time Scene Analysis
David C. Hogg
IJCAI1