VLDB 2026 Research / reviewers in the wild / expert
Andrew Gilbert
dblp:63/3257
· DBLP profile ↗
42ranked-venue papers
18as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 15 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 12 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generative Data Augmentation for Skeleton Action Recognition
Wanqing Li 0001, Anthony Adeyemi-Ejeye, Andrew Gilbert |
FG | 4 |
| 2026 | UIL-AQA: Uncertainty-Aware Clip-Level Interpretable Action Quality AssessmentabstractAbstract This work proposes UIL-AQA for long-term Action Quality Assessment AQA designed to be clip-level interpretable and uncertainty-aware. AQA evaluates the execution quality of actions in videos. However, the complexity and diversity of actions, especially in long videos, increase the difficulty of AQA. Existing AQA methods solve this by limiting themselves generally to short-term videos. These approaches lack detailed semantic interpretation for individual clips and fail to account for the impact of human biases and subjectivity in the data during model training. Moreover, although query-based Transformer networks demonstrate strong capabilities in long-term modelling, their interpretability in AQA remains insufficient. This is primarily due to a phenomenon we identified, termed Temporal Skipping , where the model skips self-attention layers to prevent output degradation. We introduce an Attention Loss function and a Query Initialization Module to enhance the modelling capability of query-based Transformer networks. Additionally, we incorporate a Gaussian Noise Injection Module to simulate biases in human scoring, mitigating the influence of uncertainty and improving model reliability. Furthermore, we propose a Difficulty-Quality Regression Module, which decomposes each clip’s action score into independent difficulty and quality components, enabling a more fine-grained and interpretable evaluation. Our extensive quantitative and qualitative analysis demonstrates that our proposed method achieves state-of-the-art performance on three long-term real-world AQA datasets. Our code is available at: https://github.com/dx199771/Interpretability-AQA Wanqing Li 0001, Anthony Adeyemi-Ejeye, Andrew Gilbert |
Int. J. Comput. Vis. | 5 |
| 2025 | Multitwine: Multi-Object Compositing with Text and Layout ControlabstractWe introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex actions requiring reposing (e.g., hugging, playing guitar). When an interaction implies additional props, like ‘taking a selfie’, our model autonomously generates these supporting objects. By jointly training for compositing and subject-driven generation, also known as customization, we achieve a more balanced integration of textual and visual inputs for text-driven object compositing. As a result, we obtain a versatile model with state-of-the-art performance in both tasks. We further present a data generation pipeline leveraging visual and language models to effortlessly synthesize multimodal, aligned training data. Gemma Canet Tarrés, Zhe Lin 0001, He Zhang 0004, Andrew Gilbert, John P. Collomosse, Soo Ye Kim |
CVPR | 5 |
| 2024 | FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
Mona Ahmadian, Frank Guerin, Andrew Gilbert |
BMVC | 3 |
| 2024 | Interpretable Long-term Action Quality Assessment
Wanqing Li 0001, Anthony Adeyemi-Ejeye, Andrew Gilbert |
BMVC | 5 |
| 2024 | Thinking Outside the BBox: Unconstrained Generative Object Compositing
Gemma Canet Tarrés, Zhe Lin 0001, Jianming Zhang 0001, Dan Ruta, Andrew Gilbert, John P. Collomosse, Soo Ye Kim |
ECCV (62) | 7 |
| 2022 | KPE: Keypoint Pose Encoding for Transformer-based Image Generation
Soon Yau Cheong, Armin Mustafa, Andrew Gilbert |
BMVC | 3 |
| 2022 | Two-Stream Transformer Architecture for Long Form Video Understanding
Edward Fish, Jon Weinbren, Andrew Gilbert |
BMVC | 3 |
| 2022 | SVS: Adversarial refinement for sparse novel view synthesis
Violeta Menéndez González, Andrew Gilbert, Graeme Phillipson, Stephen Jolly, Simon Hadfield |
BMVC | 2 |
| 2022 | StyleBabel: Artistic Style Tagging and CaptioningabstractWe present StyleBabel, a unique open access dataset of natural language captions and free-form tags describing the artistic style of over 135K digital artworks, collected via a novel participatory method from experts studying at specialist art and design schools. StyleBabel was collected via an iterative method, inspired by ‘Grounded Theory’: a qualitative approach that enables annotation while co-evolving a shared language for fine-grained artistic style attribute description. We demonstrate several downstream tasks for StyleBabel, adapting the recent ALADIN architecture for fine-grained style similarity, to train cross-modal embeddings for: 1) free-form tag generation; 2) natural language description of artistic style; 3) fine-grained text search of style. To do so, we extend ALADIN with recent advances in Visual Transformer (ViT) and cross-modal representation learning, achieving a state of the art accuracy in fine-grained style retrieval. Dan Ruta, Andrew Gilbert, Pranav Aggarwal, Naveen Marri, Ajinkya Kale, Jo Briggs, Chris Speed, Hailin Jin, Baldo Faieta, Alex Filipkowski, Zhe Lin 0001, John P. Collomosse |
ECCV (8) | 2 |
| 2022 | Light-weight Spatio-Temporal Graphs for Segmentation and Ejection Fraction Prediction in Cardiac Ultrasound
Sarina Thomas, Andrew Gilbert, Guy Ben-Yosef |
MICCAI (4) | 2 |
| 2021 | ALADIN: All Layer Adaptive Instance Normalization for Fine-grained Style SimilarityabstractWe present ALADIN (All Layer AdaIN); a novel architecture for searching images based on the similarity of their artistic style. Representation learning is critical to visual search, where distance in the learned search embedding reflects image similarity. Learning an embedding that discriminates fine-grained variations in style is hard, due to the difficulty of defining and labelling style. ALADIN takes a weakly supervised approach to learning a representation for fine-grained style similarity of digital artworks, leveraging BAM-FG, a novel large-scale dataset of user generated content groupings gathered from the web. ALADIN sets a new state of the art accuracy for style-based visual search over both coarse labelled style data (BAM) and BAM-FG; a new 2.62 million image dataset of 310,000 fine-grained style groupings also contributed by this work. Dan Ruta, Saeid Motiian, Baldo Faieta, Zhe Lin 0001, Hailin Jin, Alex Filipkowski, Andrew Gilbert, John P. Collomosse |
ICCV | 7 |
| 2021 | Rethinking Genre Classification With Fine Grained Semantic ClusteringabstractMovie genre classification is an active research area in machine learning; however, the content of movies can vary widely within a single genre label. We expand these ‘coarse’ genre labels by identifying ‘fine-grained’ contextual relationships within the multi-modal content of videos. By leveraging pre-trained ‘expert’ networks, we learn the influence of different combinations of modes for multi-label genre classification. Then, we continue to fine-tune this ‘coarse’ genre classification network self-supervised to sub-divide the genres based on the multi-modal content of the videos. Our approach is demonstrated on a new multi-moda137,866,450 frame, 8,800 movie trailer dataset, MMX-Trailer-20, which includes pre-computed audio, location, motion, and image embeddings. Edward Fish, Jon Weinbren, Andrew Gilbert |
ICIP | 3 |
| 2021 | Neural architecture search for deep image prior
Kary Ho, Andrew Gilbert, Hailin Jin, John P. Collomosse |
Comput. Graph. | 2 |
| 2021 | User-Intended Doppler Measurement Type Prediction Combining CNNs With Smart Post-ProcessingabstractSpectral Doppler measurements are an important part of the standard echocardiographic examination. These measurements give insight into myocardial motion and blood flow, providing clinicians with parameters for diagnostic decision making. Many of these measurements are performed automatically with high accuracy, increasing the efficiency of the diagnostic pipeline. However, full automation is not yet available because the user must manually select which measurement should be performed on each image. In this work, we develop a pipeline based on convolutional neural networks (CNNs) to automatically classify the measurement type from cardiac Doppler scans. We show how the multi-modal information in each spectral Doppler recording can be combined using a meta parameter post-processing mapping scheme and heatmaps to encode coordinate locations. Additionally, we experiment with several architectures to examine the tradeoff between accuracy, speed, and memory usage for resource-constrained environments. Finally, we propose a confidence metric using the values in the last fully connected layer of the network and show that our confidence metric can prevent many misclassifications. Our algorithm enables a fully automatic pipeline from acquisition to Doppler spectrum measurements. We achieve 96% accuracy on a test set drawn from separate clinical sites, indicating that the proposed method is suitable for clinical adoption. Andrew Gilbert, Marit Holden, Line Eikvil, Mariia Rakhmail, Aleksandar Babic, Svein Arne Aase, Eigil Samset, Kristin McLeod |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Generating Synthetic Labeled Data From Existing Anatomical Models: An Example With Echocardiography SegmentationabstractDeep learning can bring time savings and increased reproducibility to medical image analysis. However, acquiring training data is challenging due to the time-intensive nature of labeling and high inter-observer variability in annotations. Rather than labeling images, in this work we propose an alternative pipeline where images are generated from existing high-quality annotations using generative adversarial networks (GANs). Annotations are derived automatically from previously built anatomical models and are transformed into realistic synthetic ultrasound images with paired labels using a CycleGAN. We demonstrate the pipeline by generating synthetic 2D echocardiography images to compare with existing deep learning ultrasound segmentation datasets. A convolutional neural network is trained to segment the left ventricle and left atrium using only synthetic images. Networks trained with synthetic images were extensively tested on four different unseen datasets of real images with median Dice scores of 91, 90, 88, and 87 for left ventricle segmentation. These results match or are better than inter-observer results measured on real ultrasound datasets and are comparable to a network trained on a separate set of real images. Results demonstrate the images produced can effectively be used in place of real data for training. The proposed pipeline opens the door for automatic generation of training data for many tasks in medical imaging as the same process can be applied to other segmentation or landmark detection tasks in any modality. The source code and anatomical models are available to other researchers.11https://adgilbert.github.io/data-generation/ Andrew Gilbert, Maciej Marciniak, Cristóbal Rodero, Pablo Lamata, Eigil Samset, Kristin McLeod |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Semantic Estimation of 3D Body Shape and Pose using Minimal Cameras
Andrew Gilbert, Matthew Trumble, Adrian Hilton 0001, John P. Collomosse |
BMVC | 1 |
| 2020 | Inpainting of Wide-Baseline Multiple Viewpoint VideoabstractWe describe a non-parametric algorithm for multiple-viewpoint video inpainting. Uniquely, our algorithm addresses the domain of wide baseline multiple-viewpoint video (MVV) with no temporal look-ahead in near real time speed. A Dictionary of Patches (DoP) is built using multi-resolution texture patches reprojected from geometric proxies available in the alternate views. We dynamically update the DoP over time, and a Markov Random Field optimisation over depth and appearance is used to resolve and align a selection of multiple candidates for a given patch, this ensures the inpainting of large regions in a plausible manner conserving both spatial and temporal coherence. We demonstrate the removal of large objects (e.g., people) on challenging indoor and outdoor MVV exhibiting cluttered, dynamic backgrounds and moving cameras. Andrew Gilbert, Matthew Trumble, Adrian Hilton 0001, John P. Collomosse |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Fusing Visual and Inertial Sensors with Semantics for 3D Human Pose EstimationabstractWe propose an approach to accurately estimate 3D human pose by fusing multi-viewpoint video (MVV) with inertial measurement unit (IMU) sensor data, without optical markers, a complex hardware setup or a full body model. Uniquely we use a multi-channel 3D convolutional neural network to learn a pose embedding from visual occupancy and semantic 2D pose estimates from the MVV in a discretised volumetric probabilistic visual hull. The learnt pose stream is concurrently processed with a forward kinematic solve of the IMU data and a temporal model (LSTM) exploits the rich spatial and temporal long range dependencies among the solved joints, the two streams are then fused in a final fully connected layer. The two complementary data sources allow for ambiguities to be resolved within each sensor modality, yielding improved accuracy over prior methods. Extensive evaluation is performed with state of the art performance reported on the popular Human 3.6M dataset (Ionescu et al. in Intell IEEE Trans Pattern Anal Mach 36(7):1325–1339, 2014), the newly released TotalCapture dataset and a challenging set of outdoor videos TotalCaptureOutdoor. We release the new hybrid MVV dataset (TotalCapture) comprising of multi-viewpoint video, IMU and accurate 3D skeletal joint ground truth derived from a commercial motion capture system. The dataset is available online at http://cvssp.org/data/totalcapture/ . Andrew Gilbert, Matthew Trumble, Charles Malleson, Adrian Hilton 0001, John P. Collomosse |
Int. J. Comput. Vis. | 1 |
| 2018 | Disentangling Structure and Aesthetics for Style-Aware Image CompletionabstractContent-aware image completion or in-painting is a fundamental tool for the correction of defects or removal of objects in images. We propose a non-parametric in-painting algorithm that enforces both structural and aesthetic (style) consistency within the resulting image. Our contributions are two-fold: (1) we explicitly disentangle image structure and style during patch search and selection to ensure a visually consistent look and feel within the target image. (2) we perform adaptive stylization of patches to conform the aesthetics of selected patches to the target image, so harmonizing the integration of selected patches into the final composition. We show that explicit consideration of visual style during in-painting delivers excellent qualitative and quantitative results across the varied image styles and content, over the Places2 scene photographic dataset and a challenging new in-painting dataset of artwork derived from BAM! Andrew Gilbert, John P. Collomosse, Hailin Jin, Brian L. Price |
CVPR | 1 |
| 2018 | Volumetric Performance Capture from Minimal Camera Viewpoints
Andrew Gilbert, Marco Volino, John P. Collomosse, Adrian Hilton 0001 |
ECCV (11) | 1 |
| 2018 | Deep Autoencoder for Combined Human Pose Estimation and Body Model Upscaling
Matthew Trumble, Andrew Gilbert, Adrian Hilton 0001, John P. Collomosse |
ECCV (10) | 2 |
| 2017 | Real-Time Full-Body Motion Capture from Video and IMUsabstractA real-time full-body motion capture system is presented which uses input from a sparse set of inertial measurement units (IMUs) along with images from two or more standard video cameras and requires no optical markers or specialized infra-red cameras. A real-time optimization-based framework is proposed which incorporates constraints from the IMUs, cameras and a prior pose model. The combination of video and IMU data allows the full 6-DOF motion to be recovered including axial rotation of limbs and drift-free global position. The approach was tested using both indoor and outdoor captured data. The results demonstrate the effectiveness of the approach for tracking a wide range of human motion in real time in unconstrained indoor/outdoor scenes. Charles Malleson, Andrew Gilbert, Matthew Trumble, John P. Collomosse, Adrian Hilton 0001, Marco Volino |
3DV | 2 |
| 2017 | Total Capture: 3D Human Pose Estimation Fusing Video and Inertial Sensors
Matthew Trumble, Andrew Gilbert, Charles Malleson, Adrian Hilton 0001, John P. Collomosse |
BMVC | 2 |
| 2017 | A population-level approach to temperature robustness in neuromorphic systemsabstractWe present a novel approach to achieving temperature-robust behavior in neuromorphic systems that operates at the population level, trading an increase in silicon-neuron count for robustness across temperature. Our silicon neurons' tuning curves were highly sensitive to temperature, which could be decoded from a 400-neuron population with a precision of 0.07° C. We overcame this temperature-sensitivity by combining methods from robust optimization theory with the Neural Engineering Framework. We developed two algorithms and compared their temperature-robustness across a range of 2° C by decoding one period of a sinusoid-like function from populations with 25 to 800 neurons. We find that 560 neurons are required to achieve the same precision across this temperature range as 35 neurons achieved at a single temperature. Eric Kauderer-Abrams, Andrew Gilbert, Aaron Voelker, Ben Varkey Benjamin, Terrence C. Stewart, Kwabena Boahen 0001 |
ISCAS | 2 |
| 2017 | Image and video mining through online learning
Andrew Gilbert, Richard Bowden |
Comput. Vis. Image Underst. | 1 |
| 2017 | Guided optimisation through classification and regression for hand pose estimation
Philip Krejov, Andrew Gilbert, Richard Bowden |
Comput. Vis. Image Underst. | 2 |
| 2014 | Data Mining for Action Recognition
Andrew Gilbert, Richard Bowden |
ACCV (5) | 1 |
| 2014 | Capturing relative motion and finding modes for action recognition in the wild
Olusegun Oshin, Andrew Gilbert, Richard Bowden |
Comput. Vis. Image Underst. | 2 |
| 2012 | A Picture Is Worth a Thousand Tags: Automatic Web Based Image Tag Expansion
Andrew Gilbert, Richard Bowden |
ACCV (2) | 1 |
| 2012 | Meeting in the Middle: A top-down and bottom-up approach to detect pedestrians
Affan Shaukat, Andrew Gilbert, David Windridge, Richard Bowden |
ICPR | 2 |
| 2011 | Push and Pull: Iterative grouping of mediaabstractWe present an approach to iteratively cluster images and video in an efficient and intuitive manor. While many techniques use the traditional approach of time consuming groundtruthing large amounts of data [10, 16, 20, 23], this is increasingly infeasible as dataset size and complexity increase. Furthermore it is not applicable to the home user, who wants to intuitively group his/her own media without labelling the content. Instead we propose a solution that allows the user to select media that semantically belongs to the same class and use machine learning to this and other related content together. We introduce an signature descriptor and use min-Hash and greedy clustering to efficiently present the user with clusters of the dataset using multi-dimensional scaling. The image signatures of the dataset are then adjusted by APriori data mining identifying the common elements between a small subset of image signatures. This is able to both pull together true positive clusters and push apart false positive examples. The approach is tested on real videos harvested from the web using the state of the art YouTube dataset [18]. The accuracy of correct group label increases from 60.4% to 81.7% using 15 iterations of pulling and pushing the media around. While the process takes only 1 minute to compute the pair wise similarities of the image signatures and visualise the youtube whole dataset. © 2011. The copyright of this document resides with its authors. Andrew Gilbert, Richard Bowden |
BMVC | 1 |
| 2011 | Visualisation and prediction of conversation interest through mined social signalsabstractThis paper introduces a novel approach to social behaviour recognition governed by the exchange of non-verbal cues between people. We conduct experiments to try and deduce distinct rules that dictate the social dynamics of people in a conversation, and utilise semi-supervised computer vision techniques to extract their social signals such as laughing and nodding. Data mining is used to deduce frequently occurring patterns of social trends between a speaker and listener in both interested and not interested social scenarios. The confidence values from rules are utilised to build a Social Dynamic Model (SDM), that can then be used for classification and visualisation. By visualising the rules generated in the SDM, we can analyse distinct social trends between an interested and not interested listener in a conversation. Results show that these distinctions can be applied generally and used to accurately predict conversational interest. Dumebi Okwechime, Eng-Jon Ong, Andrew Gilbert, Richard Bowden |
FG | 3 |
| 2011 | Capturing the relative distribution of features for action recognitionabstractThis paper presents an approach to the categorisation of spatio-temporal activity in video, which is based solely on the relative distribution of feature points. Introducing a Relative Motion Descriptor for actions in video, we show that the spatio-temporal distribution of features alone (without explicit appearance information) effectively describes actions, and demonstrate performance consistent with state-of-the-art. Furthermore, we propose that for actions where noisy examples exist, it is not optimal to group all action examples as a single class. Therefore, rather than engineering features that attempt to generalise over noisy examples, our method follows a different approach: We make use of Random Sampling Consensus (RANSAC) to automatically discover and reject outlier examples within classes. We evaluate the Relative Motion Descriptor and outlier rejection approaches on four action datasets, and show that outlier rejection using RANSAC provides a consistent and notable increase in performance, and demonstrate superior performance to more complex multiple-feature based approaches. Olusegun Oshin, Andrew Gilbert, Richard Bowden |
FG | 2 |
| 2011 | iGroup: Weakly supervised image and video groupingabstractWe present a generic, efficient and iterative algorithm for interactively clustering classes of images and videos. The approach moves away from the use of large hand labelled training datasets, instead allowing the user to find natural groups of similar content based upon a handful of “seed” examples. Two efficient data mining tools originally developed for text analysis; min-Hash and APriori are used and extended to achieve both speed and scalability on large image and video datasets. Inspired by the Bag-of-Words (BoW) architecture, the idea of an image signature is introduced as a simple descriptor on which nearest neighbour classification can be performed. The image signature is then dynamically expanded to identify common features amongst samples of the same class. The iterative approach uses APriori to identify common and distinctive elements of a small set of labelled true and false positive signatures. These elements are then accentuated in the signature to increase similarity between examples and “pull” positive classes together. By repeating this process, the accuracy of similarity increases dramatically despite only a few training examples, only 10% of the labelled groundtruth is needed, compared to other approaches. It is tested on two image datasets including the caltech101 [9] dataset and on three state-of-the-art action recognition datasets. On the YouTube [18] video dataset the accuracy increases from 72% to 97% using only 44 labelled examples from a dataset of over 1200 videos. The approach is both scalable and efficient, with an iteration on the full YouTube dataset taking around 1 minute on a standard desktop machine. Andrew Gilbert, Richard Bowden |
ICCV | 1 |
| 2011 | Action Recognition Using Mined Hierarchical Compound FeaturesabstractThe field of Action Recognition has seen a large increase in activity in recent years. Much of the progress has been through incorporating ideas from single-frame object recognition and adapting them for temporal-based action recognition. Inspired by the success of interest points in the 2D spatial domain, their 3D (space-time) counterparts typically form the basic components used to describe actions, and in action recognition the features used are often engineered to fire sparsely. This is to ensure that the problem is tractable; however, this can sacrifice recognition accuracy as it cannot be assumed that the optimum features in terms of class discrimination are obtained from this approach. In contrast, we propose to initially use an overcomplete set of simple 2D corners in both space and time. These are grouped spatially and temporally using a hierarchical process, with an increasing search area. At each stage of the hierarchy, the most distinctive and descriptive features are learned efficiently through data mining. This allows large amounts of data to be searched for frequently reoccurring patterns of features. At each level of the hierarchy, the mined compound features become more complex, discriminative, and sparse. This results in fast, accurate recognition with real-time performance on high-resolution video. As the compound features are constructed and selected based upon their ability to discriminate, their speed and accuracy increase at each level of the hierarchy. The approach is tested on four state-of-the-art data sets, the popular KTH data set to provide a comparison with other state-of-the-art approaches, the Multi-KTH data set to illustrate performance at simultaneous multiaction classification, despite no explicit localization information provided during training. Finally, the recent Hollywood and Hollywood2 data sets provide challenging complex actions taken from commercial movie sequences. For all four data sets, the proposed hierarchical approach outperforms all other methods reported thus far in the literature and can achieve real-time operation. Andrew Gilbert, John Illingworth, Richard Bowden |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Social Interactive Human Video Synthesis
Dumebi Okwechime, Eng-Jon Ong, Andrew Gilbert, Richard Bowden |
ACCV (1) | 3 |
| 2009 | Fast realistic multi-action recognition using mined dense spatio-temporal featuresabstractWithin the field of action recognition, features and descriptors are often engineered to be sparse and invariant to transformation. While sparsity makes the problem tractable, it is not necessarily optimal in terms of class separability and classification. This paper proposes a novel approach that uses very dense corner features that are spatially and temporally grouped in a hierarchical process to produce an overcomplete compound feature set. Frequently reoccurring patterns of features are then found through data mining, designed for use with large data sets. The novel use of the hierarchical classifier allows real time operation while the approach is demonstrated to handle camera motion, scale, human appearance variations, occlusions and background clutter. The performance of classification, outperforms other state-of-the-art action recognition algorithms on the three datasets; KTH, multi-KTH, and Hollywood. Multiple action localisation is performed, though no groundtruth localisation data is required, using only weak supervision of class labels for each training sequence. The Hollywood dataset contain complex realistic actions from movies, the approach outperforms the published accuracy on this dataset and also achieves real time performance. Andrew Gilbert, John Illingworth, Richard Bowden |
ICCV | 1 |
| 2008 | Scale Invariant Action Recognition Using Compound Features Mined from Dense Spatio-temporal Corners
Andrew Gilbert, John Illingworth, Richard Bowden |
ECCV (1) | 1 |
| 2008 | Incremental, scalable tracking of objects inter camera
Andrew Gilbert, Richard Bowden |
Comput. Vis. Image Underst. | 1 |
| 2006 | Tracking Objects Across Cameras by Incrementally Learning Inter-camera Colour Calibration and Patterns of Activity
Andrew Gilbert, Richard Bowden |
ECCV (2) | 1 |
| 2005 | Incremental Modelling of the Posterior Distribution of Objects for Inter and Intra Camera TrackingabstractThis paper presents a scalable solution to the problem of tracking objects across spatially separated, uncalibrated, non-overlapping cameras. Unlike other approaches this technique uses an incremental learning method to create the spatio-temporal links between cameras, and thus model the posterior probability distribution of these links. This can then be used with an appearance model of the object to track across cameras. It requires no calibration or batch preprocessing and becomes more accurate over time as evidence is accumulated. 1 Andrew Gilbert, Richard Bowden |
BMVC | 1 |