VLDB 2026 Research / reviewers in the wild / expert
James Charles
dblp:91/7705
· DBLP profile ↗
20ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Applying a Quantum Annealer to the Traffic Assignment ProblemabstractThe Traffic Assignment Problem (TAP) is a complex transportation optimisation problem typically solved using meta-heuristics on classical computers. Quantum computers, despite being a nascent technology, have the potential to significantly speed up computation by exploiting quantum parallelism. A quantum annealer (QA) is a quantum computer tailored to solve combinatorial optimisation problems formulated as a Quadratic Unconstrained Binary Optimisation (QUBO). Formulating complex optimisation problems as QUBO is an open challenge. This paper derives a new QUBO formulation for TAP by employing a streamlined methodology of general applicability. It also attempts a direct comparison at solving TAP encompassing a QA (D-WAVE), a hybrid quantum-classical algorithm, and classical methods including Simulated Annealing and Genetic Algorithms. This comparison is difficult and seldom done due to the inherent differences between quantum and classic hardware. As expected from the current quantum technology, our results show that a pure QA suffers from significant noise in qubits and requires significant additional computational time, although we show that the time required solely by the QPU does not increase with problem size. We also show that the hybrid QA mitigates these noise issues and is on a par with traditional methods. Darren M. Chitty, James Charles, Alberto Moraglio, Ed Keedwell |
GECCO | 2 |
| 2024 | FOUND: Foot Optimization with Uncertain Normals for Surface Deformation Using Synthetic DataabstractSurface reconstruction from multi-view images is a challenging task, with solutions often requiring a large number of sampled images with high overlap. We seek to develop a method for few-view reconstruction, for the case of the human foot. To solve this task, we must extract rich geometric cues from RGB images, before carefully fusing them into a final 3D object. Our FOUND approach tackles this, with 4 main contributions: (i) SynFoot, a synthetic dataset of 50,000 photorealistic foot images, paired with ground truth surface normals and keypoints; (ii) an uncertainty-aware surface normal predictor trained on our synthetic dataset; (iii) an optimization scheme for fitting a generative foot model to a series of images; and (iv) a benchmark dataset of calibrated images and high resolution ground truth geometry. We show that our normal predictor outperforms all off-the-shelf equivalents significantly on real images, and our optimization scheme outperforms state-of-the-art photogrammetry pipelines, especially for a few-view setting. We release our synthetic dataset and baseline 3D scans to the research community. Oliver Boyne, Gwangbin Bae, James Charles, Roberto Cipolla |
WACV | 3 |
| 2022 | FIND: An Unsupervised Implicit 3D Model of Articulated Human Feet
Oliver Boyne, James Charles, Roberto Cipolla |
BMVC | 2 |
| 2022 | Style2NeRF: An Unsupervised One-Shot NeRF for Semantic 3D Reconstruction
James Charles, Wim Abbeloos, Daniel Olmeda Reino, Roberto Cipolla |
BMVC | 1 |
| 2022 | Discrete neural representations for explainable anomaly detectionabstractThe aim of this work is to detect and automatically generate high-level explanations of anomalous events in video. Understanding the cause of an anomalous event is crucial as the required response is dependant on its nature and severity. Recent works typically use object or action classifier to detect and provide labels for anomalous events. However, this constrains detection systems to a finite set of known classes and prevents generalisation to unknown objects or behaviours. Here we show how to robustly detect anomalies without the use of object or action classifiers yet still recover the high level reason behind the event. We make the following contributions: (1) a method using saliency maps to decouple the explanation of anomalous events from object and action classifiers, (2) show how to improve the quality of saliency maps using a novel neural architecture for learning discrete representations of video by predicting future frames and (3) beat the state-of-the-art anomaly explanation methods by 60% on a subset of the public benchmark X-MAN dataset [25]. Stanislaw Szymanowicz, James Charles, Roberto Cipolla |
WACV | 2 |
| 2021 | Scaling digital screen reading with one-shot learning and re-identificationabstractUsing only a mobile phone app, our objective is to cheaply retro-fit digital meters (e.g blood pressure, blood glucose or industrial gauges) with `smart' data transfer capabilities. Using the mobile phone camera we build an app to securely and accurately transcribe information from digital meter screens. Only a single labelled training image of a target meter is required to build a custom screen reading module. Here we show how this can scale to potentially hundreds of different meters by learning to recognising the meter type so that the reading module can be automatically selected. This makes the system very easy for a user who would need to scan multiple different meter types. To this end, we build a CNN based system which runs in real-time on mobile device with very high read accuracy and meter recognition. Our contributions include (i) a method of one-shot training by synthesis through domain shift reduction, (ii) a deep embedding network for scale, translation and rotation invariant re-identification of digital meters, (iii) a highly accurate and efficient mobile phone app for recognising and parsing digital meter screens and (iv) release of a new digital meter re-identification dataset. James Charles, Stefano Bucciarelli, Roberto Cipolla |
WACV | 1 |
| 2020 | FootNet: An Efficient Convolutional Network for Multiview 3D Foot Reconstruction
Felix Kok, James Charles, Roberto Cipolla |
ACCV (6) | 2 |
| 2020 | Real-time screen reading: reducing domain shift for one-shot learning
James Charles, Stefano Bucciarelli, Roberto Cipolla |
BMVC | 1 |
| 2020 | Who Left the Dogs Out? 3D Animal Reconstruction with Expectation Maximization in the Loop
Benjamin Biggs, Oliver Boyne, James Charles, Andrew W. Fitzgibbon, Roberto Cipolla |
ECCV (11) | 3 |
| 2017 | Latent Dirichlet Allocation for Unsupervised Activity Analysis on an Autonomous Mobile RobotabstractFor autonomous robots to collaborate on joint tasks with humans they require a shared understanding of an observed scene. We present a method for unsupervised learning of common human movements and activities on an autonomous mobile robot, which generalises and improves on recent results. Our framework encodes multiple qualitative abstractions of RGBD video from human observations and does not require external temporal segmentation. Analogously to information retrieval in text corpora, each human detection is modelled as a random mixture of latent topics. A generative probabilistic technique is used to recover topic distributions over an auto-generated vocabulary of discrete, qualitative spatio-temporal code words. We show that the emergent categories align well with human activities as interpreted by a human. This is a particularly challenging task on a mobile robot due to the varying camera viewpoints which lead to incomplete, partial and occluded human detections. Paul Duckworth, Muhannad Al-Omari, James Charles, David C. Hogg, Anthony G. Cohn 0001 |
AAAI | 3 |
| 2017 | Real-time Factored ConvNets: Extracting the X Factor in Human Parsing
James Charles, Ignas Budvytis, Roberto Cipolla |
BMVC | 1 |
| 2016 | Personalizing Human Video Pose EstimationabstractWe propose a personalized ConvNet pose estimator that automatically adapts itself to the uniqueness of a person's appearance to improve pose estimation in long videos. We make the following contributions: (i) we show that given a few high-precision pose annotations, e.g. from a generic ConvNet pose estimator, additional annotations can be generated throughout the video using a combination of image-based matching for temporally distant frames, and dense optical flow for temporally local frames, (ii) we develop an occlusion aware self-evaluation model that is able to automatically select the high-quality and reject the erroneous additional annotations, and (iii) we demonstrate that these high-quality annotations can be used to fine-tune a ConvNet pose estimator and thereby personalize it to lock on to key discriminative features of the person's appearance. The outcome is a substantial improvement in the pose estimates for the target video using the personalized ConvNet compared to the original generic ConvNet. Our method outperforms the state of the art (including top ConvNet methods) by a large margin on three standard benchmarks, as well as on a new challenging YouTube video dataset. Furthermore, we show that training from the automatically generated annotations can be used to improve the performance of a generic ConvNet on other benchmarks. James Charles, Tomas Pfister, Derek R. Magee, David C. Hogg, Andrew Zisserman |
CVPR | 1 |
| 2015 | Flowing ConvNets for Human Pose Estimation in VideosabstractThe objective of this work is human pose estimation in videos, where multiple frames are available. We investigate a ConvNet architecture that is able to benefit from temporal context by combining information across the multiple frames using optical flow. To this end we propose a network architecture with the following novelties: (i) a deeper network than previously investigated for regressing heatmaps, (ii) spatial fusion layers that learn an implicit spatial model, (iii) optical flow is used to align heatmap predictions from neighbouring frames, and (iv) a final parametric pooling layer which learns to combine the aligned heatmaps into a pooled confidence map. We show that this architecture outperforms a number of others, including one that uses optical flow solely at the input layers, one that regresses joint coordinates directly, and one that predicts heatmaps without spatial fusion. The new architecture outperforms the state of the art by a large margin on three video pose estimation datasets, including the very challenging Poses in the Wild dataset, and outperforms other deep methods that don't use a graphical model on the single-image FLIC benchmark (and also [5, 35] in the high precision region). Tomas Pfister, James Charles, Andrew Zisserman |
ICCV | 2 |
| 2014 | Deep Convolutional Neural Networks for Efficient Pose Estimation in Gesture Videos
Tomas Pfister, Karen Simonyan, James Charles, Andrew Zisserman |
ACCV (1) | 3 |
| 2014 | Upper Body Pose Estimation with Temporal Sequential Forests
James Charles, Tomas Pfister, Derek R. Magee, David C. Hogg, Andrew Zisserman |
BMVC | 1 |
| 2014 | Domain-Adaptive Discriminative One-Shot Learning of Gestures
Tomas Pfister, James Charles, Andrew Zisserman |
ECCV (6) | 2 |
| 2014 | Automatic and Efficient Human Pose Estimation for Sign Language Videos
James Charles, Tomas Pfister, Mark Everingham, Andrew Zisserman |
Int. J. Comput. Vis. | 1 |
| 2013 | Domain Adaptation for Upper Body Pose Tracking in Signed TV BroadcastsabstractThe objective of this work is to estimate upper body pose for signers in TV broadcasts. Given suitable training data, the pose is estimated using a random forest body joint detector. However, obtaining such training data can be costly. The novelty of this paper is a method of transfer learning which is able to harness existing training data and use it for new domains. Our contributions are: (i) a method for adapting existing training data to generate new training data by synthesis for signers with different appearances, and (ii) a method for personalising training data. As a case study we show how the appearance of the arms for different clothing, specifically short and long sleeved clothes, can be modelled to obtain person-specific trackers. We demonstrate that the transfer learning and person specific trackers significantly improve pose estimation performance. James Charles, Tomas Pfister, Derek R. Magee, David C. Hogg, Andrew Zisserman |
BMVC | 1 |
| 2013 | Large-scale Learning of Sign Language by Watching TV (Using Co-occurrences)
Tomas Pfister, James Charles, Andrew Zisserman |
BMVC | 2 |
| 2012 | Automatic and Efficient Long Term Arm and Hand Tracking for Continuous Sign Language TV BroadcastsabstractWe present a fully automatic arm and hand tracker that detects joint positions over continuous sign language video sequences of more than an hour in length. Our framework replicates the state-of-the-art long term tracker by Buehler et al. (IJCV 2011), but does not require the manual annotation and, after automatic initialisation, performs tracking in real-time. We cast the problem as a generic frame-by-frame random forest regressor without a strong spatial model. Our contributions are (i) a co-segmentation algorithm that automatically separates the signer from any signed TV broadcast using a generative layered model; (ii) a method of predicting joint positions given only the segmentation and a colour model using a random forest regressor; and (iii) demonstrating that the random forest can be trained from an existing semi-automatic, but computationally expensive, tracker. The method is applied to signing footage with changing background, challenging imaging conditions, and for different signers. We achieve superior joint localisation results to those obtained using the method of Buehler et al. Tomas Pfister, James Charles, Mark Everingham, Andrew Zisserman |
BMVC | 2 |