VLDB 2026 Research / reviewers in the wild / expert
Chris G. Willcocks
dblp:28/11416
· DBLP profile ↗
19ranked-venue papers
2as first author
9since 2021 · last 2024
0000-0001-6821-3924ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ∞-Diff: Infinite Resolution Diffusion with Subsampled Mollified States
Sam Bond-Taylor, Chris G. Willcocks |
ICLR | 2 |
| 2023 | Exact-NeRF: An Exploration of a Precise Volumetric Parameterization for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have attracted significant attention due to their ability to synthesize novel scene views with great accuracy. However, inherent to their underlying formulation, the sampling of points along a ray with zero width may result in ambiguous representations that lead to further rendering artifacts such as aliasing in the final scene. To address this issue, the recent variant mipNeRF proposes an Integrated Positional Encoding (IPE) based on a conical view frustum. Although this is expressed with an integral formulation, mip-NeRF instead approximates this integral as the expected value of a multivariate Gaussian distribution. This approximation is reliable for short frustums but degrades with highly elongated regions, which arises when dealing with distant scene objects under a larger depth of field. In this paper, we explore the use of an exact approach for calculating the IPE by using a pyramid-based integral formulation instead of an approximated conical-based one. We denote this formulation as Exact-NeRF and contribute the first approach to offer a precise analytical solution to the IPE within the NeRF domain. Our exploratory work illustrates that such an exact formulation (Exact-NeRF) matches the accuracy of mip-NeRF and furthermore provides a natural extension to more challenging scenarios without further modification, such as in the case of unbounded scenes. Our contribution aims to both address the hitherto unexplored issues of frustum approximation in earlier NeRF work and additionally provide insight into the potential future consideration of analytical solutions in future NeRF extensions. Brian K. S. Isaac-Medina, Chris G. Willcocks, Toby P. Breckon |
CVPR | 2 |
| 2023 | Unaligned 2D to 3D Translation with Conditional Vector-Quantized Code Diffusion using TransformersabstractGenerating 3D images of complex objects conditionally from a few 2D views is a difficult synthesis problem, compounded by issues such as domain gap and geometric misalignment. For instance, a unified framework such as Generative Adversarial Networks cannot achieve this unless they explicitly define both a domain-invariant and geometric-invariant joint latent distribution, whereas Neural Radiance Fields are generally unable to handle both issues as they optimize at the pixel level. By contrast, we propose a simple and novel 2D to 3D synthesis approach based on conditional diffusion with vector-quantized codes. Operating in an information-rich code space enables high-resolution 3D synthesis via full-coverage attention across the views. Specifically, we generate the 3D codes (e.g. for CT images) conditional on previously generated 3D codes and the entire codebook of two 2D views (e.g. 2D X-rays). Qualitative and quantitative results demonstrate state-of-the-art performance over specialized methods across varied evaluation criteria, including fidelity metrics such as density, coverage, and distortion metrics for two complex volumetric imagery datasets from in real-world scenarios. Abril Corona-Figueroa, Sam Bond-Taylor, Neelanjan Bhowmik, Yona Falinie Binti A. Gaus, Toby P. Breckon, Hubert P. H. Shum, Chris G. Willcocks |
ICCV | 7 |
| 2023 | Dynamic Unary Convolution in TransformersabstractIt is uncertain whether the power of transformer architectures can complement existing convolutional neural networks. A few recent attempts have combined convolution with transformer design through a range of structures in series, where the main contribution of this paper is to explore a parallel design approach. While previous transformed-based approaches need to segment the image into patch-wise tokens, we observe that the multi-head self-attention conducted on convolutional features is mainly sensitive to global correlations and that the performance degrades when these correlations are not exhibited. We propose two parallel modules along with multi-head self-attention to enhance the transformer. For local information, a dynamic local enhancement module leverages convolution to dynamically and explicitly enhance positive local patches and suppress the response to less informative ones. For mid-level structure, a novel unary co-occurrence excitation module utilizes convolution to actively search the local co-occurrence between patches. The parallel-designed Dynamic Unary Convolution in Transformer (DUCT) blocks are aggregated into a deep architecture, which is comprehensively evaluated across essential computer vision tasks in image-based classification, segmentation, retrieval and density estimation. Both qualitative and quantitative results show our parallel convolutional-transformer approach with dynamic and unary convolution outperforms existing series-designed structures. Haoran Duan 0001, Yang Long 0001, Haofeng Zhang 0001, Chris G. Willcocks, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes
Sam Bond-Taylor, Peter Hessey, Hiroshi Sasaki 0009, Toby P. Breckon, Chris G. Willcocks |
ECCV (23) | 5 |
| 2022 | Multi-view Vision Transformers for Object DetectionabstractObject detection has been thoroughly investigated during the last decade using deep neural networks. However, the inclusion of additional information given by multiple concurrent views of the same scene has not received much attention. In scenarios where objects may appear in obscure poses from certain view points, the use of differing simultaneous views can improve object detection. Therefore, we propose a multi-view fusion network to enrich the backbone features of standard object detection architectures across multiple source and target view points. Our method consists of a transformer decoder for the target view that combines the remaining source views feature maps. In this way, the feature representation of the target view can aggregate feature information from the source view through attention. Our architecture is detector-agnostic, meaning it can be applied across any existing detection backbone. We evaluate performance using YOLOX, Deformable DETR and Swin Transformer baseline detectors, comparing standard single view performance against the addition of our multi-view transformer architecture. Our method achieves a 3% increase of the COCO AP over a four view X-ray security dataset and a slight 0.7% increase on a seven view pedestrian dataset. We demonstrate that the integration of different views using attention-based networks improves the detection performance of multi-view datasets.1 Brian K. S. Isaac-Medina, Chris G. Willcocks, Toby P. Breckon |
ICPR | 2 |
| 2022 | Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive ModelsabstractDeep generative models are a class of techniques that train deep neural networks to model the distribution of training samples. Research has fragmented into various interconnected approaches, each of which make trade-offs including run-time, diversity, and architectural restrictions. In particular, this compendium covers energy-based models, variational autoencoders, generative adversarial networks, autoregressive models, normalizing flows, in addition to numerous hybrid approaches. These techniques are compared and contrasted, explaining the premises behind each and how they are interrelated, while reviewing current state-of-the-art advances and implementations. Sam Bond-Taylor, Adam Leach, Yang Long 0001, Chris G. Willcocks |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Gradient Origin Networks
Sam Bond-Taylor, Chris G. Willcocks |
ICLR | 2 |
| 2021 | The relationship between curvilinear structure enhancement and ridge detection methods
Haifa F. Alhasson, Chris G. Willcocks, Shuaa S. Alharbi, Adetayo Kasim, Boguslaw Obara |
Vis. Comput. | 2 |
| 2020 | Segmentation of Macular Edema Datasets with Small Residual 3D U-Net ArchitecturesabstractThis paper investigates the application of deep convolutional neural networks with prohibitively small datasets to the problem of macular edema segmentation. In particular, we investigate several different heavily regularized architectures. We find that, contrary to popular belief, neural architectures within this application setting are able to achieve close to human-level performance on unseen test images without requiring large numbers of training examples. Annotating these 3D datasets is difficult, with multiple criteria required. It takes an experienced clinician two days to annotate a single 3D image, whereas our trained model achieves similar performance in less than a second. We found that an approach which uses targeted dataset augmentation, alongside architectural simplification with an emphasis on residual design, has acceptable generalization performance- despite relying on fewer than 15 training examples. Jonathan Frawley, Chris G. Willcocks, Maged Habib, Caspar Geenen, David H. Steel, Boguslaw Obara |
BIBE | 2 |
| 2020 | Shape tracing: An extension of sphere tracing for 3D non-convex collision in protein dockingabstractThis paper presents an algorithm, similar to implicit sphere tracing, that ray marches 3D non-convex shapes for efficient collision detection. Instead of finding points on the surface where individual rays strike, an entire shape is marched in unison by a lower bound of the boundary distance, calculated at the closest point between the two surfaces. Advancing one shape towards the other by this new bound allows us to identify a contact in few steps. This method supports arbitrary nonconvex shapes, and can be run in parallel. We apply this to protein-protein docking and show that we can identify around 80 docking poses per second featuring contact but no overlap, irrespective of proteins' specific geometry. This paves the way to future fast docking algorithms, building upon implicit surface representations to quickly find a well-distributed subset of close candidate solutions for further investigation. Adam Leach, Lucas S. P. Rudden, Sam Bond-Taylor, John C. Brigham, Matteo T. Degiacomi, Chris G. Willcocks |
BIBE | 6 |
| 2020 | Multi-view Object Detection Using Epipolar Constraints within Cluttered X-ray Security ImageryabstractAutomatic detection for threat object items is an increasing emerging area of future application in X-ray security imagery. Although modern X-ray security scanners can provide two or more views, the integration of such object detectors across the views has not been widely explored with rigour. Therefore, we investigate the application of geometric constraints using the epipolar nature of multi-view imagery to improve object detection performance. Furthermore, we assume that images come from uncalibrated views, such that a method to estimate the fundamental matrix using ground truth bounding box centroids from multiple view object labels is proposed. In addition, detections are given a confidence probability based on its similarity with respect to the distribution of the distance to the epipolar line. This probability is used as confidence weights for merging duplicated predictions using non-maximum suppression. Using a standard object detector (YOLOv3), our technique increases the average precision of detection by 2.8% on a dataset composed of firearms, laptops, knives and cameras. These results indicate that the integration of images at different views significantly improves the detection performance of threat items of cluttered X-ray security images. Brian K. S. Isaac-Medina, Chris G. Willcocks, Toby P. Breckon |
ICPR | 2 |
| 2020 | Data Augmentation via Mixed Class Interpolation using Cycle-Consistent Generative Adversarial Networks Applied to Cross-Domain ImageryabstractMachine learning driven object detection and classification within non-visible imagery has an important role in many fields such as night vision, all-weather surveillance and aviation security. However, such applications often suffer due to the limited quantity and variety of non-visible spectral domain imagery, in contrast to the high data availability of visible-band imagery that readily enables contemporary deep learning driven detection and classification approaches. To address this problem, this paper proposes and evaluates a novel data augmentation approach that leverages the more readily available visible-band imagery via a generative domain transfer model. The model can synthesise large volumes of non-visible domain imagery by image-to-image (I2I) translation from the visible image domain. Furthermore, we show that the generation of interpolated mixed class (non-visible domain) image examples via our novel Conditional CycleGAN Mixup Augmentation (C2GMA) methodology can lead to a significant improvement in the quality of non-visible domain classification tasks that otherwise suffer due to limited data availability. Focusing on classification within the Synthetic Aperture Radar (SAR) domain, our approach is evaluated on a variation of the Statoil/C-CORE Iceberg Classifier Challenge dataset and achieves 75.4 % accuracy, demonstrating a significant improvement when compared against traditional data augmentation strategies (Rotation, Mixup, and MixCycleGAN). Hiroshi Sasaki 0009, Chris G. Willcocks, Toby P. Breckon |
ICPR | 2 |
| 2020 | Real Time Fencing Move Classification and Detection at Touch Time during a Fencing MatchabstractFencingis a fast-paced sport played with swords which are Épée, Foil, and Sabre. However, such fast-pace can cause referees to make wrong decisions. Review of slow-motion camera footage in tournaments helps referees' decision-making, but it interrupts the match and may not be available for every organisation. Motivated by the need for better decision-making, analysis and availability, we introduce the first fully-automated deep learning classification and detection system for fencing body moves at the moment a touch is made. This is an important step towards creating a fencing analysis system, with player profiling and decision tools that will benefit the fencing community. The proposed architecture combines You Only Look Once version three (YOLOv3) with a ResNet-34 classifier, trained on ImageNet settings, to obtain 83.0 % test accuracy on the fencing moves. These results are exciting development in the sport, providing immediate feedback and analysis along with accessibility, hence making it a valuable tool for trainers and fencing match referees. Cem Ekin Sunal, Chris G. Willcocks, Boguslaw Obara |
ICPR | 2 |
| 2018 | TMIXT: A process flow for Transcribing MIXed handwritten and machine-printed TextabstractHandling large corpuses of documents is of significant importance in many fields, no more so than in the areas of crime investigation and defence, where an organisation may be presented with a large volume of scanned documents which need to be processed in a finite time. However, this problem is exacerbated both by the volume, in terms of scanned documents and the complexity of the pages, which need to be processed. Often containing many different elements, which each need to be processed and understood. Text recognition, which is a primary task of this process, is usually dependent upon the type of text, being either handwritten or machine-printed. Accordingly, the recognition involves prior classification of the text category, before deciding on the recognition method to be applied. This poses a more challenging task if a document contains both handwritten and machine-printed text. In this work, we present a generic process flow for text recognition in scanned documents containing mixed handwritten and machine-printed text without the need to classify text in advance. We realize the proposed process flow using several open-source image processing and text recognition packages. The evaluation is performed using a specially developed variant, presented in this work, of the IAM handwriting database, where we achieve an average transcription accuracy of nearly 80% for pages containing both printed and handwritten text. Fady Medhat, Mahnaz Mohammadi, Sardar F. Jaf, Chris G. Willcocks, Toby P. Breckon, Peter Matthews, A. Stephen McGough, Georgios Theodoropoulos 0001, Boguslaw Obara |
IEEE BigData | 4 |
| 2018 | Using Deep Convolutional Neural Network Architectures for Object Classification and Detection Within X-Ray Baggage Security ImageryabstractWe consider the use of deep convolutional neural networks (CNNs) with transfer learning for the image classification and detection problems posed within the context of X-ray baggage security imagery. The use of the CNN approach requires large amounts of data to facilitate a complex end-to-end feature extraction and classification process. Within the context of X-ray security screening, limited availability of object of interest data examples can thus pose a problem. To overcome this issue, we employ a transfer learning paradigm such that a pre-trained CNN, primarily trained for generalized image classification tasks where sufficient training data exists, can be optimized explicitly as a later secondary process towards this application domain. To provide a consistent feature-space comparison between this approach and traditional feature space representations, we also train support vector machine (SVM) classifier on CNN features. We empirically show that fine-tuned CNN features yield superior performance to conventional hand-crafted features on object classification tasks within this context. Overall we achieve 0.994 accuracy based on AlexNet features trained with SVM classifier. In addition to classification, we also explore the applicability of multiple CNN driven detection paradigms, such as sliding window-based CNN (SW-CNN), Faster region-based CNNs (F-RCNNs), region-based fully convolutional networks (R-FCN), and YOLOv2. We train numerous networks tackling both single and multiple detections over SW-CNN/ F-RCNN/R-FCN/YOLOv2 variants. YOLOv2, Faster-RCNN, and R-FCN provide superior results to the more traditional SW-CNN approaches. With the use of YOLOv2, using input images of size$544\times 544$, we achieve 0.885 mean average precision (mAP) for a six-class object detection problem. The same approach with an input of size$416\times 416$yields 0.974 mAP for the two-class firearm detection problem and requires approximately 100 ms per image. Overall we illustrate the comparative performance of these techniques and show that object localization strategies cope well with cluttered X-ray security imagery, where classification techniques fail. Samet Akcay, Mikolaj E. Kundegorski, Chris G. Willcocks, Toby P. Breckon |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Multi-Scale Segmentation and Surface Fitting for Measuring 3-D Macular HolesabstractMacular holes are blinding conditions, where a hole develops in the central part of retina, resulting in reduced central vision. The prognosis and treatment options are related to a number of variables, including the macular hole size and shape. High-resolution spectral domain optical coherence tomography allows precise imaging of the macular hole geometry in three dimensions, but the measurement of these by human observers is time-consuming and prone to high inter- and intra-observer variability, being characteristically measured in 2-D rather than 3-D. We introduce several novel techniques to automatically retrieve accurate 3-D measurements of the macular hole, including: surface area, base area, base diameter, top area, top diameter, height, and minimum diameter. Specifically, we introduce a multi-scale 3-D level set segmentation approach based on a state-of-the-art level set method, and we introduce novel curvature-based cutting and 3-D measurement procedures. The algorithm is fully automatic, and we validate our extracted measurements both qualitatively and quantitatively, where our results show the method to be robust across a variety of scenarios. Our automated processes are considered a significant contribution for clinical applications. Amar Vijai Nasrulloh, Chris G. Willcocks, Philip T. G. Jackson, Caspar Geenen, Maged Habib, David H. Steel, Boguslaw Obara |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Extracting 3D Parametric Curves from 2D Images of Helical ObjectsabstractHelical objects occur in medicine, biology, cosmetics, nanotechnology, and engineering. Extracting a 3D parametric curve from a 2D image of a helical object has many practical applications, in particular being able to extract metrics such as tortuosity, frequency, and pitch. We present a method that is able to straighten the image object and derive a robust 3D helical curve from peaks in the object boundary. The algorithm has a small number of stable parameters that require little tuning, and the curve is validated against both synthetic and real-world data. The results show that the extracted 3D curve comes within close Hausdorff distance to the ground truth, and has near identical tortuosity for helical objects with a circular profile. Parameter insensitivity and robustness against high levels of image noise are demonstrated thoroughly and quantitatively. Chris G. Willcocks, Philip T. G. Jackson, Carl J. Nelson, Boguslaw Obara |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Feature-varying skeletonization - Intuitive control over the target feature size and output skeleton topology
Chris G. Willcocks, Frederick W. B. Li |
Vis. Comput. | 1 |