EDBT 2026 Demo / reviewers in the wild / expert
Toby P. Breckon
dblp:43/3031
· DBLP profile ↗
118ranked-venue papers
4as first author
43since 2021 · last 2026
0000-0003-1666-7590ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 4 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KD360-VoxelBEV: LiDAR and 360-degree Camera Cross Modality Knowledge Distillation for Bird's-Eye-View SegmentationabstractWe present the first cross-modality distillation framework specifically tailored for single-panoramic-camera Bird’s-Eye-View (BEV) segmentation. Our approach leverages a novel LiDAR image representation fused from range, intensity and ambient channels, together with a voxel-aligned view transformer that preserves spatial fidelity while enabling efficient BEV processing. During training, a high-capacity LiDAR and camera fusion Teacher network extracts both rich spatial and semantic features for cross-modality knowledge distillation into a lightweight Student network that relies solely on a single 360-degree panoramic camera image. Extensive experiments on the Dur360BEV dataset demonstrate that our teacher model significantly outperforms existing camera-based BEV segmentation methods, achieving a 25.6% IoU improvement. Meanwhile, the distilled Student network attains competitive performance with an 8.5% IoU gain and state-of-the-art inference speed of 31.2 FPS. Moreover, evaluations on KITTI-360 (two fisheye cameras) confirm that our distillation framework generalises to diverse camera setups, underscoring its feasibility and robustness. This approach reduces sensor complexity and deployment costs while providing a practical solution for efficient, low-cost BEV segmentation in real-world autonomous driving. The code is available at: https://github.com/Tom-E-Durham/KD360-VoxelBEV. Wenke E, Jiaxu Liu 0002, Hubert P. H. Shum, Amir Atapour Abarghouei, Toby P. Breckon |
WACV | 6 |
| 2026 | DDS-NAS: Dynamic data selection within neural architecture search via on-line hard example mining applied to image classificationabstractIn order to address the scalability challenge within Neural Architecture Search (NAS), we speed up NAS training via dynamic hard example mining within a curriculum learning framework. By utilizing an autoencoder that enforces an image similarity embedding in latent space, we construct an efficient kd -tree structure to order images by furthest neighbour dissimilarity in a low-dimensional embedding. From a given query image from our subsample dataset, we can identify the most dissimilar image within the global dataset in logarithmic time. Via curriculum learning, we then dynamically re-formulate an unbiased subsample dataset for NAS optimisation, upon which the current NAS solution architecture performs poorly. We show that our DDS-NAS framework speeds up gradient-based NAS strategies by up to 27 × without loss in performance. By maximising the contribution of each image sample during training, we reduce the duration of a NAS training cycle and the number of iterations required for convergence. Matt Poyser, Toby P. Breckon |
Pattern Recognit. | 2 |
| 2025 | FEVER-OOD: Free Energy Vulnerability Elimination for Robust Out-of-Distribution DetectionabstractModern machine learning models, that excel on computer vision tasks such as classification and object detection, are often overconfident in their predictions for Out-of-Distribution (OOD) examples, resulting in unpredictable behaviour for open-set environments. Recent works have demonstrated that the free energy score is an effective measure of uncertainty for OOD detection given its close relationship to the data distribution. However, despite free energy-based methods representing a significant empirical advance in OOD detection, our theoretical analysis reveals previously unexplored and inherent vulnerabilities within the free energy score formulation such that in-distribution and OOD instances can have distinct feature representations yet identical free energy scores. This phenomenon occurs when the vector direction representing the feature space difference between the in-distribution and OOD sample lies within the null space of the last layer of a neural-based classifier. To mitigate these issues, we explore lower-dimensional feature spaces to reduce the null space footprint and introduce novel regularisation to maximize the least singular value of the final linear layer, hence enhancing inter-sample free energy separation. We refer to these techniques as Free Energy Vulnerability Elimination for Robust Out-of-Distribution Detection (FEVER-OOD). Our experiments show that FEVER-OOD techniques achieve state of the art OOD detection in Imagenet-100, with average OOD false positive rate (at 95% true positive rate) of 35.83% when used with the baseline Dream-OOD model. Brian K. S. Isaac-Medina, Mauricio Che, Yona Faline A. Gaus, Samet Akcay, Toby P. Breckon |
ICCV | 5 |
| 2025 | Dur360BEV: A Real-World 360-Degree Single Camera Dataset and Benchmark for Bird-Eye View Mapping in Autonomous DrivingabstractWe present Dur360BEV, a novel spherical camera autonomous driving dataset equipped with a high-resolution 128-channel 3D LiDAR and a RTK-refined GNSS/INS system, along with a benchmark architecture designed to generate Bird-Eye-View (BEV) maps using only a single spherical camera. This dataset and benchmark address the challenges of BEV generation in autonomous driving, particularly by reducing hardware complexity through the use of a single 360-degree camera instead of multiple perspective cameras. Within our benchmark architecture, we propose a novel spherical-image-to-BEV module that leverages spherical imagery and a refined sampling strategy to project features from 2D to 3D. Our approach also includes an innovative application of focal loss, specifically adapted to address the extreme class imbalance often encountered in BEV segmentation tasks, that demonstrates improved segmentation performance on the Dur360BEV dataset. The results show that our benchmark not only simplifies the sensor setup but also achieves competitive performance. Code + Dataset: https://github.com/Tom-E-DurhamJDur360BEV Wenke E, Li Li 0092, Yona Falinie Binti A. Gaus, Amir Atapour Abarghouei, Toby P. Breckon |
ICRA | 7 |
| 2025 | TOMD: A Trail-based Off-road Multimodal Dataset for Traversable Pathway Segmentation under Challenging Illumination ConditionsabstractDetecting traversable pathways in unstructured out-door environments remains a significant challenge for autonomous robots, especially in critical applications such as wide-area search and rescue, as well as in incident management scenarios such as forest fires. Current datasets and models primarily focus on either urban environments or wide vehicle-traversable off-road tracks, leaving a substantial gap in tackling the complexities of trail-based off-road scenarios. To address this issue, we introduce the Trail-based Off-road Multimodal Dataset (TOMD), a comprehensive dataset explicitly designed for narrow and unstructured trail-like environments. Our dataset features high-fidelity multimodal sensor data — including 128-channel LiDAR, stereo imagery, GNSS, IMU, and illumination measurements — collected through repeated runs across di-verse environmental conditions. In addition, we propose a novel dynamic multiscale data fusion model for precise traversable pathway prediction in trail-like areas. The study investigates the impact of various fusion processes — early, cross, and mixed — on model performance under different illumination levels: low-light, normal ambient lighting, and bright conditions. The results highlight the effectiveness of our approach, variation in performance across illumination levels, and the potential applicability of the dataset in diverse environmental conditions.Our work provides a valuable resource for advancing trail-based off-road navigation, and we openly publish our TOMD at https://github.com/yyyxs1125/TMOD to establish a future bench-mark in this research domain. Li Li 0092, Wenke E, Amir Atapour Abarghouei, Toby P. Breckon |
IJCNN | 5 |
| 2024 | TraIL-Det: Transformation-Invariant Local Feature Networks for 3D LiDAR Object Detection with Unsupervised Pre-Training
Li Li 0092, Tanqiu Qiao, Hubert P. H. Shum, Toby P. Breckon |
BMVC | 4 |
| 2024 | Towards Open-World Object-Based Anomaly Detection via Self-Supervised Outlier Synthesis
Brian K. S. Isaac-Medina, Yona Falinie Binti A. Gaus, Neelanjan Bhowmik, Toby P. Breckon |
ECCV (71) | 4 |
| 2024 | RAPiD-Seg: Range-Aware Pointwise Distance Distribution Networks for 3D LiDAR Segmentation
Li Li 0092, Hubert P. H. Shum, Toby P. Breckon |
ECCV (7) | 3 |
| 2024 | Superpixel-based Anomaly Detection for Irregular Textures with a Focus on Pixel-level AccuracyabstractRecent anomaly detection methods achieve high performance on commonly used image and pixel-level metrics. However, due to the imbalance in the number of normal and abnormal pixels commonly encountered in anomaly detection problems, commonly adopted pixel-level performance metrics cannot effectively evaluate model performance. This paper proposes a novel approach for anomaly detection within the irregular texture domain, focusing on pixel-level accuracy metrics suitable for such imbalanced problems. The proposed Superpixel-based Coupled-hypersphere-based Feature Adaptation (Sp-CFA) method leverages the intermediate adaptive representation of superpixels to enable superior pixel-level anomaly detection performance. We demonstrate superior performance over the irregular texture classes within the MVTec AD benchmark dataset, KSDD2 dataset, and an X-ray dataset of manufactured fibrous products. Mehdi Rafiei, Toby P. Breckon, Alexandros Iosifidis |
IJCNN | 2 |
| 2024 | Disentangling Racial Phenotypes: Fine-Grained Control of Race-related Facial Phenotype CharacteristicsabstractAchieving an effective fine-grained appearance variation over 2D facial images, whilst preserving facial identity, is a challenging task due to the high complexity and entanglement of common 2D facial feature encoding spaces. Despite these challenges, such fine-grained control, by way of disentanglement is a crucial enabler for data-driven racial bias mitigation strategies across multiple automated facial analysis tasks, as it allows to analyse, characterise and synthesise human facial diversity. In this paper, we propose a novel GAN framework to enable fine-grained control over individual race-related phenotype attributes of the facial images. Our framework factors the latent (feature) space into elements that correspond to race-related facial phenotype representations, thereby separating phenotype aspects (e.g. skin, hair colour, nose, eye, mouth shapes), which are notoriously difficult to annotate robustly in real-world facial data. Concurrently, we also introduce a high quality augmented, diverse 2D face image dataset drawn from CelebA-HQ for GAN training. Unlike prior work, our framework only relies upon 2D imagery and related parameters to achieve state-of-the-art individual control over race-related phenotype attributes with improved photo-realistic output. Seyma Yucer, Amir Atapour Abarghouei, Noura Al Moubayed, Toby P. Breckon |
IJCNN | 4 |
| 2024 | U3DS3: Unsupervised 3D Semantic Scene SegmentationabstractContemporary point cloud segmentation approaches largely rely on richly annotated 3D training data. However, it is both time-consuming and challenging to obtain consistently accurate annotations for such 3D scene data. Moreover, there is still a lack of investigation into fully unsupervised scene segmentation for point clouds, especially for holistic 3D scenes. This paper presents U3DS3, as a step towards completely unsupervised point cloud segmentation for any holistic 3D scenes. To achieve this, U3DS3leverages a generalized unsupervised segmentation method for both object and background across both indoor and outdoor static 3D point clouds with no requirement for model pre-training, by leveraging only the inherent information of the point cloud to achieve full 3D scene segmentation. The initial step of our proposed approach involves generating superpoints based on the geometric characteristics of each scene. Subsequently, it undergoes a learning process through a spatial clustering-based methodology, followed by iterative training using pseudo-labels generated in accordance with the cluster centroids. Moreover, by leveraging the invariance and equivariance of the volumetric representations, we apply the geometric transformation on voxelized features to provide two sets of descriptors for robust representation learning. Finally, our evaluation provides state-of-the-art results on the ScanNet and SemanticKITTI, and competitive results on the S3DIS, benchmark datasets. Jiaxu Liu 0002, Zhengdi Yu, Toby P. Breckon, Hubert P. H. Shum |
WACV | 3 |
| 2024 | Neural architecture search: A contemporary literature review for computer vision applicationsabstractDeep Neural Networks have received considerable attention in recent years. As the complexity of network architecture increases in relation to the task complexity, it becomes harder to manually craft an optimal neural network architecture and train it to convergence. As such, Neural Architecture Search (NAS) is becoming far more prevalent within computer vision research, especially when the construction of efficient, smaller network architectures is becoming an increasingly important area of research, for which NAS is well suited. However, despite their promise, contemporary and end-to-end NAS pipeline require vast computational training resources. In this paper, we present a comprehensive overview of contemporary NAS approaches with respect to image classification, object detection, and image segmentation. We adopt consistent terminology to overcome contradictions common within existing NAS literature. Furthermore, we identify and compare current performance limitations in addition to highlighting directions for future NAS research. Matt Poyser, Toby P. Breckon |
Pattern Recognit. | 2 |
| 2023 | Exact-NeRF: An Exploration of a Precise Volumetric Parameterization for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have attracted significant attention due to their ability to synthesize novel scene views with great accuracy. However, inherent to their underlying formulation, the sampling of points along a ray with zero width may result in ambiguous representations that lead to further rendering artifacts such as aliasing in the final scene. To address this issue, the recent variant mipNeRF proposes an Integrated Positional Encoding (IPE) based on a conical view frustum. Although this is expressed with an integral formulation, mip-NeRF instead approximates this integral as the expected value of a multivariate Gaussian distribution. This approximation is reliable for short frustums but degrades with highly elongated regions, which arises when dealing with distant scene objects under a larger depth of field. In this paper, we explore the use of an exact approach for calculating the IPE by using a pyramid-based integral formulation instead of an approximated conical-based one. We denote this formulation as Exact-NeRF and contribute the first approach to offer a precise analytical solution to the IPE within the NeRF domain. Our exploratory work illustrates that such an exact formulation (Exact-NeRF) matches the accuracy of mip-NeRF and furthermore provides a natural extension to more challenging scenarios without further modification, such as in the case of unbounded scenes. Our contribution aims to both address the hitherto unexplored issues of frustum approximation in earlier NeRF work and additionally provide insight into the potential future consideration of analytical solutions in future NeRF extensions. Brian K. S. Isaac-Medina, Chris G. Willcocks, Toby P. Breckon |
CVPR | 3 |
| 2023 | Less is More: Reducing Task and Model Complexity for 3D Point Cloud Semantic SegmentationabstractWhilst the availability of 3D LiDAR point cloud data has significantly grown in recent years, annotation remains expensive and time-consuming, leading to a demand for semi-supervised semantic segmentation methods with application domains such as autonomous driving. Existing work very often employs relatively large segmentation backbone networks to improve segmentation accuracy, at the expense of computational costs. In addition, many use uniform sampling to reduce ground truth data requirements for learning needed, often resulting in sub-optimal performance. To address these issues, we propose a new pipeline that employs a smaller architecture, requiring fewer ground-truth annotations to achieve superior segmentation accuracy compared to contemporary approaches. This is facilitated via a novel Sparse Depthwise Separable Convolution module that significantly reduces the network parameter count while retaining overall task performance. To effectively sub-sample our training data, we propose a new Spatio-Temporal Redundant Frame Downsampling (ST-RFD) method that leverages knowledge of sensor motion within the environment to extract a more diverse subset of training data frame samples. To leverage the use of limited annotated data samples, we further propose a soft pseudo-label method informed by Li-DAR reflectivity. Our method outperforms contemporary semi-supervised work in terms of mIoU, using less labeled data, on the SemanticKITTI (59.5@5%) and ScribbleKITTI (58.1@5%) benchmark datasets, based on a$2.3\times reduction$in model parameters and$64l\times fewer$multiply-add operations whilst also demonstrating significant performance improvement on limited training data (i.e., Less is More). Li Li 0092, Hubert P. H. Shum, Toby P. Breckon |
CVPR | 3 |
| 2023 | ACR: Attention Collaboration-based Regressor for Arbitrary Two-Hand ReconstructionabstractReconstructing two hands from monocular RGB images is challenging due to frequent occlusion and mutual confusion. Existing methods mainly learn an entangled representation to encode two interacting hands, which are in-credibly fragile to impaired interaction, such as truncated hands, separate hands, or external occlusion. This paper presents ACR (Attention Collaboration-based Regres-sor), which makes the first attempt to reconstruct hands in arbitrary scenarios. To achieve this, ACR explicitly miti-gates interdependencies between hands and between parts by leveraging center and part-based attention for feature extraction. However, reducing interdependence helps re-lease the input constraint while weakening the mutual reasoning about reconstructing the interacting hands. Thus, based on center attention, ACR also learns cross-hand prior that handle the interacting hands better. We evaluate our method on various types of hand reconstruction datasets. Our method significantly outperforms the best interacting-hand approaches on the InterHand2.6M dataset while yielding comparable performance with the state-of-the-art single-hand methods on the FreiHand dataset. More qualitative results on in-the-wild and hand-object interaction datasets and web images/videos further demonstrate the effectiveness of our approach for arbitrary hand reconstruction. Our code is available at this link11https://github.com/ZhengdiYu/Arbitrary-Hands-3D-Reconstruction. Zhengdi Yu, Shaoli Huang, Toby P. Breckon, Jue Wang 0001 |
CVPR | 4 |
| 2023 | Unaligned 2D to 3D Translation with Conditional Vector-Quantized Code Diffusion using TransformersabstractGenerating 3D images of complex objects conditionally from a few 2D views is a difficult synthesis problem, compounded by issues such as domain gap and geometric misalignment. For instance, a unified framework such as Generative Adversarial Networks cannot achieve this unless they explicitly define both a domain-invariant and geometric-invariant joint latent distribution, whereas Neural Radiance Fields are generally unable to handle both issues as they optimize at the pixel level. By contrast, we propose a simple and novel 2D to 3D synthesis approach based on conditional diffusion with vector-quantized codes. Operating in an information-rich code space enables high-resolution 3D synthesis via full-coverage attention across the views. Specifically, we generate the 3D codes (e.g. for CT images) conditional on previously generated 3D codes and the entire codebook of two 2D views (e.g. 2D X-rays). Qualitative and quantitative results demonstrate state-of-the-art performance over specialized methods across varied evaluation criteria, including fidelity metrics such as density, coverage, and distortion metrics for two complex volumetric imagery datasets from in real-world scenarios. Abril Corona-Figueroa, Sam Bond-Taylor, Neelanjan Bhowmik, Yona Falinie Binti A. Gaus, Toby P. Breckon, Hubert P. H. Shum, Chris G. Willcocks |
ICCV | 5 |
| 2023 | On Fine-Tuned Deep Features for Unsupervised Domain AdaptationabstractPrior feature transformation based approaches to Unsupervised Domain Adaptation (UDA) employ the deep features extracted by pre-trained deep models without fine-tuning them on the specific source or target domain data for a particular domain adaptation task. In contrast, end-to-end learning based approaches optimise the pre-trained backbones and the customised adaptation modules simultaneously to learn domain-invariant features for UDA. In this work, we explore the potential of combining fine-tuned features and feature transformation based UDA methods for improved domain adaptation performance. Specifically, we integrate the prevalent progressive pseudo-labelling techniques into the fine-tuning framework to extract fine-tuned features which are subsequently used in a state-of-the-art feature transformation based domain adaptation method SPL (Selective Pseudo-Labeling). Thorough experiments with multiple deep models including ResNet-50/101 and DeiT-small/base are conducted to demonstrate the combination of fine-tuned features and SPL can achieve state-of-the-art performance on several benchmark datasets. Qian Wang 0017, Fan-Lin Meng, Toby P. Breckon |
IJCNN | 3 |
| 2023 | Generalized zero-shot domain adaptation via coupled conditional variational autoencodersabstractDomain adaptation aims to exploit useful information from the source domain where annotated training data are easier to obtain to address a learning problem in the target domain where only limited or even no annotated data are available. In classification problems, domain adaptation has been studied under the assumption all classes are available in the target domain regardless of the annotations. However, a common situation where only a subset of classes in the target domain are available has not attracted much attention. In this paper, we formulate this particular domain adaptation problem within a generalized zero-shot learning framework by treating the labelled source-domain samples as semantic representations for zero-shot learning. For this novel problem, neither conventional domain adaptation approaches nor zero-shot learning algorithms directly apply. To solve this problem, we present a novel Coupled Conditional Variational Autoencoder (CCVAE) which can generate synthetic target-domain image features for unseen classes from real images in the source domain. Extensive experiments have been conducted on three domain adaptation datasets including a bespoke X-ray security checkpoint dataset to simulate a real-world application in aviation security. The results demonstrate the effectiveness of our proposed approach both against established benchmarks and in terms of real-world applicability. Qian Wang 0017, Toby P. Breckon |
Neural Networks | 2 |
| 2023 | Data augmentation with norm-AE and selective pseudo-labelling for unsupervised domain adaptationabstractWe address the Unsupervised Domain Adaptation (UDA) problem in image classification from a new perspective. In contrast to most existing works which either align the data distributions or learn domain-invariant features, we directly learn a unified classifier for both the source and target domains in the high-dimensional homogeneous feature space without explicit domain alignment. To this end, we employ the effective Selective Pseudo-Labelling (SPL) technique to take advantage of the unlabelled samples in the target domain. Surprisingly, data distribution discrepancy across the source and target domains can be well handled by a computationally simple classifier (e.g., a shallow Multi-Layer Perceptron) trained in the original feature space. Besides, we propose a novel generative model norm-AE to generate synthetic features for the target domain as a data augmentation strategy to enhance the classifier training. Experimental results on several benchmark datasets demonstrate the pseudo-labelling strategy itself can lead to comparable performance to many state-of-the-art methods whilst the use of norm-AE for feature augmentation can further improve the performance in most cases. As a result, our proposed methods (i.e. naive-SPL and norm-AE-SPL) can achieve comparable performance with state-of-the-art methods with the average accuracy of 93.4% and 90.4% on Office-Caltech and ImageCLEF-DA datasets, and achieve competitive performance on Digits, Office31 and Office-Home datasets with the average accuracy of 97.2%, 87.6% and 68.6% respectively. Qian Wang 0017, Fan-Lin Meng, Toby P. Breckon |
Neural Networks | 3 |
| 2022 | VID-Trans-ReID: Enhanced Video Transformers for Person Re-identification
Aishah Alsehaim, Toby P. Breckon |
BMVC | 2 |
| 2022 | Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes
Sam Bond-Taylor, Peter Hessey, Hiroshi Sasaki 0009, Toby P. Breckon, Chris G. Willcocks |
ECCV (23) | 4 |
| 2022 | Does lossy image compression affect racial bias within face recognition?abstractYes - This study investigates the impact of commonplace lossy image compression on face recognition algorithms with regard to the racial characteristics of the subject. We adopt a recently proposed racial phenotype-based bias analysis methodology to measure the effect of varying levels of lossy compression across racial phenotype categories. Additionally, we determine the relationship between chroma-subsampling and race-related phenotypes for recognition performance. Prior work investigates the impact of lossy JPEG compression algorithm on contemporary face recognition performance. However, there is a gap in how this impact varies with different race-related inter-sectional groups and the cause of this impact. Via an extensive experimental setup, we demonstrate that common lossy image compression approaches have a more pronounced negative impact on facial recognition performance for specific racial phenotype categories such as darker skin tones (by up to 34.55%). Furthermore, removing chroma-subsampling during compression improves the false matching rate (up to 15.95%) across all phenotype categories affected by the compression, including darker skin tones, wide noses, big lips, and monolid eye categories. In addition, we outline the characteristics that may be attributable as the underlying cause of such phenomenon for lossy compression algorithms such as JPEG. Seyma Yucer, Matt Poyser, Noura Al Moubayed, Toby P. Breckon |
IJCB | 4 |
| 2022 | Joint Sub-component Level Segmentation and Classification for Anomaly Detection within Dual-Energy X-Ray Security ImageryabstractX-ray baggage security screening is in widespread use and crucial to maintaining transport security for threat/anomaly detection tasks. The automatic detection of anomaly, which is concealed within cluttered and complex electronics/electrical items, using 2D X-ray imagery is of primary interest in recent years. We address this task by introducing joint object sub-component level segmentation and classification strategy using deep Convolution Neural Network architecture. The performance is evaluated over a dataset of cluttered X-ray baggage security imagery, consisting of consumer electrical and electronics items using variants of dual-energy X-ray imagery (pseudo-colour, high, low, and effective-Z). The proposed joint sub-component level segmentation and classification approach achieve ∼99% true positive and ∼5% false positive for anomaly detection task. Neelanjan Bhowmik, Toby P. Breckon |
ICMLA | 2 |
| 2022 | On Depth Error from Spherical Camera Calibration within Omnidirectional Stereo VisionabstractAs a depth sensing approach, whilst stereo vision provides a good compromise between accuracy and cost, a key limitation is the limited field of view of the conventional cameras that are used within most stereo configurations. By contrast, the use of spherical cameras within a stereo configuration offers omnidirectional stereo sensing. However, despite the presence of significant image distortion in spherical camera images, only very limited attempts have been made to study and quantify omnidirectional stereo depth accuracy.In this paper we construct such an omnidirectional stereo system that is capable of real-time 360° disparity map reconstruction as the basis for such a study. We first investigate the accuracy of using a standard spherical camera model for calibration combined with a longitude-latitude projection for omnidirectional stereo, and show that the depth error increases significantly as the angle from the camera optical axis approaches the limits of the camera field of view.In contrast, we then consider an alternative calibration approach via the use of perspective undistortion with a conventional pinhole camera model allowing omnidirectional cameras to be mapped to a conventional rectilinear stereo formulation. We find that conversely this proposed approach exhibits improved depth accuracy at large angles from the camera optical axis when compared to omnidirectional stereo depth based on a spherical camera model calibration. Michael Groom, Toby P. Breckon |
ICPR | 2 |
| 2022 | Multi-view Vision Transformers for Object DetectionabstractObject detection has been thoroughly investigated during the last decade using deep neural networks. However, the inclusion of additional information given by multiple concurrent views of the same scene has not received much attention. In scenarios where objects may appear in obscure poses from certain view points, the use of differing simultaneous views can improve object detection. Therefore, we propose a multi-view fusion network to enrich the backbone features of standard object detection architectures across multiple source and target view points. Our method consists of a transformer decoder for the target view that combines the remaining source views feature maps. In this way, the feature representation of the target view can aggregate feature information from the source view through attention. Our architecture is detector-agnostic, meaning it can be applied across any existing detection backbone. We evaluate performance using YOLOX, Deformable DETR and Swin Transformer baseline detectors, comparing standard single view performance against the addition of our multi-view transformer architecture. Our method achieves a 3% increase of the COCO AP over a four view X-ray security dataset and a slight 0.7% increase on a seven view pedestrian dataset. We demonstrate that the integration of different views using attention-based networks improves the detection performance of multi-view datasets.1 Brian K. S. Isaac-Medina, Chris G. Willcocks, Toby P. Breckon |
ICPR | 3 |
| 2022 | Evaluating Gaussian Grasp Maps for Generative Grasping ModelsabstractGeneralising robotic grasping to previously unseen objects is a key task in general robotic manipulation. The current method for training many antipodal generative grasping models rely on a binary ground truth grasp map generated from the centre thirds of correctly labelled grasp rectangles. However, these binary maps do not accurately reflect the positions in which a robotic arm can correctly grasp a given object. We propose a continuous Gaussian representation of annotated grasps to generate ground truth training data which achieves a higher success rate on a simulated robotic grasping benchmark. Three modern generative grasping networks are trained with either binary or Gaussian grasp maps, along with recent advancements from the robotic grasping literature, such as discretisation of grasp angles into bins and an attentional loss function. Despite negligible difference according to the standard rectangle metric, Gaussian maps better reproduce the training data and therefore improve success rates when tested on the same simulated robot arm by avoiding collisions with the object: achieving 87.94% accuracy. Furthermore, the best performing model is shown to operate with a high success rate when transferred to a real robotic arm, at high inference speeds, without the need for transfer learning. The system is then shown to be capable of performing grasps on an antagonistic physical object dataset benchmark. William Prew, Toby P. Breckon, Magnus Bordewich, Ulrik R. Beierholm |
IJCNN | 2 |
| 2022 | Measuring Hidden Bias within Face Recognition via Racial PhenotypesabstractRecent work reports disparate performance for intersectional racial groups across face recognition tasks: face verification and identification. However, the definition of those racial groups has a significant impact on the underlying findings of such racial bias analysis. Previous studies define these groups based on either demographic information (e.g. African, Asian etc.) or skin tone (e.g. lighter or darker skins). The use of such sensitive or broad group definitions has disadvantages for bias investigation and subsequent counter-bias solutions design. By contrast, this study introduces an alternative racial bias analysis methodology via facial phenotype attributes for face recognition. We use the set of observable characteristics of an individual face where a race-related facial phenotype is hence specific to the human face and correlated to the racial profile of the subject. We propose categorical test cases to investigate the individual influence of those attributes on bias within face recognition tasks. We compare our phenotype-based grouping methodology with previous grouping strategies and show that phenotype-based groupings uncover hidden bias without reliance upon any potentially protected attributes or ill-defined grouping strategies. Furthermore, we contribute corresponding phenotype attribute category labels for two face recognition tasks: RFW for face verification and VGGFace2 (test set) for face identification. Seyma Yucer, Furkan Tektas, Noura Al Moubayed, Toby P. Breckon |
WACV | 4 |
| 2022 | Towards automatic threat detection: A survey of advances of deep learning within X-ray security imaging
Samet Akcay, Toby P. Breckon |
Pattern Recognit. | 2 |
| 2022 | Cross-domain structure preserving projection for heterogeneous domain adaptation
Qian Wang 0017, Toby P. Breckon |
Pattern Recognit. | 2 |
| 2022 | Crowd Counting via Segmentation Guided Attention Networks and Curriculum LossabstractAutomatic crowd behaviour analysis is an important task for intelligent transportation systems to enable effective flow control and dynamic route planning for varying road participants. Crowd counting is one of the keys to automatic crowd behaviour analysis. Crowd counting using deep convolutional neural networks (CNN) has achieved encouraging progress in recent years. Researchers have devoted much effort to the design of variant CNN architectures and most of them are based on the pre-trained VGG16 model. Due to the insufficient expressive capacity, the backbone network of VGG16 is usually followed by another cumbersome network specially designed for good counting performance. Although VGG models have been outperformed by Inception models in image classification tasks, the existing crowd counting networks built with Inception modules still only have a small number of layers with basic types of Inception modules. To fill in this gap, in this paper, we firstly benchmark the baseline Inception-v3 model on commonly used crowd counting datasets and achieve surprisingly good performance comparable with or better than most existing crowd counting models. Subsequently, we push the boundary of this disruptive work further by proposing a Segmentation Guided Attention Network (SGANet) with Inception-v3 as the backbone and a novel curriculum loss for crowd counting. We conduct thorough experiments to compare the performance of our SGANet with prior arts and the proposed model can achieve state-of-the-art performance with MAE of 57.6, 6.3 and 87.6 on ShanghaiTechA, ShanghaiTechB and UCF_QNRF, respectively. Qian Wang 0017, Toby P. Breckon |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Temporal and non-temporal contextual saliency analysis for generalized wide-area search within unmanned aerial vehicle (UAV) videoabstractAbstract Unmanned aerial vehicles (UAV) can be used to great effect for wide-area searches such as search and rescue operations. UAV enable search and rescue teams to cover large areas more efficiently and in less time. However, using UAV for this purpose involves the creation of large amounts of data, typically in video format, which must be analysed before any potential findings can be uncovered and actions taken. This is a slow and expensive process which can result in significant delays to the response time after a target is seen by the UAV. To solve this problem we propose a deep model architecture using a visual saliency approach to automatically analyse and detect anomalies in UAV video. Our Temporal Contextual Saliency (TeCS) approach is based on the state-of-the-art in visual saliency detection using deep Convolutional Neural Networks (CNN) and considers local and scene context, with novel additions in utilizing temporal information through a convolutional Long Short-Term Memory (LSTM) layer and modifications to the base model architecture. We additionally evaluate the impact of temporal vs non-temporal reasoning for this task. Our model achieves improved results on a benchmark dataset with the addition of temporal reasoning showing significantly improved results compared to the state-of-the-art in saliency detection. Simon G. E. Gökstorp, Toby P. Breckon |
Vis. Comput. | 2 |
| 2021 | DurLAR: A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery for Multi-Modal Autonomous Driving ApplicationsabstractWe present DurLAR, a high-fidelity 128-channel 3D LiDAR dataset with panoramic ambient (near infrared) and reflectivity imagery, as well as a sample benchmark task using depth estimation for autonomous driving applications. Our driving platform is equipped with a high resolution 128 channel LiDAR, a 2MPix stereo camera, a lux meter and a GNSSlINS system. Ambient and reflectivity images are made available along with the LiDAR point clouds to facilitate multi-modal use of concurrent ambient and reflectivity scene information. Leveraging DurLAR, with a resolution exceeding that of prior benchmarks, we consider the task of monocular depth estimation and use this increased availability of higher resolution, yet sparse ground truth scene depth information to propose a novel joint supervised/self-supervised loss formulation. We compare performance over both our new DurLAR dataset, the established KITTI benchmark and the Cityscapes dataset. Our evaluation shows our joint use supervised and self-supervised loss terms, enabled via the superior ground truth resolution and availability within DurLAR improves the quantitative and qualitative performance of leading contemporary monocular depth estimation approaches (RMSE = 3.639, SqRel = 0.936). Li Li 0092, Khalid N. Ismail, Hubert P. H. Shum, Toby P. Breckon |
3DV | 4 |
| 2021 | Re-ID-AR: Improved Person Re-identification in Videovia Joint Weakly Supervised Action Recognition
Aishah Alsehaim, Toby P. Breckon |
BMVC | 2 |
| 2021 | On the Impact of Using X-Ray Energy Response Imagery for Object Detection Via Convolutional Neural NetworksabstractAutomatic detection of prohibited items within complex and cluttered X-ray security imagery is essential to maintaining transport security, where prior work on automatic prohibited item detection focus primarily on pseudo-colour (rgb) X-ray imagery. In this work we study the impact of variant X-ray imagery, i.e., X-ray energy response (high, low) and effective-z compared to rgb, via the use of deep Convolutional Neural Networks (CNN) for the joint object detection and segmentation task posed within X-ray baggage security screening. We evaluate state-of-the-art CNN architectures (Mask R-CNN, YOLACT, CARAFE and Cascade Mask R-CNN) to explore the transferability of models trained with such ‘raw’ variant imagery between the varying X-ray security scanners that exhibits differing imaging geometries, image resolutions and material colour profiles. Overall, we observe maximal detection performance using CARAFE, attributable to training using combination of rgb, high, low, and effective-z Xray imagery, obtaining 0.7 mean Average Precision (mAP) for a six class object detection problem. Our results also exhibit a remarkable degree of generalisation capability in terms of cross-scanner transferability (AP: 0.835/0.611) for a one class object detection problem by combining rgb, high, low, and effective-z imagery. Neelanjan Bhowmik, Yona Falinie Binti A. Gaus, Toby P. Breckon |
ICIP | 3 |
| 2021 | Source Class Selection With Label Propagation For Partial Domain AdaptationabstractIn traditional unsupervised domain adaptation problems, the target domain is assumed to share the same set of classes as the source domain. In practice, there exist situations where target-domain data are from only a subset of source-domain classes and it is not known which classes the target-domain data belong to since they are unlabeled. This problem has been formulated as Partial Domain Adaptation (PDA) in the literature and is a challenging task due to the negative transfer issue (i.e. source-domain data belonging to the irrelevant classes harm the domain adaptation). We address the PDA problem by detecting the outlier classes in the source domain progressively. As a result, the PDA is boiled down to an easier unsupervised domain adaptation problem which can be solved without the issue of negative transfer. Specifically, we employ the locality preserving projection to learn a latent common subspace in which a label propagation algorithm is used to label the target-domain data. The outlier classes can be detected if no target-domain data are labeled as these classes. We remove the detected outlier classes from the source domain and repeat the process for multiple iterations until convergence. Experimental results on commonly used datasets Office31 and Office-Home demonstrate our proposed method achieves state-of-the-art performance with an average accuracy of 98.1% and 75.4% respectively. Qian Wang 0017, Toby P. Breckon |
ICIP | 2 |
| 2021 | Continuous Multi-modal Emotion Prediction in Video based on Recurrent Neural Network Variants with AttentionabstractAutomatic perception and understanding of human emotion is becoming an increasingly attractive research field in artificial intelligence and human-computer interaction. Emotion portrayal within conversation plays a significant role in the semantics of a sentence. However, emotion is not only biologically determined but is also influenced by the environment. Therefore, cultural differences exist in some aspects of emotions, and it is important for the next generation of computer systems to adapt the cross-cultural difference in order to enable more naturalistic interactions between humans and machines. In this paper, we investigate the suitability of state-of-the-art deep learning architectures based on recurrent neural network (RNN) variants with explicit attention modelling to bridge the gap across different cultures (German and Hungarian) for emotion prediction in video. Three different attention based network architectures are proposed in this work:- early attention fusion, extended multi-attention fusion and attention-based encoder-decoder. Our RNN variants with explicit attention modelling approach achieves very promising Concordance Correlation Coefficient results, which outperform the baseline on Arousal of 0.637 vs. 0.614 (baseline), for Valence of 0.689 vs. 0.615 and for Liking of 0.625 vs. 0.222. Joyal Raju, Yona Falinie Binti A. Gaus, Toby P. Breckon |
ICMLA | 3 |
| 2021 | Contraband Materials Detection Within Volumetric 3D Computed Tomography Baggage Security Screening ImageryabstractAutomatic prohibited object detection within 2D/3D X-ray Computed Tomography (CT) has been studied in literature to enhance the aviation security screening at checkpoints. Deep Convolutional Neural Networks (CNN) have demonstrated superior performance in 2D X-ray imagery. However, there exists very limited proof of how deep neural networks perform in materials detection within volumetric 3D CT baggage screening imagery. We attempt to close this gap by applying Deep Neural Networks in 3D contraband substance detection based on their material signatures. Specifically, we formulate it as a 3D semantic segmentation problem to identify material types for all voxels based on which contraband materials can be detected. To this end, we firstly investigate 3D CNN based semantic segmentation algorithms such as 3D U-Net and its variants. In contrast to the original dense representation form of volumetric 3D CT data, we propose to convert the CT volumes into sparse point clouds which allows the use of point cloud processing approaches such as PointNet++ towards more efficient processing. Experimental results on a publicly available dataset (NEU ATR) demonstrate the effectiveness of both 3D U-Net and PointNet++ in materials detection in 3D CT imagery for baggage security screening. Qian Wang 0017, Toby P. Breckon |
ICMLA | 2 |
| 2021 | Operationalizing Convolutional Neural Network Architectures for Prohibited Object Detection in X-Ray ImageryabstractThe recent advancement in deep Convolutional Neural Network (CNN) has brought insight into the automation of X-ray security screening for aviation security and beyond. Here, we explore the viability of two recent end-to-end object detection CNN architectures, Cascade R-CNN and FreeAnchor, for prohibited item detection by balancing processing time and the impact of image data compression from an operational viewpoint. Overall, we achieve maximal detection performance using a FreeAnchor architecture with a ResNet50backbone, obtaining mean Average Precision (mAP) of 87.7 and 85.8 for using the OPIXray and SIXray benchmark datasets, showing superior performance over prior work on both. With fewer parameters and less training time, FreeAnchor achieves the highest detection inference speed of $\sim 13$ fps (3.9 ms per image). Furthermore, we evaluate the impact of lossy image compression upon detector performance. The CNN models display substantial resilience to the lossy compression, resulting in only a 1.1% decrease in mAP at the JPEG compression level of 50. Additionally, a thorough evaluation of data augmentation techniques is provided, including adaptions of MixUp and CutMix strategy as well as other standard transformations, further improving the detection accuracy. Thomas W. Webb, Neelanjan Bhowmik, Yona Falinie Binti A. Gaus, Toby P. Breckon |
ICMLA | 4 |
| 2021 | Autoencoders Without Reconstruction for Textural Anomaly DetectionabstractAutomatic anomaly detection in natural textures is a key component within quality control for a range of high-speed, high-yield manufacturing industries that rely on camera-based visual inspection techniques. Targeting anomaly detection through the use of autoencoder reconstruction error readily facilitates training on an often more plentiful set of non-anomalous samples, without the explicit need for a representative set of anomalous training samples that may be difficult to source. Unfortunately, autoencoders struggle to reconstruct high-frequency visual information and therefore, such approaches often fail to achieve a low enough reconstruction error for non-anomalous pixels. In this paper, we propose a new approach in which the autoencoder is trained to directly output the desired per-pixel measure of abnormality without first having to perform reconstruction. This is achieved by corrupting training samples with noise and then predicting how pixels need to be shifted so as to remove the noise. Our direct approach enables the model to compress anomaly scores for normal pixels into a tight bound close to zero, resulting in very clean anomaly segmentations that significantly improve performance. We also introduce the Reflected ReLU output activation function that better facilitates training under this direct regime by leaving values that fall within the image dynamic range unmodified. Overall, an average area under the ROC curve of 96% is achieved on the texture classes of the MVTecAD benchmark dataset, surpassing that achieved by all current state-of-the-art methods. Philip A. Adey, Samet Akcay, Magnus Bordewich, Toby P. Breckon |
IJCNN | 4 |
| 2021 | PANDA: Perceptually Aware Neural Detection of AnomaliesabstractSemi-supervised methods of anomaly detection have seen substantial advancement in recent years. Of particular interest are applications of such methods to diverse, real-world anomaly detection problems where anomalous variations can vary from the visually obvious to the very subtle. In this work, we propose a novel fine-grained VAE-GAN architecture trained in a semi-supervised manner in order to detect both visually distinct and subtle anomalies. With the use of a residually connected dual-feature extractor, a fine-grained discriminator and a perceptual loss function, we are able to detect subtle, low inter-class (anomaly vs. normal) variant anomalies with greater detection capability and smaller margins of deviation in AUC value during inference compared to prior work whilst also remaining time-efficient during inference. We achieve state-of-the-art anomaly detection results when compared extensively with prior semi-supervised approaches across a multitude of anomaly detection benchmark tasks including trivial leave-one-out tasks (CIFAR-10 -$\mathbf{AUPRC}_{avg}$: 0.91; MNIST -$\mathbf{AUPRC}_{avg}$: 0.90) in addition to challenging real-world anomaly detection tasks (plant leaf disease - AVC: 0.776; threat item X-ray - AVC: 0.51), video frame-level anomaly detection (VCSDPedl - AVC: 0.95) and high frequency texture with object anomalous defect detection (MVTEC -$\mathbf{AUC}_{avg}$: 0.83). Jack W. Barker, Toby P. Breckon |
IJCNN | 2 |
| 2021 | On the Evaluation of Semi-Supervised 2D Segmentation for Volumetric 3D Computed Tomography Baggage Security ScreeningabstractWe address the automatic contraband material detection problem within volumetric 3D Computed Tomography (CT) data for baggage security screening. Distinct from the prohibited item detection using object detection techniques, contraband material detection is usually formulated as a segmentation problem due to the variations of their potential appearances and shapes. Previous studies have employed either morphological operation based traditional methods or 3D Convolutional Neural Networks (CNN) for 3D segmentation towards target material detection within volumetric 3D CT baggage security screening imagery. In this work, we investigate the effectiveness of 2D semantic segmentation techniques in this 3D CT segmentation problem. Specifically, we extract 2D slices from three planes of the 3D CT volumes and train a 2D segmentation model which is subsequently used to predict segmentation results for all the slices from a given test CT volume. Moreover, we also evaluate how the performance is affected when using a reduced number of annotated slices for training. As a result, it is demonstrated reasonable performance can be achieved with very limited annotated slices (1–2) per CT volume during training. Finally, we propose a semi-supervised learning framework for 3D CT segmentation. Using only 1/128 of the total number of annotated slices, our framework can achieve comparable performance with full supervision. Qian Wang 0017, Toby P. Breckon |
IJCNN | 2 |
| 2021 | Competitive Simplicity for Multi-Task Learning for Real-Time Foggy Scene Understanding via Domain AdaptationabstractAutomotive scene understanding under adverse weather conditions raises a realistic and challenging problem attributable to poor outdoor scene visibility (e.g. foggy weather). However, because most contemporary scene understanding approaches are applied under ideal-weather conditions, such approaches may not provide genuinely optimal performance when compared to established a priori insights on extreme-weather understanding. In this paper, we propose a complex but competitive multi-task learning approach capable of performing in real-time semantic scene understanding and monocular depth estimation under foggy weather conditions by leveraging both recent advances in adversarial training and domain adaptation. As an end-to-end pipeline, our model provides a novel solution to surpass degraded visibility in foggy weather conditions by transferring scenes from foggy to normal using a GAN-based model. For optimal performance in semantic segmentation, our model generates depth to be used as complementary source information with RGB in the segmentation network. We provide a robust method for foggy scene understanding by training two models (normal and foggy) simultaneously with shared weights (each model is trained on each weather condition). Our model incorporates RGB colour, depth, and luminance images via distinct encoders with dense connectivity and features fusing, and leverages skip connections to produce consistent depth and segmentation predictions. Using this architectural formulation with light computational complexity at inference time, we are able to achieve comparable performance to contemporary approaches at a fraction of the overall model complexity. Evaluation over several foggy weather condition datasets including synthetic and real-world examples illustrates our approach competitive performance compared to other contemporary state-of-the-art approaches. Naif Alshammari, Samet Akcay, Toby P. Breckon |
IV | 3 |
| 2021 | Multi-Modal Learning for Real-Time Automotive Semantic Foggy Scene Understanding via Domain AdaptationabstractRobust semantic scene segmentation for automotive applications is a challenging problem in two key aspects: (1) labelling every individual scene pixel and (2) performing this task under unstable weather and illumination changes (e.g., foggy weather), which results in poor outdoor scene visibility. Such visibility limitations lead to non-optimal performance of generalised deep convolutional neural network-based semantic scene segmentation. In this paper, we propose an efficient end-to-end automotive semantic scene understanding approach that is robust to foggy weather conditions. As an end-to-end pipeline, our proposed approach provides: (1) the transformation of imagery from foggy to clear weather conditions using a domain transfer approach (correcting for poor visibility) and (2) semantically segmenting the scene using a competitive encoder-decoder architecture with low computational complexity (en-abling real-time performance). Our approach incorporates RGB colour, depth and luminance images via distinct encoders with dense connectivity and features fusion to effectively exploit information from different inputs, which contributes to an optimal feature representation within the overall model. Using this architectural formulation with dense skip connections, our model achieves comparable performance to contemporary approaches at a fraction of the overall model complexity. Naif Alshammari, Samet Akcay, Toby P. Breckon |
IV | 3 |
| 2020 | Unsupervised Domain Adaptation via Structured Prediction Based Selective Pseudo-LabelingabstractUnsupervised domain adaptation aims to address the problem of classifying unlabeled samples from the target domain whilst labeled samples are only available from the source domain and the data distributions are different in these two domains. As a result, classifiers trained from labeled samples in the source domain suffer from significant performance drop when directly applied to the samples from the target domain. To address this issue, different approaches have been proposed to learn domain-invariant features or domain-specific classifiers. In either case, the lack of labeled samples in the target domain can be an issue which is usually overcome by pseudo-labeling. Inaccurate pseudo-labeling, however, could result in catastrophic error accumulation during learning. In this paper, we propose a novel selective pseudo-labeling strategy based on structured prediction. The idea of structured prediction is inspired by the fact that samples in the target domain are well clustered within the deep feature space so that unsupervised clustering analysis can be used to facilitate accurate pseudo-labeling. Experimental results on four datasets (i.e. Office-Caltech, Office31, ImageCLEF-DA and Office-Home) validate our approach outperforms contemporary state-of-the-art methods. Qian Wang 0017, Toby P. Breckon |
AAAI | 2 |
| 2020 | Efficient and Compact Convolutional Neural Network Architectures for Non-temporal Real-time Fire DetectionabstractAutomatic visual fire detection is used to complement traditional fire detection sensor systems (smoke/heat). In this work, we investigate different Convolutional Neural Network (CNN) architectures and their variants for the non-temporal real-time bounds detection of fire pixel regions in video (or still) imagery. Two reduced complexity compact CNN architectures (NasNet-A-OnFire and ShuffleNetV2-OnFire) are proposed through experimental analysis to optimise the computational efficiency for this task. The results improve upon the current state-of-the-art solution for fire detection, achieving an accuracy of 95% for full-frame binary classification and 97% for superpixel localisation. We notably achieve a classification speed up by a factor of 2.3× for binary classification and 1.3× for superpixel localisation, with runtime of 40 fps and 18 fps respectively, outperforming prior work in the field presenting an efficient, robust and real-time solution for fire region detection. Subsequent implementation on low-powered devices (Nvidia Xavier-NX, achieving 49 fps for full-frame classification via ShuffleNetV2-OnFire) demonstrates our architectures are suitable for various real-world deployment applications. William Thomson, Neelanjan Bhowmik, Toby P. Breckon |
ICMLA | 3 |
| 2020 | Multi-Class 3D Object Detection Within Volumetric 3D Computed Tomography Baggage Security Screening ImageryabstractAutomatic detection of prohibited objects within passenger baggage is important for aviation security. X-ray Computed Tomography (CT) based 3D imaging is widely used in airports for aviation security screening whilst prior work on automatic prohibited item detection focus primarily on 2D X-ray imagery. Whilst some prior work has proven the possibility of extending deep convolutional neural networks (CNN) based automatic prohibited item detection from 2D X-ray imagery to volumetric 3D CT baggage security screening imagery, it focuses on the detection of one specific type of objects (e.g., either bottles or handguns). As a result, multiple models are needed if more than one type of prohibited item is required to be detected in practice. In this paper, we consider the detection of multiple object categories of interest using one unified framework. To this end, we formulate a more challenging multi-class 3D object detection problem within 3D CT imagery and propose a viable solution (3D RetinaNet) to tackle this problem. To enhance the performance of detection we investigate a variety of strategies including data augmentation and varying backbone networks. Experimentation carried out to provide both quantitative and qualitative evaluations of the proposed approach to multi-class 3D object detection within 3D CT baggage security screening imagery. Experimental results demonstrate the combination of the 3D RetinaNet and a series of favorable strategies can achieve a mean Average Precision (mAP) of 65.3% over five object classes (i.e. bottles, handguns, binoculars, glock frames, iPods). The overall performance is affected by the poor performance on glock frames and iPods due to the lack of data and their resemblance with the baggage clutter. Qian Wang 0017, Neelanjan Bhowmik, Toby P. Breckon |
ICMLA | 3 |
| 2020 | Leveraging Synthetic Subject Invariant EEG Signals for Zero Calibration BCIabstractRecently, substantial progress has been made in the area of Brain-Computer Interface (BCI) using modern machine learning techniques to decode and interpret brain signals. While Electroencephalography (EEG) has provided a non-invasive method of interfacing with a human brain, the acquired data is often heavily subject and session dependent. This makes the seamless incorporation of such data into realworld applications intractable as the subject and session data variance can lead to long and tedious calibration requirements and cross-subject generalisation issues. Focusing on a Steady State Visual Evoked Potential (SSVEP) classification systems, we propose a novel means of generating highly-realistic synthetic EEG data invariant to any subject, session or other environmental conditions. Our approach, entitled the Subject Invariant SSVEP Generative Adversarial Network (SIS-GAN), produces synthetic EEG data from multiple SSVEP classes using a single network. Additionally, by taking advantage of a fixed-weight pre-trained subject classification network, we ensure that our generative model remains agnostic to subject-specific features and thus produces subject-invariant data that can be applied to new previously unseen subjects. Our extensive experimental evaluation demonstrates the efficacy of our synthetic data, leading to superior performance, with improvements of up to 16 percentage points in zero-calibration classification tasks when trained using our subject-invariant synthetic EEG signals. Nik Khadijah Nik Aznan, Amir Atapour Abarghouei, Stephen Bonner, Jason D. Connolly, Toby P. Breckon |
ICPR | 5 |
| 2020 | Not 3D Re-ID: Simple Single Stream 2D Convolution for Robust Video Re-identificationabstractVideo-based person re-identification has received increasing attention recently, as it plays an important role within surveillance video analysis. Video-based Re-Id is an expansion of earlier image-based re-identification methods by learning features from a video via multiple image frames for each person. Most contemporary video Re-ID methods utilise complex CNN-based network architectures using 3D convolution or multi-branch networks to extract spatial-temporal video features. By contrast, in this paper, we illustrate superior performance from a simple single stream 2D convolution network leveraging the ResNet50-IBN architecture to extract frame-level features followed by temporal attention for clip level features. These clip level features can be generalised to extract video level features by averaging without any significant additional cost. Our approach uses best video Re-Id practice and transfer learning between datasets to outperform existing state-of-the-art approaches on the MARS, PRID2011 and iLIDS-VID datasets with 89.62%, 97.75%, 97.33% rank-1 accuracy respectively and with 84.61% mAP for MARS, without reliance on complex and memory intensive 3D convolutions or multi-stream networks architectures as found in other contemporary work. Conversely, our work shows that global features extracted by the 2D convolution network are a sufficient representation for robust state of the art video Re-ID. Toby P. Breckon, Aishah Alsehaim |
ICPR | 1 |
| 2020 | Multi-view Object Detection Using Epipolar Constraints within Cluttered X-ray Security ImageryabstractAutomatic detection for threat object items is an increasing emerging area of future application in X-ray security imagery. Although modern X-ray security scanners can provide two or more views, the integration of such object detectors across the views has not been widely explored with rigour. Therefore, we investigate the application of geometric constraints using the epipolar nature of multi-view imagery to improve object detection performance. Furthermore, we assume that images come from uncalibrated views, such that a method to estimate the fundamental matrix using ground truth bounding box centroids from multiple view object labels is proposed. In addition, detections are given a confidence probability based on its similarity with respect to the distribution of the distance to the epipolar line. This probability is used as confidence weights for merging duplicated predictions using non-maximum suppression. Using a standard object detector (YOLOv3), our technique increases the average precision of detection by 2.8% on a dataset composed of firearms, laptops, knives and cameras. These results indicate that the integration of images at different views significantly improves the detection performance of threat items of cluttered X-ray security images. Brian K. S. Isaac-Medina, Chris G. Willcocks, Toby P. Breckon |
ICPR | 3 |
| 2020 | On the Impact of Lossy Image and Video Compression on the Performance of Deep Convolutional Neural Network ArchitecturesabstractRecent advances in generalized image understanding have seen a surge in the use of deep convolutional neural networks (CNN) across a broad range of image-based detection, classification and prediction tasks. Whilst the reported performance of these approaches is impressive, this study investigates the hitherto unapproached question of the impact of commonplace image and video compression techniques on the performance of such deep learning architectures. Focusing on the JPEG and H.264 (MPEG-4 AVC) as a representative proxy for contemporary lossy image/video compression techniques that are in common use within network-connected image/video devices and infrastructure, we examine the impact on performance across five discrete tasks: human pose estimation, semantic segmentation, object detection, action recognition, and monocular depth estimation. As such, within this study we include a variety of network architectures and domains spanning end-to-end convolution, encoder-decoder, region-based CNN (R-CNN), dual-stream, and generative adversarial networks (GAN). Our results show a non-linear and non-uniform relationship between network performance and the level of lossy compression applied. Notably, performance decreases significantly below a JPEG quality (quantization) level of 15% and a H.264 Constant Rate Factor (CRF) of 40. However, retraining said architectures on pre-compressed imagery conversely recovers network performance by up to 78.4% in some cases. Furthermore, there is a correlation between architectures employing an encoder-decoder pipeline and those that demonstrate resilience to lossy image compression. The characteristics of the relationship between input compression to output task performance can be used to inform design decisions within future image/video devices and infrastructure. Matt Poyser, Amir Atapour Abarghouei, Toby P. Breckon |
ICPR | 3 |
| 2020 | Improving Robotic Grasping on Monocular Images Via Multi-Task Learning and Positional LossabstractIn this paper we introduce two methods of improving real-time object grasping performance from monocular colour images in an end-to-end CNN architecture. The first is the addition of an auxiliary task during model training (multi-task learning). Our multi-task CNN model improves grasping performance from a baseline average of 72.04% to 78.14% on the large Jacquard grasping dataset when performing a supplementary depth reconstruction task. The second is introducing a positional loss function that emphasises loss per pixel for secondary parameters (gripper angle and width) only on points of an object where a successful grasp can take place. This increases performance from a baseline average of 72.04% to 78.92% as well as reducing the number of training epochs required. These methods can be also performed in tandem resulting in a further performance increase to 79.12%, while maintaining sufficient inference speed to afford real-time grasp processing. William Prew, Toby P. Breckon, Magnus Bordewich, Ulrik R. Beierholm |
ICPR | 2 |
| 2020 | Data Augmentation via Mixed Class Interpolation using Cycle-Consistent Generative Adversarial Networks Applied to Cross-Domain ImageryabstractMachine learning driven object detection and classification within non-visible imagery has an important role in many fields such as night vision, all-weather surveillance and aviation security. However, such applications often suffer due to the limited quantity and variety of non-visible spectral domain imagery, in contrast to the high data availability of visible-band imagery that readily enables contemporary deep learning driven detection and classification approaches. To address this problem, this paper proposes and evaluates a novel data augmentation approach that leverages the more readily available visible-band imagery via a generative domain transfer model. The model can synthesise large volumes of non-visible domain imagery by image-to-image (I2I) translation from the visible image domain. Furthermore, we show that the generation of interpolated mixed class (non-visible domain) image examples via our novel Conditional CycleGAN Mixup Augmentation (C2GMA) methodology can lead to a significant improvement in the quality of non-visible domain classification tasks that otherwise suffer due to limited data availability. Focusing on classification within the Synthetic Aperture Radar (SAR) domain, our approach is evaluated on a variation of the Statoil/C-CORE Iceberg Classifier Challenge dataset and achieves 75.4 % accuracy, demonstrating a significant improvement when compared against traditional data augmentation strategies (Rotation, Mixup, and MixCycleGAN). Hiroshi Sasaki 0009, Chris G. Willcocks, Toby P. Breckon |
ICPR | 3 |
| 2020 | On the Evaluation of Prohibited Item Classification and Detection in Volumetric 3D Computed Tomography Baggage Security Screening ImageryabstractX-ray Computed Tomography (CT) based 3D imaging is widely used in airports for aviation security screening whilst prior work on prohibited item detection focuses primarily on 2D X-ray imagery. In this paper, we aim to evaluate the possibility of extending the automatic prohibited item detection from 2D X-ray imagery to volumetric 3D CT baggage security screening imagery. To these ends, we take advantage of 3D Convolutional Neural Networks (CNN) and popular object detection frameworks such as RetinaNet and Faster R-CNN in our work. As the first attempt to use 3D CNN for volumetric 3D CT baggage security screening, we first evaluate different CNN architectures on the classification of isolated prohibited item volumes and compare against traditional methods which use hand-crafted features. Subsequently, we evaluate object detection performance of different architectures on volumetric 3D CT baggage images. The results of our experiments on Bottle and Handgun datasets demonstrate that 3D CNN models can achieve comparable performance (~ 98% true positive rate and ~1.5% false positive rate) to traditional methods but require significantly less time for inference (0.014s per volume). Furthermore, the extended 3D object detection models achieve promising performance in detecting prohibited items within volumetric 3D CT baggage imagery with ~76% mAP for bottles and ~88% mAP for handguns, which shows both the challenge and promise of such threat detection within 3D CT X-ray security imagery. Qian Wang 0017, Neelanjan Bhowmik, Toby P. Breckon |
IJCNN | 3 |
| 2019 | To Complete or to Estimate, That is the Question: A Multi-Task Approach to Depth Completion and Monocular Depth EstimationabstractRobust three-dimensional scene understanding is now an ever-growing area of research highly relevant in many real-world applications such as autonomous driving and robotic navigation. In this paper, we propose a multi-task learning-based model capable of performing two tasks:- sparse depth completion (i.e. generating complete dense scene depth given a sparse depth image as the input) and monocular depth estimation (i.e. predicting scene depth from a single RGB image) via two sub-networks jointly trained end to end using data randomly sampled from a publicly available corpus of synthetic and real-world images. The first sub-network generates a sparse depth image by learning lower level features from the scene and the second predicts a full dense depth image of the entire scene, leading to a better geometric and contextual understanding of the scene and, as a result, superior performance of the approach. The entire model can be used to infer complete scene depth from a single RGB image or the second network can be used alone to perform depth completion given a sparse depth input. Using adversarial training, a robust objective function, a deep architecture relying on skip connections and a blend of synthetic and real-world training data, our approach is capable of producing superior high quality scene depth. Extensive experimental evaluation demonstrates the efficacy of our approach compared to contemporary state-of-the-art techniques across both problem domains. Amir Atapour Abarghouei, Toby P. Breckon |
3DV | 2 |
| 2019 | On the Use of Neural Text Generation for the Task of Optical Character RecognitionabstractOptical Character Recognition (OCR), is extraction of textual data from scanned text documents to facilitate their indexing, searching, editing and to reduce storage space. Although OCR systems have improved significantly in recent years, they still suffer in situations where the OCR output does not match the text in the original document. Deep learning models have contributed positively to many problems but their full potential to many other problems are yet to be explored. In this paper we propose a post-processing approach based on the application deep learning to improve the accuracy of OCR system (minimizing the error rate). We report on the use of neural network language models to accomplish the task of correcting incorrectly predicted characters/words by OCR systems. We applied our approach to the IAM handwriting database. Our proposed approach delivers significant accuracy improvement of 20.41% in F-score, 10.86% in character level comparison using Levenshtein distance and 20.69% in document level comparison over previously reported context based OCR empirical results of IAM handwriting database. Mahnaz Mohammadi, Sardar F. Jaf, A. Stephen McGough, Toby P. Breckon, Peter Matthews, Georgios Theodoropoulos 0001, Boguslaw Obara |
AICCSA | 4 |
| 2019 | Veritatem Dies Aperit - Temporally Consistent Depth Prediction Enabled by a Multi-Task Geometric and Semantic Scene Understanding ApproachabstractRobust geometric and semantic scene understanding is ever more important in many real-world applications such as autonomous driving and robotic navigation. In this paper, we propose a multi-task learning-based approach capable of jointly performing geometric and semantic scene understanding, namely depth prediction (monocular depth estimation and depth completion) and semantic scene segmentation. Within a single temporally constrained recurrent network, our approach uniquely takes advantage of a complex series of skip connections, adversarial training and the temporal constraint of sequential frame recurrence to produce consistent depth and semantic class labels simultaneously. Extensive experimental evaluation demonstrates the efficacy of our approach compared to other contemporary state-of-the-art techniques. Amir Atapour Abarghouei, Toby P. Breckon |
CVPR | 2 |
| 2019 | Monocular Segment-Wise Depth: Monocular Depth Estimation Based on a Semantic Segmentation PriorabstractMonocular depth estimation using novel learning-based approaches has recently emerged as a promising potential alternative to more conventional 3D scene capture technologies within real-world scenarios. Many such solutions often depend on large quantities of ground truth depth data, which is rare and often intractable to obtain. Others attempt to estimate disparity as an intermediary step using a secondary supervisory signal, leading to blurring and other undesirable artefacts. In this paper, we propose a monocular depth estimation approach, which employs a jointly-trained pixel-wise semantic understanding step to estimate depth for individually-selected groups of objects (segments) within the scene. The separate depth outputs are efficiently fused to generate the final result. This creates more simplistic learning objectives for the jointly-trained individual networks, leading to more accurate overall depth. Extensive experimentation demonstrates the efficacy of the proposed approach compared to contemporary state-of-the-art techniques within the literature. Amir Atapour Abarghouei, Toby P. Breckon |
ICIP | 2 |
| 2019 | A Ranking Based Attention Approach for Visual TrackingabstractCorrelation filters (CF) combined with pre-trained convolutional neural network (CNN) feature extractors have shown an admirable accuracy and speed in visual object tracking. However, existing CNN-CF based methods still suffer from the background interference and boundary effects, even when a cosine window is introduced. This paper proposes a ranking based or guided attention approach which can reduce background interference with only forward propagation. This ranking stores several convolution kernels and scores them. Subsequently, a convolutional Long Short Time Memory network (ConvLSTM) is used to update this ranking, which makes it more robust to the variation and occlusion. Moreover, a part-based multi-channel convolutional tracker is proposed to obtain the final response map. Our extensive experiments on established benchmark datasets show comparable performance against contemporary tracking approaches. Shenhui Peng, Toby P. Breckon |
ICIP | 3 |
| 2019 | Degraf-Flow: Extending Degraf Features for Accurate and Efficient Sparse-To-Dense Optical Flow EstimationabstractModern optical flow methods make use of salient scene feature points detected and matched within the scene as a basis for sparse-to-dense optical flow estimation. Current feature detectors however either give sparse, non uniform point clouds (resulting in flow inaccuracies) or lack the efficiency for frame-rate real-time applications. In this work we use the novel Dense Gradient Based Features (DeGraF) as the input to a sparse-to-dense optical flow scheme. This consists of three stages: 1) efficient detection of uniformly distributed Dense Gradient Based Features (DeGraF) [1]; 2) feature tracking via robust local optical flow [2]; and 3) edge preserving flow interpolation [3] to recover overall dense optical flow. The tunable density and uniformity of DeGraF features yield superior dense optical flow estimation compared to other popular feature detectors within this three stage pipeline. Furthermore, the comparable speed of feature detection also lends itself well to the aim of real-time optical flow recovery. Evaluation on established real-world benchmark datasets show test performance in an autonomous vehicle setting where DeGraF-Flow shows promising results in terms of accuracy with competitive computational efficiency among non-GPU based methods, including a marked increase in speed over the conceptually similar EpicFlow approach [3]. Felix Stephenson, Toby P. Breckon, Ioannis Katramados |
ICIP | 2 |
| 2019 | A Baseline for Multi-Label Image Classification Using an Ensemble of Deep Convolutional Neural NetworksabstractRecent studies on multi-label image classification have focused on designing more complex architectures of deep neural networks such as the use of attention mechanisms and region proposal networks. Although performance gains have been reported, the backbone deep models of the proposed approaches and the evaluation metrics employed in different works vary, making it difficult to compare fairly. Moreover, due to the lack of properly investigated baselines, the advantage introduced by the proposed techniques are often ambiguous. To address these issues, we make a thorough investigation of the mainstream deep convolutional neural network architectures for multi-label image classification and present a strong baseline. With the use of proper data augmentation techniques and model ensembles, the basic deep architectures can achieve better performance than many existing more complex ones on three benchmark datasets, providing great insight for the future studies on multi-label image classification. Qian Wang 0017, Toby P. Breckon |
ICIP | 3 |
| 2019 | Experimental Exploration of Compact Convolutional Neural Network Architectures for Non-Temporal Real-Time Fire DetectionabstractIn this work we explore different Convolutional Neural Network (CNN) architectures and their variants for non-temporal binary fire detection and localization in video or still imagery. We consider the performance of experimentally defined, reduced complexity deep CNN architectures for this task and evaluate the effects of different optimization and normalization techniques applied to different CNN architectures (spanning the Inception, ResNet and EfficientNet architectural concepts). Contrary to contemporary trends in the field, our work illustrates a maximum overall accuracy of 0.96 for full frame binary fire detection and 0.94 for superpixel localization using an experimentally defined reduced CNN architecture based on the concept of InceptionV4. We notably achieve a lower false positive rate of 0.06 compared to prior work in the field presenting an efficient, robust and real-time solution for fire region detection. Ganesh Samarth C. A., Neelanjan Bhowmik, Toby P. Breckon |
ICMLA | 3 |
| 2019 | Region Based Anomaly Detection with Real-Time Training and AnalysisabstractWe present a method of anomaly detection that is capable of real-time operation on a live stream of images. The real-time performance applies to the training of the algorithm as well as subsequent analysis, and is achieved by substituting the region proposal mechanism used in [9] with one that makes the overall method more efficient. where they generate thousands of regions per image, we generate far fewer but better targeted regions. We also propose a 'convolutional' variant which does away with region extraction altogether, and propose improvements to the density estimation phase used in both variants. Philip A. Adey, Oliver K. Hamilton, Magnus Bordewich, Toby P. Breckon |
ICMLA | 4 |
| 2019 | On the Impact of Object and Sub-Component Level Segmentation Strategies for Supervised Anomaly Detection within X-Ray Security ImageryabstractX-ray security screening is in widespread use to maintain transportation security against a wide range of potential threat profiles. Of particular interest is the recent focus on the use of automated screening approaches, including the potential anomaly detection as a methodology for concealment detection within complex electronic items. Here we address this problem considering varying segmentation strategies to enable the use of both object level and sub-component level anomaly detection via the use of secondary convolutional neural network (CNN) architectures. Relative performance is evaluated over an extensive dataset of exemplar cluttered X-ray imagery, with a focus on consumer electronics items. We find that sub-component level segmentation produces marginally superior performance in the secondary anomaly detection via classification stage, with true positive of ~98% of anomalies, with a ~3% false positive. Neelanjan Bhowmik, Yona Falinie Binti A. Gaus, Samet Akcay, Jack W. Barker, Toby P. Breckon |
ICMLA | 5 |
| 2019 | Evaluating the Transferability and Adversarial Discrimination of Convolutional Neural Networks for Threat Object Detection and Classification within X-Ray Security ImageryabstractX-ray imagery security screening is essential to maintaining transport security against a varying profile of threat or prohibited items. Particular interest lies in the automatic detection and classification of weapons such as firearms and knives within complex and cluttered X-ray security imagery. Here, we address this problem by exploring various end-to-end object detection Convolutional Neural Network (CNN) architectures. We evaluate several leading variants spanning the Faster R-CNN, Mask R-CNN, and RetinaNet architectures to explore the transferability of such models between varying X-ray scanners with differing imaging geometries, image resolutions and material colour profiles. Whilst the limited availability of X-ray threat imagery can pose a challenge, we employ a transfer learning approach to evaluate whether such inter-scanner generalisation may exist over a multiple class detection problem. Overall, we achieve maximal detection performance using a Faster R-CNN architecture with a ResNet101 classification network, obtaining 0.88 and 0.86 of mean Average Precision (mAP) for a three-class and two class item from varying X-ray imaging sources. Our results exhibit a remarkable degree of generalisability in terms of cross-scanner performance (mAP: 0.87, firearm detection: 0.94 AP). In addition, we examine the inherent adversarial discriminative capability of such networks using a specifically generated adversarial dataset for firearms detection - with a variable low false positive, as low as 5%, this shows both the challenge and promise of such threat detection within X-ray security imagery. Yona Falinie Binti A. Gaus, Neelanjan Bhowmik, Samet Akcay, Toby P. Breckon |
ICMLA | 4 |
| 2019 | On the Performance of Extended Real-Time Object Detection and Attribute Estimation within Urban Scene UnderstandingabstractWhilst real-time object detection has become an increasingly important task within urban scene understanding for autonomous driving, the majority of prior work concentrates on the detection of obstacles, dynamic scene objects (pedestrians, vehicles) and road sign-age within the scene. By contrast, for an autonomous vehicle to be truly able to interact with occupants and other road users using a common semantic understanding of the environment it is traversing it requires a considerably extended scene understanding capability. In this work, we consider the performance of extended "long-list" object detection, via an extended end-to-end Region-based Convolutional Neural Network (R-CNN) architecture, over a large-scale 31 class detection problem of urban scene objects with integrated object attribute estimation for appropriate colour and primary orientation. We examine the extended performance of this multiple class object detection and attribute estimation task operating in real-time with on-vehicle processing at 10 fps. Our work is evaluated under a range of real-world automotive conditions across multiple complex and cluttered urban environments. Khalid N. Ismail, Toby P. Breckon |
ICMLA | 2 |
| 2019 | Using Variable Natural Environment Brain-Computer Interface Stimuli for Real-time Humanoid Robot NavigationabstractThis paper addresses the challenge of humanoid robot teleoperation in a natural indoor environment via a Brain-Computer Interface (BCI). We leverage deep Convolutional Neural Network (CNN) based image and signal understanding to facilitate both real-time object detection and dry-Electroencephalography (EEG) based human cortical brain bio-signals decoding. We employ recent advances in dry-EEG technology to stream and collect the cortical waveforms from subjects while they fixate on variable Steady State Visual Evoked Potential (SSVEP) stimuli generated directly from the environment the robot is navigating. To these ends, we propose the use of novel variable BCI stimuli by utilising the real-time video streamed via the on-board robot camera as visual input for SSVEP, where the CNN detected natural scene objects are altered and flickered with differing frequencies (10Hz, 12Hz and 15Hz). These stimuli are not akin to traditional stimuli - as both the dimensions of the flicker regions and their on-screen position changes depending on the scene objects detected. Onscreen object selection via such a dry-EEG enabled SSVEP methodology, facilitates the on-line decoding of human cortical brain signals, via a specialised secondary CNN, directly into teleoperation robot commands (approach object, move in a specific direction: right, left or back). This SSVEP decoding model is trained via a priori offline experimental data in which very similar visual input is present for all subjects. The resulting classification demonstrates high performance with mean accuracy of 85% for the real-time robot navigation experiment across multiple test subjects. Nik Khadijah Nik Aznan, Jason D. Connolly, Noura Al Moubayed, Toby P. Breckon |
ICRA | 4 |
| 2019 | Skip-GANomaly: Skip Connected and Adversarially Trained Encoder-Decoder Anomaly DetectionabstractDespite inherent ill-definition, anomaly detection is a research endeavour of great interest within machine learning and visual scene understanding alike. Most commonly, anomaly detection is considered as the detection of outliers within a given data distribution based on some measure of normality. The most significant challenge in real-world anomaly detection problems is that available data is highly imbalanced towards normality (i.e. non-anomalous) and contains at most a sub-set of all possible anomalous samples - hence limiting the use of well-established supervised learning methods. By contrast, we introduce an unsupervised anomaly detection model, trained only on the normal (non-anomalous, plentiful) samples in order to learn the normality distribution of the domain, and hence detect abnormality based on deviation from this model. Our proposed approach employs an encoder-decoder convolutional neural network with skip connections to thoroughly capture the multi-scale distribution of the normal data distribution in image space. Furthermore, utilizing an adversarial training scheme for this chosen architecture provides superior reconstruction both within image space and a lower-dimensional embedding vector space encoding. Minimizing the reconstruction error metric within both the image and hidden vector spaces during training aids the model to learn the distribution of normality as required. Higher reconstruction metrics during subsequent test and deployment are thus indicative of a deviation from this normal distribution, hence indicative of an anomaly. Experimentation over established anomaly detection benchmarks and challenging real-world datasets, within the context of X-ray security screening, shows the unique promise of such a proposed approach. Samet Akcay, Amir Atapour Abarghouei, Toby P. Breckon |
IJCNN | 3 |
| 2019 | Simulating Brain Signals: Creating Synthetic EEG Data via Neural-Based Generative Models for Improved SSVEP ClassificationabstractDespite significant recent progress in the area of Brain-Computer Interface (BCI), there are numerous shortcomings associated with collecting Electroencephalography (EEG) signals in real-world environments. These include, but are not limited to, subject and session data variance, long and arduous calibration processes and predictive generalisation issues across different subjects or sessions. This implies that many downstream applications, including Steady State Visual Evoked Potential (SSVEP) based classification systems, can suffer from a shortage of reliable data. Generating meaningful and realistic synthetic data can therefore be of significant value in circumventing this problem. We explore the use of modern neural-based generative models trained on a limited quantity of EEG data collected from different subjects to generate supplementary synthetic EEG signal vectors, subsequently utilised to train an SSVEP classifier. Extensive experimental analysis demonstrates the efficacy of our generated data, leading to improvements across a variety of evaluations, with the crucial task of cross-subject generalisation improving by over 35% with the use of such synthetic data. Nik Khadijah Nik Aznan, Amir Atapour Abarghouei, Stephen Bonner, Jason D. Connolly, Noura Al Moubayed, Toby P. Breckon |
IJCNN | 6 |
| 2019 | Evaluation of a Dual Convolutional Neural Network Architecture for Object-wise Anomaly Detection in Cluttered X-ray Security ImageryabstractX-ray baggage security screening is widely used to maintain aviation and transport secure. Of particular interest is the focus on automated security X-ray analysis for particular classes of object such as electronics, electrical items and liquids. However, manual inspection of such items is challenging when dealing with potentially anomalous items. Here we present a dual convolutional neural network (CNN) architecture for automatic anomaly detection within complex security X-ray imagery. We leverage recent advances in region-based (R-CNN), mask-based CNN (Mask R-CNN) and detection architectures such as RetinaNet to provide object localisation variants for specific object classes of interest. Subsequently, leveraging a range of established CNN object and fine-grained category classification approaches we formulate within object anomaly detection as a two-class problem (anomalous or benign). Whilst the best performing object localisation method is able to perform with 97.9% mean average precision (mAP) over a six-class X-ray object detection problem, subsequent two-class anomaly/benign classification is able to achieve 66% performance for within object anomaly detection. Overall, this performance illustrates both the challenge and promise of object-wise anomaly detection within the context of cluttered X-ray security imagery. Yona Falinie Binti A. Gaus, Neelanjan Bhowmik, Samet Akcay, Paolo M. Guillén-Garcia, Jack W. Barker, Toby P. Breckon |
IJCNN | 6 |
| 2019 | Unifying Unsupervised Domain Adaptation and Zero-Shot Visual RecognitionabstractUnsupervised domain adaptation aims to transfer knowledge from a source domain to a target domain so that the target domain data can be recognized without any explicit labelling information for this domain. One limitation of the problem setting is that testing data (despite no labels) from the target domain is needed during training, which prevents the trained model being directly applied to classify newly arrived test instances. We formulate a new cross-domain classification problem arising from real-world scenarios where labelled data are available for a subset of classes (known classes) in the target domain, and we expect to recognize new samples belonging to any class (known and unseen classes) once the model is learned. This is a generalized zero-shot learning problem where the side information comes from the source domain in the form of labelled samples instead of class-level semantic representations commonly used in traditional zero-shot learning. We present a unified domain adaptation framework for both unsupervised and zero-shot learning conditions. Our approach learns a joint subspace from source and target domains so that the projections of both data in the subspace can be domain invariant and easily separable. We use the supervised locality preserving projection (SLPP) as the enabling technique and conduct experiments under both unsupervised and zero-shot learning conditions, achieving state-of-the-art results on three domain adaptation benchmark datasets: Office-Caltech, Office31 and Office-Home. Qian Wang 0017, Penghui Bu, Toby P. Breckon |
IJCNN | 3 |
| 2019 | Generative adversarial framework for depth filling via Wasserstein metric, cosine transform and domain transfer
Amir Atapour Abarghouei, Samet Akcay, Grégoire Payen de La Garanderie, Toby P. Breckon |
Pattern Recognit. | 4 |
| 2019 | Discrete Curvature Representations for Noise Robust Image Corner DetectionabstractImage corner detection is very important in the fields of image analysis and computer vision. Curvature calculation techniques are used in many contour-based corner detectors. We identify that existing calculation of curvature is sensitive to local variation and noise in the discrete domain and does not perform well when corners are closely located. In this paper, discrete curvature representations of single and double corner models are investigated and obtained. A number of model properties have been discovered, which help us detect corners on contours. It is shown that the proposed method has a high corner resolution (the ability to accurately detect neighboring corners), and a corresponding corner resolution constant is also derived. Meanwhile, this method is less sensitive to any local variations and noise on the contour; and false corner detection is less likely to occur. The proposed detector is compared with seven state-of-the-art detectors. Three test images with ground truths are used to assess the detection capability and localization accuracy of these methods in cases with noise-free and different noise levels; 24 images with various scenes without ground truths are used to evaluate their repeatability under affine transformation, JPEG compression, and noise degradations. The experimental results show that our proposed detector attains a better overall performance. Changming Sun, Toby P. Breckon, Naif Alshammari |
IEEE Trans. Image Process. | 3 |
| 2018 | GANomaly: Semi-supervised Anomaly Detection via Adversarial Training
Samet Akcay, Amir Atapour Abarghouei, Toby P. Breckon |
ACCV (3) | 3 |
| 2018 | TMIXT: A process flow for Transcribing MIXed handwritten and machine-printed TextabstractHandling large corpuses of documents is of significant importance in many fields, no more so than in the areas of crime investigation and defence, where an organisation may be presented with a large volume of scanned documents which need to be processed in a finite time. However, this problem is exacerbated both by the volume, in terms of scanned documents and the complexity of the pages, which need to be processed. Often containing many different elements, which each need to be processed and understood. Text recognition, which is a primary task of this process, is usually dependent upon the type of text, being either handwritten or machine-printed. Accordingly, the recognition involves prior classification of the text category, before deciding on the recognition method to be applied. This poses a more challenging task if a document contains both handwritten and machine-printed text. In this work, we present a generic process flow for text recognition in scanned documents containing mixed handwritten and machine-printed text without the need to classify text in advance. We realize the proposed process flow using several open-source image processing and text recognition packages. The evaluation is performed using a specially developed variant, presented in this work, of the IAM handwriting database, where we achieve an average transcription accuracy of nearly 80% for pages containing both printed and handwritten text. Fady Medhat, Mahnaz Mohammadi, Sardar F. Jaf, Chris G. Willcocks, Toby P. Breckon, Peter Matthews, A. Stephen McGough, Georgios Theodoropoulos 0001, Boguslaw Obara |
IEEE BigData | 5 |
| 2018 | Real-Time Monocular Depth Estimation Using Synthetic Data With Domain Adaptation via Image Style TransferabstractMonocular depth estimation using learning-based approaches has become promising in recent years. However, most monocular depth estimators either need to rely on large quantities of ground truth depth data, which is extremely expensive and difficult to obtain, or predict disparity as an intermediary step using a secondary supervisory signal leading to blurring and other artefacts. Training a depth estimation model using pixel-perfect synthetic data can resolve most of these issues but introduces the problem of domain bias. This is the inability to apply a model trained on synthetic data to real-world scenarios. With advances in image style transfer and its connections with domain adaptation (Maximum Mean Discrepancy), we take advantage of style transfer and adversarial training to predict pixel perfect depth from a single real-world color image based on training over a large corpus of synthetic environment data. Experimental results indicate the efficacy of our approach compared to contemporary state-of-the-art techniques. Amir Atapour Abarghouei, Toby P. Breckon |
CVPR | 2 |
| 2018 | Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360 ^\circ ∘ Panoramic Imagery
Grégoire Payen de La Garanderie, Amir Atapour Abarghouei, Toby P. Breckon |
ECCV (13) | 3 |
| 2018 | Infrared Image Colorization Using a S-Shape NetworkabstractThis paper proposes a novel approach for colorizing near infrared (NIR) images using a S-shape network (SNet). The proposed approach is based on the usage of an encoder-decoder architecture followed with a secondary assistant network. The encoder-decoder consists of a contracting path to capture context and a symmetric expanding path that enables precise localization. The assistant network is a shallow encoder-decoder to enhance the edge and improve the output, which can be trained end-to-end from a few image examples. The trained model does not require any user guidance or a reference image database. Furthermore, our architecture will preserve clear edges within NIR images. Our overall architecture is trained and evaluated on a real-world dataset containing a significant amount of road scene images. This dataset was captured by a NIR camera and a corresponding RGB camera to facilitate side-by-side comparison. In the experiments, we demonstrate that our SNet works well, and outperforms contemporary state-of-the-art approaches. Ziyue Dong, Toby P. Breckon |
ICIP | 3 |
| 2018 | Experimentally Defined Convolutional Neural Network Architecture Variants for Non-Temporal Real-Time Fire DetectionabstractIn this work we investigate the automatic detection of fire pixel regions in video (or still) imagery within real-time bounds without reliance on temporal scene information. As an extension to prior work in the field, we consider the performance of experimentally defined, reduced complexity deep convolutional neural network architectures for this task. Contrary to contemporary trends in the field, our work illustrates maximal accuracy of 0.93 for whole image binary fire detection, with 0.89 accuracy within our superpixel localization framework can be achieved, via a network architecture of signficantly reduced complexity. These reduced architectures additionally offer a 3-4 fold increase in computational performance offering up to 17 fps processing on contemporary hardware independent of temporal information. We show the relative performance achieved against prior work using benchmark datasets to illustrate maximally robust real-time fire region detection. Andrew J. Dunnings, Toby P. Breckon |
ICIP | 2 |
| 2018 | On the Impact of Varying Region Proposal Strategies for Raindrop Detection and Classification Using Convolutional Neural NetworksabstractThe presence of raindrop induced image distortion has a significant negative impact on the performance of a wide range of all-weather visual sensing applications including within the increasingly important contexts of visual surveillance and vehicle autonomy. A key part of this problem is robust raindrop detection such that the potential for performance degradation in effected image regions can be identified. Here we address the problem of raindrop detection in colour video imagery by considering three varying region proposal approaches with secondary classification via a number of novel convolutional neural network architecture variants. This is verified over an extensive dataset with in-frame raindrop annotation to achieve maximal 0.95 detection accuracy with minimal false positives compared to prior work. Our approach is evaluated under a range of environmental conditions typical of all-weather automotive visual sensing applications. Tiancheng Guo, Samet Akcay, Philip A. Adey, Toby P. Breckon |
ICIP | 4 |
| 2018 | On the Impact of Illumination-Invariant Image Pre-transformation for Contemporary Automotive Semantic Scene UnderstandingabstractIllumination changes in outdoor environments under non-ideal weather conditions have a negative impact on automotive scene understanding and segmentation performance. In this paper, we present an evaluation of illuminationinvariant image transforms applied to this application domain. We compare four recent transforms for illumination invariant image representation, individually and with colour hybrid images, to show that despite assumptions to contrary such invariant pre-processing can improve the state of the art in scene understanding performance. In addition, we propose a robust approach based on using an illumination-invariant image representation, combined with the chromatic component of a perceptual colour-space to improve contemporary automotive scene understanding and segmentation. By using an illumination invariant pre-process, to reduce the impact of environmental illumination changes, we show that the performance of deep convolutional neural network based scene understanding and segmentation can yet be further improved. This illuminating result enforces the need for invariant (unbiased) training sets within such deep network training and shows that even a welltrained network may still not offer truly optimal performance (if we ignore any prior data transforms attributable to a priori insight). Our approach is demonstrated over a range of example imagery where we show a notable improvement in performance using pre-processed, illumination invariant, automotive scene imagery. Naif Alshammari, Samet Akcay, Toby P. Breckon |
Intelligent Vehicles Symposium | 3 |
| 2018 | Learning to Drive: Using Visual Odometry to Bootstrap Deep Learning for Off-Road Path PredictionabstractAutonomous driving is a fieldcurrently gaining a lot of attention, and recently ‘end to end’ approaches, whereby a machine learning algorithm learns to drive by emulating human drivers, have demonstrated significant potential. However, recent work has focused on the on-road environment, rather than the much more challenging off-road environment. In this work we propose a new approach to this problem, whereby instead of learning to predict immediate driver control inputs, we train a deep convolutional neural network (CNN) to predict the future path that a vehicle will take through an offroad environment visually, addressing several limitations inherent in existing methods. We combine a novel approach to automatic training data creation, making use of stereoscopic visual odometry, with a state of the art CNN architecture to ap a predicted route directly onto image pixels, and demonstrate the effectiveness of our approach using our own off-road data set. Christopher J. Holder, Toby P. Breckon |
Intelligent Vehicles Symposium | 2 |
| 2018 | On the Classification of SSVEP-Based Dry-EEG Signals via Convolutional Neural NetworksabstractElectroencephalography (EEG) is a common signal acquisition approach employed for Brain-Computer Interface (BCI) research. Nevertheless, the majority of EEG acquisition devices rely on the cumbersome application of conductive gel (so-called wet-EEG) to ensure a high quality signal is obtained. However, this process is unpleasant for the experimental participants and thus limits the practical application of BCI. In this work, we explore the use of a commercially available dry-EEG headset to obtain visual cortical ensemble signals. Whilst improving the usability of EEG within the BCI context, dry-EEG suffers from inherently reduced signal quality due to the lack of conduit gel, making the classification of such signals significantly more challenging. In this paper, we propose a novel Convolutional Neural Network (CNN) approach for the classification of raw dry-EEG signals without any data pre-processing. To illustrate the effectiveness of our approach, we utilise the Steady State Visual Evoked Potential (SSVEP) paradigm as our use case. SSVEP can be utilised to allow people with severe physical disabilities such as Complete Locked-In Syndrome or Amyotrophic Lateral Sclerosis to be aided via BCI applications, as it requires only the subject to fixate upon the sensory stimuli of interest. Here we utilise SSVEP flicker frequencies between 10 to 30 Hz, which we record as subject cortical waveforms via the dry-EEG headset. Our proposed end-to-end CNN allows us to automatically and accurately classify SSVEP stimulation directly from the dry-EEG waveforms. Our CNN architecture utilises a common SSVEP Convolutional Unit (SCU), comprising of a 1D convolutional layer, batch normalization and max pooling. Furthermore, We compare several deep learning neural network variants with our primary CNN architecture, in addition to traditional machine learning classification approaches. Experimental evaluation shows our CNN architecture to be significantly better than competing approaches, achieving a classification accuracy of 96% whilst demonstrating superior cross-subject performance and even being able to generalise well to unseen subjects whose data is entirely absent from the training process. Nik Khadijah Nik Aznan, Stephen Bonner, Jason D. Connolly, Noura Al Moubayed, Toby P. Breckon |
SMC | 5 |
| 2018 | A comparative review of plausible hole filling strategies in the context of scene depth image completion
Amir Atapour Abarghouei, Toby P. Breckon |
Comput. Graph. | 2 |
| 2018 | Clustering in pursuit of temporal correlation for human motion segmentation
Toby P. Breckon, Zezhong Xu |
Multim. Tools Appl. | 2 |
| 2018 | Using Deep Convolutional Neural Network Architectures for Object Classification and Detection Within X-Ray Baggage Security ImageryabstractWe consider the use of deep convolutional neural networks (CNNs) with transfer learning for the image classification and detection problems posed within the context of X-ray baggage security imagery. The use of the CNN approach requires large amounts of data to facilitate a complex end-to-end feature extraction and classification process. Within the context of X-ray security screening, limited availability of object of interest data examples can thus pose a problem. To overcome this issue, we employ a transfer learning paradigm such that a pre-trained CNN, primarily trained for generalized image classification tasks where sufficient training data exists, can be optimized explicitly as a later secondary process towards this application domain. To provide a consistent feature-space comparison between this approach and traditional feature space representations, we also train support vector machine (SVM) classifier on CNN features. We empirically show that fine-tuned CNN features yield superior performance to conventional hand-crafted features on object classification tasks within this context. Overall we achieve 0.994 accuracy based on AlexNet features trained with SVM classifier. In addition to classification, we also explore the applicability of multiple CNN driven detection paradigms, such as sliding window-based CNN (SW-CNN), Faster region-based CNNs (F-RCNNs), region-based fully convolutional networks (R-FCN), and YOLOv2. We train numerous networks tackling both single and multiple detections over SW-CNN/ F-RCNN/R-FCN/YOLOv2 variants. YOLOv2, Faster-RCNN, and R-FCN provide superior results to the more traditional SW-CNN approaches. With the use of YOLOv2, using input images of size$544\times 544$, we achieve 0.885 mean average precision (mAP) for a six-class object detection problem. The same approach with an input of size$416\times 416$yields 0.974 mAP for the two-class firearm detection problem and requires approximately 100 ms per image. Overall we illustrate the comparative performance of these techniques and show that object localization strategies cope well with cluttered X-ray security imagery, where classification techniques fail. Samet Akcay, Mikolaj E. Kundegorski, Chris G. Willcocks, Toby P. Breckon |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | DepthComp: Real-time Depth Image Completion Based on Prior Semantic Scene Segmentation
Amir Atapour Abarghouei, Toby P. Breckon |
BMVC | 2 |
| 2017 | An evaluation of region based object detection strategies within X-ray baggage security imageryabstractHere we explore the applicability of traditional sliding window based convolutional neural network (CNN) detection pipeline and region based object detection techniques such as Faster Region-based CNN (R-CNN) and Region-based Fully Convolutional Networks (R-FCN) on the problem of object detection in X-ray security imagery. Within this context, with limited dataset availability, we employ a transfer learning paradigm for network training tackling both single and multiple object detection problems over a number of R-CNN/R-FCN variants. The use of first-stage region proposal within the Faster RCNN and R-FCN provide superior results than traditional sliding window driven CNN (SWCNN) approach. With the use of Faster RCNN with VGG16, pretrained on the ImageNet dataset, we achieve 88.3 mAP for a six object class X-ray detection problem. The use of R-FCN with ResNet-101, yields 96.3 mAP for the two class firearm detection problem requiring 0.1 second computation per image. Overall we illustrate the comparative performance of these techniques as object localization strategies within cluttered X-ray security imagery. Samet Akcay, Toby P. Breckon |
ICIP | 2 |
| 2017 | Noise robust image edge detection based upon the automatic anisotropic Gaussian kernels
Yali Zhao, Toby P. Breckon |
Pattern Recognit. | 3 |
| 2016 | SMS Spam Filtering Using Probabilistic Topic Modelling and Stacked Denoising Autoencoder
Noura Al Moubayed, Toby P. Breckon, Peter Matthews, A. Stephen McGough |
ICANN (2) | 2 |
| 2016 | Transfer learning using convolutional neural networks for object classification within X-ray baggage security imageryabstractWe consider the use of transfer learning, via the use of deep Convolutional Neural Networks (CNN) for the image classification problem posed within the context of X-ray baggage security screening. The use of a deep multi-layer CNN approach, traditionally requires large amounts of training data, in order to facilitate construction of a complex complete end-to-end feature extraction, representation and classification process. Within the context of X-ray security screening, limited availability of training for particular items of interest can thus pose a problem. To overcome this issue, we employ a transfer learning paradigm such that a pre-trained CNN, primarily trained for generalized image classification tasks where sufficient training data exists, can be specifically optimized as a later secondary process that targets specific this application domain. For the classical handgun detection problem we achieve 98.92% detection accuracy outperforming prior work in the field and furthermore extend our evaluation to a multiple object classification task within this context. Samet Akcay, Mikolaj E. Kundegorski, Michael Devereux, Toby P. Breckon |
ICIP | 4 |
| 2016 | Generalized dynamic object removal for dense stereo vision based scene mapping using synthesised optical flowabstractMapping an ever changing urban environment is a challenging task as we are generally interested in mapping the static scene and not the dynamic objects, such as cars and people. We propose a novel approach to the problem of dynamic object removal within stereo based scene mapping that is both independent of the underlying stereo approach in use and applicable to varying object and camera motion. By leveraging stereo odometry, to recover camera motion in scene space, and stereo disparity, to recover synthesised optic flow over the same pixel space, we isolate regions of inconsistency in depth and image intensity. This allows us to illustrate robust dynamic object removal within the stereo mapping sequence. We show results covering objects with a range of motion dynamics and sizes of those typically observed in an urban environment. Oliver K. Hamilton, Toby P. Breckon |
ICIP | 2 |
| 2016 | Dense gradient-based features (DEGRAF) for computationally efficient and invariant feature extraction in real-time applicationsabstractWe propose a computationally efficient approach for the extraction of dense gradient-based features based on the use of localized intensity-weighted centroids within the image. Whilst prior work concentrates on sparse feature derivations or computationally expensive dense scene sensing, we show that Dense Gradient-based Features (DeGraF) can be derived based on initial multi-scale division of Gaussian preprocessing, weighted centroid gradient calculation and either local saliency (DeGraF-α) or signal-to-noise inspired (DeGraF-β) final stage filtering. We present two variants (DeGraF-α / DeGraF-β) of which the signal-to-noise based approach is shown to perform admirably against the state of the art in terms of feature density, computational efficiency and feature stability. Our approach is evaluated under a range of environmental conditions typical of automotive sensing applications with strong feature density requirements. Ioannis Katramados, Toby P. Breckon |
ICIP | 2 |
| 2016 | Constant-time bilateral filter using spectral decompositionabstractThis paper presents an efficient constant-time bilateral filter where constant-time means that computational complexity is independent of filter window size. Many state-of-the-art constant-time methods approximate the original bilateral filter by an appropriate combination of a series of convolutions. It is important for this framework to optimize the performance tradeoff between approximate accuracy and the number of convolutions. The proposed method achieves the optimal performance tradeoff in a least-squares manner by using spectral decomposition under the assumption that images consist of discrete intensities such as 8-bit images. This approach is essentially applicable to arbitrary range kernel. Experiments show that the proposed method outperforms state-of-the-art methods in terms of both computational complexity and approximate accuracy. Kenjiro Sugimoto, Toby P. Breckon |
ICIP | 2 |
| 2016 | Back to Butterworth - a Fourier basis for 3D surface relief hole filling within RGB-D imageryabstractWe address the problem of hole filling in RGB-D (color and depth) images, obtained from either active or stereo based sensing, for the purposes of object removal and missing depth estimation. This is performed independently on the low frequency depth information (surface shape) and the high frequency depth detail (relief) by way of a Fourier space transform and classical Butterworth high/low pass filtering. The high frequency detail is then filled using a texture synthesis method, whilst the low frequency shape information is inpainted using structural inpainting. Here, a classical non-parametric sampling approach is extended, using the concept of query expansion, to perform high frequency depth synthesis with the final output then recombined in Fourier space. In order to improve the overall depth relief (D) and edge detail accuracy, color information (RGB) is also used to constrain the sampling process within high frequency component completion. Experimental results demonstrate the efficacy of the proposed method outperforming prior work for generalized depth filling in the presence of high frequency surface relief detail. Amir Atapour Abarghouei, Grégoire Payen de La Garanderie, Toby P. Breckon |
ICPR | 3 |
| 2015 | Improved 3D sparse maps for high-performance SFM with low-cost omnidirectional robotsabstractWe consider the use of low-budget omnidirectional platforms for 3D mapping and self-localisation. These robots specifically permit rotational motion in the plane around a central axis, with negligible displacement. In addition, low resolution and compressed imagery, typical of the platform used, results in high level of image noise (σ ~ 10). We observe highly sparse image feature matches over narrow inter-image baselines. This particular configuration poses a challenge for epipolar geometry extraction and accurate 3D point triangulation, upon which a standard structure from motion formulation is based. We propose a novel technique for both feature filtering and tracking that solves these problems, via a novel approach to the management of feature bundles. Noisy matches are efficiently trimmed, and the scarcity of the remaining image features is adequately overcome, generating densely populated maps of highly accurate and robust 3D image features. The effectiveness of the approach is demonstrated under a variety of scenarios in experiments conducted with low-budget commercial robots. Pedro Cavestcmy, Antonio L. Rodriguez, Humberto Martínez Barberá, Toby P. Breckon |
ICIP | 4 |
| 2015 | Improved raindrop detection using combined shape and saliency descriptors with scene context isolationabstractThe presence of raindrop induced image distortion has a significant negative impact on the performance of a wide range of all-weather visual sensing applications including within the increasingly import contexts of visual surveillance and vehicle autonomy. A key part of this problem is robust raindrop detection such that the potential for performance degradation in effected image regions can be identified. Here we address the problem of raindrop detection in colour video imagery using an extended feature descriptor comprising localised shape, saliency and texture information isolated from the overall scene context. This is verified within a bag of visual words feature encoding framework using Support Vector Machine and Random Forest classification to achieve notable 86% detection accuracy with minimal false positives compared to prior work. Our approach is evaluated under a range of environmental conditions typical of all-weather automotive visual sensing applications. Dereck D. Webster, Toby P. Breckon |
ICIP | 2 |
| 2015 | Object classification in 3D baggage security computed tomography imagery using visual codebooks
Gregory T. Flitton, André Mouton, Toby P. Breckon |
Pattern Recognit. | 3 |
| 2015 | Materials-based 3D segmentation of unknown objects from dual-energy computed tomography imagery in baggage security screening
André Mouton, Toby P. Breckon |
Pattern Recognit. | 2 |
| 2014 | Improved Depth Recovery In Consumer Depth Cameras via Disparity Space Fusion within Cross-spectral Stereo
Grégoire Payen de La Garanderie, Toby P. Breckon |
BMVC | 2 |
| 2014 | 3D object classification in baggage computed tomography imagery using randomised clustering forestsabstractWe investigate the feasibility of a codebook approach for the automated classification of threats in pre-segmented 3D baggage Computed Tomography (CT) security imagery. We compare the performance of five codebook models, using various combinations of sampling strategies, feature encoding techniques and classifiers, to the current state-of-the-art 3D visual cortex approach. We demonstrate an improvement over the state-of-the-art both in terms of accuracy as well as processing time using a codebook constructed via randomised clustering forests, a dense feature sampling strategy and an SVM classifier. Correct classification rates in excess of 98% and false positive rates of less than 1%, in conjunction with a reduction of several orders of magnitude in processing time, make the proposed approach an attractive option for the automated classification of threats in security screening settings. André Mouton, Toby P. Breckon, Gregory T. Flitton, Najla Megherbi Bouallagu |
ICIP | 2 |
| 2013 | A foreground object based quantitative assessment of dense stereo approaches for use in automotive environmentsabstractThere has been significant recent interest in stereo correspondence algorithms for use in the urban automotive environment [1, 2, 3]. In this paper we evaluate a range of dense stereo algorithms, using a unique evaluation criterion which provides quantitative analysis of accuracy against range, based on ground truth 3D annotated object information. The results show that while some algorithms provide greater scene coverage, we see little differentiation in accuracy over short ranges, while the converse is shown over longer ranges. Within our long range accuracy analysis we see a distinct separation of relative algorithm performance. This study extends prior work on dense stereo evaluation of Block Matching (BM)[4], Semi-Global Block Matching (SGBM)[5], No Maximal Disparity (NoMD)[6], Cross[7], Adaptive Dynamic Programming (AdptDP)[8], Efficient Large Scale (ELAS)[9], Minimum Spanning Forest (MSF)[10] and Non-Local Aggregation (NLA)[11] using a novel quantitative metric relative to object range. Oliver K. Hamilton, Toby P. Breckon, Xuejiao Bai |
ICIP | 2 |
| 2013 | A distance driven method for metal artefact reduction in computed tomographyabstractThis paper presents an extension to a recent intensity-limiting sino-gram completion-based Metal Artefact Reduction (MAR) algorithm for Computed Tomography (CT) images containing multiple metal objects. A novel weighting scheme is introduced, whereby the intensities of the MAR-corrected pixels are modified based on their spatial locations relative to the metal objects. Pixels falling within the straight-line regions connecting multiple metal objects are subjected to less intensive intensity-limiting, thereby compensating for the characteristic dark bands occurring in these regions. Extensive experimentation is performed on a state-of-the-art numerical simulation, a clinical CT data set and a baggage security CT data set. Comprehensive performance analysis, using reference and reference-free error metrics, Bland-Altman plots and visual comparisons, demonstrate an improvement in the restoration of the underestimated intensities occurring in the regions connecting multiple metal objects. André Mouton, Najla Megherbi Bouallagu, Katrien Van Slambrouck, Johan Nuyts, Toby P. Breckon |
ICIP | 5 |
| 2013 | A comparison of 3D interest point descriptors with application to airport baggage object detection in complex CT imagery
Gregory T. Flitton, Toby P. Breckon, Najla Megherbi Bouallagu |
Pattern Recognit. | 2 |
| 2012 | On Cross-Spectral Stereo Matching using Dense Gradient FeaturesabstractWe address the problem of scene depth recovery within cross-spectral stereo imagery (each image sensed over a differing spectral range).We compare several robust matching techniques which are able to capture local similarities between the structure of crossspectral images and a range of stereo optimisation techniques for the computation of valid depth estimates in this case.Specifically we deal with the recovery of dense depth information from thermal (far infrared spectrum) and optical (visible spectrum) image pairs where large differences in the characteristics of image pairs make this task significantly more challenging than the common stereo case.We show that the use of dense gradient features, based on Histograms of Oriented Gradient (HOG) descriptors, for pixel matching in combination with a strong match optimisation approach can produce largely valid, yet coarse, dense depth estimates suitable for object localisation or environment navigation.The proposed solution is compared and shown to work favourably against prior approaches based on using Mutual Information (MI) or Local Self-Similarity (LSS) descriptors. Peter Pinggera, Toby P. Breckon, Horst Bischof |
BMVC | 2 |
| 2012 | A 3D extension to cortex like mechanisms for 3D object class recognitionabstractWe introduce a novel 3D extension to the hierarchical visual cortex model used for prior work in 2D object recognition. Prior work on the use of the visual cortex standard model for the explicit task of object class recognition has solely concentrated on 2D imagery. In this paper we discuss the explicit 3D extension of each layer in this visual cortex model hierarchy for use in object recognition in 3D volumetric imagery. We apply this extended methodology to the automatic detection of a class of threat items in Computed Tomography (CT) security baggage imagery. The CT imagery suffers from poor resolution and a large number of artefacts generated through the presence of metallic objects. In our examination of recognition performance we make a comparison to a codebook approach derived from a 3D SIFT descriptor and demonstrate that the visual cortex method out-performs in this imagery. Recognition rates in excess of 95% with minimal false positive rates are demonstrated in the detection of a range of threat items. Gregory T. Flitton, Toby P. Breckon, Najla Megherbi Bouallagu |
CVPR | 2 |
| 2012 | A comparison of classification approaches for threat detection in CT based baggage screeningabstractComputed Tomography (CT) based baggage security screening systems are of increasing use in transportation security. The ability to automatically identify potential threat item is a key aspect of current research in this area. Here we present a comparison of varying classification approaches for the automated detection of threat objects in cluttered 3D CT imagery from such security screening systems. By combining 3D medical image segmentation techniques with 3D shape classification and retrieval methods we compare five varying final classification stage approaches and present significant performance achievements in the automated detection of specified exemplar items. Najla Megherbi Bouallagu, Ji Wan Han, Toby P. Breckon, Gregory T. Flitton |
ICIP | 3 |
| 2012 | A novel intensity limiting approach to Metal Artefact Reduction in 3D CT baggage imageryabstractThis paper introduces a novel technique for Metal Artefact Reduction (MAR) in the previously unconsidered context 3D CT baggage imagery. The output of a conventional sinogram completion-based MAR approach is refined by imposing an upper limit on the intensity of the corrected images and by performing post-filtering using the non-local means filter. Furthermore, performance is evaluated using a novel quantitative analysis technique, using the ratio of noisy 3D SIFT detection points identified, as well as a standard qualitative comparison (visual quality). The objective of the quantitative analysis is to evaluate the impact of MAR on the application of computer vision techniques for automatic object recognition. The study yields encouraging results in both the qualitative and quantitative analyses. The proposed method yields a significant improvement in performance when compared to algorithms based on linear interpolation and reprojection-reconstruction; especially in terms of reducing the occurrence of new artefacts in the corrected images. The results serve as a strong indication that MAR will aid human and computerised analyses of 3D CT baggage imagery for transport security screening. André Mouton, Najla Megherbi Bouallagu, Gregory T. Flitton, Suzanne Bizot, Toby P. Breckon |
ICIP | 5 |
| 2012 | The application of support vector machine classification to detect cell nuclei for automated microscopy
Ji Wan Han, Toby P. Breckon, David A. Randell, Gabriel Landini |
Mach. Vis. Appl. | 2 |
| 2012 | Automatic real-time road marking recognition using a feature driven approach
Alireza Kheyrollahi, Toby P. Breckon |
Mach. Vis. Appl. | 2 |
| 2012 | A hierarchical extension to 3D non-parametric surface relief completion
Toby P. Breckon, Robert B. Fisher |
Pattern Recognit. | 1 |
| 2011 | A non-temporal texture driven approach to real-time fire detectionabstractHere we investigate the automatic detection of fire pixel regions in conventional video (or still) imagery within realtime bounds. As an extension to prior, established approaches within this field we specifically look to extend the primary use of threshold-driven colour spectroscopy to the combined use of colour-texture feature descriptors as an input to a trained classification approach that is independent of temporal information. We show the limitations of such spectroscopy driven approaches on simple, real-world examples and propose our novel extension as a robust, real-time solution within this field by combining simple texture descriptors to illustrate maximal ~98% fire region detection. Audrey Chenebert, Toby P. Breckon, Anna Gaszczak |
ICIP | 2 |
| 2011 | Real-time visual saliency by Division of GaussiansabstractThis paper introduces a novel method for deriving visual saliency maps in real-time without compromising the quality of the output. This is achieved by replacing the computationally expensive centre-surround filters with a simpler mathematical model named Division of Gaussians (DIVoG). The results are compared to five other approaches, demonstrating at least six times faster execution than the current state-of-the-art whilst maintaining high detection accuracy. Given the multitude of computer vision applications that make use of visual saliency algorithms such a reduction in computational complexity is essential for improving their real-time performance. Ioannis Katramados, Toby P. Breckon |
ICIP | 2 |
| 2011 | Automatic Road Environment ClassificationabstractThe ongoing development autonomous vehicles and adaptive vehicle dynamics present in many modern vehicles has generated a need for road environment classification - i.e., the ability to determine the nature of the current road or terrain environment from an onboard vehicle sensor. In this paper, we investigate the use of a low-cost camera vision solution capable of urban, rural, or off-road classification based on the analysis of color and texture features extracted from a driver's perspective camera view. A feature set based on color and texture distributions is extracted from multiple regions of interest in this forward-facing camera view and combined with a trained classifier approach to resolve two road-type classification problems of varying difficulty - {off-road, on-road} environment determination and the additional multiclass road environment problem of {off-road, urban, major/trunk road and multilane motorway/carriageway}. Two illustrative classification approaches are investigated, and the results are reported over a series of real environment data. An optimal performance of ~90% correct classification is achieved for the {off-road, on-road} problem at a near real-time classification rate of 1 Hz. Isabelle Tang, Toby P. Breckon |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2010 | Object Recognition using 3D SIFT in Complex CT VolumesabstractThe automatic detection of objects within complex volumetric imagery is becoming of increased interest due to the use of dual energy Computed Tomography (CT) scanners as an aviation security deterrent. These devices produce a volumetric image akin to that encountered in prior medical CT work but in this case we are dealing with a complex multi-object volumetric environment including significant noise artefacts. In this work we look at the application of the recent extension to the seminal SIFT approach to the 3D volumetric recognition of rigid objects within this complex volumetric environment. A detailed overview of the approach and results when applied to a set of exemplar CT volumetric imagery is presented. Gregory T. Flitton, Toby P. Breckon, Najla Megherbi Bouallagu |
BMVC | 2 |
| 2010 | A classifier based approach for the detection of potential threats in CT based Baggage ScreeningabstractRecent years have seen increased use of Computed Tomography (CT) based Unaccompanied Baggage and Package Screening (UBPS) systems for luggage examination to ensure air travel security. In this paper we present a research work on developing a system for automatic detection of potential threat items in cluttered 3D CT imagery originating from UBPS systems by combining 3D medical image segmentation techniques with 3D shape classification and retrieval methods. Najla Megherbi Bouallagu, Gregory T. Flitton, Toby P. Breckon |
ICIP | 3 |
| 2009 | Real-Time Traversable Surface Detection by Colour Space Fusion and Temporal Analysis
Ioannis Katramados, Steve Crumpler, Toby P. Breckon |
ICVS | 3 |
| 2008 | Three-Dimensional Surface Relief Completion Via Nonparametric TechniquesabstractCommon 3D acquisition techniques, such as laser scanning and stereo capture, are realistically only 2.5D in nature. Here we consider the automated completion of hidden or missing portions in 3D scenes originally acquired from 2.5D (or 3D) capture. We propose an approach based on the non-parametric propagation of available scene knowledge from the known (visible) scene areas to these unknown (invisible) 3D regions in conjunction with an initial underlying geometric surface completion. Toby P. Breckon, Robert B. Fisher |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Amodal volume completion: 3D visual completion
Toby P. Breckon, Robert B. Fisher |
Comput. Vis. Image Underst. | 1 |