Dermot Kerr

dblp:01/3591 · DBLP profile ↗
← Back
60ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0002-5077-0658ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 6 since 2021Systems, architecture and hardware · 13 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Anisotropic Optical Flow Guided Adaptive Multi-Stage Video Inpainting
abstract
Video inpainting is attracting more attention due to the potential applications of video object removal and video content restoration. Current approaches either use end-to-end methods to generate missing pixels directly or perform indirect transfer for known regions based on motion field guidance. However, such approaches cannot handle both high-resolution images and diverse degrees of scene variation between adjacent video frames, and they cannot achieve clear and accurate inpainting effects for large continuous missing areas. To this end, we propose an adaptive multi-stage interval video inpainting algorithm guided by anisotropic optical flow. First, we customize an optical flow inpainting method guided by single image inpainting, enabling optical flow to maintain a strong self-healing ability over a large range of missing areas. Then, the interval mechanism adaptively determines the required temporal neighbors for missing pixels by assessing video attributes and inpainted optical flow results. After the missing pixels complete the multi-candidate information fusion in their associated temporal neighbors, we obtain spatio-temporally consistent and accurate results. Finally, extensive experiments on the YouTubeVOS, DAVIS, A2D2, and custom datasets show that our proposed approach has achieved state-of-the-art performance with good environmental migration ability.
Lei Rong, Yunzhou Zhang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Multim.5
2025 YOLO-based in-situ Defect Monitoring System for Additive Manufacturing
abstract
In-situ monitoring of an Additive Manufacturing (AM) process is the way to enhance the quality of the components manufactured by it. However, the metal AM processes are complex to monitor because of usage of high energy-based heat source in melting the deposition material, particularly in the arc-based AM processes where arc and sparks make it difficult to capture the deposition. This paper explores the use of a High Dynamic Range (HDR) camera to capture and monitor deposition processes for μ-Plasma Transferred Arc Additive Manufacturing (μP-TAAM) process. Additionally, it proposes the YOLO-based object detection model to assess and monitor the quality of Co-Cr-Mo-4Ti depositions. The research focuses on analysing the performance of YOLOv8l, YOLOv9t, YOLOv9s, YOLOv9m and YOLOv10n models to detect and classify good and bad depositions. It has been found that YOLIv9m gave a strong balance across all evaluation metrics such as highest Recall of 0.983, high precision of 0.98, high mAP50 of 0.994 and high mAP50-95 of 0.848. These findings underscore the model's potential for deployment in in-situ monitoring scenarios.
Deepika Nikam, Maxime Hudon, Sonya A. Coleman, Dermot Kerr, Neelesh Kumar Jain 0001, Sagar Nikam
INDIN4
2024 Advancements in Industrial Visual Inspection: Harnessing Hyperspectral Imaging for Automated Solder Quality Assessment
abstract
This paper presents a groundbreaking advancement in industrial quality control through the development of an automated soldering quality assessment system for circuit boards utilizing hyperspectral imaging (USI) technology. Building upon the transformative capabilities of USI in visual inspection, our research focuses on enhancing the precision and depth of assessment in soldering processes, a critical aspect of electronics manufacturing. By leveraging the unique spectral information captured by HSI, beyond the capabilities of traditional vision systems, our automated solution offers a comprehensive evaluation of solder quality, overcoming challenges posed by similar absorption characteristics of materials. We detail the methodology, algorithms, and integration of HSI into the inspection pipeline, highlighting its effectiveness in detecting defects, ensuring uniformity, and improving overall product quality. The application of this technology extends beyond electronics manufacturing, with potential implications for various industries requiring meticulous quality control. Through this study, we contribute to the ongoing evolution of visual inspection systems, empowering industries with advanced tools for precise and reliable quality assessment.
Trishna Barman, Sonya A. Coleman, Dermot Kerr, Shane Harrigan, Justin Quinn
INDIN3
2024 A Comparative Study of Hough Transform and PCA for Bolt Orientation Detection
abstract
In the fields of manufacturing and robotics, accurately determining the orientation of manufacturing components, such as bolts, is a critical yet challenging problem due to the limitations of existing detection methods. This study introduces a novel methodology for addressing this issue, leveraging traditional computer vision techniques, by proposing a streamlined approach that exploits the inherent geometric properties of bolts for orientation detection. Two methods are presented to ascertain the initial axis angle of the bolt: the Progressive Probabilistic Hough Transform (PPHT) and Principal Component Analysis (PCA). These methods are used in conjunction with a novel tip direction detection approach. The results of the study demonstrate consistent accuracy in angle determination, with PPHT and PCA both achieving angular deviations below ±0.5° in the simple dataset, and PCA showing enhanced robustness in the dataset degraded by shadows, with a maximum error under ±1.5°. This research not only reaffirms the viability of fundamental computer vision techniques in modern robotic applications but also sets a precedent for simple, generalisable, and reliable orientation detection solutions. These methods effectively bridge the gap between highly specialised machine learning systems, which often require tailored, complex models and extensive training data, and more universally applicable, straightforward approaches.
Antonio Gambale, Sonya A. Coleman, Dermot Kerr, Philip J. Vance, Emmett Kerr, Cornelia Fermüller, Yiannis Aloimonos
INDIN3
2024 A Phased-Based Approach to Neuromorphic Audio Recognition
abstract
This paper presents two novel feature representations for neuromorphic audio data. Neuromorphic audio data are considered state-of-the-art when precise time responses are needed while also keeping energy-demands to a minimum. The approaches presented here are based on the concept of phased encoding of neuromorphic data to generate feature representations. One of the approaches enhances on this further by utilising an autoencoder to reduce the dimensionality of the feature representation allowing for increased accuracy in noise-rich environments such as industrial shop floors. The approaches are evaluated against other leading audio feature representation methods using a neuromorphic version of the TIDIGITS database and results demonstrate high accuracy for the proposed approach. We also find that the autoencoder-backed method achieves the best performance compared with the other methods as the dimensionality reduction results in a generalised representation of the feature set which is less sensitive when compared to other methods.
Shane Harrigan, Sonya A. Coleman, Dermot Kerr
INDIN3
2024 Real-Time Human Pose Estimation as a Cost-Effective Solution for the Teleoporation of a 6-Axis Cobot Arm
abstract
This paper explores the application of BlazePose, a monocular human pose estimation (HPE) model, within a teleoperation framework for a UR5 six-axis robot. Achieving teleoperation with only a single RGB camera and a device without a powerful GPU will improve accessibility and cost effectiveness of teleoperation solutions. This study evaluates the 2D pose estimation capabilities of BlazePose for robotic teleoperation tasks. Given the necessity of manipulating the UR5 in three- dimensional space, we implement a 2D-based controller that translates the teleoperator's 2D right hand position within a configurable hand workspace to the corresponding position of the robot's Tool Centre Point (TCP) within the robot's available workspace along two dimensions. The left hand is then utilised for controlling the robot's motion along the third dimension and operating the attached OnRobot RG2 gripper during the pick- and-place task. Additionally, we explore an alternative control paradigm utilising the 3D pose estimation of BlazePose for a more intuitive controller. Two experiments are conducted: the pick-and-place task to assess the 2D-based controller in common robotic tasks, and a hold position task. The hold position task aims to assess the amount of excess movement attributable to the HPE model when utilising the 3D-based controller. The results reveal that while the 2D pose estimation capabilities enable effective teleoperation, the utilisation of 3D estimation results in poor translation to robot control and significant excess motion. These findings underscore the importance of accurate depth estimation in 3D HPE models for precise and reliable teleoperation.
Benn Henderson, Sonya A. Coleman, Dermot Kerr, Justin Quinn, Shane Harrigan
INDIN3
2024 Fingerspelling Classification for Robot Control
abstract
Improvements to human-robot interaction methods could increase the ease of use of robots in manufacturing environments. Many of these environments are noisy and therefore preclude the use of audio communication between humans or in human-robot interactions. Therefore, this paper proposes using a gesture based communication system for robot control. To that end, the VGG16 and VGG19 convolutional neural network (CNN) structures are used for gesture classification along with 3 datasets of American Sign Language (ASL) fingerspelling images. The model performance is evaluated, and modifications made to their parameters to improve performance, before applying them to robot control tasks. The results show that with parameter tuning, test accuracies of up to, 100% are achievable.
Kevin McCready, Sonya A. Coleman, Dermot Kerr, Nazmul H. Siddique, Emmett Kerr, Yiannis Aloimonos, Cornelia Fermüller
INDIN3
2024 Bilateral guidance network for one-shot metal defect segmentation
Dexing Shan, Yunzhou Zhang, Xiaozheng Liu, Sonya A. Coleman, Dermot Kerr
Eng. Appl. Artif. Intell.6
2024 Multibranch Joint Representation Learning Based on Information Fusion Strategy for Cross-View Geo-Localization
abstract
Cross-view geo-localization refers to recognizing images of the same geographic target obtained from different platforms (such as drone-view, satellite-view and ground-view). However, cross-view geo-localization is challenging as image capture using different platforms coupled with extreme viewpoint variations can cause significant changes to the visual image content. Existing methods mainly focus on mining the fine-grained features or the contextual information in neighboring areas, but ignore the complete information of the entire image and the association of contextual information of adjacent regions. Therefore, a multi-branch joint representation learning network model based on information fusion strategies is proposed to solve this cross-view geo-localization problem. Firstly, we obtain feature information from the image through global information fusion branch and local information fusion branch to help the network learn the discernable information in the different images. In addition, a local-guided-global information fusion branch is introduced to make local information assist global features to enhance the learning of potential information in the images. Secondly, we introduced different information fusion strategies in each branch to increase the extraction of contextual information through expanding the global receptive field, thus improving the performance of the model. Finally, a series of experiments is carried out on four prevailing benchmark datasets, namely University-1652, SUES-200, CVUAS and CVACT datasets. The quantitative comparisons from the experiments clearly indicate that the proposed network framework has great performance. For example, compared with some state-of-the-art methods, the quantitative improvements of the R@1 and AP on the University-1652 datasets are 1.91%, 2.18% and 1.55%, 2.99% in both tasks, respectively.
Fawei Ge, Yunzhou Zhang, Yixiu Liu, Guiyuan Wang, Sonya A. Coleman, Dermot Kerr, Li Wang 0160
IEEE Trans. Geosci. Remote. Sens.6
2024 Multilevel Feedback Joint Representation Learning Network Based on Adaptive Area Elimination for Cross-View Geo-Localization
abstract
Cross-view geo-localization refers to the task of matching the same geographic target using images obtained from different platforms, such as drone-view and satellite-view. However, the view angle of images obtained through different platforms will vary greatly, which can bring great challenges to the cross-view geo-localization task. Therefore, we propose a multi-level feedback joint representation learning network based on adaptive area elimination to solve the cross-view geo-localization problem. In our network model, we first process the extracted global features to obtain part-level and patch-level features. We then utilize these features as feedback to the global features to extract the contextual information in the global features and improve the robustness of the extracted features. In addition, as images obtained from different platforms differ, there will always be some interference when matching images. Therefore, we introduce an adaptive area elimination strategy to erase the interference information in the global features and assist the model in obtaining crucial information. On this basis, the feature correlation loss function is designed to constrain learning when using global feature information, thereby eliminating the possible interference, which can improve the network model performance. Finally, a series of experiments is carried out using two well-known benchmarks, namely University-1652 and SUES-200, and the experimental results show that the proposed network model achieves competitive results, thereby demonstrating the effectiveness of proposed model.
Fawei Ge, Yunzhou Zhang, Li Wang 0160, Wei Liu 0022, Yixiu Liu, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Geosci. Remote. Sens.7
2024 DynaQuadric: Dynamic Quadric SLAM for Quadric Initialization, Mapping, and Tracking
abstract
Dynamic SLAM is a key technology for autonomous driving and robotics, and accurate pose estimation of surrounding objects is important for semantic perception tasks. Current quadric SLAM methods are based on the assumption of a static environment and can only reconstruct static quadrics in the scene, which limits their applications in complex dynamic scenarios. In this paper, we propose a visual SLAM system that is capable of reconstructing dynamic objects as quadrics, with a unified framework for jointly optimizing pose estimation, multi-object tracking (MOT), and quadric parameters. We propose a robust object-centric quadric initialization algorithm for both static and moving objects, which decouples the prior estimation of the object pose from the quadric parameters. The object is initialized with a coarse sphere, and quadric parameters are further refined. We design a novel factor graph that tightly optimizes camera pose, object pose, map points and quadric parameters within the sliding window-based optimization. To the best of our knowledge, we are the first to propose a dynamic SLAM that combines quadric representations and MOT in a tightly coupled optimization. We perform qualitative and quantitative experiments on both simulated and real-world datasets, and demonstrate the robustness and accuracy in terms of camera localization, dynamic quadric initialization, mapping and tracking. Our system demonstrates the potential application of object perception with quadric representation in complex dynamic scenes.
Rui Tian 0002, Yunzhou Zhang, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.6
2024 Fast, Robust, Accurate, Multi-Body Motion Aware SLAM
abstract
Simultaneous ego localization and surrounding object motion awareness are significant issues for the navigation capability of unmanned systems and virtual-real interaction applications. Robust and accurate data association at object and feature levels is one of the key factors in solving this problem. However, currently available solutions ignore the complementarity among different cues in the front-end object association and the negative effects of poorly tracked features on the back-end optimization. It makes them not robust enough in practical applications. Motivated by these observations, we make up rigid environment as a unified whole to assist state decoupling by integrating high-level semantic information, ultimately enabling simultaneous multi-states estimation. A filter-based multi-cues fusion object tracker is proposed for establishing more stable object-level data association. Combined with the object’s motion priors, the motion-aided feature tracking algorithm is proposed to improve the feature-level data association performance. Furthermore, a novel state estimation factor graph is designed which integrates a specific feature observation uncertainty model and the intrinsic priors of tracked object, and solved through sliding-window optimization. Our system is evaluated using the KITTI dataset and achieves comparable performance to state-of-the-art object pose estimation systems both quantitatively and qualitatively. We have also validated our system on simulation environment and a real-world dataset to confirm the potential application value in different practical scenarios.
Linghao Yang, Yunzhou Zhang, Rui Tian 0002, Shiwen Liang, You Shen, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.7
2024 CapsLoc3D: Point Cloud Retrieval for Large-Scale Place Recognition Based on 3D Capsule Networks
abstract
Point cloud-based place recognition can be used for global localization in large-scale scenes and loop-closure detection in simultaneous localization and mapping (SLAM) systems in the absence of GPS. Current learning-based approaches aim to extract global and local features from 3D point clouds to encode them as descriptors for point cloud retrieval. The key problems are that the occlusion of point clouds by dynamic objects in the scene affects the point cloud structure, a single perceptual field of the network cannot adequately extract point cloud features, and the correlation between features is not fully utilized. To overcome this, we propose a novel network called CapsLoc3D. We first use the static point cloud generation module to remove the occlusion effects of dynamic objects, and then obtain the point cloud descriptors by processing with the CapsLoc3D network which contains the point spatial transformation module, multi-scale feature fusion module, Capsnet module and a GeM Pooling layer. After validation using the Oxford RobotCar, KITTI, and NEU datasets, experiments show that our method performs better and also has good generalization performance and computational efficiency compared with current state-of-the-art algorithms.
Yunzhou Zhang, Ming Liao, Rui Tian 0002, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.6
2024 Double-Domain Adaptation Semantics for Retrieval-Based Long-Term Visual Localization
abstract
Due to seasonal and illumination variance, long-term visual localization tasks in dynamic environments is a crucial problem in the field of autonomous driving and robotics. At present, image-based retrieval is an effective method to solve this problem. However, it is difficult to completely distinguish changes in the same location over times by relying on content information alone. In order to solve these above problems, a double-domain network model combining semantic information and content information is proposed for visual localization task. In addition, this approach only needs to use the virtual KITTI 2 dataset for training. To reduce the domain difference between real scene and virtual image, the cross-predictive semantic segmentation mechanism is introduced to solve this problem. In addition, the obtained model achieves good domain adaptation and further has well generalization on other real datasets by introducing a domain loss function and a triplet semantic loss function. A series of experiments on the Extended CMU-Seasons dataset and the Oxford RobotCar-Seasons dataset demonstrates that the proposed network model outperformes the state-of-the-art baselines for retrieval-based visual localization in challenging environments.
Fawei Ge, Yunzhou Zhang, Li Wang 0160, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Multim.5
2024 LARNet: Towards Lightweight, Accurate and Real-Time Salient Object Detection
abstract
Salient object detection (SOD) has rapidly developed in recent years, and detection performance has greatly improved. However, the price of these improvements is increasingly complex networks that require more computing resources and sacrifice real-time performance. This makes it difficult to deploy these approaches on devices with limited computing resources (such as mobile phones, embedded platforms, etc.). Considering recently developed lightweight SOD models, their detection and real-time performance are always compromised in demanding practical application scenarios. To solve these problems, we propose a novel lightweight SOD method called LARNet and its corresponding extremely lightweight method LARNet$^{*}$according to application requirements. These methods balance the relationship between lightweight requirements, detection accuracy and real-time performance. First, we propose a saliency backbone network tailored for SOD, which removes the need for pre-training with ImageNet and effectively reduces feature redundancy. Subsequently, we propose a novel context gating module (CGM), which simulates the physiological mechanism of human brain neurons and visual information processing, and realizes the deep fusion of multi-level features at the global level. Finally, the saliency map is output after fusion of multi-level features. Extensive experiments on popular benchmark datasets demonstrate that the proposed LARNet (LARNet$^{*}$) achieves 98 (113) FPS on a GPU and 3 (6) FPS on a CPU. With approximately 680 K (90 K) parameters, the model has significant performance advantages over (extremely) lightweight methods, even surpassing some heavyweight models.
Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Cao Qin, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Multim.6
2023 Time Efficient Micro-Expression Recognition Using Weighted Spatio-Temporal Landmark Graphs
abstract
Micro-expressions have been shown to be effective in understanding the genuine emotions of a person. While many advances have been made in detecting micro-expressions using deep learning, previous studies in recognizing micro-expressions require pre-processing steps and the use of large feature sets resulting in large runtimes and thus have limited applicability in real-world scenarios. In this paper, we propose time-efficient end-to-end framework which uses landmark-based positional features to generate spatio-temporal graphs that can be applied to micro-expression recognition using Graph Convolutional Neural Networks (GCNs). We explore the importance of landmark features and propose a selective feature reduction approach to further improve efficiency. We perform experiments using the SMIC, CASMEII and SAMM datasets and demonstrate that our approach significantly speeds up predictions and delivers results comparable to the state-of-the-art.
Nikin Matharaarachchi, Muhammad Fermi Pasha, Sonya A. Coleman, Dermot Kerr
ICMLA4
2023 SAMLoc: Structure-Aware Constraints With Multi-Task Distillation for Long-Term Visual Localization
abstract
Real-time and robust long-term visual localization is a crucial technology for autonomous driving. Season and illumination variance make this problem more challenging. At present, most of excellent visual localization algorithms cannot run in real-time on devices with limited computing resources. In this paper, we propose SAMLoc, a structure-aware and self-supervised visual localization system, for fast and robust 6-DoF localization. To obtain structural features in the scene, we propose local and global structure-aware constraints using edge information. Then, we integrate the structure-aware constraints into the hierarchical localization network of multi-task distillation, which significantly reduces the feature extraction time while ensuring localization accuracy. As a result, real-time and robust large-scale localization can be achieved on mobile devices. Experimental results on public datasets show that our system can achieve high localization accuracy and have satisfactory real-time performance. Compared with several state-of-the-art visual localization systems, our framework achieves a competitive localization performance.
Jian Ning, Yunzhou Zhang, Sonya A. Coleman, Kunmo Li, Dermot Kerr
ICRA6
2023 BSH-Det3D: Improving 3D Object Detection with BEV Shape Heatmap
abstract
The progress of LiDAR-based 3D object detection has significantly enhanced developments in autonomous driving and robotics. However, due to the limitations of LiDAR sensors, object shapes suffer from deterioration in occluded and distant areas, which creates a fundamental challenge to 3D perception. Existing methods estimate specific 3D shapes and achieve remarkable performance. However, these methods rely on extensive computation and memory, causing imbalances between accuracy and real-time performance. To tackle this challenge, we propose a novel LiDAR-based 3D object detection model named BSH-Det3D, which applies an effective way to enhance spatial features by estimating complete shapes from a bird's eye view (BEV). Specifically, we design the Pillar-based Shape Completion (PSC) module to predict the probability of occupancy whether a pillar contains object shapes. The PSC module generates a BEV shape heatmap for each scene. After integrating with heatmaps, BSH-Det3D can provide additional information in shape deterioration areas and generate high-quality 3D proposals. We also design an attention-based densification fusion module (ADF) to adaptively associate the sparse features with heatmaps and raw points. The ADF module integrates the advantages of points and shapes knowledge with negligible overheads. Extensive experiments on the KITTI benchmark achieve state-of-the-art (SOTA) performance in terms of accuracy and speed, demonstrating the efficiency and flexibility of BSH-Det3D. The source code is available on https://github.com/mystorm16/BSH-Det3D.
You Shen, Yunzhou Zhang, Yanmin Wu, Zhenyu Wang 0010, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IROS7
2023 A novel seminar learning framework for weakly supervised salient object detection
Yan Liu 0080, Yunzhou Zhang, Zhenyu Wang 0010, Fei Yang 0007, Sonya A. Coleman, Dermot Kerr
Eng. Appl. Artif. Intell.7
2023 WUSL-SOD: Joint weakly supervised, unsupervised and supervised learning for salient object detection
Yan Liu 0080, Yunzhou Zhang, Zhenyu Wang 0010, Sonya A. Coleman, Dermot Kerr
Neural Comput. Appl.7
2023 MMPL-Net: multi-modal prototype learning for one-shot RGB-D segmentation
Dexing Shan, Yunzhou Zhang, Xiaozheng Liu, Shitong Liu, Sonya A. Coleman, Dermot Kerr
Neural Comput. Appl.6
2023 ELWNet: An Extremely Lightweight Approach for Real-Time Salient Object Detection
abstract
Existing lightweight salient object detection (SOD) methods aim to solve the problem of high computational costs that is prevalent with heavyweight methods. However, compared with heavyweight methods, the detection accuracy of lightweight methods is greatly reduced while real-time performance is not significantly improved. Therefore, we aim to establish a trade off between computational cost and detection performance by improving the network efficiency. We propose a fast and extremely lightweight end-to-end wavelet neural network (ELWNet) for real-time salient object detection. ELWNet can achieve salient object detection and segmentation at approximately 70FPS (GPU), 19FPS (CPU) with 76K parameters and 0.38G FLOPs. We introduce wavelet transform theory into a neural network, proposing a wavelet transform module (WTM), a wavelet transform fusion module (WTFM), a novel feature residual mechanism, and construct an efficient architecture. The wavelet transform theory is integrated into the neural network to realize the interaction between the features in the frequency and the time domain. Meanwhile, ELWNet does not rely on a pre-trained model, which significantly reduces redundant features. We validate the performance of ELWNet using five well-known datasets, and demonstrate state-of-the-art performance compared with 24 other SOD models in terms of being lightweight, detection accuracy and real-time capabilities. Our method maintains high detection performance while reducing the number of model parameters by approximately 99% compared with heavyweight methods.
Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Delong Zhu 0001, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Circuits Syst. Video Technol.6
2023 Unseen-Material Few-Shot Defect Segmentation With Optimal Bilateral Feature Transport Network
abstract
Industrial defect segmentation is important to ensure product quality and production safety. The main challenges in industrial applications are insufficient defect samples, large intraclass variation, and the interference of background information. However, most current texture defect segmentation methods rely on large-scale datasets and can only deal with one specific type of texture defect, which reduces the application efficiency and application scope of defect segmentation algorithms. To this end, we propose an optimal bilateral feature transport network (OBFTNet) for few-shot texture defect segmentation, which can accurately segment texture defects in multiple unseen materials (domains), such as steel, wood, and leather. OBFTNet can perform bilateral prediction for background and defect regions of unseen material by dynamically predicting task-specific semantic correspondences conditioned on a small guidance set. Specifically, we introduce background images (defect-free images) as supplementary learning information for reverse prediction and model the semantic correspondence between the guidance (support and background images) and the query images in few-shot segmentation as an optimal bilateral feature transport problem and generate a set of optimal bilateral correlation tensors. Using 4-D and 2-D convolutions, the model gradually reduces optimal bilateral correlation tensors to precise segmentation masks. Experimental results show that our proposed method outperforms several state-of-the-art techniques with very few labeled samples and the method generalizes well to industrial defects on unseen materials.
Dexing Shan, Yunzhou Zhang, Sonya A. Coleman, Dermot Kerr, Shitong Liu, Ziqiang Hu
IEEE Trans. Ind. Informatics4
2023 Object SLAM With Robust Quadric Initialization and Mapping for Dynamic Outdoors
abstract
Object SLAM is a popular approach for autonomous driving and robotics, but accurate object perception in outdoor environments remains a challenge. State-of-the-art object SLAM algorithms rely on assumptions and are sensitive to observation noise, limiting their application in real-world scenarios. To address these challenges, we propose a novel object SLAM system that utilizes a quadric initialization algorithm based on constrained quadric optimization, which does not rely on planar assumptions and is robust to partial observations. Additionally, we introduce an automatic object data association algorithm capable of detecting motion states while associating objects across frames. To further enhance the accuracy of the quadric mapping, an extra thread is used to refine the ellipsoid parameters within a local sliding window composed of keyframes. Our system utilizes a joint optimization framework that optimizes camera poses, object landmarks, and point clouds in the local mapping thread for further global optimization while maintaining a consistent map. Experimental results on the real-world KITTI dataset show that the proposed system is more robust and significantly outperforms current state-of-the-art methods in quadric initialization and mapping in outdoor scenarios. Moreover, our system achieves real-time performance, making it suitable for practical applications.
Rui Tian 0002, Yunzhou Zhang, Zhenzhong Cao, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.7
2023 Structure-Aware Feature Disentanglement With Knowledge Transfer for Appearance-Changing Place Recognition
abstract
Long-term visual place recognition (VPR) is challenging as the environment is subject to drastic appearance changes across different temporal resolutions, such as time of the day, month, and season. A wide variety of existing methods address the problem by means of feature disentangling or image style transfer but ignore the structural information that often remains stable even under environmental condition changes. To overcome this limitation, this article presents a novel structure-aware feature disentanglement network (SFDNet) based on knowledge transfer and adversarial learning. Explicitly, probabilistic knowledge transfer (PKT) is employed to transfer knowledge obtained from the Canny edge detector to the structure encoder. An appearance teacher module is then designed to ensure that the learning of appearance encoder does not only rely on metric learning. The generated content features with structural information are used to measure the similarity of images. We finally evaluate the proposed approach and compare it to state-of-the-art place recognition methods using six datasets with extreme environmental changes. Experimental results demonstrate the effectiveness and improvements achieved using the proposed framework. Source code and some trained models will be available at http://www.tianshu.org.cn.
Cao Qin, Yunzhou Zhang, Yingda Liu, Delong Zhu 0001, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Neural Networks Learn. Syst.6
2022 Deep Learning for Semiconductor Defect Classification
abstract
Automated inspection has become a vital part of quality control during semiconductor wafer production. Current processes are focused on finding defects via variation from a ‘golden' image using pixel to pixel comparisons or utilization of opaque neural network-based approaches. In this paper we present an approach which uses deep learning methods to classify defects on semiconductor die images and show the experimental steps taken in order to produce a highly accurate system based on previous models.
Terence Sweeney, Sonya A. Coleman, Dermot Kerr
INDIN3
2022 Semantic Topological Descriptor for Loop Closure Detection within 3D Point Clouds In Outdoor Environment
abstract
Loop closure detection has the potential to correct the drift of trajectories and build a global consistent map in LiDAR SLAM, however it remains a challenging problem in outdoor environment due to the sparsity of 3D point clouds data, large-scale scenes and moving objects. Inspired by the way humans perceive the environment through recognizing objects and identifying their relations, this paper presents a novel descriptor that contains semantic and topological information for loop closure detection. Unlike most existing methods that extract features from the raw point clouds or use all semantic objects, we directly discard point clouds representing pedestrians and vehicles after semantic segmentation. Then, we propose a semantic topological graph representation from the remaining point clouds and convert this graph into a descriptor. Additionally, we propose a two-stage algorithm for matching descriptors to efficiently determine the loop. Our method has been extensively evaluated using the KITTI dataset and outperforms state-of-the-art methods, especially in the challenging situations such as viewpoint changes and dynamic scenes.
Ming Liao, Yunzhou Zhang, Sonya A. Coleman, Dermot Kerr
IROS6
2022 VAC-Net: Visual Attention Consistency Network for Person Re-identification
abstract
Person re-identification (ReID) is a crucial aspect of recognising pedestrians across multiple surveillance cameras. Even though significant progress has been made in recent years, the viewpoint change and scale variations still affect model performance. In this paper, we observe that it is beneficial for the model to handle the above issues when boost the consistent feature extraction capability among different transforms (e.g., flipping and scaling) of the same image. To this end, we propose a visual attention consistency network (VAC-Net). Specifically, we propose Embedding Spatial Consistency (ESC) architecture with flipping, scaling and original forms of the same image as inputs to learn a consistent embedding space. Furthermore, we design an Input-Wise visual attention consistent loss (IW-loss) so that the class activation maps(CAMs) from the three transforms are aligned with each other to enforce their advanced semantic information remains consistent. Finally, we propose a Layer-Wise visual attention consistent loss (LW-loss) to further enforce the semantic information among different stages to be consistent with the CAMs within each branch. These two losses can effectively improve the model to address the viewpoint and scale variations. Experiments on the challenging Market-1501, DukeMTMC-reID, and MSMT17 datasets demonstrate the effectiveness of the proposed VAC-Net.
Yunzhou Zhang, Shangdong Zhu, Yixiu Liu, Sonya A. Coleman, Dermot Kerr
ICMR6
2022 Complementary characteristics fusion network for weakly supervised salient object detection
Yan Liu 0080, Yunzhou Zhang, Zhenyu Wang 0010, Fei Yang 0007, Cao Qin, Sonya A. Coleman, Dermot Kerr
Image Vis. Comput.8
2022 TF-SOD: a novel transformer framework for salient object detection
Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Sonya A. Coleman, Dermot Kerr
Neural Comput. Appl.6
2022 Data Assimilation Network for Generalizable Person Re-Identification
abstract
In this paper, a data assimilation network is proposed to tackle the challenges of domain generalization for person re-identification (ReID). Most of the existing research efforts only focus on single-dataset issues, and the trained models are difficult to generalize to unseen scenarios. This paper presents a distinctive idea to improve the generality of the model by assimilating three types of images: style-variant images, misaligned images and unlabeled images. The latter two are often ignored in the previous domain generalization ReID studies. In this paper, a non-local convolutional block attention module is designed for assimilating the misaligned images, and an attention adversary network is introduced to correct it. A progressive augmented memory is designed for assimilating the unlabeled images by progressive learning. Moreover, we propose an attention adversary difference loss for attention correction, and a labeling-guide discriminative embedding loss for progressive learning. Rather than designing a specific feature extractor that is robust to style shift as in most previous domain generalization work, we propose a data assimilation meta-learning procedure to train the proposed network, so that it learns to assimilate style-variant images. It is worth mentioning that we add an unlabeled augmented dataset to the source domain to tackle the domain generalization ReID tasks. Extensive experiments demonstrate that our approach significantly outperforms the state-of-the-art domain generalization methods.
Yixiu Liu, Yunzhou Zhang, Bir Bhanu, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Circuits Syst. Video Technol.5
2022 Editorial Biologically Learned/Inspired Methods for Sensing, Control, and Decision
Yongduan Song 0001, Jennie Si, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Neural Networks Learn. Syst.4
2021 Computational Approach to Identifying Contrast-Driven Retinal Ganglion Cells
Richard Gault, Philip J. Vance, T. Martin McGinnity, Sonya A. Coleman, Dermot Kerr
ICANN (3)5
2021 Accurate and Robust Scale Recovery for Monocular Visual Odometry Based on Plane Geometry
abstract
Scale ambiguity is a fundamental problem in monocular visual odometry. Typical solutions include loop closure detection and environment information mining. For applications like self-driving cars, loop closure is not always available, hence mining prior knowledge from the environment becomes a more promising approach. In this paper, with the assumption of a constant height of the camera above the ground, we develop a light-weight scale recovery framework leveraging an accurate and robust estimation of the ground plane. The framework includes a ground point extraction algorithm for selecting high-quality points on the ground plane, and a ground point aggregation algorithm for joining the extracted ground points in a local sliding window. Based on the aggregated data, the scale is finally recovered by solving a least-squares problem using a RANSAC-based optimizer. Sufficient data and robust optimizer enable a highly accurate scale recovery. Experiments on the KITTI dataset show that the proposed framework can achieve state-of-the-art accuracy in terms of translation errors, while maintaining competitive performance on the rotation error. Due to the light-weight design, our framework also demonstrates a high frequency of 20 Hz on the dataset.
Rui Tian 0002, Yunzhou Zhang, Delong Zhu 0001, Shiwen Liang, Sonya A. Coleman, Dermot Kerr
ICRA6
2021 Multi-level cross-view consistent feature learning for person re-identification
Yixiu Liu, Yunzhou Zhang, Bir Bhanu, Sonya A. Coleman, Dermot Kerr
Neurocomputing5
2021 A visual place recognition approach using learnable feature map filtering and graph attention networks
Cao Qin, Yunzhou Zhang, Yingda Liu, Sonya A. Coleman, Huijie Du, Dermot Kerr
Neurocomputing6
2021 MFC-Net : Multi-feature fusion cross neural network for salient object detection
Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Shichang Liu, Sonya A. Coleman, Dermot Kerr
Image Vis. Comput.6
2020 Neural Coding Strategies for Event-Based Vision Data
abstract
Neural coding schemes are powerful tools used within neuroscience. This paper introduces three different neural coding scheme formations for event-based vision data which are designed to emulate the neural behaviour exhibited by neurons under stimuli. Presented are phase-of-firing and two sparse neural coding schemes. It is determined that machine learning approaches, i.e. Convolutional Neural Network combined with a Stacked Autoencoder network, produce powerful descriptors of the patterns within events. These coding schemes are deployed in an existing action recognition template and evaluated using two popular event-based data sets.
Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001
ICASSP3
2020 Post-Stimulus Time-Dependent Event Descriptor
abstract
Event-based image processing is a relatively new domain in the field of computer vision. Much research has been carried out on adapting event-based data to comply with established techniques from frame-based computer vision. On the contrary, this paper presents a descriptor which is designed specifically for direct use with event-based data and therefore can be considered to be a pure event-based vision descriptor as it only uses events emitted from event-based vision devices without transforming the data to accommodate frame-based vision techniques. This novel descriptor is known as the Post-stimulus Time-dependent Event Descriptor (P-TED). P-TED is comprised of two features extracted from event data which describe motion and the underlying pattern of transmission respectively. Furthermore a framework is presented which leverages the P-TED descriptor to classify motions within event data. This framework is compared against another state-of-the-art event-based vision descriptor as well as an established frame-based approach.
Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001
ICIP3
2020 Technical Indicators and Prediction for Energy Market Forecasting
abstract
Machine learning usage for forecasting is popular in financial trading, particularly for stock price prediction and this is often combined with technical indicators to extract key predictive indicators from large time series trading datasets. Energy market trading data have similar characteristics to financial trading data, therefore deriving technical indicators specifically for electricity prices will help predict future prices and reduce trading costs. We have derived eight technical indicators for the Integrated Single Electricity Market (ISEM) energy market in Ireland using hourly electricity price data over the period February 2019 until November 2019. Technical indicator based models were obtained by using machine learning regression algorithms (Extreme Gradient [XG] Boost, Random Forest, and Gradient Boosting) trained with the proposed novel technical indicators. The results of the technical indicator models were compared against the baseline model (raw price data only) to see if using technical indicators as inputs improves model performance. We conclude that electricity prices can be accurately predicted using the proposed technical indicators.
Catherine McHugh, Sonya A. Coleman, Dermot Kerr
ICMLA3
2020 Reducing-Over-Time Tree for Event-based Data
abstract
This paper presents a novel Reducing-Over-Time (ROT) binary tree structure for event-based vision data and subtypes of the tree structure. A framework is presented using ROT, that takes advantage of the self-balancing and self-pruning nature of the tree structure to extract spatial-temporal information. The ROT framework is paired with an established motion classification technique and performance is evaluated against other state-of-the-art techniques using four datasets. Additionally, the ROT framework as a processing platform is compared with other event-based vision processing platforms in terms of memory usage and is found to be one of the most memory efficient platforms available.
Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001
ICPR3
2020 Towards real-time activity recognition
abstract
Activity recognition relates to the automatic visual detection and interpretation of human behaviour and is emerging as an active domain of computer vision. It has important applications such as identifying individuals who are at risk of suicide in public locations such as bridges or railway stations. These individuals are known to exhibit easily observable activities and behaviours such as pacing, looking up and down the railway tracks, and leaving objects on the platform. In order to detect these behaviours, an approach to individual person activity recognition is needed which can run in real time and monitor multiple individuals in parallel. We present a method for human activity recognition using skeletal keypoints and investigate how using varying sample rates and sequence lengths impacts accuracy. The results show that for any given sequence length, optimising the sample rate can result in an overall increase in classification accuracy and improvement in run-time. Results demonstrate that finding the optimal time period over which to sample frames is more important than simply decreasing the number of frames sampled. Further, we show that keypoint based activity recognition approaches outperform other state of the art approaches. Finally, we show that this approach is fast enough for real time activity recognition when up to 14 people are present in the image whilst maintaining a high degree of accuracy.
Shane Reid, Philip J. Vance, Sonya A. Coleman, Dermot Kerr, Siobhan O'Neill
IPAS4
2020 EAO-SLAM: Monocular Semi-Dense Object SLAM Based on Ensemble Data Association
abstract
Object-level data association and pose estimation play a fundamental role in semantic SLAM, which remain unsolved due to the lack of robust and accurate algorithms. In this work, we propose an ensemble data associate strategy for integrating the parametric and nonparametric statistic tests. By exploiting the nature of different statistics, our method can effectively aggregate the information of different measurements, and thus significantly improve the robustness and accuracy of data association. We then present an accurate object pose estimation framework, in which an outliers-robust centroid and scale estimation algorithm and an object pose initialization algorithm are developed to help improve the optimality of pose estimation results. Furthermore, we build a SLAM system that can generate semi-dense or lightweight object-oriented maps with a monocular camera. Extensive experiments are conducted on three publicly available datasets and a real scenario. The results show that our approach significantly outperforms state-of-the-art techniques in accuracy and robustness. The source code is available on https://github.com/yanmin-wu/EAO-SLAM.
Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Yonghui Feng, Sonya A. Coleman, Dermot Kerr
IROS6
2020 Multi-level and multi-scale horizontal pooling network for person re-identification
Yunzhou Zhang, Shuangwei Liu, Sonya A. Coleman, Dermot Kerr
Multim. Tools Appl.5
2019 Adversarially Erased Learning for Person Re-identification by Fully Convolutional Networks
abstract
The generalization ability of deep person re-identification networks is subject to inadequate person data and occlusions. To relieve this dilemma, we propose a feature-level augmentation strategy, Adversarially Erased Learning Module (AELM), using two adversarial classifiers. Specifically, we utilize a classifier to identify discriminative regions and erase them to increase the variant of features. Meanwhile, we input the erased feature maps to another classifier to discover new body regions, which effectively resist occlusion of key parts. To easily perform end-to-end training for AELM, we propose a novel Identity model based on Fully Convolutional Networks (IFCN) to directly obtain body response heatmap during the forward pass by selecting corresponding class-specific feature map. Thus, the discriminative regions can be identified and erased in a convenient way. Moreover, to capture discriminative region for AELM, we present a Complementary Attention Module (CoAM) combined with channel and spatial attention to automatically focus on which feature types and positions are meaningful in the feature maps. In this paper, CoAM and AELM are cascaded into one module which is applied to the outputs of different convolutional layers to integrate mid- and high-level semantic features. Experimental results on three challenging benchmarks demonstrate the effectiveness of the proposed method.
Shuangwei Liu, Yunzhou Zhang, Sonya A. Coleman, Dermot Kerr, Shangdong Zhu
IJCNN5
2019 Computational modelling of salamander retinal ganglion cells using machine learning approaches
Gautham P. Das, Philip J. Vance, Dermot Kerr, Sonya A. Coleman, T. Martin McGinnity, Jian K. Liu
Neurocomputing3
2018 Sensor-based Vital Sign Monitoring, Analysis and Visualisation for Ageing in Place
abstract
With the ever-increasing global population and average life expectancy, care homes and care at home services are continuously being stretched beyond capacity. Recent developments in tactile sensing have enabled robot systems to measure human vital signs such as beats per minute (BPM), Respiratory Rate (RR) and Capillary Refill Time (CRT). Using robotic systems to measure vital sign data in the home of an elderly or disabled person would greatly assist medical and health services. This paper proposes the use of a vital sign measuring robotic system together with Cloud computing to intelligently process big data and ascertain the current health status of the service user without the need to expose their identity or burden health professionals. Furthermore, a method that enables medical professionals to visualise the data for a complete geographical region as well as for individual patients is presented and hence we provide details of a closed loop system to support ageing-in-place.
E. P. Kerr, Sonya A. Coleman, Dermot Kerr, Philip J. Vance, Bryan Gardiner, Chengdong Wu 0001
IJCNN3
2018 Investigation into Sub-Receptive Fields of Retinal Ganglion Cells with Natural Images
abstract
Determining the receptive field of a retinal ganglion cell is critically important when formulating a computational model that maps the relationship between the stimulus and response. This process is traditionally undertaken using reverse correlation to estimate the receptive field. By stimulating the retina with artificial stimuli, such as alternating checkerboards, bars or gratings and recording the neural response it is possible to estimate the cell’s receptive field by analysing the stimuli that produced the response. Artificial stimuli such as white noise is known to not stimulate the full range of the cell’s responses. By using natural image stimuli, it is possible to estimate the receptive field and obtain a resulting model that more accurately mimics the cells’ responses to natural stimuli. This paper extends on previous work to seek further improvements in estimating a ganglion cell’s receptive field by considering that the receptive field can be divided into subunits. It is thought that these subunits may relate to receptive fields which are associated with bipolar retinal cells. The findings of this preliminary study show that by using subunits to define the receptive field we achieve a significant improvement over existing approaches when deriving computational models of the cell’s response.
Philip J. Vance, Gautham P. Das, Sonya A. Coleman, Dermot Kerr, Emmett Kerr, T. Martin McGinnity
IJCNN4
2018 Biologically Inspired Intensity and Depth Image Edge Extraction
abstract
In recent years, artificial vision research has moved from focusing on the use of only intensity images to include using depth images, or RGB-D combinations due to the recent development of low-cost depth cameras. However, depth images require a lot of storage and processing requirements. In addition, it is challenging to extract relevant features from depth images in real time. Researchers have sought inspiration from biology in order to overcome these challenges resulting in biologically inspired feature extraction methods. By taking inspiration from nature, it may be possible to reduce redundancy, extract relevant features, and process an image efficiently by emulating biological visual processes. In this paper, we present a depth and intensity image feature extraction approach that has been inspired by biological vision systems. Through the use of biologically inspired spiking neural networks, we emulate functional computational aspects of biological visual systems. The results demonstrate that the proposed bioinspired artificial vision system has increased performance over existing computer vision feature extraction approaches.
Dermot Kerr, Sonya A. Coleman, T. Martin McGinnity
IEEE Trans. Neural Networks Learn. Syst.1
2018 Bioinspired Approach to Modeling Retinal Ganglion Cells Using System Identification Techniques
abstract
The processing capabilities of biological vision systems are still vastly superior to artificial vision, even though this has been an active area of research for over half a century. Current artificial vision techniques integrate many insights from biology yet they remain far-off the capabilities of animals and humans in terms of speed, power, and performance. A key aspect to modeling the human visual system is the ability to accurately model the behavior and computation within the retina. In particular, we focus on modeling the retinal ganglion cells (RGCs) as they convey the accumulated data of real world images as action potentials onto the visual cortex via the optic nerve. Computational models that approximate the processing that occurs within RGCs can be derived by quantitatively fitting the sets of physiological data using an input-output analysis where the input is a known stimulus and the output is neuronal recordings. Currently, these input-output responses are modeled using computational combinations of linear and nonlinear models that are generally complex and lack any relevance to the underlying biophysics. In this paper, we illustrate how system identification techniques, which take inspiration from biological systems, can accurately model retinal ganglion cell behavior, and are a viable alternative to traditional linear-nonlinear approaches.
Philip J. Vance, Gautham P. Das, Dermot Kerr, Sonya A. Coleman, T. Martin McGinnity, Tim Gollisch, Jian K. Liu
IEEE Trans. Neural Networks Learn. Syst.3
2015 Modelling retinal ganglion cells using self-organising fuzzy neural networks
abstract
Even though artificial vision has been in development for over half a century it still fares poorly when compared to biological vision. The processing capabilities of biological visual systems are vastly superior in terms of power, speed, and performance. Inspired by this robust performance artificial vision systems have sought to take inspiration from biology by modeling aspects of biological vision systems. Existing computational models of visual neurons can be derived by quantitatively fitting particular sets of physiological data using an input-output analysis where a known input is given to the system and its output is recorded. These models need to capture the full spatio-temporal description of neuron behaviour under natural viewing conditions. In this work we use state-of-the-art fuzzy neural network techniques to accurately model the responses of retinal ganglion cells. We illustrate how a self-organising fuzzy neural network can accurately model ganglion cell behaviour, and are a viable alternative to traditional system identification techniques.
Scott McDonald 0003, Dermot Kerr, Sonya A. Coleman, Philip J. Vance, T. Martin McGinnity
IJCNN2
2015 Modelling of a retinal ganglion cell with simple spiking models
abstract
Modelling aspects of the human vision system, including the retina, is difficult due to insufficient knowledge about the internal components, organisation and complexity of the interactions within the system. Retinal ganglion cells are considered a core component of the human visual system as they convey the accumulated data as action potentials onto the optic nerve. Current techniques capable of mapping this input-output response involve computational combinations of linear and nonlinear models that are generally complex and lack any relevance to the underlying biophysics. This paper aims to model a retinal ganglion cell with a simple spiking neuron combined with a pre-processing method, which accounts for the preceding retinal neural structure. Performance of the models is compared with the spike responses obtained in the electrophysiological recordings from a mammalian retina subjected to visual stimulation.
Philip J. Vance, Sonya A. Coleman, Dermot Kerr, Gautham P. Das, T. Martin McGinnity
IJCNN3
2015 A biologically inspired spiking model of visual processing for image feature detection
Dermot Kerr, T. Martin McGinnity, Sonya A. Coleman, Marine Clogenson
Neurocomputing1
2013 Biologically inspired intensity and range image feature extraction
abstract
The recent development of low cost cameras that capture 3-dimensional images has changed the focus of computer vision research from using solely intensity images to the use of range images, or combinations of RGB, intensity and range images. The low cost and widespread availability of the hardware to capture these images has realised many possible applications in areas such as robotics, object recognition, surveillance, manipulation, navigation and interaction. Given the large volumes of data in range images, processing and extracting the relevant information from the images in real time becomes challenging. To achieve this, much research has been conducted in the area of bio-inspired feature extraction which aims to emulate the biological processes used to extract relevant features, reduce redundancy, and process images efficiently. Inspired by the behaviour of biological vision systems, an approach is presented for extracting important features from intensity and range images, using biologically inspired spiking neural networks in order to model aspects of the functional computational capabilities of the visual system.
Dermot Kerr, Sonya A. Coleman, T. Martin McGinnity, Marine Clogenson
IJCNN1
2012 A novel approach to robot vision using a hexagonal grid and spiking neural networks
abstract
Many robots use range data to obtain an almost 3-dimensional description of their environment. Feature driven segmentation of range images has been primarily used for 3D object recognition, and hence the accuracy of the detected features is a prominent issue. Inspired by the structure and behaviour of the human visual system, we present an approach to feature extraction in range data using spiking neural networks and a biologically plausible hexagonal pixel arrangement. Standard digital images are converted into a hexagonal pixel representation and then processed using a spiking neural network with hexagonal shaped receptive fields; this approach is a step towards developing a robotic eye that closely mimics the human eye. The performance is compared with receptive fields implemented on standard rectangular images. Results illustrate that, using hexagonally shaped receptive fields, performance is improved over standard rectangular shaped receptive fields.
Dermot Kerr, Sonya A. Coleman, T. Martin McGinnity, Qingxiang Wu, Marine Clogenson
IJCNN1
2011 Corner detection on hexagonal pixel based images
abstract
Corner detection is used in many computer vision applications that require fast and efficient feature matching. In addition, hexagonal pixel based images have been recently investigated for image capture and processing due to their ability to represent curved structures that are common in real images better than traditional rectangular pixel based images. Therefore, we present an approach to corner detection on hexagonal images and demonstrate that accuracy is comparable to well-known existing corner detectors applied to rectangular pixel based images.
Si Jing Liu, Sonya A. Coleman, Dermot Kerr, Bryan W. Scotney, Bryan Gardiner
ICIP3
2011 Biologically inspired edge detection
abstract
Inspired by the structure and behaviour of the human visual system, we present an approach to edge detection using spiking neural networks and a biologically plausible hexagonal pixel arrangement. Standard digital images are converted into a hexagonal pixel representation and then processed using a spiking neural network with hexagonal shaped receptive fields. The performance is compared with receptive fields implemented on standard rectangular images. Results illustrate that, using hexagonal shaped receptive fields, performance is improved over standard rectangular shaped receptive fields.
Dermot Kerr, Sonya A. Coleman, T. Martin McGinnity, Qingxiang Wu, Marine Clogenson
ISDA1
2008 Interest point detection on incomplete images
abstract
Use of incomplete image data has become a prominent research issue in recent years, driven by the development of space variant image sensors. Whilst image reconstruction techniques have been developed that enable the subsequent use of standard image processing algorithms, the development of image processing algorithms that can be applied directly to incomplete image data has received less attention. The problem of interest point detection for incomplete images is addressed by presenting an algorithm that can be applied directly to incomplete image data without the requirement of image reconstruction, and the accurate performance of the algorithm is illustrated through visual results and ROC curves.
Dermot Kerr, Bryan W. Scotney, Sonya A. Coleman
ICIP1
2007 Concurrent Edge and Corner Detection
abstract
To enable fast reliable feature matching or tracking in scenes, features need to be discrete and meaningful, and hence corner detection is often used for this purpose. However, to obtain a higher level description of an image, such as identification of objects, additional information such as edges is required, and more recently detectors have been proposed that find both edges and corners. We present a combined operator, enabling edge and corner detection to be achieved concurrently. We demonstrate that accuracy is comparable to well-known existing corner detectors and edge detectors, and, as standard post-smoothing of the corner map is not required, significantly reduced computation time can be achieved.
Sonya A. Coleman, Dermot Kerr, Bryan W. Scotney
ICIP (5)2
2006 A Graph Theoretic Approach to Direct Processing of Sparse Unwarped Panoramic Images
abstract
The use of omnidirectional cameras has had a significant impact on the success of vision systems for video surveillance and autonomous robot navigation. Typically images obtained from such cameras are transformed to sparse panoramic images that are interpolated prior to low level image processing. We present a graph theoretic approach that enables image processing techniques, principally feature extraction, to be performed directly on sparse panoramic images, avoiding the need for image interpolation. We thus aim to reduce the computational overheads of processing images arising from omnidirectional cameras, whilst retaining accuracy sufficient for application to real-time robot vision.
Bryan W. Scotney, Sonya A. Coleman, Dermot Kerr
ICIP3