VLDB 2026 Research / reviewers in the wild / expert
David Moloney
dblp:76/3181
· DBLP profile ↗
22ranked-venue papers
5as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Balancing Accuracy and Efficiency: Navigating The Trade-off Between Machine-Readable Code Detection and Data Size ReductionabstractDetecting machine-readable codes in industrial manufacturing is a critical yet challenging task, as it directly impacts both defect detection and broader visual anomaly detection. The complexity of visual data, coupled with high-speed production processes generating vast volumes of information, makes accurate and efficient detection essential for ensuring product quality and operational efficiency. This paper takes significant steps toward addressing these challenges by identifying and categorizing key issues related to machine-readable code defect detection, introducing the first publicly available and challenging dataset as a benchmark for future studies, and exploring techniques to enhance both detection performance and data storage efficiency. Experimental results demonstrate that leveraging the YUV color space, combined with data compression, significantly improves detection accuracy while minimizing storage requirements. This work highlights the importance of balancing data complexity, storage optimization, and detection reliability, laying a strong foundation for future advancements in defect detection, anomaly identification, and cost-efficient industrial automation. Imen Jegham, Ons Loukil, Besma Guesmi, David Moloney |
CoDIT | 4 |
| 2024 | EoFNets: EyeonFlare Networks to predict solar flare using Temporal Convolutional Network (TCN)abstractSolar Active Regions are characterised by their intense magnetic activity, which often leads to solar phenomena such as solar flares, and coronal mass ejections (CMEs). With the recent advancement of computing technologies and the huge integration of Artificial Intelligence (AI), many approaches have been proposed for forecasting solar eruptions using machine learning. In this study, we propose the use of a Temporal Convolutional Network (TCN) for predicting whether an active region will be flaring in a specific window of time and defining the flare class. The dataset is categorised into three different subsets based on the flare class and trained separately with the same TCN architecture to apply late fusion. The proposed solar flare prediction ensemble (EoFNets) is based on both the physical characteristics of the active region (EoFPhyNet) and geometric features (EoFGeoNet). Experimental results show that TCN outperforms long short-term memory (LSTM) in three cases. Our main aim is to deploy deep-learning-based approaches onboard for faster and more accurate real-time monitoring as well as leveraging the higher sampling rates for improved time-series predictions. Many major benefits can be realised if the deep learning models can be implemented onboard, including a sizeable reduction in the volume of downlinked data, and improved system latency. However, implementing deep learning models in space can be a critical task, as most approaches require high computational and memory resources, both of which are limited in typical spacecraft onboard data handling systems. Nevertheless, the EoFNets network outlined in this paper has been optimised to fit the resource constraints of a space platform deployed at the extreme edge far from Earth. Two low-power hardware targets are considered, namely the IntelMovidius MyriadX and Rockchip RK3588S. To the best of our knowledge, this is the first time that such a TCN network has been proposed for solar flare forecasting. Besma Guesmi, Jinen Daghrir, David Moloney, Carlos Urbina Ortega, Gianluca Furano, Giuseppe Mandorlo, Elena Hervas-Martin, José Luis Espinosa-Aranda |
CoDIT | 3 |
| 2024 | Hardware-aware, deep-learning approaches for image denoising and star detection for star tracker sensorabstractIn recent years, Deep Neural Networks (DNNs) approaches have outperformed traditional techniques for several computer vision problems. This has been made possible by the increase of computational resources represented by Graphical Processing Units (GPU) that allow training using large datasets and the availability of deep learning accelerators for inference. On the other hand, the attitude determination accuracy requirements for spacecraft are increasing. The most accurate attitude determination sensor for spacecraft is the so-called star sensor or star tracker. With the increase in low-cost satellite platforms such as CubeSats, research into the improvement of star sensor accuracy for low-power and low-cost sensor architectures remains a relevant subject. In this context, we examine several methods for noise reduction and star detection for improving centroiding performance. More specifically, an efficient and robust denoising method for star images using an Auto-Encoder (AE) is proposed. This method enhances the image quality for systems sensitive to noise. Furthermore, an accurate and lightweight algorithm based on an existing YOLO (You Only Look Once) architecture is proposed to detect the location of stars in the image. In this work, the YOLO bounding boxes are used to describe the space region around the stars. Subsequently, the star centroid within the bounding box is computed using the COG (Center Of Gravity) method. This method removes the need for centroiding algorithms sliding over the entire image area. An extensive comparison of the proposed denoising technique with other traditional filters confirms that the proposed method resists all noise models and reconstructs well the corrupted images. Experiments show that the proposed YOLO-based star detector achieves high accuracy with a lightweight architecture without any extra latency. Besma Guesmi, David Moloney |
CoDIT | 2 |
| 2021 | Improving Performance-Power-Programmability in Space Avionics with Edge Devices: VBN on Myriad2 SoCabstractThe advent of powerful edge devices and AI algorithms has already revolutionized many terrestrial applications; however, for both technical and historical reasons, the space industry is still striving to adopt these key enabling technologies in new mission concepts. In this context, the current work evaluates an heterogeneous multi-core system-on-chip processor for use on-board future spacecraft to support novel, computationally demanding digital signal processors and AI functionalities. Given the importance of low power consumption in satellites, we consider the Intel Movidius Myriad2 system-on-chip and focus on SW development and performance aspects. We design a methodology and framework to accommodate efficient partitioning, mapping, parallelization, code optimization, and tuning of complex algorithms. Furthermore, we propose an avionics architecture combining this commercial off-the-shelf chip with a field programmable gate array device to facilitate, among others, interfacing with traditional space instruments via SpaceWire transcoding. We prototype our architecture in the lab targeting vision-based navigation tasks. We implement a representative computer vision pipeline to track the 6D pose of ENVISAT using megapixel images during hypothetical spacecraft proximity operations. Overall, we achieve 2.6 to 4.9 FPS with only 0.8 to 1.1 W on Myriad2 , i.e., 10-fold acceleration versus modern rad-hard processors. Based on the results, we assess various benefits of utilizing Myriad2 instead of conventional field programmable gate arrays and CPUs. Vasileios Leon, George Lentaris, Evangelos Petrongonas, Dimitrios Soudris, Gianluca Furano, Antonis Tavoularis, David Moloney |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2020 | 1.2 Watt Classification of 3D Voxel Based Point-clouds using a CNN on a Neural Compute Stick
Sam Caulfield, Joao Amaro, Gabriel Falcão Paiva Fernandes, David Moloney |
Neurocomputing | 5 |
| 2020 | Tree Annotations in LiDAR Data Using Point Densities and Convolutional Neural NetworksabstractLiDAR provides highly accurate 3D point clouds. However, data needs to be manually labelled in order to provide subsequent useful information. Manual annotation of such data is time consuming, tedious and error prone, and hence in this paper we present three automatic methods for annotating trees in LiDAR data. The first method requires high density point clouds and uses certain LiDAR data attributes for the purpose of tree identification, achieving almost 90% accuracy. The second method uses a voxel-based 3D Convolutional Neural Network on low density LiDAR datasets and is able to identify most large trees accurately but struggles with smaller ones due to the voxelisation process. The third method is a scaled version of the PointNet++ method and works directly on outdoor point clouds and achieves an F_score of 82.1% on the ISPRS benchmark dataset, comparable to the state-of-the-art methods but with increased efficiency. Ananya Gupta, Jonathan Byrne, David Moloney, Simon Watson 0001, Hujun Yin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Efficient winograd-based convolution kernel implementation on edge devicesabstractThe implementation of Convolutional Neural Networks on edge Internet of Things (IoT) devices is a significant programming challenge, due to the limited computational resources and the real-time requirements of modern applications. This work focuses on the efficient implementation of the Winograd convolution, based on a set of application-independent and Winograd-specific software techniques for improving the utilization of the edge devices computational resources. The proposed techniques were evaluated in Intel/Movidius Myriad2 platform, using 4 CNNs of various computational requirements. The results show significant performance improvements, up to 54%, over other convolution algorithms. Athanasios Xygkis, Lazaros Papadopoulos, David Moloney, Dimitrios Soudris, Sofiane Yous |
DAC | 3 |
| 2017 | Eyes of ThingsabstractResponsible Research and Innovation (RRI) is an approach that anticipates and assesses potential implications and societal expectations with regard to research and innovation, with the aim to foster the design of inclusive and sustainable research and innovation. While RRI includes many aspects, in certain types of projects ethics and particularly privacy, is arguably the most sensitive topic. The objective in Horizon 2020 innovation project Eyes of Things (EoT) is to build a small high-performance, low-power, computer vision platform (similar to a smart camera) that can work independently and also embedded into all types of artefacts. In this paper, we describe the actions taken within the project related to ethics and privacy. A privacy-by-design approach has been followed, and work continues now in four platform demonstrators. Noelia Vállez, José Luis Espinosa-Aranda, Jose M. Rico-Saavedra, Javier Parra-Patino, Oscar Déniz-Suárez, Alain Pagani, Stephan Krauß, Ruben Reiser, Didier Stricker, David Moloney, Aubrey K. Dunne, Dexmont Peña, Martin Wäny, Matteo Sorci, Tim Llewellynn, Christian Fedorczak, Thierry Larmoire, Elodie Roche, Marco Herbst, Andre Seirafi, Kasra Seirafi |
IC2E | 10 |
| 2016 | MvEcho - acoustic response modelling for auralisationabstractAuralisation is the process of simulating the listening experience at a given position. While many auralisation algorithms currently exist, the majority require a manual description of the environment. In contrast, MvEcho removes the need for this manual description to provide an environment independent algorithm to approximate auralisation. Initial implementation of MvEcho concerned the acoustic response of objects, namely the pyramid El Castillo, (shown below in Fig. 1) and has now been extended to model the response of rooms. The algorithm is fully functional in C and is currently being ported to the Myriad 2. Léonie Buckley, Sam Caulfield, David Moloney |
Hot Chips Symposium | 3 |
| 2016 | MvEchoabstractPresents a collection of slides covering the following topics: auralisation; volumetric accelerator; ray casting; VR audio; and virtual reality. Léonie Buckley, Sam Caulfield, David Moloney |
Hot Chips Symposium | 3 |
| 2016 | HW acceleration for volumetric applicationsabstractPresents a collection of slides covering the following topics: SLAMbench; platform-independent Kfusion; sparse voxel tree LoD; volumetric data sharing; volumetric maps; and volumetric data accelerator. David Moloney |
Hot Chips Symposium | 1 |
| 2016 | Embedded deep neural networks: "The cost of everything and the value of nothing"abstractDeep Learning for Embedded is all about Inference. Standard Networks are designed to achieve high-accuracy. Embedded implementation on architectures such as Movidius VPU can achieve significant performance results at the network edge. Next challenge is to further optimise networks to maximise performance per Watt. David Moloney |
Hot Chips Symposium | 1 |
| 2016 | Passive dense stereo vision on the Myriad2 VPUabstractPresents a collection of slides covering the following topics: passive-dense stereo vision; Myriad2 VPU; direct memory access controller; DMA controller; image matching; and computer vision. Luca Puglia, Mircea Ionica, Giancarlo Raiconi, David Moloney |
Hot Chips Symposium | 4 |
| 2016 | Object recognition speed improvement using BITMAP-HoGabstractCommonly, HoG/SVM classifier uses rectangular images for HoG feature descriptor extraction and training. This means significant additional work has to be done to process irrelevant pixels belonging to the background surrounding the object of interest. While some objects may indeed be square or rectangular, most of objects are not easily representable by simple geometric shapes. In Bitmap-HoG approach we propose in this paper, the irregular shape of object is represented by a bitmap to avoid processing of extra background pixels. Bitmap, derived from the training dataset, encodes those portions of an image to be used to train a classifier. Experimental results show that not only the proposed algorithm decreases the workload associated with HoG/SVM classifiers by 75% compared to the state-of-the-art, but also it shows an average increase about 5% in recall and a decrease about 2% in precision in comparison with standard HoG. David Moloney, Ivan Griffin |
ICIP | 2 |
| 2015 | Mixed-length SIMD code generation for VLIW architectures with multiple native vector-widthsabstractThe degree of DLP parallelism in applications is not fixed and varies due to different computational characteristics of applications. On the contrary, most of the processors today include single-width SIMD (vector) hardware to exploit DLP. However, single-width SIMD architectures may not be optimal to serve applications with varying DLP and they may cause performance and energy inefficiency. We propose the usage of VLIW processors with multiple native vector-widths to better serve applications with changing DLP. SHAVE is an example of such VLIW processor and provides hardware support for the native 32-bit and 128-bit wide vector operations. This paper researches and implements the mixed-length SIMD code generation support for SHAVE processor. More specifically, we target generating 32-bit and 128/64-bit SIMD code for the native 32-bit and 128-bit wide vector units of SHAVE processor. In this way, we improved the performance of compiler generated SIMD code by reducing the number of overhead operations and by increasing the SIMD hardware utilization. Experimental results demonstrated that our methodology implemented in the compiler improves the performance of synthetic benchmarks up to 47%. Erkan Diken, Martin J. O'Riordan, Roel Jordans, Lech Józwiak, Henk Corporaal, David Moloney |
ASAP | 6 |
| 2014 | Myriad 2: Eye of the computational vision stormabstractThis article consists of a collection of slides from the authors' conference presentation on the special features, system design and architectures, processing capabilities, and targeted markets for Movidius' Myraid 2, mobile vision processor family of products. David Moloney, Brendan Barry, Richard Richmond, Fergal Connor, Cormac Brick, David Donohoe |
Hot Chips Symposium | 1 |
| 2014 | Level-3 BLAS on myriad multi-core media-processor SoCabstractPresents a conference poster that covers the following topic: level-3 BLAS on myriad multicore media processing system-on-chip. Tomasz Szydzik, Marius Farcas, Valeriu Ohan, David Moloney |
Hot Chips Symposium | 4 |
| 2014 | Precision refinement for media-processor SoCs: fp32 -> fp64 on myriadabstractPresents a conference poster that addresses precision refinement for media processors for system-on-chip systems. Tomasz Szydzik, David Moloney |
Hot Chips Symposium | 2 |
| 2014 | An improved interest point matching algorithm for human body trackingabstractThe interest point (IP) matching algorithms match the points either locally or spatially. We propose a local-spatial IP matching algorithm usable for articulated human body tracking. The local-based stage finds matched IP pairs of two reference and target IP lists using a local-feature-descriptors-based matching method. Then, the spatial-based stage recovers more matched pairs from the remaining unmatched IPs through based on the result of the previous stage using the Shape Contexts (SC) feature vectors. The proposed approach benefits from the speed of local matching algorithms as well as the accuracy and robustness of spatial matching methods. Experimental results show that not only the proposed algorithm increases the precision rate from 44.71% to 97.41%, but also it improves the recall rate from 80.88% to 84.96%. Alistair Sutherland, David Moloney, Dexmont Peña |
IPAS | 3 |
| 2011 | 1TOPS/W software programmable media processor
David Moloney |
Hot Chips Symposium | 1 |
| 2007 | FPGA based Sparse Matrix Vector Multiplication using Commodity DRAM MemoryabstractSparse matrix by vector multiplication (SMV) is a key operation of many scientific and engineering applications. Field Programmable Gate Arrays (FPGAs) have the potential to significantly improve the performance of computationally intensive applications which are dominated by SMV. A shortcoming of most existing FPGA SMV implementations is that they use on-chip Block RAM or external SRAM to store the matrix, which severely limits the problem size. Real applications, such as Finite Element Analysis (FEA), require large memories. Realistically this capacity can only be provided by commodity DRAM. In this paper we address the problem of SMV for large matrices using commodity memory. We implement SPAR, a special purpose architecture that was previously proposed for large SMV computations in a VLSI co-processor using cheap external memory. We present an empirical evaluation of the SPAR architecture for use on FPGAs and highlight challenges that arise when tackling realistic FEA problems. David Gregg, Colm McSweeney, Ciarán McElroy, Fergal Connor, Séamas McGettrick, David Moloney, Dermot Geraghty |
FPL | 6 |
| 2005 | Streaming Sparse Matrix Compression/Decompression
David Moloney, Dermot Geraghty, Colm McSweeney, Ciarán McElroy |
HiPEAC | 1 |