VLDB 2026 Research / reviewers in the wild / expert
Piotr Dudek
dblp:82/3359
· DBLP profile ↗
52ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-6511-6165ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 17 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inpainting of Sparse Depth Maps from Monocular Depth-from-Focus on Pixel Processor ArraysabstractDepth estimation is essential for robotics and effective navigation. While many recent methods attempt to estimate dense depth maps from a single RGB image or a combination of an RGB image and sparse depth measurements, our work leverages the in-pixel computing capabilities of a pixel processor array (PPA), combined with an electrically tunable liquid lens, to capture semi-sparse depth maps via a depth-from-focus approach. We consider the problem of reconstructing dense depth maps from simulations of such measurements. We simulate a PPA-based depth-from-focus algorithm on a synthetic focal stack derived from a monocular RGB-D dataset, demonstrating competitive dense depth map reconstruction from depth frames containing as few as 10% non-zero pixels, with 5-bit resolution. Furthermore, we enhance semi-sparse depth completion performance by fusing the depth from focus cues with concurrently acquired RGB images. We also use belief propagation, allowing for highly localized and parallel computation without access to global memory, offering a promising solution from the perspective of PPAs. We show the performance of these algorithms on semi-sparse image depth reconstruction tasks. Maciej Lewandowski, Piotr Dudek |
WACV | 2 |
| 2025 | Descriptor-In-Pixel : Point-Feature Tracking For Pixel Processor ArraysabstractThis paper presents a novel approach for joint point-feature detection and tracking, designed specifically for Pixel Processor Array (PPA) vision sensors. Instead of standard pixels, PPA sensors consist of thousands of "pixel-processors", enabling massive parallel computation of visual data at the point of light capture. Our approach performs all computation entirely in-pixel, meaning no raw image data need ever leave the sensor for external processing. We introduce a Descriptor-In-Pixel paradigm, in which a feature descriptor is held within the memory of each pixel-processor. The PPA’s architecture enables the response of every processor’s descriptor, upon the current image, to be computed in parallel. This produces a"descriptor response map" which, by generating the correct layout of descriptors across the pixel-processors, can be used for both point-feature detection and tracking. This reduces sensor output to just sparse feature locations and descriptors, read-out via an address-event interface, giving a greater than 1000× reduction in data transfer compared to raw image output. The sparse readout and complete utilization of all pixel-processors makes our approach very efficient. Our implementation upon the SCAMP-7 PPA prototype runs at over 3000 FPS (Frames Per Second), tracking point-features reliably under violent motion. This is the first work performing point-feature detection and tracking entirely in-pixel.1 Laurie Bose, Jianing Chen 0005, Piotr Dudek |
CVPR | 3 |
| 2025 | Focal Plane Visual Feature Generation and Matching on a Pixel Processor Array
Laurie Bose, Jianing Chen 0005, Piotr Dudek, Walterio W. Mayol-Cuevas |
ICCV | 4 |
| 2025 | Design and FPGA Realization of QUBO hardware accelerator for MAX-CUT problemabstractCombinatorial optimization problems are fundamental to many industrial and scientific applications, but are often NP-hard, requiring heuristic and probabilistic approaches to find good solutions fast. This paper presents the first hardware implementation of the neuromorphic-inspired NeuroSA algorithm, for accelerating Quadratic Unconstrained Binary Optimization (QUBO) problems. We show that it is possible to run a standard QUBO benchmark problem with 800 variables on a low-cost FPGA with a clock speed of approximately 144 MHz, which converges to the same solution value found by the numerical solver. Maciej Lewandowski, Piotr Dudek |
ISCAS | 2 |
| 2025 | Digital Implementation of a Spiking/Bursting Neuron Model with Low ResourcesabstractThis paper introduces a simplified digital neuron model that balances computational resources with biological accuracy and is designed for efficient hardware implementation. By reducing the complexity of the Izhikevich neuron model, the proposed design replicates diverse spiking behaviors while significantly reducing hardware complexity and power consumption. Implementing efficient hardware operations, the model achieves 16 of the Izhikevich behaviors. Field-Programmable Gate Array (FPGA) implementations demonstrate a 61% to 94% reduction in resource usage compared to previous implementations, with a 55% increase in operating frequency, making the model ideal for large-scale neuromorphic systems. Furthermore, it can emulate the LIF model with negligible additional resource requirements. Piotr Dudek, Jayawan H. B. Wijekoon |
ISCAS | 2 |
| 2024 | PixelRNN: In-pixel Recurrent Neural Networks for End-to-end-optimized Perception with Neural SensorsabstractConventional image sensors digitize high-resolution images at fast frame rates, producing a large amount of data that needs to be transmitted off the sensor for fur-ther processing. This is challenging for perception system operating on edge devices, because communication is power inefficient and induces latency. Fueled by innovations in stacked image sensor fabrication, emerging sensor-processors offer programmability and processing capabilities directly on the sensor. We exploit these capabilities by developing an efficient recurrent neural network architecture, PixelRNN, that encodes spatio-temporal features on the sensor using purely binary operations. PixelRNN reduces the amount of data to be transmitted off the sensor by factors up to 256 compared to the raw sensor data while offering competitive accuracy for hand gesture recognition and lip reading tasks. We experimentally validate PixelRNN using a prototype implementation on the SCAMP-5 sensor-processor platform. Haley M. So, Laurie Bose, Piotr Dudek, Gordon Wetzstein |
CVPR | 3 |
| 2022 | MantissaCam: Learning Snapshot High-dynamic-range Imaging with Perceptually-based In-pixel Irradiance EncodingabstractThe ability to image high-dynamic-range (HDR) scenes is crucial in many computer vision applications. The dynamic range of conventional sensors, however, is fundamentally limited by their well capacity, resulting in saturation of bright scene parts. To overcome this limitation, emerging sensors offer in-pixel processing capabilities to encode the incident irradiance. Among the most promising encoding schemes is modulo wrapping, which results in a computational photography problem where the HDR scene is computed by an irradiance unwrapping algorithm from the wrapped low-dynamic-range (LDR) sensor image. Here, we design a neural network-based algorithm that outperforms previous irradiance unwrapping methods and we design a perceptually inspired “mantissa,” or log-modulo, encoding scheme that more efficiently wraps an HDR scene into an LDR sensor. Combined with our reconstruction framework, MantissaCam achieves state-of-the-art results among modulo-type snapshot HDR imaging approaches. We demonstrate the efficacy of our method in simulation and show benefits of our algorithm on modulo images captured with a prototype implemented with a programmable sensor. Haley M. So, Julien N. P. Martel, Gordon Wetzstein, Piotr Dudek |
ICCP | 4 |
| 2021 | Agile reactive navigation for a non-holonomic mobile robot using a pixel processor arrayabstractAbstract This paper presents an agile reactive navigation strategy for driving a non‐holonomic ground vehicle around a pre‐set course of gates in a cluttered environment using a low‐cost processor array sensor. This enables machine vision tasks to be performed directly upon the sensor's image plane, rather than using a separate general‐purpose computer. The authors demonstrate a small ground vehicle running through or avoiding multiple gates at high speed using minimal computational resources. To achieve this, target tracking algorithms are developed for the Pixel Processing Array and captured images are then processed directly on the vision sensor acquiring target information for controlling the ground vehicle. The algorithm can run at up to 2000 fps outdoors and 200 fps at indoor illumination levels. Conducting image processing at the sensor level avoids the bottleneck of image transfer encountered in conventional sensors. The real‐time performance of on‐board image processing and robustness is validated through experiments. Experimental results demonstrate the algorithm's ability to enable a ground vehicle to navigate at an average speed of 2.20 m/s for passing through multiple gates and 3.88 m/s for a ‘slalom’ task in an environment featuring significant visual clutter. Laurie Bose, Colin Greatwood, Jianing Chen 0005, Rui Fan 0001, Tom Richardson 0002, Stephen J. Carey, Piotr Dudek, Walterio W. Mayol-Cuevas |
IET Image Process. | 8 |
| 2020 | High-speed Light-weight CNN Inference via Strided Convolutions on a Pixel Processor Array
Laurie Bose, Jianing Chen 0005, Stephen J. Carey, Piotr Dudek, Walterio W. Mayol-Cuevas |
BMVC | 5 |
| 2020 | Fully Embedding Fast Convolutional Networks on Pixel Processor Arrays
Laurie Bose, Piotr Dudek, Jianing Chen 0005, Stephen J. Carey, Walterio W. Mayol-Cuevas |
ECCV (29) | 2 |
| 2020 | Proximity Estimation Using Vision Features Computed On SensorabstractThis paper presents a monocular vision based proximity estimation system using abstract features, such as corner points, blobs and edges, as inputs to a neural network. An experimental vehicle was built using a vision system integrating the SCAMP-5 vision chip, a micro-controller, and an RC model car. The vision chip includes image sensor with embedded 256×256 processor SIMD array. The pixel processor array chip was programmed to capture images and run the feature algorithms directly on the focal plane, and then digest them so that only sparse feature description data were read-out in the form of 40 values. By logging the vision output and the output from three infrared proximity sensors, training data were obtained to train three fully connected layer-recurrent neural networks with fewer than 700 parameters each. The trained neural network was able to estimate the proximity to the level of accuracy sufficient for a reactive collision avoidance behaviour to be achieved. The latency of the control system, from image capture to neural network output, was under 4ms, enabling the vehicles to avoid obstacles while moving at 0.64m/s to 1.8m/s in the experiment. Jianing Chen 0005, Stephen J. Carey, Piotr Dudek |
ICRA | 4 |
| 2020 | Live Demonstration: CNN Inference on the Focal Plane with a Pixel Processor ArrayabstractWe present a novel method of CNN inference on a pixel processor array device, demonstrating it using a handwritten digit (digits 0-9) classification task, with all steps of the neural network computation performed on the focal plane. The vision chip that we deploy (SCAMP-7) has a 256×256 array of processor elements (PE) integrated within the image sensor. The algorithm runs at over 3000 frames per second (FPS) and over 90% classification accuracy, with the sensor chip only outputting ten scalar values corresponding to the classification scores. Stephen J. Carey, Laurie Bose, Tom Richardson 0002, Walterio W. Mayol-Cuevas, Jianing Chen 0005, Piotr Dudek |
ISCAS | 6 |
| 2020 | Lessons Learned the Hard Wayabstract“Fail often to succeed sooner” is a common mantra that we are told is the secret to success. When reporting research results, however, scholars rarely write about their failed attempts and only focus on the successful ones. Perhaps the source of this disconnect between what we preach and what we do can be found in the underlying assumption that published work is meant to move the field forward and failed attempts supposedly do not. The goal of the confessions presented in this paper is to show that even failed attempts are genuine and valuable contributions to our field provided that we learn from our mistakes and correct them. The 27 confessions span from planning oversights, digital and analog design errors, misunderstanding of devices, overlooked parasitics, LVS errors, and troubles in testing. Tobi Delbruck, Ibrahim M. Elfadel, Shahzad Muzaffar, Germain Haessig, Bo Wang 0012, Amine Bermak, Rui Graca, Luis A. Camuñas-Mesa, Bathiya Senevirathna, Pamela Abshire, Bernabé Linares-Barranco, Saeed Afshar, Shih-Chii Liu, Runchun Wang, Piotr Dudek, Stephen J. Carey, José M. de la Rosa 0001, Marc Dandin, Sheung Lu, Vincent Frick, Teresa Serrano-Gotarredona, Paula López Martinez 0001, Melika Payvand, Advait Madhavan, Eric R. Fossum, Juan Camilo Vasquez Tieck, Yan Liu 0016, Timothy G. Constandinou, Alexander Serb, Ricardo Carmona-Galán, Robert Nawrocki, Walter D. Leon-Salas |
ISCAS | 15 |
| 2020 | A Mixed-Signal Spatio-Temporal Signal Classifier for On-Sensor Spike SortingabstractNeuromorphic systems provide an alternative to conventional computing hardware, promising low-power operation suitable for sensory-processing and edge computing. In this paper, we present a mixed-signal processing system designed to provide on-sensor classification of signals obtained from multi-electrode array neural recordings. The designed circuits implement a real-time spike sorting algorithm, and operate on signals represented by asynchronous event streams. We combine analog circuits computation primitives (temporal surface generation, distance computation, winner-take-all) to implement a spatio-temporal clustering algorithm, classifying signals acquired by neighbouring electrodes. The prototype chip has been submitted for fabrication in a 180nm CMOS technology. The circuits are designed to fit, alongside signal conditioning and conversion circuits, in the area under the recording electrodes (below 80×80um per electrode). Circuit implementation details and simulation results are presented. The expected neural spike recognition rates of 75% in a single-layer network and 88% in a 2-layer network are comparable with a software implementation, while the system is designed to provide a low-power embedded real-time solution. This work provides a foundation towards the design of a large scale neuromorphic processing system, to be embedded in brain-machine interfaces. Germain Haessig, Daniel García-Lesta, Gregor Lenz, Ryad Benosman, Piotr Dudek |
ISCAS | 5 |
| 2020 | Neural Sensors: Learning Pixel Exposures for HDR Imaging and Video Compressive Sensing With Programmable SensorsabstractCamera sensors rely on global or rolling shutter functions to expose an image. This fixed function approach severely limits the sensors' ability to capture high-dynamic-range (HDR) scenes and resolve high-speed dynamics. Spatially varying pixel exposures have been introduced as a powerful computational photography approach to optically encode irradiance on a sensor and computationally recover additional information of a scene, but existing approaches rely on heuristic coding schemes and bulky spatial light modulators to optically implement these exposure functions. Here, we introduce neural sensors as a methodology to optimize per-pixel shutter functions jointly with a differentiable image processing method, such as a neural network, in an end-to-end fashion. Moreover, we demonstrate how to leverage emerging programmable and re-configurable sensor-processors to implement the optimized exposure functions directly on the sensor. Our system takes specific limitations of the sensor into account to optimize physically feasible optical codes and we evaluate its performance for snapshot HDR and high-speed compressive imaging both in simulation and experimentally with real scenes. Julien N. P. Martel, Lorenz K. Müller, Stephen J. Carey, Piotr Dudek, Gordon Wetzstein |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | A Camera That CNNs: Towards Embedded Neural Networks on Pixel Processor ArraysabstractWe present a convolutional neural network implementation for pixel processor array (PPA) sensors. PPA hardware consists of a fine-grained array of general-purpose processing elements, each capable of light capture, data storage, program execution, and communication with neighboring elements. This allows images to be stored and manipulated directly at the point of light capture, rather than having to transfer images to external processing hardware. Our CNN approach divides this array up into 4x4 blocks of processing elements, essentially trading-off image resolution for increased local memory capacity per 4x4 ”pixel”. We implement parallel operations for image addition, subtraction and bit-shifting images in this 4x4 block format. Using these components we formulate how to perform ternary weight convolutions upon these images, compactly store results of such convolutions, perform max-pooling, and transfer the resulting sub-sampled data to an attached micro-controller. We train ternary weight filter CNNs for digit recognition and a simple tracking task, and demonstrate inference of these networks upon the SCAMP5 PPA system. This work represents a first step towards embedding neural network processing capability directly onto the focal plane of a sensor. Laurie Bose, Piotr Dudek, Jianing Chen 0005, Stephen J. Carey, Walterio W. Mayol-Cuevas |
ICCV | 2 |
| 2018 | Perspective Correcting Visual Odometry for Agile MAVs using a Pixel Processor ArrayabstractThis paper presents a visual odometry approach using a Pixel Processor Array (PPA) camera, specifically, the SCAMP-5 vision chip. In this device, each pixel is capable of storing data and performing computation, enabling a variety of computer vision tasks to be carried out directly upon the sensor itself. In this work the PPA performs HDR edge detection, perspective correction and image alignment based odometry, allowing the position and heading of a MAV to be tracked at several hundred frames per second. We evaluate our PPA based approach by direct comparison with a motion capture system for a variety of trajectories. These include rapid accelerations that would incur significant motion blur at low frame rates, and lighting conditions that would typically lead to under or over exposure of image detail. Such challenging conditions would often lead to unusable images when relying on traditional image sensors. Colin Greatwood, Laurie Bose, Tom Richardson 0002, Walterio W. Mayol-Cuevas, Jianing Chen 0005, Stephen J. Carey, Piotr Dudek |
IROS | 7 |
| 2017 | Visual Odometry for Pixel Processor ArraysabstractWe present an approach of estimating constrained egomotion on a Pixel Processor Array (PPA). These devices embed processing and data storage capability into the pixels of the image sensor, allowing for fast and low power parallel computation directly on the image-plane. Rather than the standard visual pipeline whereby whole images are transferred to an external general processing unit, our approach performs all computation upon the PPA itself, with the camera's estimated motion as the only information output. Our approach estimates 3D rotation and a 1D scale-less estimate of translation. We introduce methods of image scaling, rotation and alignment which are performed solely upon the PPA itself and form the basis for conducting motion estimation. We demonstrate the algorithms on a SCAMP-5 vision chip, achieving frame rates >1000Hz at ~2W power consumption. Laurie Bose, Jianing Chen 0005, Stephen J. Carey, Piotr Dudek, Walterio W. Mayol-Cuevas |
ICCV | 4 |
| 2017 | Tracking control of a UAV with a parallel visual processorabstractThis paper presents a vision-based control strategy for tracking a ground target using a novel vision sensor featuring a processor for each pixel element. This enables computer vision tasks to be carried out directly on the focal plane in a highly efficient manner rather than using a separate general purpose computer. The strategy enables a small, agile quadrotor Unmanned Air Vehicle (UAV) to track the target from close range using minimal computational effort and with low power consumption. To evaluate the system we target a vehicle driven by chaotic dual-pendulum trajectories. Target proximity and the large, unpredictable accelerations of the vehicle cause challenges for the UAV in keeping it within the downward facing camera's field of view (FoV). A state observer is used to smooth out predictions of the target's location and, importantly, estimate velocity. Experimental results also demonstrate that it is possible to continue to re-acquire and follow the target during short periods of loss in target visibility. The tracking algorithm exploits the parallel nature of the visual sensor, enabling high rate image processing ahead of any communication bottleneck with the UAV controller. With the vision chip carrying out the most intense visual information processing, it is computationally trivial to compute all of the controls for tracking onboard. This work is directed toward visual agile robots that are power efficient and that ferry only useful data around the information and control pathways. Colin Greatwood, Laurie Bose, Tom Richardson 0002, Walterio W. Mayol-Cuevas, Jianing Chen 0005, Stephen J. Carey, Piotr Dudek |
IROS | 7 |
| 2017 | High-speed depth from focus on a programmable vision chip using a focus tunable lensabstractIn this paper, we present a 3D imaging system providing a semi-dense depth map, using a passive, low-power, compact, static, monocular camera. The demonstrated depth estimation system reconstructs 32 depth-levels in real-time at 25FPS drawing less than 1.9W of power. This is achieved by performing computation on an analog focal-plane processor that analyses frames captured through a vibrating liquid focus-tunable lens. The optical system provides shallow depth of focus images and fast sweeps of optical power, while the use of pixel-level processing removes the sensor-processor bandwidth limitations of depth-from-focus systems built using conventional imaging and processor technologies. All-in-focus images are also obtained. Julien N. P. Martel, Lorenz K. Müller, Stephen J. Carey, Piotr Dudek |
ISCAS | 4 |
| 2017 | Live demonstration: Depth from focus on a focal plane processor using a focus tunable liquid lensabstractWe demonstrate a 3D imaging system that produces sparse depth maps. It consists in a liquid focus-tunable lens whose focal power can be changed at high speed, placed in front of a SCAMP5 vision-chip embedding processing capabilities in each pixel. The focus-tunable lens performs focal sweeps with shallow depth of fields. These are sampled by the vision chip taking multiple images at different focus and analyzed on-chip to produce a single depth frame. The combination of the focus tunable-lens with the vision-chip, enabling near-focal plane processing, allows us to present a compact passive system that is static, monocular, real-time (> 25FPS) and low-power (<; 1.6W). Julien N. P. Martel, Lorenz K. Müller, Stephen J. Carey, Jonathan Müller, Yulia Sandamirskaya, Piotr Dudek |
ISCAS | 6 |
| 2016 | Parallel HDR tone mapping and auto-focus on a cellular processor array vision chipabstractTo improve computational efficiency, it may be advantageous to transfer part of the intelligence lying in the core of a system to its sensors. Vision sensors equipped with small programmable processors at each pixel allow us to follow this principle in so-called near-focal plane processing, which is performed on-chip directly where light is being collected. Such devices need then only to communicate relevant pre-processed visual information to other parts of the system. In this work, we demonstrate how two classical problems, namely high dynamic range imaging and auto-focus, can be solved efficiently using two simple parallel algorithms implemented on such a chip. We illustrate with these two examples that embedding uncomplicated algorithms on-chip, directly where information acquisition takes place can replace more complex dedicated post-processing. Adapting data acquisition by bringing processing at the sensor level allows us to explore solutions that would not be feasible in a conventional sensor-ADC-processor pipeline. Julien N. P. Martel, Lorenz K. Müller, Stephen J. Carey, Piotr Dudek |
ISCAS | 4 |
| 2015 | Gradient-descent-based learning in memristive crossbar arraysabstractThis paper describes techniques to implement gradient-descent-based machine learning algorithms on crossbar arrays made of memristors or other analog memory devices. We introduce the Unregulated Step Descent (USD) algorithm, which is an approximation of the steepest descent algorithm, and discuss how it addresses various hardware implementation issues. We discuss the effect of device parameters and their variability on performance of the algorithm by using artificially generated and real-world datasets. In addition to providing insights on the effect of device parameters on learning, we illustrate how the USD algorithm partially offsets the effect of device variability. Finally, we discuss how the USD algorithm can be implemented in crossbar arrays using a simple 4-phase training scheme. The method allows parallel update of crossbar memory elements and reduces the hardware cost and complexity of the training architecture significantly. Manu V. Nair, Piotr Dudek |
IJCNN | 2 |
| 2015 | Toward joint approximate inference of visual quantities on cellular processor arraysabstractThe interacting visual maps (IVM) algorithm introduced in [1] is able to perform the joint approximate inference of several visual quantities such as optic-flow, gray-level intensities and ego-motion, using a sparse input coming from a neuromorphic dynamic vision sensor (DVS). We show that features of the model such as the intrinsic parallelism and distributed nature of its computation make it a natural candidate to benefit from the cellular processor array (CPA) hardware architecture. We have now implemented the IVM algorithm on a general-purpose CPA simulator, and here we present results of our simulations and demonstrate that the IVM algorithm indeed naturally fits the CPA architecture. Our work indicates that extended versions of the IVM algorithm could benefit greatly from a dedicated hardware implementation, eventually yielding a high speed, low power visual odometry chip. Julien N. P. Martel, Miguel Chau, Piotr Dudek, Matthew Cook 0001 |
ISCAS | 3 |
| 2015 | An event-driven massively parallel fine-grained processor arrayabstractA multi-core event-driven parallel processor array design is presented. Using relatively simple 8-bit processing cores and a 2D mesh network topology, the architecture focuses on reducing the area occupation of a single processor core. A large number of these processor cores can be implemented on a single integrated chip to create a MIMD architecture capable of providing a powerful processing performance. Each processor core is an event-driven processor which can enter an idle mode when no data is changing locally. An 8 × 8 prototype processor array is implemented in a 65 nm CMOS process in 1,875 μm × 1,875 μm. This processor array is capable of performing 5.12 GOPS operating at 80 MHz with an average power consumption of 75.4 mW. Declan Walsh, Piotr Dudek |
ISCAS | 2 |
| 2014 | Live demonstration: A sensor-processor array integrated circuit for high-speed real-time machine visionabstractA demonstration is made of the high-speed real-time image processing capabilities of the SCAMP-5 vision chip. The device provides a software-programmable 256×256 pixel-parallel SIMD processor array. In the example application, the IC can determine dimensions and the location of a single object, at a sustained rate of 100,000fps. At 30,000fps, the chip can return the same metrics from 5 objects. This is accomplished by use of near-sensor processing which circumvents the requirement to digitise images; all processing is done “on the focal plane” and only high-level object information is transmitted off-chip as discrete address-events. Stephen J. Carey, David Robert Wallace Barr, Alexey Lopich, Piotr Dudek |
ISCAS | 5 |
| 2014 | Characterization of processing errors on analog fully-programmable cellular sensor-processor arraysabstractAnalog processor arrays, particularly for vision chips, have been in development for a number of years. Based normally on the SIMD computing paradigm, they achieve very high instruction parallelism through use of compact processing elements. A fundamental aspect of processor operation is the ability to copy content of a register to another. This operation carries with it a number of errors. This paper identifies these errors and provides the means by which they can be separated and measured. These techniques can then be applied to analyze any operation upon an analog processor array. The results from such an analysis reflect on the susceptibility of the circuit to process mismatch, and thermal and switch noise performance of a particular analog memory design. Stephen J. Carey, Ákos Zarándy, Piotr Dudek |
ISCAS | 3 |
| 2014 | The accuracy and scalability of continuous-time Bayesian inference in analogue CMOS circuitsabstractThis paper discusses the idea of Bayesian inference in factor graphs implemented as continuous-time current-mode analogue CMOS circuits using Gilbert multipliers for arithmetic operations. The computational accuracy, accounting for the systematic and random (fabrication mismatch) errors, and the scalability of such realisations were verified in simulations of networks consisting of 5 - 121 nodes implemented using models from a standard 90 nm CMOS technology. The obtained results show a relatively short settling time, typically below 3 μs at a power less than 7 mW, with the equivalent computational speed of over 35 arithmetic operations per nanosecond but with a limited accuracy, mainly affected by fabrication mismatch. Such realisations could be used in applications requiring fast and low power approximate Bayesian inference. Przemyslaw Mroszczyk, Piotr Dudek |
ISCAS | 2 |
| 2013 | AMBER: Adapting multi-resolution background extractorabstractIn this paper, a fast self-adapting multi-resolution background detection algorithm is introduced. A pixel-based background model is proposed, that represents not only each pixel's background values, but also their efficacies, so that new background values always replace the least effective ones. Model maintenance and global control processes ensure fast initialization, adaptation to background changes with different timescales, restrain the generation of ghosts, and adjust the decision thresholds based on noise levels. Evaluation results indicate that the proposed algorithm outperforms most other state-of-the-art algorithms not only in terms of accuracy, but also in terms of processing speed and memory requirements. Piotr Dudek |
ICIP | 2 |
| 2013 | A Smart Surface Simulation EnvironmentabstractA versatile simulation system is presented, for observing the effects of different algorithmic and mechanical approaches to physical object manipulation via "smart" surfaces -physical machines which use a distributed array of actuators algorithmically controlled to manipulate objects placed upon them. The system presented is a graphical 3D physics sandbox simulator that allows the modeller to include real-world and hardware related constraints such as inter-object collisions and trajectories, mechanical properties, friction, and sensor/actuator reliability, while conducting experiments in an interactive, repeatable and controlled environment. David Robert Wallace Barr, Declan Walsh, Piotr Dudek |
SMC | 3 |
| 2013 | Low power high-performance smart camera system based on SCAMP vision sensor
Stephen J. Carey, David Robert Wallace Barr, Piotr Dudek |
J. Syst. Archit. | 3 |
| 2012 | A field programmable array core for image processing (abstract only)abstractMassively parallel processor arrays have been shown to be an effective and suitable choice for image processing tasks [1]. More recently, some of the state of the art processor arrays have been used for real-time machine vision tasks such as intelligent transport system applications [2] or video processing on mobile applications [3] providing a much more powerful solution than a conventional processor. A number of Single Instruction Multiple Data (SIMD) processor arrays have been implemented on FPGAs [4]-[6], which are particularly suited to implementing such processor architectures because of their similarities of both being arrays of fine grained logic elements. In this work, we propose an FPGA implementation of a processor array where the processing elements (PEs) are as small as possible, while providing local memory sufficient for processing greyscale images. The PE is then replicated to form an array. A 32 × 32 PE array is implemented on a Xilinx Virtex 5 XC5VLX50 FPGA using the four-neighbour connectivity with the possibility to scale up using a larger FPGA. The processor array operates at a frequency of 96 MHz and executes a peak of 98.3 giga operations per second (GOPS) (bit-serial operations). A binary edge detection algorithm is performed in 52.08 ns. Uploading and downloading a binary image in a 32 × 32 array takes an extra 687.5 ns. Sobel edge detection of an 8-bit greyscale image is performed in 5.33 µs. Uploading and downloading an 8-bit greyscale image in a 32 × 32 array takes 5.36 µs. With larger FPGAs being available in the future, the array sizes comparable to state of the art custom designed ICs can be implemented on these FPGAs. Declan Walsh, Piotr Dudek |
FPGA | 2 |
| 2012 | Trigger-wave collision detecting asynchronous cellular logic array for fast image skeletonizationabstractThis paper presents the design of an asynchronous cellular logic array for binary image processing algorithms based on wave propagation/collision in an excitable medium. The array consists of identical logic cells enabling the propagation and detection of wave-front collisions necessary for the object skeletonization. Low power, low area and high processing speed requirements were met by employing the asynchronous dynamic logic approach resulting in a processing time less than 0.45ns/pixel and energy consumption of less than 0.15pJ/pixel. The cell consists of 19 transistors and occupies an area of 7.5×6.3μm2in 90nm CMOS technology. The proposed array could be used as a coprocessor in pixel-parallel SIMD architectures aiding the fast execution of medium-level image processing algorithms. Przemyslaw Mroszczyk, Piotr Dudek |
ISCAS | 2 |
| 2012 | Heterogeneous neurons and plastic synapses in a reconfigurable cortical neural network ICabstractThis paper presents an analogue VLSI circuit intended to be used in a neural network architecture that closely resembles the small-scale laminar micro-circuits of the neocortex. The Cortical Neural Layer (CNL) chip comprises of 120 reconfigurable cortical neurons and 7,560 synapses. The neurons can be configured to produce regular spiking, fast spiking, chattering, intrinsically bursting, and other complex activity patterns. The synaptic circuits include inhibitory/ excitatory, facilitating/depressing and spike-time dependent plasticity (STDP) dynamics. The connectivity of the neural network can be configured using off-chip spike-routing and on-chip axonal arbor connections. A pre-synaptic spike can be sent to a group of crossbar synapses simultaneously, reducing latency in the pre-synaptic spike routing, enabling a high degree of connectivity of the neural network. The device is fabricated in a 0.35 μm CMOS technology and on-chip neural dynamics are experimentally verified. Jayawan H. B. Wijekoon, Piotr Dudek |
ISCAS | 2 |
| 2011 | Self-Organizing Neural Population Coding for improving robotic visuomotor coordinationabstractWe present an extension of Kohonen's Self Organizing Map (SOM) algorithm called the Self Organizing Neural Population Coding (SONPC) algorithm. The algorithm adapts online the neural population encoding of sensory and motor coordinates of a robot according to the underlying data distribution. By allocating more neurons towards area of sensory or motor space which are more frequently visited, this representation improves the accuracy of a robot system on a visually guided reaching task. We also suggest a Mean Reflection method to solve the notorious border effect problem encountered with SOMs for the special case where the latent space and the data space dimensions are the same. Piotr Dudek, Bertram E. Shi |
IJCNN | 2 |
| 2011 | Confession session: Learning from others mistakesabstractPeople rarely put in their papers the things that didn't work, the mistakes they made, and how they found out what went wrong. Such confessions can help others learn how to avoid similar mistakes. Twenty-six confessions were collected to form the bulk of this paper. Themes that arise are errors that result from not understanding the limitations of simulation tools in modeling physical reality, chip verification errors that result from lack of clear communication between designers, and projects that are considered in their own isolated environment of technical challenges rather than the broader context of their environment or application. Pamela Abshire, Amine Bermak, Raphael Berner, Gert Cauwenberghs, Shoushun Chen, Jennifer Blain Christen, Timothy G. Constandinou, Eugenio Culurciello, Marc Dandin, Timir Datta, Tobi Delbruck, Piotr Dudek, Amir Eftekhar, Ralph Etienne-Cummings, Giacomo Indiveri, Matthew K. Law, Bernabé Linares-Barranco, Jonathan Tapson, Wei Tang 0002, Yiming Zhai |
ISCAS | 12 |
| 2011 | A processor element for a mixed signal cellular processor array vision chipabstractA combined analogue and digital processing element for a pixel-parallel vision chip has been designed in 0.18μm CMOS technology. In addition to 7 analogue registers, each pixel incorporates 14 bits of digital memory. In the analogue domain its processing capabilities include addition, subtraction and squaring, with digital domain NOT and OR operators also available. The processing element has dimensions of 32×32μm and is designed to operate at 10MHz. A test chip has been fabricated. Stephen J. Carey, Alexey Lopich, Piotr Dudek |
ISCAS | 3 |
| 2011 | Live demonstration: Real-time image processing on ASPA2 vision systemabstractThis live demonstration presents a vision system based on a digital SIMD vision chip with in-pixel processing capabilities. The system is comprised of asynchronous/synchronous processor array (ASPA2), embedded custom microcontroller with interface circuits and software development environment. Execution of a number of low and medium level image processing algorithms in real time is demonstrated. Alexey Lopich, David Robert Wallace Barr, Piotr Dudek |
ISCAS | 4 |
| 2011 | Analogue CMOS circuit implementation of a dopamine modulated synapseabstractThis paper describes a synapse circuit that approximately implements the dynamics of the dopamine (DA) modulated synapse proposed by Izhikevich (2007). The dynamics of the model, based on 'eligibility traces' generated according to a spike-timing-dependent plasticity (STDP) rule, ensure that causal pre-/postsynaptic spiking activity in the time preceding the reward, signaled by DA, leads to strengthening of the synaptic connections. The circuits are designed and fabricated in a 0.35 μm CMOS technology and the simulation results are presented. This circuit block is a good candidate for the development of neuromorphic VLSI architectures that implement brain-inspired computation using biologically plausible reinforcement learning strategies. Jayawan H. B. Wijekoon, Piotr Dudek |
ISCAS | 2 |
| 2011 | Architecture and design of a programmable 3D-integrated cellular processor array for image processingabstractIn this work we present a design of a massively-parallel cellular processor array implemented in 3D CMOS technology. The proof of concept 128×96 array device is partitioned across two custom designed layers. Additionally, three layers of DDR memory are vertically stacked and bonded underneath. The processor benefits from 358Gbit/s data rate between memory and array, as well as from high logic density, thanks to improved routing across silicon layers with Trough Silicon Vias (TSVs). Alexey Lopich, Piotr Dudek |
VLSI-SoC | 2 |
| 2010 | Using Reinforcement Learning to Guide the Development of Self-organised Feature Maps for Visual Orienting
Kevin Brohan, Kevin N. Gurney, Piotr Dudek |
ICANN (2) | 3 |
| 2010 | An 80×80 general-purpose digital vision chip in 0.18μm CMOS technologyabstractIn this paper we present an implementation of the asynchronous/synchronous processor array (ASPA2) - a digital SIMD vision chip. The chip has been fabricated in a 0.18 μm CMOS process and comprises 80×80 array of pixel processors. The architecture of the chip is overviewed, the design of the processing cell is presented and implementation issues are discussed. At 75 MHz ASPA2 demonstrates 373 GOPS/W power and 871 MOPS/mm2area efficiency, making it suitable for the design of high speed and low power vision systems. Alexey Lopich, Piotr Dudek |
ISCAS | 2 |
| 2008 | Reconfigurable platforms and the challenges for large-scale implementations of spiking neural networksabstractFPGA devices have witnessed popularity in their use for the rapid prototyping of biological Spiking Neural Network (SNNs) applications, as they offer the key requirement of reconfigurability. However, FPGAs do not efficiently realise the biological neuron/synaptic models. Also their routing structures cannot accommodate the high levels of neuron inter-connectivity inherent in complex SNNs. This paper highlights and discusses the current challenges of implementing large scale SNNs on reconfigurable FPGAs. The paper presents a novel Field Programmable Neural Network (FPNN) architecture incorporating low power analogue synapse and a network on chip architecture for SNN routing and configuration. Initial results are presented. Jim Harkin, Fearghal Morgan, Steve Hall, Piotr Dudek, Thomas Dowrick, Liam McDaid |
FPL | 4 |
| 2008 | ASPA: Focal Plane digital processor array with asynchronous processing capabilitiesabstractIn this paper we present implementation and experimental results for a digital vision chip that operates in mixed asynchronous/synchronous mode. Mixed configuration benefits from full programmability (discrete-time mode) and high operational performance in global image processing operations (continuous-time mode) thus extending the application field of smart sensors from low- to medium-level processing. A 19x22 proof-of-concept chip was fabricated and tested. At peak operational frequency (150 MHz) each cell provides 9.6 MOPS thus achieving area utilization 820.8 MOPS/mm2and power efficiency 29 GOPS/W. Alexey Lopich, Piotr Dudek |
ISCAS | 2 |
| 2008 | Focal-plane moving object segmentation for realtime video surveillanceabstractIn this paper a new technique for segmenting and tracking moving objects in a user-defined control area is presented. It is based on an active contours technique called Pixel-Level Snakes (PLS) whose capabilities to manage changes of contour topology and to introduce additional constraints in the contour evolution are used to define a control area as well as to segment and track moving objects. Furthermore, PLS can reach a very high speed of response when they are implemented on a pixel-parallel hardware platform. To illustrate the validity of the proposal some examples and results regarding the computation time achieved in the implementation of the proposed algorithm on a cellular processor array (SCAMP-3 vision chip) have been included. David López Vilariño, Piotr Dudek, Diego Cabello |
ISCAS | 2 |
| 2008 | Integrated circuit implementation of a cortical neuronabstractThis paper presents an analogue integrated circuit implementation of a cortical neuron model. The VLSI chip prototype has been implemented in a 0.35 mum CMOS technology. The single neuron cell has a compact layout and very low energy consumption, in the range of 9 pJ per spike. Experimental results demonstrate the capability of the circuit to generate a realistic spike shape and a variety of spiking and bursting firing patterns. The models of various cortical neuron types are obtained in a single circuit, through the adjustment of two biasing voltages, making the circuit suitable for applications in reconfigurable neuromorphic devices that implement biologically plausible spiking neural networks. Jayawan H. B. Wijekoon, Piotr Dudek |
ISCAS | 2 |
| 2008 | Compact silicon neuron circuit with spiking and bursting behaviour
Jayawan H. B. Wijekoon, Piotr Dudek |
Neural Networks | 2 |
| 2007 | Implementation of multi-layer leaky integrator networks on a cellular processor arrayabstractWe present an application of a massively parallel processor array VLSI circuit to the implementation of neural networks in complex architectural arrangements. The work was motivated by existing biologically plausible models of a set of sub-cortical nuclei -the basal ganglia. The model includes 5 layers, each consisting of 16384 leaky integrator neurons, with inter-layer synaptic weights forming various one-to-one and diffuse connectivity patterns. The architecture of the SIMD processor array allows all the neurons per layer to be updated simultaneously. The performance of the processor array chip in simulating the model is compared with the original model being executed on a computer workstation. It is demonstrated that in this application the chip outperforms the workstation by five orders of magnitude in terms of computational performance and seven orders of magnitude in terms of energy efficiency, providing a high-speed, low-power, compact hardware platform for possible embedded robotic applications. David Robert Wallace Barr, Piotr Dudek, Jonathan M. Chambers, Kevin N. Gurney |
IJCNN | 2 |
| 2007 | Spiking and Bursting Firing Patterns of a Compact VLSI Cortical Neuron CircuitabstractThe paper presents a silicon neuron circuit that mimics the behaviour of known classes of biological neurons. The circuit has been designed in a 0.35μm CMOS technology. The firing patterns of basic cell classes:regular spiking(RS),fast spiking(FS),chattering(CH) andintrinsic bursting(IB) are obtained with a simple adjustment of two biasing voltages. The simulations reveal the potential of the circuit to provide a wide variety of cell behaviours with required accommodation and firing frequency of a given cell type. The neuron consumes only 14 MOSFETs enabling the integration of many neurons in a small silicon area. Hence, the circuit provides a foundation for designing massively parallel analogue neuromorphic networks that closely resemble the circuits of the cortex. Jayawan H. B. Wijekoon, Piotr Dudek |
IJCNN | 2 |
| 2007 | Evolution of Pixel Level Snakes towards an efficient hardware implementationabstractSince pixel level snakes (PLS) were introduced, several algorithms implementing this cellular active contour technique have been proposed. In this paper, we review the main features of these algorithms and propose some modifications to optimize the computation performance of PLS when they are executed on fine-grain pixel-parallel processor arrays. The modified algorithm has been implemented on a focal plane cellular processor array (SCAMP-3 vision chip) and tested on several applications of practical interest. David López Vilariño, Piotr Dudek |
ISCAS | 2 |
| 2006 | Architecture of a VLSI cellular processor array for synchronous/asynchronous image processingabstractThis paper describes a new architecture for a cellular processor array integrated circuit, which operates in both discrete- and continuous-time domains. Asynchronous propagation networks, enabling trigger-wave operations, distance transform calculation, and long-distance inter-processor communication, are embedded in an SIMD processor array. The proposed approach results in an architecture that is efficient in implementing both local and global image processing algorithms Alexey Lopich, Piotr Dudek |
ISCAS | 2 |
| 2000 | A CMOS general-purpose sampled-data analogue microprocessorabstractThis paper presents a general-purpose sampled-data analogue processing element that essentially functions as an analogue microprocessor (A/spl mu/P). The A/spl mu/P executes software programs, in a way akin to a digital microprocessor, while nevertheless operating on analogue sampled data values. This enables the design of mixed-mode systems which retain the speed/area/power advantages of the analogue signal processing paradigm while being fully programmable, general-purpose systems. A proof-of-concept integrated circuit has been implemented in 0.8 /spl mu/m CMOS technology, using switched-current techniques. Experimental results and examples of the application of the A/spl mu/Ps in image processing are presented. Piotr Dudek, Peter J. Hicks |
ISCAS | 1 |