EDBT 2026 Demo / reviewers in the wild / expert
Christos-Savvas Bouganis
dblp:74/5762 · also Christos Bouganis
· DBLP profile ↗
106ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0002-4906-4510ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 80 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Small Language Models on FPGAsabstractAttention is a major bottleneck when mapping Transformer-like models to FPGAs, as its matrix multiplications and normalisation stages exhibit differing numerical requirements and are highly sensitive to accumulation error. In this work, we propose operator-wise mixed-precision schemes and configurable accumulation strategies for attention-like pipelines based on shared-exponent low-bit, block floating-point style formats. By combining custom arithmetic with FPGA-specific design optimisations, our approach improves the trade-off between model quality and hardware cost, enabling more efficient deployment of small language models on reconfigurable hardware. Filip Wojcicki, Omar Sharif, Ebby Samson, Paul H. J. Kelly, George A. Constantinides, Christos-Savvas Bouganis, Wayne Luk |
FCCM | 6 |
| 2026 | Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment
Pengtao Chen, Mingzhu Shen, Peng Ye 0006, Jianjian Cao, Chongjun Tu, Christos-Savvas Bouganis, Tao Chen 0003 |
Int. J. Comput. Vis. | 6 |
| 2025 | Harmonia: A Swift and Accurate Approximate Data Structure for Real-Time Heavy Flow Detection in High-Speed Networks
Weihe Li, Tianyue Chu, Christos-Savvas Bouganis, Paul Patras |
ADMA (3) | 3 |
| 2025 | A Resource-Aware Residual-Based Gaussian Belief Propagation Accelerator Toolflow
Omar Sharif, Christos-Savvas Bouganis |
DATE | 2 |
| 2025 | ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor DecompositionabstractRecent advancements in Large Language Models (LLMs) have demonstrated impressive capabilities as their scale expands to billions of parameters. Deploying these large-scale models on resource-constrained platforms presents significant challenges, with post-training fixed-point quantization often used as a model compression technique. However, quantization-only methods typically lead to significant accuracy degradation in LLMs when precision falls below 8 bits. This paper addresses this challenge through a software-hardware co-design framework, ITERA-LLM, which integrates sub-8-bit quantization with SVD-based iterative low-rank tensor decomposition for error compensation, leading to higher compression ratios and reduced computational complexity. The proposed approach is complemented by a hardware-aware Design Space Exploration (DSE) process that optimizes accuracy, latency, and resource utilization, tailoring the configuration to the specific requirements of the targeted LLM. Our results show that ITERA-LLM achieves linear layer latency reduction of up to 41.1%, compared to quantization-only baseline approach while maintaining similar model accuracy. Yinting Huang, Keran Zheng, Zhewen Yu, Christos-Savvas Bouganis |
FCCM | 4 |
| 2025 | Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix ItabstractLabel smoothing (LS) is a popular regularisation method for training neural networks as it is effective in improving test accuracy and is simple to implement. ''Hard'' one-hot labels are ''smoothed'' by uniformly distributing probability mass to other classes, reducing overfitting. Prior work has shown that in some cases *LS can degrade selective classification (SC)* -- where the aim is to reject misclassifications using a model's uncertainty. In this work, we first demonstrate empirically across an extended range of large-scale tasks and architectures that LS *consistently* degrades SC.
We then address a gap in existing knowledge, providing an *explanation* for this behaviour by analysing logit-level gradients: LS degrades the uncertainty rank ordering of correct vs incorrect predictions by regularising the max logit *more* when a prediction is likely to be correct, and *less* when it is likely to be wrong.
This elucidates previously reported experimental results where strong classifiers underperform in SC.
We then demonstrate the empirical effectiveness of post-hoc *logit normalisation* for recovering lost SC performance caused by LS. Furthermore, linking back to our gradient analysis, we again provide an explanation for why such normalisation is effective. Guoxuan Xia, Olivier Laurent 0002, Gianni Franchi, Christos-Savvas Bouganis |
ICLR | 4 |
| 2025 | Cached Multi-Lora Composition for Multi-Concept Image GenerationabstractLow-Rank Adaptation (LoRA) has emerged as a widely adopted technique in text-to-image models, enabling precise rendering of multiple distinct elements, such as characters and styles, in multi-concept image generation. However, current approaches face significant challenges when composing these LoRAs for multi-concept image generation, particularly as the number of LoRAs increases, resulting in diminished generated image quality.
In this paper, we initially investigate the role of LoRAs in the denoising process through the lens of the Fourier frequency domain.
Based on the hypothesis that applying multiple LoRAs could lead to "semantic conflicts", we have conducted empirical experiments and find that certain LoRAs amplify high-frequency features such as edges and textures, whereas others mainly focus on low-frequency elements, including the overall structure and smooth color gradients.
Building on these insights, we devise a frequency domain based sequencing strategy to determine the optimal order in which LoRAs should be integrated during inference. This strategy offers a methodical and generalizable solution compared to the naive integration commonly found in existing LoRA fusion techniques.
To fully leverage our proposed LoRA order sequence determination method in multi-LoRA composition tasks, we introduce a novel, training-free framework, Cached Multi-LoRA (CMLoRA), designed to efficiently integrate multiple LoRAs while maintaining cohesive image generation.
With its flexible backbone for multi-LoRA fusion and a non-uniform caching strategy tailored to individual LoRAs, CMLoRA has the potential to reduce semantic conflicts in LoRA composition and improve computational efficiency.
Our experimental evaluations demonstrate that CMLoRA outperforms state-of-the-art training-free LoRA fusion methods by a significant margin -- it achieves an average improvement of $2.19$% in CLIPScore, and $11.25%$% in MLLM win rate compared to LoraHub, LoRA Composite, and LoRA Switch. Xiandong Zou, Mingzhu Shen, Christos-Savvas Bouganis |
ICLR | 3 |
| 2024 | Budget-aware Dynamic Spatially Adaptive Inference
Georgios Zampokas, Christos-Savvas Bouganis, Dimitrios Tzovaras |
BMVC | 2 |
| 2024 | A Framework for Designing Scalable Gaussian Belief Propagation Accelerators for use in SLAMabstractGaussian Belief Propagation (GBP) is an iterative method for factor graph inference that provides an approximate solution to the probability distribution of a system. It has been shown to be a powerful tool in numerous applications including SLAM, where the estimation of the robot's position and the map of the environment is required. State-of-the-art implementations suffer from scalability issues, or exhibit performance degradation when off-chip memory access is required. This paper addresses these challenges using a streaming architecture via a chain of parameterizable Processing Elements (PE) that can be tuned to the problem's characteristics through the use of an optimizer. This work overcomes the limitations of existing GBP implementations achieving 142x-168x performance improvements over an embed-ded CPU for large graphs. Omar Sharif, Christos-Savvas Bouganis |
DATE | 2 |
| 2024 | Auto WS: Automate Weights Streaming in Layer-Wise Pipelined DNN AcceleratorsabstractWith the great success of Deep Neural Networks (DNN), the design of efficient hardware accelerators has triggered wide interest in the research community. Existing research explores two architectural strategies: sequential layer execution and layer-wise pipelining. While the former supports a wider range of models, the latter is favoured for its enhanced customization and efficiency. A challenge for the layer-wise pipelining architecture is its substantial demand for the on-chip memory for weights storage, impeding the deployment of large-scale networks on resource-constrained devices. This paper introduces AutoWs,a pioneering memory management methodology that exploits both on-chip and off-chip memory to optimize weight storage within a layer-wise pipelining architecture, taking advantage of its static schedule. Through a comprehensive investigation on both the hardware design and the Design Space Exploration, our methodology is fully automated and enables the deployment of large-scale DNN models on resource-constrained devices, which was not possible in existing works that target layer-wise pipelining architectures. AutoWS is open-source: https://github.com/Yu-Zhewen/AutoWS. Zhewen Yu, Christos-Savvas Bouganis |
DATE | 2 |
| 2024 | SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip EvictionabstractConvolutional Neural Networks (CNNs) have demonstrated their effectiveness in numerous vision tasks. However, their high processing requirements necessitate efficient hardware acceleration to meet the application's performance targets. In the space of FPGAs, streaming-based dataflow architectures are often adopted by users, as significant performance gains can be achieved through layer-wise pipelining and reduced off-chip memory access by retaining data on-chip. However, modern topologies, such as the UNet, YOLO, and X3D models, utilise long skip connections, requiring significant on-chip storage and thus limiting the performance achieved by such system architectures. The paper addresses the above limitation by introducing weight and activation eviction mechanisms to off-chip memory along the computational pipeline, taking into account the available compute and memory resources. The proposed mechanism is incorporated into an existing toolflow, expanding the design space by utilising off-chip memory as a buffer. This enables the mapping of such modern CNNs to devices with limited on-chip memory, under the streaming architecture design approach. SMOF has demonstrated the capacity to deliver competitive and, in some cases, state-of-the-art performance across a spectrum of computer vision tasks, achieving up to 10.65 × throughput improvement compared to previous works. The tool is available at https://github.com/ICIdsl/smof.git. Petros Toupas, Zhewen Yu, Christos-Savvas Bouganis, Dimitrios Tzovaras |
FCCM | 3 |
| 2024 | HASS: Hardware-Aware Sparsity Search for Dataflow DNN AcceleratorabstractDeep Neural Networks (DNNs) excel in learning hierarchical representations from raw data, such as images, audio, and text. To compute these DNN models with high performance and energy efficiency, these models are usually deployed onto customized hardware accelerators. Among various accelerator designs, dataflow architecture has shown promising performance due to its layer-pipelined structure and its scalability in data parallelism.Exploiting weights and activations sparsity can further enhance memory storage and computation efficiency. However, existing approaches focus on exploiting sparsity in non-dataflow accelerators, which cannot be applied onto dataflow accelerators because of the large hardware design space introduced. As such, this could miss opportunities to find an optimal combination of sparsity features and hardware designs.In this paper, we propose a novel approach to exploit unstructured weights and activations sparsity for dataflow accelerators, using software and hardware co-optimization. We propose a Hardware-Aware Sparsity Search (HASS) to systematically determine an efficient sparsity solution for dataflow accelerators. Over a set of models, we achieve an efficiency improvement ranging from $1.3 \times$ to $4.2 \times$ compared to existing sparse designs, which are either non-dataflow or non-hardware-aware. Particularly, the throughput of MobileNetV3 can be optimized to 4895 images per second. HASS is open-source: https://github.com/Yu-Zhewen/HASS Zhewen Yu, Sudarshan Sreeram, Krish Agrawal, Alexander Montgomerie-Corcoran, Jianyi Cheng, Christos-Savvas Bouganis |
FPL | 8 |
| 2024 | Augmenting the Softmax with Additional Confidence Scores for Improved Selective Classification with Out-of-Distribution DataabstractAbstract Detecting out-of-distribution (OOD) data is a task that is receiving an increasing amount of research attention in the domain of deep learning for computer vision. However, the performance of detection methods is generally evaluated on the task in isolation, rather than also considering potential downstream tasks in tandem. In this work, we examine selective classification in the presence of OOD data (SCOD). That is to say, the motivation for detecting OOD samples is to reject them so their impact on the quality of predictions is reduced. We show under this task specification, that existing post-hoc methods perform quite differently compared to when evaluated only on OOD detection. This is because it is no longer an issue to conflate in-distribution (ID) data with OOD data if the ID data is going to be misclassified. However, the conflation within ID data of correct and incorrect predictions becomes undesirable. We also propose a novel method for SCOD, Softmax Information Retaining Combination (SIRC), that augments a softmax-based confidence score with a secondary class-agnostic feature-based score. Thus, the ability to identify OOD samples is improved without sacrificing separation between correct and incorrect ID predictions. Experiments on a wide variety of ImageNet-scale datasets and convolutional neural network architectures show that SIRC is able to consistently match or outperform the baseline for SCOD, whilst existing OOD detection methods fail to do so. Interestingly, we find that the secondary scores investigated for SIRC do not consistently improve performance on all tested OOD datasets. To address this issue, we further extend SIRC to incorporate multiple secondary scores (SIRC+). This further improves SCOD performance, both generally, and in terms of consistency over diverse distribution shifts. Code is available at https://github.com/Guoxoug/SIRC . Guoxuan Xia, Christos-Savvas Bouganis |
Int. J. Comput. Vis. | 2 |
| 2023 | FMM-X3D: FPGA-Based Modeling and Mapping of X3D for Human Action Recognitionabstract3D Convolutional Neural Networks are gaining increasing attention from researchers and practitioners and have found applications in many domains, such as surveillance systems, autonomous vehicles, human monitoring systems, and video retrieval. However, their widespread adoption is hindered by their high computational and memory requirements, especially when resource-constrained systems are targeted. This paper addresses the problem of mapping X3D, a state-of-the-art model in Human Action Recognition that achieves accuracy of 95.5% in the UCF101 benchmark, onto any FPGA device. The proposed toolflow generates an optimised stream-based hardware system, taking into account the available resources and off-chip memory characteristics of the FPGA device. The generated designs push further the current performance-accuracy pareto front, and enable for the first time the targeting of such complex model architectures for the Human Action Recognition task. Petros Toupas, Christos-Savvas Bouganis, Dimitrios Tzovaras |
ASAP | 2 |
| 2023 | ATHEENA: A Toolflow for Hardware Early-Exit Network AutomationabstractThe continued need for improvements in accuracy, throughput, and efficiency of Deep Neural Networks has resulted in a multitude of methods that make the most of custom architectures on FPGAs. These include the creation of hand-crafted networks and the use of quantization and pruning to reduce extraneous network parameters. However, with the potential of static solutions already well exploited, we propose to shift the focus to using the varying difficulty of individual data samples to further improve efficiency and reduce average compute for classification. Input-dependent computation allows for the network to make runtime decisions to finish a task early if the result meets a confidence threshold. Early-Exit network architectures have become an increasingly popular way to implement such behaviour in software. We create A Toolflow for Hardware Early-Exit Network Automation (ATHEENA), an automated FPGA toolflow that leverages the probability of samples exiting early from such networks to scale the resources allocated to different sections of the network. The toolflow uses the data-flow model of fpgaConvNet, extended to support Early-Exit networks as well as Design Space Exploration to optimize the generated streaming architecture hardware with the goal of increasing throughput/reducing area while maintaining accuracy. Experimental results on three different networks demonstrate a throughput increase of 2.00× to 2.78× compared to an optimized baseline network implementation with no early exits. Additionally, the toolflow can achieve a throughput matching the same baseline with as low as 46% of the resources the baseline requires. Benjamin Biggs, Christos-Savvas Bouganis, George A. Constantinides |
FCCM | 2 |
| 2023 | HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA DevicesabstractFor Human Action Recognition tasks (HAR), 3D Convolutional Neural Networks have proven to be highly effective, achieving state-of-the-art results. This study introduces a novel streaming architecture-based toolflow for mapping such models onto FPGAs considering the model's inherent characteristics and the features of the targeted FPGA device. The HARFLOW3D toolflow takes as input a 3D CNN in ONNX format and a description of the FPGA characteristics, generating a design that minimises the latency of the computation. The toolflow is comprised of a number of parts, including (i) a 3D CNN parser, (ii) a performance and resource model, (iii) a scheduling algorithm for executing 3D models on the generated hardware, (iv) a resource-aware optimisation engine tailored for 3D models, (v) an automated mapping to synthesizable code for FPGAs. The ability of the toolflow to support a broad range of models and devices is shown through a number of experiments on various 3D CNN and FPGA system pairs. Furthermore, the toolflow has produced high-performing results for 3D CNN models that have not been mapped to FPGAs before, demonstrating the potential of FPGA-based systems in this space. Overall, HARFLOW3D has demonstrated its ability to deliver competitive latency compared to a range of state-of-the-art hand-tuned approaches, being able to achieve up to 5× better performance compared to some of the existing works. The tool is available at https://github.com/ptoupas/harflow3d. Petros Toupas, Alexander Montgomerie-Corcoran, Christos-Savvas Bouganis, Dimitrios Tzovaras |
FCCM | 3 |
| 2023 | PASS: Exploiting Post-Activation Sparsity in Streaming Architectures for CNN AccelerationabstractWith the ever-growing popularity of Artificial Intelligence, there is an increasing demand for more performant and efficient underlying hardware. Convolutional Neural Networks (CNN) are a workload of particular importance, which achieve high accuracy in computer vision applications. Inside CNNs, a significant number of the post-activation values are zero, resulting in many redundant computations. Recent works have explored this post-activation sparsity on instruction-based CNN accelerators but not on streaming CNN accelerators, despite the fact that streaming architectures are considered the leading design methodology in terms of performance. In this paper, we highlight the challenges associated with exploiting post-activation sparsity for performance gains in streaming CNN accelerators, and demonstrate our approach to address them. Using a set of modern CNN benchmarks, our streaming sparse accelerators achieve 1.41 x to 1.93 x efficiency (GOP/sDSP) compared to state-of-the-art instruction-based sparse accelerators. Alexander Montgomerie-Corcoran, Zhewen Yu, Jianyi Cheng, Christos-Savvas Bouganis |
FPL | 4 |
| 2023 | fpgaHART: A Toolflow for Throughput-Oriented Acceleration of 3D CNNs for HAR onto FPGAsabstractSurveillance systems, autonomous vehicles, human monitoring systems, and video retrieval are just few of the many applications in which 3D Convolutional Neural Networks are exploited. However, their extensive use is restricted by their high computational and memory requirements, especially when integrated into systems with limited resources. This study proposes a toolflow that optimises the mapping of 3D CNN models for Human Action Recognition onto FPGA devices, taking into account FPGA resources and off-chip memory characteristics. The proposed system employs Synchronous Dataflow (SDF) graphs to model the designs and introduces transformations to expand and explore the design space, resulting in high-throughput designs. A variety of 3D CNN models were evaluated using the proposed toolflow on multiple FPGA devices, demonstrating its potential to deliver competitive performance compared to earlier hand-tuned and model-specific designs. Petros Toupas, Christos-Savvas Bouganis, Dimitrios Tzovaras |
FPL | 2 |
| 2023 | Mixed-TD: Efficient Neural Network Accelerator with Layer-Specific Tensor DecompositionabstractNeural Network designs are quite diverse, from VGG-style to ResNet-style, and from Convolutional Neural Networks to Transformers. Towards the design of efficient accelerators, many works have adopted a dataflow-based, inter-layer pipelined architecture, with a customized hardware towards each layer, achieving ultra high throughput and low latency. The deployment of neural networks to such dataflow architecture accelerators is usually hindered by the available on-chip memory as it is desirable to preload the weights of neural networks on-chip to maximise the system performance. To address this, networks are usually compressed before the deployment through methods such as pruning, quantization and tensor decomposition. In this paper, a framework for mapping CNNs onto FPGAs based on a novel tensor decomposition method called Mixed-TD is proposed. The proposed method applies layer-specific Singular Value Decomposition (SVD) and Canonical Polyadic Decomposition (CPD) in a mixed manner, achieving 1.73× to 10.29× throughput per DSP to state-of-the-art CNNs. Our work is open-sourced: https://github.com/Yu-Zhewen/Mixed-TD. Zhewen Yu, Christos-Savvas Bouganis |
FPL | 2 |
| 2023 | Window-Based Early-Exit Cascades for Uncertainty Estimation: When Deep Ensembles are More Efficient than Single ModelsabstractDeep Ensembles are a simple, reliable, and effective method of improving both the predictive performance and uncertainty estimates of deep learning approaches. However, they are widely criticised as being computationally expensive, due to the need to deploy multiple independent models. Recent work has challenged this view, showing that for predictive accuracy, ensembles can be more computationally efficient (at inference) than scaling single models within an architecture family. This is achieved by cascading ensemble members via an early-exit approach. In this work, we investigate extending these efficiency gains to tasks related to uncertainty estimation. As many such tasks, e.g. selective classification, are binary classification, our key novel insight is to only pass samples within a window close to the binary decision boundary to later cascade stages. Experiments on ImageNet-scale data across a number of network architectures and uncertainty tasks show that the proposed window-based early-exit approach is able to achieve a superior uncertainty-computation trade-off compared to scaling single models. For example, a cascaded EfficientNet-B2 ensemble is able to achieve similar coverage at 5% risk as a single EfficientNet-B4 with <30% the number of MACs. We also find that cascades/ensembles give more reliable improvements on OOD data vs scaling models up. Code for this work is available at: https://github.com/Guoxoug/window-early-exit. Guoxuan Xia, Christos-Savvas Bouganis |
ICCV | 2 |
| 2023 | SVD-NAS: Coupling Low-Rank Approximation and Neural Architecture SearchabstractThe task of compressing pre-trained Deep Neural Networks has attracted wide interest of the research community due to its great benefits in freeing practitioners from data access requirements. In this domain, low-rank approximation is a promising method, but existing solutions considered a restricted number of design choices and failed to efficiently explore the design space, which lead to severe accuracy degradation and limited compression ratio achieved. To address the above limitations, this work proposes the SVD-NAS framework that couples the domains of low-rank approximation and neural architecture search. SVD-NAS generalises and expands the design choices of previous works by introducing the Low-Rank architecture space, LR-space, which is a more fine-grained design space of low-rank approximation. Afterwards, this work proposes a gradient-descent-based search for efficiently traversing the LR-space. This finer and more thorough exploration of the possible design choices results in improved accuracy as well as reduction in parameters, FLOPS, and latency of a CNN model. Results demonstrate that the SVD-NAS achieves 2.06-12.85pp higher accuracy on ImageNet than state-of-the-art methods under the data-limited problem setting. SVD-NAS is open-sourced at https://github.com/Yu-Zhewen/SVD-NAS. Zhewen Yu, Christos-Savvas Bouganis |
WACV | 2 |
| 2022 | Augmenting Softmax Information for Selective Classification with Out-of-Distribution Data
Guoxuan Xia, Christos-Savvas Bouganis |
ACCV (6) | 2 |
| 2022 | SAMO: Optimised Mapping of Convolutional Neural Networks to Streaming ArchitecturesabstractSignificant effort has been placed on the development of toolflows that map Convolutional Neural Network (CNN) models to Field Programmable Gate Arrays (FPGAs) with the aim of automating the production of high performance designs for a diverse set of applications. However, within these toolflows, the problem of finding an optimal mapping is often overlooked, with the expectation that the end user will tune their generated hardware for their desired platform. This is particularly prominent within Streaming Architecture toolflows, where there is a large design space to be explored. In this work, we establish the framework SAMO: a Streaming Architecture Mapping Optimiser. SAMO exploits the structure of CNN models and the common features that exist in Streaming Architectures, and casts the mapping optimisation problem under a unified methodology. Furthermore, SAMO explicitly explores the re-configurability property of FPGAs, allowing the methodology to overcome mapping limitations imposed by certain toolflows under resource-constrained scenarios, as well as improve on the achievable throughput. Three optimisation methods - Brute-Force, Simulated Annealing and Rule-Based - have been developed in order to generate valid, high performance designs for a range of target platforms and CNN models. Results show that SAMO-optimised designs can achieve 4x-20x better performance compared to existing hand-tuned designs. The SAMO framework is open-source: https://github.com/AlexMontgomerie/samo. Alexander Montgomerie-Corcoran, Zhewen Yu, Christos-Savvas Bouganis |
FPL | 3 |
| 2021 | Learning Boolean Circuits from Examples for Approximate Logic SynthesisabstractMany computing applications are inherently error resilient. Thus, it is possible to decrease computing accuracy to achieve greater efficiency in area, performance, and/or energy consumption. In recent years, a slew of automatic techniques for approximate computing has been proposed; however, most of these techniques require full knowledge of an exact, or 'golden' circuit description. In contrast, there has been significant recent interest in synthesizing computation from examples, a form of supervised learning. In this paper, we explore the relationship between supervised learning of Boolean circuits and existing work on synthesizing incompletely-specified functions. We show that when considered through a machine learning lens, the latter work provides a good training accuracy but poor test accuracy. We contrast this with prior work from the 1990s which uses mutual information to steer the search process, aiming for good generalization. By combining this early work with a recent approach to learning logic functions, we are able to achieve a scalable and efficient machine learning approach for Boolean circuits in terms of area/delay/test-error trade-off. Sina Boroumand, Christos-Savvas Bouganis, George A. Constantinides |
ASP-DAC | 2 |
| 2021 | DEF: Differential Encoding of Featuremaps for Low Power Convolutional Neural Network AcceleratorsabstractAs the need for the deployment of Deep Learning applications on edge-based devices becomes ever increasingly prominent, power consumption starts to become a limiting factor on the performance that can be achieved by the computational platforms. A significant source of power consumption for these edge-based machine learning accelerators is off-chip memory transactions. In the case of Convolutional Neural Network (CNN) workloads, a predominant workload in deep learning applications, those memory transactions are typically attributed to the store and recall of feature-maps. There is therefore a need to explicitly reduce the power dissipation of these transactions whilst minimising any overheads needed to do so. In this work, a Differential Encoding of Feature-maps (DEF) scheme is proposed, which aims at minimising activity on the memory data bus, specifically for CNN workloads. The coding scheme uses domain-specific knowledge, exploiting statistics of feature-maps alongside knowledge of the data types commonly used in machine learning accelerators as a means of reducing power consumption. DEF is able to out-perform recent state-of-the-art coding schemes, with significantly less overhead, achieving up to 50% reduction of activity across a number of modern CNNs. Alexander Montgomerie-Corcoran, Christos-Savvas Bouganis |
ASP-DAC | 2 |
| 2021 | POMMEL: Exploring Off-Chip Memory Energy & Power Consumption in Convolutional Neural Network AcceleratorsabstractReducing the power and energy consumption of Convolutional Neural Network (CNN) Accelerators is becoming an increasingly popular design objective for both cloud and edge-based settings. Aiming towards the design of more efficient accelerator systems, the accelerator architect must understand how different design choices impact both power and energy consumption. The purpose of this work is to enable CNN accelerator designers to explore how design choices affect the memory subsystem in particular, which is a significant contributing component. By considering high-level design parameters of CNN accelerators that affect the memory subsystem, the proposed tool returns power and energy consumption estimates for a range of networks and memory types. This allows for power and energy of the off-chip memory subsystem to be considered earlier within the design process, enabling greater optimisations at the beginning phases. Towards this, the paper introduces POMMEL, an off-chip memory subsystem modelling tool for CNN accelerators, and its evaluation across a range of accelerators, networks, and memory types is performed. Furthermore, using POMMEL, the impact of various state-of-the-art compression and activity reduction schemes on the power and energy consumption of current accelerations is also investigated. Alexander Montgomerie-Corcoran, Christos-Savvas Bouganis |
DSD | 2 |
| 2021 | StreamSVD: Low-rank Approximation and Streaming Accelerator Co-designabstractThe post-training compression of a Convolutional Neural Network (CNN) aims to produce Pareto-optimal designs on the accuracy-performance frontier when the access to training data is not possible. Low-rank approximation is one of the methods that is often utilised in such cases. However, existing work considers the low-rank approximation of the network and the optimisation of the hardware accelerator separately, leading to systems with sub-optimal performance. This work focuses on the efficient mapping of a CNN into an FPGA device, and presents StreamSVD, a model-accelerator co-design framework1. The framework considers simultaneously the compression of a CNN model through a hardware-aware low-rank approximation scheme, and the optimisation of the hardware accelerator's architecture by taking into account the approximation scheme's compute structure. Our results show that the co-designed StreamSVD outperforms existing work that utilises similar low-rank approximation schemes by providing better accuracy-throughput trade-off. The proposed framework also achieves competitive performance compared with other post-training compression methods, even outperforming them under certain cases. Zhewen Yu, Christos-Savvas Bouganis |
FPT | 2 |
| 2020 | A Throughput-Latency Co-Optimised Cascade of Convolutional Neural Network ClassifiersabstractConvolutional Neural Networks constitute a prominent AI model for classification tasks, serving a broad span of diverse application domains. To enable their efficient deployment in real-world tasks, the inherent redundancy of CNNs is frequently exploited to eliminate unnecessary computational costs. Driven by the fact that not all inputs require the same amount of computation to drive a confident prediction, multi-precision cascade classifiers have been recently introduced. FPGAs comprise a promising platform for the deployment of such input-dependent computation models, due to their enhanced customisation capabilities. Current literature, however, is limited to throughput-optimised cascade implementations, employing large batching at the expense of a substantial latency aggravation prohibiting their deployment on real-time scenarios. In this work, we introduce a novel methodology for throughput-latency co-optimised cascaded CNN classification, deployed on a custom FPGA architecture tailored to the target application and deployment platform, with respect to a set of user-specified requirements on accuracy and performance. Our experiments indicate that the proposed approach achieves comparable throughput gains with related state-of-the-art works, under substantially reduced overhead in latency, enabling its deployment on latency-sensitive applications. Alexandros Kouris, Stylianos I. Venieris, Christos-Savvas Bouganis |
DATE | 3 |
| 2020 | Caffe Barista: Brewing Caffe with FPGAs in the Training LoopabstractAs the complexity of deep learning (DL) modelsincreases, their compute requirements increase accordingly. De-ploying a Convolutional Neural Network (CNN) involves twophases: training and inference. With the inference task typicallytaking place on resource-constrained devices, a lot of research hasexplored the field of low-power inference on custom hardwareaccelerators. On the other hand, training is both more compute-and memory-intensive and is primarily performed on power-hungry GPUs in large-scale data centres. CNN training onFPGAs is a nascent field of research. This is primarily due tothe lack of tools to easily prototype and deploy various hardwareand/or algorithmic techniques for power-efficient CNN training. This work presentsBarista, an automated toolflow that providesseamless integration of FPGAs into the training of CNNs withinthe popular deep learning framework Caffe. To the best of ourknowledge, this is the only tool that allows for such versatile andrapid deployment of hardware and algorithms for the FPGA-based training of CNNs, providing the necessary infrastructurefor further research and development. Diederik Adriaan Vink, Aditya Rajagopal, Stylianos I. Venieris, Christos-Savvas Bouganis |
FPL | 4 |
| 2020 | Multi-Precision Policy Enforced Training (MuPPET) : A Precision-Switching Strategy for Quantised Fixed-Point Training of CNNsabstractLarge-scale convolutional neural networks (CNNs) suffer from very long training times, spanning from hours to weeks, limiting the productivity and experimentation of deep learning practitioners. As networks grow in size and complexity, training time can be reduced through low-precision data representations and computations, however, in doing so the final accuracy suffers due to the problem of vanishing gradients. Existing state-of-the-art methods combat this issue by means of a mixed-precision approach utilising two different precision levels, FP32 (32-bit floating-point) and FP16/FP8 (16-/8-bit floating-point), leveraging the hardware support of recent GPU architectures for FP16 operations to obtain performance gains. This work pushes the boundary of quantised training by employing a multilevel optimisation approach that utilises multiple precisions including low-precision fixed-point representations resulting in a novel training strategy MuPPET; it combines the use of multiple number representation regimes together with a precision-switching mechanism that decides at run time the transition point between precision regimes. Overall, the proposed strategy tailors the training process to the hardware-level capabilities of the target hardware architecture and yields improvements in training time and energy efficiency compared to state-of-the-art approaches. Applying MuPPET on the training of AlexNet, ResNet18 and GoogLeNet on ImageNet (ILSVRC12) and targeting an NVIDIA Turing GPU, MuPPET achieves the same accuracy as standard full-precision training with training-time speedup of up to 1.84x and an average speedup of 1.58x across the networks. Aditya Rajagopal, Diederik Adriaan Vink, Stylianos I. Venieris, Christos-Savvas Bouganis |
ICML | 4 |
| 2019 | Optimising 3D-CNN Design towards Human Pose Estimation on Low Power Devices
Manolis Vasileiadis, Christos-Savvas Bouganis, Georgios Stavropoulos, Dimitrios Tzovaras |
BMVC | 2 |
| 2019 | Informed Region Selection for Efficient UAV-based Object Detectors: Altitude-aware Vehicle Detection with CyCAR DatasetabstractDeep Learning-based object detectors enhance the capabilities of remote sensing platforms, such as Unmanned Aerial Vehicles (UAVs), in a wide spectrum of machine vision applications. However, the integration of deep learning introduces heavy computational requirements, preventing the deployment of such algorithms in scenarios that impose low-latency constraints during inference, in order to make mission-critical decisions in real-time. In this paper, we address the challenge of efficient deployment of region-based object detectors in aerial imagery, by introducing an informed methodology for extracting candidate detection regions (proposals). Our approach considers information from the UAV on-board sensors, such as flying altitude and light-weight computer vision filters, along with prior domain knowledge to intelligently decrease the number of region proposals by eliminating false-positives at an early stage of the computation, reducing significantly the computational workload while sustaining the detection accuracy. We apply and evaluate the proposed approach on the task of vehicle detection. Our experiments demonstrate that state-of-the-art detection models can achieve up to 2.6x faster inference by employing our altitude-aware data-driven methodology. Alongside, we introduce and provide to the community a novel vehicle-annotated and altitude-stamped dataset of real UAV imagery, captured at numerous flying heights under a wide span of traffic scenarios. Alexandros Kouris, Christos Kyrkou, Christos-Savvas Bouganis |
IROS | 3 |
| 2019 | Multi-person 3D pose estimation from 3D cloud data using 3D convolutional neural networks
Manolis Vasileiadis, Christos-Savvas Bouganis, Dimitrios Tzovaras |
Comput. Vis. Image Underst. | 2 |
| 2019 | Scaling Up Modulo Scheduling for High-Level SynthesisabstractHigh-Level Synthesis tools have been increasingly used within the hardware design community to bridge the gap between productivity and the need to design large and complex systems. When targeting heterogeneous systems, where the CPU and the FPGA fabric are both available to perform computations, a design space exploration is usually carried out for deciding which parts of the initial code should be mapped to the FPGA fabric such as the overall system’s performance is enhanced by accelerating its computation via dedicated processors. As the targeted systems become more complex and larger, leading to a large design space exploration, the fast estimative of the possible acceleration that can be obtained by mapping certain functionality into the FPGA fabric is of paramount importance. Loop pipelining, which is responsible for the majority of HLS compilation time, is a key optimization towards achieving high-performance acceleration kernels. A new modulo scheduling algorithm is proposed, which reformulates the classical modulo scheduling problem and leads to a reduced number of integer linear problems solved, resulting in large computational savings. Moreover, the proposed approach has a controlled trade-off between solution quality and computation time. Results show the scalability is improved efficiently from quadratic, for the state-of-the-art method, to linear, for the proposed approach, while the optimized loop suffers a 1% (geomean) increment in the total number of cycles. Leandro de Souza Rosa, Christos-Savvas Bouganis, Vanderlei Bonato |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | fpgaConvNet: Mapping Regular and Irregular Convolutional Neural Networks on FPGAsabstractSince neural networks renaissance, convolutional neural networks (ConvNets) have demonstrated a state-of-the-art performance in several emerging artificial intelligence tasks. The deployment of ConvNets in real-life applications requires power-efficient designs that meet the application-level performance needs. In this context, field-programmable gate arrays (FPGAs) can provide a potential platform that can be tailored to application-specific requirements. However, with the complexity of ConvNet models increasing rapidly, the ConvNet-to-FPGA design space becomes prohibitively large. This paper presents fpgaConvNet, an end-to-end framework for the optimized mapping of ConvNets on FPGAs. The proposed framework comprises an automated design methodology based on the synchronous dataflow (SDF) paradigm and defines a set of SDF transformations in order to efficiently navigate the architectural design space. By proposing a systematic multiobjective optimization formulation, the presented framework is able to generate hardware designs that are cooptimized for the ConvNet workload, the target device, and the application's performance metric of interest. Quantitative evaluation shows that the proposed methodology yields hardware designs that improve the performance by up to 6.65× over highly optimized graphics processing unit designs for the same power constraints and achieve up to 2.94× higher performance density compared with the state-of-the-art FPGA-based ConvNet architectures. Stylianos I. Venieris, Christos-Savvas Bouganis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | An overview of next-generation architectures for machine learning: Roadmap, opportunities and challenges in the IoT eraabstractThe number of connected Internet of Things (IoT) devices are expected to reach over 20 billion by 2020. These range from basic sensor nodes that log and report the data to the ones that are capable of processing the incoming information and taking an action accordingly. Machine learning, and in particular deep learning, is the de facto processing paradigm for intelligently processing these immense volumes of data. However, the resource inhibited environment of IoT devices, owing to their limited energy budget and low compute capabilities, render them a challenging platform for deployment of desired data analytics. This paper provides an overview of the current and emerging trends in designing highly efficient, reliable, secure and scalable machine learning architectures for such devices. The paper highlights the focal challenges and obstacles being faced by the community in achieving its desired goals. The paper further presents a roadmap that can help in addressing the highlighted challenges and thereby designing scalable, high-performance, and energy efficient architectures for performing machine learning on the edge. Muhammad Shafique 0001, Theocharis Theocharides, Christos-Savvas Bouganis, Muhammad Abdullah Hanif, Faiq Khalid, Rehan Hafiz, Semeen Rehman |
DATE | 3 |
| 2018 | DroNet: Efficient convolutional neural network detector for real-time UAV applicationsabstractUnmanned Aerial Vehicles (drones) are emerging as a promising technology for both environmental and infrastructure monitoring, with broad use in a plethora of applications. Many such applications require the use of computer vision algorithms in order to analyse the information captured from an on-board camera. Such applications include detecting vehicles for emergency response and traffic monitoring. This paper therefore, explores the trade-offs involved in the development of a single-shot object detector based on deep convolutional neural networks (CNNs) that can enable UAVs to perform vehicle detection under a resource constrained environment such as in a UAV. The paper presents a holistic approach for designing such systems; the data collection and training stages, the CNN architecture, and the optimizations necessary to efficiently map such a CNN on a lightweight embedded processing platform suitable for deployment on UAVs. Through the analysis we propose a CNN architecture that is capable of detecting vehicles from aerial UAV images and can operate between 5-18 frames-per-second for a variety of platforms with an overall accuracy of ~ 95%. Overall, the proposed architecture is suitable for UAV applications, utilizing low-power embedded processors that can be deployed on commercial UAVs. Christos Kyrkou, George Plastiras, Theocharis Theocharides, Stylianos I. Venieris, Christos-Savvas Bouganis |
DATE | 5 |
| 2018 | Cascade^CNN: Pushing the Performance Limits of Quantisation in Convolutional Neural NetworksabstractThis work presents CascadeCNN, an automated toolflow that pushes the quantisation limits of any given CNN model, aiming to perform high-throughput inference. A two-stage architecture tailored for any given CNN-FPGA pair is generated, consisting of a low-and high-precision unit in a cascade. A confidence evaluation unit is employed to identify misclassified cases from the excessively low-precision unit and forward them to the high-precision unit for re-processing. Experiments demonstrate that the proposed toolflow can achieve a performance boost up to 55% for VGG-16 and 48% for AlexNet over the baseline design for the same resource budget and accuracy, without the need of retraining the model or accessing the training data. Alexandros Kouris, Stylianos I. Venieris, Christos-Savvas Bouganis |
FPL | 3 |
| 2018 | f-CNNx: A Toolflow for Mapping Multiple Convolutional Neural Networks on FPGAsabstractThe predictive power of Convolutional Neural Networks (CNNs) has been an integral factor for emerging latency-sensitive applications, such as autonomous drones and vehicles. Such systems employ multiple CNNs, each one trained for a particular task. The efficient mapping of multiple CNNs on a single FPGA device is a challenging task as the allocation of compute resources and external memory bandwidth needs to be optimised at design time. This paper proposes f-CNNx, an automated toolflow for the optimised mapping of multiple CNNs on FPGAs, comprising a novel multi-CNN hardware architecture together with an automated design space exploration method that considers the user-specified performance requirements for each model to allocate compute resources and generate a synthesisable accelerator. Moreover, f-CNNx employs a novel scheduling algorithm that alleviates the limitations of the memory bandwidth contention between CNNs and sustains the high utilisation of the architecture. Experimental evaluation shows that f-CNNx's designs outperform contention-unaware FPGA mappings by up to 50% and deliver up to 6.8x higher performance-per-Watt over highly optimised GPU designs for multi-CNN systems. Stylianos I. Venieris, Christos-Savvas Bouganis |
FPL | 2 |
| 2018 | Scaling Up Loop Pipelining for High-Level Synthesis: A Non-iterative ApproachabstractHigh-level synthesis is a powerful tool for increasing productivity in digital hardware design. However, as digital systems become larger and more complex, designers have to consider an increased number of optimizations and directives offered by high-level synthesis tools to control the hardware generation process, resulting in a large design space to be explored. One of the most impactful optimizations is loop pipelining due to its large improvement in the hardware throughput. Nevertheless, the modulo scheduling algorithms that are used for loop pipelining are computationally expensive, and their application to the whole design space can make its exploration inviable, leading to sub-optimum solutions. Current state-of-the-art tools for modulo scheduling follow an iterative approach, which solves O(n2) optimization problems, where n is the loop code size. To address this problem, this work proposes a novel data-flow-based approach that solves exactly 2 optimization problems, independently of the loop code size. Results show orders-of-magnitude savings in the computation time, leading to significant design space exploration time savings when compared with the state-of-the-art. As such, the proposed method produces hardware designs of higher performance than the ones produced by the current state of the art for large and complex loops, maintaining a similar resource utilization. Leandro de Souza Rosa, Vanderlei Bonato, Christos-Savvas Bouganis |
FPT | 3 |
| 2018 | Learning to Fly by MySelf: A Self-Supervised CNN-Based Approach for Autonomous NavigationabstractNowadays, Unmanned Aerial Vehicles (UAVs)are becoming increasingly popular facilitated by their extensive availability. Autonomous navigation methods can act as an enabler for the safe deployment of drones on a wide range of real-world civilian applications. In this work, we introduce a self-supervised CNN-based approach for indoor robot navigation. Our method addresses the problem of real-time obstacle avoidance, by employing a regression CNN that predicts the agent's distance-to-collision in view of the raw visual input of its on-board monocular camera. The proposed CNN is trained on our custom indoor-flight dataset which is collected and annotated with real-distance labels, in a self-supervised manner using external sensors mounted on an UAV. By simultaneously processing the current and previous input frame, the proposed CNN extracts spatio-temporal features that encapsulate both static appearance and motion information to estimate the robot's distance to its closest obstacle towards multiple directions. These predictions are used to modulate the yaw and linear velocity of the UAV, in order to navigate autonomously and avoid collisions. Experimental evaluation demonstrates that the proposed approach learns a navigation policy that achieves high accuracy on real-world indoor flights, outperforming previously proposed methods from the literature. Alexandros Kouris, Christos-Savvas Bouganis |
IROS | 2 |
| 2017 | Communication-Aware MCMC Method for Big Data Applications on FPGAsabstractMarkov Chain Monte Carlo (MCMC) based methods have been the main tool for Bayesian Inference for some years now, and recently they find increasing applications in modern statistics and machine learning. Nevertheless, with the availability of large datasets and increasing complexity of Bayesian models, MCMC methods are becoming prohibitively expensive for real-world problems. At the heart of these methods, lies the computation of likelihood functions that requires access to all input data points in each iteration of the method. Current approaches, based on data subsampling, aim to accelerate these algorithms by reducing the number of the data points for likelihood evaluations at each MCMC iteration. However the existing work doesn't consider the properties of modern memory hierarchies, but treats the memory as one monolithic storage space. This paper proposes a communication-aware MCMC framework that takes into account the underlying performance of the memory subsystem. The framework is based on a novel subsampling algorithm that utilises an unbiased likelihood estimator based on Probability Proportional-to-Size (PPS) sampling, allowing information on the performance of the memory system to be taken into account during the sampling stage. The proposed MCMC sampler is mapped to an FPGA device and its performance is evaluated using the Bayesian logistic regression model on MNIST dataset. The proposed system achieves a 3.37× speed up over a highly optimised traditional FPGA design, therefore the risk in the estimates based on the generated samples is largely decreased. Shuanglong Liu, Christos-Savvas Bouganis |
FCCM | 2 |
| 2017 | fpgaConvNet: Automated Mapping of Convolutional Neural Networks on FPGAs (Abstract Only)
Stylianos I. Venieris, Christos-Savvas Bouganis |
FPGA | 2 |
| 2017 | A high-performance system-on-chip architecture for direct tracking for SLAMabstractSimultaneous Localization and Mapping or SLAM, is a family of algorithms that solve the problem of estimating an observer's position in an unknown environment while generating a map of that environment. SLAM algorithms that produce high quality dense maps require powerful hardware platforms. In the simultaneous solution of these two problems, Localization, also known as Tracking, is the one that is latency sensitive and needs a sustained high framerate. This work focuses on providing an efficient, high-performance solution for Direct Tracking using a high bandwidth streaming architecture, optimized for maximum memory throughput. At its centre is a Tracking Core that performs non-linear least-squares optimization for direct whole-image alignment. The architecture is designed to scale with the available hardware resources in order to enable its use for different performance/cost levels and platforms. An initial implementation tested with a Zynq System-on-Chip can process and track more than 22 frames/second with an embedded power budget and achieves a 5× improvement over previous work on FPGA SoCs. Konstantinos Boikos, Christos-Savvas Bouganis |
FPL | 2 |
| 2017 | Latency-driven design for FPGA-based convolutional neural networksabstractIn recent years, Convolutional Neural Networks (ConvNets) have become the quintessential component of several state-of-the-art Artificial Intelligence tasks. Across the spectrum of applications, the performance needs vary significantly, from high-throughput image recognition to the very low-latency requirements of autonomous cars. In this context, FPGAs can provide a potential platform that can be optimally configured based on different performance requirements. However, with the increasing complexity of ConvNet models, the architectural design space becomes overwhelmingly large, asking for principled design flows that address the application-level needs. This paper presents a latency-driven design methodology for mapping ConvNets on FPGAs. The proposed design flow employs novel transformations over a Synchronous Dataflow-based modelling framework together with a latency-centric optimisation procedure in order to efficiently explore the design space targeting low-latency designs. Quantitative evaluation shows large improvements in latency when latency-driven optimisation is in place yielding designs that improve the latency of AlexNet by 73.54× and VGG16 by 5.61× over throughput-optimised designs. Stylianos I. Venieris, Christos-Savvas Bouganis |
FPL | 2 |
| 2017 | Particle MCMC algorithms and architectures for accelerating inference in state-space modelsabstractParticle Markov Chain Monte Carlo (pMCMC) is a stochastic algorithm designed to generate samples from a probability distribution, when the density of the distribution does not admit a closed form expression. pMCMC is most commonly used to sample from the Bayesian posterior distribution in State-Space Models (SSMs), a class of probabilistic models used in numerous scientific applications. Nevertheless, this task is prohibitive when dealing with complex SSMs with massive data, due to the high computational cost of pMCMC and its poor performance when the posterior exhibits multi-modality. This paper aims to address both issues by: 1) Proposing a novel pMCMC algorithm (denoted ppMCMC), which uses multiple Markov chains (instead of the one used by pMCMC) to improve sampling efficiency for multi-modal posteriors, 2) Introducing custom, parallel hardware architectures, which are tailored for pMCMC and ppMCMC. The architectures are implemented on Field Programmable Gate Arrays (FPGAs), a type of hardware accelerator with massive parallelization capabilities. The new algorithm and the two FPGA architectures are evaluated using a large-scale case study from genetics. Results indicate that ppMCMC achieves 1.96x higher sampling efficiency than pMCMC when using sequential CPU implementations. The FPGA architecture of pMCMC is 12.1x and 10.1x faster than state-of-the-art, parallel CPU and GPU implementations of pMCMC and up to 53x more energy efficient; the FPGA architecture of ppMCMC increases these speedups to 34.9x and 41.8x respectively and is 173x more power efficient, bringing previously intractable SSM-based data analyses within reach. Grigorios Mingas, Leonardo Bottolo, Christos-Savvas Bouganis |
Int. J. Approx. Reason. | 3 |
| 2017 | An Unbiased MCMC FPGA-Based Accelerator in the Land of Custom Precision ArithmeticabstractMarkov Chain Monte Carlo (MCMC) based methods have been the main tool used for Bayesian Inference by practitioners and researchers due to their flexibility and theoretical properties that guarantee unbiased sampling-based estimates. Nevertheless, with the availability of large data sets and the constant need to develop more complex models that better capture the targeted problem, significant computational challenges have been presented. Current approaches, based on multi-core CPUs, GPUs, and FPGAs, aim to accelerate the execution time of the MCMC methods using subsampling techniques or custom precision arithmetic, resulting to biased estimates. In this work, a novel FPGA-based construction is proposed that utilises the custom precision support of FPGA devices in order to accelerate the computations, guaranteeing at the same time asymptotically unbiased estimates. Key to this approach is the extension of the parameter space by an extra parameter that indicates the required precision in the computation of the likelihood of a data point. The work proposes an FPGA architecture for the above algorithm, as well as discuss its tuning for maximising the performance of the system. The performance of the FPGA-mapped sampler is evaluated using two Bayesian logistic regression case studies of varying complexity, which show significant speedups compared to existing FPGAand CPU-based works that utilise double floating point arithmetic, without any bias on the sampling-based estimates. Shuanglong Liu, Grigorios Mingas, Christos-Savvas Bouganis |
IEEE Trans. Computers | 3 |
| 2016 | fpgaConvNet: A Framework for Mapping Convolutional Neural Networks on FPGAsabstractConvolutional Neural Networks (ConvNets) are a powerful Deep Learning model, providing state-of-the-art accuracy to many emerging classification problems. However, ConvNet classification is a computationally heavy task, suffering from rapid complexity scaling. This paper presents fpgaConvNet, a novel domain-specific modelling framework together with an automated design methodology for the mapping of ConvNets onto reconfigurable FPGA-based platforms. By interpreting ConvNet classification as a streaming application, the proposed framework employs the Synchronous Dataflow (SDF) model of computation as its basis and proposes a set of transformations on the SDF graph that explore the performance-resource design space, while taking into account platform-specific resource constraints. A comparison with existing ConvNet FPGA works shows that the proposed fully-automated methodology yields hardware designs that improve the performance density by up to 1.62× and reach up to 90.75% of the raw performance of architectures that are hand-tuned for particular ConvNets. Stylianos I. Venieris, Christos-Savvas Bouganis |
FCCM | 2 |
| 2016 | Semi-dense SLAM on an FPGA SoCabstractDeploying advanced Simultaneous Localisation and Mapping, or SLAM, algorithms in autonomous low-power robotics will enable emerging new applications which require an accurate and information rich reconstruction of the environment. This has not been achieved so far because accuracy and dense 3D reconstruction come with a high computational complexity. This paper discusses custom hardware design on a novel platform for embedded SLAM, an FPGA-SoC, combining an embedded CPU and programmable logic on the same chip. The use of programmable logic, tightly integrated with an efficient multicore embedded CPU stands to provide an effective solution to this problem. In this work an average framerate of more than 4 frames/second for a resolution of 320×240 has been achieved with an estimated power of less than 1 Watt for the custom hardware. In comparison to the software-only version, running on a dual-core ARM processor, an acceleration of 2× has been achieved for LSD-SLAM, without any compromise in the quality of the result. Konstantinos Boikos, Christos-Savvas Bouganis |
FPL | 2 |
| 2016 | Population-Based MCMC on Multi-Core CPUs, GPUs and FPGAsabstractMarkov Chain Monte Carlo (MCMC) is a method to draw samples from a given probability distribution. Its frequent use for solving probabilistic inference problems, where big-scale data are repeatedly processed, means that MCMC runtimes can be unacceptably large. This paper focuses on population-based MCMC, a popular family of computationally intensive MCMC samplers; we propose novel, highly optimized accelerators in three parallel hardware platforms (multi-core CPUs, GPUs and FPGAs), in order to address the performance limitations of sequential software implementations. For each platform, we jointly exploit the nature of the underlying hardware and the special characteristics of population-based MCMC. We focus particularly on the use of custom arithmetic precision, introducing two novel methods which employ custom precision in the largest part of the algorithm in order to reduce runtime, without causing sampling errors. We apply these methods to all platforms. The FPGA accelerators are up to 114x faster than multi-core CPUs and up to 53x faster than GPUs when doing inference on mixture models. Grigorios Mingas, Christos-Savvas Bouganis |
IEEE Trans. Computers | 2 |
| 2016 | Embedded Hardware-Efficient Real-Time Classification With Cascade Support Vector MachinesabstractCascade support vector machines (SVMs) are optimized to efficiently handle problems, where the majority of the data belong to one of the two classes, such as image object classification, and hence can provide speedups over monolithic (single) SVM classifiers. However, SVM classification is a computationally demanding task and existing hardware architectures for SVMs only consider monolithic classifiers. This paper proposes the acceleration of cascade SVMs through a hybrid processing hardware architecture optimized for the cascade SVM classification flow, accompanied by a method to reduce the required hardware resources for its implementation, and a method to improve the classification speed utilizing cascade information to further discard data samples. The proposed SVM cascade architecture is implemented on a Spartan-6 field-programmable gate array (FPGA) platform and evaluated for object detection on 800×600 (Super Video Graphics Array) resolution images. The proposed architecture, boosted by a neural network that processes cascade information, achieves a real-time processing rate of 40 frames/s for the benchmark face detection application. Furthermore, the hardware-reduction method results in the utilization of 25% less FPGA custom-logic resources and 20% peak power reduction compared with a baseline implementation. Christos Kyrkou, Christos-Savvas Bouganis, Theocharis Theocharides, Marios M. Polycarpou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Robust multi-image based blind face hallucinationabstractThis paper proposes a robust multi-image based blind face hallucination framework to super-resolve LR faces. The proposed framework first estimates both blurring kernel and transformations of multiple LR faces by robust deblurring and registration in PCA subspace. A patch-wise mixture of probabilistic PCA prior is then incorporated for face super-resolution. Previous work on face SR using PCA prior can be viewed as special cases of the framework. Experimental results in both simulated and real LR sequences demonstrate very promising performance of the proposed method. Yonggang Jin, Christos-Savvas Bouganis |
CVPR | 2 |
| 2015 | FPGA based nonlinear Support Vector Machine training using an ensemble learningabstractSupport Vector Machines (SVMs) are powerful supervised learning methods in machine learning. However, their applicability to large problems has been limited due to the time consuming training stage whose computational cost scales quadratically with the number of examples. In this work, a complete FPGA-based system for nonlinear SVM training using ensemble learning is presented. The proposed framework builds on the FPGA architecture and utilizes a cascaded multi-precision training flow, exploits the heterogeneity within the training problem by tuning the number representation used, and supports ensemble training tuned to each internal memory structure so to address very large datasets. Its performance evaluation shows that the proposed system achieves more than an order of magnitude better results compared to state-of-the-art CPU and GPU-based implementations. Mudhar Bin Rabieah, Christos-Savvas Bouganis |
FPL | 2 |
| 2015 | Towards heterogeneous solvers for large-scale linear systemsabstractApplying Linear Regression to systems with a massive amount of observations, a scenario which is becoming increasingly common in the era of Big Data, poses major algorithmic and computational challenges. This paper proposes a novel high-performance FPGA-based architecture for large-scale Linear Regression problems as well as a heterogeneous system comprising the custom FPGA architecture, an enhanced GPU module and a multi-core CPU for addressing the aforementioned problem. The system adaptively assigns Linear Regression workloads to the three computing devices to minimise runtime. The device with the highest performance is chosen based on an analytical framework, as well as the workload's size and structure. A quantitative comparison with existing FPGA, GPU and multi-core CPU designs yields speed-ups of up to 18.07×, 32.67× and 25.84× respectively. Stylianos I. Venieris, Grigorios Mingas, Christos-Savvas Bouganis |
FPL | 3 |
| 2015 | An exact MCMC accelerator under custom precision regimesabstractMarkov chain Monte Carlo (MCMC) is one of the most popular and important tools to generate random samples from probability distributions over many variables which occur frequently in Bayesian inference. However, MCMC cannot be practically applied to models with large data sets because of the prohibitive costly likelihood evaluations for each data point. Previous solutions propose to compute the likelihood approximately, e.g. by sub-sampling data or by using custom precision implementations. These methods involve a trade-off between bias in the output and sampling speed; therefore they cannot guarantee unbiased sampling, which is critical in many applications. This paper introduces a novel mixed precision MCMC accelerator for FPGAs, which simulates from the exact probability distribution in contrast to existing approximate MCMC samplers. An auxiliary binary variable is appended to each data point to indicate the corresponding likelihood term evaluation in full or reduced precision. The proposed method guarantees unbiased samples, while the large majority of likelihood computations are performed in reduced precision. Moreover, a tailored FPGA architecture for the algorithm is introduced, and its performance is evaluated using two Bayesian logistic regression case studies of varying complexity: a 2-dimension synthetic problem and MNIST classification with 12-dimension parameters. The achieved speedups over double-precision FPGA designs are 4.21× and 4.76× respectively. Shuanglong Liu, Grigorios Mingas, Christos-Savvas Bouganis |
FPT | 3 |
| 2015 | ARC 2014 Over-Clocking KLT Designs on FPGAs under Process, Voltage, and Temperature VariationabstractKarhunen-Loeve Transformation is a widely used algorithm in signal processing that often implemented with high-throughput requisites. This work presents a novel methodology to optimise KLT designs on FPGAs that outperform typical design methodologies, through a prior characterisation of the arithmetic units in the datapath of the circuit under various operating conditions. Limited by the ever-increasing process variation, the delay models available in synthesis tools are no longer suitable for extreme performance optimisation of designs, and as they are generic, they need to consider the worst-case performance for a given fabrication process. Hence, they heavily penalise the maximum possible achieved performance of a design by leaving safety margin. This work presents a novel unified optimisation framework which contemplates a prior characterisation of the embedded multipliers on the target FPGA device under process, voltage, and temperature variation. The proposed framework allows a design space exploration leading to designs without any latency overheads that achieve high throughput while producing less errors than typical methodologies, operating with the same throughput. Experimental results demonstrate that the proposed methodology outperforms the typical implementation in three real-life design strategies: high performance, low power, and temperature variation; and it produced circuit designs that performed up to 18dB better when over-clocked. Rui Policarpo Duarte, Christos-Savvas Bouganis |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | ARC 2014: A Multidimensional FPGA-Based Parallel DBSCAN ArchitectureabstractClustering large numbers of data points is a very computationally demanding task that often needs to be accelerated in order to be useful in practical applications. This work focuses on the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm, which is one of the state-of-the-art clustering algorithms, and targets its acceleration using an FPGA device. The article presents an optimized, scalable, and parameterizable architecture that takes advantage of the internal memory structure of modern FPGAs in order to deliver a high-performance clustering system. Post-synthesis simulation results show that the developed system can obtain mean speedups of 31× in real-world tests and 202× in synthetic tests when compared to state-of-the-art software counterparts running on a quad-core 3.4GHz Intel i7-2600k. Additionally, this implementation is also capable of clustering data with any number of dimensions without impacting the performance. Neil Scicluna, Christos-Savvas Bouganis |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2014 | Image progressive acquisition for hardware systemsabstractAs the resolution of digital images increases, accessing raw image data from memory has become a major consideration during the design of image/video processing systems. This is due to the fact that the bandwidth requirement and energy consumption of such image accessing process has increased. Inspired by the successful application of progressive image sampling techniques in many image processing tasks, this work proposes to apply similar concept within hardware systems to efficiently trade image quality for reduced memory bandwidth requirement and lower energy consumption. Based on this idea, a hardware system is proposed that is placed between the memory subsystem and the processing core of the design. The proposed system alters the conventional memory access pattern to progressively and adaptively access pixels from a target memory external to the system. The sampled pixels are used to reconstruct an approximation to the ground truth, which is stored in an internal image buffer for further processing. The system is prototyped on FPGA and its performance evaluation shows that a saving of up to 85% of memory accessing time and 33%/45% of image acquisition time/energy is achieved on the benchmark image “lena” while maintaining a PSNR of about 30 dB. Jianxiong Liu, Christos-Savvas Bouganis, Peter Y. K. Cheung |
DATE | 2 |
| 2014 | Pushing the performance boundary of linear projection designs through device specific optimisations (abstract only)abstractThe continuous scaling of the fabrication process combined with the ever increasing need of high performance designs, means that the era of treating all devices the same is about to come to an end. The presented work considers device oriented optimisations in order to further boost the performance of a Linear Projection design by focusing on the over-clocking of arithmetic operators. A methodology is proposed for the acceleration of Linear Projection designs on an FPGA, that introduces information about the performance of the hardware under over-clocking conditions to the application level. The novelty of this method is a pre-characterisation of the most prone to error arithmetic operators and the utilisation of this information in the high-level optimization process of the design. This results in a set of circuit designs that achieve higher throughput with minimum error. FPGA devices are suitable for such optimisations due to their reconfigurability feature that allows performance characterisation of the underlying fabric prior to the design of the final system. The reported results show that significant gains in the performance of the system can be achieved, i.e. up to 1.85 times speed up in the throughput compared to existing methodologies, when such device specific optimisation is considered. Rui Policarpo Duarte, Christos-Savvas Bouganis |
FPGA | 2 |
| 2014 | Parallel resampling for particle filters on FPGAsabstractParticle filters (PFs) are a set of algorithms that implement recursive Bayesian filtering, which represent the posterior distribution by a set of weighted samples. Resampling is a fundamental operation in PF algorithms. It consists of taking a population of samples and reconstructing it based on the weights attached to each sample, favouring the samples with large weights. However, resampling is computationally intensive when the number of samples is large and, most importantly, it is not inherently parallelizable like the other steps of the particle filter. Parallel computing devices such as Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs) have been proposed to accelerate resampling. In this paper, we propose novel parallel architectures that map four state-of-the-art resampling algorithms (systematic, residual systematic, Metropolis and Rejection resampling) to a FPGA. FPGA-specific optimisations are introduced to further optimize the performance of the above systems. The proposed architectures are implemented in a Virtex-6 LX240T FPGA device with half-utilization of logic resources. Compared to the respective state-of-the-art implementations on an NVIDIA K20 GPU, the achieved speedups are in the range of 1.7x-49x. Shuanglong Liu, Grigorios Mingas, Christos-Savvas Bouganis |
FPT | 3 |
| 2014 | Vision-Based Egomotion Estimation on FPGA for Unmanned Aerial Vehicle NavigationabstractThe use of unmanned aerial vehicles (UAVs) in commercial and warfare activities has intensified over the last decade. One of the main challenges is to enable UAVs to become as autonomous as possible. A vital component toward this direction is the robust and accurate estimation of the egomotion of the UAV. Egomotion estimation can be enhanced by equipping the UAV with a video camera, which enables a vision-based egomotion estimation. However, the high computational requirements of vision-based egomotion algorithms, combined with the real-time performance and low power consumption requirements that are related to such an application, cannot be met by general-purpose processing units. This paper presents a system architecture that employs a field-programmable gate array as the main processing platform connected to a low-power CPU that targets the problem of vision-based egomotion estimation in a UAV. The performance evaluation of the proposed system, using real data captured by a UAV's on-board camera, demonstrates the ability of the system to render accurate estimation of the egomotion parameters, meeting at the same time the real-time requirements imposed by the application. Maria E. Angelopoulou, Christos-Savvas Bouganis |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | On Optimizing the Arithmetic Precision of MCMC AlgorithmsabstractMarkov Chain Monte Carlo (MCMC) is an ubiquitous stochastic method, used to draw random samples from arbitrary probability distributions, such as the ones encountered in Bayesian inference. MCMC often requires forbiddingly long runtimes to give a representative sample in problems with high dimensions and large-scale data. Field-Programmable Gate Arrays (FPGAs) have proven to be a suitable platform for MCMC acceleration due to their ability to support massive parallelism. This paper introduces an automated method, which minimizes the floating point precision of the most computationally intensive part of an FPGA-mapped MCMC sampler, while keeping the precision-related bias in the output within a user-specified tolerance. The method is based on an efficient bias estimator, proposed here, which is able to estimate the bias in the output with only few random samples. The optimization process involves FPGA pre-runs, which estimate the bias and choose the optimized precision. This precision is then used to reconfigure the FPGA for the final, long MCMC run, allowing for higher sampling throughputs. The process requires no user intervention. The method is tested on two Bayesian inference case studies: Mixture models and neural network regression. The achieved speedups over double-precision FPGA designs were 3.5x-5x (including the optimization overhead). Comparisons with a sequential CPU and a GPGPU showed speedups of 223x-446x and 16x-18x respectively. Grigorios Mingas, Farhan Rahman, Christos-Savvas Bouganis |
FCCM | 3 |
| 2013 | FPGA-based acceleration of cascaded support vector machines for embedded applications (abstract only)abstractSupport Vector Machines (SVMs) are considered one of the most popular classification algorithms yielding high accuracy rates. However, SVMs often require processing a large number of support vectors, making the classification process computationally demanding, and hence it is challenging to meet real-time processing constraints imposed by many embedded applications. In order to improve SVM classification times the cascade classification scheme has been proposed. However, even in this case real-time performance is still challenging to achieve without exploiting the throughput and processing requirements of each cascade stage. Hence the design of an FPGA-based accelerator for cascaded SVM processing is proposed; in addition to a hardware reduction method in order to reduce the implementation requirements of the cascade SVM leading to significant resource savings. The accelerator was implemented on a Virtex 5 FPGA platform and evaluated using face detection as the target application on 640×480 resolution images. It was compared against FPGA implementations of the same cascade processing architecture but without using the reduction method, and a single parallel SVM classifier. The accelerator is capable an average performance of 70 frames-per-second, achieving a speed-up of 5× over the single parallel SVM classifier. Furthermore, the hardware reduction method results in the utilization of 43% less FPGA LUT resources, with only 0.7% reduction in classification accuracy. Christos Kyrkou, Christos-Savvas Bouganis, Theocharis Theocharides |
FPGA | 2 |
| 2013 | Accelerating Random Forest training process using FPGAabstractRandom Forest (RF) is one of the state-of-art supervised learning methods in Machine Learning and inherently consists of two steps: the training and the evaluation step. In applications where the system needs to be updated periodically, the training step becomes the bottleneck of the system, imposing hard constraints on its adaptability to a changing environment. In this work, a novel FPGA architecture for accelerating the RF training step is presented, exploring key features of the device. By combing a fine-grain data-flow processing at low-level and by exploiting parallelism available at high level inherent in the algorithm, significant acceleration factors are achieved. Key to the above gains is a novel FPGA FIFO based merge sorter module, a core component in the architecture, that exhibits high efficiency in memory utilisation; as well as a batch training strategy that enable full exploitation of the high memory bandwidth offered by the on-chip memory featured on FPGA devices. The proposed system achieves speed-up factors of up to 230x over a 3GHz Intel Core i5 processor when an Altera Stratix IV device is utilised under classification problems drawn from VOC2007. Chuan Cheng, Christos-Savvas Bouganis |
FPL | 2 |
| 2013 | A hardware-efficient architecture for embedded real-time cascaded support vector machines classificationabstractThis work presents an optimized architecture for cascaded SVM processing, along with a hardware reduction method for the implementation of the additional stages in the cascade, leading to significant improvements. The architecture was implemented on a Virtex 5 FPGA platform and evaluated using face detection as the target application on 640×480 resolution images. Additionally, it was compared against implementations of the same cascade processing architecture but without using the reduction method, and a single parallel SVM classifier. The proposed architecture achieves an average performance of 70 frames-per-second, demonstrating a speed-up of 5× over the single parallel SVM classifier. Furthermore, the hardware reduction method results in the utilization of 43% less hardware resources, with only 0.7% reduction in classification accuracy. Christos Kyrkou, Theocharis Theocharides, Christos-Savvas Bouganis |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | Face hallucination revisited: A joint frameworkabstractThe paper presents a joint framework for face hallucination incorporating face deblurring and registration. The joint framework not only directly hallucinates low resolution faces, but also deblurs and aligns low resolution faces iteratively to improve the performance of face hallucination. Without the need for accurate face registration and prior knowledge of blurring kernels, it is robust to errors in face registration and blurring kernel. Experimental results demonstrate the robust performance of the proposed method. Yonggang Jin, Christos-Savvas Bouganis |
ICIP | 2 |
| 2013 | High-level power and performance estimation of FPGA-based soft processors and its application to design space exploration
Adam Powell, Christos-Savvas Bouganis, Peter Y. K. Cheung |
J. Syst. Archit. | 2 |
| 2013 | Guest editorial: Workshop on Reconfigurable Computing
Ioannis Sourdis, Christos-Savvas Bouganis, Miquel Pericàs |
J. Syst. Archit. | 2 |
| 2012 | The DeSyRe Project: On-Demand System ReliabilityabstractThe DeSyRe project builds on-demand adaptive and reliable Systems-on-Chips (SoCs). As fabrication technology scales down, chips are becoming less reliable, thereby incurring increased power and performance costs for fault tolerance. To make matters worse, power density is becoming a significant limiting factor in SoC design, in general. In the face of such changes in the technological landscape, current solutions for fault tolerance are expected to introduce excessive overheads in future systems. Moreover, attempting to design and manufacture a totally defect-/fault-free system, would impact heavily, even prohibitively, the design, manufacturing, and testing costs, as well as the system performance and power consumption. In this context, DeSyRe will deliver a new generation of systems that are reliable by design at well-balanced power, performance, and design costs. Ioannis Sourdis, Christos Strydis, Christos-Savvas Bouganis, Babak Falsafi, Georgi Gaydadjiev, Alirad Malek, R. Mariani, Dionisios N. Pnevmatikatos, Dhiraj K. Pradhan, Gerard K. Rauwerda, Kim Sunesen, Stavros Tzilis |
DSD | 3 |
| 2012 | A Custom Precision Based Architecture for Accelerating Parallel Tempering MCMC on FPGAs without Introducing Sampling ErrorabstractMarkov Chain Monte Carlo (MCMC) is a method used to draw samples from probability distributions in order to estimate - otherwise intractable - integrals. When the distribution is complex, simple MCMC becomes inefficient and advanced, computationally intensive MCMC methods are employed to make sampling possible. This work proposes a novel streaming FPGA architecture to accelerate Parallel Tempering, a widely adopted MCMC method designed to sample from multimodal distributions. The proposed architecture demonstrates how custom precision can be intelligently employed without introducing sampling errors, in order to save resources and increase the sampling throughg put. Speedups of up to two orders of magnitude compared to software and 1.53x-76.88x compared to a GPGPU implementation are achieved when performing Bayesian inference for a mixture model. Grigorios Mingas, Christos-Savvas Bouganis |
FCCM | 2 |
| 2012 | High-level linear projection circuit design optimization framework for FPGAs under over-clockingabstractFrequently, the high-level algorithm parameter selection and its mapping into hardware are considered to be independent processes, often leading to suboptimal solutions. When DSP applications with real-time constraints are targeted, it is often desirable the resulting hardware system to be clocked at as high frequency as possible. Even though the trend in modern devices is to provide a fabric that can support higher frequencies, its variability makes the design tools to be pessimistic about maximum clock frequency estimates. The proposed framework optimizes and mitigates the probabilistic behaviour of digital circuits, by trying to expose the impact of variability of the fabric into high-level algorithmic specifications. FPGAs are well positioned to tackle this problem because they can be reconfigured, allowing an off-line characterization of the specific device before implementing the complete optimized circuit on the same device. Circuits generated by the proposed framework outperform typical implementations, by minimizing area, errors, and maximizing its operating clock frequency. An example of a linear projection circuit, over-clocked by 232%, shows savings up to 39% in hardware resources for the same target PSNR over traditional implementation. Rui Policarpo Duarte, Christos-Savvas Bouganis |
FPL | 2 |
| 2012 | Early performance estimation of image compression methods on soft processorsabstractThis paper presents a power and execution time estimation framework for an FPGA-based soft processor when considering the implementation of image compression techniques. Using the proposed framework, a quick power consumption and execution time estimate can be obtained early in the design phase allowing system designers to estimate these performance metrics without the need of implementing the algorithm or generating all possible soft processor architectures. This estimate is performed using both high-level algorithm parameters and soft processor architecture parameters. For system designers this can result in fast design space exploration. The model can predict the execution time of an algorithm with an average of 139% less relative error than predictions using only architecture parameters with the same framework. Adam Powell, Christos-Savvas Bouganis, Peter Y. K. Cheung |
FPL | 2 |
| 2012 | Novel Cascade FPGA Accelerator for Support Vector Machines ClassificationabstractSupport vector machines (SVMs) are a powerful machine learning tool, providing state-of-the-art accuracy to many classification problems. However, SVM classification is a computationally complex task, suffering from linear dependencies on the number of the support vectors and the problem's dimensionality. This paper presents a fully scalable field programmable gate array (FPGA) architecture for the acceleration of SVM classification, which exploits the device heterogeneity and the dynamic range diversities among the dataset attributes. An adaptive and fully-customized processing unit is proposed, which utilizes the available heterogeneous resources of a modern FPGA device in efficient way with respect to the problem's characteristics. The implementation results demonstrate the efficiency of the heterogeneous architecture, presenting a speed-up factor of 2-3 orders of magnitude, compared to the CPU implementation. The proposed architecture outperforms other proposed FPGA and graphic processor unit approaches by more than seven times. Furthermore, based on the special properties of the heterogeneous architecture, this paper introduces the first FPGA-oriented cascade SVM classifier scheme, which exploits the FPGA reconfigurability and intensifies the custom-arithmetic properties of the heterogeneous architecture. The results show that the proposed cascade scheme is able to increase the heterogeneous classifier throughput even further, without introducing any penalty on the resource utilization. Markos Papadonikolakis, Christos-Savvas Bouganis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2011 | A Run-Time Adaptive FPGA Architecture for Monte Carlo SimulationsabstractField Programmable Gate Arrays (FPGAs) are now considered to be one of the preferred computing platforms for high performance computing applications, such as Monte Carlo simulations, due to their large computational power and low power consumption. Unlike other state-of-the-art computing platforms, such as General Purpose Processors (GPPs) and General Purpose Graphics Processing Units (GPGPU), FPGAs can moreover exploit the applications' requirements with respect to the employed number representation scheme, with the potential to lead to considerable area savings and throughput increases. This work proposes a novel FPGA based architecture for Monte Carlo simulations that monitors and configures the number representation of the system during run-time in order to accommodate the dynamics of the system under investigation, resulting to a considerable boost on the overall performance of the system compared to a conventional system. In order to evaluate the efficacy of the proposed architecture, the GARCH model from the financial industry is considered as a case study. The results demonstrate that an average of ~1.35× throughput per resource unit improvement is achieved compared to conventional parallel arithmetic implementation. Christos-Savvas Bouganis |
FPL | 2 |
| 2011 | An FPGA-based object detector with dynamic workload balancingabstractIn recent years, object detection has been more frequently integrated with other vision processing functions, acting for acquisition of region of interest and is widely adopted in portable devices such as digital camera capable for automatic focusing on faces. In applications targeting those devices, limitations in both hardware resources and power supply mean an efficient utilization of hardware resource is of significance. In this paper a novel hardware architecture for Viola and Jones object detectior is proposed. The novel feature of the architecture is that it features a mechanism of dynamic workload balancing, which adaptively re-distributes the workload among available processing units, thus achieving highly efficient utilization of hardware resource. The obtained results demonstrate that the proposed system can achieve high utilisation of the dedicated resources leading to high performance over resource ratio. Chuan Cheng, Christos-Savvas Bouganis |
FPT | 2 |
| 2011 | Feature selection with geometric constraints for vision-based Unmanned Aerial Vehicle navigationabstractVision-based egomotion estimation can be employed to endow with navigation ability an Unmanned Aerial Vehicle (UAV) that is equipped with an on-board camera. The egomotion estimation block computes the 3D UAV motion, taking as an input a 2D optical flow map that is constructed for each of the captured video frames. This work considers sparse optical flow estimation, and thus the navigation system that is developed includes a feature selection unit, which initially identifies the points of the optical flow map. This paper demonstrates that the feature selection process, and in particular the geometry of the selected feature set, decisively determines the overall system performance. Various computation schedules, which combine geometric constraints with a textural quality metric for the image features, are thus investigated. This paper shows that imposing appropriate distance constraints in the feature selection process significantly increases the output precision of the egomotion estimation unit, thus enabling accurate vision-based UAV self-navigation. Maria E. Angelopoulou, Christos-Savvas Bouganis |
ICIP | 2 |
| 2010 | A Heterogeneous FPGA Architecture for Support Vector Machine TrainingabstractSupport Vector Machines is a powerful supervised learning tool. Its training phase, however, is a time-consuming task and heavily dependent on the training dataset size and dimensionality. In this work, we propose a scalable FPGA architecture for the acceleration of SVM training, which exploits the heterogeneous nature of the device and the diversities of the precision requirements among the dataset attributes. The maximum parallelization potential is obtained by maintaining the usage of DSPs and logic resources at the initial ratio of the FPGA device. The results demonstrate the efficiency of the heterogeneous architecture in both homogeneous and heterogeneous datasets. The proposed architecture outperforms other proposed designs by more than 6 times, in terms of raw computational speed. Markos Papadonikolakis, Christos-Savvas Bouganis |
FCCM | 2 |
| 2010 | GPU Versus FPGA for High Productivity ComputingabstractHeterogeneous or co-processor architectures are becoming an important component of high productivity computing systems (HPCS). In this work the performance of a GPU based HPCS is compared with the performance of a commercially available FPGA based HPC. Contrary to previous approaches that focussed on specific examples, a broader analysis is performed by considering processes at an architectural level. A set of benchmarks is employed that use different process architectures in order to exploit the benefits of each technology. These include the asynchronous pipelines common to map tasks, a partially synchronous tree common to reduce tasks and a fully synchronous, fully connected mesh. We show that the GPU is more productive than the FPGA architecture for most of the benchmarks and conclude that FPGA-based HPCS is being marginalised by GPUs. David Huw Jones, Adam Powell, Christos-Savvas Bouganis, Peter Y. K. Cheung |
FPL | 3 |
| 2010 | Mapping Multiple Multivariate Gaussian Random Number Generators on an FPGAabstractA Multivariate Gaussian random number generator (MVGRNG) is an essential block for many hardware designs, including Monte Carlo simulations. These simulations are usually used in applications such as statistical physics and financial mathematics. Field Programmable Gate Arrays (FPGAs) are often used to implement these generators as the design can be effectively optimized. Many applications require random samples from a number of multivariate Gaussian distributions leading to a problem of efficiently mapping of the required MVGRNG on an FPGA. The proposed approach presented in this paper exploits any redundancy that exists between different distributions under consideration leading to designs with improved resource usage. Experimental results demonstrate that the proposed approach outperforms the existing approaches by producing MVGRNG designs that utilize less hardware resources in comparison to existing approaches achieving up to 50\% reduction of hardware resource utilization. Chalermpol Saiprasert, Christos-Savvas Bouganis, George A. Constantinides |
FPL | 2 |
| 2010 | A novel FPGA-based SVM classifierabstractSupport Vector Machines (SVMs) are a powerful supervised learning tool, providing state-of-the-art accuracy at a cost of high computational complexity. The SVM classification suffers from linear dependencies on the number of the Support Vectors and the problem's dimensionality. In this work, we propose a scalable FPGA architecture for the acceleration of SVM classification, which exploits the device heterogeneity and the dynamic range diversities among the dataset attributes. Furthermore, this work introduces the first FPGA-oriented cascade SVM classifier scheme, which intensifies the custom-arithmetic properties of the heterogeneous architecture and boosts the classification performance even more. The implementation results demonstrate the efficiency of the heterogeneous architecture, presenting a speed-up factor of 2-3 orders of magnitude, compared to the CPU implementation, while outperforming other proposed FPGA and GPU approaches by more than 7 times. Markos Papadonikolakis, Christos-Savvas Bouganis |
FPT | 2 |
| 2010 | A Salient Region Detector for GPU Using a Cellular Automata Architecture
David Huw Jones, Adam Powell, Christos-Savvas Bouganis, Peter Y. K. Cheung |
ICONIP (2) | 3 |
| 2010 | An Optimized Hardware Architecture of a Multivariate Gaussian Random Number GeneratorabstractMonte Carlo simulation is one of the most widely used techniques for computationally intensive simulations in mathematical analysis and modeling. A multivariate Gaussian random number generator is one of the main building blocks of such a system. Field Programmable Gate Arrays (FPGAs) are gaining increased popularity as an alternative means to the traditional general purpose processors targeting the acceleration of the computationally expensive random number generator block. This article presents a novel approach for mapping a multivariate Gaussian random number generator onto an FPGA by optimizing the computational path in terms of hardware resource usage subject to an acceptable error in the approximation of the distribution of interest. The proposed approach is based on the eigenvalue decomposition algorithm which leads to a design with different precision requirements in the computational paths. An analysis on the impact of the error due to truncation/rounding operation along the computational path is performed and an analytical expression of the error inserted into the system is presented. Based on the error analysis, three algorithms that optimize the resource utilization and at the same time minimize the error in the output of the system are presented and compared. Experimental results reveal that the hardware resource usage on an FPGA as well as the error in the approximation of the distribution of interest are significantly reduced by the use of the optimization techniques introduced in the proposed approach. Chalermpol Saiprasert, Christos-Savvas Bouganis, George A. Constantinides |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2010 | Exploration of Heterogeneous FPGAs for Mapping Linear Projection DesignsabstractIn many applications, a reduction of the amount of the original data or a representation of the original data by a small set of variables is often required. Among many techniques, the linear projection is often chosen due to its computational attractiveness and good performance. For applications where real-time performance and flexibility to accommodate new data are required, the linear projection is implemented in field-programmable gate arrays (FPGAs) due to their fine-grain parallelism and reconfigurability properties. Currently, the optimization of such a design is considered as a separate problem from the basis calculation leading to suboptimal solutions. In this paper, we propose a novel approach that couples the calculation of the linear projection basis, the area optimization problem, and the heterogeneity exploration of modern FPGAs. The power of the proposed framework is based on the flexibility to insert information regarding the implementation requirements of the linear basis by assigning a proper prior distribution to the basis matrix. Results from real-life examples on modern FPGA devices demonstrate the effectiveness of our approach, where up to 48% reduction in the required area is achieved compared to the current approach, without any loss in the accuracy or throughput of the design. Christos-Savvas Bouganis, Iosifina Pournara, Peter Y. K. Cheung |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Implementation of a foveal vision mappingabstractFoveal vision reduces the data volume by utilising a spatially variant resolution. A high resolution is maintained in the fovea where it can be used by computer vision algorithms, while the resolution is reduced in the periphery where it is less important. This work focuses on the hardware architecture of a system that maps a conventional high resolution uniformly sampled sensor to a variable resolution output. The key feature of the proposed architecture is that it employs a separable forward mapping which requires a small amount of FPGA resources. This enables an efficient implementation of a continuously variable spatial resolution, requiring only 1000 LUTs on a Virtex-5, and runs at 104 MHz, enabling the mapping of a 512 × 512 window at over 300 frames per second. Donald G. Bailey, Christos-Savvas Bouganis |
FPT | 2 |
| 2009 | A sensor-based approach to linear blur identification for real-time video enhancementabstractSuper-resolution (SR) methods are largely affected by the accurate evaluation of the Point Spread Function (PSF) that is related to the input frames. When the frames are degraded by heavy motion blur, the PSFs are highly non-isotropic, which further complicates their estimation. The ill-posed nature of blur identification is usually addressed using the assumption of linear and uniform motion. However, in real-life systems, this may deviate significantly from the actual motion blur. To resolve the above, this work proposes combining a scheme that validates the initial motion assumption with the real-time reconfiguration property of an adaptive image sensor. If the linearity and uniformity assumption is invalid for a given motion region, the sensor is locally reconfigured to larger pixels that produce higher frame-rate samples with reduced blur. Once the appropriate configuration that gives rise to a valid motion assumption is applied, highly accurate PSFs are estimated, resulting to an improved SR reconstruction quality. Maria E. Angelopoulou, Christos-Savvas Bouganis, Peter Y. K. Cheung |
ICIP | 2 |
| 2009 | Robust Real-Time Super-Resolution on FPGA and an Application to Video EnhancementabstractThe high density image sensors of state-of-the-art imaging systems provide outputs with high spatial resolution, but require long exposure times. This limits their applicability, due to the motion blur effect. Recent technological advances have lead to adaptive image sensors that can combine several pixels together in real time to form a larger pixel. Larger pixels require shorter exposure times and produce high-frame-rate samples with reduced motion blur. This work proposes combining an FPGA with an adaptive image sensor to produce an output of high resolution both in space and time. The FPGA is responsible for the spatial resolution enhancement of the high-frame-rate samples using super-resolution (SR) techniques in real time. To achieve it, this article proposes utilizing the Iterative Back Projection (IBP) SR algorithm. The original IBP method is modified to account for the presence of noise, leading to an algorithm more robust to noise. An FPGA implementation of this algorithm is presented. The proposed architecture can serve as a general purpose real-time resolution enhancement system, and its performance is evaluated under various noise levels. Maria E. Angelopoulou, Christos-Savvas Bouganis, Peter Y. K. Cheung, George A. Constantinides |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2009 | Synthesis and Optimization of 2D Filter Designs for Heterogeneous FPGAsabstractMany image processing applications require fast convolution of an image with one or more 2D filters. Field-Programmable Gate Arrays (FPGAs) are often used to achieve this goal due to their fine grain parallelism and reconfigurability. However, the heterogeneous nature of modern reconfigurable devices is not usually considered during design optimization. This article proposes an algorithm that explores the space of possible implementation architectures of 2D filters, targeting the minimization of the required area, by optimizing the usage of the different components in a heterogeneous device. This is achieved by exploring the heterogeneous nature of modern reconfigurable devices using a Singular Value Decomposition based algorithm, which provides an efficient mapping of filter's implementation requirements to the heterogeneous components of modern FPGAs. In the case of multiple 2D filters, the proposed algorithm also exploits any redundancy that exists within each filter and between different filters in the set, leading to designs with minimized area. Experiments with real filter sets from computer vision applications demonstrate an average of up to 38% reduction in the required area. Christos-Savvas Bouganis, Sung-Boem Park, George A. Constantinides, Peter Y. K. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2009 | Introduction to the Special Issue ARC'08abstractNo abstract available. Katherine Compton, Roger F. Woods, Christos-Savvas Bouganis, Pedro C. Diniz |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2008 | Real-time image super resolution using an FPGAabstractImage super resolution is the process of combining a set of overlapping low-resolution images to produce a single high-resolution image. In this paper, a novel real-time super-resolution system is presented which is based on a weighted mean super-resolution algorithm combined with the existing fast and robust multi-frame super-resolution algorithm. The resource requirements of the proposed architecture scale linearly with the targeted image quality, making it ideally suited for a variety of real-time applications such as HDTV. Simulation results demonstrate a speed-up of three orders of magnitude over optimized software implementations with negligible loss to image quality. Oliver Bowen, Christos-Savvas Bouganis |
FPL | 2 |
| 2008 | Efficient FPGA mapping of Gilbert's algorithm for SVM training on large-scale classification problemsabstractSupport vector machines (SVMs) are an effective, adaptable and widely used method for supervised classification. However, training an SVM classifier on large-scale problems is proven to be a very time-consuming task for software implementations. This paper presents a scalable high-performance FPGA architecture of Gilbertpsilas Algorithm on SVM, which maximally utilizes the features of an FPGA device to accelerate the SVM training task for large-scale problems. Initial comparisons of the proposed architecture to the software approach of the algorithm show a speed-up factor range of three orders of magnitude for the SVM training time, regarding a wide range of datapsilas characteristics. Markos Papadonikolakis, Christos-Savvas Bouganis |
FPL | 2 |
| 2008 | A scalable FPGA architecture for non-linear SVM trainingabstractSupport vector machines (SVMs) is a popular supervised learning method, providing state-of-the-art accuracy in various classification tasks. However, SVM training is a time-consuming task for large-scale problems. This paper proposes a scalable FPGA architecture which targets a geometric approach to SVM training based on Gilbertpsilas algorithm using kernel functions. The architecture is partitioned into floating-point and fixed-point domains in order to efficiently exploit the FPGApsilas available resources for the acceleration of the non-linear SVM training. Implementation results present a speed-up factor up to three orders of magnitude of the most computational expensive part of the algorithm compared to the algorithmpsilas software implementation. Markos Papadonikolakis, Christos-Savvas Bouganis |
FPT | 2 |
| 2008 | Video enhancement on an adaptive image sensorabstractThe high density pixel sensors of the latest imaging systems provide images with high resolution, but require long exposure times, which limit their applicability due to the motion blur effect. Recent technological advances have lead to image sensors that can combine in real-time several pixels together to form a larger pixel. Larger pixels require shorter exposure times and produce high-frame-rate samples with reduced motion blur. This work proposes ways of configuring such a sensor to maximize the raw information collected from the environment, and methods to process that information and enhance the final output. In particular, a super-resolution and a deconvolution-based approach, for motion deblurring on an adaptive image sensor, are proposed, compared and evaluated. Maria E. Angelopoulou, Christos-Savvas Bouganis, Peter Y. K. Cheung |
ICIP | 2 |
| 2007 | Efficient Mapping of Dimensionality Reduction Designs onto Heterogeneous FPGAsabstractDimensionality reduction or feature extraction has been widely used in applications that require to reduce the amount of original data, like in image compression, or to represent the original data by a small set of variables that capture the main modes of data variation, as in face recognition and detection applications. A linear projection is often chosen due to its computational attractiveness. The calculation of the linear basis that best explains the data is usually addressed using the Karhunen-Loeve transform (KLT). Moreover, for applications where real-time performance and flexibility to accommodate new data are required, the linear projection is implemented in FPGAs due to their fine-grain parallelism and reconfigurability properties. Currently, the optimization of such a design, in terms of area usage and efficient allocation of the embedded multipliers that exist in modern FPGAs, is considered as a separate problem to the basis calculation. In this paper, we propose a novel approach that couples the calculation of the linear projection basis, the area optimization problem, and the heterogeneity exploration of modern FPGAs under a probabilistic Bayesian framework. The power of the proposed framework is based on the flexibility to insert information regarding the implementation requirements of the linear basis by assigning a proper prior distribution. Results using real-life examples demonstrate the effectiveness of our approach. Christos-Savvas Bouganis, Iosifina Pournara, Peter Y. K. Cheung |
FCCM | 1 |
| 2007 | Efficient mapping of a Kalman filter into an FPGA using Taylor ExpansionabstractThe Kalman filter is widely used as an estimator in many modern applications. In the case where its implementation in hardware is required, the computational complexity of the algorithm dictates the use of many resources. This paper presents an approximation of the conventional Kalman filter by using Taylor expansion and matrix calculus in order to remove the hardware expensive part of the algorithm. The Bierman-Thornton algorithm, as the exact counterpart of our proposed Approximate Kalman filter algorithm, is also implemented for comparison purposes. Comparing to the Bierman-Thornton algorithm, the FPGA implementation results demonstrate that our proposed Approximate Kalman filter implementation achieves one order of magnitude higher throughput using less hardware resources, obtaining similar convergence rate and accuracy. Christos-Savvas Bouganis, Peter Y. K. Cheung |
FPL | 2 |
| 2006 | Hardware efficient architectures for Eigenvalue computationabstractEigenvalue computation is essential in many fields of science and engineering. For high performance and real-time applications, this may need to be done in hardware. This paper focuses on the exploration of hardware architectures which compute eigenvalues of symmetric matrices. We propose to use the approximate Jacobi method for general case symmetric matrix eigenvalue problem. The paper illustrates that the proposed architecture is more efficient than previous architectures reported in the literature. Moreover, for the special case of 3times3 symmetric matrices, we propose to use an algebraic method. It is shown that the pipelined architecture based on the algebraic method has a significant advantage in terms of area Christos-Savvas Bouganis, Peter Y. K. Cheung, Philip H. W. Leong, Stephen J. Motley |
DATE | 2 |
| 2006 | FPGA-Accelerated Pre-Attentive Segmentation in Primary Visual CortexabstractVisual attention systems inspired by the behavior of neural architectures have attracted the attention of many researchers in the computer vision field. Of special interest is the model proposed by Li where the bottom-up saliency features of an image are detected through a mechanism that simulates the operation of the primary visual cortex (V1). Beyond its biological nature, the specific model is also of interest because it performs texture segmentation and contour enhancement using the same circuitry. The main drawback of the proposed model is its computational complexity, making it time consuming to simulate the model in software to, e.g., explore the model parameters, and also limits its applicability in real-time scenarios. In this work, we explore the inherent parallelism that exists in the model and propose a flexible hardware architecture that can accelerate the model. Moreover, the flexibility of the proposed architecture to adapt to similar models of the brain is of significant concern. Performance evaluation shows that the proposed architecture gives results close to the software model, achieving at the same time a speed up of one order of magnitude. Christos-Savvas Bouganis, Peter Y. K. Cheung, Zhaoping Li 0001 |
FPL | 1 |
| 2006 | Efficient Realtime FPGA Implementation of the Trace TransformabstractThe trace transform is a novel image transform that is able to exhibit useful properties such as scale and rotation invariance and occlusion robustness. As a result, it is particularly suited to a variety of classification and recognition tasks including image database search, token registration, activity monitoring, character recognition and face authentication. The main obstacle to the widespread use of the transform is its high computational complexity. This has precluded a detailed investigation of transform parameters. This paper presents an architecture and implementation of a trace transform engine on a Virtex-II FPGA. By exploiting the inherent parallelism in the algorithm and the use of optimised functional blocks, a huge performance gain is achieved, exceeding realtime video processing requirements for a 256 times 256 image Suhaib A. Fahmy, Christos-Savvas Bouganis, Peter Y. K. Cheung, Wayne Luk |
FPL | 2 |
| 2006 | An FPGA implementation of the simplex algorithmabstractLinear programming is applied to a large variety of scientific computing applications and industrial optimization problems. The Simplex algorithm is widely used for solving linear programs due to its robustness and scalability properties. However, application of the current software implementations of the Simplex algorithm to real-life optimization problems are time consuming when used as the bounding engine within an integer linear programming framework. This work aims to accelerate the Simplex algorithm by proposing a novel parameterizable hardware implementation of the algorithm on an FPGA. Evaluation of the proposed design using real problems demonstrates a speedup of up to 20 times over a highly optimized commercial software implementation running on a 3.4GHz Pentium 4 processor, which is itself 100 times faster than one of the main public domain solvers Samuel Bayliss, Christos-Savvas Bouganis, George A. Constantinides, Wayne Luk |
FPT | 2 |
| 2006 | A statistical framework for dimensionality reduction implementation in FPGAsabstractDimensionality reduction or feature extraction has been widely used in applications that require a set of data to be represented by a small set of variables. A linear projection is often chosen due to its computational attractiveness. The calculation of the linear basis that best explains the data is usually addressed using the Karhunen-Loeve transform (KLT). Moreover, for applications where real-time performance and flexibility to accommodate new data are required, the linear projection is implemented in FPGAs due to their fine-grain parallelism and reconfigurability properties. Currently, the optimization of such a design in terms of area usage is considered as a separate problem to the basis calculation. In this paper, we propose a novel approach that couples the calculation of the linear projection basis and the area optimization problems under a probabilistic Bayesian framework. The power of the proposed framework is based on the flexibility to insert information regarding the implementation requirements of the linear basis by assigning a proper prior distribution. Results using real-life examples demonstrate the effectiveness of our approach Christos-Savvas Bouganis, Iosifina Pournara, Peter Y. K. Cheung |
FPT | 1 |
| 2006 | A Spatiotemporal Saliency FrameworkabstractThis paper presents a novel bio-inspired spatiotemporal saliency framework. The framework incorporates spatial feature detection, feature tracking and motion prediction in order to generate a spatiotemporal saliency map. Experimental results demonstrate its ability and robustness to produce saliency responses to motion pop-up phenomena that are in line with humans responses. Moreover, the limited storage requirements permit real-time implementations of the proposed framework. Christos-Savvas Bouganis, Peter Y. K. Cheung |
ICIP | 2 |
| 2005 | A Novel 2D Filter Design Methodology for Heterogeneous DevicesabstractIn many image processing applications, fast convolution of an image with a large 2D filter is required. Field programable gate arrays (FPGAs) are often used to achieve this goal due to their fine grain parallelism and reconfigurability. However, the heterogeneous nature of modern reconfigurable devices is not usually considered during design optimization. This paper proposes an algorithm that explores the implementation architecture of 2D filters, targeting the minimization of the required area, by optimizing the usage of the different components in a heterogeneous device. Experiments show that the proposed algorithm can achieve a reduction in the required area in a range o to 70% when compared to current techniques. Christos-Savvas Bouganis, George A. Constantinides, Peter Y. K. Cheung |
FCCM | 1 |
| 2005 | Heterogeneity Exploration for Multiple 2D Filter DesignsabstractMany image processing applications require fast convolution of an image with a set of large 2D filters. Field-programmable gate arrays (FPGAs) are often used to achieve this goal due to their fine grain parallelism and reconfigurability. This paper presents a novel algorithm for the class of designs that implement a convolution with a set of 2D filters. Firstly, it explores the heterogeneous nature of modern reconfigurable devices using a singular value decomposition based algorithm, which orders the coefficients according to their impact to the filters' approximation. Secondly, it exploits any redundancy that exists within each filter and between different filters in the set, leading to designs with minimized area. Experiments with real filter sets from computer vision applications demonstrate up to 60% reduction in the required area. Christos-Savvas Bouganis, Peter Y. K. Cheung, George A. Constantinides |
FPL | 1 |
| 2005 | FPGA-Accelerated Reconstruction of Gene Regulatory NetworksabstractRapid advances in biological technologies, such as DNA microarrays, have enabled biologists to measure the expression levels of thousand of genes simultaneously under different conditions. This leads to a growing need to find methods that extract valuable information, fast and reliably, from this large amount of data. Recently, the advantages of using Bayesian networks for the reconstruction of gene regulatory networks from microarray data have been shown. However, these methods are very computationally intensive. Here, we explore the inherent parallelism of Bayesian learning and propose a hardware design that can be used for the reconstruction of such networks. The evaluation of the proposed design in a VirtexII demonstrates a speed up of the algorithm by 76 times over a software implementation in a Pentium 4. Iosifina Pournara, Christos-Savvas Bouganis, George A. Constantinides |
FPL | 2 |
| 2004 | A Steerable Complex Wavelet Construction and Its Implementation on FPGA
Christos-Savvas Bouganis, Peter Y. K. Cheung, Jeffrey Ng, Anil A. Bharath |
FPL | 1 |
| 2004 | Multiple Light Source DetectionabstractThis paper presents the V2R algorithm, a novel method for multiple light source detection using a Lambertian sphere as a calibration object. The algorithm segments the image of the sphere into regions that are each illuminated by a single virtual light and subtracts the virtual lights of adjacent regions to estimate the light source vectors. The algorithm uses all pixels within a region to form a robust estimate of the corresponding virtual light. The circumstances under which the light source detection problem lacks a unique solution are discussed in detail and the way in which the V2R algorithm resolves the ambiguity is explained. The V2R algorithm includes novel procedures for identifying the critical lines that bound the regions, for estimating the light source vectors, and for identifying opposite light pairs. Experiments are performed on synthetic and real images and the performance of the V2R algorithm is compared to that of a recent algorithm from the literature. The experimental results demonstrate that the proposed algorithm is robust and that it gives substantially improved accuracy. Christos-Savvas Bouganis, Mike Brookes |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Class-based Multiple Light Detection: An Application to FacesabstractMultiple light detection approaches have limited applicability in real life scenarios due to the need of certain calibration objects. We propose a novel approach, the “class-based ” image-based multiple light detection, that relaxes the above assumption to a “class ” of calibration objects. We formulate it as follows: Given a set of images of objects belonging to the same class, similar 3D shape and reflectance properties, and illuminated under point light sources, the purpose is to determine the light distribution of a new object of that class. This paper concentrates on the class of human faces. Six algorithms are proposed and their performance is evaluated with real images. Experiments show that a good performance is achieved for up to three lights using a small database of faces. 1 Christos-Savvas Bouganis, Mike Brookes |
BMVC | 1 |