VLDB 2026 Research / reviewers in the wild / expert
Bruno da Silva 0001
dblp:95/11194
· DBLP profile ↗
28ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-4877-9688ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 7 first-author · 20 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Equicore: Accelerating Clebsch-Gordan Tensor Product of Equivariant Neural Networks on FPGAabstractEquivariant neural networks (ENNs) are a powerful framework for modeling 3D geometric data in physical and biological systems. The Clebsch–Gordan tensor product (CGTP)—a core operation for preserving equivariance—remains the primary computational bottleneck in ENNs. Although Clebsch–Gordan (CG) coefficients exhibit pronounced structural sparsity, prior work has neither fully leveraged this property nor adopted hardware-friendly quantization, leading to limited efficiency. We present Equicore, a software–hardware co-design framework to accelerate CGTP in ENNs. Equicore introduces three key innovations: (1) a sparse-bypass strategy that exploits the CG structural sparsity together with a novel CG data format to pack the overlapping non-zeros, bypassing redundant data accesses and computations comparing to previous sparse solutions; (2) a merged-shift quantization strategy that enables full Int8 representation of irreps, weights, and CG coefficients using shift-only operations; and (3) a cascaded processing unit that tightly couples the FPGA hardware resources to achieve high operating frequency while supporting efficient sparse and quantized computation. Deployed on a AMD Virtex VCU128 platform, Equicore delivers up to 10.5× speedup and 17.4× energy-efficiency improvement over state-of-the-art GPU libraries and FPGA designs across diverse CGTP types in a benchmark of eleven ENN models. Shidi Tang, Chuanzhao Zhang, Ruiqi Chen 0001, Yuxuan Lv, Bruno da Silva 0001 |
DATE | 5 |
| 2026 | IDSPfree: An FPGA-Based Intrusion Detection System with DSP-Free Design
Abdessamad Nassihi, Ahmed Sadaqa, Muhammad Iqbal Khan, Ruiqi Chen 0001, Bruno da Silva 0001 |
ISCAS | 6 |
| 2026 | Power-Efficient Spiking Conversion of Deep Unfolded Transformers
Ahmed Sadaqa, Brent De Weerdt, Ruiqi Chen 0001, Nikos Deligiannis, Bruno da Silva 0001 |
ISCAS | 5 |
| 2026 | BenDan: Benchmarking DPU performance on FPGAs
Ahmed Sadaqa, Yanxiang Zhu, Shidi Tang, Ruiqi Chen 0001, Bruno da Silva 0001 |
Integr. | 9 |
| 2026 | FP8ApproxLib: An FPGA-based approximate multiplier library for 8-bit floating point
Ruiqi Chen 0001, Yangxintong Lyu, Shidi Tang, Jindong Li 0001, Yanxiang Zhu, Bruno da Silva 0001 |
J. Syst. Archit. | 8 |
| 2026 | FANE: FPGA-Based FP8 Approximate Neural Network EngineabstractThe 8-bit floating-point (FP8) format has gained growing interest in neural networks (NNs) for its superior dynamic range over traditional INT8. However, multiply-accumulate (MAC) operations remain a major source of power consumption during inference of NNs, which makes DSP-free design important, especially for edge FPGAs with few or no DSPs. Therefore, this brief presents FPGA-based FP8 approximate neural network engine (FANE), an FPGA-based approximate NN engine for FP8. We first introduce a novel approximation method that replaces the multiplications by linear additions. This approximate method reduces power consumption while maintaining high accuracy, outperforming the latest FP8 approximate multiplier by 53.15%. Based on this design, we construct an FP8 MAC unit and integrate it into both a convolution engine and a matrix–vector multiplication (MVM) unit. Finally, we integrate our design into a large language model (LLM). The result shows 61.5% higher efficiency (TOPS/W) than the previous design, demonstrating the superiority of FANE in terms of performance and power efficiency. The code of FANE is available on ourhttps://github.com/hanbao04/FANE-FPGA-based-FP8-Approximate-Neural-Network-Engine.git Shidi Tang, Jingdong Li, Ahmed Sadaqa, Ruiqi Chen 0001, Bruno da Silva 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | ATE-GCN: An FPGA-Based Graph Convolutional Network Accelerator with Asymmetrical Ternary QuantizationabstractTernary quantization can effectively simplify matrix multiplication, which is the primary computational operation in neural network models. It has shown success in FPGA-based accelerator designs for emerging models such as GAT and Transformer. However, existing ternary quantization methods can lead to substantial accuracy loss under certain weight distribution pat-terns, such as GCN. Furthermore, current FPGA-based ternary weight designs often focus on reducing resource consumption while neglecting full utilization of FPGA DSP blocks, limiting maximum performance. To address these challenges, we propose ATE-GCN, an FPGA-based asymmetrical ternary quantization GCN accelerator using a software-hardware co-optimization approach. First, we adopt an asymmetrical quantization strategy with specific interval divisions tailored to the bimodal distribution of GCN weights, reducing accuracy loss. Second, we design a unified processing element (PE) array on FPGA to support various matrix computation forms, optimizing FPGA resource usage while leveraging the benefits of cascade design and ternary quantization, significantly boosting performance. Finally, we implement the ATE-GCN prototype on the VCU118 FPGA board. The results show that ATE-GCN maintains an accuracy loss below 2%. Additionally, ATE-GCN achieves average performance improvements of$224.13\times$and$11.1\times$, with up to$898.82\times$and$69.9\times$energy consumption saving compared to CPU and GPU, respectively. Moreover, compared to state-of-the-art FPGA-based GCN accelerators, ATE-GCN improves DSP efficiency by 63% with an average latency reduction of 11%. Ruiqi Chen 0001, Shidi Tang, Yang Liu 0376, Yanxiang Zhu, Bruno da Silva 0001 |
DATE | 7 |
| 2025 | FPGA-Based Approximate Multiplier for FP8abstractThe 8-bit floating-point (FP8) data format has been increasingly adopted in neural network (NN) computations due to its superior dynamic range compared to traditional INT8. However, FP8-based multiplication, a core operation in NNs, still incurs significant power consumption. To address this issue, this paper presents an FPGA-based approximate multiplier design for FP8. Firstly, we conduct a bit-level analysis of the approximation method. Based on this analysis, we implement a fine-grained optimized design on mainstream FPGAs (AMD and Altera) using primitives and templates combined with physical layout constraints. Then, the accuracy and resource utilization of the FP8 approximate multiplier are evaluated and analyzed. The results indicate that, compared to previous FPGA-based 8-bit designs, our design achieves the minimal LUT consumption. Finally, we integrate the design into the inference phase of a representative NN model, demonstrating its excellent power efficiency. To the best of our knowledge, this is the first FPGA-based FP8 approximate multiplier design, which can serve as a benchmark for future designs and comparisons of FPGA-based low-precision floating-point approximate multipliers. The code of this work is available in our GitLab. Ruiqi Chen 0001, Yangxintong Lyu, Yanxiang Zhu, Shidi Tang, Bruno da Silva 0001 |
FCCM | 8 |
| 2025 | TrackGNN: A Highly Parallelized and Self-Adaptive GNN Accelerator for Track Reconstruction on FPGAsabstractReal-time track reconstruction in high energy physics imposes stringent latency constraints, hindering the deployment of graph neural networks (GNNs) on general-purpose platforms. We present TrackGNN11https//github.com/silvenachen/TrackGNN, an open-sourced GNN accelerator for track reconstruction. Using a dataflow architecture with multiple parallelism and a self-adaptive renaming mechanism, TrackGNN shows 27.6× speedup over CPUs, up to 101.1× over GPUs, and 5.7× over an FPGA overlay. Compared with FlowGNN, the renaming mechanism also reduces end-to-end latency by 1.12-1.16× with negligible resource overhead. Ruiqi Chen 0001, Bruno da Silva 0001, Giorgian Borca-Tasciuc, Dantong Yu, Cong Hao |
FCCM | 4 |
| 2025 | Diff-DiT: Temporal Differential Accelerator for Low-bit Diffusion Transformers on FPGAabstractDiffusion Transformer (DiT) models have shown superior generative capabilities in image and video synthesis, yet their high computational cost during inference remains a critical bottleneck. Temporal differential computation offers a promising solution to low-bit quantization by exploiting the temporal similarity in activations. However, applying this technique to DiT’s Attention layers introduces substantial memory and computation overheads.In this paper, we present Diff-DiT, the first FPGA accelerator designed for low-bit DiT inference with differential computation. To overcome the unique challenges of DiT quantization and hardware acceleration, we propose: (1) an approximated differential attention (ADA) method that selectively approximates attention computations across time steps using a significance score, enabling low-bit on-chip execution while minimizing memory overhead; (2) an optimal cross-cast data accessing pattern with flexible data reuse to maximize computational intensity during matrix multiplications; and (3) a half-condition splitting (HCS) dataflow optimization and fine-grained pipelining to reduce the computation and memory access latency.Extensive experiments show that Diff-DiT outperforms NVIDIA V100 GPU by 1.39× in end-to-end throughput and 5.60× in energy efficiency. When compared with state-of-the-art diffusion model accelerators, Diff-DiT also achieves 2.81× and 2.77× improvements in throughput and energy efficiency, respectively. Code is available on GitHub1. Shidi Tang, Pengwei Zheng, Ruiqi Chen 0001, Yuxuan Lv, Bruno da Silva 0001 |
ICCAD | 5 |
| 2025 | Acceleration of Fast Sample Entropy for FPGAsabstractComplexity measurement, essential in diverse fields like finance, biomedicine, climate science, and network traffic, demands real-time computation to mitigate risks and losses. Sample Entropy (SampEn) is an efficacious metric which quantifies the complexity by assessing the similarities among microscale patterns within the time-series data. Unfortunately, the conventional implementation of SampEn is computationally demanding, posing challenges for its application in real-time analysis, particularly for long time series. Field Programmable Gate Arrays (FPGAs) offer a promising solution due to their fast processing and energy efficiency, which can be customized to perform specific signal processing tasks directly in hardware. The presented work focuses on accelerating SampEn analysis on FPGAs for efficient time-series complexity analysis. A refined, fast, Lightweight SampEn architecture (LW SampEn) on FPGA, which is optimized to use sorted sequences to reduce computational complexity, is accelerated for FPGAs. Various sorting algorithms on FPGAs are assessed, and novel dynamic loop strategies and micro-architectures are proposed to tackle SampEn's undetermined search boundaries. Multi-source biomedical signals are used to profile the above design and select a proper architecture, underscoring the importance of customizing FPGA design for specific applications. Our optimized architecture achieves a 7x to 560x speedup over standard baseline architecture, enabling real-time processing of time-sensitive data. Chao Chen 0042, Chengyu Liu 0001, Jianqing Li 0002, Bruno da Silva 0001 |
IEEE Trans. Computers | 4 |
| 2025 | Optimized Modular Adder Architecture for Cryptographic Applications on FPGAsabstractModular addition is a fundamental operation in public-key cryptographic algorithms operating in finite fields, such as elliptic curve cryptography (ECC), Chebyshev polynomials, and post-quantum cryptography (PQC). The performance of these cryptographic algorithms is limited by the conventional modular adder approach, which incorporates two cascaded adders in series. This approach leads to a doubled critical path delay, ultimately causing a decrease in frequency despite utilizing a high-performance adder. This research presents a high-performance, low-area architecture for a modular adder, employing a novel approach. Specifically designed for various prime fields recommended in public key cryptography, the architecture optimally utilizes the carry chain and exploits the structural advantages of the 7-series field programmable gate array and series beyond. Implementation results demonstrate superior performance, achieving operating frequencies of 290.0 MHz for 192 bits and 205.5 MHz for 1024 bits. Notably, the proposed design performs modular addition in a single clock cycle, resulting in an approximate 57% frequency enhancement compared to the conventional approach. Consequently, this architecture stands as an optimal solution for systems demanding high-speed operations. Bachir Madani, Mohamed S. Azzaz, Said Sadoudi, Redouane Kaibou, Bruno da Silva 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | FASE: An FPGA-Based Accelerator for Lightweight Sample Entropy With Monte Carlo SamplingabstractSample entropy (SampEn) is an algorithm within information entropy that enables effective analysis of biological signals. Due to the need for extensive similarity matching operations, the SampEn calculation process is time-consuming. Although a series of fast SampEn algorithms have been proposed, they remain time-intensive when processing large data volumes. Additionally, previous field-programmable gate array (FPGA)-based hardware accelerators designed for SampEn suffer from architectural design limitations, consuming substantial on-chip memory resources and operating at low frequencies. In this article, we propose FASE, an FPGA-based accelerator for lightweight sample entropy (LW-SampEn) with Monte Carlo (MC) sampling. The FASE design comprises two main parts: algorithm and hardware optimizations. On the algorithmic side, we introduce MC sampling into the merge-sort-based LW-SampEn algorithm, named MCLW-SampEn. MCLW-SampEn effectively reduces the computation load for large data volumes while maintaining algorithmic accuracy. For hardware, we first design efficient sorting and allocation modules to address boundary localization and load imbalance issues in previous accelerator designs. Then, we replicate the computation across the main phases to enable parallel processing. Finally, we deploy the design on the Pynq-Z2 board for validation. Experimental results show that the proposed MCLW-SampEn algorithm achieves an average speed up of$3\times $over the LW-SampEn algorithm, with accuracy losses kept within 0.5%. Compared to state-of-the-art (SOTA) designs, FASE achieves an average speed up of$12.8\times $while reducing power consumption by 89.3%. Ablation studies indicate that, for the same algorithm, FASE offers a$7.4\times $speedup over related FPGA designs. Yuanhang Li, Zhengyang Huang, Chao Chen 0042, Ruiqi Chen 0001, Bruno da Silva 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | MLino bench: A comprehensive benchmarking tool for evaluating ML models on edge devicesabstractIn today's rapidly evolving technological landscape, Machine Learning (ML) has become an integral part of our daily lives, ranging from recommendation systems to advanced medical diagnostics and autonomous vehicles.As ML continues to advance, its applications extend beyond conventional boundaries.With the continuous refinement of the models and frameworks, the possibilities for leveraging these technologies into devices that traditionally lacked any form of computational autonomy are ever-expanding.This shift towards embedding ML capabilities directly into edge devices brings new challenges due to the stringent limitations these devices have in terms of memory, power consumption, and cost.The ML models implemented on such devices must find an equilibrium between memory footprint and performance, attaining a classification time that fulfills real-time demands and maintains a similar level of accuracy as the desktop version.Without automated assistance in managing these considerations, the complexity of evaluating multiple models can lead to suboptimal decisions.In this paper, we introduce MLino Bench, an open-source benchmarking tool tailored for assessing lightweight ML models on edge devices with limited resources and capabilities.The tool accommodates various models, frameworks, and platforms, presenting a meticulous design that enables a comprehensive evaluation directly on the target device.It encompasses crucial metrics such as on-target accuracy, classification time, and model size, providing a versatile framework that assists practitioners in decision-making when deploying models to such devices.The tool employs a fully streamlined benchmark flow involving training the ML model in a highlevel interpreted language, porting, compiling, flashing, and finally benchmarking on the actual target.Our experimental evaluation of the tool highlights its flexibility in assessing multiple ML models across different model hyperparameters, frameworks, datasets, and embedded platforms.Furthermore, a distinctive advantage compared to state-of-the-art ML benchmarking tools is the inclusion of classical ML models, including Random Forests, Decision Trees, Support Vector Machines, Naive Bayes, and more.This sets our tool apart from others that predominantly emphasize only neural network models.Due to this inclusive approach, our tool facilitates the evaluation of ML models across a broad spectrum of devices, ranging from resource-constrained edge devices to those with medium and advanced computational capabilities. Vlad-Eusebiu Baciu, Johan H. Stiens, Bruno da Silva 0001 |
J. Syst. Archit. | 3 |
| 2024 | FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future OpportunitiesabstractSparse matrix multiplication (SpMM) plays a critical role in high-performance computing applications, such as deep learning, image processing, and physical simulation. Field-Programmable Gate Arrays (FPGAs), with their configurable hardware resources, can be tailored to accelerate SpMMs. There has been considerable research on deploying sparse matrix multipliers across various FPGA platforms. However, the FPGA-based design of sparse matrix multipliers still presents numerous challenges. Therefore, it is necessary to summarize and organize the current work to provide a reference for further research. This article first introduces the computational method of SpMM and categorizes the different challenges of FPGA deployment. Following this, we introduce and analyze a variety of state-of-the-art FPGA-based accelerators tailored for SpMMs. In addition, a comparative analysis of these accelerators is performed, examining metrics including compression rate, throughput, and resource utilization. Finally, we propose potential research directions and challenges for further study of FPGA-based SpMM accelerators. Ruiqi Chen 0001, Bruno da Silva 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2022 | Acceleration of Fast Sample Entropy Towards Biomedical Applications on FPGAsabstractSample Entropy (SampEn) is an information en-tropy algorithm widely used for complexity analysis and chaos estimation in many applications. In particular, SampEn measures complexity of time series by the conditional probability of the inner pattern. Unfortunately, the straightforward implementation of SampEn is quadratic time complexity, restricting its real-time analysis ability for health applications and long-term data analysis. Although researchers have proposed fast versions of SampEn to avoid unnecessary comparisons, they have not been accelerated yet due to their performance bottleneck in the complex similarity pair process. In this paper, we evaluate fast SampEn algorithms by employing multi-source biomedical signals on an Field-Programmable Gate Arrays (FPGA). Since fast SampEn algorithms based of a pre-sorting stage promise to outperform other SampEn algorithms, Lightweight SampEn based on Merge Sort is here implemented and optimized. Dif-ferent type of optimizations, that can be generalized for similar Lightweight-based SampEn algorithms, are used to reduce the overall latency while the data throughput is increased. A load balancing strategy for multi similarity pair modules is also proposed to solve the unbalancing loads, a bottleneck when increasing the execution parallelism of this type of algorithms. As a result, the proposed SampEn architecture runs 10 times faster than the fastest SampEn implementation on a modern CPU. Chao Chen 0042, Bruno da Silva 0001, Jianqing Li 0002, Chengyu Liu 0001 |
FPT | 2 |
| 2022 | Effect of the Transimpedance Amplifier Topology on the Photoplethysmography SignalabstractPhotoplethysmography (PPG) is a widely used technique to extract physiological information non-invasively. Besides other factors, the transimpedance amplifier (TIA) topology choice has an impact on the PPG signal. This study presents a novel quantitative evaluation of seven different resistive TIA topologies in terms of four different Signal Quality Indexes (SQIs): AC amplitude, perfusion index, signal shape quality and signal-to-noise-ratio. Results indicate that a differential TIA configuration is preferred over the single-stage and that while the PD polarity has no impact on the PPG signal quality, the bias choice is important, being the zero bias configuration more beneficial than the reverse bias. Ángel Solé Morillo, Joan Lambert Cause, Bruno da Silva 0001, Juan Carlos García-Naranjo, Johan H. Stiens |
IECON | 3 |
| 2022 | Power Saving Techniques for Wearable Devices in Medical ApplicationsabstractMany people in the world are living with chronic diseases, demanding continuous monitoring, diagnosis, and treatment. Continuous physiological monitoring is key to providing preventive healthcare and accurate disease diagnosis, which leads to a growing demand for autonomous wearable technology. Wearable devices acquiring physiological information from the patient demand high-power efficiency to operate in a continuous acquisition mode. While power-saving techniques are applied in wearable devices for many application, very few are considered for biomedical applications. In this work, we explore existing techniques of power reduction for wearable medical devices. Our analysis addresses the power reduction of wearable medical devices and their generalization for different medical signal processing applications. In addition, we propose a taxonomy for power-saving techniques. The common categories of power-saving techniques are task scheduling, clock management, signal compression, and energy awareness. The presented analysis identifies the most appropriate and combined low-power techniques in wearable devices to reduce power consumption. Workineh Tesema, Bruno da Silva 0001, Worku Jimma, Johan H. Stiens |
IECON | 2 |
| 2022 | Frequency Evaluation of the Xilinx DPU Towards Energy EfficiencyabstractEmbedded Artificial Intelligence (AI) and the use of deep neural networks on embedded devices is becoming increasingly popular. The rise in popularity comes from the advantages gained from executing inference locally, such as privacy, and the increase in available platforms to accelerate AI. An example of such platforms are Field Programmable Gate Arrays (FPGAs). The flexibility of FPGAs resulted in the development of several accelerators for AI. One of these tools that is gaining more and more interest is Xilinx Vitis AI, which uses a configurable Deep Processing Unit (DPU) in an FPGA. The DPU is not only scalable in resource consumption, but also the frequency at which it operates. In this paper, the influence on the power and energy consumption of the DPU is investigated for different DPU configurations and frequencies. As a result, it is shown that increasing the frequency of the DPU can compensate a reduction in resource consumption. Furthermore, an increase in resources and frequency can result in an overall lower energy consumption due to a higher power consumption for a shorter time. Jurgen Vandendriessche, Bruno da Silva 0001, Abdellah Touhafi |
IECON | 2 |
| 2021 | AITIA: Embedded AI Techniques for Industrial ApplicationsabstractMotivated by an increasing interest from startups in embedded Artificial Intelligence (AI) and by their limited expertise, the AITIA Project targets the development of embedded AI techniques for industrial applications. This extended abstract presents the motivation and the solutions being developed towards four use cases: smart sensors, network intrusion detection, driver-assistance systems, and Industry 4.0. Marcelo Brandalero, Mitko Veleski, Hector Gerardo Muñoz Hernandez, Muhammad Ali 0010, Laurens Le Jeune, Toon Goedemé, Nele Mentens, Jurgen Vandendriessche, Lancelot Lhoest, Bruno da Silva 0001, Abdellah Touhafi, Diana Göhringer, Michael Hübner 0001 |
FPL | 10 |
| 2019 | Demonstration of a Multimode SoC FPGA-Based Acoustic CameraabstractThe relatively low-cost of the Micro-Electromechanical Systems (MEMS) microphones together with recent advances in the MEMS technology facilitates the construction of large MEMS microphone arrays, which are used to collect the acoustic information from certain beamed directions by applying beamforming techniques. The use of beamforming techniques to steer the microphone array response is the principle used by acoustic cameras to graphically display the acoustic information in a heatmap form. Our demonstrator exploits the heterogeneous nature of System-on-Chip (SoC) FPGA systems by generating real-time acoustic images on the FPGA component while alleviating the Wireless Sensor Networks (WSN) bandwidth limitations by performing acoustic image processing locally on the hard-core processor. As a result, a real-time acoustic heatmap is generated, enabling the visualization of the sound sources characteristics. Bruno da Silva 0001, Laurent Segers, An Braeken, Abdellah Touhafi |
FPL | 1 |
| 2018 | Efficiency analysis methodology of FPGAs based on lost frequencies, area and cycles
Jan Lemeire, Bruno da Silva 0001, An Braeken, Jan G. Cornelis, Abdellah Touhafi |
J. Parallel Distributed Comput. | 2 |
| 2017 | A partial reconfiguration based microphone array network emulatorabstractNowadays, microphone arrays are used in many applications for sound-source localization or acoustic enhancement. The current Micro-Electro-Mechanical Systems (MEMS) technology allows the development of networks of microphone arrays at a relatively low cost. Unfortunately, the evaluation of these networks requires controlled acoustic environments, such as anechoic chambers, to avoid possible distortions and acoustic artifacts. In this paper, we present a partial reconfigurable FPGA platform to emulate a network of microphone arrays. Our platform provides a controlled simulated acoustic environment, able to evaluate the impact of different network configurations such as the number of microphones per array, the network's topology or the used detection method. Data fusion techniques, combining the data collected by each node, are used in this platform. In addition, our platform is also capable to converge to the ideal network with regards to power consumption, while still maintaining the desired level of sound-source localization accuracy. A graphical user interface provides a friendly control of the network and the parameters under test during the execution of the partial reconfiguration operations. Several experiments are presented to demonstrate some of the capabilities of our platform. Bruno da Silva 0001, Federico Domínguez, An Braeken, Abdellah Touhafi |
FPL | 1 |
| 2017 | Demonstration of a partial reconfiguration based microphone array network emulatorabstractThe current Micro-Electro-Mechanical System (MEMS) technology allows to deploy relatively low-cost Wireless Sensor Networks (WSN) composed of MEMS microphone arrays for accurate sound-source localization. However, the evaluation and the selection of the most accurate and power-efficient network's topology is not trivial when considering dynamic MEMS microphone arrays. Despite software simulators are usually considered, they are high-computational intensive tasks which require hours to days to be completed. Our demonstrator is an FPGA-based network emulator, which provides a fast network design-space exploration. The user can easily evaluate a network's topology with different nodes' configurations and multiple sound sources in matter of seconds. An intuitive graphical user interface hides from the user the dynamic partial reconfigurations needed by the network emulator to set the node's configurations. As a result, a probability map generated from the fusion of the output data from the nodes and an error on the estimation of the sound-source location are graphically represented. Bruno da Silva 0001, Federico Domínguez, An Braeken, Abdellah Touhafi |
FPL | 1 |
| 2016 | Runtime reconfigurable beamforming architecture for real-time sound-source localizationabstractSound-source localization is used in many different real-time acoustic applications. Microphone arrays have the potential capability to recognize, profile and locate sound-sources in noisy environments. The quality response of such sensor arrays, however, is determined by the quantity of microphones. A higher number of microphones increases the computational demand, making real-time response challenging. In this paper, we present a scalable and runtime reconfigurable architecture to provide accurate sound-source localization in real-time. On one hand, the reconfigurable architecture is designed to be scalable in order to support a variable number of microphones. On the other hand, we use runtime reconfigurable look-up tables (CFGLUTs) to provide a dynamic response in real-time. Experiments demonstrate how an accurate sound-source localization is obtained in less than a few hundred milliseconds. As far as we are aware, it is the first time that runtime reconfiguration is applied to a reconfigurable architecture consisting of a sensor array. Bruno da Silva 0001, Laurent Segers, An Braeken, Abdellah Touhafi |
FPL | 1 |
| 2016 | A runtime reconfigurable FPGA-based microphone array for sound source localizationabstractMicrophone arrays are able to recognize, profile and locate sound-sources in noisy environments, but their quality is determined by the number of microphones. A higher number of microphones increases the computational demand, making real-time response challenging. In this demo, we present a scalable and runtime reconfigurable architecture able to support a variable number of microphones and orientations in order to provide accurate sound-source localization in real-time. Bruno da Silva 0001, Laurent Segers, An Braeken, Abdellah Touhafi |
FPL | 1 |
| 2013 | Performance and toolchain of a combined GPU/FPGA desktop (abstract only)abstractLow-power, high-performance computing nowadays relies on accelerator cards to speed up the calculations. Combining the power of GPUs with the flexibility of FPGAs enlarges the scope of problems that can be accelerated [2, 3]. We describe the performance analysis of a desktop equipped with a GPU Tesla 2050 and an FPGA Virtex-6 LX240T. First, the balance between the I/O and the raw peak performance is depicted using the roofline model [4]. Next, the performance of a number of image processing algorithms is measured and the results are mapped onto the roofline graph. This allows to compare the GPU and the FPGA and also to optimize the algorithms for both accelerators. A programming toolchain is implemented, consisting of OpenCL for the GPU and several High-Level Synthesis compilers for the FPGA. Our results show that the HLS compilers outperform handwritten code and offer a performance comparable to the GPU. In addition the FPGA compilers reduce the development time by an order of magnitude, at the expense of an increased resource consumption. The roofline model also shows that both accelerators are equally limited by the input/output bandwidth to the host. A well-tuned accelerator-based codesign, identifying the parallelism, the computation and data patterns of different classes of algorithms, will enable to maximize the performance of the combined GPU/FPGA system [1]. Bruno da Silva 0001, An Braeken, Erik H. D'Hollander, Abdellah Touhafi, Jan G. Cornelis, Jan Lemeire |
FPGA | 1 |
| 2013 | Comparing and combining GPU and FPGA accelerators in an image processing contextabstractNowadays, processors alone cannot deliver what computation hungry image processing applications demand. An alternative is to use hardware accelerators such as Graphics Processing Units (GPUs) or Field Programmable Gate Arrays (FPGAs). Applications, however, exhibit different performance characteristics depending on the accelerator. This paper describes the hybrid platform and the programming environment that allows to efficiently create programs on a combined GPU/FPGA desktop. We use the roofline model to identify the most appropriate accelerator for each application and High-Level Synthesis (HLS) tools to reduce the FPGA development time. To introduce our platform and tool chain both accelerators are compared by implementing a basic image operation. Next, a promising algorithm is explored and implemented, splitting and distributing the work between GPU, FPGA and CPU in order to validate the hybrid concept. Our results show that their combination exhibits a higher performance for computational intensive image processing applications than a GPU only. Bruno da Silva 0001, An Braeken, Erik H. D'Hollander, Abdellah Touhafi, Jan G. Cornelis, Jan Lemeire |
FPL | 1 |