Olivier Sentieys

dblp:28/4014 · DBLP profile ↗
← Back
109ranked-venue papers
1as first author
30since 2021 · last 2026
0000-0003-4334-6418ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 75 · 1 first-author · 26 since 2021Software engineering, systems software and programming languages · 22 · 11 since 2021Computer networks · 12 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KAN-SAs: Efficient Acceleration of Kolmogorov-Arnold Networks on Systolic Arrays
abstract
Kolmogorov-Arnold Networks (KANs) have garnered significant attention for their promise of improved parameter efficiency and explainability compared to traditional Deep Neural Networks (DNNs). KANs’ key innovation lies in the use of learnable non-linear activation functions, which are parametrized as splines. Splines are expressed as a linear combination of basis functions (B-splines). B-splines prove particularly challenging to accelerate due to their recursive definition. Systolic Array (SA)-based architectures have shown great promise as DNN accelerators thanks to their energy efficiency and low latency. However, their suitability and efficiency in accelerating KANs have never been assessed. Thus, in this work, we explore the use of SA architecture to accelerate the KAN inference. We show that, while SAs can be used to accelerate part of the KAN inference, their utilization can be reduced to 30%. Hence, we propose KAN-SAs, a novel SA-based accelerator that leverages intrinsic properties of B-splines to enable efficient KAN inference. By including a non-recursive B-spline implementation and leveraging the intrinsic KAN sparsity, KAN-SAs enhances conventional SAs, enabling efficient KAN inference, in addition to conventional DNNs. KANSAs achieves up to 100% SA utilization and up to 50% clock cycles reduction compared to conventional SAs of equivalent area, as shown by hardware synthesis results on a 28nm FDSOI technology. We also evaluate different configurations of the accelerator on various KAN applications, confirming the improved efficiency of KAN inference provided by KAN-SAs.
Sohaib Errabii, Olivier Sentieys, Marcello Traiola
DATE2
2026 Hardware Stencil Accelerator with Periodic Boundary Conditions
abstract
International audience
Damien Simon, Olivier Sentieys, Sylvain Lefebvre 0001
FCCM2
2026 Project Highlights - Reliability Evaluation for ARCHYTAS AI hardware accelerators
Angeliki Kritikakou, Fernando Santos 0001, Marcello Traiola, Rafael Billig Tonetto, Olivier Sentieys, Paolo Rech, Haralampos-G. D. Stratigopoulos, Georgios Keramidas
IOLTS5
2026 Energy-Efficient and Reliable Task Mapping and Offloading for Multicore Edge Devices With DVFS
abstract
Multicore platforms based on NoC are promising architectures for safety-critical applications. Application execution performance is determined by task mapping, with reliable execution, real-time response, and energy efficiency as requirements. We can perform task duplication, DVFS, and multipath routing to meet these requirements during task mapping. Furthermore, the computation platforms have limited computation capacity and energy supply in several application domains. Some complex tasks can be offloaded from the edge device to the cloud for execution. However, such task offloading influences task mapping on the edge device. Existing approaches seldom consider the correlation of task offloading to the cloud and task mapping on the edge device. To address this limitation, we jointly consider task mapping inside the NoC-based multicore edge device and task offloading to the cloud to optimize energy consumption while satisfying reliability and real-time constraints. This problem is formulated as a mixed-integer nonlinear programming and linearized to find the optimal solution. We propose a novel three-step heuristic with a feedback mechanism to enhance task schedulability and reduce computation time. We evaluate the behavior of our approaches through exhaustive simulations. The results show that our approaches outperform existing methods in terms of energy efficiency, task reliability, and schedulability.
Lei Mo, Tamim M. Al-Hasan, Angeliki Kritikakou, Xiaojun Zhai, Olivier Sentieys, Shibo He
IEEE Internet Things J.6
2026 QoS-Aware Approximate Task Mapping on Heterogeneous Multicore Platforms with DVFS and Task Migration
abstract
Heterogeneous Multicore Platforms (HMPs) have been widely adopted to execute tasks across a range of applications. Under limited system resources and diverse application requirements, allocating and executing dependent Approximate Computing (AC) tasks on these platforms to achieve high Quality-of-Service (QoS) is challenging. Dynamic Voltage and Frequency Scaling (DVFS) and task migration have proven effective for improving QoS while balancing time and energy consumption. However, existing approaches often overlook the migration overhead and the resulting dynamic changes in task dependencies, which can adversely affect mapping outcomes. To address these issues, this article presents a novel AC task mapping method that maximizes system QoS under multiple constraints on HMPs, accounting for task migration overhead, DVFS, and changes in Directed Acyclic Graph (DAG) topology. We first formulate this joint design problem as a complex nonlinear programming problem. Next, we linearize the nonlinear terms without performance loss by introducing auxiliary variables and additional constraints. Building on this formulation, we propose an optimal (OPT) and a low-complexity Heuristic Algorithm (HEU), derived from problem decomposition and a greedy strategy, which divides the Mixed-Integer Non-Linear Programming (MINLP) problem into two smaller subproblems with fewer variables and constraints, solving them sequentially. The simulation results show that the proposed OPT method achieves higher QoS performance, measured at about 2.389 times on average and up to 4.115 times, while its feasibility is increased to about 3.263 times on average and up to 9.667 times, compared to other state-of-the-art methods. In addition, the average QoS of the proposed HEU method is about 0.577 times that of the proposed method, but its computation time is over a thousand times shorter.
Hengyan Song, Lei Mo, Tamim M. Al-Hasan, Angeliki Kritikakou, Xiaojun Zhai, Shibo He, Olivier Sentieys
ACM Trans. Embed. Comput. Syst.7
2025 Hardware-Aware Training for Multiplierless Convolutional Neural Networks
abstract
Many computer vision tasks use convolutional neural networks (CNNs). These networks have a significant computational cost and complex implementations, in particular on embedded systems. A common way to implement CNNs on integrated circuits is to use low-precision quantized weights and activations instead of de facto floating-point (FP) ones. This is important to reduce the implementation cost. However, this has drawbacks regarding accuracy, and Quantization-Aware Training (QAT) is one of the most popular approaches to mitigate this issue. In this article, we introduce a multiplierless-aware training approach that significantly reduces hardware resource consumption. We propose to incrementally fix weights to their current value based on their implementation cost. To compute this cost, we base our approach on a Multiple Constant Multiplication (MCM) shift-and-add solving technique. With this idea, we show a global implementation cost reduction by around 25% w. r. t. a vanilla QAT approach without hardware usage in the loop. Compared to state-of-the-art multiplierless-aware training methods, the network accuracy of our designs is closer to that of a vanilla QAT baseline.
Rémi Garcia 0002, Léo Pradels, Silviu-Ioan Filip, Olivier Sentieys
ARITH4
2025 MPTorch-FPGA: A Custom Mixed-Precision Framework for FPGA-Based DNN Training
abstract
Training Deep Neural Networks (DNNs) is computationally demanding, leading to a growing interest in reduced precision formats to enhance hardware efficiency. Several frame-works explore custom number formats with parameterizable precision through software emulation on CPUs or GPUs. However, they lack comprehensive support for different rounding modes and struggle to accurately evaluate the impact of custom precision for FPGA-based targets. This paper introduces MPTorch-FPGA, an extension of the MPTorch framework for performing custom, multi-precision inference and training computations in CPU, GPU, and FPGA environments in PyTorch. MPTorch-FPGA can generate a model-specific accelerator for DNN training, with customizable sizes and arithmetic implementations, providing bit-level accuracy with respect to emulated low precision DNN training on GPUs or CPUs. An offline matching algorithm selects one of several pre-generated (static) FPGA configurations using a custom performance model to estimate latency. To showcase the versatility of MPTorch-FPGA, we present a series of training benchmarks using diverse DNN models, exploring a range of number format configurations and rounding modes. We report both accuracy and hardware performance metrics, verifying the precision of our performance model by comparing estimated and measured latencies across multiple benchmarks. These results highlight the flexibility and practical value of our framework.
Sami Ben Ali, Silviu-Ioan Filip, Olivier Sentieys, Guy Lemieux
DATE3
2025 SERA-Float: A Soft Error Resilient Approximate Floating-Point Computing Format
abstract
Approximate computing (AxC) reduces power consumption with minimal accuracy loss, benefiting error-tolerant, compute-intensive tasks such as machine learning, deep learning, and image processing. However, existing AxC methods often ignore the vulnerability to soft errors. Such errors can interact with approximation-induced errors, causing system failures or unexpected exceptions. To our knowledge, no work has addressed both soft error resilience and exception avoidance in approximate floating-point computing. This gap is particularly critical in deep neural network (DNN) inference, where soft error-induced errors or exceptions can significantly affect the stability and accuracy of computations.In this paper, we introduce SERA-Float, an approximate floating-point format resilient to soft errors. Specifically, it is designed to protect floating-point computations from soft error-induced errors and exception-triggering bit-flips. Unlike prior floating-point formats, SERA-Float protects the sign and exponent bits using error-correcting codes and relies on storing 8 valid bits of mantissa rather than performing coarse truncation. Additionally, by tracking critical bits in the floating-point representation, SERA-Float prevents overflow, underflow, and NaN exceptions. Our evaluation demonstrates that SERA-Float improves the reliability of floating-point operations during DNN inference by significantly reducing exceptions and ensuring the stability of computations. Moreover, it enables energy-efficient arithmetic by leveraging narrower arithmetic units, yielding up to 80.3% energy savings per multiplication with a 0.9% reduction in DNN inference accuracy.
Vishesh Mishra, Marcello Traiola, Angeliki Kritikakou, Olivier Sentieys, Urbi Chatterjee
ICCAD4
2025 Side-Channel Extraction of Dataflow AI Accelerator Hardware Parameters
abstract
Dataflow neural network accelerators efficiently process AI tasks on FPGAs, with deployment simplified by ready-to-use frameworks and pre-trained models. However, this convenience makes them vulnerable to malicious actors seeking to reverse engineer valuable Intellectual Property (IP) through Side-Channel Attacks (SCA). This paper proposes a methodology to recover the hardware configuration of dataflow accelerators generated with the FINN framework. Through unsupervised dimensionality reduction, we reduce the computational overhead compared to the state-of-the-art, enabling lightweight classifiers to recover both folding and quantization parameters. We demonstrate an attack phase requiring only 337 ms to recover the hardware parameters with an accuracy of more than 95% and 421 ms to fully recover these parameters with an averaging of 4 traces for a FINN-based accelerator running a CNN, both using a random forest classifier on side-channel traces, even with the accelerator dataflow fully loaded. This approach offers a more realistic attack scenario than existing methods, and compared to SoA attacks based on tsfresh, our method requires 940x and 110x less time for preparation and attack phases, respectively, and gives better results even without averaging traces.
Guillaume Lomet, Rubén Salvador, Brice Colombier, Vincent Grosso, Olivier Sentieys, Cédric Killian
IOLTS5
2024 Efficient Neural Networks: from SW optimization to specialized HW accelerators
abstract
Artificial Neural Networks (ANNs) appear to be one of the technological revolutions of recent human history. The capability of such systems does not come at a low cost, which led researchers to develop more and more efficient techniques to implement them. Optimization approaches have been developed, such as pruning and quantization, leading to reduced memory and computation requirements. Furthermore, such approaches are adapted to the specific hardware platform features to further increase efficiency. To improve it further, the HW programmability can be traded off in favor of more specialized custom HW ANN accelerators. In this education abstract, we illustrate how optimizing operations execution at different levels, from SW to HW, can improve the efficiency of ANN execution.
Marcello Traiola, Angeliki Kritikakou, Silviu-Ioan Filip, Olivier Sentieys
CASES4
2024 Cross-Layer Reliability Evaluation and Efficient Hardening of Large Vision Transformers Models
abstract
Vision Transformers (ViTs) are highly accurate Machine Learning (ML) models. However, their large size and complexity increase the expected error rate due to hardware faults. Measuring the error rate of large ViT models is challenging, as conventional microarchitectural fault simulations can take years to produce statistically significant data. This paper proposes a two-level evaluation based on data collected through more than 70 hours of neutron beam experiments and more than 600 hours of software fault simulation. We consider 12 ViT models executed in 2 NVIDIA GPU architectures. We first characterize the fault model in ViT's kernels to identify the faults more likely to propagate to the output. We then design dedicated procedures efficiently integrated into the ViT to locate and correct these faults. We propose Maximum corrupted Malicious values (MaxiMals), an experimentally tuned low-cost mitigation solution to reduce the impact of transient faults on ViTs. We demonstrate that MaxiMals can correct 90.7% of critical failures, with execution time overheads as low as 5.61%.
Lucas Roquet, Fernando Santos 0001, Paolo Rech, Marcello Traiola, Olivier Sentieys, Angeliki Kritikakou
DAC5
2024 A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
abstract
Training Deep Neural Networks (DNNs) can be computationally demanding, particularly when dealing with large models. Recent work has aimed to mitigate this computational challenge by introducing 8-bit floating-point (FP8) formats for multiplication. However, accumulations are still done in either half (16-bit) or single (32-bit) precision arithmetic. In this paper, we investigate lowering accumulator word length while maintaining the same model accuracy. We present a multiply-accumulate (MAC) unit with FP8 multiplier inputs and FP12 accumulations, which leverages an optimized stochastic rounding (SR) implementation to mitigate swamping errors that commonly arise during low precision accumulations. We investigate the hardware implications and accuracy impact associated with varying the number of random bits used for rounding operations. We additionally attempt to reduce MAC area and power by proposing a new scheme to support SR in floating-point MAC and by removing support for subnormal values. Our optimized eager SR unit significantly reduces delay and area when compared to a classic lazy SR design. Moreover, when compared to MACs utilizing single- or half-precision adders, our design showcases notable savings in all metrics. Furthermore, our approach consistently maintains near baseline accuracy across a diverse range of computer vision tasks, making it a promising alternative for low-precision DNN training.
Sami Ben Ali, Silviu-Ioan Filip, Olivier Sentieys
DATE3
2024 Reliability and Security of AI Hardware
abstract
In recent years, Artificial Intelligence (AI) systems have achieved revolutionary capabilities, providing intelligent solutions that surpass human skills in many cases. However, such capabilities come with power-hungry computation workloads. Therefore, the implementation of hardware acceleration becomes as fundamental as the software design to improve energy efficiency, silicon area, and latency of AI systems. Thus, innovative hardware platforms, architectures, and compiler-level approaches have been used to accelerate AI workloads. Crucially, innovative AI acceleration platforms are being adopted in application domains for which dependability must be paramount, such as autonomous driving, healthcare, banking, space exploration, and industry 4.0. Unfortunately, the complexity of both AI software and hardware makes the dependability evaluation and improvement extremely challenging. Studies have been conducted on both the security and reliability of AI systems, such as vulnerability assessments and countermeasures to random faults and analysis for side-channel attacks. This paper describes and discusses various reliability and security threats in AI systems, and presents representative case studies along with corresponding efficient countermeasures.
Dennis Gnad, Martin Gotthard, Jonas Krautter, Angeliki Kritikakou, Vincent Meyers, Paolo Rech, Josie E. Rodriguez Condia, Annachiara Ruospo, Ernesto Sánchez 0001, Fernando Santos 0001, Olivier Sentieys, Mehdi Baradaran Tahoori, Russell Tessier, Marcello Traiola
ETS11
2024 A Hardware Instruction Generation Mechanism for Energy-Efficient Computational Memories
abstract
In the Computing-In-Memory (CIM) approach, computations are directly performed within the data storage unit, which often results in energy reduction. This makes it particularly well fitted for embedded systems, highly constrained in energy efficiency. It is commonly admitted that this energy reduction comes from less data transfers between the CPU and the main memory. Nevertheless, preparing and sending instructions to the computational memory also consumes energy and time, hence limiting overall performance. In this paper, we present a hardware instruction generation mechanism integrated in computational memories and evaluate its benefit for Integer General Matrix Multiplication (IGeMM) operations. The proposed mechanism is implemented in the computational memory controller and translates macro-instructions into corresponding micro-instructions needed to execute the kernel on stored data. We modified an existing near-memory computing architecture and extracted corresponding energy consumption figures using post-layout simulations for the complete SoC. Our proposed architecture, NEar memory computing Macro-Instruction Kernel Accelerator (NeMIKA), provides an 8.2× speed-up and a 4.6× energy consumption reduction compared to a state-of-the-art CIM accelerator based on micro-instructions, while inducing an area overhead of only 0.1%.
Léo De La Fuente, Jean-Frédéric Christmann, Manuel Pezzin, Matthias Remars, Olivier Sentieys
ISCAS5
2024 Lightweight Hardware-Based Cache Side-Channel Attack Detection for Edge Devices (Edge-CaSCADe)
abstract
Cache Side-Channel Attacks (CSCAs) have been haunting most processor architectures for decades now. Existing approaches to mitigation of such attacks have certain drawbacks, namely software mishandling, performance overhead, and low throughput due to false alarms. Hence,“mitigation only when detected”should be the approach to minimize the effects of such drawbacks. We propose a novel methodology of fine-grained detection of timing-based CSCA using a hardware-based detection module. We discuss the design, implementation, and use of our proposed detection module in processor architectures. Our approach successfully detects attacks that flush secret victim information from cache memory like Flush+Reload, Flush+Flush, Prime+Probe, Evict+Probe, and Prime+Abort, commonly known as cache timing attacks. Detection is on time with minimal performance overhead. The parameterizable number of counters used in our module allows detection of multiple attacks on multiple sensitive locations simultaneously. The fine-grained nature ensures negligible false alarms, severely reducing the need for any unnecessary mitigation. The proposed work is evaluated by synthesizing the entire detection algorithm as an attack detection block, Edge-CaSCADe, in a RISC-V processor as a target example. The detection results are checked under different workload conditions with respect to the number of attackers and the number of victims having RSA-, AES-, and ECC-based encryption schemes like ECIES, and on benchmark applications like MiBench and Embench. More than 98% detection accuracy within 2% of the beginning of an attack can be achieved with negligible false alarms. The detection module has an area and power overhead of 0.9% to 2% and 1% to 2.1% for the targeted RISC-V processor core without cache for one to five counters, respectively. The detection module does not affect the processor critical path and hence has no impact on its maximum operating frequency.
Pavitra Prakash Bhade, Joseph Paturel, Olivier Sentieys, Sharad Sinha
ACM Trans. Embed. Comput. Syst.3
2024 Combining Weight Approximation, Sharing and Retraining for Neural Network Model Compression
abstract
Neural network model compression is very important to achieve model deployment based on the memory and storage available in different computing systems. Generally, the continuous drive for higher accuracy in these models increases their size and complexity, making it challenging to deploy them on resource-constrained computing environments. This article proposes various algorithms for model compression by exploiting weight characteristics and conducts an in-depth study of their performance. The algorithms involve manipulating exponents and mantissa in the floating-point representations of weights. In addition, we also present a retraining method that uses the proposed algorithms to further reduce the size of pre-trained models. The results presented in this article are mainly on BFloat16 floating-point format. The proposed weight manipulation algorithms save at least 20% of memory on state-of-the-art image classification models with very minor accuracy loss. This loss is bridged using the retraining method that saves at least 30% of memory, with potential memory savings of up to 43%. We compare the performance of the proposed methods against the state-of-the-art model compression techniques in terms of accuracy, memory savings, inference time, and energy.
Prachi Kashikar, Olivier Sentieys, Sharad Sinha
ACM Trans. Embed. Comput. Syst.2
2023 Maximizing Computing Accuracy on Resource-Constrained Architectures
abstract
With the growing complexity of applications, design-ers need to fit more and more computing kernels into a limited energy or area budget. Therefore, improving the quality of results of applications in electronic devices with a constraint on its cost is becoming a critical problem. Word Length Optimization (WLO) is the process of determining bit-width for variables or operations represented using fixed-point arithmetic to trade-off between quality and cost. State-of-the-art approaches mainly solve WLO given a quality (accuracy) constraint. In this paper, we first show that existing WLO procedures are not adapted to solve the problem of optimizing accuracy given a cost constraint. It is then interesting and challenging to propose new methods to solve this problem. Then, we propose a Bayesian optimization based algorithm to maximize the quality of computations under a cost constraint (i.e., energy in this paper). Experimental results indicate that our approach outperforms conventional WLO approaches by improving the quality of the solutions by more than 170%.
Van-Phu Ha, Olivier Sentieys
DATE2
2023 A machine-learning-guided framework for fault-tolerant DNNs
abstract
Deep Neural Networks (DNNs) show promising per-formance in several application domains. Nevertheless, DNN results may be incorrect, not only because of the network intrinsic inaccuracy, but also due to faults affecting the hardware. Ensuring the fault tolerance of DNN is crucial, but common fault tolerance approaches are not cost-effective, due to the prohibitive overheads for large DNNs. This work proposes a comprehensive framework to assess the fault tolerance of DNN parameters and cost-effectively protect them. As a first step, the proposed framework performs a statistical fault injection. The results are used in the second step with classification-based machine learning methods to obtain a bit-accurate prediction of the criticality of all network parameters. Last, Error Correction Codes (ECCs) are selectively inserted to protect only the critical parameters, hence entailing low cost. Thanks to the proposed framework, we explored and protected two Convolutional Neural Networks (CNNs), each with four different data encoding. The results show that it is possible to protect the critical network parameters with selective ECCs while saving up to 79% memory w.r.t. conventional ECC approaches.
Marcello Traiola, Angeliki Kritikakou, Olivier Sentieys
DATE3
2023 Impact of Transient Faults on Timing Behavior and Mitigation with Near-Zero WCET Overhead
abstract
As time-critical systems require timing guarantees, Worst-Case Execution Times (WCET) have to be employed. However, WCET estimation methods usually assume fault-free hardware. If proper actions are not taken, such fault-free WCET approaches become unsafe, when faults impact the hardware during execution. The majority of approaches, dealing with hardware faults, address the impact of faults on the functional behavior of an application, i.e., denial of service and binary correctness. Few approaches address the impact of faults on the application timing behavior, i.e., time to finish the application, and target faults occurring in memories. However, as the transistor size in modern technologies is significantly reduced, faults in cores cannot be considered negligible anymore. This work shows that faults not only affect the functional behavior, but they can have a significant impact on the timing behavior of applications. To expose the overall impact of faults, we enhance vulnerability analysis to include not only functional, but also timing correctness, and show that faults impact WCET estimations. As common techniques to deal with faults, such as watchdog timers and re-execution, have large timing overhead for error detection and correction, we propose a mechanism with near-zero and bounded timing overhead. A RISC-V core is used as a case study. The obtained results show that faults can lead up to almost 700% increase in the maximum observed execution time between fault-free and faulty execution without protection, affecting the WCET estimations. On the contrary, the proposed mechanism is able to restore fault-free WCET estimations with a bounded overhead of 2 execution cycles.
Pegdwende Romaric Nikiema, Angeliki Kritikakou, Marcello Traiola, Olivier Sentieys
ECRTS4
2023 harDNNing: a machine-learning-based framework for fault tolerance assessment and protection of DNNs
abstract
Deep Neural Networks (DNNs) show promising performance in several application domains, such as robotics, aerospace, smart healthcare, and autonomous driving. Nevertheless, DNN results may be incorrect, not only because of the network intrinsic inaccuracy, but also due to faults affecting the hardware. Indeed, hardware faults may impact the DNN inference process and lead to prediction failures. Therefore, ensuring the fault tolerance of DNN is crucial. However, common fault tolerance approaches are not cost-effective for DNNs protection, because of the prohibitive overheads due to the large size of DNNs and of the required memory for parameter storage. In this work, we propose a comprehensive framework to assess the fault tolerance of DNNs and cost-effectively protect them. As a first step, the proposed framework performs data-type-and-layer-based fault injection, driven by the DNN characteristics. As a second step, it uses classification-based machine learning methods in order to predict the criticality, not only of network parameters, but also of their bits. Last, dedicated Error Correction Codes (ECCs) are selectively inserted to protect the critical parameters and bits, hence protecting the DNNs with low cost. Thanks to the proposed framework, we explored and protected two Convolutional Neural Networks (CNNs), each with four different data encoding. The results show that it is possible to protect the critical network parameters with selective ECCs while saving up to 83% memory w.r.t. conventional ECC approaches.
Marcello Traiola, Angeliki Kritikakou, Olivier Sentieys
ETS3
2023 Approximation-Aware Task Deployment on Heterogeneous Multicore Platforms With DVFS
abstract
Heterogeneous (HE) multicore platforms, such as ARM big.LITTLE, are widely used to execute embedded applications under multiple and contradictory constraints, such as energy consumption and real-time (RT) execution. To fulfill these constraints and optimize system performance, application tasks should be efficiently mapped on multicore platforms. Embedded applications are usually tolerant to approximated results but acceptable quality of service (QoS). Modeling-embedded applications by using the elastic task model, namely, imprecise computation (IC) task model, can balance system QoS, energy consumption, and RT performance during task deployment. However, state-of-the-art approaches seldom consider the problem of IC task deployment on HE multicore platforms. They typically neglect task migration, which can improve the solutions due to its flexibility during the task deployment process. This article proposes a novel QoS-aware task deployment method to maximize system QoS under energy and RT constraints, where the frequency assignment (FA), task allocation (TA), scheduling, and migration are optimized simultaneously. The task deployment problem is formulated as mixed-integer nonlinear programming. Then, it is linearized to mixed-integer linear programming to find the optimal (OPT) solution. Furthermore, based on the problem structure and problem decomposition, we propose a novel heuristic (HEU) with low computational complexity. The subproblems regarding FA, TA, scheduling, and adjustment are considered and solved in sequence. Finally, the simulation results show that the proposed task deployment method improves the system QoS by 31.2% on average (up to 112.8%) compared to the state-of-the-art methods and the designed HEU achieves about 53.9% (on average) performance of the OPT solution with a negligible computing time.
Xinmei Li, Lei Mo, Angeliki Kritikakou, Olivier Sentieys
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Lossless Neural Network Model Compression Through Exponent Sharing
abstract
Artificial intelligence (AI) on the edge has emerged as an important research area in the last decade to deploy different applications in the domains of computer vision and natural language processing on tiny devices. These devices have limited on-chip memory and are battery-powered. On the other hand, neural network (NN) models require large memory to store model parameters and intermediate activation values. Thus, it is critical to make the models smaller so that their on-chip memory requirements are reduced. Various existing techniques like quantization and weight-sharing reduce model sizes at the expense of some loss in accuracy. We propose a lossless technique of model size reduction by focusing on the sharing of exponents in weights, which is different from the sharing of weights. We present results based on generalized matrix multiplication (GEMM) in NN models. Our method achieves at least a 20% reduction in memory when using Bfloat16 and around 10% reduction when using IEEE single-precision floating point, for models, in general, with a very small impact (up to 10% on the processor and less than 1% on FPGA) on the execution time with no loss in accuracy. On specific models from HLS4ML, about 20% reduction in memory is observed in single precision with little execution overhead.
Prachi Kashikar, Olivier Sentieys, Sharad Sinha
IEEE Trans. Very Large Scale Integr. Syst.2
2022 Flodam: Cross-Layer Reliability Analysis Flow for Complex Hardware Designs
abstract
Modern technologies make hardware designs more and more sensitive to radiation particles and related faults. As a result, analysing the behavior of a system under radiation-induced faults has become an essential part of the system design process. Existing approaches either focus on analysing the radiation impact at the lower hardware design layers, without further propagating any radiation-induced fault to the system execution, or analyse system reliability at higher hardware or application layers, based on fault models that are agnostic of the fabrication technology and the radiation environment. Flodam combines the benefits of existing approaches by providing a novel cross-layer reliability analysis from the semiconductor layer up to the application layer, able to quantify the risks of faults under a given context, taking into account the environmental conditions, the physical hardware design and the application under study.
Angeliki Kritikakou, Olivier Sentieys, Guillaume Hubert, Youri Helen, Jean-Francois Coulon, Patrice Deroux-Dauphin
DATE2
2022 Mixing Low-Precision Formats in Multiply-Accumulate Units for DNN Training
abstract
The most compute-intensive stage of deep neural network (DNN) training is matrix multiplication where the multiply-accumulate (MAC) operator is key. To reduce training costs, we consider using low-precision arithmetic for MAC operations. While low-precision training has been investigated in prior work, the focus has been on reducing the number of bits in weights or activations without compromising accuracy. In contrast, the focus in this paper is on implementation details beyond weight or activation width that affect area and accuracy. In particular, we investigate the impact of fixed- versus floating-point representations, multiplier rounding, and floating-point exceptional value support. Results suggest that (1) low-precision floating-point is more area-effective than fixed-point for multiplication, (2) standard IEEE-754 rules for subnormals, NaNs, and intermediate rounding serve little to no value in terms of accuracy but contribute significantly to area, (3) low-precision MACs require an adaptive loss-scaling step during training to compensate for limited representation range, and (4) fixed-point is more area-effective for accumulation, but the cost of format conversion and downstream logic can swamp the savings. Finally, we note that future work should investigate accumulation structures beyond the MAC level to achieve further gains.
Mariko Tatsumi, Silviu-Ioan Filip, Caroline White, Olivier Sentieys, Guy Lemieux
FPT4
2022 Functional and Timing Implications of Transient Faults in Critical Systems
abstract
Embedded systems in critical domains, such as auto-motive, aviation, space domains, are often required to guarantee both functional and temporal correctness. Considering transient faults, fault analysis and mitigation approaches are implemented at various levels of the system design, in order to maintain the functional correctness. However, transient faults and their mitigation methods have a timing impact, which can affect the temporal correctness of the system. In this work, we expose the functional and the timing implications of transient faults for critical systems. More precisely, we initially highlight the timing effect of transient faults occurring in the combinational and sequential logic of a processor. Furthermore, we propose a full stack vulnerability analysis that drives the design of selective hardware-based mitigation for real-time applications. Last, we study the timing impact of software-based reliability mitigation methods applied in a COTS GPU, using a fault tolerant middleware.
Angeliki Kritikakou, Panagiota Nikolaou, Ivan Rodriguez-Ferrandez, Joseph Paturel, Leonidas Kosmidis, Maria K. Michael, Olivier Sentieys, David Steenari
IOLTS7
2022 Experimental evaluation of neutron-induced errors on a multicore RISC-V platform
abstract
RISC-V architectures have gained importance in the last years due to their flexibility and open-source Instruction Set Architecture (ISA), allowing developers to efficiently adopt RISC-V processors in several domains with a reduced cost. For application domains, such as safety-critical and mission-critical, the execution must be reliable as a fault can compromise the system’s ability to operate correctly. However, the application’s error rate on RISC-V processors is not significantly evaluated, as it has been done for standard x86 processors. In this work, we investigate the error rate of a commercial RISC-V ASIC platform, the GAP8, exposed to a neutron beam. We show that for computing-intensive applications, such as classification Convolutional Neural Networks (CNN), the error rate can be $3.2 \times$ higher than the average error rate. Additionally, we find that the majority (96.12%) of the errors on the CNN do not generate misclassifications. Finally, we also evaluate the events that cause application interruption on GAP8 and show that the major source of incorrect interruptions is application hangs (i.g., due to an infinite loop or a racing condition)
Fernando Santos 0001, Angeliki Kritikakou, Olivier Sentieys
IOLTS3
2021 Leveraging Bayesian Optimization to Speed Up Automatic Precision Tuning
abstract
Using just the right amount of numerical precision is an important aspect for guaranteeing performance and energy efficiency requirements. Word-Length Optimization (WLO) is the automatic process for tuning the precision, i.e., bit-width, of variables and operations represented using fixed-point arithmetic. However, state-of-the-art precision tuning approaches do not scale well in large applications where many variables are involved. In this paper, we propose a hybrid algorithm combining Bayesian optimization (BO) and a fast local search to speed up the WLO procedure. Through experiments, we first show some evidence on how this combination can improve exploration time. Then, we propose an algorithm to automatically determine a reasonable transition point between the two algorithms. By statistically analyzing the convergence of the probabilistic models constructed during BO, we derive a stopping condition that determines when to switch to the local search phase. Experimental results indicate that our algorithm can reduce exploration time by up to 50%-80% for large benchmarks.
Van-Phu Ha, Olivier Sentieys
DATE2
2021 AdequateDL: Approximating Deep Learning Accelerators
abstract
The design and implementation of Convolutional Neural Networks (CNNs) for deep learning (DL) is currently receiving a lot of attention from both industrials and academics. However, the computational workload involved with CNNs is often out of reach for low power embedded devices and is still very costly when running on datacenters. By relaxing the need for fully precise operations, approximate computing substantially improves performance and energy efficiency. Deep learning is very relevant in this context, since playing with the accuracy to reach adequate computations will significantly enhance performance, while keeping quality of results in a user-constrained range. AdequateDL is a project aiming to explore how approximations can improve performance and energy efficiency of hardware accelerators in DL applications. This paper presents the main concepts and techniques related to approximation of CNNs and preliminary results obtained in the AdequateDL framework.
Olivier Sentieys, Silviu-Ioan Filip, David Briand, David Novo, Etienne Dupuis, Ian O'Connor, Alberto Bosio
DDECS1
2021 Real-Time Imprecise Computation Tasks Mapping for DVFS-Enabled Networked Systems
abstract
Networked systems are useful for a wide range of applications, many of which require distributed and collaborative data processing to satisfy real-time requirements. On one hand, networked systems are usually resource constrained, mainly regarding the energy supply of the nodes and their computation and communication abilities. On the other hand, many real-time applications can be executed in an imprecise way, where an approximate result is acceptable as long as the baseline Quality of Service (QoS) is satisfied. Such applications can be modeled through imprecise computation (IC) tasks. To achieve a better tradeoff between QoS and limited system resources, while meeting application requirements, the IC-tasks must be efficiently mapped to the system nodes. To tackle this problem, we first construct an IC-task mapping problem that aims to maximize system QoS subject to real-time and energy constraints. Dynamic voltage and frequency scaling (DVFS) and multipath routing are explored to further enhance real-time performance and reduce energy consumption. Second, based on the problem structure, we propose an optimal approach to perform IC-task mapping and prove its optimality. Furthermore, to enhance the scalability of the proposed approach, we present a heuristic IC-task mapping method with low computation time. Finally, the simulation results demonstrate the effectiveness of the proposed methods in terms of the solution quality and the computation time.
Lei Mo, Angeliki Kritikakou, Olivier Sentieys, Xianghui Cao
IEEE Internet Things J.3
2021 Freezer: A Specialized NVM Backup Controller for Intermittently Powered Systems
abstract
The explosion of IoT and wearable devices determined a rising attention toward energy harvesting as source for powering these systems. In this context, many applications cannot afford the presence of a battery because of size, weight, and cost issues. Therefore, due to the intermittent nature of ambient energy sources, these systems must be able to save and restore their state, in order to guarantee progress across power interruptions. In this work, we propose a specialized backup/restore controller that dynamically tracks the memory accesses during the execution of the program. The controller then commits the changes to a snapshot in a nonvolatile memory (NVM) when a power failure is detected. Our approach does not require complex hybrid memories and can be implemented with standard components. Results on a set of benchmarks show an average 8× reduction in backup size. Thanks to our dedicated controller, the backup time is further reduced by more than 100×, with an area and power overhead of only 0.4% and 0.8%, respectively, with respect to a low-end IoT node.
Davide Pala, Ivan Miro Panades, Olivier Sentieys
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Towards Generic and Scalable Word-Length Optimization
abstract
In this paper, we propose a method to improve the scalability of Word-Length Optimization (WLO) for large applications that use complex quality metrics such as Structural Similarity (SSIM). The input application is decomposed into smaller kernels to avoid uncontrolled explosion of the exploration time, which is known as noise budgeting. The main challenge addressed in this paper is how to allocate noise budgets to each kernel. This requires capturing the interactions across kernels. The main idea is to characterize the impact of approximating each kernel on accuracy/cost through simulation and regression. Our approach improves the scalability while finding better solutions for Image Signal Processor pipeline.
Van-Phu Ha, Tomofumi Yuki, Olivier Sentieys
DATE3
2019 Error Analysis of the Square Root Operation for the Purpose of Precision Tuning: A Case Study on K-means
abstract
In this paper, we propose an analytical approach to study the impact of floating point (FLP) precision variation on the square root operation, in terms of computational accuracy and performance gain. We estimate the round-off error resulting from reduced precision. We also inspect the Newton Raphson algorithm used to approximate the square root in order to bound the error caused by algorithmic deviation. Consequently, the implementation of the square root can be optimized by fittingly adjusting its number of iterations with respect to any given FLP precision specification, without the need for long simulation times. We evaluate our error analysis of the square root operation as part of approximating a classic data clustering algorithm known as K-means, for the purpose of reducing its energy footprint. We compare the resulting inexact K-means to its exact counterpart, in the context of color quantization, in terms of energy gain and quality of the output. The experimental results show that energy savings could be achieved without penalizing the quality of the output (e.g., up to 41.87% of energy gain for an output quality, measured using structural similarity, within a range of [0.95,1]).
Oumaima Matoussi, Yves Durand, Olivier Sentieys, Anca Mariana Molnos
ASAP3
2019 Accelerating Itemset Sampling using Satisfiability Constraints on FPGA
abstract
Finding recurrent patterns within a data stream is important for fields as diverse as cybersecurity or e-commerce. This requires to use pattern mining techniques. However, pattern mining suffers from two issues. The first one, known as "pattern explosion", comes from the large combinatorial space explored and is the result of too many patterns outputed to be analyzed. Recent techniques called output space sampling solve this problem by outputing only a sampled set of all the results, with a target size provided by the user. The second issue is that most algorithms are designed to operate on static datasets or low throughput streams. In this paper, we propose a contribution to tackle both issues, by designing an FPGA accelerator for pattern mining with output space sampling. We show that our accelerator can outperform a state-of-the-art implementation on a server class CPU using a modest FPGA product.
Mael Gueguen, Olivier Sentieys, Alexandre Termier
DATE2
2019 Approximation-aware Task Deployment on Asymmetric Multicore Processors
abstract
Asymmetric Multicore Processors (AMP) are a very promising architecture to deal efficiently with the wide diversity of applications. In real-time application domains, in-time approximated results are preferred than accurate - but too late - results. In this work, we propose a deployment approach that exploits the heterogeneity provided by AMP architectures and the approximation tolerance provided by the applications, so as to increase as much as possible the quality of the results under given energy and timing constraints. Initially, an optimal approach is proposed based on problem linearization and decomposition. Then, a heuristic approach is developed based on iteration relaxation of the optimal version. The obtained results show 16.3% reduction in the computation time for the optimal approach compared to the conventional optimal approaches. The proposed heuristic approach is about 100 times faster at the cost of a 29.8% QoS degradation in comparison with the optimal solution.
Lei Mo, Angeliki Kritikakou, Olivier Sentieys
DATE3
2019 Fine-Grained Hardware Mitigation for Multiple Long-Duration Transients on VLIW Function Units
abstract
Technology scaling makes hardware more susceptible to radiation, which can cause multiple transient faults with long duration. In these cases, the affected function unit is usually considered as faulty and is not further used. To reduce this performance degradation, the proposed hardware mechanism detects the faults that are still active during execution and reschedules the instructions to use the fault-free components of the affected function units. The results show multiple long-duration fault mitigation with low performance, area, and power overhead.
Rafail Psiakis, Angeliki Kritikakou, Olivier Sentieys
DATE3
2019 What You Simulate Is What You Synthesize: Designing a Processor Core from C++ Specifications
abstract
The following topics are dealt with: learning (artificial intelligence); neural nets; field programmable gate arrays; integrated circuit design; logic design; optimisation; network routing; low-power electronics; cryptography; electronic design automation.
Simon Rokicki, Davide Pala, Joseph Paturel, Olivier Sentieys
ICCAD4
2019 Multi-carrier spread-spectrum transceiver for WiNoC
abstract
In this paper, we propose a low-power, high-speed, multi-carrier reconfigurable transceiver based on Frequency Division Multiplexing to ensure data transfer in future Wireless NoCs. The proposed transceiver supports a medium access control method to sustain unicast, broadcast and multicast communication patterns, providing dynamic data exchange among wireless nodes. The proposed transceiver designed using a 28-nm FDSOI technology consumes only 2.37 mW and 4.82 mW in unicast/broadcast and multicast modes, respectively, with an area footprint of 0.0138 mm2.
Joel Ortiz Sosa, Olivier Sentieys, Christian Roland, Cédric Killian
NOCS2
2018 Zyggie: A Wireless Body Area Network platform for indoor positioning and motion tracking
abstract
Nowadays, there is a high demand for human and/or objects monitoring/localizing in the context of applications like Building Information Modeling (BIM), automated drone missions, contextual visits of museum or sports monitoring for instance. While for outdoor positioning accurate and robust solutions (i.e. GPS) exist for many years, indoor positioning is still very challenging. There is also a need of gesture/motion tracking systems that could replace video solutions. We propose in this paper a hardware/software platform named Zyggie that combines both Ultra Wide Band (UWB) technology and Received Signal Strength Indicator (RSSI) for low power accurate indoor positioning and Inertial Measurement Unit (IMU) utilization for motion tracking. Very few industrial/academic existing solutions can simultaneously perform indoor positioning and motion tracking and none of them can do both under low power, low cost and compacity constraints addressed by our platform. As Zyggie has the capability to estimate distances w.r.t other platforms in the environment and quaternions (which represents the attitude/orientation) users can test/enhance state of the art algorithms for positioning and motion tracking applications.
Antoine Courtay, Mickaël Le Gentil, Olivier Berder, Arnaud Carer, Pascal Scalart, Olivier Sentieys
ISCAS6
2018 A Diversity Scheme to Enhance the Reliability of Wireless NoC in Multipath Channel Environment
abstract
Wireless Network-on-Chip (WiNoC) is one of the most promising solutions to overcome multi-hop latency and high power consumption of modern many/multi core System-on-Chip (SoC). However, the design of efficient wireless links faces challenges to overcome multi-path propagation present in realistic WiNoC channels. In order to alleviate such channel effect, this paper presents a Time-Diversity Scheme (TDS) to enhance the reliability of on-chip wireless links using a semi-realistic channel model. First, we study the significant performance degradation of state-of-the-art wireless transceivers subject to different levels of multi-path propagation. Then we investigate the impact of using some channel correction techniques adopting standard performance metrics. Experimental results show that the proposed Time-Diversity Scheme significantly improves Bit Error Rate (BER) compared to other techniques. Moreover, our TDS allows for wireless communication links to be established in conditions where this would be impossible for standard transceiver architectures. Results on the proposed complete transceiver, designed using a 28-nm FDSOI technology, show a power consumption of 0.63mW at 1.0V and an area of 317 μm2. Full channel correction is performed in one single clock cycle.
Joel Ortiz Sosa, Olivier Sentieys, Christian Roland
NOCS2
2018 Offline Optimization of Wavelength Allocation and Laser Power in Nanophotonic Interconnects
abstract
Optical Network-on-Chip (ONoC) is a promising communication medium for large-scale multiprocessor systems-on-chips. Indeed, ONoC can outperform classical electrical NoCs in terms of energy efficiency and bandwidth density, in particular, because this medium can support multiple transactions at the same time on different wavelengths by using Wavelength Division Multiplexing (WDM). However, multiple signals sharing simultaneously the same part of a waveguide can lead to inter-channel crosstalk noise. This problem impacts the signal-to-noise ratio of the optical signals, which leads to an increase in the Bit Error Rate (BER) at the receiver side. If a specific BER is targeted, an increase of laser power should be necessary to satisfy the SNR. In this context, an important issue is to evaluate the laser power needed to satisfy the various desired communication bandwidths based on the BER performance requirements. In this article, we propose an off-line approach that concurrently optimizes the laser power scaling and execution time of a global application. A set of different levels of power is introduced for each laser, to ensure that optical signals can be emitted with just-enough power to ensure targeted BER. As a result, most promising solutions are highlighted for mapping a defined application onto a 16-core ring-based WDM ONoC.
Jiating Luo, Cédric Killian, Sébastien Le Beux, Daniel Chillet, Olivier Sentieys, Ian O'Connor
ACM J. Emerg. Technol. Comput. Syst.5
2018 Energy-Quality-Time Optimized Task Mapping on DVFS-Enabled Multicores
abstract
Multicore architectures have great potential for energy-constrained embedded systems, such as energy-harvesting wireless sensor networks. Some embedded applications, especially the real-time ones, can be modeled as imprecise computation tasks. A task is divided into a mandatory subtask that provides a baseline quality-of-service (QoS) and an optional subtask that refines the result to increase the QoS. Combining dynamic voltage and frequency scaling, task allocation, and task adjustment, we can maximize the system QoS under real-time and energy supply constraints. However, the nonlinear and combinatorial nature of this problem makes it difficult to solve. This paper first formulates a mixed-integer nonlinear programming problem to concurrently carry out task-to-processor allocation, frequency-to-task assignment and optional task adjustment. We provide a mixed-integer linear programming form of this formulation without performance degradation and we propose a novel decomposition algorithm to provide an optimal solution with reduced computation time compared to state-of-the-art optimal approaches (22.6% in average). We also propose a heuristic version that has negligible computation time.
Lei Mo, Angeliki Kritikakou, Olivier Sentieys
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Energy and Performance Trade-off in Nanophotonic Interconnects using Coding Techniques
abstract
Nanophotonic is an emerging technology considered as one of the key solutions for future generation on-chip interconnects. Indeed, this technology provides high bandwidth for data transfers and can be a very interesting alternative to bypass the bottleneck induced by classical NoC. However, their implementation in fully integrated 3D circuits remains uncertain due to the high power consumption of on-chip lasers. However, if a specific bit error rate is targeted, digital processing can be added in the electrical domain to reduce the laser power and keep the same communication reliability. This paper addresses this problem and proposesto transmit encoded data on the optical interconnect, which allows for a reduction of the laser power consumption, thus increasing nanophotonics interconnects energy efficiency. The results presented in this paper show that using simple Hamming coder and decoder permits to reduce the laser power by nearly 50% without loss in communication data rate and with a negligible hardware overhead.
Cédric Killian, Daniel Chillet, Sébastien Le Beux, Van-Dung Pham, Olivier Sentieys, Ian O'Connor
DAC5
2017 The hidden cost of functional approximation against careful data sizing - A case study
abstract
Many applications are error-resilient, allowing for the introduction of approximations in the calculations, as long as a certain accuracy target is met. Traditionally, fixed-point arithmetic is used to relax accuracy, by optimizing the bit-width. This arithmetic leads to important benefits in terms of delay, power and area. Lately, several hardware approximate operators were invented, seeking the same performance benefits. However, a fair comparison between the usage of this new class of operators and classical fixed-point arithmetic with careful truncation or rounding, has never been performed. In this paper, we first compare approximate and fixed-point arithmetic operators in terms of power, area and delay, as well as in terms of induced error, using many state-of-the-art metrics and by emphasizing the issue of data sizing. To perform this analysis, we developed a design exploration framework, APXPERF, which guarantees that all operators are compared using the same operating conditions. Moreover, operators are compared in several classical real-life applications leveraging relevant metrics. In this paper, we show that considering a large set of parameters, existing approximate adders and multipliers tend to be dominated by truncated or rounded fixed-point ones. For a given accuracy level and when considering the whole computation data-path, fixed-point operators are several orders of magnitude more accurate while spending less energy to execute the application. A conclusion of this study is that the entropy of careful sizing is always lower than approximate operators, since it require significantly less bits to be processed in the data-path and stored. Approximated data therefore always contain on average a greater amount of costly erroneous, useless information.
Benjamin Barrois, Olivier Sentieys, Daniel Ménard
DATE2
2017 Performance and energy aware wavelength allocation on ring-based WDM 3D optical NoC
abstract
Optical Network-on-Chip (ONoC) is a promising communication medium for large-scale Multiprocessor System on Chip (MPSoC). ONoC outperforms classical electrical NoC in terms of throughput and latency. The medium can support multiple transactions at the same time on different wavelengths by using Wavelength Division Multiplexing (WDM). Moreover multiple wavelengths can be used as high-bandwidth channel to reduce transmission time. However, multiple signals sharing simultaneously a waveguide can lead to inter-channel crosstalk noise. This problem impacts the Signal to Noise Ratio (SNR) of the optical signal, which leads to an increase in the Bit Error Rate (BER) at the receiver side. In this paper we first formulate the crosstalk noise and execution time models and then propose a Wavelength Allocation (WA) method in a ring-based WDM ONoC allowing to search for performance and energy trade-offs, based on the application constraints. As result, most promising WA solutions are highlighted for a defined application mapping onto 16-core WDM ONoC.
Jiating Luo, A. Elantably, Van-Dung Pham, Cédric Killian, Daniel Chillet, Sébastien Le Beux, Olivier Sentieys, Ian O'Connor
DATE7
2017 Pushing the limits of voltage over-scaling for error-resilient applications
abstract
Voltage scaling has been used as a prominent technique to improve energy efficiency in digital systems, scaling down supply voltage effects in quadratic reduction in energy consumption of the system. Reducing supply voltage induces timing errors in the system that are corrected through additional error detection and correction circuits. In this paper we are proposing voltage over-scaling based approximate operators for applications that can tolerate errors. We characterize the basic arithmetic operators using different operating triads (combination of supply voltage, body-biasing scheme and clock frequency) to generate models for approximate operators. Error-resilient applications can be mapped with the generated approximate operator models to achieve optimum trade-off between energy efficiency and error margin. Based on the dynamic speculation technique, best possible operating triad is chosen at runtime based on the user definable error tolerance margin of the application. In our experiments in 28nm FDSOI, we achieve maximum energy efficiency of 89% for basic operators like 8-bit and 16-bit adders at the cost of 20% Bit Error Rate (ratio of faulty bits over total bits) by operating them in near-threshold regime.
Rengarajan Ragavan, Benjamin Barrois, Cédric Killian, Olivier Sentieys
DATE4
2017 Fast and Energy-Driven Design Space Exploration for Heterogeneous Architectures
abstract
In the last years, the integration of specialized hardware accelerators in Multiprocessor System-on-Chip (MpSoC) led to a new kind of architectures combining both software (SW) and hardware (HW) computational resources. For these new Heterogeneous MpSoC (HMpSoC) architectures, performance and energy consumption depend on a large set of parameters such as the HW/SW partitioning, the type of HW implementation or the communication cost. Design Space Exploration (DSE) consists in adjusting these parameters while monitoring a set of metrics (execution time, power, energy efficiency) to find the best mapping of the application on the targeted architecture. With the shift from performance-aware to energy-aware designs, computer-aided design and development tools try to reduce the large design space by simplifying HW/SW mapping mechanisms. However, energy consumption is not well supported in most of DSE tools due to the difficulty to fast and accurately estimate the energy consumption. To this aim, this work introduces a DSE method based on an analytical power model to circumvent the computation time bottleneck of state-of-the-art DSE methods. This exploration method proposes to optimize the HW/SW partitioning and mapping under user-defined objectives, especially an energy constraint. It targets tiling-based parallel applications and relies on an analytical power model that provides the DSE framework with the execution time and energy of a HW/SW configuration. The power model parameters are obtained with the measurements of a tiny subset of the design space, which are then injected into two extraction functions to obtain analytical formulations of the execution time and the energy consumption of the computation kernel. The partitioning problem constraints are defined as a set of inequalities with Boolean, integer (discrete) and non-integer (continuous) variables within a Mixed Integer Linear Programming (MILP) framework. Then, the best configuration that minimizes the user objective (e.g. execution time or total energy consumption) can be efficiently determined using commercial or open source solvers within a second. This methodology was tested on a Zynq-based heterogeneous architecture with two application kernels: a matrix multiplication and a Stencil computation. The results show a minimum of 12% acceleration speed-up and energy saving compared to standard approaches. They also show that the most energy-efficient solution is application-and platform-dependent and moreover hardly predictable. Such method could be included in a complete framework with a multi-step exploration to obtain an energy-efficient mapping of a full application on HMpSoC and to open new opportunity for future computer-aided design tools.
Baptiste Roux, Matthieu Gautier, Olivier Sentieys, Jean-Philippe Delahaye
FCCM3
2017 Decomposed Task Mapping to Maximize QoS in Energy-Constrained Real-Time Multicores
abstract
Multicore architectures are now widely used in energy-constrained real-time systems, such as energy-harvesting wireless sensor networks. To take advantage of these multicores, there is a strong need to balance system energy, performance and Quality-of-Service (QoS). The Imprecise Computation (IC) model splits a task into mandatory and optional parts allowing to tradeoff QoS. The problem of mapping, i.e. allocating and scheduling, IC-tasks to a set of processors to maximize system QoS under real-time and energy constraints can be formulated as a Mixed Integer Linear Programming (MILP) problem. However, state-of-the-art solving techniques either demand high complexity or can only achieve feasible (suboptimal) solutions. In this paper, we develop an effective decomposition-based approach to achieve an optimal solution while reducing computational complexity. It decomposes the original problem into two smaller easier-to-solve problems: a master problem for IC-tasks allocation and a slave problem for IC-tasks scheduling. We also provide comprehensive optimality analysis for the proposed method. Through the simulations, we validate and demonstrate the performance of the proposed method, resulting in an average 55% QoS improvement with regards to published techniques.
Lei Mo, Angeliki Kritikakou, Olivier Sentieys
ICCD3
2017 Taking advantage of correlation in stochastic computing
abstract
In recent years, shrinking size in integrated circuits has imposed a big challenge in maintaining the reliability in conventional computing. Stochastic computing has been seen as a reliable, low-cost, and low-power alternative to overcome such issues. Stochastic Computing (SC) computes data in the form of bit streams of 1s and 0s. Therefore, SC outperforms conventional computing in terms of tolerance to soft error and uncertainty at the cost of increased computational time. Stochastic Computing with uncorrelated input streams requires streams to be highly independent for better accuracy. This results in more hardware consumption for conversion of binary numbers to stochastic streams. Correlation can be used to design Stochastic Computation Elements (SCE) with correlated input streams. These designs have higher accuracy and less hardware consumption. In this paper, we propose new SC designs to implement image processing algorithms with correlated input streams. Experimental results of proposed SC with correlated input streams show on average 37% improvement in accuracy with reduction of 50-90% in area and 20-85% in delay over existing stochastic designs.
Rahul Kumar Budhwani, Rengarajan Ragavan, Olivier Sentieys
ISCAS3
2016 Leveraging power spectral density for scalable system-level accuracy evaluation
Benjamin Barrois, Karthick Parashar, Olivier Sentieys
DATE3
2016 Effects of I/O routing through column interfaces in embedded FPGA fabrics
abstract
The emergence of 2.5D and 3D packaging technologies enables the integration of FPGA dice into more complex systems. Both heterogeneous manycore designs, which include an FPGA layer, and interposer-based multi-FPGA systems support the inclusion of reconfigurable hardware in 3D-stacked integrated circuits. In these architectures, the communication between FPGA dice or between FPGA and fixed-function layers often takes place through dedicated communication interfaces spread over the FPGA logic fabric, as opposed to an I/O ring around the fabric. In this paper, we investigate the effect of organizing FPGA fabric I/O into coarse-grained interface blocks distributed throughout the FPGA fabric. Specifically, we consider the quality of results for the placement and routing phases of the FPGA physical design flow. We evaluate the routing of I/O signals of large applications through dedicated interface blocks at various granularities in the logic fabric, and study its implications on the critical path delay of routed designs. We show that the impact of such I/O routing is limited and can improve chip routability and circuit delay in many cases.
Christophe Huriaux, Olivier Sentieys, Russell Tessier
FPL2
2016 Blind adaptive transmitter IQ imbalance compensation in M-QAM optical coherent systems
abstract
Blind adaptive source separation (BASS) based compensation for transmitter (Tx) IQ imbalance is presented for the first time in an M-QAM optical coherent system. The proposed method is numerically investigated with 4-QAM and 16-QAM signals in the presence of Tx IQ imbalance up to 30o. The robustness of the BASS method is studied after 200-km optical fiber transmission, in which the effects of chromatic dispersion (CD) and carrier frequency offset (CFO) are assumed to be dominant. It is also found that CFO, inherent to frequency difference between the transmitter and receiver lasers in optical coherent transmission, should be compensated before IQ imbalance compensation to achieve a better performance. The proposed method outperforms the Gram-Schmidt orthogonalization procedure (GSOP) in the presence of CD and CFO. We further validate experimentally the proposed method with 10-Gbaud optical 4-QAM and 16-QAM signals at 30o and 10o phase imbalance, respectively, with an emulated 200-km optical fiber transmission and 200-MHz CFO. More specifically, the optical signal-to-noise ratio (OSNR) penalty reduction of the BASS method compared to the GSOP method is 1 dB for 4-QAM at a bit-error-ratio (BER) of 2×10-3 and 2 dB for 16-QAM at a BER of 10-3. Moreover, instead of being a fully independent block and requiring statistical estimation as in GSOP, the BASS method can be integrated into an equalizer and operated at the sample rate, simplifying the operation and allowing parallel implementation.
Trung-Hien Nguyen, Pascal Scalart, Mathilde Gay, Laurent Bramerie, Christophe Peucheret, Ti Nguyen-Ti, Matthieu Gautier, Olivier Sentieys, Jean-Claude Simon, Michel Joindot
ICC8
2015 Low-complexity energy proportional posture/gesture recognition based on WBSN
abstract
This paper addresses the issue of low-power posture and gesture recognition in indoor or outdoor environments without any additional equipment. For applications based on predefined postures such as environment control and physical rehabilitation, we show that low cost and fully distributed solutions, that minimize radio communications, can be efficiently implemented. Considering that radio links provide distance information, we also demonstrate that the matrix of estimated inter-node distances offers complementary information that allows for the reduction of communication load. Our results are based on a simulator that can handle various measured input data, different algorithms and various noise models. Simulation results are useful and used for the development of real-life prototype.
Alexis Aulery, Jean-Philippe Diguet, Christian Roland, Olivier Sentieys
BSN4
2015 Design flow and run-time management for compressed FPGA configurations
Christophe Huriaux, Antoine Courtay, Olivier Sentieys
DATE3
2015 Joint simple blind IQ imbalance compensation and adaptive equalization for 16-QAM optical communications
abstract
We present a novel simple blind adaptive compensation method for in-phase/quadrature (IQ) imbalance in m-ary quadrature amplitude modulation (m-QAM) coherent optical fiber communication systems. IQ-imbalance compensation is integrated into butterfly-structured finite impulse response (FIR) filters, resulting in a significant computational effort reduction in comparison to conventional methods. A reduction in hardware complexity by a factor of about 3 is achieved by the proposed joint method. The proposed structure is experimentally validated with a 40-Gbit/s 16-QAM signal. A 7-dB power penalty reduction is experimentally achieved at a bit error rate (BER) of 10-3in the presence of a 10° phase imbalance, confirming the effectiveness of the proposed algorithm. The equalization capability remains even in the presence of group velocity dispersion along the link, which is numerically confirmed with optical fiber transmission up to 1200 km and 20° phase imbalance.
Trung-Hien Nguyen, Pascal Scalart, Michel Joindot, Mathilde Gay, Laurent Bramerie, Christophe Peucheret, Arnaud Carer, Jean-Claude Simon, Olivier Sentieys
ICC9
2015 Radio signature based posture recognition using WBSN
abstract
A body network of Inertial Measurement Units (IMUs) is a well known solution for posture recognition based on accelerometer and magnetometer data fusion. However sensors and especially the magnetometer can be disturbed by the environment. Considering a Wireless Body Sensor Network (WBSN), we propose to use available radio received power measurements as an alternative to the magnetometer. We show with simulation and real data, that the radio signal used for WBSN communications can also provide useful location information despite highly noisy Received Signal Strength Indications (RSSI). We propose a solution for the static case that leads to a very simple yet Efficient algorithm.
Alexis Aulery, Christian Roland, Jean-Philippe Diguet, Zhongwei Zheng, Olivier Sentieys, Pascal Scalart
IPSN5
2015 Energy-Neutral Design Framework for Supercapacitor-Based Autonomous Wireless Sensor Networks
abstract
To design autonomous wireless sensor networks (WSNs) with a theoretical infinite lifetime, energy harvesting (EH) techniques have been recently considered as promising approaches. Ambient sources can provide everlasting additional energy for WSN nodes and exclude their dependence on battery. In this article, an efficient energy harvesting system which is compatible with various environmental sources, such as light, heat, or wind energy, is proposed. Our platform takes advantage of double-level capacitors not only to prolong system lifetime but also to enable robust booting from the exhausting energy of the system. Simulations and experiments show that our multiple-energy-sources converter (MESC) can achive booting time in order of seconds. Although capacitors have virtual recharge cycles, they suffer higher leakage compared to rechargeable batteries. Increasing their size can decrease the system performance due to leakage energy. Therefore, an energy-neutral design framework providing a methodology to determine the minimum size of those storage devices satisfying energy-neutral operation (ENO) and maximizing system quality-of-service (QoS) in EH nodes, when using a given energy source, is proposed. Experiments validating this framework are performed on a real WSN platform with both photovoltaic cells and thermal generators in an indoor environment. Moreover, simulations on OMNET++ show that the energy storage optimized from our design framework is utilized up to 93.86%.
Trong Nhan Le, Alain Pegatoquet, Olivier Berder, Olivier Sentieys, Arnaud Carer
ACM J. Emerg. Technol. Comput. Syst.4
2014 Design Space Exploration in an FPGA-Based Software Defined Radio
abstract
The FPGA (Field Programmable Gate Array) technology is expected to play a key role in the development of Software Defined Radio (SDR) platforms. To this aim, leveraging the nascent High-Level Synthesis (HLS) tools, a design flow from high-level specifications to Register-Transfer Level (RTL) description can be thought. Based on such a flow, this paper describes the Design Space Exploration (DSE) that can be achieved using loop optimizations. The mainstream objective is to demonstrate the compile-time flexibility of an architecture when associated with a reconfigurable platform. Throughout both IEEE 802.15.4 and IEEE 802.11g waveform examples, we show how the FPGA resources can be tuned according to a targeted throughput.
Matthieu Gautier, Ganda Stéphane Ouedraogo, Olivier Sentieys
DSD3
2014 FPGA Architecture Enhancements to Support Heterogeneous Partially Reconfigurable Regions
abstract
In this work the author develop an FPGA architecture which allows for the placement of a partial FPGA design on the logic fabric even if the relative placement of heterogeneous blocks within the target region is not identical to the placement used to generate the bitstream for the partial design. This work has been conducted in the context of the European FP7 FlexTiles project in which a dynamically reconfigurable logic fabric is embedded in a 3-D stacked chip along with a manycore architecture. The reconfigurable logic fabric is used to load hardware-accelerated functions whose use is scheduled at run time. All communication between the fabric and manycore is made via dedicated I/O interface blocks in the fabric. This communication configuration increases the need for a flexible architecture which can handle the placement of a single application bitstream in multiple locations on the logic fabric.
Christophe Huriaux, Olivier Sentieys, Russell Tessier
FCCM2
2014 Low Power Reconfigurable Controllers for Wireless Sensor Network Nodes
abstract
A key concern in the design of controllers in wireless sensor network (WSN) nodes is the flexibility to execute different control tasks involving sensing, communications and computational resources of the node. In this paper, low power flexible controllers for WSN nodes based on reconfigurable microtasks composed of an FSM and datapath are presented. Coarse grain power gating opportunities are exploited in FSM and datapath for low power operation in reconfigurable microtasks. Power estimation results on typical benchmark microtasks show a 2× to 5× improvement in energy efficiency w.r.t a microcontroller at a cost of 5× relative to a microtask implemented as an ASIC with higher NRE costs.
Vivek D. Tovinakere, Olivier Sentieys, Steven Derrien, Christophe Huriaux
FCCM2
2014 FPGA architecture support for heterogeneous, relocatable partial bitstreams
abstract
The use of partial dynamic reconfiguration in FPGA-based systems has grown in recent years as the spectrum of applications which use this feature has increased. For these systems, it is desirable to create a series of partial bitstreams which represent tasks which can be located in multiple regions in the FPGA fabric. While the transferal of homogeneous collections of lookup-table based logic blocks from region to region has been shown to be relatively straightforward, it is more difficult to transfer partial bitstreams which contain fixed-function resources, such as block RAMs and DSP blocks. In this paper we consider FPGA architecture enhancements which allow for the migration of partial bitstreams including fixed-function resources from region to region even if these resources are not located in the same position in each region. Our approach does not require significant, time-consuming place-and-route during the migration process. We quantify the cost of inserting additional routing resources into the FPGA architecture to allow for easy migration of heterogeneous, fixed-function resources. Our experiments show that this flexibility can be added for a relatively low overhead and performance penalty.
Christophe Huriaux, Olivier Sentieys, Russell Tessier
FPL2
2014 Toward scalable source level accuracy analysis for floating-point to fixed-point conversion
abstract
In embedded systems, many numerical algorithms are implemented with fixed-point arithmetic to meet area cost and power constraints. Fixed-point encoding decisions can significantly affect cost and performance. To evaluate their impact on accuracy, designers resort to simulations. Their high running-time prevents thorough exploration of the design-space. To address this issue, analytical modeling techniques have been proposed, but their applicability is limited by scalability issues. In this paper, we extend these techniques to a larger class of programs. We use polyhedral methods to extract a more compact, graph-based representation of the program. We validate our approach with a several image and signal processing algorithms.
Gaël Deest, Tomofumi Yuki, Olivier Sentieys, Steven Derrien
ICCAD3
2014 RIC-MAC: A MAC protocol for low-power cooperative wireless sensor networks
abstract
In this study, a receiver initiated cooperative medium access control (RIC-MAC) protocol is proposed for cooperative communications to reduce the energy consumption of wireless sensor networks (WSNs). Considering a real WSN platform, the simulation results show that using the proposed RIC-MAC protocol in cooperative communications provides latency and energy gains as compared to multi-hop communications. However, the energy gain is shown to be reduced when the network traffic load increases. Finally, considering the impact of traffic load on energy consumption and latency, RIC-MAC is illustrated to be robust to traffic load variations in terms of latency.
Le Quang Vinh Tran, Olivier Berder, Olivier Sentieys
WCNC3
2014 Accelerated Performance Evaluation of Fixed-Point Systems With Un-Smooth Operations
abstract
The problem of accuracy evaluation is one of the most time consuming tasks during the fixed-point refinement process. Analytical techniques based on perturbation theory have been proposed in order to overcome the need for long fixed-point simulation. However, these techniques are not applicable in the presence of certain operations classified as un-smooth operations. In such circumstances, fixed-point simulation should be used. In this paper, an algorithm detailing the hybrid technique which makes use of an analytical accuracy evaluation technique used to accelerate fixed-point simulation is presented. This technique is applicable to signal processing systems with both feed-forward and feedback interconnect topology between its operations. The acceleration obtained as a result of applications of the proposed technique is consistent with fixed-point simulation, while reducing the time taken for fixed-point simulation by several orders of magnitude.
Karthick Parashar, Daniel Ménard, Olivier Sentieys
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Component-Level Datapath Merging in System-Level Design of Wireless Sensor Node Controllers for FPGA-Based Implementations
abstract
Wireless Sensor Networks (WSNs) are relatively new and challenging research area for embedded design automation. Engineering a WSN node hardware is a difficult job as the design must satisfy several constraints. Among these constraints, overall energy consumption and node size, are the two most significant constraints. WSN node platforms have until recently been designed using off-the-shelf low-power microprocessors (MCUs), even though energy profile of these MCUs is not suitable for ultra low-power sensor nodes. On the other hand, WSN-specific hardware accelerators have also been proposed that have excellent energy profile but lack in flexibility, need higher design efforts and have huge non-recurring engineering (NRE) costs. In this work, we propose an automated system level design flow for an intermediate approach, based on the concept of data path merging (DPM) where several hardware accelerators (called micro-tasks) share a common customized data path, to have an improvement in flexibility and silicon area with possible increase in dynamic power consumption for the control/processing part of the sensor node targeted for field programmable gate array (FPGA)-based implementation. Our experiments show that component-level DPM yields to savings from 20%, to 75% for various FPGA resources like I/O ports, area for combinational and sequential logic, and static power consumption.
Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys
DSD3
2013 A low-latency and energy-efficient MAC protocol for cooperative wireless sensor networks
abstract
In this paper, we investigate a novel low-latency MAC protocol (ARQ-CRI) for low-power cooperative wireless sensor networks WSNs, while preserving (i.e. in high traffic mode) or even increasing (i.e. in low traffic mode) energy-efficiency. Based on the wakeup period of relay nodes, a new relay selection technique is proposed to decrease latency. Simulations show that ARQ-CRI can get latency gain and help to increase energy-efficiency in low traffic mode as compared to CL-MAC, a recently proposed cooperative MAC protocol for WSNs. Besides, regarding traffic variation, ARQ-CRI is claimed to be more stable than CL-MAC. Finally, the effect of the number of potential relays on the performance of ARQ-CRI is also considered.
Duc-Long Nguyen, Le Quang Vinh Tran, Olivier Berder, Olivier Sentieys
GLOBECOM4
2013 A block-parallel architecture for initial and fine synchronization in OFDM systems
abstract
A novel low complexity parallel algorithm and its associated architecture are proposed for initial synchronization in orthogonal frequency division multiplexing (OFDM) systems. The method is hierarchical and uses auto-correlation for the first step and cross-correlation for the second step. The main advantage of the proposed approach is that it reduces the computational complexity by a factor of five (80%), while achieving similar mean square error (MSE) as cross-correlation based methods. The method uses block-level parallelism for auto-correlation step, which speeds up the computation significantly. After fixed-point analysis, a parallel architecture is proposed to accelerate both coarse and fine synchronization steps. This parallel architecture is scalable and provides speed-up proportional to number of parallel blocks.
Pramod P. Udupa, Olivier Sentieys, Pascal Scalart
ICC2
2013 A polynomial time algorithm for solving the word-length optimization problem
abstract
Trading off accuracy to the system costs is popularly addressed as the word-length optimization (WLO) problem. Owing to its NP-hard nature, this problem is solved using combinatorial heuristics. In this paper, a novel approach is taken by relaxing the integer constraints on the optimization variables and obtain an alternate noise-budgeting problem. This approach uses the quantization noise power introduced into the system due to fixed-point word-lengths as optimization variables instead of using the actual integer valued fixed-point word-lengths. The noise-budgeting problem is proved to be convex in the rounding mode quantization case and can therefore be solved using analytical convex optimization solvers. An algorithm with linear time complexity is provided in order to realize the actual fixed-point word-lengths from the noise budgets obtained by solving the convex noise-budgeting problem.
Karthick Parashar, Daniel Ménard, Olivier Sentieys
ICCAD3
2013 Duty-cycle power manager for thermal-powered Wireless Sensor Networks
abstract
Exploiting energy from the environment to extend the system lifetime of Wireless Sensor Network (WSN), especially thermal energy, is considered as a promising approach. When considering self-powered systems, the Power Manager (PM) plays an important role in energy harvesting WSNs. Instead of minimizing the consumed energy as in the case of battery-powered systems, it causes the harvesting node to converge to Energy Neutral Operation (ENO) in order to achieve a theoretically infinite lifetime. In this paper, a low complexity PM for a thermal-powered WSN is presented. Our PM adapts the duty cycle of the node according to the estimation of harvested energy and the consumed energy provided by a simple energy monitor for a super capacitor based WSN to achieve the ENO. Experiments are performed on a real WSN platform where harvested energy is extracted from the wasted heat of a PC adapter by two thermoelectric generators.
Trong Nhan Le, Alain Pegatoquet, Olivier Sentieys, Olivier Berder, Cécile Belleudy
PIMRC3
2013 Energy efficient reservation-based opportunistic MAC scheme in multi-hop networks
abstract
Opportunistic forward techniques can improve the energy efficiency of wireless multi-hop networks by exploiting multiuser diversity. Opportunistic forward schemes generally integrate the functions of routing and MAC layers; nevertheless, most of the works focus on the design of routing protocol while assuming an energy efficient MAC layer. However, due to the unreliable links among the source node and its relay candidates, a great amount of energy expenditure results from the synchronization and multi-relay transmission, which degrades the energy performance of opportunistic forward techniques. In this paper, an energy efficient opportunistic MAC protocol with the mechanisms of reservation and a relay candidate coordination is proposed. Moreover, the multi-relay transmission probability is analyzed. Simulation and experiment results on a real wireless sensor network platform in different channels demonstrate the the proposed scheme greatly reduces the multi-relay transmission probability and achieves about 84% improvement of energy efficiency compared with the traditional opportunistic MAC schemes.
Ruifeng Zhang 0001, Olivier Berder, Olivier Sentieys
PIMRC3
2013 GeCoS: A framework for prototyping custom hardware design flows
abstract
GeCoS is an open source framework that provide a highly productive environment for hardware design. GeCoS primarily targets custom hardware design using High Level Synthesis, distinguishing itself from classical compiler infrastructures. Compiling for custom hardware makes use of domain specific semantics that are not considered by general purpose compilers. Finding the right balance between various performance criteria, such as area, speed, and accuracy, is the goal, contrary to the typical goal in high performance context to maximize speed. The GeCoS infrastructure facilitates the prototyping of hardware design flows, going beyond compiler analyses and transformations. Hardware designers must interact with the compiler for design space exploration, and it is important to be able to give instant feedback to the users.
Antoine Floch, Tomofumi Yuki, Ali El Moussawi, Antoine Morvan, Kevin J. M. Martin, Maxime Naullet, Mythri Alle, Ludovic L'Hours, Nicolas Simon, Steven Derrien, François Charot, Christophe Wolinski, Olivier Sentieys
SCAM13
2013 An FPGA Software Defined Radio Platform with a High-Level Synthesis Design Flow
abstract
Software defined radio (SDR) opens a new door to future Internet of Things with higher degree of designing flexibility in context of wireless system development. Prototyping a remote implementation of wireless protocols on a hardware over the web requires a highly versatile software radio platform along with laid-back designing tools. To this aim, an FPGA-based SDR scheme has been proposed combining Virtex-6 Perseus 6010 platform capabilities and a design flow based on High-Level Synthesis (HLS) tools. A full IEEE 802.15.4 (ZigBee) physical layer has been implemented on the proposed platform from a C-language dataflow specification. All the results have been analyzed to lead to a fair comparison between different design flows. Although the proposed SDR has some designing issues, it shows a noticeable designing potentiality to flexible prototyping of future wireless systems.
Vaibhav Bhatnagar, Ganda Stéphane Ouedraogo, Matthieu Gautier, Arnaud Carer, Olivier Sentieys
VTC Spring5
2013 A Novel Hierarchical Low Complexity Synchronization Method for OFDM Systems
abstract
A new hierarchical synchronization method is proposed for initial timing synchronization in orthogonal frequency-division multiplexing(OFDM) systems. Based on the proposal of new training symbol, a threshold based timing metric is designed for accurate estimation of start of OFDM symbol in a frequency selective channel. Threshold is defined in terms of noise distributions and false alarm, which makes it independent of type of channel it is applied. Frequency offset estimation is also done for the proposed training symbol. The performance of the proposed timing metric is evaluated using simulation results. The proposed method achieves low mean squared error(MSE) in timing offset estimation at five(5 x) times lower computational complexity compared to cross-correlation based method in a frequency selective channel. It is also computationally efficient compared to hybrid approaches for OFDM timing synchronization.
Pramod P. Udupa, Olivier Sentieys, Pascal Scalart
VTC Spring2
2012 Latency-Energy Optimized MAC Protocol for Body Sensor Networks
abstract
This paper presents a self organized asynchronous medium access control (MAC) protocol for wireless body area sensor (WBASN). The protocol is optimized in terms of latency and energy under variable traffic. A body sensor network (BSN) exhibits a wide range of traffic variations based on different physiological data emanating from the monitored patient. For example, electrocardiogram data rate is multiple times more in comparison with body temperature rate. In this context, we exploit the traffic characteristics being observed at each sensor node and propose a novel technique for latency-energy optimization at the MAC layer. The protocol relies on dynamic adaptation of wake-up interval based on a traffic status register bank. The proposed technique allows the wake-up interval to converge to a steady state for variable traffic rates, which results in optimized energy consumption and reduced delay during the communication. A comparison with other energy efficient protocols is presented. The results show that our protocol outperforms the other protocols in terms of energy as well as latency under the variable traffic of WBASN.
Muhammad Mahtab Alam, Olivier Berder, Daniel Ménard, Olivier Sentieys
BSN4
2012 A semiempirical model for wakeup time estimation in power-gated logic clusters
abstract
Wakeup time is an important overhead that must be determined for effective power gating, particularly in logic clusters that undergo frequent mode transitions for run-time leakage power reduction. In this paper, a semiempirical model for virtual supply voltage in terms of basic parameters of the power-gated circuit is presented. Hence a closed-form expression for estimation of wakeup time of a power-gated logic cluster is derived. Experimental results of application of the model to ISCAS85 benchmark circuits show that wakeup time may be estimated within an average error of 16.3% across 22x variation in sleep transistor sizes and 13x variation in circuit sizes with significant speedup in computation time compared to SPICE level circuit simulations.
Vivek D. Tovinakere, Olivier Sentieys, Steven Derrien
DAC2
2012 From Scilab to High Performance Embedded Multicore Systems: The ALMA Approach
abstract
The mapping process of high performance embedded applications to today's multiprocessor system on chip devices suffers from a complex tool chain and programming process. The problem here is the expression of parallelism with a pure imperative programming language which is commonly C. This traditional approach limits the mapping, partitioning and the generation of optimized parallel code, and consequently the achievable performance and power consumption of applications from different domains. The Architecture oriented paraLlelization for high performance embedded Multicore systems using scilAb (ALMA) European project aims to bridge these hurdles through the introduction and exploitation of a Scilab-based toolchain which enables the efficient mapping of applications on multiprocessor platforms from high level of abstraction. This holistic solution of the toolchain allows the complexity of both the application and the architecture to be hidden, which leads to a better acceptance, reduced development cost, and shorter time-to-market. Driven by the technology restrictions in chip design, the end of exponential growth of clock speeds, and an unavoidable increasing request of computing performance, ALMA is a fundamental step forward in the necessary introduction of novel computing paradigms and methodologies.
Jürgen Becker 0001, Timo Stripf, Oliver Wolf, Michael Hübner 0001, Steven Derrien, Daniel Ménard, Olivier Sentieys, Gerard K. Rauwerda, Kim Sunesen, Nikolaos Kavvadias, Kostas Masselos, George Goulas, Panayiotis Alefragis, Nikos S. Voros, Dimitrios Kritharidis, Nikolaos Mitas, Diana Göhringer
DSD7
2012 Energy-delay tradeoff in wireless multihop networks with unreliable links
Ruifeng Zhang 0001, Olivier Berder, Jean-Marie Gorce, Olivier Sentieys
Ad Hoc Networks4
2012 System-Level Synthesis for Wireless Sensor Node Controllers: A Complete Design Flow
abstract
Wireless sensor networks (WSN) is a new and very challenging research field for embedded system design automation. Engineering a WSN node hardware platform is known to be a tough challenge, as the design must enforce many severe constraints, among which energy dissipation is by far the most important one. WSN node devices have until now been designed using off-the-shelf low-power microcontroller units (MCUs), even if their power dissipation is still an issue and hinders the widespread use of this new technology. In this work, we propose a complete system-level flow for an alternative approach based on the concept of hardware microtasks, which relies on hardware specialization and power gating to drastically improve the energy efficiency of the computational/control part of the node. Our case study shows that power savings between one to two orders of magnitude are possible w.r.t. MCU-based implementations.
Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys
ACM Trans. Design Autom. Electr. Syst.3
2011 Spectral efficiency and energy efficiency of distributed space-time relaying models
abstract
Cooperative relay techniques are exploited to reduce the transmission energy consumption which is very important for average and long range transmission in wireless sensor networks (WSNs). In this paper, the system having a two-antenna source, two one-antenna relays and a one-antenna destination is considered. Using distributed space-time code at relays, MIMO simple cooperative relay model (MSCR) and MIMO full cooperative relay model (MFCR) are proposed in comparison with MIMO normal cooperative relay model (MNCR) where the relays forward signals consecutively to destination. The outage probability analysis derives that all the models have the diversity order of 4. However, the analytic and simulation results show that MFCR has better performance than MNCR, whereas MSCR provides better spectral efficiency than MNCR. Moreover, the energy efficiency of these models is also considered by using a realistic power consumption model where the parameters are extracted from the characteristics of CC2420, a wireless sensor transceiver widely used and commercially available. For each transmission ranges, the optimal cooperative scheme in terms of energy efficiency is provided by simulation results.
Le Quang Vinh Tran, Olivier Berder, Olivier Sentieys
CCNC3
2011 Error recovery technique for coarse-grained reconfigurable architectures
abstract
This paper presents the implementation of the error recovery scheme from temporary faults, applicable for datapaths of coarse-grained reconfigurable architectures. We have chosen the DART architecture as a vehicle to study various aspects related to implementation of the instruction retry in a complex highly parallel reconfigurable system. Synthesis results have confirmed the time, hardware, and power consumption efficiency of the proposed approach, which can be applied independently on the concurrent error detection scheme actually used.
Muhammad Moazam Azeem, Stanislaw J. Piestrak, Olivier Sentieys, Sébastien Pillement
DDECS3
2011 Non-regenerative full distributed space-time codes in cooperative relaying networks
abstract
Distributed space-time codes (DSTC) are often used in cooperative relaying networks whose relays can support a single antenna due to the limited physical size. In this paper, full DSTC protocol in which there is a data exchange between relays before forwarding signals to destination is proposed to improve the performance of a cooperative relaying system. A lower bound for the average symbol error probability (ASEP) of full DSTC cooperative relaying system in a Rayleigh fading environment is provided. In the case when the Signal to Noise Ratio (SNR) of the relay-relay link is much greater than that of the source-relay link, the upper bound on ASEP of this system is also derived. From the simulations, we show that the average SNR gain of full DSTC system over DSTC system is 3.8dB and the maximum SNR gain is 5dB when the relay-relay distance is small and the relays are in the middle of the source and the destination. The effect of the distance between the relays shows that the performance does not degrade so much as the distance between relays is lower than a half of the source-destination distance. Moreover, we also show that when the error synchronization range is lower than 0.5, the impact of the transmission synchronization error of the relay-destination link on the performance is not considerable.
Le Quang Vinh Tran, Olivier Berder, Olivier Sentieys
WCNC3
2011 Real-time scheduling on heterogeneous system-on-chip architectures using an optimised artificial neural network
Daniel Chillet, Antoine Eiche, Sébastien Pillement, Olivier Sentieys
J. Syst. Archit.4
2011 Energy-Efficient Cooperative Techniques for Infrastructure-to-Vehicle Communications
abstract
In wireless distributed networks, cooperative relay and cooperative multiple-input-multiple-output (MIMO) techniques can be used to exploit the spatial and temporal diversity gains to increase the performance or reduce the transmission energy consumption. The energy efficiency of cooperative MIMO and relay techniques is then very useful for the infrastructure-to-vehicle (I2V) and infrastructure-to-infrastructure (I2I) communications in intelligent transport system (ITS) networks, where the energy consumption of wireless nodes embedded on road infrastructure is constrained. In this paper, applications of cooperation between nodes to ITS networks are proposed, and the performance and the energy consumption of cooperative relay and cooperative MIMO are investigated and compared with the traditional multihop technique. The comparison between these cooperative techniques helps us choose the optimal cooperative strategy in terms of energy consumption for energy-constrained road infrastructure networks in ITS applications.
Tuan-Duc Nguyen, Olivier Berder, Olivier Sentieys
IEEE Trans. Intell. Transp. Syst.3
2010 A complete design-flow for the generation of ultra low-power WSN node architectures based on micro-tasking
abstract
Wireless Sensor Networks (WSN) are a new and very challenging research field for embedded system design automation, as their design must enforce stringent constraints in terms of power and cost. WSN node devices have until now been designed using off-the-shelf low-power microcontroller units (MCUs), even if their power dissipation is still an issue and hinders the wide-spreading of this new technology. In this paper, we propose a new architectural model for WSN nodes (and its complete design-flow from C downto synthesizable VHDL) based on the notion of micro-tasks. Our approach combines hardware specialization and power-gating so as to provide an ultra low-power solution for WSN node design. Our first estimates show that power savings by one to two orders of magnitude are possible w.r.t. MCU-based implementations.
Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys
DAC3
2010 System Level Synthesis for Ultra Low-Power Wireless Sensor Nodes
abstract
Engineering hardware platform for a Wireless Sensor Network (WSN) node is known to be a tough challenge, as the design must enforce many severe constraints, among which energy dissipation is by far the most challenging one. Today, most of the WSN node platforms are based on low cost and low-power programmable micro controllers, even if it is acknowledged that their energy efficiency remains limited and hinders the wide-spreading of WSN to new applications. In this paper, we propose a complete system level flow for an alternative approach based on the concept of hardware micro-tasks, which relies on hardware specialization and power gating to dramatically improve the energy efficiency of the computational part of the node. Early estimates show power saving by more than one order of magnitude over MCU-based implementations.
Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys
DSD3
2010 Analytical approach for analyzing quantization noise effects on decision operators
abstract
The presence of decision operators has proved to be a serious impediment for a fully analytical noise power estimation technique. This paper proposes a generalized decision operator which can potentially capture the behavior of all possible types of decision operators and provides a fully analytical technique to handle them while performing quantization noise power estimation. The proposed method is applied to BPSK and 16-QAM decision operators. The total error rate and the PDF of the error signal are found to follow the simulation to a great degree of accuracy.
Karthick Parashar, Romuald Rocher, Daniel Ménard, Olivier Sentieys
ICASSP4
2010 Fast performance evaluation of fixed-point systems with un-smooth operators
abstract
Fixed-point refinement of signal processing systems is an essential step performed before implementation of any signal processing system. Existing analytical techniques to evaluate performance of fixed-point systems are not applicable to the errors due to quantization in the presence of un-smooth operators. Thus, it is inevitable to use simulation to evaluate performance of fixed-point systems in the presence un-smooth operators. This paper proposes a hybrid technique which can be used in place of pure simulation to accelerate the performance evaluation. The principle idea in the proposed hybrid approach is to selectively simulate parts of the system only when un-smooth errors occur but use analytical results otherwise. The acceleration thus obtained reduces the performance evaluation time which can be used to explore a wider word-length design space or speedup the optimization process. This method has been tried on a complex MIMO sphere decoding algorithm and the results obtained show several orders of magnitude improvement in terms of evaluation time.
Karthick Parashar, Daniel Ménard, Romuald Rocher, Olivier Sentieys, David Novo, Francky Catthoor
ICCAD4
2010 Cooperative MISO and Relay Comparison in Energy Constrained WSNs
abstract
In wireless distributed networks where multiple antennas can not be installed in one wireless node, cooperative relay and cooperative Multi-Input Multi-Output (MIMO) techniques can be used to exploit the spatial and temporal diversity gain in order to reduce the energy consumption. The simplicity and the energy efficiency of cooperative Multi-Input Single-Output (MISO) and relay techniques is very useful for the energy constrained Wireless Sensor Networks (WSN). In this paper, the performance and the energy consumption of the cooperative MISO and relay techniques are investigated over a Rayleigh fading channel. If under ideal conditions cooperative MISO has been proved to be better than relay, the latter is a better solution when transmission synchronization errors occur. The comparison between these two cooperative techniques helps us to choose the optimal cooperative strategy for energy constrained WSN applications.
Tuan-Duc Nguyen, Olivier Berder, Olivier Sentieys
VTC Spring3
2009 xMAML: A Modeling Language for Dynamically Reconfigurable Architectures
abstract
Constant evolution of norms and applications, usually implemented on system-on-chip (SOC), increases architecture performance and flexibility requirements. Current architectures are consequently becoming more complex and difficult to develop. One of the solutions is to develop design frameworks based on high-level architecture description languages (ADL). These ADLs are useful for a rapid description of the hardware that should be implemented on an architecture. Designers can use ADL for the development of generic frontend tools. Our framework aims at designing dynamically reconfigurable architecture with the help of an ADL. This paper presents xMAML, an architecture description language dedicated to the instantiation of dynamically reconfigurable heterogeneous computing units. From this ADL, a synthesizable model is produced after exploration, simulation and validation phases. As proof of concept, exploration for a WCDMA receiver on two dynamically reconfigurable architectures is presented.
Julien Lallet, Sébastien Pillement, Olivier Sentieys
DSD3
2009 Minimum Distance Based Precoder for MIMO-OFDM Systems Using a 16-QAM Modulation
abstract
A precoder based on the exact optimization of the minimum Euclidean distance dminbetween signal points at the receiver side is proposed for MIMO-OFDM systems using a 16-QAM modulation. Assuming that channel state information (CSI) can be made available at the transmitter, the channel is diagonalized and a precoder can be derived. A numerical approach shows that the precoder design depends on the channel characteristics, leading to 8 different precoder expressions. Comparisons with maximum signal-to-noise ratio (SNR) strategy and other precoders based on criteria, such as water-filling (WF), minimum mean square error (MMSE), and maximization of the minimum singular value of the global channel matrix, are performed to illustrate the significant bit-error-rate (BER) improvement of the proposed precoder. In order to make its implementation easier, it is shown that it can be expressed by only two ways without significant performance degradation.
Quoc-Tuong Ngo, Olivier Berder, Baptiste Vrigneau, Olivier Sentieys
ICC4
2009 Dynamic Precision Scaling for Low Power WCDMA Receiver
abstract
One of the most important applications of digital signal processing (DSP) is wireless communication. This kind of application requires low power implementation of DSP, which generally uses fixed-point arithmetic. The fixed-point architectures should be developed to maintain the energy consumption power at a reasonable level. In this paper, an approach which adapts the fixed-point specification according to the input receiver signal-to-noise ratio (SNR) is proposed. To underline our approach interest, the rake receiver of a WCDMA receiver is examined. Results show about 25% - 40% energy savings with our dynamic precision approach.
Hai-Nam Nguyen, Daniel Ménard, Olivier Sentieys
ISCAS3
2009 Ultra Low-power FSM for Control Oriented Applications
abstract
In this paper, we propose an approach combining the use of distributed hardware tasks implemented as finite state machines (FSM) and power gating techniques to obtain ultra low-power implementations. We target for control dominated applications represented as control task graphs, and propose a complete flow including a C to hardware task compiler. Our approach is validated experimentally and shows impressive improvement over software implementation on leading edge low-power microcontrollers such as the MSP430.
Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys
ISCAS3
2009 On-line Monitoring of Random Number Generators for Embedded Security
abstract
Many embedded security chips require a high-quality random number generator (RNG). Unfortunately, hardware RNG randomness can vary in time due to implementation defects or certain kinds of attacks. To overcome this issue, this paper presents the implementation of a battery of statistical test for randomness. The battery is selected for its efficient implementation, making the area and power consumption insignificant. Performance and cost of the hardware implementation are given for FPGA and VLSI targets. Results show that statistical tests can easily be implemented in low-cost embedded security circuits and can enhance on-line monitoring of RNG randomness to prevent RNG failures.
Renaud Santoro, Olivier Sentieys, Sébastien Roy 0002
ISCAS2
2008 Impact of Transmission Synchronization Error and Cooperative Reception Techniques on the Performance of Cooperative MIMO Systems
abstract
Wireless sensor networks where neighbor nodes can cooperate both at transmission and reception by using the cooperative multi-input multi-output (MIMO) technique are considered. Space-time diversity gain can be exploited to reduce the transmission energy consumption which is very important for average and long range transmission in wireless sensor network (WSN). However, differing from classical MIMO systems, cooperative systems suffer from the lack of synchronization between distributed nodes and the additive noise in the cooperative reception. At the cooperative transmission side, we investigate the effect of transmission synchronization error which generates inter-symbol interference (ISI), decreases the desired signal amplitude and makes the channel state information (CSI) more difficult to be estimated by the receiver. At the cooperative reception side, two strategies of wireless transmission techniques which are energy efficient and have good performance are also presented. The simulation of a cooperative MIMO system using Alamouti and Tarokh space-time block codes (STBC) over Rayleigh fading channels presents a performance degradation in the presence of transmission synchronization error and additional noise of cooperative reception techniques. For a small synchronization error range at cooperative transmission and a reasonable amplification factor in reception, the degradation is negligible and the cooperative MIMO performance is rather good.
Tuan-Duc Nguyen, Olivier Berder, Olivier Sentieys
ICC3
2008 Efficient Space Time Combination Technique for Unsynchronized Cooperative Miso Transmission
abstract
In the context of cooperative multi-input multi-output (MIMO) techniques for wireless sensor networks, the transmission synchronization error impact is investigated in this paper and a new space-time combination technique is proposed. Cooperative MIMO techniques have been recently studied in order to reduce the energy consumption in distributed wireless sensor networks. Differing from classical MIMO systems, the cooperative antennas are physically separated in a cooperative system which leads to the unsynchronized cooperative transmission. The transmission synchronization error generates inter-symbol interference (ISI) and decreases the desired signal amplitude at the receiver, leading to a performance degradation and an additional energy required for data transmission. For small range of synchronization error, the performance degradation is negligible and the cooperative system performance is rather tolerant. However, for large range of error, the degradation is significant. A new space time combination technique is proposed for a cooperative multi-input single-output (MISO) system using Alamouti codes. The proposed technique combines two sequences of the received signal sampled at different times in order to perform an orthogonal combination. This technique has a better tolerance to the transmission synchronization error and also a low complexity. The significant advantage of using this new combination technique over traditional Alamouti combination is illustrated by simulations over a Rayleigh fading channel.
Tuan-Duc Nguyen, Olivier Berder, Olivier Sentieys
VTC Spring3
2007 A Neural Network Model for Real-Time Scheduling on Heterogeneous SoC Architectures
abstract
With increasing embedded application complexity, designers have proposed to introduce new hardware architectures based on heterogeneous processing units on a single chip. For these architectures, the scheduling service of a realtime operating system must be able to assign tasks on different execution resources. This paper presents a model of artificial neural networks used for real-time task scheduling to heterogeneous system-on-chip architectures. Our proposition is an adaptation of the Hopfield model and the main objective concerns the minimization of the neuron number to facilitate future hardware implementation of this service. In fact, to ensure rapid convergence and low complexity, this number must be dramatically reduced. So, we propose new constructing rules to design smaller neural network and we show, through simulations, that network stabilization is obtained without reinitialisation of the network.
Daniel Chillet, Sébastien Pillement, Olivier Sentieys
IJCNN3
2007 Cooperative MIMO Schemes Optimal Selection for Wireless Sensor Networks
abstract
A cooperative MIMO scheme selection is proposed for wireless sensor networks where energy consumption is the most important design criterion. Space-time block codes are designed to achieve maximum diversity for a given number of transmit and receive antennas with very simple decoding algorithm. In radio fading channel, STBC require less transmission energy than SISO technique for the same bit error rate and can be employed practically in wireless sensor networks by using the cooperative MIMO scheme. Considering Alamouti and Tarokh space-time block codes, the number of antennas at both the transmission and the reception sides are selected with respect to the transmission distance. By using cooperative MIMO transmission instead of SISO, it is shown that the distance between nodes can be increased and a large amount of the total energy can be saved for middle and long distance transmission. The energy efficiency of cooperative MIMO over SISO and multihop SISO is proved by simulations, and a multi-hop technique for cooperative MIMO is also proposed for good energy-efficiency and limited number of available cooperative nodes
Tuan-Duc Nguyen, Olivier Berder, Olivier Sentieys
VTC Spring3
2006 An energy-efficient ternary interconnection link for asynchronous systems
abstract
We introduce a new ternary link including a binary-to-ternary encoder and a ternary-to-binary decoder in voltage-mode multiple-valued logic (MVL). This link improves the transistor count compared to existing designs and it has no DC current path. The complete link was simulated with SPICE and a 0.13mum CMOS technology. It additionally shows interesting advantages on power consumption for global interconnects compared to full-swing signaling binary systems (up to 56.4% less energy consumption). Its low propagation delay is also an advantage in the design of high-speed on-chip links for asynchronous systems
Jean-Marc Philippe, E. Kinvi-Boh, Sébastien Pillement, Olivier Sentieys
ISCAS4
2006 Fixed-point configurable hardware components for adaptive filters
abstract
To reduce the gap between the VLSI technology capability and the designer productivity, design reuse based on IP (intellectual properties) is commonly used. In terms of arithmetic accuracy, the generated architecture can generally only be configured through the input and output word-lengths. In this paper, a new kind of fixed-point arithmetic IP is presented through the LMS and delayed-LMS examples. The operator and memory word-lengths are optimized under an accuracy constraint defined by the user. To significantly reduce the optimization and design times, the architecture parameter determination is based on analytical approach
Romuald Rocher, Nicolas Hervé, Daniel Ménard, Olivier Sentieys
ISCAS4
2005 Accuracy evaluation of fixed-point APA algorithm [adaptive filter applications]
abstract
The implementation of adaptive filters with fixed-point arithmetic requires us to evaluate the computation quality. The accuracy can be determined by calculating the global quantization noise power in the system output. In this paper, a new model for evaluating analytically the global noise power in the APA (affine projection algorithm) is developed. The model is presented and applied to the NLMS-OCF. The accuracy of our model is analyzed by experimentation.
Romuald Rocher, Daniel Ménard, Olivier Sentieys, Pascal Scalart
ICASSP (5)3
2004 Accuracy evaluation of fixed-point LMS algorithm
abstract
The implementation of adaptive filters with fixed-point arithmetic requires the computation quality to be evaluated. The accuracy may be determined by calculating the global quantization noise power in the system output. A new model for evaluating analytically the global noise power in the LMS algorithm and in the NLMS algorithm is developed. Two existing models are presented, then the model is detailed and compared with the ones before. The accuracy of our model is analyzed by simulation.
Romuald Rocher, Daniel Ménard, Olivier Sentieys, Pascal Scalart
ICASSP (5)3
2004 DSP Code Generation with Optimized Data Word-Length Selection
Daniel Ménard, Olivier Sentieys
SCOPES2
2002 Automatic floating-point to fixed-point conversion for DSP code generation
abstract
The development of methodologies for the automatic implementation of floating-point algorithms in fixed-point architectures is required for the minimization of cost, power consumption and time to market of digital signal processing applications. In this paper, a new methodology of implementation in Digital Signal Processors (DSP) under accuracy constraint is presented. In comparison with the existing methodologies, the DSP architecture is completely taken into account for optimizing the execution time under accuracy constraint. The justification and the different stages of our methodology are presented.
Daniel Ménard, Daniel Chillet, François Charot, Olivier Sentieys
CASES4
2002 Automatic Evaluation of the Accuracy of Fixed-Point Algorithms
abstract
The minimization of cost, power consumption and time-to-market of DSP applications requires the development of methodologies for the automatic implementation of floating-point algorithms in fixed-point architectures. In this paper a new methodology for evaluating the quality of an implementation through the automatic determination of the Signal to Quantization Noise Ratio (SQNR) is under consideration. The theoretical concepts and the different phases of the methodology are explained. Then, the ability of our approach for computing the SQNR efficiently and its beneficial contribution in the process of data word-length minimization are shown through some examples.
Daniel Ménard, Olivier Sentieys
DATE2
2002 A Compilation Framework for a Dynamically Reconfigurable Architecture
Raphaël David, Daniel Chillet, Sébastien Pillement, Olivier Sentieys
FPL4
2002 Mapping future generation mobile telecommunication applications on a dynamically reconfigurable arcidtecture
abstract
In addition to the high performance requirements inherent to multimedia processings or to W -CDMA, future generation mobile telecommunications bring new constraints to the semiconductor design world. In fact, the traditional solutions based on the use of hardware devices (ASIC) or software ones (DSP) are unable to associate the flexibility and the high level of performance to the low energy consumption required by this application domain. In this paper, we study the efficiency of an architecture based on the use of the functional reconfiguration for such systems. Thanks to the implementation of key applications of UMTS, we will show that reconfigurable architectures can offer new compromises to associate high performances and low energy consumption in a flexible architecture and so, can be the solution to the set of problems associated with the future generation mobiles telecommunications systems.
Raphaël David, Daniel Chillet, Sébastien Pillement, Olivier Sentieys
ICASSP4
2002 A methodology for evaluating the precision of fixed-point systems
abstract
The minimization of cost, power consumption and time-to-market of DSP applications requires the development of methodologies for the automatic implementation of floating-point algorithms in fixed-point architectures. In this paper, a new methodology for evaluating the quality of an implementation through the automatic determination of the Signal to Quantization Noise Ratio (SQNR) is presented. The modelization of the system at the quantization noise level and the expression of the output noise power is detailed for linear systems. Then, the different phases of the methodology are explained and the ability of our approach for computing the SQNR efficiently is shown through examples.
Daniel Ménard, Olivier Sentieys
ICASSP2
2000 Multi-algorithm ASIP synthesis and power estimation for DSP applications
abstract
Power consumption is an increasingly important parameter in the design of mixed hardware/software systems. This work applies the high-level synthesis technique to multi-algorithms and explores its use as a means of analyzing power consumption from the high level of design. We apply a multi-algorithm synthesis technique to designing an application specific instruction set processor (ASIP) from a customized ASIC. This technique synthesizes selected time constrained algorithms to define a set of DSP applications, designs the corresponding ASIP core, and extracts the specific instruction set. Although not as effective as a DSP core solution, this technique provides much of the circuit flexibility while maintaining an available trade-off between performance and power dissipation. This technique contains three power estimators to assist algorithm integration with the view to optimizing the embedded system: the first acts during the application of usual high-level synthesis steps. The second one is triggered after the complete synthesis of the target algorithm, and the third estimator is based on the instruction set of the designed ASIP core. This technique has been implemented in our framework called BSS (Breizh Synthesis System).
Jean-Gabriel Cousin, Olivier Sentieys, Daniel Chillet
ISCAS2
1999 Memory Unit Design for Real Time DSP Applications
abstract
Today, the design complexity for new applications (such as telecommunication, multi media, internet), requires new high level tools which enable us to translate the behavioral description into hardware. All of the recent High Level Synthesis tools are able to transform high level specifications in an ASIC based on processing and control units. In general, these tools do not handle a real optimization of the memory unit. However, in many applications, the hardware solution may be challenged by the number and the complexity of memory units. This paper proposes to complete the synthesis design flow by including the memory unit synthesis. Our methodology is integrated in the BSS (Breizh Synthesis System http://www.enssat.fr/bss) project which is a framework for the design of real-time constraint applications.
Daniel Chillet, Olivier Sentieys, Michel Corazza
Great Lakes Symposium on VLSI2
1997 VLSI high level synthesis of fast exact least mean square algorithms based on fast FIR filters
abstract
This paper relates experiences of algorithmic transformations in High Level Synthesis, in the area of acoustic echo cancellation. The processing and memory units are automatically designed for various equivalent LMS algorithms, in the FIR case, with important computational load. The results obtained with different filter lengths, give an accurate prototyping of new fast versions of the LMS algorithm. It also show that a theoretical arithmetic reduction must be correlated to the associated increase of memory requirements.
Jean-Philippe Diguet, Olivier Sentieys, Daniel Chillet, Jean Luc Philippe
ICASSP2