Philippe Coussy

dblp:48/2293 · DBLP profile ↗
← Back
52ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0002-7222-5271ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 8 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Exploring the Contribution of Hardware Shuffling in Securing Low-Cost Symmetric Encryption Devices against Power-Based Side-Channel Attacks: Case Study of an AES-128 on FPGA
abstract
In the era of the Internet of Things (IoT), embedded systems are massively spreading in critical infrastructures. Low-cost and low-power components are used to build such devices, which manipulate sensitive data and communicate at continuously growing throughput. To protect these data, IoT nodes embed cryptographic primitives including countermeasures against Side-Channel Attacks (SCA). In this article, we explore the interest of hardware-based shuffling to protect AES ciphers against power-based side-channel attacks in the context of low-cost IoT devices. Shuffling is performed via a dedicated hardware module, wherein a Pseudo-Random Number Generator provides a random vector used to control a permutation network, generating a permutation which determines the AES computation sequence. The approach has been explored and evaluated through several FPGA-based design solutions in term of area, timing performance, and security. Compared to an unprotected design, the best solution leads to a minimum area overhead factor of 1.2 or a maximum throughput of 45.23 Mbit/s. Furthermore, compared to existing works that also depend on hardware shuffling, the proposed solution is up to 10.4 times faster. Results show that hardware-based shuffling solution as implemented increases the Measure-to-Disclosure metric by a factor greater than 10,000 when considering Correlation Power Analysis-based SCA.
Vianney Lapotre, Cyrille Chavet, Ghita Harcha, Philippe Coussy
ACM Trans. Reconfigurable Technol. Syst.4
2024 SplitMS: Split Modulo-Scheduling for Accelerating Loops Onto CGRAs
abstract
Coarse-Grained Reconfigurable Array (CGRA) ar-chitectures are popular for accelerating loop kernels due to a good balance between energy efficiency and flexibility. Modulo scheduling (MS) is the preferred solution for efficiently mapping loops onto CGRAs. Existing CGRA MS algorithms suffer from low resource utilization if the number of operation nodes in the Data Flow Graph (DFG) is less than the number of Processing Elements (PEs) in the CGRA. To improve instruction level parallelism (ILP), the common approaches unroll the loop before applying MS. However, finding valid MS solutions for larger DFGs becomes difficult for CGRAs with resource constraints. This paper proposes a novel Split Modulo-Scheduling (SplitMS) technique to improve the ILP by segmenting the target CGRA into clusters and mapping loop chunks. We also present a lightweight hardware approach to support the cluster execution. Experiments show that SplitMS for a$4\times[2\times 2]$CGRA cluster achieves an average speedup of$2.8\times$over MS for a$4\times 4$target CGRA with 8 Load-Store Units (LSUs). SplitMS increases an average of$2.9\times$the PE utilization and$3\times$the energy efficiency over the conventional MS approach.
Christie Sajitha Sajan, Kevin J. M. Martin, Satyajit Das, Philippe Coussy
DSD4
2024 CREPE: Concurrent Reverse-Modulo-Scheduling and Placement for CGRAs
abstract
Coarse-Grained Reconfigurable Array (CGRA) architectures are popular as high-performance and energy-efficient computing devices. Compute-intensive loop constructs of complex applications are mapped onto CGRAs by modulo-scheduling the innermost loop dataflow graph (DFG). In the state-of-the-art approaches, mapping quality is typically determined by initiation interval (II), whileschedule lengthfor one iteration is neglected. However, for nested loops,schedule lengthbecomes important. In this article, we propose CREPE, aConcurrentReverse-modulo-scheduling andPlacement technique for CGRAs that minimizes bothIIandschedule length. CREPE performs simultaneous modulo-scheduling and placement coupled with dynamic graph transformations, generating good-quality mappings with high success rates. Furthermore, we introduce a compilation flow that maps nested loops onto the CGRA and modulo-schedules the innermost loop using CREPE. Experiments show that the proposed solution outperforms the conventional approaches in mapping success rate and total execution time with no impact on the compilation time. CREPE maps all kernels considered while state-of-the-art techniques Crimson and Epimap failed to find a mapping or mapped at very highIIs. On a 2×4 CGRA, CREPE reports a 100% success rate and a speed-up up to 5.9× and 1.4× over Crimson with 78.5% and Epimap with 46.4% success rates respectively.
Chilankamol Sunny, Satyajit Das, Kevin J. M. Martin, Philippe Coussy
IEEE Trans. Parallel Distributed Syst.4
2023 An Efficient and Flexible Stochastic CGRA Mapping Approach
abstract
Coarse-Grained Reconfigurable Array (CGRA) architectures are promising high-performance and power-efficient platforms. However, mapping applications efficiently on CGRA is a challenging task. This is known to be an NP complete problem. Hence, finding good mapping solutions for a given CGRA architecture within a reasonable time is complex. Additionally, finding scalability in compilation time and memory footprint for large heterogeneous CGRAs is also a well known problem. In this article, we present a stochastic mapping approach that can efficiently explore the architecture space and allows finding best of solutions while having limited and steady use of memory footprint. Experimental results show that our compilation flow allows to reach performances with low-complexity CGRA architectures that are as good as those obtained with more complex ones thanks to the better exploration of the mapping solution space. Parameters considered in our experiments are number of tiles, Register File (RF) size, number of load/store (LS) units, network topologies, and so on. Our results demonstrate that high-quality compilation for a wide range of applications is possible within reasonable run-times. Experiments with several DSP benchmarks show that the best CGRA configuration from the architectural exploration surpasses an ultra low-power DSP optimized RISC-V CPU to achieve up to 15.28× (with an average of 6× and minimum of 3.4×) performance gain and 29.7× (with an average of 13.5× and minimum of 6.3×) energy gain with an area overhead of 1.5× only.
Satyajit Das, Kevin J. M. Martin, Thomas Peyret, Philippe Coussy
ACM Trans. Embed. Comput. Syst.4
2021 Opportunistic IP Birthmarking using Side Effects of Code Transformations on High-Level Synthesis
abstract
The increasing design and manufacturing costs are leading to globalize the semiconductor supply chain. However, a malicious attacker can resell a stolen Intellectual Property (IP) core, demanding methods to identify a relationship between a given IP and a potentially fraudulent copy. We propose a method to protect IP cores created with high-level synthesis (HLS): our method inserts a discrete birthmark in the HLS-generated designs that uses only intrinsic characteristics of the final RTL. The core of our process leverages the side effects of HLS due to specific source-code manipulations, although the method is HLS-tool agnostic. We propose two independent validation metrics, showing that our solution introduces minimal resource and delay overheads (< 6% and < 2%, respectively) and the accuracy in detecting illegal copies is above 96%.
Hannah Badier, Christian Pilato, Jean-Christophe Le Lann, Philippe Coussy, Guy Gogniat
DATE4
2020 TRANSPIRE: An energy-efficient TRANSprecision floating-point Programmable archItectuRE
abstract
In recent years, Coarse Grain Reconfigurable Architecture (CGRA) accelerators have been increasingly deployed in Internet-of-Things (IoT) end nodes. A modern CGRA has to support and efficiently accelerate both integer and floating-point (FP) operations. In this paper, we propose an ultra-low-power tunable-precision CGRA architectural template, called TRANSprecision floating-point Programmable archItectuRE (TRANSPIRE), and its associated compilation flow supporting both integer and FP operations. TRANSPIRE employs transprecision computing and multiple Single Instruction Multiple Data (SIMD) to accelerate FP operations while boosting energy efficiency as well. Experimental results show that TRANSPIRE achieves a maximum of 10.06× performance gain and consumes 12.91× less energy w.r.t. a RISC-V based CPU with an enhanced ISA supporting SIMD-style vectorization and FP data-types, while executing applications for near-sensor computing and embedded machine learning, with an area overhead of 1.25× only.
Rohit Prasad, Satyajit Das, Kevin J. M. Martin, Giuseppe Tagliavini, Philippe Coussy, Luca Benini, Davide Rossi 0001
DATE5
2020 Energy Efficient Acceleration Of Floating Point Applications Onto CGRA
abstract
In this paper, we propose a novel CGRA architecture and associated compilation flow supporting both integer and floating-point computations for energy efficient acceleration of DSP applications. Experimental results show that the proposed accelerator achieves a maximum of 4.61 × speedup compared to a DSP optimized, ultra low power RISC-V based CPU while executing seizure detection, a representative of wide range of EEG signal processing applications with an area overhead of 1.9×. The proposed CGRA achieves a maximum of 6.5× energy efficiency compared to the CPU.
Satyajit Das, Rohit Prasad, Kevin J. M. Martin, Philippe Coussy
ICASSP4
2020 Toward Secured IoT Devices: A Shuffled 8-Bit AES Hardware Implementation
abstract
In this paper, we present a lightweight secured AES hardware implementation designed to further resist to Side Channel Attacks relying on Power Analysis. The proposed architecture is based on an 8-bit data-path, and the protection is provided by shuffling computations and memory locations. Our shuffling module is based on a permutation network controlled by a Random Number Generator and leads to the best compromise between security, area, and performances compared to state-of-the-art. Implementation results on a spartan-6 FPGA show that the proposed protection mechanisms impact the area and the timing performance of the unprotected design by factors of 1.58 and 0.35 respectively. Security evaluation based on simulation results shows that the proposed secure architecture resists to a regular CPA by revealing a unique key byte when attacking with up to 1 million traces while state-of-the-art shuffled designs requires only 50000 traces to retrieve the entire secret key. Considering an integrated CPA (also called windowing attack), the proposed architecture allows increasing up to ×300 the required number of traces (Measurements to Disclosure) to retrieve 40% of the key bytes and reveals no more than 9 key bytes when attacking with up to 1 million traces.
Ghita Harcha, Vianney Lapotre, Cyrille Chavet, Philippe Coussy
ISCAS4
2019 Transient Key-based Obfuscation for HLS in an Untrusted Cloud Environment
abstract
Recent advances in cloud computing have led to the advent of Business-to-Business Software as a Service (SaaS) solutions, opening new opportunities for EDA. High-Level Synthesis (HLS) in the cloud is likely to offer great opportunities to hardware design companies. However, these companies are still reluctant to make such a transition, due to the new risks of Behavioral Intellectual Property (BIP) theft that a cloud-based solution presents. In this paper, we introduce a key-based obfuscation approach to protect BIPs during cloud-based HLS. The source-to-source transformations we propose hide functionality and make normal behavior dependent on a series of input keys. In our process, the obfuscation is transient: once an obfuscated BIP is synthesized through HLS by a service provider in the cloud, the obfuscation code can only be removed at Register Transfer Level (RTL) by the design company that owns the correct obfuscation keys. Original functionality is thus restored and design overhead is kept at a minimum. Our method significantly increases the level of security of cloud-based HLS at low performance overhead. The average area overhead after obfuscation and subsequent de-obfuscation with tests performed on ASIC and FPGA is 0.39%, and over 95% of our tests had an area overhead under 5%.
Hannah Badier, Jean-Christophe Le Lann, Philippe Coussy, Guy Gogniat
DATE3
2019 Context-memory Aware Mapping for Energy Efficient Acceleration with CGRAs
abstract
Coarse Grained Reconfigurable Arrays (CGRAs) are emerging as low power computing alternative providing a high grade of acceleration. However, the area and energy efficiency of these devices are bottlenecked by the configuration/context memory when they are made autonomous and loosely coupled with CPUs. The size of these context memories is of prime importance due to their high area and impact on the power consumption. For instance, a 64-word context memory typically represents 40% of a processing element area. In this context, since traditional mapping approaches do not take the size of the context memory into account, CGRAs often become oversized which strongly degrade their performance and interest. In this paper, we propose a context memory aware mapping for CGRAs to achieve better area and energy efficiency. This paper motivates the need of constraining the size of the context memory inside the processing element (PE) for ultra low power acceleration. It also describes the mapping approach which tries to find at least one mapping solution for a given set of constraints defined by the context memories of the PEs. Experiments show that our proposed solution achieves an average of 2.3× energy gain (with a maximum of 3.1× and a minimum of 1.4×) compared to the mapping approach without the memory constraints, while using 2× less context memory. When compared to the CPU, the proposed mapping achieves an average of 14× (with a maximum of 23× and minimum of 5×) energy gain.
Satyajit Das, Kevin J. M. Martin, Philippe Coussy
DATE3
2019 Solving Memory Access Conflicts in LTE-4G Standard
abstract
In mobile telecommunications domain, LTE-4G is currently the most advanced standard available. A major design issue of LTE-4G based systems resides in solving the memory access conflicts when connecting Rate-Matching (RM) and Error Correction Code (ECC) modules. In this paper, we first describe and analyze the problem before proposing a dedicated memory mapping approach. Results show that our method allows removing any conflicts for any block sizes and any parallelism degrees in the context of LTE-4G.
Cyrille Chavet, Fabrice Lozachmeur, T. Barguil, A. S. Hussein, Philippe Coussy
ICASSP5
2019 An Energy-Efficient Integrated Programmable Array Accelerator and Compilation Flow for Near-Sensor Ultralow Power Processing
abstract
In this paper, we give a fresh look to coarse grained reconfigurable arrays (CGRAs) as ultralow power accelerators for near-sensor processing. We present a general-purpose integrated programmable-array accelerator (IPA) exploiting a novel architecture, execution model, and compilation flow for application mapping that can handle kernels containing complex control flow, without the significant energy overhead incurred by state of the art predication approaches. To optimize the performance and energy efficiency, we explore the IPA architecture with special focus on shared memory access, with the help of the flexible compilation flow presented in this paper. We achieve a maximum energy gain of 2×, and performance gain of 1.33× and 1.8× compared with state of the art partial and full predication techniques, respectively. The proposed accelerator achieves an average energy efficiency of 1617 MOPS/mW operating at 100 MHz, 0.6 V in 28 nm UTBB FD-SOI technology, over a wide range of near-sensor processing kernels, leading to an improvement up to 18×, with an average of 9.23× (as well as a speed-up up to 20.3×, with an average of 9.7×) compared to a core specialized for ultralow power near-sensor processing.
Satyajit Das, Kevin J. M. Martin, Davide Rossi 0001, Philippe Coussy, Luca Benini
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 Clone-Based Encoded Neural Networks to Design Efficient Associative Memories
abstract
In this paper, we introduce a neural network (NN) model named clone-based neural network (CbNN) to design associative memories. Neurons in CbNN can be cloned statically or dynamically which allows to increase the number of data that can be stored and retrieved. Thanks to their plasticity, CbNN can handle correlated information more robustly than existing models and thus provides better memory capacity. We experiment this model in encoded neural networks also known as Gripon-Berrou NNs. Numerical simulations demonstrate that memory and recall abilities of CbNN outperform state of the art for the same memory footprint.
Hugues Wouafo, Cyrille Chavet, Philippe Coussy
IEEE Trans. Neural Networks Learn. Syst.3
2018 A Heterogeneous Cluster with Reconfigurable Accelerator for Energy Efficient Near-Sensor Data Analytics
abstract
IoT end-nodes require high performance and extreme energy efficiency to cope with complex near-sensor data analytics algorithms. Processing on multiple programmable processors operating in near-threshold is emerging as a promising solution to exploit the energy boost given by low-voltage operation, while recovering the related frequency degradation with parallelism. In this work, we present a heterogeneous cluster architecture extending a traditional parallel processor cluster with a reconfigurable Integrated Programmable Array (IPA) accelerator. While programmable processors guarantee programming legacy to easily manage peripherals, radio software stacks as well as the global program flow, offloading data-intensive and control-intensive kernels to the IPA leads to much higher system level performance and energy-efficiency. Experimental results show that the proposed heterogeneous cluster outperforms an 8-core homogeneous architecture by up to 4.8× in performance and 4.5× in energy efficiency when executing a mix of control-intensive and data-intensive kernels typical of near-sensor data analytics applications.
Satyajit Das, Kevin J. M. Martin, Philippe Coussy, Davide Rossi 0001
ISCAS3
2017 Efficient mapping of CDFG onto coarse-grained reconfigurable array architectures
abstract
In the approaching era of IoT, flexible and low power accelerators have become essential to meet aggressive energy efficiency targets. During the last few decades, Coarse Grain Reconfigurable Arrays (CGRA) have demonstrated high energy efficiency as accelerators, especially for high-performance streaming applications. While existing CGRAs mostly rely on partial and full predication techniques to support conditional branches, inefficient architecture and mapping support for handling control flow limits the use of CGRAs in accelerating either only inner loop bodies, or transformed loops specifically adapted to the target CGRA. This paper proposes a novel CGRA architecture with support for jump and conditional jump instructions and a lightweight global synchronization mechanism to enable complete Control Data Flow Graph (CDFG) mapping in an ultra-low-power environment. The architecture is coupled with a complete design flow that efficiently maps applications with heavy control flow starting from a generic C language description. The proposed mapping approach reduces the impact of wasteful instruction issues in the conventional approaches of predication providing an average energy improvement of 1.44× and 1.6× when compared to the state of the art partial and full predication techniques. Moreover, the proposed method achieves an average speed-up up to 21× and an energy improvement up to 50.42× while executing applications with heavy control flow with respect to sequential execution on a low-power embedded CPU, demonstrating its suitability for next generation IoT applications.
Satyajit Das, Kevin J. M. Martin, Philippe Coussy, Davide Rossi 0001, Luca Benini
ASP-DAC3
2017 A 142MOPS/mW integrated programmable array accelerator for smart visual processing
abstract
Due to increasing demand of low power computing, and diminishing returns from technology scaling, industry and academia are turning with renewed interest toward energy-efficient programmable accelerators. This paper proposes an Integrated Programmable-Array accelerator (IPA) architecture based on an innovative execution model, targeted to accelerate both data and control-flow parts of deeply embedded vision applications typical of edge-nodes of the Internet of Things (IoT). In this paper we demonstrate the performance and energy efficiency of IPA implementing a smart visual trigger application. Experimental results show that the proposed accelerator delivers 507 MOPS and 142 MOPS/mW on the target application, surpassing a low-power processor optimized for DSP applications by 6x in performance and by 10x in energy efficiency. Moreover, it surpasses performance of state of the art CGRAs only capable of implementing data-flow portion of applications by 1.6x, demonstrating the effectiveness of the proposed architecture and computational model.
Satyajit Das, Davide Rossi 0001, Kevin J. M. Martin, Philippe Coussy, Luca Benini
ISCAS4
2017 A Unified Design Flow to Automatically Generate On-Chip Monitors During High-Level Synthesis of Hardware Accelerators
abstract
Security and safety are more and more important in embedded system design. A key issue, hence lies in the ability of systems to respond safely when errors occur at runtime, to prevent unacceptable behaviors that can lead to failures or sensitive data leakage. In this paper, we propose a design approach that automatically generates on-chip monitors (OCMs) during high-level synthesis (HLS) of hardware accelerators (HWaccs). OCM checks at runtime the input/output timing behavior, the control flow execution and algorithmic properties (via American National Standards Institute C assertions) of the monitored HWacc. OCM is implemented separately from the HWacc and an original technique is introduced for their synchronization. Two synthesis options are proposed to tradeoff between performance and area. Experiment results show that error detection on the control flow is 16× better compared to the existing approaches while the cost of assertions is reduced by 17.48% on average. The impact on execution time (i.e., latency of the HWacc) is decreased by 2.76× at no area penalty and up to 4.5× with less than 10% extra-area. The clock period overhead is at worst less than 5% and the overhead on the synthesis time of the HWacc to generate OCMs is 7.44% on average.
Mohamed Ben Hammouda, Philippe Coussy, Loïc Lagadec
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 A dynamically reconfigurable ECC decoder architecture
Awais Sani, Philippe Coussy, Cyrille Chavet
DATE2
2015 In-place memory mapping approach for optimized parallel hardware interleaver architectures
Saeed Ur Reehman, Cyrille Chavet, Philippe Coussy, Awais Sani
DATE3
2015 Restricted Clustered Neural Network for Storing Real Data
abstract
Associative memories are an alternative to classical indexed memories that are capable of retrieving a message previously stored when an incomplete version of this message is presented. Recently a new model of associative memory based on binary neurons and binary links has been proposed. This model named Clustered Neural Network (CNN) offers large storage diversity (number of messages stored) and fast message retrieval when implemented in hardware. The performance of this model drops when the stored message distribution is non-uniform. In this paper, we enhance the CNN model to support non-uniform message distribution by adding features of Restricted Boltzmann Machines. In addition, we present a fully parallel hardware design of the model. The proposed implementation multiplies the performance (diversity) of Clustered Neural Networks by a factor of 3 with an increase of complexity of 40%.
Robin Danilo, Philippe Coussy, Laura Conde-Canencia, Vincent Gripon, Warren J. Gross
ACM Great Lakes Symposium on VLSI2
2015 Algorithm and implementation of an associative memory for oriented edge detection using improved clustered neural networks
abstract
Associative memories are capable of retrieving previously stored patterns given parts of them. This feature makes them good candidates for pattern detection in images. Clustered Neural Networks is a recently-introduced family of associative memories that allows a fast pattern retrieval when implemented in hardware. In this paper, we propose a new pattern retrieval algorithm that results in a dramatically lower error rate compared to that of the conventional approach when used in oriented edge detection process. This function plays an important role in image processing. Furthermore, we present the corresponding hardware architecture and implementation of the new approach in comparison with a conventional architecture in literature, and show that the proposed architecture does not significantly affect hardware complexity.
Robin Danilo, Hooman Jarollahi, Vincent Gripon, Philippe Coussy, Laura Conde-Canencia, Warren J. Gross
ISCAS4
2015 Improving storage of patterns in recurrent neural networks: Clone-based model and architecture
abstract
Artificial neural networks are used in various domains like computer science and computer engineering for tasks like image processing or design of associative memories. The goal is to mimic the impressive brain ability to process or to memorize and retrieve information. Recently a new model of neural network has been proposed and can be used to design associative memories. When considering patterns that are uniformly distributed, this model outperforms existing models like Hopfield Networks. However, when considering non-uniformly distributed patterns, its performance highly degrades. Few propositions have been made to address this problem. However, they require designing complex hardware architectures to be efficient. In this paper, we propose a new binary neural network model that allows reaching good performances at low hardware cost.
Hugues Wouafo, Cyrille Chavet, Philippe Coussy
ISCAS3
2015 Fully Binary Neural Network Model and Optimized Hardware Architectures for Associative Memories
abstract
Brain processes information through a complex hierarchical associative memory organization that is distributed across a complex neural network. The GBNN associative memory model has recently been proposed as a new class of recurrent clustered neural network that presents higher efficiency than the classical models. In this article, we propose computational simplifications and architectural optimizations of the original GBNN. This work leads to significant complexity and area reduction without affecting neither memorizing nor retrieving performance. The obtained results open new perspectives in the design of neuromorphic hardware to support large-scale general-purpose neural algorithms.
Philippe Coussy, Cyrille Chavet, Hugues Wouafo, Laura Conde-Canencia
ACM J. Emerg. Technol. Comput. Syst.1
2015 Large-Scale Spiking Neural Networks using Neuromorphic Hardware Compatible Models
abstract
Neuromorphic engineering is a fast growing field with great potential in both understanding the function of the brain, and constructing practical artifacts that build upon this understanding. For these novel chips and hardware to be useful, hardware compatible applications and simulation tools are needed. We argue that the neural circuit approach, in which networks of neuronal elements model brain circuitry are constructed, allows the development of practical applications and the exploration of brain function. At this level of abstraction, networks of 10 5 neurons or larger can be efficiently simulated, but still preserve the neuronal and synaptic dynamics that appear to be important for brain function. Because the neural circuit level supports spiking neural networks and the prevalent Addressable Event Representation (AER) communication scheme, it fits well with many existing neuromorphic hardware and simulation tools. To show how this approach can be applied, we present case studies of spiking neural networks in vision and recognition tasks based on one instantiation of a simulation environment. However, there are now many hardware options, simulation environments, and applications in this emerging field. These approaches and other considerations are discussed.
Jeffrey L. Krichmar, Philippe Coussy, Nikil Dutt
ACM J. Emerg. Technol. Comput. Syst.2
2014 Efficient application mapping on CGRAs based on backward simultaneous scheduling/binding and dynamic graph transformations
abstract
Mapping an application on a coarse grained reconfigurable architecture (CGRA) is a complex task which is still often completely or partially realized manually. This paper presents an automated synthesis flow based on simultaneous scheduling and binding steps. The proposed method uses a backward traversal of the formal model obtained after compilation and dynamically transforms it when needed. Our approach is compared with state of the art techniques and its interest is shown through the mapping of several applications from digital signal and image processing domain.
Thomas Peyret, Gwenolé Corre, Mathieu Thevenin, Kevin J. M. Martin, Philippe Coussy
ASAP5
2014 A tightly-coupled hardware controller to improve scalability and programmability of shared-memory heterogeneous clusters
abstract
Modern designs for embedded many-core systems increasingly include application-specific units to accelerate key computational kernels with orders-of-magnitude higher execution speed and energy efficiency compared to software counterparts. A promising architectural template is based on heterogeneous clusters, where simple RISC cores and specialized HW units (HWPU) communicate in a tightly-coupled manner via L1 shared memory. Efficiently integrating processors and a high number of HW Processing Units (HWPUs) in such an system poses two main challenges, namely, architectural scalability and programmability. In this paper we describe an optimized Data Pump (DP) which connects several accelerators to a restricted set of communication ports, and acts as a virtualization layer for programming, exposing FIFO queues to offload “HW tasks” to them through a set of lightweight APIs. In this work, we aim at optimizing both these mechanisms, for respectively reducing modules area and making programming sequence easier and lighter.
Paolo Burgio, Robin Danilo, Andrea Marongiu, Philippe Coussy, Luca Benini
DATE4
2014 A HLS-Based Toolflow to Design Next-Generation Heterogeneous Many-Core Platforms with Shared Memory
abstract
This work describes how we use High-Level Synthesis to support design space exploration (DSE) of heterogeneous many-core systems. Modern embedded systems increasingly couple hardware accelerators and processing cores on the same chip, to trade specialization of the platform to an application domain for increased performance and energy efficiency. However, the process of designing such a platform is complex and error-prone, and requires skills on algorithmic aspects, hardware synthesis, and software engineering. DSE can partially be automated, and thus simplified, by coupling the use of HLS tools and virtual prototyping platforms. In this paper we enable the design space exploration of heterogeneous many-cores adopting a shared-memory architecture template, where communication and synchronization between the hardware accelerators and the cores happens through L1 shared memory. This communication infrastructure leverages a "zero-copy" scheme, which simplifies both the design process of the platform and the development of applications on top of it. Moreover, the shared-memory template perfectly fits the semantics of several high-level programming models, such as OpenMP. We provide programmers with simple yet powerful abstractions to exploit accelerators from within an OpenMP application, and propose a low-cost implementation of the necessary runtime support. An HLS-based automatic design flow is set up, to quickly explore the design space using a cycle-accurate virtual platform.
Paolo Burgio, Andrea Marongiu, Philippe Coussy, Luca Benini
EUC3
2014 A design approach to automatically generate on-chip monitors during high-level synthesis of hardware accelerator
abstract
Embedded systems often implement safety critical applications making security a more and more important aspect in their design. Control-Flow Integrity (CFI) attacks are used to modify program behavior and can lead to learn valuable information directly or indirectly by perturbing a system and creating failures. Although CFI attacks are well-known in computer systems, they have been recently shown to be practical and feasible on embedded systems as well. In this context, CFI checks are mainly used to detect unintended software behaviors while very few works address non programmable hardware component monitoring. In this paper, we present a hardware-assisted paradigm to enhance embedded system security by detecting and preventing unintended hardware behavior. We propose a design approach that designs on-chip monitors (OCM) during High-Level Synthesis (HLS) of hardware accelerators (HWacc). Synthesis of OCM is introduced as a set of steps realized concurrently to the HLS flow of HWacc. Automatically generated OCM checks at runtime both the input/output timing behavior and the control flow of the monitored HWacc. Experimental results show the interest of the proposed approach: the error coverage on the control flow ranges from 99.75% to 100% while in average the OCM area overhead is less than 10%, the clock period overhead is at worst less than 5% and impact on the synthesis time is negligible.
Mohamed Ben Hammouda, Philippe Coussy, Loïc Lagadec
ACM Great Lakes Symposium on VLSI2
2014 An automated design approach to map applications on CGRAs
abstract
Coarse-Grained Reconfigurable Architectures (CGRAs) are promising high-performance and power-efficient platforms. However, their uses are still limited by the capability of mapping tools. This abstract paper outlines a new automated design flow to map applications on CGRAs. The interest of our method is shown through comparison with state of the art approaches.
Thomas Peyret, Gwenolé Corre, Mathieu Thevenin, Kevin J. M. Martin, Philippe Coussy
ACM Great Lakes Symposium on VLSI5
2014 A memory mapping approach based on network customization to design conflict-free parallel hardware architectures
abstract
Parallel hardware architectures are needed to achieve high throughput systems. Unfortunately, efficient parallel architectures often require removing memory access conflicts. This is particularly true when designing turbo-codes, channel interleaver or LDPC (Low Density Parity Check) codes architectures which are one of the most critical parts of parallel decoders. Many solutions are proposed in state of the art to find conflict free memory mapping but they are either limited to a subset of constraints, or result in high architectural cost. These drawbacks come from the interleaving law and the incompatibility between this law and the targeted interconnection network (in the coder/encoder architecture). In this paper we propose a conflict free memory mapping approach that is able to generate optimized hardware architectures by limiting these drawbacks. The proposed solution constructs a customized interconnection network by analyzing data access patterns defined in the interleaving law. Our approach is then compared to state of the art methods and its interest is shown through the design of parallel interleavers for HSPA.
Saeed Ur Reehman, Cyrille Chavet, Philippe Coussy
ACM Great Lakes Symposium on VLSI3
2014 Embedding polynomial time memory mapping and routing algorithms on-chip to design configurable decoder architectures
abstract
To fulfill the high data rate requirement of current telecommunication standards, error-correction codes decoders are implemented on parallel architectures leading to memory conflict problem. Different memory mapping approaches are proposed in the literature to solve this problem. However, these approaches can only be executed offline due to their computational complexity and resultant memory mapping is stored in dedicated ROM in order to drive the network for a particular block length. Unfortunately, to support several block lengths, multiple ROMs are required which results in huge hardware cost. In this article, we propose a novel online memory mapping architecture that consists of online mapping generator and RAM to support multiple block lengths on single chip. Online mapping generator performs two functions: First, it executes polynomial time memory mapping algorithm online and secondly, it generates command words for Benes network by using a simplified routing algorithm. Whenever new block length needs to be decoded, online mapping generator outputs addressing and command words at runtime to update the RAM. Experimental results show that significant reduction in time and memory cost is obtained while implementing polynomial time memory mapping algorithm on-chip as compared to state of the art approaches.
Saeed-ur Rehman, Awais Sani, Cyrille Chavet, Philippe Coussy
ICASSP4
2014 A design approach to automatically synthesize ANSI-C assertions during High-Level Synthesis of hardware accelerators
abstract
Evolution of Systems-On-Chip (SoC) increases the challenge of verification and post-silicon debug. Nowadays, Assertion Based Verification (ABV) is a widely used methodology. Languages like PSL (Property Specification Language) or SVA (System Verilog Assertions) allows engineers to define properties at Register Transfer Level (RTL). Properties can then be used to generate simulation/hardware assertion checkers for dynamic verification. In this paper, we propose to consider ANSI-C assertions during High-Level Synthesis (HLS) of hardware accelerators (HWacc) to automatically generate on-chip monitors (OCM). The proposed method is portable to any HLS tool and supports both static and dynamic application behaviors. OCM is implemented separately from the HWacc and an original technique is introduced for their synchronization. Two synthesis options are proposed for the OCM design i.e. speed and area. Experimental results show the interest of the proposed approach: while the cost of the OCMs mainly depends on the complexity of input assertions, setting synthesis option is area allows reducing the complexity of the OCM by 2.37x on average compared to the option for speed optimization.
Mohamed Ben Hammouda, Philippe Coussy, Loïc Lagadec
ISCAS2
2013 Dynamic branch prediction for high-level synthesis
abstract
Branch prediction is a widely used technique to optimize performances of pipelined microprocessor architectures. In High-Level Synthesis (HLS) domain, few synthesis techniques for optimizing control flows of data dominated applications have been proposed. Previous works mainly focus on using techniques like path-based scheduling algorithms, speculation techniques or static branch prediction for pipelined loops. In this paper, we present a synthesis flow that combines dynamic branch prediction and operation speculation to remove performance bottlenecks imposed by the control flow of applications. Interest of the proposed approach is shown in term of latency improvements and area overhead through a set of experiments.
Vianney Lapotre, Philippe Coussy, Cyrille Chavet, Hugues Wouafo, Robin Danilo
FPL2
2013 A memory mapping approach for network and controller optimization in parallel interleaver architectures
abstract
Recent communication standards and storage systems uses parallel architectures for error correcting codes (LDPC or Turbo-codes) to reliably transfer data between two equipments. However, parallel architectures suffer from memory access conflicts. In this paper, we present a method that finds a conflict-free memory mapping for any interleaving law and any parallelism. The proposed approach always complies with the interconnection network topology the designer wants to infer. Moreover, the resulting architecture is optimized by reducing the cost of network and controller (network and memory controller) architectures.
Aroua Briki, Cyrille Chavet, Philippe Coussy
ACM Great Lakes Symposium on VLSI3
2013 On-chip implementation of memory mapping algorithm to support flexible decoder architecture
abstract
Parallel hardware architectures are used to design turbo-like iterative decoders to meet the requirement of high data rate applications. However, parallel architectures suffer from memory conflict problem due to interleaving law used in turbo-like codes. To solve conflict problem, different memory mapping approaches have been developed. These methods automatically generate a set of control words stored in ROM to drive the architecture. These approaches are used off-chip by the designer (i.e. prior the decoder implementation) to generate different set of control words i.e. one set for each block length used in the target telecommunication standard. This requires multiple ROMs to store mapping information for multiple block lengths and results in huge hardware cost. In this article, we propose to embed memory mapping algorithms on-chip. Hence, each time word-length changes, memory mapping algorithm is executed. Command words are thus generated at runtime and stored in a RAM. This is a first attempt to embed mapping algorithms on chip and experimental results show that a significant amount of memory can be saved by using on-chip execution of mapping algorithms. Results also highlight that improvement in design and implementation of mapping algorithms are still needed to embed mapping algorithms on-chip to implement flexible decoder architectures.
Saeed-ur Rehman, Awais Sani, Philippe Coussy, Cyrille Chavet
ICASSP3
2012 OpenMP-based Synergistic Parallelization and HW Acceleration for On-Chip Shared-Memory Clusters
abstract
Modern embedded MPSoC designs increasingly couple hardware accelerators to processing cores to trade between energy efficiency and platform specialization. To assist effective design of such systems there is the need on one hand for clear methodologies to streamline accelerator definition and instantiation, on the other for architectural templates and run-time techniques that minimize processors-to-accelerator communication costs. In this paper we present an architecture featuring tightly-coupled processors and accelerators, with zero-copy communication. Efficient programming is supported by an extended OpenMP programming model, where custom directives allow to specialize code regions for execution on parallel cores, accelerators, or a mix of the two. Our integrated approach enables fast yet accurate exploration of accelerator-based HW and SW architectures.
Paolo Burgio, Andrea Marongiu, Dominique Heller, Cyrille Chavet, Philippe Coussy, Luca Benini
DSD5
2012 A design approach dedicated to network-based and conflict-free parallel interleavers
abstract
For high throughput applications, efficient parallel architectures require to avoid collision accesses, i.e. concurrent read/write accesses to the same memory bank have to be avoided. This consideration applies for example to the two main classes of turbo-like codes that are Low Density Parity Check (LDPC) and Turbo-Codes. These error correcting codes, that scramble data by using an interleaving law, are used in most of recent communication standards and storage systems like wireless access, digital video broadcasting or magnetic storage in hard disk drives. In order to optimize the architectural cost and to reduce the control complexity of such integrated circuits, designers usually use standard interconnection networks with low complexity topologies between processing elements and memory banks. However the design constraints, i.e. interleaving law, parallelism and interconnection network, often prevent mapping the data in the memory banks without any conflict. In this paper we propose a methodology which always finds a collision-free memory mapping for a given set of design constraints. The approach uses additional registers each time the design constraints forbid to use memory banks without conflict. Our approach is compared to state of the art methods and its interest is shown through the design of parallel interleavers for industrial applications: Multi Band-Orthogonal Frequency-Division Multiplexing Ultra-WideBand (MB-OFDM UWB) and non-binary LDPC decoders.
Aroua Briki, Cyrille Chavet, Philippe Coussy, Eric Martin 0001
ACM Great Lakes Symposium on VLSI3
2011 A methodology based on Transportation problem modeling for designing parallel interleaver architectures
abstract
For high-data-rate applications, turbo-like iterative decoders are implemented with parallel hardware architecture. However, to achieve high throughput, concurrent accesses to each memory bank has to be performed without any conflict. The consideration applies to the two main classes of turbo-like codes: Low Density Parity Check (LDPC) and Turbo-Codes. In this paper, we present an original approach based on Transportation problem modeling which finds conflict free memory mapping for every type of turbo codes and which optimizes the resulting interleaving architecture.
Awais Sani, Philippe Coussy, Cyrille Chavet, Eric Martin 0001
ICASSP2
2011 An approach based on edge coloring of tripartite graph for designing parallel LDPC interleaver architecture
abstract
A practical and feasible solution for LDPC decoder is to design partially-parallel hardware architecture. These architectures are efficient in terms of area, cost, flexibility and performances. However, this type of architecture is complex to design since concurrent read and write accesses to data have to be performed at each time instance without any conflict. To solve this memory mapping problem, we present in this paper, an original approach based on a tripartite graph modeling and a modified edge coloring algorithm to design parallel LDPC interleaver architecture.
Awais Sani, Philippe Coussy, Cyrille Chavet, Eric Martin 0001
ISCAS2
2010 Hierarchical and Multiple-Clock Domain High-Level Synthesis for Low-Power Design on FPGA
abstract
Power optimization has become one of the most challenging design objectives of modern digital systems. Although FPGAs are more and more used, they are however still considered as power inefficient compared to standard-cell or full-custom technologies. New dedicated design approaches are thus needed to reduce this gap. In this paper, we address low-power design on FPGA through a dedicated High-Level Synthesis (HLS) flow. The proposed approach allows to slow down the clock frequency in parts of the design, decrease the complexity of the clock-network, reduce the number of long wires and perform clock-gating. The design flow has been fully implemented and allows to automatically synthesize hierarchical and synchronous multiple-clock domain architectures. The power consumption of the architectures we generate has been investigated and compared with state-of-the-art synthesis approaches. The experiments have been realized by using a Xilinx Virtex-5 device and the power measurement results show the interest of the proposed approach.
Ghizlane Lhairech-Lebreton, Philippe Coussy, Eric Martin 0001
FPL2
2010 Static Address Generation Easing: a design methodology for parallel interleaver architectures
abstract
For high throughput applications, turbo-like iterative decoders are implemented with parallel architectures. However, to be efficient parallel architectures require to avoid collision accesses i.e. concurrent read/write accesses should not target the same memory block. This consideration applies to the two main classes of turbo-like codes which are Low Density Parity Check (LDPC) and Turbo-Codes. In this paper we propose a methodology which finds a collision-free mapping of the variables in the memory banks and which optimizes the resulting interleaving architecture. Finally, we show through a pedagogical example the interest of our approach compared to state-of-the-art techniques.
Cyrille Chavet, Philippe Coussy, Pascal Urard, Eric Martin 0001
ICASSP2
2010 A memory mapping approach for parallel interleaver design with multiples read and write accesses
abstract
For high throughput applications, turbo-like iterative decoders are implemented with parallel architectures. However, to be efficient parallel architectures require to avoid collision accesses i.e. concurrent read/write accesses should not target the same memory block. This consideration applies to the two main classes of turbo-like codes which are Low Density Parity Check (LDPC) and Turbo-Codes. In this paper we propose a methodology which always finds a collision-free mapping of the variables in the memory banks and which optimizes the resulting interleaving architecture. Finally, we show through a pedagogical example the interest our approach. This research was supported by the European project DAVINCI.
Cyrille Chavet, Philippe Coussy
ISCAS2
2010 High-Level Synthesis for Designing Multimode Architectures
abstract
This paper addresses the design of multimode architectures for digital signal and image processing applications. We present a dedicated design flow and its associated high-level synthesis tool, named GAUT. Given a unified description of a set of time-wise mutually exclusive tasks and their associated throughput constraints, a single register transfer level hardware architecture optimized in area is generated. In order to reduce the register, the steering logic, and the controller complexities, this paper proposes a joint-scheduling algorithm, which maximizes the similarities between the control steps and specific binding approaches for both operators and storage elements which maximize the similarities between the datapaths. It is shown through a set of test cases that the proposed approach offers significant area saving and low-performance penalties compared to both state-of-the-art techniques and dedicated mono-mode architectures.
Caaliph Andriamisaina, Philippe Coussy, Emmanuel Casseau, Cyrille Chavet
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 A design methodology for space-time adapter
abstract
This paper presents a solution to efficiently explore the design space of communication adapters. In most digital signal processing (DSP) applications, the overall architecture of the system is significantly affected by communication architecture, so the designers need specifically optimized adapters. By explicitly modeling these communications within an effective graph-theoretic model and analysis framework, we automatically generate an optimized architecture, named Space-Time AdapteR (STAR). Our design flow inputs a C description of Input/Output data scheduling, and user requirements (throughput, latency, parallelism&), and formalizes communication constraints through a Resource Constraints Graph (RCG). The RCG properties enable an efficient architecture space exploration in order to synthesize a STAR component. The proposed approach has been tested to design an industrial data mixing block example: an Ultra-Wideband interleaver.
Cyrille Chavet, Philippe Coussy, Pascal Urard, Eric Martin 0001
ACM Great Lakes Symposium on VLSI2
2007 A design flow dedicated to multi-mode architectures for DSP applications
abstract
This paper addresses the design of multi-mode architectures for digital signal processing applications. We present a dedicated design flow and its associated high-level synthesis tool, named GAUT. Given a unified description of a set of time-wise mutually exclusive tasks and their associated throughput constraints, a single RTL hardware architecture optimized in area is generated. In order to reduce the register, steering logic (multiplexers) and controller (decoding logic) complexities, we propose a joint-scheduling algorithm which maximizes the similarities between control steps and specific binding approaches for both functional units and storage elements which maximize the similarities between the datapaths. We show through a set of test cases that our approach offers significant area saving relative to the state-of-the-art.
Cyrille Chavet, Caaliph Andriamisaina, Philippe Coussy, Emmanuel Casseau, Emmanuel Juin, Pascal Urard, Eric Martin 0001
ICCAD3
2007 A Methodology for Efficient Space-Time Adapter Design Space Exploration: A Case Study of an Ultra Wide Band Interleaver
abstract
This paper presents a solution to efficiently explore the design space of communication adapters. In most digital signal processing (DSP) applications, the overall architecture of the system is significantly affected by communication architecture, so the designers need specifically optimized adapters. By explicitly modeling these communications within an effective graph-theoretic model and analysis framework, we automatically generate an optimized architecture, named Space-Time AdapteR (STAR). Our design flow inputs a C description of Input/Output data scheduling, and user requirements (throughput, latency, parallelism...), and formalizes communication constraints through a Resource Constraints Graph (RCG). The RCG properties enable an efficient architecture space exploration in order to synthesize a STAR component. The proposed approach has been tested to design an industrial data mixing block example: an Ultra-Wideband interleaver.
Cyrille Chavet, Philippe Coussy, Pascal Urard, Eric Martin 0001
ISCAS2
2007 Constrained algorithmic IP design for system-on-chip
Philippe Coussy, Emmanuel Casseau, Pierre Bomel, Adel Baganne, Eric Martin 0001
Integr.1
2006 A formal method for hardware IP design and integration under I/O and timing constraints
abstract
IP integration, which is one of the most important SoC design steps, requires taking into account communication and timing constraints. In that context, design and reuse can be improved using IP cores described at a high abstraction level. In this paper, we present an IP design approach that relies on three main phases: (1) constraint modeling, (2) IP constraint analysis steps for feasibility checking, and (3) synthesis. We propose a set of techniques dedicated to the digital signal processing domain that lead to an optimized IP core integration. Based on a generic architecture of components, the method we propose provides automatic generation of IP cores designed under integration constraints. We show the effectiveness of our approach with a DCT core design case study.
Philippe Coussy, Emmanuel Casseau, Pierre Bomel, Adel Baganne, Eric Martin 0001
ACM Trans. Embed. Comput. Syst.1
2005 SystemCmantic: A high level Modelling and Co-Design Framework
Lobna Kriaa, S. Adriano, Emmanuel Vaumorin, R. Nouacer, F. Blanc, S. Pajaniardja, Philippe Coussy, Eric Martin 0001, Dominique Heller, Farhat Thabet, Anne-Marie Fouilliart
FDL7
2005 A more efficient and flexible DSP design flow from Matlab-Simulink [FFT algorithm example]
abstract
The design of complex digital signal processing systems implies to minimize architectural cost and to maximize timing performances while taking into account communication and memory access constraints for the integration of dedicated hardware accelerators. Unfortunately, the traditional Matlab/Simulink design flows gather not very flexible hardware blocks. In this paper, we present a methodology and a tool that permit the high-level synthesis of DSP applications, under both I/O timing and memory constraints. Based on formal models and a generic architecture, this tool helps the designer in finding a reasonable trade-off between the circuit's latency and its architectural complexity. The efficiency of our approach is demonstrated on the case study of an FFT algorithm.
Philippe Coussy, Gwenolé Corre, Pierre Bomel, Eric Senn, Eric Martin 0001
ICASSP (5)1
2004 A methodology for IP integration into DSP SoC: a case study of a MAP algorithm for turbo decoder
abstract
The re-use of complex digital signal processing (DSP) coprocessors can be improved using IP cores described at a high abstraction level. System integration, which is a major step in SoC design, requires taking into account communication and timing constraints to design and integrate IP. In this paper, we describe an IP design approach that relies on three main phases: constraints modeling, IP constraints analysis steps for feasibility checking, and synthesis. Based on a generic architecture, the presented method provides automatic generation of IP cores designed under integration constraints. We show the effectiveness of our approach in a case study of a maximum a posteriori (MAP) algorithm for a turbo decoder.
Philippe Coussy, David Gnaedig, Amor Nafkha, Adel Baganne, Emmanuel Boutillon, Eric Martin 0001
ICASSP (5)1
2003 Communication and Timing Constraints Analysis for IP Design and Integration
Philippe Coussy, Adel Baganne, Eric Martin 0001
VLSI-SOC1