VLDB 2026 Research / reviewers in the wild / expert
Amer Baghdadi
dblp:30/4537
· DBLP profile ↗
47ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-6181-6500ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 18 · 2 first-author · 2 since 2021Computer networks · 9 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MOR-GenREH: A Multi-Stage Nonlinear Rectifier Model with Generalized Statistical Rectification for Radio Frequency Energy HarvestingabstractInternational audience Fortunatus Aabangbio Wulnye, Syma Afsha, Richard Boateng Nti, Amer Baghdadi, Sébastien Roy 0002, Samir Saoudi, Derek Kwaku Pobi Asiedu |
ICC | 4 |
| 2026 | ML-Based Hardware Trojan Detection in AI Accelerators via Power Side-Channel AnalysisabstractTo accelerate development and system integration, many companies opt to outsource the design of complex AI accelerators to third-party IP vendors rather than developing them inhouse. This practice raises security concerns, particularly the risk of hardware Trojan (HT) attacks [1]. Traditional testing methods are impractical for detecting HTs in modern AI/ML accelerators due to their hardware complexity and inability to provide insights into the inserted Trojans. In this work, we propose a methodology to detect the presence of HTs in different AI/ML accelerators and identify key Trojan characteristics using power side-channel analysis (PSCA). We present a testbed for accurate power consumption measurement, enabling the collection of real ML inference power traces. We inserted multiple HTs in several AI/ML accelerators and prepared various dataset configurations to analyze multiple HT scenarios. Then, we proposed a novel method for preprocessing PSCA data that consists of segmenting the power traces and extracting statistical features from time and frequency domains. Our proposed technique, equipped with an ML-based HT detection and identification method, achieves up to 99% accuracy. AbdelSalam Baltagi, Yehya Nasser, Amer Baghdadi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | Orthogonal Filtered Delay-Doppler Multiplexing (OF2DM) With Enhanced Robustness to Doppler and Timing Offsets
Fatima Hamdar, Jérémy Nadal, Charbel Abdel Nour, Amer Baghdadi |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | Multi-Carrier MIMO Cognitive BackCom: Adaptive Processing with Temporal Combining Improved Data Stream DecodingabstractBackscatter communication has evolved, but the research interest has predominantly focused on a single carrier radio frequency source. In this paper, we propose a joint Repetition Coding (RC) and Multi-Source Temporal Combining (TC) framework to enhance the reliability and throughput of backscatter communication. Unlike conventional approaches that rely on a single carrier source or uniform repetition coding, our method dynamically assigns position-aware repetition codes while leveraging multiple ambient RF sources to exploit temporal diversity. We developed an analytical model to evaluate the bit error rate (BER) performance under different coding strategies and source distributions. Furthermore, we derive a closed-form optimization framework for adaptive weight allocation in TC, ensuring that the contribution of each RF source is weighted based on SNR, coherence time, and estimated BER. Richard Boateng Nti, Derek Kwaku Pobi Asiedu, Amer Baghdadi, Samir Saoudi, Ji-Hoon Yun, Sébastien Roy 0002 |
GLOBECOM | 3 |
| 2025 | Efficient Filter-Bank Pilot Structures for Multi-User Massive MIMO SystemsabstractThis work proposes a novel pilot structure (PS) to improve channel estimation (CE) in massive multiple-input multiple-output filter bank multi-carrier (mMIMO-FBMC) communications using offset quadrature amplitude modulation (OQAM). Accurate CE is essential to fully benefit from the corresponding spectral efficiency (SE) advantages and robustness against dispersive channels. Unlike orthogonal frequency division multiplexing (OFDM), FBMC/OQAM shows only real-field orthogonality. Therefore, corresponding signals suffer from intrinsic interference, motivating the use of alternative CE techniques. Furthermore, typical CE methods for mMIMO-FBMC systems require large training overheads, particularly for a large number of users or for flat-fading channel conditions over each subcarrier band. To address these issues, this paper proposes a PS that reduces the training overhead by interleaving user pilots in frequency while using conventional OFDM CE techniques, particularly the least squares (LS) method in the context of single-user (SU) and multi-user (MU) scenarios. Analytical expressions for the mean square error (MSE) and the Cramer-Rao lower bound (CRLB) are derived for the estimation technique considering the proposed PSs. Furthermore, performed simulations over the 5G QuaDRiga channel show that the proposed PS improves SE while delivering optimal performance in both preamble-based MU-mMIMO and SU systems. Fatima Hamdar, Jérémy Nadal, Charbel Abdel Nour, Amer Baghdadi |
IEEE Trans. Commun. | 4 |
| 2024 | Overlap-Save FBMC Receivers for Massive MIMO SystemsabstractMassive multiple-input multiple-output (mMIMO) systems and filtered multi-carrier waveforms have recently emerged as hot research topics in next-generation wireless networks. This paper extends the use of Overlap-Save Filter Bank Multi-Carrier (FBMC) receivers to mMIMO communication systems and investigates the corresponding FBMC transceiver advantages over OFDM. To this aim, we derived the signal-to-interference ratio (SIR) expressions analytically under several channel impairments such as timing offsets and carrier frequency offsets. Furthermore, we conducted an asymptotic study on the performance of FBMC augmented with our proposed receivers in the context of massive MIMO systems. While still outperforming OFDM, results for FBMC show that increasing the number of base station (BS) antennas does not increase the SINR unboundedly. We have validated the proposed analytical study by comparing its results with those obtained from Monte Carlo simulations and confirmed the superiority of the proposed receivers over those in mMIMO literature. Moreover, analytical and simulation results confirm that the Overlap-Save FBMC receiver can support asynchronous communications in the context of mMIMO, a cornerstone for grant-free communications and massive access. Fatima Hamdar, Jérémy Nadal, Charbel Abdel Nour, Amer Baghdadi |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Novel transmission technique based on intentional overlapping for spectral efficiency enhancement in multicarrier systemsabstractA crucial element of cellular communication networks is the achievable spectral efficiency (SE). Indeed, motivated by the continuous growth in mobile data traffic in wireless networks, improving SE has represented one of the key goals of the 3rd Generation Partner Project (3GPP) along several generations of communication standards. In this paper, we propose a novel transmission scheme based on intentionally overlapping the subcarriers of adjacent users considering two state-of-the-art waveforms: the orthogonal frequency division multiplexing (OFDM) and the filter bank multi-carrier with offset quadrature amplitude modulation (FBMC/OQAM). We investigate the achievable spectral efficiency improvement of the proposed transmission scheme in the presence and absence of timing offset impairments under the tapped delay line B (TDL-B) multipath 3GPP model of the 5G QuaDRiGa channel. According to the test results, the proposed transmission technique achieves up to 20% SE improvement for FBMC/OQAM when associated to the overlap-save FBMC receiver and 12% for OFDM. Fatima Hamdar, Jérémy Nadal, Charbel Abdel Nour, Amer Baghdadi |
PIMRC | 4 |
| 2022 | Marine Objects Detection Using Deep Learning on Embedded Edge DevicesabstractArtificial Intelligence techniques based on convolution neural networks (CNNs) are now dominant in the field of object detection and classification. The deployment of CNNs on embedded edge devices targeting real-time inference sets a challenge due to the limited computing resources and power budgets. Several optimization techniques such as pruning, quantization and use of light neural networks enable the real-time inference but at the cost of precision degradation. However, using efficient approaches to apply the optimization techniques at training and inference stages enable high inference speed with limited degradation of detection performance. In this paper, we revisit the problem of detecting and classifying maritime objects. We investigate different versions of the You Only Look Once (YOLO), a state-of-the-art deep neural network, for real-time object detection and compare their performance for the specific application of detecting maritime objects. The trained YOLO networks are efficiently optimized targeting three recent edge devices: Nvidia Jetson Xavier AGX, AMD-Xilinx Kria KV260 Vision AI Kit, and Movidius Myriad X VPU. The proposed deployments demonstrate promising results with an inference speed of 90 FPS and a limited degradation of 2.4% in mean average precision. Dominique Heller, Mostafa Rizk, R. Douguet, Amer Baghdadi, Jean-Philippe Diguet |
RSP | 4 |
| 2022 | Enhancing embedded AI-based object detection using multi-view approachabstractObject detection based on convolutional neural network (CNN) is widely used in multitude emergent applications. Yet, the deployment of CNNs on embedded devices at the edge with reduced resources and power budget poses a real challenge. In this paper, we address this issue by enhancing the detection performance without impacting the inference speed. We investigate the use of multi-view for the same scene to achieve better detection performance. A novel system of distributed smart cameras is proposed where each camera integrates a CNN for detection. Implementation results show that using light networks on the distributed cameras can lead to better detection performance and a reduction in the overall consumed power. Zijie Ning, Mostafa Rizk, Amer Baghdadi, Jean-Philippe Diguet |
RSP | 3 |
| 2022 | Overlap-Save FBMC receivers for massive MIMO systems under channel impairmentsabstractMassive MIMO and filtered multi-carrier waveforms are considered as key enabling technologies for next-generation wireless networks. In this work, the Filter Bank Multi-Carrier (FBMC) waveform solution applying our previously proposed short filter and advanced receivers is extended to massive MIMO systems and evaluated in comparison to OFDM. Simulation results are presented for the non-line of sight (NLOS) 3D Urban-Macrocell (UMa) model of the 5G QuaDRiGa channel. Results show that the solution applying the proposed Overlap-Save (OS) and Overlap-Save-Block FBMC (OSB) receivers out-performs OFDM under timing offsets, carrier frequency offsets and Doppler spreads. Moreover, they confirm that the Overlap-Save FBMC receiver can support asynchronous communications in the context of massive MIMO, a cornerstone for grant-free communications and massive access. Fatima Hamdar, Jérémy Nadal, Charbel Abdel Nour, Amer Baghdadi |
VTC Spring | 4 |
| 2022 | MOL-Based In-Memory Computing of Binary Neural NetworksabstractConvolutional neural networks (CNNs) have proven very effective in a variety of practical applications involving artificial intelligence (AI). However, the layer depth of CNN deepens as user applications become more sophisticated, resulting in a huge number of operations and increased memory size. The massive amount of the produced intermediate data leads to intensive data movement between memory and computing cores causing a real bottleneck. In-memory computing (IMC) aims to address this bottleneck by directly computing inside memory, eliminating energy-intensive and time-consuming data movement. On the other hand, the emerging binary neural networks (BNNs), which is a special case of CNN, show a number of hardware-friendly properties, including memory saving. In BNN, the costly floating-point multiply-and-accumulate is replaced with lightweight bitwise XNOR and popcount operations. In this article, we propose an IMC programmable architecture targeting efficient implementation of BNN. Computational memories based on the recently introduced memristor overwrite logic (MOL) design style are employed. The architecture, which is presented in semiparallel and parallel models, efficiently executes the advanced quantization algorithm of XNOR-Net BNN. Performance evaluation based on the CIFAR-10 dataset demonstrates between$1.24\times $and$3\times $speedup and 49% and 99% energy saving compared to state-of-the-art implementations and up to 273-image/s/W throughput efficiency. Khaled Alhaj Ali, Amer Baghdadi, Elsa Dupraz, Mathieu Léonardon, Mostafa Rizk, Jean-Philippe Diguet |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Parallel and Flexible 5G LDPC Decoder Architecture Targeting FPGAabstractThe quasi-cyclic (QC) low-density parity-check (LDPC) code is a key error correction code for the fifth generation (5G) of cellular network technology. Designed to support several frame sizes and code rates, the 5G LDPC code structure allows high parallelism to deliver the high demanding data rate of 10 Gb/s. This impressive performance introduces challenging constraints on the hardware design. Particularly, allowing such high flexibility can introduce processing rate penalties on some configurations. In this context, a novel highly parallel and flexible hardware architecture for the 5G LDPC decoder is proposed, targeting field-programmable gate array (FPGA) devices. The architecture supports frame parallelism to maximize the utilization of the processing units, significantly improving the processing rate. The controller unit was carefully designed to support all 5G configurations and to avoid update conflicts. Furthermore, an efficient data scheduling is proposed to increase the processing rate. Compared to the recent related state of the art, the proposed FPGA prototype achieves a higher processing rate per hardware resource for most configurations. Jérémy Nadal, Amer Baghdadi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Memristor Overwrite Logic (MOL) for Energy-Efficient In-Memory DNNabstractIn-memory computing is a promising solution to address the memory wall challenges in future processing systems. Substantial improvement in performance and energy efficiency is expected, in particular for data intensive applications. A typical use case is neural network applications, where large amount of data should be processed and moved between memory and processing cores. Although several recent works tried to accelerate processing through dedicated parallel hardware designs, data movement cost is still a critical technical challenge. In this context, we propose a novel programmable architecture design for in-memory deep neural networks (DNN) computation. Based on a new logic design style, namely Memristor Overwrite Logic (MOL), specialized computational memory is designed. The original architecture of the proposed computational memory allows to execute multiply-accumulate operations between stored words using MOL. Outstanding features are demonstrated with respect to other recent logic design styles based on emerging non-volatile memory technologies. Khaled Alhaj Ali, Mostafa Rizk, Amer Baghdadi, Jean-Philippe Diguet, Jalal Jomaah |
ISCAS | 3 |
| 2020 | FPGA based design and prototyping of efficient 5G QC-LDPC channel decodingabstractThe Quasi-Cyclic (QC) Low-Density ParityCode (LDPC) is the key error correction code for the 5thGeneration (5G) of cellular network technology. Designed to support several frame sizes and code rates, the 5G LDPC code structure allows high parallelism to deliver the high demanding data rate of 10 Gb/s. This impressive performance introduces challenging constraints on the hardware design. Particularly, allowing such high flexibility can introduce processing rate penalties on some configurations. In this context, a novel efficient and flexible hardware architecture for the 5G LDPC decoder is proposed, targeting Field Programmable Gate Array (FPGA) devices and supporting all 5G configurations. The architecture supports frame parallelism to maximize the utilization of the processing units, significantly improving the processing rate. Compared to a recent commercial 5G LDPC decoder, the proposed FPGA prototype achieves a higher processing rate for most configurations while having similar complexity. Jérémy Nadal, Amer Baghdadi |
RSP | 2 |
| 2020 | Memristive Computational Memory Using Memristor Overwrite Logic (MOL)abstractIn this article, we present a novel logic design style, namely, memristor overwrite logic (MOL), associated with an original MOL-based computational memory. MOL relies on a fully digital representation of memristor and can operate with different memristive device technologies. Its integration in memristive crossbar arrays and computational memories allows the execution of bit and vector-level primitive logic operations in two computational steps at most. Promising features and performances are demonstrated through the implementation of N -bit full addition using the proposed MOL-based computational memory. Khaled Alhaj Ali, Mostafa Rizk, Amer Baghdadi, Jean-Philippe Diguet, Jalal Jomaah, Naoya Onizawa, Takahiro Hanyu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Overlap-Save FBMC ReceiversabstractFuture communication systems are foreseen to support several services with different requirements. Waveform designs based on filter-bank multi-carrier with offset quadrature amplitude modulation (FBMC/OQAM) can offer interesting advantages in this context, such as low out-of-band power leakage and high spectral efficiency due to the lack of guard intervals. However, downsides of FBMC/OQAM with respect to a typical orthogonal frequency-division multiplexing (OFDM) solution include higher latency, higher complexity and difficulties in adapting some existing OFDM techniques such as MIMO Alamouti. To address these issues, novel FBMC receivers suitable for short prototype filters are proposed. Based on the Overlap-Save algorithm, the proposed receivers improve error-rate performance on multipath channels and support asynchronous communication. We show that complexity can be further reduced by efficiently processing blocks of FBMC symbols jointly, and that user mobility support can be traded off for additional complexity reductions in a flexible way through polynomial decomposition of the equalizer stage. Finally, we show that a block-Alamouti scheme can be applied, and we propose a MIMO equalizer with improved error-rate performance on time-varying channels, compared to the typical FBMC block-Alamouti equalizer. Jérémy Nadal, François Leduc-Primeau, Charbel Abdel Nour, Amer Baghdadi |
IEEE Trans. Wirel. Commun. | 4 |
| 2018 | A Block FBMC Receiver Designed for Short FiltersabstractIn this paper, a new filter-bank multi-carrier (FBMC) receiver targeted at short prototype filters (PFs) is presented. In addition to the typical advantages of FBMC modulation such as improved frequency containment and support of relaxed synchronization, the proposed receiver enables accurate one-tap equalization of the signal despite the use of a short PF, while significantly reducing the complexity by merging the equalization and the standard FBMC receiver. An adapted frame structure is proposed that occupies the same radio resources as a 4G/LTE orthogonal frequency-division multiplexing (OFDM) frame while providing a similar data rate. Moreover, this frame structure is shown to readily support block-based Alamouticoded multiple-input multiple-output transmissions. Simulation results show that the proposed FBMC receiver can outperform an OFDM system on 4G/LTE channel models in both single antenna and Alamouti configurations, while having a better robustness to synchronization errors than a frequency-spread FBMC receiver. Jérémy Nadal, François Leduc-Primeau, Charbel Abdel Nour, Amer Baghdadi |
ICC | 4 |
| 2018 | Rapid Prototyping of Parameterized Rotated and Cyclic Q Delayed Constellations DemapperabstractRotated and Cyclic Q Delayed (RCQD) modulation is one of the signal processing schemes which, once used on the transmitter side, provides performance improvement on receiver side in case of fading channel conditions with erasures. However, an efficient hardware solution in terms of low area and high throughput is mandatory to get benefit from this modulation technique. In this paper we have implemented a parameterized hardware architecture of most recent RCQD demapping technique supporting different constellations. In addition to the description of the proposed hardware architecture, rapid prototyping flow based on LabVIEW FPGA is described which is useful to explore different design schemes in short period of time. The implementation results are achieved by FPGA prototyping while targeting NI-USRP RIO hardware equipment. By comparing our rapidly prototyped parameterized RCQD demapper with other state-of-the-art implementation schemes, it is shown that our prototyped demapper consumed lesser hardware resources and achieved higher throughput with very short design time. Muhammad Waqas 0003, Atif Raza Jafri, Amer Baghdadi, Muhammad Najam-ul-Islam |
RSP | 3 |
| 2018 | Networked Power-Gated MRAMs for Memory-Based ComputingabstractEmerging nonvolatile memory technologies open new perspectives for original computing architectures. In this paper, we propose a new type of flexible and energy-efficient architecture that relies on power-gated distributed magnetoresistive random access memory (MRAM). The proposed architecture uses a network-on-chip (NoC) to interconnect MRAM-based clusters, processing elements, and managers. The NoC distributes application-specific commands to MRAM devices by means of packets. Configurable network interfaces allow to transform MRAM devices into smart units able to respond to incoming commands. In this context, three types of MRAM designs are proposed with different power-gating policies and granularities. A relevant database search engine case study is considered to illustrate the benefits of this proposed architecture. It is implemented with a sparse-neural-network approach and simulated in SystemC with different scenarios including hundreds of database queries. Hardware designs and accurate power estimations have been conducted. The obtained results demonstrate important power reduction with database hit rates of about 94%. Targeting 65-nm technology, energy savings reach 87% when compared with an static random access memory-based implementation. Moreover, a new asymmetric read/write MRAM type provides from 39% to 50% energy reduction with respect to the other fixed-granularity models. This results in a low-power, highly scalable, and configurable implementation of memory-based computing. Jean-Philippe Diguet, Naoya Onizawa, Mostafa Rizk, Martha Johanna Sepúlveda, Amer Baghdadi, Takahiro Hanyu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | A Dynamically Reconfigurable Multi-ASIP Architecture for Multistandard and Multimode Turbo DecodingabstractThe multiplication of wireless communication standards is introducing the need of flexible and reconfigurable multistandard baseband receivers. In this context, multiprocessor turbo decoders have been recently developed in order to support the increasing flexibility and throughput requirements of emerging applications. However, these solutions do not sufficiently address reconfiguration performance issues, which can be a limiting factor in the future. This brief presents the design of a reconfigurable multiprocessor architecture for turbo decoding achieving very fast reconfiguration without compromising the decoding performances. Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Michael Hübner 0001, Jean-Philippe Diguet |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Caasper: providing accessible FPGA-acceleration over the networkabstractFPGA acceleration is a commonly used technology for high-performance scientific computing. It offers massive parallelism with low power requirements. One of the main issue with such an approach is interfacing accelerators implemented on FPGA fabric with a host application. This requires physical FPGA access and low-level communication interfaces. In this paper, we present Caasper, a scalable and flexible communication framework designed to provide shared access to FPGA resources over a TCP network, through high-level communication routines. Prototype implementation of the framework demonstrates its efficiency and its usability in high-throughput applications. Valentin Mena Morales, Yahia Brakni, Pierre-Henri Horrein, Amer Baghdadi |
RSP | 4 |
| 2014 | Energy-efficient FPGA implementation for binomial option pricing using OpenCLabstractEnergy efficiency of financial computations is a performance criterion that can no longer be dismissed, and is as crucial as raw acceleration and accuracy of the solution. In order to reduce the energy consumption of financial accelerators, FPGAs offer a good compromise with low power consumption and high parallelism. However, designing and prototyping an application on an FPGA-based platform are typically very time-consuming and requires significant skills in hardware design. This issue constitutes a major drawback with respect to software-centric acceleration platforms and approaches. A high-level approach has been chosen, using Altera's implementation of the OpenCL standard, to answer this issue. We present two FPGA implementations of the binomial option pricing model on American options. The results obtained on a Terasic DE4 — Stratix IV board form a solid basis to hold all the constraints necessary for a real world application. The best implementation can evaluate more than 2000 options/s with an average power of less than 20W. Valentin Mena Morales, Pierre-Henri Horrein, Amer Baghdadi, Erik Hochapfel, Sandrine Vaton |
DATE | 3 |
| 2014 | Design and prototyping flow of NISC-based flexible MIMO turbo-equalizerabstractFlexible design implementations are increasingly explored in digital communication applications to cope with diverse configurations imposed by the emerging communication standards. On the other hand, rapid hardware prototyping is a crucial requirement in system validation and performance evaluation under various use case scenarios. Adding flexibility, and hence increasing system complexity on one hand, and shrinking design time to meet with market pressure on the other hand, require a productive design approach ensuring final design quality. By eliminating the instruction set overhead, No- Instruction-Set-Computer (NISC) approach fulfills these design requirements offering static scheduling of datapath, automated RTL synthesis and allowing designer to have direct control of hardware resources. This paper presents a case study of an NISC-based implementation of a flexible low-complexity MIMO turboequalizer. The complete design and prototype flow, from architecture specification till FPGA implementation, is described in details. Using VC707 evaluation board integrating Xilinx Virtex-7 FPGA, the prototype of 2×2/4×4 spatially multiplexed MIMO system achieves a throughput of 115.8/62.4 Mega symbols per second at a clock cycle frequency of 202.67 MHz. Furthermore, the flexibility of the demonstrated prototype allows to support all communication modes defined in LTE, WiFi, WiMAX, and DVB-RCS wireless communication standards. Mostafa Rizk, Amer Baghdadi, Michel Jézéquel, Yasser Mohanna, Youssef Atat |
RSP | 2 |
| 2014 | Energy-efficient multi-standard early stopping criterion for low-density-parity-check iterative decodingabstractLow‐density‐parity‐check codes decoding relies on powerful iterative algorithms, whose implementation is often expensive in terms of complexity and power consumption. Several early stopping criteria (ESCs) have been proposed to reduce the number of iterations performed by a decoder with no (or limited) degradation of error correction performance. However, most of the existing ESCs have considered a reduced set of system parameters for validation and often have ignored the impacts related to a real hardware implementation. This study proposes a novel multi‐standard early stopping criterion (MSESC) able to adapt dynamically to changes of code parameters, quantisation and channel conditions. A dedicated hardware architecture is devised and integrated in a multi‐standard decoder, and compared with existing techniques. Post‐layout results of the proposed MSESC show a small area increment (+1.3%) and a large decrement of the average energy consumption (up to 87.2%) with respect to the same decoder implemented with no ESC. Moreover, it is shown that MSESC offers an energy consumption reduction with respect to the state‐of‐the‐art ESCs ranging from 4% [at high signal‐to‐noise ratio (SNR)] to 16% (at low SNR). Carlo Condo, Amer Baghdadi, Guido Masera |
IET Commun. | 2 |
| 2013 | Parameterized area-efficient multi-standard turbo decoderabstractEmerging wireless digital communication standards specify a large variety of channel coding options, each suitable for specific application needs. In this context, several recent efforts are being conducted to propose flexible channel decoder implementations. However, the need of optimal solutions in terms of performance, area, and power consumption is increasing and cannot be neglected against flexibility. In this paper we present a novel parameterized architecture for multi-standard Turbo decoding which illustrates how flexibility, architecture efficiency, and rapid design time can be combined. The proposed architecture supports both single-binary Turbo codes (SBTC) of 3GPP-LTE and double-binary Turbo codes (DBTC) of WiMAX and DVB-RCS standards. It achieves, in both modes, a high architecture efficiency of 4.37 bits/cycle/iteration/mm2. A major contribution of this work concerns the rapid design time allowed by the well established design concept and tools of application-specific instruction-set processors (ASIPs). Using such a tool, the paper illustrates the possibility to design application-specific parameterized cores, removing the need of the program memory and the related instruction decoder. Purushotham Murugappa, Amer Baghdadi, Michel Jézéquel |
DATE | 2 |
| 2013 | Statically-scheduled application-specific processor design: a case-study on MMSE MIMO equalizationabstractMany application-specific processor design approaches are being proposed and investigated nowadays. All of them aim to cope with the emerging flexibility requirement combined with the best performance efficiency. Application Specific Instruction-set Processor (ASIP) design approach is among the most explored, and thus in many application domains. However, this concept implies a dynamic scheduling of a set of instructions which generally lead to an overhead related to instruction decoding. To reduce this overhead, other approaches were proposed using static scheduling of datapath control signals. In this paper, we explore this last approach and illustrate its benefits through a design case-study on MMSE MIMO equalization. The proposed design has common main architectural choices as a state-of-the-art ASIP for comparison purpose. The obtained results illustrate a significant improvement in execution time while using identical computational resources and supporting same flexibility parameters. Mostafa Rizk, Amer Baghdadi, Michel Jézéquel, Yasser Mohanna, Youssef Atat |
DATE | 2 |
| 2013 | A Joint Communication and Application Simulator for NoC-Based Custom SoCs: LDPC and Turbo Codes Parallel Decoding Case StudyabstractNoCs have become a widespread paradigm in the system-on-chip design world, not only for multi-purpose SoCs, but also for application-specific ICs. The common approach in the NoC design world is to separate the design of the interconnection from the design of the processing elements: this is well suited for a large number of developments, but the need for joint application and NoC design is not uncommon, especially in the application-specific case. The correlation between processing and communication tasks can be strong, and separate or trace-based simulations fall often short of the desired precision. In this work, the OMNET++ based JANoCS simulator is presented: concurrent simulation of processing and communication allow cycle-accurate evaluation of the system. The potential of the proposed approach is illustrated through a simple application example. Furthermore, a detailed case study on LDPC and turbo codes parallel decoding is presented. Results analysis illustrates the need for joint simulations and demonstrates the effectiveness of the proposed JANoCS. Carlo Condo, Amer Baghdadi, Guido Masera |
DSD | 2 |
| 2013 | Stopping-Free Dynamic Configuration of a Multi-ASIP Turbo DecoderabstractThe multiplication of wireless standards is introducing the need of flexible and reconfigurable multistandard base band receivers. At the physical layer, multiprocessor turbo decoders have been recently developed in order to provide an answer to the increasing throughput requirement of emerging standards. However these solutions do not sufficiently address reconfiguration performance issues which can be a limiting factor in the future. This work focuses on the design of a reconfigurable multiprocessor architecture for turbo decoding achieving very fast reconfiguration without compromising decoding performances. Dynamic reconfiguration can be performed within a single frame decoding duration opening new perspective for reconfigurable multistandard base band receivers. For that purpose, optimizations at the processing element level and a novel bus-based configuration infrastructure are proposed. Results show that up to 64 processings elements can be dynamically configured in 5.352 μs. This low configuration latency corresponds to a single frame decoding duration when performing 6 decoding iterations for a throughput up to 666 Mbps. Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Michael Hübner 0001, Jean-Philippe Diguet |
DSD | 4 |
| 2013 | Optimizations for an efficient reconfiguration of an ASIP-based turbo decoderabstractThe multiplication of wireless standards is introducing the need of flexible multi-standard baseband receivers. A multi-ASIP approach for turbo decoding is an answer to reach high throughput and high flexibility. The increasing demand of throughput for new greedy application on mobile devices and the reduction of latency between two frames create the need of an efficient reconfiguration management of such multi-ASIP platforms. In this paper, we propose to tackle reconfiguration optimization of a multi-standard ASIP for turbo decoding developed during previous work. Results show that for an area overhead of 0.012 mm2in 65 nm CMOS technology, a significant reconfiguration time optimization is achieved thanks to a reduction of the ASIP configuration load of 70%. Moreover, in a multi-ASIP context in which 8 ASIPs are implemented the configuration load is divided by ten thanks to the possibility to use a multicast mechanism for ASIP configuration loading. Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Jean-Philippe Diguet, Jean-Noel Bazin, Michael Hübner 0001 |
ISCAS | 4 |
| 2013 | Rapid design and prototyping of a reconfigurable decoder architecture for QC-LDPC codesabstractMany modern and emerging designs require having efficient dynamically reconfigurable and reprogrammable processors. However, when the implemented design needs an upgrade, newly added features have to be quickly supported and validated. This is clearly noticed in modern receivers of recent wireless communication standards that feature continuously different frame lengths and code rates for the channel decoder. This paper explores with an example the possibility of realizing a flexible channel decoder to implement and validate new/incremental algorithm changes with fast turnaround time in design. An application specific instruction-set processor (ASIP) is proposed as flexible core that can decode low-density parity-check (LDPC) codes with the various block sizes and code rates as specified in WiFi and WiMAX standards. Furthermore, the proposed architecture enables quick support of other Quasi-Cyclic LDPC (QC-LDPC) codes, e.g. DVB-S2, with simple incremental hardware changes at design time. Purushotham Murugappa, Vianney Lapotre, Amer Baghdadi, Michel Jézéquel |
RSP | 3 |
| 2012 | FPGA prototyping and performance evaluation of multi-standard Turbo/LDPC Encoding and DecodingabstractHardware prototyping has been the key to system validation, once the hardware simulation matches the software model results and before the final silicon tape-out. On the other hand, flexible multi-standard implementations are being widely investigated these last years for the challenging channel decoding application. The latest contributions explore ASIP (Application-Specific Instruction-set Processor) concept and target to achieve efficient resource sharing between advanced Turbo and LDPC iterative decoders. In this paper we present an FPGA-based prototype of a multistandard Turbo/LDPC Encoding and Decoding. The functional prototype implements a full communication system including encoder, channel model, ASIP-based decoder and performance counters. All components are flexible and are dynamically configurable through a dedicated GUI (Graphical User Interface). The prototype supports all communication modes defined in LTE, WiFi, WiMAX, and DVB-RCS wireless communication standards. Purushotham Murugappa, Jean-Noel Bazin, Amer Baghdadi, Michel Jézéquel |
RSP | 3 |
| 2011 | A flexible high throughput multi-ASIP architecture for LDPC and turbo decodingabstractIn order to address the large variety of channel coding options specified in existing and future digital communication standards, there is an increasing need for flexible solutions. This paper presents a multi-core architecture which supports convolutional codes, binary/duo-binary turbo codes, and LDPC codes. The proposed architecture is based on Application Specific Instruction-set Processors (ASIP) and avoids the use of dedicated interleave/deinterleave address lookup memories. Each ASIP consists of two datapaths one optimized for turbo and the other for LDPC mode, while efficiently sharing memories and communication resources. The logic synthesis results yields an overall area of 2.6mm2using 90nm technology. Payload throughputs of up to 312Mbps in LDPC mode and of 173Mbps in Turbo mode are possible at 520MHz, fairing better than existing solutions. Purushotham Murugappa, Rachid Al-Khayat, Amer Baghdadi, Michel Jézéquel |
DATE | 3 |
| 2011 | A low complexity stopping criterion for reducing power consumption in turbo decodersabstractTurbo codes are proposed in most of the advanced digital communication standards, such as 3GPP-LTE. However, due to its computational complexity, the turbo decoder is one of the most power hungry blocks in digital baseband. To alleviate this issue, one way is to avoid surplus computing phases thanks to the early termination of the iterative decoding process. The use of stopping criteria is one of the most common algorithm level power reduction methods in literature. These methods always come with some hardware overhead. In this paper, a new trellis based stopping criterion is proposed. The novelty of this approach is the lower hardware overhead thanks to the use of trellis states as key parameter to stop the iterative process. Results are showing the importance of this added hardware in terms of method efficiency. Compared to state-of-the-art Log Likelihood Ratio (LLR) based techniques, proposed Low Complexity Trellis Based (LCTB) is demonstrating 23% less power consumption on average, for comparable performance level in terms of Bit Error Rate (BER) and Frame Error Rate (FER). Pallavi Reddy, Fabien Clermidy, Amer Baghdadi, Michel Jézéquel |
DATE | 3 |
| 2010 | Rapid design and prototyping of universal soft demapperabstractRapid advancements in wireless communication standardization is leading toward the evolution of flexible radio platforms. At the same time, the resulting severe time-to-market constraints make inevitably resorting to new design methodologies to shorten the development cycle. In this paper we are presenting the steps involved in rapid design, validation, and prototyping of the first multi standard ASIP-based universal demapper. The presented ASIP provides flexibility to support any modulation type using up to 8 bits per symbol both in turbo and non-turbo context. The rapid development flow has been described starting from ASIP modeling in LISA ADL till the FPGA implementation. Using a logic emulation board integrating Virtex 5 LX330 FPGA, the prototype achieves a throughput of 102 Mega LLR/sec for Gray mapped 16-QAM constellation at a clock frequency of 156 MHz. The highly reduced size of the ASIP, comprising of 1596 (0.7%) slice registers, 2627 (1.2%) slice LUTs and 6 DSP48Es, enables the user to achieve even higher throughputs by using multi-ASIP architecture. Atif Raza Jafri, Amer Baghdadi, Michel Jézéquel |
ISCAS | 2 |
| 2009 | ASIP-based flexible MMSE-IC Linear Equalizer for MIMO turbo-equalization applicationsabstractA novel 16-bit flexible application-specific instruction-set processor for an MMSE-IC linear equalizer, used in iterative turbo receiver, is presented in this paper. The proposed ASIP has an SIMD architecture with a specialized instruction-set and 7-stage pipeline control. It supports diverse requirements of MIMO-OFDM wireless standards such as use of QPSK, 16-QAM and 64-QAM modulation in 2times2 and 4times4 spatially multiplexed MIMO-OFDM environment. For these various operational modes, analysis of MMSE-IC LE equations and corresponding complex data representations was conducted. Efficient computational and storage resource sharing is proposed through: (1) matrix register banks (MRB) multiplexing, (2) 16-bit complex arithmetic unit (CAU) comprised of 4 combined complex adder/subtractor/multiplier units, 2 real multipliers, 5 complex adders, and 2 complex subtractors, and (3) flexible 32-bit to 16-bit data conversion at multipliers' output. With this architecture, the designed ASIP ensures, along with flexibility, high performance in terms of throughput and area. Logic synthesis results reveal a maximum clock frequency of 546 MHz and a total area of 0.37 mm2using 90 nm technology. For 2times2 spatially multiplexed MIMO system, the proposed ASIP achieves a throughput of 273 Msymbol/sec. Atif Raza Jafri, Daoud Karakolah, Amer Baghdadi, Michel Jézéquel |
DATE | 3 |
| 2009 | Flexible Architectures for LDPC Decoders Based on Network on Chip ParadigmabstractThis paper explores the possibility of building a flexible Low Density Parity Check (LDPC) decoder using a network on chip communication infrastructure. Even if this idea is not completely new, previously published works suffered from an excessive area occupation and their practical impact has been very limited. In the following we analyze two possible NOCs specifically designed for the LDPC case. From synthesis results it can be observed how the proposed networks outperform previous implementations in terms of active area with no significant bandwidth loss. Finally to prove the effectiveness of the proposed approach a complete, partially parallel LDPC decoder design is presented and characterized in terms of throughput and area occupation. Fabrizio Vacca, Guido Masera, Hazem Moussa, Amer Baghdadi, Michel Jézéquel |
DSD | 4 |
| 2009 | From Parallelism Levels to a Multi-ASIP Architecture for Turbo DecodingabstractEmerging digital communication applications and the underlying architectures encounter drastically increasing performance and flexibility requirements. In this paper, we present a novel flexible multiprocessor platform for high throughput turbo decoding. The proposed platform enables exploiting all parallelism levels of turbo decoding applications to fulfill performance requirements. In order to fulfill flexibility requirements, the platform is structured around configurable application-specific instruction-set processors (ASIP) combined with an efficient memory and communication interconnect scheme. The designed ASIP has an single instruction multiple data (SIMD) architecture with a specialized and extensible instruction-set and 6-stages pipeline control. The attached memories and communication interfaces enable its integration in multiprocessor architectures. These multiprocessor architectures benefit from the recent shuffled decoding technique introduced in the turbo-decoding field to achieve higher throughput. The major characteristics of the proposed platform are its flexibility and scalability which make it reusable for all simple and double binary turbo codes of existing and emerging standards. Results obtained for double binary WiMAX turbo codes demonstrate around 250 Mb/s throughput using 16-ASIP multiprocessor architecture. Olivier Muller, Amer Baghdadi, Michel Jézéquel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Binary de Bruijn on-chip network for a flexible multiprocessor LDPC decoderabstractThis paper proposes a novel on-chip interconnection network adapted to a flexible multiprocessor LDPC decoder based on the de Bruijn network. The main characteristics of this network -- including its logarithmic diameter, scalable aggregate bandwidth, and optimized routing technique- allow it to efficiently support the communication intensive nature of the application. We present a detailed hardware implementation of the routers and the network interfaces as well as the packet format and the routing algorithm. The latter is a parallelized version of the shortest path with deflection routing algorithm. In order to evaluate the performance of the proposed network, a generic RTL VHDL description has been developed and synthesized with CMOS STMicroelectronics 0.18 μm technology. The flexibility and the scalability of this onchip communication network enable it to be used for any kind of LDPC code. Hazem Moussa, Amer Baghdadi, Michel Jézéquel |
DAC | 2 |
| 2008 | Binary de Bruijn interconnection network for a flexible LDPC/turbo decoderabstractThis paper proposes a novel on-chip interconnection network adapted to a flexible multiprocessor LDPC/turbo decoder and based on the de Bruijn network. The main characteristics of this network -including its logarithmic diameter, scalable aggregate bandwidth, and optimized routing technique- allow it to efficiently support the communication- intensive nature of the two decoding techniques. We present a detailed hardware implementation of the routers and the network interfaces as well as the packet format and the routing algorithm. In order to evaluate the performance of the proposed network, a generic RTL VHDL description has been developed and synthesized with ST CMOS 0.18 mum technology. The flexibility and the scalability of this on-chip communication network enable it to be used in the emerging multi-code applications and standards. In addition, the results obtained for a 16-processor network demonstrate a major aggregate bandwidth of 296 Gbps with a relative small area of 3.56 mm2. Hazem Moussa, Amer Baghdadi, Michel Jézéquel |
ISCAS | 2 |
| 2007 | Butterfly and benes-based on-chip communication networks for multiprocessor turbo decodingabstractSeveral research activities have recently emerged aiming to propose multiprocessor implementations in order to achieve flexible and high throughput parallel iterative decoding. Besides application algorithm optimizations and application-specific instruction-set processor design, the on-chip communication network constitutes a major issue in this application domain. In this paper, the authors propose to use multistage interconnection networks as on-chip communication networks for parallel turbo decoding. Adapted benes and butterfly networks are proposed with detailed hardware implementation of network interfaces, routers, and topologies. In addition, appropriate packet format and routing for interleaved/deinterleaved extrinsic information exchanges are proposed. The flexibility of these on-chip communication networks enables their use for all turbo code standards and constitutes a promising feature for their reuse for any similar interleaved/deinterleaved iterative communication profile Hazem Moussa, Olivier Muller, Amer Baghdadi, Michel Jézéquel |
DATE | 3 |
| 2006 | ASIP-based multiprocessor SoC design for simple and double binary turbo decodingabstractThis paper presents a new multiprocessor platform for high throughput turbo decoding. The proposed platform is based on a new configurable ASIP combined with an efficient memory and communication interconnect scheme. This application-specific instruction-set processor has an SIMD architecture with a specialized and extensible instruction-set and 5-stages pipeline control. The attached memories and communication interfaces enable the design of efficient multiprocessor architectures. These multiprocessor architectures benefit from the recent shuffling technique introduced in the turbo-decoding field to reduce communication latency. The major characteristics of the proposed platform are its flexibility and scalability which make it reusable for various standards and operating modes. Results obtained for double binary DVB-RCS turbo codes demonstrate a 100 Mbit/s throughput using 16-ASIP multiprocessor architecture Olivier Muller, Amer Baghdadi, Michel Jézéquel |
DATE | 2 |
| 2006 | On the Parallelism of Convolutional Turbo Decoding and Interleaving InterferenceabstractIn forward error correction, convolutional turbo codes were introduced to increase error correction capability approaching the Shannon bound. Decoding of these codes, however, is an iterative process requiring high computation rate and latency. Thus, in order to achieve high throughput and to reduce latency, crucial in emerging digital communication applications, parallel implementations become mandatory. This paper explores and analyses existing parallelism techniques in convolutional turbo decoding with the BCJR algorithm. For component-decoder parallelism, we illustrate the influence of interleaving scheme and we propose new interleaving rules allowing to maximize parallelism efficiency. Olivier Muller, Amer Baghdadi, Michel Jézéquel |
GLOBECOM | 2 |
| 2004 | An efficient scalable and flexible data transfer architecture for multiprocessor SoC with massive distributed memoryabstractMassive data transfer encountered in emerging multimedia embedded applications requires architecture allowing both highly distributed memory structure and multiprocessor computation to be handled. The key issue that needs to be solved is then how to manage data transfers between large numbers of distributed memories. To overcome this issue, our paper proposes a scalable Distributed Memory Server (DMS) for multiprocessor SoC (MPSoC). The proposed DMS is composed of: (1) high-performance and flexible memory service access points (MSAPs), which execute data transfers without intervention of the processing elements, (2) data network, and (3) control network. It can handle direct massive data transfer between the distributed memories of an MPSoC. The scalability and flexibility of the proposed DMS are illustrated through the implementation of an MPEG4 video encoder for QCIF and CIF formats. The experiments show clearly how DMS can be adapted to accommodate different SoC configurations requiring various data transfer bandwidths. Synthesis results show that bandwidth can scale up to 28.8 GB/sec. Sangil Han, Amer Baghdadi, Marius Bonaciu, Soo-Ik Chae, Ahmed Amine Jerraya |
DAC | 2 |
| 2002 | Component-based design approach for multicore SoCsabstractThis paper presents a high-level component-based methodology and design environment for application-specific multicore SoC architectures. Component-based design provides primitives to build complex architectures from basic components. This bottom-up approach allows design-architects to explore efficient custom solutions with best performances. This paper presents a high-level component-based methodology and design environment for application-specific multicore SoC architectures. The system specifications are represented as a virtual architecture described in a SystemC-like model and annotated with a set of configuration parameters. Our component-based design environment provides automatic wrapper-generation tools able to synthesize hardware interfaces, device drivers, and operating systems that implement a high-level interconnect API. This approach, experimented over a VDSL system, shows a drastic design time reduction without any significant efficiency loss in the final circuit. Wander O. Cesário, Amer Baghdadi, Lovic Gauthier, Damien Lyonnard, Gabriela Nicolescu, Yanick Paviot, Sungjoo Yoo, Ahmed Amine Jerraya, Mario Diaz-Nava |
DAC | 2 |
| 2002 | Combining a Performance Estimation Methodology with a Hardware/Software Codesign Flow Supporting Multiprocessor SystemsabstractThis paper addresses performance estimation and architecture exploration issues within the context of hardware/software codesign. We introduce a new methodology to rapidly explore the large design space encountered in hardware/software systems. The proposed methodology is based on a fast and accurate estimation approach. This estimation approach takes advantage of both system and RT levels of abstraction, and combines both static and dynamic analysis techniques, in order to obtain the best trade-off between speed and accuracy. It has been implemented as an extension to a hardware/software codesign flow to enable the exploration of a large number of multiprocessor architecture solutions from the very start of the design process. The effectiveness of the proposed methodology is illustrated by a significant application example. Experimental results indicate strong advantages of the proposed methodology. Amer Baghdadi, Nacer-Eddine Zergainoh, Wander O. Cesário, Ahmed Amine Jerraya |
IEEE Trans. Software Eng. | 1 |
| 2001 | Automatic Generation of Application-Specific Architectures for Heterogeneous Multiprocessor System-on-ChipabstractWe present a design flow for the generation of application-specific multiprocessor architectures. In the flow, architectural parameters are first extracted from a high-level system specification. Parameters are used to instantiate architectural components, such as processors, memory modules and communication networks. The flow includes the automatic generation of communication coprocessor that adapts the processor to the communication network in an application-specific way. Experiments with two system examples show the effectiveness of the presented design flow. Damien Lyonnard, Sungjoo Yoo, Amer Baghdadi, Ahmed Amine Jerraya |
DAC | 3 |
| 2001 | An efficient architecture model for systematic design of application-specific multiprocessor SoCabstractIn this paper, we present a novel approach for the design of application specific multiprocessor systems-on chip. Our approach is based on a generic architecture model which is used as a template throughout the design process. The key characteristics of this model are its great modularity, flexibility and scalability which make it reusable for a large class of applications. In addition, it allows one accelerate the design cycle. This paper focuses on the definition of the architecture model and the systematic design flow that can be automated. The feasibility and effectiveness of this approach are illustrated by two significant demonstration examples. Amer Baghdadi, Damien Lyonnard, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya |
DATE | 1 |