EDBT 2026 Demo / reviewers in the wild / expert
Tarek Ould Bachir
dblp:59/9612 · also Tarek Ould-Bachir
· DBLP profile ↗
14ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-9000-5467ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synchronized CPU-FPGA Tracing for Heterogeneous PlatformsabstractInternational audience Nicolas Deloumeau, Tarek Ould Bachir, Andrew Handke, Jamie Sanderson, François Tetreault |
FPGA | 2 |
| 2025 | Accurate Real-Time Simulation of CLLLC Converters on FPGA: A Study of Sampling Resolution and Beating EffectabstractThe CLLLC resonant converter, used for high-efficiency applications such as battery chargers, presents major challenges for real-time simulation, notably by its use at high frequencies and by the involvement of natural switching. Furthermore, the resonant nature and high-frequency AC transformer waveforms characterizing such converters make the simulation highly sensitive to sampling errors introduced by the discrete nature of the simulation. This work proposes an approach based on a reconfigurable switched matrix solver associated with explicit switch handling and implemented on a low-cost FPGA target to meet these requirements. The hardware architecture achieves a computational step size as low as 25 ns, even on an affordable FPGA platform. The algorithms, validated by comparison with SIMBA software, guarantee accuracy with relative errors of less than 2%. These results demonstrate the feasibility of using low-cost FPGAs for demanding applications, offering an effective solution for real-time simulation and control of high-frequency resonant converters. The model’s limitations are also explored, and avenues for improvement are proposed. Téo Robert, Tarek Ould Bachir, Valentin Combet, Mohammed Kenzi, Romain Monthéard |
IECON | 2 |
| 2025 | OMS-CNN: Optimized Multi-Scale CNN for Lung Nodule Detection Based on Faster R-CNNabstractThe global increase in lung cancer cases, often marked by pulmonary nodules, underscores the critical importance of timely detection to mitigate cancer progression and reduce morbidity and mortality. The Faster R-CNN approach is a two-stage, high-precision nodule detection method designed for detecting small nodules, particularly in computed tomography (CT) images. This paper presents an improved Faster R-CNN by introducing an optimized multi-scale convolutional neural network (OMS-CNN) technique for feature map generation. This approach aims to achieve an optimal feature map through metaheuristic optimization by combining the last three layers of the VGG16 architecture. The advanced parameter-setting-free harmony search (PSF-HS) algorithm is utilized to implement this method, automatically adjusting the number of channels in the composite layers as a hyperparameter. The beetle antenna search (BAS) optimization algorithm is utilized to effectively initialize the kernel filter weights and biases in the composite layers, thereby enhancing training speed and detection accuracy. In the false-positive reduction stage, a combination of multiple 3D deep convolutional neural networks (3D DCNN) is designed to reduce false-positive nodules. The proposed model was evaluated using the LUNA16 and PN9 datasets. The results demonstrate that the OMS-CNN technique effectively extracted representative features of nodules at various sizes, achieving a sensitivity of 94.89% and a CPM score of 0.892. The comprehensive experiments illustrate that the proposed method can enhance detection sensitivity and manage the number of false positive nodules, thereby offering clinical utility and serving as a valuable point of reference. Yadollah Zamanidoost, Tarek Ould Bachir, Sylvain Martel |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | DMM-ESH Technique for Real-time Simulation of Three-level NPC Dual Active Bridge ConverterabstractThis paper proposes a novel approach for real-time simulation of high switching frequency power converters. Our method innovatively combines the direct mapped method (DMM) with the conventional explicit status handling (ESH) technique to accurately and efficiently model power electronic converters. Applied to the three-level neutral point clamped (NPC) dual active bridge (DAB) converter, our approach achieves a significantly reduced time-step when implemented on field-programmable gate arrays (FPGAs) while accounting for the natural commutation of diodes. Furthermore, this research demonstrates the feasibility of implementing the proposed method on low-cost FPGA systems, achieving time-steps ranging from 32 ns to 50 ns for the three-level NPC DAB test case. Tarek Ould Bachir, Karim Meddah, Téo Robert, Romain Monthéard, Emmanuel Rutovic |
IECON | 2 |
| 2024 | Enhancing P4 Syntax to Support Extended Finite State Machines as Native Stateful ObjectsabstractThe P4 language has proven to be a powerful tool for programming packet processing, but its original design did not intend to handle stateful processing effectively. This shortcoming stems from the fact that the network switches for which the language was designed have restricted memory capacities, which makes it challenging to manage complex stateful objects. As a result, P4’s syntax was not optimized for handling such objects. With contemporary networks increasingly relying on stateful processing and abstractions like Extended Finite State Machines (EFSMs), we propose extending P4’s syntax through an EFSM construct. This work aims to grant developers the ability to create streamlined and productive P4 programs that can effortlessly deal with stateful objects. This improvement holds great promise for expanding P4’s functionality and refining it to support stateful processing. Florent Allard, Tarek Ould Bachir, Yvon Savaria |
NetSoft | 2 |
| 2023 | Efficient Region Proposal Extraction of Small Lung Nodules Using Enhanced VGG16 Network ModelabstractThe efficiency of state-of-the-art convolutional networks trained to detect lung cancer nodules depends on their feature extraction model. Various feature extraction models have been proposed based on convolutional networks, such as VGG-Net, or ResNet. It has been demonstrated that such models effectively extract features from objects in an image. However, their efficacy is limited when the objects of interest are very small, such as lung nodules. One of the widely used feature extraction models for detecting small objects is the VGG16 network. The model, which has a small kernel of$\mathbf{3}\times \mathbf{3}$and optimal layers, can extract the features of small objects with reasonable accuracy. In this article, feature maps are created by combining the last three layers of the VGG16 network to extract features of various sizes of nodules. This study utilizes a Region Proposal Network (RPN) to compare the accuracy of the feature map created in the proposed method and the original VGG16. An RPN is a fully-convolutional network that simultaneously predicts object bounds and objectness scores at each position. RPNs are trained end-to-end to generate high-quality region proposals, which Faster R-CNN uses for detection. In this article, we select 300, 1, 000 and 2, 000 regions chosen by the RPN network for each method; then, we calculate the recall for different Intersection over Union (IoU) ratios with ground-truth boxes. The results show that the feature map of the proposed method works more optimally than the feature map of different layers of VGG16 for extracting various sizes of nodules. Also, by reducing the number of selected region proposals, the recall of the proposed method has fewer changes than other methods. Yadollah Zamanidoost, Nada Alami-Chentoufi, Tarek Ould Bachir, Sylvain Martel |
CBMS | 3 |
| 2023 | An Area-efficient Memory-based Architecture for P4-programmable Streaming Parsers in FPGAsabstractMoving toward software-defined networking and function virtualization, flexibility and reconfigurability of the network have become more and more critical. Packet parsing, the first processing stage of programmable switches, requires high performance and reconfigurability to allow implementing low-latency and highly flexible data networks. This paper proposes an overlay architecture for an FPGA-based P4-programmable streaming packet parser. The purpose of this architecture is to allow supporting different functionality with a fixed hardware design by changing a program stored in an embedded memory. This program is derived from the parser section of a P4 code, describing a parsing graph. This approach eliminates a pipeline of parsing blocks in favor of a single parsing block, thereby reducing the design's complexity. Our architecture offers an 11 Gb/s data rate on a Xilinx Virtex-7 XC7VX690 FPGA, while its implementation requires 312 LUTs and 1135 FFs. Parisa Mashreghi-Moghadam, Tarek Ould Bachir, Yvon Savaria |
ISCAS | 2 |
| 2022 | Real-Time Simulation of a Fast Charger Using a Low-Cost FPGA PlatformabstractAdvances in semiconductor technologies bring to the market power electronic converters (PECs) switched at high frequencies, and make their real-time simulation and testing using a hardware-in-loop configuration difficult. Various approaches have been reported in the literature that demonstrate the feasibility of simulating high switching frequency (HSF) PECs in real-time, but require high-end devices. This paper demonstrates that low-cost FPGA platforms can be leveraged for the accurate real-time simulation of HSF of PEC. The proposed platforms make use of a solver that uses state variables from the previous time-point to determine the status of uncontrolled switches. The performance of the proposed approach is assessed using a battery charger test case. The paper demonstrates that the proposed method is latency- and resource-efficient, and achieves a time-step of 32 ns. Karim Meddah, Hossein Chalangar, Tarek Ould Bachir |
IECON | 3 |
| 2022 | A Templated VHDL Architecture for Terabit/s P4-programmable FPGA-based Packet ParsingabstractThis paper proposes a templated VHDL architecture for P4-programmable packet parsing on FPGAs offering high throughput while occupying a small area footprint. The architecture comprises a multi-stage header parser unit arranged in a pipelined structure. Each header analysis unit is characterized by a set of generic parameters reflecting unique features and relations of supported protocols retrieved from the P4 code that describes each stage along the pipeline. Synthesis results of the packet parser show up to 549 Gb/s throughput on a Xilinx Virtex-7 FPGA and 1 Tb/s on a Xilinx UltraScale+ for a twelve-stage pipeline. Compared with state-of-the-art solutions, our proposed architecture performs at higher throughput with acceptable resource utilization. Parisa Mashreghi-Moghadam, Tarek Ould Bachir, Yvon Savaria |
ISCAS | 2 |
| 2015 | FPGA-based real-time simulation of a PSIM model: An indirect matrix converter case studyabstractIn this paper, an indirect matrix converter (IMC) which makes directly ac-ac power conversion, is modeled and simulated in real-time to demonstrate the capability of the new link between Opal-RT's eHS (electric Hardware Solver) and the CAD tool PSIM. An automatic methodology for the real-time simulation of power converters from PSIM circuit designs to FPGA is presented and discussed. A time-step of 250 ns is achieved on an FPGA computing engine developing up to 25.6 GFLOPS of processing power. Real-time simulation results are compared against PSIM as an offline validation tool, showing close match between the FPGA-based simulation and the reference. Tarek Ould Bachir, Asma Merdassi, Sébastien Cense, Handy Fortin-Blanchette, Jean Bélanger |
IECON | 1 |
| 2013 | Self-Alignment Schemes for the Implementation of Addition-Related Floating-Point OperatorsabstractAdvances in semiconductor technology brings to the market incredibly dense devices, capable of handling tens to hundreds floating-point operators on a single chip; so do the latest field programmable gate arrays (FPGAs). In order to alleviate the complexity of resorting to these devices in computationally intensive applications, this article proposes hardware schemes for the realization of addition-related floating-point operators based on the self-alignment technique (SAT). The article demonstrates that the schemes guarantee an accuracy as if summation was computed accurately in the precision of operator’s internal mantissa, then faithfully rounded to working precision. To achieve such performance, the article adopts the redundant high radix carry-save (HRCS) format for the rapid addition of wide mantissas. Implementation results show that combining the SAT and the HRCS format allows the implementation of complex operators with reduced area and latency, more so when a fused-path approach is adopted. The article also proposes a new hardware operator for performing endomorphic HRCS additions and presents a new technique for speeding up the conversion from the redundant HRCS to a conventional binary format. Tarek Ould Bachir, Jean-Pierre David |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2012 | General-purpose reconfigurable low-latency electric circuit and motor drive solver on FPGAabstractThis paper discusses the specifications and general structure of an electric circuit and motor drive solver on FPGA. The `Electric Hardware Solver' or eHS presented in this paper have the objective to facilitate the usage of high-fidelity FPGA solutions for Hardware-In-the-Loop simulation by avoiding the difficulties associated with the coding of such devices by simulation engineers. Christian Dufour, Sébastien Cense, Tarek Ould Bachir, Luc-André Grégoire, Jean Bélanger |
IECON | 3 |
| 2010 | Performing Floating-Point Accumulation on a Modern FPGA in Single and Double PrecisionabstractIn this paper, we discuss the feasibility of a floating-point accumulator (FPACC) on modern high-end FPGA devices. We explore different implementation scenarios and propose new FPACC architectures for both single and double precision floating-point addends. The proposed strategies can be easily adapted to the implement a multiply-accumulator (FPMAC), with one or two rounding stages, in both single and double precision as well. All the aforementioned designs are characterized by high operating frequencies (ranging from 130 to 300 MHz) and moderate occupation area (from 300 to 800 slices) when implemented on the VC5VSX50T FPGA, an entry level Virtex 5 from Xilinx. Tarek Ould Bachir, Jean-Pierre David |
FCCM | 1 |
| 2009 | FPGA-driven pseudorandom number generators aimed at accelerating Monte Carlo methodsabstractHardware acceleration in High Performance Computing (HPC) context is of growing interest, particularly in the field of Monte Carlo methods where the resort to Field Programmable Gate Array (FPGA) technology has been proven as an effective media, capable of enhancing by several orders the speed execution of stochastic processes. The spread-use of reconfigurable hardware for stochastic simulation gathered a significant effort towards effective implementations of hardware pseudorandom numbers generators (PRNGs) - these generators needed to exhibit a statistically proven random behaviour and to be charactarized by a very long period. In this paper we present the state of the art of hardware pseudorandom number generation in the context of Monte Carlo acceleration. We highlight the emerging trends over the most recent publications and suggest some insights on the forthcoming works. Furthermore, we provide a complete hardware description of a new gaussian variate generator (GVG) and an exponential variate generator (EVG) based on a decision-tree technique of ours, herein presented as well. The prototypes implemented on a Xilinx Virtex II Pro XC2VP100 FPGA occupy from 150 to 417 slices and reach 280 MHz, while exhibiting good statistical behaviours with high p-values on the x2test and offering a unitary Knuth ratio. Tarek Ould Bachir, Jean-Jules Brault |
AICCSA | 1 |