VLDB 2026 Research / reviewers in the wild / expert
M. Imtiaz Rashid
dblp:263/1093 · also Md. Imtiaz Rashid
· DBLP profile ↗
9ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0003-2144-181XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 9 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Making Legacy Hardware Robust against Side Channel Attacks via High-Level SynthesisabstractThis work introduces a complete flow to make legacy, side-channel attack (SCA) unaware, hardware given as an Register Transfer Level (RTL) description (Verilog) secure through an RTL to C compiler that generates optimized C code for High-Level Synthesis (HLS). This compiler analyzes the legacy RTL description against SCA and generates C code that can then be in turn re-synthesized into new RTL code that is security-aware. Experimental results show that our proposed flow is able to make security unaware RTL code secure introducing minimal overheads. M. Imtiaz Rashid, Benjamin Carrión Schäfer |
ASP-DAC | 1 |
| 2025 | Robust and Efficient RTL to C Compiler Optimized for High-Level SynthesisabstractDesigning hardware at the register transfer level (RTL) using low-level hardware description languages (HDLs) like Verilog or VHDL gives designers large degrees of controllability to create hardware architectures that will meet the given cost, power budget and performance requirements. The main problem with this approach is that the manually optimized architecture is fixed, which implies that future redesigns to, e.g., target other hardware platforms like field-programmable gate arrays (FPGAs) or newer technologies nodes might require the redesign and reverification of the RTL description. This is error prone and time consuming. To address this, in this work, we propose an RTL to C compiler that generates C code optimized for high-level synthesis (HLS) such that the redesign and reoptimization of new hardware design can be automated. HLS has multiple significant advantages over traditional RT-level design flows like being able to design and verify the behavioral description once and then retarget it for different hardware platforms and constraints by simply using a new technology library and synthesis constraints. Moreover, HLS allows to generate multiple functional equivalent design variants with unique tradeoffs like area, performance, and power from the same behavioral description by setting synthesis options in the form or pragmas (comments) to mainly control how to synthesize arrays (RAM or registers) and loops (unroll, partially unroll, no unroll, or pipeline). In order to leverage these advantages, in this work, we introduce an RTL to C compiler framework that we call MIRROR: maximizing the reusability of RTL through RTL to C Compiler and its new improved version MIRROR++ that is able to compile back to C different types of RTL descriptions, including pipelined circuits, finite state machines, and circuits that share functional units generating arrays and loops so that these can in turn be resynthesized (HLS) with different synthesis directives. Experimental results show the effectiveness and robustness of our approach. M. Imtiaz Rashid, Benjamin Carrión Schäfer |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | MIRROR: MaxImizing the Re-usability of RTL thrOugh RTL to C CompileRabstractThis work presents a RTL to C compiler called MIRROR that maximizes the re-usability of the generated C code for High-Level Synthesis (HLS). The uniqueness of the compiler is that it generates C code by using libraries of pre-characterized RTL micro-structures that are uniquely identifiable through perceptual hashes. This allows to quickly generate C descriptions that include arrays and loops. These are important because HLS tools extensively use synthesis directives in the form of pragmas to control how to synthesize these constructs. E.g., arrays can be synthesized as registers or RAM, and loops fully unrolled, partially unrolled, not unrolled, or pipelined. Setting different pragma combinations lead to designs with unique area vs. performance and power trade-offs. Based on this, the main goal of our compiler is to parse synthesizable RTL descriptions specified in Verilog which have a fixed micro-architecture with specific area, performance and power profile and generate C code for HLS that can then be re-synthesized with different pragma combinations generating a variety of new micro-architectures with different area vs. performance trade-offs. We call this 'maximizing the re-usability of the RTL code because it enables a path to re-target any legacy RTL description to applications with different constraints. In particular we deal with pipelined descriptions in this work due to their uniqueness. Experimental results show that our proposed compiler is very effective, opening the door to automating the re-optimization of legacy hardware designs previously manually optimized using low level Hardware Description Languages (HDLs). We aim at making this compiler framework open source and available to the research community. M. Imtiaz Rashid, Benjamin Carrión Schäfer |
DATE | 1 |
| 2023 | CERTIFY: AutomatiC MEasuRing The QualIty oF High-Level SYnthesisabstractHigh-Level Synthesis (HLS) allows to synthesis un-timed behavioral descriptions into efficient RTL (Verilog or VHDL). Although much progress has been made to improve the quality of HLS it is often reported that the generated RTL code from HLS leads to larger circuits as compared to hand optimized Verilog or VHDL. To measure the gap between hand optimized RTL and automatically generated RTL from HLS, periodic studies are presented were the authors manually optimize designs in RTL, re-write the functionality in C and then compare the quality of the generated RTL code. This is useful, but not very scalable as it is only possible to do this for a small number of designs. Moreover, the result from HLS is highly dependent on the synthesis options used, typically in the case of HLS, these have the form of pragmas (comments) that allow to control how to mainly synthesize arrays (e.g., RAM or registers), loops (e.g., unroll, partially, not unroll) and functions (e.g., inline or not). To address this, in this work we present an RTL to C compiler that generates synthesizable C code for HLS combined with an auto-tuner to automatically find HLS constraints such that the generated RTL code from HLS is as close as possible in terms of area and performance to the original manually optimized RTL code. This allows to directly compare the quality of the generated RTL code by further synthesizing these into equivalent gate netlist. M. Imtiaz Rashid, Amir H. Torabi, Benjamin Carrión Schäfer |
ISCAS | 1 |
| 2023 | Fast and Inexpensive High-Level Synthesis Design Space Exploration: Machine Learning to the RescueabstractHigh-level synthesis (HLS) has multiple significant advantages over traditional RT-level design flows. One in particular that we address in this work is the ability to generate multiple functional equivalent design variants with unique tradeoffs, such as area, performance, and power from the same behavioral description. This is typically done by setting synthesis options in the form or pragmas (comments) to mainly control how to synthesize arrays (RAM or registers), loops (unroll, partially unroll, no unroll or pipeline), and functions (inline or not). Setting different pragma combinations lead to these different design implementations. Out of all the pragma combinations the designer is typically only interested in those that lead to the Pareto-optimal designs (PODs). Fortunately, this search can be automated, but unfortunately, the search space to find these pragma combinations grows supra-linearly with the number of pragma settings. Thus, fast and efficient heuristics are needed. These heuristics generate a new pragma combination and then evaluate their effect by synthesizing (HLS) it. The most time-consuming part of this process is having to execute a full synthesis (HLS) on the behavioral description for every new pragma combination. One obvious way to accelerate the exploration is to parallelize the exploration process using a multithreaded heuristic. The theoretical speedup should match the number of parallel threads. The main problem with this approach is that every HLS invokation requires to check out an HLS tool license. This license is not released until the synthesis process has finished. This implies that the maximum number of parallel threads is restricted by the number of available licenses, which in the ASIC case are extremely expensive. On the contrary, FPGA vendors make their HLS tools free. Thus, it is tempting to investigate if FPGA HLS tools can be used to find the PODs in the ASIC case. To address this, in this work we present a dedicated multithreaded parallel HLS design space explorer (DSE) based on transfer learning that is able to accelerate HLS DSE for ASICs by targeting first FPGAs and using machine learning to convert the exploration results obtained to find the optimal ASIC equivalent. Experimental results show the effectiveness and robustness of our approach. M. Imtiaz Rashid, Benjamin Carrión Schäfer |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Improving the Quality of Hardware Accelerators through automatic Behavioral Input Language Conversion in HLSabstractHigh-Level Synthesis (HLS) is now part of most standard VLSI design flows and there are numerous commercial HLS tools available. One persistent problem of HLS is that the quality of results (QoR) still heavily depends on minor things like how the code is written. One additional observation that we have made in this work is that the input language used for the same HLS tool affects the QoR. HLS tools (commercial and academic) are built in a modular way which typically include a separate front-end (parser) for each input language supported. These front-ends parse the untimed behavioral descriptions, perform numerous technology independent optimizations and output a common intermediate representations (IR) for all different input languages supported. These optimizations also heavily depend on the synthesis directives set by the designer. These directives in the form of pragmas allow to control how to synthesize arrays (register or RAM), loops (unroll or not or pipeline) and functions (inline or not). We have observed that two functional equivalent behavioral descriptions with the same set of synthesis directives often lead to circuits with different QoR for the same HLS tool. Thus, automated approaches are needed to help designers to generate the best possible circuit independently of the input language used. To address this, in this work we propose using Graph Convolutional Networks (GCN) to determine the best language for a given new behavioral description and present an automated language converter for HLS. M. Imtiaz Rashid, Benjamin Carrión Schäfer |
ASP-DAC | 1 |
| 2022 | Fast Parallel High-Level Synthesis Design Space Explorer: Targeting FPGAs to accelerate ASIC ExplorationabstractRaising the level of VLSI design abstraction to the behavioral level allows to generate different micro-architectures from the same behavioral description by simply setting different synthesis options. These are typically synthesis directives in the form of pragmas that control how to synthesize arrays, loops, and functions. Out of all the combinations the designer is typically only interested in the synthesis directive combinations that lead to the Pareto-optimal designs. Unfortunately this multi-objective optimization problem grows supra-linearly with the number of the explorable operations. Thus, fast heuristics are needed. One additional way to accelerate the exploration process is by parallelizing the explorer tcreating multi-threaded versions. The main problem with this approach is that every time that a new pragma combination is generated the explorer requires to invoke the HLS process in order to evaluate the effect of these synthesis options on the resultant design. This tool invocation requires to check out a HLS tool license that will not be released until the HLS process has finished. This implies that the maximum number of parallel threads is limited by the number of licenses available. In the ASIC case, these licenses are extremely expensive, making it often prohibitory for some companies to have more than one. On contrary FPGA vendors provide their HLS tools free. Thus, it is tempting to investigate if FPGA HLS tools can be used to find the ASIC Pareto-optimal designs. To address this, in this work we present a dedicated multi-threaded parallel HLS DSE explorer that is able to accelerate HLS DSE for ASICs by targeting first FPGAs and using machine learning to convert the exploration results obtained to find the optimal ASIC equivalent. Experimental results show that our proposed approach is very efficient speedup up the exploration process considerably. M. Imtiaz Rashid, Benjamin Carrión Schäfer |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | Modernizing Hardware Circuits through High-Level SynthesisabstractThis works presents a design methodology to reoptimize legacy Register-Transfer level (RTL) designs specified in synthesizable Verilog or VHDL through High-Level Synthesis (HLS). The proposed methodology is based on an RTL to C compiler that converts synthesizable RTL descriptions into functional equivalent behavioral descriptions optimized to maximize its re-usability though HLS. This implies stripping off all the timing information from the RTL description and generating only C/C++ code that is functionally equivalent that has arrays, loops and functions. Generating these structures is very important in order to maximize the re-optimization potential as commercial HLS tools make extensive use of synthesis directives in the form or pragmas (comments) that allow HLS users to control how to synthesize them. E.g., loops can be fully unrolled, partially unrolled or pipelined, arrays can be synthesized as registers, memories or fully expanded into individual flip-flops and functions inline or not. Thus, generating C/C++ code with a larger number of these structures ensures that a larger variety unique implementations with different area vs. performance and power trade-offs can be generated from the converted C/C++ code. Experimental results with a variety of applications from different domains show the effectiveness of your proposed flow. M. Imtiaz Rashid, Qilin Si, Benjamin Carrión Schäfer |
ISCAS | 1 |
| 2021 | True Random Number Generation Using Latency Variations of FRAMabstractTrue random number generation (TRNG) plays an important role in security applications and protocols. In this article, we propose an effective technique to generate a robust true random number using emerging, energy-efficient, nonvolatile, consumer-off-the-shelf (COTS) ferroelectric random access memory (FRAM) chips. In the proposed method, we extract inherent randomness from internal ferroelectric capacitors by exploiting latency variation across cells within FRAM. Hardware results and subsequent National Institute of Standards and Technology (NIST) statistical test suite (STS) testing indicate that the proposed latency-based TRNG is robust over a wide range of operating conditions at speeds of 6.23 Mb/s using COTS, silicon FRAM chips from Cypress Semiconductor Corporation. M. Imtiaz Rashid, Farah Ferdaus, Bashir M. Sabquat Bahar Talukder, Paul Henny, Aubrey N. Beal, Md Tauhidur Rahman 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |