Qilin Si

dblp:274/4989 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2025
0009-0005-2885-7553ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 6 since 2021
YearPublicationVenuePosition
2025 HAMMER: Hardware-aware Runtime Program Execution Acceleration through runtime reconfigurable CGRAs
abstract
This work introduces a novel computer architecture consisting of an embedded processor with a tightly coupled Coarse-grain Re-configurable Array (CGRA) that is able to accelerate the execution of sequential programs at runtime. To accomplish this, our work pre-characterizes a large variety of different portions of code from multiple application domains that can be accelerated offline. These kernels are subsequently synthesized onto the CGRA such that the proposed architecture detects at runtime if portions of code from a new unseen application can be accelerated or not at runtime. If they can be accelerated, then the system autonomously configures the CGRA with the specific accelerator, while if not present, then the code is executed sequentially on the CPU only. This approach implies sequential code compiled for a specific CPU only does not need to be recompiled for alternative architectures like CPU+CGRA.
Qilin Si, Benjamin Carrión Schäfer
ASP-DAC1
2023 PEPA: Performance Enhancement of Embedded Processors through HW Accelerator Resource Sharing
abstract
To improve the performance while reducing the power consumption, embedded processors in Systems-on-Chip (SoC) often now include tightly coupled hardware accelerators that can execute dedicated tasks orders of magnitude more efficiently (faster and lower power). These hardware accelerators though require significant hardware resources as one of the main reason for their efficiency is that they extensively exploit the parallelism of these dedicated tasks mapped on them. The question that we address in this work is if these hardware resources can be re-used by the CPU when executing a different application.
Qilin Si, Benjamin Carrión Schäfer
ACM Great Lakes Symposium on VLSI1
2023 ADVICE: Automatic Design and Optimization of Behavioral Application Specific Processors
abstract
Application Specific Instruction Set Processor (ASIPs) have been proposed in the past to increase the performance while reducing the energy of general-purpose processors. These ASIPs are normally generated at the RT-Level (Verilog or VHDL). In this work we leverage the advantages of High-Level Synthesis (HLS) by designing the complete ASIP in ANSI-C. HLS is a single process synthesis method, thus, the key is to merge the CPU and hardware accelerator descriptions. This allows us proposed flow to synthesize the entire system together, which has numerous advantages like being able to reduce the total area, while further minimizing the power as the HLS process can now fully co-optimize the ASIP by e.g., maximizing resource sharing.
Qilin Si, Benjamin Carrión Schäfer
ACM Great Lakes Symposium on VLSI1
2023 Application Specific Approximate Behavioral Processor
abstract
Many applications require simple controllers that continuously run the same application. These applications are often found in battery operated embedded systems that require to be ultra-low power (ULP) and are very price sensitive. Some examples include IoT devices of different nature and medical devices. Currently, these systems rely on off-the-shelf general-purpose microprocessors. One of the problems of using these processors, is that not all of the resources are needed for a specific application. Furthermore, because of the regularity of the workloads running on these systems there is a large opportunity to optimize the processor by pruning those unused resources to achieve lower area (cost) and power. Moreover, these processors can be specified at the behavioral level and use High-Level Synthesis (HLS) to generate an efficient Register Transfer Level (RTL) description. This opens a window to additional optimizations as the processor implementation is fully re-optimized during the HLS process. Also, many applications running on these embedded systems tolerate imprecise outputs. These include image processing and digital signal processing (DSP) applications. This opens the door to further optimizations in the context of approximate computing. To address these issues, this work presents a methodology to customize a behavioral RISC processor automatically for a given workload such that its area and power are significantly reduced as compared to the original, general-purpose processor. First, generating a bespoke processor that leads to the exact output as compared to the original general-purpose one and then by approximating it allowing a certain level of error at the output. Compared to previous work that customizes a given processor at the gate netlist only, our proposed method shows significant benefits. In particular, this work shows that raising the level of abstraction reduces the area and power by 78.3% and 70.1% for the exact solution on average, and further reduces the area by an additional 10.0% and 16.5% for the approximate version tolerating a maximum of 10% and 20% output errors respectively.
Qilin Si, Prattay Chowdhury, Rohit Sreekumar, Benjamin Carrión Schäfer
IEEE Trans. Sustain. Comput.1
2022 Optimizing Behavioral Near On-Chip Memory Computing Systems
abstract
This work presents an automated design and optimization flow for near on-chip memory computing systems by placing dedicated hardware accelerators directly next to the onchip memory. The salient feature of our proposed flow is that it allows the design of these complex systems completely at the behavioral level, thus, allowing a much richer set of optimizations than traditional Register Transfer Level (RTL) based approaches. Moreover, raising the level of design abstraction allows to quickly evaluate the effect of different optimization on the overall area, performance and power. In addition, it allows to quickly generate system with particular area and performance trade-offs by simply setting different synthesis options' combinations. Experimental results setting different constraints show the effectiveness of our proposed approach.
Qilin Si, Benjamin Carrión Schäfer
ASAP1
2022 Modernizing Hardware Circuits through High-Level Synthesis
abstract
This works presents a design methodology to reoptimize legacy Register-Transfer level (RTL) designs specified in synthesizable Verilog or VHDL through High-Level Synthesis (HLS). The proposed methodology is based on an RTL to C compiler that converts synthesizable RTL descriptions into functional equivalent behavioral descriptions optimized to maximize its re-usability though HLS. This implies stripping off all the timing information from the RTL description and generating only C/C++ code that is functionally equivalent that has arrays, loops and functions. Generating these structures is very important in order to maximize the re-optimization potential as commercial HLS tools make extensive use of synthesis directives in the form or pragmas (comments) that allow HLS users to control how to synthesize them. E.g., loops can be fully unrolled, partially unrolled or pipelined, arrays can be synthesized as registers, memories or fully expanded into individual flip-flops and functions inline or not. Thus, generating C/C++ code with a larger number of these structures ensures that a larger variety unique implementations with different area vs. performance and power trade-offs can be generated from the converted C/C++ code. Experimental results with a variety of applications from different domains show the effectiveness of your proposed flow.
M. Imtiaz Rashid, Qilin Si, Benjamin Carrión Schäfer
ISCAS2