Shuangnan Liu

dblp:206/9096 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
0since 2021 · last 2020
0000-0001-6674-473XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Electronic design automation · 51% Reconfigurable computing and FPGAs · 32% Hardware accelerators and domain-specific architectures · 17%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA prototyping
0.822020
Predictive Compositional Method to Design and Reoptimize Complex Behavioral Dataflows · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Accelerating FPGA Prototyping through Predictive Model-Based HLS Design Space Exploration · DAC 2019
Electronic design automation
high-level synthesis
0.822020
Predictive Compositional Method to Design and Reoptimize Complex Behavioral Dataflows · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Accelerating FPGA Prototyping through Predictive Model-Based HLS Design Space Exploration · DAC 2019
Hardware accelerators and domain-specific architectures
dataflow optimization
0.412020
Predictive Compositional Method to Design and Reoptimize Complex Behavioral Dataflows · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation
design space exploration
0.412019
Accelerating FPGA Prototyping through Predictive Model-Based HLS Design Space Exploration · DAC 2019
Electronic design automation › high-level synthesis › behavioral transformation
behavioral synthesis
0.112020
Predictive Compositional Method to Design and Reoptimize Complex Behavioral Dataflows · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020

Methods — techniques the papers use, named apart from their topics

stream computing · 0.4compositional predictive model · 0.4predictive model · 0.4pareto optimization · 0.4
YearPublicationVenuePosition
2020 Predictive Compositional Method to Design and Reoptimize Complex Behavioral Dataflows
abstract
In this article, we introduce an automatic stream computing reoptimization flow from ASICs to field-programmable gate arrays (FPGAs). Complex VLSI designs need to be prototyped and/or emulated on FPGAs. The main problem that we address in this article is that configurations optimized when targeting ASICs are often, as we will show in this article, highly un-optimal when remapped onto an FPGA. Thus, this article proposes a method to first generate a variety of dataflow configurations targeting an ASIC given multiple behavioral descriptions for high-level synthesis (HLS) and then, based on a compositional predictive model, automatically reoptimize the dataflow when mapped onto an FPGA. The experimental results show that our proposed method works well and that it is very fast.
Shuangnan Liu, Francis C. M. Lau 0002, Benjamin Carrión Schäfer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Accelerating FPGA Prototyping through Predictive Model-Based HLS Design Space Exploration
abstract
One of the advantages of High-Level Synthesis (HLS), also called C-based VLSI-design, over traditional RT-level VLSI design flows, is that multiple micro-architectures of unique area vs. performance can be automatically generated by setting different synthesis options, typically in the form of synthesis directives specified as pragmas in the source code. This design space exploration (DSE) is very time-consuming and can easily take multiple days for complex designs. At the same time, and because of the complexity in designing large ASICs, verification teams now routinely make use of emulation and prototyping to test the circuit before the silicon is taped out. This also allows the embedded software designers to start their work earlier in the design process and thus, further reducing the Turn-Around-Times (TAT). In this work, we present a method to automatically re-optimize ASIC designs specified as behavioral descriptions for HLS to FPGAs for emulation and prototyping, based on the observation that synthesis directives that lead to efficient micro-architectures for ASICs, do not directly translate into optimal micro-architectures in FPGAs. This implies that the HLS DSE process would have to be completely repeated for the target FPGA. To avoid this, this work presents a predictive model-based method that takes as inputs the results of an ASIC HLS DSE and automatically, without the need to re-explore the behavioral description, finds the Pareto-optimal micro-architectures for the target FPGA. Experimental results comparing our predictive-model based method vs. completely re-exploring the search space show that our proposed method works well.
Shuangnan Liu, Francis C. M. Lau 0002, Benjamin Carrión Schäfer
DAC1
2018 Investigation and Optimization of Pin Multiplexing in High-Level Synthesis
abstract
This paper investigates the effect of pin multiplexing on the resultant micro-architecture of synthesizable behavioral descriptions for High-Level Synthesis (HLS). A method is presented to find the most efficient pin assignments by assigning multiple logic inputs and outputs to the same physical ports such that the performance degradation and area overhead is minimized. The proposed method is a fast heuristic based on the scheduling results of HLS seen as a black box and hence is flexible enough to work with any HLS tool. Experimental results show that our proposed method is very efficient compared to an exhaustive search and a simulated annealing method at a fraction of the time and much better than randomly selecting the pins to be multiplexed.
Shuangnan Liu, Francis C. M. Lau 0002, Benjamin Carrión Schäfer
ACM Great Lakes Symposium on VLSI1
2017 Learning-based interconnect-aware dataflow accelerator optimization
abstract
The interconnect is the Achilles heel of FPGAs. It currently dominates the delay and leads to high power consumption. It is thus, imperative to take it into account when designing complex FPGA systems. In this work, we propose a learning-based method for data-flow systems build out of multiple individual components directly connected and find a set of optimal configurations with unique area vs. throughput trade-offs by time-multiplexing their interconnects. These type of configurations are prevalent in FPGA designs where one block streams data to the next one, i.e., jpeg encoders. One uniqueness of this work is that it uses advanced features of state-of-the-art HLS tools that enable automatic pin multiplexing. This feature implies that logic IO ports are time multiplexed automatically, which affects not only the performance of the design, but also the area, and the interconnect complexity, leading to system configurations with unique area vs. performance trade-offs. Pin multiplexing is not feasible at the RT-level where each design has to be manually optimized. Experimental results show that the method is accurate and fast when compared to an exhaustive search as well as other state-of-the-art methods.
Shuangnan Liu, Benjamin Carrión Schäfer
FPL1