Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gai Liu

dblp:149/3915 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
3since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 8 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Electronic design automation · 53% Reconfigurable computing and FPGAs · 16% Processor architecture and microarchitecture · 16%
Theoretical computer science
3 papers
Mathematical optimization · 53% Graph algorithms and graph theory · 47%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
high-level synthesis
1.962019
Rapid Generation of High-Qality RISC-V Processors from Functional Instruction Set Specifications · DAC 2019
Rosetta: A Realistic High-Level Synthesis Benchmark Suite for Software Programmable FPGAs · FPGA 2018
A Scalable Approach to Exact Resource-Constrained Scheduling Based on a Joint SDC and SAT Formulation · FPGA 2018
High-performance computing › performance optimization
auto-tuning
0.622018
DATuner: An Extensible Distributed Autotuning Framework for FPGA Design and Design Automation: (Abstract Only) · FPGA 2018
A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017
Electronic design automation
design space exploration
0.422018
DATuner: An Extensible Distributed Autotuning Framework for FPGA Design and Design Automation: (Abstract Only) · FPGA 2018
A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017
Processor architecture and microarchitecture › instruction set architecture › instruction set customization
domain-specific ISA
0.412019
Rapid Generation of High-Qality RISC-V Processors from Functional Instruction Set Specifications · DAC 2019
Processor architecture and microarchitecture › instruction set architecture
instruction set design
0.412019
Rapid Generation of High-Qality RISC-V Processors from Functional Instruction Set Specifications · DAC 2019
Electronic design automation › high-level synthesis
RTL generation
0.412019
Rapid Generation of High-Qality RISC-V Processors from Functional Instruction Set Specifications · DAC 2019
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools
0.312018
DATuner: An Extensible Distributed Autotuning Framework for FPGA Design and Design Automation: (Abstract Only) · FPGA 2018
Electronic design automation › high-level synthesis › scheduling
resource-constrained scheduling
0.312018
A Scalable Approach to Exact Resource-Constrained Scheduling Based on a Joint SDC and SAT Formulation · FPGA 2018
Electronic design automation › design optimization
area optimization
0.312017
A Parallelized Iterative Improvement Approach to Area Optimization for LUT-Based Technology Mapping · FPGA 2017
Electronic design automation › physical design
circuit partitioning
0.312017
A Parallelized Iterative Improvement Approach to Area Optimization for LUT-Based Technology Mapping · FPGA 2017
Reconfigurable computing and FPGAs
FPGA accelerator
0.312017
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Reconfigurable computing and FPGAs
FPGA compilation
0.312017
A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017
Electronic design automation
logic synthesis
0.312017
A Parallelized Iterative Improvement Approach to Area Optimization for LUT-Based Technology Mapping · FPGA 2017
Electronic design automation › high-level synthesis › pipeline synthesis
loop pipelining
0.312017
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Memory systems › memory architecture
multi-bank memory
0.312017
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Processor architecture and microarchitecture
pipelining
0.312017
Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017
Reconfigurable computing and FPGAs › coarse-grained reconfigurable architecture
processing element array
0.312017
Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Electronic design automation › logic synthesis
technology mapping
0.312017
A Parallelized Iterative Improvement Approach to Area Optimization for LUT-Based Technology Mapping · FPGA 2017
Mathematical optimization › online optimization
bandit optimization
0.312017
A Parallel Bandit-Based Approach for Autotuning FPGA Compilation · FPGA 2017
Reconfigurable computing and FPGAs
latency-insensitive interface
0.212016
Improving high-level synthesis with decoupled data structure optimization · DAC 2016
Emerging computing paradigms
analog computing
0.212015
A reconfigurable analog substrate for highly efficient maximum flow computation · DAC 2015
Graph algorithms and graph theory
graph algorithms
0.212015
A reconfigurable analog substrate for highly efficient maximum flow computation · DAC 2015
Graph algorithms and graph theory › graph algorithms › network flow
maximum flow
0.212015
A reconfigurable analog substrate for highly efficient maximum flow computation · DAC 2015
Compilers and program optimization
loop optimization
0.212014
Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014
Processor architecture and microarchitecture › adaptive architecture
adaptive microarchitecture
0.212014
Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014
Processor architecture and microarchitecture
instruction set architecture
0.212014
Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014
Hardware accelerators and domain-specific architectures › domain-specific accelerator
application-specific processor
0.112019
Rapid Generation of High-Qality RISC-V Processors from Functional Instruction Set Specifications · DAC 2019
Distributed systems › peer-to-peer systems
distributed search
0.112018
DATuner: An Extensible Distributed Autotuning Framework for FPGA Design and Design Automation: (Abstract Only) · FPGA 2018
Mathematical optimization
discrete optimization
0.112018
A Scalable Approach to Exact Resource-Constrained Scheduling Based on a Joint SDC and SAT Formulation · FPGA 2018
Mathematical optimization
integer programming
0.112018
A Scalable Approach to Exact Resource-Constrained Scheduling Based on a Joint SDC and SAT Formulation · FPGA 2018

Methods — techniques the papers use, named apart from their topics

multi-armed bandit · 0.9system of integer difference constraints · 0.7conflict-driven learning · 0.7boolean satisfiability · 0.7functional instruction set specification · 0.4parallel search · 0.3high-level synthesis optimization · 0.3parallelization · 0.3parallel autotuning · 0.3iterative improvement · 0.3hazard resolution · 0.3memristor crossbar · 0.2analog circuit design · 0.2simulation · 0.2compiler transformation · 0.2
YearPublicationVenuePosition
2023 Linker Code Size Optimization for Native Mobile Applications
abstract
Modern mobile applications have grown rapidly in binary size, which restricts user growth and hinders updates for existing users. Thus, reducing the binary size is important for application developers. Recent studies have shown the possibility of using link-time code size optimizations by re-invoking certain compiler optimizations on the linked intermediate representation of the program. However, such methods often incur significant build time overhead and require intrusive changes to the existing build pipeline.
Gai Liu, Umar Farooq 0002, Chengyan Zhao, Nian Sun
CC1
2023 Geometric Exterior Elements Calibration of Jilin-1 Linear Array Satellites Based on Star Observation
abstract
The exterior elements of the linear array satellite will change over time, resulting in significant degradation of the geometric positioning accuracy of the image. It is necessary to conduct geometric calibration of the camera in a quick and timely manner. Therefore, this study proposes a geometric calibration method through a star observation which satellite could be implemented at any position in orbit. The linear array camera point to the deep space and aim at the star for shooting. The star coordinate extracted from the image was regarded as the control point to calibrate the exterior elements of the camera. Then the geometric positioning accuracy of the image is improved. In this study, several Jilin-1 linear array satellites have been verified. The satellites with different launch times were used for star observation, and the geometric positioning accuracy was better than 50 m after the calibration by star observation method.
Zhichao Guan 0001, Xing Zhong, Guo Zhang 0001, Yonghua Jiang 0001, Gai Liu
IGARSS5
2022 FPGA HLS Today: Successes, Challenges, and Opportunities
abstract
The year 2011 marked an important transition for FPGA high-level synthesis (HLS), as it went from prototyping to deployment. A decade later, in this article, we assess the progress of the deployment of HLS technology and highlight the successes in several application domains, including deep learning, video transcoding, graph processing, and genome sequencing. We also discuss the challenges faced by today’s HLS technology and the opportunities for further research and development, especially in the areas of achieving high clock frequency, coping with complex pragmas and system integration, legacy code transformation, building on open source HLS infrastructures, supporting domain-specific languages, and standardization. It is our hope that this article will inspire more research on FPGA HLS and bring it to a new height.
Jason Cong, Jason Lau, Gai Liu, Stephen Neuendorffer, Peichen Pan, Kees A. Vissers, Zhiru Zhang
ACM Trans. Reconfigurable Technol. Syst.3
2019 Rapid Generation of High-Qality RISC-V Processors from Functional Instruction Set Specifications
abstract
The increasing popularity of compute acceleration for emerging domains such as artificial intelligence and computer vision has led to the growing need for domain-specific accelerators, often implemented as specialized processors that execute a set of domain-optimized instructions. The ability to rapidly explore (1) various possibilities of the customized instruction set, and (2) its corresponding micro-architectural features is critical to achieve the best quality-of-results (QoRs). However, this ability is frequently hindered by the manual design process at the register transfer level (RTL). Such an RTL-based methodology is often expensive and slow to react when the design specifications change at the instruction-set level and/or micro-architectural level.
Gai Liu, Joseph Primmer, Zhiru Zhang
DAC1
2019 PIMap: A Flexible Framework for Improving LUT-Based Technology Mapping via Parallelized Iterative Optimization
abstract
Modern FPGA synthesis tools typically apply a predetermined sequence of logic optimizations on the input logic network before carrying out technology mapping. While the “known recipes” of logic transformations often lead to improved mapping results, there remains a nontrivial gap between the quality metrics driving the pre-mapping logic optimizations and those targeted by the actual technology mapping. Needless to mention, such miscorrelations would eventually result in suboptimal quality of results. In this article, we propose PIMap, which couples logic transformations and technology mapping under an iterative improvement framework for LUT-based FPGAs. In each iteration, PIMap randomly proposes a transformation on the given logic network from an ensemble of candidate optimizations; it then invokes technology mapping and makes use of the mapping result to determine the likelihood of accepting the proposed transformation. By adjusting the optimization objective and incorporating required time constraints during the iterative process, PIMap can flexibly optimize for different objectives including area minimization, delay optimization, and delay-constrained area reduction. To mitigate the runtime overhead, we further introduce parallelization techniques to decompose a large design into multiple smaller sub-netlists that can be optimized simultaneously. Experimental results show that PIMap achieves promising quality improvement over a set of commonly used benchmarks, including improving the majority of the best-known area and delay records for the EPFL benchmark suite.
Gai Liu, Zhiru Zhang
ACM Trans. Reconfigurable Technol. Syst.1
2018 A Scalable Approach to Exact Resource-Constrained Scheduling Based on a Joint SDC and SAT Formulation
abstract
Despite increasing adoption of high-level synthesis (HLS) for its design productivity advantage, success in achieving high quality-of-results out-of-the-box is often hindered by the inexactness of the common HLS optimizations. In particular, while scheduling forms the algorithmic core to HLS technology, current scheduling algorithms rely heavily on fundamentally inexact heuristics that make ad hoc local decisions and cannot accurately and globally optimize over a rich set of constraints. To tackle this challenge, we propose a scheduling formulation based on system of integer difference constraints (SDC) and Boolean satisfiability (SAT) to exactly handle a variety of scheduling constraints. We develop a specialized scheduler based on conflict-driven learning and problem-specific knowledge to optimally and efficiently solve the resource-constrained scheduling problem. By leveraging the efficiency of SDC algorithms and scalability of modern SAT solvers, our scheduling technique is able to achieve on average over 100x improvement in runtime over the integer linear programming (ILP) approach while attaining optimal latency. By integrating our scheduling formulation into a state-of-the-art open-source HLS tool, we further demonstrate the applicability of our scheduling technique with a suite of representative benchmarks targeting FPGAs.
Steve Dai, Gai Liu, Zhiru Zhang
FPGA2
2018 DATuner: An Extensible Distributed Autotuning Framework for FPGA Design and Design Automation: (Abstract Only)
abstract
Mainstream FPGA tools contain an extensive set of user-controlled compilation options and internal optimization strategies that significantly impact the design quality. These compilation and optimization parameters create a complex design space that human designers may not be able to effectively explore in a time-efficient manner. In this work we describe DATuner, an open-source extensible distributed autotuning framework for optimizing FPGA designs and design automation tools using an ensemble of search techniques managed by multi-armed bandit algorithms. DATuner is designed for a distributed environment that uses parallel searches to amortize the significant runtime overhead of the CAD tools. DATuner provides convenient interface for extension to user-supplied tools, which enables the end users to apply DATuner to design tools/flows of their interest. We demonstrate the effectiveness and extensibility of DATuner using three case studies, which include clock frequency optimization for FPGA compilation, fixed-point optimization, and autotuning logic synthesis transformations.
Gai Liu, Ecenur Ustun, Shaojie Xiang, Chang Xu 0005, Guojie Luo, Zhiru Zhang
FPGA1
2018 Rosetta: A Realistic High-Level Synthesis Benchmark Suite for Software Programmable FPGAs
abstract
Modern high-level synthesis (HLS) tools greatly reduce the turn-around time of designing and implementing complex FPGA-based accelerators. They also expose various optimization opportunities, which cannot be easily explored at the register-transfer level. With the increasing adoption of the HLS design methodology and continued advances of synthesis optimization, there is a growing need for realistic benchmarks to (1) facilitate comparisons between tools, (2) evaluate and stress-test new synthesis techniques, and (3) establish meaningful performance baselines to track progress of the HLS technology. While several HLS benchmark suites already exist, they are primarily comprised of small textbook-style function kernels, instead of complete and complex applications. To address this limitation, we introduce Rosetta, a realistic benchmark suite for software programmable FPGAs. Designs in Rosetta are fully-developed applications. They are associated with realistic performance constraints, and optimized with advanced features of modern HLS tools. We believe that Rosetta is not only useful for the HLS research community, but can also serve as a set of design tutorials for non-expert HLS users. In this paper we describe the characteristics of our benchmarks and the optimization techniques applied to them. We further report experimental results on an embedded FPGA device as well as a cloud FPGA platform.
Udit Gupta 0001, Steve Dai, Ritchie Zhao, Nitish Kumar Srivastava, Hanchen Jin, Joseph Featherston, Yi-Hsiang Lai, Gai Liu, Gustavo Angarita Velasquez, Zhiru Zhang
FPGA9
2017 Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis
Steve Dai, Ritchie Zhao, Gai Liu, Shreesha Srinath, Udit Gupta 0001, Christopher Batten, Zhiru Zhang
FPGA3
2017 A Parallelized Iterative Improvement Approach to Area Optimization for LUT-Based Technology Mapping
Gai Liu, Zhiru Zhang
FPGA1
2017 A Parallel Bandit-Based Approach for Autotuning FPGA Compilation
Chang Xu 0005, Gai Liu, Ritchie Zhao, Guojie Luo, Zhiru Zhang
FPGA2
2017 Statistically certified approximate logic synthesis
abstract
Approximate logic synthesis generates inexact implementations of logic functions in exchange for better design qualities such as area, timing and power consumption. However, the error behavior of the approximate circuits (e.g., error rate or error magnitude) depends heavily on the specific synthesis technique as well as the input vectors, hindering end users from confidently adopting approximate designs. In this paper, we propose a statistically certified approximate logic synthesis framework using techniques from stochastic optimization, and integrate it into a state-of-the-art parallelized technology mapper. During the synthesis process, our framework continuously monitors the quality of the generated designs using statistical testing, leading to approximate designs that adhere to user-specified error constraints with a high confidence level. Experimental results demonstrate up to 10x area and timing improvements over the exact counterparts with an average of 0.2% deviation from the exact outputs.
Gai Liu, Zhiru Zhang
ICCAD1
2017 Architecture and Synthesis for Area-Efficient Pipelining of Irregular Loop Nests
abstract
Modern high-level synthesis (HLS) tools commonly employ pipelining to achieve efficient loop acceleration by overlapping the execution of successive loop iterations. While existing HLS pipelining techniques obtain good performance with low complexity for regular loop nests, they provide inadequate support for effectively synthesizing irregular loop nests. For loop nests with dynamic-bound inner loops, current pipelining techniques require unrolling of the inner loops, which is either very expensive in resource or even inapplicable due to dynamic loop bounds. To address this major limitation, this paper proposes ElasticFlow, a novel architecture capable of dynamically distributing inner loops to an array of processing units (LPUs) in an area-efficient manner. The proposed LPUs can be either specialized to execute an individual inner loop or shared among multiple inner loops to balance the tradeoff between performance and area. A customized banked memory architecture is proposed to coordinate memory accesses among different LPUs to maximize memory bandwidth without significantly increasing memory footprint. We evaluate ElasticFlow using a variety of real-life applications and demonstrate significant performance improvements over a state-of-the-art commercial HLS tool for Xilinx FPGAs.
Gai Liu, Mingxing Tan, Steve Dai, Ritchie Zhao, Zhiru Zhang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 Improving high-level synthesis with decoupled data structure optimization
abstract
Existing high-level synthesis (HLS) tools are mostly effective on algorithm-dominated programs that only use primitive data structures such as fixed size arrays and queues. However, many widely used data structures such as priority queues, heaps, and trees feature complex member methods with data-dependent work and irregular memory access patterns. These methods can be inlined to their call sites, but this does not address the aforementioned issues and may further complicate conventional HLS optimizations, resulting in a low-performance hardware implementation. To overcome this deficiency, we propose a novel HLS architectural template in which complex data structures are decoupled from the algorithm using a latency-insensitive interface. This enables overlapped execution of the algorithm and data structure methods, as well as parallel and out-of-order execution of independent methods on multiple decoupled lanes. Experimental results across a variety of real-life benchmarks show our approach is capable of achieving very promising speedups without causing significant area overhead.
Ritchie Zhao, Gai Liu, Shreesha Srinath, Christopher Batten, Zhiru Zhang
DAC2
2015 A reconfigurable analog substrate for highly efficient maximum flow computation
abstract
We present the design and analysis of a novel analog reconfigurable substrate that enables fast and efficient computation of maximum flow on directed graphs. The substrate is composed of memristors and standard analog circuit components, where the on/off states of the crossbar switches encode the graph topology. We show that upon convergence, the steady-state voltages in the circuit capture the solution to the maximum flow problem. We also provide techniques to minimize the impacts of variability and non-ideal circuit components on the solution quality, enabling practical implementation of the proposed substrate. Performance evaluation demonstrates orders of magnitude improvements in speed and energy efficiency compared to a standard CPU implementation.
Gai Liu, Zhiru Zhang
DAC1
2015 ElasticFlow: A Complexity-Effective Approach for Pipelining Irregular Loop Nests
abstract
Modern high-level synthesis (HLS) tools commonly employ pipelining to achieve efficient loop acceleration by overlapping the execution of successive loop iterations. However, existing HLS techniques provide inadequate support for pipelining irregular loop nests that contain dynamic-bound inner loops, where unrolling is either very expensive or not even applicable. To overcome this major limitation, we propose ElasticFlow, a novel architectural synthesis approach capable of dynamically distributing inner loops to an array of loop processing units (LPUs) in a complexity-effective manner. These LPUs can be either specialized to execute an individual loop or shared amongst multiple inner loops for area reduction. We evaluate ElasticFlow using a variety of real-life applications and demonstrate significant performance improvements over a widely used commercial HLS tool for Xilinx FPGAs.
Mingxing Tan, Gai Liu, Ritchie Zhao, Steve Dai, Zhiru Zhang
ICCAD2
2014 CASA: correlation-aware speculative adders
abstract
Speculative adders divide addition into subgroups and execute them in parallel for higher execution speed and energy efficiency, but at the risk of generating incorrect results. In this paper, we propose a lightweight correlation-aware speculative addition (CASA) method, which exploits the correlation between input data and carry-in values observed in real-life benchmarks to improve the accuracy of speculative adders. Experimental results show that applying the CASA method leads to a significant reduction in error rate with only marginal overhead in timing, area, and power consumption.
Gai Liu, Mingxing Tan, Zhiru Zhang
ISLPED1
2014 Architectural Specialization for Inter-Iteration Loop Dependence Patterns
abstract
Hardware specialization is an increasingly common technique to enable improved performance and energy efficiency in spite of the diminished benefits of technology scaling. This paper proposes a new approach called explicit loop specialization (XLOOPS) based on the idea of elegantly encoding inter-iteration loop dependence patterns in the instruction set. XLOOPS supports a variety of inter-iteration data-and control-dependence patterns for both single and nested loops. The XLOOPS hardware/software abstraction requires only lightweight changes to a general-purpose compiler to generate XLOOPS binaries and enables executing these binaries on: (1) traditional micro architectures with minimal performance impact, (2) specialized micro architectures to improve performance and/or energy efficiency, and (3) adaptive micro architectures that can seamlessly migrate loops between traditional and specialized execution to dynamically trade-off performance vs. Energy efficiency. We evaluate XLOOPS using a vertically integrated research methodology and show compelling performance and energy efficiency improvements compared to both simple and complex general-purpose processors.
Shreesha Srinath, Berkin Ilbeyi, Mingxing Tan, Gai Liu, Zhiru Zhang, Christopher Batten
MICRO4