Qian Zhao 0001

dblp:82/4299-1 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
1since 2021 · last 2025
0000-0003-0032-1974ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 52% Cloud and datacenter computing · 35% Electronic design automation · 13%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs › cloud FPGA
cloud FPGA acceleration
0.912025
Harmonia: A Unified Framework for Heterogeneous FPGA Acceleration in the Cloud · ASPLOS (2) 2025
Reconfigurable computing and FPGAs › FPGA architecture
heterogeneous FPGA
0.312025
Harmonia: A Unified Framework for Heterogeneous FPGA Acceleration in the Cloud · ASPLOS (2) 2025
Electronic design automation › physical design › placement and routing
FPGA placement and routing
0.212013
A novel FPGA design framework with VLSI post-routing performance analysis (abstract only) · FPGA 2013
Electronic design automation
physical design
0.212013
A novel FPGA design framework with VLSI post-routing performance analysis (abstract only) · FPGA 2013

Methods — techniques the papers use, named apart from their topics

shell-role architecture · 0.9object-oriented programming · 0.2HDL template generation · 0.2
YearPublicationVenuePosition
2025 Harmonia: A Unified Framework for Heterogeneous FPGA Acceleration in the Cloud
abstract
FPGAs are gaining popularity in the cloud as accelerators for various applications. To make FPGAs more accessible for users and streamline system management, cloud providers have widely adopted the shell-role architecture on their homogeneous FPGA servers. However, the increasing heterogeneity of cloud FPGAs poses new challenges for this architecture. Previous studies either focus on homogeneous FPGAs or only partially address the portability issues for roles, while still requiring laborious shell development for providers and ad-hoc software modifications for users.
Xinchen Wan, Zilong Wang 0007, Qian Zhao 0001, Feng Ning, Qingsong Ning, Shideng Zhang, Zhenyu Li 0001, Layong Luo, Gaogang Xie
ASPLOS (2)6
2017 hCODE 2.0: An open-source toolkit for building efficient FPGA-enabled clouds
abstract
Major cloud service providers have started employing field-programmable gate arrays (FPGAs) to implement high-performance and low-power-consumption cloud capability. However, building or utilizing an FPGA-enabled cloud is still challenging due to the lack of fundamental tools. In our previous work, we proposed an hCODE base system for managing portable accelerator IPs on different hardware. In this paper, we extend the previous work and introduce the hCODE 2.0, which is an open-source toolkit for building efficient FPGA-enabled clouds. First, we provide a fundamental toolkit to simplify HW project management and FPGA management at a cluster scale. Second, we implement on-chip resource virtualization and accelerator scheduling capabilities to show possibilities of improving FPGA utilization efficiency with our tools.
Qian Zhao 0001, Hendarmawan, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPT1
2016 hCODE: An open-source platform for FPGA accelerators
abstract
Field-programmable gate arrays (FPGAs) have demonstrated great speed performance and power efficiency advantages over conventional computers in various domains. However, it is still difficult for general software engineers to employ FPGA-based hardware accelerators because of the gap between hardware and software development methods. In this paper, we propose a heterogeneous computing oriented development environment (hCODE) to simplify the creation, sharing, and project integration of hardware accelerators. The hCODE defines hardware specifications and interface design rules. Hardware developers can provide designs that follow these rules, allowing software engineers to easily search, download, and integrate accelerators in their applications without caring about the details of the hardware.
Qian Zhao 0001, Takuya Nakamichi, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPT1
2015 Architecture exploration of 3D FPGA to minimize internal layer connection
abstract
A three-dimensional (3D) integration based on wafer-to-wafer bonding using through-silicon vias (TSVs) has been developed for the fabrication of new 3D large-scale integrated chips. To balance between cost and performance, and to explore 3D field-programmable gate array (FPGA) with realistic 3D integration processes, we propose spatially distributed and functionally distributed types of 3D FPGA architectures. The functionally distributed architecture consists of two wafers, a logic layer and a routing layer, and is stacked by a face-down process technology. Since vertical wires pass through microbumps, no TSVs are needed. In contrast, the spatially distributed architecture is divided into multiple layers with the same structure, unlike in the functionally distributed type. This architecture can be expanded to more than two layers by stacking multiples of the same die. The goal of this paper is to elucidate the advantages and disadvantages of these two types of 3D FPGAs. According to our evaluation, when only two layers are used, the functionally distributed architecture is more effective. When higher performance is achieved by using more than two layers, the spatially distributed architecture achieves better performance.
Motoki Amagasaki, Yuto Takeuchi, Qian Zhao 0001, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
VLSI-SoC3
2014 A logic cell architecture exploiting the shannon expansion for the reduction of configuration memory
abstract
Most modern field-programmable gate arrays (FPGAs) employ a look-up table (LUT) as their basic logic cell. Although a k-input LUT can implement any k-input logic, its functionality relies on a large amount of configuration memory. As FPGA scales improve, the increased quantity of configuration memory cells required for FPGAs will require a larger area and consume more power. Moreover, the soft-error rate per device will also increase as more configuration memory cells are embedded. We propose scalable logic modules (SLMs), logic cells requiring less configuration memory, reducing configuration memory by making use of partial functions of Shannon expansion for frequently appearing logics. Experimental results show that SLM-based FPGAs use much less configuration memory and have smaller area than conventional LUT-based FPGAs.
Qian Zhao 0001, Kyosei Yanagida, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPL1
2014 A novel three-dimensional FPGA architecture with high-speed serial communication links
abstract
Three-dimensional (3D) integrated circuit technology is expected to offer continual improvement to very-large-scale integration performance as the process of miniaturization approaches physical limits. However, because the through-silicon vias (TSVs) that are used to create interlayer vertical connections are much larger area than transistors, there is an inherent tradeoff between connectivity and small size. Field-programmable gate arrays (FPGAs) are particularly noted for requiring a high level of routing resources, which means that it is unrealistic to make the same number of connections vertically as horizontally. In previous research, we proposed a method for creating a two-layer compact 3D FPGA with face-down integration (the base FPGA). In this paper, we discuss stacking multiple base FPGAs by the face-up method and propose a method for achieving highspeed interlayer communications with TSV serial connections. The proposed architecture improves FPGA performance by using smaller TSVs. The evaluation results show that the proposed 3D FPGA can achieve a total area that is as low as 67% the equivalent two-dimensional FPGA.
Takuya Kajiwara, Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morituro Kuga, Toshinori Sueyoshi
FPT2
2013 A novel FPGA design framework with VLSI post-routing performance analysis (abstract only)
abstract
The most widely used open-source field-programmable gate array (FPGA) placement and routing tool is VPR, which can define the target FPGA, perform placement and routing, and report area and timing information. However, it cannot be used in FPGA IP design efficiently for two reasons. First, for most newly developed FPGA architectures, VPR cannot support them directly. Modifying the C-coded VPR for using it to evaluate a number of new architectures requires a long time. Second, the accuracy of the VPR performance results is not enough for the evaluation of a complete synthesizable FPGA IP in the design that targets the productions of LSI. We propose a FPGA design framework that in particular improves FPGA IP design efficiency. A novel FPGA routing tool is developed in this framework, namely EasyRouter. EasyRouter is developed using the C# language. When an object-oriented programming method is used, the source codes are fewer and easier manage compared to VPR, which shortens the development time. By using simple HDL templates, EasyRouter can automatically generate entire chip HDL codes and the configuration bitstream. With these files, the FPGA IP can be evaluated with commercial VLSI CADs with high accuracy and reliability.
Qian Zhao 0001, Kazuki Inoue, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPGA1
2013 Defect-robust FPGA architectures for intellectual property cores in system LSI
abstract
In this paper, we propose fault-tolerant field-programmable gate array (FPGA) architectures and their computer-aid design (CAD) for intellectual property (IP) cores in system large-scale integration (LSI). Unlike discrete FPGAs, in which the integration scale can be made relatively large, programmable IP cores must correspond to arrays of various sizes. The key features of our architectures are regular tile structure, spare modules and bypass wires for fault avoidance, and configuration mechanism for single-cycle reconfiguration. In addition, we develop routing tools, namely EasyRouter for proposed architecture. This tool can handle various array sizes corresponding to developed programmable IP cores. In this evaluation, we compared the performances of conventional FPGA and the proposed fault-tolerant FPGA architectures. On average, our architectures have less than 2.2 times the area and 1.3 times the delay compared with conventional FPGA architectures. At the same time, conventional FP-GAs cannot tolerate faults, whereas our architectures perform with a 90% success rate in fault avoidance for a ratio of faulty tiles of 1% or less.
Motoki Amagasaki, Kazuki Inoue, Qian Zhao 0001, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPL3
2013 An automatic FPGA design and implementation framework
abstract
Conventional FPGA design and implementation processes involve two separate flows. The FPGA architecture is determined by academic FPGA design flow. However, in the implementation phase, commercial VLSI design flow are used. In this research, we propose an FPGA design framework in order to improve synthesizable FPGA IP design efficiency. A novel FPGA routing tool is developed in this framework, namely the EasyRouter, which can bridge the two flows efficiently. With this design flow, accurate physical information can be reported when a new FPGA IP architecture is evaluated with reliable commercial VLSI CADs.
Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPL1
2013 Three-dimensional stacking FPGA architecture using face-to-face integration
abstract
In recent years, as VLSI process scales have developed into deep sub-micrometer dimensions, routing delay problems have become critical. For reconfigurable logic devices (RLDs) like field-programmable gate arrays (FPGAs) in particular, routing resources occupy major parts of the available area and hinder performance. In order to balance cost and performance, and to explore 3D FPGA architectures with realistic 3D LSI processes, we proposed a novel two-layers 3D FPGA architecture based on 3D connections on logic block input and output pins. Evaluation shows that this novel RLD with two layers of 3D routing architecture uses 48.75% less on-board area and 30.54% less critical path delay than does a conventional 2D 4-lookup table island-style FPGA on average.
Tetsuro Hamada, Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
VLSI-SoC2
2010 A robust reconfigurable logic device based on less configuration memory logic cell
abstract
As the size of integrated circuit has reached the nanoscale, embedded memories are more sensitive to single event upset (SEU), because of their low threshold voltage. In particular field-programmable gate arrays (FPGAs), which contain large amounts of configuration memories to implement customer circuits, are more likely to suffer from soft errors caused by SEU. In this research, we first develop a Hamming code based error detect and correct (EDC) circuit that can prevent the configuration memory of a reconfigurable device from SEU. We then propose a novel reconfigurable logic element, namely COGRE, which will use much less configuration memory than the conventional FPGA 4-, 5- or 6-LUTs (lookup tables). Evaluation revealed that compared to the 6-LUT FPGAs with triple modular redundancy (TMR) configuration memory blocks, the 5- and 6-input proposed architecture save about 75.44 and 74.29% memories on average, respectively. And the dependability of the proposed architectures is about 6.8 to 10 times better than the LUTs with a tile level TMR structure on average. Moreover, with the consideration of the on the fly scrubbing advantage of the EDC, SEUs cannot be accumulated, so a much higher dependability can be achieved.
Qian Zhao 0001, Yoshihiro Ichinomiya, Yasuhiro Okamoto, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi
FPT1
2010 A Variable-Grain Logic Cell and Routing Architecture for a Reconfigurable IP Core
abstract
In the present study, we investigate the use of reconfigurable logic devices (RLDs) as intellectual properties (IPs) for system on a chip (SoC). Using RLDs, SoCs can achieve both high performance and high flexibility. However, conventional RLDs have problems related to performance, area, and power consumption. In order to resolve these problems, we investigated the features of RLD architecture. RLDs are classified into fine-grained and coarse-grained devices based on their architecture. Generally, the granularity of an RLD is limited to either type, which means that a device can only achieve high performance in applications that are suited to its architecture. Therefore, we propose a variable-grain logic cell (VGLC) architecture that can overcome the trade-off between fine-grained and coarse-grained architectures, which are required for the implementation of random and arithmetic logics, respectively. The VGLC is based on a 4-bit adder including configuration bits, which can perform arithmetic and random logic operations unlike the LUT. In the present paper, a local interconnection architecture for the VGLC is proposed. Several types of local interconnections composed of different crossbars are compared, and the trade-off between hardware resources and flexibility is discussed. Using local interconnection, the routing area is reduced by a maximum of 49%.
Kazuki Inoue, Qian Zhao 0001, Yasuhiro Okamoto, Hiroki Yosho, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi
ACM Trans. Reconfigurable Technol. Syst.2