Ke Zhang 0017

dblp:20/4152-17 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0003-0206-7401ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SCOPE: SCalable and Observable PEripheral-Borrowing Framework for FPGA-Based Prototyping
Jiarun Yan, Congwu Zhang, Ke Zhang 0017
FCCM7
2026 Chrono-Fabric: A Decoupled Hierarchical Framework for Cycle-Accurate Coordination in Multi-FPGA Systems
Congwu Zhang, Panyu Wang, Bibo Yang, Mingyu Chen 0001, Yungang Bao, Ke Zhang 0017
FPGA7
2025 DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack Communication
Xu Zhang 0033, Ke Liu 0004, Yuan Hui 0001, Yisong Chang, Yizhou Shan, Ke Zhang 0017, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005
USENIX ATC8
2024 XUNI: Virtual Machine Abstraction for Self-contained and Multi-tenant Cloud FPGAs
abstract
FPGAs have become essential infrastructural components as well as publicly rentable resources in cloud and datacenters. Although one tenant is equipped with a couple of individual physical FPGA devices to solve a single problem, cloud FPGAs are still underutilized in many scenarios. Considering the economic cost of such heterogeneous computing resources, it is essential to explore opportunities for FPGA virtualization such that multiple tenants can share one physical device. Unlike the conventional host-centric virtualization approaches considering FPGAs as I/O peripherals, we propose XUNI, an FPGA-centric and self-contained virtual machine (VM) abstraction without involving the host-side virtualization techniques. Specifically, we design a hardware-software co-designed hypervisor for resource management and provisioning of various FPGA VMs. First, XUNI partitions the FPGA fabric into a series of reconfigurable regions that can be flexibly assembled for the deployment of large-scale designs. Second, we introduce both static partition and dynamic allocation schemes for FPGA-side DRAM sharing in XUNI. Last but not least, a hierarchical multi-tenant hardware network stack is built to provide an I/O interface for each FPGA VM. We implement XUNI and conduct infrastructural evaluations on a custom cloud FPGA node populated with an AMD/Xilinx Zynq MPSoC chip. Preliminary results demonstrate that XUNI is capable of handling hundreds of thousands of FPGA-VM-initiated memory requests per second. The hardware network stack exhibits line rate (~100Gbps) when receiving packets with the default 1500-byte MTU size. Moreover, each FPGA VM boots up within hundreds of milliseconds, which is comparable to emerging lightweight host-side VMs.
Guiyuan Zhu, Yunhai Liu, Yisong Chang, Ke Zhang 0017, Mingyu Chen 0001
FPGA5
2023 Rethinking Design Paradigm of Graph Processing System with a CXL-like Memory Semantic Fabric
abstract
With the evolution of network fabrics, message-passing clusters have been promising solutions for large-scale graph processing. Alternatively, the shared-memory model is also introduced to avoid redundant copies and extra storage space of graph data. Compared to conventional network fabrics, with the capability of fine-grained, byte-addressable remote memory access, emerging memory semantic interconnects and fabrics, e.g., Intel's Compute Express Link (CXL), are intuitively more appropriate for adoption in shared-memory clusters. However, due to the latency gap between local and remote memory, it is still challenging to take advantage of the shared-memory graph processing with memory semantic fabrics. To tackle this problem, in this paper, we first investigate memory access characterizations of graph vertex propagation based on the shared-memory model. Then we propose GraCXL, a series of design paradigms to address high-frequency and long-latency of remote memory access potentially incurred in CXL-based clusters. For system adaptiveness, we elaborate GraCXL towards the general-purpose CPU cluster and the domain-specific FPGA accelerator array, respectively. We design a custom fabric with the CXL.mem protocol and leverage a couple of ARM SoC-equipped FPGAs to build an evaluation prototype in the absence of commodity CXL hardware and platforms. Experimental results show that the proposed GraCXL CPU and FPGA clusters achieve 1.33x-8.92x and 2.48x-5.01x performance improvement, respectively.
Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Zhang 0017, Mingyu Chen 0001
CCGrid4
2023 REMU: Enabling Cost-Effective Checkpointing and Deterministic Replay in FPGA-based Emulation
abstract
Left-shifted integration and evaluation of hardware and software design are increasingly crucial in pre-silicon validation of processor-centric computing systems. With the inherent cycle-accurate deployment of target processor design in programmable logic fabrics, FPGA-based emulation has attracted academic attention for early-stage performance evaluations. However, it is difficult to conduct system-level inspection and debugging within open-source academic FPGA-based emulation frameworks due to the limited HW-SW visibility at run-time.To fill such a gap, we present REMU, an FPGA-based emulation framework enabling hardware checkpointing and deterministic replay to acquire bit-accurate visibility of target processors and other system components. Specially, we first employ a cost-effective scan-chain insertion method and related implementation strategies within an optimized open-source synthesis tool for status capturing of the emulated circuit primitives. Then, we introduce mechanisms in the design of emulated memory and I/O peripheral components to precisely describe behaviors and ensure deterministic replay of system-level interactions. Experimental results show that REMU drastically speeds up the scan-chain insertion flow by 1.6x-32.7x, and the proposed mechanisms for deterministic replay in the emulated external memory introduce negligible overhead in FPGA resource utilization.
Yuxiao Chen 0009, Yisong Chang, Ke Zhang 0017, Mingyu Chen 0001, Yungang Bao
ICCD3
2023 Morpheus: An Adaptive DRAM Cache with Online Granularity Adjustment for Disaggregated Memory
abstract
Disaggregated memory introduces a cost-effective solution for improving the memory utilization rate of data centers, by sharing a distributed memory pool among several individual servers. However, latency penalty in the existing connection between a computing node and the memory pool introduces performance degradation due to frequent far memory accesses. Based on our observation, page caching in the local DRAM, despite its reductions in the number of far memory accesses, still faces severe data over-fetching problem.With a detailed analysis of far memory access traces collected via several representative real-world applications, we argue that exploiting the various page-specific preferences of caching granularity is the key point of solving the data over-fetching problem in the DRAM cache. Consequently, in this paper, we present that it is influential to enable 1) dynamic selection of caching granularity for each page to not only guarantee sufficient spatial localities compared to the conventional fine-grained cache lines but also avoid data over-fetching caused by the coarse-grained pages, as well as 2) adaptive adjustment of cache capacity during execution for each granularity to accommodate the varying proportion of pages with different granularity preferences. Specifically, we propose Morpheus, an adaptive DRAM cache architecture that determines an optimal page-specific caching granularity at run-time and dynamically adjusts capacity occupations of different caching granularities. Based on our modeling and evaluations within the DRAMSim3 simulator, Morpheus exhibits 1.17-1.34x performance speedup for a wide range of workloads against the state-of-the-art DRAM cache design.
Xu Zhang 0033, Tianyue Lu, Yisong Chang, Ke Zhang 0017, Mingyu Chen 0001
ICCD4
2022 Increasing Flexibility of Cloud FPGA Virtualization
abstract
FPGA virtualization enables multiple tenants to share programmable hardware resources for application accelerations in cloud. However, such technique is still of limited usage in commercial FPGA cloud platforms, which mainly lies in: 1) absence of direct programming interfaces of the virtualized FPGA accelerators (vFPGAs) in tenants' virtual machines (VMs), 2) a fixed VM-vFPGA data movement scheme that is inadaptive to a wide range of data sizes among different applications, and 3) performance degradation due to unregulated inter-vFPGA competitions for limited shareable external resources (e.g., off-chip DRAM bandwidth). To tackle all the above issues, we propose a flexible FPGA virtualization framework and prototype an open cloud platform with ARM SoC-equipped FPGAs. Under such framework, tenants are allowed to directly initiate FPGA partial reconfiguration in isolated VMs via a direct I/O-like vFPGA device driver with as low as 20ms overhead. A hybrid data movement approach that leverages both memory-mapped I/O and DMA is also introduced in our framework to adaptively guarantee moderate VM-vFPGA bandwidth towards various data sizes. Moreover, a lightweight priority-based hardware scheduler is elaborated to monitor and dynamically allocate off-chip DRAM bandwidth among vFPGAs. Based on our preliminary infrastructure-level evaluation results, the proposed framework and the open prototyping are of significant interests to researchers looking forward to conducting further explorations in FPGA virtualization.
Jinjie Ruan, Yisong Chang, Ke Zhang 0017, Kan Shi, Mingyu Chen 0001, Yungang Bao
FPL3
2022 FPL Demo: SERVE: Agile Hardware Development Platform with Cloud IDE and Cloud FPGAs
abstract
We introduce SERVE, a cloud platform for agile hardware software co-design, with cloud IDE and cloud FPGAs integrated. SERVE enables users to focus on logic designs, without facing the hassle of setting up FPGA tools and development environment. Users can write and simulate hardware logic in the cloud IDE and then generate bitstream files through a Continuous Integration (CI) pipeline. Finally, the bitstream files are deployed on an FPGA board. A great amount of testbenches will be executed to ensure the correctness of the hardware logic. We will demo a workflow of modifying a RISC- V processor and getting the design change quickly evaluated using SERVE.
Ke Zhang 0017, Yisong Chang, Yanlong Yin, Yuxiao Chen 0009, Songyue Wang, Mingyu Chen 0001, Yungang Bao
FPL2
2022 GraFF: A Multi-FPGA System with Memory Semantic Fabric for Scalable Graph Processing
abstract
FPGA has been a promising solution for graph processing in many scenarios. With a rapid growth in graph size, the on/off-chip memory capacity of a single FPGA is insufficient to hold large-scale graphs. To tackle such problem, in this position paper, we introduce GraFF, a Graph processing system with multiple FPGAs interconnected via a custom memory semantic Fabric. In order to efficiently exploit system parallelism, we first split the traversal of graph data into a series of independent fine-grained flits that are concurrently delivered among FPGAs as sheer memory semantic transactions. Then we relax FPGAs' synchronization from strict barrier boundaries between adjacent supersteps to fully parallelize graph traversing and computing. We build a prototype of GraFF with four custom FPGA nodes. Preliminary evaluation result based on the Breadth First Search (BFS) algorithm shows that the peak performance of GraFF reaches up to 6.23 GTEPS. Moreover, GraFF exhibits linear scalability when the number of FPGAs rises from one to four.
Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Liu 0004, Ke Zhang 0017, Mingyu Chen 0001
FPT5
2021 EdUCAS: An In-house CI/CD Platform with Cloud FPGAs for Agilely Conducting Computer Systems Course Projects
abstract
In recent years, there has been a rapidly growing recognition of the importance of conducting hands-on HW-SW co-design labs with real hardware (e.g., programmable logic chips named FPGAs) while studying computing curricula, especially the computer systems (CSys) courses. However, using FPGA is quite atime-consuming and error-prone process for students, and manipulating FPGA development tools and boards also distracts students and instructors. To overcome these obstacles and improve agility, we introduce an in-house platform named EdUCAS, in combination with the previously designed cloud FPGA servers for students to automatically conduct computer systems course projects. EdUCAS aims at enabling students to concentrate on their logic designs using hardware description language (e.g., Verilog HDL), without wasting useless time in FPGA tools and experimental environment.
Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002
ITiCSE (2)1
2019 Engaging Heterogeneous FPGAs in the Cloud
abstract
FPGA has become an essential infrastructural component in commercial cloud and datacenter for improving system performance and efficiency. Meanwhile, a heterogeneous FPGA chip (Hetero-FPGA) in which a multi-core System-on-Chip (SoC) is tightly integrated with an FPGA fabric has been successfully pioneered. Given its hardware-software co-programmability, Hetero-FPGA is supposed to become an independent and first-class cloud computing resource with networking capabilities in order to avoid involving brawny commodity x86 servers as carriers for FPGA fabrics which are usually the cases in current commercial FPGA clouds from several web vendors. Following this design paradigm, we present HeFA, a self-contained Hetero-FPGA Array architecture in cloud. We construct a high-level hardware template as well as a software stack for the Hetero-FPGA node, enabling the SoC as a primary engine to manage, coordinate and incorporate with the dominant FPGA fabric. We also propose a fully scripted design flow to make HeFA as an easy-to-use cloud infrastructure. Based on these techniques, we implement an academia prototype chassis of HeFA that includes 32 Hetero-FPGA nodes with Xilinx's Zynq MPSoC chips. By a customized cloud resource manager, the prototype is flexibly provisioned as either 32 individual FPGA nodes or multiple scalable sub-clusters to abstract arbitrary volume of reconfigurable fabrics as on-demand cloud services. In this manner, a versatile research and educational platform is delivered for agile hardware-software co-design in scenarios such as domain-specific accelerator development, open instruction set architecture-based chip design, computer system-related experimental project, and so on.
Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002
FPGA1
2019 Computer Organization and Design Course with FPGA Cloud
abstract
Computer Organization and Design (COD) is a fundamentally required early-stage undergraduate course in most computer science and engineering curricula. During the two sessions (lecture and project part) of one COD course, educational platforms play an important role in cultivating students' computational thinking, especially the ability of viewing the hardware and software in a computer system as a whole (computer system thinking ability for short in this paper). In order to improve teaching quality, in this paper, we discuss the deployment of an inexpensive in-house Field Programmable Gate Array (FPGA) cloud platform, which can provide students with hardware-software co-design methodology and practice. The platform includes 32 FPGA nodes and the scale can be dynamically changed. Each cloud node is heterogeneously composed of an ARM processor and a tightly-coupled reconfigurable fabric to provide students with hands-on hardware and software programming experiences. We illustrate our efforts to make the FPGA cloud as an easy-to-use resource pool to elastically support a class with 92 undergrads via Internet access and to monitor students' experimental behaviors. We also present key insights in our teaching activities that indicate such appliance is feasible to provide practice of both basic principles and emerging co-design techniques for students. We believe that our cost-effective FPGA cloud is of significant interests to educators looking forward to improving computer system-related courses.
Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002
SIGCSE1
2018 On Retargeting the AI Programming Framework to New Hardwares
Yisong Chang, Denghui Li, Chunwei Xia, Huimin Cui, Ke Zhang 0017, Xiaobing Feng 0002
NPC6
2018 ZyForce: An FPGA-based Cloud Platform for Experimental Curriculum of Computer System in University of Chinese Academy of Sciences (Abstract Only)
abstract
To cultivate students" capability of computer system thinking and software/hardware programming, experimental curriculum of computer system is regarded as one of the most effective methods. Some universities have set up hardware labs equipped with several or dozens of FPGA (Field Programmable Gate Array) boards for these courses. However, these lab kits are always in a relatively low utilization rate and how the students" capability is improved by these assets is not easy to be evaluated. Inspired by the merits of FPGA public cloud (e.g. Amazon AWS F1 instance), an in-house-designed FPGA-based online cloud platform (named ZyForce) is proposed to deploy in UCAS. This platform is equipped with 40 custom designed boards using Xilinx Zynq UltraScale+ MPSoC FPGAs and the utilization rate of these education resources is boosted by means of advanced cloud computing technology. With ZyForce, students remotely carry out lab assignments (e.g. MIPS, RISC-V or domain-specific architecture processor design with cache/memory, DMA, accelerator and performance counter) as using local FPGA boards; instructors can analyze the downloaded operation log file for each student and know how these kits are being used. It"s believed that this kind of online hardware lab appliances provides a novel pay-as-you-go service model for those universities in remote regions who cannot afford to set up their own hardware laboratories, and also facilitates our students, the future scientists and engineers, with this promising cloud development approach.
Ke Zhang 0017, Mingyu Chen 0001, Yungang Bao
SIGCSE1
2016 sAXI: A High-Efficient Hardware Inter-Node Link in ARM Server for Remote Memory Access
abstract
The ever-growing need for fast big-data operations has made in-memory processing increasingly important in modern datacenters. To mitigate the capacity limitation of a single server node, techniques of inner-rack cross-node memory access have drawn attention recently. However, existing proposals exhibit inefficiency in remote memory access among server nodes due to inter-protocol conversions and non-transparent coarse-grained accesses. In this study, we propose the high-performance and efficient serialized AXI (sAXI) link and its associated cross-node memory access mechanism for emerging ARM-based servers. The key idea behind sAXI is directly extending the on-chip AMBA AXI-4.0 interconnection of the SoC in a local server node to the outside, and then bringing into remote server nodes via high-speed serial lanes. As a result, natively accessing remote memory in adjacent nodes in the same manner of local assets is supported by purely using existing software. Experimental results show that, using the sAXI data-path, performance of remote memory access in the user-level micro-benchmark is very promising (min. latency: 1.16µs, max. bandwidth: 1.52GB/s on our in-house FPGA prototype). In addition, through this efficient hardware inter-node link, performance of an in-memory key-value framework, Redis, can be improved up to 1.72x and large latency overhead of database query can be effectively hidden.
Ke Zhang 0017, Yisong Chang, Lixin Zhang 0002, Mingyu Chen 0001, Zhiwei Xu 0002
CCGrid1
2016 Extending On-chip Interconnects for rack-level remote resource access
abstract
The need to perform data analytics on exploding data volumes coupled with the rapidly changing workloads in cloud computing places great pressure on data-center servers. To improve hardware resource utilization across servers within a rack, we propose Direct Extension of On-chip Interconnects (DEOI), a high-performance and efficient architecture for remote resource access among server nodes. DEOI extends an SoC server node's on-chip interconnect to access resources in adjacent nodes with no protocol changes, allowing remote memory and network resources to be used as if they were local. Our results on a four-node FPGA prototype show that the latency of user-level, cross-node, random reads to DEOI-connected remote memory is as low as 1.16µs, which beats current commercial technologies. We exploit DEOI remote access to improve performance of the Redis in-memory key-value framework by 47%. When using DEOI to access remote network resources, we observe an 8.4% average performance degradation and only a 2.52µs ping-pong latency disparity compared to using local assets. These results suggest that DEOI can be a promising mechanism for increasing both performance and efficiency in next-generation data-center servers.
Yisong Chang, Ke Zhang 0017, Sally A. McKee, Lixin Zhang 0002, Mingyu Chen 0001, Liqiang Ren, Zhiwei Xu 0002
ICCD2