EDBT 2026 Demo / reviewers in the wild / expert
Jiacheng Ma 0001
dblp:182/6469
· DBLP profile ↗
14ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-9285-422XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MemTunnel: A CXL-Based Rack-Scale Host Memory Pooling Architecture for Cloud ServiceabstractMemory underutilization poses a significant challenge in cloud services, leading to performance inefficiencies and resource wastage. The tightly coupled computing and memory resources in cloud servers are identified as the root cause of this problem. To address this issue, memory pooling has been the subject of extensive research for decades, providing centralized or distributed shared memory pools as flexible memory resources for various applications running on different servers. However, existing memory disaggregation solutions sacrifice memory resources, add extra hardware (such as memory boxes/blades/drives), and degrade memory performance to achieve flexibility. To overcome these limitations, this paper proposes MemTunnel, a rack-scale host memory pooling architecture that provides a low-cost memory pooling solution based on Compute Express Link (CXL). MemTunnel is the first hardware and software architecture to offer symmetric, memory-semantic memory pooling over CXL, with an FPGA-based platform to demonstrate its feasibility in a real implementation. MemTunnel is orthogonal to the existing CXL-based memory pool and provides an additional layer of abstraction for memory disaggregation. Evaluation results show that MemTunnel achieves comparable performance to the existing CXL-based memory pool for a single machine and provides better rack-scale performance with minor hardware overheads. Tianchan Guan, Yijin Guan, Zhaoyang Du, Jiacheng Ma 0001, Boyu Tian, Teng Ma 0006, Zheng Liu 0022, Yuan Xie 0001, Mingyu Gao 0001, Guangyu Sun 0003, Hongzhong Zheng, Dimin Niu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | Proactive Runtime Detection of Aging-Related Silent Data Corruptions: A Bottom-Up ApproachabstractRecent advancements in semiconductor process technologies have unveiled the susceptibility of hardware circuits to reliability issues, especially those related to transistor aging. Transistor aging gradually degrades gate performance, eventually causing hardware to behave incorrectly. Such misbehaving hardware can result in silent data corruptions (SDCs) in software---a type of failure that comes without logs or exceptions, but causes miscomputing instructions, bitflips, and broken cache coherency. Alas, while design efforts can be made to mitigate transistor aging, complete elimination of this problem during design and fabrication cannot be guaranteed. This emerging challenge calls for a mechanism that not only detects potentially aged hardware in the field, but also triggers software mitigations at application runtime. Jiacheng Ma 0001, Majd Ganaiem, Madeline Burbage, Theo Gregersen, Rachel McAmis, Freddy Gabbay, Baris Kasikci |
ASPLOS (4) | 1 |
| 2023 | Vidi: Record Replay for Reconfigurable HardwareabstractDevelopers are turning to heterogeneous computing devices, such as Field Programmable Gate Arrays (FPGAs), to accelerate data center and cloud computing workloads. FPGAs enable rapid prototyping and should facilitate an agile software-like development workflow to fix correctness bugs, performance issues, and security vulnerabilities. Unfortunately, hardware development still does not have a vast ecosystem of tools needed to support the agile hardware development vision. The capability to record and replay FPGA executions would constitute a key building block that will inspire the development of many tools, similar to what record/replay did for software. However, building a practical record/replay tool for FPGA is challenging; existing approaches either record too much or too little information and cannot support real-world executions. Gefei Zuo, Jiacheng Ma 0001, Andrew Quinn 0001, Baris Kasikci |
ASPLOS (3) | 2 |
| 2022 | Debugging in the brave new world of reconfigurable hardwareabstractSoftware and hardware development cycles have traditionally been quite distinct. Software allows post-deployment patches, which leads to a rapid development cycle. In contrast, hardware bugs that are found after fabrication are extremely costly to fix (and sometimes even unfixable), so the traditional hardware development cycle involves massive investment in extensive simulation and formal verification. Reconfigurable hardware, such as a Field Programmable Gate Array (FPGA), promises to propel hardware development towards an agile software-like development approach, since it enables a hardware developer to patch bugs that are detected during on-chip testing or in production. Unfortunately, FPGA programmers lack bug localization tools amenable to this rapid development cycle, since past tools mainly find bugs via simulation and verification. To develop hardware bug localization tools for a rapid development cycle, a thorough understanding of the symptoms, root causes, and fixes of hardware bugs is needed. Jiacheng Ma 0001, Gefei Zuo, Kevin Loughlin, Andrew Quinn 0001, Baris Kasikci |
ASPLOS | 1 |
| 2021 | MEGATRON: Software-Managed Device TLB for Shared-Memory FPGA VirtualizationabstractFPGAs are being virtualized to improve resource utilization in data centers. Memory access performance is essential to FPGA hypervisors for shared-memory FPGA platform, where accelerators access memory spontaneously. DMA remapping with IOMMU provides a handy solution; however, fixed IOMMU can not benefit from the reconfigurability of FPGAs. In this work, we propose MEGATRON, a hybrid address translation service consisting of a hardware TLB and a software page table walker. By integrating MEGATRON into an existing FPGA hypervisor, we conduct a comprehensive analysis of link performance of a multi-link CPU-FPGA platform, and demonstrate the competitiveness of the customizable translation service. Yanqiang Liu, Jiacheng Ma 0001, Zhengjun Zhang, Linsheng Li, Zhengwei Qi, Haibing Guan |
DAC | 2 |
| 2021 | Execution reconstruction: harnessing failure reoccurrences for failure reproductionabstractReproducing production failures is crucial for software reliability. Alas, existing bug reproduction approaches are not suitable for production systems because they are not simultaneously efficient, effective, and accurate. In this work, we survey prior techniques and show that existing approaches over-prioritize a subset of these properties, and sacrifice the remaining ones. As a result, existing tools do not enable the plethora of proposed failure reproduction use-cases (e.g., debugging, security forensics, fuzzing) for production failures. Gefei Zuo, Jiacheng Ma 0001, Andrew Quinn 0001, Pramod Bhatotia, Pedro Fonseca 0001, Baris Kasikci |
PLDI | 2 |
| 2021 | DOLMA: Securing Speculation with the Principle of Transient Non-Observability
Kevin Loughlin, Ian Neal, Jiacheng Ma 0001, Elisa Tsai, Ofir Weisse, Satish Narayanasamy, Baris Kasikci |
USENIX Security Symposium | 3 |
| 2021 | gRemote: Cloud rendering on GPU resource pool based on API-forwarding
Dongjie Tang, Linsheng Li, Jiacheng Ma 0001, Xue (Steve) Liu, Zhengwei Qi, Haibing Guan |
J. Syst. Archit. | 3 |
| 2020 | A Hypervisor for Shared-Memory FPGA PlatformsabstractCloud providers widely deploy FPGAs as application-specific accelerators for customer use. These providers seek to multiplex their FPGAs among customers via virtualization, thereby reducing running costs. Unfortunately, most virtualization support is confined to FPGAs that expose a restrictive, host-centric programming model in which accelerators cannot issue direct memory accesses (DMAs). The host-centric model incurs high runtime overhead for workloads that exhibit pointer chasing. Thus, FPGAs are beginning to support a shared-memory programming model in which accelerators can issue DMAs. However, virtualization support for shared-memory FPGAs is limited. This paper presents Optimus, the first hypervisor that supports scalable shared-memory FPGA virtualization. Optimus offers both spatial multiplexing and temporal multiplexing to provide efficient and flexible sharing of each accelerator on an FPGA. To share the FPGA-CPU interconnect at a high clock frequency, Optimus implements a multiplexer tree. To isolate each guest's address space, Optimus introduces the technique of page table slicing as a hardware-software co-design. To support preemptive temporal multiplexing, Optimus provides an accelerator preemption interface. We show that Optimus supports eight physical accelerators on a single FPGA and improves the aggregate throughput of twelve real-world benchmarks by 1.98x-7x. Jiacheng Ma 0001, Gefei Zuo, Kevin Loughlin, Xiaohe Cheng, Yanqiang Liu, Abel Mulugeta Eneyew, Zhengwei Qi, Baris Kasikci |
ASPLOS | 1 |
| 2020 | gRemote: API-Forwarding Powered Cloud RenderingabstractTraditional GPU resource allocation approaches, widely adopted in today's data centers, only focus on the server-side functions while ignoring the client-side. These approaches waste client-side hardware resources. To solve this problem, remote API-forwarding architectures appear. Through running applications on the client-side, remote API-forwarding architectures offload some workloads to the client. However, many remote API-forwarding systems suffer from one big issue: shared-resource interference, stemming from two reasons: (a) GPU resource racing caused by resource overuse for a single client, and (b) CPU resource racing caused by resource shortage among clients. This paper presents gRemote, an open-source GPU-remoting system that can address this issue. To mitigate the CPU resource shortage, gRemote improves CPU configurations by expanding CPU resources from the server-side to both server- and client-side. To maintain the reasonable GPU usage for individual tasks, we innovate a new resource-sharing mechanism called GPU throttle. gRemote supports 1,228 OpenGL commands with around 10% shared-resource interference. Dongjie Tang, Yun Wang 0039, Linsheng Li, Jiacheng Ma 0001, Xue (Steve) Liu, Zhengwei Qi, Haibing Guan |
HPDC | 4 |
| 2020 | gMig: Efficient vGPU Live Migration with Overlapped Software-Based Dirty Page VerificationabstractThis paper introduces gMig, an open-source and practical vGPU live migration solution for full virtualization. Taking the advantage of the dirty pattern of GPU workloads, gMig presents the One-Shot Pre-Copy mechanism combined with the hashing based Software Dirty Page technique to achieve efficient vGPU live migration. Particularly, we propose three core techniques for gMig: 1) Dynamic Graphics Address Remapping, which parses and manipulates GPU commands to adjust the address mapping and adapt to a different environment after migration, 2) Software Dirty Page, which utilizes a hashing based approach with sampling pre-filtering to detect page modification, overcomes the commodity GPU's hardware limitation, and speeds up the migration by only sending the dirtied pages, 3) Overlapped Migration Process, which significantly compresses the hanging overhead by overlapping the dirty page verification and transmission concurrently. Our evaluation shows that gMig achieves GPU live migration with an average downtime of 302 ms on Windows and 119 ms on Linux. With the help of Software Dirty Page, the number of GPU pages transferred during the downtime is effectively reduced by up to 80.0 percent . The design of sampling filter and overlapped processing can bring about further 30.0 and 10.0 percent improvements in page processing. Qiumin Lu, Jiacheng Ma 0001, Yaozu Dong, Zhengwei Qi, Jianguo Yao 0002, Bingsheng He, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | gMig: Efficient GPU Live Migration Optimized by Software Dirty Page for Full VirtualizationabstractThis paper introduces gMig, an open-source and practical GPU live migration solution for full virtualization. By taking advantage of the dirty pattern of GPU workloads, gMig presents the One-Shot Pre-Copy combined with the hashing based Software Dirty Page technique to achieve efficient GPU live migration. Particularly, we propose three approaches for gMig: 1) Dynamic Graphics Address Remapping, which parses and manipulates GPU commands to adjust the address mapping to adapt to a different environment after migration, 2) Software Dirty Page, which utilizes a hashing based approach to detect page modification, overcomes the commodity GPU's hardware limitation, and speeds up the migration by only sending the dirtied pages, 3) One-Shot Pre-Copy, which greatly reduces the rounds of pre-copy of graphics memory. Our evaluation shows that gMig achieves GPU live migration with an average downtime of 302 ms on Windows and 119 ms on Linux. With the help of Software Dirty Page, the number of GPU pages transferred during the downtime is effectively reduced by 80.0%. Jiacheng Ma 0001, Yaozu Dong, Wentai Li, Zhengwei Qi, Bingsheng He, Haibing Guan |
VEE | 1 |
| 2018 | Scalable GPU Virtualization with Dynamic Sharing of Graphics Memory SpaceabstractWith increasing GPU-intensive workloads deployed on cloud, cloud service providers are seeking for practical and efficient GPU virtualization solutions. However, the cutting-edge GPU virtualization techniques such as gVirt still suffer from the restriction of scalability, which constrains the number of guest virtual GPU instances. This paper presents gScale, a scalable and practical open source GPU virtualization solution based on gVirt. gScale presents a sharing mechanism which combines partition and sharing together to break the hardware limitation of global graphics memory space. Particularly, we propose two approaches for gScale: (1) the private shadow graphics translation table (GTT) , which enables global graphics memory space sharing among virtual GPUs, (2) ladder mapping and fence memory space pool, which allows CPU access host physical memory space (serving the graphics memory) to bypass global graphics memory space. Furthermore, to mitigate the performance degradation caused by switching private shadow GTT when the number of vGPUs scales up, four other mechanisms are proposed: (1) slot sharing, which improves the performance of vGPU by dividing the high global graphics memory into multiple slots, (2) fine-grained slotting, which provides a flexible virtual graphics memory configuration, (3) predictive GTT copy mechanism, which reduces the performance loss by switching private shadow GTT before context switch, (4) predictive-copy aware scheduling, which maximizes the improvement of predictive GTT copy mechanism in cloud environment. Evaluation shows that gScale scales up to 15 guest virtual GPU instances in Linux or 12 guest virtual GPU instances in Windows, which is 5x and 4x, respectively, that of gVirt. At the same time, gScale incurs a slight but acceptable runtime overhead when hosting multiple virtual GPU instances. Mochi Xue, Jiacheng Ma 0001, Wentai Li, Yaozu Dong, Zhengwei Qi, Bingsheng He, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | gScale: Scaling up GPU Virtualization with Dynamic Sharing of Graphics Memory Space
Mochi Xue, Yaozu Dong, Jiacheng Ma 0001, Zhengwei Qi, Bingsheng He, Haibing Guan |
USENIX ATC | 4 |