VLDB 2026 Research / reviewers in the wild / expert
Zhenyuan Ruan
dblp:204/5398
· DBLP profile ↗
18ranked-venue papers
7as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Systems, architecture and hardware · 4 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unleashing The Potential of Datacenter SSDs by Taming Performance Variability
Gohar Irfan Chaudhry, Ankit Bhardwaj 0002, Zhenyuan Ruan, Adam Belay |
NSDI | 3 |
| 2025 | VeriLocc: End-to-End Cross-Architecture Register Allocation via LLMabstractModern GPUs evolve rapidly, yet production compilers still rely on hand-crafted register allocation heuristics that require substantial retuning for each hardware generation.We introduce VERILOCC, a framework that combines large language models (LLMs) with formal compiler techniques to enable generalizable and verifiable register allocation across GPU architectures.VERILOCC fine-tunes an LLM to translate intermediate representations (MIRs) into target-specific register assignments, aided by static analysis for cross-architecture normalization and generalization and a verifierguided regeneration loop to ensure correctness.Evaluated on matrix multiplication (GEMM) and multi-head attention (MHA), VERILOCC achieves 85-99% single-shot accuracy and near-100% [email protected] study shows that VERILOCC discovers more performant assignments than expert-tuned libraries, outperforming rocBLAS by over 10% in runtime. Lesheng Jin, Zhenyuan Ruan, Haohui Mai, Jingbo Shang |
EMNLP | 2 |
| 2025 | Fast and Scalable Selective Retransmission for RDMA
Peihao Huang, Guo Chen 0001, Xin Zhang 0117, Huijun Shen, Ying Bian, Yuanwei Lu, Zhenyuan Ruan, Bojie Li, Jiansong Zhang 0001, Yongfeng Liu, Zhigang Chen 0001 |
INFOCOM | 9 |
| 2025 | Quicksand: Harnessing Stranded Datacenter Resources with Granular Computing
Zhenyuan Ruan, Kaiyan Fan, Seo Jin Park, Marcos K. Aguilera, Adam Belay, Malte Schwarzkopf |
NSDI | 1 |
| 2024 | Harvesting Idle Memory for Application-managed Soft State with Midas
Yifan Qiao 0002, Zhenyuan Ruan, Adam Belay, Miryung Kim, Guoqing Harry Xu |
NSDI | 2 |
| 2023 | Unleashing True Utility Computing with QuicksandabstractToday's clouds are inefficient: their utilization of resources like CPUs, GPUs, memory, and storage is low. This inefficiency occurs because applications consume resources at variable rates and ratios, while clouds offer resources at fixed rates and ratios. This mismatch of offering and consumption styles prevents fully realizing the utility computing vision. Zhenyuan Ruan, Kaiyan Fan, Marcos K. Aguilera, Adam Belay, Seo Jin Park, Malte Schwarzkopf |
HotOS | 1 |
| 2023 | Hermit: Low-Latency, High-Throughput, and Transparent Remote Memory via Feedback-Directed Asynchrony
Yifan Qiao 0002, Chenxi Wang 0005, Zhenyuan Ruan, Adam Belay, Qingda Lu, Yiying Zhang 0005, Miryung Kim, Guoqing Harry Xu |
NSDI | 3 |
| 2023 | Nu: Achieving Microsecond-Scale Resource Fungibility with Logical Processes
Zhenyuan Ruan, Seo Jin Park, Marcos K. Aguilera, Adam Belay, Malte Schwarzkopf |
NSDI | 1 |
| 2020 | Caladan: Mitigating Interference at Microsecond Timescales
Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam Belay |
OSDI | 2 |
| 2020 | AIFM: High-Performance, Application-Integrated Far Memory
Zhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, Adam Belay |
OSDI | 1 |
| 2020 | Semeru: A Memory-Disaggregated Managed Runtime
Chenxi Wang 0005, Yuanqi Li, Zhenyuan Ruan, Khanh Nguyen 0001, Michael D. Bond, Ravi Netravali, Miryung Kim, Guoqing Harry Xu |
OSDI | 5 |
| 2019 | Hardware Acceleration of Long Read Pairwise Overlapping in Genome Sequencing: A Race Between FPGA and GPUabstractIn genome sequencing, it is a crucial but time-consuming task to detect potential overlaps between any pair of the input reads, especially those that are ultra-long. The state-of-the-art overlapping tool Minimap2 outperforms other popular tools in speed and accuracy. It has a single computing hot-spot, chaining, that takes 70% of the time and needs to be accelerated. There are several crucial issues for hardware acceleration because of the nature of chaining. First, the original computation pattern is poorly parallelizable and a direct implementation will result in low utilization of parallel processing units. We propose a method to reorder the operation sequence that transforms the algorithm into a hardware-friendly form. Second, the large but variable sizes of input data make it hard to leverage task-level parallelism. Therefore, we customize a fine-grained task dispatching scheme which could keep parallel PEs busy while satisfying the on-chip memory restriction. Based on these optimizations, we map the algorithm to a fully pipelined streaming architecture on FPGA using HLS, which achieves significant performance improvement. The principles of our acceleration design apply to both FPGA and GPU. Compared to the multi-threading CPU baseline, our GPU accelerator achieves 7x acceleration, while our FPGA accelerator achieves 28x acceleration. We further conduct an architecture study to quantitatively analyze the architectural reason for the performance difference. The summarized insights could serve as a guide on choosing the proper hardware acceleration platform. Licheng Guo, Jason Lau, Zhenyuan Ruan, Peng Wei 0004, Jason Cong |
FCCM | 3 |
| 2019 | Analyzing and Modeling In-Storage Computing Workloads On EISC - An FPGA-Based System-Level Emulation PlatformabstractStorage drive technology has made continuous improvements over the last decade, shifting the bottleneck of the data processing system from the storage drive to host/drive interconnection. To overcome this “data movement wall,” people have proposed in-storage computing (ISC) architectures which add the computing unit directly into the storage drive. Rather than moving data from drive to host, it offloads computation from host to drive, thereby alleviating the interconnection bottleneck. Though existing work shows the effectiveness of ISC under some specific workloads, they have not tackled two critical issues: 1) ISC is still at the early research stage, and there is no available ISC device on the market. Researchers lack an effective way to accurately explore the benefits of ISC under different applications and different system parameters (drive performance and interconnection performance). 2) What kinds of applications can benefit from ISC, and what cannot? It is crucial to have a method to quickly discriminate between the types of applications before spending significant efforts to implement them. This paper gives a response to the above problems. First, we build a complete FPGA-based ISC emulation system to enable rapid exploration. To the best of our knowledge, it is the first open-source11https://github.com/zainryan/EISC, publicly accessible ISC emulation system. Second, we use our system to evaluate 12 common applications. The results give us the basic criteria for choosing ISC-friendly applications. By assuming a general drive program construct, we provide further insights by building an analytical model which enables an accurate quantitative analysis. Zhenyuan Ruan, Jason Cong |
ICCAD | 1 |
| 2019 | INSIDER: Designing In-Storage Computing System for Emerging High-Performance Drive
Zhenyuan Ruan, Jason Cong |
USENIX ATC | 1 |
| 2018 | ST-Accel: A High-Level Programming Platform for Streaming Applications on FPGAabstractIn recent years we have witnessed the emergence of the FPGA in many high-performance systems. This is due to FPGA's high reconfigurability and improved user-friendly programming environment. OpenCL, supported by major FPGA vendors, is a high-level programming platform that liberates hardware developers from having to deal with the complex and error-prone HDL development. While OpenCL exposes a GPU-like programming model, which is well-suited for compute-intensive tasks, in many state-of-art systems that deploy FPGA, we observe that the workloads are streaming-like, which is communication-intensive. This mismatch leads to low throughput and high end-to-end latency. In this paper, we propose ST-Accel, a new high-level programming platform for streaming applications on FPGA. It has the following advantages: (i) ST-Accel adopts the multiprocessing programming model to capture the inherent pipeline-level parallelism of streaming applications while reducing the end-to-end latency. (ii) A message-passing-based host/FPGA communication model is used to avoid the coherency issue of shared memory, thus enabling host/FPGA communication during kernel execution. (iii) ST-Accel provides a high-level abstraction for I/O devices to support direct I/O device access that eliminates the overhead of host CPU and reduces the I/O latency. (iv) ST-Accel enables the decoupled access/execute architecture to maximize the utilization of I/O devices. (v) The host/FPGA communication interface is redesigned to cater to the demands of both latency-critical and throughput-critical scenarios. The experimental results on the Amazon AWS cloud and local machine show that ST-Accel can achieve 1.6X-166X throughput and 1/3 latency for typical streaming workloads when compared to OpenCL. Zhenyuan Ruan, Bojie Li, Peipei Zhou 0001, Jason Cong |
FCCM | 1 |
| 2018 | Doppio: I/O-Aware Performance Analysis, Modeling and Optimization for In-memory Computing FrameworkabstractIn conventional Hadoop MapReduce applications, I/O used to play a heavy role in the overall system performance. More recently, a study from the Apache Spark community- state-of-the-art in-memory cluster computing framework- reports that I/O is no longer the bottleneck and has a marginal performance impact on applications like SQL processing. However, we observe that simply replacing HDDs with SSDs in a Spark cluster can have over 10x performance improvement for certain stages in large-scale production-quality genome processing. Therefore, one key question arises: How does I/O quanti- tatively impact the performance of today's big data applications developed using in-memory cluster computing frameworks like Apache Spark? In this paper we select an important yet complex application- the Spark-based Genome Analysis ToolKit (GATK4)-to guide our modeling. We first use different combinations of HDDs and SSDs to measure the I/O impact on GATK4 and change the CPU core number to discover the relation between computation and I/O access. By combining with Spark's underlying implementations, we further analyze the inherent cause of the above observations and build our model based on the analysis. Although we are building upon GATK4, our model maintains generality to other applications. Experimental results show that we can achieve a performance prediction error rate within 10% for typical Spark applications of both iterative and shuffle-heavy algorithms. Finally, we further extend our model to a broader area-that of optimal configuration selection in the public cloud. In Google Cloud, our model enables us to save 38% to 57% of cost for genome sequencing compared with its recommended default configurations. Currently, more and more companies are adopting cloud computing for specific workloads. Our proposed model can have a huge impact on their choices, while also enabling them to significantly reduce their costs. Peipei Zhou 0001, Zhenyuan Ruan, Zhenman Fang, Megan Shand, David Roazen, Jason Cong |
ISPASS | 2 |
| 2017 | Memory Efficient Loss Recovery for Hardware-based Transport in DatacenterabstractLimited by the small on-chip memory, hardware-based transport typically implements go-back-N loss recovery mechanism, which costs very few memory but is well-known to perform inferior even under small packet loss ratio. We present MELO, an efficient selective retransmission mechanism for hardware-based transport, which consumes only a constant small memory regardless of the number of concurrent connections. Specifically, MELO employs an architectural separation between data and meta data storage and uses a shared bits pool allocation mechanism to reduce meta data on-chip memory footprint. By only adding in average 23B extra on-chip states for each connection, MELO achieves up to 14.02x throughput while reduces 99% tail FCT by 3.11x compared with go-back-N under certain loss ratio. Yuanwei Lu, Guo Chen 0001, Zhenyuan Ruan, Wencong Xiao, Bojie Li, Jiansong Zhang 0001, Yongqiang Xiong, Peng Cheng 0005, Enhong Chen |
APNet | 3 |
| 2017 | KV-Direct: High-Performance In-Memory Key-Value Store with Programmable NICabstractPerformance of in-memory key-value store (KVS) continues to be of great importance as modern KVS goes beyond the traditional object-caching workload and becomes a key infrastructure to support distributed main-memory computation in data centers. Recent years have witnessed a rapid increase of network bandwidth in data centers, shifting the bottleneck of most KVS from the network to the CPU. RDMA-capable NIC partly alleviates the problem, but the primitives provided by RDMA abstraction are rather limited. Meanwhile, programmable NICs become available in data centers, enabling in-network processing. In this paper, we present KV-Direct, a high performance KVS that leverages programmable NIC to extend RDMA primitives and enable remote direct key-value access to the main host memory. Bojie Li, Zhenyuan Ruan, Wencong Xiao, Yuanwei Lu, Yongqiang Xiong, Andrew Putnam, Enhong Chen |
SOSP | 2 |