VLDB 2026 Research / reviewers in the wild / expert
Yiying Zhang 0005
dblp:64/3770-5
· DBLP profile ↗
33ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SymGPT: Auditing Smart Contracts via Combining Symbolic Execution with Large Language ModelsabstractThis paper introduces SymGPT , a tool that combines LLMs with symbolic execution to automatically verify smart contracts’ compliance with ERC rules. We begin by empirically analyzing 132 ERC rules from three major ERC standards, examining their content, security implications, and natural language descriptions. Based on this study, SymGPT instructs an LLM to translate ERC rules into a domain-specific language, synthesizes constraints from the translated rules to model potential rule violations, and performs symbolic execution for violation detection. Our evaluation shows that SymGPT identifies 5,783 ERC rule violations in 4,000 real-world contracts, including 1,375 violations with clear attack paths for financial theft. Furthermore, SymGPT outperforms six automated techniques and a security-expert auditing service, underscoring its superiority over current smart contract analysis methods. Shihao Xia, Tingting Yu 0001, Yiying Zhang 0005, Nobuko Yoshida, Linhai Song |
Proc. ACM Program. Lang. | 5 |
| 2026 | Introduction to the Invited Top Papers of USENIX ATC 2024
Saurabh Bagchi, Yiying Zhang 0005 |
ACM Trans. Comput. Syst. | 2 |
| 2025 | Preble: Efficient Distributed Prompt Scheduling for LLM ServingabstractPrompts to large language models (LLMs) have evolved beyond simple user questions.
For LLMs to solve complex problems, today’s practices are to include domain-specific
instructions, illustration of tool usages, and/or long context such as textbook chapters in
prompts. As such, many parts of prompts are repetitive across requests. Recent works
propose to cache and reuse KV state of prompts. However, they are all confined to a single-
GPU optimization, while production LLM serving systems are distributed by nature.
This paper proposes Preble, the first distributed LLM serving platform that targets and op-
timizes for prompt sharing. We designed a distributed scheduling system that co-optimizes
KV state reuse and computation load-balancing with a new scheduling algorithm and a
hierarchical scheduling mechanism. Our evaluation of Preble with real workloads and re-
quest arrival patterns on two open-source LLMs shows that Preble outperforms the SOTA
serving systems by 1.5× to 14.5× on average latency and 2× to 10× on p99 latency. Vikranth Srivatsa, Reyna Abhyankar, Yiying Zhang 0005 |
ICLR | 5 |
| 2025 | Cognify: Supercharging Gen-AI Workflows With Hierarchical AutotuningabstractToday's gen-AI workflows that involve multiple ML model calls, tool/API calls, data retrieval, or generic code execution are often tuned manually in an ad-hoc way that is both time-consuming and error-prone. In this paper, we propose a systematic approach for automatically tuning gen-AI workflows. Our key insight is that gen-AI workflows can benefit from structure, operator, and prompt changes, but unique properties of gen-AI workflows require new optimization techniques. We propose AdaSeek, an adaptive hierarchical search algorithm for autotuning gen-AI workflows. AdaSeek organizes workflow tuning methods into different layers based on the user-specified total search budget and distributes the budget across different layers based on the complexity of each layer. During its hierarchical search, AdaSeek redistributes the search budget from less useful to more promising tuning configurations based on workflow-level evaluation results. We implement AdaSeek in a workflow autotuning framework called Cognify and evaluate Cognify using six types of workflows such as RAG-based QA and text-to-SQL transformation. Overall, Cognify improves these workflows' generation quality by up to 2.8×, reduces execution monetary cost by up to 10×, and reduces end-to-end latency by 2.7×. Reyna Abhyankar, Vikranth Srivatsa, Yiying Zhang 0005 |
KDD (2) | 4 |
| 2025 | KDD 2025 Workshop on Inference Optimization for Generative AIabstractThe demand for efficient Large Language Model (LLM) inference has surged with the rising adoption of Generative AI (GenAI) applications, particularly in areas such as agents and retrieval-augmented generation. Efficient inference serves two crucial purposes: it enables the deployment of LLM-centered applications that address critical business needs, while also facilitating rapid experimentation for researchers to extract valuable insights and new understandings. However, despite the field's rapid advancement and interdisciplinary nature, there remains a limited exchange of ideas and methodologies between production-facing practitioners and researchers seeking to experiment with new GenAI concepts quickly. To bridge this gap, we are introducing the first KDD workshop on Inference Optimization for Generative AI. Our goal is to create a collaborative platform where researchers and practitioners working across various use cases and stacks of efficient inference can come together to exchange research ideas, establish connections between different disciplines, and identify challenges and research questions that will shape future work. Youngsuk Park, Lin Lee Cheong, Yida Wang 0003, Yiying Zhang 0005, George Karypis, Sherry Marcus |
KDD (2) | 5 |
| 2025 | Enabling Portable and High-Performance SmartNIC Programs with Alkali
Mihir Shah, Yiying Zhang 0005, Daehyeok Kim, Aditya Akella |
NSDI | 5 |
| 2025 | How to Save My Gas Fees: Understanding and Detecting Real-World Gas Issues in Solidity ProgramsabstractThe execution of smart contracts on Ethereum, a public blockchain system, incurs a fee called gas fee for its computation and data storage. When programmers develop smart contracts (e.g., in the Solidity programming language), they could unknowingly write code snippets that unnecessarily cause more gas fees. These issues, or what we call gas wastes, can lead to significant monetary losses for users. This paper takes the initiative in helping Ethereum users reduce their gas fees in two key steps. First, we conduct an empirical study on gas wastes in open-source Solidity programs and Ethereum transaction traces. Second, to validate our study findings, we develop a static tool called PeCatch to effectively detect gas wastes in Solidity programs, and manually examine the Solidity compiler’s code to pinpoint implementation errors causing gas wastes. Overall, we make 11 insights and four suggestions, which can foster future tool development and programmer awareness, and fixing our detected bugs can save $0.76 million in gas fees daily. Shihao Xia, Boqin Qin, Nobuko Yoshida, Tingting Yu 0001, Yiying Zhang 0005, Linhai Song |
IEEE Trans. Software Eng. | 6 |
| 2024 | SuperNIC: An FPGA-Based, Cloud-Oriented SmartNICabstractWith CPU scaling slowing down in today's data centers, more functionalities are being offloaded from the CPU to auxiliary devices. One such device is the SmartNIC, which is being increasingly adopted in data centers. In today's cloud environment, VMs on the same server can each have their own network computation (or network tasks) or workflows of network tasks to offload to a SmartNIC. These network tasks can be dynamically added/removed as VMs come and go and can be shared across VMs. Such dynamism demands that a SmartNIC not only schedules and processes packets but also manages and executes offloaded network tasks for different users. Although software solutions like an OS exist for managing software-based network tasks, such software-based SmartNICs cannot keep up with the quickly increasing data-center network speed. This paper proposes a new SmartNIC platform called SuperNIC that allows multiple tenants to efficiently and safely offload FPGA-based network computation DAGs. For efficiency and scalability, our core idea is to group network tasks into virtual chains that are dynamically mapped to different forms of physical chains depending on load and FPGA space availability. We further propose techniques to automatically scale network task chains with different types of parallelism. Moreover, we propose a fair sharing mechanism that considers both fair space sharing and fair time sharing of different types of hardware resources. Our FPGA prototype of SuperNIC achieves high bandwidth and low latency performance whilst efficiently utilizing and fairly sharing resources. Will Lin, Yizhou Shan, Ryan Kosta, Arvind Krishnamurthy, Yiying Zhang 0005 |
FPGA | 5 |
| 2024 | InferCept: Efficient Intercept Support for Augmented Large Language Model InferenceabstractLarge language models are increasingly integrated with external environments, tools, and agents like ChatGPT plugins to extend their capability beyond language-centric tasks. However, today’s LLM inference systems are designed for standalone LLMs. They treat each external interaction as the end of LLM generation and form a new request when the interaction finishes, causing unnecessary recomputation of already computed contexts, which accounts for 37-40% of total model forwarding time. This paper presents InferCept, the first LLM inference framework targeting augmented LLMs and supporting the efficient interception of LLM generation. InferCept minimizes the GPU resource waste caused by LLM interceptions and dedicates saved memory for serving more requests.InferCept improves the overall serving throughput by 1.6x-2x and completes 2x more requests per second compared to the state-of-the-art LLM inference systems. Reyna Abhyankar, Vikranth Srivatsa, Yiying Zhang 0005 |
ICML | 5 |
| 2024 | DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
Yifan Qiao 0002, Shan Yu 0001, Yuanjiang Ni, Qingda Lu, Jiesheng Wu, Yiying Zhang 0005, Miryung Kim, Guoqing Harry Xu |
OSDI | 8 |
| 2024 | Understanding and Detecting Real-World Safety Issues in RustabstractRust is a relatively new programming language designed for systems software development. Its objective is to combine the safety guarantees typically associated with high-level languages with the performance efficiency often found in executable programs implemented in low-level languages. The core design of Rust is a set of strict safety rules enforced through compile-time checks. However, to support more low-level controls, Rust also allows programmers to bypass its compiler checks by writingunsafecode. As the adoption of Rust grows in the development of safety-critical software, it becomes increasingly important to understand what safety issues may elude Rust’s compiler checks and manifest in real Rust programs.In this paper, we conduct a comprehensive, empirical study of Rust safety issues by close, manual inspection of 70 memory bugs, 100 concurrency bugs, and 110 programming errors leading to unexpected execution panics from five open-source Rust projects, five widely-used Rust libraries, and two online security databases. Our study answers three important questions: what memory-safety issues real Rust programs have, what concurrency bugs Rust programmers make, and how unexpected panics in Rust programs are caused. Our study reveals interesting real-world Rust program behaviors and highlights new issues made by Rust programmers. Building upon the findings of our study, we design and implement five static detectors. After being applied to the studied Rust programs and another 12 selected Rust projects, our checkers pinpoint 96 previously unknown bugs and report a negligible number of false positives, confirming their effectiveness and the value of our empirical study. Boqin Qin, Hua Zhang 0001, Qiaoyan Wen, Linhai Song, Yiying Zhang 0005 |
IEEE Trans. Software Eng. | 7 |
| 2023 | Hermit: Low-Latency, High-Throughput, and Transparent Remote Memory via Feedback-Directed Asynchrony
Yifan Qiao 0002, Chenxi Wang 0005, Zhenyuan Ruan, Adam Belay, Qingda Lu, Yiying Zhang 0005, Miryung Kim, Guoqing Harry Xu |
NSDI | 6 |
| 2023 | Mira: A Program-Behavior-Guided Far Memory SystemabstractFar memory, where memory accesses are non-local, has become more popular in recent years as a solution to expand memory size and avoid memory stranding. Prior far memory systems have taken two approaches: transparently swap memory pages between local and far memory, and utilizing new programming models to explicitly move fine-grained data between local and far memory. The former requires no program changes but comes with performance penalty. The latter has potentially better performance but requires significant program changes. Yiying Zhang 0005 |
SOSP | 3 |
| 2022 | Clio: a hardware-software co-designed disaggregated memory systemabstractMemory disaggregation has attracted great attention recently because of its benefits in efficient memory utilization and ease of management. So far, memory disaggregation research has all taken one of two approaches: building/emulating memory nodes using regular servers or building them using raw memory devices with no processing power. The former incurs higher monetary cost and faces tail latency and scalability limitations, while the latter introduces performance, security, and management problems. Yizhou Shan, Xuhao Luo, Yutong Huang, Yiying Zhang 0005 |
ASPLOS | 5 |
| 2021 | The future of hardware development: a perspective from systems researchersabstractAt HotOS 2011, Mogul et al. published a paper calling for "Reconnecting Architecture and OS Research" [9]. Ten years have passed. We now live in a post-x86 world where servers are no longer homogeneous. Accelerators like GPU and FPGA are dominating important workloads like machine learning and search in data centers [3, 4]. Many data centers have launched large scales of customized ASICs (e.g., Google TPU [8], AWS Nitro [1]). At the same time, HBM and NVM are promised to change the landscape of memory and storage; various programmable networking devices like SmartNIC and programmable switches are making their ways into data centers. These exciting new hardware trends have driven systems researchers to innovate on software to fit new hardware technologies. What about the hardware? Yiying Zhang 0005 |
HotOS | 1 |
| 2021 | User-defined cloudabstractSince its creation, cloud computing has always taken a provider-dictated approach, where cloud providers define and manage the cloud to accommodate the user needs they deem important. We propose "User-Defined Cloud", or UDC, a new cloud scheme that allows users to define their own "clouds", by defining hardware resource needs, system software features, and security requirements of their applications, and to do so without the need to build or manage low-level systems. Yiying Zhang 0005, Ardalan Amiri Sani, Guoqing Harry Xu |
HotOS | 1 |
| 2020 | VRLifeTime - An IDE Tool to Avoid Concurrency and Memory Bugs in RustabstractAs a young programming language designed for systems software development, Rust aims to provide safety guarantees like high-level languages and performance efficiency like low-level languages. Lifetime is a core concept in Rust, and it is key to both safety checks and automated resource management conducted by the Rust compiler. However, Rust's lifetime rules are very complex. In reality, it is not uncommon that Rust programmers fail to infer the correct lifetime, causing severe concurrency and memory bugs. In this paper, we present VRLifeTime, an IDE tool that can visualize lifetime for Rust programs and help programmers avoid lifetime-related mistakes. Moreover, VRLifeTime can help detect some lifetime-related bugs (i.e., double locks) with detailed debugging information. A demo video is available at https://youtu.be/L5F_XCOrJTQ. Boqin Qin, Linhai Song, Yiying Zhang 0005 |
CCS | 5 |
| 2020 | Understanding memory and thread safety practices and issues in real-world Rust programsabstractRust is a young programming language designed for systems software development. It aims to provide safety guarantees like high-level languages and performance efficiency like low-level languages. The core design of Rust is a set of strict safety rules enforced by compile-time checking. To support more low-level controls, Rust allows programmers to bypass these compiler checks to write unsafe code. Boqin Qin, Zeming Yu, Linhai Song, Yiying Zhang 0005 |
PLDI | 5 |
| 2020 | Disaggregating Persistent Memory and Controlling Them Remotely: An Exploration of Passive Disaggregated Key-Value Stores
Shin-Yeh Tsai, Yizhou Shan, Yiying Zhang 0005 |
USENIX ATC | 3 |
| 2019 | Understanding Real-World Concurrency Bugs in GoabstractGo is a statically-typed programming language that aims to provide a simple, efficient, and safe way to build multi-threaded software. Since its creation in 2009, Go has matured and gained significant adoption in production and open-source software. Go advocates for the usage of message passing as the means of inter-thread communication and provides several new concurrency mechanisms and libraries to ease multi-threading programming. It is important to understand the implication of these new proposals and the comparison of message passing and shared memory synchronization in terms of program errors, or bugs. Unfortunately, as far as we know, there has been no study on Go's concurrency bugs. In this paper, we perform the first systematic study on concurrency bugs in real Go programs. We studied six popular Go software including Docker, Kubernetes, and gRPC. We analyzed 171 concurrency bugs in total, with more than half of them caused by non-traditional, Go-specific problems. Apart from root causes of these bugs, we also studied their fixes, performed experiments to reproduce them, and evaluated them with two publicly-available Go bug detectors. Overall, our study provides a better understanding on Go's concurrency models and can guide future researchers and practitioners in writing better, more reliable Go software and in developing debugging and diagnosis tools for Go. Tengfei Tu, Linhai Song, Yiying Zhang 0005 |
ASPLOS | 4 |
| 2019 | Storm: a fast transactional dataplane for remote data structuresabstractRDMA technology enables a host to access the memory of a remote host without involving the remote CPU, improving the performance of distributed in-memory storage systems. Previous studies argued that RDMA suffers from scalability issues, because the NIC's limited resources are unable to simultaneously cache the state of all the concurrent network streams. These concerns led to various software-based proposals to reduce the size of this state by trading off performance. Stanko Novakovic, Yizhou Shan, Aasheesh Kolli, Michael Cui, Yiying Zhang 0005, Haggai Eran, Boris Pismenny, Liran Liss, Michael Wei, Dan Tsafrir, Marcos K. Aguilera |
SYSTOR | 5 |
| 2019 | LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation
Yizhou Shan, Yutong Huang, Yiying Zhang 0005 |
USENIX ATC | 4 |
| 2019 | Pythia: Remote Oracles for the Masses
Shin-Yeh Tsai, Mathias Payer, Yiying Zhang 0005 |
USENIX Security Symposium | 3 |
| 2018 | LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation
Yizhou Shan, Yutong Huang, Yiying Zhang 0005 |
OSDI | 4 |
| 2017 | Disaggregated operating systemabstractRecently, there is an emerging trend to move towards a disaggregated hardware architecture that breaks monolithic servers into independent hardware components that are connected to a fast, scalable network [1, 2]. The disaggregated architecture offers several benefits over traditional monolithic server model, including better resource utilization, ease of hardware deployment, and support for heterogeneity. Our vision of the future disaggregated architecture is that each component will have its own controller to manage its hardware and can communicate with other components through a fast network. Yizhou Shan, Sumukh Hallymysore, Yutong Huang, Yiying Zhang 0005 |
SoCC | 5 |
| 2017 | Distributed shared persistent memoryabstractNext-generation non-volatile memories (NVMs) will provide byte addressability, persistence, high density, and DRAM-like performance. They have the potential to benefit many datacenter applications. However, most previous research on NVMs has focused on using them in a single machine environment. It is still unclear how to best utilize them in distributed, datacenter environments. Yizhou Shan, Shin-Yeh Tsai, Yiying Zhang 0005 |
SoCC | 3 |
| 2017 | LITE Kernel RDMA Support for Datacenter ApplicationsabstractRecently, there is an increasing interest in building data-center applications with RDMA because of its low-latency, high-throughput, and low-CPU-utilization benefits. However, RDMA is not readily suitable for datacenter applications. It lacks a flexible, high-level abstraction; its performance does not scale; and it does not provide resource sharing or flexible protection. Because of these issues, it is difficult to build RDMA-based applications and to exploit RDMA's performance benefits. Shin-Yeh Tsai, Yiying Zhang 0005 |
SOSP | 2 |
| 2015 | Removing the costs and retaining the benefits of flash-based SSD virtualization with FSDVabstractWe present the design, implementation, and evaluation of the File System De-Virtualizer (FSDV), a system that dynamically removes a layer of indirection common in modern storage stacks and decreases indirection space and performance costs. FSDV is a flexible, light-weight tool that de-virtualizes data by changing file system pointers to use device physical addresses. When FSDV is not running, the file system and the device both maintain their virtualization layers and perform normal I/O operations. We implement FSDV with ext3 and an emulated flash-based SSD. Our evaluation results show that FSDV can significantly reduce indirection mapping table space in a dynamic way while preserving the foreground I/O performance. We also demonstrate that FSDV only requires small changes to existing storage systems. Yiying Zhang 0005, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau |
MSST | 1 |
| 2013 | Getting real: lessons in transitioning research simulations into hardware systems
Mohit Saxena, Yiying Zhang 0005, Michael M. Swift, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau |
FAST | 2 |
| 2013 | Warming up storage-level caches with bonfire
Yiying Zhang 0005, Gokul Soundararajan, Mark W. Storer, Lakshmi N. Bairavasundaram, Sethuraman Subbiah, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau |
FAST | 1 |
| 2013 | Warped Mirrors for flashabstractFlash-based devices are cost-competitive to traditional hard disks in both personal and industrial environments and offer the potential for large performance gains. However, as flash-based devices have a high bit-error rate and a relatively short lifetime, reliability issues remain a major problem. One possible solution is redundancy; using techniques such as mirroring, data reliability and availability can be greatly enhanced. All standard RAID approaches assume that devices do not wear out, and hence distribute work equally among them; unfortunately, for flash, this approach is not appropriate as the life of flash cell depends on the number of times it is written and cleaned. Hence, identical write patterns to mirrored flash drives introduce a failure dependency in the storage system, increasing the probability of concurrent device failure and hence data loss. We propose Warped Mirrors as a solution to this endurance problem for mirrored flash devices. By carefully inducing a slight imbalance into write traffic across devices, we intentionally increase the workload of one device in the mirror pair, and thus increase the odds that it will fail first. Thus, with our approach, device failure independence is preserved. Our simulation results show that across both synthetic and traced workloads, little performance overhead is induced. Yiying Zhang 0005, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau |
MSST | 1 |
| 2012 | FlashTier: a lightweight, consistent and durable storage cacheabstractThe availability of high-speed solid-state storage has introduced a new tier into the storage hierarchy. Low-latency and high-IOPS solid-state drives (SSDs) cache data in front of high-capacity disks. However, most existing SSDs are designed to be a drop-in disk replacement, and hence are mismatched for use as a cache. Mohit Saxena, Michael M. Swift, Yiying Zhang 0005 |
EuroSys | 3 |
| 2012 | De-indirection for flash-based SSDs with nameless writes
Yiying Zhang 0005, Leo Prasath Arulraj, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau |
FAST | 1 |