EDBT 2026 Demo / reviewers in the wild / expert
Chenxi Wang 0005
dblp:52/3121-5
· DBLP profile ↗
25ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-1451-3101ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 9 · 3 first-author · 7 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RaidenSwap: A Multi-Swap Remote System for Multi-core ApplicationsabstractKernel-based remote memory systems are gaining traction in datacenters due to their significant improvement in memory utilization and their ability to transparently provide applications with unlimited memory capacity. However, the high degree of parallelism in contemporary applications leads to a significant demand for remote access throughput, which mismatches with the state-of-the-art kernel swap path due to its inherently limited parallelism. As a result, it cannot scale up the multi-core applications. We dive into the implementation of the swap path and identify the root cause behind it - significant lock contentions and inefficient swap tasks offloading. Kefan Liu, Ke Liu 0004, Xu Zhang 0033, Ning Liu 0031, Sa Wang, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005 |
EuroSys | 11 |
| 2025 | XHarvest: Rethinking High-Performance and Cost-Efficient SSD Architecture with CXL-Driven HarvestingabstractThe occasional nature of I/O bursts in production clusters makes the substantial and expensive SSD internal hardware resources (e.g., computation and memory resources) always underutilized, resulting in cost inefficiency.Open-Channel SSD (OCSSD), as a pioneering solution, removes the SSD internal resources but rather leverages the host-side resources to serve I/O requests.Unfortunately, it faces adoption obstacles due to the heavy resource contention with user applications, hampered host-SSD collaboration, and proprietary firmware leakage risks.Tackling these challenges, we propose XHarvest, a new cost-efficient and high-performance SSD architecture, which harnesses compute express link (CXL) and trusted execution environment (TEE) to facilitate dynamic, efficient, and secure host resource harvesting.It reserves moderate SSD internal resources to isolate SSD internal tasks and applications under regular I/O loads while coping with occasional I/O bursts via dynamic host resource harvesting.To this end, XHarvest executes the firmware within the host-side TEE without disclosing sensitive Shushu Yi, Xianzhang Chen, Chenxi Wang 0005, Shengwen Liang, Zhe Wang 0017, Nong Xiao 0001, Qiao Li 0001, Mingzhe Zhang 0005, Jie Zhang 0048 |
ISCA | 5 |
| 2025 | QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code TranslationabstractThe rise of GPU-based high-performance computing (HPC) has driven the widespread adoption of parallel programming models such as CUDA. Yet, the inherent complexity of parallel programming creates a demand for the automated sequential-to-parallel approaches.
However, data scarcity poses a significant challenge for machine learning-based sequential-to-parallel code translation. Although recent back-translation methods show promise, they still fail to ensure functional equivalence in the translated code. In this paper, we propose \textbf{QiMeng-MuPa}, a novel \textbf{Mu}tual-Supervised Learning framework for Sequential-to-\textbf{Pa}rallel code translation, to address the functional equivalence issue. QiMeng-MuPa consists of two models, a Translator and a Tester. Through an iterative loop consisting of Co-verify and Co-evolve steps, the Translator and the Tester mutually generate data for each other and improve collectively. The Tester generates unit tests to verify and filter functionally equivalent translated code, thereby evolving the Translator, while the Translator generates translated code as augmented input to evolve the Tester. Experimental results demonstrate that QiMeng-MuPa significantly enhances the performance of the base models: when applied to Qwen2.5-Coder, it not only improves Pass@1 by up to 28.91\% and boosts Tester performance by 68.90\%, but also outperforms the previous state-of-the-art method CodeRosetta by 1.56 and 6.92 in BLEU and CodeBLEU scores, while achieving performance comparable to DeepSeek-R1 and GPT-4.1. Our code is available at \url{https://github.com/kcxain/mupa}. Changxin Ke, Rui Zhang 0040, Guangli Li, Yuanbo Wen 0001, Shuoming Zhang, Ruiyuan Xu, Jiaming Guo, Chenxi Wang 0005, Ling Li 0001, Qi Guo 0001, Yunji Chen |
NeurIPS | 11 |
| 2025 | Beehive: A Scalable Disaggregated Memory Runtime Exploiting Asynchrony of Multithreaded Programs
Quanxi Li, Ying Liu 0055, Yanwen Xia, Jie Zhang 0048, Mosong Zhou, Xiaobing Feng 0002, Huimin Cui, Yizhou Shan, Chenxi Wang 0005 |
NSDI | 11 |
| 2025 | Orthrus: Efficient and Timely Detection of Silent User Data Corruption in the Cloud with Resource-Adaptive Computation ValidationabstractEven with substantial endeavors to test and validate processors, computational errors may still arise post-installation. One particular category of CPU errors transpires discreetly, without crashing applications or triggering hardware warnings. These elusive errors pose a significant threat by undermining user data, and their detection is challenging. This paper introduces Orthrus, a solution for the timely detection of silent user data corruption caused by post-installation CPU errors. Orthrus safeguards user data in cloud applications by providing simple annotations and compiler support for users to identify data operators and validating these operators asynchronously across cores while maintaining a low overhead (2%–6%), making it practical for production deployment. Our evaluation, using carefully injected errors, demonstrates that Orthrus can detect 87% of data corruptions with just a single core dedicated to validation, increasing to 91% and 96% when two and four cores are used, respectively. Chenxiao Liu, Zhenting Zhu, Quanxi Li, Yanwen Xia, Yifan Qiao 0002, Xiangyun Deng, Youyou Lu, Tao Xie 0001, Huimin Cui, Zidong Du, Guoqing Harry Xu, Chenxi Wang 0005 |
SOSP | 12 |
| 2025 | DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack Communication
Xu Zhang 0033, Ke Liu 0004, Yuan Hui 0001, Yisong Chang, Yizhou Shan, Ke Zhang 0017, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005 |
USENIX ATC | 11 |
| 2025 | ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream WorkloadsabstractTransformer-based large language model (LLM) inference serving is now the backbone of many cloud services. LLM inference consists of a prefill phase and a decode phase. However, existing LLM deployment practices often overlook the distinct characteristics of these phases, leading to significant interference. To mitigate interference, our insight is to carefully schedule and group inference requests based on their characteristics. We realize this idea in ShuffleInfer through three pillars. First, it partitions prompts into fixed-size chunks so that the accelerator always runs close to its computation-saturated limit. Second, it disaggregates prefill and decode instances so each can run independently. Finally, it uses a smart two-level scheduling algorithm augmented with predicted resource usage to avoid decode scheduling hotspots. Results show that ShuffleInfer improves time-to-first-token (TTFT), job completion time (JCT), and inference efficiency in terms of performance per dollar by a large margin, e.g., it uses 38% less resources all the while lowering average TTFT and average JCT by 97% and 47%, respectively. Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Chenxi Wang 0005, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan |
ACM Trans. Archit. Code Optim. | 5 |
| 2024 | Enabling Large Dynamic Neural Network Training with Learning-based Memory ManagementabstractDynamic neural network (DyNN) enables high computational efficiency and strong representation capability. However, training DyNN can face a memory capacity problem because of increasing model size or limited GPU memory capacity. Managing tensors to save GPU memory is challenging, because of the dynamic structure of DyNN. We present DyNN-Offload, a memory management system to train DyNN. DyNN-Offload uses a learned approach (using a neural network called the pilot model) to increase predictability of tensor accesses to facilitate memory management. The key of DyNN-Offload is to enable fast inference of the pilot model in order to reduce its performance overhead, while providing high inference (or prediction) accuracy. DyNNOffload reduces input feature space and model complexity of the pilot model based on a new representation of DyNN; DyNNOffload converts the hard problem of making prediction for individual operators into a simpler problem of making prediction for a group of operators in DyNN. DyNN-Offload enables 8 × larger DyNN training on a single GPU compared with using PyTorch alone (unprecedented with any existing solution). Evaluating with AlphaFold (a production-level, large-scale DyNN), we show that DyNN-Offload outperforms unified virtual memory (UVM) and dynamic tensor rematerialization (DTR), the most advanced solutions to save GPU memory for DyNN, by 3 × and 2.1 × respectively in terms of maximum batch size. Jie Ren 0015, Dong Xu 0024, Shuangyan Yang, Christian Navasca, Chenxi Wang 0005, Guoqing Harry Xu, Dong Li 0001 |
HPCA | 7 |
| 2024 | A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
Chenxi Wang 0005, Yifan Qiao 0002, Zhe Wang 0017, Chenggang Wu 0002, Youyou Lu, Xiaobing Feng 0002, Huimin Cui, Shan Lu 0001, Guoqing Harry Xu |
OSDI | 3 |
| 2024 | ScalaCache: Scalable User-Space Page Cache Management with Software-Hardware Coordination
Yuda An, Chenxi Wang 0005, Qiao Li 0001, Chuanning Cheng, Jie Zhang 0048 |
USENIX ATC | 4 |
| 2024 | ScalaAFA: Constructing User-Space All-Flash Array Engine with Holistic Designs
Shushu Yi, Xiurui Pan, Qiao Li 0001, Chenxi Wang 0005, Bo Mao 0003, Myoungsoo Jung, Jie Zhang 0048 |
USENIX ATC | 5 |
| 2023 | Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU CoresabstractSIMD extensions are widely adopted in multi-core processors to exploit data-level parallelism. However, when co-running workloads on different cores, compute-intensive workloads cannot take advantage of the underutilized SIMD lanes allocated to memoryintensive workloads, reducing the overall performance. This paper proposes Occamy, a SIMD co-processor that can be shared by multiple CPU cores, so that their co-running workloads can spatially share its SIMD lanes. The key idea is to enable elastic spatial sharing by dynamically partitioning all the SIMD lanes across different workloads based on their phase behaviors, so that each workload may execute in variable-length SIMD mode. We also introduce an Occamy compiler to support such variable-length vectorization by analyzing such phase behaviors and generating the vectorized code that works with varying vector lengths. We demonstrate that Occamy can improve SIMD utilization, and consequently, performance over three representative SIMD architectures, with negligible chip area cost. Zhongcheng Zhang, Yan Ou, Ying Liu 0055, Chenxi Wang 0005, Yongbin Zhou, Yucheng Ouyang, Jiahao Shan, Ying Wang 0001, Jingling Xue, Huimin Cui, Xiaobing Feng 0002 |
ASPLOS (3) | 4 |
| 2023 | Skadi: Building a Distributed Runtime for Data Systems in Disaggregated Data CentersabstractData-intensive systems are the backbone of today's computing and are responsible for shaping data centers. Over the years, cloud providers have relied on three principles to maintain cost-effective data systems: use disaggregation to decouple scaling, use domain-specific computing to battle waning laws, and use serverless to lower costs. Although they work well individually, they fail to work in harmony: an issue amplified by emerging data system workloads. Cunchen Hu, Chenxi Wang 0005, Sa Wang, Ninghui Sun, Yungang Bao, Jieru Zhao, Sanidhya Kashyap, Pengfei Zuo, Xusheng Chen, Liangliang Xu, Yizhou Shan |
HotOS | 2 |
| 2023 | Hermit: Low-Latency, High-Throughput, and Transparent Remote Memory via Feedback-Directed Asynchrony
Yifan Qiao 0002, Chenxi Wang 0005, Zhenyuan Ruan, Adam Belay, Qingda Lu, Yiying Zhang 0005, Miryung Kim, Guoqing Harry Xu |
NSDI | 2 |
| 2023 | Canvas: Isolated and Adaptive Swapping for Multi-Applications on Remote Memory
Chenxi Wang 0005, Yifan Qiao 0002, Ravi Netravali, Miryung Kim, Guoqing Harry Xu |
NSDI | 1 |
| 2023 | Reinvent Cloud Software Stacks for Resource Disaggregation
Chenxi Wang 0005, Yi-Zhou Shan, Pengfei Zuo, Huimin Cui |
J. Comput. Sci. Technol. | 1 |
| 2023 | SpecBox: A Label-Based Transparent Speculation Scheme Against Transient Execution AttacksabstractSpeculative execution techniques have been a cornerstone of modern processors to improve instruction-level parallelism. However, recent studies showed that this kind of techniques could be exploited by attackers to leak secret data via transient execution attacks, such as Spectre. Many defenses are proposed to address this problem, but they all face various challenges: (1) Tracking data flow in the instruction pipeline could comprehensively address this problem, but it could cause pipeline stalls and incur high performance overhead; (2) Making side effect of speculative execution imperceptible to attackers, but it often needs additional storage components and complicated data movement operations. In this article, we propose alabel-based transparent speculationscheme calledSpecBox. It dynamically partitions the cache system to isolate speculative data and non-speculative data, which can prevent transient execution from being observed by subsequent execution. Moreover, it uses thread ownership semaphores to prevent speculative data from being accessed across cores. In addition,SpecBoxalso enhances the auxiliary components in the cache system against transient execution attacks, such as hardware prefetcher. Our security analysis shows thatSpecBoxis secure and the performance evaluation shows that the performance overhead on SPEC CPU 2006 and PARSEC-3.0 benchmarks is small. Bowen Tang 0001, Chenggang Wu 0002, Zhe Wang 0017, Lichen Jia, Pen-Chung Yew, Yueqiang Cheng, Yinqian Zhang, Chenxi Wang 0005, Guoqing Harry Xu |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2022 | MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime
Chenxi Wang 0005, Yifan Qiao 0002, Jon Eyolfson, Christian Navasca, Shan Lu 0001, Guoqing Harry Xu |
OSDI | 1 |
| 2022 | Mako: a low-pause, high-throughput evacuating collector for memory-disaggregated datacentersabstractResource disaggregation has gained much traction as an emerging datacenter architecture, as it improves resource utilization and simplifies hardware adoption. Under resource disaggregation, different types of resources (memory, CPUs, etc.) are disaggregated into dedicated servers connected by high-speed network fabrics. Memory disaggregation brings efficiency challenges to concurrent garbage collection (GC), which is widely used for latency-sensitive cloud applications, because GC and mutator threads simultaneously run and constantly compete for memory and swap resources. Chenxi Wang 0005, Yifan Qiao 0002, Michael D. Bond, Steve Blackburn, Miryung Kim, Guoqing Harry Xu |
PLDI | 3 |
| 2021 | Unified Holistic Memory Management Supporting Multiple Big Data Processing Frameworks over Hybrid MemoriesabstractTo process real-world datasets, modern data-parallel systems often require extremely large amounts of memory, which are both costly and energy inefficient. Emerging non-volatile memory (NVM) technologies offer high capacity compared to DRAM and low energy compared to SSDs. Hence, NVMs have the potential to fundamentally change the dichotomy between DRAM and durable storage in Big Data processing. However, most Big Data applications are written in managed languages and executed on top of a managed runtime that already performs various dimensions of memory management. Supporting hybrid physical memories adds a new dimension, creating unique challenges in data replacement. This article proposes Panthera, a semantics-aware, fully automated memory management technique for Big Data processing over hybrid memories. Panthera analyzes user programs on a Big Data system to infer their coarse-grained access patterns, which are then passed to the Panthera runtime for efficient data placement and migration. For Big Data applications, the coarse-grained data division information is accurate enough to guide the GC for data layout, which hardly incurs overhead in data monitoring and moving. We implemented Panthera in OpenJDK and Apache Spark. Based on Big Data applications’ memory access pattern, we also implemented a new profiling-guided optimization strategy, which is transparent to applications. With this optimization, our extensive evaluation demonstrates that Panthera reduces energy by 32–53% at less than 1% time overhead on average. To show Panthera’s applicability, we extend it to QuickCached, a pure Java implementation of Memcached. Our evaluation results show that Panthera reduces energy by 28.7% at 5.2% time overhead on average. Chenxi Wang 0005, John N. Zigman, Haris Volos 0001, Onur Mutlu, Xiaobing Feng 0002, Guoqing Harry Xu, Huimin Cui |
ACM Trans. Comput. Syst. | 3 |
| 2020 | Semeru: A Memory-Disaggregated Managed Runtime
Chenxi Wang 0005, Yuanqi Li, Zhenyuan Ruan, Khanh Nguyen 0001, Michael D. Bond, Ravi Netravali, Miryung Kim, Guoqing Harry Xu |
OSDI | 1 |
| 2020 | Systemizing Interprocedural Static Analysis of Large-scale Systems Code with GraspanabstractThere is more than a decade-long history of using static analysis to find bugs in systems such as Linux. Most of the existing static analyses developed for these systems are simple checkers that find bugs based on pattern matching. Despite the presence of many sophisticated interprocedural analyses, few of them have been employed to improve checkers for systems code due to their complex implementations and poor scalability. In this article, we revisit the scalability problem of interprocedural static analysis from a “Big Data” perspective. That is, we turn sophisticated code analysis into Big Data analytics and leverage novel data processing techniques to solve this traditional programming language problem. We propose Graspan , a disk-based parallel graph system that uses an edge-pair centric computation model to compute dynamic transitive closures on very large program graphs. We develop two backends for Graspan, namely, Graspan-C running on CPUs and Graspan-G on GPUs, and present their designs in the article. Graspan-C can analyze large-scale systems code on any commodity PC, while, if GPUs are available, Graspan-G can be readily used to achieve orders of magnitude speedup by harnessing a GPU’s massive parallelism. We have implemented fully context-sensitive pointer/alias and dataflow analyses on Graspan. An evaluation of these analyses on large codebases written in multiple languages such as Linux and Apache Hadoop demonstrates that their Graspan implementations are language-independent, scale to millions of lines of code, and are much simpler than their original implementations. Moreover, we show that these analyses can be used to uncover many real-world bugs in large-scale systems code. Zhiqiang Zuo 0002, Kai Wang 0029, Aftab Hussain 0001, Ardalan Amiri Sani, Yiyu Zhang, Shenming Lu, Wensheng Dou, Linzhang Wang, Xuandong Li, Chenxi Wang 0005, Guoqing Harry Xu |
ACM Trans. Comput. Syst. | 10 |
| 2019 | Panthera: holistic memory management for big data processing over hybrid memoriesabstractModern data-parallel systems such as Spark rely increasingly on in-memory computing that can significantly improve the efficiency of iterative algorithms. To process real-world datasets, modern data-parallel systems often require extremely large amounts of memory, which are both costly and energy-inefficient. Emerging non-volatile memory (NVM) technologies offers high capacity compared to DRAM and low energy compared to SSDs. Hence, NVMs have the potential to fundamentally change the dichotomy between DRAM and durable storage in Big Data processing. However, most Big Data applications are written in managed languages (e.g., Scala and Java) and executed on top of a managed runtime (e.g., the Java Virtual Machine) that already performs various dimensions of memory management. Supporting hybrid physical memories adds in a new dimension, creating unique challenges in data replacement and migration. Chenxi Wang 0005, Huimin Cui, John N. Zigman, Haris Volos 0001, Onur Mutlu, Xiaobing Feng 0002, Guoqing Harry Xu |
PLDI | 1 |
| 2018 | NVM Streaker: a fast and reconfigurable performance simulator for non-volatile memory-based memory architecture
Danqi Hu, Chenxi Wang 0005, Huimin Cui, Lei Wang 0004, Ying Liu 0055, Xiaobing Feng 0002 |
J. Supercomput. | 3 |
| 2016 | Efficient Management for Hybrid Memory in Managed Language Runtime
Chenxi Wang 0005, John N. Zigman, Yunquan Zhang, Xiaobing Feng 0002 |
NPC | 1 |