EDBT 2026 Demo / reviewers in the wild / expert
Xingbo Wu
dblp:145/0651
· DBLP profile ↗
19ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0003-1649-2612ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Good things come in small packages: Should we build AI clusters with Lite-GPUs?abstractTo match the blooming demand of generative AI workloads, GPU designers have so far been trying to pack more and more compute and memory into single complex and expensive packages. However, there is growing uncertainty about the scalability of individual GPUs and thus AI clusters, as state-of-the-art GPUs are already displaying packaging, yield, and cooling limitations. We propose to rethink the design and scaling of AI clusters through efficiently-connected large clusters of Lite-GPUs, GPUs with single, small dies and a fraction of the capabilities of larger GPUs. We think recent advances in co-packaged optics can enable distributing AI workloads onto many Lite-GPUs through high bandwidth and efficient communication. In this paper, we present the key benefits of Lite-GPUs on manufacturing cost, blast radius, yield, and power efficiency; and discuss systems opportunities and challenges around resource, workload, memory, and network management. Burcu Canakci, Xingbo Wu, Nathanael Cheriere, Paolo Costa, Sergey Legtchenko, Dushyanth Narayanan, Antony I. T. Rowstron |
HotOS | 3 |
| 2025 | Storage Class Memory is Dead, All Hail Managed-Retention Memory: Rethinking Memory for the AI EraabstractAI clusters today are one of the major uses of High Bandwidth Memory (HBM). However, HBM is suboptimal for AI workloads for several reasons. Analysis shows HBM is overprovisioned on write performance, but underprovisioned on density and read bandwidth, and also has significant energy per bit overheads. It is also expensive, with lower yield than DRAM due to manufacturing complexity. We propose a new memory class: Managed-Retention Memory (MRM), which is more optimized to store key data structures for AI inference workloads. We believe that MRM may finally provide a path to viability for technologies that were originally proposed to support Storage Class Memory (SCM). These technologies traditionally offered long-term persistence (10+ years) but provided poor IO performance and/or endurance. MRM makes different trade-offs, and by understanding the workload IO patterns, MRM foregoes long-term data retention and write performance for better potential performance on the metrics important for these workloads. Sergey Legtchenko, Ioan A. Stefanovici, Richard Black, Antony I. T. Rowstron, Paolo Costa, Burcu Canakci, Dushyanth Narayanan, Xingbo Wu |
HotOS | 9 |
| 2025 | Disco: A Compact Index for LSM-treesabstractMany key-value stores and database systems use log-structured merge-trees (LSM-trees) as their storage engines because of their excellent write performance. However, the read performance of LSM-trees is suboptimal due to the overlapping sorted runs. Most existing efforts rely on filters to reduce unnecessary I/Os, but filters fundamentally do not help locate items and often become the bottleneck of the system. We identify that the lack of efficient index is the root cause of subpar read performance in LSM-trees. In this paper, we propose Disco: a compact index for LSM-trees. Disco indexes all the keys in an LSM-tree, so a query does not have to search every run of the LSM-tree. It records compact key representations to minimize the number of key comparisons so as to minimize cache misses and I/Os for both point and range queries. Disco guarantees that both point queries and seeks issue at most one I/O to the underlying runs, achieving an I/O efficiency close to a B + -tree. Disco improves upon REMIX's pioneering multi-run index design with additional compact key representations to help improve read performance. The representations are compact so the cost of persisting Disco to disk is small. Moreover, while a traditional LSM-tree has to choose a more aggressive compaction policy that slows down write performance to have better read performance, a Disco-indexed LSM-tree can employ a write-efficient policy and still have good read performance. Experimental results show that Disco can save I/Os and improve point and range query performance by up to 220% over RocksDB while maintaining efficient writes. Wenshao Zhong, Chen Chen 0124, Xingbo Wu, Jakob Eriksson |
Proc. ACM Manag. Data | 3 |
| 2025 | Holographic Storage for the Cloud: advances and challengesabstractHolographic Storage is an old idea that has always promised high density and fast random access, but has never been commercially competitive with Hard Disk Drives (HDDs) and Solid State Devices (SSDs). In Project HSD at Microsoft Research we asked the question: “Does holographic storage finally make sense for cloud storage?” This article describes our journey toward answering this question. We achieved 1.8× higher density than the previous state-of-the-art, using commodity components available today and leveraging machine learning to compensate for the noise and distortions introduced by commodity components. This uncovered two new challenges which are the focus of this article: achieving high end-to-end energy efficiency without sacrificing capacity, and spatial multiplexing without mechanical movement. Improving end-to-end energy efficiency requires joint optimization across low-level media parameters and higher-level system parameters that govern background maintenance operations such as read refresh and garbage collection. We developed new physics models of the media; analytic and simulation models of the media access and background media maintenance; and workload-driven optimization to find optimal parameter combinations. These techniques resulted in a 14× improvement over the previous approach for typical workloads without sacrificing capacity. We also designed the first scalable and mechanical movement free spatial multiplexing system for holographic storage. Despite these advances, we conclude that currently, holographic storage is still far from the combination of density, capacity scaling, and energy efficiency needed to compete with the incumbent technologies. We need fundamental advances in the physical media that improve energy efficiency by another 1–2 orders of magnitude without reducing data density. Further advances in optics are also required to achieve spatial multiplexing that is simultaneously scalable, low-loss, and high-density. Nathanael Cheriere, Jiaqi Chu, Grace Brennan, Pashmina Cameron, Pedro Da Costa, Jannes Gladrow, Guilherme Ilunga, Douglas J. Kelly, Joowon Lim, Giorgio Maltese, Tony Mason, Greg O'Shea, Soujanya Ponnapalli, Michael Rudow, Alan Sanders, Theano Stavrinos, Xingbo Wu, Mengyang Yang, Dushyanth Narayanan, Benn C. Thomsen, Antony I. T. Rowstron |
ACM Trans. Storage | 18 |
| 2024 | Fast Abort-Freedom for Deterministic TransactionsabstractThe efficiency of concurrency control protocols plays a crucial role in transaction processing systems. However, when it comes to deterministic transactions (i.e., transactions with known read/write key sets), existing concurrency control protocols are not optimized to make the most of the determinism. They either force transactions to be aborted and retried, which negatively affects system throughput, or use a centralized scheduler to organize transactions in a way that avoids aborts, but with limited system scalability.In this paper, we present DecentSched, a highly efficient decentralized concurrency control protocol for deterministic transactions. DecentSched employs fine-grained queuing and a decentralized scheduling algorithm to enable serializable concurrent transaction execution with a high degree of parallelism. Extensive evaluation results show that DecentSched can outperform state-of-the-art concurrency control protocols in representative benchmarks. Chen Chen 0124, Xingbo Wu, Wenshao Zhong, Jakob Eriksson |
IPDPS | 2 |
| 2022 | Building an efficient key-value store in a flexible address spaceabstractData management applications store their data using structured files in which data are usually sorted to serve indexing and queries. However, in-place insertions and removals of data are not naturally supported in a file's address space. To avoid repeatedly rewriting existing data in a sorted file to admit changes in place, applications usually employ extra layers of indirections, such as mapping tables and logs, to admit changes out of place. However, this approach leads to increased access cost and excessive complexity. Chen Chen 0124, Wenshao Zhong, Xingbo Wu |
EuroSys | 3 |
| 2021 | REMIX: Efficient Range Query for LSM-trees
Wenshao Zhong, Xingbo Wu, Song Jiang 0001 |
FAST | 3 |
| 2021 | WipDB: A Write-in-place Key-value Store that Mimics Bucket SortabstractKey-value (KV) stores have become a major storage infrastructure on which databases, file systems, and other data management systems are built. To support efficient indexing and range search, the key-value items must be sorted. However, this sorting process can be excessively expensive. In the KV systems adopting the popular Log-Structured Merge Tree (LSM) structure or its variants, the write volume can be amplified by tens of times due to its repeated internal merge-sorting operation.In this paper we propose a KV store design that leverages relatively stable key distributions to bound the write amplification by a number as low as 4.15 in practice. The key idea is, instead of incrementally sorting KV items in the LSM's hierarchical structure, it writes KV items right in place in an approximately sorted list, much like a bucket sort algorithm does. The design also makes it possible to keep most internal data reorganization operations off the critical path of read service. The so-called Write-in-place (Wip) scheme has been implemented with its source code publicly available. Experiment results show that WipDB improves write throughput by 3 to 8× (to around 1Mops/s on one Intel PCIe SSD) over state-of-the-art KV stores. Xingsheng Zhao, Song Jiang 0001, Xingbo Wu |
ICDE | 3 |
| 2019 | Wormhole: A Fast Ordered Index for In-memory Data ManagementabstractIn-memory data management systems, such as key-value stores, have become an essential infrastructure in today's big-data processing and cloud computing. They rely on efficient index structures to access data. While unordered indexes, such as hash tables, can perform point search with O(1) time, they cannot be used in many scenarios where range queries must be supported. Many ordered indexes, such as B+ tree and skip list, have a O(log N) lookup cost, where N is number of keys in an index. For an ordered index hosting billions of keys, it may take more than 30 key-comparisons in a lookup, which is an order of magnitude more expensive than that on a hash table. With availability of large memory and fast network in today's data centers, this O(log N) time is taking a heavy toll on applications that rely on ordered indexes. Xingbo Wu, Fan Ni, Song Jiang 0001 |
EuroSys | 1 |
| 2019 | SDC: a software defined cache for efficient data indexingabstractCPU cache has been used to bridge the processor-memory performance gap to enable high-performance computing. As the cache is of limited capacity, for its maximum efficacy it should (1) avoid caching data that are less likely to be accessed and (2) identify and cache data that would otherwise cost a program multiple memory accesses to reach. Unfortunately, existing cache architectures are inadequate on these two efforts. First, to cost-effectively exploit the spatial locality, they adopt a relatively large and fixed-size cache line as the caching unit. Thus, much of the space in a cache line can be wasted when the data locality is weak. Second, for easy use, the cache is designed to be transparent to programs, which hinders programs from fully exploiting its performance potentials. Fan Ni, Song Jiang 0001, Hong Jiang 0001, Jian Huang 0006, Xingbo Wu |
ICS | 5 |
| 2018 | ThinDedup: An I/O Deduplication Scheme that Minimizes Efficiency Loss due to Metadata WritesabstractI/O deduplication is an important technique for saving I/O bandwidth and storage space for storage systems. However, it requires a new level of address mapping, and consequently needs to maintain corresponding metadata. To meet requirements on data persistency and consistency, the metadata writing is likely to make deduplication operations much fatter, in terms of amount of additional writes on the critical I/O path, than one might expect. In this paper we propose to compress the data and insert metadata into data blocks to reduce metadata writes. Assuming that performance-critical data are usually compressible, we can mostly remove separate writes of metadata out of the critical path of servicing users' requests, and make I/O deduplication much thinner. Accordingly we name the scheme ThinDedup. In addition to metadata insertion, ThinDedup also uses persistency of data fingerprints to evade enforcement of write order between data and metadata. We have implemented ThinDedup in the Linux kernel as a device mapper target to provide block-level deduplication. Experimental results show, compared to existing deduplication schemes, ThinDedup achieves (much) higher (up to 3X) I/O throughput and lower latency (reduced by up to 88%) without compromising data persistency. Fan Ni, Xingbo Wu, Song Jiang 0001 |
IPCCC | 2 |
| 2018 | WOJ: Enabling Write-Once Full-data Journaling in SSDs by using weak-hashing-based deduplication
Fan Ni, Xingbo Wu, Lei Wang 0126, Song Jiang 0001 |
Perform. Evaluation | 2 |
| 2017 | Search lookaside buffer: efficient caching for index data structuresabstractWith the ever increasing DRAM capacity in commodity computers, applications tend to store large amount of data in main memory for fast access. Accordingly, efficient traversal of index structures to locate requested data becomes crucial to their performance. The index data structures grow so large that only a fraction of them can be cached in the CPU cache. The CPU cache can leverage access locality to keep the most frequently used part of an index in it for fast access. However, the traversal on the index to a target data during a search for a data item can result in significant false temporal and spatial localities, which make CPU cache space substantially underutilized. In this paper we show that even for highly skewed accesses the index traversal incurs excessive cache misses leading to suboptimal data access performance. To address the issue, we introduce Search Lookaside Buffer (SLB) to selectively cache only the search results, instead of the index itself. SLB can be easily integrated with any index data structure to increase utilization of the limited CPU cache resource and improve throughput of search requests on a large data set. We integrate SLB with various index data structures and applications. Experiments show that SLB can improve throughput of the index data structures by up to an order of magnitude. Experiments with real-world key-value traces also show up to 73% throughput improvement on a hash table. Xingbo Wu, Fan Ni, Song Jiang 0001 |
SoCC | 1 |
| 2017 | Freewrite: creating (almost) zero-cost writes to SSD in applicationsabstractWhile flash-based SSDs have much higher access speed than hard disks, they have an Achilles heel, which is the service of write requests. Not only is writing slower than reading, but also it can incur expensive garbage collection operations and reduce SSDs' lifetime. The deduplication technique can help to avoid writing data objects whose contents have been on the disk. A typical object is the disk block, for which a block-level deduplication scheme can help identify duplicate ones and avoid their writing. For the technique to be effective, data written to the disk must not only be the same as those currently on the disk but also be block-aligned. Chunyi Liu, Fan Ni, Xingbo Wu, Xiao Zhang 0014, Song Jiang 0001 |
SYSTOR | 3 |
| 2016 | zExpander: a key-value cache with both high performance and fewer missesabstractWhile key-value (KV) cache, such as memcached, dedicates a large volume of expensive memory to holding performance-critical data, it is important to improve memory efficiency, or to reduce cache miss ratio without adding more memory. As we find that optimizing replacement algorithms is of limited effect for this purpose, a promising approach is to use a compact data organization and data compression to increase effective cache size. However, this approach has the risk of degrading the cache's performance due to additional computation cost. A common perception is that a high-performance KV cache is not compatible with use of data compacting techniques. Xingbo Wu, Li Zhang 0002, Yandong Wang 0001, Yufei Ren, Michel Hack, Song Jiang 0001 |
EuroSys | 1 |
| 2016 | FlexPoll: adaptive event polling for network-intensive applications
Xingbo Wu, Xiang Long, Lei Wang 0126 |
Frontiers Comput. Sci. | 1 |
| 2015 | Selfie: co-locating metadata and data to enable fast virtual block devicesabstractVirtual block devices are widely used to provide block interface to virtual machines (VMs). A virtual block device manages an indirection mapping from the virtual address space presented to a VM, to a storage image hosted on file system or storage volume. This indirection is recorded as metadata on the image, also known as a lookup table, which needs to be immediately updated upon each space allocation on the image for data safety (also known as image growth). This growth is common as VM templates for large-scale deployments and snapshots for fast migration of VMs are heavily used. Though each table update involves only a few bytes of data, it demands a random write of an entire block. Furthermore, data consistency demands correct order of metadata and data writes be enforced, usually by inserting the FLUSH command between them. These metadata operations compromise virtual device's efficiency. Xingbo Wu, Zili Shao, Song Jiang 0001 |
SYSTOR | 1 |
| 2015 | LSM-trie: An LSM-tree-based Ultra-Large Key-Value Store for Small Data Items
Xingbo Wu, Yuehai Xu, Zili Shao, Song Jiang 0001 |
USENIX ATC | 1 |
| 2013 | Optimizing Event Polling for Network-Intensive Applications: A Case Study on RedisabstractIn today's data centers supporting Internet-scale computing and I/O services, increasingly more network-intensive applications are deployed on the network as a service. To this end, it is critical for the applications to quickly retrieve requests from the network and send their responses to the network. To facilitate this network function, operating system usually provides an event notification mechanism so that the applications (or the library) know if the network is ready to supply data for them to read or to receive data for them to write. As a widely used and representative notification mechanism, epoll in Linux provides a scalable and high-performance implementation by allowing applications to specifically indicate which connections and what events on them need to be watched. As epoll has been used in some major systems, including KV systems, such as Redis and Memcached, and web server systems such as NGINX, we have identified a substantial performance issue in its use. For the sake of efficiency, applications usually use epoll's system calls to inform the kernel exactly of what events they are interested in and always keep the information up-to-date. However, in a system with demanding network traffic, such a rigid maintenance of the information is not necessary and the excess number of system calls for this purpose can substantially degrade the system's performance. In this paper, we use Redis as an example to explore the issue. We propose a strategy of informing the kernel of the interest events in a manner adaptive to the current network load, so that the epoll system calls can be reduced and the events can be efficiently delivered. We have implemented the strategy, named as FlexPoll, in Redis without modifying any kernel code. Our evaluation on Redis shows that the query throughput can be improved by up to 46.9% on micro benchmarks, and even up to 67.8% on workloads emulating real-world access patterns. FlexPoll can be extended to other applications and event libraries built on the epoll mechanism in a straightforward manner. Xingbo Wu, Xiang Long, Lei Wang 0126 |
ICPADS | 1 |