EDBT 2026 Demo / reviewers in the wild / expert
Jun Yuan 0006
dblp:98/4381-6
· DBLP profile ↗
14ranked-venue papers
4as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Storage systems · 98% Performance modeling and evaluation · 2% | |
| Databases, data mining, and information retrieval
1 paper |
Transaction processing and concurrency control · 44% Indexing and storage engines · 44% Data stream processing · 13% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
file systems |
3.4 | 10 | 2022 | BetrFS: a compleat file system for commodity SSDs · EuroSys 2022 Copy-on-Abundant-Write for Nimble File System Clones · ACM Trans. Storage 2021 How to Copy Files · FAST 2020 |
Storage systems › file systems
write-optimized file system |
1.2 | 5 | 2017 | Writes Wrought Right, and Other Adventures in File System Optimization · ACM Trans. Storage 2017 Optimizing Every Operation in a Write-optimized File System · USENIX ATC 2016 Optimizing Every Operation in a Write-optimized File System · FAST 2016 |
Indexing and storage engines
key-value store |
0.9 | 1 | 2025 | Focus! Fast On-disk Concurrency-control Using Sketches · Proc. ACM Manag. Data 2025 |
Transaction processing and concurrency control › concurrency control
timestamp-based concurrency control |
0.9 | 1 | 2025 | Focus! Fast On-disk Concurrency-control Using Sketches · Proc. ACM Manag. Data 2025 |
Storage systems
flash and SSD |
0.6 | 1 | 2022 | BetrFS: a compleat file system for commodity SSDs · EuroSys 2022 |
Storage systems › file systems
flash file system |
0.6 | 1 | 2022 | BetrFS: a compleat file system for commodity SSDs · EuroSys 2022 |
Storage systems › file systems
copy-on-write |
0.5 | 1 | 2021 | Copy-on-Abundant-Write for Nimble File System Clones · ACM Trans. Storage 2021 |
Storage systems
indexing |
0.3 | 1 | 2018 | The Full Path to Full-Path Indexing · FAST 2018 |
Storage systems › file systems › file system performance
file system aging |
0.3 | 1 | 2017 | File Systems Fated for Senescence? Nonsense, Says Science! · FAST 2017 |
Data stream processing › sketch
sketch-based estimation |
0.3 | 1 | 2025 | Focus! Fast On-disk Concurrency-control Using Sketches · Proc. ACM Manag. Data 2025 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2022 | BetrFS: a compleat file system for commodity SSDs · EuroSys 2022 |
Storage systems › storage reliability
file system reliability |
0.1 | 1 | 2017 | File Systems Fated for Senescence? Nonsense, Says Science! · FAST 2017 |
Storage systems › i/o optimization › write optimization
write-optimized data structure |
0.1 | 1 | 2015 | BetrFS: Write-Optimization in a Kernel File System · ACM Trans. Storage 2015 |
Methods — techniques the papers use, named apart from their topics
sketching · 0.9approximate data structure · 0.9b-epsilon tree · 0.7copy-on-abundant-write · 0.5write-optimized dictionary · 0.3range-rename · 0.3path indexing · 0.3zoning · 0.3range deletion · 0.3late-binding journaling · 0.3write-optimized data structures · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Focus! Fast On-disk Concurrency-control Using SketchesabstractConcurrency-control (CC) mechanisms are essential for ensuring consistency in large-scale key-value stores, but traditional approaches face significant challenges. Mechanisms like 2PL and OCC incur high CPU overheads. Timestamp-based mechanisms are faster but require storing timestamps for every key, resulting in substantial space overhead and numerous I/O operations in disk-based systems. We address these challenges by decomposing timestamp-based CC schemes into two components: a timestamp storage system and a CC protocol. We then show that the timestamp storage system can approximate timestamps for keys not used by ongoing transactions, substantially reducing memory requirements and I/O while maintaining correctness for various protocols (STO, MVTO, and TicToc). We introduce FPSketch, an approximate timestamp storage system, and our evaluation with SplinterDB shows that FPSketch outperforms 2PL and OCC by up to 14× on some workloads and disk-based CC systems by up to 5.9×. Remarkably, FPSketch with just 32KiB of memory yields performance comparable to an idealized in-memory implementations in our evaluation. FPSketch makes timestamp-based concurrency control mechanisms practical for disk-based key-value stores. Deukyeon Hwang, Alexander Conway 0001, Carlos Garcia-Alvarado, Jun Yuan 0006, Naama Ben-David, Rob Johnson 0001, Adriana Szekeres |
Proc. ACM Manag. Data | 4 |
| 2022 | BetrFS: a compleat file system for commodity SSDsabstractDespite the existence of file systems tailored for flash and over a decade of research into flash file systems, this paper shows that no single Linux file system performs consistently well on a commodity SSD across different workloads. We define a compleat file system as one where no workloads realize less than 30% of the best file system's performance, and most, if not all, workloads realize at least 85% of the best file system's performance, across a diverse set of microbenchmarks and applications. No file system is compleat on commodity SSDs. This paper demonstrates that one can construct a single compleat file system for commodity SSDs by introducing a set of optimizations over BetrFS. BetrFS is a compleat file system on HDDs, matching the fastest Linux file systems in its worst cases, and, in its best cases, improving performance by up to two orders of magnitude. Yizheng Jiao, Simon Bertron, Luke Zeller, Rory Bennett, Nirjhar Mukherjee, Michael A. Bender, Michael Condict, Alexander Conway 0001, Martin Farach-Colton, Xiongzi Ge, William Jannen, Rob Johnson 0001, Donald E. Porter, Jun Yuan 0006 |
EuroSys | 15 |
| 2021 | Copy-on-Abundant-Write for Nimble File System ClonesabstractMaking logical copies, or clones, of files and directories is critical to many real-world applications and workflows, including backups, virtual machines, and containers. An ideal clone implementation meets the following performance goals: (1) creating the clone has low latency; (2) reads are fast in all versions (i.e., spatial locality is always maintained, even after modifications); (3) writes are fast in all versions; (4) the overall system is space efficient. Implementing a clone operation that realizes all four properties, which we call a nimble clone , is a long-standing open problem. This article describes nimble clones in B-ϵ-tree File System (BetrFS), an open-source, full-path-indexed, and write-optimized file system. The key observation behind our work is that standard copy-on-write heuristics can be too coarse to be space efficient, or too fine-grained to preserve locality. On the other hand, a write-optimized key-value store, such as a Bε-tree or an log-structured merge-tree (LSM)-tree, can decouple the logical application of updates from the granularity at which data is physically copied. In our write-optimized clone implementation, data sharing among clones is only broken when a clone has changed enough to warrant making a copy, a policy we call copy-on-abundant-write . We demonstrate that the algorithmic work needed to batch and amortize the cost of BetrFS clone operations does not erode the performance advantages of baseline BetrFS; BetrFS performance even improves in a few cases. BetrFS cloning is efficient; for example, when using the clone operation for container creation, BetrFS outperforms a simple recursive copy by up to two orders-of-magnitude and outperforms file systems that have specialized Linux Containers (LXC) backends by 3--4×. Yang Zhan 0001, Alexander Conway 0001, Yizheng Jiao, Nirjhar Mukherjee, Ian Groombridge, Michael A. Bender, Martin Farach-Colton, William Jannen, Rob Johnson 0001, Donald E. Porter, Jun Yuan 0006 |
ACM Trans. Storage | 11 |
| 2020 | How to Copy Files
Yang Zhan 0001, Alexander Conway 0001, Yizheng Jiao, Nirjhar Mukherjee, Ian Groombridge, Michael A. Bender, Martin Farach-Colton, William Jannen, Rob Johnson 0001, Donald E. Porter, Jun Yuan 0006 |
FAST | 11 |
| 2019 | Small Refinements to the DAM Can Have Big Consequences for Data-Structure DesignabstractStorage devices have complex performance profiles, including costs to initiate IOs (e.g., seek times in hard drives), parallelism and bank conflicts (in SSDs), costs to transfer data, and firmware-internal operations. The Disk-Access Machine (DAM) model simplifies reality by assuming that storage devices transfer data in blocks of size B and that all transfers have unit cost. Despite its simplifications, the DAM model is reasonably accurate. In fact, if B is set to the half-bandwidth point, where the latency and bandwidth of the hardware are equal, the DAM approximates the IO cost on any hardware to within a factor of 2. Furthermore, the DAM explains the popularity of B-trees in the 70s and the current popularity of B-trees and log-structured merge trees. But it fails to explain why some B-trees use small nodes, whereas all B-trees use large nodes. In a DAM, all IOs, and hence all nodes, are the same size. In this paper, we show that the affine and PDAM models, which are small refinements of the DAM model, yield a surprisingly large improvement in predictability without sacrificing ease of use. We present benchmarks on a large collection of storage devices showing that the affine and PDAM models give good approximations of the performance characteristics of hard drives and SSDs, respectively. We show that the affine model explains node-size choices in B-trees and B+-trees. Furthermore, the models predict that the B-tree is highly sensitive to variations in the node size whereas B-trees are much less sensitive. These predictions are born out empirically. Finally, we show that in both the affine and PDAM models, it pays to organize data structures to exploit varying IO size. In the affine model, B-trees can be optimized so that all operations are simultaneously optimal, even up to lower order terms. In the PDAM model, B-trees (or B+-trees) can be organized so that both sequential and concurrent workloads are handled efficiently. We conclude that the DAM model is useful as a first cut when designing or analyzing an algorithm or data structure but the affine and PDAM models enable the algorithm designer to optimize parameter choices and fill in design details. Michael A. Bender, Alexander Conway 0001, Martin Farach-Colton, William Jannen, Yizheng Jiao, Rob Johnson 0001, Eric Knorr, Sara McAllister, Nirjhar Mukherjee, Prashant Pandey 0001, Donald E. Porter, Jun Yuan 0006, Yang Zhan 0001 |
SPAA | 12 |
| 2018 | The Full Path to Full-Path Indexing
Yang Zhan 0001, Alexander Conway 0001, Yizheng Jiao, Eric Knorr, Michael A. Bender, Martin Farach-Colton, William Jannen, Rob Johnson 0001, Donald E. Porter, Jun Yuan 0006 |
FAST | 10 |
| 2018 | Efficient Directory Mutations in a Full-Path-Indexed File SystemabstractFull-path indexing can improve I/O efficiency for workloads that operate on data organized using traditional, hierarchical directories, because data is placed on persistent storage in scan order. Prior results indicate, however, that renames in a local file system with full-path indexing are prohibitively expensive. This article shows how to use full-path indexing in a file system to realize fast directory scans, writes, and renames. The article introduces a range-rename mechanism for efficient key-space changes in a write-optimized dictionary. This mechanism is encapsulated in the key-value Application Programming Interface (API) and simplifies the overall file system design. We implemented this mechanism in B ε -trees File System (BetrFS), an in-kernel, local file system for Linux. This new version, BetrFS 0.4, performs recursive greps 1.5x faster and random writes 1.2x faster than BetrFS 0.3, but renames are competitive with indirection-based file systems for a range of sizes. BetrFS 0.4 outperforms BetrFS 0.3, as well as traditional file systems, such as ext4, Extents File System (XFS), and Z File System (ZFS), across a variety of workloads. Yang Zhan 0001, Yizheng Jiao, Donald E. Porter, Alexander Conway 0001, Eric Knorr, Martin Farach-Colton, Michael A. Bender, Jun Yuan 0006, William Jannen, Rob Johnson 0001 |
ACM Trans. Storage | 8 |
| 2017 | File Systems Fated for Senescence? Nonsense, Says Science!
Alexander Conway 0001, Ainesh Bakshi, Yizheng Jiao, William Jannen, Yang Zhan 0001, Jun Yuan 0006, Michael A. Bender, Rob Johnson 0001, Bradley C. Kuszmaul, Donald E. Porter, Martin Farach-Colton |
FAST | 6 |
| 2017 | Writes Wrought Right, and Other Adventures in File System OptimizationabstractFile systems that employ write-optimized dictionaries (WODs) can perform random-writes, metadata updates, and recursive directory traversals orders of magnitude faster than conventional file systems. However, previous WOD-based file systems have not obtained all of these performance gains without sacrificing performance on other operations, such as file deletion, file or directory renaming, or sequential writes. Using three techniques, late-binding journaling , zoning , and range deletion , we show that there is no fundamental trade-off in write-optimization. These dramatic improvements can be retained while matching conventional file systems on all other operations. BetrFS 0.2 delivers order-of-magnitude better performance than conventional file systems on directory scans and small random writes and matches the performance of conventional file systems on rename, delete, and sequential I/O. For example, BetrFS 0.2 performs directory scans 2.2 × faster, and small random writes over two orders of magnitude faster, than the fastest conventional file system. But unlike BetrFS 0.1, it renames and deletes files commensurate with conventional file systems and performs large sequential I/O at nearly disk bandwidth. The performance benefits of these techniques extend to applications as well. BetrFS 0.2 continues to outperform conventional file systems on many applications, such as as rsync, git-diff, and tar, but improves git-clone performance by 35% over BetrFS 0.1, yielding performance comparable to other file systems. Jun Yuan 0006, Yang Zhan 0001, William Jannen, Prashant Pandey 0001, Amogh Akshintala, Kanchan Chandnani, Pooja Deo, Zardosht Kasheff, Leif Walsh, Michael A. Bender, Martin Farach-Colton, Rob Johnson 0001, Bradley C. Kuszmaul, Donald E. Porter |
ACM Trans. Storage | 1 |
| 2016 | Optimizing Every Operation in a Write-optimized File System
Jun Yuan 0006, Yang Zhan 0001, William Jannen, Prashant Pandey 0001, Amogh Akshintala, Kanchan Chandnani, Pooja Deo, Zardosht Kasheff, Leif Walsh, Michael A. Bender, Martin Farach-Colton, Rob Johnson 0001, Bradley C. Kuszmaul, Donald E. Porter |
FAST | 1 |
| 2016 | Optimizing Every Operation in a Write-optimized File System
Jun Yuan 0006, Yang Zhan 0001, William Jannen, Prashant Pandey 0001, Amogh Akshintala, Kanchan Chandnani, Pooja Deo, Zardosht Kasheff, Leif Walsh, Michael A. Bender, Martin Farach-Colton, Rob Johnson 0001, Bradley C. Kuszmaul, Donald E. Porter |
USENIX ATC | 1 |
| 2015 | BetrFS: A Right-Optimized Write-Optimized File System
William Jannen, Jun Yuan 0006, Yang Zhan 0001, Amogh Akshintala, John Esmet, Yizheng Jiao, Ankur Mittal, Prashant Pandey 0001, Phaneendra Reddy, Leif Walsh, Michael A. Bender, Martin Farach-Colton, Rob Johnson 0001, Bradley C. Kuszmaul, Donald E. Porter |
FAST | 2 |
| 2015 | BetrFS: Write-Optimization in a Kernel File SystemabstractThe B ε -tree File System , or B e trFS (pronounced “better eff ess”), is the first in-kernel file system to use a write-optimized data structure (WODS). WODS are promising building blocks for storage systems because they support both microwrites and large scans efficiently. Previous WODS-based file systems have shown promise but have been hampered in several ways, which B e trFS mitigates or eliminates altogether. For example, previous WODS-based file systems were implemented in user space using FUSE, which superimposes many reads on a write-intensive workload, reducing the effectiveness of the WODS. This article also contributes several techniques for exploiting write-optimization within existing kernel infrastructure. B e trFS dramatically improves performance of certain types of large scans, such as recursive directory traversals, as well as performance of arbitrary microdata operations, such as file creates, metadata updates, and small writes to files. B e trFS can make small, random updates within a large file 2 orders of magnitude faster than other local file systems. B e trFS is an ongoing prototype effort and requires additional data-structure tuning to match current general-purpose file systems on some operations, including deletes, directory renames, and large sequential writes. Nonetheless, many applications realize significant performance improvements on B e trFS. For instance, an in-place rsync of the Linux kernel source sees roughly 1.6--22 × speedup over commodity file systems. William Jannen, Jun Yuan 0006, Yang Zhan 0001, Amogh Akshintala, John Esmet, Yizheng Jiao, Ankur Mittal, Prashant Pandey 0001, Phaneendra Reddy, Leif Walsh, Michael A. Bender, Martin Farach-Colton, Rob Johnson 0001, Bradley C. Kuszmaul, Donald E. Porter |
ACM Trans. Storage | 2 |
| 2012 | CAWDOR: Compiler Assisted Worm DefenseabstractThis paper explores how much the source code analysis can assist worm defense system. Previously-proposed worm defense systems have used disparate mechanisms to detect worms, analyze exploits, verify alerts, and apply mitigations. Furthermore, previous systems have not offered predictability, i.e. it is not possible to verify, in advance, that the defense system will never generate a mitigation that breaks the program. This paper describes a program transformation technique that makes collaborative worm defense systems easy to build, predictable and fast-responsive. Our transformation provides a single building block that can be used to perform worm detection, exploit analysis, alert verification, and mitigation application. In fact, our transformation makes most of these tasks trivial. Furthermore, software vendors and users can test, in advance, that the defense system will very unlikely apply a mitigation that breaks their software. Mitigations are vulnerability-specific not exploit-specific. Finally, our system can respond extremely quickly to a new worm. The exploit analysis becomes trivial so sentinel hosts can issue an alert the instant they detect a worm. We have implemented a prototype of our system based on the Jones and Kelly program transformation for memory safety. During normal operation, our system incurs only 5% overhead. We take advantage of static analysis to develop several optimizations and make the Jones and Kelly approach to memory safety efficient and practical. Jun Yuan 0006, Rob Johnson 0001 |
SCAM | 1 |