VLDB 2026 Research / reviewers in the wild / expert
Yanqi Pan
dblp:330/1367
· DBLP profile ↗
17ranked-venue papers
4as first author
17since 2021 · last 2026
0009-0007-7832-0599ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-author · 14 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast and Parallelized Crash Consistency with Opportunistic Order EliminationabstractFile systems rely on enforcing storage order for crash recovery. However, this ordering requirement limits I/O parallelism, especially in persistent memory (PM) file systems where ordered I/Os are synchronized for direct access, preventing file systems from fully exploiting PM I/O parallelism. Yanqi Pan, Wen Xia, Peixin Zeng, Yuchen Shan |
EuroSys | 2 |
| 2026 | Once Rolling Hashing is Enough: Exploiting Rolling Hash Reuse in Delta CompressionabstractIn backup storage, delta compression successfully achieves a much higher data reduction ratio than chunk-level deduplication by applying a finer granularity in redundancy detection and elimination. However, it introduces additional, intensive computation overhead and results in a 50%–70% worsening backup throughput. Haoliang Tan, Wenhao Ou, Xiangyu Zou, Yanqi Pan, Zhaoquan Gu, Wen Xia |
EuroSys | 5 |
| 2026 | Towards Condensed and Efficient Read-Only File System via Sort-Enhanced Compression
Yanqi Pan, Wen Xia, Xiangyu Zou, Darong Yang, Jubin Zhong, Hua Liao |
FAST | 3 |
| 2026 | Argus: A Precise and Efficient Resemblance Detection for Post-Deduplication Delta CompressionabstractFor data reduction techniques used in storage systems, delta compression is often implemented after deduplication, having been shown to achieve a much higher compression ratio by efficiently detecting and compressing similar data chunks. Unfortunately, existing resemblance detection approaches cannot maintain both high throughput and data reduction ratio simultaneously since they either introduce heavy calculation overhead or generate useless features that reduce the accuracy of resemblance detection. In this article, we propose Argus, a fast and precise resemblance detection approach using two primary techniques to improve data-reduction efficiency significantly. First, Argus utilizes a Bin-Wise Partitioning strategy, which separates the rolling hash values of each data chunk into different subsets according to the suffix bits of the hash value and generates features from these subsets. Thus, Argus generates features more efficiently, which both achieves high detection accuracy and improves feature generation speed. Second, based on the efficient Bin-Wise Partitioning strategy, Argus utilizes fine-grained Gear rolling hash and a Plain Feature strategy to manage the granularity of the content represented in the feature, increasing the probability of feature matching and catching as many similar chunks as possible. Consequently, Argus can achieve better detection accuracy, resulting in a much higher compression ratio than previous works while minimizing the computational overhead for resemblance detection. Our evaluation results driven by several real-world datasets suggest that, compared to the state-of-the-art approaches, Argus achieves up to 1.64 × (DeepSketch) and 2.29 × (Finesse, Odess, and N-Transform) higher delta compression ratio and achieves up to 19.9 × (N-Transform), 5.57 × (Finesse), and 1.18 × (Odess) faster feature generation speed. Xiangyu Zou, Yunsheng Dong, Philip Shilane, Yanqi Pan, Wen Xia |
ACM Trans. Storage | 5 |
| 2025 | Simplifying and Accelerating NOR Flash I/O Stack for RAM-Restricted MicrocontrollersabstractNOR flash has been increasingly popular for RAM-restricted microcontrollers due to its small package, high reliability, etc. To satisfy RAM restrictions, existing NOR flash file systems migrate their functionalities, i.e., block-level data organization and wear leveling (WL), from RAM to NOR flash. However, such fine-grained block-level management introduces frequent index updates and NOR flash scanning, leading to severe I/O amplification, which further deteriorates as they are decoupled in existing NOR flash file systems. Yanqi Pan, Wen Xia, Xiangyu Zou, Darong Yang, Liang Shi 0001, Hongwei Du 0001 |
ASPLOS (2) | 2 |
| 2025 | Garbage Collection Does Not Only Collect Garbage: Piggybacking-Style Defragmentation for Deduplicated Backup StorageabstractDeduplication is widely used in backup storage and reduces storage overhead by allowing backups to share common data chunks. However, it naturally disrupts the sequential layout of backup images, leading to fragmentation, which slows down backup restoration. Existing solutions to this issue often come with trade-offs, either reducing deduplication effectiveness or introducing significant I/O overhead. Dingbang Liu, Xiangyu Zou, Tao Lu 0014, Philip Shilane, Wen Xia, Yanqi Pan |
EuroSys | 7 |
| 2025 | Overcoming the Last Mile between Log-Structured File Systems and Persistent Memory via Scatter LoggingabstractLog-structured file system (LFS) has been a popular option for persistent memory (PM) for its high write performance and lightweight crash consistency protocols. However, even with PM's byte-addressable I/O interface, existing PM LFSes still maintain contiguous space for log locality while using heavy garbage collection (GC) to synchronously reclaim space, causing up to 50% I/O performance degradation on PM. Thus, there exists a last-mile problem between the contiguous space management of LFS (that induces GC) and the non-contiguous byte-addressability of PM. Yanqi Pan, Yuchen Shan, Wen Xia |
EuroSys | 2 |
| 2025 | Don't Maintain Twice, It's Alright: Merged Metadata Management in Deduplication File System with GogetaFS
Yanqi Pan, Wen Xia, Erci Xu, Xiangyu Zou |
FAST | 1 |
| 2025 | ALPHA: A Scalable Lock-Free Partitioned Hash Index for Persistent Memory on NUMA ArchitecturesabstractHash indexes are widely used in modern dataintensive applications to support efficient query performance. While persistent memory (PM) provides larger capacity and byteaddressability, we identify that PM-based hash indexes suffer from scalability bottlenecks under high concurrency. This stems from NUMA access imposing significant costs, exacerbating the inherently request-agnostic I/O issues in concurrent read/write operations. To this end, we present ALPHA, a highly scalable hash index designed for persistent memory. The key idea with ALPHA lies in exploiting the natural hash-based data access features of hash indexes to reduce I/O irrelevant to data requests. We propose a NUMA-friendly framework to reduce the negative impacts of remote access through the cross-thread delegation mechanism. To further improve throughput, we use two techniques to address read and write issues of the PM-oriented hash index. First, we propose a novel two-layer hash structure to minimize PM access by filtering unnecessary key retrievals during data reads. Second, we devise a partition-based finegrained data access mechanism that enables lock-free concurrent writes without excessive PM write overheads. Our comprehensive experimental evaluation shows that ALPHA outperforms state-of-the-art PM-based hash indexes by up to$22.17 \times$in throughput while reducing tail latency by up to 97.1 %. Qiyang Zheng, Hao Hu 0015, Yanqi Pan, Wen Xia |
ICCD | 4 |
| 2025 | Fast and Synchronous Crash Consistency with Metadata Write-Once File System
Yanqi Pan, Wen Xia, Xiangyu Zou, Zhenhua Li 0001, Chentao Wu |
OSDI | 1 |
| 2024 | H2C-Dedup: Reducing I/O and GC Amplification for QLC SSDs from the Deduplication Metadata PerspectiveabstractQLC SSDs have gained increasing popularity in cloud computing, PCs, and smartphones due to their low prices and high density, but they suffer from extremely limited endurance. Deduplication can convert redundant chunk writes into fine-grained metadata updates, thereby promising to alleviate QLC wear. Nevertheless, our observation shows that existing deduplication approaches cause even more I/Os than non-deduplication systems. We find that the amplification comes from two sources: (1) I/O amplification due to the mismatched granularity between SSD I/O size (e.g., 4--16 KiB page) and deduplication metadata I/O size, and (2) SSD garbage collection (GC) amplification due to deduplication metadata updates for eliminating redundant chunks (i.e., increment reference count). To address the above problem, this paper proposes H2C-Dedup, which employs two essential techniques. First, to address I/O amplification, cold2hot-heating technique utilizes a log-structured metadata I/O scheme, which ensures that the deduplication metadata is flushed until it is accumulated within I/O cache to match with the OS I/O granularity. Second, to address GC amplification, hot2cold-suppression technique divides metadata into hot (e.g., to store reference count) and cold (e.g., to store fingerprint) segments, delta-encoding the hot entries to ensure both hot and cold metadata will not be modified once they are durable. As a result, H2C-Dedup significantly reduces the deduplication-induced I/O and GC amplification. We implement H2C-Dedup based on F2FS. Extensive experiments on the FEMU platform using microbenchmarks and real-world traces indicate that H2C-Dedup can extend to at most 3.3× and 3.4× lifespan while accelerating 26% and 37% I/O performance compared to SmartDedup and HF-Dedupe. Yunsheng Dong, Boju Chen, Yanqi Pan, Xiangyu Zou, Wen Xia |
SoCC | 3 |
| 2024 | Delaying Crash Consistency for Building A High-Performance Persistent Memory File SystemabstractPersistent Memory (PM), with its low latency, byte-addressability, and non-volatility, sparks various PM file systems. Nevertheless, we observe that these file systems can waste more than 50% PM I/O bandwidth, hindering the value of PM. Our in-depth analysis reveals that this can be rooted in their strong consistency assumptions that require synchronized PM I/O for strict data integrity. Nevertheless, these I/Os, including attribute updates, space management, and journaling, are generally small, random, and fail to match with the PM I/O characteristics. We present HUNTER, a POSIX-compliant PM file system that aims to unleash PM I/O performance with strict data integrity. Our key insight is delaying crash consistency with emerging flush-on-fail (FoF) hardware. We achieve this insight by first leveraging a PM-simplified soft update mechanism to persist metadata (i.e., consistent view) in the background; while maintaining up-to-date states in DRAM (i.e., latest view) for instantaneous user responding. Compared to prior works, HUNTER carefully decouples the latest view from the consistent view, and thus, background operations will not block the foreground paths, which fully exploits soft update efficiency. Under this architecture, FoF hardware allows a residual energy window for state synchronization during system failures. However, ensuring strict data integrity is still non-trivial due to existing deficient sync mechanisms that might either hide the benefits of soft update or fail to flush all states during the FoF window. We propose collaborated synchronization that bridges the two views to minimize PM I/O to address the issue. Our extensive experiments show that HUNTER can achieve 1.35–7.51× I/O throughput compared to existing PM file systems. Under the strict workload configurations, HUNTER outperforms the fastest PM file system by 72–122%. Yanqi Pan, Wen Xia, Xiangyu Zou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | DRPTM: A Decoupled Read-efficient High-scalable Persistent Transactional MemoryabstractPersistent transactional memory (PTM) exploits transactions to provide an easy crash-consistent interface for persistent memory (PM). However, because of the substantial reader-side overhead brought on by the low bandwidth and long persistence latency of PM, present PTM research cannot scale effectively. This paper proposes a highly scalable PTM system, DRPTM, which allows nearly non-overhead reads without lowering the isolation level. DRPTM decouples persistence latency from concurrency control and traces the read-only copy maintained in logs as a lightweight read set. The evaluation shows that DRPTM significantly outperforms the state-of-the-art PTM systems for various workloads and achieves near-linear scalability. Wenkai Liang, Hao Hu 0015, Xiangyu Zou, Wen Xia, Yanqi Pan |
DAC | 5 |
| 2023 | HUNTER: Releasing Persistent Memory Write Performance with A Novel PM-DRAM Collaboration ArchitectureabstractWe present HUNTER, a POSIX-compliant persistent memory (PM) file system that fully releases PM’s write performance. Compared to state-of-the-art ones, HUNTER proposes a novel PM-DRAM collaboration architecture to significantly eliminate/reduce software overheads in the write path. Expensive in-PM metadata are updated asynchronously to hide their performance penalties. Furthermore, in-PM metadata/data are laid out separately for locality awareness, enabling collaboration with asynchronous architecture. HUNTER also adopts several lightweight in-DRAM allocators/indexes to manage PM efficiently.Experimental results suggest that HUNTER achieves 2.0–3.4× write bandwidth compared to state-of-the-art PM file systems in write-intensive workloads and shows similar write bandwidth compared to bare PM. Yanqi Pan, Wen Xia, Xiangyu Zou |
DAC | 1 |
| 2023 | Detective-Dee: A Non-Intrusive In Situ Anomaly Detection and Fault Localization FrameworkabstractMaintaining the high availability of online systems requires reliable and fast online anomaly detection and fault localization. However, existing anomaly detection methods either suffer high training costs and low generalization capabilities or are designed and evaluated using offline data with limited efficacy in online usage. Furthermore, these methods' fault localization capabilities are often inadequate due to external observability constraints. Therefore, designing a new approach to address these limitations effectively is essential. To address the aforementioned limitations, this paper proposes a novel non-intrusive in situ anomaly detection and fault lo-calization framework, Detective-Dee. The proposed framework leverages a compressed sensing method for anomaly detection, which exhibits strong generalization capabilities and eliminates extensive training. Detective-Dee further improves its performance by incorporating three optimization techniques: concurrent sub-stitution sampling, Look-Up-Table-based similarity calculation, and substitution window-based threshold selection to improve parallelism and reduce computational and comparison overheads. Additionally, the framework adopts an innovative non-intrusive fault localization strategy based on anomaly detection triggering. This approach utilizes the dynamic instrumentation capabilities of eBPF, combined with extracting vulnerable function and function call chains through source code analysis, to improve the online anomaly detection capability and achieve robust fault localization with low overhead. To validate the effectiveness of Detective-Dee, we developed a prototype system and conducted a comprehensive evaluation. The results demonstrate that, compared to the state-of-the-art anomaly detection method, Detective-Dee exhibits a 4x improve-ment in anomaly detection speed while maintaining higher online and comparable offline detection ability. Furthermore, under 33 real-world fault cases across eight popular distributed systems, Detective-Dee successfully detects 31 cases and accurately locates 26 cases with less than 1% overhead, outperforming the state-of-the-art method. Yang Man, Wen Xia, Bochun Yu, Yingchi Long, Yanqi Pan |
SRDS | 7 |
| 2023 | Light-Dedup: A Light-weight Inline Deduplication Framework for Non-Volatile Memory File Systems
Jiansheng Qiu, Yanqi Pan, Wen Xia, Xiaojia Huang, Xiangyu Zou, Yu Hua 0001 |
USENIX ATC | 2 |
| 2022 | An end-to-end tracking method for polyp detectors in colonoscopy videosabstractDeep learning based computer-aided diagnosis technology demonstrates an encouraging performance in aspect of polyp lesion detection on reducing the miss rate of polyps during colonoscopies. However, to date, few studies have been conducted for tracking polyps that have been detected in colonoscopy videos, which is an essential and intuitive issue in clinical intelligent video analysis task (e.g. lesion counting, lesion retrieval, report generation). In the paradigm of conventional tracking-by-detection system, detection task for lesion localization is separated from the tracking task for cropped lesions re-identification. In the multi object tracking problem, each target is supposed to be tracked by invoking a tracker after the detector, which introduces multiple inferences and leads to external resource and time consumption. To tackle these problems, we proposed a plug-in module named instance tracking head (ITH) for synchronous polyp detection and tracking, which can be simply inserted into object detection frameworks. It embeds a feature-based polyp tracking procedure into the detector frameworks to achieve multi-task model training. ITH and detection head share the model backbone for low level feature extraction, and then low level feature flows into the separate branches for task-driven model training. For feature maps from the same receptive field, the region of interest head assigns these features to the detection head and the ITH, respectively, and outputs the object category, bounding box coordinates, and instance feature embedding simultaneously for each specific polyp target. We also proposed a method based on similarity metric learning. The method makes full use of the prior boxes in the object detector to provide richer and denser instance training pairs, to improve the performance of the model evaluation on the tracking task. Compared with advanced tracking-by-detection paradigm methods, detectors with proposed ITH can obtain comparative tracking performance but approximate 30% faster speed. Optimized model based on Scaled-YOLOv4 detector with ITH illustrates good trade-off between detection (mAP 91.70%) and tracking (MOTA 92.50% and Rank-1 Acc 88.31%) task at the frame rate of 66 FPS. The proposed structure demonstrates the potential to aid clinicians in real-time detection with online tracking or offline retargeting of polyp instances during colonoscopies. Ne Lin, Yanqi Pan, Huiyi Hu, Wenfang Zheng, Jiquan Liu, Weiling Hu, Huilong Duan, Jianmin Si |
Artif. Intell. Medicine | 4 |