Puyuan Yang

dblp:56/9062 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Storage systems · 95% Memory systems · 5%
Artificial intelligence
1 paper
Efficient and distributed learning · 50% Generative modeling · 50%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › storage hierarchy
hybrid storage
1.022022
Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage Server · ACM Trans. Storage 2022
BCW: Buffer-Controlled Writes to HDDs for SSD-HDD Hybrid Storage Server · FAST 2020
Machine learning › Generative modeling › diffusion model
diffusion model acceleration
1.012026
DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation · AAAI 2026
Program synthesis and code generation
code generation with language models
1.012026
DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation · AAAI 2026
Storage systems › distributed storage
cloud-native storage
0.712023
Fisc: A Large-scale Cloud-native-oriented File System · FAST 2023
Storage systems › file systems
distributed file system
0.712023
Fisc: A Large-scale Cloud-native-oriented File System · FAST 2023
Storage systems › buffer management
write buffer management
0.412020
BCW: Buffer-Controlled Writes to HDDs for SSD-HDD Hybrid Storage Server · FAST 2020
Storage systems
flash and SSD
0.212016
Read/write-optimized tree indexing for solid-state drives · VLDB J. 2016
Storage systems › magnetic storage
hard disk drive
0.212022
Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage Server · ACM Trans. Storage 2022
Storage systems › flash and SSD
solid-state drive
0.212022
Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage Server · ACM Trans. Storage 2022
Storage systems › buffer management
write buffering
0.212022
Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage Server · ACM Trans. Storage 2022
Memory systems › non-volatile memory › write reliability
write endurance
0.212022
Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage Server · ACM Trans. Storage 2022
Indexing and storage engines
tree index
0.112016
Read/write-optimized tree indexing for solid-state drives · VLDB J. 2016

Methods — techniques the papers use, named apart from their topics

large language model · 2.0genetic algorithm · 2.0automated debugging · 2.0agent-based planning · 2.0mixed i/o scheduling · 0.6buffer-controlled write · 0.6
YearPublicationVenuePosition
2026 DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation
abstract
Diffusion models have achieved remarkable success in image and video generation. However, their inherently multiple step inference process imposes substantial computational overhead, hindering real-world deployment. Accelerating diffusion models is therefore essential, yet determining how to combine multiple model acceleration techniques remains a significant challenge. To address this issue, we introduce a framework driven by large language models (LLMs) for automated acceleration code generation and evaluation. First, we present DiffBench, a comprehensive benchmark that implements a three stage automated evaluation pipeline across diverse diffusion architectures, optimization combinations and deployment scenarios. Second, we propose DiffAgent, an agent that generates optimal acceleration strategies and codes for arbitrary diffusion models. DiffAgent employs a closed-loop workflow in which a planning component and a debugging component iteratively refine the output of a code generation component, while a genetic algorithm extracts performance feedback from the execution environment to guide subsequent code refinements. We provide a detailed explanation of the DiffBench construction and the design principles underlying DiffAgent. Extensive experiments show that DiffBench offers a thorough evaluation of generated codes and that DiffAgent significantly outperforms existing LLMs in producing effective diffusion acceleration strategies.
Jiajun Jiao, Haowei Zhu, Puyuan Yang, Jianghui Wang, Ziqiong Liu, Dong Li 0025, Yuejian Fang, Jun-Hai Yong, Bin Wang 0034, Emad Barsoum
AAAI3
2025 Diagno-S: Towards Verifiable Medical Reasoning via Diagnostic Decomposition
abstract
While Large Language Models (LLMs) show great promise in emulating clinicians' cognitive processes, ensuring the reliability of their reasoning remains unresolved. Existing chain-of-thought often entangles logic with supporting medical knowledge, forming an opaque “black box” that conceals errors and induces a critical failure mode we term retrospective knowledge forgery—where models fabricate facts posthoc to justify reasoning. To tackle this, we propose Diagno-S (Diagnosis-Separation), a framework that decouples reasoning and knowledge for diagnostic-oriented optimization. Diagno-S compels models to decompose reasoning into independently verifiable logical and factual steps, making hidden flaws traceable. Leveraging this transparency, we employ curriculum learning and a diagnostic module to detect and correct these errors via direct preference optimization. Diagno-S attains state-of-the-art results on seven medical QA benchmarks, including MedQA, PubMedQA, and MedXpert, and provides a systematic solution to knowledge forgery, advancing trustworthy medical AI.
Xunfei Zhu, Chenhan Wang, Shangkeng Wu, Puyuan Yang
BIBM4
2025 Bayesianly-Corrected, Bandit-Optimized Multi-agent LLMs: Rethinking Agents via Control-Theoretic Dynamics
Xunfei Zhu, Shuaizhuo Yuan, Hao Leng, Puyuan Yang, Dexin Liu
PRICAI4
2023 Fisc: A Large-scale Cloud-native-oriented File System
Qiang Li 0045, Lulu Chen, Xiaoliang Wang 0001, Qiao Xiang, Wenhui Yao, Minfei Huang, Puyuan Yang, Shanyang Liu, Zhaosheng Zhu, Huayong Wang, Haonan Qiu, Derui Liu, Shaozong Liu, Yaohui Wu, Zhiwu Wu, Zicheng Luo, Yuchao Shao, Gexiao Tian, Zhongjie Wu, Zheng Cao 0003, Jiwu Shu, Jie Wu 0003, Jiesheng Wu
FAST9
2022 Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage Server
abstract
Hybrid storage servers combining solid-state drives (SSDs) and hard-drive disks (HDDs) provide cost-effectiveness and μs-level responsiveness for applications. However, observations from cloud storage system Pangu manifest that HDDs are often underutilized while SSDs are overused, especially under intensive writes. It leads to fast wear-out and high tail latency to SSDs. On the other hand, our experimental study reveals that a series of sequential and continuous writes to HDDs exhibit a periodic, staircase-shaped pattern of write latency, i.e., low (e.g., 35 μs), middle (e.g., 55 μs), and high latency (e.g., 12 ms), resulting from buffered writes within HDD’s controller. It inspires us to explore and exploit the potential μs-level IO delay of HDDs to absorb excessive SSD writes without performance degradation. We first build an HDD writing model for describing the staircase behavior and design a profiling process to initialize and dynamically recalibrate the model parameters. Then, we propose a Buffer-Controlled Write approach (BCW) to proactively control buffered writes so that low- and mid-latency periods are scheduled with application data and high-latency periods are filled with padded data. Leveraging BCW, we design a mixed IO scheduler (MIOS) to adaptively steer incoming data to SSDs and HDDs. A multi-HDD scheduling is further designed to minimize HDD-write latency. We perform extensive evaluations under production workloads and benchmarks. The results show that MIOS removes up to 93% amount of data written to SSDs, reduces average and 99 th -percentile latencies of the hybrid server by 65% and 85%, respectively.
Shucheng Wang, Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang, Changsheng Xie 0001
ACM Trans. Storage7
2020 BCW: Buffer-Controlled Writes to HDDs for SSD-HDD Hybrid Storage Server
Shucheng Wang, Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang
FAST7
2020 SeRW: Adaptively Separating Read and Write upon SSDs of Hybrid Storage Server in Clouds
abstract
Nowadays, cloud providers embrace hybrid storage servers to reap both high IO performance of solid-state drives (SSDs) and low-cost of hard disk drives (HDDs). These hybrid storage servers generally employ SSDs as primary storage directly serving requests from front-end applications while using HDDs as the secondary storage to provide sufficient storage capacity.
Qiang Cao 0001, Shucheng Wang, Jie Yao 0001, Puyuan Yang
ICPP7
2019 Analysis of and Optimization for Write-dominated Hybrid Storage Nodes in Cloud
abstract
Cloud providers like the Alibaba cloud routinely and widely employ hybrid storage nodes composed of solid-state drives (SSDs) and hard disk drives (HDDs), reaping their respective benefits: performance from SSD and capacity from HDD. These hybrid storage nodes generally write incoming data to its SSDs and then flush them to their HDD counterparts, referred to as the SSD Write Back (SWB) mode, thereby ensuring low write latency. When comprehensively analyzing real production workloads from Pangu, a large-scale storage platform underlying the Alibaba cloud, we find that (1) there exist many write dominated storage nodes (WSNs); however, (2) under the SWB mode, the SSDs of these WSNs suffer from severely high write intensity and long tail latency. To address these unique observed problems of WSNs, we present SSD Write Redirect (SWR), a runtime IO scheduling mechanism for WSNs. SWR judiciously and selectively forwards some or all SSD-writes to HDDs, adapting to runtime conditions. By effectively offloading the right amount of write IOs from overburdened SSDs to underutilized HDDs in WSNs, SWR is able to adequately alleviate the aforementioned problems suffered by WSNs. This significantly improves overall system performance and SSD endurance. Our trace-driven evaluation of SWR, through replaying production workload traces collected from the Alibaba cloud in our cloud testbed, shows that SWR decreases the average and 99til-percentile latencies of SSD-writes by up to 13% and 47% respectively, notably improving system performance. Meanwhile the amount of data written to SSDs is reduced by up to 70%, significantly improving SSD lifetime.
Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang
SoCC8
2019 SPA-SSD: Exploit Heterogeneity and Parallelism of 3D SLC-TLC Hybrid SSD to Improve Write Performance
abstract
To address the write performance problem suffered by MLC/TLC flash, researchers have proposed hybrid SSD that aims to combine the strengths of SLC flash, used as the write-buffer zone for its superior write performance, and MLC/TLC flash, as the capacity zone for its high storage density. While leveraging SLC as a physical write-buffer zone is proven effective in traditional 2D hybrid SSDs, how to effectively incorporate SLC into a 3D-stacked TLC to form a hybrid SSD has not been studied to the best of our knowledge. Yet this is a timely and important performance issue for 3D-stacked TLC given its one-shot programming scheme that results in much worse write performance than the programming scheme in 2D TLC where pages are associated with different bits of a cell and programmed in sequence separately. We believe that naively adopting the two-physical-zone approach to 3D hybrid SSD will miss a great opportunity for performance optimization because it ignores the inherent four-level parallelism (channel/chip/die/plane) of the flash chip array. To this end, we propose in this paper an SLC and Parallelism Aware hybrid SSD (SPA-SSD) to take full advantages of SLC's superior write performance, the internal multi-level parallelism of SSD, and the high storage density of 3D-stacked TLC flash. Two novel techniques enable SPA-SSD to be highly effective: (1) Type-Parallelism Joint Page Allocation (TPJ-PA), which allocates pages for write transactions according to not only available SLC pages but also parallelism to maximize resource utilization within the hybrid SSD, and (2) Queue-length and Parallelism Constrained Data Migration (QPC-DM), which triggers data migration without degrading user write performance by analyzing the device queue length and available flash resources. To evaluate performance of SPA-SSD, a hybrid SSD simulator, called HybridSim, is developed based on MQSim. Experimental results on HybridSim show that TPJ-PA improves write throughput by 60%, while QPC-DM improves write throughput by up to 10 times. Besides, trace-driven experiments on HybridSSD demonstrate that SPA-SSD improves the write latency to the flash by up to two orders of magnitude over the state-of-the-art designs.
Wenhui Zhang 0005, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang
ICCD6
2019 VScan: Efficiently Analyzing Surveillance Videos via Model-joint Mechanism
abstract
Identifying key scenes in massive surveillance videos is extremely challenging because these scenes occur rarely while automotive identification using full-feature neural network (NN) models consumes immense computational resources. This paper proposes VScan, an efficient model-joint mechanism that adaptively schedules streams on a light-weight NN model and a full-feature NN model for analyzing videos concurrently. These two combined models with overlapped detectable objects are generic and well-developed. The former model fast scans videos to seek potential interest scenes. Only the streams with identified scenes are further analyzed by the latter model. We provide a model selection approach to select a light-weight model with an appropriate accuracy and high throughput. VScan further determines key parameters to correct predictions at runtime, thus guaranteeing the recall of target scenes. The full-feature model is responsible for ensuring output precision. To maintain a high hardware efficiency and utilization dynamically, VScan uses automatic sampling to reduce unnecessary computations, proposes stream scheduling to maximize hardware usage, and designs GPU scheduling to optimize the data processing flow. Experimental results show that benefitting from the model-joint mechanism and runtime scheduling optimizations, VScan significantly boosts the video processing throughput by up to 15x without key scene loss.
Qiang Cao 0001, Jie Yao 0001, Puyuan Yang
ICPP5
2019 BFO: Batch-File Operations on Massive Files for Consistent Performance Improvement
abstract
Existing local file systems, designed to support a typical single-file access pattern only, can lead to poor performance when accessing a batch of files, especially small files. This single-file pattern essentially serializes accesses to batched files one by one, resulting in a large number of non-sequential, random, and often dependent I/Os between file data and metadata at the storage ends. We first experimentally analyze the root cause of such inefficiency in batch-file accesses. Then, we propose a novel batch-file access approach, referred to as BFO for its set of optimized Batch-File Operations, by developing novel BFOr and BFOw operations for fundamental read and write processes respectively, using a two-phase access for metadata and data jointly. The BFO offers dedicated interfaces for batch-file accesses and additional processes integrated into existing file systems without modifying their structures and procedures. We implement a BFO prototype on ext4, one of the most popular file systems. Our evaluation results show that the batch-file read and write performances of BFO are consistently higher than those of the traditional approaches regardless of access patterns, data layouts, and storage media, with synthetic and real-world file sets. BFO improves the read performance by up to 22.4× and 1.8× with HDD and SSD respectively; and boosts the write performance by up to 111.4× and 2.9× with HDD and SSD respectively. BFO also demonstrates consistent performance advantages when applied to four representative applications, Linux cp, Tar, GridFTP, and Hadoop.
Yang Yang 0068, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang
MSST7
2016 Read/write-optimized tree indexing for solid-state drives
Peiquan Jin, Chengcheng Yang, Christian S. Jensen, Puyuan Yang, Lihua Yue
VLDB J.4
2015 Optimizing B+-tree for hybrid storage systems
Peiquan Jin, Puyuan Yang, Lihua Yue
Distributed Parallel Databases2
2014 Adaptive in-page logging for flash-memory storage systems
Peiquan Jin, Puyuan Yang, Shouhong Wan, Lihua Yue
Frontiers Comput. Sci.3