Ziyi Lu

dblp:33/2581 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SIP2Net: Situational-Aware Indoor Pathloss-Map Prediction Network for Radio Map Generation
abstract
This paper presents our indoor pathloss prediction solution to ICASSP 2025 Signal Process Grand Challenge: First Indoor Path Loss Prediction Challenge. The proposed U-Net-based network incorporates dedicated asymmetric convolutions and spatial pyramid pooling to enhance reconstruction quality. Our approach achieves a weighted root mean squared error (RMSE) of 9.411 dB on the final test set, securing the 1st place in the challenge.
Wenlihan Lu, Ziyi Lu, Jia Yan 0003, Shijian Gao
ICASSP2
2024 FluidKV: Seamlessly Bridging the Gap between Indexing Performance and Memory-Footprint on Ultra-Fast Storage
abstract
Our extensive experiments reveal that existing key-value stores (KVSs) achieve high performance at the expense of a huge memory footprint that is often impractical or unacceptable. Even with the emerging ultra-fast byte-addressable persistent memory (PM), KVSs fall far short of delivering the high performance promised by PM's superior I/O bandwidth. To find the root causes and bridge the huge performance/memory-footprint gap, we revisit the architectural features of two representative indexing mechanisms (single-stage and multi-stage) and propose a three-stage KVS called FluidKV. FluidKV effectively consolidates these indexes by fast and seamlessly running incoming key-value request stream from the write-concurrent frontend stage to the memory-efficient backend stage across an intermediate stage. FluidKV also designs important enabling techniques, such as thread-exclusive logging, PM-friendly KV-block structures, and dual-grained indexes, to fully utilize both parallel-processing and high-bandwidth capabilities of ultra-fast storage hardware while reducing the overhead. We implemented a FluidKV prototype and evaluated it under a variety of workloads. The results show that FluidKV outperforms the state-of-the-art PM-aware KVSs, including ListDB and FlatStore with different indexes, by up to 9× and 3.9× in write and read throughput respectively, while cutting up to 90% of the DRAM footprint.
Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Yuxing Chen 0003, Jie Yao 0001, Anqun Pan
Proc. VLDB Endow.1
2024 Explorations and Exploitation for Parity-based RAIDs with Ultra-fast SSDs
abstract
Following a conventional design principle that pays more fast-CPU-cycles for fewer slow-I/Os, popular software storage architecture Linux Multiple-Disk (MD) for parity-based RAID (e.g., RAID5 and RAID6) assigns one or more centralized worker threads to efficiently process all user requests based on multi-stage asynchronous control and global data structures, successfully exploiting characteristics of slow devices, e.g., Hard Disk Drives (HDDs). However, we observe that, with high-performance NVMe-based Solid State Drives (SSDs), even the recently added multi-worker processing mode in MD achieves only limited performance gain because of the severe lock contentions under intensive write workloads. In this paper, we propose a novel stripe-threaded RAID architecture, StRAID, assigning a dedicated worker thread for each stripe-write (one-for-one model) to sufficiently exploit high parallelism inherent among RAID stripes, multi-core processors, and SSDs. For the notoriously performance-punishing partial-stripe writes that induce extra read and write I/Os, StRAID presents a two-stage stripe write mechanism and a two-dimensional multi-log SSD buffer. All writes first are opportunistically batched in memory, and then are written into the primary RAID for aggregated full-stripe writes or conditionally redirected to the buffer for partial-stripe writes. These buffered data are strategically reclaimed to the primary RAID. We evaluate a StRAID prototype with a variety of benchmarks and real-world traces. StRAID is demonstrated to outperform MD by up to 5.8 times in write throughput.
Shucheng Wang, Qiang Cao 0001, Hong Jiang 0001, Ziyi Lu, Jie Yao 0001, Yuxing Chen 0003, Anqun Pan
ACM Trans. Storage4
2023 PMLDS: An LSM-Tree Direct Managed Storage for Key-Value Stores on Byte-Addressable Devices
abstract
Existing key-value stores (KVSs) based on log-structured merge-tree (LSM-tree) have been broadly deployed in practice to leverage characteristics of conventional block storage via file system, but lack effective exploitation for emerging byte-addressed persistent memory (PM). We reveal that these KVSs running upon existing PM-aware File systems cause inefficient PM I/O behaviors, including 1) numerous page faults, 2) I/O misaligned with cacheline, and 3) bandwidth wastage of concurrent I/O threads. To make full use of PM without major modification for existing LSM-based KVSs, this paper proposes PMLDS, a direct managed storage for LSM-tree-based KVSs directly running upon PM. PMLDS acts as a unified I/O layer to handle all requests from KVS to PM. PMLDS designs an LSM-tree-aware data layout to directly map the KVS’s persistent objects to the storage slots with fixed location and size, thus simplifying and replacing the file system’s functionality with a minor modification. To improve I/O efficiency, PMLDS further presents three key techniques: 1) pre-allocating reusable data slots to avoid page faults, 2) forcing cacheline-alignment for small requests, and 3) scheduling asynchronous I/O threads to harness PM’s limited parallelism. We implement PMLDS and evaluate it with popular RocksDB under a variety of workloads. The results show that compared to representative PM-aware file systems such as Ext4-DAX, XFS-DAX, NOVA, and WineFS, PMLDS improves the write performance of RocksDB by up to 2.1 × while reducing the read latency by 20%~50%.
Ziyi Lu, Qiang Cao 0001, Shucheng Wang, Jie Yao 0001, Xiangrui Yang 0001
ICPP1
2023 Incentive mechanism and path planning for Unmanned Aerial Vehicle (UAV) hitching over traffic networks
Ziyi Lu, Na Yu 0005, Xuehe Wang
Future Gener. Comput. Syst.1
2022 PATS: Taming Bandwidth Contention between Persistent and Dynamic Memories
abstract
Emerging persistent memory (PM) with fast per-sistence and byte-addressability physically shares the memory channel with DRAM-based main memory. We experimentally uncover that the throughput of application accessing DRAM collapses when multiple threads access PM due to head-of-line blockage in the memory controller within CPU. To address this problem, we design a PM-Accessing Thread Scheduling (PATS) mechanism that is guided by a contention model, to adaptively tune the maximum number of contention-free concurrent PM-threads. Experimental results show that even with 14 concurrent threads accessing PM, PATS is able to allow only up to 8% decrease in the DRAM-throughput of the front-end applications (e.g., Memcached), gaining 1.5x PM-throughput speedup over the default configuration.
Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001
DATE3
2022 p2KVS: a portable 2-dimensional parallelizing framework to improve scalability of key-value stores on SSDs
abstract
Attempts to improve the performance of key-value stores (KVS) by replacing the slow Hard Disk Drives (HDDs) with much faster Solid-State Drives (SSDs) have consistently fallen short of the performance gains implied by the large speed gap between SSDs and HDDs, especially for small KV items. We experimentally and holistically explore the root causes of performance inefficiency of existing LSM-tree based KVSs running on powerful modern hardware with multicore processors and fast SSDs. Our findings reveal that the global write-ahead-logging (WAL) and index-updating (MemTable) can become bottlenecks that are as fundamental and severe as the commonly known LSM-tree compaction bottleneck, under both the single-threaded and multi-threaded execution environments.
Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Shucheng Wang
EuroSys1
2022 Mlog: Multi-log Write Buffer upon Ultra-fast SSD RAID
abstract
Parity-based RAID suffering from partial-stripe write-penalty has to introduce write buffer to fast absorb and merge incoming writes, and then flush them to RAID array in batch. However, we experimentally observe that the popular buffering mechanism as Linux RAID journal and partial parity logging (PPL) becomes a bottleneck for ultra-fast SSD-based RAID, and we further uncover that the centralized log-buffer model is the prime cause.
Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Jie Yao 0001
ICPP3
2022 StRAID: Stripe-threaded Architecture for Parity-based RAIDs with Ultra-fast SSDs
Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001, Jie Yao 0001
USENIX ATC3
2022 Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage Server
abstract
Hybrid storage servers combining solid-state drives (SSDs) and hard-drive disks (HDDs) provide cost-effectiveness and μs-level responsiveness for applications. However, observations from cloud storage system Pangu manifest that HDDs are often underutilized while SSDs are overused, especially under intensive writes. It leads to fast wear-out and high tail latency to SSDs. On the other hand, our experimental study reveals that a series of sequential and continuous writes to HDDs exhibit a periodic, staircase-shaped pattern of write latency, i.e., low (e.g., 35 μs), middle (e.g., 55 μs), and high latency (e.g., 12 ms), resulting from buffered writes within HDD’s controller. It inspires us to explore and exploit the potential μs-level IO delay of HDDs to absorb excessive SSD writes without performance degradation. We first build an HDD writing model for describing the staircase behavior and design a profiling process to initialize and dynamically recalibrate the model parameters. Then, we propose a Buffer-Controlled Write approach (BCW) to proactively control buffered writes so that low- and mid-latency periods are scheduled with application data and high-latency periods are filled with padded data. Leveraging BCW, we design a mixed IO scheduler (MIOS) to adaptively steer incoming data to SSDs and HDDs. A multi-HDD scheduling is further designed to minimize HDD-write latency. We perform extensive evaluations under production workloads and benchmarks. The results show that MIOS removes up to 93% amount of data written to SSDs, reduces average and 99 th -percentile latencies of the hybrid server by 65% and 85%, respectively.
Shucheng Wang, Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang, Changsheng Xie 0001
ACM Trans. Storage2
2021 A Case Study of Migrating RocksDB on Intel Optane Persistent Memory
abstract
The application of product-level persistent memory (PM) presents a great opportunity for key-value stores. However, PM devices differ significantly from traditional block-based storage devices such as HDD and SSD in terms of IO characteristics and approaches. To reveal the adaptability of existing persistent key-value store on PM and to explore the potential optimization space of PM-based key-value stores, we migrate one of the most widely used persistent key-value store, RocksDB, to PM device and evaluated its performance. The results show that the performance of RocksDB is limited by the traditional IO stacks optimized for fast SSDs on PM devices. We then perform further experimental analysis on the IO methods of the two main files, log and SST, in RocksDB. Based on the results, we propose a set of optimized IO configurations for each of the two files. These configurations improve read and write performance of RocksDB by up to 3× and 2×, respectively, over the default configurations on an Intel Optane Persistent Memory.
Ziyi Lu, Qiang Cao 0001
NAS1
2020 BCW: Buffer-Controlled Writes to HDDs for SSD-HDD Hybrid Storage Server
Shucheng Wang, Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang
FAST2
2020 A Novel Multi-Stage Forest-Based Key-Value Store for Holistic Performance Improvement
abstract
Key-value (KV) stores based on multi-stage structures are widely deployed to organize massive amounts of easily searchable user data. However, current KV storage systems inevitably sacrifice at least one of the performance objectives, such as write, read, space efficiency etc., for the optimization of others. To understand the root cause of and ultimately remove such performance disparities among the representative existing KV stores, we analyze their enabling mechanisms and classify them into two fundamental models of data structures facilitating KV operations, namely, the multi-stage tree (MS-tree), and the multi-stage forest (MS-forest). We build SifrDB, a KV store on a novel split forest structure, that achieves the lowest write amplification across all workload patterns and minimizes space reservation for the compaction. To mitigate the read amplification inherent in MS-forest, we introduce a bloom filer mechanism based on Sorted String Tables (SSTs). Furthermore, we also present a highly efficient parallel search approach that fully exploits the access parallelism of modern flash-based storage devices to substantially boost the read performance. Evaluation results show that under both micro and YCSB benchmarks, SifrDB outperforms its closest competitors, i.e., the popular MS-forest implementations, making it a highly desirable choice for the modern KV stores.
Ziyi Lu, Qiang Cao 0001, Fei Mei, Hong Jiang 0001, Jingjun Li
IEEE Trans. Parallel Distributed Syst.1
2019 Analysis of and Optimization for Write-dominated Hybrid Storage Nodes in Cloud
abstract
Cloud providers like the Alibaba cloud routinely and widely employ hybrid storage nodes composed of solid-state drives (SSDs) and hard disk drives (HDDs), reaping their respective benefits: performance from SSD and capacity from HDD. These hybrid storage nodes generally write incoming data to its SSDs and then flush them to their HDD counterparts, referred to as the SSD Write Back (SWB) mode, thereby ensuring low write latency. When comprehensively analyzing real production workloads from Pangu, a large-scale storage platform underlying the Alibaba cloud, we find that (1) there exist many write dominated storage nodes (WSNs); however, (2) under the SWB mode, the SSDs of these WSNs suffer from severely high write intensity and long tail latency. To address these unique observed problems of WSNs, we present SSD Write Redirect (SWR), a runtime IO scheduling mechanism for WSNs. SWR judiciously and selectively forwards some or all SSD-writes to HDDs, adapting to runtime conditions. By effectively offloading the right amount of write IOs from overburdened SSDs to underutilized HDDs in WSNs, SWR is able to adequately alleviate the aforementioned problems suffered by WSNs. This significantly improves overall system performance and SSD endurance. Our trace-driven evaluation of SWR, through replaying production workload traces collected from the Alibaba cloud in our cloud testbed, shows that SWR decreases the average and 99til-percentile latencies of SSD-writes by up to 13% and 47% respectively, notably improving system performance. Meanwhile the amount of data written to SSDs is reduced by up to 70%, significantly improving SSD lifetime.
Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang
SoCC4
2015 Mining network traffic anomaly based on adjustable piecewise entropy
abstract
Today network traffic anomaly detection is very challenging in a big and constantly changing network, because there are millions of flows being transferred in a network at the same time, and the flow numbers change all the time. Although traditional information entropy has been proved to be an effective metric on network traffic anomaly detection, such a metric shows some limitations in large scale networks with constantly changing flow numbers, and it makes the traditional entropy inefficient for traffic anomaly detection. Another challenge is how to process large-scale traffic data in a scalable way. In this paper, we propose Adjustable Piecewise Entropy for traffic anomaly detection, and implement Adjustable Piecewise Shannon entropy in Hadoop platform with a cluster of five servers in Tsinghua University Campus Network. Furthermore, we analyze and validate Adjustable Piecewise Entropy in both mathematics and experiments. The experiment results show that Adjustable Piecewise Entropy has better performance for traffic anomaly detection.
Geng Tian, Xia Yin 0001, Zimu Li, Xingang Shi, Ziyi Lu, Yingya Guo
IWQoS6
2015 TADOOP: Mining Network Traffic Anomalies with Hadoop
Geng Tian, Xia Yin 0001, Zimu Li, Xingang Shi, Ziyi Lu
SecureComm6
2000 Visual Tracking with Subpixel Resolution using an Analog VLSI Computational Sensor
abstract
This paper describes the application of a computational vision sensor to active binocular tracking. The sensor outputs are used to control the vergence angles of the two cameras and the tilt angle of the head so that the center pixels of the sensor arrays image the same point in the environment. One distinguishing feature of the sensor used here is the possibility to resolve target motions with subpixel resolution. This is due to the use of a phase based algorithm which integrates information over multiple pixels.
Ziyi Lu, Bertram E. Shi
ICRA1
2000 Automatic speech recognition in Mandarin for embedded platforms
Fengguang Zhao, Prabhu Raghavan, Sunil K. Gupta, Ziyi Lu, Wentao Gu
INTERSPEECH4
2000 Binocular visual feedback with CNN sensors
abstract
This paper describes the use of a CNN sensor to provide visual feedback signals in an active binocular vision system. The sensor outputs are used to control the vergence angles of the two cameras and the tilt angle of the head so that the center pixels of the sensors image the same point in the environment. The sensors contain 2D arrays of phototransistors for image input and CNN-based processing circuits that convolve the input image with filters similar to even and odd Gabor filters. We use the odd and even filter outputs in a phase based algorithm, which enables target motions to be detected with sub-pixel resolution. Experimental results demonstrating the ability of the system to reconstruct target motions in 3D are described.
Ziyi Lu, Bertram E. Shi
ISCAS1
1997 Synchronization in Time-delayed Binary Oscillatory Network
Ziyi Lu, Luxi Yang, Zhenya He
Neural Process. Lett.1
1996 Synchrony in Binary-Oscillator Networks with Local Couplings
Ziyi Lu, Baoyun Wang, Luxi Yang, Zhenya He
Int. J. Neural Syst.1