VLDB 2026 Research / reviewers in the wild / expert
Zhaorui Wu
dblp:297/4813
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-8062-765XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 12 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AGCB: Adaptive Garbage Collection for Enhancing Lifetime and Performance of Bit-Alterable Flash MemoryabstractBit-alterable flash-based SSDs, offering page-level erase operation, allows individual flash pages in a block to be erased independently. The page-level erase operation alleviates the overhead of page migration during garbage collection and improves the SSD lifetime. However, when the number of invalid pages within a block exceeds a certain threshold, the latency of page-level garbage collections using page-level erase may exceed that of block-level garbage collections. In bit-alterable flash memory, existing garbage collection strategies dynamically choose between page-level and block-level garbage collections based on their latency. This often fails to fully exploit the advantage of page-level garbage collection in reducing write amplification under low-load conditions.To address this limitation, we propose an adaptive garbage collection strategy called AGCB to dynamically adjust garbage collection operations by the runtime workload of flash channels, thereby enhancing SSD performance and lifetime. Specifically, AGCB classifies flash channels as busy or idle by monitoring the depth of the transaction queue in cache. According to this classification, AGCB selectively applies page-level or block-level garbage collection operations, aiming to minimize the impact of garbage collections with host I/O requests. Meanwhile, we introduce a staged victim block selection scheme to further improve garbage collection efficiency and wear leveling. The experimental results unveil that compared with the existing schemes, AGCB reduces the number of garbage collection operations, average response time, and blocked user requests by an average of 14.6%, 14.7%, and 17.3%, respectively. Laifu Zhang, Yuhui Deng 0001, Peng Zhou 0032, Shujie Pang, Zhaorui Wu, Lin Cui 0001, Zhen Zhang 0017 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | Chrono: Efficient Serverless Analytics With Adaptive Fine-Grained Partitioning and Shadow Execution
Zhaorui Wu, Yuhui Deng 0001, Jiande Huang, Qifen Yang, Peng Zhou 0032, Geyong Min |
IEEE Trans. Cloud Comput. | 1 |
| 2025 | RDA: A Read-Request Driven Adaptive Allocation Scheme for Improving SSD PerformanceabstractThe parallel operation technology plays a pivotal role in enhancing performance of 3-D nand flash-based SSDs. High-parallel distribution of consecutive pages places the pages on different parallel units, thereby improving the parallelism and throughput of read requests. However, the high-parallel distribution generates two problems: 1) aggravating data fragmentation and 2) exacerbating the impact of garbage collection (GC) on latency. Moreover, small reads only require a few parallel units, and thus the high-parallel distribution is redundant for the requests. To address this issue, we propose a read-request driven adaptive allocation scheme called RDA to bolster SSD performance by adaptively adjusting the parallel distribution of consecutive pages. The RDA scheme employs the size of historical read requests to gauge the level of parallelism for write requests with varying sizes. Then, RDA allocates the logical pages of writes to distinct parallel units according to the parallelism of the requests. In doing so, RDA effectively mitigates the performance degradation of SSDs caused by redundant parallel distribution, while preserving the parallelism of read requests. We compare RDA with the three state-of-art schemes Amphibian, SOML, and Preemptive GC in terms of GC-blocked read requests, GC counts, and read response time under eight real-world workloads. The experimental results unveil that compared with the existing schemes, RDA revamps the GC-blocked read requests, GC counts, and read response time by averages of 20.6%, 7.8%, and 15.8%, respectively. Shujie Pang, Yuhui Deng 0001, Zhaorui Wu, Genxiong Zhang, Jie Li 0067, Xiao Qin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | An Energy-Aware Virtual Machine Scheduling Approach for Cloud Data CentersabstractThe reduction of energy consumption will be even more urgent in cloud data centers due to the explosive increase of application data. Virtual machine (VM) integration is a relatively standard technology currently applied for computing facilities of data centers. However, excessive VM consolidation can easily lead to local hot spots that lower the energy efficiency and reliability of data centers. In addition, on account of the impact of heat recirculation in data centers, the traditional VM scheduling strategy cannot comprehensively ponder optimizing the holistic data center energy, which encompasses both server energy and cooling energy. To handle these issues, we proposedEAVMS- an Energy-Aware VM Scheduling approach for minimizing the holistic energy consumption of data centers. EAVMS adopts a two-phase approach to gain energy efficiency while guaranteeing QoS. First, EAVMS leverages a Blended Genetic algorithm and Simulated Annealing algorithm (BGSA) to optimize the initial placement of VMs. Second, EAVMS utilizes a dynamic migration algorithm to achieve effective migration by setting a maximum server temperature threshold without violating the service level agreement (SLA) that cuts down energy consumption by moderating the hot spots of servers. We conducted extensive experiments using two real-world traces (i.e., PlanetLab and Google Cluster datasets) to evaluate the effectiveness of EAVMS. The experimental results unveil that our approach is capable of saving 3.23$ \%$–43.07$ \%$in the holistic energy consumption of cloud data centers with only a tiny service performance degradation compared to other state-of-the-art alternatives (e.g., MJPM, GRANITE, TAS, XINT-GA, and Random). Jie Li 0067, Yuhui Deng 0001, Zijie Zhong, Zhaorui Wu, Shujie Pang, Lin Cui 0001, Geyong Min |
IEEE Trans. Sustain. Comput. | 4 |
| 2024 | FaaSBatch: Boosting Serverless Efficiency With In-Container Parallelism and Resource MultiplexingabstractWith high scalability and flexibility, serverless computing is becoming the most promising computing model. Existing serverless computing platforms initiate a container for each function invocation, which leads to a huge waste of computing resources. Our examinations reveal that (i) executing invocations concurrently within a single container can provide comparable performance to that provided by multiple containers (i.e., traditional approaches); (ii) redundant resources generated within a container result in memory resource waste, which prolongs the execution time of function invocations. Motivated by these insightful observations, we propose FaaSBatch - a serverless framework that reduces invocation latency and saves scarce computing resources. In particular, FaaSBatch first classifies concurrent function requests into different function groups according to the invocation information. Next, FaaSBatch batches the invocations of each group, aiming to minimize resource utilization. Then, FaaSBatch utilizes an inline parallel policy to map each group of batched invocations into a single container. Finally, FaaSBatch expands and executes invocations of containers in parallel. To further reduce invocation latency and resource utilization, within each container, FaaSBatch reuses redundant resources created during function execution. We conduct extensive experiments based on Azure traces to evaluate the effectiveness and performance of FaaSBatch. We compare FaaSBatch with three state-of-the-art schedulers Vanilla, SFS, and Kraken. Our experimental results show that FaaSBatch effectively and remarkably slashes invocation latency and resource overhead. For instance, when executing I/O functions, FaaSBatch cuts back the invocation latency of Vanilla, SFS, and Kraken by up to 72.58%, 74.10%, and 72.62%, respectively; FaaSBatch also slashes the resource overhead of Vanilla, SFS, and Kraken by 70.2% to 98.40%, 67.74% to 98.12%, and 43.01% to 78.90%, respectively. Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Jie Li 0067, Shujie Pang, Xiao Qin 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Minato: A Read-Disturb-Aware Dynamic Buffer Management Scheme for NAND Flash MemoryabstractRead-disturb problem plays a pivotal factor in the performance of NAND flash memory, because it deteriorates the read-disturb errors of NAND flash. Although ECC, read retry, and read reclaim technologies are designed to correct read-disturb errors, these techniques drastically increase read latency and degrade read performance. Moreover, modern SSDs implement a buffer in the built-in DRAM to store frequently accessed data, which can cache hot read data to alleviate the read-disturb problem. Unfortunately, the buffer primarily serves write requests to curtail write operations in flash memory, and ignores the ever-increasing requirement from users for read latency. To address this issue, we propose a read-disturb-aware dynamic buffer management scheme called Minato that reduces read-disturb errors with rationally read buffer management, aiming to improve the read performance of SSDs. Minato includes two distinctive and vital features. First, Minato dynamically adjusts the size of the read buffer and write buffer through the hit situation of requests, thus increasing the size of the read buffer while maintaining the write hit of the write buffer. Second, to further reduce read-disturb errors, Minato implements a read buffer filter to preferentially cache hot read data disturbing more valid pages into the read buffer. We compare Minato with two state-of-art schemes -BPLRU and GCaR in terms of write hit ratio, read-disturb counts, and read/write response time. The experimental results derived from nine real-world workload traces show that Minato efficiently alleviates the read-disturb problem of flash memory without affecting the write hit ratio, and significantly improves read/write performance. In particular, compared with the existing schemes, Minato slashes the read/write response time by an average of 34.6%. Shujie Pang, Yuhui Deng 0001, Genxiong Zhang, Jiande Huang, Zhaorui Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | FaaSBatch: Enhancing the Efficiency of Serverless Computing by Batching and Expanding FunctionsabstractWith high scalability and flexibility, serverless computing is becoming the most promising computing model. Existing serverless computing platforms initiate a container for each function invocation, which leads to a huge waste of computing resources. Our examinations reveal that (i) executing invocations concurrently within a single container can provide comparable performance to that provided by multiple containers (i.e., traditional approaches); (ii) redundant resources generated within a container result in memory resource waste, which prolongs the execution time of function invocations. Motivated by these insightful observations, we propose FaaSBatch - a serverless framework that reduces invocation latency and saves scarce computing resources. In particular, FaaSBatch first classifies concurrent function requests into different function groups according to the invocation information. Next, FaaSBatch batches the invocations of each group, aiming to minimize resource utilization. Then, FaaSBatch utilizes an inline parallel policy to map each group of batched invocations into a single container. Finally, FaaSBatch expands and executes invocations of containers in parallel. To further reduce invocation latency and resource utilization, within each container, FaaSBatch reuses redundant resources created during function execution. We conduct extensive experiments based on Azure traces to evaluate the effectiveness and performance of FaaSBatch. We compare FaaSBatch with three state-of-the-art schedulers Vanilla, SFS, and Kraken. Our experimental results show that FaaSBatch effectively and remarkably slashes invocation latency and resource overhead. For instance, when executing I/O functions, FaaSBatch cuts back the invocation latency of Vanilla, SFS, and Kraken by up to 92.18%, 89.54%, and 90.65%, respectively; FaaSBatch also slashes the resource overhead of Vanilla, SFS, and Kraken by 58.89% to 94.77%, 43.72% to 90.39%, and 42.99% to 78.88%, respectively. Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Jie Li 0067, Shujie Pang |
ICDCS | 1 |
| 2023 | FSPDA: A Full Sequence Program Data Allocation Scheme for Boosting 3-D nand Flash Read PerformanceabstractMultibit 3-D NAND flash-based solid-state disks (SSDs), offering high storage density, contain multiple types of pages to accommodate multiple bits per physical cell. Full sequence program or FSP can program multiple pages in a word line at a time, thereby improving write throughput. Unfortunately, large-grained FSP operations coarsely aggregate consecutive logical pages on the same word line, which adversely affects the parallelism and latency of read requests. Moreover, FSP smooths the program latencies for different types of pages, whereas the pages still exhibit various read latencies. Multiple read latencies and lower read parallelism noticeably deteriorate the completion efficiency of read requests: SSD performance is degraded. To address this issue, we propose an FSP data allocation scheme called FSPDA that incorporates the physical structure characteristics of multibit 3-D NAND, aiming to bolster the read performance of 3-D NAND Flash-based SSDs. FSPDA embraces two distinctive and vital features. First, according to the distance between logical pages, FSPDA allocates logical pages to specified parallel units and stipulates that consecutive logical pages must be assigned to different planes, thus improving read parallelism and data locality. Second, to further reduce read latency, FSPDA employs cache hits to determine hot and cold data to be placed to low-latency and high-latency pages, respectively. We compare FSPDA with two state-of-the-art schemes—OSPADA and single-operation-multiple-location—in terms of multiplane read (MPR) counts, read response time, and GC counts under eight real-world workloads. The experimental results show that compared with the existing schemes, FSPDA slashes the number of MPR counts, read response time, and the number of GC counts by an average of 34.4%, 28.5%, and 13.6%, respectively. Shujie Pang, Yuhui Deng 0001, Zhaorui Wu, Genxiong Zhang, Jie Li 0067, Xiao Qin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | PcGC: A Parity-Check Garbage Collection for Boosting 3-D NAND Flash PerformanceabstractGarbage collection or GC running in the controller of 3-D NAND flash-based solid-state disks—SSDs—plays a critical role in the performance of storage systems. SSD manufacturers have developed various GC solutions based on internal data movement or IDM to mitigate the impacts of GC on request latency. Due to the circuit characteristics of flash memory, the existing IDM-based GC strategies are restricted by page parity during data movement: odd pages must be migrated to odd pages, and even pages to even pages. When migrating two consecutive pages with the same parity, the free page between the two migrated pages will be wasted after the migration is complete. This ever-increasing page waste problem inevitably deteriorates the storage space utilization of flash memory, thereby degrading the overall performance of 3-D NAND flash-based SSDs. To address this issue, we propose a parity-check GC scheme called PcGC to revamp SSD performance by alleviating page waste during GC. We build a parity-check unit in PcGC to facilitate checking the parity of migrated valid pages and destination pages. According to the parity results offered by the parity-check unit, PcGC dynamically adjusts the migration order of valid pages during the course of GC. In doing so, PcGC fundamentally averts page waste caused by the page parity restriction, thereby enhancing 3-D NAND flash performance. We quantitatively evaluate the performance of PcGC in terms of wasted pages, storage utilization, GC counts, write amplification, and average response time. We compare PcGC against the two state-of-the-art schemes—Amphibian and Tiny-tail flash (TTflash). The experimental results derived from the nine real-world workload traces unfold that compared with Amphibian and TTflash: 1) PcGC curtails the number of wasted pages by up to 91.4% with an average of 53.75%; 2) cuts back the number of GC counts by up to 52.2% with an average of 11.9%; and 3) slashes average write response time by up to 77.8% with an average of 13.0%. Shujie Pang, Yuhui Deng 0001, Genxiong Zhang, Yi Zhou 0009, Xiao Qin 0001, Zhaorui Wu, Jie Li 0067 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | TADRP: Toward Thermal-Aware Data Replica Placement in Data-Intensive Data CentersabstractWith the mushrooming growth of data volumes, data replica placement plays a key role in promoting the energy efficiency and Quality-of-Service (QoS) of data-intensive data centers. The existing data placement strategies mainly focus on storage performance improvement or QoS enhancement in data centers, but ignore the indispensable factor - heat recirculation. To bridge this gap, we propose a thermal-aware data replica placement strategy called TADRP, aiming to improve cooling efficiency and minimize the total power consumption of data-intensive data centers. TADRP leverages an ant colony optimization (ACO) algorithm coupled with Laplacian probability distribution to find a quasi-optimal disk sequence (or Disk Sequence for short), which consists of disks selected from different rack servers to place data replicas. TADRP categorizes disks of Disk Sequence into active and inactive ones, by placing hot and cold replicas on active and inactive disks, respectively. We quantitatively evaluate TADRP in terms of cooling costs, total power consumption, number of power-state transitions, and execution time. We compare TADRP with four alternative solutions, namely, Random, Hadoop, SRS, and CDP-NSGAII. Experimental results show that TADRP can reduce the cooling costs and the total power of the existing solutions by 14.7% - 61.7% and 19.2%-55.1%, respectively, without undesirable I/O performance drops. Jie Li 0067, Yuhui Deng 0001, Yi Zhou 0009, Zhaorui Wu, Shujie Pang, Geyong Min |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2023 | HashCache: Accelerating Serverless Computing by Skipping Duplicated Function ExecutionabstractServerless computing is a leading force behind deploying and managing software in cloud computing. One inherent challenge in serverless computing is the increased overall latency due to duplicate computations. Our initial investigation into the function invocations of serverless applications reveals an abundance of duplicate invocations. Inspired by this critical observation, we introduceHashCache, a system designed to cache duplicate function invocations, thereby mitigating duplicate computations. In HashCache, serverless functions are classified into three categories, namely, computational functions, stateful functions, and environment-related functions. On the grounds of such a function classification, HashCache associates the stateful functions and their states to build an adaptive synchronization mechanism. With this support, HashCache exploits the cached results of computational and stateful functions to serve upcoming invocation requests to the same functions, thereby reducing duplicate computations. Moreover, HashCache stores remote files probed by stateful functions into a local cache layer, which further curtails invocation latency. We implement HashCache within theApache OpenWhiskto forge a cache-enabled serverless computing platform. We conduct extensive experiments to quantitatively evaluate the performance of HashCache in terms of invocation latency and resource utilization. We compare HashCache against two state-of-the-art approaches -FaaSCacheandOpenWhisk. The experimental results unveil that our HashCache remarkably reduces invocation latency and resource overhead. More specifically, HashCache curbs the 99-tail latency of FaaSCache and OpenWhisk by up to 91.37% and 95.96% in real-world serverless applications. HashCache also slashes the resource utilization of FaaSCache and OpenWhisk by up to 31.62% and 35.51%, respectively. Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Lin Cui 0001, Xiao Qin 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | A Thermal-Aware Data Replica Placement Strategy for Data-Intensive Data CentersabstractIn this paper, we propose a thermal-aware data replica placement strategy called TADRP. This strategy is designed in two steps. First, the ant colony optimization (ACO) algorithm based on the Laplacian probability distribution obtains the near-optimal disk sequence with the minimum overall power consumption. Second, the near-optimal disk sequence is partitioned into the area of active and inactive disks; then, the sequence-based data placement policy places data replicas in the partitioned disk areas. Our objection is to adopt the TADRP strategy to improve cooling efficiency and reduce the overall power consumption of DDCs. To evaluate the overall power consumption of DDCs, we integrate TADRP into a thermal model that takes into account heat recirculation effects account. We apply a real dataset with different read/write ratios following the Zipf distribution to verify the effectiveness of TADRP for energy savings. Experiment results unveil that our TADRP is capable of offering about 19.2 %-55.1% for total energy savings without significantly degrading I/O performance against Random, Hadoop, SRS, CDP-NSGAIIIR schemes. Jie Li 0067, Yuhui Deng 0001, Zhaorui Wu, Shujie Pang |
PACT | 3 |
| 2022 | Blender: A Container Placement Strategy by Leveraging Zipf-Like Distribution Within Containerized Data CentersabstractInstantiated containers of an application are distributed across multiple Physical Machines (PMs) to achieve high parallel performance. Container placement plays a vital role in network traffic and the performance of containerized data centers. Existing container placement techniques are inadequate due to the ignorance of container traffic patterns. To solve this issue, we first investigate the network traffic between containers and observe that it exhibits a Zipf-like distribution. Motivated by this finding, we propose a novel container placement approach-Blender-by taking into account the Zipf-like distribution. Blender employs two algorithms calledRefineAlgandSplitAlgto divide containers of applications into blocks, and place these blocks across Virtual Machines (VMs). Blender exhibits two salient features: (i) it minimizes inter-block traffic by arranging the containers that communicate frequently in the same block. (ii) it achieves good load balancing by combining complementary blocks that request different resource types (e.g.,CPU-intensiveandmemory-intensiveblocks) and distributing these blocks across multiple VMs. The experimental results show that Blender significantly reduces communication traffic and network latency. In particular, Blender reduces the traffic of SBP and CA-WFD by 22% and 32%, respectively. Blender decreases network latency by 16% and 26% compared to SBP and CA-WFD. Furthermore, with Blender in place, the physical resources of hosting PMs are well balanced and utilized. Zhaorui Wu, Yuhui Deng 0001, Hao Feng 0010, Yi Zhou 0009, Geyong Min, Zhen Zhang 0017 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Blender: A Traffic-Aware Container Placement for Containerized Data CentersabstractInstantiated containers of an application are distributed across multiple Physical Machines (PMs) to achieve high parallel performance. Container placement plays a vital role in network traffic and the performance of containerized data centers. Existing container placement techniques do not consider the container traffic pattern, which is inadequate. To resolve this conflict, we investigate network traffic between containers and observe that it exhibits a Zipf-like distribution. We propose a novel container placement approach - Blender - by leveraging the Zipf-like distribution. Based on network traffic correlation, Blender employs RefineAlg and SplitAlg to divide containers of applications into blocks, and place these blocks across virtual machines. Blender exhibits two salient features: (i) it minimizes inter-block traffic by arranging the containers that communicate frequently in the same block. (ii) it achieves good load balancing by combining blocks according to the resource types they require and distributing them across multiple PMs. We compare Blender against two state-of-the-art methods SBP and CA-WFD. The experimental results show that Blender significantly reduces communication traffic. In particular, for the same number of PMs, Blender reduces the traffic of SBP and CA-WFD by 22% and 32%, respectively. Furthermore, with Blender in place, the physical resources of hosting PMs are well balanced and utilized. Zhaorui Wu, Yuhui Deng 0001, Hao Feng 0010, Yi Zhou 0009, Geyong Min |
DATE | 1 |