EDBT 2026 Demo / reviewers in the wild / expert
Shujie Pang
dblp:231/9111
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-1741-1319ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 11 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AGCB: Adaptive Garbage Collection for Enhancing Lifetime and Performance of Bit-Alterable Flash MemoryabstractBit-alterable flash-based SSDs, offering page-level erase operation, allows individual flash pages in a block to be erased independently. The page-level erase operation alleviates the overhead of page migration during garbage collection and improves the SSD lifetime. However, when the number of invalid pages within a block exceeds a certain threshold, the latency of page-level garbage collections using page-level erase may exceed that of block-level garbage collections. In bit-alterable flash memory, existing garbage collection strategies dynamically choose between page-level and block-level garbage collections based on their latency. This often fails to fully exploit the advantage of page-level garbage collection in reducing write amplification under low-load conditions.To address this limitation, we propose an adaptive garbage collection strategy called AGCB to dynamically adjust garbage collection operations by the runtime workload of flash channels, thereby enhancing SSD performance and lifetime. Specifically, AGCB classifies flash channels as busy or idle by monitoring the depth of the transaction queue in cache. According to this classification, AGCB selectively applies page-level or block-level garbage collection operations, aiming to minimize the impact of garbage collections with host I/O requests. Meanwhile, we introduce a staged victim block selection scheme to further improve garbage collection efficiency and wear leveling. The experimental results unveil that compared with the existing schemes, AGCB reduces the number of garbage collection operations, average response time, and blocked user requests by an average of 14.6%, 14.7%, and 17.3%, respectively. Laifu Zhang, Yuhui Deng 0001, Peng Zhou 0032, Shujie Pang, Zhaorui Wu, Lin Cui 0001, Zhen Zhang 0017 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | RDA: A Read-Request Driven Adaptive Allocation Scheme for Improving SSD PerformanceabstractThe parallel operation technology plays a pivotal role in enhancing performance of 3-D nand flash-based SSDs. High-parallel distribution of consecutive pages places the pages on different parallel units, thereby improving the parallelism and throughput of read requests. However, the high-parallel distribution generates two problems: 1) aggravating data fragmentation and 2) exacerbating the impact of garbage collection (GC) on latency. Moreover, small reads only require a few parallel units, and thus the high-parallel distribution is redundant for the requests. To address this issue, we propose a read-request driven adaptive allocation scheme called RDA to bolster SSD performance by adaptively adjusting the parallel distribution of consecutive pages. The RDA scheme employs the size of historical read requests to gauge the level of parallelism for write requests with varying sizes. Then, RDA allocates the logical pages of writes to distinct parallel units according to the parallelism of the requests. In doing so, RDA effectively mitigates the performance degradation of SSDs caused by redundant parallel distribution, while preserving the parallelism of read requests. We compare RDA with the three state-of-art schemes Amphibian, SOML, and Preemptive GC in terms of GC-blocked read requests, GC counts, and read response time under eight real-world workloads. The experimental results unveil that compared with the existing schemes, RDA revamps the GC-blocked read requests, GC counts, and read response time by averages of 20.6%, 7.8%, and 15.8%, respectively. Shujie Pang, Yuhui Deng 0001, Zhaorui Wu, Genxiong Zhang, Jie Li 0067, Xiao Qin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | DBCGM: A Granular Model for Big Data Classification Based on Data Bisection and Cascade Weighted Clustering
Jiande Huang, Yuhui Deng 0001, Yi Zhou 0009, Shujie Pang, Qifen Yang, Geyong Min |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | An Energy-Aware Virtual Machine Scheduling Approach for Cloud Data CentersabstractThe reduction of energy consumption will be even more urgent in cloud data centers due to the explosive increase of application data. Virtual machine (VM) integration is a relatively standard technology currently applied for computing facilities of data centers. However, excessive VM consolidation can easily lead to local hot spots that lower the energy efficiency and reliability of data centers. In addition, on account of the impact of heat recirculation in data centers, the traditional VM scheduling strategy cannot comprehensively ponder optimizing the holistic data center energy, which encompasses both server energy and cooling energy. To handle these issues, we proposedEAVMS- an Energy-Aware VM Scheduling approach for minimizing the holistic energy consumption of data centers. EAVMS adopts a two-phase approach to gain energy efficiency while guaranteeing QoS. First, EAVMS leverages a Blended Genetic algorithm and Simulated Annealing algorithm (BGSA) to optimize the initial placement of VMs. Second, EAVMS utilizes a dynamic migration algorithm to achieve effective migration by setting a maximum server temperature threshold without violating the service level agreement (SLA) that cuts down energy consumption by moderating the hot spots of servers. We conducted extensive experiments using two real-world traces (i.e., PlanetLab and Google Cluster datasets) to evaluate the effectiveness of EAVMS. The experimental results unveil that our approach is capable of saving 3.23$ \%$–43.07$ \%$in the holistic energy consumption of cloud data centers with only a tiny service performance degradation compared to other state-of-the-art alternatives (e.g., MJPM, GRANITE, TAS, XINT-GA, and Random). Jie Li 0067, Yuhui Deng 0001, Zijie Zhong, Zhaorui Wu, Shujie Pang, Lin Cui 0001, Geyong Min |
IEEE Trans. Sustain. Comput. | 5 |
| 2024 | FaaSBatch: Boosting Serverless Efficiency With In-Container Parallelism and Resource MultiplexingabstractWith high scalability and flexibility, serverless computing is becoming the most promising computing model. Existing serverless computing platforms initiate a container for each function invocation, which leads to a huge waste of computing resources. Our examinations reveal that (i) executing invocations concurrently within a single container can provide comparable performance to that provided by multiple containers (i.e., traditional approaches); (ii) redundant resources generated within a container result in memory resource waste, which prolongs the execution time of function invocations. Motivated by these insightful observations, we propose FaaSBatch - a serverless framework that reduces invocation latency and saves scarce computing resources. In particular, FaaSBatch first classifies concurrent function requests into different function groups according to the invocation information. Next, FaaSBatch batches the invocations of each group, aiming to minimize resource utilization. Then, FaaSBatch utilizes an inline parallel policy to map each group of batched invocations into a single container. Finally, FaaSBatch expands and executes invocations of containers in parallel. To further reduce invocation latency and resource utilization, within each container, FaaSBatch reuses redundant resources created during function execution. We conduct extensive experiments based on Azure traces to evaluate the effectiveness and performance of FaaSBatch. We compare FaaSBatch with three state-of-the-art schedulers Vanilla, SFS, and Kraken. Our experimental results show that FaaSBatch effectively and remarkably slashes invocation latency and resource overhead. For instance, when executing I/O functions, FaaSBatch cuts back the invocation latency of Vanilla, SFS, and Kraken by up to 72.58%, 74.10%, and 72.62%, respectively; FaaSBatch also slashes the resource overhead of Vanilla, SFS, and Kraken by 70.2% to 98.40%, 67.74% to 98.12%, and 43.01% to 78.90%, respectively. Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Jie Li 0067, Shujie Pang, Xiao Qin 0001 |
IEEE Trans. Computers | 5 |
| 2024 | Minato: A Read-Disturb-Aware Dynamic Buffer Management Scheme for NAND Flash MemoryabstractRead-disturb problem plays a pivotal factor in the performance of NAND flash memory, because it deteriorates the read-disturb errors of NAND flash. Although ECC, read retry, and read reclaim technologies are designed to correct read-disturb errors, these techniques drastically increase read latency and degrade read performance. Moreover, modern SSDs implement a buffer in the built-in DRAM to store frequently accessed data, which can cache hot read data to alleviate the read-disturb problem. Unfortunately, the buffer primarily serves write requests to curtail write operations in flash memory, and ignores the ever-increasing requirement from users for read latency. To address this issue, we propose a read-disturb-aware dynamic buffer management scheme called Minato that reduces read-disturb errors with rationally read buffer management, aiming to improve the read performance of SSDs. Minato includes two distinctive and vital features. First, Minato dynamically adjusts the size of the read buffer and write buffer through the hit situation of requests, thus increasing the size of the read buffer while maintaining the write hit of the write buffer. Second, to further reduce read-disturb errors, Minato implements a read buffer filter to preferentially cache hot read data disturbing more valid pages into the read buffer. We compare Minato with two state-of-art schemes -BPLRU and GCaR in terms of write hit ratio, read-disturb counts, and read/write response time. The experimental results derived from nine real-world workload traces show that Minato efficiently alleviates the read-disturb problem of flash memory without affecting the write hit ratio, and significantly improves read/write performance. In particular, compared with the existing schemes, Minato slashes the read/write response time by an average of 34.6%. Shujie Pang, Yuhui Deng 0001, Genxiong Zhang, Jiande Huang, Zhaorui Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | FaaSBatch: Enhancing the Efficiency of Serverless Computing by Batching and Expanding FunctionsabstractWith high scalability and flexibility, serverless computing is becoming the most promising computing model. Existing serverless computing platforms initiate a container for each function invocation, which leads to a huge waste of computing resources. Our examinations reveal that (i) executing invocations concurrently within a single container can provide comparable performance to that provided by multiple containers (i.e., traditional approaches); (ii) redundant resources generated within a container result in memory resource waste, which prolongs the execution time of function invocations. Motivated by these insightful observations, we propose FaaSBatch - a serverless framework that reduces invocation latency and saves scarce computing resources. In particular, FaaSBatch first classifies concurrent function requests into different function groups according to the invocation information. Next, FaaSBatch batches the invocations of each group, aiming to minimize resource utilization. Then, FaaSBatch utilizes an inline parallel policy to map each group of batched invocations into a single container. Finally, FaaSBatch expands and executes invocations of containers in parallel. To further reduce invocation latency and resource utilization, within each container, FaaSBatch reuses redundant resources created during function execution. We conduct extensive experiments based on Azure traces to evaluate the effectiveness and performance of FaaSBatch. We compare FaaSBatch with three state-of-the-art schedulers Vanilla, SFS, and Kraken. Our experimental results show that FaaSBatch effectively and remarkably slashes invocation latency and resource overhead. For instance, when executing I/O functions, FaaSBatch cuts back the invocation latency of Vanilla, SFS, and Kraken by up to 92.18%, 89.54%, and 90.65%, respectively; FaaSBatch also slashes the resource overhead of Vanilla, SFS, and Kraken by 58.89% to 94.77%, 43.72% to 90.39%, and 42.99% to 78.88%, respectively. Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Jie Li 0067, Shujie Pang |
ICDCS | 5 |
| 2023 | FSPDA: A Full Sequence Program Data Allocation Scheme for Boosting 3-D nand Flash Read PerformanceabstractMultibit 3-D NAND flash-based solid-state disks (SSDs), offering high storage density, contain multiple types of pages to accommodate multiple bits per physical cell. Full sequence program or FSP can program multiple pages in a word line at a time, thereby improving write throughput. Unfortunately, large-grained FSP operations coarsely aggregate consecutive logical pages on the same word line, which adversely affects the parallelism and latency of read requests. Moreover, FSP smooths the program latencies for different types of pages, whereas the pages still exhibit various read latencies. Multiple read latencies and lower read parallelism noticeably deteriorate the completion efficiency of read requests: SSD performance is degraded. To address this issue, we propose an FSP data allocation scheme called FSPDA that incorporates the physical structure characteristics of multibit 3-D NAND, aiming to bolster the read performance of 3-D NAND Flash-based SSDs. FSPDA embraces two distinctive and vital features. First, according to the distance between logical pages, FSPDA allocates logical pages to specified parallel units and stipulates that consecutive logical pages must be assigned to different planes, thus improving read parallelism and data locality. Second, to further reduce read latency, FSPDA employs cache hits to determine hot and cold data to be placed to low-latency and high-latency pages, respectively. We compare FSPDA with two state-of-the-art schemes—OSPADA and single-operation-multiple-location—in terms of multiplane read (MPR) counts, read response time, and GC counts under eight real-world workloads. The experimental results show that compared with the existing schemes, FSPDA slashes the number of MPR counts, read response time, and the number of GC counts by an average of 34.4%, 28.5%, and 13.6%, respectively. Shujie Pang, Yuhui Deng 0001, Zhaorui Wu, Genxiong Zhang, Jie Li 0067, Xiao Qin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | PcGC: A Parity-Check Garbage Collection for Boosting 3-D NAND Flash PerformanceabstractGarbage collection or GC running in the controller of 3-D NAND flash-based solid-state disks—SSDs—plays a critical role in the performance of storage systems. SSD manufacturers have developed various GC solutions based on internal data movement or IDM to mitigate the impacts of GC on request latency. Due to the circuit characteristics of flash memory, the existing IDM-based GC strategies are restricted by page parity during data movement: odd pages must be migrated to odd pages, and even pages to even pages. When migrating two consecutive pages with the same parity, the free page between the two migrated pages will be wasted after the migration is complete. This ever-increasing page waste problem inevitably deteriorates the storage space utilization of flash memory, thereby degrading the overall performance of 3-D NAND flash-based SSDs. To address this issue, we propose a parity-check GC scheme called PcGC to revamp SSD performance by alleviating page waste during GC. We build a parity-check unit in PcGC to facilitate checking the parity of migrated valid pages and destination pages. According to the parity results offered by the parity-check unit, PcGC dynamically adjusts the migration order of valid pages during the course of GC. In doing so, PcGC fundamentally averts page waste caused by the page parity restriction, thereby enhancing 3-D NAND flash performance. We quantitatively evaluate the performance of PcGC in terms of wasted pages, storage utilization, GC counts, write amplification, and average response time. We compare PcGC against the two state-of-the-art schemes—Amphibian and Tiny-tail flash (TTflash). The experimental results derived from the nine real-world workload traces unfold that compared with Amphibian and TTflash: 1) PcGC curtails the number of wasted pages by up to 91.4% with an average of 53.75%; 2) cuts back the number of GC counts by up to 52.2% with an average of 11.9%; and 3) slashes average write response time by up to 77.8% with an average of 13.0%. Shujie Pang, Yuhui Deng 0001, Genxiong Zhang, Yi Zhou 0009, Xiao Qin 0001, Zhaorui Wu, Jie Li 0067 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Cocktail: Mixing Data With Different Characteristics to Reduce Read Reclaims for nand Flash MemoryabstractA large number of read-disturb-induced rewrites are performed in the background [also known as Read Reclaim (RR)] to alleviate the read-disturb issue in NAND flash memory-based SSDs. RR can significantly degrade the performance and shorten the service life of SSD in read-intensive workloads. To address this issue, we propose a novel read-disturb management approach called Cocktail that mixes a small proportion of hot-read pages with a large proportion of cold-read pages, thereby avoiding clustering hot-read pages into a few blocks. Motivated by the insight that RR operations are frequently triggered by hot read-pages, Cocktail first prefills a portion of each block with cold data extracted from user requests. Then, Cocktail fills the prefilled blocks with write-back data caused by RR to create read-balanced blocks. We integrate two thresholds, write pool capacity and the ratio of RR-write data to User-write data, into Cocktail to govern the ratio of write-back data caused by RR to data of user requests in a block. Cocktail dynamically adjusts the two thresholds according to the characteristics of RR. Cocktail is conducive to decentralizing hot write-back data caused by RR across a broad range of blocks, thereby reducing the occurrence of second-time RR and the number of overall block reads. We compare Cocktail with three existing schemes baseline, redFTL, and IPR in terms of SSD service life, SSD response time, write amplification, and the number of garbage collections (GCs) under ten real-world workload conditions. Experimental results show that compared with the existing schemes, Cocktail reduces the number of RRs, the average response time, the 99-percentile tail latency, and the number of GCs by an average of 40.77%, 10.82%, 5.40%, and 12.29%, respectively. Cocktail also alleviates the write amplification of the three alternative schemes by an average of 49.57%. Genxiong Zhang, Yuhui Deng 0001, Yi Zhou 0009, Shujie Pang, Jianhui Yue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | TADRP: Toward Thermal-Aware Data Replica Placement in Data-Intensive Data CentersabstractWith the mushrooming growth of data volumes, data replica placement plays a key role in promoting the energy efficiency and Quality-of-Service (QoS) of data-intensive data centers. The existing data placement strategies mainly focus on storage performance improvement or QoS enhancement in data centers, but ignore the indispensable factor - heat recirculation. To bridge this gap, we propose a thermal-aware data replica placement strategy called TADRP, aiming to improve cooling efficiency and minimize the total power consumption of data-intensive data centers. TADRP leverages an ant colony optimization (ACO) algorithm coupled with Laplacian probability distribution to find a quasi-optimal disk sequence (or Disk Sequence for short), which consists of disks selected from different rack servers to place data replicas. TADRP categorizes disks of Disk Sequence into active and inactive ones, by placing hot and cold replicas on active and inactive disks, respectively. We quantitatively evaluate TADRP in terms of cooling costs, total power consumption, number of power-state transitions, and execution time. We compare TADRP with four alternative solutions, namely, Random, Hadoop, SRS, and CDP-NSGAII. Experimental results show that TADRP can reduce the cooling costs and the total power of the existing solutions by 14.7% - 61.7% and 19.2%-55.1%, respectively, without undesirable I/O performance drops. Jie Li 0067, Yuhui Deng 0001, Yi Zhou 0009, Zhaorui Wu, Shujie Pang, Geyong Min |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2023 | PSA-Cache: A Page-state-aware Cache Scheme for Boosting 3D NAND Flash PerformanceabstractGarbage collection (GC) plays a pivotal role in the performance of 3D NAND flash memory, where Copyback has been widely used to accelerate valid page migration during GC. Unfortunately, copyback is constrained by the parity symmetry issue: data read from an odd/even page must be written to an odd/even page. After migrating two odd/even consecutive pages, a free page between the two migrated pages will be wasted. Such wasted pages noticeably lower free space on flash memory and cause extra GCs, thereby degrading solid-state-disk (SSD) performance. To address this problem, we propose a page-state-aware cache scheme called PSA-Cache , which prevents page waste to boost the performance of NAND Flash-based SSDs. To facilitate making write-back scheduling decisions, PSA-Cache regulates write-back priorities for cached pages according to the state of pages in victim blocks. With high write-back-priority pages written back to flash chips, PSA-Cache effectively fends off page waste by breaking odd/even consecutive pages in subsequent garbage collections. We quantitatively evaluate the performance of PSA-Cache in terms of the number of wasted pages, the number of GCs, and response time. We compare PSA-Cache with two state-of-the-art schemes, GCaR and TTflash, in addition to a baseline scheme LRU. The experimental results unveil that PSA-Cache outperforms the existing schemes. In particular, PSA-Cache curtails the number of wasted pages of GCaR and TTflash by 25.7% and 62.1%, respectively. PSA-Cache immensely cuts back the number of GC counts by up to 78.7% with an average of 49.6%. Furthermore, PSA-Cache slashes the average write response time by up to 85.4% with an average of 30.05%. Shujie Pang, Yuhui Deng 0001, Genxiong Zhang, Yi Zhou 0009, Yaoqin Huang, Xiao Qin 0001 |
ACM Trans. Storage | 1 |
| 2022 | A Thermal-Aware Data Replica Placement Strategy for Data-Intensive Data CentersabstractIn this paper, we propose a thermal-aware data replica placement strategy called TADRP. This strategy is designed in two steps. First, the ant colony optimization (ACO) algorithm based on the Laplacian probability distribution obtains the near-optimal disk sequence with the minimum overall power consumption. Second, the near-optimal disk sequence is partitioned into the area of active and inactive disks; then, the sequence-based data placement policy places data replicas in the partitioned disk areas. Our objection is to adopt the TADRP strategy to improve cooling efficiency and reduce the overall power consumption of DDCs. To evaluate the overall power consumption of DDCs, we integrate TADRP into a thermal model that takes into account heat recirculation effects account. We apply a real dataset with different read/write ratios following the Zipf distribution to verify the effectiveness of TADRP for energy savings. Experiment results unveil that our TADRP is capable of offering about 19.2 %-55.1% for total energy savings without significantly degrading I/O performance against Random, Hadoop, SRS, CDP-NSGAIIIR schemes. Jie Li 0067, Yuhui Deng 0001, Zhaorui Wu, Shujie Pang |
PACT | 4 |