Yi Zhou 0009

dblp:01/1901-9 · DBLP profile ↗
← Back
38ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0002-1460-322XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 5 first-author · 16 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 eCache: A Sample-Inference-Based Intelligent Cache Scheme for High-Performance SSDs
abstract
DRAM-based cache is a practical approach to enhancing the performance of large-capacity SSDs. Due to DRAM’s limited capacity, cache sizes are significantly smaller than the data scale of workloads – and cache replacement schemes determine the cache hit ratio and SSD performance when a cache reaches full capacity. Prior work embarked on leveraging machine learning models (ML) to predict future access patterns in workloads, aiding cache replacement decisions for intelligent cache schemes inside SSDs. Unfortunately, existing ML-based cache schemes often overlook the impacts of data granularity – including page, request, and coarse granularities – of training datasets, model inference time overhead, and computational overhead on SSD performance. To address these challenges, we are motivated to propose an sample-inference-based intelligentcachescheme – eCache. eCache accurately predicts the future reuse distance of each sampled requested pages, which can enhance the accuracy of decision-making for the cache replacement. In particular, we design a parallel framework to curb the time overhead caused by model inference. We advocate for a random-group sampling inference method that utilizes the most accurate model while reducing computational overhead. Moreover, we implement eCache on the state-of-the-art SSD simulator, MQSim, and compare it against alternative cache schemes (i.e., CCache, NCache, LAC, and VS-batch). The experimental results unveil that compared with the other cache schemes, eCache significantly reduces the average response time by up to 79.68% with an average reduction of 44.08%. When compared with the page-granularity-ML-empowered cache schemes, eCache greatly curtails computational overhead by up to 85.23% with an average reduction of 66.00%.
Hui Sun 0002, Yinan Fu, Yi Zhou 0009, Xiao Qin 0001
IEEE Trans. Computers4
2026 Analyzing Request Volatility in Cloud-Based Machine Learning: Insights From Alibaba's Machine Learning as a Service Platform
abstract
With advancements in machine learning (ML) technology and the deployment of large ML-as-a-Service (MLaaS) clouds, accurately understanding request behaviors in an MLaaS cloud platform is paramount for resource scheduling and optimization. This paper sheds light on the correlation of request arrivals in a representative and dynamic MLaaS workload – Alibaba PAI (an ML platform for artificial intelligence). For requests in the PAI workloads at the job, task, instance, and machine levels, our burstiness diagnosis reveals that the request arrival processes at all levels are significantly bursty. Additionally, our Gaussianity test indicates that the bursty activities in PAI consistently appear to be non-Gaussian. Our findings show that there exists a certain degree of correlation between request arrivals at each level over long-term time scales. Moreover, we reveal the self-similar nature of request activities in the various-level wild MLaaS workloads on Alibaba PAI through visual evidence, the auto-correlation structure of the aggregated process of request sequences, and Hurst parameter estimates. Furthermore, we implement a versatile workload synthetic model to synthesize request series based on the inputs measured from the PAI trace. Experimental results demonstrate that our model outperforms typical self-similar workload models, and can improve accuracy by up to 99% compared to them.
Qiang Zou 0005, Yuhui Deng 0001, Yi Zhou 0009, Jianghe Cai, Shuibing He, Lina Ge
IEEE Trans. Netw. Serv. Manag.4
2025 Causal Pathway-Integrated Generative Adversarial Networks for Counterfactually Fair Data Generation
Haoming Mo, Yuhui Deng 0001, Qifen Yang, Jiande Huang, Yi Zhou 0009
ICIC (10)5
2025 A+Store: An Asynchronous Parallel Compaction for Multi-NDP-Enabled Key-Value Store
Hui Sun 0002, Xiaole Liu, Yi Zhou 0009, Yinliang Yue, Xiao Qin 0001
J. Syst. Archit.6
2025 MAFRO: Optimal-Granularity Fuzzy Decision Rule-Based Classification Architecture for Attribute Unlearning
abstract
Recently, many laws and regulations have granted users the right to be forgotten, i.e., the right to require data controllers to delete user data. Various methods for machine unlearning have been proposed to remove individual data points. However, they do not scale to the scenarios where larger groups of features are to be removed. To address this challenge, we propose MAFRO, an optimal-granularity fuzzy decision rule–based classifier that accelerates unlearning via influence functions. Building on granular computing (GrC), MAFRO first selects a minimal reduct of attributes, then constructs fuzzy granules with a Gaussian membership function to extract concise decision rules and realizes unlearning through the influence function. Specifically, instead of training with the full set of attributes, we use the reduct, a minimal subset of attributes that can classify the data with the same accuracy as the full set of attributes. Next, we extract fuzzy rules based on the reduct. Finally, fusing the generated rules establishes the linear model with strongly convex loss functions. In this way, MAFRO can quantify the divergence caused by attribute deleting and update the model without retraining it, thereby adapting the influence of data removal on the model and accelerating the unlearning process. We conduct extensive experiments to evaluate MAFRO on ten typical datasets in terms of performance and unlearning speed. We compare MAFRO with the state-of-theart algorithms. Experimental results demonstrate that MAFRO enhances accuracy by an average of 6.96%, and achieves up to 236× speedup for attribute unlearning tasks.
Jiande Huang, Yuhui Deng 0001, Yi Zhou 0009, Qifen Yang, Geyong Min
IEEE Trans. Fuzzy Syst.3
2025 GPDet: an anchor-free object detector based on dual center-ness and criss-cross balance for unstructured gastroscopic image data
Yuhui Deng 0001, Yi Zhou 0009, Hexian Lu, Lijuan Lu, Shun Long
J. Supercomput.3
2025 DBCGM: A Granular Model for Big Data Classification Based on Data Bisection and Cascade Weighted Clustering
Jiande Huang, Yuhui Deng 0001, Yi Zhou 0009, Shujie Pang, Qifen Yang, Geyong Min
IEEE Trans. Knowl. Data Eng.3
2025 Gecko: Efficient Sliding Window Aggregation With Granular-Based Bulk Eviction Over Big Data Streams
abstract
Sliding window aggregation, which extracts summaries from data streams, is a core operation in streaming analysis. Though existing sliding window algorithms that perform single eviction and insertion operations can achieve a worst-case time complexity of$O(1)$for in-order streams, real-world data streams often involve out-of-order data and exhibit burst data characteristics, which pose performance challenges to these sliding window algorithms. To address this challenging issue, we proposeGecko- a novel sliding window aggregation algorithm that supports bulk eviction. Gecko leverages a granular-based eviction strategy for various bulk sizes, enabling efficient bulk eviction while maintaining the performance close to that of in-order stream algorithms for single evictions. For large data bulks, Gecko performs coarse-grained eviction at the chunk level, followed by fine-grained eviction using leftward binary tree aggregation (LTA) as a complementary method. Moreover, Gecko partitions data based on chunks to prevent the impacts of out-of-order data on other chunks, thereby enabling efficient handling of out-of-order data streams. We conduct extensive experiments to evaluate the performance of Gecko. Experimental results demonstrate that Gecko exhibits superior performance over other solutions, which is consistent with theoretical expectations. In real-world data scenarios, Gecko improves the average throughput of the state-of-the-art algorithm b_FiBA by 1.7 times, with a maximum improvement of up to 3.5 times. Gecko also demonstrates the best latency performance among all compared schemes.
Jianjun Li 0012, Yuhui Deng 0001, Jiande Huang, Yi Zhou 0009, Qifen Yang, Geyong Min
IEEE Trans. Knowl. Data Eng.4
2024 FIG: Feature-Weighted Information Granules With High Consistency Rate
abstract
Information granules are effective in revealing the structure of data. Therefore, it is a common practice in data mining to use information granules for classifying datasets. In the existing granular classifiers, the information granules are often classified according to the standard membership function only without considering the influence of different feature weights on the quality of granules and label classification results. In this article, we utilize the feature weighting of data to produce the information granules with high consistency rate called FIG. Firstly, we use consistency rate and contribution scores to generate information granules. Then, we propose a granular two-stage classifier GTC based on FIG. GTC divides the data into fuzzy and fixed points and then calculates the interval matching degree to assign data points to the most suitable cluster in the second step. Finally, we compare FIG with two state-of-the-art granular models (T-GrM and FGC-rule), and classification accuracy is also compared with other classification algorithms. The extensive experiments on synthetic datasets and public datasets from UCI show that FIG has sufficient performance to describe the data structure and excellent capability under the constructed granular classifier GTC. Compared with T-GrM and FGC-rule, the time overhead required for FIG to obtain information granules is reduced by an average of 51.07%, the per unit quality of the granules is also increased by more than 14.74%. Compared with other classification algorithms, an average of 5.04% improves GTC accuracy.
Jianghe Cai, Yuhui Deng 0001, Yi Zhou 0009, Jiande Huang, Geyong Min
IEEE Trans. Big Data3
2024 FaaSBatch: Boosting Serverless Efficiency With In-Container Parallelism and Resource Multiplexing
abstract
With high scalability and flexibility, serverless computing is becoming the most promising computing model. Existing serverless computing platforms initiate a container for each function invocation, which leads to a huge waste of computing resources. Our examinations reveal that (i) executing invocations concurrently within a single container can provide comparable performance to that provided by multiple containers (i.e., traditional approaches); (ii) redundant resources generated within a container result in memory resource waste, which prolongs the execution time of function invocations. Motivated by these insightful observations, we propose FaaSBatch - a serverless framework that reduces invocation latency and saves scarce computing resources. In particular, FaaSBatch first classifies concurrent function requests into different function groups according to the invocation information. Next, FaaSBatch batches the invocations of each group, aiming to minimize resource utilization. Then, FaaSBatch utilizes an inline parallel policy to map each group of batched invocations into a single container. Finally, FaaSBatch expands and executes invocations of containers in parallel. To further reduce invocation latency and resource utilization, within each container, FaaSBatch reuses redundant resources created during function execution. We conduct extensive experiments based on Azure traces to evaluate the effectiveness and performance of FaaSBatch. We compare FaaSBatch with three state-of-the-art schedulers Vanilla, SFS, and Kraken. Our experimental results show that FaaSBatch effectively and remarkably slashes invocation latency and resource overhead. For instance, when executing I/O functions, FaaSBatch cuts back the invocation latency of Vanilla, SFS, and Kraken by up to 72.58%, 74.10%, and 72.62%, respectively; FaaSBatch also slashes the resource overhead of Vanilla, SFS, and Kraken by 70.2% to 98.40%, 67.74% to 98.12%, and 43.01% to 78.90%, respectively.
Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Jie Li 0067, Shujie Pang, Xiao Qin 0001
IEEE Trans. Computers3
2024 BTVMP: A Burst-Aware and Thermal-Efficient Virtual Machine Placement Approach for Cloud Data Centers
abstract
With the rapid growth of cloud computing, frequent workload bursts show an increasing influence on the Quality of Service (QoS) and energy efficiency of cloud-based data centers. Existing virtual machine placement schemes are expected to optimize either QoS or energy efficiency for cloud data centers running under bursty workload conditions. To bridge this gap, we propose a burst-aware and thermal-efficient virtual machine placement technique calledBTVMP. BTVMP adopts a two-step strategy to achieve energy efficiency while assuring QoS. First, BTVMP leverages a split-and-recombine algorithm – SAR – to deal with bursty workloads. SAR prioritizes critical workloads while preventing low-priority workloads from starvation, thereby assuring QoS. Second, BTVMP utilizes an enhanced simulated annealing algorithm calledESAto offer optimal thermal-efficient virtual machine placement (VMP) solutions, aiming to minimize the energy consumption of data centers. To facilitate estimating energy consumption, we integrate into BTVMP a thermal model that takes into account heat re-circulation effects. We conduct extensive experiments with a real-world trace. We compare BTVMP with the leading-edge VMP strategies, including Genetic Algorithm (XINT-GA), Power-Aware and Performance-Guaranteed Virtual Machine Placement (PPVMP), Peak Load Scheduling Control Method (PLSC), First Come First Serve (FCFS), and GReedy based scheduling Algorithm miNImizing Total Energy (GRANITE). The experimental results unveil that BTVMP not only enhances QoS but also exhibits superb energy efficiency. In particular, BTVMP reduces PLSC's workload delay and FCFS's critical workload delay by 18$\%$and 11$\%$, respectively. Moreover, BTVMP lowers the total energy consumption of the three alternative algorithms –GRANITE, XINTGA, PPVMP, and PLSC – by anywhere between 27.8$\%$and 49.4$\%$.
Jie Li 0067, Yuhui Deng 0001, Rui Wang 0001, Yi Zhou 0009, Hao Feng 0010, Geyong Min, Xiao Qin 0001
IEEE Trans. Serv. Comput.4
2023 FaaSBatch: Enhancing the Efficiency of Serverless Computing by Batching and Expanding Functions
abstract
With high scalability and flexibility, serverless computing is becoming the most promising computing model. Existing serverless computing platforms initiate a container for each function invocation, which leads to a huge waste of computing resources. Our examinations reveal that (i) executing invocations concurrently within a single container can provide comparable performance to that provided by multiple containers (i.e., traditional approaches); (ii) redundant resources generated within a container result in memory resource waste, which prolongs the execution time of function invocations. Motivated by these insightful observations, we propose FaaSBatch - a serverless framework that reduces invocation latency and saves scarce computing resources. In particular, FaaSBatch first classifies concurrent function requests into different function groups according to the invocation information. Next, FaaSBatch batches the invocations of each group, aiming to minimize resource utilization. Then, FaaSBatch utilizes an inline parallel policy to map each group of batched invocations into a single container. Finally, FaaSBatch expands and executes invocations of containers in parallel. To further reduce invocation latency and resource utilization, within each container, FaaSBatch reuses redundant resources created during function execution. We conduct extensive experiments based on Azure traces to evaluate the effectiveness and performance of FaaSBatch. We compare FaaSBatch with three state-of-the-art schedulers Vanilla, SFS, and Kraken. Our experimental results show that FaaSBatch effectively and remarkably slashes invocation latency and resource overhead. For instance, when executing I/O functions, FaaSBatch cuts back the invocation latency of Vanilla, SFS, and Kraken by up to 92.18%, 89.54%, and 90.65%, respectively; FaaSBatch also slashes the resource overhead of Vanilla, SFS, and Kraken by 58.89% to 94.77%, 43.72% to 90.39%, and 42.99% to 78.88%, respectively.
Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Jie Li 0067, Shujie Pang
ICDCS3
2023 A computational approach for real-time detection of fake news
Chaowei Zhang 0001, Ashish Gupta 0004, Xiao Qin 0001, Yi Zhou 0009
Expert Syst. Appl.4
2023 Towards Thermal-Aware Workload Distribution in Cloud Data Centers Based on Failure Models
abstract
Increasing workload conditions lead to a significant surge in power consumption and computing node failures in data centers. The existing workload distribution strategies focused on either thermal awareness or failure mitigation, overlooking the impact of node failures on the energy efficiency of cloud data centers. To address this issue, a new holistic model is built to characterize the impacts of workloads, computing and cooling costs, heat recirculation, and node failure on the energy efficiency of cloud data centers. Leveraging such a holistic model, we propose a novel thermal-aware workload distribution strategy calledHGSAthat takes node failure into accountand can improve the energy efficiency of cloud data centers. Our empirical findings confirm that (i) faulty nodes lead to a large rise in power consumption, and (ii) failure locations play a vital role in the power consumption of data centers. Experimental results unveil that HGSA is adroit at making near-optimal decisions in workload distribution strategies. In particular, HGSA cuts down the minimum inlet temperature by 5.2$\%$-15$\%$, improves the maximum air temperature of a Computer Room Air Conditioner (CRAC) model by 4.2$\%$-26.5$\%$, lowers the cooling cost by 15.4$\%$-50$\%$compared to the existing solutions. Furthermore, HGSA cuts back the total power consumption by 0.65$\%$-78$\%$.
Jie Li 0067, Yuhui Deng 0001, Yi Zhou 0009, Zhen Zhang 0017, Geyong Min, Xiao Qin 0001
IEEE Trans. Computers3
2023 PcGC: A Parity-Check Garbage Collection for Boosting 3-D NAND Flash Performance
abstract
Garbage collection or GC running in the controller of 3-D NAND flash-based solid-state disks—SSDs—plays a critical role in the performance of storage systems. SSD manufacturers have developed various GC solutions based on internal data movement or IDM to mitigate the impacts of GC on request latency. Due to the circuit characteristics of flash memory, the existing IDM-based GC strategies are restricted by page parity during data movement: odd pages must be migrated to odd pages, and even pages to even pages. When migrating two consecutive pages with the same parity, the free page between the two migrated pages will be wasted after the migration is complete. This ever-increasing page waste problem inevitably deteriorates the storage space utilization of flash memory, thereby degrading the overall performance of 3-D NAND flash-based SSDs. To address this issue, we propose a parity-check GC scheme called PcGC to revamp SSD performance by alleviating page waste during GC. We build a parity-check unit in PcGC to facilitate checking the parity of migrated valid pages and destination pages. According to the parity results offered by the parity-check unit, PcGC dynamically adjusts the migration order of valid pages during the course of GC. In doing so, PcGC fundamentally averts page waste caused by the page parity restriction, thereby enhancing 3-D NAND flash performance. We quantitatively evaluate the performance of PcGC in terms of wasted pages, storage utilization, GC counts, write amplification, and average response time. We compare PcGC against the two state-of-the-art schemes—Amphibian and Tiny-tail flash (TTflash). The experimental results derived from the nine real-world workload traces unfold that compared with Amphibian and TTflash: 1) PcGC curtails the number of wasted pages by up to 91.4% with an average of 53.75%; 2) cuts back the number of GC counts by up to 52.2% with an average of 11.9%; and 3) slashes average write response time by up to 77.8% with an average of 13.0%.
Shujie Pang, Yuhui Deng 0001, Genxiong Zhang, Yi Zhou 0009, Xiao Qin 0001, Zhaorui Wu, Jie Li 0067
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Cocktail: Mixing Data With Different Characteristics to Reduce Read Reclaims for nand Flash Memory
abstract
A large number of read-disturb-induced rewrites are performed in the background [also known as Read Reclaim (RR)] to alleviate the read-disturb issue in NAND flash memory-based SSDs. RR can significantly degrade the performance and shorten the service life of SSD in read-intensive workloads. To address this issue, we propose a novel read-disturb management approach called Cocktail that mixes a small proportion of hot-read pages with a large proportion of cold-read pages, thereby avoiding clustering hot-read pages into a few blocks. Motivated by the insight that RR operations are frequently triggered by hot read-pages, Cocktail first prefills a portion of each block with cold data extracted from user requests. Then, Cocktail fills the prefilled blocks with write-back data caused by RR to create read-balanced blocks. We integrate two thresholds, write pool capacity and the ratio of RR-write data to User-write data, into Cocktail to govern the ratio of write-back data caused by RR to data of user requests in a block. Cocktail dynamically adjusts the two thresholds according to the characteristics of RR. Cocktail is conducive to decentralizing hot write-back data caused by RR across a broad range of blocks, thereby reducing the occurrence of second-time RR and the number of overall block reads. We compare Cocktail with three existing schemes baseline, redFTL, and IPR in terms of SSD service life, SSD response time, write amplification, and the number of garbage collections (GCs) under ten real-world workload conditions. Experimental results show that compared with the existing schemes, Cocktail reduces the number of RRs, the average response time, the 99-percentile tail latency, and the number of GCs by an average of 40.77%, 10.82%, 5.40%, and 12.29%, respectively. Cocktail also alleviates the write amplification of the three alternative schemes by an average of 49.57%.
Genxiong Zhang, Yuhui Deng 0001, Yi Zhou 0009, Shujie Pang, Jianhui Yue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 UPOA: A User Preference Based Latency and Energy Aware Intelligent Offloading Approach for Cloud-Edge Systems
abstract
Task offloading has been widely used to extend the battery life of intelligent mobile devices. Existing task offloading approaches, focusing on perfecting the balance between latency and energy consumption, completely ignore the impacts of the user preference caused by low battery anxiety. The existence of low battery anxiety - mobile users’ common fear of losing battery energy, especially when the battery energy is already low - causes users to trade high latency for prolonged battery life. Taking into account the user preference impacts on task offloading, we propose a novel offloading approach calledUPOAto obtain refined offloading policies between low latency and energy consumption based on user preferences. In UPOA, we start the study by defining a user preference rule that determines users’ offloading preferences according to battery energy status. Then, we build a fine-grained task offloading model to delineate the task distribution characteristics of each node in its offloading link. Guided by this model, we develop a task prediction algorithm based on the long-short-term-memory neural network model to provide task predictions that facilitate offloading policies. Lastly, we implement a particle-swarm-optimization-based online offloading algorithm. The offloading algorithm provides the best long-term offloading policies by incorporating the user preference determined by our user preference rule and the task predictions generated by our task prediction algorithm. To quantitatively evaluate the performance of UPOA, we conduct extensive experiments in a real-world cloud-edge environment. We compare UPOA with three state-of-the-art offloading approaches, DRA, DRL-E2D, and MUDRL under various conditions. Experimental results demonstrate that UPOA can make effective policies based on user preferences compared with the existing approaches. UPOA reduces average latency by 12.49% when battery energy is sufficient and extends battery life by 20.14% when battery energy is low.
Jingling Yuan, Yao Xiang, Yuhui Deng 0001, Yi Zhou 0009, Geyong Min
IEEE Trans. Cloud Comput.4
2023 MGRM: A Multi-Segment Greedy Rewriting Method to Alleviate Data Fragmentation in Deduplication-Based Cloud Backup Systems
abstract
Data deduplication has been broadly used in Cloud due to its storage space saving ability. An issue of deduplication is the contiguous data chunks in a segment may be scattered in different containers. This phenomenon is called data fragmentation. Because of data fragmentation, a restore process must reference various containers across a wide variety of segments, thereby hurting the restore performance. Capping methods that rewrite the data chunks of low Container Reference Ratio (CRR) containers are developed to alleviate data fragmentation. We analyze and observe from real traces that a number of segments only point to lowCRRcontainers, while some others only contain highCRRcontainers. This interesting observation is ignored by the existing capping methods which sort containers from a single segment, falling short in searching multiple segments collectively. Thus, the reference count of selected containers in the existing capping methods is still high. To address this problem, we propose a multi-segment greedy rewriting method named MGRM. MGRM sorts containers of segments in a sequential way. More specifically, given thei-thsegment currently being processed, MGRM will sort all the containers in the topi-thsegments. This salient searching feature enables MGRM to select and rewrite the true low-reference container set. Moreover, to achieve a good balance between deduplication ratio and restore performance, MGRM has two working modes: an optimal rewriting mode and a radical rewriting mode. When working in the optimal rewriting mode, MGRM aims to improve the deduplication ratio; when the radical rewriting mode, MGRM strives to improve the restore performance. MGRM adaptively switches the working mode according to workload. Furthermore, unlike the existing capping methods that improve restore performance at the cost of the deduplication ratio, MGRM pays attention to both aspects. Our extensive experimental results show that MGRM achieves high restore performance, coupled with a high deduplication ratio. In particular, compared with the two state-of-art schemes FC and FLC, MGRM improves the deduplication ratio and restore performance by up to 114.83% and 99.34%, respectively.
Datong Zhang, Yuhui Deng 0001, Yi Zhou 0009, Jie Li 0067, Weiheng Zhu, Geyong Min
IEEE Trans. Cloud Comput.3
2023 TADRP: Toward Thermal-Aware Data Replica Placement in Data-Intensive Data Centers
abstract
With the mushrooming growth of data volumes, data replica placement plays a key role in promoting the energy efficiency and Quality-of-Service (QoS) of data-intensive data centers. The existing data placement strategies mainly focus on storage performance improvement or QoS enhancement in data centers, but ignore the indispensable factor - heat recirculation. To bridge this gap, we propose a thermal-aware data replica placement strategy called TADRP, aiming to improve cooling efficiency and minimize the total power consumption of data-intensive data centers. TADRP leverages an ant colony optimization (ACO) algorithm coupled with Laplacian probability distribution to find a quasi-optimal disk sequence (or Disk Sequence for short), which consists of disks selected from different rack servers to place data replicas. TADRP categorizes disks of Disk Sequence into active and inactive ones, by placing hot and cold replicas on active and inactive disks, respectively. We quantitatively evaluate TADRP in terms of cooling costs, total power consumption, number of power-state transitions, and execution time. We compare TADRP with four alternative solutions, namely, Random, Hadoop, SRS, and CDP-NSGAII. Experimental results show that TADRP can reduce the cooling costs and the total power of the existing solutions by 14.7% - 61.7% and 19.2%-55.1%, respectively, without undesirable I/O performance drops.
Jie Li 0067, Yuhui Deng 0001, Yi Zhou 0009, Zhaorui Wu, Shujie Pang, Geyong Min
IEEE Trans. Netw. Serv. Manag.3
2023 InDe: An Inline Data Deduplication Approach via Adaptive Detection of Valid Container Utilization
abstract
Inline deduplication removes redundant data in real-time as data is being sent to the storage system. However, it causes data fragmentation: logically consecutive chunks are physically scattered across various containers after data deduplication. Many rewrite algorithms aim to alleviate the performance degradation due to fragmentation by rewriting fragmented duplicate chunks as unique chunks into new containers. Unfortunately, these algorithms determine whether a chunk is fragmented based on a simple pre-set fixed value, ignoring the variance of data characteristics between data segments. Accordingly, when backups are restored, they often fail to select an appropriate set of old containers for rewrite, generating a substantial number of invalid chunks in retrieved containers. To address this issue, we propose an inline deduplication approach for storage systems, called InDe , which uses a greedy algorithm to detect valid container utilization and dynamically adjusts the number of old container references in each segment. InDe fully leverages the distribution of duplicated chunks to improve the restore performance while maintaining high backup performance. We define an effectiveness metric, valid container referenced counts (VCRC) , to identify appropriate containers for the rewrite. We design a rewrite algorithm F-greedy that detects valid container utilization to rewrite low-VCRC containers. According to the VCRC distribution of containers, F-greedy dynamically adjusts the number of old container references to only share duplicate chunks with high-utilization containers for each segment, thereby improving the restore speed. To take full advantage of the above features, we further propose another rewrite algorithm called F-greedy+ based on adaptive interval detection of valid container utilization. F-greedy+ makes a more accurate estimation of the valid utilization of old containers by detecting trends of VCRC’s change in two directions and selecting referenced containers in the global scope. We quantitatively evaluate InDe using three real-world backup workloads. The experimental results show that compared with two state-of-the-art algorithms (Capping and SMR), our scheme improves the restore speed by 1.3×–2.4× while achieving almost the same backup performance.
Lifang Lin, Yuhui Deng 0001, Yi Zhou 0009
ACM Trans. Storage3
2023 PSA-Cache: A Page-state-aware Cache Scheme for Boosting 3D NAND Flash Performance
abstract
Garbage collection (GC) plays a pivotal role in the performance of 3D NAND flash memory, where Copyback has been widely used to accelerate valid page migration during GC. Unfortunately, copyback is constrained by the parity symmetry issue: data read from an odd/even page must be written to an odd/even page. After migrating two odd/even consecutive pages, a free page between the two migrated pages will be wasted. Such wasted pages noticeably lower free space on flash memory and cause extra GCs, thereby degrading solid-state-disk (SSD) performance. To address this problem, we propose a page-state-aware cache scheme called PSA-Cache , which prevents page waste to boost the performance of NAND Flash-based SSDs. To facilitate making write-back scheduling decisions, PSA-Cache regulates write-back priorities for cached pages according to the state of pages in victim blocks. With high write-back-priority pages written back to flash chips, PSA-Cache effectively fends off page waste by breaking odd/even consecutive pages in subsequent garbage collections. We quantitatively evaluate the performance of PSA-Cache in terms of the number of wasted pages, the number of GCs, and response time. We compare PSA-Cache with two state-of-the-art schemes, GCaR and TTflash, in addition to a baseline scheme LRU. The experimental results unveil that PSA-Cache outperforms the existing schemes. In particular, PSA-Cache curtails the number of wasted pages of GCaR and TTflash by 25.7% and 62.1%, respectively. PSA-Cache immensely cuts back the number of GC counts by up to 78.7% with an average of 49.6%. Furthermore, PSA-Cache slashes the average write response time by up to 85.4% with an average of 30.05%.
Shujie Pang, Yuhui Deng 0001, Genxiong Zhang, Yi Zhou 0009, Yaoqin Huang, Xiao Qin 0001
ACM Trans. Storage4
2023 HashCache: Accelerating Serverless Computing by Skipping Duplicated Function Execution
abstract
Serverless computing is a leading force behind deploying and managing software in cloud computing. One inherent challenge in serverless computing is the increased overall latency due to duplicate computations. Our initial investigation into the function invocations of serverless applications reveals an abundance of duplicate invocations. Inspired by this critical observation, we introduceHashCache, a system designed to cache duplicate function invocations, thereby mitigating duplicate computations. In HashCache, serverless functions are classified into three categories, namely, computational functions, stateful functions, and environment-related functions. On the grounds of such a function classification, HashCache associates the stateful functions and their states to build an adaptive synchronization mechanism. With this support, HashCache exploits the cached results of computational and stateful functions to serve upcoming invocation requests to the same functions, thereby reducing duplicate computations. Moreover, HashCache stores remote files probed by stateful functions into a local cache layer, which further curtails invocation latency. We implement HashCache within theApache OpenWhiskto forge a cache-enabled serverless computing platform. We conduct extensive experiments to quantitatively evaluate the performance of HashCache in terms of invocation latency and resource utilization. We compare HashCache against two state-of-the-art approaches -FaaSCacheandOpenWhisk. The experimental results unveil that our HashCache remarkably reduces invocation latency and resource overhead. More specifically, HashCache curbs the 99-tail latency of FaaSCache and OpenWhisk by up to 91.37% and 95.96% in real-world serverless applications. HashCache also slashes the resource utilization of FaaSCache and OpenWhisk by up to 31.62% and 35.51%, respectively.
Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Lin Cui 0001, Xiao Qin 0001
IEEE Trans. Parallel Distributed Syst.3
2022 Towards Heat-Recirculation-Aware Virtual Machine Placement in Data Centers
abstract
As customers take virtual machines (VMs) as their demands, high-efficient placement of VMs is required to reduce the energy consumption in data centers. Existing Virtual Machine placement (VMP) strategies can minimize energy consumption of data centers by optimizing resource allocation in terms of multiple physical resources (e.g., memory, bandwidth, CPU, etc.). However, these strategies ignore the role of heat recirculation in the data center, which can cause a huge energy waste in cooling. To address this problem, we propose a heat-recirculation-aware VMP strategy for reducing the energy consumption of data centers. This novel VMP strategy takes into account heat recirculation coupled with multiple physical resource allocation to reduce the energy consumption of data centers. We design a simulated annealing based algorithm called SABA to lower the energy consumption of data centers where multiple VMs are deployed. SABA remarkably cuts down the energy consumption of physical resources through two salient features. First, it obtains an approximation of the optimum with much fewer iterations than simulated annealing algorithm (SA). Second, it reduces the number of activated servers required for VM tasks. We quantitatively evaluate the performance of SABA in terms of algorithm efficiency, the number of activated servers and the energy-saving. We compare the performance of SABA with state-of-art XINT-GA, PPVMP, TSTD and SA algorithms. Moreover, we evaluate the efficiency of SABA by leveraging a real-world 50 hours trace from practical IBM cloud data centers. Experimental results indicate that our heat-recirculation-aware VM placement strategy provides a powerful solution for improving the energy efficiency of data centers (SABA improves energy efficiency of cooling by up to 13.2% over TSTD, 13.8% over XINT-GA and 45% over PPVMP algorithm).
Hao Feng 0010, Yuhui Deng 0001, Yi Zhou 0009, Geyong Min
IEEE Trans. Netw. Serv. Manag.3
2022 Blender: A Container Placement Strategy by Leveraging Zipf-Like Distribution Within Containerized Data Centers
abstract
Instantiated containers of an application are distributed across multiple Physical Machines (PMs) to achieve high parallel performance. Container placement plays a vital role in network traffic and the performance of containerized data centers. Existing container placement techniques are inadequate due to the ignorance of container traffic patterns. To solve this issue, we first investigate the network traffic between containers and observe that it exhibits a Zipf-like distribution. Motivated by this finding, we propose a novel container placement approach-Blender-by taking into account the Zipf-like distribution. Blender employs two algorithms calledRefineAlgandSplitAlgto divide containers of applications into blocks, and place these blocks across Virtual Machines (VMs). Blender exhibits two salient features: (i) it minimizes inter-block traffic by arranging the containers that communicate frequently in the same block. (ii) it achieves good load balancing by combining complementary blocks that request different resource types (e.g.,CPU-intensiveandmemory-intensiveblocks) and distributing these blocks across multiple VMs. The experimental results show that Blender significantly reduces communication traffic and network latency. In particular, Blender reduces the traffic of SBP and CA-WFD by 22% and 32%, respectively. Blender decreases network latency by 16% and 26% compared to SBP and CA-WFD. Furthermore, with Blender in place, the physical resources of hosting PMs are well balanced and utilized.
Zhaorui Wu, Yuhui Deng 0001, Hao Feng 0010, Yi Zhou 0009, Geyong Min, Zhen Zhang 0017
IEEE Trans. Netw. Serv. Manag.4
2021 Blender: A Traffic-Aware Container Placement for Containerized Data Centers
abstract
Instantiated containers of an application are distributed across multiple Physical Machines (PMs) to achieve high parallel performance. Container placement plays a vital role in network traffic and the performance of containerized data centers. Existing container placement techniques do not consider the container traffic pattern, which is inadequate. To resolve this conflict, we investigate network traffic between containers and observe that it exhibits a Zipf-like distribution. We propose a novel container placement approach - Blender - by leveraging the Zipf-like distribution. Based on network traffic correlation, Blender employs RefineAlg and SplitAlg to divide containers of applications into blocks, and place these blocks across virtual machines. Blender exhibits two salient features: (i) it minimizes inter-block traffic by arranging the containers that communicate frequently in the same block. (ii) it achieves good load balancing by combining blocks according to the resource types they require and distributing them across multiple PMs. We compare Blender against two state-of-the-art methods SBP and CA-WFD. The experimental results show that Blender significantly reduces communication traffic. In particular, for the same number of PMs, Blender reduces the traffic of SBP and CA-WFD by 22% and 32%, respectively. Furthermore, with Blender in place, the physical resources of hosting PMs are well balanced and utilized.
Zhaorui Wu, Yuhui Deng 0001, Hao Feng 0010, Yi Zhou 0009, Geyong Min
DATE4
2021 Improving Restore Performance of Deduplication Systems via a Greedy Rewriting Scheme
abstract
Data deduplication has been widely used to improve storage space utilization, however, it is baffled by data fragmen-tation: logically consecutive chunks physically scattered across various containers. Many rewriting schemes, rewriting fragment-ed duplicate chunks into new containers, attempt to alleviate the restore performance degradation caused by fragmentation. Unfortunately, these schemes rely on a fixed threshold and fail to choose the appropriate set of old containers for rewriting, which leads to substantial redundant chunks existing in the retrieved containers when restoring backups. To address this issue, we propose a flexible threshold rewriting scheme to improve restore performance while maintaining high backup performance. We define an effectiveness metric - valid container reference counts (VCRC) - that facilitates identifying the appropriate containers for rewriting. We design a greedy-algorithm-based algorithm called F-greedy that dynamically adjusts the threshold according to the distribution of containers' VCRC, aiming to rewrite low-VCRC containers. We quantitatively evaluate F-greedy on three real-world backup datasets in terms of restore performance, backup performance, and storage overhead. The empirical results show that compared with two state-of-the-art schemes (Capping and SMR), our scheme improves the restore speed of the exiting algorithms by 1.3x - 2.4x while achieving similar backup performance.
Lifang Lin, Yuhui Deng 0001, Yi Zhou 0009
ICPADS3
2021 RUE: A caching method for identifying and managing hot data by leveraging resource utilization efficiency
abstract
Abstract In this study, we propose a caching method called RUE for dynamic large‐scale data streams. We define a data model to facilitate hot data identification and management. At the heart of RUE model is hot degree that takes into account two factors data resource utilization efficiency and reuse distance, aiming to quantitatively reflect data popularity in a dynamic data stream. Based on data's hot degree, RUE classifies data into four types, each of which is assigned with an associated cache residence time. Guided by RUE model, we develop HM algorithm to identify and manage hot data in a dynamic data stream. HM algorithm is implemented by four stacks, namely, new stack, short stack, long stack, and temp stack. Moreover, an eviction and a migration algorithms are integrated into HM to facilitate block replacement and migration. To evaluate the performance of HM algorithm, we quantitatively compare the performance of RUE with three state‐of‐art algorithms, namely, LRU, LIRS, and ARC under various replacement policies, operations, and workloads. Experimental results show that RUE outperforms these three existing algorithms in terms of both read and write hit rates. Furthermore, we show that with the four stacks in place, the computing overhead of HM is negligible.
Liang Ai, Yuhui Deng 0001, Yi Zhou 0009, Hao Feng 0010
Softw. Pract. Exp.3
2021 Improving the Performance of Deduplication-Based Backup Systems via Container Utilization Based Hot Fingerprint Entry Distilling
abstract
Data deduplication techniques construct an index consisting of fingerprint entries to identify and eliminate duplicated copies of repeating data. The bottleneck of disk-based index lookup and data fragmentation caused by eliminating duplicated chunks are two challenging issues in data deduplication. Deduplication-based backup systems generally employ containers storing contiguous chunks together with their fingerprints to preserve data locality for alleviating the two issues, which is still inadequate. To address these two issues, we propose a container utilization based hot fingerprint entry distilling strategy to improve the performance of deduplication-based backup systems. We divide the index into three parts: hot fingerprint entries, fragmented fingerprint entries, and useless fingerprint entries. A container with utilization smaller than a given threshold is called a sparse container . Fingerprint entries that point to non-sparse containers are hot fingerprint entries. For the remaining fingerprint entries, if a fingerprint entry matches any fingerprint of forthcoming backup chunks, it is classified as a fragmented fingerprint entry. Otherwise, it is classified as a useless fingerprint entry. We observe that hot fingerprint entries account for a small part of the index, whereas the remaining fingerprint entries account for the majority of the index. This intriguing observation inspires us to develop a hot fingerprint entry distilling approach named HID . HID segregates useless fingerprint entries from the index to improve memory utilization and bypass disk accesses. In addition, HID separates fragmented fingerprint entries to make a deduplication-based backup system directly rewrite fragmented chunks, thereby alleviating adverse fragmentation. Moreover, HID introduces a feature to treat fragmented chunks as unique chunks. This feature compensates for the shortcoming that a Bloom filter cannot directly identify certain duplicated chunks (i.e., the fragmented chunks). To take full advantage of the preceding feature, we propose an evolved HID strategy called EHID . EHID incorporates a Bloom filter, to which only hot fingerprints are mapped. In doing so, EHID exhibits two salient features: (i) EHID avoids disk accesses to identify unique chunks and the fragmented chunks; (ii) EHID slashes the false positive rate of the integrated Bloom filter. These salient features push EHID into the high-efficiency mode. Our experimental results show our approach reduces the average memory overhead of the index by 34.11% and 25.13% when using the Linux dataset and the FSL dataset, respectively. Furthermore, compared with the state-of-the-art method HAR, EHID boosts the average backup throughput by up to a factor of 2.25 with the Linux dataset, and EHID reduces the average disk I/O traffic by up to 66.21% when it comes to the FSL dataset. EHID also marginally improves the system's restore performance.
Datong Zhang, Yuhui Deng 0001, Yi Zhou 0009, Xiao Qin 0001
ACM Trans. Storage3
2020 A Heat-Recirculation-Aware VM Placement Strategy for Data Centers
abstract
Data centers consisted of a great number of IT devices (e.g., servers, switches and etc.) which generates a massive amount of heat emission. Due to the special arrangement of racks in the data center, heat-recirculation often occurs between nodes. It can cause a sharp rise in temperature of the equipment coupled with local hot spots in data centers. Existing VM placement strategies can minimize energy consumption of data centers by optimizing resource allocation in terms of multiple physical resources (e.g., memory, bandwidth, cpu and etc.). However, existing strategies ignore the role of heat-recirculation in the data center. To address this problem, in this study, we propose a heat-recirculation-aware VM placement strategy and design a Simulated Annealing Based Algorithm (SABA) to lower the energy consumption of data centers. Different from the existing SA algorithm, SABA optimize the distribution of the initial solution and the way of iteration. We quantitatively evaluate SABA’s performance in terms of algorithm efficiency, the activated servers and the energy saving against with XINT-GA algorithm (Thermal-aware task scheduling Strategy), FCFS (First-Come First-Served), and SA. Experimental results indicate that our heat-recirculation-aware VM placement strategy provides a powerful solution for improving energy efficiency of data centers.
Hao Feng 0010, Yuhui Deng 0001, Yi Zhou 0009
DATE3
2020 EDOM: Improving energy efficiency of database operations on multicore servers
Yi Zhou 0009, Shubbhi Taneja, Xiao Qin 0001, Wei-Shinn Ku, Jifu Zhang
Future Gener. Comput. Syst.1
2020 ThermoBench: A thermal efficiency benchmark for clusters in data centers
Yi Zhou 0009, Yuanqi Chen, Shubbhi Taneja, Ajit Chavan, Xiao Qin 0001, Jifu Zhang
Parallel Comput.1
2019 GreenDB: Energy-Efficient Prefetching and Caching in Database Clusters
abstract
In this study, we propose an energy-efficient database system called GreenDB running on clusters. GreenDB applies a workload-skewness strategy by managing hot nodes coupled with a set of cold nodes in a database cluster. GreenDB fetches popular data tables to hot nodes, aiming to keep cold nodes in the low-power mode in increased time periods. GreenDB is conducive to reducing the number of power-state transitions, thereby lowering energy-saving overhead. A prefetching model and an energy saving model are seamlessly integrated into GreenDB to facilitate the power management in database clusters. We quantitatively evaluate GreenDB's energy efficiency in terms of managing, fetching, and storing data. We compare GreenDB's prefetching strategy with the one implemented in Postgresql. Experimental results indicate that GreenDB conserves the energy consumption of the existing solution by up to 98.4 percent. The findings show that the energy efficiency of GreenDB can be optimized by tuning system parameters, including table size, hit rates, number of nodes, number of disks, and inter-arrival delays.
Yi Zhou 0009, Shubbhi Taneja, Chaowei Zhang 0001, Xiao Qin 0001
IEEE Trans. Parallel Distributed Syst.1
2018 Improving Energy Efficiency of Database Clusters Through Prefetching and Caching
abstract
The goal of this study is to optimize energy efficiency of database clusters through prefetching and caching strategies. We design a workload-skewness scheme to collectively manage a set of hot and cold nodes in a database cluster system. The prefetching mechanism fetches popular data tables to the hot nodes while keeping unpopular data in cold nodes. We leverage a power management module to aggressively turn cold nodes in the low-power mode to conserve energy consumption. We construct a prefetching model and an energy-saving model to govern the power management module in database lusters. The energy-efficient prefetching and caching mechanism is conducive to cutting back the number of power-state transitions, thereby offering high energy efficiency. We systematically evaluate energy conservation technique in the process of managing, fetching, and storing data on clusters supporting database applications. Our experimental results show that our prefetching/caching solution significantly improves energy efficiency of the existing PostgreSQL system.
Yi Zhou 0009, Shubbhi Taneja, Mohammed I. Alghamdi, Xiao Qin 0001
CCGrid1
2018 Thermal benchmarking and modeling for HPC using big data applications
Shubbhi Taneja, Yi Zhou 0009, Xiao Qin 0001
Future Gener. Comput. Syst.2
2018 Towards thermal-aware Hadoop clusters
Yi Zhou 0009, Shubbhi Taneja, Gautam Dudeja, Xiao Qin 0001, Jifu Zhang, Minghua Jiang, Mohammed I. Alghamdi
Future Gener. Comput. Syst.1
2017 Thermal-aware task assignments in high performance computing clusters
abstract
Summary Cluster‐level thermal management has gained much attention over the past decade due to rising cooling costs associated with data centers. In this research, we propose and implement a static scheduler called SSched and a dynamic one named DSched. These 2 algorithms schedule jobs based on CPU and disk temperatures of a Hadoop cluster's nodes. Our schedulers rely on a monitoring mechanism to keep track of CPU and disk utilization, maintaining CPU and disk temperatures below a threshold through thermal‐aware scheduling decisions. To facilitate the design of SSched and DSched, we classify jobs into the CPU‐intensive and disk‐intensive categories. When a job arrives, SSched retrieves the utilization stats from a profiled log, estimates the thermal behavior, and places the job on NodeManager to minimize thermal impacts. Unlike SSched, DSched improves thermal efficiency of Hadoop clusters through dynamic load balancing. DSched keeps track of the coolest and hottest nodes in the cluster; tasks are migrated from hot nodes into cool ones if any hot spot is detected. To evaluate the effectiveness of our schedulers, we keep track of average CPU and disk temperatures in a node, managing an optimal outlet temperature across a cluster. We demonstrate that compared with the traditional Hadoop scheduler, SSched and DSched achieve approximately 15% savings in terms of cooling cost with little performance overhead.
Shubbhi Taneja, Sanjay Kulkarni, Yi Zhou 0009, Xiao Qin 0001
Concurr. Comput. Pract. Exp.3
2017 aHDFS: An Erasure-Coded Data Archival System for Hadoop Clusters
abstract
In this paper, we propose an erasure-coded data archival system called aHDFS for Hadoop clusters, where RS(k + r; k) codes are employed to archive data replicas in the Hadoop distributed file system or HDFS. We develop two archival strategies (i.e., aHDFS-Grouping and aHDFS-Pipeline) in aHDFSto speed up the data archival process. aHDFS-Groupinga MapReduce-based data archiving scheme - keeps each mapper's intermediate output Key-Value pairs in a local key-value store. With the local store in place, aHDFS-Grouping merges all the intermediate key-value pairs with the same key into one single key-value pair, followed by shuffling the single Key-Value pair to reducers to generate final parity blocks. aHDFS-Pipeline forms a data archival pipeline using multiple data node in a Hadoop cluster. aHDFS-Pipeline delivers the merged single key-value pair to a subsequent node's local key-value store. Last node in the pipeline is responsible for outputting parity blocks. We implement aHDFS in a real-world Hadoop cluster. The experimental results show that aHDFS-Grouping and aHDFS-Pipeline speed up Baseline's shuffle and reduce phases by a factor of 10 and 5, respectively. When block size is larger than 32 MB, aHDFS improves the performance of HDFS-RAID and HDFS-EC by approximately 31.8 and 15.7 percent, respectively.
Yuanqi Chen, Yi Zhou 0009, Shubbhi Taneja, Xiao Qin 0001, Jianzhong Huang 0001
IEEE Trans. Parallel Distributed Syst.2
2016 Profiling Energy Usage of Web-Service Applications on Clusters
abstract
Energy saving is rapidly becoming one of the hottest topics in technology field within recent decades. With the development of technology, it brings a sheer increasing trend of data and the growth scale of clusters and data centers. Meanwhile, it also raises another essential issue into the path: energy cost. In this paper, we are diving into this key issue and evaluating energy- efficiency based on TPC-W benchmark: a notable web transaction e-commerce benchmark. We simulate the web transaction with different database sizes and collect the energy data by KILL-A-WATT. Also, we deploy this setup on four different cluster systems: PC nodes and wimpy nodes, and two different heterogeneous systems: using PC as front server and wimpy as Database server, and using wimpy as Web server and PC as Database server. Energy result demonstrates different characteristics among them, which can give lightening advice for future works in data center.
Mohammed I. Alghamdi, Wei-Shinn Ku, Yi Zhou 0009, Shubbhi Taneja, Xiao Qin 0001
NAS4