VLDB 2026 Research / reviewers in the wild / expert
Jianhui Yue
dblp:23/6331
· DBLP profile ↗
31ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0002-1876-6931ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 10 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gopher: Efficient Dynamic Graph Pattern Mining via DAG-Driven ExecutionabstractGraph pattern mining is essential for analyzing dynamic networks, where graphs evolve over time. To accommodate these changes, existing solutions update match sets incrementally, avoiding the need to re-mine the entire graph and achieving significant performance improvements. However, these methods suffer from inefficiencies due to redundant set intersection operations across subgraph instances, causing performance degradation. Yi Zhang 0191, Yu Huang 0013, Chaoqiang Liu, Haifeng Liu 0003, Jingrui Yuan, Jianhui Yue, Xiaofei Liao, Hai Jin 0001, Jingling Xue |
EuroSys | 7 |
| 2025 | Cheetah: Accelerating Dynamic Graph Mining with Grouping UpdatesabstractGraph pattern mining is essential for deciphering complex networks. In the real world, graphs are dynamic and evolve over time, necessitating updates in mining patterns to reflect these changes. Traditional methods use fine-grained incremental computation to avoid full re-mining after each update, which improves speed but often overlooks potential gains from examining inter-update interactions holistically, thus missing out on overall efficiency improvements. In this article, we introduce Cheetah, a dynamic graph mining system that processes updates in a coarse-grained manner by leveraging exploration domains . These domains exploit the community structure of real-world graphs to uncover data reuse opportunities typically missed by existing approaches. Exploration domains, which encapsulate extensive portions of the graph relevant to updates, allow multiple updates to explore the same regions efficiently. Cheetah dynamically constructs these domains using a management module that identifies and maintains areas of redundancy as the graph changes. By grouping updates within these domains and employing a neighbor-centric expansion strategy, Cheetah minimizes redundant data accesses. Our evaluation of Cheetah across five real-world datasets shows it outperforms current leading systems by an average factor of 2.63×. Yi Zhang 0191, Xiaomeng Yi, Yu Huang 0013, Jingrui Yuan, Chuangyi Gui, Dan Chen 0006, Long Zheng 0003, Jianhui Yue, Xiaofei Liao, Hai Jin 0001, Jingling Xue |
ACM Trans. Archit. Code Optim. | 8 |
| 2024 | FlashGNN: An In-SSD Accelerator for GNN TrainingabstractRecently, Graph Neural Networks (GNNs) have emerged as powerful tools for data analysis, surpassing traditional algorithms in various applications. However, the growing size of real-world datasets has outpaced the capabilities of centralized CPU or G PU - based systems. To address this challenge, numerous distributed systems have been proposed. However, these systems suffer from low hardware utilization due to slow network data exchange. While SSDs provide a promising alternative with large capacity and improved access latency, SSD-based G NN training on a single computer is bottlenecked by slow PCIe bus data transfer. This bottleneck leads to low CPU and G PU utilization, as confirmed by our experiments. Moreover, the design of in-SSD GNN training is hindered by slow access to flash memory. FlashGNN is a proposed solution that overcomes the PCIe bottleneck, fully utilizes I/O parallelism in flash chips, and maximizes data reuse from fetched flash memory chunks for efficient GNN training. We achieve this by designing the SSD firmware to coordinate data movements and hardware unit access. To address design challenges arising from slow flash memory and limited resources, we propose a novel node-wise GNN training method, an efficient scheduling algorithm for flash requests, and a high-performance subgraph generation method. Experimental results demonstrate that FlashGNN outperforms Ginex, a state-of-the-art SSD-based GNN training system, with a speed-up ratio ranging from 4.89× to 11.83 × and achieves energy savings of 57.14 × to 192.66 × for four typical real-world graph datasets. Additionally, FlashGNN is up to 23.17 × more efficient than the enhanced state-of-the-art in-storage accelerator, SmartSAGE+. Fuping Niu, Jianhui Yue, Jiangqiu Shen, Xiaofei Liao, Hai Jin 0001 |
HPCA | 2 |
| 2024 | A hybrid memory architecture supporting fine-grained data migration
Ye Chi, Jianhui Yue, Xiaofei Liao, Haikun Liu, Hai Jin 0001 |
Frontiers Comput. Sci. | 2 |
| 2024 | P3DC: Reducing DRAM Cache Hit Latency by Hybrid Mappings
Ye Chi, Rentong Guo, Xiaofei Liao, Haikun Liu, Jianhui Yue |
J. Comput. Sci. Technol. | 5 |
| 2023 | RACE: An Efficient Redundancy-aware Accelerator for Dynamic Graph Neural NetworkabstractDynamic Graph Neural Network (DGNN) has recently attracted a significant amount of research attention from various domains, because most real-world graphs are inherently dynamic. Despite many research efforts, for DGNN, existing hardware/software solutions still suffer significantly from redundant computation and memory access overhead, because they need to irregularly access and recompute all graph data of each graph snapshot. To address these issues, we propose an efficient redundancy-aware accelerator, RACE , which enables energy-efficient execution of DGNN models. Specifically, we propose a redundancy-aware incremental execution approach into the accelerator design for DGNN to instantly achieve the output features of the latest graph snapshot by correctly and incrementally refining the output features of the previous graph snapshot and also enable regular accesses of vertices’ input features. Through traversing the graph on the fly, RACE identifies the vertices that are not affected by graph updates between successive snapshots to reuse these vertices’ states (i.e., their output features) of the previous snapshot for the processing of the latest snapshot. The vertices affected by graph updates are also tracked to incrementally recompute their new states using their neighbors’ input features of the latest snapshot for correctness. In this way, the processing and accessing of many graph data that are not affected by graph updates can be correctly eliminated, enabling smaller redundant computation and memory access overhead. Besides, the input features, which are accessed more frequently, are dynamically identified according to graph topology and are preferentially resident in the on-chip memory for less off-chip communications. Experimental results show that RACE achieves on average 1139× and 84.7× speedups for DGNN inference, with average 2242× and 234.2× energy savings, in comparison with the state-of-the-art software DGNN running on Intel Xeon CPU and NVIDIA A100 GPU, respectively. Moreover, for DGNN inference, RACE obtains on average 13.1×, 11.7×, 10.4×, and 7.9× speedup and 14.8×, 12.9×, 11.5×, and 8.9× energy savings over the state-of-the-art Graph Neural Network accelerators, i.e., AWB-GCN, GCNAX, ReGNN, and I-GCN, respectively. Yu Zhang 0027, Jin Zhao 0003, Yujian Liao, Zhiying Huang, Donghao He, Lin Gu 0002, Hai Jin 0001, Xiaofei Liao, Haikun Liu, Bingsheng He, Jianhui Yue |
ACM Trans. Archit. Code Optim. | 12 |
| 2023 | Object Fingerprint Cache for Heterogeneous Memory SystemabstractHeterogeneous memory systems promise to provide both large storage capacity and high performance. Prior DRAM cache systems have metadata scalability issue and suffer from low cache hit rate, low DRAM space utilization, and significant data migration overhead. We observe that instances of an object type exhibit stable and predictable memory access patterns in its constituent cachelines and these patterns are referred to as object fingerprints. We propose the hardware-assisted cache that manages DRAM at the object type level and fetches data block at the granularity of a cacheline, by exploiting object fingerprint. To address its design challenges, we first present a software-hardware co-design to convey the software information to hardware. Second, we design multiple granularity sector caches that can be dynamically adjusted to adapt to changing behaviors and improve DRAM cache utilization. To address the challenges of large metadata storage overhead, we propose to bound possible sizes for each sector cache. Experimental results show our designs improve DRAM cache hit rate by 21.6%, boost IPC by 19.8%, and reduce data migration traffic by 51.6% on average, compared with state-of-art DRAM caches. More importantly, our online object fingerprint learning method is 2.3% inferior to the offline one in terms of IPC. Song Wu 0001, Jianhui Yue, Hai Jin 0001, Jiangqiu Shen |
IEEE Trans. Computers | 3 |
| 2023 | Cocktail: Mixing Data With Different Characteristics to Reduce Read Reclaims for nand Flash MemoryabstractA large number of read-disturb-induced rewrites are performed in the background [also known as Read Reclaim (RR)] to alleviate the read-disturb issue in NAND flash memory-based SSDs. RR can significantly degrade the performance and shorten the service life of SSD in read-intensive workloads. To address this issue, we propose a novel read-disturb management approach called Cocktail that mixes a small proportion of hot-read pages with a large proportion of cold-read pages, thereby avoiding clustering hot-read pages into a few blocks. Motivated by the insight that RR operations are frequently triggered by hot read-pages, Cocktail first prefills a portion of each block with cold data extracted from user requests. Then, Cocktail fills the prefilled blocks with write-back data caused by RR to create read-balanced blocks. We integrate two thresholds, write pool capacity and the ratio of RR-write data to User-write data, into Cocktail to govern the ratio of write-back data caused by RR to data of user requests in a block. Cocktail dynamically adjusts the two thresholds according to the characteristics of RR. Cocktail is conducive to decentralizing hot write-back data caused by RR across a broad range of blocks, thereby reducing the occurrence of second-time RR and the number of overall block reads. We compare Cocktail with three existing schemes baseline, redFTL, and IPR in terms of SSD service life, SSD response time, write amplification, and the number of garbage collections (GCs) under ten real-world workload conditions. Experimental results show that compared with the existing schemes, Cocktail reduces the number of RRs, the average response time, the 99-percentile tail latency, and the number of GCs by an average of 40.77%, 10.82%, 5.40%, and 12.29%, respectively. Cocktail also alleviates the write amplification of the three alternative schemes by an average of 49.57%. Genxiong Zhang, Yuhui Deng 0001, Yi Zhou 0009, Shujie Pang, Jianhui Yue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Accelerate Hardware Logging for Efficient Crash Consistency in Persistent MemoryabstractWhile logging has been adopted in persistent memory (PM) to support crash consistency, logging incurs severe performance overhead. This paper discovers two common factors that contribute to the inefficiency of logging: (1) load imbalance among memory banks, and (2) constraints of intra-record ordering. Over-loaded memory banks may significantly prolong the waiting time of log requests targeting these banks. To address this issue, we propose a novel log entry allocation scheme (LALEA) that reshapes the traffic distribution over PM banks. In addition, the intra-record ordering between a header and its log entries decreases the degree of parallelism in log operations. We design a log metadata buffering scheme (BLOM) that eliminates the intra-record ordering constraints. These two proposed log optimizations are general and can be applied to many existing designs. We evaluate our designs using both micro-benchmarks and real PM applications. Our experimental results show that LALEA and BLOM can achieve 54.04% and 17.16% higher transaction throughput on average, compared to two state-of-the-art designs, respectively. Jianhui Yue, Yifu Deng |
DATE | 2 |
| 2022 | FlashWalker: An In-Storage Accelerator for Graph Random WalksabstractGraph random walk is widely used in the graph processing as it is a fundamental component in graph analysis, ranging from vertices ranking to the graph embedding. Different from traditional graph processing workload, random walk features massive processing parallelisms and poor graph data reuse, being limited by low I/O efficiency. Prior designs for random walk mitigate slow I/O operations. However, the state-of-the-art random walk processing systems are bounded by slow disk I/O bandwidth, which is confirmed by our experiments with real-world graphs. To address this issue, we propose FlashWalker, an in-storage accelerator for random walk that moves walk updating close to graph data stored in flash memory, by exploiting significant parallelisms inside SSD. Featuring a heterogeneous and parallel processing system, FlashWalker includes a board-level accelerator, channel-level accelerators, and chip-level accelerators. To address challenges posed by the tight resource constraints for processing large-scale graphs, we propose novel designs: storing a few popular subgraphs in accelerators, the pre-walking for dense walks, two optimizations to search the subgraph mapping table, and a subgraph scheduling algorithm. We implement FlashWalker in RTL, showing small circuit area overhead. Our evaluation shows FlashWalker reduces the execution time of random walk algorithms by up to 660.50×, compared with GraphWalker, which is the state-of-the-art system for random walk algorithms. Fuping Niu, Jianhui Yue, Jiangqiu Shen, Xiaofei Liao, Haikun Liu, Hai Jin 0001 |
IPDPS | 2 |
| 2021 | Efficient Hardware-assisted Out-place Update for Persistent MemoryabstractShadow paging can guarantee crash consistency for Persistent Memory (PM). However, shadow paging requires the use of an address mapping table to track shadow pages, and frequent accesses to this table introduce significant performance overhead. In addition, maintaining crash consistency at the granularity level of a page causes a large amount of unnecessary write traffic. This paper proposes a novel hardware-assisted fine-grained out-place-update scheme at the granularity level of a cacheline to efficiently support crash consistency for PM. Our design fully leverages the Address Indirection Table (AIT) available in commodity PM to implement remapping. To ensure the atomicity and durability of AIT updates, we propose two policies: eager persisting and lazy persisting. We also employ overflow log to handle the eviction of speculative AIT cache entries upon an overflow in the AIT cache. Evaluation results based on multicore workloads demonstrate that our proposed scheme can improve the transaction throughput over the state-of-the-art design by 24.0% on average. Yifu Deng, Jianhui Yue |
DATE | 2 |
| 2021 | Efficient NVM Crash Consistency by Mitigating Resource ContentionabstractLogging is widely adopted to ensure crash consistency for Non-Volatile Memory (NVM) systems. However, the logging imposes significant performance overhead caused by the extra log operations and ordering constraints between the logging and in-place updates, degrading the system performance. There are some research efforts to reduce the logging overhead. Recently, LAD proposed that exploiting the non-volatility of Asynchronous DRAM Refresh (ADR) buffer can remove log operations for a transaction whose total amount of updated cachelines is smaller than the buffer capacity, ensuring crash consistency. However, on multi-core systems, concurrent transactions contend the scarce ADR buffer and frequently lead to the buffer overflow. Upon the buffer overflow, LAD resorts to logging operations for in-flight transactions, degrading the system performance. Our experiments show that LAD produces a significant number of log operations when multiple transactions run concurrently. To decrease log operations caused by LAD, this paper presents a new transaction execution scheme, called two-stage transaction execution(TSTE), which allows the write requests of a transaction to be in both the ADR buffer and the staging SRAM buffer. Our new scheme performs log operations for a transaction’s write requests in the SRAM buffer and executes in-place update operations for this transaction’s write requests in the ADR buffer. The introduced SRAM buffer can make the ADR buffer serve more update requests, reducing log operations.The evaluation results demonstrate that our proposed schemes can efficiently reduce log operations up to 39.29% and improve the transaction throughput up to 28.22% Jianhui Yue, Yifu Deng |
NAS | 2 |
| 2020 | Efficient Hardware-Assisted Crash Consistency in Encrypted Persistent MemoryabstractThe persistent memory (PM) requires maintaining the crash consistency and encrypting data, to ensure data recoverability and data confidentiality. The enforcement of these two goals does not only put more burden on programmers but also degrades performance. To address this issue, we propose a hardware-assisted encrypted persistent memory system. Specifically, logging and data encryption are assisted by hardware. Furthermore, we apply the counter-based encryption and the cipher feedback (CFB) mode encryption to data and log respectively, reducing the encryption overhead. Our primary experimental results show that the transaction throughput of the proposed design outperforms the baseline design by up to 34.4%. Zhan Zhang 0003, Jianhui Yue, Xiaofei Liao, Hai Jin 0001 |
DATE | 2 |
| 2020 | Improving the Performance of NVM Crash Consistency under MulticoreabstractNon- Volatile Memory (NVM) systems require logging to support crash consistency. However, log operations introduce severe performance overhead. Recently, LAD was proposed to eliminate log operations for some transactions in which the total amount of updated cachelines is smaller than the Asynchronous DRAM Refresh (ADR) buffer, without affecting crash consistency. Nevertheless, on multicore, concurrent transactions tend to exhaust the ADR resource and hence log operations have to be conducted in LAD. In this study, we observe that a significant number of log operations could be avoided if each transaction run alone. To eliminate these unnecessary log operations, this paper proposes virtual ADR buffers to decouple buffering from the ADR's reliable writing data. Specifically, only logless operations are allowed to access ADR resources. Additionally, this paper proposes to adopt redo log with DRAM cache to speed up the transaction commit speed. The evaluation results demonstrate that our proposed scheme can efficiently reduce log operation up to 94.9 % and improve the transaction throughput up to 78.6 %. Jianhui Yue, Yifu Deng |
ICCD | 2 |
| 2017 | Enhancing the Malloc System with Pollution Awareness for Better Cache PerformanceabstractCache pollution, by which weak-locality data unduly replaces strong-locality data, may notably degrade application performance in a shared-cache multicore machine. This paper presents NightWatch, a cache management subsystem that provides general, transparent and low-overhead pollution control to applications. NightWatch is based on the observation that data within the same memory chunk or chunks within the same allocation context often share similar locality property. NightWatch embodies this observation by online monitoring current cache locality to predict future behavior and restricting potential cache polluters proactively. We have integrated NightWatch into two popular allocators, tcmalloc and ptmalloc2. Experiments with SPEC CPU2006 show that NightWatch improves application performance by up to 45 percent (18 percent on average), with an average monitoring overhead of 0.57 percent (up to 3.02 percent). Xiaofei Liao, Rentong Guo, Hai Jin 0001, Jianhui Yue, Guang Tan |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Reducing Read Latency in MLC PCMabstractMulti-level-cell (MLC) phase-change memory (PCM) provides higher storage density at the cost of slower reads and writes. Since reads are latency critical, this paper propose a simple and effective bit mapping scheme, called Mapping Critical Word to MSBs (MCWM), to address slow reads in MLC. MCWM takes advantage of fast read speed of most-significant-bits (MSBs) of MLC cells and strips a cache line among MLC cells at the bit level. Taking 2- bit MLC as an example, MCWM stores the first half of each cache line at most-significant-bits (MSBs) of MLC cells, and the second half at least-significant-bits (LSBs). This design leverages the observation that most critical words are located within the first half of a cache line. Upon a cache miss, the critical word can be fetched at the same speed as single-level cell (SLC) PCM, thus reducing processor stall time. Experimental results under 4-cores SPEC CPU 2006 workloads show that MCWM can reduce memory read latency by 27.5% and IPC by 13.7% on average, compared with conventional PCM. In addition, MCWM outperforms recently proposed Striped PCM (SPCM) by 12.5% in latency and 6.1% in IPC on average. Additionally, MCWM is complementary to write optimizations. MCWM can reduce read latency of Write Pause by 25% and increase IPC by 11.7% on average. Jianhui Yue |
NAS | 1 |
| 2015 | SFMapReduce: An optimized MapReduce framework for Small FilesabstractHadoop, an open-source implementation of MapReduce, is widely used because of its ease of programming, scalability, and availability. With the explosive development of cloud computing, business and scientific applications increasingly take advantage of Hadoop. The sizes of files stored and processed in Hadoop are not bound to very large files anymore. However, Hadoop cannot provide stable and efficient services for small files at both storage and processing levels. To solve these problems, we propose an optimized MapReduce framework for small files, SFMapReduce. In SFMapReduce, we present two techniques, Small File Layout (SFLayout) and customized MapReduce (CMR). SFLayout is used to solve the memory problem and improve I/O performance in HDFS. CMR provides an interface for MapReduce so that SFMapReduce can process MapReduce with SFLayout efficiently. Our experimental results show that SFMapReduce decreases the memory pressure on the Hadoop NameNode, and provides better loading and retrieving throughput. On average, SFMapReduce achieves an improvement on MapReduce processing by 14.5 times and 20.8 times, compared with the original Hadoop and HAR layout. Hai Pham, Jianhui Yue, Weikuan Yu |
NAS | 3 |
| 2015 | NightWatch: Integrating Lightweight and Transparent Cache Pollution Control into Dynamic Memory Allocation Systems
Rentong Guo, Xiaofei Liao, Hai Jin 0001, Jianhui Yue, Guang Tan |
USENIX ATC | 4 |
| 2014 | Performance-energy adaptation of parallel programs in pervasive computing
Hai Jin 0001, Xiaofei Liao, Jianhui Yue |
J. Supercomput. | 4 |
| 2013 | Exploiting subarrays inside a bank to improve phase change memory performanceabstractEnabling subarrays reduces memory latency by allowing concurrent accesses to different subarrays within the same bank in the DRAM system. However, this technology has great challenges in the PCM system since an on-going write cannot overlap with other accesses due to large electric current draw for writes. This paper proposes two new mechanisms (PASAK and WAVAK) that leverage subarray-level parallelism to enable a bank to serve a write and multiple reads in parallel without violating power constraints. PASAK exploits the electric current difference between writing a bit 0 and a bit 1, and provides a new power allocation strategy that better utilizes the power budget to mitigate the performance degradation due to bank conflicts. WAVAK adds a simple coding method that inverts all bits to be written if there are more zeros than ones, with a goal to reduce electric current for writes and create larger power surplus to serve more reads if there is no subarray conflict. Experimental results under 4-cores SPEC CPU 2006 workloads show that our proposed mechanisms can reduce memory latency by 68.7% and running time by 34.8% on average, comparing with the standard PCM system. In addition, our mechanisms outperform Flip-N-Write 14.6% in latency and 8.5% in running time on average. Jianhui Yue |
DATE | 1 |
| 2013 | Accelerating write by exploiting PCM asymmetriesabstractTo improve the write performance of PCM, this paper proposes a new write scheme, called two-stage-write, which leverages the speed and power difference between writing a zero bit and writing a one bit. Writing a one takes longer time but less electrical current than writing a zero. We propose to divide a write into stages: in the write-0 stage all zeros are written at an accelerated speed, and in the write-1 stage stage, all ones are written with increased parallelism, without violating power constraints. We also present a new coding scheme to improve the speed of the write-1 stage by further increasing the number of bits that can be written to PCM in parallel. Based on simulation experiments of a multi-core processor under various SPEC CPU 2006 workloads, our proposed techniques can reduce the memory latency of standard PCM by 68.3% and improve the system performance by 33.9% on average. In addition, the proposed two-stage-write shows 16.5% latency reduction and 9.2% performance improvement over Flip-N-Write. Jianhui Yue |
HPCA | 1 |
| 2012 | Temporal characterization of SPEC CPU2006 workloads: Analysis and synthesisabstractSPEC CPU2006 benchmark suite has been extensively studied, with efforts focusing on the requirement understanding of memory workloads from the SPEC CPU2006 suite. However, characterizing SPEC CPU2006 workloads from a time dependence perspective has attracted little attention. This paper studies the auto-correlation functions of the arrival intervals of memory accesses in all SPEC CPU2006 traces, and concludes that correlations in memory inter-access times are inconsistent, either with evident correlations or with little and no correlation. Different with the studies focused on the prior suites, we present that self-similarity exists only in a small number of SPEC2006 workloads. In addition, we implement a memory access series generator in which the inputs are the measured properties of the available trace data. Experimental results show that this model can more accurately emulate the complex access arrival behaviors of real memory systems than the conventional self-similar and independent identically distributed methods, particularly the heavy-tail characteristics under both Gaussian and non-Gaussian workloads. Jianhui Yue, Bruce Segee |
IPCCC | 2 |
| 2012 | Making Write Less Blocking for Read Accesses in Phase Change MemoryabstractPhase-change Memory (PCM) is a promising alternative or complement to DRAM for its non-volatility, scalable bit density, and fast read performance. Nevertheless, PCM has two serious challenges including extraordinarily slow write speed and less-than-desirable write endurance. While recent research has improved the write endurance significantly, slow write speed become a more prominent issue and prevents PCM from being widely used in real systems. To improve write speed, this paper proposes a new memory micro-architecture, called Parallel Chip PCM(PC2M), which leverages the spatial locality of memory accesses and trades bank-level parallelism for larger chip-level parallelism. We also present a micro-write scheme to reduce the blocking for read accesses caused by uninterrupted serialized writes. Micro-write breaks a large write into multiple smaller writes and timely schedules newly arriving reads immediately after a small write completes. Our design is orthogonal to many existing PCM write hiding techniques, and thus can be used to further optimize PCM performance. Based on simulation experiments of a multi-core processor under SPEC CPU 2006 multi-programmed workloads, our proposed techniques can reduce the memory latency of standard PCM by 68.5% and improve the system performance by 30.3% on average. PC2M and Micro-write significantly outperform existing approaches. Jianhui Yue |
MASCOTS | 1 |
| 2011 | Hot Random Off-Loading: A Hybrid Storage System with Dynamic Data MigrationabstractRandom accesses are generally harmful to performance in hard disk drives due to more dramatic mechanical movement. This paper presents the design, implementation, and evaluation of Hot Random Off-loading (HRO), a self-optimizing hybrid storage system that uses a fast and small SSD as a by-passable cache to hard disks, with a goal to serve a majority of random I/O accesses from the fast SSD. HRO dynamically estimates the performance benefits based on history access patterns, especially the randomness and the hotness, of individual files, and then uses a 0-1 knapsack model to allocate or migrate files between the hard disks and the SSD. HRO can effectively identify files that are more frequently and randomly accessed and place these files on the SSD. We implement a prototype of HRO in Linux and our implementation is transparent to the rest of the storage stack, including applications and file systems. We evaluate its performance by directly replaying three real-world traces on our prototype. Experiments demonstrate that HRO improves the overall I/O throughput up to 39% and the latency up to 23%. Jianhui Yue, Zhao Cai, Bruce Segee |
MASCOTS | 3 |
| 2011 | Energy Efficient Buffer Cache Replacement for Data ServersabstractPower consumption is an increasingly impressing concern for data servers as it directly affects running costs and system reliability. Prior studies have shown that most memory space on data servers is used for buffer caching and thus cache replacement becomes critical. Two conflicting factors of buffer caching impacts memory energy efficiency: (1) a higher hit rate reduces memory traffic and thus saves energy, (2) temporally concentrating memory accesses to a smaller set of memory chips increases the chances of "free riding" through DMA overlapping and also makes more memory chips have opportunities to power down. This paper investigates the tradeoff between these two interacting, sometimes conflicting factors and proposes three energy-aware buffer cache replacement algorithms: On a cache miss for a new block b in a file f, evict an victim block from (1)the most recently accessed memory chip, (2) the memory chip that is accessed most recently by file f, or (3) the memory chip that is accessed most recently by file f and whose last access block belongs to the same hot or cold categories as block b. Simulation results based on three real-world I/O traces, including TPC-R, MSN-BEFS and Exchange, show that our algorithms can save up to 24.9% energy with marginal degradation in hit rates. Our algorithms show degradation in response time in some experiments. We propose an off-line energy sub optimal replacement algorithm that serves as a theortical reference. Jianhui Yue, Zhao Cai |
NAS | 1 |
| 2010 | Energy and thermal aware buffer cache replacement algorithmabstractPower consumption is an increasingly impressing concern for data servers as it directly affects running costs and system reliability. Prior studies have shown most memory space on data servers are used for buffer caching and thus cache replacement becomes critical. Temporally concentrating memory accesses to a smaller set of memory chips increases the chances of free riding through DMA overlapping and also enlarges the opportunities for other ranks to power down. This paper proposes a power and thermal-aware buffer cache replacement algorithm. It conjectures that the memory rank that holds the most amount of cold blocks are very likely to be accessed in the near future. Choosing the victim block from this rank can help reduce the number of memory ranks that are active simultaneously. We use three real-world I/O server traces, including TPC-C, LM-TBF and MSN-BEFS to evaluate our algorithm. Experimental results show that our algorithm can save up to 27% energy than LRU and reduce the temperature of memory up to 5.45°C with little or no performance degradation. Jianhui Yue, Zhao Cai |
MSST | 1 |
| 2008 | Impacts of Indirect Blocks on Buffer Cache Energy EfficiencyabstractIndirect blocks, part of a file's metadata used for locating this file's data blocks, are typically treated indistinguishably from file's data blocks in buffer cache. This paper shows that this conventional approach will significantly detriment the overall energy efficiency of memory systems. Scattering small but frequently accessed indirected blocks over allmemory chips reduce the energy saving opportunities. We propose a new energy-efficient buffer cache management scheme, named MEEP, which separates indirect and datablocks into different memory chips. Our trace-driven simulation results show that our new scheme can save memory energy up to 16.8% and 15.4% in the I/O-intensive server workloads TPC-R and TPC-H, respectively. Jianhui Yue, Zhao Cai |
ICPP | 1 |
| 2008 | An Energy-Efficient Buffer Cache Replacement
Jianhui Yue, Zhao Cai |
MASCOTS | 1 |
| 2008 | An Energy-Oriented Evaluation of Buffer Cache Algorithms Using Parallel I/O WorkloadsabstractPower consumption is an important issue for cluster supercomputers as it directly affects running cost and cooling requirements. This paper investigates the memory energy efficiency of high-end data servers used for supercomputers. Emerging memory technologies allow memory devices to dynamically adjust their power states and enable free rides by overlapping multiple DMA transfers from different I/O buses to the same memory device. To achieve maximum energy saving, the memory management on data servers needs to judiciously utilize these energy-aware devices. As we explore different management schemes under five real-world parallel I/O workloads, we find that the memory energy behavior is determined by a complex interaction among four important factors: (1) cache hit rates that may directly translate performance gain into energy saving, (2) cache populating schemes that perform buffer allocation and affect access locality at the chip level, (3) request clustering that aims to temporally align memory transfers from different buses into the same memory chips, and (4) access patterns in workloads that affect the first three factors. Jianhui Yue, Zhao Cai |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2007 | Evaluating memory energy efficiency in parallel I/O workloadsabstractPower consumption is an important issue for cluster supercomputers as it directly affects their running cost and cooling requirements. This paper investigates the memory energy efficiency of high-end data servers used for supercomputers. Emerging memory technologies allow memory devices to dynamically adjust their power states. To achieve maximum energy saving, the memory management on data servers needs to judiciously utilize these energy-aware devices. As we explore different management schemes under four real-world parallel I/O workloads, we find that the memory energy consumption is determined by a complex interaction among four important factors: (1) cache hit rates that may directly translate performance gain into energy saving, (2) cache populating schemes that perform buffer allocation and affect access locality at the chip level, (3) request clustering that aims to temporally align memory transfers from different buses into the same memory chips, and (4) access patterns in workloads that affect the first three factors. Jianhui Yue, Zhao Cai |
CLUSTER | 1 |
| 2003 | HARTs: high availability cluster architecture with redundant TCP stacksabstractImproving the availability of services of is a key issue for survivability of a cluster system. Lots of schemes are proposed for this purpose. But most of them aim at enhancing only the service-level availability or application specific. In this paper, we propose a scheme called High Availability with Redundant TCP Stacks (HARTs), providing connection-level availability by exploring the redundant TCP stacks for TCP connections at the server side. We present our performance experiment results on our HA cluster prototype. From results, we find the configuration of one primary server with one backup server running on separated 100 Mbps Ethernet has acceptable performance to support the server side applications while delivering high availability. Zhiyuan Shao, Hai Jin 0001, Jie Xu 0006, Jianhui Yue |
IPCCC | 5 |