Haixia Wang 0001

dblp:83/1575-1 · also Hai-Xia Wang 0001 · DBLP profile ↗
← Back
35ranked-venue papers
1as first author
9since 2021 · last 2026
0009-0008-0474-5030ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 1 first-author · 6 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 XcptProof: Formal Verification of CPU Exception Transient Execution Security via Leakage Contracts
abstract
Transient execution attacks triggered by CPU exceptions, such as Meltdown and MDS-type attacks, have compromised system security. Unfortunately, no prior work has conducted a formalized analysis of the CPU exception transient execution security.
Yujia Zhang 0018, Kexin Gong, Hongpeng Wang 0002, Haixia Wang 0001, Dongsheng Wang 0002
ACM Great Lakes Symposium on VLSI5
2026 OCCUPY+PROBE: Cross-Privilege Branch Target Buffer Side-Channel Attacks at Instruction Granularity
Kaiyuan Rong, Junqi Fang, Haixia Wang 0001, Dapeng Ju, Dongsheng Wang 0002
NDSS3
2026 Exploiting ARMeD Channels By Reverse Engineering ARM Memory Disambiguation Unit
abstract
ARM CPUs are widely used in both embedded systems and personal computers where security considerations are becoming important. Evidently, vulnerabilities on hardware components such as cache and translation look-aside buffer are well-documented. But there are much less studies on other components, especially those in the CPU backend, largely due to the unavailability of their design and implementation details. To address this gap, we present the first in-depth reverse engineering analysis of the Memory Disambiguation Unit (MDU) in the backend of ARM CPUs. Across four microarchitectures from ARM and Apple CPUs, we identify two different MDU designs, switch-based and counter-based. We then analyze the state machine, selection mechanism, and organization of these MDU designs. We further propose new side channels and covert channels, which we call ARMeD channels, that exploit ARM MDU to leak information. We demonstrate with three attacks using ARMeD channels: a cross-process covert channel, website fingerprinting, and a new implementation of the Spectre attack. Finally, we present a defense strategy against ARMeD Channels with less than 3% degradation on the MDU’s prediction accuracy.
Chang Liu 0117, Zhouyang Li, Haixia Wang 0001, Pengfei Qiu, Gang Qu 0001, Dongsheng Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 SCAFinder: Formal Verification of Cache Fine-Grained Features for Side Channel Detection
abstract
Recent research has unveiled numerous cache-timing side-channel attacks exploiting the side effects of fine-grained cache features, such as coherence protocol and prefetch, among others. Traditional modeling methods and verification techniques are insufficient for verifying caches with fine-grained features and detecting cache timing vulnerabilities. There is a necessity for comprehensive verification of such complex cache designs. This paper presents SCAFinder, a verification framework targeting the cache designs with fine-grained features; it identifies cache side-channel attacks through model checking techniques. Specifically, it proposes a modeling methodology for cache designs that enables us to abstract the cache’s behavior and latency characteristics. We implement a search algorithm for finding all counterexamples based on open-source model checking software. Subsequently, we add an attack scenario analysis module to discover attacks applicable to specific scenarios. We evaluate SCAFinder on Intel Skylake-X microarchitecture, demonstrating its capability to generate 7 new attack sequences exploiting coherence protocol and prefetch, and 12 new replacement policy-based side channels. As a case study, we successfully built a covert channel for one of the sequences on the real-world processor. To the best of our knowledge, we are the first to implement cross-core replacement policy-based attacks on non-inclusive caches.
Haixia Wang 0001, Pengfei Qiu, Yongqiang Lyu 0001, Hongpeng Wang 0002, Dongsheng Wang 0002
IEEE Trans. Inf. Forensics Secur.2
2024 Lightning: Leveraging DVFS-induced Transient Fault Injection to Attack Deep Learning Accelerator of GPUs
abstract
Graphics Processing Units (GPU) are widely used as deep learning accelerators because of its high performance and low power consumption. Additionally, it remains secure against hardware-induced transient fault injection attacks, a classic type of attacks that have been developed on other computing platforms. In this work, we demonstrate that well-trained machine learning models are robust against hardware fault injection attacks when the faults are generated randomly. However, we discover that these models have components, which we refer to as sensitive targets, that are vulnerable to faults. By exploiting this vulnerability, we propose the Lightning attack, which precisely strikes the model’s sensitive targets with hardware-induced transient faults based on the Dynamic Voltage and Frequency Scaling (DVFS). We design a sensitive targets search algorithm to find the most critical processing units of Deep Neural Network (DNN) models determining the inference results, and develop a genetic algorithm to automatically optimize the attack parameters for DVFS to induce faults. Experiments on three commodity Nvidia GPUs for four widely-used DNN models show that the proposed Lightning attack can reduce the inference accuracy by 69.1% on average for non-targeted attacks, and, more interestingly, achieve a success rate of 67.9% for targeted attacks.
Rihui Sun, Pengfei Qiu, Yongqiang Lyu 0001, Jian Dong 0010, Haixia Wang 0001, Dongsheng Wang 0002, Gang Qu 0001
ACM Trans. Design Autom. Electr. Syst.5
2023 Leaky MDU: ARM Memory Disambiguation Unit Uncovered and Vulnerabilities Exposed
abstract
Memory Disambiguation Unit (MDU) is widely used on modern processors to speculatively execute load instructions and improve pipeline performance. Given that the MDU design details on ARM processors are not available to the public, it is unclear whether there are any security vulnerabilities associated with its MDU. In this paper, we first reverse engineer the undocumented features of ARM MDU, then we discover three potential user-privilege attacks to leak secret data via MDU: cross-process attack that allows users to communicate through a convert channel, cross-domain attack that leaks kernel information and a new variant of inner-process and inter-processes Spectre attacks. These attacks pose serious security challenges as they can bypass both all the known countermeasures against cache side-channel attacks and those against transient execution attacks. Potential mitigation against the proposed MDU-based attacks are also discussed.
Chang Liu 0117, Yongqiang Lyu 0001, Haixia Wang 0001, Pengfei Qiu, Dapeng Ju, Gang Qu 0001, Dongsheng Wang 0002
DAC3
2022 CacheGuard: A Behavior Model Checker for Cache Timing Side-Channel Security: (Invited Paper)
abstract
Defending cache timing side-channels has become a major concern in modern secure processor designs. However, a formal method that can completely check if a given cache design can defend against timing side-channel attacks is still absent. This study presents CacheGuard, a behavior model checker for cache timing side-channel security. Compared to current state-of-the-art prose rule-based security analysis methods, CacheGuard covers the whole state space for a given cache design to discover unknown side-channel attacks. Checking results on standard cache and state-of-the-art secure cache designs discovers 5 new attack strategies, and potentially makes it possible to develop a timing side channel-safe cache with the aid of CacheGuard.
Lingfeng Yin, Yongqiang Lyu 0001, Haixia Wang 0001, Gang Qu 0001, Dongsheng Wang 0002
ASP-DAC4
2022 SSB-Tree: Making Persistent Memory B+- Trees Crash-Consistent and Concurrent by Lazy-Box
abstract
The 8-byte granularity of failure-atomicity brings two challenges to Persistent Memory (PM) B+-tree designs. The first is how to insert a key in a sorted node atomically. The second is how to atomically perform structural modification operations that involve multiple nodes, such as node splits and merges. The majority of current designs either have huge consistency cost or compromise recovery time. In this paper, we propose an in-node logging technology, named Lazy-Box, to update multiple parts of a node without exposing intermediate states. Lazy-Box aggregates successive modifications and selectively uses Copy-on-Write (CoW) to reduce consistency cost. Based on Lazy-Box, we propose a new variant of B+-tree on PM, named Side-to-Side B+-Tree (SSB- Tree). SSB- Tree relaxes the tree-structure requirement so that node splits and merges only modify one node. Taking advantage of Lazy-Box, all modification operations of SSB-Tree are committed through a single 8-byte write. Therefore, SSB- Tree not only enables efficient concurrency protocol but also achieves instant recovery. Last but most important, SSB- Tree doubles the node space and reuses the extra space to avoid expensive node allocations when performing CoW. Our experimental results show that SSB-Tree achieves up to 29%, 40%, and 14% higher throughput in Insert, Delete, and Scan benchmark respectively than other state-of-the-art PM B+-trees.
Tongliang Li, Haixia Wang 0001, Airan Shao, Dongsheng Wang 0002
IPDPS2
2022 DynaComm: Accelerating Distributed CNN Training Between Edges and Clouds Through Dynamic Communication Scheduling
abstract
To reduce uploading bandwidth and address privacy concerns, deep learning at the network edge has been an emerging topic. Typically, edge devices collaboratively train a shared model using real-time generated data through the Parameter Server framework. Although all the edge devices can share the computing workloads, the distributed training processes over edge networks are still time-consuming due to the parameters and gradients transmission procedures between parameter servers and edge devices. Focusing on accelerating distributed Convolutional Neural Networks (CNNs) training at the network edge, we present DynaComm, a novel scheduler that dynamically decomposes each transmission procedure into several segments to achieve optimal layer-wise communications and computations overlapping during run-time. Through experiments, we verify that DynaComm manages to achieve optimal layer-wise scheduling for all cases compared to competing strategies while the model accuracy remains untouched.
Shangming Cai, Dongsheng Wang 0002, Haixia Wang 0001, Yongqiang Lyu 0001, Guangquan Xu, James Xi Zheng, Athanasios V. Vasilakos
IEEE J. Sel. Areas Commun.3
2020 CRaft: An Erasure-coding-supported Version of Raft for Reducing Storage Cost and Network Cost
Zizhong Wang, Tongliang Li, Haixia Wang 0001, Airan Shao, Yunren Bai, Shangming Cai, Dongsheng Wang 0002
FAST3
2020 An Adaptive Erasure-Coded Storage Scheme with an Efficient Code-Switching Algorithm
abstract
Using erasure codes increases consumption of network traffic and disk I/O tremendously when systems recover data, resulting in high latency of degraded reads. In order to mitigate this problem, we present an adaptive storage scheme based on data access skew, a fact that most data accesses are applied in a small fraction of data. In this scheme, we use both Local Reconstruction Code (LRC), whose recovery cost is low, to store frequently accessed data, and Hitchhiker (HH) code, which guarantees minimum storage cost, to store infrequently accessed data. Besides, an efficient switching algorithm between LRC and HH code with low network and computation costs is provided. The whole system will benefit from low degraded read latency while keeping a low storage overhead, and code-switching will not become a bottleneck.
Zizhong Wang, Haixia Wang 0001, Airan Shao, Dongsheng Wang 0002
ICDCS2
2020 CARD: A Congestion-Aware Request Dispatching Scheme for Replicated Metadata Server Cluster
abstract
Replicated metadata server cluster (RMSC) is highly efficient to be used in distributed filesystems while facing data-driven scenarios (e.g., massive-scale distributed machine learning tasks). Yet, when considering cost-effectiveness and system utilization, the cluster scale is commonly restricted in practice. Within this context, servers in the cluster start to suffer from load-oscillations at higher system utilization due to clients’ congestion-unaware behaviors and unintelligent selection strategies (i.e., servers in the cluster are preferred then evaded intermittently). The consequences brought by load-oscillations degrade the overall performance of the whole system to some extent. One solution to tackle this problem is having clients share a part of the responsibility and behave more wisely for stability concerns. So in this paper, we present a Congestion-Aware Request Dispatching scheme, CARD, which is mainly conducted at clients and directed by a rate control mechanism. Through extensive experiments, we verify that CARD is highly efficient in resolving load-oscillations in RMSC. Apart from this, our results show that RMSC with our congestion-aware based optimization achieves better scalability compared to previous implementations under targeted workloads, especially in heterogeneous environments.
Shangming Cai, Dongsheng Wang 0002, Zhanye Wang, Haixia Wang 0001
ICPP4
2020 An Adaptive Erasure-Coded Storage Scheme with an Efficient Code-Switching Algorithm
abstract
Many distributed storage systems use erasure codes rather than replication for higher reliability at significantly lower storage costs. However, using traditional erasure codes increases consumption of network traffic and disk I/O tremendously when systems recover data, resulting in high latency of degraded reads. In order to mitigate this problem, we present an adaptive storage scheme based on data access skew, a fact that most data accesses are applied in a small fraction of data. In this scheme, we use both a Local Reconstruction Code (LRC) to store frequently accessed data, and a Hitchhiker (HH) code to store infrequently accessed data. Besides, an efficient switching algorithm between LRC and HH code with low network and computation costs is provided. The whole system will benefit from low degraded read latency while keeping a low storage overhead, and code-switching will not become a bottleneck. Experimental evaluation shows that this adaptive storage scheme’s performance was in line with expectations.
Zizhong Wang, Haixia Wang 0001, Airan Shao, Dongsheng Wang 0002
ICPP2
2020 Parallelizing and optimizing neural Encoder-Decoder models without padding on multi-core architecture
Yuchen Qiao, Kazuma Hashimoto, Akiko Eriguchi, Haixia Wang 0001, Dongsheng Wang 0002, Yoshimasa Tsuruoka, Kenjiro Taura
Future Gener. Comput. Syst.4
2019 Fast Recovery Techniques for Erasure-coded Clusters in Non-uniform Traffic Network
abstract
Nowadays many practical systems adopt erasure codes to ensure reliability and reduce storage overhead. However, erasure codes also bring in low recovery performance. The network links in practice, such as peer-to-peer and cross-data network, always have nonuniform bandwidth because of various reasons. To reduce recovery time, we propose Parallel Pipeline Tree (PPT) and Parallel Pipeline Cross-Tree (PPCT) to speed up single-node and multiple-node recovery in non-uniform traffic network environment, respectively. By utilizing bandwidth gap among links, PPT constructs a tree path based on bandwidth and pipelines the data in parallel. By sharing traffic pressure of requesters with helpers, PPCT constructs a tree-like path and pipelines the data in parallel without additional helpers. We also theoretically explain the effect of PPT and PPCT used in uniform network environment. The experiments implemented on geo-distributed Amazon EC2 show that the time reduction reaches up to 37.2% with PPCT over traditional technique and reaches up to 89.2%, 76.4% and 21.6% with PPT over traditional technique, Partial-Parallel-Repair and Repair Pipelining, respectively. PPT and PPCT significantly improve the performance of erasure codes' recovery.
Yunren Bai, Haixia Wang 0001, Dongsheng Wang 0002
ICPP3
2017 GDCRT: In-Memory 2D Geographical Dynamic Cascading Range Tree
Yinxing Hou, Haixia Wang 0001, Dongsheng Wang 0002
APPT2
2017 Data-centric computation mode for convolution in deep neural networks
abstract
Deep Convolutional Neural Network (CNN) based methods have shown outstanding performance in a wide range of applications. Nowadays neural networks become deeper, leading to demand of substantial computation and memory resources. Customized hardware is one option which maintains high performance in lower energy consume than general CPUs or GPUs. While hardware designing, we need to address the problem of massive data transmission, and ensure high throughput at the same time. Actually, substantial data transfer consumes more energy than computation. The larger scales of neural networks become, the harder to solve this problem. In this paper, we focus on convolution operation, which occupies nearly 90% computation and runtime in deep CNN. We propose a data-centric computation mode for convolution, which declines the total requirements of data transfer during convolution processing period efficiently, and utilizes data locality to achieve high throughput. Different from previous methods, which adopt efficient on-chip memory hierarchy or focus on partial results' movements, our proposed method concentrates on operands themselves in convolution, minimizing data transfer right from the start. Obviously, it can be combined with others to achieve higher energy efficient. Furthermore, we also simulate and analyse the hardware overhead of our data-centric convolution, corroborating its potentiality of performing high throughput in low energy consumption.
Peiqi Wang 0001, Zhenyu Liu 0001, Haixia Wang 0001, Dongsheng Wang 0002
IJCNN3
2016 Pull-off buffer: Borrowing cache space to avoid deadlock for fault-tolerant NoC routing
abstract
Advances in semiconductor technology have led to large chip multiprocessor (CMP) employing network-on-chip (NoC) to provide scalable on-chip communication. This higher integration capacity, on the other hand, increases the possibility of faults. To tackle this challenge, fault-tolerant routing in NoC becomes essential, which allows packets to be routed around faulty network components and maintains normal communication. However, to tolerate a large number of faults, the deadlock problem becomes very difficult to deal with. Existing highly fault-tolerant routing solutions employ virtual channel (VC) or topology-agnostic routing for deadlock avoidance, but at the cost of lower network performance and the demand for extra hardware. In this paper, we show that it is possible to design a novel highly fault-tolerant routing method without VC and topology-agnostic routing. We present pull-off buffer (POB), a FIFO buffer borrowing the space already present in cache, to avoid potentially existing deadlocks. POBs borrow cache space only from selected nodes and only after the occurrence of faults. The space of caches at other nodes will not be affected. Experimental results show that our solution can provide 2x to 3x higher network throughput and reduce router area and power overhead, when compared against existing highly fault-tolerant routing methods employing VC or topology-agnostic routing.
Airan Shao, Dongsheng Wang 0002, Haixia Wang 0001
ICCD3
2013 Data Access Type Aware Replacement Policy for Cache Clustering Organization of Chip Multiprocessors
Chongmin Li, Dongsheng Wang 0002, Haixia Wang 0001, Guohong Li, Yibo Xue
APPT3
2012 Proximity-Aware cache Replication
abstract
We propose Proximity-Aware cache Replication (PAR), an LLC replication technique that elegantly integrates an intelligent cache replication placement mechanism and a hierarchical directory-based coherence protocol into one cost-effective and scalable design. PAR dynamically allocates replicas of either shared or private data to a few predefined and fixed locations that are calculated at chip design time. Therefore, PAR fits well to future many-core CMPs thanks to its scalable on-chip storage and coherence design. Simulation results on a 64-core CMP show that PAR can achieve 12% speedup over the baseline shared cache design with SPLASH2 and PARSEC workloads. It also provides around 5% speedup over a couple contemporary approaches with much simpler and scalable support. Translating this speedup to cache performance, PAR achieves 40% and 70% reduction over the baseline in average L1 miss latency and on-chip network traffic, respectively. Furthermore, PAR shows good speedup with multiprogrammed workloads.
Chongmin Li, Dongsheng Wang 0002, Haixia Wang 0001, Yibo Xue, Jian Li 0059
ASP-DAC3
2012 Dynamic reusability-based replication with network address mapping in CMPs
abstract
In a Chip MultiProcessor(CMP) with shared caches, the last level cache is distributed across all the cores. This increases the on-chip communication delay and thus influence the processor's performance. Replication can be provided in shared caches to reduce the on-chip communication delay. However, current proposals do not take into account replicating blocks's access characteristics and how to make the best of replicas, which have limited performance benefit. In this paper, we observe that reusability of cache blocks influences the availability of replication scheme severely. Based on this observation, we propose Dynamic Reusability-based Replication (DRR), a novel cache design to exploit efficient replicas management using blocks's reuse pattern. DRR monitors the recent referenced cache blocks' access pattern, and replicates the blocks with high reusability to appropriate L2 slices, and the replicated copies can be shared by their nearby cores. We evaluate DRR for 16-core system using splash-2 and parsec benchmarks. DRR improves performance by 30% on average over conventional shared cache design, 16% over Victim Replication(VR), 8% over Adaptive Selected Replication (ASR), and 25% over R-NUCA.
Jinglei Wang, Dongsheng Wang 0002, Haixia Wang 0001, Yibo Xue
ASP-DAC3
2012 Wear-Resistant Hybrid Cache Architecture with Phase Change Memory
abstract
Phase-change Random Access Memory (PRAM) is one of the most promising technologies among emerging non-volatile memory technologies, which provides many benefits, such as high density, non-volatility and low leakage power. However, the limited write endurance of PRAM prevents it from being used as a drop-in replacement of SRAM cache. Moreover, the inherent high latency and power dissipation of write operations are both hindrances that PRAM faces. In this paper, we study the L2 cache write operations incurred by different types of data, and accordingly, propose Wear-Resistant Hybrid Cache Architecture (WRHCA), in which the write access behavior of the hybrid L2 cache, that is composed of SRAM and PRAM, is optimized. Through the prediction of data access patterns, the proposed WRHCA prevents write-prone data from entering PRAM L2 cache, and consequently, the wear-out issue of PRAM is alleviated efficiently. Experimental results on the basis of the trace-driven simulator demonstrate that, as compared to the baseline system with pure PRAM L2 cache, our optimized WRHCA saved 85.5% write operations to PRAM on average, and boosted the performance by the averaged 6.4% CPI reduction. Last but not least, as compared with the primitive 3-level SRAM cache with the same chip area, our WRHCA achieved 60.9% reduction in terms of power consumption.
Sanchuan Guo, Zhenyu Liu 0001, Dongsheng Wang 0002, Haixia Wang 0001, Guohong Li
NAS4
2011 Scalable Proximity-Aware Cache Replication in Chip Multiprocessors
abstract
We propose Proximity-Aware cache Replication (PAR), an LLC replication technique that elegantly integrates an intelligent cache replication placement mechanism and a hierarchical directory-based coherence protocol into one cost-effective and scalable design. Simulation results on a 64-core CMP show that PAR can achieve 12% speedup over the baseline shared cache design with SPLASH2 and PARSEC workloads. It also provides around 5% speedup over a couple contemporary approaches with much simpler and scalable support.
Chongmin Li, Haixia Wang 0001, Yibo Xue, Dongsheng Wang 0002, Jian Li 0059
PACT2
2011 Enhanced Adaptive Insertion Policy for Shared Caches
Chongmin Li, Dongsheng Wang 0002, Yibo Xue, Haixia Wang 0001, Xi Zhang 0008
APPT4
2011 A Read-Write Aware Replacement Policy for Phase Change Memory
Xi Zhang 0008, Dongsheng Wang 0002, Chongmin Li, Haixia Wang 0001
APPT5
2011 High performance cache block replication using re-reference probability in CMPs
abstract
In a Chip Multiprocessor(CMP) with shared caches, the last level cache (LLC) is distributed across all the cores. This increases the on-chip communication delay and thus influence the pr ocessor's performance. The LLC is also quite inefficient due to plenty of dead blocks. Replication can be provided in shared caches by replicating cache blocks evicted from cores to the local LLC slices to minimize access latency through utilizing the cache space of dead blocks which will not be referenced again before they are evicted. However, naively allowing all evicted blocks to be replicated have limited performance benefit as such replicating does not take into account reuse probability of replicated blocks. This paper proposes Adaptive Probability Replication (APR), a mechanism that counts each block's accesses in L2 cache slices, and monitors the number of evicted blocks with different number of accesses, to estimate the Re-Reference Probability of blocks in their lifetime at runtime. Using predicted re-reference probability, APR adopts probability replication policy and probability insertion policy to replicate blocks at corresponding probabilities, and insert them at appropriate position, according to their re-reference probability. We evaluate APR for a 16-core tiled CMP using splash-2 and parsec benchmarks. APR improves performance by 21% on average compared to conventional shared cache design, by 17% over Victim Replication (VR), by 10% over Adaptive Selective Replication (ASR), and by 15% over Reactive NUCA (R-NUCA). The additional hardware cost of APR is well under 1% of L2 cache slice.
Jinglei Wang, Dongsheng Wang 0002, Haixia Wang 0001, Yibo Xue
HiPC3
2011 Coherent Temporal Streams in PARSEC
abstract
Off-Chip miss latency remains a bottleneck even in the modern chip multiprocessors. Recent research advocates memory streaming techniques to alleviate the performance bottleneck caused by the high latencies of off-chip memory accesses. Memory streaming prefetches data by predicting recurring sequences of misses. Coherent read misses are one of the contributors in off-chip read misses for multithread workloads running on multi-chip multiprocessors. In this paper, we investigate off-chip coherent read misses using information-theoretic analysis of coherent read misses collected using execution-driven simulation of Princeton Application Repository for Shared-Memory Computers PARSEC). We found that coherent read misses recur system-wide in the same order forming sequences of two or more misses called streams. We show that 54% to 95% misses are part of streams. Our investigations have proved to be useful that streaming using multiple recurrences of the same stream cannot predict certain fraction of stream misses i.e. at least 10% to 60%. We demonstrate that 80% to 90% streams have length less than or equal to 8, 45% to 60% streams recur two times and then never repeat, finally streams recur after hundreds or thousands of misses.
Muhammad Abid Mughal, Haixia Wang 0001, Dongsheng Wang 0002
NAS2
2010 Fast Hierarchical Cache Directory: A Scalable Cache Organization for Large-Scale CMP
abstract
As more processing cores are integrated into one chip and the feature size continues to shrink, the increasing on-chip access latency complicates the design of the on-chip last-level cache for chip multiprocessors. At the same time, the overhead of maintaining on-chip directory cannot be ignored as the number of processing cores increasing. There is an urgent need for scalable organization of on-chip last-level cache. In this work, we propose fast hierarchical cache directory for tiled CMP, which divides CMP tiles into multiple regions hierarchically, and combines it with data replication. Multi-level directory is used to record the share information within a region and assist the regional home node to complete operation efficiently. Fast directory is used to get lower L2 slice access latency at the same time. Most cache requests to last-level cache can be handled within the local level-1 region. Evaluation indicates this architecture is highly scalable. Simulation results show that for a 16-core CMP, hierarchical cache directory reduces average access latency to last-level cache by 46.35% and average on-chip network traffic by 19.25% respectively. The system performance is increased by 20.82% at the same time.
Chongmin Li, Haixia Wang 0001, Yibo Xue, Xi Zhang 0008, Dongsheng Wang 0002
NAS2
2010 A Cache Replacement Policy Using Adaptive Insertion and Re-reference Prediction
abstract
Previous research shows that LRU replacement policy is not efficient when applications exhibit a distant re-reference interval. Recently proposed RRIP policy improves performance for such workloads. However, RRIP lacks of access recency information, which may confuse the replacement policy to make accurate prediction. Consequently, RRIP is not robust for recency-friendly workloads. This paper proposes an Adaptive Insertion and Re-reference Prediction (AI-RRP) policy which evicts data based on both re-reference prediction value and the access recency information. To make the replacement policy more adaptive across different workloads and different phases during execution, Dynamic AI-RRP (DAI-RRP) is proposed which adjusts the insertion position and prediction value for different access patterns. Simulation results show DAI-RRP reduces CPI over LRU and Dynamic RRIP by an average of 8.3% and 4.1% respectively on a single-core processor with a 1MB 16-way set last-level cache (LLC). Evaluations on quad-core CMP with a 4MB shared LLC show that DAI-RRP outperforms LRU and Dynamic RRIP (DRRIP) on the weighted speedup metric by an average of 13.2% and 26.7% respectively. Furthermore, compred to LRU, DAI-RRP requires similar hardware, or even less hardware for high-associativity cache.
Xi Zhang 0008, Chongmin Li, Haixia Wang 0001, Dongsheng Wang 0002
SBAC-PAD3
2010 Hierarchical Cache Directory for CMP
Songliu Guo, Haixia Wang 0001, Yibo Xue, Chong-Min Li, Dongsheng Wang 0002
J. Comput. Sci. Technol.2
2010 CCNoC: Cache-Coherent Network on Chip for Chip Multiprocessors
Jinglei Wang, Yibo Xue, Haixia Wang 0001, Chong-Min Li, Dongsheng Wang 0002
J. Comput. Sci. Technol.3
2009 An Efficient Lightweight Shared Cache Design for Chip Multiprocessors
Jinglei Wang, Dongsheng Wang 0002, Yibo Xue, Haixia Wang 0001
APPT4
2009 A Novel Cache Organization for Tiled Chip Multiprocessor
Xi Zhang 0008, Dongsheng Wang 0002, Yibo Xue, Haixia Wang 0001, Jinglei Wang
APPT4
2009 Network caching for Chip Multiprocessors
abstract
The large working sets of commercial and scientific workloads favor a shared L2 cache design that maximizes the aggregate cache capacity and minimizes off-chip memory requests in Chip Multiprocessors (CMP). There are two important hurdles that restrict the scalability of these chip multiprocessors: the on-chip memory cost of directory and the long L1 miss latencies. This work presents network caching architecture aimed at facing these two important problems. Network caching takes advantage of on-chip networks to manage shared data blocks and directory information in chip multiprocessors. The network caching architecture removes the directory structure from shared L2 caches and stores directory information for the blocks recently cached by L1 caches in the network interface components decreasing on-chip directory memory overhead and improves the scalability. The saved memory space is used as shared data caches or victim caches which are embedded into the network interface components to reduce L1 miss latencies further. This paper develops three network caching designs to reduce L1 miss latencies. The proposed architecture is evaluated based on simulations of a 16-core tiled CMP. First, we demonstrate that network caching architecture provides good scalability. Second, network caching architecture also provides robust performance. Third, different network caching designs have distinct impacts on performance of CMP. Against over the traditional shared L2 cache design, Network Victim Cache (NVC) design improves performance by 23% on average, and up to 34% at best. Network Shared Cache (NSC) design provides performance improvement by 6% on average, and up to 16% at best. Network Directory Cache (NDC) design achieves performance improvement by 4% on average, and up to 11% at best.
Jinglei Wang, Yibo Xue, Haixia Wang 0001, Dongsheng Wang 0002
IPCCC3
2007 Exploit Temporal Locality of Shared Data in SRC Enabled CMP
Haixia Wang 0001, Dongsheng Wang 0002, Peng Li 0031, Jinglei Wang, XianPing Fu
NPC1