EDBT 2026 Demo / reviewers in the wild / expert
Lixin Zhang 0002
dblp:52/5615-2
· DBLP profile ↗
50ranked-venue papers
4as first author
3since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 8Applied, interdisciplinary, general and emerging computing · 6Computer networks · 3Artificial intelligence and machine learning · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | HyBP: Hybrid Isolation-Randomization Secure Branch PredictorabstractRecently exposed vulnerabilities reveal the necessity to improve the security of branch predictors. Branch predictors record history about the execution of different processes, and such information from different processes are stored in the same structure and thus accessible to each other. This leaves the attackers with the opportunities for malicious training and malicious perception. Physical or logical isolation mechanisms such as using dedicated tables and flushing during context-switch can provide security but incur non-trivial costs in space and/or execution time. Randomization mechanisms incurs the performance cost in a different way: those with higher securities add latency to the critical path of the pipeline, while the simpler alternatives leave vulnerabilities to more sophisticated attacks.This paper proposes HyBP, a practical hybrid protection and effective mechanism for building secure branch predictors. The design applies the physical isolation and randomization in the right component to achieve the best of both worlds. We propose to protect the smaller tables with physically isolation based on (thread, privilege) combination; and protect the large tables with randomization. Surprisingly, the physical isolation also significantly enhances the security of the last-level tables by naturally filtering out accesses, reducing the information flow to these bigger tables. As a result, key changes can happen less frequently and be performed conveniently at context switches. Moreover, we propose a latency hiding design for a strong cipher by precomputing the "code book" with a validated, cryptographically strong cipher. Overall, our design incurs a performance penalty of 0.5% compared to 5.1% of physical isolation under the default context switching interval in Linux. Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Xuehai Qian, Lixin Zhang 0002, Dan Meng 0002 |
HPCA | 6 |
| 2021 | A Lightweight Isolation Mechanism for Secure Branch PredictorsabstractRecently exposed vulnerabilities reveal that branch predictors shared by different processes leave the attackers with the opportunities for malicious training and perception. Instead of flush-based or physical isolation of hardware resources, we want to achieve isolation of the content in these hardware tables with some lightweight processing using randomization as follows. (1) Content encoding. We propose to use hardware-based thread-private random numbers to encode the contents of the branch predictor tables. It achieves a similar effect of logical isolation but adds little in terms of space or time overheads. (2) Index encoding. We propose a randomized index mechanism of the branch predictor. This disrupts the correspondence between the branch instruction address and the branch predictor entry, thus increases the noise for malicious perception attacks. Our analyses using an FPGA-based RISC-V processor prototype and additional auxiliary simulations suggest that the proposed mechanisms incur a very small performance cost while providing strong protection. Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Jiazhen Li, Lixin Zhang 0002, Xuehai Qian, Dan Meng 0002 |
DAC | 6 |
| 2021 | Exploiting Security Dependence for Conditional Speculation Against Spectre AttacksabstractSpeculative execution side-channel vulnerabilities such as Spectre reveal that conventional architecture designs lack security consideration. This article proposes a software transparent defense framework, named as Conditional Speculation, against Spectre vulnerabilities found on traditional out-of-order microprocessors. It introduces the concept of security dependence to mark speculative memory instructions which could leak information with potential security risks. More specifically, security-dependent instructions are detected and marked with suspect speculation flags in the Issue Queue. All the instructions can be speculatively issued for execution in accordance with the classic out-of-order pipeline. For those instructions with suspect speculation flags, they are considered as safe instructions if their speculative execution dose not refill new cache lines with unauthorized privilege data. Otherwise, they are considered as unsafe instructions and thus not allowed to execute speculatively. To pursue a balance of performance and security, we investigate two filtering mechanisms, Cache-hit-based Hazard Filter and Trusted Page Buffer-based Hazard Filter to filter out false security hazards. As for true security hazards, we have two approaches to prevent them from changing cache states. One is to block all unsafe access, the other is to fetch them from lower-level caches or memory to a speculative buffer temporarily, and refill them after confirming that they are on the correct execution path. Our design philosophy is to speculatively execute safe instructions to maintain the performance benefits of out-of-order execution while delaying the cache updates for speculative execution of unsafe instructions for security consideration. We evaluate Conditional Speculation in terms of performance, security, and area. The experimental results show that the hardware overhead is marginal and the performance overhead is minimal. Lutan Zhao, Peinan Li, Rui Hou 0001, Michael C. Huang 0001, Peng Liu 0005, Lixin Zhang 0002, Dan Meng 0002 |
IEEE Trans. Computers | 6 |
| 2020 | Enabling Rack-scale Confidential Computing using Heterogeneous Trusted Execution EnvironmentabstractWith its huge real-world demands, large-scale confidential computing still cannot be supported by today's Trusted Execution Environment (TEE), due to the lack of scalable and effective protection of high-throughput accelerators like GPUs, FPGAs, and TPUs etc. Although attempts have been made recently to extend the CPU-like enclave to GPUs, these solutions require change to the CPU or GPU chips, may introduce new security risks due to the side-channel leaks in CPU-GPU communication and are still under the resource constraint of today's CPU TEE.To address these problems, we present the first Heterogeneous TEE design that can truly support large-scale compute or data intensive (CDI) computing, without any chip-level change. Our approach, called HETEE, is a device for centralized management of all computing units (e.g., GPUs and other accelerators) of a server rack. It is uniquely designed to work with today's data centres and clouds, leveraging modern resource pooling technologies to dynamically compartmentalize computing tasks, and enforce strong isolation and reduce TCB through hardware support. More specifically, HETEE utilizes the PCIe ExpressFabric to allocate its accelerators to the server node on the same rack for a non-sensitive CDI task, and move them back into a secure enclave in response to the demand for confidential computing. Our design runs a thin TCB stack for security management on a security controller (SC), while leaving a large set of software (e.g., AI runtime, GPU driver, etc.) to the integrated microservers that operate enclaves. An enclaves is physically isolated from others through hardware and verified by the SC at its inception. Its microserver and computing units are restored to a secure state upon termination.We implemented HETEE on a real hardware system, and evaluated it with popular neural network inference and training tasks. Our evaluations show that HETEE can easily support the CDI tasks on the real-world scale and incurred a maximal throughput overhead of 2.17% for inference and 0.95% for training on ResNet152. Rui Hou 0001, XiaoFeng Wang 0001, Wenhao Wang 0001, Jiangfeng Cao, Boyan Zhao, Zhongpu Wang, Yuhui Zhang 0011, Jiameng Ying, Lixin Zhang 0002, Dan Meng 0002 |
SP | 10 |
| 2019 | Conditional Speculation: An Effective Approach to Safeguard Out-of-Order Execution Against Spectre AttacksabstractSpeculative execution side-channel vulnerabilities such as Spectre reveal that conventional architecture designs lack security consideration. This paper proposes a software transparent defense mechanism, named as Conditional Speculation, against Spectre vulnerabilities found on traditional out-of-order microprocessors. It introduces the concept of security dependence to mark speculative memory instructions which could leak information with potential security risk. More specifically, security-dependent instructions are detected and marked with suspect speculation flags in the Issue Queue. All the instructions can be speculatively issued for execution in accordance with the classic out-of-order pipeline. For those instructions with suspect speculation flags, they are considered as safe instructions if their speculative execution will not refill new cache lines with unauthorized privilege data. Otherwise, they are considered as unsafe instructions and thus not allowed to execute speculatively. To reduce the performance impact from not executing unsafe instructions speculatively, we investigate two filtering mechanisms, Cachehit based Hazard Filter and Trusted Page Buffer based Hazard Filter to filter out false security hazards. Our design philosophy is to speculatively execute safe instructions to maintain the performance benefits of out-of-order execution while blocking the speculative execution of unsafe instructions for security consideration. We evaluate Conditional Speculation in terms of performance, security and area. The experimental results show that the hardware overhead is marginal and the performance overhead is minimal. Peinan Li, Lutan Zhao, Rui Hou 0001, Lixin Zhang 0002, Dan Meng 0002 |
HPCA | 4 |
| 2019 | RAGuard: An Efficient and User-Transparent Hardware Mechanism against ROP AttacksabstractControl-flow integrity (CFI) is a general method for preventing code-reuse attacks, which utilize benign code sequences to achieve arbitrary code execution. CFI ensures that the execution of a program follows the edges of its predefined static Control-Flow Graph: any deviation that constitutes a CFI violation terminates the application. Despite decades of research effort, there are still several implementation challenges in efficiently protecting the control flow of function returns (Return-Oriented Programming attacks). The set of valid return addresses of frequently called functions can be large and thus an attacker could bend the backward-edge CFI by modifying an indirect branch target to another within the valid return set. This article proposes RAGuard, an efficient and user-transparent hardware-based approach to prevent Return-Oreiented Programming attacks. RAGuard binds a message authentication code (MAC) to each return address to protect its integrity. To guarantee the security of the MAC and reduce runtime overhead: RAGuard (1) computes the MAC by encrypting the signature of a return address with AES-128, (2) develops a key management module based on a Physical Unclonable Function (PUF) and a True Random Number Generator (TRNG), and (3) uses a dedicated register to reduce MACs’ load and store operations of leaf functions. We have evaluated our mechanism based on the open-source LEON3 processor and the results show that RAGuard incurs acceptable performance overhead and occupies reasonable area. Rui Hou 0001, Wei Song 0002, Sally A. McKee, Zhen Jia 0001, Chen Zheng 0001, Mingyu Chen 0001, Lixin Zhang 0002, Dan Meng 0002 |
ACM Trans. Archit. Code Optim. | 8 |
| 2019 | Understanding Processors Design Decisions for Data Analytics in Homogeneous Data CentersabstractOur global economy increasingly depends on our ability to gather, analyze, link, and compare very large data sets. Keeping up with such big data poses challenges in terms of both computational performance and energy efficiency, and motivates different approaches to explore data center systems and architectures. To better understand the processor design decisions in context of data analytics in data centers, we conduct comprehensive evaluations using representative data analaytics workloads on representative conventional multi-core and many-core processors. After a comprehensive analysis of performance, power, energy efficiency and performance-cost efficiency, we have the following observations: contrasted with the conventional wisdom that uses wimpy many-core processors to improve energy-efficiency, the brawny multi-core processors with SMT (simultaneous multithreading) and dynamic overclocking technologies outperform the counterparts in terms of not only execution time, but also energy-efficiency for most of data analytics workloads in our experiments. Zhen Jia 0001, Wanling Gao, Yingjie Shi, Sally A. McKee, Zhenyan Ji, Jianfeng Zhan, Lei Wang 0004, Lixin Zhang 0002 |
IEEE Trans. Big Data | 8 |
| 2018 | XOS: An Application-Defined Operating System for Datacenter ComputingabstractRapid growth of datacenter (DC) scale, urgency of cost control, increasing workload diversity, and huge software investment protection place unprecedented demands on the operating system (OS) efficiency, scalability, performance isolation, and backward-compatibility. The traditional OSes are not built to work with deep-hierarchy software stacks, large numbers of cores, tail latency guarantee, and increasingly rich variety of applications seen in modern DCs, and thus they struggle to meet the demands of such workloads. This paper presents XOS, an application-defined OS for modern DC servers. Our design moves resource management out of the OS kernel, supports customizable kernel subsystems in user space, and enables elastic partitioning of hardware resources. Specifically, XOS leverages modern hardware support for virtualization to move resource management functionality out of the conventional kernel and into user space, which lets applications achieve near bare-metal performance. We implement XOS on top of Linux to provide backward compatibility. XOS speeds up a set of DC workloads by up to 1.6× over our baseline Linux on a 24-core server, and outperforms the state-of-the-art Dune by up to 3.3× in terms of virtual memory management. In addition, XOS demonstrates good scalability and strong performance isolation. Chen Zheng 0001, Lei Wang 0004, Sally A. McKee, Lixin Zhang 0002, Hainan Ye, Jianfeng Zhan |
IEEE BigData | 4 |
| 2018 | CVR: efficient vectorization of SpMV on x86 processorsabstractSparse Matrix-vector Multiplication (SpMV) is an important computation kernel widely used in HPC and data centers. The irregularity of SpMV is a well-known challenge that limits SpMV’s parallelism with vectorization operations. Existing work achieves limited locality and vectorization efficiency with large preprocessing overheads. To address this issue, we present the Compressed Vectorization-oriented sparse Row (CVR), a novel SpMV representation targeting efficient vectorization. The CVR simultaneously processes multiple rows within the input matrix to increase cache efficiency and separates them into multiple SIMD lanes so as to take the advantage of vector processing units in modern processors. Our method is insensitive to the sparsity and irregularity of SpMV, and thus able to deal with various scale-free and HPC matrices. We implement and evaluate CVR on an Intel Knights Landing processor and compare it with five state-of-the-art approaches through using 58 scale-free and HPC sparse matrices. Experimental results show that CVR can achieve a speedup up to 1.70 × (1.33× on average) and a speedup up to 1.57× (1.10× on average) over the best existing approaches for scale-free and HPC sparse matrices, respectively. Moreover, CVR typically incurs the lowest preprocessing overhead compared with state-of-the-art approaches. Biwei Xie, Jianfeng Zhan, Xu Liu 0001, Wanling Gao, Zhen Jia 0001, Xiwen He, Lixin Zhang 0002 |
CGO | 7 |
| 2018 | Venice: An Effective Resource Sharing Architecture for Data Center ServersabstractConsolidated server racks are quickly becoming the standard infrastructure for engineering, business, medicine, and science. Such servers are still designed much in the way when they were organized as individual, distributed systems. Given that many fields rely on big-data analytics substantially, its cost-effectiveness and performance should be improved, which can be achieved by flexibly allowing resources to be shared across nodes. Here we describe Venice, a family of data-center server architectures that includes a strong communication substrate as a first-class resource. Venice supports a diverse set of resource-joining mechanisms that enables applications to leverage non-local resources efficiently. We have constructed a hardware prototype to better understand the implications of design decisions about system support for resource sharing. We use it to measure the performance of at-scale applications and to explore performance, power, and resource-sharing transparency tradeoffs (i.e., how many programming changes are needed). We analyze these tradeoffs for sharing memory, accelerators, and NICs. We find that reducing/hiding latency is particularly important, the chosen communication channels should match the sharing access patterns of the applications, and of which we can improve performance by exploiting inter-channel collaboration. Boyan Zhao, Rui Hou 0001, Jianbo Dong, Michael C. Huang 0001, Sally A. McKee, Qianlong Zhang, Yueji Liu, Lixin Zhang 0002, Dan Meng 0002 |
ACM Trans. Comput. Syst. | 9 |
| 2017 | Efficient Regional Congestion Awareness (ERCA) for Load Balance with Aggregated Congestion InformationabstractIn this paper, we propose Efficient Regional Congestion Awareness (ERCA), a novel adaptive routing technique which utilizes both local and non-local/aggregated congestion information to estimate network congestion. To reduce the interference caused by the noises of the congestion information, two optimizations are performed: first, to minimize the noises of the congestion information, ERCA only considers the congestion status of the links adjacent to the nodes defined from current to the corresponding boundary, second, to minimize the influence of the noises, ERCA exploits dynamic instead of static weights to evaluate port's congestion. Furthermore, ERCA uses one wire per dimension direction or quadrant to transmit congestion information between two adjacent nodes, so the wiring overhead of ERCA is minimal. Compared with the DBAR, ERCA can maximally improve the saturation throughput by 11.7% and averagely improve the saturation throughput by 6.02%. Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
PDP | 5 |
| 2017 | Understanding Big Data Analytics Workloads on Modern ProcessorsabstractBig data analytics workloads are very significant ones in modern data centers, and it is more and more important to characterize their representative workloads and understand their behaviors so as to improve the performance of data center computer systems. In this paper, we embark on a comprehensive study to understand the impacts and performance implications of the big data analytics workloads on the systems equipped with modern superscalar out-of-order processors. After investigating three most important application domains in Internet services in terms of page views and daily visitors, we choose 11 representative data analytics workloads and characterize their micro-architectural behaviors by using hardware performance counters. Our study reveals that the big data analytics workloads share many inherent characteristics, which place them in a different class from the traditional workloads and the scale-out services. To further understand the characteristics of big data analytics workloads, we perform correlation analysis to identify the most key factors that affect cycles per instruction (CPI). Also, we reveal that the increasing complexity of the big data software stacks will put higher pressures on the modern processor pipelines. Zhen Jia 0001, Jianfeng Zhan, Lei Wang 0004, Chunjie Luo, Wanling Gao, Rui Han 0001, Lixin Zhang 0002 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2016 | Auto-tuning Spark Big Data Workloads on POWER8: Prediction-Based Dynamic SMT ThreadingabstractMuch research work devotes to tuning big data analytics in modern data centers, since %the truth that even a small percentage of performance improvement immediately translates to huge cost savings because of the large scale. Simultaneous multithreading (SMT) receives great interest from data center communities, as it has the potential to boost performance of big data analytics by increasing the processor resources utilization. For example, the emerging processor architectures like POWER8 support up to 8-way multithreading. However, as different big data workloads have disparate architectural characteristics, how to identify the most efficient SMT configuration to achieve the best performance is challenging in terms of both complex application behaviors and processor architectures. In this paper, we specifically focus on auto-tuning SMT configuration for Spark-based big data workloads on POWE-R8. However, our methodology could be generalized and extended to other programming software stacks and other architectures. Zhen Jia 0001, Guancheng Chen, Jianfeng Zhan, Lixin Zhang 0002, Yonghua Lin, H. Peter Hofstee |
PACT | 5 |
| 2016 | sAXI: A High-Efficient Hardware Inter-Node Link in ARM Server for Remote Memory AccessabstractThe ever-growing need for fast big-data operations has made in-memory processing increasingly important in modern datacenters. To mitigate the capacity limitation of a single server node, techniques of inner-rack cross-node memory access have drawn attention recently. However, existing proposals exhibit inefficiency in remote memory access among server nodes due to inter-protocol conversions and non-transparent coarse-grained accesses. In this study, we propose the high-performance and efficient serialized AXI (sAXI) link and its associated cross-node memory access mechanism for emerging ARM-based servers. The key idea behind sAXI is directly extending the on-chip AMBA AXI-4.0 interconnection of the SoC in a local server node to the outside, and then bringing into remote server nodes via high-speed serial lanes. As a result, natively accessing remote memory in adjacent nodes in the same manner of local assets is supported by purely using existing software. Experimental results show that, using the sAXI data-path, performance of remote memory access in the user-level micro-benchmark is very promising (min. latency: 1.16µs, max. bandwidth: 1.52GB/s on our in-house FPGA prototype). In addition, through this efficient hardware inter-node link, performance of an in-memory key-value framework, Redis, can be improved up to 1.72x and large latency overhead of database query can be effectively hidden. Ke Zhang 0017, Yisong Chang, Lixin Zhang 0002, Mingyu Chen 0001, Zhiwei Xu 0002 |
CCGrid | 3 |
| 2016 | Intra-host Rate Control with Centralized ApproachabstractToday's datacenter is shared among various applications with different QoS requirements, which poses a great challenge to deliver low delay transport with high throughput. Most of works address this challenge by reducing the in-network delay, but assumes a negligible local delay. However, we show that this assumption does not hold for a multi-tenant datacenter that a physical machine is shared by multiple tenants with virtual machines running different applications. As measured, we found that VMs in a PM competing for bandwidth resources introduce delays as high as 13 ms, resulted from the packet queueing at QDisc layer of that PM, because current VMs' rate control still operates in a distributed manner without exploiting knowledge of the QoS requirements of applications running in VMs. This work addresses this problem by proposing a centralized rate adaptation (CERA) that operates in the host PM, dynamically schedules the flows from all VMs in a centralized manner. We implemented a CERA prototype and evaluated CERA through testbed experiments. Our results show that CERA reduces the local delay significantly thus reduces the average request latency of delay sensitive applications, e.g., memcached, by a factor of 6.3, without sacrificing the throughput performance of throughput intensive applications, e.g., iperf. Ke Liu 0004, Yifan Shen 0002, Jack Y. B. Lee, Mingyu Chen 0001, Lixin Zhang 0002 |
CLUSTER | 6 |
| 2016 | Venice: Exploring server architectures for effective resource sharingabstractConsolidated server racks are quickly becoming the backbone of IT infrastructure for science, engineering, and business, alike. These servers are still largely built and organized as when they were distributed, individual entities. Given that many fields increasingly rely on analytics of huge datasets, it makes sense to support flexible resource utilization across servers to improve cost-effectiveness and performance. We introduce Venice, a family of data-center server architectures that builds a strong communication substrate as a first-class resource for server chips. Venice provides a diverse set of resource-joining mechanisms that enables user programs to efficiently leverage non-local resources. To better understand the implications of design decisions about system support for resource sharing we have constructed a hardware prototype that allows us to more accurately measure end-to-end performance of at-scale applications and to explore tradeoffs among performance, power, and resource-sharing transparency. We present results from our initial studies analyzing these tradeoffs when sharing memory, accelerators, or NICs. We find that it is particularly important to reduce or hide latency, that data-sharing access patterns should match the features of the communication channels employed, and that inter-channel collaboration can be exploited for better performance. Jianbo Dong, Rui Hou 0001, Michael C. Huang 0001, Tao Jiang 0010, Boyan Zhao, Sally A. McKee, Xiaosong Cui, Lixin Zhang 0002 |
HPCA | 9 |
| 2016 | Extending On-chip Interconnects for rack-level remote resource accessabstractThe need to perform data analytics on exploding data volumes coupled with the rapidly changing workloads in cloud computing places great pressure on data-center servers. To improve hardware resource utilization across servers within a rack, we propose Direct Extension of On-chip Interconnects (DEOI), a high-performance and efficient architecture for remote resource access among server nodes. DEOI extends an SoC server node's on-chip interconnect to access resources in adjacent nodes with no protocol changes, allowing remote memory and network resources to be used as if they were local. Our results on a four-node FPGA prototype show that the latency of user-level, cross-node, random reads to DEOI-connected remote memory is as low as 1.16µs, which beats current commercial technologies. We exploit DEOI remote access to improve performance of the Redis in-memory key-value framework by 47%. When using DEOI to access remote network resources, we observe an 8.4% average performance degradation and only a 2.52µs ping-pong latency disparity compared to using local assets. These results suggest that DEOI can be a promising mechanism for increasing both performance and efficiency in next-generation data-center servers. Yisong Chang, Ke Zhang 0017, Sally A. McKee, Lixin Zhang 0002, Mingyu Chen 0001, Liqiang Ren, Zhiwei Xu 0002 |
ICCD | 4 |
| 2016 | Isolating bandwidth guarantees from work conservation in the cloudabstractTo predict lower bounds on the performance of applications, the cloud should provide guarantees on bandwidth that each virtual machine can obtain. By competing for spare network bandwidth, current solutions that provide bandwidth guarantees can achieve work conservation as well. However, they usually fail to provide accurate bandwidth guarantees, for the interference between traffic for achieving the two objectives respectively. In order to eliminate the interference, they reserve sufficient bandwidth headroom for every link, which cannot be allocated to tenants as guarantees, incurring a decrease in the total of guarantees that each link can offer and thus a decline in the revenue of the cloud provider. To address the problem, this paper proposes a new mechanism, namely DFlow, which achieves bandwidth guarantees and work conservation simultaneously. Specifically, DFlow isolates its solutions for achieving bandwidth guarantees and work conservation from each other by splitting every flow into two subflows with distinct priorities. They are then used to achieve the two objectives respectively. Our evaluations show that DFlow can provide accurate bandwidth guarantees without reserving any bandwidth headroom while achieving work conservation to effectively utilize spare network bandwidth. Ke Liu 0004, Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
ISCC | 5 |
| 2016 | Adaptive rate control over mobile data networks with heuristic rate compensationsabstractMobile data networks exhibit highly variable data rates and stochastic non-congestion-related packet loss. These challenges result in key performance bottlenecks in current Transmission Control Protocol (TCP) implementations: bandwidth inefficiency and large end-to-end delay. This work addresses these challenges by first developing a Sliding Interval based Rate Adaptation (SIRA) that tracks bandwidths with a fixed time interval and applies them to its transmission rate periodically. Extensive experiments confirmed that SIRA achieves 96.3% bandwidth utilization and reduces the average queueing delay by a factor of 1.37, compared to TCP CUBIC, the preferred variant for Internet servers. However, the resultant end-to-end delay is still much larger for interactive applications, thus we complement SIRA with two heuristic rate compensation algorithms (SIRA-H) given that the bandwidth does not vary significantly in long time scales. Specifically, SIRA-H first reduces the transmission rate of SIRA if the estimated RTT is above a prefigured threshold. Meanwhile, it computes the amount of unsent data that would be transmitted if SIRA were used, and compensates the rate reduction with those unsent data as if their ACKs were received, when the queue is detected to be empty. We evaluated SIRA-H through a combination of trace-driven emulations and real-world experiments, and showed that it reduces the 95thpercentile queueing delay by a factor of over 3.9, while maintains a similar throughput compared to the original SIRA. In comparison to state of the art protocols such as Sprout and Verus, SIRA-H also reduces the 95thpercentile queueing delay by a factor of over 0.8. Ke Liu 0004, Jack Y. B. Lee, Mingyu Chen 0001, Lixin Zhang 0002 |
IWQoS | 5 |
| 2015 | Supporting Differentiated Services in Computers via Programmable Architecture for Resourcing-on-Demand (PARD)abstractThis paper presents PARD, a programmable architecture for resourcing-on-demand that provides a new programming interface to convey an application's high-level information like quality-of-service requirements to the hardware. PARD enables new functionalities like fully hardware-supported virtualization and differentiated services in computers. PARD is inspired by the observation that a computer is inherently a network in which hardware components communicate via packets (e.g., over the NoC or PCIe). We apply principles of software-defined networking to this intra-computer network and address three major challenges. First, to deal with the semantic gap between high-level applications and underlying hardware packets, PARD attaches a high-level semantic tag (e.g., a virtual machine or thread ID) to each memory-access, I/O, or interrupt packet. Second, to make hardware components more manageable, PARD implements programmable control planes that can be integrated into various shared resources (e.g., cache, DRAM, and I/O devices) and can differentially process packets according to tag-based rules. Third, to facilitate programming, PARD abstracts all control planes as a device file tree to provide a uniform programming interface via which users create and apply tag-based rules. Jiuyue Ma, Xiufeng Sui, Ninghui Sun, Tianni Xu, Zhicheng Yao, Lixin Zhang 0002, Yungang Bao |
ASPLOS | 11 |
| 2015 | Adapting Memory Hierarchies for Emerging Datacenter Interconnects
Tao Jiang 0010, Rui Hou 0001, Jianbo Dong, Lin Chai, Sally A. McKee, Lixin Zhang 0002, Ninghui Sun |
J. Comput. Sci. Technol. | 7 |
| 2014 | Dandelion: A locally-high-performance and globally-high-scalability hierarchical data center networkabstractThe increasing customer demand is driving modern data centers to embrace the freely-expandable network architecture. Unfortunately, state-of-the-art freely-expandable networks suffer from either the large granularity of expansion or the prohibitive implementation cost. Furthermore, a recent research showed that data center traffic tends to be highly clustered. Based on above observations, this paper proposes a freely-expandable network architecture, namely the dandelion. Dandelion is a two-level hierarchical network, where the first level aims at “high performance” and the second level aims at “high scalability”. The resulting network has two distinct advantages. First, it could arbitrarily expand with a reasonable granularity. Second, the router architecture is efficient as well as highly scalable since 1) the routing table is significantly compressed and 2) a fixed number of virtual channels per physical channel are required regardless of the network size. Finally, the traffic characteristics of four typical cloud applications are analyzed, and the generated traffic patterns are used to evaluate the proposed network architecture. Simulation results prove that the dandelion is a promising network architecture for future data centers. Binzhang Fu, Wentao Bao, Guolong Jiang, Mingyu Chen 0001, Lixin Zhang 0002, Yidong Tao, Junfeng Zhao 0003 |
ICCCN | 6 |
| 2014 | DWC: dynamic write consolidation for phase change memory systemsabstractPhase change memory (PCM) is promising to become an alternative main memory thanks to its better scalability and lower leakage than DRAM. However, the long write latency of PCM puts it at a severe disadvantage against DRAM. In this paper, we propose a Dynamic Write Consolidation (DWC) scheme to improve PCM memory system performance while reducing energy consumption. This paper is motivated by the observation that a large fraction of a cache line being written back to memory is not actually modified. DWC exploits the unnecessary burst writes of unmodified data to consolidate multiple writes targeting the same row into one write. By doing so, DWC enables multiple writes to be send within one. DWC incurs low implementation overhead and shows significant efficiency. The evaluation results show that DWC achieves up to 35.7% performance improvement, and 17.9% on average. The effective write latency are reduced by up to 27.7%, and 16.0% on average. Moreover, DWC reduces the energy consumption by up to 35.3%, and 13.9% on average. Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001, Lixin Zhang 0002, Ninghui Sun |
ICS | 5 |
| 2014 | Pipelined Compaction for the LSM-TreeabstractWrite-optimized data structures like Log-Structured Merge-tree (LSM-tree) and its variants are widely used in key-value storage systems like Big Table and Cassandra. Due to deferral and batching, the LSM-tree based storage systems need background compactions to merge key-value entries and keep them sorted for future queries and scans. Background compactions play a key role on the performance of the LSM-tree based storage systems. Existing studies about the background compaction focus on decreasing the compaction frequency, reducing I/Os or confining compactions on hot data key-ranges. They do not pay much attention to the computation time in background compactions. However, the computation time is no longer negligible, and even the computation takes more than 60% of the total compaction time in storage systems using flash based SSDs. Therefore, an alternative method to speedup the compaction is to make good use of the parallelism of underlying hardware including CPUs and I/O devices. In this paper, we analyze the compaction procedure, recognize the performance bottleneck, and propose the Pipelined Compaction Procedure (PCP) to better utilize the parallelism of CPUs and I/O devices. Theoretical analysis proves that PCP can improve the compaction bandwidth. Furthermore, we implement PCP in real system and conduct extensive experiments. The experimental results show that the pipelined compaction procedure can increase the compaction bandwidth and storage system throughput by 77% and 62% respectively. Zigang Zhang, Yinliang Yue, Bingsheng He, Jin Xiong, Mingyu Chen 0001, Lixin Zhang 0002, Ninghui Sun |
IPDPS | 6 |
| 2014 | Intelligent frame refresh for energy-aware display subsystems in mobile devicesabstractFrame refreshes, that are used to retain frame images from frame buffers for display subsystems in mobile devices, waste energy and memory bandwidth. In this paper, we propose an intelligent frame refresh mechanism to reduce redundant frame refreshes and useless data accesses to frame buffers, which bridges the semantic gap between frame buffers and frame refreshes, and exploits the knowledge of frame buffers to guide frame refreshes. Based on this mechanism, we introduce two detailed schemes to optimize refreshes by utilizing different information. The flipping-aware frame refresh scheme uses the frame buffer switching operations to detect frame image updates and triggers useful refreshes. The row-level frame refresh scheme supports to refresh only modified rows instead of the whole frame, under the guidance of pixel status information of frame buffers. Our evaluation results show that our proposed mechanism can reduce memory requests by nearly 50% and memory power consumption up to 30%, compared to conventional fixed frame refresh mechanism. Yongbing Huang, Mingyu Chen 0001, Lixin Zhang 0002, Shihai Xiao, Junfeng Zhao 0003, Zhulin Wei |
ISLPED | 3 |
| 2014 | Moby: A mobile benchmark suite for architectural simulatorsabstractMobile devices such as smartphones and tablets have become the primary consumer computing devices, and their rate of adoption continues to grow. The applications that run on these mobile platforms vary in how they use hardware resources, and their diversity is increasing. Performance and power limitations also vary widely across mobile platforms. Thus there is a growing need for tools to help computer architects design systems to meet the needs of mobile workloads. Full-system simulators are invaluable tools for designing new architectures, but we still need appropriate benchmark suites that capture the behaviors of emerging mobile applications. Current benchmark suites cover only a small range of mobile applications, and many cannot run directly in simulators due to their user interaction requirements. In this paper, we introduce and characterize Moby, a benchmark suite designed to make it easier to use full-system architectural simulators to evaluate microarchitectures for mobile processors. Moby contains popular Android applications, including a web browser, a social networking application, an email client, a music player, a video player, a document processing application, and a map program. To facilitate microarchitectural exploration, we port the Moby benchmark suite to the popular gem5 simulator. We characterize the architecture-independent features of Moby applications on the simulator and analyze the architecture-dependent features on a current-generation mobile platform. Our results show that mobile applications exhibit complex instruction execution behaviors and poor code locality, but current mobile platforms especially instruction-related components cannot meet their requirements. Yongbing Huang, Zhongbin Zha, Mingyu Chen 0001, Lixin Zhang 0002 |
ISPASS | 4 |
| 2014 | A High-Performance and Cost-Efficient Interconnection Network for High-Density Servers
Wentao Bao, Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
J. Comput. Sci. Technol. | 4 |
| 2013 | The ARMv8 simulatorabstractIn this work, we implement an ARMv8 function and performance simulator based on gem5 infrastructure, which is the first open source ARMv8 simulator. All the ARMv8 A64 instructions other than SIMD are implemented using gem5 ISA description language. The ARMv8 simulator supports multiple CPU models, multiple memory systems, and McPAT power model. Tao Jiang 0010, Rui Hou 0001, Yi Zhang 0037, Qianlong Zhang, Lin Chai, Jing Han 0011, Wuxiong Zhang, Lixin Zhang 0002 |
ICS | 10 |
| 2013 | Understanding the implications of virtual machine management on processor microarchitecture designabstractCloud computing has demonstrated tremendous capability in a wide spectrum of online services. Virtualization provides an efficient solution to the utilization of modern multicore processor systems while affording significant flexibility. The growing popularity of virtualized datacenters motivates deeper understanding of the interactions between virtual machine management and the micro-architecture behaviors of the privileged domain. We argue that these behaviors must be factored into the design of processor microarchitecture in virtualized datacenters. In this work, we use performance counters on modern servers to study the micro-architectural execution characteristics of the privileged domain while performing various VM management operations. Our study shows that today's state-of-the-art processor still has room for further optimizations when executing virtualized cloud workloads, particularly in the organization of last level caches and on-chip cache coherence protocol. Specifically, our analysis shows that: shared caches could be partitioned to eliminate interference between the privileged domain and guest domains; the cache coherence protocol could support a high degree of data sharing of the privileged domain; and cache capacity or CPU utilization occupied by the privileged domain could be effectively managed when performing management workflows to achieve high system throughput. Xiufeng Sui, Tao Li 0006, Lixin Zhang 0002 |
ISPASS | 4 |
| 2012 | CloudRank-D: benchmarking and ranking cloud computing systems for data processing applications
Chunjie Luo, Jianfeng Zhan, Zhen Jia 0001, Lei Wang 0004, Lixin Zhang 0002, Cheng-Zhong Xu 0001, Ninghui Sun |
Frontiers Comput. Sci. | 6 |
| 2012 | Active memory controller
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Ali Ibrahim, Michael A. Parker, Xiaowei Jiang |
J. Supercomput. | 2 |
| 2011 | Efficient data streaming with on-chip accelerators: Opportunities and challengesabstractThe transistor density of microprocessors continues to increase as technology scales. Microprocessors designers have taken advantage of the increased transistors by integrating a significant number of cores onto a single die. However, a large number of cores are met with diminishing returns due to software and hardware scalability issues and hence designers have started integrating on-chip special-purpose logic units (i.e., accelerators) that were previously available as PCI-attached units. It is anticipated that more accelerators will be integrated on-chip due to the increasing abundance of transistors and the fact that not all logic can be powered at all times due to power budget limits. Thus, on-chip accelerator architectures deserve more attention from the research community. There is a wide spectrum of research opportunities for design and optimization of accelerators. This paper attempts to bring out some insights by studying the data access streams of on-chip accelerators that hopefully foster some future research in this area. Specifically, this paper uses a few simple case studies to show some of the common characteristics of the data streams introduced by on-chip accelerators, discusses challenges and opportunities in exploiting these characteristics to optimize the power and performance of accelerators, and then analyzes the effectiveness of some simple optimizing extensions proposed. Rui Hou 0001, Lixin Zhang 0002, Michael C. Huang 0001, Kun Wang 0005, Hubertus Franke, Yi Ge, Xiaotao Chang |
HPCA | 2 |
| 2011 | Power shifting in Thrifty Interconnection NetworkabstractThis paper presents two complementary techniques to manage the power consumption of large-scale systems with a packet-switched interconnection network. First, we propose Thrifty Interconnection Network (TIN), where the network links are activated and de-activated dynamically with little or no overhead by using inherent system events to timely trigger link activation or de-activation. Second, we propose Network Power Shifting (NPS) that dynamically shifts the power budget between the compute nodes and their corresponding network components. TIN activates and trains the links in the interconnection network, just-in-time before the network communication is about to happen, and thriftily puts them into a low-power mode when communication is finished, hence reducing unnecessary network power consumption. Furthermore, the compute nodes can absorb the extra power budget shifted from its attached network components and increase their processor frequency for higher performance with NPS. Our simulation results on a set of real-world workload traces show that TIN can achieve on average 60% network power reduction, with the support of only one low-power mode. When NPS is enabled, the two together can achieve 12% application performance improvement and 13% overall system energy reduction. Further performance improvement is possible if the compute nodes can speed up more and fully utilize the extra power budget reinvested from the thrifty network with more aggressive cooling support. Jian Li 0059, Wei Huang 0004, Charles Lefurgy, Lixin Zhang 0002, Wolfgang E. Denzel, Richard R. Treumann, Kun Wang 0005 |
HPCA | 4 |
| 2010 | Enigma: architectural and operating system support for reducing the impact of address translationabstractMost modern microprocessors provide hardware support for rapidly translating a program logical address to a system physical address (PA). Translation typically sits on the critical path of every memory access, since an access cannot usually be performed until after it has been translated. Enigma is a novel approach to address translation that defers the bulk of the work associated with address translation until data must be retrieved from physical memory. Enigma replaces the address translation unit that exists in each conventional core with a simpler unit to translate from the logical address space to a new intermediate address (IA) space. Intermediate addresses are unique across the entire system except where sharing is required or desired, and their use sidesteps the "synonym" problem present in logically tagged caches. All cache addressing, as well as I/O and coherence traffic, is carried out using IA. Enigma translates an IA to a PA only when no cache in the entire CMP can satisfy the request and memory or I/O must be accessed. A central translation unit attached to the system bus performs translations on IA that must be resolved to a PA. Deferring the bulk of address translation work and removing it from each individual processor core in this manner affords many benefits. Lixin Zhang 0002, William Evan Speight, Ramakrishnan Rajamony, Jiang Lin |
ICS | 1 |
| 2010 | Design exploration of hybrid caches with disparate memory technologiesabstractTraditional multilevel SRAM-based cache hierarchies, especially in the context of chip multiprocessors (CMPs), present many challenges in area requirements, core--to--cache balance, power consumption, and design complexity. New advancements in technology enable caches to be built from other technologies, such as Embedded DRAM (EDRAM), Magnetic RAM (MRAM), and Phase-change RAM (PRAM), in both 2D chips or 3D stacked chips. Caches fabricated in these technologies offer dramatically different power-performance characteristics when compared with SRAM-based caches, particularly in the areas of access latency, cell density, and overall power consumption. In this article, we propose to take advantage of the best characteristics that each technology has to offer through the use of Hybrid Cache Architecture (HCA) designs. We discuss and evaluate two types of hybrid cache architectures: intercache-Level HCA (LHCA), in which the levels in a cache hierarchy can be made of disparate memory technologies; and intracache-level or cache-Region-based HCA (RHCA), where a single level of cache can be partitioned into multiple regions, each of a different memory technology. We have studied a number of different HCA architectures and explored the potential of hardware support for intracache data movement and power consumption management within HCA caches. Utilizing a full-system simulator that has been validated against real hardware, we demonstrate that an LHCA design can provide a geometric mean 6% IPC improvement over a baseline 3-level SRAM cache design under the same area constraint across a collection of 30 workloads. A more aggressive RHCA-based design provides 10% IPC improvement over the baseline. A 2-layer 3D cache stack (3DHCA) of high density memory technology within the same chip footprint gives 16% IPC improvement over the baseline. We also achieve up to a 72% reduction in power consumption over a baseline SRAM-only design. Energy-delay and thermal evaluation for 3DHCA are also presented. In addition to the fast-slow region based RHCA, we further evaluate read-write region based RHCA designs. Xiaoxia Wu, Jian Li 0059, Lixin Zhang 0002, William Evan Speight, Ramakrishnan Rajamony, Yuan Xie 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2009 | Power and performance of read-write aware Hybrid Caches with non-volatile memoriesabstractCaches made of non-volatile memory technologies, such as magnetic RAM (MRAM) and phase-change RAM (PRAM), offer dramatically different power-performance characteristics when compared with SRAM-based caches, particularly in the areas of static/dynamic power consumption, read and write access latency and cell density. In this paper, we propose to take advantage of the best characteristics that each technology has to offer through the use of read-write aware hybrid cache architecture (RWHCA) designs, where a single level of cache can be partitioned into read and write regions, each of a different memory technology with disparate read and write characteristics. We explore the potential of hardware support for intra-cache data movement within RWHCA caches. Utilizing a full-system simulator that has been validated against real hardware, we demonstrate that a RWHCA design with a conservative setup can provide a geometric mean 55% power reduction and yet 5% IPC improvement over a baseline SRAM cache design across a collection of 30 workloads. Furthermore, a 2-layer 3D cache stack (3DRWHCA) of high density memory technology with the same chip footprint still gives 10% power reduction and boost performance by 16% IPC improvement over the baseline. Xiaoxia Wu, Jian Li 0059, Lixin Zhang 0002, William Evan Speight, Yuan Xie 0001 |
DATE | 3 |
| 2009 | Lightweight predication support for out of order processorsabstractThe benefits of Out of Order (OOO) processing are well known, as is the effectiveness of predicated execution for unpredictable control flow. However, as previous research has demonstrated, these techniques are at odds with one another. One common approach to reconciling their differences is to simplify the form of predication supported by the architecture. For instance, the only form of predication supported by modern OOO processors is a simple conditional move. We argue that it is the simplicity of conditional move that has allowed its widespread adoption, but we also show that this simplicity compromises its effectiveness as a compilation target. In this paper, we introduce a generalized form of hammock predication - called predicated mutually exclusive groups - that requires few modifications to an existing processor pipeline, yet presents the compiler with abundant predication opportunities. In comparison to non-predicated code running on an aggressively clocked baseline system, our technique achieves an 8% speedup averaged across three important benchmark suites. Mark Stephenson, Lixin Zhang 0002, Ram Rangan |
HPCA | 2 |
| 2009 | Thrifty interconnection network for HPC systemsabstractWe propose Thrifty Interconnection Network (TIN), where the network links are activated and de-activated dynamically to save power with little or no overhead by using inherent system events to overlap the link activation or de-activation time. Our simulation results on a set of real world HPC workload traces show on average 35% network power reduction. Jian Li 0059, Lixin Zhang 0002, Charles Lefurgy, Richard R. Treumann, Wolfgang E. Denzel |
ICS | 2 |
| 2009 | Hybrid cache architecture with disparate memory technologiesabstractCaching techniques have been an efficient mechanism for mitigating the effects of the processor-memory speed gap. Traditional multi-level SRAM-based cache hierarchies, especially in the context of chip multiprocessors (CMPs), present many challenges in area requirements, core-to-cache balance, power consumption, and design complexity. New advancements in technology enable caches to be built from other technologies, such as Embedded DRAM (EDRAM), Magnetic RAM (MRAM), and Phase-change RAM (PRAM), in both 2D chips or 3D stacked chips. Caches fabricated in these technologies offer dramatically different power and performance characteristics when compared with SRAM-based caches, particularly in the areas of access latency, cell density, and overall power consumption. In this paper, we propose to take advantage of the best characteristics that each technology offers, through the use of Hybrid Cache Architecture (HCA) designs. We discuss and evaluate two types of hybrid cache architectures: inter cache Level HCA (LHCA), in which the levels in a cache hierarchy can be made of disparate memory technologies; and intra cache level or cache Region based HCA (RHCA), where a single level of cache can be partitioned into multiple regions, each of a different memory technology. We have studied a number of different HCA architectures and explored the potential of hardware support for intra-cache data movement and power consumption management within HCA caches. Utilizing a full-system simulator that has been validated against real hardware, we demonstrate that an LHCA design can provide a geometric mean 7% IPC improvement over a baseline 3-level SRAM cache design under the same area constraint across a collection of 25 workloads. A more aggressive RHCA-based design provides 12% IPC improvement over the baseline. Finally, a 2-layer 3D cache stack (3DHCA) of high density memory technology within the same chip footprint gives 18% IPC improvement over the baseline. Furthermore, up to 70% reduction in power consumption over a baseline SRAM-only design is achieved. Xiaoxia Wu, Jian Li 0059, Lixin Zhang 0002, William Evan Speight, Ramakrishnan Rajamony, Yuan Xie 0001 |
ISCA | 3 |
| 2007 | Active memory operationsabstractThe performance of modern microprocessors is increasingly limited by their inability to hide main memory latency. The problem is worse in large-scale shared memory systems, where remote memory latencies are hundreds, and soon thousands, of processor cycles. To mitigate this problem, we propose the use of Active Memory Operations (AMOs), in which select operations can be sent to and executed on the home memory controller of data. AMOs can eliminate significant number of coherence messages, minimize intranode and internode memory traffic, and create opportunities for parallelism. Our implementation of AMOs is cache-coherent and requires no changes to the processor core or DRAM chips. Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Ali Ibrahim, Michael A. Parker |
ICS | 2 |
| 2007 | A NUCA Substrate for Flexible CMP Cache SharingabstractWe propose an organization for the on-chip memory system of a chip multiprocessor in which 16 processors share a 16-Mbyte pool of 64 level-2 (L2) cache banks. The L2 cache is organized as a nonuniform cache architecture (NUCA) array with a switched network embedded in it for high performance. We show that this organization can support a spectrum of degrees of sharing: unshared, in which each processor owns a private portion of the cache, thus reducing hit latency, and completely shared, in which every processor shares the entire cache, thus minimizing misses, and every point in between. We measure the optimal degree of sharing for different cache bank mapping policies and also evaluate a per-application cache partitioning strategy. We conclude that a static NUCA organization with sharing degrees of 2 or 4 works best across a suite of commercial and scientific parallel workloads. We demonstrate that migratory dynamic NUCA approaches improve performance significantly for a subset of the workloads at the cost of increased complexity, especially as per-application cache partitioning strategies are applied. We also evaluate the energy efficiency of each design point in terms of network traffic, bank accesses, and external memory accesses. Jaehyuk Huh 0001, Changkyu Kim, Hazim Shafi, Lixin Zhang 0002, Doug Burger, Stephen W. Keckler |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2006 | Efficient address remapping in distributed shared-memory systemsabstractAs processor performance continues to improve at a rate much higher than DRAM and network performance, we are approaching a time when large-scale distributed shared memory systems will have remote memory latencies measured in tens of thousands of processor cycles. The Impulse memory system architecture adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to control how data is accessed and cached and thereby improve cache and bus utilization and reduce the number of memory accesses required. Previous Impulse work focuses on uniprocessor systems and relies on software to flush processor caches when necessary to ensure data coherence. In this paper, we investigate an extension of Impulse to multiprocessor systems that extends the coherence protocol to maintain data coherence without requiring software-directed cache flushing. Specifically, the multiprocessor Impulse controller can gather/scatter data across the network while its coherence protocol guarantees that each gather request gets coherent data and each scatter request updates every coherent replica in the system. Our simulation results demonstrate that the proposed system can significantly outperform conventional systems, achieving an average speedup of 9X on four memory-bound benchmarks on a 32-processor system. Lixin Zhang 0002, Michael A. Parker, John B. Carter |
ACM Trans. Archit. Code Optim. | 1 |
| 2005 | A NUCA substrate for flexible CMP cache sharingabstractWe propose an organization for the on-chip memory system of a chip multiprocessor, in which 16 processors share a 16MB pool of 256 L2 cache banks. The L2 cache is organized as a non-uniform cache architecture (NUCA) array with a switched network embedded in it for high performance. We show that this organization can support the spectrum of degrees of sharing: unshared, in which each processor has a private portion of the cache, thus reducing hit latency, completely shared, in which every processor shares the entire cache, thus minimizing misses, and every point in between. We find the optimal degree of sharing for a number of cache bank mapping policies, and also evaluate a per-application cache partitioning strategy. We conclude that a static NUCA organization with sharing degrees of two or four work best across a suite of commercial and scientific parallel workloads. We also demonstrate that migratory, dynamic NUCA approaches improve performance significantly for a subset of the workloads at the cost of increased power consumption and complexity, especially as per-application cache partitioning strategies are applied. Jaehyuk Huh 0001, Changkyu Kim, Hazim Shafi, Lixin Zhang 0002, Doug Burger, Stephen W. Keckler |
ICS | 4 |
| 2005 | Adaptive Mechanisms and Policies for Managing Cache Hierarchies in Chip MultiprocessorsabstractWith the ability to place large numbers of transistors on a single silicon chip, manufacturers have begun developing chip multiprocessors (CMPs) containing multiple processor cores, varying amounts of level 1 and level 2 caching, and on-chip directory structures for level 3 caches and memory. The level 3 cache may be used as a victim cache for both modified and clean lines evicted from on-chip level 2 caches. Efficient area and performance management of this cache hierarchy is paramount given the projected increase in access latency to off-chip memory. This paper proposes simple architectural extensions and adaptive policies for managing the L2 and L3 cache hierarchy in a CMP system. In particular, we evaluate two mechanisms that improve cache effectiveness. First, we propose the use of a small history table to provide hints to the L2 caches as to which lines are resident in the L3 cache. We employ this table to eliminate some unnecessary clean write backs to the L3 cache, reducing pressure on the L3 cache and utilization of the on-chip bus. Second, we exam-ine the performance benefits of allowing write backs from L2 caches to be placed in neighboring, on-chip L2 caches rather than forcing them to be absorbed by the L3 cache. This not only reduces the capacity pressure on the L3 cache but also makes subsequent accesses faster since L2-W-L2 cache transfers have typically lower latencies than accesses to a large L3 cache array. We evaluate the performance improvement of these two designs, and their combined effect, on four commercial workloads and observe a reduction in the overall execution time of up to 13%. William Evan Speight, Hazim Shafi, Lixin Zhang 0002, Ramakrishnan Rajamony |
ISCA | 3 |
| 2005 | Fast synchronization on shared-memory multiprocessors: An architectural approach
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Liqun Cheng, Michael A. Parker |
J. Parallel Distributed Comput. | 2 |
| 2004 | Highly Efficient Synchronization Based on Active Memory OperationsabstractSummary form only given. Synchronization is a crucial operation in many parallel applications. As network latency approaches thousands of processor cycles for large scale multiprocessors, conventional synchronization techniques are failing to keep up with the increasing demand for scalable and efficient synchronization operations. We present a mechanism that allows atomic synchronization operations to be executed on the home memory controller of the synchronization variable. By performing atomic operations near where the data resides, our proposed mechanism can significantly reduce the number of network messages required by synchronization operations. Our proposed design also enhances performance by using fine-grained updates to selectively "push " the results of offloaded synchronization operations back to processors when they complete (e.g., when a barrier count reaches the desired value). We use the proposed mechanism to optimize two of the most widely used synchronization operations, barriers and spin locks. Our simulation results show that the proposed mechanism outperforms conventional implementations based on load-linked/store-conditional, processor-centric atomic instructions, conventional memory-side atomic instructions, or active messages. It speeds up conventional barriers by up to 2.1 (4 processors) to 61.9 (256 processors) and spin locks by a factor of up to 2.0 (4 processors) to 10.4 (256 processors). Lixin Zhang 0002, Zhen Fang 0002, John B. Carter |
IPDPS | 1 |
| 2001 | Reevaluating Online Superpage Promotion with Hardware SupportabstractTypical translation lookaside buffers (TLBs) can map a far smaller region of memory than application footprints demand, and the cost of handling TLB misses therefore limits the performance of an increasing number of applications. This bottleneck can be mitigated by the use of superpages, multiple adjacent virtual memory pages that can be mapped with a single TLB entry that extend TLB reach without significantly increasing size or cost. We analyze hardware/software tradeoff for dynamically creating superpages. This study extends previous work by using execution-driven simulation to compare creating superpages via copying with remapping pages within the memory controller and by examining how the tradeoffs change when moving front a single-issue to a superscalar processor model. We find that remapping-based promotion outperforms copying-based promotion, often significantly. Copying-based promotion is slightly more effective on superscalar processors than on single-issue processors, and the relative performance of remapping-based promotion on the two platform is application-dependent. Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Wilson C. Hsieh, Sally A. McKee |
HPCA | 2 |
| 2001 | The Impulse Memory ControllerabstractImpulse is a memory system architecture that adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to remap their data structures in memory. As a result, they can control how their data is accessed and cached, which can improve cache and bus utilization. The Impulse design does not require any modification to processor, cache, or bus designs since all the functionality resides at the memory controller. As a result, Impulse can be adopted in conventional systems without major system changes. We describe the design of the Impulse architecture and how an Impulse memory system can be used in a variety of ways to improve the performance of memory-bound applications. Impulse can be used to dynamically create superpages cheaply, to dynamically recolor physical pages, to perform strided fetches, and to perform gathers and scatters through indirection vectors. Our performance results demonstrate the effectiveness of these optimizations in a variety of scenarios. Using Impulse can speed up a range of applications from 20 percent to over a factor of 5. Alternatively, Impulse can be used by the OS for dynamic superpage creation; the best policy for creating superpages using Impulse outperforms previously known superpage creation policies. Lixin Zhang 0002, Zhen Fang 0002, Michael A. Parker, Binu K. Mathew, Lambert Schaelicke, John B. Carter, Wilson C. Hsieh, Sally A. McKee |
IEEE Trans. Computers | 1 |
| 2000 | Online superpage promotion revisited (poster)abstractNo abstract available. Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Wilson C. Hsieh |
SIGMETRICS | 2 |
| 1999 | Impulse: Building a Smarter Memory ControllerabstractImpulse is a new memory system architecture that adds two important features to a traditional memory controller. First, Impulse supports application-specific optimizations through configurable physical address remapping. By remapping physical addresses, applications control how their data is accessed and cached, improving their cache and bus utilization. Second, Impulse supports prefetching at the memory controller, which can hide much of the latency of DRAM accesses. In this paper we describe the design of the Impulse architecture, and show how an Impulse memory system can be used to improve the performance of memory-bound programs. For the NAS conjugate gradient benchmark, Impulse improves performance by 67%. Because it requires no modification to processor, cache, or bus designs, Impulse can be adopted in conventional systems. In addition to scientific applications, we expect that Impulse will benefit regularly strided memory-bound applications of commercial importance, such as database and multimedia programs. John B. Carter, Wilson C. Hsieh, Leigh Stoller, Mark R. Swanson, Lixin Zhang 0002, Erik Brunvand, Al Davis, Chen-Chi Kuo, Ravindra Kuramkote, Michael A. Parker, Lambert Schaelicke, Terry Tateyama |
HPCA | 5 |