EDBT 2026 Demo / reviewers in the wild / expert
Liqun Cheng
dblp:52/2470
· DBLP profile ↗
15ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
12 papers |
Cloud and datacenter computing · 46% Hardware accelerators and domain-specific architectures · 22% Performance modeling and evaluation · 11% | |
| Artificial intelligence
4 papers |
Efficient and distributed learning · 83% Transfer learning and domain adaptation · 13% Deep learning architectures and training · 4% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
1.8 | 3 | 2023 | TripLe: Revisiting Pretrained Model Reuse and Progressive Learning for Efficient Vision Transformer Scaling and Searching · ICCV 2023 Hyperscale Hardware Optimized Neural Architecture Search · ASPLOS (3) 2023 Searching for Fast Model Families on Datacenter Accelerators · CVPR 2021 |
Cloud and datacenter computing › resource management
datacenter resource management |
1.7 | 4 | 2025 | Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage · NSDI 2025 Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019 Improving Resource Efficiency at Scale with Heracles · ACM Trans. Comput. Syst. 2016 |
Machine learning › Efficient and distributed learning › large-scale learning
model scaling |
1.2 | 2 | 2023 | TripLe: Revisiting Pretrained Model Reuse and Progressive Learning for Efficient Vision Transformer Scaling and Searching · ICCV 2023 Searching for Fast Model Families on Datacenter Accelerators · CVPR 2021 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.7 | 2 | 2020 | Autonomous Warehouse-Scale Computers · DAC 2020 Improving Resource Efficiency at Scale with Heracles · ACM Trans. Comput. Syst. 2016 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
one-shot neural architecture search |
0.7 | 1 | 2023 | Hyperscale Hardware Optimized Neural Architecture Search · ASPLOS (3) 2023 |
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pretrained model reuse |
0.7 | 1 | 2023 | TripLe: Revisiting Pretrained Model Reuse and Progressive Learning for Efficient Vision Transformer Scaling and Searching · ICCV 2023 |
Hardware accelerators and domain-specific architectures › neural architecture search
hardware-aware neural architecture search |
0.7 | 1 | 2023 | Hyperscale Hardware Optimized Neural Architecture Search · ASPLOS (3) 2023 |
Hardware accelerators and domain-specific architectures
neural architecture search |
0.7 | 1 | 2023 | Hyperscale Hardware Optimized Neural Architecture Search · ASPLOS (3) 2023 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
hardware-aware neural architecture search |
0.5 | 1 | 2021 | Searching for Fast Model Families on Datacenter Accelerators · CVPR 2021 |
Cloud and datacenter computing › job scheduling
datacenter scheduling |
0.4 | 1 | 2020 | Autonomous Warehouse-Scale Computers · DAC 2020 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2019 | Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019 |
Memory systems › memory interference
memory contention |
0.4 | 1 | 2019 | Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019 |
Cloud and datacenter computing
quality of service |
0.4 | 1 | 2019 | Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019 |
Performance modeling and evaluation › statistical analysis
statistical performance analysis |
0.3 | 1 | 2018 | WSMeter: A Performance Evaluation Methodology for Google's Production Warehouse-Scale Computers · ASPLOS 2018 |
Cloud and datacenter computing
latency-critical applications |
0.3 | 2 | 2015 | Heracles: improving resource efficiency at scale · ISCA 2015 Towards energy proportionality for large-scale latency-critical workloads · ISCA 2014 |
Cloud and datacenter computing
datacenter network |
0.3 | 1 | 2025 | Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage · NSDI 2025 |
Cloud and datacenter computing › cluster resource management and scheduling
workload colocation |
0.2 | 1 | 2016 | Improving Resource Efficiency at Scale with Heracles · ACM Trans. Comput. Syst. 2016 |
Distributed systems
resource sharing |
0.2 | 1 | 2015 | Heracles: improving resource efficiency at scale · ISCA 2015 |
Memory systems
cache coherence |
0.2 | 3 | 2008 | Extending CC-NUMA systems to support write update optimizations · SC 2008 An Adaptive Cache Coherence Protocol Optimized for Producer-Consumer Sharing · HPCA 2007 Interconnect-Aware Coherence Protocols for Chip Multiprocessors · ISCA 2006 |
Memory systems › cache coherence
cache coherence protocol |
0.2 | 3 | 2008 | Extending CC-NUMA systems to support write update optimizations · SC 2008 An Adaptive Cache Coherence Protocol Optimized for Producer-Consumer Sharing · HPCA 2007 Interconnect-Aware Coherence Protocols for Chip Multiprocessors · ISCA 2006 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.2 | 1 | 2023 | TripLe: Revisiting Pretrained Model Reuse and Progressive Learning for Efficient Vision Transformer Scaling and Searching · ICCV 2023 |
Energy-efficient computing
energy proportionality |
0.2 | 1 | 2014 | Towards energy proportionality for large-scale latency-critical workloads · ISCA 2014 |
Energy-efficient computing
power management |
0.2 | 1 | 2014 | Towards energy proportionality for large-scale latency-critical workloads · ISCA 2014 |
Cloud and datacenter computing › datacenter architecture
warehouse-scale computer |
0.2 | 1 | 2014 | Towards energy proportionality for large-scale latency-critical workloads · ISCA 2014 |
Hardware accelerators and domain-specific architectures › accelerator architecture
datacenter accelerator |
0.1 | 1 | 2021 | Searching for Fast Model Families on Datacenter Accelerators · CVPR 2021 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2020 | Autonomous Warehouse-Scale Computers · DAC 2020 |
Machine learning › Efficient and distributed learning
distributed training |
0.1 | 1 | 2019 | Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2018 | WSMeter: A Performance Evaluation Methodology for Google's Production Warehouse-Scale Computers · ASPLOS 2018 |
Performance modeling and evaluation › benchmarking
load testing |
0.1 | 1 | 2018 | WSMeter: A Performance Evaluation Methodology for Google's Production Warehouse-Scale Computers · ASPLOS 2018 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2008 | Extending CC-NUMA systems to support write update optimizations · SC 2008 |
Methods — techniques the papers use, named apart from their topics
weight sharing · 1.3multi-objective optimization · 1.3neural architecture search · 1.0latency-aware optimization · 1.0compound scaling · 1.0runtime isolation · 0.8progressive learning · 0.7knowledge distillation · 0.7machine learning · 0.4hardware-software co-optimization · 0.4statistical modeling · 0.3feedback-based control · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage
Hamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu, Ramesh Govindan, Neal Cardwell, Nandita Dukkipati, Meng-Jung Tsai, Chris DeForeest, Yuxue Jin, Charles J. Carver, Jan Kopanski, Liqun Cheng, Amin Vahdat |
NSDI | 13 |
| 2023 | Hyperscale Hardware Optimized Neural Architecture SearchabstractRecent advances in machine learning have leveraged dramatic increases in computational power, a trend expected to continue in the future. This paper introduces the first Hyperscale Hardware Optimized Neural Architecture Search (H2O-NAS) to automatically design accurate and performant machine learning models tailored to the underlying hardware architecture. H2O-NAS consists of three key components: a new massively parallel “one-shot” search algorithm with intelligent weight sharing, which can scale to search spaces of O(10280) and handle large volumes of production traffic; hardware-optimized search spaces for diverse ML models on heterogeneous hardware; and a novel two-phase hybrid performance model and a multi-objective reward function optimized for large scale deployments. Sheng Li 0007, Garrett Andersen, Tao Chen 0003, Liqun Cheng, Julian Grady, Quoc V. Le, Andrew Li, Xin Li 0082, Yang Li 0005, Yifeng Lu, Yun Ni, Ruoming Pang, Mingxing Tan, Martin Wicke, Shengqi Zhu 0003, Parthasarathy Ranganathan, Norman P. Jouppi |
ASPLOS (3) | 4 |
| 2023 | TripLe: Revisiting Pretrained Model Reuse and Progressive Learning for Efficient Vision Transformer Scaling and SearchingabstractOne promising way to accelerate transformer training is to reuse small pretrained models to initialize the transformer, as their existing representation power facilitates faster model convergence. Previous works designed expansion operators to scale up pretrained models to the target model before training. Yet, model functionality is difficult to preserve when scaling a transformer in all dimensions at once. Moreover, maintaining the pretrained optimizer states for weights is critical for model scaling, whereas the new weights added during expansion lack these states in pre-trained models. To address these issues, we propose TripLe, which partially scales a model before training, while growing the rest of the new parameters during training by copying both the warmed-up weights with the optimizer states from existing weights. As such, the new parameters introduced during training will obtain their training states. Furthermore, through serializing the scaling of model width and depth, the functionality of each expansion can be preserved. We evaluate TripLe in both single-trial model scaling and multi-trial neural architecture search (NAS). Due to the fast training convergence of TripLe, the proxy accuracy from TripLe better reveals the model quality compared to from-scratch training in multi-trial NAS. Experiments show that TripLe outperforms from-scratch training and knowledge distillation (KD) in both training time and task performance. TripLe can also be combined with KD to achieve an even higher task accuracy. For NAS, the model obtained from TripLe outperforms DeiT-B in task accuracy with 69% reduction in parameter size and FLOPs. Cheng Fu 0002, Hanxian Huang, Zixuan Jiang, Yun Ni, Lifeng Nai, Liqun Cheng, Yanqi Zhou, Sheng Li 0007, Andrew Li, Jishen Zhao |
ICCV | 7 |
| 2021 | Searching for Fast Model Families on Datacenter AcceleratorsabstractNeural Architecture Search (NAS), together with model scaling, has shown remarkable progress in designing high accuracy and fast convolutional architecture families. However, as neither NAS nor model scaling considers sufficient hardware architecture details, they do not take full advantage of the emerging datacenter (DC) accelerators. In this paper, we search for fast and accurate CNN model families for efficient inference on DC accelerators. We first analyze DC accelerators and find that existing CNNs suffer from insufficient operational intensity, parallelism, and execution efficiency and exhibit FLOPs-latency nonproportionality. These insights let us create a DC-accelerator-optimized search space, with space-to-depth, space-to-batch, hybrid fused convolution structures with vanilla and depthwise convolutions, and block-wise activation functions. We further propose a latency-aware compound scaling (LACS), the first multi-objective compound scaling method optimizing both accuracy and latency. Our LACS discovers that network depth should grow much faster than image size and network width, which is quite different from the observations from previous compound scaling. With the new search space and LACS, our search and scaling on datacenter accelerators results in a new model series named EfficientNet-X. EfficientNet-X is up to more than 2X faster than Efficient-Net (a model series with state-of-the-art trade-off on FLOPs and accuracy) on TPUv3 and GPUv100, with comparable accuracy. EfficientNet-X is also up to 7X faster than recent RegNet and ResNeSt on TPUv3 and GPUv100. Source code is at https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet/tpu Sheng Li 0007, Mingxing Tan, Ruoming Pang, Andrew Li, Liqun Cheng, Quoc V. Le, Norman P. Jouppi |
CVPR | 5 |
| 2020 | Autonomous Warehouse-Scale ComputersabstractModern Warehouse-Scale Computers (WSCs), composed of many generations of servers and a myriad of domain specific accelerators, are becoming increasingly heterogeneous. Meanwhile, WSC workloads are also becoming incredibly diverse with different communication patterns, latency requirements, and service level objectives (SLOs). Insufficient understanding of the interactions between workload characteristics and the underlying machine architecture leads to resource over-provisioning, thereby significantly impacting the utilization of WSCs. We present Autonomous Warehouse-Scale Computers, a new WSC design that leverages machine learning techniques and automation to improve job scheduling, resource management, and hardware-software co-optimization to address the increasing heterogeneity in WSC hardware and workloads. Our new design introduces two new layers in the WSC stack, namely: (a) a Software-Defined Server (SDS) Abstraction Layer which redefines the hardware-software boundary and provides greater control of the hardware to higher layers of the software stack through stable abstractions; and (b) a WSC Efficiency Layer which regularly monitors the resource usage of workloads on different hardware types, autonomously quantifies the performance sensitivity of workloads to key system configurations, and continuously improves scheduling decisions and hardware resource QoS policies to maximize cluster level performance. Our new WSC design has been successfully deployed across all WSCs at Google for several years now. The new WSC design improves throughput of workloads (by 7-10%, on average), increases utilization of hardware resources (up to 2x), and reduces performance variance for critical workloads (up to 25%). Sundar Dev, David Lo 0003, Liqun Cheng, Parthasarathy Ranganathan |
DAC | 3 |
| 2019 | Kelp: QoS for Accelerated Machine Learning SystemsabstractDevelopment and deployment of machine learning (ML) accelerators in Warehouse Scale Computers (WSCs) demand significant capital investments and engineering efforts. However, even though heavy computation can be offloaded to the accelerators, applications often depend on the host system for various supporting tasks. As a result, contention on host resources, such as memory bandwidth, can significantly discount the performance and efficiency gains of accelerators. The impact of performance interference is further amplified in distributed learning, which has become increasingly common as model sizes continue to grow. In this work, we study the performance of four production machine learning workloads on three accelerator platforms. Our experiments show that these workloads are highly sensitive to host memory bandwidth contention, which can cause 40% average performance degradation when left unmanaged. To tackle this problem, we design and implement Kelp, a software runtime that isolates high priority accelerated ML tasks from memory resource interference. We evaluate Kelp with both production and artificial aggressor workloads, and compare its effectiveness with previously proposed solutions. Our evaluation shows that Kelp is effective in mitigating performance degradation of the accelerated tasks, and improves performance by 24% on average. Compared to previous work, Kelp reduces performance degradation of ML tasks by 7% and improves system efficiency by 17%. Our results further expose opportunities in future architecture designs. Haishan Zhu, David Lo 0003, Liqun Cheng, Rama Govindaraju, Parthasarathy Ranganathan, Mattan Erez |
HPCA | 3 |
| 2018 | WSMeter: A Performance Evaluation Methodology for Google's Production Warehouse-Scale ComputersabstractEvaluating the comprehensive performance of a warehouse-scale computer (WSC) has been a long-standing challenge. Traditional load-testing benchmarks become ineffective because they cannot accurately reproduce the behavior of thousands of distinct jobs co-located on a WSC. We therefore evaluate WSCs using actual job behaviors in live production environments. From our experience of developing multiple generations of WSCs, we identify two major challenges of this approach: 1) the lack of a holistic metric that incorporates thousands of jobs and summarizes the performance, and 2) the high costs and risks of conducting an evaluation in a live environment. To address these challenges, we propose WSMeter, a cost-effective methodology to accurately evaluate a WSC's performance using a live production environment. We first define a new metric which accurately represents a WSC's overall performance, taking a wide variety of unevenly distributed jobs into account. We then propose a model to statistically embrace the performance variance inherent in WSCs, to conduct an evaluation with minimal costs and risks. We present three real-world use cases to prove the effectiveness of WSMeter. In the first two cases, WSMeter accurately discerns 7% and 1% performance improvements from WSC upgrades using only 0.9% and 6.6% of the machines in the WSCs, respectively. We emphasize that naive statistical comparisons incur much higher evaluation costs (> 4 times) and sometimes even fail to distinguish subtle differences. The third case shows that a cloud customer hosting two services on our WSC quantifies the performance benefits of software optimization (+9.3%) with minimal overheads (2.3% of the service capacity). Changkyu Kim, Liqun Cheng, Rama Govindaraju, Jangwoo Kim |
ASPLOS | 4 |
| 2016 | Improving Resource Efficiency at Scale with HeraclesabstractUser-facing, latency-sensitive services, such as websearch, underutilize their computing resources during daily periods of low traffic. Reusing those resources for other tasks is rarely done in production services since the contention for shared resources can cause latency spikes that violate the service-level objectives of latency-sensitive tasks. The resulting under-utilization hurts both the affordability and energy efficiency of large-scale datacenters. With the slowdown in technology scaling caused by the sunsetting of Moore’s law, it becomes important to address this opportunity. We present Heracles, a feedback-based controller that enables the safe colocation of best-effort tasks alongside a latency-critical service. Heracles dynamically manages multiple hardware and software isolation mechanisms, such as CPU, memory, and network isolation, to ensure that the latency-sensitive job meets latency targets while maximizing the resources given to best-effort tasks. We evaluate Heracles using production latency-critical and batch workloads from Google and demonstrate average server utilizations of 90% without latency violations across all the load and colocation scenarios that we evaluated. David Lo 0003, Liqun Cheng, Rama Govindaraju, Parthasarathy Ranganathan, Christoforos E. Kozyrakis |
ACM Trans. Comput. Syst. | 2 |
| 2015 | Heracles: improving resource efficiency at scaleabstractUser-facing, latency-sensitive services, such as websearch, underutilize their computing resources during daily periods of low traffic. Reusing those resources for other tasks is rarely done in production services since the contention for shared resources can cause latency spikes that violate the service-level objectives of latency-sensitive tasks. The resulting under-utilization hurts both the affordability and energy-efficiency of large-scale datacenters. With technology scaling slowing down, it becomes important to address this opportunity. David Lo 0003, Liqun Cheng, Rama Govindaraju, Parthasarathy Ranganathan, Christoforos E. Kozyrakis |
ISCA | 2 |
| 2014 | Towards energy proportionality for large-scale latency-critical workloadsabstractReducing the energy footprint of warehouse-scale computer (WSC) systems is key to their affordability, yet difficult to achieve in practice. The lack of energy proportionality of typical WSC hardware and the fact that important workloads (such as search) require all servers to remain up regardless of traffic intensity renders existing power management techniques ineffective at reducing WSC energy use. We present PEGASUS, a feedback-based controller that significantly improves the energy proportionality of WSC systems, as demonstrated by a real implementation in a Google search cluster. PEGASUS uses request latency statistics to dynamically adjust server power management limits in a fine-grain manner, running each server just fast enough to meet global service-level latency objectives. In large cluster experiments, PEGASUS reduces power consumption by up to 20%. We also estimate that a distributed version of PEGASUS can nearly double these savings. David Lo 0003, Liqun Cheng, Rama Govindaraju, Luiz André Barroso, Christoforos E. Kozyrakis |
ISCA | 2 |
| 2008 | Extending CC-NUMA systems to support write update optimizationsabstractProcessor stalls and protocol messages caused by coherence misses limit the performance of shared memory applications. Modern multiprocessors employ write-invalidate coherence protocols, which induce read misses to ensure consistency. Previous research has shown that an invalidate protocol is not optimal for all memory access patterns - an update protocol can significantly outperform an invalidate protocol when data is heavily shared or accessed in predictable patterns. However, update protocols can generate excessive network traffic and are difficult to build on a scalable (non-bus) interconnect. To obtain the benefits of both invalidate and update protocols, we built a speculative sequentially consistent write- update mechanism on top of a write-invalidate protocol. To ensure coherence, a processor wishing to write to a block of data uses a traditional write-invalidate protocol to obtain exclusive access to the block before modifying it. To improve performance, the writing processor can later self- downgrade the modified block to the shared state and flush it back to its home node, which forwards the new data to processors that it predicts are likely to consume the data. We present a practical and cost-effective design for extending CC-NUMA systems to support this speculative update mechanism that requires no changes to the processor core, bus interface, or memory consistency model. We also present two hardware-efficient mechanisms for detecting access patterns that benefit from the speculative update mechanism, stable reader set and stream. We evaluate our update mechanisms on a wide range of scientific benchmarks and commercial applications. Using a cycle-accurate execution-driven simulator of a future 16-node SGI multiprocessor, we find that the mechanisms proposed in this paper reduce the average remote miss rate by 30%, reduce network traffic by 15%, and improve performance by 10%, and in no case hurt performance. Liqun Cheng, John B. Carter |
SC | 1 |
| 2007 | An Adaptive Cache Coherence Protocol Optimized for Producer-Consumer SharingabstractShared memory multiprocessors play an increasingly important role in enterprise and scientific computing facilities. Remote misses limit the performance of shared memory applications, and their significance is growing as network latency increases relative to processor speeds. This paper proposes two mechanisms that improve shared memory performance by eliminating remote misses and/or reducing the amount of communication required to maintain coherence. We focus on improving the performance of applications that exhibit producer-consumer sharing. We first present a simple hardware mechanism for detecting producer-consumer sharing. We then describe a directory delegation mechanism whereby the "home node" of a cache line can be delegated to a producer node, thereby converting 3-hop coherence operations into 2-hop operations. We then extend the delegation mechanism to support speculative updates for data accessed in a producer-consumer pattern, which can convert 2-hop misses into local misses, thereby eliminating the remote memory latency. Both mechanisms can be implemented without changes to the processor. We evaluate our directory delegation and speculative update mechanisms on seven benchmark programs that exhibit producer-consumer sharing using a cycle-accurate execution-driven simulator of a future 16-node SGI multiprocessor. We find that the mechanisms proposed in this paper reduce the average remote miss rate by 40%, reduce network traffic by 15%, and improve performance by 21%. Finally, we use Murphi to verify that each mechanism is error-free and does not violate sequential consistency Liqun Cheng, John B. Carter, Donglai Dai |
HPCA | 1 |
| 2006 | Interconnect-Aware Coherence Protocols for Chip MultiprocessorsabstractImprovements in semiconductor technology have made it possible to include multiple processor cores on a single die. Chip Multi-Processors (CMP) are an attractive choice for future billion transistor architectures due to their low design complexity, high clock frequency, and high throughput. In a typical CMP architecture, the L2 cache is shared by multiple cores and data coherence is maintained among private L1s. Coherence operations entail frequent communication over global on-chip wires. In future technologies, communication between different L1s will have a significant impact on overall processor performance and power consumption. On-chip wires can be designed to have different latency, bandwidth, and energy properties. Likewise, coherence protocol messages have different latency and bandwidth needs. We propose an interconnect composed of wires with varying latency, bandwidth, and energy characteristics, and advocate intelligently mapping coherence operations to the appropriate wires. In this paper, we present a comprehensive list of techniques that allow coherence protocols to exploit a heterogeneous interconnect and evaluate a subset of these techniques to show their performance and power-efficiency potential. Most of the proposed techniques can be implemented with a minimum complexity overhead. Liqun Cheng, Naveen Muralimanohar, Karthik Ramani, Rajeev Balasubramonian, John B. Carter |
ISCA | 1 |
| 2005 | Fast Barriers for Scalable ccNUMA SystemsabstractThe contributions of this paper are threefold. First, we identify and quantify the performance deficiencies of conventional barrier implementations when they are executed on real (non-idealized) hardware. Second, we propose a queue-based barrier algorithm that has effectively O(1) time complexity as measured in round trip message latencies. Third, we demonstrate how matching the barrier implementation to the way that modern shared memory systems operate can improve performance dramatically by exploiting a hardware write-update (PUT) mechanism for signaling. The resulting barrier algorithm only costs one serialized round trip message latency to perform a barrier operation across N processors. Using a cycle-accurate execution-driven simulator of a future-generation SGI multiprocessor, we show that with no special hardware support our queue-based barrier outperforms OpenMP's LL/SC-based barrier implementation by a factor of 7.9 on 256 processors. With hardware that supports a coherent PUT operation, our queue-based barrier outperforms OpenMP barriers by a factor of 94 and outperforms barriers based on SGI's memory controller-based atomic operations by a factor of 6.5 on 256 processors. Liqun Cheng, John B. Carter |
ICPP | 1 |
| 2005 | Fast synchronization on shared-memory multiprocessors: An architectural approach
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Liqun Cheng, Michael A. Parker |
J. Parallel Distributed Comput. | 4 |