VLDB 2026 Research / reviewers in the wild / expert
Lizhong Chen
dblp:78/4756
· DBLP profile ↗
41ranked-venue papers
8as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
15 papers |
Hardware accelerators and domain-specific architectures · 26% Parallel and multicore computing · 14% Memory systems · 14% | |
| Artificial intelligence
7 papers |
Deep learning architectures and training · 51% Machine translation · 46% Language models and text generation · 3% |
Topics — the 30 heaviest of 51, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
simultaneous machine translation |
1.7 | 3 | 2024 | Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation · EMNLP 2024 Simul-LLM: A Framework for Exploring High-Quality Simultaneous Translation with Large Language Models · ACL (1) 2024 LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions · ICML 2024 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.6 | 2 | 2026 | FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer · AAAI 2026 Polymorphic Accelerators for Deep Neural Networks · IEEE Trans. Computers 2022 |
Machine learning › Deep learning architectures and training
transformer |
1.1 | 2 | 2026 | LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions · ICML 2024 FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer · AAAI 2026 |
Parallel and multicore computing
kernel optimization |
1.0 | 1 | 2026 | FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer · AAAI 2026 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
1.0 | 2 | 2022 | Polymorphic Accelerators for Deep Neural Networks · IEEE Trans. Computers 2022 Shortcut Mining: Exploiting Cross-Layer Shortcut Reuse in DCNN Accelerators · HPCA 2019 |
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.8 | 1 | 2024 | LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions · ICML 2024 |
Machine learning › Deep learning architectures and training › transformer
efficient transformer |
0.8 | 1 | 2024 | LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions · ICML 2024 |
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention |
0.8 | 1 | 2024 | LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions · ICML 2024 |
Natural language and speech › Machine translation › simultaneous machine translation
simultaneous speech translation |
0.7 | 1 | 2023 | Shiftable Context: Addressing Training-Inference Context Mismatch in Simultaneous Speech Translation · ICML 2023 |
Natural language and speech › Machine translation › neural machine translation
training-inference discrepancy |
0.7 | 1 | 2023 | Shiftable Context: Addressing Training-Inference Context Mismatch in Simultaneous Speech Translation · ICML 2023 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable accelerator |
0.6 | 1 | 2022 | Polymorphic Accelerators for Deep Neural Networks · IEEE Trans. Computers 2022 |
Energy-efficient computing
power gating |
0.6 | 3 | 2015 | Power punch: Towards non-blocking power-gating of NoC routers · HPCA 2015 MP3: Minimizing performance penalty for power-gating of Clos network-on-chip · HPCA 2014 NoRD: Node-Router Decoupling for Effective Power-gating of On-Chip Routers · MICRO 2012 |
Integrated circuit design › packaging › advanced packaging
2.5d integration |
0.4 | 1 | 2020 | EquiNox: Equivalent NoC Injection Routers for Silicon Interposer-Based Throughput Processors · HPCA 2020 |
Electronic design automation
design space exploration |
0.4 | 1 | 2020 | A Deep Reinforcement Learning Framework for Architectural Exploration: A Routerless NoC Case Study · HPCA 2020 |
Integrated circuit design › packaging › advanced packaging
silicon interposer |
0.4 | 1 | 2020 | EquiNox: Equivalent NoC Injection Routers for Silicon Interposer-Based Throughput Processors · HPCA 2020 |
Processor architecture and microarchitecture
many-core architecture |
0.4 | 2 | 2020 | Providing Balanced Mapping for Multiple Applications in Many-Core Chip Multiprocessors · IEEE Trans. Computers 2016 EquiNox: Equivalent NoC Injection Routers for Silicon Interposer-Based Throughput Processors · HPCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
DCNN accelerator |
0.4 | 1 | 2019 | Shortcut Mining: Exploiting Cross-Layer Shortcut Reuse in DCNN Accelerators · HPCA 2019 |
Memory systems › on-chip memory
on-chip buffer management |
0.4 | 1 | 2019 | Shortcut Mining: Exploiting Cross-Layer Shortcut Reuse in DCNN Accelerators · HPCA 2019 |
Memory systems › data locality
on-chip data reuse |
0.4 | 1 | 2019 | Shortcut Mining: Exploiting Cross-Layer Shortcut Reuse in DCNN Accelerators · HPCA 2019 |
Interconnection networks and networks-on-chip › router architecture
network-on-chip router |
0.4 | 2 | 2015 | Power punch: Towards non-blocking power-gating of NoC routers · HPCA 2015 NoRD: Node-Router Decoupling for Effective Power-gating of On-Chip Routers · MICRO 2012 |
Energy-efficient computing
leakage power reduction |
0.3 | 2 | 2014 | MP3: Minimizing performance penalty for power-gating of Clos network-on-chip · HPCA 2014 NoRD: Node-Router Decoupling for Effective Power-gating of On-Chip Routers · MICRO 2012 |
Machine learning › Deep learning architectures and training › feedforward neural network
kolmogorov-arnold networks |
0.3 | 1 | 2026 | FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer · AAAI 2026 |
Parallel and multicore computing › task allocation
computation-to-core mapping |
0.2 | 1 | 2016 | Providing Balanced Mapping for Multiple Applications in Many-Core Chip Multiprocessors · IEEE Trans. Computers 2016 |
Memory systems › memory bandwidth management
memory bandwidth reduction |
0.2 | 1 | 2016 | A Filtering Mechanism to Reduce Network Bandwidth Utilization of Transaction Execution · ACM Trans. Archit. Code Optim. 2016 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.2 | 2 | 2016 | In-network traffic regulation for Transactional Memory · HPCA 2013 A Filtering Mechanism to Reduce Network Bandwidth Utilization of Transaction Execution · ACM Trans. Archit. Code Optim. 2016 |
Parallel and multicore computing
transactional memory |
0.2 | 2 | 2016 | In-network traffic regulation for Transactional Memory · HPCA 2013 A Filtering Mechanism to Reduce Network Bandwidth Utilization of Transaction Execution · ACM Trans. Archit. Code Optim. 2016 |
Natural language and speech › Language models and text generation › neural language model
autoregressive language model |
0.2 | 1 | 2024 | LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions · ICML 2024 |
Natural language and speech › Machine translation › speech translation
speech-to-text translation |
0.2 | 1 | 2024 | LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions · ICML 2024 |
Energy-efficient computing
power management |
0.2 | 1 | 2015 | Power punch: Towards non-blocking power-gating of NoC routers · HPCA 2015 |
Natural language and speech › Machine translation › speech translation
streaming speech translation |
0.2 | 1 | 2023 | Shiftable Context: Addressing Training-Inference Context Mismatch in Simultaneous Speech Translation · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
gradient accumulation · 2.0GPU kernel optimization · 2.0fine-tuning · 1.5cross-layer data reuse · 1.1monte carlo tree search · 0.9position re-weighting · 0.8linear attention · 0.8evaluation pipeline · 0.8attention masking · 0.8wait-k decoding · 0.7transformer · 0.7resource partitioning · 0.6dataflow architecture · 0.6n-queen · 0.4deep reinforcement learning · 0.4full-system simulation · 0.4FPGA prototyping · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold TransformerabstractThe Kolmogorov-Arnold Network (KAN) has been gaining popularity as an alternative to the multilayer perceptron (MLP) due to its greater expressiveness and interpretability. Even so, KAN suffers from training instability and being orders of magnitude slower due to its increased computational cost, limiting its applicability to large-scale tasks. Recently, the Kolmogorov-Arnold Transformer (KAT) has been proposed, achieving FLOPs comparable to traditional Transformer models with MLPs by leveraging Group-Rational KAN (GR-KAN). Unfortunately, despite the comparable FLOPs, our testing shows that KAT remains 123x slower during training, indicating that there are other performance bottlenecks beyond FLOPs. In this paper, we conduct a series of experiments to understand the root cause of the slowdown in KAT. We uncover that the slowdown can be isolated to memory stalls, linked more specifically to inefficient gradient accumulations in the backward pass of GR-KAN. To address this memory bottleneck, we propose FlashKAT, which minimizes accesses to slow memory and the usage of atomic adds through a restructured kernel. Evaluations show that FlashKAT achieves up to an 86.5x training speedup over state-of-the-art KAT while reducing rounding errors in gradient computation. Matthew Raffel, Lizhong Chen |
AAAI | 2 |
| 2024 | Simul-LLM: A Framework for Exploring High-Quality Simultaneous Translation with Large Language ModelsabstractLarge language models (LLMs) with billions of parameters and pretrained on massive amounts of data are now capable of near or better than state-of-the-art performance in a variety of downstream natural language processing tasks.Neural machine translation (NMT) is one such task that LLMs have been applied to with great success.However, little research has focused on applying LLMs to the more difficult subset of NMT called simultaneous translation (SimulMT), where translation begins before the entire source context is available to the model.In this paper, we address key challenges facing LLMs fine-tuned for SimulMT, validate classical SimulMT concepts and practices in the context of LLMs, explore adapting LLMs that are fine-tuned for NMT to the task of SimulMT, and introduce Simul-LLM 1 , the first open-source fine-tuning and evaluation pipeline development framework for LLMs focused on SimulMT. Victor Agostinelli, Max Wild, Matthew Raffel, Kazi Ahmed Asif Fuad, Lizhong Chen |
ACL (1) | 5 |
| 2024 | Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous TranslationabstractLarge language models (LLMs) have achieved state-of-the-art performance in various language processing tasks, motivating their adoption in simultaneous translation.Current finetuning methods to adapt LLMs for simultaneous translation focus on prompting optimization strategies using either data augmentation or prompt structure modifications.However, these methods suffer from several issues, such as unnecessarily expanded training sets, computational inefficiency from dumping the key and value cache, increased prompt sizes, or restriction to a single decision policy.To eliminate these issues, in this work, we propose Simul-Mask, a new paradigm for fine-tuning LLMs for simultaneous translation.It utilizes a novel attention mask approach that models simultaneous translation during fine-tuning by masking attention for a desired decision policy.Applying the proposed SimulMask on a Falcon LLM for the IWSLT 2017 dataset, we have observed a significant translation quality improvement compared to state-of-the-art prompting optimization strategies on five language pairs while reducing the computational cost. Matthew Raffel, Victor Agostinelli, Lizhong Chen |
EMNLP | 3 |
| 2024 | LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned ProportionsabstractA promising approach to preserving model performance in linearized transformers is to employ position-based re-weighting functions. However, state-of-the-art re-weighting functions rely heavily on target sequence lengths, making it difficult or impossible to apply them to autoregressive and simultaneous tasks, where the target and sometimes even the input sequence length are unknown. To address this issue, we propose Learned Proportions (LeaP) and LeaPformers. Our contribution is built on two major components. First, we generalize the dependence on explicit positional representations and sequence lengths into dependence on sequence proportions for re-weighting. Second, we replace static positional representations with dynamic proportions derived via a compact module, enabling more flexible attention concentration patterns. We evaluate LeaPformer against eight representative efficient transformers on the Long-Range Arena benchmark, where we show that LeaPformer achieves the best quality-throughput trade-off, as well as apply LeaPformer to Wikitext-103b autoregressive language modeling and simultaneous speech-to-text translation for two language pairs, achieving competitive results in both tasks. Victor Agostinelli, Sanghyun Hong 0001, Lizhong Chen |
ICML | 3 |
| 2023 | Shiftable Context: Addressing Training-Inference Context Mismatch in Simultaneous Speech TranslationabstractTransformer models using segment-based processing have been an effective architecture for simultaneous speech translation. However, such models create a context mismatch between training and inference environments, hindering potential translation accuracy. We solve this issue by proposing Shiftable Context, a simple yet effective scheme to ensure that consistent segment and context sizes are maintained throughout training and inference, even with the presence of partially filled segments due to the streaming nature of simultaneous translation. Shiftable Context is also broadly applicable to segment-based transformers for streaming tasks. Our experiments on the English-German, English-French, and English-Spanish language pairs from the MUST-C dataset demonstrate that when applied to the Augmented Memory Transformer, a state-of-the-art model for simultaneous speech translation, the proposed scheme achieves an average increase of 2.09, 1.83, and 1.95 BLEU scores across each wait-k value for the three language pairs, respectively, with a minimal impact on computation-aware Average Lagging. Matthew Raffel, Drew Penney, Lizhong Chen |
ICML | 3 |
| 2023 | Improving Autoregressive NLP Tasks via Modular Linearized Attention
Victor Agostinelli, Lizhong Chen |
ECML/PKDD (4) | 2 |
| 2023 | PROMPT: Learning dynamic resource allocation policies for network applications
Drew Penney, Bin Li 0018, Jaroslaw J. Sydir, Lizhong Chen, Tsung-Yuan Charlie Tai, Stefan Lee, Eoin Walsh, Thomas Long |
Future Gener. Comput. Syst. | 4 |
| 2023 | RAPID: Enabling fast online policy learning in dynamic public cloud environments
Drew Penney, Bin Li 0018, Lizhong Chen, Jaroslaw J. Sydir, Anna Drewek-Ossowicka, Ramesh Illikkal, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Andrew Herdrich |
Neurocomputing | 3 |
| 2022 | Polymorphic Accelerators for Deep Neural NetworksabstractDeep neural networks (DNNs) come with many forms, such as convolutional neural networks, multilayer perceptron, and recurrent neural networks, to meet diverse needs of machine learning applications. However, existing DNN accelerator designs, when used to execute multiple neural networks, suffer from underutilization of processing elements, heavy feature map traffic, and large area overhead. In this article, we propose a novel approach,Polymorphic Accelerators, to address the flexibility issue fundamentally. We introduce the abstraction of logical accelerators to decouple the fixed mapping with physical resources. Three procedures are proposed that work collaboratively to reconfigure the accelerator for the current network that is being executed and to enable cross-layer data reuse among logical accelerators. Evaluation results show that the proposed approach achieves significant improvement in data reuse, inference latency and performance, e.g., 1.52x and 1.63x increase in throughput compared with state-of-the-art flexible dataflow approach and resource partitioning approach, respectively. This demonstrates the effectiveness and promise of polymorphic accelerator architecture. Arash AziziMazreah, Lizhong Chen |
IEEE Trans. Computers | 2 |
| 2020 | EquiNox: Equivalent NoC Injection Routers for Silicon Interposer-Based Throughput ProcessorsabstractThroughput-oriented many-core processors demand highly efficient network-on-chip (NoC) architecture for data transferring. Recent advent of silicon interposer, stacked memory and 2.5D integration have further increased data transfer rate. This greatly intensifies traffic bottleneck in the NoC but, at the same time, also brings a significant new opportunity in utilizing wiring resources in the interposer. In this paper, we propose a novel concept called Equivalent Injection Routers (EIRs) which, together with interposer links, transform the few-to-many traffic pattern to many-to-many pattern, thus fundamentally solving the bottleneck problem. We have developed EquiNox as a design example. We utilize N-Queen and Monte Carlo Tree Search (MCTS) methods to help select EIRs by considering comprehensively from topological, architectural and physical aspects. Evaluation results show that, compared with prior work, the proposed EquiNox is able to reduce execution time by 23.5%, energy consumption by 18.9%, and EDP by 32.8%, under similar hardware cost. Yunfan Li 0002, Lizhong Chen |
HPCA | 2 |
| 2020 | A Deep Reinforcement Learning Framework for Architectural Exploration: A Routerless NoC Case StudyabstractMachine learning applied to architecture design presents a promising opportunity with broad applications. Recent deep reinforcement learning (DRL) techniques, in particular, enable efficient exploration in vast design spaces where conventional design strategies may be inadequate. This paper proposes a novel deep reinforcement framework, taking routerless networks-on-chip (NoC) as an evaluation case study. The new framework successfully resolves problems with prior design approaches, which are either unreliable due to random searches or inflexible due to severe design space restrictions. The framework learns (near-)optimal loop placement for routerless NoCs with various design constraints. A deep neural network is developed using parallel threads that efficiently explore the immense routerless NoC design space with a Monte Carlo search tree. Experimental results show that, compared with conventional mesh, the proposed deep reinforcement learning (DRL) routerless design achieves a 3.25x increase in throughput, 1.6x reduction in packet latency, and 5x reduction in power. Compared with the state-of-the-art routerless NoC, DRL achieves a 1.47x increase in throughput, 1.18x reduction in packet latency, 1.14x reduction in average hop count, and 6.3% lower power consumption. Ting-Ru Lin, Drew Penney, Massoud Pedram, Lizhong Chen |
HPCA | 4 |
| 2020 | Accelerated Reply Injection for Removing NoC Bottleneck in GPGPUsabstractThe high level of parallelism in GPGPUs has resulted in significantly changed on-chip data traffic behaviors. This demands new research to identify and address the limiting factors of networks-on-chip (NoCs) in the context of GPGPUs. In this paper, we quantitatively analyze the performance of on-chip networks in GPGPUs, and address a serious NoC bottleneck where the reply data from memory controllers experience large contention when being injected to the reply network. To remove this reply injection bottleneck, we propose Accelerated Reply Injection (ARI), a very effective scheme that can supply a fast rate of data traffic from memory controllers to feed the reply injection points, and accelerates the consumption of the injected packets by quickly transferring the packets out of the injection points, thus increasing both supply and consumption of reply traffic injection. Simulation results on a wide range of benchmarks show that the proposed ARI reduces the data stall time in memory controllers by 67.8% on average, and increases IPC by more than 15.4% on average, with less than 1% area overhead. Yunfan Li 0002, Lizhong Chen |
IPDPS | 2 |
| 2019 | Shortcut Mining: Exploiting Cross-Layer Shortcut Reuse in DCNN AcceleratorsabstractOff-chip memory traffic has been a major performance bottleneck in deep learning accelerators. While reusing on-chip data is a promising way to reduce off-chip traffic, the opportunity on reusing shortcut connection data in deep networks (e.g., residual networks) have been largely neglected. Those shortcut data accounts for nearly 40% of the total feature map data. In this paper, we propose Shortcut Mining, a novel approach that “mines” the unexploited opportunity of on-chip data reusing. We introduce the abstraction of logical buffers to address the lack of flexibility in existing buffer architecture, and then propose a sequence of procedures which, collectively, can effectively reuse both shortcut and non-shortcut feature maps. The proposed procedures are also able to reuse shortcut data across any number of intermediate layers without using additional buffer resources. Experiment results from prototyping on FPGAs show that, the proposed Shortcut Mining achieves 53.3%, 58%, and 43% reduction in off-chip feature map traffic for SqueezeNet, ResNet-34, and ResNet152, respectively and a 1.93X increase in throughput compared with a state-of-the-art accelerator. Arash AziziMazreah, Lizhong Chen |
HPCA | 2 |
| 2019 | Characterizing On-Chip Traffic Patterns in General-Purpose GPUs: A Deep Learning ApproachabstractArchitectural optimizations in general-purpose graphics processing units (GPGPUs) often exploit workload characteristics to reduce power and latency while improving performance. This paper finds, however, that prevailing assumptions about GPGPU traffic pattern characterization are inaccurate. These assumptions must therefore be re-evaluated, and more appropriate new patterns must be identified. This paper proposes a methodology to classify GPGPU traffic patterns, combining a convolutional neural network (CNN) for feature extraction and a t-distributed stochastic neighbor embedding (t-SNE) algorithm to determine traffic pattern clusters. A traffic pattern dataset is generated from common GPGPU benchmarks, transformed using heat mapping, and iteratively refined to ensure appropriate and highly accurate labels. The proposed classification model achieves 98.8% validation accuracy and 94.24% test accuracy. Furthermore, traffic in 96.6% of examined kernels can be classified into the eight identified traffic pattern categories. Yunfan Li 0002, Drew Penney, Abhishek Ramamurthy, Lizhong Chen |
ICCD | 4 |
| 2019 | Express Link Placement for NoC-Based Many-Core PlatformsabstractWith the integration of up to hundreds of cores in recent general-purpose processors that can be used in parallel processing systems, it is critical to design scalable and low-latency networks-on-chip (NoCs) to support various on-chip communications. An effective way to reduce on-chip latency and improve network scalability is to add express links between pairs of non-adjacent routers. However, increasing the number of express links may result in smaller bandwidth per link due to the limited total bisection bandwidth on chip, thus leading to higher serialization latency of packets in the network. Unlike previous works on application-specific designs or on fixed placement of express links, this paper aims at finding effective placement of express links for general-purpose processors considering all the possible placement options. We formulate the problem mathematically and propose an efficient algorithm that utilizes an initial solution generation heuristic and enhanced candidate generator in simulated annealing. Evaluation on 4x4, 8x8 and 16x16 networks using multi-threaded PARSEC benchmarks and various synthetic traffic patterns shows significant reduction of average packet latency over previous works. Yunfan Li 0002, Di Zhu 0002, Lizhong Chen |
ICPP | 3 |
| 2019 | Dynamically linked MSHRs for adaptive miss handling in GPUsabstractSupporting a large number of outstanding memory requests in miss handling architecture (MHA) is critical for throughput processors such as GPUs to achieve high memory level parallelism. Conventional MHA is static in sense that it provides a fixed number of MSHR entries to track primary misses, and a fixed number of slots within each entry to track secondary misses. This leads to severe entry or slot under-utilization and poor match to practical workloads, as the number of memory requests to different cache lines can vary significantly. In this paper, we propose Dynamically Linked MSHR (DL-MSHR), a novel approach that dynamically forms MSHR entries from a pool of available slots. This approach can self-adapt to primary-miss-predominant applications by forming more entries with fewer slots, and self-adapt to secondary-miss-predominant applications by having fewer entries but more slots per entry. Evaluation results show that, compared with the conventional MSHRs, the proposed DL-MSHR is able to reduce reservation fails in MSHRs by 88.1%, improve MSHR utilization by 53.7% and increase the overall IPC of a wide range of workloads by 19.2%, on average, with only 0.6% and 0.1% area overhead on L1D and L2 cache, respectively. Yongbin Gu, Lizhong Chen |
ICS | 2 |
| 2019 | On Trade-off Between Static and Dynamic Power Consumption in NoC Power GatingabstractRecent research has proposed to minimize network-on-chip (NoC) static power by proactively power-gating selected routers when not all the cores are active. However, as more routers are powered off, on-chip packets are forced to take detours more frequently, resulting in a higher hop count and increased dynamic power that may potentially offset the static power savings. This paper investigates such a trade-off between static and dynamic power in detail, and explores how overall NoC power consumption can be minimized through proactive power-gating. Three efficient and effective algorithms are proposed to reduce NoC static power, dynamic power, and overall power consumption, respectively. Evaluation results based on PARSEC benchmarks demonstrate the importance of the trade-off and show a substantial improvement in total NoC power savings of the proposed algorithms, compared with previous work that did not give full consideration to both static and dynamic power. Di Zhu 0002, Yunfan Li 0002, Lizhong Chen |
ISLPED | 3 |
| 2018 | Routerless Network-on-ChipabstractTraditional bus-based interconnects are simple and easy to implement, but the scalability is greatly limited. While router-based networks-on-chip (NoCs) offer superior scalability, they also incur significant power and area overhead due to complex router structures. In this paper, we explore a new class of on-chip networks, referred to as Routerless NoCs, where routers are completely eliminated. We propose a novel design that utilizes on-chip wiring resources smartly to achieve comparable hop count and scalability as router-based NoCs. Several effective techniques are also proposed that significantly reduce the resource requirement to avoid new network abnormalities in routerless NoC designs. Evaluation results show that, compared with a conventional mesh, the proposed routerless NoC achieves 9.5X reduction in power, 7.2X reduction in area, 2.5X reduction in zero-load packet latency, and 1.7X increase in throughput. Compared with a state-of-the-art low-cost NoC design, the proposed approach achieves 7.7X reduction in power, 3.3X reduction in area, 1.3X reduction in zero-load packet latency, and 1.6X increase in throughput. Fawaz Alazemi, Arash AziziMazreah, Bella Bose, Lizhong Chen |
HPCA | 4 |
| 2018 | CART: Cache Access Reordering Tree for Efficient Cache and Memory Accesses in GPUsabstractGraphics processing units (GPUs) have been increasingly used to accelerate general purpose computing. Thousands of concurrently running threads in a GPU demand a highly efficient memory subsystem for data supply. A key factor that affects the memory subsystem is the order of memory accesses. While reordering memory accesses at L2 cache has large potential benefits to both cache and DRAM, little work has been conducted to exploit this. In this paper, we investigate the largely unexplored opportunity of L2 cache access reordering. We propose Cache Access Reordering Tree (CART), a novel architecture that can improve memory subsystem efficiency by actively reordering memory accesses at L2 cache to be cache-friendly and DRAM-friendly. Evaluation results using a wide range of benchmarks show that, the proposed CART is able to improve the average IPC of memory intensive benchmarks by 34.2% with only 1.7% area overhead. Yongbin Gu, Lizhong Chen |
ICCD | 2 |
| 2018 | Tolerating Soft Errors in Deep Learning Accelerators with Reliable On-Chip Memory DesignsabstractDeep learning neural network (DNN) accelerators have been increasingly deployed in many fields recently, including safety-critical applications such as autonomous vehicles and unmanned aircrafts. Meanwhile, the vulnerability of DNN accelerators to soft errors (e.g., caused by high-energy particle strikes) rapidly increases as manufacturing technology continues to scale down. A failure in the operation of DNN accelerators may lead to catastrophic consequences. Among the existing reliability techniques that can be applied to DNN accelerators, fully-hardened SRAM cells are more attractive due to their low overhead in terms of area, power and delay. However, current fully-hardened SRAM cells can only tolerate soft errors produced by single-node-upsets (SNUs), and cannot fully resist the soft errors caused by multiple-node-upsets (MNUs). In this paper, a Zero-Biased MNU-Aware SRAM Cell (ZBMA) is proposed for DNN accelerators based on two observations: first, the data (feature maps, weights) in DNNs has a strong bias towards zero; second, data flipping from zero to one is more likely to cause a failure of DNN outputs. The proposed memory cell provides a robust immunity against node upsets, and reduces the leakage current dramatically when zero is stored in the cell. Evaluation results show that when the proposed memory cell is integrated in a DNN accelerator, the total static power of the accelerator is reduced by 2.6X and 1.79X compared with the one based on the conventional and on state-of-the-art full-hardened memory cells, respectively. In terms of reliability, the DNN accelerator based on the proposed memory cell can reduce 99.99% of false outputs caused by soft errors across different DNNs. Arash AziziMazreah, Yongbin Gu, Lizhong Chen |
NAS | 4 |
| 2017 | XPro: A Cross-End Processing Architecture for Data Analytics in WearablesabstractWearable computing systems have spurred many opportunities to continuously monitor human bodies with sensors worn on or implanted in the body. These emerging platforms have started to revolutionize many fields, including healthcare and wellness applications, particularly when integrated with intelligent analytic capabilities. However, a significant challenge that computer architects are facing is how to embed sophisticated analytic capabilities in wearable computers in an energy-efficient way while not compromising system performance. In this paper, we present XPro, a novel cross-end analytic engine architecture for wearable computing systems. The proposed cross-end architecture is able to realize a generic classification design across wearable sensors and a data aggregator with high energy-efficiency. To facilitate the practical use of XPro, we also develop an Automatic XPro Generator that formally generates XPro instances according to specific design constraints. As a proof of concept, we study the design and implementation of XPro with six different health applications. Evaluation results show that, compared with state-of-the-art methods, XPro can increase the battery life of the sensor node by 1.6-2.4X while at the same time reducing system delay by 15.6-60.8% for wearable computing systems. Aosen Wang, Lizhong Chen, Wenyao Xu |
ISCA | 2 |
| 2017 | CALM: Contention-Aware Latency-Minimal Application Mapping for Flattened Butterfly On-Chip NetworksabstractWith the emergence of many-core multiprocessor system-on-chips (MPSoCs), on-chip networks are facing serious challenges in providing fast communication among various tasks and cores. One promising on-chip network design approach shown in recent studies is to add express channels to traditional mesh network as shortcuts to bypass intermediate routers, thereby reducing packet latency. This approach not only changes the packet latency models, but also greatly affects network traffic behaviors, both of which have not been fully exploited in existing mapping algorithms. In this article, we explore the opportunities in optimizing application mapping for flattened butterfly, a popular express channel-based on-chip network. Specifically, we identify the unique characteristics of flattened butterfly, analyze the opportunities and new challenges, and propose an efficient heuristic mapping algorithm. The proposed algorithm Contention-Aware Latency Minimal (CALM) is able to reduce unnecessary turns that would otherwise impose additional router pipeline latency to packets, as well as adjust forwarding traffic to reduce network contention latency. Simulation results show that the proposed algorithm can achieve, on average, 3.4X reduction in the number of turns, 24.8% reduction in contention latency, and 14.12% reduction in the overall packet latency. Di Zhu 0002, Siyu Yue, Massoud Pedram, Lizhong Chen |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2016 | Simulation of NoC power-gating: Requirements, optimizations, and the Agate simulator
Lizhong Chen, Di Zhu 0002, Massoud Pedram, Timothy M. Pinkston |
J. Parallel Distributed Comput. | 1 |
| 2016 | A Filtering Mechanism to Reduce Network Bandwidth Utilization of Transaction ExecutionabstractHardware Transactional Memory (HTM) relies heavily on the on-chip network for intertransaction communication. However, the network bandwidth utilization of transactions has been largely neglected in HTM designs. In this work, we propose a cost model to analyze network bandwidth in transaction execution. The cost model identifies a set of key factors that can be optimized through system design to reduce the communication cost of HTM. Based on the model and network traffic characterization of a representative HTM design, we identify a huge source of superfluous traffic due to failed requests in transaction conflicts. As observed in a spectrum of workloads, 39% of the transactional requests fail due to conflicts, which renders 58% of the transactional network traffic futile. To combat this pathology, a novel in-network filtering mechanism is proposed. The on-chip router is augmented to predict conflicts among transactions and proactively filter out those requests that have a high probability to fail. Experimental results show the proposed mechanism reduces total network traffic by 24% on average for a set of high-contention TM applications, thereby reducing energy consumption by an average of 24%. Meanwhile, the contention in the coherence directory is reduced by 68%, on average. These improvements are achieved with only 5% area added to a conventional on-chip router design. Lihang Zhao, Lizhong Chen, Woojin Choi, Jeffrey T. Draper |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | Providing Balanced Mapping for Multiple Applications in Many-Core Chip MultiprocessorsabstractThis paper addresses the problem of balancing the on-chip packet latencies in a chip multi-processor (CMP), which is simultaneously executing multiple applications. Specifically, this paper presents a balanced application-to-core mapping algorithm that aims to minimize the maximum on-chip packet latency of all running applications. The paper starts by formulating the balanced mapping problem for CMPs and proving its NP-completeness. Next it presents an efficient heuristic algorithm for solving the aforesaid problem, which utilizes the characteristics of on-chip cache and memory accesses in CMPs and takes into account the workload variations among applications. Simulation results on PARSEC benchmark suite show that the proposed algorithm lowers the maximum average packet latency of all applications by 11 percent while cutting the standard deviation of on-chip packet latencies by 99 percent. This is achieved by very little overhead in terms of the overall packet latency and power consumption averaged over all packets. Di Zhu 0002, Lizhong Chen, Siyu Yue, Timothy M. Pinkston, Massoud Pedram |
IEEE Trans. Computers | 2 |
| 2015 | TAPP: temperature-aware application mapping for NoC-based many-core processors
Di Zhu 0002, Lizhong Chen, Timothy M. Pinkston, Massoud Pedram |
DATE | 2 |
| 2015 | Power punch: Towards non-blocking power-gating of NoC routersabstractAs chip designs penetrate further into the dark silicon era, innovative techniques are much needed to power off idle or under-utilized system components while having minimal impact on performance. On-chip network routers are potentially good targets for power-gating, but packets in the network can be significantly delayed as their paths may be blocked by powered-off routers. In this paper, we propose Power Punch, a novel performance-aware, power reduction scheme that aims to achieve non-blocking power-gating of on-chip network routers. Two mechanisms are proposed that not only allow power control signals to utilize existing slack at source nodes to wake up powered-off routers along the first few hops before packets are injected, but also allow these signals to utilize hop count slack by staying ahead of packets to "punch through " any blocked routers along the imminent path of packets, preventing packets from having to suffer router wakeup latency or packet detour latency. Full system evaluation on PARSEC benchmarks shows Power Punch saves more than 83% of router static energy while having an execution time penalty of less than 0.4%, effectively achieving near non-blocking power-gating of on-chip network routers. Lizhong Chen, Di Zhu 0002, Massoud Pedram, Timothy M. Pinkston |
HPCA | 1 |
| 2014 | Application mapping for express channel-based networks-on-chipabstractWith the emergence of many-core multiprocessor system-on-chips (MPSoCs), the on-chip networks are facing serious challenges in providing fast communication for various tasks and cores. One promising solution shown in recent studies is to add express channels to the network as shortcuts to bypass intermediate routers, thereby reducing packet latency. However, this approach also greatly changes the packet delay estimation and traffic behaviors of the network, both of which have not yet been exploited in existing mapping algorithms. In this paper, we explore the opportunities in optimizing application mapping for express channel-based on-chip networks. Specifically, we derive a new delay model for this type of networks, identify their unique characteristics, and propose an efficient heuristic mapping algorithm that increases the bypassing opportunities by reducing unnecessary turns that would otherwise impose the entire router pipeline delay to packets. Simulation results show that the proposed algorithm can achieve a 2∼4X reduction in the number of turns and 10∼26% reduction in the average packet delay. Di Zhu 0002, Lizhong Chen, Siyu Yue, Massoud Pedram |
DATE | 2 |
| 2014 | MP3: Minimizing performance penalty for power-gating of Clos network-on-chipabstractPower-gating is a promising technique to mitigate the increasing static power of on-chip routers. Clos networks are potentially good targets for power-gating because of their path diversity and decoupling between processing elements and most of the routers. While power-gated Clos networks can perform better than power-gated direct networks such as meshes, a significant performance penalty exists when conventional power-gating techniques are used. In this paper, we propose an effective power-gating scheme, called MP3 (Minimal Performance Penalty Power-gating), which is able to achieve minimal (i.e., near-zero) performance penalty and save more static energy than conventional power-gating applied to Clos networks. MP3 is able to completely remove the wakeup latency from the critical path, reduce long-term and transient contention, and actively steer network traffic to create increased power-gating opportunities. Full system evaluation using PARSEC benchmarks shows that the proposed approach can significantly reduce the performance penalty to less than 1% (as opposed to 38% with conventional power-gating) while saving more than 47% of router static energy, with only 2.5% additional area overhead. Lizhong Chen, Lihang Zhao, Timothy M. Pinkston |
HPCA | 1 |
| 2014 | Mitigating the Mismatch between the Coherence Protocol and Conflict Detection in Hardware Transactional MemoryabstractHardware Transactional Memory (HTM) usually piggybacks onto the cache coherence protocol to detect data access conflicts between transactions. We identify an intrinsic mismatch between the typical coherence scheme and transaction execution, which causes a sizable amount of unnecessary transaction aborts. This pathological behavior is called false aborting and increases the amount of wasted computation and on-chip communication. For the TM applications we studied, 41% of the transactional write requests incur false aborting. To combat false aborting, we propose Predictive Unicast and Notification (PUNO), a novel hardware mechanism to 1) replace the inefficient coherence multicast with a unicast scheme to prevent transactions from being disrupted unnecessarily and 2) restrain transaction polling through proactive notification. PUNO reduces transaction aborts by 61% and network traffic by 32% in workloads representative of future TM applications with a VLSI implementation area overhead of 0.41%. Lihang Zhao, Lizhong Chen, Jeffrey T. Draper |
IPDPS | 2 |
| 2014 | Balancing On-Chip Network Latency in Multi-application Mapping for Chip-MultiprocessorsabstractAs the number of cores continues to grow in chip multiprocessors (CMPs), application-to-core mapping algorithms that leverage the non-uniform on-chip resource access time have been receiving increasing attention. However, existing mapping methods for reducing overall packet latency cannot meet the requirement of balanced on-chip latency when multiple applications are present. In this paper, we address the looming issue of balancing minimized on-chip packet latency with performance-awareness in the multi-application mapping of CMPs. Specifically, the proposed mapping problem is formulated, its NP-completeness is proven, and an efficient heuristic-based algorithm for solving the problem is presented. Simulation results show that the proposed algorithm is able to reduce the maximum average packet latency by 10.42% and the standard deviation of packet latency by 99.65% among concurrently running applications and, at the same time, incur little degradation in the overall performance. Di Zhu 0002, Lizhong Chen, Siyu Yue, Timothy M. Pinkston, Massoud Pedram |
IPDPS | 2 |
| 2014 | Smart butterfly: reducing static power dissipation of network-on-chip with core-state-awarenessabstractWhile power gating is a promising technique to reduce the static power consumption of network-on-chip (NoC), its effectiveness is often hindered by the requirement of maintaining network connec-tivity and the limited knowledge of traffic behaviors. In this paper, we present Smart Butterfly, a core-state-aware NoC power-gating scheme based on flattened butterfly that utilizes the active/sleep state information of processing cores to improve power-gating effectiveness. Smart Butterfly exploits the rich connectivity of the flattened butterfly topology to allow more on-chip routers to be power-gated when their attached cores are asleep. We present two heuristic algorithms to determine the set of routers to be turned on to maintain connectivity and allow tradeoff between power consumption and average packet latency. Simulation results show an average of 42.85% and 60.48% power reduction of Smart Butterfly over prior art on 4x4 and 8x8 networks, respectively. Siyu Yue, Lizhong Chen, Di Zhu 0002, Timothy M. Pinkston, Massoud Pedram |
ISLPED | 2 |
| 2014 | Futility Scaling: High-Associativity Cache PartitioningabstractAs shared last level caches are widely used in many-core CMPs to boost system performance, partitioning a large shared cache among multiple concurrently running applications becomes increasingly important in order to reduce destructive interference. However, while recent works start to show the promise of using replacement-based partitioning schemes, such existing schemes either suffer from the severe associativity degradation when the number of partitions is high, or lack the ability to precisely partition the whole cache which leads to decreased resource efficiency. In this paper, we propose Futility Scaling (FS), a novel replacement-based cache partitioning scheme that can precisely partition the whole cache while still maintaining high associativity even with a large number of partitions. The futility of a cache line represents the uselessness of this line to application performance and can be ranked in different ways by various policies, e.g., LRU and LFU. The idea of FS is to control the size of a partition by properly scaling the futility of its cache lines. We study the properties of FS on both associativity and sizing in an analytical framework, and present a feedback-based implementation of FS that incurs little overhead in practice. Simulation results show that, FS improves performance over previously proposed Vantage and Prism by up to 6.0% and 13.7%, respectively. Lizhong Chen |
MICRO | 2 |
| 2013 | Worm-Bubble Flow ControlabstractDeadlock-free flow control should be designed with minimal cost, particularly for on-chip designs where area and power resources are greatly constrained. While Bubble Flow Control, proposed a decade ago, can avoid deadlock in VCT-switched tori with only one virtual channel (VC), there has been no working solution for wormhole switching that achieves the similar objective. Wormhole switching allows the channel buffer size to be smaller than the packet size, thus is preferred by on-chip networks. However, wormhole packets can span multiple routers, thereby creating additional channel dependences and adding complexities in both deadlock and starvation avoidance. In this paper, we propose Worm-Bubble Flow Control (WBFC), a new flow control scheme that can avoid deadlock in wormhole-switched tori using minimally 1-flit-sized buffers per VC and one VC in total. Moreover, any wormhole-switched topology with embedded rings can use WBFC to avoid deadlock within each ring. Simulation results from synthetic traffic and PARSEC benchmarks show that the proposed approach can achieve significant throughput improvement and also area and energy savings compared to an optimized Dateline routing approach. Lizhong Chen, Timothy M. Pinkston |
HPCA | 1 |
| 2013 | In-network traffic regulation for Transactional MemoryabstractHardware Transactional Memory (HTM) promises to simplify parallel programming on shared-memory chip multiprocessors by providing atomic execution of code blocks. Concurrently, Networks-On-Chip (NOCs) have emerged as an efficient on-chip communication infrastructure but have been largely neglected in HTM designs. In this work, we explore the interaction between the HTM paradigm and NOCs. In the process, we find a huge source of unnecessary network traffic incurred by transactional requests that are unsuccessful. This problem is identified as false forwarding that adversely affects network performance and energy efficiency. Surprisingly, 39% (up to 79% for a specific workload) of the transactional requests have incurred false forwarding over a wide spectrum of workloads. To combat this problem, we propose TMNOC, a novel approach that exploits the co-design of HTM and NOCs to mitigate false forwarding. Transactional requests that have a high probability to fail are filtered out in-network as early as possible to save energy and improve concurrency in the memory system. Experimental results show that our design reduces total network traffic by 20% on average (up to 40%) for a set of high-contention benchmarks representative of future TM workloads, thereby reducing energy consumption by an average of 24% (up to 39%). Meanwhile, the contention in the coherence directory is reduced by 66% on average. These improvements are achieved with only 5% area overhead added to a conventional on-chip router design. Lihang Zhao, Woojin Choi, Lizhong Chen, Jeffrey T. Draper |
HPCA | 3 |
| 2013 | Bubble coloring: avoiding routing- and protocol-induced deadlocks with minimal virtual channel requirementabstractHandling routing- and protocol-induced deadlocks is a critical issue in designing a reliable communication system. Generally, to avoid these two types of deadlocks without losing routing freedom requires a large amount of virtual channels (VCs), which imposes significant negative effects on router power, energy and frequency. In this paper, we propose a virtual cut-through switched Bubble Coloring (BC) scheme, which can avoid both routing- and protocol-induced deadlocks and allow fully adaptive routing on any topology without the need for multiple virtual channels. Results from both synthetic and full-system simulation show that, compared to a conventional deadlock-free scheme with 4VCs (i.e., XY_adaptive_4VC), our BC scheme with the minimal 1VC (i.e., BC_1VC) can reduce router energy and area by up to 51.2% and 58.3%, respectively, and has comparable performance at the same time. As the proposed BC scheme does not require multiple virtual channels, it also reduces the complexity of router arbitration logic, which brings the opportunity to increase router frequency and further improve system performance. Lizhong Chen, Timothy M. Pinkston |
ICS | 2 |
| 2013 | RAIR: Interference Reduction in Regionalized Networks-on-ChipabstractWith the advent of many-core systems capable of hosting multiple concurrently running applications, the traffic characteristics of networks-on-chip (NoCs) may exhibit new regional behaviors. By recognizing and exploiting these traffic behaviors, the effectiveness of NoC interference reduction techniques can be greatly improved. However, few works have investigated these regional behaviors and their potential impact on interference, leaving the opportunity largely unexplored. In this paper, we identify and characterize regional behavior in NoC and propose RAIR, a region-aware interference reduction technique that not only removes any restrictions on the inter-region traffic patterns, but also captures and exploits regional behavior throughout the design, thus improving the effectiveness of interference reduction. Evaluation using a cycle-accurate simulator shows that RAIR can improve the average packet latency by up to 17% on synthetic traffic patterns and up to 26% on PARSEC benchmarks compared to state-of-the-art interference reduction techniques. Lizhong Chen, Kai Hwang 0001, Timothy M. Pinkston |
IPDPS | 1 |
| 2013 | An Analytical Performance Model for Partitioning Off-Chip Memory BandwidthabstractWith the emergence of multi-programmed workloads for Chip Multiprocessors (CMP), Quality of Service (QoS) of each co-scheduled application on the CMP is increasingly gaining importance. As more and more applications are consolidated into a single chip to compete for the limited off-chip memory bandwidth, off-chip memory bandwidth partitioning makes an increasing impact on system performance. Although various existing heuristic-based memory scheduling schemes have achieved significant system performance improvement by better partitioning the bandwidth, it is still not clear what are the best ways to partition off-chip bandwidth for improving different system performance objectives. The goal of this paper is to understand how off-chip memory bandwidth partitioning affects various system performance objectives. To achieve this goal, we propose an analytical model that is simple yet powerful enough to reveal the relationship between various memory bandwidth partitioning schemes and different system performance objectives. From our model, optimal memory bandwidth partitioning schemes for different system-level objectives are derived. Experimental results from a cycle-accurate full-system simulator show that, for heterogeneous workloads, performance improvements over No_partitioning/Equal_partitioning in terms of harmonic weighted speedup, minimum fairness, weighted speedup and sum of IPCs are 20.3%/2.1%, 49.8%/38.7%, 32.8%/7.6% and 64.2%/24%, on average, with our corresponding optimal partitioning schemes (i.e., Square_root, Proportional, Priority_APC, Priority_API), respectively. Lizhong Chen, Timothy M. Pinkston |
IPDPS | 2 |
| 2012 | NoRD: Node-Router Decoupling for Effective Power-gating of On-Chip RoutersabstractWhile power-gating is a promising technique to mitigate the increasing static power of a chip, a fundamental requirement is for the idle periods to be sufficiently long to compensate for the power-gating and performance overhead. On-chip routers are potentially good targets for power optimizations, but few works have explored effective ways of power-gating them due to the intrinsic dependence between the node and router -- any packet (sent, received or forwarded) must wakeup the router before being transferred, thus breaking the potentially long idle period into fragmented intervals. Simulation shows that directly applying conventional power-gating techniques would cause frequent state-transitions and significant energy and performance overhead. In this paper, we propose NoRD (Node-Router Decoupling), a novel power-aware on-chip network approach that provides for power-gating bypass to decouple the node's ability for transferring packets from the powered-on/off status of the associated router, thereby maximizing the length of router idle periods. Full system evaluation using PARSEC benchmarks shows that the proposed approach can substantially reduce the number of state-transitions, completely hide wakeup latency from the critical path of packet transport and eliminate node-network disconnection problems. Compared to an optimized conventional power-gating technique applied to on-chip routers, NoRD can further reduce the router static energy by 29.9% and improve the average packet latency by 26.3%, with only 3% additional area overhead. Lizhong Chen, Timothy M. Pinkston |
MICRO | 1 |
| 2012 | Efficient implementation of globally-aware network flow control
Lizhong Chen, Timothy M. Pinkston |
J. Parallel Distributed Comput. | 1 |
| 2011 | Critical Bubble Scheme: An Efficient Implementation of Globally Aware Network Flow ControlabstractNetwork flow control mechanisms that are aware of global conditions potentially can achieve higher performance than flow control mechanisms that are only locally aware. Owing to high implementation overhead, globally-aware flow control mechanisms in their purest form are seldom adopted in practice, leading to less efficient simplified implementations. In this paper, we propose an efficient implementation of a globally-aware flow control mechanism, called Critical Bubble Scheme, and apply it successfully to k-ary n-cube networks for the general class of buffer occupancy-based network flow control techniques. Simulation results show that the proposed scheme can reduce the buffer access portion of packet latency by as much as 77%, leading to much lower average packet latency at medium and high network loads while sustaining 11% throughput improvement after network saturation. Lizhong Chen, Timothy M. Pinkston |
IPDPS | 1 |