Jiayi Huang 0001

dblp:180/4976-1 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0003-4011-6668ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 5 first-author · 21 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 XTree on EquiMesh: Topology and Algorithm Co-Design for Collective Communication
abstract
Mesh topology is widely adopted for both on-chip and chiplet-based interconnects due to its placement-friendly physical layout. However, the low-degree nodes at the edges and corners create bandwidth bottlenecks for common collectives such as AllGather and AllReduce. We address this limitation with EquiMesh, an augmented 2D-Mesh with equivalent-degree nodes without incurring switching complexity. To fully exploit EquiMesh, we propose XTree, a topology-aware algorithm that maximizes utilization of available bandwidth, and MirrorXTree, which constructs ReduceScatter and AllReduce on top of XTree through topology mirroring. Our evaluation shows that EquiMesh with XTree and MirrorXTree achieves 2× and 1.2× higher effective bandwidth than state-of-the-art mesh-based topology-algorithm co-designs for AllGather and AllReduce, respectively.
Junwei Cui, Weilin Cai, Jiayi Huang 0001
DATE4
2026 Mapping and Communication Optimizations with Fault Tolerance for Wafer-Scale LLM Inference
Junwei Cui, Weilin Cai, Jiayi Huang 0001
ISCA4
2025 MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
abstract
As large language models continue to scale up, distributed training systems have expanded beyond 10k nodes, intensifying the importance of fault tolerance. Checkpoint has emerged as the predominant fault tolerance strategy, with extensive studies dedicated to optimizing its efficiency. However, the advent of the sparse Mixture-of-Experts (MoE) model presents new challenges due to the substantial increase in model size, despite comparable computational demands to dense models.
Weilin Cai, Jiayi Huang 0001
ASPLOS (2)3
2025 Push Multicast: A Speculative and Coherent Interconnect for Mitigating Manycore CPU Communication Bottleneck
abstract
As CPUs scale up to many cores, the bandwidth of the network-on-chip (NoC) and cache can soon become the performance bottleneck. In modern processors, the cache hierarchy plays a reactive role to supply data upon request. In parallel programs, shared data accesses from different cores at different times can consume large cache and NoC bandwidth for the same data. These same-data accesses inherently have redundancy and lead to inefficient cache and NoC bandwidth utilization. In this work, we propose Push Multicast, a speculative and coherent interconnect. We transform the last-level cache into a proactive agent to push data to other sharers upon replying to the demand requester. Pushing enables effective multicasting to reduce LLC and NoC bandwidth consumption. A coherent innetwork filter is proposed to prune the outstanding requests in the routers along the way of the pushed data delivery. Moreover, a dynamic mechanism is designed to pause and resume pushing adaptively. Compared with a system with an L1 Bingo data prefetcher and an L2 Stride prefetcher, Push Multicast achieves an average of $\mathbf{3 3 \%}$ NoC bandwidth saving, a geomean of $1.02 \times$ and a maximum of $1.56 \times$ speedup in a 16 -core system. In a 64-core system, it further achieves an average of $\mathbf{4 3 \%}$ NoC bandwidth saving, along with a geomean of $1.11 \times$ and a maximum of $2.08 \times$ speedup.
Jiayi Huang 0001, Zhe Wang 0023, Christopher J. Hughes, Yufei Ding 0001, Yuan Xie 0001
HPCA1
2025 Shortcut-connected Expert Parallelism for Accelerating Mixture of Experts
abstract
Expert parallelism has emerged as a key strategy for distributing the computational workload of sparsely-gated mixture-of-experts (MoE) models across multiple devices, enabling the processing of increasingly large-scale models. However, the All-to-All communication inherent to expert parallelism poses a significant bottleneck, limiting the efficiency of MoE models. Although existing optimization methods partially mitigate this issue, they remain constrained by the sequential dependency between communication and computation operations. To address this challenge, we propose ScMoE, a novel shortcut-connected MoE architecture integrated with an overlapping parallelization strategy. ScMoE decouples communication from its conventional sequential ordering, enabling up to 100\% overlap with computation. Compared to the prevalent top-2 MoE baseline, ScMoE achieves speedups of $1.49\times$ in training and $1.82\times$ in inference. Moreover, our experiments and analyses indicate that ScMoE not only achieves comparable but in some instances surpasses the model quality of existing approaches.
Weilin Cai, Juyong Jiang, Junwei Cui, Sunghun Kim 0001, Jiayi Huang 0001
ICML6
2025 TRACI: Network Acceleration of Input-Dynamic Communication for Large-Scale Deep Learning Recommendation Model
abstract
Large-scale deep learning recommendation models (DLRMs) rely on embedding layers with terabyte-scale embedding tables, which present significant challenges to memory capacity.In addition, these embedding layers exhibit sparse and random data access patterns, which demand high memory bandwidth.Multi-GPU systems provide a promising solution, allowing for the scaling of both memory and aggregated bandwidth.However, network communication bandwidth becomes a bottleneck for multi-GPU DLRM systems.Overcoming the communication bottleneck is crucial to unlocking the potential of multi-GPU systems for efficient and high-performance DLRM training.This paper introduces TRACI, an in-network acceleration architecture designed to optimize the communication operator in embedding layers: Aggregation.While in-network acceleration has proven successful for the All-Reduce communication collective, existing solutions do not directly apply to Aggregation due to two key challenges.Firstly, existing multi-GPU shared memory operations are designed for point-to-point communication and do not allow the network to proactively optimize communication.Secondly, in Aggregation, data transfer patterns are dynamic and dependent on input, demanding the network to dynamically discover and exploit message connections on-the-fly.To address these challenges, we propose a solution that involves a novel network transaction and switch hardware design.We introduce a new network transaction that augments messages with input reuse and output reuse identifications, and can empower the network to proactively reduce
Guyue Huang, Hao Li 0120, Jiayi Huang 0001, Yangwook Kang, Yufei Ding 0001, Yuan Xie 0001
ISCA4
2025 Chimera: Communication Fusion for Hybrid Parallelism in Large Language Models
abstract
Large Language Models (LLMs), exemplified by ChatGPT, have emerged as a predominant workload in current machine learning systems.To achieve efficient training and inference within the constraints of limited single-NPU memory capacity, deploying LLMs on multi-NPU systems typically adopt a hybrid approach that combines various parallelism patterns.This hybrid parallelism within LLMs introduces a significant amount of diverse collective communications.However, these frequent blocking communications impose a substantial burden on the multi-NPU systems.Overcoming the communication bottleneck is crucial to unlocking the potential of multi-NPU systems for efficient and scalable LLM processing.This paper introduces Chimera, a communication fusion mechanism for hybrid parallelism in LLMs.We comprehensively analyze the communication processes of each LLM parallelism pattern, identify the communication redundancy in hybrid parallelism and eliminate redundancy by fusing adjacent communication operators during parallelism transformation.By reordering operations and generating redundancy-free communication operator, Chimera effectively mitigates communication bottleneck in hybrid LLM parallelism.Our results show that Chimera achieves 1.23-7.06×network bandwidth speedup.Additionally, the end-to-end performance of LLM forward pass and backward pass on different typical multi-NPU systems achieves respective 1.32-1.58×and 1.16-1.36×speedups on average compared with those without communication fusion.
Junwei Cui, Weilin Cai, Jiayi Huang 0001
ISCA4
2025 Optimizing Heterogeneous Compute-in-Memory with Hybrid Dataflow and In-Network Reduction for Vision Transformer
abstract
Transformer models have shown remarkable performance in AI tasks. However, their large model sizes and large-scale matrix computations heavily press the memory bandwidth. Compute-in-memory (CIM) solutions have emerged to address these issues by integrating computation into memory. Nevertheless, existing scalable and heterogeneous CIM designs further suffer from inefficiencies caused by isolated mapping strategies and communication redundancy. To address these issues, we propose Cross-Type Mapped Dataflow, an architecture-dataflow co-design that enables efficient and collaborative analog-digital CIM execution.Our approach features a heterogeneous analog-digital CIM architecture with flat network-on-chip (NoC) interconnects, enabling seamless cross-type communication. The core innovations include: a Hybrid Dataflow that schedules idle CIMs for cross-type collaboration on vector-matrix multiplication, and an In-Network Reduction (INR) mechanism that eliminates redundant data transfers by embedding Reduction operations within NoC communication. Experimental results show an average 3.57× speedup and 2.78× energy efficiency improvement over Homogeneous DCIM architecture. Optimized by the proposed Hybrid Dataflow and INR, our heterogeneous CIMs further achieve an average 4.63× speedup and 2.59× energy efficiency improvement over state-of-the-art X-Former-like design across various ViT workloads.
Zexin Fu, Yihang Zuo, Yuzhe Ma, Jiayi Huang 0001
ISLPED4
2025 Software Prefetch Multicast: Sharer-Exposed Prefetching for Bandwidth Efficiency in Manycore Processors
abstract
As the core counts continue to scale in manycore processors, the increasing bandwidth pressure on the network-on-chip (NoC) and last-level cache (LLC) emerges as a critical performance bottleneck.While shared-data multicasting from the LLC can alleviate this pressure by combining responses heading to the same direction, existing solutions face fundamental challenges in accurately identifying complete sharer sets and avoiding unnecessary data responses.In this work, we propose Software Prefetch Multicast (SPM), a softwarehardware co-design that addresses these limitations through sharerexposed prefetching for bandwidth-efficient multicast.SPM introduces three key innovations: (1) New software-hardware interfaces including sharer group configuration and sharer-exposed prefetching instructions, enabling software-directed multicast initiation;(2) A corresponding microarchitecture for sharer group configuration that supports sharer-group-based multicast triggering in LLC through representative prefetch requests from leader threads; and(3) A dynamic leader thread switching algorithm to accommodate thread variation.Compared with a system with software prefetching, SPM achieves an average of 42% NoC bandwidth saving, a geometric mean of 1.28× and a maximum of 1.46× speedup in a 16-core system.In a 64-core system, it can achieve an average of 50% NoC bandwidth saving, along with a geomean of 1.38× and a maximum of 1.86× speedup.
Jiong Feng, Zhe Wang 0023, Christopher J. Hughes, Jiayi Huang 0001
MICRO5
2025 Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks
Junwei Cui, Weilin Cai, Meng Niu, Jiayi Huang 0001
MICRO6
2025 NoCFuzzer: Automating NoC Verification in UVM
abstract
Network on chip (NoC) has surfaced as a crucial interconnection strategy in modern digital systems, thereby demanding meticulous verification. Due to its multiple nodes and high concurrency, verifying an NoC is labor-intensive, making it complex to generate a multitude of test cases. Recently, hardware fuzzing has been identified as a promising automated approach for hardware verification. However, when we tried to apply these fuzzing techniques to our internally developed NoC design, we discovered that they were incompatible with the specificities of NoC. Additionally, they are also incompatible with the standard IC verification workflow and universal verification methodology (UVM) environment. In this work, we aim to automate our verification process of NoC with fuzzing. We propose a fuzzing strategy specifically tailored for industrial NoC UVM verification. We employ fuzzing in NoC verification at multiple levels, including router verification, network verification, and stress testing. As a case study we apply our approach to an open-source NoC component in OpenPiton. Remarkably, our fuzzing methods automatically achieved complete code and functional coverage in the router and mesh network. We also effectively detect injected starvation bugs with fuzzing. The evaluation results clearly demonstrate the practicability of our fuzzing approach to considerably reduce the manpower required for test case generation compared with traditional NoC verification.
Jiayi Huang 0001, Shijian Zhang, Yuan Xie 0001, Guojie Luo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 A Survey on Mixture of Experts in Large Language Models
abstract
Large language models (LLMs) have garnered unprecedented advancements across diverse fields, ranging from natural language processing to computer vision and beyond. The prowess of LLMs is underpinned by their substantial model size, extensive and diverse datasets, and the vast computational power harnessed during training, all of which contribute to the emergent abilities of LLMs (e.g., in-context learning) that are not present in small models. Within this context, the mixture of experts (MoE) has emerged as an effective method for substantially scaling up model capacity with minimal computation overhead, gaining significant attention from academia and industry. Despite its growing prevalence, there lacks a systematic and comprehensive review of the literature on MoE. This survey seeks to bridge that gap, serving as an essential resource for researchers delving into the intricacies of MoE. We first briefly introduce the structure of the MoE layer, followed by proposing a new taxonomy of MoE. Next, we overview the core designs for various MoE models including both algorithmic and systemic aspects, alongside collections of available open-source implementations, hyperparameter configurations and empirical evaluations. Furthermore, we delineate the multifaceted applications of MoE in practice, and outline some potential directions for future research.
Weilin Cai, Juyong Jiang, Fan Wang 0041, Jing Tang 0004, Sunghun Kim 0001, Jiayi Huang 0001
IEEE Trans. Knowl. Data Eng.6
2024 An Endeavor to Industrialize Hardware Fuzzing: Automating NoC Verification in UVM
abstract
We endeavor to make hardware fuzzing compatible with the standard IC development process and apply that to NoC verification in a real-world industrial environment. We systematically employ fuzzing throughout the entire NoC verification process, including router verification, network verification, and stress testing. As a case study, we apply our approach to an open-source NoC component in OpenPiton. Remarkably, our fuzzing methods automatically achieved complete code and functional coverage in the router and mesh network, and effectively detect injected starvation bugs. The evaluation results clearly demonstrate the practicability of our fuzzing approach to considerably reduce the manpower required for test case generation compared with traditional NoC verification.
Huatao Zhao, Jiayi Huang 0001, Shijian Zhang, Guojie Luo
DATE3
2024 Prefender: A Prefetching Defender Against Cache Side Channel Attacks as a Pretender
abstract
Cache side channel attacks are increasingly alarming in modern processors due to the recent emergence of Spectre and Meltdown attacks. A typical attack performs intentional cache access and manipulates cache states to leak secrets by observing the victim’s cache access patterns. Different countermeasures have been proposed to defend against both general and transient execution based attacks. Despite their effectiveness, they mostly trade some level of performance for security, or have restricted security scope. In this paper, we seek an approach to enforcing security while maintaining performance. We leverage the insight that attackers need to access cache in order to manipulate and observe cache state changes for information leakage. Specifically, we propose Prefender, a secure prefetcher that learns and predicts attack-related accesses for prefetching the cachelines to simultaneously help security and performance. Our results show that Prefenderis effective against several cache side channel attacks while maintaining or even improving performance for SPEC CPU 2006 and 2017 benchmarks.
Jiayi Huang 0001, Lang Feng 0001, Zhongfeng Wang 0001
IEEE Trans. Computers2
2023 ArchExplorer: Microarchitecture Exploration Via Bottleneck Analysis
abstract
Design space exploration (DSE) for microarchitecture parameters is an essential stage in microprocessor design to explore the trade-offs among performance, power, and area (PPA). Prior work either employs excessive expert efforts to guide microarchitecture parameter tuning or demands high computing resources to prepare datasets and train black-box prediction models for DSE.
Jiayi Huang 0001, Xuechao Wei, Yuzhe Ma, Sicheng Li 0001, Hongzhong Zheng, Bei Yu 0001, Yuan Xie 0001
MICRO2
2023 WHISTLE: CPU Abstractions for Hardware and Software Memory Safety Invariants
abstract
Memory safety invariants extracted from a program can help defend and detect against both software and hardware memory violations. For instance, by allowing only specific instructions to access certain memory locations, system can detect out-of-bound or illegal pointer dereferences that lead to correctness and security issues. In this paper, we propose CPU abstractions, called, to specify and check program invariants to provide defense mechanism against both software and hardware memory violations at runtime. ensures that the invariants must be satisfied at every memory accesses. We present a fast invariant address translation and retrieval scheme using a specialized cache. It stores and checks invariants related to global, stack and heap objects. The invariant checks can be performed synchronously or asynchronously. uses synchronous checking for high security-critical programs, while others are protected by asynchronous checking. A fast exception is proposed to alert any violations as soon as possible in order to close the gap for transient attacks. Our evaluation shows that can detect both software and hardware, spatial and temporal memory violations. incurs 53% overhead when checking synchronously, or 15% overhead when checking asynchronously.
Sungkeun Kim, Farabi Mahmud, Jiayi Huang 0001, Pritam Majumder, Chia-Che Tsai, Abdullah Muzahid, Eun Jung Kim 0001
IEEE Trans. Computers3
2022 PREFENDER: A Prefetching Defender against Cache Side Channel Attacks as A Pretender
abstract
Cache side channel attacks are increasingly alarming in modern processors due to the recent emergence of Spectre and Meltdown attacks. A typical attack performs intentional cache access and manipulates cache states to leak secrets by observing the victim's cache access patterns. Different countermeasures have been proposed to defend against both general and transient execution based attacks. Despite their effectiveness, they all trade some level of performance for security. In this paper, we seek an approach to enforcing security while maintaining performance. We leverage the insight that attackers need to access cache in order to manipulate and observe cache state changes for information leakage. Specifically, we propose PREFENDER,a secure prefetcher that learns and predicts attack-related accesses for prefetching the cachelines to simultaneously help security and performance. Our results show that PREFENDER is effective against several cache side channel attacks while maintaining or even improving performance for SPEC CPU2006 benchmarks.
Jiayi Huang 0001, Lang Feng 0001, Zhongfeng Wang 0001
DATE2
2022 RvDfi: A RISC-V Architecture With Security Enforcement by High Performance Complete Data-Flow Integrity
abstract
With the rapid revolution of open-source hardware, RISC-V architecture has been prevalent in both academic research and industrial developments. Due to the increasing threats of information leakage, it is imperative to provide a secure RISC-V ecosystem to defend against malicious software exploits. Toward this goal, data-flow integrity (DFI) is employed as a strict security policy for enforcing the legitimacy of each data access, thereby filtering out most of the attack exploits. However, due to the intensive computations needed by DFI, there are only limited proposals successfully implementing partial DFI with low performance overhead. Moreover, all the previous studies failed to enforce thecompleteDFI policy in a real hardware platform, while trading off security strength for performance efficiency. To provide RISC-V architecture with high security enforcement and low performance overhead, we leverage the open-source Rocket Chip and proposeRvDfi, the first complete DFI implementation based on RISC-V architecture with only 17.8% performance overhead on average and 3.9% in minimum, incurring much less performance loss compared to the 166.3% overhead caused by previous complete DFI implementation.
Lang Feng 0001, Jiayi Huang 0001, Zhongfeng Wang 0001
IEEE Trans. Computers2
2022 Toward Taming the Overhead Monster for Data-flow Integrity
abstract
Data-Flow Integrity (DFI) is a well-known approach to effectively detecting a wide range of software attacks. However, its real-world application has been quite limited so far because of the prohibitive performance overhead it incurs. Moreover, the overhead is enormously difficult to overcome without substantially lowering the DFI criterion. In this work, an analysis is performed to understand the main factors contributing to the overhead. Accordingly, a hardware-assisted parallel approach is proposed to tackle the overhead challenge. Simulations on SPEC CPU 2006 benchmark show that the proposed approach can completely enforce the DFI defined in the original seminal work while reducing performance overhead by 4×, on average.
Lang Feng 0001, Jiayi Huang 0001, Jeff Huang 0001, Jiang Hu 0001
ACM Trans. Design Autom. Electr. Syst.2
2021 Communication Algorithm-Architecture Co-Design for Distributed Deep Learning
abstract
Large-scale distributed deep learning training has enabled developments of more complex deep neural network models to learn from larger datasets for sophisticated tasks. In particular, distributed stochastic gradient descent intensively invokes all-reduce operations for gradient update, which dominates communication time during iterative training epochs. In this work, we identify the inefficiency in widely used all-reduce algorithms, and the opportunity of algorithm-architecture co-design. We propose MultiTree all-reduce algorithm with topology and resource utilization awareness for efficient and scalable all-reduce operations, which is applicable to different interconnect topologies. Moreover, we co-design the network interface to schedule and coordinate the all-reduce messages for contention-free communications, working in synergy with the algorithm. The flow control is also simplified to exploit the bulk data transfer of big gradient exchange. We evaluate the co-design using different all-reduce data sizes for synthetic study, demonstrating its effectiveness on various interconnection network topologies, in addition to state-of-the-art deep neural networks for real workload experiments. The results show that MultiTree achieves 2.3× and 1.56× communication speedup, as well as up to 81% and 30% training time reduction compared to ring all-reduce and state-of-the-art approaches, respectively.
Jiayi Huang 0001, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid, Ki Hwan Yum, Eun Jung Kim 0001
ISCA1
2021 A Voting Approach for Adaptive Network-on-Chip Power-Gating
abstract
Scalable Networks-on-Chip (NoCs) have become the standard interconnection mechanisms in large-scale multicore architectures. These NoCs consume a large fraction of the on-chip power budget, where the static portion is becoming dominant as technology scales down to sub-10nm node. Therefore, it is essential to reduce static power so as to achieve power- and energy-efficient computing. Power-Gating as an effective static power saving technique can be used to power off inactive routers for static power saving. However, packet deliveries in irregular power-gated networks suffer from detour or waiting time overhead to either route around or wake up power-gated routers. In this article, we proposeFly-Over (Flov), a voting approach for dynamic router power-gating in a light-weight and distributed manner, which includesFlovrouter microarchitecture, adaptive power-gating policy, and low-latency dynamic routing algorithms. We evaluateFlovusing synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. Our full-system evaluations show thatFlovreduces the power consumption of NoC by 31 and 20 percent, respectively, on average across several benchmarks, compared to the baseline and the state-of-the-art while maintaining the similar performance.
Jiayi Huang 0001, Shilpa Bhosekar, Rahul Boyapati, Byul Hur, Ki Hwan Yum, Eun Jung Kim 0001
IEEE Trans. Computers1
2021 Computing En-Route for Near-Data Processing
abstract
The data explosion and faster data analysis demand have spawned emerging applications that operate over myriads of data and exhibit large memory footprints with low data reuse rate. Such characteristics lead to enormous data movements across the memory hierarchy and pose significant pressure on modern communication fabrics and memory subsystems. To mitigate the worsening gap between high processor computation density and deficient memory bandwidth, memory networks, and near-data processing techniques are proposed to keep improving system performance and energy efficiency. In this article, we propose Active-Routing, an in-network near-data processing architecture for data-flow execution, which enables computation en-route by exploiting patterns of aggregation over intermediate results. The proposed architecture leverages the massive memory cube- and vault-level parallelism as well as network concurrency to optimize the aggregation operations along a dynamically built Active-Routing Tree. It also introduces page granular computation offloading to amortize the offloading overhead and improve the throughput. Compared to the state-of-the-art processing-in-memory architecture, the evaluations show that the baseline Active-Routing can achieve up to 7× speedup with an average of 60 percent performance improvement, and reduce the energy-delay product by 80 percent across various benchmarks. Further optimizations with vault-level parallelism and page granular offloading can achieve an extra order of magnitude improvement.
Jiayi Huang 0001, Pritam Majumder, Sungkeun Kim, Troy Fulton, Ramprakash Reddy Puli, Ki Hwan Yum, Eun Jung Kim 0001
IEEE Trans. Computers1
2021 Remote Control: A Simple Deadlock Avoidance Scheme for Modular Systems-on-Chip
abstract
Ever increasing performance demand and shrinking in the transistor size together result in complex and dense packing in large chips. That motivates designers to opt for many small specialized hardware modules in a chip to extract maximum performance benefits with relatively lower complexity and cost. These altogether opens up new directions for heterogeneous modular System-on-Chip (SoC) research, where a large system is built by assembling small independently designed chiplets (small chips). We focus on the communication aspect of such SoCs, especially the newly observed deadlock among chiplets. Even though deadlock is a classic problem in networks and many solutions are available, the modular SoC design demands customized solutions that preserves the design flexibility for chiplet designers. We proposeRemote Control (RC), a simple routing oblivious deadlock avoidance scheme based on selective injection-control mechanism. Along with guarantee on deadlock freedom,RCaims to provide a methodology to make each independently designed chiplet seamlessly integrate in any modular SoCs. We achieve up to 56.34% throughput and 15.49% zero load latency improvements on synthetic traffic and up to 20% speedup on real workloads taken from vast range of benchmark suites, over the state-of-the-art turn restriction based technique applied in the modular SoC domain.
Pritam Majumder, Sungkeun Kim, Jiayi Huang 0001, Ki Hwan Yum, Eun Jung Kim 0001
IEEE Trans. Computers3
2019 Active-Routing: Compute on the Way for Near-Data Processing
abstract
The explosion of data availability and the demand for faster data analysis have led to the emergence of applications exhibiting large memory footprint and low data reuse rate. These workloads, ranging from neural networks to graph processing, expose compute kernels that operate over myriads of data. Significant data movement requirements of these kernels impose heavy stress on modern memory subsystems and communication fabrics. To mitigate the worsening gap between high CPU computation density and deficient memory bandwidth, solutions like memory networks and near-data processing designs are being architected to improve system performance substantially. In this work, we examine the idea of mapping compute kernels to the memory network so as to leverage in-network computing in data-flow style, by means of near-data processing. We propose Active-Routing, an in-network compute architecture that enables computation on the way for near-data processing by exploiting patterns of aggregation over intermediate results of arithmetic operators. The proposed architecture leverages the massive memory-level parallelism and network concurrency to optimize the aggregation operations along a dynamically built Active-Routing Tree. Our evaluations show that Active-Routing can achieve upto 7X speedup with an average of 60% performance improvement, and reduce the energy-delay product by 80% across various benchmarks compared to the state-of-the-art processing-in-memory architecture.
Jiayi Huang 0001, Ramprakash Reddy Puli, Pritam Majumder, Sungkeun Kim, Rahul Boyapati, Ki Hwan Yum, Eun Jung Kim 0001
HPCA1
2017 Packet coalescing exploiting data redundancy in GPGPU architectures
abstract
General Purpose Graphics Processing Units (GPGPUs) are becoming a cost-effective hardware approach for parallel computing. Many executions on the GPGPUs place heavy stress on the memory system, creating network bottlenecks near memory controllers. We observe that data redundancy in communication traffic is common-place across a wide range of GPGPU applications. To exploit the data redundancy, we propose a packet coalescing mechanism to alleviate the network bottlenecks by directly reducing the traffic volume. The key idea is to coalesce multiple packets into one without increasing the packet size when they carry redundant cache blocks. To ensure that the coalesced packets are delivered to their respective destinations, we adopt multicast routing for the interconnection network of GPGPUs. Our coalescing approach yields 15% IPC improvement (up to 112%) in a large-scale GPGPU with 2D mesh across various GPGPU applications, by reducing average memory access time (AMAT) by 15.5% (up to 65.2%) and obtaining network bandwidth savings by 13% (up to 37%). Also, our coalescing approach achieves 7% IPC improvement in the NVIDIA Fermi architecture with the crossbar.
Kyung Hoon Kim, Rahul Boyapati, Jiayi Huang 0001, Yuho Jin, Ki Hwan Yum, Eun Jung Kim 0001
ICS3
2017 Fly-Over: A Light-Weight Distributed Power-Gating Mechanism for Energy-Efficient Networks-on-Chip
abstract
Scalable Networks-on-Chip (NoCs) have become the de facto interconnection mechanism in large scale Chip Multiprocessors. Not only are NoCs devouring a large fraction of the on-chip power budget but static NoC power consumption is becoming the dominant component as technology scales down. Hence reducing static NoC power consumption is critical for energy-efficient computing. Previous research has proposed to power-gate routers attached to inactive cores so as to save static power, but requires centralized control and global network knowledge. In this paper, we propose Fly-Over (FLOV), a light-weight distributed mechanism for power-gating routers, which encompasses FLOV router architecture, handshake protocols, and a partition-based dynamic routing algorithm to maintain network functionalities. With simple modifications to the baseline router architecture, FLOV can facilitate FLOV links over power-gated routers. Then we present two handshake protocols for FLOV routers, restricted FLOV that can power-gate routers under restricted conditions and generalized FLOV with more power saving capability. The proposed routing algorithm provides best-effort minimal path routing without the necessity for global network information. We evaluate our schemes using synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. Our full system evaluations show that FLOV reduces the total and static energy consumption by 18% and 22% respectively, on average across several benchmarks, compared to state-of-the-art NoC power-gating mechanism while keeping the performance degradation minimal.
Rahul Boyapati, Jiayi Huang 0001, Kyung Hoon Kim, Ki Hwan Yum, Eun Jung Kim 0001
IPDPS2
2017 APPROX-NoC: A Data Approximation Framework for Network-On-Chip Architectures
Rahul Boyapati, Jiayi Huang 0001, Pritam Majumder, Ki Hwan Yum, Eun Jung Kim 0001
ISCA2
2016 POSTER: Fly-Over: A Light-Weight Distributed Power-Gating Mechanism For Energy-Efficient Networks-on-Chip
abstract
Reducing static NoC power consumption is becoming critical for energy-efficient computing as technology scales down since NoCs are devouring a large fraction of the on-chip power budget. We propose Fly-Over (FLOV), a light-weight distributed mechanism for power-gating routers. With simple modifications to the baseline router architecture, FLOV links are facilitated over power-gated routers. A Handshake protocol that allows seamless router power-gating in addition to a dynamic routing algorithm, that provides best-effort minimal path without the necessity for global network information, maintain normal NoC functionality. We evaluate our schemes using synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. The results show that FLOV can achieve on average 19.2% latency reduction and 15.9% total energy savings.
Rahul Boyapati, Jiayi Huang 0001, Kyung Hoon Kim, Ki Hwan Yum, Eun Jung Kim 0001
PACT2