VLDB 2026 Research / reviewers in the wild / expert
Avinash Karanth
dblp:59/6384 · also Avinash Karanth Kodi, Avinash Kodi
· DBLP profile ↗
79ranked-venue papers
13as first author
26since 2021 · last 2026
0000-0002-9472-4637ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 75 · 11 first-author · 24 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MONET: A Mixture-of-Experts Accelerator with a Multicast-Optimized Two-Tier Network-on-ChipabstractThe growing complexity of Mixture-of-Experts (MoE) models in machine learning applications demands innovative hardware solutions to address their unique computational and data movement challenges. Some of the critical challenges facing MoE models include sparse activation, dynamic token routing and irregular computation patterns that lead to low utilization and higher communication latency. In this paper, we introduce MONET, a novel two-tier Network-on-Chip (NoC) architecture designed to efficiently execute MoE workloads by co-optimizing compute, memory, and interconnect subsystems. The first tier consists of a reconfigurable systolic processing element (PE) island, executing both gating and expert computations, with runtime-configurable support for sparse/dense operations, expert reordering, and activation functions. The second tier incorporates a dual mesh network connecting a grid of PE islands; one network manages input token delivery with a broadcast scheme optimized for the gating phase of MoE, while the other is tailored for efficient inter-expert communication necessary for result aggregation. Evaluated on MoE benchmarks, MONET demonstrates up to 8.5× lower latency and over 6× better energy efficiency compared to state-of-the-art MoE accelerators. Siqin Liu, Maya Roediger, Avinash Karanth |
DATE | 3 |
| 2026 | Process Identification and PUF Design Using CMOS Inverter Nonlinearity: Demonstration via 45nm & 22nm CMOSabstractWe propose harnessing the intrinsic non-linearity of a CMOS inverter to establish a strong physically unclonable function (PUF). In contrast to conventional PUFs, which typically depend on randomness of initialized SRAM memory or delay/noise in an oscillator, this study makes use of higher-order harmonic distortion in an inverter stage as physics-governed security primitive. More specifically, we utilize an efficient Integral Function Method (IFM) to extract 2nd-order (HD2), 3rd-order (HD3), and total harmonic distortion (THD) to establish new primitives for process security, and tools for a strong PUF. We illustrate the capabilities of the proposed approach via a study of inverter performance as described by a commercial PDK in 22 nm technology, and two different PDKs in the 45 nm technology. We illustrate how IFM can unravel interplay between the inherent nonlinearities and bias conditions to identify a given process technology or device family due to its relative simplicity and sensitivity to transistor operation modes. Such features and bias dependencies are then used to establish a strong PUF with a large number of challenge-response pairs as it can employ three digital-to-analog converters (DAC) to set quasistatic sweep parameters and linearity response thresholds. Therefore, we argue that a CMOS inverter’s linearity/distortion performance can become a very valuable marker to enhance root-of-trust in hardware security that can also serve as a tool to monitor age/reliability of integrated circuits. Aakriti Barat, Savas Kaya, Avinash Karanth, Sunaim Abdullah, Soumyasanta Laha |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Language semantics to support secure computation and communication in embedded systems via hardware monitorsabstractAs embedded systems with manycores and Network-on-Chips (NoCs) become ubiquitous, emerging hardware and software vulnerabilities have made it challenging to ensure system integrity especially when third-party intellectual property (IP) is used for rapid prototyping. Prior works have evaluated hardware monitors for ensuring correctness of the system by threat assessment and effective mitigation. However, none have evaluated models that combine both computation (processor pipeline) and communication (NoC) vulnerabilities simultaneously. In this paper, we propose a high-level policy language called d-GUARD that is used to define runtime security policies that can be compiled into hardware monitors. The advantage of this new language is the ability to dynamically change policies based on program’s runtime behavior. To translate high-level policies into low-level hardware monitors, we describe a compiler for d-GUARD that synthesizes policies into Verilog modules. Instead of simply evaluating the design of secure policies for processor pipelines, we extend to secure NoC microarchitectures, including policies for links and routers, as well as policies to prevent Denial-of-Service (DoS) attacks. To mitigate attacks against secure microarchitectures, we also propose fault-tolerant routing approaches to avoid rogue routers when the number of policy violations exceeds a certain threshold. Our secure policies for processor pipelines and NoC microarchitectures consume marginal area and power overhead when compared to baseline making it well suited for low-cost embedded systems. • A high-level security policy language called D-GUARD, which is expressive, modular and compatible with power-efficient hardware is proposed. • The proposed dynamic policies implemented at various levels of the embedded enable low-cost monitors that can be modified based on application demands. • The runtime costs of the processor pipeline using OptimSoC, an open-source implementation of OpenRISC1000 architecture where processor policies are evaluated, which showed less than 1% overhead in power and area when implemented in 14 nm and 45 nm technologies and no performance overhead when implemented on BEEBS benchmark suite. Garett Cunningham, Siqin Liu, Harsha Chenji, David W. Juedes, Avinash Karanth |
Integr. | 5 |
| 2025 | MERIT: A Sustainable DNN Accelerator Design With Photonic Phase-Change MemoryabstractThe growing computational demands of deep learning have driven interest in analog neural networks using resistive memory and silicon photonics. However, these technologies face inherent limitations in computing parallelism when used independently. Photonic phase-change memory (PCM), which integrates photonics with PCM, overcomes these constraints by enabling simultaneous processing of multiple inputs encoded on different wavelengths, significantly enhancing parallel computation for deep neural network (DNN) inference and training. This paper presents MERIT, a sustainable DNN accelerator that capitalizes on the non-volatility of resistive memory and the high operating speed of photonic devices. MERIT enables seamless inference and training by loading weight kernels into photonic PCM arrays and selectively supplying light encoded with input features for the forward pass and loss gradients for the backward pass. We compare MERIT with state-of-the-art digital and analog DNN accelerators including TPU, DEAP, and PTC. Simulation results demonstrate that MERIT reduces execution time by 68% and energy consumption by 64% for inference, and reduces execution time by 79% and energy consumption by 84% for training. Yuan Li 0029, Ahmed Louri, Avinash Karanth |
IEEE Trans. Sustain. Comput. | 3 |
| 2025 | GreeNX: An Energy-Efficient and Sustainable Approach to Sparse Graph Convolution Networks Accelerators Using DVFSabstractGraph convolutional networks (GCNs) have emerged as an effective approach to extend deep learning algorithms for graph-based data analytics. However, GCNs implementation over large, sparse datasets presents challenges due to irregular computation and dataflow patterns. Specialized GCN accelerators have emerged to deliver superior performance over generic processors. However, prior techniques that include specialized datapaths, optimized sparse computation, and memory access patterns, handle different phases of GCNs differently which results in excess energy consumption and reduced throughput due to sub-optimal dataflows. In this paper, we propose GreeNX, a computation and communication-aware GCN accelerator that uniformly applies three complementary techniques to all phases of GCN. First, we abstract two cascaded sparse-dense matrix multiplications that uniformly process the computation in both aggregation and combination phases of GCNs to improve throughput. Second, to mitigate the overheads of processing irregular sparse data, we develop a dynamic-voltage-and-frequency-scaling (DVFS) scheme by grouping a row of processing elements (PEs) that dynamically changes the applied voltage/frequency (V/F) to improve energyefficiency. Third, we conduct a comprehensive carbon footprint evaluation, analyzing both embodied and operational emissions for GCNs. Extensive simulation and experiments validate that our GreeNX consistently reduces memory accesses and energy consumption leading to an average 7.3× speedup and 5.6× energy savings on six real-world graph datasets over several state-of-the-art GCN accelerators including HyGCN, AWB-GCN, GCoD, GRIP, IGCN, and LW-GCN Siqin Liu, Prakash Chand Kuve, Avinash Karanth |
IEEE Trans. Sustain. Comput. | 3 |
| 2024 | d-GUARD: Thwarting Denial-of-Service Attacks via Hardware Monitoring of Information Flow using Language Semantics in Embedded SystemsabstractAs low-level embedded systems are vulnerable to attacks that exploit flaws in either hardware or software, it is essential to enforce secure policies to protect the system from malicious instructions that significantly alter program behavior. To improve efficiency of implementation, high-level secure policy languages have been defined such that the policies can be directly synthesized into hardware monitors. However, the language semantics define policies that are static throughout the program execution which limits the flexibility. Moreover, secure policies target processor pipelines and not the network-on-chip (NoC) connecting several processor where denial-of-service attacks could originate. In this paper, we enable dynamically reconfigurable security policies through a high-level language called D-GUARD that target both processor pipeline and NoC architecture in mutlicore embedded systems. Alongside static policies, D-GUARD’s semantics support policies that dynamically change behavior in response to program conditions at runtime. In addition, we also propose policies to thwart denial-of-service attacks by rate limiting the packet flow into the network using the same dynamic policies expressed by D-GUARD. We describe a Verilog compiler to support realizing policies as hardware monitors for both processor pipelines and network interfaces. D-GUARD is developed using the Coq proof assistant, enabling the formal verification of policy correctness and other properties. This approach takes advantage of the abstractions and expressiveness of a higher-level language while minimizing the overhead that comes with other general-purpose approaches implemented purely in hardware, as well as offering the groundwork for a formally verified tool chain. Garett Cunningham, Harsha Chenji, David W. Juedes, Avinash Karanth |
ASPDAC | 4 |
| 2024 | SNAC: Mitigation of Snoop-Based Attacks with Multi-Tier Security in NoC ArchitecturesabstractNetwork-on-chips (NoCs) are crucial for multicore and manycore System-on-Chip (SoC) architectures. However, the integration of third-party Intellectual Property (IP) cores in SoCs has introduced hardware vulnerabilities. Snoop-based attacks exploit these vulnerabilities by inserting malicious Hardware Trojans into routers, allowing them to extract sensitive information as packets traverse the NoC. To address these security concerns, we propose SNAC: Mitigation of Snoop-based Attacks in NoCs. SNAC employs a three-tier architecture with increasing security levels, each with proportional power and latency overheads. The first tier introduces path randomization to prevent attackers from predicting packet routes. In the second tier, we encrypt source and destination information using lightweight backward XoR encryption. The third tier combines techniques from tiers one and two, extending obfuscation along with path randomization. SNAC was evaluated using synthetic and real-world benchmarks. Our results show that SNAC incurs dynamic power overheads of 4.2%, 3.9%, and 6.1% for Tiers 1, 2, and 3 respectively, with area overheads of 6.2%, 4.2%, and 9.2%. Siqin Liu, Saumya Chauhan, Avinash Karanth |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | HSCONN: Hardware-Software Co-Optimization of Self-Attention Neural Networks for Large Language ModelsabstractSelf-attention models excel in natural language processing and computer vision by capturing contextual information but encounter several challenges such as efficient data movement, quadratic computational complexity, and excessive memory accesses. Sparse attention techniques emerge as a solution, however, their irregular or regular patterns, coupled with costly data pre-processing, diminish their hardware efficiency. This paper introduces HSCONN, an energy-efficient hardware accelerator for self-attention, mitigating computational and memory overheads. HSCONN employs dynamic voltage and frequency scaling (DVFS) along with exploiting dynamic sparsity in matrix multiplication, thereby optimizing energy efficiency. The approach includes a row-wise pruning algorithm and independent voltage/frequency islands for processing elements, exploiting additional sparsity to reduce memory access and overall energy consumption. Experiments in natural language processing showcase HSCONN’s remarkable speedups (1952 ×, 615 ×) and energy reductions (up to 820 ×, 113 ×) over CPU and GPU architectures. Compared to A3, SpAtten, and Sanger, HSCONN demonstrates superior speedup (1.71 ×, 1.25 ×, 1.47 ×) and higher energy efficiency (1.5 ×, 1.7 ×, 1.4 ×). Siqin Liu, Prakash Chand Kuve, Avinash Karanth |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | Training Photonic Mach Zehnder Meshes for Neural Network AccelerationabstractPhotonic neural networks enable faster and energy-efficient inferences for deep neural network implementations when compared to electrical counterparts. Prior training methods for photonic neural networks modify the network or decompose the inputs or weights to simplify the training, however that could impact accuracy. Electrical training uses backpropagation that needs to partially differentiate each of the weights with respect to the loss function on a number of examples to achieve high accuracy. While the traditional backpropagation algorithm can be applied to electrical networks, it is difficult to directly apply to photonic neural network with Mach-Zehnder Interferometer (MZI) meshes that have different parameters (frequency and phase). In this paper, we implement the traditional backprop-agation algorithm to train the parameters of the MZI meshes with gradient descent. Parameters of the mesh are trained by implementing a modified backpropagation algorithm on the hardware, and those operations may be performed on device. The implementation does not modify the circuits or add any additional lasers and therefore, self-trains the MZI to achieve the desired accuracy and achieve O(n2) speedup over electrical methods. We test our training algorithms on MNIST datasets and show that our method improves the training when compared to state-of-the-art electrical backpropagation. Andy Wolff, Avinash Karanth |
HiPC | 2 |
| 2024 | Editorial: EiC Farewell and Introduction of new EiCabstractTransactions on Computers (TC) draws to a close on Avinash Karanth |
IEEE Trans. Computers | 1 |
| 2024 | A High-Performance and Energy-Efficient Photonic Architecture for Multi-DNN AccelerationabstractLarge-scale deep neural network (DNN) accelerators are poised to facilitate the concurrent processing of diverse DNNs, imposing demanding challenges on the interconnection fabric. These challenges encompass overcoming performance degradation and energy increase associated with system scaling while also necessitating flexibility to support dynamic partitioning and adaptable organization of compute resources. Nevertheless, conventional metallic-based interconnects frequently confront inherent limitations in scalability and flexibility. In this paper, we leverage silicon photonic interconnects and adopt an algorithm-architecture co-design approach to develop MDA, a DNN accelerator meticulously crafted to empower high-performance and energy-efficient concurrent processing of diverse DNNs. Specifically, MDA consists of three novel components: 1) a resource allocation algorithm that assigns compute resources to concurrent DNNs based on their computational demands and priorities; 2) a dataflow selection algorithm that determines off-chip and on-chip dataflows for each DNN, with the objectives of minimizing off-chip and on-chip memory accesses, respectively; 3) a flexible silicon photonic network that can be dynamically segmented into sub-networks, each interconnecting the assigned compute resources of a certain DNN while adapting to the communication patterns dictated by the selected on-chip dataflow. Simulation results show that the proposed MDA accelerator outperforms other state-of-the-art multi-DNN accelerators, including PREMA, AI-MT, Planaria, and HDA. MDA accelerator achieves a speedup of 3.6, accompanied by substantial improvements of 7.3×, 12.7×, and 9.2× in energy efficiency, service-level agreement (SLA) satisfaction rate, and fairness, respectively. Yuan Li 0029, Ahmed Louri, Avinash Karanth |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | A Silicon Photonic Multi-DNN AcceleratorabstractIn shared environments like cloud-based datacenters, hardware accelerators are deployed to meet the scale-out computation demands of deep neural network (DNN) inference tasks. As conventional hardware accelerators optimized for single-DNN execution cannot effectively resolve the dynamic interaction of these inference-as-a-service (INFaaS) tasks, several multi-DNN hardware accelerators have been developed to improve the overall system performance while adhering to the constraints of individual tasks. Some of such multi-DNN hardware accelerators temporally schedule tasks by incorporating the preemption or load-balancing-based algorithm but suffer from resource underutilization because of unmanaged mismatch between resource demand and provision. Other multi-DNN hardware accelerators enable spatial colocation of tasks to improve resource utilization and system flexibility, but the irregular communication patterns between the fragmented resource partitions cannot be adequately supported by the metallic-based interconnects due to their rigidity and other inherent scaling limitations. We introduce a photonic multi-DNN accelerator named Aspire in this paper. The fundamental novelty of Aspire lies in the ability to adaptively create sub-accelerators for different tasks by assembling fine-grained resource partitions in the same architecture. Seamless communications between those fragmented resource partitions from a sub-accelerator are realized by exploiting photonic interconnects. Specifically, Aspire includes three novel designs: (1) a photonic network that can be adaptively partitioned into several sub-networks, each seamlessly connecting the fragmented resource partitions to construct sub-accelerators; (2) a dataflow that simultaneously leverages temporal and spatial data reuse opportunities within each resource partition and across several resource partitions, respectively; (3) an algorithm that allocates resource partitions at task granularity and derives optimal tile size and execution order at DNN layer granularity. Simulation studies show that Aspire outperforms other state-of-the-art multi-DNN accelerators, delivering 64% execution time reduction, 69% energy saving, 51% improvement in service-level agreement satisfaction rate, and 7.9× improvement in fairness. Yuan Li 0029, Ahmed Louri, Avinash Karanth |
PACT | 3 |
| 2023 | Flumen: Dynamic Processing in the Photonic InterconnectabstractIn chiplet-based heterogeneous architectures, electrical network-on-package (NoP) designs are typically over-provisioned with routers and channels to provide sufficient bandwidth during periods of high network load. Observing that there are significant periods of low/idle network utilization, prior work has proposed modified network-on-chip (NoC) architectures to enable in-network compute, especially for compute-intensive operations (e.g. linear algebra). However, electrical package-level interconnects impose fundamental energy and bandwidth scaling issues for future chiplet architectures. Kyle Shiflett, Avinash Karanth, Razvan C. Bunescu, Ahmed Louri |
ISCA | 2 |
| 2023 | Reflections of Cybersecurity Workshop for K-12 TeachersabstractIn this paper, we recount efforts in developing cybersecurity workshops for K-12 teachers, intended to learn skills to better educate the cybersecurity workers of tomorrow. In 2021, we provided two one-day virtual workshops and in 2022 we provided one two-day in-person workshop to high school teachers to increase cybersecurity awareness in three areas: general cybersecurity issues, software security, and hardware security. Both the online and in-person workshops employed Google classroom and Jupyter Notebooks, and high school teachers were provided with Raspberry Pi Zeros to use as part of the workshops. This paper describes the design and implementation of the workshops and also provides evidence demonstrating the effectiveness of the workshops, as well as commentary to provide guidance for future efforts. Chad Mourning, Harsha Chenji, Allyson Hallman-Thrasher, Savas Kaya, Nasseef Abukamail, David W. Juedes, Avinash Karanth |
SIGCSE (1) | 7 |
| 2022 | SPACX: Silicon Photonics-based Scalable Chiplet Accelerator for DNN InferenceabstractIn pursuit of higher inference accuracy, deep neural network (DNN) models have significantly increased in complexity and size. To overcome the consequent computational challenges, scalable chiplet-based accelerators have been proposed. However, data communication using metallic-based interconnects in these chiplet-based DNN accelerators is becoming a primary obstacle to performance, energy efficiency, and scalability. The photonic interconnects can provide adequate data communication support due to some superior properties like low latency, high bandwidth and energy efficiency, and ease of broadcast communication. In this paper, we propose SPACX: a Silicon Photonics-based Chiplet ACcelerator for DNN inference applications. Specifically, SPACX includes a photonic network design that enables seamless single-chiplet and cross-chiplet broadcast communications, and a tailored dataflow that promotes data broadcast and maximizes parallelism. Furthermore, we explore the broadcast granularities of the photonic network and implications on system performance and energy efficiency. A flexible bandwidth allocation scheme is also proposed to dynamically adjust communication bandwidths for different types of data. Simulation results using several DNN models show that SPACX can achieve 78% and 75% reduction in execution time and energy, respectively, as compared to other state-of-the-art chiplet-based DNN accelerators. Yuan Li 0029, Ahmed Louri, Avinash Karanth |
HPCA | 3 |
| 2022 | Reflections of Cybersecurity Workshop for K-12 Teachers and High School StudentsabstractIn this paper, we describe efforts to promote a robust cyber security workforce through a series of online workshops for K-12 teachers and grades 7-12 students. In 2021, we provided virtual workshops to high school teachers and students to increase cyber security awareness in three areas, (i) general cybersecurity issues, (ii) software security, and (iii) hardware security. The workshops employed Google classroom and Jupyter Notebooks, and high school teachers were provided with hardware (Raspberry Pi Zeros) to use as part of the workshops. We describe the design and implementation of the workshops and share evidence to demonstrate the effectiveness of the workshops and provide insights for future professional development for teachers. Our central question was: What impact do workshops have on teachers' preparation to effectively teach cybersecurity topics to their students and what do teachers report learning from their experiences in the workshop? We provide insights into workshops' effectiveness for computer science teachers and for STEM teachers who are not computer science teachers. Chad Mourning, David W. Juedes, Allyson Hallman-Thrasher, Harsha Chenji, Savas Kaya, Avinash Karanth |
SIGCSE (2) | 6 |
| 2022 | Ascend: A Scalable and Energy-Efficient Deep Neural Network Accelerator With Photonic InterconnectsabstractThe complexity and size of recent deep neural network (DNN) models have increased significantly in pursuit of high inference accuracy. Chiplet-based accelerator is considered a viable scaling approach to provide substantial computation capability and on-chip memory for efficient process of such DNN models. However, communication using metallic interconnects in prior chiplet-based accelerators poses a major challenge to system performance, energy efficiency, and scalability. Photonic interconnects can adequately support communication across chiplets due to features such as distance-independent latency, high bandwidth density, and high energy efficiency. Furthermore, the salient ease of broadcast property makes photonic interconnects suitable for DNN inference which often incurs prevalent broadcast communication. In this paper, we propose a scalable chiplet-based DNN accelerator with photonic interconnects named ASCEND. ASCEND introduces (1) a novel photonic network that supports seamless intra- and inter- chiplet broadcast communication, and flexible mapping of diverse convolution layers, and (2) a tailored dataflow that exploits the ease of broadcast property and maximizes parallelism by simultaneously processing computations with shared input data. Simulation results using multiple DNN models show that ASCEND achieves 71% and 67% reduction in execution time and energy consumption, respectively, as compared to other state-of-the-art chiplet-based DNN accelerators with metallic or photonic interconnects. Yuan Li 0029, Ke Wang 0030, Hao Zheng 0005, Ahmed Louri, Avinash Karanth |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Exploiting Wireless Technology for Energy-Efficient Accelerators With Multiple Dataflows and PrecisionabstractAs model size and the number of layers increase, Deep Neural Networks (DNNs) demand enormous computational power and throughput to meet exceedingly high prediction accuracy’s of today’s machine learning (ML) applications. Spatial hardware accelerators have been proposed that optimize the dataflow and exploit sparsity to provide a significant decrease in power consumption. As spatial architectures are traditionally designed with metallic interconnects, significant power is expended for data movement for different dataflows. In this paper, we exploit extended wireless technology to design a power-efficient and high-throughput DNN accelerator, e-WiNN, that can be configured for all representative dataflows and arithmetic precisions. We leverage novel circuit design by utilizing Dadda-algorithm based Multiply-and-Accumulate (MAC) circuits for 4-bit, 8-bit and 16-bit inputs to reduce area, power and delay constraints in 14 nm predictive technology. Our novel wireless transmitter integrates on- off keying (OOK) modulator with power amplifier that results in significant energy savings. To reduce the area overhead, we cluster wireless transceivers into groups of four such that both weights and input features can be effectively multicast to reduce the data movement. The energy efficient transceiver circuit is implemented in state-of-the-art BSIM 32 nm FinFET technology model and our link budget considers required RF power for different frequencies and inter-PE distance at three different antenna directivities including isotropic. Our detailed RTL modeling and cycle-accurate simulation results show that e-WiNN achieves 36.3% latency reduction and 76.1% energy saving when compared to state-of-art wire interconnected accelerators; 70.3% area reduction and 41.6% energy saving at the cost of 11% latency increase when compared to prior wireless accelerators on various neural networks (AlexNet, VGG16, and ResNet-9/50). Siqin Liu, Talha Furkan Canan, Harsha Chenji, Soumyasanta Laha, Savas Kaya, Avinash Karanth |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | SPRINT: A High-Performance, Energy-Efficient, and Scalable Chiplet-Based Accelerator With Photonic Interconnects for CNN InferenceabstractChiplet-based convolution neural network (CNN) accelerators have emerged as a promising solution to provide substantial processing power and on-chip memory capacity for CNN inference. The performance of these accelerators is often limited by inter-chiplet metallic interconnects. Emerging technologies such as photonic interconnects can overcome the limitations of metallic interconnects due to several superior properties including high bandwidth density and distance-independent latency. However, implementing photonic interconnects in chiplet-based CNN accelerators is challenging and requires combined effort of network architectural optimization and CNN dataflow customization. In this article, we propose SPRINT, a chiplet-based CNN accelerator that consists of a global buffer and several accelerator chiplets. SPRINT introduces two novel designs: (1) a photonic inter-chiplet network that can adapt to specific communication patterns in CNN inference through wavelength allocation and waveguide reconfiguration, and (2) a CNN dataflow that can leverage the broadcasting capability of photonic interconnects while minimizing the costly electrical-to-optical and optical-to-electrical signal conversions. Simulations using multiple CNN models show that SPRINT achieves up to 76% and 68% reduction in execution time and energy consumption, respectively, as compared to other state-of-the-art chiplet-based architectures with either metallic or photonic interconnects. Yuan Li 0029, Ahmed Louri, Avinash Karanth |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Scaling Deep-Learning Inference with Chiplet-based Architecture and Photonic InterconnectsabstractChiplet-based architectures have been proposed to scale computing systems for deep neural networks (DNNs). Prior work has shown that for the chiplet-based DNN accelerators, the electrical network connecting the chiplets poses a major challenge to system performance, energy consumption, and scalability. Some emerging interconnect technologies such as silicon photonics can potentially overcome the challenges facing electrical interconnects as photonic interconnects provide high bandwidth density, superior energy efficiency, and ease of implementing broadcast and multicast operations that are prevalent in DNN inference. In this paper, we propose a chiplet-based architecture named SPRINT for DNN inference. SPRINT uses a global buffer to simplify the data transmission between storage and computation, and includes two novel designs: (1) a reconfigurable photonic network that can support diverse communications in DNN inference with minimal implementation cost, and (2) a customized dataflow that exploits the ease of broadcast and multicast feature of photonic interconnects to support highly parallel DNN computations. Simulation studies using ResNet50 DNN model show that SPRINT achieves 46% and 61% execution time and energy consumption reduction, respectively, as compared to other state-of-the-art chiplet-based architectures with electrical or photonic interconnects. Yuan Li 0029, Ahmed Louri, Avinash Karanth |
DAC | 3 |
| 2021 | Bitwise Neural Network Acceleration Using Silicon PhotonicsabstractHardware accelerators provide significant speedup and improve energy efficiency for several demanding deep neural network (DNN) applications. DNNs have several hidden layers that perform concurrent matrix-vector multiplications (MVMs) between the network weights and input features. As MVMs are critical to the performance of DNNs, previous research has optimized the performance and energy efficiency of MVMs at both the architecture and algorithm levels. In this paper, we propose to use emerging silicon photonics technology to improve parallelism, speed and overall efficiency with the goal of providing real-time inference and fast training of neural nets. We use microring resonators (MRRs) and Mach-Zehnder interferometers (MZIs) to design two versions (all-optical and partial-optical) of hybrid matrix multiplications for DNNs. Our results indicate that our partial optical design gave the best performance in both energy efficiency and latency, with a reduction of 33.1% for energy-delay product (EDP) with conservative estimates and a 76.4% reduction for EDP with aggressive estimates. Kyle Shiflett, Avinash Karanth, Ahmed Louri, Razvan C. Bunescu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | Dynamic Voltage and Frequency Scaling to Improve Energy-Efficiency of Hardware AcceleratorsabstractNeural networks (NNs) have been used in a wide variety of artificial intelligence (AI) applications, including speech recognition, image recognition, automatic robotics, and games. State-of-the-art NNs provide high prediction accuracy at the expense of massive computation that involves large model parameters which consume substantial energy. Though sparse NN s have emerged to reduce the computation and storage overhead, existing specialized DNN accelerators cannot maximize the energy savings when exploiting both dynamic and static sparsity, especially for irregular NNs. In this paper, we propose a dynamic voltage and frequency scaling (DVFS) based hardware accelerator that effectively exploits the dynamic and static sparsity of NNs with dynamic voltage/frequency (V/F) scaling and power gating techniques to reduce both static and dynamic power. To explore the efficiency of DVFS implementation at different granularities, we evaluate both coarse-grained and fine-grained DVFS implementation with different design trade-offs. Further, our proposed DVFS model predicts the dynamic computation workloads as well as V/ F pairs to be supplied to processing elements (PEs) in the hardware intelligently through pre-trained weight vectors. The machine learning based prediction algorithm is deployed to improve the DVFS mode selection accuracy. Our simulation results on AlextNet, VGG16, and ResNet50 show that we can achieve an average dynamic energy savings of 59–66 % and an average static power reduction of 69–80 % compared to the baseline. Siqin Liu, Avinash Karanth |
HiPC | 2 |
| 2021 | CSCNN: Algorithm-hardware Co-design for CNN Accelerators using Centrosymmetric FiltersabstractConvolutional neural networks (CNNs) are at the core of many state-of-the-art deep learning models in computer vision, speech, and text processing. Training and deploying such CNN-based architectures usually require a significant amount of computational resources. Sparsity has emerged as an effective compression approach for reducing the amount of data and computation for CNNs. However, sparsity often results in computational irregularity, which prevents accelerators from fully taking advantage of its benefits for performance and energy improvement. In this paper, we propose CSCNN, an algorithm/hardware co-design framework for CNN compression and acceleration that mitigates the effects of computational irregularity and provides better performance and energy efficiency. On the algorithmic side, CSCNN uses centrosymmetric matrices as convolutional filters. In doing so, it reduces the number of required weights by nearly 50% and enables structured computational reuse without compromising regularity and accuracy. Additionally, complementary pruning techniques are leveraged to further reduce computation by a factor of $2.8-7.2\times $ with a marginal accuracy loss. On the hardware side, we propose a CSCNN accelerator that effectively exploits the structured computational reuse enabled by centrosymmetric filters, and further eliminates zero computations for increased performance and energy efficiency. Compared against a dense accelerator, SCNN and SparTen, the proposed accelerator performs $3.7\times $, $1.6\times $ and $1.3\times $ better, and improves the EDP (Energy Delay Product) by $8.9\times $, $2.8\times $ and $2.0\times $, respectively. Ahmed Louri, Avinash Karanth, Razvan C. Bunescu |
HPCA | 3 |
| 2021 | GCNAX: A Flexible and Energy-efficient Accelerator for Graph Convolutional Neural NetworksabstractGraph convolutional neural networks (GCNs) have emerged as an effective approach to extend deep learning for graph data analytics. Given that graphs are usually irregular, as nodes in a graph may have a varying number of neighbors, processing GCNs efficiently pose a significant challenge on the underlying hardware. Although specialized GCN accelerators have been proposed to deliver better performance over generic processors, prior accelerators not only under-utilize the compute engine, but also impose redundant data accesses that reduce throughput and energy efficiency. Therefore, optimizing the overall flow of data between compute engines and memory, i.e., the GCN dataflow, which maximizes utilization and minimizes data movement is crucial for achieving efficient GCN processing.In this paper, we propose a flexible and optimized dataflow for GCNs that simultaneously improves resource utilization and reduces data movement. This is realized by fully exploring the design space of GCN dataflows and evaluating the number of execution cycles and DRAM accesses through an analysis framework. Unlike prior GCN dataflows, which employ rigid loop orders and loop fusion strategies, the proposed dataflow can reconFigure the loop order and loop fusion strategy to adapt to different GCN configurations, which results in much improved efficiency. We then introduce a novel accelerator architecture called GCNAX, which tailors the compute engine, buffer structure and size based on the proposed dataflow. Evaluated on five real-world graph datasets, our simulation results show that GCNAX reduces DRAM accesses by a factor of $8.1 \times$ and $2.4 \times$, while achieving $8.9 \times, 1.6 \times$ speedup and $9.5 \times$, $2.3 \times$ energy savings on average over HyGCN and AWB-GCN, respectively. Ahmed Louri, Avinash Karanth, Razvan C. Bunescu |
HPCA | 3 |
| 2021 | WiNN: Wireless Interconnect based Neural Network AcceleratorabstractDeep Neural Networks (DNNs) have demonstrated promising performance in accuracy for several applications such as image processing, speech recognition, and autonomous systems and vehicles. Spatial accelerators have been proposed to achieve high parallelism with arrays of processing elements (PE) and energy efficient data movement using traditional Network-on-Chip (NoC) architectures. However, larger DNN models impose high bandwidth and low latency communication demands between PEs, which is a fundamental challenge for metallic NoC architectures. In this paper, we propose WiNN, a wireless and wired interconnected neural network accelerator that employs on-chip wireless links to provide high network bandwidth and single cycle multicast communication. We design separate wireless networks modulated with two different frequency bands one each for the weights and input Highly directional antennas are implemented to avoid noise and interference. We propose multicast-for-wireless (MW) dataflow for our proposed accelerator that efficiently exploits the wireless channels’ multicast capabilities to reduce the communication overheads. Our novel wireless transmitter integrates on-off keying (OOK) modulator with power amplifier that results in significant energy savings. Our simulation results show that WiNN achieves 74% latency reduction and 37.5% energy saving when compared to state-of-art metallic link-based accelerators, 38.1% latency reduction and 19.4% energy saving when compared to prior wireless accelerators for various neural networks (AlexNet, VGG16, and ResNet-50). Siqin Liu, Sushanth Karmunchi, Avinash Karanth, Soumyasanta Laha, Savas Kaya |
ICCD | 3 |
| 2021 | Albireo: Energy-Efficient Acceleration of Convolutional Neural Networks via Silicon PhotonicsabstractWith the end of Dennard scaling, highly-parallel and specialized hardware accelerators have been proposed to improve the throughput and energy-efficiency of deep neural network (DNN) models for various applications. However, collective data movement primitives such as multicast and broadcast that are required for multiply-and-accumulate (MAC) computation in DNN models are expensive, and require excessive energy and latency when implemented with electrical networks. This consequently limits the scalability and performance of electronic hardware accelerators. Emerging technology such as silicon photonics can inherently provide efficient implementation of multicast and broadcast operations, making photonics more amenable to exploit parallelism within DNN models. Moreover, when coupled with other unique features such as low energy consumption, high channel capacity with wavelength-division multiplexing (WDM), and high speed, silicon photonics could potentially provide a viable technology for scaling DNN acceleration.In this paper, we propose Albireo, an analog photonic architecture for scaling DNN acceleration. By characterizing photonic devices such as microring resonators (MRRs) and Mach-Zehnder modulators (MZM) using photonic simulators, we develop realistic device models and outline their capability for system level acceleration. Using the device models, we develop an efficient broadcast combined with multicast data distribution by leveraging parameter sharing through unique WDM dot product processing. We evaluate the energy and throughput performance of Albireo on DNN models such as ResNet18, MobileNet and VGG16. When compared to cur-rent state-of-the-art electronic accelerators, Albireo increases throughput by 110 X, and improves energy-delay product (EDP) by an average of 74 X with current photonic devices. Furthermore, by considering moderate and aggressive photonic scaling, the proposed Albireo design shows that EDP can be reduced by at least 229 X. Kyle Shiflett, Avinash Karanth, Razvan C. Bunescu, Ahmed Louri |
ISCA | 2 |
| 2020 | PIXEL: Photonic Neural Network AcceleratorabstractMachine learning (ML) architectures such as Deep Neural Networks (DNNs) have achieved unprecedented accuracy on modern applications such as image classification and speech recognition. With power dissipation becoming a major concern in ML architectures, computer architects have focused on designing both energy-efficient hardware platforms as well as optimizing ML algorithms. To dramatically reduce power consumption and increase parallelism in neural network accelerators, disruptive technology such as silicon photonics has been proposed which can improve the performance-per-Watt when compared to electrical implementation. In this paper, we propose PIXEL - Photonic Neural Network Accelerator that efficiently implements the fundamental operation in neural computation, namely the multiply and accumulate (MAC) functionality using photonic components such as microring resonators (MRRs) and Mach-Zehnder interferometer (MZI). We design two versions of PIXEL - a hybrid version that multiplies optically and accumulates electrically and a fully optical version that multiplies and accumulates optically. We perform a detailed power, area and timing analysis of the different versions of photonic and electronic accelerators for different convolution neural networks (AlexNet, VGG16, and others). Our results indicate a significant improvement in the energy-delay product for both PIXEL designs over traditional electrical designs (48.4% for OE and 73.9% for OO) while minimizing latency, at the cost of increased area over electrical designs. Kyle Shiflett, Dylan Wright, Avinash Karanth, Ahmed Louri |
HPCA | 3 |
| 2020 | DozzNoC: Reducing Static and Dynamic Energy in NoCs with Low-latency Voltage Regulators using Machine LearningabstractNetwork-on-chips (NoCs) continues to be the choice of communication fabric in multicore architectures because the NoC effectively combines the resource efficiency of the bus with the parallelizability of the crossbar. As NoC suffers from both high static and dynamic energy consumption, power-gating and dynamic voltage and frequency scaling (DVFS) have been proposed in the literature to improve energy-efficiency. In this work, we propose DozzNoC, an adaptable power management technique that effectively combines power-gating and DVFS techniques to target both static power and dynamic energy reduction with a single inductor multiple output (SIMO) voltage regulator. The proposed power management design is further enhanced by machine learning techniques that predict future traffic load for proactive DVFS mode selection. DozzNoC utilizes a SIMO voltage regulator scheme that allows for fast, low-powered, and independently power-gated or voltage scaled routers such that each router and its outgoing links share the same voltage/frequency domain. Our simulation results using PARSEC and Splash-2 benchmarks on an 8 × 8 mesh network show that for a decrease of 7% in throughput, we can achieve an average dynamic energy savings of 25% and an average static power reduction of 53%. Mark Clark, Yingping Chen, Avinash Karanth, Dongsheng Ma 0001, Ahmed Louri |
IPDPS | 3 |
| 2020 | Guest Editors' Introduction to the Special Issue on Machine Learning Architectures and AcceleratorsabstractThe twelve papers in this special section focus on machine learning architectures and accelerators. Deep learning or deep neural networks (DNNs), as one of the most powerful machine learning techniques, has achieved extraordinary performance in computer vision and surveillance, speech recognition and natural language processing, healthcare and disease diagnosis, etc. Various forms of DNNs have been proposed, including Convolutional Neural Networks, Recurrent Neural Networks, Deep Reinforcement Learning, Transformer model, etc. Deep learning exhibits an offline training phase to derive the weight parameters from an excessive training dataset, as well as an online inference phase to perform classification/prediction/perception/ control tasks based on the trained model. The paper in this section aim to find a convergence of software and hardware/architecture. It aims at DNN algorithms, parallel computing, and compiler code generation techniques that are hardware/architecture friendly, as well as computer architectures that are universal and consistently highly performant on a wide range of DNN algorithms and applications. In this co-design and co-optimization framework we can mitigate the limitation of investigating in only a single direction, shedding some light on the future of embedded, ubiquitous artificial intelligence. Xuehai Qian, Yanzhi Wang 0001, Avinash Karanth |
IEEE Trans. Computers | 3 |
| 2020 | Hardware-Level Thread Migration to Reduce On-Chip Data Movement Via Reinforcement LearningabstractAs the number of processing cores and associated threads in chip multiprocessors (CMPs) continues to scale out, on-chip memory access latency dominates application execution time due to increased data movement. Although tiled CMP architectures with distributed shared caches provide a scalable design, increased physical distance between requesting and responding cores has led to both increased on-chip memory access latency and excess energy consumption. Near data processing is a promising approach that can migrate threads closer to data, however prior hand-engineered rules for fine-grained hardware-level thread migration are either too slow to react to changes in data access patterns, or unable to exploit the large variety of data access patterns. In this article, we propose to use reinforcement learning (RL) to learn relatively complex data access patterns to improve on hardware-level thread migration techniques. By utilizing the recent history of memory access locations as input, each thread learns to recognize the relationship between prior access patterns and future memory access locations. This leads to the unique ability of the proposed technique to make fewer, more effective migrations to intermediate cores that minimize the distance to multiple distinct memory access locations. By allowing a low-overhead RL agent to learn a policy from real interaction with parallel programming benchmarks in a parallel simulator, we show that a migration policy which recognizes more complex data access patterns can be learned. The proposed approach reduces on-chip data movement and energy consumption by an average of 41%, while reducing execution time by 43% when compared to a simple baseline with no thread migration; furthermore, energy consumption and execution time are reduced by an additional 10% when compared to a hand-engineered fine-grained migration policy. Quintin Fettes, Avinash Karanth, Razvan C. Bunescu, Ahmed Louri, Kyle Shiflett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | High-performance, Energy-efficient, Fault-tolerant Network-on-Chip Design Using Reinforcement LearninabstractNetwork-on-Chips (NoCs) are becoming the standard communication fabric for multi-core and system on a chip (SoC) architectures. As technology continues to scale, transistors and wires on the chip are becoming increasingly vulnerable to various fault mechanisms, especially timing errors, resulting in exacerbation of energy efficiency and performance for NoCs. Typical techniques for handling timing errors are reactive in nature, responding to the faults after their occurrence. They rely on error detection/correction techniques which have resulted in excessive power consumption and degraded performance, since the error detection/correction hardware is constantly enabled. On the other hand, indiscriminately disabling error handling hardware can induce more errors and intrusive retransmission traffic. Therefore, the challenge is to balance the trade-offs among error rate, packet retransmission, performance, and energy. In this paper, we propose a proactive fault-tolerant mechanism to optimize energy efficiency and performance with reinforcement learning (RL). First, we propose a new proactive error handling technique comprised of a dynamic scheme for enabling per-router error detection/correction hardware and an effective retransmission mechanism. Second, we propose the use of RL to train the dynamic control policy with the goals of providing increased fault-tolerance, reduced power consumption and improved performance as compared to conventional techniques. Our evaluation indicates that, on average, end-to-end packet latency is lowered by 55%, energy efficiency is improved by 64%, and retransmission caused by faults is reduced by 48% over the reactive error correction techniques. Ke Wang 0030, Ahmed Louri, Avinash Karanth, Razvan C. Bunescu |
DATE | 3 |
| 2019 | IntelliNoC: a holistic design framework for energy-efficient and reliable on-chip communication for manycoresabstractAs technology scales, Network-on-Chips (NoCs), currently being used for on-chip communication in manycore architectures, face several problems including high network latency, excessive power consumption, and low reliability. Simultaneously addressing these problems is proving to be difficult due to the explosion of the design space and the complexity of handling many trade-offs. In this paper, we propose IntelliNoC, an intelligent NoC design framework which introduces architectural innovations and uses reinforcement learning to manage the design complexity and simultaneously optimize performance, energy-efficiency, and reliability in a holistic manner. IntelliNoC integrates three NoC architectural techniques: (1) multifunction adaptive channels (MFACs) to improve energy-efficiency; (2) adaptive error detection/correction and re-transmission control to enhance reliability; and (3) a stress-relaxing bypass feature which dynamically powers off NoC components to prevent overheating and fatigue. To handle the complex dynamic interactions induced by these techniques, we train a dynamic control policy using Q-learning, with the goal of providing improved fault-tolerance and performance while reducing power consumption and area overhead. Simulation using PARSEC benchmarks shows that our proposed IntelliNoC design improves energy-efficiency by 67% and mean-time-to-failure (MTTF) by 77%, and decreases end-to-end packet latency by 32% and area requirements by 25% over baseline NoC architecture. Ke Wang 0030, Ahmed Louri, Avinash Karanth, Razvan C. Bunescu |
ISCA | 3 |
| 2019 | Limit of Hardware Solutions for Self-Protecting Fault-Tolerant NoCsabstractWe study the ultimate limits of hardware solutions for the self-protection strategies against permanent faults in networks on chips (NoCs). NoCs reliability is improved by replacing each base router by an augmented router which includes extra protection circuitry. We compare the protection achieved by the self-test and self-protect (STAP) architectures to that of triple modular redundancy with voting (TMR). Two STAP architectures are considered. In the first one, a defective router self-disconnects from the network, while it self-heals in the second one. In practice, none of the considered architectures (STAP or TMR) can tolerate all the permanent faults, especially faults in the extra-circuitry for protection or voting, and consequently, there will always be some unidentified defective augmented routers which are going to transmit errors in an unpredictable manner. This study consists of tackling this fundamental problem. Specifically, we study and determine the average percentage of residual unidentified defective routers (UDRs) and their impact on the overall reliability of the NoC in light of self-protection strategies. Our study shows that TMR is the most efficient solution to limit the average percentage of UDRs when there are typically less than a 0.1 percent of defective base routers. However, TMR is also the most cost prohibitive and the least power efficient. Above 1% of defective base routers, the STAP approaches are more efficient although the protection efficiency decreases inexorably in the very defective technologies (e.g. when there is 10% or more of defective base routers). For instance, if the chip includes 10% of defective base routers, our study shows that there will remain on the average 1% of UDRs, which causes a major challenge for NoC reliability. Ahmed Louri, Jacques Henri Collet, Avinash Karanth |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2019 | Dynamic Voltage and Frequency Scaling in NoCs with Supervised and Reinforcement Learning TechniquesabstractNetwork-on-Chips (NoCs) are the de facto choice for designing the interconnect fabric in multicore chips due to their regularity, efficiency, simplicity, and scalability. However, NoC suffers from excessive static power and dynamic energy due to transistor leakage current and data movement between the cores and caches. Power consumption issues are only exacerbated by ever decreasing technology sizes. Dynamic Voltage and Frequency Scaling (DVFS) is one technique that seeks to reduce dynamic energy; however this often occurs at the expense of performance. In this paper, we propose LEAD Learning-enabled Energy-Aware Dynamic voltage/frequency scaling for multicore architectures using both supervised learning and reinforcement learning approaches. LEAD groups the router and its outgoing links into the same V/F domain and implements proactive DVFS mode management strategies that rely on offline trained machine learning models in order to provide optimal V/F mode selection between different voltage/frequency pairs. We present three supervised learning versions of LEAD that are based on buffer utilization, change in buffer utilization and change in energy/throughput, which allow proactive mode selection based on accurate prediction of future network parameters. We then describe a reinforcement learning approach to LEAD that optimizes the DVFS mode selection directly, obviating the need for label and threshold engineering. Simulation results using PARSEC and Splash-2 benchmarks on a 4 × 4 concentrated mesh architecture show that by using supervised learning LEAD can achieve an average dynamic energy savings of 15.4 percent for a loss in throughput of 0.8 percent with no significant impact on latency. When reinforcement learning is used, LEAD increases average dynamic energy savings to 20.3 percent at the cost of a 1.5 percent decrease in throughput and a 1.7 percent increase in latency. Overall, the more flexible reinforcement learning approach enables learning an optimal behavior for a wider range of load environments under any desired energy versus throughput tradeoff. Quintin Fettes, Mark Clark, Razvan C. Bunescu, Avinash Karanth, Ahmed Louri |
IEEE Trans. Computers | 4 |
| 2019 | Sustainability in Network-on-Chips by Exploring Heterogeneity in Emerging TechnologiesabstractWith the scaling of technology, the computing industry is experiencing a shift from multi-core to many-core architectures. However, traditional metallic-based on-chip interconnects may not scale to support many-core architectures due to high power dissipation, and increased communication latency. Attention has recently shifted to emerging technologies such as silicon-photonics and wireless interconnects to implement future on-chip communications. Although emerging technologies show promising results for power-efficient, low-latency, and scalable on-chip interconnects, the use of single technology may not be sufficient to scale future architectures. In this paper, we extend the heterogeneous architecture Optical-Wireless Network-on-Chip (OWN [1]) to Reconfigurable Optical-Wireless Network-on-Chip (R-OWN) by introducing run-time reconfigurable wireless channels. Like OWN, R-OWN is designed such that one-hop photonic interconnect is used up to 64 cores (called a cluster) and communication beyond a cluster is one-hop wireless to limit the network diameter to a maximum of three hops. The photonic bandwidth is efficiently shared using time division multiplexing (TDM) while the wireless bandwidth is shared using frequency division multiplexing (FDM). By exploiting the heterogeneity of two emerging technologies, we reduce the energy/bit, improve performance via reconfiguration, and thereby improve the sustainability of NoCs and CMPs. We propose a preliminary assessment of implementing heterogeneous technologies with the router microarchitecture. Further, we also discuss the design of horn antenna for implementing the wireless channels. Our results indicate that R-OWN improves the performance (throughput and latency) by 15 percent when compared to OWN while consuming 7 percent more energy than OWN. Further, OWN and R-OWN improve energy-efficiency by 54-61 percent when compared to WCube and CMesh architectures, respectively. It should be noted that both OWN and R-OWN require less area than state-of-the-art wired, wireless, and optical on-chip networks. Avinash Karanth, Savas Kaya, Md. Ashif I. Sikder, Daniel Carbaugh, Soumyasanta Laha, Dominic DiTomaso, Ahmed Louri, Hao Xin, JunQiang Wu |
IEEE Trans. Sustain. Comput. | 1 |
| 2018 | LEAD: learning-enabled energy-aware dynamic voltage/frequency scaling in NoCsabstractNetwork on Chips (NoCs) are the interconnect fabric of choice for multicore processors due to their superiority over traditional buses and crossbars in terms of scalability. While NoC's offer several advantages, they still suffer from high static and dynamic power consumption. Dynamic Voltage and Frequency Scaling (DVFS) is a popular technique that allows dynamic energy to be saved, but it can potentially lead to loss in throughput. In this paper, we propose LEAD - Learning-enabled Energy-Aware Dynamic voltage/frequency scaling for NoC architectures wherein we use machine learning techniques to enable energy-performance trade-offs at reduced overhead cost. LEAD enables a proactive energy management strategy that relies on an offline trained regression model and provides a wide variety of voltage/frequency pairs (modes). LEAD groups each router and the router's outgoing links locally into the same V/F domain, allowing energy management at a finer granularity without additional timing complications and overhead. Our simulation results using PARSEC and Splash-2 benchmarks on a 4 × 4 concentrated mesh architecture show an average dynamic energy savings of 17% with a minimal loss of 4% in throughput and no latency increase. Mark Clark, Avinash Karanth, Razvan C. Bunescu, Ahmed Louri |
DAC | 2 |
| 2018 | Extending the Power-Efficiency and Performance of Photonic Interconnects for Heterogeneous Multicores with Machine LearningabstractAs communication energy exceeds computation energy in future technologies, traditional on-chip electrical interconnects face fundamental challenges in the many-core era. Photonic interconnects have been proposed as a disruptive technology solution due to superior performance per Watt, distance independent energy consumption and CMOS compatibility for on-chip interconnects. Static power due to the laser being always switched on, varying link utilization due to spatial and temporal traffic fluctuations and thermal sensitivity are some of the critical challenges facing photonics interconnects. In this paper, we propose photonic interconnects for heterogeneous multicores using a checkerboard pattern that clusters CPU-GPU cores together and implements bandwidth reconfiguration using local router information without global coordination. To reduce the static power, we also propose a dynamic laser scaling technique that predicts the power level for the next epoch using the buffer occupancy of previous epoch. To further improve power-performance trade-offs, we also propose a regression-based machine learning technique for scaling the power of the photonic link. Our simulation results demonstrate a 34% performance improvement over a baseline electrical CMESH while consuming 25% less energy per bit when dynamically reallocating bandwidth. When dynamically scaling laser power, our buffer-based reactive and ML-based proactive prediction techniques show 40 - 65% in power savings with 0 - 14% in throughput loss depending on the reservation window size. Scott Van Winkle, Avinash Karanth, Razvan C. Bunescu, Ahmed Louri |
HPCA | 2 |
| 2018 | RETUNES: Reliable and Energy-Efficient Network-on-Chip ArchitectureabstractAs the number of cores integrated on the chip increases, the design of reliable and energy-efficient Network-on-Chip (NoC) to support the data movement needed by the multicores is becoming a critical challenge. Reliability of NoC is affected by several aging effects such as Hot carrier injection (HCI) and Negative Bias Temperature Instability (NBTI) which vary the threshold voltage of the transistor, causing timing errors. Dynamic Frequency and Voltage Scaling (DVFS) along with Near Threshold Voltage (NTV) scaling allows the transistor to operate close to the threshold voltage, thereby aggressively minimizing dynamic power consumption by reducing voltage/frequency and minimizing the threshold voltage variation mitigating aging process. However, the trade-off is increased latency and reduced reliability due to lower voltage margins. In this paper, we propose a unified approach called RETUNES: Reliable and Energy-Efficient NoC where NTV scaling and reliability are both achieved while improving performance. RETUNES is a five-level voltage/frequency scaling scheme, which decreases power consumption and threshold voltage variation (ΔVth) during low network load with higher reliability and increases the network performance during high network load with reduced reliability. In order to even out the wear out and minimize the impact of aging in NoC, we propose an adaptive routing algorithm in our design. RETUNES improves, total power savings by nearly 2.5× and the energy-delay product (EDP) of the NoC by 3× for Splash-2 and PARSEC benchmarks on a 4 × 4 concentrated mesh architecture. Padmaja Bhamidipati, Avinash Karanth |
ICCD | 2 |
| 2018 | Scalable Power-Efficient Kilo-Core Photonic-Wireless NoC ArchitecturesabstractAs technology scales, hundreds and thousands of cores are being integrated on a single-chip. Since metallic interconnects may not scale effectively to support thousands of cores, architects have proposed emerging technologies such as photonics and wireless for intra-chip communication. While photonics technology is limited by the complexity and thermal effects, wireless technology for on-chip communication is limited by the available bandwidth. In this paper, we combine the benefits of both technologies into novel architecture that takes advantage of the communication benefits of both technologies while circumventing their limits. We discuss the scalability of the proposed architecture to kilo-core system using wireless technology. We evaluate the power consumption, throughput and latency for 256 and 1024 core architectures when compared to photonics-only, wireless-wired, wireless-photonics and wired-only architectures on synthetic traffic traces. Our simulation results indicate that the proposed architecture and design methodology can have significant impact on the overall network power and performance. Avinash Karanth, Kyle Shifflet, Savas Kaya, Soumyasanta Laha, Ahmed Louri |
IPDPS | 1 |
| 2018 | Securing NoCs Against Timing Attacks with Non-Interference Based Adaptive RoutingabstractTiming channel attacks use interference from contending application flows to cause information leakage, and thereby either covertly transmit secrets, or create Denial-of-Service (DoS) attacks to undermine the on-chip hardware security. Protecting against timing channel attacks is very challenging since unseen vulnerabilities emerge in newer technology that can be cleverly exploited by malicious applications by intentionally gaming resources to artificially induce interference. In this paper, we propose to secure Network-on-Chips (NoCs) against timing attacks with non-Interference based adaptive routing where we efficiently separate network traffic to not only improve application performance and prevent information leakage. In our performance analysis, we show that the proposed architecture maintains non-interference between security domains, prevents DoS, and improves performance by 2-20% over existing techniques with only a 1.84% power consumption penalty. Travis Boraten, Avinash Karanth |
NOCS | 2 |
| 2018 | SHARP: Shared Heterogeneous Architecture with Reconfigurable Photonic Network-on-ChipabstractAs the relentless quest for higher throughput and lower energy cost continues in heterogenous multicores, there is a strong demand for energy-efficient and high-performance Network-on-Chip (NoC) architectures. Heterogeneous architectures that can simultaneously utilize both the serialized nature of the CPU as well as the thread level parallelism of the GPU are gaining traction in the industry. A critical issue with heterogeneous architectures is finding an optimal way to utilize the shared resources such as the last level cache and NoC without hindering the performance of either the CPU or the GPU core. Photonic interconnects are a disruptive technology solution that has the potential to increase the bandwidth, reduce latency, and improve energy-efficiency over traditional metallic interconnects. In this article, we propose a CPU-GPU heterogeneous architecture called Shared Heterogeneous Architecture with Reconfigurable Photonic Network-on-Chip (SHARP) that clusters CPU and GPU cores around the same router and dynamically allocates bandwidth between the CPU and GPU cores based on application demands. The SHARP architecture is designed as a Single-Writer Multiple-Reader (SWMR) crossbar with reservation-assist to connect CPU/GPU cores that dynamically reallocates bandwidth using buffer utilization information at runtime. As network traffic exhibits temporal and spatial fluctuations due to application behavior, SHARP can dynamically reallocate bandwidth and thereby adapt to application demands. SHARP demonstrates 34% performance (throughput) improvement over a baseline electrical CMESH while consuming 25% less energy per bit. Simulation results have also shown 6.9% to 14.9% performance improvement over other flavors of the proposed SHARP architecture without dynamic bandwidth allocation. Scott Van Winkle, Avinash Karanth |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | Mitigation of Hardware Trojan based Denial-of-Service attack for secure NoCs
Travis Boraten, Avinash Karanth |
J. Parallel Distributed Comput. | 2 |
| 2018 | Runtime Techniques to Mitigate Soft Errors in Network-on-Chip (NoC) ArchitecturesabstractAs aggressive scaling continues to push multiprocessor system-on-chips (MPSoCs) to new limits, complex hardware structures combined with stringent area and power constraints will continue to diminish reliability. Waning reliability in integrated circuits will increase the susceptibility of transient and permanent faults. There is an urgent demand for adaptive error correction coding (ECC) schemes in network-on-chips to provide fault tolerance and improve overall resiliency of MPSoC architectures. The goal of adaptive ECC schemes should be to maximize power savings when faults are infrequent and increase application speedup by boosting fault coverage when faults are frequent. In this paper, we propose runtime adaptive scrubbing (RAS), a novel multilayered error correction and detection scheme with three modes of operation enabled by an area-efficient configurable encoder for encoding packets on the switch-to-switch (s2s) layer, thus preventing faults from accumulating up the network stack and onto the end-to-end layer. As fault rates fluctuate we propose a dynamic methodology for improving fault localization and intelligently adapt fault coverage on demand to sustain graceful network degradation. RAS successfully improves network resiliency, fault localization, and fault coverage as compared to traditional static s2s schemes. Simulation results demonstrate that static RAS improves network speedup by 10% for Splash-2/PARSEC benchmarks on a $8 \times 8$ mesh network while reducing area overhead by 14% and incurring on an average 6.6% power penalty by boosting fault tolerance when fault rates increase. Further, our dynamic RAS scheme maintains 97.88% of network performance for real applications while incurring 20% power penalty. Travis Boraten, Avinash Karanth |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | GARUDA: Designing Energy-Efficient Hardware Monitors From High-Level Policies for Secure Information FlowabstractRuntime monitors detect vulnerabilities in embedded systems by running alongside untrusted software in order to detect violations of security policies as they occur, ideally with minimal overhead. Prior work has demonstrated language support for largely static security policies implemented using lattices and tag-based monitors. However, compiling high-level policies to modular hardware monitors that can implement a wide variety of security policies with minimal power has not been previously proposed. In this paper, we present a high-level security policy language, GARUDA, together with a compiler from GARUDA to Verilog, that enables the modular construction and composition of security hardware runtime monitors for a variety of security policies, including software fault isolation, secure control flow, and dynamic information flow via taint tracking. Unlike prior approaches in which the hardware monitors check all instructions, our hardware monitors are activated on-demand by the security policies which reduces the energy consumption. We perform experiments on Sniper, a full system multicore simulator, to evaluate the energy and performance tradeoffs of the security policies we have implemented so far. The policies are tested across a range of Splash-2 benchmarks. Seaghan Sefton, Taiman Siddiqui, Nathaniel St. Amour, Gordon Stewart 0001, Avinash Karanth |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Machine learning enabled power-aware Network-on-Chip designabstractAlthough Network-on-Chips (NoCs) are fast becoming pervasive as the interconnect fabric for multicore architectures and systems-on-chips, they still suffer from excessive static and dynamic power consumption. High dynamic power consumption results from switching and storing data within routers/links while excess static power is consumed when routers and links are not utilized for communication and yet have to be powered up. In this paper, we propose LESSON (Learning Enabled Sleepy Storage Links and Routers in NoCs) to reduce both static and dynamic power consumption by power-gating the links and routers at low network utilization and moving the data storage from within the routers to the links at high network utilization. As the network utilization increases from low-to-high, to accommodate more traffic, we design the same channels to flow traffic in either direction, thereby avoiding complex routing or look-ahead wake-up algorithms. Machine learning algorithms predict when to power-gate the channels and routers and when to increase the channel bandwidths such that power savings are maximized while performance penalty is minimized. Our results show that we can improve total network power consumption when compared to conventional NoC buffer designs by 85.6% and when compared with aggressive NoC buffer designs by 31.7%. Our predictor shows marginal performance penalties and by dynamically changing the direction of the links, we can improve packet latency by 14%. Dominic DiTomaso, Md. Ashif I. Sikder, Avinash Karanth, Ahmed Louri |
DATE | 3 |
| 2017 | CLAP-NET: Bandwidth adaptive optical crossbar architecture
Matthew Kennedy, Avinash Karanth |
J. Parallel Distributed Comput. | 2 |
| 2016 | Packet security with path sensitization for NoCs
Travis Boraten, Avinash Karanth |
DATE | 2 |
| 2016 | Secure Model Checkers for Network-on-Chip (NoC) ArchitecturesabstractAs chip multiprocessors (CMPs) are becoming more susceptible to process variation, crosstalk, and hard and soft errors, emerging threats from rogue employees in a compromised foundry are creating new vulnerabilities that could undermine the integrity of our chips with malicious alterations. As the Network-on-Chip (NoC) is a focal point of sensitive data transfer and critical device coordination, there is an urgent demand for secure and reliable communication. In this paper we propose Secure Model Checkers (SMCs), a real-time solution for control logic verification and functional correctness in the micro-architecture to detect Hardware Trojan (HT) induced denial-of-service attacks and improve reliability. In our evaluation, we show that SMCs provides significant security enhancements in real-time with only 1.5% power and 1.1% area overhead penalty in the micro-architecture. Travis Boraten, Dominic DiTomaso, Avinash Karanth |
ACM Great Lakes Symposium on VLSI | 3 |
| 2016 | Mitigation of Denial of Service Attack with Hardware Trojans in NoC ArchitecturesabstractAs Multiprocessor System-on-Chips (MPSoCs) continue to scale, security for Network-on-Chips (NoCs) is a growing concern as rogue agents threaten to infringe on the hardware's trust and maliciously implant Hardware Trojans (HTs) to undermine their reliability. The trustworthiness of MPSoCs will rely on our ability to detect Denial-of-Service (DoS) threats posed by the HTs and mitigate HTs in a compromised NoC to permit graceful network degradation. In this paper, we propose a new light-weight target-activated sequential payload (TASP) HT model that performs packet inspection and injects faults to create a new type of DoS attack. Faults injected are used to trigger a response from error correction code (ECC) schemes and cause repeated retransmission to starve network resources and create deadlocks capable of rendering single-application to full chip failures. To circumvent the threat of HTs, we propose a heuristic threat detection model to classify faults and discover HTs within compromised links. To prevent further disruption, we propose several switch-to-switch link obfuscation methods to avoid triggering of HTs in an effort to continue using links instead of rerouting packets with minimal overhead (1-3 cycles). Our proposed modifications complement existing fault detection and obfuscation methods and only adds 2% in area overhead and 6% in excess power consumption in the NoC micro-architecture. Travis Boraten, Avinash Karanth |
IPDPS | 2 |
| 2016 | Dynamic error mitigation in NoCs using intelligent prediction techniquesabstractNetwork-on-chips (NoCs) are quickly becoming the standard communication fabric for multi-core systems. As technology continues to scale down into the nanometer regime, device behavior will become increasingly unreliable due to a combination of aging, soft errors, aggressive transistor design, and process-voltage-temperature variations. Further, stringent timing constraints in NoCs are designed so that data can be pushed faster. The net result is an increase in errors which must be mitigated by the NoC. Typical techniques for handling faults are often reactive as they respond to faults after the error has occurred, making the recovery process inefficient in energy and time. In this paper, we take a different approach wherein we propose to use proactive, fault-tolerant schemes to be employed before the fault affects the system. We propose to utilize machine learning techniques to train a decision tree which can be used to predict faults efficiently in the network. Based on the prediction model, we dynamically mitigate these predicted faults through error correction codes (ECC) and relaxed timing transmission. Our results indicate that, on average, we can accurately predict timing errors 60.6% better than a static single error correction and double error detection (SECDED) technique resulting in an average 26.8% reduction in retransmitted packets, a average net speedup of 3.31 x, and an average energy savings of 60.0% over other designs for real traffic patterns. Dominic DiTomaso, Travis Boraten, Avinash Karanth, Ahmed Louri |
MICRO | 3 |
| 2015 | Dynamic Power Reduction Techniques in On-Chip Photonic InterconnectsabstractPhotonic interconnects is a disruptive technology solution that can overcome the power and bandwidth limitations of traditional electrical Network-on-Chips (NoCs). However, the static power dissipated in the external laser may limit the performance of future optical NoCs by dominating the stringent network power budget. From the analysis of real benchmarks for multi-cores, it is observed that high static power is consumed due to the external laser even for low channel utilization. In this paper, we propose runtime power management techniques to reduce the magnitude of laser power consumption by tuning the network in response to actual application characteristics. We scale the number of channels available for communication based on link and buffer utilization. The performance on synthetic and real traffic (PARSEC, Splash-2) for 64-cores indicate that our proposed power scaling technique can reduce optical power by about 70% with less than 1% throughput penalty for real traffic. Brian Neel, Matthew Kennedy, Avinash Karanth |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | Resilient and Power-Efficient Multi-Function Channel Buffers in Network-on-Chip ArchitecturesabstractNetwork-on-Chips (NoCs) are quickly becoming the standard communication paradigm for the growing number of cores on the chip. While NoCs can deliver sufficient bandwidth and enhance scalability, NoCs suffer from high power consumption due to the router microarchitecture and communication channels that facilitate inter-core communication. As technology keeps scaling down in the nanometer regime, unpredictable device behavior due to aging, infant mortality, design defects, soft errors, aggressive design, and process-voltage-temperature variations, will increase and will result in a significant increase in faults (both permanent and transient) and hardware failures. In this paper, we propose QORE-a fault tolerant NoC architecture with Multi-Function Channel (MFC) buffers. The use of MFC buffers and their associated control (link and fault controllers) enhance fault-tolerance by allowing the NoC to dynamically adapt to faults at the link level and reverse propagation direction to avoid faulty links. Additionally, MFC buffers reduce router power and improve performance by eliminating in-router buffering. We utilize a machine learning technique in our link controllers to predict the direction of traffic flow in order to more efficiently reverse links. Our simulation results using real benchmarks and synthetic traffic mixes show that QORE improves speedup by 1.3x and throughput by 2.3x when compared to state-of-the art fault tolerant NoCs designs such as Ariadne and Vicis. Moreover, using Synopsys Design Compiler, we also show that network power in QORE is reduced by 21 percent with minimal control overhead. Dominic DiTomaso, Avinash Karanth, Ahmed Louri, Razvan C. Bunescu |
IEEE Trans. Computers | 2 |
| 2015 | A New Frontier in Ultralow Power Wireless Links: Network-on-Chip and Chip-to-Chip InterconnectsabstractThis paper explores the general framework and prospects for on-chip and off-chip wireless interconnects implemented for high-performance computing (HPC) systems in the context of micro power wireless design. HPC interconnects demand very high (≥ 10 Gb/s) transmission rates using ultraefficient (~ 1 pJ/bit) transceivers over extremely short (≤ 100 cm) ranges. In an attempt to design such wireless interconnects, first a model for the wireless communication channel properties is developed. The use of CMOS-based energyefficient on-off keying (OOK) transceiver architectures operating in the 60-90 GHz bands is considered as a practical solution. In order to address strict performance requirements of wireless HPC interconnects, and taking advantage of the recent developments in device scaling, compact low-power and innovative circuits based on novel double-gate MOSFETs (DG-MOSFETs) are proposed in the implementation of the architecture. The performance of a compact low-noise amplifier (LNA) design using common source (CS) inductive degeneration with 32 nm DGMOSFETs is investigated by quantitative analysis and simulation. The proposed inductor-less two-stage cascode cascade LNA is optimized for 90 GHz operation and has the advantage of gain switching over its CMOS counterpart without the use of additional switching transistors, which makes it remarkably power efficient and faster. As further examples of efficient and compact DG-MOSFET circuits for OOK transceiver design, a three-stage CS 5 dB tunable power amplifier operating up to 90 GHz, and a novel 90 GHz voltage controlled oscillator are also presented. This is followed by the proposal of an array of four monopole antennas studied using full-wave EM solver. Soumyasanta Laha, Savas Kaya, David W. Matolak, William Rayess, Dominic DiTomaso, Avinash Karanth |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2015 | A-WiNoC: Adaptive Wireless Network-on-Chip Architecture for Chip MultiprocessorsabstractWith the rise of chip multiprocessors, an energy-efficient communication fabric is required to satisfy the data rate requirements of future multi-core systems. The Network-on-Chip (NoC) paradigm is fast becoming the standard communication infrastructure to provide scalable inter-core communication. However, research has shown that metallic interconnects cause high latency and consume excess energy in NoC architectures. Emerging technologies such as on-chip wireless interconnects can alleviate the power and bandwidth problems of traditional metallic NoCs. In this paper, we propose A-WiNoC, a scalable, adaptable wireless Network-on-Chip architecture that uses energy efficient wireless transceivers and improves network throughput by dynamically re-assigning channels in response to bandwidth demands from different cores. To implement such adaptability in our network at run-time, we propose an adaptable algorithm that works in the background along with a token sharing scheme to fully utilize the wireless bandwidth efficiently. Since no wireless NoC design has been completely realized with current technology, we describe technology trends in designing energy-efficient wireless transceivers with emerging technologies. We compare our proposed A-WiNoC to both wireless and wired topologies at 64 cores, with results showing a 1.4-2.6× speedup on real applications and a 54 percent improvement in throughput for synthetic traffic. Using Synopsys Design Compiler, our results indicate that A-WiNoC saves 25-35 percent energy over other state-of-the-art networks. We show that A-WiNoC can scale to 256 cores with an energy improvement of 21 percent and a saturation throughput increase of approximately 37 percent. Dominic DiTomaso, Avinash Karanth, David W. Matolak, Savas Kaya, Soumyasanta Laha, William Rayess |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | QORE: A fault tolerant network-on-chip architecture with power-efficient quad-function channel (QFC) buffersabstractNetwork-on-Chips (NoCs) are quickly becoming the standard communication paradigm for the growing number of cores on the chip. While NoCs can deliver sufficient bandwidth and enhance scalability, NoCs suffer from high power consumption due to the router microarchitecture and communication channels that facilitate inter-core communication. As technology keeps scaling down in the nanometer regime, unpredictable device behavior due to aging, infant mortality, design defects, soft errors, aggressive design, and process-voltage-temperature variations, will increase and will result in a significant increase in faults (both permanent and transient) and hardware failures. In this paper, we propose QORE - a fault tolerant NoC architecture with Quad-Function Channel (QFC) buffers. The use of QFC buffers and their associated control (link and fault controllers) enhance fault-tolerance by allowing the NoC to dynamically adapt to faults at the link level and reverse propagation direction to avoid faulty links. Additionally, QFC buffers reduce router power and improve performance by eliminating in-router buffering. Our simulation results using real benchmarks and synthetic traffic mixes show that QORE improves speedup by 1.3× and throughput by 2.3× when compared to state-of-the art fault tolerant NoCs designs such as Ariadne and Vicis. Moreover, using Synopsys Design Compiler, we also show that network power in QORE is reduced by 21% with minimal control overhead. Dominic DiTomaso, Avinash Karanth, Ahmed Louri |
HPCA | 2 |
| 2014 | Workload assignment considering NBTI degradation in multicore systemsabstractWith continuously shrinking technology, reliability issues such as Negative Bias Temperature Instability (NBTI) has resulted in considerable degradation of device performance, and eventually the short mean-time-to-failure (MTTF) of the whole multicore system. This article proposes a new workload balancing scheme based on device-level fractional NBTI model to balance the workload among active cores while relaxing stressed ones. Starting with NBTI-induced threshold voltage degradation, we define a concept of Capacity Rate (CR) as an indication of one core's ability to accept workload. Capacity rate captures core's performance variability in terms of delay and power metrics under the impact of NBTI aging. The proposed workload balancing framework employs the capacity rates as workload constraints, applies a Dynamic Zoning (DZ) algorithm to group cores into zones to process task flows, and then uses Dynamic Task Scheduling (DTS) to allocate tasks in each zone with balanced workload and minimum communication cost. Experimental results on a 64-core system show that by allowing a small part of the cores to relax over a short time period, the proposed methodology improves multicore system yield (percentage of core failures) by 20%, while extending MTTF by 30% with insignificant degradation in performance (less than 3%). Jin Sun 0006, Roman L. Lysecky, Karthik Shankar, Avinash Karanth, Ahmed Louri, Janet Roveda |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2014 | Three-Dimensional Stacked Nanophotonic Network-on-Chip Architecture with Minimal ReconfigurationabstractAs throughput, scalability, and energy efficiency in network-on-chips (NoCs) are becoming critical, there is a growing impetus to explore emerging technologies for implementing NoCs in future multicore and many-core architectures. Two disruptive technologies on the horizon are nanophotonic interconnects (NIs) and 3D stacking. NIs can deliver high on-chip bandwidth while delivering low energy/bit, thereby providing a reasonable performance-per-watt in the future. Three-dimensional stacking can reduce the interconnect distance and increase the bandwidth density by incorporating multiple communication layers. In this paper, we propose an architecture that combines NIs and 3D stacking to design an energy-efficient and reconfigurable NoC. We quantitatively compare the hardware complexity of the proposed topology to other nanophotonic networks in terms of hop count, network diameter, radix, and photonic parameters. To maximize performance, we also propose an efficient reconfiguration algorithm that dynamically reallocates channel bandwidth by adapting to traffic fluctuations. For 64-core reconfigured network, our simulation results indicate that the execution time can be reduced up to 25 percent for Splash-2, PARSEC, and SPEC CPU2006 benchmarks. Moreover, for a 256-core version of the proposed architecture, our simulation results indicate a throughput improvement of more than 25 percent and energy savings of 23 percent on synthetic traffic when compared to competitive on-chip electrical and optical networks. Randy Morris, Avinash Karanth, Ahmed Louri, Ralph D. Whaley |
IEEE Trans. Computers | 2 |
| 2014 | Extending the Performance and Energy-Efficiency of Shared Memory Multicores with Nanophotonic TechnologyabstractAs the number of cores increases exponentially on a single chip, the design and integration of both the on-chip network facilitating intercore communication, and the cache coherence protocol for enabling shared memory programming have become critical for improved energy-efficiency and overall chip performance. With traditional metal interconnects facing stringent energy constraints, researchers are currently pursuing disruptive solutions such as nanophotonics for improved energy-efficiency. Cache coherence in multicores can be enforced effectively by snoopy protocols; however, broadcasting every cache miss can limit the scalability while consuming excess energy. In this paper, we propose PULSE, a nanophotonic broadcast tree-based network for snoopy cache coherent multicores. To limit the energy-penalty from broadcasting (and thereby splitting) optical signals, we direct the optical signal from the external laser such that only the subset of requesters can receive the optical signal. Furthermore, as cache blocks are shared by a few cores, we propose a multicast version of PULSE called multi-PULSE that predicts the sharers' for each L2 miss and morphing the broadcast to a multicast network. We evaluate the energy and performance using CACTI and SIMICS on 16-core and 64-core versions of PULSE and multi-PULSE for Splash-2, PARSEC, and SPEC CPU2006 benchmarks and compare to electrical networks, optical networks, and another cache filtering techniques. Our results indicate that PULSE outperforms competitive electrical/optical networks by 60 percent in terms of execution time, and multi-PULSE reduces average energy from 10 to 80 percent even with a few mispredictions. Randy Morris, Evan Jolley, Avinash Karanth |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Energy-efficient Runtime Adaptive Scrubbing in fault-tolerant Network-on-Chips (NoCs) architecturesabstractAs Networks-on-Chips (NoCs) continue to become more susceptible to process variation, cross-talk, hard and soft errors with technology scaling to sub-nanometer, there is an urgent need for adaptive Error Correction Coding (ECC) schemes for improving the resiliency of the system. The goal of adaptive ECC schemes should be two fold; decrease power consumption when errors are infrequent, thereby maximizing power savings and increase the fault coverage when errors are frequent, thereby improving application speedup while consuming more power. In this paper, we propose Runtime Adaptive Scrubbing (RAS), a novel multi-layered error correction and detection scheme for Networks-on-Chips (NoCs) architectures that intelligently adjusts fault coverage at the physical layer using variable strength encoders to scrub (protect) flits, thereby preventing faults from accumulating and propagating up to the logical layer. RAS successfully permits graceful network degradation while improving the overall network speedup, fault granularity, and wider fault coverage than traditional static schemes. Simulation results indicate that RAS improves network latency by an average of 10% for Splash-2/PARSEC benchmarks on a 8 × 8 mesh network while incurring 6.6% power penalty per flit and saving 15% in area overhead. Travis Boraten, Avinash Karanth |
ICCD | 2 |
| 2013 | Energy-efficient adaptive wireless NoCs architectureabstractWith the increasing number of cores in chip multiprocessors, the design of an efficient communication fabric is essential to satisfy the bandwidth and energy requirements of multi-core systems. Scalable Network-on-Chip (NoC) designs are quickly becoming the standard communication framework to replace bus-based networks. However, the conventional metallic interconnects for inter-core communication consume excess energy and lower throughput which are major bottlenecks in NoC architectures. On-chip wireless interconnects can alleviate the power and bandwidth problems of traditional metallic NoCs. In this paper, we propose an adaptable wireless Network-on-Chip architecture (A-WiNoC) that uses adaptable and energy efficient wireless transceivers to improve network power and throughput by adapting channels according to traffic patterns. Our adaptable algorithm uses link utilization statistics to re-allocate wireless channels and a token sharing scheme to fully utilize the wireless bandwidth efficiently. We compare our proposed A-WiNoC to both wireless/electrical topologies with results showing a throughput improvement of 65%, a speedup between 1.4-2.6X on real benchmarks, and an energy savings of 25-35%. Dominic DiTomaso, Avinash Karanth, David W. Matolak, Savas Kaya, Soumyasanta Laha, William Rayess |
NOCS | 2 |
| 2013 | PROBE: Prediction-based optical bandwidth scaling for energy-efficient NoCsabstractOptical interconnect is a disruptive technology solution that can overcome the power and bandwidth limitations of traditional electrical Networks-on-Chip (NoCs). However, the static power dissipated in the external laser may limit the performance of future optical NoCs by dominating the stringent network power budget. From the analysis of real benchmarks for multicores, it is observed that high static power is consumed due to the external laser even for low channel utilization. In this paper, we propose PROBE: Prediction-based Optical Bandwidth Scaling for Energy-efficient NoCs by exploiting the latency/bandwidth trade-off to reduce the static power consumption by increasing the average channel utilization. With a lightweight prediction technique, we scale the bandwidth adaptively to the changing traffic demands while maintaining reasonable performance. The performance on synthetic and real traffic (PARSEC, Splash2) for 64-cores indicate that our proposed bandwidth scaling technique can reduce optical power by about 60% with at most 11% throughput penalty. Avinash Karanth |
NOCS | 2 |
| 2013 | Extending the Energy Efficiency and Performance With Channel Buffers, Crossbars, and Topology Analysis for Network-on-ChipsabstractNetwork-on-chips (NoCs) have emerged as a scalable solution to the wire delay constraints, thereby providing a high-performance communication fabric for future multicores. Research has shown that power, area, and performance of the NoC architecture are tightly integrated with the design and optimization of the link, router (buffer and crossbar), and topology. Recent work has shown that adaptive channel buffers (on-link storage) can considerably reduce power consumption and area overhead by reducing or replacing the power-hungry router buffers. However, channel buffer design can lead to head-of-line (HoL) blocking, which eventually reduces the throughput of the network. In this paper, we design channel buffers and router crossbars to improve the performance (latency, throughput) while reducing the power consumption. In addition, we implement the proposed channel buffers and crossbar organizations in a concentrated torus (CTorus) topology which is a dual network without the additional area overhead. We compare other dual networks with leading topologies such as mesh2X, concentrated mesh2X (CMesh2X), and flattened butterfly2X (FBfly2X), each implemented with channel buffers. Our proposed designs analyze the power-performance-area tradeoff in designing channel buffers for NoC architectures while alleviating HoL blocking through buffer organizations and crossbar optimizations. Results using Synopsys design compiler showed that the buffer and crossbar organizations for an 8 × 8 mesh architecture can reduce power consumption by 25%-40%, improve throughput and reduce latency by 525%, while occupying 4%-13% more area when compared to the baseline architecture for both synthetic as well as real benchmark traces such as Princeton Application Repository for Shared-Memory Computers (PARSEC) and Standard Performance Evaluation Corporation (SPEC). CPU2006. When the energy-efficient buffer and crossbar organization was inserted into our CTorus topology, we further reduced energy dissipation by 32% and area by 53%, on average, over mesh2X, CMesh2X, and FBfly2X. Dominic DiTomaso, Randy Morris, Avinash Karanth, Ashwini Sarathy, Ahmed Louri |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | 3D-NoC: Reconfigurable 3D photonic on-chip interconnect for multicoresabstractThe power dissipation of metallic interconnects in future multicore architectures is projected to be a major bottleneck as we scale to sub-nanometer regime. This has motivated researchers to develop alternate power-efficient technology solutions to the performance limitations of future multicores. Nanophotonic interconnects (NIs) is a disruptive technology solution that is capable of delivering the communication bandwidth at low power dissipation when the number of cores is scaled to large numbers. Similarly, 3D stacking is another interconnect technology solution that can lead to low energy/bit for communication. In this paper, we propose to combine NIs with with 3D stacking to develop a scalable, reconfigurable, power-efficient and high-performance interconnect for future many-core systems called 3D-NoC. We propose to develop a multi-layer NIs that can dynamically reconfigure without system intervention and allocate channel bandwidth from less utilized links to more utilized communication links. Our simulation results indicate that the performance can be further improved by 10%-25% for Splash-2, PARSEC and SPEC CPU2006 benchmarks. Randy Morris, Avinash Karanth, Ahmed Louri |
ICCD | 2 |
| 2012 | Dynamic Reconfiguration of 3D Photonic Networks-on-Chip for Maximizing Performance and Improving Fault ToleranceabstractAs power dissipation in future Networks-on-Chips (NoCs) is projected to be a major bottleneck, researchers are actively engaged in developing alternate power-efficient technology solutions. Photonic interconnects is a disruptive technology solution that is capable of delivering the communication bandwidth at low power dissipation when the number of cores is scaled to large numbers. Similarly, 3D stacking is another interconnect technology solution that can lead to low energy/bit for communication. In this paper, we propose to combine photonic interconnects with 3D stacking to develop a scalable, reconfigurable, power-efficient and high-performance interconnect for future many-core systems, called R-3PO (Reconfigurable 3D-Photonic Networks-on-Chip). We propose to develop a multi-layer photonic interconnect that can dynamically reconfigure without system intervention and allocate channel bandwidth from less utilized links to more utilized communication links. In addition to improving performance, reconfiguration can re-allocate bandwidth around faulty channels, thereby increasing the resiliency of the architecture and gracefully degrading performance. For 64-core reconfigured network, our simulation results indicate that the performance can be further improved by 10%-25% for Splash-2, PARSEC and SPEC CPU2006 benchmarks, where as simulation results for 256-core chip indicate a performance improvement of more than 25% while saving 6%-36% energy when compared to state-of-the-art on-chip electrical and optical networks. Randy Morris, Avinash Karanth, Ahmed Louri |
MICRO | 2 |
| 2011 | Co-design of channel buffers and crossbar organizations in NoCs architecturesabstractNetwork-on-Chips (NoCs) have emerged as a scalable solution to the wire delay constraints, thereby providing a high-performance communication fabric for future multicores. Research has shown that power, area and performance of Network-on-Chips (NoCs) architecture are tightly integrated with the design and optimization of the link and router (buffer and crossbar). Recent work has shown that adaptive channel buffers (on-link storage) can considerably reduce power consumption and area overhead by reducing or replacing the power hungry router buffers. However, channel buffer design can lead to Head-of-Line (HoL) blocking which eventually reduces the throughput of the network. In this paper, we explore the design space of organizing channel buffers and router crossbars to improve the performance (latency, throughput) while reducing the power consumption. Our proposed designs analyze the power-performance-area trade-off in designing channel buffers for NoC architectures while overcoming HoL blocking through crossbar optimizations. Our simulation and NoC design synthesis shows that for a 8 × 8 mesh architecture, we can reduce the power consumption by 25-40%, improve performance by 10-25% while occupying 4-13% more area when compared to the baseline architecture. Avinash Karanth, Randy Morris, Dominic DiTomaso, Ashwini Sarathy, Ahmed Louri |
ICCAD | 1 |
| 2011 | Introduction to the special issue on Networks-on-Chip (NoC) of the Journal of Parallel and Distributed Computing (JPDC)
Ahmed Louri, Avinash Karanth |
J. Parallel Distributed Comput. | 2 |
| 2010 | Workload capacity considering NBTI degradation in multi-core systemsabstractAs device feature sizes continue to shrink, long-term reliability such as Negative Bias Temperature Instability (NBTI) leads to low yields and short mean-time-to-failure (MTTF) in multi-core systems. This paper proposes a new workload balancing scheme based on device level fractional NBTI model to balance the workload among active cores while relaxing stressed ones. The proposed method employs the Capacity Rate (CR) provided by the NBTI model, applies Dynamic Zoning (DZ) algorithm to group cores into zones to process task flows, and then uses Dynamic Task Scheduling (DTS) to allocate tasks in each zone with balanced workload and minimum communication cost. Experimental results on 64-core system show that by allowing a small part of the cores to relax over a short time period (10 seconds), the proposed methodology improves multi-core system yield (percentage of core failures) by 20%, while extending MTTF by 30% with insignificant degradation in performance (less than 3%). Jin Sun 0006, Roman L. Lysecky, Karthik Shankar, Avinash Karanth, Ahmed Louri, Janet Roveda |
ASP-DAC | 4 |
| 2010 | Power-Efficient and High-Performance Multi-level Hybrid Nanophotonic Interconnect for MulticoresabstractNetwork-on-Chips (NoCs) are becoming the defacto standard for interconnecting the increasing number of cores in chip multiprocessors (CMPs) by overcoming the scalability and wire delay problems of shared buses. However, recent research has shown that future NoCs will be limited by power dissipation and reduced performance forcing architects to explore other technologies that are complementary metal oxide semiconductor (CMOS) compatible. In this paper, we propose ET-PROPEL (Extended Token based Photonic Reconfigurable On-Chip Power and Area-Efficient Links) architecture to utilize the emerging nanophotonic technology to design a high-bandwidth, low latency and low power multi-level hybrid interconnect that balances cost and performance. We develop our interconnect at three levels: at the first level (x) we design a fully connected network for exploiting locality; at the second level(y), we design a shared channel using optical tokens to reduce power while providing full connectivity and at the third level (z), we propose a novel nanophotonic crossbar that provides scalable bisection bandwidth. The first two levels are combined into T-PROPEL(token-PROPEL, 64 cores) and four separate T-PROPELs are combined into ET-PROPEL (256 cores). We have simulated both T-PROPEL and ET-PROPEL using synthetic and SPLASH-2 traffic, where our results indicate that T-PROPEL and ET-PROPEL significantly reduce power(10-fold) and increase performance (3-fold) over other well known electrical and photonic networks. Randy Wayne Morris Jr., Avinash Karanth |
NOCS | 2 |
| 2010 | Special Issue on Network-on-Chips (NoCs)
Ahmed Louri, Avinash Karanth |
J. Parallel Distributed Comput. | 2 |
| 2009 | Design of a scalable nanophotonic interconnect for future multicoresabstractAs communication-centric computing paradigm gathers momentum due to increased wire delays and excess power dissipation with technology scaling, researchers have focused their attention on developing alternate technology solutions for Network-on-Chips (NoCs) architectures. One potential solution is nanophotonics because of higher bandwidth, reduced power dissipation and increased wiring simplification. In this paper, we propose PROPEL, a balanced power and area-efficient on-chip photonic interconnect for future multicores. PROPEL overcomes two fundamental issues facing NoCs architectures, namely power dissipation and area overhead, by a combination of multiplexing techniques (wave-length and space) and by exploiting the recent advances in optical component design space. We also propose a scalable version of PROPEL, called E-PROPEL which can scale to 256 cores. Our results indicate that PROPEL and E-PROPEL are power, cost and area-effective networks when compared to competing on-chip optical topologies when the number of optical components and overall power loss in the network are considered. Simulation results on synthetic traffic indicate that PROPEL performs better (throughput and power) than electrical and optical topologies. Avinash Karanth, Randy Morris |
ANCS | 1 |
| 2009 | Adaptive inter-router links for low-power, area-efficient and reliable Network-on-Chip (NoC) architecturesabstractThe increasing wire delay constraints in deep sub-micron VLSI designs have led to the emergence of scalable and modular network-on-chip (NoC) architectures. As the power consumption, area overhead and performance of the entire NoC is influenced by the router buffers, research efforts have targeted optimized router buffer design. In this paper, we propose iDEAL - inter-router, dual-function energy and area-efficient links capable of data transmission as well as data storage when required. iDEAL enables a reduction in the router buffer size by controlling the repeaters along the links to adaptively function as link buffers during congestion, thereby achieving nearly 30% savings in overall network power and 35% reduction in area with only a marginal 1 - 3% drop in performance. In addition, aggressive speculative flow control further improves the performance of iDEAL. Moreover, the significant reduction in power consumption and area provides sufficient headroom for monitoring negative bias temperature instability (NBTI) effects in order to improve circuit reliability at reduced feature sizes. Avinash Karanth, Ashwini Sarathy, Ahmed Louri, Janet Roveda |
ASP-DAC | 1 |
| 2009 | On-Chip photonic interconnects for scalable multi-core architecturesabstractIn this paper, we propose PROPEL, a photonic network-on-chip (NoC) that improves performance and power with energy-efficient opto-electronic components for future chip multiprocessors (CMPs). Our analytical and simulation results indicate that PROPEL improves throughput and reduces power over optical and electrical networks for various traffic traces while requiring fewer photonic components and devices. Avinash Karanth, Randy Morris, Ahmed Louri |
NOCS | 1 |
| 2008 | iDEAL: Inter-router Dual-Function Energy and Area-Efficient Links for Network-on-Chip (NoC) ArchitecturesabstractNetwork-on-Chip (NoC) architectures have been adopted by a growing number of multi-core designs as a flexible and scalable solution to the increasing wire delay constraints in the deep sub-micron regime. However, the shrinking feature size limits the performance of NoCs due to power and area constraints. Research into the optimization of NoCs has shown that a reduction in the number of buffers in the NoC routers reduces the power and area overhead but degrades the network performance. In this paper, we propose iDEAL, a low-power area-efficient NoC architecture by reducing the number of buffers within the router. To overcome the performance degradation caused by the reduced buffer size, we propose to use adaptive dual-function links capable of data transmission as well as data storage when required. Simulation results for the proposed architecture show that reducing the router buffer size in half and using the adaptive dual-function links achieves nearly 40% savings in buffer power, 30% savings in overall network power and about 41% savings in the router area, with only a marginal 1-3% drop in performance. Moreover, the performance in iDEAL can be further improved by aggressive and speculative flow control techniques. Avinash Karanth, Ashwini Sarathy, Ahmed Louri |
ISCA | 1 |
| 2008 | Adaptive Channel Buffers in On-Chip Interconnection Networks - A Power and Performance AnalysisabstractOn-chip interconnection networks (OCINs) have emerged as a modular and scalable solution for wire delay constraints in deep submicron VLSI design. OCIN research has shown that the design of buffers in the router influences the energy consumption, area overhead, and overall performance of the network. In this paper, we propose a low-power low-area OCIN architecture by reducing the number of buffers within the router. To minimize the performance degradation due to the reduced buffer size, we use the already existing repeaters along the inter-router channels to double as buffers along the channel when required. At low network loads, the proposed adaptive channel buffers function as conventional repeaters, propagating the signals. At high network loads, the adaptive channel buffers function as storage elements in addition to the router buffers. The router buffers can be assigned either statically or dynamically to the incoming packets. Static allocation reserves equal buffer space partitioned among all of the incoming packets, whereas dynamic allocation reserves buffer space on a per-flit basis, enabling higher buffer occupancy. We evaluate the proposed adaptive channel buffers with both static and dynamic buffer allocation policies in the 90-nm technology node, using 8 times 8 mesh and folded torus network topologies. Simulation results using the SPLASH-2 suite benchmarks and synthetic traffic patterns show that, by reducing the router buffer size, our proposed architecture achieves nearly 40 percent savings in router buffer power, 30 percent savings in overall network power, and 41 percent savings in area, with only a marginal 1-5 percent drop in throughput under dynamic buffer allocation and about 10-20 percent drop in throughput for statically assigned buffers. Avinash Karanth, Ashwini Sarathy, Ahmed Louri |
IEEE Trans. Computers | 1 |
| 2007 | Design of adaptive communication channel buffers for low-power area-efficient network-on-chip architectureabstractNetwork-on-Chip (NoC)architectures provide a scalable solution to the wire delay constraints in deep submicron VLSI designs. Recent research into the ptimization of NoC architectures has shown that the design of buffers in the NoC routers influences the power consumption, area overhead and performance of the entire network. In this paper, we propose a low-power area-efficient NoC architecture by reducing the number of router buffers. As a reduction in the number of buffers degrades the network's performance, we propose to use the existing repeaters along the inter-router links as adaptive channel buffers for storing data when required. We evaluate the proposed adaptive communication channel buffers under static and dynamic buffer allocation in 8 x 8 mesh and folded torus network topologies. Simulation results show that reducing the router buffer size in half and using the adaptive channel buffers reduces the buffer power by 40-52% and leads to a 17-20% savings in overall network power with a 50% reduction in router area. The design with dynamic buffer allocation shows a marginal 1-5% drop in performance, while static buffer allocation shows a 10-20% drop in performance, for various traffic patterns. Avinash Karanth, Ashwini Sarathy, Ahmed Louri |
ANCS | 1 |
| 2007 | Power-Aware Bandwidth-Reconfigurable Optical Interconnects for High-Performance Computing (HPC) SystemsabstractAs communication distances and bit rates increase, opto-electronic interconnects are becoming de-facto standard/or designing high-bandwidth low-latency interconnection networks for high performance computing (HPC) systems. While bandwidth scaling with efficient multiplexing techniques (wavelengths, time and space) are available, static assignment of wavelengths can be detrimental to network performance for adversial traffic patterns. Dynamic bandwidth reconfiguration based on actual traffic pattern can lead to improved network performance by utilizing idle resources. While dynamic bandwidth re-allocation (DBR) techniques can alleviate interconnection bottlenecks, power consumption also increases considerably. In this paper, we propose a dynamically re configurable architecture called E-RAPID (extended-reconfigurable, all-photonic interconnect for distributed and parallel systems) that not only dynamically reallocates bandwidth, but also reduces the power consumption for all traffic patterns. Our proposed LS (lock-step) reconfiguration technique combines dynamic power management (DPM) with DBR techniques, achieving a reduction in power consumption of 25%-50% while degrading the throughput by less than 5%. Avinash Karanth, Ahmed Louri |
IPDPS | 1 |
| 2007 | Performance adaptive power-aware reconfigurable optical interconnects for high-performance computing (HPC) systemsabstractAs communication distances and bit rates increase, optoelectronic interconnects are being deployed for designing high-bandwidth low-latency interconnection networks for high performance computing (HPC) systems. While bandwidth scaling with efficient multiplexing techniques (wavelengths, time and space) are available, static assignment of wavelengths can be detrimental to network performance for non-uniform (adversial) workloads. Dynamic bandwidth re-allocation based on actual traffic pattern can lead to improved network performance by utilizing idle resources. While dynamic bandwidth re-allocation (DBR) techniques can alleviate interconnection bottlenecks, power consumption also increases considerably. In this paper, we propose to improve the performance of optical interconnects using DBR techniques and simultaneously optimize the power consumption using Dynamic Power Management (DPM) techniques. DBR, re-allocates idle channels to busy channels (wavelengths) for improving throughput and DPM regulates the bit rates and supply voltages for the individual channels. A reconfigurable opto-electronic architecture and a performance adaptive algorithm for implementing DBR and DPM are proposed in this paper. Our proposed reconfiguration algorithm achieves a significant reduction in power consumption and considerable improvement in throughput with a marginal increase in latency for various traffic patterns. Avinash Karanth, Ahmed Louri |
SC | 1 |
| 2004 | A Scalable Architecture for Distributed Shared Memory Multiprocessors Using Optical InterconnectsabstractSummary form only given. We describe the design and analysis of a scalable architecture suitable for large-scale DSMs (distributed shared memory) systems. The approach is based on an interconnect technology which combines optical components and a novel architecture design. In DSM systems, as the network size increases, network contention results in increasing the critical remote memory access latency, which significantly penalizes the performance of DSM systems. In our proposed architecture called RAPID (reconfigurable and scalable all-photonic interconnect for distributed-shared memory), we provide high connectivity by maximizing the channel availability for remote communication to reduce the remote memory access latency. RAPID also provides fast and efficient unicast, multicast and broadcast capabilities using a combination of aggressively designed wavelength, time and space-division multiplexing techniques. We evaluated RAPID based on network characteristics, power budget criteria and simulation using synthetic traffic workloads and compared it against other scalable electrical networks. We found that RAPID, not only outperforms other networks, but also, satisfies most of the requirements of shared memory multiprocessor design such as low latency, high bandwidth, high connectivity, and easy scalability. Avinash Karanth, Ahmed Louri |
IPDPS | 1 |
| 2004 | An Optical Interconnection Network and a Modified Snooping Protocol for the Design of Large-Scale Symmetric Multiprocessors (SMPs)abstractIn symmetric multiprocessors (SMPs), the cache coherence overhead and the speed of the shared buses limit the address/snoop bandwidth needed to broadcast transactions to all processors. As a solution, a scalable address subnetwork called symmetric multiprocessor network (SYMNET) is proposed in which address requests and snoop responses of SMPs are implemented optically. SYMNET not only uses passive optical interconnects that increases the speed of the proposed network, but also pipelines address requests at a much faster rate than electronics. This increases the address bandwidth for snooping, but the preservation of cache coherence can no longer be maintained with the usual snooping protocols. A modified coherence protocol, coherence in SYMNET (COSYM), is introduced to solve the coherence problem. COSYM was evaluated with a subset of Splash-2 benchmarks and compared with the electrical bus-based MOESI protocol. The simulation studies have shown a 5-66 percent improvement in execution time for COSYM as compared to MOESI for various applications. Simulations have also shown that the average latency for a transaction to complete using COSYM protocol was 5-78 percent better than the MOESI protocol. It is also seen that SYMNET can scale up to hundreds of processors while still using fast snooping-based cache coherence protocols, and additional performance gains may be attained with further improvement in optical device technology. Ahmed Louri, Avinash Karanth |
IEEE Trans. Parallel Distributed Syst. | 2 |