VLDB 2026 Research / reviewers in the wild / expert
Sudhakar Yalamanchili
dblp:14/5230 · also Sudha Yalamanchili
· DBLP profile ↗
82ranked-venue papers
8as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 68 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 14Artificial intelligence and machine learning · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
34 papers |
Hardware accelerators and domain-specific architectures · 18% GPUs and heterogeneous computing · 14% Energy-efficient computing · 13% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 58% Programming languages and type systems · 25% Operating systems · 17% |
Topics — the 30 heaviest of 112, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
power management |
0.8 | 4 | 2017 | Application-Specific Performance-Aware Energy Optimization on Android Mobile Devices · HPCA 2017 Harmonia: balancing compute and memory power in high-performance GPUs · ISCA 2015 Coordinated energy management in heterogeneous processors · SC 2013 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.8 | 3 | 2019 | LODESTAR: Creating Locally-Dense CNNs for Efficient Inference on Systolic Arrays · DAC 2019 DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
Memory systems
processing-in-memory |
0.6 | 2 | 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
Electronic design automation
hardware/software co-design |
0.5 | 1 | 2021 | Efficiently Solving Partial Differential Equations in a Partially Reconfigurable Specialized Hardware · IEEE Trans. Computers 2021 |
High-performance computing › scientific computing systems
partial differential equation solver |
0.5 | 1 | 2021 | Efficiently Solving Partial Differential Equations in a Partially Reconfigurable Specialized Hardware · IEEE Trans. Computers 2021 |
Reconfigurable computing and FPGAs › dynamic reconfiguration
partial reconfiguration |
0.5 | 1 | 2021 | Efficiently Solving Partial Differential Equations in a Partially Reconfigurable Specialized Hardware · IEEE Trans. Computers 2021 |
Hardware accelerators and domain-specific architectures
scientific computing accelerator |
0.5 | 1 | 2021 | Efficiently Solving Partial Differential Equations in a Partially Reconfigurable Specialized Hardware · IEEE Trans. Computers 2021 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable accelerator |
0.4 | 1 | 2020 | ALRESCHA: A Lightweight Reconfigurable Sparse-Computation Accelerator · HPCA 2020 |
High-performance computing
scientific computing |
0.4 | 1 | 2020 | ALRESCHA: A Lightweight Reconfigurable Sparse-Computation Accelerator · HPCA 2020 |
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation |
0.4 | 1 | 2020 | ALRESCHA: A Lightweight Reconfigurable Sparse-Computation Accelerator · HPCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator |
0.3 | 1 | 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
in-memory computing accelerator |
0.3 | 1 | 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Interconnection networks and networks-on-chip › network topology
low-diameter topology |
0.3 | 1 | 2018 | Slim NoC: A Low-Diameter On-Chip Network Topology for High Energy Efficiency and Scalability · ASPLOS 2018 |
Interconnection networks and networks-on-chip › network topology › network topology design
network-on-chip topology |
0.3 | 1 | 2018 | Slim NoC: A Low-Diameter On-Chip Network Topology for High Energy Efficiency and Scalability · ASPLOS 2018 |
GPUs and heterogeneous computing › GPU programming
dynamic parallelism |
0.3 | 2 | 2016 | Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUs · ISCA 2015 LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs · ISCA 2016 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.3 | 1 | 2017 | Application-Specific Performance-Aware Energy Optimization on Android Mobile Devices · HPCA 2017 |
Embedded and real-time systems
mobile devices |
0.3 | 1 | 2017 | Application-Specific Performance-Aware Energy Optimization on Android Mobile Devices · HPCA 2017 |
Hardware reliability and fault tolerance › aging
device aging |
0.2 | 1 | 2016 | Amdahl's law for lifetime reliability scaling in heterogeneous multicore processors · HPCA 2016 |
Processor architecture and microarchitecture › multicore design
heterogeneous multicore |
0.2 | 1 | 2016 | Amdahl's law for lifetime reliability scaling in heterogeneous multicore processors · HPCA 2016 |
Hardware reliability and fault tolerance › reliability analysis
lifetime reliability |
0.2 | 1 | 2016 | Amdahl's law for lifetime reliability scaling in heterogeneous multicore processors · HPCA 2016 |
Parallel and multicore computing › parallel scheduling
locality-aware scheduling |
0.2 | 1 | 2016 | LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs · ISCA 2016 |
Memory systems › processing-in-memory
memory-centric computing |
0.2 | 1 | 2016 | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
Processor architecture and microarchitecture
multicore design |
0.2 | 1 | 2016 | Amdahl's law for lifetime reliability scaling in heterogeneous multicore processors · HPCA 2016 |
Emerging computing paradigms
neuromorphic computing |
0.2 | 1 | 2016 | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
GPUs and heterogeneous computing › GPU scheduling
thread block scheduling |
0.2 | 1 | 2016 | LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUs · ISCA 2016 |
Energy-efficient computing
thermal management |
0.2 | 2 | 2015 | Cooperative boosting: needy versus greedy power management · ISCA 2013 Harmonia: balancing compute and memory power in high-performance GPUs · ISCA 2015 |
GPUs and heterogeneous computing
GPU architecture |
0.2 | 1 | 2015 | Harmonia: balancing compute and memory power in high-performance GPUs · ISCA 2015 |
GPUs and heterogeneous computing › GPU architecture
GPU execution model |
0.2 | 1 | 2015 | Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUs · ISCA 2015 |
GPUs and heterogeneous computing
GPU power management |
0.2 | 1 | 2015 | Harmonia: balancing compute and memory power in high-performance GPUs · ISCA 2015 |
Energy-efficient computing › power management › system-level power management
coordinated power management |
0.2 | 1 | 2013 | Coordinated energy management in heterogeneous processors · SC 2013 |
Methods — techniques the papers use, named apart from their topics
simulation · 1.0mathematical transformation · 0.9locally-dense storage format · 0.9structured pruning · 0.8spatial locality optimization · 0.8dynamic reconfiguration · 0.5data dependency analysis · 0.5non-prime finite field · 0.3elastic links · 0.3degree-diameter graph · 0.3producer-consumer dependence classification · 0.3offline profiling · 0.3control theory · 0.3multi-threading · 0.2SIMD · 0.2thread frontier · 0.1virtual channel flow control · 0.0coordinated scheduling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Efficiently Solving Partial Differential Equations in a Partially Reconfigurable Specialized HardwareabstractScientific computations with a wide range of applications in domains such as developing vaccines, forecasting the weather, predicting natural disasters, simulating aerodynamics of spacecraft, and exploring oil resources, create the main workloads of supercomputers. The key integration of such scientific computations is modeling physical phenomena that are done with the aid of partial differential equations (PDEs). Solving PDEs on supercomputers, even with those equipped with GPUs, consumes a large amount of power and yet is not as fast as desired. The main reason behind such slow processing is data dependency. The key challenge is that software techniques cannot resolve these dependencies, therefore, such applications cannot benefit from the parallelism provided by processors such as GPUs. Our key insight to address this challenge is that although we cannot resolve the dependencies, we can reduce their negative impacts by using hardware/software co-optimization. To this end, we propose breaking down the data-dependent operations into two groups of operations: a majority of parallelizable and the minority of data-dependent operations. We execute these two groups in the desired order: first, we put together all parallelizable operations and execute them all, subsequently; then, we switch to execute the small data-dependent part. As long as the data-dependent part is small, we can accelerate them by using fast hardware mechanisms. Besides, our proposed hardware mechanisms guarantee quickly switching between the two groups of operations. To follow the same order of execution, dictated by our software mechanism, and implemented in hardware, we also propose a new low-overhead compression format - sparsity is another attribute of PDEs that require compression. Furthermore, the core generic architecture of our proposed hardware allows the execution of other applications including sparse matrix-vector multiplication (SpMV) and graph algorithms. The key feature of the proposed hardware is partial reconfigurability, which on one hand, facilitates the execution of data-dependent computations, and on the other hand, allows executing broad application without changing the entire configuration. Our evaluations show that compared to GPUs, we achieve an average speedup of 15.6x for scientific computations while consuming 14x less energy. Bahar Asgari, Ramyad Hadidi, Tushar Krishna, Hyesoon Kim, Sudhakar Yalamanchili |
IEEE Trans. Computers | 5 |
| 2020 | Tango: An Optimizing Compiler for Just-In-Time RTL SimulationabstractWith Moore’s law coming to an end, the advent of hardware specialization presents a unique challenge for a much tighter software and hardware co-design environment to exploit domain-specific optimizations and increase design efficiency. This trend is further accentuated by rapid-pace of innovations in Machine Learning and Graph Analytic, calling for a faster product development cycle for hardware accelerators and the importance of addressing the increasing cost of hardware verification. The productivity of software-hardware co-design relies upon better integration between the software and hardware design methodologies, but more importantly in the effectiveness of the design tools and hardware simulators at reducing the development time. In this work, we developed Tango, an Optimizing compiler for Just-in-Time RTL simulation. Tango implements unique hardware-centric compiler transformations to speed up runtime code generation in a software-hardware co-design environment where hardware simulation speed is critical. Tango achieves a 6x average speedup compared to the state-of-the-art simulators. Blaise-Pascal Tine, Sudhakar Yalamanchili, Hyesoon Kim |
DATE | 2 |
| 2020 | ALRESCHA: A Lightweight Reconfigurable Sparse-Computation AcceleratorabstractSparse problems that dominate a wide range of applications fail to effectively benefit from high memory bandwidth and concurrent computations in modern high-performance computer systems. Therefore, hardware accelerators have been proposed to capture a high degree of parallelism in sparse problems. However, the unexplored challenge for sparse problems is the limited opportunity for parallelism because of data dependencies, a common computation pattern in scientific sparse problems. Our key insight is to extract parallelism by mathematically transforming the computations into equivalent forms. The transformation breaks down the sparse kernels into a majority of independent parts and a minority of data-dependent ones and reorders these parts to gain performance. To implement the key insight, we propose a lightweight reconfigurable sparse-computation accelerator (Alrescha). To efficiently run the data-dependent and parallel parts and to enable fast switching between them, Alrescha makes two contributions. First, it implements a compute engine with a fixed compute unit for the parallel parts and a lightweight reconfigurable engine for the execution of the data-dependent parts. Second, Alrescha benefits from a locally-dense storage format, with the right order of non-zero values to yield the order of computations dictated by the transformation. The combination of the lightweight reconfigurable hardware and the storage format enables uninterrupted streaming from memory. Our simulation results show that compared to GPU, Alrescha achieves an average speedup of 15.6x for scientific sparse problems, and 8x for graph algorithms. Moreover, compared to GPU, Alrescha consumes 14x less energy. Bahar Asgari, Ramyad Hadidi, Tushar Krishna, Hyesoon Kim, Sudhakar Yalamanchili |
HPCA | 5 |
| 2019 | POSTER: Tango: An Optimizing Compiler for Just-In-Time RTL SimulationabstractThe end of Moore's law with the advent of hardware specialization presents a unique challenge for a much tighter software and hardware co-design environment to exploit domain-specific optimizations and increase design efficiency. The productivity of software-hardware codesign relies not on only in better integration between the software and hardware design methodologies but more importantly in the effectiveness of the design tools at reducing the development time. In this work, we developed Tango, an Optimizing compiler for a Just-in-Time RTL simulator. Tango implements unique hardware-centric compiler transformations to speed up runtime code generation in a software-hardware codesign environment where hardware simulation speed is critical. Tango achieves a 3x average speedup compared to the state-of-the-art RTL simulators. Blaise-Pascal Tine, Sudhakar Yalamanchili, Hyesoon Kim, Jeffrey S. Vetter |
PACT | 2 |
| 2019 | LODESTAR: Creating Locally-Dense CNNs for Efficient Inference on Systolic ArraysabstractThe performance of sparse problems suffers from lack of spatial locality and low memory bandwidth utilization. However, the distribution of non-zero values in the data structures of a class of sparse problems, such as matrix operations in neural networks, is modifiable so that it can be matched with an efficient underlying hardware, such as systolic arrays. Such modification helps addressing the challenges coupled with sparsity. To efficiently execute sparse neural network inference on systolic arrays, we propose a structured pruning algorithm that increases the spatial locality in neural network models, while maintaining the accuracy of inference. Bahar Asgari, Ramyad Hadidi, Hyesoon Kim, Sudhakar Yalamanchili |
DAC | 4 |
| 2018 | Slim NoC: A Low-Diameter On-Chip Network Topology for High Energy Efficiency and ScalabilityabstractEmerging chips with hundreds and thousands of cores require networks with unprecedented energy/area efficiency and scalability. To address this, we propose Slim NoC (SN): a new on-chip network design that delivers significant improvements in efficiency and scalability compared to the state-of-the-art. The key idea is to use two concepts from graph and number theory, degree-diameter graphs combined with non-prime finite fields, to enable the smallest number of ports for a given core count. SN is inspired by state-of-the-art off-chip topologies; it identifies and distills their advantages for NoC settings while solving several key issues that lead to significant overheads on-chip. SN provides NoC-specific layouts, which further enhance area/energy efficiency. We show how to augment SN with state-of-the-art router microarchitecture schemes such as Elastic Links, to make the network even more scalable and efficient. Our extensive experimental evaluations show that SN outperforms both traditional low-radix topologies (e.g., meshes and tori) and modern high-radix networks (e.g., various Flattened Butterflies) in area, latency, throughput, and static/dynamic power consumption for both synthetic and real workloads. SN provides a promising direction in scalable and energy-efficient NoC topologies. Maciej Besta, Syed Minhaj Hassan, Sudhakar Yalamanchili, Rachata Ausavarungnirun, Onur Mutlu, Torsten Hoefler |
ASPLOS | 3 |
| 2018 | A ferroelectric FET based power-efficient architecture for data-intensive computingabstractIn this paper, we present a ferroelectric FET (FeFET) based power-efficient architecture to accelerate data-intensive applications such as deep neural networks (DNNs). We propose a cross-cutting solution combining emerging device technologies, circuit optimizations, and micro-architecture innovations. At device level, FeFET crossbar is utilized to perform vector-matrix multiplication (VMM). As a field effect device, FeFET significantly reduces the read/write energy compared with the resistive random-access memory (ReRAM). At circuit level, we propose an all-digital peripheral design, reducing the large overhead introduced by ADC and DAC in prior works. In terms of micro-architecture innovation, a dedicated hierarchical network-on-chip (H-NoC) is developed for input broadcasting and on-the-fly partial results processing, reducing the data transmission volume and latency. Speed, power, area and computing accuracy are evaluated based on detailed device characterization and system modeling. For DNN computing, our design achieves 254x and 9.7x gain in power efficiency (GOPS/W) compared to GPU and ReRAM based designs, respectively. Taesik Na, Prakshi Rastogi, Karthik Rao, Asif Islam Khan, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
ICCAD | 6 |
| 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural NetworksabstractThis paper presents, DeepTrain, an embedded platform for high-performance and energy-efficient training of deep neural network (DNN). The key architectural concept of DeepTrain is to develop a spatially homogeneous computing (and memory) fabric with temporally heterogeneous programmable data flows to optimize memory mapping and data reuse during different phases of training operation.The DeepTrain is demonstrated as an in-memory accelerator integrated in the logic layer of a 3-D memory module. A programming model and supporting architecture utilizes the flexible data flow to efficiently accelerate training of various types of DNNs. The cycle level simulation and synthesized design in 15 nm FinFET shows power efficiency of 500 GFLOPS/W, and almost similar throughput for a wide range of DNNs, including convolutional, recurrent, and mixed (CNN+RNN) networks. Duckhwan Kim 0001, Taesik Na, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Application-Specific Performance-Aware Energy Optimization on Android Mobile DevicesabstractEnergy management is a key issue for mobile devices. On current Android devices, power management relies heavily on OS modules known as governors. These modules are created for various hardware components, including the CPU, to support DVFS. They implement algorithms that attempt to balance performance and power consumption. In this paper we make the observation that the existing governors are (1) general-purpose by nature (2) focused on power reduction and (3) are not energy-optimal for many applications. We thus establish the need for an application-specific approach that could overcome these drawbacks and provide higher energy efficiency for suitable applications. We also show that existing methods manage power and performance in an independent and isolated fashion and that co-ordinated control of multiple components can save more energy. In addition, we note that on mobile devices, energy savings cannot be achieved at the expense of performance. Consequently, we propose a solution that minimizes energy consumption of specific applications while maintaining a user-specified performance target. Our solution consists of two stages: (1) offline profiling and (2) online controlling. Utilizing the offline profiling data of the target application, our control theory based online controller dynamically selects the optimal system configuration (in this paper, combination of CPU frequency and memory bandwidth) for the application, while it is running. Our energy management solution is tested on a Nexus 6 smartphone with 6 real-world applications. We achieve 4 - 31% better energy than default governors with a worst case performance loss of <; 1%. Karthik Rao, Jun Wang 0077, Sudhakar Yalamanchili, Yorai Wardi, Handong Ye |
HPCA | 3 |
| 2016 | Amdahl's law for lifetime reliability scaling in heterogeneous multicore processorsabstractHeterogeneous multicore processors have been suggested as alternative microarchitectural designs to enhance performance and energy efficiency. Using Amdahl's Law, heterogeneous models were primarily analyzed in performance and energy efficiency aspects to demonstrate its advantage over conventional homogeneous systems. In this paper, we further extend the study to understand the lifetime reliability consequences of heterogeneous multicore processors, as reliability becomes an increasingly important constraint. We present the lifetime reliability models of multicore processors based on Amdahl's Law, including compact thermal estimation that has strong correlation with device aging. Lifetime reliability is analyzed by varying i) core utilization (Amdahl's scaling factor), ii) processor composition (number of big and small cores), and iii) thread scheduling method. The study shows that the heterogeneous processor may have a serious reliability challenge. If the processor is comprised of only one big core and many small cores, stresses can be biased to the big core especially when workloads spend more time on sequential operations. Our study reveals that incorporating multiple big cores can mitigate reliability bottleneck in big cores and enhance processor lifetime, but adding too many big cores will have an adverse impact on lifetime reliability as well as performance. William J. Song, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
HPCA | 3 |
| 2016 | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D MemoryabstractThis paper presents a programmable and scalable digital neuromorphic architecture based on 3D high-density memory integrated with logic tier for efficient neural computing. The proposed architecture consists of clusters of processing engines, connected by 2D mesh network as a processing tier, which is integrated in 3D with multiple tiers of DRAM. The PE clusters access multiple memory channels (vaults) in parallel. The operating principle, referred to as the memory centric computing, embeds specialized state-machines within the vault controllers of HMC to drive data into the PE clusters. The paper presents the basic architecture of the Neurocube and an analysis of the logic tier synthesized in 28nm and 15nm process technologies. The performance of the Neurocube is evaluated and illustrated through the mapping of a Convolutional Neural Network and estimating the subsequent power and performance for both training and inference. Duckhwan Kim 0001, Jaeha Kung 0001, Sek M. Chai, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
ISCA | 4 |
| 2016 | LaPerm: Locality Aware Scheduler for Dynamic Parallelism on GPUsabstractRecent developments in GPU execution models and architectures have introduced dynamic parallelism to facilitate the execution of irregular applications where control flow and memory behavior can be unstructured, time-varying, and hierarchical. The changes brought about by this extension to the traditional bulk synchronous parallel (BSP) model also creates new challenges in exploiting the current GPU memory hierarchy. One of the major challenges is that the reference locality that exists between the parent and child thread blocks (TBs) created during dynamic nested kernel and thread block launches cannot be fully leveraged using the current TB scheduling strategies. These strategies were designed for the current implementations of the BSP model but fall short when dynamic parallelism is introduced since they are oblivious to the hierarchical reference locality. We propose LaPerm, a new locality-aware TB scheduler that exploits such parent-child locality, both spatial and temporal. LaPerm adopts three different scheduling decisions to i) prioritize the execution of the child TBs, ii) bind them to the stream multiprocessors (SMXs) occupied by their parents TBs, and iii) maintain workload balance across compute units. Experiments with a set of irregular CUDA applications executed on a cycle-level simulator employing dynamic parallelism demonstrate that LaPerm is able to achieve an average of 27% performance improvement over the baseline round-robin TB scheduler commonly used in modern GPUs. Jin Wang 0010, Norman Rubin, Albert Sidelnik, Sudhakar Yalamanchili |
ISCA | 4 |
| 2015 | Throughput Regulation in Shared Memory Multicore ProcessorsabstractPerformance scaling is now synonymous with scaling the number of cores. One of the consequences of this shift is the increasing difficulty of designing processors with predictable and controllable performance. To address this challenge this paper proposes a chip-scale throughput regulation technique that is based on dynamic tracking of instruction execution dynamics in each core. A new variable gain controller design is developed for regulating the throughput of modern out-of-order cores. The gain is adjusted based on an on-line sensitivity analysis of the core's throughput to the control parameter. We explore throughput regulation using two control paramaters - core frequency and instruction issue width and demonstrate via cycle-level, full system simulation the utility of the proposed regulator on both compute and memory intensive workloads. Performance results are presented for the application to a 16 core, cache coherent 3D multicore processor. H. Xiao, Yorai Wardi, Sudhakar Yalamanchili |
HiPC | 4 |
| 2015 | Harmonia: balancing compute and memory power in high-performance GPUsabstractIn this paper, we address the problem of efficiently managing the relative power demands of a high-performance GPU and its memory subsystem. We develop a management approach that dynamically tunes the hardware operating configurations to maintain balance between the power dissipated in compute versus memory access across GPGPU application phases. Our goal is to reduce power with minimal performance degradation. Indrani Paul, Wei Huang 0004, Manish Arora, Sudhakar Yalamanchili |
ISCA | 4 |
| 2015 | Dynamic thread block launch: a lightweight execution mechanism to support irregular applications on GPUsabstractGPUs have been proven effective for structured applications that map well to the rigid 1D-3D grid of threads in modern bulk synchronous parallel (BSP) programming languages. However, less success has been encountered in mapping data intensive irregular applications such as graph analytics, relational databases, and machine learning. Recently introduced nested device-side kernel launching functionality in the GPU is a step in the right direction, but still falls short of being able to effectively harness the GPUs performance potential. Jin Wang 0010, Norman Rubin, Albert Sidelnik, Sudhakar Yalamanchili |
ISCA | 4 |
| 2014 | Red Fox: An Execution Environment for Relational Query Processing on GPUs
Haicheng Wu, Gregory Frederick Diamos, Tim Sheard, Molham Aref, Sean Baxter, Michael Garland, Sudhakar Yalamanchili |
CGO | 7 |
| 2014 | Harmonica: An FPGA-Based Data Parallel Soft CoreabstractGeneral-purpose GPUs or GPGPUs have taken their place in the market, being present in 38 of the Top 500 supercomputers [5]. In the same way that the emergence of FPGAs in the 1980s led to a demand for soft cores with instruction sets similar to the CPUs of the day, we anticipate a similar demand in the 2010s for soft cores with GPGPU instruction sets. These architectures are distinguished by their SIMT, single-instruction-multiple-thread, execution model, acheiving throughput by running multiple threads of execution simultaneously across multiple functional units, keeping separate register values for each lane of execution. Chad D. Kersey, Sudhakar Yalamanchili, Hyojong Kim, Nimit Nigania, Hyesoon Kim |
FCCM | 2 |
| 2014 | Energy Introspector: A parallel, composable framework for integrated power-reliability-thermal modeling for multicore architecturesabstractSustaining processor performance growth is challenged by physical limitations due to increased power and heat dissipations. Power and thermal management techniques combined with inherent workload dynamics create the spatiotemporal variations of power, temperature, and degradation in processors. As industry moves to smaller feature sizes, the performance will become increasingly dominated by the physics. The challenge is in understanding how the physics is manifested at the microarchitecture level. This requires the modeling and simulation environment that can capture multiple, distinct physical phenomena and their concurrent impact on the microarchitecture. William J. Song, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
ISPASS | 3 |
| 2014 | Manifold: A parallel simulation framework for multicore systemsabstractThis paper presents Manifold, an open-source parallel simulation framework for multicore architectures. It consists of a parallel simulation kernel, a set of microarchitecture components, and an integrated library of power, thermal, reliability, and energy models. Using the components as building blocks, users can assemble multicore architecture simulation models and perform serial or parallel simulations to study the architectural and/or the physical characteristics of the models. Users can also create new components for Manifold or port existing models. Importantly, Manifold's component-based design provides the user with the ability to easily replace a component with another for efficient explorations of the design space. It also allows components to evolve independently and making it easy for simulators to incorporate new components as they become available. The distinguishing features of Manifold include i) transparent parallel execution, ii) integration of power, thermal, reliability, and energy models, iii) full system simulation, e.g., operating system and system binaries, and iv) component-based design. In this paper we provide a description of the software architecture of Manifold, and its main elements - a parallel multicore emulator front-end and a parallel component-based back-end timing model. We describe a few simulators that are built with Manifold components to illustrate its flexibility, and present test results of the scalability obtained on full-system simulation of coherent shared-memory multicore models with 16, 32, and 64 cores executing PARSEC and SPLASH-2 benchmarks. Jun Wang 0077, Jesse G. Beu, Rishiraj A. Bheda, Thomas M. Conte, Zhenjiang Dong, Chad D. Kersey, Mitchelle Rasquinha, George F. Riley, William J. Song, Sudhakar Yalamanchili |
ISPASS | 12 |
| 2014 | Bubble sharing: Area and energy efficient adaptive routers using centralized buffersabstractEdge buffers along with multiple virtual channels have traditionally been used to provide deadlock freedom guarantees in on-chip networks. The problem with such schemes is their high buffer space requirement which consumes significant power and area. In this work, we propose bubble sharing flow control to provide deadlock freedom with small, shared central buffers, eliminating edge buffers, improving buffer utilization, and decreasing router buffer requirements. The key insight involves sharing of the flit-size bubbles (free buffers) among cyclic network paths via central buffers in the router, reducing the overall router buffering space requirement. This technique effectively reconciles the trade-off between high radix and buffer space, encouraging the use of low hop count, high-radix topologies, with both deterministic and adaptive routing. Comparisons show improvement in average packet latency by 31% as compared to traditional 2VC edge buffer routers with 33% reduction in area for an 8×8 generalized hypercube topology. Syed Minhaj Hassan, Sudhakar Yalamanchili |
NOCS | 2 |
| 2014 | Power Modeling for GPU Architectures Using McPATabstractGraphics Processing Units (GPUs) are very popular for both graphics and general-purpose applications. Since GPUs operate many processing units and manage multiple levels of memory hierarchy, they consume a significant amount of power. Although several power models for CPUs are available, the power consumption of GPUs has not been studied much yet. In this article we develop a new power model for GPUs by utilizing McPAT, a CPU power tool. We generate initial power model data from McPAT with a detailed GPU configuration, and then adjust the models by comparing them with empirical data. We use the NVIDIA's Fermi architecture for building the power model, and our model estimates the GPU power consumption with an average error of 7.7% and 12.8% for the microbenchmarks and Merge benchmarks, respectively. Jieun Lim 0001, Nagesh B. Lakshminarayana, Hyesoon Kim, William J. Song, Sudhakar Yalamanchili, Wonyong Sung |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2014 | Control Principles and On-Chip Circuits for Active Cooling Using Integrated Superlattice-Based Thin-Film Thermoelectric DevicesabstractSuperlattice thin-film thermoelectric coolers (TECs) are emerging as a promising technology for hot spot mitigation in microprocessors. This paper studies the prospect of on-demand cooling with advanced TECs integrated at the back of the heat spreader inside a package (integrated TEC). Using thermal compact models of the chip and package with integrated TECs, the control principles for TEC-assisted transient cooling are presented. The control principles are implemented in a 130-nm CMOS process and cosimulated with the thermal system to show their feasibility and energy overheads. The simulation results show potential for extending the time for which a chip and package can sustain a high power load. Borislav Alexandrov, Owen Sullivan, William J. Song, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | Oncilla: A GAS runtime for efficient resource allocation and data movement in accelerated clustersabstractAccelerated and in-core implementations of Big Data applications typically require large amounts of host and accelerator memory as well as efficient mechanisms for transferring data to and from accelerators in heterogeneous clusters. Scheduling for heterogeneous CPU and GPU clusters has been investigated in depth in the high-performance computing (HPC) and cloud computing arenas, but there has been less emphasis on the management of cluster resource that is required to schedule applications across multiple nodes and devices. Previous approaches to address this resource management problem have focused on either using low-performance software layers or on adapting complex data movement techniques from the HPC arena, which reduces performance and creates barriers for migrating applications to new heterogeneous cluster architectures. This work proposes a new system architecture for cluster resource allocation and data movement built around the concept of managed Global Address Spaces (GAS), or dynamically aggregated memory regions that span multiple nodes.We propose a software layer called Oncilla that uses a simple runtime and API to take advantage of non-coherent hardware support for GAS. The Oncilla runtime is evaluated using two different high-performance networks for microkernels representative of the TPC-H data warehousing benchmark, and this runtime enables a reduction in runtime of up to 81%, on average, when compared with standard disk-based data storage techniques. The use of the Oncilla API is also evaluated for a simple breadth-first search (BFS) benchmark to demonstrate how existing applications can incorporate support for managed GAS. Jeffrey Young 0001, Se Hoon Shon, Sudhakar Yalamanchili, Alex Merritt, Karsten Schwan, Holger Fröning |
CLUSTER | 3 |
| 2013 | Cooperative boosting: needy versus greedy power managementabstractThis paper examines the interaction between thermal management techniques and power boosting in a state-of-the-art heterogeneous processor consisting of a set of CPU and GPU cores. We show that for classes of applications that utilize both the CPU and the GPU, modern boost algorithms that greedily seek to convert thermal headroom into performance can interact with thermal coupling effects between the CPU and the GPU to degrade performance. We first examine the causes of this behavior and explain the interaction between thermal coupling, performance coupling, and workload behavior. Then we propose a dynamic power-management approach called cooperative boosting (CB) to allocate power dynamically between CPU and GPU in a manner that balances thermal coupling against the needs of performance coupling to optimize performance under a given thermal constraint. Through real hardware-based measurements, we evaluate CB against a state-of-the-practice boost algorithm and show that overall application performance and power savings increase by 10% and 8% (up to 52% and 34%), respectively, resulting in average energy efficiency improvement of 25% (up to 76%) over a wide range of benchmarks. Indrani Paul, Srilatha Manne, Manish Arora, William Lloyd Bircher, Sudhakar Yalamanchili |
ISCA | 5 |
| 2013 | A Study of the Effect of Partitioning on Parallel Simulation of Multicore SystemsabstractThere has been little research that studies the effect of partitioning on parallel simulation of multicore systems. This paper presents our study of this important problem in the context of Null-message-based synchronization algorithm for parallel multicore simulation. This paper focuses on coarse grain parallel simulation where each core and its cache slices are modeled within a single logical process (LP) and different partitioning schemes are only applied to the interconnection network. In this paper we show that encapsulating the entire on-chip interconnection network into a single logical process is an impediment to scalable simulation. This baseline partitioning and two other schemes are investigated. Experiments are conducted on a subset of the PARSEC benchmarks with 16-, 32-, 64- and 128-core models. Results show that the partitioning scheme has a significant impact on simulation performance and parallel efficiency. Beyond a certain system scale, one scheme consistently outperforms the other two schemes, and the performance as well as efficiency gaps increases as the size of the model increases - with up to 4.1 times faster speed and 277% better efficiency for 128-core models. We explain the reasons for this behavior, which can be traced to the features of the Null-message-based synchronization algorithm. Because of this, we believe that, if a component has increasing number of inter-LP interactions with increasing system size, such components should be partitioned into several sub-components to achieve better performance. Zhenjiang Dong, Jun Wang 0077, George F. Riley, Sudhakar Yalamanchili |
MASCOTS | 4 |
| 2013 | Centralized buffer router: A low latency, low power router for high radix NOCsabstractWhile router buffers have been used as performance multipliers, they are also major consumers of area and power in on-chip networks. In this paper, we propose centralized elastic bubble router - a router micro-architecture based on the use of centralized buffers (CB) with elastic buffered (EB) links. At low loads, the CB is power gated, bypassed, and optimized to produce single cycle operation. A novel extension to bubble flow control enables routing deadlock and message dependent deadlock to be avoided with the same mechanism having constant buffer size per router independent of the number of message types. This solution enables end-to-end latency reduction via high radix switches with low overall buffer requirements. Comparisons made with other low latency routers across different topologies show consistent performance improvement, for example 26% improvement in no load latency of a 2D Mesh and 4X improvement in saturation throughput in a 2D-Generalized Hypercube. Syed Minhaj Hassan, Sudhakar Yalamanchili |
NOCS | 2 |
| 2013 | Optimizing parallel simulation of multicore systems using domain-specific knowledgeabstractThis paper presents two optimization techniques for the basic Null-message algorithm in the context of parallel simulation of multicore computer architectures. Unlike the general, application-independent optimization methods, these are application-specific optimizations that make use of system properties of the simulation application. We demonstrate in two aspects that the domain-specific knowledge offers great potential for optimization. First, it allows us to send Null-messages much less eagerly, thus greatly reducing the amount of Null-messages. Second, the internal state of the simulation application allows us to make conservative forecast of future outgoing events. This leads to the creation of an enhanced synchronization algorithm called Forecast Null-message algorithm, which, by combining the forecast from both sides of a link, can greatly improve the simulation look-ahead. Compared with the basic Null-message algorithm, our optimizations greatly reduce the number of Null-messages and increase simulation performance significantly as a result. On a subset of the PARSEC benchmarks, a maximum speedup of about 6 is achieved with 17 LPs. Jun Wang 0077, Zhenjiang Dong, Sudhakar Yalamanchili, George F. Riley |
SIGSIM-PADS | 3 |
| 2013 | Relational algorithms for multi-bulk-synchronous processorsabstractRelational databases remain an important application infrastructure for organizing and analyzing massive volumes of data. At the same time, processor architectures are increasingly gravitating towards Multi-Bulk-Synchronous processor (Multi-BSP) architectures employing throughput-optimized memory systems, lightweight multi-threading, and Single-Instruction Multiple-Data (SIMD) core organizations. This paper explores the mapping of primitive relational algebra operations onto such architectures to improve the throughput of data warehousing applications built on relational databases. Gregory Frederick Diamos, Haicheng Wu, Jin Wang 0010, Ashwin Sanjay Lele, Sudhakar Yalamanchili |
PPoPP | 5 |
| 2013 | Coordinated energy management in heterogeneous processorsabstractThis paper examines energy management in a heterogeneous processor consisting of an integrated CPU-GPU for high-performance computing (HPC) applications. Energy management for HPC applications is challenged by their uncompromising performance requirements and complicated by the need for coordinating energy management across distinct core types -- a new and less understood problem. Indrani Paul, Vignesh T. Ravi, Srilatha Manne, Manish Arora, Sudhakar Yalamanchili |
SC | 5 |
| 2013 | Design space exploration of on-chip ring interconnection for a CPU-GPU heterogeneous architecture
Jaekyu Lee, Hyesoon Kim, Sudhakar Yalamanchili |
J. Parallel Distributed Comput. | 4 |
| 2013 | Adaptive virtual channel partitioning for network-on-chip in heterogeneous architecturesabstractCurrent heterogeneous chip-multiprocessors (CMPs) integrate a GPU architecture on a die. However, the heterogeneity of this architecture inevitably exerts different pressures on shared resource management due to differing characteristics of CPU and GPU cores. We consider how to efficiently share on-chip resources between cores within the heterogeneous system, in particular the on-chip network. Heterogeneous architectures use an on-chip interconnection network to access shared resources such as last-level cache tiles and memory controllers, and this type of on-chip network will have a significant impact on performance. In this article, we propose a feedback-directed virtual channel partitioning (VCP) mechanism for on-chip routers to effectively share network bandwidth between CPU and GPU cores in a heterogeneous architecture. VCP dedicates a few virtual channels to CPU and GPU applications with separate injection queues. The proposed mechanism balances on-chip network bandwidth for applications running on CPU and GPU cores by adaptively choosing the best partitioning configuration. As a result, our mechanism improves system throughput by 15% over the baseline across 39 heterogeneous workloads. Jaekyu Lee, Hyesoon Kim, Sudhakar Yalamanchili |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2012 | Dynamic compilation of data-parallel kernels for vector processorsabstractModern processors enjoy augmented throughput and power efficiency through specialized functional units leveraged via instruction set extensions. These functional units accelerate performance for specific types of operations but must be programmed explicitly. Moreover, applications targeting these specialized units will not take advantage of future ISA extensions and tend not to be portable across multiple ISAs. As architecture designers increasingly rely on heterogeneity for performance improvements, the challenges of leveraging specialized functional units will only become more critical. In particular, exploiting software parallelism without sacrificing portability across the spectrum of commodity and multi-core SIMD processors remains elusive. Andrew Kerr, Gregory Frederick Diamos, Sudhakar Yalamanchili |
CGO | 3 |
| 2012 | Eiger: A framework for the automated synthesis of statistical performance modelsabstractAs processor architectures continue to evolve to increasingly heterogeneous and asymmetric designs, the construction of accurate performance models of execution time and energy consumption has become increasingly more challenging. Models that are constructed, are quickly invalidated by new features in the next generation of processors while many interactions between application and architecture parameters are often simply not obvious or even apparent. Consequently, we foresee a need for an automated methodology for the systematic construction of performance models of heterogeneous processors. The methodology should be founded on rigorous mathematical techniques yet leave room for the exploration and adaptation of a space of analytic models. Our current effort toward creating such an extensible, targeted methodology is Eiger. This paper describes the methodology implemented in Eiger, the specifics of Eiger's extensible implementation and the results of one scenario in which Eiger has been applied - the synthesis of performance models for use in the simulation-based design space exploration of Exascale architectures. Andrew Kerr, Eric Anger, Gilbert Hendry, Sudhakar Yalamanchili |
HiPC | 4 |
| 2012 | Performance impact of virtual machine placement in a datacenterabstractIn virtualized systems, several Virtual Machines (VM) running on a single hardware platform share and compete for the hardware resources such as memory, disk and network IO to meet a certain Quality of Service (QoS) requirements. It is critical to characterize and understand how the different workloads running in the VMs interact and share such resources to be able to map them efficiently onto processor cores and server hosts for optimal performance. This is especially important for resources such as memory controllers or the on-chip or inter-socket networks for which there is currently no software control. In this paper, we present a measurement-based performance analysis of server virtualization workloads from a real system using virtual machines that are part of the popular industry standard VMmark benchmark, a server consolidation benchmark. First, we characterize the relative resource contention and interference impact of VMs when multiple virtual workloads are run together. Second, we study the effects of co-locating different types of VMs under various VM to core placement schemes and discover the best placement for performance. We observe performance variations from 25-65% for Database servers and from 7-40% for File servers when compared to standalone VM depending on the placement of these VMs onto cores and the degree of sharing of resources. Finally, we propose an interference metric and regression model for the worst set of co-located VMs in our study. Based on different VM placement schemes we show that the overall server consolidation performance in a virtualized host can be improved by 8% when the VMs are placed effectively. Indrani Paul, Sudhakar Yalamanchili, Lizy Kurian John |
IPCCC | 2 |
| 2012 | Lynx: A dynamic instrumentation system for data-parallel applications on GPGPU architecturesabstractAs parallel execution platforms continue to proliferate, there is a growing need for real-time introspection tools to provide insight into platform behavior for performance debugging, correctness checks, and to drive effective resource management schemes. To address this need, we present the Lynx dynamic instrumentation system. Lynx provides the capability to write instrumentation routines that are (1) selective, instrumenting only what is needed, (2) transparent, without changes to the applications' source code, (3) customizable, and (4) efficient. Lynx is embedded into the broader GPU Ocelot system, which provides run-time code generation of CUDA programs for heterogeneous architectures. This paper describes (1) the Lynx framework and implementation, (2) its language constructs geared to the Single Instruction Multiple Data (SIMD) model of data-parallel programming used in current general-purpose GPU (GPGPU) based systems, and (3) useful performance metrics described via Lynx's instrumentation language that provide insights into the design of effective instrumentation routines for GPGPU systems. The paper concludes with a comparative analysis of Lynx with existing GPU profiling tools and a quantitative assessment of Lynx's instrumentation performance, providing insights into optimization opportunities for running instrumented GPU kernels. Naila Farooqui, Andrew Kerr, Greg Eisenhauer, Karsten Schwan, Sudhakar Yalamanchili |
ISPASS | 5 |
| 2012 | Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU ComputationabstractData warehousing applications represent an emerging application arena that requires the processing of relational queries and computations over massive amounts of data. Modern general purpose GPUs are high bandwidth architectures that potentially offer substantial improvements in throughput for these applications. However, there are significant challenges that arise due to the overheads of data movement through the memory hierarchy and between the GPU and host CPU. This paper proposes data movement optimizations to address these challenges. Inspired in part by loop fusion optimizations in the scientific computing community, we propose kernel fusion as a basis for data movement optimizations. Kernel fusion fuses the code bodies of two GPU kernels to i) reduce data footprint to cut down data movement throughout GPU and CPU memory hierarchy, and ii) enlarge compiler optimization scope. We classify producer consumer dependences between compute kernels into three types, i) fine-grained thread-to-thread dependences, ii) medium-grained thread block dependences, and iii) coarse-grained kernel dependences. Based on this classification, we propose a compiler framework, Kernel Weaver, that can automatically fuse relational algebra operators thereby eliminating redundant data movement. The experiments on NVIDIA Fermi platforms demonstrate that kernel fusion achieves 2.89x speedup in GPU computation and a 2.35x speedup in PCIe transfer time on average across the micro-benchmarks tested. We present key insights, lessons learned, measurements from our compiler implementation, and opportunities for further improvements. Haicheng Wu, Gregory Frederick Diamos, Srihari Cadambi, Sudhakar Yalamanchili |
MICRO | 4 |
| 2011 | Regulating Locality vs. Parallelism Tradeoffs in Multiple Memory Controller EnvironmentsabstractThe presence of multiple MCs and their integration into the on-chip network fabric creates a highly concurrent system that can support significant levels of memory level parallelism (MLP) across cores. This work exposes the trade-off between DRAM parameters, bank level parallelism (BLP), and row buffer hit rate that exposes the amount of effective BLP that is necessary to approximate a 100% hit rate. We further study how this trade-off can be controlled and propose a class of global (system) and local (within an MC) address mappings that can be tuned to optimize the performance across a set of multiprogrammed benchmarks. Syed Minhaj Hassan, Dhruv Choudhary, Mitchelle Rasquinha, Sudhakar Yalamanchili |
PACT | 4 |
| 2011 | SIMD re-convergence at thread frontiersabstractHardware and compiler techniques for mapping data-parallel programs with divergent control flow to SIMD architectures have recently enabled the emergence of new GPGPU programming models such as CUDA, OpenCL, and DirectX Compute. The impact of branch divergence can be quite different depending upon whether the program's control flow is structured or unstructured. In this paper, we show that unstructured control flow occurs frequently in applications and can lead to significant code expansion when executed using existing approaches for handling branch divergence. Gregory Frederick Diamos, Benjamin Ashbaugh, Subramaniam Maiyuran, Andrew Kerr, Haicheng Wu, Sudhakar Yalamanchili |
MICRO | 6 |
| 2011 | A Scalable Design Methodology for Energy Minimization of STTRAM: A Circuit and Architecture PerspectiveabstractIn this paper, we analyze the energy dissipation in spin-torque-transfer random access memory array (STTRAM). We present a methodology for exploring the design space to minimize the energy dissipation of the array while maintaining required read and write quality for a given magnetic tunnel junction technology. The proposed method shows the need for proper choice of the silicon transistor width and array operating voltage to minimize the energy dissipation of the STTRAM array. The write energy is found to be 10 × greater than read energy. Hence, read-write ratio becomes a crucial factor that determines energy for STTRAM last level caches (L2). An exploration is performed across several architectural benchmarks including shared and non-shared caches for detailed energy analysis. Subho Chatterjee, Mitchelle Rasquinha, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Ocelot: a dynamic optimization framework for bulk-synchronous applications in heterogeneous systemsabstractOcelot is a dynamic compilation framework designed to map the explicitly data parallel execution model used by NVIDIA CUDA applications onto diverse multithreaded platforms. Ocelot includes a dynamic binary translator from Parallel Thread eXecution ISA (PTX) to many-core processors that leverages the Low Level Virtual Machine (LLVM) code generator to target x86 and other ISAs. The dynamic compiler is able to execute existing CUDA binaries without recompilation from source and supports switching between execution on an NVIDIA GPU and a many-core CPU at runtime. It has been validated against over 130 applications taken from the CUDA SDK, the UIUC Parboil benchmarks [1], the Virginia Rodinia benchmarks [2], the GPU-VSIPL signal and image processing library [3], the Thrust library [4], and several domain specific applications. Gregory Frederick Diamos, Andrew Kerr, Sudhakar Yalamanchili, Nathan Clark |
PACT | 3 |
| 2010 | Speculative execution on multi-GPU systemsabstractThe lag of parallel programming models and languages behind the advance of heterogeneous many-core processors has left a gap between the computational capability of modern systems and the ability of applications to exploit them. Emerging programming models, such as CUDA and OpenCL, force developers to explicitly partition applications into components (kernels) and assign them to accelerators in order to utilize them effectively. An accelerator is a processor with a different ISA and micro-architecture than the main CPU. These static partitioning schemes are effective when targeting a system with only a single accelerator. However, they are not robust to changes in the number of accelerators or the performance characteristics of future generations of accelerators. In previous work, we presented the Harmony execution model for computing on heterogeneous systems with several CPUs and accelerators. In this paper, we extend Harmony to target systems with multiple accelerators using control speculation to expose parallelism. We refer to this technique as Kernel Level Speculation (KLS). We argue that dynamic parallelization techniques such as KLS are sufficient to scale applications across several accelerators based on the intuition that there will be fewer distinct accelerators than cores within each accelerator. In this paper, we use a complete prototype of the Harmony runtime that we developed to explore the design decisions and trade-offs in the implementation of KLS. We show that KLS improves parallelism to a sufficient degree while retaining a sequential programming model. We accomplish this by demonstrating good scaling of KLS on a highly heterogeneous system with three distinct accelerator types and ten processors. Gregory Frederick Diamos, Sudhakar Yalamanchili |
IPDPS | 2 |
| 2010 | An energy efficient cache design using spin torque transfer (STT) RAMabstractThe on-chip memory is a dominant source of power and energy consumption in modern and future processors. This paper explores the use of a new emerging non-volatile memory technology as a replacement for SRAM based lower level caches - Spin Torque Transfer(STT) RAM. While STTRAM achieves a reduction in leakage energy of 90% compared to SRAM, the dynamic energy for a write operation is 2X that of SRAM. Consequently, we propose additional microarchitectural optimizations to reduce overall dynamic energy which achieve an average reduction in dynamic energy over the base case of 30% with a range of 16% to 60% across 10 benchmarks. Mitchelle Rasquinha, Dhruv Choudhary, Subho Chatterjee, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
ISLPED | 5 |
| 2009 | A methodology for robust, energy efficient design of Spin-Torque-Transfer RAM arrays at scaled technologiesabstractIn this paper we propose a methodology for energy efficient Spin-Torque-Transfer Random Access Memory (STTRAM) array design at scaled technology nodes. We present a model to estimate and analyze the energy dissipation of an STTRAM array. The presented model shows the strong dependence of the array energy on the silicon transistor width, word line voltage and row/column organization. Using the array energy model we propose a design methodology for STTRAM arrays which minimizes the energy dissipation while maintaining the required robustness in read and write operations at scaled technologies. Subho Chatterjee, Mitchelle Rasquinha, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
ICCAD | 3 |
| 2008 | ShareStreams-V: A Virtualized QoS Packet Scheduling AcceleratorabstractThis paper introduces a virtualized FPGA-based accelerator for wire speed scheduling of packet streams under quality of service constraints. This work implements the dynamic window constrained scheduling algorithm and builds upon our previous custom accelerator by adding support for virtualization. This implementation is parametric, permitting tradeoffs between packet decision latency, decision throughput, and the number of virtual packet schedulers supported. When scheduling streams from multiple processes, ShareStreams-V 1 is able to schedule minimal size packets faster than one decision per 51.2 ns for up to 64 streams, the throughput required for 10 gbps Ethernet. The bottleneck currently is the host-accelerator HW/SW (PCIe) interface; this may be mitigated using high-speed interconnects/interfaces such as HyperTransport. Kangtao Kendall Chuang, Sudhakar Yalamanchili, Ada Gavrilovska, Karsten Schwan |
FCCM | 2 |
| 2008 | An Utilization Driven Framework for Energy Efficient Caches
Subramanian Ramaswamy, Sudhakar Yalamanchili |
HiPC | 2 |
| 2008 | Harmony: an execution model and runtime for heterogeneous many core systemsabstractThe emergence of heterogeneous many core architectures presents a unique opportunity for delivering order of magnitude performance increases to high performance applications by matching certain classes of algorithms to specifically tailored architectures. Their ubiquitous adoption, however, has been limited by a lack of programming models and management frameworks designed to reduce the high degree of complexity of software development intrinsic to heterogeneous architectures. This paper proposes Harmony, a runtime supported programming and execution model that provides: (1) semantics for simplifying parallelism management, (2) dynamic scheduling of compute intensive kernels to heterogeneous processor resources, and (3) online monitoring driven performance optimization for heterogeneous many core systems. We are particulably concerned with simplifying development and ensuring binary portability and scalability across system configurations and sizes. Initial results from ongoing development demonstrate the binary compatibility with variable number of cores, as well as dynamic adaptation of schedules to data sets. We present preliminary results of key features for some benchmark applications. Gregory Frederick Diamos, Sudhakar Yalamanchili |
HPDC | 2 |
| 2007 | Improving cache efficiency via resizing + remappingabstractIn this paper we propose techniques to dynamically downsize or upsize a cache accompanied by cache set/line shutdown to produce efficient caches. Unlike previous approaches, resizing is accompanied by a non-uniform remapping of memory into the resized cache, thus avoiding misses to sets/lines that are shut off. The paper first provides an analysis into the causes of energy inefficiencies revealing a simple model for improving efficiency. Based on this model we propose the concept of "folding" - memory regions mapping to disjoint cache resources are combined to share cache sets producing a new placement function. Folding enables powering down cache sets at the expense of possibly increasing conflict misses. Effective folding heuristics can substantially increase energy efficiency at the expense of acceptable increase in execution time. We target the 12 cache because of its larger size and greater energy consumption. Our techniques increase cache energy efficiency by 20%, and reduce the EDP (energy delay product) by up to 45% with an IPC degradation of less than 4%. The results also indicate opportunity for improving cache efficiencies further via cooperative compiler interactions. Subramanian Ramaswamy, Sudhakar Yalamanchili |
ICCD | 2 |
| 2006 | Customizable Fault Tolerant Caches for Embedded ProcessorsabstractThe continuing divergence of processor and memory speeds has led to the increasing reliance on larger caches which have become major consumers of area and power in embedded processors. Concurrently, intra-die and inter-die process variation at future technology nodes will cause defect-free yield to drop sharply unless mitigated. This paper focuses on an architectural technique to configure cache designs to be resilient to memory cell failures brought on by the effects of process variation. Profile-driven re-mapping of memory lines to cache lines is proposed to tolerate failures while minimizing degradation in average memory access time (AMAT) and thereby significantly boosting performance-based die yield beyond that which can be achieved with current techniques. For example, with 50% of the number of cache lines faulty, the performance drop quantified by increase in AMAT using our technique is 12.5% compared to 60% increase in AMAT using existing techniques. Subramanian Ramaswamy, Sudhakar Yalamanchili |
ICCD | 2 |
| 2006 | MMR: A MultiMedia Router architecture to support hybrid workloads
María Blanca Caminero, Carmen Carrión 0001, Francisco J. Quiles 0001, José Duato, Sudhakar Yalamanchili |
J. Parallel Distributed Comput. | 5 |
| 2005 | Traffic Scheduling Solutions with QoS Support for an Input-Buffered MultiMedia RouterabstractQuality of service (QoS) support in local and cluster area environments has become an issue of great interest in recent years. Most current high-performance interconnection solutions for these environments have been designed to enhance conventional best-effort traffic performance, but are not well-suited to the special requirements of the new multimedia applications. The multimedia router (MMR) aims at offering hardware-based QoS support within a compact interconnection component. One of the key elements in the MMR architecture is the algorithms used in traffic scheduling. These algorithms are responsible for the order in which information is forwarded through the internal switch. Thus, they are closely related to the QoS-provisioning mechanisms. In this paper, several traffic scheduling algorithms developed for the MMR architecture are described. Their general organization is motivated by chances for parallelization and pipelining, while providing the necessary support both to multimedia flows and to best-effort traffic. Performance evaluation results show that the QoS requirements of different connections are met, in spite of the presence of best-effort traffic, while achieving high link utilizations. María Blanca Caminero, Carmen Carrión 0001, Francisco J. Quiles 0001, José Duato, Sudhakar Yalamanchili |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2004 | ShareStreams: A Scalable Architecture and Hardware Support for High-Speed QoS Packet SchedulersabstractShareStreams (scalable hardware architectures for stream schedulers) is a unified hardware architecture for realizing a range of wire-speed packet scheduling disciplines for output link scheduling. This paper presents opportunities to exploit parallelism, design issues, tradeoffs and evaluation of the FPGA hardware architecture for use in switch network interfaces. The architecture uses processor resources for queuing and data movement and FPGA hardware resources for accelerating decisions and priority updates. The hardware architecture stores state in register base blocks, stream service attributes are compared using single-cycle decision blocks arranged in a novel single-stage recirculating network. The architecture provides effective mechanisms to trade hardware complexity for lower execution-time in a predictable manner. The hardware realized in a Virtex-I and Virtex-II FPGA can meet the packet-time requirements of 10 Gbps links for 256 stream queues with window-constrained scheduling disciplines. The hardware can schedule 1536 stream queues with priority-class/fair-queuing scheduling disciplines using 16 service-classes to meet 10 Gbps packet-times. Raj Krishnamurthy, Sudhakar Yalamanchili, Karsten Schwan, Richard West |
FCCM | 2 |
| 2002 | Algorithms for Switch-Scheduling in the Multimedia Router for LANs
Indrani Paul, Sudhakar Yalamanchili, José Duato |
HiPC | 2 |
| 2002 | The Customization Landscape for Embedded Systems
Sudhakar Yalamanchili |
HiPC | 1 |
| 2002 | A multimedia router architecture to provide high performance and QoS guarantees to mixed trafficabstractThe explosive growth in using scalable and cost-effective clusters and local area environments involve the design of high performance networks aimed at providing QoS to multimedia flows. Thus, the main goal pursued by the Multi-Media (MMR) project is to design a single-chip router able to efficiently handle multimedia flows and best-effort traffic. In this paper we focus on the performance evaluation of the MMR architecture using a mix of CBR, VBR and best effort workload. Preliminary simulation results show that, by using simple link and switch scheduling algorithms, the router is able to achieve a link bandwidth utilization of 80%, while still providing QoS guarantees to both CBR and VBR traffic in the presence of best-effort traffic. María Blanca Caminero, Carmen Carrión 0001, Francisco J. Quiles 0001, José Duato, Sudhakar Yalamanchili |
ICME (1) | 5 |
| 2001 | Tuning Buffer Size in the Multimedia Router (MMR)abstractThe primary objective of the Multimedia Router (MMR) project is the design and implementation of a compact router optimized for multimedia applications. The router is targeted for use in cluster and LAN interconnection networks, which offer different constraints and therefore differing router solutions than WANs. One of the key design parameters is the amount of buffer space, which is closely related to the silicon area required to implement the router. In this paper, the MMR performance obtained when varying the size of the input buffers is explored. Preliminary results show that buffers as small as one flit large suffice to guarantee QoS to both CBR and VBR traffic, thanks to the use of flow control and short links. María Blanca Caminero, Carmen Carrión 0001, Francisco J. Quiles 0001, José Duato, Sudhakar Yalamanchili |
IPDPS | 5 |
| 2000 | Switch Scheduling in the Multimedia Router (MMR)abstractThe primary goal of the Multimedia Router (MMR) project is the design and implementation of a router optimized for multimedia applications. The router is targeted for use in cluster and LAN interconnection networks which offer different constraints and therefore differing router solutions than WANs. This paper describes and evaluates a switch scheduling algorithm based on a priority biasing scheme for dynamically updating the priorities of the connections established through the router. Unlike existing schemes that simply use the age of a flit as its priority, the novel feature of the proposed approach is that the priority is biased using the measured quality of service (QoS) values for the connection. Furthermore, the structure of the switch scheduling algorithm is motivated by opportunities for pipelined and concurrent operation so that scheduling decisions could be made at switching speeds. The performance of two of the many possible biasing functions is evaluated. Damon S. Love, Sudhakar Yalamanchili, José Duato, María Blanca Caminero, Francisco J. Quiles 0001 |
IPDPS | 2 |
| 2000 | Software-Based Rerouting for Fault-Tolerant Pipelined CommunicationabstractThis paper presents a software-based approach to fault-tolerant routing in networks using wormhole or virtual cut-through switching. When a message encounters a faulty output link, it is removed from the network by the local router and delivered to the messaging layer of the local node's operating system. The message passing software can reroute this message, possibly along nonminimal paths. Alternatively, the message may be addressed to an intermediate node, which will forward the message to the destination. A message may encounter multiple faults and pass through multiple intermediate nodes. The proposed techniques are applicable to both obliviously and adaptively routed networks. The techniques are specifically targeted toward commercial multiprocessors where the mean time to repair (MTTR) is much smaller than the mean time between router failures (MTBF), i.e., it is sufficient to tolerate a maximum of three failures. This paper presents requirements for buffer management, deadlock freedom, and livelock freedom. Simulation results are presented to evaluate the degradation in latency and throughput as a function of the number and distribution of faults. There are several advantages of such an approach. Router designs are minimally impacted, and thus remain compact and fast. Only messages that encounter faulty components are affected, while the machine is ensured of continued operation until the faulty components can be replaced. The technique leverages existing network technology, and the concepts are portable across evolving switch and router designs. Therefore, we feel that the technique is a good candidate for incorporation into the next generation of multiprocessor networks. Young-Joo Suh, Binh Vien Dao, José Duato, Sudhakar Yalamanchili |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2000 | Configurable Algorithms for Complete Exchange in 2D MeshesabstractThe interprocessor complete exchange communication pattern can be found in many important parallel algorithms. In this paper, we present algorithms for complete exchange on 2D mesh-connected multiprocessors. The unique feature of the proposed algorithms is that they are configurable where the time for message startups can be traded against larger message sizes. At one extreme, the algorithm minimizes the number of message startups at the expense of an increased amount of time spent in message transmission. At the other extreme, the time spent in message transmission is reduced at the expense of an increased number of message startups. The structure of the algorithms is such that intermediate solutions are feasible, i.e., the number of message startups can be increased slightly and the message transmission time is correspondingly reduced. The ability to configure these algorithms enables the algorithm characteristics to be matched with machine characteristics based on specific overheads for message initiation and link speeds to minimize overall execution time. In effect, the algorithms can be configured to strike the right balance between direct and message combining approaches on a specific architecture for a given problem size. We believe these algorithms are distinguished by this ability and contribute to efficient portable implementations of complete exchange algorithms. Young-Joo Suh, Sudhakar Yalamanchili |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1999 | MMR: A High-Performance Multimedia Router - Architecture and Design Trade-OffsabstractThis paper presents the architecture of a router designed to efficiently support traffic generated by multimedia applications. The router is targeted for use in clusters and LANs rather than in WANs, the latter being served by communication substrates such as ATM. The distinguishing features of the proposed router architecture are the use of small fixed-size buffers, a large number of virtual channels, link-level virtual channel flow control, support for dynamic modification of connection bandwidth and priorities, and coordinated scheduling of connections across all output channels. The paper begins with a discussion of the design choices and architectural trade-offs made in the current MultiMedia Router (MMR) project. The performance evaluation section presents some preliminary results of the coordinated scheduling of constant bit rate (CBR) traffic streams. José Duato, Sudhakar Yalamanchili, María Blanca Caminero, Damon S. Love, Francisco J. Quiles 0001 |
HPCA | 2 |
| 1999 | Dynamically Configurable Message Flow Control for Fault-Tolerant RoutingabstractFault-tolerant routing protocols in modern interconnection networks rely heavily on the network flow control mechanisms used. Optimistic flow control mechanisms, such as wormhole switching (WS), realize very good performance, but are prone to deadlock in the presence of faults. Conservative flow control mechanisms, such as pipelined circuit switching (PCS), ensure the existence of a path to the destination prior to message transmission, achieving reliable transmission at the expense of performance. This paper proposes a general class of flow control mechanisms that can be dynamically configured to trade-off reliability and performance. Routing protocols can then be designed such that, in the vicinity of faults, protocols use a more conservative flow control mechanism, while the majority of messages that traverse fault-free portions of the network utilize a WS like flow control to maximize performance. We refer to such protocols as two-phase protocols. This ability provides new avenues for optimizing message passing performance in the presence of faults. A fully adaptive two-phase protocol is proposed, and compared via simulation to those based on WS and PCS. The architecture of a network router supporting configurable flow control is also described. Binh Vien Dao, José Duato, Sudhakar Yalamanchili |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1998 | All-To-All Communication with Minimum Start-Up Costs in 2D/3D Tori and MeshesabstractAll-to-all communication patterns occur in many important parallel algorithms. This paper presents new algorithms for all-to-all communication patterns (all-to-all broadcast and all-to-all personalized exchange) for wormhole switched 2D/3D torus- and mesh-connected multiprocessors. The algorithms use message combining to minimize message start-ups at the expense of larger message sizes. The unique feature of these algorithms is that they are the first algorithms that we know of that operate in a bottom-up fashion rather than a recursive, top-down manner. For a 2/sup d//spl times/2/sup d/ torus or mesh, the algorithms for all-to-all personalized exchange have time complexity of O(2/sup 3d/). An important property of the algorithms is the O(d) time due to message start-ups, compared with O(2/sup d/) for current algorithms. This is particularly important for modern parallel architectures where the start-up cost of message transmissions still dominates, except for very large block sizes. Finally, the 2D algorithms for all-to-all personalized exchange are extended to O(2/sup 4d/) algorithms in a 2/sup d//spl times/2/sup d//spl times/2/sup d/3D torus or mesh. These algorithms also retain the important property of O(d) time due to message start-ups. Young-Joo Suh, Sudhakar Yalamanchili |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1997 | Architectural Support for Reducing Communication Overhead in Multiprocessor Interconnection NetworksabstractModern multicomputer interconnection networks offer the delivery of messages with very low latency. However the message in-flight time is only a small portion of the total time that is required to send a message from source to destination. For fine to medium grained message sizes, the majority of time is spent in overheads for setting up and managing message transmission. It is often possible for compilers/programmers to separate inter-processor communication traffic into messages that exhibit communication locality and messages that do not. This paper proposes architectural modifications to network interfaces and routers to enable compilers/programmers to exploit known locality properties of programs in reducing the fixed overhead of transmission. These techniques work well on traffic exhibiting communication locality without unduly penalizing "ordinary" message traffic. The proposed techniques are evaluated using communication traces from 5 application program kernels. Significant reductions in average message latency are possible, and we argue that the approach can be used in the next generation of cluster interconnects. Binh Vien Dao, Sudhakar Yalamanchili, José Duato |
HPCA | 2 |
| 1997 | Power Constrained Design of Multiprocessor Interconnection NetworksabstractThe paper considers the power constrained design of orthogonal multiprocessor interconnection networks. The authors present a detailed model of message latency as a function of topology, technology architecture, and power. This model is then used to analyze a number of interesting scenarios, providing a sound engineering basis for interconnection network design in these cases. For example, they have observed that under a fixed power constraint, the network dimension which achieves minimal latency is a slowly growing function of system size. In addition, as they increase the available power per node for a fixed system size, the dimension at which message latency is minimized shifts towards higher dimensional networks. Chirag S. Patel, Sek M. Chai, Sudhakar Yalamanchili, David E. Schimmel |
ICCD | 3 |
| 1997 | On adaptive resource allocation for complex real-time applicationabstractResource allocation for high-performance real-time applications is challenging due to the applications' data-dependent nature, dynamic changes in their external environment, and limited resource availability in their target embedded system platforms. These challenges may be met by use of adaptive resource allocation (ARA) mechanisms that can promptly adjust resource allocation to changes in an application's resource needs, whenever there is a risk of failing to satisfy its timing constraints. By taking advantage of an application's adaptation capabilities, ARA eliminates the need for 'over-sizing' real-time systems to meet worst-case application needs. This paper proposes a model for describing an application's adaptation capabilities and the runtime variation of its resource needs. The paper also proposes a satisfiability-driven set of performance metrics for capturing the impact of ARA mechanisms on the performance of adaptable real-time applications. The relevance of the proposed set of metrics is demonstrated experimentally, using a synthetic application designed to represent time-critical applications in C31 systems. Daniela Rosu 0001, Karsten Schwan, Sudhakar Yalamanchili, Rakesh Jha |
RTSS | 3 |
| 1996 | Adaptive resource allocation for embedded parallel applicationsabstractParallel and distributed computer architectures are increasingly being considered for application in a wide variety of computationally intensive embedded systems. Many such applications impose highly dynamic demands for resources (processors, memory, and communication network), because their computations are data-dependent, or because the applications must constantly interact with a rapidly changing physical environment, or because the applications themselves are adaptive. This paper presents a set of dynamic resource allocation techniques aimed at maintaining high levels of application performance in the presence of varying resource demands. It focuses on a class of applications structured as multiple pipelines of data-parallel stages, as this structure is common to many sensor-based applications. We discuss the issues involved in resource management for such applications, and present preliminary results from our implementations on Intel Paragon. Our approach uses feedback control-a real-time monitoring system is used to detect significant performance shortfalls, and resources are reallocated among the application components in an attempt to improve performance. The main contribution of this work is that it combines real-time monitoring of an application's performance with dynamic resource allocation, and focuses on practical implementations rather than simulation and analysis. Rakesh Jha, Mustafa Muhammad, Sudhakar Yalamanchili, Karsten Schwan, Daniela Rosu 0001, Chris deCastro |
HiPC | 3 |
| 1996 | Distributed Deadlock-Free Routing in Faulty, Pipelined, Direct Interconnection NetworksabstractThis paper focuses on designing high performance pipelined networks that can operate in the presence of dynamic component failures. A general, rigorous framework for deadlock-free communication in faulty, pipelined networks is developed. A mechanism is also proposed for recovering from dynamic link and node failures. The recovery mechanism (1) is fully distributed, (2) does not require timeouts, (3) prevents fault-induced deadlock, and (4) is integrated into the virtual channel flow control mechanisms. This recovery mechanism is used to develop a new pipelined communication mechanism-acknowledged pipelined circuit-switching (APCS). This mechanism supports existing routing protocols that can tolerate a maximal number of static link failures, i.e., one less than the number of ports on a node. An implementation of a novel router architecture is described and the results of detailed flit level simulations are presented. Finally, the proposed recovery mechanism is shown to be applicable to existing adaptive wormhole routing protocols which are prone to deadlock in the presence of dynamic faults. Patrick T. Gaughan, Binh Vien Dao, Sudhakar Yalamanchili, David E. Schimmel |
IEEE Trans. Computers | 3 |
| 1996 | Augmented Binary Hypercube: A New Architecture for Processor ManagementabstractAugmented Binary Hypercube (AH) architecture consists of the binary hypercube processor nodes (PNs) and a hierarchy of management nodes (MNs). Several distributed algorithms maintain subcube information at the MNs to realize fault tolerant, fragmentation free processor allocation and load balancing. For efficient implementation of AH, we map MNs onto PNs, define and prove infeasibility of ideal mappings. We propose easily implementable nonoptimal mappings, having negligible overheads on performance. Extensive simulation studies and performance analysis conclude that these algorithms realize significantly better average job completion time and higher processor utilization, as compared to the best sequential allocation schemes and parallel implementation of Free List. AH algorithms can be tuned or adapt to the job and system characteristics, and resource management traffic. Hari Lalgudi, Ian F. Akyildiz, Sudhakar Yalamanchili |
IEEE Trans. Computers | 3 |
| 1995 | Software Based Fault-Tolerant Oblivious Routing in Pipelined Networks
Young-Joo Suh, Binh Vien Dao, José Duato, Sudhakar Yalamanchili |
ICPP (1) | 4 |
| 1995 | Configurable Flow Control Mechanisms for Fault-Tolerant RoutingabstractFault-tolerant routing protocols in modern interconnection networks rely heavily on the network flow control mechanisms used. Optimistic flow control mechanisms such as wormhole routing (WR) realize very good performance, but are prone to deadlock in the presence of faults. Conservative flow control mechanisms such as pipelined circuit switching (PCS) insures existence of a path to the destination prior to message transmission, but incurs increased overhead. Existing fault-tolerant routing protocols are designed with one or the other, and must accommodate their associated constraints. This paper proposes the use of configurable flow control mechanisms. Routing protocols can then be designed such that in the vicinity of faults, protocols use a more conservative flow control mechanism, while the majority of messages that traverse fault-free portions of the network utilize a WR like flow control to maximize performance. Such protocols are referred to as two-phase protocols, where routing decisions are provided some control over the operation of the virtual channels. This ability provides new avenues for optimizing message passing performance in the presence of faults. A fully adaptive two-phase protocol is proposed and compared via simulation to those based on WR and PCS. The architecture of a network router supporting configurable flow control is described, and the paper concludes with avenues for future research. Binh Vien Dao, José Duato, Sudhakar Yalamanchili |
ISCA | 3 |
| 1995 | Partitioning and mapping in embedded multiprocessor architectures in the presence of constraintsabstractAbstract The paper focuses on the problem of partitioning and mapping parallel programs onto heterogeneous embedded multiprocessor architectures for real‐time applications. Such applications present unique constraints and challenges. In addition to heterogeneity, the proposed partitioning and mapping algorithms satisfy memory, task throughput, task placement, intertask communication bandwidth, and co‐location constraints. They do so for architectures that utilize circuit‐switched (rather than packet‐switched) interprocessor communication and optimize latency and throughput in addition to load‐balancing. Finally, these mapping algorithms make use of knowledge of the local scheduling discipline to accommodate real‐time scheduling constraints. Our focus is on unstructured parallel programs that fall into one of two classes: (i) the class of computations characteristic of control applications in a real‐time environment where tasks execute concurrently, periodically exchanging information, and (ii) pipelined computation graphs found in sensor data processing applications. The algorithms are implemented in a set of tools that operate with commercial CASE tools at one end, and present an interface to multiprocessor simulators at the other end. Collectively, the algorithms form a significant component of an interactive design environment for the development and mapping of real‐time embedded parallel programs. The paper describes the algorithms, the encapsulating toolset, and presents an example of their application to an existing embedded application—an Autonomous Underwater Vehicle application. Sudhakar Yalamanchili, Lynn E. Te Winkel, David L. Perschbacher, Belle Shenoy |
Concurr. Pract. Exp. | 1 |
| 1995 | A Performance Model of Pipelined K-ary n-cubesabstractPipelined communication using virtual channels can realize low latency, high throughput, inter-processor communication. This paper presents an analytic performance model of pipelined communication in k-ary n cubes. The model contains elements intended to capture and study key performance issues. In addition to the modeling of throughput and latency, the following issues are addressed using this model: (1) the tradeoff between full-duplex vs. half-duplex links, (2) the effects of intranode delay, (3) the effects of buffer size for each virtual channel. Detailed simulation experiments under a variety of conditions establish the viability of this model.> Patrick T. Gaughan, Sudhakar Yalamanchili |
IEEE Trans. Computers | 2 |
| 1995 | A Family of Fault-Tolerant Routing Protocols for Direct Multiprocessor NetworksabstractOur goal is to reconcile the conflicting demands of performance and fault-tolerance in interprocessor communication. To this end, we propose a pipelined communication mechanism-pipelined circuit-switching (PCS)-which is a variant of the well known wormhole routing (WR) mechanism. PCS relaxes some of the routing constraints imposed by WR and as a result enables routing behavior that cannot otherwise be realized. This paper presents a new class of adaptive routing algorithms-misrouting backtracking with m misroutes (MB-m). This class of routing algorithms is made possible by PCS. We provide an analysis of the performance and static fault-tolerant properties of MB-m. The results of an experimental evaluation of PCS and MB-3 are also presented. This methodology provides performance approaching that of WR, while realizing a level of resilience to static faults that is difficult to achieve with WR.> Patrick T. Gaughan, Sudhakar Yalamanchili |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1994 | Scouting: Fully Adaptive, Deadlock-Free Routing in Faulty Pipelined NetworksabstractAdaptive routing protocols based on message pipelining using wormhole routing (WR) can provide superior performance. However, the occurrence of faults can lead to situations that may produce deadlock. Variants of adaptive WR have been introduced (P.T. Gaughan and S. Yalamanchili, 1992) that employ backtracking and misrouting to first establish a path, followed by message pipelining (pipelined circuit switching, or PCS). This scheme avoids deadlock due to faults, but is overly conservative leading to reduced performance. The paper introduces a new family of flow control mechanisms ranging from WR to PCS that offers a compromise by only decoupling the routing probe and the data fits the minimal extent required to provide deadlock-free routing in the presence of faults. José Duato, V. B. Dao, Patrick T. Gaughan, Sudhakar Yalamanchili |
ICPADS | 4 |
| 1994 | Ariadne - An Adaptive Router for Fault-Tolerant MulticomputersabstractAdaptive routing has been proposed as a means of improving performance and fault-tolerance in multicomputer networks. While a number of algorithms have been proposed, few adaptive routers have been implemented in hardware. This paper presents the design and implementation of Ariadne/spl minus/a prototype single chip, hardware router. The primary motivation is tolerance to link and router failures, while reconciling conflicting demands on performance. This is achieved by implementing the m-misroute backtracking protocol (MB-m) using the pipelined circuit-switching communication mechanism. Ariadne implements two virtual data channels and one virtual control channel per physical link. The router is self-timed with single flit buffering at the input and output ports, and is fully adaptive.> James D. Allen, Patrick T. Gaughan, David E. Schimmel, Sudhakar Yalamanchili |
ISCA | 4 |
| 1994 | Large Join Optimization on a Hypercube MultiprocessorabstractOptimizing large join queries that consist of many joins has been recognized as NP-hard. Most of the previous work focuses on a uniprocessor environment. In a multiprocessor, the location of each join adds another dimension to the complexity of the problem. In this paper, we examine the feasibility of exploiting the inherent parallelism in optimizing large join queries on a hypercube multiprocessor. This includes using the multiprocessor not only to answer the large join query but also to optimize it. We propose an algorithm to estimate the cost of a parallel large join plan. Three heuristics are provided for generating an initial solution, which is further optimized by an iterative local-improvement method. The entire process of parallel query optimization and execution is simulated on an Intel iPSC/2 hypercube machine. Our experimental results show that the performance of each heuristic depends on the characteristics of the query.> Eileen Tien Lin, Edward Omiecinski, Sudhakar Yalamanchili |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1987 | Parallel image normalization on a mesh connected array processor
Sudhakar Yalamanchili, Jake K. Aggarwal |
Pattern Recognit. | 2 |
| 1987 | A Characterization and Analysis of Parallel Processor Interconnection NetworksabstractThe permuting properties of various interconnection networks have been extensively studied. However, not too much attention has been focused on how the permuting properties interact with the mapping of tasks to processors in realizing the communication requirements between tasks. In this paper we focus on characterizing the abilities of some interconnection networks in realizing intertask communication that can be specified as permutations of the task names. From the point of view of the intertask communications requirements, the perceived permuting capabilities may depend upon the specific assignment of tasks to processors. Distinct network permutations may actually result in equivalent intertask communication patterns depending upon the mapping of tasks to processors. Characterizations of networks are developed based upon the theory of permutation groups. A number of properties as well as limitations of these networks become evident from this characterization. Finally, a class of switching networks is identified, that possess many useful properties that make them preferable to multistage interconnection networks in specific applications. Sudhakar Yalamanchili, Jake K. Aggarwal |
IEEE Trans. Computers | 1 |
| 1985 | Analysis of a model for parallel image processing
Sudhakar Yalamanchili, Jake K. Aggarwal |
Pattern Recognit. | 1 |
| 1985 | A system organization for parallel image processing
Sudhakar Yalamanchili, Jake K. Aggarwal |
Pattern Recognit. | 1 |
| 1984 | Algebraic Properties of some Parallel Processor Interconnection NetworksabstractInterconnection networks form an integral part of parallel processing systems, providing the facility for communication between several concurrently executing tasks. This paper presents an algebraic characterization of the sets of communication paths that a network can simultaneously establish between the processing elements of the system. This characterization provides a uniform framework for the analysis of interconnection networks and is shown to yield answers to several important questions concerning the capabilities, limitations and operation of these networks. The networks considered include one way and two way shift register rings, near neighbor meshes, single stage switching networks and briefly, multistage switching networks. Sudhakar Yalamanchili, Jake K. Aggarwal |
ICDE | 1 |
| 1984 | Formulation of parallel image processing tasks
Sudhakar Yalamanchili, Jake K. Aggarwal |
Pattern Recognit. Lett. | 1 |
| 1982 | Extraction of moving object descriptions via differencing
Sudhakar Yalamanchili, Worthy N. Martin, Jake K. Aggarwal |
Comput. Graph. Image Process. | 1 |