Jai Menon 0003

dblp:89/6512-3 · also Jaikrishnan Menon · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Processor architecture and microarchitecture · 39% Energy-efficient computing · 16% GPUs and heterogeneous computing · 16%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction set architecture
0.422015
ISA Wars: Understanding the Relevance of ISA being RISC or CISC to Performance, Power, and Energy on Modern Architectures · ACM Trans. Comput. Syst. 2015
Power struggles: Revisiting the RISC vs. CISC debate on contemporary ARM and x86 architectures · HPCA 2013
Processor architecture and microarchitecture › instruction set architecture › instruction set style
RISC vs CISC
0.422015
ISA Wars: Understanding the Relevance of ISA being RISC or CISC to Performance, Power, and Energy on Modern Architectures · ACM Trans. Comput. Syst. 2015
Power struggles: Revisiting the RISC vs. CISC debate on contemporary ARM and x86 architectures · HPCA 2013
GPUs and heterogeneous computing
GPU architecture
0.422015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
iGPU: Exception support and speculative execution on GPUs · ISCA 2012
Electronic design automation › hardware verification and test › hardware verification
RTL simulation
0.212015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Performance modeling and evaluation
simulation
0.212015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing › energy-efficient architecture
processor energy efficiency
0.212013
Power struggles: Revisiting the RISC vs. CISC debate on contemporary ARM and x86 architectures · HPCA 2013
Processor architecture and microarchitecture
speculative execution
0.112012
iGPU: Exception support and speculative execution on GPUs · ISCA 2012
Electronic design automation › hardware verification and test
design validation
0.112015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Electronic design automation
hardware verification and test
0.112015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Runtime systems and virtual machines
dynamic compilation
0.012012
iGPU: Exception support and speculative execution on GPUs · ISCA 2012

Methods — techniques the papers use, named apart from their topics

idempotent code regions · 0.3measurement-based study · 0.2benchmarking · 0.2RTL implementation · 0.2OpenCL · 0.2hardware measurement · 0.2benchmark analysis · 0.2
YearPublicationVenuePosition
2015 MIAOW: An open source GPGPU
Vinay Gangadhar, Raghuraman Balasubramanian, Mario Drumond, Ziliang Guo, Jai Menon 0003, Cherin Joseph, Robin Prakash, Sharath Prasad, Pradip Valathol, Karthikeyan Sankaralingam
Hot Chips Symposium5
2015 Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU
abstract
Graphic processing unit (GPU)-based general-purpose computing is developing as a viable alternative to CPU-based computing in many domains. Today’s tools for GPU analysis include simulators like GPGPU-Sim, Multi2Sim, and Barra. While useful for modeling first-order effects, these tools do not provide a detailed view of GPU microarchitecture and physical design. Further, as GPGPU research evolves, design ideas and modifications demand detailed estimates of impact on overall area and power. Fueled by this need, we introduce MIAOW (Many-core Integrated Accelerator Of Wisconsin), an open-source RTL implementation of the AMD Southern Islands GPGPU ISA, capable of running unmodified OpenCL-based applications. We present our design motivated by our goals to create a realistic, flexible, OpenCL-compatible GPGPU, capable of emulating a full system. We first explore if MIAOW is realistic and then use four case studies to show that MIAOW enables the following: physical design perspective to “traditional” microarchitecture, new types of research exploration, and validation/calibration of simulator-based characterization of hardware. The findings and ideas are contributions in their own right, in addition to MIAOW’s utility as a tool for others’ research.
Raghuraman Balasubramanian, Vinay Gangadhar, Ziliang Guo, Chen-Han Ho, Cherin Joseph, Jai Menon 0003, Mario Drumond, Robin Paul, Sharath Prasad, Pradip Valathol, Karthikeyan Sankaralingam
ACM Trans. Archit. Code Optim.6
2015 ISA Wars: Understanding the Relevance of ISA being RISC or CISC to Performance, Power, and Energy on Modern Architectures
abstract
RISC versus CISC wars raged in the 1980s when chip area and processor design complexity were the primary constraints and desktops and servers exclusively dominated the computing landscape. Today, energy and power are the primary design constraints and the computing landscape is significantly different: Growth in tablets and smartphones running ARM (a RISC ISA) is surpassing that of desktops and laptops running x86 (a CISC ISA). Furthermore, the traditionally low-power ARM ISA is entering the high-performance server market, while the traditionally high-performance x86 ISA is entering the mobile low-power device market. Thus, the question of whether ISA plays an intrinsic role in performance or energy efficiency is becoming important again, and we seek to answer this question through a detailed measurement-based study on real hardware running real applications. We analyze measurements on seven platforms spanning three ISAs (MIPS, ARM, and x86) over workloads spanning mobile, desktop, and server computing. Our methodical investigation demonstrates the role of ISA in modern microprocessors’ performance and energy efficiency. We find that ARM, MIPS, and x86 processors are simply engineering design points optimized for different levels of performance, and there is nothing fundamentally more energy efficient in one ISA class or the other. The ISA being RISC or CISC seems irrelevant.
Emily R. Blem, Jai Menon 0003, Thiruvengadam Vijayaraghavan, Karthikeyan Sankaralingam
ACM Trans. Comput. Syst.2
2014 Memory processing units
abstract
Presents a conference poster that addresses the technology of memory processing units. Some of the following topics are examined: current processing capabilities; MPU hardware; performance and energy output; and new trends in the industry.
Jai Menon 0003, Lorenzo De Carli, Vijayraghavan Thiruvengadam, Karthikeyan Sankaralingam, Cristian Estan
Hot Chips Symposium1
2013 Power struggles: Revisiting the RISC vs. CISC debate on contemporary ARM and x86 architectures
abstract
RISC vs. CISC wars raged in the 1980s when chip area and processor design complexity were the primary constraints and desktops and servers exclusively dominated the computing landscape. Today, energy and power are the primary design constraints and the computing landscape is significantly different: growth in tablets and smartphones running ARM (a RISC ISA) is surpassing that of desktops and laptops running x86 (a CISC ISA). Further, the traditionally low-power ARM ISA is entering the high-performance server market, while the traditionally high-performance x86 ISA is entering the mobile low-power device market. Thus, the question of whether ISA plays an intrinsic role in performance or energy efficiency is becoming important, and we seek to answer this question through a detailed measurement based study on real hardware running real applications. We analyze measurements on the ARM Cortex-A8 and Cortex-A9 and Intel Atom and Sandybridge i7 microprocessors over workloads spanning mobile, desktop, and server computing. Our methodical investigation demonstrates the role of ISA in modern microprocessors' performance and energy efficiency. We find that ARM and x86 processors are simply engineering design points optimized for different levels of performance, and there is nothing fundamentally more energy efficient in one ISA class or the other. The ISA being RISC or CISC seems irrelevant.
Emily R. Blem, Jai Menon 0003, Karthikeyan Sankaralingam
HPCA2
2012 iGPU: Exception support and speculative execution on GPUs
abstract
Since the introduction of fully programmable vertex shader hardware, GPU computing has made tremendous advances. Exception support and speculative execution are the next steps to expand the scope and improve the usability of GPUs. However, traditional mechanisms to support exceptions and speculative execution are highly intrusive to GPU hardware design. This paper builds on two related insights to provide a unified lightweight mechanism for supporting exceptions and speculation on GPUs. First, we observe that GPU programs can be broken into code regions that contain little or no live register state at their entry point. We then also recognize that it is simple to generate these regions in such a way that they are idempotent, allowing their entry points to function as program recovery points and enabling support for exception handling, fast context switches, and speculation, all with very low overhead. We call the architecture of GPUs executing these idempotent regions the iGPU architecture. The hardware extensions required are minimal and the construction of idempotent code regions is fully transparent under the typical dynamic compilation framework of GPUs. We demonstrate how iGPU exception support enables virtual memory paging with very low overhead (1% to 4%), and how speculation support enables circuit-speculation techniques that can provide over 25% reduction in energy.
Jai Menon 0003, Marc de Kruijf, Karthikeyan Sankaralingam
ISCA1