Ronny Krashinsky

dblp:86/4953 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
0since 2021 · last 2014
0009-0008-7949-4942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-authorComputer networks · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Processor architecture and microarchitecture · 45% Memory systems · 33% GPUs and heterogeneous computing · 11%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer networks
1 paper
Internet of things and sensor networks · 62% Transport protocols and congestion control · 19% Network measurement and analytics · 19%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › accelerator compilation
GPU compiler
0.212014
Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures · MICRO 2014
Processor architecture and microarchitecture
data-parallel architecture
0.212014
Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures · MICRO 2014
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.212014
Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures · MICRO 2014
Memory systems
cache
0.112012
Unifying Primary Cache, Scratch, and Register File Memories in a Throughput Processor · MICRO 2012
Memory systems › on-chip memory
scratchpad memory
0.112012
Unifying Primary Cache, Scratch, and Register File Memories in a Throughput Processor · MICRO 2012
Processor architecture and microarchitecture
SIMD
0.112014
Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures · MICRO 2014
Memory systems
cache design
0.012004
Cache Refill/Access Decoupling for Vector Machines · MICRO 2004
Parallel and multicore computing
parallel programming models
0.012004
The Vector-Thread Architecture · ISCA 2004
Memory systems › cache
prefetching
0.012004
Cache Refill/Access Decoupling for Vector Machines · MICRO 2004
Processor architecture and microarchitecture › vector processing
vector memory access
0.012004
Cache Refill/Access Decoupling for Vector Machines · MICRO 2004
Memory systems › cache › prefetching
vector prefetching
0.012004
Cache Refill/Access Decoupling for Vector Machines · MICRO 2004
Processor architecture and microarchitecture
vector processor
0.012004
Cache Refill/Access Decoupling for Vector Machines · MICRO 2004
Processor architecture and microarchitecture › vector processor
vector-thread architecture
0.012004
The Vector-Thread Architecture · ISCA 2004
Internet of things and sensor networks › energy management
power management
0.012002
Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff
0.012002
Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002
Energy-efficient computing › energy-efficient communication
wireless network energy management
0.012002
Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002
Embedded and real-time systems › embedded processor
low-power embedded processor
0.012004
The Vector-Thread Architecture · ISCA 2004
Transport protocols and congestion control
TCP performance
0.012002
Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002
Network measurement and analytics › web performance measurement
web access latency
0.012002
Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002

Methods — techniques the papers use, named apart from their topics

predication-based divergence management · 0.4compiler analysis · 0.4dynamic partitioning · 0.1trace-driven simulation · 0.1power management protocol · 0.1simulation · 0.0decoupled access/execute · 0.0
YearPublicationVenuePosition
2014 Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures
abstract
Data-parallel architectures must provide efficient support for complex control-flow constructs to support sophisticated applications coded in modern single-program multiple-data languages. As these architectures have wide data paths that process a single instruction across parallel threads, a mechanism is needed to track and sequence threads as they traverse potentially divergent control paths through the program. The design space for divergence management ranges from software-only approaches where divergence is explicitly managed by the compiler, to hardware solutions where divergence is managed implicitly by the micro architecture. In this paper, we explore this space and propose a new predication-based approach for handling control-flow structures in data-parallel architectures. Unlike prior predication algorithms, our new compiler analyses and hardware instructions consider the commonality of predication conditions across threads to improve efficiency. We prototype our algorithms in a production compiler and evaluate the tradeoffs between software and hardware divergence management on current GPU silicon. We show that our compiler algorithms make a predication-only architecture competitive in performance to one with hardware support for tracking divergence.
Yunsup Lee, Vinod Grover, Ronny Krashinsky, Mark Stephenson, Stephen W. Keckler, Krste Asanovic
MICRO3
2013 Convergence and scalarization for data-parallel architectures
abstract
Modern throughput processors such as GPUs achieve high performance and efficiency by exploiting data parallelism in application kernels expressed as threaded code. One draw-back of this approach compared to conventional vector architectures is redundant execution of instructions that are common across multiple threads, resulting in energy inefficiency due to excess instruction dispatch, register file accesses, and memory operations. This paper proposes to alleviate these overheads while retaining the threaded programming model by automatically detecting the scalar operations and factoring them out of the parallel code. We have developed a scalarizing compiler that employs convergence and variance analyses to statically identify values and instructions that are invariant across multiple threads. Our compiler algorithms are effective at identifying convergent execution even in programs with arbitrary control flow, identifying two-thirds of the opportunity captured by a dynamic oracle. The compile-time analysis leads to a reduction in instructions dispatched by 29%, register file reads and writes by 31% memory address counts by 47%, and data access counts by 38%.
Yunsup Lee, Ronny Krashinsky, Vinod Grover, Stephen W. Keckler, Krste Asanovic
CGO2
2012 Unifying Primary Cache, Scratch, and Register File Memories in a Throughput Processor
abstract
Modern throughput processors such as GPUs employ thousands of threads to drive high-bandwidth, long-latency memory systems. These threads require substantial on-chip storage for registers, cache, and scratchpad memory. Existing designs hard-partition this local storage, fixing the capacities of these structures at design time. We evaluate modern GPU workloads and find that they have widely varying capacity needs across these different functions. Therefore, we propose a unified local memory which can dynamically change the partitioning among registers, cache, and scratchpad on a per-application basis. The tuning that this flexibility enables improves both performance and energy consumption, and broadens the scope of applications that can be efficiently executed on GPUs. Compared to a hard-partitioned design, we show that unified local memory provides a performance benefit as high as 71% along with an energy reduction up to 33%.
Mark Gebhart, Stephen W. Keckler, Brucek Khailany, Ronny Krashinsky, William J. Dally
MICRO4
2008 Implementing the scale vector-thread processor
abstract
The Scale vector-thread processor is a complexity-effective solution for embedded computing which flexibly supports both vector and highly multithreaded processing. The 7.1-million transistor chip has 16 decoupled execution clusters, vector load and store units, and a nonblocking 32KB cache. An automated and iterative design and verification flow enabled a performance-, power-, and area-efficient implementation with two person-years of development effort. Scale has a core area of 16.6 mm 2 in 180 nm technology, and it consumes 400 mW--1.1 W while running at 260 MHz.
Ronny Krashinsky, Christopher Batten, Krste Asanovic
ACM Trans. Design Autom. Electr. Syst.1
2007 Activity-Sensitive Flip-Flop and Latch Selection for Reduced Energy
abstract
This paper presents new techniques to evaluate the energy and delay of flip-flop and latch designs and shows that no single existing design performs well across the wide range of operating regimes present in complex systems. We propose the use of a selection of flip-flop and latch designs, each tuned for different activation patterns and speed requirements. We illustrate our technique on a pipelined MIPS processor datapath running SPECint95 benchmarks, where we reduce total flip-flop and latch energy by over 60% without increasing cycle time.
Seongmoo Heo, Ronny Krashinsky, Krste Asanovic
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Minimizing Energy for Wireless Web Access with Bounded Slowdown
Ronny Krashinsky, Hari Balakrishnan
Wirel. Networks1
2004 The Vector-Thread Architecture
abstract
The vector-thread (VT) architectural paradigm unifies the vector and multithreaded compute models. The VT abstraction provides the programmer with a control processor and a vector of virtual processors (VPs). The control processor can use vector-fetch commands to broadcast instructions to all the VPs or each VP can use thread-fetches to direct its own control flow. A seamless intermixing of the vector and threaded control mechanisms allows a VT architecture to flexibly and compactly encode application parallelism and locality, and a VT machine exploits these to improve performance and efficiency. We present SCALE, an instantiation of the VT architecture designed for low-power and high-performance embedded systems. We evaluate the SCALE prototype design using detailed simulation of a broad range of embedded applications and show that its performance is competitive with larger and more complex processors.
Ronny Krashinsky, Christopher Batten, Mark Hampton, Steve Gerding, Brian Pharris, Jared Casper, Krste Asanovic
ISCA1
2004 Cache Refill/Access Decoupling for Vector Machines
abstract
Vector processors often use a cache to exploit temporal locality and reduce memory bandwidth demands, but then require expensive logic to track large numbers of outstanding cache misses to sustain peak bandwidth from memory. We present refill/access decoupling, which augments the vector processor with a Vector Refill Unit (VRU) to quickly pre-execute vector memory commands and issue any needed cache line refills ahead of regular execution. The VRU reduces costs by eliminating much of the outstanding miss state required in traditional vector architectures and by using the cache itself as a cost-effective prefetch buffer. We also introduce vector segment accesses, a new class of vector memory instructions that efficiently encode two-dimensional access patterns. Segments reduce address bandwidth demands and enable more efficient refill/access decoupling by increasing the information contained in each vector memory command. Our results show that refill/access decoupling is able to achieve better performance with less resources than more traditional decoupling methods. Even with a small cache and memory latencies as long as 800 cycles, refill/access decoupling can sustain several kilobytes of in-flight data with minimal access management state and no need for expensive reserved element buffering.
Christopher Batten, Ronny Krashinsky, Steve Gerding, Krste Asanovic
MICRO2
2002 Minimizing energy for wireless web access with bounded slowdown
abstract
On many battery-powered mobile computing devices, the wireless network is a significant contributor to the total energy consumption. In this paper, we investigate the interaction between energy-saving protocols and TCP performance for Web like transfers. We show that the popular IEEE 802.11 power-saving mode (PSM), a protocol, can harm performance by increasing fast round trip times (RTTs) to 100 ms; and that under typical Web browsing workloads, current implementations will unnecessarily spend energy waking up during long idle periods.To overcome these problems, we present the Bounded-Slowdown (BSD) protocol, a PSM that dynamically adapts to network activity. BSD is an optimal solution to the problem of minimizing energy consumption while guaranteeing that a connection's RTT does not increase by more than a factor p over its base RTT, where p is a protocol parameter that exposes the trade-off between minimizing energy and reducing latency. works by staying awake for a short period of time after the link idle. We present several trace-driven simulation results that show that, compared to a static PSM, the Bounded Slowdown protocol reduces average Web page retrieval times by 5--64%, while simultaneously reducing energy consumption by 1--14% (and by 13X compared to no power management).
Ronny Krashinsky, Hari Balakrishnan
MobiCom1