EDBT 2026 Demo / reviewers in the wild / expert
Ronny Krashinsky
dblp:86/4953
· DBLP profile ↗
9ranked-venue papers
4as first author
0since 2021 · last 2014
0009-0008-7949-4942ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-authorComputer networks · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Processor architecture and microarchitecture · 45% Memory systems · 33% GPUs and heterogeneous computing · 11% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 62% Transport protocols and congestion control · 19% Network measurement and analytics · 19% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › accelerator compilation
GPU compiler |
0.2 | 1 | 2014 | Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures · MICRO 2014 |
Processor architecture and microarchitecture
data-parallel architecture |
0.2 | 1 | 2014 | Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures · MICRO 2014 |
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution |
0.2 | 1 | 2014 | Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures · MICRO 2014 |
Memory systems
cache |
0.1 | 1 | 2012 | Unifying Primary Cache, Scratch, and Register File Memories in a Throughput Processor · MICRO 2012 |
Memory systems › on-chip memory
scratchpad memory |
0.1 | 1 | 2012 | Unifying Primary Cache, Scratch, and Register File Memories in a Throughput Processor · MICRO 2012 |
Processor architecture and microarchitecture
SIMD |
0.1 | 1 | 2014 | Exploring the Design Space of SPMD Divergence Management on Data-Parallel Architectures · MICRO 2014 |
Memory systems
cache design |
0.0 | 1 | 2004 | Cache Refill/Access Decoupling for Vector Machines · MICRO 2004 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2004 | The Vector-Thread Architecture · ISCA 2004 |
Memory systems › cache
prefetching |
0.0 | 1 | 2004 | Cache Refill/Access Decoupling for Vector Machines · MICRO 2004 |
Processor architecture and microarchitecture › vector processing
vector memory access |
0.0 | 1 | 2004 | Cache Refill/Access Decoupling for Vector Machines · MICRO 2004 |
Memory systems › cache › prefetching
vector prefetching |
0.0 | 1 | 2004 | Cache Refill/Access Decoupling for Vector Machines · MICRO 2004 |
Processor architecture and microarchitecture
vector processor |
0.0 | 1 | 2004 | Cache Refill/Access Decoupling for Vector Machines · MICRO 2004 |
Processor architecture and microarchitecture › vector processor
vector-thread architecture |
0.0 | 1 | 2004 | The Vector-Thread Architecture · ISCA 2004 |
Internet of things and sensor networks › energy management
power management |
0.0 | 1 | 2002 | Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002 |
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff |
0.0 | 1 | 2002 | Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002 |
Energy-efficient computing › energy-efficient communication
wireless network energy management |
0.0 | 1 | 2002 | Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002 |
Embedded and real-time systems › embedded processor
low-power embedded processor |
0.0 | 1 | 2004 | The Vector-Thread Architecture · ISCA 2004 |
Transport protocols and congestion control
TCP performance |
0.0 | 1 | 2002 | Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002 |
Network measurement and analytics › web performance measurement
web access latency |
0.0 | 1 | 2002 | Minimizing energy for wireless web access with bounded slowdown · MobiCom 2002 |
Methods — techniques the papers use, named apart from their topics
predication-based divergence management · 0.4compiler analysis · 0.4dynamic partitioning · 0.1trace-driven simulation · 0.1power management protocol · 0.1simulation · 0.0decoupled access/execute · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Exploring the Design Space of SPMD Divergence Management on Data-Parallel ArchitecturesabstractData-parallel architectures must provide efficient support for complex control-flow constructs to support sophisticated applications coded in modern single-program multiple-data languages. As these architectures have wide data paths that process a single instruction across parallel threads, a mechanism is needed to track and sequence threads as they traverse potentially divergent control paths through the program. The design space for divergence management ranges from software-only approaches where divergence is explicitly managed by the compiler, to hardware solutions where divergence is managed implicitly by the micro architecture. In this paper, we explore this space and propose a new predication-based approach for handling control-flow structures in data-parallel architectures. Unlike prior predication algorithms, our new compiler analyses and hardware instructions consider the commonality of predication conditions across threads to improve efficiency. We prototype our algorithms in a production compiler and evaluate the tradeoffs between software and hardware divergence management on current GPU silicon. We show that our compiler algorithms make a predication-only architecture competitive in performance to one with hardware support for tracking divergence. Yunsup Lee, Vinod Grover, Ronny Krashinsky, Mark Stephenson, Stephen W. Keckler, Krste Asanovic |
MICRO | 3 |
| 2013 | Convergence and scalarization for data-parallel architecturesabstractModern throughput processors such as GPUs achieve high performance and efficiency by exploiting data parallelism in application kernels expressed as threaded code. One draw-back of this approach compared to conventional vector architectures is redundant execution of instructions that are common across multiple threads, resulting in energy inefficiency due to excess instruction dispatch, register file accesses, and memory operations. This paper proposes to alleviate these overheads while retaining the threaded programming model by automatically detecting the scalar operations and factoring them out of the parallel code. We have developed a scalarizing compiler that employs convergence and variance analyses to statically identify values and instructions that are invariant across multiple threads. Our compiler algorithms are effective at identifying convergent execution even in programs with arbitrary control flow, identifying two-thirds of the opportunity captured by a dynamic oracle. The compile-time analysis leads to a reduction in instructions dispatched by 29%, register file reads and writes by 31% memory address counts by 47%, and data access counts by 38%. Yunsup Lee, Ronny Krashinsky, Vinod Grover, Stephen W. Keckler, Krste Asanovic |
CGO | 2 |
| 2012 | Unifying Primary Cache, Scratch, and Register File Memories in a Throughput ProcessorabstractModern throughput processors such as GPUs employ thousands of threads to drive high-bandwidth, long-latency memory systems. These threads require substantial on-chip storage for registers, cache, and scratchpad memory. Existing designs hard-partition this local storage, fixing the capacities of these structures at design time. We evaluate modern GPU workloads and find that they have widely varying capacity needs across these different functions. Therefore, we propose a unified local memory which can dynamically change the partitioning among registers, cache, and scratchpad on a per-application basis. The tuning that this flexibility enables improves both performance and energy consumption, and broadens the scope of applications that can be efficiently executed on GPUs. Compared to a hard-partitioned design, we show that unified local memory provides a performance benefit as high as 71% along with an energy reduction up to 33%. Mark Gebhart, Stephen W. Keckler, Brucek Khailany, Ronny Krashinsky, William J. Dally |
MICRO | 4 |
| 2008 | Implementing the scale vector-thread processorabstractThe Scale vector-thread processor is a complexity-effective solution for embedded computing which flexibly supports both vector and highly multithreaded processing. The 7.1-million transistor chip has 16 decoupled execution clusters, vector load and store units, and a nonblocking 32KB cache. An automated and iterative design and verification flow enabled a performance-, power-, and area-efficient implementation with two person-years of development effort. Scale has a core area of 16.6 mm 2 in 180 nm technology, and it consumes 400 mW--1.1 W while running at 260 MHz. Ronny Krashinsky, Christopher Batten, Krste Asanovic |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2007 | Activity-Sensitive Flip-Flop and Latch Selection for Reduced EnergyabstractThis paper presents new techniques to evaluate the energy and delay of flip-flop and latch designs and shows that no single existing design performs well across the wide range of operating regimes present in complex systems. We propose the use of a selection of flip-flop and latch designs, each tuned for different activation patterns and speed requirements. We illustrate our technique on a pipelined MIPS processor datapath running SPECint95 benchmarks, where we reduce total flip-flop and latch energy by over 60% without increasing cycle time. Seongmoo Heo, Ronny Krashinsky, Krste Asanovic |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Minimizing Energy for Wireless Web Access with Bounded Slowdown
Ronny Krashinsky, Hari Balakrishnan |
Wirel. Networks | 1 |
| 2004 | The Vector-Thread ArchitectureabstractThe vector-thread (VT) architectural paradigm unifies the vector and multithreaded compute models. The VT abstraction provides the programmer with a control processor and a vector of virtual processors (VPs). The control processor can use vector-fetch commands to broadcast instructions to all the VPs or each VP can use thread-fetches to direct its own control flow. A seamless intermixing of the vector and threaded control mechanisms allows a VT architecture to flexibly and compactly encode application parallelism and locality, and a VT machine exploits these to improve performance and efficiency. We present SCALE, an instantiation of the VT architecture designed for low-power and high-performance embedded systems. We evaluate the SCALE prototype design using detailed simulation of a broad range of embedded applications and show that its performance is competitive with larger and more complex processors. Ronny Krashinsky, Christopher Batten, Mark Hampton, Steve Gerding, Brian Pharris, Jared Casper, Krste Asanovic |
ISCA | 1 |
| 2004 | Cache Refill/Access Decoupling for Vector MachinesabstractVector processors often use a cache to exploit temporal locality and reduce memory bandwidth demands, but then require expensive logic to track large numbers of outstanding cache misses to sustain peak bandwidth from memory. We present refill/access decoupling, which augments the vector processor with a Vector Refill Unit (VRU) to quickly pre-execute vector memory commands and issue any needed cache line refills ahead of regular execution. The VRU reduces costs by eliminating much of the outstanding miss state required in traditional vector architectures and by using the cache itself as a cost-effective prefetch buffer. We also introduce vector segment accesses, a new class of vector memory instructions that efficiently encode two-dimensional access patterns. Segments reduce address bandwidth demands and enable more efficient refill/access decoupling by increasing the information contained in each vector memory command. Our results show that refill/access decoupling is able to achieve better performance with less resources than more traditional decoupling methods. Even with a small cache and memory latencies as long as 800 cycles, refill/access decoupling can sustain several kilobytes of in-flight data with minimal access management state and no need for expensive reserved element buffering. Christopher Batten, Ronny Krashinsky, Steve Gerding, Krste Asanovic |
MICRO | 2 |
| 2002 | Minimizing energy for wireless web access with bounded slowdownabstractOn many battery-powered mobile computing devices, the wireless network is a significant contributor to the total energy consumption. In this paper, we investigate the interaction between energy-saving protocols and TCP performance for Web like transfers. We show that the popular IEEE 802.11 power-saving mode (PSM), a protocol, can harm performance by increasing fast round trip times (RTTs) to 100 ms; and that under typical Web browsing workloads, current implementations will unnecessarily spend energy waking up during long idle periods.To overcome these problems, we present the Bounded-Slowdown (BSD) protocol, a PSM that dynamically adapts to network activity. BSD is an optimal solution to the problem of minimizing energy consumption while guaranteeing that a connection's RTT does not increase by more than a factor p over its base RTT, where p is a protocol parameter that exposes the trade-off between minimizing energy and reducing latency. works by staying awake for a short period of time after the link idle. We present several trace-driven simulation results that show that, compared to a static PSM, the Bounded Slowdown protocol reduces average Web page retrieval times by 5--64%, while simultaneously reducing energy consumption by 1--14% (and by 13X compared to no power management). Ronny Krashinsky, Hari Balakrishnan |
MobiCom | 1 |