Arun Raghavan

dblp:61/4142 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
1since 2021 · last 2025
0000-0001-5339-3352ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Hardware accelerators and domain-specific architectures · 26% Memory systems · 22% Energy-efficient computing · 15%
Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 62% Indexing and storage engines · 38%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 84% Concurrent programming · 16%

Topics — the 22 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
Wireless Hearables With Programmable Speech AI Accelerators · MobiCom 2025
Indexing and storage engines
columnar storage
0.312018
RAPID: In-Memory Analytical Query Processing Engine with Extreme Performance per Watt · SIGMOD Conference 2018
Energy-efficient computing › power management › peak power management
computational sprinting
0.322013
Computational sprinting on a hardware/software testbed · ASPLOS 2013
Computational sprinting · HPCA 2012
Energy-efficient computing
thermal management
0.322013
Computational sprinting on a hardware/software testbed · ASPLOS 2013
Computational sprinting · HPCA 2012
Memory systems
data movement
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Cloud and datacenter computing
data movement acceleration
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Hardware accelerators and domain-specific architectures › domain-specific accelerator
data processing accelerator
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Memory systems › in-memory computing
in-memory data processing
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Query processing and optimization
SIMD vectorization
0.212015
Rethinking SIMD Vectorization for In-Memory Databases · SIGMOD Conference 2015
Memory systems
cache coherence
0.222010
Token tenure and PATCH: A predictive/adaptive token-counting hybrid · ACM Trans. Archit. Code Optim. 2010
Token tenure: PATCHing token counting using directory-based cache coherence · MICRO 2008
Memory systems › cache coherence
directory-based coherence
0.222010
Token tenure and PATCH: A predictive/adaptive token-counting hybrid · ACM Trans. Archit. Code Optim. 2010
Token tenure: PATCHing token counting using directory-based cache coherence · MICRO 2008
Program synthesis and code generation › concurrent program synthesis
protocol synthesis
0.212013
TRANSIT: specifying protocols with concolic snippets · PLDI 2013
Embedded and real-time systems
mobile computing
0.212013
Computational sprinting on a hardware/software testbed · ASPLOS 2013
Processor architecture and microarchitecture
multicore design
0.112012
Computational sprinting · HPCA 2012
Distributed systems › concurrency control
optimistic concurrency control
0.112010
RETCON: transactional repair without replay · ISCA 2010
Memory systems › cache coherence
scalable coherence protocol
0.112010
Token tenure and PATCH: A predictive/adaptive token-counting hybrid · ACM Trans. Archit. Code Optim. 2010
Parallel and multicore computing
transactional memory
0.112010
RETCON: transactional repair without replay · ISCA 2010
Parallel and multicore computing › transactional memory
transactional memory scalability
0.112010
RETCON: transactional repair without replay · ISCA 2010
Processor architecture and microarchitecture
many-core architecture
0.112017
A many-core architecture for in-memory data processing · MICRO 2017
Embedded and real-time systems › mobile computing
mobile computing platforms
0.012012
Computational sprinting · HPCA 2012
Concurrent programming
synchronization
0.012010
RETCON: transactional repair without replay · ISCA 2010
Performance modeling and evaluation › parallel performance evaluation
multicore scalability
0.012008
Token tenure: PATCHing token counting using directory-based cache coherence · MICRO 2008

Methods — techniques the papers use, named apart from their topics

deep learning · 0.9hardware-software co-design · 0.7model checking · 0.5extended finite state machine · 0.5constraint solving · 0.5concolic execution · 0.5hash table · 0.4SIMD gather/scatter · 0.4hardware RPC · 0.3phase-change material thermal capacitance · 0.1workload analysis · 0.1
YearPublicationVenuePosition
2025 Wireless Hearables With Programmable Speech AI Accelerators
abstract
The conventional wisdom has been that designing ultra-compact, battery-constrained wireless hearables with on-device speech AI models is challenging due to the high computational demands of streaming deep learning models. Speech AI models require continuous, real-time audio processing, imposing strict computational and I/O constraints.
Malek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg, Shyamnath Gollakota
MobiCom3
2018 RAPID: In-Memory Analytical Query Processing Engine with Extreme Performance per Watt
abstract
Today, an ever increasing amount of transistors are packed into processor designs with extra features to support a broad range of applications. As a consequence, processors are becoming more and more complex and power hungry. At the same time, they only sustain an average performance for a wide variety of applications while not providing the best performance for specific applications. In this paper, we demonstrate through a carefully designed modern data processing system called RAPID and a simple, low-power processor specially tailored for data processing that at least an order of magnitude performance/power improvement in SQL processing can be achieved over a modern system running on today's complex processors. RAPID is designed from the ground up with hardware/software co-design in mind to provide architecture-conscious extreme performance while consuming less power in comparison to the modern database systems. The paper presents in detail the design and implementation of RAPID, a relational, columnar, in-memory query processing engine supporting analytical query workloads.
Cagri Balkesen, Nitin Kunal, Georgios Giannikis, Pit Fender, Seema Sundara, Felix Schmidt, Jarod Wen, Sandeep R. Agrawal, Arun Raghavan, Venkatanathan Varadarajan, Anand Viswanathan, Balakrishnan Chandrasekaran 0003, Sam Idicula, Nipun Agarwal, Eric Sedlar
SIGMOD Conference9
2017 A many-core architecture for in-memory data processing
abstract
For many years, the highest energy cost in processing has been data movement rather than computation, and energy is the limiting factor in processor design [21]. As the data needed for a single application grows to exabytes [56], there is clearly an opportunity to design a bandwidth-optimized architecture for big data computation by specializing hardware for data movement. We present the Data Processing Unit or DPU, a shared memory many-core that is specifically designed for high bandwidth analytics workloads. The DPU contains a unique Data Movement System (DMS), which provides hardware acceleration for data movement and partitioning operations at the memory controller that is sufficient to keep up with DDR bandwidth. The DPU also provides acceleration for core to core communication via a unique hardware RPC mechanism called the Atomic Transaction Engine. Comparison of a DPU chip fabricated in 40nm with a Xeon processor on a variety of data processing applications shows a 3× - 15× performance per watt advantage.
Sandeep R. Agrawal, Sam Idicula, Arun Raghavan, Evangelos Vlachos, Venkatraman Govindaraju, Venkatanathan Varadarajan, Cagri Balkesen, Georgios Giannikis, Charlie Roth, Nipun Agarwal, Eric Sedlar
MICRO3
2015 Rethinking SIMD Vectorization for In-Memory Databases
abstract
Analytical databases are continuously adapting to the underlying hardware in order to saturate all sources of parallelism. At the same time, hardware evolves in multiple directions to explore different trade-offs. The MIC architecture, one such example, strays from the mainstream CPU design by packing a larger number of simpler cores per chip, relying on SIMD instructions to fill the performance gap. Databases have been attempting to utilize the SIMD capabilities of CPUs. However, mainstream CPUs have only recently adopted wider SIMD registers and more advanced instructions, since they do not rely primarily on SIMD for efficiency. In this paper, we present novel vectorized designs and implementations of database operators, based on advanced SIMD operations, such as gathers and scatters. We study selections, hash tables, and partitioning; and combine them to build sorting and joins. Our evaluation on the MIC-based Xeon Phi co-processor as well as the latest mainstream CPUs shows that our vectorization designs are up to an order of magnitude faster than the state-of-the-art scalar and vector approaches. Also, we highlight the impact of efficient vectorization on the algorithmic design of in-memory database operators, as well as the architectural design and power efficiency of hardware, by making simple cores comparably fast to complex cores. This work is applicable to CPUs and co-processors with advanced SIMD capabilities, using either many simple cores or fewer complex cores.
Orestis Polychroniou, Arun Raghavan, Kenneth A. Ross
SIGMOD Conference2
2013 Computational sprinting on a hardware/software testbed
abstract
CMOS scaling trends have led to an inflection point where thermal constraints (especially in mobile devices that employ only passive cooling) preclude sustained operation of all transistors on a chip --- a phenomenon called "dark silicon." Recent research proposed computational sprinting --- exceeding sustainable thermal limits for short intervals --- to improve responsiveness in light of the bursty computation demands of many media-rich interactive mobile applications. Computational sprinting improves responsiveness by activating reserve cores (parallel sprinting) and/or boosting frequency/voltage (frequency sprinting) to power levels that far exceed the system's sustainable cooling capabilities, relying on thermal capacitance to buffer heat.
Arun Raghavan, Laurel Emurian, Marios C. Papaefthymiou, Kevin P. Pipe, Thomas F. Wenisch, Milo M. K. Martin
ASPLOS1
2013 TRANSIT: specifying protocols with concolic snippets
abstract
With the maturing of technology for model checking and constraint solving, there is an emerging opportunity to develop programming tools that can transform the way systems are specified. In this paper, we propose a new way to program distributed protocols using concolic snippets. Concolic snippets are sample execution fragments that contain both concrete and symbolic values. The proposed approach allows the programmer to describe the desired system partially using the traditional model of communicating extended finite-state-machines (EFSM), along with high-level invariants and concrete execution fragments. Our synthesis engine completes an EFSM skeleton by inferring guards and updates from the given fragments which is then automatically analyzed using a model checker with respect to the desired invariants. The counterexamples produced by the model checker can then be used by the programmer to add new concrete execution fragments that describe the correct behavior in the specific scenario corresponding to the counterexample.
Abhishek Udupa, Arun Raghavan, Jyotirmoy V. Deshmukh, Sela Mador-Haim, Milo M. K. Martin, Rajeev Alur
PLDI2
2012 Computational sprinting
abstract
Although transistor density continues to increase, voltage scaling has stalled and thus power density is increasing each technology generation. Particularly in mobile devices, which have limited cooling options, these trends lead to a utilization wall in which sustained chip performance is limited primarily by power rather than area. However, many mobile applications do not demand sustained performance; rather they comprise short bursts of computation in response to sporadic user activity. To improve responsiveness for such applications, this paper explores activating otherwise powered-down cores for sub-second bursts of intense parallel computation. The approach exploits the concept of computational sprinting, in which a chip temporarily exceeds its sustainable thermal power budget to provide instantaneous throughput, after which the chip must return to nominal operation to cool down. To demonstrate the feasibility of this approach, we analyze the thermal and electrical characteristics of a smart-phone-like system that nominally operates a single core (~1W peak), but can sprint with up to 16 cores for hundreds of milliseconds. We describe a thermal design that incorporates phase-change materials to provide thermal capacitance to enable such sprints. We analyze image recognition kernels to show that parallel sprinting has the potential to achieve the task response time of a 16W chip within the thermal constraints of a 1W mobile platform.
Arun Raghavan, Anuj Chandawalla, Marios C. Papaefthymiou, Kevin P. Pipe, Thomas F. Wenisch, Milo M. K. Martin
HPCA1
2010 RETCON: transactional repair without replay
abstract
Over the past decade there has been a surge of academic and industrial interest in optimistic concurrency, i.e. the speculative parallel execution of code regions that have the semantics of isolation. This work analyzes scalability bottlenecks of workloads that use optimistic concurrency. We find that one common bottleneck is updates to auxiliary program data in otherwise non-conflicting operations, e.g. reference count updates and hashtable occupancy field increments.
Colin Blundell, Arun Raghavan, Milo M. K. Martin
ISCA2
2010 Token tenure and PATCH: A predictive/adaptive token-counting hybrid
abstract
Traditional coherence protocols present a set of difficult trade-offs: the reliance of snoopy protocols on broadcast and ordered interconnects limits their scalability, while directory protocols incur a performance penalty on sharing misses due to indirection. This work introduces Patch (Predictive/Adaptive Token-Counting Hybrid), a coherence protocol that provides the scalability of directory protocols while opportunistically sending direct requests to reduce sharing latency. Patch extends a standard directory protocol to track tokens and use token-counting rules for enforcing coherence permissions. Token counting allows Patch to support direct requests on an unordered interconnect, while a mechanism called token tenure provides broadcast-free forward progress using the directory protocol's per-block point of ordering at the home along with either timeouts at requesters or explicit race notification messages. Patch makes three main contributions. First, Patch introduces token tenure, which provides broadcast-free forward progress for token-counting protocols. Second, Patch deprioritizes best-effort direct requests to match or exceed the performance of directory protocols without restricting scalability. Finally, Patch provides greater scalability than directory protocols when using inexact encodings of sharers because only processors holding tokens need to acknowledge requests. Overall, Patch is a “one-size-fits-all” coherence protocol that dynamically adapts to work well for small systems, large systems, and anywhere in between.
Arun Raghavan, Colin Blundell, Milo M. K. Martin
ACM Trans. Archit. Code Optim.1
2008 Token tenure: PATCHing token counting using directory-based cache coherence
abstract
Traditional coherence protocols present a set of difficult tradeoffs: the reliance of snoopy protocols on broadcast and ordered interconnects limits their scalability, while directory protocols incur a performance penalty on sharing misses due to indirection. This work introduces PATCH (Predictive/Adaptive Token Counting Hybrid), a coherence protocol that provides the scalability of directory protocols while opportunistically sending direct requests to reduce sharing latency. PATCH extends a standard directory protocol to track tokens and use token counting rules for enforcing coherence permissions. Token counting allows PATCH to support direct requests on an unordered interconnect, while a mechanism called token tenure uses local processor timeouts and the directorypsilas per-block point of ordering at the home node to guarantee forward progress without relying on broadcast. PATCH makes three main contributions. First, PATCH introduces token tenure, which provides broadcast-free forward progress for token counting protocols. Second, PATCH deprioritizes best-effort direct requests to match or exceed the performance of directory protocols without restricting scalability. Finally, PATCH provides greater scalability than directory protocols when using inexact encodings of sharers because only processors holding tokens need to acknowledge requests. Overall, PATCH is a ldquoone-size-fits-allrdquo coherence protocol that dynamically adapts to work well for small systems, large systems, and anywhere in between.
Arun Raghavan, Colin Blundell, Milo M. K. Martin
MICRO1