EDBT 2026 Demo / reviewers in the wild / expert
Arun Raghavan
dblp:61/4142
· DBLP profile ↗
10ranked-venue papers
4as first author
1since 2021 · last 2025
0000-0001-5339-3352ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Hardware accelerators and domain-specific architectures · 26% Memory systems · 22% Energy-efficient computing · 15% | |
| Databases, data mining, and information retrieval
2 papers |
Query processing and optimization · 62% Indexing and storage engines · 38% | |
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 84% Concurrent programming · 16% |
Topics — the 22 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | Wireless Hearables With Programmable Speech AI Accelerators · MobiCom 2025 |
Indexing and storage engines
columnar storage |
0.3 | 1 | 2018 | RAPID: In-Memory Analytical Query Processing Engine with Extreme Performance per Watt · SIGMOD Conference 2018 |
Energy-efficient computing › power management › peak power management
computational sprinting |
0.3 | 2 | 2013 | Computational sprinting on a hardware/software testbed · ASPLOS 2013 Computational sprinting · HPCA 2012 |
Energy-efficient computing
thermal management |
0.3 | 2 | 2013 | Computational sprinting on a hardware/software testbed · ASPLOS 2013 Computational sprinting · HPCA 2012 |
Memory systems
data movement |
0.3 | 1 | 2017 | A many-core architecture for in-memory data processing · MICRO 2017 |
Cloud and datacenter computing
data movement acceleration |
0.3 | 1 | 2017 | A many-core architecture for in-memory data processing · MICRO 2017 |
Hardware accelerators and domain-specific architectures › domain-specific accelerator
data processing accelerator |
0.3 | 1 | 2017 | A many-core architecture for in-memory data processing · MICRO 2017 |
Memory systems › in-memory computing
in-memory data processing |
0.3 | 1 | 2017 | A many-core architecture for in-memory data processing · MICRO 2017 |
Query processing and optimization
SIMD vectorization |
0.2 | 1 | 2015 | Rethinking SIMD Vectorization for In-Memory Databases · SIGMOD Conference 2015 |
Memory systems
cache coherence |
0.2 | 2 | 2010 | Token tenure and PATCH: A predictive/adaptive token-counting hybrid · ACM Trans. Archit. Code Optim. 2010 Token tenure: PATCHing token counting using directory-based cache coherence · MICRO 2008 |
Memory systems › cache coherence
directory-based coherence |
0.2 | 2 | 2010 | Token tenure and PATCH: A predictive/adaptive token-counting hybrid · ACM Trans. Archit. Code Optim. 2010 Token tenure: PATCHing token counting using directory-based cache coherence · MICRO 2008 |
Program synthesis and code generation › concurrent program synthesis
protocol synthesis |
0.2 | 1 | 2013 | TRANSIT: specifying protocols with concolic snippets · PLDI 2013 |
Embedded and real-time systems
mobile computing |
0.2 | 1 | 2013 | Computational sprinting on a hardware/software testbed · ASPLOS 2013 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2012 | Computational sprinting · HPCA 2012 |
Distributed systems › concurrency control
optimistic concurrency control |
0.1 | 1 | 2010 | RETCON: transactional repair without replay · ISCA 2010 |
Memory systems › cache coherence
scalable coherence protocol |
0.1 | 1 | 2010 | Token tenure and PATCH: A predictive/adaptive token-counting hybrid · ACM Trans. Archit. Code Optim. 2010 |
Parallel and multicore computing
transactional memory |
0.1 | 1 | 2010 | RETCON: transactional repair without replay · ISCA 2010 |
Parallel and multicore computing › transactional memory
transactional memory scalability |
0.1 | 1 | 2010 | RETCON: transactional repair without replay · ISCA 2010 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2017 | A many-core architecture for in-memory data processing · MICRO 2017 |
Embedded and real-time systems › mobile computing
mobile computing platforms |
0.0 | 1 | 2012 | Computational sprinting · HPCA 2012 |
Concurrent programming
synchronization |
0.0 | 1 | 2010 | RETCON: transactional repair without replay · ISCA 2010 |
Performance modeling and evaluation › parallel performance evaluation
multicore scalability |
0.0 | 1 | 2008 | Token tenure: PATCHing token counting using directory-based cache coherence · MICRO 2008 |
Methods — techniques the papers use, named apart from their topics
deep learning · 0.9hardware-software co-design · 0.7model checking · 0.5extended finite state machine · 0.5constraint solving · 0.5concolic execution · 0.5hash table · 0.4SIMD gather/scatter · 0.4hardware RPC · 0.3phase-change material thermal capacitance · 0.1workload analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Wireless Hearables With Programmable Speech AI AcceleratorsabstractThe conventional wisdom has been that designing ultra-compact, battery-constrained wireless hearables with on-device speech AI models is challenging due to the high computational demands of streaming deep learning models. Speech AI models require continuous, real-time audio processing, imposing strict computational and I/O constraints. Malek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg, Shyamnath Gollakota |
MobiCom | 3 |
| 2018 | RAPID: In-Memory Analytical Query Processing Engine with Extreme Performance per WattabstractToday, an ever increasing amount of transistors are packed into processor designs with extra features to support a broad range of applications. As a consequence, processors are becoming more and more complex and power hungry. At the same time, they only sustain an average performance for a wide variety of applications while not providing the best performance for specific applications. In this paper, we demonstrate through a carefully designed modern data processing system called RAPID and a simple, low-power processor specially tailored for data processing that at least an order of magnitude performance/power improvement in SQL processing can be achieved over a modern system running on today's complex processors. RAPID is designed from the ground up with hardware/software co-design in mind to provide architecture-conscious extreme performance while consuming less power in comparison to the modern database systems. The paper presents in detail the design and implementation of RAPID, a relational, columnar, in-memory query processing engine supporting analytical query workloads. Cagri Balkesen, Nitin Kunal, Georgios Giannikis, Pit Fender, Seema Sundara, Felix Schmidt, Jarod Wen, Sandeep R. Agrawal, Arun Raghavan, Venkatanathan Varadarajan, Anand Viswanathan, Balakrishnan Chandrasekaran 0003, Sam Idicula, Nipun Agarwal, Eric Sedlar |
SIGMOD Conference | 9 |
| 2017 | A many-core architecture for in-memory data processingabstractFor many years, the highest energy cost in processing has been data movement rather than computation, and energy is the limiting factor in processor design [21]. As the data needed for a single application grows to exabytes [56], there is clearly an opportunity to design a bandwidth-optimized architecture for big data computation by specializing hardware for data movement. We present the Data Processing Unit or DPU, a shared memory many-core that is specifically designed for high bandwidth analytics workloads. The DPU contains a unique Data Movement System (DMS), which provides hardware acceleration for data movement and partitioning operations at the memory controller that is sufficient to keep up with DDR bandwidth. The DPU also provides acceleration for core to core communication via a unique hardware RPC mechanism called the Atomic Transaction Engine. Comparison of a DPU chip fabricated in 40nm with a Xeon processor on a variety of data processing applications shows a 3× - 15× performance per watt advantage. Sandeep R. Agrawal, Sam Idicula, Arun Raghavan, Evangelos Vlachos, Venkatraman Govindaraju, Venkatanathan Varadarajan, Cagri Balkesen, Georgios Giannikis, Charlie Roth, Nipun Agarwal, Eric Sedlar |
MICRO | 3 |
| 2015 | Rethinking SIMD Vectorization for In-Memory DatabasesabstractAnalytical databases are continuously adapting to the underlying hardware in order to saturate all sources of parallelism. At the same time, hardware evolves in multiple directions to explore different trade-offs. The MIC architecture, one such example, strays from the mainstream CPU design by packing a larger number of simpler cores per chip, relying on SIMD instructions to fill the performance gap. Databases have been attempting to utilize the SIMD capabilities of CPUs. However, mainstream CPUs have only recently adopted wider SIMD registers and more advanced instructions, since they do not rely primarily on SIMD for efficiency. In this paper, we present novel vectorized designs and implementations of database operators, based on advanced SIMD operations, such as gathers and scatters. We study selections, hash tables, and partitioning; and combine them to build sorting and joins. Our evaluation on the MIC-based Xeon Phi co-processor as well as the latest mainstream CPUs shows that our vectorization designs are up to an order of magnitude faster than the state-of-the-art scalar and vector approaches. Also, we highlight the impact of efficient vectorization on the algorithmic design of in-memory database operators, as well as the architectural design and power efficiency of hardware, by making simple cores comparably fast to complex cores. This work is applicable to CPUs and co-processors with advanced SIMD capabilities, using either many simple cores or fewer complex cores. Orestis Polychroniou, Arun Raghavan, Kenneth A. Ross |
SIGMOD Conference | 2 |
| 2013 | Computational sprinting on a hardware/software testbedabstractCMOS scaling trends have led to an inflection point where thermal constraints (especially in mobile devices that employ only passive cooling) preclude sustained operation of all transistors on a chip --- a phenomenon called "dark silicon." Recent research proposed computational sprinting --- exceeding sustainable thermal limits for short intervals --- to improve responsiveness in light of the bursty computation demands of many media-rich interactive mobile applications. Computational sprinting improves responsiveness by activating reserve cores (parallel sprinting) and/or boosting frequency/voltage (frequency sprinting) to power levels that far exceed the system's sustainable cooling capabilities, relying on thermal capacitance to buffer heat. Arun Raghavan, Laurel Emurian, Marios C. Papaefthymiou, Kevin P. Pipe, Thomas F. Wenisch, Milo M. K. Martin |
ASPLOS | 1 |
| 2013 | TRANSIT: specifying protocols with concolic snippetsabstractWith the maturing of technology for model checking and constraint solving, there is an emerging opportunity to develop programming tools that can transform the way systems are specified. In this paper, we propose a new way to program distributed protocols using concolic snippets. Concolic snippets are sample execution fragments that contain both concrete and symbolic values. The proposed approach allows the programmer to describe the desired system partially using the traditional model of communicating extended finite-state-machines (EFSM), along with high-level invariants and concrete execution fragments. Our synthesis engine completes an EFSM skeleton by inferring guards and updates from the given fragments which is then automatically analyzed using a model checker with respect to the desired invariants. The counterexamples produced by the model checker can then be used by the programmer to add new concrete execution fragments that describe the correct behavior in the specific scenario corresponding to the counterexample. Abhishek Udupa, Arun Raghavan, Jyotirmoy V. Deshmukh, Sela Mador-Haim, Milo M. K. Martin, Rajeev Alur |
PLDI | 2 |
| 2012 | Computational sprintingabstractAlthough transistor density continues to increase, voltage scaling has stalled and thus power density is increasing each technology generation. Particularly in mobile devices, which have limited cooling options, these trends lead to a utilization wall in which sustained chip performance is limited primarily by power rather than area. However, many mobile applications do not demand sustained performance; rather they comprise short bursts of computation in response to sporadic user activity. To improve responsiveness for such applications, this paper explores activating otherwise powered-down cores for sub-second bursts of intense parallel computation. The approach exploits the concept of computational sprinting, in which a chip temporarily exceeds its sustainable thermal power budget to provide instantaneous throughput, after which the chip must return to nominal operation to cool down. To demonstrate the feasibility of this approach, we analyze the thermal and electrical characteristics of a smart-phone-like system that nominally operates a single core (~1W peak), but can sprint with up to 16 cores for hundreds of milliseconds. We describe a thermal design that incorporates phase-change materials to provide thermal capacitance to enable such sprints. We analyze image recognition kernels to show that parallel sprinting has the potential to achieve the task response time of a 16W chip within the thermal constraints of a 1W mobile platform. Arun Raghavan, Anuj Chandawalla, Marios C. Papaefthymiou, Kevin P. Pipe, Thomas F. Wenisch, Milo M. K. Martin |
HPCA | 1 |
| 2010 | RETCON: transactional repair without replayabstractOver the past decade there has been a surge of academic and industrial interest in optimistic concurrency, i.e. the speculative parallel execution of code regions that have the semantics of isolation. This work analyzes scalability bottlenecks of workloads that use optimistic concurrency. We find that one common bottleneck is updates to auxiliary program data in otherwise non-conflicting operations, e.g. reference count updates and hashtable occupancy field increments. Colin Blundell, Arun Raghavan, Milo M. K. Martin |
ISCA | 2 |
| 2010 | Token tenure and PATCH: A predictive/adaptive token-counting hybridabstractTraditional coherence protocols present a set of difficult trade-offs: the reliance of snoopy protocols on broadcast and ordered interconnects limits their scalability, while directory protocols incur a performance penalty on sharing misses due to indirection. This work introduces Patch (Predictive/Adaptive Token-Counting Hybrid), a coherence protocol that provides the scalability of directory protocols while opportunistically sending direct requests to reduce sharing latency. Patch extends a standard directory protocol to track tokens and use token-counting rules for enforcing coherence permissions. Token counting allows Patch to support direct requests on an unordered interconnect, while a mechanism called token tenure provides broadcast-free forward progress using the directory protocol's per-block point of ordering at the home along with either timeouts at requesters or explicit race notification messages. Patch makes three main contributions. First, Patch introduces token tenure, which provides broadcast-free forward progress for token-counting protocols. Second, Patch deprioritizes best-effort direct requests to match or exceed the performance of directory protocols without restricting scalability. Finally, Patch provides greater scalability than directory protocols when using inexact encodings of sharers because only processors holding tokens need to acknowledge requests. Overall, Patch is a “one-size-fits-all” coherence protocol that dynamically adapts to work well for small systems, large systems, and anywhere in between. Arun Raghavan, Colin Blundell, Milo M. K. Martin |
ACM Trans. Archit. Code Optim. | 1 |
| 2008 | Token tenure: PATCHing token counting using directory-based cache coherenceabstractTraditional coherence protocols present a set of difficult tradeoffs: the reliance of snoopy protocols on broadcast and ordered interconnects limits their scalability, while directory protocols incur a performance penalty on sharing misses due to indirection. This work introduces PATCH (Predictive/Adaptive Token Counting Hybrid), a coherence protocol that provides the scalability of directory protocols while opportunistically sending direct requests to reduce sharing latency. PATCH extends a standard directory protocol to track tokens and use token counting rules for enforcing coherence permissions. Token counting allows PATCH to support direct requests on an unordered interconnect, while a mechanism called token tenure uses local processor timeouts and the directorypsilas per-block point of ordering at the home node to guarantee forward progress without relying on broadcast. PATCH makes three main contributions. First, PATCH introduces token tenure, which provides broadcast-free forward progress for token counting protocols. Second, PATCH deprioritizes best-effort direct requests to match or exceed the performance of directory protocols without restricting scalability. Finally, PATCH provides greater scalability than directory protocols when using inexact encodings of sharers because only processors holding tokens need to acknowledge requests. Overall, PATCH is a ldquoone-size-fits-allrdquo coherence protocol that dynamically adapts to work well for small systems, large systems, and anywhere in between. Arun Raghavan, Colin Blundell, Milo M. K. Martin |
MICRO | 1 |