EDBT 2026 Demo / reviewers in the wild / expert
Sai Prashanth Muralidhara
dblp:20/6746
· DBLP profile ↗
13ranked-venue papers
5as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-authorSoftware engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Memory systems · 66% Parallel and multicore computing · 14% Storage systems · 7% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 60% Operating systems · 21% Software maintenance and evolution · 18% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › loop transformation
loop distribution |
0.2 | 2 | 2010 | Cache topology aware computation mapping for multicores · PLDI 2010 Computation mapping for multi-level storage cache hierarchies · HPDC 2010 |
Memory systems
cache |
0.2 | 2 | 2010 | Cache topology aware computation mapping for multicores · PLDI 2010 Optimizing shared cache behavior of chip multiprocessors · MICRO 2009 |
Memory systems
DRAM |
0.1 | 1 | 2011 | Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 |
Memory systems
memory controller |
0.1 | 1 | 2011 | Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 |
Memory systems
memory interference |
0.1 | 1 | 2011 | Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 |
Memory systems › memory controller
memory scheduling |
0.1 | 1 | 2011 | Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 |
Operating systems › i/o
i/o optimization |
0.1 | 1 | 2010 | Computation mapping for multi-level storage cache hierarchies · HPDC 2010 |
Memory systems
cache management |
0.1 | 1 | 2010 | Intra-application shared cache partitioning for multithreaded applications · PPoPP 2010 |
Memory systems › cache management › cache partitioning
shared cache partitioning |
0.1 | 1 | 2010 | Intra-application shared cache partitioning for multithreaded applications · PPoPP 2010 |
Compilers and program optimization › memory optimization
data locality optimization |
0.1 | 1 | 2009 | Optimizing shared cache behavior of chip multiprocessors · MICRO 2009 |
Software maintenance and evolution › software reengineering
software restructuring |
0.1 | 1 | 2009 | Optimizing shared cache behavior of chip multiprocessors · MICRO 2009 |
Parallel and multicore computing › task allocation
dynamic mapping |
0.1 | 1 | 2009 | Dynamic thread and data mapping for NoC based CMPs · DAC 2009 |
Memory systems › cache management
shared cache management |
0.1 | 1 | 2009 | Optimizing shared cache behavior of chip multiprocessors · MICRO 2009 |
Parallel and multicore computing › task allocation
thread and data mapping |
0.1 | 1 | 2009 | Dynamic thread and data mapping for NoC based CMPs · DAC 2009 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2011 | Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 Optimizing shared cache behavior of chip multiprocessors · MICRO 2009 |
Cloud and datacenter computing › resource management
shared resource management |
0.0 | 1 | 2011 | Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 |
High-performance computing
data-intensive computing |
0.0 | 1 | 2010 | Computation mapping for multi-level storage cache hierarchies · HPDC 2010 |
Processor architecture and microarchitecture
multicore design |
0.0 | 1 | 2010 | Cache topology aware computation mapping for multicores · PLDI 2010 |
Parallel and multicore computing › thread-level parallelism
multithreaded applications |
0.0 | 1 | 2010 | Intra-application shared cache partitioning for multithreaded applications · PPoPP 2010 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2loop scheduling · 0.2compiler-directed mapping · 0.2scheduling · 0.2loop iteration allocation · 0.2application-aware partitioning · 0.1cache partitioning · 0.1runtime mapping · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Bandwidth Constrained Coordinated HW/SW Prefetching for Multicores
Sai Prashanth Muralidhara, Mahmut T. Kandemir |
Euro-Par (1) | 1 |
| 2011 | Reducing memory interference in multicore systems via application-aware memory channel partitioningabstractMain memory is a major shared resource among cores in a multicore system. If the interference between different applications' memory requests is not controlled effectively, system performance can degrade significantly. Previous work aimed to mitigate the problem of interference between applications by changing the scheduling policy in the memory controller, i.e., by prioritizing memory requests from applications in a way that benefits system performance. Sai Prashanth Muralidhara, Lavanya Subramanian, Onur Mutlu, Mahmut T. Kandemir, Thomas Moscibroda |
MICRO | 1 |
| 2010 | A special-purpose compiler for look-up table and code generation for function evaluationabstractElementary functions are extensively used in computer graphics, signal and image processing, and communication systems. This paper presents a special-purpose compiler that automatically generates customized look-up tables and implementations for elementary functions under user given constraints. The generated implementations include a C/C++ code that can be used directly by applications running on multicores, as well as a MATLAB-like code that can be translated directly to a hardware module on FPGA platforms. The experimental results show that our solutions for function evaluation bring significant performance improvements to applications on multicores as well as significant resource savings to designs on FPGAs. Lanping Deng, Praveen Yedlapalli, Sai Prashanth Muralidhara, Hui Zhao 0013, Mahmut T. Kandemir, Chaitali Chakrabarti, Nikos Pitsianis, Xiaobai Sun |
DATE | 4 |
| 2010 | Code Scheduling for Optimizing Parallelism and Data Locality
Taylan Yemliha, Mahmut T. Kandemir, Ozcan Ozturk 0001, Emre Kultursay, Sai Prashanth Muralidhara |
Euro-Par (1) | 5 |
| 2010 | Computation mapping for multi-level storage cache hierarchiesabstractImproving I/O performance is an important issue for many data-intensive, large-scale parallel applications. Although storage caches are used for improving I/O latencies of parallel applications, most of the prior work has focused on the management and partitioning of cache space. In particular, the compiler's role in taking advantage of multilevel storage caches has been largely unexplored. The main contribution of this paper is a shared-storage, cache-aware loop iteration distribution (iteration-to-processor mapping) scheme for I/O-intensive applications that manipulate disk-resident data sets. The proposed scheme is compiler directed and can be tuned to target any multilevel storage cache hierarchy. At the core of our scheme lies an iterative strategy that clusters loop iterations based on the underlying storage cache hierarchy and on the way these different storage caches in the hierarchy are shared by different processors. We tested this mapping scheme using a set of eight I/O-intensive application programs. The results collected so far are promising. Our proposed scheme improves the I/O performance of the tested applications by 26.3% on average, and this improvement leads to an average 18.9% reduction in the overall execution latencies of these applications. Moreover, our scheme performs significantly better than a state-of-the-art (but storage-cache- hierarchy agnostic) data locality optimization scheme. We also present an enhancement to our baseline implementation that performs local scheduling once the loop iteration distribution is performed. We observe that applying this enhancement improves I/O latency and total execution time further by 30.7% and 21.9%, respectively. Mahmut T. Kandemir, Sai Prashanth Muralidhara, Mustafa Karaköy, Seung Woo Son 0001 |
HPDC | 2 |
| 2010 | Intra-application cache partitioningabstractEfficient management of shared on-chip resources such as the shared level 2 (L2) cache has become an important problem with the emergence of chip multiprocessors (CMPs). Partitioning the shared cache in chip multiprocessors (CMPs) among concurrently executing applications can provide important benefits such as throughput improvement, fairness guarantees, and quality of service (QoS) enhancements. In this paper, we pose an interesting related question, which is, if partitioning the shared cache space among concurrently executing threads of the same application can enhance the application performance. We address this problem by identifying and speeding up the slowest thread, also termed as the critical path thread, during each execution interval since the overall performance of a multithreaded application is determined by the critical path thread. To do so, we propose a dynamic, runtime system based, cache partitioning scheme that partitions the shared cache space dynamically among the individual threads of a given application. In a nutshell, we wish to take some cache space away from the faster threads and give it to the critical path thread at each execution interval. We show that speeding up the critical path thread this way, results in overall performance enhancement of the application execution in the long term. Our experimental evaluation indicates that, the proposed dynamic cache partitioning scheme yields benefits up to 15% over a shared cache with no partitions, up to 23% over a statically partitioned cache (private cache) and up to 20% over a throughput-oriented scheme. Sai Prashanth Muralidhara, Mahmut T. Kandemir, Padma Raghavan |
IPDPS | 1 |
| 2010 | Cache topology aware computation mapping for multicoresabstractThe main contribution of this paper is a compiler based, cache topology aware code optimization scheme for emerging multicore systems. This scheme distributes the iterations of a loop to be executed in parallel across the cores of a target multicore machine and schedules the iterations assigned to each core. Our goal is to improve the utilization of the on-chip multi-layer cache hierarchy and to maximize overall application performance. We evaluate our cache topology aware approach using a set of twelve applications and three different commercial multicore machines. In addition, to study some of our experimental parameters in detail and to explore future multicore machines (with higher core counts and deeper on-chip cache hierarchies), we also conduct a simulation based study. The results collected from our experiments with three Intel multicore machines show that the proposed compiler-based approach is very effective in enhancing performance. In addition, our simulation results indicate that optimizing for the on-chip cache hierarchy will be even more important in future multicores with increasing numbers of cores and cache levels. Mahmut T. Kandemir, Taylan Yemliha, Sai Prashanth Muralidhara, Shekhar Srikantaiah, Mary Jane Irwin |
PLDI | 3 |
| 2010 | Intra-application shared cache partitioning for multithreaded applicationsabstractIn this paper, we address the problem of partitioning a shared cache when the executing threads belong to the same application. Sai Prashanth Muralidhara, Mahmut T. Kandemir, Padma Raghavan |
PPoPP | 1 |
| 2009 | Slicing based code parallelization for minimizing inter-processor communicationabstractDate of Conference: 1 -16 October, 2009 Mahmut T. Kandemir, Sai Prashanth Muralidhara, Ozcan Ozturk 0001, Sri Hari Krishna Narayanan |
CASES | 3 |
| 2009 | Dynamic thread and data mapping for NoC based CMPsabstractThread mapping and data mapping are two important problems in the context of NoC (network-on-chip) based CMPs (chip multiprocessors). While a compiler can determine suitable mappings for data and threads, such static mappings may not work well for multithreaded applications that go through different execution phases during their execution, each phase with potentially different data access patterns than others. Instead, a dynamic mapping strategy, if its overheads can be kept low, may be a more promising option. In this work, we present dynamic (runtime) thread and data mappings for NoC based CMPs. The goal of these mappings is to reduce the distance between the location of the core that requests data and the core whose local memory contains that requested data. In our experiments, we evaluate our proposed thread mapping and data mapping in isolation as well as in an integrated manner. Mahmut T. Kandemir, Ozcan Ozturk 0001, Sai Prashanth Muralidhara |
DAC | 3 |
| 2009 | Communication Based Proactive Link Power Management
Sai Prashanth Muralidhara, Mahmut T. Kandemir |
HiPEAC | 1 |
| 2009 | Optimizing shared cache behavior of chip multiprocessorsabstractOne of the critical problems associated with emerging chip multiprocessors (CMPs) is the management of on-chip shared cache space. Unfortunately, single processor centric data locality optimization schemes may not work well in the CMP case as data accesses from multiple cores can create conflicts in the shared cache space. The main contribution of this paper is a compiler directed code restructuring scheme for enhancing locality of shared data in CMPs. The proposed scheme targets the last level shared cache that exist in many commercial CMPs and has two components, namely, allocation, which determines the set of loop iterations assigned to each core, and scheduling, which determines the order in which the iterations assigned to a core are executed. Our scheme restructures the application code such that the different cores operate on shared data blocks at the same time, to the extent allowed by data dependencies. This helps to reduce reuse distances for the shared data and improves on-chip cache performance. We evaluated our approach using the Splash-2 and Parsec applications through both simulations and experiments on two commercial multi-core machines. Our experimental evaluation indicates that the proposed data locality optimization scheme improves inter-core conflict misses in the shared cache by 67% on average when both allocation and scheduling are used. Also, the execution time improvements we achieve (29% on average) are very close to the optimal savings that could be achieved using a hypothetical scheme. Mahmut T. Kandemir, Sai Prashanth Muralidhara, Sri Hari Krishna Narayanan, Ozcan Ozturk 0001 |
MICRO | 2 |
| 2008 | Profiler and compiler assisted adaptive I/O prefetching for shared storage cachesabstractI/O prefetching has been employed in the past as one of the mechanisms to hide large disk latencies. However, I/O prefetching in parallel applications is problematic when multiple CPUs share the same set of disks due to the possibility that prefetches from different CPUs can interact on shared memory caches in the I/O nodes in complex and unpredictable ways. In this paper, we (i) quantify the impact of compiler-directed I/O prefetching - developed originally in the context of sequential execution - on shared caches at I/O nodes. The experimental data collected shows that while I/O prefetching brings benefits, its effectiveness reduces significantly as the number of CPUs is increased; (ii) identify inter-CPU misses due to harmful prefetches as one of the main sources for this reduction in performance with the increased number of CPUs; and (iii) propose and experimentally evaluate a profiler and compiler assisted adaptive I/O prefetching scheme targeting shared storage caches. The proposed scheme obtains inter-thread data sharing information using profiling and, based on the captured data sharing patterns, divides the threads into clusters and assigns a separate (customized) I/O prefetcher thread for each cluster. In our approach, the compiler generates the I/O prefetching threads automatically. We implemented this new I/O prefetching scheme using a compiler and the PVFS file system running on Linux, and the empirical data collected clearly underline the importance of adapting I/O prefetching based on program phases. Specifically, our proposed scheme improves performance, on average, by 19.9%, 11.9% and 10.3% over the cases without I/O prefetching, with independent I/O prefetching (each CPU is performing compiler-directed I/O prefetching independently), and with one CPU prefetching (one CPU is reserved for prefetching on behalf of others), respectively, when 8 CPUs are used. Seung Woo Son 0001, Sai Prashanth Muralidhara, Ozcan Ozturk 0001, Mahmut T. Kandemir, Ibrahim Kolcu, Mustafa Karaköy |
PACT | 2 |