Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sai Charan Koduru

dblp:09/1732 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2018
0000-0003-0857-6357ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 36% Distributed systems · 36% Parallel and multicore computing · 29%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel algorithms › parallel matrix algorithms
asynchronous iterative methods
0.212014
ASPIRE: exploiting asynchronous parallelism in iterative algorithms using a relaxed consistency based DSM · OOPSLA 2014
Distributed systems › consistency models
bounded staleness
0.212014
ASPIRE: exploiting asynchronous parallelism in iterative algorithms using a relaxed consistency based DSM · OOPSLA 2014
Distributed systems
consistency protocols
0.212014
ASPIRE: exploiting asynchronous parallelism in iterative algorithms using a relaxed consistency based DSM · OOPSLA 2014
Memory systems › shared memory
distributed shared memory
0.212014
ASPIRE: exploiting asynchronous parallelism in iterative algorithms using a relaxed consistency based DSM · OOPSLA 2014
Memory systems › memory consistency › memory consistency model
weak memory model
0.212014
ASPIRE: exploiting asynchronous parallelism in iterative algorithms using a relaxed consistency based DSM · OOPSLA 2014
Parallel and multicore computing › parallel algorithms
graph algorithms
0.112014
ASPIRE: exploiting asynchronous parallelism in iterative algorithms using a relaxed consistency based DSM · OOPSLA 2014
Parallel and multicore computing › graph processing
vertex-centric graph processing
0.112014
ASPIRE: exploiting asynchronous parallelism in iterative algorithms using a relaxed consistency based DSM · OOPSLA 2014

Methods — techniques the papers use, named apart from their topics

bounded staleness · 0.2best-effort refresh policy · 0.2
YearPublicationVenuePosition
2018 OMR: out-of-core MapReduce for large data sets
abstract
While single machine MapReduce systems can squeeze out maximum performance from available multi-cores, they are often limited by the size of main memory and can thus only process small datasets. Our experience shows that the state-of-the-art single-machine in-memory MapReduce system Metis frequently experiences out-of-memory crashes. Even though today's computers are equipped with efficient secondary storage devices, the frameworks do not utilize these devices mainly because disk access latencies are much higher than those for main memory. Therefore, the single-machine setup of the Hadoop system performs much slower when it is presented with the datasets which are larger than the main memory. Moreover, such frameworks also require tuning a lot of parameters which puts an added burden on the programmer. In this paper we present OMR, an Out-of-core MapReduce system that not only successfully handles datasets that are far larger than the size of main memory, it also guarantees linear scaling with the growing data sizes. OMR actively minimizes the amount of data to be read/written to/from disk via on-the-fly aggregation and it uses block sequential disk read/write operations whenever disk accesses become necessary to avoid running out of memory. We theoretically prove OMR's linear scalability and empirically demonstrate it by processing datasets that are up to 5x larger than main memory. Our experiments show that in comparison to the standalone single-machine setup of the Hadoop system, OMR delivers far higher performance. Also in contrast to Metis, OMR avoids out-of-memory crashes for large datasets as well as delivers higher performance when datasets are small enough to fit in main memory.
Gurneet Kaur, Keval Vora, Sai Charan Koduru, Rajiv Gupta 0001
ISMM3
2015 Optimizing Caching DSM for Distributed Software Speculation
abstract
Clusters with caching DSMs deliver programmability and performance by supporting shared-memory programming and tolerate remote I/O latencies via caching. The input to a data parallel program is partitioned across the cluster while the DSM transparently fetches and caches remote data as needed. Irregular applications, however, are challenging to parallelize because the input related data dependences that manifest at runtime require use of speculation for correct parallel execution. By speculating that there are no input related cross iteration dependences, private copies of the input can be processed by parallelizing the loop, the absence of dependences is validated before committing the computed results. We show that while caching helps tolerate long communication latencies in irregular data-parallel applications, using a cached values in a computation can lead to misspeculation and thus aggressive caching can degrade performance due to increased misspeculation rate. We present optimizations for distributed speculation on caching based DSMs that decrease the cost of misspeculation check and speed up the re-execution of misspeculated recomputations. Optimized distributed speculation achieves speedups of 2.24x for coloring, 1.71x for connected components, 1.88x for community detection, 1.32x for shortest path, and 1.74x for pagerank over unoptimized speculation.
Sai Charan Koduru, Keval Vora, Rajiv Gupta 0001
CLUSTER1
2014 ASPIRE: exploiting asynchronous parallelism in iterative algorithms using a relaxed consistency based DSM
abstract
Many vertex-centric graph algorithms can be expressed using asynchronous parallelism by relaxing certain read-after-write data dependences and allowing threads to compute vertex values using stale (i.e., not the most recent) values of their neighboring vertices. We observe that on distributed shared memory systems, by converting synchronous algorithms into their asynchronous counterparts, algorithms can be made tolerant to high inter-node communication latency. However, high inter-node communication latency can lead to excessive use of stale values causing an increase in the number of iterations required by the algorithms to converge. Although by using bounded staleness we can restrict the slowdown in the rate of convergence, this also restricts the ability to tolerate communication latency. In this paper we design a relaxed memory consistency model and consistency protocol that simultaneously tolerate communication latency and minimize the use of stale values. This is achieved via a coordinated use of best effort refresh policy and bounded staleness. We demonstrate that for a range of asynchronous graph algorithms and PDE solvers, on an average, our approach outperforms algorithms based upon: prior relaxed memory models that allow stale values by at least 2.27x; and Bulk Synchronous Parallel (BSP) model by 4.2x. We also show that our approach frequently outperforms GraphLab, a popular distributed graph processing framework.
Keval Vora, Sai Charan Koduru, Rajiv Gupta 0001
OOPSLA2