Arun Raman

dblp:00/6642 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
2since 2021 · last 2024
0000-0002-5510-2405ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 4 first-authorSystems, architecture and hardware · 6 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 76% Distributed systems · 8% Memory systems · 8%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 66% Concurrent programming · 34%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel programming models
0.322012
Parcae: a system for flexible parallel execution · PLDI 2012
Parallelism orchestration using DoPE: the degree of parallelism executive · PLDI 2011
Parallel and multicore computing
parallel programming runtimes
0.322012
Parcae: a system for flexible parallel execution · PLDI 2012
Parallelism orchestration using DoPE: the degree of parallelism executive · PLDI 2011
Parallel and multicore computing
speculative parallelization
0.222010
Scalable Speculative Parallelization on Commodity Clusters · MICRO 2010
Speculative parallelization using software multi-threaded transactions · ASPLOS 2010
Parallel and multicore computing › speculative parallelization
dynamic parallelization
0.112012
Parcae: a system for flexible parallel execution · PLDI 2012
Compilers and program optimization
dynamic optimization
0.112011
Sprint: speculative prefetching of remote data · OOPSLA 2011
Parallel and multicore computing › parallelization strategies
degree of parallelism management
0.112011
Parallelism orchestration using DoPE: the degree of parallelism executive · PLDI 2011
Memory systems › cache management
prefetching and caching
0.112011
Sprint: speculative prefetching of remote data · OOPSLA 2011
Distributed systems › distributed communication
remote data access
0.112011
Sprint: speculative prefetching of remote data · OOPSLA 2011
High-performance computing
cluster computing
0.112010
Scalable Speculative Parallelization on Commodity Clusters · MICRO 2010
Parallel and multicore computing › speculative parallelization
thread-level speculation
0.112010
Scalable Speculative Parallelization on Commodity Clusters · MICRO 2010
Parallel and multicore computing
transactional memory
0.112010
Scalable Speculative Parallelization on Commodity Clusters · MICRO 2010
Performance modeling and evaluation
irregular access patterns
0.012011
Sprint: speculative prefetching of remote data · OOPSLA 2011
Concurrent programming
synchronization
0.012010
Speculative parallelization using software multi-threaded transactions · ASPLOS 2010
Concurrent programming
transactional memory
0.012010
Speculative parallelization using software multi-threaded transactions · ASPLOS 2010
Parallel and multicore computing
pipeline parallelism
0.012010
Scalable Speculative Parallelization on Commodity Clusters · MICRO 2010

Methods — techniques the papers use, named apart from their topics

speculative execution · 0.2runtime system design · 0.2distributed software multi-threaded transactional memory · 0.1Spec-DSWP · 0.1DSWP · 0.1
YearPublicationVenuePosition
2024 Refute Questions for Concrete, Cluttered Specifications
abstract
Learners must be able to critique AI-generated code, which is often plausible but not always appropriate. Refute Questions have been proposed as one way to develop this ability. Each such question has two components: a task specification and a purported solution (typically, a single function) for that task. Our prior work [1, 4, 5] has focused on Refute Questions where purported solutions are inappropriate because they are logically incorrect. Given such a question, learners must demonstrate their understanding of both the buggy function and the specification by providing an input on which the function’s behavior does not satisfy the specification.
Sannidhi V. Hebbar, Sasmita Harini, Arun Raman, Viraj Kumar
ICER (2)3
2023 Helping Students Develop a Critical Eye with Refute Questions
abstract
To help students develop a critical eye for human and (increasingly) machine-generated artifacts such as code, documentation, and more, this workshop proposes Refute questions. Students are given an artifact created for a stated purpose and asked: Why does the artifact fail to serve that purpose? Students must provide evidence demonstrating this failure. After a hands-on introduction to Refute questions in their originally proposed context (an alternative to 'Explain in Plain English' questions), participants will receive and review a richer variety of Refute questions for autograded formative and summative assessments, targeting a CS1 course in Python. Then, based on their interest, they will join a subgroup to create theme-specific Refute questions and identify theme-specific challenges such as question-generation and evaluation strategies. Finally, as a group, participants will identify a wish-list of Refute-related tools to be created and research questions to be answered. The workshop caters to a diverse audience: CS instructors at all levels, designers of Intelligent Tutoring Systems and interactive e-books, and CS education researchers.
Viraj Kumar, Arun Raman
SIGCSE (2)2
2019 Sequential Synthesis of Supervisory Policies for Discrete-Event Systems Modeled by Petri Nets
abstract
It is often of interest to synthesize a supervisory policy for enforcing complex properties on the behaviour of a Discrete-Event System (DES). One way of doing this is by decomposing complex properties into simpler objectives and then synthesizing supervisors for those simpler objectives in a sequential manner. This approach is particularly convenient if the supervised-system can be represented using the same modeling framework at each stage of this sequential process. An additional desirable feature could be that the supervisory policy remain the same even if the initial-state of the DES were to change. In this paper, we consider Petri Net (PN) models of Discrete - Event Systems (DES) under a supervisory policy that enforces a desired-property B. We prove that the supervised-system can be modeled as a PN if and only if the supervisory policy is a marking-monotone-B-enforcing supervisory policy (MM-BESP) over reachable markings. In the second half of the paper we describe a software tool for the synthesis of MM-BESPs, where the desired-property B is the PN-property of liveness, for arbitrary Petri Nets. We end the paper with an example that illustrates both the contributions.
Arun Raman, Ramavarapu S. Sreenivas
SMC1
2015 Enabling Efficient Alias Speculation
abstract
Microprocessors designed using HW/SW codesign principles, such as Transmeta™ Efficeon™ and the soon-to-ship NVIDIA 64-bit Tegra® K1, use dynamic binary optimization to extract instruction-level parallelism. Many code optimizations are made significantly more effective through the use of alias speculation. The state-of-the-art alias speculation system, SMARQ, provides 40% speedup on average over a system with no alias speculation. This performance, however, comes at the cost of introducing new alias registers and increased power consumption due to new checks for validating speculation. Consequently, improving the efficiency of alias speculation by reducing alias register requirements and rationalizing speculation validation checks is critical for the viability of SMARQ. This paper presents alias coalescing, a novel technique to significantly improve the efficiency of SMARQ through a synergistic combination of compiler and microarchitectural techniques. By using a more compact encoding for memory access ranges for memory instructions, alias coalescing simultaneously reduces the alias register pressure in SMARQ by a geomean of 26.09% and 39.96%, and the dynamic alias checks by 20.73% and 33.87%, across the entire SPEC CINT2006 and SPEC CFP2006 suites respectively.
Soumyadeep Ghosh, Yongjun Park 0001, Arun Raman
LCTES3
2012 From sequential programming to flexible parallel execution
abstract
The embedded computing landscape is being transformed by three trends: growing demand for greater functionality and enriched user experience, increasing diversity and parallelism in the processing substrate, and an accelerating push for ever-greater energy efficiency. For programmers, these trends give rise to three challenges: writing code for a potentially heterogeneous architecture, extracting parallelism in software, and maximizing a multivariate (performance, power, energy, etc.) fitness function of user satisfaction which may vary with time. To meet these challenges, clarion calls have been issued for programmers to start writing software in new parallel programming models. Fundamentally, however, these proposals detract programmers from delivering new features and enriched user experience in the shortest time possible. This paper proposes to attract embedded systems programmers to a vertically integrated approach, comprising extensions to the sequential programming model, a parallelizing compiler, and an optimizing run-time system, to enable them to tackle all three challenges.
Arun Raman, Jae W. Lee, David I. August
CASES1
2012 Parcae: a system for flexible parallel execution
abstract
Workload, platform, and available resources constitute a parallel program's execution environment. Most parallelization efforts statically target an anticipated range of environments, but performance generally degrades outside that range. Existing approaches address this problem with dynamic tuning but do not optimize a multiprogrammed system holistically. Further, they either require manual programming effort or are limited to array-based data-parallel programs.
Arun Raman, Ayal Zaks, Jae W. Lee, David I. August
PLDI1
2011 A Very Fast Simulator for Exploring the Many-Core Future
abstract
Although multi-core architectures with a large number of cores ("many-cores'') are considered the future of computing systems, there are currently few practical tools to quickly explore both their design and general program scalability. In this paper, we present SiMany, a discrete-event-based many-core simulator able to support more than a thousand cores while being orders of magnitude faster than existing flexible approaches. One of the difficult challenges for a reasonably realistic many-core simulation is to model faithfully the potentially high concurrency a program can exhibit. SiMany uses a novel virtual time synchronization technique, called spatial synchronization, to achieve this goal in a completely local and distributed fashion, which diminishes interactions and preserves locality. Compared to previous simulators, it raises the level of abstraction by focusing on modeling concurrent interactions between cores, which enables fast coarse comparisons of high-level architecture design choices and parallel programs performance. Sequential pieces of code are executed natively for maximal speed. We exercise the simulator with a set of dwarf-like task-based benchmarks with dynamic control flow and irregular data structures. Scalability results are validated through comparison with a cycle-level simulator up to 64 cores. They are also shown consistent with well-known benchmark characteristics. We finally demonstrate how SiMany can be used to efficiently compare the benchmarks' behavior over a wide range of architectural organizations, such as polymorphic architectures and network of clusters.
Olivier Certner, Arun Raman, Olivier Temam
IPDPS3
2011 Sprint: speculative prefetching of remote data
abstract
Remote data access latency is a significant performance bottleneck in many modern programs that use remote databases and web services. We present Sprint - a run-time system for optimizing such programs by prefetching and caching data from remote sources in parallel to the execution of the original program. Sprint separates the concerns of exposing potentially-independent data accesses from the mechanism for executing them efficiently in parallel or in a batch. In contrast to prior work, Sprint can efficiently prefetch data in the presence of irregular or input-dependent access patterns, while preserving the semantics of the original program.
Arun Raman, Greta Yorsh, Martin T. Vechev, Eran Yahav
OOPSLA1
2011 Parallelism orchestration using DoPE: the degree of parallelism executive
abstract
In writing parallel programs, programmers expose parallelism and optimize it to meet a particular performance goal on a single platform under an assumed set of workload characteristics. In the field, changing workload characteristics, new parallel platforms, and deployments with different performance goals make the programmer's development-time choices suboptimal. To address this problem, this paper presents the Degree of Parallelism Executive (DoPE), an API and run-time system that separates the concern of exposing parallelism from that of optimizing it. Using the DoPE API, the application developer expresses parallelism options. During program execution, DoPE's run-time system uses this information to dynamically optimize the parallelism options in response to the facts on the ground. We easily port several emerging parallel applications to DoPE's API and demonstrate the DoPE run-time system's effectiveness in dynamically optimizing the parallelism for a variety of performance goals.
Arun Raman, Hanjun Kim 0001, Taewook Oh, Jae W. Lee, David I. August
PLDI1
2010 Speculative parallelization using software multi-threaded transactions
Arun Raman, Hanjun Kim 0001, Thomas R. Mason, Thomas B. Jablin, David I. August
ASPLOS1
2010 Decoupled software pipelining creates parallelization opportunities
abstract
Decoupled Software Pipelining (DSWP) is one approach to automatically extract threads from loops. It partitions loops into long-running threads that communicate in a pipelined manner via inter-core queues. This work recognizes that DSWP can also be an enabling transformation for other loop parallelization techniques. This use of DSWP, called DSWP+, splits a loop into new loops with dependence patterns amenable to parallelization using techniques that were originally either inapplicable or poorly-performing. By parallelizing each stage of the DSWP+ pipeline using (potentially) different techniques, not only is the benefit of DSWP increased, but the applicability and performance of other parallelization techniques are enhanced. This paper evaluates DSWP+ as an enabling framework for other transformations by applying it in conjunction with DOALL, LOCALWRITE, and SpecDOALL to individual stages of the pipeline. This paper demonstrates significant performance gains on a commodity 8-core multicore machine running a variety of codes transformed with DSWP+.
Jialu Huang, Arun Raman, Thomas B. Jablin, Yun Zhang 0005, Tzu-Han Hung, David I. August
CGO2
2010 Scalable Speculative Parallelization on Commodity Clusters
abstract
While clusters of commodity servers and switches are the most popular form of large-scale parallel computers, many programs are not easily parallelized for execution upon them. In particular, high inter-node communication cost and lack of globally shared memory appear to make clusters suitable only for server applications with abundant task-level parallelism and scientific applications with regular and independent units of work. Clever use of pipeline parallelism (DSWP), thread-level speculation (TLS), and speculative pipeline parallelism (Spec-DSWP) can mitigate the costs of inter-thread communication on shared memory multicore machines. This paper presents Distributed Software Multi-threaded Transactional memory (DSMTX), a runtime system which makes these techniques applicable to non-shared memory clusters, allowing them to efficiently address inter-node communication costs. Initial results suggest that DSMTX enables efficient cluster execution of a wider set of application types. For 11 sequential C programs parallelized for a 4-core 32-node (128 total core) cluster without shared memory, DSMTX achieves a geomean speedup of 49×. This compares favorably to the 15x speedup achieved by our implementation of TLS-only support for clusters.
Hanjun Kim 0001, Arun Raman, Jae W. Lee, David I. August
MICRO2
2008 Parallel-stage decoupled software pipelining
abstract
In recent years, the microprocessor industry has embraced chip multiprocessors (CMPs), also known as multi-core architectures, as the dominant design paradigm. For existing and new applications to make effective use of CMPs, it is desirable that compilers automatically extract thread-level parallelism from single-threaded applications. DOALL is a popular automatic technique for loop-level parallelization employed successfully in the domains of scientific and numeric computing. While DOALL generally scales well with the number of iterations of the loop, its applicability is limited by the presence of loop-carried dependences. A parallelization technique with greater applicability is decoupled software pipelining (DSWP), which parallelizes loops even in the presence of loop-carried dependences. However, the scalability of DSWP is limited by the size of the loop body and the number of recurrences it contains, which are usually smaller than the loop iteration count.
Easwaran Raman, Guilherme Ottoni, Arun Raman, Matthew J. Bridges, David I. August
CGO3