Sharad Singhai

dblp:48/2889 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 87% Program analysis · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines
garbage collection
0.012001
Pretenuring for Java · OOPSLA 2001
Runtime systems and virtual machines › garbage collection
pretenuring
0.012001
Pretenuring for Java · OOPSLA 2001
Program analysis › heap analysis
allocation site analysis
0.012001
Pretenuring for Java · OOPSLA 2001

Methods — techniques the papers use, named apart from their topics

build-time advice · 0.0allocation site profiling · 0.0
YearPublicationVenuePosition
2007 An integrated ARM and multi-core DSP simulator
abstract
In this paper we describe the design and implementation of a flexible, and extensible, just-in-time ARM simulator designed to run cooperatively with a multi-core DSP simulator on x86 hosts. The integrated simulator can boot ARM/Linux alongside another operating system running on DSP cores, thus truly supporting a heterogeneous multi-core operating environment. In addition, the simulator facilitates exploration of several system design parameters such as memory latencies, cache organization etc. via lightweight user-defined instrumentation.
Sharad Singhai, MingYung Ko, Sanjay Jinturkar, Mayan Moudgill, C. John Glossner
CASES1
2001 Pretenuring for Java
abstract
Pretenuring can reduce copying costs in garbage collectors by allocating long-lived objects into regions that the garbage collector with rarely, if ever, collect. We extend previous work on pretenuring as follows. (1) We produce pretenuring advice that is neutral with respect to the garbage collector algorithm and configuration. We thus can and do combine advice from different applications. We find that predictions using object lifetimes at each allocation site in Java prgroams are accurate, which simplifies the pretenuring implementation. (2) We gather and apply advice to applications and the Jalapeño JVM, a compiler and run-time system for Java written in Java. Our results demonstrate that building combined advice into Jalapeño from different application executions improves performance regardless of the application Jalapeño is compiling and executing. This build-time advice thus gives user applications some benefits of pretenuring without any application profiling. No previous work pretenures in the run-time system. (3) We find that application-only advice also improves performance, but that the combination of build-time and application-specific advice is almost always noticeably better. (4) Our same advice improves the performance of generational and Older First colleciton, illustrating that it is collector neutral.
Steve Blackburn, Sharad Singhai, Matthew Hertz, Kathryn S. McKinley, J. Eliot B. Moss
OOPSLA2
1997 A Parametrized Loop Fusion Algorithm for Improving Parallelism and Cache Locality
abstract
Loop fusion is a reordering transformation that merges multiple loops into a single loop. It can increase data locality and the granularity of parallel loops, thus improving program performance. Previous approaches to this problem have looked at these two benefits in isolation. In this work, we propose a new model which considers data locality, parallelism and register pressure together. We build a weighted directed acyclic graph in which the nodes represent program loops along with their register pressure, and the edges represent the amount of locality and parallelism present. The direction of an edge represents an execution order constraint. We then partition the graph into components such that the sum of the weights on the edges cut is minimized, subject to the constraint that the nodes in the same partition can be safely fused together, and the register pressure of the combined loop does not exceed the number of available registers. Previous work demonstrates that the general problem of finding optimal partitions is NP-hard. In restricted cases, we show that it is possible to arrive at the optimal solution. We give an algorithm for the restricted case and a heuristic for the general case. We demonstrate the effectiveness of fusion and our approach with experimental results.
Sharad Singhai, Kathryn S. McKinley
Comput. J.1