Suvas Vajracharya

dblp:29/5866 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
0since 2021 · last 2001
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
loop transformation
0.011997
Loop Re-Ordering and Pre-Fetching at Run-time · SC 1997
Memory systems › cache
cache miss reduction
0.011997
Loop Re-Ordering and Pre-Fetching at Run-time · SC 1997
Memory systems › cache
cache optimization
0.011997
Loop Re-Ordering and Pre-Fetching at Run-time · SC 1997
Compilers and program optimization
prefetching
0.011997
Loop Re-Ordering and Pre-Fetching at Run-time · SC 1997

Methods — techniques the papers use, named apart from their topics

run-time dependence-driven execution · 0.0
YearPublicationVenuePosition
2001 Asynchronous Resource Management
abstract
As the organization of high-performance computers becomes more complex, the task of managing resources on them becomes increasingly difficult. The software layers which include the operating system, the runtime system, and the compiler must now map applications to machine architectures that consist of multiple CPUs, several layers of cache, deep memory hierarchies, file systems, interconnects, network interface cards, along the logical resources defined in the software system. To fully utilize all the available resources, software systems may use multiple sequential processes or threads that act on the passive resources of the system. This paper introduces a resource-centric event-driven model, where resources are active objects. We describe an algorithm that implements this model and show that this can significantly improve the performance of a wide variety of applications.
Suvas Vajracharya, Daniel G. Chavarría-Miranda
IPDPS1
1999 SMARTS: exploiting temporal locality and parallelism through vertical execution
abstract
Article SMARTS: exploiting temporal locality and parallelism through vertical execution Share on Authors: Suvas Vajracharya Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , Steve Karmesin Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , Peter Beckman Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , James Crotinger Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , Allen Malony Dept. of Computer and Information Science, University of Oregon and Los Alamos National Laboratory, Los Alamos, NM Dept. of Computer and Information Science, University of Oregon and Los Alamos National Laboratory, Los Alamos, NMView Profile , Sameer Shende Dept. of Computer and Information Science, University of Oregon and Los Alamos National Laboratory, Los Alamos, NM Dept. of Computer and Information Science, University of Oregon and Los Alamos National Laboratory, Los Alamos, NMView Profile , Rod Oldehoeft Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile , Stephen Smith Los Alamos National Laboratory, Los Alamos, NM Los Alamos National Laboratory, Los Alamos, NMView Profile Authors Info & Claims ICS '99: Proceedings of the 13th international conference on SupercomputingJune 1999 Pages 302–310https://doi.org/10.1145/305138.305207Online:01 May 1999Publication History 16citation354DownloadsMetricsTotal Citations16Total Downloads354Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Suvas Vajracharya, Steve Karmesin, Pete Beckman, James Crotinger, Allen D. Malony, Sameer Shende, R. R. Oldehoeft, Stephen Smith 0002
International Conference on Supercomputing1
1998 Dependence Driven Execution for Multiprogrammed Multiprocessor
abstract
Abstract Barrier synchronizations can be very expensive on multiprogramming environment because no process can go past a barrier until all the processes have arrived. If a process participating at a barrier is swapped out by the operating system, the rest of participating processes end up waiting for the swapped-out process. This paper presents a compile-time/run-time system that uses a dependence-driven execution to overlap the execution of computations separated by barriers so that the processes do not spend most of the time idling at the synchronization point. Keywords: Run-time systems, multiprogramming, loop scheduling, dependence-driven execution, barrier synchronization, coarse-grain dataflow. The parallel execution of a sequence of loop nests is typically broken into phases, each phase consisting of a simple loop separated by a barrier synchronization to ensure that the execution respects
Suvas Vajracharya, Dirk Grunwald
International Conference on Supercomputing1
1997 Loop Re-Ordering and Pre-Fetching at Run-time
abstract
The order in which loop iterations are executed can have a large impact on the number of cache misses that an applications takes. A new loop order that preserves the semantics of the old order but has a better cache data re-use, improves the performance of that application. Several compiler techniques exist to transform loops such that the order of iterations reduces cache misses. This paper introduces a run-time method to determine the order based on a dependence-driven execution. In a dependence-driven execution, an execution traverses the iteration space by following the dependence arcs between the iterations.
Suvas Vajracharya, Dirk Grunwald
SC1