Ganesh Narayanaswamy

dblp:12/6271 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-authorComputer networks · 1 · 1 first-authorTheory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
4 papers
Concurrent programming · 61% Program verification · 24% Program analysis · 15%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 94% Processor architecture and microarchitecture · 6%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming
deadlock detection
0.732017
Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs · ACM Trans. Program. Lang. Syst. 2017
When truth is efficient: analysing concurrency · ISSTA 2015
Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs · FM 2014
Program analysis
static analysis
0.312017
Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs · ACM Trans. Program. Lang. Syst. 2017
Concurrent programming
memory models
0.212016
The virtues of conflict: analysing modern concurrency · PPoPP 2016
Concurrent programming
concurrency analysis
0.212015
When truth is efficient: analysing concurrency · ISSTA 2015
Program verification
model checking
0.212014
Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs · FM 2014
Parallel and multicore computing
MPI
0.122015
When truth is efficient: analysing concurrency · ISSTA 2015
Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs · FM 2014
Parallel and multicore computing
parallel programming models
0.122015
When truth is efficient: analysing concurrency · ISSTA 2015
Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs · FM 2014
Parallel and multicore computing › parallel programming models › message passing
MPI applications
0.112017
Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs · ACM Trans. Program. Lang. Syst. 2017
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling
0.112008
Asymmetric interactions in symmetric multi-core systems: analysis, enhancements and evaluation · SC 2008
Processor architecture and microarchitecture
multicore design
0.012008
Asymmetric interactions in symmetric multi-core systems: analysis, enhancements and evaluation · SC 2008

Methods — techniques the papers use, named apart from their topics

event structures · 0.7trace analysis · 0.6partial order modeling · 0.6SAT solving · 0.6true concurrency semantics · 0.4predictive analysis · 0.4dynamic analysis · 0.4symbolic semantics · 0.2partial order semantics · 0.2performance evaluation · 0.1
YearPublicationVenuePosition
2017 Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs
abstract
The Message Passing Interface (MPI) is the standard API for parallelization in high-performance and scientific computing. Communication deadlocks are a frequent problem in MPI programs, and this article addresses the problem of discovering such deadlocks. We begin by showing that if an MPI program is single path, the problem of discovering communication deadlocks is NP-complete. We then present a novel propositional encoding scheme that captures the existence of communication deadlocks. The encoding is based on modeling executions with partial orders and implemented in a tool called MOPPER . The tool executes an MPI program, collects the trace, builds a formula from the trace using the propositional encoding scheme, and checks its satisfiability. Finally, we present experimental results that quantify the benefit of the approach in comparison to other analyzers and demonstrate that it offers a scalable solution for single-path programs.
Vojtech Forejt, Saurabh Joshi 0001, Daniel Kroening, Ganesh Narayanaswamy, Subodh Sharma 0001
ACM Trans. Program. Lang. Syst.4
2016 The virtues of conflict: analysing modern concurrency
abstract
Modern shared memory multiprocessors permit reordering of memory operations for performance reasons. These reorderings are often a source of subtle bugs in programs written for such architectures. Traditional approaches to verify weak memory programs often rely on interleaving semantics, which is prone to state space explosion, and thus severely limits the scalability of the analysis. In recent times, there has been a renewed interest in modelling dynamic executions of weak memory programs using partial orders. However, such an approach typically requires ad-hoc mechanisms to correctly capture the data and control-flow choices/conflicts present in real-world programs. In this work, we propose a novel, conflict-aware, composable, truly concurrent semantics for programs written using C/C++ for modern weak memory architectures. We exploit our symbolic semantics based on general event structures to build an efficient decision procedure that detects assertion violations in bounded multi-threaded programs. Using a large, representative set of benchmarks, we show that our conflict-aware semantics outperforms the state-of-the-art partial-order based approaches.
Ganesh Narayanaswamy, Saurabh Joshi 0001, Daniel Kroening
PPoPP1
2015 When truth is efficient: analysing concurrency
abstract
Concurrent systems are hard to develop and are even harder to analyse. The usual way to analyse concurrent systems is to give them interleaving semantics and exploit automata-based methods to investigate the resultant interleaved model. Such an approach is often hard to scale without additional tools to curb the interleaving-induced state space explosion. In this work we make an alternate case: for directly capturing the behaviour of concurrent systems using true concurrency. We show how to build composable, truly concurrent models for real-world programs written using one of the most widely adopted paradigms for developing massively parallel systems, the Message Passing Interface Standard (MPI). Our method employs general event structures to symbolically capture executions of MPI programs and uses this truly concurrent model, combined with our novel deadlock characterisation, to formulate a precise, scalable decision procedure that finds communication deadlocks in large MPI programs. We show that our analysis scales to systems with hundreds of processes and strongly outperforms state of the art interleaving semantics based approaches.
Ganesh Narayanaswamy
ISSTA1
2014 Precise Predictive Analysis for Discovering Communication Deadlocks in MPI Programs
Vojtech Forejt, Daniel Kroening, Ganesh Narayanaswamy, Subodh Sharma 0001
FM3
2008 Impact of Network Sharing in Multi-Core Architectures
abstract
As commodity components continue to dominate the realm of high-end computing, two hardware trends have emerged as major contributors-high-speed networking technologies and multi-core architectures. Communication middleware such as the Message Passing Interface (MPI) uses the network technology for communicating between processes that reside on different physical nodes, while using shared memory for communicating between processes on different cores within the same node. Thus, two conflicting possibilities arise: (i) with the advent of multi-core architectures, the number of processes that reside on the same physical node and hence share the same physical network can potentially increase significantly, resulting in increased network usage, and (ii) given the increase in intra-node shared-memory communication for processes residing on the same node, the network usage can potentially decrease significantly. In this paper, we address these two conflicting possibilities and study the behavior of network usage in multi-core environments with sample scientific applications. Specifically, we analyze trends that result in increase or decrease of network usage, and we derive insights into application performance based on these. We also study the sharing of different resources in the system in multi-core environments and identify the contribution of the network in this mix. In addition, we study different process allocation strategies and analyze their impact on such network sharing.
Ganesh Narayanaswamy, Pavan Balaji, Wu-chun Feng
ICCCN1
2008 Asymmetric interactions in symmetric multi-core systems: analysis, enhancements and evaluation
abstract
Multi-core architectures have spurred the rapid growth in high-end computing systems. While the vast majority of such multi-core processors contain symmetric hardware components, their interaction with systems software, in particular the communication stack, results in a remarkable amount of asymmetry in the effective capability of the different cores. In this paper, we analyze such interactions and propose a novel management library called SyMMer (Systems Mapping Manager) that monitors these interactions and dynamically manages the mapping of processes on processor cores to transparently improve application performance. Together with a detailed description of the SyMMer library, we also present performance evaluation comparing SyMMer to a vanilla communication library using various micro-benchmarks as well as popular applications and scientific libraries. Experimental results demonstrate more than a two-fold improvement in communication time and 10-15% improvement in overall application performance.
Thomas Scogland, Pavan Balaji, Wu-chun Feng, Ganesh Narayanaswamy
SC4