Tushar Mohan

dblp:43/4751 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 3Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 67% Memory systems · 33%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.112007
METRIC: Memory tracing via dynamic binary rewriting to identify cache inefficiencies · ACM Trans. Program. Lang. Syst. 2007
Performance modeling and evaluation › tracing
memory reference tracing
0.112007
METRIC: Memory tracing via dynamic binary rewriting to identify cache inefficiencies · ACM Trans. Program. Lang. Syst. 2007
Memory systems
memory access optimization
0.012003
Identifying and Exploiting Spatial Regularity in Data Memory References · SC 2003
Performance modeling and evaluation › profiling
memory access profiling
0.012003
Identifying and Exploiting Spatial Regularity in Data Memory References · SC 2003
Performance modeling and evaluation
workload characterization
0.012003
Identifying and Exploiting Spatial Regularity in Data Memory References · SC 2003

Methods — techniques the papers use, named apart from their topics

trace compression · 0.1dynamic binary rewriting · 0.1profile-driven optimization · 0.0online parallel stream detection · 0.0
YearPublicationVenuePosition
2021 Truth and travesty intertwined: a case study of #SSR counterpublic campaign
abstract
Twitter has emerged as a prominent social media platform for activism and counterpublic narratives. The counterpublics leverage hashtags to build a diverse support network and share content on a global platform that counters the dominant narrative. This paper applies the framework of connective action on the counter-narrative campaign over the cause of death of #SushantSinghRajput. We combine descriptive network, modularity, and hashtag based topical analysis to identify three major mechanisms underlying the campaign: generative role taking, hashtag-based narratives and formation of alignment network towards a common cause. Using the case study of #SushantSinghRajput, we highlight how connective action framework can be used to identify different strategies adopted by counterpublics for the emergence of connective action.
Kumari Neha 0001, Tushar Mohan, Arun Balaji Buduru, Ponnurangam Kumaraguru
ASONAM2
2013 PAPI 5: Measuring power, energy, and the cloud
abstract
The PAPI library [1] was originally developed to provide portable access to the hardware performance counters found on a diverse collection of modern microprocessors. Rather than learning and writing to a new performance infrastructure each time code is moved to a new machine, measurement code can be written to the PAPI API which abstracts away the underlying interface. Over time, other system components besides the processor have gained performance interfaces (for example, GPUs and network interfaces). PAPI was redesigned to have a component architecture to allow modular access to these new sources of performance data [2]. In addition to incremental changes in processor support, the recent PAPI 5 release adds support for two emerging concerns in the high-performance landscape: energy consumption and cloud computing. As processor densities climb, the thermal properties and energy usage of high performance systems are becoming increasingly important. We have extended the PAPI interface to simultaneously monitor processor metrics, thermal sensors, and power meters to provide clues for correlating algorithmic activity with thermal response and energy consumption. We have also extended PAPI to provide support for running inside of Virtual Machines (VMs). This ongoing work will enable developers to use PAPI to engage in performance analysis in a virtualized cloud environment.
Vincent M. Weaver, Daniel Terpstra, Heike McCraw, Matt Johnson 0002, Kiran Kasichayanula, James Ralph, John Nelson, Philip Mucci, Tushar Mohan, Shirley Moore
ISPASS9
2010 An Open Source performance tools software suite for scientific computing
abstract
Abstract With the rapid replacement of closed, homogeneous, proprietary HPC systems by heterogeneous, Linux‐MPI cluster systems, the state of performance monitoring and analysis tools has become a cause for concern. Proprietary systems, despite their drawbacks, provided consistent tools of high quality. Modern Linux cluster systems, on the other hand, benefit from a wide variety of Open Source tools in differing stages of evolution. Recognizing that Linux clusters are here to stay, SiCortex has taken a unique approach of integrating and enhancing Open Source tools into a production‐quality suite. Further, as a tribute to the unrewarded Open Source community developers and for more pragmatic reasons, such as long‐term sustainability, changes made to the tools are fed upstream to the original tool developers. In this paper, we present an overview of the SiCortex tools' suite, and some of the challenges and successes we had in the process of realizing it. Copyright © 2009 John Wiley & Sons, Ltd.
Philip Mucci, Tushar Mohan
Concurr. Comput. Pract. Exp.2
2007 The software interface for a cluster interconnect based on the Kautz digraph
abstract
The Kautz digraph (Elspas et al., 1968) has been described as the ideal communication network for parallel computers, but it is generally unknown in the engineering community, and has never previously been used in a commercial product. We will define and characterize it in relation to HPC clusters, then discuss some of the implementation issues encountered in developing it as a large-scale cluster interconnect. The software interface to this interconnect is explained with some initial experiences with performance benchmarks.
Jud Leonard, Avi Purkayastha, Matt Reilly, Tushar Mohan
CLUSTER4
2007 METRIC: Memory tracing via dynamic binary rewriting to identify cache inefficiencies
abstract
With the diverging improvements in CPU speeds and memory access latencies, detecting and removing memory access bottlenecks becomes increasingly important. In this work we present METRIC, a software framework for isolating and understanding such bottlenecks using partial access traces. METRIC extracts access traces from executing programs without special compiler or linker support. We make four primary contributions. First, we present a framework for extracting partial access traces based on dynamic binary rewriting of the executing application. Second, we introduce a novel algorithm for compressing these traces. The algorithm generates constant space representations for regular accesses occurring in nested loop structures. Third, we use these traces for offline incremental memory hierarchy simulation. We extract symbolic information from the application executable and use this to generate detailed source-code correlated statistics including per-reference metrics, cache evictor information, and stream metrics. Finally, we demonstrate how this information can be used to isolate and understand memory access inefficiencies. This illustrates a potential advantage of METRIC over compile-time analysis for sample codes, particularly when interprocedural analysis is required.
Jaydeep Marathe, Frank Mueller 0001, Tushar Mohan, Sally A. McKee, Bronis R. de Supinski, Andy B. Yoo
ACM Trans. Program. Lang. Syst.3
2003 METRIC: Tracking Down Inefficiencies in the Memory Hierarchy via Binary Rewriting
abstract
We present METRIC, an environment for determining memory inefficiencies by examining data traces. METRIC is designed to alter the performance behavior of applications that are mostly constrained by their latency to resolve memory references. We make four primary contributions. First, we present methods to extract partial data traces from running applications by observing their memory behavior via dynamic binary rewriting. Second, we present a methodology to represent partial data traces in constant space for regular references through a novel technique for online compression of reference streams. Third, we employ offline cache simulation to derive indications about memory performance bottlenecks from partial data traces. By exploiting summarized memory metrics, by-reference metrics as well as cache evictor information, we can pin-point the sources of performance problems. Fourth, we demonstrate the ability to derive opportunities for optimizations and assess their benefits in several experiments resulting in up to 40% lower miss ratios.
Jaydeep Marathe, Frank Mueller 0001, Tushar Mohan, Bronis R. de Supinski, Sally A. McKee, Andy B. Yoo
CGO3
2003 Identifying and Exploiting Spatial Regularity in Data Memory References
abstract
The growing processor/memory performance gap causes the performance of many codes to be limited by memory accesses. If known to exist in an application, strided memory accesses forming streams can be targeted by optimizations such as prefetching, relocation, remapping, and vector loads. Undetected, they can be a significant source of memory stalls in loops. Existing stream-detection mechanisms either require special hardware, which may not gather statistics for subsequent analysis, or are limited to compile-time detection of array accesses in loops. Formally, little treatment has been accorded to the subject; the concept of locality fails to capture the existence of streams in a program's memory accesses. The contributions of this paper are as follows. First, we define spatial regularity as a means to discuss the presence and effects of streams. Second, we develop measures to quantify spatial regularity, and we design and implement an on-line, parallel algorithm to detect streams - and hence regularity - in running applications. Third, we use examples from real codes and common benchmarks to illustrate how derived stream statistics can be used to guide the application of profile-driven optimizations. Overall, we demonstrate the benefits of our novel regularity metric as an instrument to detect potential for code optimizations affecting memory performance.
Tushar Mohan, Bronis R. de Supinski, Sally A. McKee, Frank Mueller 0001, Andy B. Yoo, Martin Schulz 0001
SC1