Joshi Fullop

dblp:47/9840 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 44% Performance modeling and evaluation · 44% Distributed systems · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › profiling
application profiling
0.212014
The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014
High-performance computing
system monitoring
0.212014
The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014
Distributed systems › observability
large-scale monitoring
0.112014
The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014

Methods — techniques the papers use, named apart from their topics

lightweight distributed monitoring · 0.2
YearPublicationVenuePosition
2014 It takes a village: Monitoring the blue waters supercomputer
abstract
The performance of science applications on modern HPC equipment depends on many factors. Architectural features, individual hardware characteristics, and scheduler traits all have an impact on how a particular application performs, not only in isolation but when run in concert with other user applications. Being able to correlate system events and conditions at particular times can give insight into causes of good or bad performance. Unfortunately, the information we seek is not necessarily in a readily accessible form. The problem at hand is how to enable efficient query of the raw data and flexible graphical representation of the results. Web applications that access an underlying database serve this sort of functionality for many science applications quite well. Our scenario of data access is not very different. The data collected for a large HPC environment is complex and grows in size with time. This aspect is different from applications that deal with more static data. It is the dynamic nature of the data that make the problem interesting. In this work we present our approach for the analysis and visualization of HPC system performance data based on database access and web based graphical presentation. We discuss the details of how data is collected and processed from raw logs into the database, how queries are formulated, and how the data are graphically displayed. This process includes dynamic formulation of the queries. Finally we discuss how the system is utilized to analyze system performance.
Bart D. Semeraro, Robert Sisneros, Joshi Fullop, Gregory H. Bauer
CLUSTER3
2014 The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications
abstract
Understanding how resources of High Performance Compute platforms are utilized by applications both individually and as a composite is key to application and platform performance. Typical system monitoring tools do not provide sufficient fidelity while application profiling tools do not capture the complex interplay between applications competing for shared resources. To gain new insights, monitoring tools must run continuously, system wide, at frequencies appropriate to the metrics of interest while having minimal impact on application performance. We introduce the Lightweight Distributed Metric Service for scalable, lightweight monitoring of large scale computing systems and applications. We describe issues and constraints guiding deployment in Sandia National Laboratories' capacity computing environment and on the National Center for Supercomputing Applications' Blue Waters platform including motivations, metrics of choice, and requirements relating to the scale and specialized nature of Blue Waters. We address monitoring overhead and impact on application performance and provide illustrative profiling results.
Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Paul Cassella, Jeremy Enos, Joshi Fullop, Ann C. Gentile, Steve Monk, Nichamon Naksinehaboon, Jeff Ogden, Mahesh Rajan, Michael T. Showerman, Joel Stevenson, Narate Taerat, Thomas W. Tucker
SC6