Jennifer M. Schopf

dblp:01/6038 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0003-0726-3674ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Distributed systems · 58% Performance modeling and evaluation · 19% Parallel and multicore computing · 10%
Computer networks
2 papers
Network performance modeling · 79% Internet of things and sensor networks · 21%

Topics — the 18 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
grid computing
0.242007
Anomaly detection and diagnosis in grid environments · SC 2007
Adaptive Computing on the Grid Using AppLeS · IEEE Trans. Parallel Distributed Syst. 2003
A Performance Study of Monitoring and Information Services for Distributed Systems · HPDC 2003
Performance modeling and evaluation › performance monitoring
monitoring data analysis
0.112007
Anomaly detection and diagnosis in grid environments · SC 2007
Distributed systems › distributed scheduling
application-level scheduling
0.122003
Adaptive Computing on the Grid Using AppLeS · IEEE Trans. Parallel Distributed Syst. 2003
Application-Level Scheduling on Distributed Heterogeneous Networks · SC 1996
Performance modeling and evaluation
workload characterization
0.122001
Multi-Resolution Resource Behaviour Queries Using Wavelets · HPDC 2001
Stochastic Scheduling · SC 1999
Distributed systems › grid computing
virtual organizations
0.012004
The Inca Test Harness and Reporting Framework · SC 2004
Parallel and multicore computing › parallel scheduling
adaptive scheduling
0.012003
Adaptive Computing on the Grid Using AppLeS · IEEE Trans. Parallel Distributed Syst. 2003
Cloud and datacenter computing
resource management
0.012003
Adaptive Computing on the Grid Using AppLeS · IEEE Trans. Parallel Distributed Syst. 2003
Distributed systems › service-oriented architecture › service management
service monitoring
0.012003
A Performance Study of Monitoring and Information Services for Distributed Systems · HPDC 2003
Parallel and multicore computing
task scheduling
0.012003
Conservative Scheduling: Using Predicted Variance to Improve Scheduling Decisions in Dynamic Environments · SC 2003
Network performance modeling › performance prediction
throughput prediction
0.012002
Predicting Sporadic Grid Data Transfers · HPDC 2002
Distributed systems › grid computing
data grid
0.012002
Predicting Sporadic Grid Data Transfers · HPDC 2002
Distributed systems › replication › replica management
replica selection
0.012002
Predicting Sporadic Grid Data Transfers · HPDC 2002
Distributed systems
resource monitoring
0.012001
Multi-Resolution Resource Behaviour Queries Using Wavelets · HPDC 2001
Cloud and datacenter computing
cluster resource management and scheduling
0.011999
Stochastic Scheduling · SC 1999
Performance modeling and evaluation
scheduling policy
0.011999
Stochastic Scheduling · SC 1999
High-performance computing
data transfer
0.012002
Predicting Sporadic Grid Data Transfers · HPDC 2002
Internet of things and sensor networks › data dissemination
sensor data transmission
0.012001
Multi-Resolution Resource Behaviour Queries Using Wavelets · HPDC 2001
Parallel and multicore computing › data parallelism
data-parallel applications
0.011999
Stochastic Scheduling · SC 1999

Methods — techniques the papers use, named apart from their topics

window-based detection · 0.1throughput observation · 0.1signal processing · 0.1regression model · 0.1network load characterization · 0.1multicast · 0.1test harness · 0.0time series prediction · 0.0stochastic scheduling · 0.0benchmarking · 0.0wavelet transform · 0.0
YearPublicationVenuePosition
2019 The Engagement and Performance Operations Center: EPOC
abstract
In 2018, the US National Science Foundation (NSF) funded the Engagement and Performance Operations Center (EPOC), a joint project between Indiana University (IU) and the Department of Energy's Energy Science Network (ESnet), to work with domain scientists to accelerate the ability of distributed collaborations to share data in order to reach broader science goals. The goal of this funding was to create an operations center for engagement - including definition of formal processes, tracking of engagements, and funded staff, not simply best effort by volunteers, with a goal of enabling digital societies to better share scientific data.
Edward Moynihan, Jennifer M. Schopf, Jason Zurawski
eScience2
2012 Scalable integrated performance analysis of multi-gigabit networks
abstract
Monitoring and managing multi-gigabit networks requires dynamic adaptation to end-to-end performance characteristics. This paper presents a measurement collection and analysis framework that automates the troubleshooting of end-to-end network bottlenecks. We integrate real-time host, application, and network measurements with a common representation (compatible with perfSONAR) within a flexible and scalable architecture. Our measurement architecture is supported by a light-weight eXtensible Session Protocol (XSP), which enables context-sensitive adaptive measurement collection. We evaluate the ability of our system to analyze and detect bottleneck conditions over a series of high-speed and I/O intensive bulk data transfer experiments and find that the overhead of the system is very low and that we are able to detect and understand a variety of bottlenecks.
Ezra Kissel, Ahmed El-Hassany, Guilherme Fernandes, D. Martin Swany, Dan Gunter, Taghrid Samak, Jennifer M. Schopf
NOMS7
2009 Sustainability and the Office of CyberInfrastructure
abstract
The National Science Foundationpsilas Office of CyberInfrastructure (OCI) supports a broad set of infrastructure, including hardware, software, and services to enable computational science in a variety of disciplines. Issues of sustainability and reusability in this context have become a priority. This paper addresses several approaches to sustainability, and how they OCIs task forces are addressing this critical need.
Jennifer M. Schopf
NCA1
2007 Anomaly detection and diagnosis in grid environments
abstract
Identifying and diagnosing anomalies in application behavior is critical to delivering reliable application-level performance. In this paper we introduce a strategy to detect anomalies and diagnose the possible reasons behind them. Our approach extends the traditional window-based strategy by using signal-processing techniques to filter out recurring, background fluctuations in resource behavior. In addition, we have developed a diagnosis technique that uses standard monitoring data to determine which related changes in behavior may cause anomalies. We evaluate our anomaly detection and diagnosis technique by applying it in three contexts when we insert anomalies into the system at random intervals. The experimental results show that our strategy detects up to 96% of anomalies while reducing the false positive rate by up to 90% compared to the traditional window average strategy. In addition, our strategy can diagnose the reason for the anomaly approximately 75% of the time.
Lingyun Yang, Chuang Liu 0006, Jennifer M. Schopf, Ian T. Foster
SC3
2007 Scalability analysis of three monitoring and information systems: MDS2, R-GMA, and Hawkeye
Xuehai Zhang, Jeffrey L. Freschl, Jennifer M. Schopf
J. Parallel Distributed Comput.3
2006 Statistical Data Reduction for Efficient Application Performance Monitoring
abstract
There is a growing need for systems that can monitor and analyze application performance data automatically in order to deliver reliable and sustained performance to applications. However, the continuously growing complexity of high performance computer systems and applications makes this process difficult. We introduce a statistical data reduction method that can be used to guide the selection of system metrics that are both necessary and sufficient to describe observed application behavior, thus reducing the instrumentation perturbation and data volume to be managed. To evaluate our strategy, we applied it to one CPU-bound grid application using cluster machines and GridFTP data transfer in a wide area testbed. A comparative study shows that our strategy produces better results than other techniques. It can reduce the number of system metrics to be managed by about 80%, while still capturing enough information for performance predictions.
Lingyun Yang, Jennifer M. Schopf, Catalin Dumitrescu, Ian T. Foster
CCGRID2
2006 Monitoring the Earth System Grid with MDS4
abstract
In production Grids for scientific applications, service and resource failures must be detected and addressed quickly. In this paper, we describe the monitoring infrastructure used by the Earth System Grid (ESG) project, a scientific collaboration that supports global climate research. ESG uses the Globus Toolkit Monitoring and Discovery System (MDS4) to monitor its resources. We describe how the MDS4 Index Service collects information about ESG resources and how the MDS4 Trigger Service checks specified failure conditions and notifies system administrators when failures occur. We present monitoring statistics for May 2006 and describe our experiences using MDS4 to monitor ESG resources over the last two years.
Ann L. Chervenak, Jennifer M. Schopf, Laura Pearlman, Mei-Hui Su, Shishir Bharathi, Luca Cinquini, Mike D'Arcy, Neill Miller, David E. Bernholdt
e-Science2
2005 Improving parallel data transfer times using predicted variances in shared networks
abstract
It is increasingly common to use multiple distributed storage systems as a single data store within which large datasets may be replicated. Thus, we face the problem of how to access replicated data efficiently. Multiple-source parallel transfers can reduce access times by transferring data from several replicas in parallel. However, we then face the problem of deciding which data to fetch from which replicas. We propose a Tuned Conservative scheduling technique that uses predicted means and variances for network performance to make data selection decisions. This stochastic scheduling technique adjusts the amount of data fetched on a link according to not only the link performance but the expected variance in that performance. We incorporate our technique into the striped GridFTP server from the Globus Toolkit, and demonstrate that the technique can produce data transfer times that are significantly faster and less variable than those of other techniques.
Lingyun Yang, Jennifer M. Schopf, Ian T. Foster
CCGRID2
2004 Performance analysis of the Globus Toolkit Monitoring and Discovery Service, MDS2
abstract
Monitoring and information services form a key component of a distributed system, or grid. A quantitative study of such services can aid in understanding the performance limitations, advise in the deployment of the monitoring system, and help evaluate future development work. To this end, we examined the performance of the Globus Toolkit/spl reg/ Monitoring and Discovery Service (MDS2) by instrumenting its main services using NetLogger. Our study shows a strong advantage to caching or prefetching the data, as well as the need to have primary components at well-connected sites.
Xuehai Zhang, Jennifer M. Schopf
IPCCC2
2004 The Inca Test Harness and Reporting Framework
abstract
Virtual organizations (VOs), communities that enable coordinated resource sharing among multiple sites, are becoming more prevalent in the high-performance computing community. In order to promote cross-site resource usability, most VOs prepare service agreements that include a minimum set of common resource functionality, starting with a common software stack and evolving into more complicated service and interoperability agreements. VO service agreements are often difficult to verify and maintain, however, because the sites are dynamic and autonomous. Automated verification of service agreements is critical: manual and user tests are not practical on a large scale. The Inca test harness and reporting framework is a generic system for the automated testing, data collection, verification, and monitoring of service agreements. This paper describes Inca’s architecture, system impact, and performance. Inca is being used by the TeraGrid project to verify software installations, monitor service availability, and collect performance data.
Shava Smallen, Catherine Mills Olschanowsky, Kate Ericson, Pete Beckman, Jennifer M. Schopf
SC5
2003 Run-Time Prediction of Parallel Applications on Shared Environments
abstract
Application run-time is a fundamental component in application and job scheduling. However, accurate predictions of run times are difficult to achieve for parallel applications running in shared environments where resource capacities can change dynamically over time. In this paper, we propose a run-time prediction technique for parallel applications that uses regression methods and filtering techniques to derive the application execution time without using standard performance models. The experimental results show that our use of regression models delivers tolerable prediction accuracy and that we can improve the accuracy dramatically by using appropriate filters.
Byoung-Dai Lee, Jennifer M. Schopf
CLUSTER2
2003 A Performance Study of Monitoring and Information Services for Distributed Systems
abstract
Monitoring and information services form a key component of a distributed system, or Grid. A quantitative study of such services can aid in understanding the performance limitations, advise in the deployment of the monitoring system, and help evaluate future development work. To this end, we study the performance of three monitoring and information services for distributed systems: the Globus Toolkit/spl reg/ Monitoring and Discovery Service (MDS2), the European Data Grid Relational Grid Monitoring Architecture (R-GMA) and Hawkeye, part of the Condor project. We perform experiments to test their scalability with respect to number of users, number of resources and amount of data collected. Our study shows that each approach has different behaviors, often due to their different design goals. In the four sets of experiments we conducted to evaluate the performance of the service components under different circumstances, we found a strong advantage to caching or pre-fetching the data, as well as the need to have primary components at well-connected sites because of the high load seen by all systems.
Xuehai Zhang, Jeffrey L. Freschl, Jennifer M. Schopf
HPDC3
2003 Conservative Scheduling: Using Predicted Variance to Improve Scheduling Decisions in Dynamic Environments
abstract
In heterogeneous and dynamic environments, efficient execution of parallel computations can require mappings of tasks to processors whose performance is both irregular (because of heterogeneity) and time-varying (because of dynamicity). While adaptive domain decomposition techniques have been used to address heterogeneous resource capabilities, temporal variations in those capabilities have seldom been considered. We propose a conservative scheduling policy that uses information about expected future variance in resource capabilities to produce more efficient data mapping decisions. We first present techniques, based on time series predictors that we developed in previous work, for predicting CPU load at some future time point, average CPU load for some future time interval, and variation of CPU load over some future time interval. We then present a family of stochastic scheduling algorithms that exploit such predictions of future availability and variability when making data mapping decisions. Finally, we describe experiments in which we apply our techniques to an astrophysics application. The results of these experiments demonstrate that conservative scheduling can produce execution times that are both significantly faster and less variable than other techniques.
Lingyun Yang, Jennifer M. Schopf, Ian T. Foster
SC2
2003 Adaptive Computing on the Grid Using AppLeS
abstract
Ensembles of distributed, heterogeneous resources, also known as computational grids, have emerged as critical platforms for high-performance and resource-intensive applications. Such platforms provide the potential for applications to aggregate enormous bandwidth, computational power, memory, secondary storage, and other resources during a single execution. However, achieving this performance potential in dynamic, heterogeneous environments is challenging. Recent experience with distributed applications indicates that adaptivity is fundamental to achieving application performance in dynamic grid environments. The AppLeS (Application Level Scheduling) project provides a methodology, application software, and software environments for adaptively scheduling and deploying applications in heterogeneous, multiuser grid environments. We discuss the AppLeS project and outline our findings.
Francine Berman, Richard Wolski, Henri Casanova, Walfredo Cirne, Holly Dail, Marcio Faerman, Silvia M. Figueira, Jim Hayes, Graziano Obertelli, Jennifer M. Schopf, Gary Shao, Shava Smallen, Neil Spring, Alan Su 0001, Dmitrii Zagorodnov
IEEE Trans. Parallel Distributed Syst.10
2002 Predicting Sporadic Grid Data Transfers
abstract
The increasingly common practice of replicating datasets and using resources as distributed data stores in grid environments has led to the problem of determining which replica can be accessed most efficiently. Due diverse performance characteristics and load variations of several components in the end-to-end path linking these various locations, selecting a replica from among many requires accurate prediction information of the data transfer times between the sources and sinks. In this paper we present a prediction system that is based on combining end-to-end application throughput observations and network load variations, capturing the whole-system performance and variations in load patterns, respectively. We develop a set of regression models to derive predictions that characterize the effect of network load variations on file transfer times. We apply these techniques to the GridFTP data movement tool, part of the Globus Toolkit/spl trade/, and observe performance gains of up to 10% in prediction accuracy when compared with approaches based on past system behavior in isolation.
Sudharshan S. Vazhkudai, Jennifer M. Schopf
HPDC2
2002 Current Activities in the Scheduling and Resource Management Area of the Global Grid Forum
Bill Nitzberg, Jennifer M. Schopf
JSSPP2
2001 Multi-Resolution Resource Behaviour Queries Using Wavelets
abstract
Different adaptive applications are interested in the dynamic behavior of a resource over different fine- to coarse-grain time-scales. The resource's sensor runs at some fine-grain resource-appropriate sampling rate, producing a discrete-time resource signal. It can be very inefficient to to answer a coarse-grain application query by directly using the fine-grain resource signal. We address this gap between the sensor and its different client applications with a novel query model that explicitly incorporates time-scale as a parameter. The query model is implemented on top of an inherently multi-scale wavelet-based representation of the signal (which could be communicated over a set of multicast channels). A query uses only the wavelet coefficients necessary for its time-scale (and thus could listen to a subset of the channels), greatly reducing the data that need to be communicated. We present very promising initial results on host load signals, showing the tradeoff between compactness and query error. Finally, we describe some of the other operations that the wavelet representation enables.
Jason A. Skicewicz, Peter A. Dinda, Jennifer M. Schopf
HPDC3
1999 Stochastic Scheduling
abstract
There is a current need for scheduling policies that can leverage the performance variability of resources on multiuser clusters. We develop one solution to this problem called stochastic scheduling that utilizes a distribution of application execution performance on the target resources to determine a performance-efficient schedule. In this paper, we define a stochastic scheduling policy based on time-balancing for data parallel applications whose execution behavior can be represented as a normal distribution. Using three distributed applications on two contended platforms, we demonstrate that a stochastic scheduling policy can achieve good and predictable performance for the application as evaluated by several performance measures.
Jennifer M. Schopf, Francine Berman
SC1
1996 Application-Level Scheduling on Distributed Heterogeneous Networks
abstract
Heterogeneous networks are increasingly being used as platforms for resource-intensive distributed parallel applications. A critical contributor to the performance of such applications is the scheduling of constituent application tasks on the network. Since often the distributed resources cannot be brought under the control of a single global scheduler, the application must be scheduled by the user. To obtain the best performance, the user must take into account both application-specific and dynamic system information in developing a schedule which meets his or her performance criteria. In this paper, we define a set of principles underlying application-level scheduling and describe our work-in-progress building AppLeS (application-level scheduling) agents. We illustrate the application-level scheduling approach with a detailed description and results for a distributed 2D Jacobi application on two production heterogeneous platforms.
Francine Berman, Richard Wolski, Silvia M. Figueira, Jennifer M. Schopf, Gary Shao
SC4