EDBT 2026 Demo / reviewers in the wild / expert
Jonathan E. Cook 0001
dblp:52/4021 · also Jonathan Cook 0001
· DBLP profile ↗
30ranked-venue papers
14as first author
6since 2021 · last 2025
0009-0008-5795-9390ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 15 · 11 first-author · 1 since 2021Systems, architecture and hardware · 6 · 2 since 2021Computer networks · 6 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Storage systems · 47% Distributed systems · 33% High-performance computing · 10% | |
| Software engineering, system software, and programming languages
7 papers |
Requirements engineering and software design · 35% Empirical software engineering · 28% Software maintenance and evolution · 20% |
Topics — the 18 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › fault tolerance
failure recovery |
0.6 | 1 | 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File Systems · ACM Trans. Storage 2022 |
Storage systems › file systems › distributed file system
parallel file system |
0.6 | 1 | 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File Systems · ACM Trans. Storage 2022 |
Hardware reliability and fault tolerance
fault injection |
0.2 | 1 | 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File Systems · ACM Trans. Storage 2022 |
Storage systems
storage reliability |
0.2 | 1 | 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File Systems · ACM Trans. Storage 2022 |
Program analysis
dynamic analysis |
0.1 | 2 | 2003 | ICSE Workshop on Dynamic Analysis (WODA 2003) · ICSE 2003 Event-Base Detection of Concurrency · SIGSOFT FSE 1998 |
Requirements engineering and software design › software process
process model |
0.0 | 2 | 1999 | Software Process Validation: Quantitatively Measuring the Correspondence of a Process to a Model · ACM Trans. Softw. Eng. Methodol. 1999 Discovering Models of Software Processes from Event-Based Data · ACM Trans. Softw. Eng. Methodol. 1998 |
Requirements engineering and software design
software process |
0.0 | 2 | 1999 | Software Process Validation: Quantitatively Measuring the Correspondence of a Process to a Model · ACM Trans. Softw. Eng. Methodol. 1999 Discovering Models of Software Processes from Event-Based Data · ACM Trans. Softw. Eng. Methodol. 1998 |
Empirical software engineering
mining software repositories |
0.0 | 2 | 1998 | Cost-Effective Analysis of In-Place Software Processes · IEEE Trans. Software Eng. 1998 Discovering Models of Software Processes from Event-Based Data · ACM Trans. Softw. Eng. Methodol. 1998 |
Empirical software engineering › process mining
process discovery |
0.0 | 2 | 1998 | Discovering Models of Software Processes from Event-Based Data · ACM Trans. Softw. Eng. Methodol. 1998 Automating Process Discovery Through Event-Data Analysis · ICSE 1995 |
Storage systems › flash and SSD › flash memory management › garbage collection
partitioned garbage collection |
0.0 | 2 | 1998 | A Highly Effective Partition Selection Policy for Object Database Garbage Collection · IEEE Trans. Knowl. Data Eng. 1998 Partition Selection Policies in Object Database Garbage Collection · SIGMOD Conference 1994 |
Storage systems › storage management
storage reclamation |
0.0 | 2 | 1996 | Semi-automatic, Self-adaptive Control of Garbage Collection Rates in Object Databases · SIGMOD Conference 1996 Partition Selection Policies in Object Database Garbage Collection · SIGMOD Conference 1994 |
Software maintenance and evolution
dynamic software updating |
0.0 | 1 | 1999 | Highly Reliable Upgrading of Components · ICSE 1999 |
Empirical software engineering › software analytics
process data analysis |
0.0 | 1 | 1998 | Cost-Effective Analysis of In-Place Software Processes · IEEE Trans. Software Eng. 1998 |
Requirements engineering and software design › software architecture › software architecture analysis
software architecture recovery |
0.0 | 1 | 1998 | Event-Base Detection of Concurrency · SIGSOFT FSE 1998 |
Software maintenance and evolution
software process improvement |
0.0 | 1 | 1998 | Cost-Effective Analysis of In-Place Software Processes · IEEE Trans. Software Eng. 1998 |
Storage systems › data management
object database |
0.0 | 2 | 1996 | Semi-automatic, Self-adaptive Control of Garbage Collection Rates in Object Databases · SIGMOD Conference 1996 Partition Selection Policies in Object Database Garbage Collection · SIGMOD Conference 1994 |
Requirements engineering and software design › software process
process conformance |
0.0 | 1 | 1999 | Software Process Validation: Quantitatively Measuring the Correspondence of a Process to a Model · ACM Trans. Softw. Eng. Methodol. 1999 |
Performance modeling and evaluation › simulation
simulation-based evaluation |
0.0 | 1 | 1998 | A Highly Effective Partition Selection Policy for Object Database Garbage Collection · IEEE Trans. Knowl. Data Eng. 1998 |
Methods — techniques the papers use, named apart from their topics
black-box fault injection · 0.6trace-driven simulation · 0.0process discovery · 0.0process validation metrics · 0.0simulation · 0.0probabilistic analysis · 0.0markov methods · 0.0markov method · 0.0exploratory empirical study · 0.0case study · 0.0neural network · 0.0markov model · 0.0algorithmic grammar inference · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Machine Learning and AI Classification of HPC Job Quality Using Time-Series Heartbeat DataabstractDeciding whether or not a production job used its resources well is an ongoing concern in high performance computing (HPC). In this paper we show that low-overhead heartbeat data can be used to classify a job into a job quality level, using both machine learning and artificial intelligence techniques. Five HPC applications (one real, four mini) were used to create heartbeat datasets labeled with a job quality level, and then seven different ML and AI techniques were evaluated by training application-specific models for classifying runs into four different levels of job quality. The results show that, with very high accuracy, heartbeat data can serve as a mechanism for understanding production HPC job quality. Job quality feedback in production will help both users and administrators to better use and administer expensive HPC resources to meet computational needs. Mohammad Al-Tahat, Strahinja Trecakov, Jonathan E. Cook 0001 |
IPCCC | 3 |
| 2025 | Understanding Job Configuration Impacts on Multiple Views of HPC Job QualityabstractMost HPC resources that are available to a wide public userbase (for example XSEDE/ACCESS resources) have fairly basic user guides for building and running applications on the systems. Large performance differences, however, can occur if users are not experienced with detailed job configuration flags and their specific application behavior. This performance difference affects job quality-how well the job used its resources to accomplish the user's goals. However, a one-dimensional idea of job quality cannot capture all of the intricacies around how or why users might select a particular configuration. This paper presents four metrics for four different views of job quality, and then presents results from experiments over several applications that show the variability in quality that results from a variety of job configuration tuning parameters. This work can point the way towards guidelines for how users can potentially improve their own quality goals of their applications and thereby use the HPC resources most beneficially. One significant result we saw across applications is that high thread counts with low processes-per-node counts are generally not the best configurations; most applications do better with low thread counts and more processes per node. Strahinja Trecakov, Mohammad Al-Tahat, Jonathan E. Cook 0001 |
IPCCC | 3 |
| 2024 | Statistical and Shapelet Analysis of HPC Application Performance Using Time-Series Heartbeat DataabstractAppEKG is a high performance computing (HPC) oriented application heartbeat framework designed for low-overhead monitoring of HPC applications in production, providing a unified, understandable view of dynamic HPC application behavior by capturing time-varying behavior with little overhead (~1%). In this paper we demonstrate the value of heartbeat data collected with AppEKG, showing that it can be used for anomaly detection and run classification. This paper uses statistical modeling and time-series shapelet analyses as initial examples of heartbeat data analysis. The results of these analyses show that application heartbeats can be useful to help better understand how HPC applications are using the costly time and resources they consume, and that this can be applied in production. Mohammad Al-Tahat, Strahinja Trecakov, Jonathan E. Cook 0001 |
IPCCC | 3 |
| 2023 | Evaluation of the Performance Impact of SPM Allocation on a Novel Scratchpad MemoryabstractLocal Memory Store (LMStore) is a novel scratchpad memory (SPM) design, with recent research evaluation showing its capability for improving program performance. However, the performance of LMStore depends on its memory layout decided by its allocation scheme. In this paper, we evaluate the impact of SPM allocation on LMStore performance. Our experimental results, using benchmarks from the Malardalen WCET benchmark suite executing on LMStore architecture modeled in the PyCacheSim simulator, demonstrate that LMStore with a stack distance-based SPM allocation scheme significantly improves data movement by an average of 44.46% compared to a Cache-only architecture, and by an average of 23.89% compared to LMStore with a frequency-based SPM allocation scheme. Essa Imhmed, Edgar Eduardo Ceh-Varela, Jonathan E. Cook 0001, Caleb Parten |
COMPSAC | 3 |
| 2022 | IncProf: Efficient Source-Oriented Phase Identification for Application Behavior UnderstandingabstractLong running applications often have varying behaviors, here called phases. While considerable work in computer architecture has been done in identifying application phases based on how the hardware is being exercised, comparatively less work has been focused on identifying application phases based on regions of source code being executed. In this paper we introduce a new methodology and an efficient tool framework, IncProf, for observing and capturing the time-varying source execution behavior of applications, and for then deducing application phases from the resulting data. Uses of this capability include simply better understanding the varying behavior of long running applications, and for efficiently tracking deployed application performance in the future by providing information to identify good instrumentation points. Omar Aaziz, Mohammad Al-Tahat, Strahinja Trecakov, Jonathan E. Cook 0001 |
CLUSTER | 4 |
| 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File SystemsabstractLarge-scale parallel file systems (PFSs) play an essential role in high-performance computing (HPC). However, despite their importance, their reliability is much less studied or understood compared with that of local storage systems or cloud storage systems. Recent failure incidents at real HPC centers have exposed the latent defects in PFS clusters as well as the urgent need for a systematic analysis. To address the challenge, we perform a study of the failure recovery and logging mechanisms of PFSs in this article. First, to trigger the failure recovery and logging operations of the target PFS, we introduce a black-box fault injection tool called PFault , which is transparent to PFSs and easy to deploy in practice. PFault emulates the failure state of individual storage nodes in the PFS based on a set of pre-defined fault models and enables examining the PFS behavior under fault systematically. Next, we apply PFault to study two widely used PFSs: Lustre and BeeGFS. Our analysis reveals the unique failure recovery and logging patterns of the target PFSs and identifies multiple cases where the PFSs are imperfect in terms of failure handling. For example, Lustre includes a recovery component called LFSCK to detect and fix PFS-level inconsistencies, but we find that LFSCK itself may hang or trigger kernel panics when scanning a corrupted Lustre. Even after the recovery attempt of LFSCK, the subsequent workloads applied to Lustre may still behave abnormally (e.g., hang or report I/O errors). Similar issues have also been observed in BeeGFS and its recovery component BeeGFS-FSCK. We analyze the root causes of the abnormal symptoms observed in depth, which has led to a new patch set to be merged into the coming Lustre release. In addition, we characterize the extensive logs generated in the experiments in detail and identify the unique patterns and limitations of PFSs in terms of failure logging. We hope this study and the resulting tool and dataset can facilitate follow-up research in the communities and help improve PFSs for reliable high-performance computing. Runzhou Han, Om Rameshwar Gatla, Mai Zheng, Jinrui Cao, Di Zhang 0015, Dong Dai 0001, Yong Chen 0001, Jonathan E. Cook 0001 |
ACM Trans. Storage | 8 |
| 2018 | A Methodology for Characterizing the Correspondence Between Real and Proxy ApplicationsabstractProxy applications are a simplified means for stake-holders to evaluate how both hardware and software stacks might perform on the class of real applications that they are meant to model. However, characterizing the relationship between them and their behavior is not an easy task. We present a data-driven methodology for characterizing the relationship between real and proxy applications based on collecting runtime data from both and then using data analytics to find their correspondence and divergence. We use new capabilities for application-level monitoring within LDMS (Lightweight Distributed Monitoring System) to capture hardware performance counter and MPI-related data. To demonstrate the utility of this methodology, we present experimental evidence from two system platforms, using four proxy applications from the current ECP Proxy Application Suite and their corresponding parent applications (in the ECP application portfolio). Results show that each proxy analyzed is representative of its parent with respect to computation and memory behavior. We also analyze communication patterns separately using mpiP data and show that communication for these four proxy/parent pairs is also similar. Omar Aaziz, Jeanine E. Cook, Jonathan E. Cook 0001, Tanner Juedeman, David Richards, Courtenay T. Vaughan |
CLUSTER | 3 |
| 2018 | Modeling Expected Application Runtime for Characterizing and Assessing Job PerformanceabstractIn this paper, we present a methodology for modeling the expected runtime of a job based on historical application data and data from the job itself. This estimation model is useful for both for HPC users and administrators as a metric to compare the actual job runtime to, thus establishing a measure of performance of the job. We used job data, system data, and hardware performance counters in a near-zero overhead manner to model and assess job performance, in particular whether or not the job runtime was in line with expectations from historical application performance. We show over three proxy applications and three real applications that our estimations are within 5% of actual performance. Omar Aaziz, Jonathan E. Cook 0001, Mohammed Tanash |
CLUSTER | 2 |
| 2018 | 3D-PIM NoCs with Multiple Subnetworks: A Performance and Power EvaluationabstractThe advances in 3D circuit integration have reignited the idea of processing-in-memory (PIM). In this paper, we evaluate 3D mesh-based network on chip (NoC) for 3D-PIM systems with single and multiple network configurations. We study stacked mesh (S-Mesh), which is a mesh-bus hybrid architecture for 3D NoCs that connects vertically stacked 2D meshes through buses. Previous S-Mesh studies have not addressed the problems and modifications needed at the building blocks of the network. We explain in details the internal structure of the S-Mesh, as well as, the problems and possible solutions of connecting 2D meshes using vertical buses. Also, we evaluate the performance of 3D NoCs via two traffic patterns, one of which is a novel traffic pattern that better measures 3D-PIM systems performance. Finally, we use the Rodinia benchmarks to measure the performance under real workloads. We use DSENT to evaluate the power consumption. Our results show ~ 15% performance improvement for the S-Mesh under zero-load packet latency and ~ 11% lower average packet latency for the Rodinia benchmarks. Also, S-Mesh is the low power configuration with the router static power accountable for 90% of the total network power consumption. Abdel-Hameed A. Badawy, Jesus Gardea, Yuho Jin, Jonathan E. Cook 0001 |
IPCCC | 4 |
| 2017 | YAViT (Yet Another Viz Tool): Raising the Level of Abstraction in End-User HPC InteractionsabstractBecause data collection in HPC systems happens on the nodes and is easily related to the job running on the node, tools presenting the data and subsequent analyses to the user generally present them at the job level. Our position is that this is the wrong level of abstraction and thus limits the value of the analyses, often dissuading users from using any of the offered tools. In this paper we present the position that tools need to present analyses at the level users are interested in, which is their applications. Omar Aaziz, Ujjwal Panthi, Jonathan E. Cook 0001 |
CLUSTER | 3 |
| 2015 | Push Me Pull You: Integrating Opposing Data Transport Modes for Efficient HPC Application MonitoringabstractWhile HPC system monitoring is a necessary and accepted practice, applications are still basically opaque in the production environment. For better HPC platform management and utilization, especially as platforms push towards exascale size, HPC applications need to be more transparent in their execution in the production environment. PROMON is a framework for application monitoring in the production environment, but its design concentrated on the front end issues of offering easy to use application instrumentation. This paper presents the integration of PROMON with LDMS, a proven efficient HPC system monitoring framework. PROMON and LDMS offer a case study in integrating two disparate instrumentation and monitoring models, and the lessons are applicable to other HPC monitoring issues. Omar Aaziz, Jonathan E. Cook 0001, Hadi Sharifi |
CLUSTER | 2 |
| 2014 | Accurate statistical performance modeling and validation of out-of-order processors using Monte Carlo methodsabstractAlthough simulation is an indispensable tool in computer architecture research and development, there is a pressing need for new modeling techniques to improve simulation speeds while maintaining accuracy and robustness. It is no longer practical to use only cycle-accurate processor simulation (the dominant simulation method) for design space and performance studies due to its extremely slow speed. To address this and other problems of cycle-accurate simulation, we propose a fast and accurate statistical modeling methodology based on Monte Carlo methods to model the performance of modern out-of-order processors. Using these statistical models, simulation and performance prediction can be achieved in seconds regardless of the modeled application's size. This paper presents the proposed methodology and its first application to model a modern out-of-order execution processor. We present a statistical model for the Opteron (Magny-Cours) processor and validate it against real hardware. Using SPEC CPU2006 and Mantevo benchmarks, the model can predict performance in terms of cycles-per-instruction within 4.79% of actual on average. We also present a novel method for generating CPI stacks which are CPI representations that quantify the contribution of individual performance components to the total CPI. To further validate these CPI stacks, we use a detailed processor simulator, build a statistical model of the simulator architecture, validate the model against the simulator, and then proceed to validate the CPI stacks predicted by our statistical model. The average CPI prediction error is 5.6%, and the average difference between the predicted and measured CPI components is 1.3% with a maximum difference of 5.4%. Waleed Alkohlani, Jeanine E. Cook, Jonathan E. Cook 0001 |
IPCCC | 3 |
| 2012 | Compositional Verification of Sensor Software Using UppallabstractVerification of wireless sensor networks has long been performed for communication protocols and for network-level behavior over multiple nodes, but not for the basic properties that should hold at a single node. Testing sensor networks, however, is extremely hard due to the lack of controllability, and complex simulation setups are often too expensive to undertake. Thus, verification of properties for a sensor node is desirable. We created a verification methodology that extracts timed models of the high-level behavior of a wireless sensor and then uses UPPAAL to verify both functional and non-functional (timed) properties for the sensor. This verification capability will enhance the trustworthiness of deployed sensor networks. Mustafa Hammad, Jonathan E. Cook 0001 |
ISSRE | 2 |
| 2012 | Towards More Generic Aspect-Oriented Programming: Rethinking the AOP Joinpoint Concept
Jonathan E. Cook 0001, Amjad Nusayr |
SEKE | 1 |
| 2010 | A service-based runtime environment for native applicationsabstractAbstract Unlike interpreted application programs that run within the rich runtime environment of their own interpreters, natively compiled application programs are typically thought of as executing on the barebones runtime environment provided by their own operating systems and dynamic linkers. This article presents the notion of a service‐based runtime (SBRT) environment for natively compiled application programs, in which the current dynamic linker is only a core service and other extensive middleware‐type services are available to the running application program. It then describes a dynamic‐linker‐based implementation of such an environment, calleddlSBRT.dlSBRTimplements a unified interface that allows it to be extensible by means of extension dynamic services. Implemented examples of such services are presented and the performance ofdlSBRT, when such example services are deployed, is evaluated. Copyright © 2009 John Wiley & Sons, Ltd. Abdulmalik Al-Gahmi, Jonathan E. Cook 0001 |
Softw. Pract. Exp. | 2 |
| 2009 | Lightweight Deployable Software Monitoring for Sensor NetworksabstractMost efforts at analyzing the behavior and performance of sensor network software have been simulation-based, but deployment in the field can bring about conditions that were unexpected, or just too hard to simulate. Thus monitoring the software in the field would be a valuable capability for engineering of sensor networks. We have developed a lightweight instrumentation tool for TinyOS, called LlMOW, that is deploy- able onto sensor hardware, is low cost in overhead, and is configurable and controllable in terms of how much it monitors and how much overhead is tolerable. We show both behavioral and timing results from our tool, and compare results with other monitoring approaches. Mustafa Hammad, Jonathan E. Cook 0001 |
ICCCN | 2 |
| 2009 | Extending AOP to Support Broad Runtime Monitoring Needs
Amjad Nusayr, Jonathan E. Cook 0001 |
SEKE | 2 |
| 2005 | MonDe: safe updating through monitored deployment of new component versionsabstractSafely updating software at remote sites is a cautious balance of enabling new functionality and avoiding adverse effects on existing functionality. A useful first step in this process would be to evaluate the performance of a new version of a component on the current workload before enabling its functionality. This step would let the engineers assess the component's performance over more (and more realistic) data points than by simply performing regression testing in-house.In this paper we propose to evaluate the performance of a new version of a component by (1) deploying it to remote sites, (2) running it in a controlled environment with the actual workloads being generated at that site, and (3) reporting the results back to the development engineers. Running the new version can either be done on-line, alongside the current system, or offline, using capture-replay techniques. By running at the remote site and reporting concise results, issues of data security, protection, and confidentiality are diminished, yet the new version can be evaluated on real workloads. Jonathan E. Cook 0001, Alessandro Orso |
PASTE | 1 |
| 2005 | Discovering thread interactions in a concurrent system
Jonathan E. Cook 0001, Zhidian Du |
J. Syst. Softw. | 1 |
| 2003 | ICSE Workshop on Dynamic Analysis (WODA 2003)abstractDynamic analysis of software systems has long proven to be a practical approach to gain understanding of the operational behavior of the system. This workshop will bring together researchers in the field of dynamic analysis to discuss the breadth of the field, order the field along logical dimensions, expose common issues and approaches, and stimulate synergistic collaborations among the participants. Jonathan E. Cook 0001, Michael D. Ernst |
ICSE | 1 |
| 2001 | Measuring Behavioral Correspondence to a Timed Concurrent ModelabstractResearch in formal methods has produced fruitful techniques that can verify global properties of a design of a real-time system, or exact behavioral correspondence to the design. Exactness is often not achieved, however, and yet understanding how close the design and system correspond still would be very valuable to direct further efforts in achieving exactness or to modify the design where the system simply cannot achieve the requirements. The paper describes a method and tool that fills this niche, by quantitatively measuring how closely the behavior of a real-time system corresponds to its specification, given in a timed, concurrent model. Jonathan E. Cook 0001, Cha He, Changjun Ma |
ICSM | 1 |
| 1999 | Highly Reliable Upgrading of ComponentsabstractArticle Free Access Share on Highly reliable upgrading of components Authors: Jonathan E. Cook Department of Computer Science, New Mexico State University, Las Cruces, NM Department of Computer Science, New Mexico State University, Las Cruces, NMView Profile , Jeffrey A. Dage Department of Computer Science, New Mexico State University, Las Cruces, NM Department of Computer Science, New Mexico State University, Las Cruces, NMView Profile Authors Info & Claims ICSE '99: Proceedings of the 21st international conference on Software engineeringMay 1999 Pages 203–212https://doi.org/10.1145/302405.302466Published:16 May 1999Publication History 75citation678DownloadsMetricsTotal Citations75Total Downloads678Last 12 Months14Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Jonathan E. Cook 0001, Jeffrey A. Dage |
ICSE | 1 |
| 1999 | Software Process Validation: Quantitatively Measuring the Correspondence of a Process to a ModelabstractTo a great extent, the usefulness of a formal model of a software process lies in its ability to accurately predict the behavior of the executing process. Similarly, the usefulness of an executing process lies largely in its ability to fulfill the requirements embodied in a formal model of the process. When process models and process executions diverge, something significant is happening. We have developed techniques for uncovering and measuring the discrepancies between models and executions, which we call process validation . Process validation takes a process execution and a process model, and measures the level of correspondence between the two. Our metrics are tailorable and give process engineers control over determining the severity of different types of discrepancies. The techniques provide detailed information once a high-level measurement indicates the presence of a problem. We have applied our processes validation methods in an industrial case study, of which a portion is described in this article. Jonathan E. Cook 0001, Alexander L. Wolf |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 1998 | Event-Base Detection of ConcurrencyabstractUnderstanding the behavior of a system is crucial in being able to modify, maintain, and improve the system. A particularly difficult aspect of some system behaviors is concurrency. While there are many techniques to specify intended concurrent behavior, there are few, if any, techniques to capture and model actual concurrent behavior. This paper presents a technique to discover patterns of concurrent behavior from traces of system events. The technique is based on a probabilistic analysis of the event traces. Using metrics for the number, frequency, and regularity of event occurrences, a determination is made of the likely concurrent behavior being manifested by the system. The technique is useful in a wide variety of software engineering tasks, including architecture discovery, reengineering, user interaction modeling, and software process improvement. Jonathan E. Cook 0001, Alexander L. Wolf |
SIGSOFT FSE | 1 |
| 1998 | A Highly Effective Partition Selection Policy for Object Database Garbage CollectionabstractWe investigate methods to improve the performance of algorithms for automatic storage reclamation of object databases. These algorithms are based on a technique called partitioned garbage collection, in which a subset of the entire database is collected independently of the rest. We evaluate how different application, database system, and garbage collection implementation parameters affect the performance of garbage collection in object database systems. We focus specifically on investigating the policy that is used to select which partition in the database should be collected. Three of the policies that we investigate are based on the intuition that the values of overwritten pointers provide good hints about where to find garbage. A fourth policy investigated chooses the partition with the greatest presence in the I/O buffer. Using simulations based on a synthetic database, we show that one of our policies requires less I/O to collect more garbage than any existing implementable policy. Furthermore, that policy performs close to a locally optimal policy over a wide range of simulation parameters, including database size, collection rate, and database connectivity. We also show what impact these simulation parameters have on application performance and investigate the expected costs and benefits of garbage collection in such systems. Jonathan E. Cook 0001, Alexander L. Wolf, Benjamin G. Zorn |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1998 | Discovering Models of Software Processes from Event-Based DataabstractMany software process methods and tools presuppose the existence of a formal model of a process. Unfortunately, developing a formal model for an on-going, complex process can be difficult, costly, and error prone. This presents a practical barrier to the adoption of process technologies, which would be lowered by automated assistance in creating formal models. To this end, we have developed a data analysis technique that we term process discovery. Under this technique, data describing process events are first captured from an on-going process and then used to generate a formal model of the behavior of that process. In this article we describe a Markov method that we developed specifically for process discovery, as well as describe two additional methods that we adopted from other domains and augmented for our purposes. The three methods range from the purely algorithmic to the purely statistical. We compare the methods and discuss their application in an industrial case study. Jonathan E. Cook 0001, Alexander L. Wolf |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 1998 | Cost-Effective Analysis of In-Place Software ProcessesabstractProcess studies and improvement efforts typically call for new instrumentation on the process in order to collect the data they have deemed necessary. This can be intrusive and expensive, and resistance to the extra workload often foils the study before it begins. The result is neither interesting new knowledge nor an improved process. In many organizations, however, extensive historical process and product data already exist. Can these existing data be used to empirically explore what process factors might be affecting the outcome of the process? If they can, organizations would have a cost-effective method for quantitatively, if not causally, understanding their process and its relationship to the product. We present a case study that analyzes an in-place industrial process and takes advantage of existing data sources. In doing this, we also illustrate and propose a methodology for such exploratory empirical studies. The case study makes use of several readily-available repositories of process data in the industrial organization. Our results show that readily available data can be used to correlate both simple aggregate metrics and complex process metrics with defects in the product. Through the case study, we give evidence supporting the claim that exploratory empirical studies can provide significant results and benefits while being cost-effective in their demands on the organization. Jonathan E. Cook 0001, Lawrence G. Votta, Alexander L. Wolf |
IEEE Trans. Software Eng. | 1 |
| 1996 | Semi-automatic, Self-adaptive Control of Garbage Collection Rates in Object DatabasesabstractA fundamental problem in automating object database storage reclamation is determining how often to perform garbage collection. We show that the choice of collection rate can have a significant impact on application performance and that the "best" rate depends on the dynamic behavior of the application, tempered by the particular performance goals of the user. We describe two semi-automatic, selfadaptive policies for controlling collection rate that we have developed to address the problem. Using tracedriven simulations, we evaluate the performance of the policies on a test database application that demonstrates two distinct reclustering behaviors. Our results show that the policies are effective at achieving user-specified levels of I/O operations and database garbage percentage. We also investigate the sensitivity of the policies over a range of object connectivities. The evaluation demonstrates that semi-automatic, self-adaptive policies are a practical means for flexibly controllin... Jonathan E. Cook 0001, Artur Klauser, Alexander L. Wolf, Benjamin G. Zorn |
SIGMOD Conference | 1 |
| 1995 | Automating Process Discovery Through Event-Data AnalysisabstractMany software process methods and tools presuppose the existence of a formal model of a process.Unfortunately, developing a formal model for an on-going, complex process can be dificult, costly, and err-or prone.This presents a practical barrier to the adoption of process technologies.The barrier would be lowered by automatmg the creation of formal models.We are currently exploring techniques that can use basic event data captured from an on-going process to generate a formal model of process behavior.We term this kind of data analysis process discovery.Thts paper descr~bes and illustrates three methods with whzch we have been experimenting: algorithmic grammar inference, Markov models, and neural networks. Jonathan E. Cook 0001, Alexander L. Wolf |
ICSE | 1 |
| 1994 | Partition Selection Policies in Object Database Garbage CollectionabstractThe automatic reclamation of storage for unreferenced objects is very important in object databases. Existing language system algorithms for automatic storage reclamation have been shown to be inappropriate. In this paper, we investigate methods to improve the performance of algorithms for automatic for automatic storage reclamation of object databases. These algorithms are based on a technique called partitioned garbage collection, in which a subset of the entire database is collected independently of the rest. Specifically, we investigate the policy that is used to select what partition in the database should be collected. The policies that we propose and investigate are based on the intuition that the values of overwritten pointers provide good hints about where to find garbage. Using trace-driven simulation, we show that one of our policies requires less I/O to collect more garbage than any existing implementable policy and performs close to a near-optimal policy over a wide range of database sizes and object connectivities. Jonathan E. Cook 0001, Alexander L. Wolf, Benjamin G. Zorn |
SIGMOD Conference | 1 |