VLDB 2026 Research / reviewers in the wild / expert
Madeline Diep
dblp:77/3716
· DBLP profile ↗
12ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0002-9908-0367ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
5 papers |
Program analysis · 41% Program verification · 29% Software testing · 17% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program verification › dynamic verification
runtime verification |
0.2 | 2 | 2011 | Lattice-Based Sampling for Path Property Monitoring · ACM Trans. Softw. Eng. Methodol. 2011 Reducing the Cost of Path Property Monitoring Through Sampling · ASE 2008 |
Program analysis › dynamic analysis
runtime monitoring |
0.1 | 1 | 2008 | Reducing the Cost of Path Property Monitoring Through Sampling · ASE 2008 |
Empirical software engineering › software analytics
deployed software analysis |
0.1 | 1 | 2007 | Analysis of a deployed software · ESEC/SIGSOFT FSE 2007 |
Program analysis
dynamic analysis |
0.1 | 1 | 2007 | Reducing irrelevant trace variations · ASE 2007 |
Program analysis › dynamic analysis
instrumentation |
0.1 | 1 | 2007 | Analysis of a deployed software · ESEC/SIGSOFT FSE 2007 |
Program analysis › dynamic analysis
trace analysis |
0.1 | 1 | 2007 | Reducing irrelevant trace variations · ASE 2007 |
Software testing › regression testing
test suite augmentation |
0.1 | 1 | 2005 | Profiling Deployed Software: Assessing Strategies and Testing Opportunities · IEEE Trans. Software Eng. 2005 |
Debugging and program repair
fault localization |
0.0 | 1 | 2007 | Reducing irrelevant trace variations · ASE 2007 |
Software testing
test suite evaluation |
0.0 | 1 | 2005 | Profiling Deployed Software: Assessing Strategies and Testing Opportunities · IEEE Trans. Software Eng. 2005 |
Methods — techniques the papers use, named apart from their topics
property composition · 0.2lattice-based sampling · 0.2weighting scheme · 0.1trace segmentation · 0.1dynamic analysis · 0.1user session analysis · 0.1software profiling · 0.1remote data collection · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | SeqScreen: a biocuration platform for robust taxonomic and biological process characterization of nucleic acid sequences of interestabstractRapid advancements in synthetic biology and nucleic acid synthesis, in particular concerns about its intentional or accidental misuse, call for more sophisticated screening tools to identify genes of interest within short sequence fragments. One major gap in predicting genes of concern is the inadequacy of current tools and ontologies to describe the specific biological processes of pathogenic proteins. The objective of this work is to design software that sensitively assigns taxonomic classifications, functional annotations, and biological processes of interest to short nucleotide sequences of unknown origin (50bp-1,000bp). The overarching goal is to perform sensitive characterization of short sequences and highlight specific pathogenic biological processes of interest (BPoIs). The SeqScreen software executes these tasks in analytical workflows with Nextflow and outputs results in a tab-delimited report. Local and global alignments differentiate hits to taxonomically-related sequences from similar but unrelated sequences, and an ensemble approach leverages multiple tools and databases to assign a variety of functional terms to each query sequence. Final biological process assessments are made from the predicted functional annotations, which leverage information in pre-existing databases, as well as new custom biocurations. Machine learning models predict each biological process of interest on large protein databases before incorporation into the SeqScreen framework to streamline computational efficiency, ensure reproducible results, allow for version control, and facilitate the review of the automated predictions by expert biocurators. The SeqScreen source code is available at https://gitlab.com/treangenlab/seqscreen. Dreycey Albin, Pravin Muthu, Gene Godbold, Mikael Lindvall, Madeline Diep, Adam A. Porter, Mihai Pop, Krista Ternus, Todd J. Treangen, Dan Nasko, Ryan A. Leo Elworth, Jacob Lu, Advait Balaji, Christian Diaz, Nidhi Shah, Jeremy D. Selengut, Chris Hulme-Lowe |
BIBM | 5 |
| 2017 | Safety-Focused Security Requirements Elicitation for Medical Device SoftwareabstractSecurity attacks on medical devices have been shown to have potential safety concerns. Because of this, stakeholders (device makers, regulators, users, etc.) have increasing interest in enhancing security in medical devices. An effective means to approach this objective is to integrate systematic security requirements elicitation and analysis into the design and evaluation of medical device software. This paper extends the sequence-based enumeration approach, a systematic approach for defining the behavior of embedded software, to analyze the requirement documents of a medical device for the purpose of eliciting security requirements. As a proof of concept, we apply our approach on a concrete case study, which shows that the extended approach is useful for identifying sequences of medical device events that might be harmful to the patient, for example because the events are initiated by an active adversary trying to use the device in a malicious way. We then show how security requirements may be formulated based on the identified threats. By exploring these sequences systematically, the developers can reliably assess what, where, and how the security threats may manifest in their system, what the safety implications are, and finally they can evaluate the resulting requirements and mitigations. Mikael Lindvall, Madeline Diep, Michele Klein, Paul L. Jones, Yi Zhang 0051, Eugene Y. Vasserman |
RE | 2 |
| 2013 | Debugging Revisited: Toward Understanding the Debugging Needs of Contemporary Software DevelopersabstractWe know surprisingly little about how professional developers define debugging and the challenges they face in industrial environments. To begin exploring professional debugging challenges and needs, we conducted and analyzed interviews with 15 professional software engineers at Microsoft. The goals of this study are: 1) to understand how professional developers currently use information and tools to debug, 2) to identify new challenges in debugging in contemporary software development domains (web services, multithreaded/multicore programming), and 3) to identify the improvements in debugging support desired by these professionals that are needed from research. The interviews were coded to identify the most common information resources, techniques, challenges, and needs for debugging as articulated by the developers. The study reveals several debugging challenges faced by professionals, including: 1) the interaction of hypothesis instrumentation and software environment as a source of debugging difficulty, 2) the impact of log file information on accurate debugging of web services, and 3) the mismatch between the sequential human thought process and the non-sequential execution of multithreaded environments as source of difficulty. The interviewees also describe desired improvements to tools to support debugging, many of which have been discussed in research but not transitioned to practice. Lucas Layman, Madeline Diep, Meiyappan Nagappan, Janice Singer, Robert DeLine, Gina Venolia |
ESEM | 2 |
| 2012 | Analyzing inspection data for heuristic effectivenessabstractA significant body of knowledge concerning software inspection practice indicates that the value of inspections varies widely both within and across organizations. Inspection effectiveness and efficiency may be affected by a variety of factors such as inspection planning, the type of software, the developing organization, and many others. In the early 1990's, a governmental organization developing complex and highly critical software systems formulated heuristics for inspection planning based on best practices and their early inspection data. Since the development context at the organization has changed in some ways since the heuristics were proposed, it is important to assess whether the heuristics are still a suitable guideline to use. To investigate this question, we statistically evaluated the differences in effectiveness and efficiency between inspections that adhered to the heuristics and ones that did not. Our analysis revealed no significant difference in effectiveness or efficiency for most heuristics. We also learned that compliance with the heuristics is diminishing over time. Forrest Shull, Carolyn B. Seaman, Madeline Diep |
ESEM | 3 |
| 2011 | Lattice-Based Sampling for Path Property MonitoringabstractRuntime monitoring can provide important insights about a program’s behavior and, for simple properties, it can be done efficiently. Monitoring properties describing sequences of program states and events, however, can result in significant runtime overhead. This is particularly critical when monitoring programs deployed at user sites that have low tolerance for overhead. In this paper we present a novel approach to reducing the cost of runtime monitoring of path properties. A set of original properties are composed to form a single integrated property that is then systematically decomposed into a set of properties that encode necessary conditions for property violations. The resulting set of properties forms a lattice whose structure is exploited to select a sample of properties that can lower monitoring cost, while preserving violation detection power relative to the original properties. The lattice is then complemented with a weighting scheme that assigns each property a different priority that can be adjusted continuously to better drive the property sampling process. Our evaluation using the Hibernate API reveals that our approach produces a rich, structured set of properties that enables control of monitoring overhead, while detecting more violations more quickly than alternative techniques. Madeline Diep, Matthew B. Dwyer, Sebastian G. Elbaum |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2010 | Obtaining valid safety data for software safety measurement and process improvementabstractArt. 46 Victor R. Basili, Marvin V. Zelkowitz, Lucas Layman, Kathleen Coleman Dangle, Madeline Diep |
ESEM | 5 |
| 2008 | Trace NormalizationabstractIdentifying truly distinct traces is crucial for the performance of many dynamic analysis activities. For example, given a set of traces associated with a program failure, identifying a subset of unique traces can reduce the debugging effort by producing a smaller set of candidate fault locations. The process of identifying unique traces, however, is subject to the presence of irrelevant variations in the sequence of trace events, which can make a trace appear unique when it is not. In this paper we present an approach to reduce inconsequential and potentially detrimental trace variations. The approach decomposes traces into segments on which irrelevant variations caused by event ordering or repetition can be identified, and then used to normalize the traces in the pool. The approach is investigated on two well-known client dynamic analyses by replicating the conditions under which they were originally assessed, revealing that the clients can deliver more precise results with the normalized traces. Madeline Diep, Sebastian G. Elbaum, Matthew B. Dwyer |
ISSRE | 1 |
| 2008 | Reducing the Cost of Path Property Monitoring Through SamplingabstractRun-time monitoring can provide important insights about a program's behavior and, for simple properties, it can be done efficiently. Monitoring properties describing sequences of program states and events, however, can result in significant run-time overhead. In this paper we present a novel approach to reducing the cost of run-time monitoring of path properties. Properties are composed to form a single integrated property that is then systematically decomposed into a set of properties that encode necessary conditions for property violations. The resulting set of properties forms a lattice whose structure is exploited to select a sample of properties that can lower monitoring cost, while preserving violation detection power relative to the original properties. Preliminary studies for a widely used Java API reveal that our approach produces a rich, structured set of properties that enables control of monitoring overhead, while detecting more violations than alternative techniques. Matthew B. Dwyer, Madeline Diep, Sebastian G. Elbaum |
ASE | 2 |
| 2007 | Reducing irrelevant trace variationsabstractIdentifying truly distinct traces is crucial for the performance and practicality of many dynamic analysis activities. For example, given a trace pool resulting from program failures, identifying the set of distinct traces can reduce the debugging effort by more quickly producing a smaller set of candidate fault locations. The process of discriminating valuable traces, however, is subject to the presence of irrelevant variations in the trace constitution, i.e., the sequence of events in a trace, that can make a trace appear unique when it is not, leading to the retention of a trace that adds no value. In this paper we present an approach to address inconsequential and potentially detrimental trace variations. The approach decomposes traces into segments on which irrelevant variations caused by event ordering or repetition can be detected and removed. The approach is illustrated on two well-known client dynamic analyses and is supported by an infrastructure to explore the approach Madeline Diep, Sebastian G. Elbaum, Matthew B. Dwyer |
ASE | 1 |
| 2007 | Analysis of a deployed softwareabstractAnalyzing a deployed software provides a means to characterize and leverage the software's runtime behavior as it is employed by its intended users. Preliminary studies have shown that leveraging the information obtained from the field provides engineers an opportunity to improve their software testing activities. The analysis of a deployed software can be performed in three stages: (1) the analysis to determine, before the software is deployed, where the instrumentation probes should be inserted into the software and what information that they should capture, (2) the analysis to determine when the field data should be sent back to the company during deployment, and (3) the analysis to leverage the field information after deployment. To make the analysis activities more feasible, we need to take into consideration that there are distinct characteristic differences between the development and the deployed environment. Deployed environment allows for less overhead,provides less control for the engineers, and requires highlyscalable techniques due to the high volume of information. Hence, the existing approaches for in-house analysis may become ineffective, inefficient, or even useless when they are directly applied to the deployed environment. Existing approaches for analyzing deployed software also need to be more aware that a technique in one analysis stage may affect the performance of a technique in other analysis stage. This research proposal details the challenges that arise when analyzing a deployed software and seeks to develop a set techniques to address these challenges that can be applied to each stage or across the analysis stages. Madeline Diep |
ESEC/SIGSOFT FSE | 1 |
| 2006 | Probe Distribution Techniques to Profile Events in Deployed SoftwareabstractProfiling deployed software provides valuable insights for quality improvement activities. The probes required for profiling, however, can cause an unacceptable performance overhead for users. In previous work we have shown that significant overhead reduction can be achieved, with limited information loss, through the distribution of probes across deployed instances. However, existing techniques for probe distribution are designed to profile simple events. In this paper we present a set of techniques for probe distribution to enable the profiling of complex events that require multiple probes and that share probes with other events. Our evaluation of the new techniques indicates that, for tight overhead bounds, techniques that produce a balanced event allocation can retain significantly more field information Madeline Diep, Myra B. Cohen, Sebastian G. Elbaum |
ISSRE | 1 |
| 2005 | Profiling Deployed Software: Assessing Strategies and Testing OpportunitiesabstractAn understanding of how software is employed in the field can yield many opportunities for quality improvements. Profiling released software can provide such an understanding. However, profiling released software is difficult due to the potentially large number of deployed sites that must be profiled, the transparency requirements at a user's site, and the remote data collection and deployment management process. Researchers have recently proposed various approaches to tap into the opportunities offered by profiling deployed systems and overcome those challenges. Initial studies have illustrated the application of these approaches and have shown their feasibility. Still, the proposed approaches, and the tradeoffs between overhead, accuracy, and potential benefits for the testing activity have been barely quantified. This paper aims to overcome those limitations. Our analysis of 1,200 user sessions on a 155 KLOC deployed system substantiates the ability of field data to support test suite improvements, assesses the efficiency of profiling techniques for released software, and the effectiveness of testing efforts that leverage profiled field data. Sebastian G. Elbaum, Madeline Diep |
IEEE Trans. Software Eng. | 2 |