VLDB 2026 Research / reviewers in the wild / expert
Harvey P. Siy
dblp:77/1325
· DBLP profile ↗
28ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-4482-5712ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 26 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Reflection on "Advances in Software Inspections"abstractMichael Fagan’s work on the software inspection process defined a practical approach for improving software quality that was broadly adopted throughout the software industry. This work sparked numerous follow-on efforts aimed at understanding and improving the approach. Finally, the process has flexibly evolved over roughly 50 years, incorporating new languages, tools, techniques, and new understandings of its goals and benefits. Adam A. Porter, Harvey P. Siy, Lawrence G. Votta |
IEEE Trans. Software Eng. | 2 |
| 2023 | Exploring the Potential of Frama-C in IoT Static AnalysisabstractIn this research, we investigated the feasibility of using static analysis for IoT applications with Frama-C. We looked at different kinds of possible IoT vulnerabilities and how static analysis specifically could be used to identify them. With certain Frama-C plugins such as Eva, we were able to run static analysis on most IoT code without modifying the code itself and catch errors that could potentially be exploited in real-world applications that would have otherwise been missed. Additionally, we created a simple IoT device, by utilizing Raspberry Pi 4 hardware with a set of different SunFounder sensors, and ran our created code for it through Frama-C to find any errors. The static analysis done gave a significant amount of potential vulnerabilities in our code, mostly consisting of integer overflows. We learned how we could use static analysis tools, like Frama-C, as a powerful way to find potential vulnerabilities with minimal changes to code. Minh Le Kim Tran, William King, Harvey P. Siy |
MobiHoc | 3 |
| 2023 | A Statistical Method for API Usage Learning and API Misuse Violation FindingabstractA large corpus of software repositories enables an opportunity for using machine learning (ML) approaches to create new software engineering tools. In this paper, we propose a novel technique which leverages ML approaches for automating software engineering tasks and thus improves software quality. Our concrete goal is to (1) explore the abundance of predictable repetitive regularities of such a massive codebase, (2) develop an ML approach for training a statistical model to identify common patterns in software corpora, and then (3) use these patterns to statistically detect anomalous, likely buggy, program behavior that significantly deviates from these typical patterns. These internal regularities and repetitive properties of software can be captured as patterns to detect violations of these common patterns. Such violations have a critical impact on program behavior such as bugs, security vulnerabilities, or even program crashes. Our approach focuses on usage patterns of application programming interfaces (APIs). API usage patterns are commonly recurring, representative examples of how real-world applications use APIs in software corpora. These desirable patterns of API usage are learnable to validate or improve developers' implementations. This paper shows preliminary results that we use standard cross-entropy and perplexity to measure how surprising a test subject application is to a statistical model estimated from a software corpus. We continue to develop our approach and evaluate the effectiveness to focus on the following research questions. Are our ML models effectively trainable on large code corpora to learn desirable API usage patterns? How does the performance of our ML-based approach compare to state-of-the-art language models for software when learning API usage for detecting API misuse violations? Deepak Panda, Piyush Basia, Kushal Nallavolu, Xin Zhong 0001, Harvey P. Siy, Myoungkyu Song |
SERA | 5 |
| 2023 | SSDTutor: A feedback-driven intelligent tutoring system for secure software development
Dip Kiran Pradhan Newar, Rui Zhao 0005, Harvey P. Siy, Leen-Kiat Soh, Myoungkyu Song |
Sci. Comput. Program. | 3 |
| 2022 | A Feasibility Study of Using Code Clone Detection for Secure Programming EducationabstractSecure library reuse is critical for modern ap-plications to protect private information in software security engineering. Teaching secure programming is also more critical to tackle the challenges of new and evolving threats. However, novice students often make mistakes by API misuses due to a lack of understanding of secure libraries or a false sense of security. In this paper, we study the feasibility of applying code clone detection (CCD) for finding relevant examples to effectively teach secure programming to computer science students. CCD is an emerging new technology that extracts syntactically or semantically similar code fragments to support many software engineering tasks, such as program understanding, code quality analysis, software evolution analysis, and bug detection. We have developed a prototype implementation ExTUTOR that allows students to search for relevant examples as feedback when they want to fix their programming issues or vulnerabilities. In our evaluation, we applied ExTUTOR to open source subject applications in the security domain. Our approach should help novice students gain benefits from feedback and identify how to effectively make use of APIs, encouraging students to fix their own security violations in their own applications. Michael Menard, Tommy Nelson, Milan Shahi, Hugh Morton, Adam DeTavernier, Harvey P. Siy, Rui Zhao 0005, Myoungkyu Song |
COMPSAC | 6 |
| 2022 | An Intelligent Tutoring System for API Misuse Correction by Instant Quality FeedbackabstractComputer science students have difficulty understanding correct usages of an Application Programming Interface (API) and programming violations that cause compilation or runtime errors. Despite high-quality documentation for programming, the students typically need an instructor's feedback when their programs cause bugs, crashes, and vulnerabilities. This paper presents a pedagogical approach that is based on an Intelligent Tutoring System called INTTuToR. Briefly, INTTUTORprovides novice students with instant feedback to fix their programming issues or vulnerabilities. We have implemented our approach as a plug-in application in the Integrated Development Environment (IDE) for an interactive educational environment. In our proposed evaluation, we plan to perform empirical studies with CS students to assess how effectively INTTUTORimproves their ability to identify and fix potential bugs or vulnerabilities in the cryptography-related programming assignments. Rui Zhao 0005, Harvey P. Siy, Chulwoo Pack, Leen-Kiat Soh, Myoungkyu Song |
COMPSAC | 2 |
| 2021 | FireBugs: Finding and Repairing Cryptography API Misuses in Mobile ApplicationsabstractIn this paper, we present FireBugs for Finding and Repairing Bugs based on security patterns. For the common misuse patterns of cryptography APIs (crypto APIs), we encode common cryptography rules into the pattern representations for bug detection and program repair regarding cryptography rule violations. In the evaluation, we conducted a case study to assess the bug detection capability by applying FireBugs to datasets mined from both open source and commercial projects. Also, we conducted a user study with professional software engineers at Mutual of Omaha Insurance Company to estimate the program repair capability. This evaluation showed that FireBugs can help professional engineers develop various cryptographic requirements in a resilient application. Larry Singleton, Rui Zhao 0005, Harvey P. Siy, Myoungkyu Song |
COMPSAC | 3 |
| 2020 | Modular norm models: practical representation and analysis of contractual rights and obligationsabstractCompliance analysis requires legal counsel but is generally unavailable in many software projects. Analysis of legal text using logic-based models can help developers understand requirements for the development and use of software-intensive systems throughout its lifecycle. We outline a practical modeling process for norms in legally binding agreements that include contractual rights and obligations. A computational norm model analyzes available rights and required duties based on the satisfiability of situations, a state of affairs, in a given scenario. Our method enables modular norm model extraction, representation, and reasoning. For norm extraction, using the theory of frame semantics, we construct two foundational norm templates for linguistic guidance. These templates correspond to Hohfeld’s concepts of claim-right and its jural correlative, duty. Each template instantiation results in a norm model, encapsulated in a modular unit which we call a super-situation that corresponds to an atomic fragment of law. For hierarchical modularity, super-situations contain a primary norm that participates in relationships with other norm models. Norm compliance values are logically derived from its related situations and propagated to the norm’s containing super-situation, which in turn participates in other super-situations. This modularity allows on-demand incremental modeling and reasoning using simpler model primitives than previous approaches. While we demonstrate the usefulness of our norm models through empirical studies with contractual statements in open source software and privacy domains, its grounding in theories of law and linguistics allows wide applicability. Sayonnha Mandal, Robin A. Gandhi, Harvey P. Siy |
Requir. Eng. | 3 |
| 2020 | Correction to: Modular norm models: practical representation and analysis of contractual rights and obligationsabstractThe article “Modular norm models: practical representation and analysis of contractual rights and obligations. Sayonnha Mandal, Robin A. Gandhi, Harvey P. Siy |
Requir. Eng. | 3 |
| 2014 | Architectural reliability analysis of framework-intensive applications: A web service case study
M. Rahmani, Azad H. Azadmanesh, Harvey P. Siy |
J. Syst. Softw. | 3 |
| 2013 | Test-driven learning in high school computer scienceabstractTest-driven development (TDD) is an accepted practice in the software development industry. Although computer science teaching programs have been slower to adopt test-driven practices, test-driven learning has been used in a number of universities with generally positive results. The use of test-driven learning at the high school level is less studied. We introduce and assess the benefits of using test-driven learning in a high school Advanced Placement (AP) computer science course. This course is a strong candidate for the introduction of TDD. The Java language used in AP computer science is well-supported by TDD tools, and the concepts of TDD show promise in helping students develop the ability to analyze problem statements and develop programs. Preliminary results indicate that students respond well to the use of TDD tools to complement other teaching techniques in AP CS. Ryan Stejskal, Harvey P. Siy |
CSEE&T | 2 |
| 2012 | Petri Net Modeling of Application Server Performance for Web Services
M. Rahmani, Azad H. Azadmanesh, Harvey P. Siy |
SEKE | 3 |
| 2011 | Empirical results on the study of software vulnerabilitiesabstractWhile the software development community has put a significant effort to capture the artifacts related to a discovered vulnerability in organized repositories, much of this information is not amenable to meaningful analysis and requires a deep and manual inspection. In the software assurance community a body of knowledge that provides an enumeration of common weaknesses has been developed, but it is not readily usable for the study of vulnerabilities in specific projects and user environments. We propose organizing the information in project repositories around semantic templates. In this paper, we present preliminary results of an experiment conducted to evaluate the effectiveness of using semantic templates as an aid to studying software vulnerabilities. Yan Wu 0018, Harvey P. Siy, Robin A. Gandhi |
ICSE | 2 |
| 2011 | Measuring disruption from software evolution activities using graph-based metricsabstractIn this paper, we investigate how class relationships are disrupted after large scale changes. We use graphs to represent different software versions and study changes to graph properties. We explore different combinatorial metrics to measure the extent of disruption after perfective maintenance activities. Our early results, on JHotDraw, demonstrate that combinatorial metrics can provide a good indicator to the degree to which relationships are disrupted or preserved across different versions. Prashant Paymal, Rajvardhan Patil, Sanjukta Bhowmick, Harvey P. Siy |
ICSM | 4 |
| 2009 | Assessing the impact of framework changes using component rankingabstractMost of today's software applications are built on top of libraries or frameworks. Just as applications evolve, libraries and frameworks also evolve. Upgrading is straightforward when the framework changes preserve the API and behavior of the offered services. However, in most cases, major changes are introduced with the new framework release, which can have a significant impact on the application. Hence, a common question a framework user might ask is, ldquoIs it worth upgrading to the new framework version?rdquo In this paper, we study the evolution of an application and its underlying framework to understand the information we can get through a multi-version use relation analysis. We use component rank changes to measure this impact. Component rank measurement is a way of quantifying the importance of a component by its usage. As framework components are used by applications, the rankings of the components are changed. We use component ranking to identify the core components in each framework version. We also confirm that upgrading to the new framework version has an impact to a component rank of the entire system and the framework, and this impact not only involves components which use the framework directly, but also other indirectly-related components. Finally, we also confirm that there is a difference in the growth of use relations between application and framework. Reishi Yokomori, Harvey P. Siy, Masami Noro, Katsuro Inoue |
ICSM | 2 |
| 2008 | Summarizing developer work history using time series segmentation: challenge reportabstractTemporal segmentation partitions time series data with the intent of producing more homogeneous segments. It is a technique used to preprocess data so that subsequent time series analysis on individual segments can detect trends that may not be evident when performing time series analysis on the entire dataset. Harvey P. Siy, Parvathi Chundi, Mahadevan Subramaniam |
MSR | 1 |
| 2008 | Discovering Meaningful Clusters from Mining the Software Engineering Literature
Yan Wu 0018, Harvey P. Siy |
SEKE | 2 |
| 2008 | A segmentation-based approach for temporal analysis of software version repositoriesabstractAbstract Time series segmentation is a promising approach to discover temporal patterns from time‐stamped numeric data. A novel approach to apply time series segmentation to discern temporal information from software version repositories is proposed. Data from such repositories, both numeric and non‐numeric, are represented as item‐set time series data. A dynamic programming algorithm for optimal segmentation is presented. The algorithm automatically produces a compacted item‐set time series that can be analyzed to identify temporal patterns. The effectiveness of the approach is illustrated by analyzing version control repositories of several open‐source projects to identify time‐varying patterns of developer activity. The experimental results show that the segmentation algorithm produces segments that capture meaningful information and is superior to the information content obtained by arbitrarily segmenting software history into regular time intervals. Copyright © 2008 John Wiley & Sons, Ltd. Harvey P. Siy, Parvathi Chundi, Daniel J. Rosenkrantz, Mahadevan Subramaniam |
J. Softw. Maintenance Res. Pract. | 1 |
| 2007 | Discovering Dynamic Developer Relationships from Software Version Histories by Time Series SegmentationabstractTime series analysis is a promising approach to discover temporal patterns from time stamped, numeric data. A novel approach to apply time series analysis to discern temporal information from software version repositories is proposed. Version logs containing numeric as well as non-numeric data are represented as an item-set time series. A dynamic programming based algorithm to optimally segment an item-set time series is presented. The algorithm automatically produces a compacted item-set time series that can be analyzed to discern temporal patterns. The effectiveness of the approach is illustrated by applying to the Mozilla data set to study the change frequency and developer activity profiles. The experimental results show that the segmentation algorithm produces segments that capture meaningful information and is superior to the information content obtaining by arbitrarily segmenting time period into regular time intervals. Harvey P. Siy, Parvathi Chundi, Daniel J. Rosenkrantz, Mahadevan Subramaniam |
ICSM | 1 |
| 2001 | Does the Modern Code Inspection Have Value?abstractFor years, it was believed that the value of inspections is in finding and fixing defects early in the development process. Otherwise, the cost to find and fix them later is much higher However in examining code inspection data, we are finding that inspections are beneficial for an additional reason. They make the code easier to understand and change. An analysis of data from a recent code inspection experiment shows that 60% of all issues raised in the code inspections are not problems that could have been uncovered by latter phases of testing or field usage because they have little or nothing to do with the visible execution behavior of the software. Rather they improve the maintainability of the code by making the code conform to coding standards, minimizing redundancies, improving language proficiency, improving safety and portability, and raising the quality of the documentation. We conclude that even if advances in software technology have diminished the value of inspections as a defect detection tool, in most cases, it continues to be of value as a maintenance tool. Harvey P. Siy, Lawrence G. Votta |
ICSM | 1 |
| 2001 | Parallel changes in large-scale software development: an observational case studyabstractAn essential characteristic of large-scale software development is parallel development by teams of developers. How this parallel development is structured and supported has a profound effect on both the quality and timeliness of the product. We conduct an observational case study in which we collect and analyze the change and configuration management history of a legacy system to delineate the boundaries of, and to understand the nature of, the problems encountered in parallel development. The results of our studies are (1) that the degree of parallelism is very highhigher than considered by tool builders; (2) there are multiple levels of parallelism, and the data for some important aspects are uniform and consistent for all levels; (3) the tails of the distributions are long, indicating the tail, rather than the mean, must receive serious attention in providing solutions for these problems; and (4) there is a significant correlation between the degree of parallel work on a given component and the number of quality problems it has. Thus, the results of this study are important both for tool builders and for process and project engineers. Dewayne E. Perry, Harvey P. Siy, Lawrence G. Votta |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2000 | Software product lines: a case studyabstractA software product line is a family of products that share common features to meet the needs of a market area. Systematic processes have been developed to dramatically reduce the cost of a product line. Such product-line engineering processes have proven practical and effective in industrial use, but are not widely understood. The Family-Oriented Abstraction, Specification and Translation (FAST) process has been used successfully at Lucent Technologies in over 25 domains, providing productivity improvements of as much as four to one. In this paper, we show how to use FAST to document precisely the key abstractions in a domain, exploit design patterns in a generic product-line architecture, generate documentation and Java code, and automate testing to reduce costs. The paper is based on a detailed case study covering all aspects from domain analysis through testing. Copyright © 2000 John Wiley & Sons, Ltd. Mark A. Ardis, Nigel Daley, Daniel Hoffman, Harvey P. Siy, David M. Weiss 0001 |
Softw. Pract. Exp. | 4 |
| 2000 | Predicting Fault Incidence Using Software Change HistoryabstractThis paper is an attempt to understand the processes by which software ages. We define code to be aged or decayed if its structure makes it unnecessarily difficult to understand or change and we measure the extent of decay by counting the number of faults in code in a period of time. Using change management data from a very large, long-lived software system, we explore the extent to which measurements from the change history are successful in predicting the distribution over modules of these incidences of faults. In general, process measures based on the change history are more useful in predicting fault rates than product metrics of the code: For instance, the number of times code has been changed is a better indication of how many faults it will contain than is its length. We also compare the fault rates of code of various ages, finding that if a module is, on the average, a year older than an otherwise similar module, the older module will have roughly a third fewer faults. Our most successful model measures the fault potential of a module as the sum of contributions from all of the times the module has been changed, with large, recent changes receiving the most weight. Todd L. Graves, Alan F. Karr, J. S. Marron, Harvey P. Siy |
IEEE Trans. Software Eng. | 4 |
| 1998 | Parallel Changes in Large Scale Software Development: An Observational Case StudyabstractAn essential characteristic of large scale software development is parallel development by teams of developers. How this parallel development is structured and supported has a profound effect on both the quality and timeliness of the product. We conduct an observational case study in which me collect and analyze the change and configuration management history of a legacy system to delineate the boundaries of, and to understand the nature of, the problems encountered in parallel development. The results of our studies are: 1) that the degree of parallelism is very high-higher than considered by tool builders; 2) there are multiple levels of parallelism and the data for some important aspects are uniform and consistent for all levels and 3) the tails of the distributions are long, indicating the tail, rather than the mean, must receive serious attention in providing solutions for these problems. Dewayne E. Perry, Harvey P. Siy, Lawrence G. Votta |
ICSE | 2 |
| 1998 | Understanding the Sources of Variation in Software InspectionsabstractIn a previous experiment, we determined how various changes in three structural elements of the software inspection process (team size and the number and sequencing of sessions) altered effectiveness and interval. Our results showed that such changes did not significantly influence the defect detection rate, but that certain combinations of changes dramatically increased the inspection interval. We also observed a large amount of unexplained variance in the data, indicating that other factors must be affecting inspection performance. The nature and extent of these other factors now have to be determined to ensure that they had not biased our earlier results. Also, identifying these other factors might suggest additional ways to improve the efficiency of inspections. Acting on the hypothesis that the “inputs” into the inspection process (reviewers, authors, and code units) were significant sources of variation, we modeled their effects on inspection performance. We found that they were responsible for much more variation in detect detection than was process structure. This leads us to conclude that better defect detection techniques, not better process structures, are the key to improving inspection effectiveness. The combined effects of process inputs and process structure on the inspection interval accounted for only a small percentage of the variance in inspection interval. Therefore, there must be other factors which need to be identified. Adam A. Porter, Harvey P. Siy, Audris Mockus, Lawrence G. Votta |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 1997 | Understanding the Effects of Developer Activities on Inspection IntervalabstractWe have conducted an industrial experiment to assess the cost-benefit tradeoffs of several software inspection processes. Our results to date explain the variation in observed effectiveness very well, but are unable to satisfactorily explain variation in inspection interval. In this article we examine the effect of a new factor - process environment - on inspection interval (calendar time needed to complete the inspection). Our analysis suggests that process environment does indeed influence inspection interval. in particular, we found that non-uniform work priorities, time-varying workloads, and deadlines have significant effects. Moreover, these experiences suggest that regression models are inherently inadequate for interval modeling, and that queueing models may be more effective. (Also cross-referenced as UMIACS-TR-97-19) Adam A. Porter, Harvey P. Siy, Lawrence G. Votta |
ICSE | 2 |
| 1997 | An Experiment ot Assess the Cost-Benefits of Code Inspections in Large Scale Software DevelopmentabstractWe conducted a long term experiment to compare the costs and benefits of several different software inspection methods. These methods were applied by professional developers to a commercial software product they were creating. Because the laboratory for this experiment was a live development effort, we took special care to minimize cost and risk to the project, while maximizing our ability to gather useful data. The article has several goals: (1) to describe the experiment's design and show how we used simulation techniques to optimize it; (2) to present our results and discuss their implications for both software practitioners and researchers; and (3) to discuss several new questions raised by our findings. For each inspection, we randomly assigned three independent variables: (1) the number of reviewers on each inspection team (1, 2, or 4); (2) the number of teams inspecting the code unit (1 or 2); and (3) the requirement that defects be repaired between the first and second team's inspections. The reviewers for each inspection were randomly selected without replacement from a pool of 11 experienced software developers. The dependent variables for each inspection included inspection interval (elapsed time), total effort, and the defect detection rate. Our results showed that these treatments did not significantly influence the defect detection effectiveness, but that certain combinations of changes dramatically increased the inspection interval. Adam A. Porter, Harvey P. Siy, Carol A. Toman, Lawrence G. Votta |
IEEE Trans. Software Eng. | 2 |
| 1995 | An Experiment to Assess the Cost-Benefits of Code Inspections in Large Scale Software DevelopmentabstractWe are conducting a long-term experiment (in progress) to compare the costs and benefits of several different software inspection methods.These methods are being applied by beliefs about the most cost-effective ways to conduct inspections and raise some questions about the feasibility of recently proposed methods.professional developers to a commercial software product Adam A. Porter, Harvey P. Siy, Carol A. Toman, Lawrence G. Votta |
SIGSOFT FSE | 2 |