James Walden

dblp:21/5259 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
2since 2021 · last 2022
0000-0002-7272-7538ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1
YearPublicationVenuePosition
2022 OpenSSL 3.0.0: An exploratory case study
abstract
Context: The OpenSSL project released version 3.0.0 in September 2021. This release was a departure from previous versions of OpenSSL in several ways, including a new versioning system and the first use of public software design documents.
James Walden
MSR1
2021 An Exploratory Study of Project Activity Changepoints in Open Source Software Evolution
abstract
To explore the prevalence of abrupt changes (changepoints) in open source project activity, we assembled a dataset of 8,919 projects from the World of Code. Projects were selected based on age, number of commits, and number of authors. Using the nonparametric PELT algorithm, we identified changepoints in project activity time series, finding that more than 90% of projects had between one and six changepoints. Increases and decreases in project activity occurred with roughly equal frequency. While most changes are relatively small, on the order of a few authors or few dozen commits per month, there were long tails of much larger project activity changes. In future work, we plan to focus on larger changes to search for common open source lifecycle patterns as well as common responses to external events.
James Walden, Noah Burgin, Kuljit Kaur
MSR1
2020 The Impact of a Major Security Event on an Open Source Project: The Case of OpenSSL
abstract
Context: The Heartbleed vulnerability brought OpenSSL to international attention in 2014. The almost moribund project was a key security component in public web servers and over a billion mobile devices. This vulnerability led to new investments in OpenSSL.
James Walden
MSR1
2017 The Effect of Dimensionality Reduction on Software Vulnerability Prediction Models
abstract
Statistical prediction models can be an effective technique to identify vulnerable components in large software projects. Two aspects of vulnerability prediction models have a profound impact on their performance: 1) the features (i.e., the characteristics of the software) that are used as predictors and 2) the way those features are used in the setup of the statistical learning machinery. In a previous work, we compared models based on two different types of features: software metrics and term frequencies (text mining features). In this paper, we broaden the set of models we compare by investigating an array of techniques for the manipulation of said features. These techniques fall under the umbrella of dimensionality reduction and have the potential to improve the ability of a prediction model to localize vulnerabilities. We explore the role of dimensionality reduction through a series of cross-validation and cross-project prediction experiments. Our results show that in the case of software metrics, a dimensionality reduction technique based on confirmatory factor analysis provided an advantage when performing cross-project prediction, yielding the best F-measure for the predictions in five out of six cases. In the case of text mining, feature selection can make the prediction computationally faster, but no dimensionality reduction technique provided any other notable advantage.
Jeff Stuckman, James Walden, Riccardo Scandariato
IEEE Trans. Reliab.2
2014 Predicting Vulnerable Components: Software Metrics vs Text Mining
abstract
Building secure software is difficult, time-consuming, and expensive. Prediction models that identify vulnerability prone software components can be used to focus security efforts, thus helping to reduce the time and effort required to secure software. Several kinds of vulnerability prediction models have been proposed over the course of the past decade. However, these models were evaluated with differing methodologies and datasets, making it difficult to determine the relative strengths and weaknesses of different modeling techniques. In this paper, we provide a high-quality, public dataset, containing 223 vulnerabilities found in three web applications, to help address this issue. We used this dataset to compare vulnerability prediction models based on text mining with models using software metrics as predictors. We found that text mining models had higher recall than software metrics based models for all three applications.
James Walden, Jeff Stuckman, Riccardo Scandariato
ISSRE1
2014 Predicting Vulnerable Software Components via Text Mining
abstract
This paper presents an approach based on machine learning to predict which components of a software application contain security vulnerabilities. The approach is based on text mining the source code of the components. Namely, each component is characterized as a series of terms contained in its source code, with the associated frequencies. These features are used to forecast whether each component is likely to contain vulnerabilities. In an exploratory validation with 20 Android applications, we discovered that a dependable prediction model can be built. Such model could be useful to prioritize the validation activities, e.g., to identify the components needing special scrutiny.
Riccardo Scandariato, James Walden, Aram Hovsepyan, Wouter Joosen
IEEE Trans. Software Eng.2
2013 Accelerating e-commerce sites in the cloud
abstract
E-commerce web sites typically have large fluctuations in their IT resource usage, while rapid elasticity is an essential characteristic of cloud computing. These characteristics make the cloud a good fit for hosting e-commerce web sites. Cloud providers deploy their cloud in their data centers. However, cloud providers usually have a limited number of data center locations around the world. Thus, e-commerce web sites in the cloud may be far away from their customers. Long client-perceived response latency may cause e-commerce web sites to lose business. To solve this problem, we propose a virtual proxy solution to reduce the response latency of the e-commerce site in the cloud. In our approach, a virtual proxy platform is designed to cache applications and data of e-commerce sites. A k-means based table partitioning algorithm is designed to select frequently used data from the database in the cloud. We have used an industrial e-commerce benchmark TPC-W to evaluate the performance of our approach. The experimental results show that our approach can significantly reduce the client-perceived response time.
Wei Hao 0001, James Walden, Chris Trenkamp
CCNC2
2013 Static analysis versus penetration testing: A controlled experiment
abstract
Suppose you have to assemble a security team, which is tasked with performing the security analysis of your organization's latest applications. After researching how to assess your applications, you find that the most popular techniques (also offered by most security consultancies) are automated static analysis and black box penetration testing. Under time and budget constraints, which technique would you use first? This paper compares these two techniques by means of an exploratory controlled experiment, in which 9 participants analyzed the security of two open source blogging applications. Despite its relative small size, this study shows that static analysis finds more vulnerabilities and in a shorter time than penetration testing.
Riccardo Scandariato, James Walden, Wouter Joosen
ISSRE2
2013 An informatics perspective on computational thinking
abstract
In this paper, we examine computational thinking and its connections to critical thinking from the perspective of in- formatics. We developed an introductory course for students in our College of Informatics, which includes majors rang- ing from journalism to computer science. The course cov- ered a set of principles of informatics, using both lectures and active learning sessions designed to develop informat- ics and computational thinking skills. The set of principles was drawn from a wide set of sources, and included broad principles like those of Denning and Loidl, as well as more limited principles related to topics like universal computa- tion and undecidability. We evaluated the change in both computational and critical thinking skills over the course of the semester, using a well-known validated critical thinking test and a computational thinking test of our own devising.
James Walden, Maureen Doyle, Rudy Garns, Zachary Hart
ITiCSE1
2011 Profiling file repository access patterns for identifying data exfiltration activities
abstract
Studies show that a significant number of employees steal data when changing jobs. Insider attackers who have the authorization to access the best-kept secrets of organizations pose a great challenge for organizational security. Although increasing efforts have been spent on identifying insider attacks, little research concentrates on detecting data exfiltration activities. This paper proposes a model for identifying data exfiltration activities by insiders. It uses statistical methods to profile legitimate uses of file repositories by authorized users. By analyzing legitimate file repository access logs, user access profiles are created and can be employed to detect a large set of data exfiltration activities. The effectiveness of the proposed model was tested with file access histories from the subversion logs of the popular open source project KDE.
Charles E. Frank, James Walden, Emily Crawford, Dhanuja Kasturiratna
CICS3
2011 Work in progress - Does maintenance first improve student's understanding and appreciation of clean code and documentation
abstract
The ACM's “Computer Science Curriculum 2008: An Interim Revision of CS 2001” suggests 31 core hours of Software Engineering covering the standard phases of software development. At many universities, including ours, Software Engineering is a capstone course with a semester-long team project. While the course covers the software development lifecycle there is not enough time in a single semester for students to gain an appreciation for all phases, especially the importance of testing and documentation to the maintenance phase which is the longest phase in successful software projects. To address this, a new course, Software Maintenance and Testing, was added as a pre-requisite to Software Engineering. The course will be implemented in fall, 2011. Students will be surveyed to determine if their appreciation for different development models, documentation, recording of design decisions, and clean code are improved by maintaining code prior to developing a brand-new or greenfield project. A Likert scale instrument will be develop to evaluate students' attitudes. The research instrument will be tuned until validity and reliability are achieved. Preliminary and baseline results will be presented.
Maureen Doyle, Brooke Buckley, Wei Hao 0001, James Walden
FIE4
2010 An effective log mining approach for database intrusion detection
abstract
Organizations spend a significant amount of resources securing their servers and network perimeters. However, these mechanisms are not sufficient for protecting databases. In this paper, we present a new technique for identifying malicious database transactions. Compared to many existing approaches which profile SQL query structures and database user activities to detect intrusions, the novelty of this approach is the automatic discovery and use of essential data dependencies, namely, multi-dimensional and multi-level data dependencies, for identifying anomalous database transactions. Since essential data dependencies reflect semantic relationships among data items and are less likely to change than SQL query structures or database user behaviors, they are ideal for profiling data correlations for identifying malicious database activities.
Alina Campan, James Walden, Irina Vorobyeva, Justin Shelton
SMC3
2009 Security of open source web applications
abstract
In an empirical study of fourteen widely used open source PHP Web applications, we found that the vulnerability density of the aggregate code base decreased from 8.88 vulnerabilities/KLOC to 3.30 from Summer 2006 to Summer 2008. Individual web applications varied widely, with vulnerability densities ranging from 0 to 121.4 at the beginning of the study. While the total number of security problems decreased, vulnerability density increased in eight of the fourteen applications over the analysis period. We developed a security resources indicator metric, which we found to be strongly correlated (rho = 0.67, p < 0.05) with change in vulnerability density over time. Traditional software metrics, such as code size, cyclomatic complexity, nesting complexity, and churn, had significant (p < 0.05) but much smaller correlations (rho = 0.31 at best) with vulnerability density. Vulnerability density was measured using the fortify source code analyzer static analysis tool.
James Walden, Maureen Doyle, Grant A. Welch, Michael Whelan
ESEM1
2005 A real-time information warfare exercise on a virtual network
abstract
Information warfare exercises, such as "Capture the Flag," serve as a capstone experience for a computer security class, giving students the opportunity to apply and integrate the security skills they learned during the class. However, many information security classes don't offer such exercises, because they can be difficult, expensive, time-consuming, and risky to organize and implement. This paper describes a real-time "Capture the Flag" exercise, implemented using a virtual network with free, open-source software to reduce the risk and effort of conducting such an exercise.
James Walden
SIGCSE1