VLDB 2026 Research / reviewers in the wild / expert
E. James Whitehead Jr.
dblp:w/EJWhiteheadJr · also Jim Whitehead
· DBLP profile ↗
80ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0002-6887-7330ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 40 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 23 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 12 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorComputer networks · 2Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SketchTiler: Co-Creative Sketch-to-Tilemap Conversion Using Wave Function CollapseabstractIn 2D game development, tilesets are commonly used to create environments and levels, forming what is known as a tilemap. The unique workflow and constraints of tile-based art can be tedious to work with, especially for designers accustomed to more traditional artistic methods. Although AI-assisted systems have been used across various artistic contexts to translate high-level instructions into low-level visual representations, their application to tilemap design remains largely unexplored. To address this gap, we developed SketchTiler, an application that interprets freehand sketch lines and converts them into tilebased representations. Using the Wave Function Collapse (WFC) algorithm, the tool is able to generate contextually appropriate structural representations from user sketches. Suggestions are provided for unfilled areas, informed by surrounding tile context and the overall map structure. Blythe Chen, Raven Cruz-James, Justin Lam, E. James Whitehead Jr. |
CoG | 4 |
| 2025 | Conversational Interactions with Procedural Generators using Large Language ModelsabstractThis paper explores the potential of Large Language Models (LLMs) to facilitate conversational natural language interactions to aid humans in mixed-initiative generation of game worlds.This paper explores the issues in creating a system that allows for rapid iteration in a turn-based user-LLM design software.We identify key E. James Whitehead Jr., Thomas Wessel, Blythe Chen, Raven Cruz-James, Luc Harnist, William Klunder, Justin Lam, Ethan Lin, Roman Luo, Naitik Poddar, Shiva Ravinutula, Alejandro Montoreano, Logan Shehane, Yazmyn Sims, Jarod Spangler, Michelle Tan, Zosia Trela |
FDG | 1 |
| 2025 | Pedestrian Archetypes - The Must-Have Pedestrian Models for Autonomous Vehicle Safety TestingabstractSimulation models of law-abiding pedestrians are all too prevalent. However, the models of dangerous pedestrians are frighteningly limited. This leads to the unreliability of autonomous driving in human habitats as they learn to deal with pedestrians exhaustively in a simulated world. There exists a shortage of data and analysis on dangerous pedestrians. Decision-making factors and behavior identification are not sufficient; a dangerous pedestrian model should consist of a collection of behaviors that represent their natural behavior pattern. On this basis, we propose to define these collections of behaviors in the form of pedestrian archetypes. Each archetype suggests a distinct pedestrian personality with their own special take on crossing the road, giving autonomous driving an opportunity to significantly improve their testing strategies and reliability against pedestrians. Golam Md Muktadir, Taorui Huang, Ritvik Bansal, Namita Gaidhani, S. M. Jubaer, Michael Lin, E. James Whitehead Jr. |
IV | 7 |
| 2025 | A Reflection on Change Classification in the Era of Large Language ModelsabstractChange classification, today known as Just-in-Time Defect Prediction, is a technique for predicting software bugs at the change level of granularity. Several ideas came together to form change classification: predictions on code changes, using word-level textual features, use of machine learning classifiers, and leveraging open source code repositories. While change classification has led to a robust line of research, it has not yet had significant industrial adoption. A key recommendation is to explore explainability features so developers can better understand why a prediction is being made. We explore how large language models can advance this work by providing prediction explanations and bug fix suggestions. Sunghun Kim 0001, Shivkumar Shivaji, E. James Whitehead Jr. |
IEEE Trans. Software Eng. | 3 |
| 2024 | Language-Driven Play: Large Language Models as Game-Playing Agents in Slay the SpireabstractOne of the major challenges in procedural generation of game rules is evaluating the generated content. Since the effect of a rule on game balance and complexity might not be immediately apparent, one way to evaluate such a content is to simulate the gameplay. To achieve this, it is necessary to create an agent capable of both playing the game and adjusting to alterations in the game’s design, which is often referred to as a general game-playing agent. Bahar Bateni, E. James Whitehead Jr. |
FDG | 2 |
| 2024 | Evaluating Exploratory Reading Groups for Supporting Undergraduate Research Pipelines in ComputingabstractThis paper reports on a summative analysis of Exploratory Reading Groups (ERGs), a low time-commitment, relational, student-led reading group program designed to provide students from any background and year with a broad exploration of computing research. Since prior work, the program was institutionalized as a 1-credit course with a greater emphasis on strengthening pipelines into research labs. In analyzing 3 quarters of data from 136 participants, we found diverse indicators of impact. Surprisingly, despite the lightweight nature of the program (∼ 2 hours/week), we observed a statistically significant increase in satisfaction with their intellectual development at the university; confidence in reading, presenting, and communicating about their field; sense of belonging for women and minoritized ethnic groups; alignment with faculty goals in joining research labs (greater desire to make a research contribution and publish, decreased desire to join for the purpose of exploration); and engagement in the ‘reconsideration’ dimension of career identity formation. Over 70% of the participants continued on into group research projects for undergraduate students. The effectiveness of this scalable, lightweight initiative shows the promise of ERGs as a tool to support students in computing when connected to group research projects and points to future research directions on designing other lightweight, relational, scalable learning experiences. David M. Torres-Mendoza, Saba Kheirinejad, Mustafa Ajmal, Ashwin Chembu, Dustin Palea, E. James Whitehead Jr., David Lee 0002 |
ICER (1) | 6 |
| 2024 | Adaptive Pedestrian Agent Modeling for Scenario-based Testing of Autonomous Vehicles through Behavior RetargetingabstractThis work proposes a new representation of pedestrian crossing scenarios and a hybrid modeling approach, RePed, that facilitates transferring microscopic behavior models from behavior research to higher-level trajectories. With this, real-world trajectory-based scenarios can be augmented with a diverse set of human crossing maneuvers, producing a wealth of new scenarios and addressing the scarcity of rare case data that existing works struggle to deal with. Leveraging the controllability of this modeling approach, perturbation-based augmentation can be applied to enrich scenarios further. In addition, the representation is rooted in the Ego vehicle’s coordinate system with a logical representation of roads. This design enables scenario retargeting to various road structures, traffic conditions, and ego vehicle behaviors. Thus, it strongly supports scenario-based testing by forcing pedestrians to produce certain situations in simulation even when the Ego Vehicle tries to evade them. Golam Md Muktadir, E. James Whitehead Jr. |
ICRA | 2 |
| 2024 | PedAnalyze - Pedestrian Behavior Annotator and OntologyabstractDeveloping safer autonomous vehicles necessitates extensive testing of pedestrian behavior, particularly in atypical situations. Existing datasets lack consistent annotations, with text-based explanations and per-frame annotations causing redundancy and obscuring temporal relationships. To address these issues, we propose PedAnalyze, a Python-based annotator that focuses on pedestrian and vehicle behavior and facilitates structured datasets with pre-defined tags. Our approach allows for both single-frame and multi-frame annotations, which reduces the number of repetitive tasks. In addition, we focus on curating datasets from dash-cam videos on platforms such as YouTube, capturing valuable and rare pedestrian-vehicle incidents. We aim to create a comprehensive pedestrian behavior ontology and dataset to advance autonomous driving system research and development. Taorui Huang, Golam Md Muktadir, Srishti Sripada, Rishi Saravanan, Amelia Yuan, E. James Whitehead Jr. |
IV | 6 |
| 2024 | HyGenPed: A Hybrid Procedural Generation Approach in Pedestrian Trajectory Modeling in Arbitrary Crosswalk AreaabstractWe propose a new method to create plausible pedestrian crossing trajectories that cover a given arbitrarily shaped crosswalk area for simulation-based testing of autonomous vehicles. This method addresses the crossing area coverage problem where the trajectories produced by the generative methods do not cover the entire area that pedestrians may possibly walk on. The actual area covered by pedestrians often differs from marked crosswalks on the road. Furthermore, in the case of jaywalking, the area can take a variety of shapes based on the road structure and surrounding places of interest. Our method is a constructive process that generates trajectories conditioned on an area defined with polygons. We demonstrate that the method can generate trajectories that cover a wide range of crossing areas, including ones from the InD dataset. Golam Md Muktadir, Xuyuan Cai, E. James Whitehead Jr. |
IV | 3 |
| 2023 | Structure and Coherence in City Road Network GenerationabstractWe analyze two popular tile-based Procedural Content Generation (PCG) methods, WaveFunctionCollapse and transformers, to explore different notions of conditional probability and their significant effect on the quality of results. Specifically, we seek to answer the question of which method produces higher-quality results for generating city road networks as 2D images. Road networks are an interesting domain since they have large scale structures which require consistency across many tiles. To experiment with this, first we use OpenStreetMap to create a dataset of 2D tiled images with additional information on the type of each road encoded in the tile tokens. We use this dataset to train two WFC models with different decision heuristics, and an additional transformer model with a sliding window inference process. We then compare the results of these models using a set of tile-based metrics (e.g. tile-frequency resemblance and edge-frequency resemblance) and urban-planning metrics (e.g. node density, road connectivity, etc.). Our results show that the transformer model outperforms the WFC methods in terms of generating high-quality city road networks, and demonstrate the potential of transformers for tile-based PCG methods, especially when a considerable amount of data is available. Bahar Bateni, E. James Whitehead Jr. |
CoG | 2 |
| 2023 | Challenges of End-to-End Testing with Selenium WebDriver and How to Face Them: A SurveyabstractModern web applications are complex and used for tasks of primary importance, so their quality must be guaranteed at the highest levels. For this reason, testing techniques (e.g., end-to-end) are required to validate the overall behavior of web applications. One of the most popular tools for testing web applications is Selenium WebDriver. Selenium WebDriver automates the browser to mimic real user actions on the web.While Selenium has made testing easier for many Teams worldwide, it still has its share of challenges. To better understand the challenges and the corresponding solutions adopted we decided to undertake a personal opinion survey from the industry (in total with 78 highly skilled participants) with a focus on the Selenium ecosystem.The results allow understanding which challenges are consid-ered more relevant by professionals in their daily practice and which are the techniques, approaches, and tools they adopt to face them. Therefore, this study is useful to (1) practitioners interested in understanding how to solve the problems they face every day and (2) researchers interested in proposing innovative solutions to problems having a solid industrial impact. Maurizio Leotta, Boni García, Filippo Ricca, E. James Whitehead Jr. |
ICST | 4 |
| 2023 | Procedural Generation of Complex Roundabouts for Autonomous Vehicle TestingabstractHigh-definition roads are an essential component of realistic driving scenario simulation for autonomous vehicle testing. Roundabouts are one of the key road segments that have not been thoroughly investigated. Based on the geometric constraints of the nearby road structure, this work presents a novel method for procedurally building roundabouts. The suggested method can result in roundabout lanes that are not perfectly circular and resemble real-world roundabouts by allowing approaching roadways to be connected to a roundabout at any angle. One can easily incorporate the roundabout in their HD road generation process or use the standalone roundabouts in scenario-based testing of autonomous driving. Zarif Ikram, Golam Md Muktadir, E. James Whitehead Jr. |
IV | 3 |
| 2022 | CogMod: Simulating Human Information Processing Limitation While DrivingabstractWe develop a human driver behavior model (CogMod) based on two complementary cognitive architectures; Queueing Network-Model Human Processor (QN-MHP) and Adaptive Control of Thought - Rational (ACT-R), to represent human cognition while driving. The proposed model can integrate different task-specific analytical driver models under a similar cognitive procedure. The model can simulate variable cognitive processing ability, resulting in different stopping distances in a scenario where the front vehicle brakes sharply when it enters a trigger distance. We evaluate the model based on the distribution of stopping distance with varying cognitive processing time. This approach is useful for modeling non-ego vehicles in scenario-based testing of automated vehicles (AVs). Abdul Jawad, E. James Whitehead Jr. |
IV | 2 |
| 2022 | Adversarial jaywalker modeling for simulation-based testing of Autonomous Vehicle SystemsabstractWe present an approach for creating adversarial jaywalkers, autonomous pedestrian models which intentionally act to create unsafe situations involving other vehicles. An adversarial jaywalker employs a hybrid state-model with social forces and state transition rules. The parameters (for social forces and state transitions) of this model are tuned via reinforcement learning to create risky situations faster with synthetic yet plausible behavior. The resulting jaywalkers are capable of realistic behavior while still engaging in sufficiently risky actions to be useful for testing. These adversarial pedestrian models are useful in a wide range of scenario-based tests for autonomous vehicles. Golam Md Muktadir, E. James Whitehead Jr. |
IV | 2 |
| 2022 | Just-in-time defect prediction for software hunksabstractAbstract Just‐in‐time defect prediction can remind software developers and managers to verify and fix bugs at the moment they appeared, thus improving the effectiveness and validity of bug fixing. Existing studies mainly focus on just‐in‐time prediction for software files (JIT‐F). JIT‐F is a binary classification problem, which classifies (hence predicts) a file change as buggy or clean. This article provides a detailed analysis of just‐in‐time defect prediction for software hunks (JIT‐H), which predicts bugs at a finer level of granularity, and hence further improves the efficiency of bug fixing. Classification is performed using the ensemble technique of bagging—aggregated combinations of random under sampling plus multiple classifiers (J48 and Random Forest). An empirical study with 10 open source projects was conducted to validate the effectiveness of JIT‐H. Experimental results show that JIT‐H is effective at predicting defects in software hunk changes. Compared with JIT‐F, JIT‐H is more cost effective. Additionally, analysis on the change features indicates that Text Vector features and hunk change level features are of more importance than features in other groups and levels. Xiaoyan Zhu 0003, Chenyu Yan, E. James Whitehead Jr., Binbin Niu, Lei Zhu 0011, Long Pan |
Softw. Pract. Exp. | 3 |
| 2020 | Scheherazade's Tavern: A Prototype For Deeper NPC InteractionsabstractIn many games, NPC-player interactions play a vital role in gameplay. Previous literature has successfully shown how NPC interaction focused on social simulation is an effective means for creating dynamic characters such as in the games Prom Week and Versu. We believe that social simulation is a key element in the creation of complex characters, which is further aided by natural language interaction and knowledge modeling. In this work, we propose an architecture for player-NPC interactions built on top of the Ensemble engine that additionally incorporates chatbots and knowledge modeling technology, with the objective of making craftable and interesting NPCs more easily authorable. Rehaf Aljammaz, Elisabeth Oliver, E. James Whitehead Jr., Michael Mateas |
FDG | 3 |
| 2020 | Spatial Layout of Procedural Dungeons Using Linear Constraints and SMT SolversabstractDungeon generation is among the oldest problems in procedural content generation. Creating the spatial aspects of a dungeon requires three steps: random generation of rooms and sizes, placement of these rooms inside a fixed area, and connecting rooms with passageways. This paper uses a series of integer linear constraints, solved by a satisfiability modulo theories (SMT) solver, to perform the placement step. Separation constraints ensure dungeon rooms do not intersect and maintain a minimum fixed separation. Designers can specify control lines, and dungeon rooms will be placed within a fixed distance of these control lines. Generation times vary with number of rooms and constraints, but are often very fast. Spatial distribution of solutions tend to have hot spots, but is surprisingly uniform given the underlying complexity of the solver. The approach demonstrates the effectiveness of a declarative approach to dungeon layout generation, where designers can express desired intent, and the SMT solver satisfies this if possible. E. James Whitehead Jr. |
FDG | 1 |
| 2020 | A Modular Architecture for Procedural Generation of Towns, Intersections and Scenarios for Testing Autonomous VehiclesabstractSimulation-based testing is critical for ensuring safety of autonomous vehicles. Autonomous vehicles are enabled by deep learning techniques which require a large quantity of data. With simulation testing, we can create rare events for testing and training of autonomous vehicles. Procedural generation of roads and modeling of driving behaviors in an easily extendable architecture ensures that we are able to create rare scenarios at scale with minimal artistic burden. In this paper, we present CruzWay, a system that both supports and creates these scenarios. With CruzWay, we are able to procedurally generate town sized road networks or road intersections. CruzWay supports generation of road meshes as well as navigation meshes from SUMO road network files. CruzWay can generate cars as well as pedestrians run by behavior trees (BTs) in this environment. The self-contained, modular nature of BTs in combination with procedural roads allows us to create a large number of scenarios. Ishaan Paranjape, Abdul Jawad, Yanwen Xu, Asiiah Song, E. James Whitehead Jr. |
IV | 5 |
| 2019 | A methodology for designing natural language interfaces for procedural content generationabstractProcedural Content Generation (PCG) uses algorithmic techniques to create a wide variety of content for games. These generators often have a large number of parameters, making it difficult for non-technical designers to explore the design space of generated artifacts. Natural language interfaces for generators can map natural language keywords to parameter space changes spanning multiple simultaneous parameters and afford use of expressive language. This way, designers can navigate to interesting points in the design space of a generator by describing desired properties of the artifact using a series of natural language descriptors. We present a design methodology that designers can use to develop natural language interfaces for procedural content generation systems. This design methodology begins by defining a design vocabulary that can describe the output of a generator, mapping the vocabulary to a series of parameters, and translating natural language queries to movements in the generator's design space. We further address issues around designer intent understanding, design space exploration and workflows using natural language interfaces in PCG. An example and implementation of our methodology is provided demonstrating its application to existing plug-ins for content creation in the Unity3D engine Afshin Mobramaein, E. James Whitehead Jr. |
FDG | 2 |
| 2019 | TownSim: agent-based city evolution for naturalistic road network generationabstractWe describe an agent-based city evolution algorithm creating road networks over time, and explore several approaches for analyzing the malleability of the algorithm to exposed parameters. In addition to qualitatively assessing the generated content, we look at the directionality, connectivity, and curvature of the generated road networks. Asiiah Song, E. James Whitehead Jr. |
FDG | 2 |
| 2019 | Identifying the Within-Statement Changes to Facilitate Change UnderstandingabstractAs current tree-differencing approaches ignore changes that occur within a statement or do not present them in an abstract way, it is difficult to automatically understand revisions involving statement updates. We propose a tree-differencing approach to identifying the within-statement changes. It calculates edit operations based on an element-sensitive strategy and the longest common sequence algorithm. Then, it generates the metadata for each edit operation. Meta-data include the type of operation, the type of entity and the name of the element part, the content, the content pattern and all references involved. We have implemented the approach as a free accessible tool. It is built upon ChangeDistiller and refines its statement-update type. Finally, to demonstrate how to use the proposed approach for change understanding, we studied the condition-expression changes in four open projects. We analyzed the non-essential condition changes, the effective changes that definitely affect the condition, and other changes. The results show that for revisions with condition-expression changes, nearly 20% contain non-essential changes, while more than 60% have effective changes. Furthermore, we found many common patterns. For example, we found that half of the revisions with effective changes were caused by adding or removing expressions in logical expressions. And, in these revisions, 47% enhanced the condition, while 49% weakened it. Chunhua Yang 0004, E. James Whitehead Jr. |
ICSME | 2 |
| 2019 | Pruning the AST with Hunks to Speed up Tree DifferencingabstractInefficiency is a problem in tree-differencing approaches. As textual-differencing approaches are highly efficient and the hunks returned by them reflect the line range of the modified code, we propose a novel approach to prune the AST with hunks. We define the pruning strategies at the declaration level and the statement level, respectively. And, we have designed an algorithm to implement the strategies. Furthermore, we have integrated the algorithm to Change Distiller and GumTree. Through an evaluation on four open source projects, the results show that the approach is very effective in reducing the number of nodes and shortening the running time. On average, with declaration-level pruning, the number of nodes in the two tools is reduced by at least 64%. With statement-level pruning, the number of nodes in both tools is reduced by at least 74%. By using the declaration-level pruning and the statement-level pruning, GumTree's runtime is reduced by at least 70% and 75%, respectively. Chunhua Yang 0004, E. James Whitehead Jr. |
SANER | 2 |
| 2018 | Cognitive and Experiential Interestingness in Abstract Visual Narrative
Morteza Behrooz, Afshin Mobramaein, Arnav Jhala, E. James Whitehead Jr. |
CogSci | 4 |
| 2018 | A Taxonomy of Code Changes Occurring within a Statement or a SignatureabstractWe propose a taxonomy of code changes at a granularity finer than the statement level. It classifies changes that occur within a statement or signature. We firstly define the changes based on the proposed operations on the tree of the statement or signature. Then, we classify the changes according to action type, entity kind, element kind and pattern kind. Based on the current implementation, we applied it to four open Java projects. Through the case study, we validated that the taxonomy can classify changes in all the modified statements and signatures. And, we checked the proportions of change patterns. Furthermore, we demonstrated that it is easy to find out the rename-induced statement updates with the help of the taxonomy. As a result, this taxonomy can be used for further change understanding and change analysis. Chunhua Yang 0004, E. James Whitehead Jr. |
TASE | 2 |
| 2018 | An empirical study of software change classification with imbalance data-handling methodsabstractSummary Bug prediction in software code changes can help developers to find out and fix bugs immediately when they are introduced, thus to improve the effectiveness and validity of bug fixing. In data mining, this problem can be regarded as a change classification task. However, one of its key characteristics, ie, class‐imbalance, holds back the performance of standard classification methods. In this paper, we consider a quantity of imbalance data‐handling methods and extract a more comprehensive groups of change features, aiming to achieve better change classification performance. Two different types of imbalance data‐handling methods, namely, resampling and ensemble learning methods, are employed. Especially, we explore the performance of their combination. To compare the performance of different imbalance data‐handling methods, an experiment with 10 open source projects is conducted. Four classification methods, including J48, Naïve Bayes, SMO, and Random Forest, are used as standard classifiers and as the base classifiers, respectively. Moreover, contribution of different groups of change features are evaluated. Experimental results show that imbalance data‐handling methods can improve the performance of change classification and the combination methods, which take advantage of both ensemble learning and resampling, perform better than using ensemble learning methods or resampling methods individually. Of the studied imbalance data‐handling methods, the combination of Bagging and random undersampling with J48 as the base classifier yields out better prediction results than those achieved by other methods. Additionally, of the collected change features, text vector features accounts for a larger proportion than others. Xiaoyan Zhu 0003, Binbin Niu, E. James Whitehead Jr., Zhongbin Sun |
Softw. Pract. Exp. | 3 |
| 2017 | Towards generative emotions in games based on cognitive modelingabstractProcedural Content Generation (PCG) accomplishes feats that once would be considered magic, creating near-infinite amounts of unique levels, worlds, objects and other content for games. Yet, despite this near magical quality, generated content is often found to have a sameness to it. After a short time it loses the interest of players. Many procedurally generated games, such as No Man's Sky, have disappointed customers in this sense. We argue that one important cause is generated content fails to create an emotional connection with players. Emotions help in keeping players engaged with game content [7] and thereby improve the gameplay experience. Among the many emotions players experience in games, surprise is perhaps the most important for PCG. Surprise intensifies other emotions [9] and it lies at the origin of humor, strategy and problem solving [15]. Thus, surprise helps to increase player enjoyment and engagement with games. So far, PCG in games has produced surprise by accident of chance. PCG systems which intentionally create surprising moments, in a controllable way, can play an important role in increasing engagement and interest in games. Chandranil Chakraborttii, Lucas Ferreira, E. James Whitehead Jr. |
FDG | 3 |
| 2017 | Solusforge: controlling the generation of the 3D models with spatial relation graphsabstractIn this paper, we propose Solusforge, a system for automatically generating Lego models from a graph of the components' spatial relationships. The system uses a two step constraint solving approach in which the spatial layout is solved for first, followed by the specific pieces that make up the model, thereby allowing us to explore two separate solution spaces independently. This technology has many uses, including in games featuring a system of snap-together pieces, including Kerbal Space Program, Beseiged, and Spore. While many of these games involve procedurally augmenting human generated design, none of them feature a fully procedural system for generating the artifacts within that space. Stella Mazeika, E. James Whitehead Jr. |
FDG | 2 |
| 2017 | Art and science of engineered design: what kind of discipline is PCG?abstractWhat kind of discipline is PCG? PCG research can be viewed as science, engineering, design, and art. PCG is thus a multidiscipline, drawing from a broad set of epistemic traditions. E. James Whitehead Jr. |
FDG | 1 |
| 2017 | Reflections on the REST architectural style and "principled design of the modern web architecture" (impact paper award)abstractSeventeen years after its initial publication at ICSE 2000, the Representational State Transfer (REST) architectural style continues to hold significance as both a guide for understanding how the World Wide Web is designed to work and an example of how principled design, through the application of architectural styles, can impact the development and understanding of large-scale software architecture. However, REST has also become an industry buzzword: frequently abused to suit a particular argument, confused with the general notion of using HTTP, and denigrated for not being more like a programming methodology or implementation framework. Roy T. Fielding, Richard N. Taylor, Justin R. Erenkrantz, Michael Martin Gorlick, E. James Whitehead Jr., Rohit Khare, Peyman Oreizy |
ESEC/SIGSOFT FSE | 5 |
| 2016 | Crowdsourcing program preconditions via a classification gameabstractInvariant discovery is one of the central problems in software verification. This paper reports on an approach that addresses this problem in a novel way; it crowdsources logical expressions for likely invariants by turning invariant discovery into a computer game. The game, called Binary Fission, employs a classification model. In it, players compose preconditions by separating program states that preserve or violate program assertions. The players have no special expertise in formal methods or programming, and are not specifically aware they are solving verification tasks. We show that Binary Fission players discover concise, general, novel, and human readable program preconditions. Our proof of concept suggests that crowdsourcing offers a feasible and promising path towards the practical application of verification technology. Daniel S. Fava, Daniel G. Shapiro, Joseph C. Osborn, Martin Schäf, E. James Whitehead Jr. |
ICSE | 5 |
| 2016 | Multistaging to understand: Distilling the essence of java code examplesabstractProgrammers commonly search the Web to find code examples that can help them solve a specific programming task. While some novice programmers may be willing to spend as much time as needed to understand a found code example, more experienced ones want to spend as little time as possible. They want to get a quick overview of the example's operation, so they can start working with it immediately. Getting this overview is often non-trivial and requires a tedious and manual inspection process. In this paper, we introduce a technique called Multi-staging to Understand, which streamlines this inspection process by distilling the essence of code examples. The essence of a code example conveys the most important aspects of the example's intended function. Our technique automatically decomposes the code in an example into code stages that can be explored non-sequentially; enabling fast exploratory learning. We discuss the key components of our technique and describe empirical results based on actual code examples on StackOverflow. Huascar Sanchez, E. James Whitehead Jr., Martin Schäf |
ICPC | 2 |
| 2015 | Playing with Recipes
Johnathan Pagnutti, E. James Whitehead Jr. |
FDG | 2 |
| 2015 | Generative Mixology: An Engine for Creating Cocktails
Johnathan Pagnutti, E. James Whitehead Jr. |
ICCC | 2 |
| 2015 | 4th International Workshop on Games and Software Engineering (GAS 2015)abstractWe present a summary of the 4th ICSE Workshop on Games and Software Engineering. The full day workshop is planned to include a keynote speaker, game-jam demonstration session, and paper presentations on game software engineering topics related to software engineering education, frameworks for game development and infrastructure, quality assurance, and model-based game development. The accepted papers are overviewed here. Judith Bishop, Kendra M. L. Cooper, Walt Scacchi, E. James Whitehead Jr. |
ICSE (2) | 4 |
| 2015 | Source Code Curation on StackOverflow: The Vesperin SystemabstractThe past few years have witnessed the rise of software question and answer sites like StackOverflow, where developers can pose detailed coding questions and receive quality answers. Developers using these sites engage in a complex code foraging process of understanding and adapting the code snippets they encounter. We introduce the notion of source code curation to cover the act of discovering some source code of interest, cleaning and transforming (refining) it, and then presenting it in a meaningful and organized way. In this paper, we present Vesperin, a source code curation system geared towards curating Java code examples on StackOverflow. Huascar Sanchez, E. James Whitehead Jr. |
ICSE (2) | 2 |
| 2015 | Why Power Laws? An Explanation from Fine-Grained Code ChangesabstractThroughout the years, empirical studies have found power law distributions in various measures across many software systems. However, surprisingly little is known about how they are produced. What causes these power law distributions? We offer an explanation from the perspective of fine-grained code changes. A model based on preferential attachment and self-organized criticality is proposed to simulate software evolution. The experiment shows that the simulation is able to render power law distributions out of fine-grained code changes, suggesting preferential attachment and self-organized criticality are the underlying mechanism causing the power law distributions in software systems. Zhongpeng Lin, E. James Whitehead Jr. |
MSR | 2 |
| 2015 | An analysis of programming language statement frequency in C, C++, and Java source codeabstractSummary Statement frequency data can inform programming language research and provide a solid basis for frequency‐based code analysis. This paper presents an analysis of programming language statement frequency in a large corpus of C, C++, and Java source code, comprised of more than 54 million lines of code. Across these languages, the top four work‐performing statement types are Method/Function Call, Assignment, If, and Return. As compared to studies of Formula Translating System, Common Business Oriented Language and Programming Language One in the 1970s, the main change is the prevalence of method/function calls. Statement use frequency across languages is remarkably similar, and within each individual language, most statement types have a frequency distribution that occupies a small range. A more detailed examination of assignment and looping statement types shows that many assignments simply involve copying of data and that C++/Java useforstatements more than C. Copyright © 2014 John Wiley & Sons, Ltd. Xiaoyan Zhu 0003, E. James Whitehead Jr., Caitlin Sadowski, Qinbao Song |
Softw. Pract. Exp. | 2 |
| 2014 | Xylem: The Code of Plants
Heather Logas, E. James Whitehead Jr., Michael Mateas, Richard Vallejos, Lauren Scott, John T. Murray, Kate Compton, Joseph C. Osborn, Orlando Salvatore, Daniel G. Shapiro, Zhongpeng Lin, Huascar Sanchez, Michael Shavlovsky, Chris Lewis 0002, Daniel Cetina, Shayne Clementi |
FDG | 2 |
| 2014 | Software verification games: Designing Xylem, The Code of Plants
Heather Logas, E. James Whitehead Jr., Michael Mateas, Richard Vallejos, Lauren Scott, Daniel G. Shapiro, John T. Murray, Kate Compton, Joseph C. Osborn, Orlando Salvatore, Zhongpeng Lin, Huascar Sanchez, Michael Shavlovsky, Daniel Cetina, Shayne Clementi, Chris Lewis 0002 |
FDG | 2 |
| 2013 | Does bug prediction support human developers? findings from a google case studyabstractWhile many bug prediction algorithms have been developed by academia, they're often only tested and verified in the lab using automated means. We do not have a strong idea about whether such algorithms are useful to guide human developers. We deployed a bug prediction algorithm across Google, and found no identifiable change in developer behavior. Using our experience, we provide several characteristics that bug prediction algorithms need to meet in order to be accepted by human developers and truly change how developers evaluate their code. Chris Lewis 0002, Zhongpeng Lin, Caitlin Sadowski, Xiaoyan Zhu 0003, Rong Ou, E. James Whitehead Jr. |
ICSE | 6 |
| 2013 | Reducing Features to Improve Code Change-Based Bug PredictionabstractMachine learning classifiers have recently emerged as a way to predict the introduction of bugs in changes made to source code files. The classifier is first trained on software history, and then used to predict if an impending change causes a bug. Drawbacks of existing classifier-based bug prediction techniques are insufficient performance for practical use and slow prediction times due to a large number of machine learned features. This paper investigates multiple feature selection techniques that are generally applicable to classification-based bug prediction methods. The techniques discard less important features until optimal classification performance is reached. The total number of features used for training is substantially reduced, often to less than 10 percent of the original. The performance of Naive Bayes and Support Vector Machine (SVM) classifiers when using this technique is characterized on 11 software projects. Naive Bayes using feature selection provides significant improvement in buggy F-measure (21 percent improvement) over prior change classification bug prediction results (by the second and fourth authors [28]). The SVM's improvement in buggy F-measure is 9 percent. Interestingly, an analysis of performance for varying numbers of features shows that strong performance is achieved at even 1 percent of the original number of features. Shivkumar Shivaji, E. James Whitehead Jr., Ram Akella, Sunghun Kim 0001 |
IEEE Trans. Software Eng. | 2 |
| 2012 | Motivational game design patterns of 'ville gamesabstractThe phenomenal growth of social network games in the last five years has left many game designers, game scholars, and long-time game players wondering how these games so effectively engage their audiences. Without a strong understanding of the sources of appeal of social network games, and how they relate to the appeal of past games and other human activities, it has proven difficult to interpret the phenomenon accurately or build upon its successes. In this paper we propose and employ a particular approach to this challenge, analyzing the motivational game design patterns in the popular 'Ville style of game using the lenses of behavioral economics and behavioral psychology, explaining ways these games engage and retain players. We show how such games employ strategies in central, visible ways that are also present (if perhaps harder to perceive) in games with very different mechanics and audiences. Our conclusions point to lessons for game design, game interpretation, and the design of engaging software of any type. Chris Lewis 0002, Noah Wardrip-Fruin, E. James Whitehead Jr. |
FDG | 3 |
| 2012 | PCG-based game design: creating Endless WebabstractThis paper describes the creation of the game Endless Web, a 2D platforming game in which the player's actions determine the ongoing creation of the world she is exploring. Endless Web is an example of a PCG-based game: it uses procedural content generation (PCG) as a mechanic, and its PCG system, Launchpad, greatly influenced the aesthetics of the game. All of the player's strategies for the game revolve around the use of procedural content generation. Many design challenges were encountered in the design and creation of Endless Web, for both the game and modifications that had to be made to Launchpad. These challenges arise largely from a loss of fine-grained control over the player's experience; instead of being able to carefully craft each element the player can interact with, the designer must instead craft algorithms to produce a range of content the player might experience. In this paper we provide a definition of PCG-based game design and describe the challenges faced in creating a PCG-based game. We offer our solutions, which impacted both the game and the underlying level generator, and identify issues which may be particularly important as this area matures. Gillian Smith 0001, Alexei Othenin-Girard, E. James Whitehead Jr., Noah Wardrip-Fruin |
FDG | 3 |
| 2012 | Introduction to the special issue on software repository mining in 2009
Michael W. Godfrey, E. James Whitehead Jr. |
Empir. Softw. Eng. | 2 |
| 2012 | Introduction to the Special Issue on Mining Software Repositories in 2010
E. James Whitehead Jr., Thomas Zimmermann 0001 |
Empir. Softw. Eng. | 1 |
| 2011 | Situating Quests: Design Patterns for Quest and Level Design in Role-Playing Games
Gillian Smith 0001, Brian Kopleck, Zach Lindblad, Lauren Scott, Adam Wardell, E. James Whitehead Jr., Michael Mateas |
ICIDS | 7 |
| 2011 | Workshop on games and software engineering: (GAS 2011)abstractAt the core of video games are complex interactions leading to emergent behaviors. This complexity creates difficulties architecting components, predicting their behaviors and testing the results. The Workshop on Games and Software Engineering (GAS 2011) provides an opportunity for software engineering researchers and practitioners who work with games to come together and discuss how these two areas can be intertwined. E. James Whitehead Jr., Chris Lewis 0002 |
ICSE | 1 |
| 2011 | An empirical analysis of the FixCache algorithmabstractThe FixCache algorithm, introduced in 2007, effectively identifies files or methods which are likely to contain bugs by analyzing source control repository history. However, many open questions remain about the behaviour of this algorithm. What is the variation in the hit rate over time? How long do files stay in the cache? Do buggy files tend to stay buggy, or can they be redeemed? This paper analyzes the behaviour of the FixCache algorithm on four open source projects. FixCache hit rate is found to generally increase over time for three of the four projects; file duration in cache follows a Zipf distribution; and topmost bug-fixed files go through periods of greater and lesser stability over a project's history. Caitlin Sadowski, Chris Lewis 0002, Zhongpeng Lin, Xiaoyan Zhu 0003, E. James Whitehead Jr. |
MSR | 5 |
| 2011 | Fantasy, farms, and freemium: what game data mining teaches us about retention, conversion, and virality (keynote abstract)abstractIn December of 2010, the new game CityVille achieved 6 million daily active users in just 8 days. Clearly the success of CityVille owes something to the fun gameplay experience it provides. That said, it was far from the best game released in 2010. Why did it grow so fast? In this talk the key factors behind the dramatic success of social network games are explained. Social network games build word-of-mouth player acquisition directly into the gameplay experience via friend invitations and game mechanics that require contributions by friends to succeed. Software analytics (mined data about player sessions) yield detailed models of factors that affect player retention and engagement. Player engagement is directly related to conversion, shifting a free player into a paying player, the critical move in a freemium business model. Analytics also permit tracking of player virality, the degree to which one player invites other players into the game. Social network games offer multiple lessons for software engineers in general, and software mining researchers in particular. Since software is in competition for people's attention along with a wide range of other media and software, it is important to design software for high engagement and retention. Retention engineering requires constant attention to mined user experience data, and this data is easiest to acquire with web-based software. Building user acquisition directly into software provides powerful benefits, especially when it is integrated deeply into the experience delivered by the software. Since retention engineering and viral user acquisition are much easier with web-based software, the trend of software applications migrating to the web will accelerate. E. James Whitehead Jr. |
MSR | 1 |
| 2011 | Tanagra: Reactive Planning and Constraint Solving for Mixed-Initiative Level DesignabstractTanagra is a mixed-initiative tool for level design, allowing a human and a computer to work together to produce a level for a 2-D platformer. An underlying, reactive level generator ensures that all levels created in the environment are playable, and provides the ability for a human designer to rapidly view many different levels that meet their specifications. The human designer can iteratively refine the level by placing and moving level geometry, as well as through directly manipulating the pacing of the level. This paper presents the design environment, its underlying architecture that integrates reactive planning and numerical constraint solving, and an evaluation of Tanagra's expressive range. Gillian Smith 0001, E. James Whitehead Jr., Michael Mateas |
IEEE Trans. Comput. Intell. AI Games | 2 |
| 2011 | Launchpad: A Rhythm-Based Level Generator for 2-D PlatformersabstractLaunchpad is an autonomous level generator that is based on a formal model of 2-D platformer level design. Levels are built out of small segments called “rhythm groups,” which are generated using a two-tiered, grammar-based approach. These segments are pieced together into complete levels that are then rated according to a set of design heuristics. Generation can be controlled using a set of parameters that influence the level pacing and geometry. The approach minimizes the amount of content that must be manually authored: instead of piecing together large segments of a level, Launchpad uses base components that are commonly found in a number of 2-D platformers. Launchpad produces an impressive variety of levels which are all guaranteed to be playable. Gillian Smith 0001, E. James Whitehead Jr., Michael Mateas, Mike Treanor, Jameka March, Mee Cha |
IEEE Trans. Comput. Intell. AI Games | 2 |
| 2011 | Guest Editorial: Procedural Content Generation in GamesabstractThe eight papers in this special issue focus on procedural content generation in games. They present a good combination of surveys, conceptual frameworks, innovative methods, and applications. Julian Togelius, E. James Whitehead Jr., Rafael Bidarra |
IEEE Trans. Comput. Intell. AI Games | 2 |
| 2010 | Design patterns in FPS levelsabstractLevel designers create gameplay through geometry, AI scripting, and item placement. There is little formal understanding of this process, but rather a large body of design lore and rules of thumb. As a result, there is no accepted common language for describing the building blocks of level design and the gameplay they create. This paper presents level design patterns for first-person shooter (FPS) games, providing cause-effect relationships between level design elements and gameplay. These patterns allow designers to create more interesting and varied levels. Kenneth Hullett, E. James Whitehead Jr. |
FDG | 2 |
| 2010 | What went wrong: a taxonomy of video game bugsabstractVideo games are complex, emergent systems that are difficult to design and test. This difficulty invariably leads to failures being present in the game, negatively impacting the play experience of some. We present a taxonomy of possible failures, divided into temporal and non-temporal failures. The taxonomy can guide the thinking of designers and testers alike, helping them expose bugs in the game. This will lead to games being better tested and designed, with fewer failures when released. Chris Lewis 0002, E. James Whitehead Jr., Noah Wardrip-Fruin |
FDG | 2 |
| 2010 | Tanagra: a mixed-initiative level design toolabstractTanagra is a prototype mixed-initiative design tool for 2D platformer level design, in which a human and computer can work together to produce a level. The human designer can place constraints on a continuously running level generator, in the form of exact geometry placement and manipulation of the level's pacing. The computer then fills in the rest of the level with geometry that guarantees playability, or informs the designer that there is no level that meets their requirements. This paper presents the design of Tanagra, a discussion of the editing operations it provides to the designer, and an evaluation of the expressivity of its generator. Gillian Smith 0001, E. James Whitehead Jr., Michael Mateas |
FDG | 2 |
| 2010 | Runtime repair of software faults using event-driven monitoringabstractIn software with emergent properties, despite the best efforts to remove faults before execution, there is a high likelihood that faults will occur during runtime. These faults can lead to unacceptable program behavior during execution, even leading to the program terminating unexpectedly. Using a distributed event-driven runtime software-fault monitor to repair faulty states creates an enforceable runtime specification. Using such an architecture can help ensure that emergent systems operate within specification, increasing the reliability of such software. Chris Lewis 0002, E. James Whitehead Jr. |
ICSE (2) | 2 |
| 2009 | Rhythm-based level generation for 2D platformersabstractWe present a rhythm-based method for the automatic generation of levels for 2D platformers, where the rhythm is that which the player feels with his hands while playing. Levels are created using a grammar-based method: first generating rhythms, then generating geometry based on those rhythms. Generation is constrained by a set of style parameters tweakable by a human designer. The approach also minimizes the amount of content that must be manually authored, instead relying on geometry components that are included in the level designer's tileset and a set of jump types. Our results show that this method produces an impressive variety of levels, all of which are fully playable. Gillian Smith 0001, Mike Treanor, E. James Whitehead Jr., Michael Mateas |
FDG | 3 |
| 2009 | Kenyon-web: Reconfigurable web-based feature extractorabstractResearch on Mining Software Repositories (MSR) has yielded fruitful results in many Software Engineering areas including software change comprehension, bug prediction, and developer network recovery. When performing MSR research, the first task is to extract features corresponding to source code details from repositories. Since reusable feature extraction tools are not available, each MSR research group builds their own extraction tool, a duplication of effort. We introduce a reusable feature extractor, Kenyon-web, for MSR research. Kenyon-web is fully reconfigurable, pluggable, and serves most MSR related tasks. In this report, we show the architecture of Kenyon-web and demonstrate its utility by showcasing a sample MSR task. Sunghun Kim 0001, Shivkumar Shivaji, E. James Whitehead Jr. |
ICPC | 3 |
| 2009 | Reducing Features to Improve Bug PredictionabstractRecently, machine learning classifiers have emerged as a way to predict the existence of a bug in a change made to a source code file. The classifier is first trained on software history data, and then used to predict bugs. Two drawbacks of existing classifier-based bug prediction are potentially insufficient accuracy for practical use, and use of a large number of features. These large numbers of features adversely impact scalability and accuracy of the approach. This paper proposes a feature selection technique applicable to classification-based bug prediction. This technique is applied to predict bugs in software changes, and performance of Naive Bayes and Support Vector Machine (SVM) classifiers is characterized. Shivkumar Shivaji, E. James Whitehead Jr., Ram Akella, Sunghun Kim 0001 |
ASE | 2 |
| 2009 | Toward an understanding of bug fix patterns
Kai Pan, Sunghun Kim 0001, E. James Whitehead Jr. |
Empir. Softw. Eng. | 3 |
| 2008 | Rhizome: A Feature Modeling and Generation PlatformabstractRhizome is an end-to-end feature modeling and code generation platform that includes a feature modeling language (FeatureML), a template language (MarkerML) and a template-based code generator. A software designer creates feature models using FeatureML by selecting and defining design choices. These design choices can be automatically associated with code templates and interpreted as parameter values for code generation. The code generator then replaces markers embedded in the code templates with dynamically generated code blocks to produce source code. Guozheng Ge, E. James Whitehead Jr. |
ASE | 2 |
| 2008 | Understanding bug fix patterns in verilogabstractToday, many electronic systems are developed using a hardware description language, a kind of software that can be converted into integrated circuits or programmable logic devices. Like traditional software projects, hardware projects have bugs, and significant developer time is spent fixing them. A useful first step toward reducing bugs in hardware is developing an understanding of the frequency of different types of errors. Once the most common types are known, it is then possible to focus attention on eliminating them. As most hardware projects use software configuration management repositories, these can be mined for the textual bug fix changes. In this project, we analyze the bug fix history of four hardware projects written in Verilog and manually define 25 bug fix patterns. The frequency of each bug type is then computed for all projects. We find that 29 -- 55% of the bug fix pattern instances in Verilog involve assignment statements, while 18 -- 25% are related to if statements. Sangeetha Sudhakrishnan, Janaki T. Madhavan, E. James Whitehead Jr., Jose Renau |
MSR | 3 |
| 2008 | Classifying Software Changes: Clean or Buggy?abstractThis paper introduces a new technique for finding latent software bugs called change classification. Change classification uses a machine learning classifier to determine whether a new software change is more similar to prior buggy changes, or clean changes. In this manner, change classification predicts the existence of bugs in software changes. The classifier is trained using features (in the machine learning sense) extracted from the revision history of a software project, as stored in its software configuration management repository. The trained classifier can classify changes as buggy or clean with 78% accuracy and 65% buggy change recall (on average). Change classification has several desirable qualities: (1) the prediction granularity is small (a change to a single file), (2) predictions do not require semantic information about the source code, (3) the technique works for a broad array of project types and programming languages, and (4) predictions can be made immediately upon completion of a change. Contributions of the paper include a description of the change classification approach, techniques for extracting features from source code and change histories, a characterization of the performance of change classification across 12 open source projects, and evaluation of the predictive power of different groups of features. Sunghun Kim 0001, E. James Whitehead Jr., Yi Zhang 0001 |
IEEE Trans. Software Eng. | 2 |
| 2007 | Predicting Faults from Cached HistoryabstractWe analyze the version history of 7 software systems to predict the most fault prone entities and files. The basic assumption is that faults do not occur in isolation, but rather in bursts of several related faults. Therefore, we cache locations that are likely to have faults: starting from the location of a known (fixed) fault, we cache the location itself, any locations changed together with the fault, recently added locations, and recently changed locations. By consulting the cache at the moment a fault is fixed, a developer can detect likely fault-prone locations. This is useful for prioritizing verification and validation resources on the most fault prone files or entities. In our evaluation of seven open source projects with more than 200,000 revisions, the cache selects 10% of the source code files; these files account for 73%-95% of faults - a significant advance beyond the state of the art. Sunghun Kim 0001, Thomas Zimmermann 0001, E. James Whitehead Jr., Andreas Zeller |
ICSE | 3 |
| 2006 | Properties of Signature Change PatternsabstractUnderstanding function signature change properties and evolution patterns is important for researchers concerned with alleviating signature change impacts, understanding software evolution, and predicting future evolution patterns. We provide detailed signature change properties by analyzing seven software project histories to reveal multiple properties of signature changes, including their kind, frequency, correlation with other changes, number of parameter changes, and evolution patterns of signature change kinds. We show that signature changes can be used as measurement aid for software evolution analysis Sunghun Kim 0001, E. James Whitehead Jr. |
ICSM | 2 |
| 2006 | Automatic Identification of Bug-Introducing ChangesabstractBug-fixes are widely used for predicting bugs or finding risky parts of software. However, a bug-fix does not contain information about the change that initially introduced a bug. Such bug-introducing changes can help identify important properties of software bugs such as correlated factors or causalities. For example, they reveal which developers or what kinds of source code changes introduce more bugs. In contrast to bug-fixes that are relatively easy to obtain, the extraction of bugintroducing changes is challenging. In this paper, we present algorithms to automatically and accurately identify bug-introducing changes. We remove false positives and false negatives by using annotation graphs, by ignoring non-semantic source code changes, and outlier fixes. Additionally, we validated that the fixes we used are true fixes by a manual inspection. Altogether, our algorithms can remove about 38%~51% of false positives and 14%~15% of false negatives compared to the previous algorithm. Finally, we show applications of bug-introducing changes that demonstrate their value for research. Sunghun Kim 0001, Thomas Zimmermann 0001, Kai Pan, E. James Whitehead Jr. |
ASE | 4 |
| 2006 | Memories of bug fixesabstractThe change history of a software project contains a rich collection of code changes that record previous development experience. Changes that fix bugs are especially interesting, since they record both the old buggy code and the new fixed code. This paper presents a bug finding algorithm using bug fix memories: a project-specific bug and fix knowledge base developed by analyzing the history of bug fixes. A bug finding tool, BugMem, implements the algorithm. The approach is different from bug finding tools based on theorem proving or static model checking such as Bandera, ESC/Java, FindBugs, JLint, and PMD. Since these tools use pre-defined common bug patterns to find bugs, they do not aim to identify project-specific bugs. Bug fix memories use a learning process, so the bug patterns are project-specific, and project-specific bugs can be detected. The algorithm and tool are assessed by evaluating if real bugs and fixes in project histories can be found in the bug fix memories. Analysis of five open source projects shows that, for these projects, 19.3%-40.3% of bugs appear repeatedly in the memories, and 7.9%-15.5% of bug and fix pairs are found in memories. The results demonstrate that project-specific bug fix patterns occur frequently enough to be useful as a bug detection technique. Furthermore, for the bug and fix pairs, it is possible to both detect the bug and provide a strong suggestion for the fix. However, there is also a high false positive rate, with 20.8%-32.5% of non-bug containing changes also having patterns found in the memories. A comparison of BugMem with a bug finding tool, PMD, shows that the bug sets identified by both tools are mostly exclusive, indicating that BugMem complements other bug finding tools. Copyright ACM 2006. Sunghun Kim 0001, Kai Pan, E. James Whitehead Jr. |
SIGSOFT FSE | 3 |
| 2005 | Automatic generation of rule-based software configuration management systemsabstractWe propose a model-driven methodology and toolset for automatic SCM system repository creation and feature composition using code generation and rule engine technologies. Guozheng Ge, E. James Whitehead Jr. |
ICSE | 2 |
| 2005 | Bamboo: an architecture modeling and code generation framework for configuration management systemsabstractWe describe an architecture modeling and code generation framework called Bamboo. Using Bamboo, engineers design SCM repository and feature models, and then generate a running SCM system from the models. Guozheng Ge, E. James Whitehead Jr. |
ASE | 2 |
| 2005 | Facilitating software evolution research with kenyonabstractSoftware evolution research inherently has several resource-intensive logistical constraints. Archived project artifacts, such as those found in source code repositories and bug tracking systems, are the principal source of input data. Analysis-specific facts, such as commit metadata or the location of design patterns within the code, must be extracted for each change or configuration of interest. The results of this resource-intensive "fact extraction" phase must be stored efficiently, for later use by more experimental types of research tasks, such as algorithm or model refinement. In order to perform any type of software evolution research, each of these logistical issues must be addressed and an implementation to manage it created. In this paper, we introduce Kenyon, a system designed to facilitate software evolution research by providing a common set of solutions to these common logistical problems. We have used Kenyon for processing source code data from 12 systems of varying sizes and domains, archived in 3 different types of software configuration management systems. We present our experiences using Kenyon with these systems, and also describe Kenyon's usage by students in a graduate seminar class. Jennifer Bevan, E. James Whitehead Jr., Sunghun Kim 0001, Michael W. Godfrey |
ESEC/SIGSOFT FSE | 2 |
| 2004 | The WebDAV property designabstractAbstract This paper provides a detailed description of the general design space for metadata storage capabilities. The design space considers issues of metadata identification, typing and representation, dynamic behavior, predefined and user‐defined metadata, schema discovery/update, operations, API packaging/marshalling, searching, and versioning. The design space is used to structure a retrospective analysis of the three major alternative metadata designs considered during the design of the WebDAV distributed authoring protocol. Deployment experience with WebDAV properties is also discussed, with the most successful use occurring in custom client/server pairs and in protocol extensions. Copyright © 2004 John Wiley & Sons, Ltd. E. James Whitehead Jr., Yaron Y. Goland |
Softw. Pract. Exp. | 1 |
| 2002 | A Proposed Curriculum for a Masters in Web Engineering
E. James Whitehead Jr. |
J. Web Eng. | 1 |
| 2001 | Supporting integrated voice and data traffic over EGPRSabstractWe investigate the performance of voice and best-effort data traffic over Enhanced General Packet Radio Services (EGPRS) using tight reuse configurations (e.g. 1/3 or 1/1), with the aim of identifying suitable bearer designs and optimal system configurations for these services. We find that voice and data have somewhat conflicting preferences: voice capacity is optimized using random frequency hopping and quality-driven power control: data throughput is optimized using no frequency hopping and no power control. These somewhat contradictory requirements impose a challenge in the system design. We investigate two simple traffic integration strategies: (1) segregating the resource for the two services and using the optimal configuration in each segment; and (2) integrating the traffic by allowing voice and data to share the same resource, and use a compromised configuration for the entire system. We find that if the majority of traffic is voice, traffic integration has an advantage since data can use any residual voice capacity. On the other hand, if the majority of traffic is data, segregation provides higher spectrum efficiency. Furthermore, the integrated traffic option becomes increasingly sub-optimal with increasing data traffic load. We conclude that the system design should depend upon the expected mix of traffic (including traffic types in addition to voice and best-effort data) and the desired engineering complexity, and that further work is needed to understand the possible design options. One possible design option is to soft segregate the resource, e.g., to allow data to use the voice segment on an overflow basis. Similar strategies may be good avenues for future work. Xiaoxin Qiu, Li-Fung Chang, Kapil K. Chawla, Justin C.-I. Chuang, Nelson Sollenberger, E. James Whitehead Jr. |
ICC | 6 |
| 2001 | Panel: Perspectives on Software Engineering
David Notkin, Marc Donner, Michael D. Ernst, Michael M. Gorlick, E. James Whitehead Jr. |
ICSE | 5 |
| 2000 | Supporting voice over EGPRS: system design and capacity evaluationabstractWe investigate the possibility of supporting voice over EGPRS as a potential evolutionary path for 2/sup nd/ generation wireless TDMA systems. We discuss different voice bearer designs for EGPRS radio access networks. We demonstrate how the operating mode can be chosen at each layer in the protocol stack to form appropriate voice bearers, and quantify the performance of these bearers. Different system control options are examined, including frequency reuse factor, MAC layer operating mode, frequency hopping, power control, and dynamic channel assignment. We find that a baseline EGPRS system using random channel assignment and no power control can support up to 27 Erlang/Site/MHz, which is significantly higher than the capacity of current 2/sup nd/ generation TDMA systems. In addition, the packet-switched nature of EGPRS makes it a very flexible system. We conclude that EGPRS offers an excellent evolutionary path for 2/sup nd/ generation TDMA systems. Xiaoxin Qiu, Li-Fung Chang, Kapil K. Chawla, Justin C.-I. Chuang, Nelson Sollenberger, E. James Whitehead Jr. |
PIMRC | 6 |
| 2000 | Chimera: hypermedia for heterogeneous software development enviromentsabstractEmerging software development environments are characterized by heterogeneity: they are composed of diverse object stores, user interfaces, and tools. This paper presents an approach for providing hypermedia services in this heterogeneous setting. Central notions of the approach include the following: anchors are established with respect to interactiveviewsof objects, rather than the objects themselves; composable,n-ary links can be established between anchors on different views of objects which may be stored in distinct object bases; viewers may be implemented in different programming languages; and, hypermedia services are provided to multiple, concurrently active, viewers. The paper describes the approach, supporting architecture, and lessons learned. Related work in the areas of supporing heterogeneity and hypermedia data modeling is discussed. The system has been employed in a variety of contexts including research, development, and education. Kenneth M. Anderson, Richard N. Taylor, E. James Whitehead Jr. |
ACM Trans. Inf. Syst. | 3 |
| 1999 | WebDAV: A network protocol for remote collaborative authoring on the Web
E. James Whitehead Jr., Yaron Y. Goland |
ECSCW | 1 |
| 1999 | Performance comparison of link adaptation and incremental redundancy in wireless data networksabstractLink adaptation and incremental redundancy are two link level techniques proposed to enhance data rate and increase throughput in wireless data networks. Link adaptation pro-actively adjusts the modulation and coding scheme based upon estimated channel conditions. Incremental redundancy, on the other hand, adjusts the code rate to actual channel conditions by incrementally transmitting redundancy information until decoding is successful. In this study, we investigate the system performance of these two techniques. Our results show that incremental redundancy achieves significantly higher throughput compared to link adaptation, especially under high traffic loading conditions. This gain is achieved at the expense of higher terminal and base station complexities, as well as a higher packet delay. On the other hand, in contrast to link adaptation, incremental redundancy does not require any channel quality measurements, and can operate based solely upon ACKs and NACKs. Xiaoxin Qiu, Justin C.-I. Chuang, Kapil K. Chawla, E. James Whitehead Jr. |
WCNC | 4 |
| 1996 | A Component- and Message-Based Architectural Style for GUI SoftwareabstractWhile a large fraction of application code is devoted to graphical user interface (GUI) functions, support for reuse in this domain has largely been confined to the creation of GUI toolkits ("widgets"). We present a novel architectural style directed at supporting larger grain reuse and flexible system composition. Moreover, the style supports design of distributed, concurrent applications. Asynchronous notification messages and asynchronous request messages are the sole basis for intercomponent communication. A key aspect of the style is that components are not built with any dependencies on what typically would be considered lower-level components, such as user interface toolkits. Indeed, all components are oblivious to the existence of any components to which notification messages are sent. While our focus has been on applications involving graphical user interfaces, the style has the potential for broader applicability. Several trial applications using the style are described. Richard N. Taylor, Nenad Medvidovic, Kenneth M. Anderson, E. James Whitehead Jr., Jason E. Robbins, Kari A. Nies, Peyman Oreizy, Deborah L. Dubrow |
IEEE Trans. Software Eng. | 4 |
| 1995 | A Component- and Message-Based Architectural Style for GUI SoftwareabstractWhile a large j7action of application system code is devoted to user interface (U[)fi.mctions,support for reuse in this domain has largely been conjined to creation of UI toolkits ("widgets").We present a novel architectural style directed at supporting larger grain reuse andjexible system composition.Moreoveq the style supports design of distributed, concurrent, applications.A key aspect of the style is that components are not built with any dependencies on what typically would be considered lower-level components, such as user interface toolkits.Indeed, all components are oblivious to the existence of any components to which notijcation messages are sent.Asynchronous notification messages and asynchronous request messages are the sole basis for inter-component communication.While our focus has been on applications involving graphical user interfaces, the style has the potential for broader applicability.Several trial applications using the style are described~. Richard N. Taylor, Nenad Medvidovic, Kenneth M. Anderson, E. James Whitehead Jr., Jason E. Robbins |
ICSE | 4 |