VLDB 2026 Research / reviewers in the wild / expert
Lori L. Pollock
dblp:p/LLPollock
· DBLP profile ↗
123ranked-venue papers
18as first author
9since 2021 · last 2024
0000-0002-4388-4892ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 92 · 8 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 20 · 10 first-author · 6 since 2021Systems, architecture and hardware · 11Databases, data management, data science and information retrieval · 9Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Customizing ChatGPT to Help Computer Science Principles Students Learn Through ConversationabstractThis paper explores leveraging conversational agents, specifically ChatGPT, to enhance the introduction of computing, focused on the Advanced Placement Computer Science Principles (CSP) course in secondary schools. Despite the potential benefits for diverse student audiences, little research has investigated their effectiveness and engagement in this context. We examine the customization of ChatGPT for secondary school CSP students, assessing its impact on exploratory searches for learning CSP concepts. Results from 20 high school students in grades 10-12 (ages 15-18) in a CSP course indicate that students preferred a customized ChatGPT, with its terminology more suitable to secondary school level, examples more understandable, and better connections to personal experiences compared to standard ChatGPT. Matthew Frazier, Kostadin Damevski, Lori L. Pollock |
ITiCSE (1) | 3 |
| 2023 | Message from the ICSE 2023 Program Co-ChairsabstractWelcome to ICSE 2023! It is our great pleasure to introduce the program of the 45th IEEE/ACM International Conference on Software Engineering (ICSE 2023), which will be held in Melbourne, Australia, on May 14-20, 2023. Lori L. Pollock, Massimiliano Di Penta |
ICSE | 1 |
| 2023 | Using Domain-Specific, Immediate Feedback to Support Students Learning Computer Programming to Make MusicabstractBroadening participation in computer science has been widely studied, creating many different techniques to attract, motivate, and engage students. A common meta-strategy is to use an outside domain as a hook, using the concepts in that domain to teach computer science. These domains are selected to interest the student, but students often lack a strong background in these domains. Therefore, a strategy designed to increase students' interest, motivation, and engagement could actually create more barriers for students, who now are faced with learning two new topics. To reduce this potential barrier in the domain of music, this paper presents the use of automated, immediate feedback during programming activities at a summer camp that uses music to teach foundational programming concepts. The feedback guides students musically, correcting notes that are out-of-key or rhythmic phrases that are too long or short, allowing students to focus their learning on the computer science concepts. This paper compares the correctness of students that received automated feedback with students that did not, which shows the effectiveness of the feedback. Follow up focus groups with students confirmed this quantitative data, with students claiming that the feedback was not only useful but that the activities would be much more challenging without the feedback. Douglas Lusa Krug, Chrystalla Mouza, Taylor Barnett, Lori L. Pollock, David C. Shepherd |
ITiCSE (1) | 5 |
| 2023 | Experiences Piloting a Diversity and Inclusion in Computing Innovations CourseabstractWith society's increasing dependence on computing innovations---especially technologies that impact decision-making in fields such as healthcare, financial services, child welfare, hiring, safety, and policing---it is increasingly important for the future creators of these innovations to learn how technologies can potentially negatively impact people of different identities and backgrounds. Unfortunately, few universities offer courses designed specifically for Computer Science and Engineering students to explore the issues of diversity, equity and inclusion of computing innovations. Minji Kong, Lori L. Pollock |
SIGCSE (1) | 2 |
| 2022 | A Case Study of Middle Schoolers' Use of Computational Thinking Concepts and Practices during Coded Music CompositionabstractResearchers and practitioners have demonstrated various benefits of introducing computational thinking (CT) through music composition coding. While researchers have studied the impacts on participant attitudes towards CT and their learning of CT concepts, more case studies are needed on both learning CT concepts as well as CT practices, i.e., the processes of constructing music coding projects. This paper presents a case study of middle schoolers in an informal learning environment focused on integrating music composition with coding in TunePad. Specifically, we collected and analyzed logs of coding events, final code products, and surveys to explore both CT concept use and CT practices exhibited by the participants as they completed open-ended music coding activities to create their own melodies with specific music and CT requirements and recommendations. Douglas Lusa Krug, Chrystalla Mouza, David C. Shepherd, Lori L. Pollock |
ITiCSE (1) | 5 |
| 2021 | Automatic Extraction of Opinion-based Q&A from Online Developer ChatsabstractVirtual conversational assistants designed specifically for software engineers could have a huge impact on the time it takes for software engineers to get help. Research efforts are focusing on virtual assistants that support specific software development tasks such as bug repair and pair programming. In this paper, we study the use of online chat platforms as a resource towards collecting developer opinions that could potentially help in building opinion Q&A systems, as a specialized instance of virtual assistants and chatbots for software engineers. Opinion Q&A has a stronger presence in chats than in other developer communications, thus mining them can provide a valuable resource for developers in quickly getting insight about a specific development topic (e.g., What is the best Java library for parsing JSON?). We address the problem of opinion Q&A extraction by developing automatic identification of opinion-asking questions and extraction of participants' answers from public online developer chats. We evaluate our automatic approaches on chats spanning six programming communities and two platforms. Our results show that a heuristic approach to opinion-asking questions works well (.87 precision), and a deep learning approach customized to the software domain outperforms heuristics-based, machine-learning-based and deep learning for answer extraction in community question answering. Preetha Chatterjee, Kostadin Damevski, Lori L. Pollock |
ICSE | 3 |
| 2021 | Code Beats: A Virtual Camp for Middle Schoolers Coding Hip HopabstractIn spite of the efforts to provide computer science education for all, the percentage of Black and Latino Americans entering the computer science (CS) field has been stagnant for years. In an effort to attract and engage students many summer camps and after-school clubs use robotics, video-games, and even IoT devices, but these approaches seem to only attract those already considering STEM careers, a population low in Black and Latino students. To attract Black and Latino students to computer science a promising approach is to engage with their culture, making CS relevant to them personally. To this end, we present an approach that teaches middle school students to program using hip hop beats, intentionally leveraging a genre of music that appeals to a wide array of urban youth of color. This approach, called Code Beats, uses extensive scaffolding to support beginning students, authentic-sounding beats to engage students, and a expressive programming environment to support creative freedom. We present the results of our pilot camp, where students clearly showed an increase in computing enjoyment, confidence, belonging, and persistence. By the end of this course, all students were able to create their own, original beat from scratch, suggesting their progression to the Create phase of the Use-Modify-Create framework. Douglas Lusa Krug, Edtwuan Bowman, Taylor Barnett, Lori L. Pollock, David C. Shepherd |
SIGCSE | 4 |
| 2021 | Exploring Computational Thinking Across Disciplines Through Student-Generated Artifact AnalysisabstractTo meet the demands of 21st century societies, it is essential that faculty across disciplines engage students with course activities and assignments that foster the development of computational thinking (CT). In this study, we address two pertinent questions: (1) What types of artifacts do students develop across different disciplines in response to CT-driven problem prompts' and (2) What types of CT skills do these artifacts demonstrate? To answer the questions, we examined 273 artifacts developed by undergraduate students across seven course assignments from four disciplines: mathematics, sociology, music, and English using a rubric developed to evaluate the following CT skills: abstraction, decomposition, data analysis, and algorithmic thinking. We found that a range of skills were reflected across student artifacts. Amanda Mohammad Mirzaei, Lori L. Pollock, Chrystalla Mouza, Kevin R. Guidry |
SIGCSE | 3 |
| 2021 | Automatically Identifying the Quality of Developer Chats for Post Hoc UseabstractSoftware engineers are crowdsourcing answers to their everyday challenges on Q&A forums (e.g., Stack Overflow) and more recently in public chat communities such as Slack, IRC, and Gitter. Many software-related chat conversations contain valuable expert knowledge that is useful for both mining to improve programming support tools and for readers who did not participate in the original chat conversations. However, most chat platforms and communities do not contain built-in quality indicators (e.g., accepted answers, vote counts). Therefore, it is difficult to identify conversations that contain useful information for mining or reading, i.e., conversations of post hoc quality. In this article, we investigate automatically detecting developer conversations of post hoc quality from public chat channels. We first describe an analysis of 400 developer conversations that indicate potential characteristics of post hoc quality, followed by a machine learning-based approach for automatically identifying conversations of post hoc quality. Our evaluation of 2,000 annotated Slack conversations in four programming communities (python, clojure, elm, and racket) indicates that our approach can achieve precision of 0.82, recall of 0.90, F-measure of 0.86, and MCC of 0.57. To our knowledge, this is the first automated technique for detecting developer conversations of post hoc quality. Preetha Chatterjee, Kostadin Damevski, Nicholas A. Kraft, Lori L. Pollock |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2020 | Achieving Reliable Sentiment Analysis in the Software Engineering Domain using BERTabstractResearchers have shown that sentiment analysis of software artifacts can potentially improve various software engineering tools, including API and library recommendation systems, code suggestion tools, and tools for improving communication among software developers. However, sentiment analysis techniques applied to software artifacts still have not yet yielded very high accuracy. Recent adaptations of sentiment analysis tools to the software domain have reported some improvements, but the f-measures for the positive and negative sentences still remain in the 0.4-0.64 range, which deters their practical usefulness for software engineering tools.In this paper, we explore the potential effectiveness of customizing BERT, a language representation model, which has recently achieved very good results on various Natural Language Processing tasks on English texts, for the task of sentiment analysis of software artifacts. We describe our application of BERT to analyzing sentiments of sentences in Stack Overflow posts and compare the impact of a BERT sentiment classifier to state-of-the-art sentiment analysis techniques when used on a domain-specific data set created from Stack Overflow posts. We also investigate how the performance of sentiment analysis changes when using a much (3 times) larger data set than previous studies. Our results show that the BERT classifier achieves reliable performance for sentiment analysis of software engineering texts. BERT combined with the larger data set achieves an overall f-measure of 0.87, with the f-measures for the negative and positive sentences reaching 0.91 and 0.78 respectively, a significant improvement over the state-of-the-art. Eeshita Biswas, Mehmet Efruz Karabulut, Lori L. Pollock, K. Vijay-Shanker |
ICSME | 3 |
| 2020 | Software-related Slack Chats with Disentangled ConversationsabstractMore than ever, developers are participating in public chat communities to ask and answer software development questions. With over ten million daily active users, Slack is one of the most popular chat platforms, hosting many active channels focused on software development technologies, e.g., python, react. Prior studies have shown that public Slack chat transcripts contain valuable information, which could provide support for improving automatic software maintenance tools or help researchers understand developer struggles or concerns. Preetha Chatterjee, Kostadin Damevski, Nicholas A. Kraft, Lori L. Pollock |
MSR | 4 |
| 2020 | Finding help with programming errors: An exploratory study of novice software engineers' focus in stack overflow posts
Preetha Chatterjee, Minji Kong, Lori L. Pollock |
J. Syst. Softw. | 3 |
| 2019 | Exploring word embedding techniques to improve sentiment analysis of software engineering textsabstractSentiment analysis (SA) of text-based software artifacts is increasingly used to extract information for various tasks including providing code suggestions, improving development team productivity, giving recommendations of software packages and libraries, and recommending comments on defects in source code, code quality, possibilities for improvement of applications. Studies of state-of-the-art sentiment analysis tools applied to software-related texts have shown varying results based on the techniques and training approaches. In this paper, we investigate the impact of two potential opportunities to improve the training for sentiment analysis of SE artifacts in the context of the use of neural networks customized using the Stack Overflow data developed by Lin et al. We customize the process of sentiment analysis to the software domain, using software domain-specific word embeddings learned from Stack Overflow (SO) posts, and study the impact of software domain-specific word embeddings on the performance of the sentiment analysis tool, as compared to generic word embeddings learned from Google News. We find that the word embeddings learned from the Google News data performs mostly similar and in some cases better than the word embeddings learned from SO posts. We also study the impact of two machine learning techniques, oversampling and undersampling of data, on the training of a sentiment classifier for handling small SE datasets with a skewed distribution. We find that oversampling alone, as well as the combination of oversampling and undersampling together, helps in improving the performance of a sentiment classifier. Eeshita Biswas, K. Vijay-Shanker, Lori L. Pollock |
MSR | 3 |
| 2019 | Exploratory study of slack Q&A chats as a mining source for software engineering toolsabstractModern software development communities are increasingly social. Popular chat platforms such as Slack host public chat communities that focus on specific development topics such as Python or Ruby-on-Rails. Conversations in these public chats often follow a Q&A format, with someone seeking information and others providing answers in chat form. In this paper, we describe an exploratory study into the potential use-fulness and challenges of mining developer Q&A conversations for supporting software maintenance and evolution tools. We designed the study to investigate the availability of information that has been successfully mined from other developer communications, particularly Stack Overflow. We also analyze characteristics of chat conversations that might inhibit accurate automated analysis. Our results indicate the prevalence of useful information, including API mentions and code snippets with descriptions, and several hurdles that need to be overcome to automate mining that information. Preetha Chatterjee, Kostadin Damevski, Lori L. Pollock, Vinay Augustine, Nicholas A. Kraft |
MSR | 3 |
| 2019 | A Collaborative Practicum Targeting Communication Skills for Computer Science ResearchersabstractComputer science researchers spend significant time writing and presenting their research ideas to conference audiences, reviewers, funding agencies, collaborators, and other technical and general audiences. Good communication skills can increase a researcher's effectiveness and efficiency, while currently most students rely on their PhD mentor, technical writing courses, tutors, or self-teaching to improve their communication skills. This paper details a graduate level course to provide graduate students with a collaborative practicum in building strong communication skills for a technical research career in computer science. Each class meeting is activity-based, using collaborative small-group analysis of writing and presentation samples and peer reviewing with rubrics created by the students and refined by the instructor. Assignments require students to stretch their thinking and practice their communication skills. Students write weekly reflect blog entries guided by prompts to promote reflection about their communication experiences, challenges and successes. We summarize and reflect on initial outcomes from two instantiations of the course. Lori L. Pollock |
SIGCSE | 1 |
| 2019 | Infusing Computational Thinking Across Disciplines: Reflections & Lessons LearnedabstractIn this work, we describe our effort to develop, pilot, and evaluate a model for infusing computational thinking into undergraduate curricula across a variety of disciplines using multiple methods that previously have been individually tried and tested, including: (1) multiple pathways of computational thinking, (2) faculty professional development, (3) undergraduate peer mentors, and (4) formative assessment. We present pilot instantiations of computational thinking integration in three different disciplines including sociology, mathematics and music. We also present our professional development approach, which is based on faculty support rather than a co-teaching model. Further, we discuss formative assessment during the pilot implementation, including data focusing on undergraduate students' understanding and dispositions towards computational thinking. Finally, we reflect on what worked, what did not work and why, and identify lessons learned. Our work is relevant to higher education institutions across the nation interested in preparing students who can utilize computational principles to address discipline-specific problems. Lori L. Pollock, Chrystalla Mouza, Kevin R. Guidry, Kathleen L. Pusecker |
SIGCSE | 1 |
| 2019 | A statistics-based performance testing methodology for cloud applicationsabstractThe low cost of resource ownership and flexibility have led users to increasingly port their applications to the clouds. To fully realize the cost benefits of cloud services, users usually need to reliably know the execution performance of their applications. However, due to the random performance fluctuations experienced by cloud applications, the black box nature of public clouds and the cloud usage costs, testing on clouds to acquire accurate performance results is extremely difficult. In this paper, we present a novel cloud performance testing methodology called PT4Cloud. By employing non-parametric statistical approaches of likelihood theory and the bootstrap method, PT4Cloud provides reliable stop conditions to obtain highly accurate performance distributions with confidence bands. These statistical approaches also allow users to specify intuitive accuracy goals and easily trade between accuracy and testing cost. We evaluated PT4Cloud with 33 benchmark configurations on Amazon Web Service and Chameleon clouds. When compared with performance data obtained from extensive performance tests, PT4Cloud provides testing results with 95.4% accuracy on average while reducing the number of test runs by 62%. We also propose two test execution reduction techniques for PT4Cloud, which can reduce the number of test runs by 90.1% while retaining an average accuracy of 91%. We compared our technique to three other techniques and found that our results are much more accurate. Sen He 0002, Glenna Manns, John Saunders, Wei Wang 0054, Lori L. Pollock, Mary Lou Soffa |
ESEC/SIGSOFT FSE | 5 |
| 2019 | Supporting software evolution through feedback on executing/skipping energy tests for proposed source code changesabstractAbstract With the increasing use of battery‐powered devices comes the need to test mobile applications for energy consumption and energy issues. Unfortunately, energy testing is expensive because it is a manual, labor‐intensive process that often requires multiple, separate, energy‐measuring devices to collect energy usage data. The high costs of energy testing can negatively affect the planning process of application evolution. For example, developers might be limited in the number of changes they can include in a release because they must conservatively plan to conduct energy testing after each change. In this paper, we present a new approach to provide developers with feedback on executing/skipping energy tests for proposed code changes. Our technique leverages change impact analysis and precomputed API energy usage information. More specifically, for a proposed change, the technique predicts whether energy testing will be required, and if so, which energy tests will need to be run. Such information may allow developers to avoid spending unnecessary time for energy testing and develop an effective application evolution timeline. To investigate the feasibility of our technique, we implemented a prototype for Android applications and conducted three case studies at different granularity levels on 10 Android applications. Cagri Sahin, Lori L. Pollock, James Clause |
J. Softw. Evol. Process. | 2 |
| 2018 | Predicting future developer behavior in the IDE using topic modelsabstractInteraction data, gathered from developers' daily clicks and key presses in the IDE, has found use in both empirical studies and in recommendation systems for software engineering. We observe that this data has several characteristics, common across IDEs: Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Nicholas A. Kraft, Lori L. Pollock |
ICSE | 5 |
| 2018 | Testing Cloud Applications under Cloud-Uncertainty Performance EffectsabstractThe paradigm shift of deploying applications to the cloud has introduced both opportunities and challenges. Although clouds use elasticity to scale resource usage at runtime to help meet an application's performance requirements, developers are still challenged by unpredictable performance, little control of execution environment, and differences among cloud service providers, all while being charged for their cloud usages. Application performance stability is particularly affected by multi-tenancy in which the hardware is shared among varying applications and virtual machines. Developers porting their applications need to meet performance requirements, but testing on the cloud under the effects of performance uncertainty is difficult and expensive, due to high cloud usage costs. This paper presents a first approach to testing an application with typical inputs for how its performance will be affected by performance uncertainty, without incurring undue costs of brute force testing in the cloud. We specify cloud uncertainty testing criteria, design a test-based strategy to characterize the black box cloud's performance distributions using these testing criteria, and support execution of tests to characterize the resource usage and cloud baseline performance of the application to be deployed. Importantly, we developed a smart test oracle that estimates the application's performance with certain confidence levels using the above characterization test results and determines whether it will meet its performance requirements. We evaluated our testing approach on both the Chameleon cloud and Amazon web services; results indicate that this testing strategy shows promise as a cost-effective approach to test for performance effects of cloud uncertainty when porting an application to the cloud. Wei Wang 0054, Ningjing Tian, Sunzhou Huang, Sen He 0002, Abhijeet Srivastava, Mary Lou Soffa, Lori L. Pollock |
ICST | 7 |
| 2018 | A Computer Science Study Abroad with Service Learning: Design and ReflectionsabstractStudy abroad offers students the opportunity to experience other cultures, languages, and environments while obtaining credits toward their degree. Students are also taught to appreciate the diversity of people and culture, such that they may dismiss stereotypes and learn to communicate and collaborate cross-culturally in a global economy. Unfortunately, few universities offer study abroad programs directed specifically to computer science and particularly in combining student technical learning with service learning for broadening participation in computing throughout the world. In this paper, we describe a service-learning-based model for computer science students and other university students with minimal prior computer science experience to engage and inspire themselves and the next generation of computational thinkers through learning, teaching and creating web-based learning games along with local children and teachers in a foreign country. We describe the model focusing on learning objectives, curriculum, field component, planning, and partnership building. We describe the products that undergraduates were able to create in four weeks and their CS education service learning field experiences. Finally, we investigate the impact of the study abroad model on undergraduates' content knowledge, and their career and personal development. Lori L. Pollock, James Atlas, Timothy C. Bell, Tracy Henderson |
SIGCSE | 1 |
| 2018 | Customizing a Field Experience for CS Undergrads in Teaching Computer Science for Your School Context: (Abstract Only)abstractThis workshop's goal is to help faculty who want to establish a course (or alternate vehicle) for mentoring undergraduates with some CS background to participate in K-12 teaching CS in local schools with engaging pedagogy. The workshop leverages the experiences and lessons learned from ten semesters of the organizers leading a course that meets once a week on campus for mentoring to support the undergraduates' field experience in local schools and libraries. The workshop will dive deep into logistics including how to establish and maintain strong teacher partnerships, establishing student-teacher matches and weekly field experience schedules, weekly in-class activities and assignments to support the field experience, weekly student reflective journal prompts, and surveys for formative evaluation. Participants will actively reflect on their own contexts with potential opportunities and challenges, and organizers will facilitate small group discussions of how to address the challenges, different models for different contexts, and how to get started. Participants should leave with a plan for next steps toward offering a mentored undergraduate field experience in teaching computer science and access to a community of faculty who are working to help to broaden participation in computer science in K-12 while providing opportunities for undergraduates to hone their communication and leadership skills, increase their self confidence, and participate in giving back to the community using their technical skills. The activities do not require a laptop, only pens and handouts provided by the organizers. Lori L. Pollock, Terry Harvey, James Atlas, Chrystalla Mouza |
SIGCSE | 1 |
| 2018 | Exploring Evolutionary Search Strategies to Improve Applications' Energy EfficiencyabstractEnergy consumption have become an important non-functional requirement for applications running on battery powered devices through data centers. Despite the increased interest on detecting and understanding what causes an application to be energy inefficient, few works focus on helping developers to automatically make their applications more energy efficient based on developers’ design and implementation decisions. This paper explores how search strategies based on genetic algorithms can help developers automatically find an energy efficient version of an application based on transformations corresponding to developers’ high level decisions (e.g., selecting API implementations). Our results show how different search strategies can help to improve the energy efficiency for nine Java applications. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Irene Manotas, James Clause, Lori L. Pollock |
SSBSE | 3 |
| 2018 | Predicting Future Developer Behavior in the IDE Using Topic ModelsabstractWhile early software command recommender systems drew negative user reaction, recent studies show that users of unusually complex applications will accept and utilize command recommendations. Given this new interest, more than a decade after first attempts, both the recommendation generation (backend) and the user experience (frontend) should be revisited. In this work, we focus on recommendation generation. One shortcoming of existing command recommenders is that algorithms focus primarily on mirroring the short-term past,-i.e., assuming that a developer who is currently debugging will continue to debug endlessly. We propose an approach to improve on the state of the art by modeling future task context to make better recommendations to developers. That is, the approach can predict that a developer who is currently debugging may continue to debug OR may edit their program. To predict future development commands, we applied Temporal Latent Dirichlet Allocation, a topic model used primarily for natural language, to software development interaction data (i.e., command streams). We evaluated this approach on two large interaction datasets for two different IDEs, Microsoft Visual Studio and ABB Robot Studio. Our evaluation shows that this is a promising approach for both predicting future IDE commands and producing empirically-interpretable observations. Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Nicholas A. Kraft, Lori L. Pollock |
IEEE Trans. Software Eng. | 5 |
| 2017 | Behavior Metrics for Prioritizing Investigations of ExceptionsabstractMany software development teams collect product defect reports, which can either be manually submitted or automatically created from product logs. Periodically, the teams use the collected defect reports to prioritize which defect to address next. We present a set of behavior-based metrics that can be used in this process. These metrics are based on the insight that development teams can estimate user inconvenience from user and application behavior in interaction logs. To estimate user inconvenience, the behavior metrics capture important user and application behavior after exceptions (the defects of interest in our case). We validated these metrics through a survey of how developers would incorporate the behavior metrics into their prioritization decisions. We found that developers change their priority of investigating an exception about 31% of the time after including the behavior metrics in the priority decision. These findings provide evidence that behavior metrics provide a promising advance towards prioritizing application exceptions. Zack Coker, Kostadin Damevski, Claire Le Goues, Nicholas A. Kraft, David C. Shepherd, Lori L. Pollock |
ICSME | 6 |
| 2017 | Extracting code segments and their descriptions from research articlesabstractThe availability of large corpora of online software-related documents today presents an opportunity to use machine learning to improve integrated development environments by first automatically collecting code examples along with associated descriptions. Digital libraries of computer science research and education conference and journal articles can be a rich source for code examples that are used to motivate or explain particular concepts or issues. Because they are used as examples in an article, these code examples are accompanied by descriptions of their functionality, properties, or other associated information expressed in natural language text. Identifying code segments in these documents is relatively straightforward, thus this paper tackles the problem of extracting the natural language text that is associated with each code segment in an article. We present and evaluate a set of heuristics that address the challenges of the text often not being colocated with the code segment as in developer communications such as online forums. Preetha Chatterjee, Benjamin Gause, Hunter Hedinger, Lori L. Pollock |
MSR | 4 |
| 2017 | From Professional Development to the Classroom: Findings from CS K-12 TeachersabstractThe CS for All initiative places increased emphasis on the need to prepare K-12 teachers of computer science (CS). Professional development (PD) programs continue to be an essential mechanism for preparing in-service teachers who have little formal background in CS content, skills, and teaching pedagogy. While increased investment by federal agencies and the industry has raised the number of CS PD opportunities for K-12 teachers, there has been limited study of how teachers apply what they learn back in their classroom. This paper describes an in-depth qualitative study through interviews of 28 elementary, middle and high school teachers who participated in summer PD in preparation of teaching a full CS course or integrate CS modules into existing courses (e.g., science, engineering, business, technology, etc). The interview protocol focused on educators' involvement in the PD, specific skills and strategies they learned, whether and how they have been able to apply these new skills in the classroom, what facilitated or impeded this application, and how students have responded. Lori L. Pollock, Chrystalla Mouza, Amanda Czik, Alexis Little, Debra Coffey, Joan Buttram |
SIGCSE | 1 |
| 2017 | What information about code snippets is available in different software-related documents? An exploratory studyabstractA large corpora of software-related documents is available on the Web, and these documents offer the unique opportunity to learn from what developers are saying or asking about the code snippets that they are discussing. For example, the natural language in a bug report provides information about what is not functioning properly in a particular code snippet. Previous research has mined information about code snippets from bug reports, emails, and Q&A forums. This paper describes an exploratory study into the kinds of information that is embedded in different software-related documents. The goal of the study is to gain insight into the potential value and difficulty of mining the natural language text associated with the code snippets found in a variety of software-related documents, including blog posts, API documentation, code reviews, and public chats. Preetha Chatterjee, Manziba Akanda Nishi, Kostadin Damevski, Vinay Augustine, Lori L. Pollock, Nicholas A. Kraft |
SANER | 5 |
| 2017 | Automatically generating natural language descriptions for object-related statement sequencesabstractCurrent source code analyses driving software maintenance tools treat methods as either a single unit or a set of individual statements or words. They often leverage method names and any existing internal comments. However, internal comments are rare, and method names do not typically capture the method's multiple high-level algorithmic steps that are too small to be a single method, but require more than one statement to implement. Previous work demonstrated feasibility of identifying high level actions automatically for loops; however, many high level actions remain unaddressed and undocumented, particularly sequences of consecutive statements that are associated with each other primarily by object references. We call these object-related action units. In this paper, we present an approach to automatically generate natural language descriptions of object-related action units within methods. We leverage the available, large source of high-quality open source projects to learn the templates of object-related actions, identify the statement that can represent the main action, and generate natural language descriptions for these actions. Our evaluation study of a set of 100 object-related statement sequences showed promise of our approach to automatically identify the action and arguments and generate natural language descriptions. Lori L. Pollock, K. Vijay-Shanker |
SANER | 2 |
| 2017 | Mining Sequences of Developer Interactions in Visual Studio for Usage SmellsabstractIn this paper, we present a semi-automatic approach for mining a large-scale dataset of IDE interactions to extract usage smells, i.e., inefficient IDE usage patterns exhibited by developers in the field. The approach outlined in this paper first mines frequent IDE usage patterns, filtered via a set of thresholds and by the authors, that are subsequently supported (or disputed) using a developer survey, in order to form usage smells. In contrast with conventional mining of IDE usage data, our approach identifies time-ordered sequences of developer actions that are exhibited by many developers in the field. This pattern mining workflow is resilient to the ample noise present in IDE datasets due to the mix of actions and events that these datasets typically contain. We identify usage patterns and smells that contribute to the understanding of the usability of Visual Studio for debugging, code search, and active file navigation, and, more broadly, to the understanding of developer behavior during these software development activities. Among our findings is the discovery that developers are reluctant to use conditional breakpoints when debugging, due to perceived IDE performance problems as well as due to the lack of error checking in specifying the conditional. Kostadin Damevski, David C. Shepherd, Johannes Schneider 0002, Lori L. Pollock |
IEEE Trans. Software Eng. | 4 |
| 2016 | An empirical study of practitioners' perspectives on green software engineeringabstractThe energy consumption of software is an increasing concern as the use of mobile applications, embedded systems, and data center-based services expands. While research in green software engineering is correspondingly increasing, little is known about the current practices and perspectives of software engineers in the field. This paper describes the first empirical study of how practitioners think about energy when they write requirements, design, construct, test, and maintain their software. We report findings from a quantitative, targeted survey of 464 practitioners from ABB, Google, IBM, and Microsoft, which was motivated by and supported with qualitative data from 18 in-depth interviews with Microsoft employees. The major findings and implications from the collected data contextualize existing green software engineering research and suggest directions for researchers aiming to develop strategies and tools to help practitioners improve the energy usage of their applications. Irene Manotas, Christian Bird, David C. Shepherd, Ciera Jaspan, Caitlin Sadowski, Lori L. Pollock, James Clause |
ICSE | 7 |
| 2016 | A case study of program comprehension effort and technical debt estimationsabstractThis paper describes a case study of using developer activity logs as indicators of a program comprehension effort by analyzing temporal sequences of developer actions (e.g., navigation and edit actions). We analyze developer activity data spanning 109,065 events and 69 hours of work on a medium-sized industrial application. We examine potential correlations between different measures of developer activity, code change metrics and code smells to gain insight into questions that could direct future technical debt interest estimation. To gain more insights into the data, we follow our analysis with commit message analysis and a developer interview. Our results indicate that developer activity as an estimate of program comprehension effort is correlated with both change proneness and static metrics for code smells. Vallary Singh, Lori L. Pollock, Will Snipes, Nicholas A. Kraft |
ICPC | 2 |
| 2016 | Interactive exploration of developer interaction traces using a hidden markov modelabstractUsing IDE usage data to analyze the behavior of software developers in the field, during the course of their daily work, can lend support to (or dispute) laboratory studies of developers. This paper describes a technique that leverages Hidden Markov Models (HMMs) as a means of mining high-level developer behavior from low-level IDE interaction traces of many developers in the field. HMMs use dual stochastic processes to model higher-level hidden behavior using observable input sequences of events. We propose an interactive approach of mining interpretable HMMs, based on guiding a human expert in building a high quality HMM in an iterative, one state at a time, manner. The final result is a model that is both representative of the field data and captures the field phenomena of interest. We apply our HMM construction approach to study debugging behavior, using a large IDE interaction dataset collected from nearly 200 developers at ABB, Inc. Our results highlight the different modes and constituent actions in debugging, exhibited by the developers in our dataset. Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Lori L. Pollock |
MSR | 4 |
| 2016 | Implementation and Outcomes of a Three-Pronged Approach to Professional Development for CS PrinciplesabstractOne of the greatest challenges in broadening participation in computer science is teacher preparation, as few middle and high school teachers have a formal background in computing. Further, without a credentialing program, there are limited ways to learn content and pedagogical strategies for effective computer science instruction. As a result, professional development is key to successful reform in the teaching of computer science. In this paper, we describe our three-pronged approach to the design of a professional development model for middle and high school teachers interested in implementing the Computer Science Principles (CSP) curriculum in their classrooms or infusing CSP modules into STEM curricula. We describe our model focusing on content, pedagogical strategies and follow-up classroom support during the academic year. We subsequently report on participating teacher outcomes, in terms of self-rated understandings, attitudes and implementation practices. We share lessons learned and offer recommendations for professional development designers. Chrystalla Mouza, Lori L. Pollock, Kathleen L. Pusecker, Kevin R. Guidry, Ching-Yi Yeh, James Atlas, Terry Harvey |
SIGCSE | 2 |
| 2016 | A field study of how developers locate features in source code
Kostadin Damevski, David C. Shepherd, Lori L. Pollock |
Empir. Softw. Eng. | 3 |
| 2016 | From benchmarks to real apps: Exploring the energy impacts of performance-directed changes
Cagri Sahin, Lori L. Pollock, James Clause |
J. Syst. Softw. | 2 |
| 2016 | Introduction to the special issue on software maintenance and evolutionabstractIt is our pleasure to introduce you to the papers in this Special Issue based on the 30th International Conference on Software Maintenance and Evolution (ICSME 2014). ICSME is the premier international venue in software maintenance and evolution, where participants from academia, government, and industry gather to share and discuss their ideas on and experiences with solving critical software maintenance problems. In response to the call for research papers, we received 267 abstracts and 210 full paper submissions. Each submitted paper was reviewed by at least three members of the program committee (PC), who were selected through bidding to create a good match between paper topic and PC member expertise; each PC member reviewed 9–10 papers over several weeks, with collectively 632 reviews submitted. During the week-long discussion period, reviewers submitted over 1000 comments. In the end, 40 high-quality papers were accepted for publication in the conference proceedings, yielding an acceptance rate of 19%. The accepted papers covered a broad range of topics in software maintenance and evolution, including developer knowledge, evolving systems, developer support, technical debt, managing change, empirical studies, fault localization, software quality, patches, recommender systems, and software clones. Guided by the reviews and discussions, ICSME 2014 Program Co-Chairs Leon Moonen and Lori Pollock carefully selected nine outstanding papers of the 40 accepted for the conference and invited their authors to submit a significantly extended version of their conference paper to this special issue in the Journal of Software: Evolution and Process (JSEP). Six of the nine invited papers were extended by their authors and subjected to the rigorous JSEP reviewing process, thus undergoing additional rounds of reviews and revisions. Eventually, the following four papers successfully completed the review process and are contained in this special issue. The paper ‘A Simple, Efficient, Context Sensitive Approach for Code Completion’ by Muhammad Asaduzzaman, Chanchal K. Roy, Kevin A. Schneider, and Daqing Hou describes a technique for method call completion that uses the type name and context to search for method calls whose contexts match with that of the receiver object. A database of context–method pairs is created by collecting code examples from repositories. The proposed approach was shown to either outperform or perform as well as state-of-the-art techniques for code completion based on statistical language models. The paper ‘An Empirical Study on How Expert Knowledge Affects Bug Reports’ by Da Huo, Tao Ding, Collin McMillan, and Malcom Gethers describes an empirical study of the textual difference between bug reports written by experts and non-experts. The study showed that experts and non-experts wrote bug reports differently. The findings support the hypothesis that expert knowledge affects the way in which people write bug reports. The paper ‘How Does Code Obfuscation Impact Energy Usage?’ by Cagri Sahin, Philip Tornquist, Ryan Mckenna, Zachary Pearson, and James Clause describes an empirical study into the energy impacts of applying different code obfuscations. In addition to investigating how different obfuscations in four obfuscation tools alter the overall energy usage of an application, the paper also studies whether the impacts of obfuscations are likely to be meaningful for mobile application users. The results support the notion that developers can protect their applications without impacting the battery life of the devices where their applications execute. The paper ‘Empirical Analysis of the Relationship between CC and SLOC in a Large Corpus of Java Methods and C Functions’ by Davy Landman, Alexander Serebrenik, and Jurgen Vinju describes an extensive literature study of the Cyclomatic Complexity (CC) and Source Lines of Code (SLOC) correlation results, followed by a correlation study of CC/SLOC on large Java and C corpora. In contrast to the majority of the previous studies, this study did not observe a strong linear correlation between CC and SLOC of Java methods and C functions. We hope that readers will enjoy this special issue and gain useful insights from the four papers presented. We would like to thank all the authors who submitted papers to the conference and to this special issue. In addition, we would like to thank the members of the ICSME 2014 program committee and the external reviewers for their time, careful reviews, and active discussions of the submitted papers, which helped make this special issue special. This kind of service is important to the health of the community and the quality of its publications. Finally, we would like to thank the editorial board of the Journal of Software: Evolution and Process and the publisher Wiley for providing us with the opportunity to devote this issue to the best of ICSME 2014. We also thank JSEP Editor Gerardo Canfora for providing expert guidance and important advice throughout the process. Enjoy! Leon Moonen, Lori L. Pollock |
J. Softw. Evol. Process. | 2 |
| 2015 | How and When to Transfer Software Engineering Research via ExtensionsabstractIt is often reported that there is a large gap between software engineering research and practice, with little transfer from research to practice. While this is true in general, one transfer technique is increasingly breaking down this barrier: extensions to integrated development environments (IDEs). With the proliferation of app stores for IDEs and increasing transfer effort from researchers several research-based extensions have seen significant adoption. In this talk we'll discuss our experience transferring code search research, which currently is in the top 5% of Visual Studio extensions with over 13,000 downloads, as well as other research techniques transferred via extensions such as NCrunch, FindBugs, Code Recommenders, Mylyn, and Instasearch. We'll use the lessons learned from our transfer experience to provide case study evidence as to best practices for successful transfer, supplementing it with the quantitative evidence offered by app store and usage data across the broader set of extensions. The goal of this 30 minute talk is to provide researchers with a realistic view on which research techniques can be transferred to practice as well as concrete steps to execute such a transfer. David C. Shepherd, Kostadin Damevski, Lori L. Pollock |
ICSE (2) | 3 |
| 2015 | Developing a model of loop actions by mining loop characteristics from a large code corpusabstractSome high level algorithmic steps require more than one statement to implement, but are not large enough to be a method on their own. Specifically, many algorithmic steps (e.g., count, compare pairs of elements, find the maximum) are implemented as loop structures, which lack the higher level abstraction of the action being performed, and can negatively affect both human readers and automatic tools. Additionally, in a study of 14,317 projects, we found that less than 20% of loops are documented to help readers. In this paper, we present a novel automatic approach to identify the high level action implemented by a given loop. We leverage the available, large source of high-quality open source projects to mine loop characteristics and develop an action identification model. We use the model and feature vectors extracted from loop code to automatically identify the high level actions implemented by loops. We have evaluated the accuracy of the loop action identification and coverage of the model over 7159 open source programs. The results show great promise for this approach to automatically insert internal comments and provide additional higher level naming for loop actions to be used by tools such as code search. Lori L. Pollock, K. Vijay-Shanker |
ICSME | 2 |
| 2015 | Exploring the use of concern element role information in feature location evaluationabstractBefore making changes, programmers need to locate and understand source code that corresponds to specific functionality, i.e., Perform concern or feature location. Numerous concern and feature location techniques have been proposed, but to the best of our knowledge, no existing techniques or evaluations report information on what role a code element plays in the larger concern. In this paper, we report on two case studies that investigate two hypotheses on how evaluation studies of concern location techniques can be strengthened by utilizing concern role information: (1) by increasing agreement among human annotators for gold set establishment and (2) by providing richer information about the elements ranked as relevant by concern location techniques, which could help further improve the tools. We conducted a case study of 6 Java developers annotating 3 concerns with role information. When the developers understood the task description, pair wise agreement increased by 20%, 25%, and 135% for the 3 concerns over a prior concern location study without role information. Our findings also suggest that there may be core element roles that need to be annotated by humans, but that the remaining roles may be automatically derived, which could facilitate more reliable concern location benchmarks in the future. We also conducted an exploratory study of the element roles represented in results returned by a state of the art feature location tool. The results of these two studies suggest that integrating concern element role information into evaluations can help to strengthen both the gold set establishment and the analysis of results returned by various tools. Emily Hill 0001, David C. Shepherd, Lori L. Pollock |
ICPC | 3 |
| 2015 | Field Experiences in Teaching Computer Science: Course Organization and ReflectionsabstractA major challenge for broadening participation in computing within K-12 settings is the lack of trained teachers. While professional development programs provide opportunities for the development of knowledge, skills, and pedagogy in teaching computing, teachers need ongoing support throughout the academic year. In this paper, we describe a course-based model for partnering undergraduates with teachers and students in a field experience model. We describe the model focusing on learning objectives, curriculum, field component and partnership building. We subsequently report on the products that undergraduates were able to create with their partner teachers. Finally, we investigate the impact of the field experience model on undergraduates' content knowledge, pedagogical skills and career development. Lori L. Pollock, Chrystalla Mouza, James Atlas, Terry Harvey |
SIGCSE | 1 |
| 2015 | Scaling up evaluation of code search tools through developer usage metricsabstractCode search is a fundamental part of program understanding and software maintenance and thus researchers have developed many techniques to improve its performance, such as corpora preprocessing and query reformulation. Unfortunately, to date, evaluations of code search techniques have largely been in lab settings, while scaling and transitioning to effective practical use demands more empirical feedback from the field. This paper addresses that need by studying metrics based on automatically-gathered anonymous field data from code searches to infer user satisfaction. We describe techniques for addressing important concerns, such as how privacy is retained and how the overhead on the interactive system is minimized. We perform controlled user and field studies which identify metrics that correlate with user satisfaction, enabling the future evaluation of search tools through anonymous usage data. In comparing our metrics to similar metrics used in Internet search we observe differences in the relationship of some of the metrics to user satisfaction. As we further explore the data, we also present a predictive multi-metric model that achieves accuracy of over 70% in determining query satisfaction. Kostadin Damevski, David C. Shepherd, Lori L. Pollock |
SANER | 3 |
| 2014 | How do code refactorings affect energy usage?abstractContext: Code refactoring's benefits to understandability, maintainability and extensibility are well known enough that automated support for refactoring is now common in IDEs. However, the decision to apply such transformations is currently performed without regard to the impacts of the refactorings on energy consumption. This is primarily due to a lack of information and tools to provide such relevant information to developers. Unfortunately, concerns about energy efficiency are rapidly becoming a high priority concern in many environments, including embedded systems, laptops, mobile devices, and data centers. Cagri Sahin, Lori L. Pollock, James Clause |
ESEM | 2 |
| 2014 | SEEDS: a software engineer's energy-optimization decision support frameworkabstractReducing the energy usage of software is becoming more important in many environments, in particular, battery-powered mobile devices, embedded systems and data centers. Recent empirical studies indicate that software engineers can support the goal of reducing energy usage by making design and implementation decisions in ways that take into consideration how such decisions impact the energy usage of an application. However, the large number of possible choices and the lack of feedback and information available to software engineers necessitates some form of automated decision-making support. This paper describes the first known automated support for systematically optimizing the energy usage of applications by making code-level changes. It is effective at reducing energy usage while freeing developers from needing to deal with the low-level, tedious tasks of applying changes and monitoring the resulting impacts to the energy usage of their application. We present a general framework, SEEDS, as well as an instantiation of the framework that automatically optimizes Java applications by selecting the most energy-efficient library implementations for Java's Collections API. Our empirical evaluation of the framework and instantiation show that it is possible to improve the energy usage of an application in a fully automated manner for a reasonable cost. Irene Lizeth Manotas Gutiérrez, Lori L. Pollock, James Clause |
ICSE | 2 |
| 2014 | A comparison of two hands-on laboratory experiences in computers, networks and cyber security for 10th-12th graders (abstract only)abstractIn this poster, we describe our experience of designing and executing two different weeklong programs for 10th -- 12th grade students. The goal of our program is to attract students to the field of computing, increase their computing confidence and familiarize them with ways that computing impacts our community. Student groups consist of 32 students for each week with a 1:0.8 male-to-female ratio. No prerequisite knowledge is required to attend. We compare the different facets of the curriculums by evaluating the impact of each week both quantitatively and qualitatively. Our evaluation implies the attraction of cyber security topics to this age group, particularly male students, and presents a curriculum that may help increase confidence in computing concepts, particularly for female students. We solicit feedback and welcome input on our curriculum and evaluation method. Lisa M. Marvel, Stephen Raio, Lori L. Pollock, David Arty, Gerard Chaney, Giorgio Bertoli, Christopher Paprcka, Wendy Choi, Erica Bertoli, Sandra K. Young |
SIGCSE | 3 |
| 2014 | An empirical study of identifier splitting techniques
Emily Hill 0001, Dave W. Binkley, Dawn J. Lawrie, Lori L. Pollock, K. Vijay-Shanker |
Empir. Softw. Eng. | 4 |
| 2014 | Automatic Segmentation of Method Code into Meaningful Blocks: Design and EvaluationabstractSUMMARY Good programming practice and guidelines suggest that programmers use both vertical and horizontal spacing to visibly delineate between code segments that represent different algorithmic steps or high‐level actions. Unfortunately, programmers do not always follow these guidelines. Editors and integrated development environments (IDEs) can easily indent codes based on syntax, but they do not currently support automatic blank line insertion, which presents more significant challenges involving the semantics. This paper presents and evaluates a heuristic solution to the automatic blank line insertion problem by leveraging both program structure and naming information to identify ‘meaningful blocks’, consecutive statements that logically implement a high‐level action. Our tool, SEGMENT, takes as input a Java method and outputs a segmented version that separates meaningful blocks by vertical spacing. We report on several studies involving human judgments to evaluate the effectiveness of the automatic blank line insertion algorithm, for different size methods and for different levels of programmer expertise. The results indicate strong positive overall opinion of SEGMENT's effectiveness in comparison with both developer‐written blank lines and blank lines inserted by newcomers to the code. The results vary only slightly among short and long methods, and among novice and advanced programmers. SEGMENT assists in making users obtain an overall picture of a method's actions and comprehend it quicker as well as provides hints for internal documentation placement. Copyright © 2013 John Wiley & Sons, Ltd. Lori L. Pollock, K. Vijay-Shanker |
J. Softw. Evol. Process. | 2 |
| 2013 | 1st international workshop on natural language analysis in software engineering (NaturaLiSE 2013)abstractSoftware engineers produce code that has formal syntax and semantics, which establishes its formal meaning. However, the code also includes significant natural language found primarily in identifier names and comments. Furthermore, the code is surrounded by non-source artifacts, predominantly written in natural language. The NaturaLiSE workshop focuses on natural language analysis of software. The workshop brings together researchers and practitioners interested in exploiting natural language information to create improved software engineering tools. Participants will explore natural language analysis applied to software artifacts, combining natural language and traditional program analysis, integration of natural language analyses into client tools, mining natural language data, and empirical studies focused on evaluating the usefulness of natural language analysis. Lori L. Pollock, Dave W. Binkley, Dawn J. Lawrie, Emily Hill 0001, Rocco Oliveto, Gabriele Bavota, Alberto Bacchelli |
ICSE | 1 |
| 2013 | Differentiating Roles of Program Elements in Action-Oriented ConcernsabstractMany techniques have been developed to help programmers locate source code that corresponds to specific functionality, i.e., concern or feature location, as it is a frequent software maintenance activity. This paper proposes operational definitions for differentiating the roles that each program element of a concern plays with respect to the concern's implementation. By identifying the respective roles, we enable evaluations that provide more insight into comparative performance of concern location techniques. To provide definitions that are specific enough to be useful in practice, we focus on the subset of concerns that are action-oriented. We also conducted a case study that compares concern mappings derived from our role definitions with three developers' mappings across three concerns. The results suggest that our definitions capture the majority of developer-identified elements and that control-flow islands (i.e., groups of elements with little to no control flow connections) can cause developers to omit relevant elements. Emily Hill 0001, David C. Shepherd, Lori L. Pollock, K. Vijay-Shanker |
ICSM | 3 |
| 2013 | Part-of-speech tagging of program identifiers for improved text-based software engineering toolsabstractTo aid program comprehension, programmers choose identifiers for methods, classes, fields and other program elements primarily by following naming conventions in software. These software “naming conventions” follow systematic patterns which can convey deep natural language clues that can be leveraged by software engineering tools. For example, they can be used to increase the accuracy of software search tools, improve the ability of program navigation tools to recommend related methods, and raise the accuracy of other program analyses. After splitting multi-word names into their component words, the next step to extracting accurate natural language information is tagging each word with its part of speech (POS) and then chunking the name into natural language phrases. State-of-theart approaches, most of which rely on “traditional POS taggers” trained on natural language documents, do not capture the syntactic structure of program elements. In this paper, we present a POS tagger and syntactic chunker for source code names that takes into account programmers' naming conventions to understand the regular, systematic ways a program element is named. We studied the naming conventions used in Object Oriented Programming and identified different grammatical constructions that characterize a large number of program identifiers. This study then informed the design of our POS tagger and chunker. Our evaluation results show a significant improvement in accuracy(11%-20%) of POS tagging of identifiers, over the current approaches. With this improved accuracy, both automated software engineering tools and developers will be able to better capture and understand the information available in code. Samir Gupta, Sana Malik, Lori L. Pollock, K. Vijay-Shanker |
ICPC | 3 |
| 2013 | Automatic generation of natural language summaries for Java classesabstractMost software engineering tasks require developers to understand parts of the source code. When faced with unfamiliar code, developers often rely on (internal or external) documentation to gain an overall understanding of the code and determine whether it is relevant for the current task. Unfortunately, the documentation is often absent or outdated. This paper presents a technique to automatically generate human readable summaries for Java classes, assuming no documentation exists. The summaries allow developers to understand the main goal and structure of the class. The focus of the summaries is on the content and responsibilities of the classes, rather than their relationships with other classes. The summarization tool determines the class and method stereotypes and uses them, in conjunction with heuristics, to select the information to be included in the summaries. Then it generates the summaries using existing lexicalization tools. A group of programmers judged a set of generated summaries for Java classes and determined that they are readable and understandable, they do not include extraneous information, and, in most cases, they are not missing essential information. Laura Moreno, Jairo Aponte, Giriprasad Sridhara, Andrian Marcus, Lori L. Pollock, K. Vijay-Shanker |
ICPC | 5 |
| 2013 | JSummarizer: An automatic generator of natural language summaries for Java classesabstractJSummarizer is an Eclipse plug-in for automatically generating natural language summaries of Java classes. The summary is based on the stereotype of the class, which implicitly encodes the design intent of the class and is automatically inferred by JSummarizer. The tool uses a set of predefined heuristics to determine what information will be reflected in the summary, and it uses natural language processing and generation techniques to form the summary. The generated summaries can be used to re-document the code and to help developers to easier understand large and complex classes. Laura Moreno, Andrian Marcus, Lori L. Pollock, K. Vijay-Shanker |
ICPC | 3 |
| 2013 | A dataset for evaluating identifier splittersabstractSoftware engineering and evolution techniques have recently started to exploit the natural language information in source code. A key step in doing so is splitting identifiers into their constituent words. While simple in concept, identifier splitting raises several challenging issues, leading to a range of splitting techniques. Consequently, the research community would benefit from a dataset (i.e., a gold set) that facilitates comparative studies of identifier splitting techniques. A gold set of 2,663 split identifiers was constructed from 8,522 individual human splitting judgements and can be obtained from www.cs.loyola.edu/~binkley/ludiso. This set's construction and observations aimed at its effective use are described. Dave W. Binkley, Dawn J. Lawrie, Lori L. Pollock, Emily Hill 0001, K. Vijay-Shanker |
MSR | 3 |
| 2013 | Automatically mining software-based, semantically-similar words from comment-code mappingsabstractMany software development and maintenance tools involve matching between natural language words in different software artifacts (e.g., traceability) or between queries submitted by a user and software artifacts (e.g., code search). Because different people likely created the queries and various artifacts, the effectiveness of these tools is often improved by expanding queries and adding related words to textual artifact representations. Synonyms are particularly useful to overcome the mismatch in vocabularies, as well as other word relations that indicate semantic similarity. However, experience shows that many words are semantically similar in computer science situations, but not in typical natural language documents. In this paper, we present an automatic technique to mine semantically similar words, particularly in the software context. We leverage the role of leading comments for methods and programmer conventions in writing them. Our evaluation of our mined related comment-code word mappings that do not already occur in WordNet are indeed viewed as computer science, semantically-similar word pairs in high proportions. Samir Gupta, Lori L. Pollock, K. Vijay-Shanker |
MSR | 3 |
| 2013 | Making the most of undergraduate research (abstract only)abstractInvolving undergraduates in Computer Science research has many benefits. It's an exciting way for students to gain independent problem solving skills. It exposes them to interesting projects and the research process, thereby keeping them in computer science, even encouraging them to go to graduate school. And especially in primarily teaching institutions, it's a rewarding way for faculty to remain engaged in their own research. In this workshop we will (1) present best practices for mentoring undergraduate research, (2) equip participants with resources for mentoring their own students, and (3) further develop (1) and (2) through breakout sessions on concerns of interest to attendees. For more, please see www.cs.williams.edu/~andrea/SIGCSE2013. This workshop is intended for all college level computer science educators. Laptop Optional. Andrea Pohoreckyj Danyluk, Nancy M. Amato, Ran Libeskind-Hadas, Lori L. Pollock, Susan H. Rodger |
SIGCSE | 4 |
| 2013 | Configuring effective navigation models and abstract test cases for web applications by analysing user behaviourabstractSUMMARY As web applications become more complex and are used more pervasively, testing demands are increasing without corresponding automated support. One promising approach to automatic test generation is statistical model‐based testing, where logged user behaviour is used to build a usage‐based model of web application navigation, from which abstract test cases are generated. Executable test cases are then created by adding parameter values to the abstract test cases. Several researchers have proposed variations of this approach; however, no one has empirically examined the tradeoffs and implications of the different ways to represent user behaviour in a navigation model and the characteristics of the test cases automatically generated from different models. This paper reports on our exploratory study of automatically generated abstract test cases and the underlying usage‐based navigation models constructed from over 19,000 user sessions across five publicly deployed web applications. Our results suggest how web testers can easily configure statistical model‐based automatic test case generators for web applications toward generating tests closely related to user behaviour or toward new navigations without using large additional test resources. Copyright © 2013 John Wiley & Sons, Ltd. Sara Sprenkle, Lori L. Pollock, Lucy Simko |
Softw. Test. Verification Reliab. | 2 |
| 2012 | Leveraging natural language analysis of software: Achievements, challenges, and opportunitiesabstractSummary form only given. Studies continue to report that more time is spent reading, locating, and comprehending code than actually writing code. The increasing size and complexity of software systems makes it significantly more challenging for humans to perform maintenance tasks on software without automated and semi-automated tools to support them, especially in the error-prone tasks. Thus, software engineers increasingly rely on software engineering tools to automate maintenance tasks as much as possible. The program analyses that drive today's software engineering tools have historically focused on analyzing the program's data and control flow, dependencies, and other structural information about the program to uncover and prove program properties. Yet, a software system is more than just the source code and its structure. To build effective software tools, the underlying automated analyses need to use all the information available to make the tools as intelligent and useful as possible. By adapting natural language processing (NLP) to source code analysis, and integrating information retrieval (IR), NLP, and traditional program analyses, we can expect significant improvement in automated and semi-automated software engineering tools for many different software engineering tasks. In this talk, I will overview research in text analysis of software and discuss our achievements to date, the challenges faced in text analysis, and the opportunities for text analysis of software in the future. Lori L. Pollock |
ICSM | 1 |
| 2012 | Leveraging User-Privilege Classification to Customize Usage-based Statistical Models of Web ApplicationsabstractAutomatically creating test cases from statistical models of web application usage is an effective approach to generating test cases that represent actual usage. The models are typically generated from all collected user sessions. In this paper, we consider how grouping the user sessions -- specifically by the user's privilege -- creates different statistical models and the testing implications of those differences. We performed a study of user-privilege-specific navigation models and the resulting abstract test cases generated from over 19,000 user sessions to four deployed web applications. Our results suggest that grouping user sessions by the users' privileges results in smaller navigation models, which yield realistic test cases that represent users with that privilege well while also exploring navigations not seen in the input user sessions. In some cases, the user-privilege-specific models are significantly smaller, which allows the tester to either (a) generate relatively few test cases and still represent the user type well or (b) create test cases from a less abstract model -- without exorbitant model space costs or the need for additional models to generate executable test cases. However, the benefits are not universal for all applications, thus, we present guidance to testers on metrics to determine whether creating user-privilege-specific test cases will be advantageous. Sara Sprenkle, Camille Cobb, Lori L. Pollock |
ICST | 3 |
| 2012 | Integrating hard and soft skills: software engineers serving middle school teachersabstractWe have developed and implemented, over four semesters, a model for engaging computer science majors in service learning for teachers of grades 6-8 at a K-8 school in an underserved community. This paper describes the design of a course focused on interweaving software engineering practice, service learning, and development of "soft" professional skills. CS student teams partner with middle school teacher teams to create learning games for classrooms, and then conduct classroom instruction and observation. We report on our results from evaluating the experience of CS students and middle school teachers through pre-post surveys, evaluator observation of student demo presentations and classroom instruction, focus groups, and student reflective journals. Richard Burns, Lori L. Pollock, Terry Harvey |
SIGCSE | 2 |
| 2011 | Automatically detecting and describing high level actions within methodsabstractOne approach to easing program comprehension is to reduce the amount of code that a developer has to read. Describing the high level abstract algorithmic actions associated with code fragments using succinct natural language phrases potentially enables a newcomer to focus on fewer and more abstract concepts when trying to understand a given method. Unfortunately, such descriptions are typically missing because it is tedious to create them manually. Giriprasad Sridhara, Lori L. Pollock, K. Vijay-Shanker |
ICSE | 2 |
| 2011 | A Study of Usage-Based Navigation Models and Generated Abstract Test Cases for Web ApplicationsabstractWhile web applications expand in usage and complexity, testing demands are growing without corresponding automated support. One promising approach to automatic test generation is statistical model-based testing, where logged user behavior is used to build a usage-based model of web application navigation, from which abstract test cases are generated. Executable test cases are then created by adding parameter values to the abstract test cases. Several researchers have proposed variations of this approach, however, no one has empirically examined the tradeoffs and implications of the different ways to represent user behavior in a navigation model and the characteristics of the automatically generated test cases from different models. We report on our exploratory study of automatically generated abstract test cases and the underlying usage-based navigation models constructed from over 3500 user sessions across five publicly deployed web applications. Our results suggest how web testers can easily tune statistical model-based automatic test case generators for web applications toward generating tests closely related to user behavior or toward new navigations without using large additional test resources. Sara Sprenkle, Lori L. Pollock, Lucy Simko |
ICST | 2 |
| 2011 | Combining multiple pedagogies to boost learning and enthusiasmabstractThis paper describes the pedagogy we applied in a 5-week class, in which students taught themselves (and each other) a new language, new OS, GUI programming, and simple networking for collaborative games. They learned communication, negotiation, collaboration, presentation and teamwork skills; and project design and iterative development. We had four goals: increased learning, enthusiasm about CS, confidence in technical ability and communication skills. To achieve these goals, we decided to rely solely on the integration of teaching techniques that we believed would be highly effective: collaborative teams, student presentations, student critique of work, open-ended projects of student design, iterative process, journal re-ection, and motivation through helping others. The students had to learn about each technique through discussion, modeling, and moderated practice. We focused on this process learning and trusted that the technical material would come from solving the (unspecified) assignments. This focus left no time for traditional teaching activities. We present quantitative and qualitative results from a student survey and the students' re-ective journals. Students reported learning at a greater rate than in other CS courses while maintaining (and in some cases acquiring) a high level of enthusiasm and confidence. Lori L. Pollock, Terry Harvey |
ITiCSE | 1 |
| 2011 | Generating Parameter Comments and Integrating with Method SummariesabstractAn important part of the leading comments for a method are the comments for the formal parameters of the method. According to the Java documentation writing guidelines, developers should write a summary of the method'sactions followed by comments for each parameter. In this paper, we describe a novel technique to automatically generate descriptive comments for parameters of Java methods. Such generated comments can help alleviate the lack of developer written parameter comments. In addition, they can help a programmer in ensuring that a parameter comment is current with the code. We present heuristics to generate comments that provide a high-level overview of the role of a parameter in a method. We ensure that sufficient context is provided such that a developer can understand the role of the parameter in achieving the computational intent of the method. In the opinion of nine experienced developers, the automatically generated parameter comments for methods are accurate and provide a quick synopsis of the role of the parameter in achieving the desired functionality of the method. Giriprasad Sridhara, Lori L. Pollock, K. Vijay-Shanker |
ICPC | 2 |
| 2011 | Improving source code search with natural language phrasal representations of method signaturesabstractAs software continues to grow, locating code for maintenance tasks becomes increasingly difficult. Software search tools help developers find source code relevant to their maintenance tasks. One major challenge to successful search tools is locating relevant code when the user's query contains words with multiple meanings or words that occur frequently throughout the program. Traditional search techniques, which treat each word individually, are unable to distinguish relevant and irrelevant methods under these conditions. In this paper, we present a novel search technique that uses information such as the position of the query word and its semantic role to calculate relevance. Our evaluation shows that this approach is more consistently effective than three other state of the art search techniques. Emily Hill 0001, Lori L. Pollock, K. Vijay-Shanker |
ASE | 2 |
| 2011 | PASTE'11: Proceedings of the 10th ACM sigplan-sigsoft workshop on program analysis for software tools and engineeringabstractNo abstract available. Jeff Foster, Lori L. Pollock |
SIGSOFT FSE | 2 |
| 2010 | Towards automatically generating summary comments for Java methodsabstractStudies have shown that good comments can help programmers quickly understand what a method does, aiding program comprehension and software maintenance. Unfortunately, few software projects adequately comment the code. One way to overcome the lack of human-written summary comments, and guard against obsolete comments, is to automatically generate them. In this paper, we present a novel technique to automatically generate descriptive summary comments for Java methods. Given the signature and body of a method, our automatic comment generator identifies the content for the summary and generates natural language text that summarizes the method's overall actions. According to programmers who judged our generated comments, the summaries are accurate, do not miss important content, and are reasonably concise. Giriprasad Sridhara, Emily Hill 0001, Divya Muppaneni, Lori L. Pollock, K. Vijay-Shanker |
ASE | 4 |
| 2009 | MPI-aware compiler optimizations for improving communication-computation overlapabstractSeveral existing compiler transformations can help improve communication-computation overlap in MPI applications. However, traditional compilers treat calls to the MPI library as a black box with unknown side effects and thus miss potential optimizations. This paper's contributions enable the development of an MPI-aware optimizing compiler that can perform transformations exploiting knowledge of MPI call effects to increase communication-computa-tion overlap. We formulate a set of data flow equations and rules to describe the side effects of key MPI functions so an MPI-aware compiler can automatically assess the safety of transformations. After categorizing existing compiler transformations based on their effect on the application code, we present an optimization algorithm that specifies when and how to apply these optimizing transformations to achieve improved communication-computation overlap. By manually applying the optimization algorithm to kernels extracted from HYCOM and the NAS benchmarks, we show that even when transforming these highly optimized codes, execution time can be decreased by an average of over 30%. Anthony Danalis, Lori L. Pollock, D. Martin Swany, John Cavazos |
ICS | 2 |
| 2009 | Automatically capturing source code context of NL-queries for software maintenance and reuseabstractAs software systems continue to grow and evolve, locating code for maintenance and reuse tasks becomes increasingly difficult. Existing static code search techniques using natural language queries provide little support to help developers determine whether search results are relevant, and few recommend alternative words to help developers reformulate poor queries. In this paper, we present a novel approach that automatically extracts natural language phrases from source code identifiers and categorizes the phrases and search results in a hierarchy. Our contextual search approach allows developers to explore the word usage in a piece of software, helping them to quickly identify relevant program elements for investigation or to quickly recognize alternative words for query reformulation. An empirical evaluation of 22 developers reveals that our contextual search approach significantly outperforms the most closely related technique in terms of effort and effectiveness. Emily Hill 0001, Lori L. Pollock, K. Vijay-Shanker |
ICSE | 2 |
| 2009 | Mining source code to automatically split identifiers for software analysisabstractAutomated software engineering tools (e.g., program search, concern location, code reuse, quality assessment, etc.) increasingly rely on natural language information from comments and identifiers in code. The first step in analyzing words from identifiers requires splitting identifiers into their constituent words. Unlike natural languages, where space and punctuation are used to delineate words, identifiers cannot contain spaces. One common way to split identifiers is to follow programming language naming conventions. For example, Java programmers often use camel case, where words are delineated by uppercase letters or non-alphabetic characters. However, programmers also create identifiers by concatenating sequences of words together with no discernible delineation, which poses challenges to automatic identifier splitting. In this paper, we present an algorithm to automatically split identifiers into sequences of words by mining word frequencies in source code. With these word frequencies, our identifier splitter uses a scoring technique to automatically select the most appropriate partitioning for an identifier. In an evaluation of over 8000 identifiers from open source Java programs, our Samurai approach outperforms the existing state of the art techniques. Eric Enslen, Emily Hill 0001, Lori L. Pollock, K. Vijay-Shanker |
MSR | 3 |
| 2008 | Introducing gravel: An MPI companion libraryabstractA non-trivial challenge in high performance, cluster computing is the communication overhead introduced by the cluster interconnect. A common strategy for addressing this challenge is the use of communication-computation overlapping. In this paper, we introduce Gravel, a portable communication library designed to inter-operate with MPI for improving communication-computation overlapping. Selected MPI calls are semi-automatically replaced by Gravel calls only in key locations in an application where performance is critical. Gravel separates the data transfers from the handshake messages, enabling application developers to utilize Remote Data Memory Access (RDMA) directly, whether the data exchange scheme of their application is pure one-sided, or a more traditional two-sided. The Gravel API is much simpler than existing low level libraries and details (i.e., pointers) are hidden from the application layer so Gravel is usable in FORTRAN applications. This paper presents an overview of Gravel. Anthony Danalis, Aaron Brown, Lori L. Pollock, D. Martin Swany |
IPDPS | 3 |
| 2008 | RUGRAT: Runtime Test Case Generation Using Dynamic CompilersabstractThe testing of error handling and dynamic security mechanisms often depends on reproducing specific conditions outside the realm of an application's normal program state. We present RUGRAT, a novel technique to automatically generate tests for these challenging test situations. RUGRAT uses a dynamic compiler to add instructions to the program during execution, and thus dynamically generates tests to exercise code designed to handle uncommon situations during program execution. The RUGRAT testing approach is independent of the source language, requires no modification to the source orbinary program under test and generates runtime tests automatically based on a simple test specification. We demonstrate RUGRAT's capabilities by targeting two particular uncommon situations: handling errors from system and application calls, and testing security mechanisms that protect a program against attacks on function pointers. Both code coverage and failure detection results indicate that RUGRAT is a cost effective approach that reduces the number of required test inputs and need for vulnerable programs. Ben Breech, Lori L. Pollock, John Cavazos |
ISSRE | 2 |
| 2008 | Identifying Word Relations in Software: A Comparative Study of Semantic Similarity ToolsabstractModern software systems are typically large and complex, making comprehension of these systems extremely difficult. Experienced programmers comprehend code by seamlessly processing synonyms and other word relations. Thus, we believe that automated comprehension and software tools can be significantly improved by leveraging word relations in software. In this paper, we perform a comparative study of six state of the art, English-based semantic similarity techniques and evaluate their effectiveness on words from the comments and identifiers in software. Our results suggest that applying English-based semantic similarity techniques to software without any customization could be detrimental to the performance of the client software tools. We propose strategies to customize the existing semantic similarity techniques to software, and describe how various program comprehension tools can benefit from word relation information. Giriprasad Sridhara, Emily Hill 0001, Lori L. Pollock, K. Vijay-Shanker |
ICPC | 3 |
| 2008 | AMAP: automatically mining abbreviation expansions in programs to enhance software maintenance toolsabstractWhen writing software, developers often employ abbreviations in identifier names. In fact, some abbreviations may never occur with the expanded word, or occur more often in the code. However, most existing program comprehension and search tools do little to address the problem of abbreviations, and therefore may miss meaningful pieces of code or relationships between software artifacts. In this paper, we present an automated approach to mining abbreviation expansions from source code to enhance software maintenance tools that utilize natural language information. Our scoped approach uses contextual information at the method, program, and general software level to automatically select the most appropriate expansion for a given abbreviation. We evaluated our approach on a set of 250 potential abbreviations and found that our scoped approach provides a 57% improvement in accuracy over the current state of the art. Emily Hill 0001, Zachary P. Fry, Haley Boyd, Giriprasad Sridhara, Yana Novikova, Lori L. Pollock, K. Vijay-Shanker |
MSR | 6 |
| 2007 | Automatic MPI application transformation with ASPhALTabstractThis paper describes a source to source compilation tool for optimizing MPI-based parallel applications. This tool is able to automatically apply a "prepushing" transformation that causes MPI programs to aggressively send data as soon as it is available, thus improving communication-computation overlap and improving application performance. In this paper we present asphalt_transformer; the Open64-based component of our framework, ASPhALT, responsible for automatically performing the prepushing transformation. We also present an extensive study of the performance gains witnessed from automatically transformed codes. In particular, we demonstrate how different levels of aggregation affect the performance of parallel programs executing various computation kernels on different clusters. Furthermore, we discuss the differences in performance improvement between the hand-optimized and automatically optimized codes, as well as the effect of automation on time-to-solution. Anthony Danalis, Lori L. Pollock, D. Martin Swany |
IPDPS | 2 |
| 2007 | Automated Oracle Comparators for TestingWeb ApplicationsabstractSoftware developers need automated techniques to maintain the correctness of complex, evolving Web applications. While there has been success in automating some of the testing process for this domain, there exists little automated support for verifying that the executed test cases produce expected results. We assist in this tedious task by presenting a suite of automated oracle comparators for testing Web applications. To effectively identify failures, each comparator is specialized to particular characteristics of the possibly nondeterministic Web applications' output in the form of HTML responses. We also describe combinations of comparators designed to achieve both high precision and recall in failure detection and a tool for helping testers to analyze the output of multiple oracles in detail. We present results from an evaluation of the effectiveness and costs of the oracle comparators. We also provide recommendations to testers on applying effective oracle comparators based on their application's characteristics. Sara Sprenkle, Lori L. Pollock, Holly Esquivel, Barbara Hazelwood, Stacey Ecott |
ISSRE | 2 |
| 2007 | Exploring the neighborhood with dora to expedite software maintenanceabstractCompleting software maintenance and evolution tasks for today's large, complex software systems can be difficult, often requiring considerable time to understand the system well enough to make correct changes. Despite evidence that successful programmers use program structure as well as identifier names to explore software, most existing program exploration techniques use either structural or lexical identifier information. By using only one type of information, automated tools ignore valuable clues about a developer's intentions - clues critical to the human program comprehension process. In this paper, we present and evaluate a technique that exploits both program structure and lexical information to help programmers more effectively explore programs. Our approach uses structural information to focus automated program exploration and lexical information to prune irrelevant structure edges from consideration. For the important program exploration step of expanding from a seed, our experimental results demonstrate that an integrated lexical-and structural-based approach is significantly more effective than a state-of-the-art structural program exploration technique Emily Hill 0001, Lori L. Pollock, K. Vijay-Shanker |
ASE | 2 |
| 2007 | Introducing natural language program analysisabstractThis research group presentation focuses on our work in extracting and utilizing natural language clues from source code to improve software maintenance tools. We demonstrate the valuable information that can be gained from a software system's identifiers, literals, and comments. We then present an overview of our extraction process, program representation, and a set of tools we have developedusing this natural language program analysis. Lori L. Pollock, K. Vijay-Shanker, David C. Shepherd, Emily Hill 0001, Zachary P. Fry, Kishen Maloor |
PASTE | 1 |
| 2007 | Case study: supplementing program analysis with natural language analysis to improve a reverse engineering taskabstractSoftware maintainers often use reverse engineering tools to aid in the extremely difficult task of understanding unfamiliar code, especially within large, complex software systems. While traditional program analysis can provide detailed information for reverse engineering, often this information is not sufficient to assist the user with high-level program understanding tasks. To bridge the gap between current reverse engineering tools and the high-level questions that software maintainers want answered, we propose supplementing traditional program analysis with natural language analysis of program source code. This paper presents a case study where we have augmented an existing reverse engineering tool, an aspect miner, to complement the existing traditional program analysis-based miner with natural language analysis of method names, class names, and comments. Our quantitative and qualitative results strongly suggest that supplementing traditional program analysis with natural language analysis is a promising approach to raising the level of effectiveness of reverse engineering tools. David C. Shepherd, Lori L. Pollock, K. Vijay-Shanker |
PASTE | 2 |
| 2007 | Applying Concept Analysis to User-Session-Based Testing of Web ApplicationsabstractThe continuous use of the web for daily operations by businesses, consumers, and the government has created a great demand for reliable web applications. One promising approach to testing the functionality of web applications leverages user-session data collected by web servers. User-session-based testing automatically generates test cases based on real user profiles. The key contribution of this paper is the application of concept analysis for clustering user sessions and a set of heuristics for test case selection. Existing incremental concept analysis algorithms are exploited to avoid collecting and maintaining large user-session data sets and thus to provide scalability. We have completely automated the process from user session collection and test suite reduction through test case replay. Our incremental test suite update algorithm coupled with our experimental study indicate that concept analysis provides a promising means for incrementally updating reduced test suites in response to newly captured user sessions with little loss in fault detection capability and program coverage. Sreedevi Sampath, Sara Sprenkle, Emily Hill 0001, Lori L. Pollock, Amie Souter Greenwald |
IEEE Trans. Software Eng. | 4 |
| 2006 | Integrating Influence Mechanisms into Impact Analysis for Increased PrecisionabstractSoftware change impact analysis is the process of determining the potential effects, or impacts, of a change to a program. Strategies for impact analysis vary in their approach toward the opposing goals of high precision and low analysis time. Fine-grained techniques, such as slicing, can be used to gain very precise knowledge of a change's impact, but may be prohibitively expensive. Coarse-grained techniques such as method-level impact analyses sacrifice precision for faster analysis. In this paper, we present static and dynamic method-level impact analysis algorithms that utilize value propagation information from the source code to increase precision and keep analysis times low. We experimentally compare the results of our analyses with common static and dynamic impact analysis techniques. Our results show that the precision of the common method-level analyses can be improved with very little added overhead Ben Breech, Mike Tegtmeyer, Lori L. Pollock |
ICSM | 3 |
| 2006 | An automated approach to improve communication-computation overlap in clustersabstractApplications that execute on parallel clusters face scalability concerns due to the high communication overhead that is usually associated with such environments. Modern network technologies that support remote direct memory access (RDMA) can offer true zero copy communication and reduce communication overhead by overlapping it with computation. For this approach to be effective the parallel application using the cluster must be structured in a way that enables communication computation overlapping. Unfortunately, the trade-off between maintainability and performance often leads to a structure that prevents exploiting the potential for communication computation overlapping. This paper describes a source-to-source optimizing transformation that can be performed by an automatic (or semi-automatic) system in order to restructure MPI codes towards maximizing communication-computation overlapping. Lewis Fishgold, Anthony Danalis, Lori L. Pollock, D. Martin Swany |
IPDPS | 3 |
| 2006 | An Attack Simulator for Systematically Testing Program-based Security MechanismsabstractThe use of insecure programming practices has led to a large number of vulnerable programs that can be exploited for malicious purposes. These vulnerabilities are often difficult to find during traditional software testing. In response to these difficulties, various program-based security mechanisms have been proposed to help protect potentially vulnerable programs. Testing these security mechanisms, however, also can be difficult and is currently rather ad hoc. In this paper, we describe the design, implementation, and evaluation of an attack simulator that enables the systematic and semi-automatic testing and evaluation of the effectiveness of current and future security mechanisms by automatically providing numerous contexts for testing the reliability of the mechanisms. Capable of automatically creating attacks on running programs by dynamically adding code (but not modifying existing code), the attack simulator can run in different modes and simulate attacks at various program points systematically. Through a case study, we demonstrate how our tool can be used to test two well-known security mechanisms for stack smashing attacks in several different testing modes Ben Breech, Mike Tegtmeyer, Lori L. Pollock |
ISSRE | 3 |
| 2006 | Web Application Testing with Customized Test Requirements - An Experimental Comparison StudyabstractTest suite reduction uses test requirement coverage to determine if the reduced test suite maintains the original suite's requirement coverage. Based on observations from our previous experimental studies on test suite reduction, we believe there is a need for customized test requirements for Web applications. In this paper, we examine usage-based customized test requirements for the test suite reduction problem in Web application testing. We conduct an extensive experimental study to evaluate the tradeoffs between five classes of customized requirements with respect to reduced test suite size, program coverage and fault detection effectiveness. Our results show that the reduced suites' program coverage and fault detection effectiveness increases with the context or data associated with the reduction requirement. Based on our experimental results, we provide guidance to testers on the most useful test requirement for Web applications in general and provide intuition on factors testers need to consider when selecting test requirements Sreedevi Sampath, Sara Sprenkle, Emily Hill 0001, Lori L. Pollock |
ISSRE | 4 |
| 2005 | Third international workshop on dynamic analysis(WODA 2005)abstractDynamic analysis techniques reason over program executions and show promise in aiding the development of robust and reliable large-scale systems. It has become increasingly clear that limitations of static analysis can be overcome by integrating static and dynamic analyses, and that the performance and value of dynamic analysis can be improved by static analysis. Hence, a key focus of the workshop will be on hybrid analyses that involve both static and dynamic components. James H. Andrews, Lori L. Pollock |
ICSE | 2 |
| 2005 | An Empirical Comparison of Test Suite Reduction Techniques for User-Session-Based Testing of Web ApplicationsabstractAutomated cost-effective test strategies are needed to provide reliable, secure, and usable Web applications. As a software maintainer updates an application, test cases must accurately reflect usage to expose faults that users are most likely to encounter. User-session-based testing is an automated approach to enhancing an initial test suite with real user data, enabling additional testing during maintenance as well as adding test data that represents usage as operational profiles evolve. Test suite reduction techniques are critical to the cost effectiveness of user-session-based testing because a key issue is the cost of collecting, analyzing, and replaying the large number of test cases generated from user-session data. We performed an empirical study comparing the test suite size, program coverage, fault detection capability, and costs of three requirements-based reduction techniques and three variations of concept analysis reduction applied to two Web applications. The statistical analysis of our results indicates that concept analysis-based reduction is a cost-effective alternative to requirements-based approaches. Sara Sprenkle, Sreedevi Sampath, Emily Hill 0001, Lori L. Pollock, Amie L. Souter |
ICSM | 4 |
| 2005 | Timna: a framework for automatically combining aspect mining analysesabstractTo realize the benefits of Aspect Oriented Programming (AOP), developers must refactor active and legacy code bases into an AOP language. When refactoring, developers first need to identify refactoring candidates, a process called aspect mining. Humans perform mining by using a variety of clues to determine which code to refactor. However, existing approaches to automating the aspect mining process focus on developing analyses of a single program characteristic. Each analysis often finds only a subset of possible refactoring candidates and is unlikely to find candidates which humans find by combining analyses. In this paper, we present Timna, a framework for enabling the automatic combination of aspect mining analyses. The key insight is the use of machine learning to learn when to refactor, from vetted examples. Experimental evaluation of the cost-effectiveness of Timna in comparison to Fan-in, a leading aspect mining analysis, indicates that such a framework for automatically combining analyses is very promising. David C. Shepherd, Jeffrey Palm, Lori L. Pollock, Mark Chu-Carroll |
ASE | 3 |
| 2005 | Automated replay and failure detection for web applicationsabstractUser-session-based testing of web applications gathers user sessions to create and continually update test suites based on real user input in the field. To support this approach during maintenance and beta testing phases, we have built an automated framework for testing web-based software that focuses on scalability and evolving the test suite automatically as the application's operational profile changes. This paper reports on the automation of the replay and oracle components for web applications, which pose issues beyond those in the equivalent testing steps for traditional, stand-alone applications. Concurrency, nondeterminism, dependence on persistent state and previous user sessions, a complex application infrastructure, and a large number of output formats necessitate developing different replay and oracle comparator operators, which have tradeoffs in fault detection effectiveness, precision of analysis, and efficiency. We have designed, implemented, and evaluated a set of automated replay techniques and oracle comparators for user-session-based testing of web applications. This paper describes the issues, algorithms, heuristics, and an experimental case study with user sessions for two web applications. From our results, we conclude that testers performing user-session-based testing should consider their expectations for program coverage and fault detection when choosing a replay and oracle technique. Sara Sprenkle, Emily Hill 0001, Sreedevi Sampath, Lori L. Pollock |
ASE | 4 |
| 2005 | Transformations to Parallel Codes for Communication-Computation OverlapabstractThis paper presents program transformations directed toward improving communication-computation overlap in parallel programs that use MPI’s collective operations. Our transformations target a wide variety of applications focusing on scientific codes with computation loops that exhibit limited dependence among iterations. We include guidance for developers for transforming an application code in order to exploit the communicationcomputation overlap available in the underlying cluster, as well as a discussion of the performance improvements achieved by our transformations. We present results from a detailed study of the effect of the problem and message size, level of communication-computation overlap, and amount of communication aggregation on runtime performance in a cluster environment based on an RDMA-enabled network. The targets of our study are two scientific codes written by domain scientists, but the applicability of our work extends far beyond the scope of these two applications. Anthony Danalis, Ki-Yong Kim, Lori L. Pollock, D. Martin Swany |
SC | 3 |
| 2004 | Online Impact Analysis via Dynamic Compilation TechnologyabstractDynamic impact analysis based on whole path profiling of method calls and returns has been shown to provide more useful predictions of software change impacts than method-level static slicing and to avoid the overhead of expensive dependency analysis needed for dynamic slicing-based impact analysis. This work presents the design, implementation, and evaluation of an online approach to dynamic impact analysis as an extension to the DynamoRIO binary code modification system and to the Jikes Research Virtual Machine. Storage and postmortem analysis of program traces, even compressed, are avoided. Ben Breech, Anthony Danalis, Stacey A. Shindo, Lori L. Pollock |
ICSM | 4 |
| 2004 | Composing a Framework to Automate Testing of Operational Web-Based SoftwareabstractLow reliability in Web-based applications can result in detrimental effects for business, government, and consumers as they become increasingly dependent on the Internet for routine operations. A short time to market, large user community, demand for continuous availability, and frequent updates motivate automated, cost-effective testing strategies. To investigate the practical tradeoffs of different automated strategies for key components of the Web-based software testing process, we have designed a framework for Web-based software testing that focuses on scalability and evolving the test suite automatically as the application's operational profile changes. We have developed an initial prototype that not only demonstrates how existing tools can be used together but provides insight into the cost effectiveness of the overall approach. This paper describes the testing framework, discusses the issues in building and reusing tools in an integrated manner, and presents a case study that exemplifies the usability, costs, and scalability of the approach. Sreedevi Sampath, Valentin Mihaylov, Amie L. Souter, Lori L. Pollock |
ICSM | 4 |
| 2004 | Scalable Approach to User-Session based Testing of Web Applications through Concept Analysis
Sreedevi Sampath, Valentin Mihaylov, Amie L. Souter, Lori L. Pollock |
ASE | 4 |
| 2004 | Increasing high school girls' self confidence and awareness of CS through a positive summer experience
Lori L. Pollock, Kathleen F. McCoy, Sandra Carberry, Namratha Hundigopal, Xiaoxin You |
SIGCSE | 1 |
| 2003 | Testing with Respect to ConcernsabstractOften the code regions that are assigned for a maintenance task do not follow the modularization of the original application program, but instead include parts of code from many different units scattered throughout the application. In this paper, we investigate an approach to testing which we call concern-based testing, which leverages existing tools to help software maintainers identify the relevant code for their assigned task, their concern. The main contribution is a demonstration of the possible savings in test suite execution overhead and the increased precision in coverage information that can be obtained for a software maintainer if testing tasks are performed with respect to concerns. Based on a concern graph representation of the concern, a framework for guiding selective instrumentation for scalable coverage analysis is also presented. Amie L. Souter, David C. Shepherd, Lori L. Pollock |
ICSM | 3 |
| 2003 | A Framework for Tamper Detection Marking of Mobile ApplicationsabstractToday's applications are highly mobile; we download software from the Internet, machine executable code arrives attached to electronic mail, and Java applets increase the functionality and appearance of Web pages. This movement has stirred a great deal of research in the area of mobile code security. The fact remains that a newly arrived program to a local host has the potential to inflict significant damage to the local host and local resources. Perhaps the new program originated from a charlatan host masquerading as a trusted server, or has been modified by a malicious party during transit from the trusted server to the local host. In light of this risk, security models that address mobile code are in high demand. We have developed a framework named SECRYT, which enables users of a mobile application to validate the application with integrity and authentication data while simplifying the management and distribution of the authentication data. Mike Jochen, Lisa M. Marvel, Lori L. Pollock |
ISSRE | 3 |
| 2003 | All-uses testing of shared memory parallel programsabstractAbstract Parallelism has become a way of life for many scientific programmers. A significant challenge in bringing the power of parallel machines to these programmers is providing them with a suite of software tools similar to the tools that sequential programmers currently utilize. Unfortunately, writing correct parallel programs remains a challenging task.In particular, automatic or semi‐automatic testing tools for parallel programs are lacking. This paper takes a first step in developing an approach to providing all‐uses coverage for parallel programs. A testing framework and theoretical foundations for structural testing are presented, including test data adequacy criteria and hierarchy, formulation and illustration of all‐uses testing problems, classification of all‐uses test cases for parallel programs, and both theoretical and empirical results with regard to what can be achieved with all‐uses coverage for parallel programs. Copyright © 2003 John Wiley & Sons, Ltd. Cheer-Sun D. Yang, Lori L. Pollock |
Softw. Test. Verification Reliab. | 2 |
| 2003 | The Construction of Contextual Def-Use Associations for Object-Oriented SystemsabstractThis paper describes a program representation and algorithms for realizing a novel structural testing methodology that not only focuses on addressing the complex features of object-oriented languages, but also incorporates the structure of object-oriented software into the approach. The testing methodology is based on the construction of contextual def-use associations, which provide context to each definition and use of an object. Testing based on contextual def-use associations can provide increased test coverage by identifying multiple unique contextual def-use associations for the same context-free association. Such a testing methodology promotes more thorough and focused testing of the manipulation of objects in object-oriented programs. This paper presents a technique for the construction of contextual def-use associations, as well as detailed examples illustrating their construction, an analysis of the cost of constructing contextual def-use associations with this approach, and a description of a prototype testing tool that shows how the theoretical contributions of this work can be useful for structural test coverage. Amie L. Souter, Lori L. Pollock |
IEEE Trans. Software Eng. | 2 |
| 2002 | Putting Escape Analysis to Work for Software TestingabstractDeveloped primarily for optimization of functional and object-oriented software, escape analysis discerns information to determine whether the lifetime of data exceeds its static scope. We demonstrate how to apply escape analysis to software engineering tasks. In particular we present novel software testing and retesting techniques for object-oriented software which utilize escape analysis. We exploit a combined pointer and escape analysis that is able to identify how individual objects allocated in one region of a program interact with other regions of a program. The analysis framework increases flexibility and scalability as testing coverage can be targeted to a specific arbitrary region of a program, followed by integration testing that can be focused on particular sets of objects escaping the region. We demonstrate how regression testing can be performed utilizing this framework. We believe such a flexible framework becomes increasingly beneficial as applications become more component-oriented. Amie L. Souter, Lori L. Pollock |
ICSM | 2 |
| 2002 | Characterization and automatic identification of type infeasible call chains
Amie L. Souter, Lori L. Pollock |
Inf. Softw. Technol. | 2 |
| 2001 | Incremental Call Graph Reanalysis for Object-Oriented Software MaintenanceabstractA program's call graph is an essential underlying structure for performing the various interprocedural analyses used in software development tools for object-oriented software systems. For interactive software development tools and software maintenance activities, the call graph needs to remain fairly precise and be updated quickly in response to software changes. The paper presents incremental algorithms for updating a call graph that has been initially constructed using the Cartesian Product Algorithm, which computes a highly precise call graph in the presence of dynamically dispatched message sends. Templates are exploited to reduce unnecessary reanalysis as software component changes occur. The preliminary empirical results from our implementation within a Java environment are encouraging. Significant time savings were observed for the incremental algorithm in comparison with an exhaustive analysis, with no loss in precision. Amie L. Souter, Lori L. Pollock |
ICSM | 2 |
| 2001 | Contextual def-use associations for object aggregationabstractThis paper presents a novel formulation of definitions, uses, and def-use associations for objects in object-oriented programs by exploiting the relations that occur between classes and their instantiated objects due to aggregation. Contextual def-use associations are computed by generating a partial call sequence for each def and use based on object aggregation relations. By extending an escape points-to graph representation of the program, we have developed and implemented three strategies for achieving different levels of context for contextual def-use associations. Our experiments reveal that with all three strategies, multiple unique contextual def-use associations related to the same traditional (context-free) association are often generated. Contextual def-use associations are particularly useful for increasing test coverage and focusing the testing on critical method invocation sequences of object-oriented programs. Amie L. Souter, Lori L. Pollock |
PASTE | 2 |
| 2001 | Integrating an intensive experience with communication skills development into a computer science courseabstractThis paper describes how a technical computer science course was transformed into an intensive communication skills course without sacrificing the technical content of the course. By integrating this experience into existing technical courses, the acquired skills are specific to the CS context without requiring an additional course. The main contribution of this paper is a set of activities which are targeted to building communications skills required for successful research in computer science at any level, but also generally useful for computer science students entering careers not involving basic research. We describe the specific methods and tools implemented in a way to provide considerable support, guidance, and feedback to students without a large investment by the professor. Lori L. Pollock |
SIGCSE | 1 |
| 2001 | Making parallel programming accessible to inexperienced programmers through cooperative learningabstractThis paper describes how we utilized cooperative learning to meet the practical challenges of teaching parallel programming in the early college years, as well as to provide a more real world context to the course. Our main contribution is a set of cooperative group activities for both inside and outside the classroom, which are targeted to the computer science discipline, have received very positive student feedback, are easy to implement, and achieve a number of learning objectives beyond knowledge of the specific topic. These activities can be applied directly or be easily adapted to other computer science courses, particularly programming, systems, and experimental computer science courses. Lori L. Pollock, Mike Jochen |
SIGCSE | 1 |
| 2001 | TATOO: Testing and Analysis Tool for Object- Oriented Software
Amie L. Souter, Tiffany M. Wong, Stacey A. Shindo, Lori L. Pollock |
TACAS | 4 |
| 2000 | Automatic compiler techniques for thread coarsening for multithreaded architecturesabstractMultithreaded architectures are emerging as an important class of parallel machines. By allowing fast context switching between threads on the same processor, these systems hide communication and synchronization latencies and allow scalable parallelism for dynamic and irregular applications. Thread partitioning is the most important task in compiling high-level languages for multithreaded architectures. Non-preemptive multithreaded architectures, which can be built from off-the-shelf components, require that if a thread issues a potentially remote memory request, then any statement that is dependent upon this request must be in a separate thread. Gary M. Zoppetti, Gagan Agrawal, Lori L. Pollock, José Nelson Amaral, Xinan Tang, Guang R. Gao |
ICS | 3 |
| 2000 | OMEN: A strategy for testing object-oriented softwareabstractThis paper presents a strategy for structural testing of object-oriented software systems with possibly unknown clients and unknown information about invoked methods. By exploiting the combined points-to and escape analysis developed for compiler optimization, our testing paradigm does not require a whole program representation to be in memory simultaneously for testing analysis. Potential effects from outside the component under test are easily identified and reported to the tester. As client and server methods become known, the graph representation of object relationships is easily extended, allowing the computation of test tuples to be performed in a demand-driven manner, without requiring unnecessary computation of test tuples based on predictions of potential clients. Amie L. Souter, Lori L. Pollock |
ISSTA | 2 |
| 2000 | Porting and performance evaluation of irregular codes using OpenMPabstractIn the last two years, OpenMP has been gaining popularity as a standard for developing portable shared memory parallel programs. With the improvements in centralized shared memory technologies and the emergence of distributed shared memory (DSM) architectures, several medium-to-large physical and logical shared memory configurations are now available. Thus, OpenMP stands to be a promising medium for developing scalable and portable parallel programs. In this paper, we focus on evaluating the suitability of OpenMP for developing scalable and portable irregular applications. We examine the programming paradigms supported by OpenMP that are suitable for this important class of applications, the performance and scalability achieved with these applications, the achieved locality and uniprocessor cache performance and the factors behind imperfect scalability. We have used two irregular applications and one NAS irregular code as the benchmarks for our study. Our experiments have been conducted on a 64-processor SGI Origin 2000. Our experiments show that reasonably good scalability is possible using OpenMP if careful attention is paid to locality and load balancing issues. Particularly, using the Single Program Multiple Data (SPMD) paradigm for programming is a significant win over just using loop parallelization directives. As expected, the cost of remote accesses is the major factor behind imperfect speedups of SPMD OpenMP programs. Copyright © 2000 John Wiley & Sons, Ltd. Dixie Hisley, Gagan Agrawal, Punyam Satya-narayana, Lori L. Pollock |
Concurr. Pract. Exp. | 4 |
| 1999 | Inter-Class Def-Use Analysis with Partial Class RepresentationsabstractObject-oriented program design promotes the reuse of code not only through inheritance and polymorphism, but also through building server classes which can be used by many different client classes. Research on static analysis of object-oriented software has focused on addressing the new features of classes, inheritance, polymorphism, and dynamic binding. This paper demonstrates how exploiting the nature of object-oriented design principles can enable development of scalable static analyses. We present an algorithm for computing def-use information for a single class's manipulation of objects of other classes, which requires that only partial representations of server classes be constructed. This information is useful for data flow testing and debugging. Amie L. Souter, Lori L. Pollock, Dixie Hisley |
PASTE | 2 |
| 1998 | All-du-path Coverage for Parallel ProgramsabstractOne significant challenge in bringing the power of parallel machines to application programmers is providing them with a suite of software tools similar to the tools that sequential programmers currently utilize. In particular, automatic or semi-automatic testing tools for parallel programs are lacking. This paper describes our work in automatic generation of all-du-paths for testing parallel programs. Our goal is to demonstrate that, with some extension, sequential test data adequacy criteria are still applicable to parallel program testing. The concepts and algorithms in this paper have been incorporated as the foundation of our DELaware PArallel Software Testing Aid, della pasta. Cheer-Sun D. Yang, Amie L. Souter, Lori L. Pollock |
ISSTA | 3 |
| 1998 | The Design and Implementation of RAP: A PDG-Based Register AllocatorabstractThis paper describes the design and implementation of a register allocator that performs the allocation over the Program Dependence Graph (PDG) representation of a routine. The PDG representation has been used successfully as the basis for various scalar optimizations, as well as for detecting and improving parallelization for vector machines, multiple processor machines, and architectures that exhibit instruction level parallelism. Variations of the PDG have also been used for debugging and integrating different versions of a program via program-slicing, and to enable translation of imperative programs for data-flow machines and demand-driven graph reducers. By basing register allocation on the PDG, the register allocation phase may be more easily integrated and intertwined with other optimization analyses and transformations. In addition, the advantages of a hierarchical approach to global register allocation can be attained without constructing an additional structure used solely for register allocation. © 1998 John Wiley & Sons, Ltd. Cindy Norris, Lori L. Pollock |
Softw. Pract. Exp. | 2 |
| 1997 | An Algorithm for All-du-path Testing Coverage of Shared Memory Parallel ProgramsabstractLittle attention has focused on applying traditional testing methodology to parallel programs. This paper discusses issues involved in providing all-du-path coverage in shared memory parallel programs, and describes an algorithm for finding a set of paths covering all define-use pairs. To our knowledge, this is the first effort of this kind. Cheer-Sun D. Yang, Lori L. Pollock |
Asian Test Symposium | 2 |
| 1997 | Issues and Experiences in Implementing a Distributed TuplespaceabstractDistributed memory multiprocessors and network clusters are being used increasingly as parallel computing resources due to their scalability and cost/performance advantages. However, it is generally believed that shared memory parallel programming is easier than explicit message passing programming. Although the generative communication model provides scalability like message passing and the simplicity of shared memory programming, it is a challenge to effectively implement this model on machines with physically distributed memories. This paper describes the issues involved in implementing the essential component of generative communication, the shared data space abstraction called tuplespace, on a distributed memory machine. The paper gives a detailed description of Deli, a UNIX-based distributed tuplespace implementation for a network of workstations. This description, along with discussions of implementation alternatives, provides a detailed basis for designers and implementors of shared data spaces, not currently available in the literature. © 1997 John Wiley & Sons, Ltd. James B. Fenwick Jr., Lori L. Pollock |
Softw. Pract. Exp. | 2 |
| 1996 | Towards a Structural Load Testing ToolabstractLoad sensitive faults cause a program to fail when it is executed under a heavy load or over a long period of time, but may have no detrimental effect under small loads or short executions. In addition to testing the functionality of these programs, testing how well they perform under stress is very important. Current approaches to stress, or load, testing treat the system as a black box, generating test data based on parameters specified by the tester within an operational profile. In this paper, we advocate a structural approach to load testing. There exist many structural testing methods; however, their main goal is generating test data for executing all statements, branches, definition-use pairs, or paths of a program at least once, without consideration for executing any particular path extensively.Our initial work has focused on the identification of potentially load sensitive modules based on a static analysis of the module's code, and then limiting the stress testing to the regions of the modules that could be the potential causes of the load sensitivity. This analysis will be incorporated into a testing tool for structural load testing which takes a program as input, and automatically determines whether that program needs to be load tested, and if so, automatically generates test data for structural load testing of the program. Cheer-Sun D. Yang, Lori L. Pollock |
ISSTA | 2 |
| 1996 | On the Optimality of Change Propagation for Incremental Evaluation of Hierarchical Attribute GrammarsabstractSeveral new attribute grammar dialects have recently been developed, all with the common goal of allowing large, complex language translators to be specified through a modular composition of smaller attribute grammars. We refer to the class of dialects as hierarchical attribute grammars . In this short article, we present a characterization of optimal incremental evaluation that indicates the unsuitability of change propagation as the basis of an optimal incremental evaluator for hierarchical attribute grammars. This result lends strong support to the use of incremental evaluators based on more applicative approaches to attribute evaluation, such as Carle and Pollock's evaluator based on more applicative approaches to attribute evaluation, such as Carle and Pollock's evaluator based on caching of partially attributed subtree, Pugh's evaluator based on function caching of semantic functions, and Swierstra and Vogt's evaluator based on functions, and Swierstra and Vogt's evaluator based on function caching of visit sequences. Alan Carle, Lori L. Pollock |
ACM Trans. Program. Lang. Syst. | 2 |
| 1995 | Register allocation sensitive region scheduling
Cindy Norris, Lori L. Pollock |
PACT | 2 |
| 1995 | An experimental study of several cooperative register allocation and instruction scheduling strategiesabstractCompile-time reordering of low level instructions is successful in achieving large increases in performance of programs on fine-grain parallel machines. However, because of the interdependences between instruction scheduling rand register allocation, a lack of cooperation between the schedules and register allocator can result in generating code that contains excess register spills and/or a lower degree of parallelism than actually achievable. This paper describes a strategy for providing cooperation between register allocation and both global and local instruction scheduling. We experimentally compare this strategy with other cooperative and uncooperative scenarios. Our experiments indicate that the greatest speedups are obtained by performing either cooperative or uncooperative global instruction scheduling with cooperative register allocation and local instruction scheduling. Cindy Norris, Lori L. Pollock |
MICRO | 2 |
| 1995 | Matching-Based Incremental Evaluators for Hierarchical Attribute Grammar DialectsabstractAlthough attribute grammars have been very effective for defining individual modules of language translators, they have been rather ineffective for specifying large program-transformational systems. Recently, several new attribute grammar “dialects” have been developed that support the modular specification of these systems by allowing modules, each described by an attribute grammar, to be composed to form a complete system. Acceptance of these new hierarchical attribute grammar dialects requires the availability of efficient batch and incremental evaluators for hierarchical specifications. This paper addresses the problem of developing efficient incremental evaluators for hierarchical specifications. A matching-based approach is taken in order to exploit existing optimal change propagation algorithms for nonhierarchical attribute grammars. A sequence of four new matching algorithms is presented, each increasing the number of previously computed attribute values that are made available for reuse during the incremental update. Alan Carle, Lori L. Pollock |
ACM Trans. Program. Lang. Syst. | 2 |
| 1994 | Debugging Optimized Code Via Tailoring (Abstract)abstractNo abstract available. Lori L. Pollock, Mary P. Bivens, Mary Lou Soffa |
ISSTA | 1 |
| 1994 | register Allocation over the Program Dependence GraphabstractThis paper describes RAP, a Register Allocator that allocates registers over the Program Dependence Graph (PDG) representation of a program in a hierarchical manner. The PDG program representation has been used successfully for scalar optimizations, the detection and improvement of parallelism for vector machines, multiple processor machines, and machines that exhibit instruction level parallelism, as well as debugging, the integration of different versions of a program, and translation of imperative programs for data flow machines. By basing register allocation on the PDG, the register allocation phase may be more easily integrated and intertwined with other optimization analyses and transformations. In addition, the advantages of a hierarchical approach to global register allocation can be attained without constructing an additional structure used solely for register allocation. Our experimental results have shown that on average, code allocated registers via RAP executed 2.7% faster than code allocated registers via a standard global register allocator. Cindy Norris, Lori L. Pollock |
PLDI | 2 |
| 1993 | A scheduler-sensitive global register allocatorabstractCompile-time reordering of machine-level instructions has been very successful at achieving large increases in performance of programs on machines offering fine-grained parallelism. However, because of the interdependences between instruction scheduling and register allocation, it is not clear which of these two phases of the compiler should run first to generate the most efficient final code. The authors describe their investigation into slight modifications to key phases of a successful global register allocator to create a scheduler-sensitive register allocator, which is then followed by an off-the-shelf instruction scheduler. These experimental studies reveal that this approach achieves speedups comparable and increasingly better than previous cooperative approaches with an increasing number of available registers without the complexities of the previous approaches. Cindy Norris, Lori L. Pollock |
SC | 2 |
| 1992 | Incremental Global Reoptimization of ProgramsabstractAlthough optimizing compilers have been quite successful in producing excellent code, two factors that limit their usefulness are the accompanying long compilation times and the lack of good symbolic debuggers for optimized code. One approach to attaining faster recompilations is to reduce the redundant analysis that is performed for optimization in response to edits, and in particulars, small maintenance changes, without affecting the quality of the generated code. Although modular programming with separate compilation aids in eliminating unnecessary recompilation and reoptimization, recent studies have discovered that more efficient code can be generated by collapsing a modular program through procedure inlining. To avoid having to reoptimize the resultant large procedures, this paper presents techniques for incrementally incorporating changes into globally optimized code. An algorithm is given for determining which optimizations are no longer safe after a program change, and for discovering which new optimizations can be performed in order to maintain a high level of optimization. An intermediate representation is incrementally updated to reflect the current optimizations in the program. Analysis is performed in response to changes rather than in preparation for possible changes, so analysis is not wasted if an edit has no far-reaching effects. The techniques developed in this paper have also been exploited to improve on the current techniques for symbolic debugging of optimized code. Lori L. Pollock, Mary Lou Soffa |
ACM Trans. Program. Lang. Syst. | 1 |
| 1989 | Modular Specification of Incremental Program Transformation SystemsabstractArticle Modular specification of incremental program transformation systems Share on Authors: Alan Carle Department of Computer Science, Rice University, P.O. Box 1892, Houston, Texas Department of Computer Science, Rice University, P.O. Box 1892, Houston, TexasView Profile , Lori Pollock Department of Computer Science, Rice University, P.O. Box 1892, Houston, Texas Department of Computer Science, Rice University, P.O. Box 1892, Houston, TexasView Profile Authors Info & Claims ICSE '89: Proceedings of the 11th international conference on Software engineeringMay 1989 Pages 178–187https://doi.org/10.1145/74587.74612Online:15 May 1989Publication History 3citation238DownloadsMetricsTotal Citations3Total Downloads238Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Alan Carle, Lori L. Pollock |
ICSE | 2 |
| 1989 | An Incremental Version of Iterative Data Flow AnalysisabstractA technique is presented for incrementally updating solutions to both union and intersection data-flow problems in response to program edits and transformations. For generality, the technique is based on the iterative approach to computing data-flow information. The authors show that for both union and intersection problems, some changes can be incrementally incorporated immediately into the data-flow sets while others are handled by a two-phase approach. The first phase updates the data-flow sets to overestimate the effect of the program change, enabling the second phase to incrementally update the affected data-flow sets to reflect the actual program change. An important problem that is addressed is the computation of the data-flow changes that need to be propagated throughout a program, based on different local code changes. The technique is compared to other approaches to incremental data-flow analysis.> Lori L. Pollock, Mary Lou Soffa |
IEEE Trans. Software Eng. | 1 |
| 1985 | Incremental Compilation of Locally Optimized CodeabstractAlthough optimizing compilers have successfully been used to reduce the size and running times of compiled programs, present incremental compilers only support the incremental update of unoptimized code. In this work, we extend the notion of incremental compilation to include optimized code. Techniques to incrementally compile locally optimized code, given intermediate code modifications are developed using a program representation based on flow graphs and dags. A model is designed to represent both unoptimized and optimized code and to maintain an optimizing history. Changes to the optimized code which either destroy optimizations or create conditions for further optimizations are incorporated into the model and the optimized code without recompiling unaffected optimizations. Lori L. Pollock, Mary Lou Soffa |
POPL | 1 |