EDBT 2026 Demo / reviewers in the wild / expert
Stephen H. Edwards
dblp:34/4539
· DBLP profile ↗
99ranked-venue papers
29as first author
15since 2021 · last 2026
0000-0002-5162-9314ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 72 · 19 first-author · 14 since 2021Software engineering, systems software and programming languages · 18 · 8 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Call for Critical Technology to Enable Innovative and Alternative Grading PracticesabstractThe call for alternative grading practices has been made both inside and outside the computing education community. Various practices exist to provide assessment and feedback to students that do not rely strictly on points out of one hundred percent, weighted averages, high stakes assignments, and grading for behaviors instead of learning. However, modern classrooms, especially computer science classrooms, rely on a myriad of digital tools to organize and maintain the course structure. Tools like learning management systems, automatic grading systems, submission systems, and practice systems all exist for computing students and faculty to use to help support the learning of programming concepts. By and large, these systems all rely on an underlying mechanism of points and aggregating points for scoring. In the face of such technological choices, adopting alternative grading practices can prove challenging for instructors and confusing for students. In this position paper, we advocate addressing key research problems to make these systems easier to use with alternative grading practices. These include comprehensive support for categorical grading, comprehensive support for rework and resubmission, and improved protocols for communication of scores and feedback. We propose an extension to LTI to support the needs of alternative grading practices, and we provide an initial design for this LTI extension. We discuss current problems and potential solutions and challenge the community to work on these problems and consider the design of future systems to embrace grading approaches that go beyond just points-based scoring. Adrienne Decker, Stephen H. Edwards, Bob Edmison, Manuel A. Pérez-Quiñones, Audrey Rorrer |
SIGCSE (1) | 2 |
| 2026 | Adoption of Alternative Grading Practices in CS ClassroomsabstractAlternative grading practices seek to improve learning and make education more equitable. For the past few years, our community has been exploring these techniques and contextualizing them for use in computing courses. In this panel, we will present various grading techniques with a particular attention to the experience implementing them and pitfalls to be avoided. The presenters all have extensive experience exploring alternative grading techniques. Furthermore, we hope to engage the audience in a conversation about the nuance of implementing these particular techniques in the variety of courses that are typical in computing programs. Adrienne Decker, Stephen H. Edwards, David L. Largent, Kevin Lin 0001, Nadia Najjar, Manuel A. Pérez-Quiñones |
SIGCSE (2) | 2 |
| 2026 | BOF: Learner-Centered Grading in Computer Science CoursesabstractThe field of computer science has a problem of representation–many groups are not represented in our classroom at levels approaching their composition in society. Unfortunately, the representation issue is a larger societal issue and begins well before students enter our institutions. Though we acknowledge that building inclusive and learner-centered classroom environments cannot increase representation by itself, it can have an impact on retention and inclusion for members of marginalized communities. David L. Largent, Adrienne Decker, Stephen H. Edwards, Manuel A. Pérez-Quiñones, Christian Roberson |
SIGCSE (2) | 3 |
| 2025 | BOF: Grading for Equity in Computer Science CoursesabstractThe field of computer science has a problem of representation-many groups are not represented in our classroom at levels approaching their composition in society. Unfortunately, the representation issue is a larger societal issue and begins well before students enter our institutions. Though we acknowledge that building inclusive and equitable classroom environments cannot increase representation by itself, it can have an impact on retention and inclusion for members of marginalized communities. David L. Largent, Adrienne Decker, Stephen H. Edwards, Manuel A. Pérez-Quiñones, Christian Roberson |
SIGCSE (2) | 3 |
| 2024 | Early Experiences with Specifications Grading in Introductory CS CoursesabstractThis innovative practice paper describes our experiences with alternative grading practices in introductory computing courses and two large public universities in the United States. Computing classrooms often use traditional grading practices involving allocating points to assignments, deducting points for mistakes and tardiness, and combining assignment scores using a weighted average to determine grades. Recent research suggests that these practices may diminish achievement, discourage students, and suppress effort to such an extent that they are considered by some as detrimental. We approach our work as an exploratory case study, without predefined research questions or hypotheses. Our experiences began with the adoption of specifications grading. We outline the grading scheme applied to traditional programming assignments and exams/quizzes, and discuss the initial integration of these schemes with conventional auto-grading tools. We delve into student perceptions of alternative grading, their utilization of flexible deadlines, and resubmission opportunities. We conclude with a discussion of two challenges encountered during our exploration: student acceptance of a novel grading form, and the adaptation of tools designed for traditional grading to support alternative grading mechanisms. Our early exploration aims to inspire further research on the use of alternative grading in computing. It is clear from our observations that simply implementing the practices does not ensure the equitable and inclusive outcomes that can be achieved with these practices. If students are not prepared to use these practices, they find them difficult to understand and can feel that they are not being treated fairly. Additionally, we wish to foster a community of practice to assist faculty members exploring these changes, with the goal of creating more equitable and inclusive classrooms. Stephen H. Edwards, Manuel A. Pérez-Quiñones, Adrienne Decker, Bob Edmison, Audrey Rorrer, Anmol Shukla |
FIE | 1 |
| 2024 | Transforming Grading Practices in the Computing Education CommunityabstractIt is often the case that computer science classrooms use traditional grading practices where points are allocated to assignments, mistakes result in point deductions, and assignment scores are combined using some form of weighted averaging to determine grades. Unfortunately, traditional grading practices have been shown to reduce achievement, discourage students, and suppress effort to such an extent that some common elements of traditional grading practices have been termed toxic. Using grades to reward or punish student behavior does not encourage learning and instead increases anxiety and stress. These toxic elements are present throughout computing education and computer science classrooms in the form of late penalties, lack of credit for code that doesn't compile or pass certain unit tests, among others. These types of metrics, that evaluate behavior are often influenced by implicit bias, factors outside of the classrooms (e.g., part-time employment), and family life situations (e.g., students who are caregivers). Often, students in these situations are disproportionately from low-socioeconomic backgrounds and predominantly students of color. Through this paper, we will present a case for adoption of equitable grading practices and a call for additional support in classroom and teaching technologies as well as support from administrations both at the department and university level. By adopting a community of practice approach, we argue that we can support new faculty making these changes, which would be more equitable and inclusive. Further, these practices have been shown to better support student learning and can help increase student learning gains and retention. Adrienne Decker, Stephen H. Edwards, Brian McSkimming, Bob Edmison, Audrey Rorrer, Manuel A. Pérez-Quiñones |
SIGCSE (1) | 2 |
| 2024 | BOF: Grading for Equity in Computer Science CoursesabstractThe field of computer science has a problem of representation-many groups are not represented in our classroom at levels approaching their composition in society. Unfortunately, the representation issue is a larger societal issue and begins well before students enter our institutions. Though we acknowledge that building inclusive and equitable classroom environments cannot increase representation by itself, it can have an impact on retention and inclusion for members of marginalized communities. David L. Largent, Manuel A. Pérez-Quiñones, Christian Roberson, Linda F. Wilson, Adrienne Decker, Stephen H. Edwards |
SIGCSE (2) | 6 |
| 2023 | Toward a New State-level Framework for Sharing Computer Science ContentabstractThe process of sharing content among instructors at different institutions is not straightforward. In most contexts, "shared" material is unidirectional: a more experienced instructor shares their materials with a more novice instructor; a book publisher provides resources to instructors who have adopted their textbook. In a fully-realized sharing ecosystem, this flow is bi-directional. Materials can be shared, modified, corrected or edited, and then the changes are committed back to the repository for use by everyone who has adopted the material. There are many issues that must be addressed regarding the sharing process, related to both the actual content, as well as the context in which the sharing occurs. In fact, there is a wide opinion on what exactly "sharing" means. As such, we endeavoured to identify sharing opportunities and address issues that sharing course content raises. The primary goal of this initiative was to inform the development of ways in which shared courses and content can be made available online to share across higher education institutions in Virginia, as a potential model to be expanded to other localities and disciplines. Bob Edmison, Stephen H. Edwards, Lujean Babb, Margaret Ellis 0001, Chris Mayfield, Youna Jung, Marthe Honts |
SIGCSE (1) | 2 |
| 2023 | The Programming Exercise Markup Language: Towards Reducing the Effort Needed to Use Automated Grading ToolsabstractAutomated programming assignment grading tools have become integral to CS courses at introductory as well as advanced levels. However such tools have their own custom approaches to setting up assignments and describing how solutions should be tested, requiring instructors to make a significant learning investment to begin using a new tool. In addition, differences between tools mean that initial investment must be repeated when switching tools or adding a new one. Worse still, tool-specific strategies further reduce the ability of educators to share and reuse their assignments. This paper describes an early experiences with PEML, the Programming Exercise Markup Language, which provides an easy to use, instructor friendly approach for writing programming assignments. Unlike tool-oriented data interchange formats, PEML is designed to provide a human friendly authoring format that has been developed to be intuitive, expressive and not be a technological or notational barrier to instructors. We describe the design and implementation of PEML, both as a programming library and also a public-access web microservice that provides full parsing and rendering capabilities for easy integration into any tools or scripting libraries. We also describe experiences using PEML to describe a full range of programming assignments, laboratory exercises, and small coding questions of varying complexity in demonstrating the practicality of the notation. The aim is to develop PEML as a community resource to reduce the barriers to entry for automated assignment tools while widening the scope of programming assignment sharing and reuse across courses and institutions. Divyansh S. Mishra, Stephen H. Edwards |
SIGCSE (1) | 2 |
| 2022 | Piecing Together the Next 15 Years of Computing Education Research Workshop ReportabstractThe session will present an overview of findings of a recently funded NSF workshop that set out to examine the pressing issues for computing education research for the next 15 years. Based on dialogs for scholars working in this area, the workshop participants developed a series of themes and topics that they felt should become the focus of computing education research efforts for the next 15 years. Main themes that emerged and will be discussed in this session include: diversity, equity, inclusion, ethics, broadening participation, teaching, learning, K-12, research to practice, computing's connection to other fields, and computing education research disciplinary issues. The session will focus on interesting research questions from each of these areas that are ripe to be explored as well as enablers and blockers to the progress of this work. Adrienne Decker, Mark Allen Weiss, Brett A. Becker, John P. Dougherty, Stephen H. Edwards, Joanna Goode, Amy J. Ko, Monica McGill, Briana B. Morrison, Manuel A. Pérez-Quiñones, Yolanda A. Rankin, Monique Ross, Jan Vahrenhold, David Weintrop, Aman Yadav |
SIGCSE (2) | 5 |
| 2022 | Factors Influencing Student Performance and Persistence in CS2abstractPerformance in CS1 and introductory CS courses has been an area of active research in the CS education research community for more than four decades, but studies related to student performance in CS2 are not as widely available. Past studies have examined the impact of CS1 grade, prior math preparation, and other factors such as homework, test, and project grades, on the overall performance in CS2. In this work, we will build upon the existing research related to CS2 performance with an emphasis on a few factors that have not been previously considered for this course. In addition to typical factors studied by others (i.e. gender, race, CS1 performance), our work also takes into account the impact of various CS1 pathways to CS2 and the number of previous college CS courses (including transfer credits) on student performance in CS2. We also look into both persistence, by distinguishing students who stay in the course versus those who drop from the class before the mid-semester drop deadline, and performance. Gender and race were not significant factors in determining performance in CS2 but undeclared engineering majors stood out as high performers and students' CS pathway leading to CS2 was also significant. Notably, students with CS1 transfer credit had significantly lower pass rates. Students with only 1 previous CS course credit were less likely to drop or not pass the course. Sara Hooshangi, Margaret Ellis 0001, Stephen H. Edwards |
SIGCSE (1) | 3 |
| 2022 | Helping Student Programmers Through Industrial-Strength Static Analysis: A Replication StudyabstractStatic analysis tools evaluate source code to identify problems beyond typical compiler errors. Prior work has shown a statistically significant relationship between correctness and static analysis results. This paper replicates and extends a prior study on FindBugs, a static analysis tool aimed at professional Java programmers. The prior study showed a strong link between certain FindBugs issues and problems with program correctness. It also showed they were significantly associated with struggling, as indicated by taking more time, making more submissions, and receiving lower scores. However, the study used small programming exercises involving only a handful of lines of code from one semester of a CS1 course. This replication study uses the same experimental approach, but applies it to full-scale programming assignments from hundreds to thousands of lines in length, across all sections of CS1, CS2, and CS3 at a large public university over a period of 4 academic semesters, and involving 4,244 students in 109 laboratory sections completing 255,222 submission attempts. The goal of this replication study is to confirm how prior results hold up, and to explore how the results apply to full-sized programming assignments. We find a set of FindBugs warnings that are inversely correlated with correctness and confirm that their presence is still significantly associated with struggling on larger programming assignments. However, a larger number of FindBugs issues were identified as valuable. We also discuss how student-friendly messages have been added to provided student-readable feedback for a tool where the native messages were written for professional programmers. Allyson Senger, Stephen H. Edwards, Margaret Ellis 0001 |
SIGCSE (1) | 2 |
| 2021 | Experience Report: Exploring the Use of CTF-based Co-Curricular Instruction to Increase Student Comfort and Success in ComputingabstractVery little research exists on the incorporation of co-curricular opportunities, specifically capture-the-flag (CTF) activities, and the impact on computer science students. However, some literature on underrepresented groups in the field of computer science indicates that co-curricular experiences, such as CTF activities, have a significant impact on efficacy and sense of belonging. This experience report seeks to identify and share teaching strategies that impact student learning, efficacy, and sense of self/belonging in the field of Computer Science, for all students. In an effort to design learning opportunities to help bridge the gap between the haves and have nots, and maximize learning for all students, this strategy uses CTF-based instruction as a way to generate and support culture-based, computer science experiences. The aim is to identify instructional strategies that build students' computing skills and improve their comfort participating in computing activities thus broadening participation in computing and cybersecurity. As students' comfort is related to their self-efficacy and sense of belonging, we explore the integration of co-curricular opportunities, specifically a capture-the-flag (CTF) approach to computer science instruction, as means of increasing student success, with regard to learning, efficacy, and sense of belonging/identity. This report involves a description of the class activities as well as data from a series of pre- and post- surveys of the student experience, attitudes, and gain in knowledge. Initial analysis indicates that low-stakes, entry-level exposure to computer science concepts has a positive effect on student comfort and skill acquisition in computing domains external to coursework. Margaret Ellis 0001, Liesl Baum, Kimberly Filer, Stephen H. Edwards |
ITiCSE (1) | 4 |
| 2021 | Automated Feedback, the Next Generation: Designing Learning ExperiencesabstractThe SIGCSE Technical Symposium is a wonderful venue for learning about innovative tools of all sorts for teaching and learning programming. We are at the cusp of a new focus for tool research and development, and this presentation proposes a vision for the next generation of teaching tools, using automated program assessment as an example for how our tools should evolve. Stephen H. Edwards |
SIGCSE | 1 |
| 2021 | Fast and accurate incremental feedback for students' software tests using selective mutation analysisabstractAs incorporating software testing into programming assignments becomes routine, educators have begun to assess not only the correctness of students’ software, but also the adequacy of their tests. In practice, educators rely on code coverage measures, though its shortcomings are widely known. Mutation analysis is a stronger measure of test adequacy, but it is too costly to be applied beyond the small programs developed in introductory programming courses. We demonstrate how to adapt mutation analysis to provide rapid automated feedback on software tests for complex projects in large programming courses. We study a dataset of 1389 student software projects ranging from trivial to complex. We begin by showing that although the state-of-the-art in mutation analysis is practical for providing rapid feedback on projects in introductory courses, it is prohibitively expensive for the more complex projects in subsequent courses. To reduce this cost, we use a statistical procedure to select a subset of mutation operators that maintains accuracy while minimizing cost. We show that with only 2 operators, costs can be reduced by a factor of 2–3 with negligible loss in accuracy. Finally, we evaluate our approach on open-source software and report that our findings may generalize beyond our educational context. Ayaan M. Kazerouni, James C. Davis 0001, Arinjoy Basak, Clifford A. Shaffer, Francisco Servant, Stephen H. Edwards |
J. Syst. Softw. | 6 |
| 2020 | ProgSnap2: A Flexible Format for Programming Process DataabstractIn this paper, we introduce ProgSnap2, a standardized format for logging programming process data. ProgSnap2 is a tool for computing education researchers, with the goal of enabling collaboration by helping them to collect and share data, analysis code, and data-driven tools to support students. We give an overview of the format, including how events, event attributes, metadata, code snapshots and external resources are represented. We also present a case study to evaluate how ProgSnap2 can facilitate collaborative research. We investigated three metrics designed to quantify students' difficulty with compiler errors - the Error Quotient, Repeated Error Density and Watwin score - and compared their distributions and ability to predict students' performance. We analyzed five different ProgSnap2 datasets, spanning a variety of contexts and programming languages. We found that each error metric is mildly predictive of students' performance. We reflect on how the common data format allowed us to more easily investigate our research questions. Thomas W. Price, David Hovemeyer, Kelly Rivers, Austin Cory Bart, Ayaan M. Kazerouni, Brett A. Becker, Andrew Petersen 0001, Luke Gusukuma, Stephen H. Edwards, David S. Babcock |
ITiCSE | 10 |
| 2020 | Auto-Grading Jupyter NotebooksabstractJupyter Notebooks are becoming more widely used, both for data science applications and as a convenient environment for learning Python. Currently, grading of assignments done in Jupyter Notebooks is typically done manually. Manual grading results in students receiving feedback only long after the assignment is complete. We implemented support for auto-grading programs written in Jupyter Notebooks within the Web-CAT auto-grading system. Scores received are directly reported to the Canvas gradebook. A Jupyter notebook extension allows students to upload their notebook files to Web-CAT directly. Survey results from class use show that 80% of students believe that getting immediate feedback from Web-CAT improved their performance. Instructors report that this implementation has significantly reduced their workload. Hamza Manzoor, Amit Naik, Clifford A. Shaffer, Chris North 0001, Stephen H. Edwards |
SIGCSE | 5 |
| 2019 | Experiences Using Heat Maps to Help Students Find Their Bugs: Problems and SolutionsabstractAutomated grading systems provide feedback to students in a variety of ways, but usually focus on identifying incorrect program behaviors. Such systems provide notices of test case failures or runtime errors, but without debugging skills, students often become frustrated when they don't know where to start. They know their code has defects, but finding the problem may be beyond their experience, especially for beginners. An additional concern is balancing the need to provide enough direction to be useful, without giving the student so much direction that you effectively give them the answer. This paper presents our experience using heat maps to visually guide student attention to parts of their code that are most likely to contain problems. These visualizations are generated using existing tools that capture execution traces from instructor-written tests to identify which portions of the code are executed during tests that pass, and which portions are executed during tests that fail. Superimposing execution footprints allows statistical identification of locations in the student's code that are most likely to contain faults. This paper describes the results of using this feedback approach to help guide student attention with heat map visualizations over two semesters of CS1 involving over 700 students. Based on this experience, we analyze the utility of the heat maps, describe student perceptions of their helpfulness, and describe the unexpected challenges arising from students attempts to understand and apply this style of feedback. We conclude with concrete solutions proposed to improve how guiding feedback is presented to students. Bob Edmison, Stephen H. Edwards |
SIGCSE | 2 |
| 2019 | Approaches for Coordinating eTextbooks, Online Programming Practice, Automated Grading, and More into One CourseabstractWe share approaches for coordinating the use of many online educational tools within a CS2 course, including an eTextbook, automated grading system, programming practice website, diagramming tool, and debugger. These work with other commonly used tools such as a response system, forum, version control system, and our learning management system. We describe a number of approaches to deal with the potential negative effects of adopting so many tools. To improve student success we scaffold tool use by staging the addition of tools and by introducing individual tools in phases, we test tool assignments before student use, and we adapt tool use based on student feedback and performance. We streamline course management by consulting mentors who have used the tools before, starting small with room to grow, and choosing tools that simplify student account and grade management across multiple tools. Margaret Ellis 0001, Clifford A. Shaffer, Stephen H. Edwards |
SIGCSE | 3 |
| 2019 | Student Debugging Practices and Their Relationships with Project OutcomesabstractDebugging is an important part of the software development process, studied by both the CS education and software engineering communities. Most prior work has focused either on novice or professional programmers. Intermediate-to-advanced students (such as those enrolled in post-CS2 Data Structures courses) who are working on large and complex projects have largely been ignored. We present results from an empirical observational study that examined junior-level undergraduate students' debugging practices on relatively large (4-week lifecycle) projects, using IDE clickstream data collected by a custom Eclipse plugin. Specifically, we hypothesize that there are differing debugging behaviors exhibited, and that differing behaviors lead to differing project out-comes. For example, how often do students use the symbolic debugger available in modern IDEs, versus how often do they use diagnostic print statements, or both? What triggers a debugging session? What follows a debugging session? Does it matter when in the project lifecycle that debugging takes place? We have a number of interesting preliminary results. When using the debugger, there was a negative relationship between step-over and step-into actions versus final course grades, indicating that when students "spin their wheels" while debugging, they tend to perform more poorly. Students also tend to perform better on the project when debugging takes place earlier in the overall project life-cycle. We developed an algorithm to identify diagnostic print statements in the students' projects. We found that over 90% used at least one diagnostic print statement, and about 75% used the symbolic debugger, at least once in any given project. Ayaan M. Kazerouni, Rifat Sabbir Mansur, Stephen H. Edwards, Clifford A. Shaffer |
SIGCSE | 3 |
| 2019 | Assessing Incremental Testing Practices and Their Impact on Project OutcomesabstractSoftware testing is an important aspect of the development process, one that has proven to be a challenge to formally introduce into the typical undergraduate CS curriculum. Unfortunately, existing assessment of testing in student software projects tends to focus on evaluation of metrics like code coverage over the finished software product, thus eliminating the possibility of giving students early feedback as they work on the project. Furthermore, assessing and teaching the process of writing and executing software tests is also important, as shown by the multiple variants proposed and disseminated by the software engineering community, e.g., test-driven development (TDD) or incremental test-last (ITL). We present a family of novel metrics for assessment of testing practices for increments of software development work, thus allowing early feedback before the software project is finished. Our metrics measure the balance and sequence of effort spent writing software tests in a work increment. We performed an empirical study using our metrics to evaluate the test-writing practices of 157 advanced undergraduate students, and their relationships with project outcomes over multiple projects for a whole semester. We found that projects where more testing effort was spent per work session tended to be more semantically correct and have higher code coverage. The percentage of method-specific testing effort spent before production code did not contribute to semantic correctness, and had a negative relationship with code coverage. These novel metrics will enable educators to give students early, incremental feedback about their testing practices as they work on their software projects. Ayaan M. Kazerouni, Clifford A. Shaffer, Stephen H. Edwards, Francisco Servant |
SIGCSE | 3 |
| 2019 | RecurTutor: An Interactive Tutorial for Learning RecursionabstractRecursion is one of the most important and hardest topics in lower division computer science courses. As it is an advanced programming skill, the best way to learn it is through targeted practice exercises. But the best practice problems are time consuming to manually grade by an instructor. As a consequence, students historically have completed only a small number of recursion programming exercises as part of their coursework. We present a new way for teaching such programming skills. Students view examples and visualizations, then practice a wide variety of automatically assessed, small-scale programming exercises that address the sub-skills required to learn recursion. The basic recursion tutorial (RecurTutor) teaches material typically encountered in CS2 courses. Students who used RecurTutor had significantly better grades on recursion exam questions than did students who used typical instruction. Students who experienced RecurTutor spent significantly more time on solving recursive programming exercises than students who experienced typical instruction, and came out with a significantly higher confidence level. Sally Hamouda, Stephen H. Edwards, Hicham G. Elmongui, Jeremy V. Ernst, Clifford A. Shaffer |
ACM Trans. Comput. Educ. | 2 |
| 2018 | Pedagogical Agent as a Teaching Assistant for Programming Assignments: (Abstract Only)abstractPedagogical agents have received a large amount of interest in the recent years. Equipped with the ability to express emotions, these agents can influence the user attitudes, perceptions and behaviour. In our study, we are leveraging these emotionally-intelligent pedagogical agents to deliver effective and efficient feedback to students about their programming assignments and also act as a teaching assistant for any general programming related queries. We have integrated the pedagogical agent as part of Web-CAT - an automated online grading tool for students' programs. One of our main objectives is to communicate clearly the feedback about student programs while motivating them to perform better. Displaying the feedback and motivational messages to students all the time can quickly become noise and students tend to ignore them. Our study is to strategically have the pedagogical agent communicate with the student to provide them growth mindset feedback and also provide motivation to improve upon their work. The feedback would be based upon few indicators which would be triggered based on the student's program and the agent would guide the students to the correct solution by providing appropriate suggestions. The students can also voluntarily ask the agent for feedback and areas of improvement in their work. In addition, the agent can also help the students with any programming related queries or ways to fix a specific error encountered in the student's program. We will conduct a user study to gather feedback from students about the influence of the agent in helping them achieve their goal Stephen H. Edwards, Mukund B. M. Rajagopal, Nischel Kandru |
SIGCSE | 1 |
| 2018 | CS Education Infrastructure for All: Interoperability for Tools and Data Analytics (Abstract Only)abstractCS Education makes heavy use of online educational tools like IDEs, Learning Management Systems, eTextbooks, interactive programming environments, and other smart content. Instructors and students would benefit from greater interoperability between tools. CS Ed researchers increasingly make use of the large collections of data generated by click streams coming from them. However, we all face barriers that slow progress: (1) Educational tools do not integrate well. (2) Information about CS learning process and outcome data generated by one system is not compatible with that from other systems. (3) CS problem solving and learning (e.g., coding solutions) is different from the type of data (discrete answers to questions or verbal responses) that current educational data mining focuses on. This BOF will discuss ways that we might support and better coordinate efforts to build community and capacity among CS Ed researchers, data scientists, and learning scientists toward reducing these barriers. CS Ed infrastructure should support broader re-use of innovative learning content that is instrumented for rich data collection, formats and tools for analysis of learner data, and best practices to make large collections of learner data available to researchers. Achieving these goals requires engaging a large community of researchers to define, develop, and use critical elements of this infrastructure to address specific data-intensive research questions. Clifford A. Shaffer, Peter Brusilovsky, Kenneth R. Koedinger, Stephen H. Edwards |
SIGCSE | 4 |
| 2018 | Peer Review in CS2: Conceptual Learning and High-Level ThinkingabstractIn computer science, students could benefit from exposure to critical programming concepts from multiple perspectives. Peer review is one method to allow students to experience authentic uses of the concepts in an activity that is not itself programming. In this work, we examine how to implement the peer review process in early, object-oriented computer science courses as a way to increase the students’ knowledge of programming concepts, specifically Abstraction, Decomposition, and Encapsulation, and to develop their higher-level thinking skills. We are exploring the peer review process, the effects of the type of review on the reviewers , and the results this has on the students’ learning. To study these ideas, we used peer review activities in CS2 classes at two universities over the course of a semester. Using three groups (one reviewing their peers, one reviewing the instructor, and one completing small design or coding assignments), we measured the students’ conceptual understanding throughout the semester with concept maps and the reviews they completed. We found that reviewing helped students learn Decomposition, especially those reviewing the instructor's programs, but we did not find that it improved the students’ level of thinking. Overall, reviews (peer or otherwise) are beneficial for teaching Decomposition to CS2 students and can be used as an alternative method for teaching other object-oriented programming concepts. Scott A. Turner, Manuel A. Pérez-Quiñones, Stephen H. Edwards |
ACM Trans. Comput. Educ. | 3 |
| 2017 | Investigating Static Analysis Errors in Student Java ProgramsabstractResearch on students learning to program has produced studies on both compile-time errors (syntax errors) and run-time errors (exceptions). Both of these types of errors are natural targets, since detection is built into the programming language. In this paper, we present an empirical investigation of static analysis errors present in syntactically correct code. Static analysis errors can be revealed by tools that examine a program's source code, but this error detection is typically not built into common programming languages and instead requires separate tools. Static analysis can be used to check formatting or commenting expectations, but it also can be used to identify problematic code or to find some kinds of conceptual or logic errors. We study nearly 10 million static analysis errors found in over 500 thousand program submissions made by students over a five-semester period. The study includes data from four separate courses, including a non-majors introductory course as well as the CS1/CS2/CS3 sequence for CS majors. We examine the differences between the error rates of CS major and non-major beginners, and also examine how these patterns change over time as students progress through the CS major course sequence. Our investigation shows that while formatting and Javadoc issues are the most common, static checks that identify coding flaws that are likely to be errors are strongly correlated with producing correct programs, even when students eventually fix the problems. With experience, students produce fewer errors, but the errors that are most frequent are consistent between both computer science majors and non-majors, and across experience levels. These results can highlight student struggles or misunderstandings that have escaped past analyses focused on syntax or run-time errors. Stephen H. Edwards, Nischel Kandru, Mukund B. M. Rajagopal |
ICER | 1 |
| 2017 | Quantifying Incremental Development Practices and Their Relationship to ProcrastinationabstractWe present quantitative analyses performed on character-level program edit and execution data, collected in a junior-level data structures and algorithms course. The goal of this research is to determine whether proposed measures of student behaviors such as incremental development and procrastination during their program development process are significantly related to the correctness of final solutions, the time when work is completed, or the total time spent working on a solution. A dataset of 6.3 million fine-grained events collected from each student's local Eclipse environment is analyzed, including the edits made and events such as running the program or executing software tests. We examine four primary metrics proposed as part of previous work, and also examine variants and refinements that may be more effective. We quantify behaviors such as working early and often, frequency of program and test executions, and incremental writing of software tests. Projects where the author had an earlier mean time of edits were more likely to submit their projects earlier and to earn higher scores for correctness. Similarly earlier median time of edits to software tests was also associated with higher correctness scores. No significant relationships were found with incremental test writing or incremental checking of work using either interactive program launches or running of software tests, contrary to expectations. A preliminary prediction model with 69% accuracy suggests that the underlying metrics may support early prediction of student success on projects. Such metrics also can be used to give targeted feedback to help students improve their development practices. Ayaan M. Kazerouni, Stephen H. Edwards, Clifford A. Shaffer |
ICER | 2 |
| 2017 | CodeWorkout: Short Programming Exercises with Built-in Data CollectionabstractLearning programming techniques can be challenging and frustrating for many students. Many instructors use drill-and-practice strategies to help students develop basic programming techniques and improve their confidence. Online systems that provide short programming exercises with immediate, automated feedback are often seen as a valuable approach to drill-and-practice. However, the relationship between practicing with short programming exercises and performance on larger programming assignments or exams is unclear. This paper describes CodeWorkout, an open-source drill-and-practice system that supports short programming exercises and multiple choice questions. CodeWorkout combines an open, gradual engagement model that allows any student to practice exercises, whether or not they have an account or are enrolled in a course, together with powerful course management features that support graded assignments, quizzes, and even practicum style exams. It also provides a rich data collection and evaluation infrastructure for educational research purposes. We report on initial experiences using CodeWorkout in a CS1 course, including student perceptions of the tool and its benefits. Stephen H. Edwards, Krishnan Panamalai Murali |
ITiCSE | 1 |
| 2017 | DevEventTracker: Tracking Development Events to Assess Incremental Development and ProcrastinationabstractGood project management practices are hard to teach, and hard for novices to learn. Procrastination and bad project management practice occur frequently, and may interfere with successfully completing major programming projects in mid-level programming courses. Students often see these as abstract concepts that do not need to be actively applied in practice. Changing student behavior requires changing how this material is taught, and more importantly, changing how learning and practice are assessed. To provide proper assessment, we need to collect detailed data about how each student conducts their project development as they work on solutions. We present DevEventTracker, a system that continuously collects data from the Eclipse IDE as students program, giving us in-depth insight into students' programming habits. We report on data collected using DevEventTracker over the course of four programming projects involving 370 students in five sections of a Data Structures and Algorithms course over two semesters. These data support a new measure for how well students apply "incremental development" practices. We present a detailed description of the system, our methodology, and an initial evaluation of our ability to accurately assess incremental development on the part of the students. The goal is to help students improve their programming habits, with an emphasis on incremental development and time management. Ayaan M. Kazerouni, Stephen H. Edwards, T. Simin Hall, Clifford A. Shaffer |
ITiCSE | 2 |
| 2015 | The Effects of Procrastination Interventions on Programming Project SuccessabstractIn computer science, procrastination and related problems with managing programming projects are viewed as primary causes of student attrition. Unfortunately, the most successful techniques for reducing procrastination (such as courses in study skills) are resource-intensive and do not scale to large classrooms. In this paper, we describe three course interventions that are designed to be scalable for large classrooms and require few resources to implement. Reflective writing assignments require students to consciously consider how their time management choices impact their classroom performance. Schedule sheets force students to actively plan out the time required to solve a programming project. Email alerts inform students of their progress relative to their peers as they work on an assignment, and suggest ways to improve behavior if their progress is found to be unsatisfactory. We implemented these interventions in a junior-level data structures course and analyzed data from 330 students over two semesters. Separate analyses of reflective writing responses, schedule sheet contents, and e-mail alert contents are discussed, along with student opinions about the value and effectiveness of each treatment. We found a statistically significant relationship between the time when work is completed and its quality, with late work being of lower quality. We found that one of the three interventions had a statistically significant effect on reducing late work: e-mail alerts sent to students to make them more aware of how they were doing with respect to expectations were associated with both a reduction in assignments completed late, and an increase in assignments completed at least one day early. This result was found despite the fact that students reported subjectively that e-mail alerts were of marginal utility. Joshua Martin, Stephen H. Edwards, Clifford A. Shaffer |
ICER | 2 |
| 2015 | Examining Classroom Interventions to Reduce ProcrastinationabstractProcrastination is a common problem for students. Many believe procrastination may keep otherwise competent students from succeeding. However, the most effective interventions for procrastination are resource-intensive---providing supplemental training or courses in study skills and self-regulation. These techniques do not scale to large courses. This paper investigates three new classroom interventions designed to be low-cost and low-effort to implement. Reflective writing assignments ask students to reflect on how their time management choices affect their work. Project schedule sheets require students to plan out and schedule specific tasks on their projects. E-mail situational awareness alerts give students feedback on how their progress compares to others, and to expectations. 353 students over two semesters of a junior-level advanced data structures course participated in a study where these interventions were investigated. While neither reflective writing assignments nor schedule sheets produced any significant effect, e-mail alerts were associated with both significantly reduced rates of late program submissions, and increased rates of early program submissions. As a result, this intervention shows promise for further investigation as a potential strategy for reducing late submissions among students. Stephen H. Edwards, Joshua Martin, Clifford A. Shaffer |
ITiCSE | 1 |
| 2015 | Reconsidering Automated Feedback: A Test-Driven ApproachabstractWriting meaningful software tests requires students to think critically about a problem and consider a variety of cases that might break the solution code. Consequently, to overcome bugs in their code, it would be beneficial for students to reflect over their work and write robust tests rather than relying on trial-and-error techniques. Automated grading systems provide students with prompt feedback on their programming assignments and may help them identify where their interpretation of requirements do not match the instructor's expectations. Kevin Buffardi, Stephen H. Edwards |
SIGCSE | 2 |
| 2015 | Checked Coverage and Object Branch Coverage: New Alternatives for Assessing Student-Written TestsabstractMany educators currently use code coverage metrics to assess student-written software tests. While test adequacy criteria such as statement or branch coverage can also be used to measure the thoroughness of a test suite, they have limitations. Coverage metrics assess what percentage of code has been exercised, but do not depend on whether a test suite adequately checks that the expected behavior is achieved. This paper evaluates checked coverage, an alternative measure of test thoroughness aimed at overcoming this limitation, along with object branch coverage, a structure code coverage metric that has received little discussion in educational assessment. Checked coverage works backwards from behavioral assertions in test cases, measuring the dynamic slice of the executed code that actually influences the outcome of each assertion. Object branch coverage (OBC) is a stronger coverage criterion similar to weak variants of modified condition/decision coverage. We experimentally compare checked coverage and OBC against statement coverage, branch coverage, mutation analysis, and all-pairs testing to evaluate which is the best predictor of how likely a test suite is to detect naturally occurring defects. While checked coverage outperformed other coverage measures in our experiment, followed closely by OBC, both were only weakly correlated with a test suite's ability to detect naturally occurring defects produced by students in the final versions of their programs. Still, OBC appears to be an improved and practical alternative to existing statement and branch coverage measures, while achieving nearly the same benefits as checked coverage. Zalia Shams, Stephen H. Edwards |
SIGCSE | 2 |
| 2014 | Responses to adaptive feedback for software testingabstractAs students learn to program they also learn basic software development methods and techniques, but educators do not often directly assess students' development processes or evaluate their adherence to specific techniques. However, automated grading systems provide opportunities to evaluate students' programming and provide feedback while the student is still in the process of developing. Consequently, automated adaptive feedback may help reinforce effective techniques and processes. Kevin Buffardi, Stephen H. Edwards |
ITiCSE | 2 |
| 2014 | Do student programmers all tend to write the same software tests?abstractWhile many educators have added software testing practices to their programming assignments, assessing the effectiveness of student-written tests using statement coverage or branch coverage has limitations. While researchers have begun investigating alternative approaches to assessing student-written tests, this paper reports on an investigation of the quality of student written tests in terms of the number of authentic, human-written defects those tests can detect. An experiment was conducted using 101 programs written for a CS2 data structures assignment where students implemented a queue two ways, using both an array-based and a link-based representation. Students were required to write their own software tests and graded in part on the branch coverage they achieved. Using techniques from prior work, we were able to approximate the number of bugs present in the collection of student solutions, and identify which of these were detected by each student-written test suite. The results indicate that, while students achieved an average branch coverage of 95.4% on their own solutions, their test suites were only able to detect an average of 13.6% of the faults present in the entire program population. Further, there was a high degree of similarity among 90% of the student test suites. Analysis of the suites suggest that students were following naïve, "happy path" testing, writing basic test cases covering mainstream expected behavior rather than writing tests designed to detect hidden bugs. These results suggest that educators should strive to reinforce test design techniques intended to find bugs, rather than simply confirming that features work as expected. Stephen H. Edwards, Zalia Shams |
ITiCSE | 1 |
| 2014 | Adaptive and social mechanisms for automated improvement of eLearning materialsabstractOnline environments introduce unprecedented scale for formal and informal learning communities. In these environments, user-contributed content enables social constructivist approaches to education. In particular, students can help each other by providing hints and suggestions on how to approach problems, by rating each other's suggestions, and by engaging in discussions about the questions. In addition, students can also learn through composing their own questions. Furthermore, with grounding in Item Response Theory, data mining and statistical student models can assess questions and hints for their quality and effectiveness. As a result, internet-scale learning environments allow us to move from simple, canned quizzing systems to a new model where automated, data-driven analysis continuously assesses and refines the quality of teaching material. Our poster describes a framework and prototype of an online drill-and-practice system that leverages user-contributed content and large-scale data to organically improve itself. Kevin Buffardi, Stephen H. Edwards |
L@S | 2 |
| 2014 | Work-in-progress: program grading and feedback generation with Web-CATabstractWeb-CAT, the Web-based Center for Automated Testing, is the most widely used open-source automated grading system for programming assignments in the world. Web-CAT is customizable and extensible, allowing it to support a wide variety of programming languages and assessment strategies. Web-CAT is most well known as the system that "grades students on how well they test their own code," with experimental evidence that it offers greater learning benefits than more traditional output-comparison grading. This work-in-progress demonstration will show how Web-CAT can be used to automatically grade student work, assess conformance with coding style guidelines, provide students with feedback on how well they have tested their own code, and allow instructors to provide directed hints to students on where to focus their attention for improvements. Stephen H. Edwards |
L@S | 1 |
| 2014 | A formative study of influences on student testing behaviorsabstractWhile Computer Science curricula teach students strategic software development processes, assessment is often product-instead of process-oriented. Test-Driven Development (TDD) has gained popularity in computing education, but evaluating students' adherence to TDD requires analyzing their development processes instead of only their final product. Consequently, we designed an adaptive feedback system for reinforcing incremental testing behaviors. In this paper, we compare the results of the system with different reinforcement schedules and with- or without- visually salient testing goals. We analyzed snapshots of students' programming projects gathered during development and interviewed students at the end of the academic term. From our findings, we identify potential for influencing student development behaviors and suggest future direction for designing adaptive reinforcement. Kevin Buffardi, Stephen H. Edwards |
SIGCSE | 2 |
| 2014 | Introducing CodeWorkout: an adaptive and social learning environment (abstract only)abstractRudimentary programming skills are essential to developing fundamental proficiency in computer science. However, learning programming techniques can be challenging and frustrating for many students. CodeWorkout is an online learning environment that offers drill-and-practice exercises with novel social and adaptive scaffolding. Learners can track their progress on an assortment of computer science areas and skills while taking advantage of social features to discuss questions and help teach each other. Meanwhile, objective measurements of questions and teaching hints help promote the best, most effective content for learning. Our poster demonstrates how both computer science students and teachers benefit from joining the CodeWorkout community and taking advantage of its unique features. Kevin Buffardi, Stephen H. Edwards |
SIGCSE | 2 |
| 2014 | The absolute beginner's guide to JUnit in the classroom (abstract only)abstractSoftware testing has become popular in introductory courses, but many educators are unfamiliar with how to write software tests or how they might be used in the classroom. This workshop provides a practical introduction to JUnit for educators. JUnit is the Java testing framework that is most commonly used in the classroom. Participants will learn how to write and run JUnit test cases; how-to's for common classroom uses (as a behavioral addition to an assignment specification, as part of manual grading, as part of automated grading, as a student-written activity, etc.); and common solutions to tricky classroom problems (testing standard input/output, randomness, main programs, assignments with lots of design freedom, assertions, and code that calls exit()). Stephen H. Edwards, Manuel A. Pérez-Quiñones |
SIGCSE | 1 |
| 2014 | Adaptively identifying non-terminating code when testing student programsabstractInfinite looping problems that keep student programs from termi-nating may occur in many kinds of programming assignments. While non-terminating code is easier to diagnose interactively, it poses different concerns when software tests are being run auto-matically in batch. The common strategy of using a timeout to preemptively kill test runs that execute for too long has limita-tions, however. When one test case gets stuck in an infinite loop, forcible termination prevents any later test cases from running. Worse, when test results are buffered inside a test execution framework, forcible termination may prevent any information from being produced. Further, overly generous timeouts can de-lay the availability of results, and when tests are executed on a shared server, one non-terminating program can delay results for many people. This paper describes an alternative strategy that uses a fine-grained timeout on the execution of each individual test case in a test suite, and that adaptively adjusts this timeout dynamically based on the termination behavior of test cases com-pleted so far. By avoiding forcible termination of test runs, this approach allows all test cases an opportunity to run and produce results, even when infinite looping behaviors occur. Such fine-grained timeouts also result in faster completion of entire test runs when non-terminating code is present. Experimental results from applying this strategy to 4,214 student-written programs are dis-cussed, along with experiences from live deployment in the class-room where 8,926 non-termination events were detected over the course of one academic year. Stephen H. Edwards, Zalia Shams, Craig Estep |
SIGCSE | 1 |
| 2014 | Pythy: improving the introductory python programming experienceabstractPythy is a web-based programming environment for Python that eliminates software-related barriers to entry for novice programmers, such as installing an IDE or the Python runtime. Using only a web browser, within minutes students can begin writing code, watch it run, and access support materials and tutorials. While there are a number of web-based Python teaching tools, Pythy differs in several respects: it manages student assignment work, including deadlines, turn-in, and grading; it supports live, interactive code examples that instructors can write and students can explore; it provides auto-saving of student work in the cloud, with full, transparent version control; and it supports media-computation-style projects that manipulate images and sounds. Pythy provides a complete ecosystem for student learning, with a user interface that follows a more familiar web browsing model, rather than a developer-focused IDE interface. An evaluation compares student perceptions of Pythy in relation to JES, another student-friendly beginner Python environment. Classroom experiences indicate that Pythy does reduce the novice obstacles that it aims to address. Stephen H. Edwards, Daniel S. Tilden, Anthony Allevato |
SIGCSE | 1 |
| 2014 | Using and sharing programming exercises to improve introductory courses (abstract only)abstractShort, automatically-assessed programming exercises, and other types of short practice problems, are a useful way to introduce and reinforce concepts and techniques in introductory programming courses. When delivered over the web, they allow students to learn and practice, with immediate feedback, at any time and place where they have access to a web browser. However, such exercises do not seem to be as widely used as they could be. Similarly, there is not a lot of literature on the effectiveness of these types of problems. The purpose of this BOF is to bring together users (and potential users) of programming exercises with developers of programming exercise systems to discuss how exercises could be used more widely and effectively. Possible discussion topics include: What features are absolutely essential for faculty to consider adoption? What are the major obstacles preventing more widespread adoption? Are faculty willing to share their exercises under an open/non-commercial license? Should exercises best used for extra practice, as graded assignments, or both? David Hovemeyer, Jaime Spacco, Robert C. Duvall, Stephen H. Edwards, Amruth N. Kumar, Andrew Petersen 0001, Daniel Zingaro |
SIGCSE | 4 |
| 2014 | Open source software and the algorithm visualization communityabstractAlgorithm visualizations are widely viewed as having the potential for major impact on computer science education, but their quality is highly variable. We report on the software development practices used by creators of algorithm visualizations, based on data that can be inferred from a catalog of over 600 algorithm visualizations. Since nearly all are free for use and many provide source code, they might be construed as being open source software. Yet many AV developers do not appear to have used open source best practices. We discuss how such development practices might be employed by the algorithm visualization community, and how they might lead to improved algorithm visualizations in the future. We conclude with a discussion of OpenDSA, an open-source project that builds on earlier progress in the field of algorithm visualization and hopes to use open-source procedures to gain users and contributors. Matthew Cooper 0002, Clifford A. Shaffer, Stephen H. Edwards, Sean P. Ponce |
Sci. Comput. Program. | 3 |
| 2014 | Dereferee: instrumenting C++ pointers with meaningful runtime diagnosticsabstractSUMMARY Proper memory management and pointer usage often prove to be the most difficult concepts for students learning C++ to grasp. Compounding this problem is the fact that the compilers and runtime environments traditionally used to introduce these concepts leave much to be desired with regard to generating meaningful diagnostics to assist students in tracking down and fixing memory‐related logical errors. To alleviate this, we have developed Dereferee, an advanced yet thin wrapper around C++ pointers that greatly increases the quality of these runtime diagnostics, but with only a small amount of intrusion into the development process. With regard to performance, memory‐intensive programs will experience execution times approximately 20–30 times slower when using Dereferee, which is comparable with other similar tools. Furthermore, the library has been designed to be customizable and easily disabled to transition codes from development to production.Copyright © 2013 John Wiley & Sons, Ltd. Anthony Allevato, Stephen H. Edwards |
Softw. Pract. Exp. | 2 |
| 2014 | Open source software-defined radio tools for education, research, and rapid prototyping
Jason Snyder, Deepan Seeralan, Shereef Sayed, Jeffery Wilson, Carl B. Dietrich, Stephen H. Edwards, Jeffrey H. Reed |
Int. J. Softw. Tools Technol. Transf. | 6 |
| 2013 | Adding software testing to programming assignmentsabstractThis tutorial provides a practical introduction to how one can incorporate software testing activities as a regular part of programming assignments, supported by live demonstrations, with a special focus on early introduction in CS1 and/or CS2 courses. It presents five different models for how one can incorporate testing into assignments, provides examples of each technique, and discusses the corresponding advantages and disadvantages. The focus is on unit testing, test-driven development, and incremental testing, all of which work well in a classroom environment. Examples will use Java, although participant discussion regarding support in other languages such as Python and C++ is welcome. Approaches to assessment- using testing to assess student code, assessing tests that students write, and automated grading-are all discussed. A live demonstration of automatic assignment grading based on student-written tests is included. Advice for writing “testable” assignments is given. Participant discussions are encouraged. Stephen H. Edwards |
CSEE&T | 1 |
| 2013 | Automatically Generating Tests from Natural Language Descriptions of Software BehaviorabstractBehavior-Driven Development (BDD) is an emerging agile development approach where all stakeholders (including developers and customers) work together to write user stories in structured natural language to capture a software application's functionality in terms of re- quired "behaviors". Developers then manually write "glue" code so that these scenarios can be executed as software tests. This glue code represents individual steps within unit and acceptance test cases, and tools exist that automate the mapping from scenario descriptions to manually written code steps (typically using regular expressions). Instead of requiring programmers to write manual glue code, this thesis investigates a practical approach to con- vert natural language scenario descriptions into executable software tests fully automatically. To show feasibility, we developed a tool called Kirby that uses natural language processing techniques, code information extraction and probabilistic matching to automatically gener- ate executable software tests from structured English scenario descriptions. Kirby relieves the developer from the laborious work of writing code for the individual steps described in scenarios, so that both developers and customers can both focus on the scenarios as pure behavior descriptions (understandable to all, not just programmers). Results from assessing the performance and accuracy of this technique are presented. Sunil Kamalakar, Stephen H. Edwards, Tung M. Dao |
ENASE | 2 |
| 2013 | The effects of extra credit opportunities on student procrastinationabstractMany techniques have been attempted to encourage students to exercise better time management on class projects, such as staging an assignment into multiple deliverables, requiring students to keep records of the time they spend, and offering extra credit for early completion. This paper reports on a study of the effects of offering extra credit for early completion. Students in an introductory course completed four programming assignments throughout the term. For two assignments, no extra credit was offered. For the other two, students were offered a 10% bonus if they finished at least three days before the deadline. While one might expect this incentive to encourage students to shift their work habits, we found that there was no positive change in their time management. In fact, students started on the assignments where extra credit was offered later than on those where it was not offered. This leads us to believe that there were other pressures or concerns that outweigh the possibility of earning a bonus on an assignment, so that this kind of incentive only helps students who already manage their time well. Anthony Allevato, Stephen H. Edwards |
FIE | 2 |
| 2013 | Effective and ineffective software testing behaviors by novice programmersabstractThis data-driven paper quantitatively evaluates software testing behaviors that students exhibited in introductory computer science courses. The evaluation includes data collected over five years (10 semesters) from 49,980 programming assignment submissions by 883 different students. To examine the effectiveness of software testing behaviors, we investigate the quality of their testing at different stages of their development. We partition testing behaviors into four groups according to when in their development they first achieve substantial (at least 85%) test coverage. Kevin Buffardi, Stephen H. Edwards |
ICER | 2 |
| 2013 | Toward practical mutation analysis for evaluating the quality of student-written software testsabstractSoftware testing is being added to programming courses at many schools, but current assessment techniques for evaluating student-written tests are imperfect. Code coverage measures are typically used in practice, but they have limitations and sometimes overestimate the true quality of tests. Others have proposed using mutation analysis instead, but mutation analysis poses a number of practical obstacles to classroom use. This paper describes a new approach to mutation analysis of student-written tests that is more practical for educational use, especially in an automated grading context. This approach combines several techniques to produce a novel solution that addresses the shortcomings raised by more traditional mutation analysis. An evaluation of this approach in the context of both CS1 and CS2 courses illustrates how it differs from code coverage analysis. At the same time, however, the evaluation results also raise questions of concern for CS educators regarding the relative value of more comprehensive assessment of test quality, the value of more open-ended assignments that offer significant design freedom for students, the cost of providing higher-quality reference solutions in order to support better quality assessment, and the cost of supporting assignments that require more intensive testing, such as GUI assignments. Zalia Shams, Stephen H. Edwards |
ICER | 2 |
| 2013 | Sofia: the simple open framework for inventive android applicationsabstractMobile application development in general, and the Android platform in particular, are hot topics among educators because of their power to motivate and engage students. Unfortunately, Android's software API is not designed for beginners and presents a number of stumbling blocks to classroom use. Sofia is a new abstraction layer over the Android API that provides a cleaner, simpler, easier to use API for beginners and professionals alike. It includes a novel event dispatch design that eliminates the glue code required by more conventional frameworks, provides a powerful 2D shape package with declarative animation support and physics simulation, streamlines the process of writing multi-activity apps for Android, and addresses a number of other issues that make Android hard to use in introductory courses. Stephen H. Edwards, Anthony Allevato |
ITiCSE | 1 |
| 2013 | A new event dispatch strategy to eliminate dispatch "glue"abstractIn statically typed object-oriented languages such as Java, GUI event handling is traditionally handled through listener interfaces or similar types of polymorphic delegation. In the case of events that pass information about their source to the handling method, the programmer is required to perform runtime type checks to determine the true types of the components involved. This produces poorly designed code that contains a second layer of hand-written type-based dispatch before events can actually be handled. In this paper we present an alternative approach that builds this second dispatch layer into the underlying framework. The approach uses run-time reflection and overload resolution to automatically distinguish events based on method argument types, and to implicitly bind them to the event publishers. This approach combines the type safety of a statically typed language with the run-time flexibility of modern dynamic languages and enhances the readability of event handling code. Anthony Allevato, Stephen H. Edwards |
RCIS | 2 |
| 2013 | Impacts of adaptive feedback on teaching test-driven developmentabstractStudies have found that following Test-Driven Development (TDD) can improve code and testing quality. However, a preliminary investigation was consistent with concerns raised by other educators about programmers resisting TDD. In this paper, we describe an adaptive, pedagogical system for tracking and encouraging students' adherence to TDD. Along with an empirical evaluation of the system, we discuss challenges and opportunities for persuading student behavior through adaptive technology. Kevin Buffardi, Stephen H. Edwards |
SIGCSE | 2 |
| 2013 | Re-imagining CS1/CS2 with Android using the Sofia framework (abstract only)abstractAndroid has seen increased use in introductory CS courses to motivate and excite students about their programming assignments, but using the standard Android libraries as a GUI platform in CS2 presents numerous challenges and using it in CS1 is nearly impossible. This workshop introduces participants to Sofia, the Simplified Open Framework for Innovative Android Applications, developed by the Web-CAT team at Virginia Tech. Sofia abstracts out many of the advanced concepts normally required to develop interesting applications, using a unique approach to event handling, binding GUI elements to Java code, and user interaction. The goal is to allow students to focus entirely on using Java programming skills to solve problems in the application domain, instead of writing monotonous glue code typically required to construct an Android application. Laptop optional. Stephen H. Edwards |
SIGCSE | 1 |
| 2013 | An experiment to test bug density in students' code (abstract only)abstractA normal industry standard measure, bug density (bugs per thousand non-commented source line of code), is a through mechanism to assess code quality. If it is used for evaluating students' code, students will realize their ability to write bug free code from professional context. The main issues of using bug density for object oriented languages are creating a comprehensive test suit, and running them against all solutions as the test cases are written as part of solutions may fail to compile against other codes. We provide a novel four phase Java specific solution: 1) developing a comprehensive master test suit by collecting all the students written valid test cases; 2) transforming the test cases to use late binding so that they can run against any solution; 3) running the entire tests against all the programs and removing redundant test suits; and 4) estimating bugs/KSLOC by determining the relationship between test case failures in the master suite and latent bugs hidden in student programs. The first two phases of this ongoing research are applied to two programming assignments in two different courses encompassing 147 student programs and 240,158 individual test cases. Experimental results show that we have indeed removed compile-time dependencies from test cases using late binding and thus, have resolved the main technical challenge of using bug density for accessing students' code. Our experimental results will help students to realize the quality of their code in terms of industry standard. Zalia Shams, Stephen H. Edwards |
SIGCSE | 2 |
| 2012 | Exploring influences on student adherence to test-driven developmentabstractTest-Driven Development (TDD) is a software development process with a test-first approach that shows promise for improving code quality. Our research addresses concerns raised in both academia and industry about a lack of motivation or acceptance in adopting TDD. In a CS2 class, we used an automated testing tool and post-class surveys to observe patterns of behavior in testing as well as changes in attitudes. We found significant positive outcomes for students following TDD. We also identified obstacles deterring students from adhering to TDD and discuss reasons and possible remedies. Kevin Buffardi, Stephen H. Edwards |
ITiCSE | 2 |
| 2012 | RoboLIFT: engaging CS2 students with testable, automatically evaluated android applicationsabstractMaking computer science assignments interesting and relevant is a constant challenge for instructors of introductory courses. Android has become popular in these courses to take advantage of the increasing popularity of smartphones and mobile "apps." This has been shown to increase student engagement but it is only the first step, and we must continue to provide support for teaching methodologies that we have used in the past, such as test-driven development and automated assessment. We have developed RoboLIFT, a library that makes unit testing of Android applications approachable for students. Furthermore, by supporting existing automated grading techniques, we are able to sustain large student enrollments, and we evaluate the effects that using Android has had on student performance. Anthony Allevato, Stephen H. Edwards |
SIGCSE | 2 |
| 2012 | RoboLIFT: simple GUI-based unit testing of student-written android applications (abstract only)abstractMany computer science educators have adopted test-driven development practices in their introductory computer science courses, as a way of encouraging incremental development and decreasing defects in student code. This practice is straightforward for basic data-driven objects, but making unit testing of GUI applications approachable for students poses a larger challenge. We have previously addressed this problem for Swing applications by developing LIFT, a library that allows students to easily write JUnit tests for Swing interfaces. Since then, we have transitioned away from Swing to Android as the development platform in CS2 to better motivate and excite our students about their assignments. To fully support this change, we had to ensure that our students could fully test the GUI portions of their solutions on that platform as well. The Android operating system has significant built-in support for GUI testing, but the standard API is too complex for students to use. In order to address this, we developed RoboLIFT, a framework that eases the task of writing concise and complete unit tests for Android applications. Furthermore, RoboLIFT also has support for automated grading on the Web-CAT automated assessment system, so even if instructors do not require their students to follow test-driven development practices, they can still enjoy the benefits of automated grading by writing correctness tests that use RoboLIFT to exercise the students' graphical user interfaces. Anthony Allevato, Stephen H. Edwards |
SIGCSE | 2 |
| 2012 | Web-CAT user group (abstract only)abstractWeb-CAT is the most widely used open-source automated grading system, with about 10,000 users at over 65 institutions worldwide. Its plug-in architecture supports extensibility, with plug-ins for Java (including Objectdraw, JTF, Swing, and Android), C++, Python, Haskell, and more. It is also a powerful tool for educational research data collection. It supports a wide variety of assessment strategies, but is famous for "grading students on how well they test their own code". Web-CAT won the 2006 Premier Award, recognizing high-quality, non-commercial courseware for engineering education. This BOF will allow existing users and new adopters to meet, share experiences, and talk about what works and what doesn't. Information on getting started quickly with Web-CAT will also be provided. Stephen H. Edwards |
SIGCSE | 1 |
| 2012 | The absolute beginner's guide to JUnit in the classroom (abstract only)abstractSoftware testing has become popular in introductory courses, but many educators are unfamiliar with how to write software tests or how they might be used in the classroom. This workshop provides a practical introduction to JUnit for educators. JUnit is the Java testing framework that is most commonly used in the classroom. Participants will learn how to write and run JUnit test cases; how-to's for common classroom uses (as a behavioral addition to an assignment specification, as part of manual grading, as part of automated grading, as a student-written activity, etc.); and common solutions to tricky classroom problems (testing standard input/output, randomness, main programs, assignments with lots of design freedom, assertions, and code that calls exit()). Laptop recommended. Stephen H. Edwards, Manuel A. Pérez-Quiñones |
SIGCSE | 1 |
| 2012 | Running students' software tests against each others' code: new life for an old "gimmick"abstractAt SIGCSE 2002, Michael Goldwasser suggested a strategy for adding software testing practices to programming courses by requiring students to turn in tests along with their solutions, and then running every student's tests against every other student's program. This approach provides a much more robust environment for assessing the quality of student-written tests, and also provides more thorough testing of student solutions. Although software testing is included as a regular part of many more programming courses today, the all-pairs model of executing tests is still a rarity. This is because student-written tests, such as JUnit tests written for Java programs, are now more commonly written in the form of program code themselves, and they may depend on virtually any aspect of their author's own solution. These dependencies may keep one student's tests from even compiling against another student's program. This paper discusses the problem and presents a novel solution for Java that uses bytecode rewriting to transform a student's tests into a form that uses reflection to run against any other solution, regardless of any compile-time dependencies that may have been present in the original tests. Results of applying this technique to two assignments, encompassing 147 student programs and 240,158 individual test case runs, shows the feasibility of the approach and provides some insight into the quality of both student tests and student programs. An analysis of these results is presented. Stephen H. Edwards, Zalia Shams, Michael Cogswell, Robert C. Senkbeil |
SIGCSE | 1 |
| 2012 | Motivating CS1/2 students with the android platform (abstract only)abstractThe use of Android in computing courses is growing. Students find it engaging because it offers a unique opportunity to develop Java apps for mobile devices. Android offers opportunities and challenges in a teaching environment, especially in CS1 and CS2. As a professional-level platform, it incorporates many design idioms that may require students to learn advanced language features earlier. It also introduces logistical complications in setting up development tools and code projects. Existing approaches to software testing and automated grading also must be adapted. This BOF will gather educators interested in using Android in their courses, focusing on issues that arise when balancing the need to teach fundamental concepts with the complexities required to accomplish basic tasks on the Android platform. We look forward to sharing assignments, resources, techniques, and experiences with others interested in Android. Anthony Allevato, Stephen H. Edwards |
SIGCSE | 3 |
| 2012 | A better API for Java reflection (abstract only)abstractInstructors often write reference tests to evaluate student programs. In Java, reference tests should be independent of submitted solutions as they are run against all student submissions. Otherwise, they may even fail to compile against some solutions. Reflection is a useful feature for writing code without compile-time dependencies, which is valuable for writing software tools that inspect code. However, educators avoid using reflection as code written using Java's Reflection API is complex, unintuitive and verbose. We present ReflectionSupport, a library that enables one to write reflection-based code in concise, simple and readable fashion. It helps educators write reference tests without compile-time dependencies of solutions and develop educational tools such as automated graders. Zalia Shams, Stephen H. Edwards |
SIGCSE | 2 |
| 2011 | Scheduling and student performanceabstractWe present data showing strong correlation between students' time management and a successful outcome on programming assignments. Students who spread their work over more time will produce a better result without additional expenditure of total effort. We examined performance of students who sometimes did well and sometimes did poorly, and found that their good performance occurred on the projects where they displayed better time management. While these results will not surprise most instructors, hard data is more compelling than intuition when trying to train students to use good time management. Clifford A. Shaffer, Stephen H. Edwards |
ITiCSE | 2 |
| 2011 | Getting algorithm visualizations into the classroomabstractAlgorithm visualizations (AVs) are widely viewed as having the potential for improving computer science education. However, the rate of AV use and overall impact on education does not match the positive interest in their use that instructors report. Surveys of CS faculty show that impediments to successful use of AVs in the classroom include difficulties in finding quality AVs on desired topics, difficulties in adapting AVs to a given classroom setting, and lack of knowledge on the best way to deploy AVs. This indicates a need for better support for instructors, to get them past these barriers. We seek to provide this support through an online educational community that relies on a new model based less on the "digital library" approach of information gained by going to a site and searching. Instead, the focus is on community-added content through members' discussions, reviews, and ratings of content items. The AlgoViz community effort will better focus the future direction of AV development and use. Clifford A. Shaffer, Monika Akbar, Alexander Joel D. Alon, Michael Stewart 0001, Stephen H. Edwards |
SIGCSE | 5 |
| 2011 | LIFT: taking GUI unit testing to new heightsabstractThe Library for Interface Testing (LIFT) supports writing unit tests for Java applications with graphical user interfaces (GUIs). Current frameworks for GUI testing provide the necessary tools, but are complicated and difficult to use for beginners, often requiring a significant amount of time to learn. LIFT takes the approach that unit testing GUIs should be no different than testing any other type of code. By providing a set of frequently used filters for identifying GUI components and a set of operations for acting on those components, LIFT lets programmers quickly and easily test their GUI applications. Jason Snyder, Stephen H. Edwards, Manuel A. Pérez-Quiñones |
SIGCSE | 2 |
| 2011 | Student attitudes and motivation for peer review in CS2abstractComputer science students need experience with essential concepts and professional activities. Peer review is one way to meet these goals. In this work, we examine the students' attitudes towards and engagement in the peer review process, in early, object-oriented, computer science courses. To do this, we used peer review exercises in two CS2 classes at neighboring universities over the course of a semester. Using three groups (one reviewing their peers, one reviewing the instructor, and one completing small design or coding exercises), we measured the students' attitudes, their perceptions of their abilities, and how many of the reviews they completed. We found moderately positive attitudes that generally increased over time but were not significantly different between groups. We also saw a lower completion rate for students reviewing peers than for the other groups. The students' internal motivation, as measured by their need for cognition, was not shown to be strongly related to their attitudes nor to the number of assignments completed. Overall, our results show a strong need for external motivation to help engage students in peer reviews. Scott A. Turner, Manuel A. Pérez-Quiñones, Stephen H. Edwards, Joseph Chase |
SIGCSE | 3 |
| 2010 | Building an online educational community for algorithm visualizationabstractNo abstract available. Clifford A. Shaffer, Thomas L. Naps, Susan H. Rodger, Stephen H. Edwards |
SIGCSE | 4 |
| 2010 | Peer review in CS2: conceptual learningabstractIn computer science, students could benefit from exposure to critical programming concepts from multiple perspectives. Peer review is one method to allow students to experience authentic uses of the concepts in a non-programming manner. In this work, we examine the use of the peer review process in early, object-oriented, computer science courses as a way to develop the reviewers' knowledge of object-oriented programming concepts, specifically Abstraction, Decomposition, and Encapsulation. Scott A. Turner, Manuel A. Pérez-Quiñones, Stephen H. Edwards, Joseph Chase |
SIGCSE | 3 |
| 2010 | Algorithm Visualization: The State of the FieldabstractWe present findings regarding the state of the field of Algorithm Visualization (AV) based on our analysis of a collection of over 500 AVs. We examine how AVs are distributed among topics, who created them and when, their overall quality, and how they are disseminated. There does exist a cadre of good AVs and active developers. Unfortunately, we found that many AVs are of low quality, and coverage is skewed toward a few easier topics. This can make it hard for instructors to locate what they need. There are no effective repositories of AVs currently available, which puts many AVs at risk for being lost to the community over time. Thus, the field appears in need of improvement in disseminating materials, propagating known best practices, and informing developers about topic coverage. These concerns could be mitigated by building community and improving communication among AV users and developers. Clifford A. Shaffer, Matthew Cooper 0002, Alexander Joel D. Alon, Monika Akbar, Michael Stewart 0001, Sean P. Ponce, Stephen H. Edwards |
ACM Trans. Comput. Educ. | 7 |
| 2009 | Comparing effective and ineffective behaviors of student programmersabstractThis paper reports on a quantitative evaluation of five years of data collected in the first three programming courses at Virginia Tech. The dataset involves a total of 89,879 assignment submissions by 1,101 different students. Assignment results were partitioned into two groups: scores above 80% (A/B) and scores below 80% (C/D/F). To investigate student behaviors that result in differing levels of achievement, all students who consistently received A/B scores and all students who consistently received C/D/F scores were removed from the dataset. A within-subjects comparison of the scores received by the remaining individuals was performed. Further, time and code-size data that is difficult to compare directly between different courses was normalized.This study revealed several significant results. When students received A/B scores, they started earlier and finished earlier than when the same students received C/D/F scores. They also wrote slightly more program code. They did not appear to spend any more time on their work, however. Approximately two-thirds of the A/B scores were received by individuals who started more than a day in advance of the deadline, while approximately two-thirds of the C/D/F scores were received by individuals who started on the last day or later. One possible explanation is that students who start earlier simply have more time to seek assistance when they get stuck. Stephen H. Edwards, Jason Snyder, Manuel A. Pérez-Quiñones, Anthony Allevato, Dongkwan Kim 0003, Betsy Tretola |
ICER | 1 |
| 2009 | Dereferee: exploring pointer mismanagement in student codeabstractDynamic memory management and the use of pointers are critical topics in teaching the C++ language. They are also some of the most difficult for students to grasp properly. The responsibility of ensuring that students understand these concepts does not end with the instructor's lectures---a library enhanced with diagnostics beyond those provided by the language's run-time system itself is a useful tool for giving students more detailed information when their code fails. Anthony Allevato, Stephen H. Edwards, Manuel A. Pérez-Quiñones |
SIGCSE | 2 |
| 2008 | Mining Data from an Automated Grading and Testing System by Adding Rich Reporting Capabilities
Anthony Allevato, Matthew Thornton, Stephen H. Edwards, Manuel A. Pérez-Quiñones |
EDM | 3 |
| 2008 | DCER: sharing empirical computer science education dataabstractData sharing is common, and sometimes even required, in other disciplines. Creating a mechanism for data sharing in computer science education research will benefit both individual researchers and the community. While it is easy to say that data sharing is desirable, it is much more difficult to make it a practical reality. Kate Sanders 0001, Brad Richards, Jan Erik Moström, Vicki L. Almstrum, Stephen H. Edwards, Sally Fincher, Katherine Gunion, Mark S. Hall, Brian Hanks, Stephen Lonergan, Robert McCartney, Briana B. Morrison, Jaime Spacco, Lynda Thomas |
ICER | 5 |
| 2008 | Web-CAT: automatically grading programming assignmentsabstractThis demonstration introduces participants to using Web-CAT, an open-source automated grading system. Web-CAT is customizable and extensible, allowing it to support a wide variety of programming languages and assessment strategies. Web-CAT is most well-known as the system that grades students on how well they test their own code, with experimental evidence that it offers greater learning benefits than more traditional output-comparison grading. Participants will learn how to set up courses, prepare reference tests, set up assignments, and allow graders to manually grade for design. Stephen H. Edwards, Manuel A. Pérez-Quiñones |
ITiCSE | 1 |
| 2008 | A data type to exploit online data sourcesabstractRecent work in developing student assignments has involved making use of online data resources to make them more interesting and to give students real world information to interact with in some manner. While definitely a practical approach, the work that has been done so far is either for "CS0" courses targeted at non-majors, often using tools like Microsoft Excel, or courses that require a level of skill at programming from the students. Additionally, existing tools are specific to a particular structure of the data (CSV, XML, and others). As a result, these constraints make on-line real-world data sets difficult to use in typical introductory programming courses for majors. Matthew Thornton, Stephen H. Edwards |
ITiCSE | 2 |
| 2008 | Supporting student-written tests of gui programsabstractTools like JUnit and its relatives are making software testing reachable even for introductory students. At the same time, however, many introductory computer sciences courses use graphical interfaces as an "attention grabber" for students and as a metaphor for teaching object-oriented programming. Unfortunately, developing software tests for programs that have significant graphical user interfaces is beyond the abilities of typical students (and, for that matter, many educators). This paper describes a framework for combining readily available tools to create an infrastructure for writing tests for Java programs that have graphical user interfaces. These tests are level-appropriate for introductory students and fit in with current approaches in computer science education that incorporate testing in programming assignments. An analysis of data collected during actual student use of the framework in a CS1 course is presented. Matthew Thornton, Stephen H. Edwards, Roy Patrick Tan, Manuel A. Pérez-Quiñones |
SIGCSE | 2 |
| 2008 | Misunderstandings about object-oriented design: experiences using code reviewsabstractIn this paper we present our experience using code reviews in a CS2 course. In particular, we highlight a series of misunderstandings of object-oriented (OO) concepts we observed as a by-product of the code review exercise. In our activity, we asked students to review code, rate it using a rubric, and to justify their explanation. The students were asked to review two solutions to a project from a previous year. Through examples of their explanations, we found that students had a number of basic misunderstandings of object-oriented principles. In this paper, we present our observations of the misunderstandings, and present some general observations of how code reviews can be used as an assessment tool in CS2. Scott A. Turner, Ricardo Quintana-Castillo, Manuel A. Pérez-Quiñones, Stephen H. Edwards |
SIGCSE | 4 |
| 2007 | It seemed like a good idea at the timeabstractNo abstract available. Jonas Boustedt, Robert McCartney, Josh Tenenberg, Titus Winters, Stephen H. Edwards, Briana B. Morrison, David R. Musicant, Ian Utting, Carol Zander |
SIGCSE | 5 |
| 2007 | Algorithm visualization: a report on the state of the fieldabstractWe present our findings on the state of the field of algorithm visualization, based on extensive search and analysis of links to hundreds of visualizations. We seek to answer questions such as how content is distributed among topics, who created algorithm visualizations and when, the overall quality of available visualizations, and how visualizations are disseminated. We have built a wiki that currently catalogs over 350 algorithm visualizations, contains the beginnings of an annotated bibliography on algorithm visualization literature, and provides information about researchers and projects. Unfortunately, we found that most existing algorithm visualizations are of low quality, and the content coverage is skewed heavily toward easier topics. There are no effective repositories or organized collections of algorithm visualizations currently available. Thus, the field appears in need of improvement in dissemination of materials, informing potential developers about what is needed, and propagating known best practices for creating new visualizations. Clifford A. Shaffer, Matthew Cooper 0003, Stephen H. Edwards |
SIGCSE | 3 |
| 2007 | A Flexible Strategy for Embedding and Configuring Run-Time Contract Checks in .Net ComponentsabstractIn component-based systems, there are several obstacles to using Design by Contract (DbC), particularly with respect to third-party components. Contracts are particularly valuable when debugging or testing composite software structures that include third-party components. However, existing approaches have critical weaknesses. First, existing approaches typically require a component's source code to be available if you wish to strip (or re-insert) checks. Second, documentation of the contract is either distributed separately from the component or embedded in the component's source code. Third, enabling and disabling specific kinds of checks on separate components from independent vendors can be a significant challenge. This paper describes an approach to representing contracts for .NET components using attributes. This contract information can be retrieved from the compiled component's metadata and used for many purposes. The paper also describes nContract, a tool that automatically generates run-time checks from embedded contracts. Such run-time checks can be generated and added to a system without requiring source code access or recompilation. Further, when checks for a given component are excluded, they impose no run-time overhead. Finally, a highly expressive, fine-grained mechanism for controlling user preferences about which specific checks are enabled or disabled is presented. Stephen H. Edwards, Westley Haggard |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2006 | Designing an adaptive learning module to teach software testingabstractAdaptive learning systems aim to precisely tailor education and training to the individual needs of learners. Such systems use an internal model of a user's current knowledge to adjust the navigational affordances and presentation order of material. The user model is incrementally built and updated as the user demonstrates mastery by completing exercises and tests. Designing courses that are delivered adaptively involves addressing many complexities. This paper describes experiences designing the first adaptive module in a series intended to teach software testing skills. Experiences in using the first module and a preliminary evaluation of its effectiveness are presented. Rahul Agarwal, Stephen H. Edwards, Manuel A. Pérez-Quiñones |
SIGCSE | 2 |
| 2005 | minimUML: A minimalist approach to UML diagramming for early computer science educationabstractIn introductory computer science courses, the Unified Modeling Language (UML) is commonly used to teach basic object-oriented design. However, there appears to be a lack of suitable software to support this task. Many of the available programs that support UML focus on developing code and not on enhancing learning. Programs designed for educational use sometimes have poor interfaces or are missing common and important features such as multiple selection and undo/redo. Hence the need for software that is tailored to an instructional environment and that has all the useful and needed functionality for that specific task. This is the purpose of minimUML. It provides a minimum amount of UML, just what is commonly used in beginning programming classes, and a simple, usable interface. In particular, minimUML is designed to support abstract design while supplying features for exploratory learning and error avoidance. It supports functionality that includes multiple selection, undo/redo, flexible printing, cut and paste, and drag and drop. In addition, it allows for the annotation of diagrams, through text or free-form drawings, so students can receive feedback on their work. minimUML was developed with the goals of supporting ease of use, of supporting novice students, and of requiring no prior training for its use. This article presents the rationale behind the minimUML design, a description of the tool, and the results of usability evaluations and student feedback on the use of the tool. Scott A. Turner, Manuel A. Pérez-Quiñones, Stephen H. Edwards |
ACM J. Educ. Resour. Comput. | 3 |
| 2005 | Model variables: cleanly supporting abstraction in design by contractabstractIn design by contract (DBC), assertions are typically written using program variables and query methods. The lack of separation between program code and assertions is confusing, because readers do not know what code is intended for use in the program and what code is only intended for specification purposes. This lack of separation also creates a potential runtime performance penalty, even when runtime assertion checks are disabled, due to both the increased memory footprint of the program and the execution of code maintaining that part of the program's state intended for use in specifications. To solve these problems, we present a new way of writing and checking DBC assertions without directly referring to concrete program states, using ‘model’, i.e. specification-only, variables and methods. The use of model variables and methods does not incur the problems mentioned above, but it also allow one to write more easily assertions that are abstract, concise, and independent of representation details, and hence more readable and maintainable. We implemented these features in the runtime assertion checker for the Java Modeling Language (JML), but the approach could also be implemented in other DBC tools. Copyright © 2005 John Wiley & Sons, Ltd. Yoonsik Cheon, Gary T. Leavens, Murali Sitaraman, Stephen H. Edwards |
Softw. Pract. Exp. | 4 |
| 2004 | Using software testing to move students from trial-and-error to reflection-in-actionabstractIntroductory computer science students rely on a trial and error approach to fixing errors and debugging for too long. Moving to a reflection in action strategy can help students become more successful. Traditional programming assignments are usually assessed in a way that ignores the skills needed for reflection in action, but software testing promotes the hypothesis-forming and experimental validation that are central to this mode of learning. By changing the way assignments are assessed--where students are responsible for demonstrating correctness through testing, and then assessed on how well they achieve this goal--it is possible to reinforce desired skills. Automated feedback can also play a valuable role in encouraging students while also showing them where they can improve. Stephen H. Edwards |
SIGCSE | 1 |
| 2004 | Contract-Checking Wrappers for C++ ClassesabstractTwo kinds of interface contract violations can occur in component-based software: A client component can fail to satisfy a requirement of a component it is using, or a component implementation can fail to fulfill its obligations to the client. The traditional approach to detecting and reporting such violations is to embed assertion checks into component source code, with compile-time control over whether they are enabled. This works well for the original component developers, but it fails to meet the needs of component clients who do not have access to source code for such components. A wrapper-based approach, in which contract checking is not hard-coded into the underlying component but is "layered" on top of it, offers several relative advantages. It is practical and effective for C++ classes. Checking code can be distributed in binary form along with the underlying component, it can be installed or removed without requiring recompilation of either the underlying component or the client code, it can be selectively enabled or disabled by the component client on a per-component basis, and it does not require the client to have access to any special tools (which might have been used by the component developer) to support wrapper installation and control. Experimental evidence indicates that wrappers in C++ impose-modest additional overhead compared to inlining assertion checks. Stephen H. Edwards, Murali Sitaraman, Bruce W. Weide, Joseph E. Hollingsworth |
IEEE Trans. Software Eng. | 1 |
| 2003 | Improving student performance by evaluating how well students test their own programsabstractStudents need to learn more software testing skills. This paper presents an approach to teaching software testing in a way that will encourage students to practice testing skills in many classes and give them concrete feedback on their testing performance, without requiring a new course, any new faculty resources, or a significant number of lecture hours in each course where testing will be practiced. The strategy is to give students basic exposure to test-driven development, and then provide an automated tool that will assess student submissions on-demand and provide feedback for improvement. This approach has been demonstrated in an undergraduate programming languages course using a prototype tool. The results have been positive, with students expressing appreciation for the practical benefits of test-driven development on programming assignments. Experimental analysis of student programs shows a 28% reduction in defects per thousand lines of code. Stephen H. Edwards |
ACM J. Educ. Resour. Comput. | 1 |
| 2001 | A framework for practical, automated black-box testing of component-based software
Stephen H. Edwards |
Softw. Test. Verification Reliab. | 1 |
| 2000 | Palette: A Reuse-Oriented Specification Language for Real-Time Systems
Binoy Ravindran, Stephen H. Edwards |
ICSR | 2 |
| 2000 | Black-box testing using flowgraphs: an experimental assessment of effectiveness and automation potentialabstractA black-box testing strategy based on Zweben et al.'s specification-based test data adequacy criteria is explored. The approach focuses on generating a flowgraph from a component's specification and applying analogues of white-box strategies to it. An experimental assessment of the fault-detecting ability of test sets generated using this approach was performed for three of Zweben et al.'s criteria using mutation analysis. By using precondition, postcondition and invariant checking wrappers around the component under test, fault detection ratios competitive with white-box techniques were achieved. Experience with a prototype test set generator used in the experiment suggests that practical automation may be feasible. Copyright © 2000 John Wiley & Sons, Ltd. Stephen H. Edwards |
Softw. Test. Verification Reliab. | 1 |
| 1998 | A framework for detecting interface violations in component-based softwareabstractTwo kinds of interface contract violations can occur in component based software: a client component may fail to satisfy a requirement of a component it is using, or a component implementation may fail to fulfil its obligations to the client. The paper proposes a systematic approach for detecting both kinds of violations, so that violation detection is not hard coded into base level components, but is "layered" on top of them, and so that it can be turned "on" or "off" selectively for one or more components, with practically no change to executable code (limiting changes to a few declarations). Among the salient features of this approach are its use of formal specifications, the ability to handle parameterized (i.e., generic, or template) components, and the automatic generation of routine aspects of violation detection. We have designed, built, and experimented with a generator of checking components for C++ templates. Stephen H. Edwards, Gulam Shakir, Murali Sitaraman, Bruce W. Weide, Joseph E. Hollingsworth |
ICSR | 1 |
| 1998 | Providing intellectual focus to CS1/CS2abstractFirst-year computer science students need to see clearly that computer science as a discipline has an important intellectual role to play and that it offers deep philosophical questions, much like the other hard sciences and mathematics; that CS is not "just programming". An appropriate intellectual focus for CS1/CS2 can be built on the foundations of systems thinking and mathematical modeling, as these principles are manifested in a component-based software paradigm. We outline some of the main technical features of this approach to CS1/CS2 and report preliminary observations from our experience with it. Timothy J. Long, Bruce W. Weide, Paolo Bucci, David S. Gibson, Joseph E. Hollingsworth, Murali Sitaraman, Stephen H. Edwards |
SIGCSE | 7 |
| 1997 | Representation Inheritance: A Safe Form of "White Box'' Code InheritanceabstractThere are two approaches to using code inheritance for defining new component implementations in terms of existing implementations. Black box code inheritance allows subclasses to reuse superclass implementations as-is, without direct access to their internals. Alternatively, white box code inheritance allows subclasses to have direct access to superclass implementation details, which may be necessary for the efficiency of some subclass operations and to prevent unnecessary duplication of code. Unfortunately, white box code inheritance violates the protection that encapsulation affords superclasses, opening up the possibility of a subclass interfering with the correct operation of its superclass methods. Representation inheritance is proposed as a restricted form of white box code inheritance where subclasses have direct access to superclass implementation details, but are required to respect the representation invariant(s) and abstraction relation(s) of their ancestor(s). This preserves the protection that encapsulation provides, while allowing the freedom of access that component implementers sometimes desire. Stephen H. Edwards |
IEEE Trans. Software Eng. | 1 |
| 1996 | Representation inheritance: a safe form of "white box" code inheritanceabstractThere are two approaches to using code inheritance for defining new component implementations in terms of existing implementations. Black-box code inheritance allows subclasses to reuse superclass implementations as-is, without direct access to their internals. Alternatively, white-box code inheritance allows subclasses to have direct access to superclass implementation details, which may be necessary for the efficiency of some subclass operations. Unfortunately, white-box code inheritance violates the protection that encapsulation affords to superclasses, opening up the possibility of a subclass interfering with the correct operation of its superclass methods. Representation inheritance is proposed as a restricted form of white-box code inheritance where subclasses have direct access to superclass implementation details, but are required to respect the representation invariant(s) and abstraction relation(s) of their ancestor(s). This preserves the protection that encapsulation provides, while allowing the freedom of access that component implementers sometimes desire. Stephen H. Edwards |
ICSR | 1 |
| 1996 | Characterizing observability and controllability of software componentsabstractTwo important objectives when designing a specification for a reusable software component are understandability and utility. For a typical component defining a new abstract data type, a significant common factor affecting both of these objectives is the choice of a mathematical model of the (state space of the) ADT, which is used to explain the behavior of the ADT's operations to potential clients. There are subtle connections between the expressiveness of this mathematical model and the functions computable using the operations provided with the ADT, giving rise to interesting issues involving the two complementary system theoretic principles of "observability" and "controllability". The paper discusses problems associated with formalizing intuitively stated observability and controllability principles in accordance with these tests. Although the example we use for illustration is simple, the analysis has implications for the design of reusable software components of every scale and conceptual complexity. Bruce W. Weide, Stephen H. Edwards, Wayne D. Heym, Timothy J. Long, William F. Ogden |
ICSR | 2 |
| 1995 | The Effects of Layering and Encapsulation on Software Development Cost and QualityabstractSoftware engineers often espouse the importance of using abstraction and encapsulation in developing software components. They advocate the "layering" of new components on top of existing components, using only information about the functionality and interfaces provided by the existing components. This layering approach is in contrast to a "direct implementation" of new components, utilizing unencapsulated access to the representation data structures and code present in the existing components. By increasing the reuse of existing components, the layering approach intuitively should result in reduced development costs, and in increased quality for the new components. However, there is no empirical evidence that indicates whether the layering approach improves developer productivity or component quality. We discuss three controlled experiments designed to gather such empirical evidence. The results support the contention that layering significantly reduces the effort required to build new components. Furthermore, the quality of the components, in terms of the number of defects introduced during their development, is at least as good using the layered approach. Experiments such as these illustrate a number of interesting and important issues in statistical analysis. We discuss these issues because, in our experience, they are not well known to software engineers.> Stuart H. Zweben, Stephen H. Edwards, Bruce W. Weide, Joseph E. Hollingsworth |
IEEE Trans. Software Eng. | 2 |
| 1994 | Design and Specification of Iterators Using the Swapping ParadigmabstractHow should iterators be abstracted and encapsulated in modern imperative languages? We consider the combined impact of several factors on this question: the need for a common interface model for user defined iterator abstractions, the importance of formal methods in specifying such a model, and problems involved in modular correctness proofs of iterator implementations and clients. A series of iterator designs illustrates the advantages of the swapping paradigm over the traditional copying paradigm. Specifically, swapping based designs admit more efficient implementations while offering relatively straightforward formal specifications and the potential for modular reasoning about program behavior. The final proposed design schema is a common interface model for an iterator for any generic collection.> Bruce W. Weide, Stephen H. Edwards, Douglas E. Harms, David Alex Lamb |
IEEE Trans. Software Eng. | 2 |
| 1993 | Common Interface Models for Reusable SoftwareabstractRelated reusable components are often based on different conceptual models of behavior. This may unduly restrict the ways in which they can be composed. The conceptual model underlying the specification of one module’s parameter requirements may differ significantly from the model underlying the specification of another module’s exported features, even if the two modules intuitively seem compatible. There is no well understood groundwork of common models for component interaction, and the lack of guidance for applying these models exacerbates the composability problem. This article describes how varying interface models and techniques for describing a component’s interface requirements affect composability. These problems are illustrated in the context of common interface properties that are exhibited even in simple ADT components. Stephen H. Edwards |
Int. J. Softw. Eng. Knowl. Eng. | 1 |