Joel Coffman

dblp:51/2518 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
5since 2021 · last 2023
0000-0002-5500-4450ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Visual vs. Textual Programming Languages in CS0.5: Comparing Student Learning with and Student Perception of RAPTOR and Python
abstract
Much debate surrounds the choice of programming language for teaching computer science. Our institution's replacement of a visual programming language (RAPTOR) with a textual programming language (Python) provided a novel opportunity to explore the impacts of the programming language on students' learning and perception of programming. We conducted a randomized comparative study that involved 1083 students who took our introductory computing course in the 2019-2020 academic year. A unique aspect of our work stems from our course being a general education requirement; thus, our study includes students with a wide variety of backgrounds and majors. This report presents a comparison of student performance in each version of the course, including the impact of the programming language on underrepresented groups, and provides a summary of student feedback. Our results show that students in our introductory course performed similarly overall, but overwhelmingly perceived Python to be more valuable.
Joel Coffman, Adrian A. de Freitas, Justin M. Hill, Troy Weingart
SIGCSE (1)1
2023 FalconCode: A Multiyear Dataset of Python Code Samples from an Introductory Computer Science Course
abstract
The lack of large and diverse datasets of student code samples limits some forms of computer science education research. To address this problem, we created FalconCode, a novel collection of over 1.5 million Python programs from over two thousand undergraduate students at the United States Air Force Academy. FalconCode captures over five semesters worth of code samples from our introduction to computing course, which is taken by every student regardless of their academic major. The dataset contains student code submissions for over 800 programming assignments, as well as additional metadata such as the prompt for each assignment, the testcase(s) used to evaluate student submissions, and the specific skills needed to solve each problem. In this paper, we describe the methodology used to create FalconCode and the steps taken to anonymize the data. We then describe FalconCode's data schema, and show how it can support a wide range of research---including those utilizing machine learning (ML) and artificial intelligence (AI). FalconCode is provided free-of-charge, and is available upon request for computer science education research.
Adrian A. de Freitas, Joel Coffman, Michelle M. de Freitas, Justin C. Wilson, Troy Weingart
SIGCSE (1)2
2022 Good Students are Good Students Student Achievement with Visual versus Textual Programming
abstract
In this full research paper, we compare the impact of learning a visual versus textual programming language in an introductory computing course that is a general education requirement at our institution. We conducted a randomized comparative study with "experimental" sections that were taught using Python instead of RAPTOR, a flowchart-based programming language. The populations of students learning each programming language were similar with respect to gender, race, and predicted performance based upon standardized test scores and prior post-secondary education. Although students' performance on the whole was similar regardless of the programming language taught, predicted performance is correlated with SAT Math scores, grades in mathematics courses (specifically Calculus II), and, for lower-performing students, grades in other courses that satisfy general education requirements. That is, students from these groups who had lower predicted performance and learned Python performed worse on average than their peers who learned RAPTOR, and students with higher predicted performance outperformed (on average) their peers who learned RAPTOR. In addition, students' performance in subsequent computer science courses was not correlated with their performance and the language they learned in our introductory computing course. Our results raise important questions about the role of an introductory computing course in promoting equity and engaging students from historically underrepresented groups in computing fields.
Joel Coffman, Justin M. Hill, Shannon Beck, Adrian A. de Freitas, Troy Weingart
FIE1
2022 Introducing Software Development Process, Software Engineering, and Artificial Intelligence in a CS0.5 Course Project
abstract
Our CS0.5 course is required for all students and tasked to develop and assess the system development process proficiencies of an engineering-based institutional outcome. To achieve this tasking, we created a course-wide project that simulates NASA's Mars Ingenuity helicopter using an approach that emphasized our software system development process. With this project, our first-year college students created a 2-dimensional simulation of the Ingenuity helicopter flying through the thin Martian atmosphere with the goal of maximizing the area mapped subject to flight dynamics, available battery, landing proximity, and impact constraints. Students created their Ingenuity simulator using Python in three spirals: Spiral 1 - rendering of the simulation view with some initial movement, Spiral 2 - manual flight operations via thrust and roll keyboard inputs, and Spiral 3 - full auto-pilot. The students utilized a software system development process called "UDIT" (pronounced, "U Did IT") which stands for Understand - Design - Implement - Test. The assignment document was purposefully organized based on this process. The Understand and Design steps were presented via storyboards, enumerated requirements, a recommended structure chart, pseudocode, and suggested variables. As the Understand and Design steps address higher order objectives on Bloom's Taxonomy, we strived to model effective approaches for these steps. Most of our novice programmer's efforts involved the Implementation and Test steps emphasizing a build-a-little, test-a-little strategy. Forty percent of points come from testing via test procedures that the students created. The remaining points were earned based on code correctness and quality. The course also introduced the students to Artificial Intelligence to contribute to another proficiency of the engineering institutional outcome. For this, the project introduced students to genetic algorithms. They learned how the algorithm's parameters can be configured to train a more sophisticated version of the autopilot that needed to deal with additional Ingenuity features, including altitude-dependent mapping, as well as randomness in the form of varying winds at different altitudes. Currently being used with 450 students across 22 sections, the project is being assessed by sub-score tracking across the spirals; students' self-assessments of learning, interest, and self-efficacy; and collection of instructors' experiences and perceptions on the project.
Steven M. Hadfield, Alexander C. Roosma, Adrian A. de Freitas, Kimberly A. Braun, Steven Fulton, Joel Coffman, David T. Merritt, Kenneth R. Sample, Justin C. Wilson, Bobby D. Birrer
ITiCSE (2)6
2021 Election Security in the Cloud: A CTF Activity to Teach Cloud and Web Security
abstract
In this innovative practice work in progress (WIP) paper, we present a novel capture the flag (CTF) activity to teach students about the potential pitfalls and consequences of cloud misconfiguration. While cloud computing has proved an attractive option in terms of pricing, availability, and scalability, potential cloud consumers must equally weigh the security concerns of a cloud environment. The real-world consequences of misconfigurations are self-evident; cloud consuming companies that suffer a misconfiguration-related breach lose data, time, money, and trust from their customers. However, breaches due to misconfiguration are common, and this prevalence starts with inadequate education. Existing resources in cloud computing courses do not provide sufficient urgency, depth, or engagement when covering cloud security. Consequently, we created a CTF activity that has students pose as malicious actors who seek to compromise an election application running on a cloud environment. We believe that students who complete our CTF activity will have a deeper understanding of the potential pitfalls and consequences of cloud misconfiguration and a better understanding of how to protect against such issues in their own applications, and we are currently evaluating the extent to which our CTF activity achieves these goals.
Zachary Romano, Jennifer Windsor, Mathew VanDerPol, Joel Coffman
FIE4
2018 Low-Cost Distributed Key Management
abstract
Key management is one of the biggest challenges in cryptography. Traditionally, organizations stored cryptographic keys using file-based storage, which is insecure due to the lack of sufficient authentication. To overcome this issue, industry has moved towards using Hardware Security Modules (HSMs) for storing cryptographic keys. However, storing keys on HSMs does not ensure high availability if they fail due to network outages or lack of sufficient resources. Major cloud offerings provide high-availability key management solutions, but their cost may be prohibitively high for small-and mid-sized organizations. In this paper, we propose a system that combines distributed object storage with Trusted Platform Modules (TPMs) to ensure secure storage of keys, high availability of sensitive data, and ease of deployment. We envision this system as an attractive alternative for key management in private and public cloud settings.
Venkatesh Gopal, Shikha Fadnavis, Joel Coffman
SERVICES3
2017 Data Protection in OpenStack
abstract
As cloud computing becomes increasingly pervasive, it is critical for cloud providers to support basic security controls. Although major cloud providers tout such features, relatively little is known in many cases about their design and implementation. In this paper, we describe several security features in OpenStack, a widely-used, open source cloud computing platform. Our contributions to OpenStack range from key management and storage encryption to guaranteeing the integrity of virtual machine (VM) images prior to boot. We describe the design and implementation of these features in detail and provide a security analysis that enumerates the threats that each mitigates. Our performance evaluation shows that these security features have an acceptable cost-in some cases, within the measurement error observed in an operational cloud deployment. Finally, we highlight lessons learned from our real-world development experiences from contributing these features to OpenStack as a way to encourage others to transition their research into practice.
Bruce Benjamin, Joel Coffman, Hadi Esiely-Barrera, Kaitlin Farr, Dane Fichter, Daniel Genin, Laura Glendenning, Peter A. Hamilton, Shaku Harshavardhana, Rosalind Hom, Brianna Poulos, Nathan Reller
CLOUD2
2014 An Empirical Performance Evaluation of Relational Keyword Search Techniques
abstract
Extending the keyword search paradigm to relational data has been an active area of research within the database and IR community during the past decade. Many approaches have been proposed, but despite numerous publications, there remains a severe lack of standardization for the evaluation of proposed search techniques. Lack of standardization has resulted in contradictory results from different evaluations, and the numerous discrepancies muddle what advantages are proffered by different approaches. In this paper, we present the most extensive empirical performance evaluation of relational keyword search techniques to appear to date in the literature. Our results indicate that many existing search techniques do not provide acceptable performance for realistic retrieval tasks. In particular, memory consumption precludes many search techniques from scaling beyond small data sets with tens of thousands of vertices. We also explore the relationship between execution time and factors varied in previous evaluations; our analysis indicates that most of these factors have relatively little impact on performance. In summary, our work confirms previous claims regarding the unacceptable performance of these search techniques and underscores the need for standardization in evaluations--standardization exemplified by the IR community.
Joel Coffman, Alfred C. Weaver
IEEE Trans. Knowl. Data Eng.1
2011 Learning to rank results in relational keyword search
abstract
Keyword search within databases has become a hot topic within the research community as databases store increasing amounts of information. Users require an effective method to retrieve information from these databases without learning complex query languages (viz. SQL). Despite the recent research interest, performance and search effectiveness have not received equal attention, and scoring functions in particular have become increasingly complex while providing only modest benefits with regards to the quality of search results. An analysis of the factors appearing in existing scoring functions suggests that some factors previously deemed critical to search effectiveness are at best loosely correlated with relevance. We consider a number of these different scoring factors and use machine learning to create a new scoring function that provides significantly better results than existing approaches. We simplify our scoring function by systematically removing the factors with the lowest weight and show that this version still outperforms the previous state-of-the-art in this area.
Joel Coffman, Alfred C. Weaver
CIKM1
2010 A framework for evaluating database keyword search strategies
abstract
With regard to keyword search systems for structured data, research during the past decade has largely focused on performance. Researchers have validated their work using ad hoc experiments that may not reflect real-world workloads. We illustrate the wide deviation in existing evaluations and present an evaluation framework designed to validate the next decade of research in this field. Our comparison of 9 state-of-the-art keyword search systems contradicts the retrieval effectiveness purported by existing evaluations and reinforces the need for standardized evaluation. Our results also suggest that there remains considerable room for improvement in this field. We found that many techniques cannot scale to even moderately-sized datasets that contain roughly a million tuples. Given that existing databases are considerably larger than this threshold, our results motivate the creation of new algorithms and indexing techniques that scale to meet both current and future workloads.
Joel Coffman, Alfred C. Weaver
CIKM1
2010 Electronic commerce virtual laboratory
abstract
Website security is essential for successful e-commerce ventures, but the vital "how-to" components of security are often lacking in academic courses. This paper describes our attempt to instill an awareness of security concerns and techniques by having the students develop an Artist eXchange website, a social networking site that permits the posting and sharing of pictures, music, and text, including an end-user rating system. The six-homework set progresses through HTML, JavaScript, PHP, MySQL, file uploads, and security testing. An innovative feature is that each assignment is evaluated via automated testing, which guides the student toward detecting and correcting mistakes, especially with regard to common attack vectors.
Joel Coffman, Alfred C. Weaver
SIGCSE1
2007 Generalizing parametric timing analysis
abstract
In the design of real-time and embedded systems, it is important to establish a bound on the worst-case execution time (WCET) of programs to assure via schedulability analysis that deadlines are not missed. Static WCET analysis is performed by a timing analysis tool. This paper describes novel improvements to such a tool, allowing parametric timing analysis to be performed. Parametric timing analyzers receive an upper bound on the number of loop iterations in terms of an expression which is used to create a parametric formula. This parametric formula is later evaluated to determine the WCET based on input values only known at runtime. Effecting a transformation from a numeric to a parametric timing analyzer requires two innovations: 1) a summation solver capable of summation non-constant expressions and 2) a polynomial data structure which can replace integers as the basis for all calculations. Both additions permit other methods of analysis (e.g. caching, pipeline, constraint) to occur simultaneously. Combining these techniques allows our tool to statically bound the WCET for a larger class of benchmarks.
Joel Coffman, Christopher A. Healy, Frank Mueller 0001, David B. Whalley
LCTES1