Jaromír Savelka

dblp:05/11075 · DBLP profile ↗
← Back
52ranked-venue papers
12as first author
40since 2021 · last 2026
0000-0002-3674-5456ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 9 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 23 · 4 first-author · 23 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 MCQ Difficulty Prediction via Modeling Learner Heterogeneity Using Data-Driven Cognitive Profiling
Dhriti Krishnan, Jaromír Savelka
AIED (3)2
2026 Toward Automated Curriculum Design: Ensuring Alignment Among LLM-generated Course Components
abstract
We introduce a system that leverages large language models (LLMs) to produce learning objectives, instructional text, assessments, pro­jects, and auto-graders from a minimal course description. Quality is ensured through a hierarchical generation strategy, schema enforcement with JSON function-calling, and iterative refinement via instructor and student feedback. Tested on a graduate-level data science course and delivered on an online platform with auto-graded coding projects, the system produced full modules in under an hour.
Karthik Mittal, Jaromír Savelka
ITiCSE (2)2
2026 Teaching Debugging in Cloud-Native Application Development: A Course Intervention Study
abstract
We report on a debugging intervention introduced in the Cloud Native course---part of the AI Technicians program, a collaboration between the U.S. Army's AI Integration Center (AI2C) and Carnegie Mellon University. Learners take an introductory programming course prior to Cloud Native, yet their debugging skills did not transfer to this more advanced context. Across two cohorts, students consistently struggled to identify root causes of issues they faced. We describe the intervention we designed to address this gap and the instruments we are using to evaluate its impact.
Michael Sambou, Can Kultur, Christopher Bogart, Jaromír Savelka
ITiCSE (2)4
2026 Changes in Coding Behavior and Performance Since the Introduction of LLMs
abstract
The widespread availability of large language models (LLMs) has changed how students engage with coding and problem-solving. While these tools may increase student productivity, they also make it more difficult for instructors to assess students’ learning and effort. In this quasi-longitudinal study, we analyze five years of student source code submissions in a graduate-level cloud computing course, focusing on an assignment that remained unchanged and examining students’ behavior during the period spanning five semesters before the release of ChatGPT and five semesters after.
Jaromír Savelka, Seth Goldstein, Michael Conway
LAK2
2025 Auto-Grader Feedback Utilization and Its Impacts: An Observational Study Across Five Community Colleges
abstract
Automated grading systems, or auto-graders, have become ubiquitous in programming education, and the way they generate feedback has become increasingly automated as well. However, there is insufficient evidence regarding auto-grader feedback's effectiveness in improving student learning outcomes, in a way that differentiates students who utilized the feedback and students who did not. In this study, we fill this critical gap. Specifically, we analyze students' interactions with auto-graders in an introductory Python programming course, offered at five community colleges in the United States. Our results show that students checking the feedback more frequently tend to get higher scores from their programming assignments overall. Our results also show that a submission that follows a student checking the feedback tends to receive a higher score than a submission that follows a student ignoring the feedback. Our results provide evidence on auto-grader feedback's effectiveness, encourage their increased utilization, and call for future work to continue their evaluation in this age of automation
Adam Zhang, Heather Burte, Jaromír Savelka, Christopher Bogart, Majd F. Sakr
CSEDU (1)3
2025 Are Students' Evaluations of Auto-Graders Biased by Their Grades?
Jaromír Savelka, Heather Burte, Christopher Bogart, Seth Copen Goldstein, Majd F. Sakr
EC-TEL (2)2
2025 Generating Legal Arguments with Automatically Identified Factor Magnitudes
abstract
Computational models of legal reasoning often employ factors to reason about cases. Factors can be used to analogize a current factual scenario and precedents and to make arguments for or against a conclusion. Courts not only determine whether a factor applies to a case or not, but often how strongly the factor applies, that is, the factor’s magnitude in the case. Previous methods for automatically extracting factors from cases cannot identify factors’ magnitudes. We present and evaluate a method employing Large Language Models (LLMs) to identify factor magnitudes using few-shot prompts with or without Wordnet definitions. We also show how the extracted magnitudes can be used in constructing legal arguments that employ factors and magnitudes the way judges and lawyers do.
Morgan A. Gray, Jaromír Savelka, Wesley M. Oliver, Kevin D. Ashley
ICAIL2
2025 Automated Mapping of Legal Criteria to the Texts of Adjudicatory Decisions Using LLMs
abstract
Legal reasoning, argumentation and decision making employ various structures, containing logically connected criteria when determining specific outcomes. Here, we examine whether large language models (LLMs) can automatically map decision texts to such criteria, by determining whether the decision maker found the criteria to be satisfied or not, and by providing explanations as to why the decision maker arrived at a particular decision. We introduce a Web-browser interface (CATLEX) to present the user with the LLM-generated determinations and explanations. We evaluate the ability of LLMs to generate such determinations and explanations, based on an experiment on 50 cases in the domain of military veterans’ disability appeals in the United States. A quantitative and qualitative analysis shows promising results, with the best model achieving an F1-score of 0.97 and explanations being generally accurate and useful. The results suggest new ways of reading decisions and developing legal arguments, with significant potential in, e.g., access to justice.
Hannes Westermann, Vern R. Walker, Jaromír Savelka
ICAIL3
2025 Show Me the Mastery Learning! Obstacles to Adoption and Opportunities for New Solutions
Claudio Alvarez, Nick Falkner, Päivi Kinnunen, Jaromír Savelka, Lisa Zhang 0003
ITiCSE (1)4
2025 Can GPT4 Generate Effective Feedback on Code Readability?
abstract
Effective feedback is often timely and consistent but, with large cohorts, this is not always achievable. This study explored the potential of GPT4 to generate feedback on code readability for students enrolled in a CS1 Java course. We developed rubrics based on three readability criteria: naming, commenting, and formatting. We defined feedback criteria and incorporated them into GPT4 prompts to guide feedback generation. Results were mixed: while some feedback messages closely aligned with the rubrics, offering valuable insights, others fell short in providing corrective guidance. This highlights the potential and limitations of using LLMs to generate feedback on code readability. Future research could refine these methods to improve feedback consistency and quality.
Xiaotian Su 0001, Yajie Song, Marcus Messer, Jaromír Savelka, Maria Cutumisu, April Yi Wang
ITiCSE (2)4
2025 Can LLMs Create Legally Relevant Summaries and Analyses of Videos?
abstract
Understanding the legally relevant factual basis of an event and conveying it through text is a key skill of legal professionals. This skill is important for preparing forms (e.g., insurance claims) or other legal documents (e.g., court claims), but often presents a challenge for laypeople. Current AI approaches aim to bridge this gap, but mostly rely on the user to articulate what has happened in text, which may be challenging for many. Here, we investigate the capability of large language models (LLMs) to understand and summarize events occurring in videos. We ask an LLM to summarize and draft legal letters, based on 120 YouTube videos showing legal issues in various domains. Overall, 71.7% of the summaries were rated as of high or medium quality, which is a promising result, opening the door to a number of applications in e.g. access to justice.
Lyra Hoeben-Kuil, Gijs van Dijck, Jaromír Savelka, Johanna Gunawan, Konrad Kollnig, Marta Kolacz, Mindy Duffourc, Shashank Chakravarthy, Hannes Westermann
JURIX3
2025 Do LLMs Truly "Understand" when a Precedent Is Overruled?
abstract
Large language models (LLMs) with extended context windows show promise for complex legal reasoning tasks, yet their ability to understand long legal documents remains insufficiently evaluated. Developing long-context benchmarks that capture realistic, high-stakes tasks remains a significant challenge in the field, as most existing evaluations rely on simplified synthetic tasks that fail to represent the complexity of real-world document understanding. Overruling relationships are foundational to common-law doctrine and commonly found in judicial opinions. They provide a focused and important testbed for long-document legal understanding that closely resembles what legal professionals actually do. We present an assessment of state-of-the-art LLMs on identifying overruling relationships from U.S. Supreme Court cases using a dataset of 236 case pairs. Our evaluation reveals three critical limitations: (1) era sensitivity – the models show degraded performance on historical cases compared to modern ones, revealing fundamental temporal bias in their training; (2) shallow reasoning – models rely on shallow logical heuristics rather than deep legal comprehension; and (3) context-dependent reasoning failures – models produce temporally impossible relationships in complex open-ended tasks despite maintaining basic temporal awareness in simple contexts. Our work contributes a benchmark that addresses the critical gap in realistic long-context evaluation, providing an environment that mirrors the complexity and stakes of actual legal reasoning tasks. The full dataset can be accessed at https://github.com/lizhang-AIandLaw/Do-LLMs-Truly-Understand-When-a-Precedent-Is-Overruled
Jaromír Savelka, Kevin D. Ashley
JURIX2
2025 Generating Effective Distractors for Introductory Programming Challenges: LLMs vs Humans
Mohammad Hassany, Peter Brusilovsky, Jaromír Savelka, Arun Balajiee Lekshmi Narayanan, Kamil Akhuseyinoglu, Arav Agarwal, Rully Agus Hendrawan
LAK3
2025 AI Technicians: Developing Rapid Occupational Training Methods for a Competitive AI Workforce
abstract
The accelerating pace of developments in Artificial Intelligence (AI) and the increasing role that technology plays in society necessitates substantial changes in the structure of the workforce. Besides scientists and engineers, there is a need for a very large workforce of competent AI technicians (i.e., maintainers, integrators) and users (i.e., operators). As traditional 4-year and 2-year degree-based education cannot fill this quickly opening gap, alternative training methods have to be developed. We present the results of the first four years of the AI Technicians program which is a unique collaboration between the U.S. Army's Artificial Intelligence Integration Center (AI2C) and Carnegie Mellon University to design, implement and evaluate novel rapid occupational training methods to create a competitive AI workforce at the technicians level. Through this multi-year effort we have already trained 59 AI Technicians. A key observation is that ongoing frequent updates to the training are necessary as the adoption of AI in the U.S. Army and within the society at large is evolving rapidly. A tight collaboration among the stakeholders from the army and the university is essential for successful development and maintenance of the training for the evolving role. Our findings can be leveraged by large organizations that face the challenge of developing a competent AI workforce as well as educators and researchers engaged in solving the challenge.
Jaromír Savelka, Can Kultur, Arav Agarwal, Christopher Bogart, Heather Burte, Adam Zhang, Majd F. Sakr
SIGCSE (1)1
2024 Examining the Trade-Offs Between Simplified and Realistic Coding Environments in an Introductory Python Programming Class
Huy Anh Nguyen, Christopher Bogart, Jaromír Savelka, Adam Zhang, Majd F. Sakr
EC-TEL (1)3
2024 Desirable Characteristics for AI Teaching Assistants in Programming Education
abstract
Providing timely and personalized feedback to large numbers of students is a long-standing challenge in programming courses. Relying on human teaching assistants (TAs) has been extensively studied, revealing a number of potential shortcomings. These include inequitable access for students with low confidence when needing support, as well as situations where TAs provide direct solutions without helping students to develop their own problem-solving skills. With the advent of powerful large language models (LLMs), digital teaching assistants configured for programming contexts have emerged as an appealing and scalable way to provide instant, equitable, round-the-clock support. Although digital TAs can provide a variety of help for programming tasks, from high-level problem solving advice to direct solution generation, the effectiveness of such tools depends on their ability to promote meaningful learning experiences. If students find the guardrails implemented in digital TAs too constraining, or if other expectations are not met, they may seek assistance in ways that do not help them learn. Thus, it is essential to identify the features that students believe make digital teaching assistants valuable. We deployed an LLM-powered digital assistant in an introductory programming course and collected student feedback ($n=813$) on the characteristics of the tool they perceived to be most important. Our results highlight that students value such tools for their ability to provide instant, engaging support, particularly during peak times such as before assessment deadlines. They also expressed a strong preference for features that enable them to retain autonomy in their learning journey, such as scaffolding that helps to guide them through problem-solving steps rather than simply being shown direct solutions.
Paul Denny 0001, Stephen MacNeil, Jaromír Savelka, Leo Porter 0001, Andrew Luxton-Reilly
ITiCSE (1)3
2024 Course Delivery Methods, Student Success, and Self-efficacy in Introductory Programming
abstract
Self-efficacy has been claimed to be a predictor of students' motivation and learning [1]. It has been found to be sensitive to students' success, and to affect their academic achievement. In the CS/IT education context, where the drop rates are high, it is important that students not only gain knowledge and skills, but also self-efficacy, so that they persist in the program. In this study, we investigate 602 students taking an introductory Python course via different delivery methods: (i) traditional in-person; (ii) cohort in-person; (iii) synchronous online; and (iv) asynchronous online. Although modality predicted retention and success, we found no apparent links among learning, student retention, and self-efficacy. However we found evidence that cohort learning may in particular help struggling students catch up with their peers.
Christopher Bogart, Can Kultur, Eric Keylor, Jaromír Savelka, Majd F. Sakr
ITiCSE (2)4
2024 Designing Modular Auto-graded Programming Projects
abstract
In this poster we propose an approach to designing auto-graded programming course projects that are modular and easily manageable by an instructor. Based on our experiences with the Sail() platform which supports auto-grading and feedback generation in multiple contexts, we design the approach to overcome the challenges we observed. The approach is especially focused on designing projects that can be utilized by multiple instructors who may have various scopes or students with varying backgrounds. The approach enables differentiated learning-thereby improving learning experiences and outcomes. We also discuss challenges of using such a modular approach to auto-graded projects.
Can Kultur, Jaromír Savelka, Christopher Bogart, Majd F. Sakr
ITiCSE (2)2
2024 How Instructors Incorporate Generative AI into Teaching Computing
abstract
Generative AI (GenAI) has seen great advancements in the past two years and the conversation around adoption is increasing. Widely available GenAI tools are disrupting classroom practices as they can write and explain code with minimal student prompting. While most acknowledge that there is no way to stop students from using such tools, a consensus has yet to form on how students should use them if they choose to do so. At the same time, researchers have begun to introduce new pedagogical tools that integrate GenAI into computing curricula. These new tools offer students personalized help or attempt to teach prompting skills without undercutting code comprehension. This working group aims to detail the current landscape of education-focused GenAI tools and teaching approaches, present gaps where new tools or approaches could appear, identify good practice-examples, and provide a guide for instructors to utilize GenAI as they continue to adapt to this new era.
James Prather, Juho Leinonen 0001, Natalie Kiesler, Jamie Gorson Benario, Sam Lau, Stephen MacNeil, Narges Norouzi, Simone Opel, Virginia Pettit, Leo Porter 0001, Brent N. Reeves, Jaromír Savelka, David H. Smith IV, Sven Strickroth, Daniel Zingaro
ITiCSE (2)12
2024 Using LLMs to Discover Legal Factors
abstract
Factors are a foundational component of legal analysis and computational models of legal reasoning. These factor-based representations enable lawyers, judges, and AI and Law researchers to reason about legal cases. In this paper, we introduce a methodology that leverages large language models (LLMs) to discover lists of factors that effectively represent a legal domain. Our method takes as input raw court opinions and produces a set of factors and associated definitions. We demonstrate that a semi-automated approach, incorporating minimal human involvement, produces factor representations that can predict case outcomes with moderate success, if not yet as well as expert-defined factors can.
Morgan A. Gray, Jaromír Savelka, Wesley M. Oliver, Kevin D. Ashley
JURIX2
2024 Robots in the Middle: Evaluating LLMs in Dispute Resolution
abstract
Mediation is a dispute resolution method featuring a neutral third-party (mediator) who intervenes to help the individuals resolve their dispute. In this paper, we investigate to what extent large language models (LLMs) are able to act as mediators. We investigate whether LLMs are able to analyze dispute conversations, select suitable intervention types, and generate appropriate intervention messages. Using a novel, manually created dataset of 50 dispute scenarios, we conduct a blind evaluation comparing LLMs with human annotators across several key metrics. Overall, the LLMs showed strong performance, even outperforming our human annotators across key dimensions. Specifically, in 62% of the cases, the LLMs chose intervention types that were rated as better than or equivalent to those chosen by humans. Moreover, in 84% of the cases, the intervention messages generated by the LLMs were rated as better than or equal to the intervention messages written by humans. LLMs likewise performed favourably on metrics such as impartiality, understanding and contextualization. Our results demonstrate the potential of integrating AI in online dispute resolution (ODR) platforms.
Jinzhe Tan, Hannes Westermann, Nikhil Reddy Pottanigari, Jaromír Savelka, Sébastien Meeùs, Mia Godet, Karim Benyekhlef
JURIX4
2024 Understanding the Role of Temperature in Diverse Question Generation by GPT-4
abstract
We conduct a preliminary study of the effect of GPT's temperature parameter on the diversity of GPT4-generated questions. We find that using higher temperature values leads to significantly higher diversity, with different temperatures exposing different types of similarity between generated sets of questions. We also demonstrate that diverse question generation is especially difficult for questions targeting lower levels of Bloom's Taxonomy.
Arav Agarwal, Karthik Mittal, Aidan Doyle, Pragnya Sridhar, Zipiao Wan, Jacob Doughty, Jaromír Savelka, Majd F. Sakr
SIGCSE (2)7
2024 What Factors Influence Persistence in Project-based Programming Courses at Community Colleges?
abstract
The rapid adoption of emergent technologies is creating significant shortfall in the CS/IT workforce. With not enough students in the educational pipeline to meet the forthcoming demand over the next decade, community colleges are making the effort to train confident, knowledgeable, and self-driven workers in this field. Project-based learning (PBL) has been shown to be effective for these ends, but it poses distinct challenges in resource-limited community college contexts since it may require more time, preparation, and motivation than other teaching modalities, from both the student and the instructor. We studied fifteen sections of an introductory project-based Python course taught at six community colleges, investigating several features of PBL theorized to be particular barriers to student persistence, particularly among women and other identities traditionally underrepresented in technical fields. We describe successes and challenges faced by students in these areas and suggest implications for project-based learning curriculum and platform design.
Christopher Bogart, Marshall An, Eric Keylor, Pawanjeet Singh, Jaromír Savelka, Majd F. Sakr
SIGCSE (1)5
2024 Assessing the Efficacy of Goal-Based Scenarios in Scaling AI Literacy for Non-Technical Learners
abstract
AI's pervasive role in various fields highlights the imperative for the workforce to adeptly leverage its potential. While numerous courses cater to developers, there exists a discernible void for the wider community of AI users. To address this, our study introduces 'AI User'-a suite of interactive modules hosted on the Sail() platform, designed specifically for non-technical individuals utilizing Goal-Based Scenario (GBS) learning. We conducted a controlled experiment to ascertain whether GBS offers superior learning gains in AI literacy compared to traditional deliberate practice using multiple choice questions.
Ying-Jui Tseng, Ruiwei Xiao, Christopher Bogart, Jaromír Savelka, Majd F. Sakr
SIGCSE (2)4
2023 Large Language Models (GPT) Struggle to Answer Multiple-Choice Questions About Code
Jaromír Savelka, Arav Agarwal, Christopher Bogart, Majd F. Sakr
CSEDU (2)1
2023 Automatic Identification and Empirical Analysis of Legally Relevant Factors
abstract
This research addresses how to automatically identify certain factors in the texts of legal decisions and analyze their role in courts' decisions. It focuses on drug interdiction auto stop cases in which courts decide whether police officers have reasonable suspicion to detain a motorist. It illustrates how the methods to identify factors automatically can support empirical legal research in the domain and how machine learning methods of different accuracy and interpretability can be harnessed to explain case outcomes in terms legal professionals can understand.
Morgan A. Gray, Jaromír Savelka, Wesley M. Oliver, Kevin D. Ashley
ICAIL2
2023 Unlocking Practical Applications in Legal Domain: Evaluation of GPT for Zero-Shot Semantic Annotation of Legal Texts
abstract
We evaluated the capability of a state-of-the-art generative pretrained transformer (GPT) model to perform semantic annotation of short text snippets (one to few sentences) coming from legal documents of various types. Discussions of potential uses (e.g., document drafting, summarization) of this emerging technology in legal domain have intensified, but to date there has not been a rigorous analysis of these large language models' (LLM) capacity in sentence-level semantic annotation of legal texts in zero-shot learning settings. Yet, this particular type of use could unlock many practical applications (e.g., in contract review) and research opportunities (e.g., in empirical legal studies). We fill the gap with this study. We examined if and how successfully the model can semantically annotate small batches of short text snippets (10-50) based exclusively on concise definitions of the semantic types. We found that the GPT model performs surprisingly well in zero-shot settings on diverse types of documents (F1 = .73 on a task involving court opinions, .86 for contracts, and .54 for statutes and regulations). These findings can be leveraged by legal scholars and practicing lawyers alike to guide their decisions in integrating LLMs in wide range of workflows involving semantic annotation of legal texts.
Jaromír Savelka
ICAIL1
2023 Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses
abstract
This paper studies recent developments in large language models’ (LLM) abilities to pass assessments in introductory and intermediate Python programming courses at the postsecondary level. The emergence of ChatGPT resulted in heated debates of its potential uses (e.g., exercise generation, code explanation) as well as misuses in programming classes (e.g., cheating). Recent studies show that while the technology performs surprisingly well on diverse sets of assessment instruments employed in typical programming classes the performance is usually not sufficient to pass the courses. The release of GPT-4 largely emphasized notable improvements in the capabilities related to handling assessments originally designed for human test-takers. This study is the necessary analysis in the context of this ongoing transition towards mature generative AI systems. Specifically, we report the performance of GPT-4, comparing it to the previous generations of GPT models, on three Python courses with assessments ranging from simple multiple-choice questions (no code involved) to complex programming projects with code bases distributed into multiple files (599 exercises overall). Additionally, we analyze the assessments that were not handled well by GPT-4 to understand the current limitations of the model, as well as its capabilities to leverage feedback provided by an auto-grader. We found that the GPT models evolved from completely failing the typical programming class’ assessments (the original GPT-3) to confidently passing the courses with no human involvement (GPT-4). While we identified certain limitations in GPT-4’s handling of MCQs and coding exercises, the rate of improvement across the recent generations of GPT models strongly suggests their potential to handle almost any type of assessment widely used in higher education programming courses. These findings could be leveraged by educators and institutions to adapt the design of programming assessments as well as to fuel the necessary discussions into how programming classes should be updated to reflect the recent technological developments. This study provides evidence that programming instructors need to prepare for a world in which there is an easy-to-use widely accessible technology that can be utilized by learners to collect passing scores, with no effort whatsoever, on what today counts as viable programming knowledge and skills assessments.
Jaromír Savelka, Arav Agarwal, Marshall An, Christopher Bogart, Majd F. Sakr
ICER (1)1
2023 Transformed by Transformers: Navigating the AI Coding Revolution for Computing Education: An ITiCSE Working Group Conducted by Humans
abstract
The recent advent of highly accurate and scalable large language models (LLMs) has taken the world by storm. From art to essays to computer code, LLMs are producing novel content that until recently was thought only humans could produce. Recent work in computing education has sought to understand the capabilities of LLMs for solving tasks such as writing code, explaining code, creating novel coding assignments, interpreting programming error messages, and more. However, these technologies continue to evolve at an astonishing rate leaving educators little time to adapt. This working group seeks to document the state-of-the-art for code generation LLMs, detail current opportunities and challenges related to their use, and present actionable approaches to integrating them into computing curricula.
James Prather, Paul Denny 0001, Juho Leinonen 0001, Brett A. Becker, Ibrahim Albluwi, Michael E. Caspersen, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, Andrew Luxton-Reilly, Stephen MacNeil, Andrew Petersen 0001, Raymond Pettit, Brent N. Reeves, Jaromír Savelka
ITiCSE (2)16
2023 Can Generative Pre-trained Transformers (GPT) Pass Assessments in Higher Education Programming Courses?
abstract
We evaluated the capability of generative pre-trained transformers (GPT), to pass assessments in introductory and intermediate Python programming courses at the postsecondary level. Discussions of potential uses (e.g., exercise generation, code explanation) and misuses (e.g., cheating) of this emerging technology in programming education have intensified, but to date there has not been a rigorous analysis of the models' capabilities in the realistic context of a full-fledged programming course with diverse set of assessment instruments. We evaluated GPT on three Python courses that employ assessments ranging from simple multiple-choice questions (no code involved) to complex programming projects with code bases distributed into multiple files (599 exercises overall). Further, we studied if and how successfully GPT models leverage feedback provided by an auto-grader. We found that the current models are not capable of passing the full spectrum of assessments typically involved in a Python programming course (<70% on even entry-level modules). Yet, it is clear that a straightforward application of these easily accessible models could enable a learner to obtain a non-trivial portion of the overall available score (>55%) in introductory and intermediate courses alike. While the models exhibit remarkable capabilities, including correcting solutions based on auto-grader's feedback, some limitations exist (e.g., poor handling of exercises requiring complex chains of reasoning steps). These findings can be leveraged by instructors wishing to adapt their assessments so that GPT becomes a valuable assistant for a learner as opposed to an end-to-end solution.
Jaromír Savelka, Arav Agarwal, Christopher Bogart, Yifan Song 0007, Majd F. Sakr
ITiCSE (1)1
2023 Using Large Language Models to Support Thematic Analysis in Empirical Legal Studies
abstract
Thematic analysis and other variants of inductive coding are widely used qualitative analytic methods within empirical legal studies (ELS). We propose a novel framework facilitating effective collaboration of a legal expert with a large language model (LLM) for generating initial codes (phase 2 of thematic analysis), searching for themes (phase 3), and classifying the data in terms of the themes (to kick-start phase 4). We employed the framework for an analysis of a dataset (n = 785) of facts descriptions from criminal court opinions regarding thefts. The goal of the analysis was to discover classes of typical thefts. Our results show that the LLM, namely OpenAI’s GPT-4, generated reasonable initial codes, and it was capable of improving the quality of the codes based on expert feedback. They also suggest that the model performed well in zero-shot classification of facts descriptions in terms of the themes. Finally, the themes autonomously discovered by the LLM appear to map fairly well to the themes arrived at by legal experts. These findings can be leveraged by legal researchers to guide their decisions in integrating LLMs into their thematic analyses, as well as other inductive coding projects.
Jakub Drápal, Hannes Westermann, Jaromír Savelka
JURIX3
2023 Can GPT Alleviate the Burden of Annotation?
abstract
Manual annotation is just as burdensome as it is necessary for some legal text analytic tasks. Given the promising performance of Generative Pretrained Transformers (GPT) on a number of different tasks in the legal domain, it is natural to ask if it can help with text annotation. Here we report a series of experiments using GPT-4 and GPT 3.5 as a pre-annotation tool to determine whether a sentence in a legal opinion describes a legal factor. These GPT models assign labels that human annotators subsequently confirm or reject. To assess the utility of pre-annotating sentences at scale, we examine the agreement among gold-standard annotations, GPT's pre-annotations, and law students' annotations. The agreements among these groups support that using GPT-4 as a pre-annotation tool is a useful starting point for large-scale annotation of factors.
Morgan A. Gray, Jaromír Savelka, Wesley M. Oliver, Kevin D. Ashley
JURIX2
2023 From Text to Structure: Using Large Language Models to Support the Development of Legal Expert Systems
abstract
Encoding legislative text in a formal representation is an important prerequisite to different tasks in the field of AI & Law. For example, rule-based expert systems focused on legislation can support laypeople in understanding how legislation applies to them and provide them with helpful context and information. However, the process of analyzing legislation and other sources to encode it in the desired formal representation can be time-consuming and represents a bottleneck in the development of such systems. Here, we investigate to what degree large language models (LLMs), such as GPT-4, are able to automatically extract structured representations from legislation. We use LLMs to create pathways from legislation, according to the JusticeBot methodology for legal decision support systems, evaluate the pathways and compare them to manually created pathways. The results are promising, with 60% of generated pathways being rated as equivalent or better than manually created ones in a blind comparison. The approach suggests a promising path to leverage the capabilities of LLMs to ease the costly development of systems based on symbolic approaches that are transparent and explainable.
Samyar Janatian, Hannes Westermann, Jinzhe Tan, Jaromír Savelka, Karim Benyekhlef
JURIX4
2022 Toward Automatically Identifying Legally Relevant Factors
abstract
In making legal decisions, courts apply relevant law to facts. While the law typically changes slowly over time, facts vary from case to case. Nevertheless, underlying patterns of fact may emerge. This research focuses on underlying fact patterns commonly present in cases where motorists are stopped for a traffic violation and subsequently detained while a police officer conducts a canine sniff of the vehicle for drugs. We present a set of underlying patterns of fact, that is, factors of suspicion, that police and courts apply in determining reasonable suspicion. We demonstrate how these fact patterns can be identified and annotated in legal cases and how these annotations can be employed to fine-tune a transformer model to identify the factors in previously unseen legal opinions.
Morgan A. Gray, Jaromír Savelka, Wesley M. Oliver, Kevin D. Ashley
JURIX2
2022 Toward an Intelligent Tutoring System for Argument Mining in Legal Texts
abstract
We propose an adaptive environment (CABINET) to support caselaw analysis (identifying key argument elements) based on a novel cognitive computing framework that carefully matches various machine learning (ML) capabilities to the proficiency of a user. CABINET supports law students in their learning as well as professionals in their work. The results of our experiments focused on the feasibility of the proposed framework are promising. We show that the system is capable of identifying a potential error in the analysis with very low false positives rate (2.0–3.5%), as well as of predicting the key argument element type (e.g., an issue or a holding) with a reasonably high F1-score (0.74).
Hannes Westermann, Jaromír Savelka, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
JURIX2
2021 Lex Rosetta: transfer of predictive models across languages, jurisdictions, and legal domains
abstract
In this paper, we examine the use of multi-lingual sentence embeddings to transfer predictive models for functional segmentation of adjudicatory decisions across jurisdictions, legal systems (common and civil law), languages, and domains (i.e. contexts). Mechanisms for utilizing linguistic resources outside of their original context have significant potential benefits in AI & Law because differences between legal systems, languages, or traditions often block wider adoption of research outcomes. We analyze the use of Language-Agnostic Sentence Representations in sequence labeling models using Gated Recurrent Units (GRUs) that are transferable across languages. To investigate transfer between different contexts we developed an annotation scheme for functional segmentation of adjudicatory decisions. We found that models generalize beyond the contexts on which they were trained (e.g., a model trained on administrative decisions from the US can be applied to criminal law decisions from Italy). Further, we found that training the models on multiple contexts increases robustness and improves overall performance when evaluating on previously unseen contexts. Finally, we found that pooling the training data from all the contexts enhances the models' in-context performance.
Jaromír Savelka, Hannes Westermann, Karim Benyekhlef, Charlotte Alexander, Jayla C. Grant, David Restrepo Amariles, Rajaa El Hamdani, Sébastien Meeùs, Aurore Clément Troussel, Michal Araszkiewicz, Kevin D. Ashley, Alexandra Ashley, Karl Branting, Mattia Falduti, Matthias Grabmair, Jakub Harasta, Tereza Novotná, Elizabeth Tippett, Shiwanni Johnson
ICAIL1
2021 Toward summarizing case decisions via extracting argument issues, reasons, and conclusions
abstract
In this paper, we assess the use of several deep learning classification algorithms as a step toward automatically preparing succinct summaries of legal decisions. Short case summaries that tease out the decision's argument structure by making explicit its issues, conclusions, and reasons (i.e., argument triples) could make it easier for the lay public and legal professionals to gain an insight into what the case is about. We have obtained a sizeable dataset of expert-crafted case summaries paired with full texts of the decisions issued by various Canadian courts. As the manual annotation of the full texts is prohibitively expensive, we explore various ways of leveraging the existing longer summaries which are much less time-consuming to annotate. We compare the performance of the systems trained on the annotations that are manually ported to the full texts from the summaries to the performance of the same systems trained on annotations that are projected from the summaries automatically. The results show the possibility of pursuing the automatic annotation in the future.
Jaromír Savelka, Kevin D. Ashley
ICAIL2
2021 Are Working Habits Different Between Well-Performing and at-Risk Students in Online Project-Based Courses?
abstract
We analyze differences in working habits between well-performing and at-risk students using highly-granular data collected from two semesters of an online project-based, upper-level course on cloud computing at a US institution of higher education. Such differentiating metrics may provide deeper insights than interim grades, which are oftentimes the only quantifiable data that is captured and available to an instructor as a proxy for students' learning. Interim grades provide little insight into students' broader work habits and may mask unsustainable learning strategies that result in shallow learning or quickly-forgotten skills/knowledge. The adoption of technology-enhanced learning tools for course delivery, automatic feedback, and grading enable data-informed insight and reflection into students' working habits. This data could allow the detection of early signs of under-prepared students or students in crisis. We empirically assess what working habits, if any, differ among well-performing and at-risk students. From clickstream and other activity data, we derive 22 metrics such as time spent reading project write-ups, timing of starting and finishing work, or break-taking. We also calculate two measures of consistency of each metric measured by a coefficient of variance and a variance of ranking over the semester as well as outlier behavior of a student. Using Z-test and Kolmogorov-Smirnov test, we confirm differences in multiple behavior patterns. Notably, our data suggest that well-performing students start and finish working on a project earlier than at-risk students but they also tend to have fewer submissions which indicate they are more thoughtful about feedback.
Mingxiao An, Jaromír Savelka, Christopher Bogart, Majd F. Sakr
ITiCSE (1)3
2021 Data-Centric Machine Learning: Improving Model Performance and Understanding Through Dataset Analysis
abstract
Machine learning research typically starts with a fixed data set created early in the process. The focus of the experiments is finding a model and training procedure that result in the best possible performance in terms of some selected evaluation metric. This paper explores how changes in a data set influence the measured performance of a model. Using three publicly available data sets from the legal domain, we investigate how changes to their size, the train/test splits, and the human labelling accuracy impact the performance of a trained deep learning classifier. Our experiments suggest that analyzing how data set properties affect performance can be an important step in improving the results of trained classifiers, and leads to better understanding of the obtained results.
Hannes Westermann, Jaromír Savelka, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
JURIX2
2021 Accounting for Sentence Position and Legal Domain Sentence Embedding in Learning to Classify Case Sentences
abstract
In this paper, we treat sentence annotation as a classification task. We employ sequence-to-sequence models to take sentence position information into account in identifying case law sentences as issues, conclusions, or reasons. We also compare the legal domain specific sentence embedding with other general purpose sentence embeddings to gauge the effect of legal domain knowledge, captured during pre-training, on text classification. We deployed the models on both summaries and full-text decisions. We found that the sentence position information is especially useful for full-text sentence classification. We also verified that legal domain specific sentence embeddings perform better, and that meta-sentence embedding can further enhance performance when sentence position information is included.
Jaromír Savelka, Kevin D. Ashley
JURIX2
2020 Sentence Embeddings and High-Speed Similarity Search for Fast Computer Assisted Annotation of Legal Documents
abstract
Human-performed annotation of sentences in legal documents is an important prerequisite to many machine learning based systems supporting legal tasks. Typically, the annotation is done sequentially, sentence by sentence, which is often time consuming and, hence, expensive. In this paper, we introduce a proof-of-concept system for annotating sentences “laterally.” The approach is based on the observation that sentences that are similar in meaning often have the same label in terms of a particular type system. We use this observation in allowing annotators to quickly view and annotate sentences that are semantically similar to a given sentence, across an entire corpus of documents. Here, we present the interface of the system and empirically evaluate the approach. The experiments show that lateral annotation has the potential to make the annotation process quicker and more consistent.
Hannes Westermann, Jaromír Savelka, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
JURIX2
2020 Using Argument Mining for Legal Text Summarization
abstract
Argument mining, a subfield of natural language processing and text mining, is a process of extracting argumentative text portions and identifying the role the selected texts play. Legal argument mining targets the argumentative parts of a legal text. In order to better understand how to apply legal argument mining as a step toward improving case summarization, we have assembled a sizeable set of cases and human-expert-prepared summaries annotated in terms of legal argument triples that capture the most important skeletal argument structures in a case. We report the results of applying multiple machine learning techniques to demonstrate and analyze the advantages and disadvantages of different methods to identify sentence components of these legal argument triples.
Jaromír Savelka, Kevin D. Ashley
JURIX2
2019 Improving Sentence Retrieval from Case Law for Statutory Interpretation
abstract
Statutory texts employ vague terms that are difficult to understand. Here we study and evaluate methods for retrieving useful sentences from court opinions that elaborate on the meaning of a vague statutory term. Retrieving sentences instead of whole cases may spare a user the need to review long lists of cases in search of useful explanations. We assembled a data set of 4,635 sentences that were responses to three statutory queries and labeled them in terms of their usefulness for interpretation. We have run a series of experiments on this data set, which we have made public, assessing different techniques to solve the task. These include techniques that measure the similarity between the sentence and the query, utilize the context of a sentence, expand queries, or assess the novelty of a sentence with respect to a statutory provision from which the interpreted term comes. Based on a detailed error analysis we propose a specialized sentence retrieval framework that mitigates the challenges of retrieving case law sentences for interpreting statutory terms. The results of evaluating different implementations of the framework are promising (.725 for NDGC at 10, .662 at 100).
Jaromír Savelka, Kevin D. Ashley
ICAIL1
2019 Computer-Assisted Creation of Boolean Search Rules for Text Classification in the Legal Domain
Hannes Westermann, Jaromír Savelka, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
JURIX2
2018 Segmenting U.S. Court Decisions into Functional and Issue Specific Parts
abstract
In common law jurisdictions, legal research often involves an analysis of relevant case law. Court opinions comprise several high-level parts with different functions. A statement's membership in one of the parts is a key factor influencing how the statement should be understood. In this paper we present a number of experiments in automatically segmenting court opinions into the functional and the issue specific parts. We defined a set of seven types including Background, Analysis, and Conclusions. We used the types to annotate a sizable corpus of US trade secret and cyber crime decisions. We used the data set to investigate the feasibility of recognizing the parts automatically. The proposed framework based on conditional random fields proved to be very promising in this respect. To support research in automatic case law analysis we plan to release the data set to the public.
Jaromír Savelka, Kevin D. Ashley
JURIX1
2017 Toward Linking Heterogenous References in Czech Court Decisions to Content
Jakub Harasta, Jaromír Savelka
JURIX2
2017 Detecting Agent Mentions in U.S. Court Decisions
abstract
Case law analysis is a significant component of research on almost any legal issue and understanding which agents are involved and mentioned in a decision is integral part of the analysis. In this paper we present a first experiment in detecting mentions of different agents in court decisions automatically. We defined a light-weight and easily extensible hierarchy of agents that play important roles in the decisions. We used the types from the hierarchy to annotate a corpus of US court decisions. The resulting data set enabled us to test the hypothesis that the mentions of agents in the decisions could be detected automatically. Conditional random fields models trained on the data set were shown to be very promising in this respect. To support research in automatic case-law analysis we release the agent mentions data set with this paper.
Jaromír Savelka, Kevin D. Ashley
JURIX1
2015 Transfer of predictive models for classification of statutory texts in multi-jurisdictional settings
abstract
In this paper we use statistical machine learning to classify statutory texts in terms of highly specific functional categories. We focus on regulatory provisions from multiple US state jurisdictions, all dealing with the same general topic of public health system emergency preparedness and response. In prior work we have established that one can improve classification performance on one jurisdiction's statutory texts using texts from another jurisdiction. Here we describe a framework facilitating transfer of predictive models for classification of statutory texts among multiple state jurisdictions. Our results show that the classification performance improves as we employ an increasing number of models trained on data coming from different states.
Jaromír Savelka, Kevin D. Ashley
ICAIL1
2015 Applying an Interactive Machine Learning Approach to Statutory Analysis
abstract
Statutory analysis is a significant component of research on almost any legal issue and determining if a statutory provision applies is an integral part of the analysis. In this paper we present the initial results from an attempt to support the applicability assessment in situations where the number of statutory provisions to be considered is large. We propose the use of a framework in which a single human expert cooperates with a machine learning text classification algorithm. Our experiments show that an adoption of the approach leads to a better performance during the relevance assessment. In addition, we suggest how to re-use a classification model trained during one statutory analysis for another related analysis. This points to a new way of capturing and re-using knowledge produced in the course of statutory analysis. Our experiments confirm the viability of this approach.
Jaromír Savelka, Gaurav Trivedi, Kevin D. Ashley
JURIX1
2014 Mining Information from Statutory Texts in Multi-Jurisdictional Settings
abstract
In this paper we mine statutory texts for highly-specific functional information using NLP techniques and a supervised ML approach. We focus on regulatory provisions from multiple state jurisdictions (Pennsylvania and Florida), all dealing with the same general topic (i.e., public health system emergency preparedness and response). While the number of annotated provisions from any one jurisdiction is not large, we are investigating whether one can improve classification performance on one jurisdiction's statutory texts by including other jurisdictions' annotated statutory texts dealing with the same general topic. Our experiments suggest that data from one jurisdiction can be used to boost the performance of the classifiers trained for different jurisdictions.
Jaromír Savelka, Matthias Grabmair, Kevin D. Ashley
JURIX1
2012 Refined Coherence as Constraint Satisfaction Framework for Representing Judicial Reasoning
abstract
In this paper we present a refined coherence as constraint satisfaction framework as a potent tool for representation of judicial reasoning. We demonstrate usefulness of the framework on a model of the famous Popov v Hayashi case. Although we do not claim that the presented framework can be already considered fully developed we believe that the account constitutes a major improvement over those that have been published previously. The resulting representation is strongly anchored in a raw text of the decision itself and by means of formal logic can be transformed to a graphical representation which is a surprisingly intuitive and transparent account of application of rules in legal cases.
Michal Araszkiewicz, Jaromír Savelka
JURIX2
2011 Two Methods for Representing Judicial Reasoning in the Framework of Coherence as Constraint Satisfaction
Michal Araszkiewicz, Jaromír Savelka
JURIX2