Hannes Westermann

dblp:245/3643 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0002-4527-7316ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2025 A Demonstration of a Semi-Structured Legal Reasoning Framework
abstract
Legal reasoning strikes a balance between structured elements (i.e. legal criteria and their connections) and unstructured elements (i.e. about of the applicability of open-textured legal terms). Expert systems perform well on the former task, while large language models offer significant potential in the latter style of reasoning. In this demonstration, we present the DALLMA framework, which enables the combination of these two approaches, for semi-structured reasoning. This approach has the potential to assist in many legal reasoning and drafting tasks, while reducing the risks of hallucinations, with important implications for e.g. access to justice.
Hannes Westermann
ICAIL1
2025 Automated Mapping of Legal Criteria to the Texts of Adjudicatory Decisions Using LLMs
abstract
Legal reasoning, argumentation and decision making employ various structures, containing logically connected criteria when determining specific outcomes. Here, we examine whether large language models (LLMs) can automatically map decision texts to such criteria, by determining whether the decision maker found the criteria to be satisfied or not, and by providing explanations as to why the decision maker arrived at a particular decision. We introduce a Web-browser interface (CATLEX) to present the user with the LLM-generated determinations and explanations. We evaluate the ability of LLMs to generate such determinations and explanations, based on an experiment on 50 cases in the domain of military veterans’ disability appeals in the United States. A quantitative and qualitative analysis shows promising results, with the best model achieving an F1-score of 0.97 and explanations being generally accurate and useful. The results suggest new ways of reading decisions and developing legal arguments, with significant potential in, e.g., access to justice.
Hannes Westermann, Vern R. Walker, Jaromír Savelka
ICAIL1
2025 Can LLMs Create Legally Relevant Summaries and Analyses of Videos?
abstract
Understanding the legally relevant factual basis of an event and conveying it through text is a key skill of legal professionals. This skill is important for preparing forms (e.g., insurance claims) or other legal documents (e.g., court claims), but often presents a challenge for laypeople. Current AI approaches aim to bridge this gap, but mostly rely on the user to articulate what has happened in text, which may be challenging for many. Here, we investigate the capability of large language models (LLMs) to understand and summarize events occurring in videos. We ask an LLM to summarize and draft legal letters, based on 120 YouTube videos showing legal issues in various domains. Overall, 71.7% of the summaries were rated as of high or medium quality, which is a promising result, opening the door to a number of applications in e.g. access to justice.
Lyra Hoeben-Kuil, Gijs van Dijck, Jaromír Savelka, Johanna Gunawan, Konrad Kollnig, Marta Kolacz, Mindy Duffourc, Shashank Chakravarthy, Hannes Westermann
JURIX9
2025 A Hohfeldian Knowledge Base for LLM-Assisted Legal Information Retrieval in Marine Biodiversity Law
abstract
Governance of marine genetic resources is fragmented across overlapping treaties, creating uncertainty about which obligations apply in specific situations. We address this by mapping treaty provisions into a normative structure based on Hohfelds framework, the Hohfeld-Structured Normative Knowledge Base (HSNKB), and by constructing a dataset of 15 fact-pattern questions with expert gold answers. We evaluate four recent large language models (LLMs) on retrieving rows that contain the relevant rules to answer the fact-pattern questions. The results indicate that reasoning LLMs achieved modest precision and middling recall. Hohfeldian representations help avoid false positives, but improving recall without degrading precision remains an open problem for cross-treaty retrieval.
Rohan Nanda, Henrique Marcos, Hannes Westermann, Julia Schutz Veiga
JURIX3
2024 AI in Healthcare: Navigating Legal Risk Assessment with JusticeBot
abstract
The increasing prevalence of technology in healthcare is leading to a corresponding rise in the complexity of the challenges facing hospitals. One such challenge is the difficulty of ensuring legal compliance for projects based on artificial intelligence and data processing. JusticeBot Laso is a tool designed to provide a comprehensive mapping of legal risks and a rating system to assist decision-makers in selecting AI projects. Our evaluation shows the tool improves risk identification accuracy from 20% to over 85%, and enhances risk scoring accuracy from 48% to 92%, compared to manual assessments.
Sébastien Meeùs, Valentina Dalla Giovanna, Samyar Janatian, Hannes Westermann, Karim Benyekhlef, Gregory Lewkowicz
JURIX4
2024 Getting in the Door: Streamlining Intake in Civil Legal Services with Large Language Models
abstract
Legal intake, the process of finding out if an applicant is eligible for help from a free legal aid program, takes significant time and resources. In part this is because eligibility criteria are nuanced, open-textured, and require frequent revision as grants start and end. In this paper, we investigate the use of large language models (LLMs) to reduce this burden. We describe a digital intake platform that combines logical rules with LLMs to offer eligibility recommendations, and we evaluate the ability of 8 different LLMs to perform this task. We find promising results for this approach to help close the access to justice gap, with the best model reaching an F1 score of .82, while minimizing false negatives.
Quinten Steenhuis, Hannes Westermann
JURIX2
2024 Robots in the Middle: Evaluating LLMs in Dispute Resolution
abstract
Mediation is a dispute resolution method featuring a neutral third-party (mediator) who intervenes to help the individuals resolve their dispute. In this paper, we investigate to what extent large language models (LLMs) are able to act as mediators. We investigate whether LLMs are able to analyze dispute conversations, select suitable intervention types, and generate appropriate intervention messages. Using a novel, manually created dataset of 50 dispute scenarios, we conduct a blind evaluation comparing LLMs with human annotators across several key metrics. Overall, the LLMs showed strong performance, even outperforming our human annotators across key dimensions. Specifically, in 62% of the cases, the LLMs chose intervention types that were rated as better than or equivalent to those chosen by humans. Moreover, in 84% of the cases, the intervention messages generated by the LLMs were rated as better than or equal to the intervention messages written by humans. LLMs likewise performed favourably on metrics such as impartiality, understanding and contextualization. Our results demonstrate the potential of integrating AI in online dispute resolution (ODR) platforms.
Jinzhe Tan, Hannes Westermann, Nikhil Reddy Pottanigari, Jaromír Savelka, Sébastien Meeùs, Mia Godet, Karim Benyekhlef
JURIX2
2023 JusticeBot: A Methodology for Building Augmented Intelligence Tools for Laypeople to Increase Access to Justice
abstract
Laypeople (i.e. individuals without legal training) may often have trouble resolving their legal problems. In this work, we present the JusticeBot methodology. This methodology can be used to build legal decision support tools, that support laypeople in exploring their legal rights in certain situations, using a hybrid case-based and rule-based reasoning approach. The system ask the user questions regarding their situation and provides them with legal information, references to previous similar cases and possible next steps. This information could potentially help the user resolve their issue, e.g. by settling their case or enforcing their rights in court. We present the methodology for building such tools, which consists of discovering typically applied legal rules from legislation and case law, and encoding previous cases to support the user. We also present an interface to build tools using this methodology and a case study of the first deployed JusticeBot version, focused on landlord-tenant disputes, which has been used by thousands of individuals.
Hannes Westermann, Karim Benyekhlef
ICAIL1
2023 Using Large Language Models to Support Thematic Analysis in Empirical Legal Studies
abstract
Thematic analysis and other variants of inductive coding are widely used qualitative analytic methods within empirical legal studies (ELS). We propose a novel framework facilitating effective collaboration of a legal expert with a large language model (LLM) for generating initial codes (phase 2 of thematic analysis), searching for themes (phase 3), and classifying the data in terms of the themes (to kick-start phase 4). We employed the framework for an analysis of a dataset (n = 785) of facts descriptions from criminal court opinions regarding thefts. The goal of the analysis was to discover classes of typical thefts. Our results show that the LLM, namely OpenAI’s GPT-4, generated reasonable initial codes, and it was capable of improving the quality of the codes based on expert feedback. They also suggest that the model performed well in zero-shot classification of facts descriptions in terms of the themes. Finally, the themes autonomously discovered by the LLM appear to map fairly well to the themes arrived at by legal experts. These findings can be leveraged by legal researchers to guide their decisions in integrating LLMs into their thematic analyses, as well as other inductive coding projects.
Jakub Drápal, Hannes Westermann, Jaromír Savelka
JURIX2
2023 From Text to Structure: Using Large Language Models to Support the Development of Legal Expert Systems
abstract
Encoding legislative text in a formal representation is an important prerequisite to different tasks in the field of AI & Law. For example, rule-based expert systems focused on legislation can support laypeople in understanding how legislation applies to them and provide them with helpful context and information. However, the process of analyzing legislation and other sources to encode it in the desired formal representation can be time-consuming and represents a bottleneck in the development of such systems. Here, we investigate to what degree large language models (LLMs), such as GPT-4, are able to automatically extract structured representations from legislation. We use LLMs to create pathways from legislation, according to the JusticeBot methodology for legal decision support systems, evaluate the pathways and compare them to manually created pathways. The results are promising, with 60% of generated pathways being rated as equivalent or better than manually created ones in a blind comparison. The approach suggests a promising path to leverage the capabilities of LLMs to ease the costly development of systems based on symbolic approaches that are transparent and explainable.
Samyar Janatian, Hannes Westermann, Jinzhe Tan, Jaromír Savelka, Karim Benyekhlef
JURIX2
2022 Conditional Abstractive Summarization of Court Decisions for Laymen and Insights from Human Evaluation
abstract
Legal text summarization is generally formalized as an extractive text summarization task applied to court decisions from which the most relevant sentences are identified and returned as a gist meant to be read by legal experts. However, such summaries are not suitable for laymen seeking intelligible legal information. In the scope of the JusticeBot, a question-answering system in French that provides information about housing law, we intend to generate summaries of court decisions that are, on the one hand, conditioned by a question-answer-decision triplet, and on the other hand, intelligible for ordinary citizens not familiar with legal documents. So far, our best model, a further pre-trained BARThez, achieves an average ROUGE-1 score of 37.7 and a deepened manual evaluation of summaries reveals that there is still room for improvement.
Olivier Salaün, Aurore Clément Troussel, Sylvain Longhais, Hannes Westermann, Philippe Langlais, Karim Benyekhlef
JURIX4
2022 Toward an Intelligent Tutoring System for Argument Mining in Legal Texts
abstract
We propose an adaptive environment (CABINET) to support caselaw analysis (identifying key argument elements) based on a novel cognitive computing framework that carefully matches various machine learning (ML) capabilities to the proficiency of a user. CABINET supports law students in their learning as well as professionals in their work. The results of our experiments focused on the feasibility of the proposed framework are promising. We show that the system is capable of identifying a potential error in the analysis with very low false positives rate (2.0–3.5%), as well as of predicting the key argument element type (e.g., an issue or a holding) with a reasonably high F1-score (0.74).
Hannes Westermann, Jaromír Savelka, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
JURIX1
2021 Lex Rosetta: transfer of predictive models across languages, jurisdictions, and legal domains
abstract
In this paper, we examine the use of multi-lingual sentence embeddings to transfer predictive models for functional segmentation of adjudicatory decisions across jurisdictions, legal systems (common and civil law), languages, and domains (i.e. contexts). Mechanisms for utilizing linguistic resources outside of their original context have significant potential benefits in AI & Law because differences between legal systems, languages, or traditions often block wider adoption of research outcomes. We analyze the use of Language-Agnostic Sentence Representations in sequence labeling models using Gated Recurrent Units (GRUs) that are transferable across languages. To investigate transfer between different contexts we developed an annotation scheme for functional segmentation of adjudicatory decisions. We found that models generalize beyond the contexts on which they were trained (e.g., a model trained on administrative decisions from the US can be applied to criminal law decisions from Italy). Further, we found that training the models on multiple contexts increases robustness and improves overall performance when evaluating on previously unseen contexts. Finally, we found that pooling the training data from all the contexts enhances the models' in-context performance.
Jaromír Savelka, Hannes Westermann, Karim Benyekhlef, Charlotte Alexander, Jayla C. Grant, David Restrepo Amariles, Rajaa El Hamdani, Sébastien Meeùs, Aurore Clément Troussel, Michal Araszkiewicz, Kevin D. Ashley, Alexandra Ashley, Karl Branting, Mattia Falduti, Matthias Grabmair, Jakub Harasta, Tereza Novotná, Elizabeth Tippett, Shiwanni Johnson
ICAIL2
2021 Data-Centric Machine Learning: Improving Model Performance and Understanding Through Dataset Analysis
abstract
Machine learning research typically starts with a fixed data set created early in the process. The focus of the experiments is finding a model and training procedure that result in the best possible performance in terms of some selected evaluation metric. This paper explores how changes in a data set influence the measured performance of a model. Using three publicly available data sets from the legal domain, we investigate how changes to their size, the train/test splits, and the human labelling accuracy impact the performance of a trained deep learning classifier. Our experiments suggest that analyzing how data set properties affect performance can be an important step in improving the results of trained classifiers, and leads to better understanding of the obtained results.
Hannes Westermann, Jaromír Savelka, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
JURIX1
2021 Extracting Facts from Case Rulings Through Paragraph Segmentation of Judicial Decisions
Andrés Lou, Olivier Salaün, Hannes Westermann, Leila Kosseim
NLDB3
2020 Sentence Embeddings and High-Speed Similarity Search for Fast Computer Assisted Annotation of Legal Documents
abstract
Human-performed annotation of sentences in legal documents is an important prerequisite to many machine learning based systems supporting legal tasks. Typically, the annotation is done sequentially, sentence by sentence, which is often time consuming and, hence, expensive. In this paper, we introduce a proof-of-concept system for annotating sentences “laterally.” The approach is based on the observation that sentences that are similar in meaning often have the same label in terms of a particular type system. We use this observation in allowing annotators to quickly view and annotate sentences that are semantically similar to a given sentence, across an entire corpus of documents. Here, we present the interface of the system and empirically evaluate the approach. The experiments show that lateral annotation has the potential to make the annotation process quicker and more consistent.
Hannes Westermann, Jaromír Savelka, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
JURIX1
2020 Analysis and Multilabel Classification of Quebec Court Decisions in the Domain of Housing Law
Olivier Salaün, Philippe Langlais, Andrés Lou, Hannes Westermann, Karim Benyekhlef
NLDB4
2019 Using Factors to Predict and Analyze Landlord-Tenant Decisions to Increase Access to Justice
abstract
This paper reports results from the JusticeBot Project, in which we analyzed two datasets drawn from 1 million written decisions from the Régie du logement du Québec. Using an empirical methodology, we identified 44 factors that occur in disputes where the tenant seeks a remedy due to problems with the rented apartment, such as the existence of bedbugs, high noise levels or problems with insulation. In the first dataset, we used these factors to tag 149 cases. We found a correlation between how many factors are found in a case and how likely the judge is to award rent reduction to a tenant; the amount of reduction was also higher in cases with more factors. For the second dataset (39 cases with bedbugs, drawn from the first dataset), we developed in-depth factors and used them to tag the cases. We found a number of plausible correlations, such as the average damage award being higher in cases with infestations of high intensity. Finally, in predicting the decision of the judge using the factors present in a case, the results were similar to the baselines or slightly above. We discuss the possible reasons for this, and why the approach shows promise in providing useful information to lay people and lawyers.
Hannes Westermann, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
ICAIL1
2019 Computer-Assisted Creation of Boolean Search Rules for Text Classification in the Legal Domain
Hannes Westermann, Jaromír Savelka, Vern R. Walker, Kevin D. Ashley, Karim Benyekhlef
JURIX1