VLDB 2026 Research / reviewers in the wild / expert
Katherine Metcalf
dblp:141/6401
· DBLP profile ↗
13ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 34% Reinforcement learning · 29% Language models and text generation · 26% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
1.5 | 2 | 2024 | Hindsight PRIORs for Reward Learning from Human Preferences · ICLR 2024 Can You Rely on Synthetic Labellers in Preference-Based Reinforcement Learning? It's Complicated · AAAI 2024 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Aligning LLMs by Predicting Preferences from User Writing Samples · ICML 2025 |
Machine learning › Representation and self-supervised learning › representation matching › feature alignment › embedding alignment
cross-lingual alignment |
0.9 | 1 | 2025 | Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models · ACL (1) 2025 |
Machine learning › Representation and self-supervised learning
embedding space |
0.9 | 1 | 2025 | Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models · ACL (1) 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning › fairness
fairness evaluation |
0.9 | 1 | 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs · ICML 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs · ICML 2025 |
Natural language and speech › Language models and text generation
multilingual language models |
0.9 | 1 | 2025 | Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
personalization |
0.9 | 1 | 2025 | Aligning LLMs by Predicting Preferences from User Writing Samples · ICML 2025 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.9 | 1 | 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial natural language processing
adversarial prompt defense |
0.8 | 1 | 2024 | Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models · ICML 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment |
0.8 | 1 | 2024 | Hindsight PRIORs for Reward Learning from Human Preferences · ICLR 2024 |
Machine learning › Reinforcement learning
reward learning |
0.8 | 1 | 2024 | Hindsight PRIORs for Reward Learning from Human Preferences · ICLR 2024 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.8 | 1 | 2024 | Hindsight PRIORs for Reward Learning from Human Preferences · ICLR 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models · ICML 2024 |
Machine learning › Trustworthy machine learning
toxicity reduction |
0.8 | 1 | 2024 | Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models · ICML 2024 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction |
0.4 | 1 | 2019 | Unsupervised Hierarchical Temporal Abstraction by Simultaneously Learning Expectations and Representations · IJCAI 2019 |
Natural language and speech › Language models and text generation
in-context learning |
0.3 | 1 | 2025 | Aligning LLMs by Predicting Preferences from User Writing Samples · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
uncertainty calibration · 0.9preference verification · 0.9model interventions · 0.9iterative refinement · 0.9in-context learning · 0.9equalized odds · 0.9preference-based reinforcement learning · 0.8forward dynamics world model · 0.8auxiliary return redistribution · 0.8activation editing · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language ModelsabstractAnirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald, Skyler Seto, Masha Fedzechkina. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Anirudh S. Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald, Skyler Seto, Masha Fedzechkina |
ACL (1) | 3 |
| 2025 | Aligning LLMs by Predicting Preferences from User Writing SamplesabstractAccommodating human preferences is essential for creating aligned LLM agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs acting as writing agents to infer a description of user preferences. Agent alignment then comes from conditioning on the inferred preference description. However, existing methods often produce generic preference descriptions that fail to capture the unique and individualized nature of human preferences. This paper introduces PROSE, a method designed to enhance the precision of preference descriptions inferred from user writing samples. PROSE incorporates two key elements: (1) iterative refinement of inferred preferences, and (2) verification of inferred preferences across multiple user writing samples. We evaluate PROSE with several LLMs (i.e., Qwen2.5 7B and 72B Instruct, GPT-mini, and GPT-4o) on a summarization and an email writing task. We find that PROSE more accurately infers nuanced human preferences, improving the quality of the writing agent’s generations over CIPHER (a state-of-the-art method for inferring preferences) by 33%. Lastly, we demonstrate that ICL and PROSE are complementary methods, and combining them provides up to a 9% improvement over ICL alone. Code: https://github.com/apple/ml-predict Stephane Aroca-Ouellette, Natalie Mackraz, Barry-John Theobald, Katherine Metcalf |
ICML | 4 |
| 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMsabstractThe recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accuracy-based evaluations (i.e., prediction correctness), fail to capture the implicit impact of model uncertainty (e.g., higher model confidence about one group over another despite similar accuracy). To address this limitation, we propose an uncertainty-aware fairness metric, UCerf, to enable a fine-grained evaluation of model fairness that is more reflective of the internal bias in model decisions. Furthermore, observing data size, diversity, and clarity issues in current datasets, we introduce a new gender-occupation fairness evaluation dataset with 31,756 samples for co-reference resolution, offering a more diverse and suitable benchmark for modern LLMs. Combining our metric and dataset, we provide insightful comparisons of eight open-source LLMs. For example, Mistral-8B exhibits suboptimal fairness due to high confidence in incorrect predictions, a detail overlooked by Equalized Odds but captured by UCerF. Overall, this work provides a holistic framework for LLM evaluation by jointly assessing fairness and uncertainty, enabling the development of more transparent and accountable AI systems. Yinong Wang 0001, Nivedha Sivakumar, Falaah Arif Khan, Katherine Metcalf, Adam Golinski, Natalie Mackraz, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff |
ICML | 4 |
| 2024 | Can You Rely on Synthetic Labellers in Preference-Based Reinforcement Learning? It's ComplicatedabstractPreference-based Reinforcement Learning (PbRL) enables non-experts to train Reinforcement Learning models using preference feedback. However, the effort required to collect preference labels from real humans means that PbRL research primarily relies on synthetic labellers. We validate the most common synthetic labelling strategy by comparing against labels collected from a crowd of humans on three Deep Mind Control (DMC) suite tasks: stand, walk, and run. We find that: (1) the synthetic labels are a good proxy for real humans under some circumstances, (2) strong preference label agreement between human and synthetic labels is not necessary for similar policy performance, (3) policy performance is higher at the start of training from human feedback and is higher at the end of training from synthetic feedback, and (4) training on only examples with high levels of inter-annotator agreement does not meaningfully improve policy performance. Our results justify the use of synthetic labellers to develop and ablate PbRL methods, and provide insight into how human labelling changes over the course of policy training. Katherine Metcalf, Miguel Sarabia, Masha Fedzechkina, Barry-John Theobald |
AAAI | 1 |
| 2024 | Hindsight PRIORs for Reward Learning from Human PreferencesabstractPreference based Reinforcement Learning (PbRL) removes the need to hand specify a reward function by learning one from preference feedback over policy behaviors. Current approaches to PbRL do not address the credit assignment problem inherent in determining which parts of a behavior most contributed to a preference resulting in data intensive approaches and subpar reward models. We address such limitations by introducing a credit assignment strategy (PRIOR) that uses a forward dynamics world model to approximate state importance within a trajectory and then guides rewards to be proportional to state importance through an auxiliary predicted return redistribution objective. Incorporating state importance into reward learning improves the speed of policy learning, overall policy performance, and reward recovery on both locomotion and manipulation tasks. For example, PRIOR achieves 80% success rate with half the amount of data compared to baselines. The performance gains and our ablations demonstrate the benefits even a simple credit assignment strategy can have on reward learning and that state importance in forward dynamics prediction is a strong proxy for a state's contribution to a preference decision. Mudit Verma, Katherine Metcalf |
ICLR | 2 |
| 2024 | Whispering Experts: Neural Interventions for Toxicity Mitigation in Language ModelsabstractAn important issue with Large Language Models (LLMs) is their undesired ability to generate toxic language. In this work, we show that the neurons responsible for toxicity can be determined by their power to discriminate toxic sentences, and that toxic language can be mitigated by reducing their activation levels proportionally to this power. We propose AUROC adaptation (AurA), an intervention that can be applied to any pre-trained LLM to mitigate toxicity. As the intervention is proportional to the ability of each neuron to discriminate toxic content, it is free of any model-dependent hyperparameters. We show that AurA can achieve up to $2.2\times$ reduction in toxicity with only a $0.72$ perplexity increase. We also show that AurA is effective with models of different scale (from 1.5B to 40B parameters), and its effectiveness in mitigating toxic language, while preserving common-sense zero-shot abilities, holds across all scales. AurA can be combined with pre-prompting strategies, boosting its average mitigation potential from $1.28\times$ to $2.35\times$. Moreover, AurA can counteract adversarial pre-prompts that maliciously elicit toxic content, making it an effective method for deploying safer and less toxic models. Xavier Suau, Pieter Delobelle, Katherine Metcalf, Armand Joulin, Nicholas Apostoloff, Luca Zappella, Pau Rodríguez |
ICML | 3 |
| 2023 | On the Role of LIP Articulation in Visual Speech PerceptionabstractGenerating realistic lip motion from audio to simulate speech production is critical for driving natural character animation. Previous research has shown that traditional metrics used to optimize and assess models for generating lip motion from speech are not a good indicator of subjective opinion of animation quality. Devising metrics that align with subjective opinion first requires understanding what impacts human perception of quality. In this work, we focus on the degree of articulation and run a series of experiments to study how articulation strength impacts human perception of lip motion accompanying speech. Specifically, we study how increasing under-articulated (dampened) and over-articulated (exaggerated) lip motion affects human perception of quality. We examine the impact of articulation strength on human perception when considering only lip motion, where viewers are presented with talking faces represented by landmarks, and in the context of embodied characters, where viewers are presented with photo-realistic videos. Our results show that viewers prefer over-articulated lip motion consistently more than under-articulated lip motion and that this preference generalizes across different speakers and embodiments. Zakaria Aldeneh, Masha Fedzechkina, Skyler Seto, Katherine Metcalf, Miguel Sarabia, Nicholas Apostoloff, Barry-John Theobald |
ICASSP | 4 |
| 2019 | Unsupervised Hierarchical Temporal Abstraction by Simultaneously Learning Expectations and RepresentationsabstractThis paper presents ENHAnCE, an algorithm that simultaneously learns a predictive model of the input stream and generates representations of the concepts being observed. Following cognitively-inspired models of event segmentation, ENHAnCE uses expectation violations to identify boundaries between temporally extended patterns. It applies its expectation-driven process at multiple levels of temporal granularity to produce a hierarchy of predictive models that enable it to identify concepts at multiple levels of temporal abstraction. Evaluations show that the temporal abstraction hierarchies generated by ENHAnCE closely match hand-coded hierarchies for the test data streams. Given language data streams, ENHAnCE learns a hierarchy of predictive models that capture basic units of both spoken and written language: morphemes, lexemes, phonemes, syllables, and words. Katherine Metcalf, David B. Leake |
IJCAI | 1 |
| 2019 | Mirroring to Build Trust in Digital AssistantsabstractWe describe experiments towards building a conversational digital assistant that considers the preferred conversational style of the user. In particular, these experiments are designed to measure whether users prefer and trust an assistant whose conversational style matches their own. To this end we conducted a user study where subjects interacted with a digital assistant that responded in a way that either matched their conversational style, or did not. Using self-reported personality attributes and subjects' feedback on the interactions, we built models that can reliably predict a user's preferred conversational style. Katherine Metcalf, Barry-John Theobald, Garrett Weinberg, Robert Lee, Ing-Marie Jonsson, Russell Webb, Nicholas Apostoloff |
INTERSPEECH | 1 |
| 2018 | Embedded Word Representations for Rich Indexing: A Case Study for Medical Records
Katherine Metcalf, David B. Leake |
ICCBR | 1 |
| 2017 | Modelling Unsupervised Event Segmentation: Learning Event Boundaries from Prediction Errors
Katherine Metcalf, David B. Leake |
CogSci | 1 |
| 2016 | A Computational Method for Extracting, Representing, and Predicting Social ClosenessabstractIdentifying the social closeness between two individuals is a key skill in any situation where interpersonal relations play a role. Consequently, such capacity is needed to enable AI systems to understand human interactions and interact naturally in social situations. However, current research has relied on simple proxies for social closeness, and richer models are difficult to achieve. This paper presents an approach to predicting social closeness based on linguistic interactions and demonstrates an ability to predict social closeness with high accuracy. Katherine Metcalf, David B. Leake |
ECAI | 1 |
| 2013 | Classification of Regulatory Paragraphs by Discourse Structure, Reference Structure, and Regulation TypeabstractAs part of a larger effort to explore the feasibility of automating the translation of regulatory text to formal, executable rules, we have developed automated methods to classify regulatory paragraphs into various categories of interest, including regulation type and reference structure. We achieve F1 scores of at least .8 for most categories. Alan Buabuchachart, Katherine Metcalf, Nina Charness, Leora Morgenstern |
JURIX | 2 |