EDBT 2026 Demo / reviewers in the wild / expert
Eleni Straitouri
dblp:302/4619
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 76% Language models and text generation · 12% Multi-agent systems · 12% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-AI interaction · 100% | |
| Theoretical computer science
1 paper |
Computational complexity · 100% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
2.4 | 4 | 2024 | Controlling Counterfactual Harm in Decision Support Systems Based on Prediction Sets · NeurIPS 2024 Designing Decision Support Systems using Counterfactual Prediction Sets · ICML 2024 Improving Expert Predictions with Conformal Prediction · ICML 2023 |
Human-AI interaction
decision support |
1.4 | 2 | 2024 | Designing Decision Support Systems using Counterfactual Prediction Sets · ICML 2024 Improving Expert Predictions with Conformal Prediction · ICML 2023 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Controlling Counterfactual Harm in Decision Support Systems Based on Prediction Sets · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction › human-AI decision making
human-AI complementarity |
0.8 | 1 | 2024 | Towards Human-AI Complementarity with Prediction Sets · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction sets |
0.8 | 1 | 2024 | Towards Human-AI Complementarity with Prediction Sets · NeurIPS 2024 |
Computational complexity
hardness of approximation |
0.8 | 1 | 2024 | Towards Human-AI Complementarity with Prediction Sets · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
conformal prediction · 4.4online learning · 1.5greedy algorithm · 1.5bandit algorithms · 1.5structural causal model · 0.8statistical framework · 0.8pairwise comparison · 0.8conformal risk control · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Designing Decision Support Systems using Counterfactual Prediction SetsabstractDecision support systems for classification tasks are predominantly designed to predict the value of the ground truth labels. However, since their predictions are not perfect, these systems also need to make human experts understand when and how to use these predictions to update their own predictions. Unfortunately, this has been proven challenging. In this context, it has been recently argued that an alternative type of decision support systems may circumvent this challenge. Rather than providing a single label prediction, these systems provide a set of label prediction values constructed using a conformal predictor, namely a prediction set, and forcefully ask experts to predict a label value from the prediction set. However, the design and evaluation of these systems have so far relied on stylized expert models, questioning their promise. In this paper, we revisit the design of this type of systems from the perspective of online learning and develop a methodology that does not require, nor assumes, an expert model. Our methodology leverages the nested structure of the prediction sets provided by any conformal predictor and a natural counterfactual monotonicity assumption to achieve an exponential improvement in regret in comparison to vanilla bandit algorithms. We conduct a large-scale human subject study ($n = 2{,}751$) to compare our methodology to several competitive baselines. The results show that, for decision support systems based on prediction sets, limiting experts’ level of agency leads to greater performance than allowing experts to always exercise their own agency. Eleni Straitouri, Manuel Gomez-Rodriguez |
ICML | 1 |
| 2024 | Prediction-Powered Ranking of Large Language ModelsabstractLarge language models are often ranked according to their level of alignment with human preferences---a model is better than other models if its outputs are more frequently preferred by humans. One of the popular ways to elicit human preferences utilizes pairwise comparisons between the outputs provided by different models to the same inputs. However, since gathering pairwise comparisons by humans is costly and time-consuming, it has become a common practice to gather pairwise comparisons by a strong large language model---a model strongly aligned with human preferences. Surprisingly, practitioners cannot currently measure the uncertainty that any mismatch between human and model preferences may introduce in the constructed rankings. In this work, we develop a statistical framework to bridge this gap. Given a (small) set of pairwise comparisons by humans and a large set of pairwise comparisons by a model, our framework provides a rank-set---a set of possible ranking positions---for each of the models under comparison. Moreover, it guarantees that, with a probability greater than or equal to a user-specified value, the rank-sets cover the true ranking consistent with the distribution of human pairwise preferences asymptotically. Using pairwise comparisons made by humans in the LMSYS Chatbot Arena platform and pairwise comparisons made by three strong large language models, we empirically demonstrate the effectivity of our framework and show that the rank-sets constructed using only pairwise comparisons by the strong large language models are often inconsistent with (the distribution of) human pairwise preferences. Ivi Chatzi, Eleni Straitouri, Suhas Thejaswi, Manuel Gomez-Rodriguez |
NeurIPS | 2 |
| 2024 | Controlling Counterfactual Harm in Decision Support Systems Based on Prediction SetsabstractDecision support systems based on prediction sets help humans solve multiclass classification tasks by narrowing down the set of potential label values to a subset of them, namely a prediction set, and asking them to always predict label values from the prediction sets. While this type of systems have been proven to be effective at improving the average accuracy of the predictions made by humans, by restricting human agency, they may cause harm---a human who has succeeded at predicting the ground-truth label of an instance on their own may have failed had they used these systems. In this paper, our goal is to control how frequently a decision support system based on prediction sets may cause harm, by design. To this end, we start by characterizing the above notion of harm using the theoretical framework of structural causal models. Then, we show that, under a natural, albeit unverifiable, monotonicity assumption, we can estimate how frequently a system may cause harm using only predictions made by humans on their own. Further, we also show that, under a weaker monotonicity assumption, which can be verified experimentally, we can bound how frequently a system may cause harm again using only predictions made by humans on their own. Building upon these assumptions, we introduce a computational framework to design decision support systems based on prediction sets that are guaranteed to cause harm less frequently than a user-specified value
using conformal risk control. We validate our framework using real human predictions from two different human subject studies and show that, in decision support systems based on prediction sets, there is a trade-off between accuracy and counterfactual harm. Eleni Straitouri, Suhas Thejaswi, Manuel Gomez-Rodriguez |
NeurIPS | 1 |
| 2024 | Towards Human-AI Complementarity with Prediction SetsabstractDecision support systems based on prediction sets have proven to be effective at helping human experts solve classification tasks. Rather than providing single-label predictions, these systems provide sets of label predictions constructed using conformal prediction, namely prediction sets, and ask human experts to predict label values from these sets. In this paper, we first show that the prediction sets constructed using conformal prediction are, in general, suboptimal in terms of average accuracy. Then, we show that the problem of finding the optimal prediction sets under which the human experts achieve the highest average accuracy is NP-hard. More strongly, unless P = NP, we show that the problem is hard to approximate to any factor less than the size of the label set. However, we introduce a simple and efficient greedy algorithm that, for a large class of expert models and non-conformity scores, is guaranteed to find prediction sets that provably offer equal or greater performance than those constructed using conformal prediction. Further, using a simulation study with both synthetic and real expert predictions, we demonstrate that, in practice, our greedy algorithm finds near-optimal prediction sets offering greater performance than conformal prediction. Giovanni De Toni, Nastaran Okati, Suhas Thejaswi, Eleni Straitouri, Manuel Gomez-Rodriguez |
NeurIPS | 4 |
| 2023 | Improving Expert Predictions with Conformal PredictionabstractAutomated decision support systems promise to help human experts solve multiclass classification tasks more efficiently and accurately. However, existing systems typically require experts to understand when to cede agency to the system or when to exercise their own agency. Otherwise, the experts may be better off solving the classification tasks on their own. In this work, we develop an automated decision support system that, by design, does not require experts to understand when to trust the system to improve performance. Rather than providing (single) label predictions and letting experts decide when to trust these predictions, our system provides sets of label predictions constructed using conformal prediction—prediction sets—and forcefully asks experts to predict labels from these sets. By using conformal prediction, our system can precisely trade-off the probability that the true label is not in the prediction set, which determines how frequently our system will mislead the experts, and the size of the prediction set, which determines the difficulty of the classification task the experts need to solve using our system. In addition, we develop an efficient and near-optimal search method to find the conformal predictor under which the experts benefit the most from using our system. Simulation experiments using synthetic and real expert predictions demonstrate that our system may help experts make more accurate predictions and is robust to the accuracy of the classifier the conformal predictor relies on. Eleni Straitouri, Lequn Wang, Nastaran Okati, Manuel Gomez-Rodriguez |
ICML | 1 |