EDBT 2026 Demo / reviewers in the wild / expert
Guy Horowitz
dblp:322/0129
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0001-5093-7235ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 100% | |
| Artificial intelligence
4 papers |
Trustworthy machine learning · 82% Language models and text generation · 18% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
evaluation |
1.9 | 2 | 2026 | LiveRAG: A Diverse Q&A Dataset with Varying Difficulty Level for RAG Evaluation · SIGIR 2026 The LiveRAG Challenge at SIGIR 2025 · SIGIR 2025 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
1.4 | 2 | 2024 | Classification Under Strategic Self-Selection · ICML 2024 Causal Strategic Classification: A Tale of Two Shifts · ICML 2023 |
Machine learning › Trustworthy machine learning › performative prediction
strategic classification |
1.4 | 2 | 2024 | Classification Under Strategic Self-Selection · ICML 2024 Causal Strategic Classification: A Tale of Two Shifts · ICML 2023 |
Information retrieval › evaluation
benchmark dataset |
1.0 | 1 | 2026 | LiveRAG: A Diverse Q&A Dataset with Varying Difficulty Level for RAG Evaluation · SIGIR 2026 |
Information retrieval › evaluation › text generation evaluation
retrieval-augmented generation evaluation |
1.0 | 1 | 2026 | LiveRAG: A Diverse Q&A Dataset with Varying Difficulty Level for RAG Evaluation · SIGIR 2026 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | Do RAG Systems Really Suffer From Positional Bias? · EMNLP 2025 |
Information retrieval › evaluation › benchmark evaluation
shared task |
0.9 | 1 | 2025 | The LiveRAG Challenge at SIGIR 2025 · SIGIR 2025 |
Machine learning › Trustworthy machine learning
strategic behavior |
0.7 | 1 | 2023 | Causal Strategic Classification: A Tale of Two Shifts · ICML 2023 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.6 | 1 | 2022 | In the Eye of the Beholder: Robust Prediction with Causal User Modeling · NeurIPS 2022 |
Information retrieval › ranking
relevance prediction |
0.6 | 1 | 2022 | In the Eye of the Beholder: Robust Prediction with Causal User Modeling · NeurIPS 2022 |
Information retrieval
question answering |
0.6 | 2 | 2026 | LiveRAG: A Diverse Q&A Dataset with Varying Difficulty Level for RAG Evaluation · SIGIR 2026 The LiveRAG Challenge at SIGIR 2025 · SIGIR 2025 |
Information retrieval
retrieval-augmented generation |
0.6 | 2 | 2026 | LiveRAG: A Diverse Q&A Dataset with Varying Difficulty Level for RAG Evaluation · SIGIR 2026 The LiveRAG Challenge at SIGIR 2025 · SIGIR 2025 |
Information retrieval › ranking › text ranking
passage ranking |
0.3 | 1 | 2025 | Do RAG Systems Really Suffer From Positional Bias? · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
causal graph · 1.1bounded rationality modeling · 1.1item response theory · 1.0LLM-as-a-judge · 0.9strategic awareness · 0.8differentiable framework · 0.8end-to-end training · 0.7causal modeling · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiveRAG: A Diverse Q&A Dataset with Varying Difficulty Level for RAG EvaluationabstractWith Retrieval-Augmented Generation (RAG) becoming more and more prominent in generative AI solutions, there is an emerging need for systematically evaluating its effectiveness. We introduce the LiveRAG benchmark, a publicly available dataset of 895 synthetic questions and answers designed to support systematic evaluation of RAG-based Q&A systems. This synthetic benchmark is derived from the one used during the SIGIR'2025 LiveRAG challenge, where competitors were evaluated under strict time constraints. It is augmented with information that was not made available to competitors during the challenge, such as the ground-truth answers, together with their associated supporting claims which were used for evaluating competitors' answers. In addition, each question is associated with estimated difficulty and discriminability scores, derived from applying an Item Response Theory model to competitors' responses. Our analysis highlights the benchmark's question diversity, the wide range of difficulty levels, and their usefulness in differentiating between system capabilities. The LiveRAG benchmark will hopefully help the community advance RAG research, conduct systematic evaluation, and develop more robust Q&A systems. David Carmel, Simone Filice, Guy Horowitz, Yoelle Maarek, Alex Shtoff, Oren Somekh, Ran Tavory |
SIGIR | 3 |
| 2025 | Do RAG Systems Really Suffer From Positional Bias?abstractRetrieval Augmented Generation enhances LLM accuracy by adding passages retrieved from an external corpus to the LLM prompt.This paper investigates how positional bias-the tendency of LLMs to weight information differently based on its position in the promptaffects not only the LLM's capability to capitalize on relevant passages, but also its susceptibility to distracting passages.Through extensive experiments on three benchmarks, we show how state-of-the-art retrieval pipelines, while attempting to retrieve relevant passages, systematically bring highly distracting ones to the top ranks, with over 60% of queries containing at least one highly distracting passage among the top-10 retrieved passages.As a result, the impact of the LLM positional bias, which in controlled settings is often reported as very prominent by related works, is actually marginal in real scenarios since both relevant and distracting passages are, in turn, penalized.Indeed, our findings reveal that sophisticated strategies that attempt to rearrange the passages based on LLM positional preferences do not perform better than random shuffling. Florin Cuconasu, Simone Filice, Guy Horowitz, Yoelle Maarek, Fabrizio Silvestri |
EMNLP | 3 |
| 2025 | The LiveRAG Challenge at SIGIR 2025abstractThe LiveRAG Challenge at SIGIR 2025 provides a competitive platform for advancing Retrieval-Augmented Generation (RAG) technologies. Participants from academia and industry have been invited to build a RAG-based question answering system using a fixed corpus (Fineweb-10BT) and a common open-source LLM (Falcon3-10B-Instruct). The goal is to enable fair, focused comparisons on retrieval and prompting strategies. During the Live Challenge Day, the competing teams must provide answers and supportive information to 500 unseen questions within a strict two-hour window. Evaluation is conducted in two stages: automated LLM-as-a-judge scoring mechanism for correctness and faithfulness, followed by a manual review of top ranked submissions. The winners will be announced and prizes awarded during the LiveRAG Workshop at SIGIR 2025 in Padua, Italy. David Carmel, Simone Filice, Guy Horowitz, Yoelle Maarek, Oren Somekh, Ran Tavory |
SIGIR | 3 |
| 2024 | Classification Under Strategic Self-SelectionabstractWhen users stand to gain from certain predictive outcomes, they are prone to act strategically to obtain predictions that are favorable. Most current works consider strategic behavior that manifests as users modifying their features; instead, we study a novel setting in which users decide whether to even participate (or not), this in response to the learned classifier. Considering learning approaches of increasing strategic awareness, we investigate the effects of user self-selection on learning, and the implications of learning on the composition of the self-selected population. Building on this, we propose a differentiable framework for learning under self-selective behavior, which can be optimized effectively. We conclude with experiments on real data and simulated behavior that complement our analysis and demonstrate the utility of our approach. Guy Horowitz, Yonatan Sommer, Moran Koren, Nir Rosenfeld |
ICML | 1 |
| 2023 | Causal Strategic Classification: A Tale of Two ShiftsabstractWhen users can benefit from certain predictive outcomes, they may be prone to act to achieve those outcome, e.g., by strategically modifying their features. The goal in strategic classification is therefore to train predictive models that are robust to such behavior. However, the conventional framework assumes that changing features does not change actual outcomes, which depicts users as "gaming" the system. Here we remove this assumption, and study learning in a causal strategic setting where true outcomes do change. Focusing on accuracy as our primary objective, we show how strategic behavior and causal effects underlie two complementing forms of distribution shift. We characterize these shifts, and propose a learning algorithm that balances between these two forces and over time, and permits end-to-end training. Experiments on synthetic and semi-synthetic data demonstrate the utility of our approach. Guy Horowitz, Nir Rosenfeld |
ICML | 1 |
| 2022 | In the Eye of the Beholder: Robust Prediction with Causal User ModelingabstractAccurately predicting the relevance of items to users is crucial to the success of many social platforms. Conventional approaches train models on logged historical data; but recommendation systems, media services, and online marketplaces all exhibit a constant influx of new content---making relevancy a moving target, to which standard predictive models are not robust. In this paper, we propose a learning framework for relevance prediction that is robust to changes in the data distribution. Our key observation is that robustness can be obtained by accounting for \emph{how users causally perceive the environment}. We model users as boundedly-rational decision makers whose causal beliefs are encoded by a causal graph, and show how minimal information regarding the graph can be used to contend with distributional changes. Experiments in multiple settings demonstrate the effectiveness of our approach. Amir Feder, Guy Horowitz, Yoav Wald, Roi Reichart, Nir Rosenfeld |
NeurIPS | 2 |