EDBT 2026 Demo / reviewers in the wild / expert
Nour Jedidi
dblp:355/1243
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0007-0189-9678ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Question answering and dialogue systems · 50% Knowledge representation and reasoning · 25% Trustworthy machine learning · 25% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability › attribution methods
knowledge attribution |
1.0 | 1 | 2026 | NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026 |
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering › knowledge-grounded question answering
open book question answering |
1.0 | 1 | 2026 | NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
parametric knowledge |
1.0 | 1 | 2026 | NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026 |
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering › knowledge-grounded question answering
retrieval-augmented question answering |
1.0 | 1 | 2026 | NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026 |
Information retrieval › relevance feedback
pseudo-relevance feedback |
1.0 | 1 | 2026 | Revisiting BM25 Feedback Models using HyDE · SIGIR 2026 |
Information retrieval
query processing |
1.0 | 1 | 2026 | Revisiting BM25 Feedback Models using HyDE · SIGIR 2026 |
Information retrieval › evaluation › benchmark dataset
benchmark dataset construction |
0.3 | 1 | 2026 | NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026 |
Information retrieval
evaluation |
0.3 | 1 | 2026 | NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026 |
Information retrieval › query reformulation
query expansion |
0.3 | 1 | 2026 | Revisiting BM25 Feedback Models using HyDE · SIGIR 2026 |
Methods — techniques the papers use, named apart from their topics
pre-training data partitioning · 2.0closed-book evaluation · 2.0large language model · 1.0HyDE · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NanoKnow: How to Know What Your Language Model KnowsabstractHow do large language models (LLMs) know what they know? Answering this question has been difficult because pre-training data is often a ''black box'' – unknown or inaccessible. The recent release of nanochat – a family of small LLMs with fully open pre-training data – addresses this as it provides a transparent view into where a model's parametric knowledge comes from. Towards the goal of understanding how knowledge is encoded by LLMs, we release NanoKnow, a benchmark dataset that partitions questions from Natural Questions and SQuAD into splits based on whether their answers are present in nanochat's pre-training corpus. Using these splits, we can now properly disentangle the sources of knowledge that LLMs rely on when producing an output. To demonstrate NanoKnow's utility, we conduct experiments using eight nanochat checkpoints. Our findings show: (1) closed-book accuracy is strongly influenced by answer frequency in the pre-training data, (2) providing external evidence can mitigate this frequency dependence, (3) even with external evidence, models are more accurate when answers were seen during pre-training, demonstrating that parametric and external knowledge are complementary, and (4) non-relevant information is harmful, with accuracy decreasing based on both the position and the number of non-relevant contexts. We release all NanoKnow artifacts at https://github.com/castorini/NanoKnow. Lingwei Gu, Nour Jedidi, Jimmy Lin |
SIGIR | 2 |
| 2026 | Revisiting BM25 Feedback Models using HyDEabstractRecent approaches that leverage large language models (LLMs) for pseudo-relevance feedback (PRF) have generally not utilized well-established feedback models like Rocchio and RM3 when expanding queries for BM25. Instead, they often opt for a simple string concatenation of the query and LLM-generated expansion content. In this paper, we revisit and systematically evaluate traditional BM25 feedback models in the context of HyDE, a popular method that enriches query representations using LLM-generated hypothetical answer documents. Experiments demonstrate a mutually beneficial relationship between BM25 feedback models and HyDE. Feedback models are more effective when provided LLM-generated documents versus top-ranked documents retrieved by BM25, while HyDE benefits from feedback algorithms that select and weight expansion terms. In fact, by incorporating BM25 feedback models within the HyDE setup, we can further narrow the gap between BM25 and a strong ''single-shot'' dense retriever, without incurring the costs associated with such embedding-based retrieval methods. Ultimately, our results demonstrate that traditional BM25 feedback models can play an important role within modern LLM PRF methods and provide a simple approach to further enhance the accuracy of HyDE. Our code is available at https://github.com/nourj98/hyde-feedback. Nour Jedidi, Jimmy Lin |
SIGIR | 1 |
| 2025 | Study on LLMs for Promptagator-Style Dense Retriever TrainingabstractPromptagator demonstrated that Large Language Models (LLMs) with few-shot prompts can be used as task-specific query generators for fine-tuning domain-specialized dense retrieval models. However, the original Promptagator approach relied on proprietary and large-scale LLMs which users may not have access to or may be prohibited from using with sensitive data. In this work, we study the impact of open-source LLMs at accessible scales (≤14B parameters) as an alternative. Our results demonstrate that open-source LLMs as small as 3B parameters can serve as effective Promptagator-style query generators. We hope our work will inform practitioners with reliable alternatives for synthetic data generation and give insights to maximize fine-tuning results for domain-specific applications. Our code is available at https://www.github.com/mitll/promptodile Daniel Gwon, Nour Jedidi, Jimmy Lin |
CIKM | 2 |