Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Nour Jedidi

dblp:355/1243 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0007-0189-9678ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Question answering and dialogue systems · 50% Knowledge representation and reasoning · 25% Trustworthy machine learning · 25%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability › attribution methods
knowledge attribution
1.012026
NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering › knowledge-grounded question answering
open book question answering
1.012026
NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
parametric knowledge
1.012026
NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering › knowledge-grounded question answering
retrieval-augmented question answering
1.012026
NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026
Information retrieval › relevance feedback
pseudo-relevance feedback
1.012026
Revisiting BM25 Feedback Models using HyDE · SIGIR 2026
Information retrieval
query processing
1.012026
Revisiting BM25 Feedback Models using HyDE · SIGIR 2026
Information retrieval › evaluation › benchmark dataset
benchmark dataset construction
0.312026
NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026
Information retrieval
evaluation
0.312026
NanoKnow: How to Know What Your Language Model Knows · SIGIR 2026
Information retrieval › query reformulation
query expansion
0.312026
Revisiting BM25 Feedback Models using HyDE · SIGIR 2026

Methods — techniques the papers use, named apart from their topics

pre-training data partitioning · 2.0closed-book evaluation · 2.0large language model · 1.0HyDE · 1.0
YearPublicationVenuePosition
2026 NanoKnow: How to Know What Your Language Model Knows
abstract
How do large language models (LLMs) know what they know? Answering this question has been difficult because pre-training data is often a ''black box'' – unknown or inaccessible. The recent release of nanochat – a family of small LLMs with fully open pre-training data – addresses this as it provides a transparent view into where a model's parametric knowledge comes from. Towards the goal of understanding how knowledge is encoded by LLMs, we release NanoKnow, a benchmark dataset that partitions questions from Natural Questions and SQuAD into splits based on whether their answers are present in nanochat's pre-training corpus. Using these splits, we can now properly disentangle the sources of knowledge that LLMs rely on when producing an output. To demonstrate NanoKnow's utility, we conduct experiments using eight nanochat checkpoints. Our findings show: (1) closed-book accuracy is strongly influenced by answer frequency in the pre-training data, (2) providing external evidence can mitigate this frequency dependence, (3) even with external evidence, models are more accurate when answers were seen during pre-training, demonstrating that parametric and external knowledge are complementary, and (4) non-relevant information is harmful, with accuracy decreasing based on both the position and the number of non-relevant contexts. We release all NanoKnow artifacts at https://github.com/castorini/NanoKnow.
Lingwei Gu, Nour Jedidi, Jimmy Lin
SIGIR2
2026 Revisiting BM25 Feedback Models using HyDE
abstract
Recent approaches that leverage large language models (LLMs) for pseudo-relevance feedback (PRF) have generally not utilized well-established feedback models like Rocchio and RM3 when expanding queries for BM25. Instead, they often opt for a simple string concatenation of the query and LLM-generated expansion content. In this paper, we revisit and systematically evaluate traditional BM25 feedback models in the context of HyDE, a popular method that enriches query representations using LLM-generated hypothetical answer documents. Experiments demonstrate a mutually beneficial relationship between BM25 feedback models and HyDE. Feedback models are more effective when provided LLM-generated documents versus top-ranked documents retrieved by BM25, while HyDE benefits from feedback algorithms that select and weight expansion terms. In fact, by incorporating BM25 feedback models within the HyDE setup, we can further narrow the gap between BM25 and a strong ''single-shot'' dense retriever, without incurring the costs associated with such embedding-based retrieval methods. Ultimately, our results demonstrate that traditional BM25 feedback models can play an important role within modern LLM PRF methods and provide a simple approach to further enhance the accuracy of HyDE. Our code is available at https://github.com/nourj98/hyde-feedback.
Nour Jedidi, Jimmy Lin
SIGIR1
2025 Study on LLMs for Promptagator-Style Dense Retriever Training
abstract
Promptagator demonstrated that Large Language Models (LLMs) with few-shot prompts can be used as task-specific query generators for fine-tuning domain-specialized dense retrieval models. However, the original Promptagator approach relied on proprietary and large-scale LLMs which users may not have access to or may be prohibited from using with sensitive data. In this work, we study the impact of open-source LLMs at accessible scales (≤14B parameters) as an alternative. Our results demonstrate that open-source LLMs as small as 3B parameters can serve as effective Promptagator-style query generators. We hope our work will inform practitioners with reliable alternatives for synthetic data generation and give insights to maximize fine-tuning results for domain-specific applications. Our code is available at https://www.github.com/mitll/promptodile
Daniel Gwon, Nour Jedidi, Jimmy Lin
CIKM2