VLDB 2026 Research / reviewers in the wild / expert
Maria Movin
dblp:345/7846
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0001-5759-7846ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › ranking
learning to rank |
0.9 | 1 | 2025 | Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025 |
Information retrieval › retrieval models › neural retrieval › neural ranking model
LLM-based ranking |
0.9 | 1 | 2025 | Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025 |
Information retrieval
reranking |
0.9 | 1 | 2025 | Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025 |
Information retrieval › reranking
zero-shot re-ranking |
0.9 | 1 | 2025 | Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025 |
Information retrieval
ranking |
0.3 | 1 | 2025 | Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
large language model prompting · 0.9LambdaMART · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | You Say Search, I Say Recs: A Scalable Agentic Approach to Query Understanding and Exploratory Search at SpotifyabstractOn online content platforms, users often aim to explore the catalog and discover new, personalized content through exploratory searches-such as "new releases for me." Traditional search systems, which prioritize lexical and semantic matching over personalized retrieval, have historically struggled to support this type of intent.In contrast, recommendation services that leverage user-item and item-item signals tend to be more effective for addressing exploratory queries.Agentic technologies offer a promising opportunity to enhance exploratory search by harnessing large language models (LLMs) to interpret complex query intents and route them to the most suitable downstream services.However, deploying such Enrico Palumbo, Marcus Isaksson, Alexandre Tamborrino, Maria Movin, Catalin Dincu, Ali Vardasbi, Lev Nikeshkin, Oksana Gorobets, Anders Nyman, Poppy Newdick, Hugues Bouchard, Paul N. Bennett, Mounia Lalmas-Roelleke, Dani Doro, Christine Doig Cardet, Ziad Sultan |
RecSys | 4 |
| 2025 | Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and LimitationsabstractLLMs have been explored for their use in IR as end-to-end rankers, rerankers and assessors. Recently, the exploration of the prompt-and-predict paradigm for reranking in combination with highly performant LLMs have drawn the attention of researchers. Instead of training or fine-tuning a reranker, LLMs are prompted in a zero-shot manner to produce relevance scores, pairwise preferences, or reranked lists. Existing research, though, has been confined to unstructured text corpora, leaving a gap in our understanding: to what extent do the findings of zero-shot LLM rerankers established on plain text corpora hold for datasets containing predominantly precomputed ranking features as is common in industrial settings? We explore this question via an empirical study on one public learning-to-rank dataset (MSLR-WEB10K) and two datasets collected from an audio streaming platform's search logs. Our results paint a differentiated picture: On average, there remains a significant performance gap: prompting the high-capacity LLM GPT-4 results in up to 16% lower NDCG@10 compared to the traditional supervised learning-to-rank (LTR) approach LambdaMART on the public MSLR-WEB10K dataset. However, when focusing only on a subset of hard queries-i.e. queries where the LTR approach ranks a non-relevant document at the top-the zero-shot LLM reranking outperforms the LTR baseline. We confirm the same trends on two proprietary audio search datasets. We also provide insights into prompt design choices and their impact on LLM reranking. We show that LLMs remain brittle, with the same strategies sometimes helping or hurting depending on the model size and dataset. Maria Movin, Claudia Hauff |
SIGIR | 1 |
| 2023 | Explaining Black Box Reinforcement Learning Agents Through Counterfactual Policies
Maria Movin, Guilherme Dinis Junior, Jaakko Hollmén, Panagiotis Papapetrou |
IDA | 1 |