Maria Movin

dblp:345/7846 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0001-5759-7846ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › ranking
learning to rank
0.912025
Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025
Information retrieval › retrieval models › neural retrieval › neural ranking model
LLM-based ranking
0.912025
Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025
Information retrieval
reranking
0.912025
Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025
Information retrieval › reranking
zero-shot re-ranking
0.912025
Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025
Information retrieval
ranking
0.312025
Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

large language model prompting · 0.9LambdaMART · 0.9
YearPublicationVenuePosition
2025 You Say Search, I Say Recs: A Scalable Agentic Approach to Query Understanding and Exploratory Search at Spotify
abstract
On online content platforms, users often aim to explore the catalog and discover new, personalized content through exploratory searches-such as "new releases for me." Traditional search systems, which prioritize lexical and semantic matching over personalized retrieval, have historically struggled to support this type of intent.In contrast, recommendation services that leverage user-item and item-item signals tend to be more effective for addressing exploratory queries.Agentic technologies offer a promising opportunity to enhance exploratory search by harnessing large language models (LLMs) to interpret complex query intents and route them to the most suitable downstream services.However, deploying such
Enrico Palumbo, Marcus Isaksson, Alexandre Tamborrino, Maria Movin, Catalin Dincu, Ali Vardasbi, Lev Nikeshkin, Oksana Gorobets, Anders Nyman, Poppy Newdick, Hugues Bouchard, Paul N. Bennett, Mounia Lalmas-Roelleke, Dani Doro, Christine Doig Cardet, Ziad Sultan
RecSys4
2025 Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations
abstract
LLMs have been explored for their use in IR as end-to-end rankers, rerankers and assessors. Recently, the exploration of the prompt-and-predict paradigm for reranking in combination with highly performant LLMs have drawn the attention of researchers. Instead of training or fine-tuning a reranker, LLMs are prompted in a zero-shot manner to produce relevance scores, pairwise preferences, or reranked lists. Existing research, though, has been confined to unstructured text corpora, leaving a gap in our understanding: to what extent do the findings of zero-shot LLM rerankers established on plain text corpora hold for datasets containing predominantly precomputed ranking features as is common in industrial settings? We explore this question via an empirical study on one public learning-to-rank dataset (MSLR-WEB10K) and two datasets collected from an audio streaming platform's search logs. Our results paint a differentiated picture: On average, there remains a significant performance gap: prompting the high-capacity LLM GPT-4 results in up to 16% lower NDCG@10 compared to the traditional supervised learning-to-rank (LTR) approach LambdaMART on the public MSLR-WEB10K dataset. However, when focusing only on a subset of hard queries-i.e. queries where the LTR approach ranks a non-relevant document at the top-the zero-shot LLM reranking outperforms the LTR baseline. We confirm the same trends on two proprietary audio search datasets. We also provide insights into prompt design choices and their impact on LLM reranking. We show that LLMs remain brittle, with the same strategies sometimes helping or hurting depending on the model size and dataset.
Maria Movin, Claudia Hauff
SIGIR1
2023 Explaining Black Box Reinforcement Learning Agents Through Counterfactual Policies
Maria Movin, Guilherme Dinis Junior, Jaakko Hollmén, Panagiotis Papapetrou
IDA1