EDBT 2026 Demo / reviewers in the wild / expert
Mónica Marrero
dblp:38/6948
· DBLP profile ↗
12ranked-venue papers
5as first author
3since 2021 · last 2026
0000-0002-2359-6340ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Neural IR in Europeana
Suhaib Basir, Mónica Marrero, Julián Urbano |
ECIR (4) | 2 |
| 2022 | Implementation and Evaluation of a Multilingual Search Pilot in the Europeana Digital Library
Mónica Marrero, Antoine Isaac |
TPDL | 1 |
| 2021 | Automatic Translation and Multilingual Cultural Heritage Retrieval: A Case Study with Transcriptions in Europeana
Mónica Marrero, Antoine Isaac, Nuno Freire 0001 |
TPDL | 1 |
| 2018 | A/B Testing with APONEabstractIn order to improve long-term retention, ad conversion rates, and so on, A/B testing has become the norm within Web portals, enabling efficient large-scale experimentation. While A/B testing is also increasingly used by academic researchers (with crowd-working platforms offering a large pool of artificial users), few platforms are freely available to this end. Academic researchers usually develop adhoc solutions, leading to many duplicated efforts and time spent on work not directly related to one's research. As an alternative, we have developed and open sourced APONE, an A cademic P latform for ON line Experiments. APONE uses PlanOut, a framework and high-level language, to specify online experiments, and offers Web services and a Web GUI to easily create, manage and monitor them. By building a user friendly Web application, we enable not only experts to conduct valid A/B experiments. In particular as a secondary use case, we envision large classrooms to also benefit from the deployment of APONE, a vision we put into practice in a graduate Information Retrieval course. We open-source APONE at https://marrerom.github.io/APONE. A demo version is running at http://ireplatform.ewi.tudelft.nl:8080/APONE. Mónica Marrero, Claudia Hauff |
SIGIR | 1 |
| 2018 | A Semi-automatic and low-cost method to learn patterns for named entity recognitionabstractAbstract Named Entity Recognition is a basic task in Information Extraction that aims at identifying entities of interest within full text documents. The patterns used to recognize entities can be rule based, as in the popular JAPE system. However, hand-crafting effective patterns is often difficult, and yet there is little research devoted to methods capable of learning human-readable patterns, possibly with arbitrary sets of features. In this paper, we present a semi-automatic method to generate both regular expressions and a subset of the JAPE language. It does not need a corpus annotated beforehand. Instead, it employs active learning and combines clustering with an algorithm that finds alignments between symbols present in the entities discovered during the learning process. The method currently supports a fixed set of character features and an arbitrary set of token features, but it can incorporate other kinds of features as well. Through several experiments with an English corpus, we show the ability of the method to generate effective patterns at a low annotation cost, and how it can successfully help in the annotation of brand new corpora. Mónica Marrero, Julián Urbano |
Nat. Lang. Eng. | 1 |
| 2016 | Toward Estimating the Rank Correlation between the Test Collection Results and the True System PerformanceabstractThe Kendall ? and AP rank correlation coefficients have become mainstream in Information Retrieval research for comparing the rankings of systems produced by two different evaluation conditions, such as different effectiveness measures or pool depths. However, in this paper we focus on the expected rank correlation between the mean scores observed with a test collection and the true, unobservable means under the same conditions. In particular, we propose statistical estimators of ? and AP correlations following both parametric and non-parametric approaches, and with special emphasis on small topic sets. Through large scale simulation with TREC data, we study the error and bias of the estimators. In general, such estimates of expected correlation with the true ranking may accompany the results reported from an evaluation experiment, as an easy to understand figure of reliability. All the results in this paper are fully reproducible with data and code available online Julián Urbano, Mónica Marrero |
SIGIR | 2 |
| 2015 | Information Extraction Grammars
Mónica Marrero, Julián Urbano |
ECIR | 1 |
| 2015 | How Do Gain and Discount Functions Affect the Correlation between DCG and User Satisfaction?
Julián Urbano, Mónica Marrero |
ECIR | 2 |
| 2013 | On the measurement of test collection reliabilityabstractThe reliability of a test collection is proportional to the number of queries it contains. But building a collection with many queries is expensive, so researchers have to find a balance between reliability and cost. Previous work on the measurement of test collection reliability relied on data-based approaches that contemplated random what if scenarios, and provided indicators such as swap rates and Kendall tau correlations. Generalizability Theory was proposed as an alternative founded on analysis of variance that provides reliability indicators based on statistical theory. However, these reliability indicators are hard to interpret in practice, because they do not correspond to well known indicators like Kendall tau correlation. We empirically established these relationships based on data from over 40 TREC collections, thus filling the gap in the practical interpretation of Generalizability Theory. We also review the computation of these indicators, and show that they are extremely dependent on the sample of systems and queries used, so much that the required number of queries to achieve a certain level of reliability can vary in orders of magnitude. We discuss the computation of confidence intervals for these statistics, providing a much more reliable tool to measure test collection reliability. Reflecting upon all these results, we review a wealth of TREC test collections, arguing that they are possibly not as reliable as generally accepted and that the common choice of 50 queries is insufficient even for stable rankings. Julián Urbano, Mónica Marrero, Diego Martín 0001 |
SIGIR | 2 |
| 2013 | A comparison of the optimality of statistical significance tests for information retrieval evaluationabstractPrevious research has suggested the permutation test as the theoretically optimal statistical significance test for IR evaluation, and advocated for the discontinuation of the Wilcoxon and sign tests. We present a large-scale study comprising nearly 60 million system comparisons showing that in practice the bootstrap, t-test and Wilcoxon test outperform the permutation test under different optimality criteria. We also show that actual error rates seem to be lower than the theoretically expected 5%, further confirming that we may actually be underestimating significance. Julián Urbano, Mónica Marrero, Diego Martín 0001 |
SIGIR | 2 |
| 2011 | Bringing undergraduate students closer to a real-world information retrieval setting: methodology and resourcesabstractWe describe a pilot experiment to update the program of an Information Retrieval course for Computer Science undergraduates. We have engaged the students in the development of a search engine from scratch, and they have been involved in the elaboration, also from scratch, of a complete test collection to evaluate their systems. With this methodology they get a whole vision of the Information Retrieval process as they would find it in a real-world setting, and their direct involvement in the evaluation makes them realize the importance of these laboratory experiments in Computer Science. We show that this methodology is indeed reliable and feasible, and so we plan on improving and keep using it in the next years, leading to a public repository of resources for Information Retrieval courses. Julián Urbano, Mónica Marrero, Diego Martín 0001, Jorge Morato |
ITiCSE | 2 |
| 2010 | Crawling the web for structured documentsabstractStructured Information Retrieval is gaining a lot of interest in recent years, as this kind of information is becoming an invaluable asset for professional communities such as Software Engineering. Most of the research has focused on XML documents, with initiatives like INEX to bring together and evaluate new techniques focused on structured information. Despite the use of XML documents is the immediate choice, the Web is filled with several other types of structured information, which account for millions of other documents. These documents may be collected directly using standard Web search engines like Google and Yahoo, or following specific search patterns in online repositories like SourceForge. This demo describes a distributed and focused web crawler for any kind of structured documents, and we show with it how to exploit general-purpose resources to gather large amounts of real-world structured documents off the Web. This kind of tool could help building large test collections of other types of documents, such as Java source code for software-oriented search engines or RDF for semantic searching. Julián Urbano, Juan Llorens Morillo, Yorgos Andreadakis, Mónica Marrero |
CIKM | 4 |